runningqilin commited on
Commit
019cd9a
·
verified ·
1 Parent(s): 5b90cbe

dataset page: fact-check corrections

Browse files
Files changed (2) hide show
  1. README.md +17 -14
  2. decks/README.md +1 -1
README.md CHANGED
@@ -64,30 +64,31 @@ Autoregressive next-step surrogate of a 2D SPH copper bar under Taylor impact ag
64
 
65
  ## Manifest and input decks (Hugging Face mirror)
66
 
67
- - `cases.csv` — one row per `.h5`: `case_id`, `split` (`held_aside` for files shipped outside the protocol splits), the loading/geometry parameters parsed from the id, `n_nodes`, `n_frames`, `file_bytes`, `sha256` (integrity manifest; also what the Dataset Viewer shows).
68
  - `decks/<case_id>.k` — the LS-DYNA input deck of every case (also embedded verbatim in each file's `metadata/source_deck`); re-running a deck regenerates the raw output the adapter converts to canonical HDF5.
69
  - Case ids: `T-20-<L>-<V>` — 20 mm bar width, bar length L mm, impact speed V m/s (`T-20-80-Convergence` is the held-aside mesh-convergence run).
70
 
71
  ## HDF5 layout
72
 
73
- One HDF5 file per case, readable with `h5py` or any HDF5 tool. Every quantity is stored in strict SI (m, s, kg, Pa, J) regardless of the solver's `g-mm-ms` source convention. Small scalars are HDF5 attributes; arrays are datasets (float64 geometry and time, float32 response, int64 ids); response arrays are gzip-compressed and chunked along the frame axis, so one frame or transition can be read without loading the whole trajectory. Shapes below use N nodes, P SPH particles, E elements, T frames and d = `metadata.dimension`; the exact schema version is the `schema_version` attribute (ADR-0013 — 0.2.0 readers read 0.1.0 files unchanged, ADR-0042).
74
 
75
  | Path | Shape | Dtype | Content |
76
  |---|---|---|---|
77
  | `metadata` (attrs) | — | — | `case_id`, `dataset_id`, `dimension`, `schema_version`, `source_units`, `units_convention` (= `SI`) |
78
  | `metadata/provenance` (attrs) | — | — | `solver_name`, `solver_version`, `generation_date` |
79
- | `metadata/source_deck` | scalar | str | the complete solver input deck, verbatim |
80
  | `nodes/coords` | (N, d) | f64 | initial node coordinates [m] |
81
  | `nodes/node_id` | (N,) | i64 | solver node ids |
82
- | `materials/{canonical_model, source_model, source_params, material_id}` | (M,) | str / i64 | material models; `source_params` is the solver's material card as JSON |
83
- | `response/time/t` | (T,) | f64 | output times [s]; frame 0 is the initial state |
84
  | `elements/sph/connectivity` | (P, 1) | i64 | particle → node index (0-based) |
85
  | `elements/sph/{element_id, part_id}` | (P,) | i64 | solver element id, part id |
86
- | `elements/<other>/…` | (E, n), (E,) | i64 | any further element group (e.g. a single rigid-wall / boundary `shell`) follows the same connectivity, element_id, part_id pattern |
87
  | `response/node/{displacement, velocity, acceleration}` | (T, N, d) | f32 | [m], [m/s], [m/s²] |
88
  | `response/element/sph/{stress, strain, strain_rate}` | (T, P, 6) | f32 | Voigt (xx, yy, zz, xy, yz, zx): [Pa], [–], [1/s] — six components even for 2D cases |
89
- | `response/element/sph/{pressure, density, mass, internal_energy}` | (T, P) | f32 | [Pa], [kg/m³], [kg], [J] |
90
- | `response/element/sph/{effective_plastic_strain, radius, n_neighbors, deletion}` | (T, P) | f32 | [–], smoothing length [m], neighbour count, 0/1 deletion flag |
 
91
  | `response/element/<other>/…` | (T, E, …) | f32 | per-element response of any further element group |
92
  | `response/global/{kinetic_energy, internal_energy, total_energy}` | (T,) | f32 | [J] |
93
 
@@ -108,20 +109,20 @@ with h5py.File("<case_id>.h5") as f:
108
  sig = f["response/element/sph/stress"][:] # (T, P, 6) Pa, Voigt
109
  ```
110
 
111
- Or through StructBench's loader, which returns the ML working frame (positions in mm; `von_mises_stress` in MPa) with the auxiliary target derived on the fly:
112
 
113
  ```python
114
  from structbench.datasets import load_case_trajectory
115
 
116
  traj = load_case_trajectory("<case_id>.h5", aux_field="von_mises_stress")
117
- traj.positions # (T, P, d) float32, mm
118
- traj.aux # (T, P) float32, MPa
119
- traj.time # (T,) float64, s
120
  ```
121
 
122
  ## Splits
123
 
124
- The benchmark's fixed split assignment, by case id (the file name without `.h5`):
125
 
126
  - **train** (21): `T-20-60-100`, `T-20-80-100`, `T-20-100-100`, `T-20-60-110`, `T-20-80-110`, `T-20-100-110`, `T-20-60-120`, `T-20-80-120`, `T-20-100-120`, `T-20-60-140`, `T-20-80-140`, `T-20-100-140`, `T-20-60-160`, `T-20-80-160`, `T-20-100-160`, `T-20-60-180`, `T-20-80-180`, `T-20-100-180`, `T-20-60-190`, `T-20-80-190`, `T-20-100-190`
127
  - **val** (3): `T-20-60-150`, `T-20-80-150`, `T-20-100-150`
@@ -133,7 +134,7 @@ The benchmark's fixed split assignment, by case id (the file name without `.h5`)
133
  This archive backs the **Taylor2D-Impact** benchmark in StructBench. Task: autoregressive transition (ADR-0019); auxiliary target `von_mises_stress` (MPa); 6 input frames, horizon full, scored at native output times; quantities of interest: final_length, mushroom_width, peak_von_mises, t_peak_von_mises. The full evaluation protocol and its rationale, the baseline recipes and checkpoints, and the current leaderboard live on the benchmark page in the code repository — <https://github.com/qilinli/StructBench/blob/main/docs/benchmarks/taylor_impact_2d.md> — so the numbers have a single home. To train a baseline on this archive:
134
 
135
  ```bash
136
- pip install structbench # or: pip install -e . from the repo
137
  structbench-train --mode train --config configs/taylor_impact_2d/cgn.toml \
138
  --data-root /path/to/this/folder --out runs/taylor_impact_2d-cgn
139
  ```
@@ -159,3 +160,5 @@ The data and the code are released together — cite the software
159
  url = {https://github.com/qilinli/StructBench},
160
  }
161
  ```
 
 
 
64
 
65
  ## Manifest and input decks (Hugging Face mirror)
66
 
67
+ - `cases.csv` — one row per `.h5`: `case_id`, `split` (`held_aside` for files shipped outside the protocol splits), the loading/geometry parameters parsed from the id, `n_nodes` (rows of `nodes/coords`, so including any boundary-shell nodes), `n_frames` (stored frames), `file_bytes`, `sha256` (integrity manifest; also what the Dataset Viewer shows).
68
  - `decks/<case_id>.k` — the LS-DYNA input deck of every case (also embedded verbatim in each file's `metadata/source_deck`); re-running a deck regenerates the raw output the adapter converts to canonical HDF5.
69
  - Case ids: `T-20-<L>-<V>` — 20 mm bar width, bar length L mm, impact speed V m/s (`T-20-80-Convergence` is the held-aside mesh-convergence run).
70
 
71
  ## HDF5 layout
72
 
73
+ One HDF5 file per case, readable with `h5py` or any HDF5 tool. Every quantity is stored in strict SI (m, s, kg, Pa, J) regardless of the solver's `g-mm-ms` source convention. Small scalars are HDF5 attributes; arrays are datasets (float64 geometry and time, float32 response, int64 ids, variable-length UTF-8 strings — h5py returns those as `bytes`); response arrays are gzip-compressed and chunked in blocks of frames, so slicing along the frame axis reads only the chunks it touches. Shapes below use N nodes, P SPH particles, E elements, T stored frames and d = `metadata.dimension`; the exact schema version is the `schema_version` attribute (ADR-0013 — 0.2.0 readers read 0.1.0 files unchanged, ADR-0042). `ADR-NNNN` refers to the decision records under `decisions/` in the code repository.
74
 
75
  | Path | Shape | Dtype | Content |
76
  |---|---|---|---|
77
  | `metadata` (attrs) | — | — | `case_id`, `dataset_id`, `dimension`, `schema_version`, `source_units`, `units_convention` (= `SI`) |
78
  | `metadata/provenance` (attrs) | — | — | `solver_name`, `solver_version`, `generation_date` |
79
+ | `metadata/source_deck` | scalar | str | the complete solver input deck, verbatim (solver-ingested cases) |
80
  | `nodes/coords` | (N, d) | f64 | initial node coordinates [m] |
81
  | `nodes/node_id` | (N,) | i64 | solver node ids |
82
+ | `materials/{canonical_model, source_model, source_params, material_id}` | (M,) | str / i64 | material models; `source_params` is the solver's material card as JSON; `canonical_model` is empty when the source model has no canonical mapping |
83
+ | `response/time/t` | (T,) | f64 | the solver's actual output times [s], nominally every 0.002 ms; frame 0 is the initial state; the last stored frame is a terminal solver-output artifact that the loader drops (ADR-0028) |
84
  | `elements/sph/connectivity` | (P, 1) | i64 | particle → node index (0-based) |
85
  | `elements/sph/{element_id, part_id}` | (P,) | i64 | solver element id, part id |
86
+ | `elements/<other>/…` | (E, n), (E,) | i64 | any further element group (e.g. a single rigid-wall / boundary `shell`, whose nodes are counted in N but are not particles) follows the same connectivity, element_id, part_id pattern |
87
  | `response/node/{displacement, velocity, acceleration}` | (T, N, d) | f32 | [m], [m/s], [m/s²] |
88
  | `response/element/sph/{stress, strain, strain_rate}` | (T, P, 6) | f32 | Voigt (xx, yy, zz, xy, yz, zx): [Pa], [–], [1/s] — six components even for 2D cases |
89
+ | `response/element/sph/{pressure, density, mass, internal_energy}` | (T, P) | f32 | [Pa] (positive in compression, = −tr σ / 3), [kg/m³], [kg], [J] |
90
+ | `response/element/sph/effective_plastic_strain` | (T, P) | f32 | whatever the material model writes to LS-DYNA's plastic-strain history slot: equivalent plastic strain [–] for elastoplastic models, the K&C concrete model's scaled damage measure (0–2) for `*MAT_CONCRETE_DAMAGE_REL3`, and an unrelated history variable for purely elastic materials (treat as unused) |
91
+ | `response/element/sph/{radius, n_neighbors, deletion}` | (T, P) | f32 | smoothing length [m], neighbour count, 0/1 deletion flag |
92
  | `response/element/<other>/…` | (T, E, …) | f32 | per-element response of any further element group |
93
  | `response/global/{kinetic_energy, internal_energy, total_energy}` | (T,) | f32 | [J] |
94
 
 
109
  sig = f["response/element/sph/stress"][:] # (T, P, 6) Pa, Voigt
110
  ```
111
 
112
+ Or through StructBench's loader, which returns the ML working frame (positions in mm; `von_mises_stress` in MPa) with the auxiliary target derived on the fly — P SPH particles only (boundary-shell nodes are dropped); T′ = T − 1: the terminal solver-output frame is dropped (ADR-0028):
113
 
114
  ```python
115
  from structbench.datasets import load_case_trajectory
116
 
117
  traj = load_case_trajectory("<case_id>.h5", aux_field="von_mises_stress")
118
+ traj.positions # (T′, P, d) float32, mm
119
+ traj.aux # (T′, P) float32, MPa
120
+ traj.time # (T′,) float64, s
121
  ```
122
 
123
  ## Splits
124
 
125
+ The benchmark's fixed split assignment, by case id (the file name without `.h5`). `train` fits the model and `val` selects it; every other split is held out for reporting — `test_interp` sits inside the training parameter range, `test_extrap` beyond it, `probe` cases are off-grid stress tests. Files shipped outside these splits (`held_aside` in the manifest) are not part of the protocol.
126
 
127
  - **train** (21): `T-20-60-100`, `T-20-80-100`, `T-20-100-100`, `T-20-60-110`, `T-20-80-110`, `T-20-100-110`, `T-20-60-120`, `T-20-80-120`, `T-20-100-120`, `T-20-60-140`, `T-20-80-140`, `T-20-100-140`, `T-20-60-160`, `T-20-80-160`, `T-20-100-160`, `T-20-60-180`, `T-20-80-180`, `T-20-100-180`, `T-20-60-190`, `T-20-80-190`, `T-20-100-190`
128
  - **val** (3): `T-20-60-150`, `T-20-80-150`, `T-20-100-150`
 
134
  This archive backs the **Taylor2D-Impact** benchmark in StructBench. Task: autoregressive transition (ADR-0019); auxiliary target `von_mises_stress` (MPa); 6 input frames, horizon full, scored at native output times; quantities of interest: final_length, mushroom_width, peak_von_mises, t_peak_von_mises. The full evaluation protocol and its rationale, the baseline recipes and checkpoints, and the current leaderboard live on the benchmark page in the code repository — <https://github.com/qilinli/StructBench/blob/main/docs/benchmarks/taylor_impact_2d.md> — so the numbers have a single home. To train a baseline on this archive:
135
 
136
  ```bash
137
+ pip install git+https://github.com/qilinli/StructBench # or: pip install -e . from a clone
138
  structbench-train --mode train --config configs/taylor_impact_2d/cgn.toml \
139
  --data-root /path/to/this/folder --out runs/taylor_impact_2d-cgn
140
  ```
 
160
  url = {https://github.com/qilinli/StructBench},
161
  }
162
  ```
163
+
164
+ Licence: CC BY 4.0 — when redistributing or building on the data, credit Qilin Li (Curtin University) / StructBench and link this dataset repository.
decks/README.md CHANGED
@@ -1,5 +1,5 @@
1
  LS-DYNA input decks, one per case (`<case_id>.k`), in the deck's
2
- native kg-mm-ms unit system — the provenance root: re-running a
3
  deck in LS-DYNA regenerates the case's raw output, which the
4
  repository's adapter converts to the canonical HDF5 shipped here
5
  (strict SI, ADR-0012/0016).
 
1
  LS-DYNA input decks, one per case (`<case_id>.k`), in the deck's
2
+ native g-mm-ms unit system — the provenance root: re-running a
3
  deck in LS-DYNA regenerates the case's raw output, which the
4
  repository's adapter converts to the canonical HDF5 shipped here
5
  (strict SI, ADR-0012/0016).