Dataset Viewer
The dataset viewer is not available for this split.
Cannot load the dataset split (in streaming mode) to extract the first rows.
Error code: StreamingRowsError
Exception: CastError
Message: Couldn't cast
schema_version: int64
created_utc: string
repository_commit: string
prompts: struct<path: string, rows: int64, sha256: string>
child 0, path: string
child 1, rows: int64
child 2, sha256: string
shape_per_prompt: list<item: int64>
child 0, item: int64
dtype: string
lam_checkpoint: struct<path: string, sha256: string>
child 0, path: string
child 1, sha256: string
validation: struct<cosine_min: double, coverage: string, live_samples: int64, log: string, mae: double, max_abs: (... 30 chars omitted)
child 0, cosine_min: double
child 1, coverage: string
child 2, live_samples: int64
child 3, log: string
child 4, mae: double
child 5, max_abs: double
child 6, relative_mae: double
archives: struct<droid_old78k_text_cache_shard0of2.tar.zst: struct<bytes: int64, sha256: string>, droid_old78k (... 68 chars omitted)
child 0, droid_old78k_text_cache_shard0of2.tar.zst: struct<bytes: int64, sha256: string>
child 0, bytes: int64
child 1, sha256: string
child 1, droid_old78k_text_cache_shard1of2.tar.zst: struct<bytes: int64, sha256: string>
child 0, bytes: int64
child 1, sha256: string
query_shape_per_row: list<item: int64>
child 0, item: int64
query_storage: string
manifest_sha256: string
rows: int64
rows_per_shard: list<item: int64>
child 0, item: int64
query_dtype: string
to
{'schema_version': Value('int64'), 'created_utc': Value('timestamp[s]'), 'repository_commit': Value('string'), 'manifest_sha256': Value('string'), 'rows': Value('int64'), 'rows_per_shard': List(Value('int64')), 'query_shape_per_row': List(Value('int64')), 'query_dtype': Value('string'), 'query_storage': Value('string'), 'archives': {'droid_old78k_query_s29_shard0of2.tar.zst': {'bytes': Value('int64'), 'sha256': Value('string')}, 'droid_old78k_query_s29_shard1of2.tar.zst': {'bytes': Value('int64'), 'sha256': Value('string')}}}
because column names don't match
Traceback: Traceback (most recent call last):
File "/src/services/worker/src/worker/utils.py", line 147, in get_rows_or_raise
return get_rows(
dataset=dataset,
...<4 lines>...
column_names=column_names,
)
File "/src/libs/libcommon/src/libcommon/utils.py", line 272, in decorator
return func(*args, **kwargs)
File "/src/services/worker/src/worker/utils.py", line 127, in get_rows
rows_plus_one = list(itertools.islice(safe_iter(ds, dataset=dataset), rows_max_number + 1))
File "/src/services/worker/src/worker/utils.py", line 483, in safe_iter
yield from ds.decode(False) if ds.features else ds
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2840, in __iter__
for key, example in ex_iterable:
^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2373, in __iter__
for key, pa_table in self._iter_arrow():
~~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2398, in _iter_arrow
for key, pa_table in self.ex_iterable._iter_arrow():
~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 536, in _iter_arrow
for key, pa_table in iterator:
^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 419, in _iter_arrow
for key, pa_table in self.generate_tables_fn(**gen_kwags):
~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 343, in _generate_tables
self._cast_table(pa_table, json_field_paths=json_field_paths),
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 132, in _cast_table
pa_table = table_cast(pa_table, self.info.features.arrow_schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2378, in table_cast
return cast_table_to_schema(table, schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2306, in cast_table_to_schema
raise CastError(
...<3 lines>...
)
datasets.table.CastError: Couldn't cast
schema_version: int64
created_utc: string
repository_commit: string
prompts: struct<path: string, rows: int64, sha256: string>
child 0, path: string
child 1, rows: int64
child 2, sha256: string
shape_per_prompt: list<item: int64>
child 0, item: int64
dtype: string
lam_checkpoint: struct<path: string, sha256: string>
child 0, path: string
child 1, sha256: string
validation: struct<cosine_min: double, coverage: string, live_samples: int64, log: string, mae: double, max_abs: (... 30 chars omitted)
child 0, cosine_min: double
child 1, coverage: string
child 2, live_samples: int64
child 3, log: string
child 4, mae: double
child 5, max_abs: double
child 6, relative_mae: double
archives: struct<droid_old78k_text_cache_shard0of2.tar.zst: struct<bytes: int64, sha256: string>, droid_old78k (... 68 chars omitted)
child 0, droid_old78k_text_cache_shard0of2.tar.zst: struct<bytes: int64, sha256: string>
child 0, bytes: int64
child 1, sha256: string
child 1, droid_old78k_text_cache_shard1of2.tar.zst: struct<bytes: int64, sha256: string>
child 0, bytes: int64
child 1, sha256: string
query_shape_per_row: list<item: int64>
child 0, item: int64
query_storage: string
manifest_sha256: string
rows: int64
rows_per_shard: list<item: int64>
child 0, item: int64
query_dtype: string
to
{'schema_version': Value('int64'), 'created_utc': Value('timestamp[s]'), 'repository_commit': Value('string'), 'manifest_sha256': Value('string'), 'rows': Value('int64'), 'rows_per_shard': List(Value('int64')), 'query_shape_per_row': List(Value('int64')), 'query_dtype': Value('string'), 'query_storage': Value('string'), 'archives': {'droid_old78k_query_s29_shard0of2.tar.zst': {'bytes': Value('int64'), 'sha256': Value('string')}, 'droid_old78k_query_s29_shard1of2.tar.zst': {'bytes': Value('int64'), 'sha256': Value('string')}}}
because column names don't matchNeed help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
old78k latent-query caches
This is the canonical repository for old78k [32, 2048] latent-action query
caches and their row-aligned text/manifest metadata. First-frame VAE latents
are stored separately in
Jiiiiiisoo/old78k_firstframe_vae_s29_v1.
Layout
droid/: DROID query and text caches.molmoact/additional/: the non-overlapping additional MolmoAct tranche (458,332 clips).molmoact/selected/: the original selected MolmoAct cache (458,332 clips: 219,817available+ 238,515remainder).
The Bridge, AgiBot, EgoDex, and EgoVerse query caches remain in the original
public source repository:
chyun/old78k_human_s29_q32_step078000_c056967f_v1.
MolmoAct additional and selected are disjoint sets. Keep the generation
identifier (additional, selected/available, or selected/remainder) when
constructing a combined index; row numbers are local to their own manifests.
- Downloads last month
- 128