Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| assets | 164 items | ||
| calibration | 21 items | ||
| checkpoints | 3 items | ||
| fastercache_humanoid_sv | 432 items | ||
| scores | 1,067 items | ||
| train | 38 items | ||
| web | 160 items | ||
| BOOTSTRAP.md | 5.02 kB xet | 1dc69261 | |
| README.md | 3.06 kB xet | 5993bf09 | |
| ledger.json | 6.67 kB xet | 8b1ad6a0 | |
| report.json | 4.83 kB xet | 5bc8e81c |
video-bench-model
Pose readers and their benchmark scores. A reader is
<robot>.<corpus>.<view_id>; a cell is <embodiment>.<view_id>.<model>.
A reader is trained once and scores every cell that names it.
Readers
| reader | head | train ep | val ep | train mm | val mm | best step | steps | train_log |
|---|---|---|---|---|---|---|---|---|
| airbot_mmk2.humanoid_mv.mv4_grid | diffusion | 227 | 21 | 18.51 | 98.32 | 500 | 6000 | yes |
| airbot_mmk2.humanoid_mv.mv4_row | diffusion | 227 | 21 | 22.15 | 78.64 | 4500 | 6000 | - |
| fourier_gr1.humanoid_sv.sv1_16x9 | diffusion | 234 | 26 | 17.29 | 33.00 | 4500 | 6000 | yes |
| fourier_gr1.humanoid_sv.sv1_4x3 | diffusion | 234 | 26 | 49.19 | 64.15 | 6000 | 6000 | yes |
train mm / val mm are keypoint error in millimetres against
forward-kinematics targets.
Violations
Verdicts are per segment of 16 frames, not per frame: rigidity reduces by median over the segment, every other detector by worst.
| cell | clips | segments | rigidity | jerk | teleport | joint_limit | self_collision |
|---|---|---|---|---|---|---|---|
| augment.mv4_row.ctrlworld_4view_grid | 20 | 71 | 0/71 | 0/71 | 0/71 | 64/71 | 9/71 |
| humanoid.mv4_grid.dreamgen | 34 | 343 | 43/343 | 0/343 | 29/343 | 251/343 | 169/343 |
| humanoid.mv4_row.ctrlworld_4view_grid | 34 | 118 | 0/118 | 0/118 | 0/118 | 97/118 | 21/118 |
| humanoid.sv1_16x9.dreamgen | 34 | 486 | 32/486 | 93/486 | 246/486 | 106/486 | 112/486 |
| humanoid.sv1_4x3.dreamdojo | 34 | 334 | 24/334 | 0/334 | 77/334 | 84/334 | 56/334 |
Thresholds
The 95th percentile of real motion, fitted per reader. A cell's counts are only as sharp as the reader behind them: a threshold near the noise floor of its reader flags almost every segment and separates nothing.
| cell | rigidity (mm) | jerk (mm/s^3) | teleport (mm/s) | joint_limit (deg) | self_collision (mm) |
|---|---|---|---|---|---|
| augment.mv4_row.ctrlworld_4view_grid | 76.15 | 2.817e+06 | 1093 | 4.07 | 190.2 |
| humanoid.mv4_grid.dreamgen | 51.81 | 2.25e+06 | 864.8 | 4.92 | 200.3 |
| humanoid.mv4_row.ctrlworld_4view_grid | 76.61 | 2.809e+06 | 1114 | 4.29 | 190.3 |
| humanoid.sv1_16x9.dreamgen | 43.75 | 4.549e+05 | 450.8 | 3 | 164.4 |
| humanoid.sv1_4x3.dreamdojo | 30.55 | 4.631e+05 | 365.4 | 3.53 | 163.7 |
Layout
train/<reader_id>/<head>/ checkpoint.pt meta.json train_log.jsonl
scores/<cell_id>/<head>/ summary.json results.jsonl metrics.csv
segments.csv run_manifest.json
render/<path in bench>.mp4
render/reel/<tree>.mp4
assets/ URDF tree
calibration/<reader_id>/ real clips the thresholds are fitted on
segments.csv is the file to read a number off: one row per segment, carrying
<detector>_reduce, _value, _threshold, _violated beside the clip's
bench path. render/ mirrors the bench tree so a row traces to its video, and
render/reel/ holds one reel per tree so a population can be watched alone.
- Total size
- 8.52 GB
- Files
- 1,889
- Last updated
- Aug 29
- Pre-warmed CDN
- US EU US EU