pusht-simulator / docs /html /dataset_definition.html
desertmouse's picture
pages
473cc8a verified
Raw History Blame Contribute Delete
11.7 kB
<!doctype html><html><head><meta charset='utf-8'><title>Dataset definition: `pusht-simulator-wall05-contact2`</title><style>
body{max-width:980px;margin:2rem auto;padding:0 1.25rem;font:15px/1.55 -apple-system,Segoe UI,Roboto,sans-serif;color:#1c2330;background:#fafbfc}
h1{font-size:1.7rem;border-bottom:2px solid #d8dee8;padding-bottom:.4rem}h2{margin-top:2.2rem;border-bottom:1px solid #e3e8ef;padding-bottom:.25rem}
table{border-collapse:collapse;margin:1rem 0;font-size:13.5px;font-variant-numeric:tabular-nums}th,td{border:1px solid #d8dee8;padding:.35rem .6rem;text-align:left}
th{background:#eef2f7}td:not(:first-child){text-align:right}code{background:#eef2f7;padding:.1rem .3rem;border-radius:3px;font-size:.92em}
pre{background:#1c2330;color:#e6edf5;padding:.8rem 1rem;border-radius:6px;overflow:auto}pre code{background:none;color:inherit}
img{max-width:100%;border:1px solid #d8dee8;border-radius:4px;margin:.5rem 0}blockquote{border-left:3px solid #6aa9ff;margin:0;padding:.2rem 1rem;color:#4b5563}
nav{font-size:13px;margin-bottom:1.5rem}nav a{margin-right:1rem}
</style></head><body><nav><a href="index.html">index</a> <a href="brain3_vs_3a_events.html">brain3 vs 3a events</a> <a href="contact_friction_study.html">contact friction study</a> <a href="convergence_study.html">convergence study</a> <a href="dataset_definition.html">dataset definition</a> <a href="ds08_distribution_at_1000.html">ds08 distribution at 1000</a> <a href="ds_distribution_at_1000.html">ds distribution at 1000</a> <a href="ds_distribution_report.html">ds distribution report</a> <a href="ds_mu03_brain3b_distribution.html">ds mu03 brain3b distribution</a> <a href="episode_store.html">episode store</a> <a href="human_pushing_analysis.html">human pushing analysis</a> <a href="lewm_parity.html">lewm parity</a> <a href="reference_parity.html">reference parity</a> <a href="session_fidelity_and_behaviour.html">session fidelity and behaviour</a></nav><h1 id="dataset-definition-pusht-simulator-wall05-contact2">Dataset definition: <code>pusht-simulator-wall05-contact2</code></h1>
<p>The specification every published pusht-simulator dataset follows, fixed
here before the first 100 episodes of the pymunk-wall dataset were
generated. It is the LeWM-compatible layout the two reference datasets
already use (<code>scripts/publish_hf.py</code>), plus sharding for size and the
provenance fields the friction study introduced. A dataset that deviates
from this file says so in its README.</p>
<h2 id="identity">Identity</h2>
<table>
<thead>
<tr>
<th>field</th>
<th>value</th>
</tr>
</thead>
<tbody>
<tr>
<td>repo</td>
<td><code>desertmouse/pusht-simulator-wall05-contact2</code> (moved to the team org when it exists)</td>
</tr>
<tr>
<td>engine</td>
<td><code>pymunk-wall</code>: the published gym-pusht environment (pixels, kinematic PD pusher <code>k_p=100, k_v=20</code>, mass 1, moment 3000, <code>damping=0</code>, frictionless walls) plus <strong>Coulomb friction 0.5 between the pusher and the T</strong> (<code>contact_friction</code>, set on both shapes; pymunk's pair rule is the geometric mean)</td>
</tr>
<tr>
<td>brain</td>
<td><strong>Brain-contact2 + trim</strong>: rule planner with the friction-study corrections (<code>ContactLoopPlanner</code>, <code>planner: contact2</code>) and the executor's near-goal trim release. Deterministic; no model call; <code>model: null</code> in every plan record</td>
</tr>
<tr>
<td>horizon</td>
<td>200 control steps at 10 Hz; episodes are not cut short on success (the executor backs off and holds)</td>
</tr>
<tr>
<td>goal</td>
<td>the published fixed goal: block centre (256, 256), angle π/4; the pusher's goal position follows the <em>goal convention</em> below</td>
</tr>
<tr>
<td>seeds</td>
<td>drawn once, deterministically: <code>random.Random("pusht-wall05-contact2").sample(range(100_000, 1_000_000), 3000)</code>; the first 100 are the sample, positions 100–2999 the extension. Disjoint from every eval and study batch (those are &lt; 100 000)</td>
</tr>
<tr>
<td>target size</td>
<td>100 (sample, reviewed) → 3000</td>
</tr>
</tbody>
</table>
<h2 id="per-episode-metaepisodesparquet-one-row">Per episode (<code>meta/episodes.parquet</code>, one row)</h2>
<p>The full episode-store record (33 columns: seed, situation labels, start
and final errors, coverage, first-success and solved step, bouts, speeds,
phase shares, planner, git commit, render style, ...) plus:</p>
<ul>
<li><code>episode_index</code> (0-based, dataset order), <code>video</code>, <code>length</code></li>
<li><code>goal_state</code> (7): <code>[agent_x, agent_y, block_x, block_y, angle % 2π, agent_vx, agent_vy]</code> at the goal</li>
<li><code>goal_proprio</code> (4), <code>goal_pose</code> (3) = (256, 256, π/4)</li>
<li><code>goal_agent_convention</code>: <code>final_agent_position</code> when the episode was solved, else <code>goal_cog_front</code> (pusher centre on the outward normal of the bar's top edge through the CoG, 25 px outside the outline = radius 15 + clear gap 10; gap exactly 10 px) - and <code>goal_agent_gap_px</code></li>
<li><strong>goal images</strong> <code>meta/goals/episode_NNNNNN_{512,224}.png</code>: <code>render_state</code> of the block at the goal pose and the pusher at the goal agent position. This is LeWM's <code>goal</code> info key, the evaluator's conditioning target</li>
<li><code>contact_friction</code>, <code>sim_hz</code>, <code>control_hz</code> (the honored engine parameters, repeated per row so a merged dataset stays self-describing)</li>
</ul>
<h2 id="per-step-dataconfigframes-nnnnnparquet-one-row-per-control-step">Per step (<code>data/&lt;config&gt;/frames-NNNNN.parquet</code>, one row per control step)</h2>
<p>Row <em>k</em> is the observation <strong>before</strong> action <em>k</em> (video frame <em>k</em>; frame 0
is the reset state). <code>reward_*</code>, <code>next.*</code>, <code>terminated_*</code>, <code>n_contacts</code>
describe the state <strong>after</strong> action <em>k</em>.</p>
<table>
<thead>
<tr>
<th>column</th>
<th>dtype / shape</th>
<th>meaning</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>episode_index</code>, <code>frame_index</code>, <code>timestamp</code>, <code>index</code></td>
<td>int, int, float (s), int (global, gapless)</td>
<td></td>
</tr>
<tr>
<td><code>pixels</code></td>
<td>PNG bytes</td>
<td>512×512×3 in <code>native512</code>, 224×224×3 (<code>cv2.INTER_AREA</code>, LeWM's size) in <code>lewm224</code>; decoded from the episode mp4</td>
</tr>
<tr>
<td><code>state</code></td>
<td>float[7]</td>
<td><code>[agent_x, agent_y, block_x, block_y, angle % 2π, agent_vx, agent_vy]</code></td>
</tr>
<tr>
<td><code>proprio</code></td>
<td>float[4]</td>
<td>agent position and velocity</td>
</tr>
<tr>
<td><code>observation.state</code></td>
<td>float[5]</td>
<td>LeRobot's 5-D state (raw angle), kept for LeRobot readers</td>
</tr>
<tr>
<td><code>pos_agent</code>, <code>vel_agent</code>, <code>block_pose</code></td>
<td>float[2], float[2], float[3]</td>
<td>LeWM's info keys</td>
</tr>
<tr>
<td><code>action</code></td>
<td>float[2]</td>
<td><strong>relative</strong>, LeWM's <code>clip((target − agent) / 100, −1, 1)</code></td>
</tr>
<tr>
<td><code>action_abs</code></td>
<td>float[2]</td>
<td>the absolute pusher target in px the brain actually issued</td>
</tr>
<tr>
<td><code>reward_lewm</code></td>
<td>float</td>
<td><code>−‖goal_state − state_after‖</code> over 7 dims</td>
</tr>
<tr>
<td><code>reward_coverage</code></td>
<td>float</td>
<td><code>clip(coverage / 0.95, 0, 1)</code>, LeRobot style</td>
</tr>
<tr>
<td><code>next.coverage</code>, <code>next.gap_px</code></td>
<td>float</td>
<td>after the step</td>
</tr>
<tr>
<td><code>terminated_lewm</code></td>
<td>bool</td>
<td><code>‖goal[:4] − state[:4]‖ &lt; 20 px</code> and wrapped angle diff <code>&lt; π/9</code></td>
</tr>
<tr>
<td><code>terminated_ours</code></td>
<td>bool</td>
<td>coverage ≥ 0.95 <strong>and</strong> pusher gap ≥ 10 px (the workbench's success)</td>
</tr>
<tr>
<td><code>truncated</code></td>
<td>bool</td>
<td>last step</td>
</tr>
<tr>
<td><code>n_contacts</code></td>
<td>int</td>
<td>geometric proxy: 1 if post-step gap ≤ 1 px (the logs do not hold gym-pusht's per-substep collision count)</td>
</tr>
<tr>
<td><code>phase</code>, <code>plan_index</code>, <code>contact</code>, <code>gap_px</code></td>
<td>str, int, bool, float</td>
<td>executor phase, which plan was active, in contact, pusher-to-outline gap</td>
</tr>
</tbody>
</table>
<p><strong>Velocities are exact, not estimated.</strong> The pusher is kinematic under the
published PD law, so <code>agent_vx, agent_vy</code> come from replaying that law
from the recorded initial position and actions; the export <strong>asserts</strong> the
replayed positions match the logged ones to &lt; 1e-3 px on every step of
every episode and aborts otherwise.</p>
<p><strong>Sharding.</strong> 100 episodes (20 000 rows) per parquet file, written as
episodes are processed; nothing is held in memory across shards. HF's
config globs <code>data/native512/*.parquet</code> and <code>data/lewm224/*.parquet</code> read
all shards as one table.</p>
<h2 id="per-plan-metaplansparquet">Per plan (<code>meta/plans.parquet</code>)</h2>
<p>Every plan the brain made: <code>episode_index</code>, <code>plan_index</code>, the plan (edge,
lever_px, orbit, release_alignment_deg, intent, reasoning - which for
contact2 includes the online gain and trim notes), <code>fallback</code>, <code>latency_s</code>,
and <code>outcome.*</code> (angle and position error before/after, coverage
before/after, bout frames, touched, contact turn/travel, why it released).</p>
<h2 id="narratives-metanarrativesparquet">Narratives (<code>meta/narratives.parquet</code>)</h2>
<p>The generated <code>auto</code> narrative per episode (situation, plan sequence,
outcome, what the last bout did), built from measured facts only. No model
narratives.</p>
<h2 id="videos-videosepisode_nnnnnnmp4">Videos (<code>videos/episode_NNNNNN.mp4</code>)</h2>
<p>512×512, 10 fps, 201 frames (reset + 200 steps), H.264 yuv420p, rendered
in <code>pygame-parity</code> style; <strong>no target cross</strong> (the commanded target is in
<code>action_abs</code>).</p>
<h2 id="metainfojson"><code>meta/info.json</code></h2>
<p><code>codebase</code>, <code>source_commit</code>, <code>tag</code>, <code>backend</code>, <code>planner</code>, <code>brain</code>
("Brain-contact2 + trim"), <code>contact_friction</code>, <code>fps</code>, <code>total_episodes</code>,
<code>total_frames</code>, <code>solved_episodes</code>, <code>seed_rule</code> (the line above), <code>features</code>
(every column with dtype/shape/units/notes), <code>episode_features</code>, <code>configs</code>,
<code>goals</code>, <code>goal_convention</code>, <code>success_criterion</code>, <code>lewm_success_criterion</code>,
<code>velocity_source</code>, <code>pd_law</code>, <code>render_note</code> (OpenCV frames match the
published pygame frames on 99.7 % of pixels, not bit-identical),
<code>frame_convention</code>, <code>shards</code>.</p>
<h2 id="acceptance-before-extending-100-3000">Acceptance before extending 100 → 3000</h2>
<p>Run <code>scripts/contact_study.py</code>-style checks on the 100 and compare with
the 50-seed eval of the same engine + brain: solved rate, coverage
distribution, bouts per episode, situation mix, and release-reason mix
within the ranges seen there; the PD-replay assertion passes; every
episode has 201 video frames, a goal image at both sizes, ≥ 1 plan, a
narrative; no NaN; indices gapless. The sample is uploaded to HF <em>before</em>
review so nothing waits on it.</p></body></html>