pusht-simulator / docs /html /session_fidelity_and_behaviour.html
desertmouse's picture
pages
473cc8a verified
Raw History Blame Contribute Delete
22.4 kB
<!doctype html><html><head><meta charset='utf-8'><title>Fidelity, banging, and what humans actually do: the reference-engine investigation</title><style>
body{max-width:980px;margin:2rem auto;padding:0 1.25rem;font:15px/1.55 -apple-system,Segoe UI,Roboto,sans-serif;color:#1c2330;background:#fafbfc}
h1{font-size:1.7rem;border-bottom:2px solid #d8dee8;padding-bottom:.4rem}h2{margin-top:2.2rem;border-bottom:1px solid #e3e8ef;padding-bottom:.25rem}
table{border-collapse:collapse;margin:1rem 0;font-size:13.5px;font-variant-numeric:tabular-nums}th,td{border:1px solid #d8dee8;padding:.35rem .6rem;text-align:left}
th{background:#eef2f7}td:not(:first-child){text-align:right}code{background:#eef2f7;padding:.1rem .3rem;border-radius:3px;font-size:.92em}
pre{background:#1c2330;color:#e6edf5;padding:.8rem 1rem;border-radius:6px;overflow:auto}pre code{background:none;color:inherit}
img{max-width:100%;border:1px solid #d8dee8;border-radius:4px;margin:.5rem 0}blockquote{border-left:3px solid #6aa9ff;margin:0;padding:.2rem 1rem;color:#4b5563}
nav{font-size:13px;margin-bottom:1.5rem}nav a{margin-right:1rem}
</style></head><body><nav><a href="index.html">index</a> <a href="brain3_vs_3a_events.html">brain3 vs 3a events</a> <a href="contact_friction_study.html">contact friction study</a> <a href="convergence_study.html">convergence study</a> <a href="dataset_definition.html">dataset definition</a> <a href="ds08_distribution_at_1000.html">ds08 distribution at 1000</a> <a href="ds_distribution_at_1000.html">ds distribution at 1000</a> <a href="ds_distribution_report.html">ds distribution report</a> <a href="ds_mu03_brain3b_distribution.html">ds mu03 brain3b distribution</a> <a href="episode_store.html">episode store</a> <a href="human_pushing_analysis.html">human pushing analysis</a> <a href="lewm_parity.html">lewm parity</a> <a href="reference_parity.html">reference parity</a> <a href="session_fidelity_and_behaviour.html">session fidelity and behaviour</a></nav><h1 id="fidelity-banging-and-what-humans-actually-do-the-reference-engine-investigation">Fidelity, banging, and what humans actually do: the reference-engine investigation</h1>
<p>This document consolidates one thread of the workbench's development: the user reported that the scripted
pusher "bangs" into the T on <code>pymunk</code>, that <code>superdex</code> bangs too, and that <code>pymunk-reference</code> — although
much gentler — still goes back and forth. The underlying question was sharper than the symptom: <em>if the
original code is in hand, how can the values and the behaviour be a bit off?</em> Answering that required
separating three things that had been tangled together: how faithful the reference engine is, what the
scripted policy was doing, and what the human demonstrations that everyone trains on actually contain.</p>
<p>The sections below follow the order in which the questions were settled.</p>
<hr />
<h2 id="1-is-pymunk-reference-a-100-replica">1. Is <code>pymunk-reference</code> a 100% replica?</h2>
<p><strong>Verdict: it was 95%; it is now ~98%. The missing 2% is deliberate and written down.</strong></p>
<p>The audit went element by element against the published gym-pusht source, with each row backed by a
measurement rather than a reading.</p>
<table>
<thead>
<tr>
<th>element</th>
<th>grade</th>
<th>evidence</th>
</tr>
</thead>
<tbody>
<tr>
<td>PD law, kinematic pusher, <code>dt</code>, substeps</td>
<td>100</td>
<td>1e-9 against an independent re-implementation of the control loop</td>
</tr>
<tr>
<td>Block mass, inertia typo (3000), CoG (0, 45)</td>
<td>100</td>
<td>pinned by tests</td>
</tr>
<tr>
<td>Frictionless shapes, <code>damping = 0</code>, <code>iterations = 10</code>, walls, settle step</td>
<td>100</td>
<td>pinned</td>
</tr>
<tr>
<td>Full trajectories under identical actions</td>
<td>100</td>
<td>0.000e+00 over 120 steps on contact-heavy seeds</td>
</tr>
<tr>
<td>5 recorded human episodes replayed</td>
<td>100</td>
<td>≤ 0.002 px over 738 frames</td>
</tr>
<tr>
<td>Angle <code>% 2π</code>, keypoints, contact count, <code>block_cog</code> / <code>damping</code> options</td>
<td>100</td>
<td>pinned</td>
</tr>
<tr>
<td>RNG draw for seed → sampled state</td>
<td>100</td>
<td>identical numbers</td>
</tr>
<tr>
<td><strong>Start pose the workbench actually used</strong></td>
<td><strong>was wrong; fixed</strong></td>
<td>see below</td>
</tr>
<tr>
<td>Observation dtype (float32 vs float64)</td>
<td>deviation, ~1.5e-5</td>
<td>documented</td>
</tr>
<tr>
<td>Success <code>&gt;=</code> vs <code>&gt;</code>; truncation inside vs a 300-step <code>TimeLimit</code> outside</td>
<td>deviation</td>
<td>documented</td>
</tr>
<tr>
<td>Rendering (OpenCV vs pygame)</td>
<td>deviation</td>
<td>pixel observations are not the dataset's pixels</td>
</tr>
</tbody>
</table>
<h3 id="the-defect-the-user-was-sensing-was-real">The defect the user was sensing was real</h3>
<p>gym-pusht's <code>reset(seed=783958)</code> places the block at <strong>(231.16, 332.93)</strong>. The workbench placed it at
<strong>(271, 267)</strong>. The published <em>draw</em> had been reproduced exactly, but the sampled pose was then applied
through the workbench's own round-tripping setter (angle first) instead of gym-pusht's legacy setter
(position, <em>then</em> angle). Because pymunk rotates a body about its centre of gravity, that order shifts the
block by roughly 75 px. Every reference episode watched up to that point had therefore started from a pose
that seed never produces upstream — a genuine "the values are a bit off".</p>
<p>The fix routes a <em>sampled</em> start through the legacy setter, exactly as upstream does, while an explicit
<code>reset(state=...)</code> keeps round-tripping (that path is for everything that is not a published seed).
The proof is on the live server, not only in tests:</p>
<pre><code>live pymunk-reference, seed 783958 -&gt; block (231.162, 332.927) angle 4.229
gym-pusht reset(seed=783958) -&gt; block (231.162, 332.927) angle 4.229
</code></pre>
<p>Two tests that had pinned the old, wrong behaviour were rewritten to assert the corrected semantics. The
cross-version parity harness (ours on pymunk 7.3.0, the published env on 6.11.1 in its own virtualenv —
it cannot run on pymunk 7 because it calls a <code>Space</code> API that was removed) still reports 0.000e+00.</p>
<h3 id="the-banging-is-0-the-reference">The banging is 0% the reference</h3>
<p>Neither gym-pusht nor LeWM ships a controller that approaches the block. gym-pusht's data was driven by
human teleoperation; LeWM's <code>WeakPolicy</code> is seeded random relative offsets clipped to the block's
neighbourhood. The retreat-and-ram cycle is entirely the workbench's own <code>ScriptedPushPolicy</code>. Measured on
seed 783958 over 150 control steps:</p>
<table>
<thead>
<tr>
<th>engine</th>
<th>separate contact events</th>
<th>peak approach speed</th>
</tr>
</thead>
<tbody>
<tr>
<td>pymunk</td>
<td>25</td>
<td>348 px/s</td>
</tr>
<tr>
<td>superdex</td>
<td>17</td>
<td>364 px/s</td>
</tr>
<tr>
<td>pymunk-reference</td>
<td>10</td>
<td>234 px/s</td>
</tr>
<tr>
<td>lewm</td>
<td>11</td>
<td>239 px/s</td>
</tr>
</tbody>
</table>
<p><code>pymunk-reference</code> only <em>looks</em> gentler because a kinematic pusher cannot bounce off the block. So "you
have the original code but the behaviour is off" resolves cleanly: the <em>environment</em> is a replica; the
<em>thing pushing in it</em> was never part of the original.</p>
<hr />
<h2 id="2-does-a-proper-approach-policy-exist-anywhere-public">2. Does a proper approach policy exist anywhere public?</h2>
<p>The rule set for the search was strict: a policy counts only if its code is readable or its generated data is
inspectable. Hearsay and README claims without artefacts do not count.</p>
<p><strong>Verdict: no public, working scripted approach-and-push controller for the pymunk Push-T exists.</strong></p>
<table>
<thead>
<tr>
<th>candidate</th>
<th>kind</th>
<th>code</th>
<th>data</th>
<th>holds contact?</th>
<th>status</th>
</tr>
</thead>
<tbody>
<tr>
<td>LeRobot <code>pusht</code> (206 episodes)</td>
<td>human teleop</td>
<td>—</td>
<td>yes</td>
<td>yes, human</td>
<td>verified</td>
</tr>
<tr>
<td>DINO-WM <code>pusht_noise</code> (18.5k)</td>
<td><strong>human demos + action noise</strong></td>
<td>—</td>
<td>yes</td>
<td>inherited</td>
<td>verified (paper App. A.1)</td>
</tr>
<tr>
<td>LeWM <code>lewm-pusht</code> (20k)</td>
<td><strong>the same DINO-WM data</strong></td>
<td>—</td>
<td>yes</td>
<td>inherited</td>
<td>verified (paper App. E)</td>
</tr>
<tr>
<td>stable-worldmodel <code>WeakPolicy</code></td>
<td>random, clipped near the block</td>
<td>yes</td>
<td>yes</td>
<td>no</td>
<td>verified</td>
</tr>
<tr>
<td><code>jimchen2/pushT-dataset-example</code></td>
<td>scripted heuristic</td>
<td>yes</td>
<td>10 episodes</td>
<td>no</td>
<td>verified — 0/10 success, max coverage 0.47</td>
</tr>
<tr>
<td>Xie / Chen / Goldberg, <em>Revisiting Push-T with Agentic Robotics</em> (arXiv 2608.18227)</td>
<td>state-machine controller, 100% over 200 seeds</td>
<td>no</td>
<td>no</td>
<td>claimed</td>
<td><strong>unverified</strong> — "will be posted online"; nothing found on GitHub</td>
</tr>
<tr>
<td><code>interactive_world_sim</code> planner</td>
<td>scripted, MuJoCo ALOHA, random-direction pushes</td>
<td>yes</td>
<td>yes</td>
<td>yes</td>
<td>verified — different environment, not goal-directed</td>
</tr>
<tr>
<td>Model-Based Diffusion, CRISP</td>
<td>trajectory optimisers on their own physics</td>
<td>yes</td>
<td>—</td>
<td>—</td>
<td>planners, not heuristics</td>
</tr>
</tbody>
</table>
<p>Two things are worth stating plainly:</p>
<ol>
<li><strong>Every "expert" Push-T dataset in the world-model literature is the same 206 human demonstrations</strong>,
replayed with noise. There is no scripted expert behind any of them. The gentle, contact-holding motion
people have in mind is a human moving a mouse.</li>
<li>The one scripted heuristic that does exist is the same naive rule the workbench's policy started from —
go behind the block on the goal→block axis and push through — and it fails every episode.</li>
</ol>
<p>The only claimed working controller (the Goldberg-lab paper) describes plan → approach → push → retreat,
a contact library, a quasi-static model, and greedy contact selection by dominant pose error. No code or
data yet, so by the rule above it does not exist — but its architecture is the informed design, and it
explicitly includes a retreat phase, so "never retreat" is not the target either.</p>
<p>Three options were put to the user: build the controller from the paper's description, fix the
retreat-and-ram in the existing heuristic, or replay the human demonstrations as the policy. The user's
answer was a fourth: <em>measure the humans first</em>.</p>
<hr />
<h2 id="3-how-humans-push-the-t">3. How humans push the T</h2>
<h3 id="data-and-a-provenance-decision">Data and a provenance decision</h3>
<p>The LeWM paper's dataset is the DINO-WM set, which is the Diffusion Policy human teleop episodes replayed
with injected action noise to reach 18.5k–20k trajectories. The pushing skill in that data is a human's;
the noise is synthetic. For velocity, acceleration and jerk statistics the noised copies would be actively
misleading, so the analysis uses the <strong>206 clean demonstrations</strong> (LeRobot <code>pusht_keypoints</code>, 8 keypoints
per frame, 10 Hz). Block pose is recovered per frame by a rigid fit to the keypoints (max residual 3.8e-05
px); the dataset's reward column was checked against the repository's <code>coverage()</code> to 2.2e-07.</p>
<p>All 206 episodes contain contact; the first 50 were analysed and the remaining 156 held out to score the
contact-selection rule. Contact is defined as <code>distance(agent centre, T outline) - 15 &lt;= 1 px</code>; the gap
histogram has a spike in (−1, 0.5] px from the collision slop and a flat tail beyond, so the threshold is
not sensitive.</p>
<h3 id="the-headline-humans-do-not-retreat-and-ram">The headline: humans do not retreat and ram</h3>
<p>Per-frame roles were assigned from contact state, the agent's velocity split into normal and tangential
components at the nearest outline point, and block motion.</p>
<table>
<thead>
<tr>
<th>role</th>
<th>share of frames</th>
<th>median run</th>
<th>median speed (px/s)</th>
<th>median |accel| (px/s²)</th>
</tr>
</thead>
<tbody>
<tr>
<td>APPROACH</td>
<td>13%</td>
<td>4 frames</td>
<td>69</td>
<td>147</td>
</tr>
<tr>
<td>TOUCH</td>
<td>2.5%</td>
<td>1</td>
<td>57</td>
<td>—</td>
</tr>
<tr>
<td><strong>PUSH</strong></td>
<td><strong>39%</strong></td>
<td><strong>7</strong></td>
<td>48</td>
<td>88</td>
</tr>
<tr>
<td>SLIDE</td>
<td>7%</td>
<td>2</td>
<td>50</td>
<td>145</td>
</tr>
<tr>
<td>RETREAT</td>
<td><strong>2.7%</strong></td>
<td>2</td>
<td>106</td>
<td>567</td>
</tr>
<tr>
<td>REPLAN / orbit</td>
<td>35%</td>
<td>8</td>
<td>122</td>
<td>305</td>
</tr>
</tbody>
</table>
<p>Retreat is 2.7% of all motion. What replaces it is <strong>orbiting</strong>: 72% of no-contact motion is tangential to
the outline at a gap of about 37 px, followed by a four-frame radial approach. A typical episode has three
contact bouts of about 17 frames, with 16 frames between them. The dominant role sequence is
idle → approach → touch → push → slide → push.</p>
<h3 id="touching">Touching</h3>
<ul>
<li><strong>Head-on</strong>: incidence 17° to the inward normal; 72% of touches within 30°; only 7% glancing.</li>
<li><strong>From behind, as seen from the goal</strong>: 24° off the goal→block axis at touch.</li>
<li><strong>Slowing in</strong>: arrival speed 67 px/s, about half the orbit speed.</li>
<li><strong>Edge chosen by goal geometry, not proximity</strong>: the edge nearest the pusher at the previous release is
used only 9% of the time.</li>
</ul>
<h3 id="pushing-the-strongest-result">Pushing — the strongest result</h3>
<p>Block rotation follows the lever arm of the push about the CoG: correlation <strong>r = 0.87</strong>, slope 0.56 °/s
per px, and the sign is correct in <strong>100%</strong> of frames in which the block rotates faster than 2 °/s. The
lever is chosen deliberately, scaled to the angle error:</p>
<table>
<thead>
<tr>
<th>|angle error|</th>
<th>median lever from the CoG line</th>
<th>corrective sign</th>
</tr>
</thead>
<tbody>
<tr>
<td>0–5°</td>
<td><strong>5.7 px</strong> (through the CoG)</td>
<td>57%</td>
</tr>
<tr>
<td>10–20°</td>
<td>34 px</td>
<td>92%</td>
</tr>
<tr>
<td>30–45°</td>
<td><strong>42.5 px</strong></td>
<td>99%</td>
</tr>
</tbody>
</table>
<p>A released bout takes the position error from 96 px to 48 px and the angle error from 53° to <strong>5°</strong>, with
0.95 of the CoG displacement along the goal direction.</p>
<h3 id="release-when-to-let-go">Release — when to let go</h3>
<p>Release is triggered by geometry, not by distance. Push alignment (the angle between the push normal and
the CoG→goal direction) is 26° at touch and <strong>89° at release</strong> — above 45° in 91% of releases. The block is
stopped first (0.4 px/s on the release step); the pusher then swings about 89° around the block and touches
again roughly 16 frames later.</p>
<h3 id="choosing-the-next-contact">Choosing the next contact</h3>
<p>The unfitted rule <em>"push where the inward normal points at the goal, as close to the CoG line as possible"</em>
predicts the human's edge in <strong>57% of touches under cross-validation and 49% on the external hold-out</strong>
(540 touches), against 12.5% chance and 23% for always choosing the majority edge. A fitted contact-point
model (alignment, torque × angle error, off-centre penalty; weights 1.79, 2.97, 1.92) reaches 56% / 51% and
halves the perimeter error (24 px vs 32 px). The torque term does not change <em>which</em> edge is chosen; it
moves the contact point <em>along</em> it. A mirrored decision table (goal direction in the block frame × angle
error) is in the full report.</p>
<h3 id="what-this-means-for-a-controller">What this means for a controller</h3>
<p>Orbit at ~120 px/s at a 37 px gap → four-frame normal approach at ~70 px/s → push at ~48 px/s with a lever
of 0 px below 5° error and ~42 px above 10° → slide (two frames) to re-aim while alignment stays below 45°
→ release when alignment exceeds 45–60° with the block stopped → swing ~90° → next contact from the
goal-normal rule.</p>
<hr />
<h2 id="4-which-findings-are-physics-and-which-are-the-human">4. Which findings are physics and which are the human</h2>
<p>This was the user's sharpest question: of everything in the analysis, what is contributed by the already
defined elements — PD law, kinematic pusher, <code>dt</code>, inertia, damping — and what by the operator?</p>
<h3 id="the-split-is-measurable-not-arguable">The split is measurable, not arguable</h3>
<p>The environment is a deterministic function of the action stream, so the two contributions can be
separated by computation. The pure PD law was re-simulated in free space from each recorded start, fed only
the recorded actions, and compared with the recorded agent path:</p>
<table>
<thead>
<tr>
<th>frames</th>
<th>median error</th>
<th>p90</th>
</tr>
</thead>
<tbody>
<tr>
<td>no contact (n = 3065)</td>
<td><strong>0.00 px</strong></td>
<td>0.00</td>
</tr>
<tr>
<td>in contact (n = 3120)</td>
<td><strong>0.00 px</strong></td>
<td>0.00</td>
</tr>
</tbody>
</table>
<p>Zero <em>in contact</em> is the decisive number: the pusher is kinematic, the block never acts on it, and the
entire agent trajectory is the PD law applied to the mouse. Every agent-side statistic therefore factors as</p>
<p>$$\text{agent path} = \underbrace{\text{PD}(k_p{=}100,\ k_v{=}20,\ dt{=}0.01)}<em>{\text{physics, fixed}} \circ \underbrace{\text{mouse target stream}}</em>{\text{human}}$$</p>
<h3 id="what-the-physics-contributes-fixed-by-the-reference-would-change-on-any-other-engine">What the physics contributes (fixed by the reference; would change on any other engine)</h3>
<p>The PD response to a 100 px step command is 29.6 → 61.6 → 80.8 → 95.6 px over four control steps, with no
overshoot, 90% settled at step four, and a peak of 320 px/s: a critically damped lag of about one control
step. Its consequences are not human decisions:</p>
<table>
<thead>
<tr>
<th>element</th>
<th>what it puts into the data</th>
</tr>
</thead>
<tbody>
<tr>
<td>PD law, <code>dt</code>, substeps</td>
<td>a one-step lag and a ~10 px lead (median |target − agent| = 10.1 px); <strong>halves the human's acceleration and quarters the jerk</strong> — mouse |accel| medians are 224 / 300 / 500 / 1153 px/s² for push / approach / orbit / retreat, the agent's are 88 / 147 / 305 / 567. About half of the "gentle" look is filter.</td>
</tr>
<tr>
<td>kinematic pusher</td>
<td>"push depth" carries no force. The block is displaced by a moving wall; "push at 48 px/s, block follows at 32" is kinematic slip, not contact mechanics.</td>
</tr>
<tr>
<td>inertia typo (3000 vs 7875 geometric), CoG</td>
<td>the 0.56 °/s per px rotation gain is rigid-body kinematics under the understated moment; with the correct inertia the same lever would give roughly 40% of the rotation rate.</td>
</tr>
<tr>
<td><code>damping = 0</code></td>
<td>the block stops on the release step (0.4 px/s) because velocity is annihilated every substep — not a skill.</td>
</tr>
<tr>
<td>frictionless shapes</td>
<td>no tangential drag, so slides are cheap.</td>
</tr>
</tbody>
</table>
<h3 id="what-the-human-contributes-portable-across-engines">What the human contributes (portable across engines)</h3>
<p>Measured on the raw commanded targets before any filtering. In every role the agent goes exactly where the
mouse points (<code>cos(lead, motion) = 1.00</code>), so these are genuinely the operator's decisions:</p>
<table>
<thead>
<tr>
<th>decision</th>
<th>evidence from the mouse stream</th>
</tr>
</thead>
<tbody>
<tr>
<td>speed by phase: slow in contact, fast when free</td>
<td>45 px/s (push), 41 (touch), 51 (slide) vs 114 (orbit), 184 (retreat)</td>
</tr>
<tr>
<td>orbit rather than retreat</td>
<td>72% tangential free motion; retreat 2.7% of frames; standoff 37 px</td>
</tr>
<tr>
<td>head-on aim</td>
<td>17° incidence, 72% within 30°</td>
</tr>
<tr>
<td>edge selection by goal geometry</td>
<td>goal-normal rule 57% / 49%; nearest edge 9%</td>
</tr>
<tr>
<td>lever offset scaled to angle error</td>
<td>5.7 px at &lt; 5°, 42.5 px at 30–45°, corrective 99%</td>
</tr>
<tr>
<td>release on geometry, not distance</td>
<td>alignment 26° at touch, 89° at release</td>
</tr>
<tr>
<td>pacing</td>
<td>~3 bouts of ~17 frames, 16 frames between</td>
</tr>
</tbody>
</table>
<h3 id="consequences-for-the-controller">Consequences for the controller</h3>
<ul>
<li>The speeds in the design section are <em>agent</em> speeds. On a kinematic pusher the controller must emit
<em>target-stream</em> speeds — the mouse numbers — or it double-filters. On a dynamic pusher (<code>pymunk</code>,
<code>superdex</code>) the mapping is different again.</li>
<li>The rotation gain is reference-only. Do not feed forward 0.56 °/s per px; close the loop on the measured
angle error.</li>
<li>Edge, lever, orbit, incidence and release rules are the transferable content. Lag, smoothing,
stop-on-release and the rotation gain are not.</li>
</ul>
<hr />
<h2 id="5-where-this-leaves-the-work">5. Where this leaves the work</h2>
<table>
<thead>
<tr>
<th>item</th>
<th>state</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>pymunk-reference</code> start pose</td>
<td>fixed, proven on the live server, pinned on seed 783958</td>
</tr>
<tr>
<td>gym-pusht parity</td>
<td>0.000e+00 across pymunk majors; deviations documented</td>
</tr>
<tr>
<td>public approach policy</td>
<td>none exists; the only claimed one has no code or data</td>
</tr>
<tr>
<td>human behaviour analysis</td>
<td><code>docs/human_pushing_analysis.md</code>, <code>scripts/analyze_human_pushing.py</code>, 15 figures</td>
</tr>
<tr>
<td>physics-vs-human decomposition</td>
<td>§8 of that report</td>
</tr>
<tr>
<td>scripted policy rewrite</td>
<td><strong>not started</strong> — the design above is the spec; the decision to build it is the user's</td>
</tr>
</tbody>
</table></body></html>