pusht-simulator / docs /html /ds_distribution_report.html
desertmouse's picture
pages
473cc8a verified
Raw History Blame Contribute Delete
8.53 kB
<!doctype html><html><head><meta charset='utf-8'><title>Dataset `wall05-contact2`: distribution of the first 1000, the gap check, and the last 1000</title><style>
body{max-width:980px;margin:2rem auto;padding:0 1.25rem;font:15px/1.55 -apple-system,Segoe UI,Roboto,sans-serif;color:#1c2330;background:#fafbfc}
h1{font-size:1.7rem;border-bottom:2px solid #d8dee8;padding-bottom:.4rem}h2{margin-top:2.2rem;border-bottom:1px solid #e3e8ef;padding-bottom:.25rem}
table{border-collapse:collapse;margin:1rem 0;font-size:13.5px;font-variant-numeric:tabular-nums}th,td{border:1px solid #d8dee8;padding:.35rem .6rem;text-align:left}
th{background:#eef2f7}td:not(:first-child){text-align:right}code{background:#eef2f7;padding:.1rem .3rem;border-radius:3px;font-size:.92em}
pre{background:#1c2330;color:#e6edf5;padding:.8rem 1rem;border-radius:6px;overflow:auto}pre code{background:none;color:inherit}
img{max-width:100%;border:1px solid #d8dee8;border-radius:4px;margin:.5rem 0}blockquote{border-left:3px solid #6aa9ff;margin:0;padding:.2rem 1rem;color:#4b5563}
nav{font-size:13px;margin-bottom:1.5rem}nav a{margin-right:1rem}
</style></head><body><nav><a href="index.html">index</a> <a href="brain3_vs_3a_events.html">brain3 vs 3a events</a> <a href="contact_friction_study.html">contact friction study</a> <a href="convergence_study.html">convergence study</a> <a href="dataset_definition.html">dataset definition</a> <a href="ds08_distribution_at_1000.html">ds08 distribution at 1000</a> <a href="ds_distribution_at_1000.html">ds distribution at 1000</a> <a href="ds_distribution_report.html">ds distribution report</a> <a href="ds_mu03_brain3b_distribution.html">ds mu03 brain3b distribution</a> <a href="episode_store.html">episode store</a> <a href="human_pushing_analysis.html">human pushing analysis</a> <a href="lewm_parity.html">lewm parity</a> <a href="reference_parity.html">reference parity</a> <a href="session_fidelity_and_behaviour.html">session fidelity and behaviour</a></nav><h1 id="dataset-wall05-contact2-distribution-of-the-first-1000-the-gap-check-and-the-last-1000">Dataset <code>wall05-contact2</code>: distribution of the first 1000, the gap check, and the last 1000</h1>
<p>The plan was: generate the first 1000 episodes, measure how their starts
are distributed, find any under-represented region, and if one exists,
choose the last 1000 seeds to fill it. This report records what was found
at each step. <strong>Conclusion up front: the first 1000 are a uniform draw to
within sampling noise on every axis tested, so no correction was applied;
the last 1000 were run on the seeds as originally drawn, and they match
the first 1000.</strong> Full per-bin tables: <code>ds_distribution_at_1000.md</code>.</p>
<h2 id="1-what-correctly-distributed-means-here">1. What "correctly distributed" means here</h2>
<p>The seeds are <code>random.Random("pusht-wall05-contact2").sample(range(100_000,
1_000_000), 3000)</code>, and each seed produces a gym-pusht reset: block
position integer-uniform on [100, 400)², block angle uniform on [−π, π),
pusher uniform on [50, 450)². The goal is fixed. So the <em>target</em>
distribution is not a guess - it is exactly computable, and that is what
the first 1000 were tested against:</p>
<table>
<thead>
<tr>
<th>quantity</th>
<th>analytic expectation</th>
</tr>
</thead>
<tbody>
<tr>
<td>signed angle error (goal − block)</td>
<td>uniform, 1/24 per 15° bin</td>
</tr>
<tr>
<td>goal direction in the block frame</td>
<td>uniform, 1/24 per 15° bin (independent of position)</td>
</tr>
<tr>
<td>angle band</td>
<td>aligned 5.6 %, mixed 44.4 %, flipped 38.9 %, inverted 11.1 %</td>
</tr>
<tr>
<td>position band (CoG to goal CoG)</td>
<td>near 5.6 %, moderate 29.3 %, far 65.1 %</td>
</tr>
<tr>
<td>near a wall (&lt; 115 px)</td>
<td>24.8 %</td>
</tr>
</tbody>
</table>
<p>Tests: chi-square goodness-of-fit per histogram, plus a one-sided binomial
test per bin for <em>under</em>-representation, Bonferroni-corrected across the
bins of each histogram.</p>
<h2 id="2-the-first-1000">2. The first 1000</h2>
<table>
<thead>
<tr>
<th>histogram</th>
<th>bins</th>
<th>chi-square p</th>
<th>any bin under-represented?</th>
</tr>
</thead>
<tbody>
<tr>
<td>signed angle error</td>
<td>24</td>
<td>0.20</td>
<td>no (smallest one-sided p 0.003 vs threshold 0.002)</td>
</tr>
<tr>
<td>|angle error|</td>
<td>12</td>
<td>0.30</td>
<td>no</td>
</tr>
<tr>
<td>position error</td>
<td>10</td>
<td>0.77</td>
<td>no</td>
</tr>
<tr>
<td>goal direction in block frame</td>
<td>24</td>
<td>0.78</td>
<td>no</td>
</tr>
<tr>
<td>wall distance</td>
<td>8</td>
<td>0.78</td>
<td>no</td>
</tr>
<tr>
<td>angle band</td>
<td>4</td>
<td>0.16</td>
<td>no</td>
</tr>
<tr>
<td>position band</td>
<td>3</td>
<td>0.73</td>
<td>no</td>
</tr>
<tr>
<td>near-wall share</td>
<td>2</td>
<td>0.48</td>
<td>no</td>
</tr>
<tr>
<td>angle band × position band</td>
<td>12</td>
<td>0.55</td>
<td>no</td>
</tr>
<tr>
<td>situation class (16 labels)</td>
<td>16</td>
<td>0.37</td>
<td>no</td>
</tr>
</tbody>
</table>
<p>Observed vs expected on the bands: aligned 4.8 % (5.6), mixed 46.9 %
(44.4), flipped 36.2 % (38.9), inverted 12.1 % (11.1); near a wall 23.8 %
(24.8). The one bin that looks low - signed angle error in [120°, 135°),
25 observed vs 41.7 expected, one-sided p = 0.003 - is inside the
Bonferroni threshold for 24 bins (0.0021) and has no neighbour bin low
with it; it is what one expects to see once in ~116 bins.</p>
<p>Data integrity on the same 1000: every episode has 200 logged steps, at
least one plan, a non-empty video, finite record fields, a seed from the
planned list, and no seed repeats.</p>
<p><img alt="" src="../figures/ds_distribution/start_signed_angle.png" />
<img alt="" src="../figures/ds_distribution/start_goal_direction_block_frame.png" />
<img alt="" src="../figures/ds_distribution/angle_x_position.png" /></p>
<h2 id="3-the-gap">3. The gap</h2>
<p><strong>None beyond sampling noise.</strong> Every histogram is consistent with the
analytic reset distribution; no band, cell or situation label is
under-represented after correction. The sampler does what it says.</p>
<p>The correction step therefore had nothing to correct. It was prepared
anyway: a rebalancing routine that would draw candidate seeds from a
disjoint range (1 000 000-10 000 000), instantiate only the <em>start</em> of
each (<code>make_env("pymunk-wall", seed).get_state()</code>, no rollout), classify
it, and pick 1000 that bring the 3000 closest to the target. It was not
invoked, because invoking it on a distribution already at target would
have replaced a uniform draw with a hand-shaped one and made the dataset
<em>less</em> faithful to the published reset, not more.</p>
<h2 id="4-the-last-1000">4. The last 1000</h2>
<p>Run on seeds 2000-2999 of the original list. Compared with the first 1000:</p>
<table>
<thead>
<tr>
<th></th>
<th>first 1000</th>
<th>middle 1000</th>
<th>last 1000</th>
</tr>
</thead>
<tbody>
<tr>
<td>near a wall</td>
<td>23.8 %</td>
<td>25.5 %</td>
<td>22.9 %</td>
</tr>
<tr>
<td>|angle error| uniform, chi-square p</td>
<td>0.28</td>
<td>0.73</td>
<td>0.75</td>
</tr>
<tr>
<td>solved</td>
<td>8.9 %</td>
<td>11.9 %</td>
<td>10.9 %</td>
</tr>
<tr>
<td>final coverage mean</td>
<td>0.627</td>
<td>0.621</td>
<td>0.632</td>
</tr>
</tbody>
</table>
<p>Two-sample KS, first 1000 vs last 1000: start angle error D = 0.030,
p = 0.76; start position error D = 0.030, p = 0.76; final coverage
D = 0.049, p = 0.18. The thirds are draws from the same distribution.</p>
<h2 id="5-the-whole-3000">5. The whole 3000</h2>
<p>Acceptance check against the 50-seed eval of the same engine and brain:
solved 10.6 % (eval 12.0 %), coverage mean 0.627 (0.584), median 0.678
(0.623), bouts 3.69 (3.72), plans 4.29 (4.34); KS on final coverage
p = 0.25; situation mix mixed 45.7 / flipped 37.7 / inverted 11.6 /
aligned 4.6 / already-nearly-solved 0.3 %; <strong>0 artifact problems</strong>.</p>
<p>What the dataset is therefore good for, and what it is not: it is an
unbiased sample of the published reset distribution on the friction
engine with the rule brain - the right thing for training or for
measuring a policy against the same starts the published benchmark uses.
It is <em>not</em> balanced by difficulty: 11.6 % of starts are inverted and the
brain solves none of them, so a model trained on it sees ~350 inverted
failures and few inverted successes. A difficulty-balanced companion set
would be a different sampler, stated as such, not a correction of this one.</p></body></html>