loglens-learnability / EXPECTED_RESULTS.md
resoajoe's picture
Upload folder using huggingface_hub
021d920 verified
|
Raw
History Blame Contribute Delete
3.13 kB

Expected results — what a successful replication looks like

Our measured values (AGX Orin, 2026-07; single scene family unless noted). Your rig/scene should reproduce the STRUCTURE (the PASS bars), not the exact numbers — absolute R and the learnability threshold are scene-local by design.

The three PASS bars (what replicate.py checks)

# test bar our value(s)
P1 Spearman(R, feat-R) >= 0.90 0.981 (home), 0.969 (2nd scene)
P2 Spearman(R, learnability) — THE LAW >= 0.85 0.988 (10-cell blind), 0.884 (feat-R), 0.80 (2nd scene, weak)
P3 ignition: learnable cell escapes (loss < log 4 = 1.386), floor cell pinned at log 4 present confirmed on CNN + ViT, home + 2nd scene

What we also established (context for interpreting your run)

  • MECHANISM: R = init feature separability. The untrained encoder preserves & amplifies the pixel signal/clutter ratio into feature space (that's P1). When a model learns, the SIGNAL component of feat-R amplifies ~1e5x and the head's margin ramps; when it collapses, signal stays ~0 and loss sits at log 4.
  • CROSS-ARCH: the ignition fires in a ViT too (not a CNN quirk).
  • CROSS-SCENE: R's ORDERING is scene-universal (Spearman .996 between two scenes); the THRESHOLD is scene-local (a low-R cell can be learnable on an easier scene). So compare ORDERING, never absolute thresholds, across rigs.

CAVEATS that will bite you if ignored (hard-won this session)

  1. THE WALL IS SEED-BIMODAL / SLOW. A cell "collapses" at 8k steps and ESCAPES at 18k. Escape times spread 6k-18k. => ALWAYS multi-seed (>=2, ideally 3) and train long (--full = 20k). Single-seed near the wall is noise, not signal.
  2. COLLAPSE = loss ~ log(K) = 1.386 exactly (uniform-guess basin), acc ~ 1/K. Use this to detect collapse with zero eval and early-abort to save compute.
  3. feat-R (the ratio) is CONFOUNDED near collapse: it can spike when the encoder degenerates (clutter->0), NOT from signal. Track the SIGNAL component, or just the head margin/loss, to judge escape — not the raw ratio.
  4. A cell that floors on our scene may learn on yours (easier scene, lower threshold). That is EXPECTED (scene-local threshold), not a contradiction.

Verdict rules

  • P1 & P2 & P3 PASS on your independent rig+scene -> the law + mechanism replicate. This is the result we most want confirmed or killed.
  • P2 fails (ordering doesn't predict) -> R is scene-specific / overfit. A real, publishable negative. Send it.
  • P1 passes, P3 fails -> feature separability holds but learning doesn't ignite -> optimization/scale issue on your rig; note steps, seeds, lr. Send back results_replicate.json + your scene_hash. Disagreement is a RESULT.

Note on synthetic timestamps (kit internals)

replicate.py spreads your images over a synthetic ~96-hour timeline, because the task dates changes over a 24-hour age horizon and needs memory slots at ages up to 24h. If you build a bank yourself, give it a timestamp array spanning well beyond 86,400,000 ms — 1-second spacing silently breaks episode sampling.