BurnyCoder's picture
Publish fact-checked research documentation and independent audit evidence
76b9910 verified
|
Raw
History Blame Contribute Delete
22.9 kB

Experiment register and learnings

Current evidence

The documentation audit records the later fact-check corrections and their evidence. Locally dated protocol documents and saved phase specifications are distinguished from independent registration. Generated numerical reports remain historical snapshots; narrative terminology is corrected here without changing results.

Featured: Ant experiment 005. The locked vector at βˆ’0.1 changes world-ground lateral velocity by +0.40238 m/s, paired 95% interval [+0.30803, +0.49127], on thirty fresh confirmation episodes while retaining 96.83% of forward speed. Failure counts are 4/30 in both arms. This is a measured causal effect in the tested frozen policy; it is not a claim of body-frame sidestepping or a newly discovered internal concept.

Its recorded status is nevertheless confirmation_failed, solely because an additional strict rule forbids any episode ending before steering: seed 520025 ends identically at step 34 in both arms. The effect, movement preservation, and allowed quality increases pass. Equal aggregate failure counts do not mean identical failed-seed identities. No replication or application was run. See 005 findings and three attached videos, its complete bundle, the separate within-project causal/code audit, and the standalone research paper / PDF.

Watch the tracking, top-down, and fixed far-camera views of the same predetermined validation pair. The Ant vectors and controls are published on Hugging Face, with pinned provenance and numerical results. Videos illustrate existing data and do not add independent evaluation trials.

The first experiment, hc-classic-001, was rejected at fitting review before validation. Ten near-stationary windows, only 1.9% of the initially eligible fitting windows, substantially changed all three extracted directions when removed. No causal steering outcome has been measured for this experiment. See its artifact bundle and the separate diagnostic analysis.

Follow-up hc-running-002 completed 120 validation conditions after correcting an episode-failure measurement problem. None of the 24 primary settings passed every gate for its registered target. The height vector at +0.05 raised the torso 4.5048 cm without measured physical failure, contact, or inversion, but its 12.85% slowdown exceeded the height target's 10% movement allowance. Eight control settings passed exploratory gates. No selection, confirmation, or replication was performed. See the findings and complete bundle.

hc-height-speed-003 tested that frozen height-derived direction for speed control on 40 fresh validation conditions. All eight primary strengths failed physical quality; +0.05 produced two physical failures in ten trials. Two settings of the reused random control passed exploratory gates. No primary candidate was selected and no held-out data were consumed. No confirmed useful vector is claimed from this experiment.

hc-height-speed-004 completed with failed confirmation. Six primary strengths passed calibration; the locked βˆ’0.02 setting slowed the policy by 31.89% on thirty fresh episodes but physically failed in ten. Two controls passed confirmation pilot gates; none received replication. See 004 findings and its complete bundle.

The 002 numerical bundle was published at Hugging Face commit 9fa4bb9. The 003 raw run is at commit 238bd2d, with numerical reports at commit d4d4e76. The 004 run data and numerical reports are at commit dc0748d. Those pins lack the local 002 quality-analysis and 003/004 vector-import audit files; the documentation audit tracks their publication refresh. Publication status is separate from experimental success; later narrative corrections have their own version history.

Each research run retains its numerical evidence; the separate report command writes its Markdown/PDF report. The table below records decisions, including rejected extractions, and links the detailed reasoning instead of duplicating complete results.

Experiment Question / hypothesis Observed outcome Decision and learning
hc-classic-001 Do health-filtered speed, effort, and height contrasts describe sustained running? 16 diagnostic and 64 fitting episodes; 517/576 fitting windows initially eligible. Removing 10 near-stationary windows rotated the speed and height directions strongly and changed the effort direction to negative cosine similarity. No validation was performed. Reject the initial construction. Upright/contact checks and broad episode contributor counts do not prevent unusual observation magnitudes from influencing activation means. See sensitivity evidence.
hc-running-002 Do sustained-running contrasts yield useful control of their speed, effort, or height target? 64 fresh fitting episodes; 564/576 eligible windows, with zero additional speed-floor exclusions. All 24 primary settings fail at least one gate; height +0.05 gives a clean +4.5048 cm change but slows 12.85%. Eight control settings pass validation only. Record a causal posture effect, with no confirmed useful intervention. A mid-validation physical-failure measurement correction catches a collapse hidden by the across-episode mean; all 120 conditions were reanalyzed and the old 74 estimates preserved. See the separate inspection chronology.
hc-height-speed-003 Can the exact 002 height-derived direction provide useful forward-speed control, regardless of its extraction label? 40 fresh validation conditions on seeds 210000–210009. All eight primary strengths fail quality; +0.05 slows 17.01% but has two physical failures among ten trials. Two source height_random0 strengths pass validation. No primary candidate, selection, or held-out evaluation. Unchanged vectors can fail physically on fresh resets even after a clean earlier sample. Preserve positive control evidence and distinguish a causal slowdown from reliable useful control.
hc-height-speed-004 Can smaller signed strengths of the unchanged 002 height vector retain a useful speed effect while avoiding physical failures? 72 validation conditions, including action bias; six primary strengths pass. Locked βˆ’0.02 confirmation has βˆ’31.89% speed change but 10/30 physical failures. Two confirmation controls pass pilot gates; no replication. Record confirmation_failed and preserve selection. The largest eligible validation effect did not retain physical quality on fresh resets; other positive calibration settings were not substituted after the failure. The locally dated protocol retains the original grid and decision rules with timestamp qualifications.
ant-classic-005 Does a frozen Ant policy support useful lateral or turning steering from sustained natural contrasts? 88 validation and six confirmation conditions. Locked lateral βˆ’0.1 has +0.40238 m/s held-out effect, 96.83% forward-speed retention, and 4/30 failures in each arm. No turning primary passes validation. Record the positive causal lateral effect and confirmation_failed separately. One shared pre-intervention failure alone violates an additional strict criterion; no thresholds or exclusions changed. No replication or application. The locally dated protocol and fitting diagnostics retain the prior reasoning.

Recorded partitions for the latest experiments

These seed sets were assigned before the corresponding phase outcomes. Assignment does not mean a phase has completed. The full per-run manifests remain the machine-readable record.

Experiment Diagnostic and fitting provenance Fresh validation Fresh confirmation Fresh replication Conditional application
hc-height-speed-004 Reuse 002: diagnostic 100000–100015 and fitting 101000–101063; no refitting 310000–310009, completed 320000–320029, completed 330000–330029, unused 340000–340009, unused
ant-classic-005 Diagnostic 500000–500015; fitting 501000–501063 510000–510009, completed 520000–520029, completed 530000–530029, unused No application run; a lateral application requires successful replication and its own locked specification

Policies considered

Policy Actor architecture Training algorithm and objective Training data Curriculum Steering accuracy / useful effect Episode completion / failure Domain Evidence status
Farama-Minari HalfCheetah-v5 TQC medium Two 256-unit ReLU hidden layers, verified on load TQC online RL; environment return Upstream policy-environment interactions; no action demonstrations used here Not established by this audit 001 rejected before validation; 002/003 have no passing primary selection. 004's smaller additions pass calibration, but the locked primary fails confirmation. Some controls pass pilot gates 004 selected βˆ’0.02 has zero failures in ten validation episodes, then ten failures in thirty confirmation episodes Simulated planar locomotion No replicated primary intervention or application result; complete negative evidence retained
Farama-Minari Ant-v5 SAC medium Two 256-unit ReLU hidden layers, verified on load SAC online RL; environment return Upstream policy-environment interactions; no action demonstrations used here Not established by this audit 005 lateral βˆ’0.1 gives +0.40238 m/s on fresh held-out episodes and retains 96.83% forward speed Confirmation failure counts 4/30 in both arms; one identical pair ends before intervention. Fitting completion 56/64 Simulated quadruped locomotion Causal lateral effect measured; strict pre-onset criterion blocks formal acceptance, replication and application

Checkpoint links and training provenance are in the source audit. Classification accuracy and token fill/completion rates do not apply to these continuous-control policies. Reports instead measure behavioral change, episode completion/failure, quality, and uncertainty.

Initial design learning

The closest directly reusable intervention component found was Baukit Trace: it edits arbitrary PyTorch layer outputs and works with ordinary actor prediction calls. NNsight and pyvene support generic models but add integration machinery; the steering-vectors package assumes language-model token dimensions. This motivated reusing Baukit for hook mechanics and writing only locomotion measurements, contrasts, and evaluation. A plain-function adapter resolved Baukit's handling of bound-method signatures during integration; the prediction and cleanup behavior subsequently passed runtime tests. The exact source audit and fix rationale are in sources.md.

HalfCheetah's initial eligible fitting windows had a median speed of 14.593 m/s and an interquartile range of 13.899-15.192 m/s. Variation was present. The important obstacle was contrast composition: low-speed recovery histories could have unusually large hidden activations even when the current health proxies passed. The revised protocol preserves the original trajectories and separates fitting exclusions from evaluation, where failed episodes must remain counted.

The second review showed why averaging contact fractions across episodes can hide a collapse: one trial's long failure is diluted by nine clean trials. The corrected analysis first labels physically failed episodes, then averages those labels while retaining episode-averaged contact/inversion fractions. It averages each episode's available timesteps first, then weights episodes equally; this is not general pooled timestep averaging. The inspection and old/new analysis audit preserve the validation-informed correction before held-out evaluation. A low across-episode mean does not establish competence in every rollout.

Experiment 002 also separates an extraction label from its causal effects. Its height contrast was associated with lower speed in fitting, and its positive intervention changed both posture and speed. Strong split-half extraction agreement did not ensure a useful intervention for the original target. A random direction also changed height within the original pilot gates, so semantic specificity is unsupported. Experiment 003 targeted the observed speed effect explicitly on new episodes, then rejected the coarse strengths for physical failures. A clean ten-episode sample was not sufficient evidence of reliable control across resets. Both studies retain random-direction successes as exploratory evidence, without relabeling a failed primary experiment as confirmed success.

Experiment 004 made this failure of generalization explicit after a locked selection. Several smaller positive strengths passed calibration, but the larger eligible slowdown at βˆ’0.02 won the specified selection rule and then failed physical quality on confirmation. The favorable alternatives remain calibration evidence, not replacement confirmations. One confirmed shuffled-control gate also tolerated a single failed episode because 1/30 is within the five-percentage-point allowance; a pilot pass must not be described as zero failures.

Experiment 005 exposes a different issue: a pre-existing failure before intervention can block a strict acceptance procedure even when a held-out causal effect and every allowed quality increase pass. All thirty pairs remain counted; the early pair receives zero behavioral/time-fraction summaries under the project convention but still counts as failed, while other episode summaries use their available post-onset steps. This is not a general missing-data correction. The rule remains unchanged after confirmation. A future protocol could distinguish pre-intervention baseline competence from intervention-induced failure, but it would not retroactively pass this experiment. Random and shuffled directions also cause lateral effects; some increase failure, and no isolated internal concept is established. The first three fixed validation seeds supply paired movies; the first seed supplies the paper's ground-plane plot. These existing trajectories add explanatory visuals, not new evaluation data.

Before the scientific run, the core implementation passed 23 tests; both checkpoints were exercised on native Windows/Python 3.12. A 1,000-step HalfCheetah smoke rollout took about 1.10 seconds under the tested CPU setup, and an Ant video was checked. These are development checks, not steering-effect evidence or performance guarantees for another machine.

Terms in plain language

Term Meaning here
Policy / actor The neural network that maps the robot's observations to motor commands
Reinforcement learning Learning actions from interaction and rewards, rather than copying demonstration actions
SAC Soft Actor-Critic, a reinforcement-learning algorithm for continuous actions
TQC Truncated Quantile Critics, a reinforcement-learning variant using a distributional value estimate
Hidden activation An intermediate numerical representation inside the actor
Steering vector A fixed list of numbers added to a hidden activation to change the actor's behavior
Contrastive activation addition Subtracting mean hidden activations of contrasting groups and injecting the resulting direction
Strength / alpha The multiplier controlling how much of a direction is added
Frozen policy The actor's learned weights do not change during this experiment
Paired episode Baseline and intervention runs sharing a reset seed and the same state before steering starts
Validation Data used to choose a candidate and its strength
Confirmation / replication Fresh episode sets checking a choice after it has been locked
Confidence interval A bootstrap estimate of uncertainty in the mean paired behavioral change
Median / interquartile range The middle observed value / the interval covering the middle half of observed values
Cosine similarity How similarly two vectors point; 1 means aligned, 0 means perpendicular, and a negative value means opposing components dominate
Fitting speed floor A minimum forward-window speed used only to choose extraction data, not to discard evaluation failures
Episode-averaged contact fraction Average the available post-onset contact indicator within each episode, then weight episode summaries equally; a low average can hide a long failure in one trial
Episode physical-failure rate The fraction of episodes classified as failed before averaging, retaining individual collapses in the quality assessment
Prefix-only episode An episode with no post-onset steps; its behavioral/time-fraction summaries are assigned zero but it still counts as failed, the pair remains counted, and the original strict rule blocks the full quality gate
Forward-speed retention The ratio of steered to baseline episode-averaged forward speeds; it is not distance retention, completion rate, or evidence of full-horizon persistence
Lateral velocity Speed along the world's ground-plane y axis; it is distinct from torso height and does not establish sideways motion relative to the body's heading
Effort proxy Mean squared normalized actions, which is not a measurement of physical energy
Useful gate A predeclared threshold combining behavioral effect, uncertainty, and locomotion preservation
Causal effect A behavioral difference caused by the controlled activation intervention in the tested system

Entry requirements

For each experiment, record its question and hypothesis, reason for choosing the policy/descriptor/parameters, exact Git revision and artifact path, outcome including controls, interpretation, problems and fixes, supporting sources, and the next hypothesis. For a future training experiment also record architecture, objective, data generation, curriculum, hyperparameters, training progression, checkpoint selection, and evaluation splits. Distinguish an untested explanation from an observed cause.