# Experiments and Published Results This document describes the results published by `registry/registry.json`. Aggregate numbers are recomputed from the referenced artifacts by the full release gate. ## Common protocol - Splits are fight-disjoint. - L1 consumes five causal RGB frames at offsets `[-8,-4,-2,-1,0]`. - L2 uses the compact belief interface and anonymous option IDs. - The L2 model is Gemma-4-E2B-it revision `70af34e...`, temperature 0. - balanced-v3 selects source states across boss, fight, and distance strata before any L2 response is read. - The prompt conditions are full mechanics, no mechanics, and swapped mechanics. The word `full` in a run name refers to the full-description prompt, not the deprecated full belief view. ## L1 official evaluation | Game | Test rows | Distance MAE | dp@2 | Angle MAE | angle@20° | Distance-bin acc. | Angle-bin acc. | |---|---:|---:|---:|---:|---:|---:|---:| | V Rising | 8,075 | 0.2620 | 0.9960 | 11.7599° | 0.8293 | 0.8320 | 0.7201 | | Hollow Knight | 8,360 | 0.4680 | 0.9728 | 3.6527° | 0.9788 | 0.7261 | 0.9646 | | Isaac | 1,784 | 9.0263 px | 0.1519 | 29.6388° | 0.6211 | 0.9036 | 0.8352 | `dp@2` uses native distance units and is therefore not comparable between Isaac and the other games. ## L1 matched-data visual comparison | Game | Method | Train rows | Distance MAE | Angle MAE | Distance-bin acc. | Angle-bin acc. | |---|---|---:|---:|---:|---:|---:| | V Rising | PACT-L1 | 18,068 | **0.2620** | **11.7599°** | **0.8320** | **0.7201** | | | frozen DINOv2-S + probe | 18,068 | 0.6900 | 20.1532° | 0.6405 | 0.6041 | | | frozen VideoMAE-B + probe | 18,068 | 0.6672 | 17.6564° | 0.6821 | 0.5948 | | Hollow Knight | PACT-L1 | 10,000 | **0.5671** | **4.0899°** | **0.6940** | **0.9560** | | | frozen DINOv2-S + probe | 10,000 | 0.8641 | 20.0433° | 0.5542 | 0.8276 | | | frozen VideoMAE-B + probe | 10,000 | 1.0017 | 27.5987° | 0.4544 | 0.7100 | | Isaac | PACT-L1 | 9,484 | **15.0472** | **39.6855°** | **0.8358** | **0.7848** | | | frozen DINOv2-S + probe | 9,484 | 35.2852 | 77.7867° | 0.6015 | 0.6048 | | | frozen VideoMAE-B + probe | 9,484 | 47.5451 | 84.9615° | 0.4608 | 0.5067 | Every three-way comparison uses the same ordered training rows, official test rows, resolution, and five-frame history within a game. PACT-L1 is task-trained end to end; the two foundation baselines freeze their pretrained backbone and train task probes. These results compare complete training schemes rather than isolating backbone architecture. ### V Rising end-to-end DINOv2-S L1 comparison experiment This separate comparison makes the DINOv2-S backbone task-trainable and applies the same published V Rising rows, five-frame history, task heads/losses, seed, and two-stage 10+10 epoch schedule used by PACT-L1. Both methods are evaluated on the same ordered 8,075 official test rows. | Method | Parameters | Distance MAE | Angle MAE | Distance-bin acc. | Angle-bin acc. | A800 batch-1 latency | |---|---:|---:|---:|---:|---:|---:| | PACT-L1 | 13,040,742 | **0.2620** | **11.7599°** | **0.8320** | **0.7201** | **5.8497 ms** | | DINOv2-S end to end | 23,443,622 | 0.8566 | 24.8192° | 0.5771 | 0.5845 | 8.7039 ms | Under this matched recipe, DINOv2-S has 3.2695 times the distance error, 2.1105 times the angle error, and 48.79% higher model-forward latency. The latency protocol uses AMP, 256 pre-decoded histories, 10 warm-up passes, and 20 measured repeats; it includes CPU-to-GPU transfer and model-specific preprocessing but excludes mp4 decoding. The measured 8.7039 ms is slower than PACT-L1 but does not establish failure of a 60 Hz real-time budget on this A800. The result does not mean that a larger pretrained model must be more accurate after fine-tuning. In this implementation, the most likely explanation is a training-recipe and spatial-interface mismatch: - The PACT learning-rate schedule uses one shared learning rate for the pretrained DINO backbone and the newly initialized task heads. The frozen DINO probe obtains 0.6900 distance MAE and 20.1532° angle MAE, while the first end-to-end stage at learning rate 3e-4 degrades to 0.9128 and 24.7542°; the second stage recovers distance only to 0.8566. This is evidence that the copied schedule over-updates the current DINO adaptation, although it is not a one-factor causal ablation. - The adapter retains one global CLS vector per frame and discards DINO patch tokens. That removes localized spatial evidence that is useful for distance and angle estimation. - The native 192x336 frame is stretched to DINO's 224x224 square input, changing the geometry seen by the model. Accordingly, this experiment supports the conclusion that the evaluated end-to-end DINOv2-S configuration is worse than PACT-L1 under the matched PACT recipe. It is not a general upper bound on a DINOv2 system with separately validated backbone/head learning rates, patch-token readout, and aspect-ratio-preserving preprocessing. The machine-readable record is `results/vrising/eval_runs/L1-DINOV2-END-TO-END-MATCHED20/comparison.json`. ## L2 semantic intervention Description-follow records whether the swapped-prompt decision follows the mechanics moved to another anonymous option rather than retaining the original option identity. ### Natural-ready menus | Game | Valid pairs | Mean K | Chance | No-description change | Swap change | Description-follow | |---|---:|---:|---:|---:|---:|---:| | V Rising | 200 | 4.875 | 21.93% | 46.0% | 79.5% | **67.0%** | | Hollow Knight | 200 | 5.370 | 18.78% | 90.0% | 95.0% | **79.5%** | | Isaac | 122 | 2.238 | 46.04% | 56.6% | 85.2% | **83.6%** | The natural-ready menu is a deterministic benchmark tracker rule based on logged re-fire lower bounds and previous-action constraints. It is not an engine legality oracle. Isaac has fewer skills and smaller natural menus than the other games. Its 122 valid pairs are the predefined `menu>=2` subset of the 200 source states, not a selection based on model success. ### Fixed two-choice control | Game | Pairs | Chance | No-description change | Swap change | Description-follow | |---|---:|---:|---:|---:|---:| | Hollow Knight | 200 | 50.0% | 31.5% | 60.5% | 60.5% | | Isaac | 200 | 50.0% | 39.0% | 54.0% | 54.0% | The two-choice experiment is a fixed high-chance stress control. Hollow Knight is above its 50% reference; Isaac is close to it. The cross-game result is the eligible natural-ready intervention, while the two-choice table shows that the effect depends on candidate-space construction. ## Image-modality diagnostic Text and text-plus-current-image conditions use the same states, prompt text, and inference parameters. | Game | Effective paired states | Text follow | Text + image follow | Original decision changed | |---|---:|---:|---:|---:| | V Rising | 60 | 60.0% | 65.0% | 21.7% | | Hollow Knight | 60 | 78.3% | 80.0% | 3.3% | | Isaac | 40 | 90.0% | 70.0% | 5.0% | The observed direction is not uniform across games. This diagnostic measures modality sensitivity and does not establish that adding an image always improves semantic control. ## Prompt-visible rationale correspondence The deterministic audit checks whether the visible reason and mechanics cards support reconstruction of the model's selected option. It does not inspect a hidden chain of thought. | Game | Valid pairs | Normal support | No-description support | Shuffled-reason support | |---|---:|---:|---:|---:| | V Rising | 200 | 93.5% | 3.5% | 3.7% | | Hollow Knight | 200 | 94.5% | 0.0% | 2.6% | | Isaac | 122 | 74.6% | 31.1% | 48.2% | ## Game-specific completed analyses - V Rising model ablation: description-follow is 67.0% for Gemma-4 E2B, 64.5% for Gemma-3n E2B, and 56.0% for Gemma-4 E4B. Model size is not monotonic with this metric. - V Rising batch-8 latency on pre-decoded histories: PACT-L1 1.1440 ms, frozen DINOv2-S + probe 4.3040 ms, and frozen VideoMAE-B + probe 20.4469 ms per decision. This benchmark excludes video decoding. - Hollow Knight current-frame ablation: on the fixed paired subset, distance MAE changes from 0.5160 with five frames to 0.5478 with the current frame only. ## Decision-quality judge The semantic intervention measures whether decisions respond to visible skill mechanics. Tactical decision quality is a different question and is evaluated by the external menu-level judge described in `pipelines/l2/llm_judge/README.md`. External judge artifacts are not part of the local deterministic registry until explicitly reviewed and synchronized.