Experiments and Published Results
This document describes the results published by
registry/registry.json. Aggregate numbers are recomputed from the referenced
artifacts by the full release gate.
Common protocol
- Splits are fight-disjoint.
- L1 consumes five causal RGB frames at offsets
[-8,-4,-2,-1,0]. - L2 uses the compact belief interface and anonymous option IDs.
- The L2 model is Gemma-4-E2B-it revision
70af34e..., temperature 0. - balanced-v3 selects source states across boss, fight, and distance strata before any L2 response is read.
- The prompt conditions are full mechanics, no mechanics, and swapped
mechanics. The word
fullin a run name refers to the full-description prompt, not the deprecated full belief view.
L1 official evaluation
| Game | Test rows | Distance MAE | dp@2 | Angle MAE | angle@20° | Distance-bin acc. | Angle-bin acc. |
|---|---|---|---|---|---|---|---|
| V Rising | 8,075 | 0.2620 | 0.9960 | 11.7599° | 0.8293 | 0.8320 | 0.7201 |
| Hollow Knight | 8,360 | 0.4680 | 0.9728 | 3.6527° | 0.9788 | 0.7261 | 0.9646 |
| Isaac | 1,784 | 9.0263 px | 0.1519 | 29.6388° | 0.6211 | 0.9036 | 0.8352 |
dp@2 uses native distance units and is therefore not comparable between
Isaac and the other games.
L1 matched-data visual comparison
| Game | Method | Train rows | Distance MAE | Angle MAE | Distance-bin acc. | Angle-bin acc. |
|---|---|---|---|---|---|---|
| V Rising | PACT-L1 | 18,068 | 0.2620 | 11.7599° | 0.8320 | 0.7201 |
| frozen DINOv2-S + probe | 18,068 | 0.6900 | 20.1532° | 0.6405 | 0.6041 | |
| frozen VideoMAE-B + probe | 18,068 | 0.6672 | 17.6564° | 0.6821 | 0.5948 | |
| Hollow Knight | PACT-L1 | 10,000 | 0.5671 | 4.0899° | 0.6940 | 0.9560 |
| frozen DINOv2-S + probe | 10,000 | 0.8641 | 20.0433° | 0.5542 | 0.8276 | |
| frozen VideoMAE-B + probe | 10,000 | 1.0017 | 27.5987° | 0.4544 | 0.7100 | |
| Isaac | PACT-L1 | 9,484 | 15.0472 | 39.6855° | 0.8358 | 0.7848 |
| frozen DINOv2-S + probe | 9,484 | 35.2852 | 77.7867° | 0.6015 | 0.6048 | |
| frozen VideoMAE-B + probe | 9,484 | 47.5451 | 84.9615° | 0.4608 | 0.5067 |
Every three-way comparison uses the same ordered training rows, official test rows, resolution, and five-frame history within a game. PACT-L1 is task-trained end to end; the two foundation baselines freeze their pretrained backbone and train task probes. These results compare complete training schemes rather than isolating backbone architecture.
V Rising end-to-end DINOv2-S L1 comparison experiment
This separate comparison makes the DINOv2-S backbone task-trainable and applies the same published V Rising rows, five-frame history, task heads/losses, seed, and two-stage 10+10 epoch schedule used by PACT-L1. Both methods are evaluated on the same ordered 8,075 official test rows.
| Method | Parameters | Distance MAE | Angle MAE | Distance-bin acc. | Angle-bin acc. | A800 batch-1 latency |
|---|---|---|---|---|---|---|
| PACT-L1 | 13,040,742 | 0.2620 | 11.7599° | 0.8320 | 0.7201 | 5.8497 ms |
| DINOv2-S end to end | 23,443,622 | 0.8566 | 24.8192° | 0.5771 | 0.5845 | 8.7039 ms |
Under this matched recipe, DINOv2-S has 3.2695 times the distance error, 2.1105 times the angle error, and 48.79% higher model-forward latency. The latency protocol uses AMP, 256 pre-decoded histories, 10 warm-up passes, and 20 measured repeats; it includes CPU-to-GPU transfer and model-specific preprocessing but excludes mp4 decoding. The measured 8.7039 ms is slower than PACT-L1 but does not establish failure of a 60 Hz real-time budget on this A800.
The result does not mean that a larger pretrained model must be more accurate after fine-tuning. In this implementation, the most likely explanation is a training-recipe and spatial-interface mismatch:
- The PACT learning-rate schedule uses one shared learning rate for the pretrained DINO backbone and the newly initialized task heads. The frozen DINO probe obtains 0.6900 distance MAE and 20.1532° angle MAE, while the first end-to-end stage at learning rate 3e-4 degrades to 0.9128 and 24.7542°; the second stage recovers distance only to 0.8566. This is evidence that the copied schedule over-updates the current DINO adaptation, although it is not a one-factor causal ablation.
- The adapter retains one global CLS vector per frame and discards DINO patch tokens. That removes localized spatial evidence that is useful for distance and angle estimation.
- The native 192x336 frame is stretched to DINO's 224x224 square input, changing the geometry seen by the model.
Accordingly, this experiment supports the conclusion that the evaluated
end-to-end DINOv2-S configuration is worse than PACT-L1 under the matched PACT
recipe. It is not a general upper bound on a DINOv2 system with separately
validated backbone/head learning rates, patch-token readout, and
aspect-ratio-preserving preprocessing. The machine-readable record is
results/vrising/eval_runs/L1-DINOV2-END-TO-END-MATCHED20/comparison.json.
L2 semantic intervention
Description-follow records whether the swapped-prompt decision follows the mechanics moved to another anonymous option rather than retaining the original option identity.
Natural-ready menus
| Game | Valid pairs | Mean K | Chance | No-description change | Swap change | Description-follow |
|---|---|---|---|---|---|---|
| V Rising | 200 | 4.875 | 21.93% | 46.0% | 79.5% | 67.0% |
| Hollow Knight | 200 | 5.370 | 18.78% | 90.0% | 95.0% | 79.5% |
| Isaac | 122 | 2.238 | 46.04% | 56.6% | 85.2% | 83.6% |
The natural-ready menu is a deterministic benchmark tracker rule based on
logged re-fire lower bounds and previous-action constraints. It is not an
engine legality oracle. Isaac has fewer skills and smaller natural menus than
the other games. Its 122 valid pairs are the predefined menu>=2 subset of the
200 source states, not a selection based on model success.
Fixed two-choice control
| Game | Pairs | Chance | No-description change | Swap change | Description-follow |
|---|---|---|---|---|---|
| Hollow Knight | 200 | 50.0% | 31.5% | 60.5% | 60.5% |
| Isaac | 200 | 50.0% | 39.0% | 54.0% | 54.0% |
The two-choice experiment is a fixed high-chance stress control. Hollow Knight is above its 50% reference; Isaac is close to it. The cross-game result is the eligible natural-ready intervention, while the two-choice table shows that the effect depends on candidate-space construction.
Image-modality diagnostic
Text and text-plus-current-image conditions use the same states, prompt text, and inference parameters.
| Game | Effective paired states | Text follow | Text + image follow | Original decision changed |
|---|---|---|---|---|
| V Rising | 60 | 60.0% | 65.0% | 21.7% |
| Hollow Knight | 60 | 78.3% | 80.0% | 3.3% |
| Isaac | 40 | 90.0% | 70.0% | 5.0% |
The observed direction is not uniform across games. This diagnostic measures modality sensitivity and does not establish that adding an image always improves semantic control.
Prompt-visible rationale correspondence
The deterministic audit checks whether the visible reason and mechanics cards support reconstruction of the model's selected option. It does not inspect a hidden chain of thought.
| Game | Valid pairs | Normal support | No-description support | Shuffled-reason support |
|---|---|---|---|---|
| V Rising | 200 | 93.5% | 3.5% | 3.7% |
| Hollow Knight | 200 | 94.5% | 0.0% | 2.6% |
| Isaac | 122 | 74.6% | 31.1% | 48.2% |
Game-specific completed analyses
- V Rising model ablation: description-follow is 67.0% for Gemma-4 E2B, 64.5% for Gemma-3n E2B, and 56.0% for Gemma-4 E4B. Model size is not monotonic with this metric.
- V Rising batch-8 latency on pre-decoded histories: PACT-L1 1.1440 ms, frozen DINOv2-S + probe 4.3040 ms, and frozen VideoMAE-B + probe 20.4469 ms per decision. This benchmark excludes video decoding.
- Hollow Knight current-frame ablation: on the fixed paired subset, distance MAE changes from 0.5160 with five frames to 0.5478 with the current frame only.
Decision-quality judge
The semantic intervention measures whether decisions respond to visible skill
mechanics. Tactical decision quality is a different question and is evaluated
by the external menu-level judge described in
pipelines/l2/llm_judge/README.md. External judge artifacts are not part of
the local deterministic registry until explicitly reviewed and synchronized.