pactbench / docs /EXPERIMENTS.md
BBoran's picture
Publish current portable PACTBench release
f1fc3a0 verified
|
Raw
History Blame Contribute Delete
8.55 kB
# Experiments and Published Results
This document describes the results published by
`registry/registry.json`. Aggregate numbers are recomputed from the referenced
artifacts by the full release gate.
## Common protocol
- Splits are fight-disjoint.
- L1 consumes five causal RGB frames at offsets `[-8,-4,-2,-1,0]`.
- L2 uses the compact belief interface and anonymous option IDs.
- The L2 model is Gemma-4-E2B-it revision `70af34e...`, temperature 0.
- balanced-v3 selects source states across boss, fight, and distance strata
before any L2 response is read.
- The prompt conditions are full mechanics, no mechanics, and swapped
mechanics. The word `full` in a run name refers to the full-description
prompt, not the deprecated full belief view.
## L1 official evaluation
| Game | Test rows | Distance MAE | dp@2 | Angle MAE | angle@20° | Distance-bin acc. | Angle-bin acc. |
|---|---:|---:|---:|---:|---:|---:|---:|
| V Rising | 8,075 | 0.2620 | 0.9960 | 11.7599° | 0.8293 | 0.8320 | 0.7201 |
| Hollow Knight | 8,360 | 0.4680 | 0.9728 | 3.6527° | 0.9788 | 0.7261 | 0.9646 |
| Isaac | 1,784 | 9.0263 px | 0.1519 | 29.6388° | 0.6211 | 0.9036 | 0.8352 |
`dp@2` uses native distance units and is therefore not comparable between
Isaac and the other games.
## L1 matched-data visual comparison
| Game | Method | Train rows | Distance MAE | Angle MAE | Distance-bin acc. | Angle-bin acc. |
|---|---|---:|---:|---:|---:|---:|
| V Rising | PACT-L1 | 18,068 | **0.2620** | **11.7599°** | **0.8320** | **0.7201** |
| | frozen DINOv2-S + probe | 18,068 | 0.6900 | 20.1532° | 0.6405 | 0.6041 |
| | frozen VideoMAE-B + probe | 18,068 | 0.6672 | 17.6564° | 0.6821 | 0.5948 |
| Hollow Knight | PACT-L1 | 10,000 | **0.5671** | **4.0899°** | **0.6940** | **0.9560** |
| | frozen DINOv2-S + probe | 10,000 | 0.8641 | 20.0433° | 0.5542 | 0.8276 |
| | frozen VideoMAE-B + probe | 10,000 | 1.0017 | 27.5987° | 0.4544 | 0.7100 |
| Isaac | PACT-L1 | 9,484 | **15.0472** | **39.6855°** | **0.8358** | **0.7848** |
| | frozen DINOv2-S + probe | 9,484 | 35.2852 | 77.7867° | 0.6015 | 0.6048 |
| | frozen VideoMAE-B + probe | 9,484 | 47.5451 | 84.9615° | 0.4608 | 0.5067 |
Every three-way comparison uses the same ordered training rows, official test
rows, resolution, and five-frame history within a game. PACT-L1 is task-trained
end to end; the two foundation baselines freeze their pretrained backbone and
train task probes. These results compare complete training schemes rather than
isolating backbone architecture.
### V Rising end-to-end DINOv2-S L1 comparison experiment
This separate comparison makes the DINOv2-S backbone task-trainable and applies
the same published V Rising rows, five-frame history, task heads/losses, seed,
and two-stage 10+10 epoch schedule used by PACT-L1. Both methods are evaluated
on the same ordered 8,075 official test rows.
| Method | Parameters | Distance MAE | Angle MAE | Distance-bin acc. | Angle-bin acc. | A800 batch-1 latency |
|---|---:|---:|---:|---:|---:|---:|
| PACT-L1 | 13,040,742 | **0.2620** | **11.7599°** | **0.8320** | **0.7201** | **5.8497 ms** |
| DINOv2-S end to end | 23,443,622 | 0.8566 | 24.8192° | 0.5771 | 0.5845 | 8.7039 ms |
Under this matched recipe, DINOv2-S has 3.2695 times the distance error, 2.1105
times the angle error, and 48.79% higher model-forward latency. The latency
protocol uses AMP, 256 pre-decoded histories, 10 warm-up passes, and 20 measured
repeats; it includes CPU-to-GPU transfer and model-specific preprocessing but
excludes mp4 decoding. The measured 8.7039 ms is slower than PACT-L1 but does
not establish failure of a 60 Hz real-time budget on this A800.
The result does not mean that a larger pretrained model must be more accurate
after fine-tuning. In this implementation, the most likely explanation is a
training-recipe and spatial-interface mismatch:
- The PACT learning-rate schedule uses one shared learning rate for the
pretrained DINO backbone and the newly initialized task heads. The frozen
DINO probe obtains 0.6900 distance MAE and 20.1532° angle MAE, while the first
end-to-end stage at learning rate 3e-4 degrades to 0.9128 and 24.7542°; the
second stage recovers distance only to 0.8566. This is evidence that the
copied schedule over-updates the current DINO adaptation, although it is not
a one-factor causal ablation.
- The adapter retains one global CLS vector per frame and discards DINO patch
tokens. That removes localized spatial evidence that is useful for distance
and angle estimation.
- The native 192x336 frame is stretched to DINO's 224x224 square input, changing
the geometry seen by the model.
Accordingly, this experiment supports the conclusion that the evaluated
end-to-end DINOv2-S configuration is worse than PACT-L1 under the matched PACT
recipe. It is not a general upper bound on a DINOv2 system with separately
validated backbone/head learning rates, patch-token readout, and
aspect-ratio-preserving preprocessing. The machine-readable record is
`results/vrising/eval_runs/L1-DINOV2-END-TO-END-MATCHED20/comparison.json`.
## L2 semantic intervention
Description-follow records whether the swapped-prompt decision follows the
mechanics moved to another anonymous option rather than retaining the original
option identity.
### Natural-ready menus
| Game | Valid pairs | Mean K | Chance | No-description change | Swap change | Description-follow |
|---|---:|---:|---:|---:|---:|---:|
| V Rising | 200 | 4.875 | 21.93% | 46.0% | 79.5% | **67.0%** |
| Hollow Knight | 200 | 5.370 | 18.78% | 90.0% | 95.0% | **79.5%** |
| Isaac | 122 | 2.238 | 46.04% | 56.6% | 85.2% | **83.6%** |
The natural-ready menu is a deterministic benchmark tracker rule based on
logged re-fire lower bounds and previous-action constraints. It is not an
engine legality oracle. Isaac has fewer skills and smaller natural menus than
the other games. Its 122 valid pairs are the predefined `menu>=2` subset of the
200 source states, not a selection based on model success.
### Fixed two-choice control
| Game | Pairs | Chance | No-description change | Swap change | Description-follow |
|---|---:|---:|---:|---:|---:|
| Hollow Knight | 200 | 50.0% | 31.5% | 60.5% | 60.5% |
| Isaac | 200 | 50.0% | 39.0% | 54.0% | 54.0% |
The two-choice experiment is a fixed high-chance stress control. Hollow Knight
is above its 50% reference; Isaac is close to it. The cross-game result is the
eligible natural-ready intervention, while the two-choice table shows that the
effect depends on candidate-space construction.
## Image-modality diagnostic
Text and text-plus-current-image conditions use the same states, prompt text,
and inference parameters.
| Game | Effective paired states | Text follow | Text + image follow | Original decision changed |
|---|---:|---:|---:|---:|
| V Rising | 60 | 60.0% | 65.0% | 21.7% |
| Hollow Knight | 60 | 78.3% | 80.0% | 3.3% |
| Isaac | 40 | 90.0% | 70.0% | 5.0% |
The observed direction is not uniform across games. This diagnostic measures
modality sensitivity and does not establish that adding an image always
improves semantic control.
## Prompt-visible rationale correspondence
The deterministic audit checks whether the visible reason and mechanics cards
support reconstruction of the model's selected option. It does not inspect a
hidden chain of thought.
| Game | Valid pairs | Normal support | No-description support | Shuffled-reason support |
|---|---:|---:|---:|---:|
| V Rising | 200 | 93.5% | 3.5% | 3.7% |
| Hollow Knight | 200 | 94.5% | 0.0% | 2.6% |
| Isaac | 122 | 74.6% | 31.1% | 48.2% |
## Game-specific completed analyses
- V Rising model ablation: description-follow is 67.0% for Gemma-4 E2B,
64.5% for Gemma-3n E2B, and 56.0% for Gemma-4 E4B. Model size is not
monotonic with this metric.
- V Rising batch-8 latency on pre-decoded histories: PACT-L1 1.1440 ms,
frozen DINOv2-S + probe 4.3040 ms, and frozen VideoMAE-B + probe
20.4469 ms per decision. This benchmark excludes video decoding.
- Hollow Knight current-frame ablation: on the fixed paired subset, distance
MAE changes from 0.5160 with five frames to 0.5478 with the current frame
only.
## Decision-quality judge
The semantic intervention measures whether decisions respond to visible skill
mechanics. Tactical decision quality is a different question and is evaluated
by the external menu-level judge described in
`pipelines/l2/llm_judge/README.md`. External judge artifacts are not part of
the local deterministic registry until explicitly reviewed and synchronized.