pactbench / docs /EXPERIMENTS.md
BBoran's picture
Publish current portable PACTBench release
f1fc3a0 verified
|
Raw
History Blame Contribute Delete
8.55 kB

Experiments and Published Results

This document describes the results published by registry/registry.json. Aggregate numbers are recomputed from the referenced artifacts by the full release gate.

Common protocol

  • Splits are fight-disjoint.
  • L1 consumes five causal RGB frames at offsets [-8,-4,-2,-1,0].
  • L2 uses the compact belief interface and anonymous option IDs.
  • The L2 model is Gemma-4-E2B-it revision 70af34e..., temperature 0.
  • balanced-v3 selects source states across boss, fight, and distance strata before any L2 response is read.
  • The prompt conditions are full mechanics, no mechanics, and swapped mechanics. The word full in a run name refers to the full-description prompt, not the deprecated full belief view.

L1 official evaluation

Game Test rows Distance MAE dp@2 Angle MAE angle@20° Distance-bin acc. Angle-bin acc.
V Rising 8,075 0.2620 0.9960 11.7599° 0.8293 0.8320 0.7201
Hollow Knight 8,360 0.4680 0.9728 3.6527° 0.9788 0.7261 0.9646
Isaac 1,784 9.0263 px 0.1519 29.6388° 0.6211 0.9036 0.8352

dp@2 uses native distance units and is therefore not comparable between Isaac and the other games.

L1 matched-data visual comparison

Game Method Train rows Distance MAE Angle MAE Distance-bin acc. Angle-bin acc.
V Rising PACT-L1 18,068 0.2620 11.7599° 0.8320 0.7201
frozen DINOv2-S + probe 18,068 0.6900 20.1532° 0.6405 0.6041
frozen VideoMAE-B + probe 18,068 0.6672 17.6564° 0.6821 0.5948
Hollow Knight PACT-L1 10,000 0.5671 4.0899° 0.6940 0.9560
frozen DINOv2-S + probe 10,000 0.8641 20.0433° 0.5542 0.8276
frozen VideoMAE-B + probe 10,000 1.0017 27.5987° 0.4544 0.7100
Isaac PACT-L1 9,484 15.0472 39.6855° 0.8358 0.7848
frozen DINOv2-S + probe 9,484 35.2852 77.7867° 0.6015 0.6048
frozen VideoMAE-B + probe 9,484 47.5451 84.9615° 0.4608 0.5067

Every three-way comparison uses the same ordered training rows, official test rows, resolution, and five-frame history within a game. PACT-L1 is task-trained end to end; the two foundation baselines freeze their pretrained backbone and train task probes. These results compare complete training schemes rather than isolating backbone architecture.

V Rising end-to-end DINOv2-S L1 comparison experiment

This separate comparison makes the DINOv2-S backbone task-trainable and applies the same published V Rising rows, five-frame history, task heads/losses, seed, and two-stage 10+10 epoch schedule used by PACT-L1. Both methods are evaluated on the same ordered 8,075 official test rows.

Method Parameters Distance MAE Angle MAE Distance-bin acc. Angle-bin acc. A800 batch-1 latency
PACT-L1 13,040,742 0.2620 11.7599° 0.8320 0.7201 5.8497 ms
DINOv2-S end to end 23,443,622 0.8566 24.8192° 0.5771 0.5845 8.7039 ms

Under this matched recipe, DINOv2-S has 3.2695 times the distance error, 2.1105 times the angle error, and 48.79% higher model-forward latency. The latency protocol uses AMP, 256 pre-decoded histories, 10 warm-up passes, and 20 measured repeats; it includes CPU-to-GPU transfer and model-specific preprocessing but excludes mp4 decoding. The measured 8.7039 ms is slower than PACT-L1 but does not establish failure of a 60 Hz real-time budget on this A800.

The result does not mean that a larger pretrained model must be more accurate after fine-tuning. In this implementation, the most likely explanation is a training-recipe and spatial-interface mismatch:

  • The PACT learning-rate schedule uses one shared learning rate for the pretrained DINO backbone and the newly initialized task heads. The frozen DINO probe obtains 0.6900 distance MAE and 20.1532° angle MAE, while the first end-to-end stage at learning rate 3e-4 degrades to 0.9128 and 24.7542°; the second stage recovers distance only to 0.8566. This is evidence that the copied schedule over-updates the current DINO adaptation, although it is not a one-factor causal ablation.
  • The adapter retains one global CLS vector per frame and discards DINO patch tokens. That removes localized spatial evidence that is useful for distance and angle estimation.
  • The native 192x336 frame is stretched to DINO's 224x224 square input, changing the geometry seen by the model.

Accordingly, this experiment supports the conclusion that the evaluated end-to-end DINOv2-S configuration is worse than PACT-L1 under the matched PACT recipe. It is not a general upper bound on a DINOv2 system with separately validated backbone/head learning rates, patch-token readout, and aspect-ratio-preserving preprocessing. The machine-readable record is results/vrising/eval_runs/L1-DINOV2-END-TO-END-MATCHED20/comparison.json.

L2 semantic intervention

Description-follow records whether the swapped-prompt decision follows the mechanics moved to another anonymous option rather than retaining the original option identity.

Natural-ready menus

Game Valid pairs Mean K Chance No-description change Swap change Description-follow
V Rising 200 4.875 21.93% 46.0% 79.5% 67.0%
Hollow Knight 200 5.370 18.78% 90.0% 95.0% 79.5%
Isaac 122 2.238 46.04% 56.6% 85.2% 83.6%

The natural-ready menu is a deterministic benchmark tracker rule based on logged re-fire lower bounds and previous-action constraints. It is not an engine legality oracle. Isaac has fewer skills and smaller natural menus than the other games. Its 122 valid pairs are the predefined menu>=2 subset of the 200 source states, not a selection based on model success.

Fixed two-choice control

Game Pairs Chance No-description change Swap change Description-follow
Hollow Knight 200 50.0% 31.5% 60.5% 60.5%
Isaac 200 50.0% 39.0% 54.0% 54.0%

The two-choice experiment is a fixed high-chance stress control. Hollow Knight is above its 50% reference; Isaac is close to it. The cross-game result is the eligible natural-ready intervention, while the two-choice table shows that the effect depends on candidate-space construction.

Image-modality diagnostic

Text and text-plus-current-image conditions use the same states, prompt text, and inference parameters.

Game Effective paired states Text follow Text + image follow Original decision changed
V Rising 60 60.0% 65.0% 21.7%
Hollow Knight 60 78.3% 80.0% 3.3%
Isaac 40 90.0% 70.0% 5.0%

The observed direction is not uniform across games. This diagnostic measures modality sensitivity and does not establish that adding an image always improves semantic control.

Prompt-visible rationale correspondence

The deterministic audit checks whether the visible reason and mechanics cards support reconstruction of the model's selected option. It does not inspect a hidden chain of thought.

Game Valid pairs Normal support No-description support Shuffled-reason support
V Rising 200 93.5% 3.5% 3.7%
Hollow Knight 200 94.5% 0.0% 2.6%
Isaac 122 74.6% 31.1% 48.2%

Game-specific completed analyses

  • V Rising model ablation: description-follow is 67.0% for Gemma-4 E2B, 64.5% for Gemma-3n E2B, and 56.0% for Gemma-4 E4B. Model size is not monotonic with this metric.
  • V Rising batch-8 latency on pre-decoded histories: PACT-L1 1.1440 ms, frozen DINOv2-S + probe 4.3040 ms, and frozen VideoMAE-B + probe 20.4469 ms per decision. This benchmark excludes video decoding.
  • Hollow Knight current-frame ablation: on the fixed paired subset, distance MAE changes from 0.5160 with five frames to 0.5478 with the current frame only.

Decision-quality judge

The semantic intervention measures whether decisions respond to visible skill mechanics. Tactical decision quality is a different question and is evaluated by the external menu-level judge described in pipelines/l2/llm_judge/README.md. External judge artifacts are not part of the local deterministic registry until explicitly reviewed and synchronized.