sra-trajectory-code / GAMEFORMER_SRA.md
po03087's picture
Add GAMEFORMER_SRA.md (negative result) + standalone GameFormer/C2F/SRA-GameFormer scripts
4f3fc40 verified
|
Raw
History Blame Contribute Delete
6.66 kB
# GameFormer + SRA — a negative result
**Question.** SRA improves three *iterative denoisers* (MID, LED, MoFlow). Do its ideas also help
a **feedforward, non-diffusion** predictor? If so, SRA generalizes beyond denoising.
**Answer. No.** Injecting SRA's three components into GameFormer is **neutral-to-slightly-negative**.
| Model (NBA, min-ADE₂₀ / min-FDE₂₀ @ 4.0 s) | ADE₄ | FDE₄ | params |
|---|---|---|---|
| Vanilla GameFormer (control) | **0.8496** | **0.9557** | 11.85 M |
| GameFormer **+ SRA** (`GF_SRA=1`) | 0.8554 | 0.9564 | 11.90 M |
| **Δ (SRA − vanilla)** | **+0.0058 (worse)** | +0.0007 (≈neutral) | +0.05 M |
Both peaked at epoch 48 of 50. Same file, same seed (3407), same schedule — **only the three flags differ**.
---
## 1. Experimental design
The comparison is a **within-file ablation**, not a cross-codebase one:
* One script, `gameformer_sra_nba_standalone.py`, contains both arms.
* Three env-gated components; **with all flags off the code path is exactly vanilla GameFormer.**
* Verified: flags-off reports **11.85 M params** — identical to the separate vanilla
implementation, which scores 0.848/0.951. The refactor perturbs nothing.
* SRA adds only **+0.05 M params (+0.4 %)**, so any effect is attributable to the *mechanism*,
not to capacity.
Config: level-k = 3, K = 20 modes, 6 encoder layers, AdamW lr 1e-4, wd 1e-4, grad-clip 5,
MultiStepLR ×0.5, batch 128, 50 epochs, seed 3407. Standalone predictor — history only,
**no denoiser and no external future signal**.
---
## 2. What was injected
All three live inside GameFormer's `InteractionDecoder` (one per level-k stage).
### 2.1 Uncertainty-gated message passing — `GF_SRA_SIGMA`
GameFormer's GMM head already *emits* a per-step log-σ but never **uses** it for interaction.
SRA turns it into a message gate: uncertain agents send weaker interaction messages.
```python
sig = torch.exp(last_traj[..., 2:].clamp(0, 5)).mean(-1) # [B,K,A,T] from prev level's GMM
sig_agent = (sig * w[..., None]).sum(2).mean(-1) # [B,A] score-weighted, mean over T
...
futures = futures * self.sigma_gate(sig_agent[..., None]) # MLP(1→32→1)+Sigmoid ⇒ gate ∈ (0,1)
```
### 2.2 Sparse RAG neighbor selection — `GF_SRA_TOPN`
GameFormer attends **densely** over all agents. SRA restricts each agent to its top-N (=5)
neighbours, chosen by *future* proximity and biased toward confident agents.
```python
prox = (rep[:, :, None] - rep[:, None]).norm(dim=-1).mean(-1) # [B,A,A] mean future distance
if sig_agent is not None:
prox = prox + SIG_SEL_W * sig_agent[:, None, :] # prefer confident neighbours
prox = prox.masked_fill(eye[None], float('inf')) # exclude self
nbr_idx = prox.topk(N, largest=False, dim=-1).indices
```
The resulting mask is applied **both** to the interaction self-attention (`attn_mask`) and to the
cross-attention key mask, so sparsity is enforced everywhere, not just in one place.
### 2.3 Explicit relational future features — `GF_SRA_RELFUT`
GameFormer's `FutureEncoder` encodes each agent's future **individually**. SRA adds explicit
**pairwise-relative** future geometry (Δposition, Δvelocity), max-pooled over time.
```python
dpos = rep[:, :, None] - rep[:, None] # [B,A,A,T,2]
dvel = F.pad(torch.diff(dpos, dim=-2), (0, 0, 1, 0))
relf = self.rel_mlp(torch.cat([dpos, dvel], -1)).max(dim=-2).values # MLP(4→64→D)
futures = futures + relf.sum(2) / denom # neighbour-mean
```
---
## 3. Reproducing
```bash
# control (vanilla GameFormer)
python gameformer_sra_nba_standalone.py --gpu 0 --exp van --level 3 --epochs 50 --batch_size 128
# + SRA (all three components)
GF_SRA=1 python gameformer_sra_nba_standalone.py --gpu 0 --exp sra --level 3 --epochs 50 --batch_size 128
```
| Variable | Effect |
|---|---|
| `GF_SRA=1` | enable all three components |
| `GF_SRA_SIGMA` / `GF_SRA_TOPN` / `GF_SRA_RELFUT` | enable individually |
| `GF_TOPN` | neighbour budget (default 5) |
| `GF_SIG_SEL_W` | weight of σ in neighbour selection (default 0.5) |
The script prints its configuration so every run is self-documenting:
```
[SRA-GF] SRA components: neighbor-topN=True(N=5) sigma-gate=True relational-future=True
```
---
## 4. Why it does not transfer
1. **The future-interaction idea is already native to GameFormer.** Its level-k reasoning
*is* "condition each agent on the others' predicted futures". SRA's relational features
are therefore largely **redundant** here, adding parameters and noise rather than new signal.
In MID/LED/MoFlow nothing plays this role, which is exactly why SRA helps there.
2. **The uncertainty signal is weaker.** In a denoiser, σ is *principled* — a noise schedule or
the leapfrog initializer's variance estimate, tied to how much the model should trust its
current iterate. GameFormer's log-σ is just a learned regression head; gating on it does not
carry comparable information.
3. **Sparsity costs more than it saves at A = 11.** Dense attention over 11 agents is cheap and
fully informative; discarding all but 5 neighbours removes usable context. SRA's sparse
selection pays off on 23-agent sport scenes and on hosts where dense interaction is unstable.
4. **One-shot vs iterative (the core reason).** SRA is built to *refine an evolving estimate*: it
is applied at every denoising step, so its gated residual is corrected repeatedly and errors
are damped. A feedforward decoder gets **one** pass — no iterative correction, no damping.
The regime SRA was designed for is absent.
---
## 5. How to read this result
This is a **scope finding, not a failure**, and it is consistent with the paper's framing:
SRA is a *denoiser* adapter. Its gains come from repeatedly reshaping an iteratively refined
future estimate — a mechanism a single-pass predictor does not have.
Stated honestly:
> SRA's gains are specific to the iterative-denoising setting. Applied to a strong feedforward
> predictor (GameFormer), it is neutral-to-slightly-negative, which supports rather than
> contradicts the paper's positioning.
It mirrors the MoFlow observation from the other direction: the *host's* update mechanism
determines what an interaction module can contribute.
**Caveats.** Single seed; the ADE gap (0.006, 0.7 %) is small enough that multi-seed runs could
place it within noise — i.e. "neutral" rather than "worse". It is clearly **not** an improvement
either way. No per-component ablation was run, so it remains open whether one component alone
transfers even though the combination does not.