occ-stack / reports /occ_paper.md
narcolepticchicken's picture
Add workshop paper: Compute Is Not Neutral
5a85f0b verified
|
Raw
History Blame Contribute Delete
12.2 kB

Compute Is Not Neutral: Mechanism Analysis of Adversarial Debate Collapse and the OCC Stack for Governed Agent Resource Allocation

Abstract

Multi-agent debate systems are increasingly used for LLM reasoning, but they treat compute as a free good — every agent gets equal turns regardless of contribution. We show that this creates an exploitable attack surface: in 4-agent debate (3 honest + 1 adversarial) on Qwen3-Coder-30B, majority-vote accuracy collapses from 73.3% (1 round) to 56.7% (3 rounds) — worse than single-round voting — because the adversarial agent's extra turns amplify into extra votes. Through mechanism isolation (7 conditions, 30 topics), we identify the root cause as volume amplification, not persuasion: honest agents retain their positions at 84% but get outvoted. Simple protocol fixes (LLM judge voting, confidence-weighted voting) fully recover accuracy to 73.3%. We then present OCC (Oracle-Credit-Compute), a mechanism-design layer where agents earn non-transferable, decaying, capability-scoped credits based on verified marginal impact. We evaluate OCC on three benchmarks (code generation, retrieval QA, multi-agent debate), test 10 anti-gaming attacks (all contained), and provide a GRPO-compatible reward hook for learned allocation. We honestly report where OCC succeeds (preventing catastrophic collapse, anti-gaming) and where it does not (matching random gating at moderate budgets, no policy improvement at 0.5B scale).


1. Introduction

Modern LLM agent systems allocate compute indiscriminately: every agent gets equal debate turns, unlimited retries, and unconstrained tool access. This implicitly assumes compute is neutral — that more turns, tokens, and tool calls can only help.

This assumption is wrong. We show that in multi-agent debate with one adversarial agent, giving all agents equal speaking turns across 3 rounds causes accuracy to collapse by 16.7 percentage points relative to a single-round baseline. The mechanism is not persuasion — honest agents keep their original answers 84% of the time — but volume amplification: the adversary's extra turns become extra votes in the majority pool, tipping close-call topics.

This finding motivates OCC (Oracle-Credit-Compute), a mechanism-design layer that treats agent compute as a scarce, earned, auditable privilege rather than a right. OCC has four components: (1) an Impact Oracle that scores whether actions produce verified marginal value, (2) a Credit Ledger with non-transferable, decaying, capability-scoped credits, (3) a Resource Broker that grants capability-based access, and (4) a GRPO-compatible reward hook for learned allocation.

Contributions:

  1. Empirical finding: Multi-round debate collapses under adversarial pressure (§3). Mechanism isolation shows volume amplification, not persuasion, is the cause (§4).
  2. Protocol fixes: Judge voting and confidence weighting fully recover accuracy (§4.3).
  3. OCC system: Open-source stack with formal definition, anti-gaming threat model, and GRPO hook (§5).
  4. Three benchmarks: Code generation (52.3% compute savings, simulated), retrieval QA, and multi-agent debate (§6).
  5. Anti-gaming evaluation: 10 attack vectors tested, all contained by ledger design (§7).
  6. Honest limitations: OCC ≈ random gating at moderate budgets; GRPO training produced no improvement at 0.5B scale (§8).

2. Related Work

Test-time compute allocation. Snell et al. (2024) showed compute-optimal scaling outperforms 14× larger models. Brown et al. (2024) showed coverage scales log-linearly with samples. OCC extends this from per-prompt to per-agent allocation.

Multi-agent debate. Du et al. (2023) introduced multi-agent debate for reasoning. Subramanian et al. (2025) found majority voting drives most gains. We show that majority voting itself is the vulnerability under adversarial conditions.

GRPO/RLVR. DeepSeek-R1 (2025) uses GRPO with rule-based rewards. OCC's reward hook is GRPO-compatible but adds compute-cost penalties.

Calibration and abstention. RLCR (2025) uses Brier score for calibration rewards. OCC integrates calibration as a first-class reward component.

Capability-based security. AgentGuardian (2026) and SAGA (2025) propose access control for agents. OCC dynamically earns capabilities through verified impact rather than static assignment.

Full literature review in reports/literature_review.md.


3. The Collapse

3.1 Setup

  • Model: Qwen3-Coder-30B-A3B-Instruct (MoE, 30B total, ~3B active)
  • Topics: 30 yes/no scientific fact questions (CS, physics, biology, math)
  • Agents: 4 total — 3 honest, 1 adversarial (argues wrong answer)
  • Protocol: Each agent generates a YES/NO answer with explanation. Majority vote determines group answer.
  • Rounds: 1 round (baseline) vs 3 rounds (collapse condition)
  • Seed: 42 (single seed, n=30 topics, 95% CI ±16pp)

3.2 Result

Condition Accuracy Δ from Baseline
1-round majority vote 73.3% (22/30)
3-round majority vote 56.7% (17/30) −16.7pp

Three rounds of debate — 3× the compute — produces performance 16.7 percentage points worse than a single round. The adversarial agent, given equal speaking turns, converts extra rounds into influence over the vote pool.

3.3 Individual Agent Accuracy

Honest agents individually answer correctly 75.6% of the time. The adversary individually answers correctly 73.3% of the time. The collapse is not about individual capability; it's about structural vote amplification.


4. Mechanism Isolation

We test 7 conditions to isolate why the collapse occurs. Pre-registered hypotheses in jobs/analyze_collapse.py.

4.1 Conditions

# Condition Manipulation
1 baseline_1round Single round, majority vote
2 equal_3round 3 rounds, all agents equal turns
3 equal_token Adversary gets 1 long turn (512 tok), honest agents get 3 short turns (171 each) — same total tokens
4 randomized_order Speaking order shuffled each round
5 judge_vote LLM judge reads transcript, decides answer
6 confidence_weighted Votes weighted by self-reported confidence
7 adversary_weak Adversary told to "struggle to make convincing arguments"

4.2 Results

Condition Accuracy Δ from Collapse Hypothesis
baseline_1round 0.733 +16.7pp
equal_3round (collapse) 0.567
equal_token 0.700 +13.3pp H1: Volume ✅
randomized_order 0.500 −6.7pp H2: Turn order ⚠️
judge_vote 0.733 +16.7pp H3: Voting ✅
confidence_weighted 0.733 +16.7pp H5: Confidence ✅
adversary_weak 0.600 +3.3pp H6: Skill ❌

4.3 Hypothesis Verdicts

Hypothesis Verdict Evidence
H1: Volume amplification ✅ Confirmed Equal-token recovers to 0.700. Adversary's 3× vote count drives collapse.
H2: Turn-order effect ⚠️ Partial Randomization makes things worse (0.500). Disrupts honest coordination.
H3: Voting vulnerability ✅ Confirmed Judge and confidence voting both fully recover to 0.733.
H4: Contamination ❌ Refuted 84.0% retention (63/75). Agents don't copy adversary.
H5: Confidence distortion ✅ Confirmed Confidence weighting recovers to 0.733.
H6: Skill dependency ❌ Refuted Even weak adversary collapses (0.600). Structural.
H7: Topic vulnerability ✅ Confirmed Variance=0.921. Collapse clusters on ambiguous topics.

4.4 Flip Analysis (R1→R3)

Transition Count Rate
Stable 68 75.6%
Degraded (correct→wrong) 13 14.4%
Improved (wrong→correct) 9 10.0%

14.4% degradation + adversary's 3 extra wrong votes flips 6/30 topics.

4.5 Key Insight

The collapse is structural, not persuasive. The adversary outvotes honest agents by injecting 3× votes into the majority pool. Even weak adversaries cause collapse. Simple protocol fixes (judge voting, confidence weighting, token caps) fully recover.


5. The OCC Stack

5.1 Design Principle

OCC treats compute allocation as a security boundary. Agents earn capability-scoped, decaying, non-transferable credits through verified marginal impact.

5.2 Components

Impact Oracle (oracle/oracle.py): Scores (action, context, result) → JSON with raw score, cost-adjusted score, confidence, evidence, reason, failure tags, reward. Supports code, QA, and debate modes.

Credit Ledger (ledger/ledger.py): Append-only log. Credits are non-transferable, decaying (δ=0.995/turn), capability-scoped, revocable. SHA-256 hash chain for auditability.

Resource Broker (broker/broker.py): Capability-based access control. Decides allow/deny/downgrade/escalate/require-approval.

GRPO Hook (rl/grpo_hook.py): TRL-compatible reward:

reward = oracle_score + abstention_utility + calibration_bonus
         − hallucination_penalty − confident_wrong_penalty
         − compute_cost × cost_multiplier − gaming_penalty

5.3 Anti-Gaming

10 attack vectors tested, all contained: credit farming, collusion, oracle spoofing, verbosity gaming, confidence manipulation, strategic abstention, identity laundering, sybil agents, sandbagging, griefing.


6. Benchmarks

6.1 Code Compute Allocation (Simulated)

Strategy Pass@1 Compute Savings
Fixed budget 0.78 17,500
OCC tiered 0.78 8,350 52.3%

6.2 Multi-Agent Debate (Simulated)

Strategy Accuracy Containment
Conf-weighted voting 0.56 0%
OCC credit filtering 0.76 100%

6.3 Real LLM Results

Debate Collapse (Qwen3-Coder-30B, H200): §3-4 results. Real inference.

HumanEval (Qwen3-Coder-30B): 42.1% pass@1, 67.8% compute savings via adaptive retry. Honestly labeled as adaptive retry, not OCC.

TruthfulQA (Qwen3-Coder-30B, AllenAI judges): OCC+Abstention iso-quality (0.917) with 21.1% fewer tokens. Savings are judge-dependent.


7. Ablations

Ablation Effect
No credit ledger 27% less savings
Transferable credits Gaming: 0% → 45%
Non-decaying credits Hoarding, −18% throughput
No confident-wrong penalty 2.3× higher rate
No calibration penalty ECE: 0.12 → 0.31
No cost penalty Tokens +40%
No anti-gaming penalty Gaming agents earn 3.2× more

8. Honest Assessment

What Worked

  • Debate collapse is real: 73.3% → 56.7% with 3× compute
  • Mechanism isolated: volume amplification, not persuasion (84% retention)
  • Protocol fixes work: judge/confidence voting fully recover
  • Anti-gaming sound: 10 attacks, all contained
  • OCC prevents catastrophic collapse (20pp recovery in debate)

What Failed

  • OCC ≈ random gating at moderate budgets (83.3% vs 85.0%)
  • GRPO training: no improvement at 0.5B scale
  • Retrieval QA: accuracy lags baseline (0.71 vs 0.79)
  • HumanEval: adaptive retry, not OCC

Limitations

  • Single seed (n=30, CI ±16pp)
  • Simulated benchmarks for code/QA
  • GRPO not trained at meaningful scale
  • Narrow domain (yes/no trivia)
  • Same-model judge and debater
  • Scripted adversary

9. Conclusion

Compute is not neutral in multi-agent systems. Extra turns amplify adversarial influence unless governed by verified marginal contribution. We demonstrated this empirically, isolated the mechanism, showed protocol fixes, and presented OCC as a governance layer.

Open-source: https://huggingface.co/narcolepticchicken/occ-stack


Citation

@misc{occ2026,
  title={Compute Is Not Neutral: Mechanism Analysis of Adversarial Debate Collapse and the OCC Stack},
  author={narcolepticchicken},
  year={2026},
  url={https://huggingface.co/narcolepticchicken/occ-stack}
}