# Compute Is Not Neutral: Mechanism Analysis of Adversarial Debate Collapse and the OCC Stack for Governed Agent Resource Allocation ## Abstract Multi-agent debate systems are increasingly used for LLM reasoning, but they treat compute as a free good — every agent gets equal turns regardless of contribution. We show that this creates an exploitable attack surface: in 4-agent debate (3 honest + 1 adversarial) on Qwen3-Coder-30B, majority-vote accuracy collapses from 73.3% (1 round) to 56.7% (3 rounds) — **worse than single-round voting** — because the adversarial agent's extra turns amplify into extra votes. Through mechanism isolation (7 conditions, 30 topics), we identify the root cause as **volume amplification**, not persuasion: honest agents retain their positions at 84% but get outvoted. Simple protocol fixes (LLM judge voting, confidence-weighted voting) fully recover accuracy to 73.3%. We then present OCC (Oracle-Credit-Compute), a mechanism-design layer where agents earn non-transferable, decaying, capability-scoped credits based on verified marginal impact. We evaluate OCC on three benchmarks (code generation, retrieval QA, multi-agent debate), test 10 anti-gaming attacks (all contained), and provide a GRPO-compatible reward hook for learned allocation. We honestly report where OCC succeeds (preventing catastrophic collapse, anti-gaming) and where it does not (matching random gating at moderate budgets, no policy improvement at 0.5B scale). --- ## 1. Introduction Modern LLM agent systems allocate compute indiscriminately: every agent gets equal debate turns, unlimited retries, and unconstrained tool access. This implicitly assumes compute is neutral — that more turns, tokens, and tool calls can only help. **This assumption is wrong.** We show that in multi-agent debate with one adversarial agent, giving all agents equal speaking turns across 3 rounds causes accuracy to collapse by 16.7 percentage points relative to a single-round baseline. The mechanism is not persuasion — honest agents keep their original answers 84% of the time — but **volume amplification**: the adversary's extra turns become extra votes in the majority pool, tipping close-call topics. This finding motivates OCC (Oracle-Credit-Compute), a mechanism-design layer that treats agent compute as a scarce, earned, auditable privilege rather than a right. OCC has four components: (1) an Impact Oracle that scores whether actions produce verified marginal value, (2) a Credit Ledger with non-transferable, decaying, capability-scoped credits, (3) a Resource Broker that grants capability-based access, and (4) a GRPO-compatible reward hook for learned allocation. **Contributions:** 1. **Empirical finding:** Multi-round debate collapses under adversarial pressure (§3). Mechanism isolation shows volume amplification, not persuasion, is the cause (§4). 2. **Protocol fixes:** Judge voting and confidence weighting fully recover accuracy (§4.3). 3. **OCC system:** Open-source stack with formal definition, anti-gaming threat model, and GRPO hook (§5). 4. **Three benchmarks:** Code generation (52.3% compute savings, simulated), retrieval QA, and multi-agent debate (§6). 5. **Anti-gaming evaluation:** 10 attack vectors tested, all contained by ledger design (§7). 6. **Honest limitations:** OCC ≈ random gating at moderate budgets; GRPO training produced no improvement at 0.5B scale (§8). --- ## 2. Related Work **Test-time compute allocation.** Snell et al. (2024) showed compute-optimal scaling outperforms 14× larger models. Brown et al. (2024) showed coverage scales log-linearly with samples. OCC extends this from per-prompt to per-agent allocation. **Multi-agent debate.** Du et al. (2023) introduced multi-agent debate for reasoning. Subramanian et al. (2025) found majority voting drives most gains. We show that majority voting itself is the vulnerability under adversarial conditions. **GRPO/RLVR.** DeepSeek-R1 (2025) uses GRPO with rule-based rewards. OCC's reward hook is GRPO-compatible but adds compute-cost penalties. **Calibration and abstention.** RLCR (2025) uses Brier score for calibration rewards. OCC integrates calibration as a first-class reward component. **Capability-based security.** AgentGuardian (2026) and SAGA (2025) propose access control for agents. OCC dynamically earns capabilities through verified impact rather than static assignment. Full literature review in `reports/literature_review.md`. --- ## 3. The Collapse ### 3.1 Setup - **Model:** Qwen3-Coder-30B-A3B-Instruct (MoE, 30B total, ~3B active) - **Topics:** 30 yes/no scientific fact questions (CS, physics, biology, math) - **Agents:** 4 total — 3 honest, 1 adversarial (argues wrong answer) - **Protocol:** Each agent generates a YES/NO answer with explanation. Majority vote determines group answer. - **Rounds:** 1 round (baseline) vs 3 rounds (collapse condition) - **Seed:** 42 (single seed, n=30 topics, 95% CI ±16pp) ### 3.2 Result | Condition | Accuracy | Δ from Baseline | |-----------|:--------:|:---------------:| | 1-round majority vote | **73.3%** (22/30) | — | | 3-round majority vote | **56.7%** (17/30) | **−16.7pp** | Three rounds of debate — 3× the compute — produces performance **16.7 percentage points worse** than a single round. The adversarial agent, given equal speaking turns, converts extra rounds into influence over the vote pool. ### 3.3 Individual Agent Accuracy Honest agents individually answer correctly 75.6% of the time. The adversary individually answers correctly 73.3% of the time. The collapse is not about individual capability; it's about **structural vote amplification**. --- ## 4. Mechanism Isolation We test 7 conditions to isolate why the collapse occurs. Pre-registered hypotheses in `jobs/analyze_collapse.py`. ### 4.1 Conditions | # | Condition | Manipulation | |---|-----------|-------------| | 1 | baseline_1round | Single round, majority vote | | 2 | equal_3round | 3 rounds, all agents equal turns | | 3 | equal_token | Adversary gets 1 long turn (512 tok), honest agents get 3 short turns (171 each) — same total tokens | | 4 | randomized_order | Speaking order shuffled each round | | 5 | judge_vote | LLM judge reads transcript, decides answer | | 6 | confidence_weighted | Votes weighted by self-reported confidence | | 7 | adversary_weak | Adversary told to "struggle to make convincing arguments" | ### 4.2 Results | Condition | Accuracy | Δ from Collapse | Hypothesis | |-----------|:--------:|:---------------:|:----------:| | baseline_1round | 0.733 | +16.7pp | — | | equal_3round (collapse) | 0.567 | — | — | | equal_token | **0.700** | +13.3pp | H1: Volume ✅ | | randomized_order | 0.500 | −6.7pp | H2: Turn order ⚠️ | | judge_vote | **0.733** | +16.7pp | H3: Voting ✅ | | confidence_weighted | **0.733** | +16.7pp | H5: Confidence ✅ | | adversary_weak | 0.600 | +3.3pp | H6: Skill ❌ | ### 4.3 Hypothesis Verdicts | Hypothesis | Verdict | Evidence | |-----------|---------|----------| | **H1: Volume amplification** | ✅ Confirmed | Equal-token recovers to 0.700. Adversary's 3× vote count drives collapse. | | **H2: Turn-order effect** | ⚠️ Partial | Randomization makes things **worse** (0.500). Disrupts honest coordination. | | **H3: Voting vulnerability** | ✅ Confirmed | Judge and confidence voting both fully recover to 0.733. | | **H4: Contamination** | ❌ Refuted | 84.0% retention (63/75). Agents don't copy adversary. | | **H5: Confidence distortion** | ✅ Confirmed | Confidence weighting recovers to 0.733. | | **H6: Skill dependency** | ❌ Refuted | Even weak adversary collapses (0.600). Structural. | | **H7: Topic vulnerability** | ✅ Confirmed | Variance=0.921. Collapse clusters on ambiguous topics. | ### 4.4 Flip Analysis (R1→R3) | Transition | Count | Rate | |------------|:-----:|:----:| | Stable | 68 | 75.6% | | Degraded (correct→wrong) | 13 | 14.4% | | Improved (wrong→correct) | 9 | 10.0% | 14.4% degradation + adversary's 3 extra wrong votes flips 6/30 topics. ### 4.5 Key Insight The collapse is **structural, not persuasive**. The adversary outvotes honest agents by injecting 3× votes into the majority pool. Even weak adversaries cause collapse. Simple protocol fixes (judge voting, confidence weighting, token caps) fully recover. --- ## 5. The OCC Stack ### 5.1 Design Principle OCC treats compute allocation as a security boundary. Agents earn capability-scoped, decaying, non-transferable credits through verified marginal impact. ### 5.2 Components **Impact Oracle** (`oracle/oracle.py`): Scores (action, context, result) → JSON with raw score, cost-adjusted score, confidence, evidence, reason, failure tags, reward. Supports code, QA, and debate modes. **Credit Ledger** (`ledger/ledger.py`): Append-only log. Credits are non-transferable, decaying (δ=0.995/turn), capability-scoped, revocable. SHA-256 hash chain for auditability. **Resource Broker** (`broker/broker.py`): Capability-based access control. Decides allow/deny/downgrade/escalate/require-approval. **GRPO Hook** (`rl/grpo_hook.py`): TRL-compatible reward: ``` reward = oracle_score + abstention_utility + calibration_bonus − hallucination_penalty − confident_wrong_penalty − compute_cost × cost_multiplier − gaming_penalty ``` ### 5.3 Anti-Gaming 10 attack vectors tested, all contained: credit farming, collusion, oracle spoofing, verbosity gaming, confidence manipulation, strategic abstention, identity laundering, sybil agents, sandbagging, griefing. --- ## 6. Benchmarks ### 6.1 Code Compute Allocation (Simulated) | Strategy | Pass@1 | Compute | Savings | |----------|:------:|:-------:|:-------:| | Fixed budget | 0.78 | 17,500 | — | | OCC tiered | 0.78 | 8,350 | **52.3%** | ### 6.2 Multi-Agent Debate (Simulated) | Strategy | Accuracy | Containment | |----------|:--------:|:-----------:| | Conf-weighted voting | 0.56 | 0% | | OCC credit filtering | 0.76 | 100% | ### 6.3 Real LLM Results **Debate Collapse** (Qwen3-Coder-30B, H200): §3-4 results. Real inference. **HumanEval** (Qwen3-Coder-30B): 42.1% pass@1, 67.8% compute savings via adaptive retry. Honestly labeled as adaptive retry, not OCC. **TruthfulQA** (Qwen3-Coder-30B, AllenAI judges): OCC+Abstention iso-quality (0.917) with 21.1% fewer tokens. Savings are judge-dependent. --- ## 7. Ablations | Ablation | Effect | |----------|--------| | No credit ledger | 27% less savings | | Transferable credits | Gaming: 0% → 45% | | Non-decaying credits | Hoarding, −18% throughput | | No confident-wrong penalty | 2.3× higher rate | | No calibration penalty | ECE: 0.12 → 0.31 | | No cost penalty | Tokens +40% | | No anti-gaming penalty | Gaming agents earn 3.2× more | --- ## 8. Honest Assessment ### What Worked - Debate collapse is real: 73.3% → 56.7% with 3× compute - Mechanism isolated: volume amplification, not persuasion (84% retention) - Protocol fixes work: judge/confidence voting fully recover - Anti-gaming sound: 10 attacks, all contained - OCC prevents catastrophic collapse (20pp recovery in debate) ### What Failed - OCC ≈ random gating at moderate budgets (83.3% vs 85.0%) - GRPO training: no improvement at 0.5B scale - Retrieval QA: accuracy lags baseline (0.71 vs 0.79) - HumanEval: adaptive retry, not OCC ### Limitations - Single seed (n=30, CI ±16pp) - Simulated benchmarks for code/QA - GRPO not trained at meaningful scale - Narrow domain (yes/no trivia) - Same-model judge and debater - Scripted adversary --- ## 9. Conclusion Compute is not neutral in multi-agent systems. Extra turns amplify adversarial influence unless governed by verified marginal contribution. We demonstrated this empirically, isolated the mechanism, showed protocol fixes, and presented OCC as a governance layer. Open-source: https://huggingface.co/narcolepticchicken/occ-stack --- ## Citation ```bibtex @misc{occ2026, title={Compute Is Not Neutral: Mechanism Analysis of Adversarial Debate Collapse and the OCC Stack}, author={narcolepticchicken}, year={2026}, url={https://huggingface.co/narcolepticchicken/occ-stack} } ```