Papers
arxiv:2605.22949

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination

Published on Oct 8
· Submitted by
Joss Armstrong
on Oct 9
Authors:

Abstract

When a coordinator compares answers from heterogeneous foundation models, self-reported confidence may have different meanings across responders and changing workloads. This paper presents MARGIN (Multi-Agent Runtime Grading via Incremental Normalisation), a runtime calibration method that learns model-specific confidence corrections from observed answer outcomes without retraining the models or requiring a held-out calibration set. MARGIN tracks recent accuracy and stated confidence within confidence bands, uses their ratio to correct reported confidence, and blends sparse-band corrections toward a model-level estimate. The corrected scores weight candidate answers in a collective decision. Evaluation covers code generation, question answering, and mathematics, using an 18-model pool and a nine-model subset for distribution-shift experiments. On BigCodeBench, model-mean confidence is negatively related to accuracy; among correct/incorrect response pairs, choosing the more confident responder performs below chance. Against five online calibration baselines receiving identical feedback and retaining their learned state across each transition, MARGIN achieves lower post-shift expected calibration error than all five in two code-generation transitions and than four in a question-answering transition; the remaining question-answering comparison is inconclusive. In separate code-generation coordination experiments, calibration improves the ranking of correct responses and increases answer-selection accuracy by 4.3 and 14.0 percentage points on two of three benchmarks relative to uncalibrated confidence weighting. These results support model-specific runtime calibration for coordination under changing workloads when correctness feedback is available for the participating responders.

Community

Paper author Paper submitter

Self-reported confidence is the wrong signal for combining answers from multiple LLMs. On BigCodeBench, across 18 models, confidence runs backwards. Models with higher average confidence pass fewer tests. When one of two answers is right and the other wrong, the more confident model is right only 42% of the time, worse than a coin toss.
MARGIN fixes this at runtime. It learns what each model's confidence actually means from that model's own track record, band by band, with no retraining, no model internals and no calibration set. When the workload shifts, it beats all five online calibration baselines on both coding transitions.
The coordinator's choices improve. Faced with one right and one wrong answer, it now picks the right one 81% of the time on HumanEval and MBPP, up from about half. On BigCodeBench, the share of tasks where the selected answer is correct rises from 14% to 28%.
Method, worked example and results: https://jossarmstrong.org/works/margin/

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2605.22949 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2605.22949 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2605.22949 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.