hv-quorum-inference
Enforced-diversity consensus with collapse detection for LLM outputs.
NumPy-only core; works with any backend that implements sample() and
embed().
Keywords: collapse detection, quorum, ensemble, mode collapse, LLM safety, numpy-only.
Size: ~14 KB source, no weights. Runtime: ~9 s for the full benchmark on CPU. Dependencies: NumPy only.
The primitive
Standard majority vote runs N candidates and picks the most common.
When the underlying model has a biased mode collapse, majority vote
confidently returns the wrong answer. The quorum algorithm detects the
collapse and returns collapsed=True instead of a wrong answer.
Three methods in one file:
| method | output |
|---|---|
majority_vote |
most common answer, no collapse detection |
quorum |
most common answer or a collapsed=True flag |
The consumer decides what to do with a flag: re-prompt, use a different model, route to a human, or simply refuse.
Headline numbers
Collapse detection at n_samples=64, dominance_threshold=0.70:
| collapse rate | majority acc | quorum flag | quorum useful |
|---|---|---|---|
| 0.00 | 0.965 | 0.030 | 0.970 |
| 0.30 | 0.455 | 0.000 | 0.455 |
| 0.50 | 0.010 | 0.000 | 0.025 |
| 0.60 | 0.000 | 0.015 | 0.015 |
| 0.70 | 0.000 | 0.525 | 0.525 |
| 0.80 | 0.000 | 0.980 | 0.980 |
useful = quorum accuracy + quorum flag rate. This is the fraction of
prompts that either got a correct answer or were correctly flagged as
collapsed.
The detection floor is 0.70 on this synthetic model. Below that threshold, the answer distribution cannot exceed the dominance threshold, so no flag is raised.
The detection floor is structural
A dominance-based detector cannot distinguish:
- "60% collapse with 40% random noise" from
- "correct convergence with 60% share"
Both look the same in the answer-count distribution. The threshold must be set above the natural answer dominance of the target distribution to avoid false positives, and below the collapse signal to catch real collapse.
The floor is not fixed. It depends on the ratio between the model's natural confidence and the collapse signal:
| model | natural dominance | threshold | floor |
|---|---|---|---|
| synthetic (ambiguous prompt, p_top β 0.35) | 0.35 | 0.70 | 0.70 collapse |
| real LLM (typical QA) | 0.40β0.55 | 0.60β0.70 | 0.60β0.70 collapse |
| real LLM (open-ended) | 0.20β0.35 | 0.50 | 0.50 collapse |
The pattern: the floor is roughly threshold + (threshold - natural)
because the collapse signal and natural confidence add linearly.
Threshold sweep at n=64
| dominance threshold | flag(0%) | flag(50%) | flag(70%) |
|---|---|---|---|
| 0.55 | 0.190 | 0.150 | 0.995 |
| 0.60 | 0.115 | 0.025 | 0.955 |
| 0.65 | 0.070 | 0.005 | 0.775 |
| 0.70 | 0.030 | 0.000 | 0.525 |
| 0.75 | 0.010 | 0.000 | 0.105 |
| 0.80 | 0.000 | 0.000 | 0.000 |
The tradeoff:
- 0.55β0.65: catches moderate collapse, flags 5β20% of clean prompts.
- 0.70: flags 3% of clean prompts, catches collapse above 70%.
- 0.75+: never fires except at extreme collapse.
The default 0.70 is a good balance for the synthetic benchmark. For production, tune on a held-out set of clean prompts.
How to use
from hv_quorum_inference import (
QuorumConfig, quorum, majority_vote,
)
# Define your backend (any class with sample() and embed())
class MyBackend:
def sample(self, prompt, n):
# return a list of n candidate answers
...
def embed(self, candidates):
# return (n, d) float array of embeddings
...
def is_correct(self, candidate, prompt):
# optional; used by benchmarks
...
backend = MyBackend()
# Configure
cfg = QuorumConfig(
n_samples=64,
max_resamples=1,
dominance_threshold=0.70,
)
# Use quorum for collapse detection
r = quorum(my_prompt, backend, cfg)
if r.collapsed:
print("Model collapsed on this prompt. Re-prompt or escalate.")
else:
print(f"Answer: {r.candidate}")
print(f"Dominance: {r.final_dominance:.3f}")
# Or majority vote for baseline comparison
r = majority_vote(my_prompt, backend, n_samples=64)
print(f"Answer: {r.candidate}")
- Downloads last month
- 138