hv-quorum-inference

Enforced-diversity consensus with collapse detection for LLM outputs. NumPy-only core; works with any backend that implements sample() and embed().

Keywords: collapse detection, quorum, ensemble, mode collapse, LLM safety, numpy-only.

Size: ~14 KB source, no weights. Runtime: ~9 s for the full benchmark on CPU. Dependencies: NumPy only.

The primitive

Standard majority vote runs N candidates and picks the most common. When the underlying model has a biased mode collapse, majority vote confidently returns the wrong answer. The quorum algorithm detects the collapse and returns collapsed=True instead of a wrong answer.

Three methods in one file:

method output
majority_vote most common answer, no collapse detection
quorum most common answer or a collapsed=True flag

The consumer decides what to do with a flag: re-prompt, use a different model, route to a human, or simply refuse.

Headline numbers

Collapse detection at n_samples=64, dominance_threshold=0.70:

collapse rate majority acc quorum flag quorum useful
0.00 0.965 0.030 0.970
0.30 0.455 0.000 0.455
0.50 0.010 0.000 0.025
0.60 0.000 0.015 0.015
0.70 0.000 0.525 0.525
0.80 0.000 0.980 0.980

useful = quorum accuracy + quorum flag rate. This is the fraction of prompts that either got a correct answer or were correctly flagged as collapsed.

The detection floor is 0.70 on this synthetic model. Below that threshold, the answer distribution cannot exceed the dominance threshold, so no flag is raised.

The detection floor is structural

A dominance-based detector cannot distinguish:

  • "60% collapse with 40% random noise" from
  • "correct convergence with 60% share"

Both look the same in the answer-count distribution. The threshold must be set above the natural answer dominance of the target distribution to avoid false positives, and below the collapse signal to catch real collapse.

The floor is not fixed. It depends on the ratio between the model's natural confidence and the collapse signal:

model natural dominance threshold floor
synthetic (ambiguous prompt, p_top β‰ˆ 0.35) 0.35 0.70 0.70 collapse
real LLM (typical QA) 0.40–0.55 0.60–0.70 0.60–0.70 collapse
real LLM (open-ended) 0.20–0.35 0.50 0.50 collapse

The pattern: the floor is roughly threshold + (threshold - natural) because the collapse signal and natural confidence add linearly.

Threshold sweep at n=64

dominance threshold flag(0%) flag(50%) flag(70%)
0.55 0.190 0.150 0.995
0.60 0.115 0.025 0.955
0.65 0.070 0.005 0.775
0.70 0.030 0.000 0.525
0.75 0.010 0.000 0.105
0.80 0.000 0.000 0.000

The tradeoff:

  • 0.55–0.65: catches moderate collapse, flags 5–20% of clean prompts.
  • 0.70: flags 3% of clean prompts, catches collapse above 70%.
  • 0.75+: never fires except at extreme collapse.

The default 0.70 is a good balance for the synthetic benchmark. For production, tune on a held-out set of clean prompts.

How to use

from hv_quorum_inference import (
    QuorumConfig, quorum, majority_vote,
)

# Define your backend (any class with sample() and embed())
class MyBackend:
    def sample(self, prompt, n):
        # return a list of n candidate answers
        ...
    def embed(self, candidates):
        # return (n, d) float array of embeddings
        ...
    def is_correct(self, candidate, prompt):
        # optional; used by benchmarks
        ...

backend = MyBackend()

# Configure
cfg = QuorumConfig(
    n_samples=64,
    max_resamples=1,
    dominance_threshold=0.70,
)

# Use quorum for collapse detection
r = quorum(my_prompt, backend, cfg)
if r.collapsed:
    print("Model collapsed on this prompt. Re-prompt or escalate.")
else:
    print(f"Answer: {r.candidate}")
    print(f"Dominance: {r.final_dominance:.3f}")

# Or majority vote for baseline comparison
r = majority_vote(my_prompt, backend, n_samples=64)
print(f"Answer: {r.candidate}")
Downloads last month
138
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support