FINAL-Bench Quantum Quantum Olympics

An open, neutral benchmark suite for quantum-computing methods — five events, one fair yardstick.

양자 컴퓨팅 방법론을 공정하게 비교하는 공개 벤치마크 (5종목). FINAL-Bench=중립 주최, 양자우위 주장 아님.

What is this?

Quantum computers are powerful but noisy. Making them useful needs better methods — decoders, optimizers, chemistry solvers, memories, simulators. FINAL-Bench Quantum measures and ranks these on fixed, public test sets. FINAL-Bench is a neutral organizer; every method — incl. VIDRAFT / QuantumOS — is evaluated under the same protocol and ranked honestly. A method benchmark, not a quantum-advantage claim.

The Five Events

#EventWhat it measuresMetricStatus
QEC DecoderHow well a decoder fixes quantum errors from syndromes.Logical Error RateLIVE
Quantum OptimizationHow well/fast a solver cracks hard combinatorial problems.approx-ratio · timeplanned
Quantum Chemistry (VQE)How accurately a circuit estimates ground-state energy.error (mHa)planned
QRAM / Quantum MemoryHow faithfully data is stored & retrieved.fidelityplanned
Quantum-Inspired SimHow accurately/quickly classical HW simulates quantum circuits.accuracy · timeplanned

① QEC Decoder Leaderboard LIVE · beta

A. Head-to-head ranking same test set · d=5 · rotated surface code

All four decoders run on the identical Stim test set. Ranked by Logical Error Rate at p=0.005 (lower = better).

RankDecoderLER @ p=0.005LER @ p=0.01
1BeliefMatching (BP+MWPM)0.009080.06378
2BP+OSD (stimbposd)0.009150.06443
3VIDRAFT QEC-AI Decoder v3 · meta-ensemble0.010650.07863
4PyMatching (MWPM)0.013130.08023

Honest note. Our VIDRAFT QEC-AI Decoder (v3) — a learned meta-ensemble of matching + belief-propagation + a neural network — beats the standard MWPM decoder (0.0107 vs 0.0131) and trails the two strongest specialized decoders (BeliefMatching, BP+OSD) by ~17%. We list it at its true rank (3rd of 4). Learned decoders' biggest advantage appears under real-hardware noise — see the upcoming IBM QPU track.

A2. Scaling reference — PyMatching (MWPM), d = 3 / 5 / 7

pd=3d=5d=7
0.0017.8e-41.25e-42.0e-5
0.0051.68e-21.41e-21.01e-2
0.0106.0e-28.3e-21.04e-1

Threshold ≈ 0.7% — surface-code scaling sanity check.

B. Published reference results from literature — setups differ

MethodReported resultSetupSource
Google Willow (surface code, hardware)0.143% ± 0.003% logical error / cycle (d=7); Λ=2.14superconducting HW, real-timeNature 2024
AlphaQubit (neural decoder)-30% errors vs correlated matching; -6% vs tensor-networkSycamore d3-d5 + sim to d11Nature 2024
Astra (graph neural network)higher threshold & lower LER than BP+OSDsurface to d11; BB to d18npj QI 2025

Codes/noise/hardware differ — references, not a head-to-head with track A.

Roadmap

IBM QPU real-hardware track (where learned decoders beat matching) · more baselines · distance-scaling to d=35 · open submissions · then events ② – ⑤.

Honesty. FINAL-Bench is a neutral host; all methods (incl. our VIDRAFT QEC-AI Decoder) are scored on the identical public test set and ranked honestly — listed at true rank, never a fake #1. Track B numbers are quoted from sources (setups differ). Decoder benchmark — no quantum-advantage claims. © FINAL-Bench · powered by QuantumOS.