An open, neutral benchmark suite for quantum-computing methods — five events, one fair yardstick.
양자 컴퓨팅 방법론을 공정하게 비교하는 공개 벤치마크 (5종목). FINAL-Bench=중립 주최, 양자우위 주장 아님.
Quantum computers are powerful but noisy. Making them useful needs better methods — decoders, optimizers, chemistry solvers, memories, simulators. FINAL-Bench Quantum measures and ranks these on fixed, public test sets. FINAL-Bench is a neutral organizer; every method — incl. VIDRAFT / QuantumOS — is evaluated under the same protocol and ranked honestly. A method benchmark, not a quantum-advantage claim.
| # | Event | What it measures | Metric | Status |
|---|---|---|---|---|
| ① | QEC Decoder | How well a decoder fixes quantum errors from syndromes. | Logical Error Rate | LIVE |
| ② | Quantum Optimization | How well/fast a solver cracks hard combinatorial problems. | approx-ratio · time | planned |
| ③ | Quantum Chemistry (VQE) | How accurately a circuit estimates ground-state energy. | error (mHa) | planned |
| ④ | QRAM / Quantum Memory | How faithfully data is stored & retrieved. | fidelity | planned |
| ⑤ | Quantum-Inspired Sim | How accurately/quickly classical HW simulates quantum circuits. | accuracy · time | planned |
All four decoders run on the identical Stim test set. Ranked by Logical Error Rate at p=0.005 (lower = better).
| Rank | Decoder | LER @ p=0.005 | LER @ p=0.01 |
|---|---|---|---|
| 1 | BeliefMatching (BP+MWPM) | 0.00908 | 0.06378 |
| 2 | BP+OSD (stimbposd) | 0.00915 | 0.06443 |
| 3 | VIDRAFT QEC-AI Decoder v3 · meta-ensemble | 0.01065 | 0.07863 |
| 4 | PyMatching (MWPM) | 0.01313 | 0.08023 |
Honest note. Our VIDRAFT QEC-AI Decoder (v3) — a learned meta-ensemble of matching + belief-propagation + a neural network — beats the standard MWPM decoder (0.0107 vs 0.0131) and trails the two strongest specialized decoders (BeliefMatching, BP+OSD) by ~17%. We list it at its true rank (3rd of 4). Learned decoders' biggest advantage appears under real-hardware noise — see the upcoming IBM QPU track.
| p | d=3 | d=5 | d=7 |
|---|---|---|---|
| 0.001 | 7.8e-4 | 1.25e-4 | 2.0e-5 |
| 0.005 | 1.68e-2 | 1.41e-2 | 1.01e-2 |
| 0.010 | 6.0e-2 | 8.3e-2 | 1.04e-1 |
Threshold ≈ 0.7% — surface-code scaling sanity check.
| Method | Reported result | Setup | Source |
|---|---|---|---|
| Google Willow (surface code, hardware) | 0.143% ± 0.003% logical error / cycle (d=7); Λ=2.14 | superconducting HW, real-time | Nature 2024 |
| AlphaQubit (neural decoder) | -30% errors vs correlated matching; -6% vs tensor-network | Sycamore d3-d5 + sim to d11 | Nature 2024 |
| Astra (graph neural network) | higher threshold & lower LER than BP+OSD | surface to d11; BB to d18 | npj QI 2025 |
Codes/noise/hardware differ — references, not a head-to-head with track A.
IBM QPU real-hardware track (where learned decoders beat matching) · more baselines · distance-scaling to d=35 · open submissions · then events ② – ⑤.
Honesty. FINAL-Bench is a neutral host; all methods (incl. our VIDRAFT QEC-AI Decoder) are scored on the identical public test set and ranked honestly — listed at true rank, never a fake #1. Track B numbers are quoted from sources (setups differ). Decoder benchmark — no quantum-advantage claims. © FINAL-Bench · powered by QuantumOS.