Instructions to use WestQuantStudio/WQT20M-Beta with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WestQuantStudio/WQT20M-Beta with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WestQuantStudio/WQT20M-Beta")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("WestQuantStudio/WQT20M-Beta") model = AutoModelForCausalLM.from_pretrained("WestQuantStudio/WQT20M-Beta", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WestQuantStudio/WQT20M-Beta with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WestQuantStudio/WQT20M-Beta" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WestQuantStudio/WQT20M-Beta", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/WestQuantStudio/WQT20M-Beta
- SGLang
How to use WestQuantStudio/WQT20M-Beta with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WestQuantStudio/WQT20M-Beta" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WestQuantStudio/WQT20M-Beta", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WestQuantStudio/WQT20M-Beta" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WestQuantStudio/WQT20M-Beta", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use WestQuantStudio/WQT20M-Beta with Docker Model Runner:
docker model run hf.co/WestQuantStudio/WQT20M-Beta
WQT20M-Beta
WestQuant Transformer 20M β Quantum Representation Scheduler
AI schedules. Deterministic mathematics executes. Independent verification certifies.
WQT20M-Beta is a 19M-parameter Transformer that ranks quantum transformations. Given a quantum optimization state, it predicts which transformation is most promising. It does not execute transformations β deterministic engines do that.
WQT20M-Beta is a proof-of-concept that a 19M-parameter Transformer can learn structured quantum optimization preferences from synthetic data. Researchers building better compiler-optimization models can use this as a baseline.
This is WestQuant Open's first step toward a production-ready transformer with calibrated value predictions (wider range, objective sensitivity, real QPU training data) that could replace heuristic transpiler defaults in large complex Quantum Computing production pipelines.
What It Does β 3 Tested Examples
Example 1: Preference Comparison (96% accuracy)
Given two candidate transformations with their costs, the model predicts which is better:
State: <DOMAIN:quantum_annealing> <LEVEL:ISING> <N_QUBITS:18> ...
<RES:n_q=18 D=8 G1=12 G2=4 T=0 M=3 A=0 E=0.0100 C=1.0000>
<OBJ_TYPE:balanced>
Candidate A: QUENCH (cost 1278.25)
Candidate B: SET_BIAS (cost 1245.57)
Model predicts: B>A (SET_BIAS is better)
Correct answer: B>A β
Example 2: Value Prediction (mean error ~8)
Given a state, the model predicts the cost-to-go (remaining optimization cost):
State: <DOMAIN:graph_optimization> <LEVEL:GRAPH> <N_NODES:21> ...
<RES:n_q=21 D=5 G1=10 G2=3 T=0 M=2 A=0 E=0.0100 C=1.0000>
<OBJ_TYPE:balanced>
Model predicts cost-to-go: 1275.93
Actual cost-to-go: 1264.63
Error: 11.30 (0.9%)
Example 3: Ranked Policy (73% Top-1 via value ranking)
Given a state and all legal candidate actions, the model ranks them by predicted cost and picks the best:
State: <DOMAIN:hardware_mapping> <LEVEL:COMPILED> <N_QUBITS:12> ...
<RES:n_q=12 D=6 G1=8 G2=2 T=0 M=1 A=0 E=0.0050 C=1.0000>
<OBJ_TYPE:2q_focused>
Candidates ranked by predicted cost:
1. NOISE_AWARE predicted=620.83 β model picks this
2. LAYOUT_SCORE predicted=631.83
3. DENSE_PLACE predicted=639.83
Oracle (actual best): NOISE_AWARE β Correct!
Quick Start
pip install transformers torch
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained("WestQuantStudio/WQT20M-Beta")
tokenizer = AutoTokenizer.from_pretrained("WestQuantStudio/WQT20M-Beta")
device = "mps" if torch.backends.mps.is_available() else "cpu"
model = model.to(device).eval()
text = ("<SOLVE> <DOMAIN:quantum_annealing> <LEVEL:ISING> "
"<N_QUBITS:18> <N_GROUND:14> <COUPLING_STRENGTH:0.8000> "
"<RES:n_q=18 D=8 G1=12 G2=4 T=0 M=3 A=0 E=0.0100 C=1.0000> "
"<OBJ_TYPE:balanced> "
"<CAND_A> QUENCH <COST_A> 1278.25 "
"<CAND_B> SET_BIAS <COST_B> 1245.57 "
"<PREF>")
ids = tokenizer.encode(text, add_special_tokens=False, return_tensors="pt").to(device)
with torch.no_grad():
for _ in range(5):
out = model(input_ids=ids)
nxt = out.logits[0, -1].argmax().unsqueeze(0).unsqueeze(0)
ids = torch.cat([ids, nxt], dim=1)
if nxt.item() == tokenizer.eos_token_id:
break
print(tokenizer.decode(ids[0].tolist()).split("<PREF>")[-1].strip())
# β "B>A" (SET_BIAS is better because cost 1245 < 1278)
What WQT20M-Beta Predicts
| Task | Input | Output | Accuracy |
|---|---|---|---|
| Preference | State + 2 candidates with costs | Which candidate is better | 96.1% |
| Value | State | Predicted cost-to-go | Spearman 0.98, mean error ~8 |
| Ranked Policy | State + all legal actions | Best action (via value ranking) | 73% Top-1 |
| Legality | State + action | Is this action legal? | 76.9% |
| Hardware | State + backend | Is this feasible? | 91.2% |
Model
| Config | Value |
|---|---|
| Architecture | Llama-style decoder Transformer |
| Parameters | 19.06M |
| Layers | 8 |
| Hidden size | 384 |
| Attention heads | 6 (2 KV heads, GQA) |
| FFN | 1536 (SwiGLU) |
| Context length | 2048 |
| Vocab | 4561 (quantum-native structured tokens + BPE) |
| Model size | 76 MB |
Standard HuggingFace LlamaForCausalLM. Compatible with AutoModelForCausalLM,
AutoTokenizer, and SafeTensors.
Validation
All 6 release gates passed:
| Gate | Result |
|---|---|
| Data integrity | PASS |
| Structural learning (no cheating) | PASS |
| Search improvement over random | PASS (+35% at budget=100) |
| Generalization to unseen problems | PASS (96%) |
| No regression | PASS |
| Reproduction | PASS |
Anti-cheating
The model reads structure, not identifiers:
- ID-only accuracy: 24% (domain token alone can't predict the answer)
- Shuffled structure: 74% (scrambling the structure hurts performance)
Search improvement
| Search budget | Random | WQT20M-Beta | Improvement |
|---|---|---|---|
| 10 evals | baseline | β | +3.5% |
| 50 evals | baseline | β | +17.5% |
| 100 evals | baseline | β | +35.0% |
QFT & Grover experiments
Independent experiments on QFT and Grover circuits across 18 configurations:
- 0.0 regret in all 18 configs β the model consistently picked CANCEL, the best action
- CANCEL ranked at position 1/15 (median model rank of true-best action)
- Top-3 rate: 63% β the true-best action is in the model's top-3 63% of the time
CALL_PYZXis always correctly identified as the worst action
Honest Assessment
Where WQT20M-Beta has narrow real-world value
The core limitation: exhaustive search over 15 transpiler configs takes <0.01s per circuit. Any user who can afford to run Qiskit at all can afford to try all candidates and pick the best. In experiments, many configurations were ties (model and random both found the optimum), and the model did not consistently beat random on small search spaces.
Where it would add genuine value
1. Scaling to large search spaces where exhaustive is infeasible
If the candidate space grows to hundreds or thousands of configurations (e.g., combining routing methods Γ layout methods Γ synthesis engines Γ basis gates Γ opt levels Γ seeds), exhaustive evaluation becomes expensive. The model's value-ranking approach becomes meaningful when each evaluation costs real QPU time or long simulation. A user compiling 10,000 circuits Γ 500 candidate configs would benefit from model-guided pruning.
2. Multi-step optimization pipelines
The model predicts cost-to-go, not just immediate cost. In a sequential optimization pipeline (apply transformation A, then B, then C), the model could guide intermediate decisions where the full tree is exponentially large. This is the model's actual design intent β it's a scheduler, not a single-shot selector.
3. Researchers studying AI-guided quantum compilation
The model is a research artifact. Its value is as a proof-of-concept that a 19M-parameter Transformer can learn structured quantum optimization preferences from synthetic data. Researchers building better compiler-optimization models would use this as a baseline.
4. WestQuant iterating toward a production model
The Beta is explicitly a stepping stone. A future version with calibrated value predictions (wider range, objective sensitivity, real QPU training data) could replace heuristic transpiler defaults in production pipelines.
Who would NOT benefit
- Practitioners running small circuits β exhaustive search is free and optimal
- Users needing guaranteed optimality β the model has nonzero regret and 76.9% legality accuracy
- Anyone with objective-dependent optimization β the model has 0% objective sensitivity
- Users on real QPUs today β the model was trained on synthetic data, not hardware
Bottom line
The model adds real value only when the candidate search space is large enough that exhaustive evaluation is expensive, and when approximate guidance is acceptable. In its current Beta form, that threshold is well above what Qiskit's built-in transpiler exposes. The honest framing is: this is a research prototype demonstrating that learned quantum-structure preferences are feasible, not a production tool that outperforms brute force on practical workloads.
Roadmap to Production
| Gap | Current (Beta) | Production Target |
|---|---|---|
| Training data | Synthetic surrogate | Real Qiskit/TKET transpilation outputs |
| Value calibration | Narrow range (6.0β6.6) | Actual cost magnitudes (50β25,000) |
| Objective sensitivity | 0% | Changes ranking with objective |
| Legality | 76.9% | >95% |
| Optimization steps | Single-step | Multi-step trajectory planning |
| Backend awareness | Topology tokens only | Per-edge error rates, per-qubit T1/T2 |
| Model size | 19M | 100β200M |
| Feedback loop | None | Active learning from user compilations |
| Deployment | Standalone script | Qiskit transpiler plugin with budget parameter |
Key changes needed
1. Real compiler outputs, not synthetic β Train on actual transpilation trajectories from Qiskit/TKET/PyZX across 10,000+ circuits from MQT Bench, QASMBench, and circuit libraries, with 20-50 real backend coupling maps.
2. Calibrated value prediction β Add a regression head that predicts cost in the actual cost range of the training data, not a narrow band.
3. Objective conditioning that works β Train with objective-dependent labels where the same (state, action) pair has different costs depending on the objective (depth-focused, 2q-focused, fidelity-focused, time-focused).
4. Multi-step trajectory planning β Train on real optimization trajectories with cumulative cost at each step, so the model learns which sequences are promising, not just which single action looks good.
5. Legality prediction β A binary classification head trained on real constraint data (which actions are valid for a given backend, circuit family, and optimization level). Target >95% accuracy.
6. Real backend calibration β Include actual per-edge error rates, per-qubit T1/T2, gate durations, and queue/wait times in the state encoding.
7. Larger model β 100-200M parameters, 12-16 layers, 768 hidden, 4096+ context length to encode full circuit metrics + backend calibration + action history.
8. Active learning β A feedback loop where the model improves from real user compilations: user compiles, model suggests, Qiskit applies, real cost is measured, (state, action, real_cost) is logged, model is fine-tuned.
9. Production deployment β A Qiskit transpiler plugin where the model ranks candidates, top-K are evaluated by real Qiskit, and the user specifies a budget parameter (how many real evaluations they can afford).
The model's architecture and approach (value-prediction-based ranking, quantum-native tokenization, structured state encoding) are sound. The gap is entirely in training data quality, calibration, and multi-step capability. Fix those three and it becomes a model someone would use in production.
Coverage
15 quantum optimization domains:
| Domain | Example Actions |
|---|---|
| Quantum chemistry | UCCSD_ANSATZ, ADAPT_VQE, FROZEN_CORE |
| Graph optimization | MAXCUT_ROUND, COLOR_GRAPH, SDP_RELAX |
| Error correction | SYNDROME_MEASURE, DECODE_SURFACE, MAGIC_STATE_DISTILL |
| Quantum annealing | REVERSE_ANNEAL, MINOR_EMBED, HYBRID_SOLVE |
| Variational | QAOA_P1/P2/P3, WARM_START, ADAPTIVE_LAYER |
| Hamiltonian simulation | TROTTER_STEP, LCU_DECOMPOSE, QUBITIZE |
| Circuit optimization | CANCEL_GATES, ZX_SIMPLIFY, FUSE_ROTATIONS |
| Hardware mapping | SABRE_ROUTE, PULSE_OPTIMIZE, NOISE_AWARE |
| Neutral atom | SET_RYDBERG, PULSE_SHAPE, ADIABATIC_PASS |
| Photonic | KLM_CNOT, CLUSTER_STATE, HERALD |
| Topological | BRAID_ANYON, FIBONACCI_BRAID |
| Quantum walk | COINED_WALK, GROVER_COIN, AMPLIFY |
| State preparation | MPS_PREP, TENSOR_NETWORK, COMPRESS_STATE |
| Amplitude amplification | GROVER_ITER, QSEARCH, ITERATIVE_QPE |
| Quantum ML | IQP_KERNEL, DATA_REUPLOADING, FIDELITY_KERNEL |
Limitations
This is a Beta release:
- Policy Top-1 (generation): 0.8% β The model cannot reliably generate the best action name. However, ranked policy via value prediction achieves 73% Top-1 β use the value-ranking approach, not generation.
- Legality: 76.9% β Below target. The model sometimes marks illegal actions as valid.
- Objective sensitivity: 0% β The model does not yet change its predictions when the optimization objective changes.
- Value calibration β Predicted values cluster in a narrow range (6.0β6.6), while actual costs range from 50 to 25,000. The model has good rank correlation (Spearman 0.98) but poor absolute calibration.
- Not a circuit compiler β The model works on its own structured state representation, not on raw Qiskit/TKET circuits. The plugins translate between the model's representation and real frameworks.
License
Apache 2.0
Citation
@misc{westquant2026wqt20m,
title={WQT20M-Beta: A 19M-Parameter Transformer for Quantum Representation Scheduling},
author={Vesterlund, David},
year={2026},
url={https://huggingface.co/WestQuantStudio/WQT20M-Beta},
note={WestQuant Open Source Project}
}
- Downloads last month
- -