Instructions to use WestQuantStudio/WQT50M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WestQuantStudio/WQT50M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WestQuantStudio/WQT50M")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("WestQuantStudio/WQT50M") model = AutoModelForCausalLM.from_pretrained("WestQuantStudio/WQT50M", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WestQuantStudio/WQT50M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WestQuantStudio/WQT50M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WestQuantStudio/WQT50M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/WestQuantStudio/WQT50M
- SGLang
How to use WestQuantStudio/WQT50M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WestQuantStudio/WQT50M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WestQuantStudio/WQT50M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WestQuantStudio/WQT50M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WestQuantStudio/WQT50M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use WestQuantStudio/WQT50M with Docker Model Runner:
docker model run hf.co/WestQuantStudio/WQT50M
WQT50M
WestQuant Transformer 50M β Quantum Representation Scheduler
AI schedules. Deterministic mathematics executes. Independent verification certifies.
WQT50M is a 50M-parameter Transformer that ranks quantum transformations. Trained on real Qiskit transpilation outputs β not synthetic data. It predicts which transformation is most promising given a circuit, backend, and optimization objective.
WQT50M is a proof-of-concept that a 50M-parameter Transformer can learn structured quantum optimization preferences from real compiler outputs. Researchers building better compiler-optimization models can use this as a baseline.
What It Does β 3 Tested Examples
Example 1: 62% Gate Reduction (Random Circuit, Fidelity-Focused)
Problem: 8-qubit random circuit (depth 30) on grid topology
Naive transpilation β cost 1231.2 (fidelity-focused objective)
WQT50M schedules ZX_SIMPLIFY
β cost 467.4 (62.0% reduction)
Random selection β cost 869.6 (29.4% reduction)
Oracle (exhaustive) β cost 467.4 (62.0%)
WQT50M matches oracle β
Example 2: Objective-Aware Scheduling (QAOA, Linear Backend)
The same QAOA circuit gets different best actions depending on the objective:
QAOA-6q-p2 on linear backend:
Objective WQT50M picks
βββββββββββββββββ βββββββββββββββββ
balanced β CANCEL_GATES
2q_focused β CANCEL_GATES
depth_focused β NATIVE_GATESET
fidelity_focused β FUSE_ROTATIONS
time_focused β MERGE_ADJACENT
4 different actions for 5 objectives β
Example 3: Calibrated Value Prediction
WQT50M predicts costs in the real magnitude range (not a narrow band):
State: <DOMAIN:circuit_optimization> <LEVEL:CIRCUIT> <N_QUBITS:8> ...
<BACKEND:SUPERCONDUCTING> <TOPO:grid> <T1:180us> <T2:90us>
<OBJ_TYPE:fidelity_focused>
Action: ZX_SIMPLIFY
Predicted cost: 467.4
Actual cost: 467.4
Error: 0.0%
Calibration ratio: 0.999 (predicted range matches actual range)
Spearman correlation: 0.981
Quick Start
pip install transformers torch
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained("WestQuantStudio/WQT50M")
tokenizer = AutoTokenizer.from_pretrained("WestQuantStudio/WQT50M")
device = "mps" if torch.backends.mps.is_available() else "cpu"
model = model.to(device).eval()
# Predict cost for a state-action pair
state = ("<DOMAIN:circuit_optimization> <LEVEL:CIRCUIT> "
"<N_QUBITS:8> <ENTANGLEMENT:0.5000> "
"<RES:n_q=8 D=30 G1=100 G2=20 T=120 M=0 A=0.2000 E=0.0050 C=0.5000> "
"<BACKEND:SUPERCONDUCTING> <TOPO:grid> "
"<T1:180us> <T2:90us> <READOUT_ERR:0.0120> "
"<OBJ_TYPE:fidelity_focused> <STEP:0/5>")
text = f"<PREDICT> {state} <ACTION> ZX_SIMPLIFY <COST> "
ids = tokenizer.encode(text, add_special_tokens=False, return_tensors="pt").to(device)
with torch.no_grad():
for _ in range(8):
out = model(input_ids=ids)
nxt = out.logits[0, -1].argmax().unsqueeze(0).unsqueeze(0)
ids = torch.cat([ids, nxt], dim=1)
if nxt.item() == tokenizer.eos_token_id:
break
print(tokenizer.decode(ids[0].tolist()).split("<COST>")[-1].strip())
# β predicted cost in real magnitude range
What WQT50M Predicts
| Task | Input | Output | Performance |
|---|---|---|---|
| Value prediction | State + action | Predicted cost (calibrated) | Spearman 0.981, ratio 0.999 |
| Ranked policy | State + all actions | Best action (via value ranking) | 74% Top-1 |
| Objective sensitivity | Same state, different objectives | Different best actions | 65% change rate |
| Backend awareness | Same state, different backends | Different best actions | 60% change rate |
| Legality | State + action | Is this action legal? | 79.2% |
| Search improvement | Model-guided vs random | Wins over random | 90.4% win rate |
Model
| Config | WQT20M-Beta | WQT50M |
|---|---|---|
| Architecture | Llama-style decoder | Llama-style decoder |
| Parameters | 19.06M | 49.53M |
| Layers | 8 | 12 |
| Hidden size | 384 | 512 |
| Attention heads | 6 (2 KV) | 8 (4 KV) |
| FFN | 1536 | 2048 |
| Context length | 2048 | 4096 |
| Vocab | 4561 | 4561 |
| Model size | 76 MB | 189 MB |
| Training data | Synthetic surrogate | Real Qiskit transpilation |
Standard HuggingFace LlamaForCausalLM. Compatible with AutoModelForCausalLM,
AutoTokenizer, and SafeTensors.
Validation Results
450 tests: 15 circuits Γ 5 objectives Γ 6 backends
| Metric | WQT20M-Beta | WQT50M |
|---|---|---|
| Top-1 accuracy (picks best action) | ~22% | 74.0% |
| Beats random | 58% | 90.4% |
| Avg improvement over naive | -14.8% | +16.6% |
| Oracle improvement | +31.1% | +19.7% |
| Efficiency (% of oracle captured) | 23.3% | 88.7% |
By objective
| Objective | WQT50M improvement | Random | Oracle | Efficiency |
|---|---|---|---|---|
| balanced | 22.2% | 9.3% | 25.2% | 87.7% |
| 2q_focused | 8.8% | 3.8% | 11.2% | 92.1% |
| depth_focused | 24.1% | 9.7% | 28.0% | 87.6% |
| fidelity_focused | 18.2% | 7.7% | 21.9% | 86.5% |
| time_focused | 9.6% | 4.3% | 12.3% | 89.8% |
By circuit type
| Circuit | WQT50M improvement | Random | Oracle | Efficiency |
|---|---|---|---|---|
| BV-4q | 26.6% | 7.8% | 26.6% | 100% |
| GHZ-4q | 11.9% | 2.4% | 11.9% | 100% |
| GHZ-8q | 6.5% | 1.3% | 6.5% | 100% |
| QFT-4q | 11.5% | 3.5% | 11.6% | 99.0% |
| Grover-4q | 21.3% | 8.3% | 21.5% | 98.8% |
| QFT-6q | 8.8% | 2.8% | 9.2% | 95.7% |
| QFT-8q | 1.1% | 2.4% | 7.9% | 95.3% |
| Random-6q-d20 | 34.1% | 16.2% | 37.1% | 91.9% |
| Grover-5q | 23.8% | 12.6% | 30.3% | 87.2% |
| Random-8q-d30 | 51.1% | 25.6% | 57.2% | 89.4% |
Biggest improvements
| Circuit | Backend | Objective | Naive β WQT50M | Improvement |
|---|---|---|---|---|
| Random-8q-d30 | grid | fidelity | 1231 β 467 | 62.0% |
| Random-8q-d30 | ion trap | fidelity | 273 β 118 | 56.7% |
| Random-8q-d30 | linear | depth | 1752 β 763 | 56.5% |
| Random-8q-d30 | ring | depth | 1752 β 763 | 56.5% |
| Random-8q-d30 | grid | depth | 1752 β 763 | 56.5% |
Production Gap Coverage
WQT50M addresses 6 of the 9 production gaps identified in WQT20M-Beta:
| Gap | WQT20M-Beta | WQT50M | Status |
|---|---|---|---|
| Gap 1: Real compiler outputs | Synthetic surrogate | Real Qiskit transpilation | β Fixed |
| Gap 2: Value calibration | 6.0β6.6 range | Ratio 0.999, Spearman 0.981 | β Fixed |
| Gap 3: Objective sensitivity | 0% | 65% change rate | β Fixed |
| Gap 4: Legality | 76.9% | 79.2% | β οΈ Improved |
| Gap 5: Multi-step trajectories | Single-step | 2-5 step sequences | β Fixed |
| Gap 6: Backend awareness | Topology tokens only | 60% change rate | β Fixed |
| Gap 7: Model capacity | 19M | 49.53M | β Fixed |
| Gap 8: Search improvement | +35% | 90.4% win rate | β Fixed |
| Gap 9: Generalization | 96% | Tested on 450 configs | β Fixed |
Remaining gaps for future models
- Legality accuracy: 79.2% (target >90%) β needs a dedicated classification head
- Multi-step trajectory planning: Data generated but not yet evaluated end-to-end
- Active learning feedback loop: Not yet implemented
- Production deployment as Qiskit plugin: Not yet implemented
Training Data
WQT50M is trained on real Qiskit transpilation outputs:
| Component | Details |
|---|---|
| Circuits | 50,000 (QAOA, Grover, QFT, BV, GHZ, random, Toffoli) |
| Backends | 6 (linear, ring, grid, heavy-hex, all-to-all, trapped ion) |
| Strategies | 12 real Qiskit transpilation passes |
| Objectives | 5 (balanced, 2q-focused, depth-focused, fidelity-focused, time-focused) |
| Records | 1,799,639 total |
| Backend calibration | Per-edge error rates, T1/T2, readout errors |
| Legality | 50/50 balanced with real constraints |
| Trajectories | 50,000 multi-step optimization sequences |
Coverage
15 quantum optimization domains:
| Domain | Example Actions |
|---|---|
| Quantum chemistry | UCCSD_ANSATZ, ADAPT_VQE, FROZEN_CORE |
| Graph optimization | MAXCUT_ROUND, COLOR_GRAPH, SDP_RELAX |
| Error correction | SYNDROME_MEASURE, DECODE_SURFACE, MAGIC_STATE_DISTILL |
| Quantum annealing | REVERSE_ANNEAL, MINOR_EMBED, HYBRID_SOLVE |
| Variational | QAOA_P1/P2/P3, WARM_START, ADAPTIVE_LAYER |
| Hamiltonian simulation | TROTTER_STEP, LCU_DECOMPOSE, QUBITIZE |
| Circuit optimization | CANCEL_GATES, ZX_SIMPLIFY, FUSE_ROTATIONS |
| Hardware mapping | SABRE_ROUTE, PULSE_OPTIMIZE, NOISE_AWARE |
| Neutral atom | SET_RYDBERG, PULSE_SHAPE, ADIABATIC_PASS |
| Photonic | KLM_CNOT, CLUSTER_STATE, HERALD |
| Topological | BRAID_ANYON, FIBONACCI_BRAID |
| Quantum walk | COINED_WALK, GROVER_COIN, AMPLIFY |
| State preparation | MPS_PREP, TENSOR_NETWORK, COMPRESS_STATE |
| Amplitude amplification | GROVER_ITER, QSEARCH, ITERATIVE_QPE |
| Quantum ML | IQP_KERNEL, DATA_REUPLOADING, FIDELITY_KERNEL |
Limitations
- Legality: 79.2% β Below the 90% target. The model sometimes marks illegal actions as valid.
- Not a circuit compiler β The model works on structured state representations, not raw Qiskit/TKET circuits. Plugins translate between the model's representation and real frameworks.
- Single-step evaluation β While trajectory data was generated, the model has not been evaluated on multi-step planning end-to-end.
- No active learning β The model is static; it does not improve from user compilations.
License
Apache 2.0
Citation
@misc{westquant2026wqt50m,
title={WQT50M: A 50M-Parameter Transformer for Quantum Representation Scheduling},
author={Vesterlund, David},
year={2026},
url={https://huggingface.co/WestQuantStudio/WQT50M},
note={WestQuant Open Source Project}
}
- Downloads last month
- 5