Blitz for Boltz-2
Blitz distills Boltz-2 structure generation into deterministic 8- and 16-step samplers. Both checkpoints are self-contained and include the Boltz-2 trunk, structure model, and confidence model.
Results
Results on Boltz's roughly 2,300-target benchmark:
| system | model | NFE | complex lDDT โ | RF-valid โ | DockQ โ | ligand RMSD (ร ) โ | any violation โ |
|---|---|---|---|---|---|---|---|
| AlphaFold 3 | teacher | 200 | 0.8615 | 41.5% | 0.4198 | 6.63 | 9.8% |
| AlphaFold 3 | Blitz K8 | 8 | 0.8635 | 40.5% | 0.4207 | 6.44 | 11.3% |
| AlphaFold 3 | Blitz K16 | 16 | 0.8634 | 42.3% | 0.4167 | 6.51 | 6.8% |
| Boltz-2 | teacher | 600 | 0.8499 | 98.5% | 0.3929 | 9.00 | 30.3% |
| Boltz-2 | Blitz K8 | 8 | 0.8492 | 93.7% | 0.3897 | 8.75 | 41.9% |
| Boltz-2 | Blitz K16 | 16 | 0.8533 | 94.1% | 0.3932 | 8.78 | 26.7% |
Every row uses the same five-candidate, reference-free reranker. Bold values are student point estimates that improve on the corresponding teacher. Across both systems, Blitz retains teacher-level complex accuracy with far fewer denoiser evaluations. K16 also lowers the overall violation rate for both model families. NFE counts denoiser evaluations; the Boltz-2 teacher uses three recycle updates.
Running inference
boltz predict input.yaml \
--model boltz2 \
--checkpoint /path/to/blitz-boltz2-k16.ckpt \
--blitz_policy k16 \
--recycling_steps 5 \
--diffusion_samples 5 \
--seed 1 \
--out_dir predictions
For K8, use blitz-boltz2-k8.ckpt with --blitz_policy k8.
Checksums are provided in SHA256SUMS. Training and data preparation are
documented in the Blitz repository.