--- language: - en - ko tags: - speculative-decoding - solar license: other license_name: solar-open2 license_link: LICENSE --- # Solar-Open2 250B + Speculative Draft (combined, 500K draft) This repository bundles **Upstage Solar Open2 250B-A15B** (original weights at the repository root, unchanged — see `LICENSE` and the tech report) together with our **block-draft speculative-decoding draft model** in `draft/`. - **Root**: Solar Open2 250B-A15B — 48 layers (12 NoPE GQA + 36 linear-attention), 320-expert MoE top-8, vocab 196,608. Original Upstage distribution files. - **`draft/`**: 5-layer draft (core 0.90B, total 2.51B with frozen target embedding/LM head), conditioned on five taps of the target's hidden states (layers 3/15/27/39/47). Proposes up to γ=7 tokens per step; the target verifies the block in one forward. ## Draft results (500K-scale, 2 epochs) | Metric | Value | |---|---| | val accept_len (γ=7) | 3.431 | | Serving accept_len γ=3 / γ=7 | 2.75 / 3.54 | | gsm8k (5-shot greedy, spec on) | 0.930–0.940 (lossless band) | Training: 500K-scale multi-domain collection (415K-prompt manifest), **~1.2B tokens** through the target for hidden-state distillation. 2 epochs, chunked KL, 2048 stratified anchors per batch. ## Usage Serve the root model as usual for non-speculative inference. For speculative decoding, point your serving stack's block-draft speculative option at the `draft/` folder (block size 4–7 recommended, draft unquantized). Requires a serving build with block-draft speculative decoding support for this draft format. ## License Root model weights follow the Upstage Solar Open2 license (`LICENSE`). The draft contains the target's frozen embedding/LM-head weights and follows the same terms.