Solar-Open2-250B-DSpark

A DSpark speculative-decoding draft model for Solar-Open2-250B-W4A8.

Model Detail

Field Value
Method DSpark (block-diffusion draft, block rejection sampling)
Target model Solar-Open2-250B-W4A8
Block size / num_speculative_tokens 4 (exported ns4 variant)
Hidden layers 5
Attention geometry, RoPE, KV-head count, rms_norm_eps inherited from target
Auxiliary hidden-state taps layers 4, 12, 24, 36, 44 of the target — all full (softmax) attention layers; Solar-Open2-250B interleaves 12 full-attention layers among 36 linear-attention (KDA) layers, and only the full-attention layers are tapped
Training on-policy (responses regenerated by the target itself); prompt mixture = public instruction/code/multilingual datasets + a private dataset (not published)

Usage

vllm serve vessl/Solar-Open2-250B-W4A8 \
  --served-model-name solar-open2-250b \
  --tensor-parallel-size 4 --enable-expert-parallel \
  --kv-cache-dtype fp8 \
  --enable-prefix-caching \
  --speculative-config '{"method":"dspark","model":"vessl/Solar-Open2-250B-DSpark","num_speculative_tokens":4}' \
  --reasoning-parser solar_open2 --tool-call-parser solar_open2 \
  --enable-auto-tool-choice \
  --port 8000

License

Distributed under the same Upstage Solar License as the Solar Open 2 base model and its W4A8 derivative, per the license's requirements for derivative AI models (name prefixed with "Solar," "Built with Solar" attribution).

Downloads last month
27
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vessl/Solar-Open2-250B-DSpark

Finetuned
(1)
this model