How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "pyromind/PyroDash-4B-SFT" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "pyromind/PyroDash-4B-SFT",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "pyromind/PyroDash-4B-SFT" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "pyromind/PyroDash-4B-SFT",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Quick Links

PyroDash-4B-SFT


This repository hosts PyroDash-4B-SFT — the offload cold-start (Stage 2) checkpoint of PyroDash, fine-tuned from Qwen/Qwen3.5-4B.

Companion models: GRPO λ=0.05 · GRPO λ=0.6


Inference Architecture

We propose PyroDash, a token-level dynamic reasoning paradigm for collaborative inference between small and large language models. PyroDash enables the small model to autonomously emit the control token <|llm_offload|> during autoregressive streaming decoding; the collaboration engine then dynamically offloads the local reasoning chain to a large model based on this control signal. This approach requires neither an additional router model nor retraining of the large model, and is naturally compatible with closed-source LLM services.

During training, PyroDash follows a three-stage progressive optimization pipeline: (1) train the control-token embedding layer so the small model acquires basic offloading expressiveness; (2) cold-start the offload capability (this checkpoint) to establish a collaboration pattern between the small and large models; and (3) apply GRPO reinforcement learning that jointly optimizes the dynamic offloading policy with a task-accuracy reward and a large-model call-cost penalty, achieving an adaptive balance between reasoning quality and compute cost.

Three-stage progressive training pipeline

Model Details

Item Value
Base model Qwen/Qwen3.5-4B
Stage (2) Offload cold-start SFT (LoRA → merged)
Control token <|llm_offload|>
SFT dataset EasyHard-24K
Expert LLM (eval) GLM-5.2-FP8
Precision bfloat16

Quick Start

1. Setup

git clone https://github.com/PyroMind-Dynamics/pyroDash.git
cd pyroDash
pip install -r requirements.txt

2. Run evaluation (evaluation/math_eval.sh)

Edit placeholders in evaluation/math_eval.sh, then:

bash evaluation/math_eval.sh

The script (1) starts a local vLLM server for the small model on port 8001, (2) runs math_eval.py, and (3) stops vLLM on exit.

Parameters

Variable / flag Meaning Example
MODEL Local merged model path (vLLM serve + tokenizer) /path/to/your/merged_model
--glm-base-url OpenAI-compatible API for the large/relay model http://your-glm-host:8000/v1
--glm-api-key API key for that endpoint your-glm-api-key
--glm-model Served model name on the GLM side your-glm-model
--output-dir Per-dataset JSON output directory ./results_500
--datasets Benchmarks (space-separated) gsm8k minerva olympiad aime2024 aime2025

Tokenizer must include the special token <|llm_offload|>.

Results

Cost–Accuracy Pareto

Method Avg. Acc. (%) LLM Token Ratio (%) Avg. LLM Calls Cost ($)
Qwen3.5-4B 28.36 0.00 0.000 2.26
Qwen3.5-4B (+SFT) ← this 46.25 0.00 0.000 1.32
RouteLLM (~75% GLM-5.2-FP8) 52.74 77.37 0.808 44.62
GlimpRouter (τ=0.9) 54.20 75.11 1.20 31.61
PyroDash (λ=0.1) 55.29 8.19 0.058 4.71
PyroDash (λ=0.6) 54.55 1.90 0.012 1.78
PyroDash (λ=0.05) 64.04 95.34 0.975 39.29
GLM-5.2-FP8 57.68 100.00 1.000 49.36

Resources

Citation

@misc{pyrodash2026,
  title        = {PyroDash: Cost-Efficient Token-Level Small-Large Model Collaborative Inference},
  author       = {{PyroMind Dynamics}},
  year         = {2026},
  note         = {Preprint}
}

@misc{pyromind2026easyhard24k,
  title        = {{EasyHard-24K} v0.02},
  author       = {{PyroMind Dynamics}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/pyromind/easyhard-24k}}
}

License

Apache 2.0 (derived from Qwen3.5-4B).

Downloads last month
33
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pyromind/PyroDash-4B-SFT

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(500)
this model