--- license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3.5-4B/blob/main/LICENSE library_name: transformers pipeline_tag: image-text-to-text base_model: Qwen/Qwen3.5-4B tags: - pyrodash - collaborative-decoding - llm-offload - qwen3.5 - grpo - reinforcement-learning language: - en - zh --- # PyroDash-4B-GRPO-Lambda-0.6 --- This repository hosts **PyroDash-4B-GRPO-Lambda-0.6** — the **GRPO Stage 3** checkpoint with efficiency penalty **λ = 0.6** (cost-oriented), fine-tuned from [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) via [PyroDash-4B-SFT](https://huggingface.co/pyromind/PyroDash-4B-SFT). Companion models: [SFT](https://huggingface.co/pyromind/PyroDash-4B-SFT) · [GRPO λ=0.05](https://huggingface.co/pyromind/PyroDash-4B-GRPO-Lambda-0.05) ---
Inference Architecture We propose **PyroDash**, a token-level dynamic reasoning paradigm for collaborative inference between small and large language models. PyroDash enables the small model to autonomously emit the control token `<|llm_offload|>` during autoregressive streaming decoding; the collaboration engine then dynamically offloads the local reasoning chain to a large model based on this control signal. This approach requires neither an additional router model nor retraining of the large model, and is naturally compatible with closed-source LLM services.
During training, PyroDash follows a three-stage progressive optimization pipeline: (1) train the control-token embedding layer so the small model acquires basic offloading expressiveness; (2) cold-start the offload capability to establish a collaboration pattern between the small and large models; and (3) **apply GRPO** (this checkpoint, **λ=0.6**) that jointly optimizes the dynamic offloading policy with a task-accuracy reward and a large-model call-cost penalty, achieving an adaptive balance between reasoning quality and compute cost.

Three-stage progressive training pipeline

## Model Details | Item | Value | | --- | --- | | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) | | Init | [PyroDash-4B-SFT](https://huggingface.co/pyromind/PyroDash-4B-SFT) | | Stage | (3) GRPO · **λ = 0.6** (cost-oriented) | | Control token | <\|llm_offload\|> | | SFT dataset | [EasyHard-24K](https://huggingface.co/datasets/pyromind/easyhard-24k) | | GRPO dataset | [BytedTsinghua-SIA/DAPO-Math-17k](https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k) | | Expert LLM (train/eval) | GLM-5.2-FP8 (frozen) | | Precision | bfloat16 | ## Quick Start ### 1. Setup ```bash git clone https://github.com/PyroMind-Dynamics/pyroDash.git cd pyroDash pip install -r requirements.txt ``` ### 2. Run evaluation (`evaluation/math_eval.sh`) Edit placeholders in [`evaluation/math_eval.sh`](https://github.com/PyroMind-Dynamics/pyroDash/blob/main/evaluation/math_eval.sh), then: ```bash bash evaluation/math_eval.sh ``` The script (1) starts a local **vLLM** server for the small model on port `8001`, (2) runs `math_eval.py`, and (3) stops vLLM on exit. #### Parameters | Variable / flag | Meaning | Example | |-----------------|---------|---------| | `MODEL` | Local merged model path (vLLM serve + tokenizer) | `/path/to/your/merged_model` | | `--glm-base-url` | OpenAI-compatible API for the large/relay model | `http://your-glm-host:8000/v1` | | `--glm-api-key` | API key for that endpoint | `your-glm-api-key` | | `--glm-model` | Served model name on the GLM side | `your-glm-model` | | `--output-dir` | Per-dataset JSON output directory | `./results_500` | | `--datasets` | Benchmarks (space-separated) | `gsm8k minerva olympiad aime2024 aime2025` | Tokenizer must include the special token `<|llm_offload|>`. ## Results

Cost–Accuracy Pareto

| Method | Avg. Acc. (%) | LLM Token Ratio (%) | Avg. LLM Calls | Cost ($) | | --- | ---: | ---: | ---: | ---: | | Qwen3.5-4B | 28.36 | 0.00 | 0.000 | 2.26 | | Qwen3.5-4B (+SFT) | 46.25 | 0.00 | 0.000 | 1.32 | | RouteLLM (~75% GLM-5.2-FP8) | 52.74 | 77.37 | 0.808 | 44.62 | | GlimpRouter (τ=0.9) | 54.20 | 75.11 | 1.20 | 31.61 | | PyroDash (λ=0.1) | 55.29 | 8.19 | 0.058 | 4.71 | | **PyroDash (λ=0.6) ← this** | **54.55** | **1.90** | **0.012** | **1.78** | | PyroDash (λ=0.05) | 64.04 | 95.34 | 0.975 | 39.29 | | GLM-5.2-FP8 | 57.68 | 100.00 | 1.000 | 49.36 | λ=0.6 is the **cost** operating point (~**1.9%** LLM tokens; ~$1.78 vs $49.36 for GLM-only). For peak accuracy, use [λ=0.05](https://huggingface.co/pyromind/PyroDash-4B-GRPO-Lambda-0.05). ## Resources | Resource | Link | | --- | --- | | Project website | [PyroMind-Dynamics.github.io/pyroDash](https://PyroMind-Dynamics.github.io/pyroDash/) | | Code | [github.com/PyroMind-Dynamics/pyroDash](https://github.com/PyroMind-Dynamics/pyroDash) | | SFT dataset (EasyHard-24K) | [huggingface.co/datasets/pyromind/easyhard-24k](https://huggingface.co/datasets/pyromind/easyhard-24k) | | GRPO dataset (DAPO-Math-17k) | [huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k](https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k) | | Hugging Face org | [huggingface.co/pyromind](https://huggingface.co/pyromind) | ## Citation ```bibtex @misc{pyrodash2026, title = {PyroDash: Cost-Efficient Token-Level Small-Large Model Collaborative Inference}, author = {{PyroMind Dynamics}}, year = {2026}, note = {Preprint} } @misc{pyromind2026easyhard24k, title = {{EasyHard-24K} v0.02}, author = {{PyroMind Dynamics}}, year = {2026}, howpublished = {\url{https://huggingface.co/datasets/pyromind/easyhard-24k}} } ``` GRPO training data: [BytedTsinghua-SIA/DAPO-Math-17k](https://huggingface.co/datasets/BytedTsinghua-SIA/DAPO-Math-17k). ## License Apache 2.0 (derived from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)).