--- library_name: transformers pipeline_tag: text-generation base_model: Chenyu-Zhou/StepORLM-Qwen3-8B tags: - solid-opt - operations-research - mathematical-optimization - self-distillation - grpo - qwen3 --- # SOLID-StepORLM This is the checkpoint of **SOLID (Solver-Informed Self-Distillation)** built from `Chenyu-Zhou/StepORLM-Qwen3-8B` for operations-research modeling and solver-backed answer generation. The model was trained with GRPO and solver-informed token-level KL supervision. It uses the COPT-style StepORLM response template. ## Evaluation Each problem was sampled 64 times. `maj@64` is majority-vote accuracy; `pass@k` uses the unbiased pass-at-k estimator. | Dataset | maj@64 | pass@1 | pass@2 | pass@4 | |---|---:|---:|---:|---:| | OptMATH | 31.33 | 18.25 | 24.40 | 30.28 | | MAMO-Complex | 70.44 | 66.43 | 71.58 | 74.79 | | InOR | 48.00 | 39.81 | 46.07 | 50.59 | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "AIOR-Research/SOLID-StepORLM" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype="auto", device_map="auto", ) ``` The generated optimization code expects a compatible COPT environment for execution. ## Citation If you use SOLID in your research, please cite: ```bibtex @misc{zhu2026verifiedanswerssolverinformedselfdistillation, title={Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models}, author={Rui Zhu and Minglong Cao and Chenyu Zhou and Jianghao Lin and Dongdong Ge}, year={2026}, eprint={2609.09957}, archivePrefix={arXiv}, primaryClass={math.OC}, url={https://arxiv.org/abs/2609.09957}, } ```