SOLID-StepORLM / README.md
JR-James-0125's picture
Update README.md
896f8ff verified
|
Raw
History Blame Contribute Delete
1.79 kB
---
library_name: transformers
pipeline_tag: text-generation
base_model: Chenyu-Zhou/StepORLM-Qwen3-8B
tags:
- solid-opt
- operations-research
- mathematical-optimization
- self-distillation
- grpo
- qwen3
---
# SOLID-StepORLM
This is the checkpoint of **SOLID (Solver-Informed Self-Distillation)** built from `Chenyu-Zhou/StepORLM-Qwen3-8B` for operations-research modeling and solver-backed answer generation.
The model was trained with GRPO and solver-informed token-level KL supervision. It uses the COPT-style StepORLM response template.
## Evaluation
Each problem was sampled 64 times. `maj@64` is majority-vote accuracy; `pass@k` uses the unbiased pass-at-k estimator.
| Dataset | maj@64 | pass@1 | pass@2 | pass@4 |
|---|---:|---:|---:|---:|
| OptMATH | 31.33 | 18.25 | 24.40 | 30.28 |
| MAMO-Complex | 70.44 | 66.43 | 71.58 | 74.79 |
| InOR | 48.00 | 39.81 | 46.07 | 50.59 |
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "AIOR-Research/SOLID-StepORLM"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
```
The generated optimization code expects a compatible COPT environment for execution.
## Citation
If you use SOLID in your research, please cite:
```bibtex
@misc{zhu2026verifiedanswerssolverinformedselfdistillation,
title={Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models},
author={Rui Zhu and Minglong Cao and Chenyu Zhou and Jianghao Lin and Dongdong Ge},
year={2026},
eprint={2609.09957},
archivePrefix={arXiv},
primaryClass={math.OC},
url={https://arxiv.org/abs/2609.09957},
}
```