SPX-CD-Flash / README.md
Surd-AI's picture
Initial model release
9d648e6
|
Raw History Blame Contribute Delete
4.32 kB
---
library_name: peft
base_model: Qwen/Qwen3.6-35B-A3B
base_model_relation: adapter
license: apache-2.0
language:
- en
- zh
tags:
- lora
- decision-model
- multimodal
- multi-select
---
# SPX-CD-Flash
[CMDB-1500 leaderboard](https://surdai.com/zh/blog/cmdb-1500#results) · [Release overview](https://surdai.com/zh/blog/simplex-cd-release) · [Dataset](https://huggingface.co/datasets/SurdAI/CMDB-1500) · [SPX-CD-Omni](https://huggingface.co/SurdAI/SPX-CD-Omni)
**SPX-CD-Flash (Simplex Calibrated Decision Model)** by **SurdAI** is a LoRA adapter for **Qwen/Qwen3.6-35B-A3B** (35B-A3B mixture-of-experts). It supports text and image decisions: single-choice, multi-select, binary judgments, and ordinal ratings. The runner returns candidate probabilities without generating a reasoning trace.
LoRA rank/alpha **64/128** · Adapter **292.6 MiB** · Context **65,536 tokens** · Up to **1,024 candidates**.
Download the base model separately; this repository provides the LoRA adapter and decision runner.
## Download
```bash
pip install huggingface_hub
hf download SurdAI/SPX-CD-Flash --local-dir SPX-CD-Flash
cd SPX-CD-Flash
```
## Run inference
Use separate environments for the two backends. Install CUDA-enabled PyTorch for Transformers; install the vLLM requirements in a clean environment.
**Transformers, BF16 base:**
```bash
pip install -r requirements-transformers.txt
python infer.py --backend transformers --adapter . \
--input examples/text.json --output predictions.jsonl
```
**vLLM, compatible 4-bit base:**
```bash
pip install -r requirements-vllm.txt
python infer.py --backend vllm --base /path/to/compatible-4bit-base \
--adapter . --effort 2 --input examples/text.json --output predictions.jsonl
```
Use a compatible 4-bit quantization of **Qwen/Qwen3.6-35B-A3B** with its matching processor. The inference backend must support both that quantization format and this LoRA adapter; 4-bit formats are not interchangeable. The exact evaluated configuration is recorded in [EVALUATION.md](EVALUATION.md).
For images or multi-select, use `examples/image.json` or `examples/multi_select.json`. See these files for the input format and `python infer.py --help` for options. Image paths are relative to the input file, or to `--image-root` when specified.
## Results
### CMDB-1500
Completed CMDB-1500 evaluation, **effort 2**, **4-bit base**, vLLM:
| Scope | Questions | Accuracy |
| --- | ---: | ---: |
| Text | 1,200 | 77.58% |
| Images | 300 | 87.33% |
| All | 1,500 | 79.53% |
All 1,500 questions produced valid predictions. Evaluation settings are in [EVALUATION.md](EVALUATION.md); machine-readable results are in [evaluation_results.json](evaluation_results.json).
### JevBench and Decision Index
The [SurdAI release page](https://surdai.com/zh/blog/simplex-cd-release#cd-latest-results) reports the following results:
| Effort | JevBench Public (231) | JevBench Hard (111) | Decision Index 0.2.1 ↑ |
| --- | ---: | ---: | ---: |
| 1 | 88.31% (204/231) | 77.48% (86/111) | 54.65 |
| 2 | 89.18% (206/231) | 77.48% (86/111) | — |
Hard 111 is a subset of Public 231. Decision Index uses its own score scale. These results use the serving configurations described on the linked release page.
## Settings
- `--max-context`: default **65,536 total tokens**. Each scoring pass reserves one output token, leaving at most 65,535 prompt tokens, including image tokens, candidates, and selected-label prefixes. Reduce this setting if needed for available memory.
- `--effort 1–5`: average distinct candidate orderings, capped at two for binary questions. Default 1; the CMDB results use 2.
- `--temperature`: candidate-softmax temperature, default 1.0.
- `--prompt-format`: `cmdb` for text and image decisions; `open-format` for text single-choice tasks.
- Multi-select uses sequential labels and STOP. Images: up to five per request, with a processor pixel budget of 524,288 per image. Inputs exceeding the context limit are rejected.
## License and integrity
Apache 2.0 for this adapter and included code. The base model is distributed separately under its own [model card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B). Verify release files with [checksums.json](checksums.json). Release identity and original weight SHA-256 are recorded in [release.json](release.json).