File size: 4,274 Bytes
04f53c5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---
library_name: peft
base_model: Qwen/Qwen3.8-27B
base_model_relation: adapter
license: apache-2.0
language:
- en
- zh
tags:
- lora
- decision-model
- multimodal
- multi-select
---

# SPX-CD-Pro

[CMDB-1500 leaderboard](https://surdai.com/zh/blog/cmdb-1500#results) · [Release overview](https://surdai.com/zh/blog/simplex-cd-release) · [Dataset](https://huggingface.co/datasets/SurdAI/CMDB-1500) · [SPX-CD-Omni](https://huggingface.co/SurdAI/SPX-CD-Omni)

**SPX-CD-Pro (Simplex Calibrated Decision Model)** by **SurdAI** is a LoRA adapter for **Qwen/Qwen3.8-27B** (27B dense). It supports text and image decisions: single-choice, multi-select, binary judgments, and ordinal ratings. The runner returns candidate probabilities without generating a reasoning trace.

LoRA rank/alpha **32/64** · Adapter **890.7 MiB** · Context **65,536 tokens** · Up to **1,024 candidates**.

Download the base model separately; this repository provides the LoRA adapter and decision runner.

## Download

```bash
pip install huggingface_hub
hf download SurdAI/SPX-CD-Pro --local-dir SPX-CD-Pro
cd SPX-CD-Pro
```

## Run inference

Use separate environments for the two backends. Install CUDA-enabled PyTorch for Transformers; install the vLLM requirements in a clean environment.

**Transformers, BF16 base:**

```bash
pip install -r requirements-transformers.txt
python infer.py --backend transformers --adapter . \
  --input examples/text.json --output predictions.jsonl
```

**vLLM, compatible 4-bit base:**

```bash
pip install -r requirements-vllm.txt
python infer.py --backend vllm --base /path/to/compatible-4bit-base \
  --adapter . --effort 2 --input examples/text.json --output predictions.jsonl
```

Use a compatible 4-bit quantization of **Qwen/Qwen3.8-27B** with its matching processor. The inference backend must support both that quantization format and this LoRA adapter; 4-bit formats are not interchangeable. The exact evaluated configuration is recorded in [EVALUATION.md](EVALUATION.md).

For images or multi-select, use `examples/image.json` or `examples/multi_select.json`. See these files for the input format and `python infer.py --help` for options. Image paths are relative to the input file, or to `--image-root` when specified.

## Results

### CMDB-1500

Completed CMDB-1500 evaluation, **effort 2**, **4-bit base**, vLLM:

| Scope | Questions | Accuracy |
| --- | ---: | ---: |
| Text | 1,200 | 79.33% |
| Images | 300 | 88.67% |
| All | 1,500 | 81.20% |

All 1,500 questions produced valid predictions. Evaluation settings are in [EVALUATION.md](EVALUATION.md); machine-readable results are in [evaluation_results.json](evaluation_results.json).

### JevBench and Decision Index

The [SurdAI release page](https://surdai.com/zh/blog/simplex-cd-release#cd-latest-results) reports the following results:

| Effort | JevBench Public (231) | JevBench Hard (111) | Decision Index 0.2.1 ↑ |
| --- | ---: | ---: | ---: |
| 1 | 89.61% (207/231) | 79.28% (88/111) | 58.30 |
| 2 | 90.04% (208/231) | 80.18% (89/111) | — |

Hard 111 is a subset of Public 231. Decision Index uses its own score scale. These results use the serving configurations described on the linked release page.

## Settings

- `--max-context`: default **65,536 total tokens**. Each scoring pass reserves one output token, leaving at most 65,535 prompt tokens, including image tokens, candidates, and selected-label prefixes. Reduce this setting if needed for available memory.
- `--effort 1–5`: average distinct candidate orderings, capped at two for binary questions. Default 1; the CMDB results use 2.
- `--temperature`: candidate-softmax temperature, default 1.0.
- `--prompt-format`: `cmdb` for text and image decisions; `open-format` for text single-choice tasks.
- Multi-select uses sequential labels and STOP. Images: up to five per request, with a processor pixel budget of 524,288 per image. Inputs exceeding the context limit are rejected.

## License and integrity

Apache 2.0 for this adapter and included code. The base model is distributed separately under its own [model card](https://huggingface.co/Qwen/Qwen3.8-27B). Verify release files with [checksums.json](checksums.json). Release identity and original weight SHA-256 are recorded in [release.json](release.json).