# Evaluation ## CMDB-1500 | Setting | Value | | --- | --- | | Backend | vLLM | | Base precision | NVFP4 | | Effort | 2 | | Candidate-softmax temperature | 1.0 | | Context | 16,384 tokens | | Image processing | Native processor, 524,288-pixel budget per image | | Coverage | 1,200 text + 300 image questions; all 1,500 valid | Effort 2 averages probabilities from the original and a deterministic shuffled candidate order after mapping them back to the original candidates. Multi-select uses greedy sequential labels and STOP. Reference answers are used only for scoring. Frozen input SHA-256: `2e2f03485bf8811570d4b3f72fe5d3b7e68c7a4dda9cc20e317eeeba34110242`. Adapter SHA-256: `5688f89bc2dd8e71d376d7c3037805ae82ad918a9cfe077750a8ab4ed6a87812`. ## JevBench and Decision Index The additional results in [README.md](README.md) are reported on the [SurdAI release page](https://surdai.com/zh/blog/simplex-cd-release#cd-latest-results). Public 231 includes Hard 111. Decision Index 0.2.1 has a separate score scale. The linked results use their respective serving configurations; they do not establish numerical equivalence with the included CLI. ## Reproduction Use the matching base model, precision, effort, and context setting when comparing results. BF16 and quantized backends are not guaranteed to produce identical predictions. The release defaults to 65,536 tokens. The reported CMDB evaluation used 16,384 tokens; use the explicit setting below to match that evaluation. ```bash python infer.py --backend vllm --base /path/to/compatible-4bit-base \ --adapter . --effort 2 --max-context 16384 --prompt-format cmdb \ --input /path/to/inputs.jsonl --image-root /path/to/images \ --output predictions.jsonl ``` Use the JSON schema in `examples/`. Test the decision logic with `python -m unittest test_decision`, or validate inputs with `--dry-run` before loading weights.