SPX-CD-Flash / EVALUATION.md
Surd-AI's picture
Initial model release
9d648e6
|
Raw History Blame Contribute Delete
1.89 kB

Evaluation

CMDB-1500

Setting Value
Backend vLLM
Base precision NVFP4
Effort 2
Candidate-softmax temperature 1.0
Context 16,384 tokens
Image processing Native processor, 524,288-pixel budget per image
Coverage 1,200 text + 300 image questions; all 1,500 valid

Effort 2 averages probabilities from the original and a deterministic shuffled candidate order after mapping them back to the original candidates. Multi-select uses greedy sequential labels and STOP. Reference answers are used only for scoring.

Frozen input SHA-256: 2e2f03485bf8811570d4b3f72fe5d3b7e68c7a4dda9cc20e317eeeba34110242. Adapter SHA-256: 5688f89bc2dd8e71d376d7c3037805ae82ad918a9cfe077750a8ab4ed6a87812.

JevBench and Decision Index

The additional results in README.md are reported on the SurdAI release page. Public 231 includes Hard 111. Decision Index 0.2.1 has a separate score scale. The linked results use their respective serving configurations; they do not establish numerical equivalence with the included CLI.

Reproduction

Use the matching base model, precision, effort, and context setting when comparing results. BF16 and quantized backends are not guaranteed to produce identical predictions. The release defaults to 65,536 tokens. The reported CMDB evaluation used 16,384 tokens; use the explicit setting below to match that evaluation.

python infer.py --backend vllm --base /path/to/compatible-4bit-base \
  --adapter . --effort 2 --max-context 16384 --prompt-format cmdb \
  --input /path/to/inputs.jsonl --image-root /path/to/images \
  --output predictions.jsonl

Use the JSON schema in examples/. Test the decision logic with python -m unittest test_decision, or validate inputs with --dry-run before loading weights.