#!/usr/bin/env python3 """Run the Piko-9b custom regression suite. Every check is deterministic and executed in Python. There is no judge model anywhere in this suite, so results are reproducible and carry no LLM-grader bias. The trade-off is recorded in README.md: string- and structure-based checks can mark a correct-but-differently-phrased answer as a failure. python evaluation/custom_suite/run_custom_eval.py \ --model --quantization 4bit \ --output evaluation/results/custom_suite_