feat: add benchmark.py (Qwen3 vs baseline) and multi-model eval script 922932a Surya-sj commited on Apr 26