Add FinCalc-NLI eval (financial arithmetic) + results for all checkpoints

#4
by anespo28 - opened

Hi Alex, thanks for openjev. I use it as the first stage of a judge cascade for banking agents, and it's excellent on lookups and faithfulness. I found one gap: claims that need arithmetic (ratios, covenant tests, net figures).

This PR adds code/eval_fincalc.py, in the style of eval_extra.py and built on your NLIScorer, unchanged. It evaluates any checkpoint on FinCalc-NLI (MIT, synthetic, labels computed in code). It also adds one line to "What's inside".

Results on the 4,000-item test split (run on HF Jobs with your code/eval.py as-is):

checkpoint FinCalc acc arithmetic slips caught wrong-entity figures caught covenant pass/breach
4B v5 63.2% 30% 33% 65%
4B v2 58.2% 34% 39% 17%
0.8B v2s-long 50.5% 12% 7% 32%
4B original 31.9% 0% 0% 0%

A LoRA on the original 4B with general-NLI replay reaches 94.4% (openjev-fincalc-4b, demo). If useful, the generator could also join your data mix for a future version. Happy to adapt anything to your conventions.

Tony Esposito

AlexWortega changed pull request status to merged

Sign up or log in to comment