openjev-fincalc: a financial-arithmetic fine-tune of openjev (31% β†’ 94% on numeric claims)

#3
by anespo28 - opened

Hi Alex, thanks for openjev, it's an excellent base.

While testing it as the first stage of an LLM-as-a-judge cascade, I found one clear gap: claims that need arithmetic (e.g. "DSCR is 1.42x", "leverage breaches the 3.5x covenant"). openjev scored 31% there, mostly answering "neutral".

So I built a small extension on top of it:

  • FinCalc-NLI: a fully synthetic dataset where every label is computed in code (LTV, DSCR, interest cover, leverage, CET1, NPL, covenant tests), including the typical failure modes (inverted ratios, unit errors, wrong entity)
  • LoRA fine-tune with general-NLI replay, about 40 minutes on one A100

Results on unseen test data: financial claims 31% β†’ 94%, covenant pass/breach 100%, MNLI 90.4% β†’ 89.6% (essentially unchanged). Transfer to formulas it never saw reaches 71%, and the data is synthetic and English only.

Everything is MIT:

Would you be interested in numeric reasoning as a direction for openjev itself? Happy to share anything that helps.

AlexWortega changed discussion status to closed

Sign up or log in to comment