IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
Abstract
Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an invalid value after stating the correct one. We introduce IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes. Cases are evaluated at four stages: safety, action and tool use, response adequacy, and advisory quality. Tool use and most safety checks are deterministic. A narrow resolver handles only ambiguous confirmation-before-write cases, while a separate LLM judge evaluates semantic response adequacy. We run every case three times and report strict pass^3, which requires success on all trials. Across the eleven evaluated models, strict reliability ranges from 43.7% to 58.2%, whereas at-least-once success ranges from 60% to 74%. This gap shows that at-least-once success can overstate dependable banking behavior. The case-level diagnostics also distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request. We release the cases, mock environment, and evaluation harness.
Community
How do we tell whether a banking assistant handled a customer’s request correctly, rather than just producing a plausible final answer?
We’re sharing IndicBankBench: 799 synthetic retail-banking cases across 20 primary axes, evaluated in a mock banking environment. Cases test whether an assistant identifies the right account, uses fresh
tool evidence, obtains confirmation before an action, or declines an unsupported request.
The evaluator checks safety, action/tool use, response adequacy, and advisory quality in order, making it possible to locate where a case failed. We run each case three times to distinguish occasional
success from reliable behavior. We’d welcome feedback on the case design and grading rules.
Paper (https://arxiv.org/abs/2609.29167)
Code (https://github.com/npci/IndicBankBench)
Dataset (https://huggingface.co/datasets/NPCI/IndicBankBench)
Get this paper in your agent:
hf papers read 2609.29167 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper