Attributed Decision Model β decider (stage 1b)
Demo: Ask your contract, where every answer is shown with the quotes it was made from Β· Companion model: evidence model
The second half of a two-model pipeline whose decisions are faithful by construction:
- an evidence model (ModernBERT-base token classifier, 8k context) quotes verbatim spans of a document that bear on a question;
- this decider (Laya, ModernBERT-large with an option-scoring head) answers the question from those quotes only β it never sees the rest of the document.
Because the decider's input is nothing but the quotes, every answer comes with the exact text it was based on, and an answer cannot rest on text that was not shown. If the quotes do not settle the question the decider is trained to say so ("insufficient" / "absent"); if the evidence model quotes nothing, the pipeline answers that way without calling the decider.
Fine-tuned from convaiinnovations/laya (Apache-2.0) in two stages: stage 1 on quoted-evidence decisions (Tier A, below), stage 1b adding free typed-decision data so Laya's general choice / score / yes-no abilities are kept and improved.
Usage
from rl_agent_api import RLAgent # shipped in this repository
agent = RLAgent("<this repo>", device="cuda")
quotes = "\n\n".join(spans) # the evidence model's spans, in document order
q = {"type": "choice",
"instructions": "Based only on the quoted text, is this statement supported, contradicted, or not settled? "
"Statement: The receiving party may share confidential information with its employees.",
"criteria": {"supported": "the quoted text states or implies this",
"contradicted": "the quoted text states the opposite",
"insufficient": "the quoted text does not settle it"}}
ans = agent.system_one(quotes, {"q": q})["answers"]["q"] # {"choice": ..., "probabilities": {...}, ...}
Inputs: at most 1,024 tokens, of which at most 384 for the question and options. Pack quotes by evidence score until the budget is full, then put them in document order (this is how all numbers below were produced).
Evaluation
All end-to-end numbers use the evidence model at threshold 0.5. "Gold evidence" replaces the evidence model's quotes with the annotated evidence and shows the decider's ceiling.
| Benchmark | Laya + v2 | this decider + evidence model | this decider, gold evidence |
|---|---|---|---|
| ContractNLI test (123 contracts Γ 17 hypotheses), macro-F1 | 0.464 | 0.826 (ev-s1c2) / 0.828 (ev-s1) | 0.887 |
| Evidence Inference test (never trained on), macro-F1 | 0.555 | 0.589 (ev-s1c2) / 0.596 (ev-s1) | 0.748 |
Forgetting checks (no quotes involved; the decider reads the whole input):
| Laya | this decider | |
|---|---|---|
| typed-decisions test, accuracy | 0.360 | 0.734 |
| AG News (1,000), accuracy | 0.917 | 0.928 |
| Banking77 (1,000), accuracy | 0.385 | 0.656 |
Calibration: temperatures were refitted on held-out dev data (per question type and number of options); on 46.6k dev decisions, NLL 0.591 β 0.409 and ECE 0.077 β 0.017. End to end with ev-s1c2, ECE 0.117 β 0.094 (ContractNLI) and 0.222 β 0.093 (Evidence Inference); accuracy is unchanged.
Training data
Stage 1 (Tier A): decisions with gold evidence, converted to span sets (gold; gold + distractor sentences; distractors only β insufficient; quotes from a document that does not decide the question β insufficient; whole short documents). Sources: ContractNLI, CUAD, MAUD, BoolQ, MultiRC, Movie Rationales, SciFact, FEVER, HotpotQA, Natural Questions, e-SNLI, VitaminC, IIRC, MuSiQue, QASPER, MASH-QA, 2WikiMultiHopQA, FEVEROUS. Their test splits and Evidence Inference are not used for training.
Stage 1b: tasksource-jev-typed-decisions, procedural-typed-decisions, jevhome-decisions, jev-bench (data configs), jev-distill-corpus-v3, LocalLLaMA typed-decisions (train), SWAG β rows overlapping any held-out or Tier A dataset removed.
Licence
CC BY-NC 4.0: research and non-commercial use only. The base model is Apache-2.0, but the training data includes
non-commercial sources: jevhome-decisions (CC BY-NC 4.0), about 11% of the sampled tasksource-jev rows (CC BY-NC /
research-only upstream sets), jev-bench yelp5 (Yelp Dataset License), MultiRC (CogComp Research and Academic Use
License) and Movie Rationales (no licence).
- Permissive: ContractNLI, CUAD, MAUD, SciFact, QASPER, MuSiQue (CC BY 4.0); 2WikiMultiHopQA, MASH-QA, procedural-typed-decisions, jev-distill-corpus-v3, LocalLLaMA typed-decisions (Apache-2.0); e-SNLI explanations (MIT).
- Share-alike: BoolQ, Natural Questions, FEVER, FEVEROUS, VitaminC (CC BY-SA 3.0); HotpotQA, SNLI (CC BY-SA 4.0).
- Licence not stated or unclear: IIRC (Hugging Face mirror), SWAG (card says unknown; captions from LSMDC / ActivityNet), about 30% of the sampled tasksource-jev rows.
- Model-generated labels: jev-distill-corpus-v3 (
yuri_v3, distilled from Jev 1.13 via OpenRouter), jevhome and LocalLLaMA typed-decisions (Qwen-class teachers).
Source-by-source details: LICENSE_AUDIT.md in the project repository.
Limitations
- The evidence model is the bottleneck. With gold evidence the decider reaches 0.887 (ContractNLI) and 0.748 (Evidence Inference); with predicted evidence 0.828 and 0.596. About a third of Evidence Inference items get no gold evidence among the quotes at all.
- Extra quoted text hurts. Quoting more to raise recall (lower thresholds, whole sentences, filling the budget) raises gold recall but does not raise accuracy yet; keep the default threshold.
- Contradiction is the weakest class (recall 0.86 on ContractNLI with gold evidence).
- English only; tested on contracts, biomedical abstracts and one set of medical consultation transcripts.
- The medical consultation set (390 questions) was used to choose thresholds and checkpoints; its numbers are validation numbers, not test numbers.
Model tree for ricardocs99/attributed-decision-decider
Base model
convaiinnovations/laya