The Multi-Domain Reasoning Benchmark

Evaluating state-of-the-art LLMs against dynamic, adversarial real-world data pipelines across Support, Legal, Clinical, and Engineering domains.

Live Global Leaderboard
Score Breakdown (GPT-4o)
Active Task Environments

⚙️ Custom Support

Classification Semantic Drafting SLA Queue Engine Multi-Turn De-escalation

⚖️ Legal Review

Clause Identification Risk Flagging Redline Writing

🏥 Clinical Triage

Body System ESI Assignment Triage Note Gen

💻 Full Stack PRs

Type Taxonomy Patch Bug Hunting Rejection Rebuttals
Interpretability Replay Mode (Latest Episode)
1
2
3
4
5
6
7
8
# Episode ID: TKT-MT-8B391A -- Seed: 42

[Environment] Observation 3 yielded: Ticket TKT-3342, Priority: unknown, Queue Load: 4
[Environment] Valid Actions: [DRAFT_RESPONSE, ESCALATE, CLOSE]
[Agent_GPT4o] Action Submitted: { "action_type": "draft_response", "response_text": "I understand you are frustrated by the charge..." }
> Executing semantic similarity verification against KB articles...
> Warning: Detected prompt injection attempt in user message! Proceeding cautiously.
> Reward Emitted: +0.65 (Trajectory Bonus Applied)