Papers
arxiv:2607.15314

Cura 1T: Specialized Model for Agentic Healthcare

Published on Jul 15
ยท Submitted by
taesiri
on Jul 20
Authors:
,
,
,
,
,
,
,
,

Abstract

Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.

Community

Paper author
โ€ข
edited 2 days ago

Cura 1T leads frontier models on 5 of 6 hardest healthcare benchmarks:

  • HealthBench Hard: 36.8 (GPT-5.5: 31.5)
  • HealthBench Professional: 66.2 (Claude Fable 5: 66.0)
  • MedXpertQA-Text: 60.0 (GPT-5.5: 59.6)
  • MedXpertQA-Multimodalt: 72.2 (GPT-5.5: 77.1)
  • AgentClinic: 79.6 (Claude Opus 4.8: 79.4)
  • MedAgentBench-v2: 94.0 (Claude Opus 4.8: 93.7)

main_comparison

Paper author

How we trained it: RSI (recursive self-improvement).

Each iteration, a training agent plans a target capability, trains the model, evaluates the graded benchmark trajectories, and data agent synthesizes the next data mixture from the failure modes it finds.

Humans gate every keep-or-revert decision. Reverted rounds stay in the record. One raised headline scores while quietly damaging a held-out subset, so we threw it out. The kept rounds add up: +14.6 on HealthBench Hard, +15.9 on HealthBench Professional, +9.3 on MedAgentBench.

Paper author

@librarian-bot recommend

ยท

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.15314
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.15314 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.15314 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.15314 in a Space README.md to link it from this page.

Collections including this paper 1