| license: cc-by-nc-4.0 | |
| language: | |
| - en | |
| task_categories: | |
| - question-answering | |
| - text-generation | |
| tags: | |
| - medical | |
| - clinical | |
| - healthcare | |
| - llm | |
| - sft | |
| size_categories: | |
| - 100K<n<1M | |
| pretty_name: Fully Open Meditron Corpus | |
| configs: | |
| - config_name: default | |
| data_files: | |
| - split: train | |
| path: data/*/train-*.parquet | |
| - config_name: curated_qa | |
| data_files: | |
| - split: train | |
| path: data/curated_qa/train-*.parquet | |
| - config_name: synthetic_curated_qa | |
| data_files: | |
| - split: train | |
| path: data/synthetic_curated_qa/train-*.parquet | |
| - config_name: guidelines_qa | |
| data_files: | |
| - split: train | |
| path: data/guidelines_qa/train-*.parquet | |
| - config_name: synthetic_moove | |
| data_files: | |
| - split: train | |
| path: data/synthetic_moove/train-*.parquet | |
| # Fully Open Meditron Corpus | |
| A clinician-vetted training corpus for medical large language models, accompanying the paper *Fully Open Meditron: An Auditable Pipeline for Clinical LLMs* (anonymous submission to NeurIPS 2026 Evaluations & Datasets Track). | |
| The corpus combines eight aggregated public medical QA datasets with three clinician-vetted synthetic components, totaling approximately 601k examples (~150M tokens). It is designed to support supervised fine-tuning of large language models for clinical decision support and medical question answering, with full transparency over data provenance, generation prompts, and decontamination. | |
| ## Quick Start | |
| ```python | |
| from datasets import load_dataset | |
| # Load the full merged corpus (default) | |
| ds = load_dataset("meditron-fo-anon/fully-open-meditron") | |
| # Load a single component (e.g. for ablations) | |
| ds = load_dataset("meditron-fo-anon/fully-open-meditron", "synthetic_moove") | |
| ``` | |
| ## Components | |
| | Config | Examples | Description | | |
| |---|---|---| | |
| | `curated_qa` | 216,546 | Aggregated public medical QA training splits (MedQA, MedMCQA, PubMedQA, MedExpQA, HealthSearchQA, LiveQA, AfriMed-QA v1/v2), normalized into (system, user, assistant) conversational format. 173 items removed by system-wide decontamination. | | |
| | `synthetic_curated_qa` | 214,654 | Novel exam-style QA generated by gpt-oss-120b, seeded from the curated pool, stratified by question type with continuous answer-position monitoring to prevent label bias. | | |
| | `guidelines_qa` | 145,681 | QA grounded in 46,469 clinical practice guidelines from 16 global institutions. | | |
| | `synthetic_moove` | 24,465 | Open-ended clinical vignette prompts seeded from an expert-written vignette pool, designed to elicit complex diagnostic reasoning. | | |
| | **Total** | **601,346** | | | |
| The `default` config concatenates all four. | |
| ## Schema | |
| | Field | Type | Description | | |
| |---|---|---| | |
| | `id` | string | Unique identifier | | |
| | `messages` | list of `{role, content}` | Conversation in OpenAI-style format. Roles: `system`, `user`, `assistant`. | | |
| | `source_component` | string | One of `curated_qa`, `synthetic_curated_qa`, `guidelines_qa`, `synthetic_moove`. | | |
| | `is_synthetic` | bool | Whether the row was generated by an LLM teacher. | | |
| | `teacher_model` | string | Teacher model used for generation (`gpt-oss-120b` or null for source items). | | |
| | `source_dataset` | string | Original public dataset name (curated_qa rows only). | | |
| | `gold_label` | string | Multiple-choice gold answer letter, where applicable. | | |
| | `label_text` | string | Multiple-choice gold answer text, where applicable. | | |
| | `exact_match` | bool | Whether teacher prediction matched the gold label after rejection-sampling resampling. | | |
| | `try_count` | int | Number of resampling attempts required (1–8). | | |
| ## Construction | |
| The corpus was constructed in three stages: | |
| 1. **Aggregation.** Eight public medical QA datasets were normalized into a unified conversational schema. Items that could not be unambiguously mapped were discarded. | |
| 2. **Clinician-vetted synthetic generation.** A four-physician panel reviewed three sampled outputs per few-shot generation prompt template, with disagreements resolved by panel discussion. The audit produced four structural changes to the generation pipeline: tightening overbroad constraints on "controversial" and "outdated" content; requiring explicit disease progression and geographic context; decoupling stems from answers; and excluding overly US-centric phrasing. Synthetic components were then generated by gpt-oss-120b. | |
| 3. **Hallucination mitigation.** For every multiple-choice item carrying a labeled answer, the predicted letter was extracted via dataset-specific regex and resampled independently up to 8 times at temperature 0.7 until the extracted letter matched the gold label. | |
| ## Decontamination | |
| System-wide two-stage n-gram and token-alignment decontamination was applied against all evaluation references used in the accompanying paper: MedQA, MedMCQA, PubMedQA, MedXpertQA test splits, an open-ended clinical evaluation split, HealthBench Hard, MMLU-Pro, IFEval, and ARC-Challenge. Decontamination is syntactic rather than semantic. | |
| ## Licensing | |
| The synthetic components (`synthetic_curated_qa`, `guidelines_qa`, `synthetic_moove`) are released under CC BY-NC 4.0 for research use. | |
| The `curated_qa` component is a derived aggregation of publicly released datasets, each retaining its original license. Users redistributing this component should consult the original licenses of MedQA, MedMCQA, PubMedQA, MedExpQA, HealthSearchQA, LiveQA, and AfriMed-QA v1/v2. | |
| ## Limitations | |
| - Approximately 64% of items (71% of tokens) are synthetic, generated by a single teacher (gpt-oss-120b), which introduces model-specific stylistic and reasoning biases. | |
| - Clinician audit covered three sampled QA pairs per generation prompt template, bounding systematic but not item-level errors. | |
| - Decontamination is syntactic (n-gram and token alignment) rather than semantic, leaving open the possibility of paraphrased leakage from teacher generations. | |
| - Coverage of non-English clinical content is limited. | |
| - Inherits geographic and demographic biases of the source datasets (predominantly North American and European clinical contexts), partially mitigated by AfriMed-QA v1/v2. | |
| ## Intended Use | |
| For research, including reproducibility, auditing, and red-teaming of medical LLMs. **Not intended as a substitute for clinical judgment.** Models trained on this corpus should not be deployed without independent domain-specific safety evaluation. | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{anonymous2026meditron, | |
| title={Fully Open Meditron: An Auditable Pipeline for Clinical LLMs}, | |
| author={Anonymous}, | |
| booktitle={NeurIPS 2026 Evaluations and Datasets Track}, | |
| year={2026}, | |
| note={Under review} | |
| } | |
| ``` |
Xet Storage Details
- Size:
- 6.65 kB
- Xet hash:
- 897bac5cd8cab4eccd8a92b8961654073c4fa9153f49cdf0db5b6db2c06694cd
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.