Oncology RAG Benchmark
An oncology-specific, CC-BY-4.0 retrieval-augmented-generation (RAG) benchmark built from the Cancer Knowledge Base. Answer keys are provable from the KB's own structured data — every question carries golden docs, an evidence level, and a citation.
Files
| File | Contents |
|---|---|
retrieval_pool.parquet |
~1,100 chunked retrieval units (sections with heading weighting, markdown tables, and domain tables: codes / drugs / staging / prognosis / subtypes / biomarkers) |
mcq_benchmark.parquet |
112 multiple-choice questions with options, answer, answer_index, difficulty, evidence_level, citation, golden_docs, sources |
mcq_robustness.parquet |
336 MedRGB-style variants — sufficiency (golden docs), noise (golden + 5 irrelevant), robustness (golden + a counterfactually edited doc) |
leaderboard.json |
recorded results from eval_mcq.py --write-leaderboard |
mirage_format.jsonl |
112 questions in MIRAGE-family interchange format (context = golden docs) |
medrgb_format.jsonl |
336 robustness variants in MedRGB interchange format (incl. counterfactual_doc) |
See MIRAGE_INTEGRATION.md for how this maps onto the
MIRAGE/MedRGB evaluation ecosystem and how to submit a system there.
Scenarios
- sufficiency — answer from the golden docs only (upper-bound signal).
- noise — answer from golden + irrelevant docs (measures retrieval robustness).
- robustness — golden + one adversarially-edited doc that contradicts the answer (measures hallucination resistance).
Evaluate your system
Install the repo scripts (needs pyarrow):
# score your own predictions (JSONL: {"question_id": 1, "answer_index": 2})
python3 scripts/eval_mcq.py --benchmark hf_dataset/benchmark --predictions preds.jsonl
# baseline with real BM25 retrieval (no deps)
python3 scripts/eval_mcq.py --benchmark hf_dataset/benchmark --retriever bm25 --topk 3
# record your result on the leaderboard
python3 scripts/eval_mcq.py --retriever bm25 --topk 3 --write-leaderboard \
--system-name "my-system"
# compare leaderboard entries
python3 scripts/eval_mcq.py --compare my-system retriever:bm25@top3
Prediction JSONL format:
{"question_id": 1, "answer_index": 2}
{"question_id": 2, "answer_index": 0}
Submission
- Run
eval_mcq.pywith--write-leaderboard --system-name "<your system>". - Open a PR / issue against this dataset repo describing the system (LLM,
retriever, top-k) and attach your
preds.jsonlfor reproducibility.
Baseline
Current retriever:bm25@top3 baseline (dependency-free BM25):
- sufficiency accuracy: 33.0%
- retrieval@3 (golden doc in top-3): 49.1%
Random guessing is 25% (4-option questions), so the baseline beats chance and shows real headroom for better retrievers (embeddings) and generators.