Cancer Knowledge Base
Educational knowledge base on cancer types, classification, case studies, treatment regimens, monitoring, clinical trials, and guidelines.
Contents (parquet files)
| File | Rows | Description |
|---|---|---|
documents.parquet |
57 | De-identified documents: metadata + body text |
sections.parquet |
396 | Markdown headings + verbatim section content |
markdown_tables.parquet |
209 | All markdown tables (header/rows as JSON) |
markdown_links.parquet |
145 | Internal links between documents |
prognosis.parquet |
168 | Lethality/survival tables by cancer type and stage (dated, educational) |
biomarkers.parquet |
72 | Mutations, receptors, genetic syndromes |
guidelines.parquet |
142 | Guideline-grade treatment statements per cancer (NCCN/ESMO/ASCO) |
subtypes.parquet |
110 | Histologic subtypes by cancer |
drugs.parquet |
59 | Drugs, class, mechanism, toxicity |
staging.parquet |
31 | Staging systems (WHO, BCLC, Gleason, NCCN, Clark, Breslow) |
regimens.parquet |
85 | Treatment regimens per case and phase |
monitoring.parquet |
109 | Factors to monitor during chemo/targeted/surgery/RT |
diagnostics.parquet |
96 | Recommended diagnostics per case |
labs.parquet |
72 | Baseline labs (synthetic, clearly flagged) |
imaging_findings.parquet |
59 | PET-CT / imaging findings (synthetic, clearly flagged) |
consultations.parquet |
59 | Consultation plan evolution (synthetic, clearly flagged) |
case_facts.parquet |
218 | Structured case fields (diagnosis, staging, prognosis) |
evidence_levels.parquet |
101 | Evidence grading (guidelines + trials) per entity |
codes.parquet |
82 | Interop codes: ICD-10 for sites/types, HGNC for biomarkers, LOINC for labs, ICD-O-3 for breast subtypes |
regimen_drugs.parquet |
100 | Regimen-to-drug join table |
mcq_questions.parquet |
124 | Oncology multiple-choice eval questions with rationale (112 total) |
prognosis_sources.parquet |
32 | Prognosis rows linked to verified SEER references (vintage + value) |
locales.parquet |
9 | Registered target languages (BCP-47 codes, review requirement, priority) |
glossary.parquet |
421 | Code-anchored localization terminology registry: concept -> code (ICD/LOINC/HGNC/ICD-O-3) -> local term per locale |
translations.parquet |
4 | Per-document localized bodies with master_hash/content_hash provenance and review status |
RAG Benchmark
benchmark/ contains an oncology RAG evaluation suite built from this knowledge base:
retrieval_pool.parquetβ ~1,092 chunked retrieval units (sections, markdown tables, domain tables)mcq_benchmark.parquetβ 112 multiple-choice questions with golden docs, difficulty, evidence level, and citationsmcq_robustness.parquetβ 336 MedRGB-style variants (sufficiency / noise / counterfactual-robustness)mirage_format.jsonl/medrgb_format.jsonlβ interchange exports for MIRAGE-family leaderboards and MedRGB harnesses (seebenchmark/MIRAGE_INTEGRATION.md)leaderboard.jsonβ recorded system results
Evaluate your own system with scripts/eval_mcq.py (BM25 or embedding retrieval, or score your predictions).
Ground-truth validation scorecards
scorecards/ holds the results of screening LLM-synthetic oncology datasets against the KB's cited
statements (scripts/validate_synthetic.py). Each claim is tiered supported / partial / unsupported
by lexical coverage of site-gated evidence, and every item links to the matching evidence chunk.
cancerguide.all.summary.md+.scorecard.json+.claims.parquetβ CancerGUIDE (316 synthetic patient notes): 238 scorable claims, 87.0% supported, 12.6% partial, 0.4% unsupported.
Localization (l10n)
The KB is being localized with provenance β every translation keeps its citation (via the master
document), a master_hash so a change in the English source flags the locale stale, and an explicit
review status (auto β reviewed β approved).
locales.parquetβ registered languages (de, zh, es, hi + pipeline candidates)glossary.parquetβ code-anchored terminology registry: each concept is linked to its language-neutral code (ICD-10 / ICD-O-3 / LOINC / HGNC) and maps to the local clinical term, so structured rows localize for free while codes stay the join key across languages.translations.parquetβ localized document bodies (pilot:know/breast-ref.mdat 100% translatable coverage in de/zh/es/hi) withmaster_hash/content_hashandstatus.
Coverage is not clinical approval: status='auto' rows are glossary-generated drafts that must
be clinically reviewed (reviewed) before use. See l10n-moat.md and
scripts/localize_pilot.py --coverage.
Privacy
This is the de-identified public subset. The real patient case (cases/specific/ β PHI) and all
content derived from it are excluded. Documents are filtered by phi = 0.
Synthetic rows in labs, imaging_findings, and consultations are clearly flagged with
synthetic = 1 and are illustrative public-case data β they are not real patient values.
Disclaimer
Educational reference only β not medical advice. Survival figures are population averages and may
be outdated (prognosis.as_of marks their review date); consult current guidelines and clinicians for
patient decisions.