# knowledge_parsing Document parsing — the first half of the knowledge pipeline. PDF/DOCX in, a versioned `ParsedDocument` out. The second half (`src/knowledge_extraction/`) consumes that artifact and never the parser itself, which is what keeps MinerU swappable: a Tesseract or Azure Document Intelligence path emits the same artifact and extraction never learns which one ran. Additive and flag-gated. The existing unstructured path (`src/knowledge/`, Tesseract OCR → chunk → pgvector) is untouched and keeps running as-is. ## Files - `contracts.py` — **the seam.** `Chunk`, `ParsedDocument`, and the declared `Mention` / `TermRecord` handoff shapes. The one file both halves must agree on - `config.py` — all settings, and the single place the GPU backend flips - `parse.py` — MinerU wrapper with a content-addressed parse cache - `normalize.py` — MinerU's flat item list → meaningful chunks - `render.py` — LaTeX/HTML → readable prose for `Chunk.text` - `checks.py` — quality warnings, including silent-corruption detection - `manifest.py` / `report.py` — per-run record and throughput extrapolation - `run.py` — batch CLI: one folder in, artifacts + manifest out ## Use ```bash python -m src.knowledge_parsing.run --input data/knowledge_docs/ python -m src.knowledge_parsing.report --scale 6800 ``` MinerU is an optional extra, not a main dependency — the agent service never parses documents at request time, so the deployed Space does not ship it: ```bash uv sync --extra knowledge-parsing ``` Importing this package does not import MinerU or torch; only `parse.py` does, at call time. ## Three things that are easy to get wrong **`Chunk.text` must stay verbatim.** Extraction's guardrail locates quoted spans literally inside it. Reflowing or normalising the text makes the lookup fail and the field goes silently null — which presents as a bad model, not as a parser bug. **`page_idx` is 0-based**, exactly as MinerU reports it, with no conversion anywhere in the pipeline. Converting to 1-based is the UI's job, done once at display time, so the artifact always matches the raw MinerU output kept beside it in the cache. **Markup never goes in `text`.** The term filter is an NER model reading prose; MinerU writes formulas character-spaced (`P u r c h a s i n g ~ c o s t s`) and tables as HTML, and neither yields a single mention. Measured on one document and gold set, changing only the parse: raw markup scored recall 0.7561 against 0.8537 for plain text; rendering it back recovered 0.8293. The markup is preserved in `Chunk.latex` and `Chunk.table_html`, because the formula branch needs that form. ## Output layout ``` data/knowledge_cache/parse/--/ untouched MinerU output data/knowledge_runs// ├── manifest.json backend, versions, timings, quality warnings ├── failures.jsonl documents that failed, with traces └── /chunks.json the artifact ``` Both are gitignored — they are data, not code. See `KNOWLEDGE_PIPELINE_TODO.md` (root) for the task breakdown and the seam discussion, and `KNOWLEDGE_PIPELINE_CALIBRATION.md` §4 for the chunking constants this module follows.