sofhiaazzhr's picture
[NOTICKET] docs(knowledge-parsing): module README + refresh the parsing/seam statuses
f0cf168
|
Raw
History Blame
3.21 kB

knowledge_parsing

Document parsing β€” the first half of the knowledge pipeline. PDF/DOCX in, a versioned ParsedDocument out.

The second half (src/knowledge_extraction/) consumes that artifact and never the parser itself, which is what keeps MinerU swappable: a Tesseract or Azure Document Intelligence path emits the same artifact and extraction never learns which one ran.

Additive and flag-gated. The existing unstructured path (src/knowledge/, Tesseract OCR β†’ chunk β†’ pgvector) is untouched and keeps running as-is.

Files

  • contracts.py β€” the seam. Chunk, ParsedDocument, and the declared Mention / TermRecord handoff shapes. The one file both halves must agree on
  • config.py β€” all settings, and the single place the GPU backend flips
  • parse.py β€” MinerU wrapper with a content-addressed parse cache
  • normalize.py β€” MinerU's flat item list β†’ meaningful chunks
  • render.py β€” LaTeX/HTML β†’ readable prose for Chunk.text
  • checks.py β€” quality warnings, including silent-corruption detection
  • manifest.py / report.py β€” per-run record and throughput extrapolation
  • run.py β€” batch CLI: one folder in, artifacts + manifest out

Use

python -m src.knowledge_parsing.run --input data/knowledge_docs/
python -m src.knowledge_parsing.report --scale 6800

MinerU is an optional extra, not a main dependency β€” the agent service never parses documents at request time, so the deployed Space does not ship it:

uv sync --extra knowledge-parsing

Importing this package does not import MinerU or torch; only parse.py does, at call time.

Three things that are easy to get wrong

Chunk.text must stay verbatim. Extraction's guardrail locates quoted spans literally inside it. Reflowing or normalising the text makes the lookup fail and the field goes silently null β€” which presents as a bad model, not as a parser bug.

page_idx is 0-based, exactly as MinerU reports it, with no conversion anywhere in the pipeline. Converting to 1-based is the UI's job, done once at display time, so the artifact always matches the raw MinerU output kept beside it in the cache.

Markup never goes in text. The term filter is an NER model reading prose; MinerU writes formulas character-spaced (P u r c h a s i n g ~ c o s t s) and tables as HTML, and neither yields a single mention. Measured on one document and gold set, changing only the parse: raw markup scored recall 0.7561 against 0.8537 for plain text; rendering it back recovered 0.8293. The markup is preserved in Chunk.latex and Chunk.table_html, because the formula branch needs that form.

Output layout

data/knowledge_cache/parse/<hash>-<config>-<version>/   untouched MinerU output
data/knowledge_runs/<run_id>/
β”œβ”€β”€ manifest.json     backend, versions, timings, quality warnings
β”œβ”€β”€ failures.jsonl    documents that failed, with traces
└── <doc>/chunks.json the artifact

Both are gitignored β€” they are data, not code.

See KNOWLEDGE_PIPELINE_TODO.md (root) for the task breakdown and the seam discussion, and KNOWLEDGE_PIPELINE_CALIBRATION.md Β§4 for the chunking constants this module follows.