--- license: apache-2.0 base_model: answerdotai/ModernBERT-base pipeline_tag: token-classification tags: - text-segmentation - semantic-chunking - rag - chunking datasets: - saeedabc/wiki727k language: - en --- # Semantic chunker (ModernBERT-base) Predicts, for each sentence, whether a topic boundary follows it. Intended as a drop-in chunker for RAG pipelines. ## What is different here - **2048-token training context.** The closest existing model trains at 1024. Median Wikipedia article is ~980 tokens, so 2048 keeps ~82% of documents in a single window instead of ~53%. - **Supervision at sentence-final tokens only.** Labelling every token puts the positive rate near 1% and the model collapses toward never predicting a split. Scoring only sentence-final positions raises it to ~12%. - **Trained on Wikipedia section structure**, not book paragraphing. ## Results Both models are ModernBERT-base token taggers sharing a tokenizer, so they are scored at identical positions: P(boundary) read at each sentence's final token. Threshold swept 0.05–0.95; the best operating point is reported. ### wiki727k test — in-distribution for this model | Model | F1 | Precision | Recall | Threshold | |---|---|---|---|---| | **this model** | 0.8156 | 0.8390 | 0.7934 | 0.70 | | `mirth/chonky_modernbert_base_1` | 0.4986 | 0.4543 | 0.5524 | 0.90 | | baseline:every-k (best k, oracle-tuned) (k=2) | 0.1984 | 0.1228 | 0.5159 | — | | baseline:base-rate random | 0.1177 | 0.1166 | 0.1187 | — | ### PubMed-RCT — out-of-distribution for both models Non-Wikipedia scientific prose with human-assigned section labels. Boundary rate is 35.2% here versus ~11.7% on Wikipedia, so scores are **not** comparable across the two tables — only across models within one table. | Model | F1 | Precision | Recall | Threshold | |---|---|---|---|---| | **this model** | 0.6932 | 0.7293 | 0.6605 | 0.55 | | `mirth/chonky_modernbert_base_1` | 0.5024 | 0.5095 | 0.4954 | 0.85 | | baseline:every-k (best k, oracle-tuned) (k=2) | 0.4034 | 0.3500 | 0.4759 | — | | baseline:base-rate random | 0.3490 | 0.3521 | 0.3460 | — | ### Retrieval — does better boundary detection actually help? Macro-average over 3 corpora (gov_report, qmsum, stackoverflow), embedder `BAAI/bge-small-en-v1.5`, identical 512-token cap on every strategy so chunk size cannot confound the comparison. | Strategy | nDCG@10 | R@1 | R@5 | R@10 | mean chunk tokens | |---|---|---|---|---|---| | `fixed-512-ovl64` | 0.8625 | 0.8254 | 0.9218 | 0.9500 | 388 | | `chonky-min256` | 0.8615 | 0.7995 | 0.9353 | 0.9635 | 274 | | **ours-min256** (this model) | 0.8583 | 0.7914 | 0.9328 | 0.9659 | 283 | | **ours** (this model) | 0.8575 | 0.7854 | 0.9321 | 0.9685 | 215 | | `sentence-8` | 0.8536 | 0.7956 | 0.9267 | 0.9536 | 182 | | `fixed-512` | 0.8531 | 0.8019 | 0.9279 | 0.9512 | 385 | | `recursive-512` | 0.8528 | 0.7903 | 0.9218 | 0.9598 | 308 | | `chonky` | 0.8498 | 0.7702 | 0.9321 | 0.9684 | 146 | Per corpus (nDCG@10): | Strategy | gov_report | qmsum | stackoverflow | |---|---|---|---| | `fixed-512-ovl64` | **0.9423** | 0.6988 | 0.9464 | | `chonky-min256` | 0.9285 | 0.7099 | 0.9461 | | **ours-min256** | 0.9225 | 0.7056 | 0.9467 | | **ours** | 0.9191 | **0.7254** | 0.9280 | | `sentence-8` | 0.9272 | 0.6865 | **0.9472** | | `fixed-512` | 0.9348 | 0.6782 | 0.9463 | | `recursive-512` | 0.9214 | 0.6910 | 0.9459 | | `chonky` | 0.9046 | 0.7118 | 0.9329 | ### Reading these numbers honestly `baseline:every-k` splits every k sentences with k chosen to maximise its own score — a model-free floor with oracle tuning in its favour. The wiki727k comparison flatters this model: Wikipedia section boundaries are its training distribution and are out-of-distribution for chonky, which trained on BookCorpus paragraphs. **The PubMed table is the fair comparison.** **Better boundaries did not produce a uniform retrieval win, and the retrieval table above should be read before adopting this.** That result is consistent with the published critique of semantic chunking, and it is reported here rather than omitted. Boundary F1 and retrieval quality are different things; this model is much better at the first and situationally better at the second. ## When this helps, and when it does not **Use it** for long documents whose topic genuinely shifts partway through — transcripts, reports, articles, manuals. That is where fixed-size splitting cuts through the middle of an idea and where boundary detection pays. **Do not use it** for corpora whose documents are already shorter than your embedder's window. Splitting a 150-token document into two 75-token fragments makes retrieval worse, not better, no matter how correct the boundary is. On such corpora fixed-size chunking is competitive, faster, and has no dependencies — use that instead. **`min_chunk_tokens` is a corpus-dependent knob, not a default.** Setting it to 256 recovered most of the short-document deficit (+0.019 nDCG@10 on stackoverflow) but cost almost exactly as much on long transcripts (−0.020 on qmsum), by merging away the boundaries that made the model useful there. Set it high for short documents, leave it at 0 for long ones, and measure on your own corpus rather than trusting either default. **Cost.** Chunking with this model is roughly one to two orders of magnitude slower than a fixed-size splitter (per-strategy wall time is recorded in `retrieval_eval.json`). On corpora where it does not win, that cost buys nothing. ## Usage Nothing to download by hand in any of these — the weights resolve from this repo on first use. ### With transformers alone The custom pipeline is registered in this repo's config, so this needs no install beyond `transformers`: ```python from transformers import pipeline chunk = pipeline("semantic-chunking", model="0xKitkat/semantic-chunker-modernbert-base", trust_remote_code=True) for c in chunk(document_text): print(c["n_tokens"], c["boundary_score"], c["text"][:80]) ``` Pass `return_text=True` for plain strings, or any of `threshold`, `max_chunk_tokens`, `min_chunk_tokens` to tune it. ### With the pip package ```bash pip install "boundary-chunker[onnx]" # CPU, no torch pip install "boundary-chunker[torch]" # GPU ``` ```python from boundary_chunker import SemanticChunker chunker = SemanticChunker(backend="onnx") # or backend="torch" for c in chunker.split(document_text): print(c.n_tokens, c.text[:80]) ``` There is a CLI too: ```bash boundary-chunk --backend onnx --file report.txt --json ``` ### Straight from this repo, no install ```bash hf download 0xKitkat/semantic-chunker-modernbert-base chunk.py --local-dir . ``` ```python from chunk import SemanticChunker chunker = SemanticChunker(threshold=0.5, max_chunk_tokens=512) ``` `max_chunk_tokens` force-splits over-long chunks at the model's lowest-confidence interior boundary rather than at an arbitrary offset. ### CPU-only, without torch An ONNX export lives in `onnx/` and needs `onnxruntime` instead of the full torch stack — roughly 40MB of dependencies rather than a CUDA install. It is self-contained, with its own config and tokenizer beside the graph, and `backend="onnx"` fetches it for you. The export is verified against the torch model on held-out wiki727k documents: identical chunk boundaries on 24/24 documents (739–65,468 chars, including multi-window ones), and max |P(boundary) delta| 8.1e-06 on padded batches. The pipeline is likewise checked against the direct API — identical boundaries on 12/12 documents — so all three entry points above give the same answer. Both checks are reproducible with `onnx_parity.py` in the training repo. ## Training | | | |---|---| | Base | `answerdotai/ModernBERT-base` (149M) | | Data | wiki727k, 582,160 train docs / 30,580,099 sentences | | Boundary rate | 11.7% | | Max length | 2048 | | Effective batch | 32 | | LR / schedule | 3e-5, linear, 6% warmup | | Epochs | 1 | | Loss | class-weighted CE (positive weight 3.0) | | Hardware | 1x RTX 4070 SUPER (12GB) | Section titles are dropped from the training text: with headings present the task partly degenerates into "a heading follows", which does not transfer to the unformatted prose a chunker sees in production. ## Limitations - English only. - Trained on encyclopedic prose; conversational transcripts, code, and tabular documents are out of distribution. - Boundary F1 is a proxy. If your goal is retrieval quality, measure retrieval.