0xKitkat's picture
Add transformers custom pipeline, auto-resolving weights, expanded usage
2ef9bc2 verified
|
Raw
History Blame Contribute Delete
8.72 kB
---
license: apache-2.0
base_model: answerdotai/ModernBERT-base
pipeline_tag: token-classification
tags:
- text-segmentation
- semantic-chunking
- rag
- chunking
datasets:
- saeedabc/wiki727k
language:
- en
---
# Semantic chunker (ModernBERT-base)
Predicts, for each sentence, whether a topic boundary follows it. Intended as a
drop-in chunker for RAG pipelines.
## What is different here
- **2048-token training context.** The closest existing model trains at 1024.
Median Wikipedia article is ~980 tokens, so 2048 keeps ~82% of documents in a
single window instead of ~53%.
- **Supervision at sentence-final tokens only.** Labelling every token puts the
positive rate near 1% and the model collapses toward never predicting a split.
Scoring only sentence-final positions raises it to ~12%.
- **Trained on Wikipedia section structure**, not book paragraphing.
## Results
Both models are ModernBERT-base token taggers sharing a tokenizer, so they are
scored at identical positions: P(boundary) read at each sentence's final token.
Threshold swept 0.05–0.95; the best operating point is reported.
### wiki727k test β€” in-distribution for this model
| Model | F1 | Precision | Recall | Threshold |
|---|---|---|---|---|
| **this model** | 0.8156 | 0.8390 | 0.7934 | 0.70 |
| `mirth/chonky_modernbert_base_1` | 0.4986 | 0.4543 | 0.5524 | 0.90 |
| baseline:every-k (best k, oracle-tuned) (k=2) | 0.1984 | 0.1228 | 0.5159 | β€” |
| baseline:base-rate random | 0.1177 | 0.1166 | 0.1187 | β€” |
### PubMed-RCT β€” out-of-distribution for both models
Non-Wikipedia scientific prose with human-assigned section labels. Boundary rate
is 35.2% here versus ~11.7% on Wikipedia, so
scores are **not** comparable across the two tables β€” only across models within
one table.
| Model | F1 | Precision | Recall | Threshold |
|---|---|---|---|---|
| **this model** | 0.6932 | 0.7293 | 0.6605 | 0.55 |
| `mirth/chonky_modernbert_base_1` | 0.5024 | 0.5095 | 0.4954 | 0.85 |
| baseline:every-k (best k, oracle-tuned) (k=2) | 0.4034 | 0.3500 | 0.4759 | β€” |
| baseline:base-rate random | 0.3490 | 0.3521 | 0.3460 | β€” |
### Retrieval β€” does better boundary detection actually help?
Macro-average over 3 corpora (gov_report, qmsum, stackoverflow), embedder `BAAI/bge-small-en-v1.5`, identical 512-token cap on every strategy so chunk size cannot confound the comparison.
| Strategy | nDCG@10 | R@1 | R@5 | R@10 | mean chunk tokens |
|---|---|---|---|---|---|
| `fixed-512-ovl64` | 0.8625 | 0.8254 | 0.9218 | 0.9500 | 388 |
| `chonky-min256` | 0.8615 | 0.7995 | 0.9353 | 0.9635 | 274 |
| **ours-min256** (this model) | 0.8583 | 0.7914 | 0.9328 | 0.9659 | 283 |
| **ours** (this model) | 0.8575 | 0.7854 | 0.9321 | 0.9685 | 215 |
| `sentence-8` | 0.8536 | 0.7956 | 0.9267 | 0.9536 | 182 |
| `fixed-512` | 0.8531 | 0.8019 | 0.9279 | 0.9512 | 385 |
| `recursive-512` | 0.8528 | 0.7903 | 0.9218 | 0.9598 | 308 |
| `chonky` | 0.8498 | 0.7702 | 0.9321 | 0.9684 | 146 |
Per corpus (nDCG@10):
| Strategy | gov_report | qmsum | stackoverflow |
|---|---|---|---|
| `fixed-512-ovl64` | **0.9423** | 0.6988 | 0.9464 |
| `chonky-min256` | 0.9285 | 0.7099 | 0.9461 |
| **ours-min256** | 0.9225 | 0.7056 | 0.9467 |
| **ours** | 0.9191 | **0.7254** | 0.9280 |
| `sentence-8` | 0.9272 | 0.6865 | **0.9472** |
| `fixed-512` | 0.9348 | 0.6782 | 0.9463 |
| `recursive-512` | 0.9214 | 0.6910 | 0.9459 |
| `chonky` | 0.9046 | 0.7118 | 0.9329 |
### Reading these numbers honestly
`baseline:every-k` splits every k sentences with k chosen to maximise its own
score β€” a model-free floor with oracle tuning in its favour. The wiki727k
comparison flatters this model: Wikipedia section boundaries are its training
distribution and are out-of-distribution for chonky, which trained on BookCorpus
paragraphs. **The PubMed table is the fair comparison.**
**Better boundaries did not produce a uniform retrieval win, and the retrieval
table above should be read before adopting this.** That result is consistent
with the published critique of semantic chunking, and it is reported here rather
than omitted. Boundary F1 and retrieval quality are different things; this model
is much better at the first and situationally better at the second.
## When this helps, and when it does not
**Use it** for long documents whose topic genuinely shifts partway through β€”
transcripts, reports, articles, manuals. That is where fixed-size splitting cuts
through the middle of an idea and where boundary detection pays.
**Do not use it** for corpora whose documents are already shorter than your
embedder's window. Splitting a 150-token document into two 75-token fragments
makes retrieval worse, not better, no matter how correct the boundary is. On
such corpora fixed-size chunking is competitive, faster, and has no
dependencies β€” use that instead.
**`min_chunk_tokens` is a corpus-dependent knob, not a default.** Setting it to
256 recovered most of the short-document deficit (+0.019 nDCG@10 on
stackoverflow) but cost almost exactly as much on long transcripts (βˆ’0.020 on
qmsum), by merging away the boundaries that made the model useful there. Set it
high for short documents, leave it at 0 for long ones, and measure on your own
corpus rather than trusting either default.
**Cost.** Chunking with this model is roughly one to two orders of magnitude
slower than a fixed-size splitter (per-strategy wall time is recorded in
`retrieval_eval.json`). On corpora where it does not win, that cost buys
nothing.
## Usage
Nothing to download by hand in any of these β€” the weights resolve from this
repo on first use.
### With transformers alone
The custom pipeline is registered in this repo's config, so this needs no
install beyond `transformers`:
```python
from transformers import pipeline
chunk = pipeline("semantic-chunking", model="0xKitkat/semantic-chunker-modernbert-base",
trust_remote_code=True)
for c in chunk(document_text):
print(c["n_tokens"], c["boundary_score"], c["text"][:80])
```
Pass `return_text=True` for plain strings, or any of `threshold`,
`max_chunk_tokens`, `min_chunk_tokens` to tune it.
### With the pip package
```bash
pip install "boundary-chunker[onnx]" # CPU, no torch
pip install "boundary-chunker[torch]" # GPU
```
```python
from boundary_chunker import SemanticChunker
chunker = SemanticChunker(backend="onnx") # or backend="torch"
for c in chunker.split(document_text):
print(c.n_tokens, c.text[:80])
```
There is a CLI too:
```bash
boundary-chunk --backend onnx --file report.txt --json
```
### Straight from this repo, no install
```bash
hf download 0xKitkat/semantic-chunker-modernbert-base chunk.py --local-dir .
```
```python
from chunk import SemanticChunker
chunker = SemanticChunker(threshold=0.5, max_chunk_tokens=512)
```
`max_chunk_tokens` force-splits over-long chunks at the model's lowest-confidence
interior boundary rather than at an arbitrary offset.
### CPU-only, without torch
An ONNX export lives in `onnx/` and needs `onnxruntime` instead of the full
torch stack β€” roughly 40MB of dependencies rather than a CUDA install. It is
self-contained, with its own config and tokenizer beside the graph, and
`backend="onnx"` fetches it for you.
The export is verified against the torch model on held-out wiki727k documents:
identical chunk boundaries on 24/24 documents (739–65,468 chars, including
multi-window ones), and max |P(boundary) delta| 8.1e-06 on padded batches. The
pipeline is likewise checked against the direct API β€” identical boundaries on
12/12 documents β€” so all three entry points above give the same answer. Both
checks are reproducible with `onnx_parity.py` in the training repo.
## Training
| | |
|---|---|
| Base | `answerdotai/ModernBERT-base` (149M) |
| Data | wiki727k, 582,160 train docs / 30,580,099 sentences |
| Boundary rate | 11.7% |
| Max length | 2048 |
| Effective batch | 32 |
| LR / schedule | 3e-5, linear, 6% warmup |
| Epochs | 1 |
| Loss | class-weighted CE (positive weight 3.0) |
| Hardware | 1x RTX 4070 SUPER (12GB) |
Section titles are dropped from the training text: with headings present the
task partly degenerates into "a heading follows", which does not transfer to the
unformatted prose a chunker sees in production.
## Limitations
- English only.
- Trained on encyclopedic prose; conversational transcripts, code, and tabular
documents are out of distribution.
- Boundary F1 is a proxy. If your goal is retrieval quality, measure retrieval.