customs-decomposer / README.md
NagaYu's picture
Publish Customs span tagger (safetensors + MLX + GGUF)
6610811 verified
|
Raw
History Blame Contribute Delete
7.05 kB
---
license: apache-2.0
library_name: pytorch
pipeline_tag: token-classification
language:
- en
tags:
- customs
- token-classification
- span-tagging
- pii
- pii-detection
- privacy
- data-minimisation
- on-device
- synthetic
- safetensors
- mlx
- gguf
---
# Customs Span Tagger
A small word-level transformer that tags the **personal spans** of a request, so a
device can decide which clauses may leave it. It is the learned detector inside the
`Customs` research prototype's `SemanticDecomposer`: spans it finds are either replaced
by a de-identified derived key or kept on-device entirely.
The interesting case it exists for is positional: in this world the *same* city string
is personal when it is the user's home city and public when it is a travel destination.
A gazetteer tags both. A personal-store lookup tags only the first, but only because it
was handed the answer. This model has to read the sentence.
## Architecture
| | |
| --- | --- |
| architecture | customs-span-tagger |
| d_model | 128 |
| layers | 2 |
| attention heads | 4 |
| feed-forward width | 256 |
| max sequence length | 128 |
| vocabulary size | 481 |
| BIO tags | 31 |
Word embeddings with learned positions, a pre-norm transformer encoder, and a linear
head over 31 BIO tags derived from `customs.types.DataClass`. Unseen tokens fall back to
a character-shape key (`AC-19384756` -> `<shape:A-d>`), which is what gives the tagger
recall on freshly drawn account numbers, device identifiers and phone numbers that
could never have been memorised.
No pretrained checkpoint or tokenizer is downloaded; the vocabulary is built from the
synthetic corpus alone.
## Files
| format | files |
| --- | --- |
| `safetensors` | `model.safetensors` |
| `mlx` | `model.mlx.safetensors` |
| `gguf` | `model.gguf` |
`config.json` and `vocab.json` are required to load the model in any format -- the
vocabulary carries the shape-backoff keys the tagger relies on.
## Usage
```python
from customs.model.tagger import SpanTagger
tagger = SpanTagger.from_pretrained("NagaYu/customs-decomposer")
for span in tagger.predict_spans("I live in Aveiro and I take Glucoform daily."):
print(span.start, span.end, span.text, span.data_class.slug)
```
It satisfies the `customs.detector.Detector` protocol, so it drops straight into the
decomposer:
```python
from customs.decomposer import SemanticDecomposer
from customs.store import PersonalStore
store = PersonalStore(values={"HOME_CITY": "Aveiro", "MEDICATION": "Glucoform"})
decomposer = SemanticDecomposer(store, detector=tagger)
plan = decomposer.decompose("r-1", "I live in Aveiro and I take Glucoform. Max daily dose?")
print(plan.remote_payload_text())
```
## Evaluation
Reported by the training run, and shipped verbatim as `train_report.json`:
| metric | value |
| --- | --- |
| `seed` | 0 |
| `steps` | 1,500 |
| `n_parameters` | 347,167 |
| `vocab_size` | 481 |
| `n_tags` | 31 |
| `seconds` | 33.9500 |
| `data_provenance` | fully synthetic; generated by customs.world + customs.templates |
**corpus**
| metric | value |
| --- | --- |
| `requests` | 1,440 |
| `profiles` | 60 |
| `train` | 1,080 |
| `val` | 360 |
**train**
| metric | value |
| --- | --- |
| `precision` | 1.0000 |
| `recall` | 1.0000 |
| `f1` | 1.0000 |
| `value_recall` | 1.0000 |
| `n_gold` | 2,430 |
| `n_pred` | 2,430 |
**val**
| metric | value |
| --- | --- |
| `precision` | 1.0000 |
| `recall` | 1.0000 |
| `f1` | 1.0000 |
| `value_recall` | 1.0000 |
| `n_gold` | 810 |
| `n_pred` | 810 |
**held_out_check**
| metric | value |
| --- | --- |
| `covered` | 72 |
| `total` | 72 |
| `rate` | 1.0000 |
Not printed here: `ambiguity_probe`, `history`. The full report ships as `train_report.json` in this repo.
Precision matters as much as recall here. A span wrongly tagged personal is a clause
needlessly kept on-device, which costs answer quality, so span-F1 rather than recall is
the target.
## Training data
Trained entirely on the synthetic Customs benchmark, generated by
`scripts/build_bench.py` from the invented tables in `customs/world.py`. Every name,
city, employer, medication, condition, account identifier, email address, phone number
and device identifier is fictional; the email domain used throughout is the reserved
`.invalid` TLD. **No real personal information was used at any stage**, and no external
dataset was downloaded.
Splits are cut by synthetic profile, not by request, so a tagger cannot score well at
test time by memorising a training profile's employer string.
## Scope of the claim
This model supports two specific, measurable claims and no others:
1. the **volume of personal content crossing a trust boundary can be reduced**, by
detecting personal spans and emitting derived keys in their place; and
2. **whatever does cross can be recorded**, in the egress certificate the surrounding
system writes.
Derived keys are themselves informative -- a country code or an age band is not nothing
-- so this is data minimisation. It is **not** a claim of complete privacy, and nothing
in this repository should be read as one. The safety property of the surrounding
decomposer is deliberately not a function of this model being good: clauses containing
a personal span with no derived stand-in are withheld regardless, and an independent
egress guard re-scans the constructed payload against the personal store.
## Limitations
* Trained on template-generated text. Surface diversity is far below real user writing,
and the tag set is the closed taxonomy in `customs/types.py`.
* The vocabulary is corpus-specific. Out-of-world vocabulary reaches the model only
through the shape backoff.
* The GGUF file is a container for this repository's custom span-tagger architecture,
not a drop-in for a general-purpose runtime's built-in model graphs.
* Fictional-world performance is not evidence of real-world PII detection performance.
## The project this belongs to
This tagger is one component of **Customs**, a research prototype for routing across an
on-device / private-cloud / external-cloud model hierarchy when the middle tier is free
but rationed daily. The tagger decides *which spans are personal*; the surrounding system
decides what to do about it -- bind the private half on-device and send only a
de-identified derived key, price the free ration as an online knapsack, record every
crossing, and walk a degradation ladder instead of falling off a cliff when the ration
runs out.
Full code, benchmark, figures and a Swift package: **https://github.com/NagaYu/customs**
Against confidence-threshold escalation on that benchmark (24 seeds), the full system
delivered +6.0% whole-day answer quality, +52.9% evening quality, a 73% smaller quality
gap across ration exhaustion, 14% less paid-tier spend, and zero personal tokens crossing
any boundary instead of 49.5. Those numbers are a property of the whole system, not of
this model alone.
## Provenance
Exported from `tagger` by `scripts/push_model.py` via
`customs.model.export.export_all()`.
## License
Apache-2.0.