token-provenance
Every token count carries the encoding that produced it.
A prompt is budgeted with len(text) // 4. The number is plausible.
It propagates into a cost estimate. A user sees a bill. The bill is
wrong. Nothing in the pipeline flagged it.
Concrete example: a prompt was budgeted with len(text) // 4 = 44
tokens. The serving stack used a ChatML template that adds 11 special
tokens, and the actual BPE tokenizer produced 43 tokens for the body.
Real count: 54. Under the heuristic the count fit; under the real
encoding it did not. The server truncated the prompt silently. The
answer was wrong.
token-provenance fixes this by making the encoding a first-class
property of the count.
The claim in one sentence
Attach encoding metadata to every token count at the point of production; propagate it through every operation; refuse to format or consume it unless the caller explicitly opts in.
What it produces
Every tokenizer output is a TCount carrying:
valueโ the raw intencoding.tokenizerโ which tokenizer produced the countencoding.templateโ which chat template was appliedencoding.special_tokensโ how many special tokens the template addsencoding.exactโ True iff the count came from a real tokenizer, False for heuristics likelen // 4encoding.versionโ tokenizer version string
Install
pip install token-provenance
Or run the module directly:
python token_provenance.py
Pure stdlib. No dependencies.
Usage
Basic
from token_provenance import (
tokenizer, serving, report, assert_serving,
TCount, Encoding, InexactError, ServingMismatchError,
)
SERVING = Encoding(tokenizer="gpt2", template="chatml",
special_tokens=11, version="v1", exact=True)
@tokenizer("gpt2", template="chatml", special_tokens=11,
version="v1", exact=True)
def count_gpt2_chatml(text):
"""Return (count, info)."""
ids = gpt2_tokenizer.encode(text)
return len(ids) + 11, {"notes": ("chatml template applied",)}
t = count_gpt2_chatml(prompt)
# Sanctioned reporting. Refuses inexact or mismatched encodings.
print(report(t, expected=SERVING))
# Sanctioned assertion. Fails loudly on mismatch.
assert_serving(t, SERVING, "startup")
# Arithmetic propagates encoding.
result = (t + 100) * 2
print(report(result, expected=SERVING))
Strict consumers
@serving(SERVING)
def cost_estimate(count, price_per_1k=0.002):
"""Refuses counts from the wrong encoding."""
return count.value * price_per_1k / 1000.0
try:
cost_estimate(naive_count)
except ServingMismatchError as e:
print(f"refused: {e}")
Arrays
from token_provenance import TArray
# A batch of independent token counts.
counts = TArray([count_gpt2_chatml(p) for p in prompts])
# Elementwise arithmetic preserves the mixed state.
doubled = counts * 2 + 1
# Reductions merge provenance.
total = counts.sum()
if total.exact:
print(f"total = {total.value}")
else:
print(f"refused: {report(total)}")
Cross-checks
from token_provenance import cross_check, Mismatch
result = cross_check([count_gpt2, count_llama, count_heuristic],
prompt, rel_tol=0.05)
if isinstance(result, TCount):
print(f"all agree: {report(result)}")
else: # Mismatch
print(f"mismatch: {report(result)}")
idx, name, dist = result.outlier()
print(f"outlier: {name}")
Adaptive budgets
from token_provenance import safe_context_budget
# Instead of hardcoding 4096 - 512 = 3584:
budget = safe_context_budget(4096, SERVING, reserve_for_output=512)
# = 4096 - 11 - 512 = 3573
The six defenses
The module provides six levels of escalation for the same bug class:
| level | mechanism | action |
|---|---|---|
| 1. tag | @tokenizer |
every output is a TCount, not an int |
| 2. propagate | arithmetic | taint spreads through + - * // |
| 3. refuse | report() |
refuses to format inexact or mismatched counts |
| 4. raise | @serving |
raises ServingMismatchError on wrong encoding |
| 5. assert | assert_serving() |
AssertionError at pipeline startup |
| 6. bypass | .unwrap(allow_inexact=True) |
requires explicit opt-in |
Any pipeline can choose its level. The choice is explicit at every boundary.
Benchmarks
The demo (python token_provenance.py)
| part | concept | outcome |
|---|---|---|
| 0 | naive pipeline | prints 36 with no flag |
| 1 | TCount |
carries exact=False, tokenizer=heuristic:chars//4 |
| 2 | taint propagation | (naive + 100) * 2 = 272, still inexact |
| 3 | report() |
refuses inexact, formats exact as 54 [toy-bpe+chatml(+11)] |
| 4 | @serving |
raises on heuristic and on llama2; accepts chatml |
| 5 | escape hatch | .unwrap(allow_inexact=True) = 36; .unwrap() raises |
| 6 | exact path | 54 with tokenizer=toy-bpe, special_tokens=11 |
| 7 | assert_serving |
fails on llama2 and heuristic; returns raw 54 on chatml |
| 8 | TArray |
mixed batch 3/4 exact, sum=72; exact batch sum=95 |
| 9 | adaptive budget | none=3584, llama2=3577, chatml=3573 |
| 10 | cross_check |
A=mixed TCount, B/C=Mismatch, D=clean TCount |
| 11 | discipline | the rule set, in one place |
Total runtime: <1 second on a laptop CPU.
Template overhead and context budget
safe_context_budget(model_max, encoding, reserve_for_output):
| model_max | reserve | template | special | budget |
|---|---|---|---|---|
| 4096 | 512 | none | 0 | 3584 |
| 4096 | 512 | llama2 | 7 | 3577 |
| 4096 | 512 | chatml | 11 | 3573 |
The naive answer 4096 - 512 = 3584 silently overshoots by 11 tokens
under ChatML. That is enough to truncate the last few tokens of a
prompt โ often the JSON closing brace or the instruction suffix.
The outlier rules
When cross_check returns a Mismatch, the outlier is identified by
three rules in order:
- Inexact wins. A heuristic encoder is the outlier by construction, regardless of where its value landed.
- Most special tokens wins. If all encoders are exact but templates differ, the one with the largest template overhead is less trustworthy. Ties break by distance from the min-special-tokens group mean.
- Distance from mean. If all encodings are identical, fall back to the value farthest from the mean. The mean is unambiguous for both odd and even counts, unlike the median.
The order matters. A rule based purely on distance-from-mean gets the outlier wrong in two cases:
- A heuristic that happens to land near the exact encoders is not flagged, even though it was produced by a different process.
- A chatml count that happens to land near the raw token count is flagged as the outlier when the real problem is the missing template on the other solvers.
Rules 1 and 2 fix both cases. Rule 3 is the fallback for the case where every encoder used the same encoding and still disagreed โ which means the disagreement is real, not a provenance artifact.
Value agreement is not encoding equivalence
This is the subtlest point in the module.
cross_check returning a TCount means the values agreed within
tolerance. It does not mean the counts are interchangeable with
the serving config. A merged count of {toy-bpe, toy-bpe+chatml, toy-bpe+llama2} is not a chatml count, even if its mean value happens
to land near one.
assert_serving enforces this: a merged encoding matches the serving
config only if every constituent matches. This is why demo Case
A is refused even though the three encoders agreed within
rel_tol=0.30:
A (refused) : AssertionError
A: encoding 'mix[toy-bpe, toy-bpe+chatml(+11), toy-bpe+llama2(+7)]'
!= serving 'toy-bpe+chatml(+11)'
Case D passes because all three tokenizers used the same chatml template, so the merged encoding is trivially a chatml encoding:
D (passes) : 54
Value agreement is necessary but not sufficient.
When to use it
- Any pipeline that bills or truncates by token count. Cost estimation, context-window budgeting, rate limiting. If a count decides what the user sees or pays, wrap it.
- Multi-tokenizer pipelines. When two encoders should agree,
cross_checkmakes disagreement an output rather than a silent pass-through. - Serving-config-aware pipelines. When the model, the template,
or the special-token set can drift between training and serving,
@servingcatches the drift at the first call site. - Publication-grade methodology. A token count that goes into a
paper should carry its encoding.
report()refuses to format anything else.
When not to use it
- When the encoder is always the same and always exact. A single fixed tokenizer at a single fixed template has no provenance to track. The overhead is small but nonzero.
- When performance is critical and profiling shows the wrapper
dominates. The
TCountlayer is cheap (a few microseconds per operation), but arithmetic on many small counts in a tight loop will pay for the provenance. Unwrap at the boundary and rewrap the result. - When downstream code cannot accept the wrapper. Third-party
libraries will coerce
TCountto a plainintat the call boundary. Unwrap before calling, rewrap after. - As a substitute for using the real tokenizer. The module makes the heuristic visible; it does not make it accurate. If your budget allows, run the real tokenizer.
Honest limitations
- Trusts the tokenizer's self-report. A
@tokenizerfunction can claimexact=Truewhile doinglen // 4. The module has no way to verify the claim. The decorator'sexactargument is a declaration, not a proof. - Integers only, plus
TArray. No float counts (subword fractions, weighted averages). Costs must be computed on the unwrapped int. - Arithmetic covers
+ - * //. No%, no**, nomath.logfor cost formulas that need it. Add the ones you need by extending_arith. - Comparisons strip provenance.
tc < 4096returns a plainbool. If you need to propagate the fact that a comparison was made on an inexact count, useassert_servingfirst. cross_checkruns tokenizers sequentially. No parallelism. Tokenizers are usually fast enough that this does not matter; for a remote tokenizer API, run them yourself and constructMismatchmanually.safe_context_budgetis an accounting formula, not a guarantee. Real serving stacks may inject additional tokens (system prompts, tool schemas, RAG prefixes) that are invisible to the tokenizer used to produce the count. The budget is a floor on what fits, not a ceiling.- No calibration against real tokenizers. The demo uses
toy-bpe, a synthetic counter. The numbers in the demo are illustrative, not measured against GPT-2, Llama, or any production tokenizer.
The bug it prevents
The silent truncation from the introduction would have looked like this in a provenance-aware pipeline:
naive budget : 4096 - 512 = 3584 tokens available
real encoding : toy-bpe+chatml(+11)
real count : 54 tokens for a 143-char prompt
report(naive) : <refused: 'heuristic:chars//4' is a heuristic;
no exact tokenizer was used>
report(exact) : 54 [toy-bpe+chatml(+11)]
assert_serving : AssertionError: encoding 'heuristic:chars//4'
!= serving 'toy-bpe+chatml(+11)'
The heuristic count never reaches a budget. The prompt is never truncated silently. The failure is caught at the point of production, not at the point of billing.
Version history
| version | change |
|---|---|
| 0.1.0 | Encoding, TCount, @tokenizer, report(), assert_serving() |
| 0.2.0 | added TArray for sequences |
| 0.3.0 | added cross_check and Mismatch |
| 0.4.0 | safe_context_budget accounts for template overhead |
| 0.4.1 | Mismatch.outlier() fixed: inexact first, then most-special-tokens, then distance from mean |
| 0.4.2 | Encoding.merge flattens nested merges; serving_matches requires every constituent to match |
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
hv-manifold |
a corpus | the geometry of style space |
hv-reader |
one text | how it reads |
anomaly-or-bug |
a number and a matrix | is this a bug or a discovery? |
frontier-check |
one claim | where does it sit relative to the frontier? |
numerical-provenance |
a numeric pipeline | can I trust this number? |
token-provenance |
an LLM pipeline | can I trust this token count? |
The design principle โ that a token count should carry its own
encoding rather than relying on the caller to remember which
tokenizer produced it โ came out of a session in which a
len(text) // 4 heuristic silently disagreed with a ChatML-wrapped
BPE tokenizer by 18 tokens. The tool was built to make that class
of bug impossible.
License
Apache-2.0