tessera
A tiny model for extracting contact details from text, locally.
Tessera finds people, organizations, postal addresses, emails, and phone numbers. It returns their exact positions in the original text, splits addresses into components, and groups related details into contacts. The same Rust implementation runs natively and through WebAssembly in browsers, Node, and Bun.
Email and phone extraction uses validating rules. Names, organizations, and addresses use an int8 neural network; a separate network parses address components. Inference stays in your process or browser worker. No text is sent to an inference service.
This release is experimental. The learned detector currently focuses on English United States material. Organization detection and unfamiliar prose need review; fresh unseen release accuracy has not been established. See results and limitations.
Usage 路 Model 路 Results 路 License
Usage
| Operation | Use it to | Returns |
|---|---|---|
detect |
Find individual entities in a document | Kind, source offsets, confidence, and origin |
parse_address / parseAddress |
Split text already known to be one address | Labeled address components |
extract_contacts / extractContacts |
Associate extracted details with people or organizations | Contacts and unassigned entities |
Rust and CLI offsets use UTF-8 bytes. JavaScript offsets use UTF-16 code units, so text.slice(start, end) returns the entity. Every end is exclusive. Keep the original string when using offsets.
Rust
For an application beside a Tessera checkout, add the runtime as a local dependency:
[dependencies]
tessera = { path = "../tessera/tessera" }
Load the bundle and its checksum once, then reuse the extractor:
use tessera::{Config, Kind, Query, Tessera};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let bytes = std::fs::read("models/tessera-v1.safetensors")?;
let checksum = std::fs::read_to_string("models/tessera-v1.sha256")?;
let extractor = Tessera::load(&bytes, Config {
kinds: Kind::all(),
expected_checksum: Some(checksum.trim()),
})?;
let query = Query { country_hint: &["US"], ..Query::default() };
let text = "Jordan Avery, Project Coordinator\nAcme Corporation\n500 Main St, Springfield, IL 62701\n+1 202-555-0199\njordan@acme.example";
let contacts = extractor.extract_contacts(text, &query)?;
println!("{} contacts, {} unassigned entities",
contacts.contacts.len(), contacts.unassigned.len());
let address = extractor.parse_address("400 Broad St, Seattle, WA 98109", &query)?;
println!("{} address components", address.components.len());
Ok(())
}
unassigned is part of a successful extraction: the grouper leaves details there when it cannot confidently choose a contact. Experimental detection may omit or mislabel entities.
JavaScript
Serve the bundle and checksum with your application:
import { createTessera } from "tessera";
const response = await fetch("/models/tessera-v1.sha256");
if (!response.ok) throw new Error("Could not load model checksum");
const integrity = (await response.text()).trim();
const extractor = await createTessera({
modelUrl: "/models/tessera-v1.safetensors",
integrity,
worker: true,
});
try {
const text = "Contact hello@example.org or +1 202-555-0123.";
const entities = await extractor.detect(text, { countryHint: ["US"] });
for (const entity of entities) {
console.log(entity.kind, text.slice(entity.start, entity.end));
}
const contacts = await extractor.extractContacts(text, {
countryHint: ["US"],
});
console.log(contacts.contacts, contacts.unassigned);
} finally {
extractor.dispose();
}
For Node or Bun, read the bundle from disk and pass modelBytes instead of modelUrl. worker: true runs inference off the browser's main thread when workers are available; it is ignored outside that environment. Rules alone need no model fetch: createTessera({ kinds: ["email", "phone"] }).
The complete JavaScript API is in index.d.ts.
Options
country_hint/countryHinthelps interpret national phone numbers; it does not change the detector's training coverage.include_uncertain/includeUncertainreturns additional uncertain results. Internal detector thresholds still apply.parse_addressaccepts one address of at most 256 non-whitespace tokens and 8 KiB.- Rust's optional
markdownfeature scans prose and contact link destinations while retaining source offsets. The default JavaScript package omits it and rejectsformat: "markdown". - Rust failures use
tessera::Error; JavaScript failures expose aTesseraErrorcode. An empty successful result means nothing was returned.
Model
The distributed Safetensors bundle contains both the entity detector and address parser. It uses Tessera's own runtime: it is not a Transformers AutoModel checkpoint, and no hosted inference endpoint is included.
| Property | Value |
|---|---|
| Model version / bundle format | 0.3.0 / 2 |
| Runtime version | 0.1.0 |
| Bundle size | 3,489,616 bytes (3.49 MB; about 3.09 MB with gzip) |
| Weight encoding | Per-channel symmetric int8; float32 scales and biases |
| Detector architecture | detector-context96-rms-v2 |
| Hidden channels / convolution kernel | 96 / 3 |
| Detector dilations | [1, 2, 4, 8, 16, 1, 64] |
| Normalization | Channel RMS after each residual block, epsilon 1e-5 |
| Context radius / window / overlap | 96 / 2,048 / 448 retained tokens |
The detector combines hashed character n-grams of lengths 2, 3, and 4 with script, shape, and contextual features. It predicts BIO labels for PERSON, ORG, and ADDRESS. Email and phone spans found by rules are protected from competing model labels. The parser labels components of supplied or detected addresses, and the grouper associates related entities.
The Safetensors header contains 39 metadata fields, including graph and decoder identifiers, dimensions, labels, feature settings, licensing, experimental status, and source checkpoint hashes. bundle.json exposes the same metadata in readable form. Load tessera-v1.sha256 with the bundle to verify its integrity.
Format 2 binds the context graph and decoder explicitly. The runtime also reads legacy format 1 bundles using their original graph. The asset filename is independent of its embedded model version.
Results and limitations
The detector completed 4,000 optimizer updates. It failed the automatic quality-promotion requirements, including the full 95% training-seen gate, and was accepted for an experimental release. The bounded run finished, but the full configured training schedule did not; metadata retains detector_training_complete=false.
These are native checkpoint exact entity F1 scores, measured before int8 packaging. Both the kind and boundaries must match:
| Evaluation set | Documents | PERSON | ORG | ADDRESS |
|---|---|---|---|---|
| Real training-seen documents | 1,799 | 96.47% | 75.39% | 97.33% |
| Reused real development documents | 196 | 76.86% | 53.76% | 89.22% |
Training-seen results measure fitting. The development set was reused during model selection. Neither establishes fresh unseen release accuracy. These scores do not measure email/phone rules, address parsing, or contact grouping. Aggregate PERSON results also conceal weaker biography coverage; organization names, aliases, and boundaries remain a substantial source of error.
Separately, the packaged address parser scored 99.05% component F1 and 95.23% completely correct parses on its 3,000-row US test shard through parse_address. This tests parsing supplied addresses, not detecting them in prose.
Eight of thirteen small public detector fixture cases remain documented failures, including missed people in prose and tables. Their expected labels are retained. These examples expose limitations; they are not an independent accuracy benchmark.
Confidence and review_recommended are signals for inspection, not guarantees. Check names, organizations, and contact assignments before using them in important records. Other countries, languages, and unfamiliar document layouts have no demonstrated general accuracy from this evaluation.
License
Code is available under MIT or Apache-2.0. Trained weights are licensed under CC BY 4.0; retain NOTICE when redistributing them. It preserves the underlying source attributions, including OpenStreetMap and Open Addresses UK. Training combines synthetic contacts and annotated real US documents; the parser has its own training history. Raw corpora and local run evidence are not part of the model upload.