Instructions to use episod/tt-tnt-1024 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use episod/tt-tnt-1024 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="episod/tt-tnt-1024") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("episod/tt-tnt-1024") model = AutoModelForCausalLM.from_pretrained("episod/tt-tnt-1024", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use episod/tt-tnt-1024 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "episod/tt-tnt-1024" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "episod/tt-tnt-1024", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/episod/tt-tnt-1024
- SGLang
How to use episod/tt-tnt-1024 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "episod/tt-tnt-1024" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "episod/tt-tnt-1024", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "episod/tt-tnt-1024" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "episod/tt-tnt-1024", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use episod/tt-tnt-1024 with Docker Model Runner:
docker model run hf.co/episod/tt-tnt-1024
tt-tnt-1024
tt-tnt-1024 is a 123M-parameter Llama-3-style model with a 512-token context. It was trained
from random initialization on Tenstorrent Blackhole with ttml (tt-train), in three passes:
- over a ten-source curated corpus that includes a small
databricks-dolly-15kquestion-answer slice; - over 2.53B tokens of FineWeb-Edu;
- over the curated blend again.
It is served through the Tenstorrent vLLM plugin across all four Blackhole chips of a P300x2,
as a (1, 4) ring mesh with FABRIC_2D_TORUS_XY. Its chat endpoint uses the plain
Question: … Answer: … format it saw in training. It is small on purpose, and useful as an
instrument rather than a product. It is more fluent than earlier checkpoints, but not more
knowledgeable.
Status: Experimental. An earlier 4-chip output-quality regression didn't reproduce on these weights, but its cause was never identified (see Limitations).
It is the larger sibling of episod/tt-tnt, a project
first published as tt-nanollama3.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 6, v6 thin).
At a glance
| Architecture | Llama-3 style decoder: 122,962,944 parameters, hidden 1024, 8 layers, 16 heads / 4 KV heads, SwiGLU (intermediate 2816), RoPE θ=500000, tied embeddings |
| Hardware | P300x2: 4 Blackhole chips (two p300c cards), (1, 4) ring, FABRIC_2D_TORUS_XY |
| Context | 512 tokens (max_model_len 512). Longer prompts get HTTP 400. The chat template keeps only the last 5 messages |
| Vocabulary | 32,000 (BPE, trained on this project's corpus) |
| License | Apache-2.0 (weights and code). The training data includes share-alike and attribution-licensed sources; see Licensing |
| Status | Experimental: serves on 4 chips and is measured at concurrency 1 and 8. The historical 4-chip regression wasn't reproduced, and its cause is unexplained |
| Model CI v0 | not yet run |
Intended use
Direct use: a hardware-and-tooling research instrument. Use it for short story continuation,
short single questions in the Question: … Answer: … format, and experiments on training,
packaging, and serving on Tenstorrent hardware.
Out-of-scope use:
- Anything that needs facts. MMLU sits below chance at every checkpoint, and it answers confidently when wrong.
- A system role. The template drops system messages, because nothing in the training data had one.
- Tool calling. These weights emit no tool calls.
- Long contexts. Prompts longer than 512 tokens are rejected.
Quickstart
uv tool install tenstorrent # once — the Tenstorrent CLI, `tt`
tt model pull episod/tt-tnt-1024
tt serve episod/tt-tnt-1024
tt model pull installs the bundle into its own venv: Python 3.12, ttnn==0.77.0,
empty-target vLLM 0.25.1, and the TT vLLM plugin. It downloads the weights by default (tt
has no --with-weights flag). The server listens on port 20000 by default and walks upward if
that port is busy.
The first serve does the following:
- initializes the 4-chip fabric (
Fabric initialized on 4 devicesin the log); - captures prefill and decode traces;
- converts the weights into a tensor cache inside the install folder.
On a TT-QuietBox 2 the first serve reached a ready endpoint in 64 s, and a restart is faster.
The server is ready when the log prints Application startup complete.
Without tt-cli, tt-model alone does the whole job:
tt-model pull episod/tt-tnt-1024 --with-weights
tt-model serve episod/tt-tnt-1024
Stop it with tt-model stop episod/tt-tnt-1024. The bundle needs four chips:
- If you set
TT_METAL_VISIBLE_DEVICES, that choice wins. - Otherwise, if a chip grant sets
TT_VISIBLE_DEVICES,run.shuses the grant's first 4 devices. It refuses to start if the grant has fewer than 4. - Otherwise it opens chips
0,1,2,3.
Serve profiles
| profile | hardware | mesh | max_num_seqs | block_size | max_model_len |
|---|---|---|---|---|---|
| P300x2 (only profile) | 4 Blackhole chips, two p300c | (1, 4) ring (mesh-1x4-ring.textproto), FABRIC_2D_TORUS_XY |
32 | 64 | 512 |
4 KV heads divide 1, 2, and 4, so the architecture can shard across any of those. Only the 4-chip profile is packaged. On 2026-09-27, one chip (TP1) decoded at the same concurrency-1 speed as four (2.93 against 2.88 ms/token) during tuning. A 1-chip profile isn't packaged.
Use a vLLM TT plugin at or after c127c17. Earlier builds have a decode defect that degrades
free-running generation into repetition within a few tokens. The plugin reports version 0.1.0
either way, so a version check can't detect this; the bundle's adapter warns structurally.
Using it
The tokenizer's chat template (dolly_qa, since 2026-09-27) renders a conversation as the exact
token sequence of a databricks-dolly-15k document in the pretraining corpus:
- Each closed exchange becomes
Question: q Answer: a</s>. - The last user turn becomes
Question: q Answer:, and the model writes the answer. - System messages are dropped.
- Each line loses trailing spaces and tabs, blank lines are dropped, and lines are joined with a single space.
- Only the last 5 messages are kept.
- Reference text goes in the user message after a blank line, where dolly's
contextfield sat.
curl -s http://localhost:20000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "episod/tt-tnt-1024",
"messages": [{"role": "user", "content": "What is the capital of France?"}],
"max_tokens": 48, "temperature": 0}'
On 2026-09-27, this bundle on 4 chips returned the following (greedy; each reply ended on its
own with finish_reason: stop):
What is the capital of France? → The capital of France is Paris.
Who wrote Romeo and Juliet? → Juliet is the daughter of King Henry VIII of England.
(after the France exchange as history) What is the capital of Italy? → The capital of Italy is Rome.
The second answer is fluent and wrong, which is typical of this model. Every prompt token vLLM received matched a local render of the template, in 3 of 3 requests (one with a system message and a history).
Why this format. The previous template rendered Q: …\nAnswer:. That sequence occurs zero
times in the training tokens: the pipeline encoded the corpus line by line, so newlines were
never tokens. The results of the switch:
- The new render matches the training pipeline's own ids on 15,006 of 15,006 consecutive dolly document pairs.
- On 30 fresh questions (CPU), answers end cleanly with
</s>77% of the time, against 7% with the old template. - The old prompt cost 0.41 nats/token of answer likelihood (paired over 300 documents, t = 14.9).
The proof is in
docs/measurements/chat-template-proof.json
and scripts/verify_chat_templates.py.
Why the 5-message cap exists. A growing conversation otherwise crashes the vLLM engine
outright. That is a generic tt-metal/vLLM defect, not this model: it reproduces identically on
stock meta-llama/Llama-3.2-1B-Instruct at the same context size (entry 6 in
docs/upstream-tt-metal-asks.md).
- With the guard active at 512 context, 14 growing turns ran with zero crashes. Prompt tokens plateaued at ~102 while the number of messages sent grew to 27.
- Without the guard, the engine crashed hard at turn 3.
- That was measured on the prior (dialogue) weights; the guard is unchanged since.
/v1/completions works for plain story continuation.
Expected performance
Serving throughput and latency (on device, 4 chips, measured 2026-09-27):
| sampling | ISL / OSL | concurrency | N | TTFT median (p99) | TPOT median (p99) | tok/s per user | output tok/s |
|---|---|---|---|---|---|---|---|
| greedy | 128 / 128 | 1 | 16 | 6.52 ms (7.45) | 3.16 ms (4.05) | 317 | 310 |
| greedy | 128 / 128 | 8 | 64 | 22.10 ms (105.50) | 3.40 ms (3.91) | 294 | 2,174 |
| greedy | 384 / 128 | 1 | 8 | 14.00 ms (20.78) | 2.96 ms (3.07) | 338 | 328 |
| default (temperature 1.0) | 128 / 128 | 1 | 16 | 7.46 ms (7.92) | 4.63 ms (5.71) | 216 | 212 |
| default (temperature 1.0) | 128 / 128 | 8 | 64 | 24.78 ms (46.45) | 7.88 ms (8.63) | 127 | 995 |
Methodology.
- Hardware. A TT-QuietBox 2 (2 × p300c, 4 Blackhole chips),
(1, 4)mesh,FABRIC_2D_TORUS_XY. - Bundle and stack. The 2026-09-27 v6 thin build, identical to this one except for the weights-revision string; ttnn 0.77.0, vLLM 0.25.1.
- Tool.
vllm bench serveover streaming/v1/completions, with random-token prompts (seed 0,--ignore-eos) and 2 warmup requests. - Passes. Each row is the second of two identical passes, so compile time is excluded.
- Sampling. "Default" sends no temperature, so vLLM's default of 1.0 applies. Most chat clients do the same.
How to read these numbers:
- ISL 384 is the longest prompt that fits 128 output tokens in the 512 context. Before 2026-09-27, any prompt longer than 128 tokens killed the server.
- Run-to-run variation is about 10%. The tuning run earlier the same day measured 2.88 ms c1 TPOT and 2,487 tok/s at concurrency 8 with the same settings and bundle contents.
- Higher concurrency. The tuning run reached 7,779 tok/s at concurrency 32 and 7,567 at 64,
because the batch is capped at 32
(
docs/measurements/serving-tune-2026-09-27.md). - Sampling costs speed. Default sampling at concurrency 8 cost 2.3× in TPOT here (7.88 against 3.40 ms). A single-pass tuning measurement showed 15.3 ms.
Accuracy (CPU, on these exact weights, sha256 b2958652e9cc3308…):
| benchmark | this checkpoint (Stage B) | prior checkpoint (dialogue) | chance | source |
|---|---|---|---|---|
| wikitext bits/byte (lower is better) | 1.2551 | 1.4584 | n/a | docs/measurements/external-tt-tnt-1024-stageb.md |
| lambada_openai last-word acc | 0.2135 | 0.0980 | ~0 | same |
| piqa acc | 0.5925 | 0.5484 | 0.50 | same |
| arc_easy acc | 0.4272 | 0.3106 | 0.25 | same |
| hellaswag acc | 0.2745 | not recorded | 0.25 | same |
| winogrande acc | 0.4988 (at chance) | not recorded | 0.50 | same |
| arc_challenge acc | 0.2133 (below chance) | 0.1783 (below chance) | 0.25 | same |
| mmlu acc | 0.2292 (below chance) | 0.2295 | 0.25 | same |
| StoryCloze (1,511 items) | 0.6161 (context-blind 0.5248) | 0.6062 | 0.5281 | docs/measurements/storycloze-tt-tnt-1024-stageb.json, storycloze-tt-tnt-1024-vs-stageb.json |
Accuracy sources and methodology.
- The Stage B column comes from EleutherAI lm-evaluation-harness 0.4.9 on CPU (fp32, a
512-token window, batch 16) through
scripts/benchmark_external.py. StoryCloze comes fromscripts/eval_storycloze.py(xstory_clozeen, mean-per-token normalization). - The prior-checkpoint column comes from the same instruments at the same window, as recorded in
docs/current_model.json.
How to read the accuracy rows:
- StoryCloze is parity with the prior checkpoint, not a win.
- Stage A alone significantly regressed it (0.6062 → 0.5725, paired sign test p=0.00504).
- Stage B significantly reversed that regression (p=6.2e-05 vs Stage A; 166 of 266 discordant pairs favour Stage B).
- Against the prior checkpoint, the difference isn't significant (p=0.326; 109 of 203 discordant pairs).
- MMLU never moves. It reads 0.2295, then 0.2297, then 0.2292 across a 7.2× increase in training tokens. ARC-Challenge stays at or below chance throughout. The gains are in fluency and commonsense (wikitext, LAMBADA, PIQA, ARC-Easy), not knowledge.
- Stage A's composition and token count moved together. Stage A replaced a share of the curated blend with 100% web-educational text. So "more tokens helped" and "the register shifted" are confounded in the external-benchmark gains.
- These numbers use a 512-token window. They aren't comparable to
episod/tt-tnt's 2048-window figures.
On-device agreement with CPU (greedy, 10 prompts × 64 tokens, 2026-09-27; tokens that match CPU fp32 before the first divergence):
| topology | mean matching tokens | prompts identical to CPU |
|---|---|---|
4 chips (TP4, FABRIC_2D_TORUS_XY) |
17.4 | 4/10 |
| 1 chip (TP1) | 15.1 | 3/10 |
TP4 and TP1 are token-identical to each other on 4 of 10 prompts. Every divergence is ordinary bf16 drift into different, fluent English.
Limitations
Context is 512 tokens, enforced.
max_model_lenis 512.- A 513-token prompt gets a clean HTTP 400 ("maximum context length is 512 tokens"), and the server keeps serving. A 496-token prompt succeeds.
- The published bundle before 2026-09-27 accepted up to 512 tokens. It crashed the whole server on any prompt over 128 tokens, because prefill was padded to a 1024 bucket. The adapter now clamps that bucket to 512.
max_num_seqs is 32, and that is a ceiling of the stack, not a tuning choice.
tt_transformers 0.77 supports batch sizes 1, 2, 4, 8, 16, and 32 only. With 64 or 128 the server fails to start (
ValueError: Batch size 64 not supported). KV-cache capacity was not the limit.The only way past 32 is vLLM data parallelism. Four 1-chip replicas (DP4, 4 × 32 sequences) were measured:
concurrency DP4 TP4 (shipped) 1 (TPOT) 3.25 ms (13% slower) 2.88 ms 8 1,018 tok/s 2,487 tok/s 32 3,713 tok/s 7,779 tok/s 128 12,366 tok/s not measured DP4 only pays off at concurrency 128 and above, so it isn't shipped (
docs/measurements/serving-tune-2026-09-27.md).
4-chip output quality: a documented regression that did not reproduce.
- The earlier record. A controlled A/B on the prior (dialogue) weights produced invented
non-words on 4 chips (
FABRIC_2D_TORUS_XY), such as "Tryburg", "Alexandary", and "Higheriq". The same prompt and sampling settings on 2 chips gave rough but recognizable English (docs/serving-with-tt-kernel.md§8).- A stale output buffer behind a live trace was the working hypothesis. It was never confirmed
(entry 7 in
docs/upstream-tt-metal-asks.md).
- A stale output buffer behind a live trace was the working hypothesis. It was never confirmed
(entry 7 in
- The 2026-09-27 comparison. The regression did not reproduce on the current Stage B weights.
- 1-chip vs 4-chip, greedy, 10 prompts: fluent English on both topologies, with no invented non-words.
- 4 chips agreed with CPU slightly longer than 1 chip did (table above).
- The three chat answers above, served on 4 chips, are also fluent English.
- What this does and doesn't show. Ten prompts show that 4 chips are not measurably worse
than 1 chip here. They can't rule the regression out.
- The two comparisons differ in weights, in comparison topology (2 chips then, 1 chip now),
and in decoding (then, a sampled short-generation task
from
scripts/story_tools.py; now, greedy decoding). - Nothing has explained the original observation, so it isn't declared fixed.
- The two comparisons differ in weights, in comparison topology (2 chips then, 1 chip now),
and in decoding (then, a sampled short-generation task
from
The tokenizer is pinned; the on-device weights load from main.
- The manifest pins
weights.revisionto the commit that added thedolly_qachat template, and vLLM loads the tokenizer and template from that commit. - The tt_transformers model code loads the model config and weights from
episod/tt-tnt-1024without a revision, so it readsmain. The two are identical today. - On 2026-09-25 these weights replaced the dialogue checkpoint on
mainwithout a bundle rebuild. Any future weight upload would again change what installs serve, and it will come with a repackage and a changelog entry. - The tensor cache is keyed by commit, so stale weights are never served. A card-only edit triggers one re-conversion on the next serve.
It repeats. Greedy decoding often falls into a repetition loop after the first sentence,
though the new template lets short answers stop on </s>.
- 2.5B+ tokens of additional pretraining (Stage A) didn't fix it.
- On the prior checkpoint, the 4-gram repeat rate was measurably worse than
tt-tnt-1024a's, at 3.32× the seed floor:
| signal | delta | vs seed floor | verdict |
|---|---|---|---|
| 4-gram repeat rate | +0.0074 | 3.32× | worse |
| termination rate | −0.0076 | 0.52× | not interpretable |
| genre collapse | −0.0035 | 0.06× | not interpretable |
| loss at matched window | +0.0102 | — | no floor for this instrument |
- Nine of ten behavioural signals came back NOT INTERPRETABLE against the project's 0.1944-nat seed-only noise floor.
- That comparison was measured on the prior (dialogue) checkpoint and is kept for the record
(
docs/measurements/evaluation-tt-tnt-1024a-vs-tt-tnt-1024-dialogue.md).
It isn't instruction-tuned beyond a 2% slice of databricks-dolly-15k.
- It has no system role. That role is structurally absent from anything this regimen could produce, not merely unimplemented.
- It is more fluent than earlier checkpoints and not more knowledgeable.
- Treat it as an artifact of a hardware-and-tooling project, not as an assistant.
Tool calling isn't supported on these weights. They emit no tool calls. A separate, unpublished continued-training checkpoint does (see History).
Risks and safety considerations
- Confident wrong answers. It answers in the shape of an answer whether or not it knows. On 2026-09-27 it said Juliet "is the daughter of King Henry VIII of England". Don't rely on it for facts.
- 4-chip numerical path. The historical invented-word regression (see Limitations) is unexplained. It is a TT-specific serving-path deviation, not a property of the weights. Watch for non-words in long free-running output.
- Repetition loops under greedy decoding. Use sampling for open-ended text, and accept the throughput cost shown above.
- Unbounded conversations crash the engine without the shipped chat-template guard. Don't replace the template with one that renders the full history.
Licensing
- Weights and this project's code: Apache-2.0.
- Stage B / curated blend (400M-token budget, shipped only as a
recipe):
- 46% of the blend is share-alike, under two mutually incompatible copyleft terms.
- TinyStories is CDLA-Sharing-1.0.
- Wikipedia is CC-BY-SA-3.0.
- The
databricks-dolly-15kdialogue slice (2%, rendered as plainQuestion: … Answer: …prose with no role markers) is CC-BY-SA-3.0. - Per-source terms are in
docs/corpus_licensing.md.
- Stage A: 2,529,270,500 tokens of
HuggingFaceFW/fineweb-edu(sample-10BT, ODC-By 1.0, pinned revision87f09149ef4734204d70ed1d046ddc9ca3f2b8f9).- It was fetched directly rather than through the project's corpus-blend registry, so it isn't
part of the
episod/tt-tnt-corpusrecipe. - By a wide margin, it is the majority of what these weights have read.
- See
docs/current_model.json'scorpus.note.
- It was fetched directly rather than through the project's corpus-blend registry, so it isn't
part of the
Training
| stage | data | steps | notes |
|---|---|---|---|
| dialogue (base) | curated ten-source blend (tokens-v4, 399,486,992 emitted tokens) |
10,764 | batch 64, seq 512 |
| Stage A | FineWeb-Edu, 2,529,270,500 tokens (the Chinchilla-matched budget for 123M) | 76,503 | cosine lr, --ddp 4 |
| Stage B | curated blend again (tokens-v4), one epoch |
10,761 | cosine lr warm restart, stochastic_rounding: True, final step 87,264 |
Training ran as 4-chip DDP across both p300c cards of a TT-QuietBox 2. The final validation loss
is 2.5373 (Stage B, real held-out loss over the full validation split). The ten-source blend is
episod/tt-tnt's nine sources plus the dolly slice.
History: experiments on the prior checkpoint
⚠️ Everything in this section was measured against the prior (2026-08-20 to 2026-08-29 dialogue) checkpoint, not against the currently published Stage A/B weights. It records this model's lineage and makes no claims about the current weights.
Routing by physical die address is nearly free. Tokens can be routed to experts by where they live on the harvested 11×10 Tensix grid rather than by a learned gate:
- Freezing the gate to that geography costs only 0.0118 nats against a freely learned gate (|t| 5.1, 14/15 signs).
- Source-characteristic tokens occupy measurably distinct die regions: cell purity 0.546 against a 0.231 permutation floor.
- Seeding the gate from the die map and then letting it move buys nothing measurable (+0.0044, signs 8+/7−), even though the seeding works as a classifier (61.2% region recovery against a 10% chance floor).
Sparse routing (Mixture of Enthusiasts) beats dense from scratch. Both arms trained one epoch from init, paired on seed 5489:
- Validation was 2.8098 for MoE against 2.8748 for dense (mean delta +0.0481, |t| 7.3, 20 of 22 signs).
- It replicated at seed 8191 (+0.0354, |t| 4.5, 19/22). Pooled, the gain is +0.0417 over 44 points, with the same late-separating trajectory in both runs.
- Read it as the ordinary MoE bargain: 3.62× total parameters at 0.989× active compute. It isn't evidence about die-region routing.
Tool calling works structurally, and only structurally. A continued-training run teaches
four tools (factual_response, witty_response, absurdist_response,
misunderstood_question). They are emitted as <tool_call> blocks, which vLLM's hermes parser
turns into structured tool_calls:
| gate | trained | control |
|---|---|---|
| emits a tool call | 100% | 0% |
| parses | 85.9% | 0% |
| schema-valid | 75.0% | 0% |
| distinct tools used | 4/4 | 0/4 |
- Unseen questions score higher than seen ones (78% vs 72%).
- 3/5 chat requests came back as genuine structured
tool_callsend-to-end through the server. - The content inside the calls is poor ("The capital of France is the capital of France"; Portugal answered "Madrid").
- That checkpoint is not published. See
docs/measurements/tool-calling-stage{1,2}.json. ItsQ: …\nAnswer:format was the chat template this repo shipped until 2026-09-27.
A five-slot think-block can be learned, but it doesn't help yet. The fine-tuned model emits
offer / accept / add / stakes / handback blocks:
- Blocks are well formed in 98% of generations (784/800; the control produces them 0% of the time).
- Substituting another story's block changes 100% of continuations.
- It moves none of the four failure-mode scores at α = 0.01. The block is context the model conditions on, not an instruction it obeys.
- An earlier pass reported 0% adherence from a run in which all 17 RMSNorm gammas were frozen
(
stochastic_roundingdefaults off on the SFT path). With the gammas free, 0% became 98%. - Close reading is in
episod-log.md, 2026-08-21.
The reach dial (2026-08-23/24) works, is small, and is paused. Forcing a reach slot to
near / mid / far moves the realised semantic distance of the add word monotonically,
scene-paired over 826 held-out scenes:
| contrast | raw | frequency-residualised |
|---|---|---|
near < mid |
+0.0839 (t 16.2) | +0.0324 (t 7.4) |
mid < far |
+0.0456 (t 13.9) | +0.0281 (t 9.0) |
near < far |
+0.1295 (t 23.3) | +0.0604 (t 12.5) |
- About 53% of the raw effect is word frequency and ~47% survives the frequency control.
- A control arm never shown a
reachslot shows nothing. - The pre-declared EUREKA criterion failed on one gate:
addslot-hit shortfall 0.0896, anear-side dip. - Three explanations were eliminated:
- More training doesn't help: at 9000 steps, +0.060392 vs +0.060438.
- The arms aren't undertrained: the adherence gate got worse at 9000 steps.
- The vocabulary isn't fixable by filtering: a validated content-word filter made the slot more concentrated.
- The constraint is the corpus. TinyStories has 13,777 distinct words, and the top 1,000 cover 90.9% of tokens.
- Everything re-derives from
docs/measurements/reach-dial.jsonthroughscripts/eval_reach.py --rescore-from, with no model, tokenizer, or device.
Changelog
| date | change |
|---|---|
| 2026-08-18 | First published (dialogue-slice 512-context weights); first 1024-size tt-model bundle |
| 2026-08-29/30 | A 2048-context retrain was published (HF 038d6c6a8d), found worse at Q&A, and reverted within the hour to the 512 dialogue weights (57224f4b87). The chat template's 5-message guard was added |
| 2026-09-01 | Bundle manifest republished at schema v5 |
| 2026-09-05 | First v6 thin package |
| 2026-09-15 | Briefly a v5 fat package, then repackaged v6 thin (136d240427); stale v5-fat tree removed |
| 2026-09-25 | Weights replaced with the Stage A + Stage B checkpoint (9abed7c785), designated 2026-09-24. Bundle not rebuilt |
| 2026-09-27 | Repackaged for serving; weights unchanged. Changes: the chat template becomes dolly_qa (Question: … Answer: …</s>, the trained format; it replaces Q:\nAnswer:); max_model_len is set to 512 and adapter 1.1.0 clamps prefill buckets, so prompts of 129–512 tokens no longer crash the server and longer ones get HTTP 400; run.sh honours a chip grant; the manifest records tt_metal_version 0.77.0 and pins the weights revision. Hardware-verified on 4 chips; measured at concurrency 1 and 8, greedy and default sampling; 1-chip vs 4-chip output compared |
Related packages
episod/tt-tntis the 22M sibling in the same from-scratch family. It is a different model, not a hardware variant: hidden 384, 3 KV heads, 2048 context, no FineWeb-Edu, and no dialogue data, so its chat template is plain continuation. It serves on one chip. Neither package supersedes the other.episod/tt-tnt-corpusis the curated blend recipe. Stage A's FineWeb-Edu isn't part of it.
Feedback
- Questions or problems with this package: open a discussion at
https://huggingface.co/episod/tt-tnt-1024/discussions. That is the one channel that reaches the bundle's author. - A problem with the
tttooling itself: usett report issue. It opens a prefilled issue against tenstorrent/tt-cli, not this package. - Product feedback: send it to support@tenstorrent.com.
The full build log and every measurement are at tsingletaryTT/tt-tnt.
Provenance
| component | built from |
|---|---|
| tt-metal | ttnn==0.77.0 (PyPI pin in requirements.txt; the manifest records tt_metal_version: 0.77.0) |
| tt-metal models code | tt-tnt-models-closure==0.77.0, a vendored models/common + models/tt_transformers/tt (not upstream tt-metal-models) |
| vLLM | 0.25.1, empty target |
| vllm-tt-plugin | 0.1.0 wheel built 2026-09-15 (sha256 3990ec4d…). Its source commit isn't recorded in the wheel |
| mesh | mesh-1x4-ring.textproto (dims: [1, 4], [LINE, RING]), FABRIC_2D_TORUS_XY |
| weights | model.safetensors sha256 b2958652e9cc3308…, Stage B step 87,264, first published in HF commit 9abed7c7855d39583bb757a8bbf679ce569c40f6. The manifest pins weights.revision to the commit that added the dolly_qa chat template |
| adapter | tt_tnt_adapter.py 1.1.0 = bundle/tt_tnt_adapter.py at tsingletaryTT/tt-tnt commit 2830e43 |
| designation | docs/current_model.json, commit 44d8496 |
| Model CI v0 | not yet run |
| build | 2026-09-27 · tt-model-manager 0.1.0, integration build of the v6 thin fixes · schema 6 |
- Downloads last month
- 2,236