Tessera-1B-Nano-Base

The matched NTP-only baseline of the Tessera-1B-Nano comparison from Paragon Intelligence Labs: the unchanged HuggingFaceTB/SmolLM2-360M backbone, continued-pretrained without a concept path. It exists so the concept-arm result is interpretable.

TL;DR

This arm is the control of a token-matched comparison: it consumed the same tokens of the same packed corpus, in the same order, from the same initialization as the concept arm, with no restarts and no NaNs.

  • Token loss is the reference: the final numbers below are the baseline the concept arm is compared against.
  • The comparison outcome and the intervention readings live in the whitepaper at docs/whitepaper.md in the ncp-smol repository.

Evaluation

Metric Value
Held-out NTP loss 2.5135
Held-out perplexity 12.3485
Training tokens 999,948,288
Tracked compute estimate $24.88

Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.

Architecture

  • chunk size: 4
  • product code: 15 segments x 64 entries
  • causal concept blocks: 2
  • injection point: before token decoder block 2
  • NCP target: next continuous concept
  • loss: L_ntp + 1 L_ncp + 1 L_vq

This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.

Training data and provenance

The matched corpus and the training record are pinned in the ncp-smol repository: packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw intervention evaluation JSON (also shipped in this repository as eval.json and metrics.jsonl).

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-1B-Nano-Base")
model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano-Base")

References

Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yava-code/Tessera-1B-Nano-Base

Finetuned
(121)
this model

Papers for yava-code/Tessera-1B-Nano-Base