Instructions to use yava-code/Tessera-1B-Nano with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yava-code/Tessera-1B-Nano with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="yava-code/Tessera-1B-Nano", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use yava-code/Tessera-1B-Nano with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yava-code/Tessera-1B-Nano" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/yava-code/Tessera-1B-Nano
- SGLang
How to use yava-code/Tessera-1B-Nano with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-1B-Nano" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-1B-Nano" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use yava-code/Tessera-1B-Nano with Docker Model Runner:
docker model run hf.co/yava-code/Tessera-1B-Nano
Tessera-1B-Nano
A Next Concept Prediction checkpoint from Paragon Intelligence Labs: an independent,
small-scale replication of ConceptLM on HuggingFaceTB/SmolLM2-360M. The model keeps
ordinary next-token generation and adds a causal, product-quantized concept path over
4-token chunks: token states are pooled into concepts,
product-quantized, processed by causal concept blocks, and the predicted next concept
is fed back into the token decoder.
TL;DR
This checkpoint comes from a token-matched comparison against the unchanged HuggingFaceTB/SmolLM2-360M backbone: both arms consumed the same tokens of the same packed corpus, in the same order, from the same initialization, with no restarts and no NaNs.
- Token loss is neutral: the matched baseline is within run noise (see the
whitepaper and
results/fineweb-edu/README.mdin the ncp-smol repository for the exact numbers). - The concept channel is causally used: zeroing predicted concept feedback at inference costs +0.1050 nats of held-out NTP loss.
- The codebook is rich: no low-entropy shortcut, usage grows monotonically over the run (growth curves are in the repository whitepaper).
- Full reading of the interventions:
docs/whitepaper.mdin the ncp-smol repo.
Evaluation
| Metric | Value |
|---|---|
| Held-out NTP loss | 2.5142 |
| Held-out perplexity | 12.3568 |
| Zero feedback delta | 0.1050 |
| Shuffled feedback delta | 0.0008 |
| Codebook perplexity / usage | 7.5477 / 85.6% |
| Training tokens | 999,948,288 |
| Tracked compute estimate | $30.27 |
Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.
Architecture
- chunk size: 4
- product code: 15 segments x 64 entries
- causal concept blocks: 2
- injection point: before token decoder block 2
- NCP target: next continuous concept
- loss:
L_ntp + 1 L_ncp + 1 L_vq
This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.
Training data and provenance
The matched corpus and the training record are pinned in the ncp-smol repository:
packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw
intervention evaluation JSON (also shipped in this repository as eval.json and
metrics.jsonl).
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-1B-Nano")
model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano", trust_remote_code=True)
References
- ConceptLM: https://arxiv.org/abs/2602.08984
- NCP-ArchPreview: https://arxiv.org/abs/2609.10715
- Downloads last month
- -
Model tree for yava-code/Tessera-1B-Nano
Base model
HuggingFaceTB/SmolLM2-360M