Instructions to use yava-code/Tessera-1B-Nano-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yava-code/Tessera-1B-Nano-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="yava-code/Tessera-1B-Nano-Base")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-1B-Nano-Base") model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano-Base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use yava-code/Tessera-1B-Nano-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yava-code/Tessera-1B-Nano-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/yava-code/Tessera-1B-Nano-Base
- SGLang
How to use yava-code/Tessera-1B-Nano-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-1B-Nano-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yava-code/Tessera-1B-Nano-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yava-code/Tessera-1B-Nano-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use yava-code/Tessera-1B-Nano-Base with Docker Model Runner:
docker model run hf.co/yava-code/Tessera-1B-Nano-Base
Tessera-1B-Nano-Base
The matched NTP-only baseline of the Tessera-1B-Nano comparison from Paragon
Intelligence Labs: the unchanged HuggingFaceTB/SmolLM2-360M backbone,
continued-pretrained without a concept path. It exists so the concept-arm result
is interpretable.
TL;DR
This arm is the control of a token-matched comparison: it consumed the same tokens of the same packed corpus, in the same order, from the same initialization as the concept arm, with no restarts and no NaNs.
- Token loss is the reference: the final numbers below are the baseline the concept arm is compared against.
- The comparison outcome and the intervention readings live in the whitepaper at
docs/whitepaper.mdin the ncp-smol repository.
Evaluation
| Metric | Value |
|---|---|
| Held-out NTP loss | 2.5135 |
| Held-out perplexity | 12.3485 |
| Training tokens | 999,948,288 |
| Tracked compute estimate | $24.88 |
Intervention deltas are increases in held-out NTP loss relative to normal predicted concept feedback, evaluated on identical batches.
Architecture
- chunk size: 4
- product code: 15 segments x 64 entries
- causal concept blocks: 2
- injection point: before token decoder block 2
- NCP target: next continuous concept
- loss:
L_ntp + 1 L_ncp + 1 L_vq
This is a compact ConceptLM-style implementation, not an 8.9B NCP-ArchPreview replica. It omits iterative residual coding, cross-scale residual connections, and the large-scale training recipe.
Training data and provenance
The matched corpus and the training record are pinned in the ncp-smol repository:
packed-cache SHA256 hashes, the complete metric log, trainer state, and the raw
intervention evaluation JSON (also shipped in this repository as eval.json and
metrics.jsonl).
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("yava-code/Tessera-1B-Nano-Base")
model = AutoModelForCausalLM.from_pretrained("yava-code/Tessera-1B-Nano-Base")
References
- ConceptLM: https://arxiv.org/abs/2602.08984
- NCP-ArchPreview: https://arxiv.org/abs/2609.10715
- Downloads last month
- -
Model tree for yava-code/Tessera-1B-Nano-Base
Base model
HuggingFaceTB/SmolLM2-360M