Tercet-R-1.0 / README.md
kerzgrr's picture
Link Tercet-R-1.1 in the model family table
6f2f485 verified
|
Raw
History Blame Contribute Delete
5.72 kB
metadata
license: apache-2.0
language:
  - en
tags:
  - text-generation
  - causal-lm
  - pytorch
  - sft
  - instruction-tuned
  - chat
  - reasoning
  - tool-use
  - hybrid
  - gated-deltanet
  - gqa
  - tercet
pipeline_tag: text-generation
library_name: tiny_gdn
datasets:
  - HuggingFaceFW/fineweb-edu
  - HuggingFaceTB/smoltalk2
  - HuggingFaceTB/smoltalk
  - nvidia/Llama-Nemotron-Post-Training-Dataset
  - open-thoughts/OpenThoughts3-1.2M
  - teknium/OpenHermes-2.5
  - open-r1/Mixture-of-Thoughts
  - CohereLabs/aya_dataset
  - NousResearch/hermes-function-calling-v1
  - simplescaling/s1K-1.1
  - allenai/tulu-3-sft-personas-instruction-following
  - Salesforce/xlam-function-calling-60k
  - THUDM/LongAlign-10k
base_model: kerzgrr/Tercet-base

Tercet-R-1.0

Reasoning + tool-call chat model (~502M) — Tercet-R family

Model Stage License Base

A ~502M hybrid GDN-2 + GQA model, supervised fine-tuned for thinking and tools


What this is

Tercet-R-1.0 is the first public reasoning / tool-use checkpoint in the Tercet-R line.


Chat contract

Thinking

Each assistant turn is prefixed with a zero-loss control token:

Mode Prefix Typical body
think <|think|>\n <think>…</think> then the answer
no-think <|no_think|>\n answer only

inference.py and the local scripts/chat.py stream the <think> region live (dim yellow) and hide the control tokens.

Tool calls (SmolTalk JSON)

<tool_call>
{"name": "web_search", "arguments": {"query": "…"}}
</tool_call>

Tool results

Each observation is a tool (or user) turn prefixed with the special token:

<|tool_response|>
{observation}

inference.py --tools web_search pauses after a <tool_call>, you paste the search result, and generation continues.


Install & run

pip install torch safetensors tokenizers huggingface_hub
hf download kerzgrr/Tercet-R-1.0 inference.py --local-dir .
python inference.py
python inference.py --prompt "What is the capital of France?"
python inference.py --tools web_search
python inference.py --no-think --prompt "Reply in one sentence."

inference.py auto-downloads weights / tokenizer / tiny_gdn/ and auto-installs pinned flash-linear-attention. Git is required on PATH.

Flag Default Description
--prompt One-shot user message
--system System prompt, used verbatim
--think / --no-think think Assistant control prefix
--tools Built-in tools (web_search)
--temperature 0.7 Sampling temperature
--max-new-tokens 4096 Max generation length
--device cuda if available cuda / cpu

Interactive commands: /think /no_think /system … /reset /exit.


Model architecture

Same TinyGDN hybrid as Tercet-base (501,635,264 parameters):

Layers 32 (GDN-2 ×3 + GQA every 4th)
Hidden 1,024
MLP SwiGLU 2,624
Attention 8 Q / 2 KV, head dim 128, partial RoPE
Linear Gated DeltaNet-2, 8 heads × 128
Vocab 49,152 BPE
Context 16,384

Training

Stage Details
Base HuggingFaceFW/fineweb-eduTercet-base
Mid-SFT HuggingFaceTB/smoltalk2 Mid: Llama-Nemotron-Post-Training-Dataset + OpenThoughts3-1.2M
Instruct SFT HuggingFaceTB/smoltalk2 SFT (SmolTalk, OpenHermes-2.5, OpenThoughts3, Aya, Hermes function calling, s1K, Tulu-3 personas IF, xLAM, LongAlign, Mixture-of-Thoughts, …), seq 16,384, AdamW 5×10⁻⁵, 61.1 hours
Checkpoint optimizer step 4,500 (latest complete instruct snapshot)
Weights EMA (this repo's model.safetensors)
Val loss (EMA) 1.744 (ppl 5.72)

Limitations

  • Scale: ~502M is a research / edge model, not a frontier system
  • Requires flash-linear-attention; not GGUF / llama.cpp compatible today

Model family

Model Stage Hub
Tercet-base Pretrain kerzgrr/Tercet-base
Tercet SFT chat kerzgrr/Tercet
Tercet-R-1.0 SFT reasoning + tools this repo
Tercet-R-1.1 Stage-3 SFT (latest) kerzgrr/Tercet-R-1.1

Citation

@misc{tercetr2026,
  title={Tercet-R-1.0: A 502M Hybrid GDN-2 + GQA Reasoning Model},
  author={kerzgrr},
  year={2026},
  url={https://huggingface.co/kerzgrr/Tercet-R-1.0}
}

R is for reasoning.