Arcee Trinity Nano

Trinity Nano Base Pre Anneal

Trinity-Nano-Base-Pre-Anneal is an Arcee AI 6B MoE model with 1B active parameters. It is the small-sized model in our new Trinity family, a series of open-weight models for enterprise and tinkerers alike.

This base model is a pre-anneal checkpoint captured at Adam LR: 0.002, Muon LR: 0.001 before starting learning rate decay on a high-quality data mix. While this checkpoint was not exposed to the anneal phase mix containing high proportions of math and code content, it has been trained on significant amounts of such data. This checkpoint is not suitable for chatting or general use without further finetuning and should be trained for your specific domain before use.


Trinity-Nano-Base-Pre-Anneal is trained on 8.8T tokens gathered and curated through a key partnership with Datology, building upon the excellent dataset we used on AFM-4.5B with additional math and code.

Training was performed on a cluster of 512 H200 GPUs powered by Prime Intellect using HSDP parallelism.

More details, including key architecture decisions, can be found on our blog here


Model Details

  • Model Architecture: AfmoeForCausalLM
  • Parameters: 6B, 1B active
  • Experts: 128 total, 8 active, 1 shared
  • Context length: 4K
  • Learning rate during pretraining:
    • adam_lr = 0.0002
    • muon_lr = 0.001
  • Training Tokens: 8.8T
  • License: Apache 2.0

Powered by Datology

Try out our reasoning tune of our medium-sized Trinity Mini model

Trinity Mini is available today on openrouter:

https://openrouter.ai/arcee-ai/trinity-mini

curl -X POST "https://openrouter.ai/v1/chat/completions" \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arcee-ai/trinity-mini",
    "messages": [
      {
        "role": "user",
        "content": "What are some fun things to do in New York?"
      }
    ]
  }'

License

Trinity-Nano-Base-Pre-Anneal is released under the Apache-2.0 license.

Downloads last month
-
Safetensors
Model size
6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arcee-ai/Trinity-Nano-Base-Pre-Anneal

Finetunes
1 model

Collection including arcee-ai/Trinity-Nano-Base-Pre-Anneal