trinity-nano-stheno โ€” GGUF

realoperator42/trinity-nano-stheno-lora (attention-only LoRA, r=16, alpha=32) merged into arcee-ai/Trinity-Nano-Preview and converted to GGUF.

Architecture is afmoe, so you need a llama.cpp build that includes AFMOE support.

The embedded chat template is the adapter's corrected ChatML, with the training-only {% generation %} markers removed so --jinja parses cleanly.

llama-cli -m trinity-nano-stheno-Q4_K_M.gguf --jinja -cnv

Files

File Size
trinity-nano-stheno-Q4_K_M.gguf 3.8 GB
trinity-nano-stheno-Q8_0.gguf 6.5 GB
Downloads last month
38
GGUF
Model size
6B params
Architecture
afmoe
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for realoperator42/trinity-nano-stheno-GGUF