TR-hash-tiny / README.md
Pacific-i64's picture
Serve chat without forced thinking
efc50ab verified
|
Raw
History Blame Contribute Delete
1.87 kB
---
title: TR-Hash Tiny
emoji: 🧭
colorFrom: purple
colorTo: green
sdk: docker
app_port: 7860
pinned: false
license: apache-2.0
models:
- AETHORIA-AI/TR-HASH-MoE-200M-160B-SFT
---
# TR-Hash Tiny · TR-HASH MoE 200M Full SFT
OpenAI-compatible chat API serving the PIQA-selected epoch-3 root weights from
[`AETHORIA-AI/TR-HASH-MoE-200M-160B-SFT`](https://huggingface.co/AETHORIA-AI/TR-HASH-MoE-200M-160B-SFT)
with [`Complexity-ML/TR-Hash-i64`](https://github.com/Complexity-ML/TR-Hash-i64).
This is a full-parameter SFT, not a LoRA adapter. All 201.2M parameters were
trained for three epochs on the audited 300K SFT v2 mixture. Epoch 3 scores
**68.01% PIQA accuracy** and **69.21% length-normalized accuracy** on the full
1,838-example validation set in the MLX FP16 cross-check.
The launcher intentionally downloads only the release files at the model-repo
root. The three large resumable `step_*` training folders are excluded. The
bundled `chat_template.jinja` and tokenizer are loaded beside the single root
`model.safetensors`.
```bash
curl https://pacific-i64-tr-hash-tiny.hf.space/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "tr-hash-moe-200m",
"messages": [{"role": "user", "content": "Explain why 17 × 23 − 14 equals 377."}],
"max_tokens": 1024,
"temperature": 0.15,
"top_k": 10,
"top_p": 0.85,
"repetition_penalty": 1.02,
"frequency_penalty": 0.0,
"stream": false
}'
```
The CPU deployment uses continuous batching, paged KV caching, prefix caching,
streaming and rolling context compaction at 1,024 prompt tokens. It also
supports optional dynamic INT8 packing of linear layers. Set the Space variable
`CPU_INT8=false` to serve floating-point weights.
Endpoints: `/health`, `/v1/models`, `/v1/completions`,
`/v1/chat/completions`, `/v1/metrics`, `/v1/monitor`, `/v1/experts`.