ThoxRustCoder / README.md
tommytracx's picture
deploy thoxfast 058d0e08b664
f73e85a verified
|
Raw
History Blame Contribute Delete
3.03 kB
metadata
title: THOX Interactive and Rust Coder
emoji: 
colorFrom: indigo
colorTo: purple
sdk: docker
app_port: 7860
license: apache-2.0

THOX interactive and specialist model service

This Space exposes one OpenAI-compatible endpoint:

  • thox-fast-chat: eagerly loaded, bounded, full-CPU Qwen2.5 0.5B target with a six-second queue ceiling;
  • thox-rust-coder: lazily loaded 25B/3B-active Rust-specialist target.

GET /healthz is ready only after the interactive context is loaded and has completed a one-token startup warmup. The default single-context topology gives that request the full measured container CPU quota. A bounded six-second queue absorbs overlap; operators can still opt into at most two contexts with THOX_FAST_POOL_SIZE when throughput is more important than the 15-second cold turn contract. CPU allocation comes from the cgroup quota, so it does not repeat the historical host-count oversubscription regression. The same quota is passed to llama.cpp's generation and prompt-evaluation thread pools; the batch pool is never allowed to infer a larger host CPU count and oversubscribe the constrained Space during a cold persona prefix. Streaming sends a role chunk before model evaluation so upstream time-to-first-byte watchdogs do not abandon healthy CPU work; two-second SSE comments keep the stream live during synchronous prompt evaluation. The interactive provider itself caps every request at 16 output tokens, so malformed or older callers cannot leave minutes of abandoned work. Every interactive request starts with the same neutral provider-owned system prefix warmed during startup. Caller-owned system messages follow it and remain the authoritative persona and identity. This preserves llama prefix-cache reuse across otherwise unrelated personas without replacing caller instructions; the provider does not add this fast-chat prefix to specialist requests. The interactive model is pinned to immutable model revision 9217f5db79a29953eb74d5343926648285ec7e67 and baked into the image. The health response also reports the exact 40-character ThoxRoute Git revision baked into the Space image. Production acceptance must match this value rather than inferring deployment identity from a RUNNING stage alone.

Atomic production release

From a clean checkout at the exact validated commit, supply a newly issued fine-grained token scoped to write only Thox-ai/thoxrustcoder:

export THOX_FAST_HF_DEPLOY_TOKEN='set-in-protected-shell-environment'
python ops/deploy_thoxfast_space.py \
  --expected-source "$(git rev-parse HEAD)" \
  --expected-parent 45b26461379624b63430c7c3477eff63868cfaac

The helper ignores generic cached credentials, validates token metadata, rejects the known compromised token name, stages only the runtime allowlist, injects the owning Git SHA, and refuses to upload if either source cleanliness or the exact Space parent has changed. Never place the token in the repository, command arguments, logs, or a generic HF_TOKEN variable.