fable5-inference / README.md
Leon4gr45's picture
Upload folder using huggingface_hub
c7271de verified
|
Raw
History Blame Contribute Delete
3.88 kB
metadata
title: Ornith 1.0 Inference API (CPU Agentic)
emoji: πŸ¦…
colorFrom: green
colorTo: blue
sdk: docker
pinned: false
license: mit

Qwythos 9B Claude Mythos 5 1M MTP Inference API β€” CPU Agentic (v2.0.0)

CPU-only LLM inference for agentic / tool-calling workloads. A thin FastAPI reverse proxy in front of a single, fully-optimized llama.cpp server that speaks the native OpenAI API β€” including function/tool calling, streaming, and the model's real chat template.

Model: Qwythos 9B Claude Mythos 5 1M MTP (Q6_K GGUF) β€” a reasoning model (<think>).

What changed from v1.0 (and why)

v1.0 ran two llama.cpp instances and hand-rolled the prompt. On a shared CPU that was a net loss: 2 instances Γ— 4 threads oversubscribed the cores (per-token speed collapsed ~4–5Γ— under load) while doubling RAM, and the hand-built System:/User:/Assistant: prompt bypassed the chat template (looping output, no tool calling).

v2.0:

  • One instance, all cores, concurrency via llama.cpp continuous-batching slots (--parallel). No thread oversubscription, one model in RAM.
  • Native OpenAI passthrough β€” tools, tool_choice, response_format, streaming and the model's jinja chat template all handled by llama.cpp (grammar-constrained tool calls instead of brittle string parsing).
  • CPU tuning: flash-attention, q4_0 K/V cache, --threads-batch, native SIMD.
  • Pooled HTTP client; no per-request health pre-flight; real SSE streaming.

Performance (honest, CPU basic tier β‰ˆ 2 vCPU)

CPU generation of a 9B model is memory-bandwidth bound. Expect ~1–2 tok/s single-request on the free 2-vCPU tier. For snappy agents, use CPU-Upgrade (8 vCPU) or a GPU tier β€” agent latency multiplies by step count.

Configuration (env vars)

Variable Default Description
NUM_INSTANCES 1 llama.cpp processes (keep 1 on shared CPU)
CPU_THREADS all vCPUs generation threads
PARALLEL 2 continuous-batching slots (concurrent requests/subagents)
CONTEXT_SIZE 32768 total context (per-slot = ctx / parallel)
KV_CACHE_QUANT_K q8_0 K cache type β€” K is more quality-sensitive than V
KV_CACHE_QUANT_V q4_0 V cache type (needs flash-attn when quantized)
KV_CACHE_QUANT – legacy: overrides both K and V
CACHE_REUSE 256 min tokens for prompt-prefix reuse (0 = off)
FLASH_ATTN true flash attention (enables V-cache quant)
MMAP_ENABLED / MLOCK_ENABLED true / false memory mapping / locking
REASONING_FORMAT auto how <think> reasoning is surfaced

API

Chat completion (OpenAI-compatible)

curl -X POST https://Leon4gr45-fable5-inference.hf.space/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwythos-9b","messages":[{"role":"user","content":"Hello"}],"max_tokens":64}'

Tool / function calling (native)

curl -X POST https://Leon4gr45-fable5-inference.hf.space/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model":"qwythos-9b",
    "messages":[{"role":"user","content":"What is the weather in Paris?"}],
    "tools":[{"type":"function","function":{
      "name":"get_weather",
      "description":"Get current weather for a city",
      "parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}],
    "tool_choice":"auto"
  }'

The response contains a proper tool_calls array with grammar-constrained JSON arguments β€” feed the tool result back as a role:"tool" message to continue the chain.

Streaming

Add "stream": true to receive SSE token deltas.

Health / config

curl https://Leon4gr45-fable5-inference.hf.space/health
curl https://Leon4gr45-fable5-inference.hf.space/v1/config

License

MIT