flow2 / docs /OMNIROUTE_PROVIDER.md
AndrianBalanescu
fix: require auth only when FLOW_API_KEY is explicitly set
434c049
|
Raw
History Blame Contribute Delete
2.12 kB

A newer version of the Gradio SDK is available: 6.25.0

Upgrade

OmniRoute Provider Runbook

The Space is an OpenAI-compatible upstream:

Base URL: https://abalanescu-flow.hf.space/v1
Models:   GET /v1/models
Chat:     POST /v1/chat/completions
Auth:     Authorization: Bearer <FLOW_API_KEY>

Discover the exact model ID

curl -sS "$FLOW_BASE/v1/models" \
  -H "Authorization: Bearer $FLOW_API_KEY" | jq -r '.data[].id'

Use the exact returned ID. The verified models are:

Qwen3.8-27B-Q6_K.gguf
gemma-4-26B-A4B-it-ultra-uncensored-heretic.i1-Q4_K_M.gguf
gemma-4-26B-A4B-it-ultra-uncensored-heretic.i1-Q6_K.gguf

Smoke test

Test Qwen3.8-27B Q6

curl -sS --max-time 180 "$FLOW_BASE/v1/chat/completions" \
  -H "Authorization: Bearer $FLOW_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-27B-Q6_K.gguf",
    "messages": [{"role":"user","content":"Tell me a joke."}],
    "temperature": 0.7,
    "max_tokens": 64
  }'

Test Gemma 4 26B Heretic

curl -sS --max-time 180 "$FLOW_BASE/v1/chat/completions" \
  -H "Authorization: Bearer $FLOW_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26B-A4B-it-ultra-uncensored-heretic.i1-Q4_K_M.gguf",
    "messages": [{"role":"user","content":"Reply exactly OMNIROUTE_FLOW_OK"}],
    "temperature": 0.1,
    "max_tokens": 16
  }'

OmniRoute configuration checklist

Add a provider named hf-flow-zerogpu in the OmniRoute provider configuration. The exact UI/API differs by the deployed OmniRoute version, so do not invent a config endpoint. Set:

  • Base URL: https://abalanescu-flow.hf.space/v1
  • API key: reference a homelab secret, never a literal token in git
  • Model: exact ID from /v1/models
  • Timeout: 180 seconds for cold starts and model loading
  • Concurrency: 1 initially, because one cached llama.cpp model is shared
  • Retries: transport/5xx only; do not replay a completed GPU request
  • Health check: /v1/models with the bearer token

After adding it, run the smoke test through OmniRoute and record the request ID, latency, status, and selected upstream. Only then enable it in a combo.