Spaces:
Runtime error
A newer version of the Gradio SDK is available: 6.22.0
Infrastructure Recommendations for Sanyu
This document captures the current thinking on hosting, model sizing, cost, and local testing options for the Sanyu peer-support chatbot. It is a living note for the team to review before any code or Space changes are made.
1. Current architecture
- LLM:
Qwen/Qwen2.5-72B-Instructis called remotely via Hugging Face Inference Providers (currently routed through Novita). - Space workload: The Hugging Face Space only runs the Gradio UI, a small
sentence-transformersembedding model, and the RAG index. The LLM itself does not run on the Space. - Space hardware: Free CPU basic (2 vCPU / 16 GB RAM) is sufficient for the current setup.
Important: Upgrading the Space GPU will not make chatbot replies faster. The response time is dominated by the remote API round-trip and the time it takes the 72B model to generate tokens.
2. Should we upgrade the Space hardware?
Only if the goal is to run the LLM inside the Space instead of via API.
| Model size | Minimum Space GPU | Notes |
|---|---|---|
| 7B / 8B | Nvidia T4 (16 GB VRAM) | We already saw 8B quality is too weak for the full v2 prompt. |
| 14B | Nvidia L4 (24 GB VRAM) or L40S (48 GB) | Borderline for the v2 prompt; acceptable for testing. |
| 32B | L40S (48 GB) or A100 80GB | Much better quality, but costly. |
| 72B (current) | Nvidia A100 80GB | Very expensive ($2.50+/hr) and usually overkill. |
Recommendation: Stay on the free CPU basic Space for the API-based setup.
3. Reducing API cost without changing hardware
On the Hugging Face Inference Providers page for
Qwen/Qwen2.5-72B-Instruct, only two providers are listed for this exact
model:
| Provider | Input $/1M | Output $/1M | Context | Latency | Throughput | Tools | Structured |
|---|---|---|---|---|---|---|---|
| DeepInfra | $0.36 | $0.40 | 32,768 | 0.31 s | 31 t/s | Yes | No |
| Novita (current) | $0.38 | $0.40 | 32,000 | 2.00 s | 33 t/s | Yes | No |
Recommendation: Switch the provider from Novita to DeepInfra. It is slightly cheaper (~5% less on input) and has dramatically lower latency (0.31 s vs 2.0 s). Throughput is similar, and the context window is slightly larger.
Note: Prices like Hyperbolic (
$0.12/$0.30) and Nebius ($0.13/$0.40) are available through those providers’ direct APIs, but they are not listed on the HF Inference Providers page for this model, so using them would require a separate API key and client setup.
Other credit-saving tips:
- HF PRO ($9/month) gives more monthly inference credits than the free tier.
- Never test 500B+ models (e.g.
zai-org/GLM-5.2) on a free account — they burn the $0.10 free allowance in minutes. - Keep the system prompt efficient; the current v2 prompt is ~9K tokens, which is the largest cost driver per request.
4. Local testing option
If you want to test without using API credits or while offline:
- Use Ollama (Windows/Mac/Linux) with a local GGUF model.
- Suggested starting points: Qwen2.5-7B or Qwen2.5-14B.
- Hardware needs:
- 7B: ~6 GB RAM (very slow on CPU; a GPU is strongly recommended).
- 14B: ~10 GB RAM / 24 GB VRAM for comfortable speed.
- Caveat: a 7B/8B model will not reproduce the 72B quality. We already
observed that
Llama-3.1-8B-Instructfailed to follow the full v2 prompt, outputting literal placeholders like[name]. Local testing is best for UI, logging, and RAG plumbing checks, not for prompt-quality validation.
To wire a local Ollama instance into the app, you would point the client to its OpenAI-compatible endpoint:
client = InferenceClient(model="http://localhost:11434/v1/")
and pass the local model name (e.g. model="qwen2.5:7b") in the chat call.
5. Streaming
- Streaming is implemented and pushed.
- It does not reduce total generation time, but it improves perceived responsiveness because the user sees tokens appear as they are generated.
6. Open issues to close before scaling
- Crash after two chats: We still need the Space log output to diagnose the root cause. Upgrade hardware is unlikely to fix this unless it is an OOM crash.
- Template/placeholder leakage in responses: If the bot still prints literal
[name], XML tags, or system-prompt fragments, capture the exact response and the Space logs. This is usually a model-size or provider-routing issue, not a Space hardware issue.
Last updated: 2026-07-07