caps-chatbot-internal / recommendations.md
atwine's picture
Update recommendations.md with DeepInfra vs Novita comparison
de369d1
|
Raw
History Blame Contribute Delete
4.67 kB
# Infrastructure Recommendations for Sanyu
This document captures the current thinking on hosting, model sizing, cost, and
local testing options for the Sanyu peer-support chatbot. It is a living note
for the team to review before any code or Space changes are made.
---
## 1. Current architecture
- **LLM:** `Qwen/Qwen2.5-72B-Instruct` is called remotely via Hugging Face
Inference Providers (currently routed through **Novita**).
- **Space workload:** The Hugging Face Space only runs the Gradio UI, a small
`sentence-transformers` embedding model, and the RAG index. The LLM itself
does **not** run on the Space.
- **Space hardware:** Free **CPU basic** (2 vCPU / 16 GB RAM) is sufficient for
the current setup.
> **Important:** Upgrading the Space GPU will **not** make chatbot replies
> faster. The response time is dominated by the remote API round-trip and the
> time it takes the 72B model to generate tokens.
---
## 2. Should we upgrade the Space hardware?
Only if the goal is to run the LLM **inside** the Space instead of via API.
| Model size | Minimum Space GPU | Notes |
|------------|---------------------|-------|
| 7B / 8B | Nvidia T4 (16 GB VRAM) | We already saw 8B quality is too weak for the full v2 prompt. |
| 14B | Nvidia L4 (24 GB VRAM) or L40S (48 GB) | Borderline for the v2 prompt; acceptable for testing. |
| 32B | L40S (48 GB) or A100 80GB | Much better quality, but costly. |
| 72B (current) | Nvidia A100 80GB | Very expensive ($2.50+/hr) and usually overkill. |
**Recommendation:** Stay on the free **CPU basic** Space for the API-based
setup.
---
## 3. Reducing API cost without changing hardware
On the Hugging Face Inference Providers page for
`Qwen/Qwen2.5-72B-Instruct`, only two providers are listed for this exact
model:
| Provider | Input $/1M | Output $/1M | Context | Latency | Throughput | Tools | Structured |
|------------|------------|-------------|---------|---------|------------|-------|------------|
| **DeepInfra** | **$0.36** | **$0.40** | 32,768 | **0.31 s** | 31 t/s | Yes | No |
| Novita (current) | $0.38 | $0.40 | 32,000 | 2.00 s | 33 t/s | Yes | No |
**Recommendation:** Switch the provider from **Novita** to **DeepInfra**.
It is slightly cheaper (~5% less on input) and has dramatically lower latency
(0.31 s vs 2.0 s). Throughput is similar, and the context window is slightly
larger.
> Note: Prices like Hyperbolic (~$0.12/$0.30) and Nebius (~$0.13/$0.40) are
> available through those providers’ direct APIs, but they are not listed on
> the HF Inference Providers page for this model, so using them would require
> a separate API key and client setup.
Other credit-saving tips:
- **HF PRO** ($9/month) gives more monthly inference credits than the free tier.
- Never test 500B+ models (e.g. `zai-org/GLM-5.2`) on a free account — they burn
the $0.10 free allowance in minutes.
- Keep the system prompt efficient; the current v2 prompt is ~9K tokens, which
is the largest cost driver per request.
---
## 4. Local testing option
If you want to test without using API credits or while offline:
- Use **Ollama** (Windows/Mac/Linux) with a local GGUF model.
- Suggested starting points: **Qwen2.5-7B** or **Qwen2.5-14B**.
- Hardware needs:
- 7B: ~6 GB RAM (very slow on CPU; a GPU is strongly recommended).
- 14B: ~10 GB RAM / 24 GB VRAM for comfortable speed.
- Caveat: a 7B/8B model will **not** reproduce the 72B quality. We already
observed that `Llama-3.1-8B-Instruct` failed to follow the full v2 prompt,
outputting literal placeholders like `[name]`. Local testing is best for UI,
logging, and RAG plumbing checks, not for prompt-quality validation.
To wire a local Ollama instance into the app, you would point the client to its
OpenAI-compatible endpoint:
```python
client = InferenceClient(model="http://localhost:11434/v1/")
```
and pass the local model name (e.g. `model="qwen2.5:7b"`) in the chat call.
---
## 5. Streaming
- Streaming is implemented and pushed.
- It does **not** reduce total generation time, but it improves perceived
responsiveness because the user sees tokens appear as they are generated.
---
## 6. Open issues to close before scaling
- **Crash after two chats:** We still need the Space log output to diagnose the
root cause. Upgrade hardware is unlikely to fix this unless it is an OOM
crash.
- **Template/placeholder leakage in responses:** If the bot still prints literal
`[name]`, XML tags, or system-prompt fragments, capture the exact response
and the Space logs. This is usually a model-size or provider-routing issue,
not a Space hardware issue.
---
*Last updated: 2026-07-07*