caps-chatbot-internal / recommendations.md
atwine's picture
Update recommendations.md with DeepInfra vs Novita comparison
de369d1
|
Raw
History Blame Contribute Delete
4.67 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade

Infrastructure Recommendations for Sanyu

This document captures the current thinking on hosting, model sizing, cost, and local testing options for the Sanyu peer-support chatbot. It is a living note for the team to review before any code or Space changes are made.


1. Current architecture

  • LLM: Qwen/Qwen2.5-72B-Instruct is called remotely via Hugging Face Inference Providers (currently routed through Novita).
  • Space workload: The Hugging Face Space only runs the Gradio UI, a small sentence-transformers embedding model, and the RAG index. The LLM itself does not run on the Space.
  • Space hardware: Free CPU basic (2 vCPU / 16 GB RAM) is sufficient for the current setup.

Important: Upgrading the Space GPU will not make chatbot replies faster. The response time is dominated by the remote API round-trip and the time it takes the 72B model to generate tokens.


2. Should we upgrade the Space hardware?

Only if the goal is to run the LLM inside the Space instead of via API.

Model size Minimum Space GPU Notes
7B / 8B Nvidia T4 (16 GB VRAM) We already saw 8B quality is too weak for the full v2 prompt.
14B Nvidia L4 (24 GB VRAM) or L40S (48 GB) Borderline for the v2 prompt; acceptable for testing.
32B L40S (48 GB) or A100 80GB Much better quality, but costly.
72B (current) Nvidia A100 80GB Very expensive ($2.50+/hr) and usually overkill.

Recommendation: Stay on the free CPU basic Space for the API-based setup.


3. Reducing API cost without changing hardware

On the Hugging Face Inference Providers page for Qwen/Qwen2.5-72B-Instruct, only two providers are listed for this exact model:

Provider Input $/1M Output $/1M Context Latency Throughput Tools Structured
DeepInfra $0.36 $0.40 32,768 0.31 s 31 t/s Yes No
Novita (current) $0.38 $0.40 32,000 2.00 s 33 t/s Yes No

Recommendation: Switch the provider from Novita to DeepInfra. It is slightly cheaper (~5% less on input) and has dramatically lower latency (0.31 s vs 2.0 s). Throughput is similar, and the context window is slightly larger.

Note: Prices like Hyperbolic ($0.12/$0.30) and Nebius ($0.13/$0.40) are available through those providers’ direct APIs, but they are not listed on the HF Inference Providers page for this model, so using them would require a separate API key and client setup.

Other credit-saving tips:

  • HF PRO ($9/month) gives more monthly inference credits than the free tier.
  • Never test 500B+ models (e.g. zai-org/GLM-5.2) on a free account — they burn the $0.10 free allowance in minutes.
  • Keep the system prompt efficient; the current v2 prompt is ~9K tokens, which is the largest cost driver per request.

4. Local testing option

If you want to test without using API credits or while offline:

  • Use Ollama (Windows/Mac/Linux) with a local GGUF model.
  • Suggested starting points: Qwen2.5-7B or Qwen2.5-14B.
  • Hardware needs:
    • 7B: ~6 GB RAM (very slow on CPU; a GPU is strongly recommended).
    • 14B: ~10 GB RAM / 24 GB VRAM for comfortable speed.
  • Caveat: a 7B/8B model will not reproduce the 72B quality. We already observed that Llama-3.1-8B-Instruct failed to follow the full v2 prompt, outputting literal placeholders like [name]. Local testing is best for UI, logging, and RAG plumbing checks, not for prompt-quality validation.

To wire a local Ollama instance into the app, you would point the client to its OpenAI-compatible endpoint:

client = InferenceClient(model="http://localhost:11434/v1/")

and pass the local model name (e.g. model="qwen2.5:7b") in the chat call.


5. Streaming

  • Streaming is implemented and pushed.
  • It does not reduce total generation time, but it improves perceived responsiveness because the user sees tokens appear as they are generated.

6. Open issues to close before scaling

  • Crash after two chats: We still need the Space log output to diagnose the root cause. Upgrade hardware is unlikely to fix this unless it is an OOM crash.
  • Template/placeholder leakage in responses: If the bot still prints literal [name], XML tags, or system-prompt fragments, capture the exact response and the Space logs. This is usually a model-size or provider-routing issue, not a Space hardware issue.

Last updated: 2026-07-07