metadata
title: Qwen 7B Hosted Chat
emoji: 🤖
colorFrom: blue
colorTo: indigo
sdk: docker
app_port: 7860
Qwen 2.5 7B - 4-bit Quantized
This Space runs a fully hosted Qwen 2.5 7B model using llama.cpp for CPU-friendly 4-bit quantization.
Features
- UI: Gradio ChatInterface with session memory.
- API: FastAPI endpoint available at
/api/chat. - Memory Footprint: ~4.3GB (Fits in Free Tier).
API Usage
You can send POST requests to /api/chat.
import requests
url = "https://YOUR_SPACE_URL.hf.space/api/chat"
payload = {
"messages": [
{"role": "user", "content": "Hello!"}
]
}
response = requests.post(url, json=payload)
print(response.json())