A newer version of the Gradio SDK is available: 6.22.0
title: Qwen2.5-Coder ZeroGPU API
emoji: π§βπ»
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 5.6.0
app_file: app.py
pinned: false
license: apache-2.0
Qwen2.5-Coder-7B β OpenAI-compatible API on ZeroGPU
Serves Qwen/Qwen2.5-Coder-7B-Instruct (swappable) behind an OpenAI-compatible
/v1/chat/completions endpoint, protected by an API key and rate-limited to
5 requests/minute by default.
1. Deploy
- Create a new Space β SDK: Gradio β Hardware: ZeroGPU.
- Upload
app.pyandrequirements.txt(and thisREADME.md) to the Space repo. - Go to Settings β Variables and secrets β New secret:
- Name:
API_KEY - Value: a long random string, e.g. generate one locally with:
python3 -c "import secrets; print(secrets.token_urlsafe(32))" - Do not use anything guessable. This is the only thing standing between "just your app" and "anyone on the internet" using your GPU quota.
- Name:
- (Optional) Add more secrets/variables if you want to override defaults:
RATE_LIMIT_PER_MINUTE(default5)MODEL_ID(defaultQwen/Qwen2.5-Coder-7B-Instruct)GPU_DURATION_SECONDS(default90)
- The Space will build and start. Since the Space is public, the page
is visible to anyone, but the API itself refuses every request that
doesn't carry your
API_KEYas a Bearer token β see the "Why public Space + secret API key is safe" section below.
2. Use it with the OpenAI SDK
from openai import OpenAI
client = OpenAI(
base_url="https://<your-username>-<your-space-name>.hf.space/v1",
api_key="<the same value you put in the API_KEY secret>",
)
resp = client.chat.completions.create(
model="Qwen/Qwen2.5-Coder-7B-Instruct", # informational; server ignores and uses whatever is loaded
messages=[{"role": "user", "content": "Write a Rust function that reverses a linked list."}],
max_tokens=512,
)
print(resp.choices[0].message.content)
Find your exact Space URL under the Space's Settings tab or in the
address bar when viewing the Space (it's https://<owner>-<space-name>.hf.space).
Note: stream=True is not implemented in this version β if you need
token-by-token streaming, that's a follow-up (SSE handling adds real
complexity with ZeroGPU's process-per-call model, so it was left out of a
first version by design, not by accident).
3. Admin panel (switch the active model)
Visit https://<your-space-url>/admin in a browser. Enter your API_KEY
in the password field to unlock the panel, then pick a model from the
dropdown (or type any Hugging Face model ID) and click Load this model.
The swap is in-memory only β it resets to MODEL_ID (or the default)
if the Space restarts or goes to sleep. If you want a swap to survive
restarts, change the MODEL_ID secret/variable instead.
4. Where to put the API key in your own application
- Never put it directly in client-side/browser JavaScript, a public GitHub repo, or a Colab notebook you might share β anything that ships to a browser or gets committed is effectively public.
- Local scripts / backend services: put it in an environment variable
and read it with
os.environ, e.g. a.envfile that is in your.gitignore:# .env (add this file to .gitignore!) QWEN_API_KEY=the-long-random-string-you-generatedimport os from dotenv import load_dotenv load_dotenv() api_key = os.environ["QWEN_API_KEY"] - A server you control (VPS, another HF Space, etc.): use that
platform's secret manager (HF Spaces secrets, systemd
EnvironmentFile, Docker--env-file, etc.) β same principle, never in source code. - If it ever leaks (committed by accident, pasted somewhere, etc.): go back to Settings β Variables and secrets, delete it, and create a new one. The old key stops working immediately once removed and the Space restarts.
5. Why "public Space + secret API key" is safe
The Space repo/page being public just means anyone can see the code and
visit the page. It does not mean anyone can call the API: every request
to /v1/chat/completions is checked against the API_KEY secret before
any GPU time is spent (see require_api_key in app.py). Requests without
a valid key get 401 before generate() β and therefore before ZeroGPU β
is ever invoked. Rate limiting is applied per-key on top of that, so even
if your key were somehow guessed, usage is capped at
RATE_LIMIT_PER_MINUTE requests/minute.
6. Notes / limitations
- ZeroGPU only works with the Gradio SDK, so the OpenAI-compatible layer is implemented as a FastAPI app mounted alongside Gradio in the same process, not as a separate Docker/FastAPI Space.
usage.prompt_tokens/completion_tokensin responses are placeholder-1values, not real counts β the server doesn't currently tokenize for counting purposes outside generation. If your client code does billing math off these fields, don't rely on them yet.- Rate limiting is in-memory and per-process. If the Space restarts, the window resets. This is intentional for a single-instance hobby deployment and is not a distributed rate limiter.