--- title: Qwen2.5-Coder ZeroGPU API emoji: 🧑‍💻 colorFrom: indigo colorTo: blue sdk: gradio sdk_version: 5.6.0 app_file: app.py pinned: false license: apache-2.0 --- # Qwen2.5-Coder-7B — OpenAI-compatible API on ZeroGPU Serves `Qwen/Qwen2.5-Coder-7B-Instruct` (swappable) behind an OpenAI-compatible `/v1/chat/completions` endpoint, protected by an API key and rate-limited to 5 requests/minute by default. ## 1. Deploy 1. Create a new Space → SDK: **Gradio** → Hardware: **ZeroGPU**. 2. Upload `app.py` and `requirements.txt` (and this `README.md`) to the Space repo. 3. Go to **Settings → Variables and secrets → New secret**: - Name: `API_KEY` - Value: a long random string, e.g. generate one locally with: ```bash python3 -c "import secrets; print(secrets.token_urlsafe(32))" ``` - **Do not use anything guessable.** This is the only thing standing between "just your app" and "anyone on the internet" using your GPU quota. 4. (Optional) Add more secrets/variables if you want to override defaults: - `RATE_LIMIT_PER_MINUTE` (default `5`) - `MODEL_ID` (default `Qwen/Qwen2.5-Coder-7B-Instruct`) - `GPU_DURATION_SECONDS` (default `90`) 5. The Space will build and start. Since the Space is **public**, the *page* is visible to anyone, but the API itself refuses every request that doesn't carry your `API_KEY` as a Bearer token — see the "Why public Space + secret API key is safe" section below. ## 2. Use it with the OpenAI SDK ```python from openai import OpenAI client = OpenAI( base_url="https://-.hf.space/v1", api_key="", ) resp = client.chat.completions.create( model="Qwen/Qwen2.5-Coder-7B-Instruct", # informational; server ignores and uses whatever is loaded messages=[{"role": "user", "content": "Write a Rust function that reverses a linked list."}], max_tokens=512, ) print(resp.choices[0].message.content) ``` Find your exact Space URL under the Space's **Settings** tab or in the address bar when viewing the Space (it's `https://-.hf.space`). Note: `stream=True` is **not** implemented in this version — if you need token-by-token streaming, that's a follow-up (SSE handling adds real complexity with ZeroGPU's process-per-call model, so it was left out of a first version by design, not by accident). ## 3. Admin panel (switch the active model) Visit `https:///admin` in a browser. Enter your `API_KEY` in the password field to unlock the panel, then pick a model from the dropdown (or type any Hugging Face model ID) and click **Load this model**. The swap is **in-memory only** — it resets to `MODEL_ID` (or the default) if the Space restarts or goes to sleep. If you want a swap to survive restarts, change the `MODEL_ID` secret/variable instead. ## 4. Where to put the API key in *your own* application - **Never** put it directly in client-side/browser JavaScript, a public GitHub repo, or a Colab notebook you might share — anything that ships to a browser or gets committed is effectively public. - **Local scripts / backend services:** put it in an environment variable and read it with `os.environ`, e.g. a `.env` file that is in your `.gitignore`: ``` # .env (add this file to .gitignore!) QWEN_API_KEY=the-long-random-string-you-generated ``` ```python import os from dotenv import load_dotenv load_dotenv() api_key = os.environ["QWEN_API_KEY"] ``` - **A server you control (VPS, another HF Space, etc.):** use that platform's secret manager (HF Spaces secrets, systemd `EnvironmentFile`, Docker `--env-file`, etc.) — same principle, never in source code. - **If it ever leaks** (committed by accident, pasted somewhere, etc.): go back to Settings → Variables and secrets, delete it, and create a new one. The old key stops working immediately once removed and the Space restarts. ## 5. Why "public Space + secret API key" is safe The Space repo/page being public just means anyone can *see the code and visit the page*. It does not mean anyone can *call the API*: every request to `/v1/chat/completions` is checked against the `API_KEY` secret before any GPU time is spent (see `require_api_key` in `app.py`). Requests without a valid key get `401` before `generate()` — and therefore before ZeroGPU — is ever invoked. Rate limiting is applied per-key on top of that, so even if your key were somehow guessed, usage is capped at `RATE_LIMIT_PER_MINUTE` requests/minute. ## 6. Notes / limitations - ZeroGPU only works with the Gradio SDK, so the OpenAI-compatible layer is implemented as a FastAPI app mounted alongside Gradio in the same process, not as a separate Docker/FastAPI Space. - `usage.prompt_tokens` / `completion_tokens` in responses are placeholder `-1` values, not real counts — the server doesn't currently tokenize for counting purposes outside generation. If your client code does billing math off these fields, don't rely on them yet. - Rate limiting is in-memory and per-process. If the Space restarts, the window resets. This is intentional for a single-instance hobby deployment and is not a distributed rate limiter.