A newer version of the Gradio SDK is available: 6.22.0
title: Hy3
emoji: π»
colorFrom: purple
colorTo: yellow
sdk: gradio
sdk_version: 6.19.0
python_version: '3.13'
app_file: app.py
pinned: false
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Hy3 Chat β Gradio Server app
A drop-in Gradio app that wraps the OpenRouter Quick Start for tencent/hy3:free, using gradio.Server β the pattern from HF's "Any Custom Frontend with Gradio's Backend" β to pair a hand-written HTML/JS chat UI with Gradio's backend engine (queuing, gradio_client access, SSE streaming).
What it does, matching the Quick Start:
- Calls OpenRouter's OpenAI-compatible API (
https://openrouter.ai/api/v1). - Enables reasoning:
extra_body={"reasoning": {"enabled": True}}. - Streams the response (reasoning + answer tokens arrive live).
- Preserves
reasoning_detailsacross turns β each assistant turn stores it and sends it back unmodified, so the model continues reasoning from where it left off (the "Are you sure? Think carefully." flow from the Quick Start).
Files
app.py # gradio.Server backend: @app.api() chat() streams JSON events to the frontend
index.html # vanilla HTML/CSS/JS chat UI, talks to the backend via the Gradio JS Client
requirements.txt
Setup
pip install -r requirements.txt
export OPENROUTER_API_KEY=sk-or-v1-... # from https://openrouter.ai/keys
python app.py
Open http://127.0.0.1:7860.
How it works
Backend (app.py) β a gradio.Server (FastAPI subclass) with two routes:
@app.api() def chat(messages, model, temperature, max_tokens)β goes through Gradio's queue (concurrency-managed,gradio_client-callable). It rebuilds the OpenRouter message list, forwarding each assistant turn'sreasoning_detailsunmodified, callsclient.chat.completions.create(..., stream=True, extra_body={"reasoning": {"enabled": True}}), and yields one JSON string per event:{"type": "reasoning", "text": "..."} {"type": "content", "text": "..."} {"type": "done", "content": "...", "reasoning": "...", "reasoning_details": [...]} {"type": "error", "text": "..."}Robustness notes: the streaming delta's OpenRouter-specific fields (
reasoning/reasoning_content) live in Pydantic extra storage, so_delta_fields()mergesmodel_dump()with__pydantic_extra__to read them across SDK versions. Thedoneevent reconstructs areasoning_detailssummary object from the streamed reasoning text so the next turn can carry it forward like the non-streaming Quick Start does.@app.get("/")β plain FastAPI route servingindex.html.
Frontend (index.html) β vanilla, no build step. Connects to the backend with the Gradio JS Client:
import { Client } from "https://cdn.jsdelivr.net/npm/@gradio/client@1.7.1/dist/index.min.js";
const client = await Client.connect(window.location.origin);
const result = await client.predict("/chat", { messages, model, temperature, max_tokens });
for await (const raw of result.data) { const evt = JSON.parse(raw); /* render */ }
Because it goes through client.predict (not raw fetch), requests hit Gradio's queue β concurrency is managed, not a bare POST. The UI renders:
- user / assistant bubbles,
- a collapsible Reasoning panel that streams thinking tokens live (auto-collapses once the answer starts),
- a blinking caret while streaming,
- temperature and
max_tokenscontrols, - example chips, new-chat, and stop.
The frontend keeps its own messages array, and after each done event writes content, reasoning, and reasoning_details onto the last assistant turn so the next send forwards them back β exactly the continuation the Quick Start describes.
Notes
- Reasoning across turns: OpenRouter streams reasoning as text deltas;
app.pypackages the accumulated reasoning back into areasoning_detailssummary shape so continuation works without a non-streaming round-trip. - API key: never reaches the frontend β the backend reads
OPENROUTER_API_KEYfrom the environment and proxies requests. gradio_clientaccess: sincechatis an@app.api()endpoint, other scripts can call it too:from gradio_client import Client c = Client("http://127.0.0.1:7860") for evt in c.predict("/chat", {"messages": [...], "model": "tencent/hy3:free", "temperature": 0.9, "max_tokens": 2048}): print(evt)- Set
HTTP-Referer/X-Titleheaders on the OpenAI client if you want your app to appear on OpenRouter's leaderboards:client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key=OPENROUTER_API_KEY, default_headers={"HTTP-Referer": "https://huggingface.co/spaces/you/hy3", "X-Title": "Hy3 Chat"})