QuickferenceV1 / README.md
dumbbutt0's picture
Add task templates, searchable model catalog, and social links
8f03aec verified
|
Raw History Blame Contribute Delete
4.63 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade
metadata
title: Quickference
emoji: ⚡
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
python_version: '3.12'
pinned: false
short_description: OAuth-authenticated streaming chat with Hugging Face.
hf_oauth: true
hf_oauth_scopes:
  - inference-api
hf_oauth_expiration_minutes: 60

Quickference

Sign in with Hugging Face, authorize inference, then select a hosted chat model. Press Check Model to commit the selection. Typing alone does not change it. Editable system prompts, temperature, streaming, and three task templates are included. The searchable model dropdown loads the router catalog; Check Model commits your choice.

Authentication

OAuth uses a one-time state tied to a browser cookie, PKCE S256, and the HF userinfo endpoint. Credentials and authenticated identity are kept on the server. The browser receives an opaque HttpOnly, Secure, SameSite=Lax login cookie. Sign-in opens the Space directly in a new tab to avoid third-party cookie issues. Gradio is mounted at /app and protected by authentication plus ownership checks on session hashes and queued event IDs. Neither a username header nor a Gradio session hash establishes identity. Unauthenticated requests receive 401; cross-login session/event access receives 403. The login expires in at most one hour.

Sign out invalidates that login and clears its model/client entries. Requests already sent to a provider cannot be recalled; the stream stops exposing results after local logout/expiry and provider work is closed when the next chunk arrives. HF revocation can cause inference to fail; browser sessions are not continuously introspected against HF between requests. No automatic refresh token is retained.

The Space needs the OAuth settings above. HF supplies OAUTH_CLIENT_ID, OAUTH_CLIENT_SECRET and SPACE_HOST. The app does not use the deployment HF_TOKEN as a visitor credential and does not provide a production fake-login fallback. All web requests must reach the same server process: this in-memory implementation is for a single replica; restarts invalidate logins. Distributed session storage, rate limiting, and broader security review remain separate production work.

Run and test

Install requirements.txt in the deploy tool environment as well as the Space.

python deploy.py --test
python deploy.py --check-zerogpu

Local tests use an explicitly injected fake OAuth provider and fake inference; they test real HTTP routes without reading real credentials or making paid calls. They do not validate the real HF OAuth callback, consent, private-Space access, or actual inference billing. Verify those on a private Space with two real accounts.

The normal local app starts a sign-in landing page; without Spaces OAuth configuration it refuses login instead of pretending authentication succeeded.

Build and deploy

python deploy.py
python deploy.py --upload

The first command only generates app.py, requirements.txt, README.md. Upload creates private Spaces by default; --public explicitly opts into a public new Space. Existing visibility is unchanged. Upload credentials are read only by the deployment command from HF_TOKEN or a local getpass prompt. Generated files never contain that credential. No deployment happens during tests.

Programmatic smoke tests

python deploy.py --test-space your-name/QuickferenceV1

This opt-in command sends up to three real prompts and may use inference credits. Supply an OAuth access token via QUICKFERENCE_TEST_TOKEN or the masked prompt. The API validates it using HF userinfo on every HTTP request. A separate QUICKFERENCE_SPACE_TOKEN may be needed to access a private Space itself. The application header X-Quickference-OAuth must not be logged by a reverse proxy. Do not pass a deployment/write token as the inference OAuth token.

ZeroGPU

This app calls remote Inference Providers. It does not load a model on the Space, and selecting ZeroGPU does not turn those calls into local GPU inference. --check-zerogpu is a static assessment, not certification of hosted GPU execution. See https://huggingface.co/docs/hub/spaces-zerogpu for a local-model implementation.

Troubleshooting

Use python deploy.py --troubleshoot for the reference. An OAuth configuration error means the Space metadata/environment needs checking. 401 means sign-in is missing, expired, or invalid; 403 means session/event ownership or origin checks failed. Clear the browser tab and sign in again when switching accounts. Successful login does not guarantee gated-model access, inference credits, or provider availability.