QuickferenceV1 / README.md
dumbbutt0's picture
Add task templates, searchable model catalog, and social links
8f03aec verified
|
Raw History Blame Contribute Delete
4.63 kB
---
title: Quickference
emoji: "⚡"
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: "6.28.0"
app_file: app.py
python_version: "3.12"
pinned: false
short_description: OAuth-authenticated streaming chat with Hugging Face.
hf_oauth: true
hf_oauth_scopes:
- inference-api
hf_oauth_expiration_minutes: 60
---
# Quickference
Sign in with Hugging Face, authorize inference, then select a hosted chat model.
Press Check Model to commit the selection. Typing alone does not change it.
Editable system prompts, temperature, streaming, and three task templates are included.
The searchable model dropdown loads the router catalog; Check Model commits your choice.
## Authentication
OAuth uses a one-time state tied to a browser cookie, PKCE S256, and the HF
userinfo endpoint. Credentials and authenticated identity are kept on the server.
The browser receives an opaque HttpOnly, Secure, SameSite=Lax login cookie.
Sign-in opens the Space directly in a new tab to avoid third-party cookie issues.
Gradio is mounted at /app and protected by authentication plus ownership checks
on session hashes and queued event IDs. Neither a username header nor a Gradio
session hash establishes identity. Unauthenticated requests receive 401;
cross-login session/event access receives 403. The login expires in at most one hour.
Sign out invalidates that login and clears its model/client entries. Requests
already sent to a provider cannot be recalled; the stream stops exposing results
after local logout/expiry and provider work is closed when the next chunk arrives.
HF revocation can cause inference to fail; browser sessions are not continuously
introspected against HF between requests. No automatic refresh token is retained.
The Space needs the OAuth settings above. HF supplies OAUTH_CLIENT_ID,
OAUTH_CLIENT_SECRET and SPACE_HOST. The app does not use the deployment HF_TOKEN
as a visitor credential and does not provide a production fake-login fallback.
All web requests must reach the same server process: this in-memory implementation
is for a single replica; restarts invalidate logins. Distributed session storage,
rate limiting, and broader security review remain separate production work.
## Run and test
Install requirements.txt in the deploy tool environment as well as the Space.
```bash
python deploy.py --test
python deploy.py --check-zerogpu
```
Local tests use an explicitly injected fake OAuth provider and fake inference;
they test real HTTP routes without reading real credentials or making paid calls.
They do not validate the real HF OAuth callback, consent, private-Space access,
or actual inference billing. Verify those on a private Space with two real accounts.
The normal local app starts a sign-in landing page; without Spaces OAuth
configuration it refuses login instead of pretending authentication succeeded.
## Build and deploy
```bash
python deploy.py
python deploy.py --upload
```
The first command only generates app.py, requirements.txt, README.md.
Upload creates private Spaces by default; --public explicitly opts into a public
new Space. Existing visibility is unchanged. Upload credentials are read only
by the deployment command from HF_TOKEN or a local getpass prompt.
Generated files never contain that credential. No deployment happens during tests.
## Programmatic smoke tests
```bash
python deploy.py --test-space your-name/QuickferenceV1
```
This opt-in command sends up to three real prompts and may use inference credits.
Supply an OAuth access token via QUICKFERENCE_TEST_TOKEN or the masked prompt.
The API validates it using HF userinfo on every HTTP request. A separate
QUICKFERENCE_SPACE_TOKEN may be needed to access a private Space itself.
The application header X-Quickference-OAuth must not be logged by a reverse proxy.
Do not pass a deployment/write token as the inference OAuth token.
## ZeroGPU
This app calls remote Inference Providers. It does not load a model on the Space,
and selecting ZeroGPU does not turn those calls into local GPU inference.
--check-zerogpu is a static assessment, not certification of hosted GPU execution.
See https://huggingface.co/docs/hub/spaces-zerogpu for a local-model implementation.
## Troubleshooting
Use python deploy.py --troubleshoot for the reference.
An OAuth configuration error means the Space metadata/environment needs checking.
401 means sign-in is missing, expired, or invalid; 403 means session/event ownership
or origin checks failed. Clear the browser tab and sign in again when switching accounts.
Successful login does not guarantee gated-model access, inference credits, or provider availability.