--- title: Quickference emoji: "⚡" colorFrom: yellow colorTo: gray sdk: gradio sdk_version: "6.28.0" app_file: app.py python_version: "3.12" pinned: false short_description: OAuth-authenticated streaming chat with Hugging Face. hf_oauth: true hf_oauth_scopes: - inference-api hf_oauth_expiration_minutes: 60 --- # Quickference Sign in with Hugging Face, authorize inference, then select a hosted chat model. Press Check Model to commit the selection. Typing alone does not change it. Editable system prompts, temperature, streaming, and three task templates are included. The searchable model dropdown loads the router catalog; Check Model commits your choice. ## Authentication OAuth uses a one-time state tied to a browser cookie, PKCE S256, and the HF userinfo endpoint. Credentials and authenticated identity are kept on the server. The browser receives an opaque HttpOnly, Secure, SameSite=Lax login cookie. Sign-in opens the Space directly in a new tab to avoid third-party cookie issues. Gradio is mounted at /app and protected by authentication plus ownership checks on session hashes and queued event IDs. Neither a username header nor a Gradio session hash establishes identity. Unauthenticated requests receive 401; cross-login session/event access receives 403. The login expires in at most one hour. Sign out invalidates that login and clears its model/client entries. Requests already sent to a provider cannot be recalled; the stream stops exposing results after local logout/expiry and provider work is closed when the next chunk arrives. HF revocation can cause inference to fail; browser sessions are not continuously introspected against HF between requests. No automatic refresh token is retained. The Space needs the OAuth settings above. HF supplies OAUTH_CLIENT_ID, OAUTH_CLIENT_SECRET and SPACE_HOST. The app does not use the deployment HF_TOKEN as a visitor credential and does not provide a production fake-login fallback. All web requests must reach the same server process: this in-memory implementation is for a single replica; restarts invalidate logins. Distributed session storage, rate limiting, and broader security review remain separate production work. ## Run and test Install requirements.txt in the deploy tool environment as well as the Space. ```bash python deploy.py --test python deploy.py --check-zerogpu ``` Local tests use an explicitly injected fake OAuth provider and fake inference; they test real HTTP routes without reading real credentials or making paid calls. They do not validate the real HF OAuth callback, consent, private-Space access, or actual inference billing. Verify those on a private Space with two real accounts. The normal local app starts a sign-in landing page; without Spaces OAuth configuration it refuses login instead of pretending authentication succeeded. ## Build and deploy ```bash python deploy.py python deploy.py --upload ``` The first command only generates app.py, requirements.txt, README.md. Upload creates private Spaces by default; --public explicitly opts into a public new Space. Existing visibility is unchanged. Upload credentials are read only by the deployment command from HF_TOKEN or a local getpass prompt. Generated files never contain that credential. No deployment happens during tests. ## Programmatic smoke tests ```bash python deploy.py --test-space your-name/QuickferenceV1 ``` This opt-in command sends up to three real prompts and may use inference credits. Supply an OAuth access token via QUICKFERENCE_TEST_TOKEN or the masked prompt. The API validates it using HF userinfo on every HTTP request. A separate QUICKFERENCE_SPACE_TOKEN may be needed to access a private Space itself. The application header X-Quickference-OAuth must not be logged by a reverse proxy. Do not pass a deployment/write token as the inference OAuth token. ## ZeroGPU This app calls remote Inference Providers. It does not load a model on the Space, and selecting ZeroGPU does not turn those calls into local GPU inference. --check-zerogpu is a static assessment, not certification of hosted GPU execution. See https://huggingface.co/docs/hub/spaces-zerogpu for a local-model implementation. ## Troubleshooting Use python deploy.py --troubleshoot for the reference. An OAuth configuration error means the Space metadata/environment needs checking. 401 means sign-in is missing, expired, or invalid; 403 means session/event ownership or origin checks failed. Clear the browser tab and sign in again when switching accounts. Successful login does not guarantee gated-model access, inference credits, or provider availability.