Spaces:
Sleeping
Sleeping
|
Download README.md from dumbbutt0/QuickferenceV1: direct link, hf CLI and curl.
- Browser
- Download file 4.63 kB
-
https://huggingface.co/spaces/dumbbutt0/QuickferenceV1/resolve/main/README.md
- Command line
-
hf download hf://spaces/dumbbutt0/QuickferenceV1/README.md
-
curl -L -o README.md https://huggingface.co/spaces/dumbbutt0/QuickferenceV1/resolve/main/README.md
4.63 kB
| title: Quickference | |
| emoji: "⚡" | |
| colorFrom: yellow | |
| colorTo: gray | |
| sdk: gradio | |
| sdk_version: "6.28.0" | |
| app_file: app.py | |
| python_version: "3.12" | |
| pinned: false | |
| short_description: OAuth-authenticated streaming chat with Hugging Face. | |
| hf_oauth: true | |
| hf_oauth_scopes: | |
| - inference-api | |
| hf_oauth_expiration_minutes: 60 | |
| # Quickference | |
| Sign in with Hugging Face, authorize inference, then select a hosted chat model. | |
| Press Check Model to commit the selection. Typing alone does not change it. | |
| Editable system prompts, temperature, streaming, and three task templates are included. | |
| The searchable model dropdown loads the router catalog; Check Model commits your choice. | |
| ## Authentication | |
| OAuth uses a one-time state tied to a browser cookie, PKCE S256, and the HF | |
| userinfo endpoint. Credentials and authenticated identity are kept on the server. | |
| The browser receives an opaque HttpOnly, Secure, SameSite=Lax login cookie. | |
| Sign-in opens the Space directly in a new tab to avoid third-party cookie issues. | |
| Gradio is mounted at /app and protected by authentication plus ownership checks | |
| on session hashes and queued event IDs. Neither a username header nor a Gradio | |
| session hash establishes identity. Unauthenticated requests receive 401; | |
| cross-login session/event access receives 403. The login expires in at most one hour. | |
| Sign out invalidates that login and clears its model/client entries. Requests | |
| already sent to a provider cannot be recalled; the stream stops exposing results | |
| after local logout/expiry and provider work is closed when the next chunk arrives. | |
| HF revocation can cause inference to fail; browser sessions are not continuously | |
| introspected against HF between requests. No automatic refresh token is retained. | |
| The Space needs the OAuth settings above. HF supplies OAUTH_CLIENT_ID, | |
| OAUTH_CLIENT_SECRET and SPACE_HOST. The app does not use the deployment HF_TOKEN | |
| as a visitor credential and does not provide a production fake-login fallback. | |
| All web requests must reach the same server process: this in-memory implementation | |
| is for a single replica; restarts invalidate logins. Distributed session storage, | |
| rate limiting, and broader security review remain separate production work. | |
| ## Run and test | |
| Install requirements.txt in the deploy tool environment as well as the Space. | |
| ```bash | |
| python deploy.py --test | |
| python deploy.py --check-zerogpu | |
| ``` | |
| Local tests use an explicitly injected fake OAuth provider and fake inference; | |
| they test real HTTP routes without reading real credentials or making paid calls. | |
| They do not validate the real HF OAuth callback, consent, private-Space access, | |
| or actual inference billing. Verify those on a private Space with two real accounts. | |
| The normal local app starts a sign-in landing page; without Spaces OAuth | |
| configuration it refuses login instead of pretending authentication succeeded. | |
| ## Build and deploy | |
| ```bash | |
| python deploy.py | |
| python deploy.py --upload | |
| ``` | |
| The first command only generates app.py, requirements.txt, README.md. | |
| Upload creates private Spaces by default; --public explicitly opts into a public | |
| new Space. Existing visibility is unchanged. Upload credentials are read only | |
| by the deployment command from HF_TOKEN or a local getpass prompt. | |
| Generated files never contain that credential. No deployment happens during tests. | |
| ## Programmatic smoke tests | |
| ```bash | |
| python deploy.py --test-space your-name/QuickferenceV1 | |
| ``` | |
| This opt-in command sends up to three real prompts and may use inference credits. | |
| Supply an OAuth access token via QUICKFERENCE_TEST_TOKEN or the masked prompt. | |
| The API validates it using HF userinfo on every HTTP request. A separate | |
| QUICKFERENCE_SPACE_TOKEN may be needed to access a private Space itself. | |
| The application header X-Quickference-OAuth must not be logged by a reverse proxy. | |
| Do not pass a deployment/write token as the inference OAuth token. | |
| ## ZeroGPU | |
| This app calls remote Inference Providers. It does not load a model on the Space, | |
| and selecting ZeroGPU does not turn those calls into local GPU inference. | |
| --check-zerogpu is a static assessment, not certification of hosted GPU execution. | |
| See https://huggingface.co/docs/hub/spaces-zerogpu for a local-model implementation. | |
| ## Troubleshooting | |
| Use python deploy.py --troubleshoot for the reference. | |
| An OAuth configuration error means the Space metadata/environment needs checking. | |
| 401 means sign-in is missing, expired, or invalid; 403 means session/event ownership | |
| or origin checks failed. Clear the browser tab and sign in again when switching accounts. | |
| Successful login does not guarantee gated-model access, inference credits, or provider availability. | |