LukeFP's picture
Run on ZeroGPU: add @spaces.GPU entry point
49a8dc7
|
Raw History Blame Contribute Delete
2.8 kB

Deploying to LukeFP/Physh_Classification

The Space repo lives at ~/code/2026.7/Physh_Classification.

1. Add the token secret

google/embeddinggemma-300m is gated. Accept the Gemma license while signed in, create a read token, then on the Space page: Settings β†’ Variables and secrets β†’ New secret, name HF_TOKEN, value the token. Without it the Space boots fine but the first classification fails with a 401.

2. Hardware

On the free tier, Gradio Spaces run on ZeroGPU, which stops the container at startup unless it finds at least one @spaces.GPU function β€” the No @spaces.GPU function detected during startup error. infer() in app.py carries that decorator, so ZeroGPU is satisfied.

Constraints ZeroGPU imposes, and how app.py meets them:

Constraint Handling
import spaces must precede import torch It is the first import in app.py
Nothing may touch CUDA outside a @GPU function Models load with device="cpu"; .to(device) happens inside infer()
Return values cross a process boundary infer() returns plain list[float], never CUDA tensors
One GPU allocation per call, with a duration budget @GPU(duration=60); the model is already resident, so only the encode runs

CPU basic (a PRO perk) also works with this code unchanged β€” spaces is an optional import and the device is chosen from torch.cuda.is_available().

3. Push

cd ~/code/2026.7/Physh_Classification
git push origin main

The build takes a few minutes, most of it pip install torch.

4. First checks

  • Predictions look like noise, or nothing clears the threshold. Almost certainly the embedding prompt. Open Advanced and try the other two formats; the one matching your training pipeline gives confident, coherent labels. Once you know which, set DEFAULT_PROMPT at the top of app.py. (~/code/2026/embedding_title_abstract likely has the answer.)
  • Error mentioning a gated repo, or a 401. HF_TOKEN is missing, wrong, or the account behind it hasn't accepted the Gemma license.
  • First request is slow, later ones fast. Expected β€” EmbeddingGemma loads lazily on first use so the Space boots quickly. Cached after that.

Updating later

Retraining only needs a push to LukeFP/physh_topic_supervised_classifier; the Space picks up new weights on its next restart. Only change this repo if the filenames change β€” they're the constants at the top of app.py.

Local smoke test

Runs the real checkpoints through the full chain with a stubbed embedder, so it needs no token and no model download:

PHYSH_WEIGHTS_DIR=~/code/2026.7/physh_topic_supervised_classifier python test_local.py