RandomZ / DEPLOY_HF_SPACES.md
StormShadow308's picture
Document TestBunny Space deploy and replacing HF Docker template.
c27181b
|
Raw
History Blame Contribute Delete
14.4 kB
# Deploying to Hugging Face Spaces (free tier, Docker SDK)
This document is the canonical guide for the **free-of-cost** deployment
target the project is configured for. The image is built from the repo's
existing `Dockerfile`, runs as UID 1000 (Spaces requirement), and listens on
port 7860 (Spaces convention). All other configuration is via environment
variables / Space secrets — the source code is identical across deployments.
> **Cost reality check**
> The Space itself is free. **OpenAI API calls are not.** Every report
> generation, embedding, vision call, and notes upload bills against the
> `OPENAI_API_KEY` you provide. Typical cost per generated report on
> `gpt-4o-mini` + `text-embedding-3-small` is roughly £0.04 – £0.25.
---
## 1. Prerequisites
| Item | Where |
|---|---|
| Hugging Face account | <https://huggingface.co/join> |
| Personal **OpenAI API key** | <https://platform.openai.com/api-keys> |
| Locally installed `git` | already present in this repo |
You don't need a paid HF subscription. The free CPU Space tier is enough.
**Spaces in this project:**
| Space | Git remote | URL |
|-------|------------|-----|
| `StormShadow308/RICS` | `hf` or `space` | https://huggingface.co/spaces/StormShadow308/RICS |
| `StormShadow308/TestBunny` | `testbunny` | https://huggingface.co/spaces/StormShadow308/TestBunny |
If HF created **TestBunny** with the default `app.py` + tiny `Dockerfile` template, **replace it entirely** by pushing this repo (see §2b). Do not merge the template with Report Genius — use our root `Dockerfile` and `app.main:app`.
---
## 2. Create the Space (one time)
1. Go to <https://huggingface.co/new-space>.
2. Fill in:
- **Owner / Space name** — anything, e.g. `your-handle/rics-report-genius`.
- **License** — match the repo (`mit`).
- **SDK** — choose **Docker****Blank**.
- **Space hardware**`CPU basic` (free).
- **Visibility** — Public is fine; Private also works (no behaviour change).
3. Click **Create Space**. HF will create an empty Git repo at
`https://huggingface.co/spaces/<owner>/<space-name>`.
### 2b. Deploy Report Genius to an existing Space (e.g. TestBunny)
From your machine (HF token as password when prompted):
```powershell
cd "D:\RICS UK Project"
git remote add testbunny https://huggingface.co/spaces/StormShadow308/TestBunny
git push testbunny experiment-lab:main
```
Or with the [HF CLI](https://huggingface.co/docs/huggingface_hub/guides/cli):
```powershell
hf auth login
git push testbunny HEAD:main
```
The Space must use **Docker** SDK (not Gradio). Our `README.md` already has the required YAML (`sdk: docker`, `app_port: 7860`). The image listens on **7860** and runs `uvicorn app.main:app` — not the template `app.py` world handler.
---
## 3. Add the YAML frontmatter to `README.md`
HF Spaces reads metadata from the **first** YAML block of `README.md` in the
Space repo. Open the project's `README.md` and prepend exactly this block at
the very top (above the existing `# Report Genius AI` heading):
```yaml
---
title: RICS Report Genius
emoji: 🏠
colorFrom: blue
colorTo: indigo
sdk: docker
app_port: 7860
pinned: false
license: mit
short_description: RICS-style RAG report generator (FAISS + GPT-4o-mini).
---
```
That YAML must be the **first** characters of the file — no blank line, no
BOM, nothing before the opening `---`.
> Do this in a separate commit on a `huggingface` branch (or directly on the
> branch you push to the Space) so the GitHub README isn't littered with HF
> metadata permanently. Locally the file renders fine on both platforms.
---
## 3b. Pre-flight before a demo
After the Space is running:
```bash
python scripts/preflight_hf.py --base https://<your-space-subdomain>.hf.space --tenant your-tenant-id
```
All checks must pass (especially `parallel_multi_section` and `indexed_chunk_count` after upload).
---
## 4. Set the Space secrets
In the Space's **Settings → Variables and secrets**, add at minimum:
| Type | Name | Value |
|---|---|---|
| **Secret** | `OPENAI_API_KEY` | your OpenAI key (sk-…) — **required** |
| **Secret** | `TENANT_SECRET_KEY` | any 32+ char random string (recommended) |
**Production AI profile (recommended on Spaces):** the app auto-detects `SPACE_ID` and
disables ingest-time LLM sanitisation so uploads do not burn your OpenAI quota before
report generation. Set these **Variables** so inspector + photo vision run on the main
`/generate` path:
| Type | Name | Value |
|---|---|---|
| Variable | `INSPECTOR_TOOL_AGENT` | `true` |
| Variable | `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY` | `true` |
| Variable | `SECTION_PHOTO_VISION_ENABLED` | `true` |
| Variable | `SECTION_PHOTO_VISION_MODEL` | `gpt-4o` (optional; better vision than mini) |
| Variable | `PRIMARY_GENERATE_PIPELINE` | `agentic` |
| Variable | `ENABLE_RAG_UPLOAD_SANITISATION` | `true` (regex redaction; LLM sanitisation stays off on Space) |
| Variable | `RAG_SANITISATION_USE_LLM` | `false` |
| Secret | `HF_TOKEN` | optional — faster `all-MiniLM-L6-v2` download on cold start |
> **Required vs recommended.** The container ships with `DEV_MODE=false`,
> which engages a **startup guard** in `app/main.py`. If `OPENAI_API_KEY` is
> empty/whitespace, the container will **refuse to start** with a clear
> `RuntimeError` in the Logs tab — that is intentional. The agentic inspector
> would otherwise silently fall back to crude keyword heuristics and reports
> would be low quality without you knowing.
>
> `TENANT_SECRET_KEY` is currently inert (no signing path consumes it) but
> the guard logs a loud `WARNING` if it equals one of the sentinel
> placeholders (`""`, `"dev-secret-change-me"`, `"change-me-in-production"`).
> Set it now so the day someone wires it into tenant-token signing, you don't
> have to redeploy.
Optional — override Dockerfile defaults if you want different behaviour:
| Type | Name | Default | Notes |
|---|---|---|---|
| Variable | `KNOWLEDGE_BASE_ENABLED` | `false` | Set `true` *only* if you also bake KB PDFs into the image. The startup guard logs a `WARNING` when this is `false` in production — that's expected on the HF free tier (KB source PDFs are gitignored), so the warning is informational. |
| Variable | `INSPECTOR_TOOL_AGENT` | `true` | Disable to force the legacy fixed pipeline. |
| Variable | `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY` | `true` | Inspector + vision on `POST /generate`. Slower than standard-only batch; use `false` if you need sub-10m full reports without vision. |
| Variable | `GENERATION_SLA_SECONDS` | `600` | Product target (10 min). `SPACE_ID` auto-enables async parallel sections. |
| Variable | `SECTION_PHOTO_VISION_ENABLED` | `true` | Vision analysis of section photos. |
| Variable | `SECTION_PHOTO_ANALYZE_ON_UPLOAD` | `true` | Run vision after each photo upload (cached per section). |
| Variable | `PERSONALISED_STYLE_RAG_ENABLED` | `true` | Private per-tenant style library (default product mode). |
| Variable | `ENABLE_RAG_UPLOAD_SANITISATION` | `true` | Regex redaction before indexing (forced on when personalised RAG is on). **Do not** set `RAG_SANITISATION_USE_LLM=true` on HF. |
| Variable | `RAG_SANITISATION_USE_LLM` | `false` | Auto-disabled when `SPACE_ID` is set. |
| Variable | `UVICORN_WORKERS` | `1` | Correct for the 1-vCPU HF Spaces free tier. On bigger hosts, set to `(2 * CPU cores) + 1`. |
| Variable | `RATE_LIMIT_GENERATE_RPM` | `20` | Per-tenant /generate rate limit. |
| Variable | `RATE_LIMIT_READ_RPM` | `120` | Per-tenant read rate limit. |
You should **never** put `OPENAI_API_KEY` in a `Variable` — that exposes it
in the public Space metadata. Always use `Secret`.
---
## 5. Push the code
The Space is just a git remote. From the project root:
```bash
# one-time: add the Space as a remote (use https or ssh)
git remote add space https://huggingface.co/spaces/<owner>/<space-name>
# push (force-with-lease the first time if the Space already has an empty README)
git push space HEAD:main
```
If you forked the repo on GitHub and want a 1-click sync, you can also
configure the Space to **Sync from GitHub** in **Settings → Repository**
instead of pushing manually.
### Pre-push image size sanity check (optional)
HF Spaces enforces a hard image-size cap (~10 GB on free Docker Spaces). With
the current Dockerfile the image lands around **1.2–1.5 GB**. To verify
locally before pushing (requires Docker with ≥ 5 GB free disk):
```bash
docker build --target production -t rics-genius:check .
docker images rics-genius:check --format '{{.Size}}'
# expected: ~1.2-1.5 GB. If >5 GB, rebuild with `--no-cache` and check that
# .dockerignore is being honoured (Behrang RICS Documents/, *.db, etc.).
```
---
## 6. First boot — what to expect
HF builds the Docker image (5–8 min the first time). When it's ready you'll
see logs in the **Logs** tab roughly like:
```
INFO: Started server process [1]
WARNING app.main — KNOWLEDGE_BASE_ENABLED=false in production: ... (expected on HF free tier)
INFO app.main — Initialising database…
INFO app.main — Pre-warming vector store…
INFO app.vectorstore.faiss_wrapper — Created new LangChain FAISS index at /home/user/.report_genius/faiss_index
INFO app.main — Pre-warming embedding model…
INFO app.main — RICS inspector runtime: mode=openai_tool_agent key=True flag=True
INFO app.main — Startup complete.
INFO: Uvicorn running on http://0.0.0.0:7860
```
The `KNOWLEDGE_BASE_ENABLED=false` warning is **expected** on the free tier —
the source PDFs are gitignored, so there is nothing to seed. Tenants supply
their own corpus via uploads.
If instead you see:
```
RuntimeError: OPENAI_API_KEY is required when DEV_MODE=false. ...
```
…the startup guard fired because `OPENAI_API_KEY` is missing or whitespace —
re-save it in **Settings → Secrets**, the Space restarts automatically.
The Space's public URL (`https://<owner>-<space-name>.hf.space`) now serves
the FastAPI app. `/health` returns 200, `/` returns the frontend. The
container also has a Docker-level `HEALTHCHECK` that polls `/health` every
30 s, so the Space dashboard reflects readiness independently of port-binding.
---
## 7. Persistence — the trade-off you implicitly made
**Free CPU Spaces have an ephemeral filesystem.** Every restart (idle,
crash, image rebuild, manual restart) wipes:
- `dev.db` — tenants, reports, document rows, section text.
- `~/.report_genius/faiss_index` — the embedding index.
- `~/.report_genius/uploads` — uploaded PDFs/DOCX.
What survives across restarts: **only the source code in the image**.
That matches the **hybrid** option you picked: tenant uploads are
expected to be ephemeral on the free tier. If you later want true
persistence, you have three options, in order of effort:
1. **Pay for HF Persistent Storage** ($5/mo for 20 GB on `/data`).
Add `Variable: HF_HOME=/data` and change `DATABASE_URL`,
`FAISS_INDEX_PATH`, `UPLOAD_DIR` to point under `/data`. Smallest code
change, fastest deploy.
2. **Sync FAISS + DB to a private HF Dataset on shutdown.**
Requires writing a small `app/services/hf_index_sync.py` and hooking it
into FastAPI's `lifespan` shutdown. ~50 lines + `huggingface_hub` dep.
Free, but you lose anything written between snapshots.
3. **Move state out of the container** — use a hosted Postgres (Supabase
free tier, Neon) for `DATABASE_URL` and an external object store for
uploads. Most robust, biggest plumbing change.
If you want any of these wired up, ask and I'll do it as a follow-up.
---
## 8. Troubleshooting
### "Space exited with error" during build
Check the **Logs** tab. The most common causes are:
- **Build context too big** — make sure `.dockerignore` is committed.
`Behrang RICS Documents/`, `RAW Context/`, `*.db`, `.claude/`, `app/tests/`
and similar must NOT enter the build context.
- **Image too large** — see §5 "Pre-push image size sanity check". The
current Dockerfile installs the project **non-editable, without `[dev]`
extras**, so a clean image is ~1.2–1.5 GB. If yours is much bigger,
something is leaking into the context.
### "RuntimeError: OPENAI_API_KEY is required when DEV_MODE=false"
This is the **production startup guard** doing its job. Re-add the
`OPENAI_API_KEY` secret in **Settings → Secrets** (whitespace doesn't count
— it must be a real `sk-…` key). Save; the Space restarts automatically.
### App boots but `/` 502s
HF's reverse proxy expects port `7860`. Confirm:
- `app_port: 7860` in the README YAML.
- Container CMD listens on `7860` (the default in the Dockerfile when
`PORT` is unset).
### "Internal Server Error" on first generation
With the new startup guard the container won't even boot if
`OPENAI_API_KEY` is missing — so a 500 here is more likely an OpenAI rate
limit, expired key, or model availability issue. Check `/health` under
`rics_inspector.openai_api_key_configured` (should be `true`) and watch the
Logs tab for the actual upstream OpenAI error.
### KB-related logs you can ignore
```
WARNING app.main — KNOWLEDGE_BASE_ENABLED=false in production: ...
```
This is the production guard reminding you that the agentic inspector's
KB-grounding tool returns nothing in this deployment. Expected on the HF
free tier; see §6.
### Container restart drops in-flight requests
Shouldn't happen with the current Dockerfile — `CMD` uses an `exec`-form
shell wrapper so `uvicorn` becomes PID 1 and receives `SIGTERM` directly. If
you see clients getting connection-resets on every redeploy, verify the
running image's CMD with `docker inspect <image> --format '{{.Config.Cmd}}'`:
it must contain `exec uvicorn`.
### Out of memory
The free CPU tier has 16 GB RAM but a slow CPU. The app idles around
~600 MB plus the FAISS index. If you hit OOM after heavy usage, restart
the Space (it clears the index). The deeper fix is option (1) or (3) in
§7.
---
## 9. Updating the Space
```bash
git push space HEAD:main
```
HF auto-rebuilds and reboots. Build time is faster after the first deploy
because the base image and venv are cached.
To force a clean rebuild, use **Settings → Factory rebuild**.
---
## 10. Reverting the deployment
In **Settings → Delete this Space**. The container, image cache, and
ephemeral data are removed. Your repo is untouched.