Spaces:
Sleeping
Sleeping
File size: 5,552 Bytes
f65e025 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 | # Deployment guide
Architecture: **frontend β Vercel**, **backend β a Docker container host** (Render
blueprint included). The two talk over HTTPS; the frontend proxies `/api/*` to the
backend via a Next.js rewrite, so the browser only ever sees the Vercel origin.
```
Browser ββ> Vercel (Next.js) ββ/api/* rewriteββ> Backend container (FastAPI + Docling)
βββ persistent disk /data
βββ models baked into image
```
Why the backend can't be serverless: it bundles Docling + (CPU) torch + ~850 MB of
layout/table/OCR models and needs real RAM/disk. Plan for **2 GB minimum, 4 GB
recommended**.
---
## 1. Backend β Render (recommended)
The image bakes the models in, runs as non-root, exposes a health check, and binds
`$PORT`. Everything is described in [`render.yaml`](../render.yaml).
1. **Push this repo to GitHub** (Render deploys from Git):
```bash
git remote add origin https://github.com/<you>/document-agent.git
git push -u origin main
```
2. In the **Render dashboard β New β Blueprint**, select the repo. Render reads
`render.yaml` and provisions a Docker web service + a 10 GB disk at `/data`.
3. When prompted, set the secret env vars:
- `OPENAI_API_KEY` β your key
- `GEMINI_API_KEY` β optional (leave blank to disable Gemini)
4. After the first deploy, note the URL (e.g. `https://document-agent-backend.onrender.com`)
and verify: `GET <url>/api/health` β `{"status":"ok", ...}`.
5. Set `CORS_ORIGINS` to your Vercel domain (see step 2 below) and redeploy.
First build is slow (downloads torch + models). Subsequent deploys reuse layers.
### Free option β Hugging Face Spaces (no card required)
Render's free plan is 512 MB RAM with no disk and can't run this image. **HF Spaces
free CPU tier gives 16 GB RAM** and runs Docker, so it fits β the only loss is
persistence (free Spaces have no disk, so uploads/DB reset on restart/rebuild).
This repo ships an HF-ready **root `Dockerfile`** (builds the backend with the repo
root as context, runs as UID 1000) and the HF config in the **`README.md`
frontmatter** (`sdk: docker`, `app_port: 8000`).
1. Create a free account at https://huggingface.co, then **New β Space**:
- **SDK: Docker** (blank template), name e.g. `document-agent`, visibility your choice.
2. Push this repo to the Space's git remote:
```bash
git remote add space https://huggingface.co/spaces/<hf-user>/document-agent
git push space main # auth: HF username + an HF access token as password
```
(Create a write token at https://huggingface.co/settings/tokens.)
3. In the Space: **Settings β Variables and secrets**, add as **secrets**:
- `OPENAI_API_KEY` (freshly rotated) and/or `GEMINI_API_KEY`
- `DEFAULT_LLM_PROVIDER` (`openai` or `gemini`), `ENABLE_OCR` (`true`/`false`)
4. The Space builds (slow first time: torch + ~850 MB models). When live, the URL is
`https://<hf-user>-document-agent.hf.space`; verify `GET <url>/api/health`.
5. Use that URL as `NEXT_PUBLIC_API_BASE` on Vercel (step 2 below).
> Tip: if free RAM is tight on large scans, set `ENABLE_OCR=false` to cut memory.
### Other alternatives
- **Fly.io**: `fly launch` from `./backend` (it detects the Dockerfile), add a volume
(`fly volumes create data`) mounted at `/data`, set secrets with `fly secrets set`.
- **Railway**: new service from the Dockerfile, add a volume at `/data`, set the env vars.
---
## 2. Frontend β Vercel
The frontend is already env-driven β it needs exactly one variable pointing at the
backend.
```bash
cd frontend
vercel # first run: link/create the project
# set the backend URL for all environments:
vercel env add NEXT_PUBLIC_API_BASE
# value: https://document-agent-backend.onrender.com (no trailing slash)
vercel --prod # deploy
```
`NEXT_PUBLIC_API_BASE` drives the `/api/:path*` rewrite in
[`next.config.mjs`](../frontend/next.config.mjs), so uploads, page images, exports and
the chat stream all flow through the same origin β no CORS needed in the common case.
> Streaming note: chat uses SSE through the Vercel rewrite. If a host buffers the
> stream, switch to direct calls by setting `BASE = process.env.NEXT_PUBLIC_API_BASE`
> in `frontend/lib/api.ts` and make sure `CORS_ORIGINS` on the backend lists the
> Vercel domain.
---
## 3. Verify locally first (optional but recommended)
Build and run the real production image before shipping:
```bash
docker compose up --build
# backend: http://localhost:8000/api/health
```
`docker-compose.yml` mounts a named volume at `/data` and reads `backend/.env`, so it
mirrors production (persistent uploads + sqlite + rendered pages).
---
## Operational notes
- **Persistence**: `/data` holds uploads, the SQLite DB, the numpy vector store and
rendered page images. Keep it on a mounted disk/volume β without one, data resets on
every deploy.
- **Models**: baked into the image (not in `/data`), so they survive redeploys with no
re-download.
- **Secrets**: only ever via host env vars. `backend/.env` is git-ignored.
- **Scaling**: state lives on local disk + SQLite, so run a **single instance**. To
scale horizontally, move the DB to Postgres and the vector store to pgvector
(`services/vectorstore.py` is a drop-in seam) and uploads to object storage.
- **Cost guard**: the backend holds your LLM key behind a public URL. Consider adding
auth or a rate limit before sharing widely.
|