File size: 15,307 Bytes
732b14f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
# Production Report β€” RICS Report Genius AI

**Document purpose:** Single reference for production deployment, scale, AI operations, and ownership.  
**Scope:** Current codebase behaviour (FastAPI, SQLite default, FAISS, LangChain, OpenAI), target production on **AWS** with **OpenAI** retained.  
**Audience:** Engineering leads, DevOps, AI engineers, backend engineers, frontend engineers, security/compliance.

---

## Table of contents

1. [Executive summary](#1-executive-summary)  
2. [System architecture (as implemented)](#2-system-architecture-as-implemented)  
3. [Data & state](#3-data--state)  
4. [AI & LLM pipeline](#4-ai--llm-pipeline)  
5. [APIs, concurrency & failure modes](#5-apis-concurrency--failure-modes)  
6. [Security & tenancy](#6-security--tenancy)  
7. [Performance & scale](#7-performance--scale)  
8. [AWS production stack (OpenAI retained)](#8-aws-production-stack-openai-retained)  
9. [Observability & operations](#9-observability--operations)  
10. [Role responsibilities & deliverables](#10-role-responsibilities--deliverables)  
11. [Pre-launch checklist](#11-pre-launch-checklist)  
12. [Post-launch & continuous improvement](#12-post-launch--continuous-improvement)  

---

## 1. Executive summary

The application is a **FastAPI** service serving a **static frontend**, with **per-tenant RAG** (FAISS + LangChain embeddings), **OpenAI** for generation and several post-processing passes, and **async SQLite** by default. It is **fit for pilot or single-tenant–style** deployments out of the box. For **thousands of concurrent users**, production requires:

| Pillar | Gap in default / local shape | Production requirement |
|--------|------------------------------|-------------------------|
| **Data** | SQLite + local files | **PostgreSQL (RDS)** + **S3** for uploads |
| **Coordination** | In-process rate limits, per-pod FAISS/cache | **Redis (ElastiCache)** + **SQS** workers + clear vector strategy |
| **Tenancy** | `X-Tenant-ID` header only | **Authenticated identity** mapped to tenant |
| **AI reliability** | OpenAI rate limits, retries, cost | **Queues, concurrency caps**, timeouts, observability |
| **Ops** | Ephemeral disk on some hosts | **Backups**, alarms, stuck-job handling |

**OpenAI remains** the LLM provider; AWS provides **compute, data, queue, cache, secrets, and logs** β€” not a replacement for OpenAI’s API limits or model behaviour.

---

## 2. System architecture (as implemented)

### 2.1 Runtime

- **Framework:** FastAPI (`app/main.py`).
- **Entry:** Uvicorn; Docker image targets port **7860** (HF Spaces) or **8000** (local/docker-compose) per `Dockerfile` / env.
- **Startup:** DB init, vector store pre-warm, embedding client pre-warm; optional KB ingest in background; **production guard** requires `OPENAI_API_KEY` when `DEV_MODE=false`.

### 2.2 Major components

| Component | Location / behaviour |
|-----------|----------------------|
| **HTTP API** | `app/api/*` β€” reports, generate, upload, status, export, photos, etc. |
| **Generation orchestration** | `app/services/generation.py` β€” RAG, style, LLM, verify, retries, persist. |
| **Vector store** | `app/vectorstore/factory.py` β€” **FAISS** singleton, LangChain-backed. |
| **Embeddings** | `app/embeddings/factory.py` β€” default **HuggingFace** local; optional **OpenAI** via config. |
| **Prompts** | `app/generator/prompts.py` β€” LCEL vars, interference levels, token budgets. |
| **LLM adapter** | `app/generator/adapter.py` β€” OpenAI / mock. |
| **Post-process** | `app/generator/postprocess.py` β€” `async_enforce_verify`, etc. |
| **Section cache** | `app/cache/section_cache.py` β€” **filesystem** JSON keyed by hash of inputs. |
| **Rate limiting** | `app/api/rate_limit.py` β€” **in-process** sliding window per `tenant_id` from header. |

### 2.3 Generation modes (API)

- **`generate`** β€” Full pipeline (default product path; `notes_only_generation` forces standard pipeline in current code paths).
- **`proofread`** / **`enhance`** β€” Separate pipelines reusing parts of generation.

### 2.4 AI interference (product)

- **Qualitative levels:** `minimum` | `medium` | `maximum` (`interference_level` on `GenerateRequest`; maps to internal `ai_percent` bands in `app/models/schemas.py`).
- **Prompts** branch on interference level in `app/generator/prompts.py`.

---

## 3. Data & state

### 3.1 Database

- **Default:** `sqlite+aiosqlite:///./dev.db` (`app/config.py`).
- **Production risk:** SQLite **single-writer** semantics, file locking, no horizontal write scale β€” **unsuitable** for high concurrent write load.
- **Target:** **Amazon RDS PostgreSQL** (or Aurora PostgreSQL) with connection pooling (app uses `pool_size` / `max_overflow` for non-SQLite in `app/db/database.py`).

### 3.2 Uploads & artifacts

- **Default:** Local `upload_dir` on instance filesystem.
- **Production risk:** **Ephemeral** disks lose data on redeploy/restart; multi-replica **inconsistent** views of files.
- **Target:** **Amazon S3** + DB rows storing object keys; presigned URLs for client download where needed.

### 3.3 Vector index (FAISS)

- **Current:** In-process singleton; persisted under app-configured paths (e.g. user home / cache dirs).
- **Production risk:** **Not shared** across replicas; each pod has its own index unless you mount shared storage (generally discouraged for FAISS at scale).
- **Target options:**  
  - **RDS `pgvector`** β€” vectors in Postgres; one operational system.  
  - **Amazon OpenSearch** kNN β€” separate service; strong scale/filter story.  
  - **Single replica** only β€” acceptable only for early internal launch, not β€œthousands concurrent.”

### 3.4 Section result cache

- **Current:** File-based cache in `settings.cache_dir`; atomic write via temp + rename.
- **Production risk:** **Per-node** only; no cross-instance deduplication.
- **Target:** **Redis** for hot cache keys (optional) or accept cache miss on each pod.

---

## 4. AI & LLM pipeline

### 4.1 OpenAI usage (representative chain per section)

Depending on settings and branch, a **generate** path may include:

- **Notes expansion** β€” LLM when OpenAI key present and expansion not skipped (`app/generator/notes_expander.py`).
- **Main section generation** β€” LangChain LCEL + ChatOpenAI (`app/generator/adapter.py`).
- **`async_enforce_verify`** β€” Additional LLM pass when key present (`app/generator/postprocess.py`).
- **Retries** β€” Verbatim ratio, identity consistency, survey-tier hints β€” can trigger **extra** `generate_section` calls.
- **Optional LLM section validator** β€” `llm_section_validator_enabled` in `app/config.py` (default **false**); when true, extra validate + possible regenerate.
- **Photo / vision** β€” When photos exist and policy triggers (`app/services/photo_vision.py`).

**Implication:** Wall-clock and **cost scale with number of LLM round-trips**, not only β€œone GPT call per section.”

### 4.2 Embeddings

- **Default:** `prefer_local_embeddings=True` β†’ **HuggingFace** sentence-transformers on **CPU** per process (`app/embeddings/factory.py`).
- **Production trade-off:** Each replica loads model into **RAM**; cold start and memory multiply with replica count.
- **Mitigation:** Set **`prefer_local_embeddings=false`** to use **OpenAI embeddings** API (cost vs RAM); or keep HF on **GPU** workers dedicated to ingest (heavier ops design).

### 4.3 Quality & guardrails (behavioural, not infra)

- **Notes-only mode** (`notes_only_generation`, default **true**): bullets are factual authority; uploads are **reference-only** in that design; provenance suppressed accordingly.
- **Interference levels** encode how much drafting vs mapping is allowed (see `app/generator/prompts.py` and schemas).
- **Production AI engineering:** golden-set evaluation, pinned models, documented limits of each guard (what verify does / does not guarantee).

### 4.4 Cost & limits (OpenAI)

- **Account-level rate limits** and **token pricing** apply regardless of AWS scale.
- **Mitigation:** **SQS-backed workers** with **global + per-tenant concurrency caps**, exponential backoff on **429**, timeouts, circuit breaker patterns.

---

## 5. APIs, concurrency & failure modes

### 5.1 Report generation locking

- `POST /reports/{report_id}/generate` uses a **compare-and-swap** so `Report.status` is only flipped to **`generating`** if not already `generating` (`app/api/generate.py`).
- **Second concurrent** generate for same report β†’ **409** until the first completes or fails.
- **Worker crash** after CAS but before completion can leave status **`generating`** until the **stale-generation sweeper** marks the report failed (`generation_stale_sweep_seconds`, `generation_started_at` on `reports`). Runs on the API process and on the Temporal worker when enabled.

### 5.2 Background execution

- Generation runs as **`asyncio.create_task`** from the API process; errors in `run_generation` set report **`failed`** and `error_message` (`app/services/generation.py`).
- **Risk:** Process kill loses in-flight task; see stuck state above.

### 5.3 Rate limiting

- **Generate:** default **20 RPM per tenant id** (header); **Read:** **120 RPM** (`app/api/rate_limit.py`, env `RATE_LIMIT_GENERATE_RPM`, `RATE_LIMIT_READ_RPM`).
- **Limitation:** **In-process** β€” incorrect under **multiple workers/pods** (each has its own counter). Production should use **Redis** or **API Gateway** throttling.

### 5.4 HTTP status summary (client expectations)

| Code | Typical cause |
|------|----------------|
| **401** | Missing / empty `X-Tenant-ID` on non-exempt paths (`app/api/middleware.py`). |
| **404** | Report or document not found / wrong tenant. |
| **409** | Report already `generating`; or business rule conflicts elsewhere. |
| **413** | Payload too large (e.g. section text PATCH limit). |
| **422** | Validation (e.g. invalid `interference_level`). |
| **429** | Rate limit exceeded. |

---

## 6. Security & tenancy

### 6.1 Current tenant model

- **`TenantAuthMiddleware`:** requires non-empty **`X-Tenant-ID`** header; no cryptographic proof of ownership (`app/api/middleware.py`).

### 6.2 Production requirement

- **Authenticate** users (Cognito, Auth0, custom JWT, API keys behind API Gateway, etc.) and **derive** `tenant_id` from trusted claims β€” do not treat the raw header as the security boundary.

### 6.3 Secrets

- **`OPENAI_API_KEY`:** required in production when `DEV_MODE=false` (`app/main.py`).
- **`TENANT_SECRET_KEY`:** defaults trigger **warnings**; should be non-sentinel before any future signing use (`app/main.py`).

### 6.4 CORS

- `allowed_origins` default `["*"]` with credentials false in `app/main.py` β€” tighten for production web origins.

---

## 7. Performance & scale

### 7.1 Dominant bottlenecks

1. **Sequential LLM calls** per section (expand β†’ generate β†’ verify β†’ retries).  
2. **OpenAI latency + rate limits.**  
3. **Database** under concurrent writes if SQLite remains.  
4. **Per-replica FAISS + HF embeddings** β€” memory and duplicate ingest work.

### 7.2 Application-level levers (documented in config)

- Retrieval pool sizes (`hierarchical_k_*`, `retrieval_top_k`, `rerank_top_n` in `app/config.py`).
- Optional **LLM section validator** (off by default).
- **Interference level** β€” Maximum mode increases output tokens and time.

### 7.3 Architecture-level levers

- **Queue + workers** (SQS) for generate jobs.  
- **Bounded concurrency** per tenant and globally toward OpenAI.  
- **Shared vector store** (pgvector or OpenSearch) + **Postgres** OLTP.  
- **Redis** for rate limits and optional response cache.

---

## 8. AWS production stack (OpenAI retained)

**Minimal recommended set** (OpenAI unchanged):

| Service | Role |
|---------|------|
| **Amazon RDS PostgreSQL** | Primary OLTP; optional **`pgvector`** for embeddings index. |
| **Amazon S3** | Uploads, exports, large artifacts. |
| **Amazon ElastiCache Redis** | Distributed rate limits, locks, optional cache. |
| **Amazon SQS** (+ DLQ) | Async generation jobs, backoff, worker scaling. |
| **Amazon ECS on Fargate** (or EKS) | API service + worker service. |
| **Application Load Balancer** | Routing, health checks. |
| **AWS Secrets Manager** | OpenAI key, DB credentials. |
| **Amazon CloudWatch** | Logs, metrics, alarms. |

**Defer until needed:** WAF, Step Functions, X-Ray, OpenSearch (if pgvector suffices), EventBridge (can add for stuck-report sweeper).

---

## 9. Observability & operations

### 9.1 Minimum viable

- **Structured logs** (JSON): `request_id`, `tenant_id`, `report_id`, `section_code`, latency, OpenAI error codes.  
- **Alarms:** 5xx rate, p95 latency, queue depth, DB connections, **OpenAI 429** spikes, **stuck `generating` %**.

### 9.2 Backups & DR

- **RDS** automated backups + restore drill.  
- **S3** versioning / lifecycle as per compliance.  
- **RPO/RTO** targets documented.

### 9.3 Runbooks

- OpenAI key rotation.  
- Queue poison message (DLQ replay).  
- Stuck report reset procedure until automated sweeper exists.

---

## 10. Role responsibilities & deliverables

| Role | Owns | Key deliverables |
|------|------|------------------|
| **DevOps / SRE** | AWS accounts, networking, secrets, CI/CD, scaling, backups | RDS, S3, Redis, SQS, ECS/Fargate, ALB, Secrets Manager, CloudWatch alarms, TLS |
| **Backend** | Application integration with AWS; job lifecycle; vector adapter | Postgres schema, S3 ingest, SQS consumer, Redis limiter, pgvector/OpenSearch `VectorStore`, stuck-job handling, real auth β†’ tenant |
| **AI engineering** | Models, prompts, evals, cost/latency policy | Golden sets, pinned models, retry policy, document guardrail limits, Bedrock **not** required if staying on OpenAI |
| **Frontend** | UX for auth, long jobs, errors | Bearer/session auth, 429/409 handling, progress against real job status |

---

## 11. Pre-launch checklist

- [ ] **RDS Postgres** live; app `DATABASE_URL` pointed; migrations applied.  
- [ ] **S3** for uploads; no critical path depends on local disk alone.  
- [ ] **Redis** rate limiter (or gateway limits) β€” verified under **multi-task** load.  
- [ ] **SQS** (or equivalent) β€” worker processes tested; DLQ monitored.  
- [ ] **Vector strategy** chosen β€” pgvector **or** OpenSearch **or** explicit single-replica limitation documented.  
- [ ] **Auth** β€” tenant cannot be spoofed via header alone.  
- [ ] **Secrets** β€” OpenAI + DB in Secrets Manager; no keys in images.  
- [ ] **Stuck `generating`** β€” sweeper or manual runbook tested once.  
- [ ] **Load test** β€” representative: upload β†’ ingest β†’ N sections generate; measure p95 and error rate.  
- [ ] **AI eval** β€” regression set run on candidate prod config.  
- [ ] **CORS** β€” restricted to real frontends.  

---

## 12. Post-launch & continuous improvement

- Review **OpenAI spend** and **429** dashboards weekly after launch.  
- Tune **worker concurrency** vs **OpenAI tier**.  
- Expand **golden sets** with production-redacted examples.  
- Revisit **retrieval k** and **interference defaults** based on measured quality vs latency.

---

## Document control

| Field | Value |
|-------|--------|
| **Repository path** | `docs/PRODUCTION_REPORT.md` |
| **Intent** | Living reference β€” update when architecture or AWS choices change. |

*End of report.*