Spaces:
Sleeping
Production Report β RICS Report Genius AI
Document purpose: Single reference for production deployment, scale, AI operations, and ownership.
Scope: Current codebase behaviour (FastAPI, SQLite default, FAISS, LangChain, OpenAI), target production on AWS with OpenAI retained.
Audience: Engineering leads, DevOps, AI engineers, backend engineers, frontend engineers, security/compliance.
Table of contents
- Executive summary
- System architecture (as implemented)
- Data & state
- AI & LLM pipeline
- APIs, concurrency & failure modes
- Security & tenancy
- Performance & scale
- AWS production stack (OpenAI retained)
- Observability & operations
- Role responsibilities & deliverables
- Pre-launch checklist
- Post-launch & continuous improvement
1. Executive summary
The application is a FastAPI service serving a static frontend, with per-tenant RAG (FAISS + LangChain embeddings), OpenAI for generation and several post-processing passes, and async SQLite by default. It is fit for pilot or single-tenantβstyle deployments out of the box. For thousands of concurrent users, production requires:
| Pillar | Gap in default / local shape | Production requirement |
|---|---|---|
| Data | SQLite + local files | PostgreSQL (RDS) + S3 for uploads |
| Coordination | In-process rate limits, per-pod FAISS/cache | Redis (ElastiCache) + SQS workers + clear vector strategy |
| Tenancy | X-Tenant-ID header only |
Authenticated identity mapped to tenant |
| AI reliability | OpenAI rate limits, retries, cost | Queues, concurrency caps, timeouts, observability |
| Ops | Ephemeral disk on some hosts | Backups, alarms, stuck-job handling |
OpenAI remains the LLM provider; AWS provides compute, data, queue, cache, secrets, and logs β not a replacement for OpenAIβs API limits or model behaviour.
2. System architecture (as implemented)
2.1 Runtime
- Framework: FastAPI (
app/main.py). - Entry: Uvicorn; Docker image targets port 7860 (HF Spaces) or 8000 (local/docker-compose) per
Dockerfile/ env. - Startup: DB init, vector store pre-warm, embedding client pre-warm; optional KB ingest in background; production guard requires
OPENAI_API_KEYwhenDEV_MODE=false.
2.2 Major components
| Component | Location / behaviour |
|---|---|
| HTTP API | app/api/* β reports, generate, upload, status, export, photos, etc. |
| Generation orchestration | app/services/generation.py β RAG, style, LLM, verify, retries, persist. |
| Vector store | app/vectorstore/factory.py β FAISS singleton, LangChain-backed. |
| Embeddings | app/embeddings/factory.py β default HuggingFace local; optional OpenAI via config. |
| Prompts | app/generator/prompts.py β LCEL vars, interference levels, token budgets. |
| LLM adapter | app/generator/adapter.py β OpenAI / mock. |
| Post-process | app/generator/postprocess.py β async_enforce_verify, etc. |
| Section cache | app/cache/section_cache.py β filesystem JSON keyed by hash of inputs. |
| Rate limiting | app/api/rate_limit.py β in-process sliding window per tenant_id from header. |
2.3 Generation modes (API)
generateβ Full pipeline (default product path;notes_only_generationforces standard pipeline in current code paths).proofread/enhanceβ Separate pipelines reusing parts of generation.
2.4 AI interference (product)
- Qualitative levels:
minimum|medium|maximum(interference_levelonGenerateRequest; maps to internalai_percentbands inapp/models/schemas.py). - Prompts branch on interference level in
app/generator/prompts.py.
3. Data & state
3.1 Database
- Default:
sqlite+aiosqlite:///./dev.db(app/config.py). - Production risk: SQLite single-writer semantics, file locking, no horizontal write scale β unsuitable for high concurrent write load.
- Target: Amazon RDS PostgreSQL (or Aurora PostgreSQL) with connection pooling (app uses
pool_size/max_overflowfor non-SQLite inapp/db/database.py).
3.2 Uploads & artifacts
- Default: Local
upload_diron instance filesystem. - Production risk: Ephemeral disks lose data on redeploy/restart; multi-replica inconsistent views of files.
- Target: Amazon S3 + DB rows storing object keys; presigned URLs for client download where needed.
3.3 Vector index (FAISS)
- Current: In-process singleton; persisted under app-configured paths (e.g. user home / cache dirs).
- Production risk: Not shared across replicas; each pod has its own index unless you mount shared storage (generally discouraged for FAISS at scale).
- Target options:
- RDS
pgvectorβ vectors in Postgres; one operational system. - Amazon OpenSearch kNN β separate service; strong scale/filter story.
- Single replica only β acceptable only for early internal launch, not βthousands concurrent.β
- RDS
3.4 Section result cache
- Current: File-based cache in
settings.cache_dir; atomic write via temp + rename. - Production risk: Per-node only; no cross-instance deduplication.
- Target: Redis for hot cache keys (optional) or accept cache miss on each pod.
4. AI & LLM pipeline
4.1 OpenAI usage (representative chain per section)
Depending on settings and branch, a generate path may include:
- Notes expansion β LLM when OpenAI key present and expansion not skipped (
app/generator/notes_expander.py). - Main section generation β LangChain LCEL + ChatOpenAI (
app/generator/adapter.py). async_enforce_verifyβ Additional LLM pass when key present (app/generator/postprocess.py).- Retries β Verbatim ratio, identity consistency, survey-tier hints β can trigger extra
generate_sectioncalls. - Optional LLM section validator β
llm_section_validator_enabledinapp/config.py(default false); when true, extra validate + possible regenerate. - Photo / vision β When photos exist and policy triggers (
app/services/photo_vision.py).
Implication: Wall-clock and cost scale with number of LLM round-trips, not only βone GPT call per section.β
4.2 Embeddings
- Default:
prefer_local_embeddings=Trueβ HuggingFace sentence-transformers on CPU per process (app/embeddings/factory.py). - Production trade-off: Each replica loads model into RAM; cold start and memory multiply with replica count.
- Mitigation: Set
prefer_local_embeddings=falseto use OpenAI embeddings API (cost vs RAM); or keep HF on GPU workers dedicated to ingest (heavier ops design).
4.3 Quality & guardrails (behavioural, not infra)
- Notes-only mode (
notes_only_generation, default true): bullets are factual authority; uploads are reference-only in that design; provenance suppressed accordingly. - Interference levels encode how much drafting vs mapping is allowed (see
app/generator/prompts.pyand schemas). - Production AI engineering: golden-set evaluation, pinned models, documented limits of each guard (what verify does / does not guarantee).
4.4 Cost & limits (OpenAI)
- Account-level rate limits and token pricing apply regardless of AWS scale.
- Mitigation: SQS-backed workers with global + per-tenant concurrency caps, exponential backoff on 429, timeouts, circuit breaker patterns.
5. APIs, concurrency & failure modes
5.1 Report generation locking
POST /reports/{report_id}/generateuses a compare-and-swap soReport.statusis only flipped togeneratingif not alreadygenerating(app/api/generate.py).- Second concurrent generate for same report β 409 until the first completes or fails.
- Worker crash after CAS but before completion can leave status
generatinguntil the stale-generation sweeper marks the report failed (generation_stale_sweep_seconds,generation_started_atonreports). Runs on the API process and on the Temporal worker when enabled.
5.2 Background execution
- Generation runs as
asyncio.create_taskfrom the API process; errors inrun_generationset reportfailedanderror_message(app/services/generation.py). - Risk: Process kill loses in-flight task; see stuck state above.
5.3 Rate limiting
- Generate: default 20 RPM per tenant id (header); Read: 120 RPM (
app/api/rate_limit.py, envRATE_LIMIT_GENERATE_RPM,RATE_LIMIT_READ_RPM). - Limitation: In-process β incorrect under multiple workers/pods (each has its own counter). Production should use Redis or API Gateway throttling.
5.4 HTTP status summary (client expectations)
| Code | Typical cause |
|---|---|
| 401 | Missing / empty X-Tenant-ID on non-exempt paths (app/api/middleware.py). |
| 404 | Report or document not found / wrong tenant. |
| 409 | Report already generating; or business rule conflicts elsewhere. |
| 413 | Payload too large (e.g. section text PATCH limit). |
| 422 | Validation (e.g. invalid interference_level). |
| 429 | Rate limit exceeded. |
6. Security & tenancy
6.1 Current tenant model
TenantAuthMiddleware: requires non-emptyX-Tenant-IDheader; no cryptographic proof of ownership (app/api/middleware.py).
6.2 Production requirement
- Authenticate users (Cognito, Auth0, custom JWT, API keys behind API Gateway, etc.) and derive
tenant_idfrom trusted claims β do not treat the raw header as the security boundary.
6.3 Secrets
OPENAI_API_KEY: required in production whenDEV_MODE=false(app/main.py).TENANT_SECRET_KEY: defaults trigger warnings; should be non-sentinel before any future signing use (app/main.py).
6.4 CORS
allowed_originsdefault["*"]with credentials false inapp/main.pyβ tighten for production web origins.
7. Performance & scale
7.1 Dominant bottlenecks
- Sequential LLM calls per section (expand β generate β verify β retries).
- OpenAI latency + rate limits.
- Database under concurrent writes if SQLite remains.
- Per-replica FAISS + HF embeddings β memory and duplicate ingest work.
7.2 Application-level levers (documented in config)
- Retrieval pool sizes (
hierarchical_k_*,retrieval_top_k,rerank_top_ninapp/config.py). - Optional LLM section validator (off by default).
- Interference level β Maximum mode increases output tokens and time.
7.3 Architecture-level levers
- Queue + workers (SQS) for generate jobs.
- Bounded concurrency per tenant and globally toward OpenAI.
- Shared vector store (pgvector or OpenSearch) + Postgres OLTP.
- Redis for rate limits and optional response cache.
8. AWS production stack (OpenAI retained)
Minimal recommended set (OpenAI unchanged):
| Service | Role |
|---|---|
| Amazon RDS PostgreSQL | Primary OLTP; optional pgvector for embeddings index. |
| Amazon S3 | Uploads, exports, large artifacts. |
| Amazon ElastiCache Redis | Distributed rate limits, locks, optional cache. |
| Amazon SQS (+ DLQ) | Async generation jobs, backoff, worker scaling. |
| Amazon ECS on Fargate (or EKS) | API service + worker service. |
| Application Load Balancer | Routing, health checks. |
| AWS Secrets Manager | OpenAI key, DB credentials. |
| Amazon CloudWatch | Logs, metrics, alarms. |
Defer until needed: WAF, Step Functions, X-Ray, OpenSearch (if pgvector suffices), EventBridge (can add for stuck-report sweeper).
9. Observability & operations
9.1 Minimum viable
- Structured logs (JSON):
request_id,tenant_id,report_id,section_code, latency, OpenAI error codes. - Alarms: 5xx rate, p95 latency, queue depth, DB connections, OpenAI 429 spikes, stuck
generating%.
9.2 Backups & DR
- RDS automated backups + restore drill.
- S3 versioning / lifecycle as per compliance.
- RPO/RTO targets documented.
9.3 Runbooks
- OpenAI key rotation.
- Queue poison message (DLQ replay).
- Stuck report reset procedure until automated sweeper exists.
10. Role responsibilities & deliverables
| Role | Owns | Key deliverables |
|---|---|---|
| DevOps / SRE | AWS accounts, networking, secrets, CI/CD, scaling, backups | RDS, S3, Redis, SQS, ECS/Fargate, ALB, Secrets Manager, CloudWatch alarms, TLS |
| Backend | Application integration with AWS; job lifecycle; vector adapter | Postgres schema, S3 ingest, SQS consumer, Redis limiter, pgvector/OpenSearch VectorStore, stuck-job handling, real auth β tenant |
| AI engineering | Models, prompts, evals, cost/latency policy | Golden sets, pinned models, retry policy, document guardrail limits, Bedrock not required if staying on OpenAI |
| Frontend | UX for auth, long jobs, errors | Bearer/session auth, 429/409 handling, progress against real job status |
11. Pre-launch checklist
- RDS Postgres live; app
DATABASE_URLpointed; migrations applied. - S3 for uploads; no critical path depends on local disk alone.
- Redis rate limiter (or gateway limits) β verified under multi-task load.
- SQS (or equivalent) β worker processes tested; DLQ monitored.
- Vector strategy chosen β pgvector or OpenSearch or explicit single-replica limitation documented.
- Auth β tenant cannot be spoofed via header alone.
- Secrets β OpenAI + DB in Secrets Manager; no keys in images.
- Stuck
generatingβ sweeper or manual runbook tested once. - Load test β representative: upload β ingest β N sections generate; measure p95 and error rate.
- AI eval β regression set run on candidate prod config.
- CORS β restricted to real frontends.
12. Post-launch & continuous improvement
- Review OpenAI spend and 429 dashboards weekly after launch.
- Tune worker concurrency vs OpenAI tier.
- Expand golden sets with production-redacted examples.
- Revisit retrieval k and interference defaults based on measured quality vs latency.
Document control
| Field | Value |
|---|---|
| Repository path | docs/PRODUCTION_REPORT.md |
| Intent | Living reference β update when architecture or AWS choices change. |
End of report.