Spaces:
Sleeping
Sleeping
File size: 6,542 Bytes
be9fd4a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | # AI feature phases (strict — no backend infra)
Backend scale (Redis job queue, distributed rate limits, Temporal, stale-job sweepers) is **not** an AI phase. See `OPTIMIZATION_SUMMARY.md` → **Backend infrastructure**.
| Old doc label | Correct bucket |
|---------------|----------------|
| Phase 3 (Redis) | **Backend** — not AI |
| Phase 2 (Temporal) | **Backend** — not AI |
| `SCALE_OPTIMIZATION_PROFILE` | **Backend + mixed** — also toggles AI Phase 3 flags |
`GET /health` exposes `ai_phases` (this document) separately from `infrastructure` (Redis, job queue).
---
## Phase 1 — Core RAG generation
**Purpose:** Turn surveyor notes + tenant library into RICS section prose via retrieval-augmented generation.
| Capability | Module / API | Requires `OPENAI_API_KEY` |
|------------|--------------|---------------------------|
| Section **generate** / **proofread** / **enhance** | `app/services/generation.py`, `POST …/generate` | Yes |
| Notes expansion | `app/generator/notes_expander.py` | Yes |
| Writing style analysis | `app/generator/style_analyzer.py` | Yes |
| Main LLM adapt + verify pass | `app/generator/adapter.py`, `postprocess.py` | Yes |
| Embeddings + vector index | `app/embeddings/`, `app/vectorstore/`, `app/ingest/` | No (HF local default) |
| Retrieval + rerank + hierarchical RAG | `app/retrieval/` | No |
| Personalised tenant style RAG | `app/services/personalised_rag.py` | No (index); Yes (generate) |
| Upload sanitisation before index | `app/services/document_sanitiser.py` | Optional LLM; regex default |
| Interference levels (`minimum` / `medium` / `maximum`) | `app/generator/prompts.py` | Yes |
| Standard paragraphs injection | `app/services/standard_paragraphs.py` | No |
| Optional LLM section validator | `LLM_SECTION_VALIDATOR_ENABLED` | Yes |
**Default product path:** `NOTES_ONLY_GENERATION=true` → standard pipeline on `POST /generate` (bullets are facts; uploads reference-only).
**Key env (Phase 1):**
```env
OPENAI_API_KEY=…
CHAT_MODEL=gpt-4o-mini
PERSONALISED_STYLE_RAG_ENABLED=true
ENABLE_RAG_UPLOAD_SANITISATION=true
RAG_SANITISATION_USE_LLM=false
NOTES_ONLY_GENERATION=true
```
---
## Phase 2 — Agentic inspector & multimodal
**Purpose:** OpenAI tool-calling inspector, full agentic report orchestration, and section photo vision.
| Capability | Module / API | Gate |
|------------|--------------|------|
| Inspector tool loop | `app/agentic/inspector_loop.py`, `HeadAgent` | `INSPECTOR_TOOL_AGENT=true` + API key |
| Multi-agent report (`generate_full_report`) | `app/agentic/agents.py` | Same |
| Agentic HTTP API | `POST /reports/{id}/agentic/generate` | Same |
| Inspector on main generate | `POST /reports/{id}/generate` | `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true` or `NOTES_ONLY_GENERATION=false` |
| Section photo vision | `app/services/photo_vision.py` | `SECTION_PHOTO_VISION_ENABLED=true` + API key |
| Vision on upload (cache) | `SECTION_PHOTO_ANALYZE_ON_UPLOAD` | Same |
| LLM RAG sanitisation at ingest | `RAG_SANITISATION_USE_LLM` | Dev/staging; off on HF `production_ai_profile` |
**HF / production AI (recommended):**
```env
INSPECTOR_TOOL_AGENT=true
INSPECTOR_BODY_MODEL=gpt-4o-mini
AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true
SECTION_PHOTO_VISION_ENABLED=true
SECTION_PHOTO_VISION_MODEL=gpt-4o
PRIMARY_GENERATE_PIPELINE=agentic
PRODUCTION_AI_PROFILE=true
```
---
## Phase 3 — AI latency & retrieval quality
**Purpose:** Same AI outputs, faster or cheaper — **not** separate workers or queues.
| Capability | Env | Notes |
|------------|-----|--------|
| Async OpenAI + parallel sections | `ENABLE_ASYNC_PIPELINE=true` | `app/llm/async_llm_adapter.py`, `generation_facade.py` |
| Speculative inspector tool prefetch | `ENABLE_SPECULATIVE_EXECUTOR=true` | Requires async + Phase 2 inspector path |
| OpenAI prompt cache keys | `ENABLE_PROMPT_CACHING=true` | Requires async |
| Global LLM concurrency cap | `MAX_CONCURRENT_LLM_CALLS` | Throttle / 429 retry |
| Hybrid BM25 + vector (RRF) | `ENABLE_HYBRID_RETRIEVAL` + `VECTORSTORE_BACKEND=qdrant` | Better retrieval |
| Semantic retrieval cache | `SEMANTIC_CACHE_ENABLED` + Qdrant | Caches embedding-neighbour hits |
| Qdrant vector backend | `VECTORSTORE_BACKEND=qdrant` | Re-ingest after switch |
**Not Phase 3 (backend):** `REDIS_URL`, `ENABLE_JOB_QUEUE`, `ENABLE_TEMPORAL_WORKFLOW`, `GENERATION_STALE_SWEEP_SECONDS`, `SCALE_OPTIMIZATION_PROFILE` (mixed).
**Example (local AI perf only):**
```env
ENABLE_ASYNC_PIPELINE=true
ENABLE_SPECULATIVE_EXECUTOR=true
ENABLE_PROMPT_CACHING=true
AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true
# Optional stronger retrieval:
# VECTORSTORE_BACKEND=qdrant
# QDRANT_URL=http://localhost:6333
```
---
## Full-report time target (10 minutes)
| Setting | Default | Role |
|---------|---------|------|
| `GENERATION_SLA_SECONDS` | `600` | Product target for batch / full-report jobs |
| `GENERATION_TIMEOUT_SECONDS` | `720` | Sweeper fails stuck jobs (SLA + 2 min grace) |
**Requirements to hit SLA** (typical 15–25 sections):
1. **AI Phase 3 on:** `ENABLE_ASYNC_PIPELINE=true` and parallel multi-section (auto with `PRODUCTION_AI_PROFILE` or HF `SPACE_ID`).
2. **Phase 2 on main path:** `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true` only if you need the inspector; standard pipeline is faster for bulk.
3. **Interference:** prefer `medium` or `minimum` for batch — `maximum` adds tokens and retries.
4. **Bullets:** empty sections stay blank; fill bullets per section before batch generate.
5. **UI:** batch poll waits at least 10 minutes before showing “still generating”.
Sequential generation (async off) often exceeds 10 minutes — not supported for full-report SLA.
---
## Ship checklist by environment
| Environment | Phase 1 | Phase 2 | Phase 3 |
|-------------|---------|---------|---------|
| **HF Space pilot** | API key + personalised RAG | Inspector + vision + `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY` | Usually **off** (single process; async optional) |
| **Docker single node** | Same | Same | Optional async + speculation |
| **Multi-replica AWS** | Same | Same | Phase 3 AI + **backend** Redis queue (see `PRODUCTION_REPORT.md`) |
---
## Code map
| Phase | Health JSON | Python |
|-------|-------------|--------|
| 1–3 | `GET /health` → `ai_phases` | `app/optimization/ai_phases.py` |
| Warnings | `ai_phase_warnings` | `collect_ai_phase_warnings()` |
| User-facing summary | `ai_features` | `app/optimization/ai_readiness.py` |
| Backend only | `infrastructure` | `app/optimization/scale_status.py`, Redis, jobs, Temporal |
|