File size: 6,542 Bytes
be9fd4a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
# AI feature phases (strict — no backend infra)

Backend scale (Redis job queue, distributed rate limits, Temporal, stale-job sweepers) is **not** an AI phase. See `OPTIMIZATION_SUMMARY.md`**Backend infrastructure**.

| Old doc label | Correct bucket |
|---------------|----------------|
| Phase 3 (Redis) | **Backend** — not AI |
| Phase 2 (Temporal) | **Backend** — not AI |
| `SCALE_OPTIMIZATION_PROFILE` | **Backend + mixed** — also toggles AI Phase 3 flags |

`GET /health` exposes `ai_phases` (this document) separately from `infrastructure` (Redis, job queue).

---

## Phase 1 — Core RAG generation

**Purpose:** Turn surveyor notes + tenant library into RICS section prose via retrieval-augmented generation.

| Capability | Module / API | Requires `OPENAI_API_KEY` |
|------------|--------------|---------------------------|
| Section **generate** / **proofread** / **enhance** | `app/services/generation.py`, `POST …/generate` | Yes |
| Notes expansion | `app/generator/notes_expander.py` | Yes |
| Writing style analysis | `app/generator/style_analyzer.py` | Yes |
| Main LLM adapt + verify pass | `app/generator/adapter.py`, `postprocess.py` | Yes |
| Embeddings + vector index | `app/embeddings/`, `app/vectorstore/`, `app/ingest/` | No (HF local default) |
| Retrieval + rerank + hierarchical RAG | `app/retrieval/` | No |
| Personalised tenant style RAG | `app/services/personalised_rag.py` | No (index); Yes (generate) |
| Upload sanitisation before index | `app/services/document_sanitiser.py` | Optional LLM; regex default |
| Interference levels (`minimum` / `medium` / `maximum`) | `app/generator/prompts.py` | Yes |
| Standard paragraphs injection | `app/services/standard_paragraphs.py` | No |
| Optional LLM section validator | `LLM_SECTION_VALIDATOR_ENABLED` | Yes |

**Default product path:** `NOTES_ONLY_GENERATION=true` → standard pipeline on `POST /generate` (bullets are facts; uploads reference-only).

**Key env (Phase 1):**

```env
OPENAI_API_KEY=…
CHAT_MODEL=gpt-4o-mini
PERSONALISED_STYLE_RAG_ENABLED=true
ENABLE_RAG_UPLOAD_SANITISATION=true
RAG_SANITISATION_USE_LLM=false
NOTES_ONLY_GENERATION=true
```

---

## Phase 2 — Agentic inspector & multimodal

**Purpose:** OpenAI tool-calling inspector, full agentic report orchestration, and section photo vision.

| Capability | Module / API | Gate |
|------------|--------------|------|
| Inspector tool loop | `app/agentic/inspector_loop.py`, `HeadAgent` | `INSPECTOR_TOOL_AGENT=true` + API key |
| Multi-agent report (`generate_full_report`) | `app/agentic/agents.py` | Same |
| Agentic HTTP API | `POST /reports/{id}/agentic/generate` | Same |
| Inspector on main generate | `POST /reports/{id}/generate` | `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true` or `NOTES_ONLY_GENERATION=false` |
| Section photo vision | `app/services/photo_vision.py` | `SECTION_PHOTO_VISION_ENABLED=true` + API key |
| Vision on upload (cache) | `SECTION_PHOTO_ANALYZE_ON_UPLOAD` | Same |
| LLM RAG sanitisation at ingest | `RAG_SANITISATION_USE_LLM` | Dev/staging; off on HF `production_ai_profile` |

**HF / production AI (recommended):**

```env
INSPECTOR_TOOL_AGENT=true
INSPECTOR_BODY_MODEL=gpt-4o-mini
AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true
SECTION_PHOTO_VISION_ENABLED=true
SECTION_PHOTO_VISION_MODEL=gpt-4o
PRIMARY_GENERATE_PIPELINE=agentic
PRODUCTION_AI_PROFILE=true
```

---

## Phase 3 — AI latency & retrieval quality

**Purpose:** Same AI outputs, faster or cheaper — **not** separate workers or queues.

| Capability | Env | Notes |
|------------|-----|--------|
| Async OpenAI + parallel sections | `ENABLE_ASYNC_PIPELINE=true` | `app/llm/async_llm_adapter.py`, `generation_facade.py` |
| Speculative inspector tool prefetch | `ENABLE_SPECULATIVE_EXECUTOR=true` | Requires async + Phase 2 inspector path |
| OpenAI prompt cache keys | `ENABLE_PROMPT_CACHING=true` | Requires async |
| Global LLM concurrency cap | `MAX_CONCURRENT_LLM_CALLS` | Throttle / 429 retry |
| Hybrid BM25 + vector (RRF) | `ENABLE_HYBRID_RETRIEVAL` + `VECTORSTORE_BACKEND=qdrant` | Better retrieval |
| Semantic retrieval cache | `SEMANTIC_CACHE_ENABLED` + Qdrant | Caches embedding-neighbour hits |
| Qdrant vector backend | `VECTORSTORE_BACKEND=qdrant` | Re-ingest after switch |

**Not Phase 3 (backend):** `REDIS_URL`, `ENABLE_JOB_QUEUE`, `ENABLE_TEMPORAL_WORKFLOW`, `GENERATION_STALE_SWEEP_SECONDS`, `SCALE_OPTIMIZATION_PROFILE` (mixed).

**Example (local AI perf only):**

```env
ENABLE_ASYNC_PIPELINE=true
ENABLE_SPECULATIVE_EXECUTOR=true
ENABLE_PROMPT_CACHING=true
AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true
# Optional stronger retrieval:
# VECTORSTORE_BACKEND=qdrant
# QDRANT_URL=http://localhost:6333
```

---

## Full-report time target (10 minutes)

| Setting | Default | Role |
|---------|---------|------|
| `GENERATION_SLA_SECONDS` | `600` | Product target for batch / full-report jobs |
| `GENERATION_TIMEOUT_SECONDS` | `720` | Sweeper fails stuck jobs (SLA + 2 min grace) |

**Requirements to hit SLA** (typical 15–25 sections):

1. **AI Phase 3 on:** `ENABLE_ASYNC_PIPELINE=true` and parallel multi-section (auto with `PRODUCTION_AI_PROFILE` or HF `SPACE_ID`).
2. **Phase 2 on main path:** `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true` only if you need the inspector; standard pipeline is faster for bulk.
3. **Interference:** prefer `medium` or `minimum` for batch — `maximum` adds tokens and retries.
4. **Bullets:** empty sections stay blank; fill bullets per section before batch generate.
5. **UI:** batch poll waits at least 10 minutes before showing “still generating”.

Sequential generation (async off) often exceeds 10 minutes — not supported for full-report SLA.

---

## Ship checklist by environment

| Environment | Phase 1 | Phase 2 | Phase 3 |
|-------------|---------|---------|---------|
| **HF Space pilot** | API key + personalised RAG | Inspector + vision + `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY` | Usually **off** (single process; async optional) |
| **Docker single node** | Same | Same | Optional async + speculation |
| **Multi-replica AWS** | Same | Same | Phase 3 AI + **backend** Redis queue (see `PRODUCTION_REPORT.md`) |

---

## Code map

| Phase | Health JSON | Python |
|-------|-------------|--------|
| 1–3 | `GET /health``ai_phases` | `app/optimization/ai_phases.py` |
| Warnings | `ai_phase_warnings` | `collect_ai_phase_warnings()` |
| User-facing summary | `ai_features` | `app/optimization/ai_readiness.py` |
| Backend only | `infrastructure` | `app/optimization/scale_status.py`, Redis, jobs, Temporal |