File size: 3,961 Bytes
be9fd4a
732b14f
be9fd4a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
732b14f
 
 
be9fd4a
 
 
 
 
 
 
 
732b14f
be9fd4a
732b14f
 
 
be9fd4a
 
 
 
 
 
732b14f
be9fd4a
732b14f
be9fd4a
732b14f
 
 
be9fd4a
 
 
 
 
 
 
 
 
 
 
 
 
732b14f
be9fd4a
732b14f
be9fd4a
732b14f
be9fd4a
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
# Optimization summary β€” AI phases vs backend

**Authoritative AI taxonomy:** [`docs/AI_FEATURES_PHASES.md`](docs/AI_FEATURES_PHASES.md)

| Layer | Phases | Scope |
|-------|--------|--------|
| **AI** | Phase 1–3 | OpenAI, RAG, inspector, vision, async LLM, speculation, prompt cache, Qdrant retrieval |
| **Backend** | (not numbered as AI) | Redis rate limits, job queue, Temporal, stale sweeps, SQLite parallel writes |

`GET /health` returns `ai_phases` + `ai_phase_warnings` separately from `infrastructure`.

---

## AI Phase 1 β€” Core RAG generation

| Feature | Env / code | Notes |
|---------|------------|--------|
| Generate / proofread / enhance | `OPENAI_API_KEY`, `CHAT_MODEL` | `app/services/generation.py` |
| Notes expansion, style analyzer, verify | (same key) | `notes_expander`, `style_analyzer`, `postprocess` |
| Embeddings | `PREFER_LOCAL_EMBEDDINGS`, `LOCAL_EMBEDDING_MODEL` | HF default; optional OpenAI |
| FAISS / Qdrant index | `VECTORSTORE_BACKEND`, `FAISS_INDEX_PATH` | Qdrant optional upgrade |
| Hierarchical RAG | `HIERARCHICAL_RAG_ENABLED` | Document β†’ section β†’ paragraph |
| Personalised style RAG | `PERSONALISED_STYLE_RAG_ENABLED` | Tenant library retrieval |
| Upload sanitisation | `ENABLE_RAG_UPLOAD_SANITISATION`, `RAG_SANITISATION_USE_LLM` | Regex default in prod |
| Notes-only mode | `NOTES_ONLY_GENERATION` | Standard pipeline default on `POST /generate` |

---

## AI Phase 2 β€” Agentic inspector & vision

| Feature | Env flags | Notes |
|---------|-----------|--------|
| Inspector tool loop | `INSPECTOR_TOOL_AGENT=true` | `app/agentic/inspector_loop.py` |
| Inspector on main generate | `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true` | Or `NOTES_ONLY_GENERATION=false` |
| Agentic API | `POST …/agentic/generate` | `app/api/agentic.py` |
| Section photo vision | `SECTION_PHOTO_VISION_ENABLED` | `app/services/photo_vision.py` |
| LLM section validator | `LLM_SECTION_VALIDATOR_ENABLED` | Optional extra pass |
| Production AI profile | `PRODUCTION_AI_PROFILE` / `SPACE_ID` | Disables ingest LLM sanitisation only |

---

## AI Phase 3 β€” Latency & retrieval quality (still AI, not Redis)

| Feature | Env flags | Notes |
|---------|-----------|--------|
| Async LLM | `ENABLE_ASYNC_PIPELINE=true` | Parallel multi-section when safe |
| Speculative tools | `ENABLE_SPECULATIVE_EXECUTOR=true` | Inspector prefetch; needs async + Phase 2 path |
| Prompt cache | `ENABLE_PROMPT_CACHING=true` | OpenAI `prompt_cache_key` |
| LLM throttle | `MAX_CONCURRENT_LLM_CALLS` | 429 backoff |
| Hybrid BM25 + RRF | `ENABLE_HYBRID_RETRIEVAL` + `VECTORSTORE_BACKEND=qdrant` | |
| Semantic retrieval cache | `SEMANTIC_CACHE_ENABLED` + Qdrant | |

---

## Backend infrastructure (not AI phases)

| Feature | Env flags | Notes |
|---------|-----------|--------|
| Distributed rate limits | `REDIS_URL` | `app/rate_limit/redis_limiter.py` |
| Generation job queue | `ENABLE_JOB_QUEUE` + `REDIS_URL` | `jobs_worker.py` |
| Temporal workflows | `ENABLE_TEMPORAL_WORKFLOW` | `worker.py`, durable orchestration |
| Stale generation recovery | `GENERATION_TIMEOUT_SECONDS`, `GENERATION_STALE_SWEEP_SECONDS` | DB sweeper |
| Scale profile (mixed) | `SCALE_OPTIMIZATION_PROFILE` | Auto-enables AI Phase 3 flags + job queue when Redis set; **off on HF** |

`POST /generate` dispatch order: Temporal (if enabled) β†’ Redis queue β†’ in-process asyncio.

```bash
# Backend only (no extra AI beyond flags the profile sets)
SCALE_OPTIMIZATION_PROFILE=true REDIS_URL=redis://redis:6379/0
docker compose --profile redis --profile jobs up -d
```

---

## Health fields

| Field | Content |
|-------|---------|
| `ai_phases` | Phase 1–3 status (`app/optimization/ai_phases.py`) |
| `ai_phase_warnings` | AI misconfiguration only |
| `ai_features` | Compact summary for dashboards |
| `infrastructure` | Redis, job queue, legacy `backend_scale` blob |
| `optimization_warnings` | Mixed ops hints (includes Qdrant/FAISS migration notes) |