Spaces:
Sleeping
Sleeping
| # Optimization summary β AI phases vs backend | |
| **Authoritative AI taxonomy:** [`docs/AI_FEATURES_PHASES.md`](docs/AI_FEATURES_PHASES.md) | |
| | Layer | Phases | Scope | | |
| |-------|--------|--------| | |
| | **AI** | Phase 1β3 | OpenAI, RAG, inspector, vision, async LLM, speculation, prompt cache, Qdrant retrieval | | |
| | **Backend** | (not numbered as AI) | Redis rate limits, job queue, Temporal, stale sweeps, SQLite parallel writes | | |
| `GET /health` returns `ai_phases` + `ai_phase_warnings` separately from `infrastructure`. | |
| --- | |
| ## AI Phase 1 β Core RAG generation | |
| | Feature | Env / code | Notes | | |
| |---------|------------|--------| | |
| | Generate / proofread / enhance | `OPENAI_API_KEY`, `CHAT_MODEL` | `app/services/generation.py` | | |
| | Notes expansion, style analyzer, verify | (same key) | `notes_expander`, `style_analyzer`, `postprocess` | | |
| | Embeddings | `PREFER_LOCAL_EMBEDDINGS`, `LOCAL_EMBEDDING_MODEL` | HF default; optional OpenAI | | |
| | FAISS / Qdrant index | `VECTORSTORE_BACKEND`, `FAISS_INDEX_PATH` | Qdrant optional upgrade | | |
| | Hierarchical RAG | `HIERARCHICAL_RAG_ENABLED` | Document β section β paragraph | | |
| | Personalised style RAG | `PERSONALISED_STYLE_RAG_ENABLED` | Tenant library retrieval | | |
| | Upload sanitisation | `ENABLE_RAG_UPLOAD_SANITISATION`, `RAG_SANITISATION_USE_LLM` | Regex default in prod | | |
| | Notes-only mode | `NOTES_ONLY_GENERATION` | Standard pipeline default on `POST /generate` | | |
| --- | |
| ## AI Phase 2 β Agentic inspector & vision | |
| | Feature | Env flags | Notes | | |
| |---------|-----------|--------| | |
| | Inspector tool loop | `INSPECTOR_TOOL_AGENT=true` | `app/agentic/inspector_loop.py` | | |
| | Inspector on main generate | `AGENTIC_INSPECTOR_WHEN_NOTES_ONLY=true` | Or `NOTES_ONLY_GENERATION=false` | | |
| | Agentic API | `POST β¦/agentic/generate` | `app/api/agentic.py` | | |
| | Section photo vision | `SECTION_PHOTO_VISION_ENABLED` | `app/services/photo_vision.py` | | |
| | LLM section validator | `LLM_SECTION_VALIDATOR_ENABLED` | Optional extra pass | | |
| | Production AI profile | `PRODUCTION_AI_PROFILE` / `SPACE_ID` | Disables ingest LLM sanitisation only | | |
| --- | |
| ## AI Phase 3 β Latency & retrieval quality (still AI, not Redis) | |
| | Feature | Env flags | Notes | | |
| |---------|-----------|--------| | |
| | Async LLM | `ENABLE_ASYNC_PIPELINE=true` | Parallel multi-section when safe | | |
| | Speculative tools | `ENABLE_SPECULATIVE_EXECUTOR=true` | Inspector prefetch; needs async + Phase 2 path | | |
| | Prompt cache | `ENABLE_PROMPT_CACHING=true` | OpenAI `prompt_cache_key` | | |
| | LLM throttle | `MAX_CONCURRENT_LLM_CALLS` | 429 backoff | | |
| | Hybrid BM25 + RRF | `ENABLE_HYBRID_RETRIEVAL` + `VECTORSTORE_BACKEND=qdrant` | | | |
| | Semantic retrieval cache | `SEMANTIC_CACHE_ENABLED` + Qdrant | | | |
| --- | |
| ## Backend infrastructure (not AI phases) | |
| | Feature | Env flags | Notes | | |
| |---------|-----------|--------| | |
| | Distributed rate limits | `REDIS_URL` | `app/rate_limit/redis_limiter.py` | | |
| | Generation job queue | `ENABLE_JOB_QUEUE` + `REDIS_URL` | `jobs_worker.py` | | |
| | Temporal workflows | `ENABLE_TEMPORAL_WORKFLOW` | `worker.py`, durable orchestration | | |
| | Stale generation recovery | `GENERATION_TIMEOUT_SECONDS`, `GENERATION_STALE_SWEEP_SECONDS` | DB sweeper | | |
| | Scale profile (mixed) | `SCALE_OPTIMIZATION_PROFILE` | Auto-enables AI Phase 3 flags + job queue when Redis set; **off on HF** | | |
| `POST /generate` dispatch order: Temporal (if enabled) β Redis queue β in-process asyncio. | |
| ```bash | |
| # Backend only (no extra AI beyond flags the profile sets) | |
| SCALE_OPTIMIZATION_PROFILE=true REDIS_URL=redis://redis:6379/0 | |
| docker compose --profile redis --profile jobs up -d | |
| ``` | |
| --- | |
| ## Health fields | |
| | Field | Content | | |
| |-------|---------| | |
| | `ai_phases` | Phase 1β3 status (`app/optimization/ai_phases.py`) | | |
| | `ai_phase_warnings` | AI misconfiguration only | | |
| | `ai_features` | Compact summary for dashboards | | |
| | `infrastructure` | Redis, job queue, legacy `backend_scale` blob | | |
| | `optimization_warnings` | Mixed ops hints (includes Qdrant/FAISS migration notes) | | |