Spaces:
Sleeping
Sleeping
File size: 5,060 Bytes
49f0cfb 32c4506 49f0cfb 32c4506 49f0cfb 837762d 49f0cfb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 | # Experiment Lab β LangChain RAG Implementation
This branch rewrites the core RAG pipeline using **LangChain** and its
supporting frameworks, replacing the custom implementations in `master`.
## What Changed
| Component | master branch | experiment-lab branch |
|---|---|---|
| **DOCX loading** | `python-docx` + manual XML block extraction | `langchain_community` `Docx2txtLoader` |
| **PDF loading** | `pymupdf` + font-size heading heuristics | `langchain_community` `PyMuPDFLoader` |
| **Text splitting** | Custom recursive token splitter | `langchain_text_splitters` `RecursiveCharacterTextSplitter` (tiktoken) |
| **Embeddings** | Custom `EmbeddingClient` ABC + manual OpenAI/SentenceTransformer wrappers | `langchain_openai.OpenAIEmbeddings` / `langchain_huggingface.HuggingFaceEmbeddings` |
| **Vector store** | Manual `chromadb` + FAISS with custom metadata filtering | `langchain_community.vectorstores.FAISS` only β local, no Chroma |
| **Retrieval** | Manual query embedding β `vs.search(vector, tenant_id)` | `vs.search(text, tenant_id)` β store embeds internally |
| **LLM generation** | Manual `openai.OpenAI().chat.completions.create()` | **LCEL chain**: `RICS_PROMPT \| ChatOpenAI \| StrOutputParser()` |
| **Retry logic** | None | Built-in via `ChatOpenAI(max_retries=3)` |
## What Stayed the Same
These components contain business logic specific to this system and were
**not** replaced with LangChain equivalents:
- `app/api/` β FastAPI endpoints, tenant middleware, upload/generate/status
- `app/db/` β SQLAlchemy ORM models and session management
- `app/cache/section_cache.py` β SHA-256 keyed JSON section cache
- `app/generator/postprocess.py` β `[VERIFY]` non-invention enforcement
- `app/generator/reranker.py` β Custom Jaccard + numeric overlap reranker
- `app/templates/rics_templates.py` β RICS section skeleton templates
- `app/models/schemas.py` β Pydantic API schemas
- `app/config.py` β Settings management
## New Dependencies
```toml
langchain>=0.3.0
langchain-core>=0.3.0
langchain-openai>=0.2.0
langchain-community>=0.3.0
faiss-cpu>=1.8.0
langchain-huggingface>=0.1.0
langchain-text-splitters>=0.3.0
docx2txt>=0.8
```
## Key Design Decisions
### Why keep a custom reranker?
LangChain's built-in rerankers (`FlashrankRerank`, `CrossEncoderReranker`)
require additional model downloads or paid APIs. The custom Jaccard+numeric
reranker is lightweight, free, and tailored to RICS numeric fact verification.
### Why keep `[VERIFY]` post-processing?
This is a domain-specific safety feature with no LangChain equivalent.
It enforces the non-invention guarantee that is core to RICS report integrity.
### LCEL Chain Structure
```python
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
from langchain_core.output_parsers import StrOutputParser
chain = RICS_PROMPT | ChatOpenAI(model="gpt-4o-mini", temperature=0.2) | StrOutputParser()
text = chain.invoke({"skeleton": ..., "bullets": ..., "examples": ...})
```
### Vectorstore Pipeline
```python
# Ingestion β store handles embedding internally
vectorstore.add_documents(langchain_docs)
# Retrieval β store embeds query internally
results = vectorstore.search(query_text, tenant_id=tenant_id, k=10)
```
## Merged product features (AI transparency + tenant RAG)
The following were merged from the main-line feature work and are **in addition**
to the LangChain stack above:
- **`retrieve_for_report()`** β tenant-scoped retrieval with the reportβs source
document prioritised.
- **Provenance enrichment** β filenames, snippet previews, and section hints for
citations in the API and UI.
- **`compute_ai_transparency()`** β heuristic AI-involvement estimates stored in
section metadata.
- **Frontend** β aggregate AI involvement bar, per-section transparency,
citation chips, and fixed proofread/enhance re-run handling.
- **Generate mode** uses **LCEL** (`RICS_PROMPT | ChatOpenAI | StrOutputParser`)
with the same rich `build_user_prompt` / style-profile content as before
(single `user_content` variable in the chat template).
Proofread and enhance modes still use the OpenAI Chat Completions client for
stable, predictable behaviour.
## Running This Branch
```bash
git checkout experiment-lab
pip install -e ".[dev]"
cp .env.example .env # add OPENAI_API_KEY if available
uvicorn app.main:app --reload
```
Open http://localhost:8000 β the same UI as master.
## Trade-offs vs Master Branch
| Aspect | Master (custom) | Experiment-lab (LangChain) |
|---|---|---|
| Code volume | ~800 lines of RAG pipeline code | ~400 lines (LangChain handles the rest) |
| Structural fidelity | Heading/list/table block types preserved | Plain text only (heading levels lost) |
| Debuggability | Full control, easy to step through | LangChain abstractions can obscure flow |
| Observability | Manual logging | LangSmith tracing available |
| Upgradability | Manual updates to each component | Single `pip upgrade langchain` |
| Vendor lock-in | None | LangChain API surface |
|