Spaces:
Paused
Paused
| # **Software Requirements Specification (SRS)** | |
| ## **1. Introduction** | |
| This document outlines the technical specifications for an Agentic RAG system built on Hugging Face **SmolAgents**, using **pgvector** for storage and **TruLens** for evaluation. | |
| ## **2. System Architecture** | |
| The system follows a modular "Code-First" Agent architecture: | |
| 1. **Orchestrator:** smolagents.CodeAgent (interprets user query and writes execution code). | |
| 2. **Tool Layer:** Custom pgvector retriever tool with integrated reranking logic. | |
| 3. **Storage Layer:** PostgreSQL with the vector extension. | |
| 4. **Evaluation Layer:** TruCustomApp recorder with OpenAI/HF/Gemini feedback providers. | |
| ## **3. Detailed Component Specifications** | |
| ### **3.1 Data Layer (PostgreSQL + pgvector)** | |
| * **Schema:** | |
| * id: UUID (Primary Key) | |
| * content: TEXT (Document chunk) | |
| * embedding: VECTOR(384) | |
| * metadata: JSONB (Source URL, page number, etc.) | |
| * **Search Type:** HNSW or IVFFlat index for sub-100ms vector similarity lookups. | |
| ### **3.2 Retrieval & Reranking Tool** | |
| * **First Stage:** Dense retrieval via pgvector (Cosine Similarity) returning Top 20 results. | |
| * **Second Stage:** Reranking via CrossEncoder('cross-encoder/ms-marco-MiniLM-L-6-v2') — ~85 MB, CPU-friendly, no GPU required. | |
| * **Logic:** The tool takes a string query, computes the embedding, queries PG, reranks the 20 results, and returns a formatted string of the top 5. | |
| ### **3.3 Agentic Logic (SmolAgents)** | |
| The system MUST provide a configuration toggle to switch between the following Model Providers: | |
| * **Option A: OpenAI Ecosystem** | |
| * Includes standard OpenAI (GPT-4o, GPT-4o-mini) and **Azure OpenAI** instances. | |
| * Uses OpenAIServerModel or LiteLLMModel for integration. | |
| * **Option B: Google Gemini Family** | |
| * Includes gemini-3.1-flash-lite-preview and gemini-flash-lite-latest variants. | |
| * Uses native Google Generative AI integration or LiteLLMModel. | |
| * **Reasoning:** The agent is configured with max_steps=5 to allow for query expansion or multi-step searching if the first attempt fails. | |
| ### **3.4 Evaluation Framework (TruLens 2.x OTEL)** | |
| The backend uses **TruApp** with OTEL-instrumented replay, **RAG Triad** metrics (Context Relevance, Groundedness, Answer Relevance), and an LLM-as-judge whose provider is configured separately from the agent (`TRULENS_PROVIDER`: OpenAI, Azure OpenAI, Gemini, or none). Evaluation runs **asynchronously** after the HTTP response; the worker synchronizes on TruLens feedback completion up to **`TRULENS_FEEDBACK_TIMEOUT`** before persisting to `evaluation_results`. Retrieval context for scoring prefers the **authoritative reranked top-5 tool string** when the retriever ran. | |
| * **Metrics (The RAG Triad):** | |
| * Context relevance: User query vs. reranked retrieval context. | |
| * Groundedness: Final answer vs. reranked retrieval context. | |
| * Answer relevance: User query vs. final answer. | |
| ## **4. Technical Constraints** | |
| * **Environment:** Python 3.10+. | |
| * **API-First Architecture:** Inference for the primary Orchestrator and Evaluation Judge is handled via external APIs. Local compute is reserved for Embedding and Reranking models. | |
| * **Credential Management:** Secure handling of API keys via environment variables: | |
| * OPENAI_API_KEY / AZURE_OPENAI_API_KEY | |
| * AZURE_OPENAI_ENDPOINT | |
| * GOOGLE_API_KEY | |
| * **Connectivity:** The execution environment requires outbound HTTPS access to api.openai.com, \*.openai.azure.com, and generativelanguage.googleapis.com. | |
| * **Memory (Hybrid):** Reranking model (Stage 2) is CPU-only (~85 MB RAM); no GPU required. LLM reasoning is offloaded to the cloud. | |
| * **Security:** Agent code execution must be sandboxed. API keys must not be logged or exposed in the TruLens dashboard traces. | |
| ## **5. Performance Requirements** | |
| * **Reranking Latency:** \< 500ms for 20 document pairs (local inference). | |
| * **Vector Query:** \< 50ms for 100k vectors. | |
| * **Inference Latency:** Target \< 3s for LLM generation steps; depends on provider (OpenAI vs. Gemini) and network overhead. | |
| * **Evaluation Overhead:** TruLens recording should be asynchronous or sampled to avoid blocking the user response. |