# **Software Requirements Specification (SRS)** ## **1. Introduction** This document outlines the technical specifications for an Agentic RAG system built on Hugging Face **SmolAgents**, using **pgvector** for storage and **TruLens** for evaluation. ## **2. System Architecture** The system follows a modular "Code-First" Agent architecture: 1. **Orchestrator:** smolagents.CodeAgent (interprets user query and writes execution code). 2. **Tool Layer:** Custom pgvector retriever tool with integrated reranking logic. 3. **Storage Layer:** PostgreSQL with the vector extension. 4. **Evaluation Layer:** TruCustomApp recorder with OpenAI/HF/Gemini feedback providers. ## **3. Detailed Component Specifications** ### **3.1 Data Layer (PostgreSQL + pgvector)** * **Schema:** * id: UUID (Primary Key) * content: TEXT (Document chunk) * embedding: VECTOR(384) * metadata: JSONB (Source URL, page number, etc.) * **Search Type:** HNSW or IVFFlat index for sub-100ms vector similarity lookups. ### **3.2 Retrieval & Reranking Tool** * **First Stage:** Dense retrieval via pgvector (Cosine Similarity) returning Top 20 results. * **Second Stage:** Reranking via CrossEncoder('cross-encoder/ms-marco-MiniLM-L-6-v2') — ~85 MB, CPU-friendly, no GPU required. * **Logic:** The tool takes a string query, computes the embedding, queries PG, reranks the 20 results, and returns a formatted string of the top 5. ### **3.3 Agentic Logic (SmolAgents)** The system MUST provide a configuration toggle to switch between the following Model Providers: * **Option A: OpenAI Ecosystem** * Includes standard OpenAI (GPT-4o, GPT-4o-mini) and **Azure OpenAI** instances. * Uses OpenAIServerModel or LiteLLMModel for integration. * **Option B: Google Gemini Family** * Includes gemini-3.1-flash-lite-preview and gemini-flash-lite-latest variants. * Uses native Google Generative AI integration or LiteLLMModel. * **Reasoning:** The agent is configured with max_steps=5 to allow for query expansion or multi-step searching if the first attempt fails. ### **3.4 Evaluation Framework (TruLens 2.x OTEL)** The backend uses **TruApp** with OTEL-instrumented replay, **RAG Triad** metrics (Context Relevance, Groundedness, Answer Relevance), and an LLM-as-judge whose provider is configured separately from the agent (`TRULENS_PROVIDER`: OpenAI, Azure OpenAI, Gemini, or none). Evaluation runs **asynchronously** after the HTTP response; the worker synchronizes on TruLens feedback completion up to **`TRULENS_FEEDBACK_TIMEOUT`** before persisting to `evaluation_results`. Retrieval context for scoring prefers the **authoritative reranked top-5 tool string** when the retriever ran. * **Metrics (The RAG Triad):** * Context relevance: User query vs. reranked retrieval context. * Groundedness: Final answer vs. reranked retrieval context. * Answer relevance: User query vs. final answer. ## **4. Technical Constraints** * **Environment:** Python 3.10+. * **API-First Architecture:** Inference for the primary Orchestrator and Evaluation Judge is handled via external APIs. Local compute is reserved for Embedding and Reranking models. * **Credential Management:** Secure handling of API keys via environment variables: * OPENAI_API_KEY / AZURE_OPENAI_API_KEY * AZURE_OPENAI_ENDPOINT * GOOGLE_API_KEY * **Connectivity:** The execution environment requires outbound HTTPS access to api.openai.com, \*.openai.azure.com, and generativelanguage.googleapis.com. * **Memory (Hybrid):** Reranking model (Stage 2) is CPU-only (~85 MB RAM); no GPU required. LLM reasoning is offloaded to the cloud. * **Security:** Agent code execution must be sandboxed. API keys must not be logged or exposed in the TruLens dashboard traces. ## **5. Performance Requirements** * **Reranking Latency:** \< 500ms for 20 document pairs (local inference). * **Vector Query:** \< 50ms for 100k vectors. * **Inference Latency:** Target \< 3s for LLM generation steps; depends on provider (OpenAI vs. Gemini) and network overhead. * **Evaluation Overhead:** TruLens recording should be asynchronous or sampled to avoid blocking the user response.