# AI Codebase Assistant An AI-powered developer tool that lets you interact with an entire code repository using **RAG (Retrieval-Augmented Generation)** and LLMs. It can answer questions about your codebase, detect potential bugs, analyze code complexity, explain functions, generate documentation, create new files, and interact with GitHub through the GitHub MCP API. --- ## Features ### Codebase Question Answering Ask natural-language questions about your repository. **Example:** ```text How does authentication work? ``` The assistant searches the relevant parts of the codebase and generates an answer with source references. --- ### Bug Detection Analyze a source file and identify potential issues using an LLM-powered code review. Example: ```text [HIGH] Line 34 SQL query uses string concatenation Recommendation: Use parameterized queries to prevent SQL injection. ``` The system returns the issue severity, location, and recommendation. --- ### Cyclomatic Complexity Analysis Analyze the complexity of Python functions using **Radon**. Example: ```text [A] process_order complexity = 2 [B] validate_cart complexity = 5 [C] apply_discounts complexity = 8 [F] handle_edge_cases complexity = 18 ``` This helps identify functions that may be difficult to maintain or test. --- ### Function Explanation Select a function and get a plain-English explanation of what it does. For Python code, the project uses the `ast` module to extract the function and provides relevant repository context to the LLM. --- ### Documentation Generation Generate documentation automatically for your codebase. The assistant can generate: - Module documentation - Function explanations - Docstrings - Repository README files --- ### AI File Generation Describe the file you want to create and let the AI generate it. Example: ```text Create a Rectangle class in JavaScript that calculates area and perimeter. ``` The generated file is shown before it is written to the repository, allowing the user to approve it first. --- ### GitHub MCP Integration The project integrates with the **GitHub Copilot MCP API**. It currently supports access to **44 GitHub MCP tools**, including: - `search_code` - `search_repositories` - `get_file_contents` - `list_issues` - `create_branch` - `create_pull_request` - `push_files` - `fork_repository` --- ## RAG-Based Code Search The project uses **Retrieval-Augmented Generation (RAG)** to work with repository-level code. Repository files are: 1. Loaded and filtered 2. Split into smaller chunks 3. Converted into embeddings 4. Stored in ChromaDB 5. Retrieved based on semantic similarity when a question is asked The retrieved code is then provided to the LLM as context. Source metadata such as file path and line range is preserved during this process. --- ## Supported Languages The repository ingestion system currently supports: | Extension | Language | |---|---| | `.py` | Python | | `.js` | JavaScript | | `.ts` | TypeScript | | `.java` | Java | | `.go` | Go | | `.md` | Markdown | --- ## Tech Stack ### AI / LLM - Groq - Llama 3.3 70B - Google Gemini ### RAG - LangChain - Google Gemini Embeddings - ChromaDB ### Backend - Python - FastAPI - Pydantic Settings ### Code Analysis - Python AST - Radon - LLM-based code analysis ### Integration - GitHub MCP - GitHub Copilot MCP API --- ## LLM Providers | Provider | Model | Usage | |---|---|---| | Groq | `llama-3.3-70b-versatile` | LLM generation | | Gemini | Configurable | LLM generation | | Gemini Embeddings | `models/gemini-embedding-001` | Code embeddings | The project supports switching between Groq and Gemini for LLM generation. Gemini Embeddings are used for semantic retrieval. --- # Installation ## 1. Clone the Repository ```bash git clone cd "Codebase Assistant" ``` ## 2. Create a Virtual Environment ### Windows ```bash python -m venv .venv .venv\Scripts\Activate.ps1 ``` ### Linux / macOS ```bash python3 -m venv .venv source .venv/bin/activate ``` ## 3. Install Dependencies ```bash pip install -r requirements.txt ``` --- ## 🏗️ Architecture ![AI Codebase Assistant Architecture](./architecture.png) The system combines a RAG pipeline, LLM providers, code analysis services, FastAPI, and GitHub MCP integration to provide repository-level AI assistance. # Configuration Create a `.env` file in the project root. ```env LLM_PROVIDER=groq GROQ_API_KEY=your_groq_api_key GROQ_MODEL=llama-3.3-70b-versatile GEMINI_API_KEY=your_gemini_api_key EMBEDDING_MODEL=models/gemini-embedding-001 VECTOR_DB_PATH=./vector_db GITHUB_MCP_TOKEN=your_github_token ``` ### Required API Keys **Groq** Used for LLM generation when: ```env LLM_PROVIDER=groq ``` **Gemini** Required for embeddings because the project uses Gemini Embeddings for repository indexing. **GitHub Token** Required for GitHub MCP functionality. --- # Running the CLI Start the interactive CLI: ```bash python cli.py ``` The application will ask for the repository you want to analyze. ```text ▶ Enter the path to the repository you want to analyse: ``` After the repository is indexed, you can choose from: ```text 1 Ask a question about the codebase 2 Detect bugs in a file 3 Cyclomatic complexity analysis 4 Explain a function 5 Generate module documentation 6 Generate README 7 Propose & create a new file 8 List GitHub MCP tools 9 Re-ingest repository 0 Exit ``` --- # REST API The project also provides a FastAPI REST API. Start the server: ```bash python -m uvicorn app:app --reload --port 8000 ``` Open the interactive API documentation: ```text http://localhost:8000/docs ``` ## API Endpoints | Method | Endpoint | Description | |---|---|---| | `GET` | `/health` | Health check | | `POST` | `/api/query` | Ask questions about the codebase | | `POST` | `/api/analyze/bugs` | Detect potential bugs | | `POST` | `/api/analyze/complexity` | Analyze cyclomatic complexity | | `POST` | `/api/analyze/explain` | Explain a function | | `POST` | `/api/docs/module` | Generate module documentation | | `POST` | `/api/docs/readme` | Generate README | | `POST` | `/api/files/propose` | Generate a file proposal | | `POST` | `/api/files/approve` | Write an approved file | --- ## Example API Request ```bash curl -X POST http://localhost:8000/api/query \ -H "Content-Type: application/json" \ -d '{"query": "How does authentication work?", "k": 5}' ``` Example response: ```json { "answer": "Authentication is handled by ...", "sources": [ { "file_path": "src/auth.py", "start_line": 10, "end_line": 45 } ] } ``` --- # 📁 Project Structure ```text Codebase Assistant/ │ ├── cli.py ├── app.py ├── config.py ├── llm.py ├── mcp_client.py ├── requirements.txt │ ├── rag/ │ ├── repository_loader.py │ ├── splitter.py │ ├── embedding.py │ ├── retriever.py │ └── rag_chain.py │ ├── services/ │ ├── code_analysis.py │ ├── documentation.py │ └── file_creator.py │ ├── api/ │ └── routes.py │ └── vector_db/ └── chroma.sqlite3 ``` --- # Configuration Options | Setting | Default | Description | |---|---|---| | `LLM_PROVIDER` | `groq` | LLM provider | | `GROQ_MODEL` | `llama-3.3-70b-versatile` | Groq model | | `EMBEDDING_MODEL` | `models/gemini-embedding-001` | Embedding model | | `VECTOR_DB_PATH` | `./vector_db` | ChromaDB storage | | `CHUNK_SIZE` | `1200` | Chunk size | | `CHUNK_OVERLAP` | `100` | Chunk overlap | | `MAX_FILE_SIZE_KB` | `1042` | Maximum file size | | `ALLOWED_EXTENSIONS` | `.py,.js,.ts,.go,.java,.md` | Supported files | --- # Limitations - Repository ingestion currently runs locally. - External API keys are required for LLM and embedding services. - Gemini embedding quotas may limit large repositories. - Only the currently supported file types are indexed. - LLM-generated code and bug reports should be reviewed before use. - The vector database currently focuses on the actively ingested repository. --- # Future Improvements Planned or possible improvements include: - More programming language support - Incremental repository indexing - GitHub repository ingestion - AST-based code chunking - Hybrid keyword + vector search - Retrieval reranking - Dependency graph analysis - Automated test generation - Automated pull request review - Multi-repository search - Web-based developer interface --- # Project Goal The goal of this project is to build an AI developer assistant that can understand and interact with an entire codebase rather than only individual code snippets. It combines: **RAG + LLMs + Vector Search + Code Analysis + FastAPI + MCP** into a single developer-focused tool. --- # Author **Armaan Alam** AI Engineer & Software Developer Interested in: - Generative AI - RAG Systems - Backend Engineering - LLM Applications - AI Developer Tools - Machine Learning