|
Download README.md from armaanalam/CodeBase-Agent: direct link, hf CLI and curl.
- Browser
- Download file 9.06 kB
-
https://huggingface.co/armaanalam/CodeBase-Agent/resolve/refs%2Fpr%2F1/README.md
- Command line
-
hf download hf://armaanalam/CodeBase-Agent@refs/pr/1/README.md
-
curl -L -o README.md https://huggingface.co/armaanalam/CodeBase-Agent/resolve/refs%2Fpr%2F1/README.md
9.06 kB
| # AI Codebase Assistant | |
| An AI-powered developer tool that lets you interact with an entire code repository using **RAG (Retrieval-Augmented Generation)** and LLMs. | |
| It can answer questions about your codebase, detect potential bugs, analyze code complexity, explain functions, generate documentation, create new files, and interact with GitHub through the GitHub MCP API. | |
| --- | |
| ## Features | |
| ### Codebase Question Answering | |
| Ask natural-language questions about your repository. | |
| **Example:** | |
| ```text | |
| How does authentication work? | |
| ``` | |
| The assistant searches the relevant parts of the codebase and generates an answer with source references. | |
| --- | |
| ### Bug Detection | |
| Analyze a source file and identify potential issues using an LLM-powered code review. | |
| Example: | |
| ```text | |
| [HIGH] Line 34 | |
| SQL query uses string concatenation | |
| Recommendation: | |
| Use parameterized queries to prevent SQL injection. | |
| ``` | |
| The system returns the issue severity, location, and recommendation. | |
| --- | |
| ### Cyclomatic Complexity Analysis | |
| Analyze the complexity of Python functions using **Radon**. | |
| Example: | |
| ```text | |
| [A] process_order complexity = 2 | |
| [B] validate_cart complexity = 5 | |
| [C] apply_discounts complexity = 8 | |
| [F] handle_edge_cases complexity = 18 | |
| ``` | |
| This helps identify functions that may be difficult to maintain or test. | |
| --- | |
| ### Function Explanation | |
| Select a function and get a plain-English explanation of what it does. | |
| For Python code, the project uses the `ast` module to extract the function and provides relevant repository context to the LLM. | |
| --- | |
| ### Documentation Generation | |
| Generate documentation automatically for your codebase. | |
| The assistant can generate: | |
| - Module documentation | |
| - Function explanations | |
| - Docstrings | |
| - Repository README files | |
| --- | |
| ### AI File Generation | |
| Describe the file you want to create and let the AI generate it. | |
| Example: | |
| ```text | |
| Create a Rectangle class in JavaScript | |
| that calculates area and perimeter. | |
| ``` | |
| The generated file is shown before it is written to the repository, allowing the user to approve it first. | |
| --- | |
| ### GitHub MCP Integration | |
| The project integrates with the **GitHub Copilot MCP API**. | |
| It currently supports access to **44 GitHub MCP tools**, including: | |
| - `search_code` | |
| - `search_repositories` | |
| - `get_file_contents` | |
| - `list_issues` | |
| - `create_branch` | |
| - `create_pull_request` | |
| - `push_files` | |
| - `fork_repository` | |
| --- | |
| ## RAG-Based Code Search | |
| The project uses **Retrieval-Augmented Generation (RAG)** to work with repository-level code. | |
| Repository files are: | |
| 1. Loaded and filtered | |
| 2. Split into smaller chunks | |
| 3. Converted into embeddings | |
| 4. Stored in ChromaDB | |
| 5. Retrieved based on semantic similarity when a question is asked | |
| The retrieved code is then provided to the LLM as context. | |
| Source metadata such as file path and line range is preserved during this process. | |
| --- | |
| ## Supported Languages | |
| The repository ingestion system currently supports: | |
| | Extension | Language | | |
| |---|---| | |
| | `.py` | Python | | |
| | `.js` | JavaScript | | |
| | `.ts` | TypeScript | | |
| | `.java` | Java | | |
| | `.go` | Go | | |
| | `.md` | Markdown | | |
| --- | |
| ## Tech Stack | |
| ### AI / LLM | |
| - Groq | |
| - Llama 3.3 70B | |
| - Google Gemini | |
| ### RAG | |
| - LangChain | |
| - Google Gemini Embeddings | |
| - ChromaDB | |
| ### Backend | |
| - Python | |
| - FastAPI | |
| - Pydantic Settings | |
| ### Code Analysis | |
| - Python AST | |
| - Radon | |
| - LLM-based code analysis | |
| ### Integration | |
| - GitHub MCP | |
| - GitHub Copilot MCP API | |
| --- | |
| ## LLM Providers | |
| | Provider | Model | Usage | | |
| |---|---|---| | |
| | Groq | `llama-3.3-70b-versatile` | LLM generation | | |
| | Gemini | Configurable | LLM generation | | |
| | Gemini Embeddings | `models/gemini-embedding-001` | Code embeddings | | |
| The project supports switching between Groq and Gemini for LLM generation. | |
| Gemini Embeddings are used for semantic retrieval. | |
| --- | |
| # Installation | |
| ## 1. Clone the Repository | |
| ```bash | |
| git clone <your-repository-url> | |
| cd "Codebase Assistant" | |
| ``` | |
| ## 2. Create a Virtual Environment | |
| ### Windows | |
| ```bash | |
| python -m venv .venv | |
| .venv\Scripts\Activate.ps1 | |
| ``` | |
| ### Linux / macOS | |
| ```bash | |
| python3 -m venv .venv | |
| source .venv/bin/activate | |
| ``` | |
| ## 3. Install Dependencies | |
| ```bash | |
| pip install -r requirements.txt | |
| ``` | |
| --- | |
| ## ποΈ Architecture | |
|  | |
| The system combines a RAG pipeline, LLM providers, code analysis services, | |
| FastAPI, and GitHub MCP integration to provide repository-level AI assistance. | |
| # Configuration | |
| Create a `.env` file in the project root. | |
| ```env | |
| LLM_PROVIDER=groq | |
| GROQ_API_KEY=your_groq_api_key | |
| GROQ_MODEL=llama-3.3-70b-versatile | |
| GEMINI_API_KEY=your_gemini_api_key | |
| EMBEDDING_MODEL=models/gemini-embedding-001 | |
| VECTOR_DB_PATH=./vector_db | |
| GITHUB_MCP_TOKEN=your_github_token | |
| ``` | |
| ### Required API Keys | |
| **Groq** | |
| Used for LLM generation when: | |
| ```env | |
| LLM_PROVIDER=groq | |
| ``` | |
| **Gemini** | |
| Required for embeddings because the project uses Gemini Embeddings for repository indexing. | |
| **GitHub Token** | |
| Required for GitHub MCP functionality. | |
| --- | |
| # Running the CLI | |
| Start the interactive CLI: | |
| ```bash | |
| python cli.py | |
| ``` | |
| The application will ask for the repository you want to analyze. | |
| ```text | |
| βΆ Enter the path to the repository you want to analyse: | |
| ``` | |
| After the repository is indexed, you can choose from: | |
| ```text | |
| 1 Ask a question about the codebase | |
| 2 Detect bugs in a file | |
| 3 Cyclomatic complexity analysis | |
| 4 Explain a function | |
| 5 Generate module documentation | |
| 6 Generate README | |
| 7 Propose & create a new file | |
| 8 List GitHub MCP tools | |
| 9 Re-ingest repository | |
| 0 Exit | |
| ``` | |
| --- | |
| # REST API | |
| The project also provides a FastAPI REST API. | |
| Start the server: | |
| ```bash | |
| python -m uvicorn app:app --reload --port 8000 | |
| ``` | |
| Open the interactive API documentation: | |
| ```text | |
| http://localhost:8000/docs | |
| ``` | |
| ## API Endpoints | |
| | Method | Endpoint | Description | | |
| |---|---|---| | |
| | `GET` | `/health` | Health check | | |
| | `POST` | `/api/query` | Ask questions about the codebase | | |
| | `POST` | `/api/analyze/bugs` | Detect potential bugs | | |
| | `POST` | `/api/analyze/complexity` | Analyze cyclomatic complexity | | |
| | `POST` | `/api/analyze/explain` | Explain a function | | |
| | `POST` | `/api/docs/module` | Generate module documentation | | |
| | `POST` | `/api/docs/readme` | Generate README | | |
| | `POST` | `/api/files/propose` | Generate a file proposal | | |
| | `POST` | `/api/files/approve` | Write an approved file | | |
| --- | |
| ## Example API Request | |
| ```bash | |
| curl -X POST http://localhost:8000/api/query \ | |
| -H "Content-Type: application/json" \ | |
| -d '{"query": "How does authentication work?", "k": 5}' | |
| ``` | |
| Example response: | |
| ```json | |
| { | |
| "answer": "Authentication is handled by ...", | |
| "sources": [ | |
| { | |
| "file_path": "src/auth.py", | |
| "start_line": 10, | |
| "end_line": 45 | |
| } | |
| ] | |
| } | |
| ``` | |
| --- | |
| # π Project Structure | |
| ```text | |
| Codebase Assistant/ | |
| β | |
| βββ cli.py | |
| βββ app.py | |
| βββ config.py | |
| βββ llm.py | |
| βββ mcp_client.py | |
| βββ requirements.txt | |
| β | |
| βββ rag/ | |
| β βββ repository_loader.py | |
| β βββ splitter.py | |
| β βββ embedding.py | |
| β βββ retriever.py | |
| β βββ rag_chain.py | |
| β | |
| βββ services/ | |
| β βββ code_analysis.py | |
| β βββ documentation.py | |
| β βββ file_creator.py | |
| β | |
| βββ api/ | |
| β βββ routes.py | |
| β | |
| βββ vector_db/ | |
| βββ chroma.sqlite3 | |
| ``` | |
| --- | |
| # Configuration Options | |
| | Setting | Default | Description | | |
| |---|---|---| | |
| | `LLM_PROVIDER` | `groq` | LLM provider | | |
| | `GROQ_MODEL` | `llama-3.3-70b-versatile` | Groq model | | |
| | `EMBEDDING_MODEL` | `models/gemini-embedding-001` | Embedding model | | |
| | `VECTOR_DB_PATH` | `./vector_db` | ChromaDB storage | | |
| | `CHUNK_SIZE` | `1200` | Chunk size | | |
| | `CHUNK_OVERLAP` | `100` | Chunk overlap | | |
| | `MAX_FILE_SIZE_KB` | `1042` | Maximum file size | | |
| | `ALLOWED_EXTENSIONS` | `.py,.js,.ts,.go,.java,.md` | Supported files | | |
| --- | |
| # Limitations | |
| - Repository ingestion currently runs locally. | |
| - External API keys are required for LLM and embedding services. | |
| - Gemini embedding quotas may limit large repositories. | |
| - Only the currently supported file types are indexed. | |
| - LLM-generated code and bug reports should be reviewed before use. | |
| - The vector database currently focuses on the actively ingested repository. | |
| --- | |
| # Future Improvements | |
| Planned or possible improvements include: | |
| - More programming language support | |
| - Incremental repository indexing | |
| - GitHub repository ingestion | |
| - AST-based code chunking | |
| - Hybrid keyword + vector search | |
| - Retrieval reranking | |
| - Dependency graph analysis | |
| - Automated test generation | |
| - Automated pull request review | |
| - Multi-repository search | |
| - Web-based developer interface | |
| --- | |
| # Project Goal | |
| The goal of this project is to build an AI developer assistant that can understand and interact with an entire codebase rather than only individual code snippets. | |
| It combines: | |
| **RAG + LLMs + Vector Search + Code Analysis + FastAPI + MCP** | |
| into a single developer-focused tool. | |
| --- | |
| # Author | |
| **Armaan Alam** | |
| AI Engineer & Software Developer | |
| Interested in: | |
| - Generative AI | |
| - RAG Systems | |
| - Backend Engineering | |
| - LLM Applications | |
| - AI Developer Tools | |
| - Machine Learning |