CodeBase-Agent / README.md
armaanalam's picture
Update README.md
4deadc0 verified
|
Raw History Blame
9.06 kB
# AI Codebase Assistant
An AI-powered developer tool that lets you interact with an entire code repository using **RAG (Retrieval-Augmented Generation)** and LLMs.
It can answer questions about your codebase, detect potential bugs, analyze code complexity, explain functions, generate documentation, create new files, and interact with GitHub through the GitHub MCP API.
---
## Features
### Codebase Question Answering
Ask natural-language questions about your repository.
**Example:**
```text
How does authentication work?
```
The assistant searches the relevant parts of the codebase and generates an answer with source references.
---
### Bug Detection
Analyze a source file and identify potential issues using an LLM-powered code review.
Example:
```text
[HIGH] Line 34
SQL query uses string concatenation
Recommendation:
Use parameterized queries to prevent SQL injection.
```
The system returns the issue severity, location, and recommendation.
---
### Cyclomatic Complexity Analysis
Analyze the complexity of Python functions using **Radon**.
Example:
```text
[A] process_order complexity = 2
[B] validate_cart complexity = 5
[C] apply_discounts complexity = 8
[F] handle_edge_cases complexity = 18
```
This helps identify functions that may be difficult to maintain or test.
---
### Function Explanation
Select a function and get a plain-English explanation of what it does.
For Python code, the project uses the `ast` module to extract the function and provides relevant repository context to the LLM.
---
### Documentation Generation
Generate documentation automatically for your codebase.
The assistant can generate:
- Module documentation
- Function explanations
- Docstrings
- Repository README files
---
### AI File Generation
Describe the file you want to create and let the AI generate it.
Example:
```text
Create a Rectangle class in JavaScript
that calculates area and perimeter.
```
The generated file is shown before it is written to the repository, allowing the user to approve it first.
---
### GitHub MCP Integration
The project integrates with the **GitHub Copilot MCP API**.
It currently supports access to **44 GitHub MCP tools**, including:
- `search_code`
- `search_repositories`
- `get_file_contents`
- `list_issues`
- `create_branch`
- `create_pull_request`
- `push_files`
- `fork_repository`
---
## RAG-Based Code Search
The project uses **Retrieval-Augmented Generation (RAG)** to work with repository-level code.
Repository files are:
1. Loaded and filtered
2. Split into smaller chunks
3. Converted into embeddings
4. Stored in ChromaDB
5. Retrieved based on semantic similarity when a question is asked
The retrieved code is then provided to the LLM as context.
Source metadata such as file path and line range is preserved during this process.
---
## Supported Languages
The repository ingestion system currently supports:
| Extension | Language |
|---|---|
| `.py` | Python |
| `.js` | JavaScript |
| `.ts` | TypeScript |
| `.java` | Java |
| `.go` | Go |
| `.md` | Markdown |
---
## Tech Stack
### AI / LLM
- Groq
- Llama 3.3 70B
- Google Gemini
### RAG
- LangChain
- Google Gemini Embeddings
- ChromaDB
### Backend
- Python
- FastAPI
- Pydantic Settings
### Code Analysis
- Python AST
- Radon
- LLM-based code analysis
### Integration
- GitHub MCP
- GitHub Copilot MCP API
---
## LLM Providers
| Provider | Model | Usage |
|---|---|---|
| Groq | `llama-3.3-70b-versatile` | LLM generation |
| Gemini | Configurable | LLM generation |
| Gemini Embeddings | `models/gemini-embedding-001` | Code embeddings |
The project supports switching between Groq and Gemini for LLM generation.
Gemini Embeddings are used for semantic retrieval.
---
# Installation
## 1. Clone the Repository
```bash
git clone <your-repository-url>
cd "Codebase Assistant"
```
## 2. Create a Virtual Environment
### Windows
```bash
python -m venv .venv
.venv\Scripts\Activate.ps1
```
### Linux / macOS
```bash
python3 -m venv .venv
source .venv/bin/activate
```
## 3. Install Dependencies
```bash
pip install -r requirements.txt
```
---
## πŸ—οΈ Architecture
![AI Codebase Assistant Architecture](./architecture.png)
The system combines a RAG pipeline, LLM providers, code analysis services,
FastAPI, and GitHub MCP integration to provide repository-level AI assistance.
# Configuration
Create a `.env` file in the project root.
```env
LLM_PROVIDER=groq
GROQ_API_KEY=your_groq_api_key
GROQ_MODEL=llama-3.3-70b-versatile
GEMINI_API_KEY=your_gemini_api_key
EMBEDDING_MODEL=models/gemini-embedding-001
VECTOR_DB_PATH=./vector_db
GITHUB_MCP_TOKEN=your_github_token
```
### Required API Keys
**Groq**
Used for LLM generation when:
```env
LLM_PROVIDER=groq
```
**Gemini**
Required for embeddings because the project uses Gemini Embeddings for repository indexing.
**GitHub Token**
Required for GitHub MCP functionality.
---
# Running the CLI
Start the interactive CLI:
```bash
python cli.py
```
The application will ask for the repository you want to analyze.
```text
β–Ά Enter the path to the repository you want to analyse:
```
After the repository is indexed, you can choose from:
```text
1 Ask a question about the codebase
2 Detect bugs in a file
3 Cyclomatic complexity analysis
4 Explain a function
5 Generate module documentation
6 Generate README
7 Propose & create a new file
8 List GitHub MCP tools
9 Re-ingest repository
0 Exit
```
---
# REST API
The project also provides a FastAPI REST API.
Start the server:
```bash
python -m uvicorn app:app --reload --port 8000
```
Open the interactive API documentation:
```text
http://localhost:8000/docs
```
## API Endpoints
| Method | Endpoint | Description |
|---|---|---|
| `GET` | `/health` | Health check |
| `POST` | `/api/query` | Ask questions about the codebase |
| `POST` | `/api/analyze/bugs` | Detect potential bugs |
| `POST` | `/api/analyze/complexity` | Analyze cyclomatic complexity |
| `POST` | `/api/analyze/explain` | Explain a function |
| `POST` | `/api/docs/module` | Generate module documentation |
| `POST` | `/api/docs/readme` | Generate README |
| `POST` | `/api/files/propose` | Generate a file proposal |
| `POST` | `/api/files/approve` | Write an approved file |
---
## Example API Request
```bash
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-d '{"query": "How does authentication work?", "k": 5}'
```
Example response:
```json
{
"answer": "Authentication is handled by ...",
"sources": [
{
"file_path": "src/auth.py",
"start_line": 10,
"end_line": 45
}
]
}
```
---
# πŸ“ Project Structure
```text
Codebase Assistant/
β”‚
β”œβ”€β”€ cli.py
β”œβ”€β”€ app.py
β”œβ”€β”€ config.py
β”œβ”€β”€ llm.py
β”œβ”€β”€ mcp_client.py
β”œβ”€β”€ requirements.txt
β”‚
β”œβ”€β”€ rag/
β”‚ β”œβ”€β”€ repository_loader.py
β”‚ β”œβ”€β”€ splitter.py
β”‚ β”œβ”€β”€ embedding.py
β”‚ β”œβ”€β”€ retriever.py
β”‚ └── rag_chain.py
β”‚
β”œβ”€β”€ services/
β”‚ β”œβ”€β”€ code_analysis.py
β”‚ β”œβ”€β”€ documentation.py
β”‚ └── file_creator.py
β”‚
β”œβ”€β”€ api/
β”‚ └── routes.py
β”‚
└── vector_db/
└── chroma.sqlite3
```
---
# Configuration Options
| Setting | Default | Description |
|---|---|---|
| `LLM_PROVIDER` | `groq` | LLM provider |
| `GROQ_MODEL` | `llama-3.3-70b-versatile` | Groq model |
| `EMBEDDING_MODEL` | `models/gemini-embedding-001` | Embedding model |
| `VECTOR_DB_PATH` | `./vector_db` | ChromaDB storage |
| `CHUNK_SIZE` | `1200` | Chunk size |
| `CHUNK_OVERLAP` | `100` | Chunk overlap |
| `MAX_FILE_SIZE_KB` | `1042` | Maximum file size |
| `ALLOWED_EXTENSIONS` | `.py,.js,.ts,.go,.java,.md` | Supported files |
---
# Limitations
- Repository ingestion currently runs locally.
- External API keys are required for LLM and embedding services.
- Gemini embedding quotas may limit large repositories.
- Only the currently supported file types are indexed.
- LLM-generated code and bug reports should be reviewed before use.
- The vector database currently focuses on the actively ingested repository.
---
# Future Improvements
Planned or possible improvements include:
- More programming language support
- Incremental repository indexing
- GitHub repository ingestion
- AST-based code chunking
- Hybrid keyword + vector search
- Retrieval reranking
- Dependency graph analysis
- Automated test generation
- Automated pull request review
- Multi-repository search
- Web-based developer interface
---
# Project Goal
The goal of this project is to build an AI developer assistant that can understand and interact with an entire codebase rather than only individual code snippets.
It combines:
**RAG + LLMs + Vector Search + Code Analysis + FastAPI + MCP**
into a single developer-focused tool.
---
# Author
**Armaan Alam**
AI Engineer & Software Developer
Interested in:
- Generative AI
- RAG Systems
- Backend Engineering
- LLM Applications
- AI Developer Tools
- Machine Learning