Download README.md from armaanalam/CodeBase-Agent: direct link, hf CLI and curl.
- Browser
- Download file 9.06 kB
-
https://huggingface.co/armaanalam/CodeBase-Agent/resolve/refs%2Fpr%2F1/README.md
- Command line
-
hf download hf://armaanalam/CodeBase-Agent@refs/pr/1/README.md
-
curl -L -o README.md https://huggingface.co/armaanalam/CodeBase-Agent/resolve/refs%2Fpr%2F1/README.md
AI Codebase Assistant
An AI-powered developer tool that lets you interact with an entire code repository using RAG (Retrieval-Augmented Generation) and LLMs.
It can answer questions about your codebase, detect potential bugs, analyze code complexity, explain functions, generate documentation, create new files, and interact with GitHub through the GitHub MCP API.
Features
Codebase Question Answering
Ask natural-language questions about your repository.
Example:
How does authentication work?
The assistant searches the relevant parts of the codebase and generates an answer with source references.
Bug Detection
Analyze a source file and identify potential issues using an LLM-powered code review.
Example:
[HIGH] Line 34
SQL query uses string concatenation
Recommendation:
Use parameterized queries to prevent SQL injection.
The system returns the issue severity, location, and recommendation.
Cyclomatic Complexity Analysis
Analyze the complexity of Python functions using Radon.
Example:
[A] process_order complexity = 2
[B] validate_cart complexity = 5
[C] apply_discounts complexity = 8
[F] handle_edge_cases complexity = 18
This helps identify functions that may be difficult to maintain or test.
Function Explanation
Select a function and get a plain-English explanation of what it does.
For Python code, the project uses the ast module to extract the function and provides relevant repository context to the LLM.
Documentation Generation
Generate documentation automatically for your codebase.
The assistant can generate:
- Module documentation
- Function explanations
- Docstrings
- Repository README files
AI File Generation
Describe the file you want to create and let the AI generate it.
Example:
Create a Rectangle class in JavaScript
that calculates area and perimeter.
The generated file is shown before it is written to the repository, allowing the user to approve it first.
GitHub MCP Integration
The project integrates with the GitHub Copilot MCP API.
It currently supports access to 44 GitHub MCP tools, including:
search_codesearch_repositoriesget_file_contentslist_issuescreate_branchcreate_pull_requestpush_filesfork_repository
RAG-Based Code Search
The project uses Retrieval-Augmented Generation (RAG) to work with repository-level code.
Repository files are:
- Loaded and filtered
- Split into smaller chunks
- Converted into embeddings
- Stored in ChromaDB
- Retrieved based on semantic similarity when a question is asked
The retrieved code is then provided to the LLM as context.
Source metadata such as file path and line range is preserved during this process.
Supported Languages
The repository ingestion system currently supports:
| Extension | Language |
|---|---|
.py |
Python |
.js |
JavaScript |
.ts |
TypeScript |
.java |
Java |
.go |
Go |
.md |
Markdown |
Tech Stack
AI / LLM
- Groq
- Llama 3.3 70B
- Google Gemini
RAG
- LangChain
- Google Gemini Embeddings
- ChromaDB
Backend
- Python
- FastAPI
- Pydantic Settings
Code Analysis
- Python AST
- Radon
- LLM-based code analysis
Integration
- GitHub MCP
- GitHub Copilot MCP API
LLM Providers
| Provider | Model | Usage |
|---|---|---|
| Groq | llama-3.3-70b-versatile |
LLM generation |
| Gemini | Configurable | LLM generation |
| Gemini Embeddings | models/gemini-embedding-001 |
Code embeddings |
The project supports switching between Groq and Gemini for LLM generation.
Gemini Embeddings are used for semantic retrieval.
Installation
1. Clone the Repository
git clone <your-repository-url>
cd "Codebase Assistant"
2. Create a Virtual Environment
Windows
python -m venv .venv
.venv\Scripts\Activate.ps1
Linux / macOS
python3 -m venv .venv
source .venv/bin/activate
3. Install Dependencies
pip install -r requirements.txt
ποΈ Architecture
The system combines a RAG pipeline, LLM providers, code analysis services, FastAPI, and GitHub MCP integration to provide repository-level AI assistance.
Configuration
Create a .env file in the project root.
LLM_PROVIDER=groq
GROQ_API_KEY=your_groq_api_key
GROQ_MODEL=llama-3.3-70b-versatile
GEMINI_API_KEY=your_gemini_api_key
EMBEDDING_MODEL=models/gemini-embedding-001
VECTOR_DB_PATH=./vector_db
GITHUB_MCP_TOKEN=your_github_token
Required API Keys
Groq
Used for LLM generation when:
LLM_PROVIDER=groq
Gemini
Required for embeddings because the project uses Gemini Embeddings for repository indexing.
GitHub Token
Required for GitHub MCP functionality.
Running the CLI
Start the interactive CLI:
python cli.py
The application will ask for the repository you want to analyze.
βΆ Enter the path to the repository you want to analyse:
After the repository is indexed, you can choose from:
1 Ask a question about the codebase
2 Detect bugs in a file
3 Cyclomatic complexity analysis
4 Explain a function
5 Generate module documentation
6 Generate README
7 Propose & create a new file
8 List GitHub MCP tools
9 Re-ingest repository
0 Exit
REST API
The project also provides a FastAPI REST API.
Start the server:
python -m uvicorn app:app --reload --port 8000
Open the interactive API documentation:
http://localhost:8000/docs
API Endpoints
| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Health check |
POST |
/api/query |
Ask questions about the codebase |
POST |
/api/analyze/bugs |
Detect potential bugs |
POST |
/api/analyze/complexity |
Analyze cyclomatic complexity |
POST |
/api/analyze/explain |
Explain a function |
POST |
/api/docs/module |
Generate module documentation |
POST |
/api/docs/readme |
Generate README |
POST |
/api/files/propose |
Generate a file proposal |
POST |
/api/files/approve |
Write an approved file |
Example API Request
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-d '{"query": "How does authentication work?", "k": 5}'
Example response:
{
"answer": "Authentication is handled by ...",
"sources": [
{
"file_path": "src/auth.py",
"start_line": 10,
"end_line": 45
}
]
}
π Project Structure
Codebase Assistant/
β
βββ cli.py
βββ app.py
βββ config.py
βββ llm.py
βββ mcp_client.py
βββ requirements.txt
β
βββ rag/
β βββ repository_loader.py
β βββ splitter.py
β βββ embedding.py
β βββ retriever.py
β βββ rag_chain.py
β
βββ services/
β βββ code_analysis.py
β βββ documentation.py
β βββ file_creator.py
β
βββ api/
β βββ routes.py
β
βββ vector_db/
βββ chroma.sqlite3
Configuration Options
| Setting | Default | Description |
|---|---|---|
LLM_PROVIDER |
groq |
LLM provider |
GROQ_MODEL |
llama-3.3-70b-versatile |
Groq model |
EMBEDDING_MODEL |
models/gemini-embedding-001 |
Embedding model |
VECTOR_DB_PATH |
./vector_db |
ChromaDB storage |
CHUNK_SIZE |
1200 |
Chunk size |
CHUNK_OVERLAP |
100 |
Chunk overlap |
MAX_FILE_SIZE_KB |
1042 |
Maximum file size |
ALLOWED_EXTENSIONS |
.py,.js,.ts,.go,.java,.md |
Supported files |
Limitations
- Repository ingestion currently runs locally.
- External API keys are required for LLM and embedding services.
- Gemini embedding quotas may limit large repositories.
- Only the currently supported file types are indexed.
- LLM-generated code and bug reports should be reviewed before use.
- The vector database currently focuses on the actively ingested repository.
Future Improvements
Planned or possible improvements include:
- More programming language support
- Incremental repository indexing
- GitHub repository ingestion
- AST-based code chunking
- Hybrid keyword + vector search
- Retrieval reranking
- Dependency graph analysis
- Automated test generation
- Automated pull request review
- Multi-repository search
- Web-based developer interface
Project Goal
The goal of this project is to build an AI developer assistant that can understand and interact with an entire codebase rather than only individual code snippets.
It combines:
RAG + LLMs + Vector Search + Code Analysis + FastAPI + MCP
into a single developer-focused tool.
Author
Armaan Alam
AI Engineer & Software Developer
Interested in:
- Generative AI
- RAG Systems
- Backend Engineering
- LLM Applications
- AI Developer Tools
- Machine Learning
