armaanalam commited on
Commit
4deadc0
Β·
verified Β·
1 Parent(s): 8515d12

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +277 -561
README.md CHANGED
@@ -1,638 +1,337 @@
1
- # πŸ€– AI Codebase Assistant β€” Complete Documentation
2
 
3
- > An AI-powered developer tool that ingests any local code repository, answers questions about it using RAG (Retrieval-Augmented Generation), detects bugs, measures complexity, generates docs, proposes new files, and integrates with the GitHub MCP API β€” all from an interactive CLI.
 
 
4
 
5
  ---
6
 
7
- ## Table of Contents
8
 
9
- 1. [Project Overview](#1-project-overview)
10
- 2. [Repository Structure](#2-repository-structure)
11
- 3. [Architecture & Flow](#3-architecture--flow)
12
- 4. [Component Deep-Dive](#4-component-deep-dive)
13
- 5. [Setup & Installation](#5-setup--installation)
14
- 6. [Configuration (.env)](#6-configuration-env)
15
- 7. [Running the CLI](#7-running-the-cli)
16
- 8. [CLI Features β€” All 9 Options](#8-cli-features--all-9-options)
17
- 9. [REST API](#9-rest-api)
18
- 10. [Key Design Decisions & Bug Fixes](#10-key-design-decisions--bug-fixes)
19
- 11. [Extending the Project](#11-extending-the-project)
20
 
21
- ---
22
 
23
- ## 1. Project Overview
24
 
25
- The **AI Codebase Assistant** is a local developer tool that:
 
 
26
 
27
- - πŸ“‚ **Ingests** any code repository (Python, JS, TS, Java, Go, Markdown)
28
- - πŸ” **Answers questions** about the code using RAG + LLM
29
- - πŸ› **Detects bugs** via LLM-powered code review
30
- - πŸ“Š **Measures cyclomatic complexity** using Radon
31
- - πŸ“ **Explains functions** in plain English
32
- - πŸ“„ **Generates module docs and READMEs** automatically
33
- - πŸ› οΈ **Proposes & creates new files** using AI, saved to your chosen directory
34
- - πŸ”— **Lists 44 GitHub MCP tools** via the GitHub Copilot MCP API
35
 
36
- **LLM Providers supported:**
37
- | Provider | Model | Notes |
38
- |---|---|---|
39
- | Groq | `llama-3.3-70b-versatile` | Free tier, recommended |
40
- | Gemini | Configurable | Paid, stricter quota |
41
 
42
- **Embedding:** Google Gemini Embedding API (`models/gemini-embedding-001`)
43
- **Vector Store:** ChromaDB (local persistent)
44
 
45
- ---
46
 
47
- ## 2. Repository Structure
48
 
 
 
 
 
 
 
49
  ```
50
- Codebase Assistant/
51
- β”‚
52
- β”œβ”€β”€ cli.py # ← Main entry point (interactive CLI)
53
- β”œβ”€β”€ app.py # ← FastAPI REST server (optional)
54
- β”œβ”€β”€ config.py # ← All settings via pydantic-settings + .env
55
- β”œβ”€β”€ llm.py # ← LLM abstraction (Gemini / Groq) + prompt templates
56
- β”œβ”€β”€ mcp_client.py # ← GitHub MCP async context-manager client
57
- β”œβ”€β”€ requirements.txt # ← All Python dependencies
58
- β”œβ”€β”€ .env # ← API keys and configuration (not committed)
59
- β”‚
60
- β”œβ”€β”€ rag/ # ── RAG Pipeline ──────────────────────────────
61
- β”‚ β”œβ”€β”€ repository_loader.py # Walk directory, read files β†’ CodeDocument
62
- β”‚ β”œβ”€β”€ splitter.py # Split CodeDocuments into chunks with metadata
63
- β”‚ β”œβ”€β”€ embedding.py # Embed chunks/queries via Gemini Embeddings
64
- β”‚ β”œβ”€β”€ retriever.py # ChromaDB vector store: build / load / query
65
- β”‚ └── rag_chain.py # Orchestrate retrieve β†’ build context β†’ LLM
66
- β”‚
67
- β”œβ”€β”€ services/ # ── Feature Services ──────────────────────────
68
- β”‚ β”œβ”€β”€ code_analysis.py # Bug detection, complexity, function explainer
69
- β”‚ β”œβ”€β”€ documentation.py # Module docs, README generator
70
- β”‚ └── file_creator.py # AI file proposal + local disk write
71
- β”‚
72
- β”œβ”€β”€ api/ # ── REST API (FastAPI) ────────────────────────
73
- β”‚ └── routes.py # All HTTP endpoints (mirrors CLI features)
74
- β”‚
75
- └── vector_db/ # ── ChromaDB Storage (auto-created) ───────────
76
- └── chroma.sqlite3 # Persisted vector embeddings
77
- ```
78
 
79
  ---
80
 
81
- ## 3. Architecture & Flow
82
-
83
- ### 3.1 Full System Architecture
84
-
85
- ```mermaid
86
- graph TB
87
- subgraph USER["User Interface"]
88
- CLI["cli.py\n(Interactive CLI)"]
89
- API["app.py\n(FastAPI REST API)"]
90
- end
91
-
92
- subgraph RAG["RAG Pipeline"]
93
- RL["repository_loader.py\nWalk & read files"]
94
- SP["splitter.py\nChunk by language"]
95
- EM["embedding.py\nGemini Embeddings"]
96
- RT["retriever.py\nChromaDB"]
97
- RC["rag_chain.py\nOrchestrator"]
98
- end
99
-
100
- subgraph SERVICES["Feature Services"]
101
- CA["code_analysis.py\nBugs Β· Complexity Β· Explain"]
102
- DOC["documentation.py\nModule Docs Β· README"]
103
- FC["file_creator.py\nAI File Creator"]
104
- end
105
-
106
- subgraph EXTERNAL["External APIs"]
107
- GROQ["Groq API\nllama-3.3-70b"]
108
- GEM["Gemini API\nEmbeddings"]
109
- MCP["GitHub MCP\n44 tools"]
110
- end
111
-
112
- LLM["llm.py\nLLM Abstraction Layer"]
113
- CFG["config.py\n.env Settings"]
114
- DB[("vector_db/\nChromaDB")]
115
-
116
- CLI --> RAG
117
- CLI --> SERVICES
118
- API --> RAG
119
- API --> SERVICES
120
-
121
- RL --> SP --> EM --> RT
122
- RT --> DB
123
- RC --> RT
124
- RC --> LLM
125
-
126
- CA --> LLM
127
- DOC --> LLM
128
- FC --> LLM
129
-
130
- LLM --> GROQ
131
- LLM --> GEM
132
- EM --> GEM
133
- CLI --> MCP
134
-
135
- CFG -.->|settings| RAG
136
- CFG -.->|settings| LLM
137
- CFG -.->|settings| SERVICES
138
- ```
139
 
140
- ### 3.2 Ingest Flow (Step-by-step)
141
-
142
- ```mermaid
143
- flowchart LR
144
- A["User provides\nrepo path"] --> B["repository_loader.py\nwalk directories\nskip: .git venv __pycache__"]
145
- B --> C["Filter by\nallowed extensions\n.py .js .ts .java .go .md"]
146
- C --> D["Read each file\n→ CodeDocument\n(content, path, language, size)"]
147
- D --> E["splitter.py\nLanguage-aware chunking\nRecursiveCharacterTextSplitter"]
148
- E --> F["Each chunk gets\nmetadata: file_path\nlanguage Β· start_line Β· end_line"]
149
- F --> G["embedding.py\nGemini embed_documents\n→ float vectors"]
150
- G --> H["retriever.py\nDelete old collection\nCreate fresh ChromaDB\ncollection.add(...)"]
151
- H --> I["βœ“ Vector store ready"]
152
- ```
153
 
154
- ### 3.3 RAG Query Flow
155
 
156
- ```mermaid
157
- flowchart LR
158
- Q["User question"] --> E["embed_query\n(Gemini)"]
159
- E --> S["ChromaDB\ncollection.query\ntop-k chunks"]
160
- S --> C["build_context\nformat chunks\nwith file/line headers"]
161
- C --> P["build_prompt\nqa template\ncontext + question"]
162
- P --> L["LLM\n(Groq / Gemini)"]
163
- L --> A["Answer + Sources\n(file_path, line range)"]
164
  ```
165
 
166
- ### 3.4 CLI Menu Flow
167
-
168
- ```mermaid
169
- flowchart TD
170
- START([Start cli.py]) --> REPO["β–Ά Enter repo path"]
171
- REPO --> INGEST["Ingest Repository\nload β†’ split β†’ embed β†’ store"]
172
- INGEST --> MENU["Show Menu\nOptions 1–9"]
173
-
174
- MENU --> O1["1 Ask a question\n→ RAG Query"]
175
- MENU --> O2["2 Detect bugs\n→ LLM code review"]
176
- MENU --> O3["3 Cyclomatic complexity\n→ Radon"]
177
- MENU --> O4["4 Explain function\n→ AST + LLM"]
178
- MENU --> O5["5 Module docs\n→ LLM"]
179
- MENU --> O6["6 Generate README\n→ LLM"]
180
- MENU --> O7["7 Propose & create file\n→ LLM + local write"]
181
- MENU --> O8["8 List GitHub MCP tools\n→ GitHub Copilot MCP"]
182
- MENU --> O9["9 Re-ingest repo\n→ new repo path"]
183
- MENU --> O0["0 Exit"]
184
-
185
- O1 & O2 & O3 & O4 & O5 & O6 & O7 & O8 & O9 --> MENU
186
- O0 --> END([Goodbye!])
187
- ```
188
 
189
  ---
190
 
191
- ## 4. Component Deep-Dive
192
-
193
- ### 4.1 `config.py` β€” Settings
194
 
195
- All configuration lives in one `pydantic-settings` class loaded from `.env`:
196
 
197
- | Setting | Default | Description |
198
- |---|---|---|
199
- | `llm_provider` | `"groq"` | `"groq"` or `"gemini"` |
200
- | `groq_api_key` | `""` | From `.env` |
201
- | `groq_model` | `"llama-3.3-70b-versatile"` | Free Groq model |
202
- | `gemini_api_key` | `""` | From `.env` |
203
- | `embedding_model` | `"models/gemini-embedding-001"` | Always Gemini for embeddings |
204
- | `vector_db_path` | `"./vector_db"` | ChromaDB storage directory |
205
- | `chunk_size` | `1200` | Characters per chunk |
206
- | `chunk_overlap` | `100` | Overlap between chunks |
207
- | `allowed_extensions` | `[.py .js .ts .go .java .md]` | File types to load |
208
- | `max_file_size_kb` | `1042` | Skip files larger than this |
209
- | `github_mcp_url` | GitHub Copilot MCP endpoint | For option 8 |
210
- | `github_mcp_token` | `""` | GitHub PAT from `.env` |
211
-
212
- ### 4.2 `rag/repository_loader.py` β€” File Ingestion
213
 
214
- ```
215
- load_repository(root_path)
216
- └── os.walk(root_path)
217
- β”œβ”€β”€ Skip: .git, node_modules, __pycache__, venv, .venv, dist, build
218
- β”œβ”€β”€ should_include(fpath) β†’ checks extension + file size
219
- └── read_file_with_metadata(fpath) β†’ CodeDocument(content, file_path, language, size_bytes)
220
- ```
221
 
222
- **Supported languages detected by extension:**
223
 
224
- | Extension | Language |
225
- |---|---|
226
- | `.py` | python |
227
- | `.js` | javascript |
228
- | `.ts` | typescript |
229
- | `.java` | java |
230
- | `.go` | go |
231
- | `.md` | markdown |
232
 
233
- ### 4.3 `rag/splitter.py` β€” Chunking
234
 
235
- Uses **LangChain's `RecursiveCharacterTextSplitter`** with language-aware splitting:
236
- - For Python/JS/TS/Java/Go β€” uses syntax-aware boundaries (functions, classes)
237
- - For Markdown/unknown β€” falls back to generic character splitting
238
- - Each chunk carries: `content`, `file_path`, `language`, `chunk_index`, `start_line`, `end_line`
239
 
240
- ### 4.4 `rag/embedding.py` β€” Embeddings
241
 
242
- - Uses `GoogleGenerativeAIEmbeddings` (`models/gemini-embedding-001`)
243
- - `embed_document(chunks)` β€” batch embeds all chunks, mutates dicts in-place
244
- - `embed_query(query)` β€” single query vector for similarity search
245
 
246
- ### 4.5 `rag/retriever.py` β€” Vector Store
247
 
248
- > **Critical fix applied:** collection is deleted and recreated on every ingest to prevent stale data from previous repos bleeding through.
249
 
250
- ```python
251
- build_vector_store(chunks) # delete β†’ create β†’ add (fresh each ingest)
252
- load_vector_store() # get_or_create for reading
253
- retrieve_relevant_chunks(q, k) # embed query β†’ cosine similarity β†’ top-k
254
  ```
255
 
256
- ### 4.6 `rag/rag_chain.py` β€” Orchestration
257
-
258
- ```python
259
- run_rag_query(query, k=5)
260
- 1. retrieve_relevant_chunks(query, k)
261
- 2. build_context(chunks) # format with file/line headers
262
- 3. build_prompt(query, context) # fill qa template
263
- 4. llm.generate(prompt)
264
- 5. return { answer, sources }
265
- ```
266
 
267
- ### 4.7 `llm.py` β€” LLM Abstraction
268
 
269
- Abstract `BaseLLM` with two concrete providers:
270
 
271
- | Class | Provider | API |
272
- |---|---|---|
273
- | `GeminiLLM` | Google Gemini | `google-genai` SDK |
274
- | `GroqLLM` | Groq | `groq` SDK, chat completions |
275
 
276
- **Prompt templates (`build_prompt`):**
277
 
278
- | `task_type` | Used by |
279
- |---|---|
280
- | `"qa"` | RAG query, function explain, module docs, README |
281
- | `"bug_finding"` | Bug detection (returns JSON) |
282
- | `"docstring"` | Docstring generation |
283
- | `"file_creation"` | AI file proposal |
 
 
284
 
285
- ### 4.8 `mcp_client.py` β€” GitHub MCP Client
286
 
287
- Async context-manager pattern using `AsyncExitStack` to keep all `anyio` cancel scopes in the **same task**:
288
 
289
- ```python
290
- async with get_github_mcp_client() as client:
291
- tools = await client.list_tools()
292
- result = await client.call_tool("create_branch", {...})
293
- ```
294
 
295
- Available GitHub MCP tools (44 total) include: `search_code`, `list_issues`, `create_pull_request`, `get_file_contents`, `push_files`, `create_branch`, `fork_repository`, `search_repositories`, and many more.
296
 
297
- ### 4.9 `services/code_analysis.py`
 
 
 
 
298
 
299
- | Function | How it works |
300
- |---|---|
301
- | `explain_function(file, name)` | Python `ast` extracts the function source β†’ RAG context β†’ LLM |
302
- | `detect_bugs(file)` | Read file β†’ `bug_finding` prompt β†’ LLM returns JSON list |
303
- | `analyze_complexity(file)` | `radon.cc_visit` β†’ cyclomatic complexity + rank A–F |
304
-
305
- ### 4.10 `services/file_creator.py`
306
 
307
- ```
308
- propose_new_file(description, context_query)
309
- β”œβ”€β”€ RAG retrieve relevant context
310
- β”œβ”€β”€ build_prompt(description, context, "file_creation")
311
- β”œβ”€β”€ LLM generates complete file content
312
- └── infer_file_path(description) β†’ LLM suggests relative path
313
-
314
- apply_approved_file(proposal, confirmed, base_dir)
315
- β”œβ”€β”€ Strip leading / or \ from LLM path
316
- β”œβ”€β”€ os.path.join(base_dir, relative_path)
317
- β”œβ”€β”€ os.makedirs(parent_dirs, exist_ok=True)
318
- └── open(abs_path, "w").write(content)
319
- ```
320
 
321
  ---
322
 
323
- ## 5. Setup & Installation
324
 
325
- ### Prerequisites
326
 
327
- | Requirement | Version |
328
  |---|---|
329
- | Python | 3.11+ |
330
- | pip | Latest |
331
- | Internet | For API calls |
 
 
 
332
 
333
- ### Step 1 β€” Clone / Download the project
334
 
335
- ```bash
336
- git clone <your-repo-url>
337
- cd "Codebase Assistant"
338
- ```
339
 
340
- ### Step 2 β€” Create a virtual environment
341
 
342
- ```bash
343
- # Windows (PowerShell)
344
- python -m venv .venv
345
- .venv\Scripts\Activate.ps1
346
 
347
- # macOS / Linux
348
- python3 -m venv .venv
349
- source .venv/bin/activate
350
- ```
351
 
352
- ### Step 3 β€” Install dependencies
 
 
353
 
354
- ```bash
355
- pip install -r requirements.txt
356
- ```
357
 
358
- > [!TIP]
359
- > If you hit permission issues on Windows, use:
360
- > `pip install -r requirements.txt --user`
361
 
362
- ### Step 4 β€” Create your `.env` file
363
 
364
- Create a file named `.env` in the project root:
 
 
365
 
366
- ```env
367
- # Choose your LLM provider: "groq" (free) or "gemini"
368
- LLM_PROVIDER=groq
369
 
370
- # Groq β€” free at https://console.groq.com
371
- GROQ_API_KEY=gsk_your_key_here
372
 
373
- # Gemini β€” get at https://aistudio.google.com
374
- GEMINI_API_KEY=your_gemini_key_here
375
 
376
- # GitHub PAT for MCP tools β€” create at https://github.com/settings/tokens
377
- # Required scopes: repo, read:org
378
- GITHUB_MCP_TOKEN=github_pat_your_token_here
379
- ```
380
 
381
- > [!IMPORTANT]
382
- > `GEMINI_API_KEY` is **always required** regardless of LLM provider, because embeddings always use Gemini.
 
 
 
383
 
384
- ### Step 5 β€” Run the CLI
385
 
386
- ```bash
387
- .venv\Scripts\python.exe cli.py # Windows
388
- python cli.py # macOS / Linux
389
- ```
390
 
391
  ---
392
 
393
- ## 6. Configuration (.env)
394
-
395
- ```env
396
- # ─── LLM Provider ────────────────────────────────────────
397
- LLM_PROVIDER=groq # "groq" | "gemini"
398
-
399
- # ─── Groq (recommended β€” free tier) ─────────────────────
400
- GROQ_API_KEY=gsk_...
401
- GROQ_MODEL=llama-3.3-70b-versatile
402
-
403
- # ─── Gemini ──────────────────────────────────────────────
404
- GEMINI_API_KEY=...
405
- LLM_MODEL=gemini-1.5-flash # only used if LLM_PROVIDER=gemini
406
- EMBEDDING_MODEL=models/gemini-embedding-001
407
-
408
- # ─── RAG / Vector DB ─────────────────────────────────────
409
- VECTOR_DB_PATH=./vector_db
410
- CHUNK_SIZE=1200
411
- CHUNK_OVERLAP=100
412
 
413
- # ─── File loader ─────────────────────────────────────────
414
- # comma-separated extensions
415
- ALLOWED_EXTENSIONS=[".py",".js",".ts",".go",".java",".md"]
416
- MAX_FILE_SIZE_KB=1042
417
 
418
- # ─── GitHub MCP ──────────────────────────────────────────
419
- GITHUB_MCP_URL=https://api.githubcopilot.com/mcp/
420
- GITHUB_MCP_TOKEN=github_pat_...
421
  ```
422
 
423
- ---
424
 
425
- ## 7. Running the CLI
426
 
 
 
 
427
  ```
428
- ╔══════════════════════════════════════════╗
429
- β•‘ AI Codebase Assistant CLI β•‘
430
- β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
431
-
432
- β–Ά Enter the path to the repository you want to analyse: D:\MyProject
433
-
434
- ── Ingesting Repository ──────────────────────
435
- Loading files from: D:\MyProject
436
- βœ“ Loaded 12 files
437
- βœ“ Split into 47 chunks
438
- Embedding chunks (this may take a moment)...
439
- βœ“ Embedded 47 chunks
440
- βœ“ Vector store ready (ChromaDB)
441
-
442
- ── What would you like to do? ────────────────
443
- 1 Ask a question about the codebase
444
- 2 Detect bugs in a file
445
- 3 Cyclomatic complexity analysis
446
- 4 Explain a function
447
- 5 Generate module documentation
448
- 6 Generate README for a repository
449
- 7 Propose & create a new file (AI)
450
- 8 List GitHub MCP tools
451
- 9 Re-ingest a repository
452
- 0 Exit
453
-
454
- β–Ά Choose an option [0–9]:
455
- ```
456
-
457
- ---
458
 
459
- ## 8. CLI Features β€” All 9 Options
460
-
461
- ### Option 1 β€” Ask a Question (RAG Query)
462
 
 
 
 
463
  ```
464
- β–Ά Your question: How does the authentication work?
465
- β–Ά Number of source chunks to retrieve? [default: 5] 3
466
 
467
- ── Answer ────────────────────────────────────
468
- The authentication uses JWT tokens ...
469
 
470
- ── Sources ───────────────────────────────────
471
- β€’ src/auth/middleware.py lines 12–45
472
- β€’ src/auth/tokens.py lines 1–30
473
  ```
474
 
475
- **Pipeline:** Embed question β†’ ChromaDB similarity search β†’ Build context β†’ LLM β†’ Answer + cited sources
476
-
477
  ---
 
478
 
479
- ### Option 2 β€” Detect Bugs
480
 
481
- ```
482
- β–Ά File path to analyse: src/payment.py
483
 
484
- ── Found 2 issue(s) ──────────────────────────
485
- [HIGH] Line 34: SQL query uses string concatenation
486
- β†’ Use parameterized queries to prevent SQL injection
487
 
488
- [MEDIUM] Line 67: Exception swallowed silently
489
- β†’ Log or re-raise the exception
490
- ```
491
 
492
- **Pipeline:** Read file β†’ `bug_finding` prompt β†’ LLM returns JSON β†’ parsed and displayed
493
 
494
- ---
495
-
496
- ### Option 3 β€” Cyclomatic Complexity
497
-
498
- ```
499
- β–Ά File path to analyse: src/processor.py
500
-
501
- ── 4 function(s) ─────────────────────────────
502
- [A] process_order complexity=2 β–ˆβ–ˆ
503
- [B] validate_cart complexity=5 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
504
- [C] apply_discounts complexity=8 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
505
- [F] handle_edge_cases complexity=18 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
506
- ```
507
-
508
- **Rank scale:** A (1–5, simple) β†’ F (26+, untestable)
509
 
510
- ---
 
511
 
512
- ### Option 4 β€” Explain a Function
 
513
 
514
- ```
515
- β–Ά File path: src/utils.py
516
- β–Ά Function name: parse_date_range
517
 
518
- ── Explanation ───────────────────────────────
519
- parse_date_range takes a string like "2024-01-01:2024-12-31"
520
- and returns a tuple of (start_date, end_date) as datetime objects ...
521
  ```
522
 
523
- **Pipeline:** Python `ast` module extracts exact function source β†’ RAG retrieves usages β†’ LLM explains
524
 
525
- ---
526
 
527
- ### Option 5 β€” Generate Module Documentation
528
 
 
 
529
  ```
530
- β–Ά File path: src/database.py
531
 
532
- ── Documentation ─────────────────────────────
533
- ## database.py
534
 
535
- ### Overview
536
- This module provides the database connection layer ...
537
 
538
- ### Functions
539
- - `connect(url)` β€” Establishes a connection ...
540
- - `execute(query, params)` β€” Runs a parameterized query ...
541
- ```
542
-
543
- ---
544
 
545
- ### Option 6 β€” Generate README
546
 
547
- ```
548
- β–Ά Repository root path: D:\MyProject
549
 
550
- ── README Preview ────────────────────────────
551
- # MyProject
552
 
553
- ## Overview
554
- A FastAPI application that ...
555
 
556
- β–Ά Save to README.md in that directory? [y/N] y
557
- βœ“ Saved to D:\MyProject\README.md
558
  ```
559
 
560
- ---
561
-
562
- ### Option 7 β€” Propose & Create a New File
563
 
 
 
564
  ```
565
- β–Ά Describe the file you want to create: Rectangle area calculator in JavaScript
566
- β–Ά Optional context query (or press Enter to skip):
567
 
568
- ── Proposed File ─────────────────────────────
569
- Path: /src/geometry/rectangle.js
570
 
571
- class Rectangle { ...full generated code... }
572
-
573
- β–Ά Write this file to disk? [y/N] y
574
- β–Ά Save under which directory? [default: D:\MyProject]
575
-
576
- βœ“ File written: D:\MyProject\src\geometry\rectangle.js
 
 
 
 
 
577
  ```
578
 
579
- **Pipeline:** RAG context (optional) β†’ `file_creation` prompt β†’ LLM generates code β†’ user confirms β†’ write to `base_dir/relative_path`
580
-
581
  ---
582
 
583
- ### Option 8 β€” List GitHub MCP Tools
584
 
585
- ```
586
- ── List GitHub MCP Tools ─────────────────────
587
- 44 tools available:
588
- β€’ create_branch Create a new branch in a GitHub repository
589
- β€’ create_pull_request Create a new pull request ...
590
- β€’ search_code Fast and precise code search ...
591
- β€’ list_issues List issues in a GitHub repository ...
592
- ...
593
- ```
594
-
595
- **Connection:** Uses `streamable_http_client` β†’ `ClientSession` β†’ GitHub Copilot MCP endpoint
596
 
597
- ---
598
 
599
- ### Option 9 β€” Re-ingest a Repository
600
-
601
- ```
602
- β–Ά New repository path to ingest: D:\AnotherProject
603
- ── Ingesting Repository ──────────────────────
604
- βœ“ Loaded 8 files ...
605
  ```
606
 
607
- Wipes the ChromaDB collection and ingests the new repo fresh. All subsequent queries answer from the new repo only.
608
-
609
- ---
610
-
611
- ## 9. REST API
612
-
613
- Start the FastAPI server:
614
 
615
- ```bash
616
- .venv\Scripts\python.exe -m uvicorn app:app --reload --port 8000
617
  ```
618
 
619
- Open docs at: `http://localhost:8000/docs`
620
-
621
- ### Endpoints
622
 
623
- | Method | Path | Description |
624
  |---|---|---|
625
  | `GET` | `/health` | Health check |
626
- | `POST` | `/api/query` | RAG question answering |
627
- | `POST` | `/api/analyze/bugs?file_path=...` | Bug detection |
628
- | `POST` | `/api/analyze/complexity?file_path=...` | Cyclomatic complexity |
629
  | `POST` | `/api/analyze/explain` | Explain a function |
630
- | `POST` | `/api/docs/module?file_path=...` | Module documentation |
631
- | `POST` | `/api/docs/readme?root_path=...` | README generation |
632
- | `POST` | `/api/files/propose` | Propose a new file |
633
- | `POST` | `/api/files/approve` | Write approved file to disk |
634
 
635
- ### Example β€” Query
 
 
636
 
637
  ```bash
638
  curl -X POST http://localhost:8000/api/query \
@@ -640,107 +339,124 @@ curl -X POST http://localhost:8000/api/query \
640
  -d '{"query": "How does authentication work?", "k": 5}'
641
  ```
642
 
 
 
643
  ```json
644
  {
645
  "answer": "Authentication is handled by ...",
646
  "sources": [
647
- {"file_path": "src/auth.py", "start_line": 10, "end_line": 45}
 
 
 
 
648
  ]
649
  }
650
  ```
651
 
652
  ---
653
 
654
- ## 10. Key Design Decisions & Bug Fixes
655
-
656
- ### Bug Fix 1 β€” Collection Isolation (retriever.py)
657
-
658
- **Problem:** ChromaDB used `upsert` on a shared `"codebase"` collection β€” stale chunks from previously ingested repos leaked into new queries.
659
-
660
- **Fix:** On every ingest, `delete_collection` + `create_collection` ensures a clean slate:
661
-
662
- ```python
663
- # Before (bug)
664
- collection = client.get_or_create_collection("codebase")
665
- collection.upsert(...)
666
 
667
- # After (fix)
668
- client.delete_collection("codebase") # wipe old repo
669
- collection = client.create_collection("codebase")
670
- collection.add(...)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
671
  ```
672
 
673
  ---
674
 
675
- ### Bug Fix 2 β€” GitHub MCP Cancel Scope (mcp_client.py)
676
-
677
- **Problem:** Manually calling `__aenter__`/`__aexit__` on `anyio`-backed context managers across tasks causes `RuntimeError: Attempted to exit cancel scope in a different task`.
678
 
679
- **Fix:** Use `AsyncExitStack` to nest all context managers inside a single `async with`:
680
-
681
- ```python
682
- async with get_github_mcp_client() as client:
683
- tools = await client.list_tools()
684
- ```
 
 
 
 
685
 
686
  ---
687
 
688
- ### Bug Fix 3 β€” Wrong Unpack Count (mcp_client.py)
689
-
690
- **Problem:** `streamable_http_client` yields 2 values, not 3. Unpacking 3 caused `ValueError: not enough values to unpack`.
691
 
692
- ```python
693
- # Before (bug)
694
- read, write, _ = await ctx.__aenter__()
695
-
696
- # After (fix)
697
- read, write = await stack.enter_async_context(streamable_http_client(...))
698
- ```
699
 
700
  ---
701
 
702
- ### Bug Fix 4 β€” File Save Path (file_creator.py)
703
 
704
- **Problem:** Generated files saved to the current working directory regardless of user input.
705
 
706
- **Fix:**
707
- - Strip leading `/\` from LLM path to make it always relative
708
- - Accept `base_dir` defaulting to `_current_repo` (the ingested repo path)
709
- - Create all parent directories automatically
 
 
 
 
 
 
 
710
 
711
  ---
712
 
713
- ## 11. Extending the Project
714
 
715
- ### Add a new LLM provider
716
 
717
- 1. Add a new class in `llm.py` extending `BaseLLM`
718
- 2. Add the provider name to `get_llm_client()` factory
719
- 3. Add matching settings in `config.py`
720
 
721
- ### Add support for new file types
722
 
723
- 1. Add extension β†’ language in `LANGUAGE_BY_EXT` in `repository_loader.py`
724
- 2. Add extension to `allowed_extensions` in `config.py`
725
- 3. If LangChain has a `Language` enum for it, add to `LANGUAGE_MAP` in `splitter.py`
726
 
727
- ### Use a GitHub MCP tool in a feature
728
 
729
- ```python
730
- async with get_github_mcp_client() as client:
731
- result = await client.call_tool("create_branch", {
732
- "owner": "myuser",
733
- "repo": "myrepo",
734
- "branch": "feature/new-branch",
735
- "from_branch": "main"
736
- })
737
- ```
738
 
739
- ### Add a new CLI option
740
 
741
- 1. Write a `feature_xxx()` function in `cli.py`
742
- 2. Add `("Label", feature_xxx)` to the `MENU` list
743
- 3. Add a matching FastAPI endpoint in `api/routes.py`
744
 
745
- ---
 
 
746
 
 
 
 
 
 
 
 
1
+ # AI Codebase Assistant
2
 
3
+ An AI-powered developer tool that lets you interact with an entire code repository using **RAG (Retrieval-Augmented Generation)** and LLMs.
4
+
5
+ It can answer questions about your codebase, detect potential bugs, analyze code complexity, explain functions, generate documentation, create new files, and interact with GitHub through the GitHub MCP API.
6
 
7
  ---
8
 
9
+ ## Features
10
 
11
+ ### Codebase Question Answering
 
 
 
 
 
 
 
 
 
 
12
 
13
+ Ask natural-language questions about your repository.
14
 
15
+ **Example:**
16
 
17
+ ```text
18
+ How does authentication work?
19
+ ```
20
 
21
+ The assistant searches the relevant parts of the codebase and generates an answer with source references.
 
 
 
 
 
 
 
22
 
23
+ ---
 
 
 
 
24
 
25
+ ### Bug Detection
 
26
 
27
+ Analyze a source file and identify potential issues using an LLM-powered code review.
28
 
29
+ Example:
30
 
31
+ ```text
32
+ [HIGH] Line 34
33
+ SQL query uses string concatenation
34
+
35
+ Recommendation:
36
+ Use parameterized queries to prevent SQL injection.
37
  ```
38
+
39
+ The system returns the issue severity, location, and recommendation.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40
 
41
  ---
42
 
43
+ ### Cyclomatic Complexity Analysis
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
+ Analyze the complexity of Python functions using **Radon**.
 
 
 
 
 
 
 
 
 
 
 
 
46
 
47
+ Example:
48
 
49
+ ```text
50
+ [A] process_order complexity = 2
51
+ [B] validate_cart complexity = 5
52
+ [C] apply_discounts complexity = 8
53
+ [F] handle_edge_cases complexity = 18
 
 
 
54
  ```
55
 
56
+ This helps identify functions that may be difficult to maintain or test.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
57
 
58
  ---
59
 
60
+ ### Function Explanation
 
 
61
 
62
+ Select a function and get a plain-English explanation of what it does.
63
 
64
+ For Python code, the project uses the `ast` module to extract the function and provides relevant repository context to the LLM.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65
 
66
+ ---
 
 
 
 
 
 
67
 
68
+ ### Documentation Generation
69
 
70
+ Generate documentation automatically for your codebase.
 
 
 
 
 
 
 
71
 
72
+ The assistant can generate:
73
 
74
+ - Module documentation
75
+ - Function explanations
76
+ - Docstrings
77
+ - Repository README files
78
 
79
+ ---
80
 
81
+ ### AI File Generation
 
 
82
 
83
+ Describe the file you want to create and let the AI generate it.
84
 
85
+ Example:
86
 
87
+ ```text
88
+ Create a Rectangle class in JavaScript
89
+ that calculates area and perimeter.
 
90
  ```
91
 
92
+ The generated file is shown before it is written to the repository, allowing the user to approve it first.
 
 
 
 
 
 
 
 
 
93
 
94
+ ---
95
 
96
+ ### GitHub MCP Integration
97
 
98
+ The project integrates with the **GitHub Copilot MCP API**.
 
 
 
99
 
100
+ It currently supports access to **44 GitHub MCP tools**, including:
101
 
102
+ - `search_code`
103
+ - `search_repositories`
104
+ - `get_file_contents`
105
+ - `list_issues`
106
+ - `create_branch`
107
+ - `create_pull_request`
108
+ - `push_files`
109
+ - `fork_repository`
110
 
111
+ ---
112
 
113
+ ## RAG-Based Code Search
114
 
115
+ The project uses **Retrieval-Augmented Generation (RAG)** to work with repository-level code.
 
 
 
 
116
 
117
+ Repository files are:
118
 
119
+ 1. Loaded and filtered
120
+ 2. Split into smaller chunks
121
+ 3. Converted into embeddings
122
+ 4. Stored in ChromaDB
123
+ 5. Retrieved based on semantic similarity when a question is asked
124
 
125
+ The retrieved code is then provided to the LLM as context.
 
 
 
 
 
 
126
 
127
+ Source metadata such as file path and line range is preserved during this process.
 
 
 
 
 
 
 
 
 
 
 
 
128
 
129
  ---
130
 
131
+ ## Supported Languages
132
 
133
+ The repository ingestion system currently supports:
134
 
135
+ | Extension | Language |
136
  |---|---|
137
+ | `.py` | Python |
138
+ | `.js` | JavaScript |
139
+ | `.ts` | TypeScript |
140
+ | `.java` | Java |
141
+ | `.go` | Go |
142
+ | `.md` | Markdown |
143
 
144
+ ---
145
 
146
+ ## Tech Stack
 
 
 
147
 
148
+ ### AI / LLM
149
 
150
+ - Groq
151
+ - Llama 3.3 70B
152
+ - Google Gemini
 
153
 
154
+ ### RAG
 
 
 
155
 
156
+ - LangChain
157
+ - Google Gemini Embeddings
158
+ - ChromaDB
159
 
160
+ ### Backend
 
 
161
 
162
+ - Python
163
+ - FastAPI
164
+ - Pydantic Settings
165
 
166
+ ### Code Analysis
167
 
168
+ - Python AST
169
+ - Radon
170
+ - LLM-based code analysis
171
 
172
+ ### Integration
 
 
173
 
174
+ - GitHub MCP
175
+ - GitHub Copilot MCP API
176
 
177
+ ---
 
178
 
179
+ ## LLM Providers
 
 
 
180
 
181
+ | Provider | Model | Usage |
182
+ |---|---|---|
183
+ | Groq | `llama-3.3-70b-versatile` | LLM generation |
184
+ | Gemini | Configurable | LLM generation |
185
+ | Gemini Embeddings | `models/gemini-embedding-001` | Code embeddings |
186
 
187
+ The project supports switching between Groq and Gemini for LLM generation.
188
 
189
+ Gemini Embeddings are used for semantic retrieval.
 
 
 
190
 
191
  ---
192
 
193
+ # Installation
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
194
 
195
+ ## 1. Clone the Repository
 
 
 
196
 
197
+ ```bash
198
+ git clone <your-repository-url>
199
+ cd "Codebase Assistant"
200
  ```
201
 
202
+ ## 2. Create a Virtual Environment
203
 
204
+ ### Windows
205
 
206
+ ```bash
207
+ python -m venv .venv
208
+ .venv\Scripts\Activate.ps1
209
  ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
210
 
211
+ ### Linux / macOS
 
 
212
 
213
+ ```bash
214
+ python3 -m venv .venv
215
+ source .venv/bin/activate
216
  ```
 
 
217
 
218
+ ## 3. Install Dependencies
 
219
 
220
+ ```bash
221
+ pip install -r requirements.txt
 
222
  ```
223
 
 
 
224
  ---
225
+ ## πŸ—οΈ Architecture
226
 
227
+ ![AI Codebase Assistant Architecture](./architecture.png)
228
 
229
+ The system combines a RAG pipeline, LLM providers, code analysis services,
230
+ FastAPI, and GitHub MCP integration to provide repository-level AI assistance.
231
 
 
 
 
232
 
233
+ # Configuration
 
 
234
 
235
+ Create a `.env` file in the project root.
236
 
237
+ ```env
238
+ LLM_PROVIDER=groq
 
 
 
 
 
 
 
 
 
 
 
 
 
239
 
240
+ GROQ_API_KEY=your_groq_api_key
241
+ GROQ_MODEL=llama-3.3-70b-versatile
242
 
243
+ GEMINI_API_KEY=your_gemini_api_key
244
+ EMBEDDING_MODEL=models/gemini-embedding-001
245
 
246
+ VECTOR_DB_PATH=./vector_db
 
 
247
 
248
+ GITHUB_MCP_TOKEN=your_github_token
 
 
249
  ```
250
 
251
+ ### Required API Keys
252
 
253
+ **Groq**
254
 
255
+ Used for LLM generation when:
256
 
257
+ ```env
258
+ LLM_PROVIDER=groq
259
  ```
 
260
 
261
+ **Gemini**
 
262
 
263
+ Required for embeddings because the project uses Gemini Embeddings for repository indexing.
 
264
 
265
+ **GitHub Token**
 
 
 
 
 
266
 
267
+ Required for GitHub MCP functionality.
268
 
269
+ ---
 
270
 
271
+ # Running the CLI
 
272
 
273
+ Start the interactive CLI:
 
274
 
275
+ ```bash
276
+ python cli.py
277
  ```
278
 
279
+ The application will ask for the repository you want to analyze.
 
 
280
 
281
+ ```text
282
+ β–Ά Enter the path to the repository you want to analyse:
283
  ```
 
 
284
 
285
+ After the repository is indexed, you can choose from:
 
286
 
287
+ ```text
288
+ 1 Ask a question about the codebase
289
+ 2 Detect bugs in a file
290
+ 3 Cyclomatic complexity analysis
291
+ 4 Explain a function
292
+ 5 Generate module documentation
293
+ 6 Generate README
294
+ 7 Propose & create a new file
295
+ 8 List GitHub MCP tools
296
+ 9 Re-ingest repository
297
+ 0 Exit
298
  ```
299
 
 
 
300
  ---
301
 
302
+ # REST API
303
 
304
+ The project also provides a FastAPI REST API.
 
 
 
 
 
 
 
 
 
 
305
 
306
+ Start the server:
307
 
308
+ ```bash
309
+ python -m uvicorn app:app --reload --port 8000
 
 
 
 
310
  ```
311
 
312
+ Open the interactive API documentation:
 
 
 
 
 
 
313
 
314
+ ```text
315
+ http://localhost:8000/docs
316
  ```
317
 
318
+ ## API Endpoints
 
 
319
 
320
+ | Method | Endpoint | Description |
321
  |---|---|---|
322
  | `GET` | `/health` | Health check |
323
+ | `POST` | `/api/query` | Ask questions about the codebase |
324
+ | `POST` | `/api/analyze/bugs` | Detect potential bugs |
325
+ | `POST` | `/api/analyze/complexity` | Analyze cyclomatic complexity |
326
  | `POST` | `/api/analyze/explain` | Explain a function |
327
+ | `POST` | `/api/docs/module` | Generate module documentation |
328
+ | `POST` | `/api/docs/readme` | Generate README |
329
+ | `POST` | `/api/files/propose` | Generate a file proposal |
330
+ | `POST` | `/api/files/approve` | Write an approved file |
331
 
332
+ ---
333
+
334
+ ## Example API Request
335
 
336
  ```bash
337
  curl -X POST http://localhost:8000/api/query \
 
339
  -d '{"query": "How does authentication work?", "k": 5}'
340
  ```
341
 
342
+ Example response:
343
+
344
  ```json
345
  {
346
  "answer": "Authentication is handled by ...",
347
  "sources": [
348
+ {
349
+ "file_path": "src/auth.py",
350
+ "start_line": 10,
351
+ "end_line": 45
352
+ }
353
  ]
354
  }
355
  ```
356
 
357
  ---
358
 
359
+ # πŸ“ Project Structure
 
 
 
 
 
 
 
 
 
 
 
360
 
361
+ ```text
362
+ Codebase Assistant/
363
+ β”‚
364
+ β”œβ”€β”€ cli.py
365
+ β”œβ”€β”€ app.py
366
+ β”œβ”€β”€ config.py
367
+ β”œβ”€β”€ llm.py
368
+ β”œβ”€β”€ mcp_client.py
369
+ β”œβ”€β”€ requirements.txt
370
+ β”‚
371
+ β”œβ”€β”€ rag/
372
+ β”‚ β”œβ”€β”€ repository_loader.py
373
+ β”‚ β”œβ”€β”€ splitter.py
374
+ β”‚ β”œβ”€β”€ embedding.py
375
+ β”‚ β”œβ”€β”€ retriever.py
376
+ β”‚ └── rag_chain.py
377
+ β”‚
378
+ β”œβ”€β”€ services/
379
+ β”‚ β”œβ”€β”€ code_analysis.py
380
+ β”‚ β”œβ”€β”€ documentation.py
381
+ β”‚ └── file_creator.py
382
+ β”‚
383
+ β”œβ”€β”€ api/
384
+ β”‚ └── routes.py
385
+ β”‚
386
+ └── vector_db/
387
+ └── chroma.sqlite3
388
  ```
389
 
390
  ---
391
 
392
+ # Configuration Options
 
 
393
 
394
+ | Setting | Default | Description |
395
+ |---|---|---|
396
+ | `LLM_PROVIDER` | `groq` | LLM provider |
397
+ | `GROQ_MODEL` | `llama-3.3-70b-versatile` | Groq model |
398
+ | `EMBEDDING_MODEL` | `models/gemini-embedding-001` | Embedding model |
399
+ | `VECTOR_DB_PATH` | `./vector_db` | ChromaDB storage |
400
+ | `CHUNK_SIZE` | `1200` | Chunk size |
401
+ | `CHUNK_OVERLAP` | `100` | Chunk overlap |
402
+ | `MAX_FILE_SIZE_KB` | `1042` | Maximum file size |
403
+ | `ALLOWED_EXTENSIONS` | `.py,.js,.ts,.go,.java,.md` | Supported files |
404
 
405
  ---
406
 
407
+ # Limitations
 
 
408
 
409
+ - Repository ingestion currently runs locally.
410
+ - External API keys are required for LLM and embedding services.
411
+ - Gemini embedding quotas may limit large repositories.
412
+ - Only the currently supported file types are indexed.
413
+ - LLM-generated code and bug reports should be reviewed before use.
414
+ - The vector database currently focuses on the actively ingested repository.
 
415
 
416
  ---
417
 
418
+ # Future Improvements
419
 
420
+ Planned or possible improvements include:
421
 
422
+ - More programming language support
423
+ - Incremental repository indexing
424
+ - GitHub repository ingestion
425
+ - AST-based code chunking
426
+ - Hybrid keyword + vector search
427
+ - Retrieval reranking
428
+ - Dependency graph analysis
429
+ - Automated test generation
430
+ - Automated pull request review
431
+ - Multi-repository search
432
+ - Web-based developer interface
433
 
434
  ---
435
 
436
+ # Project Goal
437
 
438
+ The goal of this project is to build an AI developer assistant that can understand and interact with an entire codebase rather than only individual code snippets.
439
 
440
+ It combines:
 
 
441
 
442
+ **RAG + LLMs + Vector Search + Code Analysis + FastAPI + MCP**
443
 
444
+ into a single developer-focused tool.
 
 
445
 
446
+ ---
447
 
 
 
 
 
 
 
 
 
 
448
 
449
+ # Author
450
 
451
+ **Armaan Alam**
 
 
452
 
453
+ AI Engineer & Software Developer
454
+
455
+ Interested in:
456
 
457
+ - Generative AI
458
+ - RAG Systems
459
+ - Backend Engineering
460
+ - LLM Applications
461
+ - AI Developer Tools
462
+ - Machine Learning