armaanalam commited on
Commit
e9b3659
Β·
verified Β·
1 Parent(s): 43f8812

Upload 10 files

Browse files
Files changed (10) hide show
  1. Readme.md +746 -0
  2. app.py +34 -0
  3. cli.py +323 -0
  4. config.py +41 -0
  5. example.py +21 -0
  6. llm.py +130 -0
  7. mcp_client.py +68 -0
  8. requirements.txt +38 -0
  9. test_github_mcp.py +18 -0
  10. test_runner.py +164 -0
Readme.md ADDED
@@ -0,0 +1,746 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # πŸ€– AI Codebase Assistant β€” Complete Documentation
2
+
3
+ > An AI-powered developer tool that ingests any local code repository, answers questions about it using RAG (Retrieval-Augmented Generation), detects bugs, measures complexity, generates docs, proposes new files, and integrates with the GitHub MCP API β€” all from an interactive CLI.
4
+
5
+ ---
6
+
7
+ ## Table of Contents
8
+
9
+ 1. [Project Overview](#1-project-overview)
10
+ 2. [Repository Structure](#2-repository-structure)
11
+ 3. [Architecture & Flow](#3-architecture--flow)
12
+ 4. [Component Deep-Dive](#4-component-deep-dive)
13
+ 5. [Setup & Installation](#5-setup--installation)
14
+ 6. [Configuration (.env)](#6-configuration-env)
15
+ 7. [Running the CLI](#7-running-the-cli)
16
+ 8. [CLI Features β€” All 9 Options](#8-cli-features--all-9-options)
17
+ 9. [REST API](#9-rest-api)
18
+ 10. [Key Design Decisions & Bug Fixes](#10-key-design-decisions--bug-fixes)
19
+ 11. [Extending the Project](#11-extending-the-project)
20
+
21
+ ---
22
+
23
+ ## 1. Project Overview
24
+
25
+ The **AI Codebase Assistant** is a local developer tool that:
26
+
27
+ - πŸ“‚ **Ingests** any code repository (Python, JS, TS, Java, Go, Markdown)
28
+ - πŸ” **Answers questions** about the code using RAG + LLM
29
+ - πŸ› **Detects bugs** via LLM-powered code review
30
+ - πŸ“Š **Measures cyclomatic complexity** using Radon
31
+ - πŸ“ **Explains functions** in plain English
32
+ - πŸ“„ **Generates module docs and READMEs** automatically
33
+ - πŸ› οΈ **Proposes & creates new files** using AI, saved to your chosen directory
34
+ - πŸ”— **Lists 44 GitHub MCP tools** via the GitHub Copilot MCP API
35
+
36
+ **LLM Providers supported:**
37
+ | Provider | Model | Notes |
38
+ |---|---|---|
39
+ | Groq | `llama-3.3-70b-versatile` | Free tier, recommended |
40
+ | Gemini | Configurable | Paid, stricter quota |
41
+
42
+ **Embedding:** Google Gemini Embedding API (`models/gemini-embedding-001`)
43
+ **Vector Store:** ChromaDB (local persistent)
44
+
45
+ ---
46
+
47
+ ## 2. Repository Structure
48
+
49
+ ```
50
+ Codebase Assistant/
51
+ β”‚
52
+ β”œβ”€β”€ cli.py # ← Main entry point (interactive CLI)
53
+ β”œβ”€β”€ app.py # ← FastAPI REST server (optional)
54
+ β”œβ”€β”€ config.py # ← All settings via pydantic-settings + .env
55
+ β”œβ”€β”€ llm.py # ← LLM abstraction (Gemini / Groq) + prompt templates
56
+ β”œβ”€β”€ mcp_client.py # ← GitHub MCP async context-manager client
57
+ β”œβ”€β”€ requirements.txt # ← All Python dependencies
58
+ β”œβ”€β”€ .env # ← API keys and configuration (not committed)
59
+ β”‚
60
+ β”œβ”€β”€ rag/ # ── RAG Pipeline ──────────────────────────────
61
+ β”‚ β”œβ”€β”€ repository_loader.py # Walk directory, read files β†’ CodeDocument
62
+ β”‚ β”œβ”€β”€ splitter.py # Split CodeDocuments into chunks with metadata
63
+ β”‚ β”œβ”€β”€ embedding.py # Embed chunks/queries via Gemini Embeddings
64
+ β”‚ β”œβ”€β”€ retriever.py # ChromaDB vector store: build / load / query
65
+ β”‚ └── rag_chain.py # Orchestrate retrieve β†’ build context β†’ LLM
66
+ β”‚
67
+ β”œβ”€β”€ services/ # ── Feature Services ──────────────────────────
68
+ β”‚ β”œβ”€β”€ code_analysis.py # Bug detection, complexity, function explainer
69
+ β”‚ β”œβ”€β”€ documentation.py # Module docs, README generator
70
+ β”‚ └── file_creator.py # AI file proposal + local disk write
71
+ β”‚
72
+ β”œβ”€β”€ api/ # ── REST API (FastAPI) ────────────────────────
73
+ β”‚ └── routes.py # All HTTP endpoints (mirrors CLI features)
74
+ β”‚
75
+ └── vector_db/ # ── ChromaDB Storage (auto-created) ───────────
76
+ └── chroma.sqlite3 # Persisted vector embeddings
77
+ ```
78
+
79
+ ---
80
+
81
+ ## 3. Architecture & Flow
82
+
83
+ ### 3.1 Full System Architecture
84
+
85
+ ```mermaid
86
+ graph TB
87
+ subgraph USER["User Interface"]
88
+ CLI["cli.py\n(Interactive CLI)"]
89
+ API["app.py\n(FastAPI REST API)"]
90
+ end
91
+
92
+ subgraph RAG["RAG Pipeline"]
93
+ RL["repository_loader.py\nWalk & read files"]
94
+ SP["splitter.py\nChunk by language"]
95
+ EM["embedding.py\nGemini Embeddings"]
96
+ RT["retriever.py\nChromaDB"]
97
+ RC["rag_chain.py\nOrchestrator"]
98
+ end
99
+
100
+ subgraph SERVICES["Feature Services"]
101
+ CA["code_analysis.py\nBugs Β· Complexity Β· Explain"]
102
+ DOC["documentation.py\nModule Docs Β· README"]
103
+ FC["file_creator.py\nAI File Creator"]
104
+ end
105
+
106
+ subgraph EXTERNAL["External APIs"]
107
+ GROQ["Groq API\nllama-3.3-70b"]
108
+ GEM["Gemini API\nEmbeddings"]
109
+ MCP["GitHub MCP\n44 tools"]
110
+ end
111
+
112
+ LLM["llm.py\nLLM Abstraction Layer"]
113
+ CFG["config.py\n.env Settings"]
114
+ DB[("vector_db/\nChromaDB")]
115
+
116
+ CLI --> RAG
117
+ CLI --> SERVICES
118
+ API --> RAG
119
+ API --> SERVICES
120
+
121
+ RL --> SP --> EM --> RT
122
+ RT --> DB
123
+ RC --> RT
124
+ RC --> LLM
125
+
126
+ CA --> LLM
127
+ DOC --> LLM
128
+ FC --> LLM
129
+
130
+ LLM --> GROQ
131
+ LLM --> GEM
132
+ EM --> GEM
133
+ CLI --> MCP
134
+
135
+ CFG -.->|settings| RAG
136
+ CFG -.->|settings| LLM
137
+ CFG -.->|settings| SERVICES
138
+ ```
139
+
140
+ ### 3.2 Ingest Flow (Step-by-step)
141
+
142
+ ```mermaid
143
+ flowchart LR
144
+ A["User provides\nrepo path"] --> B["repository_loader.py\nwalk directories\nskip: .git venv __pycache__"]
145
+ B --> C["Filter by\nallowed extensions\n.py .js .ts .java .go .md"]
146
+ C --> D["Read each file\n→ CodeDocument\n(content, path, language, size)"]
147
+ D --> E["splitter.py\nLanguage-aware chunking\nRecursiveCharacterTextSplitter"]
148
+ E --> F["Each chunk gets\nmetadata: file_path\nlanguage Β· start_line Β· end_line"]
149
+ F --> G["embedding.py\nGemini embed_documents\n→ float vectors"]
150
+ G --> H["retriever.py\nDelete old collection\nCreate fresh ChromaDB\ncollection.add(...)"]
151
+ H --> I["βœ“ Vector store ready"]
152
+ ```
153
+
154
+ ### 3.3 RAG Query Flow
155
+
156
+ ```mermaid
157
+ flowchart LR
158
+ Q["User question"] --> E["embed_query\n(Gemini)"]
159
+ E --> S["ChromaDB\ncollection.query\ntop-k chunks"]
160
+ S --> C["build_context\nformat chunks\nwith file/line headers"]
161
+ C --> P["build_prompt\nqa template\ncontext + question"]
162
+ P --> L["LLM\n(Groq / Gemini)"]
163
+ L --> A["Answer + Sources\n(file_path, line range)"]
164
+ ```
165
+
166
+ ### 3.4 CLI Menu Flow
167
+
168
+ ```mermaid
169
+ flowchart TD
170
+ START([Start cli.py]) --> REPO["β–Ά Enter repo path"]
171
+ REPO --> INGEST["Ingest Repository\nload β†’ split β†’ embed β†’ store"]
172
+ INGEST --> MENU["Show Menu\nOptions 1–9"]
173
+
174
+ MENU --> O1["1 Ask a question\n→ RAG Query"]
175
+ MENU --> O2["2 Detect bugs\n→ LLM code review"]
176
+ MENU --> O3["3 Cyclomatic complexity\n→ Radon"]
177
+ MENU --> O4["4 Explain function\n→ AST + LLM"]
178
+ MENU --> O5["5 Module docs\n→ LLM"]
179
+ MENU --> O6["6 Generate README\n→ LLM"]
180
+ MENU --> O7["7 Propose & create file\n→ LLM + local write"]
181
+ MENU --> O8["8 List GitHub MCP tools\n→ GitHub Copilot MCP"]
182
+ MENU --> O9["9 Re-ingest repo\n→ new repo path"]
183
+ MENU --> O0["0 Exit"]
184
+
185
+ O1 & O2 & O3 & O4 & O5 & O6 & O7 & O8 & O9 --> MENU
186
+ O0 --> END([Goodbye!])
187
+ ```
188
+
189
+ ---
190
+
191
+ ## 4. Component Deep-Dive
192
+
193
+ ### 4.1 `config.py` β€” Settings
194
+
195
+ All configuration lives in one `pydantic-settings` class loaded from `.env`:
196
+
197
+ | Setting | Default | Description |
198
+ |---|---|---|
199
+ | `llm_provider` | `"groq"` | `"groq"` or `"gemini"` |
200
+ | `groq_api_key` | `""` | From `.env` |
201
+ | `groq_model` | `"llama-3.3-70b-versatile"` | Free Groq model |
202
+ | `gemini_api_key` | `""` | From `.env` |
203
+ | `embedding_model` | `"models/gemini-embedding-001"` | Always Gemini for embeddings |
204
+ | `vector_db_path` | `"./vector_db"` | ChromaDB storage directory |
205
+ | `chunk_size` | `1200` | Characters per chunk |
206
+ | `chunk_overlap` | `100` | Overlap between chunks |
207
+ | `allowed_extensions` | `[.py .js .ts .go .java .md]` | File types to load |
208
+ | `max_file_size_kb` | `1042` | Skip files larger than this |
209
+ | `github_mcp_url` | GitHub Copilot MCP endpoint | For option 8 |
210
+ | `github_mcp_token` | `""` | GitHub PAT from `.env` |
211
+
212
+ ### 4.2 `rag/repository_loader.py` β€” File Ingestion
213
+
214
+ ```
215
+ load_repository(root_path)
216
+ └── os.walk(root_path)
217
+ β”œβ”€β”€ Skip: .git, node_modules, __pycache__, venv, .venv, dist, build
218
+ β”œβ”€β”€ should_include(fpath) β†’ checks extension + file size
219
+ └── read_file_with_metadata(fpath) β†’ CodeDocument(content, file_path, language, size_bytes)
220
+ ```
221
+
222
+ **Supported languages detected by extension:**
223
+
224
+ | Extension | Language |
225
+ |---|---|
226
+ | `.py` | python |
227
+ | `.js` | javascript |
228
+ | `.ts` | typescript |
229
+ | `.java` | java |
230
+ | `.go` | go |
231
+ | `.md` | markdown |
232
+
233
+ ### 4.3 `rag/splitter.py` β€” Chunking
234
+
235
+ Uses **LangChain's `RecursiveCharacterTextSplitter`** with language-aware splitting:
236
+ - For Python/JS/TS/Java/Go β€” uses syntax-aware boundaries (functions, classes)
237
+ - For Markdown/unknown β€” falls back to generic character splitting
238
+ - Each chunk carries: `content`, `file_path`, `language`, `chunk_index`, `start_line`, `end_line`
239
+
240
+ ### 4.4 `rag/embedding.py` β€” Embeddings
241
+
242
+ - Uses `GoogleGenerativeAIEmbeddings` (`models/gemini-embedding-001`)
243
+ - `embed_document(chunks)` β€” batch embeds all chunks, mutates dicts in-place
244
+ - `embed_query(query)` β€” single query vector for similarity search
245
+
246
+ ### 4.5 `rag/retriever.py` β€” Vector Store
247
+
248
+ > **Critical fix applied:** collection is deleted and recreated on every ingest to prevent stale data from previous repos bleeding through.
249
+
250
+ ```python
251
+ build_vector_store(chunks) # delete β†’ create β†’ add (fresh each ingest)
252
+ load_vector_store() # get_or_create for reading
253
+ retrieve_relevant_chunks(q, k) # embed query β†’ cosine similarity β†’ top-k
254
+ ```
255
+
256
+ ### 4.6 `rag/rag_chain.py` β€” Orchestration
257
+
258
+ ```python
259
+ run_rag_query(query, k=5)
260
+ 1. retrieve_relevant_chunks(query, k)
261
+ 2. build_context(chunks) # format with file/line headers
262
+ 3. build_prompt(query, context) # fill qa template
263
+ 4. llm.generate(prompt)
264
+ 5. return { answer, sources }
265
+ ```
266
+
267
+ ### 4.7 `llm.py` β€” LLM Abstraction
268
+
269
+ Abstract `BaseLLM` with two concrete providers:
270
+
271
+ | Class | Provider | API |
272
+ |---|---|---|
273
+ | `GeminiLLM` | Google Gemini | `google-genai` SDK |
274
+ | `GroqLLM` | Groq | `groq` SDK, chat completions |
275
+
276
+ **Prompt templates (`build_prompt`):**
277
+
278
+ | `task_type` | Used by |
279
+ |---|---|
280
+ | `"qa"` | RAG query, function explain, module docs, README |
281
+ | `"bug_finding"` | Bug detection (returns JSON) |
282
+ | `"docstring"` | Docstring generation |
283
+ | `"file_creation"` | AI file proposal |
284
+
285
+ ### 4.8 `mcp_client.py` β€” GitHub MCP Client
286
+
287
+ Async context-manager pattern using `AsyncExitStack` to keep all `anyio` cancel scopes in the **same task**:
288
+
289
+ ```python
290
+ async with get_github_mcp_client() as client:
291
+ tools = await client.list_tools()
292
+ result = await client.call_tool("create_branch", {...})
293
+ ```
294
+
295
+ Available GitHub MCP tools (44 total) include: `search_code`, `list_issues`, `create_pull_request`, `get_file_contents`, `push_files`, `create_branch`, `fork_repository`, `search_repositories`, and many more.
296
+
297
+ ### 4.9 `services/code_analysis.py`
298
+
299
+ | Function | How it works |
300
+ |---|---|
301
+ | `explain_function(file, name)` | Python `ast` extracts the function source β†’ RAG context β†’ LLM |
302
+ | `detect_bugs(file)` | Read file β†’ `bug_finding` prompt β†’ LLM returns JSON list |
303
+ | `analyze_complexity(file)` | `radon.cc_visit` β†’ cyclomatic complexity + rank A–F |
304
+
305
+ ### 4.10 `services/file_creator.py`
306
+
307
+ ```
308
+ propose_new_file(description, context_query)
309
+ β”œβ”€β”€ RAG retrieve relevant context
310
+ β”œβ”€β”€ build_prompt(description, context, "file_creation")
311
+ β”œβ”€β”€ LLM generates complete file content
312
+ └── infer_file_path(description) β†’ LLM suggests relative path
313
+
314
+ apply_approved_file(proposal, confirmed, base_dir)
315
+ β”œβ”€β”€ Strip leading / or \ from LLM path
316
+ β”œβ”€β”€ os.path.join(base_dir, relative_path)
317
+ β”œβ”€β”€ os.makedirs(parent_dirs, exist_ok=True)
318
+ └── open(abs_path, "w").write(content)
319
+ ```
320
+
321
+ ---
322
+
323
+ ## 5. Setup & Installation
324
+
325
+ ### Prerequisites
326
+
327
+ | Requirement | Version |
328
+ |---|---|
329
+ | Python | 3.11+ |
330
+ | pip | Latest |
331
+ | Internet | For API calls |
332
+
333
+ ### Step 1 β€” Clone / Download the project
334
+
335
+ ```bash
336
+ git clone <your-repo-url>
337
+ cd "Codebase Assistant"
338
+ ```
339
+
340
+ ### Step 2 β€” Create a virtual environment
341
+
342
+ ```bash
343
+ # Windows (PowerShell)
344
+ python -m venv .venv
345
+ .venv\Scripts\Activate.ps1
346
+
347
+ # macOS / Linux
348
+ python3 -m venv .venv
349
+ source .venv/bin/activate
350
+ ```
351
+
352
+ ### Step 3 β€” Install dependencies
353
+
354
+ ```bash
355
+ pip install -r requirements.txt
356
+ ```
357
+
358
+ > [!TIP]
359
+ > If you hit permission issues on Windows, use:
360
+ > `pip install -r requirements.txt --user`
361
+
362
+ ### Step 4 β€” Create your `.env` file
363
+
364
+ Create a file named `.env` in the project root:
365
+
366
+ ```env
367
+ # Choose your LLM provider: "groq" (free) or "gemini"
368
+ LLM_PROVIDER=groq
369
+
370
+ # Groq β€” free at https://console.groq.com
371
+ GROQ_API_KEY=gsk_your_key_here
372
+
373
+ # Gemini β€” get at https://aistudio.google.com
374
+ GEMINI_API_KEY=your_gemini_key_here
375
+
376
+ # GitHub PAT for MCP tools β€” create at https://github.com/settings/tokens
377
+ # Required scopes: repo, read:org
378
+ GITHUB_MCP_TOKEN=github_pat_your_token_here
379
+ ```
380
+
381
+ > [!IMPORTANT]
382
+ > `GEMINI_API_KEY` is **always required** regardless of LLM provider, because embeddings always use Gemini.
383
+
384
+ ### Step 5 β€” Run the CLI
385
+
386
+ ```bash
387
+ .venv\Scripts\python.exe cli.py # Windows
388
+ python cli.py # macOS / Linux
389
+ ```
390
+
391
+ ---
392
+
393
+ ## 6. Configuration (.env)
394
+
395
+ ```env
396
+ # ─── LLM Provider ────────────────────────────────────────
397
+ LLM_PROVIDER=groq # "groq" | "gemini"
398
+
399
+ # ─── Groq (recommended β€” free tier) ─────────────────────
400
+ GROQ_API_KEY=gsk_...
401
+ GROQ_MODEL=llama-3.3-70b-versatile
402
+
403
+ # ─── Gemini ──────────────────────────────────────────────
404
+ GEMINI_API_KEY=...
405
+ LLM_MODEL=gemini-1.5-flash # only used if LLM_PROVIDER=gemini
406
+ EMBEDDING_MODEL=models/gemini-embedding-001
407
+
408
+ # ─── RAG / Vector DB ─────────────────────────────────────
409
+ VECTOR_DB_PATH=./vector_db
410
+ CHUNK_SIZE=1200
411
+ CHUNK_OVERLAP=100
412
+
413
+ # ─── File loader ─────────────────────────────────────────
414
+ # comma-separated extensions
415
+ ALLOWED_EXTENSIONS=[".py",".js",".ts",".go",".java",".md"]
416
+ MAX_FILE_SIZE_KB=1042
417
+
418
+ # ─── GitHub MCP ──────────────────────────────────────────
419
+ GITHUB_MCP_URL=https://api.githubcopilot.com/mcp/
420
+ GITHUB_MCP_TOKEN=github_pat_...
421
+ ```
422
+
423
+ ---
424
+
425
+ ## 7. Running the CLI
426
+
427
+ ```
428
+ ╔══════════════════════════════════════════╗
429
+ β•‘ AI Codebase Assistant CLI β•‘
430
+ β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
431
+
432
+ β–Ά Enter the path to the repository you want to analyse: D:\MyProject
433
+
434
+ ── Ingesting Repository ──────────────────────
435
+ Loading files from: D:\MyProject
436
+ βœ“ Loaded 12 files
437
+ βœ“ Split into 47 chunks
438
+ Embedding chunks (this may take a moment)...
439
+ βœ“ Embedded 47 chunks
440
+ βœ“ Vector store ready (ChromaDB)
441
+
442
+ ── What would you like to do? ────────────────
443
+ 1 Ask a question about the codebase
444
+ 2 Detect bugs in a file
445
+ 3 Cyclomatic complexity analysis
446
+ 4 Explain a function
447
+ 5 Generate module documentation
448
+ 6 Generate README for a repository
449
+ 7 Propose & create a new file (AI)
450
+ 8 List GitHub MCP tools
451
+ 9 Re-ingest a repository
452
+ 0 Exit
453
+
454
+ β–Ά Choose an option [0–9]:
455
+ ```
456
+
457
+ ---
458
+
459
+ ## 8. CLI Features β€” All 9 Options
460
+
461
+ ### Option 1 β€” Ask a Question (RAG Query)
462
+
463
+ ```
464
+ β–Ά Your question: How does the authentication work?
465
+ β–Ά Number of source chunks to retrieve? [default: 5] 3
466
+
467
+ ── Answer ────────────────────────────────────
468
+ The authentication uses JWT tokens ...
469
+
470
+ ── Sources ───────────────────────────────────
471
+ β€’ src/auth/middleware.py lines 12–45
472
+ β€’ src/auth/tokens.py lines 1–30
473
+ ```
474
+
475
+ **Pipeline:** Embed question β†’ ChromaDB similarity search β†’ Build context β†’ LLM β†’ Answer + cited sources
476
+
477
+ ---
478
+
479
+ ### Option 2 β€” Detect Bugs
480
+
481
+ ```
482
+ β–Ά File path to analyse: src/payment.py
483
+
484
+ ── Found 2 issue(s) ──────────────────────────
485
+ [HIGH] Line 34: SQL query uses string concatenation
486
+ β†’ Use parameterized queries to prevent SQL injection
487
+
488
+ [MEDIUM] Line 67: Exception swallowed silently
489
+ β†’ Log or re-raise the exception
490
+ ```
491
+
492
+ **Pipeline:** Read file β†’ `bug_finding` prompt β†’ LLM returns JSON β†’ parsed and displayed
493
+
494
+ ---
495
+
496
+ ### Option 3 β€” Cyclomatic Complexity
497
+
498
+ ```
499
+ β–Ά File path to analyse: src/processor.py
500
+
501
+ ── 4 function(s) ─────────────────────────────
502
+ [A] process_order complexity=2 β–ˆβ–ˆ
503
+ [B] validate_cart complexity=5 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
504
+ [C] apply_discounts complexity=8 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
505
+ [F] handle_edge_cases complexity=18 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
506
+ ```
507
+
508
+ **Rank scale:** A (1–5, simple) β†’ F (26+, untestable)
509
+
510
+ ---
511
+
512
+ ### Option 4 β€” Explain a Function
513
+
514
+ ```
515
+ β–Ά File path: src/utils.py
516
+ β–Ά Function name: parse_date_range
517
+
518
+ ── Explanation ───────────────────────────────
519
+ parse_date_range takes a string like "2024-01-01:2024-12-31"
520
+ and returns a tuple of (start_date, end_date) as datetime objects ...
521
+ ```
522
+
523
+ **Pipeline:** Python `ast` module extracts exact function source β†’ RAG retrieves usages β†’ LLM explains
524
+
525
+ ---
526
+
527
+ ### Option 5 β€” Generate Module Documentation
528
+
529
+ ```
530
+ β–Ά File path: src/database.py
531
+
532
+ ── Documentation ─────────────────────────────
533
+ ## database.py
534
+
535
+ ### Overview
536
+ This module provides the database connection layer ...
537
+
538
+ ### Functions
539
+ - `connect(url)` β€” Establishes a connection ...
540
+ - `execute(query, params)` β€” Runs a parameterized query ...
541
+ ```
542
+
543
+ ---
544
+
545
+ ### Option 6 β€” Generate README
546
+
547
+ ```
548
+ β–Ά Repository root path: D:\MyProject
549
+
550
+ ── README Preview ────────────────────────────
551
+ # MyProject
552
+
553
+ ## Overview
554
+ A FastAPI application that ...
555
+
556
+ β–Ά Save to README.md in that directory? [y/N] y
557
+ βœ“ Saved to D:\MyProject\README.md
558
+ ```
559
+
560
+ ---
561
+
562
+ ### Option 7 β€” Propose & Create a New File
563
+
564
+ ```
565
+ β–Ά Describe the file you want to create: Rectangle area calculator in JavaScript
566
+ β–Ά Optional context query (or press Enter to skip):
567
+
568
+ ── Proposed File ─────────────────────────────
569
+ Path: /src/geometry/rectangle.js
570
+
571
+ class Rectangle { ...full generated code... }
572
+
573
+ β–Ά Write this file to disk? [y/N] y
574
+ β–Ά Save under which directory? [default: D:\MyProject]
575
+
576
+ βœ“ File written: D:\MyProject\src\geometry\rectangle.js
577
+ ```
578
+
579
+ **Pipeline:** RAG context (optional) β†’ `file_creation` prompt β†’ LLM generates code β†’ user confirms β†’ write to `base_dir/relative_path`
580
+
581
+ ---
582
+
583
+ ### Option 8 β€” List GitHub MCP Tools
584
+
585
+ ```
586
+ ── List GitHub MCP Tools ─────────────────────
587
+ 44 tools available:
588
+ β€’ create_branch Create a new branch in a GitHub repository
589
+ β€’ create_pull_request Create a new pull request ...
590
+ β€’ search_code Fast and precise code search ...
591
+ β€’ list_issues List issues in a GitHub repository ...
592
+ ...
593
+ ```
594
+
595
+ **Connection:** Uses `streamable_http_client` β†’ `ClientSession` β†’ GitHub Copilot MCP endpoint
596
+
597
+ ---
598
+
599
+ ### Option 9 β€” Re-ingest a Repository
600
+
601
+ ```
602
+ β–Ά New repository path to ingest: D:\AnotherProject
603
+ ── Ingesting Repository ──────────────────────
604
+ βœ“ Loaded 8 files ...
605
+ ```
606
+
607
+ Wipes the ChromaDB collection and ingests the new repo fresh. All subsequent queries answer from the new repo only.
608
+
609
+ ---
610
+
611
+ ## 9. REST API
612
+
613
+ Start the FastAPI server:
614
+
615
+ ```bash
616
+ .venv\Scripts\python.exe -m uvicorn app:app --reload --port 8000
617
+ ```
618
+
619
+ Open docs at: `http://localhost:8000/docs`
620
+
621
+ ### Endpoints
622
+
623
+ | Method | Path | Description |
624
+ |---|---|---|
625
+ | `GET` | `/health` | Health check |
626
+ | `POST` | `/api/query` | RAG question answering |
627
+ | `POST` | `/api/analyze/bugs?file_path=...` | Bug detection |
628
+ | `POST` | `/api/analyze/complexity?file_path=...` | Cyclomatic complexity |
629
+ | `POST` | `/api/analyze/explain` | Explain a function |
630
+ | `POST` | `/api/docs/module?file_path=...` | Module documentation |
631
+ | `POST` | `/api/docs/readme?root_path=...` | README generation |
632
+ | `POST` | `/api/files/propose` | Propose a new file |
633
+ | `POST` | `/api/files/approve` | Write approved file to disk |
634
+
635
+ ### Example β€” Query
636
+
637
+ ```bash
638
+ curl -X POST http://localhost:8000/api/query \
639
+ -H "Content-Type: application/json" \
640
+ -d '{"query": "How does authentication work?", "k": 5}'
641
+ ```
642
+
643
+ ```json
644
+ {
645
+ "answer": "Authentication is handled by ...",
646
+ "sources": [
647
+ {"file_path": "src/auth.py", "start_line": 10, "end_line": 45}
648
+ ]
649
+ }
650
+ ```
651
+
652
+ ---
653
+
654
+ ## 10. Key Design Decisions & Bug Fixes
655
+
656
+ ### Bug Fix 1 β€” Collection Isolation (retriever.py)
657
+
658
+ **Problem:** ChromaDB used `upsert` on a shared `"codebase"` collection β€” stale chunks from previously ingested repos leaked into new queries.
659
+
660
+ **Fix:** On every ingest, `delete_collection` + `create_collection` ensures a clean slate:
661
+
662
+ ```python
663
+ # Before (bug)
664
+ collection = client.get_or_create_collection("codebase")
665
+ collection.upsert(...)
666
+
667
+ # After (fix)
668
+ client.delete_collection("codebase") # wipe old repo
669
+ collection = client.create_collection("codebase")
670
+ collection.add(...)
671
+ ```
672
+
673
+ ---
674
+
675
+ ### Bug Fix 2 β€” GitHub MCP Cancel Scope (mcp_client.py)
676
+
677
+ **Problem:** Manually calling `__aenter__`/`__aexit__` on `anyio`-backed context managers across tasks causes `RuntimeError: Attempted to exit cancel scope in a different task`.
678
+
679
+ **Fix:** Use `AsyncExitStack` to nest all context managers inside a single `async with`:
680
+
681
+ ```python
682
+ async with get_github_mcp_client() as client:
683
+ tools = await client.list_tools()
684
+ ```
685
+
686
+ ---
687
+
688
+ ### Bug Fix 3 β€” Wrong Unpack Count (mcp_client.py)
689
+
690
+ **Problem:** `streamable_http_client` yields 2 values, not 3. Unpacking 3 caused `ValueError: not enough values to unpack`.
691
+
692
+ ```python
693
+ # Before (bug)
694
+ read, write, _ = await ctx.__aenter__()
695
+
696
+ # After (fix)
697
+ read, write = await stack.enter_async_context(streamable_http_client(...))
698
+ ```
699
+
700
+ ---
701
+
702
+ ### Bug Fix 4 β€” File Save Path (file_creator.py)
703
+
704
+ **Problem:** Generated files saved to the current working directory regardless of user input.
705
+
706
+ **Fix:**
707
+ - Strip leading `/\` from LLM path to make it always relative
708
+ - Accept `base_dir` defaulting to `_current_repo` (the ingested repo path)
709
+ - Create all parent directories automatically
710
+
711
+ ---
712
+
713
+ ## 11. Extending the Project
714
+
715
+ ### Add a new LLM provider
716
+
717
+ 1. Add a new class in `llm.py` extending `BaseLLM`
718
+ 2. Add the provider name to `get_llm_client()` factory
719
+ 3. Add matching settings in `config.py`
720
+
721
+ ### Add support for new file types
722
+
723
+ 1. Add extension β†’ language in `LANGUAGE_BY_EXT` in `repository_loader.py`
724
+ 2. Add extension to `allowed_extensions` in `config.py`
725
+ 3. If LangChain has a `Language` enum for it, add to `LANGUAGE_MAP` in `splitter.py`
726
+
727
+ ### Use a GitHub MCP tool in a feature
728
+
729
+ ```python
730
+ async with get_github_mcp_client() as client:
731
+ result = await client.call_tool("create_branch", {
732
+ "owner": "myuser",
733
+ "repo": "myrepo",
734
+ "branch": "feature/new-branch",
735
+ "from_branch": "main"
736
+ })
737
+ ```
738
+
739
+ ### Add a new CLI option
740
+
741
+ 1. Write a `feature_xxx()` function in `cli.py`
742
+ 2. Add `("Label", feature_xxx)` to the `MENU` list
743
+ 3. Add a matching FastAPI endpoint in `api/routes.py`
744
+
745
+ ---
746
+
app.py ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from fastapi import FastAPI
2
+ from fastapi.middleware.cors import CORSMiddleware
3
+
4
+ from api.routes import router
5
+ from mcp_client import get_github_mcp_client
6
+
7
+ app = FastAPI(title="AI Codebase Assistant")
8
+
9
+ app.add_middleware(
10
+ CORSMiddleware,
11
+ allow_origins=["*"],
12
+ allow_credentials=True,
13
+ allow_methods=["*"],
14
+ allow_headers=["*"],
15
+ )
16
+
17
+ app.include_router(router, prefix="/api")
18
+
19
+ mcp_client_instance = get_github_mcp_client()
20
+
21
+
22
+ @app.on_event("startup")
23
+ async def startup_event():
24
+ await mcp_client_instance.connect()
25
+
26
+
27
+ @app.on_event("shutdown")
28
+ async def shutdown_event():
29
+ await mcp_client_instance.disconnect()
30
+
31
+
32
+ @app.get("/health")
33
+ def health_check():
34
+ return {"status": "ok"}
cli.py ADDED
@@ -0,0 +1,323 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ Codebase Assistant β€” Interactive CLI
3
+ Run with: .venv\Scripts\python.exe cli.py
4
+ """
5
+
6
+ import os
7
+ import sys
8
+ import asyncio
9
+
10
+ # ── styling ───────────────────────────────────────────────────────────────────
11
+
12
+ CYAN = "\033[96m"
13
+ GREEN = "\033[92m"
14
+ YELLOW = "\033[93m"
15
+ RED = "\033[91m"
16
+ BOLD = "\033[1m"
17
+ DIM = "\033[2m"
18
+ RESET = "\033[0m"
19
+
20
+ # ── session state ─────────────────────────────────────────────────────────────
21
+ # Tracks the repo path provided at startup (or via re-ingest) so that option 7
22
+ # can save generated files directly into that directory without asking again.
23
+ _current_repo: str = "."
24
+
25
+ def banner():
26
+ print(f"""
27
+ {CYAN}{BOLD}
28
+ ╔══════════════════════════════════════════╗
29
+ β•‘ AI Codebase Assistant CLI β•‘
30
+ β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
31
+ {RESET}""")
32
+
33
+ def section(title: str):
34
+ print(f"\n{CYAN}{BOLD}── {title} {'─' * (42 - len(title))}{RESET}")
35
+
36
+ def ok(msg: str):
37
+ print(f" {GREEN}βœ“{RESET} {msg}")
38
+
39
+ def err(msg: str):
40
+ print(f" {RED}βœ— {msg}{RESET}")
41
+
42
+ def info(msg: str):
43
+ print(f" {DIM}{msg}{RESET}")
44
+
45
+ def ask(prompt: str) -> str:
46
+ return input(f"\n{YELLOW}β–Ά {prompt}{RESET} ").strip()
47
+
48
+ # ── ingest ────────────────────────────────────────────────────────────────────
49
+
50
+ def ingest_repository(root_path: str):
51
+ global _current_repo
52
+ section("Ingesting Repository")
53
+ from rag.repository_loader import load_repository
54
+ from rag.splitter import split_code
55
+ from rag.embedding import embed_document
56
+ from rag.retriever import build_vector_store
57
+
58
+ print(f" {DIM}Loading files from: {root_path}{RESET}")
59
+ docs = load_repository(root_path)
60
+ if not docs:
61
+ err("No supported files found. Check allowed_extensions in config.py.")
62
+ return False
63
+ ok(f"Loaded {len(docs)} files")
64
+
65
+ all_chunks = []
66
+ for doc in docs:
67
+ all_chunks.extend(split_code(doc))
68
+ ok(f"Split into {len(all_chunks)} chunks")
69
+
70
+ print(f" {DIM}Embedding chunks (this may take a moment)...{RESET}")
71
+ embedded = embed_document(all_chunks)
72
+ ok(f"Embedded {len(embedded)} chunks")
73
+
74
+ build_vector_store(embedded)
75
+ ok("Vector store ready (ChromaDB)\n")
76
+ _current_repo = root_path # remember for option 7
77
+ return True
78
+
79
+ # ── feature handlers ──────────────────────────────────────────────────────────
80
+
81
+ def feature_query():
82
+ section("Ask a Question [RAG Query]")
83
+ query = ask("Your question:")
84
+ if not query:
85
+ return
86
+ k_str = ask("Number of source chunks to retrieve? [default: 5]")
87
+ k = int(k_str) if k_str.isdigit() else 5
88
+
89
+ from rag.rag_chain import run_rag_query
90
+ print(f"\n {DIM}Thinking...{RESET}")
91
+ try:
92
+ result = run_rag_query(query, k=k)
93
+ section("Answer")
94
+ print(f"\n{result['answer']}")
95
+ section("Sources")
96
+ for s in result["sources"]:
97
+ print(f" {DIM}β€’ {s.get('file_path')} lines {s.get('start_line')}–{s.get('end_line')}{RESET}")
98
+ except Exception as e:
99
+ err(str(e))
100
+
101
+
102
+ def feature_bugs():
103
+ section("Detect Bugs [LLM Code Review]")
104
+ file_path = ask("File path to analyse:")
105
+ if not os.path.isfile(file_path):
106
+ err(f"File not found: {file_path}")
107
+ return
108
+
109
+ from services.code_analysis import detect_bugs
110
+ print(f"\n {DIM}Analysing...{RESET}")
111
+ try:
112
+ bugs = detect_bugs(file_path)
113
+ if not bugs:
114
+ ok("No issues found.")
115
+ return
116
+ section(f"Found {len(bugs)} issue(s)")
117
+ for b in bugs:
118
+ sev = b.get("severity", "?").upper()
119
+ color = RED if sev in ("HIGH", "CRITICAL") else YELLOW
120
+ print(f"\n {color}[{sev}]{RESET} Line {b.get('line', '?')}: {b.get('issue', '')}")
121
+ print(f" {DIM} β†’ {b.get('suggestion', '')}{RESET}")
122
+ except Exception as e:
123
+ err(str(e))
124
+
125
+
126
+ def feature_complexity():
127
+ section("Cyclomatic Complexity Analysis")
128
+ file_path = ask("File path to analyse:")
129
+ if not os.path.isfile(file_path):
130
+ err(f"File not found: {file_path}")
131
+ return
132
+
133
+ from services.code_analysis import analyze_complexity
134
+ try:
135
+ result = analyze_complexity(file_path)
136
+ fns = result["functions"]
137
+ if not fns:
138
+ info("No functions found (or file is not Python).")
139
+ return
140
+ section(f"{len(fns)} function(s)")
141
+ for fn in fns:
142
+ rank = fn["rank"]
143
+ color = GREEN if rank == "A" else (YELLOW if rank in ("B", "C") else RED)
144
+ bar = "β–ˆ" * fn["complexity"]
145
+ print(f" {color}[{rank}]{RESET} {fn['name']:<30} complexity={fn['complexity']} {DIM}{bar}{RESET}")
146
+ except Exception as e:
147
+ err(str(e))
148
+
149
+
150
+ def feature_explain():
151
+ section("Explain a Function")
152
+ file_path = ask("File path:")
153
+ if not os.path.isfile(file_path):
154
+ err(f"File not found: {file_path}")
155
+ return
156
+ func_name = ask("Function name:")
157
+ if not func_name:
158
+ return
159
+
160
+ from services.code_analysis import explain_function
161
+ print(f"\n {DIM}Thinking...{RESET}")
162
+ try:
163
+ explanation = explain_function(file_path, func_name)
164
+ section("Explanation")
165
+ print(f"\n{explanation}")
166
+ except Exception as e:
167
+ err(str(e))
168
+
169
+
170
+ def feature_module_docs():
171
+ section("Generate Module Documentation")
172
+ file_path = ask("File path:")
173
+ if not os.path.isfile(file_path):
174
+ err(f"File not found: {file_path}")
175
+ return
176
+
177
+ from services.documentation import generate_module_docs
178
+ print(f"\n {DIM}Generating docs...{RESET}")
179
+ try:
180
+ docs = generate_module_docs(file_path)
181
+ section("Documentation")
182
+ print(f"\n{docs}")
183
+ except Exception as e:
184
+ err(str(e))
185
+
186
+
187
+ def feature_readme():
188
+ section("Generate README.md")
189
+ root = ask("Repository root path:")
190
+ if not os.path.isdir(root):
191
+ err(f"Directory not found: {root}")
192
+ return
193
+
194
+ from services.documentation import generate_readme
195
+ print(f"\n {DIM}Generating README...{RESET}")
196
+ try:
197
+ readme = generate_readme(root)
198
+ section("README Preview")
199
+ print(f"\n{readme[:1500]}{'...' if len(readme) > 1500 else ''}")
200
+ save = ask("Save to README.md in that directory? [y/N]")
201
+ if save.lower() == "y":
202
+ out = os.path.join(root, "README.md")
203
+ with open(out, "w", encoding="utf-8") as f:
204
+ f.write(readme)
205
+ ok(f"Saved to {out}")
206
+ except Exception as e:
207
+ err(str(e))
208
+
209
+
210
+ def feature_propose_file():
211
+ section("Propose a New File [AI File Creator]")
212
+ description = ask("Describe the file you want to create:")
213
+ if not description:
214
+ return
215
+ context_query = ask("Optional context query (or press Enter to skip):")
216
+
217
+ from services.file_creator import propose_new_file
218
+ print(f"\n {DIM}Generating...{RESET}")
219
+ try:
220
+ proposal = propose_new_file(description, context_query or None)
221
+ section("Proposed File")
222
+ print(f"\n {BOLD}Path:{RESET} {proposal.path}")
223
+ print(f"\n{DIM}{'─'*50}{RESET}")
224
+ print(proposal.content[:1000] + ("..." if len(proposal.content) > 1000 else ""))
225
+ print(f"{DIM}{'─'*50}{RESET}")
226
+
227
+ confirm = ask("Write this file to disk? [y/N]")
228
+ if confirm.lower() == "y":
229
+ default_dir = _current_repo
230
+ typed = ask(f"Save under which directory? [default: {default_dir}]").strip()
231
+ base_dir = typed if typed else default_dir
232
+ if not os.path.isdir(base_dir):
233
+ err(f"Directory not found: {base_dir}")
234
+ return
235
+ async def _write():
236
+ from services.file_creator import apply_approved_file
237
+ proposal.approved = True
238
+ return await apply_approved_file(proposal, True, base_dir=base_dir)
239
+ result = asyncio.run(_write())
240
+ if result.get("status") == "written":
241
+ ok(f"File written: {result.get('path', proposal.path)}")
242
+ else:
243
+ err(f"Write failed: {result}")
244
+ except Exception as e:
245
+ err(str(e))
246
+
247
+
248
+ def feature_mcp_tools():
249
+ section("List GitHub MCP Tools")
250
+ async def _list():
251
+ from mcp_client import get_github_mcp_client
252
+ async with get_github_mcp_client() as client:
253
+ return await client.list_tools()
254
+ try:
255
+ tools = asyncio.run(_list())
256
+ ok(f"{len(tools)} tools available:")
257
+ for t in tools:
258
+ print(f" {DIM}β€’ {t['name']:<30}{RESET} {t.get('description','')[:60]}")
259
+ except Exception as e:
260
+ err(str(e))
261
+
262
+
263
+ def feature_reingest():
264
+ root = ask("New repository path to ingest:")
265
+ if os.path.isdir(root):
266
+ ingest_repository(root)
267
+ else:
268
+ err(f"Directory not found: {root}")
269
+
270
+ # ── menu ──────────────────────────────────────────────────────────────────────
271
+
272
+ MENU = [
273
+ ("Ask a question about the codebase", feature_query),
274
+ ("Detect bugs in a file", feature_bugs),
275
+ ("Cyclomatic complexity analysis", feature_complexity),
276
+ ("Explain a function", feature_explain),
277
+ ("Generate module documentation", feature_module_docs),
278
+ ("Generate README for a repository", feature_readme),
279
+ ("Propose & create a new file (AI)", feature_propose_file),
280
+ ("List GitHub MCP tools", feature_mcp_tools),
281
+ ("Re-ingest a repository", feature_reingest),
282
+ ]
283
+
284
+ def show_menu():
285
+ section("What would you like to do?")
286
+ for i, (label, _) in enumerate(MENU, 1):
287
+ print(f" {CYAN}{i:>2}{RESET} {label}")
288
+ print(f" {DIM} 0 Exit{RESET}")
289
+
290
+ # ── main ──────────────────────────────────────────────────────────────────────
291
+
292
+ def main():
293
+ # Enable ANSI colours on Windows
294
+ os.system("")
295
+
296
+ banner()
297
+
298
+ # Step 1 β€” ask for repo path
299
+ while True:
300
+ root = ask("Enter the path to the repository you want to analyse:")
301
+ if os.path.isdir(root):
302
+ break
303
+ err(f"Directory not found: {root}")
304
+
305
+ # Step 2 β€” ingest
306
+ if not ingest_repository(root):
307
+ sys.exit(1)
308
+
309
+ # Step 3 β€” interactive menu loop
310
+ while True:
311
+ show_menu()
312
+ choice = ask("Choose an option [0–9]:")
313
+ if choice == "0" or choice.lower() in ("exit", "quit", "q"):
314
+ print(f"\n{DIM}Goodbye!{RESET}\n")
315
+ break
316
+ if choice.isdigit() and 1 <= int(choice) <= len(MENU):
317
+ MENU[int(choice) - 1][1]()
318
+ else:
319
+ err("Invalid choice. Enter a number from the menu.")
320
+
321
+
322
+ if __name__ == "__main__":
323
+ main()
config.py ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from pydantic_settings import BaseSettings
2
+ from functools import lru_cache
3
+
4
+
5
+ class Settings(BaseSettings):
6
+ # ── LLM Provider ──────────────────────────────────────────────────────────
7
+ # Set LLM_PROVIDER to "gemini" or "groq" in your .env file.
8
+ # Groq is free and has no strict daily quota β€” recommended when Gemini
9
+ # free-tier is exhausted.
10
+ llm_provider: str = "groq"
11
+
12
+ # ── Gemini settings ───────────────────────────────────────────────────────
13
+ gemini_api_key: str = ""
14
+ llm_model: str = "llama-3.3-70b-versatile" # overridden per provider below
15
+ embedding_model: str = "models/gemini-embedding-001"
16
+
17
+ # ── Groq settings ─────────────────────────────────────────────────────────
18
+ groq_api_key: str = ""
19
+ groq_model: str = "llama-3.3-70b-versatile" # free, 6k tokens/min on Groq
20
+
21
+ # ── RAG / Vector DB ───────────────────────────────────────────────────────
22
+ vector_db_path: str = "./vector_db"
23
+ chunk_size: int = 1200
24
+ chunk_overlap: int = 100
25
+
26
+ # ── MCP ───────────────────────────────────────────────────────────────────
27
+ mcp_server_url: str = "http://localhost:3333"
28
+ github_mcp_url: str = "https://api.githubcopilot.com/mcp/"
29
+ github_mcp_token: str = ""
30
+
31
+ # ── Loader ────────────────────────────────────────────────────────────────
32
+ allowed_extensions: list[str] = [".py", ".js", ".ts", ".go", ".java", ".md"]
33
+ max_file_size_kb: int = 1042
34
+
35
+ class Config:
36
+ env_file = ".env"
37
+
38
+
39
+ @lru_cache()
40
+ def get_settings() -> Settings:
41
+ return Settings()
example.py ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ class Area:
2
+ # Function to calculate rectangle area
3
+ def rectangle_area(self, length, width):
4
+ return length * width
5
+
6
+ # Function to calculate square area
7
+ def square_area(self, side):
8
+ return side * side
9
+
10
+
11
+ # Create object
12
+ obj = Area()
13
+
14
+ # Input
15
+ length = int(input("Enter length of rectangle: "))
16
+ width = int(input("Enter width of rectangle: "))
17
+ side = int(input("Enter side of square: "))
18
+
19
+ # Output
20
+ print("Rectangle Area =", obj.rectangle_area(length, width))
21
+ print("Square Area =", obj.square_area(side))
llm.py ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from abc import ABC, abstractmethod
2
+ from config import get_settings
3
+
4
+
5
+ class BaseLLM(ABC):
6
+ @abstractmethod
7
+ def generate(self, prompt: str, system: str = "") -> str:
8
+ ...
9
+
10
+ @abstractmethod
11
+ def stream(self, prompt: str, system: str = ""):
12
+ ...
13
+
14
+
15
+ # ── Gemini provider ────────────────────────────────────────────────────────────
16
+
17
+ class GeminiLLM(BaseLLM):
18
+ def __init__(self, api_key: str, model: str):
19
+ from google import genai
20
+ self.client = genai.Client(api_key=api_key)
21
+ self.model = model
22
+
23
+ def generate(self, prompt: str, system: str = "") -> str:
24
+ full_prompt = f"{system}\n\n{prompt}" if system else prompt
25
+ response = self.client.models.generate_content(
26
+ model=self.model,
27
+ contents=full_prompt,
28
+ )
29
+ return response.text
30
+
31
+ def stream(self, prompt: str, system: str = ""):
32
+ full_prompt = f"{system}\n\n{prompt}" if system else prompt
33
+ for chunk in self.client.models.generate_content_stream(
34
+ model=self.model,
35
+ contents=full_prompt,
36
+ ):
37
+ if chunk.text:
38
+ yield chunk.text
39
+
40
+
41
+ # ── Groq provider (free tier, no daily quota issues) ──────────────────────────
42
+
43
+ class GroqLLM(BaseLLM):
44
+ """
45
+ Uses Groq's free API β€” no per-day quota, just a per-minute rate limit
46
+ that is far more generous than Gemini's free tier.
47
+
48
+ Sign up at https://console.groq.com and set GROQ_API_KEY in your .env.
49
+ Recommended model: llama-3.3-70b-versatile (free)
50
+ """
51
+
52
+ def __init__(self, api_key: str, model: str):
53
+ from groq import Groq
54
+ self.client = Groq(api_key=api_key)
55
+ self.model = model
56
+
57
+ def _messages(self, prompt: str, system: str) -> list[dict]:
58
+ messages = []
59
+ if system:
60
+ messages.append({"role": "system", "content": system})
61
+ messages.append({"role": "user", "content": prompt})
62
+ return messages
63
+
64
+ def generate(self, prompt: str, system: str = "") -> str:
65
+ response = self.client.chat.completions.create(
66
+ model=self.model,
67
+ messages=self._messages(prompt, system),
68
+ )
69
+ return response.choices[0].message.content
70
+
71
+ def stream(self, prompt: str, system: str = ""):
72
+ stream = self.client.chat.completions.create(
73
+ model=self.model,
74
+ messages=self._messages(prompt, system),
75
+ stream=True,
76
+ )
77
+ for chunk in stream:
78
+ delta = chunk.choices[0].delta.content
79
+ if delta:
80
+ yield delta
81
+
82
+
83
+ # ── Factory ───────────────────────────────────────────────────────────────────
84
+
85
+ def get_llm_client() -> BaseLLM:
86
+ settings = get_settings()
87
+
88
+ if settings.llm_provider == "gemini":
89
+ return GeminiLLM(settings.gemini_api_key, settings.llm_model)
90
+
91
+ if settings.llm_provider == "groq":
92
+ return GroqLLM(settings.groq_api_key, settings.groq_model)
93
+
94
+ raise ValueError(
95
+ f"Unsupported LLM_PROVIDER '{settings.llm_provider}'. "
96
+ "Valid options: 'gemini', 'groq'"
97
+ )
98
+
99
+
100
+ # ── Prompt templates ──────────────────────────────────────────────────────────
101
+
102
+ def build_prompt(query: str, context: str, task_type: str) -> str:
103
+ templates = {
104
+ "qa": (
105
+ "You are a codebase assistant. Use the context below to answer the question.\n\n"
106
+ "Context:\n{context}\n\nQuestion: {query}\n\nAnswer clearly, citing file names when relevant."
107
+ ),
108
+ "bug_finding": (
109
+ "Analyze the following code for potential bugs, security issues, or bad practices.\n\n"
110
+ "Code:\n{context}\n\n"
111
+ "Return a JSON list of objects with keys: line, issue, severity, suggestion."
112
+ ),
113
+ "docstring": (
114
+ "Write a clear, concise docstring for the following function. "
115
+ "Follow standard conventions for its language.\n\nFunction:\n{query}"
116
+ ),
117
+ "file_creation": (
118
+ "Generate complete, production-ready code for the following request.\n\n"
119
+ "Request: {query}\n\nRelevant existing code for context:\n{context}\n\n"
120
+ "Return only the code, no explanations."
121
+ ),
122
+ }
123
+ template = templates.get(task_type, templates["qa"])
124
+ return template.format(query=query, context=context)
125
+
126
+
127
+ def count_tokens(text: str) -> int:
128
+ import tiktoken
129
+ encoder = tiktoken.get_encoding("cl100k_base")
130
+ return len(encoder.encode(text))
mcp_client.py ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from contextlib import asynccontextmanager
2
+
3
+ from mcp.client.streamable_http import streamable_http_client
4
+ from mcp.shared._httpx_utils import create_mcp_http_client
5
+ from mcp import ClientSession
6
+ from config import get_settings
7
+
8
+
9
+ class RemoteMCPClient:
10
+ """
11
+ Async context-manager wrapper around an MCP streamable-HTTP session.
12
+
13
+ Usage (all work must happen inside the `async with` block so that every
14
+ anyio cancel scope is entered and exited in the *same* task):
15
+
16
+ async with RemoteMCPClient(url, token) as client:
17
+ tools = await client.list_tools()
18
+ """
19
+
20
+ def __init__(self, server_url: str, api_key: str):
21
+ self.server_url = server_url
22
+ self.api_key = api_key
23
+ self.session: ClientSession | None = None
24
+ self._exit_stack = None
25
+
26
+ # ── async context manager ─────────────────────────────────────────────────
27
+
28
+ async def __aenter__(self):
29
+ from contextlib import AsyncExitStack
30
+ self._exit_stack = AsyncExitStack()
31
+ await self._exit_stack.__aenter__()
32
+
33
+ headers = {"Authorization": f"Bearer {self.api_key}"}
34
+ http_client = create_mcp_http_client(headers=headers)
35
+ await self._exit_stack.enter_async_context(http_client)
36
+
37
+ # streamable_http_client yields (read_stream, write_stream) β€” 2 values
38
+ read_stream, write_stream = await self._exit_stack.enter_async_context(
39
+ streamable_http_client(self.server_url, http_client=http_client)
40
+ )
41
+
42
+ self.session = await self._exit_stack.enter_async_context(
43
+ ClientSession(read_stream, write_stream)
44
+ )
45
+ await self.session.initialize()
46
+ return self
47
+
48
+ async def __aexit__(self, *exc_info):
49
+ await self._exit_stack.__aexit__(*exc_info)
50
+ self.session = None
51
+
52
+ # ── public API ────────────────────────────────────────────────────────────
53
+
54
+ async def list_tools(self) -> list[dict]:
55
+ result = await self.session.list_tools()
56
+ return [{"name": t.name, "description": t.description} for t in result.tools]
57
+
58
+ async def call_tool(self, tool_name: str, params: dict) -> dict:
59
+ try:
60
+ result = await self.session.call_tool(tool_name, params)
61
+ return {"status": "success", "data": result.content}
62
+ except Exception as e:
63
+ return {"status": "error", "message": str(e)}
64
+
65
+
66
+ def get_github_mcp_client() -> RemoteMCPClient:
67
+ settings = get_settings()
68
+ return RemoteMCPClient(settings.github_mcp_url, settings.github_mcp_token)
requirements.txt ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Core LLM providers
2
+ google-genai==2.14.0
3
+ google-generativeai==0.8.6
4
+ groq
5
+
6
+ # LangChain + embeddings
7
+ langchain==1.3.14
8
+ langchain-chroma==1.1.0
9
+ langchain-community==0.4.2
10
+ langchain-core==1.5.2
11
+ langchain-google-genai==4.3.2
12
+ langchain-text-splitters==1.1.2
13
+
14
+ # Vector store
15
+ chromadb==1.5.9
16
+
17
+ # API / Web
18
+ fastapi==0.116.1
19
+ uvicorn==0.35.0
20
+ python-multipart==0.0.32
21
+
22
+ # MCP client
23
+ mcp==2.0.0
24
+ httpx==0.28.1
25
+ httpx-sse==0.4.3
26
+
27
+ # Config & utilities
28
+ pydantic==2.13.4
29
+ pydantic-settings==2.14.2
30
+ python-dotenv==1.2.2
31
+ tiktoken==0.13.0
32
+ rich==14.3.3
33
+ typer==0.25.1
34
+ requests==2.34.2
35
+
36
+ # Code analysis
37
+ radon==6.0.1
38
+ GitPython==3.1.45
test_github_mcp.py ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import asyncio
2
+ from mcp_client import get_github_mcp_client
3
+ from config import get_settings
4
+
5
+
6
+ async def main():
7
+ client = get_github_mcp_client()
8
+ await client.connect()
9
+
10
+ tools = await client.list_tools()
11
+ print(f"Connected. {len(tools)} tools available:\n")
12
+ for t in tools:
13
+ print(f"- {t['name']}: {t['description']}")
14
+
15
+ await client.disconnect()
16
+
17
+
18
+ asyncio.run(main())
test_runner.py ADDED
@@ -0,0 +1,164 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ Manual test runner for the Codebase Assistant.
3
+ Run with: .venv\Scripts\python.exe test_runner.py [test_name]
4
+
5
+ Available tests:
6
+ mcp - Test GitHub MCP connection & list tools
7
+ embed - Test embedding a small text snippet
8
+ ingest - Index this repo into the local vector store
9
+ query - Run a RAG query against the vector store
10
+ bugs - Detect bugs in a given file
11
+ complexity - Analyse cyclomatic complexity of a file
12
+ all - Run embed -> ingest -> query -> bugs -> complexity (no MCP)
13
+ """
14
+
15
+ import asyncio
16
+ import sys
17
+ import os
18
+
19
+ # ── helpers ──────────────────────────────────────────────────────────────────
20
+
21
+ def header(title: str):
22
+ print(f"\n{'='*60}")
23
+ print(f" {title}")
24
+ print(f"{'='*60}")
25
+
26
+ def ok(msg: str): print(f" OK {msg}")
27
+ def fail(msg: str): print(f" FAIL {msg}")
28
+
29
+ ROOT = os.path.dirname(__file__)
30
+
31
+ # ── individual tests ──────────────────────────────────────────────────────────
32
+
33
+ async def test_mcp():
34
+ """Connect to GitHub MCP and list available tools."""
35
+ header("Test: GitHub MCP Connection")
36
+ from mcp_client import get_github_mcp_client
37
+ client = get_github_mcp_client()
38
+ try:
39
+ await client.connect()
40
+ tools = await client.list_tools()
41
+ ok(f"Connected. {len(tools)} tools found:")
42
+ for t in tools[:5]:
43
+ print(f" - {t['name']}: {t['description'][:60]}")
44
+ if len(tools) > 5:
45
+ print(f" ... and {len(tools)-5} more")
46
+ except Exception as e:
47
+ fail(f"MCP connection failed: {e}")
48
+ finally:
49
+ await client.disconnect()
50
+
51
+
52
+ def test_embed():
53
+ """Embed a small text snippet and verify the vector shape."""
54
+ header("Test: Embedding")
55
+ from rag.embedding import embed_query, embed_document
56
+ try:
57
+ vec = embed_query("def hello(): pass")
58
+ assert isinstance(vec, list) and len(vec) > 0, "embed_query returned empty"
59
+ ok(f"embed_query -> vector of {len(vec)} dims")
60
+
61
+ fake_chunks = [{"content": "def add(a, b): return a+b"}]
62
+ chunks = embed_document(fake_chunks)
63
+ assert "embedding" in chunks[0], "embed_document did not add 'embedding' key"
64
+ ok(f"embed_document -> chunk embedding of {len(chunks[0]['embedding'])} dims")
65
+ except Exception as e:
66
+ fail(f"Embedding failed: {e}")
67
+
68
+
69
+ def test_ingest():
70
+ """Load and index this repository into the local ChromaDB."""
71
+ header("Test: Repository Ingest")
72
+ from rag.repository_loader import load_repository
73
+ from rag.splitter import split_code
74
+ from rag.embedding import embed_document
75
+ from rag.retriever import build_vector_store
76
+ try:
77
+ docs = load_repository(ROOT)
78
+ ok(f"Loaded {len(docs)} files from repo")
79
+
80
+ all_chunks = []
81
+ for doc in docs:
82
+ all_chunks.extend(split_code(doc))
83
+ ok(f"Split into {len(all_chunks)} chunks")
84
+
85
+ embedded = embed_document(all_chunks)
86
+ ok(f"Embedded {len(embedded)} chunks")
87
+
88
+ build_vector_store(embedded)
89
+ ok("Vector store built / updated (ChromaDB)")
90
+ except Exception as e:
91
+ fail(f"Ingest failed: {e}")
92
+
93
+
94
+ def test_query():
95
+ """Run a RAG query. Requires the vector store to be populated first."""
96
+ header("Test: RAG Query")
97
+ from rag.rag_chain import run_rag_query
98
+ try:
99
+ result = run_rag_query("How does the embedding module work?", k=3)
100
+ ok(f"Answer received ({len(result['answer'])} chars)")
101
+ ok(f"Sources returned: {len(result['sources'])}")
102
+ for s in result["sources"]:
103
+ print(f" - {s.get('file_path','?')} lines {s.get('start_line')}-{s.get('end_line')}")
104
+ print(f"\n Answer preview:\n {result['answer'][:300]}...")
105
+ except Exception as e:
106
+ fail(f"RAG query failed: {e}")
107
+
108
+
109
+ def test_bugs():
110
+ """Run bug detection on llm.py."""
111
+ header("Test: Bug Detection")
112
+ from services.code_analysis import detect_bugs
113
+ target = os.path.join(ROOT, "llm.py")
114
+ try:
115
+ bugs = detect_bugs(target)
116
+ ok(f"Bug detection returned {len(bugs)} item(s) for llm.py")
117
+ for b in bugs[:3]:
118
+ print(f" Line {b.get('line','?')}: [{b.get('severity','?')}] {b.get('issue','?')}")
119
+ except Exception as e:
120
+ fail(f"Bug detection failed: {e}")
121
+
122
+
123
+ def test_complexity():
124
+ """Analyse cyclomatic complexity of llm.py."""
125
+ header("Test: Complexity Analysis")
126
+ from services.code_analysis import analyze_complexity
127
+ target = os.path.join(ROOT, "llm.py")
128
+ try:
129
+ result = analyze_complexity(target)
130
+ ok(f"Complexity analysis returned {len(result['functions'])} function(s):")
131
+ for fn in result["functions"]:
132
+ print(f" {fn['name']}: complexity={fn['complexity']}, rank={fn['rank']}")
133
+ except Exception as e:
134
+ fail(f"Complexity analysis failed: {e}")
135
+
136
+
137
+ # ── entry point ───────────────────────────────────────────────────────────────
138
+
139
+ TESTS = {
140
+ "mcp": lambda: asyncio.run(test_mcp()),
141
+ "embed": test_embed,
142
+ "ingest": test_ingest,
143
+ "query": test_query,
144
+ "bugs": test_bugs,
145
+ "complexity": test_complexity,
146
+ }
147
+
148
+ def run_all_non_mcp():
149
+ """Run embed -> ingest -> query -> bugs -> complexity in sequence."""
150
+ test_embed()
151
+ test_ingest()
152
+ test_query()
153
+ test_bugs()
154
+ test_complexity()
155
+
156
+ if __name__ == "__main__":
157
+ arg = sys.argv[1] if len(sys.argv) > 1 else "all"
158
+ if arg == "all":
159
+ run_all_non_mcp()
160
+ elif arg in TESTS:
161
+ TESTS[arg]()
162
+ else:
163
+ print(f"Unknown test '{arg}'. Choose from: {', '.join(TESTS)} or 'all'")
164
+ sys.exit(1)