Spaces:
Sleeping
Sleeping
File size: 7,899 Bytes
3786a3f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 | # GraphRAG Knowledge AI β Implementation Plan
## Project Overview
- **Assignment #10:** Neo4j graph DB, entity/relationship extraction, Cypher traversal, hybrid search, community detection, multi-hop reasoning, evaluation dashboard
- **Stack:** Django 4.2 + Neo4j + ChromaDB + LangChain LLM (Groq/Gemini)
- **118 unit tests** β all PASS
- **68 manual endpoint tests** β 62 pass, 6 failures explained
- **All frontend files exist but are 0 bytes** β frontend will be built separately with Next.js
## Strategy β Phase-Based Approach
### Phase 1: Backend Fixes (Current Phase)
**Goal:** Make backend bulletproof β handle all worst-case scenarios
**Approach:** All 10 fixes at once, then full testing
**Database:** SQLite with WAL mode (good enough for assignment)
**Background:** Simple polling (raw threads + DB progress fields)
### Phase 2: Deploy to HuggingFace
**Goal:** Get backend running on HuggingFace Spaces
**Changes:** Switch to PostgreSQL, add proper CORS, configure for HF deployment
### Phase 3: Frontend Build
**Goal:** Complete professional frontend with real-time processing status
**Framework:** Next.js (as mentioned in assignment)
**Features:** Real-time progress polling, graph visualization, multi-hop reasoning UI
### Phase 4: Full Celery + Redis Migration
**Goal:** Production-grade background task handling
**When:** Only after Phase 1-3 work perfectly end-to-end
**Changes:** Add Redis, Celery worker, Celery beat, proper task queue
---
## Phase 1: Backend Fixes β Implementation Plan
### Fix List (All 10 β Implement Together)
| # | Fix | File(s) | Time | Status |
|---|-----|---------|------|--------|
| 1 | Startup recovery β reset stuck PROCESSING docs | `graphrag/apps.py` (create) | 30 min | PENDING |
| 2 | Gunicorn timeout 120β300s | `Dockerfile:45` | 1 min | PENDING |
| 3 | LLM extractors raise on failure (not return []) | `entity_extractor.py`, `relationship_extractor.py` | 1 hour | PENDING |
| 4 | Check API key before starting ingestion thread | `views.py` (trigger_ingestion_background) | 30 min | PENDING |
| 5 | Add tenacity retry with exponential backoff | `entity_extractor.py`, `relationship_extractor.py`, `neo4j_client.py` | 2 hours | PENDING |
| 6 | UUID-based ChromaDB IDs | `vector_retriever.py:61` | 30 min | PENDING |
| 7 | Entity resolver LLM disambiguation | `entity_resolver.py` | 1 hour | PENDING |
| 8 | Login empty body returns 400 not 401 | `views.py:140-145` | 15 min | PENDING |
| 9 | Empty relationship type fallback to RELATED_TO | `neo4j_client.py:114,119` | 5 min | PENDING |
| 10 | Community endpoints graceful degradation | `community_detector.py:38` | 15 min | PENDING |
### Progress Tracking (Simple Polling)
**Add to Document model:**
- `processing_progress` (IntegerField, 0-100)
- `processing_step` (CharField, max 200 chars)
**Update graph_builder.py** to set progress at each step:
1. Parsing document β 5%
2. Indexing vectors β 20%
3. Extracting entities β 20-80% (per section)
4. Extracting relationships β 20-80% (per section)
5. Resolving duplicates β 85%
6. Building knowledge graph β 95%
7. Complete β 100%
**Frontend polling:** `GET /api/documents/{id}/` every 3 seconds
### Implementation Order
**Batch 1 β Quick Wins (30 min total):**
1. Fix #2 β Gunicorn timeout (1 min)
2. Fix #9 β Empty rel type fallback (5 min)
3. Fix #8 β Login 400 (15 min)
4. Fix #10 β Community graceful degradation (15 min)
**Batch 2 β Medium Complexity (1.5 hours):**
5. Fix #6 β UUID ChromaDB IDs (30 min)
6. Fix #1 β Startup recovery apps.py (30 min)
7. Fix #4 β API key check (30 min)
**Batch 3 β Complex (4 hours):**
8. Fix #3 β Extractors raise instead of swallow (1 hour)
9. Fix #5 β Tenacity retry (2 hours)
10. Fix #7 β Entity resolver LLM disambiguation (1 hour)
**Batch 4 β Progress Tracking (1.5 hours):**
11. Add progress fields to Document model (30 min)
12. Update graph_builder.py with progress steps (1 hour)
**Batch 5 β Testing (2 hours):**
13. Run all 118 existing tests (5 min)
14. Write new tests for all fixes (2 hours)
**Total: ~9.5 hours**
### Testing Strategy
**After each batch:**
1. Run `python manage.py test graphrag.tests graphrag.tests_comprehensive` (all 118 tests)
2. Run manual endpoint tests (verify nothing regressed)
3. Test worst-case scenarios:
- Kill server mid-ingestion β restart β verify doc reset to FAILED
- No API keys β upload β verify immediate FAILED
- Empty login body β verify 400 response
**After all batches:**
1. Full 68-test manual verification
2. New tests for each fix
3. Load test with concurrent uploads (if possible)
---
## Worst-Case Scenarios Handled
| Scenario | Current Behavior | After Fix | Fix # |
|----------|-----------------|-----------|-------|
| Server restarts mid-ingestion | Document stuck PROCESSING forever | Auto-reset to FAILED on startup | #1 |
| Gunicorn kills worker at 120s | Ingestion thread dies | Timeout increased to 300s | #2 |
| LLM API returns 500 | Empty entities, document "completed" | Exception raised, document FAILED | #3 |
| LLM API returns 429 (rate limit) | Section skipped silently | Retry 3 times with backoff | #5 |
| No API keys configured | Document stuck PENDING | Immediate FAILED with error | #4 |
| Same filename uploaded twice | ChromaDB vectors overwritten | UUID-based IDs prevent overwrite | #6 |
| Neo4j goes down mid-ingestion | Partial graph data | Cleanup on failure + FAILED status | #3 |
| Neo4j down during query | 500 error | Graceful empty response | #10 |
| LLM returns garbage relationship type | Invalid Cypher | Fallback to RELATED_TO | #9 |
| "J. Smith" vs "John Smith" | Not merged (borderline score) | LLM disambiguation | #7 |
| Frontend refreshes during processing | Shows stale status | Polling every 3s gets new progress | New |
---
## File Change List
| File | Change | Lines Changed |
|------|--------|---------------|
| `graphrag/apps.py` | **CREATE** β Startup recovery | ~25 lines |
| `graphrag/models.py` | Add `processing_progress`, `processing_step` fields | ~5 lines |
| `graphrag/views.py` | Fix #4 (API key check), Fix #8 (login 400) | ~15 lines |
| `graphrag/services/entity_extractor.py` | Fix #3 (raise), Fix #5 (tenacity retry) | ~15 lines |
| `graphrag/services/relationship_extractor.py` | Fix #3 (raise), Fix #5 (tenacity retry) | ~15 lines |
| `graphrag/services/graph_builder.py` | Add progress tracking at each step | ~30 lines |
| `graphrag/services/neo4j_client.py` | Fix #9 (empty rel type), Fix #5 (tenacity) | ~10 lines |
| `graphrag/services/entity_resolver.py` | Fix #7 (LLM disambiguation) | ~30 lines |
| `graphrag/services/community_detector.py` | Fix #10 (graceful degradation) | ~5 lines |
| `graphrag/services/vector_retriever.py` | Fix #6 (UUID-based IDs) | ~3 lines |
| `graphrag/serializers.py` | Add new fields to DocumentSerializer | ~3 lines |
| `Dockerfile` | Fix #2 (timeout 300s) | ~1 line |
**Total: 12 files, ~160 lines of changes**
---
## Commands to Run
**Run all existing tests:**
```bash
cd /home/creator/Desktop/ExcellenceTechnology/07.graphrag-knowledge-ai/backend
python manage.py test graphrag.tests graphrag.tests_comprehensive
```
**Run manual endpoint tests:**
```bash
cd /home/creator/Desktop/ExcellenceTechnology/07.graphrag-knowledge-ai
python manual_test.py
```
**Check Docker services:**
```bash
docker compose ps
docker compose logs backend --tail=50
```
---
## Notes
- `tenacity` is already in requirements.txt but never imported/used
- No Redis/Celery yet β will add in Phase 4
- Frontend is 0 bytes across all 14 pages β will build with Next.js in Phase 3
- Deploy to HuggingFace in Phase 2 (after Phase 1 fixes work perfectly)
- Always test fixes against real Neo4j and ChromaDB when possible
- Keep all 118 existing tests passing β never break them
- Document every fix with before/after code in commit messages
|