Atlas / tasks.md
findEthics
feat: add comprehensive search optimization and ChromaDB caching system
4b28fb0
|
Raw
History Blame Contribute Delete
11.5 kB
# Atlas Intelligent Search Management - Implementation Tasks
## Project Overview
Optimize the web search functionality in Atlas to avoid unnecessary searches when conversation history already contains sufficient context to answer user questions.
**Current Problem:** Web search is performed regardless of conversation history, leading to increased latency, API costs, and poor user experience for follow-up questions.
**Goal:** Reduce unnecessary web searches by 40-60% while maintaining response quality.
---
## Phase 1: Rule-Based Context Analysis (Quick Win)
**Target Timeline:** 1-2 weeks
**Expected Impact:** 40% reduction in unnecessary searches
### Core Implementation
- [x] **Create search decision function**
- [x] Add `should_perform_search()` function in `app.py`
- [x] Implement pattern detection for follow-up questions
- [x] Add referential question detection logic
- [x] Create question type classification
- [x] **Define detection patterns**
- [x] Elaboration patterns: "elaborate", "explain more", "tell me more", "expand on"
- [x] Clarification patterns: "what do you mean", "can you clarify", "I don't understand"
- [x] Referential patterns: "this", "that", "it", "the previous", "above mentioned"
- [x] Continuation patterns: "and what about", "what else", "continue"
- [x] **Integrate with chat endpoint**
- [x] Modify chat endpoint (`app.py:479-487`) to call search decision function
- [x] Add conditional search logic before `search_web_combined()`
- [x] Preserve existing search behavior as fallback
- [x] Add logging for search decision tracking
### Configuration & Control
- [x] **Add configuration options**
- [x] Search sensitivity level (conservative/balanced/aggressive)
- [x] Pattern matching thresholds
- [x] Fallback behavior settings
- [x] Debug mode for search decisions
- [x] **Request model updates**
- [x] Add optional `force_search` parameter to `ChatRequest`
- [x] Add `search_decision_mode` parameter
- [ ] Update API documentation
### Testing & Validation
- [x] **Unit tests for search decision logic**
- [x] Test common follow-up question patterns
- [x] Test referential question detection
- [x] Test edge cases and false positives
- [x] Test with various conversation histories
- [x] **Integration testing**
- [x] Test full chat flow with search decisions
- [x] Validate response quality maintenance
- [x] Test fallback mechanisms
- [x] Performance impact measurement
### Analytics Integration
- [x] **Track search decision metrics**
- [x] Add search_decision_made field to message tracking
- [x] Track search_skipped_reason
- [ ] Update analytics dashboard with search optimization metrics
- [ ] Monitor false positive/negative rates
---
## Phase 2: AI-Based Search Decision Engine (Enhanced Intelligence)
**Target Timeline:** 2-3 weeks
**Expected Impact:** Additional 15-20% optimization + better edge case handling
### AI Decision Engine
- [x] **Implement AI-based search analyzer**
- [x] Create `analyze_search_necessity()` function
- [x] Design prompt template for search decision
- [x] Implement lightweight Gemini call for decision making
- [x] Add confidence scoring for search decisions
- [x] **Context analysis enhancement**
- [x] Analyze conversation history relevance
- [x] Implement topic continuity detection
- [x] Add semantic similarity analysis between current question and history
- [x] Create information sufficiency assessment
### Smart Query Classification
- [x] **Question type classification**
- [x] New information requests vs. clarifications
- [x] Factual questions vs. opinion/analysis requests
- [x] Time-sensitive vs. evergreen information needs
- [x] Broad topics vs. specific details
- [x] **Context sufficiency analysis**
- [x] Analyze if conversation history contains answer
- [x] Detect information gaps that require search
- [x] Assess recency requirements for information
- [x] Evaluate completeness of existing context
### Hybrid Decision Logic
- [x] **Combine rule-based and AI decisions**
- [x] Use rules for obvious cases (performance)
- [x] Use AI for ambiguous cases (accuracy)
- [x] Implement decision confidence thresholds
- [x] Add override mechanisms
- [x] **Fallback and error handling**
- [x] Handle AI decision timeouts
- [x] Implement graceful degradation to rule-based
- [x] Add decision audit logging
- [x] Create manual override capabilities
### Performance Optimization
- [x] **Optimize AI decision calls**
- [x] Implement decision result caching
- [x] Use minimal token prompts for decisions
- [x] Add async processing for decision analysis
- [x] Batch decision calls where possible
---
## Phase 3: Universal Search Caching with Vector Database (Resource Optimization)
**Target Timeline:** 3-4 weeks
**Expected Impact:** Additional 10-15% optimization + improved response times + cache persistence
### Phase 3a: Cache-First Architecture Optimization (COMPLETED)
- [x] **Universal cache implementation**
- [x] Replace session-based caches with single universal cache
- [x] Implement cache data structure with TTL and LRU eviction
- [x] Add semantic similarity matching using spaCy
- [x] Create cache analytics and monitoring endpoints
- [x] **Cache integration**
- [x] Integrate cache with search flow in chat endpoint
- [x] Add cache hit/miss tracking and statistics
- [x] Implement cache clearing and maintenance endpoints
- [x] Add cache information to API responses
### Phase 3b: Context-Aware Request Flow Optimization (COMPLETED)
- [x] **Implement conversation history detection**
- [x] Add logic to detect if request has meaningful conversation history
- [x] Handle edge cases (empty history, malformed history entries)
- [x] Create helper function for history validation (`has_meaningful_conversation_history`)
- [x] Add history detection to cache_info tracking
- [x] **Implement dual flow paths**
- [x] **First Message (No History)**: Cache-first approach
- [x] Check cache immediately after search term extraction
- [x] Skip search decision logic for performance
- [x] Perform web search only on cache miss
- [x] **Follow-up Messages (Has History)**: Search decision first
- [x] Run hybrid search decision analysis
- [x] Check cache only if search is determined necessary
- [x] Skip cache entirely if search not needed
- [x] **Enhanced request handling**
- [x] Preserve force_search override functionality
- [x] Add context-aware performance metrics to analytics (flow_type tracking)
- [x] Enhanced logging for monitoring flow paths
- [x] Test performance improvements for both scenarios
### Phase 3c: ChromaDB Vector Database Migration
- [ ] **ChromaDB integration setup**
- [ ] Add ChromaDB dependency to requirements.txt
- [ ] Design vector database schema for search cache
- [ ] Implement embedding function configuration (SentenceTransformer)
- [ ] Create persistent storage directory structure
- [ ] **Vector cache implementation**
- [ ] Replace hash-based cache with ChromaDB collection
- [ ] Implement unified semantic search (eliminates dual lookup paths)
- [ ] Add TTL filtering in vector queries
- [ ] Store search results as separate JSON files with metadata
- [ ] **Migration and testing**
- [ ] Create cache migration script from current to ChromaDB
- [ ] Implement performance benchmarking (O(log n) vs O(n))
- [ ] Test persistent cache across server restarts
- [ ] Validate semantic similarity improvements
### Phase 3d: System Design Documentation
- [x] **Architecture documentation**
- [x] Create current system design diagram
- [x] Create proposed system architecture diagram (with context-aware flow)
- [ ] Document performance characteristics and trade-offs
- [ ] Add vector database operational guide
### Advanced Features and Monitoring
- [ ] **Enhanced cache analytics**
- [ ] Popular queries tracking and visualization
- [ ] Cache effectiveness scoring and recommendations
- [ ] Memory usage optimization and reporting
- [ ] Cross-restart cache persistence validation
- [ ] **Performance optimization**
- [ ] Batch query operations for ChromaDB
- [ ] Cache warming strategies for popular queries
- [ ] Background cleanup and maintenance tasks
- [ ] Load testing and scalability validation
---
## Cross-Phase Implementation Tasks
### Code Quality & Maintenance
- [ ] **Documentation updates**
- [ ] Update API documentation with new parameters
- [ ] Add developer documentation for search decision logic
- [ ] Create troubleshooting guides
- [ ] Update README with optimization features
- [ ] **Code organization**
- [ ] Create separate module for search optimization (`search_optimizer.py`)
- [ ] Refactor search-related functions into dedicated module
- [ ] Add type hints and docstrings
- [ ] Implement proper error handling throughout
### Monitoring & Analytics
- [ ] **Enhanced analytics dashboard**
- [ ] Add search optimization metrics section
- [ ] Create search decision breakdown charts
- [ ] Add cache performance monitoring
- [ ] Implement A/B testing capabilities for optimization
- [ ] **Performance monitoring**
- [ ] Track response time improvements
- [ ] Monitor API cost reductions
- [ ] Add search decision accuracy metrics
- [ ] Create performance regression alerts
### Configuration Management
- [ ] **Environment configuration**
- [ ] Add optimization settings to environment variables
- [ ] Create configuration profiles (development/production)
- [ ] Implement runtime configuration updates
- [ ] Add feature flags for gradual rollout
### Deployment & Rollout
- [ ] **Gradual rollout strategy**
- [ ] Implement feature flags for each phase
- [ ] Create rollback mechanisms
- [ ] Add canary deployment support
- [ ] Plan staged user group rollouts
---
## Success Metrics
### Performance Metrics
- **Search Reduction:** Target 40-60% reduction in unnecessary searches
- **Response Time:** Improve average response time by 20-30% for follow-up questions
- **API Cost:** Reduce search API costs by 35-50%
- **User Experience:** Improve conversation flow satisfaction
### Quality Metrics
- **Response Accuracy:** Maintain >95% response quality
- **False Negatives:** Keep search-skipped-but-needed rate <5%
- **Cache Hit Rate:** Achieve >60% cache hit rate in Phase 3
- **User Satisfaction:** Maintain or improve user satisfaction scores
### Technical Metrics
- **Code Coverage:** Maintain >80% test coverage
- **Error Rate:** Keep optimization-related errors <1%
- **Performance Impact:** Add <50ms overhead for decision making
- **Memory Usage:** Keep cache memory usage <100MB per session
---
## Implementation Notes
### Development Priorities
1. **Start with Phase 1** for immediate impact and user feedback
2. **Validate thoroughly** before moving to next phase
3. **Monitor metrics continuously** during each phase
4. **Maintain backward compatibility** throughout implementation
### Risk Mitigation
- Implement comprehensive fallback mechanisms
- Add detailed logging for troubleshooting
- Create feature flags for quick rollback
- Plan gradual user rollout to minimize impact
### Future Enhancements
- Machine learning models for search decision optimization
- User behavior-based search prediction
- Advanced semantic analysis for context understanding
- Multi-language support for search optimization