| # Redis Caching Integration - Product Requirements Document |
|
|
| ## Executive Summary |
|
|
| This PRD outlines the integration of Redis caching into the CLI Proxy API to enable horizontal scaling, improve performance, and add distributed state management capabilities. |
|
|
| ## Background |
|
|
| ### Current State |
| The CLI Proxy API currently uses: |
| - **In-memory caching**: Signature cache, Codex prompt cache, OAuth sessions |
| - **File-based storage**: Authentication tokens, configuration |
| - **Single-instance limitations**: No shared state across multiple server instances |
| - **No persistence**: Usage statistics and cache data lost on restart |
|
|
| ### Problems |
| 1. **Cannot scale horizontally**: Multiple instances don't share state |
| 2. **Data loss on restart**: Usage statistics and cache data are ephemeral |
| 3. **No distributed rate limiting**: Each instance tracks limits independently |
| 4. **Session management issues**: OAuth sessions not shared across instances |
| 5. **Performance bottlenecks**: Repeated expensive operations (token counting, model lookups) |
|
|
| ## Goals |
|
|
| ### Primary Goals |
| 1. **Enable Horizontal Scaling**: Allow multiple server instances to share state via Redis |
| 2. **Improve Performance**: Cache expensive operations (API responses, token counts, model registry) |
| 3. **Add Persistence**: Preserve usage statistics and cache data across restarts |
| 4. **Distributed Rate Limiting**: Implement per-API-key and per-model rate limiting |
|
|
| ### Secondary Goals |
| 1. **Backward Compatibility**: Maintain existing file-based storage as fallback |
| 2. **Optional Redis**: System should work without Redis (degraded mode) |
| 3. **Monitoring**: Add Redis health checks and metrics |
| 4. **Configuration**: Simple Redis configuration via config.yaml |
|
|
| ## Target Users |
|
|
| 1. **Self-hosted users**: Running multiple instances for high availability |
| 2. **Enterprise users**: Need distributed rate limiting and usage tracking |
| 3. **Cloud deployments**: Kubernetes/Docker environments with multiple replicas |
| 4. **Development teams**: Need persistent cache during development |
|
|
| ## Requirements |
|
|
| ### Functional Requirements |
|
|
| #### FR-1: Redis Connection Management |
| - Support Redis standalone, Sentinel, and Cluster modes |
| - Connection pooling with configurable pool size |
| - Automatic reconnection on connection loss |
| - Health check endpoint for Redis status |
|
|
| #### FR-2: Distributed Caching Layer |
| - **Cache-Aside Pattern**: Application checks cache first, loads from source on miss |
| - **Write-Through Pattern**: Update cache and source simultaneously for critical data |
| - **TTL Management**: Configurable expiration for different cache types |
| - **Namespace Support**: Prefix keys by instance/environment |
|
|
| #### FR-3: Authentication Token Caching |
| - Cache OAuth tokens with automatic refresh |
| - Share tokens across instances |
| - Fallback to file-based storage if Redis unavailable |
| - Secure storage with encryption support |
|
|
| #### FR-4: Model Registry Caching |
| - Cache model availability across instances |
| - Real-time updates when models added/removed |
| - TTL: 5 minutes (configurable) |
| - Pub/Sub for instant registry updates |
|
|
| #### FR-5: Usage Statistics Persistence |
| - Persist aggregated statistics to Redis |
| - Support statistics export/import |
| - Hourly/daily aggregation with Redis sorted sets |
| - Retention policy (e.g., 30 days) |
|
|
| #### FR-6: Response Caching |
| - Cache identical API requests |
| - Configurable TTL per model/endpoint |
| - Cache key based on: model, messages hash, parameters |
| - Respect streaming vs non-streaming |
| - Optional feature (disabled by default) |
|
|
| #### FR-7: Rate Limiting |
| - Distributed rate limiting per API key |
| - Per-model rate limiting |
| - Sliding window algorithm using Redis sorted sets |
| - Configurable limits: requests/minute, requests/hour, requests/day |
| - Return proper 429 responses with Retry-After header |
|
|
| #### FR-8: OAuth Session Management |
| - Store OAuth sessions in Redis |
| - Share sessions across instances |
| - TTL: 10 minutes (existing behavior) |
| - Automatic cleanup of expired sessions |
|
|
| #### FR-9: Signature Cache (Thinking Blocks) |
| - Migrate existing in-memory signature cache to Redis |
| - TTL: 3 hours (existing behavior) |
| - Support for Claude thinking block signatures |
|
|
| #### FR-10: Codex Prompt Cache |
| - Migrate existing Codex cache to Redis |
| - TTL: 1 hour (existing behavior) |
| - Share prompt cache IDs across instances |
|
|
| ### Non-Functional Requirements |
|
|
| #### NFR-1: Performance |
| - Redis operations should add < 5ms latency |
| - Connection pool should handle 1000+ concurrent requests |
| - Cache hit rate > 80% for frequently accessed data |
|
|
| #### NFR-2: Reliability |
| - Graceful degradation when Redis unavailable |
| - Automatic failover in Sentinel/Cluster mode |
| - No data loss for critical operations |
|
|
| #### NFR-3: Security |
| - Support Redis AUTH password |
| - TLS/SSL connection support |
| - Encryption for sensitive cached data |
| - No credentials in cache keys |
|
|
| #### NFR-4: Scalability |
| - Support 10+ server instances |
| - Handle 10,000+ requests/second aggregate |
| - Redis memory usage < 2GB for typical workload |
|
|
| #### NFR-5: Maintainability |
| - Clear logging for Redis operations |
| - Metrics for cache hit/miss rates |
| - Configuration validation on startup |
| - Migration path from in-memory to Redis |
|
|
| ## Architecture |
|
|
| ### High-Level Design |
|
|
| ``` |
| ┌─────────────────────────────────────────────────────────┐ |
| │ Load Balancer │ |
| └─────────────────────────────────────────────────────────┘ |
| │ |
| ┌─────────────────┼─────────────────┐ |
| │ │ │ |
| ┌───────▼──────┐ ┌──────▼───────┐ ┌──────▼───────┐ |
| │ Instance 1 │ │ Instance 2 │ │ Instance 3 │ |
| │ │ │ │ │ │ |
| │ ┌──────────┐ │ │ ┌──────────┐ │ │ ┌──────────┐ │ |
| │ │ App Code │ │ │ │ App Code │ │ │ │ App Code │ │ |
| │ └────┬─────┘ │ │ └────┬─────┘ │ │ └────┬─────┘ │ |
| │ │ │ │ │ │ │ │ │ |
| │ ┌────▼─────┐ │ │ ┌────▼─────┐ │ │ ┌────▼─────┐ │ |
| │ │ Cache │ │ │ │ Cache │ │ │ │ Cache │ │ |
| │ │ Layer │ │ │ │ Layer │ │ │ │ Layer │ │ |
| │ └────┬─────┘ │ │ └────┬─────┘ │ │ └────┬─────┘ │ |
| └──────┼───────┘ └──────┼───────┘ └──────┼───────┘ |
| │ │ │ |
| └─────────────────┼─────────────────┘ |
| │ |
| ┌────────▼────────┐ |
| │ Redis Cluster │ |
| │ │ |
| │ ┌───────────┐ │ |
| │ │ Cache │ │ |
| │ │ Storage │ │ |
| │ └───────────┘ │ |
| │ ┌───────────┐ │ |
| │ │ Pub/Sub │ │ |
| │ └───────────┘ │ |
| └─────────────────┘ |
| ``` |
|
|
| ### Cache Layer Structure |
|
|
| ```go |
| // Proposed package structure |
| internal/cache/ |
| ├── redis/ |
| │ ├── client.go // Redis client wrapper |
| │ ├── config.go // Configuration |
| │ ├── health.go // Health checks |
| │ └── pool.go // Connection pooling |
| ├── store/ |
| │ ├── interface.go // Cache store interface |
| │ ├── redis_store.go // Redis implementation |
| │ ├── memory_store.go // In-memory fallback |
| │ └── hybrid_store.go // Redis + memory hybrid |
| ├── auth_cache.go // Authentication caching |
| ├── model_cache.go // Model registry caching |
| ├── response_cache.go // Response caching |
| ├── rate_limiter.go // Distributed rate limiting |
| ├── session_cache.go // OAuth session caching |
| ├── signature_cache.go // Thinking signature caching (migrate existing) |
| └── usage_cache.go // Usage statistics caching |
| ``` |
|
|
| ### Configuration Schema |
|
|
| ```yaml |
| # config.yaml additions |
| redis: |
| # Enable Redis caching (default: false) |
| enabled: true |
| |
| # Redis connection |
| host: "localhost" |
| port: 6379 |
| password: "" |
| db: 0 |
| |
| # TLS configuration |
| tls: |
| enabled: false |
| cert: "" |
| key: "" |
| ca: "" |
| |
| # Connection pool |
| pool: |
| max-idle: 10 |
| max-active: 100 |
| idle-timeout: 300s |
| max-conn-lifetime: 3600s |
| |
| # Cluster/Sentinel mode |
| mode: "standalone" # standalone, sentinel, cluster |
| sentinel: |
| master-name: "mymaster" |
| addresses: |
| - "sentinel1:26379" |
| - "sentinel2:26379" |
| cluster: |
| addresses: |
| - "node1:6379" |
| - "node2:6379" |
| - "node3:6379" |
| |
| # Key prefix for namespacing |
| key-prefix: "cliproxy:" |
| |
| # Feature-specific settings |
| caching: |
| # Authentication tokens |
| auth-tokens: |
| enabled: true |
| ttl: 3600s |
| |
| # Model registry |
| model-registry: |
| enabled: true |
| ttl: 300s |
| |
| # Response caching |
| responses: |
| enabled: false # Disabled by default |
| ttl: 60s |
| max-size: 1000 # Max cached responses |
| |
| # OAuth sessions |
| oauth-sessions: |
| enabled: true |
| ttl: 600s |
| |
| # Signature cache (thinking blocks) |
| signatures: |
| enabled: true |
| ttl: 10800s # 3 hours |
| |
| # Codex prompt cache |
| codex-cache: |
| enabled: true |
| ttl: 3600s |
| |
| # Usage statistics |
| usage-stats: |
| enabled: true |
| retention-days: 30 |
| |
| # Rate limiting |
| rate-limiting: |
| enabled: false # Disabled by default |
| rules: |
| - api-key: "default" |
| requests-per-minute: 60 |
| requests-per-hour: 1000 |
| requests-per-day: 10000 |
| - model: "claude-opus-4" |
| requests-per-minute: 10 |
| requests-per-hour: 100 |
| ``` |
|
|
| ## Cache Strategies by Component |
|
|
| ### 1. Authentication Tokens (Write-Through) |
| - **Pattern**: Write-Through + Cache-Aside |
| - **TTL**: 1 hour (or token expiry) |
| - **Key**: `cliproxy:auth:{provider}:{auth_id}` |
| - **Fallback**: File-based storage |
| - **Invalidation**: On token refresh or logout |
|
|
| ### 2. Model Registry (Cache-Aside + Pub/Sub) |
| - **Pattern**: Cache-Aside with Pub/Sub updates |
| - **TTL**: 5 minutes |
| - **Key**: `cliproxy:models:{provider}:{client_id}` |
| - **Pub/Sub**: `cliproxy:events:models` |
| - **Invalidation**: On model registration/unregistration |
|
|
| ### 3. Response Caching (Cache-Aside) |
| - **Pattern**: Cache-Aside (optional feature) |
| - **TTL**: 60 seconds (configurable) |
| - **Key**: `cliproxy:response:{hash(model+messages+params)}` |
| - **Conditions**: Only cache successful, non-streaming responses |
| - **Invalidation**: TTL expiry only |
|
|
| ### 4. Rate Limiting (Sliding Window) |
| - **Pattern**: Sliding Window with Sorted Sets |
| - **Key**: `cliproxy:ratelimit:{api_key}:{window}` |
| - **Windows**: minute, hour, day |
| - **Algorithm**: ZREMRANGEBYSCORE + ZADD + ZCARD |
| - **Cleanup**: Automatic via TTL |
|
|
| ### 5. OAuth Sessions (Cache-Aside) |
| - **Pattern**: Cache-Aside |
| - **TTL**: 10 minutes |
| - **Key**: `cliproxy:oauth:session:{state}` |
| - **Invalidation**: On completion or TTL expiry |
|
|
| ### 6. Usage Statistics (Write-Behind) |
| - **Pattern**: Write-Behind (async aggregation) |
| - **Key**: `cliproxy:usage:{api_key}:{date}:{hour}` |
| - **Storage**: Sorted Sets for time-series data |
| - **Aggregation**: Hourly background job |
| - **Retention**: 30 days |
|
|
| ## Implementation Phases |
|
|
| ### Phase 1: Foundation (Week 1) |
| - Redis client wrapper and connection management |
| - Configuration schema and validation |
| - Health check endpoints |
| - Basic cache interface and Redis store implementation |
|
|
| ### Phase 2: Core Caching (Week 2) |
| - Migrate signature cache to Redis |
| - Migrate Codex cache to Redis |
| - OAuth session caching |
| - Model registry caching with Pub/Sub |
|
|
| ### Phase 3: Advanced Features (Week 3) |
| - Authentication token caching with fallback |
| - Usage statistics persistence |
| - Response caching (optional) |
| - Monitoring and metrics |
|
|
| ### Phase 4: Rate Limiting (Week 4) |
| - Distributed rate limiting implementation |
| - Per-API-key and per-model limits |
| - Admin API for rate limit management |
| - Testing and optimization |
|
|
| ## Success Metrics |
|
|
| ### Performance Metrics |
| - Cache hit rate > 80% for model registry |
| - Cache hit rate > 60% for responses (if enabled) |
| - Redis operation latency < 5ms (p95) |
| - Overall request latency increase < 10ms |
|
|
| ### Reliability Metrics |
| - Zero data loss for authentication tokens |
| - 99.9% uptime with Redis failover |
| - Graceful degradation when Redis unavailable |
|
|
| ### Scalability Metrics |
| - Support 10+ instances without performance degradation |
| - Handle 10,000+ requests/second aggregate |
| - Redis memory usage < 2GB for typical workload |
|
|
| ## Risks and Mitigations |
|
|
| ### Risk 1: Redis Unavailability |
| - **Impact**: High - System cannot share state |
| - **Mitigation**: Fallback to in-memory/file-based storage, health checks, automatic reconnection |
|
|
| ### Risk 2: Cache Stampede |
| - **Impact**: Medium - Multiple instances fetch same data simultaneously |
| - **Mitigation**: Use locking mechanisms, staggered TTLs, cache warming |
|
|
| ### Risk 3: Memory Exhaustion |
| - **Impact**: High - Redis runs out of memory |
| - **Mitigation**: Eviction policies (LRU), memory limits, monitoring, TTL enforcement |
|
|
| ### Risk 4: Data Inconsistency |
| - **Impact**: Medium - Cache and source out of sync |
| - **Mitigation**: Proper invalidation strategies, short TTLs for critical data, Pub/Sub updates |
|
|
| ### Risk 5: Security Concerns |
| - **Impact**: High - Sensitive data exposed in cache |
| - **Mitigation**: Encryption, Redis AUTH, TLS, no credentials in keys, secure key generation |
|
|
| ## Dependencies |
|
|
| ### External Dependencies |
| - **Redis**: v6.0+ (for ACL support) or v7.0+ (recommended) |
| - **Go Redis Client**: github.com/redis/go-redis/v9 |
| - **Existing**: All current dependencies remain |
|
|
| ### Internal Dependencies |
| - Modify: `internal/cache/signature_cache.go` |
| - Modify: `internal/runtime/executor/cache_helpers.go` |
| - Modify: `internal/api/handlers/management/oauth_sessions.go` |
| - Modify: `internal/usage/logger_plugin.go` |
| - Modify: `sdk/cliproxy/auth/conductor.go` |
| - Modify: `internal/config/config.go` |
|
|
| ## Open Questions |
|
|
| 1. Should response caching be enabled by default or opt-in? |
| 2. What should be the default rate limits per API key? |
| 3. Should we support Redis Cluster from day one or add later? |
| 4. How to handle cache warming on startup? |
| 5. Should we add a cache admin API for manual invalidation? |
|
|
| ## Appendix |
|
|
| ### A. Redis Key Naming Convention |
| ``` |
| cliproxy:{component}:{identifier}:{sub-key} |
| |
| Examples: |
| - cliproxy:auth:claude:abc123 |
| - cliproxy:models:gemini:client-1 |
| - cliproxy:response:hash-xyz |
| - cliproxy:ratelimit:api-key-1:minute |
| - cliproxy:oauth:session:state-xyz |
| - cliproxy:usage:api-1:2026-02-10:14 |
| ``` |
|
|
| ### B. Cache TTL Guidelines |
| - **Authentication tokens**: 1 hour or token expiry |
| - **Model registry**: 5 minutes |
| - **Responses**: 60 seconds (if enabled) |
| - **OAuth sessions**: 10 minutes |
| - **Signatures**: 3 hours |
| - **Codex cache**: 1 hour |
| - **Rate limit windows**: 1 day (for daily limits) |
| - **Usage stats**: 30 days |
|
|
| ### C. Redis Memory Estimation |
| For 1000 concurrent users: |
| - Auth tokens: ~1MB (1KB per token) |
| - Model registry: ~500KB |
| - OAuth sessions: ~100KB |
| - Signatures: ~5MB |
| - Responses (1000 cached): ~50MB |
| - Rate limiting: ~10MB |
| - Usage stats (30 days): ~100MB |
| - **Total**: ~170MB + overhead = ~250MB |
|
|
| For 10,000 concurrent users: ~2.5GB |
|
|