A newer version of the Streamlit SDK is available: 1.62.0
DRL Trading System - Project Architecture
Last Updated: March 12, 2026 System Version: Ultimate Agent (Whale-Fused PPO-LSTM)
π Executive Summary
This is an autonomous Deep Reinforcement Learning (DRL) Trading System that uses PPO-LSTM agents to trade cryptocurrency on Binance Testnet. The system features real-time whale wallet tracking, multi-timeframe analysis, advanced feature engineering (150+ features), and a Streamlit dashboard for monitoring.
Key Capabilities
- π§ PPO-LSTM Agent with 150+ advanced features (Wyckoff, SMC, whale patterns)
- π Whale Pattern Prediction using ML models trained on verified whale wallets
- π Multi-Asset Trading (BTC, ETH, SOL, XRP)
- π Self-Improvement Loop (fine-tunes on successful trades every 24 hours)
- π Real-time Dashboard with TradingView-style charts
- π‘οΈ Advanced Risk Management (circuit breaker, adaptive SL/TP, regime detection)
- β‘ Live Trading on Binance Testnet with dry-run mode
ποΈ System Architecture
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β USER INTERFACE β
β Streamlit Dashboard (app.py) + API Server (api_server.py) β
β - Real-time charts (TradingView-style) β
β - Whale analytics dashboard β
β - Bot status monitoring β
β - Market analysis cards (Whale, Funding, Order Flow) β
ββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββ
β
ββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββββ
β LIVE TRADING ORCHESTRATOR β
β live_trading_multi.py (MultiAssetTradingBot) β
β - Multi-threaded bot execution per asset β
β - Global portfolio management β
β - 30-min cooldown after losses β
β - 4-hour minimum hold time β
βββββββ¬ββββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββ
β β
βββββββΌββββββββββββββββββ ββββββββββΌβββββββββββββββββββββββ
β DRL BRAIN β β MARKET INTELLIGENCE β
β (PPO-LSTM Agent) β β β
β β β 1. Whale Tracker β
β β’ UltimateFeatureEng β β - Pattern Predictor β
β β’ VecNormalize β β - Wallet Collector β
β β’ 150+ features β β - Registry (ETH/SOL/XRP) β
β β’ Model inference β β β
β β β 2. Order Flow Analyzer β
β Models: β β - CVD, OI, funding rates β
β ββ ultimate_agent.zip β β β
β ββ vec_normalize.pkl β β 3. Multi-Timeframe Analyzer β
β β β - 4h, 1d, 1w timeframes β
β β β β
β β β 4. Regime Detector β
β β β - Trending/Ranging β
β β β - ADX/ATR-based β
β β β β
β β β 5. TFT Price Forecaster β
β β β - Neural price prediction β
β β β - Confidence scoring β
βββββββββββββββββββββββββ βββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββ
β DATA PIPELINE β
β β
β 1. Historical Data (Multi-Asset Fetcher) β
β ββ CCXT β Binance API β CSV cache β
β β
β 2. Whale Wallet Data β
β ββ ETH: 8 verified wallets (Binance, Bitfinex, etc.) β
β ββ SOL: 11 verified wallets β
β ββ XRP: 13 verified wallets β
β β
β 3. Alternative Data β
β ββ Fear & Greed Index β
β ββ BTC Dominance β
β β
β 4. Storage Layer β
β ββ SQLite (trading.db) - trades, state β
β ββ JSON - whale data, backtest reports β
β ββ CSV - historical OHLCV β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
π Directory Structure
drl-trading-system/
βββ .agent/ # Agent workflow definitions
β βββ workflows/
β βββ feature.md # Feature development workflow
β βββ fix.md # Bug fix workflow
β βββ train.md # Model training workflow
β βββ deploy.md # Deployment workflow
β
βββ config/
β βββ config.yaml # System configuration (exchange, risk, model params)
β
βββ data/ # All data storage
β βββ historical/ # Cached OHLCV data (CSV)
β βββ models/ # Trained DRL models
β β βββ ultimate_agent.zip # Main PPO model
β β βββ ultimate_agent_vec_normalize.pkl
β β βββ tft/ # TFT price forecaster models
β β βββ multi_asset/ # Asset-specific fine-tuned models
β βββ whale_wallets/ # Whale wallet transaction data
β β βββ eth/ # 8 verified ETH whale wallets
β β βββ sol/ # 11 verified SOL whale wallets
β β βββ xrp/ # 13 verified XRP whale wallets
β βββ alternative_cache/ # Fear/Greed, BTC dominance
β βββ checkpoints/ # Training checkpoints
β βββ backtest_report.json # Latest backtest results
β βββ trading.db # SQLite database (trades, positions, state)
β
βββ src/ # Source code
β βββ env/ # Gymnasium trading environments
β β βββ ultimate_env.py # Main env (150+ features)
β β βββ advanced_env.py # Advanced features env
β β βββ trading_env.py # Base trading env
β β βββ rewards.py # Reward functions (Sharpe/Sortino)
β β
β βββ brain/ # DRL agent & training
β β βββ agent.py # PPO-LSTM wrapper
β β βββ trainer.py # Training loops
β β βββ replay_buffer.py # Experience replay
β β
β βββ features/ # Feature engineering modules
β β βββ ultimate_features.py # 150+ feature engine (Wyckoff, SMC, etc.)
β β βββ whale_tracker.py # Real-time whale monitoring (46KB)
β β βββ whale_pattern_predictor.py # ML-based whale signal generator
β β βββ whale_wallet_collector.py # Scrapes whale wallet data
β β βββ whale_wallet_registry.py # Verified wallet addresses
β β βββ order_flow.py # CVD, OI, funding rate analysis (25KB)
β β βββ mtf_analyzer.py # Multi-timeframe analysis
β β βββ regime_detector.py # Market regime classification
β β βββ risk_manager.py # Adaptive risk management
β β βββ correlation_engine.py # Multi-asset correlation
β β βββ on_chain_whales.py # On-chain whale watcher
β β
β βββ models/ # ML models (non-DRL)
β β βββ whale_pattern_learner.py # Random Forest for whale patterns (30KB)
β β βββ price_forecaster.py # TFT (Temporal Fusion Transformer)
β β βββ confidence_engine.py # Signal confidence scoring
β β βββ regime_classifier.py # Regime classification model
β β βββ ensemble_orchestrator.py # Model ensemble coordination
β β
β βββ data/ # Data fetching & storage
β β βββ multi_asset_fetcher.py # CCXT data fetcher
β β βββ storage.py # Database abstraction layer
β β βββ candle_stream.py # Real-time candle streaming
β β βββ whale_stream.py # Real-time whale data streaming
β β
β βββ api/ # Exchange integration & execution
β β βββ binance.py # Binance API wrapper
β β βββ executor.py # Order execution engine
β β βββ risk_manager.py # Pre-execution risk checks
β β βββ portfolio_manager.py # Global portfolio coordination
β β
β βββ backtest/ # Backtesting engine
β β βββ engine.py # Main backtest executor
β β βββ data_loader.py # Historical data loader
β β
β βββ ui/ # User interface
β βββ app.py # Streamlit dashboard (2470 lines)
β βββ api_server.py # Flask API server (21KB)
β βββ charts.py # TradingView chart components
β βββ components.py # Reusable UI components
β
βββ logs/ # All logs
β βββ trading_log.json # Trade history
β βββ multi_asset_state.json # Bot state per asset
β βββ tensorboard/ # TensorBoard training logs
β βββ training/ # Training session logs
β
βββ Main Scripts # Top-level executable scripts
β βββ run.py # Legacy single-asset runner
β βββ live_trading_multi.py # Multi-asset live trading (68KB)
β βββ train_ultimate.py # Ultimate agent training (17KB)
β βββ train_whale_patterns.py # Train whale ML models
β βββ backtest_strategy.py # Strategy backtester (28KB)
β βββ launch_dashboard.sh # Start Streamlit UI
β βββ start.sh # Production startup script
β
βββ requirements.txt # Python dependencies
βββ Dockerfile # Docker deployment config
βββ .env # Environment variables (API keys)
βββ README.md # User-facing documentation
π§© Core Components Deep Dive
1. DRL Agent (PPO-LSTM)
Location: src/brain/agent.py, src/env/ultimate_env.py
- Algorithm: Proximal Policy Optimization (PPO)
- Architecture: LSTM policy network for temporal dependencies
- Observation Space: 153 dimensions (150 features + 3 position state variables)
- Action Space: Discrete(3) β [0: Hold, 1: Buy/Long, 2: Sell/Short]
- Reward Function: Sharpe/Sortino ratio-based (risk-adjusted returns)
Training Pipeline:
- Fetch historical data (1-year, 1h timeframe)
- Compute 150+ features via UltimateFeatureEngine
- Create Gymnasium env (UltimateTradingEnv)
- Wrap with VecNormalize for observation scaling
- Train PPO agent (500k-2M timesteps)
- Save model + VecNormalize stats
Key Training Files:
train_ultimate.py- Main training scripttrain_multi_asset.py- Multi-asset transfer learningtrain_whale_patterns.py- Whale pattern ML training
2. Feature Engineering (150+ Features)
Location: src/features/ultimate_features.py
Feature Categories:
Technical Indicators (40+ features)
- RSI, MACD, Bollinger Bands, ATR, ADX
- EMA crossovers (9/21, 50/200)
- Volume indicators (OBV, MFI)
Wyckoff Analysis (20+ features)
- Accumulation/Distribution phases
- Spring/Upthrust detection
- Volume spread analysis
Smart Money Concepts (SMC) (15+ features)
- Order blocks
- Fair value gaps
- Liquidity zones
Multi-Timeframe (30+ features)
- 4h, 1d, 1w trend alignment
- Higher timeframe support/resistance
Whale Patterns (20+ features)
- Exchange dump ratio
- Accumulator hoard ratio
- Flow velocity, momentum
- Wallet-specific hit rates
Market Regime (10+ features)
- Trending/Ranging classification
- Volatility regime
Correlation Features (15+ features)
- BTC dominance
- Multi-asset correlation matrix
- Fear & Greed Index
3. Whale Tracking System
Location: src/features/whale_tracker.py, src/models/whale_pattern_learner.py
Verified Whale Wallets:
- ETH (8 wallets): Binance, Bitfinex, Kraken hot/cold wallets
- SOL (11 wallets): Major exchange wallets + accumulators
- XRP (13 wallets): Ripple, exchanges, known whales
Data Collection:
whale_wallet_collector.pyscrapes blockchain APIs (Etherscan, Solscan, XRPScan)- Stores transaction history in JSON files (
data/whale_wallets/) - Updates every 1 hour (configurable)
Pattern Learning:
whale_pattern_learner.pytrains Random Forest models per chain- Features: flow velocity, exchange dump ratio, accumulator hoard ratio, wallet-specific patterns
- Predicts price impact (momentum signal: -1 to +1)
- Wallets weighted by historical hit rate (>55% = 2x weight)
Real-time Prediction:
whale_pattern_predictor.pyloads trained models- Fetches recent wallet data from cache
- Computes flow features
- Returns aggregated signal with confidence score
4. Live Trading Pipeline
Location: live_trading_multi.py
Flow:
1. Initialize MultiAssetTradingBot for each asset (BTC, ETH, SOL, XRP)
2. Load ultimate_agent.zip + vec_normalize.pkl
3. Initialize feature engines (UltimateFeatureEngine, WhaleTracker, etc.)
4. Loop every 5 minutes:
a. Fetch latest OHLCV data
b. Compute 150+ features
c. Normalize observation with VecNormalize
d. Get PPO model prediction (action 0/1/2)
e. Compute confidence score (TFT forecast + whale signals + regime)
f. Execute trade if confidence > 0.6
g. Manage position (trailing SL, TP, min hold time)
h. Update state to database
5. Self-improvement: Every 24h, fine-tune on high-reward trades
Risk Management:
- Circuit Breaker: Stops trading at 5% daily loss
- Position Sizing: 25% of balance per trade (was 50%)
- Stop Loss: 2.5% (adaptive based on regime)
- Take Profit: 5% (2:1 R:R ratio)
- Trailing Stop: 60% of max profit
- Min Hold Time: 4 hours
- Cooldown: 30 minutes after stop loss hit
5. Backtesting System
Location: backtest_strategy.py, src/backtest/engine.py
Process:
- Load historical data (default: 1 year)
- Replay EXACT live trading pipeline:
- Same feature computation
- Same VecNormalize scaling
- Same PPO model
- Same risk management rules
- Track all trades, equity curve, Sharpe ratio
- Save report to
data/backtest_report.json
Metrics:
- Total Return (%)
- Sharpe Ratio
- Sortino Ratio
- Max Drawdown (%)
- Win Rate (%)
- Average Trade Duration
- Total Trades
6. Streamlit Dashboard
Location: src/ui/app.py (2470 lines)
Pages:
Bot Status
- Current position, PnL, equity
- Model info (last trained, confidence)
- Start/Stop bot controls
Charts
- TradingView-style candlestick charts
- Buy/Sell signal markers
- Support/Resistance levels
- Volume bars
Whale Analytics
- Real-time whale flow signals
- Top wallets by accuracy
- Exchange dump ratio trends
Market Analysis
- Funding rates
- Order flow (CVD, OI)
- Multi-timeframe alignment
- Regime detection
Trade History
- All executed trades
- PnL breakdown
- Performance metrics
API Server: src/ui/api_server.py
- Flask REST API for real-time data
- 60-second caching for external APIs
- Endpoints:
/whale,/funding,/order_flow,/trades
π Data Flow
Training Flow
Historical Data (CSV)
β DataLoader
β UltimateTradingEnv
β VecNormalize
β PPO.learn()
β Save model + VecNormalize stats
Live Trading Flow
Binance API (real-time)
β MultiAssetDataFetcher
β Feature Engines (Whale, MTF, OrderFlow, etc.)
β UltimateFeatureEngine (150+ features)
β VecNormalize
β PPO.predict()
β Risk Manager
β Order Executor
β Database (trading.db)
Whale Tracking Flow
Blockchain APIs (Etherscan, Solscan, XRPScan)
β WhaleWalletCollector
β JSON cache (data/whale_wallets/)
β WhalePatternLearner (Random Forest)
β WhalePatternPredictor
β Signal aggregation
β Live Trading Bot
π οΈ Technology Stack
Core Frameworks
- DRL:
stable-baselines3(PPO),gymnasium(env) - ML:
scikit-learn(Random Forest),hmmlearn(regime detection) - Neural Networks:
torch(TFT forecaster) - UI:
streamlit,streamlit-lightweight-charts,plotly - API:
flask,flask-cors
Data & Exchange
- Exchange:
ccxt(Binance API) - Data Processing:
pandas,numpy - Database: SQLite (
sqlite3), JSON, CSV
Utilities
- Config:
pyyaml,python-dotenv - Testing:
pytest,pytest-asyncio - Deployment: Docker, Hugging Face Spaces
π Security & Configuration
Environment Variables (.env)
BINANCE_TESTNET_API_KEY=<key>
BINANCE_TESTNET_API_SECRET=<secret>
BINANCE_PROXY=<optional_proxy>
HF_TOKEN=<huggingface_token>
ETHERSCAN_API_KEY=<etherscan_key>
SOLSCAN_API_KEY=<solscan_key>
Risk Parameters (config/config.yaml)
- Max Daily Loss: 5%
- Max Drawdown: 20%
- Stop Loss: 2.5%
- Take Profit: 5%
- Position Size: 25%
π Model Performance
Ultimate Agent (Latest Training)
- Training Data: 1 year (2024-2025), 1h timeframe
- Assets: BTC, ETH, SOL, XRP
- Timesteps: 2M+ per asset
- Validation Sharpe: ~1.2-1.8 (asset-dependent)
- Backtest Win Rate: 55-65%
Whale Pattern Models
- ETH Whale Model: 62% hit rate (top wallets)
- SOL Whale Model: 58% hit rate
- XRP Whale Model: 60% hit rate
π Deployment
Hugging Face Spaces
- Space:
chen470/drl-trading-bot - Runtime: Docker container
- Auto-deploy: Triggered by
git push origin main - Build Time: ~90 seconds
- Logs: Via HF API +
api_server.log
Local Development
# Install dependencies
pip install -r requirements.txt
# Run backtest
python backtest_strategy.py
# Start dashboard
streamlit run src/ui/app.py
# Run live trading (dry-run)
python live_trading_multi.py --assets BTCUSDT ETHUSDT --dry-run
π Key Files Reference
Configuration
config/config.yaml- All system parameters.env- API keys & secretsrequirements.txt- Python dependencies
Training
train_ultimate.py- Main DRL trainingtrain_whale_patterns.py- Whale ML trainingtrain_multi_asset.py- Multi-asset transfer learning
Live Trading
live_trading_multi.py- Multi-asset trading bot (68KB)src/api/executor.py- Order executionsrc/api/portfolio_manager.py- Portfolio coordination
Backtesting
backtest_strategy.py- Strategy backtester (28KB)src/backtest/engine.py- Backtest engine
UI
src/ui/app.py- Streamlit dashboard (2470 lines)src/ui/api_server.py- Flask API server (21KB)
Feature Engineering
src/features/ultimate_features.py- Main feature enginesrc/features/whale_tracker.py- Whale tracking (46KB)src/features/whale_pattern_predictor.py- Whale ML predictor
π Known Issues & Limitations
- scikit-learn 1.7.2 pinned - Prevents pickle OOM crash on Hugging Face
- Proxy required for some APIs - Binance Futures API, Etherscan (rate limits)
- TFT forecaster optional - Falls back gracefully if model not trained
- Whale data collection blocking - Should be async/background cron
- VecNormalize dependency - Model predictions fail without proper normalization stats
π― Future Roadmap
- Async whale data collection - Background scraper instead of blocking API calls
- Advanced ensemble methods - Combine DRL + TFT + whale patterns more intelligently
- Real money trading - Migrate from Testnet to production (with proper safeguards)
- More chains - Add BTC-native whale tracking (currently uses ETH proxy)
- Improved UI - Add more charts, alerts, mobile responsiveness
- Paper trading mode - Simulate trades without Binance API
π Learning Resources
Understanding the System
- Start with
README.mdfor high-level overview - Read
.agent/workflows/*.mdfor development workflows - Study
config/config.yamlfor all parameters - Explore
src/env/ultimate_env.pyto understand the environment - Dive into
live_trading_multi.pyto see how it all connects
Key Concepts
- PPO (Proximal Policy Optimization): DRL algorithm that balances exploration/exploitation
- VecNormalize: Critical for normalizing observations to prevent gradient explosion
- Sharpe Ratio: Risk-adjusted return metric (reward function)
- Whale Tracking: Monitor large holders to predict price movements
- Wyckoff Analysis: Market phase detection (accumulation, distribution)
- Smart Money Concepts: Institutional order flow analysis
π€ Contributing
See .agent/workflows/ for development workflows:
feature.md- Adding new featuresfix.md- Bug fixestrain.md- Model trainingdeploy.md- Deployment process
π License
MIT License - See LICENSE file
Document Version: 1.0 Last Reviewed: March 12, 2026 Maintainer: DRL Trading System Team