DockerSpace / HANDOFF.md
DennisChan0909's picture
feat: integrate local architecture with HF Space
e610a2f
|
Raw
History Blame Contribute Delete
9.59 kB

Stock Predictor β€” Codex Handoff

Project in one line

Taiwan stock ML predictor (FastAPI + React/Vite). Runs locally on a Mac mini, also deployed on HuggingFace Space. Mac computes predictions; HF Space reads them from a shared HF Dataset queue.


Repo layout

stock-predictor/
β”œβ”€β”€ main.py                     # FastAPI app entry, lifespan warmup
β”œβ”€β”€ routers/
β”‚   β”œβ”€β”€ stock.py                # /api/stock/* endpoints
β”‚   └── line_webhook.py         # LINE Bot webhook
β”œβ”€β”€ services/
β”‚   β”œβ”€β”€ predictor_service.py    # get_prediction(), caches, quick estimate
β”‚   β”œβ”€β”€ hf_queue_service.py     # HF Dataset queue read/write
β”‚   β”œβ”€β”€ backtest_service.py     # walk-forward backtest
β”‚   └── news_service.py         # Gemini news overlay
β”œβ”€β”€ models/
β”‚   └── predictor.py            # StockPredictor, FEATURE_COLUMNS, _build_features
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ fetcher.py              # yfinance + TWSE history
β”‚   β”œβ”€β”€ institutional_flow.py   # 三倧法人
β”‚   β”œβ”€β”€ margin_flow.py          # θžθ³‡θžεˆΈ
β”‚   β”œβ”€β”€ monthly_revenue.py      # ζœˆη‡Ÿζ”Ά
β”‚   β”œβ”€β”€ ptt_sentiment.py        # PTT [ζ¨™ηš„] sentiment (currently returns zeros)
β”‚   └── ...
β”œβ”€β”€ indicators/
β”‚   └── technical.py            # add_all_indicators(), add_cross_asset_tw()
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ hf_queue_worker.py      # Mac launchd job: poll queue, push predictions
β”‚   β”œβ”€β”€ backtest_3way.py        # baseline backtest runner
β”‚   β”œβ”€β”€ backtest_c*.py          # per-experiment backtest scripts (C1–C12)
β”‚   └── improvement_harness.py  # runs backtest on 5 stocks, checks thresholds
β”œβ”€β”€ frontend/
β”‚   └── src/
β”‚       └── components/         # React components (StockPage, RecommendPage, etc.)
β”œβ”€β”€ static/                     # built frontend (served by FastAPI)
└── docs/
    └── improvements_log.md     # history of model experiments

Deployment

Mac local (port 7861)

  • Managed by launchd: com.dennis.stock-predictor
  • Restart: launchctl stop com.dennis.stock-predictor && launchctl start com.dennis.stock-predictor
  • Logs: logs/server.log

HF Space

  • URL: https://dennischan0909-dockerspace.hf.space/
  • Rebuilds automatically on push to main branch
  • HF Dataset queue: DennisChan0909/stock-predictor-queue (stores predictions as JSON)
  • Current runbook: docs/huggingface_integration.md
  • HF is the lightweight serving layer. Mac remains the durable ML/precompute worker; HF returns quick_rules while waiting for Mac queue results.

Mac worker (launchd, every 5 min)

  • scripts/hf_queue_worker.py β€” computes predictions for WATCHLIST (20 stocks) + any pending queue items, pushes to HF Dataset
  • Commands:
    • ./venv/bin/python scripts/hf_queue_worker.py status β€” show queue
    • ./venv/bin/python scripts/hf_queue_worker.py push 2330 0050 β€” force-push specific stocks

Frontend build + deploy

cd frontend && npm run build
cp frontend/dist/index.html static/index.html
cp frontend/dist/assets/* static/assets/
git add -u && git commit -m "..." && git push origin main

HF Space prediction flow (as of 2026-05-12)

  1. /api/stock/<code>/predict called
  2. get_prediction() checks in-memory pred cache (30 min TTL on HF)
  3. Cache miss: checks HF Dataset queue synchronously (get_cached_result)
    • Queue hit β†’ return Mac ML result immediately, cache it for 30 min
    • Queue miss β†’ call enqueue() + fall through to quick estimate
  4. Quick estimate: rule-based signal from technical indicators (<100ms), NOT cached on HF Space
  5. Background ML training is skipped on HF Space (Mac worker handles it)
  6. Startup warmup (_warmup_from_queue in main.py) pre-loads all "done" queue results before accepting requests

Model

Stack

RF + XGB + LGBM ensemble (CatBoost compiled but optional). 29 features after SHAP pruning.

Labels

Triple barrier (C3): upper/lower barrier = close Γ— (1 Β± vol_20d Γ— 1.0), vertical = 5 trading days. Label = first barrier hit.

Current baseline (5-stock average, 24-month walk-forward)

  • dir_accuracy β‰ˆ 42%
  • up_precision β‰ˆ 53%
  • FEATURE_COLUMNS: 29 features in models/predictor.py:153

Meta-label filter (C4)

meta_filtered field in prediction output. Secondary RF trained on (features + buy_prob β†’ correct?). Threshold 0.50.


Experiments history (docs/improvements_log.md)

ID What Result
C1 fracdiff price features FAILED
C3 Triple barrier labels PASSED β€” dir_acc +9.8pp
C4 Meta-label filter PASSED β€” up_prec +4.1pp
C5 SHAP pruning (removed 11 low-signal features) PASSED (+1pp), now in FEATURE_COLUMNS
C6 HMM regime detection FAILED (<1.5pp)
C2 PTT sentiment (ptt_sentiment_1d, _5d_ma) FAILED β€” data pipeline returns 0 rows for all stocks
C7 Calendar features + 36m lookback FAILED β€” up_prec dropped -5.8pp
C8 Optuna RF tuning FAILED β€” params already near-optimal
C9 Multi-stock training universe FAILED
C10 Asymmetric triple barrier grid FAILED
C11 Full ensemble walk-forward FAILED
C12 1-day label horizon FAILED catastrophically

All C-experiments are self-contained in scripts/backtest_c*.py. New experiments should follow the same pattern: run scripts/improvement_harness.py to validate before touching models/predictor.py.


Known issues / pending work

Hermes / agent-team integration audit (2026-05-21)

Hermes added standalone agent-team utilities: Antigravity notes, W&B logging, data quality validation, and Telegram notification wrappers. These are intended to stay outside the core predictor path unless separately promoted.

Hermes hierarchy entry point: docs/hermes/README.md. Read docs/hermes/rules.md before changing Hermes utilities; it records that generic Hermes helper cleanup must not modify predictor/recommendation/precompute logic.

Codex audit fixes applied:

  • scripts/telegram_bot_wrapper.py: fixed .env.telegram parsing, and made CLI arguments mode-specific so --update, --validation, and experiment modes can run without a dummy --message.
  • scripts/test_agent_integration.py: added missing argparse import and avoided importing Python's built-in antigravity easter-egg module. The check now uses the CLI/config path instead.
  • .gitignore: ignore local .env.telegram and .env.wandb files; keep .env.telegram.example as the shareable template.

Verification:

./venv/bin/python -m py_compile scripts/telegram_bot_wrapper.py scripts/send_experiment_notification.py scripts/data_quality_validator.py scripts/test_agent_integration.py scripts/wandb_experiment_tracker.py
./venv/bin/python scripts/telegram_bot_wrapper.py --message "Hermes side-effect smoke" --success
./venv/bin/python scripts/telegram_bot_wrapper.py --update --step 1 --total 2
./venv/bin/python scripts/data_quality_validator.py --check features
./venv/bin/python scripts/test_agent_integration.py --quiet

Residual risk: hermes_tools is not importable in the normal venv, so Telegram wrapper uses fallback console mode unless Hermes injects that module at runtime. Do not wire these utilities into production precompute or recommendation publishing until a real Telegram send and no-lookahead validation gate pass.

PTT sentiment (data/ptt_sentiment.py)

add_ptt_sentiment() exists but returns 0 rows for all stocks. PTT scraper (scripts/survey_ptt.py) exists. The merge pipeline needs debugging β€” likely a date alignment or cache miss issue. Features ptt_sentiment_1d, ptt_sentiment_5d_ma are NOT in FEATURE_COLUMNS yet.

Improve harness threshold

Current pass gate: mean(dir_accuracy) > 31.5% β€” this was set before C3/C4/C5 raised the baseline to ~42%. Should be updated to > 43.5% and up_precision > 54% for any new experiment to be meaningful.

Market Scan (RecommendPage)

Custom stock input added (2026-05-12). Backend /api/stock/recommend accepts ?watchlist=2330,0050 query param. State and handlers are wired.


Key env vars

HF_TOKEN (or .hf_token file)       β€” HF Hub write access
SPACE_ID / HF_SPACE_ID / SPACE_HOST β€” set automatically on HF, detect HF vs Mac
WATCHLIST_STOCKS                   β€” override default 20-stock watchlist in worker
LIGHTWEIGHT_MODE=1                 β€” use 12-month training window (HF)
LINE_CHANNEL_SECRET / ACCESS_TOKEN β€” LINE Bot

How to run a new backtest experiment

# 1. Write scripts/backtest_cN.py following backtest_c3.py as template
# 2. Run improvement_harness.py which tests on 5 stocks:
./venv/bin/python scripts/improvement_harness.py --script scripts/backtest_cN.py

# Pass gate: mean dir_accuracy > 43.5% AND up_precision > 54%
# If PASSED: update FEATURE_COLUMNS in models/predictor.py, add row to docs/improvements_log.md
# If FAILED: do NOT touch models/predictor.py, log result only

Quick test commands

# Predict
curl http://localhost:7861/api/stock/2330/predict | python3 -m json.tool | head -20

# Backtest
curl "http://localhost:7861/api/stock/2330/backtest?months=6" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d['stats'])"

# Queue status
./venv/bin/python scripts/hf_queue_worker.py status

# Force-push to queue
./venv/bin/python scripts/hf_queue_worker.py push 2330 0050

# Run frontend dev server
cd frontend && npm run dev