--- license: cc-by-nc-4.0 library_name: xgboost tags: - finance - stock-market - indian-stock-market - nse - bse - tabular - time-series - xgboost - classification - indian-market --- # XGBoost — Indian Stock Market Prediction XGBoost binary classification model trained on historical Indian equity-market data to predict whether a stock reaches a specified positive-return threshold within a **5-trading-day horizon**. The current model targets a **+3.0% return threshold** and was trained across thousands of Indian equities using technical, market, cross-sectional, and macroeconomic features. > **Research status:** This repository contains a research model and its out-of-sample evaluation artifacts. The reported results should not be interpreted as a guarantee of future market performance or profitability. --- ## Model Overview | Property | Value | | ----------------------- | ------------------------------------ | | Model | XGBoost | | Task | Binary classification | | Prediction target | ≥ +3.0% return within 5 trading days | | Dataset size | ~6.92 million rows | | Symbols | 2,656 | | Historical period | 2008-09-02 → 2026-06-30 | | Features | 221 | | Positive class | 36.8% | | Negative class | 63.2% | | `scale_pos_weight` | 1.49 | | Training method | Histogram-based XGBoost | | Hardware | CUDA GPU | | Maximum boosting rounds | 2,000 | | Learning rate | 0.02 | | Maximum tree depth | 6 | | Early stopping | 75 rounds | --- # Dataset The training dataset contains approximately **6.92 million observations** covering **2,656 Indian stock symbols** from September 2008 through June 2026. The model uses 221 input features derived from historical market data and broader market conditions. The feature groups include: * Price returns * Momentum * Trend indicators * Volatility * Technical indicators * Relative strength * Market breadth * Cross-sectional statistics * Nifty relationships * Gold relationships * Brent crude relationships * USD/INR relationships * US-market indicators The target is a binary label representing whether the specified positive-return threshold is reached within the prediction horizon. --- # Temporal Dataset Split The dataset was divided chronologically rather than randomly. | Split | Period | Approx. Share | | ----------------- | ----------------------- | ------------: | | Training | 2008-09-02 → 2021-02-15 | 70% | | Validation | 2021-02-16 → 2024-08-30 | 20% | | Walk-forward test | 2024-09-01 → 2026-06-30 | 10% | This temporal separation is intended to reduce leakage from randomly mixing observations from different points in time. The final test period was not used for model fitting. --- # Validation Performance The model achieved the following results on the held-out validation period: **2021-02-16 → 2024-08-30** | Metric | Score | | ---------------------------------- | ---------: | | ROC-AUC | **0.9472** | | PR-AUC | **0.9042** | | Precision | **0.7981** | | Recall | **0.8364** | | Information Coefficient (Spearman) | **0.7140** | | Brier Score | **0.0835** | | Brier Skill Score | **0.6068** | | Mean Calibration Error | **0.0749** | Validation observations: **1,698,058 rows** The validation results indicate strong separation between the positive and negative classes on this historical period. However, validation performance alone should not be treated as evidence of future profitability. --- # Validation Decile Analysis Predictions were divided into ten probability-ranked groups. | Decile | Hit Rate | Samples | | -----: | -------: | ------: | | D10 | 98.4% | 169,806 | | D9 | 88.8% | 169,806 | | D8 | 60.5% | 169,806 | | D7 | 31.1% | 169,805 | | D6 | 14.6% | 169,806 | | D5 | 6.6% | 169,806 | | D4 | 3.1% | 169,805 | | D3 | 1.7% | 169,806 | | D2 | 0.9% | 169,806 | | D1 | 0.4% | 169,806 | The probability ranking shows strong separation across the validation sample, with substantially different observed positive rates between the lowest- and highest-ranked groups. These figures describe historical classification performance and do not account for transaction costs, slippage, liquidity constraints, portfolio construction, or market impact. --- # Out-of-Sample Walk-Forward Evaluation The model was subsequently evaluated chronologically over the final portion of the dataset. The test period was divided into four chronological folds of approximately six months each. | Fold | Period | PR-AUC | IC | Brier | Samples | | ---- | ----------------------- | ---------: | ---------: | -----: | ------: | | 1 | 2024-09-01 → 2025-02-28 | **0.9336** | **0.7619** | 0.0850 | 262,062 | | 2 | 2025-03-01 → 2025-08-31 | **0.8987** | **0.7044** | 0.0829 | 266,786 | | 3 | 2025-09-01 → 2026-02-28 | **0.8921** | **0.7011** | 0.0845 | 282,585 | | 4 | 2026-03-01 → 2026-06-30 | **0.8962** | **0.7157** | 0.0925 | 180,837 | ### Walk-Forward Summary | Statistic | Value | | ---------------- | ---------: | | Mean PR-AUC | **0.9052** | | Mean IC | **0.7208** | | Mean Brier Score | **0.0862** | | Minimum IC | **0.7011** | | Maximum IC | **0.7619** | The model maintained relatively strong classification performance across all four chronological folds. Performance did, however, vary between periods. The difference between the highest and lowest observed IC is approximately **0.061**. This suggests that model performance is not completely invariant across market regimes and warrants additional robustness testing. --- # Overfitting & Generalization The model was evaluated using a chronological validation set followed by a later out-of-sample test period. The results provide evidence that the model retains predictive separation outside its training period. However, the current experiment does **not** establish that the model is free from overfitting. In particular: * Validation performance is very strong. * Walk-forward test performance remains strong. * Performance varies between chronological folds. * The model reached the maximum configured **2,000 boosting rounds**, with the best iteration at **1,999**. * Early stopping therefore did not activate before the configured maximum number of rounds. The last point means that additional experiments with a larger maximum number of boosting rounds should be evaluated carefully rather than assuming that more trees will improve generalization. Future experiments should compare performance across additional temporal folds and assess whether increasing model complexity improves out-of-sample performance or simply improves the validation period. --- # Calibration The validation set produced: * Brier Score: **0.0835** * Brier Skill Score: **0.6068** * Mean Calibration Error: **0.0749** The model therefore produces probability scores that contain useful information, but the raw probabilities should not automatically be interpreted as perfectly calibrated real-world probabilities. Probability calibration should be evaluated separately if the model's output is to be interpreted probabilistically. --- # Feature Set The model uses 221 features. Examples of macro and cross-market features include: ```text usd_inr_sma20_ratio usd_inr_sma200_ratio usd_inr_momentum_10 usd_inr_momentum_20 usd_inr_volatility_20d usd_inr_trend usd_inr_zscore usd_inr_vol_zscore usd_inr_high_vol usd_inr_appreciation us_ret_1d us_ret_5d us_ret_20d us_ret_60d us_breadth_1d us_breadth_5d us_breadth_20d us_dispersion_1d us_dispersion_5d us_avg_vol_ratio us_pct_above_sma20 us_avg_rsi us_avg_dist_52w_high corr_nifty_20d corr_nifty_60d beta_nifty_60d is_high_beta is_low_beta corr_breakdown rel_strength_5d rel_strength_20d corr_gold_60d corr_gold_rising corr_brent_60d corr_brent_rising corr_usd_inr_60d corr_usd_inr_rising corr_us_60d ``` The complete feature schema is available in `model_metadata.json`. --- # Model Artifacts The repository contains the following artifacts: | File | Description | | ----------------------------- | --------------------------------------------------------------------------------- | | `xgb_model.json` | Trained XGBoost model | | `model_metadata.json` | Training configuration, feature schema, split information and evaluation metadata | | `feature_importance.csv` | XGBoost feature-importance statistics | | `val_predictions.parquet` | Validation predictions | | `test_wf_predictions.parquet` | Chronological walk-forward test predictions | The prediction files are provided to make the reported evaluation results independently inspectable and reproducible. --- # Reproducibility The model metadata records: * Feature names * Feature count * Dataset information * Training period * Validation period * Model parameters * Class weighting * Best iteration * Evaluation metrics * Training timestamp The model uses a fixed random seed of `42`. --- # Important Limitations This model should be considered a **research artifact**, not a guaranteed trading system. The reported classification metrics do not directly measure: * Portfolio returns * Sharpe ratio * Maximum drawdown * Transaction costs * Brokerage * Taxes * Slippage * Bid/ask spread * Market impact * Position sizing * Portfolio concentration * Liquidity constraints * Capacity * Execution latency A model can achieve strong classification metrics while producing poor investment returns after these factors are considered. The target is also a classification threshold rather than a direct optimization of portfolio returns. Additional research is therefore required before using the model for live decision-making. --- # Current Research Status The current experiment establishes a useful baseline for the **+3.0% / 5-trading-day** prediction task. The next stages of research should include: 1. Additional temporal walk-forward experiments. 2. Comparison of model complexity and boosting-round counts. 3. Testing alternative return thresholds and horizons. 4. Feature ablation and importance stability analysis. 5. Leakage and feature-timing audits. 6. Performance analysis by market regime. 7. Calibration analysis on completely unseen periods. 8. Portfolio-level backtesting. 9. Transaction-cost and slippage modelling. 10. Paper-trading validation before considering live deployment. --- # License This model is published under the **Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)** license. The license permits non-commercial use subject to the terms of the license. Commercial use requires appropriate licensing. --- # Author [**Sayantan Basu**](https://sayantan-basu.vercel.app)