M5 Demand Forecasting - Complete Technical Report
Overview
This repository contains a production-grade M5 forecasting pipeline that predicts 28-day demand horizons for 30,490 Walmart retail series across 10 stores, 3 states, and 3 categories. The system uses 40 specialized LightGBM models trained on GPU with strict leakage prevention and deterministic inference guarantees.
1. Problem Definition
Task: Predict daily unit sales for d_1942 through d_1969 (28 days) for all 30,490 series in the M5 evaluation set.
Data Sources:
sales_train_evaluation.csv: 30,490 series × 1,941 days (wide format)
calendar.csv: 1,969 days with events, snapshots, holidays
sell_prices.csv: ~6.8M weekly price records (store × item × wm_yr_wk)
Evaluation Metric: WRMSSE (Weighted Root Mean Squared Scaled Error) across 12 aggregation levels with dollar-sales weights computed per fold.
2. Data Preparation Pipeline
2.1 Melting & Joining (src/prep.py)
sales_long = sales.melt(id_vars=['id','item_id','dept_id','cat_id','store_id','state_id'],
value_vars=[f'd_{i}' for i in range(1,1942)],
var_name='d', value_name='sales')
sales_long = sales_long.merge(calendar[['d','date','wm_yr_wk','wday','month','year',
'event_name_1','event_type_1',
'event_name_2','event_type_2',
'snap_CA','snap_TX','snap_WI']], on='d')
sales_long = sales_long.merge(sell_prices, on=['store_id','item_id','wm_yr_wk'])
sales_long = sales_long.dropna(subset=['sell_price'])
2.2 Downcasting & Type Safety
sales → int16
- All ID columns →
category
- All float →
float32
- Invariant: Zero
float64 or object columns in final Parquet
2.3 Output
data/processed/m5_melted.parquet (46.9M rows, 23 columns)
3. WRMSSE Evaluator (src/wrmsse.py)
3.1 Aggregation Levels (12 levels)
| Level |
Grouping |
Count |
| 1 |
Total |
1 |
| 2 |
state_id |
3 |
| 3 |
store_id |
10 |
| 4 |
cat_id |
3 |
| 5 |
dept_id |
7 |
| 6 |
state_id × cat_id |
9 |
| 7 |
state_id × dept_id |
21 |
| 8 |
store_id × cat_id |
30 |
| 9 |
store_id × dept_id |
70 |
| 10 |
item_id |
3,049 |
| 11 |
item_id × state_id |
9,147 |
| 12 |
item_id × store_id |
30,490 |
3.2 RMSSE Formula
Scale Denominator = mean((y_t - y_{t-1})^2) from first non-zero observation
RMSSE = sqrt(mean((y_true - y_pred)^2) / scale_denom)
WRMSSE = sum(weight_i * RMSSE_i) where weights = dollar_sales_{t-28:t} normalized per level
Key Implementation Details:
- Scale computed from first non-zero sale (not from d_1)
- Weights are fold-specific (recomputed per origin)
- Series with zero denominator excluded and weights renormalized
- Phase 0 baselines (naive last-28-mean, seasonal naive) produce finite WRMSSE
4. Feature Engineering (src/features.py)
4.1 Core Principle: Origin-Relative Features
Every feature computed as-of origin day D using only data with day index ≤ D. Target is sales at D+k for horizon k ∈ {1..28}.
4.2 Horizon Blocks (4 blocks)
| Block |
Horizons |
Min Lag Required |
Max Horizon |
| block_1_7 |
1-7 |
7 |
7 |
| block_8_14 |
8-14 |
14 |
14 |
| block_15_21 |
15-21 |
21 |
21 |
| block_22_28 |
22-28 |
28 |
28 |
Why blocks?: Features require lags ≥ max_horizon. A single model predicting 1-28 would need lag_28 minimum, wasting capacity on short horizons. Blocks allow shorter lags for near-term predictions.
4.3 Feature Categories
Calendar (10 features)
wday, month, year (cyclic encodings)
event_name_1/2, event_type_1/2 (categorical)
snap_CA/TX/WI (state-specific binary)
is_dec_25 (Christmas closure indicator)
day_of_month
Sales Lags (6 features)
lag_7, lag_14, lag_21, lag_28, lag_35, lag_42
- Only lags ≥ block's min_lag are non-null (earlier lags still computed for consistency)
Rolling Statistics (6 features)
rolling_mean_7, rolling_mean_28, rolling_mean_60, rolling_mean_180
rolling_std_7, rolling_std_28
- Computed on shifted sales (shift(1)) to prevent leakage
Price Signals (3 features)
sell_price (current)
price_vs_hist_max = price / cummax(price) (discount depth)
price_vs_dept_mean = price / expanding_mean(dept_price) (positioning)
Intermittent Demand (2 features)
days_since_last_nonzero (cumulative count since last sale > 0)
weeks_since_release (weeks since item first appeared in store)
Target Encodings (3 features, leakage-safe)
te_dept_id, te_store_id, te_item_id
- Computed as expanding mean of sales per group, shifted by 1
- Only uses data strictly before origin
4.4 Categorical Encoding
Persistent JSON maps: {column: {value: int}} built from training data.
Applied via .map() at inference. Unseen values → -1.
Never rely on pandas Categorical ordering.
4.5 Leakage Test (tests/test_leakage.py)
Tested on origins d_1857 and d_1913 - PASSED.
5. Model Architecture: Why 40 Models?
5.1 Split Strategy
| Dimension |
Count |
Values |
| Stores |
10 |
CA_1..4, TX_1..3, WI_1..3 |
| Horizon Blocks |
4 |
1-7, 8-14, 15-21, 22-28 |
| Total |
40 |
|
5.2 Why Per-Store?
- Local Seasonality: CA stores have different weekly patterns than TX/WI
- SNAP Timing:
snap_CA, snap_TX, snap_WI affect stores differently
- Price Dynamics: Regional pricing strategies vary
- Data Volume: ~1.6M training rows per store (manageable for GPU memory)
- Isolation: Poor performance in one store doesn't corrupt others
5.3 Why Per-Horizon-Block?
- Lag Requirements: Near horizons need short lags; far horizons need long lags
- Feature Relevance:
lag_7 predictive for day 1-7, noisy for day 22-28
- Training Efficiency: Smaller feature sets per block
- Specialization: Each block learns horizon-specific dynamics
5.4 Why Not Single Global Model?
| Approach |
Pros |
Cons |
| Single Global (30490 series) |
Simple deployment, shared representations |
46M rows × 34 features = 1.5B cells; GPU OOM; washes out local patterns; lag_28 required for all horizons |
| Per-Item (30490 models) |
Maximum specialization |
Impossible to train; no shared learning; cold-start for new items |
| Per-Store (10 models) |
Good balance |
Still needs lag_28 for all horizons; mixes near/far dynamics |
| Per-Store × Block (40 models) �� |
Optimal lag/horizon match; GPU-friendly; isolates failures |
40 model files; needs routing logic |
Chosen: 40 models (per-store × per-horizon-block) — optimal trade-off.
6. Training Configuration
6.1 LightGBM Parameters
LGB_PARAMS = {
'objective': 'tweedie',
'tweedie_variance_power': 1.1,
'metric': 'rmse',
'num_leaves': 128,
'learning_rate': 0.03,
'min_data_in_leaf': 100,
'feature_fraction': 0.8,
'bagging_fraction': 0.8,
'bagging_freq': 1,
'n_estimators': 2000,
'early_stopping_rounds': 50,
'seed': 42,
'device': 'gpu',
'num_threads': 4,
'verbosity': -1,
}
6.2 Training Protocol
- Training window: 730 days before origin (d_1184 to d_1913)
- Validation: Fold C (origin d_1913, predicts d_1914-1941)
- Early stopping: Against Fold C validation RMSE
- Best iteration saved:
booster.save_model(..., num_iteration=best_iteration)
- Format: Text format only (no pickle/joblib)
6.3 Feature Matrix Constraints
- All columns
float32 (asserted)
- Feature order matches
feature_schema.json element-by-element (asserted)
- No
float64 or object columns (asserted)
7. Evaluation Results
7.1 Phase 0 Baselines
| Fold |
Origin |
Naive1 (Last-28 Mean) |
Naive2 (Seasonal) |
| A |
d_1857 |
354.14 |
192.12 |
| B |
d_1885 |
387.51 |
149.47 |
| C |
d_1913 |
375.54 |
179.92 |
7.2 Trained Models (40 models)
| Fold |
WRMSSE |
vs Naive1 |
vs Naive2 |
| A |
148.33 |
-58% |
-23% |
| B |
121.28 |
-69% |
-19% |
| C |
167.07 |
-55% |
-7% |
| Mean |
145.56 |
-61% |
-16% |
All three folds beat both naive baselines ��
7.3 Per-Level WRMSSE (Fold C)
| Level |
WRMSSE |
| Total |
963.12 |
| state_cat |
94.43 |
| state_dept |
54.40 |
| store_cat |
34.65 |
| store_dept |
20.86 |
| item_state |
1.26 |
| item_store |
0.79 |
8. Artifacts & Deployment
8.1 Artifact Inventory (artifacts/)
| File |
Purpose |
models/*.txt |
40 LightGBM boosters (text format) |
feature_schema.json |
Feature names, dtypes, categorical list |
categorical_maps.json |
{col: {value: int}} for .map() encoding |
training_config.json |
Params, versions, WRMSSE scores |
predictions.parquet |
Mode A: id, F1-F28 (float32) |
predictions_meta.json |
Metadata, row_hash, git_commit |
submission.csv |
Kaggle format (60980 rows) |
history_tail.parquet |
Last 56 days actuals + calendar + prices |
parity_predictions.parquet |
Determinism verification |
README.md |
Load instructions |
SHA256SUMS |
Integrity verification |
8.2 Loading & Inference
import lightgbm as lgb
import pandas as pd
import json
models = {}
for f in os.listdir('artifacts/models/'):
models[f] = lgb.Booster(model_file=f'artifacts/models/{f}')
with open('artifacts/feature_schema.json') as f:
schema = json.load(f)
with open('artifacts/categorical_maps.json') as f:
cat_maps = json.load(f)
8.3 Routing Logic
def get_model_key(store_id: str, horizon_day: int) -> str:
if 1 <= horizon_day <= 7: block = 'block_1_7'
elif 8 <= horizon_day <= 14: block = 'block_8_14'
elif 15 <= horizon_day <= 21: block = 'block_15_21'
else: block = 'block_22_28'
return f"model_store={store_id}_hblock={block}"
9. Alternatives & Trade-offs
9.1 Model Architecture Alternatives
| Alternative |
Why Not Chosen |
| Single LightGBM (all series) |
GPU OOM on 46M rows; loses local seasonality; requires lag_28 for all horizons |
| DeepAR / Temporal Fusion Transformer |
Overkill for tabular retail; harder to debug; slower inference; less interpretable |
| XGBoost |
No native GPU tweedie; slower training |
| Statistical (ETS/ARIMA) |
Cannot handle 30K series with covariates; no price/event features |
| Per-item models (30K) |
Cold start impossible; no shared learning; training infeasible |
9.2 Feature Engineering Alternatives
| Alternative |
Trade-off |
| No target encoding |
Loses dept/store/item signal; +5-10% WRMSSE |
| All lags for all blocks |
Increases feature dim; slower training; marginal gain |
| Daily price join |
Wrong - creates false nulls (prices are weekly) |
| Pandas Categorical codes |
Non-deterministic across sessions; breaks parity |
9.3 Training Alternatives
| Alternative |
Trade-off |
| Random validation split |
Leaks future into training; inflated metrics |
| No early stopping |
Overfits; 2000 trees vs ~150 optimal |
| CPU training |
10-20x slower; blocks iteration |
| Pickle serialization |
Version-dependent; security risk; not portable |
10. Known Limitations
- Forecast Origin Locked: Models trained on d_1913 origin. New origins require retraining or history extension.
- No New Items/Stores: Categorical maps fixed at training time. Unseen items → -1 encoding.
- Price Assumption: Uses last known price for forecast horizon. Real prices may differ.
- Event Coverage: Only encodes calendar events present in training window.
- Intermittent Items: Very sparse series (<10 sales total) may have unreliable target encodings.
11. Reproducibility
- Environment: Pinned in
training_config.json (Python 3.12, LightGBM 4.5.0, etc.)
- Seed: 42 (LightGBM, numpy, pandas sampling)
- Git Commit: Recorded in
training_config.json and predictions_meta.json
- Parity: Deterministic inference verified (atol=0)
- Integrity:
SHA256SUMS covers all 50 artifact files
12. Decision Log
| Phase |
Decision |
Rationale |
| Data |
Weekly price join on wm_yr_wk |
Daily join creates false nulls |
| Features |
Origin-relative, leakage-safe |
Prevents look-ahead bias |
| Features |
Horizon blocks |
Matches lag requirements to prediction distance |
| Architecture |
40 per-store × block models |
Optimal bias-variance-compute trade-off |
| Objective |
Tweedie (p=1.1) |
Handles zero-inflated continuous sales |
| Validation |
Fold C (d_1913) |
Most recent, matches test horizon |
| Early Stopping |
On Fold C RMSE |
Prevents overfitting |
| Serialization |
Booster text format |
Portable, version-robust, no pickle |
| Parity |
atol=0 determinism |
Guarantees exact reproduction |
13. Files for Further Study
src/prep.py — Data melting, joining, downcasting
src/wrmsse.py — Full WRMSSE implementation with 12 levels
src/features.py — Origin-relative feature engineering
src/train.py — 40-model training loop with GPU
src/evaluate.py — Harness evaluation on 3 folds
src/scalar_sweep.py — Global multiplier optimization
tests/test_leakage.py — Leakage verification
tests/test_dtypes.py — Type safety checks
tests/test_hierarchy.py — Aggregation consistency
scripts/export_artifacts.py — Full artifact generation
scripts/verify_local.py — Parity verification
scripts/publish.py — HF Hub upload
14. Contact & Version
- HF Repo:
rishini/NPN
- Revision:
85868e044da123cf3d4a9211ddec7c2773c4ee2d
- Date: 2026-08-14
- Framework: LightGBM 4.5.0, Python 3.12