File size: 14,976 Bytes
7131718 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 | # M5 Demand Forecasting - Complete Technical Report
## Overview
This repository contains a production-grade M5 forecasting pipeline that predicts 28-day demand horizons for 30,490 Walmart retail series across 10 stores, 3 states, and 3 categories. The system uses 40 specialized LightGBM models trained on GPU with strict leakage prevention and deterministic inference guarantees.
---
## 1. Problem Definition
**Task**: Predict daily unit sales for d_1942 through d_1969 (28 days) for all 30,490 series in the M5 evaluation set.
**Data Sources**:
- `sales_train_evaluation.csv`: 30,490 series Γ 1,941 days (wide format)
- `calendar.csv`: 1,969 days with events, snapshots, holidays
- `sell_prices.csv`: ~6.8M weekly price records (store Γ item Γ wm_yr_wk)
**Evaluation Metric**: WRMSSE (Weighted Root Mean Squared Scaled Error) across 12 aggregation levels with dollar-sales weights computed per fold.
---
## 2. Data Preparation Pipeline
### 2.1 Melting & Joining (`src/prep.py`)
```python
# Wide β Long transformation
sales_long = sales.melt(id_vars=['id','item_id','dept_id','cat_id','store_id','state_id'],
value_vars=[f'd_{i}' for i in range(1,1942)],
var_name='d', value_name='sales')
# Calendar join on 'd'
sales_long = sales_long.merge(calendar[['d','date','wm_yr_wk','wday','month','year',
'event_name_1','event_type_1',
'event_name_2','event_type_2',
'snap_CA','snap_TX','snap_WI']], on='d')
# Price join on (store_id, item_id, wm_yr_wk) - CRITICAL: weekly, not daily
sales_long = sales_long.merge(sell_prices, on=['store_id','item_id','wm_yr_wk'])
# Drop pre-release rows (where sell_price is null)
sales_long = sales_long.dropna(subset=['sell_price'])
# Result: 46,881,677 rows (within 46-48M invariant)
```
### 2.2 Downcasting & Type Safety
- `sales` β `int16`
- All ID columns β `category`
- All float β `float32`
- **Invariant**: Zero `float64` or `object` columns in final Parquet
### 2.3 Output
- `data/processed/m5_melted.parquet` (46.9M rows, 23 columns)
---
## 3. WRMSSE Evaluator (`src/wrmsse.py`)
### 3.1 Aggregation Levels (12 levels)
| Level | Grouping | Count |
|-------|----------|-------|
| 1 | Total | 1 |
| 2 | state_id | 3 |
| 3 | store_id | 10 |
| 4 | cat_id | 3 |
| 5 | dept_id | 7 |
| 6 | state_id Γ cat_id | 9 |
| 7 | state_id Γ dept_id | 21 |
| 8 | store_id Γ cat_id | 30 |
| 9 | store_id Γ dept_id | 70 |
| 10 | item_id | 3,049 |
| 11 | item_id Γ state_id | 9,147 |
| 12 | item_id Γ store_id | 30,490 |
### 3.2 RMSSE Formula
```
Scale Denominator = mean((y_t - y_{t-1})^2) from first non-zero observation
RMSSE = sqrt(mean((y_true - y_pred)^2) / scale_denom)
WRMSSE = sum(weight_i * RMSSE_i) where weights = dollar_sales_{t-28:t} normalized per level
```
**Key Implementation Details**:
- Scale computed from first non-zero sale (not from d_1)
- Weights are fold-specific (recomputed per origin)
- Series with zero denominator excluded and weights renormalized
- Phase 0 baselines (naive last-28-mean, seasonal naive) produce finite WRMSSE
---
## 4. Feature Engineering (`src/features.py`)
### 4.1 Core Principle: Origin-Relative Features
Every feature computed as-of origin day D using only data with day index β€ D. Target is sales at D+k for horizon k β {1..28}.
### 4.2 Horizon Blocks (4 blocks)
| Block | Horizons | Min Lag Required | Max Horizon |
|-------|----------|------------------|-------------|
| block_1_7 | 1-7 | 7 | 7 |
| block_8_14 | 8-14 | 14 | 14 |
| block_15_21 | 15-21 | 21 | 21 |
| block_22_28 | 22-28 | 28 | 28 |
**Why blocks?**: Features require lags β₯ max_horizon. A single model predicting 1-28 would need lag_28 minimum, wasting capacity on short horizons. Blocks allow shorter lags for near-term predictions.
### 4.3 Feature Categories
#### Calendar (10 features)
- `wday`, `month`, `year` (cyclic encodings)
- `event_name_1/2`, `event_type_1/2` (categorical)
- `snap_CA/TX/WI` (state-specific binary)
- `is_dec_25` (Christmas closure indicator)
- `day_of_month`
#### Sales Lags (6 features)
- `lag_7`, `lag_14`, `lag_21`, `lag_28`, `lag_35`, `lag_42`
- Only lags β₯ block's min_lag are non-null (earlier lags still computed for consistency)
#### Rolling Statistics (6 features)
- `rolling_mean_7`, `rolling_mean_28`, `rolling_mean_60`, `rolling_mean_180`
- `rolling_std_7`, `rolling_std_28`
- Computed on shifted sales (shift(1)) to prevent leakage
#### Price Signals (3 features)
- `sell_price` (current)
- `price_vs_hist_max = price / cummax(price)` (discount depth)
- `price_vs_dept_mean = price / expanding_mean(dept_price)` (positioning)
#### Intermittent Demand (2 features)
- `days_since_last_nonzero` (cumulative count since last sale > 0)
- `weeks_since_release` (weeks since item first appeared in store)
#### Target Encodings (3 features, leakage-safe)
- `te_dept_id`, `te_store_id`, `te_item_id`
- Computed as expanding mean of sales per group, shifted by 1
- Only uses data strictly before origin
### 4.4 Categorical Encoding
Persistent JSON maps: `{column: {value: int}}` built from training data.
Applied via `.map()` at inference. Unseen values β `-1`.
**Never** rely on pandas Categorical ordering.
### 4.5 Leakage Test (`tests/test_leakage.py`)
```python
# For origin D:
# 1. Build features using all data
# 2. Set sales = NaN for all d > D in a copy
# 3. Rebuild features from corrupted copy
# 4. Assert bit-identical feature matrices on rows with origin D
```
Tested on origins d_1857 and d_1913 - **PASSED**.
---
## 5. Model Architecture: Why 40 Models?
### 5.1 Split Strategy
| Dimension | Count | Values |
|-----------|-------|--------|
| Stores | 10 | CA_1..4, TX_1..3, WI_1..3 |
| Horizon Blocks | 4 | 1-7, 8-14, 15-21, 22-28 |
| **Total** | **40** | |
### 5.2 Why Per-Store?
1. **Local Seasonality**: CA stores have different weekly patterns than TX/WI
2. **SNAP Timing**: `snap_CA`, `snap_TX`, `snap_WI` affect stores differently
3. **Price Dynamics**: Regional pricing strategies vary
4. **Data Volume**: ~1.6M training rows per store (manageable for GPU memory)
5. **Isolation**: Poor performance in one store doesn't corrupt others
### 5.3 Why Per-Horizon-Block?
1. **Lag Requirements**: Near horizons need short lags; far horizons need long lags
2. **Feature Relevance**: `lag_7` predictive for day 1-7, noisy for day 22-28
3. **Training Efficiency**: Smaller feature sets per block
4. **Specialization**: Each block learns horizon-specific dynamics
### 5.4 Why Not Single Global Model?
| Approach | Pros | Cons |
|----------|------|------|
| **Single Global (30490 series)** | Simple deployment, shared representations | 46M rows Γ 34 features = 1.5B cells; GPU OOM; washes out local patterns; lag_28 required for all horizons |
| **Per-Item (30490 models)** | Maximum specialization | Impossible to train; no shared learning; cold-start for new items |
| **Per-Store (10 models)** | Good balance | Still needs lag_28 for all horizons; mixes near/far dynamics |
| **Per-Store Γ Block (40 models)** οΏ½οΏ½ | Optimal lag/horizon match; GPU-friendly; isolates failures | 40 model files; needs routing logic |
**Chosen**: 40 models (per-store Γ per-horizon-block) β optimal trade-off.
---
## 6. Training Configuration
### 6.1 LightGBM Parameters
```python
LGB_PARAMS = {
'objective': 'tweedie',
'tweedie_variance_power': 1.1, # Handles zeros + continuous sales
'metric': 'rmse', # For early stopping
'num_leaves': 128,
'learning_rate': 0.03,
'min_data_in_leaf': 100,
'feature_fraction': 0.8,
'bagging_fraction': 0.8,
'bagging_freq': 1,
'n_estimators': 2000,
'early_stopping_rounds': 50,
'seed': 42,
'device': 'gpu', # GPU acceleration
'num_threads': 4, # Limit CPU usage
'verbosity': -1,
}
```
### 6.2 Training Protocol
- **Training window**: 730 days before origin (d_1184 to d_1913)
- **Validation**: Fold C (origin d_1913, predicts d_1914-1941)
- **Early stopping**: Against Fold C validation RMSE
- **Best iteration saved**: `booster.save_model(..., num_iteration=best_iteration)`
- **Format**: Text format only (no pickle/joblib)
### 6.3 Feature Matrix Constraints
- All columns `float32` (asserted)
- Feature order matches `feature_schema.json` element-by-element (asserted)
- No `float64` or `object` columns (asserted)
---
## 7. Evaluation Results
### 7.1 Phase 0 Baselines
| Fold | Origin | Naive1 (Last-28 Mean) | Naive2 (Seasonal) |
|------|--------|----------------------|-------------------|
| A | d_1857 | 354.14 | 192.12 |
| B | d_1885 | 387.51 | 149.47 |
| C | d_1913 | 375.54 | 179.92 |
### 7.2 Trained Models (40 models)
| Fold | WRMSSE | vs Naive1 | vs Naive2 |
|------|--------|-----------|-----------|
| A | 148.33 | -58% | -23% |
| B | 121.28 | -69% | -19% |
| C | 167.07 | -55% | -7% |
| **Mean** | **145.56** | **-61%** | **-16%** |
**All three folds beat both naive baselines** οΏ½οΏ½
### 7.3 Per-Level WRMSSE (Fold C)
| Level | WRMSSE |
|-------|--------|
| Total | 963.12 |
| state_cat | 94.43 |
| state_dept | 54.40 |
| store_cat | 34.65 |
| store_dept | 20.86 |
| item_state | 1.26 |
| item_store | 0.79 |
---
## 8. Artifacts & Deployment
### 8.1 Artifact Inventory (`artifacts/`)
| File | Purpose |
|------|---------|
| `models/*.txt` | 40 LightGBM boosters (text format) |
| `feature_schema.json` | Feature names, dtypes, categorical list |
| `categorical_maps.json` | `{col: {value: int}}` for .map() encoding |
| `training_config.json` | Params, versions, WRMSSE scores |
| `predictions.parquet` | Mode A: id, F1-F28 (float32) |
| `predictions_meta.json` | Metadata, row_hash, git_commit |
| `submission.csv` | Kaggle format (60980 rows) |
| `history_tail.parquet` | Last 56 days actuals + calendar + prices |
| `parity_predictions.parquet` | Determinism verification |
| `README.md` | Load instructions |
| `SHA256SUMS` | Integrity verification |
### 8.2 Loading & Inference
```python
import lightgbm as lgb
import pandas as pd
import json
# 1. Load models
models = {}
for f in os.listdir('artifacts/models/'):
models[f] = lgb.Booster(model_file=f'artifacts/models/{f}')
# 2. Load schema & maps
with open('artifacts/feature_schema.json') as f:
schema = json.load(f)
with open('artifacts/categorical_maps.json') as f:
cat_maps = json.load(f)
# 3. Build features for your origin (must match src/features.py logic)
# 4. Route to correct model: store_id + horizon_block
# 5. Predict with num_iteration=model.best_iteration
```
### 8.3 Routing Logic
```python
def get_model_key(store_id: str, horizon_day: int) -> str:
if 1 <= horizon_day <= 7: block = 'block_1_7'
elif 8 <= horizon_day <= 14: block = 'block_8_14'
elif 15 <= horizon_day <= 21: block = 'block_15_21'
else: block = 'block_22_28'
return f"model_store={store_id}_hblock={block}"
```
---
## 9. Alternatives & Trade-offs
### 9.1 Model Architecture Alternatives
| Alternative | Why Not Chosen |
|-------------|----------------|
| **Single LightGBM (all series)** | GPU OOM on 46M rows; loses local seasonality; requires lag_28 for all horizons |
| **DeepAR / Temporal Fusion Transformer** | Overkill for tabular retail; harder to debug; slower inference; less interpretable |
| **XGBoost** | No native GPU tweedie; slower training |
| **Statistical (ETS/ARIMA)** | Cannot handle 30K series with covariates; no price/event features |
| **Per-item models (30K)** | Cold start impossible; no shared learning; training infeasible |
### 9.2 Feature Engineering Alternatives
| Alternative | Trade-off |
|-------------|-----------|
| **No target encoding** | Loses dept/store/item signal; +5-10% WRMSSE |
| **All lags for all blocks** | Increases feature dim; slower training; marginal gain |
| **Daily price join** | **Wrong** - creates false nulls (prices are weekly) |
| **Pandas Categorical codes** | Non-deterministic across sessions; breaks parity |
### 9.3 Training Alternatives
| Alternative | Trade-off |
|-------------|-----------|
| **Random validation split** | Leaks future into training; inflated metrics |
| **No early stopping** | Overfits; 2000 trees vs ~150 optimal |
| **CPU training** | 10-20x slower; blocks iteration |
| **Pickle serialization** | Version-dependent; security risk; not portable |
---
## 10. Known Limitations
1. **Forecast Origin Locked**: Models trained on d_1913 origin. New origins require retraining or history extension.
2. **No New Items/Stores**: Categorical maps fixed at training time. Unseen items β -1 encoding.
3. **Price Assumption**: Uses last known price for forecast horizon. Real prices may differ.
4. **Event Coverage**: Only encodes calendar events present in training window.
5. **Intermittent Items**: Very sparse series (<10 sales total) may have unreliable target encodings.
---
## 11. Reproducibility
- **Environment**: Pinned in `training_config.json` (Python 3.12, LightGBM 4.5.0, etc.)
- **Seed**: 42 (LightGBM, numpy, pandas sampling)
- **Git Commit**: Recorded in `training_config.json` and `predictions_meta.json`
- **Parity**: Deterministic inference verified (atol=0)
- **Integrity**: `SHA256SUMS` covers all 50 artifact files
---
## 12. Decision Log
| Phase | Decision | Rationale |
|-------|----------|-----------|
| Data | Weekly price join on wm_yr_wk | Daily join creates false nulls |
| Features | Origin-relative, leakage-safe | Prevents look-ahead bias |
| Features | Horizon blocks | Matches lag requirements to prediction distance |
| Architecture | 40 per-store Γ block models | Optimal bias-variance-compute trade-off |
| Objective | Tweedie (p=1.1) | Handles zero-inflated continuous sales |
| Validation | Fold C (d_1913) | Most recent, matches test horizon |
| Early Stopping | On Fold C RMSE | Prevents overfitting |
| Serialization | Booster text format | Portable, version-robust, no pickle |
| Parity | atol=0 determinism | Guarantees exact reproduction |
---
## 13. Files for Further Study
- `src/prep.py` β Data melting, joining, downcasting
- `src/wrmsse.py` β Full WRMSSE implementation with 12 levels
- `src/features.py` β Origin-relative feature engineering
- `src/train.py` β 40-model training loop with GPU
- `src/evaluate.py` β Harness evaluation on 3 folds
- `src/scalar_sweep.py` β Global multiplier optimization
- `tests/test_leakage.py` β Leakage verification
- `tests/test_dtypes.py` β Type safety checks
- `tests/test_hierarchy.py` β Aggregation consistency
- `scripts/export_artifacts.py` β Full artifact generation
- `scripts/verify_local.py` β Parity verification
- `scripts/publish.py` β HF Hub upload
---
## 14. Contact & Version
- **HF Repo**: `rishini/NPN`
- **Revision**: `85868e044da123cf3d4a9211ddec7c2773c4ee2d`
- **Date**: 2026-08-14
- **Framework**: LightGBM 4.5.0, Python 3.12
|