restockiq / DECISIONS.md
RV302001's picture
RestockIQ v1: M5 quantile forecasting + reorder decision dashboard
3d4c9ec
|
Raw
History Blame Contribute Delete
6.03 kB
# DECISIONS
Running log of design decisions for RestockIQ. Newest at the bottom.
## Dataset: M5 Forecasting - Accuracy (not the Store Item Demand playground set)
The "Store Item Demand Forecasting Challenge" set only has date/store/item/sales β€” no
prices, no events, no SNAP flags β€” which makes price and calendar features impossible and
turns the project into a toy. M5 (Walmart) has 30,490 real series over ~5.4 years
(2011-01-29 to 2016-06-19), weekly sell prices per item/store, and a calendar with named
events and SNAP eligibility flags per state. That supports genuinely useful features and
an honest global-model story.
## Why quantile forecasts (P10/P50/P90) instead of a point forecast
A point forecast tells you nothing about how wrong it might be, and inventory decisions
are entirely about the cost of being wrong (stockouts vs overstock). Training LightGBM
with quantile loss at alpha = 0.1 / 0.5 / 0.9 gives a calibrated uncertainty band per
(store, item, day). The band width feeds directly into the safety-stock formula (below),
so a SKU with volatile demand automatically gets a bigger buffer than a stable one, even
if both have the same average demand.
## Safety-stock formula derivation
```
safety_stock = z * demand_std * sqrt(lead_time_days)
reorder_point = avg_daily_demand * lead_time_days + safety_stock
suggested_order_qty = max(0, reorder_point - current_inventory)
```
- `avg_daily_demand` = mean of the P50 forecast over the lead-time horizon.
- `demand_std β‰ˆ (p90 - p50) / 1.2816`, averaged over the lead-time horizon. If daily
demand were normal, the 90th percentile sits 1.2816 standard deviations above the
median, so the P90βˆ’P50 spread divided by 1.2816 recovers the daily demand std. This is
the line where the model's uncertainty band becomes an inventory number β€” a wide band
(uncertain SKU) yields a large std and therefore a large safety stock.
- `z` = normal inverse CDF of the target service level (95% β†’ 1.6449). Configurable.
- `sqrt(lead_time_days)` scales daily std to std over the lead-time window, assuming
independent daily demands.
- `current_inventory` and `lead_time_days` are illustrative inputs (the dataset has no
inventory column); both are API parameters with documented defaults and are labeled as
assumptions in the UI. No fake inventory time series is invented.
## Why no LLM
The tasks here β€” forecasting counts, computing reorder points, simulating policies β€” are
numeric problems with exact, cheap, classical solutions. An LLM adds latency, cost, and
nondeterminism and improves nothing. Per spec, the project is deliberately LLM-free.
## Naive backtest policy = trailing historical mean, no buffer (user-confirmed)
The spec's "naive fixed-reorder" is defined as: reorder point = trailing historical
average daily sales x lead_time_days, zero safety stock. This isolates exactly what the
quantile policy adds (the uncertainty buffer); both policies share the same review/order
mechanics in the simulation.
## Holdout = final 28 days (user-confirmed)
Spec allows 28-56. 28 matches the M5 competition horizon and the lag_28 feature
availability. Train on d_1 .. d_1913, validate/backtest on d_1914 .. d_1941.
## Multi-step forecasting = recursive with P50 feedback (user-confirmed)
lag_7/lag_14 are unknown beyond 7/14 days into the horizon. Forecast day-by-day and feed
the P50 prediction back as the "observed" sales for future lag/rolling computations.
Using actual holdout sales as lags would leak future information and inflate the
backtest. Keeps the spec's feature list and 3-global-models structure intact.
## Split serving requirements from pipeline requirements
`requirements.txt` (pipeline) pins everything including Prophet; the Docker image
installs `requirements-serve.txt` only (pandas/numpy/lightgbm/fastapi/uvicorn/pyarrow).
Prophet pulls in cmdstan (~hundreds of MB) and is only used offline for baselines β€”
shipping it in the Space image would bloat it for zero runtime benefit.
## Ship a 56-day history tail for the chart
`data/history_tail.parquet` (last 56 pre-holdout days of actuals per series) is written
by the backtest stage and shipped with the app, so the dashboard chart can show trailing
history before the forecast band without shipping the full 236MB sales matrix.
## Simulation start condition
Both backtest policies start at their own reorder point with nothing on order. Each
policy is judged on the inventory level it would itself choose to hold β€” starting both
at some common arbitrary level would bias the comparison toward whichever policy that
level accidentally favors.
## Training scoped to d_1000 onward (the spec's accepted fallback) + pandas-categorical fix
The first full-history training run (46.9M rows) was killed after ~3.7 hours wall time:
process inspection showed it running effectively single-threaded (CPU-time β‰ˆ wall-time on
a 32-logical-core machine) despite LightGBM's thread pool existing. A 3M-row synthetic
benchmark trained 20 rounds in 2.9s, proving tree building itself parallelizes fine β€” the
bottleneck was LightGBM's pandas-categorical ingestion path on ~40M rows. Two changes:
1. Categorical columns are converted to plain int32 codes before `lgb.Dataset`
(`encode_categoricals`), taking the fast numeric path; the code<->category mapping
comes from the shared parquet dtypes so training and inference agree.
2. Training rows restricted to `d_1000` onward β€” the back half of history β€” which is the
tradeoff the spec explicitly accepts for slow machines. Recent history dominates
short-horizon retail forecasting; 2011-2013 data adds little for a 28-day horizon.
Models are also checkpointed to models/ immediately after each quantile finishes, so a
long run's progress is observable and survives interruption.
## Python 3.13 works as-is
All pinned packages (lightgbm 4.6.0, prophet 1.3.0, fastapi 0.139.0, pandas 2.3.3,
numpy 2.3.4) have Python 3.13 wheels; installed and imported cleanly. No 3.12 fallback
needed.