File size: 24,286 Bytes
749bffa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
# 02 - Feature engineering

[Back to 01: Dataset and EDA](01_dataset_and_eda.md) | [README](../README.md)

Covers **Phase 4**, the scientific core: 57 features across 7 groups, built at
each of 5 diagnostic budgets under a structurally enforced no-leakage rule.

The feature tables below are **generated from the code registry**
([`outputs/reports/feature_registry.json`](../outputs/reports/feature_registry.json)),
so this document cannot drift from what is actually computed.

---

## 1. The leakage rule is enforced by construction

For budget *N*, features may use only cycles 1..*N*. That is guaranteed
structurally, not by discipline:

1. [`builder.py`](../src/features/builder.py) slices **all three** per-cell
   frames to `cycle_measured <= max_cycle` **once, at the top**.
2. It **asserts** the slice maximum does not exceed the budget.
3. It passes **only those slices** to the group modules.

Group modules accept DataFrames -- never a cell id, never a path -- so they have
no route back to unsliced data. `build_cell_features` is therefore a pure
function of its frames, which is what makes the test below possible.

### The test has demonstrated teeth

[`tests/test_leakage.py`](../tests/test_leakage.py) computes a cell's features,
then **destroys every cycle beyond N** -- shuffling rows, rescaling values and
adding noise, all three together because a statistic can be invariant to any one
alone -- and demands bit-identical output.

A passing test proves nothing unless it can fail, so the suite carries two
controls:

- **Negative control.** A deliberately leaky statistic *does* change under the
  corruption, confirming the corruption is effective.
- **Mutation test.** The builder's slice is deliberately weakened to `3 x N`;
  **45 of 57 features change**, so the test detects the realistic regression.

**A limitation found while building this, and closed.** The invariance test can
only catch leaks that flow through the frames passed in. An early mutation
attempt that read from a captured global was **not** detected: because the
corruption produces copies, the leaky path saw identical data both times. A
group reading from disk would evade the test entirely. That vector is now closed
by a **static guard** asserting no feature module references `read_parquet`,
`open(`, `CELLS_DIR` or `load_cell_frames`.

## 2. Cycle numbering: `cycle_measured`, not the raw index



Phase 2 established that the all-zero placeholder row exists in batch 1 (46/46

cells) but not in batches 2-3 (0/89). Slicing "cycles 1..N" on the raw index

would hand batch-1 cells one *fewer* measured cycle at every budget -- a

systematic, batch-correlated bias in exactly the variable RQ4 tests for.



Every feature therefore indexes on `cycle_measured`, which counts only measured
cycles from 1, identically in every batch. It is attached to all three frames at
parse time and re-derived after the continuation join.

**One documented consequence.** The literature's "discharge capacity at cycle 2"
is a raw index, which in batch 1 is the *first* measured cycle. Here
`q_at_cycle_2` is the second *measured* cycle for every cell. That is a
one-cycle offset from a naive raw reading in batch 1 only, it is uniform across
the cohort, and it removes a batch-dependent inconsistency rather than
inheriting one. It does not affect Gate 1, which reproduced at R^2 = 0.859 using
measured cycles 10 and 100.

## 3. The plausibility mask

Physically impossible readings become `NaN` **before any feature sees them**,
using the same bounds the Phase 2 validator applied, applied identically to
every cell. Masking by *physics* rather than by cell name is what makes this
defensible: it is a stated rule about what a real measurement can be, not a
patch for an inconvenient cell. Imputation happens **inside** the CV pipeline,
on training-fold statistics only.

### What the mask caught, and why it mattered

Phase 4 found a **third** corrupted channel on b1c18, beyond the two Phase 2
identified: its *median* cycle duration reads **2236 minutes (37 hours)** against
a cohort median of **51.8 min**.

This was not cosmetic. Thermal exposure is an *integral over time*, so a broken
clock propagates straight into Group D. b1c18's `thermal_exposure` came out
**71x the cohort median**, and that one feature drove a cross-validated ridge
prediction to **-50.9** against a true value of 2.84, destroying an entire fold
(RMSE 10.75 against ~0.07 for the others).

The bound was derived from the data, not guessed:

| Measurement (cycles 1-100, 135 cells) | Value |
|---|---|
| Median cycle duration, median across cells | 51.8 min |
| 118 of 135 cells never exceed | 78.6 min |
| 16 cells with one isolated long cycle | ~460 min (plausibly a real rest) |
| **b1c18 median** | **2236 min** |

`max_cycle_duration_minutes: 300` sits ~6x the median and ~4x above the ceiling
of the 118 clean cells. It is applied **per cycle**, so a cell with one genuine
long rest keeps its other 99.

After the fix, cross-validated RMSE across budgets fell from 0.13-2.21 to
**0.075-0.16** with consistent folds. *(Sanity check only -- Phase 6 owns

modelling.)*

## 4. Corrections to the Phase 4 plan, from measurement

Two flags raised in the plan did not survive contact with the data. Recorded
rather than quietly dropped:

| Plan said | Measurement says |
|---|---|
| Group A is **most exposed** to b1c18, whose voltage spans 0.736-6.606 V | **Wrong.** b1c18's `Qdlin` spans [-0.0004, 1.0523] Ah, indistinguishable from peers, and its `log_abs_dq_var` sits at the **60th percentile**. `Qdlin` derives from the *discharge* segment; b1c18's corruption is in the charge phase. |
| The `fig04` DeltaQ(V) spike near 3.09 V is b1c18 | **Wrong.** It is **b1c41**, the only cell with a localized DeltaQ(V) spike (8.7 sigma). Corrected in [docs/01](01_dataset_and_eda.md). |

Groups C and E were confirmed **unaffected** by b1c18: its internal resistance
(0.0158-0.0199 Ohm) and charge time (9.87-9.97 min) are plausible and
peer-consistent.

### A known fragility, left honest

`dq_kurtosis` for **b1c41** is 112.9 against a cohort inter-quartile range of
[-1.10, -0.30], because its DeltaQ(V) contains one localized spike. The feature
is **correct** -- it faithfully reports a heavy-tailed curve -- and it has
**not** been clipped. Suppressing it would be fitting the data to the method.

The consequence is that *linear* models extrapolate badly from this cell.
Phase 6 must therefore use robust preprocessing or outlier-insensitive
estimators. `configs/models.yaml` already declares Huber regression and tree
ensembles alongside the linear models, and this is a concrete reason to compare
them rather than a stylistic preference.

## 5. Feature groups

| Group | Module | Features | In-line | Process recipe |
|---|---|---:|:--:|:--:|
| **A_curve** | `curve_features.py` | 12 | yes | no |

| **B_degradation** | `degradation.py` | 10 | yes | no |

| **C_resistance** | `resistance.py` | 7 | yes | no |

| **D_thermal** | `thermal.py` | 8 | NO | no |

| **E_charge_dynamics** | `charge_dynamics.py` | 6 | yes | no |
| **F_protocol** | `protocol.py` | 8 | yes | **yes** |

| **G_interactions** | `interactions.py` | 6 | yes | no |
| | **total** | **57** | | |

### Measurable in-line on a production line

45 features requiring no instrumentation beyond what a production cycler already records.

| Feature | Group | Formula | Rationale |
|---|---|---|---|
| `dq_integral` | A_curve | sum(DeltaQ(V)) over the voltage grid | Total accumulated capacity difference across the discharge window; a magnitude complement to the shape statistics above. |

| `dq_kurtosis` | A_curve | kurtosis(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |

| `dq_mean` | A_curve | mean(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |

| `dq_min` | A_curve | min(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |

| `dq_skew` | A_curve | skew(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |

| `dq_var` | A_curve | var(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |

| `log_abs_dq_kurtosis` | A_curve | log10(\|kurtosis(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |

| `log_abs_dq_mean` | A_curve | log10(\|mean(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |

| `log_abs_dq_min` | A_curve | log10(\|min(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |

| `log_abs_dq_skew` | A_curve | log10(\|skew(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |

| `log_abs_dq_var` | A_curve | log10(\|var(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |

| `q_discharge_curve_area_N` | A_curve | sum(Q_N(V)) over the voltage grid | Absolute area under the interpolated discharge curve at the budget cycle, capturing overall accessible capacity independently of the difference curve. |

| `cycle_at_q_max` | B_degradation | measured cycle at which Q_discharge peaks | When the trajectory turns over. An early peak means net fade has already overtaken activation. |
| `q_at_cycle_2` | B_degradation | Q_discharge at measured cycle 2 | Early-life capacity, effectively the as-manufactured usable capacity once formation has settled. The reference point for all fade measures. |
| `q_at_cycle_N` | B_degradation | Q_discharge at the budget cycle N | Capacity at the moment the QC decision must be made. |
| `q_diff_N_2` | B_degradation | Q(N) - Q(2) | Absolute capacity change over the observation window. Small and noisy this early, which is precisely why curve-shape features outperform it. |

| `q_fade_rate` | B_degradation | (Q(N) - Q(2)) / (N - 2) | Capacity loss per cycle, normalising the above by window length so budgets are comparable. |
| `q_intercept` | B_degradation | least-squares intercept of Q(cycle) over the window | Fitted capacity at cycle 0; a noise-robust proxy for initial capacity. |

| `q_max` | B_degradation | max Q_discharge over the window | Peak capacity. Many cells gain capacity for tens of cycles as formation completes before net fade begins. |
| `q_max_minus_q2` | B_degradation | max(Q) - Q(2) | Size of that early capacity rise. A larger rise indicates continuing electrode wetting and activation. |

| `q_slope` | B_degradation | least-squares slope of Q(cycle) over the window | Trend estimate using every cycle rather than just the endpoints, so it is far less sensitive to noise in any single measurement. |

| `q_var` | B_degradation | variance of Q_discharge over the window | Trajectory roughness; elevated variance indicates unstable cycling behaviour. |
| `ir_at_cycle_2` | C_resistance | internal resistance at measured cycle 2 | Post-formation baseline resistance, set by the interphase built during formation. A high starting value indicates lithium already consumed. |

| `ir_at_cycle_N` | C_resistance | internal resistance at the budget cycle N | Resistance at the QC decision point. |

| `ir_diff_N_2` | C_resistance | IR(N) - IR(2) | Resistance growth over the window: the direct signature of continuing interphase growth and contact loss. |

| `ir_mean` | C_resistance | mean internal resistance over the window | Window-average level, robust to single-cycle measurement noise. |

| `ir_min` | C_resistance | min internal resistance over the window | Resistance floor. Cells typically fall to a minimum as wetting completes before rising; the floor is a cleaner baseline than any single cycle. |

| `ir_ratio_N_min` | C_resistance | IR(N) / min(IR) | Relative rise from the floor. Being dimensionless it is comparable across cells with different absolute resistance. |

| `ir_slope` | C_resistance | least-squares slope of IR(cycle) over the window | Rate of resistance rise, using every cycle rather than two endpoints. |

| `cc_time_fraction` | E_charge_dynamics | mean over cycles of (time in CC) / (total charge time) | Share of charging spent at constant current. As impedance rises the cell hits the voltage limit earlier and this fraction falls, making it a more specific impedance probe than total charge time. |

| `charge_time_diff_N_2` | E_charge_dynamics | charge_time(N) - charge_time(2) | Endpoint change in charge duration over the window. |

| `charge_time_mean` | E_charge_dynamics | mean charge time over measured cycles 2..N | Typical charge duration under a fixed protocol; a direct impedance proxy. |

| `charge_time_slope` | E_charge_dynamics | least-squares slope of charge time vs cycle | Whether charge time is drifting upward, the signature of growing polarization resistance. |

| `charge_time_var` | E_charge_dynamics | variance of charge time over measured cycles 2..N | Cycle-to-cycle instability in charge acceptance. |

| `cv_time_fraction` | E_charge_dynamics | mean over cycles of (time in CV) / (total charge time) | Complement of the above; the constant-voltage tail lengthens as the cell becomes harder to charge. |

| `policy_c1` | F_protocol | first-step charging C-rate | How aggressively the cell is charged at low state of charge, where lithium plating risk is highest. |

| `policy_c2` | F_protocol | second-step charging C-rate | Charging rate through the middle of the state-of-charge range. |

| `policy_c3` | F_protocol | third-step charging C-rate where the recipe defines one | Present only for multi-step recipes; NaN otherwise, and imputed inside the CV pipeline rather than filled with a fabricated rate. |

| `policy_c_max` | F_protocol | maximum C-rate used anywhere in the recipe | Peak stress the cell experiences during charging. |
| `policy_c_mean` | F_protocol | SOC-weighted mean C-rate across the recipe steps | Single summary of overall charging aggressiveness, weighting each rate by the fraction of charge delivered at it. |

| `policy_c_range` | F_protocol | policy_c_max - min C-rate in the recipe | Spread between the gentlest and harshest step; a flat recipe and a steeply stepped one can share a mean but stress the cell very differently. |
| `policy_n_steps` | F_protocol | number of steps in the recipe | Recipe complexity, distinguishing simple two-step from multi-step profiles. |

| `policy_q1_pct` | F_protocol | state of charge (%) at which the protocol switches step | Where the recipe hands over between rates; determines how much of the high-plating-risk region is traversed at the first rate. |
| `charge_time_slope_x_ir_slope` | G_interactions | charge_time_slope * ir_slope | Both are proxies for growing polarization resistance measured through different channels. Their product is large only when the two agree, suppressing the case where either drifts through measurement noise alone. |
| `ir_growth_per_cycle` | G_interactions | ir_diff_N_2 / (N - 2) | Resistance growth normalised by window length so that budgets are directly comparable; without it the raw difference grows mechanically with N. |

### NOT reliably measurable in-line - deployability caveat

12 features needing per-cell instrumentation that many lines do not have. A screening rule depending on these may not be deployable without additional hardware, which is why the distinction is tracked in the registry rather than left implicit.

| Feature | Group | Formula | Rationale |
|---|---|---|---|
| `temp_avg_slope` | D_thermal | least-squares slope of per-cycle mean temperature vs cycle | Whether the cell is progressively self-heating, which indicates rising internal impedance dissipating more energy. |

| `temp_max` | D_thermal | max of per-cycle maximum temperature over the window | Worst-case thermal excursion, which drives the fastest side reactions. |

| `temp_max_slope` | D_thermal | least-squares slope of per-cycle maximum temperature vs cycle | The same trend measured at the thermal peak, where it appears earliest. |
| `temp_mean` | D_thermal | mean of per-cycle average temperature over the window | Typical operating temperature of the cell. |

| `temp_min` | D_thermal | min of per-cycle minimum temperature over the window | Lower bound of the thermal envelope; with the maximum it gives the range. |

| `temp_range` | D_thermal | temp_max - temp_min | Thermal swing per cycle, a proxy for heat generated under load relative to the chamber setpoint. |

| `thermal_exposure` | D_thermal | sum over cycles of the integral of T dt within each cycle | Cumulative time-at-temperature, the quantity an Arrhenius rate law integrates. Computed from the within-cycle traces rather than from per-cycle averages so that cycles of differing duration are weighted correctly. |

| `thermal_exposure_per_cycle` | D_thermal | thermal_exposure / number of measured cycles | Rate rather than accumulation, so budgets are comparable. |
| `fade_per_thermal_exposure` | G_interactions | q_fade_rate / thermal_exposure | Capacity loss normalised by thermal load. A cell fading fast despite low thermal exposure is degrading for a non-thermal reason, which is diagnostically distinct. |
| `log_dq_var_x_ir_ratio` | G_interactions | log_abs_dq_var * ir_ratio_N_min | Combines loss of accessible lithium inventory (curve shape) with rising interfacial impedance, distinguishing cells losing capacity through inventory loss from those losing it through impedance rise. |

| `recipe_stress_x_thermal` | G_interactions | policy_c_max * thermal_exposure | Higher charging C-rates drive both ohmic heating and lithium plating. This links the commanded process stress to the thermal response it actually produced. PROCESS-RECIPE DERIVED: excluded from the no_recipe feature set with Group F. |

| `thermal_x_ir_growth` | G_interactions | thermal_exposure * ir_diff_N_2 | Side reactions are Arrhenius-activated AND proceed at the interphase. A cell that is both hot and building resistance is degrading by both mechanisms at once, which neither term captures alone. |



### Degenerate columns are reported, not dropped



Two columns are degenerate in this corpus: `policy_c3` is entirely `NaN` and
`policy_n_steps` is constant, because **every one of the 124 cells uses a

two-step charging recipe**. The three-step parsing path is retained because it is
correct and general, not because this dataset exercises it.

They are deliberately **left in the matrix**. Dropping a column for being
constant across the whole cohort is a decision taken with sight of every cell,
including those destined for test folds -- a mild form of fitting on all the
data. The variance filter in [`selector.py`](../src/features/selector.py)
removes them **per fold**, on training-fold statistics only, which is where that
decision belongs.

## 6. Selection is confined to CV folds - structurally

With n = 124, selecting on the full dataset is the fastest way to manufacture a
result that does not replicate. Every selector is an sklearn transformer with a
`fit`/`transform` split, composed by `build_selection_pipeline()`. As **Pipeline

steps**, scikit-learn guarantees `fit` sees only the training fold -- the
guarantee does not depend on this project remembering to arrange it. There is
deliberately **no** module-level `select_features(X, y)` helper, because that is
the API shape that invites full-data use.

`configs/features.yaml` sets `selection.fit_inside_cv_only: true`, and the schema
**rejects** the config if it is ever false.

| Step | Purpose | Supervised? |
|---|---|---|
| `SimpleImputer` | Median fill. `keep_empty_features=True` so all-NaN columns survive to the variance filter instead of being silently dropped, which would misalign every downstream feature name. | no |
| `VarianceFilter` | Remove constant and all-missing columns. | no |
| `CorrelationFilter` | Drop one of each pair above rho = 0.95. Curve features are collinear by construction, and collinearity destabilises both linear coefficients and the Phase 9 SHAP attributions that must support an auditable scrap decision. | no |
| `StabilitySelector` | Keep features chosen in at least 60% of 100 bootstrap resamples of the training fold. | **yes** |
| `MutualInfoSelector` | Rank by mutual information, capturing monotone-but-nonlinear relationships. | **yes** |
| `StandardScaler` | Scale. | no |

Both supervised selectors **raise** if fitted without labels, so they cannot be
quietly misused on unlabelled full data.

### A silent bug found and fixed here

`StabilitySelector` originally ranked features with
`pd.Series(sub[c]).corr(pd.Series(sub_y))`. **pandas correlates on the index**,
and the bootstrap subset carries the original row labels while the target gets a
fresh `RangeIndex` -- so the two were silently mis-paired.

The effect: `dq_var`, whose true |rho| is **0.908**, scored **0.023**. The single
strongest predictor in the project was demoted to noise, and selection returned
an all-*thermal* feature set -- the group Phase 3 measured as the **weakest**
(|rho| ~ 0.22).

It raised no error and produced a plausible-looking result. It was caught only
because the output contradicted a Phase 3 measurement. The statistic is now
computed on NumPy arrays via `scipy.stats.spearmanr`, and after the fix selection
returns `dq_var` first, as it should.

## 7. Outputs

| File | Contents |
|---|---|
| `data/processed/features_budget{005,010,020,050,100}.parquet` | 124 cells x 57 features per budget |
| `outputs/reports/feature_registry.json` | Every feature with formula, rationale and in-line flag |
| `outputs/reports/feature_mask_report.json` | Exactly what the plausibility mask removed, per cell per budget |
| `data/processed/features_provenance.json` | Config hash and package versions |

```bash

python -m src.features.builder       # build all budgets

python -m src.features.make_docs     # regenerate this page

python -m pytest tests/test_leakage.py -q

```

---

[← Dataset and EDA](01_dataset_and_eda.md) · [README](../README.md) · [Modelling →](03_modeling.md)