Spaces:
Running on Zero
Running on Zero
File size: 24,286 Bytes
749bffa | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 | # 02 - Feature engineering
[Back to 01: Dataset and EDA](01_dataset_and_eda.md) | [README](../README.md)
Covers **Phase 4**, the scientific core: 57 features across 7 groups, built at
each of 5 diagnostic budgets under a structurally enforced no-leakage rule.
The feature tables below are **generated from the code registry**
([`outputs/reports/feature_registry.json`](../outputs/reports/feature_registry.json)),
so this document cannot drift from what is actually computed.
---
## 1. The leakage rule is enforced by construction
For budget *N*, features may use only cycles 1..*N*. That is guaranteed
structurally, not by discipline:
1. [`builder.py`](../src/features/builder.py) slices **all three** per-cell
frames to `cycle_measured <= max_cycle` **once, at the top**.
2. It **asserts** the slice maximum does not exceed the budget.
3. It passes **only those slices** to the group modules.
Group modules accept DataFrames -- never a cell id, never a path -- so they have
no route back to unsliced data. `build_cell_features` is therefore a pure
function of its frames, which is what makes the test below possible.
### The test has demonstrated teeth
[`tests/test_leakage.py`](../tests/test_leakage.py) computes a cell's features,
then **destroys every cycle beyond N** -- shuffling rows, rescaling values and
adding noise, all three together because a statistic can be invariant to any one
alone -- and demands bit-identical output.
A passing test proves nothing unless it can fail, so the suite carries two
controls:
- **Negative control.** A deliberately leaky statistic *does* change under the
corruption, confirming the corruption is effective.
- **Mutation test.** The builder's slice is deliberately weakened to `3 x N`;
**45 of 57 features change**, so the test detects the realistic regression.
**A limitation found while building this, and closed.** The invariance test can
only catch leaks that flow through the frames passed in. An early mutation
attempt that read from a captured global was **not** detected: because the
corruption produces copies, the leaky path saw identical data both times. A
group reading from disk would evade the test entirely. That vector is now closed
by a **static guard** asserting no feature module references `read_parquet`,
`open(`, `CELLS_DIR` or `load_cell_frames`.
## 2. Cycle numbering: `cycle_measured`, not the raw index
Phase 2 established that the all-zero placeholder row exists in batch 1 (46/46
cells) but not in batches 2-3 (0/89). Slicing "cycles 1..N" on the raw index
would hand batch-1 cells one *fewer* measured cycle at every budget -- a
systematic, batch-correlated bias in exactly the variable RQ4 tests for.
Every feature therefore indexes on `cycle_measured`, which counts only measured
cycles from 1, identically in every batch. It is attached to all three frames at
parse time and re-derived after the continuation join.
**One documented consequence.** The literature's "discharge capacity at cycle 2"
is a raw index, which in batch 1 is the *first* measured cycle. Here
`q_at_cycle_2` is the second *measured* cycle for every cell. That is a
one-cycle offset from a naive raw reading in batch 1 only, it is uniform across
the cohort, and it removes a batch-dependent inconsistency rather than
inheriting one. It does not affect Gate 1, which reproduced at R^2 = 0.859 using
measured cycles 10 and 100.
## 3. The plausibility mask
Physically impossible readings become `NaN` **before any feature sees them**,
using the same bounds the Phase 2 validator applied, applied identically to
every cell. Masking by *physics* rather than by cell name is what makes this
defensible: it is a stated rule about what a real measurement can be, not a
patch for an inconvenient cell. Imputation happens **inside** the CV pipeline,
on training-fold statistics only.
### What the mask caught, and why it mattered
Phase 4 found a **third** corrupted channel on b1c18, beyond the two Phase 2
identified: its *median* cycle duration reads **2236 minutes (37 hours)** against
a cohort median of **51.8 min**.
This was not cosmetic. Thermal exposure is an *integral over time*, so a broken
clock propagates straight into Group D. b1c18's `thermal_exposure` came out
**71x the cohort median**, and that one feature drove a cross-validated ridge
prediction to **-50.9** against a true value of 2.84, destroying an entire fold
(RMSE 10.75 against ~0.07 for the others).
The bound was derived from the data, not guessed:
| Measurement (cycles 1-100, 135 cells) | Value |
|---|---|
| Median cycle duration, median across cells | 51.8 min |
| 118 of 135 cells never exceed | 78.6 min |
| 16 cells with one isolated long cycle | ~460 min (plausibly a real rest) |
| **b1c18 median** | **2236 min** |
`max_cycle_duration_minutes: 300` sits ~6x the median and ~4x above the ceiling
of the 118 clean cells. It is applied **per cycle**, so a cell with one genuine
long rest keeps its other 99.
After the fix, cross-validated RMSE across budgets fell from 0.13-2.21 to
**0.075-0.16** with consistent folds. *(Sanity check only -- Phase 6 owns
modelling.)*
## 4. Corrections to the Phase 4 plan, from measurement
Two flags raised in the plan did not survive contact with the data. Recorded
rather than quietly dropped:
| Plan said | Measurement says |
|---|---|
| Group A is **most exposed** to b1c18, whose voltage spans 0.736-6.606 V | **Wrong.** b1c18's `Qdlin` spans [-0.0004, 1.0523] Ah, indistinguishable from peers, and its `log_abs_dq_var` sits at the **60th percentile**. `Qdlin` derives from the *discharge* segment; b1c18's corruption is in the charge phase. |
| The `fig04` DeltaQ(V) spike near 3.09 V is b1c18 | **Wrong.** It is **b1c41**, the only cell with a localized DeltaQ(V) spike (8.7 sigma). Corrected in [docs/01](01_dataset_and_eda.md). |
Groups C and E were confirmed **unaffected** by b1c18: its internal resistance
(0.0158-0.0199 Ohm) and charge time (9.87-9.97 min) are plausible and
peer-consistent.
### A known fragility, left honest
`dq_kurtosis` for **b1c41** is 112.9 against a cohort inter-quartile range of
[-1.10, -0.30], because its DeltaQ(V) contains one localized spike. The feature
is **correct** -- it faithfully reports a heavy-tailed curve -- and it has
**not** been clipped. Suppressing it would be fitting the data to the method.
The consequence is that *linear* models extrapolate badly from this cell.
Phase 6 must therefore use robust preprocessing or outlier-insensitive
estimators. `configs/models.yaml` already declares Huber regression and tree
ensembles alongside the linear models, and this is a concrete reason to compare
them rather than a stylistic preference.
## 5. Feature groups
| Group | Module | Features | In-line | Process recipe |
|---|---|---:|:--:|:--:|
| **A_curve** | `curve_features.py` | 12 | yes | no |
| **B_degradation** | `degradation.py` | 10 | yes | no |
| **C_resistance** | `resistance.py` | 7 | yes | no |
| **D_thermal** | `thermal.py` | 8 | NO | no |
| **E_charge_dynamics** | `charge_dynamics.py` | 6 | yes | no |
| **F_protocol** | `protocol.py` | 8 | yes | **yes** |
| **G_interactions** | `interactions.py` | 6 | yes | no |
| | **total** | **57** | | |
### Measurable in-line on a production line
45 features requiring no instrumentation beyond what a production cycler already records.
| Feature | Group | Formula | Rationale |
|---|---|---|---|
| `dq_integral` | A_curve | sum(DeltaQ(V)) over the voltage grid | Total accumulated capacity difference across the discharge window; a magnitude complement to the shape statistics above. |
| `dq_kurtosis` | A_curve | kurtosis(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
| `dq_mean` | A_curve | mean(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
| `dq_min` | A_curve | min(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
| `dq_skew` | A_curve | skew(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
| `dq_var` | A_curve | var(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
| `log_abs_dq_kurtosis` | A_curve | log10(\|kurtosis(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
| `log_abs_dq_mean` | A_curve | log10(\|mean(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
| `log_abs_dq_min` | A_curve | log10(\|min(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
| `log_abs_dq_skew` | A_curve | log10(\|skew(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
| `log_abs_dq_var` | A_curve | log10(\|var(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
| `q_discharge_curve_area_N` | A_curve | sum(Q_N(V)) over the voltage grid | Absolute area under the interpolated discharge curve at the budget cycle, capturing overall accessible capacity independently of the difference curve. |
| `cycle_at_q_max` | B_degradation | measured cycle at which Q_discharge peaks | When the trajectory turns over. An early peak means net fade has already overtaken activation. |
| `q_at_cycle_2` | B_degradation | Q_discharge at measured cycle 2 | Early-life capacity, effectively the as-manufactured usable capacity once formation has settled. The reference point for all fade measures. |
| `q_at_cycle_N` | B_degradation | Q_discharge at the budget cycle N | Capacity at the moment the QC decision must be made. |
| `q_diff_N_2` | B_degradation | Q(N) - Q(2) | Absolute capacity change over the observation window. Small and noisy this early, which is precisely why curve-shape features outperform it. |
| `q_fade_rate` | B_degradation | (Q(N) - Q(2)) / (N - 2) | Capacity loss per cycle, normalising the above by window length so budgets are comparable. |
| `q_intercept` | B_degradation | least-squares intercept of Q(cycle) over the window | Fitted capacity at cycle 0; a noise-robust proxy for initial capacity. |
| `q_max` | B_degradation | max Q_discharge over the window | Peak capacity. Many cells gain capacity for tens of cycles as formation completes before net fade begins. |
| `q_max_minus_q2` | B_degradation | max(Q) - Q(2) | Size of that early capacity rise. A larger rise indicates continuing electrode wetting and activation. |
| `q_slope` | B_degradation | least-squares slope of Q(cycle) over the window | Trend estimate using every cycle rather than just the endpoints, so it is far less sensitive to noise in any single measurement. |
| `q_var` | B_degradation | variance of Q_discharge over the window | Trajectory roughness; elevated variance indicates unstable cycling behaviour. |
| `ir_at_cycle_2` | C_resistance | internal resistance at measured cycle 2 | Post-formation baseline resistance, set by the interphase built during formation. A high starting value indicates lithium already consumed. |
| `ir_at_cycle_N` | C_resistance | internal resistance at the budget cycle N | Resistance at the QC decision point. |
| `ir_diff_N_2` | C_resistance | IR(N) - IR(2) | Resistance growth over the window: the direct signature of continuing interphase growth and contact loss. |
| `ir_mean` | C_resistance | mean internal resistance over the window | Window-average level, robust to single-cycle measurement noise. |
| `ir_min` | C_resistance | min internal resistance over the window | Resistance floor. Cells typically fall to a minimum as wetting completes before rising; the floor is a cleaner baseline than any single cycle. |
| `ir_ratio_N_min` | C_resistance | IR(N) / min(IR) | Relative rise from the floor. Being dimensionless it is comparable across cells with different absolute resistance. |
| `ir_slope` | C_resistance | least-squares slope of IR(cycle) over the window | Rate of resistance rise, using every cycle rather than two endpoints. |
| `cc_time_fraction` | E_charge_dynamics | mean over cycles of (time in CC) / (total charge time) | Share of charging spent at constant current. As impedance rises the cell hits the voltage limit earlier and this fraction falls, making it a more specific impedance probe than total charge time. |
| `charge_time_diff_N_2` | E_charge_dynamics | charge_time(N) - charge_time(2) | Endpoint change in charge duration over the window. |
| `charge_time_mean` | E_charge_dynamics | mean charge time over measured cycles 2..N | Typical charge duration under a fixed protocol; a direct impedance proxy. |
| `charge_time_slope` | E_charge_dynamics | least-squares slope of charge time vs cycle | Whether charge time is drifting upward, the signature of growing polarization resistance. |
| `charge_time_var` | E_charge_dynamics | variance of charge time over measured cycles 2..N | Cycle-to-cycle instability in charge acceptance. |
| `cv_time_fraction` | E_charge_dynamics | mean over cycles of (time in CV) / (total charge time) | Complement of the above; the constant-voltage tail lengthens as the cell becomes harder to charge. |
| `policy_c1` | F_protocol | first-step charging C-rate | How aggressively the cell is charged at low state of charge, where lithium plating risk is highest. |
| `policy_c2` | F_protocol | second-step charging C-rate | Charging rate through the middle of the state-of-charge range. |
| `policy_c3` | F_protocol | third-step charging C-rate where the recipe defines one | Present only for multi-step recipes; NaN otherwise, and imputed inside the CV pipeline rather than filled with a fabricated rate. |
| `policy_c_max` | F_protocol | maximum C-rate used anywhere in the recipe | Peak stress the cell experiences during charging. |
| `policy_c_mean` | F_protocol | SOC-weighted mean C-rate across the recipe steps | Single summary of overall charging aggressiveness, weighting each rate by the fraction of charge delivered at it. |
| `policy_c_range` | F_protocol | policy_c_max - min C-rate in the recipe | Spread between the gentlest and harshest step; a flat recipe and a steeply stepped one can share a mean but stress the cell very differently. |
| `policy_n_steps` | F_protocol | number of steps in the recipe | Recipe complexity, distinguishing simple two-step from multi-step profiles. |
| `policy_q1_pct` | F_protocol | state of charge (%) at which the protocol switches step | Where the recipe hands over between rates; determines how much of the high-plating-risk region is traversed at the first rate. |
| `charge_time_slope_x_ir_slope` | G_interactions | charge_time_slope * ir_slope | Both are proxies for growing polarization resistance measured through different channels. Their product is large only when the two agree, suppressing the case where either drifts through measurement noise alone. |
| `ir_growth_per_cycle` | G_interactions | ir_diff_N_2 / (N - 2) | Resistance growth normalised by window length so that budgets are directly comparable; without it the raw difference grows mechanically with N. |
### NOT reliably measurable in-line - deployability caveat
12 features needing per-cell instrumentation that many lines do not have. A screening rule depending on these may not be deployable without additional hardware, which is why the distinction is tracked in the registry rather than left implicit.
| Feature | Group | Formula | Rationale |
|---|---|---|---|
| `temp_avg_slope` | D_thermal | least-squares slope of per-cycle mean temperature vs cycle | Whether the cell is progressively self-heating, which indicates rising internal impedance dissipating more energy. |
| `temp_max` | D_thermal | max of per-cycle maximum temperature over the window | Worst-case thermal excursion, which drives the fastest side reactions. |
| `temp_max_slope` | D_thermal | least-squares slope of per-cycle maximum temperature vs cycle | The same trend measured at the thermal peak, where it appears earliest. |
| `temp_mean` | D_thermal | mean of per-cycle average temperature over the window | Typical operating temperature of the cell. |
| `temp_min` | D_thermal | min of per-cycle minimum temperature over the window | Lower bound of the thermal envelope; with the maximum it gives the range. |
| `temp_range` | D_thermal | temp_max - temp_min | Thermal swing per cycle, a proxy for heat generated under load relative to the chamber setpoint. |
| `thermal_exposure` | D_thermal | sum over cycles of the integral of T dt within each cycle | Cumulative time-at-temperature, the quantity an Arrhenius rate law integrates. Computed from the within-cycle traces rather than from per-cycle averages so that cycles of differing duration are weighted correctly. |
| `thermal_exposure_per_cycle` | D_thermal | thermal_exposure / number of measured cycles | Rate rather than accumulation, so budgets are comparable. |
| `fade_per_thermal_exposure` | G_interactions | q_fade_rate / thermal_exposure | Capacity loss normalised by thermal load. A cell fading fast despite low thermal exposure is degrading for a non-thermal reason, which is diagnostically distinct. |
| `log_dq_var_x_ir_ratio` | G_interactions | log_abs_dq_var * ir_ratio_N_min | Combines loss of accessible lithium inventory (curve shape) with rising interfacial impedance, distinguishing cells losing capacity through inventory loss from those losing it through impedance rise. |
| `recipe_stress_x_thermal` | G_interactions | policy_c_max * thermal_exposure | Higher charging C-rates drive both ohmic heating and lithium plating. This links the commanded process stress to the thermal response it actually produced. PROCESS-RECIPE DERIVED: excluded from the no_recipe feature set with Group F. |
| `thermal_x_ir_growth` | G_interactions | thermal_exposure * ir_diff_N_2 | Side reactions are Arrhenius-activated AND proceed at the interphase. A cell that is both hot and building resistance is degrading by both mechanisms at once, which neither term captures alone. |
### Degenerate columns are reported, not dropped
Two columns are degenerate in this corpus: `policy_c3` is entirely `NaN` and
`policy_n_steps` is constant, because **every one of the 124 cells uses a
two-step charging recipe**. The three-step parsing path is retained because it is
correct and general, not because this dataset exercises it.
They are deliberately **left in the matrix**. Dropping a column for being
constant across the whole cohort is a decision taken with sight of every cell,
including those destined for test folds -- a mild form of fitting on all the
data. The variance filter in [`selector.py`](../src/features/selector.py)
removes them **per fold**, on training-fold statistics only, which is where that
decision belongs.
## 6. Selection is confined to CV folds - structurally
With n = 124, selecting on the full dataset is the fastest way to manufacture a
result that does not replicate. Every selector is an sklearn transformer with a
`fit`/`transform` split, composed by `build_selection_pipeline()`. As **Pipeline
steps**, scikit-learn guarantees `fit` sees only the training fold -- the
guarantee does not depend on this project remembering to arrange it. There is
deliberately **no** module-level `select_features(X, y)` helper, because that is
the API shape that invites full-data use.
`configs/features.yaml` sets `selection.fit_inside_cv_only: true`, and the schema
**rejects** the config if it is ever false.
| Step | Purpose | Supervised? |
|---|---|---|
| `SimpleImputer` | Median fill. `keep_empty_features=True` so all-NaN columns survive to the variance filter instead of being silently dropped, which would misalign every downstream feature name. | no |
| `VarianceFilter` | Remove constant and all-missing columns. | no |
| `CorrelationFilter` | Drop one of each pair above rho = 0.95. Curve features are collinear by construction, and collinearity destabilises both linear coefficients and the Phase 9 SHAP attributions that must support an auditable scrap decision. | no |
| `StabilitySelector` | Keep features chosen in at least 60% of 100 bootstrap resamples of the training fold. | **yes** |
| `MutualInfoSelector` | Rank by mutual information, capturing monotone-but-nonlinear relationships. | **yes** |
| `StandardScaler` | Scale. | no |
Both supervised selectors **raise** if fitted without labels, so they cannot be
quietly misused on unlabelled full data.
### A silent bug found and fixed here
`StabilitySelector` originally ranked features with
`pd.Series(sub[c]).corr(pd.Series(sub_y))`. **pandas correlates on the index**,
and the bootstrap subset carries the original row labels while the target gets a
fresh `RangeIndex` -- so the two were silently mis-paired.
The effect: `dq_var`, whose true |rho| is **0.908**, scored **0.023**. The single
strongest predictor in the project was demoted to noise, and selection returned
an all-*thermal* feature set -- the group Phase 3 measured as the **weakest**
(|rho| ~ 0.22).
It raised no error and produced a plausible-looking result. It was caught only
because the output contradicted a Phase 3 measurement. The statistic is now
computed on NumPy arrays via `scipy.stats.spearmanr`, and after the fix selection
returns `dq_var` first, as it should.
## 7. Outputs
| File | Contents |
|---|---|
| `data/processed/features_budget{005,010,020,050,100}.parquet` | 124 cells x 57 features per budget |
| `outputs/reports/feature_registry.json` | Every feature with formula, rationale and in-line flag |
| `outputs/reports/feature_mask_report.json` | Exactly what the plausibility mask removed, per cell per budget |
| `data/processed/features_provenance.json` | Config hash and package versions |
```bash
python -m src.features.builder # build all budgets
python -m src.features.make_docs # regenerate this page
python -m pytest tests/test_leakage.py -q
```
---
[← Dataset and EDA](01_dataset_and_eda.md) · [README](../README.md) · [Modelling →](03_modeling.md)
|