# 02 - Feature engineering [Back to 01: Dataset and EDA](01_dataset_and_eda.md) | [README](../README.md) Covers **Phase 4**, the scientific core: 57 features across 7 groups, built at each of 5 diagnostic budgets under a structurally enforced no-leakage rule. The feature tables below are **generated from the code registry** ([`outputs/reports/feature_registry.json`](../outputs/reports/feature_registry.json)), so this document cannot drift from what is actually computed. --- ## 1. The leakage rule is enforced by construction For budget *N*, features may use only cycles 1..*N*. That is guaranteed structurally, not by discipline: 1. [`builder.py`](../src/features/builder.py) slices **all three** per-cell frames to `cycle_measured <= max_cycle` **once, at the top**. 2. It **asserts** the slice maximum does not exceed the budget. 3. It passes **only those slices** to the group modules. Group modules accept DataFrames -- never a cell id, never a path -- so they have no route back to unsliced data. `build_cell_features` is therefore a pure function of its frames, which is what makes the test below possible. ### The test has demonstrated teeth [`tests/test_leakage.py`](../tests/test_leakage.py) computes a cell's features, then **destroys every cycle beyond N** -- shuffling rows, rescaling values and adding noise, all three together because a statistic can be invariant to any one alone -- and demands bit-identical output. A passing test proves nothing unless it can fail, so the suite carries two controls: - **Negative control.** A deliberately leaky statistic *does* change under the corruption, confirming the corruption is effective. - **Mutation test.** The builder's slice is deliberately weakened to `3 x N`; **45 of 57 features change**, so the test detects the realistic regression. **A limitation found while building this, and closed.** The invariance test can only catch leaks that flow through the frames passed in. An early mutation attempt that read from a captured global was **not** detected: because the corruption produces copies, the leaky path saw identical data both times. A group reading from disk would evade the test entirely. That vector is now closed by a **static guard** asserting no feature module references `read_parquet`, `open(`, `CELLS_DIR` or `load_cell_frames`. ## 2. Cycle numbering: `cycle_measured`, not the raw index Phase 2 established that the all-zero placeholder row exists in batch 1 (46/46 cells) but not in batches 2-3 (0/89). Slicing "cycles 1..N" on the raw index would hand batch-1 cells one *fewer* measured cycle at every budget -- a systematic, batch-correlated bias in exactly the variable RQ4 tests for. Every feature therefore indexes on `cycle_measured`, which counts only measured cycles from 1, identically in every batch. It is attached to all three frames at parse time and re-derived after the continuation join. **One documented consequence.** The literature's "discharge capacity at cycle 2" is a raw index, which in batch 1 is the *first* measured cycle. Here `q_at_cycle_2` is the second *measured* cycle for every cell. That is a one-cycle offset from a naive raw reading in batch 1 only, it is uniform across the cohort, and it removes a batch-dependent inconsistency rather than inheriting one. It does not affect Gate 1, which reproduced at R^2 = 0.859 using measured cycles 10 and 100. ## 3. The plausibility mask Physically impossible readings become `NaN` **before any feature sees them**, using the same bounds the Phase 2 validator applied, applied identically to every cell. Masking by *physics* rather than by cell name is what makes this defensible: it is a stated rule about what a real measurement can be, not a patch for an inconvenient cell. Imputation happens **inside** the CV pipeline, on training-fold statistics only. ### What the mask caught, and why it mattered Phase 4 found a **third** corrupted channel on b1c18, beyond the two Phase 2 identified: its *median* cycle duration reads **2236 minutes (37 hours)** against a cohort median of **51.8 min**. This was not cosmetic. Thermal exposure is an *integral over time*, so a broken clock propagates straight into Group D. b1c18's `thermal_exposure` came out **71x the cohort median**, and that one feature drove a cross-validated ridge prediction to **-50.9** against a true value of 2.84, destroying an entire fold (RMSE 10.75 against ~0.07 for the others). The bound was derived from the data, not guessed: | Measurement (cycles 1-100, 135 cells) | Value | |---|---| | Median cycle duration, median across cells | 51.8 min | | 118 of 135 cells never exceed | 78.6 min | | 16 cells with one isolated long cycle | ~460 min (plausibly a real rest) | | **b1c18 median** | **2236 min** | `max_cycle_duration_minutes: 300` sits ~6x the median and ~4x above the ceiling of the 118 clean cells. It is applied **per cycle**, so a cell with one genuine long rest keeps its other 99. After the fix, cross-validated RMSE across budgets fell from 0.13-2.21 to **0.075-0.16** with consistent folds. *(Sanity check only -- Phase 6 owns modelling.)* ## 4. Corrections to the Phase 4 plan, from measurement Two flags raised in the plan did not survive contact with the data. Recorded rather than quietly dropped: | Plan said | Measurement says | |---|---| | Group A is **most exposed** to b1c18, whose voltage spans 0.736-6.606 V | **Wrong.** b1c18's `Qdlin` spans [-0.0004, 1.0523] Ah, indistinguishable from peers, and its `log_abs_dq_var` sits at the **60th percentile**. `Qdlin` derives from the *discharge* segment; b1c18's corruption is in the charge phase. | | The `fig04` DeltaQ(V) spike near 3.09 V is b1c18 | **Wrong.** It is **b1c41**, the only cell with a localized DeltaQ(V) spike (8.7 sigma). Corrected in [docs/01](01_dataset_and_eda.md). | Groups C and E were confirmed **unaffected** by b1c18: its internal resistance (0.0158-0.0199 Ohm) and charge time (9.87-9.97 min) are plausible and peer-consistent. ### A known fragility, left honest `dq_kurtosis` for **b1c41** is 112.9 against a cohort inter-quartile range of [-1.10, -0.30], because its DeltaQ(V) contains one localized spike. The feature is **correct** -- it faithfully reports a heavy-tailed curve -- and it has **not** been clipped. Suppressing it would be fitting the data to the method. The consequence is that *linear* models extrapolate badly from this cell. Phase 6 must therefore use robust preprocessing or outlier-insensitive estimators. `configs/models.yaml` already declares Huber regression and tree ensembles alongside the linear models, and this is a concrete reason to compare them rather than a stylistic preference. ## 5. Feature groups | Group | Module | Features | In-line | Process recipe | |---|---|---:|:--:|:--:| | **A_curve** | `curve_features.py` | 12 | yes | no | | **B_degradation** | `degradation.py` | 10 | yes | no | | **C_resistance** | `resistance.py` | 7 | yes | no | | **D_thermal** | `thermal.py` | 8 | NO | no | | **E_charge_dynamics** | `charge_dynamics.py` | 6 | yes | no | | **F_protocol** | `protocol.py` | 8 | yes | **yes** | | **G_interactions** | `interactions.py` | 6 | yes | no | | | **total** | **57** | | | ### Measurable in-line on a production line 45 features requiring no instrumentation beyond what a production cycler already records. | Feature | Group | Formula | Rationale | |---|---|---|---| | `dq_integral` | A_curve | sum(DeltaQ(V)) over the voltage grid | Total accumulated capacity difference across the discharge window; a magnitude complement to the shape statistics above. | | `dq_kurtosis` | A_curve | kurtosis(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. | | `dq_mean` | A_curve | mean(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. | | `dq_min` | A_curve | min(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. | | `dq_skew` | A_curve | skew(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. | | `dq_var` | A_curve | var(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. | | `log_abs_dq_kurtosis` | A_curve | log10(\|kurtosis(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. | | `log_abs_dq_mean` | A_curve | log10(\|mean(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. | | `log_abs_dq_min` | A_curve | log10(\|min(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. | | `log_abs_dq_skew` | A_curve | log10(\|skew(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. | | `log_abs_dq_var` | A_curve | log10(\|var(DeltaQ(V))\|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. | | `q_discharge_curve_area_N` | A_curve | sum(Q_N(V)) over the voltage grid | Absolute area under the interpolated discharge curve at the budget cycle, capturing overall accessible capacity independently of the difference curve. | | `cycle_at_q_max` | B_degradation | measured cycle at which Q_discharge peaks | When the trajectory turns over. An early peak means net fade has already overtaken activation. | | `q_at_cycle_2` | B_degradation | Q_discharge at measured cycle 2 | Early-life capacity, effectively the as-manufactured usable capacity once formation has settled. The reference point for all fade measures. | | `q_at_cycle_N` | B_degradation | Q_discharge at the budget cycle N | Capacity at the moment the QC decision must be made. | | `q_diff_N_2` | B_degradation | Q(N) - Q(2) | Absolute capacity change over the observation window. Small and noisy this early, which is precisely why curve-shape features outperform it. | | `q_fade_rate` | B_degradation | (Q(N) - Q(2)) / (N - 2) | Capacity loss per cycle, normalising the above by window length so budgets are comparable. | | `q_intercept` | B_degradation | least-squares intercept of Q(cycle) over the window | Fitted capacity at cycle 0; a noise-robust proxy for initial capacity. | | `q_max` | B_degradation | max Q_discharge over the window | Peak capacity. Many cells gain capacity for tens of cycles as formation completes before net fade begins. | | `q_max_minus_q2` | B_degradation | max(Q) - Q(2) | Size of that early capacity rise. A larger rise indicates continuing electrode wetting and activation. | | `q_slope` | B_degradation | least-squares slope of Q(cycle) over the window | Trend estimate using every cycle rather than just the endpoints, so it is far less sensitive to noise in any single measurement. | | `q_var` | B_degradation | variance of Q_discharge over the window | Trajectory roughness; elevated variance indicates unstable cycling behaviour. | | `ir_at_cycle_2` | C_resistance | internal resistance at measured cycle 2 | Post-formation baseline resistance, set by the interphase built during formation. A high starting value indicates lithium already consumed. | | `ir_at_cycle_N` | C_resistance | internal resistance at the budget cycle N | Resistance at the QC decision point. | | `ir_diff_N_2` | C_resistance | IR(N) - IR(2) | Resistance growth over the window: the direct signature of continuing interphase growth and contact loss. | | `ir_mean` | C_resistance | mean internal resistance over the window | Window-average level, robust to single-cycle measurement noise. | | `ir_min` | C_resistance | min internal resistance over the window | Resistance floor. Cells typically fall to a minimum as wetting completes before rising; the floor is a cleaner baseline than any single cycle. | | `ir_ratio_N_min` | C_resistance | IR(N) / min(IR) | Relative rise from the floor. Being dimensionless it is comparable across cells with different absolute resistance. | | `ir_slope` | C_resistance | least-squares slope of IR(cycle) over the window | Rate of resistance rise, using every cycle rather than two endpoints. | | `cc_time_fraction` | E_charge_dynamics | mean over cycles of (time in CC) / (total charge time) | Share of charging spent at constant current. As impedance rises the cell hits the voltage limit earlier and this fraction falls, making it a more specific impedance probe than total charge time. | | `charge_time_diff_N_2` | E_charge_dynamics | charge_time(N) - charge_time(2) | Endpoint change in charge duration over the window. | | `charge_time_mean` | E_charge_dynamics | mean charge time over measured cycles 2..N | Typical charge duration under a fixed protocol; a direct impedance proxy. | | `charge_time_slope` | E_charge_dynamics | least-squares slope of charge time vs cycle | Whether charge time is drifting upward, the signature of growing polarization resistance. | | `charge_time_var` | E_charge_dynamics | variance of charge time over measured cycles 2..N | Cycle-to-cycle instability in charge acceptance. | | `cv_time_fraction` | E_charge_dynamics | mean over cycles of (time in CV) / (total charge time) | Complement of the above; the constant-voltage tail lengthens as the cell becomes harder to charge. | | `policy_c1` | F_protocol | first-step charging C-rate | How aggressively the cell is charged at low state of charge, where lithium plating risk is highest. | | `policy_c2` | F_protocol | second-step charging C-rate | Charging rate through the middle of the state-of-charge range. | | `policy_c3` | F_protocol | third-step charging C-rate where the recipe defines one | Present only for multi-step recipes; NaN otherwise, and imputed inside the CV pipeline rather than filled with a fabricated rate. | | `policy_c_max` | F_protocol | maximum C-rate used anywhere in the recipe | Peak stress the cell experiences during charging. | | `policy_c_mean` | F_protocol | SOC-weighted mean C-rate across the recipe steps | Single summary of overall charging aggressiveness, weighting each rate by the fraction of charge delivered at it. | | `policy_c_range` | F_protocol | policy_c_max - min C-rate in the recipe | Spread between the gentlest and harshest step; a flat recipe and a steeply stepped one can share a mean but stress the cell very differently. | | `policy_n_steps` | F_protocol | number of steps in the recipe | Recipe complexity, distinguishing simple two-step from multi-step profiles. | | `policy_q1_pct` | F_protocol | state of charge (%) at which the protocol switches step | Where the recipe hands over between rates; determines how much of the high-plating-risk region is traversed at the first rate. | | `charge_time_slope_x_ir_slope` | G_interactions | charge_time_slope * ir_slope | Both are proxies for growing polarization resistance measured through different channels. Their product is large only when the two agree, suppressing the case where either drifts through measurement noise alone. | | `ir_growth_per_cycle` | G_interactions | ir_diff_N_2 / (N - 2) | Resistance growth normalised by window length so that budgets are directly comparable; without it the raw difference grows mechanically with N. | ### NOT reliably measurable in-line - deployability caveat 12 features needing per-cell instrumentation that many lines do not have. A screening rule depending on these may not be deployable without additional hardware, which is why the distinction is tracked in the registry rather than left implicit. | Feature | Group | Formula | Rationale | |---|---|---|---| | `temp_avg_slope` | D_thermal | least-squares slope of per-cycle mean temperature vs cycle | Whether the cell is progressively self-heating, which indicates rising internal impedance dissipating more energy. | | `temp_max` | D_thermal | max of per-cycle maximum temperature over the window | Worst-case thermal excursion, which drives the fastest side reactions. | | `temp_max_slope` | D_thermal | least-squares slope of per-cycle maximum temperature vs cycle | The same trend measured at the thermal peak, where it appears earliest. | | `temp_mean` | D_thermal | mean of per-cycle average temperature over the window | Typical operating temperature of the cell. | | `temp_min` | D_thermal | min of per-cycle minimum temperature over the window | Lower bound of the thermal envelope; with the maximum it gives the range. | | `temp_range` | D_thermal | temp_max - temp_min | Thermal swing per cycle, a proxy for heat generated under load relative to the chamber setpoint. | | `thermal_exposure` | D_thermal | sum over cycles of the integral of T dt within each cycle | Cumulative time-at-temperature, the quantity an Arrhenius rate law integrates. Computed from the within-cycle traces rather than from per-cycle averages so that cycles of differing duration are weighted correctly. | | `thermal_exposure_per_cycle` | D_thermal | thermal_exposure / number of measured cycles | Rate rather than accumulation, so budgets are comparable. | | `fade_per_thermal_exposure` | G_interactions | q_fade_rate / thermal_exposure | Capacity loss normalised by thermal load. A cell fading fast despite low thermal exposure is degrading for a non-thermal reason, which is diagnostically distinct. | | `log_dq_var_x_ir_ratio` | G_interactions | log_abs_dq_var * ir_ratio_N_min | Combines loss of accessible lithium inventory (curve shape) with rising interfacial impedance, distinguishing cells losing capacity through inventory loss from those losing it through impedance rise. | | `recipe_stress_x_thermal` | G_interactions | policy_c_max * thermal_exposure | Higher charging C-rates drive both ohmic heating and lithium plating. This links the commanded process stress to the thermal response it actually produced. PROCESS-RECIPE DERIVED: excluded from the no_recipe feature set with Group F. | | `thermal_x_ir_growth` | G_interactions | thermal_exposure * ir_diff_N_2 | Side reactions are Arrhenius-activated AND proceed at the interphase. A cell that is both hot and building resistance is degrading by both mechanisms at once, which neither term captures alone. | ### Degenerate columns are reported, not dropped Two columns are degenerate in this corpus: `policy_c3` is entirely `NaN` and `policy_n_steps` is constant, because **every one of the 124 cells uses a two-step charging recipe**. The three-step parsing path is retained because it is correct and general, not because this dataset exercises it. They are deliberately **left in the matrix**. Dropping a column for being constant across the whole cohort is a decision taken with sight of every cell, including those destined for test folds -- a mild form of fitting on all the data. The variance filter in [`selector.py`](../src/features/selector.py) removes them **per fold**, on training-fold statistics only, which is where that decision belongs. ## 6. Selection is confined to CV folds - structurally With n = 124, selecting on the full dataset is the fastest way to manufacture a result that does not replicate. Every selector is an sklearn transformer with a `fit`/`transform` split, composed by `build_selection_pipeline()`. As **Pipeline steps**, scikit-learn guarantees `fit` sees only the training fold -- the guarantee does not depend on this project remembering to arrange it. There is deliberately **no** module-level `select_features(X, y)` helper, because that is the API shape that invites full-data use. `configs/features.yaml` sets `selection.fit_inside_cv_only: true`, and the schema **rejects** the config if it is ever false. | Step | Purpose | Supervised? | |---|---|---| | `SimpleImputer` | Median fill. `keep_empty_features=True` so all-NaN columns survive to the variance filter instead of being silently dropped, which would misalign every downstream feature name. | no | | `VarianceFilter` | Remove constant and all-missing columns. | no | | `CorrelationFilter` | Drop one of each pair above rho = 0.95. Curve features are collinear by construction, and collinearity destabilises both linear coefficients and the Phase 9 SHAP attributions that must support an auditable scrap decision. | no | | `StabilitySelector` | Keep features chosen in at least 60% of 100 bootstrap resamples of the training fold. | **yes** | | `MutualInfoSelector` | Rank by mutual information, capturing monotone-but-nonlinear relationships. | **yes** | | `StandardScaler` | Scale. | no | Both supervised selectors **raise** if fitted without labels, so they cannot be quietly misused on unlabelled full data. ### A silent bug found and fixed here `StabilitySelector` originally ranked features with `pd.Series(sub[c]).corr(pd.Series(sub_y))`. **pandas correlates on the index**, and the bootstrap subset carries the original row labels while the target gets a fresh `RangeIndex` -- so the two were silently mis-paired. The effect: `dq_var`, whose true |rho| is **0.908**, scored **0.023**. The single strongest predictor in the project was demoted to noise, and selection returned an all-*thermal* feature set -- the group Phase 3 measured as the **weakest** (|rho| ~ 0.22). It raised no error and produced a plausible-looking result. It was caught only because the output contradicted a Phase 3 measurement. The statistic is now computed on NumPy arrays via `scipy.stats.spearmanr`, and after the fix selection returns `dq_var` first, as it should. ## 7. Outputs | File | Contents | |---|---| | `data/processed/features_budget{005,010,020,050,100}.parquet` | 124 cells x 57 features per budget | | `outputs/reports/feature_registry.json` | Every feature with formula, rationale and in-line flag | | `outputs/reports/feature_mask_report.json` | Exactly what the plausibility mask removed, per cell per budget | | `data/processed/features_provenance.json` | Config hash and package versions | ```bash python -m src.features.builder # build all budgets python -m src.features.make_docs # regenerate this page python -m pytest tests/test_leakage.py -q ``` --- [← Dataset and EDA](01_dataset_and_eda.md) · [README](../README.md) · [Modelling →](03_modeling.md)