Spaces:
Running on Zero
02 - Feature engineering
Back to 01: Dataset and EDA | README
Covers Phase 4, the scientific core: 57 features across 7 groups, built at each of 5 diagnostic budgets under a structurally enforced no-leakage rule.
The feature tables below are generated from the code registry
(outputs/reports/feature_registry.json),
so this document cannot drift from what is actually computed.
1. The leakage rule is enforced by construction
For budget N, features may use only cycles 1..N. That is guaranteed structurally, not by discipline:
builder.pyslices all three per-cell frames tocycle_measured <= max_cycleonce, at the top.- It asserts the slice maximum does not exceed the budget.
- It passes only those slices to the group modules.
Group modules accept DataFrames -- never a cell id, never a path -- so they have
no route back to unsliced data. build_cell_features is therefore a pure
function of its frames, which is what makes the test below possible.
The test has demonstrated teeth
tests/test_leakage.py computes a cell's features,
then destroys every cycle beyond N -- shuffling rows, rescaling values and
adding noise, all three together because a statistic can be invariant to any one
alone -- and demands bit-identical output.
A passing test proves nothing unless it can fail, so the suite carries two controls:
- Negative control. A deliberately leaky statistic does change under the corruption, confirming the corruption is effective.
- Mutation test. The builder's slice is deliberately weakened to
3 x N; 45 of 57 features change, so the test detects the realistic regression.
A limitation found while building this, and closed. The invariance test can
only catch leaks that flow through the frames passed in. An early mutation
attempt that read from a captured global was not detected: because the
corruption produces copies, the leaky path saw identical data both times. A
group reading from disk would evade the test entirely. That vector is now closed
by a static guard asserting no feature module references read_parquet,
open(, CELLS_DIR or load_cell_frames.
2. Cycle numbering: cycle_measured, not the raw index
Phase 2 established that the all-zero placeholder row exists in batch 1 (46/46 cells) but not in batches 2-3 (0/89). Slicing "cycles 1..N" on the raw index would hand batch-1 cells one fewer measured cycle at every budget -- a systematic, batch-correlated bias in exactly the variable RQ4 tests for.
Every feature therefore indexes on cycle_measured, which counts only measured
cycles from 1, identically in every batch. It is attached to all three frames at
parse time and re-derived after the continuation join.
One documented consequence. The literature's "discharge capacity at cycle 2"
is a raw index, which in batch 1 is the first measured cycle. Here
q_at_cycle_2 is the second measured cycle for every cell. That is a
one-cycle offset from a naive raw reading in batch 1 only, it is uniform across
the cohort, and it removes a batch-dependent inconsistency rather than
inheriting one. It does not affect Gate 1, which reproduced at R^2 = 0.859 using
measured cycles 10 and 100.
3. The plausibility mask
Physically impossible readings become NaN before any feature sees them,
using the same bounds the Phase 2 validator applied, applied identically to
every cell. Masking by physics rather than by cell name is what makes this
defensible: it is a stated rule about what a real measurement can be, not a
patch for an inconvenient cell. Imputation happens inside the CV pipeline,
on training-fold statistics only.
What the mask caught, and why it mattered
Phase 4 found a third corrupted channel on b1c18, beyond the two Phase 2 identified: its median cycle duration reads 2236 minutes (37 hours) against a cohort median of 51.8 min.
This was not cosmetic. Thermal exposure is an integral over time, so a broken
clock propagates straight into Group D. b1c18's thermal_exposure came out
71x the cohort median, and that one feature drove a cross-validated ridge
prediction to -50.9 against a true value of 2.84, destroying an entire fold
(RMSE 10.75 against ~0.07 for the others).
The bound was derived from the data, not guessed:
| Measurement (cycles 1-100, 135 cells) | Value |
|---|---|
| Median cycle duration, median across cells | 51.8 min |
| 118 of 135 cells never exceed | 78.6 min |
| 16 cells with one isolated long cycle | ~460 min (plausibly a real rest) |
| b1c18 median | 2236 min |
max_cycle_duration_minutes: 300 sits ~6x the median and ~4x above the ceiling
of the 118 clean cells. It is applied per cycle, so a cell with one genuine
long rest keeps its other 99.
After the fix, cross-validated RMSE across budgets fell from 0.13-2.21 to 0.075-0.16 with consistent folds. (Sanity check only -- Phase 6 owns modelling.)
4. Corrections to the Phase 4 plan, from measurement
Two flags raised in the plan did not survive contact with the data. Recorded rather than quietly dropped:
| Plan said | Measurement says |
|---|---|
| Group A is most exposed to b1c18, whose voltage spans 0.736-6.606 V | Wrong. b1c18's Qdlin spans [-0.0004, 1.0523] Ah, indistinguishable from peers, and its log_abs_dq_var sits at the 60th percentile. Qdlin derives from the discharge segment; b1c18's corruption is in the charge phase. |
The fig04 DeltaQ(V) spike near 3.09 V is b1c18 |
Wrong. It is b1c41, the only cell with a localized DeltaQ(V) spike (8.7 sigma). Corrected in docs/01. |
Groups C and E were confirmed unaffected by b1c18: its internal resistance (0.0158-0.0199 Ohm) and charge time (9.87-9.97 min) are plausible and peer-consistent.
A known fragility, left honest
dq_kurtosis for b1c41 is 112.9 against a cohort inter-quartile range of
[-1.10, -0.30], because its DeltaQ(V) contains one localized spike. The feature
is correct -- it faithfully reports a heavy-tailed curve -- and it has
not been clipped. Suppressing it would be fitting the data to the method.
The consequence is that linear models extrapolate badly from this cell.
Phase 6 must therefore use robust preprocessing or outlier-insensitive
estimators. configs/models.yaml already declares Huber regression and tree
ensembles alongside the linear models, and this is a concrete reason to compare
them rather than a stylistic preference.
5. Feature groups
| Group | Module | Features | In-line | Process recipe |
|---|---|---|---|---|
| A_curve | curve_features.py |
12 | yes | no |
| B_degradation | degradation.py |
10 | yes | no |
| C_resistance | resistance.py |
7 | yes | no |
| D_thermal | thermal.py |
8 | NO | no |
| E_charge_dynamics | charge_dynamics.py |
6 | yes | no |
| F_protocol | protocol.py |
8 | yes | yes |
| G_interactions | interactions.py |
6 | yes | no |
| total | 57 |
Measurable in-line on a production line
45 features requiring no instrumentation beyond what a production cycler already records.
| Feature | Group | Formula | Rationale |
|---|---|---|---|
dq_integral |
A_curve | sum(DeltaQ(V)) over the voltage grid | Total accumulated capacity difference across the discharge window; a magnitude complement to the shape statistics above. |
dq_kurtosis |
A_curve | kurtosis(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
dq_mean |
A_curve | mean(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
dq_min |
A_curve | min(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
dq_skew |
A_curve | skew(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
dq_var |
A_curve | var(DeltaQ(V)) where DeltaQ(V) = Q_N(V) - Q_base(V) on the common voltage grid | Shape of the capacity-voltage difference curve. Variance is the single strongest published early-life predictor; minimum locates the largest local capacity loss; skewness and kurtosis describe whether that loss is concentrated at particular voltages, which distinguishes degradation modes. |
log_abs_dq_kurtosis |
A_curve | log10(|kurtosis(DeltaQ(V))|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
log_abs_dq_mean |
A_curve | log10(|mean(DeltaQ(V))|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
log_abs_dq_min |
A_curve | log10(|min(DeltaQ(V))|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
log_abs_dq_skew |
A_curve | log10(|skew(DeltaQ(V))|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
log_abs_dq_var |
A_curve | log10(|var(DeltaQ(V))|) | Log scale. These quantities span several orders of magnitude across cells; the source literature models them in log space and an unlogged feature lets a single high-variance cell dominate a linear fit. |
q_discharge_curve_area_N |
A_curve | sum(Q_N(V)) over the voltage grid | Absolute area under the interpolated discharge curve at the budget cycle, capturing overall accessible capacity independently of the difference curve. |
cycle_at_q_max |
B_degradation | measured cycle at which Q_discharge peaks | When the trajectory turns over. An early peak means net fade has already overtaken activation. |
q_at_cycle_2 |
B_degradation | Q_discharge at measured cycle 2 | Early-life capacity, effectively the as-manufactured usable capacity once formation has settled. The reference point for all fade measures. |
q_at_cycle_N |
B_degradation | Q_discharge at the budget cycle N | Capacity at the moment the QC decision must be made. |
q_diff_N_2 |
B_degradation | Q(N) - Q(2) | Absolute capacity change over the observation window. Small and noisy this early, which is precisely why curve-shape features outperform it. |
q_fade_rate |
B_degradation | (Q(N) - Q(2)) / (N - 2) | Capacity loss per cycle, normalising the above by window length so budgets are comparable. |
q_intercept |
B_degradation | least-squares intercept of Q(cycle) over the window | Fitted capacity at cycle 0; a noise-robust proxy for initial capacity. |
q_max |
B_degradation | max Q_discharge over the window | Peak capacity. Many cells gain capacity for tens of cycles as formation completes before net fade begins. |
q_max_minus_q2 |
B_degradation | max(Q) - Q(2) | Size of that early capacity rise. A larger rise indicates continuing electrode wetting and activation. |
q_slope |
B_degradation | least-squares slope of Q(cycle) over the window | Trend estimate using every cycle rather than just the endpoints, so it is far less sensitive to noise in any single measurement. |
q_var |
B_degradation | variance of Q_discharge over the window | Trajectory roughness; elevated variance indicates unstable cycling behaviour. |
ir_at_cycle_2 |
C_resistance | internal resistance at measured cycle 2 | Post-formation baseline resistance, set by the interphase built during formation. A high starting value indicates lithium already consumed. |
ir_at_cycle_N |
C_resistance | internal resistance at the budget cycle N | Resistance at the QC decision point. |
ir_diff_N_2 |
C_resistance | IR(N) - IR(2) | Resistance growth over the window: the direct signature of continuing interphase growth and contact loss. |
ir_mean |
C_resistance | mean internal resistance over the window | Window-average level, robust to single-cycle measurement noise. |
ir_min |
C_resistance | min internal resistance over the window | Resistance floor. Cells typically fall to a minimum as wetting completes before rising; the floor is a cleaner baseline than any single cycle. |
ir_ratio_N_min |
C_resistance | IR(N) / min(IR) | Relative rise from the floor. Being dimensionless it is comparable across cells with different absolute resistance. |
ir_slope |
C_resistance | least-squares slope of IR(cycle) over the window | Rate of resistance rise, using every cycle rather than two endpoints. |
cc_time_fraction |
E_charge_dynamics | mean over cycles of (time in CC) / (total charge time) | Share of charging spent at constant current. As impedance rises the cell hits the voltage limit earlier and this fraction falls, making it a more specific impedance probe than total charge time. |
charge_time_diff_N_2 |
E_charge_dynamics | charge_time(N) - charge_time(2) | Endpoint change in charge duration over the window. |
charge_time_mean |
E_charge_dynamics | mean charge time over measured cycles 2..N | Typical charge duration under a fixed protocol; a direct impedance proxy. |
charge_time_slope |
E_charge_dynamics | least-squares slope of charge time vs cycle | Whether charge time is drifting upward, the signature of growing polarization resistance. |
charge_time_var |
E_charge_dynamics | variance of charge time over measured cycles 2..N | Cycle-to-cycle instability in charge acceptance. |
cv_time_fraction |
E_charge_dynamics | mean over cycles of (time in CV) / (total charge time) | Complement of the above; the constant-voltage tail lengthens as the cell becomes harder to charge. |
policy_c1 |
F_protocol | first-step charging C-rate | How aggressively the cell is charged at low state of charge, where lithium plating risk is highest. |
policy_c2 |
F_protocol | second-step charging C-rate | Charging rate through the middle of the state-of-charge range. |
policy_c3 |
F_protocol | third-step charging C-rate where the recipe defines one | Present only for multi-step recipes; NaN otherwise, and imputed inside the CV pipeline rather than filled with a fabricated rate. |
policy_c_max |
F_protocol | maximum C-rate used anywhere in the recipe | Peak stress the cell experiences during charging. |
policy_c_mean |
F_protocol | SOC-weighted mean C-rate across the recipe steps | Single summary of overall charging aggressiveness, weighting each rate by the fraction of charge delivered at it. |
policy_c_range |
F_protocol | policy_c_max - min C-rate in the recipe | Spread between the gentlest and harshest step; a flat recipe and a steeply stepped one can share a mean but stress the cell very differently. |
policy_n_steps |
F_protocol | number of steps in the recipe | Recipe complexity, distinguishing simple two-step from multi-step profiles. |
policy_q1_pct |
F_protocol | state of charge (%) at which the protocol switches step | Where the recipe hands over between rates; determines how much of the high-plating-risk region is traversed at the first rate. |
charge_time_slope_x_ir_slope |
G_interactions | charge_time_slope * ir_slope | Both are proxies for growing polarization resistance measured through different channels. Their product is large only when the two agree, suppressing the case where either drifts through measurement noise alone. |
ir_growth_per_cycle |
G_interactions | ir_diff_N_2 / (N - 2) | Resistance growth normalised by window length so that budgets are directly comparable; without it the raw difference grows mechanically with N. |
NOT reliably measurable in-line - deployability caveat
12 features needing per-cell instrumentation that many lines do not have. A screening rule depending on these may not be deployable without additional hardware, which is why the distinction is tracked in the registry rather than left implicit.
| Feature | Group | Formula | Rationale |
|---|---|---|---|
temp_avg_slope |
D_thermal | least-squares slope of per-cycle mean temperature vs cycle | Whether the cell is progressively self-heating, which indicates rising internal impedance dissipating more energy. |
temp_max |
D_thermal | max of per-cycle maximum temperature over the window | Worst-case thermal excursion, which drives the fastest side reactions. |
temp_max_slope |
D_thermal | least-squares slope of per-cycle maximum temperature vs cycle | The same trend measured at the thermal peak, where it appears earliest. |
temp_mean |
D_thermal | mean of per-cycle average temperature over the window | Typical operating temperature of the cell. |
temp_min |
D_thermal | min of per-cycle minimum temperature over the window | Lower bound of the thermal envelope; with the maximum it gives the range. |
temp_range |
D_thermal | temp_max - temp_min | Thermal swing per cycle, a proxy for heat generated under load relative to the chamber setpoint. |
thermal_exposure |
D_thermal | sum over cycles of the integral of T dt within each cycle | Cumulative time-at-temperature, the quantity an Arrhenius rate law integrates. Computed from the within-cycle traces rather than from per-cycle averages so that cycles of differing duration are weighted correctly. |
thermal_exposure_per_cycle |
D_thermal | thermal_exposure / number of measured cycles | Rate rather than accumulation, so budgets are comparable. |
fade_per_thermal_exposure |
G_interactions | q_fade_rate / thermal_exposure | Capacity loss normalised by thermal load. A cell fading fast despite low thermal exposure is degrading for a non-thermal reason, which is diagnostically distinct. |
log_dq_var_x_ir_ratio |
G_interactions | log_abs_dq_var * ir_ratio_N_min | Combines loss of accessible lithium inventory (curve shape) with rising interfacial impedance, distinguishing cells losing capacity through inventory loss from those losing it through impedance rise. |
recipe_stress_x_thermal |
G_interactions | policy_c_max * thermal_exposure | Higher charging C-rates drive both ohmic heating and lithium plating. This links the commanded process stress to the thermal response it actually produced. PROCESS-RECIPE DERIVED: excluded from the no_recipe feature set with Group F. |
thermal_x_ir_growth |
G_interactions | thermal_exposure * ir_diff_N_2 | Side reactions are Arrhenius-activated AND proceed at the interphase. A cell that is both hot and building resistance is degrading by both mechanisms at once, which neither term captures alone. |
Degenerate columns are reported, not dropped
Two columns are degenerate in this corpus: policy_c3 is entirely NaN and
policy_n_steps is constant, because every one of the 124 cells uses a
two-step charging recipe. The three-step parsing path is retained because it is
correct and general, not because this dataset exercises it.
They are deliberately left in the matrix. Dropping a column for being
constant across the whole cohort is a decision taken with sight of every cell,
including those destined for test folds -- a mild form of fitting on all the
data. The variance filter in selector.py
removes them per fold, on training-fold statistics only, which is where that
decision belongs.
6. Selection is confined to CV folds - structurally
With n = 124, selecting on the full dataset is the fastest way to manufacture a
result that does not replicate. Every selector is an sklearn transformer with a
fit/transform split, composed by build_selection_pipeline(). As Pipeline
steps, scikit-learn guarantees fit sees only the training fold -- the
guarantee does not depend on this project remembering to arrange it. There is
deliberately no module-level select_features(X, y) helper, because that is
the API shape that invites full-data use.
configs/features.yaml sets selection.fit_inside_cv_only: true, and the schema
rejects the config if it is ever false.
| Step | Purpose | Supervised? |
|---|---|---|
SimpleImputer |
Median fill. keep_empty_features=True so all-NaN columns survive to the variance filter instead of being silently dropped, which would misalign every downstream feature name. |
no |
VarianceFilter |
Remove constant and all-missing columns. | no |
CorrelationFilter |
Drop one of each pair above rho = 0.95. Curve features are collinear by construction, and collinearity destabilises both linear coefficients and the Phase 9 SHAP attributions that must support an auditable scrap decision. | no |
StabilitySelector |
Keep features chosen in at least 60% of 100 bootstrap resamples of the training fold. | yes |
MutualInfoSelector |
Rank by mutual information, capturing monotone-but-nonlinear relationships. | yes |
StandardScaler |
Scale. | no |
Both supervised selectors raise if fitted without labels, so they cannot be quietly misused on unlabelled full data.
A silent bug found and fixed here
StabilitySelector originally ranked features with
pd.Series(sub[c]).corr(pd.Series(sub_y)). pandas correlates on the index,
and the bootstrap subset carries the original row labels while the target gets a
fresh RangeIndex -- so the two were silently mis-paired.
The effect: dq_var, whose true |rho| is 0.908, scored 0.023. The single
strongest predictor in the project was demoted to noise, and selection returned
an all-thermal feature set -- the group Phase 3 measured as the weakest
(|rho| ~ 0.22).
It raised no error and produced a plausible-looking result. It was caught only
because the output contradicted a Phase 3 measurement. The statistic is now
computed on NumPy arrays via scipy.stats.spearmanr, and after the fix selection
returns dq_var first, as it should.
7. Outputs
| File | Contents |
|---|---|
data/processed/features_budget{005,010,020,050,100}.parquet |
124 cells x 57 features per budget |
outputs/reports/feature_registry.json |
Every feature with formula, rationale and in-line flag |
outputs/reports/feature_mask_report.json |
Exactly what the plausibility mask removed, per cell per budget |
data/processed/features_provenance.json |
Config hash and package versions |
python -m src.features.builder # build all budgets
python -m src.features.make_docs # regenerate this page
python -m pytest tests/test_leakage.py -q