Spaces:
Running
Running
File size: 6,324 Bytes
3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 964a5ed 3499a42 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 | <!-- Source: https://scikit-learn.org/stable/modules/preprocessing.html β fetched 2026-07-01 -->
# Feature Engineering (sklearn.preprocessing)
## What It Does
`sklearn.preprocessing` transforms raw feature vectors into representations suitable for ML estimators. Most algorithms benefit from standardized data; some require bounded inputs or specific distributions. All transformers follow the fit/transform pattern β always fit on training data only.
---
## Scaling Transformers
| Transformer | Formula | When to Use | Caveats |
|---|---|---|---|
| **StandardScaler** | `(x - mean) / std` | Linear models, SVMs, neural nets, PCA | Cannot center sparse matrices (`with_mean=False` for sparse) |
| **MinMaxScaler** | `(x - min) / (max - min)` β [0,1] | Neural nets, bounded input needed | Sensitive to outliers; new data can fall outside [0,1] |
| **MaxAbsScaler** | `x / max(|x|)` β [-1,1] | Sparse data (preserves sparsity) | Assumes data already centered near zero |
| **RobustScaler** | `(x - median) / IQR` | Data with significant outliers | Cannot fit sparse inputs (can transform them) |
| **Normalizer** | Row-wise L1/L2/max norm | Text classification, clustering, vector space models | Operates per-sample, not per-feature |
**Key parameter:** `with_mean=True / with_std=True` on StandardScaler β disable either independently.
---
## Non-linear Transformations
### PowerTransformer
Maps data to approximately Gaussian using parametric power transforms.
- **`method='yeo-johnson'`** β works for any values including negatives. Default.
- **`method='box-cox'`** β only for strictly positive data.
- By default applies zero-mean, unit-variance normalization to the output.
- Use when a column has skewed distribution and you need Gaussian-like input for a linear model or neural net.
### QuantileTransformer
Non-parametric rank-based transform to uniform [0,1] or normal distribution.
- `output_distribution='uniform'` (default) or `'normal'`.
- Robust to outliers β extreme values are compressed to the boundary.
- **Caveat:** Distorts correlations and distances between features; do not use before PCA or distance-based models if feature relationships matter.
---
## Categorical Encoders
### OrdinalEncoder
Maps categories to integers 0..n-1. Implies an ordering β use only for genuinely ordered categories (e.g., education level, satisfaction rating).
- `handle_unknown='use_encoded_value'` with `unknown_value=-1` to handle unseen categories at inference.
- Never apply to nominal (unordered) categories.
### OneHotEncoder
Creates one binary column per category. Suitable for nominal, low-cardinality features (< 15β20 unique values).
- `drop='first'` or `'if_binary'` β drop one column to avoid perfect multicollinearity in non-regularized linear models.
- `handle_unknown='infrequent_if_exist'` β encodes unseen categories as all-zeros or as infrequent bucket.
- `max_categories` β groups infrequent categories into a single "infrequent" bucket to control output width.
- `min_frequency` β define the cardinality threshold for infrequency.
### TargetEncoder (sklearn >= 1.3)
Encodes categories using target mean for that category. Handles high-cardinality automatically with a single dense output column.
- Uses k-fold cross-fitting in `fit_transform()` to prevent target leakage.
- **Critical:** `fit(X,y).transform(X)` is NOT equal to `fit_transform(X,y)` β always use `fit_transform` on training data.
- For test data: `fit(X_train, y_train).transform(X_test)`.
---
## Discretization
### KBinsDiscretizer
Partitions continuous features into discrete bins.
- `strategy='uniform'` β constant-width bins.
- `strategy='quantile'` β equally populated bins. More robust for skewed data.
- `strategy='kmeans'` β bins by k-means clustering per feature.
- `encode='ordinal'` (integer bins) or `'onehot'` (one-hot encoded output).
- **When to use:** Introduce non-linearity to linear models; handle non-monotonic relationships.
---
## Feature Generation
### PolynomialFeatures
Generates polynomial and interaction terms from numeric features.
- `degree=2` on [a, b] β [1, a, b, aΒ², ab, bΒ²].
- `interaction_only=True` β only cross-products, no squared terms.
- `include_bias=False` β omit the bias column (intercept) when the model adds its own.
- **Caveat:** Feature count grows as O(n^degree). With 20 features and degree=2, output is 231 columns. Use carefully; requires strong regularization.
- **Tree-based models do not need polynomial features** β they discover interactions natively.
### SplineTransformer
Generates B-spline basis functions (piecewise polynomials).
- `degree` β polynomial degree per piece (typically 3).
- `n_knots` β number of knot points.
- Advantages over PolynomialFeatures: no Runge's oscillation at boundaries, low condition number, more flexible with fixed low degree.
- Treats each feature independently β no interaction terms.
### FunctionTransformer
Wraps an arbitrary Python function as a sklearn transformer.
- `func=np.log1p` β apply log1p to all features.
- `validate=True` β ensure numeric input.
- `check_inverse=True` β verify `func` and `inverse_func` are inverses.
---
## Missing Value Handling
### MissingIndicator
Creates binary indicator columns marking which values were missing. Use alongside an imputer to give the model explicit signal about missingness patterns.
---
## Best Practices and Caveats
| Issue | Solution |
|---|---|
| Data leakage | Always `.fit()` on training data only; use `Pipeline` to enforce |
| Sparse data centering | Use `MaxAbsScaler` or `StandardScaler(with_mean=False)` |
| Outliers distorting scaling | Use `RobustScaler` instead of `StandardScaler` |
| Target encoder leakage | Use `fit_transform(X_train, y_train)` for training data |
| Collinearity in OHE | Use `OneHotEncoder(drop='first')` for non-regularized linear models |
| Unknown categories at inference | Set `handle_unknown='infrequent_if_exist'` in OneHotEncoder |
| High-cardinality features | Use `TargetEncoder` instead of one-hot encoding |
### Pipeline pattern (enforce correct fit/transform discipline)
```python
from sklearn.pipeline import make_pipeline
pipe = make_pipeline(StandardScaler(), LogisticRegression())
pipe.fit(X_train, y_train)
pipe.score(X_test, y_test)
```
|