Spaces:
Running
Hyperparameter Ranges β Practical Reference
Practical starting ranges for tree-based and linear models on tabular data. These are empirically validated defaults β tune from here rather than searching blind.
XGBoost
| Parameter | Typical Range | Notes |
|---|---|---|
n_estimators |
100β1000 | Use early stopping to find optimal |
learning_rate |
0.01β0.3 | Lower LR needs more trees; 0.05β0.1 is a good start |
max_depth |
3β8 | Deeper trees overfit; 4β6 is usually optimal |
min_child_weight |
1β10 | Higher = more conservative, reduces overfitting |
subsample |
0.6β1.0 | Row sampling per tree; 0.8 is a safe default |
colsample_bytree |
0.5β1.0 | Feature sampling per tree; 0.8 is a safe default |
gamma |
0β5 | Minimum loss reduction to split; 0 means always split |
reg_alpha (L1) |
0β1 | Useful for sparse feature sets |
reg_lambda (L2) |
0β10 | Default 1; increase to regularize |
scale_pos_weight |
neg/pos |
For imbalanced classification; set to ratio of negative/positive samples |
Search order: Fix n_estimators high with early stopping β tune max_depth + min_child_weight β tune subsample + colsample_bytree β tune learning_rate (lower it, add more trees).
LightGBM
| Parameter | Typical Range | Notes |
|---|---|---|
n_estimators |
100β2000 | With early stopping |
learning_rate |
0.01β0.2 | Start at 0.05 |
num_leaves |
20β300 | Main complexity control; < 2^max_depth |
max_depth |
-1 (unlimited) or 6β12 | Prefer controlling via num_leaves |
min_child_samples |
10β100 | Minimum samples per leaf; prevents overfitting on small data |
subsample |
0.6β1.0 | Row sampling (bagging_fraction) |
colsample_bytree |
0.5β1.0 | Feature sampling (feature_fraction) |
reg_alpha |
0β1 | L1 regularization |
reg_lambda |
0β10 | L2 regularization |
is_unbalance |
True/False | Auto-adjusts weights for imbalance |
Key difference from XGBoost: num_leaves is the primary complexity control. With max_depth=6, XGBoost has max 64 leaves; LightGBM can have 300+ at the same depth setting.
CatBoost
| Parameter | Typical Range | Notes |
|---|---|---|
iterations |
100β2000 | Equivalent to n_estimators |
learning_rate |
0.01β0.3 | Auto-tuned if not set |
depth |
4β10 | CatBoost trees are symmetric (oblivious); depth 6β8 is typical |
l2_leaf_reg |
1β10 | L2 regularization on leaves |
border_count |
32β255 | Quantization borders for numeric features |
bagging_temperature |
0β1 | 0 = no bagging, 1 = full Bayesian bootstrap |
random_strength |
0β10 | Noise for split selection; helps avoid overfitting |
CatBoost advantage: Pass categorical column indices via cat_features β no manual encoding needed. It handles them via target statistics internally.
Random Forest
| Parameter | Typical Range | Notes |
|---|---|---|
n_estimators |
100β500 | More trees = better, diminishing returns after 300 |
max_depth |
None or 10β30 | None (fully grown) is often best; limit for speed |
min_samples_split |
2β20 | Minimum samples to split a node |
min_samples_leaf |
1β10 | Minimum samples in a leaf |
max_features |
"sqrt", "log2", 0.5 |
sqrt is standard for classification, 1/3 for regression |
max_samples |
0.5β1.0 | Bootstrap sample size (if bootstrap=True) |
class_weight |
"balanced" |
For imbalanced classification |
Note: Random Forest is less sensitive to hyperparameters than gradient boosting. Default settings often work well. Focus tuning budget on max_features and min_samples_leaf.
Logistic Regression / Ridge / Lasso
| Parameter | Typical Range | Notes |
|---|---|---|
C (LogReg) |
0.001β100 | Inverse of regularization strength; log scale search |
alpha (Ridge/Lasso) |
0.001β100 | Regularization strength; log scale search |
solver |
"lbfgs", "saga" |
saga for L1 or large datasets |
max_iter |
100β1000 | Increase if convergence warnings appear |
penalty |
"l1", "l2", "elasticnet" |
L1 gives sparsity; L2 is default |
Always scale features before fitting linear models. Use StandardScaler or MinMaxScaler.
SVM
| Parameter | Typical Range | Notes |
|---|---|---|
C |
0.01β100 | Regularization; higher = less regularization |
gamma |
"scale", "auto", 0.001β1 |
RBF kernel width; scale is a good default |
kernel |
"rbf", "linear", "poly" |
rbf for nonlinear, linear for high-dim sparse |
Scale features before SVM. SVMs do not scale to large datasets (>50k rows) β use SGDClassifier instead.
Optuna Search Spaces
def objective(trial):
params = {
"n_estimators": trial.suggest_int("n_estimators", 100, 1000),
"learning_rate": trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
"max_depth": trial.suggest_int("max_depth", 3, 8),
"subsample": trial.suggest_float("subsample", 0.6, 1.0),
"colsample_bytree": trial.suggest_float("colsample_bytree", 0.5, 1.0),
"min_child_weight": trial.suggest_int("min_child_weight", 1, 10),
}
Use log=True for parameters that span orders of magnitude (learning_rate, C, alpha).
General Tuning Principles
Use early stopping for gradient boosting β set
n_estimatorshigh (1000+) and let early stopping find the right count. Avoids overfitting and saves compute.Search in log space for regularization parameters and learning rates β the difference between 0.001 and 0.01 matters as much as 0.1 vs 1.0.
Fix one group at a time β tune tree structure (
max_depth,num_leaves) first, then sampling params, then regularization, then learning rate.More data beats more tuning β hyperparameter tuning typically gives 1β3% improvement; getting more/better data often gives 5β20%.
Don't tune on test set β use cross-validation or a held-out validation set for tuning. Report test set performance only once, at the end.