license: mit
tags:
- soil-science
- pedotransfer-function
- water-retention
- attention-mechanism
- ensemble
- tensorflow
- geoscience
pipeline_tag: tabular-regression
library_name: tensorflow
datasets:
- custom
language:
- en
HABIT — Hierarchical Attention-Based Inference with Transfer Learning
A pre-trained ensemble model for predicting soil water retention curves from basic soil properties.
Paper: Ghezzehei TA (2025). Interpretable Soil Water Retention Prediction Using Hierarchical Attention Networks with Uncertainty Quantification. Water Resources Research. [DOI pending]
Training data & code: Dryad repository
What HABIT does
HABIT predicts volumetric water content (θ) at any water potential (ψ) from soil properties. Unlike parameter-based approaches (e.g., van Genuchten fitting), HABIT predicts retention curves directly — no ill-posed inverse problem.
Input → Output
| Input | Required? | Units |
|---|---|---|
| Sand, Silt, Clay | Yes | fraction (0–1) |
| Bulk density | Optional | g/cm³ |
| Organic carbon | Optional | fraction (0–1) |
| Saturated hydraulic conductivity | Optional | cm/day |
| Water potential(s) | Yes | kPa |
Output: Volumetric water content (cm³/cm³) with ensemble uncertainty (mean ± std from 20 models).
Automatic input adaptation
HABIT uses a single model architecture trained hierarchically. Provide whatever soil properties you have — the model adapts automatically:
- Texture only → equivalent to Stage 0 (R² ≈ 0.78)
- Texture + BD → equivalent to Stage 1 (R² ≈ 0.85)
- Texture + BD + OC → equivalent to Stage 2 (R² ≈ 0.86)
- All properties → Stage 3 (R² ≈ 0.92)
More properties = better predictions, but texture alone already works.
Quick start
from habit_inference import HABITPredictor, load_ensemble
# Load the 20-member ensemble
predictor = load_ensemble(stage=3)
# Predict for a single soil
import pandas as pd
soil = pd.DataFrame({
'soil_id': ['my_soil'],
'sand': [0.40],
'silt': [0.35],
'clay': [0.25],
'bd': [1.35], # optional
'oc': [0.012], # optional
'ksat': [25.0] # optional, cm/day
})
predictions = predictor.predict(soil)
# Returns: soil_id, water_potential_kPa, water_content_mean, water_content_std
From CSV
predictor.predict_from_csv('input.csv', 'output.csv')
Command line
habit-predict --input soils.csv --output predictions.csv
Model architecture
HABIT uses property-specific encoders, cross-attention layers for property interactions, multi-head attention over soil properties and water potentials, and a monotonic output layer that enforces physically correct behavior (water content decreases with increasing tension).
Key design: Hierarchical transfer learning trains the model sequentially — texture first, then adding bulk density, organic carbon, and Ksat — so a single set of weights handles any combination of available inputs via masking.
Architecture parameters
| Parameter | Value |
|---|---|
| Embedding dimension | 192 |
| Attention heads | 4 |
| Monotonic basis functions | 40 |
| Dropout rate | 0.15 |
| Ensemble members | 20 |
Repository contents
weights/
member_01.h5 ... member_20.h5 # Hierarchical ensemble weights (19 MB each)
model/
habit.py # Model architecture
layers.py # Custom layers (attention, monotonic)
config/
model_config.json # Architecture hyperparameters
scaler_params.json # Feature scaling parameters (robust scaling)
Performance
Evaluated on held-out test set (cluster bootstrap, 1000 iterations):
| Configuration | R² | RMSE (cm³/cm³) | MAE (cm³/cm³) |
|---|---|---|---|
| HABIT Stage 0 (texture) | 0.779 [0.737, 0.817] | 0.067 [0.060, 0.074] | 0.049 [0.044, 0.055] |
| HABIT Stage 1 (+BD) | 0.846 [0.748, 0.906] | 0.056 [0.044, 0.070] | 0.039 [0.033, 0.049] |
| HABIT Stage 2 (+OC) | 0.862 [0.781, 0.920] | 0.052 [0.040, 0.066] | 0.038 [0.030, 0.047] |
| HABIT Stage 3 (+Ksat) | 0.923 [0.899, 0.944] | 0.043 [0.036, 0.050] | 0.030 [0.026, 0.035] |
| Rosetta Model 2 (texture) | 0.009 | 0.141 | 0.113 |
| Rosetta Model 3 (+BD) | 0.511 | 0.099 | 0.075 |
Training data
2,577 soil samples compiled from two international databases (Hohenbrink et al. 2023; Gupta et al. 2022), with 44,636 water retention measurements after adaptive thinning. See the Dryad repository for the full dataset and training code.
License
MIT (code and weights). Training data: CC BY 4.0.
Citation
@article{ghezzehei2025habit,
title={Interpretable Soil Water Retention Prediction Using Hierarchical Attention Networks with Uncertainty Quantification},
author={Ghezzehei, Teamrat A.},
journal={Water Resources Research},
year={2025},
note={DOI pending}
}
Contact
Teamrat A. Ghezzehei — taghezzehei@ucmerced.edu Life and Environmental Sciences, University of California, Merced