--- license: mit tags: - soil-science - pedotransfer-function - water-retention - attention-mechanism - ensemble - tensorflow - geoscience pipeline_tag: tabular-regression library_name: tensorflow datasets: - custom language: - en --- # HABIT — Hierarchical Attention-Based Inference with Transfer Learning A pre-trained ensemble model for predicting soil water retention curves from basic soil properties. **Paper:** Ghezzehei TA (2025). Interpretable Soil Water Retention Prediction Using Hierarchical Attention Networks with Uncertainty Quantification. *Water Resources Research*. [DOI pending] **Training data & code:** [Dryad repository](https://datadryad.org/share/LINK_NOT_FOR_PUBLICATION/iXOXxn5h-Hdq9m7aA3cTLgGMAHiHMre2HxZItyI-Ykc) ## What HABIT does HABIT predicts volumetric water content (θ) at any water potential (ψ) from soil properties. Unlike parameter-based approaches (e.g., van Genuchten fitting), HABIT predicts retention curves **directly** — no ill-posed inverse problem. ### Input → Output | Input | Required? | Units | |---|---|---| | Sand, Silt, Clay | **Yes** | fraction (0–1) | | Bulk density | Optional | g/cm³ | | Organic carbon | Optional | fraction (0–1) | | Saturated hydraulic conductivity | Optional | cm/day | | Water potential(s) | **Yes** | kPa | **Output:** Volumetric water content (cm³/cm³) with ensemble uncertainty (mean ± std from 20 models). ### Automatic input adaptation HABIT uses a single model architecture trained hierarchically. Provide whatever soil properties you have — the model adapts automatically: - **Texture only** → equivalent to Stage 0 (R² ≈ 0.78) - **Texture + BD** → equivalent to Stage 1 (R² ≈ 0.85) - **Texture + BD + OC** → equivalent to Stage 2 (R² ≈ 0.86) - **All properties** → Stage 3 (R² ≈ 0.92) More properties = better predictions, but texture alone already works. ## Quick start ```python from habit_inference import HABITPredictor, load_ensemble # Load the 20-member ensemble predictor = load_ensemble(stage=3) # Predict for a single soil import pandas as pd soil = pd.DataFrame({ 'soil_id': ['my_soil'], 'sand': [0.40], 'silt': [0.35], 'clay': [0.25], 'bd': [1.35], # optional 'oc': [0.012], # optional 'ksat': [25.0] # optional, cm/day }) predictions = predictor.predict(soil) # Returns: soil_id, water_potential_kPa, water_content_mean, water_content_std ``` ### From CSV ```python predictor.predict_from_csv('input.csv', 'output.csv') ``` ### Command line ```bash habit-predict --input soils.csv --output predictions.csv ``` ## Model architecture HABIT uses property-specific encoders, cross-attention layers for property interactions, multi-head attention over soil properties and water potentials, and a monotonic output layer that enforces physically correct behavior (water content decreases with increasing tension). **Key design:** Hierarchical transfer learning trains the model sequentially — texture first, then adding bulk density, organic carbon, and Ksat — so a single set of weights handles any combination of available inputs via masking. ### Architecture parameters | Parameter | Value | |---|---| | Embedding dimension | 192 | | Attention heads | 4 | | Monotonic basis functions | 40 | | Dropout rate | 0.15 | | Ensemble members | 20 | ## Repository contents ``` weights/ member_01.h5 ... member_20.h5 # Hierarchical ensemble weights (19 MB each) model/ habit.py # Model architecture layers.py # Custom layers (attention, monotonic) config/ model_config.json # Architecture hyperparameters scaler_params.json # Feature scaling parameters (robust scaling) ``` ## Performance Evaluated on held-out test set (cluster bootstrap, 1000 iterations): | Configuration | R² | RMSE (cm³/cm³) | MAE (cm³/cm³) | |---|---|---|---| | HABIT Stage 0 (texture) | 0.779 [0.737, 0.817] | 0.067 [0.060, 0.074] | 0.049 [0.044, 0.055] | | HABIT Stage 1 (+BD) | 0.846 [0.748, 0.906] | 0.056 [0.044, 0.070] | 0.039 [0.033, 0.049] | | HABIT Stage 2 (+OC) | 0.862 [0.781, 0.920] | 0.052 [0.040, 0.066] | 0.038 [0.030, 0.047] | | HABIT Stage 3 (+Ksat) | 0.923 [0.899, 0.944] | 0.043 [0.036, 0.050] | 0.030 [0.026, 0.035] | | Rosetta Model 2 (texture) | 0.009 | 0.141 | 0.113 | | Rosetta Model 3 (+BD) | 0.511 | 0.099 | 0.075 | ## Training data 2,577 soil samples compiled from two international databases (Hohenbrink et al. 2023; Gupta et al. 2022), with 44,636 water retention measurements after adaptive thinning. See the [Dryad repository](https://datadryad.org/share/LINK_NOT_FOR_PUBLICATION/iXOXxn5h-Hdq9m7aA3cTLgGMAHiHMre2HxZItyI-Ykc) for the full dataset and training code. ## License MIT (code and weights). Training data: CC BY 4.0. ## Citation ```bibtex @article{ghezzehei2025habit, title={Interpretable Soil Water Retention Prediction Using Hierarchical Attention Networks with Uncertainty Quantification}, author={Ghezzehei, Teamrat A.}, journal={Water Resources Research}, year={2025}, note={DOI pending} } ``` ## Contact Teamrat A. Ghezzehei — [taghezzehei@ucmerced.edu](mailto:taghezzehei@ucmerced.edu) Life and Environmental Sciences, University of California, Merced