| import streamlit as st |
| from pathlib import Path |
|
|
| st.set_page_config(page_title="Calibration Benchmark", page_icon="π ", layout="wide") |
|
|
| |
| st.sidebar.title("Navigation") |
| st.sidebar.page_link("streamlit_app.py", label="Home", icon="π ") |
| st.sidebar.page_link("pages/OptimizationLeaderboard.py", label="Optimization Leaderboard", icon="π") |
| st.sidebar.page_link("pages/UQLeaderboard.py", label="UQ Leaderboard", icon="π―") |
| st.sidebar.page_link("pages/MethodDetails.py", label="Methods", icon="π") |
| st.sidebar.page_link("pages/RawData.py", label="Get Data", icon="π§Ύ") |
|
|
| st.title("Calibration Benchmark") |
| st.markdown( |
| "A benchmark comparing parameter-calibration methods on chaotic dynamical systems. " |
| "Methods are ranked by **forward-model run efficiency** β how many forward-model " |
| "evaluations are needed, on average across random seeds, to reach a target accuracy." |
| ) |
|
|
| st.divider() |
|
|
| col1, col2 = st.columns(2, gap="large") |
|
|
| with col1: |
| st.subheader("π Optimization Leaderboard") |
| st.markdown( |
| "Ranks methods by mean forward-model runs to reach an **RMSE target** on " |
| "Lorenz-63 and Lorenz-96 benchmarks. " |
| "Lower is better; failed runs are tracked separately as a failure rate." |
| ) |
| st.page_link("pages/OptimizationLeaderboard.py", label="Go to Optimization Leaderboard β") |
|
|
| with col2: |
| st.subheader("π― UQ Leaderboard") |
| st.markdown( |
| "Ranks methods by mean forward-model runs to reach an **uncertainty quantification " |
| "target**. Same metric and benchmarks as the Optimization Leaderboard, evaluated " |
| "at a UQ-specific convergence criterion." |
| ) |
| st.page_link("pages/UQLeaderboard.py", label="Go to UQ Leaderboard β") |
|
|
| st.divider() |
|
|
| st.subheader("Benchmarks") |
| st.markdown( |
| "Results are reported on four benchmark configurations of the [Lorenz system]" |
| "(https://en.wikipedia.org/wiki/Lorenz_system), a standard testbed for " |
| "data-assimilation and calibration algorithms:" |
| ) |
|
|
| _media = Path(__file__).parent / "media" |
|
|
| st.markdown("- **L63** β Lorenz-63 β 3-variable chaotic attractor; learn 2 parameters, strongly nonlinear.") |
| st.image( |
| str(_media / "posterior_ribbons_20_13_k5.png"), |
| caption="Example prior-posterior & truth. L: Difference to true parameter. R: data-sample/output predictived distribution (state-mean [1:3], state-covariance (diag [4:6] and off-diag [7:9]))", |
| width=800, |
| ) |
|
|
|
|
| st.markdown("- **L96** β Lorenz-96 (40-variable); learn 1-parameter constant forcing.") |
| st.image( |
| str(_media / "posterior_ribbons_const-force_12_1_k3.png"), |
| caption="Example prior-posterior & truth. L: Difference to true parameter. C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:40] state-mean [41:80] state-std)", |
| width=1200, |
| ) |
|
|
| st.markdown("- **L96_SPATIAL_FORCING** β Lorenz-96 (40-variable) with spatially-varying forcing; learn 40 parameters; moderately correlated prior.") |
| st.image( |
| str(_media / "posterior_ribbons_vec-force_65_1_k3.png"), |
| caption="Example prior-posterior & truth. L: Difference to true parameters. C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:40] state-mean [41:80] state-std)", |
| width=1200, |
| ) |
|
|
|
|
| st.markdown("- **L96_NN_FORCING** β Lorenz-96 (100-variable) with a neural-network forcing; Learn 61 parameters (weights and biases of the network). Reasonable prior given.") |
| st.image( |
| str(_media / "posterior_ribbons_flux-force_80_1_k3.png"), |
| caption="Example prior-posterior & truth. L: Difference to true parameters (weights). C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:100] state-mean [101:200] state-std)", |
| width=1200, |
| ) |
|
|
| st.subheader("Method taxonomy") |
| st.markdown( |
| "Every method carries four independent tags β how it searches, what update " |
| "mechanism drives each step, what it's built to report, and whether/when it uses " |
| "a surrogate model. See the **π Methods** page for citations and per-method " |
| "performance charts." |
| ) |
| st.markdown( |
| """ |
| - **Parallelism** β how the search explores parameter space |
| - **Serial** β `ADAM`, `LM` β a single point estimate advanced step by step. |
| - **Parallel-independent** β `ABC`, `HM` β a population of candidates updated |
| with no coupling between members (accepted samples / per-wave resampling). |
| - **Parallel-interacting** β `TEKI`, `ETKI`, `IEKF`, `UKI`, `CES-EKI-DMC`, `CES-EKI-CONST`, |
| `CES-IEKF-CONST` β an ensemble whose members are coupled through a shared update |
| each iteration. |
| - **Update type** β the mechanism driving each update step |
| - **Gradient** β `ADAM`, `LM` β follow the loss gradient (or a Gauss-Newton |
| approximation of it) directly. |
| - **Kalman** β `TEKI`, `ETKI`, `IEKF`, `UKI`, `CES-EKI-DMC`, `CES-EKI-CONST`, |
| `CES-IEKF-CONST` β a (possibly linearized or unscented) Kalman-style ensemble update. |
| - **General** β `ABC`, `HM` β neither gradient- nor Kalman-based (rejection |
| sampling, implausibility cuts). |
| - **Method goal** β what the method is built to report |
| - **Optimization** β `TEKI`, `ETKI`, `UKI`, `ADAM`, `LM` β a single best-fit |
| parameter estimate. |
| - **UQ** β `IEKF`, `ABC`, `HM`, `CES-EKI-DMC`, `CES-EKI-CONST`, `CES-IEKF-CONST` β the |
| full posterior / parameter uncertainty. Can still be scored on the Optimization |
| leaderboard, but tends to be less competitive there since it's not optimizing for |
| speed-to-target. |
| - **Emulator use** β when/whether a surrogate model of the forward model is used |
| - **None** β `TEKI`, `ETKI`, `IEKF`, `UKI`, `ADAM`, `LM`, `ABC` β samples/evaluates |
| the true forward model throughout. |
| - **Within-optimize** β `HM` β refits a surrogate at each iteration (wave) of the |
| search itself. |
| - **After-optimize** β `CES-EKI-DMC`, `CES-EKI-CONST`, `CES-IEKF-CONST` β fits a |
| surrogate (e.g. a GP) once, after calibration finishes, and samples the posterior |
| through it. |
| """ |
| ) |
| st.caption( |
| "Note: Kalman methods are Bayesian in spirit too (they're approximate Gaussian " |
| "posterior updates) β update type is about mechanism (gradient vs. Kalman vs. " |
| "general), not whether a method is 'Bayesian'." |
| ) |
|
|
| st.subheader("Key metric") |
| st.markdown( |
| "The reported metric is the **mean number of forward-model evaluations** required " |
| "to reach the target, averaged over random seeds. " |
| "A value of **β1** indicates a failed run (did not reach the target); " |
| "the **failure rate** shows the fraction of seeds that failed." |
| ) |
|
|