File size: 4,851 Bytes
53f9f3e 43604b6 53f9f3e 43604b6 53f9f3e ad36a58 fce6c09 43604b6 fce6c09 43604b6 ad36a58 43604b6 fce6c09 43604b6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 | import streamlit as st
from pathlib import Path
st.set_page_config(page_title="Calibration Benchmark", page_icon="π ", layout="wide")
# Sidebar navigation
st.sidebar.title("Navigation")
st.sidebar.page_link("streamlit_app.py", label="Home", icon="π ")
st.sidebar.page_link("pages/OptimizationLeaderboard.py", label="Optimization Leaderboard", icon="π")
st.sidebar.page_link("pages/UQLeaderboard.py", label="UQ Leaderboard", icon="π―")
st.sidebar.page_link("pages/MethodDetails.py", label="Methods", icon="π")
st.sidebar.page_link("pages/RawData.py", label="Get Data", icon="π§Ύ")
st.title("Calibration Benchmark")
st.markdown(
"A benchmark comparing parameter-calibration methods on chaotic dynamical systems. "
"Methods are ranked by **forward-model run efficiency** β how many forward-model "
"evaluations are needed, on average across random seeds, to reach a target accuracy."
)
st.divider()
col1, col2 = st.columns(2, gap="large")
with col1:
st.subheader("π Optimization Leaderboard")
st.markdown(
"Ranks methods by mean forward-model runs to reach an **RMSE target** on "
"Lorenz-63 and Lorenz-96 benchmarks. "
"Lower is better; failed runs are tracked separately as a failure rate."
)
st.page_link("pages/OptimizationLeaderboard.py", label="Go to Optimization Leaderboard β")
with col2:
st.subheader("π― UQ Leaderboard")
st.markdown(
"Ranks methods by mean forward-model runs to reach an **uncertainty quantification "
"target**. Same metric and benchmarks as the Optimization Leaderboard, evaluated "
"at a UQ-specific convergence criterion."
)
st.page_link("pages/UQLeaderboard.py", label="Go to UQ Leaderboard β")
st.divider()
st.subheader("Benchmarks")
st.markdown(
"Results are reported on four benchmark configurations of the [Lorenz system]"
"(https://en.wikipedia.org/wiki/Lorenz_system), a standard testbed for "
"data-assimilation and calibration algorithms:"
)
_media = Path(__file__).parent / "media"
st.markdown("- **L63** β Lorenz-63 β 3-variable chaotic attractor; learn 2 parameters, strongly nonlinear.")
st.image(
str(_media / "posterior_ribbons_20_13_k5.png"),
caption="Example prior-posterior & truth. L: Difference to true parameter. R: data-sample/output predictived distribution (state-mean [1:3], state-covariance (diag [4:6] and off-diag [7:9]))",
width=800,
)
st.markdown("- **L96** β Lorenz-96 (40-variable); learn 1-parameter constant forcing.")
st.image(
str(_media / "posterior_ribbons_const-force_12_1_k3.png"),
caption="Example prior-posterior & truth. L: Difference to true parameter. C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:40] state-mean [41:80] state-std)",
width=1200,
)
st.markdown("- **L96_SPATIAL_FORCING** β Lorenz-96 (40-variable) with spatially-varying forcing; learn 40 parameters; moderately correlated prior.")
st.image(
str(_media / "posterior_ribbons_vec-force_65_1_k3.png"),
caption="Example prior-posterior & truth. L: Difference to true parameters. C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:40] state-mean [41:80] state-std)",
width=1200,
)
st.markdown("- **L96_NN_FORCING** β Lorenz-96 (100-variable) with a neural-network forcing; Learn 61 parameters (weights and biases of the network). Reasonable prior given.")
st.image(
str(_media / "posterior_ribbons_flux-force_80_1_k3.png"),
caption="Example prior-posterior & truth. L: Difference to true parameters (weights). C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:100] state-mean [101:200] state-std)",
width=1200,
)
st.subheader("Method families")
st.markdown(
"Methods are grouped into three families. See the **π Methods** page for "
"citations and per-method performance charts."
)
families = {
"Kalman": "Ensemble Kalman variants (TEKI, ETKI, IEKF, UKI) β update an ensemble of parameter guesses via a linearized observation operator.",
"Bayesian": "Sampling-based approaches (ABC, HM) β explore parameter space without requiring gradient information.",
"Calibrate-then-emulate": "Two-stage pipelines (CES-EKI-DMC) β use an initial calibration phase to build a cheap emulator, then sample the posterior via MCMC.",
}
for family, desc in families.items():
st.markdown(f"- **{family}** β {desc}")
st.subheader("Key metric")
st.markdown(
"The reported metric is the **mean number of forward-model evaluations** required "
"to reach the target, averaged over random seeds. "
"A value of **β1** indicates a failed run (did not reach the target); "
"the **failure rate** shows the fraction of seeds that failed."
)
|