File size: 4,851 Bytes
53f9f3e
43604b6
53f9f3e
43604b6
53f9f3e
ad36a58
fce6c09
 
43604b6
 
fce6c09
 
 
43604b6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ad36a58
43604b6
 
 
 
 
 
 
 
fce6c09
43604b6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
import streamlit as st
from pathlib import Path

st.set_page_config(page_title="Calibration Benchmark", page_icon="🏠", layout="wide")

# Sidebar navigation
st.sidebar.title("Navigation")
st.sidebar.page_link("streamlit_app.py", label="Home", icon="🏠")
st.sidebar.page_link("pages/OptimizationLeaderboard.py", label="Optimization Leaderboard", icon="πŸ“Š")
st.sidebar.page_link("pages/UQLeaderboard.py", label="UQ Leaderboard", icon="🎯")
st.sidebar.page_link("pages/MethodDetails.py", label="Methods", icon="πŸ“˜")
st.sidebar.page_link("pages/RawData.py", label="Get Data", icon="🧾")

st.title("Calibration Benchmark")
st.markdown(
    "A benchmark comparing parameter-calibration methods on chaotic dynamical systems. "
    "Methods are ranked by **forward-model run efficiency** β€” how many forward-model "
    "evaluations are needed, on average across random seeds, to reach a target accuracy."
)

st.divider()

col1, col2 = st.columns(2, gap="large")

with col1:
    st.subheader("πŸ“Š Optimization Leaderboard")
    st.markdown(
        "Ranks methods by mean forward-model runs to reach an **RMSE target** on "
        "Lorenz-63 and Lorenz-96 benchmarks. "
        "Lower is better; failed runs are tracked separately as a failure rate."
    )
    st.page_link("pages/OptimizationLeaderboard.py", label="Go to Optimization Leaderboard β†’")

with col2:
    st.subheader("🎯 UQ Leaderboard")
    st.markdown(
        "Ranks methods by mean forward-model runs to reach an **uncertainty quantification "
        "target**. Same metric and benchmarks as the Optimization Leaderboard, evaluated "
        "at a UQ-specific convergence criterion."
    )
    st.page_link("pages/UQLeaderboard.py", label="Go to UQ Leaderboard β†’")

st.divider()

st.subheader("Benchmarks")
st.markdown(
    "Results are reported on four benchmark configurations of the [Lorenz system]"
    "(https://en.wikipedia.org/wiki/Lorenz_system), a standard testbed for "
    "data-assimilation and calibration algorithms:"
)

_media = Path(__file__).parent / "media"

st.markdown("- **L63** β€” Lorenz-63 β€” 3-variable chaotic attractor; learn 2 parameters, strongly nonlinear.")
st.image(
    str(_media / "posterior_ribbons_20_13_k5.png"),
    caption="Example prior-posterior & truth. L: Difference to true parameter. R: data-sample/output predictived distribution (state-mean [1:3],  state-covariance (diag [4:6] and off-diag [7:9]))",
    width=800,
)


st.markdown("- **L96** β€” Lorenz-96 (40-variable); learn 1-parameter constant forcing.")
st.image(
    str(_media / "posterior_ribbons_const-force_12_1_k3.png"),
    caption="Example prior-posterior & truth. L: Difference to true parameter. C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:40] state-mean [41:80] state-std)",
    width=1200,
)

st.markdown("- **L96_SPATIAL_FORCING** β€” Lorenz-96 (40-variable) with spatially-varying forcing; learn 40 parameters; moderately correlated prior.")
st.image(
    str(_media / "posterior_ribbons_vec-force_65_1_k3.png"),
    caption="Example prior-posterior & truth. L: Difference to true parameters. C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:40] state-mean [41:80] state-std)",
    width=1200,
)


st.markdown("- **L96_NN_FORCING** β€” Lorenz-96 (100-variable) with a neural-network forcing; Learn 61 parameters (weights and biases of the network). Reasonable prior given.")
st.image(
    str(_media / "posterior_ribbons_flux-force_80_1_k3.png"),
    caption="Example prior-posterior & truth. L: Difference to true parameters (weights). C: parameter-induced forcing of the L96 system. R: data-sample/output predictived distribution ([1:100] state-mean [101:200] state-std)",
    width=1200,
)

st.subheader("Method families")
st.markdown(
    "Methods are grouped into three families. See the **πŸ“˜ Methods** page for "
    "citations and per-method performance charts."
)
families = {
    "Kalman": "Ensemble Kalman variants (TEKI, ETKI, IEKF, UKI) β€” update an ensemble of parameter guesses via a linearized observation operator.",
    "Bayesian": "Sampling-based approaches (ABC, HM) β€” explore parameter space without requiring gradient information.",
    "Calibrate-then-emulate": "Two-stage pipelines (CES-EKI-DMC) β€” use an initial calibration phase to build a cheap emulator, then sample the posterior via MCMC.",
}
for family, desc in families.items():
    st.markdown(f"- **{family}** β€” {desc}")

st.subheader("Key metric")
st.markdown(
    "The reported metric is the **mean number of forward-model evaluations** required "
    "to reach the target, averaged over random seeds. "
    "A value of **βˆ’1** indicates a failed run (did not reach the target); "
    "the **failure rate** shows the fraction of seeds that failed."
)