File size: 3,392 Bytes
c29fb7f
 
 
60bd556
 
 
c99f474
78bc9ab
 
 
c99f474
 
 
c29fb7f
 
 
9ebdf13
 
fcf25f2
 
 
 
 
9ebdf13
fcf25f2
 
 
9ebdf13
fcf25f2
 
 
 
 
 
5063745
66928b3
c29fb7f
6e22424
c99f474
 
 
 
 
 
7ce86b7
 
 
6e22424
 
c99f474
 
 
c29fb7f
 
 
 
 
 
 
 
eda9bad
c29fb7f
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
TITLE = """ TSFM Realworld Bench Leaderboard """

INTRODUCTION_TEXT = """
**TSFM Realworld Bench** evaluates time series foundation models on **live TS-Bench
real-world data** with **zero-shot API inference** via [TSFM.ai](https://tsfm.ai/).
Following [GIFT-Eval](https://huggingface.co/spaces/Salesforce/GIFT-Eval), the leaderboard
reports **absolute metric values** and **per-dataset ranks**. The Overall tab also reports
the latest overall snapshot plus full 24-hour, 7-day, and 30-day metric tables and a
dataset-balanced **pairwise historical ranking** over all shared releases through the
current cutoff. The
**GIFT-style Aggregates** tab provides Seasonal-Naive-normalized MSE, CRPS,
and mean CRPS rank grouped by actual prediction length, domain, and frequency. Each
subsequent domain tab retains the original absolute per-dataset results.
"""

LLM_BENCHMARKS_TEXT = """
## How to participate (Submit a New Model)

Community models are evaluated through owner-operated HTTPS endpoints. Copy the
portable FastAPI template in `examples/community_endpoint/`, replace its
`forecast_one` function, and deploy it on your own inference server or existing
cloud service. A paid Hugging Face Space is optional and a Static Space cannot
run the forecasting API.

The endpoint receives only causal history, prediction length, frequency, and
requested quantiles. It never receives future labels. Servers behind NAT can use
a stable named HTTPS tunnel; temporary Quick Tunnel URLs are not accepted.

Before submitting, run `scripts/validate_external_model_endpoint.py` against the
public `/forecast` URL and generate `community_model.yaml` with
`scripts/generate_community_model_metadata.py`. Submit the public model card,
endpoint-code URL, endpoint URL, metadata, and successful validator receipt
through the community model request form. Accepted entries are evaluated only
on future live releases.

## Metrics

- **MSE** β€” Mean Squared Error on the mean forecast (absolute)
- **RMSE** β€” Root Mean Squared Error on the mean forecast
- **MAPE** β€” Mean Absolute Percentage Error, reported only away from zero
- **CRPS** β€” quantile approximation of the Continuous Ranked Probability Score
- **RTG** β€” normalized real-time MSE gain over causal Seasonal-Naive (higher is better)
- **Stability** β€” standard deviation of release-level MSE (lower is better)
- **Improvement** β€” Kendall trend statistic over release-level MSE (more negative is better)
- **Pairwise Win Rate** β€” dataset-balanced MSE/CRPS wins over shared future releases;
  official ranks require sufficient shared releases, datasets, time span, opponents,
  and membership in the main comparison component
- **RankScore** β€” Elo-style aggregate from per-dataset MSE and CRPS ranks (higher is better)
- **MSE_Rank** / **CRPS_Rank** β€” per-dataset rank (lower is better)
- **Grouped MSE / CRPS** β€” geometric mean after per-configuration normalization against
  Seasonal-Naive (lower is better; 1.0 equals the baseline)
- **Grouped Rank** β€” mean per-configuration CRPS rank (lower is better)
"""

CITATION_BUTTON_LABEL = "Copy citation"
CITATION_BUTTON_TEXT = r"""
@misc{tsfm_realworld_bench,
  title={TSFM Realworld Bench: A Benchmark for Time Series Foundation Models},
  author={TSFM Realworld Bench Team},
  year={2026},
  howpublished={\url{https://huggingface.co/spaces/CityMindDev/TSFM-Realworld-Bench}}
}
"""