tsfm-onnx: t0-alpha for the browser

Browser-ready ONNX exports of t0-alpha, the zero-shot time-series foundation model by The Forecasting Company. These graphs run a full probabilistic forecast in about 100 ms under onnxruntime-web (WASM), with no server and no Python.

This is an unofficial conversion, not affiliated with or endorsed by The Forecasting Company. The conversion pipeline, a browser app using these files, an export guide, and a debugging logbook live at github.com/siddharth7113/tsfm-onnx.

Files

File Size Graph Precision
t0-alpha-ctx512-h64.onnx 411 MB univariate fp32
t0-alpha-ctx512-h64-int8.onnx 108 MB univariate int8 (dynamic, per-channel)
t0-alpha-ctx512-h64-mv.onnx 411 MB grouped (multivariate) fp32
t0-alpha-ctx512-h64-mv-int8.onnx 108 MB grouped (multivariate) int8 (dynamic, per-channel)

Graph contract

Univariate graph:

input   context    float32 [batch, 512]    NaN marks missing values
output  quantiles  float32 [batch, 64, 5]  levels [0.1, 0.25, 0.5, 0.75, 0.9]

Grouped (multivariate) graph adds one input:

input   group_ids  int64   [rows]

Rows that share a group id are forecast jointly as variates of one system (they inform each other through the model's group attention); rows with distinct ids are independent series. Flatten [B, V, T] to [B*V, T] and pass ids like [0, 0, 0, 1, 1, 1].

Practical notes:

  • Series shorter than 512 points: LEFT-pad with NaN. The model treats NaN as "missing" and was trained on gappy data, so this is a valid input, not an approximation. Series longer than 512: pass the most recent 512.
  • The context length (512) and horizon (64) are baked into the graphs. Other sizes require re-export (see the GitHub repo).
  • Known future covariates are not exported.

Usage

JavaScript (onnxruntime-web):

const session = await ort.InferenceSession.create(modelUrl, { executionProviders: ["wasm"] });
const ctx = new Float32Array(512).fill(NaN);
ctx.set(series.slice(-512), 512 - Math.min(series.length, 512));
const out = await session.run({ context: new ort.Tensor("float32", ctx, [1, 512]) });
// out.quantiles.data is row-major [batch, step, level]; element (s, l) is at s * 5 + l

Python (onnxruntime):

import numpy as np, onnxruntime as ort

sess = ort.InferenceSession("t0-alpha-ctx512-h64.onnx")
context = np.full((1, 512), np.nan, dtype=np.float32)
context[0, -len(series):] = series[-512:]
(quantiles,) = sess.run(None, {"context": context})   # (1, 64, 5)

Fidelity

Check Result
fp32 ONNX vs library model.predict(), identical input max abs diff 1.7e-05
grouped graph vs multivariate predict(), both id modes max abs diff 1.7e-05
onnxruntime-web (WASM) vs native ONNX Runtime at most 1.5e-05
int8 vs fp32 forecast drift about 1% mean of forecast spread on long series; up to 3% mean / 17% max on short NaN-padded series

Every export is validated against the original tfc-t0 library at export time; the numbers above are reproducible from the scripts in the GitHub repo. Prefer fp32 when fidelity matters more than the download size.

How these were made

Exported with the PyTorch dynamo exporter from tfc-t0 0.2.3 (torch 2.8), wrapping the library's single-forward-pass inference path with branch-free equivalents of its data-dependent Python. Quantization is ONNX Runtime dynamic int8 with per-channel weights. The full worked case study (including every failure and fix) is in the repo's export guide and logbook.

License and attribution

The t0-alpha weights are released by The Forecasting Company under Apache-2.0 (with gated access on the original repository); these files are a converted redistribution of those weights under the same license, with attribution. Per the tfc-t0 source headers, parts of the architecture derive from Toto (Datadog) and Chronos (Amazon Science), both Apache-2.0. If you use these files, please credit The Forecasting Company and consider citing their model card.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Siddharth899/tsfm-onnx

Quantized
(1)
this model