File size: 9,430 Bytes
807c1c3 6184b45 807c1c3 6184b45 807c1c3 6184b45 807c1c3 6184b45 807c1c3 6184b45 807c1c3 6184b45 807c1c3 6184b45 807c1c3 6184b45 807c1c3 6184b45 807c1c3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 | ---
license: apache-2.0
library_name: tinycast
pipeline_tag: time-series-forecasting
datasets:
- Salesforce/GiftEvalPretrain
tags:
- time-series
- time-series-forecasting
- foundation-model
- attention-free
- gift-eval
- edge
- streaming
- quantization
---
# TinyCast
**Probabilistic zero-shot forecasting at 146,505 parameters: the smallest GIFT-Eval entry that
publishes per-configuration results and declares no leakage.**
TinyCast forecasts a series it has never seen, with no fitting and no fine-tuning, and returns nine
quantiles rather than a single number. It is attention-free, and it computes each context's
periodicity instead of learning it, so no capacity is spent rediscovering seasonality.
**Below 1.4 M parameters it is the only zero-shot model on the GIFT-Eval board that emits a
predictive distribution**, and it fits on a microcontroller.
- **Paper:** https://arxiv.org/abs/2608.15767
- **Code, training and replication:** https://github.com/raws-labs/tinycast
- **Parameters:** 146,505 (fp32 weights, about 0.6 MB)
- **GIFT-Eval, zero-shot:** 0.774 nGMASE, 0.545 nWQL, 0.554 nMSIS over 97 configurations
- **On a Cortex-M7:** 138.1 KiB INT8 weights, 730.7 KiB peak RAM, 4.08 s per call
([firmware profile](#on-device), 0.833 nGMASE)
- **License:** Apache-2.0
## GIFT-Eval results (zero-shot)

Point accuracy (left) and probabilistic accuracy (right) against parameter count, log scale, over
the 97 GIFT-Eval configurations. Filled marker is the host profile, hollow is the firmware profile.
Dashed outlines mark models that emit a point forecast, whose nWQL column is a point error and is
not comparable.
| Metric | Value |
|--------|-------|
| nGMASE (point accuracy) | 0.774 |
| nWQL (probabilistic accuracy) | 0.545 |
| nMSIS (interval score) | 0.554 |
All three are geometric means, over the 97 benchmark configurations, of the ratio between the
model's metric and the seasonal-naive reference's; 1.000 is parity with seasonal naive. Every number
on this page comes from one profile: bf16 autocast at compute, flip-invariance symmetrization and
period-alignment downsampling, which the reproduction command below runs.
Against every zero-shot model on the board up to 10 M parameters with a public per-configuration
result and no declared leakage, recomputed from one pinned snapshot against the same reference:
| Model | Params | nGMASE | nWQL |
|---|---:|---:|---:|
| **TinyCast** | 146 K | 0.774 | 0.545 |
| Reverso-Nano | 200 K | 0.760 | (0.661) |
| Reverso-Small | 550 K | 0.726 | (0.626) |
| TTM-R3 | 1.4 M | 0.724 | 0.520 |
| Reverso | 2.6 M | 0.711 | (0.610) |
| Toto-2.0-4m | 4.1 M | 0.757 | 0.524 |
| YingLong-6m | 7.3 M | 0.880 | 0.609 |
| FlowState-9.1M | 9.1 M | 0.726 | 0.502 |
| Kairos-10m | 9.9 M | 0.753 | 0.554 |
Parenthesized nWQL means the model emits a point forecast rather than a predictive distribution, so
the figure is a point error and is not comparable with the rest. The two Reverso models are the only
others here below 1.4 M parameters, and both are in that category.
**Other benchmarks.** On Chronos-ZS (27 tasks) TinyCast reaches relative MASE 0.880 and relative WQL
0.722; on fev-bench (100 tasks) relative MASE 0.819, relative WQL 0.658 and a skill score of 0.304.
On both, every neural model ahead of it carries at least 28 times its parameters. Each benchmark
normalizes over its own dataset set, so these aggregates are not comparable with the GIFT-Eval
figures above. Disjointness is not established on fev-bench: twelve of its hundred tasks name corpus
subsets we train on.
## Model description
| Component | Design |
|---|---|
| Backbone | 10 dilated causal Conv1d blocks, kernel 3, dilations doubling from 1 to 512 (receptive field 2047 over a context of 2048) |
| Efficiency | depthwise-separable convolutions, one SwiGLU feed-forward ALBERT-tied across all ten blocks |
| Structural prior | Fisher's significance test for harmonic analysis on the normalized periodogram, alpha 0.05, up to four periods, zero parameters, then a 16-bin phase fold |
| Decoder | pooled summary, phase-gather seasonal profile and a causal future-conv correction; nine decile quantiles |
| Normalization | per-context min-max over the observed history, inverted on the output |
| Context / horizon | context 2048; forecasts in blocks of 48 steps, rolled out autoregressively (evaluated to 720) |
| Working memory | bounded per-layer ring buffers that do not grow as the model runs; causal padding, so a per-step streaming variant is exact rather than approximate |
Every learned operation is a convolution, a matrix multiplication or a normalization, so the model
exports to a static INT8 graph and runs a forecast end to end on an embedded device.
### On device
Deployed on an STM32H753 (Arm Cortex-M7) development board as a static W8A8 graph with quantization
scales calibrated once and frozen. Single core, no neural accelerator, no off-chip memory.
| Measurement | Value |
|---|---|
| INT8 matrix and convolution coefficients | 138.1 KiB |
| Complete firmware image, including an 8 KiB context | 365.5 KiB |
| Peak RAM (statics, heap, stack high-water) | 730.7 KiB |
| One core call at a 2048-step context | 4.08 s |
The board runs the **firmware configuration**, which scores **0.833 nGMASE and 0.581 nWQL**, not the
0.774 / 0.545 headlined above: it executes INT8 and drops both inference-time strategies.
Quantization alone, with the strategies kept, costs 2.1% of aggregate point accuracy over all 97
configurations.
**Training data.** Pretrained once on GIFT-Eval-Pretrain, Chronos KernelSynth and four synthetic
shards, for 36,621 steps in about 7.8 h on eight RTX 3090s, roughly 62 accelerator-hours. Before
caching, the corpus is filtered on two exclusion lists: of its 71 top-level directories two are
removed, leaving 69. GIFT-Eval-Pretrain is published already disjoint from the GIFT-Eval test set,
so the overlap list removes nothing here and exists for the secondary benchmark. No per-dataset
fine-tuning; every result above is zero-shot.
## Files
- `model.safetensors`: 146,505 fp32 parameters, about 0.6 MB. The weight-tied feed-forward is stored
once, so counting parameters by summing a loaded `state_dict()` overcounts; instantiate from the
config instead.
- `config.json`: the `TinyCastConfig` the loader rebuilds the architecture from.
## Usage
There is no PyPI package. Install from this repository:
```bash
git clone https://github.com/raws-labs/tinycast.git
cd tinycast && pip install -e .
```
Load the weights:
```python
from huggingface_hub import hf_hub_download
from tinycast import load_model
weights = hf_hub_download("raws-labs/tinycast", "model.safetensors")
hf_hub_download("raws-labs/tinycast", "config.json") # sibling, picked up automatically
model, config = load_model(weights) # TinyCastForPrediction, 146,505 params
```
Forecast with the gluonts predictor these results were produced with:
```python
from tinycast import TinyCastPredictor
predictor = TinyCastPredictor(prediction_length=48, checkpoint_path=weights,
freq="H", domain="Energy", device="cpu")
forecasts = predictor.predict(test_input) # QuantileForecasts, nine deciles
```
Reproduce the table above. The GIFT-Eval data loader is a separate install, needed only by the
benchmark driver; note that its distribution name and its import name differ:
```bash
pip install "salesforce-gift-eval @ git+https://github.com/SalesforceAIResearch/gift-eval.git"
# the loader does not fetch the data; download it once first
hf download Salesforce/GiftEval --repo-type dataset --local-dir /path/to/gift-eval
export GIFT_EVAL=/path/to/gift-eval
python -m tinycast.eval --ckpt model.safetensors --flip \
--device cuda --output all_results.csv
```
The repository README documents the thread caps that keep this from running many times slower, and
everything else that moves the numbers. It also covers the training recipe, checkpoint averaging and
export, and the synthetic-corpus builder, all of which ship in the same package.
## Intended use and limitations
TinyCast is built for forecasting a univariate signal on hardware that was not chosen for machine
learning, where a per-signal model would have to be fitted and maintained for every deployment.
- It is univariate. It uses no covariates and no cross-series structure, so a task that supplies
either will be forecast without them.
- The model emits no signal when its input leaves the regime its pretraining covers, so degradation
there is silent.
- The period comes from a rounded FFT bin, so its resolution falls as the ratio of window to period
falls: about seven percent for a weekly cycle in hourly data.
- Min-max normalization is per context window and therefore sensitive to a single extreme value.
- Fresh deployments degrade toward the seasonal-naive baseline rather than failing: parity at 64
observed samples, and two thirds of the way back to full-context accuracy by 512.
## License
Apache-2.0. Third-party attributions are in the code repository's `NOTICE`.
## Citation
```bibtex
@misc{tinycast2026,
title = {TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity},
author = {Armin Steinhauser},
year = {2026},
eprint = {2608.15767},
archivePrefix = {arXiv},
primaryClass = {cs.LG}
}
```
|