QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities
East China Normal University Β· Huawei Technologies Co., Ltd.
Overview
Existing multimodal time series foundation models typically model heterogeneous modalities through largely shared mechanisms, overlooking their distinct roles in forecasting.
We propose QiYao-M, a role-aware multimodal time series foundation model that separately models:
- Endogenous modalities, which evolve together with the underlying temporal dynamics.
- Exogenous modalities, which provide external information and vary substantially across domains in modality type and number.
QiYao-M introduces dedicated modeling and training strategies for these two types of multimodal information, enabling effective forecasting both with and without exogenous modalities.
Highlights
π§© Endogenous Multimodal Modeling
We introduce:
- Endo-Multimodal Fusion
- Endo-Multimodal Predictor
- Endo-Multimodal Supervision
These components explicitly model how endogenous modalities evolve from the historical window to the forecasting horizon.
π Exogenous Multimodal Retrieval
We propose an Exo-Multimodal Retrieval Enhancer that retrieves historical cases with similar multimodal conditions and uses their future responses as forecasting evidence.
The retrieval mechanism supports various exogenous modality types and numbers without updating the TSFM parameters.
π§ Endo-Modality Proxy Training
To address the scarcity of exogenous multimodal pretraining data, we introduce Endo-Modality Proxy Training, which dynamically samples endogenous modalities as retrieval proxies during pretraining.
This enables the retrieval module to generalize to unseen domains and unseen modality configurations.
π Performance
QiYao-M demonstrates strong forecasting performance across both unimodal and multimodal benchmarks.
Overall Performance
Unimodal Forecasting
GIFT-Eval
TIME
Multimodal Forecasting
Citation
If you find this work useful, please consider citing:
@article{cheng2026qiyaom,
title = {{QiYao-M}: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities},
author = {Hanyin Cheng and Linfeng Wang and Zhengbo Qu and Yang Shu and Zhongwen Rao and Meng Wang and Yijie Li and Xin Jiang and Bin Yang and Chenjuan Guo},
journal = {arXiv preprint},
year = {2026}
}