Timer S1 Timer S1 is a time series foundation model with 8.3B total parameters, 0.75B activated parameters per token, and a context length of 11,520 . The model supports zero shot forecasting (predicting without dataset specific training) at different quantile levels. For more details, please refer to our technical report. Architecture : Timer S1 is a decoder only Mixture of Experts (MoE) Transformer. For time series forecasting (a sequential problem where each step depends on previous ones), we propose TimeSTP , enabling multi step prediction with cost effective serial computations . Performance : Timer S1 achieves state of the art results on GIFT Eval. The model excels particularly at medium term and long term forecasting tasks. Post Training : Timer S1 undergoes post training, including continued pre training ( CPT ) and long context extension ( LCE ), which improves short term and long context performance. Quickstart This model support inference using either CPU or GPU. To load this model on GPU, we recommend a GPU with at least 40GB VRAM (e.g., A100 40GB/80GB, or H100). Encounter out of memory at runtime? Try the following options: Specification Architecture : decoder only Tra…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy