Time-Series Foundation Models (Chronos)

Summary

Chronos (Ansari et al., Amazon, TMLR 2024) asks: what if forecasting is just language modelling over a different vocabulary? Real values are mean-scaled and quantised into uniform bins on , giving a token sequence; an off-the-shelf T5 encoder-decoder (20M-710M parameters; also GPT-2) is trained with ordinary categorical cross-entropy to predict the next token. Forecasts are obtained by autoregressive sampling, dequantising, and unscaling. The model uses no time features, no covariates, no time-series-specific architecture. Trained on 28 public datasets (~890K series, ~84B observations) augmented by TSMixup (convex mixtures of real series) and KernelSynth (Gaussian-process samples from randomly composed kernels), it beats local statistical and task-specific deep models in-domain and is on par with the best trained deep models zero-shot on 27 unseen datasets.

Overview

Chronos sits at the end of the local → global → pretrained progression (Local vs Global Forecasting Models). DeepAR shares parameters across the series of one dataset; Chronos shares them across all datasets and is then applied, frozen, to data it has never seen — an “inference-only alternative to the conventional approach involving training and tuning a model on individual tasks” (Sec. 7). The authors’ design philosophy is minimalism: since transformers (Transformers and LLM Foundations - Overview) excel on token sequences, change the data representation, not the model.

The key conceptual point (Sec. 2-3): a language model predicts a categorical distribution over a finite vocabulary; a time series is real-valued and unbounded. Tokenisation bridges this gap, and the consequence is regression via classification — the predictive distribution is a free-form histogram over bins, able to represent skewed or multimodal futures without choosing a likelihood family.

Main Content

Tokenisation: scaling then quantisation (Sec. 3.1) ^def-chronos-tokenisation

Given context and horizon :

Mean scaling. with (the affine map with ). It preserves zeros, which are “often semantically meaningful, such as zero sales.”

Quantisation. Choose bin centres and edges ; then

Chronos uses uniform binning (centres equally spaced on , ) rather than quantile binning, because downstream data distributions differ from training. The vocabulary has 4096 entries: 4094 bins plus PAD (padding and missing values) and EOS.

Objective (Sec. 3.2) ^def-chronos-objective

With the token sequence, the model outputs a categorical over and minimises

This loss is not distance-aware: it does not know bin is closer to than to ; the model must learn bin topology from data. In scoring-rule terms it is the log score on the discretised outcome.

Zero-shot forecasting (Sec. 3.3) ^alg-chronos-forecast

  1. Take the last observations; compute ; scale and quantise.
  2. Autoregressively sample tokens from ; repeat for sample paths (20 in the paper’s evaluation).
  3. Dequantise and multiply by .
  4. Summarise: median for a point forecast; empirical quantiles for intervals.

TSMixup (Sec. 4.1, Alg. 1) ^alg-tsmixup

Draw () and length ; sample series of length from random training datasets; mean-scale each; draw weights ; return

Original series appear with probability ().

KernelSynth (Sec. 4.2, Alg. 2) ^alg-kernelsynth

From a kernel bank (constant, white noise, linear for trend, RBF for smooth local variation, rational quadratic, periodic kernels at typical seasonal periods) draw kernels with replacement; combine them left to right with random binary operators in into ; sample . This inverts the Automatic Statistician: instead of searching kernel compositions to explain a series, random compositions generate series.

Training setup (Sec. 5.2). T5 Mini (20M), Small (46M), Base (200M), Large (710M) and GPT-2 (90M); 10M TSMixup augmentations plus 1M KernelSynth series sampled 9:1; context 512, prediction length 64; 200K steps of AdamW, learning rate linearly annealed, batch 256; 8×A100. Training cost ranges from 7.7 h ($252, Mini) to 63 h ($2,066, Large) (Table 6).

Results (Sec. 5.5; Tables 7-10)

Metrics are WQL on quantiles (probabilistic) and MASE (point), each divided by Seasonal Naive’s score and aggregated across datasets by geometric mean (Forecast Evaluation and Backtesting). Lower is better; Seasonal Naive = 1.000.

Aggregate relative scoreT5-LargeT5-BaseT5-SmallT5-MiniGPT-2
Benchmark I (15 in-domain datasets) — WQL0.5640.5800.6030.5980.623
Benchmark I — MASE0.6950.7060.7270.7320.741
Benchmark II (27 zero-shot datasets) — WQL0.6450.6620.6670.6780.687
Benchmark II — MASE0.8230.8320.8410.8500.852
  • In-domain, the larger Chronos models beat local models (AutoETS, AutoARIMA, AutoTheta), task-specific deep models (DeepAR, TFT, PatchTST, N-BEATS, N-HiTS, DLinear, WaveNet) and other pretrained models (Lag-Llama, Moirai); even Chronos-Mini (20M) beats Moirai-Large (311M).
  • Zero-shot, Chronos “significantly outperform[s] standalone local statistical models,” takes 2nd-4th place on WQL and 2nd on MASE among all methods, and clearly beats LLMTime and ForecastPFN. Naive scores 1.152 (WQL) / 1.188 (MASE) on this benchmark.
  • Fine-tuning Chronos-T5-Small for 1,000 steps per dataset moves it to first place on Benchmark II.

Ablations (Sec. 5.6)

  • Size: loss and downstream scores improve monotonically from 20M to 710M.
  • LLM initialisation: starting from text-pretrained T5 weights gives no benefit over random initialisation (slightly higher final loss) — the transferable asset is the architecture, not the language knowledge.
  • TSMixup leaves in-domain performance unchanged but improves zero-shot; KernelSynth helps both, best at ≈10% synthetic data, degrading beyond. A model trained on synthetic data only is still better than ForecastPFN and several baselines.
  • Context helps up to 1024; vocabulary size trades discretisation error against sparsely populated bins (MASE improves with more bins; WQL is non-monotone).

Limitations (Sec. 5.7, 6.1) ^chronos-limitations

  • Bounded range: representable values lie in . Sparse spiky series (small ) overflow; strong or exponential trends are under-predicted (suggested fix: log-transform first). Short contexts lead to underestimated trend.
  • Precision: token spacing is ; a large-mean, small-variance signal collapses into a few tokens (suggested fix: standardise instead of mean-scale).
  • Univariate, no covariates: prices, promotions, media or holidays cannot be injected. The authors suggest task-specific adaptors or stacking Chronos with a covariate model such as LightGBM.
  • Inference cost: slower than task-specific deep models, comparable to some local statistical models.
  • Leakage: with public benchmark corpora, a truly clean zero-shot split requires test data after the pretraining data ends (footnote 5).

Examples

Synthetic diagnostics (Sec. 5.7, Figs. 12-14). On i.i.d. and noise the 80% interval matches the true interval. Linear trend, single and triple seasonality, and additive or multiplicative trend×season composites are forecast accurately; exponential trend is not. On stationary AR() data a correctly specified AR model beats Chronos for , but for AR(3)-AR(4) Chronos-Base beats AutoARIMA and matches the correctly-ordered fitted AR model — it has learned generic autoregressive structure.

Tokenisation arithmetic. Weekly sales with mean absolute value : bins are spaced units apart and the ceiling is . A Black-Friday week of 3,500 units is unrepresentable and clips to 3,000 — a concrete reason to backtest zero-shot models around promotional peaks.

import torch
from chronos import ChronosPipeline
pipe = ChronosPipeline.from_pretrained("amazon/chronos-t5-small")
samples = pipe.predict(torch.tensor(y_hist[-512:]), prediction_length=13, num_samples=20)
q10, q50, q90 = np.quantile(samples[0].numpy(), [0.1, 0.5, 0.9], axis=0)

Connections

See Also