// Insight
Nine foundation models against HAR. The smallest one wins
Realized volatility is the fairest fight you can pick with a foundation model. The target persists. The benchmark is public. Corsi’s HAR model from 2009 has outlived seventeen years of people trying to beat it. Alessio Brini at Duke runs the fight properly. Nine zero-shot foundation models go against eight econometric specifications on 50 assets spanning equities, foreign exchange, and futures, at three horizons, with formal model-comparison tests.
Pooled across assets a broad tier of foundation models looks competitive with the HAR family. Weight every asset equally and one survivor is left out of nine.That gap between two ways of averaging carries the whole paper. Its lesson belongs to evaluation as much as to volatility. Pooled losses let a few high-volatility assets set the ranking. The leaders then sit so close together that the gaps rarely exceed two or three hundredths of a QLIKE point. A model that does well on the volatile names posts a good average while losing to the benchmark on most of the panel. Brini reports both aggregations, which is why the paper earns its afternoon.
Equal-weight each asset’s loss ratio against a well-specified Log-HAR and the field thins fast.
Tiny Time Mixers wins at all three horizons by roughly 1.3 to 1.8 percent. It is also the smallest model in the evaluation, under a million parameters. Chronos-Bolt, Moirai, TimesFM, and Toto all lose to Log-HAR on the typical asset at every horizon. Four of the nine enter the Model Confidence Set for zero of the 50 assets at the daily horizon.
Do not read that as classical safety. The worst model in the whole study is econometric. HARQ posts a loss ratio of 5.132 at the daily horizon, its realized-quarticity correction amplifying the noise it was built to tame. Specification quality decides these fights. It always has.
Then Brini takes the win apart. A Mincer-Zarnowitz recalibration strips level and scale bias from every forecast symmetrically. Much of the short-horizon advantage turns out to be better scaling rather than better prediction of the dynamics. At the daily horizon the econometric benchmarks forecast more efficiently. It is not close. Log-HAR rejects the joint efficiency null for 2 percent of assets. TTM rejects for 88 percent, despite a slope near unity. Only at the monthly horizon does a genuine informational gain survive. There the picture inverts. Log-HAR’s calibration slope collapses to 0.532 while TTM holds 0.670.
Blending beats both ingredients. An equal-weight average of TTM and Log-HAR posts loss ratios of 0.977, 0.984, and 0.982, under TTM’s own numbers at every horizon. It enters the Model Confidence Set for 98 to 100 percent of assets and beats Log-HAR significantly on 36 of the 50 names. You have to stop choosing.
I have watched three generations of this argument, from neural nets through gradient boosting to pretrained transformers. The shape repeats. A new class arrives and the pooled numbers look decisive. Then somebody reweights the panel. Brini’s most durable sentence is that variation across foundation-model architectures exceeds the difference between foundation models and econometrics as a class. The winner here carries fewer parameters than a mid-sized random forest.
A desk gets nothing glamorous from this. Log-HAR remains the thing to beat. A foundation model can add a point or two if you pick the right one and blend it. The burden of proof stays with the newcomer, which is roughly where this blog has landed all month. A specification choice reversed the sign of a causal driver. A search manufactured a backtest out of data holding nothing. Generators passed every fidelity check while destroying the patterns a fraud system reads. Here an average hides a panel. Look first at the unit you actually trade, then at the number that pools it away.
Pooled across 50 assets the foundation models look like they beat HAR. Weight each asset once and a seventeen-year-old benchmark holds off eight of nine, losing narrowly to the smallest model in the field.
Working on AI that needs to ship?
I help funds, fintechs, and data teams take AI from prototype to production.