How Wrong Can a Good Predictor Be?
Diverging Updates with
Vanishing Predictive KL

Qifu Wen*  ·  Shuaijun Liu*  ·  Zihan Zhou  ·  Xi Zeng  ·  Ningxin Su
* Co-first authors
arXiv PDF Code pending BibTeX

A predictor can get the answer right without copying Bayes inside. For every fixed finite hidden-state count K ≥ 2, we construct a filter whose internal update gap grows without bound, yet its decoded predictions become arbitrarily close to Bayes.

Diverging updatesAt explicit inputs, the two update rules move farther apart as switches become rarer.
Typical pathsBoth filters become confident in the same state on typical rare-switching windows.
Vanishing predictive KLSoftmax is insensitive there, so the remaining internal error becomes predictively cheap.

Large internal differences need not mean poor predictions. What matters is how those differences affect the probabilities we use.

The paper's current finite example

For K = 2, 4, and 8, Exact Bayes and the radial filter see the same observations from equally spaced Gaussian states. Across 4,096 paired stationary paths per setting, centered internal distance grows while exact-to-radial categorical KL falls.

2 hidden statesThe binary case makes the geometry easiest to inspect.Internal gap up · predictive KL down
4 hidden statesThe same opposing trends appear beyond the binary case.Internal gap up · predictive KL down
8 hidden statesThe trend remains visible, but convergence is slower at finite scale.Internal gap up · predictive KL down

This is a measured illustration, not the proof. The theorem covers every fixed finite state count as switching becomes rarer; it does not give a uniform rate as the state count grows.

Binary microscope: what happens around a switch?

This separate 2-hidden-state visualization uses two imposed switches to make adaptation visible. Every method sees the same evidence. It illustrates a long-horizon recovery question and is not a proved recovery comparison or the fixed-state theorem.

The whole idea

Why can different models make the same prediction?

The score does not inspect every internal difference. It sees the final prediction on the paths that actually occur.

Exact update Restricted update

Different inside

The restricted model cannot copy the exact internal update rule.

Typical evidence pushes both right

Same confident side

On the paths that dominate, both models become confident in the same answer.

Internal gap Prediction gap

Small difference outside

The probability decoder is flat there, so the remaining internal gap costs little.

The paper proves a diverging update gap with vanishing predictive KL at selected inputs, and expected terminal KL convergence along stationary Gaussian HMM paths. As switches become rarer, the theorem's growing observation windows contain almost no switches. The result is not uniform as the number of states grows.