A predictor can get the answer right without copying Bayes inside. For every fixed finite hidden-state count K ≥ 2, we construct a filter whose internal update gap grows without bound, yet its decoded predictions become arbitrarily close to Bayes.
Large internal differences need not mean poor predictions. What matters is how those differences affect the probabilities we use.
For K = 2, 4, and 8, Exact Bayes and the radial filter see the same observations from equally spaced Gaussian states. Across 4,096 paired stationary paths per setting, centered internal distance grows while exact-to-radial categorical KL falls.
This is a measured illustration, not the proof. The theorem covers every fixed finite state count as switching becomes rarer; it does not give a uniform rate as the state count grows.
This separate 2-hidden-state visualization uses two imposed switches to make adaptation visible. Every method sees the same evidence. It illustrates a long-horizon recovery question and is not a proved recovery comparison or the fixed-state theorem.
The score does not inspect every internal difference. It sees the final prediction on the paths that actually occur.
The restricted model cannot copy the exact internal update rule.
On the paths that dominate, both models become confident in the same answer.
The probability decoder is flat there, so the remaining internal gap costs little.
The paper proves a diverging update gap with vanishing predictive KL at selected inputs, and expected terminal KL convergence along stationary Gaussian HMM paths. As switches become rarer, the theorem's growing observation windows contain almost no switches. The result is not uniform as the number of states grows.