Learning over time

Predicting a preference before measuring it.

Corpus word-sequence statistics forecast the model’s preference for one continuation over another. The forecast is fixed before evaluation on these fresh examples.

Loading forecasts…

The same learning, seen in two coordinates.

Follow a shared set of 300 continuation preferences across three model sizes. Change the axis to compare training time with held-out predictive loss.

Deeper evidence makes learning arrive later.

Evidence depth is assigned at random across three architectures and two tasks. Compare the resulting delay with fixed-depth and shuffled-label controls.

Loading effects…