What different models share
Different models agree more about their outputs.
The same 200 text prefixes are shown to each model. Each dot compares the ordering of distances between those prefixes in two models.
Loading comparisons…
Predictive errors explain held-out agreement.
Loading recorded measurements…
How close are models to human completions?
Loading recorded measurements…
A fresh test against human responses.
Five previously unused stories provide 640 positions and 64,000 human responses. Compare model scale, training progress and a calibration learned before this test.
Loading human comparisons…