Extended Data 2b–d; SI: human protocols
Fresh human replication and calibration
The scale and training trends replicate on 640 positions. A source-trained full 65-outcome calibration improves expected human log score in six of seven families; Qwen2 has a small negative gain.
A fresh test against human responses.
Five previously unused stories provide 640 positions and 64,000 human responses. Compare model scale, training progress and a calibration learned before this test.
Loading human comparisons…
Predicting human alignment in a held-out family
Loading recorded comparison…
Methods and interpretation
64,000 responses from five fixed stories. Distance estimates have confidence intervals; calibration gains are reported for each story.
Procedure in the paper: Acquire the fresh human responses and frozen positions, reconstruct native model samples and the common 65-outcome chart, then apply the earlier fixed calibration without fitting to the new human responses.
Code and data
v0.1.0 · 12.4 MB · View source on GitHub ↗
Source calibration fitting from the included 1,726-context training laws, fresh evaluation of seven frozen family maps on 640 contexts, and the separate 21-model native/conditional human-geometry replication; fresh model probabilities and samples are regenerated.
v0.1.0 · 1.4 MB · View source on GitHub ↗
Direct human geometry, calibration decomposition, random-subspace controls and crossed bootstrap intervals on 1,726 contexts.
Each standalone package includes code, shared helpers, required small inputs and reference results, with setup and commands in its README. Model weights and public datasets are obtained separately where needed.
Download complete source (v0.1.0) for all experiments and the companion website.