Figure 4a; Extended Data 1b,c

Shared output structure

Mean output rank agreement is 0.88 across 29 cross-family, cross-tokenizer dyads, compared with 0.62 and 0.61 for activation geometry. First-byte coarse-graining gives agreement of 0.91 on its matched subset.

Different models agree more about their outputs.

The same 200 text prefixes are shown to each model. Each dot compares the ordering of distances between those prefixes in two models.

Loading comparisons…

Loading recorded comparison…

Does the comparison depend on the measurement?

Loading recorded comparison…

Text and reference controls

Loading recorded comparison…

What changes when the readout is centred?

Loading recorded comparison…

A shared structure across words and images

Loading recorded comparison…

Methods and interpretation

Relational agreement compares distances within each model. The common-byte control uses 21 dyads among eight models, a subset distinct from the main 29-pair panel.

Procedure in the paper: Prepare the natural context battery, extract model output laws and activation baselines, construct relational matrices and apply the frozen dyad selection, resampling and common-byte map.

Code and data

Shared geometry, dense gauges and common bytesZIP

v0.1.0 · 71 KB · View source on GitHub ↗

Ten-model output/mid/last convergence with corrected 5,000-draw context bootstrap, the complete eight-model 27-cell dense gauge sweep, and complete common-byte contraction, covariance, risk-floor and 1,000-subsample stability analyses.

Shared decision axesZIP

v0.1.0 · 69 KB · View source on GitHub ↗

Fixed-prompt and natural-text axis sharing, salience-matched nulls, the corrected natural-context bootstrap and all six rank bands.

Geometry comparison controlsZIP

v0.1.0 · 91 KB · View source on GitHub ↗

Representation metrics, coordinate changes, natural-text and n-gram baselines, corpus mediation, reference scale, output distances and permutation calibration.

Unembedding centering controlsZIP

v0.1.0 · 66 KB · View source on GitHub ↗

Vocabulary centering, word-frequency controls and the decision-axis anisotropy comparison.

Language–vision geometryZIP

v0.1.0 · 56 KB · View source on GitHub ↗

The 100-concept DINOv2 comparison, text co-occurrence controls and larger-vision-model sweep.

Each standalone package includes code, shared helpers, required small inputs and reference results, with setup and commands in its README. Model weights and public datasets are obtained separately where needed.

Download complete source (v0.1.0) for all experiments and the companion website.