Figure 7a
Predicted intervention advantage
The finite-damping prediction tracks 3,515 usable matched-response observations with median measured-to-predicted ratio 0.967 and log–log slope 0.982.
The geometry predicts the cost of ignoring it.
Two interventions reach the same calibrated target. Each dot compares their predicted cost ratio with the ratio measured after changing the model.
Loading interventions…
Methods and interpretation
98.6% of planned observations are usable. The regularised metric selects the direction; the output metric measures its predicted behavioural cost. Damping and reference coordinates remain fixed parts of the experiment.
Procedure in the paper: Compute the declared model/layer pullback operators, freeze the finite-damping ratio, solve matched target responses for both directions, and measure the stable off-target KL endpoint.
Code and data
v0.1.0 · 77 KB · View source on GitHub ↗
Predictions, reachability calibration, matched finite responses and primary and magnitude analyses across 1,188 evaluation cells in eleven models.
v0.1.0 · 77 KB · View source on GitHub ↗
Sycophancy, truth and style frontiers; weighted objective composition; the 20-question capability check; and six-size factual preferences.
v0.1.0 · 66 KB · View source on GitHub ↗
Four-model, six-objective matched-effect steering and its model/domain robustness analysis.
v0.1.0 · 62 KB · View source on GitHub ↗
SAE matching pursuit, interpretable token partitions and token-grouping versus CG flow fidelity.
Each standalone package includes code, shared helpers, required small inputs and reference results, with setup and commands in its README. Model weights and public datasets are obtained separately where needed.
Download complete source (v0.1.0) for all experiments and the companion website.