Table 1; Extended Data 5a,b and 8

Steering, composition and magnitude

Instruction, factual and style frontiers show lower off-target divergence at matched mean objective change. The advantage varies with injection depth and attenuates as intervention magnitude grows.

Loading recorded comparison…

The full cubic correction

Loading recorded comparison…

Curvature and erosion depend on the regime

Loading recorded comparison…

Distributional change and coarse accuracy

Loading recorded comparison…

The six-size factual preference sweep

Loading recorded comparison…

Instruction-model objective frontiers

Loading recorded comparison…

Cost of combined instruction objectives

Loading recorded comparison…

Methods and interpretation

Frontiers match aggregate mean changes under their stated protocols. The layer and magnitude studies use their specified prompt sets and intervention strengths.

Procedure in the paper: Reproduce the fixed prompt batteries, objective definitions, target grids and layer settings. Apply each frontier’s matching and aggregation rules separately.

Code and data

Full internal-layer cubic correctionZIP

v0.1.0 · 56 KB · View source on GitHub ↗

Full fixed four-prompt Pythia-410m audit at layers 4, 12, 20 and the final readout, with separate AC, map and full coefficients and 64 parity-cancelling finite differences.

Cubic geometry across model sizesZIP

v0.1.0 · 62 KB · View source on GitHub ↗

Expanded-range measurements across five Pythia sizes, with complete first-crossing, monotone and partial-rank analyses.

Instruction and preference controlZIP

v0.1.0 · 77 KB · View source on GitHub ↗

Sycophancy, truth and style frontiers; weighted objective composition; the 20-question capability check; and six-size factual preferences.

Objective cost, sharing and transferZIP

v0.1.0 · 66 KB · View source on GitHub ↗

Four-model, six-objective matched-effect steering and its model/domain robustness analysis.

Each standalone package includes code, shared helpers, required small inputs and reference results, with setup and commands in its README. Model weights and public datasets are obtained separately where needed.

Download complete source (v0.1.0) for all experiments and the companion website.