Results; SI: identification
Assigning the language law
Changing the assigned language recovers more than 99% of the imposed geometric separation. Architecture changes within a language leave geometry nearly unchanged at matched accuracy.
Loading recorded measurements…
Methods and interpretation
Eight synthetic language pairs with matched token frequencies and conditional entropy, three architecture families and two capacities. This is a controlled language experiment.
Procedure in the paper: Generate the two fixed pilot language pairs and select each architecture/capacity schedule using pilot risk. Train on all eight evaluation pairs, retain every evaluation pair, and compare their root-probability distance matrices.
Assigned versus alternate law
96 trained models: eight language pairs × two assigned laws × three architectures × two capacities. Errors are normalised by the known separation within each pair. Language assignment is randomised while unigram frequency and entropy are matched.
Code and data
v0.1.0 · 70 KB · View source on GitHub ↗
The complete paired synthetic-language study, including pilot schedule selection and every evaluation pair.
Each standalone package includes code, shared helpers, required small inputs and reference results, with setup and commands in its README. Model weights and public datasets are obtained separately where needed.
Download complete source (v0.1.0) for all experiments and the companion website.