--- library_name: transformers tags: [model-merging, mergeability, training-free, quotient-merge-distance] --- # beetle-humanscale-nld-eng__average__aligned Training-free merged checkpoint from the **Mergeability** sweep (`benchmark/emit_lm.py --real`), produced by weight-space merging of two independently trained parents. No gradient steps were taken. | field | value | |---|---| | pair_id | `beetle-humanscale-nld-eng` | | parent_a | `Beetle-HumanScale/beetle-monolingual-humanscale-nld` | | parent_b | `Beetle-HumanScale/beetle-monolingual-humanscale-eng` | | ceiling | `Beetle-HumanScale/beetle-bilingual-l2-50-simultaneous-b2-humanscale-nld-eng` | | operator | `average` | | alignment | `aligned` | | align_method | `permutation` | | regime | `shared_base` | | eval_langs | `eng+nld` | | nll_merge | `7.6155` | | nll_floor | `6.3583` | | param_coverage | `1.0` | | MS | `-4.1098` | ## How it was made Parents were loaded, activations extracted on a shared calibration corpus, and the merge applied either **naive** (parents combined in their own coordinates) or **aligned** (parent B carried into parent A's residual-stream basis via `common.alignment.residual_basis_map` before merging — permutation for same-width pairs, orthogonal/rectangular for cross-width). `MS` is the recovery score from `common.eval.mergeability_score` (merged vs. floor vs. ceiling), the same normalisation used by Zhou et al., so it is comparable across rows of the sweep. ## Caveats Sub-1B merges are noisy; an aligned signal where the naive one is noise is the finding, not a bug. Rows without a joint ceiling are floor-relative and must not be read as absolute recovery. Generated automatically — see the [mergeability repo](https://github.com/suchirsalhan/mergeability).