Qwen-2.5-7B CM-Align (EN-pivot DPO)

Qwen/Qwen2.5-7B post-trained with CM-Align, an English-pivot preference-alignment method, as a baseline in a study of cross-lingual factual recall and consistency.

This is the Qwen counterpart of jvonrad/OLMo-2-7B-CM-Align. Weights are merged (LoRA adapter folded into the base model), so it loads like any standard Qwen2.5 checkpoint.

Method

Adaptation of CM-Align (Zhang et al., EMNLP 2025 Findings, arXiv:2509.08541) to a parallel multilingual factual-QA setting. The procedure is self-supervised — it never uses gold answer labels:

  1. Preference construction. For each fact, sample K=4 free-text answers per language (temperature 0.9, top-p 0.95). Embed all candidates with sentence-transformers/LaBSE. Choose as pivot the most self-consistent English candidate (highest mean cosine similarity to the other English candidates). Then, for every other language, take chosen = argmax and rejected = argmin cosine similarity to that English pivot.
  2. DPO training. Hand-rolled DPO on those preference pairs. The reference distribution is the same model with the LoRA adapter disabled, so no second copy of the model is held in memory. Objective: L_DPO + gamma * L_NLL.

CM-Align's original embedder was gte-multilingual-base; LaBSE is used here because it is the cross-lingual encoder used throughout this project.

Training details

Base model Qwen/Qwen2.5-7B
Data 40,000 facts from jvonrad/WIKI-FACT (train)
Languages 12 — en, de, id, pt, ar, bn, sw, es, ru, fr, ja, zh
Candidates / language 4 (temperature 0.9, top-p 0.95)
Embedder sentence-transformers/LaBSE
DPO beta 0.1
NLL gamma 0.0
Learning rate 5e-6
Epochs 1
Batch size 4 x 4 grad-accum
Precision bf16
LoRA r=64, alpha=128, dropout=0.05, on q/k/v/o/gate/up/down projections
Reference model same model, adapter disabled (disable_adapter())

Evaluation

Answers are scored by length-normalised log-likelihood over the four options, with the plain prompt Question: {question}\nAnswer:. Metrics beyond accuracy:

  • TotCons (Total Consistency) — fraction of facts answered correctly in all languages.
  • RankC — cross-lingual agreement between full option rankings (Qi et al., EMNLP 2023).
  • AnsAgr — pairwise answer agreement across language pairs.

PolyFact test (2,523 facts, 12 languages):

Model Acc TotCons RankC AnsAgr
Qwen-2.5-7B (base) 61.89 7.09 61.66 55.48
Qwen-2.5-7B CM-Align 64.82 10.07 64.03 58.65

Global-MMLU-Lite (400 facts, 11 languages — Lite has no Russian config):

Model Acc TotCons RankC AnsAgr
Qwen-2.5-7B (base) 63.20 13.50 69.70 64.49
Qwen-2.5-7B CM-Align 62.07 11.00 69.52 64.17

Per-language accuracy on PolyFact test:

en de id pt ar bn sw es ru fr ja zh
77.1 70.5 71.3 71.9 54.9 50.7 48.1 71.3 61.2 71.2 63.6 66.0

Reading these numbers

CM-Align gives consistent in-domain gains over the base model on PolyFact (all four metrics, e.g. Total Consistency 7.09 -> 10.07). Out of domain on Global-MMLU-Lite it does not transfer: every metric is flat or slightly below base. This model is published as a baseline for comparison, not as a recommended general-purpose checkpoint.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "jvonrad/Qwen-2.5-CM-Align"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

prompt = "Question: What is the capital of Poland?\nAnswer:"
out = model.generate(**tok(prompt, return_tensors="pt").to(model.device), max_new_tokens=16)
print(tok.decode(out[0], skip_special_tokens=True))

This is a base (non-instruct) model. Use plain completion-style prompts such as the Question: ... \nAnswer: form above rather than chat formatting.

Limitations

  • Trained only on Wikidata-derived factual QA in 12 languages; it is not a general instruction-following model.
  • Preference pairs are built from the model's own samples via embedding similarity, so they inherit LaBSE's similarity biases and can reward fluent-but-wrong answers.
  • Improved cross-lingual consistency can make incorrect factual associations more uniform across languages as well as correct ones.
  • Evaluation is multiple-choice log-likelihood scoring; it does not measure free-form generation quality.

Citation

The CM-Align method this baseline implements:

@inproceedings{zhang2025cmalign,
  title     = {CM-Align: Consistency-based Multilingual Alignment for Large Language Models},
  author    = {Zhang, Xue and others},
  booktitle = {Findings of EMNLP},
  year      = {2025}
}

The RankC consistency metric used above:

@inproceedings{qi2023crosslingual,
  title     = {Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models},
  author    = {Qi, Jirui and Fern{\'a}ndez, Raquel and Bisazza, Arianna},
  booktitle = {EMNLP},
  year      = {2023}
}
Downloads last month
28
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jvonrad/Qwen-2.5-CM-Align

Base model

Qwen/Qwen2.5-7B
Finetuned
(942)
this model

Datasets used to train jvonrad/Qwen-2.5-CM-Align

Paper for jvonrad/Qwen-2.5-CM-Align