Instructions to use jvonrad/Qwen-2.5-CM-Align with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jvonrad/Qwen-2.5-CM-Align with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jvonrad/Qwen-2.5-CM-Align") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jvonrad/Qwen-2.5-CM-Align") model = AutoModelForCausalLM.from_pretrained("jvonrad/Qwen-2.5-CM-Align", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jvonrad/Qwen-2.5-CM-Align with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jvonrad/Qwen-2.5-CM-Align" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jvonrad/Qwen-2.5-CM-Align", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jvonrad/Qwen-2.5-CM-Align
- SGLang
How to use jvonrad/Qwen-2.5-CM-Align with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jvonrad/Qwen-2.5-CM-Align" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jvonrad/Qwen-2.5-CM-Align", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jvonrad/Qwen-2.5-CM-Align" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jvonrad/Qwen-2.5-CM-Align", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jvonrad/Qwen-2.5-CM-Align with Docker Model Runner:
docker model run hf.co/jvonrad/Qwen-2.5-CM-Align
Qwen-2.5-7B CM-Align (EN-pivot DPO)
Qwen/Qwen2.5-7B post-trained with CM-Align, an English-pivot preference-alignment method,
as a baseline in a study of cross-lingual factual recall and consistency.
This is the Qwen counterpart of jvonrad/OLMo-2-7B-CM-Align.
Weights are merged (LoRA adapter folded into the base model), so it loads like any standard
Qwen2.5 checkpoint.
Method
Adaptation of CM-Align (Zhang et al., EMNLP 2025 Findings, arXiv:2509.08541) to a parallel multilingual factual-QA setting. The procedure is self-supervised — it never uses gold answer labels:
- Preference construction. For each fact, sample
K=4free-text answers per language (temperature 0.9, top-p 0.95). Embed all candidates withsentence-transformers/LaBSE. Choose as pivot the most self-consistent English candidate (highest mean cosine similarity to the other English candidates). Then, for every other language, takechosen = argmaxandrejected = argmincosine similarity to that English pivot. - DPO training. Hand-rolled DPO on those preference pairs. The reference distribution is
the same model with the LoRA adapter disabled, so no second copy of the model is held in
memory. Objective:
L_DPO + gamma * L_NLL.
CM-Align's original embedder was gte-multilingual-base; LaBSE is used here because it is the
cross-lingual encoder used throughout this project.
Training details
| Base model | Qwen/Qwen2.5-7B |
| Data | 40,000 facts from jvonrad/WIKI-FACT (train) |
| Languages | 12 — en, de, id, pt, ar, bn, sw, es, ru, fr, ja, zh |
| Candidates / language | 4 (temperature 0.9, top-p 0.95) |
| Embedder | sentence-transformers/LaBSE |
| DPO beta | 0.1 |
| NLL gamma | 0.0 |
| Learning rate | 5e-6 |
| Epochs | 1 |
| Batch size | 4 x 4 grad-accum |
| Precision | bf16 |
| LoRA | r=64, alpha=128, dropout=0.05, on q/k/v/o/gate/up/down projections |
| Reference model | same model, adapter disabled (disable_adapter()) |
Evaluation
Answers are scored by length-normalised log-likelihood over the four options, with the plain
prompt Question: {question}\nAnswer:. Metrics beyond accuracy:
- TotCons (Total Consistency) — fraction of facts answered correctly in all languages.
- RankC — cross-lingual agreement between full option rankings (Qi et al., EMNLP 2023).
- AnsAgr — pairwise answer agreement across language pairs.
PolyFact test (2,523 facts, 12 languages):
| Model | Acc | TotCons | RankC | AnsAgr |
|---|---|---|---|---|
| Qwen-2.5-7B (base) | 61.89 | 7.09 | 61.66 | 55.48 |
| Qwen-2.5-7B CM-Align | 64.82 | 10.07 | 64.03 | 58.65 |
Global-MMLU-Lite (400 facts, 11 languages — Lite has no Russian config):
| Model | Acc | TotCons | RankC | AnsAgr |
|---|---|---|---|---|
| Qwen-2.5-7B (base) | 63.20 | 13.50 | 69.70 | 64.49 |
| Qwen-2.5-7B CM-Align | 62.07 | 11.00 | 69.52 | 64.17 |
Per-language accuracy on PolyFact test:
| en | de | id | pt | ar | bn | sw | es | ru | fr | ja | zh |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 77.1 | 70.5 | 71.3 | 71.9 | 54.9 | 50.7 | 48.1 | 71.3 | 61.2 | 71.2 | 63.6 | 66.0 |
Reading these numbers
CM-Align gives consistent in-domain gains over the base model on PolyFact (all four metrics, e.g. Total Consistency 7.09 -> 10.07). Out of domain on Global-MMLU-Lite it does not transfer: every metric is flat or slightly below base. This model is published as a baseline for comparison, not as a recommended general-purpose checkpoint.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "jvonrad/Qwen-2.5-CM-Align"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
prompt = "Question: What is the capital of Poland?\nAnswer:"
out = model.generate(**tok(prompt, return_tensors="pt").to(model.device), max_new_tokens=16)
print(tok.decode(out[0], skip_special_tokens=True))
This is a base (non-instruct) model. Use plain completion-style prompts such as the
Question: ... \nAnswer: form above rather than chat formatting.
Limitations
- Trained only on Wikidata-derived factual QA in 12 languages; it is not a general instruction-following model.
- Preference pairs are built from the model's own samples via embedding similarity, so they inherit LaBSE's similarity biases and can reward fluent-but-wrong answers.
- Improved cross-lingual consistency can make incorrect factual associations more uniform across languages as well as correct ones.
- Evaluation is multiple-choice log-likelihood scoring; it does not measure free-form generation quality.
Citation
The CM-Align method this baseline implements:
@inproceedings{zhang2025cmalign,
title = {CM-Align: Consistency-based Multilingual Alignment for Large Language Models},
author = {Zhang, Xue and others},
booktitle = {Findings of EMNLP},
year = {2025}
}
The RankC consistency metric used above:
@inproceedings{qi2023crosslingual,
title = {Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models},
author = {Qi, Jirui and Fern{\'a}ndez, Raquel and Bisazza, Arianna},
booktitle = {EMNLP},
year = {2023}
}
- Downloads last month
- 28
Model tree for jvonrad/Qwen-2.5-CM-Align
Base model
Qwen/Qwen2.5-7B