--- license: apache-2.0 library_name: transformers tags: - dllm - diffusion - llm - text_generation --- # LLaDA2.2-mini **LLaDA2.2-mini** is the lightweight variant of the agentic diffusion language model in the LLaDA2 series. Built upon the LLaDA2.0-mini architecture, it inherits the core innovations of the LLaDA2.2 series — **Levenshtein Editing** (introducing `DELETE` and `INSERT` control tokens) — enabling long-context tool calling, multi-turn interaction, and robust error correction, while maintaining a smaller parameter footprint and lower inference cost. For more details, please refer to our [technical report](https://github.com/inclusionAI/LLaDA2.X/blob/main/LLaDA2_2_tech_report.pdf). --- ## 📊 Benchmarks The following table compares **LLaDA2.0-mini**, **LLaDA2.1-mini**, and **LLaDA2.2-mini** across General and Agentic capabilities.
Category Benchmark LLaDA2.0-mini LLaDA2.1-mini LLaDA2.2-mini
General
Function CallingBFCL v425.0528.4447.68
BFCL v370.7272.0669.02
MathAIME 202637.7140.3735.05
OlympiadBench67.7064.3061.11
CodingLiveCodeBench v627.7028.8028.14
MultiPL-E67.4664.1665.26
Instruction FollowingIFBench32.3331.6024.93
Multi-IF60.5858.4357.03
ReasoningKOR-Bench49.9246.6443.60
KnowledgeGPQA-Diamond47.7648.3644.41
Long ContextLongBench v215.5112.1334.99
General Average45.6845.0346.47
Agentic
Agentτ²-Bench--57.50
Claw-Eval--57.16
PinchBench--62.33
Agentic Average--59.00
--- ## 🚀 Key Features + **Efficient 128K Diffusion Infrastructure**: LLaDA2.2-mini extends the context window to **128K** and introduces the **Block Routing** mechanism, which restricts MoE expert activation at the diffusion block level, enabling efficient long-context agentic tasks. + **Levenshtein Editing**: Introduces **DELETE** and **INSERT** control tokens, enabling diffusion decoding to edit sequence structure, remove redundant content, and create insertion points during parallel generation. + **Agentic Reinforcement Learning**: Proposes **Levenshtein Editing ELBO-based Block-level Policy Optimization (L-EBPO)**, leveraging agentic environment rewards to train Levenshtein editing and error correction capabilities in multi-turn tool-use scenarios. + **Lightweight & Efficient**: With a total of **16B** parameters and only **1.4B** activated during inference, it significantly reduces computational cost while maintaining strong capabilities. --- ## 📦 Model Variants | Model ID | Description | Hugging Face Link | | --- | --- | --- | | `inclusionAI/LLaDA2.2-flash` | Agentic MoE Diffusion Language Model (100B) with Levenshtein editing capabilities. | [🤗 Model Card](https://huggingface.co/inclusionAI/LLaDA2.2-flash) | | `inclusionAI/LLaDA2.2-mini` | Lightweight Agentic MoE Diffusion Language Model (16B) with Levenshtein editing capabilities. | [🤗 Model Card](https://huggingface.co/inclusionAI/LLaDA2.2-mini) | --- ## 🔍 Model Overview Key specifications of **LLaDA2.2-mini**: + **Type**: Mixture-of-Experts (MoE) Diffusion Language Model with Levenshtein Editing + **Context Length**: 128K tokens + **Levenshtein Editing Control Tokens**: `DELETE`, `INSERT` + **Total Parameters (excl. Embedding)**: 16B + **Layers**: 20 + **Attention Heads**: 16 + **KV Heads**: 4 + **Experts**: 256 (8 activated per token) + **Positional Encoding**: Rotary Position Embedding (RoPE) + **Vocabulary Size**: 157,184 --- ## 🤗 Hugging Face Transformers Usage Please ensure `transformers>=5.2.0` and related dependencies are installed. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_path = "inclusionAI/LLaDA2.2-mini" model = AutoModelForCausalLM.from_pretrained( model_path, trust_remote_code=True, device_map="auto", ) model = model.to(torch.bfloat16) model.eval() tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True) prompt = "Calculate 1+5-28*0.5-200=?" input_ids = tokenizer.apply_chat_template( [{"role": "user", "content": prompt}], add_generation_prompt=True, tokenize=True, return_tensors="pt", ).input_ids generated_tokens = model.generate( inputs=input_ids, eos_early_stop=True, gen_length=512, block_length=32, threshold=0.5, editing_threshold=0.0, temperature=0.0, ) generated_answer = tokenizer.decode( generated_tokens[0], skip_special_tokens=True, ) print(generated_answer) ``` ### Best Practices For optimal performance, we recommend the following configurations: 1. **Sampling Parameters**: Use `block_length=32`, `temperature=0.0`, `top_p=None`, `top_k=None` as stable defaults. 2. **Denoising Threshold**: Adjust `threshold`, `editing_threshold`, and `max_post_steps` based on the speed-quality trade-off for your use case. Lower thresholds can improve inference speed but may lead to repetitive or unstable outputs. 3. **Output Length**: For most queries, an output length of 32768 tokens is recommended. 4. **Long-Context Agentic Tasks**: For long-context tool calling and multi-turn agentic applications, we recommend using **SGLang** as the serving backend. Ensure the server configuration supports a 128K context window and the model's MoE diffusion inference requirements. --- ## 🤖 ModelScope If you are in mainland China, we strongly recommend accessing our models via 🤖 [ModelScope](https://modelscope.cn/models/inclusionAI/LLaDA2.2-mini). --- ## 🌐 License This project is licensed under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). --- ## 🤝 Contact & Collaboration For any questions, collaboration opportunities, or feedback, please reach out to us via [Hugging Face](https://huggingface.co/inclusionAI/LLaDA2.2-mini) or submit an issue on our [GitHub repository](https://github.com/inclusionAI). Join us in advancing open, efficient, and intelligent diffusion language models for agentic applications! ---