PATCH

qwen2.5_0.5b-MaskLLM-50Sparse

This checkpoint (Qwen-2.5 0.5B, MaskLLM, 50% sparsity): 39.33% average zero-shot accuracy, 15.22 WikiText2 perplexity.

Paper GitHub

This repository hosts a mask only release for the paper PATCH: Learnable Tile-level Hybrid Sparsity for LLMs. This is a MaskLLM 2:4 baseline, trained with our re-implementation of MaskLLM (learnable semi-structured 2:4 sparsity via Gumbel-Softmax over frozen weights). It is released as a baseline for the PATCH paper.

Because PATCH/MaskLLM keep the base weights frozen, we distribute only the binary keep/prune mask (bit-packed in mask.npz) - no weight values. You recover the sparse model by downloading the original base model and applying the mask.

  • Base model: Qwen/Qwen2.5-0.5B
  • Method: MaskLLM  |  Target sparsity: 50%  |  Pattern: 2:4
  • Measured mask sparsity: 50.00% over 168 Linear layers (357,826,560 weights).
PATCH pipeline

Results (Qwen-2.5 0.5B)

Sparsity Method Pattern Avg Acc (% ↑) WikiText2 PPL (↓)
0% Dense - 46.00 12.08
50% Magnitude 2:4 30.16 6734.97
50% Wanda 2:4 32.97 72.48
50% SparseGPT 2:4 34.81 36.59
50% Thanos 2:4 34.85 37.32
50% ProxSparse 2:4 32.05 111.05
50% MaskLLM 2:4 39.33 15.22
45% PATCH-Joint Dense/2:4 40.29 14.57
35% PATCH-Joint Dense/2:4 41.15 13.84
25% PATCH-Joint Dense/2:4 42.39 13.47

Per-task zero-shot accuracy (%) for this checkpoint:

MMLU PIQA ARC-E ARC-C WinoG. OBQA RACE HellaS. Average
25.11 67.03 56.57 23.98 52.57 20.20 33.30 35.90 39.33

All numbers are from the PATCH paper (arXiv:2509.23410); accuracy is the average over MMLU, PIQA, ARC-Easy, ARC-Challenge, Winogrande, OpenBookQA, RACE and HellaSwag, evaluated with the LM-Evaluation-Harness. PPL is WikiText2.

Training hyper-parameters

Hyper-parameter Value
Fine-tuning dataset SlimPajama (2B tokens)
Training steps 2000
Global batch size 256
Sequence length 4096
Mask tile size 128 x 128 (hardware tiles: 128x128 / 128x64 / 64x128 / 64x64)
Logits init. N(0, 0.014)
Tile-logit prior SparseGPT (strength 3)
Regularization scope Global (single target density)
Evaluation LM-Eval-Harness (8 zero-shot tasks) + WikiText2 PPL @ seqlen 4096
Hardware 1 node x 4 GPUs, data parallel (HuggingFace Trainer)
Optimizer Adam
Learning rate 1e-3
Gumbel scaling (kappa) 100 -> 500
Gumbel temp (tau) 4 -> 0.05
2:4 prior strength 3
Weight reg. (lambda2) 1e-5
Sparsity pattern 2:4 (fixed 50%)

How to use

from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM
import torch
from load_patch_mask import apply_patch_mask  # shipped in this repo

npz = hf_hub_download(repo_id="mohammad-mozaffari/qwen2.5_0.5b-MaskLLM-50Sparse", filename="mask.npz")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B", torch_dtype=torch.bfloat16)
apply_patch_mask(model, npz)   # zeroes the pruned weights in place

Or from the command line:

python load_patch_mask.py --base_model Qwen/Qwen2.5-0.5B --mask_repo mohammad-mozaffari/qwen2.5_0.5b-MaskLLM-50Sparse

Speedup on real hardware requires a 2:4-aware / hybrid sparse kernel; see the GitHub repository and STOICC.

License

The released mask is a derivative of the base model and is distributed under the base model's license (apache-2.0). You must comply with that license and obtain access to the base model separately.

The mask-generation code is released under the MIT license (see the PATCH repository).

Citation

@article{hourri2025patch,
    title  = {PATCH: Learnable Tile-level Hybrid Sparsity for LLMs},
    author = {Hourri, Younes and Mozaffari, Mohammad and Mehri Dehnavi, Maryam},
    year   = 2025,
    journal = {arXiv preprint arXiv:2509.23410}
}

This checkpoint is a MaskLLM baseline; please also cite MaskLLM:

@inproceedings{fang2024maskllm,
    title     = {MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models},
    author    = {Fang, Gongfan and Yin, Hongxu and Muralidharan, Saurav and Heinrich, Greg and Pool, Jeff and Kautz, Jan and Molchanov, Pavlo and Wang, Xinchao},
    booktitle = {NeurIPS},
    year      = {2024}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mohammad-mozaffari/qwen2.5_0.5b-MaskLLM-50Sparse

Finetuned
(687)
this model

Collection including mohammad-mozaffari/qwen2.5_0.5b-MaskLLM-50Sparse

Paper for mohammad-mozaffari/qwen2.5_0.5b-MaskLLM-50Sparse