File size: 5,035 Bytes
c4bd8e0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
---
base_model:
- meta-llama/Llama-2-7b-hf
library_name: transformers
license: llama2
tags:
- pruning
- sparsity
- 2:4-sparsity
- patch
- maskllm
- mask
pipeline_tag: text-generation
---

<div align="center">
<img src="./PATCH-Logo.png" alt="PATCH" width="360">
</div>

# llama_2_7b-PATCH-25Sparse

> **This checkpoint** (LLaMA-2 7B, PATCH-Tile, 25% sparsity): **51.58%** average zero-shot accuracy, **5.86** WikiText2 perplexity.

[![Paper](https://img.shields.io/badge/arXiv-2509.23410-b31b1b.svg)](https://arxiv.org/abs/2509.23410) [![GitHub](https://img.shields.io/badge/GitHub-Paramathic%2Fpatch-black.svg?logo=github)](https://github.com/Paramathic/patch)

This repository hosts a **mask only** release for the paper
**[PATCH: Learnable Tile-level Hybrid Sparsity for LLMs](https://arxiv.org/abs/2509.23410)**.
PATCH (Pruning with a Learnable Tile-level Configuration for Hybrid Sparsity) learns a structured mask on **frozen** pretrained weights, assigning each tile as dense (0% sparsity) or 2:4 sparse (50% sparsity) to hit a flexible global sparsity target while staying hardware-friendly.

Because PATCH/MaskLLM keep the base weights **frozen**, we distribute *only the
binary keep/prune mask* (bit-packed in `mask.npz`) - **no weight values**. You
recover the sparse model by downloading the original base model and applying the
mask.

- Base model: [`meta-llama/Llama-2-7b-hf`](https://huggingface.co/meta-llama/Llama-2-7b-hf)
- Method: **PATCH-Tile** &nbsp;|&nbsp; Target sparsity: **25%** &nbsp;|&nbsp; Pattern: **Dense / 2:4 tiles**
- Measured mask sparsity: **25.00%** over 224 Linear layers (6,476,005,376 weights).

<div align="center">
<img src="./PATCH-Pipeline.svg" alt="PATCH pipeline" width="760">
</div>

## Results (LLaMA-2 7B)

| Sparsity | Method | Pattern | Avg Acc (% ↑) | WikiText2 PPL (↓) |
|---|---|---|---|---|
| 0% | Dense | - | 54.61 | 5.12 |
| 50% | Magnitude | 2:4 | 43.44 | 54.39 |
| 50% | Wanda | 2:4 | 44.30 | 11.15 |
| 50% | SparseGPT | 2:4 | 45.09 | 10.12 |
| 50% | Thanos | 2:4 | 44.80 | 11.19 |
| 50% | ProxSparse | 2:4 | 45.92 | 9.18 |
| 50% | MaskLLM | 2:4 | 48.62 | 6.78 |
| 45% | PATCH-Tile | Dense/2:4 | 48.99 | 6.55 |
| 35% | PATCH-Tile | Dense/2:4 | 50.08 | 6.18 |
| 25% | **PATCH-Tile** ⭐ | Dense/2:4 | **51.58** | **5.86** |


Per-task zero-shot accuracy (%) for this checkpoint:

| MMLU | PIQA | ARC-E | ARC-C | WinoG. | OBQA | RACE | HellaS. | **Average** |
|---|---|---|---|---|---|---|---|---|
| 32.33 | 76.99 | 72.81 | 38.57 | 68.27 | 29.80 | 39.52 | 54.34 | **51.58** |

All numbers are from the PATCH paper ([arXiv:2509.23410](https://arxiv.org/abs/2509.23410)); accuracy
is the average over MMLU, PIQA, ARC-Easy, ARC-Challenge, Winogrande, OpenBookQA,
RACE and HellaSwag, evaluated with the LM-Evaluation-Harness. PPL is WikiText2.

## Training hyper-parameters

| Hyper-parameter | Value |
|---|---|
| Fine-tuning dataset | SlimPajama (2B tokens) |
| Training steps | 2000 |
| Global batch size | 256 |
| Sequence length | 4096 |
| Mask tile size | 128 x 128 (hardware tiles: 128x128 / 128x64 / 64x128 / 64x64) |
| Logits init. | N(0, 0.014) |
| Tile-logit prior | SparseGPT (strength 3) |
| Regularization scope | Global (single target density) |
| Evaluation | LM-Eval-Harness (8 zero-shot tasks) + WikiText2 PPL @ seqlen 4096 |
| Hardware | 1 node x 4 GPUs, data parallel (HuggingFace Trainer) |
| Optimizer | Adam |
| Learning rate | 1e-4 |
| Gumbel scaling (kappa) | 100 -> 500 |
| Gumbel temp (tau) | 2 -> 0.05 |
| Sparsity reg. (lambda1) | 3 |
| Weight reg. (lambda2) | 0.1 |

## How to use

```python
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM
import torch
from load_patch_mask import apply_patch_mask  # shipped in this repo

npz = hf_hub_download(repo_id="mohammad-mozaffari/llama_2_7b-PATCH-25Sparse", filename="mask.npz")
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-2-7b-hf", torch_dtype=torch.bfloat16)
apply_patch_mask(model, npz)   # zeroes the pruned weights in place
```

Or from the command line:

```bash
python load_patch_mask.py --base_model meta-llama/Llama-2-7b-hf --mask_repo mohammad-mozaffari/llama_2_7b-PATCH-25Sparse
```

Speedup on real hardware requires a 2:4-aware / hybrid sparse kernel; see the
[GitHub repository](https://github.com/Paramathic/patch) and [STOICC](https://github.com/Paramathic/stoicc).

## License

The released mask is a derivative of the base model and is distributed under the
base model's license (**`llama2`**). You must comply with that license and
obtain access to the base model separately.

> Use governed by the Llama 2 Community License.

The mask-generation code is released under the MIT license (see the
[PATCH repository](https://github.com/Paramathic/patch)).

## Citation

```bibtex
@article{hourri2025patch,
    title  = {PATCH: Learnable Tile-level Hybrid Sparsity for LLMs},
    author = {Hourri, Younes and Mozaffari, Mohammad and Mehri Dehnavi, Maryam},
    year   = 2025,
    journal = {arXiv preprint arXiv:2509.23410}
}
```