Link model card to paper, project page, and code repository

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +24 -9
README.md CHANGED
@@ -1,21 +1,23 @@
1
  ---
2
- license: apache-2.0
3
  base_model: Qwen/Qwen2.5-VL-7B-Instruct
 
 
4
  library_name: transformers
 
5
  pipeline_tag: image-text-to-text
6
  tags:
7
- - vision-language
8
- - safety
9
- - guardrail
10
- - policy-conditioned
11
- - qwen2.5-vl
12
- - policyshiftguard
13
- datasets:
14
- - PolicyShiftBench/PolicyShiftBench
15
  ---
16
 
17
  # PolicyShiftGuard-7B-RP-SFT
18
 
 
 
19
  This repository releases the **Stage-1 Randomized Policy SFT (RP-SFT)** checkpoint for the 7B PolicyShiftGuard model.
20
 
21
  RP-SFT is the first training stage in PolicyShiftGuard. It trains a Qwen2.5-VL guardrail model to read policy bundles under randomized policy identifiers and randomized policy ordering. This checkpoint is provided for reproducibility and ablation use. The final public model after the second-stage adaptation is available at [`PolicyShiftGuard/PolicyShiftGuard-7B`](https://huggingface.co/PolicyShiftGuard/PolicyShiftGuard-7B).
@@ -39,3 +41,16 @@ The model is trained with PolicyShiftBench supervision:
39
  - This is an intermediate checkpoint, not the final model reported as the main PolicyShiftGuard model.
40
  - This checkpoint corresponds to the randomized-policy no-think Stage-1 SFT setting.
41
  - Training-state files such as optimizer states are intentionally not included.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
 
2
  base_model: Qwen/Qwen2.5-VL-7B-Instruct
3
+ datasets:
4
+ - PolicyShiftBench/PolicyShiftBench
5
  library_name: transformers
6
+ license: apache-2.0
7
  pipeline_tag: image-text-to-text
8
  tags:
9
+ - vision-language
10
+ - safety
11
+ - guardrail
12
+ - policy-conditioned
13
+ - qwen2.5-vl
14
+ - policyshiftguard
 
 
15
  ---
16
 
17
  # PolicyShiftGuard-7B-RP-SFT
18
 
19
+ [\ud83d\udcc3 Paper](https://arxiv.org/abs/2607.05910) | [\ud83c\udf10 Project Page](https://policyshiftguard.github.io/) | [\ud83d\udcbb GitHub](https://github.com/ssmisya/PolicyShiftGuard)
20
+
21
  This repository releases the **Stage-1 Randomized Policy SFT (RP-SFT)** checkpoint for the 7B PolicyShiftGuard model.
22
 
23
  RP-SFT is the first training stage in PolicyShiftGuard. It trains a Qwen2.5-VL guardrail model to read policy bundles under randomized policy identifiers and randomized policy ordering. This checkpoint is provided for reproducibility and ablation use. The final public model after the second-stage adaptation is available at [`PolicyShiftGuard/PolicyShiftGuard-7B`](https://huggingface.co/PolicyShiftGuard/PolicyShiftGuard-7B).
 
41
  - This is an intermediate checkpoint, not the final model reported as the main PolicyShiftGuard model.
42
  - This checkpoint corresponds to the randomized-policy no-think Stage-1 SFT setting.
43
  - Training-state files such as optimizer states are intentionally not included.
44
+
45
+ ## Citation
46
+
47
+ If you find this work helpful, please cite the paper:
48
+
49
+ ```bibtex
50
+ @article{song2026policyshiftguard,
51
+ title = {PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails},
52
+ author = {Song, Mingyang and Xu, Luxin and Sun, Haoyu and Pan, Minzhou and Cheng, Yu and Li, Bo},
53
+ journal = {arXiv preprint arXiv:2607.05910},
54
+ year = {2026}
55
+ }
56
+ ```