Add paper info and metadata

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +35 -11
README.md CHANGED
@@ -1,22 +1,46 @@
1
  ---
2
  library_name: transformers
3
- tags: []
 
 
 
 
 
 
4
  ---
5
 
6
- # Model Card for Model ID
7
 
8
- <!-- Provide a quick summary of what the model is/does. -->
9
- ## UserMirrorrer-Llama-DPO
10
 
11
- This is the fine-tuned checkpoint using our constructed dataset [UserMirrorer](https://huggingface.co/datasets/MirrorUser/UserMirrorer).
12
-
13
- Please refer to our paper: "Mirroring Users: Towards Building Preference-aligned User Simulator with Recommendation Feedback".
14
 
15
  ## Model Details
16
 
17
- The finetuned model is based on [Llama-3.2-3B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct).
 
 
 
 
 
 
 
 
 
 
 
 
18
 
19
- The fine-tuning process involves 2 stage:
20
- 1. Supervised Finetuning for 1 epoch
21
- 2. DPO Finetuning for 2 epochs
22
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  library_name: transformers
3
+ pipeline_tag: text-generation
4
+ base_model: meta-llama/Llama-3.2-3B-Instruct
5
+ datasets:
6
+ - Joinn/UserMirrorer
7
+ tags:
8
+ - recommender-system
9
+ - user-simulation
10
  ---
11
 
12
+ # UserMirrorrer-Llama-DPO
13
 
14
+ This is a preference-aligned user simulator for recommendation systems, fine-tuned from [Llama-3.2-3B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct) using the [UserMirrorer](https://huggingface.co/datasets/Joinn/UserMirrorer) framework.
 
15
 
16
+ The model was introduced in the paper [Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation](https://huggingface.co/papers/2508.18142).
 
 
17
 
18
  ## Model Details
19
 
20
+ UserMirrorer is designed to simulate user behavior and preferences in recommender systems by leveraging extensive user feedback. The framework generates decision-making processes as explanatory rationales to enhance alignment with human preferences.
21
+
22
+ The fine-tuning process involved two stages:
23
+ 1. **Supervised Fine-tuning (SFT):** 1 epoch.
24
+ 2. **Direct Preference Optimization (DPO):** 2 epochs.
25
+
26
+ ## Resources
27
+
28
+ - **Paper:** [Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation](https://huggingface.co/papers/2508.18142)
29
+ - **GitHub Repository:** [Joinn99/UserMirrorer](https://github.com/Joinn99/UserMirrorer)
30
+ - **Dataset:** [UserMirrorer](https://huggingface.co/datasets/Joinn/UserMirrorer)
31
+
32
+ ## Citation
33
 
34
+ If you find this work useful, please consider citing:
 
 
35
 
36
+ ```bibtex
37
+ @misc{wei2025mirroringusersbuildingpreferencealigned,
38
+ title={Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation},
39
+ author={Tianjun Wei and Huizhong Guo and Yingpeng Du and Zhu Sun and Huang Chen and Dongxia Wang and Jie Zhang},
40
+ year={2025},
41
+ eprint={2508.18142},
42
+ archivePrefix={arXiv},
43
+ primaryClass={cs.HC},
44
+ url={https://arxiv.org/abs/2508.18142},
45
+ }
46
+ ```