Improve language tag

#1
by lbourdois - opened
Files changed (1) hide show
  1. README.md +188 -174
README.md CHANGED
@@ -1,175 +1,189 @@
1
- ---
2
- base_model:
3
- - Qwen/Qwen2.5-1.5B-Instruct
4
- - Qwen/Qwen2.5-1.5B-Instruct
5
- - Qwen/Qwen2.5-1.5B-Instruct
6
- - Qwen/Qwen2.5-1.5B-Instruct
7
- - Qwen/Qwen2.5-1.5B-Instruct
8
- - Qwen/Qwen2.5-1.5B-Instruct
9
- - Qwen/Qwen2.5-1.5B-Instruct
10
- - Qwen/Qwen2.5-1.5B-Instruct
11
- - Qwen/Qwen2.5-1.5B-Instruct
12
- - Qwen/Qwen2.5-1.5B-Instruct
13
- - Qwen/Qwen2.5-1.5B-Instruct
14
- - Qwen/Qwen2.5-1.5B-Instruct
15
- - Qwen/Qwen2.5-1.5B-Instruct
16
- - Qwen/Qwen2.5-1.5B-Instruct
17
- - Qwen/Qwen2.5-1.5B-Instruct
18
- - Qwen/Qwen2.5-1.5B-Instruct
19
- - Qwen/Qwen2.5-1.5B-Instruct
20
- - Qwen/Qwen2.5-1.5B-Instruct
21
- - Qwen/Qwen2.5-1.5B-Instruct
22
- - Qwen/Qwen2.5-1.5B-Instruct
23
- - Qwen/Qwen2.5-1.5B-Instruct
24
- - Qwen/Qwen2.5-1.5B-Instruct
25
- - Qwen/Qwen2.5-1.5B-Instruct
26
- - Qwen/Qwen2.5-1.5B-Instruct
27
- - Qwen/Qwen2.5-1.5B-Instruct
28
- - Qwen/Qwen2.5-1.5B-Instruct
29
- tags:
30
- - merge
31
- - mergekit
32
- - lazymergekit
33
- - Qwen/Qwen2.5-1.5B-Instruct
34
- ---
35
-
36
- # Qwen2.5-2B-Instruct
37
-
38
- Qwen2.5-2B-Instruct is a merge of the following models using [LazyMergekit](https://colab.research.google.com/drive/1obulZ1ROXHjYLn6PPZJwRR6GzgQogxxb?usp=sharing):
39
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
40
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
41
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
42
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
43
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
44
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
45
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
46
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
47
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
48
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
49
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
50
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
51
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
52
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
53
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
54
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
55
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
56
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
57
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
58
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
59
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
60
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
61
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
62
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
63
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
64
- * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
65
-
66
- ## 🧩 Configuration
67
-
68
- ```yaml
69
- dtype: bfloat16
70
- merge_method: passthrough
71
- slices:
72
- - sources:
73
- - layer_range: [0, 2]
74
- model: Qwen/Qwen2.5-1.5B-Instruct
75
- - sources:
76
- - layer_range: [1, 3]
77
- model: Qwen/Qwen2.5-1.5B-Instruct
78
- - sources:
79
- - layer_range: [2, 4]
80
- model: Qwen/Qwen2.5-1.5B-Instruct
81
- - sources:
82
- - layer_range: [3, 5]
83
- model: Qwen/Qwen2.5-1.5B-Instruct
84
- - sources:
85
- - layer_range: [4, 6]
86
- model: Qwen/Qwen2.5-1.5B-Instruct
87
- - sources:
88
- - layer_range: [5, 7]
89
- model: Qwen/Qwen2.5-1.5B-Instruct
90
- - sources:
91
- - layer_range: [6, 8]
92
- model: Qwen/Qwen2.5-1.5B-Instruct
93
- - sources:
94
- - layer_range: [7, 9]
95
- model: Qwen/Qwen2.5-1.5B-Instruct
96
- - sources:
97
- - layer_range: [8, 10]
98
- model: Qwen/Qwen2.5-1.5B-Instruct
99
- - sources:
100
- - layer_range: [9, 11]
101
- model: Qwen/Qwen2.5-1.5B-Instruct
102
- - sources:
103
- - layer_range: [10, 12]
104
- model: Qwen/Qwen2.5-1.5B-Instruct
105
- - sources:
106
- - layer_range: [11, 13]
107
- model: Qwen/Qwen2.5-1.5B-Instruct
108
- - sources:
109
- - layer_range: [12, 14]
110
- model: Qwen/Qwen2.5-1.5B-Instruct
111
- - sources:
112
- - layer_range: [13, 15]
113
- model: Qwen/Qwen2.5-1.5B-Instruct
114
- - sources:
115
- - layer_range: [14, 16]
116
- model: Qwen/Qwen2.5-1.5B-Instruct
117
- - sources:
118
- - layer_range: [16, 18]
119
- model: Qwen/Qwen2.5-1.5B-Instruct
120
- - sources:
121
- - layer_range: [17, 19]
122
- model: Qwen/Qwen2.5-1.5B-Instruct
123
- - sources:
124
- - layer_range: [18, 20]
125
- model: Qwen/Qwen2.5-1.5B-Instruct
126
- - sources:
127
- - layer_range: [19, 21]
128
- model: Qwen/Qwen2.5-1.5B-Instruct
129
- - sources:
130
- - layer_range: [20, 22]
131
- model: Qwen/Qwen2.5-1.5B-Instruct
132
- - sources:
133
- - layer_range: [21, 23]
134
- model: Qwen/Qwen2.5-1.5B-Instruct
135
- - sources:
136
- - layer_range: [22, 24]
137
- model: Qwen/Qwen2.5-1.5B-Instruct
138
- - sources:
139
- - layer_range: [23, 25]
140
- model: Qwen/Qwen2.5-1.5B-Instruct
141
- - sources:
142
- - layer_range: [24, 26]
143
- model: Qwen/Qwen2.5-1.5B-Instruct
144
- - sources:
145
- - layer_range: [25, 27]
146
- model: Qwen/Qwen2.5-1.5B-Instruct
147
- - sources:
148
- - layer_range: [26, 28]
149
- model: Qwen/Qwen2.5-1.5B-Instruct
150
- ```
151
-
152
- ## 💻 Usage
153
-
154
- ```python
155
- !pip install -qU transformers accelerate
156
-
157
- from transformers import AutoTokenizer
158
- import transformers
159
- import torch
160
-
161
- model = "win10/Qwen2.5-2B-Instruct"
162
- messages = [{"role": "user", "content": "What is a large language model?"}]
163
-
164
- tokenizer = AutoTokenizer.from_pretrained(model)
165
- prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
166
- pipeline = transformers.pipeline(
167
- "text-generation",
168
- model=model,
169
- torch_dtype=torch.float16,
170
- device_map="auto",
171
- )
172
-
173
- outputs = pipeline(prompt, max_new_tokens=256, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
174
- print(outputs[0]["generated_text"])
 
 
 
 
 
 
 
 
 
 
 
 
 
 
175
  ```
 
1
+ ---
2
+ base_model:
3
+ - Qwen/Qwen2.5-1.5B-Instruct
4
+ - Qwen/Qwen2.5-1.5B-Instruct
5
+ - Qwen/Qwen2.5-1.5B-Instruct
6
+ - Qwen/Qwen2.5-1.5B-Instruct
7
+ - Qwen/Qwen2.5-1.5B-Instruct
8
+ - Qwen/Qwen2.5-1.5B-Instruct
9
+ - Qwen/Qwen2.5-1.5B-Instruct
10
+ - Qwen/Qwen2.5-1.5B-Instruct
11
+ - Qwen/Qwen2.5-1.5B-Instruct
12
+ - Qwen/Qwen2.5-1.5B-Instruct
13
+ - Qwen/Qwen2.5-1.5B-Instruct
14
+ - Qwen/Qwen2.5-1.5B-Instruct
15
+ - Qwen/Qwen2.5-1.5B-Instruct
16
+ - Qwen/Qwen2.5-1.5B-Instruct
17
+ - Qwen/Qwen2.5-1.5B-Instruct
18
+ - Qwen/Qwen2.5-1.5B-Instruct
19
+ - Qwen/Qwen2.5-1.5B-Instruct
20
+ - Qwen/Qwen2.5-1.5B-Instruct
21
+ - Qwen/Qwen2.5-1.5B-Instruct
22
+ - Qwen/Qwen2.5-1.5B-Instruct
23
+ - Qwen/Qwen2.5-1.5B-Instruct
24
+ - Qwen/Qwen2.5-1.5B-Instruct
25
+ - Qwen/Qwen2.5-1.5B-Instruct
26
+ - Qwen/Qwen2.5-1.5B-Instruct
27
+ - Qwen/Qwen2.5-1.5B-Instruct
28
+ - Qwen/Qwen2.5-1.5B-Instruct
29
+ tags:
30
+ - merge
31
+ - mergekit
32
+ - lazymergekit
33
+ - Qwen/Qwen2.5-1.5B-Instruct
34
+ language:
35
+ - zho
36
+ - eng
37
+ - fra
38
+ - spa
39
+ - por
40
+ - deu
41
+ - ita
42
+ - rus
43
+ - jpn
44
+ - kor
45
+ - vie
46
+ - tha
47
+ - ara
48
+ ---
49
+
50
+ # Qwen2.5-2B-Instruct
51
+
52
+ Qwen2.5-2B-Instruct is a merge of the following models using [LazyMergekit](https://colab.research.google.com/drive/1obulZ1ROXHjYLn6PPZJwRR6GzgQogxxb?usp=sharing):
53
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
54
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
55
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
56
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
57
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
58
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
59
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
60
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
61
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
62
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
63
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
64
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
65
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
66
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
67
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
68
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
69
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
70
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
71
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
72
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
73
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
74
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
75
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
76
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
77
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
78
+ * [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
79
+
80
+ ## 🧩 Configuration
81
+
82
+ ```yaml
83
+ dtype: bfloat16
84
+ merge_method: passthrough
85
+ slices:
86
+ - sources:
87
+ - layer_range: [0, 2]
88
+ model: Qwen/Qwen2.5-1.5B-Instruct
89
+ - sources:
90
+ - layer_range: [1, 3]
91
+ model: Qwen/Qwen2.5-1.5B-Instruct
92
+ - sources:
93
+ - layer_range: [2, 4]
94
+ model: Qwen/Qwen2.5-1.5B-Instruct
95
+ - sources:
96
+ - layer_range: [3, 5]
97
+ model: Qwen/Qwen2.5-1.5B-Instruct
98
+ - sources:
99
+ - layer_range: [4, 6]
100
+ model: Qwen/Qwen2.5-1.5B-Instruct
101
+ - sources:
102
+ - layer_range: [5, 7]
103
+ model: Qwen/Qwen2.5-1.5B-Instruct
104
+ - sources:
105
+ - layer_range: [6, 8]
106
+ model: Qwen/Qwen2.5-1.5B-Instruct
107
+ - sources:
108
+ - layer_range: [7, 9]
109
+ model: Qwen/Qwen2.5-1.5B-Instruct
110
+ - sources:
111
+ - layer_range: [8, 10]
112
+ model: Qwen/Qwen2.5-1.5B-Instruct
113
+ - sources:
114
+ - layer_range: [9, 11]
115
+ model: Qwen/Qwen2.5-1.5B-Instruct
116
+ - sources:
117
+ - layer_range: [10, 12]
118
+ model: Qwen/Qwen2.5-1.5B-Instruct
119
+ - sources:
120
+ - layer_range: [11, 13]
121
+ model: Qwen/Qwen2.5-1.5B-Instruct
122
+ - sources:
123
+ - layer_range: [12, 14]
124
+ model: Qwen/Qwen2.5-1.5B-Instruct
125
+ - sources:
126
+ - layer_range: [13, 15]
127
+ model: Qwen/Qwen2.5-1.5B-Instruct
128
+ - sources:
129
+ - layer_range: [14, 16]
130
+ model: Qwen/Qwen2.5-1.5B-Instruct
131
+ - sources:
132
+ - layer_range: [16, 18]
133
+ model: Qwen/Qwen2.5-1.5B-Instruct
134
+ - sources:
135
+ - layer_range: [17, 19]
136
+ model: Qwen/Qwen2.5-1.5B-Instruct
137
+ - sources:
138
+ - layer_range: [18, 20]
139
+ model: Qwen/Qwen2.5-1.5B-Instruct
140
+ - sources:
141
+ - layer_range: [19, 21]
142
+ model: Qwen/Qwen2.5-1.5B-Instruct
143
+ - sources:
144
+ - layer_range: [20, 22]
145
+ model: Qwen/Qwen2.5-1.5B-Instruct
146
+ - sources:
147
+ - layer_range: [21, 23]
148
+ model: Qwen/Qwen2.5-1.5B-Instruct
149
+ - sources:
150
+ - layer_range: [22, 24]
151
+ model: Qwen/Qwen2.5-1.5B-Instruct
152
+ - sources:
153
+ - layer_range: [23, 25]
154
+ model: Qwen/Qwen2.5-1.5B-Instruct
155
+ - sources:
156
+ - layer_range: [24, 26]
157
+ model: Qwen/Qwen2.5-1.5B-Instruct
158
+ - sources:
159
+ - layer_range: [25, 27]
160
+ model: Qwen/Qwen2.5-1.5B-Instruct
161
+ - sources:
162
+ - layer_range: [26, 28]
163
+ model: Qwen/Qwen2.5-1.5B-Instruct
164
+ ```
165
+
166
+ ## 💻 Usage
167
+
168
+ ```python
169
+ !pip install -qU transformers accelerate
170
+
171
+ from transformers import AutoTokenizer
172
+ import transformers
173
+ import torch
174
+
175
+ model = "win10/Qwen2.5-2B-Instruct"
176
+ messages = [{"role": "user", "content": "What is a large language model?"}]
177
+
178
+ tokenizer = AutoTokenizer.from_pretrained(model)
179
+ prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
180
+ pipeline = transformers.pipeline(
181
+ "text-generation",
182
+ model=model,
183
+ torch_dtype=torch.float16,
184
+ device_map="auto",
185
+ )
186
+
187
+ outputs = pipeline(prompt, max_new_tokens=256, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
188
+ print(outputs[0]["generated_text"])
189
  ```