dongjiang1989 commited on
Commit
ea14618
·
0 Parent(s):

Initial release of Spark-X2.5-4B

Browse files

Co-authored-by: chao <linchao-2026@users.noreply.huggingface.co>
Co-authored-by: Dong Jiang <dongjiang1989@users.noreply.huggingface.co>
Co-authored-by: Lei Yang <lyang01@users.noreply.huggingface.co>
Co-authored-by: Marx Benjamin <Flyingmarx@users.noreply.huggingface.co>
Co-authored-by: TylerZhang <Tylerwtzhang3@users.noreply.huggingface.co>
Co-authored-by: Wenzhi Peng <pengwenzhi@users.noreply.huggingface.co>

Signed-off-by: chao <linchao-2026@users.noreply.huggingface.co>
Signed-off-by: Dong Jiang <dongjiang1989@users.noreply.huggingface.co>
Signed-off-by: Lei Yang <lyang01@users.noreply.huggingface.co>
Signed-off-by: Marx Benjamin <Flyingmarx@users.noreply.huggingface.co>
Signed-off-by: TylerZhang <Tylerwtzhang3@users.noreply.huggingface.co>
Signed-off-by: Wenzhi Peng <pengwenzhi@users.noreply.huggingface.co>

.gitattributes ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ images/spark3-hybrid-architecture-light.png filter=lfs diff=lfs merge=lfs -text
38
+ images/xhtoken-wechat.jpg filter=lfs diff=lfs merge=lfs -text
39
+ images/spark25-hybrid-architecture-light.png filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright 2026 XHToken
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
README.md ADDED
@@ -0,0 +1,349 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - zh
6
+ library_name: transformers
7
+ pipeline_tag: text-generation
8
+ tags:
9
+ - llm
10
+ - sparkx2_5
11
+ base_model:
12
+ - XHToken/Spark-X2.5-4B-Base
13
+ ---
14
+
15
+
16
+ # Spark-X2.5
17
+ <div align="center">
18
+
19
+ [![Slack](https://img.shields.io/badge/Slack-Join-4A154B?logo=slack&logoColor=white)](https://join.slack.com/t/tokenspark/shared_invite/zt-432qf8l2f-5~dLyXv8uETr0P0UuC07nw)
20
+ [![Discord](https://img.shields.io/badge/Discord-Join-5865F2?logo=discord&logoColor=white)](https://discord.gg/kTDE2Hg8aw)
21
+ [![YouTube](https://img.shields.io/badge/YouTube-Subscribe-FF0000?logo=youtube&logoColor=white)](https://www.youtube.com/@SparkLLM)
22
+ [![dev.to](https://img.shields.io/badge/dev.to-Follow-0A0A0A?logo=devdotto&logoColor=white)](https://dev.to/sparkllm)
23
+ [![Bluesky](https://img.shields.io/badge/Bluesky-Follow-0285FF?logo=bluesky&logoColor=white)](https://bsky.app/profile/sparkllm.bsky.social)
24
+ [![X](https://img.shields.io/badge/X-Follow-000000?logo=x&logoColor=white)](https://x.com/sparkllm)
25
+ [![Zhihu](https://img.shields.io/badge/Zhihu-Follow-0084FF?logo=zhihu&logoColor=white)](https://www.zhihu.com/people/zhiikz7qh7m)
26
+ [![WeChat](https://img.shields.io/badge/WeChat-Join-07C160?logo=wechat&logoColor=white)](images/xhtoken-wechat.jpg)
27
+
28
+ </div>
29
+
30
+ > [!Note]
31
+ > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
32
+
33
+ ## Introduction
34
+
35
+ We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages.
36
+
37
+ **Technical Highlights**:
38
+ - **Efficient Architecture and Native 1M-token Context**: The models use a hybrid attention architecture that combines one full-attention layer with three sliding-window attention layers. This design substantially reduces the computational overhead typically associated with long-context models while natively supporting a context window of up to 1M tokens.
39
+ - **Strong Coding and Agent Capabilities**: The models are deeply integrated with popular agent harnesses, including Codex, Claude Code, OpenClaw, and Hermes. They deliver state-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction-following tasks.
40
+ - **Broad Hardware and Software Compatibility**: The models support a wide range of hardware platforms, including NVIDIA, Huawei, Hygon, HOUMO.AI, etc. It is compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, MLX, and can be deployed quickly through platforms including Ollama and LM Studio. The models can also be customized using popular fine-tuning frameworks such as LLaMA-Factory. Across multiple hardware platforms, they deliver superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
41
+ - **Advanced Training Algorithms**: The models were trained on Huawei Ascend clusters. Large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance its reasoning, coding, agentic, and instruction-following capabilities.
42
+
43
+ <p align="center">
44
+ <img src="./images/model-benchmark-comparison.svg" alt="Spark-X2.5 benchmark comparison" width="1010">
45
+ </p>
46
+
47
+
48
+ ## Model Overview
49
+
50
+ For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture. This approach leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size—thereby improving its practicality and effectiveness across real-world deployment scenarios.
51
+
52
+ <p align="center">
53
+ <img src="./images/spark25-hybrid-architecture-light.png" alt="Spark-X2.5 hybrid architecture" width="610">
54
+ </p>
55
+
56
+
57
+
58
+ ## Training Methods
59
+
60
+ Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies are conducted to determine an effective balance among mathematics, logic, code, and other high-value domains. This enables the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.
61
+
62
+ Post-training begins with supervised fine-tuning on a carefully curated corpus. This stage establishes robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning. We subsequently apply large-scale reinforcement learning across several capability domains, including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. This process yields a set of domain-specialized teacher policies, whose complementary strengths are consolidated into a single deployable model through MOPD.
63
+
64
+ <p align="center">
65
+ <img src="./images/post_training_pipeline.svg" alt="Spark-X2.5 hybrid architecture" width="610">
66
+ </p>
67
+
68
+ ## Benchmarks
69
+
70
+ We evaluate our models and compare them with leading on-device models of similar size across a broad range of tasks, including agent, code, math, general and knowledge.
71
+
72
+ <div style="overflow-x: auto; width: 100%;">
73
+ <table style="width: 100%; min-width: 1080px; border-collapse: collapse; text-align: center;">
74
+ <colgroup>
75
+ <col style="width: 190px;">
76
+ <col span="8" style="width: 110px;">
77
+ </colgroup>
78
+ <thead>
79
+ <tr>
80
+ <th align="center">Benchmark</th>
81
+ <th align="center"><span style="white-space: nowrap;">Spark‑X2.5‑4B</span></th>
82
+ <th align="center"><span style="white-space: nowrap;">Spark‑X2.5‑1.7B</span></th>
83
+ <th align="center"><span style="white-space: nowrap;">Qwen3.5‑9B</span></th>
84
+ <th align="center"><span style="white-space: nowrap;">Qwen3.5‑4B</span></th>
85
+ <th align="center"><span style="white-space: nowrap;">Qwen3.5‑2B</span></th>
86
+ <th align="center"><span style="white-space: nowrap;">Gemma4‑12B</span></th>
87
+ <th align="center"><span style="white-space: nowrap;">Gemma4‑E4B</span></th>
88
+ <th align="center"><span style="white-space: nowrap;">Gemma4‑E2B</span></th>
89
+ </tr>
90
+ </thead>
91
+ <tbody>
92
+ <tr><th colspan="9" align="left">Agent</th></tr>
93
+ <tr><td align="center">BFCL‑V4</td><td align="center">65.1</td><td align="center">46.9</td><td align="center"><strong>66.1*</strong></td><td align="center">50.3*</td><td align="center">43.6*</td><td align="center">37.4</td><td align="center">36.9</td><td align="center">30.2</td></tr>
94
+ <tr><td align="center">τ²‑bench</td><td align="center">75.1</td><td align="center">65.3</td><td align="center">79.1*</td><td align="center"><strong>79.9*</strong></td><td align="center">48.8*</td><td align="center">69.0*</td><td align="center">42.2*</td><td align="center">24.5*</td></tr>
95
+ <tr><td align="center">τ³‑bench</td><td align="center"><strong>30.4</strong></td><td align="center">20.1</td><td align="center">9.3</td><td align="center">6.7</td><td align="center">4.1</td><td align="center">13.3</td><td align="center">10.1</td><td align="center">8.8</td></tr>
96
+ <tr><td align="center">MCP‑Atlas</td><td align="center"><strong>54.6</strong></td><td align="center">23.4</td><td align="center">47.4*</td><td align="center">40.8*</td><td align="center">14.8</td><td align="center">30.5*</td><td align="center">15.0*</td><td align="center">12.6</td></tr>
97
+ <tr><td align="center">MCP‑Mark</td><td align="center"><strong>14.2</strong></td><td align="center">2.3</td><td align="center">13.4</td><td align="center">12.5</td><td align="center">–</td><td align="center">–</td><td align="center">–</td><td align="center">–</td></tr>
98
+ <tr><td align="center"><span style="white-space: nowrap;">Workspace Bench</span></td><td align="center"><strong>31.2</strong></td><td align="center">18.9</td><td align="center">25.5</td><td align="center">21.3</td><td align="center">7.7</td><td align="center">–</td><td align="center">–</td><td align="center">–</td></tr>
99
+ <tr><td align="center">VitaBench2.0</td><td align="center"><strong>25.2</strong></td><td align="center">8.3</td><td align="center">15.6</td><td align="center">18.2</td><td align="center">5.2</td><td align="center">12.4</td><td align="center">4.8</td><td align="center">4.4</td></tr>
100
+ <tr><td align="center">BrowseComp</td><td align="center"><strong>40.9</strong></td><td align="center">29.7</td><td align="center">8.3</td><td align="center">14.3</td><td align="center">3.1</td><td align="center">10.0</td><td align="center">8.3</td><td align="center">3.7</td></tr>
101
+ <tr><th colspan="9" align="left">Code</th></tr>
102
+ <tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Pro</span></td><td align="center"><strong>44.4</strong></td><td align="center">10.4</td><td align="center">33.8*</td><td align="center">29.4*</td><td align="center">1.9</td><td align="center">21.9*</td><td align="center">4.0*</td><td align="center">–</td></tr>
103
+ <tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Verified</span></td><td align="center">41.6</td><td align="center">28.3</td><td align="center"><strong>53.1*</strong></td><td align="center">38.8*</td><td align="center">6.8</td><td align="center">44.2*</td><td align="center">14.0*</td><td align="center">–</td></tr>
104
+ <tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Multilingual</span></td><td align="center"><strong>53.3</strong></td><td align="center">23.3</td><td align="center">43.3</td><td align="center">27.7</td><td align="center">5.0</td><td align="center">32.5*</td><td align="center">–</td><td align="center">–</td></tr>
105
+ <tr><td align="center">SciCode</td><td align="center">34.7</td><td align="center">18.2</td><td align="center">32.7*</td><td align="center">24.0</td><td align="center">6.0</td><td align="center"><strong>39.8</strong></td><td align="center">27.5</td><td align="center">20.5</td></tr>
106
+ <tr><th colspan="9" align="left">Math</th></tr>
107
+ <tr><td align="center"><span style="white-space: nowrap;">Gaokao 2026</span></td><td align="center">133.4</td><td align="center">114.8</td><td align="center"><strong>135.5</strong></td><td align="center">130.3</td><td align="center">94.0</td><td align="center">130.6</td><td align="center">102.4</td><td align="center">81.8</td></tr>
108
+ <tr><td align="center"><span style="white-space: nowrap;">AIME 2026</span></td><td align="center"><strong>90.7</strong></td><td align="center">69.4</td><td align="center">88.2</td><td align="center">83.0</td><td align="center">30.8</td><td align="center">82.1*</td><td align="center">42.5*</td><td align="center">37.5*</td></tr>
109
+ <tr><td align="center"><span style="white-space: nowrap;">HMMT Feb 2026</span></td><td align="center"><strong>81.2</strong></td><td align="center">48.4</td><td align="center">70.8</td><td align="center">69.7</td><td align="center">21.5</td><td align="center">65.6</td><td align="center">34.2</td><td align="center">20.5</td></tr>
110
+ <tr><td align="center"><span style="white-space: nowrap;">IMO‑AnswerBench</span></td><td align="center"><strong>74.2</strong></td><td align="center">45.4</td><td align="center">69.8</td><td align="center">68.5</td><td align="center">–</td><td align="center">57.2</td><td align="center">26.9</td><td align="center">22.6</td></tr>
111
+ <tr><th colspan="9" align="left">General &amp; Knowledge</th></tr>
112
+ <tr><td align="center">IFEval</td><td align="center">93.0</td><td align="center">89.5</td><td align="center">91.5*</td><td align="center">89.8*</td><td align="center">78.6*</td><td align="center"><strong>94.8</strong></td><td align="center">45.3</td><td align="center">34.8</td></tr>
113
+ <tr><td align="center">IFBench</td><td align="center"><strong>75.0</strong></td><td align="center">66.3</td><td align="center">64.5</td><td align="center">59.2</td><td align="center">41.3*</td><td align="center">73.5*</td><td align="center">44.0*</td><td align="center">22.7</td></tr>
114
+ <tr><td align="center">AA‑LCR</td><td align="center">56.3</td><td align="center">24.3</td><td align="center"><strong>63.0*</strong></td><td align="center">57.0*</td><td align="center">25.6*</td><td align="center">55.3*</td><td align="center">34.7</td><td align="center">18.3</td></tr>
115
+ <tr><td align="center">HLE</td><td align="center">12.3</td><td align="center">6.3</td><td align="center"><strong>14.3</strong></td><td align="center">8.6</td><td align="center">2.1</td><td align="center">13.1</td><td align="center">3.9</td><td align="center">2.5</td></tr>
116
+ <tr><td align="center">GPQA</td><td align="center">67.4</td><td align="center">43.8</td><td align="center"><strong>77.2</strong></td><td align="center">67.2</td><td align="center">44.6</td><td align="center">72.8</td><td align="center">54.5</td><td align="center">43.8</td></tr>
117
+ </tbody>
118
+ </table>
119
+ </div>
120
+
121
+ - \* denotes reported results from publicly‑released model cards / papers and - denotes scores not yet available.
122
+ - All evaluations are conducted in thinking mode. The recommended sampling parameters for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1.
123
+ - Gaokao 2026 consists of the five 2026 Chinese GAOKAO examinations (National I,National II, Beijing, Shanghai, Tianjin), each graded out of 150 points.
124
+
125
+ ## Quickstart
126
+
127
+ ### SGLang
128
+
129
+ #### Install SGLang
130
+
131
+ Use the pre-built image that tracks the Spark-X2.5 runtime:
132
+
133
+ ```bash
134
+ docker pull lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1
135
+ ```
136
+
137
+ #### Run Inference
138
+
139
+ The following command can be used to start an OpenAI-compatible API server on a single GPU with maximum context length 1,048,576 tokens.
140
+
141
+ #### Server
142
+
143
+ ```bash
144
+ docker run -it \
145
+ --gpus '"device=0"' \
146
+ --ipc=host \
147
+ -p 30000:30000 \
148
+ -v "$MODEL_PATH":/root/Spark-X2.5-4B \
149
+ lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 \
150
+ python -m sglang.launch_server \
151
+ --model-path /root/Spark-X2.5-4B \
152
+ --served-model-name spark2.5 \
153
+ --tool-call-parser spark25 \
154
+ --reasoning-parser qwen3 \
155
+ --tp-size 1 \
156
+ --mem-fraction-static 0.8 \
157
+ --context-length 1048576 \
158
+ --chat-template /root/Spark-X2.5-4B/chat_template.jinja \
159
+ --host 0.0.0.0 \
160
+ --port 30000
161
+ ```
162
+
163
+ #### Client
164
+
165
+ Thinking is enabled by default by both the chat template and the qwen3 reasoning parser. To disable thinking for a specific request, set "chat_template_kwargs": {"enable_thinking": false}.
166
+
167
+ ```bash
168
+ curl -s http://localhost:30000/v1/chat/completions \
169
+ -H "Content-Type: application/json" \
170
+ -d '{
171
+ "model": "spark2.5",
172
+ "messages": [
173
+ {
174
+ "role": "user",
175
+ "content": "安徽的省会在哪里?"
176
+ }
177
+ ],
178
+ "max_tokens": 131072,
179
+ "temperature": 1,
180
+ "top_k": -1,
181
+ "top_p": 0.95,
182
+ "repetition_penalty": 1,
183
+ "presence_penalty": 0,
184
+ "frequency_penalty": 0
185
+ }'
186
+ ```
187
+
188
+ ### vLLM
189
+
190
+ #### Install vLLM
191
+
192
+ ```bash
193
+ pip install uv
194
+ uv venv ~/spark2_5
195
+ source ~/spark2_5/bin/activate
196
+ git clone https://github.com/XHToken/Spark-plugin.git
197
+ cd ./Spark-plugin
198
+ uv pip install .
199
+ ```
200
+
201
+ #### Server
202
+
203
+ ```bash
204
+ vllm serve "./Spark-X2.5-4B" \
205
+ --port "30000" \
206
+ --trust-remote-code \
207
+ --served-model-name spark25 \
208
+ --tensor-parallel-size 1 \
209
+ --gpu-memory-utilization 0.7 \
210
+ --enable-prefix-caching \
211
+ --chat-template Spark-X2.5-4B/chat_template.jinja
212
+ ```
213
+
214
+ #### Client
215
+
216
+ ```bash
217
+ curl -s http://127.0.0.1:30000/v1/chat/completions \
218
+ -H "Content-Type: application/json" \
219
+ -d '{
220
+ "model": "spark25",
221
+ "messages": [{"role": "user", "content": "安徽的省会在哪里?"}],
222
+ "temperature": 1.0,
223
+ "top_k": -1,
224
+ "top_p": 0.95
225
+ }'
226
+ ```
227
+
228
+ ### MLX
229
+ Spark-MLX-LLM runs the original Spark-X2.5 Hugging Face checkpoints locally. It supports Apple silicon GPU, Linux CPU, and NVIDIA CUDA on Linux. No GGUF conversion is required.
230
+
231
+ #### Installation
232
+
233
+ ```bash
234
+ git clone https://github.com/XHToken/Spark-MLX-LLM.git
235
+ cd Spark-MLX-LLM
236
+
237
+ python3 -m venv .venv
238
+ source .venv/bin/activate
239
+
240
+ # Apple silicon
241
+ python -m pip install -e .
242
+ # Linux cpu
243
+ python -m pip install -e '.[cpu]'
244
+ # Linux with cuda12
245
+ python -m pip install -e '.[cuda12]'
246
+ # Linux with cuda13
247
+ python -m pip install -e '.[cuda13]'
248
+ ```
249
+
250
+ Run Spark-X2.5
251
+
252
+ ```bash
253
+ spark-mlx-generate \
254
+ --device gpu \
255
+ --dtype bfloat16 \
256
+ --model XHToken/Spark-X2.5-4B \
257
+ --prompt "安徽的省会在哪里?" \
258
+ --max-tokens 512 \
259
+ --temp 0
260
+ ```
261
+
262
+ ### Ollama
263
+
264
+ #### Build
265
+
266
+ ```bash
267
+ git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
268
+ git clone https://github.com/ollama/ollama.git ollama-spark
269
+ cd ollama-spark
270
+ export OLLAMA_LLAMA_CPP_SOURCE="$(cd ../llama.cpp-spark && pwd)"
271
+ cmake -S . -B build
272
+ cmake --build build --parallel 8
273
+ ```
274
+
275
+ #### Create and Run
276
+
277
+ ```bash
278
+ printf 'FROM /absolute/path/to/your.gguf\n' > ./Modelfile.spark
279
+ ./ollama serve
280
+ ./ollama create Spark-X2.5-4B -f ./Modelfile.spark
281
+ ./ollama run Spark-X2.5-4B
282
+ ```
283
+
284
+ ### LM Studio
285
+
286
+ #### Build
287
+
288
+ ```bash
289
+ git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
290
+ cd llama.cpp-spark
291
+ cmake -S . -B build
292
+ cmake --build build --parallel 8
293
+ ```
294
+
295
+ #### Set Up LM Studio
296
+
297
+ 1. Close LM Studio.
298
+
299
+ 2. Back up the selected runtime directory:
300
+
301
+ ```text
302
+ <LM_STUDIO_HOME>/extensions/backends/<selected-runtime>/
303
+ ```
304
+
305
+ 3. Copy the `llama.cpp-spark` build output into the selected runtime directory, overwriting the existing files.
306
+
307
+ 4. Place the GGUF model in the following directory:
308
+
309
+ ```text
310
+ <LM_STUDIO_HOME>/models/<org>/<name>/
311
+ ```
312
+
313
+ Example runtime directory on macOS:
314
+
315
+ ```text
316
+ ./build/bin/* -> ~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-<version>/
317
+ ```
318
+
319
+ #### Run with LM Studio
320
+
321
+ Open My Models, select the Spark-X2.5 model, click Load, then start a new Chat.
322
+
323
+ #### Run with lms cli
324
+
325
+ ```bash
326
+ # Replace `<model>` with a model listed by `lms ls`
327
+ lms load <model>
328
+ lms chat <model>
329
+ ```
330
+
331
+ ### Finetuning
332
+
333
+ We advise you to use [Llama-Factory](https://github.com/XHToken/LlamaFactory) to finetune your models.
334
+
335
+
336
+ ## License
337
+
338
+ The Spark-X2.5 model series is licensed under the [Apache 2.0 License](https://huggingface.co/XHToken/Spark-X2.5-4B/blob/main/LICENSE).
339
+
340
+ ## Citation
341
+ If you find our work helpful, feel free to give us a cite.
342
+
343
+ ```bibtex
344
+ @misc{sparkx2.5,
345
+ title = {Spark-X2.5 4B&1.7B: Pushing the Limits of Agentic Capabilities in On-Device Models},
346
+ author = {SparkLLM Team},
347
+ year = {2026}
348
+ }
349
+ ```
chat_template.jinja ADDED
@@ -0,0 +1,110 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {#- 0826版本 -#}
2
+ {%- if not messages %}
3
+ {{- raise_exception('No messages provided.') }}
4
+ {%- endif %}
5
+ {%- set enable_thinking = enable_thinking | default(true) %}
6
+
7
+ {#- Render a string or a list of text blocks. -#}
8
+ {%- macro render_content(content, context_name) %}
9
+ {%- if content is string %}
10
+ {{- content }}
11
+ {%- elif content is none or content is undefined %}
12
+ {{- '' }}
13
+ {%- elif content is iterable and content is not mapping %}
14
+ {%- for block in content %}
15
+ {%- if block.type == 'text' %}
16
+ {{- block.text }}
17
+ {%- else %}
18
+ {{- raise_exception('Unsupported ' ~ context_name ~ ' content block type: ' ~ (block.type | string)) }}
19
+ {%- endif %}
20
+ {%- endfor %}
21
+ {%- else %}
22
+ {{- raise_exception(context_name ~ ' content must be a string or a list of text blocks') }}
23
+ {%- endif %}
24
+ {%- endmacro %}
25
+
26
+ {#- Default system prompt -#}
27
+ {%- set default_system = "you are a helpful assistant." %}
28
+
29
+ {#- The first message-level system is placed in the initial system block. -#}
30
+ {%- set ns = namespace(initial_system='') %}
31
+ {%- if messages[0].role == "system" %}
32
+ {%- set ns.initial_system = render_content(messages[0].content, 'system') %}
33
+ {%- endif %}
34
+
35
+ {#- System block -#}
36
+ {{- '<|start▁of▁sentence|><|System|>' + '\n' + default_system }}
37
+ {%- if tools %}
38
+ {{- '## Tools' + '\n' + 'You have access to the following functions:' + '\n' + '<tools>' }}
39
+ {%- for tool in tools %}
40
+ {{- '\n' + tool.function | tojson}}
41
+ {%- endfor %}
42
+ {{- '\n' + '</tools>' }}
43
+ {%- endif %}
44
+ {%- if ns.initial_system %}
45
+ {{- '\n\n' + ns.initial_system }}
46
+ {%- endif %}
47
+ {{- '<|end▁of▁sentence|>'}}
48
+
49
+ {#- Conversation turns -#}
50
+ {%- for message in messages %}
51
+ {%- if message.role == "system" %}
52
+ {#- The first system message was consumed by the initial block. -#}
53
+ {%- if not loop.first %}
54
+ {{- '<|start▁of▁sentence|><|System|>\n' + render_content(message.content, 'system') + '<|end▁of▁sentence|>' }}
55
+ {%- endif %}
56
+ {%- elif message.role == "user" %}
57
+ {{- '<|start▁of▁sentence|><|User|>' + render_content(message.content, 'user') + '<|end▁of▁sentence|>' }}
58
+ {%- elif message.role == "assistant" %}
59
+ {%- set assistant_content = render_content(message.content, 'assistant') %}
60
+ {%- if message.reasoning_content is defined and message.reasoning_content %}
61
+ {%- set reasoning_content = message.reasoning_content %}
62
+ {%- else %}
63
+ {%- set reasoning_content = '' %}
64
+ {%- endif %}
65
+ {{- '<|start▁of▁sentence|><|Bot|>'}}
66
+ {%- if reasoning_content %}
67
+ {{- '<think>' + reasoning_content + '</think>'}}
68
+ {%- else %}
69
+ {{- '</think>' }}
70
+ {%- endif %}
71
+ {%- if assistant_content %}
72
+ {{- assistant_content }}
73
+ {%- endif %}
74
+ {%- if message.tool_calls is defined and message.tool_calls is not none %}
75
+ {%- for tool_call in message.tool_calls %}
76
+ {%- if tool_call.function.arguments is not mapping %}
77
+ {{- raise_exception('tool_call.function.arguments must be a dictionary; normalize JSON strings before apply_chat_template') }}
78
+ {%- endif %}
79
+ {%- set args = tool_call.function.arguments %}
80
+ {{- '<tool_call>' + tool_call.function.name }}
81
+ {%- for k, v in args.items() %}
82
+ {{- '<arg_key>' ~ k ~ '</arg_key><arg_value>' ~ (v if v is string else v | tojson) ~ '</arg_value>' }}
83
+ {%- endfor %}
84
+ {{- '</tool_call>' }}
85
+ {%- endfor %}
86
+ {%- endif %}
87
+ {{- '<|end▁of▁sentence|>' }}
88
+ {%- elif message.role == "tool" %}
89
+ {%- if loop.previtem is undefined or loop.previtem.role != "tool" %}
90
+ {{- '<|start▁of▁sentence|><|Tool|>' }}
91
+ {%- endif %}
92
+ {{- '<tool_response>' ~ message.content ~ '</tool_response>' }}
93
+ {%- if loop.nextitem is undefined or loop.nextitem.role != "tool" %}
94
+ {{- '<|end▁of▁sentence|>' }}
95
+ {%- endif %}
96
+ {%- else %}
97
+ {{- raise_exception('Unsupported message role: ' ~ message.role) }}
98
+ {%- endif %}
99
+ {%- endfor %}
100
+
101
+ {#- Generation prompt -#}
102
+ {%- if add_generation_prompt %}
103
+ {{- '<|start▁of▁sentence|><|Bot|>' }}
104
+ {%- if enable_thinking is defined and enable_thinking %}
105
+ {{- '<think>' }}
106
+ {%- endif %}
107
+ {%- if enable_thinking is defined and not enable_thinking %}
108
+ {{- '</think>' }}
109
+ {%- endif %}
110
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,83 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Spark2_5ForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "auto_map": {
8
+ "AutoConfig": "configuration_spark.Spark2_5Config",
9
+ "AutoModel": "modeling_spark.Spark2_5Model",
10
+ "AutoModelForCausalLM": "modeling_spark.Spark2_5ForCausalLM"
11
+ },
12
+ "bos_token_id": 0,
13
+ "dtype": "bfloat16",
14
+ "eos_token_id": 1,
15
+ "gate_attn_act_mode": "sigmoid",
16
+ "head_dim": 256,
17
+ "headwise_attn_output_gate": true,
18
+ "hidden_act": "gelu",
19
+ "hidden_size": 2560,
20
+ "initializer_range": 0.01976,
21
+ "intermediate_size": 10240,
22
+ "layer_types": [
23
+ "sliding_attention",
24
+ "sliding_attention",
25
+ "sliding_attention",
26
+ "full_attention",
27
+ "sliding_attention",
28
+ "sliding_attention",
29
+ "sliding_attention",
30
+ "full_attention",
31
+ "sliding_attention",
32
+ "sliding_attention",
33
+ "sliding_attention",
34
+ "full_attention",
35
+ "sliding_attention",
36
+ "sliding_attention",
37
+ "sliding_attention",
38
+ "full_attention",
39
+ "sliding_attention",
40
+ "sliding_attention",
41
+ "sliding_attention",
42
+ "full_attention",
43
+ "sliding_attention",
44
+ "sliding_attention",
45
+ "sliding_attention",
46
+ "full_attention",
47
+ "sliding_attention",
48
+ "sliding_attention",
49
+ "sliding_attention",
50
+ "full_attention",
51
+ "sliding_attention",
52
+ "sliding_attention",
53
+ "sliding_attention",
54
+ "full_attention",
55
+ "sliding_attention",
56
+ "sliding_attention",
57
+ "sliding_attention",
58
+ "full_attention"
59
+ ],
60
+ "max_position_embeddings": 1048576,
61
+ "mlp_bias": false,
62
+ "model_type": "spark2_5",
63
+ "num_attention_heads": 16,
64
+ "num_hidden_layers": 36,
65
+ "num_key_value_heads": 4,
66
+ "pad_token_id": 2,
67
+ "rms_norm_eps": 1e-06,
68
+ "rope_parameters": {
69
+ "full_attention": {
70
+ "partial_rotary_factor": 0.25,
71
+ "rope_theta": 5000000
72
+ },
73
+ "sliding_attention": {
74
+ "partial_rotary_factor": 1.0,
75
+ "rope_theta": 10000
76
+ }
77
+ },
78
+ "sliding_window": 512,
79
+ "tie_word_embeddings": true,
80
+ "transformers_version": "4.57.1",
81
+ "use_cache": true,
82
+ "vocab_size": 131072
83
+ }
configuration_spark.py ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # coding=utf-8
2
+ # Copyright 2026 The XHToken team and the HuggingFace Inc. team. All rights reserved.
3
+ #
4
+ # Licensed under the Apache License, Version 2.0 (the "License");
5
+ # you may not use this file except in compliance with the License.
6
+ # You may obtain a copy of the License at
7
+ #
8
+ # http://www.apache.org/licenses/LICENSE-2.0
9
+ #
10
+ # Unless required by applicable law or agreed to in writing, software
11
+ # distributed under the License is distributed on an "AS IS" BASIS,
12
+ # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
13
+ # See the License for the specific language governing permissions and
14
+ # limitations under the License.
15
+
16
+ from transformers import PretrainedConfig
17
+
18
+
19
+ class Spark2_5Config(PretrainedConfig):
20
+ model_type = "spark2_5"
21
+ keys_to_ignore_at_inference = ["past_key_values"]
22
+
23
+ base_model_tp_plan = {
24
+ "layers.*.self_attn.q_k_v_proj": "colwise",
25
+ "layers.*.self_attn.g_proj": "colwise",
26
+ "layers.*.self_attn.out_proj": "rowwise",
27
+ "layers.*.mlp.gate_proj": "colwise",
28
+ "layers.*.mlp.up_proj": "colwise",
29
+ "layers.*.mlp.down_proj": "rowwise",
30
+ }
31
+ base_model_pp_plan = {
32
+ "embedding": (["input_ids"], ["inputs_embeds"]),
33
+ "layers": (["hidden_states", "attention_mask"], ["hidden_states"]),
34
+ "norm": (["hidden_states"], ["hidden_states"]),
35
+ }
36
+
37
+ def __init__(
38
+ self,
39
+ vocab_size=32000,
40
+ hidden_size=4096,
41
+ intermediate_size=11008,
42
+ num_hidden_layers=32,
43
+ num_attention_heads=32,
44
+ num_key_value_heads=None,
45
+ hidden_act="gelu",
46
+ max_position_embeddings=2048,
47
+ initializer_range=0.02,
48
+ rms_norm_eps=1e-6,
49
+ use_cache=True,
50
+ pad_token_id=None,
51
+ bos_token_id=1,
52
+ eos_token_id=2,
53
+ tie_word_embeddings=False,
54
+ rope_parameters=None,
55
+ attention_bias=False,
56
+ attention_dropout=0.0,
57
+ mlp_bias=False,
58
+ head_dim=None,
59
+ headwise_attn_output_gate=False,
60
+ gate_attn_act_mode="sigmoid",
61
+ sliding_window=None,
62
+ layer_types=None,
63
+ **kwargs,
64
+ ):
65
+ self.vocab_size = vocab_size
66
+ self.max_position_embeddings = max_position_embeddings
67
+ self.hidden_size = hidden_size
68
+ self.intermediate_size = intermediate_size
69
+ self.num_hidden_layers = num_hidden_layers
70
+ self.num_attention_heads = num_attention_heads
71
+
72
+ if num_key_value_heads is None:
73
+ num_key_value_heads = num_attention_heads
74
+ if num_attention_heads % num_key_value_heads != 0:
75
+ raise ValueError(
76
+ f"num_attention_heads ({num_attention_heads}) must be divisible by num_key_value_heads ({num_key_value_heads})"
77
+ )
78
+ self.num_key_value_heads = num_key_value_heads
79
+
80
+ self.hidden_act = hidden_act
81
+ self.initializer_range = initializer_range
82
+ self.rms_norm_eps = rms_norm_eps
83
+ self.use_cache = use_cache
84
+ self.attention_bias = attention_bias
85
+ self.attention_dropout = attention_dropout
86
+ self.mlp_bias = mlp_bias
87
+ self.head_dim = head_dim if head_dim is not None else self.hidden_size // self.num_attention_heads
88
+ self.headwise_attn_output_gate = headwise_attn_output_gate
89
+ self.gate_attn_act_mode = gate_attn_act_mode
90
+ self.sliding_window = sliding_window
91
+ self.rope_parameters = rope_parameters
92
+
93
+ if layer_types is None:
94
+ layer_types = ["full_attention"] * num_hidden_layers
95
+ if len(layer_types) != num_hidden_layers:
96
+ raise ValueError(
97
+ f"layer_types length ({len(layer_types)}) must match num_hidden_layers ({num_hidden_layers})"
98
+ )
99
+ self.layer_types = layer_types
100
+
101
+ super().__init__(
102
+ pad_token_id=pad_token_id,
103
+ bos_token_id=bos_token_id,
104
+ eos_token_id=eos_token_id,
105
+ tie_word_embeddings=tie_word_embeddings,
106
+ **kwargs,
107
+ )
108
+
109
+ def get_rope_theta(self, layer_type):
110
+ params = self.rope_parameters.get(layer_type, {})
111
+ return params.get("rope_theta", 10000)
112
+
113
+ def get_partial_rotary_factor(self, layer_type):
114
+ params = self.rope_parameters.get(layer_type, {})
115
+ return params.get("partial_rotary_factor", 1.0)
116
+
117
+
118
+ __all__ = ["Spark2_5Config"]
generation_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 0,
3
+ "eos_token_id": 1,
4
+ "pad_token_id": 2,
5
+ "max_tokens": 1048576,
6
+ "temperature": 1.0,
7
+ "top_p": 0.95,
8
+ "top_k": -1,
9
+ "repetition_penalty": 1.0,
10
+ "presence_penalty":0,
11
+ "frequency_penalty":0,
12
+ "do_sample": true,
13
+ "transformers_version": "4.57.1"
14
+ }
images/model-benchmark-comparison.svg ADDED
images/post_training_pipeline.svg ADDED
images/spark25-hybrid-architecture-light.png ADDED

Git LFS Details

  • SHA256: 5eddac6baf8e8460cc2a75b96070eff4074428674f3c33665bc2ff9479756bbd
  • Pointer size: 131 Bytes
  • Size of remote file: 429 kB
images/xhtoken-wechat.jpg ADDED

Git LFS Details

  • SHA256: 0c3fbd25c987237b383461fdaa96bf6881af615c23ba3baf19e51a7879354ddb
  • Pointer size: 131 Bytes
  • Size of remote file: 160 kB
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model-00001-of-00005.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cc36bfe9ca2b839df89fd8c0c94f566f6717bd07ae859dfad074b303f90bdf5b
3
+ size 1982449560
model-00002-of-00005.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d98ea6b9059433181042c1ff9fbec167ae2f0850029787c82520d04a04942952
3
+ size 1993132424
model-00003-of-00005.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0ba22891378cf0ef3d9bf1f73c3159df3fc3dab8aa6e17679a9ffe7e9d867526
3
+ size 1993225080
model-00004-of-00005.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6aae0314fdf8fc3801677c017f296bcb14bc16095286c3bd1e7ce27e0398a113
3
+ size 1993132448
model-00005-of-00005.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ffe48ea1c2ee668407dd6ded0c02529da9c54580a8c38f6e53cf902b85a1bd95
3
+ size 262252896
model.safetensors.index.json ADDED
@@ -0,0 +1,298 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata": {
3
+ "total_parameters": 4112079360,
4
+ "total_size": 8224158720
5
+ },
6
+ "weight_map": {
7
+ "model.embedding.weight": "model-00001-of-00005.safetensors",
8
+ "model.layers.0.input_layernorm.weight": "model-00001-of-00005.safetensors",
9
+ "model.layers.0.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
10
+ "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
11
+ "model.layers.0.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
12
+ "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
13
+ "model.layers.0.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
14
+ "model.layers.0.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
15
+ "model.layers.0.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
16
+ "model.layers.1.input_layernorm.weight": "model-00001-of-00005.safetensors",
17
+ "model.layers.1.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
18
+ "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
19
+ "model.layers.1.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
20
+ "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
21
+ "model.layers.1.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
22
+ "model.layers.1.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
23
+ "model.layers.1.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
24
+ "model.layers.10.input_layernorm.weight": "model-00002-of-00005.safetensors",
25
+ "model.layers.10.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
26
+ "model.layers.10.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
27
+ "model.layers.10.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
28
+ "model.layers.10.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
29
+ "model.layers.10.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
30
+ "model.layers.10.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
31
+ "model.layers.10.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
32
+ "model.layers.11.input_layernorm.weight": "model-00002-of-00005.safetensors",
33
+ "model.layers.11.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
34
+ "model.layers.11.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
35
+ "model.layers.11.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
36
+ "model.layers.11.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
37
+ "model.layers.11.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
38
+ "model.layers.11.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
39
+ "model.layers.11.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
40
+ "model.layers.12.input_layernorm.weight": "model-00002-of-00005.safetensors",
41
+ "model.layers.12.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
42
+ "model.layers.12.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
43
+ "model.layers.12.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
44
+ "model.layers.12.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
45
+ "model.layers.12.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
46
+ "model.layers.12.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
47
+ "model.layers.12.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
48
+ "model.layers.13.input_layernorm.weight": "model-00002-of-00005.safetensors",
49
+ "model.layers.13.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
50
+ "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
51
+ "model.layers.13.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
52
+ "model.layers.13.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
53
+ "model.layers.13.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
54
+ "model.layers.13.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
55
+ "model.layers.13.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
56
+ "model.layers.14.input_layernorm.weight": "model-00002-of-00005.safetensors",
57
+ "model.layers.14.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
58
+ "model.layers.14.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
59
+ "model.layers.14.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
60
+ "model.layers.14.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
61
+ "model.layers.14.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
62
+ "model.layers.14.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
63
+ "model.layers.14.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
64
+ "model.layers.15.input_layernorm.weight": "model-00003-of-00005.safetensors",
65
+ "model.layers.15.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
66
+ "model.layers.15.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
67
+ "model.layers.15.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
68
+ "model.layers.15.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
69
+ "model.layers.15.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
70
+ "model.layers.15.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
71
+ "model.layers.15.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
72
+ "model.layers.16.input_layernorm.weight": "model-00003-of-00005.safetensors",
73
+ "model.layers.16.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
74
+ "model.layers.16.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
75
+ "model.layers.16.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
76
+ "model.layers.16.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
77
+ "model.layers.16.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
78
+ "model.layers.16.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
79
+ "model.layers.16.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
80
+ "model.layers.17.input_layernorm.weight": "model-00003-of-00005.safetensors",
81
+ "model.layers.17.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
82
+ "model.layers.17.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
83
+ "model.layers.17.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
84
+ "model.layers.17.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
85
+ "model.layers.17.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
86
+ "model.layers.17.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
87
+ "model.layers.17.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
88
+ "model.layers.18.input_layernorm.weight": "model-00003-of-00005.safetensors",
89
+ "model.layers.18.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
90
+ "model.layers.18.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
91
+ "model.layers.18.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
92
+ "model.layers.18.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
93
+ "model.layers.18.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
94
+ "model.layers.18.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
95
+ "model.layers.18.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
96
+ "model.layers.19.input_layernorm.weight": "model-00003-of-00005.safetensors",
97
+ "model.layers.19.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
98
+ "model.layers.19.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
99
+ "model.layers.19.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
100
+ "model.layers.19.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
101
+ "model.layers.19.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
102
+ "model.layers.19.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
103
+ "model.layers.19.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
104
+ "model.layers.2.input_layernorm.weight": "model-00001-of-00005.safetensors",
105
+ "model.layers.2.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
106
+ "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
107
+ "model.layers.2.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
108
+ "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
109
+ "model.layers.2.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
110
+ "model.layers.2.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
111
+ "model.layers.2.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
112
+ "model.layers.20.input_layernorm.weight": "model-00003-of-00005.safetensors",
113
+ "model.layers.20.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
114
+ "model.layers.20.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
115
+ "model.layers.20.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
116
+ "model.layers.20.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
117
+ "model.layers.20.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
118
+ "model.layers.20.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
119
+ "model.layers.20.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
120
+ "model.layers.21.input_layernorm.weight": "model-00003-of-00005.safetensors",
121
+ "model.layers.21.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
122
+ "model.layers.21.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
123
+ "model.layers.21.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
124
+ "model.layers.21.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
125
+ "model.layers.21.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
126
+ "model.layers.21.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
127
+ "model.layers.21.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
128
+ "model.layers.22.input_layernorm.weight": "model-00003-of-00005.safetensors",
129
+ "model.layers.22.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
130
+ "model.layers.22.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
131
+ "model.layers.22.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
132
+ "model.layers.22.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
133
+ "model.layers.22.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
134
+ "model.layers.22.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
135
+ "model.layers.22.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
136
+ "model.layers.23.input_layernorm.weight": "model-00003-of-00005.safetensors",
137
+ "model.layers.23.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
138
+ "model.layers.23.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
139
+ "model.layers.23.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
140
+ "model.layers.23.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
141
+ "model.layers.23.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
142
+ "model.layers.23.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
143
+ "model.layers.23.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
144
+ "model.layers.24.input_layernorm.weight": "model-00003-of-00005.safetensors",
145
+ "model.layers.24.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
146
+ "model.layers.24.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
147
+ "model.layers.24.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
148
+ "model.layers.24.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
149
+ "model.layers.24.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
150
+ "model.layers.24.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
151
+ "model.layers.24.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
152
+ "model.layers.25.input_layernorm.weight": "model-00004-of-00005.safetensors",
153
+ "model.layers.25.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
154
+ "model.layers.25.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
155
+ "model.layers.25.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
156
+ "model.layers.25.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
157
+ "model.layers.25.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
158
+ "model.layers.25.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
159
+ "model.layers.25.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
160
+ "model.layers.26.input_layernorm.weight": "model-00004-of-00005.safetensors",
161
+ "model.layers.26.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
162
+ "model.layers.26.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
163
+ "model.layers.26.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
164
+ "model.layers.26.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
165
+ "model.layers.26.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
166
+ "model.layers.26.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
167
+ "model.layers.26.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
168
+ "model.layers.27.input_layernorm.weight": "model-00004-of-00005.safetensors",
169
+ "model.layers.27.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
170
+ "model.layers.27.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
171
+ "model.layers.27.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
172
+ "model.layers.27.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
173
+ "model.layers.27.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
174
+ "model.layers.27.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
175
+ "model.layers.27.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
176
+ "model.layers.28.input_layernorm.weight": "model-00004-of-00005.safetensors",
177
+ "model.layers.28.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
178
+ "model.layers.28.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
179
+ "model.layers.28.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
180
+ "model.layers.28.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
181
+ "model.layers.28.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
182
+ "model.layers.28.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
183
+ "model.layers.28.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
184
+ "model.layers.29.input_layernorm.weight": "model-00004-of-00005.safetensors",
185
+ "model.layers.29.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
186
+ "model.layers.29.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
187
+ "model.layers.29.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
188
+ "model.layers.29.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
189
+ "model.layers.29.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
190
+ "model.layers.29.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
191
+ "model.layers.29.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
192
+ "model.layers.3.input_layernorm.weight": "model-00001-of-00005.safetensors",
193
+ "model.layers.3.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
194
+ "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
195
+ "model.layers.3.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
196
+ "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
197
+ "model.layers.3.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
198
+ "model.layers.3.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
199
+ "model.layers.3.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
200
+ "model.layers.30.input_layernorm.weight": "model-00004-of-00005.safetensors",
201
+ "model.layers.30.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
202
+ "model.layers.30.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
203
+ "model.layers.30.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
204
+ "model.layers.30.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
205
+ "model.layers.30.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
206
+ "model.layers.30.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
207
+ "model.layers.30.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
208
+ "model.layers.31.input_layernorm.weight": "model-00004-of-00005.safetensors",
209
+ "model.layers.31.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
210
+ "model.layers.31.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
211
+ "model.layers.31.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
212
+ "model.layers.31.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
213
+ "model.layers.31.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
214
+ "model.layers.31.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
215
+ "model.layers.31.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
216
+ "model.layers.32.input_layernorm.weight": "model-00004-of-00005.safetensors",
217
+ "model.layers.32.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
218
+ "model.layers.32.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
219
+ "model.layers.32.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
220
+ "model.layers.32.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
221
+ "model.layers.32.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
222
+ "model.layers.32.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
223
+ "model.layers.32.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
224
+ "model.layers.33.input_layernorm.weight": "model-00004-of-00005.safetensors",
225
+ "model.layers.33.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
226
+ "model.layers.33.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
227
+ "model.layers.33.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
228
+ "model.layers.33.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
229
+ "model.layers.33.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
230
+ "model.layers.33.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
231
+ "model.layers.33.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
232
+ "model.layers.34.input_layernorm.weight": "model-00005-of-00005.safetensors",
233
+ "model.layers.34.mlp.down_proj.weight": "model-00005-of-00005.safetensors",
234
+ "model.layers.34.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
235
+ "model.layers.34.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
236
+ "model.layers.34.post_attention_layernorm.weight": "model-00005-of-00005.safetensors",
237
+ "model.layers.34.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
238
+ "model.layers.34.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
239
+ "model.layers.34.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
240
+ "model.layers.35.input_layernorm.weight": "model-00005-of-00005.safetensors",
241
+ "model.layers.35.mlp.down_proj.weight": "model-00005-of-00005.safetensors",
242
+ "model.layers.35.mlp.gate_proj.weight": "model-00005-of-00005.safetensors",
243
+ "model.layers.35.mlp.up_proj.weight": "model-00005-of-00005.safetensors",
244
+ "model.layers.35.post_attention_layernorm.weight": "model-00005-of-00005.safetensors",
245
+ "model.layers.35.self_attn.g_proj.weight": "model-00005-of-00005.safetensors",
246
+ "model.layers.35.self_attn.out_proj.weight": "model-00005-of-00005.safetensors",
247
+ "model.layers.35.self_attn.q_k_v_proj.weight": "model-00005-of-00005.safetensors",
248
+ "model.layers.4.input_layernorm.weight": "model-00001-of-00005.safetensors",
249
+ "model.layers.4.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
250
+ "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
251
+ "model.layers.4.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
252
+ "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
253
+ "model.layers.4.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
254
+ "model.layers.4.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
255
+ "model.layers.4.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
256
+ "model.layers.5.input_layernorm.weight": "model-00001-of-00005.safetensors",
257
+ "model.layers.5.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
258
+ "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
259
+ "model.layers.5.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
260
+ "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
261
+ "model.layers.5.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
262
+ "model.layers.5.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
263
+ "model.layers.5.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
264
+ "model.layers.6.input_layernorm.weight": "model-00002-of-00005.safetensors",
265
+ "model.layers.6.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
266
+ "model.layers.6.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
267
+ "model.layers.6.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
268
+ "model.layers.6.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
269
+ "model.layers.6.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
270
+ "model.layers.6.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
271
+ "model.layers.6.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
272
+ "model.layers.7.input_layernorm.weight": "model-00002-of-00005.safetensors",
273
+ "model.layers.7.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
274
+ "model.layers.7.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
275
+ "model.layers.7.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
276
+ "model.layers.7.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
277
+ "model.layers.7.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
278
+ "model.layers.7.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
279
+ "model.layers.7.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
280
+ "model.layers.8.input_layernorm.weight": "model-00002-of-00005.safetensors",
281
+ "model.layers.8.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
282
+ "model.layers.8.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
283
+ "model.layers.8.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
284
+ "model.layers.8.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
285
+ "model.layers.8.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
286
+ "model.layers.8.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
287
+ "model.layers.8.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
288
+ "model.layers.9.input_layernorm.weight": "model-00002-of-00005.safetensors",
289
+ "model.layers.9.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
290
+ "model.layers.9.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
291
+ "model.layers.9.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
292
+ "model.layers.9.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
293
+ "model.layers.9.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
294
+ "model.layers.9.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
295
+ "model.layers.9.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
296
+ "model.norm.weight": "model-00005-of-00005.safetensors"
297
+ }
298
+ }
modeling_spark.py ADDED
@@ -0,0 +1,483 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import math
2
+
3
+ import torch
4
+ import torch.nn.functional as F
5
+ from torch import nn
6
+ from transformers.activations import ACT2FN
7
+ from transformers.cache_utils import Cache, DynamicCache
8
+ from transformers.generation import GenerationMixin
9
+ from transformers.masking_utils import create_causal_mask, create_sliding_window_causal_mask
10
+ from transformers.modeling_outputs import (
11
+ BaseModelOutputWithPast,
12
+ CausalLMOutputWithPast,
13
+ )
14
+ from transformers.modeling_utils import PreTrainedModel
15
+ from transformers.processing_utils import Unpack
16
+ from transformers.pytorch_utils import ALL_LAYERNORM_LAYERS
17
+ from transformers.utils import TransformersKwargs, can_return_tuple, logging
18
+
19
+ from .configuration_spark import Spark2_5Config
20
+
21
+ logger = logging.get_logger(__name__)
22
+
23
+ _CONFIG_FOR_DOC = "Spark2_5Config"
24
+
25
+ def rotate_half(x):
26
+ x1 = x[..., : x.shape[-1] // 2]
27
+ x2 = x[..., x.shape[-1] // 2 :]
28
+ return torch.cat((-x2, x1), dim=-1)
29
+
30
+
31
+ def compute_rope_cos_sin(positions, head_dim, rope_theta, partial_rotary_factor=1.0, device="cpu"):
32
+ rope_head_dim = int(head_dim * partial_rotary_factor)
33
+ inv_freq = 1.0 / (rope_theta ** (torch.arange(0, rope_head_dim, 2, dtype=torch.int64).to(device="cpu", dtype=torch.float) / rope_head_dim))
34
+ inv_freq = inv_freq.to(device)
35
+ t = positions.to(device=device, dtype=torch.float32)
36
+ freqs = torch.outer(t, inv_freq)
37
+ freqs = torch.cat([freqs, freqs], dim=-1)
38
+ cos = freqs.cos()
39
+ sin = freqs.sin()
40
+ return cos, sin
41
+
42
+
43
+ def apply_rotary_pos_emb(x, cos, sin):
44
+ rope_head_dim = cos.shape[-1]
45
+ x_f32 = x.float()
46
+ if x_f32.shape[-1] > rope_head_dim:
47
+ x_rot = x_f32[..., :rope_head_dim]
48
+ x_pass = x_f32[..., rope_head_dim:]
49
+ c = cos.unsqueeze(0).unsqueeze(0)
50
+ s = sin.unsqueeze(0).unsqueeze(0)
51
+ x_rot = x_rot * c + rotate_half(x_rot) * s
52
+ result = torch.cat([x_rot, x_pass], dim=-1)
53
+ else:
54
+ c = cos.unsqueeze(0).unsqueeze(0)
55
+ s = sin.unsqueeze(0).unsqueeze(0)
56
+ result = x_f32 * c + rotate_half(x_f32) * s
57
+ return result.to(x.dtype)
58
+
59
+
60
+ def repeat_kv(hidden_states: torch.Tensor, n_rep: int) -> torch.Tensor:
61
+ batch, num_key_value_heads, slen, head_dim = hidden_states.shape
62
+ if n_rep == 1:
63
+ return hidden_states
64
+ hidden_states = hidden_states[:, :, None, :, :].expand(batch, num_key_value_heads, n_rep, slen, head_dim)
65
+ return hidden_states.reshape(batch, num_key_value_heads * n_rep, slen, head_dim)
66
+
67
+
68
+ def eager_attention_forward(
69
+ module: nn.Module,
70
+ query: torch.Tensor,
71
+ key: torch.Tensor,
72
+ value: torch.Tensor,
73
+ attention_mask: torch.Tensor | None = None,
74
+ scaling: float | None = None,
75
+ dropout: float = 0.0,
76
+ **kwargs: Unpack[TransformersKwargs],
77
+ ):
78
+ key = repeat_kv(key, module.num_key_value_groups)
79
+ value = repeat_kv(value, module.num_key_value_groups)
80
+
81
+ if scaling is None:
82
+ scaling = 1.0 / math.sqrt(query.shape[-1])
83
+
84
+ attn_weights = torch.matmul(query, key.transpose(2, 3)) * scaling
85
+ if attention_mask is not None:
86
+ causal_mask = attention_mask[:, :, :, : key.shape[-2]]
87
+ attn_weights = attn_weights + causal_mask
88
+
89
+ attn_weights = attn_weights - attn_weights.max(dim=-1, keepdim=True).values
90
+ attn_weights = F.softmax(attn_weights, dim=-1, dtype=torch.float32).to(query.dtype)
91
+ attn_weights = nn.functional.dropout(attn_weights, p=dropout, training=module.training)
92
+ attn_output = torch.matmul(attn_weights, value)
93
+ return attn_output, attn_weights
94
+
95
+
96
+ class Spark2_5RMSNorm(nn.Module):
97
+ def __init__(self, hidden_size, eps=1e-6):
98
+ super().__init__()
99
+ self.weight = nn.Parameter(torch.ones(hidden_size))
100
+ self.variance_epsilon = eps
101
+
102
+ def forward(self, hidden_states):
103
+ input_dtype = hidden_states.dtype
104
+ hidden_states = hidden_states.to(torch.float32)
105
+ variance = hidden_states.pow(2).mean(-1, keepdim=True)
106
+ hidden_states = hidden_states * torch.rsqrt(variance + self.variance_epsilon)
107
+ return (self.weight.float() * hidden_states).to(input_dtype)
108
+
109
+ def extra_repr(self):
110
+ return f"{tuple(self.weight.shape)}, eps={self.variance_epsilon}"
111
+
112
+
113
+ ALL_LAYERNORM_LAYERS.append(Spark2_5RMSNorm)
114
+
115
+
116
+ class Spark2_5MLP(nn.Module):
117
+ def __init__(self, config):
118
+ super().__init__()
119
+ self.config = config
120
+ self.hidden_size = config.hidden_size
121
+ self.intermediate_size = config.intermediate_size
122
+ self.gate_proj = nn.Linear(self.hidden_size, self.intermediate_size, bias=config.mlp_bias)
123
+ self.up_proj = nn.Linear(self.hidden_size, self.intermediate_size, bias=config.mlp_bias)
124
+ self.down_proj = nn.Linear(self.intermediate_size, self.hidden_size, bias=config.mlp_bias)
125
+
126
+ if config.hidden_act != "gelu":
127
+ raise ValueError(f"只支持hidden_act='gelu',当前传入:{config.hidden_act}")
128
+
129
+ self.act_fn = ACT2FN[config.hidden_act]
130
+
131
+ def forward(self, x):
132
+ return self.down_proj(self.act_fn(self.gate_proj(x)) * self.up_proj(x))
133
+
134
+
135
+ class Spark2_5Attention(nn.Module):
136
+ def __init__(self, config: Spark2_5Config, layer_idx: int | None = None):
137
+ super().__init__()
138
+ self.config = config
139
+ self.layer_idx = layer_idx
140
+ self.attention_dropout = config.attention_dropout
141
+ self.hidden_size = config.hidden_size
142
+ self.num_heads = config.num_attention_heads
143
+ self.head_dim = config.head_dim
144
+ self.num_key_value_heads = config.num_key_value_heads
145
+ self.num_key_value_groups = self.num_heads // self.num_key_value_heads
146
+ self.scaling = 1.0 / math.sqrt(self.head_dim)
147
+ self.headwise_attn_output_gate = config.headwise_attn_output_gate
148
+ self.gate_attn_act_mode = config.gate_attn_act_mode
149
+ self.q_dim = self.num_heads * self.head_dim
150
+ self.kv_dim = self.num_key_value_heads * self.head_dim
151
+
152
+ qkv_out_dim = self.q_dim + 2 * self.kv_dim
153
+ self.q_k_v_proj = nn.Linear(self.hidden_size, qkv_out_dim, bias=config.attention_bias)
154
+ self.g_proj = nn.Linear(self.hidden_size, self.num_heads, bias=config.attention_bias) if self.headwise_attn_output_gate else None
155
+ self.out_proj = nn.Linear(self.num_heads * self.head_dim, self.hidden_size, bias=config.attention_bias)
156
+ self.sliding_window = None
157
+
158
+ def forward(
159
+ self,
160
+ hidden_states: torch.Tensor,
161
+ position_embeddings: tuple[torch.Tensor, torch.Tensor],
162
+ attention_mask: torch.Tensor | None = None,
163
+ past_key_values: Cache | None = None,
164
+ cache_position: torch.LongTensor | None = None,
165
+ **kwargs: Unpack[TransformersKwargs],
166
+ ) -> tuple[torch.Tensor, torch.Tensor]:
167
+ input_shape = hidden_states.shape[:-1]
168
+ bsz, seq_len = input_shape
169
+
170
+ qkv = self.q_k_v_proj(hidden_states)
171
+ q = qkv[..., :self.q_dim]
172
+ k = qkv[..., self.q_dim:self.q_dim + self.kv_dim]
173
+ v = qkv[..., self.q_dim + self.kv_dim:]
174
+ gate_score = self.g_proj(hidden_states) if self.g_proj is not None else None
175
+
176
+ q = q.view(bsz, seq_len, self.num_heads, self.head_dim).transpose(1, 2)
177
+ k = k.view(bsz, seq_len, self.num_key_value_heads, self.head_dim).transpose(1, 2)
178
+ v = v.view(bsz, seq_len, self.num_key_value_heads, self.head_dim).transpose(1, 2)
179
+ if gate_score is not None:
180
+ gate_score = gate_score.view(bsz, seq_len, self.num_heads, 1).transpose(1, 2)
181
+
182
+ cos, sin = position_embeddings
183
+ q = apply_rotary_pos_emb(q, cos, sin)
184
+ k = apply_rotary_pos_emb(k, cos, sin)
185
+
186
+
187
+ if past_key_values is not None:
188
+ cache_kwargs = {"sin": sin, "cos": cos, "cache_position": cache_position}
189
+ k, v = past_key_values.update(k, v, self.layer_idx, cache_kwargs)
190
+
191
+ attn_output, attn_weights = eager_attention_forward(
192
+ self, q, k, v,
193
+ attention_mask=attention_mask,
194
+ scaling=self.scaling,
195
+ dropout=self.attention_dropout if self.training else 0.0,
196
+ )
197
+
198
+ if gate_score is not None:
199
+ if self.gate_attn_act_mode == "sigmoid":
200
+ gate = torch.sigmoid(gate_score.float())
201
+ elif self.gate_attn_act_mode == "silu":
202
+ gate = F.silu(gate_score.float())
203
+ else:
204
+ raise ValueError(f"Unsupported gate_attn_act_mode: {self.gate_attn_act_mode}")
205
+ gate = gate.to(attn_output.dtype)
206
+ attn_output = attn_output * gate
207
+
208
+ attn_output = attn_output.transpose(1, 2).contiguous().view(bsz, seq_len, -1)
209
+ attn_output = self.out_proj(attn_output)
210
+
211
+ return attn_output, attn_weights
212
+
213
+
214
+ class Spark2_5DecoderLayer(nn.Module):
215
+ def __init__(self, config: Spark2_5Config, layer_idx: int):
216
+ super().__init__()
217
+ self.hidden_size = config.hidden_size
218
+
219
+ self.self_attn = Spark2_5Attention(config=config, layer_idx=layer_idx)
220
+ self.mlp = Spark2_5MLP(config)
221
+ self.input_layernorm = Spark2_5RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
222
+ self.post_attention_layernorm = Spark2_5RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
223
+
224
+ self.layer_type = config.layer_types[layer_idx] if layer_idx < len(config.layer_types) else "full_attention"
225
+ if self.layer_type == "sliding_attention" and config.sliding_window is not None:
226
+ self.self_attn.sliding_window = config.sliding_window
227
+ else:
228
+ self.self_attn.sliding_window = None
229
+ self.self_attn.partial_rotary_factor = config.get_partial_rotary_factor(self.layer_type)
230
+
231
+ def forward(
232
+ self,
233
+ hidden_states: torch.Tensor,
234
+ position_embeddings: tuple[torch.Tensor, torch.Tensor],
235
+ attention_mask: torch.Tensor | None = None,
236
+ past_key_values: Cache | None = None,
237
+ cache_position: torch.LongTensor | None = None,
238
+ position_ids: torch.LongTensor | None = None,
239
+ **kwargs: Unpack[TransformersKwargs]
240
+ ) -> torch.Tensor:
241
+
242
+ residual = hidden_states
243
+ hidden_states = self.input_layernorm(hidden_states)
244
+ hidden_states = hidden_states.to(self.mlp.gate_proj.weight.dtype)
245
+
246
+ hidden_states, _ = self.self_attn(
247
+ hidden_states=hidden_states,
248
+ position_embeddings=position_embeddings,
249
+ attention_mask=attention_mask,
250
+ past_key_values=past_key_values,
251
+ cache_position=cache_position,
252
+ position_ids=position_ids,
253
+ )
254
+ hidden_states = residual + hidden_states
255
+
256
+ residual = hidden_states
257
+ hidden_states = self.post_attention_layernorm(hidden_states)
258
+ hidden_states = hidden_states.to(self.mlp.gate_proj.weight.dtype)
259
+
260
+
261
+ hidden_states = self.mlp(hidden_states)
262
+ hidden_states = residual + hidden_states
263
+
264
+ return hidden_states
265
+
266
+
267
+ class Spark2_5PreTrainedModel(PreTrainedModel):
268
+ config_class = Spark2_5Config
269
+ base_model_prefix = "model"
270
+ supports_gradient_checkpointing = True
271
+ _no_split_modules = ["Spark2_5DecoderLayer"] # noqa: RUF012
272
+ _skip_keys_device_placement = ["past_key_values"] # noqa: RUF012
273
+
274
+ def _init_weights(self, module):
275
+ std = self.config.initializer_range
276
+ if isinstance(module, nn.Linear):
277
+ module.weight.data.normal_(mean=0.0, std=std)
278
+ if module.bias is not None:
279
+ module.bias.data.zero_()
280
+ elif isinstance(module, nn.Embedding):
281
+ module.weight.data.normal_(mean=0.0, std=std)
282
+ if module.padding_idx is not None:
283
+ module.weight.data[module.padding_idx].zero_()
284
+
285
+
286
+ class Spark2_5Model(Spark2_5PreTrainedModel):
287
+ def __init__(self, config: Spark2_5Config):
288
+ super().__init__(config)
289
+ self.padding_idx = config.pad_token_id
290
+ self.vocab_size = config.vocab_size
291
+
292
+ self.embedding = nn.Embedding(config.vocab_size, config.hidden_size, self.padding_idx)
293
+ self.layers = nn.ModuleList(
294
+ [Spark2_5DecoderLayer(config, layer_idx) for layer_idx in range(config.num_hidden_layers)]
295
+ )
296
+ self.norm = Spark2_5RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
297
+ self.gradient_checkpointing = False
298
+ self.has_sliding_layers = "sliding_attention" in config.layer_types
299
+
300
+ self.post_init()
301
+
302
+ def get_input_embeddings(self):
303
+ return self.embedding
304
+
305
+ def set_input_embeddings(self, value):
306
+ self.embedding = value
307
+ def forward(
308
+ self,
309
+ input_ids: torch.LongTensor = None,
310
+ attention_mask: torch.Tensor | None = None,
311
+ position_ids: torch.LongTensor | None = None,
312
+ past_key_values: Cache | list[torch.FloatTensor] | None = None,
313
+ inputs_embeds: torch.FloatTensor | None = None,
314
+ use_cache: bool | None = None,
315
+ cache_position: torch.LongTensor | None = None,
316
+ token_type_ids: torch.LongTensor | None = None,
317
+ **kwargs: Unpack[TransformersKwargs],
318
+ ) -> BaseModelOutputWithPast:
319
+ use_cache = use_cache if use_cache is not None else self.config.use_cache
320
+ if (input_ids is None) ^ (inputs_embeds is not None):
321
+ raise ValueError(
322
+ "You cannot specify both input_ids and inputs_embeds at the same time, and must specify either one"
323
+ )
324
+
325
+ if self.gradient_checkpointing and self.training and use_cache:
326
+ logger.warning_once(
327
+ "`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`."
328
+ )
329
+ use_cache = False
330
+
331
+ if inputs_embeds is None:
332
+ inputs_embeds = self.embedding(input_ids)
333
+
334
+ if use_cache and past_key_values is None:
335
+ past_key_values = DynamicCache(config=self.config)
336
+
337
+ if cache_position is None:
338
+ past_seen_tokens = past_key_values.get_seq_length() if past_key_values is not None else 0
339
+ cache_position = torch.arange(
340
+ past_seen_tokens, past_seen_tokens + inputs_embeds.shape[1], device=inputs_embeds.device
341
+ )
342
+
343
+ if position_ids is None:
344
+ position_ids = cache_position.unsqueeze(0)
345
+
346
+ if not isinstance(attention_mask, dict):
347
+ mask_kwargs = {
348
+ "config": self.config,
349
+ "input_embeds": inputs_embeds,
350
+ "attention_mask": attention_mask,
351
+ "cache_position": cache_position,
352
+ "past_key_values": past_key_values,
353
+ "position_ids": position_ids,
354
+ }
355
+ causal_mask_mapping = {
356
+ "full_attention": create_causal_mask(**mask_kwargs),
357
+ }
358
+ if self.has_sliding_layers:
359
+ causal_mask_mapping["sliding_attention"] = create_sliding_window_causal_mask(**mask_kwargs)
360
+ else:
361
+ causal_mask_mapping = attention_mask
362
+
363
+ hidden_states = inputs_embeds.float()
364
+
365
+ device = hidden_states.device
366
+ dtype = self.embedding.weight.dtype
367
+
368
+ head_dim = self.config.head_dim
369
+ rope_cache = {}
370
+ for lt in set(self.config.layer_types):
371
+ rope_theta = self.config.get_rope_theta(lt)
372
+ prf = self.config.get_partial_rotary_factor(lt)
373
+ cos, sin = compute_rope_cos_sin(cache_position, head_dim, rope_theta, partial_rotary_factor=prf, device=device)
374
+ rope_cache[lt] = (cos, sin)
375
+
376
+ for decoder_layer in self.layers:
377
+ layer_type = decoder_layer.layer_type
378
+ position_embeddings = rope_cache.get(layer_type, rope_cache.get("full_attention"))
379
+ layer_attention_mask = causal_mask_mapping.get(layer_type, causal_mask_mapping.get("full_attention"))
380
+
381
+ if self.gradient_checkpointing and self.training:
382
+ layer_outputs = self._gradient_checkpointing_func(
383
+ decoder_layer.__call__,
384
+ hidden_states,
385
+ position_embeddings,
386
+ layer_attention_mask,
387
+ )
388
+ hidden_states = layer_outputs[0] if isinstance(layer_outputs, tuple) else layer_outputs
389
+ else:
390
+ hidden_states = decoder_layer(
391
+ hidden_states,
392
+ position_embeddings=position_embeddings,
393
+ attention_mask=layer_attention_mask,
394
+ past_key_values=past_key_values,
395
+ cache_position=cache_position,
396
+ position_ids=position_ids,
397
+ )
398
+
399
+ hidden_states = self.norm(hidden_states)
400
+ hidden_states = hidden_states.to(dtype)
401
+
402
+ return BaseModelOutputWithPast(
403
+ last_hidden_state=hidden_states,
404
+ past_key_values=past_key_values if use_cache else None,
405
+ )
406
+
407
+ class Spark2_5ForCausalLM(Spark2_5PreTrainedModel, GenerationMixin):
408
+ _tied_weights_keys = ["lm_head.weight"] # noqa: RUF012
409
+
410
+ def __init__(self, config):
411
+ super().__init__(config)
412
+ self.model = Spark2_5Model(config)
413
+ self.vocab_size = config.vocab_size
414
+ self.lm_head = nn.Linear(config.hidden_size, config.vocab_size, bias=False)
415
+ self.post_init()
416
+
417
+ def get_input_embeddings(self):
418
+ return self.model.embedding
419
+
420
+ def set_input_embeddings(self, value):
421
+ self.model.embedding = value
422
+
423
+ def get_output_embeddings(self):
424
+ return self.lm_head
425
+
426
+ def set_output_embeddings(self, new_embeddings):
427
+ self.lm_head = new_embeddings
428
+
429
+ def set_decoder(self, decoder):
430
+ self.model = decoder
431
+
432
+ def get_decoder(self):
433
+ return self.model
434
+
435
+ @can_return_tuple
436
+ def forward(
437
+ self,
438
+ input_ids: torch.LongTensor = None,
439
+ attention_mask: torch.Tensor | None = None,
440
+ position_ids: torch.LongTensor | None = None,
441
+ past_key_values: Cache | list[torch.FloatTensor] | None = None,
442
+ inputs_embeds: torch.FloatTensor | None = None,
443
+ labels: torch.LongTensor | None = None,
444
+ use_cache: bool | None = None,
445
+ cache_position: torch.LongTensor | None = None,
446
+ logits_to_keep: int = 0,
447
+ token_type_ids: torch.LongTensor | None = None,
448
+ **kwargs: Unpack[TransformersKwargs],
449
+ ) -> CausalLMOutputWithPast:
450
+ outputs: BaseModelOutputWithPast = self.model(
451
+ input_ids=input_ids,
452
+ attention_mask=attention_mask,
453
+ position_ids=position_ids,
454
+ past_key_values=past_key_values,
455
+ inputs_embeds=inputs_embeds,
456
+ use_cache=use_cache,
457
+ cache_position=cache_position,
458
+ )
459
+
460
+ hidden_states = outputs.last_hidden_state
461
+ slice_indices = slice(-logits_to_keep, None) if isinstance(logits_to_keep, int) else logits_to_keep
462
+ hidden_states = hidden_states[:, slice_indices, :]
463
+
464
+ if self.config.tie_word_embeddings:
465
+ embed_weight = self.model.embedding.weight
466
+ logits = F.linear(hidden_states, embed_weight)
467
+ else:
468
+ logits = self.lm_head(hidden_states)
469
+
470
+ loss = None
471
+ if labels is not None:
472
+ loss = self.loss_function(logits=logits, labels=labels, vocab_size=self.config.vocab_size, **kwargs)
473
+
474
+ return CausalLMOutputWithPast(
475
+ loss=loss,
476
+ logits=logits,
477
+ past_key_values=outputs.past_key_values,
478
+ hidden_states=outputs.hidden_states,
479
+ attentions=outputs.attentions,
480
+ )
481
+
482
+
483
+ __all__ = ["Spark2_5Config", "Spark2_5ForCausalLM", "Spark2_5Model"]
special_tokens_map.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<|start▁of▁sentence|>",
3
+ "eos_token": "<|end▁of▁sentence|>",
4
+ "unk_token": "<unk>",
5
+ "pad_token": "<|▁pad▁|>"
6
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:710cce15cf3565674c499f9413997c6e8101f2bdd96245cff8f0311fb501248c
3
+ size 10115786
tokenizer_config.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_eos_token": false,
4
+ "bos_token": {
5
+ "__type": "AddedToken",
6
+ "content": "<|start▁of▁sentence|>",
7
+ "lstrip": false,
8
+ "normalized": true,
9
+ "rstrip": false,
10
+ "single_word": false
11
+ },
12
+ "clean_up_tokenization_spaces": false,
13
+ "eos_token": {
14
+ "__type": "AddedToken",
15
+ "content": "<|end▁of▁sentence|>",
16
+ "lstrip": false,
17
+ "normalized": true,
18
+ "rstrip": false,
19
+ "single_word": false
20
+ },
21
+ "legacy": true,
22
+ "model_max_length": 131072,
23
+ "pad_token": {
24
+ "__type": "AddedToken",
25
+ "content": "<|▁pad▁|>",
26
+ "lstrip": false,
27
+ "normalized": true,
28
+ "rstrip": false,
29
+ "single_word": false
30
+ },
31
+ "sp_model_kwargs": {},
32
+ "unk_token": {
33
+ "__type": "AddedToken",
34
+ "content": "<unk>",
35
+ "lstrip": false,
36
+ "normalized": true,
37
+ "rstrip": false,
38
+ "single_word": false
39
+ },
40
+ "tokenizer_class": "PreTrainedTokenizerFast",
41
+ "chat_template": "{%- if not messages %}{{- raise_exception('No messages provided.') }}{%- endif %}{%- set enable_thinking = enable_thinking | default(true) %}{#- Render a string or a list of text blocks. -#}{%- macro render_content(content, context_name) %}{%- if content is string %}{{- content }}{%- elif content is none or content is undefined %}{{- '' }}{%- elif content is iterable and content is not mapping %}{%- for block in content %}{%- if block.type == 'text' %}{{- block.text }}{%- else %}{{- raise_exception('Unsupported ' ~ context_name ~ ' content block type: ' ~ (block.type | string)) }}{%- endif %}{%- endfor %}{%- else %}{{- raise_exception(context_name ~ ' content must be a string or a list of text blocks') }}{%- endif %}{%- endmacro %}{#- Default system prompt -#}{%- set default_system = 'you are a helpful assistant.' %}{#- The first message-level system is placed in the initial system block. -#}{%- set ns = namespace(initial_system='') %}{%- if messages[0].role == 'system' %}{%- set ns.initial_system = render_content(messages[0].content, 'system') %}{%- endif %}{#- System block -#}{{- '<|start▁of▁sentence|><|System|>' + '\n' + default_system }}{%- if tools %}{{- '## Tools' + '\n' + 'You have access to the following functions:' + '\n' + '<tools>' }}{%- for tool in tools %}{{- '\n' + tool.function | tojson}}{%- endfor %}{{- '\n' + '</tools>' }}{%- endif %}{%- if ns.initial_system %}{{- '\n\n' + ns.initial_system }}{%- endif %}{{- '<|end▁of▁sentence|>'}}{#- Conversation turns -#}{%- for message in messages %}{%- if message.role == 'system' %}{#- The first system message was consumed by the initial block. -#}{%- if not loop.first %}{{- '<|start▁of▁sentence|><|System|>\n' + render_content(message.content, 'system') + '<|end▁of▁sentence|>' }}{%- endif %}{%- elif message.role == 'user' %}{{- '<|start▁of▁sentence|><|User|>' + render_content(message.content, 'user') + '<|end▁of▁sentence|>' }}{%- elif message.role == 'assistant' %}{%- set assistant_content = render_content(message.content, 'assistant') %}{%- if message.reasoning_content is defined and message.reasoning_content %}{%- set reasoning_content = message.reasoning_content %}{%- else %}{%- set reasoning_content = '' %}{%- endif %}{{- '<|start▁of▁sentence|><|Bot|>'}}{%- if reasoning_content %}{{- '<think>' + reasoning_content + '</think>'}}{%- else %}{{- '</think>' }}{%- endif %}{%- if assistant_content %}{{- assistant_content }}{%- endif %}{%- if message.tool_calls is defined and message.tool_calls is not none %}{%- for tool_call in message.tool_calls %}{%- if tool_call.function.arguments is not mapping %}{{- raise_exception('tool_call.function.arguments must be a dictionary; normalize JSON strings before apply_chat_template') }}{%- endif %}{%- set args = tool_call.function.arguments %}{{- '<tool_call>' + tool_call.function.name }}{%- for k, v in args.items() %}{{- '<arg_key>' ~ k ~ '</arg_key><arg_value>' ~ (v if v is string else v | tojson) ~ '</arg_value>' }}{%- endfor %}{{- '</tool_call>' }}{%- endfor %}{%- endif %}{{- '<|end▁of▁sentence|>' }}{%- elif message.role == 'tool' %}{%- if loop.previtem is undefined or loop.previtem.role != 'tool' %}{{- '<|start▁of▁sentence|><|Tool|>' }}{%- endif %}{{- '<tool_response>' ~ message.content ~ '</tool_response>' }}{%- if loop.nextitem is undefined or loop.nextitem.role != 'tool' %}{{- '<|end▁of▁sentence|>' }}{%- endif %}{%- else %}{{- raise_exception('Unsupported message role: ' ~ message.role) }}{%- endif %}{%- endfor %}{#- Generation prompt -#}{%- if add_generation_prompt %}{{- '<|start▁of▁sentence|><|Bot|>' }}{%- if enable_thinking is defined and enable_thinking %}{{- '<think>' }}{%- endif %}{%- if enable_thinking is defined and not enable_thinking %}{{- '</think>' }}{%- endif %}{%- endif %}"
42
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff