Instructions to use XHToken/Spark-X2.5-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use XHToken/Spark-X2.5-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="XHToken/Spark-X2.5-4B", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("XHToken/Spark-X2.5-4B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use XHToken/Spark-X2.5-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "XHToken/Spark-X2.5-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XHToken/Spark-X2.5-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/XHToken/Spark-X2.5-4B
- SGLang
How to use XHToken/Spark-X2.5-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "XHToken/Spark-X2.5-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XHToken/Spark-X2.5-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "XHToken/Spark-X2.5-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XHToken/Spark-X2.5-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use XHToken/Spark-X2.5-4B with Docker Model Runner:
docker model run hf.co/XHToken/Spark-X2.5-4B
Initial release of Spark-X2.5-4B
Browse filesCo-authored-by: chao <linchao-2026@users.noreply.huggingface.co>
Co-authored-by: Dong Jiang <dongjiang1989@users.noreply.huggingface.co>
Co-authored-by: Lei Yang <lyang01@users.noreply.huggingface.co>
Co-authored-by: Marx Benjamin <Flyingmarx@users.noreply.huggingface.co>
Co-authored-by: TylerZhang <Tylerwtzhang3@users.noreply.huggingface.co>
Co-authored-by: Wenzhi Peng <pengwenzhi@users.noreply.huggingface.co>
Signed-off-by: chao <linchao-2026@users.noreply.huggingface.co>
Signed-off-by: Dong Jiang <dongjiang1989@users.noreply.huggingface.co>
Signed-off-by: Lei Yang <lyang01@users.noreply.huggingface.co>
Signed-off-by: Marx Benjamin <Flyingmarx@users.noreply.huggingface.co>
Signed-off-by: TylerZhang <Tylerwtzhang3@users.noreply.huggingface.co>
Signed-off-by: Wenzhi Peng <pengwenzhi@users.noreply.huggingface.co>
- .gitattributes +39 -0
- LICENSE +202 -0
- README.md +349 -0
- chat_template.jinja +110 -0
- config.json +83 -0
- configuration_spark.py +118 -0
- generation_config.json +14 -0
- images/model-benchmark-comparison.svg +654 -0
- images/post_training_pipeline.svg +123 -0
- images/spark25-hybrid-architecture-light.png +3 -0
- images/xhtoken-wechat.jpg +3 -0
- merges.txt +0 -0
- model-00001-of-00005.safetensors +3 -0
- model-00002-of-00005.safetensors +3 -0
- model-00003-of-00005.safetensors +3 -0
- model-00004-of-00005.safetensors +3 -0
- model-00005-of-00005.safetensors +3 -0
- model.safetensors.index.json +298 -0
- modeling_spark.py +483 -0
- special_tokens_map.json +6 -0
- tokenizer.json +3 -0
- tokenizer_config.json +42 -0
- vocab.json +0 -0
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
*.7z filter=lfs diff=lfs merge=lfs -text
|
| 2 |
+
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
+
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
+
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
+
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
+
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
+
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
+
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
+
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
+
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
+
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
+
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
+
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
+
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
+
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
+
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
+
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
+
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
+
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
+
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
+
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
+
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
+
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
+
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
+
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
+
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
| 27 |
+
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
| 28 |
+
*.tar filter=lfs diff=lfs merge=lfs -text
|
| 29 |
+
*.tflite filter=lfs diff=lfs merge=lfs -text
|
| 30 |
+
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
+
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
+
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
+
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
+
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
+
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
images/spark3-hybrid-architecture-light.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
images/xhtoken-wechat.jpg filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
images/spark25-hybrid-architecture-light.png filter=lfs diff=lfs merge=lfs -text
|
|
@@ -0,0 +1,202 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
|
| 2 |
+
Apache License
|
| 3 |
+
Version 2.0, January 2004
|
| 4 |
+
http://www.apache.org/licenses/
|
| 5 |
+
|
| 6 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 7 |
+
|
| 8 |
+
1. Definitions.
|
| 9 |
+
|
| 10 |
+
"License" shall mean the terms and conditions for use, reproduction,
|
| 11 |
+
and distribution as defined by Sections 1 through 9 of this document.
|
| 12 |
+
|
| 13 |
+
"Licensor" shall mean the copyright owner or entity authorized by
|
| 14 |
+
the copyright owner that is granting the License.
|
| 15 |
+
|
| 16 |
+
"Legal Entity" shall mean the union of the acting entity and all
|
| 17 |
+
other entities that control, are controlled by, or are under common
|
| 18 |
+
control with that entity. For the purposes of this definition,
|
| 19 |
+
"control" means (i) the power, direct or indirect, to cause the
|
| 20 |
+
direction or management of such entity, whether by contract or
|
| 21 |
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
| 22 |
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
| 23 |
+
|
| 24 |
+
"You" (or "Your") shall mean an individual or Legal Entity
|
| 25 |
+
exercising permissions granted by this License.
|
| 26 |
+
|
| 27 |
+
"Source" form shall mean the preferred form for making modifications,
|
| 28 |
+
including but not limited to software source code, documentation
|
| 29 |
+
source, and configuration files.
|
| 30 |
+
|
| 31 |
+
"Object" form shall mean any form resulting from mechanical
|
| 32 |
+
transformation or translation of a Source form, including but
|
| 33 |
+
not limited to compiled object code, generated documentation,
|
| 34 |
+
and conversions to other media types.
|
| 35 |
+
|
| 36 |
+
"Work" shall mean the work of authorship, whether in Source or
|
| 37 |
+
Object form, made available under the License, as indicated by a
|
| 38 |
+
copyright notice that is included in or attached to the work
|
| 39 |
+
(an example is provided in the Appendix below).
|
| 40 |
+
|
| 41 |
+
"Derivative Works" shall mean any work, whether in Source or Object
|
| 42 |
+
form, that is based on (or derived from) the Work and for which the
|
| 43 |
+
editorial revisions, annotations, elaborations, or other modifications
|
| 44 |
+
represent, as a whole, an original work of authorship. For the purposes
|
| 45 |
+
of this License, Derivative Works shall not include works that remain
|
| 46 |
+
separable from, or merely link (or bind by name) to the interfaces of,
|
| 47 |
+
the Work and Derivative Works thereof.
|
| 48 |
+
|
| 49 |
+
"Contribution" shall mean any work of authorship, including
|
| 50 |
+
the original version of the Work and any modifications or additions
|
| 51 |
+
to that Work or Derivative Works thereof, that is intentionally
|
| 52 |
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
| 53 |
+
or by an individual or Legal Entity authorized to submit on behalf of
|
| 54 |
+
the copyright owner. For the purposes of this definition, "submitted"
|
| 55 |
+
means any form of electronic, verbal, or written communication sent
|
| 56 |
+
to the Licensor or its representatives, including but not limited to
|
| 57 |
+
communication on electronic mailing lists, source code control systems,
|
| 58 |
+
and issue tracking systems that are managed by, or on behalf of, the
|
| 59 |
+
Licensor for the purpose of discussing and improving the Work, but
|
| 60 |
+
excluding communication that is conspicuously marked or otherwise
|
| 61 |
+
designated in writing by the copyright owner as "Not a Contribution."
|
| 62 |
+
|
| 63 |
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
| 64 |
+
on behalf of whom a Contribution has been received by Licensor and
|
| 65 |
+
subsequently incorporated within the Work.
|
| 66 |
+
|
| 67 |
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
| 68 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 69 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 70 |
+
copyright license to reproduce, prepare Derivative Works of,
|
| 71 |
+
publicly display, publicly perform, sublicense, and distribute the
|
| 72 |
+
Work and such Derivative Works in Source or Object form.
|
| 73 |
+
|
| 74 |
+
3. Grant of Patent License. Subject to the terms and conditions of
|
| 75 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 76 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 77 |
+
(except as stated in this section) patent license to make, have made,
|
| 78 |
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
| 79 |
+
where such license applies only to those patent claims licensable
|
| 80 |
+
by such Contributor that are necessarily infringed by their
|
| 81 |
+
Contribution(s) alone or by combination of their Contribution(s)
|
| 82 |
+
with the Work to which such Contribution(s) was submitted. If You
|
| 83 |
+
institute patent litigation against any entity (including a
|
| 84 |
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
| 85 |
+
or a Contribution incorporated within the Work constitutes direct
|
| 86 |
+
or contributory patent infringement, then any patent licenses
|
| 87 |
+
granted to You under this License for that Work shall terminate
|
| 88 |
+
as of the date such litigation is filed.
|
| 89 |
+
|
| 90 |
+
4. Redistribution. You may reproduce and distribute copies of the
|
| 91 |
+
Work or Derivative Works thereof in any medium, with or without
|
| 92 |
+
modifications, and in Source or Object form, provided that You
|
| 93 |
+
meet the following conditions:
|
| 94 |
+
|
| 95 |
+
(a) You must give any other recipients of the Work or
|
| 96 |
+
Derivative Works a copy of this License; and
|
| 97 |
+
|
| 98 |
+
(b) You must cause any modified files to carry prominent notices
|
| 99 |
+
stating that You changed the files; and
|
| 100 |
+
|
| 101 |
+
(c) You must retain, in the Source form of any Derivative Works
|
| 102 |
+
that You distribute, all copyright, patent, trademark, and
|
| 103 |
+
attribution notices from the Source form of the Work,
|
| 104 |
+
excluding those notices that do not pertain to any part of
|
| 105 |
+
the Derivative Works; and
|
| 106 |
+
|
| 107 |
+
(d) If the Work includes a "NOTICE" text file as part of its
|
| 108 |
+
distribution, then any Derivative Works that You distribute must
|
| 109 |
+
include a readable copy of the attribution notices contained
|
| 110 |
+
within such NOTICE file, excluding those notices that do not
|
| 111 |
+
pertain to any part of the Derivative Works, in at least one
|
| 112 |
+
of the following places: within a NOTICE text file distributed
|
| 113 |
+
as part of the Derivative Works; within the Source form or
|
| 114 |
+
documentation, if provided along with the Derivative Works; or,
|
| 115 |
+
within a display generated by the Derivative Works, if and
|
| 116 |
+
wherever such third-party notices normally appear. The contents
|
| 117 |
+
of the NOTICE file are for informational purposes only and
|
| 118 |
+
do not modify the License. You may add Your own attribution
|
| 119 |
+
notices within Derivative Works that You distribute, alongside
|
| 120 |
+
or as an addendum to the NOTICE text from the Work, provided
|
| 121 |
+
that such additional attribution notices cannot be construed
|
| 122 |
+
as modifying the License.
|
| 123 |
+
|
| 124 |
+
You may add Your own copyright statement to Your modifications and
|
| 125 |
+
may provide additional or different license terms and conditions
|
| 126 |
+
for use, reproduction, or distribution of Your modifications, or
|
| 127 |
+
for any such Derivative Works as a whole, provided Your use,
|
| 128 |
+
reproduction, and distribution of the Work otherwise complies with
|
| 129 |
+
the conditions stated in this License.
|
| 130 |
+
|
| 131 |
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
| 132 |
+
any Contribution intentionally submitted for inclusion in the Work
|
| 133 |
+
by You to the Licensor shall be under the terms and conditions of
|
| 134 |
+
this License, without any additional terms or conditions.
|
| 135 |
+
Notwithstanding the above, nothing herein shall supersede or modify
|
| 136 |
+
the terms of any separate license agreement you may have executed
|
| 137 |
+
with Licensor regarding such Contributions.
|
| 138 |
+
|
| 139 |
+
6. Trademarks. This License does not grant permission to use the trade
|
| 140 |
+
names, trademarks, service marks, or product names of the Licensor,
|
| 141 |
+
except as required for reasonable and customary use in describing the
|
| 142 |
+
origin of the Work and reproducing the content of the NOTICE file.
|
| 143 |
+
|
| 144 |
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
| 145 |
+
agreed to in writing, Licensor provides the Work (and each
|
| 146 |
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
| 147 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
| 148 |
+
implied, including, without limitation, any warranties or conditions
|
| 149 |
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
| 150 |
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
| 151 |
+
appropriateness of using or redistributing the Work and assume any
|
| 152 |
+
risks associated with Your exercise of permissions under this License.
|
| 153 |
+
|
| 154 |
+
8. Limitation of Liability. In no event and under no legal theory,
|
| 155 |
+
whether in tort (including negligence), contract, or otherwise,
|
| 156 |
+
unless required by applicable law (such as deliberate and grossly
|
| 157 |
+
negligent acts) or agreed to in writing, shall any Contributor be
|
| 158 |
+
liable to You for damages, including any direct, indirect, special,
|
| 159 |
+
incidental, or consequential damages of any character arising as a
|
| 160 |
+
result of this License or out of the use or inability to use the
|
| 161 |
+
Work (including but not limited to damages for loss of goodwill,
|
| 162 |
+
work stoppage, computer failure or malfunction, or any and all
|
| 163 |
+
other commercial damages or losses), even if such Contributor
|
| 164 |
+
has been advised of the possibility of such damages.
|
| 165 |
+
|
| 166 |
+
9. Accepting Warranty or Additional Liability. While redistributing
|
| 167 |
+
the Work or Derivative Works thereof, You may choose to offer,
|
| 168 |
+
and charge a fee for, acceptance of support, warranty, indemnity,
|
| 169 |
+
or other liability obligations and/or rights consistent with this
|
| 170 |
+
License. However, in accepting such obligations, You may act only
|
| 171 |
+
on Your own behalf and on Your sole responsibility, not on behalf
|
| 172 |
+
of any other Contributor, and only if You agree to indemnify,
|
| 173 |
+
defend, and hold each Contributor harmless for any liability
|
| 174 |
+
incurred by, or claims asserted against, such Contributor by reason
|
| 175 |
+
of your accepting any such warranty or additional liability.
|
| 176 |
+
|
| 177 |
+
END OF TERMS AND CONDITIONS
|
| 178 |
+
|
| 179 |
+
APPENDIX: How to apply the Apache License to your work.
|
| 180 |
+
|
| 181 |
+
To apply the Apache License to your work, attach the following
|
| 182 |
+
boilerplate notice, with the fields enclosed by brackets "[]"
|
| 183 |
+
replaced with your own identifying information. (Don't include
|
| 184 |
+
the brackets!) The text should be enclosed in the appropriate
|
| 185 |
+
comment syntax for the file format. We also recommend that a
|
| 186 |
+
file or class name and description of purpose be included on the
|
| 187 |
+
same "printed page" as the copyright notice for easier
|
| 188 |
+
identification within third-party archives.
|
| 189 |
+
|
| 190 |
+
Copyright 2026 XHToken
|
| 191 |
+
|
| 192 |
+
Licensed under the Apache License, Version 2.0 (the "License");
|
| 193 |
+
you may not use this file except in compliance with the License.
|
| 194 |
+
You may obtain a copy of the License at
|
| 195 |
+
|
| 196 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 197 |
+
|
| 198 |
+
Unless required by applicable law or agreed to in writing, software
|
| 199 |
+
distributed under the License is distributed on an "AS IS" BASIS,
|
| 200 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 201 |
+
See the License for the specific language governing permissions and
|
| 202 |
+
limitations under the License.
|
|
@@ -0,0 +1,349 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
- zh
|
| 6 |
+
library_name: transformers
|
| 7 |
+
pipeline_tag: text-generation
|
| 8 |
+
tags:
|
| 9 |
+
- llm
|
| 10 |
+
- sparkx2_5
|
| 11 |
+
base_model:
|
| 12 |
+
- XHToken/Spark-X2.5-4B-Base
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
|
| 16 |
+
# Spark-X2.5
|
| 17 |
+
<div align="center">
|
| 18 |
+
|
| 19 |
+
[](https://join.slack.com/t/tokenspark/shared_invite/zt-432qf8l2f-5~dLyXv8uETr0P0UuC07nw)
|
| 20 |
+
[](https://discord.gg/kTDE2Hg8aw)
|
| 21 |
+
[](https://www.youtube.com/@SparkLLM)
|
| 22 |
+
[](https://dev.to/sparkllm)
|
| 23 |
+
[](https://bsky.app/profile/sparkllm.bsky.social)
|
| 24 |
+
[](https://x.com/sparkllm)
|
| 25 |
+
[](https://www.zhihu.com/people/zhiikz7qh7m)
|
| 26 |
+
[](images/xhtoken-wechat.jpg)
|
| 27 |
+
|
| 28 |
+
</div>
|
| 29 |
+
|
| 30 |
+
> [!Note]
|
| 31 |
+
> This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
|
| 32 |
+
|
| 33 |
+
## Introduction
|
| 34 |
+
|
| 35 |
+
We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages.
|
| 36 |
+
|
| 37 |
+
**Technical Highlights**:
|
| 38 |
+
- **Efficient Architecture and Native 1M-token Context**: The models use a hybrid attention architecture that combines one full-attention layer with three sliding-window attention layers. This design substantially reduces the computational overhead typically associated with long-context models while natively supporting a context window of up to 1M tokens.
|
| 39 |
+
- **Strong Coding and Agent Capabilities**: The models are deeply integrated with popular agent harnesses, including Codex, Claude Code, OpenClaw, and Hermes. They deliver state-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction-following tasks.
|
| 40 |
+
- **Broad Hardware and Software Compatibility**: The models support a wide range of hardware platforms, including NVIDIA, Huawei, Hygon, HOUMO.AI, etc. It is compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, MLX, and can be deployed quickly through platforms including Ollama and LM Studio. The models can also be customized using popular fine-tuning frameworks such as LLaMA-Factory. Across multiple hardware platforms, they deliver superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
|
| 41 |
+
- **Advanced Training Algorithms**: The models were trained on Huawei Ascend clusters. Large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance its reasoning, coding, agentic, and instruction-following capabilities.
|
| 42 |
+
|
| 43 |
+
<p align="center">
|
| 44 |
+
<img src="./images/model-benchmark-comparison.svg" alt="Spark-X2.5 benchmark comparison" width="1010">
|
| 45 |
+
</p>
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
## Model Overview
|
| 49 |
+
|
| 50 |
+
For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture. This approach leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size—thereby improving its practicality and effectiveness across real-world deployment scenarios.
|
| 51 |
+
|
| 52 |
+
<p align="center">
|
| 53 |
+
<img src="./images/spark25-hybrid-architecture-light.png" alt="Spark-X2.5 hybrid architecture" width="610">
|
| 54 |
+
</p>
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
## Training Methods
|
| 59 |
+
|
| 60 |
+
Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies are conducted to determine an effective balance among mathematics, logic, code, and other high-value domains. This enables the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.
|
| 61 |
+
|
| 62 |
+
Post-training begins with supervised fine-tuning on a carefully curated corpus. This stage establishes robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning. We subsequently apply large-scale reinforcement learning across several capability domains, including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. This process yields a set of domain-specialized teacher policies, whose complementary strengths are consolidated into a single deployable model through MOPD.
|
| 63 |
+
|
| 64 |
+
<p align="center">
|
| 65 |
+
<img src="./images/post_training_pipeline.svg" alt="Spark-X2.5 hybrid architecture" width="610">
|
| 66 |
+
</p>
|
| 67 |
+
|
| 68 |
+
## Benchmarks
|
| 69 |
+
|
| 70 |
+
We evaluate our models and compare them with leading on-device models of similar size across a broad range of tasks, including agent, code, math, general and knowledge.
|
| 71 |
+
|
| 72 |
+
<div style="overflow-x: auto; width: 100%;">
|
| 73 |
+
<table style="width: 100%; min-width: 1080px; border-collapse: collapse; text-align: center;">
|
| 74 |
+
<colgroup>
|
| 75 |
+
<col style="width: 190px;">
|
| 76 |
+
<col span="8" style="width: 110px;">
|
| 77 |
+
</colgroup>
|
| 78 |
+
<thead>
|
| 79 |
+
<tr>
|
| 80 |
+
<th align="center">Benchmark</th>
|
| 81 |
+
<th align="center"><span style="white-space: nowrap;">Spark‑X2.5‑4B</span></th>
|
| 82 |
+
<th align="center"><span style="white-space: nowrap;">Spark‑X2.5‑1.7B</span></th>
|
| 83 |
+
<th align="center"><span style="white-space: nowrap;">Qwen3.5‑9B</span></th>
|
| 84 |
+
<th align="center"><span style="white-space: nowrap;">Qwen3.5‑4B</span></th>
|
| 85 |
+
<th align="center"><span style="white-space: nowrap;">Qwen3.5‑2B</span></th>
|
| 86 |
+
<th align="center"><span style="white-space: nowrap;">Gemma4‑12B</span></th>
|
| 87 |
+
<th align="center"><span style="white-space: nowrap;">Gemma4‑E4B</span></th>
|
| 88 |
+
<th align="center"><span style="white-space: nowrap;">Gemma4‑E2B</span></th>
|
| 89 |
+
</tr>
|
| 90 |
+
</thead>
|
| 91 |
+
<tbody>
|
| 92 |
+
<tr><th colspan="9" align="left">Agent</th></tr>
|
| 93 |
+
<tr><td align="center">BFCL‑V4</td><td align="center">65.1</td><td align="center">46.9</td><td align="center"><strong>66.1*</strong></td><td align="center">50.3*</td><td align="center">43.6*</td><td align="center">37.4</td><td align="center">36.9</td><td align="center">30.2</td></tr>
|
| 94 |
+
<tr><td align="center">τ²‑bench</td><td align="center">75.1</td><td align="center">65.3</td><td align="center">79.1*</td><td align="center"><strong>79.9*</strong></td><td align="center">48.8*</td><td align="center">69.0*</td><td align="center">42.2*</td><td align="center">24.5*</td></tr>
|
| 95 |
+
<tr><td align="center">τ³‑bench</td><td align="center"><strong>30.4</strong></td><td align="center">20.1</td><td align="center">9.3</td><td align="center">6.7</td><td align="center">4.1</td><td align="center">13.3</td><td align="center">10.1</td><td align="center">8.8</td></tr>
|
| 96 |
+
<tr><td align="center">MCP‑Atlas</td><td align="center"><strong>54.6</strong></td><td align="center">23.4</td><td align="center">47.4*</td><td align="center">40.8*</td><td align="center">14.8</td><td align="center">30.5*</td><td align="center">15.0*</td><td align="center">12.6</td></tr>
|
| 97 |
+
<tr><td align="center">MCP‑Mark</td><td align="center"><strong>14.2</strong></td><td align="center">2.3</td><td align="center">13.4</td><td align="center">12.5</td><td align="center">–</td><td align="center">–</td><td align="center">–</td><td align="center">–</td></tr>
|
| 98 |
+
<tr><td align="center"><span style="white-space: nowrap;">Workspace Bench</span></td><td align="center"><strong>31.2</strong></td><td align="center">18.9</td><td align="center">25.5</td><td align="center">21.3</td><td align="center">7.7</td><td align="center">–</td><td align="center">–</td><td align="center">–</td></tr>
|
| 99 |
+
<tr><td align="center">VitaBench2.0</td><td align="center"><strong>25.2</strong></td><td align="center">8.3</td><td align="center">15.6</td><td align="center">18.2</td><td align="center">5.2</td><td align="center">12.4</td><td align="center">4.8</td><td align="center">4.4</td></tr>
|
| 100 |
+
<tr><td align="center">BrowseComp</td><td align="center"><strong>40.9</strong></td><td align="center">29.7</td><td align="center">8.3</td><td align="center">14.3</td><td align="center">3.1</td><td align="center">10.0</td><td align="center">8.3</td><td align="center">3.7</td></tr>
|
| 101 |
+
<tr><th colspan="9" align="left">Code</th></tr>
|
| 102 |
+
<tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Pro</span></td><td align="center"><strong>44.4</strong></td><td align="center">10.4</td><td align="center">33.8*</td><td align="center">29.4*</td><td align="center">1.9</td><td align="center">21.9*</td><td align="center">4.0*</td><td align="center">–</td></tr>
|
| 103 |
+
<tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Verified</span></td><td align="center">41.6</td><td align="center">28.3</td><td align="center"><strong>53.1*</strong></td><td align="center">38.8*</td><td align="center">6.8</td><td align="center">44.2*</td><td align="center">14.0*</td><td align="center">–</td></tr>
|
| 104 |
+
<tr><td align="center"><span style="white-space: nowrap;">SWE‑Bench Multilingual</span></td><td align="center"><strong>53.3</strong></td><td align="center">23.3</td><td align="center">43.3</td><td align="center">27.7</td><td align="center">5.0</td><td align="center">32.5*</td><td align="center">–</td><td align="center">–</td></tr>
|
| 105 |
+
<tr><td align="center">SciCode</td><td align="center">34.7</td><td align="center">18.2</td><td align="center">32.7*</td><td align="center">24.0</td><td align="center">6.0</td><td align="center"><strong>39.8</strong></td><td align="center">27.5</td><td align="center">20.5</td></tr>
|
| 106 |
+
<tr><th colspan="9" align="left">Math</th></tr>
|
| 107 |
+
<tr><td align="center"><span style="white-space: nowrap;">Gaokao 2026</span></td><td align="center">133.4</td><td align="center">114.8</td><td align="center"><strong>135.5</strong></td><td align="center">130.3</td><td align="center">94.0</td><td align="center">130.6</td><td align="center">102.4</td><td align="center">81.8</td></tr>
|
| 108 |
+
<tr><td align="center"><span style="white-space: nowrap;">AIME 2026</span></td><td align="center"><strong>90.7</strong></td><td align="center">69.4</td><td align="center">88.2</td><td align="center">83.0</td><td align="center">30.8</td><td align="center">82.1*</td><td align="center">42.5*</td><td align="center">37.5*</td></tr>
|
| 109 |
+
<tr><td align="center"><span style="white-space: nowrap;">HMMT Feb 2026</span></td><td align="center"><strong>81.2</strong></td><td align="center">48.4</td><td align="center">70.8</td><td align="center">69.7</td><td align="center">21.5</td><td align="center">65.6</td><td align="center">34.2</td><td align="center">20.5</td></tr>
|
| 110 |
+
<tr><td align="center"><span style="white-space: nowrap;">IMO‑AnswerBench</span></td><td align="center"><strong>74.2</strong></td><td align="center">45.4</td><td align="center">69.8</td><td align="center">68.5</td><td align="center">–</td><td align="center">57.2</td><td align="center">26.9</td><td align="center">22.6</td></tr>
|
| 111 |
+
<tr><th colspan="9" align="left">General & Knowledge</th></tr>
|
| 112 |
+
<tr><td align="center">IFEval</td><td align="center">93.0</td><td align="center">89.5</td><td align="center">91.5*</td><td align="center">89.8*</td><td align="center">78.6*</td><td align="center"><strong>94.8</strong></td><td align="center">45.3</td><td align="center">34.8</td></tr>
|
| 113 |
+
<tr><td align="center">IFBench</td><td align="center"><strong>75.0</strong></td><td align="center">66.3</td><td align="center">64.5</td><td align="center">59.2</td><td align="center">41.3*</td><td align="center">73.5*</td><td align="center">44.0*</td><td align="center">22.7</td></tr>
|
| 114 |
+
<tr><td align="center">AA‑LCR</td><td align="center">56.3</td><td align="center">24.3</td><td align="center"><strong>63.0*</strong></td><td align="center">57.0*</td><td align="center">25.6*</td><td align="center">55.3*</td><td align="center">34.7</td><td align="center">18.3</td></tr>
|
| 115 |
+
<tr><td align="center">HLE</td><td align="center">12.3</td><td align="center">6.3</td><td align="center"><strong>14.3</strong></td><td align="center">8.6</td><td align="center">2.1</td><td align="center">13.1</td><td align="center">3.9</td><td align="center">2.5</td></tr>
|
| 116 |
+
<tr><td align="center">GPQA</td><td align="center">67.4</td><td align="center">43.8</td><td align="center"><strong>77.2</strong></td><td align="center">67.2</td><td align="center">44.6</td><td align="center">72.8</td><td align="center">54.5</td><td align="center">43.8</td></tr>
|
| 117 |
+
</tbody>
|
| 118 |
+
</table>
|
| 119 |
+
</div>
|
| 120 |
+
|
| 121 |
+
- \* denotes reported results from publicly‑released model cards / papers and - denotes scores not yet available.
|
| 122 |
+
- All evaluations are conducted in thinking mode. The recommended sampling parameters for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1.
|
| 123 |
+
- Gaokao 2026 consists of the five 2026 Chinese GAOKAO examinations (National I,National II, Beijing, Shanghai, Tianjin), each graded out of 150 points.
|
| 124 |
+
|
| 125 |
+
## Quickstart
|
| 126 |
+
|
| 127 |
+
### SGLang
|
| 128 |
+
|
| 129 |
+
#### Install SGLang
|
| 130 |
+
|
| 131 |
+
Use the pre-built image that tracks the Spark-X2.5 runtime:
|
| 132 |
+
|
| 133 |
+
```bash
|
| 134 |
+
docker pull lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1
|
| 135 |
+
```
|
| 136 |
+
|
| 137 |
+
#### Run Inference
|
| 138 |
+
|
| 139 |
+
The following command can be used to start an OpenAI-compatible API server on a single GPU with maximum context length 1,048,576 tokens.
|
| 140 |
+
|
| 141 |
+
#### Server
|
| 142 |
+
|
| 143 |
+
```bash
|
| 144 |
+
docker run -it \
|
| 145 |
+
--gpus '"device=0"' \
|
| 146 |
+
--ipc=host \
|
| 147 |
+
-p 30000:30000 \
|
| 148 |
+
-v "$MODEL_PATH":/root/Spark-X2.5-4B \
|
| 149 |
+
lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 \
|
| 150 |
+
python -m sglang.launch_server \
|
| 151 |
+
--model-path /root/Spark-X2.5-4B \
|
| 152 |
+
--served-model-name spark2.5 \
|
| 153 |
+
--tool-call-parser spark25 \
|
| 154 |
+
--reasoning-parser qwen3 \
|
| 155 |
+
--tp-size 1 \
|
| 156 |
+
--mem-fraction-static 0.8 \
|
| 157 |
+
--context-length 1048576 \
|
| 158 |
+
--chat-template /root/Spark-X2.5-4B/chat_template.jinja \
|
| 159 |
+
--host 0.0.0.0 \
|
| 160 |
+
--port 30000
|
| 161 |
+
```
|
| 162 |
+
|
| 163 |
+
#### Client
|
| 164 |
+
|
| 165 |
+
Thinking is enabled by default by both the chat template and the qwen3 reasoning parser. To disable thinking for a specific request, set "chat_template_kwargs": {"enable_thinking": false}.
|
| 166 |
+
|
| 167 |
+
```bash
|
| 168 |
+
curl -s http://localhost:30000/v1/chat/completions \
|
| 169 |
+
-H "Content-Type: application/json" \
|
| 170 |
+
-d '{
|
| 171 |
+
"model": "spark2.5",
|
| 172 |
+
"messages": [
|
| 173 |
+
{
|
| 174 |
+
"role": "user",
|
| 175 |
+
"content": "安徽的省会在哪里?"
|
| 176 |
+
}
|
| 177 |
+
],
|
| 178 |
+
"max_tokens": 131072,
|
| 179 |
+
"temperature": 1,
|
| 180 |
+
"top_k": -1,
|
| 181 |
+
"top_p": 0.95,
|
| 182 |
+
"repetition_penalty": 1,
|
| 183 |
+
"presence_penalty": 0,
|
| 184 |
+
"frequency_penalty": 0
|
| 185 |
+
}'
|
| 186 |
+
```
|
| 187 |
+
|
| 188 |
+
### vLLM
|
| 189 |
+
|
| 190 |
+
#### Install vLLM
|
| 191 |
+
|
| 192 |
+
```bash
|
| 193 |
+
pip install uv
|
| 194 |
+
uv venv ~/spark2_5
|
| 195 |
+
source ~/spark2_5/bin/activate
|
| 196 |
+
git clone https://github.com/XHToken/Spark-plugin.git
|
| 197 |
+
cd ./Spark-plugin
|
| 198 |
+
uv pip install .
|
| 199 |
+
```
|
| 200 |
+
|
| 201 |
+
#### Server
|
| 202 |
+
|
| 203 |
+
```bash
|
| 204 |
+
vllm serve "./Spark-X2.5-4B" \
|
| 205 |
+
--port "30000" \
|
| 206 |
+
--trust-remote-code \
|
| 207 |
+
--served-model-name spark25 \
|
| 208 |
+
--tensor-parallel-size 1 \
|
| 209 |
+
--gpu-memory-utilization 0.7 \
|
| 210 |
+
--enable-prefix-caching \
|
| 211 |
+
--chat-template Spark-X2.5-4B/chat_template.jinja
|
| 212 |
+
```
|
| 213 |
+
|
| 214 |
+
#### Client
|
| 215 |
+
|
| 216 |
+
```bash
|
| 217 |
+
curl -s http://127.0.0.1:30000/v1/chat/completions \
|
| 218 |
+
-H "Content-Type: application/json" \
|
| 219 |
+
-d '{
|
| 220 |
+
"model": "spark25",
|
| 221 |
+
"messages": [{"role": "user", "content": "安徽的省会在哪里?"}],
|
| 222 |
+
"temperature": 1.0,
|
| 223 |
+
"top_k": -1,
|
| 224 |
+
"top_p": 0.95
|
| 225 |
+
}'
|
| 226 |
+
```
|
| 227 |
+
|
| 228 |
+
### MLX
|
| 229 |
+
Spark-MLX-LLM runs the original Spark-X2.5 Hugging Face checkpoints locally. It supports Apple silicon GPU, Linux CPU, and NVIDIA CUDA on Linux. No GGUF conversion is required.
|
| 230 |
+
|
| 231 |
+
#### Installation
|
| 232 |
+
|
| 233 |
+
```bash
|
| 234 |
+
git clone https://github.com/XHToken/Spark-MLX-LLM.git
|
| 235 |
+
cd Spark-MLX-LLM
|
| 236 |
+
|
| 237 |
+
python3 -m venv .venv
|
| 238 |
+
source .venv/bin/activate
|
| 239 |
+
|
| 240 |
+
# Apple silicon
|
| 241 |
+
python -m pip install -e .
|
| 242 |
+
# Linux cpu
|
| 243 |
+
python -m pip install -e '.[cpu]'
|
| 244 |
+
# Linux with cuda12
|
| 245 |
+
python -m pip install -e '.[cuda12]'
|
| 246 |
+
# Linux with cuda13
|
| 247 |
+
python -m pip install -e '.[cuda13]'
|
| 248 |
+
```
|
| 249 |
+
|
| 250 |
+
Run Spark-X2.5
|
| 251 |
+
|
| 252 |
+
```bash
|
| 253 |
+
spark-mlx-generate \
|
| 254 |
+
--device gpu \
|
| 255 |
+
--dtype bfloat16 \
|
| 256 |
+
--model XHToken/Spark-X2.5-4B \
|
| 257 |
+
--prompt "安徽的省会在哪里?" \
|
| 258 |
+
--max-tokens 512 \
|
| 259 |
+
--temp 0
|
| 260 |
+
```
|
| 261 |
+
|
| 262 |
+
### Ollama
|
| 263 |
+
|
| 264 |
+
#### Build
|
| 265 |
+
|
| 266 |
+
```bash
|
| 267 |
+
git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
|
| 268 |
+
git clone https://github.com/ollama/ollama.git ollama-spark
|
| 269 |
+
cd ollama-spark
|
| 270 |
+
export OLLAMA_LLAMA_CPP_SOURCE="$(cd ../llama.cpp-spark && pwd)"
|
| 271 |
+
cmake -S . -B build
|
| 272 |
+
cmake --build build --parallel 8
|
| 273 |
+
```
|
| 274 |
+
|
| 275 |
+
#### Create and Run
|
| 276 |
+
|
| 277 |
+
```bash
|
| 278 |
+
printf 'FROM /absolute/path/to/your.gguf\n' > ./Modelfile.spark
|
| 279 |
+
./ollama serve
|
| 280 |
+
./ollama create Spark-X2.5-4B -f ./Modelfile.spark
|
| 281 |
+
./ollama run Spark-X2.5-4B
|
| 282 |
+
```
|
| 283 |
+
|
| 284 |
+
### LM Studio
|
| 285 |
+
|
| 286 |
+
#### Build
|
| 287 |
+
|
| 288 |
+
```bash
|
| 289 |
+
git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
|
| 290 |
+
cd llama.cpp-spark
|
| 291 |
+
cmake -S . -B build
|
| 292 |
+
cmake --build build --parallel 8
|
| 293 |
+
```
|
| 294 |
+
|
| 295 |
+
#### Set Up LM Studio
|
| 296 |
+
|
| 297 |
+
1. Close LM Studio.
|
| 298 |
+
|
| 299 |
+
2. Back up the selected runtime directory:
|
| 300 |
+
|
| 301 |
+
```text
|
| 302 |
+
<LM_STUDIO_HOME>/extensions/backends/<selected-runtime>/
|
| 303 |
+
```
|
| 304 |
+
|
| 305 |
+
3. Copy the `llama.cpp-spark` build output into the selected runtime directory, overwriting the existing files.
|
| 306 |
+
|
| 307 |
+
4. Place the GGUF model in the following directory:
|
| 308 |
+
|
| 309 |
+
```text
|
| 310 |
+
<LM_STUDIO_HOME>/models/<org>/<name>/
|
| 311 |
+
```
|
| 312 |
+
|
| 313 |
+
Example runtime directory on macOS:
|
| 314 |
+
|
| 315 |
+
```text
|
| 316 |
+
./build/bin/* -> ~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-<version>/
|
| 317 |
+
```
|
| 318 |
+
|
| 319 |
+
#### Run with LM Studio
|
| 320 |
+
|
| 321 |
+
Open My Models, select the Spark-X2.5 model, click Load, then start a new Chat.
|
| 322 |
+
|
| 323 |
+
#### Run with lms cli
|
| 324 |
+
|
| 325 |
+
```bash
|
| 326 |
+
# Replace `<model>` with a model listed by `lms ls`
|
| 327 |
+
lms load <model>
|
| 328 |
+
lms chat <model>
|
| 329 |
+
```
|
| 330 |
+
|
| 331 |
+
### Finetuning
|
| 332 |
+
|
| 333 |
+
We advise you to use [Llama-Factory](https://github.com/XHToken/LlamaFactory) to finetune your models.
|
| 334 |
+
|
| 335 |
+
|
| 336 |
+
## License
|
| 337 |
+
|
| 338 |
+
The Spark-X2.5 model series is licensed under the [Apache 2.0 License](https://huggingface.co/XHToken/Spark-X2.5-4B/blob/main/LICENSE).
|
| 339 |
+
|
| 340 |
+
## Citation
|
| 341 |
+
If you find our work helpful, feel free to give us a cite.
|
| 342 |
+
|
| 343 |
+
```bibtex
|
| 344 |
+
@misc{sparkx2.5,
|
| 345 |
+
title = {Spark-X2.5 4B&1.7B: Pushing the Limits of Agentic Capabilities in On-Device Models},
|
| 346 |
+
author = {SparkLLM Team},
|
| 347 |
+
year = {2026}
|
| 348 |
+
}
|
| 349 |
+
```
|
|
@@ -0,0 +1,110 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{#- 0826版本 -#}
|
| 2 |
+
{%- if not messages %}
|
| 3 |
+
{{- raise_exception('No messages provided.') }}
|
| 4 |
+
{%- endif %}
|
| 5 |
+
{%- set enable_thinking = enable_thinking | default(true) %}
|
| 6 |
+
|
| 7 |
+
{#- Render a string or a list of text blocks. -#}
|
| 8 |
+
{%- macro render_content(content, context_name) %}
|
| 9 |
+
{%- if content is string %}
|
| 10 |
+
{{- content }}
|
| 11 |
+
{%- elif content is none or content is undefined %}
|
| 12 |
+
{{- '' }}
|
| 13 |
+
{%- elif content is iterable and content is not mapping %}
|
| 14 |
+
{%- for block in content %}
|
| 15 |
+
{%- if block.type == 'text' %}
|
| 16 |
+
{{- block.text }}
|
| 17 |
+
{%- else %}
|
| 18 |
+
{{- raise_exception('Unsupported ' ~ context_name ~ ' content block type: ' ~ (block.type | string)) }}
|
| 19 |
+
{%- endif %}
|
| 20 |
+
{%- endfor %}
|
| 21 |
+
{%- else %}
|
| 22 |
+
{{- raise_exception(context_name ~ ' content must be a string or a list of text blocks') }}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endmacro %}
|
| 25 |
+
|
| 26 |
+
{#- Default system prompt -#}
|
| 27 |
+
{%- set default_system = "you are a helpful assistant." %}
|
| 28 |
+
|
| 29 |
+
{#- The first message-level system is placed in the initial system block. -#}
|
| 30 |
+
{%- set ns = namespace(initial_system='') %}
|
| 31 |
+
{%- if messages[0].role == "system" %}
|
| 32 |
+
{%- set ns.initial_system = render_content(messages[0].content, 'system') %}
|
| 33 |
+
{%- endif %}
|
| 34 |
+
|
| 35 |
+
{#- System block -#}
|
| 36 |
+
{{- '<|start▁of▁sentence|><|System|>' + '\n' + default_system }}
|
| 37 |
+
{%- if tools %}
|
| 38 |
+
{{- '## Tools' + '\n' + 'You have access to the following functions:' + '\n' + '<tools>' }}
|
| 39 |
+
{%- for tool in tools %}
|
| 40 |
+
{{- '\n' + tool.function | tojson}}
|
| 41 |
+
{%- endfor %}
|
| 42 |
+
{{- '\n' + '</tools>' }}
|
| 43 |
+
{%- endif %}
|
| 44 |
+
{%- if ns.initial_system %}
|
| 45 |
+
{{- '\n\n' + ns.initial_system }}
|
| 46 |
+
{%- endif %}
|
| 47 |
+
{{- '<|end▁of▁sentence|>'}}
|
| 48 |
+
|
| 49 |
+
{#- Conversation turns -#}
|
| 50 |
+
{%- for message in messages %}
|
| 51 |
+
{%- if message.role == "system" %}
|
| 52 |
+
{#- The first system message was consumed by the initial block. -#}
|
| 53 |
+
{%- if not loop.first %}
|
| 54 |
+
{{- '<|start▁of▁sentence|><|System|>\n' + render_content(message.content, 'system') + '<|end▁of▁sentence|>' }}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{%- elif message.role == "user" %}
|
| 57 |
+
{{- '<|start▁of▁sentence|><|User|>' + render_content(message.content, 'user') + '<|end▁of▁sentence|>' }}
|
| 58 |
+
{%- elif message.role == "assistant" %}
|
| 59 |
+
{%- set assistant_content = render_content(message.content, 'assistant') %}
|
| 60 |
+
{%- if message.reasoning_content is defined and message.reasoning_content %}
|
| 61 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 62 |
+
{%- else %}
|
| 63 |
+
{%- set reasoning_content = '' %}
|
| 64 |
+
{%- endif %}
|
| 65 |
+
{{- '<|start▁of▁sentence|><|Bot|>'}}
|
| 66 |
+
{%- if reasoning_content %}
|
| 67 |
+
{{- '<think>' + reasoning_content + '</think>'}}
|
| 68 |
+
{%- else %}
|
| 69 |
+
{{- '</think>' }}
|
| 70 |
+
{%- endif %}
|
| 71 |
+
{%- if assistant_content %}
|
| 72 |
+
{{- assistant_content }}
|
| 73 |
+
{%- endif %}
|
| 74 |
+
{%- if message.tool_calls is defined and message.tool_calls is not none %}
|
| 75 |
+
{%- for tool_call in message.tool_calls %}
|
| 76 |
+
{%- if tool_call.function.arguments is not mapping %}
|
| 77 |
+
{{- raise_exception('tool_call.function.arguments must be a dictionary; normalize JSON strings before apply_chat_template') }}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- set args = tool_call.function.arguments %}
|
| 80 |
+
{{- '<tool_call>' + tool_call.function.name }}
|
| 81 |
+
{%- for k, v in args.items() %}
|
| 82 |
+
{{- '<arg_key>' ~ k ~ '</arg_key><arg_value>' ~ (v if v is string else v | tojson) ~ '</arg_value>' }}
|
| 83 |
+
{%- endfor %}
|
| 84 |
+
{{- '</tool_call>' }}
|
| 85 |
+
{%- endfor %}
|
| 86 |
+
{%- endif %}
|
| 87 |
+
{{- '<|end▁of▁sentence|>' }}
|
| 88 |
+
{%- elif message.role == "tool" %}
|
| 89 |
+
{%- if loop.previtem is undefined or loop.previtem.role != "tool" %}
|
| 90 |
+
{{- '<|start▁of▁sentence|><|Tool|>' }}
|
| 91 |
+
{%- endif %}
|
| 92 |
+
{{- '<tool_response>' ~ message.content ~ '</tool_response>' }}
|
| 93 |
+
{%- if loop.nextitem is undefined or loop.nextitem.role != "tool" %}
|
| 94 |
+
{{- '<|end▁of▁sentence|>' }}
|
| 95 |
+
{%- endif %}
|
| 96 |
+
{%- else %}
|
| 97 |
+
{{- raise_exception('Unsupported message role: ' ~ message.role) }}
|
| 98 |
+
{%- endif %}
|
| 99 |
+
{%- endfor %}
|
| 100 |
+
|
| 101 |
+
{#- Generation prompt -#}
|
| 102 |
+
{%- if add_generation_prompt %}
|
| 103 |
+
{{- '<|start▁of▁sentence|><|Bot|>' }}
|
| 104 |
+
{%- if enable_thinking is defined and enable_thinking %}
|
| 105 |
+
{{- '<think>' }}
|
| 106 |
+
{%- endif %}
|
| 107 |
+
{%- if enable_thinking is defined and not enable_thinking %}
|
| 108 |
+
{{- '</think>' }}
|
| 109 |
+
{%- endif %}
|
| 110 |
+
{%- endif %}
|
|
@@ -0,0 +1,83 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Spark2_5ForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"attention_bias": false,
|
| 6 |
+
"attention_dropout": 0.0,
|
| 7 |
+
"auto_map": {
|
| 8 |
+
"AutoConfig": "configuration_spark.Spark2_5Config",
|
| 9 |
+
"AutoModel": "modeling_spark.Spark2_5Model",
|
| 10 |
+
"AutoModelForCausalLM": "modeling_spark.Spark2_5ForCausalLM"
|
| 11 |
+
},
|
| 12 |
+
"bos_token_id": 0,
|
| 13 |
+
"dtype": "bfloat16",
|
| 14 |
+
"eos_token_id": 1,
|
| 15 |
+
"gate_attn_act_mode": "sigmoid",
|
| 16 |
+
"head_dim": 256,
|
| 17 |
+
"headwise_attn_output_gate": true,
|
| 18 |
+
"hidden_act": "gelu",
|
| 19 |
+
"hidden_size": 2560,
|
| 20 |
+
"initializer_range": 0.01976,
|
| 21 |
+
"intermediate_size": 10240,
|
| 22 |
+
"layer_types": [
|
| 23 |
+
"sliding_attention",
|
| 24 |
+
"sliding_attention",
|
| 25 |
+
"sliding_attention",
|
| 26 |
+
"full_attention",
|
| 27 |
+
"sliding_attention",
|
| 28 |
+
"sliding_attention",
|
| 29 |
+
"sliding_attention",
|
| 30 |
+
"full_attention",
|
| 31 |
+
"sliding_attention",
|
| 32 |
+
"sliding_attention",
|
| 33 |
+
"sliding_attention",
|
| 34 |
+
"full_attention",
|
| 35 |
+
"sliding_attention",
|
| 36 |
+
"sliding_attention",
|
| 37 |
+
"sliding_attention",
|
| 38 |
+
"full_attention",
|
| 39 |
+
"sliding_attention",
|
| 40 |
+
"sliding_attention",
|
| 41 |
+
"sliding_attention",
|
| 42 |
+
"full_attention",
|
| 43 |
+
"sliding_attention",
|
| 44 |
+
"sliding_attention",
|
| 45 |
+
"sliding_attention",
|
| 46 |
+
"full_attention",
|
| 47 |
+
"sliding_attention",
|
| 48 |
+
"sliding_attention",
|
| 49 |
+
"sliding_attention",
|
| 50 |
+
"full_attention",
|
| 51 |
+
"sliding_attention",
|
| 52 |
+
"sliding_attention",
|
| 53 |
+
"sliding_attention",
|
| 54 |
+
"full_attention",
|
| 55 |
+
"sliding_attention",
|
| 56 |
+
"sliding_attention",
|
| 57 |
+
"sliding_attention",
|
| 58 |
+
"full_attention"
|
| 59 |
+
],
|
| 60 |
+
"max_position_embeddings": 1048576,
|
| 61 |
+
"mlp_bias": false,
|
| 62 |
+
"model_type": "spark2_5",
|
| 63 |
+
"num_attention_heads": 16,
|
| 64 |
+
"num_hidden_layers": 36,
|
| 65 |
+
"num_key_value_heads": 4,
|
| 66 |
+
"pad_token_id": 2,
|
| 67 |
+
"rms_norm_eps": 1e-06,
|
| 68 |
+
"rope_parameters": {
|
| 69 |
+
"full_attention": {
|
| 70 |
+
"partial_rotary_factor": 0.25,
|
| 71 |
+
"rope_theta": 5000000
|
| 72 |
+
},
|
| 73 |
+
"sliding_attention": {
|
| 74 |
+
"partial_rotary_factor": 1.0,
|
| 75 |
+
"rope_theta": 10000
|
| 76 |
+
}
|
| 77 |
+
},
|
| 78 |
+
"sliding_window": 512,
|
| 79 |
+
"tie_word_embeddings": true,
|
| 80 |
+
"transformers_version": "4.57.1",
|
| 81 |
+
"use_cache": true,
|
| 82 |
+
"vocab_size": 131072
|
| 83 |
+
}
|
|
@@ -0,0 +1,118 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# coding=utf-8
|
| 2 |
+
# Copyright 2026 The XHToken team and the HuggingFace Inc. team. All rights reserved.
|
| 3 |
+
#
|
| 4 |
+
# Licensed under the Apache License, Version 2.0 (the "License");
|
| 5 |
+
# you may not use this file except in compliance with the License.
|
| 6 |
+
# You may obtain a copy of the License at
|
| 7 |
+
#
|
| 8 |
+
# http://www.apache.org/licenses/LICENSE-2.0
|
| 9 |
+
#
|
| 10 |
+
# Unless required by applicable law or agreed to in writing, software
|
| 11 |
+
# distributed under the License is distributed on an "AS IS" BASIS,
|
| 12 |
+
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 13 |
+
# See the License for the specific language governing permissions and
|
| 14 |
+
# limitations under the License.
|
| 15 |
+
|
| 16 |
+
from transformers import PretrainedConfig
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
class Spark2_5Config(PretrainedConfig):
|
| 20 |
+
model_type = "spark2_5"
|
| 21 |
+
keys_to_ignore_at_inference = ["past_key_values"]
|
| 22 |
+
|
| 23 |
+
base_model_tp_plan = {
|
| 24 |
+
"layers.*.self_attn.q_k_v_proj": "colwise",
|
| 25 |
+
"layers.*.self_attn.g_proj": "colwise",
|
| 26 |
+
"layers.*.self_attn.out_proj": "rowwise",
|
| 27 |
+
"layers.*.mlp.gate_proj": "colwise",
|
| 28 |
+
"layers.*.mlp.up_proj": "colwise",
|
| 29 |
+
"layers.*.mlp.down_proj": "rowwise",
|
| 30 |
+
}
|
| 31 |
+
base_model_pp_plan = {
|
| 32 |
+
"embedding": (["input_ids"], ["inputs_embeds"]),
|
| 33 |
+
"layers": (["hidden_states", "attention_mask"], ["hidden_states"]),
|
| 34 |
+
"norm": (["hidden_states"], ["hidden_states"]),
|
| 35 |
+
}
|
| 36 |
+
|
| 37 |
+
def __init__(
|
| 38 |
+
self,
|
| 39 |
+
vocab_size=32000,
|
| 40 |
+
hidden_size=4096,
|
| 41 |
+
intermediate_size=11008,
|
| 42 |
+
num_hidden_layers=32,
|
| 43 |
+
num_attention_heads=32,
|
| 44 |
+
num_key_value_heads=None,
|
| 45 |
+
hidden_act="gelu",
|
| 46 |
+
max_position_embeddings=2048,
|
| 47 |
+
initializer_range=0.02,
|
| 48 |
+
rms_norm_eps=1e-6,
|
| 49 |
+
use_cache=True,
|
| 50 |
+
pad_token_id=None,
|
| 51 |
+
bos_token_id=1,
|
| 52 |
+
eos_token_id=2,
|
| 53 |
+
tie_word_embeddings=False,
|
| 54 |
+
rope_parameters=None,
|
| 55 |
+
attention_bias=False,
|
| 56 |
+
attention_dropout=0.0,
|
| 57 |
+
mlp_bias=False,
|
| 58 |
+
head_dim=None,
|
| 59 |
+
headwise_attn_output_gate=False,
|
| 60 |
+
gate_attn_act_mode="sigmoid",
|
| 61 |
+
sliding_window=None,
|
| 62 |
+
layer_types=None,
|
| 63 |
+
**kwargs,
|
| 64 |
+
):
|
| 65 |
+
self.vocab_size = vocab_size
|
| 66 |
+
self.max_position_embeddings = max_position_embeddings
|
| 67 |
+
self.hidden_size = hidden_size
|
| 68 |
+
self.intermediate_size = intermediate_size
|
| 69 |
+
self.num_hidden_layers = num_hidden_layers
|
| 70 |
+
self.num_attention_heads = num_attention_heads
|
| 71 |
+
|
| 72 |
+
if num_key_value_heads is None:
|
| 73 |
+
num_key_value_heads = num_attention_heads
|
| 74 |
+
if num_attention_heads % num_key_value_heads != 0:
|
| 75 |
+
raise ValueError(
|
| 76 |
+
f"num_attention_heads ({num_attention_heads}) must be divisible by num_key_value_heads ({num_key_value_heads})"
|
| 77 |
+
)
|
| 78 |
+
self.num_key_value_heads = num_key_value_heads
|
| 79 |
+
|
| 80 |
+
self.hidden_act = hidden_act
|
| 81 |
+
self.initializer_range = initializer_range
|
| 82 |
+
self.rms_norm_eps = rms_norm_eps
|
| 83 |
+
self.use_cache = use_cache
|
| 84 |
+
self.attention_bias = attention_bias
|
| 85 |
+
self.attention_dropout = attention_dropout
|
| 86 |
+
self.mlp_bias = mlp_bias
|
| 87 |
+
self.head_dim = head_dim if head_dim is not None else self.hidden_size // self.num_attention_heads
|
| 88 |
+
self.headwise_attn_output_gate = headwise_attn_output_gate
|
| 89 |
+
self.gate_attn_act_mode = gate_attn_act_mode
|
| 90 |
+
self.sliding_window = sliding_window
|
| 91 |
+
self.rope_parameters = rope_parameters
|
| 92 |
+
|
| 93 |
+
if layer_types is None:
|
| 94 |
+
layer_types = ["full_attention"] * num_hidden_layers
|
| 95 |
+
if len(layer_types) != num_hidden_layers:
|
| 96 |
+
raise ValueError(
|
| 97 |
+
f"layer_types length ({len(layer_types)}) must match num_hidden_layers ({num_hidden_layers})"
|
| 98 |
+
)
|
| 99 |
+
self.layer_types = layer_types
|
| 100 |
+
|
| 101 |
+
super().__init__(
|
| 102 |
+
pad_token_id=pad_token_id,
|
| 103 |
+
bos_token_id=bos_token_id,
|
| 104 |
+
eos_token_id=eos_token_id,
|
| 105 |
+
tie_word_embeddings=tie_word_embeddings,
|
| 106 |
+
**kwargs,
|
| 107 |
+
)
|
| 108 |
+
|
| 109 |
+
def get_rope_theta(self, layer_type):
|
| 110 |
+
params = self.rope_parameters.get(layer_type, {})
|
| 111 |
+
return params.get("rope_theta", 10000)
|
| 112 |
+
|
| 113 |
+
def get_partial_rotary_factor(self, layer_type):
|
| 114 |
+
params = self.rope_parameters.get(layer_type, {})
|
| 115 |
+
return params.get("partial_rotary_factor", 1.0)
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
__all__ = ["Spark2_5Config"]
|
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 0,
|
| 3 |
+
"eos_token_id": 1,
|
| 4 |
+
"pad_token_id": 2,
|
| 5 |
+
"max_tokens": 1048576,
|
| 6 |
+
"temperature": 1.0,
|
| 7 |
+
"top_p": 0.95,
|
| 8 |
+
"top_k": -1,
|
| 9 |
+
"repetition_penalty": 1.0,
|
| 10 |
+
"presence_penalty":0,
|
| 11 |
+
"frequency_penalty":0,
|
| 12 |
+
"do_sample": true,
|
| 13 |
+
"transformers_version": "4.57.1"
|
| 14 |
+
}
|
|
|
|
|
|
Git LFS Details
|
|
Git LFS Details
|
|
The diff for this file is too large to render.
See raw diff
|
|
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:cc36bfe9ca2b839df89fd8c0c94f566f6717bd07ae859dfad074b303f90bdf5b
|
| 3 |
+
size 1982449560
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d98ea6b9059433181042c1ff9fbec167ae2f0850029787c82520d04a04942952
|
| 3 |
+
size 1993132424
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0ba22891378cf0ef3d9bf1f73c3159df3fc3dab8aa6e17679a9ffe7e9d867526
|
| 3 |
+
size 1993225080
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6aae0314fdf8fc3801677c017f296bcb14bc16095286c3bd1e7ce27e0398a113
|
| 3 |
+
size 1993132448
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ffe48ea1c2ee668407dd6ded0c02529da9c54580a8c38f6e53cf902b85a1bd95
|
| 3 |
+
size 262252896
|
|
@@ -0,0 +1,298 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"metadata": {
|
| 3 |
+
"total_parameters": 4112079360,
|
| 4 |
+
"total_size": 8224158720
|
| 5 |
+
},
|
| 6 |
+
"weight_map": {
|
| 7 |
+
"model.embedding.weight": "model-00001-of-00005.safetensors",
|
| 8 |
+
"model.layers.0.input_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 9 |
+
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
|
| 10 |
+
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
|
| 11 |
+
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
|
| 12 |
+
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 13 |
+
"model.layers.0.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
|
| 14 |
+
"model.layers.0.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
|
| 15 |
+
"model.layers.0.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
|
| 16 |
+
"model.layers.1.input_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 17 |
+
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
|
| 18 |
+
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
|
| 19 |
+
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
|
| 20 |
+
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 21 |
+
"model.layers.1.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
|
| 22 |
+
"model.layers.1.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
|
| 23 |
+
"model.layers.1.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
|
| 24 |
+
"model.layers.10.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 25 |
+
"model.layers.10.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 26 |
+
"model.layers.10.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 27 |
+
"model.layers.10.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 28 |
+
"model.layers.10.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 29 |
+
"model.layers.10.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 30 |
+
"model.layers.10.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 31 |
+
"model.layers.10.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 32 |
+
"model.layers.11.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 33 |
+
"model.layers.11.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 34 |
+
"model.layers.11.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 35 |
+
"model.layers.11.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 36 |
+
"model.layers.11.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 37 |
+
"model.layers.11.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 38 |
+
"model.layers.11.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 39 |
+
"model.layers.11.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 40 |
+
"model.layers.12.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 41 |
+
"model.layers.12.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 42 |
+
"model.layers.12.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 43 |
+
"model.layers.12.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 44 |
+
"model.layers.12.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 45 |
+
"model.layers.12.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 46 |
+
"model.layers.12.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 47 |
+
"model.layers.12.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 48 |
+
"model.layers.13.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 49 |
+
"model.layers.13.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 50 |
+
"model.layers.13.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 51 |
+
"model.layers.13.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 52 |
+
"model.layers.13.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 53 |
+
"model.layers.13.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 54 |
+
"model.layers.13.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 55 |
+
"model.layers.13.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 56 |
+
"model.layers.14.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 57 |
+
"model.layers.14.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 58 |
+
"model.layers.14.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 59 |
+
"model.layers.14.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 60 |
+
"model.layers.14.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 61 |
+
"model.layers.14.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 62 |
+
"model.layers.14.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 63 |
+
"model.layers.14.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 64 |
+
"model.layers.15.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 65 |
+
"model.layers.15.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 66 |
+
"model.layers.15.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 67 |
+
"model.layers.15.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 68 |
+
"model.layers.15.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 69 |
+
"model.layers.15.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 70 |
+
"model.layers.15.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 71 |
+
"model.layers.15.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 72 |
+
"model.layers.16.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 73 |
+
"model.layers.16.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 74 |
+
"model.layers.16.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 75 |
+
"model.layers.16.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 76 |
+
"model.layers.16.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 77 |
+
"model.layers.16.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 78 |
+
"model.layers.16.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 79 |
+
"model.layers.16.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 80 |
+
"model.layers.17.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 81 |
+
"model.layers.17.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 82 |
+
"model.layers.17.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 83 |
+
"model.layers.17.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 84 |
+
"model.layers.17.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 85 |
+
"model.layers.17.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 86 |
+
"model.layers.17.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 87 |
+
"model.layers.17.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 88 |
+
"model.layers.18.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 89 |
+
"model.layers.18.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 90 |
+
"model.layers.18.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 91 |
+
"model.layers.18.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 92 |
+
"model.layers.18.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 93 |
+
"model.layers.18.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 94 |
+
"model.layers.18.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 95 |
+
"model.layers.18.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 96 |
+
"model.layers.19.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 97 |
+
"model.layers.19.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 98 |
+
"model.layers.19.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 99 |
+
"model.layers.19.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 100 |
+
"model.layers.19.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 101 |
+
"model.layers.19.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 102 |
+
"model.layers.19.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 103 |
+
"model.layers.19.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 104 |
+
"model.layers.2.input_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 105 |
+
"model.layers.2.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
|
| 106 |
+
"model.layers.2.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
|
| 107 |
+
"model.layers.2.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
|
| 108 |
+
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 109 |
+
"model.layers.2.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
|
| 110 |
+
"model.layers.2.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
|
| 111 |
+
"model.layers.2.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
|
| 112 |
+
"model.layers.20.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 113 |
+
"model.layers.20.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 114 |
+
"model.layers.20.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 115 |
+
"model.layers.20.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 116 |
+
"model.layers.20.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 117 |
+
"model.layers.20.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 118 |
+
"model.layers.20.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 119 |
+
"model.layers.20.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 120 |
+
"model.layers.21.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 121 |
+
"model.layers.21.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 122 |
+
"model.layers.21.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 123 |
+
"model.layers.21.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 124 |
+
"model.layers.21.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 125 |
+
"model.layers.21.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 126 |
+
"model.layers.21.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 127 |
+
"model.layers.21.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 128 |
+
"model.layers.22.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 129 |
+
"model.layers.22.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 130 |
+
"model.layers.22.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 131 |
+
"model.layers.22.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 132 |
+
"model.layers.22.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 133 |
+
"model.layers.22.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 134 |
+
"model.layers.22.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 135 |
+
"model.layers.22.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 136 |
+
"model.layers.23.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 137 |
+
"model.layers.23.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 138 |
+
"model.layers.23.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 139 |
+
"model.layers.23.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 140 |
+
"model.layers.23.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 141 |
+
"model.layers.23.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 142 |
+
"model.layers.23.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 143 |
+
"model.layers.23.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 144 |
+
"model.layers.24.input_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 145 |
+
"model.layers.24.mlp.down_proj.weight": "model-00003-of-00005.safetensors",
|
| 146 |
+
"model.layers.24.mlp.gate_proj.weight": "model-00003-of-00005.safetensors",
|
| 147 |
+
"model.layers.24.mlp.up_proj.weight": "model-00003-of-00005.safetensors",
|
| 148 |
+
"model.layers.24.post_attention_layernorm.weight": "model-00003-of-00005.safetensors",
|
| 149 |
+
"model.layers.24.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 150 |
+
"model.layers.24.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 151 |
+
"model.layers.24.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 152 |
+
"model.layers.25.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 153 |
+
"model.layers.25.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 154 |
+
"model.layers.25.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 155 |
+
"model.layers.25.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 156 |
+
"model.layers.25.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 157 |
+
"model.layers.25.self_attn.g_proj.weight": "model-00003-of-00005.safetensors",
|
| 158 |
+
"model.layers.25.self_attn.out_proj.weight": "model-00003-of-00005.safetensors",
|
| 159 |
+
"model.layers.25.self_attn.q_k_v_proj.weight": "model-00003-of-00005.safetensors",
|
| 160 |
+
"model.layers.26.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 161 |
+
"model.layers.26.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 162 |
+
"model.layers.26.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 163 |
+
"model.layers.26.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 164 |
+
"model.layers.26.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 165 |
+
"model.layers.26.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 166 |
+
"model.layers.26.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 167 |
+
"model.layers.26.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 168 |
+
"model.layers.27.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 169 |
+
"model.layers.27.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 170 |
+
"model.layers.27.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 171 |
+
"model.layers.27.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 172 |
+
"model.layers.27.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 173 |
+
"model.layers.27.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 174 |
+
"model.layers.27.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 175 |
+
"model.layers.27.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 176 |
+
"model.layers.28.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 177 |
+
"model.layers.28.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 178 |
+
"model.layers.28.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 179 |
+
"model.layers.28.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 180 |
+
"model.layers.28.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 181 |
+
"model.layers.28.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 182 |
+
"model.layers.28.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 183 |
+
"model.layers.28.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 184 |
+
"model.layers.29.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 185 |
+
"model.layers.29.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 186 |
+
"model.layers.29.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 187 |
+
"model.layers.29.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 188 |
+
"model.layers.29.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 189 |
+
"model.layers.29.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 190 |
+
"model.layers.29.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 191 |
+
"model.layers.29.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 192 |
+
"model.layers.3.input_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 193 |
+
"model.layers.3.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
|
| 194 |
+
"model.layers.3.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
|
| 195 |
+
"model.layers.3.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
|
| 196 |
+
"model.layers.3.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 197 |
+
"model.layers.3.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
|
| 198 |
+
"model.layers.3.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
|
| 199 |
+
"model.layers.3.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
|
| 200 |
+
"model.layers.30.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 201 |
+
"model.layers.30.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 202 |
+
"model.layers.30.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 203 |
+
"model.layers.30.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 204 |
+
"model.layers.30.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 205 |
+
"model.layers.30.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 206 |
+
"model.layers.30.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 207 |
+
"model.layers.30.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 208 |
+
"model.layers.31.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 209 |
+
"model.layers.31.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 210 |
+
"model.layers.31.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 211 |
+
"model.layers.31.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 212 |
+
"model.layers.31.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 213 |
+
"model.layers.31.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 214 |
+
"model.layers.31.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 215 |
+
"model.layers.31.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 216 |
+
"model.layers.32.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 217 |
+
"model.layers.32.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 218 |
+
"model.layers.32.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 219 |
+
"model.layers.32.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 220 |
+
"model.layers.32.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 221 |
+
"model.layers.32.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 222 |
+
"model.layers.32.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 223 |
+
"model.layers.32.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 224 |
+
"model.layers.33.input_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 225 |
+
"model.layers.33.mlp.down_proj.weight": "model-00004-of-00005.safetensors",
|
| 226 |
+
"model.layers.33.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 227 |
+
"model.layers.33.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 228 |
+
"model.layers.33.post_attention_layernorm.weight": "model-00004-of-00005.safetensors",
|
| 229 |
+
"model.layers.33.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 230 |
+
"model.layers.33.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 231 |
+
"model.layers.33.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 232 |
+
"model.layers.34.input_layernorm.weight": "model-00005-of-00005.safetensors",
|
| 233 |
+
"model.layers.34.mlp.down_proj.weight": "model-00005-of-00005.safetensors",
|
| 234 |
+
"model.layers.34.mlp.gate_proj.weight": "model-00004-of-00005.safetensors",
|
| 235 |
+
"model.layers.34.mlp.up_proj.weight": "model-00004-of-00005.safetensors",
|
| 236 |
+
"model.layers.34.post_attention_layernorm.weight": "model-00005-of-00005.safetensors",
|
| 237 |
+
"model.layers.34.self_attn.g_proj.weight": "model-00004-of-00005.safetensors",
|
| 238 |
+
"model.layers.34.self_attn.out_proj.weight": "model-00004-of-00005.safetensors",
|
| 239 |
+
"model.layers.34.self_attn.q_k_v_proj.weight": "model-00004-of-00005.safetensors",
|
| 240 |
+
"model.layers.35.input_layernorm.weight": "model-00005-of-00005.safetensors",
|
| 241 |
+
"model.layers.35.mlp.down_proj.weight": "model-00005-of-00005.safetensors",
|
| 242 |
+
"model.layers.35.mlp.gate_proj.weight": "model-00005-of-00005.safetensors",
|
| 243 |
+
"model.layers.35.mlp.up_proj.weight": "model-00005-of-00005.safetensors",
|
| 244 |
+
"model.layers.35.post_attention_layernorm.weight": "model-00005-of-00005.safetensors",
|
| 245 |
+
"model.layers.35.self_attn.g_proj.weight": "model-00005-of-00005.safetensors",
|
| 246 |
+
"model.layers.35.self_attn.out_proj.weight": "model-00005-of-00005.safetensors",
|
| 247 |
+
"model.layers.35.self_attn.q_k_v_proj.weight": "model-00005-of-00005.safetensors",
|
| 248 |
+
"model.layers.4.input_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 249 |
+
"model.layers.4.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
|
| 250 |
+
"model.layers.4.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
|
| 251 |
+
"model.layers.4.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
|
| 252 |
+
"model.layers.4.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 253 |
+
"model.layers.4.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
|
| 254 |
+
"model.layers.4.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
|
| 255 |
+
"model.layers.4.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
|
| 256 |
+
"model.layers.5.input_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 257 |
+
"model.layers.5.mlp.down_proj.weight": "model-00001-of-00005.safetensors",
|
| 258 |
+
"model.layers.5.mlp.gate_proj.weight": "model-00001-of-00005.safetensors",
|
| 259 |
+
"model.layers.5.mlp.up_proj.weight": "model-00001-of-00005.safetensors",
|
| 260 |
+
"model.layers.5.post_attention_layernorm.weight": "model-00001-of-00005.safetensors",
|
| 261 |
+
"model.layers.5.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
|
| 262 |
+
"model.layers.5.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
|
| 263 |
+
"model.layers.5.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
|
| 264 |
+
"model.layers.6.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 265 |
+
"model.layers.6.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 266 |
+
"model.layers.6.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 267 |
+
"model.layers.6.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 268 |
+
"model.layers.6.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 269 |
+
"model.layers.6.self_attn.g_proj.weight": "model-00001-of-00005.safetensors",
|
| 270 |
+
"model.layers.6.self_attn.out_proj.weight": "model-00001-of-00005.safetensors",
|
| 271 |
+
"model.layers.6.self_attn.q_k_v_proj.weight": "model-00001-of-00005.safetensors",
|
| 272 |
+
"model.layers.7.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 273 |
+
"model.layers.7.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 274 |
+
"model.layers.7.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 275 |
+
"model.layers.7.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 276 |
+
"model.layers.7.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 277 |
+
"model.layers.7.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 278 |
+
"model.layers.7.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 279 |
+
"model.layers.7.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 280 |
+
"model.layers.8.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 281 |
+
"model.layers.8.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 282 |
+
"model.layers.8.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 283 |
+
"model.layers.8.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 284 |
+
"model.layers.8.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 285 |
+
"model.layers.8.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 286 |
+
"model.layers.8.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 287 |
+
"model.layers.8.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 288 |
+
"model.layers.9.input_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 289 |
+
"model.layers.9.mlp.down_proj.weight": "model-00002-of-00005.safetensors",
|
| 290 |
+
"model.layers.9.mlp.gate_proj.weight": "model-00002-of-00005.safetensors",
|
| 291 |
+
"model.layers.9.mlp.up_proj.weight": "model-00002-of-00005.safetensors",
|
| 292 |
+
"model.layers.9.post_attention_layernorm.weight": "model-00002-of-00005.safetensors",
|
| 293 |
+
"model.layers.9.self_attn.g_proj.weight": "model-00002-of-00005.safetensors",
|
| 294 |
+
"model.layers.9.self_attn.out_proj.weight": "model-00002-of-00005.safetensors",
|
| 295 |
+
"model.layers.9.self_attn.q_k_v_proj.weight": "model-00002-of-00005.safetensors",
|
| 296 |
+
"model.norm.weight": "model-00005-of-00005.safetensors"
|
| 297 |
+
}
|
| 298 |
+
}
|
|
@@ -0,0 +1,483 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import math
|
| 2 |
+
|
| 3 |
+
import torch
|
| 4 |
+
import torch.nn.functional as F
|
| 5 |
+
from torch import nn
|
| 6 |
+
from transformers.activations import ACT2FN
|
| 7 |
+
from transformers.cache_utils import Cache, DynamicCache
|
| 8 |
+
from transformers.generation import GenerationMixin
|
| 9 |
+
from transformers.masking_utils import create_causal_mask, create_sliding_window_causal_mask
|
| 10 |
+
from transformers.modeling_outputs import (
|
| 11 |
+
BaseModelOutputWithPast,
|
| 12 |
+
CausalLMOutputWithPast,
|
| 13 |
+
)
|
| 14 |
+
from transformers.modeling_utils import PreTrainedModel
|
| 15 |
+
from transformers.processing_utils import Unpack
|
| 16 |
+
from transformers.pytorch_utils import ALL_LAYERNORM_LAYERS
|
| 17 |
+
from transformers.utils import TransformersKwargs, can_return_tuple, logging
|
| 18 |
+
|
| 19 |
+
from .configuration_spark import Spark2_5Config
|
| 20 |
+
|
| 21 |
+
logger = logging.get_logger(__name__)
|
| 22 |
+
|
| 23 |
+
_CONFIG_FOR_DOC = "Spark2_5Config"
|
| 24 |
+
|
| 25 |
+
def rotate_half(x):
|
| 26 |
+
x1 = x[..., : x.shape[-1] // 2]
|
| 27 |
+
x2 = x[..., x.shape[-1] // 2 :]
|
| 28 |
+
return torch.cat((-x2, x1), dim=-1)
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
def compute_rope_cos_sin(positions, head_dim, rope_theta, partial_rotary_factor=1.0, device="cpu"):
|
| 32 |
+
rope_head_dim = int(head_dim * partial_rotary_factor)
|
| 33 |
+
inv_freq = 1.0 / (rope_theta ** (torch.arange(0, rope_head_dim, 2, dtype=torch.int64).to(device="cpu", dtype=torch.float) / rope_head_dim))
|
| 34 |
+
inv_freq = inv_freq.to(device)
|
| 35 |
+
t = positions.to(device=device, dtype=torch.float32)
|
| 36 |
+
freqs = torch.outer(t, inv_freq)
|
| 37 |
+
freqs = torch.cat([freqs, freqs], dim=-1)
|
| 38 |
+
cos = freqs.cos()
|
| 39 |
+
sin = freqs.sin()
|
| 40 |
+
return cos, sin
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
def apply_rotary_pos_emb(x, cos, sin):
|
| 44 |
+
rope_head_dim = cos.shape[-1]
|
| 45 |
+
x_f32 = x.float()
|
| 46 |
+
if x_f32.shape[-1] > rope_head_dim:
|
| 47 |
+
x_rot = x_f32[..., :rope_head_dim]
|
| 48 |
+
x_pass = x_f32[..., rope_head_dim:]
|
| 49 |
+
c = cos.unsqueeze(0).unsqueeze(0)
|
| 50 |
+
s = sin.unsqueeze(0).unsqueeze(0)
|
| 51 |
+
x_rot = x_rot * c + rotate_half(x_rot) * s
|
| 52 |
+
result = torch.cat([x_rot, x_pass], dim=-1)
|
| 53 |
+
else:
|
| 54 |
+
c = cos.unsqueeze(0).unsqueeze(0)
|
| 55 |
+
s = sin.unsqueeze(0).unsqueeze(0)
|
| 56 |
+
result = x_f32 * c + rotate_half(x_f32) * s
|
| 57 |
+
return result.to(x.dtype)
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
def repeat_kv(hidden_states: torch.Tensor, n_rep: int) -> torch.Tensor:
|
| 61 |
+
batch, num_key_value_heads, slen, head_dim = hidden_states.shape
|
| 62 |
+
if n_rep == 1:
|
| 63 |
+
return hidden_states
|
| 64 |
+
hidden_states = hidden_states[:, :, None, :, :].expand(batch, num_key_value_heads, n_rep, slen, head_dim)
|
| 65 |
+
return hidden_states.reshape(batch, num_key_value_heads * n_rep, slen, head_dim)
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
def eager_attention_forward(
|
| 69 |
+
module: nn.Module,
|
| 70 |
+
query: torch.Tensor,
|
| 71 |
+
key: torch.Tensor,
|
| 72 |
+
value: torch.Tensor,
|
| 73 |
+
attention_mask: torch.Tensor | None = None,
|
| 74 |
+
scaling: float | None = None,
|
| 75 |
+
dropout: float = 0.0,
|
| 76 |
+
**kwargs: Unpack[TransformersKwargs],
|
| 77 |
+
):
|
| 78 |
+
key = repeat_kv(key, module.num_key_value_groups)
|
| 79 |
+
value = repeat_kv(value, module.num_key_value_groups)
|
| 80 |
+
|
| 81 |
+
if scaling is None:
|
| 82 |
+
scaling = 1.0 / math.sqrt(query.shape[-1])
|
| 83 |
+
|
| 84 |
+
attn_weights = torch.matmul(query, key.transpose(2, 3)) * scaling
|
| 85 |
+
if attention_mask is not None:
|
| 86 |
+
causal_mask = attention_mask[:, :, :, : key.shape[-2]]
|
| 87 |
+
attn_weights = attn_weights + causal_mask
|
| 88 |
+
|
| 89 |
+
attn_weights = attn_weights - attn_weights.max(dim=-1, keepdim=True).values
|
| 90 |
+
attn_weights = F.softmax(attn_weights, dim=-1, dtype=torch.float32).to(query.dtype)
|
| 91 |
+
attn_weights = nn.functional.dropout(attn_weights, p=dropout, training=module.training)
|
| 92 |
+
attn_output = torch.matmul(attn_weights, value)
|
| 93 |
+
return attn_output, attn_weights
|
| 94 |
+
|
| 95 |
+
|
| 96 |
+
class Spark2_5RMSNorm(nn.Module):
|
| 97 |
+
def __init__(self, hidden_size, eps=1e-6):
|
| 98 |
+
super().__init__()
|
| 99 |
+
self.weight = nn.Parameter(torch.ones(hidden_size))
|
| 100 |
+
self.variance_epsilon = eps
|
| 101 |
+
|
| 102 |
+
def forward(self, hidden_states):
|
| 103 |
+
input_dtype = hidden_states.dtype
|
| 104 |
+
hidden_states = hidden_states.to(torch.float32)
|
| 105 |
+
variance = hidden_states.pow(2).mean(-1, keepdim=True)
|
| 106 |
+
hidden_states = hidden_states * torch.rsqrt(variance + self.variance_epsilon)
|
| 107 |
+
return (self.weight.float() * hidden_states).to(input_dtype)
|
| 108 |
+
|
| 109 |
+
def extra_repr(self):
|
| 110 |
+
return f"{tuple(self.weight.shape)}, eps={self.variance_epsilon}"
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
ALL_LAYERNORM_LAYERS.append(Spark2_5RMSNorm)
|
| 114 |
+
|
| 115 |
+
|
| 116 |
+
class Spark2_5MLP(nn.Module):
|
| 117 |
+
def __init__(self, config):
|
| 118 |
+
super().__init__()
|
| 119 |
+
self.config = config
|
| 120 |
+
self.hidden_size = config.hidden_size
|
| 121 |
+
self.intermediate_size = config.intermediate_size
|
| 122 |
+
self.gate_proj = nn.Linear(self.hidden_size, self.intermediate_size, bias=config.mlp_bias)
|
| 123 |
+
self.up_proj = nn.Linear(self.hidden_size, self.intermediate_size, bias=config.mlp_bias)
|
| 124 |
+
self.down_proj = nn.Linear(self.intermediate_size, self.hidden_size, bias=config.mlp_bias)
|
| 125 |
+
|
| 126 |
+
if config.hidden_act != "gelu":
|
| 127 |
+
raise ValueError(f"只支持hidden_act='gelu',当前传入:{config.hidden_act}")
|
| 128 |
+
|
| 129 |
+
self.act_fn = ACT2FN[config.hidden_act]
|
| 130 |
+
|
| 131 |
+
def forward(self, x):
|
| 132 |
+
return self.down_proj(self.act_fn(self.gate_proj(x)) * self.up_proj(x))
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
class Spark2_5Attention(nn.Module):
|
| 136 |
+
def __init__(self, config: Spark2_5Config, layer_idx: int | None = None):
|
| 137 |
+
super().__init__()
|
| 138 |
+
self.config = config
|
| 139 |
+
self.layer_idx = layer_idx
|
| 140 |
+
self.attention_dropout = config.attention_dropout
|
| 141 |
+
self.hidden_size = config.hidden_size
|
| 142 |
+
self.num_heads = config.num_attention_heads
|
| 143 |
+
self.head_dim = config.head_dim
|
| 144 |
+
self.num_key_value_heads = config.num_key_value_heads
|
| 145 |
+
self.num_key_value_groups = self.num_heads // self.num_key_value_heads
|
| 146 |
+
self.scaling = 1.0 / math.sqrt(self.head_dim)
|
| 147 |
+
self.headwise_attn_output_gate = config.headwise_attn_output_gate
|
| 148 |
+
self.gate_attn_act_mode = config.gate_attn_act_mode
|
| 149 |
+
self.q_dim = self.num_heads * self.head_dim
|
| 150 |
+
self.kv_dim = self.num_key_value_heads * self.head_dim
|
| 151 |
+
|
| 152 |
+
qkv_out_dim = self.q_dim + 2 * self.kv_dim
|
| 153 |
+
self.q_k_v_proj = nn.Linear(self.hidden_size, qkv_out_dim, bias=config.attention_bias)
|
| 154 |
+
self.g_proj = nn.Linear(self.hidden_size, self.num_heads, bias=config.attention_bias) if self.headwise_attn_output_gate else None
|
| 155 |
+
self.out_proj = nn.Linear(self.num_heads * self.head_dim, self.hidden_size, bias=config.attention_bias)
|
| 156 |
+
self.sliding_window = None
|
| 157 |
+
|
| 158 |
+
def forward(
|
| 159 |
+
self,
|
| 160 |
+
hidden_states: torch.Tensor,
|
| 161 |
+
position_embeddings: tuple[torch.Tensor, torch.Tensor],
|
| 162 |
+
attention_mask: torch.Tensor | None = None,
|
| 163 |
+
past_key_values: Cache | None = None,
|
| 164 |
+
cache_position: torch.LongTensor | None = None,
|
| 165 |
+
**kwargs: Unpack[TransformersKwargs],
|
| 166 |
+
) -> tuple[torch.Tensor, torch.Tensor]:
|
| 167 |
+
input_shape = hidden_states.shape[:-1]
|
| 168 |
+
bsz, seq_len = input_shape
|
| 169 |
+
|
| 170 |
+
qkv = self.q_k_v_proj(hidden_states)
|
| 171 |
+
q = qkv[..., :self.q_dim]
|
| 172 |
+
k = qkv[..., self.q_dim:self.q_dim + self.kv_dim]
|
| 173 |
+
v = qkv[..., self.q_dim + self.kv_dim:]
|
| 174 |
+
gate_score = self.g_proj(hidden_states) if self.g_proj is not None else None
|
| 175 |
+
|
| 176 |
+
q = q.view(bsz, seq_len, self.num_heads, self.head_dim).transpose(1, 2)
|
| 177 |
+
k = k.view(bsz, seq_len, self.num_key_value_heads, self.head_dim).transpose(1, 2)
|
| 178 |
+
v = v.view(bsz, seq_len, self.num_key_value_heads, self.head_dim).transpose(1, 2)
|
| 179 |
+
if gate_score is not None:
|
| 180 |
+
gate_score = gate_score.view(bsz, seq_len, self.num_heads, 1).transpose(1, 2)
|
| 181 |
+
|
| 182 |
+
cos, sin = position_embeddings
|
| 183 |
+
q = apply_rotary_pos_emb(q, cos, sin)
|
| 184 |
+
k = apply_rotary_pos_emb(k, cos, sin)
|
| 185 |
+
|
| 186 |
+
|
| 187 |
+
if past_key_values is not None:
|
| 188 |
+
cache_kwargs = {"sin": sin, "cos": cos, "cache_position": cache_position}
|
| 189 |
+
k, v = past_key_values.update(k, v, self.layer_idx, cache_kwargs)
|
| 190 |
+
|
| 191 |
+
attn_output, attn_weights = eager_attention_forward(
|
| 192 |
+
self, q, k, v,
|
| 193 |
+
attention_mask=attention_mask,
|
| 194 |
+
scaling=self.scaling,
|
| 195 |
+
dropout=self.attention_dropout if self.training else 0.0,
|
| 196 |
+
)
|
| 197 |
+
|
| 198 |
+
if gate_score is not None:
|
| 199 |
+
if self.gate_attn_act_mode == "sigmoid":
|
| 200 |
+
gate = torch.sigmoid(gate_score.float())
|
| 201 |
+
elif self.gate_attn_act_mode == "silu":
|
| 202 |
+
gate = F.silu(gate_score.float())
|
| 203 |
+
else:
|
| 204 |
+
raise ValueError(f"Unsupported gate_attn_act_mode: {self.gate_attn_act_mode}")
|
| 205 |
+
gate = gate.to(attn_output.dtype)
|
| 206 |
+
attn_output = attn_output * gate
|
| 207 |
+
|
| 208 |
+
attn_output = attn_output.transpose(1, 2).contiguous().view(bsz, seq_len, -1)
|
| 209 |
+
attn_output = self.out_proj(attn_output)
|
| 210 |
+
|
| 211 |
+
return attn_output, attn_weights
|
| 212 |
+
|
| 213 |
+
|
| 214 |
+
class Spark2_5DecoderLayer(nn.Module):
|
| 215 |
+
def __init__(self, config: Spark2_5Config, layer_idx: int):
|
| 216 |
+
super().__init__()
|
| 217 |
+
self.hidden_size = config.hidden_size
|
| 218 |
+
|
| 219 |
+
self.self_attn = Spark2_5Attention(config=config, layer_idx=layer_idx)
|
| 220 |
+
self.mlp = Spark2_5MLP(config)
|
| 221 |
+
self.input_layernorm = Spark2_5RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
|
| 222 |
+
self.post_attention_layernorm = Spark2_5RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
|
| 223 |
+
|
| 224 |
+
self.layer_type = config.layer_types[layer_idx] if layer_idx < len(config.layer_types) else "full_attention"
|
| 225 |
+
if self.layer_type == "sliding_attention" and config.sliding_window is not None:
|
| 226 |
+
self.self_attn.sliding_window = config.sliding_window
|
| 227 |
+
else:
|
| 228 |
+
self.self_attn.sliding_window = None
|
| 229 |
+
self.self_attn.partial_rotary_factor = config.get_partial_rotary_factor(self.layer_type)
|
| 230 |
+
|
| 231 |
+
def forward(
|
| 232 |
+
self,
|
| 233 |
+
hidden_states: torch.Tensor,
|
| 234 |
+
position_embeddings: tuple[torch.Tensor, torch.Tensor],
|
| 235 |
+
attention_mask: torch.Tensor | None = None,
|
| 236 |
+
past_key_values: Cache | None = None,
|
| 237 |
+
cache_position: torch.LongTensor | None = None,
|
| 238 |
+
position_ids: torch.LongTensor | None = None,
|
| 239 |
+
**kwargs: Unpack[TransformersKwargs]
|
| 240 |
+
) -> torch.Tensor:
|
| 241 |
+
|
| 242 |
+
residual = hidden_states
|
| 243 |
+
hidden_states = self.input_layernorm(hidden_states)
|
| 244 |
+
hidden_states = hidden_states.to(self.mlp.gate_proj.weight.dtype)
|
| 245 |
+
|
| 246 |
+
hidden_states, _ = self.self_attn(
|
| 247 |
+
hidden_states=hidden_states,
|
| 248 |
+
position_embeddings=position_embeddings,
|
| 249 |
+
attention_mask=attention_mask,
|
| 250 |
+
past_key_values=past_key_values,
|
| 251 |
+
cache_position=cache_position,
|
| 252 |
+
position_ids=position_ids,
|
| 253 |
+
)
|
| 254 |
+
hidden_states = residual + hidden_states
|
| 255 |
+
|
| 256 |
+
residual = hidden_states
|
| 257 |
+
hidden_states = self.post_attention_layernorm(hidden_states)
|
| 258 |
+
hidden_states = hidden_states.to(self.mlp.gate_proj.weight.dtype)
|
| 259 |
+
|
| 260 |
+
|
| 261 |
+
hidden_states = self.mlp(hidden_states)
|
| 262 |
+
hidden_states = residual + hidden_states
|
| 263 |
+
|
| 264 |
+
return hidden_states
|
| 265 |
+
|
| 266 |
+
|
| 267 |
+
class Spark2_5PreTrainedModel(PreTrainedModel):
|
| 268 |
+
config_class = Spark2_5Config
|
| 269 |
+
base_model_prefix = "model"
|
| 270 |
+
supports_gradient_checkpointing = True
|
| 271 |
+
_no_split_modules = ["Spark2_5DecoderLayer"] # noqa: RUF012
|
| 272 |
+
_skip_keys_device_placement = ["past_key_values"] # noqa: RUF012
|
| 273 |
+
|
| 274 |
+
def _init_weights(self, module):
|
| 275 |
+
std = self.config.initializer_range
|
| 276 |
+
if isinstance(module, nn.Linear):
|
| 277 |
+
module.weight.data.normal_(mean=0.0, std=std)
|
| 278 |
+
if module.bias is not None:
|
| 279 |
+
module.bias.data.zero_()
|
| 280 |
+
elif isinstance(module, nn.Embedding):
|
| 281 |
+
module.weight.data.normal_(mean=0.0, std=std)
|
| 282 |
+
if module.padding_idx is not None:
|
| 283 |
+
module.weight.data[module.padding_idx].zero_()
|
| 284 |
+
|
| 285 |
+
|
| 286 |
+
class Spark2_5Model(Spark2_5PreTrainedModel):
|
| 287 |
+
def __init__(self, config: Spark2_5Config):
|
| 288 |
+
super().__init__(config)
|
| 289 |
+
self.padding_idx = config.pad_token_id
|
| 290 |
+
self.vocab_size = config.vocab_size
|
| 291 |
+
|
| 292 |
+
self.embedding = nn.Embedding(config.vocab_size, config.hidden_size, self.padding_idx)
|
| 293 |
+
self.layers = nn.ModuleList(
|
| 294 |
+
[Spark2_5DecoderLayer(config, layer_idx) for layer_idx in range(config.num_hidden_layers)]
|
| 295 |
+
)
|
| 296 |
+
self.norm = Spark2_5RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
|
| 297 |
+
self.gradient_checkpointing = False
|
| 298 |
+
self.has_sliding_layers = "sliding_attention" in config.layer_types
|
| 299 |
+
|
| 300 |
+
self.post_init()
|
| 301 |
+
|
| 302 |
+
def get_input_embeddings(self):
|
| 303 |
+
return self.embedding
|
| 304 |
+
|
| 305 |
+
def set_input_embeddings(self, value):
|
| 306 |
+
self.embedding = value
|
| 307 |
+
def forward(
|
| 308 |
+
self,
|
| 309 |
+
input_ids: torch.LongTensor = None,
|
| 310 |
+
attention_mask: torch.Tensor | None = None,
|
| 311 |
+
position_ids: torch.LongTensor | None = None,
|
| 312 |
+
past_key_values: Cache | list[torch.FloatTensor] | None = None,
|
| 313 |
+
inputs_embeds: torch.FloatTensor | None = None,
|
| 314 |
+
use_cache: bool | None = None,
|
| 315 |
+
cache_position: torch.LongTensor | None = None,
|
| 316 |
+
token_type_ids: torch.LongTensor | None = None,
|
| 317 |
+
**kwargs: Unpack[TransformersKwargs],
|
| 318 |
+
) -> BaseModelOutputWithPast:
|
| 319 |
+
use_cache = use_cache if use_cache is not None else self.config.use_cache
|
| 320 |
+
if (input_ids is None) ^ (inputs_embeds is not None):
|
| 321 |
+
raise ValueError(
|
| 322 |
+
"You cannot specify both input_ids and inputs_embeds at the same time, and must specify either one"
|
| 323 |
+
)
|
| 324 |
+
|
| 325 |
+
if self.gradient_checkpointing and self.training and use_cache:
|
| 326 |
+
logger.warning_once(
|
| 327 |
+
"`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`."
|
| 328 |
+
)
|
| 329 |
+
use_cache = False
|
| 330 |
+
|
| 331 |
+
if inputs_embeds is None:
|
| 332 |
+
inputs_embeds = self.embedding(input_ids)
|
| 333 |
+
|
| 334 |
+
if use_cache and past_key_values is None:
|
| 335 |
+
past_key_values = DynamicCache(config=self.config)
|
| 336 |
+
|
| 337 |
+
if cache_position is None:
|
| 338 |
+
past_seen_tokens = past_key_values.get_seq_length() if past_key_values is not None else 0
|
| 339 |
+
cache_position = torch.arange(
|
| 340 |
+
past_seen_tokens, past_seen_tokens + inputs_embeds.shape[1], device=inputs_embeds.device
|
| 341 |
+
)
|
| 342 |
+
|
| 343 |
+
if position_ids is None:
|
| 344 |
+
position_ids = cache_position.unsqueeze(0)
|
| 345 |
+
|
| 346 |
+
if not isinstance(attention_mask, dict):
|
| 347 |
+
mask_kwargs = {
|
| 348 |
+
"config": self.config,
|
| 349 |
+
"input_embeds": inputs_embeds,
|
| 350 |
+
"attention_mask": attention_mask,
|
| 351 |
+
"cache_position": cache_position,
|
| 352 |
+
"past_key_values": past_key_values,
|
| 353 |
+
"position_ids": position_ids,
|
| 354 |
+
}
|
| 355 |
+
causal_mask_mapping = {
|
| 356 |
+
"full_attention": create_causal_mask(**mask_kwargs),
|
| 357 |
+
}
|
| 358 |
+
if self.has_sliding_layers:
|
| 359 |
+
causal_mask_mapping["sliding_attention"] = create_sliding_window_causal_mask(**mask_kwargs)
|
| 360 |
+
else:
|
| 361 |
+
causal_mask_mapping = attention_mask
|
| 362 |
+
|
| 363 |
+
hidden_states = inputs_embeds.float()
|
| 364 |
+
|
| 365 |
+
device = hidden_states.device
|
| 366 |
+
dtype = self.embedding.weight.dtype
|
| 367 |
+
|
| 368 |
+
head_dim = self.config.head_dim
|
| 369 |
+
rope_cache = {}
|
| 370 |
+
for lt in set(self.config.layer_types):
|
| 371 |
+
rope_theta = self.config.get_rope_theta(lt)
|
| 372 |
+
prf = self.config.get_partial_rotary_factor(lt)
|
| 373 |
+
cos, sin = compute_rope_cos_sin(cache_position, head_dim, rope_theta, partial_rotary_factor=prf, device=device)
|
| 374 |
+
rope_cache[lt] = (cos, sin)
|
| 375 |
+
|
| 376 |
+
for decoder_layer in self.layers:
|
| 377 |
+
layer_type = decoder_layer.layer_type
|
| 378 |
+
position_embeddings = rope_cache.get(layer_type, rope_cache.get("full_attention"))
|
| 379 |
+
layer_attention_mask = causal_mask_mapping.get(layer_type, causal_mask_mapping.get("full_attention"))
|
| 380 |
+
|
| 381 |
+
if self.gradient_checkpointing and self.training:
|
| 382 |
+
layer_outputs = self._gradient_checkpointing_func(
|
| 383 |
+
decoder_layer.__call__,
|
| 384 |
+
hidden_states,
|
| 385 |
+
position_embeddings,
|
| 386 |
+
layer_attention_mask,
|
| 387 |
+
)
|
| 388 |
+
hidden_states = layer_outputs[0] if isinstance(layer_outputs, tuple) else layer_outputs
|
| 389 |
+
else:
|
| 390 |
+
hidden_states = decoder_layer(
|
| 391 |
+
hidden_states,
|
| 392 |
+
position_embeddings=position_embeddings,
|
| 393 |
+
attention_mask=layer_attention_mask,
|
| 394 |
+
past_key_values=past_key_values,
|
| 395 |
+
cache_position=cache_position,
|
| 396 |
+
position_ids=position_ids,
|
| 397 |
+
)
|
| 398 |
+
|
| 399 |
+
hidden_states = self.norm(hidden_states)
|
| 400 |
+
hidden_states = hidden_states.to(dtype)
|
| 401 |
+
|
| 402 |
+
return BaseModelOutputWithPast(
|
| 403 |
+
last_hidden_state=hidden_states,
|
| 404 |
+
past_key_values=past_key_values if use_cache else None,
|
| 405 |
+
)
|
| 406 |
+
|
| 407 |
+
class Spark2_5ForCausalLM(Spark2_5PreTrainedModel, GenerationMixin):
|
| 408 |
+
_tied_weights_keys = ["lm_head.weight"] # noqa: RUF012
|
| 409 |
+
|
| 410 |
+
def __init__(self, config):
|
| 411 |
+
super().__init__(config)
|
| 412 |
+
self.model = Spark2_5Model(config)
|
| 413 |
+
self.vocab_size = config.vocab_size
|
| 414 |
+
self.lm_head = nn.Linear(config.hidden_size, config.vocab_size, bias=False)
|
| 415 |
+
self.post_init()
|
| 416 |
+
|
| 417 |
+
def get_input_embeddings(self):
|
| 418 |
+
return self.model.embedding
|
| 419 |
+
|
| 420 |
+
def set_input_embeddings(self, value):
|
| 421 |
+
self.model.embedding = value
|
| 422 |
+
|
| 423 |
+
def get_output_embeddings(self):
|
| 424 |
+
return self.lm_head
|
| 425 |
+
|
| 426 |
+
def set_output_embeddings(self, new_embeddings):
|
| 427 |
+
self.lm_head = new_embeddings
|
| 428 |
+
|
| 429 |
+
def set_decoder(self, decoder):
|
| 430 |
+
self.model = decoder
|
| 431 |
+
|
| 432 |
+
def get_decoder(self):
|
| 433 |
+
return self.model
|
| 434 |
+
|
| 435 |
+
@can_return_tuple
|
| 436 |
+
def forward(
|
| 437 |
+
self,
|
| 438 |
+
input_ids: torch.LongTensor = None,
|
| 439 |
+
attention_mask: torch.Tensor | None = None,
|
| 440 |
+
position_ids: torch.LongTensor | None = None,
|
| 441 |
+
past_key_values: Cache | list[torch.FloatTensor] | None = None,
|
| 442 |
+
inputs_embeds: torch.FloatTensor | None = None,
|
| 443 |
+
labels: torch.LongTensor | None = None,
|
| 444 |
+
use_cache: bool | None = None,
|
| 445 |
+
cache_position: torch.LongTensor | None = None,
|
| 446 |
+
logits_to_keep: int = 0,
|
| 447 |
+
token_type_ids: torch.LongTensor | None = None,
|
| 448 |
+
**kwargs: Unpack[TransformersKwargs],
|
| 449 |
+
) -> CausalLMOutputWithPast:
|
| 450 |
+
outputs: BaseModelOutputWithPast = self.model(
|
| 451 |
+
input_ids=input_ids,
|
| 452 |
+
attention_mask=attention_mask,
|
| 453 |
+
position_ids=position_ids,
|
| 454 |
+
past_key_values=past_key_values,
|
| 455 |
+
inputs_embeds=inputs_embeds,
|
| 456 |
+
use_cache=use_cache,
|
| 457 |
+
cache_position=cache_position,
|
| 458 |
+
)
|
| 459 |
+
|
| 460 |
+
hidden_states = outputs.last_hidden_state
|
| 461 |
+
slice_indices = slice(-logits_to_keep, None) if isinstance(logits_to_keep, int) else logits_to_keep
|
| 462 |
+
hidden_states = hidden_states[:, slice_indices, :]
|
| 463 |
+
|
| 464 |
+
if self.config.tie_word_embeddings:
|
| 465 |
+
embed_weight = self.model.embedding.weight
|
| 466 |
+
logits = F.linear(hidden_states, embed_weight)
|
| 467 |
+
else:
|
| 468 |
+
logits = self.lm_head(hidden_states)
|
| 469 |
+
|
| 470 |
+
loss = None
|
| 471 |
+
if labels is not None:
|
| 472 |
+
loss = self.loss_function(logits=logits, labels=labels, vocab_size=self.config.vocab_size, **kwargs)
|
| 473 |
+
|
| 474 |
+
return CausalLMOutputWithPast(
|
| 475 |
+
loss=loss,
|
| 476 |
+
logits=logits,
|
| 477 |
+
past_key_values=outputs.past_key_values,
|
| 478 |
+
hidden_states=outputs.hidden_states,
|
| 479 |
+
attentions=outputs.attentions,
|
| 480 |
+
)
|
| 481 |
+
|
| 482 |
+
|
| 483 |
+
__all__ = ["Spark2_5Config", "Spark2_5ForCausalLM", "Spark2_5Model"]
|
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token": "<|start▁of▁sentence|>",
|
| 3 |
+
"eos_token": "<|end▁of▁sentence|>",
|
| 4 |
+
"unk_token": "<unk>",
|
| 5 |
+
"pad_token": "<|▁pad▁|>"
|
| 6 |
+
}
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:710cce15cf3565674c499f9413997c6e8101f2bdd96245cff8f0311fb501248c
|
| 3 |
+
size 10115786
|
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_bos_token": false,
|
| 3 |
+
"add_eos_token": false,
|
| 4 |
+
"bos_token": {
|
| 5 |
+
"__type": "AddedToken",
|
| 6 |
+
"content": "<|start▁of▁sentence|>",
|
| 7 |
+
"lstrip": false,
|
| 8 |
+
"normalized": true,
|
| 9 |
+
"rstrip": false,
|
| 10 |
+
"single_word": false
|
| 11 |
+
},
|
| 12 |
+
"clean_up_tokenization_spaces": false,
|
| 13 |
+
"eos_token": {
|
| 14 |
+
"__type": "AddedToken",
|
| 15 |
+
"content": "<|end▁of▁sentence|>",
|
| 16 |
+
"lstrip": false,
|
| 17 |
+
"normalized": true,
|
| 18 |
+
"rstrip": false,
|
| 19 |
+
"single_word": false
|
| 20 |
+
},
|
| 21 |
+
"legacy": true,
|
| 22 |
+
"model_max_length": 131072,
|
| 23 |
+
"pad_token": {
|
| 24 |
+
"__type": "AddedToken",
|
| 25 |
+
"content": "<|▁pad▁|>",
|
| 26 |
+
"lstrip": false,
|
| 27 |
+
"normalized": true,
|
| 28 |
+
"rstrip": false,
|
| 29 |
+
"single_word": false
|
| 30 |
+
},
|
| 31 |
+
"sp_model_kwargs": {},
|
| 32 |
+
"unk_token": {
|
| 33 |
+
"__type": "AddedToken",
|
| 34 |
+
"content": "<unk>",
|
| 35 |
+
"lstrip": false,
|
| 36 |
+
"normalized": true,
|
| 37 |
+
"rstrip": false,
|
| 38 |
+
"single_word": false
|
| 39 |
+
},
|
| 40 |
+
"tokenizer_class": "PreTrainedTokenizerFast",
|
| 41 |
+
"chat_template": "{%- if not messages %}{{- raise_exception('No messages provided.') }}{%- endif %}{%- set enable_thinking = enable_thinking | default(true) %}{#- Render a string or a list of text blocks. -#}{%- macro render_content(content, context_name) %}{%- if content is string %}{{- content }}{%- elif content is none or content is undefined %}{{- '' }}{%- elif content is iterable and content is not mapping %}{%- for block in content %}{%- if block.type == 'text' %}{{- block.text }}{%- else %}{{- raise_exception('Unsupported ' ~ context_name ~ ' content block type: ' ~ (block.type | string)) }}{%- endif %}{%- endfor %}{%- else %}{{- raise_exception(context_name ~ ' content must be a string or a list of text blocks') }}{%- endif %}{%- endmacro %}{#- Default system prompt -#}{%- set default_system = 'you are a helpful assistant.' %}{#- The first message-level system is placed in the initial system block. -#}{%- set ns = namespace(initial_system='') %}{%- if messages[0].role == 'system' %}{%- set ns.initial_system = render_content(messages[0].content, 'system') %}{%- endif %}{#- System block -#}{{- '<|start▁of▁sentence|><|System|>' + '\n' + default_system }}{%- if tools %}{{- '## Tools' + '\n' + 'You have access to the following functions:' + '\n' + '<tools>' }}{%- for tool in tools %}{{- '\n' + tool.function | tojson}}{%- endfor %}{{- '\n' + '</tools>' }}{%- endif %}{%- if ns.initial_system %}{{- '\n\n' + ns.initial_system }}{%- endif %}{{- '<|end▁of▁sentence|>'}}{#- Conversation turns -#}{%- for message in messages %}{%- if message.role == 'system' %}{#- The first system message was consumed by the initial block. -#}{%- if not loop.first %}{{- '<|start▁of▁sentence|><|System|>\n' + render_content(message.content, 'system') + '<|end▁of▁sentence|>' }}{%- endif %}{%- elif message.role == 'user' %}{{- '<|start▁of▁sentence|><|User|>' + render_content(message.content, 'user') + '<|end▁of▁sentence|>' }}{%- elif message.role == 'assistant' %}{%- set assistant_content = render_content(message.content, 'assistant') %}{%- if message.reasoning_content is defined and message.reasoning_content %}{%- set reasoning_content = message.reasoning_content %}{%- else %}{%- set reasoning_content = '' %}{%- endif %}{{- '<|start▁of▁sentence|><|Bot|>'}}{%- if reasoning_content %}{{- '<think>' + reasoning_content + '</think>'}}{%- else %}{{- '</think>' }}{%- endif %}{%- if assistant_content %}{{- assistant_content }}{%- endif %}{%- if message.tool_calls is defined and message.tool_calls is not none %}{%- for tool_call in message.tool_calls %}{%- if tool_call.function.arguments is not mapping %}{{- raise_exception('tool_call.function.arguments must be a dictionary; normalize JSON strings before apply_chat_template') }}{%- endif %}{%- set args = tool_call.function.arguments %}{{- '<tool_call>' + tool_call.function.name }}{%- for k, v in args.items() %}{{- '<arg_key>' ~ k ~ '</arg_key><arg_value>' ~ (v if v is string else v | tojson) ~ '</arg_value>' }}{%- endfor %}{{- '</tool_call>' }}{%- endfor %}{%- endif %}{{- '<|end▁of▁sentence|>' }}{%- elif message.role == 'tool' %}{%- if loop.previtem is undefined or loop.previtem.role != 'tool' %}{{- '<|start▁of▁sentence|><|Tool|>' }}{%- endif %}{{- '<tool_response>' ~ message.content ~ '</tool_response>' }}{%- if loop.nextitem is undefined or loop.nextitem.role != 'tool' %}{{- '<|end▁of▁sentence|>' }}{%- endif %}{%- else %}{{- raise_exception('Unsupported message role: ' ~ message.role) }}{%- endif %}{%- endfor %}{#- Generation prompt -#}{%- if add_generation_prompt %}{{- '<|start▁of▁sentence|><|Bot|>' }}{%- if enable_thinking is defined and enable_thinking %}{{- '<think>' }}{%- endif %}{%- if enable_thinking is defined and not enable_thinking %}{{- '</think>' }}{%- endif %}{%- endif %}"
|
| 42 |
+
}
|
|
The diff for this file is too large to render.
See raw diff
|
|
|