mishrahimanshuedu ceadar-ie commited on
Commit
d4bb8b0
·
0 Parent(s):

Duplicate from ceadar-ie/FinanceConnect-13B

Browse files

Co-authored-by: CeADAR <ceadar-ie@users.noreply.huggingface.co>

.gitattributes ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,196 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language: en
3
+ datasets:
4
+ - FinTalk-19k
5
+ tags:
6
+ - summarization
7
+ - classification
8
+ - translation
9
+ - NLP
10
+ - finance
11
+ - domain specific llm
12
+ license: apache-2.0
13
+ pipeline_tag: text-generation
14
+ ---
15
+
16
+ # FinanceConnect
17
+
18
+ FinanceConnect is a state-of-the-art, open-source chat model tailored for finance and economic discussions. Built on the robust Llama2-13B architecture, this model has been fine-tuned on a combination of FinTalk-19k and Alpaca datasets, making it a valuable resource for finance professionals, researchers, and enthusiasts.
19
+
20
+ ## Model Details
21
+
22
+ - Architecture: Llama2-13B
23
+ - Training Dataset: [FinTalk-19k](https://huggingface.co/datasets/ceadar-ie/FinTalk-19k), [Alpaca](https://huggingface.co/datasets/tatsu-lab/alpaca)
24
+
25
+ ## Dataset Utilized: FinTalk-19k and Alpaca
26
+
27
+ Drawing strength from the FinTalk-19k and Alpaca dataset, a curated collection focused on financial knowledge, this model provides insights and information related to the finance industry. For a deeper dive into the dataset, visit: [FinTalk-19k](https://huggingface.co/datasets/ceadar-ie/FinTalk-19k), [Alpaca](https://huggingface.co/datasets/tatsu-lab/alpaca)
28
+
29
+ ## Model Specification
30
+
31
+ - **Developed by:** CeADAR Connect Group
32
+ - **Model type:** Large Language Model
33
+ - **Language(s):** en
34
+ - **Finetuned from model:** Llama2-13B
35
+
36
+ ## Key Features and Functionalities
37
+
38
+ - **Domain Specialization:** The FinanceConnect model is specialized in Finance conversations, serving as a resource for financial researchers, and enthusiasts.
39
+ - **Model API Accessibility:** Offers a straightforward Python integration for generating financial content insights.
40
+ - **Performance Optimisation:** Efficient performance across both CPU and GPU platforms.
41
+ - **Data Representation:** Utilises a combination of comprehensive Finance dataset, enabling content generation to professional standards.
42
+
43
+ ## Benchmarks
44
+ | **Benchmark** | **BloombergGPT 50B** | **FinanceConnect 13B** |
45
+ |--------------|--------------|--------------|
46
+ | MMLU | 39.8 | 52.08 |
47
+ | FPB | 51.1 | 57.2 |
48
+ | **Cost**| **$2.67 Million** | **$27** |
49
+
50
+ | **Benchmark** | **FinanceConnect 13B** |
51
+ |--------------|--------------
52
+ | MMLU | 52.08 |
53
+ | ARC | 55.12 |
54
+ | HellaSwag | 77.73 |
55
+ | TruthfulQA | 38.80 |
56
+ | Winogrande | 71.82 |
57
+ | GSM8K | 1.6 |
58
+
59
+ ## Model Usage
60
+ Experience the capabilities of the FinanceConnect model through a well-structured Python interface. To kick-start your exploration, follow the steps and snippets given below:
61
+
62
+ ### Prerequisites
63
+ #### 1. Ensure required packages are available
64
+
65
+ ```python
66
+ import torch
67
+ from typing import Any, Dict
68
+ from transformers import (
69
+ AutoModelForCausalLM,
70
+ AutoTokenizer,
71
+ BitsAndBytesConfig,
72
+ HfArgumentParser,
73
+ TrainingArguments,
74
+ PreTrainedTokenizerFast,
75
+ pipeline,
76
+ logging,
77
+ )
78
+ import time
79
+ ```
80
+
81
+ #### 2. Initiate the model and tokenizer
82
+
83
+ ```python
84
+ model_name = "ceadar-ie/FinanceConnect-13B"
85
+ tokenizer = AutoTokenizer.from_pretrained(model_name)
86
+ model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16, load_in_8bit = True, device_map = "auto", trust_remote_code=True)
87
+ ```
88
+
89
+ #### 3. Create a function for generating text
90
+
91
+ ```python
92
+ def generate_text(input_text):
93
+ tokenizer.padding_side = "right"
94
+ num_words_input = len(input_text.split())
95
+
96
+ start_time = time.time()
97
+
98
+ pipe = pipeline(task="text-generation", model=model, tokenizer=tokenizer, max_length = 1000)
99
+ prompt = f"<s>[INST] {input_text} [/INST]"
100
+ #prompt = input_text
101
+ result = pipe(prompt)
102
+ generated_text = result[0]['generated_text'].replace(prompt, "").strip()
103
+
104
+ end_time = time.time()
105
+ time_taken = end_time - start_time
106
+
107
+ num_words_generated = len(generated_text.split())
108
+
109
+ print("Generated Output:")
110
+ print(generated_text)
111
+ print('=' * 60 )
112
+ print("\nNumber of Words in Input Text:", num_words_input)
113
+ print("Number of Words in Generated Output:", num_words_generated)
114
+ print("Time Taken to Generate Output: {:.2f} seconds".format(time_taken))
115
+ ```
116
+
117
+ ## Example Evaluation and Use
118
+ ### Example Prompt 1:
119
+ List in detail ten key factors influencing the current state of the global economy.
120
+
121
+ ### Generated Output:
122
+ The current state of the global economy is influenced by several factors, including:
123
+
124
+ 1. Globalization: The increasing interconnectedness of economies through trade, investment, and the movement of people.
125
+ 2. Technological advancements: The rapid pace of technological innovation and its impact on industries and jobs.
126
+ 3. Political factors: Government policies and decisions that affect economic growth and stability.
127
+ 4. Demographics: Changes in population demographics, such as aging populations and youth bulges, can impact economic growth and stability.
128
+ 5. Environmental factors: Climate change, natural disasters, and other environmental factors can disrupt economic activity and impact growth.
129
+ 6. International trade: The rise of protectionism and trade tensions between countries can impact economic growth and stability.
130
+ 7. Financial factors: The performance of financial markets and the availability of credit can impact economic growth and stability.
131
+ 8. Infrastructure: The availability and quality of infrastructure, such as transportation and communication networks, can impact economic growth and stability.
132
+ 9. Education: The availability and quality of education can impact economic growth and stability by influencing the supply of skilled workers and the ability of countries to adapt to changing economic conditions.
133
+ 10. Health: The availability and quality of healthcare can impact economic growth and stability by influencing the productivity of the workforce and the ability of countries to attract and retain talent.
134
+
135
+ ### Example Prompt 2:
136
+ Explain the concept of quantitative easing and its impact on financial markets.
137
+
138
+ ### Generated Output:
139
+ Quantitative easing is a monetary policy tool used by central banks to stimulate economic growth and combat inflation. It involves the purchase of assets such as government bonds, mortgage-backed securities, and corporate bonds by a central bank. This increases the amount of money in circulation, lowers interest rates, and encourages banks to lend more, which can boost economic growth. However, quantitative easing can also lead to inflation if it is not implemented carefully, as it can cause prices to rise more than wages can keep up with.
140
+
141
+ ### Example Prompt 3:
142
+ I've recently started working and want to manage my finances. Give me a detailed answer in 10 points on how do I set up a budget?
143
+
144
+ ### Generated Output:
145
+ Setting up a budget can help manage finances and save money. Here are some tips:
146
+
147
+ 1. Determine your income and expenses.
148
+ 2. Create a budget that includes all necessary expenses, such as rent, utilities, and groceries.
149
+ 3. Set aside a portion of your income for savings and investments.
150
+ 4. Monitor your spending regularly and adjust your budget as needed.
151
+ 5. Use cash or debit cards instead of credit cards to avoid overspending.
152
+ 6. Cut back on unnecessary expenses, such as eating out or subscriptions.
153
+ 7. Consider using a budgeting app or tool to help track your spending.
154
+ 8. Prioritize your spending based on your financial goals.
155
+ 9. Review your budget regularly and make adjustments as needed.
156
+ 10. Be flexible and willing to make changes to your budget as your financial situation changes.
157
+
158
+ ## Training Details
159
+ ### Training Hyperparameters
160
+ - per_device_train_batch_size = 10
161
+ - gradient_accumulation_steps = 4
162
+ - optim = "paged_adamw_32bit"
163
+ - learning_rate = 2e-4
164
+ - max_grad_norm = 0.3
165
+ - warmup_ratio = 0.03
166
+
167
+ ## Licensing
168
+ The FinanceConnect model, developed by CeADAR Connect Group, combines the licensing frameworks of Llama2, FinTalk-8k and Alpaca. Under Meta's terms, users are granted a non-exclusive, worldwide, non-transferable, royalty-free limited license for the use and modification of Llama Materials, inclusive of the Llama2 model and its associated documentation. When redistributing, the provided Agreement and a specific attribution notice must be included. Further, in alignment with the FinTalk dataset's(Apache 2.0) licensing and Alpaca dataset's(cc-by-nc-4.0) licensing, the model is distributed under the umbrella of all three licenses.
169
+
170
+ ## Model Limitations
171
+ ### Out-of-Scope Use
172
+ FinanceConnect is specifically tailored for finanical discussions and knowledge. It is not optimized for:
173
+ - General conversations.
174
+ - Domain-specific tasks outside financial tasks.
175
+ - Direct interfacing with physical devices or applications.
176
+
177
+ ### Bias, Risks, and Limitations
178
+ - Dataset Biases: The FinTalk-19k and Alpaca dataset may contain inherent biases that influence the model's outputs.
179
+ - Over-reliance: The model is an aid, not a replacement for human expertise. Decisions should be made with careful consideration.
180
+ - Content Understanding: The model lacks human-like understanding and cannot judge the veracity of knowledge.
181
+ - Language Limitations: The model's primary language is English. Performance may decrease with other languages.
182
+ - Knowledge Cut-off: The model may not be aware of events or trends post its last training update.
183
+
184
+ ## Citation
185
+ ```
186
+ @misc {ceadar_2023,
187
+ author = { {CeADAR} },
188
+ title = { FinanceConnect-13B (Revision 5f7841d) },
189
+ year = 2023,
190
+ url = { https://huggingface.co/ceadar-ie/FinanceConnect-13B },
191
+ doi = { 10.57967/hf/1405 },
192
+ publisher = { Hugging Face }
193
+ }
194
+ ```
195
+ ## Contact
196
+ For any further inquiries or feedback concerning FinanceConnect, please forward your communications to ahtsham.zafar@ucd.ie
config.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_name_or_path": "meta-llama/Llama-2-13b-chat-hf",
3
+ "architectures": [
4
+ "LlamaForCausalLM"
5
+ ],
6
+ "attention_bias": false,
7
+ "attention_dropout": 0.0,
8
+ "bos_token_id": 1,
9
+ "eos_token_id": 2,
10
+ "hidden_act": "silu",
11
+ "hidden_size": 5120,
12
+ "initializer_range": 0.02,
13
+ "intermediate_size": 13824,
14
+ "max_position_embeddings": 4096,
15
+ "model_type": "llama",
16
+ "num_attention_heads": 40,
17
+ "num_hidden_layers": 40,
18
+ "num_key_value_heads": 40,
19
+ "pretraining_tp": 1,
20
+ "rms_norm_eps": 1e-05,
21
+ "rope_scaling": null,
22
+ "rope_theta": 10000.0,
23
+ "tie_word_embeddings": false,
24
+ "torch_dtype": "float16",
25
+ "transformers_version": "4.36.0.dev0",
26
+ "use_cache": true,
27
+ "vocab_size": 32000
28
+ }
generation_config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 1,
3
+ "do_sample": true,
4
+ "eos_token_id": 2,
5
+ "max_length": 4096,
6
+ "pad_token_id": 0,
7
+ "temperature": 0.6,
8
+ "top_p": 0.9,
9
+ "transformers_version": "4.36.0.dev0"
10
+ }
model-00001-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:528e1360af52277e06311226d16fce53e114585c62dfbe54c0fe77182cc64626
3
+ size 4978265728
model-00002-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3accbf9480df22c74233ef2ad0041ecd8d10b5207da723c67036eb6cdec97cba
3
+ size 4970422160
model-00003-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a9710afd61d457f31dbd0c12522e33371df4f6d4caba21e041f850dd8d17208e
3
+ size 4970422184
model-00004-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7c41bc9549d9de5c799e811e1291cf4a1dac2d5bf826f07d0a0762f12c72caf7
3
+ size 4933701432
model-00005-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:be64c9db127c7e55cc952b39ded0abc83b60548c507e18d4e6dc8d0d67d9c130
3
+ size 4933722144
model-00006-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:77ce8e32f9ea28e2a32865f6a3f446f242cedda9db9db20d0080962cef46a22e
3
+ size 1245236904
model.safetensors.index.json ADDED
@@ -0,0 +1,370 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata": {
3
+ "total_size": 26031728640
4
+ },
5
+ "weight_map": {
6
+ "lm_head.weight": "model-00006-of-00006.safetensors",
7
+ "model.embed_tokens.weight": "model-00001-of-00006.safetensors",
8
+ "model.layers.0.input_layernorm.weight": "model-00001-of-00006.safetensors",
9
+ "model.layers.0.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
10
+ "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
11
+ "model.layers.0.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
12
+ "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
13
+ "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
14
+ "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
15
+ "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
16
+ "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
17
+ "model.layers.1.input_layernorm.weight": "model-00001-of-00006.safetensors",
18
+ "model.layers.1.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
19
+ "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
20
+ "model.layers.1.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
21
+ "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
22
+ "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
23
+ "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
24
+ "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
25
+ "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
26
+ "model.layers.10.input_layernorm.weight": "model-00002-of-00006.safetensors",
27
+ "model.layers.10.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
28
+ "model.layers.10.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
29
+ "model.layers.10.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
30
+ "model.layers.10.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
31
+ "model.layers.10.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
32
+ "model.layers.10.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
33
+ "model.layers.10.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
34
+ "model.layers.10.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
35
+ "model.layers.11.input_layernorm.weight": "model-00002-of-00006.safetensors",
36
+ "model.layers.11.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
37
+ "model.layers.11.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
38
+ "model.layers.11.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
39
+ "model.layers.11.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
40
+ "model.layers.11.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
41
+ "model.layers.11.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
42
+ "model.layers.11.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
43
+ "model.layers.11.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
44
+ "model.layers.12.input_layernorm.weight": "model-00002-of-00006.safetensors",
45
+ "model.layers.12.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
46
+ "model.layers.12.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
47
+ "model.layers.12.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
48
+ "model.layers.12.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
49
+ "model.layers.12.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
50
+ "model.layers.12.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
51
+ "model.layers.12.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
52
+ "model.layers.12.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
53
+ "model.layers.13.input_layernorm.weight": "model-00002-of-00006.safetensors",
54
+ "model.layers.13.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
55
+ "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
56
+ "model.layers.13.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
57
+ "model.layers.13.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
58
+ "model.layers.13.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
59
+ "model.layers.13.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
60
+ "model.layers.13.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
61
+ "model.layers.13.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
62
+ "model.layers.14.input_layernorm.weight": "model-00002-of-00006.safetensors",
63
+ "model.layers.14.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
64
+ "model.layers.14.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
65
+ "model.layers.14.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
66
+ "model.layers.14.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
67
+ "model.layers.14.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
68
+ "model.layers.14.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
69
+ "model.layers.14.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
70
+ "model.layers.14.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
71
+ "model.layers.15.input_layernorm.weight": "model-00003-of-00006.safetensors",
72
+ "model.layers.15.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
73
+ "model.layers.15.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
74
+ "model.layers.15.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
75
+ "model.layers.15.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
76
+ "model.layers.15.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
77
+ "model.layers.15.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
78
+ "model.layers.15.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
79
+ "model.layers.15.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
80
+ "model.layers.16.input_layernorm.weight": "model-00003-of-00006.safetensors",
81
+ "model.layers.16.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
82
+ "model.layers.16.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
83
+ "model.layers.16.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
84
+ "model.layers.16.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
85
+ "model.layers.16.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
86
+ "model.layers.16.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
87
+ "model.layers.16.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
88
+ "model.layers.16.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
89
+ "model.layers.17.input_layernorm.weight": "model-00003-of-00006.safetensors",
90
+ "model.layers.17.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
91
+ "model.layers.17.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
92
+ "model.layers.17.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
93
+ "model.layers.17.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
94
+ "model.layers.17.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
95
+ "model.layers.17.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
96
+ "model.layers.17.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
97
+ "model.layers.17.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
98
+ "model.layers.18.input_layernorm.weight": "model-00003-of-00006.safetensors",
99
+ "model.layers.18.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
100
+ "model.layers.18.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
101
+ "model.layers.18.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
102
+ "model.layers.18.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
103
+ "model.layers.18.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
104
+ "model.layers.18.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
105
+ "model.layers.18.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
106
+ "model.layers.18.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
107
+ "model.layers.19.input_layernorm.weight": "model-00003-of-00006.safetensors",
108
+ "model.layers.19.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
109
+ "model.layers.19.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
110
+ "model.layers.19.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
111
+ "model.layers.19.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
112
+ "model.layers.19.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
113
+ "model.layers.19.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
114
+ "model.layers.19.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
115
+ "model.layers.19.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
116
+ "model.layers.2.input_layernorm.weight": "model-00001-of-00006.safetensors",
117
+ "model.layers.2.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
118
+ "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
119
+ "model.layers.2.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
120
+ "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
121
+ "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
122
+ "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
123
+ "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
124
+ "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
125
+ "model.layers.20.input_layernorm.weight": "model-00003-of-00006.safetensors",
126
+ "model.layers.20.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
127
+ "model.layers.20.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
128
+ "model.layers.20.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
129
+ "model.layers.20.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
130
+ "model.layers.20.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
131
+ "model.layers.20.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
132
+ "model.layers.20.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
133
+ "model.layers.20.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
134
+ "model.layers.21.input_layernorm.weight": "model-00003-of-00006.safetensors",
135
+ "model.layers.21.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
136
+ "model.layers.21.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
137
+ "model.layers.21.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
138
+ "model.layers.21.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
139
+ "model.layers.21.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
140
+ "model.layers.21.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
141
+ "model.layers.21.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
142
+ "model.layers.21.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
143
+ "model.layers.22.input_layernorm.weight": "model-00003-of-00006.safetensors",
144
+ "model.layers.22.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
145
+ "model.layers.22.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
146
+ "model.layers.22.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
147
+ "model.layers.22.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
148
+ "model.layers.22.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
149
+ "model.layers.22.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
150
+ "model.layers.22.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
151
+ "model.layers.22.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
152
+ "model.layers.23.input_layernorm.weight": "model-00004-of-00006.safetensors",
153
+ "model.layers.23.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
154
+ "model.layers.23.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
155
+ "model.layers.23.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
156
+ "model.layers.23.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
157
+ "model.layers.23.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
158
+ "model.layers.23.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
159
+ "model.layers.23.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
160
+ "model.layers.23.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
161
+ "model.layers.24.input_layernorm.weight": "model-00004-of-00006.safetensors",
162
+ "model.layers.24.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
163
+ "model.layers.24.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
164
+ "model.layers.24.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
165
+ "model.layers.24.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
166
+ "model.layers.24.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
167
+ "model.layers.24.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
168
+ "model.layers.24.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
169
+ "model.layers.24.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
170
+ "model.layers.25.input_layernorm.weight": "model-00004-of-00006.safetensors",
171
+ "model.layers.25.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
172
+ "model.layers.25.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
173
+ "model.layers.25.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
174
+ "model.layers.25.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
175
+ "model.layers.25.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
176
+ "model.layers.25.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
177
+ "model.layers.25.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
178
+ "model.layers.25.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
179
+ "model.layers.26.input_layernorm.weight": "model-00004-of-00006.safetensors",
180
+ "model.layers.26.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
181
+ "model.layers.26.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
182
+ "model.layers.26.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
183
+ "model.layers.26.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
184
+ "model.layers.26.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
185
+ "model.layers.26.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
186
+ "model.layers.26.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
187
+ "model.layers.26.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
188
+ "model.layers.27.input_layernorm.weight": "model-00004-of-00006.safetensors",
189
+ "model.layers.27.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
190
+ "model.layers.27.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
191
+ "model.layers.27.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
192
+ "model.layers.27.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
193
+ "model.layers.27.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
194
+ "model.layers.27.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
195
+ "model.layers.27.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
196
+ "model.layers.27.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
197
+ "model.layers.28.input_layernorm.weight": "model-00004-of-00006.safetensors",
198
+ "model.layers.28.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
199
+ "model.layers.28.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
200
+ "model.layers.28.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
201
+ "model.layers.28.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
202
+ "model.layers.28.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
203
+ "model.layers.28.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
204
+ "model.layers.28.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
205
+ "model.layers.28.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
206
+ "model.layers.29.input_layernorm.weight": "model-00004-of-00006.safetensors",
207
+ "model.layers.29.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
208
+ "model.layers.29.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
209
+ "model.layers.29.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
210
+ "model.layers.29.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
211
+ "model.layers.29.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
212
+ "model.layers.29.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
213
+ "model.layers.29.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
214
+ "model.layers.29.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
215
+ "model.layers.3.input_layernorm.weight": "model-00001-of-00006.safetensors",
216
+ "model.layers.3.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
217
+ "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
218
+ "model.layers.3.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
219
+ "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
220
+ "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
221
+ "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
222
+ "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
223
+ "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
224
+ "model.layers.30.input_layernorm.weight": "model-00005-of-00006.safetensors",
225
+ "model.layers.30.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
226
+ "model.layers.30.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
227
+ "model.layers.30.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
228
+ "model.layers.30.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
229
+ "model.layers.30.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
230
+ "model.layers.30.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
231
+ "model.layers.30.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
232
+ "model.layers.30.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
233
+ "model.layers.31.input_layernorm.weight": "model-00005-of-00006.safetensors",
234
+ "model.layers.31.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
235
+ "model.layers.31.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
236
+ "model.layers.31.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
237
+ "model.layers.31.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
238
+ "model.layers.31.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
239
+ "model.layers.31.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
240
+ "model.layers.31.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
241
+ "model.layers.31.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
242
+ "model.layers.32.input_layernorm.weight": "model-00005-of-00006.safetensors",
243
+ "model.layers.32.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
244
+ "model.layers.32.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
245
+ "model.layers.32.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
246
+ "model.layers.32.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
247
+ "model.layers.32.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
248
+ "model.layers.32.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
249
+ "model.layers.32.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
250
+ "model.layers.32.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
251
+ "model.layers.33.input_layernorm.weight": "model-00005-of-00006.safetensors",
252
+ "model.layers.33.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
253
+ "model.layers.33.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
254
+ "model.layers.33.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
255
+ "model.layers.33.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
256
+ "model.layers.33.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
257
+ "model.layers.33.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
258
+ "model.layers.33.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
259
+ "model.layers.33.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
260
+ "model.layers.34.input_layernorm.weight": "model-00005-of-00006.safetensors",
261
+ "model.layers.34.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
262
+ "model.layers.34.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
263
+ "model.layers.34.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
264
+ "model.layers.34.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
265
+ "model.layers.34.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
266
+ "model.layers.34.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
267
+ "model.layers.34.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
268
+ "model.layers.34.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
269
+ "model.layers.35.input_layernorm.weight": "model-00005-of-00006.safetensors",
270
+ "model.layers.35.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
271
+ "model.layers.35.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
272
+ "model.layers.35.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
273
+ "model.layers.35.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
274
+ "model.layers.35.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
275
+ "model.layers.35.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
276
+ "model.layers.35.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
277
+ "model.layers.35.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
278
+ "model.layers.36.input_layernorm.weight": "model-00005-of-00006.safetensors",
279
+ "model.layers.36.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
280
+ "model.layers.36.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
281
+ "model.layers.36.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
282
+ "model.layers.36.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
283
+ "model.layers.36.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
284
+ "model.layers.36.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
285
+ "model.layers.36.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
286
+ "model.layers.36.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
287
+ "model.layers.37.input_layernorm.weight": "model-00005-of-00006.safetensors",
288
+ "model.layers.37.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
289
+ "model.layers.37.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
290
+ "model.layers.37.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
291
+ "model.layers.37.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
292
+ "model.layers.37.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
293
+ "model.layers.37.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
294
+ "model.layers.37.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
295
+ "model.layers.37.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
296
+ "model.layers.38.input_layernorm.weight": "model-00006-of-00006.safetensors",
297
+ "model.layers.38.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
298
+ "model.layers.38.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
299
+ "model.layers.38.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
300
+ "model.layers.38.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
301
+ "model.layers.38.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
302
+ "model.layers.38.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
303
+ "model.layers.38.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
304
+ "model.layers.38.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
305
+ "model.layers.39.input_layernorm.weight": "model-00006-of-00006.safetensors",
306
+ "model.layers.39.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
307
+ "model.layers.39.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
308
+ "model.layers.39.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
309
+ "model.layers.39.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
310
+ "model.layers.39.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
311
+ "model.layers.39.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
312
+ "model.layers.39.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
313
+ "model.layers.39.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
314
+ "model.layers.4.input_layernorm.weight": "model-00001-of-00006.safetensors",
315
+ "model.layers.4.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
316
+ "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
317
+ "model.layers.4.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
318
+ "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
319
+ "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
320
+ "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
321
+ "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
322
+ "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
323
+ "model.layers.5.input_layernorm.weight": "model-00001-of-00006.safetensors",
324
+ "model.layers.5.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
325
+ "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
326
+ "model.layers.5.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
327
+ "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
328
+ "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
329
+ "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
330
+ "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
331
+ "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
332
+ "model.layers.6.input_layernorm.weight": "model-00001-of-00006.safetensors",
333
+ "model.layers.6.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
334
+ "model.layers.6.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
335
+ "model.layers.6.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
336
+ "model.layers.6.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
337
+ "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
338
+ "model.layers.6.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
339
+ "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
340
+ "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
341
+ "model.layers.7.input_layernorm.weight": "model-00002-of-00006.safetensors",
342
+ "model.layers.7.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
343
+ "model.layers.7.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
344
+ "model.layers.7.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
345
+ "model.layers.7.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
346
+ "model.layers.7.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
347
+ "model.layers.7.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
348
+ "model.layers.7.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
349
+ "model.layers.7.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
350
+ "model.layers.8.input_layernorm.weight": "model-00002-of-00006.safetensors",
351
+ "model.layers.8.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
352
+ "model.layers.8.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
353
+ "model.layers.8.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
354
+ "model.layers.8.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
355
+ "model.layers.8.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
356
+ "model.layers.8.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
357
+ "model.layers.8.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
358
+ "model.layers.8.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
359
+ "model.layers.9.input_layernorm.weight": "model-00002-of-00006.safetensors",
360
+ "model.layers.9.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
361
+ "model.layers.9.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
362
+ "model.layers.9.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
363
+ "model.layers.9.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
364
+ "model.layers.9.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
365
+ "model.layers.9.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
366
+ "model.layers.9.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
367
+ "model.layers.9.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
368
+ "model.norm.weight": "model-00006-of-00006.safetensors"
369
+ }
370
+ }
special_tokens_map.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "eos_token": {
10
+ "content": "</s>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "unk_token": {
17
+ "content": "<unk>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ }
23
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347
3
+ size 499723
tokenizer_config.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": true,
3
+ "add_eos_token": false,
4
+ "added_tokens_decoder": {
5
+ "0": {
6
+ "content": "<unk>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "1": {
14
+ "content": "<s>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "2": {
22
+ "content": "</s>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ }
29
+ },
30
+ "bos_token": "<s>",
31
+ "chat_template": "{% if messages[0]['role'] == 'system' %}{% set loop_messages = messages[1:] %}{% set system_message = messages[0]['content'] %}{% else %}{% set loop_messages = messages %}{% set system_message = false %}{% endif %}{% for message in loop_messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if loop.index0 == 0 and system_message != false %}{% set content = '<<SYS>>\\n' + system_message + '\\n<</SYS>>\\n\\n' + message['content'] %}{% else %}{% set content = message['content'] %}{% endif %}{% if message['role'] == 'user' %}{{ bos_token + '[INST] ' + content.strip() + ' [/INST]' }}{% elif message['role'] == 'assistant' %}{{ ' ' + content.strip() + ' ' + eos_token }}{% endif %}{% endfor %}",
32
+ "clean_up_tokenization_spaces": false,
33
+ "eos_token": "</s>",
34
+ "legacy": false,
35
+ "model_max_length": 1000000000000000019884624838656,
36
+ "pad_token": null,
37
+ "padding_side": "right",
38
+ "sp_model_kwargs": {},
39
+ "tokenizer_class": "LlamaTokenizer",
40
+ "unk_token": "<unk>",
41
+ "use_default_system_prompt": false
42
+ }