Instructions to use arcee-ai/Trinity-Large-TrueBase with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use arcee-ai/Trinity-Large-TrueBase with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="arcee-ai/Trinity-Large-TrueBase", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)

# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("arcee-ai/Trinity-Large-TrueBase", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("arcee-ai/Trinity-Large-TrueBase", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))

Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use arcee-ai/Trinity-Large-TrueBase with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "arcee-ai/Trinity-Large-TrueBase"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "arcee-ai/Trinity-Large-TrueBase",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/arcee-ai/Trinity-Large-TrueBase

SGLang

How to use arcee-ai/Trinity-Large-TrueBase with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "arcee-ai/Trinity-Large-TrueBase" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "arcee-ai/Trinity-Large-TrueBase",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "arcee-ai/Trinity-Large-TrueBase" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "arcee-ai/Trinity-Large-TrueBase",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Docker Model Runner
How to use arcee-ai/Trinity-Large-TrueBase with Docker Model Runner:
```
docker model run hf.co/arcee-ai/Trinity-Large-TrueBase
```

Alyosha11

sapiosaturn

lckr commited on Jan 27

Commit

18a66cf

verified ·

0 Parent(s):

Super-squash branch 'main' using huggingface_hub

Browse files

Co-authored-by: sapiosaturn <sapiosaturn@users.noreply.huggingface.co>
Co-authored-by: lckr <lckr@users.noreply.huggingface.co>

Files changed (39) hide show

.gitattributes +36 -0
README.md +151 -0
chat_template.jinja +1 -0
config.json +108 -0
configuration_afmoe.py +133 -0
model-00001-of-00031.safetensors +3 -0
model-00002-of-00031.safetensors +3 -0
model-00003-of-00031.safetensors +3 -0
model-00004-of-00031.safetensors +3 -0
model-00005-of-00031.safetensors +3 -0
model-00006-of-00031.safetensors +3 -0
model-00007-of-00031.safetensors +3 -0
model-00008-of-00031.safetensors +3 -0
model-00009-of-00031.safetensors +3 -0
model-00010-of-00031.safetensors +3 -0
model-00011-of-00031.safetensors +3 -0
model-00012-of-00031.safetensors +3 -0
model-00013-of-00031.safetensors +3 -0
model-00014-of-00031.safetensors +3 -0
model-00015-of-00031.safetensors +3 -0
model-00016-of-00031.safetensors +3 -0
model-00017-of-00031.safetensors +3 -0
model-00018-of-00031.safetensors +3 -0
model-00019-of-00031.safetensors +3 -0
model-00020-of-00031.safetensors +3 -0
model-00021-of-00031.safetensors +3 -0
model-00022-of-00031.safetensors +3 -0
model-00023-of-00031.safetensors +3 -0
model-00024-of-00031.safetensors +3 -0
model-00025-of-00031.safetensors +3 -0
model-00026-of-00031.safetensors +3 -0
model-00027-of-00031.safetensors +3 -0
model-00028-of-00031.safetensors +3 -0
model-00029-of-00031.safetensors +3 -0
model-00030-of-00031.safetensors +3 -0
model-00031-of-00031.safetensors +3 -0
model.safetensors.index.json +0 -0
tokenizer.json +3 -0
tokenizer_config.json +271 -0

.gitattributes ADDED Viewed

	@@ -0,0 +1,36 @@

+*.7z filter=lfs diff=lfs merge=lfs -text
+*.arrow filter=lfs diff=lfs merge=lfs -text
+*.bin filter=lfs diff=lfs merge=lfs -text
+*.bz2 filter=lfs diff=lfs merge=lfs -text
+*.ckpt filter=lfs diff=lfs merge=lfs -text
+*.ftz filter=lfs diff=lfs merge=lfs -text
+*.gz filter=lfs diff=lfs merge=lfs -text
+*.h5 filter=lfs diff=lfs merge=lfs -text
+*.joblib filter=lfs diff=lfs merge=lfs -text
+*.lfs.* filter=lfs diff=lfs merge=lfs -text
+*.mlmodel filter=lfs diff=lfs merge=lfs -text
+*.model filter=lfs diff=lfs merge=lfs -text
+*.msgpack filter=lfs diff=lfs merge=lfs -text
+*.npy filter=lfs diff=lfs merge=lfs -text
+*.npz filter=lfs diff=lfs merge=lfs -text
+*.onnx filter=lfs diff=lfs merge=lfs -text
+*.ot filter=lfs diff=lfs merge=lfs -text
+*.parquet filter=lfs diff=lfs merge=lfs -text
+*.pb filter=lfs diff=lfs merge=lfs -text
+*.pickle filter=lfs diff=lfs merge=lfs -text
+*.pkl filter=lfs diff=lfs merge=lfs -text
+*.pt filter=lfs diff=lfs merge=lfs -text
+*.pth filter=lfs diff=lfs merge=lfs -text
+*.rar filter=lfs diff=lfs merge=lfs -text
+*.safetensors filter=lfs diff=lfs merge=lfs -text
+saved_model/**/* filter=lfs diff=lfs merge=lfs -text
+*.tar.* filter=lfs diff=lfs merge=lfs -text
+*.tar filter=lfs diff=lfs merge=lfs -text
+*.tflite filter=lfs diff=lfs merge=lfs -text
+*.tgz filter=lfs diff=lfs merge=lfs -text
+*.wasm filter=lfs diff=lfs merge=lfs -text
+*.xz filter=lfs diff=lfs merge=lfs -text
+*.zip filter=lfs diff=lfs merge=lfs -text
+*.zst filter=lfs diff=lfs merge=lfs -text
+*tfevents* filter=lfs diff=lfs merge=lfs -text
+tokenizer.json filter=lfs diff=lfs merge=lfs -text

README.md ADDED Viewed

	@@ -0,0 +1,151 @@

+---
+license: apache-2.0
+language:
+- en
+- es
+- fr
+- de
+- it
+- pt
+- ru
+- ar
+- hi
+- ko
+- zh
+library_name: transformers
+---
+<!-- markdownlint-disable first-line-h1 -->
+<!-- markdownlint-disable html -->
+<!-- markdownlint-disable no-duplicate-header -->
+<div align="center">
+  <picture>
+    <img
+      src="https://cdn-uploads.huggingface.co/production/uploads/6435718aaaef013d1aec3b8b/i-v1KyAMOW_mgVGeic9WJ.png"
+      alt="Arcee Trinity Large"
+      style="max-width: 100%; height: auto;"
+    >
+  </picture>
+</div>
+<hr>
+# Trinity-Large-TrueBase
+## Introduction
+Trinity-Large-TrueBase is a base pretraining checkpoint from Arcee AI's Trinity Large training run. It is a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. The checkpoint was captured after 10 trillion tokens of pretraining, prior to learning-rate annealing and before any instruction tuning or reinforcement learning.
+This checkpoint is intended for research, probing, ablation studies, and downstream fine-tuning and comes without any pre-baked alignment, instruction formatting, or preference optimization.
+More details on the training of Trinity Large are available in the [technical report](https://github.com/arcee-ai/trinity-large-tech-report/).
+## Model Variants
+The Trinity Large family consists of three checkpoints from the same training run:
+- **Trinity-Large-TrueBase** (this release): 10T-token pre-anneal checkpoint with no instruction data
+- **[Trinity-Large-Base](https://huggingface.co/arcee-ai/Trinity-Large-Base)**: Full 17T-token pretrained foundation model with mid-training anneals
+- **[Trinity-Large-Preview](https://huggingface.co/arcee-ai/Trinity-Large-Preview)**: Lightly post-trained, chat-ready model undergoing active RL
+## Architecture
+Trinity-Large-TrueBase uses a sparse MoE configuration designed to maximize efficiency while maintaining large-scale capacity.
+| Hyperparameter | Value |
+|:---|:---:|
+| Total parameters | ~398B |
+| Active parameters per token | ~13B |
+| Experts | 256 |
+| Active experts | 4 |
+| Routing strategy | 4-of-256 (1.56% sparsity) |
+| Dense layers | 6 |
+| Pretraining context length | 8,192 |
+| Architecture | Sparse MoE (AfmoeForCausalLM) |
+Note: Extended context support (e.g., 512k) was introduced after this checkpoint and is not available in TrueBase.
+## Benchmark Results
+| Benchmark                     | N-shot | Metric                        | Score  | Stderr  |
+|-------------------------------|--------|-------------------------------|--------|---------|
+| arc_challenge_0shot           | 0      | acc_norm,none                 | 0.6237 | ±0.0142 |
+| bbh_fewshot                   | 3      | exact_match,remove_whitespace | 0.5784 | ±0.0054 |
+| gpqa_diamond_5shot            | 5      | acc_norm,none                 | 0.4091 | ±0.0350 |
+| gpqa_diamond_generative_5shot | 5      | exact_match,flexible-extract  | 0.3788 | ±0.0346 |
+| gsm8k_8shot                   | 8      | exact_match,flexible-extract  | 0.8036 | ±0.0109 |
+| gsm8k_cot                     | 8      | exact_match,flexible-extract  | 0.8044 | ±0.0109 |
+| hellaswag_5shot               | 5      | acc_norm,none                 | 0.8813 | ±0.0032 |
+| humaneval_plus                | 0      | pass@1,create_test            | 0.5183 | ±0.0391 |
+| leaderboard_math_hard         | 4      | exact_match,none              | 0.2696 | ±0.0113 |
+| mbpp_plus                     | 3      | pass_at_1,none                | 0.8095 | ±0.0202 |
+| minerva_math500               | 4      | math_verify,none              | 0.4820 | ±0.0224 |
+| mmlu_5shot                    | 5      | acc,none                      | 0.7845 | ±0.0033 |
+| mmlu_generative_5shot         | 5      | exact_match,get_response      | 0.7848 | ±0.0033 |
+| mmlu_pro                      | 5      | exact_match,custom-extract    | 0.5160 | ±0.0044 |
+| triviaqa_5shot                | 5      | exact_match,remove_whitespace | 0.8096 | ±0.0029 |
+| winogrande_5shot              | 5      | acc,none                      | 0.8145 | ±0.0109 |
+## Training Configuration
+### Pretraining
+- Training tokens: 10 trillion
+- Checkpoint type: Pre-anneal
+- Instruction data: None
+- RLHF or post-training: None
+This checkpoint branches from the main Trinity Large run at the 10T-token mark, prior to learning-rate decay or post-training phases.
+### Optimizers
+Optimizer learning rates after WSD warm-up:
+- Adam learning rate: 2e-4
+- Muon learning rate: 8e-4
+Muon was used to support larger critical batch sizes in a highly sparse MoE regime.
+### Infrastructure
+- Hardware: 2,048 NVIDIA B300 GPUs
+- Parallelism: HSDP + Expert Parallelism
+- Compute partner: [Prime Intellect](https://www.primeintellect.ai/)
+- Data partner: [Datology](https://www.datologyai.com/)
+<div align="center">
+  <picture>
+      <img src="https://cdn-uploads.huggingface.co/production/uploads/6435718aaaef013d1aec3b8b/sSVjGNHfrJKmQ6w8I18ek.png" style="background-color:ghostwhite;padding:5px;" width="17%" alt="Powered by Datology">
+  </picture>
+</div>
+<div align="center">
+  <picture>
+      <img src="https://cdn-avatars.huggingface.co/v1/production/uploads/61e020e4a343274bb132e138/H2mcdPRWtl4iKLd-OYYBc.jpeg" style="background-color:ghostwhite;padding:5px;" width="17%" alt="Powered by Datology">
+  </picture>
+</div>
+## Intended Use
+- Studying emergent behavior from large-scale pretraining
+- Sparse MoE routing and load-balancing research
+- Interpretability, probing, and ablation studies
+- Domain-specific fine-tuning from a clean base
+- Academic and industrial foundation model research
+## Rationale for Release
+Most base model releases include instruction data, annealed training dynamics, or early alignment stages. Trinity-Large-TrueBase excludes these, providing an opportunity to study what large-scale models learn from pretraining data alone. This checkpoint is intended as a foundation for research rather than as a finished conversational assistant.
+## Known Limitations
+- Not aligned for safety, helpfulness, or conversational tone
+- Requires substantial compute and expertise to fine-tune
+- May exhibit raw or unstable behaviors typical of unaligned models
+- No extended-context tuning beyond the 8K pretraining window
+## License
+Trinity-Large-TrueBase is released under the Apache License, Version 2.0.

chat_template.jinja ADDED Viewed

	@@ -0,0 +1 @@


1	+ {{ bos_token }}{% for message in messages %}{{ message['content'] }}{% endfor %}

config.json ADDED Viewed

	@@ -0,0 +1,108 @@

+{
+  "architectures": [
+    "AfmoeForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "auto_map": {
+    "AutoConfig": "configuration_afmoe.AfmoeConfig",
+    "AutoModel": "modeling_afmoe.AfmoeModel",
+    "AutoModelForCausalLM": "modeling_afmoe.AfmoeForCausalLM"
+  },
+  "dtype": "bfloat16",
+  "global_attn_every_n_layers": 4,
+  "head_dim": 128,
+  "hidden_act": "silu",
+  "hidden_size": 3072,
+  "initializer_range": 0.02,
+  "intermediate_size": 12288,
+  "layer_types": [
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention"
+  ],
+  "load_balance_coeff": 0.00005,
+  "max_position_embeddings": 8192,
+  "model_type": "afmoe",
+  "moe_intermediate_size": 3072,
+  "mup_enabled": true,
+  "n_group": 1,
+  "num_attention_heads": 48,
+  "num_dense_layers": 6,
+  "num_expert_groups": 1,
+  "num_experts": 256,
+  "num_experts_per_tok": 4,
+  "num_hidden_layers": 60,
+  "num_key_value_heads": 8,
+  "num_limited_groups": 1,
+  "num_shared_experts": 1,
+  "rms_norm_eps": 1e-05,
+  "rope_scaling": null,
+  "rope_theta": 10000,
+  "route_norm": true,
+  "route_scale": 2.448,
+  "score_func": "sigmoid",
+  "sliding_window": 4096,
+  "tie_word_embeddings": false,
+  "topk_group": 1,
+  "transformers_version": "4.57.1",
+  "use_cache": true,
+  "use_grouped_mm": true,
+  "vocab_size": 200192
+}

configuration_afmoe.py ADDED Viewed

	@@ -0,0 +1,133 @@

+# coding=utf-8
+# Copyright 2022 EleutherAI and the HuggingFace Inc. team. All rights reserved.
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+#     http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+from transformers.configuration_utils import PretrainedConfig
+from transformers.modeling_rope_utils import rope_config_validation
+from transformers.configuration_utils import layer_type_validation
+from transformers.utils import logging
+logger = logging.get_logger(__name__)
+class AfmoeConfig(PretrainedConfig):
+    """
+    n_group (`int`, *optional*, defaults to 1):
+            Number of groups for routed experts.
+    topk_group (`int`, *optional*, defaults to 1):
+        Number of selected groups for each token(for each token, ensuring the selected experts is only within `topk_group` groups).
+    """
+    model_type = "afmoe"
+    base_model_pp_plan = {
+        "embed_tokens": (["input_ids"], ["inputs_embeds"]),
+        "layers": (["hidden_states", "attention_mask"], ["hidden_states"]),
+        "norm": (["hidden_states"], ["hidden_states"]),
+    }
+    def __init__(
+        self,
+        num_hidden_layers: int = 32,
+        vocab_size: int = 200192,
+        hidden_size: int = 2048,
+        intermediate_size: int = 6144,
+        moe_intermediate_size=1408,
+        num_dense_layers=1,
+        num_attention_heads=16,
+        num_key_value_heads=None,
+        head_dim=128,
+        hidden_act="silu",
+        max_position_embeddings=16384,
+        initializer_range=0.02,
+        rms_norm_eps=1e-5,
+        use_cache=True,
+        tie_word_embeddings=False,
+        rope_theta=10000.0,
+        rope_scaling=None,
+        num_experts=64,
+        num_experts_per_tok=6,
+        num_shared_experts=2,
+        num_expert_groups=1,
+        num_limited_groups=1,
+        score_func="sigmoid",
+        route_norm=True,
+        route_scale=1.0,
+        global_attn_every_n_layers=4,
+        sliding_window=1024,
+        mup_enabled=False,
+        layer_types=None,
+        attention_dropout: float = 0.0,
+        n_group: int = 1,
+        topk_group: int = 1,
+        **kwargs,
+    ):
+        self.vocab_size = vocab_size
+        self.max_position_embeddings = max_position_embeddings
+        self.hidden_size = hidden_size
+        self.intermediate_size = intermediate_size
+        self.num_hidden_layers = num_hidden_layers
+        self.num_dense_layers = num_dense_layers
+        self.num_attention_heads = num_attention_heads
+        self.head_dim = head_dim
+        self.hidden_act = hidden_act
+        self.initializer_range = initializer_range
+        self.rms_norm_eps = rms_norm_eps
+        self.use_cache = use_cache
+        self.rope_theta = rope_theta
+        self.rope_scaling = rope_scaling
+        # MoE specific
+        self.moe_intermediate_size = moe_intermediate_size
+        self.num_experts_per_tok = num_experts_per_tok
+        self.n_group = n_group
+        self.topk_group = topk_group
+        self.num_experts = num_experts
+        self.num_shared_experts = num_shared_experts
+        self.num_expert_groups = num_expert_groups
+        self.num_limited_groups = num_limited_groups
+        self.score_func = score_func
+        self.route_norm = route_norm
+        self.route_scale = route_scale
+        # Attention specific
+        self.attention_dropout = attention_dropout
+        self.global_attn_every_n_layers = global_attn_every_n_layers
+        self.sliding_window = sliding_window
+        self.layer_types = layer_types
+        if self.layer_types is None:
+            self.layer_types = [
+                "sliding_attention" if bool((i + 1) % global_attn_every_n_layers) else "full_attention" for i in range(self.num_hidden_layers)
+            ]
+        layer_type_validation(self.layer_types)
+        # muP specific
+        self.mup_enabled = mup_enabled
+        if num_key_value_heads is None:
+            num_key_value_heads = num_attention_heads
+        self.num_key_value_heads = num_key_value_heads
+        # Validate rope configs
+        if self.rope_scaling is not None and "type" in self.rope_scaling:
+            self.rope_scaling["rope_type"] = self.rope_scaling["type"]
+        rope_config_validation(self)
+        super().__init__(
+            tie_word_embeddings=tie_word_embeddings,
+            **kwargs,
+        )
+__all__ = ["AfmoeConfig"]

model-00001-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d8ef47e43a3bd45ed35a7bfa029d41a9664041c408739912b1de400d10388245
+size 2459965736

model-00002-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c349e1581c0929d54b54b07344372ae9cb9bf3ed0aeb4e88b8c11314cd57b7bd
+size 704696408

model-00003-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f598dfbe3a990a39d7fc2a7e8dd05593c85943423cd4a65d43e740c2083bafb5
+size 704696408

model-00004-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7658299483daeb115c2b1ccc3c408640d8b405c49ec9c13c82ebf0662ce24220
+size 704696408

model-00005-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:ad75328510e97233f0ad61a4cefad7e691cd1be1aac8a94a0535c19a7be42bbf
+size 29359328160

model-00006-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:926d0493fd9a188d60f4672e0927d91a9d08ec44cd0308652f1ca3847ba0f075
+size 29359328160

model-00007-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e6e070cc896a4b44e3ac9f7ea0102de543844ab95cc378d97981fd3bc310d261
+size 29359329728

model-00008-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:358a46bdc004dd43ca1369f39736b21ac6039205fcac26a77bb86c9e43d253f4
+size 29359329728

model-00009-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d937d038548864681b462bdf0c719a093385c28f70e601aeec4db07772bdd006
+size 29359329728

model-00010-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7a3491020cc1ea7adf57b8921bf7e3b6f959f58da86b72b5cbfa11ef390baf2f
+size 29359329728

model-00011-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7825668e95ecf1c70d764d6d93d4b958322290598e92ce430c885bfbc5ce02d1
+size 29359329728

model-00012-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d594b9b247190857751d505709d91f631655f233f2d5af109299d086a1f04a0d
+size 29359329728

model-00013-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7194268ada9198556d2a7af652fb801742dc93526e6349faf456ecc477ff4be9
+size 29359329728

model-00014-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f5998370a186aa28a17efe298504fe08905e7207a43f37188280f84609cf286d
+size 29359329728

model-00015-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:68f969d915e20ca17deb3d815d68ce7115475689b79e6dbb03f254a6c365e838
+size 29359329728

model-00016-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5abbc0ab1bf566db036285f39f74fc19fc6a0d010184111a9cb9a63c11a3eb4d
+size 29359329728

model-00017-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:b095cd23be4d33e44778dfd2fa10f8c602c765373da9f6c23b8ab7573dbec777
+size 29359329728

model-00018-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:593f073e5de6b0c758204eae716064de9b994dfb2b5d88224bc7389d7851bf99
+size 29359329728

model-00019-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e35cdb8c4cf59a5741037ffe688ab7d32812b5b31cf16129500904976150b986
+size 29359329728

model-00020-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f225b62691f6095bb6dcd1819ca38f5c93946813849e285e0e5b7145835a56b9
+size 29359329728

model-00021-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:cafc496484c88020509890bb48e126a632652dd751d295952f14b64827f52957
+size 29359329728

model-00022-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:0c3e735fb88e3ba9849940eeb92c2e27a4a0552228f40b5cf6236838fd77f4d1
+size 29359329728

model-00023-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5dc529b8b359f5a01d731ff857282a38cf08b37414cbf60f6923ad021058545e
+size 29359329728

model-00024-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:3e67e816b792801ecaa119be4ae7c2ea5c5d02525e5b28e0b2d73e6dbc8e5743
+size 29359329728

model-00025-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e3218fe985aae8fad4aabfded78454efe80f3b2b7d5a0bcce9997b3e6785820d
+size 29359329728

model-00026-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8011c5a9b86d216bdf6721c1960e70f7fd21259e9651ecaf36775909e3fa1720
+size 29359329728

model-00027-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:67ada14883072cde0392e01cae9314691bc512e17122eab4508b4b51a2f9f17e
+size 29359329728

model-00028-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:253e535e5a4409b2e62ca2785e744aa862db4d65619a49237e809238cc6f7fa9
+size 29359329728

model-00029-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4d3c2a1eeb215d09087965623f66f9a1c9c8fb3671ee23d4f1d598a533db44bf
+size 29359329728

model-00030-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:324d716716027b07141103a136269d5441153d8814be7d4126d44a53efa02384
+size 29359329728

model-00031-of-00031.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:20e47fc5a288fd634d952422c22838604d655e2f19ea709cff7f608d93668122
+size 29359329728

model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer.json ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:606b0740c5d8940dbee3268e05beaa3e34f4e43bc7df153681122f8fcc78e134
+size 14614897

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,271 @@

+{
+  "add_bos_token": true,
+  "add_eos_token": false,
+  "add_prefix_space": null,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<|begin_of_text|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<|end_of_text|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "3": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "4": {
+      "content": "<|eot_id|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "5": {
+      "content": "<|start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "6": {
+      "content": "<|channel|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "7": {
+      "content": "<|message|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "8": {
+      "content": "<|end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "9": {
+      "content": "<|fitm_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "10": {
+      "content": "<|fitm_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "11": {
+      "content": "<|fitm_hole|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "12": {
+      "content": "<|pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "13": {
+      "content": "<|reserved_special_0|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "14": {
+      "content": "<|reserved_special_1|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "15": {
+      "content": "<|reserved_special_2|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "16": {
+      "content": "<|reserved_special_3|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "17": {
+      "content": "<|reserved_special_4|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "18": {
+      "content": "<|reserved_special_5|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "19": {
+      "content": "<|reserved_special_6|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "20": {
+      "content": "<|reserved_special_7|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "21": {
+      "content": "<|reserved_special_8|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "22": {
+      "content": "<|reserved_special_9|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "23": {
+      "content": "<|reserved_special_10|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "24": {
+      "content": "<|reserved_special_11|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "25": {
+      "content": "<|reserved_special_12|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "26": {
+      "content": "<|reserved_special_13|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "27": {
+      "content": "<|reserved_special_14|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "28": {
+      "content": "<|reserved_special_15|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "29": {
+      "content": "<|reserved_special_16|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "30": {
+      "content": "<|reserved_special_17|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "31": {
+      "content": "<|reserved_special_18|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|begin_of_text|>",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "extra_special_tokens": {},
+  "model_max_length": 65536,
+  "pad_token": "<|pad|>",
+  "tokenizer_class": "PreTrainedTokenizerFast",
+  "use_default_system_prompt": false
+}