Instructions to use arcee-ai/Trinity-Large-Thinking-FP8-Block with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use arcee-ai/Trinity-Large-Thinking-FP8-Block with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="arcee-ai/Trinity-Large-Thinking-FP8-Block", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)

# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("arcee-ai/Trinity-Large-Thinking-FP8-Block", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("arcee-ai/Trinity-Large-Thinking-FP8-Block", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))

Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use arcee-ai/Trinity-Large-Thinking-FP8-Block with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "arcee-ai/Trinity-Large-Thinking-FP8-Block"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "arcee-ai/Trinity-Large-Thinking-FP8-Block",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/arcee-ai/Trinity-Large-Thinking-FP8-Block

SGLang

How to use arcee-ai/Trinity-Large-Thinking-FP8-Block with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "arcee-ai/Trinity-Large-Thinking-FP8-Block" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "arcee-ai/Trinity-Large-Thinking-FP8-Block",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "arcee-ai/Trinity-Large-Thinking-FP8-Block" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "arcee-ai/Trinity-Large-Thinking-FP8-Block",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Docker Model Runner
How to use arcee-ai/Trinity-Large-Thinking-FP8-Block with Docker Model Runner:
```
docker model run hf.co/arcee-ai/Trinity-Large-Thinking-FP8-Block
```

Anneketh Vij commited on Apr 1

Commit

6c33626

0 Parent(s):

Super-squash branch 'main' using huggingface_hub

Browse files

This view is limited to 50 files because it contains too many changes. See raw diff

Files changed (50) hide show

.gitattributes +36 -0
README.md +142 -0
__init__.py +4 -0
chat_template.jinja +159 -0
config.json +515 -0
configuration_afmoe.py +133 -0
generation_config.json +10 -0
model-00001-of-00081.safetensors +3 -0
model-00002-of-00081.safetensors +3 -0
model-00003-of-00081.safetensors +3 -0
model-00004-of-00081.safetensors +3 -0
model-00005-of-00081.safetensors +3 -0
model-00006-of-00081.safetensors +3 -0
model-00007-of-00081.safetensors +3 -0
model-00008-of-00081.safetensors +3 -0
model-00009-of-00081.safetensors +3 -0
model-00010-of-00081.safetensors +3 -0
model-00011-of-00081.safetensors +3 -0
model-00012-of-00081.safetensors +3 -0
model-00013-of-00081.safetensors +3 -0
model-00014-of-00081.safetensors +3 -0
model-00015-of-00081.safetensors +3 -0
model-00016-of-00081.safetensors +3 -0
model-00017-of-00081.safetensors +3 -0
model-00018-of-00081.safetensors +3 -0
model-00019-of-00081.safetensors +3 -0
model-00020-of-00081.safetensors +3 -0
model-00021-of-00081.safetensors +3 -0
model-00022-of-00081.safetensors +3 -0
model-00023-of-00081.safetensors +3 -0
model-00024-of-00081.safetensors +3 -0
model-00025-of-00081.safetensors +3 -0
model-00026-of-00081.safetensors +3 -0
model-00027-of-00081.safetensors +3 -0
model-00028-of-00081.safetensors +3 -0
model-00029-of-00081.safetensors +3 -0
model-00030-of-00081.safetensors +3 -0
model-00031-of-00081.safetensors +3 -0
model-00032-of-00081.safetensors +3 -0
model-00033-of-00081.safetensors +3 -0
model-00034-of-00081.safetensors +3 -0
model-00035-of-00081.safetensors +3 -0
model-00036-of-00081.safetensors +3 -0
model-00037-of-00081.safetensors +3 -0
model-00038-of-00081.safetensors +3 -0
model-00039-of-00081.safetensors +3 -0
model-00040-of-00081.safetensors +3 -0
model-00041-of-00081.safetensors +3 -0
model-00042-of-00081.safetensors +3 -0
model-00043-of-00081.safetensors +3 -0

.gitattributes ADDED Viewed

	@@ -0,0 +1,36 @@

+*.7z filter=lfs diff=lfs merge=lfs -text
+*.arrow filter=lfs diff=lfs merge=lfs -text
+*.bin filter=lfs diff=lfs merge=lfs -text
+*.bz2 filter=lfs diff=lfs merge=lfs -text
+*.ckpt filter=lfs diff=lfs merge=lfs -text
+*.ftz filter=lfs diff=lfs merge=lfs -text
+*.gz filter=lfs diff=lfs merge=lfs -text
+*.h5 filter=lfs diff=lfs merge=lfs -text
+*.joblib filter=lfs diff=lfs merge=lfs -text
+*.lfs.* filter=lfs diff=lfs merge=lfs -text
+*.mlmodel filter=lfs diff=lfs merge=lfs -text
+*.model filter=lfs diff=lfs merge=lfs -text
+*.msgpack filter=lfs diff=lfs merge=lfs -text
+*.npy filter=lfs diff=lfs merge=lfs -text
+*.npz filter=lfs diff=lfs merge=lfs -text
+*.onnx filter=lfs diff=lfs merge=lfs -text
+*.ot filter=lfs diff=lfs merge=lfs -text
+*.parquet filter=lfs diff=lfs merge=lfs -text
+*.pb filter=lfs diff=lfs merge=lfs -text
+*.pickle filter=lfs diff=lfs merge=lfs -text
+*.pkl filter=lfs diff=lfs merge=lfs -text
+*.pt filter=lfs diff=lfs merge=lfs -text
+*.pth filter=lfs diff=lfs merge=lfs -text
+*.rar filter=lfs diff=lfs merge=lfs -text
+*.safetensors filter=lfs diff=lfs merge=lfs -text
+saved_model/**/* filter=lfs diff=lfs merge=lfs -text
+*.tar.* filter=lfs diff=lfs merge=lfs -text
+*.tar filter=lfs diff=lfs merge=lfs -text
+*.tflite filter=lfs diff=lfs merge=lfs -text
+*.tgz filter=lfs diff=lfs merge=lfs -text
+*.wasm filter=lfs diff=lfs merge=lfs -text
+*.xz filter=lfs diff=lfs merge=lfs -text
+*.zip filter=lfs diff=lfs merge=lfs -text
+*.zst filter=lfs diff=lfs merge=lfs -text
+*tfevents* filter=lfs diff=lfs merge=lfs -text
+tokenizer.json filter=lfs diff=lfs merge=lfs -text

README.md ADDED Viewed

	@@ -0,0 +1,142 @@

+---
+license: apache-2.0
+language:
+- en
+- es
+- fr
+- de
+- it
+- pt
+- ru
+- ar
+- hi
+- ko
+- zh
+library_name: transformers
+base_model:
+- arcee-ai/Trinity-Large-Thinking
+base_model_relation: quantized
+tags:
+- reasoning
+- agentic
+- tool-calling
+- thinking
+---
+<!-- markdownlint-disable first-line-h1 -->
+<!-- markdownlint-disable html -->
+<!-- markdownlint-disable no-duplicate-header -->
+<div align="center">
+  <picture>
+    <img
+      src="https://cdn-uploads.huggingface.co/production/uploads/6435718aaaef013d1aec3b8b/i-v1KyAMOW_mgVGeic9WJ.png"
+      alt="Arcee Trinity Large Thinking"
+      style="max-width: 100%; height: auto;"
+    >
+  </picture>
+</div>
+<hr>
+# Trinity-Large-Thinking-FP8-Block
+## Introduction
+Trinity-Large-Thinking is a reasoning-optimized variant of Arcee AI's Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token, post-trained with extended chain-of-thought reasoning and agentic RL.
+**This repository contains the FP8 block-quantized weights of Trinity-Large-Thinking (FP8 weights and activations with per-block scaling).**
+For full model details, benchmarks, and usage guidance, see the main [Trinity-Large-Thinking](https://huggingface.co/arcee-ai/Trinity-Large-Thinking) model card.
+## Quantization Details
+- **Scheme:** `FP8 Block` (FP8 weights and activations, per-block scaling with E8M0 scale format)
+- **Format:** `compressed-tensors`
+- **Intended use:** High-throughput FP8 deployment with near-lossless quality, optimized for NVIDIA Hopper/Blackwell GPUs
+- **Supported backends:** [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM), vLLM CUTLASS, Triton
+## Usage
+### Inference tested on
+- 8x NVIDIA H100 80GB (tensor parallel = 8)
+- vLLM 0.18.0+
+### vLLM
+Supported in vLLM 0.18.0+ with DeepGEMM FP8 MoE acceleration.
+```bash
+pip install "vllm>=0.18.0"
+```
+Serving with DeepGEMM enabled (recommended):
+```bash
+VLLM_USE_DEEP_GEMM=1 vllm serve arcee-ai/Trinity-Large-Thinking-FP8-Block \
+  --trust-remote-code \
+  --tensor-parallel-size 8 \
+  --enable-reasoning \
+  --reasoning-parser deepseek_r1 \
+  --enable-auto-tool-choice \
+  --tool-call-parser qwen3_coder
+```
+Without DeepGEMM (falls back to CUTLASS/Triton):
+```bash
+vllm serve arcee-ai/Trinity-Large-Thinking-FP8-Block \
+  --trust-remote-code \
+  --tensor-parallel-size 8 \
+  --enable-reasoning \
+  --reasoning-parser deepseek_r1 \
+  --enable-auto-tool-choice \
+  --tool-call-parser qwen3_coder
+```
+### Transformers
+```python
+from transformers import AutoTokenizer, AutoModelForCausalLM
+model_id = "arcee-ai/Trinity-Large-Thinking-FP8-Block"
+tokenizer = AutoTokenizer.from_pretrained(model_id)
+model = AutoModelForCausalLM.from_pretrained(
+    model_id,
+    device_map="auto",
+    trust_remote_code=True
+)
+messages = [{"role": "user", "content": "Who are you?"}]
+input_ids = tokenizer.apply_chat_template(
+    messages, add_generation_prompt=True, return_tensors="pt"
+).to(model.device)
+outputs = model.generate(input_ids, max_new_tokens=4096, do_sample=True, temperature=0.6, top_k=50, top_p=0.95)
+print(tokenizer.decode(outputs[0], skip_special_tokens=True))
+```
+### API
+Works out of the box on [OpenRouter](https://openrouter.ai/) as `arcee-ai/trinity-large-thinking`.
+## License
+Trinity-Large-Thinking-FP8-Block is released under the Apache License, Version 2.0.
+## Citation
+If you use this model, please cite:
+```bibtex
+@misc{singh2026arceetrinity,
+  title        = {Arcee Trinity Large Technical Report},
+  author       = {Varun Singh and Lucas Krauss and Sami Jaghouar and Matej Sirovatka and Charles Goddard and Fares Obied and Jack Min Ong and Jannik Straube and Fern and Aria Harley and Conner Stewart and Colin Kealty and Maziyar Panahi and Simon Kirsten and Anushka Deshpande and Anneketh Vij and Arthur Bresnu and Pranav Veldurthi and Raghav Ravishankar and Hardik Bishnoi and DatologyAI Team and Arcee AI Team and Prime Intellect Team and Mark McQuade and Johannes Hagemann and Lucas Atkins},
+  year         = {2026},
+  eprint       = {2602.17004},
+  archivePrefix= {arXiv},
+  primaryClass = {cs.LG},
+  doi          = {10.48550/arXiv.2602.17004},
+  url          = {https://arxiv.org/abs/2602.17004}
+}
+```

__init__.py ADDED Viewed

	@@ -0,0 +1,4 @@

+from .configuration_afmoe import AfmoeConfig
+from .modeling_afmoe     import AfmoeForCausalLM
+__all__ = ["AfmoeConfig", "AfmoeForCausalLM"]

chat_template.jinja ADDED Viewed

	@@ -0,0 +1,159 @@

+<|begin_of_text|>{%- macro render_extra_keys(json_dict, handled_keys) -%}
+    {%- if json_dict is mapping %}
+        {%- for json_key in json_dict if json_key not in handled_keys %}
+            {%- if json_dict[json_key] is mapping or (json_dict[json_key] is sequence and json_dict[json_key] is not string) %}
+                {{- '\n<' ~ json_key ~ '>' ~ (json_dict[json_key] | tojson | safe) ~ '</' ~ json_key ~ '>' }}
+            {%- else %}
+                {{- '\n<' ~ json_key ~ '>' ~ (json_dict[json_key] | string) ~ '</' ~ json_key ~ '>' }}
+            {%- endif %}
+        {%- endfor %}
+    {%- endif %}
+{%- endmacro -%}
+{%- macro render_tool_call(raw_tool_call) -%}
+    {%- if raw_tool_call.function is defined and raw_tool_call.function is mapping %}
+        {%- set tool_call = raw_tool_call.function %}
+    {%- else %}
+        {%- set tool_call = raw_tool_call %}
+    {%- endif %}
+    {{- '<tool_call>\n<function=' + (tool_call.name | default('') | string) + '>\n' }}
+    {%- if tool_call.arguments is defined and tool_call.arguments is mapping %}
+        {%- for args_name, args_value in tool_call.arguments.items() %}
+            {{- '<parameter=' + (args_name | string) + '>\n' }}
+            {%- if args_value is mapping or (args_value is sequence and args_value is not string) %}
+                {{- args_value | tojson | safe }}
+            {%- else %}
+                {{- args_value | string }}
+            {%- endif %}
+            {{- '\n</parameter>\n' }}
+        {%- endfor %}
+    {%- endif %}
+    {{- '</function>\n</tool_call>' }}
+{%- endmacro -%}
+{%- set system_message = none %}
+{%- if messages and messages[0]["role"] == "system" %}
+    {%- set system_message = messages[0]["content"] %}
+    {%- set loop_messages = messages[1:] %}
+{%- else %}
+    {%- set loop_messages = messages %}
+{%- endif %}
+{%- if not tools is defined %}
+    {%- set tools = [] %}
+{%- endif %}
+{%- set has_tools = tools is iterable and tools is not string and tools | length > 0 %}
+{%- if system_message is not none or has_tools %}
+    {{- '<|im_start|>system\n' }}
+    {%- if system_message is not none %}
+        {{- system_message }}
+    {%- else %}
+        {{- "You are Trinity Large, a helpful assistant developed by Arcee AI, that can interact with a computer to solve tasks." }}
+    {%- endif %}
+    {%- if has_tools %}
+        {{- "\n\n# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
+        {%- for tool in tools %}
+            {%- if tool.function is defined and tool.function is mapping %}
+                {%- set tool = tool.function %}
+            {%- endif %}
+            {{- '\n<function>\n<name>' ~ (tool.name | default('') | string) ~ '</name>' }}
+            {%- if tool.description is defined and tool.description is not none %}
+                {{- '\n<description>' ~ (tool.description | string | trim) ~ '</description>' }}
+            {%- endif %}
+            {{- '\n<parameters>' }}
+            {%- if tool.parameters is defined and tool.parameters is mapping and tool.parameters.properties is defined and tool.parameters.properties is mapping %}
+                {%- for param_name, param_fields in tool.parameters.properties.items() %}
+                    {{- '\n<parameter>\n<name>' ~ (param_name | string) ~ '</name>' }}
+                    {%- if param_fields is mapping and param_fields.type is defined and param_fields.type is not none %}
+                        {{- '\n<type>' ~ (param_fields.type | string) ~ '</type>' }}
+                    {%- endif %}
+                    {%- if param_fields is mapping and param_fields.description is defined and param_fields.description is not none %}
+                        {{- '\n<description>' ~ (param_fields.description | string | trim) ~ '</description>' }}
+                    {%- endif %}
+                    {%- if param_fields is mapping %}
+                        {%- set handled_keys = ['name', 'type', 'description'] %}
+                        {{- render_extra_keys(param_fields, handled_keys) }}
+                    {%- endif %}
+                    {{- '\n</parameter>' }}
+                {%- endfor %}
+            {%- endif %}
+            {%- if tool.parameters is defined %}
+                {%- set handled_keys = ['type', 'properties'] %}
+                {{- render_extra_keys(tool.parameters, handled_keys) }}
+            {%- endif %}
+            {{- '\n</parameters>' }}
+            {%- set handled_keys = ['type', 'name', 'description', 'parameters'] %}
+            {{- render_extra_keys(tool, handled_keys) }}
+            {{- '\n</function>' }}
+        {%- endfor %}
+        {{- "\n</tools>" }}
+        {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
+    {%- endif %}
+    {{- '<|im_end|>\n' }}
+{%- endif %}
+{%- for message in loop_messages %}
+    {%- set role = message.role | default('') %}
+    {%- if role == "assistant" %}
+        {%- set content_str = '' if message.content is none else (message.content | string) %}
+        {%- set trimmed_content = content_str | trim %}
+        {%- set has_reasoning_content = message.reasoning_content is defined %}
+        {%- set has_reasoning = has_reasoning_content or (message.reasoning is defined) %}
+        {%- if has_reasoning_content %}
+            {%- set reasoning_value = message.reasoning_content %}
+        {%- elif message.reasoning is defined %}
+            {%- set reasoning_value = message.reasoning %}
+        {%- else %}
+            {%- set reasoning_value = none %}
+        {%- endif %}
+        {%- set has_tool_calls = message.tool_calls is defined and message.tool_calls is iterable and message.tool_calls is not string and message.tool_calls | length > 0 %}
+        {{- '<|im_start|>assistant\n' }}
+        {%- if has_reasoning %}
+            {%- if reasoning_value %}
+                {{- '<think>' + (reasoning_value | string | trim) + '</think>' }}
+            {%- else %}
+                {{- '<think></think>' }}
+            {%- endif %}
+            {%- if trimmed_content %}
+                {{- '\n' + trimmed_content }}
+            {%- endif %}
+        {%- elif has_tool_calls %}
+            {%- if trimmed_content %}
+                {{- trimmed_content }}
+            {%- endif %}
+        {%- else %}
+            {{- content_str }}
+        {%- endif %}
+        {%- if has_tool_calls %}
+            {%- for tool_call in message.tool_calls %}
+                {%- set separator = '\n' if ((loop.first and (has_reasoning or trimmed_content)) or (not loop.first)) else '' -%}
+                {{- separator + render_tool_call(tool_call) }}
+            {%- endfor %}
+        {%- endif %}
+        {{- '<|im_end|>\n' }}
+    {%- elif role == "tool" or role == "observation" or role == "function" %}
+        {%- if loop.first or loop.previtem.role not in ["tool", "observation", "function"] %}
+            {{- '<|im_start|>user\n' }}
+        {%- endif %}
+        {{- '<tool_response>\n' }}
+        {{- '' if message.content is none else (message.content | string) }}
+        {{- '\n</tool_response>\n' }}
+        {%- if loop.last or loop.nextitem.role not in ["tool", "observation", "function"] %}
+            {{- '<|im_end|>\n' }}
+        {%- endif %}
+    {%- else %}
+        {{- '<|im_start|>' + (role | string) }}
+        {{- '\n' + ('' if message.content is none else (message.content | string)) }}
+        {{- '<|im_end|>\n' }}
+    {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+    {{- '<|im_start|>assistant\n<think>' }}
+{%- endif %}

config.json ADDED Viewed

	@@ -0,0 +1,515 @@

+{
+  "architectures": [
+    "AfmoeForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "auto_map": {
+    "AutoConfig": "configuration_afmoe.AfmoeConfig",
+    "AutoModel": "modeling_afmoe.AfmoeModel",
+    "AutoModelForCausalLM": "modeling_afmoe.AfmoeForCausalLM"
+  },
+  "dtype": "bfloat16",
+  "global_attn_every_n_layers": 4,
+  "head_dim": 128,
+  "hidden_act": "silu",
+  "hidden_size": 3072,
+  "initializer_range": 0.02,
+  "intermediate_size": 12288,
+  "layer_types": [
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "sliding_attention",
+    "full_attention"
+  ],
+  "load_balance_coeff": 5e-05,
+  "max_position_embeddings": 262144,
+  "model_type": "afmoe",
+  "moe_intermediate_size": 3072,
+  "mup_enabled": true,
+  "n_group": 1,
+  "num_attention_heads": 48,
+  "num_dense_layers": 6,
+  "num_expert_groups": 1,
+  "num_experts": 256,
+  "num_experts_per_tok": 4,
+  "num_hidden_layers": 60,
+  "num_key_value_heads": 8,
+  "num_limited_groups": 1,
+  "num_shared_experts": 1,
+  "quantization_config": {
+    "config_groups": {
+      "group_0": {
+        "format": "float-quantized",
+        "input_activations": {
+          "actorder": null,
+          "block_structure": null,
+          "dynamic": true,
+          "group_size": 128,
+          "num_bits": 8,
+          "observer": null,
+          "observer_kwargs": {},
+          "scale_dtype": null,
+          "strategy": "group",
+          "symmetric": true,
+          "type": "float",
+          "zp_dtype": null
+        },
+        "output_activations": null,
+        "targets": [
+          "Linear"
+        ],
+        "weights": {
+          "actorder": null,
+          "block_structure": [
+            128,
+            128
+          ],
+          "dynamic": false,
+          "group_size": null,
+          "num_bits": 8,
+          "observer": "memoryless_minmax",
+          "observer_kwargs": {},
+          "scale_dtype": null,
+          "strategy": "block",
+          "symmetric": true,
+          "type": "float",
+          "zp_dtype": null
+        }
+      }
+    },
+    "format": "float-quantized",
+    "global_compression_ratio": null,
+    "ignore": [
+      "model.layers.0.self_attn.q_proj",
+      "model.layers.0.self_attn.k_proj",
+      "model.layers.0.self_attn.v_proj",
+      "model.layers.0.self_attn.o_proj",
+      "model.layers.0.self_attn.gate_proj",
+      "model.layers.1.self_attn.q_proj",
+      "model.layers.1.self_attn.k_proj",
+      "model.layers.1.self_attn.v_proj",
+      "model.layers.1.self_attn.o_proj",
+      "model.layers.1.self_attn.gate_proj",
+      "model.layers.2.self_attn.q_proj",
+      "model.layers.2.self_attn.k_proj",
+      "model.layers.2.self_attn.v_proj",
+      "model.layers.2.self_attn.o_proj",
+      "model.layers.2.self_attn.gate_proj",
+      "model.layers.3.self_attn.q_proj",
+      "model.layers.3.self_attn.k_proj",
+      "model.layers.3.self_attn.v_proj",
+      "model.layers.3.self_attn.o_proj",
+      "model.layers.3.self_attn.gate_proj",
+      "model.layers.4.self_attn.q_proj",
+      "model.layers.4.self_attn.k_proj",
+      "model.layers.4.self_attn.v_proj",
+      "model.layers.4.self_attn.o_proj",
+      "model.layers.4.self_attn.gate_proj",
+      "model.layers.5.self_attn.q_proj",
+      "model.layers.5.self_attn.k_proj",
+      "model.layers.5.self_attn.v_proj",
+      "model.layers.5.self_attn.o_proj",
+      "model.layers.5.self_attn.gate_proj",
+      "model.layers.6.self_attn.q_proj",
+      "model.layers.6.self_attn.k_proj",
+      "model.layers.6.self_attn.v_proj",
+      "model.layers.6.self_attn.o_proj",
+      "model.layers.6.self_attn.gate_proj",
+      "model.layers.6.mlp.router.gate",
+      "model.layers.7.self_attn.q_proj",
+      "model.layers.7.self_attn.k_proj",
+      "model.layers.7.self_attn.v_proj",
+      "model.layers.7.self_attn.o_proj",
+      "model.layers.7.self_attn.gate_proj",
+      "model.layers.7.mlp.router.gate",
+      "model.layers.8.self_attn.q_proj",
+      "model.layers.8.self_attn.k_proj",
+      "model.layers.8.self_attn.v_proj",
+      "model.layers.8.self_attn.o_proj",
+      "model.layers.8.self_attn.gate_proj",
+      "model.layers.8.mlp.router.gate",
+      "model.layers.9.self_attn.q_proj",
+      "model.layers.9.self_attn.k_proj",
+      "model.layers.9.self_attn.v_proj",
+      "model.layers.9.self_attn.o_proj",
+      "model.layers.9.self_attn.gate_proj",
+      "model.layers.9.mlp.router.gate",
+      "model.layers.10.self_attn.q_proj",
+      "model.layers.10.self_attn.k_proj",
+      "model.layers.10.self_attn.v_proj",
+      "model.layers.10.self_attn.o_proj",
+      "model.layers.10.self_attn.gate_proj",
+      "model.layers.10.mlp.router.gate",
+      "model.layers.11.self_attn.q_proj",
+      "model.layers.11.self_attn.k_proj",
+      "model.layers.11.self_attn.v_proj",
+      "model.layers.11.self_attn.o_proj",
+      "model.layers.11.self_attn.gate_proj",
+      "model.layers.11.mlp.router.gate",
+      "model.layers.12.self_attn.q_proj",
+      "model.layers.12.self_attn.k_proj",
+      "model.layers.12.self_attn.v_proj",
+      "model.layers.12.self_attn.o_proj",
+      "model.layers.12.self_attn.gate_proj",
+      "model.layers.12.mlp.router.gate",
+      "model.layers.13.self_attn.q_proj",
+      "model.layers.13.self_attn.k_proj",
+      "model.layers.13.self_attn.v_proj",
+      "model.layers.13.self_attn.o_proj",
+      "model.layers.13.self_attn.gate_proj",
+      "model.layers.13.mlp.router.gate",
+      "model.layers.14.self_attn.q_proj",
+      "model.layers.14.self_attn.k_proj",
+      "model.layers.14.self_attn.v_proj",
+      "model.layers.14.self_attn.o_proj",
+      "model.layers.14.self_attn.gate_proj",
+      "model.layers.14.mlp.router.gate",
+      "model.layers.15.self_attn.q_proj",
+      "model.layers.15.self_attn.k_proj",
+      "model.layers.15.self_attn.v_proj",
+      "model.layers.15.self_attn.o_proj",
+      "model.layers.15.self_attn.gate_proj",
+      "model.layers.15.mlp.router.gate",
+      "model.layers.16.self_attn.q_proj",
+      "model.layers.16.self_attn.k_proj",
+      "model.layers.16.self_attn.v_proj",
+      "model.layers.16.self_attn.o_proj",
+      "model.layers.16.self_attn.gate_proj",
+      "model.layers.16.mlp.router.gate",
+      "model.layers.17.self_attn.q_proj",
+      "model.layers.17.self_attn.k_proj",
+      "model.layers.17.self_attn.v_proj",
+      "model.layers.17.self_attn.o_proj",
+      "model.layers.17.self_attn.gate_proj",
+      "model.layers.17.mlp.router.gate",
+      "model.layers.18.self_attn.q_proj",
+      "model.layers.18.self_attn.k_proj",
+      "model.layers.18.self_attn.v_proj",
+      "model.layers.18.self_attn.o_proj",
+      "model.layers.18.self_attn.gate_proj",
+      "model.layers.18.mlp.router.gate",
+      "model.layers.19.self_attn.q_proj",
+      "model.layers.19.self_attn.k_proj",
+      "model.layers.19.self_attn.v_proj",
+      "model.layers.19.self_attn.o_proj",
+      "model.layers.19.self_attn.gate_proj",
+      "model.layers.19.mlp.router.gate",
+      "model.layers.20.self_attn.q_proj",
+      "model.layers.20.self_attn.k_proj",
+      "model.layers.20.self_attn.v_proj",
+      "model.layers.20.self_attn.o_proj",
+      "model.layers.20.self_attn.gate_proj",
+      "model.layers.20.mlp.router.gate",
+      "model.layers.21.self_attn.q_proj",
+      "model.layers.21.self_attn.k_proj",
+      "model.layers.21.self_attn.v_proj",
+      "model.layers.21.self_attn.o_proj",
+      "model.layers.21.self_attn.gate_proj",
+      "model.layers.21.mlp.router.gate",
+      "model.layers.22.self_attn.q_proj",
+      "model.layers.22.self_attn.k_proj",
+      "model.layers.22.self_attn.v_proj",
+      "model.layers.22.self_attn.o_proj",
+      "model.layers.22.self_attn.gate_proj",
+      "model.layers.22.mlp.router.gate",
+      "model.layers.23.self_attn.q_proj",
+      "model.layers.23.self_attn.k_proj",
+      "model.layers.23.self_attn.v_proj",
+      "model.layers.23.self_attn.o_proj",
+      "model.layers.23.self_attn.gate_proj",
+      "model.layers.23.mlp.router.gate",
+      "model.layers.24.self_attn.q_proj",
+      "model.layers.24.self_attn.k_proj",
+      "model.layers.24.self_attn.v_proj",
+      "model.layers.24.self_attn.o_proj",
+      "model.layers.24.self_attn.gate_proj",
+      "model.layers.24.mlp.router.gate",
+      "model.layers.25.self_attn.q_proj",
+      "model.layers.25.self_attn.k_proj",
+      "model.layers.25.self_attn.v_proj",
+      "model.layers.25.self_attn.o_proj",
+      "model.layers.25.self_attn.gate_proj",
+      "model.layers.25.mlp.router.gate",
+      "model.layers.26.self_attn.q_proj",
+      "model.layers.26.self_attn.k_proj",
+      "model.layers.26.self_attn.v_proj",
+      "model.layers.26.self_attn.o_proj",
+      "model.layers.26.self_attn.gate_proj",
+      "model.layers.26.mlp.router.gate",
+      "model.layers.27.self_attn.q_proj",
+      "model.layers.27.self_attn.k_proj",
+      "model.layers.27.self_attn.v_proj",
+      "model.layers.27.self_attn.o_proj",
+      "model.layers.27.self_attn.gate_proj",
+      "model.layers.27.mlp.router.gate",
+      "model.layers.28.self_attn.q_proj",
+      "model.layers.28.self_attn.k_proj",
+      "model.layers.28.self_attn.v_proj",
+      "model.layers.28.self_attn.o_proj",
+      "model.layers.28.self_attn.gate_proj",
+      "model.layers.28.mlp.router.gate",
+      "model.layers.29.self_attn.q_proj",
+      "model.layers.29.self_attn.k_proj",
+      "model.layers.29.self_attn.v_proj",
+      "model.layers.29.self_attn.o_proj",
+      "model.layers.29.self_attn.gate_proj",
+      "model.layers.29.mlp.router.gate",
+      "model.layers.30.self_attn.q_proj",
+      "model.layers.30.self_attn.k_proj",
+      "model.layers.30.self_attn.v_proj",
+      "model.layers.30.self_attn.o_proj",
+      "model.layers.30.self_attn.gate_proj",
+      "model.layers.30.mlp.router.gate",
+      "model.layers.31.self_attn.q_proj",
+      "model.layers.31.self_attn.k_proj",
+      "model.layers.31.self_attn.v_proj",
+      "model.layers.31.self_attn.o_proj",
+      "model.layers.31.self_attn.gate_proj",
+      "model.layers.31.mlp.router.gate",
+      "model.layers.32.self_attn.q_proj",
+      "model.layers.32.self_attn.k_proj",
+      "model.layers.32.self_attn.v_proj",
+      "model.layers.32.self_attn.o_proj",
+      "model.layers.32.self_attn.gate_proj",
+      "model.layers.32.mlp.router.gate",
+      "model.layers.33.self_attn.q_proj",
+      "model.layers.33.self_attn.k_proj",
+      "model.layers.33.self_attn.v_proj",
+      "model.layers.33.self_attn.o_proj",
+      "model.layers.33.self_attn.gate_proj",
+      "model.layers.33.mlp.router.gate",
+      "model.layers.34.self_attn.q_proj",
+      "model.layers.34.self_attn.k_proj",
+      "model.layers.34.self_attn.v_proj",
+      "model.layers.34.self_attn.o_proj",
+      "model.layers.34.self_attn.gate_proj",
+      "model.layers.34.mlp.router.gate",
+      "model.layers.35.self_attn.q_proj",
+      "model.layers.35.self_attn.k_proj",
+      "model.layers.35.self_attn.v_proj",
+      "model.layers.35.self_attn.o_proj",
+      "model.layers.35.self_attn.gate_proj",
+      "model.layers.35.mlp.router.gate",
+      "model.layers.36.self_attn.q_proj",
+      "model.layers.36.self_attn.k_proj",
+      "model.layers.36.self_attn.v_proj",
+      "model.layers.36.self_attn.o_proj",
+      "model.layers.36.self_attn.gate_proj",
+      "model.layers.36.mlp.router.gate",
+      "model.layers.37.self_attn.q_proj",
+      "model.layers.37.self_attn.k_proj",
+      "model.layers.37.self_attn.v_proj",
+      "model.layers.37.self_attn.o_proj",
+      "model.layers.37.self_attn.gate_proj",
+      "model.layers.37.mlp.router.gate",
+      "model.layers.38.self_attn.q_proj",
+      "model.layers.38.self_attn.k_proj",
+      "model.layers.38.self_attn.v_proj",
+      "model.layers.38.self_attn.o_proj",
+      "model.layers.38.self_attn.gate_proj",
+      "model.layers.38.mlp.router.gate",
+      "model.layers.39.self_attn.q_proj",
+      "model.layers.39.self_attn.k_proj",
+      "model.layers.39.self_attn.v_proj",
+      "model.layers.39.self_attn.o_proj",
+      "model.layers.39.self_attn.gate_proj",
+      "model.layers.39.mlp.router.gate",
+      "model.layers.40.self_attn.q_proj",
+      "model.layers.40.self_attn.k_proj",
+      "model.layers.40.self_attn.v_proj",
+      "model.layers.40.self_attn.o_proj",
+      "model.layers.40.self_attn.gate_proj",
+      "model.layers.40.mlp.router.gate",
+      "model.layers.41.self_attn.q_proj",
+      "model.layers.41.self_attn.k_proj",
+      "model.layers.41.self_attn.v_proj",
+      "model.layers.41.self_attn.o_proj",
+      "model.layers.41.self_attn.gate_proj",
+      "model.layers.41.mlp.router.gate",
+      "model.layers.42.self_attn.q_proj",
+      "model.layers.42.self_attn.k_proj",
+      "model.layers.42.self_attn.v_proj",
+      "model.layers.42.self_attn.o_proj",
+      "model.layers.42.self_attn.gate_proj",
+      "model.layers.42.mlp.router.gate",
+      "model.layers.43.self_attn.q_proj",
+      "model.layers.43.self_attn.k_proj",
+      "model.layers.43.self_attn.v_proj",
+      "model.layers.43.self_attn.o_proj",
+      "model.layers.43.self_attn.gate_proj",
+      "model.layers.43.mlp.router.gate",
+      "model.layers.44.self_attn.q_proj",
+      "model.layers.44.self_attn.k_proj",
+      "model.layers.44.self_attn.v_proj",
+      "model.layers.44.self_attn.o_proj",
+      "model.layers.44.self_attn.gate_proj",
+      "model.layers.44.mlp.router.gate",
+      "model.layers.45.self_attn.q_proj",
+      "model.layers.45.self_attn.k_proj",
+      "model.layers.45.self_attn.v_proj",
+      "model.layers.45.self_attn.o_proj",
+      "model.layers.45.self_attn.gate_proj",
+      "model.layers.45.mlp.router.gate",
+      "model.layers.46.self_attn.q_proj",
+      "model.layers.46.self_attn.k_proj",
+      "model.layers.46.self_attn.v_proj",
+      "model.layers.46.self_attn.o_proj",
+      "model.layers.46.self_attn.gate_proj",
+      "model.layers.46.mlp.router.gate",
+      "model.layers.47.self_attn.q_proj",
+      "model.layers.47.self_attn.k_proj",
+      "model.layers.47.self_attn.v_proj",
+      "model.layers.47.self_attn.o_proj",
+      "model.layers.47.self_attn.gate_proj",
+      "model.layers.47.mlp.router.gate",
+      "model.layers.48.self_attn.q_proj",
+      "model.layers.48.self_attn.k_proj",
+      "model.layers.48.self_attn.v_proj",
+      "model.layers.48.self_attn.o_proj",
+      "model.layers.48.self_attn.gate_proj",
+      "model.layers.48.mlp.router.gate",
+      "model.layers.49.self_attn.q_proj",
+      "model.layers.49.self_attn.k_proj",
+      "model.layers.49.self_attn.v_proj",
+      "model.layers.49.self_attn.o_proj",
+      "model.layers.49.self_attn.gate_proj",
+      "model.layers.49.mlp.router.gate",
+      "model.layers.50.self_attn.q_proj",
+      "model.layers.50.self_attn.k_proj",
+      "model.layers.50.self_attn.v_proj",
+      "model.layers.50.self_attn.o_proj",
+      "model.layers.50.self_attn.gate_proj",
+      "model.layers.50.mlp.router.gate",
+      "model.layers.51.self_attn.q_proj",
+      "model.layers.51.self_attn.k_proj",
+      "model.layers.51.self_attn.v_proj",
+      "model.layers.51.self_attn.o_proj",
+      "model.layers.51.self_attn.gate_proj",
+      "model.layers.51.mlp.router.gate",
+      "model.layers.52.self_attn.q_proj",
+      "model.layers.52.self_attn.k_proj",
+      "model.layers.52.self_attn.v_proj",
+      "model.layers.52.self_attn.o_proj",
+      "model.layers.52.self_attn.gate_proj",
+      "model.layers.52.mlp.router.gate",
+      "model.layers.53.self_attn.q_proj",
+      "model.layers.53.self_attn.k_proj",
+      "model.layers.53.self_attn.v_proj",
+      "model.layers.53.self_attn.o_proj",
+      "model.layers.53.self_attn.gate_proj",
+      "model.layers.53.mlp.router.gate",
+      "model.layers.54.self_attn.q_proj",
+      "model.layers.54.self_attn.k_proj",
+      "model.layers.54.self_attn.v_proj",
+      "model.layers.54.self_attn.o_proj",
+      "model.layers.54.self_attn.gate_proj",
+      "model.layers.54.mlp.router.gate",
+      "model.layers.55.self_attn.q_proj",
+      "model.layers.55.self_attn.k_proj",
+      "model.layers.55.self_attn.v_proj",
+      "model.layers.55.self_attn.o_proj",
+      "model.layers.55.self_attn.gate_proj",
+      "model.layers.55.mlp.router.gate",
+      "model.layers.56.self_attn.q_proj",
+      "model.layers.56.self_attn.k_proj",
+      "model.layers.56.self_attn.v_proj",
+      "model.layers.56.self_attn.o_proj",
+      "model.layers.56.self_attn.gate_proj",
+      "model.layers.56.mlp.router.gate",
+      "model.layers.57.self_attn.q_proj",
+      "model.layers.57.self_attn.k_proj",
+      "model.layers.57.self_attn.v_proj",
+      "model.layers.57.self_attn.o_proj",
+      "model.layers.57.self_attn.gate_proj",
+      "model.layers.57.mlp.router.gate",
+      "model.layers.58.self_attn.q_proj",
+      "model.layers.58.self_attn.k_proj",
+      "model.layers.58.self_attn.v_proj",
+      "model.layers.58.self_attn.o_proj",
+      "model.layers.58.self_attn.gate_proj",
+      "model.layers.58.mlp.router.gate",
+      "model.layers.59.self_attn.q_proj",
+      "model.layers.59.self_attn.k_proj",
+      "model.layers.59.self_attn.v_proj",
+      "model.layers.59.self_attn.o_proj",
+      "model.layers.59.self_attn.gate_proj",
+      "model.layers.59.mlp.router.gate",
+      "lm_head"
+    ],
+    "kv_cache_scheme": null,
+    "quant_method": "compressed-tensors",
+    "quantization_status": "compressed",
+    "sparsity_config": {},
+    "transform_config": {},
+    "version": "0.14.0"
+  },
+  "rms_norm_eps": 1e-05,
+  "rope_scaling": null,
+  "rope_theta": 10000,
+  "route_norm": true,
+  "route_scale": 2.448,
+  "score_func": "sigmoid",
+  "sliding_window": 4096,
+  "tie_word_embeddings": false,
+  "topk_group": 1,
+  "transformers_version": "4.57.6",
+  "use_cache": true,
+  "use_grouped_mm": true,
+  "vocab_size": 200192
+}

configuration_afmoe.py ADDED Viewed

	@@ -0,0 +1,133 @@

+# coding=utf-8
+# Copyright 2022 EleutherAI and the HuggingFace Inc. team. All rights reserved.
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+#     http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+from transformers.configuration_utils import PretrainedConfig
+from transformers.modeling_rope_utils import rope_config_validation
+from transformers.configuration_utils import layer_type_validation
+from transformers.utils import logging
+logger = logging.get_logger(__name__)
+class AfmoeConfig(PretrainedConfig):
+    """
+    n_group (`int`, *optional*, defaults to 1):
+            Number of groups for routed experts.
+    topk_group (`int`, *optional*, defaults to 1):
+        Number of selected groups for each token(for each token, ensuring the selected experts is only within `topk_group` groups).
+    """
+    model_type = "afmoe"
+    base_model_pp_plan = {
+        "embed_tokens": (["input_ids"], ["inputs_embeds"]),
+        "layers": (["hidden_states", "attention_mask"], ["hidden_states"]),
+        "norm": (["hidden_states"], ["hidden_states"]),
+    }
+    def __init__(
+        self,
+        num_hidden_layers: int = 32,
+        vocab_size: int = 200192,
+        hidden_size: int = 2048,
+        intermediate_size: int = 6144,
+        moe_intermediate_size=1408,
+        num_dense_layers=1,
+        num_attention_heads=16,
+        num_key_value_heads=None,
+        head_dim=128,
+        hidden_act="silu",
+        max_position_embeddings=16384,
+        initializer_range=0.02,
+        rms_norm_eps=1e-5,
+        use_cache=True,
+        tie_word_embeddings=False,
+        rope_theta=10000.0,
+        rope_scaling=None,
+        num_experts=64,
+        num_experts_per_tok=6,
+        num_shared_experts=2,
+        num_expert_groups=1,
+        num_limited_groups=1,
+        score_func="sigmoid",
+        route_norm=True,
+        route_scale=1.0,
+        global_attn_every_n_layers=4,
+        sliding_window=1024,
+        mup_enabled=False,
+        layer_types=None,
+        attention_dropout: float = 0.0,
+        n_group: int = 1,
+        topk_group: int = 1,
+        **kwargs,
+    ):
+        self.vocab_size = vocab_size
+        self.max_position_embeddings = max_position_embeddings
+        self.hidden_size = hidden_size
+        self.intermediate_size = intermediate_size
+        self.num_hidden_layers = num_hidden_layers
+        self.num_dense_layers = num_dense_layers
+        self.num_attention_heads = num_attention_heads
+        self.head_dim = head_dim
+        self.hidden_act = hidden_act
+        self.initializer_range = initializer_range
+        self.rms_norm_eps = rms_norm_eps
+        self.use_cache = use_cache
+        self.rope_theta = rope_theta
+        self.rope_scaling = rope_scaling
+        # MoE specific
+        self.moe_intermediate_size = moe_intermediate_size
+        self.num_experts_per_tok = num_experts_per_tok
+        self.n_group = n_group
+        self.topk_group = topk_group
+        self.num_experts = num_experts
+        self.num_shared_experts = num_shared_experts
+        self.num_expert_groups = num_expert_groups
+        self.num_limited_groups = num_limited_groups
+        self.score_func = score_func
+        self.route_norm = route_norm
+        self.route_scale = route_scale
+        # Attention specific
+        self.attention_dropout = attention_dropout
+        self.global_attn_every_n_layers = global_attn_every_n_layers
+        self.sliding_window = sliding_window
+        self.layer_types = layer_types
+        if self.layer_types is None:
+            self.layer_types = [
+                "sliding_attention" if bool((i + 1) % global_attn_every_n_layers) else "full_attention" for i in range(self.num_hidden_layers)
+            ]
+        layer_type_validation(self.layer_types)
+        # muP specific
+        self.mup_enabled = mup_enabled
+        if num_key_value_heads is None:
+            num_key_value_heads = num_attention_heads
+        self.num_key_value_heads = num_key_value_heads
+        # Validate rope configs
+        if self.rope_scaling is not None and "type" in self.rope_scaling:
+            self.rope_scaling["rope_type"] = self.rope_scaling["type"]
+        rope_config_validation(self)
+        super().__init__(
+            tie_word_embeddings=tie_word_embeddings,
+            **kwargs,
+        )
+__all__ = ["AfmoeConfig"]

generation_config.json ADDED Viewed

	@@ -0,0 +1,10 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 0,
+  "do_sample": true,
+  "eos_token_id": 3,
+  "pad_token_id": 12,
+  "temperature": 0.8,
+  "top_p": 0.8,
+  "transformers_version": "4.57.6"
+}

model-00001-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:17d8032fafe6ea1bbc1a3965138b55f922b6408bc488fd2c27c478e46b5df905
+size 4991297952

model-00002-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:39ec8ac900fe6451abaa9d2780dd9e743974515cadf00e8f1304b39ab8b77dc0
+size 4993010240

model-00003-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:fdc9bbc9d2455c73731ccc13921db2c71794e242f2aa39417e009d5b3814a36b
+size 4997737720

model-00004-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:62c0c5b1b44912f21e2b5425f156077a088ab9216b8fa5dc7b4247d04a02a7d6
+size 4997737816

model-00005-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4383e929baa767a1e9b8cc792d7cef0f1abc1a0084a9ac6a990e93a6250a0c14
+size 4986706032

model-00006-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:99121544dece1dd95fbe4fa4905545459ee41b40786aba22dbb636f8e5b3f6d7
+size 4994603392

model-00007-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5a40a85706d10934eb1675d0e0e4bd88ff08cd1e882d3908232cd3bc45799884
+size 4997739824

model-00008-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:a96a8307c1006fea6de61ed1138fdf16a5b08de7169edd857b171b93fe70cbe2
+size 4997739304

model-00009-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:3bd40717a9ec81cce1932357e865d129283dc32ddb9ac1e6b44dd8bb14a41f94
+size 4993010824

model-00010-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:aaba8ad2a72a27261b35ef0a9e50e5814e2398c964c6d97d0d03dc2d2ac4ba40
+size 4997738792

model-00011-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8b68dd29426c4fc0ebfc87a531d914d095ae819627bd7ede838399973db45722
+size 4997739272

model-00012-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8a9feb4614755537a72b6322eed3c97a000ae1b93b087aeda7ab9b5727f11854
+size 4993010896

model-00013-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:9ede719658d9234a7cb0d658db8012147b69a23b4bdf0a63ff1b63d4c3f814d4
+size 4997738760

model-00014-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:cdb3ad46440df9b3e463d33039025a166a8627536c12eeb7f740a23444dd46df
+size 4997739264

model-00015-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:75fad5dc16ce440814451503324a4133ac95a8622f9f31780c74bedbaf2d39c0
+size 4993010936

model-00016-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8c6bf34ab60420a9f66bdce6095e27a76ff416e0eb6af9e715c160a5f42027b0
+size 4997738760

model-00017-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:162fea404b03d47710535c13fcd1d20ad3d58aa84479b0da6f65b65a763022a0
+size 4997739224

model-00018-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:b02f27369b041f7d0b12a2c0234bdb6e5b4c098308c9f19939f17b38ebfe1db5
+size 4993010976

model-00019-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:07a35f6ca6504457e972d90dd22902c6a0096767e3bdf37ccd018ab6cb73b563
+size 4997738760

model-00020-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:cd803f7dd68083b8b3a565d54caa5189cd789456c2944d4c7ce1d5b15bc42c12
+size 4997739184

model-00021-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6a5178453852c3d2bbab8133727481e7b19bd1af69ed7a475892c250ae0679ea
+size 4993011008

model-00022-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:cadae97c99fd231eacc79239370348b217d57025d0fe097a9b2db3f805b3abdb
+size 4997738760

model-00023-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:39f5fbdd7fedfae928386b256af0fc836170a594842426e52694c9078ecc6017
+size 4997739144

model-00024-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:aaf5184aa4e40926044e5c557f1e2f626094a02f5e8ff2a78c8b910b47520b1f
+size 4993011048

model-00025-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e581269653588c4164376bdc1333752b03e76aadb75dfaf98d11df6e4be1818d
+size 4997738760

model-00026-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:fd20a36515acfa8db4377377839ba990a10151d4f5add0678889dce2ea04a423
+size 4997739112

model-00027-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c80e9cf5599b8c2bcee50bf55063e87fd6f6119c0375d7ec9119ed63e766d231
+size 4993011088

model-00028-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e0c44491a88db4dba47b7484d856e1f6c747e574c71698ca0ac20e9a54ea0a2d
+size 4997738760

model-00029-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:2bef43f5d4d5dd21adf28358117d771ed045b8823377ae0e980955ca3bc9859e
+size 4997739072

model-00030-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:9cfb0626fa9f84397c4a407f42c11d6b87288e992cd96afde05aa2d7908a938a
+size 4993011120

model-00031-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:30935b28748c0063919074cefe1ded91bcd3d56b399418831cbde18ce96a11af
+size 4997738760

model-00032-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d465e3f331560851ace35d861977e8feb7119261f2a920d172dc6006d6ba7e27
+size 4997739032

model-00033-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7673aad9fdceb1adc0a78c4cf237e51bb743287fa9e863dbe7a7e4051510f88f
+size 4993011160

model-00034-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:47824d2fc6f13f2f031da0b62154fb2db120f77317a02059172b1069c3109844
+size 4997738768

model-00035-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7cadabe2c5219bbeb220fb0f792a5e8883c747087935179effb19d9d1f3375c3
+size 4997738992

model-00036-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:013dcb25788a5cea80af5fbf571baae030f54ca0b12d4fa20af45d7f96d9b18c
+size 4993011200

model-00037-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d50335896c003d70ed9e5b1d02a6631b49921ac2ef9976a31824b8867eaf2cfa
+size 4997738768

model-00038-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8d6acdce3dc22c7b71928378e0c0777e378282f0dc012d215127ee433de13740
+size 4997738960

model-00039-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:b6da2e35531d607dd42f45ad79b3b85ada4e9b61f04a48c613e59ff8324ce9a1
+size 4993011232

model-00040-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:acf7f08c463a6f6a52105671316fcece7b0fdde757ec4b3ea94bee9e03d5161f
+size 4997738768

model-00041-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e7e7ce29a81ce5fd7fe6e802a9358c9ab9f5eb158bf1467d5446be39c69ac83b
+size 4997738920

model-00042-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:a55b030a9c3f082fb853e8917ce33aa41de2f1b1521089eb768851693aa4acb2
+size 4993011280

model-00043-of-00081.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:b46484b4cd664c8da178e790b2cfbfbb5b82a5d42e38add0007c205d7a9dc3d8
+size 4997738768