--- license: other license_name: openmdw-1.1 license_link: LICENSE base_model: arcee-ai/Trinity-Large-Base library_name: transformers pipeline_tag: text-generation --- This is a non-reasoning instruction and creative-writing fine-tune of [Arcee AI's Trinity-Large-Base](https://huggingface.co/arcee-ai/Trinity-Large-Base). This repository contains the merged BF16 checkpoint, not a standalone LoRA adapter. ## Training - LoRA rank: 64 - LoRA alpha: 128 - Context length: 32,768 - Learning rate: 8e-6 - Schedule: cosine - Weight decay: 0.0001 - Maximum gradient norm: 1.0 - Final checkpoint: step 523 The training mix intentionally contains non-reasoning instruction, roleplay, and creative-writing data: - `PocketDoc/Dans-Kinomaxx-VanillaBackrooms` - `PocketDoc/Dans-Personamaxx-Logs-2` - `PocketDoc/Dans-Prosemaxx-RepRemover-1` - `PocketDoc/Dans-Failuremaxx-Adventure-3` - `Delta-Vector/Hydrus-Claude-Instruct-2.7K` - `Delta-Vector/Hydrus-Claude-Instruct-5K` - `anthracite-org/kalo-opus-instruct-22k-no-refusal` - `anthracite-org/nopm_claude_writing_fixed` - `anthracite-org/kalo_opus_misc_240827` - `anthracite-org/kalo_misc_part2` - `Epiculous/SynthRP-Gens-v1.1-Filtered-n-Cleaned` - `Epiculous/Synthstruct-Gens-v1.1-Filtered-n-Cleaned` - `Delta-Vector/Orion-Sonnet-CharCard` ## Format The checkpoint uses the tokenizer and ChatML template saved by the final training checkpoint. The template supports `system`, `user`, and `assistant` messages and uses `<|im_end|>` as EOS and padding. Trinity-Large is a 398B-parameter sparse mixture-of-experts model with roughly 13B active parameters per token. Serving the BF16 checkpoint requires multiple GPUs. ## License This retains the base model's OpenMDW 1.1 license. See `LICENSE`.