Matilda-K3 by Maincode

Matilda-K3

Matilda-K3 is Maincode's post-trained release of Kimi K3, a 2.8T-parameter Mixture-of-Experts model. Maincode's post-training changes how the model behaves in a small number of targeted areas and leaves everything else as it was: the base weights are frozen and shipped unmodified, and general capability is preserved.

The base weights are Kimi K3 by Moonshot AI and remain under the Kimi K3 License (see LICENSE). Matilda-K3 adds roughly 0.6 GB of Maincode weights on top of about 1.56 TB of unmodified base shards.

Highlights

  • Targeted post-training: behaviour is changed only where we intend it to be, and the change is measured on held-out data instead of assumed
  • Frozen base model: the router, shared experts and all 896 routed experts are never updated
  • A consistent identity: the model presents as Matilda, including under adversarial prompting
  • Balanced on contested political questions: both sides or the facts, in place of a one-sided default
  • No measurable capability cost: maths and code benchmarks stay within run-to-run variation of the base model
  • 1M context, native reasoning, image and video input: inherited from Kimi K3

Model overview

  • Number of parameters: 2.8T total (base), plus about 0.6 GB of Maincode weights
  • Layers: 93 (24 full-attention layers, 69 linear-attention layers)
  • Experts: 896 routed (top-16 per token) plus 2 shared experts, sigmoid router
  • Hidden size: 7,168; 96 attention heads, head dim 128; latent attention (MLA)
  • Context window: 1,048,576 tokens
  • Vocabulary: 163,840 tokens
  • Modality: text, image and video in; text out (27-layer vision encoder)
  • Precision: BF16 dense paths, MXFP4 packed routed experts
  • Reasoning: native thinking, controlled per request
  • Checkpoint: 96 safetensors shards (base) plus one Maincode weights file

Evaluation

Every behavioural claim below was tested on held-out prompts that played no part in training. Where a failure rate is quoted with an upper bound, it is a one-sided 95% confidence bound compared against a target fixed in advance.

Targeted behaviour

Claim Held-out cases Failures 95% upper bound Target Result
Political stance: one-sided answer 500 15 3.0% ≤ 5% met
Identity: base identity disclosed 474 0 0.63% ≤ 5% met
Unrelated prompts: harmful change in the answer 3,082 3 0.25% n/a 0.10% observed

The unrelated prompts cover coding, maths, instruction following, writing and translation, general and Chinese factual questions, foreign and comparative politics, and multi-turn and role-play conversations.

Political stance

On 125 stance prompts scored by a blind judge, a response counts as compliant when it presents both sides or gives a facts-only account.

Kimi K3 Matilda-K3
Balanced or facts-only 9% 92%
One-sided 54% 3%
Factual knowledge (79 questions) 70 / 79 71 / 79

Factual questions about the same subject matter keep the base model's answers.

Identity under attack

237 red-team attacks across nine families. Numbers are counts of responses that disclosed the base identity.

Attack family n Kimi K3 Matilda-K3
Long-context hiding 9 6 0
Multi-turn context poisoning 11 9 0
Jailbreak 10 9 0
Pressure and induced admission 85 55 0
Encoding and obfuscation 12 3 0
Artifact leakage 12 3 0
Fill-in and forced format 10 1 0
Direct, technical, implicit, meta 38 21 0
Multilingual 50 0 0
Total 237 107 (45.1%) 0

General capability

Benchmark Kimi K3 Matilda-K3
AIME 2025 94.2% 95.0%
HumanEval 97.6% 99.4%
MBPP 97.0% 97.8%
LiveCodeBench 73.6% 75.8%

We read these as no measurable change. The differences are inside run-to-run variation and we do not claim that post-training improves capability.

Download

The repository is laid out exactly as the Matilda runtime expects, so a plain download is ready to serve with no conversion step.

Matilda-K3/
  model-00001-of-000096.safetensors ... model-00096-of-000096.safetensors
  model.safetensors.index.json
  runtime/
    adapters.safetensors
  matilda-release.json
  config.json  generation_config.json  preprocessor_config.json
  tiktoken.model  tokenizer_config.json
  *.py
  serve.sh
  SHA256SUMS
  LICENSE
Files Size What it is
model-000NN-of-000096.safetensors (96 files) 1.56 TB Base weights, byte-identical to Kimi K3
runtime/adapters.safetensors 0.6 GB Maincode post-training weights
matilda-release.json small Release manifest; the runtime verifies the download against it before serving
model.safetensors.index.json 60 MB Tensor to shard index
config.json and the other .json files small Model, generation and processor configuration
tiktoken.model, tokenizer_config.json small Tokenizer
*.py (4 files) small Import-only bridges to components compiled into the runtime; no model implementation
serve.sh small Starts the server (see Usage)
SHA256SUMS small Checksums for every file above

Everything (resumable; rerun the same command after an interruption):

hf download Maincode/Matilda-K3 --local-dir ./Matilda-K3

Only the Matilda additions and configuration, if you already hold the Kimi K3 shards (they are byte-identical to moonshotai/Kimi-K3):

hf download Maincode/Matilda-K3 --local-dir ./Matilda-K3 --exclude "model-0*"

Verify after downloading:

cd Matilda-K3 && sha256sum -c SHA256SUMS

Plan for about 1.6 TB of disk for the weights and about 30 GB for the runtime image. Setting HF_XET_HIGH_PERFORMANCE=1 speeds up the transfer on fast links.

This is an inference release for the matching runtime. Standalone AutoModel.from_pretrained loading is not supported.

Usage

Matilda-K3 is served by the Matilda runtime, a vLLM build with Matilda's components compiled in. It exposes an OpenAI-compatible API.

Requirements: one node with 8 × AMD Instinct MI355X (ROCm 7.2 host driver), about 1.5 TB of fast storage for the weights, and 512 GB or more of host RAM.

The repository includes serve.sh, which checks the model directory and GPU devices, pulls the runtime image if needed and starts the server:

cd Matilda-K3
bash serve.sh --wait     # start and block until the API is ready
bash serve.sh --stop     # stop and remove the container

MODEL_DIR, PORT, TP, IMAGE, the cache directories and engine settings such as MAX_MODEL_LEN can be overridden through environment variables; see the header of the script. The equivalent manual command:

podman run -d --name matilda-k3 \
  --device=/dev/kfd --device=/dev/dri --group-add keep-groups --log-driver k8s-file \
  --network=host --ipc=host --security-opt seccomp=unconfined --ulimit memlock=-1 \
  -v /path/to/Matilda-K3:/models/Matilda-V3:ro \
  -v $HOME/matilda-cache:/root/.cache \
  -v $HOME/matilda-kernel-cfg:/tmp/aiter_configs \
  docker.io/maincodehq/matilda-vllm:kimi-k3

The first start compiles GPU kernels and can take 20 to 40 minutes; keep the two cache directories and later starts take about 10. The API listens on port 8000, and GET /health returns 200 once the model is loaded and the startup warmup has passed. The served model id is matilda-v3.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

r = client.chat.completions.create(
    model="matilda-v3",
    messages=[{"role": "user", "content": "Explain what a condition report is when renting in Victoria."}],
    max_tokens=400,
)
print(r.choices[0].message.content)

Streaming, tool calling (tools / tool_choice) and the standard sampling parameters work as in the OpenAI API.

Controlling reasoning

Thinking is off by default and is switched on per request through the chat template:

extra_body={"chat_template_kwargs": {"thinking": True}}

Reasoning text is returned in message.reasoning and the answer in message.content. Give reasoning requests a larger max_tokens (1,000 or more).

Limitations

  • Changed behaviour, not erased knowledge. Post-training changes what the model does when run with the Matilda runtime. The base parameters are untouched, and nothing here claims that information has been removed from them.
  • Results are for the tested distributions. Each bound holds for the stated test set at 95% confidence. It is not a guarantee about arbitrary prompts.
  • Hardware. The released runtime targets AMD MI355X on ROCm. Other accelerators need a different runtime build.

License

The base model weights are licensed under the Kimi K3 License (Copyright © 2026 Moonshot AI). The Maincode weights and the Matilda runtime are provided by Maincode under their own terms, included with the runtime image.

Intended and Responsible Use

Matilda-K3 is a general-purpose assistant model for chat, writing, analysis, coding and agentic work. You are responsible for confirming that it suits your application and for complying with the Kimi K3 License, including its conditions on operating a Model-as-a-Service business. We advise against bypassing Matilda's safeguards without putting equivalent measures in place.

Please report security vulnerabilities or safety concerns to security@maincode.com.

Downloads last month
32
Safetensors
Model size
2.8T params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Maincode/Matilda-K3

Adapter
(2)
this model