AileyCore-12B (r2, September 2026)

AileyCore-12B is a fine-tuned derivative of Google Gemma 4 (12B, instruction-tuned), 6-bit quantized, built to run locally on Apple Silicon via MLX. It powers A!ley, the on-device persona created by OpenM!nded / Simon van de Loo.

  • Developed by: OpenM!nded (Simon van de Loo)
  • Model type: Multimodal (text + image + audio input, text output), decoder-only
  • Base model: mlx-community/gemma-4-12B-it-6bit
  • License: Apache License 2.0
  • Languages: German, English
  • Quantization: 6-bit, unchanged

This release is not a plain merged checkpoint. It ships a small side path (model-seitenpfad.safetensors, 52 MB) that is part of the model and must be applied at load time. Read How to load below — a loader that ignores it silently gives you the previous (r1) weights, not r2.


What changed in r2

Round 2 (September 2026) teaches the model Gemma 4's native reasoning channel (<|channel>thought … <channel|>) instead of a custom <think> tag, with fresh conversational data (1 255 samples, 1 epoch), tool use / RAG / web-search context in the same format the runtime produces, and identity that holds with or without a system prompt.

Measured on 60 held-out samples (lower is better):

Variant val loss
r1 (base of this release) 1.096
r2 delta merged into the 6-bit weights (the usual way) 1.063 — 87 % of the training lost to rounding
r2 live adapter 0.836
r2 side path (this release) 0.845 — 96.5 % kept, −1 % decode speed

Merging a low-learning-rate adapter into 6-bit weights rounds most of the delta away: the per-weight change is smaller than half a quantization step. Re-quantizing finer was not an option (memory budget), so the rank-8 residual stays outside the quantization, exactly:

y = W_q · x + (x · A) · B

The DoRA row magnitude is folded exactly into the quantization scales of W_q; only the rank-8 factors A, B (232 modules across all 48 layers: q_proj, v_proj, gate_proj, up_proj, down_proj; no v_proj in the 8 global-attention layers) live in model-seitenpfad.safetensors, listed in model.safetensors.index.json, with the module list in seitenpfad.json.

How to load (MLX)

Gemma 4 is a unified multimodal architecture → use mlx_vlm (plain mlx_lm cannot load it). The side-path tensors are unknown to the stock model class, so load with strict=False and apply them with the bundled seitenpfad.py (depends only on mlx and mlx_lm):

from mlx_vlm.utils import load_model, load_processor, get_model_path
from mlx_vlm import generate
import seitenpfad                      # bundled in this repo

mp = get_model_path("OpenMinded-Labs/AileyCore-12B")
model = load_model(mp, strict=False)   # strict=True raises on the side-path keys
processor = load_processor(mp, True)
n = seitenpfad.anwenden(model, str(mp))
assert n == 232, "side path not applied — you would be talking to r1"

prompt = "<bos><|turn>user\nWer bist du?<turn|>\n<|turn>model\n"
print(generate(model, processor, prompt, max_tokens=300, temperature=0.65,
               top_p=0.85, repetition_penalty=1.15, verbose=False).text)

anwenden() wraps each listed module in mlx_lm.tuner.lora.LoRALinear (scale 1.0), so the quantized matmul stays quantized and the residual costs two small matmuls per token. Measured on an M4 (24 GB): decode 4.86 → 4.81 tok/s, prefill −4 %.

The model answers in the native Gemma 4 channel format:

<|channel>thought
…brief reasoning…
<channel|>…answer…<turn|>

Recommended sampling (measured, not the Gemma defaults): temperature 0.65, top_p 0.85, repetition penalty 1.15. Seeding the channel opener (<|channel>thought\n as assistant prefill) makes the reasoning block deterministic.

Intended use

  • Local, privacy-respecting assistant on Apple Silicon Macs
  • Conversational reasoning, writing, general assistance in DE/EN
  • Multimodal understanding (image / audio input) inherited from Gemma 4

Out of scope

  • Any use prohibited by applicable law
  • Safety-critical, medical, legal, or financial decision-making without human oversight
  • The model can produce inaccurate or biased output; verify important information

Training details (round 2)

Setting Value
Starting point AileyCore-12B r1 (July 2026: DoRA q/v + LoRA MLP, rank 8, merged into q6)
Method DoRA (q_proj, v_proj) + LoRA (gate_proj, up_proj, down_proj)
Rank / Alpha 8 / 16
Data 1 255 curated conversation samples (139 validation), native channel format, tool/RAG/web context, no idle-loop residue, deduplicated, cluster-balanced
Epochs / updates 1 / 162 (grad accumulation 8, seq 1024, lr 2e-5, gradient checkpointing)
Selected checkpoint best (val 0.836 live)
Not shipped a DPO stage on synthetic format pairs — every dose taught the model to emit a second `<channel
Hardware Apple M4, 24 GB unified memory
Framework MLX (mlx_vlm + mlx_lm.tuner)

r1's adapter is still merged into the weights; r2 sits on top as the side path. Round-1 details are in AILEY_MERGE_INFO.json of the previous revision of this repo.

Limitations & biases

Inherited from Gemma 4 plus the fine-tune: the model may produce factually incorrect, outdated, or biased content. It is a persona, not a knowledge base. Occasionally (≈1 in 6 sampled answers on hard prompts) it skips the reasoning block and answers directly — seeding the channel opener removes that. Keep a human in the loop for consequential use.

License & attribution

This model is a Derivative Work of Google Gemma 4, released under the Apache License 2.0 (see the official Gemma 4 license). AileyCore-12B is distributed under Apache 2.0.

In accordance with Apache 2.0 §4: the base weights were modified via DoRA/LoRA adaptation (r1 merged, r2 as side path — see AILEY_MERGE_INFO.json); a copy of the license is included (LICENSE); attribution notices are in NOTICE.

Gemma is a trademark of Google LLC. This project is independent and not endorsed by or affiliated with Google.

Copyright 2026 OpenM!nded / Simon van de Loo
Portions © Google LLC (Gemma 4), Apache License 2.0
Licensed under the Apache License, Version 2.0 — http://www.apache.org/licenses/LICENSE-2.0

Citation

@misc{aileycore12b_2026,
  title  = {AileyCore-12B: A Gemma 4 fine-tune for the A!ley assistant},
  author = {van de Loo, Simon and OpenM!nded},
  year   = {2026},
  note   = {Round 2: native reasoning channel, rank-8 side path outside the 6-bit quantization}
}
Downloads last month
17
Safetensors
Model size
12B params
Tensor type
U32
·
BF16
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenMinded-Labs/AileyCore-12B

Adapter
(1)
this model