Instructions to use OpenMinded-Labs/AileyCore-12B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use OpenMinded-Labs/AileyCore-12B with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("OpenMinded-Labs/AileyCore-12B") config = load_config("OpenMinded-Labs/AileyCore-12B") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use OpenMinded-Labs/AileyCore-12B with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "OpenMinded-Labs/AileyCore-12B"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "OpenMinded-Labs/AileyCore-12B" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use OpenMinded-Labs/AileyCore-12B with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "OpenMinded-Labs/AileyCore-12B"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default OpenMinded-Labs/AileyCore-12B
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use OpenMinded-Labs/AileyCore-12B with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "OpenMinded-Labs/AileyCore-12B"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "OpenMinded-Labs/AileyCore-12B" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
AileyCore-12B (r2, September 2026)
AileyCore-12B is a fine-tuned derivative of Google Gemma 4 (12B, instruction-tuned), 6-bit quantized, built to run locally on Apple Silicon via MLX. It powers A!ley, the on-device persona created by OpenM!nded / Simon van de Loo.
- Developed by: OpenM!nded (Simon van de Loo)
- Model type: Multimodal (text + image + audio input, text output), decoder-only
- Base model:
mlx-community/gemma-4-12B-it-6bit - License: Apache License 2.0
- Languages: German, English
- Quantization: 6-bit, unchanged
This release is not a plain merged checkpoint. It ships a small side path (
model-seitenpfad.safetensors, 52 MB) that is part of the model and must be applied at load time. Read How to load below — a loader that ignores it silently gives you the previous (r1) weights, not r2.
What changed in r2
Round 2 (September 2026) teaches the model Gemma 4's native reasoning channel
(<|channel>thought … <channel|>) instead of a custom <think> tag, with fresh
conversational data (1 255 samples, 1 epoch), tool use / RAG / web-search context in the
same format the runtime produces, and identity that holds with or without a system prompt.
Measured on 60 held-out samples (lower is better):
| Variant | val loss |
|---|---|
| r1 (base of this release) | 1.096 |
| r2 delta merged into the 6-bit weights (the usual way) | 1.063 — 87 % of the training lost to rounding |
| r2 live adapter | 0.836 |
| r2 side path (this release) | 0.845 — 96.5 % kept, −1 % decode speed |
Merging a low-learning-rate adapter into 6-bit weights rounds most of the delta away: the per-weight change is smaller than half a quantization step. Re-quantizing finer was not an option (memory budget), so the rank-8 residual stays outside the quantization, exactly:
y = W_q · x + (x · A) · B
The DoRA row magnitude is folded exactly into the quantization scales of W_q; only the
rank-8 factors A, B (232 modules across all 48 layers: q_proj, v_proj, gate_proj,
up_proj, down_proj; no v_proj in the 8 global-attention layers) live in model-seitenpfad.safetensors, listed in
model.safetensors.index.json, with the module list in seitenpfad.json.
How to load (MLX)
Gemma 4 is a unified multimodal architecture → use mlx_vlm (plain mlx_lm cannot load it).
The side-path tensors are unknown to the stock model class, so load with strict=False and
apply them with the bundled seitenpfad.py (depends only on mlx and mlx_lm):
from mlx_vlm.utils import load_model, load_processor, get_model_path
from mlx_vlm import generate
import seitenpfad # bundled in this repo
mp = get_model_path("OpenMinded-Labs/AileyCore-12B")
model = load_model(mp, strict=False) # strict=True raises on the side-path keys
processor = load_processor(mp, True)
n = seitenpfad.anwenden(model, str(mp))
assert n == 232, "side path not applied — you would be talking to r1"
prompt = "<bos><|turn>user\nWer bist du?<turn|>\n<|turn>model\n"
print(generate(model, processor, prompt, max_tokens=300, temperature=0.65,
top_p=0.85, repetition_penalty=1.15, verbose=False).text)
anwenden() wraps each listed module in mlx_lm.tuner.lora.LoRALinear (scale 1.0), so
the quantized matmul stays quantized and the residual costs two small matmuls per token.
Measured on an M4 (24 GB): decode 4.86 → 4.81 tok/s, prefill −4 %.
The model answers in the native Gemma 4 channel format:
<|channel>thought
…brief reasoning…
<channel|>…answer…<turn|>
Recommended sampling (measured, not the Gemma defaults): temperature 0.65, top_p 0.85,
repetition penalty 1.15. Seeding the channel opener (<|channel>thought\n as assistant
prefill) makes the reasoning block deterministic.
Intended use
- Local, privacy-respecting assistant on Apple Silicon Macs
- Conversational reasoning, writing, general assistance in DE/EN
- Multimodal understanding (image / audio input) inherited from Gemma 4
Out of scope
- Any use prohibited by applicable law
- Safety-critical, medical, legal, or financial decision-making without human oversight
- The model can produce inaccurate or biased output; verify important information
Training details (round 2)
| Setting | Value |
|---|---|
| Starting point | AileyCore-12B r1 (July 2026: DoRA q/v + LoRA MLP, rank 8, merged into q6) |
| Method | DoRA (q_proj, v_proj) + LoRA (gate_proj, up_proj, down_proj) |
| Rank / Alpha | 8 / 16 |
| Data | 1 255 curated conversation samples (139 validation), native channel format, tool/RAG/web context, no idle-loop residue, deduplicated, cluster-balanced |
| Epochs / updates | 1 / 162 (grad accumulation 8, seq 1024, lr 2e-5, gradient checkpointing) |
| Selected checkpoint | best (val 0.836 live) |
| Not shipped | a DPO stage on synthetic format pairs — every dose taught the model to emit a second `<channel |
| Hardware | Apple M4, 24 GB unified memory |
| Framework | MLX (mlx_vlm + mlx_lm.tuner) |
r1's adapter is still merged into the weights; r2 sits on top as the side path. Round-1
details are in AILEY_MERGE_INFO.json of the previous revision of this repo.
Limitations & biases
Inherited from Gemma 4 plus the fine-tune: the model may produce factually incorrect, outdated, or biased content. It is a persona, not a knowledge base. Occasionally (≈1 in 6 sampled answers on hard prompts) it skips the reasoning block and answers directly — seeding the channel opener removes that. Keep a human in the loop for consequential use.
License & attribution
This model is a Derivative Work of Google Gemma 4, released under the Apache License 2.0 (see the official Gemma 4 license). AileyCore-12B is distributed under Apache 2.0.
In accordance with Apache 2.0 §4: the base weights were modified via DoRA/LoRA adaptation
(r1 merged, r2 as side path — see AILEY_MERGE_INFO.json); a copy of the license is included
(LICENSE); attribution notices are in NOTICE.
Gemma is a trademark of Google LLC. This project is independent and not endorsed by or affiliated with Google.
Copyright 2026 OpenM!nded / Simon van de Loo
Portions © Google LLC (Gemma 4), Apache License 2.0
Licensed under the Apache License, Version 2.0 — http://www.apache.org/licenses/LICENSE-2.0
Citation
@misc{aileycore12b_2026,
title = {AileyCore-12B: A Gemma 4 fine-tune for the A!ley assistant},
author = {van de Loo, Simon and OpenM!nded},
year = {2026},
note = {Round 2: native reasoning channel, rank-8 side path outside the 6-bit quantization}
}
- Downloads last month
- 17
6-bit
Model tree for OpenMinded-Labs/AileyCore-12B
Base model
mlx-community/gemma-4-12B-it-6bit