How to use from
Docker Model Runner
# Gated model: Login with a HF token with gated access permission
hf auth login
docker model run hf.co/baselquant/Inkling-GGUF:IQ1_M
Quick Links

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Atomic Chat Join Discord GitHub

Inkling

Inkling, self-quantized to GGUF by Atomic Chat. Built straight from Thinking Machines Lab's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.

Highlights

  • 952.4B parameters: the weights this repo quantizes.
  • 66 layers: Mixture-of-Experts.
  • Modalities: the base model handles Text, Image, Audio; this repo ships text-only quants, it carries no vision projector.
  • Full imatrix ladder: every quant is calibrated with an importance matrix, published here alongside the quants.

These GGUFs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.

Always pass --jinja so the Inkling chat template is applied. Without it the model can emit malformed turns.

Model Overview

Property Value
Base model thinkingmachines/Inkling
Parameters 952.4B
Layers 66
Experts 256 routed (top-6)
Context length not stated
Vocabulary 201,024
Modalities Text, Image, Audio in the base model; text only in this repo, it ships no vision projector
Architecture Mixture-of-Experts, 256 experts (top-6), 64 attention heads over 8 KV heads, InklingForConditionalGeneration
This repo GGUF quants (imatrix); the importance matrix is published here as imatrix/imatrix-code-at_128.gguf
Inkling benchmark scores

Scores are Thinking Machines Lab's published results for the base thinkingmachines/Inkling, not our own measurements. Quantization preserves the large majority of this; Q4_K_M and up stay close to full precision.

Get started

Run Inkling locally with:

  • Atomic Chat: the easiest path. Open the app, search AtomicChat/Inkling-GGUF, pick a quant, hit Use this model.
  • llama.cpp: llama-server -hf AtomicChat/Inkling-GGUF:None --jinja -c 8192
  • Ollama: ollama run hf.co/AtomicChat/Inkling-GGUF:None
  • LM Studio / Jan: search the repo id, download any quant.

Best practices

Parameter Value
sampling defaults not stated

The base model card does not state sampling defaults.

Run in llama.cpp

git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
./llama.cpp/build/bin/llama-server \
    -hf AtomicChat/Inkling-GGUF:None \
    --jinja -ngl 99 -c 8192 -fa on

How these were made

  1. Download thinkingmachines/Inkling (original weights).
  2. Convert to f16 GGUF with llama.cpp.
  3. Build an importance matrix over our calibration corpus, published here as imatrix/imatrix-code-at_128.gguf.
  4. Quantize the ladder with --imatrix.

License

Original model by Thinking Machines Lab, released under the Apache 2.0 license. Full terms: Apache 2.0. Quantized by Atomic Chat.

Downloads last month
-
GGUF
Model size
947B params
Architecture
inkling
Hardware compatibility
Log In to add your hardware

1-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for baselquant/Inkling-GGUF

Quantized
(24)
this model