Inkling-Small-Hybrid-4layer

A 4-layer slice of thinkingmachines/Inkling-Small, made for miles CI tests. It is not meant for inference quality: the layers were cut without any retraining, so its outputs are not meaningful.

  • Text layers 0, 1, 2 and 5 of the original 42, renumbered 0-3: two dense layers (0, 1), one local (sliding-window) attention MoE layer (2) and one global attention MoE layer (3, originally 5). Unlike a slice of layers 0-3, it keeps a global attention layer, as the full model does.
  • config.json is the original config with only text_config.num_hidden_layers = 4 and text_config.local_layer_ids = [0, 1, 2] changed.
  • Every tensor is a byte-identical copy from the original checkpoint, in its original dtype (BF16, plus a few F32 router tensors); original layer 5 is stored as layer 3. That covers the embeddings, final norm and unembedding, the four text layers, all MTP layers, and the vision and audio adapters. The tokenizer, chat template and processor files are copied unchanged.
Downloads last month
16
Safetensors
Model size
18B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support