ilo toki — MiLMMT-46 1B

A translator between Toki Pona and English, Russian and Vietnamese. Small enough to run on a phone: it powers ilo toki, which does all of its translation on device.

This repository holds both the merged weights and GGUF builds, so there is one place to look rather than a repository per format.

Prompt format

The model keeps the prompt format of its base, and there is no chat template — do not wrap the input in one.

Translate this from Toki Pona to English:
Toki Pona: jan li moku e kili
English:

The translation follows the final <target language>: line and ends at the model's end-of-generation token. Either side can be the source:

Translate this from Russian to Toki Pona:
Russian: Я тебя люблю.
Toki Pona:

Language names are written out in full — Toki Pona, English, Russian, Vietnamese. Getting the format wrong does not fail loudly: the model keeps producing fluent text while silently ignoring the requested target language.

Toki Pona is written in lower case; capitalization in the input is not something the model expects.

Which file to use

File Size Notes
ilo-toki-MiLMMT-46-1b-Q4_K_M.gguf 0.94 GB Smallest.
ilo-toki-MiLMMT-46-1b-Q5_K_M.gguf 1.00 GB
ilo-toki-MiLMMT-46-1b-Q6_K.gguf 1.24 GB
ilo-toki-MiLMMT-46-1b-Q8_0.gguf 1.29 GB What the app ships — see below.
model.safetensors 2.48 GB Merged weights, bf16, for transformers.

The quantizations sit unusually close together because the 262k-token embedding matrix is about a third of the model and quantizes the same way in all of them. Q8_0 therefore costs only 0.05 GB more than Q6_K and 0.35 GB more than Q4_K_M, which is why the app ships it: on a phone the difference between these files is small, while the difference between fitting in RAM and not is enormous.

Running it

With llama.cpp:

llama-cli -m ilo-toki-MiLMMT-46-1b-Q8_0.gguf --no-cnv \
  -p "Translate this from Toki Pona to English:
Toki Pona: jan li moku e kili
English:"

With transformers:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "NetherQuartz/ilo-toki-MiLMMT-46-1b-merged"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "Translate this from Toki Pona to English:\nToki Pona: jan li moku e kili\nEnglish:"
inputs = tokenizer(prompt, return_tensors="pt")
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=64)[0]))

How it was built

A LoRA adapter (NetherQuartz/ilo-toki-MiLMMT-46-1b) trained with TRL SFT — rank 64, targeting the attention and MLP projections as well as the token embeddings — merged into MiLMMT-46-1B-v0.1 and quantized with llama.cpp.

The base is a 46-language translation model from Xiaomi Research, so the fine-tune starts from a model that already translates rather than from a general purpose one.

A note for anyone re-merging this adapter

The adapter puts a LoRA on embed_tokens, and gemma3 ties lm_head to that same tensor. A plain merge_and_unload() produces a model that repeats a single token forever: a LoRA on an embedding changes what the lookup returns, not the stored weights, so during training the tied output head read the base embeddings — merging writes the delta into the tensor and the head suddenly sees an update it never saw while training.

The weights here were merged with the output head untied and left at the original embeddings, which reproduces training exactly. That is also why model.safetensors carries a separate lm_head.weight and tie_word_embeddings is false.

Licence

Gemma Terms of Use, inherited through the base model.

Downloads last month
93
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NetherQuartz/ilo-toki-MiLMMT-46-1b-merged