Laya FI: Finnish + multilingual fine-tune of Laya

This is a fine-tune of the multilingual checkpoint of convaiinnovations/laya. Laya is a typed-decision model: an mmBERT encoder plus the RLCD decision head. You give it a state and typed questions (choice, noul, score) and it returns calibrated decisions. It does not generate text.

The fine-tune focuses on Finnish. It also trains on about 20 other languages and on the original English typed-decisions data, so the model keeps its multilingual behaviour.

Results (held-out eval, 20,642 decisions)

slice n base fine-tuned delta
overall 20642 62.3% 86.7% +24.4
Finnish 7604 67.6% 91.9% +24.3
other languages 13038 59.2% 83.7% +24.5
Full breakdown by task and language
slice n base fine-tuned delta
all 20642 0.623 0.867 +0.244
lang_group / fi 7604 0.676 0.919 +0.243
lang_group / other 13038 0.592 0.837 +0.245
task_lang / massive / fi 2400 0.539 0.921 +0.383
task_lang / massive / other 5700 0.598 0.896 +0.298
task_lang / reviews / other 1590 0.488 0.725 +0.236
task_lang / sentiment / fi 2000 0.702 0.917 +0.216
task_lang / sib200 / fi 204 0.775 0.843 +0.069
task_lang / sib200 / other 2448 0.781 0.875 +0.094
task_lang / toxicity / fi 3000 0.762 0.922 +0.160
task_lang / toxicity / other 1800 0.617 0.831 +0.213
task_lang / typed_decisions / other 1500 0.343 0.681 +0.338
lang / ar 654 0.572 0.813 +0.242
lang / da 300 0.597 0.893 +0.297
lang / de 919 0.613 0.839 +0.226
lang / en 2419 0.475 0.758 +0.283
lang / es 919 0.652 0.820 +0.169
lang / et 204 0.725 0.838 +0.113
lang / fi 7604 0.676 0.919 +0.243
lang / fr 919 0.624 0.850 +0.226
lang / he 150 0.553 0.773 +0.220
lang / hi 654 0.609 0.869 +0.260
lang / hu 504 0.621 0.867 +0.246
lang / it 450 0.662 0.884 +0.222
lang / ja 919 0.646 0.831 +0.185
lang / ko 300 0.503 0.890 +0.387
lang / nb 300 0.563 0.913 +0.350
lang / nl 300 0.640 0.900 +0.260
lang / pl 300 0.593 0.907 +0.313
lang / pt 300 0.640 0.900 +0.260
lang / ru 654 0.650 0.878 +0.228
lang / sv 504 0.665 0.901 +0.236
lang / tr 300 0.580 0.893 +0.313
lang / uk 150 0.400 0.820 +0.420
lang / zh 619 0.601 0.769 +0.168
lang / zh-CN 300 0.677 0.913 +0.237
ilang / fi 5583 0.654 0.898 +0.245
ilang / en 15059 0.612 0.856 +0.244

Usage

from laya import Agent
from huggingface_hub import snapshot_download
agent = Agent(snapshot_download("misukisu/laya-fi-multilingual"))
out = agent.predict("Laita olohuoneen valot pois päältä", {
    "scenario": {"type": "choice", "criteria": {"iot": "smart home devices", "music": "music", "weather": "weather"}},
    "is_question": {"type": "noul", "criteria": "the user is asking a question"},
})
print(out["answers"])

ONNX (CPU, no torch needed at inference)

from laya.onnx_agent import ONNXAgent
from huggingface_hub import snapshot_download
d = snapshot_download("misukisu/laya-fi-multilingual")
agent = ONNXAgent(d, onnx_path=f"{d}/onnx/laya.onnx")

The ONNX export (fp32, opset 18) uses Laya's official export_onnx.py. On 104 sampled decisions it matched PyTorch on 104 of them (agree 104/104 max_conf_diff 0.0000).

GGUF: there is no GGUF build. llama.cpp and Ollama only run autoregressive LMs and cannot run Laya's decision head. Use ONNX for lightweight or edge deployment.

Training data (~107k typed decisions)

  • Finnish: MASSIVE intent and scenario (fi-FI), SIB-200 topic (fin_Latn), Finnish sentiment, and the Finnish Jigsaw toxicity set (toxic, obscene, threat, insult, identity attack). Label descriptions are randomly in Finnish or English, so the model works with either.
  • Multilingual: MASSIVE in many locales, SIB-200, textdetox multilingual toxicity, and Amazon reviews (rating plus would-recommend).
  • Anti-forgetting: the original LocalLLaMA/typed-decisions (English).

Training setup

  • Laya's laya.train.finetune pipeline (RLCD loss, choice options shuffled, temperature calibration), with a data-parallel loop across 2 x Tesla T4.
  • fp16 autocast, 2 epochs, effective batch 64 (16 per GPU x 2 accumulation x 2 GPUs), length-bucketed batches, max length 512.
  • Embeddings frozen to protect the 100+ language vocabulary; all encoder layers trained.
  • Training time: 1.05 h.

Limitations

  • The eval set uses the same sources as the training data (held-out rows only), so it measures in-domain gains. Results on other domains will vary.
  • Toxicity labels in the Finnish Jigsaw set are machine-translated.
Downloads last month
11
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for misukisu/laya-fi-multilingual

Quantized
(62)
this model

Datasets used to train misukisu/laya-fi-multilingual