Decision models (Laya and CLM), compiled for the SiMa.ai MLSoC Modalix
Laya by Convai Innovations is a "System-1" decision model: an encoder (ModernBERT-large, or mmBERT-base for the multilingual one) with a decision head that answers a typed question about a piece of text in one forward pass. It picks one of several options, scores on a scale, or says yes or no.
These are those checkpoints compiled for the Modalix MLA. Every graph is a single MLA stage, with no layers on the CPU, and a decision takes about 19 ms at 128 tokens (8 ms for the multilingual model).
| folder | model | precision | latency per decision, by tokens | same decision as PyTorch | size |
|---|---|---|---|---|---|
general |
Laya | BF16 | 128: 19.5 ms, 256: 35.2 ms, 512: 87.5 ms | 100 / 100 | 3.27 GB |
general-int8 |
Laya INT8 | A_BF16_W_INT8 | 128: 14 ms | 98 / 100 | 0.67 GB |
typed-decisions |
Laya typed-decisions | BF16 | 128: 19.4 ms, 512: 87.5 ms, 1024: 282 ms | 99 / 100 | 4.35 GB |
multilingual |
Laya multilingual | BF16 | 128: 8.1 ms, 256: 15.3 ms, 512: 33.9 ms, 1024: 116.5 ms | 100 / 100 | 2.61 GB |
clm |
CLM v0.1 8B | A_BF16_W_INT8 | 128: 254 ms | 9 of 13 reference decisions (int8 weights; an approximation) | 9.12 GB |
dino |
Laya-dino | BF16 | 64: 18 ms | 30 / 30 game states | 0.95 GB |
clm is a second kind of decision model, CLM v0.1 8B
by Contrastive-LM: a frozen Qwen3-8B embeds the state
and each option, and two small projection heads compare them. Its encoder is compiled as a
chain of 18 graphs with int8 weights (7.25 GB on the MLA), and its latency above is one pass
through them, which holds a question's state and all its options: about 0.3 s for a new
question, nothing for one seen before. With int8 weights it is an approximation of the
PyTorch model (embedding cosine about 0.99; 9 of 13 reference decisions the same), and a
one-token option is embedded poorly.
Latency is one decision on a Modalix DevKit, measured inside the runtime. "Same decision" is
against the fp32 PyTorch model on 100 decisions with the 128-token graph. Precision BF16 is
weights and activations in bfloat16; A_BF16_W_INT8 keeps bfloat16 activations and stores
the weights as int8.
What is in a folder
| file | what it is |
|---|---|
laya_s<N>_stage1_mla.elf |
the compiled graph for sequences of up to N tokens |
token_embeddings.bf16 |
the embedding table, looked up on the CPU |
act_tail.f32 |
the last layer of the act head, run on the CPU |
tokenizer.json |
the checkpoint's tokenizer |
laya_config.json |
token ids, calibration temperatures, and which graphs there are |
and in clm:
| file | what it is |
|---|---|
qwen_s128_l<NN>n2_stage1_mla.elf |
two layers of the encoder, from layer NN: 18 graphs that run one after another |
token_embeddings.bf16 |
Qwen3-8B's embedding table, looked up on the CPU |
heads.f32, heads.json |
CLM's state and action heads, run on the CPU, and their layout |
tokenizer.json |
Qwen3-8B's tokenizer |
clm_config.json |
which graphs there are, in order |
models.json lists every folder with file sizes and SHA-256 sums.
Using them
The files run on a Modalix board with the runtime and web app from neat-decision-studio: its Models page downloads a model from this repository onto the board and loads it on the MLA. By hand:
hf download TDoSiMa/sima-laya --include "general/*" --local-dir .
./laya run general --state "We were billed twice for March." \
--question '{"type": "noul", "instructions": "Is this a billing problem?"}'
They are not usable with PyTorch or ONNX Runtime; for that, use the original checkpoints at convaiinnovations/laya.
License and attribution
Apache-2.0. The models are Convai Innovations' Laya checkpoints
(https://github.com/NandhaKishorM/laya, Apache-2.0), compiled without retraining. The
exception is dino, whose decision head was fine-tuned on top of the English checkpoint for
a browser game. Zero-shot quality is the checkpoints': the port reproduces PyTorch's answers,
including its wrong ones.
clm holds the projection heads of Contrastive-LM's CLM-v0.1-8B
(https://github.com/Contrastive-LM/CLM, Apache-2.0), unchanged, and Alibaba Cloud's Qwen3-8B
(https://huggingface.co/Qwen/Qwen3-8B, Apache-2.0): its embedding table and tokenizer
unchanged, and its 36 decoder layers compiled for the MLA with their weights rounded to int8
and no language-model head.