laya-MLX-8bit

MLX conversion of convaiinnovations/laya for Apple Silicon: the root (English) checkpoint and the multilingual and typed-decisions checkpoints.

  • Converted from the base model at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851.
  • Not affiliated with or endorsed by convaiinnovations.
  • Converted by Sasan Sotoodehfar, CAVI AI (https://cavi-ai.xyz).

Contents

Path Checkpoint Encoder Context (max_len / head_max_len) Size
/ laya (English) ModernBERT-large, 28 layers 512 / 192 524 MB
multilingual/ laya-multilingual mmBERT-base, 22 layers, 256k vocabulary 1024 / 256 569 MB
typed-decisions/ laya-typed-decisions ModernBERT-large, 28 layers 1024 / 256 524 MB
  • Each folder holds model.safetensors, config.json, tokenizer.json, tokenizer_config.json.
  • Quantized to 8-bit (affine, group size 32): the linear layers of the encoder and of the decision head.
  • Kept in float16: token embeddings, norms, type embedding, option scorer, act head.
  • config.json carries the decision settings and the shipped temperatures (temperature, temperature_by_options).
  • laya/: MLX model code (ModernBERT encoder from mlx-embeddings, decision head, request rendering, calibrated answers). mlx-embeddings 0.1.0 does not include this model.
  • 8-bit only: at 4 bits the multilingual checkpoint changed 3 of 25 reference answers.

Requirements

  • Apple Silicon Mac.
  • Python 3.12.
  • pip install mlx-embeddings==0.1.0

Usage

import sys
from huggingface_hub import snapshot_download
path = snapshot_download("cavi-ai/laya-MLX-8bit")
sys.path.insert(0, path)
from laya import load, predict

model, tokenizer = load(path)  # or f"{path}/multilingual", f"{path}/typed-decisions"
state = {"document": "I was charged twice for my subscription this month. Please refund the duplicate charge."}
questions = {
    "department": {"type": "choice", "instructions": "Which team should handle this ticket?",
                   "criteria": {"billing": "payments, invoices, refunds", "technical": "bugs, outages, errors", "sales": "new purchases and upgrades"}},
    "urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["not urgent", "somewhat urgent", "very urgent"]},
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
}
print(predict(model, tokenizer, state, questions)["answers"])
  • state: text or a JSON object.
  • Question types: choice (criteria as a list or an object of option: description), score (ordered list of levels), noul (optional true / false texts).
  • Answers: choice with probabilities and confidence; score (expected level) with probabilities; noul (probability of true).
  • All questions of a request run in one forward pass.

Measured results

Hardware: Apple M5 Max.

Check root multilingual typed-decisions
Reference questions 17 (English) 25 (7 languages) 17 (English)
Same answer as the PyTorch fp32 reference 17/17 25/25 17/17
Mean / max probability change vs the reference 0.0009 / 0.011 0.0026 / 0.021 0.0008 / 0.003
Usage example above: department, refund probability billing 0.95, 0.89 billing 1.00, 0.99 billing 0.82, 0.73
  • Reference: transformers ModernBertModel plus torch.nn decision-head layers, float32, on CPU, loaded from the original checkpoints; tokenization through AutoTokenizer.
  • Port check: the MLX code in float32 on CPU matches the reference within 2.5e-5 on every option logit, with identical input ids.
  • One request with three questions answers in 12–46 ms after loading.

License

  • Model weights, tokenizer, and configuration files: Apache-2.0, inherited from the base model. See LICENSE.
  • Code in laya/: MIT. See laya/LICENSE.

Links

Downloads last month
21
Safetensors
Model size
0.4B params
Tensor type
F16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cavi-ai/laya-MLX-8bit

Finetuned
(132)
this model