Fonles Decision 1 9B

Fonles Decision 1 9B is a decision model built for Spanish, made by Fonles. It also works in English. It reads a state (text or JSON) and a set of typed questions (noul, choice, score) and returns a calibrated probability for every option in one forward pass. Nothing is generated, so there is no output to parse. It was trained on Spanish as people actually write it, from Mexico to Argentina to Spain, with regional slang, voseo and typos.

Fonles Decision 1 9B

Fonles Decision 1 9B ranks #1 in Spanish (89.6 on human-labeled Spanish) among the decision models we tested, including models three times its size. Its lead over Perplexity Decider 27B (89.3) is within noise; over the rest it is clear.

fonles.com · Benchmarks · Quickstart · En español · License

Model

Architecture 9B hybrid transformer (Gated-DeltaNet + attention) with a joint decision head
Parameters about 9B, bf16 weights
Head joint schema head (joint_head.safetensors), one score per option
Languages Spanish (main target), English
Question types noul (yes/no), choice (one of several options), score (ordinal rubric)
Calibration one temperature per question type in calibration.json
API shape compatible with SystemOne / Jev (POST /v1/systemone)
License CC BY-NC 4.0; commercial use under a Fonles license
Release date 2026-10-10

Benchmarks

We compared Fonles Decision 1 9B with six other decision models: Perplexity Decider 27B, Microsoft-Decision-1 9B, Clef 27B, Jev 1.13, GPT-6 Luna Decisions and Clef-flash 9B. Every model saw the same questions and was scored by the same code. None of these questions were used for training or checkpoint selection.

Spanish: the same 6,299 human-labeled questions for every model (PAWS-X, XNLI, MASSIVE, Belebele, SIB-200). English: the same 5,874 questions (MASSIVE en-US, PAWS-X en, Belebele eng). Accuracy in %; ECE on the Spanish set, lower is better.

Model Spanish PAWS-X XNLI MASSIVE Belebele SIB-200 ECE English ES + EN
Fonles Decision 1 9B 89.6 91.3 89.4 88.3 91.8 90.2 0.006 92.3 90.9
Perplexity Decider 27B 89.3 89.5 90.3 86.1 94.9 82.4 0.016 91.9 90.6
Microsoft-Decision-1 9B 85.8 84.4 84.1 87.3 96.2 89.2 0.054 90.6 88.1
Clef 27B 85.1 77.8 85.0 87.4 94.2 89.2 0.018 88.3 86.7
Jev 1.13 83.2 78.7 85.7 76.3 95.3 91.2 0.034 84.1 83.6
Clef-flash 9B (base) 82.6 72.7 82.0 87.3 93.1 87.3 0.035 86.4 84.5
GPT-6 Luna Decisions 76.9 78.0 75.6 73.8 91.8 87.3 0.065 80.2 78.5

Fonles Decision 1 9B is first in Spanish with a 9B model, ahead of models three times its size, and has the lowest calibration error of the group. Its lead over Perplexity Decider 27B is within noise; over every other model it is clear. English reading comprehension (Belebele) is its weakest spot. Competitors were queried through OpenRouter Decisions on 2026-10-10/11 (UTC) with the same items and the same scoring code.

Improvement from Fonles training

The full held-out benchmark: human-labeled MASSIVE es-ES, PAWS-X es, XNLI-es, Belebele (spa) and SIB-200 (spa), and the English counterparts for retention. All training data was filtered against it, exact and near-duplicate (5-word shingle Jaccard ≥ 0.6). Intervals are 95% CIs, paired on the same items.

Spanish, human-labeled (12,599 questions): 82.6% → 89.8% accuracy (+7.3 points) over the base model.
English (5,874 questions): 86.4% → 92.3% (+5.9 points). English improved as well; it did not regress.

Spanish (held-out)

Benchmark Clef-flash 9B (base) Fonles Decision 1 9B Δ accuracy Δ NLL vs base [95% CI]
PAWS-X (paraphrase) 72.4 91.0 +18.6 -0.360 [-0.391, -0.328] better
XNLI (entailment) 81.9 89.7 +7.8 -0.169 [-0.184, -0.154] better
MASSIVE (intent) 87.6 88.5 +0.9 -0.046 [-0.069, -0.024] better
SIB-200 (topic) 87.7 91.2 +3.4 -0.075 [-0.119, -0.031] better
Belebele (reading) 92.3 92.3 +0.0 +0.030 [-0.004, +0.064] ≈
Human-labeled average 82.6 89.8 +7.3

English retention

Benchmark Clef-flash 9B (base) Fonles Decision 1 9B Δ accuracy Δ NLL vs base [95% CI]
PAWS-X (paraphrase) 77.3 94.8 +17.4 -0.340 [-0.364, -0.317] better
MASSIVE (intent) 90.0 90.0 +0.0 -0.020 [-0.035, -0.004] better
Belebele (reading) 94.9 94.6 -0.3 +0.005 [-0.023, +0.032] ≈
Average 86.4 92.3 +5.9

Bold marks the higher accuracy in each row. "better": significantly better than the base (paired 95% CI on NLL entirely below 0) · ≈: no significant difference. NLL CIs are clustered by record. Probabilities are calibrated with the per-type temperatures shipped in calibration.json.

Quickstart

Tested with torch 2.11 and transformers 5.10.2. accelerate and torchvision are required even for text-only use.

pip install torch==2.11.0 torchvision==0.26.0 transformers==5.10.2 accelerate==1.15.0 pillow
import json, sys
import torch
from huggingface_hub import snapshot_download

path = snapshot_download("fonles-studios/fonles-decision-1")  # pin revision= in production
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")
temperature = json.load(open(f"{path}/calibration.json"))["temperature"]

record = {
    "state": "Che, me cobraron dos veces la cuota de este mes y nadie me responde.",
    "questions": {
        "equipo": {"type": "choice", "instructions": "¿Qué equipo debe atender el caso?",
                   "criteria": {"facturacion": "Cobros, pagos y reembolsos",
                                "soporte": "Fallas técnicas del servicio",
                                "ventas": "Planes nuevos y cambios de plan"}},
        "reembolso": {"type": "noul", "instructions": "¿El cliente pide que le devuelvan dinero?"},
    },
}
encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]
for q, q_logits in zip(encoded.questions, logits):
    t = temperature[record["questions"][q.question_id]["type"]]
    print(q.question_id, dict(zip(q.option_ids, (q_logits.float() / t).softmax(-1).tolist())))

You can ask many questions about the same state in one call; they all share the same forward pass.

joint_schema_model.py is Python code shipped with the repo and runs on import. Pin revision= to the commit you reviewed. The systemone() helper in that file comes from the base model and does not apply the temperatures; use the code above. Temperatures for this release: choice: T=0.900, noul: T=1.019, score: T=0.713. Calibrated probabilities are softmax(logits / T) with the T of the question type.

To serve the model over HTTP with the SystemOne / Jev request shape (POST /v1/systemone), use the server in server/.

How it was built

We built a Spanish decision dataset of tens of millions of tokens. Each case was designed from settings fixed in advance (country, industry, channel, register, difficulty, question shape and correct answer), so the answer is known before the text exists. Cases cover Latin America, the Caribbean, Spain and the US, and dozens of industries: banking, health administration, legal, e-commerce, logistics, telecom, government, insurance, HR and moderation. Every example was checked against its schema and target, near-duplicates were clustered, and any question a text-only classifier could answer without reading the case was removed.

That dataset was mixed with human-labeled data: MASSIVE es-ES, PAWS-X es, XNLI-es and BANKING77 retranslated into regional Spanish varieties, plus English replay so English would not regress.

The decision head was trained jointly with a rank-256 LoRA over the whole backbone on one H100, using label-smoothed cross-entropy with a Brier term, ordinal partial credit for score questions, a KL penalty toward the base model and random question order. Per-source early stopping guarded PAWS-X, MASSIVE and English. The LoRA is merged into the released weights, so this is a single checkpoint. Temperatures were then fitted on Spanish dev data, and the release was only published because it beat the base on the held-out Spanish benchmark with no significant regression on any human-labeled source.

Training data

Source What it adds License / origin
Fonles Spanish decision data Decision cases across Spanish varieties and dozens of industries Fonles; schema- and answer-validated
MASSIVE es-ES and en-US (train) Human-translated assistant intents CC BY 4.0, © Amazon.com (FitzGerald et al., 2022)
PAWS-X es and en (train) Adversarial paraphrase detection PAWS-X license (Yang et al., 2019)
BANKING77 (train) Banking intents. The Spanish side was machine-translated into regional varieties; the English side is original CC BY 4.0, PolyAI (Casanueva et al., 2020). Adapted material
XNLI es (train split only) Textual entailment CC BY-NC 4.0 per the XNLI repo (Conneau et al., 2018). Test and validation were not used for training
Fonles templates Invoices, triage, priority, moderation, out-of-scope Fonles

Intended use and limitations

Good fits: ticket routing, intent classification, triage and prioritization, moderation, policy and eligibility checks, NLI-style verification, guardrails for AI agents and rubric scoring, especially over Spanish text. The probabilities are meant to be thresholded: act, defer or send to a person.

Not a fit: high-stakes decisions (credit, health, employment, legal) without human review, or open-ended text generation.

  • English reading comprehension is the weak spot. On Belebele eng it scores 94.6, slightly below the base (94.9, not significant) and below Perplexity, Microsoft, Jev and Clef 27B (95.8–97.8).
  • Evaluation was text-only. Image inputs are inherited from the base but were not tested after fine-tuning.
  • Part of the training data is machine-translated, and training labels can contain errors (a low single-digit share of questions in our manual audit).
  • Calibration was fitted on Spanish dev data. Outside that domain, probabilities may be over- or under-confident.
  • Only Spanish and English were evaluated.
  • Competitor numbers come from OpenRouter Decisions on 2026-10-10/11 (UTC) and may change as those models are updated.

En español

Fonles Decision 1 9B es el modelo de decisiones de Fonles. Le pasas un state (texto o JSON) y preguntas tipadas (sí/no, opción múltiple o puntaje) y te devuelve una probabilidad calibrada para cada opción en una sola pasada, sin generar texto. Lo entrenamos con decenas de millones de tokens de casos de decisión en español de toda Latinoamérica, el Caribe, España y EE. UU., más datos etiquetados por personas.

En nuestras pruebas quedó primero en español (89,6), prácticamente empatado con Perplexity Decider 27B (89,3), un modelo tres veces más grande. Su punto más débil es la comprensión lectora en inglés (Belebele).

Sirve para enrutar tickets, clasificar intenciones, priorizar, moderar, revisar políticas y poner límites a agentes de IA. Puedes usar los pesos para investigación y uso no comercial. Para usarlo en producción, ajustarlo a tu dominio o tener un endpoint privado, escríbenos en fonles.com.

License and commercial use

Fonles Decision 1 9B is released under CC BY-NC 4.0: free for research, evaluation and other non-commercial use, with attribution to Fonles. Commercial use (production systems, paid products or services, hosted APIs) requires a license from Fonles: fonles.com.

Fonles Decision 1 9B was built starting from Cloudflare's open Clef-flash checkpoint. It is a modified derivative of Cloudflare/clef-flash and Qwen/Qwen3.5-9B (© Alibaba Cloud), both under Apache-2.0, reproduced in LICENSE-APACHE-2.0. Changes from the base model (Apache-2.0 §4(b)):

  • Modified backbone: a LoRA (rank 256) was trained on the backbone and merged into the weights, so the model-*.safetensors differ from the base.
  • Retrained joint schema head: joint_head.safetensors (same architecture, joint_head_config.json).
  • New calibration.json: choice: T=0.900, noul: T=1.019, score: T=0.713.
  • Unchanged: joint_schema_model.py, tokenizer, processor and chat_template.jinja are copied from the pinned base revision.

Training data keeps its own licenses and attributions (see the table above). Fonles is not affiliated with Cloudflare or Alibaba Cloud.

Citation

@misc{fonles2026decision1,
  title        = {Fonles Decision 1 9B},
  author       = {Fonles},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/fonles-studios/fonles-decision-1}},
  note         = {https://fonles.com}
}
Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for fonles-studios/fonles-decision-1

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(7)
this model

Datasets used to train fonles-studios/fonles-decision-1