⚑ Tasksource-JEV-Nano (tasksource-jev-nano-v0)

SOTA High-Throughput Native Decision Model (<200M Parameters)
Decoupled Multi-Vector Late Interaction for Fast, Calibrated System-One Decisions

Decision Index Decision Index Skill JevBench Composite Classifier Benchmark v2 License Parameters Context Length


What is Tasksource-JEV-Nano?

Tasksource-JEV-Nano is a compact (~149M parameter) decision model that picks the optimal action, category, or verdict from a candidate list using token-level multi-vector late interaction rather than conventional classification heads.

Built on answerdotai/ModernBERT-base and lightonai/LateOn, it handles all three fundamental System-One decision types:

  1. Choice: Selecting the best candidate among variable numbers of options ($K=2$ to $K=150+$).
  2. NOUL: Yes/no uncertainty decisions.
  3. Score: Quantitative ranking and graded scales.

Key Architectural Strengths

  • Reusable State Caching: The situation context (state + question) is encoded once into token vectors. Evaluating 1, 10, or 100 candidate options reuses this cached state without re-encoding the context, delivering up to 5.45x speedup.
  • Exact Candidate Permutation Equivariance: Unlike encoder-concatenation classifiers where candidate order biases logits, late-interaction candidate scoring is mathematically independent across candidates. Permuting the option list guarantees an identically permuted output (max diff: 0.00e+00).
  • Sub-40ms Low-Latency Inference: Operates in ~30–37 ms per decision on standard GPUs, requiring only ~1.5 GB VRAM.

πŸ“Š Benchmark Evaluations

All evaluations below are measured on official benchmark suites under zero-shot out-of-distribution conditions:

πŸ† Sub-200M Peer Comparison

Model Size Decision Index Raw Classifier-Benchmark v2 JevBench v1.3 Banking77 ($K=77$) CLINC150 ($K=151$) Latency (p50)
Tasksource-JEV-Nano-v0 149M 30.65 πŸ† 55.78% πŸ† 63.14 πŸ† 44.22% πŸ† 67.04% πŸ† 37.6 ms
GLiClass Modern-Large 340M β€” 53.80% ~55.9 40.20% ~58.4% ~85 ms
GLiNER2.5 Base 200M β€” 54.20% ~60.0 41.50% ~61.0% ~62 ms
GLiClass Base 150M β€” 50.70% ~51.2 37.30% ~54.1% ~45 ms
Laya Multilingual 150M β€” 46.10% ~51.7 1.80% (collapses) ~38.2% ~50 ms
GLiClass Edge 35M β€” 35.20% ~35.3 24.10% ~31.0% ~18 ms

Official Decision Index 0.2.1 (Full 100% Suite: 150,759 Requests)

Evaluated on the full official Decision Index 0.2.1 test suite covering all 38 scored benchmarks with 100% suite coverage:

  • Balanced Raw Index: 30.65 / 100
  • Balanced Skill Index (Chance-Corrected): 9.58 / 100
  • Breadth Skill Index: 8.39 / 100
  • Suite Coverage: 100.0% (150,759 / 150,759 requests completed, 0 errors)

Area Breakdown (Official 0.2.1 Weights):

  • Retrieval & Intent Classification: 40.37% (Skill: 20.84%)
    • CLINC150+OOS: 67.04% macro-F1 (151 classes, 5,500 requests)
    • BANKING77: 40.79% macro-F1 / 44.22% accuracy (77 classes, 3,080 requests)
  • Arts & Human Taste: 33.73% (Skill: 9.32%)
    • BPoMP: 55.54% accuracy (5,000 requests)
    • Humicroedit: 55.52% accuracy (2,628 requests)
    • New Yorker Caption Matching: 34.66% accuracy
  • Language Understanding: 32.21% (Skill: 4.34%)
    • WinoGrande: 51.85% accuracy
    • ANLI: 32.63% macro-F1 (3,200 requests)
    • HellaSwag: 30.69% accuracy
  • Tools & Automation: 26.70% (Skill: 15.18%)
    • BFCL: 62.60% field accuracy / 34.95% case-exact (1,694 requests)
    • API-Bank: 31.30% accuracy (53 API tools, 508 requests)
    • When2Call: 30.70% accuracy
  • Knowledge & Reasoning: 23.16% (Skill: 2.23%)
    • MMLU: 29.72% accuracy (all 14,033 questions across 57 subjects)
    • RAGTruth: 47.8% hallucination detection F1

πŸš€ Usage Guide

Installation

pip install -U pylate torch

Python Inference with PyLate

from pylate import models
import torch

# Load the trained Decision Index champion model
model = models.ColBERT("tasksource/tasksource-jev-nano-v0")

# 1. Provide the decision context (State + Question)
state = "Customer reports unauthorized international wire transfer of $4,500."
question = "Select the appropriate fraud mitigation routing:"
context = [state + "\nQuestion: " + question]

# 2. Provide candidate decision options
options = [
    "approve and monitor silently",
    "challenge with a push notification",
    "freeze the account and call the customer",
    "decline and file a report",
]

# 3. State is encoded once; each option is scored via token MaxSim
ctx = model.encode(context, is_query=False, convert_to_tensor=True)[0]
opts = model.encode(options, is_query=True, convert_to_tensor=True)
scores = torch.stack([(o @ ctx.T).max(dim=1).values.sum() for o in opts])

# 4. Calibrated temperature softmax probabilities
temperature = 0.3236  # Calibrated model temperature
probs = torch.softmax(scores / temperature, dim=0)

for opt, p in zip(options, probs):
    print(f"  {opt:45s} -> {p:.4f}")

πŸ“œ Citation

@misc{sileo2026jevnanov0,
  title={Tasksource-JEV-Nano-v0: Decoupled Multi-Vector Late Interaction for Typed Decisions},
  author={Sileo, Damien},
  year={2026},
  howpublished={\url{https://huggingface.co/tasksource/tasksource-jev-nano-v0}},
}

Base model LateOn by LightOn (Apache 2.0). Built on ModernBERT-base.

Downloads last month
148
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for tasksource/tasksource-jev-nano-v0

Base model

lightonai/LateOn
Finetuned
(1)
this model