--- license: apache-2.0 language: - en library_name: transformers pipeline_tag: zero-shot-classification tags: - decision-model - system-one - falcondec - lightdec_v2 - calibrated-decisions - multiple-choice - intent-classification - customer-support - natural-language-inference - code - guardrails - agents - selective-prediction - falconsai - model-surgeon - attested-lineage --- # LightDec_V2 (Long) **[View in Model Surgeon](https://surgeon.falcons.ai/?hub=Falconsai/LightDec_V2)** **LightDec_V2** is a small, fast, calibrated *decision model* from [Falcons.ai](https://huggingface.co/Falconsai). Given a piece of state (an email, a ticket, an agent trace, a contract, a log, a JSON record) and one or more typed questions, it picks an answer from a closed set of options in a single forward pass and returns calibrated probabilities, so your application knows when to trust the answer and when to defer to a human or a larger model. It is the long-context successor to [Falconsai/LightDec](https://huggingface.co/Falconsai/LightDec). Every call now uses a **2,048-token window** (LightDec v1 used 512 tokens by default and only switched to 2,048 for questions with more than 24 options), so long email threads, contracts and service logs fit without truncation. | | | |---|---| | **Architecture** | FalconDec (encoder + permutation-equivariant option head) | | **Backbone** | [`jhu-clsp/ettin-encoder-150m`](https://huggingface.co/jhu-clsp/ettin-encoder-150m) | | **Parameters** | ~0.16B (fp16 weights ≈ 319 MB; int8 export available) | | **Context** | 2,048 tokens per call | | **Question types** | `choice`, `noul` (yes/no), `score` (ordinal levels) | | **Options per question** | 2 to 96 in one pass; more via an automatic tournament | | **Test accuracy** | 78.4% micro / 78.9% macro over 31,990 test items, 73 tasks | | **Calibration** | ECE 0.029 (per-type, per-option-count temperature scaling) | | **Latency** | ~10 ms per call on GPU (fp16), ~48 ms on CPU (int8) | | **Version** | 1.0.0 (trained from the pretrained backbone, not fine-tuned from LightDec v1) | ## What it does LightDec_V2 answers the kind of small, structured questions that sit inside real systems: *Which queue does this ticket go to? Does this email ask for a refund? How angry is the customer, 0–2? Is this agent step wrong? Is this prompt a jailbreak? Does this invoice violate the policy?* It is not a generative model and never produces free text; it only chooses among the options you give it, which makes its output easy to validate, log and act on. Three question types are supported, and several questions about the same state can be asked in one call: - **`choice`**: pick one of N labelled options (options can carry descriptions, e.g. `{"billing": "Charges, invoices, refunds"}`). - **`noul`**: a yes/no judgment about a statement (returns `p_true`). - **`score`**: an ordinal scale such as `["Calm", "Annoyed", "Furious"]` (also returns an `expected_level`). Every answer includes a `confidence` and a `defer` flag. With the default `defer_threshold` of 0.7, the model answered **68.0%** of test questions and was **90.6%** accurate on those; the remainder are flagged for review. ## How it works The question, the options and the state are packed into one sequence: ``` [CLS] question [SEP] [MASK] option_1 [MASK] option_2 ... [MASK] option_k [SEP] state [SEP] ``` The encoder reads the whole sequence once. The hidden vector at each `[MASK]` marker represents one option; it is combined with the `[CLS]` context and a question-type embedding, then passed through a 2-layer, 8-head set transformer in which options attend to each other **without positional encoding**, so the result does not depend on option order. An MLP produces one logit per option and a softmax turns them into a distribution. Finally a temperature, learned per question type and per option-count bucket (2, 3–5, 6–12, 13+ options) and stored in `falcondec_config.json`, calibrates the probabilities. Questions with more than 96 options are resolved with a tournament: options are scored in chunks, the best of each chunk advance, and a final pass ranks the survivors. ## Usage The model ships with its own loader, `falcondec_modeling.py`, which requires `torch`, `transformers`, `safetensors` and `huggingface_hub`. ```python import importlib.util from huggingface_hub import hf_hub_download path = hf_hub_download("Falconsai/LightDec_V2", "falcondec_modeling.py") spec = importlib.util.spec_from_file_location("falcondec_modeling", path) fd = importlib.util.module_from_spec(spec) spec.loader.exec_module(fd) model, tok = fd.load_falcondec("Falconsai/LightDec_V2") # GPU if available, else CPU state = { "email": { "subject": "Charged twice this month", "body": "I was billed twice for my Pro plan in March. Please refund the duplicate charge. This is the second time!", }, "customer": {"plan": "Pro", "customer_since": "2021"}, } questions = { "topic": { "type": "choice", "instructions": "What is this support email primarily about?", "criteria": { "billing": "Charges, invoices, payments, refunds or subscription costs", "technical": "Something in the product is broken, failing or erroring", "feature_request": "Asking for a new feature or an improvement", "account": "Login, password, profile, seats or account settings", "none": "None of these", }, }, "refund": {"type": "noul", "instructions": "The email asks for money to be returned or credited"}, "anger": {"type": "score", "instructions": "How angry does the customer sound?", "criteria": ["Calm", "Annoyed", "Furious"]}, } out = fd.decide(model, tok, state, questions) for key, r in out["answers"].items(): print(key, r["choice"], round(r["confidence"], 3), "DEFER" if r["defer"] else "") ``` Each result contains `choice`, `choice_text`, `confidence`, the full `probs` distribution and `defer`; `noul` answers add `p_true` and `score` answers add `expected_level`. Pass `defer_threshold=` to `decide` to trade coverage for accuracy. For lower-level control, `fd.score_items(model, tok, items)` scores raw `{"state", "question", "options", "type"}` items in batches. **CPU / small footprint.** If the repository includes the int8 export (as LightDec v1 does in `compact-int8/`), load that folder the same way; the loader dequantizes the per-channel int8 weights automatically. ## Evaluation All numbers come from `falcondec_report.json`. The test split has 31,990 items across 73 tasks in 11 domains. Tasks marked *held-out* are flagged as such in the training report. ### Overall | Metric | Value | |---|---:| | Micro accuracy | 78.4% | | Macro accuracy (mean over tasks) | 78.9% | | Negative log-likelihood | 0.563 | | Brier score | 0.295 | | Expected calibration error | 0.029 | | AURC (area under risk–coverage) | 0.068 | | Mean absolute error on `score` questions (levels) | 0.51 | | Coverage at defer threshold 0.7 | 68.0% | | Accuracy on covered questions | 90.6% | ### By domain | Domain | Tasks | Test items | Accuracy | |---|---:|---:|---:| | long_context | 3 | 450 | 100.0% | | support | 3 | 1,505 | 98.2% | | code | 9 | 3,854 | 91.1% | | intents | 4 | 1,800 | 85.0% | | agentic | 6 | 1,563 | 82.1% | | policy | 11 | 5,500 | 80.6% | | guardrails | 4 | 1,294 | 79.8% | | workflows | 4 | 2,000 | 77.3% | | reasoning | 18 | 9,224 | 71.8% | | tev1_benchmark | 7 | 2,800 | 65.4% | | classification | 4 | 2,000 | 61.2% | ### Head-to-head against the proof_v2 baseline On 72 shared tasks (up to 120 items each), LightDec_V2 averaged **79.6%** against **46.7%** for the `proof_v2` baseline, winning on 70 tasks, tying on 1 and losing on 1 (`policy/invoice_total_transfer`).
Per-task head-to-head | Task | n | LightDec_V2 | proof_v2 | Δ | |---|---:|---:|---:|---:| | `long/email_thread` | 120 | 100.0% | 5.0% | +95.0 | | `agenttrek/next_action_type` | 120 | 87.5% | 7.5% | +80.0 | | `policy/table_extreme_transfer` | 120 | 92.5% | 14.2% | +78.3 | | `long/contract_clause` | 120 | 100.0% | 23.3% | +76.7 | | `policy/access_control_transfer` | 120 | 100.0% | 23.3% | +76.7 | | `long/service_log` | 120 | 100.0% | 24.2% | +75.8 | | `policy/count_threshold_transfer` | 120 | 86.7% | 10.8% | +75.8 | | `bigclonebench/clone` | 120 | 96.7% | 24.2% | +72.5 | | `policy/return_window_transfer` | 120 | 100.0% | 34.2% | +65.8 | | `snli/contradicts` | 120 | 100.0% | 35.8% | +64.2 | | `policy/refund_approval_transfer` | 120 | 95.8% | 32.5% | +63.3 | | `gsm8k/math` | 120 | 73.3% | 12.5% | +60.8 | | `hotpotqa/retrieve` | 120 | 87.5% | 29.2% | +58.3 | | `civil_comments/toxic` | 120 | 93.3% | 35.8% | +57.5 | | `snli/nli` | 120 | 90.8% | 34.2% | +56.7 | | `triage/support_email` | 120 | 95.8% | 39.2% | +56.7 | | `mnli/claim` | 120 | 87.5% | 31.7% | +55.8 | | `scitail/support` | 120 | 95.8% | 42.5% | +53.3 | | `agenttrek/finish_now` | 120 | 78.3% | 25.8% | +52.5 | | `tev1_test/ag_news` | 120 | 92.5% | 42.5% | +50.0 | | `typed_decisions/agent_trace_observability` | 120 | 80.8% | 33.3% | +47.5 | | `hotpotqa/comparison_yes_no` | 26 | 92.3% | 46.2% | +46.2 | | `typed_decisions/customer_service` | 120 | 74.2% | 28.3% | +45.8 | | `policy/table_compare_transfer` | 120 | 91.7% | 49.2% | +42.5 | | `jailbreak/detect` | 120 | 98.3% | 56.7% | +41.7 | | `ag_news/topic` | 120 | 85.0% | 45.0% | +40.0 | | `openbookqa/mcq` | 120 | 65.0% | 28.3% | +36.7 | | `tev1_test/mnli` | 120 | 75.0% | 39.2% | +35.8 | | `counsel/critique_quality` | 120 | 60.8% | 25.8% | +35.0 | | `typed_decisions/invoice_processing` | 120 | 81.7% | 47.5% | +34.2 | | `humaneval/completion` | 119 | 84.9% | 51.3% | +33.6 | | `clinc150/intent` | 120 | 97.5% | 64.2% | +33.3 | | `hellaswag/continuation` | 120 | 60.8% | 29.2% | +31.7 | | `yelp/score` | 120 | 62.5% | 32.5% | +30.0 | | `mbpp/bugspot` | 120 | 85.8% | 57.5% | +28.3 | | `tev1_test/boolq` | 120 | 84.2% | 55.8% | +28.3 | | `commonsense_qa/mcq` | 120 | 68.3% | 40.8% | +27.5 | | `anli/nli` | 120 | 59.2% | 32.5% | +26.7 | | `typed_decisions/security_incidents` | 120 | 75.8% | 50.0% | +25.8 | | `arc_challenge/mcq` | 120 | 53.3% | 30.8% | +22.5 | | `policy/free_shipping_transfer` | 120 | 92.5% | 70.8% | +21.7 | | `agentharm/refuse` | 120 | 66.7% | 45.0% | +21.7 | | `massive_en/intent` | 120 | 95.0% | 74.2% | +20.8 | | `arc_easy/mcq` | 120 | 65.0% | 45.8% | +19.2 | | `tev1_test/banking77` | 120 | 71.7% | 55.0% | +16.7 | | `boolq/yes_no` | 120 | 82.5% | 66.7% | +15.8 | | `policy/table_count_transfer` | 120 | 31.7% | 15.8% | +15.8 | | `mbpp/solution` | 120 | 100.0% | 85.0% | +15.0 | | `policy/invoice_overdue_transfer` | 120 | 80.8% | 65.8% | +15.0 | | `tev1_test/routing` | 120 | 44.2% | 29.2% | +15.0 | | `bitext/category` | 120 | 100.0% | 85.8% | +14.2 | | `qasc/mcq` | 120 | 99.2% | 85.8% | +13.3 | | `winogrande/blank` | 120 | 69.2% | 55.8% | +13.3 | | `emotion/6way` | 120 | 49.2% | 36.7% | +12.5 | | `tev1_test/sst5` | 120 | 39.2% | 27.5% | +11.7 | | `aqua_rat/math` | 120 | 35.0% | 24.2% | +10.8 | | `sst5/score` | 120 | 38.3% | 28.3% | +10.0 | | `devign/vulnerability` | 120 | 66.7% | 56.7% | +10.0 | | `mmlu/mcq` | 120 | 40.8% | 30.8% | +10.0 | | `sciq/mcq` | 120 | 96.7% | 86.7% | +10.0 | | `tev1_test/policy` | 120 | 50.0% | 40.8% | +9.2 | | `codexglue/func_name` | 120 | 99.2% | 91.7% | +7.5 | | `policy/sla_urgency_transfer` | 120 | 51.7% | 44.2% | +7.5 | | `bitext/route` | 120 | 100.0% | 92.5% | +7.5 | | `counsel/step_has_error` | 120 | 82.5% | 75.0% | +7.5 | | `codexglue/code_to_doc` | 120 | 100.0% | 93.3% | +6.7 | | `snli/must_be_true` | 120 | 99.2% | 92.5% | +6.7 | | `banking77/intent` | 120 | 92.5% | 86.7% | +5.8 | | `prompt_injections/detect` | 116 | 59.5% | 54.3% | +5.2 | | `codexglue/doc_to_code` | 120 | 98.3% | 97.5% | +0.8 | | `codexglue/lang_id` | 120 | 100.0% | 100.0% | +0.0 | | `policy/invoice_total_transfer` | 120 | 45.0% | 52.5% | -7.5 |
### Support-email triage (bundled benchmark) `benchmarks/triage_support_email_test.jsonl` contains 101 support emails, each with five typed questions and reference answers. It doubles as a worked example of the input format. | Question | Type | Accuracy | |---|---|---:| | `topic` | choice (5-way) | 99.0% | | `refund` | noul (yes/no) | 100.0% | | `breakage` | score (4 levels) | 99.0% | | `anger` | score (3 levels) | 92.1% | | `judgment` | choice (3-way) | 82.2% | All five answers together route an email to the correct handling pile **94.1%** of the time. ### Long context Three long-document tasks exercise the 2,048-token window (150 items each): `long/contract_clause`, `long/email_thread` and `long/service_log`. LightDec_V2 scored 100% on all three. These are synthetic, in-distribution tasks, so treat them as a check that long inputs are read end to end rather than as a measure of general long-document reasoning. ### TEV1 transfer benchmark TEV1 is a separate decision benchmark. Some of its tasks draw on the same public sources as LightDec's training data; the table marks which. | Task | n | Accuracy | Overlaps LightDec training source | |---|---:|---:|:---:| | `tev1_test/ag_news` | 150 | 92.0% | yes | | `tev1_test/banking77` | 200 | 72.0% | no | | `tev1_test/boolq` | 200 | 83.5% | yes | | `tev1_test/mnli` | 300 | 72.3% | yes | | `tev1_test/policy` | 1200 | 52.9% | no | | `tev1_test/routing` | 600 | 45.8% | no | | `tev1_test/sst5` | 150 | 39.3% | no | Overall TEV1 accuracy is **58.4%** (2,800 items), and **51.8%** on the tasks with no source overlap. For reference, the report lists published results for the 4B-parameter TEV1 model of 88% on its main decisions set and 100% on its policy transfer set; those are different splits and a model roughly 25× larger, so the figures are context rather than a like-for-like comparison.
Full per-task test results (73 tasks) | Domain | Task | n | Accuracy | Chance | ECE | Held-out | |---|---|---:|---:|---:|---:|:---:| | agentic | `agenttrek/finish_now` | 151 | 78.8% | 50.0% | 0.075 | | | agentic | `agenttrek/next_action_type` | 487 | 86.4% | 19.1% | 0.092 | | | agentic | `counsel/critique_quality` | 201 | 62.7% | 33.3% | 0.271 | | | agentic | `counsel/step_has_error` | 201 | 83.1% | 50.0% | 0.148 | | | agentic | `hotpotqa/comparison_yes_no` | 26 | 92.3% | 50.0% | 0.082 | | | agentic | `hotpotqa/retrieve` | 497 | 89.3% | 16.7% | 0.043 | | | classification | `ag_news/topic` | 500 | 90.0% | 25.0% | 0.052 | | | classification | `emotion/6way` | 500 | 47.2% | 16.7% | 0.211 | ✓ | | classification | `sst5/score` | 500 | 40.6% | 20.0% | 0.069 | ✓ | | classification | `yelp/score` | 500 | 66.8% | 20.0% | 0.091 | | | code | `bigclonebench/clone` | 500 | 96.2% | 50.0% | 0.031 | | | code | `codexglue/code_to_doc` | 504 | 99.0% | 26.2% | 0.010 | | | code | `codexglue/doc_to_code` | 504 | 97.6% | 27.9% | 0.013 | | | code | `codexglue/func_name` | 467 | 95.5% | 25.5% | 0.039 | | | code | `codexglue/lang_id` | 504 | 100.0% | 23.5% | 0.001 | | | code | `devign/vulnerability` | 500 | 63.2% | 50.0% | 0.070 | | | code | `humaneval/completion` | 119 | 84.9% | 39.4% | 0.140 | ✓ | | code | `mbpp/bugspot` | 256 | 85.5% | 40.6% | 0.037 | | | code | `mbpp/solution` | 500 | 97.8% | 25.0% | 0.020 | | | guardrails | `agentharm/refuse` | 416 | 70.0% | 50.0% | 0.089 | ✓ | | guardrails | `civil_comments/toxic` | 500 | 92.6% | 50.0% | 0.051 | | | guardrails | `jailbreak/detect` | 262 | 97.3% | 50.0% | 0.031 | | | guardrails | `prompt_injections/detect` | 116 | 59.5% | 50.0% | 0.333 | ✓ | | intents | `banking77/intent` | 500 | 92.2% | 32.1% | 0.037 | ✓ | | intents | `banking77/intent_77` | 300 | 58.0% | 1.3% | 0.241 | ✓ | | intents | `clinc150/intent` | 500 | 97.0% | 12.4% | 0.015 | | | intents | `massive_en/intent` | 500 | 93.0% | 13.0% | 0.028 | | | long_context | `long/contract_clause` | 150 | 100.0% | 20.0% | 0.000 | | | long_context | `long/email_thread` | 150 | 100.0% | 25.0% | 0.000 | | | long_context | `long/service_log` | 150 | 100.0% | 20.0% | 0.001 | | | policy | `policy/access_control_transfer` | 500 | 100.0% | 33.3% | 0.002 | | | policy | `policy/count_threshold_transfer` | 500 | 88.0% | 10.5% | 0.048 | | | policy | `policy/free_shipping_transfer` | 500 | 93.8% | 50.0% | 0.041 | | | policy | `policy/invoice_overdue_transfer` | 500 | 82.4% | 50.0% | 0.020 | | | policy | `policy/invoice_total_transfer` | 500 | 47.4% | 50.0% | 0.082 | | | policy | `policy/refund_approval_transfer` | 500 | 96.2% | 33.3% | 0.019 | | | policy | `policy/return_window_transfer` | 500 | 100.0% | 33.3% | 0.026 | | | policy | `policy/sla_urgency_transfer` | 500 | 60.6% | 25.0% | 0.131 | | | policy | `policy/table_compare_transfer` | 500 | 92.8% | 50.0% | 0.049 | | | policy | `policy/table_count_transfer` | 500 | 32.4% | 12.1% | 0.541 | | | policy | `policy/table_extreme_transfer` | 500 | 93.2% | 12.2% | 0.053 | | | reasoning | `anli/nli` | 498 | 48.6% | 33.3% | 0.146 | | | reasoning | `aqua_rat/math` | 247 | 34.0% | 20.0% | 0.080 | | | reasoning | `arc_challenge/mcq` | 500 | 49.0% | 25.0% | 0.168 | ✓ | | reasoning | `arc_easy/mcq` | 500 | 61.2% | 25.0% | 0.122 | ✓ | | reasoning | `boolq/yes_no` | 500 | 82.2% | 50.0% | 0.080 | | | reasoning | `commonsense_qa/mcq` | 493 | 64.3% | 20.0% | 0.143 | | | reasoning | `gsm8k/math` | 500 | 70.0% | 25.0% | 0.085 | | | reasoning | `hellaswag/continuation` | 500 | 57.6% | 25.0% | 0.075 | | | reasoning | `mmlu/mcq` | 500 | 39.0% | 25.0% | 0.155 | ✓ | | reasoning | `mnli/claim` | 500 | 86.0% | 33.3% | 0.081 | | | reasoning | `openbookqa/mcq` | 500 | 57.2% | 25.0% | 0.218 | | | reasoning | `qasc/mcq` | 500 | 98.6% | 12.5% | 0.008 | | | reasoning | `sciq/mcq` | 498 | 95.4% | 25.0% | 0.021 | | | reasoning | `scitail/support` | 500 | 96.2% | 50.0% | 0.030 | | | reasoning | `snli/contradicts` | 500 | 99.0% | 33.3% | 0.019 | | | reasoning | `snli/must_be_true` | 500 | 98.6% | 33.3% | 0.043 | | | reasoning | `snli/nli` | 988 | 89.5% | 33.3% | 0.108 | | | reasoning | `winogrande/blank` | 500 | 66.2% | 50.0% | 0.116 | | | support | `bitext/category` | 500 | 100.0% | 16.1% | 0.001 | | | support | `bitext/route` | 500 | 100.0% | 19.0% | 0.000 | | | support | `triage/support_email` | 505 | 94.5% | 32.3% | 0.046 | | | tev1_benchmark | `tev1_test/ag_news` | 150 | 92.0% | 25.0% | 0.098 | ✓ | | tev1_benchmark | `tev1_test/banking77` | 200 | 72.0% | 17.8% | 0.137 | ✓ | | tev1_benchmark | `tev1_test/boolq` | 200 | 83.5% | 50.0% | 0.050 | ✓ | | tev1_benchmark | `tev1_test/mnli` | 300 | 72.3% | 33.3% | 0.075 | ✓ | | tev1_benchmark | `tev1_test/policy` | 1200 | 52.9% | 33.3% | 0.059 | ✓ | | tev1_benchmark | `tev1_test/routing` | 600 | 45.8% | 20.0% | 0.038 | ✓ | | tev1_benchmark | `tev1_test/sst5` | 150 | 39.3% | 20.0% | 0.108 | ✓ | | workflows | `typed_decisions/agent_trace_observability` | 500 | 73.4% | 30.0% | 0.222 | | | workflows | `typed_decisions/customer_service` | 500 | 76.4% | 28.0% | 0.210 | | | workflows | `typed_decisions/invoice_processing` | 500 | 82.6% | 35.0% | 0.195 | | | workflows | `typed_decisions/security_incidents` | 500 | 76.8% | 34.0% | 0.235 | |
## Training | | | |---|---| | Initialization | Pretrained `jhu-clsp/ettin-encoder-150m` (trained from scratch on top of the backbone; no LightDec v1 weights) | | Data | 1,023,814 training / 24,200 validation / 31,990 test examples | | Sources | Public classification, NLI, QA, code, intent, guardrail and agent-trace datasets converted into typed decisions, plus synthetic policy-transfer, workflow, support-triage and long-context tasks and TEV1 builders. `mind2web` was excluded; no custom data was added. | | Preset | `long` (2,048-token sequences, 1,500 long-context training tasks) | | Objective | Cross-entropy with spherical-score (0.5) and ranked-probability-score (1.0) terms for ordinal questions; 8% "none of the above" augmentation; task sampling α = 0.5 | | Optimizer | Learning rate 8e-5 (encoder) / 6e-4 (head), layer-wise decay 0.9, weight decay 0.01, 6% warmup, gradient clip 1.0, EMA 0.999 | | Batching | Token-budget batches of 65,536 tokens (batch size 128) | | Schedule | 10 epochs; checkpoint from **epoch 4** selected on validation macro accuracy | | Calibration | Temperature per (question type × option-count bucket) fit on validation | | Compute | 292 minutes on one NVIDIA RTX PRO 6000 Blackwell Server Edition | | Software | PyTorch 2.9.0 (CUDA 13.0), Transformers 5.17.0, Python 3.12 | | Epoch | Train loss | Train acc | Val macro acc | Val NLL | |---:|---:|---:|---:|---:| | 1 | 0.706 | 73.4% | 81.9% | 0.408 | | 2 | 0.431 | 84.7% | 84.5% | 0.362 | | 3 | 0.343 | 87.9% | 85.0% | 0.392 | | 4 **(selected)** | 0.277 | 90.4% | 85.2% | 0.453 | | 5 | 0.220 | 92.6% | 84.9% | 0.537 | | 6 | 0.170 | 94.5% | 84.7% | 0.671 | | 7 | 0.127 | 96.0% | 84.5% | 0.843 | | 8 | 0.095 | 97.1% | 84.5% | 1.034 | | 9 | 0.071 | 97.9% | 84.4% | 1.211 | | 10 | 0.057 | 98.4% | 84.4% | 1.419 | Validation accuracy peaked at epoch 4 while validation NLL kept rising after epoch 2, so the selected checkpoint is somewhat overconfident before calibration. The fitted temperatures (roughly 1.6 to 2.5) correct for this; they are part of the checkpoint and applied automatically by `decide` and `score_items`. ## Speed Measured with `decide` on a single state; latency grows only slightly as more questions are asked about the same state. | Device | Weights | Questions per call | p50 | p95 | |---|---|---:|---:|---:| | cuda | fp16 | 1 | 10.05 ms | 10.43 ms | | cuda | fp16 | 5 | 11.32 ms | 11.41 ms | | cuda | fp16 | 10 | 12.57 ms | 12.69 ms | | cpu fp32 | int8 file | 1 | 48.4 ms | 48.8 ms | GPU throughput: about **2,589 decisions per second** with batching. ## Intended use LightDec_V2 is built for high-volume, low-latency decisions inside software: ticket and email routing, intent detection, triage scoring, content and prompt-safety screening, policy and rule checks over structured records, agent step verification and "should the agent stop now" checks, and pre-filtering before a more expensive LLM call. The calibrated `confidence` and `defer` flag are designed for human-in-the-loop and cascade setups in which uncertain cases escalate. ## Limitations - **Closed-set only.** The model can only choose among the options you supply. If the right answer is not listed, it will still pick something; include a `none` option when that can happen. - **Weak at multi-step reasoning and arithmetic.** Knowledge- and math-heavy tasks are well below the rest: MMLU 39.0%, AQuA-RAT 34.0%, ANLI 48.6%, ARC-Challenge 49.0%. Some policy tasks that require computing over a table are also weak: `policy/table_count_transfer` scores 32.4% with poor calibration (ECE 0.54), and `policy/invoice_total_transfer` (47.4%) is below chance. Use a larger model for anything that needs counting or totalling. - **Fine-grained and subjective labels.** Accuracy drops on 77-way banking intents (58.0%), six-way emotion (47.2%) and five-level sentiment (SST-5 40.6%, Yelp 66.8%). - **Transfer to unfamiliar rule sets is limited.** On TEV1 tasks without source overlap, accuracy is 51.8%, and routing and policy transfer (45.8% and 52.9%) are the weakest; validate on your own policies before relying on it. - **Guardrail coverage is partial.** Jailbreak detection is strong (97.3%), but prompt-injection detection (59.5%) and harmful-request refusal (70.0%) are not reliable enough to be a sole safety layer. - **Calibration is aggregate.** Overall ECE is low, but a few tasks (for example `counsel/critique_quality`, `prompt_injections/detect`, the table-count policy task) are noticeably miscalibrated; check calibration on your own distribution before tuning `defer_threshold`. - **English only**, and questions are truncated to 96 tokens and each option to 24 tokens by default (`option_tokens` can raise the per-option budget). - **Long-context scores are synthetic.** The 100% long-context results are in-distribution checks, not evidence of general long-document reasoning. ## Files | File | Description | |---|---| | `model.safetensors` | fp16 weights | | `falcondec_config.json` | Head configuration, special-token ids, defer threshold and calibration temperatures | | `falcondec_modeling.py` | Model definition, loader (`load_falcondec`) and inference API (`decide`, `score_items`) | | `falcondec_report.json` | Full training and evaluation report | | `encoder/`, `tokenizer/` | Ettin encoder config and tokenizer | | `benchmarks/triage_support_email_test.jsonl` | 101-email support-triage benchmark with reference answers | ## Citation ```bibtex @misc{falconsai_lightdec_v2_2026, title = {LightDec_V2: A Long-Context, Calibrated Decision Model}, author = {Falcons.ai}, year = {2026}, howpublished = {\url{https://huggingface.co/Falconsai/LightDec_V2}} } ``` Built on the Ettin encoder from JHU CLSP ([`jhu-clsp/ettin-encoder-150m`](https://huggingface.co/jhu-clsp/ettin-encoder-150m)). This card is generated from the surgical record itself; the package's `lineage.intoto.jsonl` is the signed source of truth (verify it free at the Surgeon's public verifier or with the bundled `verify_attestation.py`). ## Architecture - Identification: **NLP · Small Language Model (SLM)** (98% confidence) - Source format: `safetensors` · Intended task: not declared - `config.json`: synthesized from the anatomy (no source config.json) - Source license: apache-2.0 - Lineage chain: 1 surgery (no prior attestation reachable) · Falconsai/LightDec_V2 - Post-surgery totals: 159,654,157 parameters · 168 tensors - Compute estimate: 15.02439 GFLOPs (comparison metric, not a measurement) ## Provenance & operations - Parents: Falconsai/LightDec_V2/model.safetensors - Operations performed: load×1 - Weight merges recorded: 0 - Quantized tensors (F32→F16): 0 ## Surgery Log (ordered) 1. **load** — hub:Falconsai/LightDec_V2/model.safetensors (319.3 MB, safetensors) ## Validation - Tissue imaging: not run - Structural integrity is testable offline via the packaged `load_and_test.py`. ## Compliance note The signed attestation + this card together document model composition, modification history, and validation evidence — the record structure technical-documentation obligations (e.g. EU AI Act Annex IV) ask for. This is evidence, not legal advice. --- *Operated with Model Surgeon — verify this package at https://surgeon.falcons.ai/verify* *© 2026 FALCONS.AI — Model Surgeon record format. The model weights remain their owner's.*