Instructions to use FINAL-Bench/Darwin-27B-ZTC-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FINAL-Bench/Darwin-27B-ZTC-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="FINAL-Bench/Darwin-27B-ZTC-v2")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("FINAL-Bench/Darwin-27B-ZTC-v2") model = AutoModel.from_pretrained("FINAL-Bench/Darwin-27B-ZTC-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Darwin-27B-ZTC-v2
A zero-token decision engine from the Darwin family, second version. Darwin-27B-ZTC-v2 reads a piece of state and a typed question (noul yes or no, choice one of N labels, score an ordered rubric) and returns a probability for every option. It uses one forward pass per question and generates no tokens.
v2 adds control and workflow decisions (games, drones, retrieval control, entity alignment, customer and incident workflows) on top of FINAL-Bench/Darwin-27B-ZTC.
Results
S1MB (System One Mosaic Benchmark), english-v1
S1MB, all 137 benchmarks, scored with the S1MB evaluator and its autojev adapter. Numbers are the mean baseline-adjusted score (x100).
| Model | Choice (57) | Noul (59) | Score (21) | Mean of the 3 types | Mean of 137 |
|---|---|---|---|---|---|
| Darwin-27B-ZTC-v2 | 71.51 | 67.66 | 60.21 | 66.46 | 68.12 |
| Darwin-27B-ZTC (v1) | 70.12 | 63.54 | 43.97 | 59.21 | 63.28 |
The table is from our run of the merged checkpoint before upload; the S1MB submission uses this repository at a pinned revision.
Training overlap (disclosed). v2 was trained on the train split of ZefanCai/Open-Jev (CC0). S1MB includes 22 benchmarks built from the test split of the same Open-Jev tasks. We used no S1MB test case: an exact-match filter over every S1MB test state, query and document removed zero training rows. v2 was not trained on any S1MB data or on Typed Decisions data.
How it was made
- Start from Darwin-27B-ZTC (v1).
- Continue full-weight training on the Open-Jev
trainsplit mixed with v1's original training data, selecting the checkpoint on held-out development rows only. - Average the weights of v1 and the continued model (50/50). The average keeps v1's skills on its original tasks and adds the new control and workflow skills.
Files and usage
The layout is the same as v1: backbone weights (BF16, about 54 GB), readout.safetensors, decision_config.json, tokenizer, and the inference code in autojev/ (from autojev, MIT, with text-only backbone support added).
ztc_server.py: aPOST /v1/systemoneserver (ZTC_MODEL=<path> PORT=8000 python ztc_server.py).ztc_engine.py: an in-process engine for the Decision Index kit.
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("FINAL-Bench/Darwin-27B-ZTC-v2")
sys.path.insert(0, path)
from autojev.model import DecisionModel
model = DecisionModel(checkpoint=path)
row = {"state": {"ticket": "Charged twice for one order."},
"question": {"type": "choice", "instructions": "What should support do?",
"criteria": {"refund": "Refund the duplicate charge.", "escalate": "Send to billing.", "close": "No action."}}}
print(model.predict([row])[0])
Citation
@misc{darwin27bztcv2,
title = {Darwin-27B-ZTC-v2: a zero-token decision engine},
author = {VIDRAFT and FINAL-Bench},
year = {2026},
url = {https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2}
}
- Downloads last month
- -