physical ai
masterset-ai
AI & ML interests
https://www.masterset.ai
Recent Activity
reacted to SeaWolf-AI's post with โค๏ธ about 1 hour ago
๐ Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB).
S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score.
๐ Top of the board
๐ฅ Darwin ZTC v2 (FINAL-Bench) 89.58
๐ฅ OpenJev-27B 87.50
๐ฅ AutoJev-27B 87.07
4๏ธโฃ Eikos 27B 85.43
5๏ธโฃ Jev 1.13 85.05
๐ Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline.
โ๏ธ Why a zero-token judge wins here
๐น It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass.
๐น Zero generated tokens, no decoding loop, so latency and cost stay low.
๐น Holds up out of distribution too: General Noul 96.00, General Choice 99.34.
It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call.
๐ Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2
๐ Leaderboard: https://huggingface.co/spaces/hotchpotch/S1MB-leaderboard
Standings move as new models are added. Numbers reflect the board at the time of writing. ๐
reacted to SeaWolf-AI's post with ๐ค about 1 hour ago
๐ Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB).
S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score.
๐ Top of the board
๐ฅ Darwin ZTC v2 (FINAL-Bench) 89.58
๐ฅ OpenJev-27B 87.50
๐ฅ AutoJev-27B 87.07
4๏ธโฃ Eikos 27B 85.43
5๏ธโฃ Jev 1.13 85.05
๐ Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline.
โ๏ธ Why a zero-token judge wins here
๐น It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass.
๐น Zero generated tokens, no decoding loop, so latency and cost stay low.
๐น Holds up out of distribution too: General Noul 96.00, General Choice 99.34.
It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call.
๐ Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2
๐ Leaderboard: https://huggingface.co/spaces/hotchpotch/S1MB-leaderboard
Standings move as new models are added. Numbers reflect the board at the time of writing. ๐
upvoted an article about 9 hours ago
Leading the System One Mosaic Benchmark: What Darwin-27B-ZTC-v2's #1 MeansOrganizations
None yet