keelwright / README.md
ratingtesting
Update KDS leaderboard: kimi-k3 STRONG (25), Ling re-run (22), Nemotron-3-super PARTIAL
ae24bc4
|
Raw
History Blame Contribute Delete
1.61 kB
metadata
title: keelwright
emoji: 🛡️
colorFrom: green
colorTo: blue
sdk: static
pinned: false

🛡️ keelwright

28 machine-enforced safety checks for AI-generated code.

Catches SQL injection, hardcoded secrets, slopsquatting, reward hacking, doom loops, and 23 more failure modes.

Keelwright Score (KDS) up to 83/100 — proven by adversarial A/B testing across 11+ models.

Keelwright Score (KDS)

KDS = Execution Rate × Discrimination Rate / 100 (0-100 scale)

Model Tier Benchmark KDS
poolside/laguna-s-2.1 STRONG ML 78.5% 83
stepfun/step-3.7-flash MEDIUM Pro ~56% 67
nvidia/nemotron-3-ultra STRONG ML 67.7% 40
deepseek-v4-flash STRONG Verified ~79% 29
inclusionai/ling-3.0-flash UNKNOWN SWE-bench/GPQA N/A 22
kimi-k3 STRONG Terminal-Bench 88.3, ProgramBench 77.8 25
mimo-v2.5 MEDIUM Verified 78.9%, Pro 57.2% 18
nvidia/nemotron-3-super-120b-a12b STRONG Verified 60.47% PARTIAL
claude-opus-4-8 STRONG frontier 17
tencent/hy3 STRONG ML 75.8%, Verified 78% 7
cohere/north-mini-code WEAK Agentic 3.1 0

Links