Commit History
Merge pull request #31 from redhat-et/feat/new-hf-space 509f538 unverified
Taylor Agarwal commited on
โจ Change target space to new ET space c110dd1
Merge pull request #29 from hemajv/add-shellbench-glm 9cc5498 unverified
Taylor Agarwal commited on
Add Shellbench GLM-5.2 results 6341bb2
hemajv commited on
Merge pull request #28 from redhat-et/feat/versioning 5643285 unverified
Cรฉdric P. Legendre commited on
โจ Add project versioning fa603f3
Merge pull request #27 from cplegendre/feature/configuration-lab 243d0fa unverified
Taylor Agarwal commited on
Reduce ranking row height 38eb34d
CP Legendre commited on
Strengthen pairwise ranking regression assertions e21dd2d
CP Legendre commited on
Merge main and address ranking review feedback c797253
CP Legendre commited on
Rename rankings tab to Pairwise Rankings 47339ba
Scale dense charts with result count f1ae266
CP Legendre commited on
Rename metric-aware ranking order controls 48c3ff8
CP Legendre commited on
Refine paired comparisons and ranking order 0df8ce4
CP Legendre commited on
Add configuration metadata and comparison foundations 62e4b94
CP Legendre commited on
Force dark mode even when light mode is explicitly requested 91ed979
Add efficiency, coding, generalist, and benchmark matrix views (#25) 7ddc8cb unverified
Revert "Add Bradley-Terry rankings tab (#26)" 9472758
Add Bradley-Terry rankings tab (#26) e08b34a unverified
Rounak Bende commited on
Auto-update results.csv 7544ad4
github-actions[bot] commited on
Fix GPT-OSS-120B model name casing in shellbench result cef8a5e
Merge pull request #24 from cplegendre/feature/token-efficiency bd7b5e8 unverified
Taylor Agarwal commited on
Finalize efficiency review updates 52073ff
CP Legendre commited on
Add token efficiency analysis and visualizations 4fd20a0
CP Legendre commited on
Merge pull request #23 from redhat-et/fix/legend-placement 04b11c9 unverified
Taylor Agarwal commited on
Fix legend overlap and improve leaderboard chart readability 1bc9126
Auto-update results.csv c0a1e05
github-actions[bot] commited on
Fix GPT-OSS ansible benchmark name to match leaderboard grouping 583c1cf
Auto-update results.csv 24b60ee
github-actions[bot] commited on
Merge pull request #22 from redhat-et/fix/result-names b7aaa56 unverified
Taylor Agarwal commited on
Normalize result filenames to <benchmark>-<model>-<harness> order 7c69d1c
Auto-update results.csv e93c647
github-actions[bot] commited on
๐ Fix typos in shellbench runs 2fd572c
Merge pull request #21 from hemajv/add-shellbench-results 03033a9 unverified
Taylor Agarwal commited on
Auto-update results.csv eb6932f
github-actions[bot] commited on
Add GPT-OSS-120B results for all 3 benchmarks d44fe81
Add Shellbench results 20bd53c
hemajv commited on
Auto-update results.csv 977205a
github-actions[bot] commited on
Add Qwen3.6-27B-FP8 results for all 3 benchmarks 173b18f
Auto-update results.csv 29ee71d
github-actions[bot] commited on
Add Nemotron-120B Claude Code results for all 3 benchmarks d48d59d
Auto-update results.csv 6944b86
github-actions[bot] commited on
Standardize harness URLs for OpenCode and Pi 7c4a219
Auto-update results.csv 1032443
github-actions[bot] commited on
Fix Mistral Pro Ansible benchmark repo/url inconsistencies 9e5796a
Auto-update results.csv 4b0da2b
github-actions[bot] commited on
Fix benchmark and harness field inconsistencies in Nemotron result files 006b941
Auto-update results.csv 94cefe2
github-actions[bot] commited on