Spaces:
Running
Running
Position FlavourBench as frontier culinary benchmark
Browse files
README.md
CHANGED
|
@@ -8,12 +8,12 @@ sdk_version: 6.9.0
|
|
| 8 |
app_file: app.py
|
| 9 |
pinned: false
|
| 10 |
license: other
|
| 11 |
-
short_description:
|
| 12 |
---
|
| 13 |
|
| 14 |
# FlavourBench: An Executable Benchmark for Culinary Reasoning Without a Model Judge
|
| 15 |
|
| 16 |
-
An
|
| 17 |
executable culinary answer keys without a human or model judge.
|
| 18 |
|
| 19 |
[Paper](https://github.com/josefchen/flavourbench/blob/main/paper/build/flavourbench.pdf) 路
|
|
|
|
| 8 |
app_file: app.py
|
| 9 |
pinned: false
|
| 10 |
license: other
|
| 11 |
+
short_description: Culinary reasoning for frontier LLMs, without a model judge.
|
| 12 |
---
|
| 13 |
|
| 14 |
# FlavourBench: An Executable Benchmark for Culinary Reasoning Without a Model Judge
|
| 15 |
|
| 16 |
+
An executable benchmark and evidence explorer for 20 current frontier language-model endpoints, scored against
|
| 17 |
executable culinary answer keys without a human or model judge.
|
| 18 |
|
| 19 |
[Paper](https://github.com/josefchen/flavourbench/blob/main/paper/build/flavourbench.pdf) 路
|