ChihHsuan-Yang's picture
Release v1.0.0 accompanying the EMNLP 2026 camera-ready paper (arXiv:2608.14927)
0d3ef4a verified
Raw History Blame Contribute Delete
2.59 kB
NOTICE
======
This repository is the public research artifact for the EMNLP 2026 paper
"LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration
Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks"
(arXiv:2608.14927).
Code in this repository is released under the MIT License (see LICENSE).
--------------------------------------------------------------------------
Upstream benchmarks
--------------------------------------------------------------------------
This artifact reports measured outcomes of language-model protocol executions
on problems drawn from the benchmarks below. We do NOT redistribute upstream
problem text or gold answers. We release stable problem identifiers, derived
per-protocol outcomes, and instructions for reconstructing the inputs from the
pinned upstream releases. Each benchmark remains subject to its own license.
* Omni-MATH - https://huggingface.co/datasets/KbsdJames/Omni-MATH
License: Apache-2.0.
Used as a filtered subset of 4,181 competition-mathematics problems.
* JEEBench - https://huggingface.co/datasets/daman1209arora/jeebench
License: MIT. 515 problems.
* SciBench - https://huggingface.co/datasets/xw27/scibench
License: MIT (see the upstream repository's LICENSE file; the Hugging Face
dataset card does not currently declare a license tag). 565 problems.
* LAB-Bench - https://huggingface.co/datasets/futurehouse/lab-bench
License: CC-BY-SA-4.0. FutureHouse.
Pinned upstream revision: 5c77cec648430f30611808808861eb86f81d5eaa.
Two text-only slices are used: a strict slice (741 problems) and a broader
text-no-tool slice (1,542 problems). Image-dependent subsets (FigQA, TableQA)
are excluded because the protocols studied here are text-only.
Please cite the original benchmark authors when using these identifiers.
--------------------------------------------------------------------------
Foundation models
--------------------------------------------------------------------------
The solver families evaluated here (openai/gpt-oss-120b and
google/gemma-4-31B-it) are third-party models. This artifact does not
redistribute, retrain, or claim ownership of those model weights.
--------------------------------------------------------------------------
Acknowledgments
--------------------------------------------------------------------------
This research used resources of the Argonne Leadership Computing Facility, a
U.S. Department of Energy (DOE) Office of Science user facility at Argonne
National Laboratory (ANL) operated under Contract No. DE-AC02-06CH11357.