StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Paper • 2505.20139 • Published • 19
StructEval and related public resources for evaluating LLM generation and conversion of structured, renderable, and non-renderable outputs.
Note StructEval paper and benchmark website: https://arxiv.org/abs/2505.20139 · code: https://github.com/TIGER-AI-Lab/StructEval · project: https://structeval.github.io/
Note StructEval dataset: 2,035 examples across 18 structured formats and 44 task types. Dataset: https://huggingface.co/datasets/TIGER-Lab/StructEval · code: https://github.com/TIGER-AI-Lab/StructEval