When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation Paper • 2602.16763 • Published Jun 29
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation Paper • 2406.11580 • Published Oct 18, 2024
QE4PE: Word-level Quality Estimation for Human Post-Editing Paper • 2503.03044 • Published Mar 4, 2025 • 6
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement Paper • 2505.23183 • Published May 29, 2025 • 1
Can Large Language Models Capture Human Annotator Disagreements? Paper • 2506.19467 • Published Jun 24, 2025 • 18
COMET-poly: Machine Translation Metric Grounded in Other Candidates Paper • 2508.18549 • Published Aug 25, 2025
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Paper • 2512.16378 • Published Dec 18, 2025 • 8
Biased Tales: Cultural and Topic Bias in Generating Children's Stories Paper • 2509.07908 • Published Sep 9, 2025 • 1
TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models Paper • 2509.13395 • Published Sep 16, 2025 • 2