BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data Paper • 2510.10159 • Published Oct 11, 2025 • 3
Measuring what Matters: Construct Validity in Large Language Model Benchmarks Paper • 2511.04703 • Published Nov 3, 2025 • 8
platzi/platzi-distilbert-base-uncased-mrpc-glue-juan-garcia-bauza Text Classification • 82.1M • Updated Dec 1, 2025 • 7
platzi/platzi-distilbert-base-uncased-mrpc-glue-juan-garcia-bauza Text Classification • 82.1M • Updated Dec 1, 2025 • 7
Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings Paper • 2509.14405 • Published Sep 17, 2025 • 2
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Paper • 2506.22439 • Published May 29, 2025 • 3
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments Paper • 2509.14233 • Published Sep 17, 2025 • 23
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America Paper • 2507.00999 • Published Jul 1, 2025 • 2
Query Attribute Modeling: Improving search relevance with Semantic Search and Meta Data Filtering Paper • 2508.04683 • Published Aug 6, 2025
DSBC : Data Science task Benchmarking with Context engineering Paper • 2507.23336 • Published Jul 31, 2025 • 2
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation Paper • 2504.07072 • Published Apr 9, 2025 • 9
Uncovering Cultural Representation Disparities in Vision-Language Models Paper • 2505.14729 • Published May 20, 2025 • 1
It's the same but not the same: Do LLMs distinguish Spanish varieties? Paper • 2504.20049 • Published Apr 8, 2025