Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Paper • 2606.13603 • Published Jun 11
Distilling Formal Logic into Neural Spaces: A Kernel Alignment Approach for Signal Temporal Logic Paper • 2603.05198 • Published Mar 5
Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers Paper • 2507.07808 • Published Jul 10, 2025
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German Paper • 2406.06131 • Published Jun 10, 2024
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Paper • 2506.06275 • Published Jun 6, 2025
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios Paper • 2606.06177 • Published Jun 4
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese Paper • 2603.26511 • Published Mar 27
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering Paper • 2510.09351 • Published Oct 10, 2025
AgREE: Agentic Reasoning for Knowledge Graph Completion on Emerging Entities Paper • 2508.04118 • Published Aug 6, 2025
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents Paper • 2602.08964 • Published Feb 9 • 1
EAGER: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling Paper • 2510.11170 • Published Oct 13, 2025 • 3
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering Paper • 2503.14996 • Published Mar 19, 2025 • 3
Steering Large Language Models for Machine Translation Personalization Paper • 2505.16612 • Published May 22, 2025 • 6
Mergenetic: a Simple Evolutionary Model Merging Library Paper • 2505.11427 • Published May 16, 2025 • 15
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results Paper • 2504.13677 • Published Apr 18, 2025 • 1
Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces Paper • 2503.05283 • Published Mar 7, 2025 • 4