Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry Paper • 2601.22588 • Published Jan 30 • 6
FineEdit: Unlock Instruction-Based Text Editing for LLMs Paper • 2502.13358 • Published Feb 19, 2025 • 1
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 8 days ago • 72