Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding Paper • 2606.21906 • Published Jun 20 • 26
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security Paper • 2605.29801 • Published May 28 • 150
SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation Paper • 2606.03348 • Published Jun 2 • 2
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety Paper • 2604.12710 • Published Apr 13 • 5
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety Paper • 2604.12710 • Published Apr 13 • 5
CritiqueLLM: Scaling LLM-as-Critic for Effective and Explainable Evaluation of Large Language Model Generation Paper • 2311.18702 • Published Nov 30, 2023
AlignBench: Benchmarking Chinese Alignment of Large Language Models Paper • 2311.18743 • Published Nov 30, 2023 • 1
EVA: An Open-Domain Chinese Dialogue System with Large-Scale Generative Pre-Training Paper • 2108.01547 • Published Aug 3, 2021
CharacterBench: Benchmarking Character Customization of Large Language Models Paper • 2412.11912 • Published Dec 16, 2024
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Paper • 2508.06471 • Published Aug 8, 2025 • 213