Trainable Dynamic Mask Sparse Attention: Bridging Efficiency and Effectiveness in Long-Context Language Models wubingheng • Aug 5, 2025 • 7
AG-BPE v4: Enhanced Attention-Guided Byte-Pair Encoding with Weighted Layer Aggregation RDTvlokip • Aug 4, 2025 • 2
What Open-Source Developers Need to Know about the EU AI Act's Rules for GPAI Models yjernite • Aug 4, 2025 • 29
Unsupervised Model Improvement via Internal Coherence Maximization: Outperforming Human-Supervised Methods Through Self-Elicitation codelion • Aug 3, 2025 • 7
🚀 Wan 2.2 & FLUX Krea Full Tutorial — Automated Install & Perfect Presets MonsterMMORPG • Aug 2, 2025 • 1
AG-BPE: Attention-Guided Byte-Pair Encoding for Semantic-Aware Tokenization RDTvlokip • Aug 2, 2025 • 1