Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 3 days ago • 64
ai-sage/GigaChat3.5-432B-A28B-Reasoning-GGUF Text Generation • 438B • Updated 17 days ago • 2.08k • 10
GigaChat 3.5 Collection GigaChat 3.5 is a large-scale Mixture-of-Experts (MoE) Hybrid model with 432B total parameters • 8 items • Updated 17 days ago • 20
ai-sage/GigaChat3.5-432B-A28B-Reasoning-GGUF Text Generation • 438B • Updated 17 days ago • 2.08k • 10
ai-sage/GigaChat3.5-432B-A28B-Reasoning-GGUF Text Generation • 438B • Updated 17 days ago • 2.08k • 10
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning Paper • 2608.14229 • Published Aug 14 • 17