Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 2 days ago • 55
ai-sage/GigaChat3.5-432B-A28B-Reasoning-GGUF Text Generation • 438B • Updated 16 days ago • 2.07k • 10
GigaChat 3.5 Collection GigaChat 3.5 is a large-scale Mixture-of-Experts (MoE) Hybrid model with 432B total parameters • 8 items • Updated 16 days ago • 20
ai-sage/GigaChat3.5-432B-A28B-Reasoning-GGUF Text Generation • 438B • Updated 16 days ago • 2.07k • 10
ai-sage/GigaChat3.5-432B-A28B-Reasoning-GGUF Text Generation • 438B • Updated 16 days ago • 2.07k • 10
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning Paper • 2608.14229 • Published Aug 14 • 17