Running Featured 57 Gemma 4 - Vision Token Budget πΌ 57 Resize images for visual token budgets while keeping aspect ratio
Running 3 Minimal Conversation (S2S backend, WebSocket) π 3 Voice chat over WebSocket against a HF speech-to-speech
Artificial Hippocampus Networks for Efficient Long-Context Modeling Paper β’ 2510.07318 β’ Published Oct 8, 2025 β’ 32
view article Article SmolLM3: smol, multilingual, long-context reasoner +21 eliebak, cmpatino, anton-l, edbeeching, m-ric, nouamanetazi, akseljoonas, guipenedo, hynky, clefourrier, SaylorTwift, kashif, qgallouedec, hlarcher, glutamatt, Xenova, reach-vb, ngxson, craffel, lewtun, loubnabnl, lvwerra, thomwolf β’ Jul 8, 2025 β’ 787
view article Article Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM +2 ariG23498, merve, pcuenq, reach-vb β’ Mar 12, 2025 β’ 498
view article Article SigLIP 2: A better multilingual vision language encoder +1 ariG23498, merve, qubvel-hf β’ Feb 21, 2025 β’ 224
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features Paper β’ 2502.14786 β’ Published Feb 20, 2025 β’ 169
view article Article PaliGemma 2 Mix - New Instruction Vision Language Models by Google +1 merve, ariG23498, andsteing β’ Feb 19, 2025 β’ 74
view article Article PaliGemma 2 Mix - New Instruction Vision Language Models by Google +1 merve, ariG23498, andsteing β’ Feb 19, 2025 β’ 74
view article Article Introducing smolagents: simple agents that write actions in code. +1 m-ric, merve, thomwolf β’ Dec 31, 2024 β’ 1.21k
view article Article Welcome PaliGemma 2 β New vision language models by Google +2 merve, andsteing, pcuenq, ariG23498 β’ Dec 5, 2024 β’ 168