SWE-chat: Coding Agent Interactions From Real Users in the Wild Paper • 2604.20779 • Published Apr 22 • 21
view article Article Training Design for Text-to-Image Models: Lessons from Ablations Photoroom • Feb 3 • 77
view article Article There is no such thing as a tokenizer-free lunch catherinearnett • Sep 25, 2025 • 103
Running 4.05k The Ultra-Scale Playbook 🌌 4.05k The ultimate guide to training LLM on large GPU Clusters
AMDGPU onnx Collection optimized image generation ONNX models for AMD Ryzen (TM) AI GPUs and Radeon Discrete GPUs • 18 items • Updated Aug 10 • 15
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training Paper • 2505.11594 • Published May 16, 2025 • 77
Elucidating the Design Space of Diffusion-Based Generative Models Paper • 2206.00364 • Published Jun 1, 2022 • 18
PaliGemma 2 Release Collection Vision-Language Models available in multiple 3B, 10B and 28B variants. • 32 items • Updated Jul 21 • 154