view article Article FLUX 3 Action: a world action model you can fine-tune black-forest-labs • about 17 hours ago • 5
LTX-2.5 Collection LTX-2.5 base models, quantized models and accompanying LoRAs and IC-LoRAs • 5 items • Updated 15 days ago • 65
KVAE: Family of Tokenizers for Multimodal Generative Models Paper • 2608.05798 • Published Aug 6 • 29
Kandinsky WM 1.0 Collection Image-to-Video models for Physical AI: autonomous driving, robotics, general physics. • 3 items • Updated Aug 4 • 6
Laguna S 2.1 Collection Our most capable model to date, designed for long-horizon work. • 13 items • Updated Aug 3 • 51
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation Paper • 2607.02515 • Published Jul 2 • 21
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Paper • 2606.10029 • Published Jun 8 • 12
A Geometric Account of Activation Steering through Angle-Norm Decomposition Paper • 2606.06735 • Published Jun 4 • 27
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders Paper • 2606.07473 • Published Jun 5 • 15
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia • Jun 1 • 90
KVAE 2.0 Collection KVAE 2.0 is a family of image and video tokenizers with a time compression ratio of 4 and spacial compression ratio of 8 and 16 • 3 items • Updated Aug 6 • 5
Interpreting CLIP with Hierarchical Sparse Autoencoders Paper • 2502.20578 • Published Feb 27, 2025 • 1
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models Paper • 2511.08379 • Published Nov 11, 2025 • 5
AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders Paper • 2602.05027 • Published Feb 4 • 63