Pre-training Dataset Samples Collection A collection of pre-training datasets samples of sizes 10M, 100M and 1B tokens. Ideal for use in quick experimentation and ablations. • 15 items • Updated Apr 2 • 18
view article Article Meta is back with Muse Glimmer: local, agentic, multimodal, and open source +2 pcuenq, merve, burtenshaw, ariG23498 • 10 days ago • 104
📚 LLM pretraining datasets Collection A collection of datasets for LLM pretraining • 9 items • Updated May 5, 2025 • 27
LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling Paper • 2606.04438 • Published Jun 3 • 2
EMO: Pretraining Mixture of Experts for Emergent Modularity Paper • 2605.06663 • Published May 7 • 13
LTX-2.3 Creative Lab Collection LoRAs and IC-LoRAs, trained on the LTX-2.3 model • 26 items • Updated 23 days ago • 89
Ornith-1.0 Collection Ornith-1.0 is a family of open-source LLMs specialized for agentic coding. • 8 items • Updated about 20 hours ago • 382
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer Paper • 2503.07027 • Published Mar 10, 2025 • 30
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model Paper • 2603.21986 • Published Mar 23 • 125
view article Article 🎨 VAE (Variational Autoencoders) — When AI learns to dream! 🌈🧠 RDTvlokip • Oct 19, 2025 • 4