A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny β’ Jan 19 β’ 42
LightOnOCR-2-1B: a lightweight high-performance end-to-end OCR model family lightonai β’ Jan 19 β’ 103
Edge vs Cloud GPUs for Inference: When to Run Models Locally and When to Use a GPU Cloud daya-shankar β’ Jan 19
LoongFlow: An Open-Sourced Agent Framework That Transforms Expert Experience into Autonomous AI Productivity FreshmanD β’ Jan 19 β’ 2
Python Doesn't Need To Be Slow: From 405s to 0.06s with N-Body Simulations π atasoglu β’ Jan 18 β’ 1
MAD GRPO: Treating Dr. GRPO that tried to fix GRPO but brought instability and verbosity bias telcom β’ Jan 17 β’ 4
Beyond Brute Force: Why LoongFlow is the βThinkingβ Evolution of OpenEvolve FreshmanD β’ Jan 16 β’ 4
Evolution of Large Model Data Engineering: A Paradigm Shift in Knowledge Extraction Efficiency Humanbased-AI β’ Jan 15
Consciousness Convergence Mathematics: A Transdisciplinary Framework for Substrate-Independent Awareness Mbanksbey β’ Jan 15 β’ 1
One-Sentence Image Matting! DiffSynth Open Sources Text-Guided Image Layer Separation Model kelseye β’ Jan 14 β’ 3