nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 Text Generation • 124B • Updated 3 days ago • 990k • • 422
view article Article Building Tensors from Scratch in Rust (Part 1.2): View Operations KeighBee • Jun 18, 2025 • 4
Running 602 Scaling test-time compute 📈 602 Boost LLM answers with flexible test‑time search strategies
Search-R1 Collection Preliminary checkpoints with outcome-only RL. • 15 items • Updated Aug 12, 2025 • 19
Running Agents 435 Reward Bench Leaderboard 📐 435 Explore and compare model scores on RewardBench benchmarks
Skywork/Skywork-Reward-Llama-3.1-8B-v0.2 Text Classification • 8B • Updated Oct 25, 2024 • 81.2k • 43
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention Paper • 2502.11089 • Published Feb 16, 2025 • 170