etomoscow/tool-calling-hallucination-modernbert-large-crf-best Token Classification • 0.4B • Updated Aug 21
view article Article A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes ybelkada, timdettmers • Aug 17, 2022 • 140
view article Article DualPipe Explained: A Comprehensive Guide to DualPipe That Anyone Can Understand—Even Without a Distributed Training Background NormalUhr • Feb 28, 2025 • 22
view article Article Efficient LLM Pretraining: Packed Sequences and Masked Attention sirluk • Oct 7, 2024 • 74
Running 4.05k The Ultra-Scale Playbook 🌌 4.05k The ultimate guide to training LLM on large GPU Clusters
The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar Paper • 2606.26015 • Published Jun 24 • 10
etomoscow/tool-calling-hallucination-modernbert-large-crf-best Token Classification • 0.4B • Updated Aug 21