-
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
Paper • 2402.12030 • Published • 4 -
mistralai/Mistral-7B-Instruct-v0.2
Text Generation • 7B • Updated • 1.17M • • 3.2k -
meta-llama/Llama-2-7b-chat-hf
Text Generation • 7B • Updated • 539k • 4.82k -
EleutherAI/pythia-160m-deduped
Text Generation • 0.2B • Updated • 433k • 4
Collections
Discover the best community collections!
Collections trending this week
-
How to Train Data-Efficient LLMs
Paper • 2402.09668 • Published • 43 -
An Introduction to Vision-Language Modeling
Paper • 2405.17247 • Published • 91 -
Agency Is Frame-Dependent
Paper • 2502.04403 • Published • 23 -
Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
Paper • 2502.04404 • Published • 25
-
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
Paper • 2402.12030 • Published • 4 -
mistralai/Mistral-7B-Instruct-v0.2
Text Generation • 7B • Updated • 1.17M • • 3.2k -
meta-llama/Llama-2-7b-chat-hf
Text Generation • 7B • Updated • 539k • 4.82k -
EleutherAI/pythia-160m-deduped
Text Generation • 0.2B • Updated • 433k • 4
-
How to Train Data-Efficient LLMs
Paper • 2402.09668 • Published • 43 -
An Introduction to Vision-Language Modeling
Paper • 2405.17247 • Published • 91 -
Agency Is Frame-Dependent
Paper • 2502.04403 • Published • 23 -
Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
Paper • 2502.04404 • Published • 25