Lewis Tunstall PRO
AI & ML interests
LLMs, LLMs, LLMs
Recent Activity
upvoted an article about 14 hours ago
Security incident disclosure β July 2026 liked a Space about 15 hours ago
ICML-2026-agent-repro/challenge liked a Space 5 days ago
joelniklaus/harness-optimizationOrganizations
Awesome RLHF
A curated collection of datasets, models, Spaces, and papers on Reinforcement Learning from Human Feedback (RLHF).
- RunningAgents203
MT Bench
π203Explore and compare AI model answers on benchmark questions
-
garage-bAInd/Open-Platypus
Viewer β’ Updated β’ 24.9k β’ 7.87k β’ 422 -
meta-llama/Llama-2-7b-chat-hf
Text Generation β’ 7B β’ Updated β’ 285k β’ 4.8k -
meta-llama/Llama-2-70b-chat-hf
Text Generation β’ 69B β’ Updated β’ 8.41k β’ 2.21k
Hub tools
β Long-context post-training π§Ά β
Resources for post-training LLMs with long-context samples
Mistral 7B + UltraChat + Arithmo checkpoints
A collection of Mistral 7B fine-tunes on UltraChat and Arithmo to boost the math capabilities of chat models. See https://x.com/_lewtun/status/1715652
-
lewtun/mistral-7b-sft-ultrachat-arithmo-full
Text Generation β’ Updated β’ 10 β’ β’ 1 -
lewtun/mistral-7b-sft-ultrachat-arithmo-50
Text Generation β’ Updated β’ 10 β’ β’ 1 -
lewtun/mistral-7b-sft-ultrachat-arithmo-25
Text Generation β’ Updated β’ 11 β’ -
openbmb/UltraChat
Viewer β’ Updated β’ 949k β’ 3.62k β’ 501
Gemma RLAIF
-
lewtun/gemma-7b-sft-full-ultrachat-v0
Text Generation β’ 9B β’ Updated β’ 2 β’ 1 -
lewtun/gemma-7b-sft-full-dolly-v3
Text Generation β’ 9B β’ Updated β’ 10 -
lewtun/gemma-7b-sft-full-deita-10k-v0
Text Generation β’ 9B β’ Updated β’ 5 -
lewtun/gemma-7b-dpo-full-ultrafeedback-v0
Text Generation β’ Updated β’ 4
β Awesome RL datasets π β
β Long-context post-training π§Ά β
Resources for post-training LLMs with long-context samples
Awesome RLHF
A curated collection of datasets, models, Spaces, and papers on Reinforcement Learning from Human Feedback (RLHF).
- RunningAgents203
MT Bench
π203Explore and compare AI model answers on benchmark questions
-
garage-bAInd/Open-Platypus
Viewer β’ Updated β’ 24.9k β’ 7.87k β’ 422 -
meta-llama/Llama-2-7b-chat-hf
Text Generation β’ 7B β’ Updated β’ 285k β’ 4.8k -
meta-llama/Llama-2-70b-chat-hf
Text Generation β’ 69B β’ Updated β’ 8.41k β’ 2.21k
Mistral 7B + UltraChat + Arithmo checkpoints
A collection of Mistral 7B fine-tunes on UltraChat and Arithmo to boost the math capabilities of chat models. See https://x.com/_lewtun/status/1715652
-
lewtun/mistral-7b-sft-ultrachat-arithmo-full
Text Generation β’ Updated β’ 10 β’ β’ 1 -
lewtun/mistral-7b-sft-ultrachat-arithmo-50
Text Generation β’ Updated β’ 10 β’ β’ 1 -
lewtun/mistral-7b-sft-ultrachat-arithmo-25
Text Generation β’ Updated β’ 11 β’ -
openbmb/UltraChat
Viewer β’ Updated β’ 949k β’ 3.62k β’ 501
Hub tools
Gemma RLAIF
-
lewtun/gemma-7b-sft-full-ultrachat-v0
Text Generation β’ 9B β’ Updated β’ 2 β’ 1 -
lewtun/gemma-7b-sft-full-dolly-v3
Text Generation β’ 9B β’ Updated β’ 10 -
lewtun/gemma-7b-sft-full-deita-10k-v0
Text Generation β’ 9B β’ Updated β’ 5 -
lewtun/gemma-7b-dpo-full-ultrafeedback-v0
Text Generation β’ Updated β’ 4