-
nuprl/MultiPL-E
Viewer • Updated • 12.7k • 40.7k • 71 -
openai/openai_humaneval
Viewer • Updated • 164 • 292k • 402 -
Big Code Models Leaderboard
📈1.52kExplore and compare code model performance on a leaderboard
-
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
Paper • 2402.14261 • Published • 10
Shaun
drgitt
AI & ML interests
None yet
Recent Activity
liked a dataset 5 days ago
saidutta69/fable-5-premium liked a dataset 5 days ago
TIGER-Lab/MMLU-Pro liked a model 5 months ago
mistralai/Voxtral-4B-TTS-2603Organizations
None yet