| # HumanEval — merged | |
| - Model: `/workspace/checkpoints/qwen36_a100_stage1_merged` | |
| - pass@1: **81.71%** (134/164) | |
| ## Reference pass@1 (published, approx) | |
| - Qwen2.5-Coder-32B-Instruct: 92% | |
| - GPT-4o: 90% | |
| - DeepSeek-V3: 91% | |
| - Llama-3.1-70B-Instruct: 80% | |
| - Qwen2.5-7B-Instruct: 84% | |
| - CodeLlama-34B: 51% | |