mossez-systems's picture
Publish Mossez-100M-Coder-Instruct
8dc5e05 verified
|
Raw
History Blame Contribute Delete
1.48 kB

Evaluation

Evaluation used frozen, source-group-disjoint validation and test splits with assistant-only loss. Lower is better.

Model Validation loss Test loss
Source Coder-Base 2.008803 2.023115
One-epoch Coder-Instruct 0.029855 0.032257

Per-task loss

Task Source val Candidate val Source test Candidate test
bug_fix 0.666634 0.002418 0.685901 0.002504
completion 1.098412 0.072219 1.184431 0.074582
explain 3.757432 0.003248 3.861036 0.003465
fim_repair 2.926159 0.171918 2.934412 0.188182
json_yaml_conversion 1.304223 0.002704 1.310993 0.002706
refactor 1.088105 0.002698 1.089293 0.002557
shell 2.545998 0.001775 2.548895 0.001855
short_function 1.503674 0.003699 1.520041 0.003951
sql 2.761549 0.018859 2.715693 0.021215
traceback 3.734223 0.002768 3.769516 0.002884
unit_test 1.751150 0.075700 1.721415 0.083690

Overall and all 11 per-task losses improved against both the source and the 100-step SFT pilot. A deterministic 11-prompt generation smoke passed all local syntax/format/safety validators; 9 outputs exactly matched references.

These results come from a small, project-authored, template-heavy corpus. They do not establish performance on HumanEval, MBPP, SWE-bench, security tasks, repository-level work, or arbitrary real-world prompts.