I reproduced the test-set results using the optimal hyperparameters, but the scores are lower than official report, is my parameters wrong?

#27
by pengwenzhi - opened

20260820-093705

OpenBMB org

Hi! Try increase the max output tokens to 65536? For some deep think datasets, it would be better to set max output tokens to 80k.

image
I reproduced the test‑set results using 64k, but there is a large gap between my results and those reported.

Sign up or log in to comment