Very low acceptance rate compared to eagle-mla
Hello!
My real test are very different from the evalutation from the model page. I get only 10-25% acceptance rate on this model with n=7 while i was getting 50%+ acceptance rate with https://huggingface.co/novita/kimi-k2.7-code-eagle3-mla
Main usage - coding agents
I am doing something wrong?
(APIServer pid=1) INFO 07-22 07:30:57 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.05, Accepted throughput: 120.49 tokens/s, Drafted throughput: 803.50 tokens/s, Accepted: 1205 tokens, Drafted: 8036 tokens, Per-position acceptance rate: 0.365, 0.214, 0.151, 0.106, 0.086, 0.071, 0.057, Avg Draft acceptance rate: 15.0%
(APIServer pid=1) INFO 07-22 07:31:07 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.31, Accepted throughput: 113.99 tokens/s, Drafted throughput: 611.06 tokens/s, Accepted: 1140 tokens, Drafted: 6111 tokens, Per-position acceptance rate: 0.391, 0.259, 0.197, 0.158, 0.129, 0.096, 0.076, Avg Draft acceptance rate: 18.7%
(APIServer pid=1) INFO 07-22 07:31:17 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.84, Accepted throughput: 111.18 tokens/s, Drafted throughput: 930.84 tokens/s, Accepted: 1112 tokens, Drafted: 9310 tokens, Per-position acceptance rate: 0.286, 0.173, 0.123, 0.092, 0.063, 0.056, 0.044, Avg Draft acceptance rate: 11.9%
(APIServer pid=1) INFO 07-22 07:31:27 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.62, Accepted throughput: 153.69 tokens/s, Drafted throughput: 664.97 tokens/s, Accepted: 1537 tokens, Drafted: 6650 tokens, Per-position acceptance rate: 0.453, 0.317, 0.253, 0.200, 0.165, 0.128, 0.102, Avg Draft acceptance rate: 23.1%
(APIServer pid=1) INFO 07-22 07:31:37 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.00, Accepted throughput: 109.49 tokens/s, Drafted throughput: 766.42 tokens/s, Accepted: 1095 tokens, Drafted: 7665 tokens, Per-position acceptance rate: 0.337, 0.220, 0.154, 0.105, 0.079, 0.061, 0.043, Avg Draft acceptance rate: 14.3%
(APIServer pid=1) INFO 07-22 07:31:47 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.90, Accepted throughput: 100.23 tokens/s, Drafted throughput: 775.75 tokens/s, Accepted: 1003 tokens, Drafted: 7763 tokens, Per-position acceptance rate: 0.327, 0.204, 0.142, 0.105, 0.062, 0.039, 0.025, Avg Draft acceptance rate: 12.9%
(APIServer pid=1) INFO 07-22 07:31:57 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.02, Accepted throughput: 123.19 tokens/s, Drafted throughput: 843.41 tokens/s, Accepted: 1232 tokens, Drafted: 8435 tokens, Per-position acceptance rate: 0.376, 0.229, 0.148, 0.110, 0.074, 0.052, 0.033, Avg Draft acceptance rate: 14.6%
(APIServer pid=1) INFO 07-22 07:32:07 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.09, Accepted throughput: 198.84 tokens/s, Drafted throughput: 1281.32 tokens/s, Accepted: 1989 tokens, Drafted: 12817 tokens, Per-position acceptance rate: 0.475, 0.283, 0.152, 0.085, 0.046, 0.028, 0.016, Avg Draft acceptance rate: 15.5%
(APIServer pid=1) INFO 07-22 07:32:17 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.28, Accepted throughput: 247.16 tokens/s, Drafted throughput: 1353.58 tokens/s, Accepted: 2472 tokens, Drafted: 13538 tokens, Per-position acceptance rate: 0.480, 0.297, 0.205, 0.131, 0.086, 0.049, 0.031, Avg Draft acceptance rate: 18.3%