utdawn commited on
Commit
f1d89a1
·
verified ·
1 Parent(s): c5428b3

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -34,7 +34,7 @@ The following tables compare **LLaDA2.2-flash** and **Ling-2.6-flash** in terms
34
 
35
  > **LLaDA2.2-flash evaluation setup:** The SWE-bench series was evaluated using the Claude Code scaffold. Across all benchmarks, we used a 128K context window with `temperature=1.0`, `block_length=32`, `threshold=0.5`, and `editing_threshold=0.0`. Each score represents the average of five runs.
36
 
37
- <sup>†</sup> The Ling-2.6-flash score on SWE-bench Verified is taken from the <a href="https://arxiv.org/abs/2606.15079">Ling and Ring 2.6 Technical Report</a>, where it was obtained using the OpenHands scaffold. The Ling-2.6-flash scores on τ²-Bench, Claw-Eval, and PinchBench are also sourced from the technical report, whereas its SWE-bench Pro and SWE-bench Multilingual scores were evaluated by us using the same Claude Code scaffold as LLaDA2.2-flash.
38
 
39
 
40
  **Throughput (TPS)**
 
34
 
35
  > **LLaDA2.2-flash evaluation setup:** The SWE-bench series was evaluated using the Claude Code scaffold. Across all benchmarks, we used a 128K context window with `temperature=1.0`, `block_length=32`, `threshold=0.5`, and `editing_threshold=0.0`. Each score represents the average of five runs.
36
 
37
+ <sup>†</sup> The Ling-2.6-flash score on SWE-bench Verified is taken from the Ling and Ring 2.6 Technical Report, where it was obtained using the OpenHands scaffold. The Ling-2.6-flash scores on τ²-Bench, Claw-Eval, and PinchBench are also sourced from the technical report, whereas its SWE-bench Pro and SWE-bench Multilingual scores were evaluated by us using the same Claude Code scaffold as LLaDA2.2-flash.
38
 
39
 
40
  **Throughput (TPS)**