Add ExtractBench evaluation results

#28
Files changed (1) hide show
  1. .eval_results/extractbench.yaml +40 -0
.eval_results/extractbench.yaml ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: llamaindex/ExtractBench
3
+ task_id: mean
4
+ value: 45.19
5
+ date: '2026-08-26'
6
+ source:
7
+ url: https://huggingface.co/datasets/llamaindex/ExtractBench
8
+ name: ExtractBench
9
+ user: boyang-runllama
10
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"
11
+ - dataset:
12
+ id: llamaindex/ExtractBench
13
+ task_id: short
14
+ value: 47.28
15
+ date: '2026-08-26'
16
+ source:
17
+ url: https://huggingface.co/datasets/llamaindex/ExtractBench
18
+ name: ExtractBench
19
+ user: boyang-runllama
20
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"
21
+ - dataset:
22
+ id: llamaindex/ExtractBench
23
+ task_id: medium
24
+ value: 44.52
25
+ date: '2026-08-26'
26
+ source:
27
+ url: https://huggingface.co/datasets/llamaindex/ExtractBench
28
+ name: ExtractBench
29
+ user: boyang-runllama
30
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"
31
+ - dataset:
32
+ id: llamaindex/ExtractBench
33
+ task_id: long
34
+ value: 22.17
35
+ date: '2026-08-26'
36
+ source:
37
+ url: https://huggingface.co/datasets/llamaindex/ExtractBench
38
+ name: ExtractBench
39
+ user: boyang-runllama
40
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"