boyang-runllama commited on
Commit
43b79f0
·
verified ·
1 Parent(s): 2426fa6

Update ExtractBench results (shared 32k-budget fairness re-run)

Browse files
Files changed (1) hide show
  1. .eval_results/extractbench.yaml +12 -12
.eval_results/extractbench.yaml CHANGED
@@ -1,40 +1,40 @@
1
  - dataset:
2
  id: llamaindex/ExtractBench
3
  task_id: mean
4
- value: 41.76
5
- date: '2026-08-24'
6
  source:
7
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
8
  name: ExtractBench
9
  user: boyang-runllama
10
- notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file"
11
  - dataset:
12
  id: llamaindex/ExtractBench
13
  task_id: short
14
- value: 47.31
15
- date: '2026-08-24'
16
  source:
17
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
18
  name: ExtractBench
19
  user: boyang-runllama
20
- notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file"
21
  - dataset:
22
  id: llamaindex/ExtractBench
23
  task_id: medium
24
- value: 32.71
25
- date: '2026-08-24'
26
  source:
27
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
28
  name: ExtractBench
29
  user: boyang-runllama
30
- notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file"
31
  - dataset:
32
  id: llamaindex/ExtractBench
33
  task_id: long
34
- value: 16.16
35
- date: '2026-08-24'
36
  source:
37
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
38
  name: ExtractBench
39
  user: boyang-runllama
40
- notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file"
 
1
  - dataset:
2
  id: llamaindex/ExtractBench
3
  task_id: mean
4
+ value: 45.19
5
+ date: '2026-08-26'
6
  source:
7
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
8
  name: ExtractBench
9
  user: boyang-runllama
10
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"
11
  - dataset:
12
  id: llamaindex/ExtractBench
13
  task_id: short
14
+ value: 47.28
15
+ date: '2026-08-26'
16
  source:
17
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
18
  name: ExtractBench
19
  user: boyang-runllama
20
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"
21
  - dataset:
22
  id: llamaindex/ExtractBench
23
  task_id: medium
24
+ value: 44.52
25
+ date: '2026-08-26'
26
  source:
27
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
28
  name: ExtractBench
29
  user: boyang-runllama
30
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"
31
  - dataset:
32
  id: llamaindex/ExtractBench
33
  task_id: long
34
+ value: 22.17
35
+ date: '2026-08-26'
36
  source:
37
  url: https://huggingface.co/datasets/llamaindex/ExtractBench
38
  name: ExtractBench
39
  user: boyang-runllama
40
+ notes: "Pipeline name: glm_4_6v_flash_vllm_extract_oneshot_structured_output_file; 64k output budget (32k answer + 32k reasoning allowance for the thinking mode), single generation"