SaylorTwift HF Staff commited on
Commit
2d894c6
·
verified ·
1 Parent(s): 84c6a6a

Add evaluation results

Browse files

# PR Description: Add Evaluation Results for zai-org/GLM-5.3-Flash

## Summary

This PR adds evaluation results extracted from the model card's benchmark graph for `zai-org/GLM-5.3-Flash` to the `.eval_results/` directory, following the [Hugging Face Hub evaluation-results specification](https://huggingface.co/docs/hub/eval-results).

## Benchmarks Added

- [Terminal-Bench 2.1](https://huggingface.co/datasets/harborframework/terminal-bench-2.1?eval_result=zai-org/GLM-5.3-Flash&leaderboard_task_id=terminalbench_2_1) — 84.3
- [DeepSWE](https://huggingface.co/datasets/datacurve/deep-swe?eval_result=zai-org/GLM-5.3-Flash&leaderboard_task_id=deep_swe) — 63.4
- [Humanity's Last Exam](https://huggingface.co/datasets/cais/hle?eval_result=zai-org/GLM-5.3-Flash&leaderboard_task_id=hle) — 55.3

## Benchmarks Skipped (Not Registered on Hub)

The following benchmarks were present in the model card but could not be added because they do not have a registered `eval.yaml` on the Hugging Face Hub:

- **Agent's Last Exam**: 26.3 — no registered `eval.yaml` found on the Hub.
- **AutomationBench v1.0.6**: 48.8 — no registered `eval.yaml` found on the Hub.
- **GDPval-AA v2**: 1773 — no registered `eval.yaml` found on the Hub.

These can be added once the benchmark authors register their `eval.yaml` on the Hub.

## Source

- Model card: https://huggingface.co/zai-org/GLM-5.3-Flash
- Note: the model card cites `arxiv:2602.15763` as "the GLM-5 Technical report", but that paper (published Feb 2026) is the original GLM-5 report and predates this model — it does not contain GLM-5.3-Flash-specific numbers, so it wasn't used as a source.

## Files Added

- `.eval_results/GLM-5.3-Flash.yaml`

## Verification

These results were extracted from the model card's published benchmark chart (`bench_53.png`, an image embedded in the README, read visually — the announcement blog is a client-rendered SPA and wasn't fetchable). No verified token is provided as these were not run via HF Jobs with inspect-ai.

---

**To upload this file to the Hub, run:**

```bash
hf upload zai-org/GLM-5.3-Flash --type model --include .eval_results/*.yaml --commit-message "Add evaluation results"
```

Files changed (1) hide show
  1. .eval_results/GLM-5.3-Flash.yaml +28 -0
.eval_results/GLM-5.3-Flash.yaml ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: harborframework/terminal-bench-2.1
3
+ task_id: terminalbench_2_1
4
+ value: 84.3
5
+ date: "2026-08-26"
6
+ source:
7
+ url: https://huggingface.co/zai-org/GLM-5.3-Flash
8
+ name: "GLM-5.3-Flash model card"
9
+
10
+ - dataset:
11
+ id: datacurve/deep-swe
12
+ task_id: deep_swe
13
+ value: 63.4
14
+ date: "2026-08-26"
15
+ source:
16
+ url: https://huggingface.co/zai-org/GLM-5.3-Flash
17
+ name: "GLM-5.3-Flash model card"
18
+ notes: "Reported as DeepSWE v1.1 on the model card, run via the mini-swe-agent harness with 400K context."
19
+
20
+ - dataset:
21
+ id: cais/hle
22
+ task_id: hle
23
+ value: 55.3
24
+ date: "2026-08-26"
25
+ source:
26
+ url: https://huggingface.co/zai-org/GLM-5.3-Flash
27
+ name: "GLM-5.3-Flash model card"
28
+ notes: "HLE with tools (full set) and a 300K-context management strategy, not the no-tools default; judged by GPT-5.6-luna (medium)."