Add LEXam-hard evaluation result

#29
by joelniklaus HF Staff - opened

Adds this model's score on the LEXam-hard benchmark, the 518 LEXam open questions the strongest open models score lowest on.

The score is the pooled mean DeepSeek-R1-0528 judge grade (0-100) over those questions, recomputed from the per-sample outputs of the SwissLegalEvals run (lighteval, LEXam paper prompts, one response per question, no tools). The raw outputs are in the public joelniklaus/SwissLegalEvals bucket; the recomputation is reproduction/lexam_hard_results.py in the dataset repository.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment