OpenBookQA full-test result: 32.8% for SmolLM2-135M under a pinned zero-shot setup

#14
by GerhardBalz - opened

Hi SmolLM team,

I ran the complete 500-item OpenBookQA test split for HuggingFaceTB/SmolLM2-135M to compare a pinned local configuration with the model card’s reported 34.6% result.

Configuration

  • Model revision: 1e728f6bee47d2fd840b70f44e58d4ed97c0c32d
  • Weight SHA-256: 80521b40281d6ce74e35c9282c22539e75aa0ac8578892b2a59955ef78d55da1
  • Dataset revision: allenai/openbookqa@388097ea7776314e93a529163e0fea805b8a6454
  • Dataset configuration/split: main, all 500 test rows in pinned source order
  • Task: zero-shot, four answer-text continuations prefixed with a space
  • Metric: directly audited pinned LightEval loglikelihood_acc_norm_nospace semantics—summed continuation loglikelihood divided by the continuation’s character count excluding its leading space, followed by argmax
  • Runtime: LightEval 0.6.0.dev0, Transformers 4.45.2, Torch 2.4.0+cpu, Accelerate 0.34.2
  • Effective execution: Windows CPU, BF16 parameters, FP32 logits, batch size 1, maximum length 2,048, model seed 1234, Torch intra-op/inter-op threads both 1

Result

  • 164/500 correct
  • acc_norm = 0.328 (32.8%)
  • Difference from the model-card reference: −1.8 percentage points
  • All 500 sample identities and all 2,000 planned likelihood-request identities matched
  • All 2,000 likelihoods were finite
  • Zero truncation, token-count, character-denominator, normalization, prediction, correctness, or independent-scorer discrepancies

The 95% Wilson interval is approximately 28.83%–37.03%, which contains 34.6%. That interval is only an inferential summary under a binomial/superpopulation interpretation; 32.8% is the exact observed score on this fixed test split.

This is therefore numerically compatible with 34.6% in that limited sense, but it is not an exact reproduction, confirmation, refutation, or equivalence result. The model card identifies LightEval but does not pin the evaluator version, split, backend, or complete configuration used for 34.6%.

Execution and evidence limitations

The run was offline. One expected requests→urllib3 IPv6 capability probe was intercepted before operating-system bind; prohibited socket operations and delegated network operations were zero, and 537 process-tree endpoint samples found no endpoints or audit errors.

A general loaded-module safety gate remained failed because four runtime DLLs coexisted: libiomp5md.dll, libiompstubs5md.dll, libomp140.x86_64.dll, and vcomp140.dll. The run proceeded only under a bounded exact-four-runtime, single-thread waiver. I make no general mixed-runtime safety or cross-machine reproducibility claim.

The sealed local evidence bundle contains 46,045 manifest entries:

  • Manifest SHA-256: 4bd24d34a7de4e2b61712bec22c28339683aa70005b982ca26e9f2bf07a9541d
  • Completion record SHA-256: 9ce0c3e49c59f7fabf8113190a21788b67805ac34491c53c09eef6b826f9a678
  • Controller result SHA-256: 8454ba68fbd71462683f8faf2ec2d192d3ac9631ddc643b06bb129a8739f2358
  • Independent score-audit SHA-256: ddac2ee76de0640d0154b36c16b88e3536debc6340560f03df94bf39f1dcba71

No raw artifacts are attached. The bundle includes dataset/model payloads, third-party runtime components, and machine-local paths; its redistribution, dependency-license, and privacy review is not complete.

If the evaluator revision, split, backend, or configuration used for the reported 34.6% is available, I would appreciate those details so the remaining configuration delta can be assessed.

Sign up or log in to comment