Slayer149 Balanced β ARC-Easy fine-tune (seed 45500)
A full-weight ARC-Easy fine-tune of the 149,333,081-parameter Slayer149 Balanced checkpoint at step 45,500. This repository contains the predeclared primary seed (45500), selected at epoch 3 by validation accuracy. It is a research artifact for inspection and reproduction, not a general-purpose instruction model.
Evaluation
| Result | Accuracy | Count |
|---|---|---|
| ARC-Easy validation, selected epoch | 59.47% | 339 / 570 |
| ARC-Easy test, evaluated once after selection | 60.23% | 1,431 / 2,376 |
The test split did not select the epoch or recipe. Two additional fixed-recipe seeds scored 60.52% (1,438/2,376) and 59.93% (1,424/2,376); see replication-summary.json. Do not treat the best seed's test score as a model-selection criterion.
This score uses the Glint-1.3 ARC protocol: zero-shot raw summed log-likelihood of question + space + choice, first 256 tokens, no added BOS, with raw accuracy. The exact dataset revision, tokenizer hash, evaluator hash, and checkpoint provenance are in test-manifest.json. It is not an EleutherAI lm-evaluation-harness arc_easy result: that task uses a Question: ...\nAnswer: prompt and reports both acc and length-normalized acc_norm. Values from these protocols should not be compared as identical measurements.
The fine-tuning set was ARC-Easy train (2,251 questions); validation was ARC-Easy validation (570). The held-out test set (2,376) was not used for training or selection. Dataset revision: 210d026faf9955653af8916fad021475a3f00453.
Load locally
from load_model import load_model
model, tokenizer = load_model(".", device="cpu")
ids = tokenizer.encode("What is the result?", add_special_tokens=False).ids
import torch
with torch.inference_mode():
logits, _ = model(torch.tensor([ids]))
The model uses a custom PyTorch architecture; this package does not register Transformers remote code. model.safetensors is FP32 and was round-trip checked against the selected training checkpoint.
Provenance
- Base:
SlayerLab/Slayer149-balanced, checkpointcheckpoints/000045500/training-state.pt; SHA-2567b63994219a35281cdb4a69eba3fb861c8f0bf3e58a293883fc5a98c3776d21a. - Fine-tuning: listwise cross-entropy over raw answer likelihoods; 5 epochs maximum, AdamW at 5e-6, batch of 8 questions; epoch selected by validation accuracy.
- Seed: 45500. Selected epoch: 3.
- Fine-tuned checkpoint SHA-256:
8e62ae4f586db75b8c76396f716676f462fb3b89ec1ed80b34620d1a1e54c939. - FP32 safetensors SHA-256:
c3dcdcb84dd709d8c1cd17bbe35ba2caed01df57a68b61266eb2ee9d4a91d7b7. balanced_model.py,config.json, andtokenizer.jsonprovide the architecture and tokenizer required to load it.
The base model is Apache-2.0. ARC is credited to AI2 and the Hugging Face dataset card lists CC-BY-SA-4.0; the dataset itself is not redistributed here.
- Downloads last month
- 15
Model tree for SlayerLab/Slayer149-Balanced-ARC-E-ft-45500
Base model
SlayerLab/Slayer149-balanced