Slayer149 Balanced β€” ARC-Easy fine-tune (seed 45500)

A full-weight ARC-Easy fine-tune of the 149,333,081-parameter Slayer149 Balanced checkpoint at step 45,500. This repository contains the predeclared primary seed (45500), selected at epoch 3 by validation accuracy. It is a research artifact for inspection and reproduction, not a general-purpose instruction model.

Evaluation

Result Accuracy Count
ARC-Easy validation, selected epoch 59.47% 339 / 570
ARC-Easy test, evaluated once after selection 60.23% 1,431 / 2,376

The test split did not select the epoch or recipe. Two additional fixed-recipe seeds scored 60.52% (1,438/2,376) and 59.93% (1,424/2,376); see replication-summary.json. Do not treat the best seed's test score as a model-selection criterion.

This score uses the Glint-1.3 ARC protocol: zero-shot raw summed log-likelihood of question + space + choice, first 256 tokens, no added BOS, with raw accuracy. The exact dataset revision, tokenizer hash, evaluator hash, and checkpoint provenance are in test-manifest.json. It is not an EleutherAI lm-evaluation-harness arc_easy result: that task uses a Question: ...\nAnswer: prompt and reports both acc and length-normalized acc_norm. Values from these protocols should not be compared as identical measurements.

The fine-tuning set was ARC-Easy train (2,251 questions); validation was ARC-Easy validation (570). The held-out test set (2,376) was not used for training or selection. Dataset revision: 210d026faf9955653af8916fad021475a3f00453.

Load locally

from load_model import load_model
model, tokenizer = load_model(".", device="cpu")
ids = tokenizer.encode("What is the result?", add_special_tokens=False).ids
import torch
with torch.inference_mode():
    logits, _ = model(torch.tensor([ids]))

The model uses a custom PyTorch architecture; this package does not register Transformers remote code. model.safetensors is FP32 and was round-trip checked against the selected training checkpoint.

Provenance

  • Base: SlayerLab/Slayer149-balanced, checkpoint checkpoints/000045500/training-state.pt; SHA-256 7b63994219a35281cdb4a69eba3fb861c8f0bf3e58a293883fc5a98c3776d21a.
  • Fine-tuning: listwise cross-entropy over raw answer likelihoods; 5 epochs maximum, AdamW at 5e-6, batch of 8 questions; epoch selected by validation accuracy.
  • Seed: 45500. Selected epoch: 3.
  • Fine-tuned checkpoint SHA-256: 8e62ae4f586db75b8c76396f716676f462fb3b89ec1ed80b34620d1a1e54c939.
  • FP32 safetensors SHA-256: c3dcdcb84dd709d8c1cd17bbe35ba2caed01df57a68b61266eb2ee9d4a91d7b7.
  • balanced_model.py, config.json, and tokenizer.json provide the architecture and tokenizer required to load it.

The base model is Apache-2.0. ARC is credited to AI2 and the Hugging Face dataset card lists CC-BY-SA-4.0; the dataset itself is not redistributed here.

Downloads last month
15
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for SlayerLab/Slayer149-Balanced-ARC-E-ft-45500

Finetuned
(1)
this model

Spaces using SlayerLab/Slayer149-Balanced-ARC-E-ft-45500 2