OLMo 3 1B Baseline โ€” Stage 4 Instruct SFT

This repository contains the pure OLMo 3 1B baseline checkpoint from o3b1b-instruct-sft-dolci-s32768-g32-m1-tp1-cp8-dp32-hsdp32-b2-lr3e5-min0-wd5e2-wu10pct-2ep-256npu-share-20260802-v1 at iteration 3252. SiameseNorm and Depth-Attention are disabled.

  • Training sequence length: 32,768
  • Model context capacity: 65,536
  • Sliding-window size: 4,096
  • Attention pattern: [SWA, SWA, SWA, Full]
  • Vocabulary: 100,278 real tokens; 74 Megatron padding rows removed

Stage 3/4 apply YaRN only to Full-Attention layers. OLMo 3 SWA layers use the original RoPE and retain their 4,096-token local window.

Loading

transformers>=4.57.6,<5 is required.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "ArchSpace-Collection/OLMo3-1B-SiameseNorm-DepthAttention-baseline-stage4-instruct"
tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    use_fast=True,
    fix_mistral_regex=False,
)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    dtype=torch.bfloat16,
    attn_implementation="sdpa",
)

fix_mistral_regex=False preserves the tokenizer behavior used for training. The checkpoint uses the official Transformers Olmo3ForCausalLM implementation and does not require remote code.

Downloads last month
14
Safetensors
Model size
1B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including ArchSpace-Collection/OLMo3-1B-stage4-instruct