--- library_name: transformers pipeline_tag: text-generation tags: - olmo3 - baseline - safetensors - sliding-window-attention --- # OLMo 3 1B Baseline — Stage 4 Instruct SFT This repository contains the pure OLMo 3 1B baseline checkpoint from `o3b1b-instruct-sft-dolci-s32768-g32-m1-tp1-cp8-dp32-hsdp32-b2-lr3e5-min0-wd5e2-wu10pct-2ep-256npu-share-20260802-v1` at iteration `3252`. SiameseNorm and Depth-Attention are disabled. - Training sequence length: 32,768 - Model context capacity: 65,536 - Sliding-window size: 4,096 - Attention pattern: `[SWA, SWA, SWA, Full]` - Vocabulary: 100,278 real tokens; 74 Megatron padding rows removed Stage 3/4 apply YaRN only to Full-Attention layers. OLMo 3 SWA layers use the original RoPE and retain their 4,096-token local window. ## Loading `transformers>=4.57.6,<5` is required. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "ArchSpace-Collection/OLMo3-1B-SiameseNorm-DepthAttention-baseline-stage4-instruct" tokenizer = AutoTokenizer.from_pretrained( repo_id, use_fast=True, fix_mistral_regex=False, ) model = AutoModelForCausalLM.from_pretrained( repo_id, dtype=torch.bfloat16, attn_implementation="sdpa", ) ``` `fix_mistral_regex=False` preserves the tokenizer behavior used for training. The checkpoint uses the official Transformers `Olmo3ForCausalLM` implementation and does not require remote code.