--- library_name: transformers pipeline_tag: text-generation tags: - olmo3 - baseline - safetensors - sliding-window-attention --- # OLMo 3 1B Baseline — Stage 3 Long-context Training This repository contains the pure OLMo 3 1B baseline checkpoint from `o3b1b-s3-longmino50b-s65536-g64-m1-ga1-tp1-cp8-dp64-hsdp32-b2-lr2p5e4-w200-save1000-identity-512npu-share-20260802-v1` at iteration `11921`. SiameseNorm and Depth-Attention are disabled. - Training sequence length: 65,536 - Model context capacity: 65,536 - Sliding-window size: 4,096 - Attention pattern: `[SWA, SWA, SWA, Full]` - Vocabulary: 100,278 real tokens; 74 Megatron padding rows removed Stage 3/4 apply YaRN only to Full-Attention layers. OLMo 3 SWA layers use the original RoPE and retain their 4,096-token local window. ## Loading `transformers>=4.57.6,<5` is required. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "ArchSpace-Collection/OLMo3-1B-SiameseNorm-DepthAttention-baseline-stage3" tokenizer = AutoTokenizer.from_pretrained( repo_id, use_fast=True, fix_mistral_regex=False, ) model = AutoModelForCausalLM.from_pretrained( repo_id, dtype=torch.bfloat16, attn_implementation="sdpa", ) ``` `fix_mistral_regex=False` preserves the tokenizer behavior used for training. The checkpoint uses the official Transformers `Olmo3ForCausalLM` implementation and does not require remote code.