--- library_name: transformers pipeline_tag: text-generation tags: - olmo3 - baseline - safetensors - sliding-window-attention --- # OLMo 3 1B Baseline — Stage 1 Pretraining This repository contains the pure OLMo 3 1B baseline checkpoint from `o3b1b-s8192-g8192-m2-ga4-tp1-cp1-dp1024-i1024-b4-lr1e3-fresh-1024npu-share-20260726-v1` at iteration `89407`. SiameseNorm and Depth-Attention are disabled. - Training sequence length: 8,192 - Model context capacity: 8,192 - Sliding-window size: 4,096 - Attention pattern: `[SWA, SWA, SWA, Full]` - Vocabulary: 100,278 real tokens; 74 Megatron padding rows removed Stage 3/4 apply YaRN only to Full-Attention layers. OLMo 3 SWA layers use the original RoPE and retain their 4,096-token local window. ## Loading `transformers>=4.57.6,<5` is required. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "ArchSpace-Collection/OLMo3-1B-SiameseNorm-DepthAttention-baseline-stage1" tokenizer = AutoTokenizer.from_pretrained( repo_id, use_fast=True, fix_mistral_regex=False, ) model = AutoModelForCausalLM.from_pretrained( repo_id, dtype=torch.bfloat16, attn_implementation="sdpa", ) ``` `fix_mistral_regex=False` preserves the tokenizer behavior used for training. The checkpoint uses the official Transformers `Olmo3ForCausalLM` implementation and does not require remote code.