--- library_name: transformers pipeline_tag: text-generation tags: - olmo3 - baseline - safetensors - sliding-window-attention --- # OLMo 3 1B Baseline — Stage 2 Mid-training This repository contains the pure OLMo 3 1B baseline checkpoint from `o3b1b-s2-s8192-g256-m1-tp1-cp2-dp256-hsdp32-b8-halo-norecomp-lr2p071235e4-save10000-512npu-share-20260801-v1` at iteration `47684`. SiameseNorm and Depth-Attention are disabled. - Training sequence length: 8,192 - Model context capacity: 8,192 - Sliding-window size: 4,096 - Attention pattern: `[SWA, SWA, SWA, Full]` - Vocabulary: 100,278 real tokens; 74 Megatron padding rows removed Stage 3/4 apply YaRN only to Full-Attention layers. OLMo 3 SWA layers use the original RoPE and retain their 4,096-token local window. ## Loading `transformers>=4.57.6,<5` is required. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "ArchSpace-Collection/OLMo3-1B-SiameseNorm-DepthAttention-baseline-stage2" tokenizer = AutoTokenizer.from_pretrained( repo_id, use_fast=True, fix_mistral_regex=False, ) model = AutoModelForCausalLM.from_pretrained( repo_id, dtype=torch.bfloat16, attn_implementation="sdpa", ) ``` `fix_mistral_regex=False` preserves the tokenizer behavior used for training. The checkpoint uses the official Transformers `Olmo3ForCausalLM` implementation and does not require remote code.