--- license: apache-2.0 datasets: - HuggingFaceFW/fineweb-edu - EleutherAI/the_pile_deduplicated - HuggingFaceTB/dclm-edu - HuggingFaceTB/finemath - HuggingFaceTB/smollm-corpus - wikimedia/wikipedia - Harley-ml/lesswrong - Harley-ml/HFMC - AxiomicLabs/NPset-2-Python-Edu language: - en tags: - er - fromziro - fromzero - harley-ml - lyjonathan - small - slm - orez - sfz --- **Note**: This model belongs to the **Er** SLM family. All models in the Er family are trained using the same tokenizer, dataset, and token count. # Er-Large ## Summary ``` Task: Text-Generation Total training time: 116 hours Inputs: text Outputs: text Params: 31,944,632 Final Loss: 2.137 Important Benchmark Scores: 1. ARC Easy - 37.54% 2. HellaSwag - 30.52% 3. ArithMark-2.0 - 32.56% Framework: PyTorch, transformers Author: Paul Courneya, Jonathon LY ``` ## Description ‘Er-Large’ is a 32M-parameter Small Language Model trained on 34.8B tokens from a nine-source dataset. Its name, “Er,” is the reverse of “Re,” the prefix of Re:Zero – Starting Life in Another World, the light novel series that inspired the organization’s name. ## Model Details - Architecture: Qwen3.5 - Hidden Size: 408 - Number of Layers: 16 - Intermediate Size: 1101 (a 2.7x expansion) - Number of Attention Heads: 12 - Number of KV Heads: 3 - Head Dim: 34 - Vocab Size: 2564 - Max Position Embeddings: 768 - Total Parameters: 31,944,632 ## Training ### Dataset | Source | Bytes (GB) | Share (%) | What it is | | ---------------- | ---------: | --------: | ----------------------------------------------- | | FineWeb-edu | 35.0 | 28.2% | Educational-filtered Common Crawl | | DCLM-Edu | 20.0 | 16.1% | Educational-filtered webtext | | The Pile Deduped | 20.0 | 16.1% | Broad, diverse 23-source dataset | | FineWeb-HQ | 20.0 | 16.1% | Knowledge-filtered webtext | | FineMath | 13.0 | 10.5% | Math-filtered Common Crawl | | Cosmopedia-v2 | 7.0 | 5.6% | Synthetic textbooks | | Wikipedia | 5.0 | 4.0% | Wikipedia articles | | NpSetPython-Edu | 3.5 | 2.8% | Normalized Python code | | Misc | 0.6 | 0.5% | LessWrong + HF configs + HF dataset/model cards | ### Training Details - Maximum Learning Rate: 2.3e-3 - Minimum Learning Rate: 0 - Number of Epochs: 1 - Sequence Length: 768 - Global Batch Size: 64 - Local Batch Size: 32 - Eval Split Ratio: 0.0025 - Gradient Accumulation Steps: 2 - Gradient Checkpointing: True - Gradient Clipping: 1.0 - Torch Compile: False - Torch Compile Mode: None - AdamW Betas: `(0.9, 0.95)` - WSD Warmup Ratio: 0.02 - WSD Stable Ratio: 0.73 - WSD Decay Ratio: 0.25 - DType: `bfloat16` ### Final Eval and Train Loss - Train: 2.137 - Val: 1.844 ### Hardware - GPU: Two NVIDIA RTX 5070s (used for training) - CPU: AMD Ryzen 5 2600 (used for tokenization) ## Benchmark scores | Task | Value | Random Chance | | ------------- | -----: | ------------: | | ARC Challenge | 22.17% | 25.00% | | ARC Easy | 37.54% | 25.00% | | BoolQ | 58.32% | 50.00% | | HellaSwag | 30.52% | 25.00% | | PiQA | 60.12% | 50.00% | | SciQ | 64.00% | 25.00% | | SWAG | 46.75% | 25.00% | | OpenbookQA | 28.60% | 25.00% | | Winogrande | 51.30% | 50.00% | ArithMark-2.0: | Category | Accuracy | | -------- | -------: | | ops = 1 | 33.52% | | ops = 2 | 36.53% | | ops = 3 | 24.20% | | Avg | 32.56% | For a comparison with other small language models like this one, go [here](https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard). ## Generation Sample ```text ``` ## Use Cases 1. Educational work and research 2. Fine-tuning for downstream use 3. Deployment on edge devices 4. Or just for fun. ## Limitations 1. Cannot chat, reason, code, or answer questions 2. Almost always unfactual 3. No long-context handling ## License Before using, distributing, selling, or modifying this software, you must read the license [here](https://huggingface.co/fromziro/Er-Large-30M/blob/main/LICENSE.txt). ## Inference ```python #!/usr/bin/env python3 MODEL_DIR = "fromziro/Er-Large-30M" TOKENIZER_PATH = MODEL_DIR PROMPT = "Artificial intelligence is" MAX_NEW_TOKENS = 256 TEMPERATURE = 0.7 TOP_P = 0.95 TOP_K = 30 REPETITION_PENALTY = 1.2 DO_SAMPLE = True import torch from pathlib import Path from transformers import AutoModelForCausalLM, AutoTokenizer, PreTrainedTokenizerFast device = ( "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu" ) print(f"Device : {device}") def load_tokenizer(path_or_repo: str): p = Path(path_or_repo) if p.exists() and p.is_file() and p.suffix.lower() == ".json": tok = PreTrainedTokenizerFast(tokenizer_file=str(p.resolve())) else: tok = AutoTokenizer.from_pretrained(path_or_repo, use_fast=True) if tok.bos_token is None: tok.add_special_tokens({"bos_token": "<|bos|>"}) if tok.eos_token is None: tok.add_special_tokens({"eos_token": "<|eos|>"}) if tok.unk_token is None: tok.add_special_tokens({"unk_token": "<|unk|>"}) if tok.pad_token is None: tok.pad_token = tok.eos_token if tok.eos_token is not None else "<|pad|>" tok.padding_side = "left" return tok print("Loading tokenizer...") tokenizer = load_tokenizer(TOKENIZER_PATH) print(f" Vocab size : {len(tokenizer)}") print(f" BOS : {tokenizer.bos_token!r}") print(f" EOS : {tokenizer.eos_token!r}") print(f" PAD : {tokenizer.pad_token!r} (id={tokenizer.pad_token_id})") print(f"\nLoading model from {MODEL_DIR} ...") model = AutoModelForCausalLM.from_pretrained( MODEL_DIR, torch_dtype=torch.float16 if device == "cuda" else torch.float32, low_cpu_mem_usage=True, ) model.eval() model.to(device) model.config.use_cache = False if hasattr(model, "generation_config") and model.generation_config is not None: model.generation_config.use_cache = False total_params = sum(p.numel() for p in model.parameters()) print(f" Parameters : {total_params:,}") def generate( prompt: str = PROMPT, max_new_tokens: int = MAX_NEW_TOKENS, temperature: float = TEMPERATURE, top_p: float = TOP_P, top_k: int = TOP_K, repetition_penalty: float = REPETITION_PENALTY, do_sample: bool = DO_SAMPLE, ) -> str: bos = tokenizer.bos_token or "" full_prompt = bos + prompt inputs = tokenizer( full_prompt, return_tensors="pt", add_special_tokens=False, ).to(device) inputs.pop("token_type_ids", None) gen_kwargs = dict( max_new_tokens=max_new_tokens, do_sample=do_sample, repetition_penalty=repetition_penalty, eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.pad_token_id, use_cache=False, ) if do_sample: gen_kwargs["temperature"] = temperature gen_kwargs["top_p"] = top_p gen_kwargs["top_k"] = top_k with torch.inference_mode(): output_ids = model.generate(**inputs, **gen_kwargs) prompt_len = inputs["input_ids"].shape[-1] new_ids = output_ids[0][prompt_len:] return tokenizer.decode(new_ids, skip_special_tokens=True) if __name__ == "__main__": print(f"\nPrompt : {PROMPT!r}") print("-" * 60) output = generate(PROMPT) print("Generated:") print(output) ``` ## Copyright ``` Copyright (c) 2026 FromZero Copyright (c) 2026 Paul Courneya Copyright (c) 2026 Jonathon LY ``` ## Citation ```bibtex @misc{er-large-30m, title = {Er-Large-30M}, organization = [FromZero], authors = {Paul Courneya, Jonathon LY}, year = {2026}, url = {https://huggingface.co/fromziro/Er-Large-30M] } ```