--- license: apache-2.0 language: - en pipeline_tag: text-generation tags: - pkai - text-generation - base-model - safetensors --- # PKaI Nano 1.2 PKaI Nano 1.2 is a 300M-class base language model from PowderKeg Intelligence. It is a compact LLaMA-style decoder model trained from scratch with a Mistral tokenizer and released as a PKaI-native artifact. It succeeds PKaI Nano 1.1 with a broader, higher-quality training mix and improves on it across nearly every benchmark in our evaluation suite. This repository contains PKaI-native weights and metadata, not a drop-in `transformers.AutoModelForCausalLM` package. 📖 For the full writeup and evaluation details, see the announcement post: [Introducing PKaI Nano 1.2: Better Data, Stronger Results](https://powderkeg.info/ai-research.html?post=introducing-pkai-nano-1-2). ## Benchmarks Evaluated in our standard benchmark suite against the previous Nano releases. Accuracy values are percentages; higher is better. WikiText is word perplexity; lower is better. | Benchmark | PKaI Nano 1.2 | PKaI Nano 1.1 | PKaI Nano 1 | |---|---|---|---| | HellaSwag | **39.71** | 36.52 | 31.02 | | SciQ | **81.50** | 80.00 | 73.00 | | PIQA | **65.94** | 64.96 | 59.74 | | WinoGrande | **52.96** | 52.72 | 52.57 | | ARC-Easy | **49.92** | 45.45 | 43.35 | | ARC-Challenge | 27.73 | **28.24** | 24.83 | | LAMBADA OpenAI | **44.07** | 33.82 | 23.33 | | WikiText (ppl ↓) | **24.13** | 35.47 | 57.37 | Across the seven accuracy tasks, PKaI Nano 1.2 averages **51.69**, up from 48.82 for PKaI Nano 1.1 and 43.98 for PKaI Nano 1. See the announcement post for the full writeup and evaluation notes. PKaI Nano 1 and PKaI Nano 1.1 were re-scored under our updated evaluation methodology for this comparison; their figures may differ slightly from their original release notes. ## Files - `model.safetensors`: PKaI base model weights. - `config.json`: PKaI model architecture configuration. - `tokenizer.json`: PKaI tokenizer metadata. - `tokenizer.model`: SentencePiece tokenizer model from `mistralai/Mistral-7B-v0.1`. - `THIRD_PARTY_NOTICES.txt`: tokenizer and training-data provenance notices. - `LICENSE`: Apache License, Version 2.0. ## Architecture - Parameters: 311,218,176 - Vocabulary size: 32,000 - Context length: 1024 - Layers: 24 - Attention heads: 16 - KV heads: 8 - Embedding size: 1024 - Tied embeddings: yes - QK normalization: yes ## Training Data Training data included the following publicly disclosed sources, each processed with best-effort in-house decontamination and deduplication by PowderKeg Intelligence prior to training: - `HuggingFaceFW/fineweb`, released under the Open Data Commons Attribution License (ODC-By) v1.0 and subject to the Common Crawl Terms of Use as noted by the dataset card. - `HuggingFaceFW/fineweb-edu`, released under the Open Data Commons Attribution License (ODC-By) v1.0 and subject to the Common Crawl Terms of Use as noted by the dataset card. - `HuggingFaceTB/smollm-corpus` (`cosmopedia-v2` subset), released under the Open Data Commons Attribution License (ODC-By) v1.0. Cosmopedia v2 is synthetic text generated with `mistralai/Mixtral-8x7B-Instruct-v0.1`. - `wikimedia/wikipedia` (20231101.en snapshot), released under the Creative Commons Attribution-ShareAlike 4.0 License (CC BY-SA 4.0). - `open-web-math/open-web-math`, released under the Open Data Commons Attribution License (ODC-By) v1.0 and subject to the Common Crawl Terms of Use as noted by the dataset card. The training mix also included public-domain book text and other public text sources. See `THIRD_PARTY_NOTICES.txt` for source URLs, licenses, and attribution notes. ## License Copyright 2026 PowderKeg Intelligence LLC. The PKaI Nano 1.2 model artifact is released under the Apache License, Version 2.0. The bundled tokenizer and training-data sources have their own provenance and notices listed in `THIRD_PARTY_NOTICES.txt`. ## Limitations This is a small base model and has not been instruction-tuned or safety-tuned. It may produce inaccurate, unsafe, biased, or otherwise unsuitable text. Users are responsible for evaluating fitness, safety, and legal compliance for their own use cases.