PicoPolite240M-base

Model Summary:

PicoPolite240M is a small-scale language model created using extremely limited computational resources (a single GPU with 32 GB memory) to research training methods for LLMs and SLMs. We will publish a book detailing the entire learning methodology. Furthermore, we will release the complete synthetic dataset created for this purpose.

Training Pipeline:

Using the data below, I trained the model for one epoch with a token window of 2000 and fp32 precision; the RoPE θ was set to 3200.

Subsequently, using the data below, I trained it for one epoch with a token window of 6400 and bf16 precision; the RoPE θ was set to 10000.

Afterward, I applied LongRoPE to extrapolate the token window.

License:

This model utilizes google gemma-4-E2B-it for its tokenizer and activation engineering. Therefore, the Gemma license applies.

Intended use:

For research and the development of new methods. The objective is to compare the performance of the entire learning pipeline by standardizing the computational resources used for training—rather than the number of parameters—across the systems being compared.

Resources:

  • ⭐️ Comprehensive guide to learning methods ➡️ Coming soon
  • 📄 Models created under the same conditions ➡️ Scratch’n 32GB-GPU Challenge
Downloads last month
38
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Reiwa-AI/PicoPolite240M-base