clanker-creator's picture
Update README.md
b5a5657 verified
|
Raw
History Blame Contribute Delete
6.15 kB
---
license: other
license_name: april
license_link: LICENSE
language:
- en
tags:
- gated
- sparse
- language-model
- cpu
- efficient
- recurrent
- acrs
- gsn
pipeline_tag: text-generation
library_name: custom
---
# GSN β€” Gated Sparse Network
**Acid Research (ACRS)**
GSN is a language model architecture designed around one principle: **do not spend compute you do not need.**
Unlike transformers β€” which run the full model on every input regardless of complexity β€” GSN gates compute dynamically. Simple inputs take a shallow path. Complex inputs go deeper. The network decides, not the configuration.
---
## Architecture
GSN is built from three core ideas stacked together:
**1. Gated Depth**
A lightweight gate network evaluates each input and decides how many layers to activate. Layers that are not needed do not run. Their compute cost is exactly zero.
**2. Sparse Activation**
Within each active layer, only the top-k neurons fire. The rest are masked to zero. A 4-layer GSN running at 25% sparsity uses a fraction of the compute a dense model would.
**3. Recurrent Encoder**
Input tokens are processed sequentially through a GRU encoder before the gate sees anything. This replaces mean pooling β€” token order matters, context accumulates, and the gate receives a hidden state that actually encodes sequence structure.
**4. Compute Penalty in Training**
The loss function penalizes wasted compute. The model is trained to be cheap, not just accurate. Over time the gate learns: if I can answer correctly with one layer, using two is a mistake.
**No attention. No transformers. O(T) encoding. O(1) per sparse layer.**
---
## Why GSN
| Property | Transformer | RSM | GSN |
|----------|-------------|-----|-----|
| Attention cost | O(nΒ²) | none | none |
| Compute per input | fixed | fixed | dynamic |
| Sparse activation | no | no | yes |
| CPU viable | marginal | yes | yes |
| Trains on 2 cores | no | hours | minutes |
GSN was developed and trained entirely on a consumer CPU with 2 physical cores. No GPU. No cloud compute. That is not a limitation β€” it is the point.
---
## Usage
### Requirements
```bash
pip install numpy
```
No PyTorch. No CUDA. No framework dependency. Pure NumPy.
### Training
```bash
python train.py your_corpus.txt 20000
```
Point it at any plain text file. The tokenizer trains from scratch on your corpus. Checkpoints save every 500 steps. Resume is automatic.
### Inference
```bash
python infer.py "your prompt here" --max_new 100
```
The inference report shows complexity score, gate depth decision, and compute saved per generation.
### Configuration
All hyperparameters live in `config.py`. Key settings:
```python
GSNConfig(
vocab_size = 1024, # BPE vocabulary size
dim = 256, # embedding and hidden dimension
enc_dim = 256, # GRU encoder hidden dimension
n_layers = 4, # total sparse layers available
compute_penalty = 0.001, # Ξ» β€” weight of compute cost in loss
lr = 3e-4, # learning rate
seq_len = 128, # context length
)
```
---
## Training Details
**Corpus:** Shakespeare complete works (~1.1M characters)
**Vocabulary:** 1024 BPE tokens
**Steps:** 20,000
**Hardware:** 2 physical CPU cores (Debian Linux)
**Training time:** ~25 minutes
**Final loss:** ~3.5
**Average gate depth:** 1.44 / 4 layers
**Average compute saved:** ~88%
The gate learned that Shakespeare β€” structured, repetitive, consistent vocabulary β€” is mostly shallow complexity. On a more diverse corpus the gate is expected to show greater depth variation.
---
## Repository Structure
```
config.py β€” all hyperparameters
tokenizer.py β€” BPE tokenizer, trains from scratch
model.py β€” GSN model, forward pass, analytical backprop
train.py β€” training loop, Adam optimizer, checkpointing
infer.py β€” inference CLI with compute report
```
---
## Limitations
- **Undertrained on small corpus.** 20k steps on Shakespeare is a proof of concept. Coherent generation requires significantly more training on a larger and more diverse corpus.
- **Small vocabulary.** 1024 tokens is minimal. Real deployments should use 8k-32k.
- **No pretrained weights included.** This release is an architecture and training framework, not a ready-to-use model. Train your own.
- **Gate behavior is corpus-dependent.** The gate learns complexity relative to the training distribution. A model trained on Shakespeare will gate differently than one trained on code or web text.
- **Single-sequence inference only.** Batched inference is not yet implemented.
---
## Known Issues and Community Contributions Welcome
- Batched inference
- Larger vocabulary and longer context experiments
- Perplexity benchmarking against RSM and small transformers
- Gate visualization tooling
- Training on FineWeb, OpenWebText, or code corpora
This is an incomplete release by design. The architecture is sound. The community is invited to take it further.
---
## License
Licensed under the **Acid Research Protected Interests License (APRIL) v1.0**.
- Free for personal, academic, and non-commercial use.
- Derivatives must be open sourced under APRIL.
- Commercial use requires written permission from ACRS.
- Attribution to Acid Research (ACRS) is required in all derivatives.
See [LICENSE](LICENSE) for full terms.
---
## Citation
If you use GSN in research or build on this architecture, please cite:
```
@misc{gsn2025,
title = {GSN: Gated Sparse Network},
author = {Acid Research (ACRS)},
year = {2025},
url = {https://huggingface.co/AcidAI/Acid-GSN-Architecture}
}
```
---
## About Acid Research
Acid Research (ACRS) is an independent AI research organization building CPU-native, economically viable alternatives to transformer-based architectures.
Current architecture portfolio:
- **HAM** β€” Hebbian Architecture Model
- **RSM** β€” Recurrent State Machine
- **RDM** β€” Recurrent Depth Machine
- **IMA** β€” Intent Machine Architecture
- **GSN** β€” Gated Sparse Network
*"Make it linear, or else your cost ain't going to be."*