GSN-Open-Weights / README.md
clanker-creator's picture
Update README.md
51d4d02 verified
|
Raw
History Blame Contribute Delete
4.87 kB
---
license: other
license_name: april
license_link: LICENSE.md
language:
- en
tags:
- language-model
- cpu-native
- sparse
- gated
- numpy
- no-gpu
- acid-research
- gsn
library_name: custom
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
---
# GSN β€” Gated Sparse Network (Open Weights)
**Acid Research (ACRS)**
This repository contains trained weights for GSN, a language model architecture designed around one principle: **do not spend compute you do not need.**
Unlike transformers β€” which run the full model on every input regardless of complexity β€” GSN gates compute dynamically. Simple inputs take a shallow path. Complex inputs go deeper. The network decides, not the configuration.
For the full architecture, training code, and inference CLI, see [AcidAI/Acid-GSN-Architecture](https://huggingface.co/AcidAI/Acid-GSN-Architecture).
---
## Architecture
GSN is built from three core ideas stacked together:
**1. Gated Depth** A lightweight gate network evaluates each input and decides how many layers to activate. Layers that are not needed do not run. Their compute cost is exactly zero.
**2. Sparse Activation** Within each active layer, only the top-k neurons fire. The rest are masked to zero. A 4-layer GSN running at 25% sparsity uses a fraction of the compute a dense model would.
**3. Recurrent Encoder** Input tokens are processed sequentially through a GRU encoder before the gate sees anything. This replaces mean pooling β€” token order matters, context accumulates, and the gate receives a hidden state that actually encodes sequence structure.
**4. Compute Penalty in Training** The loss function penalizes wasted compute. The model is trained to be cheap, not just accurate. Over time the gate learns: if I can answer correctly with one layer, using two is a mistake.
**No attention. No transformers. O(T) encoding. O(1) per sparse layer.**
---
## Files
| File | Description |
| --- | --- |
| `gsn_fw_weights.npz` | Model weights (~19.7 MB) |
| `gsn_fw_opt.npz` | Adam optimizer state β€” m1/m2 moments and global step for resuming training (~39.4 MB) |
| `LICENSE.md` | APRIL license |
---
## Usage
### Requirements
```bash
pip install numpy
```
No PyTorch. No CUDA. No framework dependency. Pure NumPy.
### Loading the Weights
```python
import numpy as np
# Load weights
weights = np.load("gsn_fw_weights.npz")
print(list(weights.keys())) # inspect layer names
# Load optimizer state (for resuming training)
opt_state = np.load("gsn_fw_opt.npz")
```
For training and inference code, clone the architecture repository:
```bash
# Clone the architecture repo alongside these weights
# python infer.py "your prompt here" --max_new 100
```
See [AcidAI/Acid-GSN-Architecture](https://huggingface.co/AcidAI/Acid-GSN-Architecture) for the full usage guide.
---
## Training Details
**Corpus:** FineWeb-Edu (high-quality educational web text)
**Vocabulary:** 1024 BPE tokens
**Hardware:** 2 physical CPU cores (Debian Linux)
**Training time:** ~25 minutes
**Best loss:** ~3.72
**Average gate depth:** ~2 / 4 layers
The gate settled at adaptive depth 2, confirming the compute penalty is functioning as intended. Depth variation is expected to increase on more diverse corpora.
---
## Limitations
- **Undertrained checkpoint.** These weights are a research-stage snapshot, not a production model. Coherent generation requires significantly more training on a larger corpus.
- **Small vocabulary.** 1024 tokens is minimal. Real deployments should use 8k–32k.
- **Gate behavior is corpus-dependent.** A model trained on FineWeb-Edu will gate differently than one trained on code or conversational text.
- **Single-sequence inference only.** Batched inference is not yet implemented.
---
## License
Licensed under the **Acid Research Protected Interests License (APRIL) v1.0**.
- Free for personal, academic, and non-commercial use.
- Derivatives must be open sourced under APRIL.
- Commercial use requires written permission from ACRS.
- Attribution to Acid Research (ACRS) is required in all derivatives.
See [LICENSE.md](LICENSE.md) for full terms.
---
## Citation
If you use GSN in research or build on this architecture, please cite:
```
@misc{gsn2025,
title = {GSN: Gated Sparse Network},
author = {Acid Research (ACRS)},
year = {2025},
url = {https://huggingface.co/AcidAI/GSN-Open-Weights}
}
```
---
## About Acid Research
Acid Research (ACRS) is an independent AI research organization building CPU-native, economically viable alternatives to transformer-based architectures.
Current architecture portfolio:
- **HAM** β€” Hebbian Architecture Model
- **RSM** β€” Recurrent State Machine
- **RDM** β€” Recurrent Depth Machine
- **IMA** β€” Intent Machine Architecture
- **GSN** β€” Gated Sparse Network
*"Make it linear, or else your cost ain't going to be."*