ARTN - Adaptive Ternary Recurrent Network

1.58-bit Ternary Quantization for Language Models

Experimental implementation of ternary (1.58-bit) quantization for training small language models on extremely limited hardware.

Status: EXPERIMENTAL - Not Production Ready


Project Overview

This project explores whether ternary quantization (weights constrained to -1, 0, +1) can be used to train small language models on hardware with:

Constraint Value
RAM 4GB
CPU 2-core (no GPU)
Storage ~20GB

Model Architecture

Parameter Value
Model type BitLinearGPT
Parameters 6,856,608 (~6.8M)
Hidden size (d_model) 256
Layers 6
Attention heads 8
Feed-forward size 1024
Max sequence length 256
Vocabulary size 4,000
Quantization Ternary (1.58-bit)

Checkpoints Included

1. checkpoint-step5000-experiment.pt

Attribute Value
Steps 5,000
Final Loss 8.45
Entropy Not monitored
Status SUCCESS (pipeline verified)
Description First test training, proved architecture works
Duration ~4 hours

Purpose: Verified that training pipeline works correctly.


2. checkpoint-step50000-failed-collapsed.pt

Attribute Value
Steps 50,000
Final Loss 3.16
Final Entropy 8.2% (COLLAPSED)
Top Token "time" at 94%
Status FAILED (collapsed)
Description Beta1=None caused model collapse
Duration 79 hours
Root Cause Missing momentum (beta1=None)

Warning: This checkpoint is COLLAPSED. Model outputs repetitive text like "time time time...". Use only as a learning example of what NOT to do.


3. checkpoint-step50000-experiment-gibberish.pt

Attribute Value
Steps 50,000
Final Loss 7.56
Final Entropy 96.2% (HEALTHY)
Top Token "the" at 0.56%
Status HEALTHY but NOT USEFUL
Description Beta1=0.9 prevented collapse but learning too slow
Duration 83 hours
Issue Beta1=0.9 too conservative

Note: This checkpoint is HEALTHY (entropy normal) but produces gibberish. Model learned slowly due to overly conservative optimizer settings.


Training Configuration

Setting Value
Dataset TinyStories (roneneldan/TinyStories)
Samples 50,000
Batch size 4
Gradient accumulation 8
Effective batch size 32
Optimizer Adafactor
Max Learning Rate 0.0003
Min Learning Rate 0.00003
Warmup steps 1,000
Weight decay 0.01
Gradient clipping 1.0

Key Findings

Beta1 Trade-off

Beta1 Value Learning Speed Stability Result
None (0.0) FAST UNSTABLE COLLAPSED at 50k
0.9 SLOW STABLE HEALTHY but not useful
0.5 - 0.7 UNKNOWN UNKNOWN Need to test

Critical Lesson: Beta1=None causes collapse. Beta1=0.9 too conservative. Need to find optimal value.


Training Metrics Comparison

Training Loss @ 50k Entropy Top Token Status
Attempt 2 (beta1=None) 3.16 8.2% 94% "time" COLLAPSED
Attempt 4 (beta1=0.9) 7.56 96.2% 0.56% "the" HEALTHY (gibberish)

Paradox: Lower loss does NOT mean better model. Attempt 2 had lower loss but was collapsed.


Total Time Invested

Attempt Duration Outcome
1 (5k test) 4 hours SUCCESS (pipeline verified)
2 (50k collapsed) 79 hours FAILED (beta1=None)
3 (stopped) 1 hour STOPPED (wrong accumulation)
4 (50k gibberish) 83 hours COMPLETED (not useful)
Total 167 hours

Lessons Learned

  1. Entropy monitoring is CRITICAL - Loss alone doesn't indicate model health
  2. Beta1=None causes collapse - Model falls into repetitive patterns
  3. Beta1=0.9 is too conservative - Learning becomes extremely slow
  4. Lower loss ≠ better model - Collapsed model had lower loss (3.16 vs 7.56)
  5. Need optimal beta1 - Testing 0.5 and 0.7 as candidates
  6. Ternary quantization is challenging - Requires careful optimizer tuning

How to Use These Checkpoints

For Research/Learning

import torch
from src.model.bitlinear import BitLinearGPT
from src.tokenizer.bpe_tokenizer import BPETokenizer

# Load model architecture
model = BitLinearGPT(
    vocab_size=4000,
    d_model=256,
    n_heads=8,
    n_layers=6,
    d_ff=1024,
    max_seq_len=256
)

# Load checkpoint (choose one)
checkpoint = torch.load('checkpoint-step50000-experiment-gibberish.pt', map_location='cpu')
model.load_state_dict(checkpoint)
model.eval()

# Load tokenizer
tokenizer = BPETokenizer.load('tokenizer.json')

# Generate text
input_ids = torch.tensor([tokenizer.encode("Once upon a time")])
# ... generation code ...

Not For Production

These checkpoints are experimental. They will NOT produce coherent text. Use them to:

  • Learn about ternary quantization challenges
  • Understand model collapse patterns
  • Compare optimizer settings
  • Train your own models with different configs

Tokenizers

File Description
tokenizer.json Main tokenizer (BPE, vocab size 4000)
tokenizer-backup-50k.json Backup from collapsed training

Source Code

Full source code available at GitHub:

https://github.com/kishxrx/ARTN

Includes:

  • Complete training scripts
  • Model architecture
  • Training history documentation
  • Monitoring tools

Training History

Complete documentation of all experiments:

https://github.com/kishxrx/ARTN/blob/main/TRAINING_HISTORY.md


Next Steps

Phase Plan
1 Test beta1=0.5 (10k steps)
2 Test beta1=0.7 (10k steps)
3 Compare entropy, loss, samples
4 Select optimal beta1
5 Full 50k training with best config
6 Release v1.0 if successful

Citation

If you use this work for learning or research:

@misc{artn-ternary-2026,
  author = {kishxrx},
  title = {ARTN: Adaptive Ternary Recurrent Network - 1.58-bit Quantization Experiments},
  year = {2026},
  publisher = {HuggingFace},
  url = {https://huggingface.co/kishxrx/artn-ternary-experimental}
}

License

MIT License


Disclaimer

THIS IS EXPERIMENTAL SOFTWARE.

  • Models do NOT produce coherent text
  • Checkpoints are for learning purposes only
  • Not suitable for any production use
  • No guarantees of any kind

Contact


Last Updated: August 2026

Version: v0.1-experimental

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support