SpliNet

SpliNet: A Zero-Parameter B-Spline Transformer with Linear Complexity

SpliNet replaces learned self-attention token mixing with a fixed order-2 single-sided cardinal B-spline operator.

This model was pretrained from scratch on exactly 2,000,000,000 C4 tokens using the dedicated SpliNet tokenizer.

Architecture

  • Layers: 12
  • Hidden size: 768
  • Heads: 12
  • FFN width: 3072
  • Sequence length: 512
  • Vocabulary: 32000
  • Spline order: 2
  • Spline radius: 16
  • Trainable mixer parameters: 0
  • Total parameters: 82,894,592
  • Trainable parameters: 82,894,592

Pretraining

  • Training tokens: 2,000,000,000
  • Validation tokens: 5,120,000
  • Objective: masked language modeling
  • Masked positions: 77/512
  • Optimizer: AdamW
  • Precision: BF16
  • Hardware: NVIDIA A100-SXM4-80GB

Final validation

  • MLM loss: 4.614502
  • MLM perplexity: 100.937575

Load with trust_remote_code=True.

OpenReview: https://openreview.net/forum?id=nWHnuiEF3C GitHub: https://github.com/AngshulMajumdar/SpliNet

Downloads last month
17
Safetensors
Model size
82.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Angshul/SpliNet