inference.py

Model Overview

A small-scale implementation of the mixer architecture, built for multitask tasks.

Architecture

  • Architecture: mixer
  • Scale: small
  • Attention: dilated
  • Fusion strategy: low rank
  • Task head: multitask
  • Activation: swish
  • Normalization: rmsnorm
  • Initialization: xavier uniform

Training

  • Optimizer: adafactor
  • LR scheduler: step

Files

  • inference.py — main artifact of this repository

License

See the license field above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support