model_633190436_mae_nano.py
Model Overview
A nano-scale implementation of the mae architecture, built for generation tasks.
Architecture
- Architecture: mae
- Scale: nano
- Attention: sparse
- Fusion strategy: tucker
- Task head: generation
- Activation: gelu
- Normalization: batchnorm
- Initialization: orthogonal
Training
- Optimizer: adamw
- LR scheduler: polynomial
Files
model_633190436_mae_nano.py— main artifact of this repository
License
See the license field above.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support