metadata
license: cc-by-4.0
tags:
- co-attention
- groupnorm
- huge
- mae
- mish
- multi-query
- multitask
- orthogonal
- rmsprop
- step
train.py
Model Overview
A huge-scale implementation of the mae architecture, built for multitask tasks.
Architecture
- Architecture: mae
- Scale: huge
- Attention: multi query
- Fusion strategy: co attention
- Task head: multitask
- Activation: mish
- Normalization: groupnorm
- Initialization: orthogonal
Training
- Optimizer: rmsprop
- LR scheduler: step
Files
train.py— main artifact of this repository
License
See the license field above.