--- license: cc-by-4.0 tags: - co-attention - groupnorm - huge - mae - mish - multi-query - multitask - orthogonal - rmsprop - step --- # train.py ## Model Overview A **huge**-scale implementation of the **mae** architecture, built for **multitask** tasks. ## Architecture - **Architecture**: mae - **Scale**: huge - **Attention**: multi query - **Fusion strategy**: co attention - **Task head**: multitask - **Activation**: mish - **Normalization**: groupnorm - **Initialization**: orthogonal ## Training - **Optimizer**: rmsprop - **LR scheduler**: step ## Files - `train.py` — main artifact of this repository ## License See the license field above.