| ---
|
| license: cc-by-4.0
|
| tags:
|
| - co-attention
|
| - groupnorm
|
| - huge
|
| - mae
|
| - mish
|
| - multi-query
|
| - multitask
|
| - orthogonal
|
| - rmsprop
|
| - step
|
| ---
|
|
|
| # train.py
|
|
|
| ## Model Overview
|
|
|
| A **huge**-scale implementation of the **mae** architecture, built for **multitask** tasks.
|
|
|
| ## Architecture
|
|
|
| - **Architecture**: mae
|
| - **Scale**: huge
|
| - **Attention**: multi query
|
| - **Fusion strategy**: co attention
|
| - **Task head**: multitask
|
| - **Activation**: mish
|
| - **Normalization**: groupnorm
|
| - **Initialization**: orthogonal
|
|
|
| ## Training
|
|
|
| - **Optimizer**: rmsprop
|
| - **LR scheduler**: step
|
|
|
| ## Files
|
|
|
| - `train.py` — main artifact of this repository
|
|
|
| ## License
|
|
|
| See the license field above.
|
|
|