| ---
|
| license: mit
|
| tags:
|
| - adafactor
|
| - constant-warmup
|
| - cross-attention
|
| - groupnorm
|
| - linear
|
| - matching
|
| - swish
|
| - tiny-transformer
|
| - trunc-normal
|
| - xlarge
|
| ---
|
|
|
| # model_321092553_tiny_transformer_xlarge.py
|
|
|
| ## Model Overview
|
|
|
| A **xlarge**-scale implementation of the **tiny transformer** architecture, built for **matching** tasks.
|
|
|
| ## Architecture
|
|
|
| - **Architecture**: tiny transformer
|
| - **Scale**: xlarge
|
| - **Attention**: linear
|
| - **Fusion strategy**: cross attention
|
| - **Task head**: matching
|
| - **Activation**: swish
|
| - **Normalization**: groupnorm
|
| - **Initialization**: trunc normal
|
|
|
| ## Training
|
|
|
| - **Optimizer**: adafactor
|
| - **LR scheduler**: constant warmup
|
|
|
| ## Files
|
|
|
| - `model_321092553_tiny_transformer_xlarge.py` — main artifact of this repository
|
|
|
| ## License
|
|
|
| See the license field above.
|
|
|