| # Multi-Head Self-Attention (Vaswani 2017) | |
| CRNN with Transformer-style Multi-Head Self-Attention between BiLSTM layers 2 and 3. | |
| ## Architecture | |
| - Family: **MHSA** | |
| - CNN backbone: 6 convolutional blocks (max 256 channels) | |
| - Recurrent block: 3 BiLSTM layers, hidden = 160 per direction | |
| - Decoding: Connectionist Temporal Classification (CTC) | |
| - Parameters: 4,044,625 | |
| ## Test Performance (DASTNUS, seed 42) | |
| | Metric | Value | | |
| |--------|-------| | |
| | Test CER | 0.0383 | | |
| | Test WER | 0.1655 | | |
| ## Files | |
| - `model.safetensors` — model weights | |
| - `config.json` — architecture and training configuration | |
| - `vocab.json` — character-to-index mapping (CTC blank at index 0) | |
| - `idx_to_char.json` — reverse mapping for decoding | |
| ## License | |
| Released for non-commercial scientific research purposes only. | |