KHWR / MHSA-Word-Model /README.md
Karez's picture
Update MHSA-Word-Model/README.md
ba768de verified
|
Raw
History Blame Contribute Delete
821 Bytes
# Multi-Head Self-Attention (Vaswani 2017)
CRNN with Transformer-style Multi-Head Self-Attention between BiLSTM layers 2 and 3.
## Architecture
- Family: **MHSA**
- CNN backbone: 6 convolutional blocks (max 256 channels)
- Recurrent block: 3 BiLSTM layers, hidden = 160 per direction
- Decoding: Connectionist Temporal Classification (CTC)
- Parameters: 4,044,625
## Test Performance (DASTNUS, seed 42)
| Metric | Value |
|--------|-------|
| Test CER | 0.0383 |
| Test WER | 0.1655 |
## Files
- `model.safetensors` — model weights
- `config.json` — architecture and training configuration
- `vocab.json` — character-to-index mapping (CTC blank at index 0)
- `idx_to_char.json` — reverse mapping for decoding
## License
Released for non-commercial scientific research purposes only.