KHWR / MHSA-Word-Model /README.md
Karez's picture
Update MHSA-Word-Model/README.md
ba768de verified
|
Raw
History Blame Contribute Delete
821 Bytes

Multi-Head Self-Attention (Vaswani 2017)

CRNN with Transformer-style Multi-Head Self-Attention between BiLSTM layers 2 and 3.

Architecture

  • Family: MHSA
  • CNN backbone: 6 convolutional blocks (max 256 channels)
  • Recurrent block: 3 BiLSTM layers, hidden = 160 per direction
  • Decoding: Connectionist Temporal Classification (CTC)
  • Parameters: 4,044,625

Test Performance (DASTNUS, seed 42)

Metric Value
Test CER 0.0383
Test WER 0.1655

Files

  • model.safetensors — model weights
  • config.json — architecture and training configuration
  • vocab.json — character-to-index mapping (CTC blank at index 0)
  • idx_to_char.json — reverse mapping for decoding

License

Released for non-commercial scientific research purposes only.