File size: 4,042 Bytes
9ab41bf cf31005 db9e598 9ab41bf cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 cf31005 db9e598 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 | ---
license: cc-by-nc-4.0
language:
- en
- vi
tags:
- translation
- machine-translation
- english-to-vietnamese
- onnx
- seq2seq
- bart
pipeline_tag: translation
widget:
- text: "Machine translation is useful for multilingual applications."
example_title: "EN → VI"
---
# AnViMt
**AnViMt** — **An**=English **Vi**=Vietnamese **Mt**=Machine Translation — dịch **EN → VI** cho văn bản thường + kỹ thuật/hệ thống/lập trình/hội thoại.
Export sang **ONNX** để chạy chỉ với `onnxruntime + sentencepiece + numpy` (không cần `torch/transformers`).
## Architecture
**Base:** `transformers.BartForConditionalGeneration` train từ đầu, không phải pretrained BART.
- `vocab_size=32000` (SentencePiece BPE, `byte_fallback=True`)
- `d_model=512` | `encoder_layers=6` | `decoder_layers=6`
- `encoder_ffn_dim=2048` | `decoder_ffn_dim=2048`
- `encoder_attention_heads=8` | `decoder_attention_heads=8`
- `max_position_embeddings=256` | `dropout 0.1`
- `tie_word_embeddings=True` | `is_encoder_decoder=True`
- `pad_token_id=0` (`<pad>`) | `unk_token_id=1` (`<unk>`) | `bos_token_id=2` (`<s>`) | `eos_token_id=3` (`</s>`) | `decoder_start_token_id=2`
- **Params:** ~60.79M
```
EN text
↓ spm.model [2] + BPE ids + [3] (pad 0, bos 2, eos 3, max 256)
↓
encoder_model.onnx (input_ids, attention_mask → encoder_hidden_states [batch, seq, 512])
↓
decoder_model.onnx / decoder_with_past_model.onnx
(decoder_input_ids [2] + encoder_hidden_states + encoder_attention_mask + past_key_values → logits)
↓ autoregressive greedy (argmax) đến eos 3
↓ sp.decode → VI text
```
## Files
Export bằng `optimum` `task=seq2seq-lm` `opset=14`:
```
AnViMt/
├── encoder_model.onnx # 136M
├── decoder_model.onnx # 223M
├── decoder_with_past_model.onnx # KV cache (tùy chọn, tăng tốc)
├── spm.model # 752K, vocab 32000
└── spm.vocab
```
Runtime tối thiểu chỉ cần `encoder_model.onnx` + `decoder_model.onnx` + `spm.model`. `decoder_with_past_model.onnx` dùng khi muốn KV-cache.
Không cần: `config.json`, `pytorch_model.bin`, `safetensors`, `torch`, `transformers`.
## Installation
```bash
pip install onnxruntime sentencepiece numpy
# Termux: apt install python3-onnxruntime python3-sentencepiece python3-numpy
```
## Quick Start
```
project/
├── main.py
├── encoder_model.onnx
├── decoder_model.onnx
└── spm.model # (decoder_with_past_model.onnx nếu có)
```
`main.py` thực tế (đã test, xử lý đúng BOS/EOS/pad):
```python
import numpy as np, onnxruntime as ort, sentencepiece as spm
sp = spm.SentencePieceProcessor(); sp.load("spm.model")
enc = ort.InferenceSession("encoder_model.onnx", providers=["CPUExecutionProvider"])
dec = ort.InferenceSession("decoder_model.onnx", providers=["CPUExecutionProvider"])
BOS, EOS, PAD = 2, 3, 0
def translate(text, max_len=256):
ids = [BOS] + sp.encode(text, out_type=int)[:254] + [EOS]
input_ids = np.array([ids], dtype=np.int64)
attn = np.ones_like(input_ids)
enc_hs = enc.run(None, {"input_ids": input_ids, "attention_mask": attn})[0]
gen = [BOS]
for _ in range(max_len):
dec_ids = np.array([gen], dtype=np.int64)
logits = dec.run(None, {"input_ids": dec_ids, "encoder_hidden_states": enc_hs, "encoder_attention_mask": attn})[0]
nxt = int(np.argmax(logits[0, -1]))
gen.append(nxt)
if nxt == EOS: break
out = [x for x in gen[1:] if x not in (BOS, EOS, PAD)]
if EOS in out: out = out[:out.index(EOS)]
return sp.decode(out)
```
## Training Data
Pretrain 1.13M EN-VI (PhoMT + TED + TECH) → finetune 200k TECH → finetune 27k v3 (hội thoại + kỹ thuật). Tokenizer BPE 32k train trên 400k pairs reservoir sampled, `nmt_nfkc`, `split_digits`, `byte_fallback`.
## Limitations
EN → VI only. Câu >254 BPE tokens bị cắt. Thuật ngữ mới có thể giữ nguyên EN.
## License
CC BY-NC 4.0 — non-commercial.
Author: Phitran21
|