File size: 4,042 Bytes
9ab41bf
cf31005
 
 
 
 
 
 
 
 
 
 
 
db9e598
 
 
9ab41bf
cf31005
 
 
db9e598
cf31005
db9e598
cf31005
db9e598
cf31005
db9e598
cf31005
db9e598
 
 
 
 
 
 
 
cf31005
db9e598
 
 
 
 
 
 
 
 
 
 
cf31005
db9e598
cf31005
db9e598
cf31005
db9e598
cf31005
db9e598
 
 
 
 
 
cf31005
db9e598
cf31005
db9e598
cf31005
db9e598
cf31005
db9e598
 
 
 
cf31005
db9e598
cf31005
db9e598
cf31005
 
 
 
db9e598
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
---
license: cc-by-nc-4.0
language:
- en
- vi
tags:
- translation
- machine-translation
- english-to-vietnamese
- onnx
- seq2seq
- bart
pipeline_tag: translation
widget:
- text: "Machine translation is useful for multilingual applications."
  example_title: "EN → VI"
---

# AnViMt

**AnViMt****An**=English **Vi**=Vietnamese **Mt**=Machine Translation — dịch **EN → VI** cho văn bản thường + kỹ thuật/hệ thống/lập trình/hội thoại.

Export sang **ONNX** để chạy chỉ với `onnxruntime + sentencepiece + numpy` (không cần `torch/transformers`).

## Architecture

**Base:** `transformers.BartForConditionalGeneration` train từ đầu, không phải pretrained BART.

- `vocab_size=32000` (SentencePiece BPE, `byte_fallback=True`)
- `d_model=512` | `encoder_layers=6` | `decoder_layers=6`
- `encoder_ffn_dim=2048` | `decoder_ffn_dim=2048`
- `encoder_attention_heads=8` | `decoder_attention_heads=8`
- `max_position_embeddings=256` | `dropout 0.1`
- `tie_word_embeddings=True` | `is_encoder_decoder=True`
- `pad_token_id=0` (`<pad>`) | `unk_token_id=1` (`<unk>`) | `bos_token_id=2` (`<s>`) | `eos_token_id=3` (`</s>`) | `decoder_start_token_id=2`
- **Params:** ~60.79M

```
EN text
  ↓ spm.model [2] + BPE ids + [3] (pad 0, bos 2, eos 3, max 256)

encoder_model.onnx (input_ids, attention_mask → encoder_hidden_states [batch, seq, 512])

decoder_model.onnx / decoder_with_past_model.onnx
  (decoder_input_ids [2] + encoder_hidden_states + encoder_attention_mask + past_key_values → logits)
  ↓ autoregressive greedy (argmax) đến eos 3
  ↓ sp.decode → VI text
```

## Files

Export bằng `optimum` `task=seq2seq-lm` `opset=14`:

```
AnViMt/
├── encoder_model.onnx          # 136M
├── decoder_model.onnx          # 223M
├── decoder_with_past_model.onnx # KV cache (tùy chọn, tăng tốc)
├── spm.model                   # 752K, vocab 32000
└── spm.vocab
```

Runtime tối thiểu chỉ cần `encoder_model.onnx` + `decoder_model.onnx` + `spm.model`. `decoder_with_past_model.onnx` dùng khi muốn KV-cache.

Không cần: `config.json`, `pytorch_model.bin`, `safetensors`, `torch`, `transformers`.

## Installation

```bash
pip install onnxruntime sentencepiece numpy
# Termux: apt install python3-onnxruntime python3-sentencepiece python3-numpy
```

## Quick Start

```
project/
├── main.py
├── encoder_model.onnx
├── decoder_model.onnx
└── spm.model   # (decoder_with_past_model.onnx nếu có)
```

`main.py` thực tế (đã test, xử lý đúng BOS/EOS/pad):

```python
import numpy as np, onnxruntime as ort, sentencepiece as spm
sp = spm.SentencePieceProcessor(); sp.load("spm.model")
enc = ort.InferenceSession("encoder_model.onnx", providers=["CPUExecutionProvider"])
dec = ort.InferenceSession("decoder_model.onnx", providers=["CPUExecutionProvider"])
BOS, EOS, PAD = 2, 3, 0
def translate(text, max_len=256):
    ids = [BOS] + sp.encode(text, out_type=int)[:254] + [EOS]
    input_ids = np.array([ids], dtype=np.int64)
    attn = np.ones_like(input_ids)
    enc_hs = enc.run(None, {"input_ids": input_ids, "attention_mask": attn})[0]
    gen = [BOS]
    for _ in range(max_len):
        dec_ids = np.array([gen], dtype=np.int64)
        logits = dec.run(None, {"input_ids": dec_ids, "encoder_hidden_states": enc_hs, "encoder_attention_mask": attn})[0]
        nxt = int(np.argmax(logits[0, -1]))
        gen.append(nxt)
        if nxt == EOS: break
    out = [x for x in gen[1:] if x not in (BOS, EOS, PAD)]
    if EOS in out: out = out[:out.index(EOS)]
    return sp.decode(out)
```

## Training Data

Pretrain 1.13M EN-VI (PhoMT + TED + TECH) → finetune 200k TECH → finetune 27k v3 (hội thoại + kỹ thuật). Tokenizer BPE 32k train trên 400k pairs reservoir sampled, `nmt_nfkc`, `split_digits`, `byte_fallback`.

## Limitations

EN → VI only. Câu >254 BPE tokens bị cắt. Thuật ngữ mới có thể giữ nguyên EN.

## License

CC BY-NC 4.0 — non-commercial.
Author: Phitran21