Instructions to use DanVP/vp2vi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DanVP/vp2vi with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="DanVP/vp2vi")# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("DanVP/vp2vi") model = AutoModelForSeq2SeqLM.from_pretrained("DanVP/vp2vi", device_map="auto") - Notebooks
- Google Colab
- Kaggle
license: cc-by-4.0
language:
- vi
base_model: DanVP/MoxhiMT-30-QT
library_name: transformers
pipeline_tag: translation
tags:
- marian
- vietnamese
- text-rewriting
- post-editing
- web-novel
vp2vi — VietPhrase/Convert to Natural Vietnamese
vp2vi is a Vietnamese-to-Vietnamese text-rewriting model for polishing VietPhrase/QT-style Chinese web-novel drafts into more natural Vietnamese. It is a 36.5M-trainable-parameter Marian model with 8 encoder layers, 2 decoder layers, a 448-dimensional hidden size, and a shared 24,000-token vocabulary.
The model continues from DanVP/MoxhiMT-30-QT, retaining its full tokenizer, vocabulary, encoder, and two-layer decoder initialization.
Quick start
import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model_id = "DanVP/vp2vi"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
draft = "Hắn đối với chuyện này cũng không có biện pháp nào."
inputs = tokenizer(draft, return_tensors="pt", add_special_tokens=False)
with torch.inference_mode():
output_ids = model.generate(
**inputs,
do_sample=False,
num_beams=1,
max_new_tokens=224,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
This is the standard Transformers compatibility path. Plain generate() can
still mishandle digits, repetition, or rare names; the
hosted browser application uses
additional decoding constraints. Its runtime is distributed separately and is
not part of this model repository.
Intended use
- Sentence- or line-level polishing of Vietnamese VietPhrase/QT drafts from xianxia, xuanhuan, and related web-novel genres.
- Local reading tools, post-editing experiments, and further fine-tuning.
This is not a Chinese-to-Vietnamese translator: the input is already a Vietnamese machine-converted draft. It is not intended for factual, legal, medical, or other high-stakes text.
Training data and procedure
The model was initialized from DanVP/MoxhiMT-30-QT and continued on
owner-curated Vietnamese draft→edited-text pairs. Training categories included
a filtered QT-convert pool (about 749k unique pairs), teacher-silver data
(about 49.6k pairs), a small teacher-gold seed (899 pairs), identity examples,
and about 150k synthetic name-swap pairs. Cumulative exposure reached
1,237,021,456 target labels. Raw training and holdout corpora are not
redistributed in this repository.
The published model.safetensors is a deterministic reverse mapping of the
training checkpoint into standard MarianMTModel tensor names. The release
gate verifies complete tensor coverage, exact expected key layout, successful
AutoTokenizer/AutoModelForSeq2SeqLM loading, and logits/argmax parity against
the training implementation. Exact hashes are recorded in
release_manifest.json.
Limitations
- Training and validation focus on VietPhrase/QT drafts from Chinese web novels. News, technical, legal, medical, academic, business, conversational, and other genres are out of domain. The model may rewrite already-correct prose, corrupt specialist terminology, or alter the meaning; avoid using it as a general-purpose Vietnamese editor.
- Rare or out-of-distribution name inventories can cause large quality drops.
- Bare decoding may invent a digit, repeat text, omit content, or damage a proper name. The guarded application path reduces these failures but does not eliminate semantic errors.
- The checker still has a known blind spot for name-boundary resegmentation: individually attested words can be recombined into the wrong name span.
- Earlier checker revisions also conflated book-level name recap with invention and could flag legitimate repeated proper nouns; those two cases were corrected in checker v2.
- The model works one line at a time, so cross-sentence consistency depends on the input draft and surrounding application.
License and attribution
The model repository is released under
CC BY 4.0. Commercial use,
modification, and redistribution are permitted with attribution. Suggested
credit: “vp2vi by DanVP”, linking to
https://huggingface.co/DanVP/vp2vi.
The project owner authorized this public release on 2026-08-11. The model repository contains model artifacts and documentation only; the separately deployed browser runtime/Space has its own release boundary.