Instructions to use febilly/Hy-MT2-1.8B-StreamRevise-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use febilly/Hy-MT2-1.8B-StreamRevise-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
Use Docker
docker model run hf.co/febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use febilly/Hy-MT2-1.8B-StreamRevise-GGUF with Ollama:
ollama run hf.co/febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
- Unsloth Studio
How to use febilly/Hy-MT2-1.8B-StreamRevise-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for febilly/Hy-MT2-1.8B-StreamRevise-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for febilly/Hy-MT2-1.8B-StreamRevise-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for febilly/Hy-MT2-1.8B-StreamRevise-GGUF to start chatting
- Docker Model Runner
How to use febilly/Hy-MT2-1.8B-StreamRevise-GGUF with Docker Model Runner:
docker model run hf.co/febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
- Lemonade
How to use febilly/Hy-MT2-1.8B-StreamRevise-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull febilly/Hy-MT2-1.8B-StreamRevise-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Hy-MT2-1.8B-StreamRevise-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Hy-MT2-1.8B-StreamRevise — GGUF
4-bit llama.cpp build of StreamRevise, a translation model for live subtitles. Every time the ASR hypothesis updates, you pass the model its own previous translation of the sentence in progress; it decides whether to keep that text and extend it, or rewrite part of it because the meaning changed. Subtitles grow smoothly instead of flickering.
1.07 GB on disk, and about 1.4 GB resident at n_ctx 2048. Small enough to ship with a desktop app.
See Footprint for how that scales.
Adapter, full prompt spec, training details:
Hy-MT2-1.8B-StreamRevise-LoRA ·
Base model: tencent/Hy-MT2-1.8B
中文简介:实时字幕翻译模型的 4-bit 本地部署版。语音边说边识别,源文一直在变——每次更新时把模型上一版 译文一起喂回去,它自己决定保留已显示的部分继续接,还是因为意思变了而改写,让字幕平滑增长不闪烁。 1.07 GB,纯 CPU 可跑(约 2.1 GB 内存),有显卡则单次请求 <0.1 秒。
Files
| file | size | what it is |
|---|---|---|
Hy-MT2-1.8B-StreamRevise-Q4_K_M.gguf |
1.07 GB | the model. Q4_K_M with imatrix calibration, token embeddings at Q4_K |
Hy-MT2-1.8B-StreamRevise.imatrix.gguf |
2.3 MB | importance matrix, only needed if you want to re-quantize at another bit width |
Only one quantization is published. The alternatives were measured and the differences sit inside the noise floor, so the smallest one shipped:
| variant | size | verdict |
|---|---|---|
Q4_K_M, embeddings Q4_K |
1.07 GB | published here |
Q4_K_M, embeddings Q8_0 |
1.19 GB | +0.5pp exact match — noise |
Q4_K_M, embeddings F16 |
1.42 GB | no measurable gain |
| Q4_K_M, defaults | 1.92 GB | llama-quantize promotes embeddings to F32. No quality gain, 470 MB wasted |
Embeddings are tied to lm_head in this architecture, which is why their bit width gets its own column.
imatrix calibration used real streaming-translation prompts, which is what makes 4-bit embeddings safe here.
Run it
llama-server -m Hy-MT2-1.8B-StreamRevise-Q4_K_M.gguf -c 4096 --port 8080 -ngl 99 --no-webui
Drop -ngl 99 (or set 0) on CPU-only machines. Then POST to /completion:
{
"prompt": "<|hy_begin▁of▁sentence|><|hy_User|>{PROMPT}<|hy_Assistant|>",
"n_predict": 256,
"temperature": 0,
"cache_prompt": true
}
Three things to get right:
temperature: 0— greedy, not the base model's0.7. Consecutive updates hand the model nearly the same prompt and you want nearly the same output back; sampling introduces changes that have nothing to do with the source changing, and those land on screen as flicker. All numbers below are greedy.cache_prompt: true— updates within one utterance share a long prefix. KV reuse is what makes per-update latency negligible.{PROMPT}must use the StreamRevise layout, not a plain translation instruction.
A warning you can ignore
llama.cpp prints this on load:
load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect
It is harmless here. Hy-MT2 declares a single-value eos_token_id, and this build carries an explicit
tokenizer.ggml.eom_token_id = 120020 that llama.cpp does not fold into its special_eog_ids set — hence
the complaint. Generation still terminates correctly: expect stop_type: eos and a predicted_n far below
your n_predict. If instead the model never stops, you are running a GGUF converted without the
end-of-message fix — see below.
Prompt format
Full spec and a copy-paste renderer are in the LoRA repo. Short version.
First chunk of a new sentence, nothing to revise yet:
Translate the following text into {TARGET_LANGUAGE}. Note that you should only output the translated result without any additional explanation:
{CURRENT_SOURCE}
Every update after that:
[Background Information]
Recent source utterances:
{UP_TO_10_PREVIOUS_SOURCE_SENTENCES}
Previous version of the current source:
{PREVIOUS_SOURCE}
Previous translation of the current source:
{PREVIOUS_TRANSLATION}
When the source meaning has not changed, preserve the still-correct prefix of the previous translation whenever possible. When content is added or corrected, accuracy and completeness take priority.
Please translate the following text into {TARGET_LANGUAGE}, taking the provided background information into consideration.
[Source Text]
{CURRENT_SOURCE}
Background blocks are each optional, joined by a blank line, always in that order. {TARGET_LANGUAGE} is a
full English language name. Recent source utterances carries source text only, never the translations.
The model keeps no state between requests — the caller owns the revision chain and sends it in full every time. That means you can retry, reorder, or drop requests freely.
Speed
| short prompt (33 tok) | long prompt (137 tok) | |
|---|---|---|
GPU (A40, CUDA, -ngl 99) |
0.06 s, 256 tok/s decode | 0.06 s, 337 tok/s decode |
| CPU (56-thread server) | 0.53 s, 17 tok/s decode | 1.16 s, 17 tok/s decode |
Rough expectations elsewhere: 8-core desktop CPU ~10–20 tok/s; Apple Silicon (Metal) ~30–60+ tok/s; any discrete GPU with ≥2 GB VRAM (CUDA or Vulkan) ~50–300 tok/s.
For live subtitles: under 100 ms per update on a GPU, 0.5–1.2 s on CPU alone — usable, if not instant.
Footprint
Weights are ~1.02 GiB once loaded. The KV cache is the only part that scales with your settings: this is a 32-layer model with 4 KV heads at head_dim 128, so 64 KiB per token at f16.
n_ctx |
KV cache | ballpark resident |
|---|---|---|
| 1024 | 64 MiB | ~1.3 GB |
| 2048 | 128 MiB | ~1.4 GB |
| 4096 | 256 MiB | ~1.5 GB |
| 8192 | 512 MiB | ~1.8 GB |
Measured: raising -c from 2048 to 4096 costs exactly 135 MB. --parallel costs nothing extra — -c is the
total KV budget, divided among slots, not multiplied by them.
Two things push the number above these estimates, and neither is the model's doing: the CUDA context and
allocator pool (often 300–500 MB, and nvidia-smi reports it against your process), and compute buffers
sized by -b/-ub. If you are counting VRAM on a small card, lower -c first — that is the part you
control.
What 4-bit costs
843 trajectories / 2,438 states, same prompts and greedy decoding on both sides.
| bf16 | this build | |
|---|---|---|
| state exact match | 0.287 | 0.282 |
| mean state similarity | 0.814 | 0.801 |
| final exact match | 0.199 | 0.211 |
| characters erased per append | 2.56 | 2.32 |
| prefix preserved (append transition) | 0.798 | 0.824 |
| prefix preserved (final transition) | 0.917 | 0.925 |
| empty-output rate | 0.000 | 0.000 |
About 1.3pp of similarity and half a point of exact match. The stability rows come out level or slightly ahead, which is within noise. Full table in the LoRA repo.
If you convert Hy-MT2 yourself
llama.cpp's HunYuan converter doesn't write an eog/eom token when the source model has a single-valued
eos_token_id — and Hy-MT2 does. Without a fix, generation never stops and you get endless repetition.
This build already has the fix applied; its metadata carries tokenizer.ggml.eom_token_id = 120020. If
you're converting from scratch, append this to _fix_special_tokens() in conversion/hunyuan.py:
eos = self.hparams.get("eos_token_id")
if eos is not None:
ids = eos if isinstance(eos, list) else [eos]
self.gguf_writer.add_eom_token_id(int(ids[0]))
Converted and quantized with llama.cpp b10442.
Limitations
- Stability is a tendency, not a guarantee. Nothing enforces prefix preservation; worst case a single update rewrites the whole line. If your UI can't tolerate that, only display the prefix that has held steady for N updates.
- Prompt format matters a lot — off-format prompts lose quality and stability.
- Language coverage is uneven: zh/en/ja are what it was trained on and measured on. Other directions fall back to base-model behaviour, untested.
- Greedy decoding assumed.
- Metric definitions are project-internal; don't compare them against numbers from simultaneous-MT papers.
License
Apache 2.0, same as the base model tencent/Hy-MT2-1.8B.
- Downloads last month
- 72
4-bit
Model tree for febilly/Hy-MT2-1.8B-StreamRevise-GGUF
Base model
tencent/Hy-MT2-1.8B