nllb200-coreml-1024 / README.md
aoiandroid's picture
docs: neutralize TranslateBlue references in README
8955063 verified
|
Raw
History Blame Contribute Delete
4.17 kB
---
language:
- multilingual
language_details: "NLLB-200 supports 200 languages (see facebook/nllb-200-distilled-600M)"
license: "cc-by-nc-4.0"
tags:
- translation
- coreml
- nllb
- apple
- ios
- mlpackage
pipeline_tag: translation
datasets:
- flores-200
base_model: facebook/nllb-200-distilled-600M
---
# NLLB-200 CoreML 1024 (Float16)
CoreML conversion of [facebook/nllb-200-distilled-600M](https://huggingface.co/facebook/nllb-200-distilled-600M) (NLLB-200 / M2M100, distilled 600M parameters, 200 languages) for on-device inference on Apple platforms. Max sequence length **1024** matches the official model `max_position_embeddings`.
## Base model (official)
This CoreML model is derived from **[facebook/nllb-200-distilled-600M](https://huggingface.co/facebook/nllb-200-distilled-600M)** (NLLB-200 distilled 600M on the Hub). Key official specs:
| Item | Official value |
|------|----------------|
| Model type | M2M100 (encoder-decoder) |
| Parameters | ~600M (distilled) |
| Languages | 200 (Flores-200 coverage) |
| **Max position embeddings** | **1024** ([config.json](https://huggingface.co/facebook/nllb-200-distilled-600M/raw/main/config.json)) |
| Vocab size | 256,206 |
| License | [CC-BY-NC-4.0](https://creativecommons.org/licenses/by-nc/4.0/) |
From the [base model card](https://huggingface.co/facebook/nllb-200-distilled-600M): research and non-commercial use; single-sentence translation; not for production, medical, or legal domain; training used input lengths not exceeding 512 tokens (longer sequences may degrade). Not for certified translation. For full intended use, limitations, metrics (BLEU, spBLEU, chrF++), and ethical considerations, see the official model page.
*Note: Other NLLB variants (e.g. larger or MoE) may support longer context (e.g. 128k); this CoreML conversion follows the distilled 600M config, which has max 1024.*
## Model description
- **Base model:** [facebook/nllb-200-distilled-600M](https://huggingface.co/facebook/nllb-200-distilled-600M)
- **Precision:** Float16
- **Max length:** 1024 tokens (encoder and decoder), aligned to official `max_position_embeddings`
- **Architecture:** Encoder + KV-cache decoder (init + step), compatible with variable sequence length 1..1024.
## Files
| File / folder | Description |
|---------------|-------------|
| `NLLB_Encoder_1024.mlpackage` | Encoder (source encoding) |
| `NLLB_Decoder_1024_init.mlpackage` | Decoder first step (with encoder outputs) |
| `NLLB_Decoder_1024_step.mlpackage` | Decoder subsequent steps (with past KV cache) |
| `tokenizer/` | SentencePiece tokenizer (same as base model) |
| `config.json` | Model config |
## Usage
1. Download this repo (or clone via Hugging Face Hub).
2. Load the three CoreML packages: encoder, decoder init, decoder step.
3. Encode source with the included tokenizer (use source/target language codes, e.g. `eng_Latn`, `jpn_Jpan`).
4. Run decoder in a loop: call init once with encoder outputs, then step with the last token and updated past key/value tensors until EOS or max length.
For a smaller, 8-bit quantized variant (same 1024 length), see [aoiandroid/nllb200-coreml-1024-palettized](https://huggingface.co/aoiandroid/nllb200-coreml-1024-palettized).
## Conversion
Produced with the notebook [nllb200_coreml_colab_v3_fixed.ipynb](local iOS project (private)) (CoreML encoder + KV decoder trace and export). Alignment: 1024 matches the official `config.json` of the base model.
## License and intended use
- **License:** [CC-BY-NC-4.0](https://creativecommons.org/licenses/by-nc/4.0/) (same as the base model).
- **Intended use:** Research and non-commercial use. Not for production deployment. General domain only (not for medical or legal domain). See the [base model card](https://huggingface.co/facebook/nllb-200-distilled-600M) for full intended use, limitations, and ethical considerations.
## Citation
If you use this model, please cite the NLLB-200 paper and the base model:
```bibtex
@article{nllb2022,
title={No Language Left Behind: Scaling Human-Centered Machine Translation},
author={{NLLB Team} and others},
journal={arXiv preprint arXiv:2207.04672},
year={2022}
}
```