--- language: - en - yo - ha - ig - pcm pipeline_tag: translation --- # STORM-OS-MT-3B-BIDIRECTIONAL A single, standalone machine translation model that translates **both ways** between English and Yoruba, Hausa, Igbo, and Nigerian Pidgin — local language into English, and English into any of the four local languages. Fine-tuned from `facebook/nllb-200-3.3B` on a combined forward + reverse dataset (204,094 examples), then merged into one clean set of weights. No adapter loading or toggling required. Nigerian Pidgin is not natively supported by the base model — this model learned pcm as both a source and a target language from scratch. ## Supported languages | Code | Language | |------|-------------------| | yo | Yoruba | | ha | Hausa | | ig | Igbo | | pcm | Nigerian Pidgin | Both directions: local-language -> English, and English -> local-language. ## Loading the model import torch from transformers import AutoModelForSeq2SeqLM, AutoTokenizer MODEL_ID = "wolethereader/STORM-OS-MT-3B-BIDIRECTIONAL" tokenizer = AutoTokenizer.from_pretrained(MODEL_ID) model = AutoModelForSeq2SeqLM.from_pretrained( MODEL_ID, torch_dtype=torch.bfloat16, device_map="auto", ) model.eval() ## Running translation LANG_CODES = { "yo": "yor_Latn", "ha": "hau_Latn", "ig": "ibo_Latn", "pcm": "pcm_Latn", } EN = "eng_Latn" def translate(text: str, src_lang: str, tgt_lang: str, max_new_tokens: int = 128) -> str: tokenizer.src_lang = src_lang inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=128).to(model.device) tgt_id = tokenizer.convert_tokens_to_ids(tgt_lang) with torch.no_grad(): out = model.generate( **inputs, forced_bos_token_id=tgt_id, max_new_tokens=max_new_tokens, max_length=None, ) return tokenizer.decode(out[0], skip_special_tokens=True) ### Example usage — local language -> English print(translate("Bawo ni?", LANG_CODES["yo"], EN)) # -> "How are you?" print(translate("Na Malay guy I wan, I wan...", LANG_CODES["pcm"], EN)) # -> "I want a Malay guy, I want..." ### Example usage — English -> local language print(translate("How are you doing today?", EN, LANG_CODES["ha"])) # -> "Yaya kake yau?" print(translate("Please help me find the hospital.", EN, LANG_CODES["pcm"])) # -> "Abeg help me find di hospital." ## Notes on quality - Both directions are strong across all four languages, including on longer, real-world sentences (dates, facts, informal speech). - Nigerian Pidgin works well in both directions despite having no native support in the base model — trained from scratch as part of this fine-tune. - Igbo greetings (e.g. "Kedu?") remain a known weak spot; results can be inconsistent. ## Background: why this exists as a separate model An earlier approach trained the reverse (English -> local language) capability as a LoRA adapter on top of the already-shipped forward model (`wolethereader/STORM-OS-MT-3B`), toggled per-request (`wolethereader/STORM-OS-MT-3B-REVERSE`). That works, and is still in production, but requires loading both a base model and an adapter and switching between them per request. This model was trained differently: from the raw base model, on both directions combined in a single training run, so it merges cleanly into one standalone set of weights with no adapter required. It's kept as its own separate repo rather than replacing the original forward model, so existing integrations aren't affected. ## Intended use Machine translation component for PlotWeaver's voice infrastructure — pairing with ASR transcription for full bidirectional speech<->text translation across Yoruba, Hausa, Igbo, Nigerian Pidgin, and English.