Brunei Malay Normalizer V3

2026-09-07: release qualification withdrawn. Independent testing with the exact released weights found severe Standard Malay and other-language preservation failures. Examples: Saya membawa bayi ke kandang untuk melihat kambing. changes infant bayi to pig babi; Sakit parut pembedahan ini masih berterusan. changes scar parut to abdomen perut; Spanish No tengo nada. becomes No tengo tiada.. Removing forced semantic rules alone leaves four of six new controls failing. Retained for research/diagnostic comparison; do not deploy as an approved normalizer. Historical synthetic scores below are reproducible regression results, not evidence of real-world readiness.

Product V3 normalizes Brunei Malay (Bahasa Melayu Brunei, BMB) expressions to Standard Malay using constrained, source-conditioned local edits. It is not a chat model and must not answer, summarize or freely paraphrase the input.

The release preserves meaning, voice, participant roles, polarity and scope, time/aspect/modality, quantities, punctuation and non-target content. Already standard Malay and other languages are copied exactly. Necessary local word order is allowed, for example Inda ku tahu to Saya tidak tahu; active/passive conversion is forbidden.

Release identity

  • Product version: 3
  • Run: release-candidate-v3-full-z-20260904
  • Base: FacebookAI/xlm-roberta-base
  • Pinned base revision: e73636d4f797dec63c3081bb6ed5c7b0bb3f2089
  • Architecture: XLMRobertaForTokenClassification plus deterministic renderer
  • Training: full-parameter BF16 token classification; no QLoRA and no system prompt
  • Selected checkpoint: step 300
  • Model-weight SHA256: 43723616637e4e9a00119fa9303cd6a50c968837a0b206109fb6722e86107e8d
  • Training-data manifest SHA256: 449d81def1006f26d825453e31e061fed8f701346c8383e12136c89103162baa

Do not run the raw Transformers token-classification pipeline as if it produced translated text. labels.json, candidate-index.json and the code under bmb_normalizer/ are part of the model behavior. The supported service entry point is serve.py.

Automated evidence

  • Frozen held-out test: 9,803/9,803 exact
  • Critical semantic cases: 42/42 exact
  • Retained independent probes: 952/952 exact
  • Post-candidate frozen M20: 108/108 exact
  • Retained M17–M19 regressions: 442/442 exact
  • Live Kaggle artifact replay: 1,502/1,502 exact
  • Post-deployment authored M21: 90/90 exact

These are provisional/synthetic engineering gates. They do not prove 99.99% accuracy on arbitrary real users. Current status is automated_release_gate_passed_native_review_pending; native Brunei Malay and medical-domain review remains required before production approval.

Quick deployment

See DEPLOYMENT.md. In summary:

python3 -m venv /opt/bmb-v3/venv
source /opt/bmb-v3/venv/bin/activate
pip install -r requirements.txt
hf download ssh2025/brunei-malay-normalizer-v3 --local-dir /opt/bmb-v3/model

export BMB_MODEL_DIR=/opt/bmb-v3/model
export BMB_API_KEY='replace-with-a-long-random-secret'
export BMB_DEVICE=cuda:0
python /opt/bmb-v3/model/serve.py

The OpenAI-compatible endpoint is POST /v1/chat/completions; use model ID bmb-normalizer-v3. Only the final user message is normalized. External system prompts are ignored rather than allowed to change the normalization contract.

Distribution notice

The base model is MIT-licensed. The fine-tuned release contains product-authorized provisional training material and locally implemented policies whose final redistribution/licensing terms have not been separately certified; the repository therefore uses license: other. Public availability is not a medical-device or linguistic-quality certification.

Downloads last month
34
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ssh2025/brunei-malay-normalizer-v3

Finetuned
(4223)
this model