OpenTextShield / README.md
ajamous's picture
docs: link the company site at www.telecomsxchange.com
9e839db verified
|
Raw History Blame Contribute Delete
7.39 kB
---
license: mit
library_name: transformers
pipeline_tag: text-classification
base_model: google-bert/bert-base-multilingual-cased
language:
- multilingual
- en
- es
- fr
- de
- pt
- it
- nl
- ar
- he
- hi
- id
- ja
- ru
- tr
- zh
tags:
- sms
- spam
- phishing
- smishing
- sms-spam-detection
- phishing-detection
- fraud-detection
- sms-firewall
- a2p-messaging
- telecom
- cybersecurity
- bert
widget:
- text: "Your account has been suspended. Verify now at http://secure-login-check.xyz"
example_title: Phishing
- text: "Running about 15 min late, order me the usual? I'll grab the bill."
example_title: Legitimate
- text: "CONGRATULATIONS! Your number was picked for a $1,000 gift card. Reply YES to claim before midnight!"
example_title: Spam
---
# OpenTextShield: open-source multilingual SMS spam and phishing detection
**OpenTextShield is an open-source machine-learning model that detects SMS spam and phishing (smishing).** It is a fine-tuned multilingual BERT (about 180M parameters) that labels a text message as `ham` (legitimate), `spam` or `phishing` in around 150 ms on a small CPU instance. It is used by telecom carriers to screen live SMS traffic and protect subscribers in real networks, and it runs entirely on your own infrastructure — as a REST API, an SMPP proxy in front of your SMSC, or a plain Transformers model. No third-party AI service is involved and no message ever leaves your servers.
- **Try it now:** [Hugging Face Space](https://huggingface.co/spaces/telecomsxchange/OpenTextShield) · [ots.telecomsxchange.com](https://ots.telecomsxchange.com)
- **Source, REST API and SMPP proxy:** [github.com/TelecomsXChangeAPi/OpenTextShield](https://github.com/TelecomsXChangeAPi/OpenTextShield)
- **Docker image (API + model included):** [`telecomsxchange/opentextshield`](https://hub.docker.com/r/telecomsxchange/opentextshield)
## At a glance
| | |
|---|---|
| Task | SMS / text-message classification: `ham`, `spam`, `phishing` |
| Model | Fine-tuned `bert-base-multilingual-cased`, ~180M parameters |
| Current version | 2.7 |
| Languages | Multilingual (mBERT base); strongest where training data is richest |
| Latency | ~150 ms per message on a small CPU instance; hundreds of messages/s on one GPU with batching |
| Deployment | `transformers` pipeline, Docker, REST API, SMPP proxy |
| Used in | Live carrier SMS traffic (SMSC-side screening via SMPP) |
| License | MIT — free for commercial use |
## How do I classify an SMS with OpenTextShield?
```python
from transformers import pipeline
classifier = pipeline("text-classification", model="telecomsxchange/OpenTextShield")
classifier("USPS: Your parcel could not be delivered because of an unpaid customs fee. "
"Settle it within 24h to avoid return: http://usps-redelivery.top/pay")
# [{'label': 'phishing', 'score': 0.9999}]
```
| Label | Meaning |
|---|---|
| `ham` | A normal, legitimate message |
| `spam` | Unwanted promotional or bulk content |
| `phishing` | An attempt to steal credentials, money or personal data |
**Note on normalisation:** the production OpenTextShield API normalises text before classification, so zero-width, full-width, homoglyph and leetspeak disguises (`Paypal`, `раураl`, `paypa1`) are classified as the text they imitate. If you load the model directly as above, you get the raw model without that step. The normaliser is a single dependency-free method — [`EnhancedPreprocessor.normalize_unicode`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/src/api_interface/services/enhanced_preprocessing.py) — and is worth applying in front of the model if your traffic may be adversarial.
## How accurate is OpenTextShield?
Numbers below are for model 2.7, measured through the same text normalisation the production API applies. Full method, caveats and model-to-model comparisons are in [`evals/REPORT.md`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/evals/REPORT.md).
| Benchmark | Messages | Block rate | Phishing recall |
|---|---|---|---|
| UCI SMS Spam Collection (classic spam) | 5,574 | 99.5% | n/a (no phishing class) |
| Mishra & Soni SMS phishing | 5,971 | 99.3% | 6.1% |
| IMC 2025 smishing (modern, multilingual) | 8,007 | 72.4% | 45.7% |
| In-house adversarial suite | 127 | 96.9% | 80.6% |
"Block rate" counts a spam or phishing message as blocked whichever of the two labels it received. UCI and Mishra & Soni overlap the training corpus and serve as regression gates; IMC 2025 is the most independent signal. The spam/phishing boundary is the hardest part of the task: many scams are blocked but under the other label, which is why block rate and phishing recall are reported separately.
## What languages does it support?
The base model, `bert-base-multilingual-cased`, covers a broad range of languages, so OpenTextShield accepts SMS in essentially any major language — English, Spanish, French, German, Portuguese, Arabic, Hebrew, Hindi, Indonesian, Japanese, Russian, Turkish, Chinese and many more. Accuracy is strongest in the languages best represented in the training corpus; contributions of labelled SMS data in more languages are the most useful thing you can send to the [GitHub project](https://github.com/TelecomsXChangeAPi/OpenTextShield).
## How do I run it in production?
The same model ships inside the OpenTextShield platform, which adds dynamic batching, text normalisation, Prometheus metrics, audit logging, a TM Forum TMF922 interface and an SMPP proxy that screens `submit_sm` traffic in front of your SMSC — the configuration telecom operators use to protect subscribers on live networks:
```bash
docker pull telecomsxchange/opentextshield:latest
docker run -d -p 8002:8002 -p 8080:8080 telecomsxchange/opentextshield:latest
curl -X POST "http://localhost:8002/predict/" \
-H "Content-Type: application/json" \
-d '{"text":"Your account has been suspended. Verify now at http://secure-login-check.xyz","model":"ots-mbert"}'
```
## How is it different from a cloud SMS-filtering API?
OpenTextShield is self-hosted and MIT-licensed: there are no per-message fees, no vendor lock-in, and message content never leaves your network — which matters for subscriber privacy and for regulators. The model, training scripts, datasets tooling, evaluation harness and deployment stack are all open source.
## Training
- **Base model:** `bert-base-multilingual-cased`
- **Task:** 3-class sequence classification (`ham` = 0, `spam` = 1, `phishing` = 2)
- **Data:** public SMS spam corpora plus an in-house multilingual corpus labelled `ham` / `spam` / `phishing`; deduplicated across train and test
- **Input length:** SMS-sized; the production API truncates at 96 tokens
Training scripts, dataset tooling and the labelling guide live in the [GitHub repository](https://github.com/TelecomsXChangeAPi/OpenTextShield/tree/main/src/mBERT/training/model-training).
## Citation
```bibtex
@software{opentextshield,
title = {OpenTextShield: open-source SMS spam and phishing detection},
author = {{TelecomsXChange (TCXC)}},
url = {https://github.com/TelecomsXChangeAPi/OpenTextShield},
license = {MIT}
}
```
## About
OpenTextShield is built by [TelecomsXChange (TCXC)](https://www.telecomsxchange.com) and released under the [MIT License](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/LICENSE).