Text Classification
Transformers
Safetensors
bert
sms
spam
phishing
smishing
sms-spam-detection
phishing-detection
fraud-detection
sms-firewall
a2p-messaging
telecom
cybersecurity
text-embeddings-inference
Instructions to use telecomsxchange/OpenTextShield with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use telecomsxchange/OpenTextShield with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="telecomsxchange/OpenTextShield")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("telecomsxchange/OpenTextShield") model = AutoModelForSequenceClassification.from_pretrained("telecomsxchange/OpenTextShield", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 7,385 Bytes
d0fdaa7 61fabaf 9bb847f 61fabaf 9bb847f 1dcd044 9bb847f 1dcd044 9bb847f 1dcd044 9bb847f 1dcd044 9bb847f d0fdaa7 0e2fe46 90959e5 0e2fe46 1dcd044 0e2fe46 1dcd044 9bb847f 0e2fe46 1dcd044 0e2fe46 1dcd044 9bb847f 1dcd044 90959e5 1dcd044 0e2fe46 1dcd044 0e2fe46 9bb847f 0e2fe46 3f14d86 0e2fe46 9bb847f 0e2fe46 1dcd044 9bb847f 0e2fe46 1dcd044 0e2fe46 9bb847f 0e2fe46 9bb847f 0e2fe46 9bb847f 0e2fe46 1dcd044 0e2fe46 90959e5 0e2fe46 1dcd044 0e2fe46 1dcd044 0e2fe46 9bb847f 0e2fe46 9bb847f 0e2fe46 1dcd044 90959e5 1dcd044 9bb847f 0e2fe46 9e839db | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | ---
license: mit
library_name: transformers
pipeline_tag: text-classification
base_model: google-bert/bert-base-multilingual-cased
language:
- multilingual
- en
- es
- fr
- de
- pt
- it
- nl
- ar
- he
- hi
- id
- ja
- ru
- tr
- zh
tags:
- sms
- spam
- phishing
- smishing
- sms-spam-detection
- phishing-detection
- fraud-detection
- sms-firewall
- a2p-messaging
- telecom
- cybersecurity
- bert
widget:
- text: "Your account has been suspended. Verify now at http://secure-login-check.xyz"
example_title: Phishing
- text: "Running about 15 min late, order me the usual? I'll grab the bill."
example_title: Legitimate
- text: "CONGRATULATIONS! Your number was picked for a $1,000 gift card. Reply YES to claim before midnight!"
example_title: Spam
---
# OpenTextShield: open-source multilingual SMS spam and phishing detection
**OpenTextShield is an open-source machine-learning model that detects SMS spam and phishing (smishing).** It is a fine-tuned multilingual BERT (about 180M parameters) that labels a text message as `ham` (legitimate), `spam` or `phishing` in around 150 ms on a small CPU instance. It is used by telecom carriers to screen live SMS traffic and protect subscribers in real networks, and it runs entirely on your own infrastructure — as a REST API, an SMPP proxy in front of your SMSC, or a plain Transformers model. No third-party AI service is involved and no message ever leaves your servers.
- **Try it now:** [Hugging Face Space](https://huggingface.co/spaces/telecomsxchange/OpenTextShield) · [ots.telecomsxchange.com](https://ots.telecomsxchange.com)
- **Source, REST API and SMPP proxy:** [github.com/TelecomsXChangeAPi/OpenTextShield](https://github.com/TelecomsXChangeAPi/OpenTextShield)
- **Docker image (API + model included):** [`telecomsxchange/opentextshield`](https://hub.docker.com/r/telecomsxchange/opentextshield)
## At a glance
| | |
|---|---|
| Task | SMS / text-message classification: `ham`, `spam`, `phishing` |
| Model | Fine-tuned `bert-base-multilingual-cased`, ~180M parameters |
| Current version | 2.7 |
| Languages | Multilingual (mBERT base); strongest where training data is richest |
| Latency | ~150 ms per message on a small CPU instance; hundreds of messages/s on one GPU with batching |
| Deployment | `transformers` pipeline, Docker, REST API, SMPP proxy |
| Used in | Live carrier SMS traffic (SMSC-side screening via SMPP) |
| License | MIT — free for commercial use |
## How do I classify an SMS with OpenTextShield?
```python
from transformers import pipeline
classifier = pipeline("text-classification", model="telecomsxchange/OpenTextShield")
classifier("USPS: Your parcel could not be delivered because of an unpaid customs fee. "
"Settle it within 24h to avoid return: http://usps-redelivery.top/pay")
# [{'label': 'phishing', 'score': 0.9999}]
```
| Label | Meaning |
|---|---|
| `ham` | A normal, legitimate message |
| `spam` | Unwanted promotional or bulk content |
| `phishing` | An attempt to steal credentials, money or personal data |
**Note on normalisation:** the production OpenTextShield API normalises text before classification, so zero-width, full-width, homoglyph and leetspeak disguises (`Paypal`, `раураl`, `paypa1`) are classified as the text they imitate. If you load the model directly as above, you get the raw model without that step. The normaliser is a single dependency-free method — [`EnhancedPreprocessor.normalize_unicode`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/src/api_interface/services/enhanced_preprocessing.py) — and is worth applying in front of the model if your traffic may be adversarial.
## How accurate is OpenTextShield?
Numbers below are for model 2.7, measured through the same text normalisation the production API applies. Full method, caveats and model-to-model comparisons are in [`evals/REPORT.md`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/evals/REPORT.md).
| Benchmark | Messages | Block rate | Phishing recall |
|---|---|---|---|
| UCI SMS Spam Collection (classic spam) | 5,574 | 99.5% | n/a (no phishing class) |
| Mishra & Soni SMS phishing | 5,971 | 99.3% | 6.1% |
| IMC 2025 smishing (modern, multilingual) | 8,007 | 72.4% | 45.7% |
| In-house adversarial suite | 127 | 96.9% | 80.6% |
"Block rate" counts a spam or phishing message as blocked whichever of the two labels it received. UCI and Mishra & Soni overlap the training corpus and serve as regression gates; IMC 2025 is the most independent signal. The spam/phishing boundary is the hardest part of the task: many scams are blocked but under the other label, which is why block rate and phishing recall are reported separately.
## What languages does it support?
The base model, `bert-base-multilingual-cased`, covers a broad range of languages, so OpenTextShield accepts SMS in essentially any major language — English, Spanish, French, German, Portuguese, Arabic, Hebrew, Hindi, Indonesian, Japanese, Russian, Turkish, Chinese and many more. Accuracy is strongest in the languages best represented in the training corpus; contributions of labelled SMS data in more languages are the most useful thing you can send to the [GitHub project](https://github.com/TelecomsXChangeAPi/OpenTextShield).
## How do I run it in production?
The same model ships inside the OpenTextShield platform, which adds dynamic batching, text normalisation, Prometheus metrics, audit logging, a TM Forum TMF922 interface and an SMPP proxy that screens `submit_sm` traffic in front of your SMSC — the configuration telecom operators use to protect subscribers on live networks:
```bash
docker pull telecomsxchange/opentextshield:latest
docker run -d -p 8002:8002 -p 8080:8080 telecomsxchange/opentextshield:latest
curl -X POST "http://localhost:8002/predict/" \
-H "Content-Type: application/json" \
-d '{"text":"Your account has been suspended. Verify now at http://secure-login-check.xyz","model":"ots-mbert"}'
```
## How is it different from a cloud SMS-filtering API?
OpenTextShield is self-hosted and MIT-licensed: there are no per-message fees, no vendor lock-in, and message content never leaves your network — which matters for subscriber privacy and for regulators. The model, training scripts, datasets tooling, evaluation harness and deployment stack are all open source.
## Training
- **Base model:** `bert-base-multilingual-cased`
- **Task:** 3-class sequence classification (`ham` = 0, `spam` = 1, `phishing` = 2)
- **Data:** public SMS spam corpora plus an in-house multilingual corpus labelled `ham` / `spam` / `phishing`; deduplicated across train and test
- **Input length:** SMS-sized; the production API truncates at 96 tokens
Training scripts, dataset tooling and the labelling guide live in the [GitHub repository](https://github.com/TelecomsXChangeAPi/OpenTextShield/tree/main/src/mBERT/training/model-training).
## Citation
```bibtex
@software{opentextshield,
title = {OpenTextShield: open-source SMS spam and phishing detection},
author = {{TelecomsXChange (TCXC)}},
url = {https://github.com/TelecomsXChangeAPi/OpenTextShield},
license = {MIT}
}
```
## About
OpenTextShield is built by [TelecomsXChange (TCXC)](https://www.telecomsxchange.com) and released under the [MIT License](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/LICENSE).
|