File size: 7,385 Bytes
d0fdaa7
61fabaf
9bb847f
 
 
61fabaf
9bb847f
1dcd044
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9bb847f
 
 
 
 
1dcd044
 
9bb847f
1dcd044
 
9bb847f
1dcd044
9bb847f
 
 
 
 
 
 
 
d0fdaa7
0e2fe46
90959e5
0e2fe46
1dcd044
0e2fe46
1dcd044
 
9bb847f
0e2fe46
1dcd044
0e2fe46
1dcd044
9bb847f
1dcd044
 
 
90959e5
1dcd044
 
 
 
0e2fe46
1dcd044
0e2fe46
9bb847f
 
0e2fe46
3f14d86
0e2fe46
9bb847f
 
 
0e2fe46
 
1dcd044
 
 
 
 
 
9bb847f
0e2fe46
1dcd044
0e2fe46
9bb847f
0e2fe46
9bb847f
 
 
 
 
 
0e2fe46
9bb847f
0e2fe46
1dcd044
0e2fe46
90959e5
0e2fe46
1dcd044
0e2fe46
1dcd044
0e2fe46
 
9bb847f
 
0e2fe46
9bb847f
 
 
0e2fe46
 
1dcd044
 
 
 
 
 
90959e5
1dcd044
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9bb847f
0e2fe46
9e839db
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
---
license: mit
library_name: transformers
pipeline_tag: text-classification
base_model: google-bert/bert-base-multilingual-cased
language:
- multilingual
- en
- es
- fr
- de
- pt
- it
- nl
- ar
- he
- hi
- id
- ja
- ru
- tr
- zh
tags:
- sms
- spam
- phishing
- smishing
- sms-spam-detection
- phishing-detection
- fraud-detection
- sms-firewall
- a2p-messaging
- telecom
- cybersecurity
- bert
widget:
- text: "Your account has been suspended. Verify now at http://secure-login-check.xyz"
  example_title: Phishing
- text: "Running about 15 min late, order me the usual? I'll grab the bill."
  example_title: Legitimate
- text: "CONGRATULATIONS! Your number was picked for a $1,000 gift card. Reply YES to claim before midnight!"
  example_title: Spam
---

# OpenTextShield: open-source multilingual SMS spam and phishing detection

**OpenTextShield is an open-source machine-learning model that detects SMS spam and phishing (smishing).** It is a fine-tuned multilingual BERT (about 180M parameters) that labels a text message as `ham` (legitimate), `spam` or `phishing` in around 150 ms on a small CPU instance. It is used by telecom carriers to screen live SMS traffic and protect subscribers in real networks, and it runs entirely on your own infrastructure — as a REST API, an SMPP proxy in front of your SMSC, or a plain Transformers model. No third-party AI service is involved and no message ever leaves your servers.

- **Try it now:** [Hugging Face Space](https://huggingface.co/spaces/telecomsxchange/OpenTextShield) · [ots.telecomsxchange.com](https://ots.telecomsxchange.com)
- **Source, REST API and SMPP proxy:** [github.com/TelecomsXChangeAPi/OpenTextShield](https://github.com/TelecomsXChangeAPi/OpenTextShield)
- **Docker image (API + model included):** [`telecomsxchange/opentextshield`](https://hub.docker.com/r/telecomsxchange/opentextshield)

## At a glance

| | |
|---|---|
| Task | SMS / text-message classification: `ham`, `spam`, `phishing` |
| Model | Fine-tuned `bert-base-multilingual-cased`, ~180M parameters |
| Current version | 2.7 |
| Languages | Multilingual (mBERT base); strongest where training data is richest |
| Latency | ~150 ms per message on a small CPU instance; hundreds of messages/s on one GPU with batching |
| Deployment | `transformers` pipeline, Docker, REST API, SMPP proxy |
| Used in | Live carrier SMS traffic (SMSC-side screening via SMPP) |
| License | MIT — free for commercial use |

## How do I classify an SMS with OpenTextShield?

```python
from transformers import pipeline

classifier = pipeline("text-classification", model="telecomsxchange/OpenTextShield")

classifier("USPS: Your parcel could not be delivered because of an unpaid customs fee. "
           "Settle it within 24h to avoid return: http://usps-redelivery.top/pay")
# [{'label': 'phishing', 'score': 0.9999}]
```

| Label | Meaning |
|---|---|
| `ham` | A normal, legitimate message |
| `spam` | Unwanted promotional or bulk content |
| `phishing` | An attempt to steal credentials, money or personal data |

**Note on normalisation:** the production OpenTextShield API normalises text before classification, so zero-width, full-width, homoglyph and leetspeak disguises (`Paypal`, `раураl`, `paypa1`) are classified as the text they imitate. If you load the model directly as above, you get the raw model without that step. The normaliser is a single dependency-free method — [`EnhancedPreprocessor.normalize_unicode`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/src/api_interface/services/enhanced_preprocessing.py) — and is worth applying in front of the model if your traffic may be adversarial.

## How accurate is OpenTextShield?

Numbers below are for model 2.7, measured through the same text normalisation the production API applies. Full method, caveats and model-to-model comparisons are in [`evals/REPORT.md`](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/evals/REPORT.md).

| Benchmark | Messages | Block rate | Phishing recall |
|---|---|---|---|
| UCI SMS Spam Collection (classic spam) | 5,574 | 99.5% | n/a (no phishing class) |
| Mishra & Soni SMS phishing | 5,971 | 99.3% | 6.1% |
| IMC 2025 smishing (modern, multilingual) | 8,007 | 72.4% | 45.7% |
| In-house adversarial suite | 127 | 96.9% | 80.6% |

"Block rate" counts a spam or phishing message as blocked whichever of the two labels it received. UCI and Mishra & Soni overlap the training corpus and serve as regression gates; IMC 2025 is the most independent signal. The spam/phishing boundary is the hardest part of the task: many scams are blocked but under the other label, which is why block rate and phishing recall are reported separately.

## What languages does it support?

The base model, `bert-base-multilingual-cased`, covers a broad range of languages, so OpenTextShield accepts SMS in essentially any major language — English, Spanish, French, German, Portuguese, Arabic, Hebrew, Hindi, Indonesian, Japanese, Russian, Turkish, Chinese and many more. Accuracy is strongest in the languages best represented in the training corpus; contributions of labelled SMS data in more languages are the most useful thing you can send to the [GitHub project](https://github.com/TelecomsXChangeAPi/OpenTextShield).

## How do I run it in production?

The same model ships inside the OpenTextShield platform, which adds dynamic batching, text normalisation, Prometheus metrics, audit logging, a TM Forum TMF922 interface and an SMPP proxy that screens `submit_sm` traffic in front of your SMSC — the configuration telecom operators use to protect subscribers on live networks:

```bash
docker pull telecomsxchange/opentextshield:latest
docker run -d -p 8002:8002 -p 8080:8080 telecomsxchange/opentextshield:latest

curl -X POST "http://localhost:8002/predict/" \
  -H "Content-Type: application/json" \
  -d '{"text":"Your account has been suspended. Verify now at http://secure-login-check.xyz","model":"ots-mbert"}'
```

## How is it different from a cloud SMS-filtering API?

OpenTextShield is self-hosted and MIT-licensed: there are no per-message fees, no vendor lock-in, and message content never leaves your network — which matters for subscriber privacy and for regulators. The model, training scripts, datasets tooling, evaluation harness and deployment stack are all open source.

## Training

- **Base model:** `bert-base-multilingual-cased`
- **Task:** 3-class sequence classification (`ham` = 0, `spam` = 1, `phishing` = 2)
- **Data:** public SMS spam corpora plus an in-house multilingual corpus labelled `ham` / `spam` / `phishing`; deduplicated across train and test
- **Input length:** SMS-sized; the production API truncates at 96 tokens

Training scripts, dataset tooling and the labelling guide live in the [GitHub repository](https://github.com/TelecomsXChangeAPi/OpenTextShield/tree/main/src/mBERT/training/model-training).

## Citation

```bibtex
@software{opentextshield,
  title   = {OpenTextShield: open-source SMS spam and phishing detection},
  author  = {{TelecomsXChange (TCXC)}},
  url     = {https://github.com/TelecomsXChangeAPi/OpenTextShield},
  license = {MIT}
}
```

## About

OpenTextShield is built by [TelecomsXChange (TCXC)](https://www.telecomsxchange.com) and released under the [MIT License](https://github.com/TelecomsXChangeAPi/OpenTextShield/blob/main/LICENSE).