Instructions to use TinyModels/Setfit-Banking-Spam with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- setfit
How to use TinyModels/Setfit-Banking-Spam with setfit:
from setfit import SetFitModel model = SetFitModel.from_pretrained("TinyModels/Setfit-Banking-Spam") - sentence-transformers
How to use TinyModels/Setfit-Banking-Spam with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("TinyModels/Setfit-Banking-Spam") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
๐ง Tiny model. Tiny dataset. Real classification.
A compact few-shot spam classifier built with SetFit + BAAI/bge-small-en-v1.5.
This model is a small experiment in few-shot text classification.
It learns to separate:
โ๏ธ HAM โ legitimate email ๐จ SPAM โ unwanted / suspicious email
The interesting part?
Only 16 labeled training examples.
โโโโโโโโโโโโโโโโโโโโโโโโ
โ 16 EXAMPLES โ
โ โ
โ 8 HAM + 8 SPAM โ
โโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโ
โ BAAI/bge-small-en-v1.5โ
โโโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโ
โ SetFit โ
โ Few-shot NLP โ
โโโโโโโโโฌโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโ
โ Logistic Regression โ
โโโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โโโโโโโโดโโโโโโโ
โผ โผ
HAM SPAM
๐ก Model Status
| โ๏ธ Component | ๐ง Configuration |
|---|---|
| Task | Spam Classification |
| Classes | 2 |
| Backbone | BAAI/bge-small-en-v1.5 |
| Framework | SetFit |
| Dataset | SetFit/enron_spam |
| Training Examples | 16 |
| Examples / Class | 8 |
| Iterations | 20 |
| Epochs | 1 |
| Batch Size | 16 |
| Classification Head | Logistic Regression |
| Language | English |
๐ Results
91.4% Accuracy
91.4% Macro F1
Accuracy โโโโโโโโโโโโโโโโโโโโ 91.4%
Macro F1 โโโโโโโโโโโโโโโโโโโโ 91.4%
These results come from a very small few-shot training setup. They should not be interpreted as a benchmark against production spam-filtering systems.
๐งฌ The TinyModels Recipe
โโโโโโโโโโโโโโโโโโโโโ
โ SetFit/enron โ
โ _spam โ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
16 examples
โ
โผ
โโโโโโโโโโโโโโโโโโโโโ
โ BGE-small โ
โ text encoder โ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
embeddings
โ
โผ
โโโโโโโโโโโโโโโโโโโโโ
โ SetFit โ
โ contrastive loss โ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโ
โ LogisticRegressionโ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
โโโโโโโโดโโโโโโโ
โผ โผ
โ๏ธ HAM ๐จ SPAM
Training configuration
backbone: BAAI/bge-small-en-v1.5
dataset: SetFit/enron_spam
examples:
total: 16
per_class: 8
setfit:
iterations: 20
epochs: 1
batch_size: 16
loss: CosineSimilarityLoss
distance_metric: cosine_distance
classifier:
type: LogisticRegression
seed: 42
๐ Run It
pip install -q setfit
from setfit import SetFitModel
model = SetFitModel.from_pretrained(
"TinyModels/setfit-banking-spam"
)
text = """
Congratulations! You have won $1,000,000.
Click here immediately to claim your prize.
"""
prediction = model(text)
print(prediction)
Example output
1
๐งช Quick Examples
๐จ Spam
CONGRATULATIONS!!!
You have been selected to receive
$1,000,000. Click the link below
to claim your prize immediately.
โ SPAM
โ๏ธ Legitimate
Please find attached the global markets
monitor for the week ending 12 January 2001.
โ HAM
๐ง Why SetFit?
Traditional supervised classification can require a large labeled dataset.
SetFit takes a different route:
Large Dataset
โ
โ
โ
โโโโโโโผโโโโโโ
โ SetFit โ
โโโโโโโฌโโโโโโ
โ
โผ
Few labeled examples
โ
โผ
Useful classifier
This makes the experiment useful for exploring:
- โก Few-shot learning
- ๐ง Sentence embeddings
- ๐ Text classification
- ๐จ Spam detection
- ๐ฌ Efficient training
- ๐ค SetFit
๐งฉ Architecture
INPUT
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ BAAI/bge-small-en-v1.5 โ
โ โ
โ Sentence Transformer โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โ
โผ
Dense Embedding
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Logistic Regression โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โ
โโโโโโโโดโโโโโโโ
โผ โผ
HAM SPAM
โ๏ธ ๐จ
โ ๏ธ Limitations
This is intentionally a tiny experimental model.
Because only 16 examples were used for training:
- Performance can vary on unseen data.
- Domain shift can significantly affect predictions.
- Unusual spam may be missed.
- Enron-style email does not represent every modern spam pattern.
- The reported score comes from a lightweight few-shot experiment.
Do not use this model as the sole component of a security-critical email filtering system.
๐ฆ Model Identity
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ TINYMODEL โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ
โ MODEL SetFit Banking Spam โ
โ BACKBONE BGE-small โ
โ TASK Binary Classification โ
โ DATA Enron Spam โ
โ EXAMPLES 16 โ
โ RESULT 91.4% Accuracy โ
โ โ
โ STATUS โ EXPERIMENTAL โ
โ โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
- Downloads last month
- 4
Model tree for TinyModels/Setfit-Banking-Spam
Base model
BAAI/bge-small-en-v1.5