TinyModels animated header



๐Ÿง  Tiny model. Tiny dataset. Real classification.

A compact few-shot spam classifier built with SetFit + BAAI/bge-small-en-v1.5.


โšก TINY โ†’ FAST โ†’ USEFUL

This model is a small experiment in few-shot text classification.

It learns to separate:

โœ‰๏ธ HAM โ€” legitimate email ๐Ÿšจ SPAM โ€” unwanted / suspicious email

The interesting part?

Only 16 labeled training examples.

                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ”‚      16 EXAMPLES     โ”‚
                 โ”‚                      โ”‚
                 โ”‚   8 HAM  +  8 SPAM   โ”‚
                 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
                โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                โ”‚ BAAI/bge-small-en-v1.5โ”‚
                โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
                   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                   โ”‚     SetFit     โ”‚
                   โ”‚  Few-shot NLP  โ”‚
                   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                           โ–ผ
                โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                โ”‚ Logistic Regression   โ”‚
                โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
                     โ–ผ             โ–ผ
                   HAM           SPAM

๐Ÿ“ก Model Status

โš™๏ธ Component ๐Ÿ”ง Configuration
Task Spam Classification
Classes 2
Backbone BAAI/bge-small-en-v1.5
Framework SetFit
Dataset SetFit/enron_spam
Training Examples 16
Examples / Class 8
Iterations 20
Epochs 1
Batch Size 16
Classification Head Logistic Regression
Language English

๐Ÿ“Š Results

91.4% Accuracy

91.4% Macro F1


Accuracy   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘  91.4%
Macro F1   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘  91.4%

These results come from a very small few-shot training setup. They should not be interpreted as a benchmark against production spam-filtering systems.


๐Ÿงฌ The TinyModels Recipe

              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚   SetFit/enron    โ”‚
              โ”‚      _spam        โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                  16 examples
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚     BGE-small     โ”‚
              โ”‚   text encoder    โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                   embeddings
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚      SetFit       โ”‚
              โ”‚ contrastive loss  โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚ LogisticRegressionโ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ–ผ             โ–ผ
              โœ‰๏ธ HAM        ๐Ÿšจ SPAM

Training configuration

backbone: BAAI/bge-small-en-v1.5
dataset: SetFit/enron_spam

examples:
  total: 16
  per_class: 8

setfit:
  iterations: 20
  epochs: 1
  batch_size: 16
  loss: CosineSimilarityLoss
  distance_metric: cosine_distance

classifier:
  type: LogisticRegression

seed: 42

๐Ÿš€ Run It

pip install -q setfit
from setfit import SetFitModel

model = SetFitModel.from_pretrained(
    "TinyModels/setfit-banking-spam"
)

text = """
Congratulations! You have won $1,000,000.
Click here immediately to claim your prize.
"""

prediction = model(text)

print(prediction)

Example output

1

๐Ÿงช Quick Examples

๐Ÿšจ Spam

CONGRATULATIONS!!!

You have been selected to receive
$1,000,000. Click the link below
to claim your prize immediately.
โ†’ SPAM

โœ‰๏ธ Legitimate

Please find attached the global markets
monitor for the week ending 12 January 2001.
โ†’ HAM

๐Ÿง  Why SetFit?

Traditional supervised classification can require a large labeled dataset.

SetFit takes a different route:

        Large Dataset
             โœ•
             โ”‚
             โ”‚
       โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”
       โ”‚  SetFit   โ”‚
       โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜
             โ”‚
             โ–ผ
     Few labeled examples
             โ”‚
             โ–ผ
       Useful classifier

This makes the experiment useful for exploring:

  • โšก Few-shot learning
  • ๐Ÿง  Sentence embeddings
  • ๐Ÿ“š Text classification
  • ๐Ÿšจ Spam detection
  • ๐Ÿ”ฌ Efficient training
  • ๐Ÿค— SetFit

๐Ÿงฉ Architecture

INPUT
  โ”‚
  โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  BAAI/bge-small-en-v1.5     โ”‚
โ”‚                             โ”‚
โ”‚  Sentence Transformer       โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
               โ”‚
               โ–ผ
        Dense Embedding
               โ”‚
               โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚    Logistic Regression      โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
               โ”‚
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ–ผ             โ–ผ
      HAM           SPAM
       โœ‰๏ธ             ๐Ÿšจ

โš ๏ธ Limitations

This is intentionally a tiny experimental model.

Because only 16 examples were used for training:

  • Performance can vary on unseen data.
  • Domain shift can significantly affect predictions.
  • Unusual spam may be missed.
  • Enron-style email does not represent every modern spam pattern.
  • The reported score comes from a lightweight few-shot experiment.

Do not use this model as the sole component of a security-critical email filtering system.


๐Ÿ“ฆ Model Identity

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚              TINYMODEL               โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                      โ”‚
โ”‚  MODEL      SetFit Banking Spam      โ”‚
โ”‚  BACKBONE   BGE-small                โ”‚
โ”‚  TASK       Binary Classification    โ”‚
โ”‚  DATA       Enron Spam               โ”‚
โ”‚  EXAMPLES   16                       โ”‚
โ”‚  RESULT     91.4% Accuracy           โ”‚
โ”‚                                      โ”‚
โ”‚  STATUS     โ— EXPERIMENTAL           โ”‚
โ”‚                                      โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โšก 16 Examples.

๐Ÿง  One Small Model.

๐Ÿšจ One Real Task.




TinyModels โ€” building small models that actually do things.

Downloads last month
4
Safetensors
Model size
33.4M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for TinyModels/Setfit-Banking-Spam

Finetuned
(413)
this model

Dataset used to train TinyModels/Setfit-Banking-Spam