How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-classification", model="specific-AI/email-agent-phishing-detection")
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.from_pretrained("specific-AI/email-agent-phishing-detection")
model = AutoModelForSequenceClassification.from_pretrained("specific-AI/email-agent-phishing-detection", device_map="auto")
Quick Links

specific-AI/email-agent-phishing-detection

A compact BERT phishing detector distilled with Specific AI. It classifies email content as phishing or not, for use in email agents and security-aware inbox workflows.

Task Single-label text classification
Base model bert-base-uncased
Training data ~15,000 examples
License MIT

Input format

Examples were trained on emails formatted as plain text with From, Subject, and body (blank line between the headers and the body):

From: <from>
Subject: <subject>

<body>

Pass inputs in this same shape at inference time for best results.

Labels

Label Meaning
True Phishing detected
False Phishing was not detected

Evaluation

Compared against gpt-5.4-mini as a teacher / baseline on the same evaluation set:

Metric gpt-5.4-mini SpecificAI
Accuracy 0.971 0.975
Precision 0.976 0.975
Recall 0.971 0.975
F1 score 0.972 0.975

Repository contents

This card ships both a full Hugging Face checkpoint and GGUF-ready artifacts:

  • Full BertForSequenceClassification weights (model.safetensors) + tokenizer
  • Head layers as NumPy files (pooler_*.npy, classifier_*.npy) for GGUF / Lemonade fusion
  • Encoder GGUF: bert-base-only.gguf (CLS pooling; use with raw / unnormalized embeddings)

Quick start β€” Transformers

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_id = "specific-AI/email-agent-phishing-detection"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()

text = """From: security@paypa1-support.com
Subject: Your account will be locked

Verify your password at http://example-phish.test/login to keep access."""
inputs = tokenizer(text, return_tensors="pt", truncation=True)
with torch.no_grad():
    logits = model(**inputs).logits
pred = model.config.id2label[int(logits.argmax(-1))]
print(pred)  # "True" or "False"

Quick start β€” Lemonade + specific-ai-tools

When running the GGUF encoder through Lemonade Server:

pip install specific-ai-tools
from specific_ai_tools.embedding_heads import LemonadeEmbeddingClassifier

classifier = LemonadeEmbeddingClassifier(
    lemonade_model_name="user.email-agent-phishing-detection",
    checkpoint="specific-AI/email-agent-phishing-detection:bert-base-only.gguf",
    lemonade_base_url="http://localhost:13305",
)

text = """From: noreply@secure-mail-alert.com
Subject: Reset your password now

Click here to reset your password immediately."""
result = classifier.predict_one(text)
print(result.predicted_labels, result.predicted_confidences)

See the Specific AI toolkit docs for llama-cpp and other embedding backends.

Intended use

  • Email / inbox agents that need a fast on-device or CPU phishing signal
  • Pre-filter or assistive scoring alongside other security controls

Out of scope: sole authority for blocking, quarantine, or legal determinations. Treat outputs as a high-throughput classifier signal and keep human / policy review in the loop for high-impact actions.

About Us

Specific AI is the automatic SLM distillation platform that turns task prompts into production-grade small language models in days β€” not weeks β€” so your subject matter experts can ship models without waiting on scarce data-science bandwidth.

We help enterprises move agentic AI from prototype to production with SLMs that are typically 1,000×–10,000Γ— smaller than teacher LLMs, run in milliseconds on CPUs or edge devices, and deliver the same or better task quality at a fraction of the cost β€” self-hosted on your cloud or downloaded for your own inference stack.

Prompt β†’ Distill β†’ Deploy. Bring your prompt and data, drop them into Specific AI, and get a validated small model ready to test and ship.

Ready to create SLMs at scale? Visit specific.ai.

License

MIT β€” see LICENSE.

Copyright (C) 2026 Specific AI Inc. All rights reserved.

Downloads last month
130
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for specific-AI/email-agent-phishing-detection

Quantized
(29)
this model