TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Abstract
Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distinguish fraud from lawful, near-domain calls rather than relying on topic-separated negative examples. We present TeleAntiFraud 2.0, constructed with our Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol. The pipeline transforms online fraud-case abstracts into profile-grounded scenarios, expands them through mixed-tree generation, realizes fraud and non-fraud dialogue paths under shared contexts, renders validated dialogues as role-matched speech, and freezes the resulting audio, labels, prompts, manifests, and provenance records for each monthly evaluation set. Each frozen set contains 900 Chinese calls, comprising 600 fraud and 300 near-domain non-fraud cases. Controlled text experiments show that three classifiers achieve perfect macro-averaged F1 (Macro-F1) when evaluated against unrelated or ordinary negatives, but drop to 0.65-0.68 with near-domain sibling negatives. Full-set audio and automatic-speech-recognition plus large-language-model (ASR+LLM) evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity. Together, these findings establish near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions. The accompanying research artifact includes the construction code, evaluation scripts, manifests, and documentation. Our dataset and code are available at https://anonymous.4open.science/r/TeleAntiFraud-2_0-EEB2/.
Community
Hi HF community! 👋 Co-first author here, sharing TeleAntiFraud 2.0, a refreshable benchmark for audio-based telecom fraud detection.
Can a model distinguish a scam from a legitimate call when both start with the same scenario and suspicious-sounding language? We construct paired fraud and non-fraud conversations that share context but diverge at the actions that determine whether fraud occurs.
A key finding: three text classifiers reach 1.00 Macro-F1 with unrelated or ordinary negatives, but drop to 0.65–0.68 with these closely matched negatives. Evaluations of audio models and ASR+LLM pipelines also reveal false-positive bias and prediction collapse.
Each monthly snapshot contains 900 synthetic Chinese calls and stays fixed for reproducible comparisons, while new snapshots can incorporate emerging scam patterns.
We’d love feedback from the audio and evaluation communities, especially on building harder legitimate-call negatives and measuring robustness across snapshots!
Get this paper in your agent:
hf papers read 2609.18748 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper