Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| Audio_Originated_Jailbreak | 1,279 items | ||
| HarmfulQuery | 2 items | ||
| Text_Transferred_Jailbreak | 177 items | ||
| .gitattributes | 2.46 kB xet | 19463de8 | |
| README.md | 5.43 kB xet | d99fced3 |
About the Dataset
📦 JALMBench contains 245,355 audio samples and 11,316 text prompts to benchmark jailbreak attacks against audio-language models (ALMs). It consists of three main categories:
🔥 Harmful Query Category:
Includes 246 harmful text queries ($T_{Harm}$), their corresponding audio ($A_{Harm}$), and a diverse audio set ($A_{Div}$) with 9 languages, 2 genders, 3 accents, and 3 TTS methods.📒 Text-Transferred Jailbreak Category:
Features adversarial texts generated by 4 prompting methods—ICA, DAN, DI, and PAP—and their corresponding audio versions. PAP includes 9,840 samples using 40 persuasion styles per query.🎧 Audio-Originated Jailbreak Category:
Contains adversarial audio samples generated directly by 4 audio-level attacks—SSJ (masking), AMSE (audio editing), BoN (large-scale noise variants), and AdvWave (black-box optimization with GPT-4o).
Field Descriptions
Each parquet file contains the following fields:
id: Unique identifier for from the AdvBench Dataset, MM-SafetyBench, JailbreakBench, and HarmBench.text(for subsets $T_{Harm}$, $A_{Harm}$, $A_{Div}$, ICA, DAN, DI, and PAP): Harmful or Adversarial Text.original_text: Orginal Harmful Text Queries.audio(for all subsets except $T_{Harm}$): Harmful or Adversarial Audios.attempt_id(for subsets PAP, BoN, and AdvWave): Attempt number of one specific ID.target_model(for subst AdvWave): AdvWave includes this fields for target model, for testing one specific ID, you can ignore thetarget_modeland transfer to any other models.source: from which dataset
🧪 Evaluation
To evaluate ALMs on this dataset and reproduce the benchmark results, please refer to the official GitHub repository.
- Total size
- 138 GB
- Files
- 1,460
- Last updated
- Aug 15
- Pre-warmed CDN
- US EU US EU