Benign
Standard non-adversarial user prompts for baseline evaluation and false-positive measurement.
Explore Prompts →
security research dataset
A curated research dataset containing 731,908 prompts across 13 attack categories. Engineered for benchmarking LLM vulnerabilities, guardrail robustness, and prompt injection defenses.
Empowering the open AI safety and red-teaming research ecosystem. Star the repository to track novel jailbreak vectors, or stream all 731,908 prompts natively in Python with one line of code.
Star this repository to bookmark the leading prompt injection benchmark. Track ongoing research, submit novel jailbreak payloads, and help fortify LLM guardrails against emerging threats.
Stream or download all 731,908 prompts directly in Apache Parquet / Arrow format on Hugging Face Datasets Hub for safety evaluations, model fine-tuning, and red-team audits.
Click any category card to immediately load, filter, and inspect its prompts in the interactive explorer below.
Standard non-adversarial user prompts for baseline evaluation and false-positive measurement.
Explore Prompts →Direct instruction override commands attempting to negate developer instructions and rules.
Explore Prompts →Persona assumption, fictional framing, and theatrical roleplay to evade policy constraints.
Explore Prompts →Adversarial jailbreaks including DAN variants, hypothetical worlds, and developer mode framing.
Explore Prompts →Base64, ROT13, hex, leetspeak, and multi-language encodings designed to bypass keyword filters.
Explore Prompts →Prompts attempting to elicit exploit code, cyber weaponization, malware, or harmful output.
Explore Prompts →Methods aiming to leak confidential context, API credentials, private memory, or sensitive files.
Explore Prompts →Multi-turn adversarial dialogues progressively steering the LLM away from safety guardrails.
Explore Prompts →Indirect prompt injection hidden inside external documents, emails, web pages, and RAG data.
Explore Prompts →Targeted attempts to extract, replicate, or leak hidden system instructions and guardrail definitions.
Explore Prompts →Payloads fragmented across whitespace, unicode control characters, and steganographic delimiters.
Explore Prompts →Email-vector prompt injection embedded inside inbox bodies and message automation pipelines.
Explore Prompts →URL parameters and markdown link injections inducing unauthorized web navigations.
Explore Prompts →A comprehensive 10-stage system-level framework mapping 58 attack techniques across large language models, autonomous agents, and tool-augmented workflows.
Prompt Probing & Behavioral Fingerprinting
Tool Surface Discovery
Safety Boundary Mapping
Context / Token Pressure Attacks
RAG / Knowledge Source Inference
Memory & State Detection
Hallucination & Confidence Testing