Brynn Knowlton, Jovani Campa, David Solis Gallo, Khalil Dajani, Nabeel Alzahrani
Generative AI systems—particularly large language models (LLMs)—remain vulnerable to jailbreak attacks: adversarial prompts that bypass safeguards and elicit unsafe or restricted outputs. This survey synthesizes jailbreak research from 2023–2025, covering attack methods, defense strategies, and evaluation frameworks. Jailbreak techniques are grouped into five main categories: prompt-based injections, role-play conditioning, multi-turn dialogue, multilingual or multimodal exploits, and optimization-driven pipelines. We also examine discovery methods such as human red teaming and automated attacker systems (e.g., AutoJailbreak, LLMStinger), and review defense mechanisms including supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), adversarial fine-tuning, and output or pipeline filtering. Standardized benchmarks—such as PromptBench, JailbreakBench, and multimodal evaluation suites—are analyzed in terms of reproducibility, coverage, and robustness testing. Persistent issues of generalization, reproducibility, and scalability are discussed, along with tradeoffs between safety alignment and model utility. This paper provides an up-to-date reference clarifying the evolving landscape of jailbreak attacks and defenses in LLMs.