科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Artificial Intelligence2026-02-17· Adversarial system

Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses

Brynn Knowlton, Jovani Campa, David Solis Gallo, Khalil Dajani, Nabeel Alzahrani

原始摘要(英文原文)· Original abstract
Generative AI systems—particularly large language models (LLMs)—remain vulnerable to jailbreak attacks: adversarial prompts that bypass safeguards and elicit unsafe or restricted outputs. This survey synthesizes jailbreak research from 2023–2025, covering attack methods, defense strategies, and evaluation frameworks. Jailbreak techniques are grouped into five main categories: prompt-based injections, role-play conditioning, multi-turn dialogue, multilingual or multimodal exploits, and optimization-driven pipelines. We also examine discovery methods such as human red teaming and automated attacker systems (e.g., AutoJailbreak, LLMStinger), and review defense mechanisms including supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), adversarial fine-tuning, and output or pipeline filtering. Standardized benchmarks—such as PromptBench, JailbreakBench, and multimodal evaluation suites—are analyzed in terms of reproducibility, coverage, and robustness testing. Persistent issues of generalization, reproducibility, and scalability are discussed, along with tradeoffs between safety alignment and model utility. This paper provides an up-to-date reference clarifying the evolving landscape of jailbreak attacks and defenses in LLMs.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses — 科研速览 Science Skim