💥Derail Yourself: Multi-turn LLM Jailbreak Attack through Self-discovered Clues
The main features of renqibing/actorattack are: Multi Turn Attacks.
Open-source alternatives to renqibing/actorattack include: fmmarkmq/sema — The official repository for the paper: SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks. jinxiaolong1129/foot-in-the-door-jailbreak — Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A… ragib-amin-nihal/pe-coa — Code Implementation of "Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large… salman-lui/x-teaming — by Salman Rahman\, Liwei Jiang\, James Shiffer\, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, Hamid Palangi, Kai-Wei… xxiqiao/trojail — TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards. yancykahn/coa — Large language models (LLMs) have achieved remarkable performance in various natural language processing tasks,…
The official repository for the paper: SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks.
Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A key challenge is jailbreak, where adversarial prompts bypass built-in safeguards to elicit harmful disallowed outputs. Inspired by psychological foot-in-the-door principles, we…
Code Implementation of "Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models"
by Salman Rahman\, Liwei Jiang\, James Shiffer\, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, Hamid Palangi, Kai-Wei Chang, Yejin Choi, Saadia Gabriel*