How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
The official implementation of our NAACL 2024 paper "A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily".
The main features of njunlp/renellm are: Defense Strategies, Jailbreak Attack Methods.
Open-source alternatives to njunlp/renellm include: damo-nlp-sg/multilingual-safety-for-llms — 📄 Paper • 🤗 Dataset. aounon/certified-llm-safety — This repository contains code for the paper Certifying LLM Safety against Adversarial Prompting. arobey1/smooth-llm — This is the official source code for "SmoothLLM: Defending LLMs Against Jailbreaking Attacks" by Alex Robey, Eric… cam-fss/jailbreak-langchain — In this paper, we conduct the first work to propose the concept of indirect jailbreak and achieve Retrieval-Augmented… chawins/pal — Chawin Sitawarin 1 Norman Mu 1 David Wagner 1 Alexandre Araujo 2. agencyenterprise/promptinject — PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the…
This repository contains code for the paper Certifying LLM Safety against Adversarial Prompting.
This is the official source code for "SmoothLLM: Defending LLMs Against Jailbreaking Attacks" by Alex Robey, Eric Wong, Hamed Hassani, and George J. Pappas. To learn more about our work, see our blog post.
PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022