How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
This repository contains the code for the paper "Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks" by Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury, accepted at LREC-CoLING 2024
The main features of aetherprior/trickllm are: Evaluation Benchmarks.
Open-source alternatives to aetherprior/trickllm include: pyspur-dev/pyspur. datawhalechina/prompt-engineering-for-developers — This project is a technical curriculum and development guide focused on large language model prompt engineering,… aifeg/benchlmm — [ECCV 2024] BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models. albertwy/gpt-4v-evaluation — Data for evaluating GPT-4V. alexandrasouly/strongreject — This repository is no longer maintained and deprecated in favour of the repository at… ailab-cvc/seed-bench — (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap
ECCV 2024 BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.