How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
π° Latest News π° - π‘οΈ What is HarmBench π‘οΈ - π Overview π - β Quick Start β - βοΈ Installation - π οΈ Running the Evaluation Pipeline - β Using your own models in HarmBench - β Using your own red teaming methods in HarmBench - π€ Classifiers - β Documentation β - π± HarmBench's Roadmap π± -β¦
The main features of centerforaisafety/harmbench are: Evaluation Benchmarks, Guardrails and AI Safety.
Open-source alternatives to centerforaisafety/harmbench include: pyspur-dev/pyspur. datawhalechina/prompt-engineering-for-developers β This project is a technical curriculum and development guide focused on large language model prompt engineering,β¦ ai45lab/openrt β Open-source red teaming framework for MLLMs with 42+ attack methods. ailab-cvc/seed-bench β (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions. albertwy/gpt-4v-evaluation β Data for evaluating GPT-4V. aifeg/benchlmm β [ECCV 2024] BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models.
This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap
Open-source red teaming framework for MLLMs with 42+ attack methods
ECCV 2024 BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models