13 रिपॉजिटरी
Automated adversarial testing and vulnerability scanning to detect biases and safety failures in AI models.
Distinct from Vulnerability Scanning: Existing candidates focus on source code, containers, or network vulnerabilities, not adversarial model-specific red-teaming
Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Model Red-Teaming. Refine with filters or upvote what's useful.
This project is a Docker educational resource and a collection of practical examples designed for learning containerization technologies. It serves as a guide for understanding container fundamentals, including the creation and management of custom images and the use of registries. The repository provides specialized references for container security hardening, such as managing kernel privileges and implementing supply chain security. It also includes tutorials for multi-container orchestration and a DevOps guide focused on CI/CD automation and image optimization. The material covers a broad
Provides resources for adversarial testing and vulnerability scanning to detect safety failures in AI models.
RagaAI-Catalyst is a suite of software implementation tools providing an SDK, dashboard, and platform for monitoring, debugging, red-teaming, and evaluating agentic AI workflows. It serves as an observability framework for tracing the execution paths of large language models and multi-agent systems. The project distinguishes itself through a security suite for automated red-teaming and vulnerability scanning to detect biases, alongside a centralized prompt registry that decouples templates from application code. It further provides an evaluation platform that combines synthetic data generatio
Provides a security suite for automated red-teaming and vulnerability scanning to detect model biases.
G0DM0D3 is a static web client and multi-model chat gateway designed for AI research, prompt optimization, and red teaming. It provides a unified interface to query numerous AI models in parallel, allowing for the simultaneous evaluation of different prompt variations and sampling parameters to identify the most successful outputs. The project features specialized tooling for probing safety filters and bypassing model constraints through an input perturbation engine that applies text obfuscation and character substitution. It includes a composite scoring system to rank model performance and a
Implements adversarial testing techniques to probe and bypass model safety filters.
Garak AI विश्वसनीयता को मापने, कमजोरियों के लिए स्कैन करने और अनुकूली जांच (adaptive probing) के माध्यम से सुरक्षा आकलन को स्वचालित करने के लिए टूल्स का एक सूट है। यह एक जेनरेटिव AI भेद्यता स्कैनर और मूल्यांकन टूल के रूप में कार्य करता है जिसे भाषा मॉडल्स में सुरक्षा अंतराल, मतिभ्रम (hallucinations) और विफलता मोड की पहचान करने के लिए डिज़ाइन किया गया है। यह फ्रेमवर्क रेड-टीमिंग और सुरक्षा आकलन के लिए एक टूलकिट प्रदान करता है, जो विफलता दरों की गणना करने के लिए जांच और डिटेक्टरों की एक संरचित प्रणाली का उपयोग करता है। यह विशेष रूप से प्रतिकूल इनपुट्स के लिए मॉडल प्रतिक्रियाओं को रिकॉर्ड करके डेटा रिसाव और प्रॉम्प्ट इंजेक्शन जैसे जोखिमों के लिए स्कैन करता है। इस प्रोजेक्ट में टेस्ट सूट का विस्तार करने के लिए कस्टम जनरेटर और डिटेक्टर बनाने के लिए एक प्लगइन सिस्टम शामिल है। यह क्लाउड APIs, लोकल फाइलों और वेब एंडपॉइंट्स के माध्यम से विभिन्न AI मॉडल प्रोवाइडर्स के साथ एकीकरण का समर्थन करता है, जिसके परिणाम सुरक्षा आकलन लॉगिंग के माध्यम से बनाए रखे जाते हैं।
Implements automated adversarial testing and vulnerability scanning to identify safety failures in generative AI models.
Garak is an AI model evaluation tool and vulnerability scanner designed for red teaming large language models and auditing the security of retrieval-augmented generation pipelines. It identifies behavioral weaknesses, such as jailbreaks, hallucinations, and data leakage, by simulating adversarial attacks and executing automated testing vectors. The framework utilizes an adaptive probing loop where prompts can react to previous model behavior and be modified in flight via middleware. To ensure consistent analysis, it employs a provider-agnostic interface to interact with various model APIs and
Simulates adversarial attacks to identify behavioral weaknesses, jailbreaks, and misinformation risks in AI models.
Superagent is a framework for AI assistant orchestration and agent security. It provides the tools to build intelligent assistants that integrate external APIs and maintain conversation memory to automate complex tasks. The project focuses on AI agent security through adversarial testing, red teaming, and the detection of prompt injections and malicious tool calls. It includes automated vulnerability patching, which scans codebases and configurations for security flaws and generates pull requests with fixes. The platform supports retrieval augmented generation by connecting language models t
Provides automated adversarial testing and vulnerability scanning specifically designed to detect safety failures in AI models.
Giskard एक AI गुणवत्ता आश्वासन सूट और मूल्यांकन फ्रेमवर्क है जिसे लार्ज लैंग्वेज मॉडल और AI एजेंटों के प्रदर्शन, पूर्वाग्रह और सुरक्षा जोखिमों को मापने के लिए डिज़ाइन किया गया है। यह सुरक्षा खामियों और प्रदर्शन रिग्रेशन का पता लगाने के लिए एक भेद्यता स्कैनर के रूप में कार्य करता है। यह प्रोजेक्ट स्वचालित रेड-टीमिंग और प्रतिकूल परीक्षण वर्कफ़्लो प्रदान करता है। ये उपकरण सुरक्षा अंतराल और कमजोरियों की पहचान करने के लिए सिस्टम विवरण के आधार पर प्रॉम्प्ट-इंजेक्शन प्रोब और प्रतिकूल हमले उत्पन्न करते हैं। यह प्लेटफ़ॉर्म AI एजेंट ऑडिटिंग और RAG गुणवत्ता वैलिडेशन को कवर करता है, जो तथ्यात्मक सटीकता को सत्यापित करने के लिए नॉलेज-बेस ग्राउंडिंग और सिंथेटिक डेटा जनरेशन का उपयोग करता है। यह गैर-निर्धारित आउटपुट को मान्य करने के लिए दावा-आधारित मूल्यांकन और सिमेंटिक समानता मिलान के माध्यम से रिग्रेशन परीक्षण को भी संभालता है।
Implements automated adversarial testing and vulnerability scanning to detect biases and safety failures in AI models.
Giskard लार्ज लैंग्वेज मॉडल्स (LLMs) और AI एजेंट्स के लिए एक इवैल्यूएशन फ्रेमवर्क, टेस्टिंग लाइब्रेरी और क्वालिटी मॉनिटरिंग सिस्टम है। यह मॉडल के प्रदर्शन और विश्वसनीयता को मापने के लिए एक टूलकिट के रूप में कार्य करता है, जो RAG (retrieval-augmented generation) पाइपलाइन्स को वैलिडेट करने के लिए विशेष क्षमताएं प्रदान करता है। यह प्रोजेक्ट अपने ऑटोमेटेड रेड टीमिंग टूल और सिक्योरिटी स्कैनर के माध्यम से खुद को अलग बनाता है, जिसे कमजोरियों, प्रॉम्प्ट इंजेक्शन और सुरक्षा जोखिमों की पहचान करने के लिए डिज़ाइन किया गया है। यह मॉडल की मजबूती को मापने और जानकारी लीक होने का पता लगाने के लिए एडवर्सरियल प्रोबिंग और सिंथेटिक एज केस जनरेशन का उपयोग करता है। यह प्लेटफॉर्म तथ्यात्मक सटीकता और मतिभ्रम (hallucination) का पता लगाने, तर्क और लॉजिक बेंचमार्किंग, और बायस डिटेक्शन जैसी व्यापक क्षमताएं प्रदान करता है। यह रिग्रेशन टेस्टिंग, RAG कंपोनेंट असेसमेंट और नॉलेज बेस से टेस्ट केस जनरेट करने के लिए टूल्स प्रदान करता है। सिस्टम में सहयोगी टीम वर्कस्पेस, रोल-आधारित एक्सेस कंट्रोल और समय के साथ प्रदर्शन में गिरावट की निगरानी के लिए शेड्यूल्ड इवैल्यूएशन पाइपलाइन्स जैसी मैनेजमेंट सुविधाएं शामिल हैं।
Provides automated adversarial testing and vulnerability scanning to identify safety failures and biases in AI models.
This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.
Implements automated adversarial testing to detect safety failures and bypass safety filters in LLMs.
lmms-eval is a benchmarking system and performance analysis suite designed to measure the capabilities of large multimodal models. It provides a framework for evaluating models across text, image, audio, and video datasets, serving as a multimodal dataset orchestrator and benchmarking tool to quantify accuracy and efficiency. The project distinguishes itself through a unified multimodal message protocol that structures diverse media inputs for consistent model consumption. It features specialized benchmarking for audio, video, visual, document, and spatial reasoning, alongside tools for model
Implements adversarial testing using red-teaming datasets to detect hallucinations, biases, and jailbreak vulnerabilities.
This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and security testing. It provides a library of advanced templates and frameworks designed to optimize the quality and specificity of model responses. The project includes resources for red teaming and security research, featuring a repository of prompts designed to bypass safety filters and operational constraints. It also provides techniques for system prompt extraction to reveal the internal instructions and configurations of AI personas. The collection covers a broader surface
Offers a framework for adversarial testing to identify safety failures and prompt-based vulnerabilities.
PyRIT is an AI vulnerability assessment tool and security scanner designed to detect risks in large language model applications. It functions as a generative AI red teaming framework used to simulate adversarial attacks and identify weaknesses in system guardrails. The tool automates AI risk assessment by scanning generative AI components for security vulnerabilities. It utilizes automated testing and analysis to identify security gaps and prevent potential exploits through a consistent, repeatable process. The system incorporates asynchronous model orchestration to compare security postures
Provides a framework for adversarial testing and vulnerability scanning to detect safety failures in AI models.
AI-Infra-Guard is a security scanning platform designed to detect vulnerabilities across large language model deployments, AI agent skills, and the underlying infrastructure. It functions as a security toolset for auditing source code, evaluating model robustness, and identifying insecure network configurations. The project provides a red teaming framework that uses curated attack datasets to test for jailbreak vulnerabilities and prompt injections. It also includes an infrastructure auditor that employs network fingerprinting and asset discovery to match running components against known comm
Evaluates model robustness using curated attack datasets to detect potential jailbreak vulnerabilities.