14 个仓库
Creation of edge-case and adversarial inputs to stress-test AI models for safety and brand risks.
Distinct from Adversarial Robustness Testing: Focuses specifically on the generation of inputs for AI model testing rather than general network or security vulnerability research.
Explore 14 awesome GitHub repositories matching security & cryptography · Adversarial Input Generation. Refine with filters or upvote what's useful.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Generates adversarial inputs and edge-case scenarios to perform safety evaluations and brand risk stress-testing on AI models.
Superagent is an AI safety platform that protects applications from prompt injections, data leaks, and harmful outputs through built-in guardrails. It functions as a prompt injection detection system, data redaction tool, and red team testing tool, automatically removing personally identifiable information and protected health data from AI inputs and outputs while scanning image uploads with vision AI to detect visual prompt injection attacks before processing. The platform routes every prompt through a sequential pipeline of safety checks including injection detection, data redaction, and co
Detects and blocks prompt injection attacks, jailbreaks, and malicious instructions before they reach the language model.
NeMo-Guardrails is a toolkit for adding programmable safety constraints and dialogue boundaries to large language model conversational systems. It functions as security middleware that intercepts inputs and outputs to block prompt injections, jailbreaks, and sensitive data leaks, while providing a conversational dialogue manager to define structured interaction flows through configuration files. The framework includes a hallucination filter to screen model outputs for factual accuracy and a specialized modeling language for defining conversational flows and constraints. It provides capabiliti
Protects models from jailbreak attempts and malicious instructions using input inspection.
Cleverhans 是一个 TensorFlow 对抗性机器学习库,既是攻击框架,也是鲁棒性基准测试和防御库。它提供了一系列工具来生成对抗样本、测试神经网络的安全性,并实施保护机制以提高模型对恶意输入的抵御能力。 该项目专注于创建旨在欺骗机器学习模型并使其做出错误预测的扰动输入。它能够评估深度学习模型在受到对抗性噪声干扰时的稳定性和准确性,并提供已知攻击方法的参考实现以识别安全弱点。 该工具包涵盖了对抗样本生成、机器学习模型防御以及神经网络鲁棒性基准测试。它利用模型无关的接口和可微分的攻击实现来执行基于梯度的扰动和迭代优化循环。
Generates malicious input perturbations using reference methods to deceive machine learning models.
The Adversarial Robustness Toolbox (ART) is an open-source library that provides a unified framework for evaluating, defending, and certifying machine learning models against adversarial threats. It wraps models from any framework behind a common estimator interface, enabling composable pipelines for attack generation, defense application, robustness certification, and privacy auditing across evasion, poisoning, and extraction threats. The library distinguishes itself by covering the full adversarial ML security lifecycle within a single toolkit. It supports gradient-based adversarial example
Analyzes inputs and activations to flag samples crafted to deceive the model.
Helicone is an AI gateway and observability platform designed to intercept, manage, and monitor interactions with large language models. By acting as a reverse-proxy, it provides a centralized layer for routing requests across multiple AI providers, allowing developers to maintain consistent application logic while gaining deep visibility into model performance, usage, and costs. The platform distinguishes itself through a robust suite of traffic management and prompt engineering tools. It enables policy-driven control, including automatic failover between providers, rate limiting, and edge-b
Analyzes user messages to identify jailbreak attempts and malicious instructions across multiple languages and blocks the request.
Giskard 是一个用于大语言模型和 AI 智能体的评估框架、测试库及质量监控系统。它作为量化模型性能和可靠性的工具包,为验证检索增强生成(RAG)流水线提供了专门的功能。 该项目通过自动化的红队测试工具和安全扫描器脱颖而出,旨在识别漏洞、提示词注入和安全风险。它利用对抗性探测和合成边缘案例生成来量化模型的鲁棒性并检测信息泄露。 该平台涵盖了广泛的功能,包括事实准确性与幻觉检测、推理与逻辑基准测试以及偏见检测。它提供了回归测试、RAG 组件评估以及从知识库自动生成测试用例的工具。 该系统包含协作团队工作区管理、基于角色的访问控制以及用于监控性能漂移的定时评估流水线等功能。
Generates synthetic edge cases and adversarial inputs to stress-test model resilience and robustness.
该项目是一个全面的教育资源和技术手册,专注于可解释机器学习和可解释 AI(XAI)。它作为一本教科书和参考资料,用于实现使复杂的机器学习模型对人类透明且易于理解的技术。 该资源提供了关于构建本质上透明的模型(如决策树和稀疏线性模型)以及将事后解释方法应用于黑盒系统的指导。它详细介绍了量化特征重要性、为单个预测生成理由以及使用代理模型近似复杂决策过程的具体方法。 内容涵盖了广泛的分析功能,包括全局和局部特征影响分析、计算机视觉可解释性以及使用 Shapley 值等博弈论贡献。它还通过可解释性评估、识别模型捷径的调试工作流以及透明算法结构的设计来解决模型评估问题。 该项目以 Jupyter Notebooks 集合的形式实现。
Generates adversarial inputs to stress-test AI models and identify vulnerabilities in their decision logic.
agent-governance-toolkit 是一个用于执行安全策略、管理零信任身份以及沙箱化自主 AI 代理执行的框架。它提供了一个治理层,旨在通过使用安全策略引擎、加密身份管理和运行时执行沙箱来控制代理的行为。 该项目通过多级特权环系统和加密身份网格脱颖而出,该网格保护自主实体之间的通信。它实现了基于衰减的信任评分机制来跟踪实体可靠性,并利用哈希链式、防篡改审计日志来维护可验证的执行历史。 该工具包涵盖了广泛的能力领域,包括防御注入攻击的提示词安全、针对监管标准的自动化合规性映射,以及使用 Saga 模式的自主工作流编排。它还具有用于跟踪健康状况和支出限额的舰队监控,以及用于限制未经授权资源访问的工具执行沙箱。 提供了一个命令行界面,用于执行控制信号、验证治理策略以及管理扩展的安装。
Uses a multi-vector evaluation system to detect and block prompt injection and jailbreak attempts.
This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.
Generates gradient-based adversarial inputs to stress-test AI model safety alignments.
Mimic is a unicode homoglyph generator and text obfuscation tool. It functions as a character substitutor that replaces standard ASCII characters with visually similar Unicode symbols to create text that appears correct to humans but is functionally different. The project is used for source code obfuscation by inserting subtle syntax errors into code to hide intent or break automated analysis. It also serves as a tool for textual adversarial testing to evaluate the resilience of software filters against maliciously crafted input. The utility achieves these results through a mapping system th
Generates maliciously crafted input using Unicode substitutions to test the resilience of software filters.
PyRIT is an AI vulnerability assessment tool and security scanner designed to detect risks in large language model applications. It functions as a generative AI red teaming framework used to simulate adversarial attacks and identify weaknesses in system guardrails. The tool automates AI risk assessment by scanning generative AI components for security vulnerabilities. It utilizes automated testing and analysis to identify security gaps and prevent potential exploits through a consistent, repeatable process. The system incorporates asynchronous model orchestration to compare security postures
Generates adversarial inputs through iterative prompt refinement to bypass safety filters.
Identifies jailbreak attempts and prompt injections in real time to prevent unauthorized model behavior.
LLM Guard is a security firewall and guardrail framework designed to scan and sanitize inputs and outputs for large language models. It functions as a proxy gateway and security layer to block prompt injections, toxicity, and sensitive data leakage while ensuring that model interactions remain compliant with organizational policies. The system distinguishes itself through a modular scanner pipeline that utilizes local model orchestration to eliminate external network dependencies. It supports real-time security filtering via streaming chunk analysis and implements a fail-fast execution model
Detects and blocks prompt injection and jailbreak attempts to prevent malicious hijacking of model behavior.