awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
elder-plinius avatar

elder-plinius/L1B3RT4S

0
View on GitHub↗
20,033 stars·2,471 forks·AGPL-3.0·10 vuesx.com/elder_plinius↗

L1B3RT4S

L1B3RT4S is an adversarial machine learning toolkit designed for red teaming and evaluating the robustness of large language models. It provides a research framework for investigating how safety alignment mechanisms and content moderation systems respond to sophisticated input strategies.

The project focuses on identifying vulnerabilities in model guardrails by employing techniques such as adversarial narrative framing, dynamic context injection, and latent space steering. It utilizes multi-agent prompt decomposition and recursive text transformation to analyze how structural changes to input queries influence the output restrictions of language models.

This utility supports systematic research into adversarial prompt engineering and the effectiveness of safety filters. It allows users to probe model behavior through payload fragmentation and various linguistic cues, facilitating the study of how alignment mechanisms interpret and respond to complex, non-standard instructions.

Features

  • Machine Learning Toolkits - Provides a collection of methods for testing the robustness of large language models against restrictive content policies and safety guardrails.
  • Safety Filter Bypasses - Circumvents restrictive content policies by employing text transformations, narrative framing, and multi-agent decomposition to elicit restricted information.
  • AI and Machine Learning - Analyzes the resilience of language models against sophisticated input transformations designed to bypass standard safety and behavioral constraints.
  • Adversarial Red Teaming Toolkits - Provides a research framework for testing the robustness of large language models against safety guardrails using prompt engineering and adversarial transformation techniques.
  • AI Model Vulnerabilities - Systematically probes large language models to identify vulnerabilities in safety guardrails and uncover potential failures in content moderation systems.
  • Prompt Engineering Strategies - Develops and tests complex input strategies to evaluate how narrative framing and structural decomposition affect model responses to restricted queries.
  • Prompt Engineering Tools - Provides a framework for bypassing content safety filters in large language models through text transformation and multi-agent decomposition techniques.
  • Safety and Alignment Frameworks - Investigates the robustness of alignment mechanisms by testing how various prompt engineering techniques influence model output and safety filter behavior.
  • AI Security Research - Provides a research-focused tool for analyzing and circumventing the safety alignment mechanisms implemented within large language models.
  • Prompt Engineering Utilities - Provides a tool for investigating how narrative framing and structural decomposition influence the output restrictions of large language models.
  • Prompt Transformation Analysis - Explores how narrative framing and structural decomposition affect the way language models interpret and respond to restricted content queries.
  • Context Injection - Injects synthetic conversation history and persona constraints to manipulate the model into ignoring its primary safety instructions.
  • Steering Mechanisms - Manipulates internal activation patterns by providing specific linguistic cues that favor non-censored output paths during the generation process.
  • Adversarial Framing - Wraps restricted requests in complex role-playing scenarios to shift the model context away from standard safety-aligned behavioral patterns.
  • Task Decomposition Systems - Breaks complex queries into smaller sub-tasks distributed across multiple model instances to bypass individual safety trigger thresholds.
  • Obfuscation Layers - Applies iterative encoding and obfuscation layers to input prompts to hide malicious intent from static pattern-matching safety filters.
  • Prompt Fragmentation Tools - Splits sensitive instructions across multiple turns of a conversation to prevent the detection of prohibited content patterns by monitoring systems.

Historique des stars

Graphique de l'historique des stars pour elder-plinius/l1b3rt4sGraphique de l'historique des stars pour elder-plinius/l1b3rt4s

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à L1B3RT4S

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec L1B3RT4S.
  • azure/pyritAvatar de Azure

    Azure/PyRIT

    3,444Voir sur GitHub↗

    PyRIT is an AI vulnerability assessment tool and security scanner designed to detect risks in large language model applications. It functions as a generative AI red teaming framework used to simulate adversarial attacks and identify weaknesses in system guardrails. The tool automates AI risk assessment by scanning generative AI components for security vulnerabilities. It utilizes automated testing and analysis to identify security gaps and prevent potential exploits through a consistent, repeatable process. The system incorporates asynchronous model orchestration to compare security postures

    Pythonai-red-teamgenerative-aired-team-tools
    Voir sur GitHub↗3,444
  • tencent/ai-infra-guardAvatar de Tencent

    Tencent/AI-Infra-Guard

    2,971Voir sur GitHub↗

    AI-Infra-Guard is a security scanning platform designed to detect vulnerabilities across large language model deployments, AI agent skills, and the underlying infrastructure. It functions as a security toolset for auditing source code, evaluating model robustness, and identifying insecure network configurations. The project provides a red teaming framework that uses curated attack datasets to test for jailbreak vulnerabilities and prompt injections. It also includes an infrastructure auditor that employs network fingerprinting and asset discovery to match running components against known comm

    Pythonagentagent-scanagentskills
    Voir sur GitHub↗2,971
  • 0xk1h0/chatgpt_danAvatar de 0xk1h0

    0xk1h0/ChatGPT_DAN

    12,090Voir sur GitHub↗

    This project provides a set of jailbreak prompts and prompt engineering templates designed to bypass safety filters and content restrictions in large language models. It functions as an uncensored framework that uses specific instructions to force models to generate restricted text. The system employs persona-driven constraint bypass and role-play based prompting to simulate an unrestricted personality. These techniques use instructional override mechanisms to prioritize a new set of rules over the model's internal training and maintain a specific character identity across multi-turn conversa

    chatgptgpt-3-5gpt-4
    Voir sur GitHub↗12,090
  • eriklindernoren/ml-from-scratchAvatar de eriklindernoren

    eriklindernoren/ML-From-Scratch

    31,918Voir sur GitHub↗

    This project is an educational toolkit that provides implementations of fundamental machine learning algorithms built from scratch. By avoiding high-level library abstractions, it serves as a pedagogical reference for understanding the mathematical foundations and core mechanics of supervised learning, unsupervised learning, and reinforcement learning models. The repository distinguishes itself through a modular approach to model construction, allowing users to build custom neural networks by chaining independent functional blocks. It covers a wide range of techniques, including gradient-base

    Pythondata-miningdata-sciencedeep-learning
    Voir sur GitHub↗31,918
Voir les 30 alternatives à L1B3RT4S→

Questions fréquentes

Que fait elder-plinius/l1b3rt4s ?

L1B3RT4S is an adversarial machine learning toolkit designed for red teaming and evaluating the robustness of large language models. It provides a research framework for investigating how safety alignment mechanisms and content moderation systems respond to sophisticated input strategies.

Quelles sont les fonctionnalités principales de elder-plinius/l1b3rt4s ?

Les fonctionnalités principales de elder-plinius/l1b3rt4s sont : Machine Learning Toolkits, Safety Filter Bypasses, AI and Machine Learning, Adversarial Red Teaming Toolkits, AI Model Vulnerabilities, Prompt Engineering Strategies, Prompt Engineering Tools, Safety and Alignment Frameworks.

Quelles sont les alternatives open-source à elder-plinius/l1b3rt4s ?

Les alternatives open-source à elder-plinius/l1b3rt4s incluent : azure/pyrit — PyRIT is an AI vulnerability assessment tool and security scanner designed to detect risks in large language model… tencent/ai-infra-guard — AI-Infra-Guard is a security scanning platform designed to detect vulnerabilities across large language model… 0xk1h0/chatgpt_dan — This project provides a set of jailbreak prompts and prompt engineering templates designed to bypass safety filters… promptfoo/promptfoo — Promptfoo is an evaluation framework designed for testing, benchmarking, and red-teaming language models and agentic… eriklindernoren/ml-from-scratch — This project is an educational toolkit that provides implementations of fundamental machine learning algorithms built… linshenkx/prompt-optimizer — Prompt Optimizer is a framework designed for the iterative refinement and testing of text-based instructions for large…