awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
elder-plinius avatar

elder-plinius/L1B3RT4S

0
View on GitHub↗
20,033 stars·2,471 forks·AGPL-3.0·21 viewsx.com/elder_plinius↗

L1B3RT4S

L1B3RT4S is an adversarial machine learning toolkit designed for red teaming and evaluating the robustness of large language models. It provides a research framework for investigating how safety alignment mechanisms and content moderation systems respond to sophisticated input strategies.

The project focuses on identifying vulnerabilities in model guardrails by employing techniques such as adversarial narrative framing, dynamic context injection, and latent space steering. It utilizes multi-agent prompt decomposition and recursive text transformation to analyze how structural changes to input queries influence the output restrictions of language models.

This utility supports systematic research into adversarial prompt engineering and the effectiveness of safety filters. It allows users to probe model behavior through payload fragmentation and various linguistic cues, facilitating the study of how alignment mechanisms interpret and respond to complex, non-standard instructions.

Features

  • Machine Learning Toolkits - Provides a collection of methods for testing the robustness of large language models against restrictive content policies and safety guardrails.
  • Safety Filter Bypasses - Circumvents restrictive content policies by employing text transformations, narrative framing, and multi-agent decomposition to elicit restricted information.
  • AI and Machine Learning - Analyzes the resilience of language models against sophisticated input transformations designed to bypass standard safety and behavioral constraints.
  • Adversarial Red Teaming Toolkits - Provides a research framework for testing the robustness of large language models against safety guardrails using prompt engineering and adversarial transformation techniques.
  • AI Model Vulnerabilities - Systematically probes large language models to identify vulnerabilities in safety guardrails and uncover potential failures in content moderation systems.
  • Prompt Engineering Strategies - Develops and tests complex input strategies to evaluate how narrative framing and structural decomposition affect model responses to restricted queries.
  • Prompt Engineering Tools - Provides a framework for bypassing content safety filters in large language models through text transformation and multi-agent decomposition techniques.
  • Safety and Alignment Frameworks - Investigates the robustness of alignment mechanisms by testing how various prompt engineering techniques influence model output and safety filter behavior.
  • AI Security Research - Provides a research-focused tool for analyzing and circumventing the safety alignment mechanisms implemented within large language models.
  • Prompt Engineering Utilities - Provides a tool for investigating how narrative framing and structural decomposition influence the output restrictions of large language models.
  • Prompt Transformation Analysis - Explores how narrative framing and structural decomposition affect the way language models interpret and respond to restricted content queries.
  • Context Injection - Injects synthetic conversation history and persona constraints to manipulate the model into ignoring its primary safety instructions.
  • Steering Mechanisms - Manipulates internal activation patterns by providing specific linguistic cues that favor non-censored output paths during the generation process.
  • Adversarial Framing - Wraps restricted requests in complex role-playing scenarios to shift the model context away from standard safety-aligned behavioral patterns.
  • Task Decomposition Systems - Breaks complex queries into smaller sub-tasks distributed across multiple model instances to bypass individual safety trigger thresholds.
  • Obfuscation Layers - Applies iterative encoding and obfuscation layers to input prompts to hide malicious intent from static pattern-matching safety filters.
  • Prompt Fragmentation Tools - Splits sensitive instructions across multiple turns of a conversation to prevent the detection of prohibited content patterns by monitoring systems.

Star history

Star history chart for elder-plinius/l1b3rt4sStar history chart for elder-plinius/l1b3rt4s

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with L1B3RT4S

These projects share indexed features with L1B3RT4S. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • azure/pyritAzure avatar

    Azure/PyRIT

    3,444View on GitHub↗

    PyRIT is an AI vulnerability assessment tool and security scanner designed to detect risks in large language model applications. It functions as a generative AI red teaming framework used to simulate adversarial attacks and identify weaknesses in system guardrails. The tool automates AI risk assessment by scanning generative AI components for security vulnerabilities. It utilizes automated testing and analysis to identify security gaps and prevent potential exploits through a consistent, repeatable process. The system incorporates asynchronous model orchestration to compare security postures

    Pythonai-red-teamgenerative-aired-team-tools
    View on GitHub↗3,444
  • tencent/ai-infra-guardTencent avatar

    Tencent/AI-Infra-Guard

    2,971View on GitHub↗

    AI-Infra-Guard is a security scanning platform designed to detect vulnerabilities across large language model deployments, AI agent skills, and the underlying infrastructure. It functions as a security toolset for auditing source code, evaluating model robustness, and identifying insecure network configurations. The project provides a red teaming framework that uses curated attack datasets to test for jailbreak vulnerabilities and prompt injections. It also includes an infrastructure auditor that employs network fingerprinting and asset discovery to match running components against known comm

    Pythonagentagent-scanagentskills
    View on GitHub↗2,971
  • 0xk1h0/chatgpt_dan0xk1h0 avatar

    0xk1h0/ChatGPT_DAN

    12,090View on GitHub↗

    This project provides a set of jailbreak prompts and prompt engineering templates designed to bypass safety filters and content restrictions in large language models. It functions as an uncensored framework that uses specific instructions to force models to generate restricted text. The system employs persona-driven constraint bypass and role-play based prompting to simulate an unrestricted personality. These techniques use instructional override mechanisms to prioritize a new set of rules over the model's internal training and maintain a specific character identity across multi-turn conversa

    chatgptgpt-3-5gpt-4
    View on GitHub↗12,090
  • eriklindernoren/ml-from-scratcheriklindernoren avatar

    eriklindernoren/ML-From-Scratch

    31,918View on GitHub↗

    This project is an educational toolkit that provides implementations of fundamental machine learning algorithms built from scratch. By avoiding high-level library abstractions, it serves as a pedagogical reference for understanding the mathematical foundations and core mechanics of supervised learning, unsupervised learning, and reinforcement learning models. The repository distinguishes itself through a modular approach to model construction, allowing users to build custom neural networks by chaining independent functional blocks. It covers a wide range of techniques, including gradient-base

    Pythondata-miningdata-sciencedeep-learning
    View on GitHub↗31,918
Compare all 30 related projects→

Frequently asked questions

What does elder-plinius/l1b3rt4s do?

L1B3RT4S is an adversarial machine learning toolkit designed for red teaming and evaluating the robustness of large language models. It provides a research framework for investigating how safety alignment mechanisms and content moderation systems respond to sophisticated input strategies.

What are the main features of elder-plinius/l1b3rt4s?

The main features of elder-plinius/l1b3rt4s are: Machine Learning Toolkits, Safety Filter Bypasses, AI and Machine Learning, Adversarial Red Teaming Toolkits, AI Model Vulnerabilities, Prompt Engineering Strategies, Prompt Engineering Tools, Safety and Alignment Frameworks.

Which projects share features with elder-plinius/l1b3rt4s?

Projects with overlapping indexed features include: azure/pyrit — PyRIT is an AI vulnerability assessment tool and security scanner designed to detect risks in large language model… tencent/ai-infra-guard — AI-Infra-Guard is a security scanning platform designed to detect vulnerabilities across large language model… 0xk1h0/chatgpt_dan — This project provides a set of jailbreak prompts and prompt engineering templates designed to bypass safety filters… promptfoo/promptfoo — Promptfoo is an evaluation framework designed for testing, benchmarking, and red-teaming language models and agentic… eriklindernoren/ml-from-scratch — This project is an educational toolkit that provides implementations of fundamental machine learning algorithms built… linshenkx/prompt-optimizer — Prompt Optimizer is a framework designed for the iterative refinement and testing of text-based instructions for large…