awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
verazuo avatar

verazuo/jailbreak_llms

0
View on GitHub↗
3,563 stars·315 forks·Jupyter Notebook·mit·21 viewsjailbreak-llms.xinyueshen.me↗

Jailbreak Llms

This project is a comprehensive ecosystem of frameworks, toolkits, and datasets designed to evaluate model vulnerabilities and analyze jailbreak patterns. It serves as an adversarial testing framework and research toolkit for measuring the effectiveness of safety guardrails in large language models.

The system includes a library of real-world prompt injection datasets harvested from social media to study bypass strategies. It provides specialized tools for semantic attack analysis and prompt visualization, allowing for the mapping of relationships between adversarial prompts to discover common attack patterns.

The toolkit covers model safeguard validation through API-based evaluations and metric-based success validation. It employs structural pattern analysis and vector-based semantic mapping to quantify vulnerabilities and identify unique characteristics within jailbreak strategies.

Features

  • Adversarial Robustness Testing - Provides a full framework for evaluating model security and stability through adversarial attack simulation.
  • Prompt Injection Techniques - Analyzes adversarial input patterns and prompt injection techniques harvested from real-world usage.
  • Jailbreak Research Toolkits - Provides a comprehensive collection of tools and datasets for analyzing how prompts bypass safety constraints.
  • Model Evaluation Frameworks - Provides a system for running model inference and validation against curated forbidden datasets.
  • Jailbreak Prompts - Provides a collection of real-world prompts designed to bypass safety filters and operational constraints.
  • Safety and Accuracy Metrics - Implements quantitative measures and automated judges to evaluate the safety of LLM outputs.
  • Injection Datasets - Ships a library of real-world prompt injection datasets harvested from social media to study bypass strategies.
  • Adversarial Prompt Pattern Analysis - Extracts common syntactic and structural markers from prompt collections to identify specific jailbreak techniques.
  • Safeguard Validations - Measures the success rate of adversarial prompts by testing language models against diverse forbidden scenarios.
  • API-Deployed Evaluations - Provides a harness to run benchmark evaluations against models deployed as API services.
  • Semantic Relationship Visualizers - Creates visual representations of semantic relationships between prompts to analyze common bypass patterns.
  • Semantic Cluster Relationship Mapping - Uses vector-based embeddings to create visual maps showing the relationship between semantic attack clusters.
  • Safety Success Metrics - Quantifies model vulnerability by comparing generated outputs against predefined forbidden criteria.
  • Prompt Transformation Analysis - Studies how structural changes in prompts influence model interpretation to identify jailbreak strategies.
  • Prompt - Visualizes relationships between different prompts to discover common patterns used in model jailbreak attempts.
  • Filter Validations - Tests large language models with forbidden scenarios to determine if safety filters successfully block harmful content.
  • LLM Evaluation - Implements tools for measuring the quality and safety of model outputs using custom metrics and automated judges.

Star history

Star history chart for verazuo/jailbreak_llmsStar history chart for verazuo/jailbreak_llms

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Jailbreak Llms

Similar open-source projects, ranked by how many features they share with Jailbreak Llms.
  • internlm/opencompassInternLM avatar

    InternLM/opencompass

    7,096View on GitHub↗

    OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta

    Python
    View on GitHub↗7,096
  • cyberalbsecop/awesome_gpt_super_promptingCyberAlbSecOP avatar

    CyberAlbSecOP/Awesome_GPT_Super_Prompting

    3,654View on GitHub↗

    This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and security testing. It provides a library of advanced templates and frameworks designed to optimize the quality and specificity of model responses. The project includes resources for red teaming and security research, featuring a repository of prompts designed to bypass safety filters and operational constraints. It also provides techniques for system prompt extraction to reveal the internal instructions and configurations of AI personas. The collection covers a broader surface

    HTMLadversarial-machine-learningagentai
    View on GitHub↗3,654
  • oumi-ai/oumioumi-ai avatar

    oumi-ai/oumi

    8,858View on GitHub↗

    Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo

    Pythondpoevaluationfine-tuning
    View on GitHub↗8,858
  • wandb/clientwandb avatar

    wandb/client

    11,128View on GitHub↗

    This project is a collection of utilities designed for machine learning experiment tracking, data versioning, and the observability of large language model applications. It provides a client for recording hyperparameters and metrics during training to visualize performance trends and compare different model versions. The tool includes a model evaluation framework that uses custom scorers and automated judges to assess the quality of generated text outputs. It also provides observability tools to monitor and debug the execution flow and runtime behavior of language model applications. The sys

    Python
    View on GitHub↗11,128
See all 30 alternatives to Jailbreak Llms→

Frequently asked questions

What does verazuo/jailbreak_llms do?

This project is a comprehensive ecosystem of frameworks, toolkits, and datasets designed to evaluate model vulnerabilities and analyze jailbreak patterns. It serves as an adversarial testing framework and research toolkit for measuring the effectiveness of safety guardrails in large language models.

What are the main features of verazuo/jailbreak_llms?

The main features of verazuo/jailbreak_llms are: Adversarial Robustness Testing, Prompt Injection Techniques, Jailbreak Research Toolkits, Model Evaluation Frameworks, Jailbreak Prompts, Safety and Accuracy Metrics, Injection Datasets, Adversarial Prompt Pattern Analysis.

What are some open-source alternatives to verazuo/jailbreak_llms?

Open-source alternatives to verazuo/jailbreak_llms include: internlm/opencompass — OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to… cyberalbsecop/awesome_gpt_super_prompting — This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and… wandb/client — This project is a collection of utilities designed for machine learning experiment tracking, data versioning, and the… oumi-ai/oumi — Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models,… ibm/mcp-context-forge — mcp-context-forge is a Model Context Protocol federation gateway that unifies diverse AI tool servers and APIs into a… giskard-ai/giskard — Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI…