awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to linlt-leon/self-eval

Projects sharing features with Self Eval

22 open-source projects similar to linlt-leon/self-eval, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • aounon/certified-llm-safetyaounon avatar

    aounon/certified-llm-safety

    53View on GitHub↗

    This repository contains code for the paper Certifying LLM Safety against Adversarial Prompting.

    Python
    View on GitHub↗53
  • arobey1/smooth-llmarobey1 avatar

    arobey1/smooth-llm

    134View on GitHub↗

    This is the official source code for "SmoothLLM: Defending LLMs Against Jailbreaking Attacks" by Alex Robey, Eric Wong, Hamed Hassani, and George J. Pappas. To learn more about our work, see our blog post.

    Python
    View on GitHub↗134
  • chuhac/reasoning-to-defendchuhac avatar

    chuhac/Reasoning-to-Defend

    12View on GitHub↗

    Code for paper

    Python
    View on GitHub↗12
  • crystaleye42/eval-safetyCrystalEye42 avatar

    CrystalEye42/eval-safety

    9View on GitHub↗

    This is a repository for replicating the experiments from our paper: Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning .

    Jupyter Notebook
    View on GitHub↗9
  • damo-nlp-sg/multilingual-safety-for-llmsDAMO-NLP-SG avatar

    DAMO-NLP-SG/multilingual-safety-for-LLMs

    105View on GitHub↗

    📄 Paper • 🤗 Dataset

    View on GitHub↗105

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • devoallen/industDevoAllen avatar

    DevoAllen/INDust

    8View on GitHub↗

    We have reorganized INDust, aligning evidence with three types of inductive instructions and implementing stricter quality control measures.

    Python
    View on GitHub↗8
  • ed-zh/pardenEd-Zh avatar

    Ed-Zh/PARDEN

    12View on GitHub↗
    HTML
    View on GitHub↗12
  • guoruic/fjdGuoruiC avatar

    GuoruiC/FJD

    4View on GitHub↗

    2025/09 We have released our code. - 2025/08 Our paper is accepted by EMNLP 2025.

    Python
    View on GitHub↗4
  • idea-xl/g4dIDEA-XL avatar

    IDEA-XL/G4D

    9View on GitHub↗

    Weidi Luo†, Cao He†, Yu Wang, Zijing Liu, Bin Feng, Yao Yuan, Yu Li

    Python
    View on GitHub↗9
  • matthew-pisano/bergeronmatthew-pisano avatar

    matthew-pisano/Bergeron

    7View on GitHub↗

    The goal of this project is to create a framework that protects models against both natural language adversarial attacks and its own bias toward mis-alignment. This is done through the usage of a secondary model that judges the prompts to and responses from that primary model. This leaves the…

    Python
    View on GitHub↗7
  • njunlp/renellmNJUNLP avatar

    NJUNLP/ReNeLLM

    162View on GitHub↗

    The official implementation of our NAACL 2024 paper "A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily".

    Python
    View on GitHub↗162
  • poloclub/llm-self-defensepoloclub avatar

    poloclub/llm-self-defense

    52View on GitHub↗

    LLM Self Defense: By Self Examination, LLMs know they are being tricked. Mansi Phute, Alec Heibling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, Duen Horng Chau. In ICLR 2024 TinyPaper, 2024.

    Python
    View on GitHub↗52
  • rapidresponsebench/rapidresponsebenchrapidresponsebench avatar

    rapidresponsebench/rapidresponsebench

    35View on GitHub↗

    Setting up

    Jupyter Notebook
    View on GitHub↗35
  • safeailab/rainSafeAILab avatar

    SafeAILab/RAIN

    98View on GitHub↗

    RAIN is an innovative inference method that, by integrating self-evaluation and rewind mechanisms, enables frozen large language models to directly produce responses consistent with human preferences without requiring additional alignment data or model fine-tuning, thereby offering an effective…

    Python
    View on GitHub↗98
  • ucsb-nlp-chang/semanticsmoothUCSB-NLP-Chang avatar

    UCSB-NLP-Chang/SemanticSmooth

    24View on GitHub↗

    This is the official implementation for the paper Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing.

    Python
    View on GitHub↗24
  • uw-nsl/safedecodinguw-nsl avatar

    uw-nsl/SafeDecoding

    154View on GitHub↗

    This is the official repository for "SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding" (Accepted by ACL 2024).

    Jupyter Notebook
    View on GitHub↗154
  • weiyezhimeng/prefix-guidanceweiyezhimeng avatar

    weiyezhimeng/Prefix-Guidance

    5View on GitHub↗

    This is the official code repository for Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks.

    Python
    View on GitHub↗5
  • xhmy/autodefenseXHMY avatar

    XHMY/AutoDefense

    67View on GitHub↗

    Blog

    Python
    View on GitHub↗67
  • xyq7/gradsafexyq7 avatar

    xyq7/GradSafe

    68View on GitHub↗

    Official Code for ACL 2024 paper "GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis" https://arxiv.org/abs/2402.13494

    Python
    View on GitHub↗68
  • yihanwang617/llm-jailbreaking-defense-backtranslationYihanWang617 avatar

    YihanWang617/LLM-Jailbreaking-Defense-Backtranslation

    35View on GitHub↗

    Defending LLMs against Jailbreaking Attacks via Backtranslation

    Python
    View on GitHub↗35
  • yjw1029/self-reminderyjw1029 avatar

    yjw1029/Self-Reminder

    57View on GitHub↗

    Overview - Repo Contents - System Requirements - Installation Guide - Demo - Results - License

    Python
    View on GitHub↗57
  • zichuan-liu/ib4llmszichuan-liu avatar

    zichuan-liu/IB4LLMs

    26View on GitHub↗

    Arxiv Paper Slides 中文版 Website Page

    Python
    View on GitHub↗26