awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to safeailab/rain

Projects sharing features with RAIN

30 open-source projects similar to safeailab/rain, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • aounon/certified-llm-safetyaounon avatar

    aounon/certified-llm-safety

    53View on GitHub↗

    This repository contains code for the paper Certifying LLM Safety against Adversarial Prompting.

    Python
    View on GitHub↗53
  • arobey1/smooth-llmarobey1 avatar

    arobey1/smooth-llm

    134View on GitHub↗

    This is the official source code for "SmoothLLM: Defending LLMs Against Jailbreaking Attacks" by Alex Robey, Eric Wong, Hamed Hassani, and George J. Pappas. To learn more about our work, see our blog post.

    Python
    View on GitHub↗134
  • chuhac/reasoning-to-defendchuhac avatar

    chuhac/Reasoning-to-Defend

    12View on GitHub↗

    Code for paper

    Python
    View on GitHub↗12
  • crystaleye42/eval-safetyCrystalEye42 avatar

    CrystalEye42/eval-safety

    9View on GitHub↗

    This is a repository for replicating the experiments from our paper: Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning .

    Jupyter Notebook
    View on GitHub↗9
  • damo-nlp-sg/multilingual-safety-for-llmsDAMO-NLP-SG avatar

    DAMO-NLP-SG/multilingual-safety-for-LLMs

    105View on GitHub↗

    📄 Paper • 🤗 Dataset

    View on GitHub↗105

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • devoallen/industDevoAllen avatar

    DevoAllen/INDust

    8View on GitHub↗

    We have reorganized INDust, aligning evidence with three types of inductive instructions and implementing stricter quality control measures.

    Python
    View on GitHub↗8
  • ed-zh/pardenEd-Zh avatar

    Ed-Zh/PARDEN

    12View on GitHub↗
    HTML
    View on GitHub↗12
  • ezelikman/starezelikman avatar

    ezelikman/STaR

    227View on GitHub↗

    1. STaR 2. Mesh Transformer JAX 1. Updates 3. Pretrained Models 1. GPT-J-6B 1. Links 2. Acknowledgments 3. License 4. Model Details 5. Zero-Shot Evaluations 4. Architecture and Usage 1. Fine-tuning 2. JAX Dependency 5. TODO

    Python
    View on GitHub↗227
  • facebookresearch/rlcdfacebookresearch avatar

    facebookresearch/rlcd

    70View on GitHub↗

    This repo contains code and instructions for reproducing the experiments in the paper "RLCD: Reinforcement Learning from Contrast Distillation for Language Model Alignment" (https://arxiv.org/abs/2307.12950), by Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. RLCD is a…

    Python
    View on GitHub↗70
  • guoruic/fjdGuoruiC avatar

    GuoruiC/FJD

    4View on GitHub↗

    2025/09 We have released our code. - 2025/08 Our paper is accepted by EMNLP 2025.

    Python
    View on GitHub↗4
  • ibm/dromedaryIBM avatar

    IBM/Dromedary

    1,138View on GitHub↗

    Dromedary: towards helpful, ethical and reliable LLMs.

    Python
    View on GitHub↗1,138
  • idea-xl/g4dIDEA-XL avatar

    IDEA-XL/G4D

    9View on GitHub↗

    Weidi Luo†, Cao He†, Yu Wang, Zijing Liu, Bin Feng, Yao Yuan, Yu Li

    Python
    View on GitHub↗9
  • jaehunjung1/impossible-distillationjaehunjung1 avatar

    jaehunjung1/impossible-distillation

    18View on GitHub↗

    This repository is the official code repository for our paper Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing.

    Python
    View on GitHub↗18
  • linlt-leon/self-evalLinlt-leon avatar

    Linlt-leon/self-eval

    4View on GitHub↗

    This is the official repository for "Self-Evaluation as a Defense Against Adversarial Attacks on LLMs" by Hannah Brown, Leon Lin, Kenji Kawaguchi, Michael Shieh.

    Python
    View on GitHub↗4
  • lucidrains/self-rewarding-lm-pytorchlucidrains avatar

    lucidrains/self-rewarding-lm-pytorch

    1,410View on GitHub↗

    Implementation of the training framework proposed in Self-Rewarding Language Model , from MetaAI

    Python
    View on GitHub↗1,410
  • matthew-pisano/bergeronmatthew-pisano avatar

    matthew-pisano/Bergeron

    7View on GitHub↗

    The goal of this project is to create a framework that protects models against both natural language adversarial attacks and its own bias toward mis-alignment. This is done through the usage of a secondary model that judges the prompts to and responses from that primary model. This leaves the…

    Python
    View on GitHub↗7
  • njunlp/renellmNJUNLP avatar

    NJUNLP/ReNeLLM

    162View on GitHub↗

    The official implementation of our NAACL 2024 paper "A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily".

    Python
    View on GitHub↗162
  • poloclub/llm-self-defensepoloclub avatar

    poloclub/llm-self-defense

    52View on GitHub↗

    LLM Self Defense: By Self Examination, LLMs know they are being tricked. Mansi Phute, Alec Heibling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, Duen Horng Chau. In ICLR 2024 TinyPaper, 2024.

    Python
    View on GitHub↗52
  • project-baize/baize-chatbotproject-baize avatar

    project-baize/baize-chatbot

    3,156View on GitHub↗

    Let ChatGPT teach your own chatbot in hours with a single GPU!

    Python
    View on GitHub↗3,156
  • rapidresponsebench/rapidresponsebenchrapidresponsebench avatar

    rapidresponsebench/rapidresponsebench

    35View on GitHub↗

    Setting up

    Jupyter Notebook
    View on GitHub↗35
  • spico197/humbackSpico197 avatar

    Spico197/Humback

    138View on GitHub↗

    An unofficial implementation of Self-Alignment with Instruction Backtranslation .

    Python
    View on GitHub↗138
  • thunlp-mt/skrTHUNLP-MT avatar

    THUNLP-MT/SKR

    27View on GitHub↗

    Self-Knowledge Guided Retrieval Augmentation for Large Language Models (EMNLP Findings 2023)

    Python
    View on GitHub↗27
  • uclaml/spinuclaml avatar

    uclaml/SPIN

    1,245View on GitHub↗

    The official implementation of Self-Play Fine-Tuning (SPIN)

    Pythondeep-learningfine-tuninglarge-language-models
    View on GitHub↗1,245
  • ucsb-nlp-chang/semanticsmoothUCSB-NLP-Chang avatar

    UCSB-NLP-Chang/SemanticSmooth

    24View on GitHub↗

    This is the official implementation for the paper Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing.

    Python
    View on GitHub↗24
  • uw-nsl/safedecodinguw-nsl avatar

    uw-nsl/SafeDecoding

    154View on GitHub↗

    This is the official repository for "SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding" (Accepted by ACL 2024).

    Jupyter Notebook
    View on GitHub↗154
  • weiyezhimeng/prefix-guidanceweiyezhimeng avatar

    weiyezhimeng/Prefix-Guidance

    5View on GitHub↗

    This is the official code repository for Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks.

    Python
    View on GitHub↗5
  • xhmy/autodefenseXHMY avatar

    XHMY/AutoDefense

    67View on GitHub↗

    Blog

    Python
    View on GitHub↗67
  • xyq7/gradsafexyq7 avatar

    xyq7/GradSafe

    68View on GitHub↗

    Official Code for ACL 2024 paper "GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis" https://arxiv.org/abs/2402.13494

    Python
    View on GitHub↗68
  • yihanwang617/llm-jailbreaking-defense-backtranslationYihanWang617 avatar

    YihanWang617/LLM-Jailbreaking-Defense-Backtranslation

    35View on GitHub↗

    Defending LLMs against Jailbreaking Attacks via Backtranslation

    Python
    View on GitHub↗35
  • yizhongw/self-instructyizhongw avatar

    yizhongw/self-instruct

    4,602View on GitHub↗

    Self-instruct is a framework for generating synthetic instruction datasets and fine-tuning large language models to improve their instruction-following capabilities. It provides a pipeline for aligning pretrained models with human intentions through a supervised fine-tuning workflow. The system utilizes a synthetic data generator that uses a seed set of tasks to prompt a model to create new instructional data. It includes an instruction dataset curator to remove redundant or low-quality entries, maintaining dataset diversity through a filtered task pool. The framework covers the full alignme

    Python
    View on GitHub↗4,602