awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to pku-alignment/safe-rlhf

Projects sharing features with Safe Rlhf

30 open-source projects similar to pku-alignment/safe-rlhf, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • cyberalbsecop/awesome_gpt_super_promptingCyberAlbSecOP avatar

    CyberAlbSecOP/Awesome_GPT_Super_Prompting

    3,654View on GitHub↗

    This repository is a collection of specialized toolsets and libraries for large language model prompt engineering and security testing. It provides a library of advanced templates and frameworks designed to optimize the quality and specificity of model responses. The project includes resources for red teaming and security research, featuring a repository of prompts designed to bypass safety filters and operational constraints. It also provides techniques for system prompt extraction to reveal the internal instructions and configurations of AI personas. The collection covers a broader surface

    HTMLadversarial-machine-learningagentai
    View on GitHub↗3,654
  • allenai/rl4lmsallenai avatar

    allenai/RL4LMs

    2,390View on GitHub↗

    A modular RL library to fine-tune language models to human preferences

    Python
    View on GitHub↗2,390
  • ganjinzero/rrhfGanjinZero avatar

    GanjinZero/RRHF

    806View on GitHub↗

    Arxiv

    Python
    View on GitHub↗806
  • anthropics/constitutionalharmlessnesspaperanthropics avatar

    anthropics/ConstitutionalHarmlessnessPaper

    263View on GitHub↗

    This repository provides supplementary material for our paper Constitutional AI: Harmlessness from AI Feedback.

    View on GitHub↗263

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • carperai/trlxcarperai avatar

    carperai/trlx

    4,749View on GitHub↗

    trlx is a reinforcement learning library and training framework designed to align large language models using human feedback. It serves as a distributed trainer and compute orchestrator for scaling high-parameter models across multiple GPUs and nodes. The project provides tools for reinforcement learning from human feedback and model alignment. It implements reward-model-based optimization and proximal policy optimization to refine model behavior based on goal-oriented rewards or human-labeled datasets. The framework covers distributed training strategies, including model parallelism, parame

    Python
    View on GitHub↗4,749
  • cascip/chatalpacacascip avatar

    cascip/ChatAlpaca

    176View on GitHub↗

    ChatAlpaca is a chat dataset that aims to help researchers develop models for instruction-following in multi-turn conversations. The dataset is an extension of the Stanford Alpaca data, which contains multi-turn instructions and their corresponding responses.

    Python
    View on GitHub↗176
  • chujiezheng/llm-safeguardchujiezheng avatar

    chujiezheng/LLM-Safeguard

    108View on GitHub↗

    Official repository for our ICML 2024 paper "On Prompt-Driven Safeguarding for Large Language Models"

    Python
    View on GitHub↗108
  • cornell-rl/drpoC

    Cornell-RL/drpo

    0View on GitHub↗
    View on GitHub↗0
  • haoxiang-wang/directional-preference-alignmentH

    Haoxiang-Wang/directional-preference-alignment

    0View on GitHub↗
    View on GitHub↗0
  • declare-lab/red-instructdeclare-lab avatar

    declare-lab/red-instruct

    111View on GitHub↗

    Paper | Github | Dataset | Model

    Python
    View on GitHub↗111
  • dunzeng/moreD

    dunzeng/MORE

    0View on GitHub↗
    View on GitHub↗0
  • eit-nlp/accuracyparadox-rlhfE

    EIT-NLP/AccuracyParadox-RLHF

    0View on GitHub↗
    View on GitHub↗0
  • ernie-research/ma-rlhfE

    ernie-research/MA-RLHF

    0View on GitHub↗
    View on GitHub↗0
  • exlaw/dlmaE

    exlaw/DLMA

    0View on GitHub↗
    View on GitHub↗0
  • amadeuszhao/qmllmA

    Amadeuszhao/QMLLM

    0View on GitHub↗
    View on GitHub↗0
  • gururise/alpacadatacleanedgururise avatar

    gururise/AlpacaDataCleaned

    1,602View on GitHub↗

    Alpaca dataset from Stanford, cleaned and curated

    Python
    View on GitHub↗1,602
  • gximinglu/quarkG

    gximinglu/quark

    0View on GitHub↗
    View on GitHub↗0
  • halfrot/alarmH

    halfrot/ALaRM

    0View on GitHub↗
    View on GitHub↗0
  • allenai/finegrainedrlhfA

    allenai/FineGrainedRLHF

    0View on GitHub↗

    Fine-Grained RLHF

    View on GitHub↗0
  • instruction-tuning-with-gpt-4/gpt-4-llmInstruction-Tuning-with-GPT-4 avatar

    Instruction-Tuning-with-GPT-4/GPT-4-LLM

    4,335View on GitHub↗

    This project is an instruction tuning framework and synthetic data generator that uses high-capacity teacher models to produce instruction-following pairs for training smaller student models. It provides datasets and tools for supervised instruction tuning and reinforcement learning from human feedback. The framework specializes in cross-lingual tuning, offering high-quality instruction-following examples in English and Chinese to improve model generalization across different scripts. It includes a reward modeling tool for creating preference datasets and comparative ratings used to train rew

    HTMLalpacachatgptgpt-4
    View on GitHub↗4,335
  • jaearly/mil-for-non-markovian-reward-modellingJ

    JAEarly/MIL-for-Non-Markovian-Reward-Modelling

    0View on GitHub↗
    View on GitHub↗0
  • jayfeather1024/backdoor-enhanced-alignmentJayfeather1024 avatar

    Jayfeather1024/Backdoor-Enhanced-Alignment

    24View on GitHub↗

    This is the official code repository for the paper BackdoorAlign: Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment.

    Python
    View on GitHub↗24
  • jhejna/few-shot-preference-rlJ

    jhejna/few-shot-preference-rl

    0View on GitHub↗
    View on GitHub↗0
  • jhejna/inverse-preference-learningJ

    jhejna/inverse-preference-learning

    0View on GitHub↗
    View on GitHub↗0
  • kwai-yuanqi/mm-rlhfKwai-YuanQi avatar

    Kwai-YuanQi/MM-RLHF

    200View on GitHub↗

    The Next Step Forward in Multimodal LLM Alignment

    Python
    View on GitHub↗200
  • ledllm/ledllmledllm avatar

    ledllm/ledllm

    24View on GitHub↗

    This repository contains the code for the paper "Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing". The focus of this work is to enhance the resilience of large language models (LLMs) against jailbreak attacks through a novel method termed Layer-specific…

    Jupyter Notebook
    View on GitHub↗24
  • lianjiatech/belleLianjiaTech avatar

    LianjiaTech/BELLE

    8,273View on GitHub↗

    BELLE is a specialized implementation of Chinese conversational large language models, encompassing a full instruction tuning framework. It provides a pipeline for training, evaluating, and deploying models optimized for natural language understanding and dialogue tasks in the Chinese language. The project is distinguished by its integrated approach to model refinement, combining the curation of multi-million entry instruction datasets with a distributed training pipeline. This pipeline supports both full fine-tuning and low-rank adaptation to optimize conversational performance. The system

    HTMLbloomchinese-nlpgpt-evaluation
    View on GitHub↗8,273
  • linear95/apoL

    Linear95/APO

    0View on GitHub↗
    View on GitHub↗0
  • llava-rlhf/llava-rlhfllava-rlhf avatar

    llava-rlhf/LLaVA-RLHF

    396View on GitHub↗

    Aligning LMMs with Factually Augmented RLHF

    Python
    View on GitHub↗396
  • alibabaresearch/damo-convaiAlibabaResearch avatar

    AlibabaResearch/DAMO-ConvAI

    1,561View on GitHub↗

    DAMO-ConvAI: The official repository which contains the codebase for Alibaba DAMO Conversational AI.

    Pythonconversational-aideep-learningdialog
    View on GitHub↗1,561