awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to anthropics/constitutionalharmlessnesspaper

Open-source alternatives to ConstitutionalHarmlessnessPaper

30 open-source projects similar to anthropics/constitutionalharmlessnesspaper, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best ConstitutionalHarmlessnessPaper alternative.

  • ganjinzero/rrhfالصورة الرمزية لـ GanjinZero

    GanjinZero/RRHF

    806عرض على GitHub↗

    Arxiv

    Python
    عرض على GitHub↗806
  • volcengine/verlالصورة الرمزية لـ volcengine

    volcengine/verl

    22,015عرض على GitHub↗

    verl is a distributed training system designed for large language model alignment and reinforcement learning. It provides a framework for executing post-training pipelines, including supervised fine-tuning and reinforcement learning from human feedback, to refine model behavior and agentic capabilities. The system utilizes a hybrid training and inference engine that optimizes memory and communication when switching between model generation and gradient updates. It supports multi-modal reinforcement learning for models processing both image and text data, and implements algorithms such as PPO

    Python
    عرض على GitHub↗22,015
  • alibabaresearch/damo-convaiالصورة الرمزية لـ AlibabaResearch

    AlibabaResearch/DAMO-ConvAI

    1,561عرض على GitHub↗

    DAMO-ConvAI: The official repository which contains the codebase for Alibaba DAMO Conversational AI.

    Pythonconversational-aideep-learningdialog
    عرض على GitHub↗1,561
  • openai/following-instructions-human-feedbackالصورة الرمزية لـ openai

    openai/following-instructions-human-feedback

    1,258عرض على GitHub↗

    Paper linkLINKTOPAPER

    عرض على GitHub↗1,258
  • rucaibox/rlmecR

    RUCAIBox/RLMEC

    0عرض على GitHub↗

    This repo provides the source code & data of our paper: Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint (arXiv 2024)

    عرض على GitHub↗0

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • uclaml/spinالصورة الرمزية لـ uclaml

    uclaml/SPIN

    1,245عرض على GitHub↗

    The official implementation of Self-Play Fine-Tuning (SPIN)

    Pythondeep-learningfine-tuninglarge-language-models
    عرض على GitHub↗1,245
  • openrlhf/openrlhfالصورة الرمزية لـ OpenRLHF

    OpenRLHF/OpenRLHF

    9,675عرض على GitHub↗

    OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project

    Pythonlarge-language-modelsopenai-o1proximal-policy-optimization
    عرض على GitHub↗9,675
  • alexis-jacq/pytorch-dppoA

    alexis-jacq/Pytorch-DPPO

    0عرض على GitHub↗
    عرض على GitHub↗0
  • alessiodm/drl-zhالصورة الرمزية لـ alessiodm

    alessiodm/drl-zh

    2,291عرض على GitHub↗

    Welcome to drlzh.ai: a hands-on deep reinforcement learning course where you build the algorithms, not just read about them.

    Jupyter Notebook
    عرض على GitHub↗2,291
  • ai4finance-foundation/finrlالصورة الرمزية لـ AI4Finance-Foundation

    AI4Finance-Foundation/FinRL

    13,964عرض على GitHub↗

    FinRL is a reinforcement learning framework designed for the development, training, and backtesting of automated trading strategies. It functions as a quantitative finance toolkit that integrates deep learning algorithms with financial market simulations to address complex portfolio management and asset allocation tasks. The platform provides an end-to-end pipeline for transforming raw market data into actionable trading models. The project distinguishes itself through a layered, modular architecture that separates data processing, environment simulation, and agent training. This design allow

    Jupyter Notebookalgorithmic-tradingdeep-reinforcement-learningdrl-algorithms
    عرض على GitHub↗13,964
  • allenai/finegrainedrlhfA

    allenai/FineGrainedRLHF

    0عرض على GitHub↗

    Fine-Grained RLHF

    عرض على GitHub↗0
  • allenai/rl4lmsالصورة الرمزية لـ allenai

    allenai/RL4LMs

    2,390عرض على GitHub↗

    A modular RL library to fine-tune language models to human preferences

    Python
    عرض على GitHub↗2,390
  • amenra/ranxالصورة الرمزية لـ AmenRa

    AmenRa/ranx

    681عرض على GitHub↗

    ⚡️A Blazing-Fast Python Library for Ranking Evaluation, Comparison, and Fusion 🐍

    Python
    عرض على GitHub↗681
  • applieddatasciencepartners/deepreinforcementlearningالصورة الرمزية لـ AppliedDataSciencePartners

    AppliedDataSciencePartners/DeepReinforcementLearning

    2,032عرض على GitHub↗

    A replica of the AlphaZero methodology for deep reinforcement learning in Python

    Jupyter Notebook
    عرض على GitHub↗2,032
  • alihassanijr/discernالصورة الرمزية لـ alihassanijr

    alihassanijr/DISCERN

    0عرض على GitHub↗

    This repository contains the implementation of DISCERN in Python. You can download the manuscript from my website or arXiv.

    Python
    عرض على GitHub↗0
  • awjuliani/deeprl-agentsالصورة الرمزية لـ awjuliani

    awjuliani/DeepRL-Agents

    2,277عرض على GitHub↗

    A set of Deep Reinforcement Learning Agents implemented in Tensorflow.

    Jupyter Notebookreinforcement-learningtensorflow
    عرض على GitHub↗2,277
  • airlab-polimi/mushroomA

    AIRLab-POLIMI/mushroom

    0عرض على GitHub↗
    عرض على GitHub↗0
  • aunum/goldالصورة الرمزية لـ aunum

    aunum/gold

    351عرض على GitHub↗

    Reinforcement Learning in Go

    Go
    عرض على GitHub↗351
  • atgambardella/pytorch-esA

    atgambardella/pytorch-es

    0عرض على GitHub↗
    عرض على GitHub↗0
  • breakend/deepreinforcementlearningthatmattersB

    Breakend/DeepReinforcementLearningThatMatters

    0عرض على GitHub↗
    عرض على GitHub↗0
  • breakend/reproducibilityincontinuouspolicygradientmethodsB

    Breakend/ReproducibilityInContinuousPolicyGradientMethods

    0عرض على GitHub↗
    عرض على GitHub↗0
  • carpedm20/deep-rl-tensorflowالصورة الرمزية لـ carpedm20

    carpedm20/deep-rl-tensorflow

    1,581عرض على GitHub↗

    TensorFlow implementation of Deep Reinforcement Learning papers

    Pythondeep-reinforcement-learningdqntensorflow
    عرض على GitHub↗1,581
  • carperai/trlxالصورة الرمزية لـ carperai

    carperai/trlx

    4,749عرض على GitHub↗

    trlx is a reinforcement learning library and training framework designed to align large language models using human feedback. It serves as a distributed trainer and compute orchestrator for scaling high-parameter models across multiple GPUs and nodes. The project provides tools for reinforcement learning from human feedback and model alignment. It implements reward-model-based optimization and proximal policy optimization to refine model behavior based on goal-oriented rewards or human-labeled datasets. The framework covers distributed training strategies, including model parallelism, parame

    Python
    عرض على GitHub↗4,749
  • catalyst-team/catalyst-rlالصورة الرمزية لـ catalyst-team

    catalyst-team/catalyst-rl

    48عرض على GitHub↗

    Accelerated RL

    Python
    عرض على GitHub↗48
  • chenmientan/rl2الصورة الرمزية لـ ChenmienTan

    ChenmienTan/RL2

    1,293عرض على GitHub↗
    Python
    عرض على GitHub↗1,293
  • coax-dev/coaxالصورة الرمزية لـ coax-dev

    coax-dev/coax

    185عرض على GitHub↗

    |tests| |pypi| |docs| |License|

    Python
    عرض على GitHub↗185
  • contextualai/halosالصورة الرمزية لـ ContextualAI

    ContextualAI/HALOs

    906عرض على GitHub↗

    A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).

    Pythonalignmentdpohalos
    عرض على GitHub↗906
  • coreylynch/async-rlالصورة الرمزية لـ coreylynch

    coreylynch/async-rl

    1,006عرض على GitHub↗

    Tensorflow Keras OpenAI Gym implementation of 1-step Q Learning from "Asynchronous Methods for Deep Reinforcement Learning"

    Python
    عرض على GitHub↗1,006
  • cornell-rl/drpoC

    Cornell-RL/drpo

    0عرض على GitHub↗
    عرض على GitHub↗0
  • alibaba/rollالصورة الرمزية لـ alibaba

    alibaba/ROLL

    2,844عرض على GitHub↗

    ROLL is a distributed reinforcement learning framework and model alignment toolkit designed for large language models. It serves as a scalable training pipeline and GPU cluster manager, providing the infrastructure to align model behavior using reinforcement learning algorithms and preference optimization techniques. The project distinguishes itself through an agentic rollout orchestrator that generates and collects multi-turn interaction trajectories between AI agents and simulated environments. It supports specialized alignment methods including Direct Preference Optimization, reinforcement

    Pythonagenticrlhfrlvr
    عرض على GitHub↗2,844