awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 dépôts

Awesome GitHub RepositoriesReinforcement Learning Alignment

Techniques for aligning generative models using reward-based feedback to improve output quality and instruction adherence.

Distinguishing note: None of the candidates matched; this is a specific training methodology for model alignment.

Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Reinforcement Learning Alignment. Refine with filters or upvote what's useful.

Awesome Reinforcement Learning Alignment GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • fishaudio/fish-speechAvatar de fishaudio

    fishaudio/fish-speech

    24,928Voir sur GitHub↗

    This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a two-stage autoregressive transformer architecture that separates semantic token prediction from acoustic detail reconstruction to balance linguistic accuracy with audio quality. The system is designed to support multilingual output and conversational AI development, enabling the generation of context-aware speech that maintains flow across multiple dialogue turns. The platform distinguishes itself through a production-ready inference server that employs continuous batching to

    Refines speech models using reward-based evaluation of semantic accuracy and acoustic quality.

    Pythonllamatransformertts
    Voir sur GitHub↗24,928
  • verl-project/verlAvatar de verl-project

    verl-project/verl

    22,000Voir sur GitHub↗

    This project is a distributed training infrastructure designed for aligning large language models through reinforcement learning. It functions as an end-to-end engine for complex alignment tasks, including proximal policy optimization, direct preference optimization, and iterative self-play. By providing a unified framework for multi-turn interactions and tool-use scenarios, it enables the development of models capable of reasoning and external environment engagement. The framework distinguishes itself through a decoupled architecture that separates model training from sample generation. This

    Provides a distributed training infrastructure for aligning large language models using reinforcement learning techniques like PPO, GRPO, and online DPO.

    Python
    Voir sur GitHub↗22,000
  • funaudiollm/cosyvoiceAvatar de FunAudioLLM

    FunAudioLLM/CosyVoice

    21,673Voir sur GitHub↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Refines model outputs using reward signals to optimize for naturalness and human-perceived speech quality.

    Pythonaudio-generationcantonesechatbot
    Voir sur GitHub↗21,673
  • tensorflow/magentaAvatar de tensorflow

    tensorflow/magenta

    19,797Voir sur GitHub↗

    Magenta is an AI creative suite and TensorFlow generative art framework used to train and deploy models for the production of artistic media. It functions as a generative music library and a deep learning art generator, providing tools to automate the creation of original musical compositions and visual artwork. The project covers AI music composition and generative visual art through neural art generation and machine learning creativity. It enables the training of generative models to produce original songs, images, and drawings based on learned patterns.

    Optimizes generative outputs through reward-based feedback to align compositions with human aesthetic preferences.

    Python
    Voir sur GitHub↗19,797
  • huggingface/trlAvatar de huggingface

    huggingface/trl

    18,653Voir sur GitHub↗

    This library provides a comprehensive framework for fine-tuning, aligning, and distilling transformer-based language models. It serves as a toolkit for adapting models to specialized domains through supervised learning, while offering advanced methodologies to improve output quality and reasoning capabilities. The project distinguishes itself through specialized alignment and optimization techniques, including direct preference optimization and reinforcement learning, which allow models to be tuned against human preferences without complex reward modeling. It further supports training efficie

    Aligns language models with human preferences using reinforcement learning techniques to improve output quality and safety.

    Python
    Voir sur GitHub↗18,653
  • alibaba-nlp/deepresearchAvatar de Alibaba-NLP

    Alibaba-NLP/DeepResearch

    18,251Voir sur GitHub↗

    DeepResearch is an autonomous research agent framework designed to orchestrate multi-step information gathering and complex reasoning tasks. The platform functions as an agent orchestration system that manages the entire lifecycle of autonomous research, from initial planning and web navigation to the synthesis of evidence-backed reports. The framework distinguishes itself through a specialized training pipeline that supports the development and fine-tuning of autonomous models using reinforcement learning and structured knowledge graph synthesis. By employing parallel agent coordination, the

    Refines agent decision-making through reinforcement learning feedback loops aligned with research goals.

    Pythonagentalibabaartificial-intelligence
    Voir sur GitHub↗18,251
  • xiaolincoder/cs-baseAvatar de xiaolincoder

    xiaolincoder/CS-Base

    18,024Voir sur GitHub↗

    CS-Base is a comprehensive educational platform and technical repository designed to support software engineers in mastering backend architecture, artificial intelligence engineering, and career development. It functions as a centralized knowledge hub that combines illustrated theoretical tutorials with practical, project-based learning to bridge the gap between foundational computer science concepts and professional industry requirements. The project distinguishes itself by integrating a robust career mentorship framework with advanced AI engineering resources. It provides users with tools f

    Implements reinforcement learning alignment algorithms to optimize model performance.

    ccppgolang
    Voir sur GitHub↗18,024
  • nvidia-nemo/nemoAvatar de NVIDIA-NeMo

    NVIDIA-NeMo/NeMo

    17,389Voir sur GitHub↗

    NeMo is a comprehensive framework designed for the development, training, and deployment of large-scale conversational and generative artificial intelligence models. It provides an integrated platform for building multimodal systems, encompassing speech processing, language modeling, and reinforcement learning alignment. The framework is built to handle the entire lifecycle of AI development, from data curation and model pretraining to production-ready service deployment. The platform distinguishes itself through advanced distributed training capabilities, including tensor and pipeline parall

    Refines model behavior using reinforcement learning and post-training techniques to improve output quality and safety.

    Pythonasrdeeplearninggenerative-ai
    Voir sur GitHub↗17,389
  • modelscope/ms-swiftAvatar de modelscope

    modelscope/ms-swift

    14,597Voir sur GitHub↗

    This project is a comprehensive toolkit designed for the full lifecycle management of large language and multimodal models. It functions as a unified orchestrator that handles the entire development process, ranging from dataset preparation and supervised fine-tuning to advanced reinforcement learning alignment and production-ready inference deployment. The platform distinguishes itself through a specialized reinforcement learning library that supports complex optimization algorithms, including group relative policy optimization and leave-one-out techniques, to improve model instruction-follo

    Optimizes model behavior through iterative policy updates and reward-based feedback loops to improve instruction following and safety.

    Pythondeepseek-r1embeddinggrpo
    Voir sur GitHub↗14,597
  • axolotl-ai-cloud/axolotlAvatar de axolotl-ai-cloud

    axolotl-ai-cloud/axolotl

    12,059Voir sur GitHub↗

    Axolotl is a configuration-driven framework designed for the fine-tuning, evaluation, and quantization of large language models. It functions as a comprehensive orchestrator for distributed training, enabling users to manage complex workflows across multi-node and multi-GPU environments. By utilizing structured configuration files, the platform streamlines the setup of training parameters, dataset paths, and hardware distribution strategies. The project distinguishes itself through its support for diverse training methodologies, including full-parameter tuning, parameter-efficient adaptation,

    Optimizes language models against reward signals and human preferences through policy optimization and iterative generation pipelines.

    Pythonfine-tuningllm
    Voir sur GitHub↗12,059
  • datawhalechina/so-large-lmAvatar de datawhalechina

    datawhalechina/so-large-lm

    7,400Voir sur GitHub↗

    This project is a comprehensive educational curriculum and structured learning path covering the full lifecycle of large language models. It provides a guided progression through the theory, architecture, training, and deployment of these models. The curriculum includes specialized guides on transformer architecture, model training tutorials, and frameworks for designing autonomous agents. It also provides dedicated resources for studying model safety and ethics. The material covers a wide range of technical capabilities, including distributed training strategies, parameter-efficient fine-tu

    Details reinforcement learning alignment techniques to reduce toxicity and improve human value adherence.

    Voir sur GitHub↗7,400
  • ymcui/chinese-llama-alpaca-2Avatar de ymcui

    ymcui/Chinese-LLaMA-Alpaca-2

    7,136Voir sur GitHub↗

    This project provides a Chinese large language model based on the LLaMA architecture. It is an instruction-tuned model optimized for natural language processing and multi-turn conversations in Chinese. The system includes a framework for parameter-efficient fine-tuning using low-rank adaptation and quantization to reduce memory requirements. It also implements retrieval augmented generation for local document question answering and supports long-context processing for sequences up to 64K tokens. The project covers a broad set of capabilities including supervised instruction tuning, reinforce

    Optimizes model outputs using RLHF and reward models to align responses with safety guidelines.

    Python64kalpacaalpaca-2
    Voir sur GitHub↗7,136
  • hiyouga/chatglm-efficient-tuningAvatar de hiyouga

    hiyouga/ChatGLM-Efficient-Tuning

    3,720Voir sur GitHub↗

    ChatGLM-Efficient-Tuning is a fine-tuning framework and toolkit designed to optimize large language models using parameter-efficient fine-tuning techniques. It provides a pipeline for adjusting model behavior and reducing the memory and compute requirements necessary for training. The project features a web-based trainer and orchestration interface for configuring and executing the fine-tuning process on a single GPU. It supports quantized training in lower precision formats to enable fine-tuning on hardware with limited memory, as well as reinforcement learning from human feedback for model

    Implements reinforcement learning alignment to adjust model behavior according to human safety and quality standards.

    Pythonalpacachatglmchatglm2
    Voir sur GitHub↗3,720
  • harderthenharder/transformers_tasksAvatar de HarderThenHarder

    HarderThenHarder/transformers_tasks

    2,420Voir sur GitHub↗

    Transformers Tasks is a collection of toolkits and scripts dedicated to language model fine-tuning, natural language processing tasks, and transformer-based pipelines. The project functions as a natural language processing toolkit and transformer pipeline library, providing Python scripts and algorithms designed to adapt foundational language models and route text inputs through modular processing workflows. The repository covers supervised fine-tuning pipelines and reinforcement learning alignment procedures that optimize generative text outputs through reward modeling and policy gradient lo

    Optimizes generative text outputs through reward modeling and policy gradient reinforcement learning loops.

    Jupyter Notebookinformation-extractionnlpreinforcement-learning
    Voir sur GitHub↗2,420
  1. Home
  2. Artificial Intelligence & ML
  3. Reinforcement Learning Alignment