How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Open-r1 is a framework designed for the large-scale training, distillation, and optimization of language models focused on complex reasoning and programming tasks. It provides a comprehensive suite of tools for managing distributed training jobs across multi-node clusters, enabling the development of high-performance models through reinforcement learning and supervised fine-tuning. The project distinguishes itself by integrating secure, containerized code execution environments directly into the training and evaluation lifecycle. By allowing models to run and verify code snippets against test
π Reinforcement Learning for Language Agentsπ
β Unleashing the Power of Reinforcement Learning for Math and Code Reasoners π€
The main features of skyworkai/skywork-or1 are: Code and Formal Reasoning, Frontier Reasoning Models, Reasoning Datasets, Reasoning Models, Regularization Objectives.
Open-source alternatives to skyworkai/skywork-or1 include: huggingface/open-r1 β Open-r1 is a framework designed for the large-scale training, distillation, and optimization of language modelsβ¦ open-reasoner-zero/open-reasoner-zero β An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model. ganler/code-r1 β This repository includes implementations to reproduce the R1 pipeline for code generation:. agentica-project/rllm β π Reinforcement Learning for Language Agentsπ. deepseek-ai/deepseek-r1 β DeepSeek-R1 is an open-weights large language model focused on advanced reasoning. It uses chain-of-thought processingβ¦ gair-nlp/limo β π Paper | π Dataset (v2) | π Model (v2).