24 open-source projects similar to zwhe99/deepmath, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best DeepMath alternative.
π Reinforcement Learning for Language Agentsπ
ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Community | Paper | Cookbook | Datasets | Loong Blog | Contributing | CAMEL-AI
Z1: Efficient Test-time Scaling with Code Train Large Language Model to Reason with Shifted Thinking
This repository includes implementations to reproduce the R1 pipeline for code generation:
Open-r1 is a framework designed for the large-scale training, distillation, and optimization of language models focused on complex reasoning and programming tasks. It provides a comprehensive suite of tools for managing distributed training jobs across multi-node clusters, enabling the development of high-performance models through reinforcement learning and supervised fine-tuning. The project distinguishes itself by integrating secure, containerized code execution environments directly into the training and evaluation lifecycle. By allowing models to run and verify code snippets against test
SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
This repository contains the official code and resources for the paper "Synthesizing Sheet Music Problems for Evaluation and Reinforcement Learning".
Training performance of MiroMind-M1-RL-7B on AIME24 and AIME25.
LeetCodeDataset is a dataset comprising Python LeetCode problems designed for training and evaluating Large Language Models (LLMs).
Nemo-Skills is a collection of pipelines to improve "skills" of large language models (LLMs). We support everything needed for LLM development, from synthetic data generation, to model training, to evaluation on a wide range of benchmarks. Start developing on a local workstation and move to aβ¦
An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
COLMβ25 DeepRetrieval β π₯ Training Search Agent by RLVR with Retrieval Outcome
Search-R1 is a distributed training system and reinforcement learning framework designed to create search-augmented language models. It provides an architecture for scaling model workloads across head and worker nodes while optimizing how models interleave internal reasoning with external tool calls. The system focuses on refining model behavior through custom reward signals and reinforcement learning to improve tool-use formatting and information retrieval. It implements an interleaved reasoning-search loop that allows models to alternate between internal thought generation and external data
β Unleashing the Power of Reinforcement Learning for Math and Code Reasoners π€
Training Software Engineering Agents and Verifiers with SWE-Gym