awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
Β© 2026 Bringes Technology SRLΒ·VAT RO45896025Β·hello@awesome-repositories.com
TsinghuaC3I avatar

TsinghuaC3I/Unify-Post-Training

0
View on GitHub↗
211 starsΒ·11 forksΒ·PythonΒ·MITΒ·8 viewsarxiv.org/pdf/2509.04419β†—

Unify Post Training

πŸ“ Unified Policy Gradient Estimator β€’ ✨ Hybrid Post-Training πŸš€ Getting Started β€’ πŸ“Š Main Results β€’ πŸ’– Acknowledgements β€’ πŸ“¨ Contact β€’ 🎈 Citation

Features

  • Critic-Free Algorithms - Unified post-training frameworks for large language models.

Star history

Star history chart for tsinghuac3i/unify-post-trainingStar history chart for tsinghuac3i/unify-post-training

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English β€” the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Unify Post Training

These projects share indexed features with Unify Post Training. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • deepseek-ai/deepseek-mathdeepseek-ai avatar

    deepseek-ai/DeepSeek-Math

    3,346View on GitHub↗

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Python
    View on GitHub↗3,346
  • liziniu/remaxliziniu avatar

    liziniu/ReMax

    202View on GitHub↗

    ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.

    Python
    View on GitHub↗202
  • mcgill-nlp/vineppoMcGill-NLP avatar

    McGill-NLP/VinePPO

    192View on GitHub↗

    Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation

    Python
    View on GitHub↗192
  • bytedtsinghua-sia/dapoBytedTsinghua-SIA avatar

    BytedTsinghua-SIA/DAPO

    1,831View on GitHub↗

    DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR

    Python
    View on GitHub↗1,831
Compare all 11 related projects→

Frequently asked questions

What does tsinghuac3i/unify-post-training do?

πŸ“ Unified Policy Gradient Estimator β€’ ✨ Hybrid Post-Training πŸš€ Getting Started β€’ πŸ“Š Main Results β€’ πŸ’– Acknowledgements β€’ πŸ“¨ Contact β€’ 🎈 Citation

What are the main features of tsinghuac3i/unify-post-training?

The main features of tsinghuac3i/unify-post-training are: Critic-Free Algorithms.

Which projects share features with tsinghuac3i/unify-post-training?

Projects with overlapping indexed features include: bytedtsinghua-sia/dapo β€” DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR. deepseek-ai/deepseek-math β€” DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. liziniu/remax β€” ReMax is a reinforcement learning method, tailored for reward maximization in RLHF. mcgill-nlp/vineppo β€” Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single… minimax-ai/minimax-m1 β€” MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. modalminds/mm-eureka β€” MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.