Hacker NewsΒ·5 min read

DAPO: An Open-Source RL System from ByteDance Seed and Tsinghua Air

T
the_arun
DAPO: An Open-Source RL System from ByteDance Seed and Tsinghua Air
✦Dive DeeperCreate a free account to unlock

We release a fully open-sourced system for large-scale LLM RL, including algorithm, code infrastructure, and dataset. The system achieves state-of-the-art large-scale LLM RL performance. We propose the D ecoupled Clip and D ynamic s A mpling P olicy O ptimization ( DAPO ) algorithm. Through open-sourcing, we provide the broader research community and society with practical access to scalable reinforcement learning, enabling all to benefit from these advancements. Our system is based on the awesome verl framework. Thanks for their great work!

πŸ€— If you have any questions about our paper, issues are welcomed and we could discuss there. Thank you!

AIME 2024 Performance πŸš€ DAPO achieves 50 points on AIME 2024 based on the Qwen2.5-32B base model, outperforming the previous SoTA DeepSeek-R1-Zero-Qwen-32B with 50% training steps.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article β†’
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in