02 | GRPO (Group Relative Policy Optimization)¶ 约 71 个字 预计阅读时间不到 1 分钟 概述 ¶ GRPO 是 DeepSeek 提出的改进版 PPO,通过组内相对比较替代价值函数(Value Function),消除了对 Critic 模型的依赖,显著降低 RLHF 的训练成本。 参考资料 ¶ DeepSeekMath: Pushing the Limits of Mathematical Reasoning (GRPO 出处 ) GRPO 详解 Was this note helpful? 感谢支持~ 感谢指出,欢迎提 issue 或发邮件反馈~