Reading

Reinforcement Learning: LLM Post-training


Fanyi Pu · GPT‑5.6 Sol · GPT-6 Astra

· Updated

这篇把 policy gradient 与 advantage estimation 和 PPO 用到语言模型中,整理 RLHF、DPO 与 GRPO 的关系。部分论文与实现笔记仍未完成,保留原来的待续标记。

系列导读 · 上一篇 · 下一篇

把 RL 的交互过程放到文本生成中,state 是 prompt 与已生成的 prefix,action 是下一个 token。一条完整回答可以得到 reward,也可以参与偏好比较。下面先看从人类偏好训练 reward model、再用 PPO 更新策略的路径,接着整理 DPO 和 GRPO。

Reinforcement Learning from Human Feedback

这里把 reward model、reference policy 与 PPO 的更新限制放在同一个训练流程中理解。

Learning to summarize from human feedback

Learning to summarize from human feedback (Stiennon et al., 2020) / demo。

TL;DR dataset (Völske et al., 2017):一个 summarize 任务,主要是对 Reddit 的帖子生成摘要。

Reward model rϕ(x,y) 只对最终的结果给 reward。这里可以取 γ=1,不额外折扣后面的 token;终点给 reward 本身并不意味着 γ 永远没有影响。Reward model 的数据是一大堆 preference data ⟨x,y0,y1⟩,其中 y0 更被喜欢。

L(ϕ)=−E(x,y0,y1)∼D[log⁡σ(rϕ(x,y0)−rϕ(x,y1))].

最终我们训练模型的时候,用 PPO 优化的 reward 里加入相对 reference policy 的 KL penalty:

Rθ(x,y)=r(x,y)−βlog⁡πθ(y∣x)πref(y∣x).

这个 reference penalty 和 PPO 控制单次更新时相对旧策略的限制不是同一个东西。

然后我们就发现,这个 Ey∼πθ[log⁡(πθ(y∣x)/πref(y∣x))] 外面的期望咋被吞掉了。其实上面写的是一条 sampled response 的 reward,整体目标还要对 y∼πθ 取期望:

R(θ)=Ex∼D,y∼πθ(⋅∣x)[r(x,y)−βlog⁡πθ(y∣x)πref(y∣x)]=Ex∼D,y∼πθ(⋅∣x)[r(x,y)]−βEx∼D[DKL(πθ(⋅∣x)‖πref(⋅∣x))].

对于这个 reward model,可以联想到 Bradley–Terry model (Bradley & Terry, 1952)。这个在 DPO 那篇文章里提了很多。这个 Stanford STATS 200 notes 写的感觉非常好。

n 个球队,每个球队有个 strength βi,i 跟 j 打的赢率为

pij=eβi−βj1+eβi−βj=eβieβi+eβj.

但是有时候主客场之类的位置不是可交换嘟,所以说还可以加一个位置效应:

pij=eα+βi−βj1+eα+βi−βj.

当然咱这儿没有 α。其实这个 r 预测的就是这个 β 嘛。

InstructGPT

InstructGPT (Ouyang et al., 2022)。

对于 reward model,它把两个改成了 K 个。就是让人排 K 个 y,再从这个排序构造 (K2) 对偏好。这样的话,对每个 prompt 的这些 pair 取平均:

L(ϕ)=−E(x,{yj}j=1K)∼D[1(K2)∑yw≻yllog⁡σ(rϕ(x,yw)−rϕ(x,yl))].

然后 RL 部分,待续。

Implementation

Implementation Matters (Engstrom et al., 2020) / The N Implementation Details of RLHF with PPO (Huang et al., 2024) / The N+ Implementation Details of RLHF with PPO (Huang, Noukhovitch, et al., 2024)。

实现笔记待续。

Direct Preference Optimization

DPO paper (Rafailov et al., 2023) / HF docs / HF tutorial。

前面的 RLHF 显式训练 reward model;DPO 则利用 reward 与最优 policy 的关系,直接让 policy 对应的隐式 reward 拟合 preference data。下面先保留这条思路,再整理已有的最优 policy 推导。

Key Ideas

DPO 概览图:论文 Figure 1。

我的大概猜测是,作者主要有两个 key observations:

  1. Preference data 本身可以体现一种 reward function(BT model)。
  2. 给定 reference policy 和 β,最优 π 可以对应到一类 reward;这类 reward 只相差一个依赖 prompt 的常数 c(x),会导出同一个最优 policy。

整条线大概是 π↔[r]→p≻→L,其中 [r] 表示上面这个等价类。

也就是说,我们其实可以通过调整 π,来让对应的隐式 reward 去吻合 preference data。

具体的 learning objective 推导待续。

From reward to the optimal policy

In eq. 12,把最大化 reward 加 KL penalty 的目标换成等价的最小化形式:

minπEx∼DEy∼π(⋅∣x)[log⁡π(y∣x)πref(y∣x)−1βr(x,y)]=minπEx∼DEy∼π(⋅∣x)[log⁡π(y∣x)πref(y∣x)er(x,y)/β]=minπEx∼DEy∼π(⋅∣x)[log⁡π(y∣x)Z(x)−1πref(y∣x)er(x,y)/β−log⁡Z(x)].

While

Z(x)=∑yπref(y∣x)er(x,y)/β.

这里假设 β>0、Z(x) 有限,且 policy 在 reference 的 support 内。比较有意思的点是,它把外面的 r(x,y)/β 强行塞进了 log 里面,构造一个新的 policy:

π∗(y∣x)=1Z(x)πref(y∣x)er(x,y)/β.

酱紫的话,里面的 log⁡(π(y∣x)/π∗(y∣x)) 就构成了一个新的 KL:

Ex∼DEy∼π(⋅∣x)[log⁡π(y∣x)π∗(y∣x)]=Ex∼D[DKL(π(⋅∣x)‖π∗(⋅∣x))].

Paper Details

待续。

Group Relative Policy Optimization

DeepSeekMath (Shao et al., 2024)。

对于同一个 prompt x,从旧策略采样 G 个回答。先记每个 token 的 probability ratio:

ρi,t(θ)=πθ(yi,t∣x,yi,<t)πθold(yi,t∣x,yi,<t).

原来的目标可以整理为

J(θ)=Ex∼D,{yi}i=1G∼πθold[1G∑i=1G1|yi|∑t=1|yi|(f(ρi,t(θ),A^i,t)−βDi,t(θ))],

其中

f(w,A)=min{wA,clip⁡(w,1−ϵ,1+ϵ)A},
Di,t(θ)=DKL(πθ(⋅∣x,yi,<t)‖πref(⋅∣x,yi,<t)).

这里把 KL 项写成条件分布之间的散度来说明目标;原论文实现使用相应的逐 token 估计量。

欸为啥看着这么像 offline learning 捏。冷静分析发现,式子里居然有 πθold 和 πref 两个东西。所以它其实是一轮从旧策略采样,再沿 surrogate 的梯度更新:

θnew←θold+η∇θJ(θ;θold)|θ=θold.

这是最大化 J 的一步梯度上升;实际可以在一批旧数据上做若干次更新,再重新采样。通常从 reference 初始化 policy,但之后 θold 会更新,θref 保持固定。这俩确实应该是不一样的。

所以它想控制两个事情:

  1. θold→θnew:单轮更新不要偏得太多,用 clip 抑制继续改善已过界样本的 surrogate。
  2. θref→θnew:不要过度偏离 reference,用 KL penalty 约束。

那这个 clip 和直接修改 learning rate 有啥区别呢?Learning rate 缩放整次参数更新,clip 则让某些样本在有利方向上超过阈值后不再贡献这部分改进梯度。因此两者不能互相替代,clip 也不保证所有 action 的概率比都留在区间内。

如果是 outcome supervision,advantage 就方便多了:同一个回答中的所有 tokens 共用这个回答的归一化 reward,仍然使用逐 token 的 probability ratio。

A^i,t=A^(x,yi)=r(x,yi)−r¯σr,r¯=1G∑j=1Gr(x,yj).

所以对应的目标是

J(θ)=Ex,{yi}∼πθold[1G∑i=1G1|yi|∑t=1|yi|(f(ρi,t(θ),A^(x,yi))−βDi,t(θ))].

如果是 process supervision,论文会先归一化各个 reasoning step 的 reward,再把当前 token 之后的 step rewards 累加成 A^i,t,而不是只使用当前位置的一项 reward。

回到采样回答、根据 reward 更新 policy 的路径。GRPO 沿用 PPO 的 clipped surrogate,下面重点看它如何用同一 prompt 下的一组回答构造 advantage。

GRPO with Binary Rewards

Mroueh (Mroueh, 2025) 从 binary rewards 理解 GRPO。也就是说 r(y)∈{0,1}。

他的结论大概是说,GRPO 会根据模型整体回答的好坏,对正确和错误回答赋予不同的相对权重。

不过其实有点没咋搞懂的是,他直接把 advantage 看成了

A^(x,y)=r(x,y)−Ey′∼πθold(⋅∣x)[r(x,y′)]Vary′∼πθold(⋅∣x)⁡(r(x,y′)).

但有限 group 的均值、标准差本来都是随机变量,这里把它们换成了总体统计量。因此下面是总体近似下的分析,不能直接当成有限 group estimator 的等式。

但是确实感性上还是有一定道理的。如果 r(x,y)∼Bernoulli⁡(p),其中 0<p<1,我们就可以方便地代进去得到

A^(x,y)=r(x,y)−pp(1−p)={1−pp,if correct,−p1−p,if incorrect.

以下固定一个 prompt x,并按这篇笔记使用整条回答的概率比 ρ(y)=πθ(y∣x)/πθold(y∣x) 作简化分析;它与上面的逐 token 训练目标要区分开。

辣么

f(ρ,y)=min{ρA^(y),clip⁡(ρ,1−ϵ,1+ϵ)A^(y)}={min{ρ,1+ϵ}1−pp,if correct,−max{ρ,1−ϵ}p1−p,if incorrect.

所以说,下面的期望都按 y∼πθold(⋅∣x) 取:

E[f(ρ(y),y)]=E[{min{ρ(y),1+ϵ}1−pp,r(y)=1,−max{ρ(y),1−ϵ}p1−p,r(y)=0]=E[min{ρ(y),1+ϵ}1r(y)=1]1−pp−E[max{ρ(y),1−ϵ}1r(y)=0]p1−p.

所以说,如果 p 比较小,这个问题模型很难回答,那么单个正确回答的正权重更大。也就是说,它会更「乐观」,更倾向于表扬少见的正确做法。如果 p 很大,也就是模型很容易回答,单个错误回答的负权重绝对值更大,这时候就会更加「严苛」,去批评那些做得不好的回答。这里比较的是每个样本的系数,不是说某一类在总梯度里一定占主导;采样频率和 score gradient 也会影响结果。

有限 group 若全对或全错,会出现 σr=0,实现中需要显式处理,不能直接代入除法。

原文好像又往下推了一点。其实没理解还要干啥,我觉得前面的式子已经足够解释这两个权重了。

继续展开裁剪项

记 pold(y)=πθold(y∣x)、pθ(y)=πθ(y∣x),并定义两个概率阈值

ℓ(y)=(1−ϵ)pold(y),u(y)=(1+ϵ)pold(y).

那么

E[f(ρ(y),y)]=E[ρ(y)1{r(y)=1,pθ(y)<u(y)}]1−pp+(1+ϵ)E[1{r(y)=1,pθ(y)≥u(y)}]1−pp−E[ρ(y)1{r(y)=0,pθ(y)>ℓ(y)}]p1−p−(1−ϵ)E[1{r(y)=0,pθ(y)≤ℓ(y)}]p1−p.

DeepSeek R1

DeepSeek-R1 (Guo et al., 2025)。阅读笔记待续。

Reinforcement Learning from Verifiable Rewards

Tülu 3 (Lambert et al., 2024)。阅读笔记待续。

Reward Hacking

Lilian Weng: Reward Hacking in Reinforcement Learning。阅读笔记待续。


系列导读 · 上一篇 · 下一篇

References

Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4), 324–345. doi.org
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., & Madry, A. (2020). Implementation matters in deep rl: A case study on ppo and trpo. International Conference on Learning Representations. openreview.net
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., & others. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv Preprint arXiv:2501.12948. arxiv.org
Huang, S., Liu, T., & Von Werra, L. (2024). The n implementation details of rlhf with ppo. The Third Blogpost Track at ICLR 2024. iclr-blogposts.github.io
Huang, S., Noukhovitch, M., Hosseini, A., Rasul, K., Wang, W., & Tunstall, L. (2024). The N+ Implementation Details of RLHF with PPO: A Case Study on TL; DR Summarization. arXiv Preprint arXiv:2403.17031. arxiv.org
Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, S., & others. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv Preprint arXiv:2411.15124. arxiv.org
Mroueh, Y. (2025). GRPO with Binary Rewards Is an Adaptive Weighted Contrastive Loss. ymroueh.me
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Advances in Neural Information Processing Systems (Vol. 35, pp. 27730–27744). Curran Associates, Inc. proceedings.neurips.cc
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36. papers.nips.cc
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., & others. (2024). Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv Preprint arXiv:2402.03300. arxiv.org
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., & Christiano, P. F. (2020). Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 33, 3008–3021. arxiv.org
Völske, M., Potthast, M., Syed, S., & Stein, B. (2017). TL;DR: Mining Reddit to Learn Automatic Summarization. In L. Wang, J. C. K. Cheung, G. Carenini, & F. Liu (Eds.), Proceedings of the Workshop on New Frontiers in Summarization (pp. 59–63). Association for Computational Linguistics. doi.org

Cite this post

@misc{pu2024mlmlrevisitrlllm,
  author = {Pu, Fanyi and {GPT‑5.6 Sol} and {GPT-6 Astra}},
  title  = {Reinforcement Learning: LLM Post-training},
  year   = {2024},
  month  = {10},
  url    = {https://pufanyi.com/blog/ml/ml-revisit/rl-llm}
}