# Reinforcement Learning: 从旧策略数据到 PPO

Authors: Fanyi Pu, GPT‑5.6 Sol, GPT-6 Astra

Published: 2024-10-16

Updated: 2026-09-29

Canonical: <https://pufanyi.com/blog/ml/ml-revisit/rl-policy-updates>

Reinforcement Learning: 从旧策略数据到 PPO：RL 系列第 2 篇。

[上一篇](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-gradient) 从采样轨迹构造了 policy gradient 和 advantage estimate。这篇接着回答：策略已经变化后，旧数据还能怎么用，以及一轮更新应该走多远。沿用 $\pi_\theta(a\mid s)$ 表示当前策略，$\hat{\mathcal A}(s,a)$ 表示求导时固定的 advantage estimate；轨迹记为 $\tau$，总 reward 为 $r(\tau)$。

[系列导读](https://pufanyi.com/blog/ml/ml-revisit/rl) · [上一篇](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-gradient) · [下一篇](https://pufanyi.com/blog/ml/ml-revisit/rl-llm)

有了 advantage estimate，还要决定如何使用一批轨迹更新策略。先看旧数据与当前策略之间的分布差异，再看 TRPO 和 PPO 如何限制单轮更新的幅度。

## Off-Policy Policy Gradients

[上一篇的梯度估计](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-gradient) 使用 on-policy 数据。但真实在训练的时候，我们很难做到稍稍改一点 $\theta$，就重新生成一堆新的 $\tau$。这样是非常 inefficient 的。

所以现在可能的问题是，我们没有关于 $\tau\sim\pi_\theta$ 的数据，但是我们可能有一个其他的 distribution，通过这个 distribution 来 sample 出的数据。也就是 $\tau\sim\overline{\pi}$。

我们需要使用的一个 trick 叫做 importance sampling：

$$
\begin{aligned}
\mathbb{E}_{x\sim p(x)}[f(x)]
&= \int p(x)f(x)\,\mathrm{d}x \\
&= \int q(x)\frac{p(x)}{q(x)}f(x)\,\mathrm{d}x \\
&= \mathbb{E}_{x\sim q(x)}\left[\frac{p(x)}{q(x)}f(x)\right]
\end{aligned}
$$

所以说我们的 RL objective 可以改成：

$$
\begin{aligned}
\mathcal{J}(\theta)
&= \mathbb{E}_{\tau\sim\overline{\pi}}\left[
\frac{\pi_\theta(\tau)}{\overline{\pi}(\tau)}r(\tau)
\right] \\
&= \mathbb{E}_{\tau\sim\overline{\pi}}\left[
\frac{p(s_1)\prod_{t=1}^T\pi_\theta(a_t\mid s_t)p(s_{t+1}\mid s_t,a_t)}
{p(s_1)\prod_{t=1}^T\overline{\pi}(a_t\mid s_t)p(s_{t+1}\mid s_t,a_t)}
r(\tau)
\right] \\
&= \mathbb{E}_{\tau\sim\overline{\pi}}\left[
r(\tau)\prod_{t=1}^T\frac{\pi_\theta(a_t\mid s_t)}{\overline{\pi}(a_t\mid s_t)}
\right]
\end{aligned}
$$

所以说我们可以推导梯度

$$
\begin{aligned}
\nabla_\theta\mathcal{J}(\theta)
&= \mathbb{E}_{\tau\sim\overline{\pi}}\left[
\frac{\pi_\theta(\tau)}{\overline{\pi}(\tau)}\nabla_\theta\log\pi_\theta(\tau)r(\tau)
\right] \\
&= \mathbb{E}_{\tau\sim\overline{\pi}}\left[
\left(\prod_{t=1}^T\frac{\pi_\theta(a_t\mid s_t)}{\overline{\pi}(a_t\mid s_t)}\right)
\left(\sum_{t=1}^T\nabla_\theta\log\pi_\theta(a_t\mid s_t)\right)
\left(\sum_{t=1}^T r(s_t,a_t)\right)
\right] \\
&= \mathbb{E}_{\tau\sim\overline{\pi}}\Biggl[
\sum_{t=1}^T\nabla_\theta\log\pi_\theta(a_t\mid s_t)
\left(\prod_{t'=1}^t\frac{\pi_\theta(a_{t'}\mid s_{t'})}{\overline{\pi}(a_{t'}\mid s_{t'})}\right) \\
&\qquad\qquad\cdot\left(
\sum_{t'=t}^T r(s_{t'},a_{t'})
\prod_{t''=t+1}^{t'}\frac{\pi_\theta(a_{t''}\mid s_{t''})}{\overline{\pi}(a_{t''}\mid s_{t''})}
\right)
\Biggr]
\end{aligned}
$$

完整 trajectory ratio 给出了复用数据的理论起点，但长轨迹上的乘积权重可能难以使用。接下来转向旧策略访问的 state 上的 surrogate，并限制新旧策略之间的变化。

## Trust Region Policy Optimization

TRPO ([Schulman et al., 2015](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-updates#bib-schulman2015trust)) / [Spinning Up tutorial](https://spinningup.openai.com/en/latest/algorithms/trpo.html) ([Achiam, 2018](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-updates#bib-spinningup2018))。

$$
\begin{aligned}
\theta_{k+1}&=\arg\max_\theta\mathcal L_{\theta_k}(\theta)\\
\text{s.t.}\quad
\overline D_{\mathrm{KL}}(\theta_k\|\theta)&\le\delta.
\end{aligned}
$$

其中

$$
\mathcal L_{\theta_k}(\theta)
=\mathbb E_{(s,a)\sim\pi_{\theta_k}}\left[
\frac{\pi_\theta(a\mid s)}{\pi_{\theta_k}(a\mid s)}
\hat{\mathcal A}^{\pi_{\theta_k}}(s,a)\right],
$$

$$
\overline D_{\mathrm{KL}}(\theta_k\|\theta)
=\mathbb E_{s\sim\pi_{\theta_k}}\left[
D_{\mathrm{KL}}\!\left(
\pi_{\theta_k}(\cdot\mid s)\|\pi_\theta(\cdot\mid s)
\right)\right].
$$

这里的 state 分布来自旧策略的访问分布；一轮优化中，采样数据、旧策略与 advantage estimates 都固定。

## Proximal Policy Optimization

PPO paper ([Schulman et al., 2017](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-updates#bib-schulman2017proximal)) / [OpenAI tutorial](https://spinningup.openai.com/en/latest/algorithms/ppo.html) ([Achiam, 2018](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-updates#bib-spinningup2018)) / [Hugging Face tutorial](https://huggingface.co/blog/deep-rl-ppo) ([Simonini & Sanseviero, 2023](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-updates#bib-deep-rl-course))。

TRPO 这么复杂，主要是因为它的限制和 $\mathcal L$ 是分开的。文章提出了两种把更新限制放进 surrogate objective 的方法：PPO-Penalty 和 PPO-Clip。

### PPO-Clip

我们对式子略加修改。先记概率比

$$
\rho_\theta(s,a)=\frac{\pi_\theta(a\mid s)}{\pi_{\theta_k}(a\mid s)}.
$$

那么

$$
\theta_{k+1}=\arg\max_\theta
\mathbb E_{(s,a)\sim\pi_{\theta_k}}\left[
\min\left\{
\rho_\theta(s,a)\hat{\mathcal A}(s,a),
\operatorname{clip}(\rho_\theta(s,a),1-\epsilon,1+\epsilon)
\hat{\mathcal A}(s,a)
\right\}\right].
$$

其中

$$
\operatorname{clip}(x,l,u)=
\begin{cases}
l,&x<l,\\
x,&l\le x\le u,\\
u,&x>u.
\end{cases}
$$

大概感性理解是这样的：$\hat{\mathcal A}>0$ 时，我们希望增大这个 action 的概率；$\hat{\mathcal A}<0$ 时，希望减小它的概率。但是当概率比沿着有利方向变化太多，就不再继续奖励这个变化。具体地，正 advantage 在 $\rho_\theta>1+\epsilon$ 时截平，负 advantage 在 $\rho_\theta<1-\epsilon$ 时截平。

它不是把所有概率比硬限制在区间里；共享参数下，其他样本的梯度仍可能继续改变这个 action 的概率。

### PPO-Penalty

非常直觉的式子：

$$
\theta_{k+1}=\arg\max_\theta\left\{
\mathbb E_{(s,a)\sim\pi_{\theta_k}}\left[
\rho_\theta(s,a)\hat{\mathcal A}(s,a)\right]
-\beta\overline D_{\mathrm{KL}}(\theta_k\|\theta)
\right\}.
$$

但是捏，固定一个 $\beta$ 不一定合适。我们既想优化前面那部分，又想控制每一步离旧策略多远。那咋搞呢，就是让 $\beta$ 动起来。跑一段优化，就去康康这个

$$
d=\widehat{\mathbb E}_{s\sim\pi_{\theta_k}}\left[
D_{\mathrm{KL}}(\pi_{\theta_k}(\cdot\mid s)\|\pi_\theta(\cdot\mid s))
\right].
$$

直觉上就是这玩意儿如果大了，就说明太远了，那我们得稍微调高一下 $\beta$；如果太小，就可以把 $\beta$ 搞低，允许更大的更新。

所以说我们可以定一个阈值 $d_{\mathrm{targ}}$：

$$
\beta\leftarrow
\begin{cases}
\beta/2,&d<d_{\mathrm{targ}}/1.5,\\
\beta,&d_{\mathrm{targ}}/1.5\le d\le1.5d_{\mathrm{targ}},\\
2\beta,&d>1.5d_{\mathrm{targ}}.
\end{cases}
$$

***

[系列导读](https://pufanyi.com/blog/ml/ml-revisit/rl) · [上一篇](https://pufanyi.com/blog/ml/ml-revisit/rl-policy-gradient) · [下一篇](https://pufanyi.com/blog/ml/ml-revisit/rl-llm)

## References

Achiam, J. (2018). *Spinning Up in Deep Reinforcement Learning*. [spinningup.openai.com](https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html "https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html")

Schulman, J., Levine, S., Moritz, P., Jordan, M., & Abbeel, P. (2015). Trust region policy optimization. *Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37*, 1889–1897. [proceedings.mlr.press](https://proceedings.mlr.press/v37/schulman15.html "https://proceedings.mlr.press/v37/schulman15.html")

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. *arXiv Preprint arXiv:1707.06347*. [arxiv.org](https://arxiv.org/abs/1707.06347 "https://arxiv.org/abs/1707.06347")

Simonini, T., & Sanseviero, O. (2023). The Hugging Face Deep Reinforcement Learning Class. In *GitHub repository*. GitHub. [github.com](https://github.com/huggingface/deep-rl-class "https://github.com/huggingface/deep-rl-class")
