# Auto Regressive

Author: Fanyi Pu

Published: 2026-09-04

Canonical: <https://pufanyi.com/blog/ml/ml-revisit/ar>

Notes for Auto Regressive Models

## Early Explorations

主要参考了 ([Ermon, 2023](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-ermon2023autoregressive); [Grover, 2018](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-grover2018autoregressive); [He, 2024](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-he2024autoregressive))。

正常情况下，auto regressive 是

$$
p(x_1, \dots, x_n) = p(x_1)\prod_{i=2}^np(x_i\mid x_1, \dots, x_{i-1})
$$

但其实 $x_i$ 不一定依赖全部的 $x_{1}\sim x_{i-1}$，可能就是一个子集 $\mathcal{P}_i$。那么其实我们可以搞一个 DAG，然后 $j\in \mathcal{P}_i$ 那么 $j$ 做 $i$ 的 parent。这样

$$
p(x_1, \dots, x_n)=\prod_{i=1}^n p(x_i\mid x_{\mathcal{P}_i})
$$

就是 Bayesian Networks ([Kyburg Jr, 1991](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-kyburg1991probabilistic); [Peal, 1985](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-peal1985bayesian))。

Fully Visible Sigmoid Belief Network (FVSBN) ([Frey et al., 1995](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-nips1995_55b1927f)) 把每个像素看成 $0$ 或 $1$（黑白）然后跑 Logistic regression:

$$
f_i(x_1, \dots, x_{i-1})=\sigma\left(\alpha_0^{(i)}+\sum_{j=1}^{i-1}\alpha_j^{(i)}x_j\right)
$$

Neural Autoregressive Density Estimator (NADE) ([Larochelle & Murray, 2011](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-larochelle2011neural)) 把 Logistic regression 改成了 MLP，RNADE ([Uria et al., 2014](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-uria2014rnaderealvaluedneuralautoregressive)) 把离散变量改成了连续分布，用多个 normal distribution 表示。EoNADE ([Uria et al., 2014a](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-uria2014deeptractabledensityestimator)) 消除了标准 NADE 对 order 的依赖。训练的时候随机选一个顺序进行生成。

Masked Autoencoder for Distribution Estimation (MADE) ([Germain et al., 2015](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-germain2015mademaskedautoencoderdistribution)) 将 autoencoder 做成了 auto regressive 的形式，将每个 node 标 degree，小 degree 朝着大的 degree 连边（隐藏层允许 degree 相等，输出层要求严格小于），别的都 mask 掉。

MADE masking example

A three-variable example with node degrees three, one, two at the input. Binary masks remove connections so the outputs model p of x2, p of x3 given x2, and p of x1 given x2 and x3.

Autoencoder

$x_1$

$x_2$

$x_3$

$\hat{x}_1$

$\hat{x}_2$

$\hat{x}_3$

Binary masks

$M^{3}$

$M^{2}$

$M^{1}$

1: keep

0: remove

$W \odot M$

MADE

3

1

2

$p(x_1 \mid x_2,x_3)$

$p(x_2)$

$p(x_3 \mid x_2)$

depends on 1 input

depends on 2 inputs

[View diagram in the original article](https://pufanyi.com/blog/ml/ml-revisit/ar)

Example order: $x_2 \to x_3 \to x_1$, giving $p(x)=p(x_2)p(x_3\mid x_2)p(x_1\mid x_2,x_3)$.

RNN 就是把历史压成一个 latent，不解释。

GPT ([Brown et al., 2020](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-brown2020languagemodelsfewshotlearners); [Radford et al., 2018](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-radford2018improving), [2019](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-radford2019language)) 用了 transformer ([Vaswani et al., 2017](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-vaswani2017attention)) 做 AR。

PixelRNN ([van den Oord et al., 2016](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-oord2016pixelrecurrentneuralnetworks)) 用 RNN 逐行填 pixel，PixelCNN 一开始也是在 ([van den Oord et al., 2016](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-oord2016pixelrecurrentneuralnetworks)) 提出的，思想大概是用 CNN 每次推理做 AR。当然因为一些 CNN 的问题后续还有很多别的修改 ([van den Oord, Kalchbrenner, Vinyals, et al., 2016](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-oord2016conditionalimagegenerationpixelcnn))，这个有空再来补。

## AR for Multimodal Generation

参考了 ([Jin, 2026](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-jin2026nativear))。

VideoGPT ([Yan et al., 2021](https://pufanyi.com/blog/ml/ml-revisit/ar#bib-yan2021videogptvideogenerationusing)) 使用 3D VQ-VAE 将视频 down-sampling 成 latent，然后用 auto regressive 的方法进行视频生成。

## References

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). *Language Models are Few-Shot Learners*. [arxiv.org](https://arxiv.org/abs/2005.14165 "https://arxiv.org/abs/2005.14165")

Ermon, S. (2023). *Autoregressive Models*. Stanford CS236: Deep Generative Models, Lecture 3 slides. [deepgenerativemodels.github.io](https://deepgenerativemodels.github.io/assets/slides/cs236_lecture3.pdf "https://deepgenerativemodels.github.io/assets/slides/cs236_lecture3.pdf")

Frey, B. J., Hinton, G. E., & Dayan, P. (1995). Does the Wake-sleep Algorithm Produce Good Density Estimators? In D. Touretzky, M. C. Mozer, & M. Hasselmo (Eds.), *Advances in Neural Information Processing Systems* (Vol. 8). MIT Press. [proceedings.neurips.cc](https://proceedings.neurips.cc/paper_files/paper/1995/hash/55b1927fdafef39c48e5b73b5d61ea60-Abstract.html "https://proceedings.neurips.cc/paper_files/paper/1995/hash/55b1927fdafef39c48e5b73b5d61ea60-Abstract.html")

Germain, M., Gregor, K., Murray, I., & Larochelle, H. (2015). *MADE: Masked Autoencoder for Distribution Estimation*. [arxiv.org](https://arxiv.org/abs/1502.03509 "https://arxiv.org/abs/1502.03509")

Grover, A. (2018). *Autoregressive Models*. Stanford CS236: Deep Generative Models Lecture Notes. [deepgenerativemodels.github.io](https://deepgenerativemodels.github.io/notes/autoregressive/ "https://deepgenerativemodels.github.io/notes/autoregressive/")

He, K. (2024). *Autoregressive Models*. MIT 6.S978: Deep Generative Models, Lecture 3 slides. [mit-6s978.github.io](https://mit-6s978.github.io/assets/pdfs/lec3_ar.pdf "https://mit-6s978.github.io/assets/pdfs/lec3_ar.pdf")

Jin, W. (2026). *How Far Are We from a Native Autoregressive Video Model?* The University of Hong Kong. [waynejin0918.github.io](https://waynejin0918.github.io/how-far-native-ar-video/ "https://waynejin0918.github.io/how-far-native-ar-video/")

Kyburg Jr, H. E. (1991). *Probabilistic reasoning in intelligent systems: networks of plausible inference*. JSTOR. [doi.org](https://doi.org/10.2307/2026705 "https://doi.org/10.2307/2026705")

Larochelle, H., & Murray, I. (2011). The neural autoregressive distribution estimator. *Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics*, 29–37. [proceedings.mlr.press](https://proceedings.mlr.press/v15/larochelle11a.html "https://proceedings.mlr.press/v15/larochelle11a.html")

Peal, J. (1985). Bayesian networks: A model of self-activated memory for evidential reasoning. *Proceedings of the Annual Meeting of the Cognitive Science Society*, *7*. [escholarship.org](https://escholarship.org/uc/item/0vr7830n "https://escholarship.org/uc/item/0vr7830n")

Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., & others. (2018). *Improving language understanding by generative pre-training*.

Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., & others. (2019). Language models are unsupervised multitask learners. *OpenAI Blog*, *1*(8), 9.

Uria, B., Murray, I., & Larochelle, H. (2014a). *A Deep and Tractable Density Estimator*. [arxiv.org](https://arxiv.org/abs/1310.1757 "https://arxiv.org/abs/1310.1757")

Uria, B., Murray, I., & Larochelle, H. (2014b). *RNADE: The real-valued neural autoregressive density-estimator*. [arxiv.org](https://arxiv.org/abs/1306.0186 "https://arxiv.org/abs/1306.0186")

van den Oord, A., Kalchbrenner, N., & Kavukcuoglu, K. (2016). *Pixel Recurrent Neural Networks*. [arxiv.org](https://arxiv.org/abs/1601.06759 "https://arxiv.org/abs/1601.06759")

van den Oord, A., Kalchbrenner, N., Vinyals, O., Espeholt, L., Graves, A., & Kavukcuoglu, K. (2016). *Conditional Image Generation with PixelCNN Decoders*. [arxiv.org](https://arxiv.org/abs/1606.05328 "https://arxiv.org/abs/1606.05328")

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. *Advances in Neural Information Processing Systems*, *30*. [arxiv.org](https://arxiv.org/abs/1706.03762 "https://arxiv.org/abs/1706.03762")

Yan, W., Zhang, Y., Abbeel, P., & Srinivas, A. (2021). *VideoGPT: Video Generation using VQ-VAE and Transformers*. [arxiv.org](https://arxiv.org/abs/2104.10157 "https://arxiv.org/abs/2104.10157")
