ML Revisit: Auto Regressive


2026-09-04

Early Explorations

主要参考了 (Ermon, 2023; Grover, 2018; He, 2024)

正常情况下,auto regressive 是

\[ p(x_1, \dots, x_n) = p(x_1)\prod_{i=2}^np(x_i\mid x_1, \dots, x_{i-1}) \]

但其实 \(x_i\) 不一定依赖全部的 \(x_{1}\sim x_{i-1}\),可能就是一个子集 \(\mathcal{P}_i\)。那么其实我们可以搞一个 DAG,然后 \(j\in \mathcal{P}_i\) 那么 \(j\)\(i\) 的 parent。这样

\[ p(x_1, \dots, x_n)=\prod_{i=1}^n p(x_i\mid x_{\mathcal{P}_i}) \]

就是 Bayesian Networks (Kyburg Jr, 1991; Peal, 1985)

Fully Visible Sigmoid Belief Network (FVSBN) (Frey et al., 1995) 把每个像素看成 \(0\)\(1\)(黑白)然后跑 Logistic regression:

\[ f_i(x_1, \dots, x_{i-1})=\sigma\left(\alpha_0^{(i)}+\sum_{j=1}^{i-1}\alpha_j^{(i)}x_j\right) \]

Neural Autoregressive Density Estimator (NADE) (Larochelle & Murray, 2011) 把 Logistic regression 改成了 MLP,RNADE (Uria et al., 2014) 把离散变量改成了连续分布,用多个 normal distribution 表示。EoNADE (Uria et al., 2014a) 消除了标准 NADE 对 order 的依赖。训练的时候随机选一个顺序进行生成。

Masked Autoencoder for Distribution Estimation (MADE) (Germain et al., 2015) 将 autoencoder 做成了 auto regressive 的形式,将每个 node 标 degree,小 degree 朝着大的 degree 连边(隐藏层允许 degree 相等,输出层要求严格小于),别的都 mask 掉。

MADE masking exampleA three-variable example with node degrees three, one, two at the input. Binary masks remove connections so the outputs model p of x2, p of x3 given x2, and p of x1 given x2 and x3.AutoencoderBinary masks1: keep0: removeMADE31221221221312depends on 1 inputdepends on 2 inputs
Example order: \(x_2 \to x_3 \to x_1\), giving \(p(x)=p(x_2)p(x_3\mid x_2)p(x_1\mid x_2,x_3)\).

RNN 就是把历史压成一个 latent,不解释。

GPT (Brown et al., 2020; Radford et al., 2018, 2019) 用了 transformer (Vaswani et al., 2017) 做 AR。

PixelRNN (van den Oord et al., 2016) 用 RNN 逐行填 pixel,PixelCNN 一开始也是在 (van den Oord et al., 2016) 提出的,思想大概是用 CNN 每次推理做 AR。当然因为一些 CNN 的问题后续还有很多别的修改 (van den Oord, Kalchbrenner, Vinyals, et al., 2016),这个有空再来补。

AR for Multimodal Generation

参考了 (Jin, 2026)

VideoGPT (Yan et al., 2021) 使用 3D VQ-VAE 将视频 down-sampling 成 latent,然后用 auto regressive 的方法进行视频生成。

References

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few-Shot Learners. arxiv.org
Ermon, S. (2023). Autoregressive Models. Stanford CS236: Deep Generative Models, Lecture 3 slides. deepgenerativemodels.github.io
Frey, B. J., Hinton, G. E., & Dayan, P. (1995). Does the Wake-sleep Algorithm Produce Good Density Estimators? In D. Touretzky, M. C. Mozer, & M. Hasselmo (Eds.), Advances in Neural Information Processing Systems (Vol. 8). MIT Press. proceedings.neurips.cc
Germain, M., Gregor, K., Murray, I., & Larochelle, H. (2015). MADE: Masked Autoencoder for Distribution Estimation. arxiv.org
Grover, A. (2018). Autoregressive Models. Stanford CS236: Deep Generative Models Lecture Notes. deepgenerativemodels.github.io
He, K. (2024). Autoregressive Models. MIT 6.S978: Deep Generative Models, Lecture 3 slides. mit-6s978.github.io
Jin, W. (2026). How Far Are We from a Native Autoregressive Video Model? The University of Hong Kong. waynejin0918.github.io
Kyburg Jr, H. E. (1991). Probabilistic reasoning in intelligent systems: networks of plausible inference. JSTOR. doi.org
Larochelle, H., & Murray, I. (2011). The neural autoregressive distribution estimator. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, 29–37. proceedings.mlr.press
Peal, J. (1985). Bayesian networks: A model of self-activated memory for evidential reasoning. Proceedings of the Annual Meeting of the Cognitive Science Society, 7. escholarship.org
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., & others. (2018). Improving language understanding by generative pre-training.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., & others. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 9.
Uria, B., Murray, I., & Larochelle, H. (2014a). A Deep and Tractable Density Estimator. arxiv.org
Uria, B., Murray, I., & Larochelle, H. (2014b). RNADE: The real-valued neural autoregressive density-estimator. arxiv.org
van den Oord, A., Kalchbrenner, N., & Kavukcuoglu, K. (2016). Pixel Recurrent Neural Networks. arxiv.org
van den Oord, A., Kalchbrenner, N., Vinyals, O., Espeholt, L., Graves, A., & Kavukcuoglu, K. (2016). Conditional Image Generation with PixelCNN Decoders. arxiv.org
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. arxiv.org
Yan, W., Zhang, Y., Abbeel, P., & Srinivas, A. (2021). VideoGPT: Video Generation using VQ-VAE and Transformers. arxiv.org

Cite this post

@misc{pu2026mlrevisitar,
  author = {Pu, Fanyi},
  title  = {ML Revisit: Auto Regressive},
  year   = {2026},
  month  = {9},
  url    = {https://pufanyi.com/blog/ml-revisit-ar}
}