# Transfusion

Author: Fanyi Pu

Published: 2026-08-31

Canonical: <https://pufanyi.com/blog/ml/ml-revisit/umm>

Notes for Transfusion

论文：([Zhou et al., 2024](https://pufanyi.com/blog/ml/ml-revisit/umm#bib-zhou2024transfusionpredicttokendiffuse))

大概就是用一个 Transfusion 去同时做生成和理解两件事情：

Transfusion model diagram

A single Transformer predicts text tokens autoregressively and denoises a sequence of four image patches.

Transformer

A

cute

cat

.

\<BOI>

\<EOI>

What

color

is

its

nose

?

[View diagram in the original article](https://pufanyi.com/blog/ml/ml-revisit/umm)

Attention mask 大概是这样，就是文字看到前面的，然后图片互相都能看到。

|        | A       | cute    | cat     | \<BOI>  |         |         |         |         | \<EOI>  | What    |
| ------ | ------- | ------- | ------- | ------- | ------- | ------- | ------- | ------- | ------- | ------- |
| A      | Allowed | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  |
| cute   | Allowed | Allowed | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  |
| cat    | Allowed | Allowed | Allowed | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  |
| \<BOI> | Allowed | Allowed | Allowed | Allowed | Masked  | Masked  | Masked  | Masked  | Masked  | Masked  |
|        | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Masked  | Masked  |
|        | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Masked  | Masked  |
|        | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Masked  | Masked  |
|        | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Masked  | Masked  |
| \<EOI> | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Masked  |
| What   | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed |

## References

Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., & Levy, O. (2024). *Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model*. [arxiv.org](https://arxiv.org/abs/2408.11039 "https://arxiv.org/abs/2408.11039")
