[๋ ผ๋ฌธ๋ฆฌ๋ทฐ] Classifier-Free Diffusion Guidance
๐ NeurIPS Workshop 2021
Introduction
๊ธฐ์กด ์ฐ๊ตฌ๋ค์ ํ๊ณ์
Classifier guidance๋ ์๋์ ๋ฌธ์ ์ ์ด ์๋ค.
- ์ถ๊ฐ์ ์ธ classifier ํ์ต์ด ํ์ํ๊ธฐ ๋๋ฌธ์ diffusion ๋ชจ๋ธ์ ํ์ต ํ์ดํ๋ผ์ธ์ ๋ณต์กํ๊ฒ ๋ง๋ ๋ค.
- ์ฌ์ฉํ๋ classifier๋ ๋ ธ์ด์ฆ๊ฐ ์์ธ ๋ฐ์ดํฐ ์์์ ํ์ต๋์ด์ผ ํ๋ฏ๋ก, ์ผ๋ฐ์ ์ธ ์ฌ์ ํ์ต๋ classifier๋ฅผ ๋ฐ๋ก ์ฌ์ฉํ ์ ์๋ค.
FID, IS ๋ฑ์ metric์ ์ฌ์ ํ์ต๋ classifier๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ํ๊ฐํ๋๋ฐ, ์ด ๋ฐฉ์์ score๋ฅผ ์ง์ ์ ์ผ๋ก ๋ํ๋ฉด์ ํด๋์ค ๋ฐฉํฅ์ผ๋ก ์ํ๋ง์ ์ ๋ํ๊ธฐ ๋๋ฌธ์ classifier๊ฐ ์ ๋ง์ถ๊ฒ๋ ์ด๋ฏธ์ง์ ์์ ์กฐ์์ ๋ฐ๋ณตํ๋ ๊ฒ์ผ๋ก ๋ณผ ์ ์๋ค.
๋ฐ๋ผ์ ์ง์ง ์ข์ ์ด๋ฏธ์ง๋ฅผ ๋ง๋ ๊ฒ์ธ์ง, classifier๋ฅผ ์ ์์ธ ๊ฒฐ๊ณผ์ธ์ง๊ฐ ๋ถ๋ถ๋ช ํ๋ค.
์ ์ํ๋ ๋ฐฉ๋ฒ
๋ถ๋ฅ๊ธฐ๋ฅผ ์ฌ์ฉํ์ง ์๋ guidance ๊ธฐ๋ฒ์ ์ ์ํ๋ค.
- Conditional ๋ํจ์ ๋ชจ๋ธ์ socre ์ถ์ ๊ฐ๊ณผ unconditional ๋ํจ์ ๋ชจ๋ธ์ score ์ถ์ ๊ฐ์ ํผํฉํ๋ค.
Methods
GAN์ด๋ flow-based ๋ชจ๋ธ๊ณผ ๊ฐ์ ์์ฑ ๋ชจ๋ธ์ ์ํ๋ง ๋ ์ ๋ ฅ ๋ ธ์ด์ฆ์ ๋ถ์ฐ ๋๋ ๋ฒ์๋ฅผ ์ค์์ผ๋ก์จ truncated sampling ๋๋ low temperature sampling์ ์ํํ ์ ์๋ค.
์ด๋ฌํ ๋ฐฉ์์ ์ํ์ ๋ค์์ฑ์ ์ค์ด์ง๋ง, ๊ฐ ์ํ์ ํ์ง์ ๋์ผ ์ ์๋ค.
๊ทธ๋ฌ๋ ์์ ๋ฐฉ์์ ๋ํจ์ ๋ชจ๋ธ์ ์ ์ฉํ๊ธฐ ์ํด, score๋ฅผ ์ค์ผ์ผ๋งํ๊ฑฐ๋ reverse process์์ ๊ฐ์ฐ์์ ๋ ธ์ด์ฆ์ ๋ถ์ฐ์ ์ค์ด๋ ๋ฐฉ์์ ์ฌ์ฉํ๋ฉด ์คํ๋ ค ์ ํ์ง์ ์ํ์ ์์ฑํ๊ฒ ๋ง๋ ๋ค.
1. Classifier Guidance
๋ํจ์ ๋ชจ๋ธ์์ truncation๊ณผ ์ ์ฌํ ํจ๊ณผ๋ฅผ ์ป๊ธฐ ์ํด [Diffusion models beat GANs on image synthesis] ๋ ผ๋ฌธ์์๋ classifier guidance ๊ธฐ๋ฒ์ ๋์ ํ์๋ค.
์ด ๊ธฐ๋ฒ์์๋ ๋ํจ์ ๋ชจ๋ธ์ score ${\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda,\mathbf{c})\approx\sigma_\lambda\nabla_{\mathbf{z}_\lambda}\log p_\theta(\mathbf{z}_\lambda\mid\mathbf{c})$์ auxiliary classifier ๋ชจ๋ธ $p_\theta(\mathbf{c}\mid\mathbf{z}_\lambda)$์ log-likelihood์ gradient๋ฅผ ํฌํจ์ํจ๋ค.
\[\tilde{\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda,\mathbf{c}) ={\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda,\mathbf{c})-w\sigma_\lambda\nabla_{\mathbf{z}_\lambda}\log p_\theta(\mathbf{c}\mid\mathbf{z}_\lambda) \approx\sigma_\lambda\nabla_{\mathbf{z}_\lambda}[\log p_\theta(\mathbf{z}_\lambda\mid\mathbf{c})+w\log p_\theta(\mathbf{c}\mid\mathbf{z}_\lambda)]\]์์ ์์์์ $w$๋ classifier guidance์ ๊ฐ๋๋ฅผ ์กฐ์ ํ๋ ํ์ดํผํ๋ผ๋ฏธํฐ์ด๋ค.
์ํ๋ง ๋๋ ๊ธฐ์กด score ${\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda,\mathbf{c})$ ๋์ ๋ณ๊ฒฝ๋ score $\tilde{\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda,\mathbf{c})$๊ฐ ์ฌ์ฉ๋๋ฉฐ, ์ด๋ ์๋์ ๊ฐ์ ๋ถํฌ๋ก๋ถํฐ ์ํ๋ง์ ํ๋ ๊ฒ๊ณผ ๋์ผํ๋ค.
\[\tilde{p}_\theta(\mathbf{z}_\lambda \mid \mathbf{c}) \propto p_\theta(\mathbf{z}_\lambda \mid \mathbf{c}) \, p_\theta(\mathbf{c} \mid \mathbf{z}_\lambda)^w\]์ด๋, ๋ถ๋ฅ๊ธฐ $p_\theta(\mathbf{c} \mid \mathbf{z}_\lambda)$๊ฐ ํด๋นํ๋ ํด๋์ค $c$์ ๋ํด ๋์ likelihood๋ฅผ ๋ถ์ฌํ๋๋ก ๊ฐ์ค์น๋ฅผ ์กฐ์ ํ๋ฉฐ, $w>0$์ผ๋ก ์ค์ ํ๋ฉด ์์ฑ๋ ์ํ์ ๋ค์์ฑ์ด ๊ฐ์ํ์ง๋ง IS๊ฐ ํฅ์๋๋ค.
์๋์ ๊ด๊ณ ๋๋ฌธ์ ์ด๋ก ์ ์ผ๋ก, unconditional ๋ชจ๋ธ์ ๊ฐ์ค์น $w+1$์ classifier guidance๋ฅผ ์ ์ฉํ๋ฉด, conditional ๋ชจ๋ธ์ $w$์ guidance๋ฅผ ์ ์ฉํ ๊ฒ๊ณผ ๋์ผํ ๊ฒฐ๊ณผ๋ฅผ ์ป์ ์ ์๋ค.
\[p_\theta(\mathbf{z}_\lambda \mid \mathbf{c}) \, p_\theta(\mathbf{c} \mid \mathbf{z}_\lambda)^w =p_\theta(\mathbf{z}_\lambda)p_\theta(\mathbf{c} \mid \mathbf{z}_\lambda)^{w+1}\]Score ๊ด์ ์์๋ ์๋์ ๊ฐ๋ค.
\[\epsilon_\theta(\mathbf{z}_\lambda) - (w + 1) \, \sigma_\lambda \, \nabla_{\mathbf{z}_\lambda} \log p_\theta(\mathbf{c} \mid \mathbf{z}_\lambda) \approx -\sigma_\lambda \, \nabla_{\mathbf{z}_\lambda} \left[ \log p(\mathbf{z}_\lambda) + (w + 1) \log p_\theta(\mathbf{c} \mid \mathbf{z}_\lambda) \right] \\ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ = -\sigma_\lambda \, \nabla_{\mathbf{z}_\lambda} \left[ \log p(\mathbf{z}_\lambda \mid \mathbf{c}) + w \log p_\theta(\mathbf{c} \mid \mathbf{z}_\lambda) \right]\]ํ์ง๋ง Diffusion models beat GANs on image synthesis์ ์ ์๋ค์ ์ด๋ฏธ unconditional ๋ชจ๋ธ์ guidance๋ฅผ ์ ์ฉํ๋ ๊ฒ๋ณด๋ค conditional๋ก ํ์ต๋ ๋ชจ๋ธ์ guidance๋ฅผ ์ ์ฉํ๋ ๊ฒฝ์ฐ์ ์ฑ๋ฅ์ด ๋ ์ฐ์ํจ์ ๋ฐ๊ฒฌํ์๊ณ , ๋ฐ๋ผ์ ๋ณธ ์ฐ๊ตฌ์์๋ conditional ๋ชจ๋ธ์ guidance๋ฅผ ์ ์ฉํ๋ ์ค์ ์ ์ค์ฌ์ผ๋ก ๋ ผ์๋ฅผ ํ๋ค.
2. Classifier-Free Guidance
Classifier guidance๋ ์ด๋ฏธ์ง classifier์ graident์ ์์กดํ๋ค๋ ํ๊ณ์ ์ด ์๋ค.
๋ณ๋์ classifier๋ฅผ ํ์ตํ๋ ๋์ , ์๋์ ๋ ๋ชจ๋ธ์ ํจ๊ป ํ์ตํ๋ค.
- Unconditional denoising ๋ํจ์ ๋ชจ๋ธ $p_\theta(\mathbf{z})$: ์ด ๋ชจ๋ธ์ score ์ถ์ ๊ฐ์ $\boldsymbol\epsilon_\theta(\mathbf{z}_\lambda)$
- Conditional ๋ํจ์ ๋ชจ๋ธ $p_\theta(\mathbf{z}\mid\mathbf{c})$: ์ด ๋ชจ๋ธ์ score ์ถ์ ๊ฐ์ $\boldsymbol\epsilon_\theta(\mathbf{z}_\lambda,\mathbf{c})$
ํ๋ผ๋ฏธํฐ ์ฆ๊ฐ ์์ด ์ ๋ ๋ชจ๋ธ์ ๋์์ ํ์ตํ๊ธฐ ์ํด ์๋์ ์ ๋ต์ ์ฌ์ฉํ์๋ค.
- Unconditional ๋ชจ๋ธ์ ๊ฒฝ์ฐ conditional ๋ชจ๋ธ์ class identifier $\mathbf{c}$๋ก null token $\varnothing$์ ์ ๋ ฅ์ผ๋ก ์ค๋ค. ์ฆ, $\boldsymbol\epsilon_\theta(\mathbf{z}_\lambda)=\boldsymbol\epsilon_\theta(\mathbf{z}_\lambda,\mathbf{c}=\varnothing)$์ด๋ค.
- ํ์ดํผํ๋ผ๋ฏธํฐ $p_{\text{uncond}}$๋ฅผ ์ค์ ํด $p_{\text{uncond}}$์ ํ๋ฅ ๋ก $\mathbf{c}$๋ฅผ $\varnothing$์ผ๋ก ๋ง๋ค์ด ๋ชจ๋ธ์ ์ ๋ ฅํ๋ค.
์ดํ ์๋์ ๊ฐ์ด conditional๊ณผ unconditional score ์ถ์ ๊ฐ์ ์ ํ ๊ฒฐํฉํ์ฌ ์ํ๋งํ๋ค.
\[\tilde{\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda,\mathbf{c})= (1+w){\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda,\mathbf{c})-w{\boldsymbol\epsilon}_\theta(\mathbf{z}_\lambda)\]Classifier-Free Guidance์ ํ์ต๊ณผ ์ํ๋ง ์๊ณ ๋ฆฌ์ฆ์ ์๋์ ๊ฐ๋ค.
Experiments
Varying the Classifier-Free Guidance Strength
์์ ๊ทธ๋ฆผ์์ ์ผ์ชฝ์ $w=0$ (non-guided)์ผ ๋ ์์ฑ๋ ์ํ, ์ค๋ฅธ์ชฝ์ $w=3$์ผ ๋ ์์ฑ๋ ์ํ ๊ฒฐ๊ณผ์ด๋ค.
Classifier-free guidance strenght๊ฐ ํด์๋ก fidelity๋ ์ฆ๊ฐํ์ง๋ง ๋ค์์ฑ์ด ๊ฐ์ํ๋ค.
์์ ํ์ ๊ทธ๋ฆผ์ classifier-free guidance strenght์ ๋ฐ๋ฅธ FID์ IS๋ฅผ ๋น๊ตํ ๊ฒฐ๊ณผ๋ฅผ ๋ํ๋ธ๋ค.





