Hi, Thank you for sharing this work. It has been very inspiring for me. However, I still have some questions regarding the SD-VAE part. Is the SD-VAE directly using the existing Stable Diffusion VAE? The paper seems to mention that the SD-VAE has three corresponding upsampling and downsampling layers. Was this part designed independently? If it needs to be designed independently, how are the SD pretrained weights utilized? Looking forward to your response.
Hi, Thank you for sharing this work. It has been very inspiring for me. However, I still have some questions regarding the SD-VAE part. Is the SD-VAE directly using the existing Stable Diffusion VAE? The paper seems to mention that the SD-VAE has three corresponding upsampling and downsampling layers. Was this part designed independently? If it needs to be designed independently, how are the SD pretrained weights utilized? Looking forward to your response.