TY - GEN
T1 - Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
AU - Yi, Qiaosi
AU - Li, Shuai
AU - Wu, Rongyuan
AU - Sun, Lingchen
AU - Wu, Yuhui
AU - Zhang, Lei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025/10
Y1 - 2025/10
N2 - Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggressive resolution reduction of the VAE (e.g., 8 × downsampling) in the SD model. One solution is to employ a VAE with a lower downsampling rate for diffusion; however, adapting its latent features with the pretrained UNet while mitigating the increased computational cost poses new challenges. To address these issues, we propose a Transfer VAE Training (TVT) strategy to transfer the 8 × downsampled VAE into a 4 × one while adapting to the pre-trained UNet. Specifically, we first train a 4 × decoder based on the output features of the original VAE encoder, then train a 4 × encoder while keeping the newly trained decoder fixed. Such a TVT strategy aligns the new encoder-decoder pair with the original VAE latent space while enhancing image fine details. Additionally, we introduce a compact VAE and compute-efficient UNet by optimizing their network architectures, reducing the computational cost while capturing high-resolution fine-scale features. Experimental results demonstrate that our TVT method significantly improves fine-structure preservation, which is often compromised by other SD-based methods, while requiring fewer FLOPs than state-of-the-art onestep diffusion models. The official code can be found at https://github.com/Joyies/TVT.
AB - Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggressive resolution reduction of the VAE (e.g., 8 × downsampling) in the SD model. One solution is to employ a VAE with a lower downsampling rate for diffusion; however, adapting its latent features with the pretrained UNet while mitigating the increased computational cost poses new challenges. To address these issues, we propose a Transfer VAE Training (TVT) strategy to transfer the 8 × downsampled VAE into a 4 × one while adapting to the pre-trained UNet. Specifically, we first train a 4 × decoder based on the output features of the original VAE encoder, then train a 4 × encoder while keeping the newly trained decoder fixed. Such a TVT strategy aligns the new encoder-decoder pair with the original VAE latent space while enhancing image fine details. Additionally, we introduce a compact VAE and compute-efficient UNet by optimizing their network architectures, reducing the computational cost while capturing high-resolution fine-scale features. Experimental results demonstrate that our TVT method significantly improves fine-structure preservation, which is often compromised by other SD-based methods, while requiring fewer FLOPs than state-of-the-art onestep diffusion models. The official code can be found at https://github.com/Joyies/TVT.
UR - https://www.scopus.com/pages/publications/105044257271
U2 - 10.1109/ICCV51701.2025.01154
DO - 10.1109/ICCV51701.2025.01154
M3 - Conference article published in proceeding or book
AN - SCOPUS:105044257271
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 12415
EP - 12426
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Y2 - 19 October 2025 through 23 October 2025
ER -