TY - GEN
T1 - Pretrained Reversible Generation as Unsupervised Visual Representation Learning
AU - Xue, Rongkun
AU - Zhang, Jinouwen
AU - Niu, Yazhe
AU - Shen, Dazhong
AU - Ma, Bingqi
AU - Liu, Yu
AU - Yang, Jing
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Recent generative models based on score matching and flow matching have significantly advanced generation tasks, but their potential in discriminative tasks remains underexplored. Previous approaches, such as generative classifiers, have not fully leveraged the capabilities of these models for discriminative tasks due to their intricate designs. We propose Pretrained Reversible Generation (PRG), which extracts unsupervised representations by reversing the generative process of a pretrained continuous generation model. PRG effectively reuses unsupervised generative models, leveraging their high capacity to serve as robust and generalizable feature extractors for downstream tasks. This framework enables the flexible selection of feature hierarchies tailored to specific downstream tasks. Our method consistently outperforms prior approaches across multiple benchmarks, achieving state-of-the-art performance among generative model based methods, including 78% top-1 accuracy on ImageNet at a resolution of 64 × 64. Extensive ablation studies, including out-of-distribution evaluations, further validate the effectiveness of our approach. PRG is available at https://opendilab.github.io/PRG/.
AB - Recent generative models based on score matching and flow matching have significantly advanced generation tasks, but their potential in discriminative tasks remains underexplored. Previous approaches, such as generative classifiers, have not fully leveraged the capabilities of these models for discriminative tasks due to their intricate designs. We propose Pretrained Reversible Generation (PRG), which extracts unsupervised representations by reversing the generative process of a pretrained continuous generation model. PRG effectively reuses unsupervised generative models, leveraging their high capacity to serve as robust and generalizable feature extractors for downstream tasks. This framework enables the flexible selection of feature hierarchies tailored to specific downstream tasks. Our method consistently outperforms prior approaches across multiple benchmarks, achieving state-of-the-art performance among generative model based methods, including 78% top-1 accuracy on ImageNet at a resolution of 64 × 64. Extensive ablation studies, including out-of-distribution evaluations, further validate the effectiveness of our approach. PRG is available at https://opendilab.github.io/PRG/.
KW - computer vision
KW - diffusion model
KW - flow model
KW - generation and understanding
KW - generative model
KW - unsupervised visual representation learning
UR - https://www.scopus.com/pages/publications/105044210231
U2 - 10.1109/ICCV51701.2025.01786
DO - 10.1109/ICCV51701.2025.01786
M3 - 会议稿件
AN - SCOPUS:105044210231
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 19216
EP - 19226
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Y2 - 19 October 2025 through 23 October 2025
ER -