VLDB 2026 Research / reviewers in the wild / expert
Fuwei Zhao
dblp:298/8129
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
7 papers |
Visual content generation and editing · 93% Rendering · 4% Image and video processing · 2% | |
| Artificial intelligence
7 papers |
Generative modeling · 72% 3D vision · 14% Efficient and distributed learning · 10% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing
virtual try-on |
2.9 | 5 | 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Dressing in the Wild by Watching Dance Videos · CVPR 2022 Towards Scalable Unpaired Virtual Try-On via Patch-Routed Spatially-Adaptive GAN · NeurIPS 2021 |
Machine learning › Generative modeling
diffusion model |
1.5 | 2 | 2025 | DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder · AAAI 2025 GP-VTON: Towards General Purpose Virtual Try-On via Collaborative Local-Flow Global-Parsing Learning · CVPR 2023 |
Machine learning › Generative modeling
generative adversarial network |
1.0 | 2 | 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Towards Scalable Unpaired Virtual Try-On via Patch-Routed Spatially-Adaptive GAN · NeurIPS 2021 |
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis |
0.9 | 1 | 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Visual content generation and editing › virtual try-on
3d virtual try-on |
0.9 | 1 | 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Visual content generation and editing › virtual try-on
image-based virtual try-on |
0.7 | 1 | 2023 | GP-VTON: Towards General Purpose Virtual Try-On via Collaborative Local-Flow Global-Parsing Learning · CVPR 2023 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.5 | 1 | 2021 | WAS-VTON: Warping Architecture Search for Virtual Try-on Network · ACM Multimedia 2021 |
Visual content generation and editing › virtual try-on
garment transfer |
0.5 | 1 | 2021 | Towards Scalable Unpaired Virtual Try-On via Patch-Routed Spatially-Adaptive GAN · NeurIPS 2021 |
Computer vision › 3D vision
3d human reconstruction |
0.4 | 2 | 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2025 M3D-VTON: A Monocular-to-3D Virtual Try-On Network · ICCV 2021 |
Computer vision › 3D vision › 3d shape representation › implicit surface representation
signed distance function |
0.3 | 1 | 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Visual content generation and editing
image generation |
0.3 | 1 | 2025 | DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder · AAAI 2025 |
Rendering
multi-view rendering |
0.3 | 1 | 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Visual content generation and editing › image generation
text-to-image generation |
0.3 | 1 | 2025 | DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder · AAAI 2025 |
Image and video processing
image warping |
0.1 | 1 | 2021 | WAS-VTON: Warping Architecture Search for Virtual Try-on Network · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
signed distance function · 1.7large multimodal model · 1.7inverse skinning · 1.7deferred pose guidance · 1.7conditional 3D-GAN · 1.7adaptive attention · 1.7LoRA · 1.7local-flow global-parsing warping · 1.3dynamic gradient truncation · 1.3cycle optimization · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing EncoderabstractDiffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetune-based methods involve high training costs and struggle to maintain the generalization capabilities of pretrained diffusion models, limiting their performance across diverse scenarios. To address these challenges, we propose DreamFit, which incorporates a lightweight Anything-Dressing Encoder specifically tailored for the garment-centric human generation. DreamFit has three key advantages: (1) Lightweight training: with the proposed adaptive attention and LoRA modules, DreamFit significantly minimizes the model complexity to 83.4M trainable parameters. (2) Anything-Dressing: Our model generalizes surprisingly well to a wide range of (non-)garments, creative styles, and prompt instructions, consistently delivering high-quality results across diverse scenarios. (3) Plug-and-play: DreamFit is engineered for smooth integration with any community control plugins for diffusion models, ensuring easy compatibility and minimizing adoption barriers. To further enhance generation quality, DreamFit leverages pretrained large multi-modal models (LMMs) to enrich the prompt with fine-grained garment descriptions, thereby reducing the prompt gap between training and inference. We conduct comprehensive experiments on both 768 x 512 high-resolution benchmarks and in-the-wild images. DreamFit surpasses all existing methods, highlighting its state-of-the-art capabilities of garment-centric human generation. Ente Lin, Xujie Zhang, Fuwei Zhao, Yuxuan Luo 0002, Long Zeng 0001, Xiaodan Liang |
AAAI | 3 |
| 2025 | Monocular-to-3D Virtual Try-On With Generative Semantic Articulated FieldsabstractWe introduce a monocular-to-3D virtual try-on network based on a conditional 3D-aware Generative Adversarial Network (3D-GAN) for synthesizing multi-view try-on results from single monocular images. In contrast to previous 3D virtual try-on methods that rely on costly scanned meshes or pseudo-depth maps for supervision, our approach utilizes a conditional 3D-GAN trained solely on 2D images, greatly simplifying dataset construction and enhancing model scalability. Specifically, we propose a Generative monocular-to-3D Virtual Try-ON network (G3D-VTON) that integrates a 3D-aware conditional Parsing Module (3DPM), a U-Net Refinement Module (URM), and a Flow-based 2D Virtual Try-On Module (FTM). In our framework, the 3DPM is designed to generate a 3D representation of the virtual try-on result, thereby enabling multi-view rendering. To accomplish this, it is implemented using conditional generative semantic articulated fields, which leverage the 3D SMPL prior via inverse skinning to learn the Signed Distance Function (SDF) of the try-on results in a canonical pose space. This learned SDF enables the rendering of both a coarse human parsing map and a preliminary try-on output with explicit camera control. Furthermore, within 3DPM, we introduce deferred pose guidance to decouple style and pose conditions during training, thereby facilitating view controllable generation during inference. However, the rendered human parsing and try-on results exhibit imprecise shapes and blurry textures. To address these issues, the URM subsequently refines these rendered outputs using a refinement U-Net, and the FTM integrates the refined results with the 2D warped garment to generate the final try-on output with more accurate and realistic appearance details. Extensive experiments demonstrate that the proposed G3D-VTON effectively manipulates and generates faithful 3D human appearances wearing the desired garment, outperforming both 3D-GAN and depth-based 3D approaches while delivering superior visual results in 2D. Zhenyu Xie, Fuwei Zhao, Jun Zheng 0020, Feida Zhu 0002, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | GP-VTON: Towards General Purpose Virtual Try-On via Collaborative Local-Flow Global-Parsing LearningabstractImage-based Virtual Try-ON aims to transfer an in-shop garment onto a specific person. Existing methods employ a global warping module to model the anisotropic deformation for different garment parts, which fails to preserve the semantic information of different parts when receiving challenging inputs (e.g, intricate human poses, difficult garments). Moreover, most of them directly warp the input garment to align with the boundary of the preserved region, which usually requires texture squeezing to meet the boundary shape constraint and thus leads to texture distortion. The above inferior performance hinders existing methods from real-world applications. To address these problems and take a step towards real-world virtual try-on, we propose a General-Purpose Virtual Try-ON framework, named GP-VTON, by developing an innovative Local-Flow Global-Parsing (LFGP) warping module and a Dynamic Gradient Truncation (DGT) training strategy. Specifically, compared with the previous global warping mechanism, LFGP employs local flows to warp garments parts individually, and assembles the local warped results via the global garment parsing, resulting in reasonable warped parts and a semantic-correct intact garment even with challenging inputs. On the other hand, our DGT training strategy dynamically truncates the gradient in the overlap area and the warped garment is no more required to meet the boundary constraint, which effectively avoids the texture squeezing problem. Furthermore, our GP-VTON can be easily extended to multi-category scenario and jointly trained by using data from different garment categories. Extensive experiments on two high-resolution benchmarks demonstrate our superiority over the existing state-of-the-art methods.11Code is available at gp-vton. Zhenyu Xie, Zaiyu Huang, Fuwei Zhao, Haoye Dong, Xijin Zhang, Feida Zhu 0002, Xiaodan Liang |
CVPR | 4 |
| 2022 | Dressing in the Wild by Watching Dance VideosabstractWhile significant progress has been made in garment transfer, one of the most applicable directions of human-centric image generation, existing works overlook the in-the-wild imagery, presenting severe garment-person mis-alignment as well as noticeable degradation in fine texture details. This paper, therefore, attends to virtual try-on in real-world scenes and brings essential improvements in authenticity and naturalness especially for loose garment (e.g., skirts, formal dresses), challenging poses (e.g., cross arms, bent legs), and cluttered backgrounds. Specifically, we find that the pixel flow excels at handling loose gar-ments whereas the vertex flow is preferred for hard poses, and by combining their advantages we propose a novel generative network called wFlow that can effectively push up garment transfer to in-the-wild context. Moreover, former approaches require paired images for training. Instead, we cut down the laboriousness by working on a newly constructed large-scale video dataset named Dance50k with self-supervised cross-frame training and an online cycle op-timization. The proposed Dance50k can boost real-world virtual dressing by covering a wide variety of garments under dancing poses. Extensive experiments demonstrate the superiority of our w Flow in generating realistic garment transfer results for in-the-wild images without resorting to expensive paired datasets.11Xiaodan Liang is the corresponding author. The project page of wFlow is https://awesome-wflow.github.io. Fuwei Zhao, Zhenyu Xie, Xijin Zhang, Daniel K. Du, Xiang Long, Xiaodan Liang, Jianchao Yang |
CVPR | 2 |
| 2021 | M3D-VTON: A Monocular-to-3D Virtual Try-On NetworkabstractVirtual 3D try-on can provide an intuitive and realistic view for online shopping and has a huge potential commercial value. However, existing 3D virtual try-on methods mainly rely on annotated 3D human shapes and garment templates, which hinders their applications in practical scenarios. 2D virtual try-on approaches provide a faster alternative to manipulate clothed humans, but lack the rich and realistic 3D representation. In this paper, we propose a novel Monocular-to-3D Virtual Try-On Network (M3D-VTON) that builds on the merits of both 2D and 3D approaches. By integrating 2D information efficiently and learning a mapping that lifts the 2D representation to 3D, we make the first attempt to reconstruct a 3D try-on mesh only taking the target clothing and a person image as inputs. The proposed M3D-VTON includes three modules: 1) The Monocular Prediction Module (MPM) that estimates an initial full-body depth map and accomplishes 2D clothes-person alignment through a novel two-stage warping procedure; 2) The Depth Refinement Module (DRM) that refines the initial body depth to produce more detailed pleat and face characteristics; 3) The Texture Fusion Module (TFM) that fuses the warped clothing with the non-target body part to refine the results. We also construct a high-quality synthesized Monocular-to-3D virtual try-on dataset, in which each person image is associated with a front and a back depth map. Extensive experiments demonstrate that the proposed M3D-VTON can manipulate and reconstruct the 3D human body wearing the given clothing with compelling details and is more efficient than other 3D approaches.1 Fuwei Zhao, Zhenyu Xie, Michael Kampffmeyer, Haoye Dong, Songfang Han, Tianxiang Zheng 0001, Tao Zhang 0042, Xiaodan Liang |
ICCV | 1 |
| 2021 | WAS-VTON: Warping Architecture Search for Virtual Try-on NetworkabstractDespite recent progress on image-based virtual try-on, current methods are constraint by shared warping networks and thus fail to synthesize natural try-on results when faced with clothing categories that require different warping operations. In this paper, we address this problem by finding clothing category-specific warping networks for the virtual try-on task via Neural Architecture Search (NAS). We introduce a NAS-Warping Module and elaborately design a bilevel hierarchical search space to identify the optimal network-level and operation-level flow estimation architecture. Given the network-level search space, containing different numbers of warping blocks, and the operation-level search space with different convolution operations, we jointly learn a combination of repeatable warping cells and convolution operations specifically for the clothing-person alignment. Moreover, a NAS-Fusion Module is proposed to synthesize more natural final try-on results, which is realized by leveraging particular skip connections to produce better-fused features that are required for seamlessly fusing the warped clothing and the unchanged person part. We adopt an efficient and stable one-shot searching strategy to search the above two modules. Extensive experiments demonstrate that our WAS-VTON significantly outperforms the previous fixed-architecture try-on methods with more natural warping results and virtual try-on results. Zhenyu Xie, Xujie Zhang, Fuwei Zhao, Haoye Dong, Michael Kampffmeyer, Haonan Yan, Xiaodan Liang |
ACM Multimedia | 3 |
| 2021 | Towards Scalable Unpaired Virtual Try-On via Patch-Routed Spatially-Adaptive GANabstractImage-based virtual try-on is one of the most promising applications of human-centric image generation due to its tremendous real-world potential. Yet, as most try-on approaches fit in-shop garments onto a target person, they require the laborious and restrictive construction of a paired training dataset, severely limiting their scalability. While a few recent works attempt to transfer garments directly from one person to another, alleviating the need to collect paired datasets, their performance is impacted by the lack of paired (supervised) information. In particular, disentangling style and spatial information of the garment becomes a challenge, which existing methods either address by requiring auxiliary data or extensive online optimization procedures, thereby still inhibiting their scalability. To achieve a scalable virtual try-on system that can transfer arbitrary garments between a source and a target person in an unsupervised manner, we thus propose a texture-preserving end-to-end network, the PAtch-routed SpaTially-Adaptive GAN (PASTA-GAN), that facilitates real-world unpaired virtual try-on. Specifically, to disentangle the style and spatial information of each garment, PASTA-GAN consists of an innovative patch-routed disentanglement module for successfully retaining garment texture and shape characteristics. Guided by the source person's keypoints, the patch-routed disentanglement module first decouples garments into normalized patches, thus eliminating the inherent spatial information of the garment, and then reconstructs the normalized patches to the warped garment complying with the target person pose. Given the warped garment, PASTA-GAN further introduces novel spatially-adaptive residual blocks that guide the generator to synthesize more realistic garment details. Extensive comparisons with paired and unpaired approaches demonstrate the superiority of PASTA-GAN, highlighting its ability to generate high-quality try-on images when faced with a large variety of garments(e.g. vests, shirts, pants), taking a crucial step towards real-world scalable try-on. Zhenyu Xie, Zaiyu Huang, Fuwei Zhao, Haoye Dong, Michael Kampffmeyer, Xiaodan Liang |
NeurIPS | 3 |