Zhengwentai Sun

dblp:284/1032 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-3884-137XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Robust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-On
abstract
This paper tackles the emerging challenge of multi-view virtual try-on, utilizing both front- and back-view clothing images as inputs. Extending frontal try-on methods to a multi-view context is not straightforward. Simply concatenating the two input views or encoding their features for a generative model, such as a diffusion model, often fails to produce satisfactory results. The main challenge lies in effectively extracting and fusing meaningful clothing features from these input views. Existing explicit warping-based methods, which establish direct correspondence between input and target views, tend to introduce artifacts, particularly when there is a significant disparity between the input and target views. Conversely, implicit encoding-based methods often lose spatial information about clothing, resulting in outputs that lack detail. To overcome these challenges, we propose Robust-MVTON, an end-to-end method for robust and high-quality multi-view try-ons. Our approach introduces a novel cross-pose feature alignment technique to guide the fusion of clothing features and incorporates a newly designed loss function for training. With the fused multi-scale clothing features, we employ a coarse-to-fine diffusion model to generate realistic and detailed results. Extensive experiments conducted on the Deepfashion and MPV datasets affirm the superiority of our method, achieving state-of-the-art performance.
Yijiang Li, Dong Du 0002, Zheng Chong, Zhengwentai Sun, Jianhao Zeng, Yusheng Dai, Zhengyu Xie, Hairui Zhu, Xiaoguang Han 0001
CVPR5
2025 MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis
abstract
Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current methods primarily concentrate on redirecting camera trajectory within the front view while struggling to generate 360-degree viewpoint changes. In this paper, we focus on human-centric subdomain and present MV-Performer, an innovative framework for creating synchronized novel view videos from monocular full-body captures. To achieve a 360-degree synthesis, we extensively leverage the MVHumanNet dataset and incorporate an informative condition signal. Specifically, we use the camera-dependent normal maps rendered from oriented partial point clouds, which effectively alleviate the ambiguity between seen and unseen observations. To maintain synchronization in the generated videos, we propose a multi-view human-centric video diffusion model that fuses information from the reference video, partial rendering, and different viewpoints. Additionally, we provide a robust inference procedure for in-the-wild video cases, which greatly mitigates the artifacts induced by imperfect monocular depth estimation. Extensive experiments on three datasets demonstrate our MV-Performer’s state-of-the-art effectiveness and robustness, setting a strong model for human-centric 4D novel view synthesis. Code is available at https://github.com/zyhbili/MV-Performer.
Yihao Zhi, Chenghong Li, Hongjie Liao, Xihe Yang, Zhengwentai Sun, Xiaodong Cun, Wensen Feng, Xiaoguang Han 0001
SIGGRAPH Asia5
2025 TiPGAN: High-quality tileable textures synthesis with intrinsic priors for cloth digitization applications
abstract
Seamless textures play an important role in 3D modeling, animation, video games, and Augmented Reality/Virtual Reality, enhancing the realism and aesthetics of the digital environments. Despite its significance, generating seamless textures is not trivial, requiring the edges of the synthesized texture image to represent a continuous pattern when tiled. Although traditional methods and deep learning models have made good progress in texture synthesis, they often fail in ensuring the seamless property of the synthesized textures. In this paper, we report on TiPGAN, a Generative Adversarial Network (GAN) model, which we developed to generate seamless textures. Leveraging the inherent intrinsics of seamless textures as priors , our model introduces two novel modules: a Patch Swapping Module, for maintaining texture continuity through diagonal patch swapping, and a Patch Tiling Module, for ensuring seamless repetition across tiles. To overcome the limitations of existing image quality metrics in evaluating tileability, we introduce a new metric, termed Relative Total Variation (RTV), for assessing the smoothness and continuity of the synthesized textures. Our experimental results demonstrate that TiPGAN outperforms existing methods in generating high-quality, seamless textures, as validated by both conventional image quality metrics and our newly proposed RTV metric. This research represents a significant advancement in texture generation, offering valuable applications in graphic design, virtual reality, and digital art. Our code and dataset are available at https://github.com/VickyHEHonghong/TiPGAN . • TiPGAN leverages intrinsic priors to improve seamless texture synthesis. • RTV metric is introduced to quantitatively evaluates the texture tileability. • TiPGAN outperforms existing methods as verified by image quality metrics and RTV. • TiPGAN contributes to cloth digitization, providing value to digital art.
Honghong He, Zhengwentai Sun, Jintu Fan, P. Y. Mok 0001
Comput. Aided Des.2
2025 CoDE-GAN: Content Decoupled and Enhanced GAN for Sketch-guided Flexible Fashion Editing
abstract
Rapid advancements in generative models, including Generative Adversarial Networks (GANs) and diffusion models, have made possible automated image editing through the use of text descriptions, semantic segmentation, and/or reference style images. Nevertheless, in terms of fashion image editing, it often requires more flexible, and typically iterative, modifications to the image content that existing methods struggle to achieve. This article proposes a new model called Content Decoupled and Enhanced GAN (CoDE-GAN), which is formulated and trained for the task of image editing, drawing on methods from image reconstruction, more specifically, image inpainting with sketch guidance. Through this proxy task, the trained model can be used for flexible image editing, generating new images with consistent colors and required textures based on sketch inputs. In this new model, a content decoupling block is introduced including specially designed dual encoders, which pre-process inputs and transform into separated structure and texture representations. Moreover, a content enhancing module is designed and applied to the decoder, improving the color consistency and refining the texture of the generated images. The proposed CoDE-GAN can achieve coarse-to-fine results in one single stage. Extensive experiments on three datasets, covering human, garment-only, and scene images, show that CoDE-GAN outperforms other state-of-the-art methods in terms of both generated image quality and editing flexibility. The code and dataset are available at https://github.com/Taited/CoDE-GAN .
Zhengwentai Sun, Yanghong Zhou, P. Y. Mok 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2025 ViTon-GUN: Person-to-Person Virtual Try-on via Garment Unwrapping
abstract
The image-based Person-to-Person (P2P) virtual try-on, involving the direct transfer of garments from one person to another, is one of the most promising applications of human-centric image generation. However, existing approaches struggle to accurately learn the clothing deformation when directly warping the garment from the source pose onto the target pose. To address this, we propose Person-to-Person virtual try-on via Garment UNwrapping, a novel framework dubbed as ViTon-GUN. Specifically, we divide the P2P task into two subtasks: Person-to-Garment (P2G) and Garment-to-Person (G2P). The P2G aims to unwrap the target garment from a source pose to a canonical representation based on A-Pose. In the P2G stage, we enable the implementation of a flow-based P2G scheme by introducing an A-Pose estimator and establishing comprehensive training conditions. Building upon this step-wise strategy, we introduce a novel pipeline for P2P try-on. Once trained, the P2G strategy can serve as a "plug-and-play" module, which efficiently adapts existing diffusion-based pre-trained G2P models to P2P try-on without further training. Quantitative and qualitative experiments demonstrate that our ViTon-GUN performs remarkably well on P2P try-on, even for dresses with intricate design details.
Zhenyu Xie, Zhengwentai Sun, Hairui Zhu, Zirong Jin, Xiaoguang Han 0001
IEEE Trans. Vis. Comput. Graph.3
2024 A 3D Virtual Try-On Method with Global-Local Alignment and Diffusion Model
abstract
3D virtual try-on has recently received more attention due to its great practical and commercial value. However, there remains the problems that the garment cannot accurately correspond to a human body by geometric transformation and abnormal textures may be produced in the synthesis result. To address these issues, a 3D virtual try-on method with global-local alignment and diffusion model is proposed. The global-local alignment module is designed for more accurate garment warping, combined with the guidance of an "after-try-on" semantic map alignment. Then, a diffusion model is introduced to design a try-on synthesizer that can not only avoid producing abnormal textures but also enhance the quality of textures produced in edge regions. Experiments on existing datasets show that our method outperforms the state-of-the-art methods. Code is available at https://github.com/Breaveh/VTON-GD.
Shougan Pan, Zhengwentai Sun, Chenxing Wang 0002
ICASSP2
2023 SGDiff: A Style Guided Diffusion Model for Fashion Synthesis
abstract
This paper reports on the development of a novel style guided diffusion model (SGDiff) which overcomes certain weaknesses inherent in existing models for image synthesis. The proposed SGDiff combines image modality with a pretrained text-to-image diffusion model to facilitate creative fashion image synthesis. It addresses the limitations of text-to-image diffusion models by incorporating supplementary style guidance, substantially reducing training costs, and overcoming the difficulties of controlling synthesized styles with text-only inputs. This paper also introduces a new dataset -- SG-Fashion, specifically designed for fashion image synthesis applications, offering high-resolution images and an extensive range of garment categories. By means of comprehensive ablation study, we examine the application of classifier-free guidance to a variety of conditions and validate the effectiveness of the proposed model for generating fashion images of the desired categories, product attributes, and styles. The contributions of this paper include a novel classifier-free guidance method for multi-modal feature fusion, a comprehensive dataset for fashion image synthesis application, a thorough investigation on conditioned text-to-image synthesis, and valuable insights for future research in the text-to-image synthesis domain. The code and dataset are available at: https://github.com/taited/SGDiff.
Zhengwentai Sun, Yanghong Zhou, Honghong He, P. Y. Mok 0001
ACM Multimedia1
2021 Automatic segmentation of organs-at-risk from head-and-neck CT using separable convolutional neural network with hard-region-weighted loss
Wenhui Lei, Haochen Mei, Zhengwentai Sun, Shan Ye, Ran Gu, Huan Wang 0015, Rui Huang 0001, Shichuan Zhang, Shaoting Zhang 0001, Guotai Wang
Neurocomputing3
2021 Automatic segmentation of gross target volume of nasopharynx cancer using ensemble of multiscale deep neural networks with spatial attention
Haochen Mei, Wenhui Lei, Ran Gu, Shan Ye, Zhengwentai Sun, Shichuan Zhang, Guotai Wang
Neurocomputing5