VLDB 2026 Research / reviewers in the wild / expert
Chenghu Du
dblp:309/4642
· DBLP profile ↗
26ranked-venue papers
14as first author
26since 2021 · last 2026
0000-0001-7275-5064ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 11 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SLOcc: Selective interaction and long-range modelling for occupancy prediction
Junyin Wang, Chenghu Du, Tongao Ge, Shengwu Xiong 0001 |
Expert Syst. Appl. | 2 |
| 2026 | OV-Pro: Enhancing open-vocabulary 3D object detection by prototype contrastive distillation
Tongao Ge, Junyin Wang, Chenghu Du, Hui Li 0010, Huikai Liu, Shengwu Xiong 0001 |
Pattern Recognit. | 3 |
| 2026 | D3PD: Dual distillation and dynamic fusion for camera-radar 3D perception
Junyin Wang, Chenghu Du, Tongao Ge, Bingyi Liu, Shengwu Xiong 0001 |
Pattern Recognit. | 2 |
| 2026 | RDNet: Rotate-Groundtruth Augmentation and Decoupled Attention HEAD for 3D Object DetectionabstractLiDAR is one of the most important sensors in the field of autonomous driving and allows for better and more accurate perception of the changes in the surrounding environment. Most of the existing 3D object detection methods use data augmentation and feature fusion enhancement to improve the performance of detection, but the majority of the methods ignore the handling of sample imbalance problems during data augmentation. Also, the designed feature fusion and enhancement methods were not well suited to work with the enhancement methods. To this end, we developed a combined method involving data augmentation and feature enhancement. The designed approach has two main objectives: 1) to address the problem of unbalanced sample distribution in detection scenes through data augmentation, and 2) to enhance feature perception using a special feature enhancement module. Our proposed method solves the problem of class imbalance by directly increasing the number of pedestrian samples in the scene through mixed data augmentation, i.e., RG-Aug. In addition, we introduce the Decoupling and Attention Fusion module (DAF), which combines classification headers with high-level features and prediction branches with low-level features. Leverage data features between different layers of features to get a more robust feature representation. Finally, the multi-scale pyramid attention enhancement module is designed to achieve feature enhancement of multi-scale features by means of attention to improve the detection ability of small objects in the scene, especially the detection ability of pedestrians. Our method can achieve 1.57%, 2.16%, and 2.05% performance improvement on the KITTI dataset for Easy, Mod, and Hard samples, respectively. Furthermore, for the detection of pedestrians, our method has a significant competitive advantage over other state-of-the-art techniques with a mAP of 73.42%. Zhenchang Xia, Guanqun Zheng, Shengwu Xiong 0001, Junyin Wang, Jianqun Cui, Yanan Chang, Chenghu Du, Jia Wu 0001 |
IEEE Trans. Big Data | 7 |
| 2025 | Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label GenerationabstractEfficiently applying fully supervised learning to virtual try-on tasks is challenging due to the lack of paired ground truth in available training samples. Recent works have achieved virtual try-ons by employing self-supervised learning-based inpainting paradigms. However, this approach is heavily dependent on the constraints of inpainting masks. An incorrect mask can mislead the generated results, while overly large mask areas can lose essential original information, thereby hindering the synthesis of high-quality results. To address these problems, we propose a latent diffusion model-based virtual try-on network that achieves fully supervised learning using the concept of cycle consistency and knowledge distillation. Specifically, we divide our approach into pretext and downstream tasks. In the pretext task, we generate a pseudo-label (pseudo-person image) to form paired training samples, which enables the downstream task to achieve fully supervised learning. To prevent the unreliable pseudo-person image from introducing irresponsible prior knowledge, we propose a noise-covering strategy, which aims at fully optimizing the pseudo-label to eliminate the impact of the incorrect inpainting mask as much as possible. Additionally, we propose a skin refinement loss to further enhance the generation of details in the skin region. Extended experiments demonstrate that our proposed method is superior to state-of-the-art methods. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 1 |
| 2025 | GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free ApproachabstractA good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-decoder, ensuring faster speeds. To tackle the problem of lower image quality typically generated by lighter models, we present GarFast, a simplified, parser-free framework that optimizes the same lightweight network through a two-stage transformation of real data roles (from input to supervision), thereby greatly promoting model convergence. Specifically, first, we propose a correction strategy to prevent the difficulty of convergence caused by the lack of ground truth in the first stage. Second, we propose a fine-grained domain consistency to ensure that the results generated in the unsupervised first stage are highly realistic clothed human images. Finally, we propose a skin-variant refinement loss and a skinMix regularization to amplify texture differences and enhance the realism of skin-variant regions, thereby improving the quality of the generated skin. Extensive experiments thoroughly demonstrate that our method achieves high resolution, near real-time performance, and superior reconstruction quality compared to state-of-the-art approaches, with processing times of less than 0.03 seconds on an Nvidia A100. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 1 |
| 2025 | ROME: Radar Sparsity Improvement and Omnimodal Enhancement for 3D Object Detection in Bird's Eye ViewsabstractCombining omnimodal feature interaction using LiDAR, surround-view camera, and Radar to form a network has a great guarantee for the safety of autonomous driving, but most of the current omnimodal fusion methods focus on the interaction enhancement of LiDAR and surround-view camera, ignoring the focus on Radar. Enhancing the contextual representation of Radar can ensure better all-weather capability of the perceptual network. To this end, we design the ROME method based on Radar sparsity improvement to better enhance the performance and robustness of the model in terms of alleviating Radar sparsity shortcomings. Firstly, we design the Autocorrelation Point Enhancement (APE) module to improve Radar sparsity leveraging the point-to-point autocorrelation of Radar. Moreover, for omnimodal Bird’s Eye View (BEV) features, an Omnimodal Adaptive Fusion (OAF) module is designed to improve the robustness of BEV features. With the improved Radar modality, the performance of BEV features for the whole driving scene is further improved. Comprehensive experiments on the nuScenes dataset and comparisons with state-of-the-art methods demonstrate the advantages of our proposed method. Yilong Guo, Junyin Wang, Chenghu Du, Shengwu Xiong 0001, Yaxiong Chen |
ICASSP | 3 |
| 2025 | All Parts Matter: A Unified Mask-Free Virtual Try-On Framework
Chenghu Du, Shengwu Xiong 0001 |
ICCV | 1 |
| 2025 | Enhancing Virtual Try-On with Text-Image Fusion Guidance
Jingyi Guo, Pengfei Duan 0005, Chenghu Du, Shengwu Xiong 0001 |
ICIC (21) | 3 |
| 2025 | Mask Does Not Matter: A Unified Latent Diffusion-Enhanced Framework for Mask-Free Virtual Try-OnabstractA good virtual try-on model should introduce minimal redundant conditional information to avoid instability and increase inference efficiency. Existing methods rely on inpainting masks to guide the generation of the object, but the masks, generated by unstable human parsers, often produce unreliable results with fabric residues due to wrong segmentation. Moreover, large mask regions can lose spatial structure and identity information, requiring extra conditional inputs to compensate, which increases model instability and reduces efficiency. To tackle the problem, we present a novel Mask-Free virtual Try-ON (MFTON) framework. Specifically, we propose a mask-free strategy to eliminate all denoising conditions except for clothing and person images, thereby directly extracting spatial structure and identity information from the person image to improve efficiency and reduce instability. Additionally, to optimize the generated clothing regions, we propose a clothing texture-aware attention mechanism to enable the model to focus on texture generation with significant visual differences. We then introduce a geometric detail capture loss to further enable the model to capture more high-frequency information. Finally, we propose an appearance consistency inference method to reduce the initial randomness of the sampling process significantly. Extensive experiments on popular datasets demonstrate that our method outperforms state-of-the-art virtual try-on methods. Chenghu Du, Junyin Wang, Shengwu Xiong 0001 |
IJCAI | 1 |
| 2025 | Mitigating Occlusions in Virtual Try-On via A Simple-Yet-Effective Mask-Free FrameworkabstractThis paper investigates the occlusion problems in virtual try-on (VTON) tasks. According to how they affect the try-on results, the occlusion issues of existing VTON methods can be grouped into two categories: (1) Inherent Occlusions, which are the ghosts of the clothing from reference input images that exist in the try-on results. (2) Acquired Occlusions, where the spatial structures of the generated human body parts are disrupted and appear unreasonable. To this end, we analyze the causes of these two types of occlusions, and propose a novel mask-free VTON framework based on our analysis to deal with these occlusions effectively. In this framework, we develop two simple-yet-powerful operations: (1) The background pre-replacement operation prevents the model from confusing the target clothing information with the human body or image background, thereby mitigating inherent occlusions. (2) The covering-and-eliminating operation enhances the model's ability of understanding and modeling human semantic structures, leading to more realistic human body generation and thus reducing acquired occlusions. Moreover, our method is highly generalizable, which can be applied in in-the-wild scenarios, and our proposed operations can also be easily integrated into different generative network architectures (e.g., GANs and diffusion models) in a plug-and-play manner. Extensive experiments on three VTON datasets validate the effectiveness and generalization ability of our method. Both qualitative and quantitative results demonstrate that our method outperforms recently proposed VTON benchmarks. Chenghu Du, Shengwu Xiong 0001, Junyin Wang, Shili Xiong |
NeurIPS | 1 |
| 2025 | GLV: Geometric Correlation Distillation for Latent Diffusion-Enhanced Parser-Free Virtual Try-OnabstractApplying knowledge distillation to virtual try-on tasks is challenging because current methods fail to fully and efficiently exploit responsible teacher knowledge. In other words, existing approaches merely transfer prior knowledge to the student model via pseudo-labels generated by the teacher model, resulting in shallow knowledge representation and low training efficiency. To address these limitations, we propose a novel teacher-student architecture for parser-free virtual try-on, named GLV, which generates high-quality try-on results with realistic body details. Specifically, we propose a deformation-related prior distillation method to effectively leverage the valuable deformation information contained in the teacher warpage model. This enhances the convergence efficiency of the student warpage model, preventing it from getting stuck in a local minima. Moreover, we are the first to propose a geometric correlation distillation, which models the underlying geometric relationship between clothing and the person and transfers this relationship from the teacher to the student. This enables the student warpage model to reduce the entanglement of deformation-irrelevant features, such as color and texture. Finally, we propose a clothing-body retouching method for try-on result synthesis, which refines the denoising process in the latent space of a well-trained diffusion model, thereby preventing catastrophic forgetting. This method seamlessly transforms the parser-based inpainting synthesis paradigm into a parser-free synthesis paradigm and enables efficient convergence of the diffusion model with only fine-tuning. Extensive experiments demonstrate the generality of our approach and highlight its superiority over previous methods. Chenghu Du, Junyin Wang, Shengwu Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | HybridBEV: Hybrid Encode and Distillation for Improved BEV 3D Object DetectionabstractThe development of surround-view cameras is crucial for the advancement of autonomous driving. Utilizing depth information and image features to simulate LiDAR bird’s-eye-view (BEV) features can accomplish efficient 3D object detection tasks. Existing dense BEV generation methods heavily rely on the use of depth features, however, the suboptimal exploitation of these features often results in ambiguity in object location and feature representation during the BEV generation process. To address this, we have designed a hybrid encode and distillation method to enhance 3D object detection performance, termed HybridBEV. Initially, we designed the HybridEncode module, which employs a resampling strategy of depth features in voxel space to obtain BEV features that more accurately reflect the distribution of objects. Subsequently, we introduced multiple distillation methods to supervise the network’s voxel features and BEV feature representations, assisting the student network in learning critical features from the teacher model and ensuring that BEV features can more distinctly represent object distribution. Furthermore, during network training, we loaded pre-trained weights from the teacher network to guide network optimization and accelerate training. Extensive experiments on the nuScenes benchmark demonstrate that HybridBEV can effectively improve the performance of the student network and outperform previous state-of-the-art methods based on surround-view cameras. The code will be published athttps://github.com/wjyxx/HybridBEV Junyin Wang, Chenghu Du, Huikai Liu, Zhenchang Xia, Bingyi Liu, Shengwu Xiong 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | CycleVTON: A Cycle Mapping Framework for Parser-Free Virtual Try-OnabstractImage-based virtual try-on aims to transfer a target clothing onto a specific person. A significant challenge is arbitrarily matched clothing and person lack corresponding ground truth to supervised learning. A recent pioneering work leveraged an improved cycleGAN to enable one network to generate the desired image for another network during training. However, there is no difference in the result distribution before and after the clothing changes. Therefore, using two different networks is unnecessary and may even increase the difficulty of convergence. Furthermore, the introduced human parsing used to provide body structure information in the input also have a negative impact on the try-on result. How to employ a single network for supervised learning while eliminating human parsing? To tackle these issues, we present a Cycle mapping Virtual Try-On Network (CycleVTON), which can produce photo-realistic try-on results by using a cycle mapping framework without the parser. In particular, we introduce a flow constraint loss to achieve supervised learning of arbitrarily matched clothing and person as inputs to the deformer, thus naturally mimicking the interaction between clothing and the human body. Additionally, we design a skin generation strategy that can adapt to the shape of the target clothing by dynamically adjusting the skin region, i.e., by first removing and then filling skin areas. Extensive experiments conducted on challenging benchmarks demonstrate that our proposed method exhibits superior performance compared to state-of-the-art methods. Chenghu Du, Junyin Wang, Shuqing Liu, Shengwu Xiong 0001 |
AAAI | 1 |
| 2024 | IFNET: Integrating Data Augmentation and Decoupled Attention Fusion for 3D Object DetectionabstractLiDAR is a key sensor for accurately sensing of the environment in autonomous driving. While existing 3D object detection methods generally rely on data augmentation and feature fusion to improve performance, the challenge of dealing with sample imbalance is often overlooked. We design a novel 3D detection network, IFNet, that tackles these issues by introducing mutually reinforcing data augmentation and feature enhancement strategies. It aims to achieve a dual purpose: 1) correcting the category imbalance by directly enhancing pedestrian samples using mixed data augmentation, i.e., RG-Aug; and 2) enhancing feature perception by introducing the decoupling and attention fusion module (DAF). DAF enables robust feature representations across different layers, improving the detection performance, especially for small objects in the scene. Comprehensive experiments on the KITTI dataset and comparisons with state-of-the-art methods demonstrate the superiority of our proposed approach. Zhenchang Xia, Guanqun Zheng, Shengwu Xiong 0001, Jia Wu 0001, Junyin Wang, Chenghu Du |
ICASSP | 6 |
| 2024 | SterAF: A Scene Text Recognizer with Appearance-Flow RectificationabstractAs the demand for recognizing irregular text in natural scenes increases, people are increasingly realizing the value of such applications, such as license plate recognition systems, image search, handwriting recognition, and autonomous driving, which are profoundly changing our lives in the field of text recognition. Recent studies have shown that the recognition of curved text and perspective text has become an important challenge in the field of text recognition, and the correction of curved text is a key step to achieve accurate recognition. However, current methods use strained text image correction methods, resulting in poor recognition accuracy when recognizing curved text. Therefore, we propose an end-to-end framework called Scene Text Recognizer with Appearance-Flow rectification (SterAF), which includes a correction network and a recognition network. Specifically, the framework’s steps are as follows: first, the input text image is deformed through an appearance flow-based correction network to adaptively warp the text image, to prevent irregular and unnatural deformations of the text image. Second, a sequence-to-sequence recognition network predicts the sequence of characters in the corrected text image to accurately recognize the text in the image. Through subjective and objective experiments, our SterAF model has shown excellent performance in both qualitative and quantitative experiments. Chunyan Liao, Chenghu Du, Yanbao Tan |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | Intelligent Wearable System With Motion and Emotion Recognition Based on Digital Twin TechnologyabstractIntelligent wearable systems have been widely used in health monitoring, motion tracking, and engineering safety. However, the single function of current wearable systems cannot satisfy the requirements of complex scenarios, and the wearable systems cannot establish a relationship with the virtual 3D visualization platform. To address these issues, this paper proposes a novel intelligent wearable system with motion and emotion recognition. Multiple sensors are integrated into the system to collect motion and emotion information. In order to achieve accurate classification and recognition of multiple sensor information, we propose a novel human action recognition network called the three-branch spatial-temporal feature extraction network (TB-SFENet), which can obtain more robust features and achieve an accuracy of 97.04% on the UCI-HAR dataset and 92.68% on the UniMiB SHAR dataset. To establish the relationship between the real entity and virtual space, we use digital twin (DT) technology to establish the 3D display DT platform. The platform enables real-time information interaction, such as activity, emotion, location, and monitoring information. Additionally, we establish the TGAM electroencephalogram emotion classification (TEEC) dataset, which contains 120,000 pieces of data, for the proposed system. Experimental results indicate that the proposed system realizes virtual reality information interaction between the personal digital human and actual person based on the intelligent wearable system, which has great potential for applications in intelligent healthcare, virtual reality, and other fields. Feng Yu 0017, Chenyu Yu, Zhangyuan Tian, Jiacheng Cao, Li Liu 0047, Chenghu Du, Minghua Jiang |
IEEE Internet Things J. | 7 |
| 2023 | COCCI: Context-Driven Clothing Classification Network
Minghua Jiang, Shuqing Liu, Yankang Shi, Chenghu Du, Guangyu Tang, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017 |
CGI (1) | 4 |
| 2023 | CF-VTON: Multi-Pose Virtual Try-on with Cross-Domain FusionabstractThe multi-pose virtual try-on technology aims to seamlessly fit an in-shop garment onto a reference person in various poses. This technology has attracted considerable attention from researchers due to its potential commercial and practical applications. Previous works in this field have encountered issues such as unnatural garment alignment and difficulty in preserving the person’s identity, arising from the weak mapping relationship between different feature crosses. To address these challenges, this paper proposes a novel multi-pose virtual try-on network named CF-VTON. Our approach involves predicting the "after-try-on" semantic map to guide garment alignment and try-on synthesis, warping the garment using an improved garment alignment network (GANet) to optimize unnatural alignment, synthesizing a coarse result with our proposed try-on synthesis network (TSN), and refining the output to reconstruct the virtual try-on result with rich facial identity and garment details. Qualitative and quantitative experiments demonstrate the superiority of our approach, outperforming state-of-the-art methods in an efficient manner. Chenghu Du, Shengwu Xiong 0001 |
ICASSP | 1 |
| 2023 | DLFusion: Painting-Depth Augmenting-LiDAR for Multimodal Fusion 3D Object DetectionabstractSurround-view cameras combined with image depth transformation to 3D feature space and fusion with point cloud features are highly regarded. The transformation of 2D features into 3D feature space by means of predefined sampling points and depth distribution happens throughout the scene, and this process generates a large number of redundant features. In addition, multimodal feature fusion unified in 3D space often happens in the previous step of the downstream task, ignoring the interactive fusion between different scales. To this end, we design a new framework, focusing on the design that can give 3D geometric perception information to images and unify them into voxel space to accomplish multi-scale interactive fusion, and we mitigate feature alignment between modal features by geometric relationships between voxel features. The method has two main designs. First, a Segmentation-guided Image View Transformation module is used to accurately transform the pixel region containing the object into a 3D pseudo-point voxel space with the help of a depth distribution. This allows subsequent feature fusion to be performed in a unified voxel feature. Secondly, a Voxel-centric Consistent Fusion module is used to alleviate the errors caused by depth estimation, as well as to achieve better feature fusion between unified modalities. Through extensive experiments on the KITTI and nuScenes datasets, we validate the effectiveness of our camera-LIDAR fusion method. Our proposed approach shows competitive performance on both datasets and outperforms state-of-the-art methods in certain classes of 3D object detection benchmarks. https://github.com/no-Name128/DLFusion [code release] Junyin Wang, Chenghu Du, Hui Li 0010, Shengwu Xiong 0001 |
ACM Multimedia | 2 |
| 2023 | Greatness in Simplicity: Unified Self-Cycle Consistency for Parser-Free Virtual Try-OnabstractImage-based virtual try-on tasks remain challenging, primarily due to inherent complexities associated with non-rigid garment deformation modeling and strong feature entanglement of clothing within human body. Recent groundbreaking formulations, such as in-painting, cycle consistency, and knowledge distillation, have facilitated self-supervised generation of try-on images. However, these paradigms necessitate the disentanglement of garment features within human body features through auxiliary tasks, such as leveraging 'teacher knowledge' and dual generators. The potential presence of irresponsible prior knowledge in the auxiliary task can serve as a significant bottleneck for the main generator (e.g., 'student model') in the downstream task. Moreover, existing garment deformation methods lack the ability to perceive the correlation between the garment and the human body in the real world, leading to unrealistic alignment effects. To tackle these limitations, we present a new parser-free virtual try-on network based on unified self-cycle consistency (USC-PFN), which enables robust translation between different garments using just a single generator, faithfully replicating non-rigid geometric deformation of garments in real-life scenarios. Specifically, we first propose a self-cycle consistency architecture with a circular mode. It utilizes real unpaired garment-person images exclusively as input for training, effectively eliminating the impact of irresponsible prior knowledge at the model input end. Additionally, we formulate a Markov Random Field to simulate a more natural and realistic garment deformation. Furthermore, USC-PFN can leverage a general generator for self-supervised cycle training. Experiments demonstrate that our method achieves state-of-the-art performance on a popular virtual try-on benchmark. Chenghu Du, Junyin Wang, Shuqing Liu, Shengwu Xiong 0001 |
NeurIPS | 1 |
| 2023 | VTON-SCFA: A Virtual Try-On Network Based on the Semantic Constraints and Flow AlignmentabstractAn image-based virtual try-on system transfers an in-shop garment to the corresponding garment region of a reference person, which has huge application potential and commercial value in online clothing shopping. Existing methods have difficulty preserving garment texture and body details because of rough garment alignment and imperfect detail-retention strategies. To address this problem, we propose a virtual try-on network based on semantic constraints and flow alignment. The key idea of the framework is as follows: 1) a global-local semantic predictor (GLSP) is proposed to generate a reasonable target semantic map, which clearly guides the correct alignment of the in-shop garment with the body and the generation of try-on result; and 2) a novel appearance flow-based garment alignment network (AFGAN) is proposed to align the in-shop garment with the body, which is important to preserve maximum garment detail and ensure natural and realistic warping; and 3) we propose a synthesis strategy to integrate the aligned garment and the human body to preserve maximum body detail for generating a realistic result and preventing cross-occlusion and pixel confusion between different body parts. Experiments on the existing benchmark dataset demonstrate that the proposed method achieves the best performance on qualitative and quantitative experiments among the state-of-the-art virtual try-on techniques. Chenghu Du, Feng Yu 0017, Minghua Jiang, Ailing Hua, Tao Peng 0006, Xinrong Hu |
IEEE Trans. Multim. | 1 |
| 2022 | Multi-Pose Virtual Try-On Via Self-Adaptive Feature FilteringabstractWith the growing trend of virtual try-on, multi-pose tasks attract researchers due to their higher commercial value. Prior methods lack an effective geometric deformation to maintain the original image details resulting in many details loss in the head and garment. To address this problem, we propose a new multi-pose virtual try-on network, which can fit a garment to the corresponding area of a person in arbitrary poses. First, the target pose’s body-semantic distribution is predicted by the target pose point. Second, the in-shop garment and human body are warped based on a human pose to solve the unnatural alignment and the lack of body details by the Deformation Module (DM). Finally, the human body in the given pose and garment is fine generated by the Filtering Synthesis Network (FSN). Compared to state-of-the-art methods with objective experiments on the MPV dataset, the proposed method achieves the best performance in metrics and the rich details in visual results. Chenghu Du, Feng Yu 0017, Minghua Jiang, Tao Peng 0006, Xinrong Hu |
ICASSP | 1 |
| 2022 | Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Captureabstract3D virtual try-on receives widespread attention from scholars due to its great practical and commercial values. In prior methods, the fundamental problems lie in the limitations on texture retention during garment deformation and the lack of feature context capture during depth estimation. To address these problems, we propose a new 3D virtual try-on network via multi-scale characteristic capture (VTON-MC), which can produce an exact 3D model with the generated photo-realistic monocular image. The main processes are as follows: 1) predicting the human semantic-map and aligning the in-shop garment in the human pose using the appearance flow method, 2) synthesizing the human body and the warped garment to gain the image try-on result, and 3) estimating the human double-depth map of the image try-on result to reconstruct desired 3D try-on mesh by designed Depth Estimation Network (DEN). Extensive experiments on existing benchmark datasets demonstrate that VTON-MC outperforms state-of-the-art approaches efficiently. Chenghu Du, Feng Yu 0017, Minghua Jiang, Yaxin Zhao, Tao Peng 0006, Xinrong Hu |
ICASSP | 1 |
| 2022 | High fidelity virtual try-on network via semantic adaptation and distributed componentizationabstractImage-based virtual try-on systems have significant commercial value in online garment shopping. However, prior methods fail to appropriately handle details, so are defective in maintaining the original appearance of organizational items including arms, the neck, and in-shop garments. We propose a novel high fidelity virtual try-on network to generate realistic results. Specifically, a distributed pipeline is used for simultaneous generation of organizational items. First, the in-shop garment is warped using thin plate splines (TPS) to give a coarse shape reference, and then a corresponding target semantic map is generated, which can adaptively respond to the distribution of different items triggered by different garments. Second, organizational items are componentized separately using our novel semantic map-based image adjustment network (SMIAN) to avoid interference between body parts. Finally, all components are integrated to generate the overall result by SMIAN. A priori dual-modal information is incorporated in the tail layers of SMIAN to improve the convergence rate of the network. Experiments demonstrate that the proposed method can retain better details of condition information than current methods. Our method achieves convincing quantitative and qualitative results on existing benchmark datasets. Chenghu Du, Feng Yu 0017, Minghua Jiang, Ailing Hua, Yaxin Zhao, Tao Peng 0006, Xinrong Hu |
Comput. Vis. Media | 1 |
| 2021 | VTON-HF: High Fidelity Virtual Try-on Network via Semantic AdaptationabstractThe image-based virtual try-on network transfers the target garment item to the corresponding region of the human body. Due to its commercial value in online garment shopping, it has attracted extensive attention from researchers. However, the previous virtual try-on methods are interfered heavily by garments in reference images, so they have defects in maintaining details of human upper limbs, neck, and given garment. Therefore, a novel High Fidelity Virtual Try-on Network via Semantic Adaptation (VTON-HF) is proposed to generate a result with better details. The main processes are as follows: 1) Thin Plate Spline (TPS) warps the target garment coarsely, 2) parsing network generates a target semantic map with the coarse warped garment, 3) our novel Semantic Map-based Image Adjustment Network (SMIAN) generates components separately to avoid interference between image parts with different semantics, 4) SMIAN fuses all components to generate the final result. VTON-HF can retain the maximum amount of detail in the reference garment than previous methods. Our novel architecture generates desired results by fusing separately generated components (garment, upper limb, and neck) and unchanging parts of the reference image. Moreover, our SMIAN incorporates a priori multimodal information in the tail layer, which effectively improves the convergence efficiency of the network. Our method achieves state-of-the-art quantitative results on IS, SSIM, PSNR, and FID using the VITON dataset. (see Fig. 1). Chenghu Du, Feng Yu 0017, Minghua Jiang, Tao Peng 0006, Xinrong Hu |
ICTAI | 1 |