EDBT 2026 Demo / reviewers in the wild / expert
Feng Yu 0017
dblp:28/1708-17
· DBLP profile ↗
40ranked-venue papers
10as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 29 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Computer networks · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency domain feature enhancement network for clothing semantic segmentation
Feng Yu 0017, Jianhang Zhu, Jiaolong Wan, Li Liu 0047, Minghua Jiang |
Expert Syst. Appl. | 1 |
| 2026 | DFFENet: Dual-Branch Frequency Domain Feature Enhancement Network for Skin Lesion Classification
Feng Yu 0017, Yuyu Jin, Li Liu 0047, Minghua Jiang |
Image Vis. Comput. | 1 |
| 2026 | DFENet: dual-frequency feature enhancement network for breast tumor classification
Jiacheng Cao, Yuyu Jin, Ziheng Cai, Li Liu 0047, Feng Yu 0017, Minghua Jiang |
Vis. Comput. | 6 |
| 2025 | Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label GenerationabstractEfficiently applying fully supervised learning to virtual try-on tasks is challenging due to the lack of paired ground truth in available training samples. Recent works have achieved virtual try-ons by employing self-supervised learning-based inpainting paradigms. However, this approach is heavily dependent on the constraints of inpainting masks. An incorrect mask can mislead the generated results, while overly large mask areas can lose essential original information, thereby hindering the synthesis of high-quality results. To address these problems, we propose a latent diffusion model-based virtual try-on network that achieves fully supervised learning using the concept of cycle consistency and knowledge distillation. Specifically, we divide our approach into pretext and downstream tasks. In the pretext task, we generate a pseudo-label (pseudo-person image) to form paired training samples, which enables the downstream task to achieve fully supervised learning. To prevent the unreliable pseudo-person image from introducing irresponsible prior knowledge, we propose a noise-covering strategy, which aims at fully optimizing the pseudo-label to eliminate the impact of the incorrect inpainting mask as much as possible. Additionally, we propose a skin refinement loss to further enhance the generation of details in the skin region. Extended experiments demonstrate that our proposed method is superior to state-of-the-art methods. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 3 |
| 2025 | GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free ApproachabstractA good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-decoder, ensuring faster speeds. To tackle the problem of lower image quality typically generated by lighter models, we present GarFast, a simplified, parser-free framework that optimizes the same lightweight network through a two-stage transformation of real data roles (from input to supervision), thereby greatly promoting model convergence. Specifically, first, we propose a correction strategy to prevent the difficulty of convergence caused by the lack of ground truth in the first stage. Second, we propose a fine-grained domain consistency to ensure that the results generated in the unsupervised first stage are highly realistic clothed human images. Finally, we propose a skin-variant refinement loss and a skinMix regularization to amplify texture differences and enhance the realism of skin-variant regions, thereby improving the quality of the generated skin. Extensive experiments thoroughly demonstrate that our method achieves high resolution, near real-time performance, and superior reconstruction quality compared to state-of-the-art approaches, with processing times of less than 0.03 seconds on an Nvidia A100. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 4 |
| 2025 | SuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor SegmentationabstractMultimodal 3D segmentation involves a significant number of 3D convolution operations, which requires substantial computational resources and high-performance computing devices in MRI multimodal brain tumor segmentation. The key challenge in multimodal 3D segmentation is how to minimize network computational load while maintaining high accuracy. To address the issue, a novel lightweight parameter aggregation network (SuperLightNet) is proposed to realize the efficient encoder and decoder for the high accurate and low computation. A random multiview drop encoder is designed to learn the spatial structure of multimodal images through a random multi-view approach for solving the high computational time complexity that has arisen in recent years with methods relying on transformers and Mamba. A learnable residual skip decoder is designed to incorporate learnable residual and group skip weights for addressing the reduced computational efficiency caused by the use of overly heavy convolution and deconvolution decoders. Experimental results demonstrate that the proposed method achieves a leading reduction in parameter count by 95.59%, the 96.78% improvement in computational efficiency, the 96.86% enhancement in memory access performance, and the average performance gain of 0.21% on the BraTS2019 and BraTS2021 datasets in comparison with the state-of-the-art methods. Code is available at https://github.com/WTU-MIS-Laboratory/SuperLightNet. Feng Yu 0017, Jiacheng Cao, Li Liu 0047, Minghua Jiang |
CVPR | 1 |
| 2025 | BiaCanDet: Bioelectrical impedance analysis for breast cancer detection with space-time attention neural network
Feng Yu 0017, Zhiyong Xiao 0003, Li Liu 0047, Man Tang, Minghua Jiang, Jinxuan Hou |
Expert Syst. Appl. | 1 |
| 2025 | Multimodal Wearable System With Dual-Frequency Enhancement Network for Risk RecognitionabstractSmart wearable systems can monitor users’ physiological data in real time, detect anomalies promptly through risk recognition technologies, provide early warnings, and assist users in taking preventive measures. However, single modal information is difficult to accurately recognize the behavioral state, expression state, and environmental conditions. Furthermore, multimodal data are often affected by noise and interference, complicating the accurate identification of risky behaviors. To address these challenges, we propose a smart wearable system based on the dual-frequency enhancement network (DFENet): 1) the multimodal sensor system is designed to combine behavioral recognition, expression recognition, and environmental recognition for comprehensive monitoring and recognition of multidimensional risk factors in complex scenarios; 2) the DFENet is proposed to overcome challenges in feature extraction and accurate classification in complex environments; and 3) the behavioral recognition dataset and the expression recognition dataset are built to verify the effectiveness of the designed smart wearable system. Experimental results indicate that the proposed system can real-time achieve risk recognition across physical activity, expression state, and environmental conditions, and the proposed DFENet achieves excellent performance in accuracy, parameters, and floating-point operations (FLOPs) metrics on the three datasets. The algorithm and datasets can be downloaded athttps://github.com/wtu1020/Multimodal-Wearable. Feng Yu 0017, Hanchen Yu, Li Liu 0047, Minghua Jiang |
IEEE Internet Things J. | 1 |
| 2025 | ArmBCIsys: Robot Arm BCI System With Time-Frequency Network for Multiobject GraspingabstractBrain-computer interface (BCI) offers a direct communication and control channel between the human brain and external devices, presenting new pathways for individuals with physical disabilities to operate robotic arms for complex tasks. However, achieving multiobject grasping tasks under low signal-to-noise ratio (SNR) consumer-grade EEG signals is a significant challenge due to the lack of robust decoding algorithms and precise visual tracking methods. This article proposes, ArmBCIsys, an integrated robotic arm system that combines a novel dual-branch frequency-enhanced network (DBFENet) to robustly decode EEG signals under noisy conditions with the high-precision vision-guided grasping module. The proposed DBFENet designs the scaling temporal convolution block (STCB) to extract multiscale spatiotemporal features from the time domain, while the designed DropScale projected Transformer (DSPT) utilizes discrete cosine transform (DCT) to capture key frequency-domain features, significantly improving decoding robustness. We fine-tune the masked-attention mask Transformer (Mask2Former) model on the Jacquard dataset and incorporate the multiframe centroid-intersection over union (IoU) tracking algorithm to build visual grasp segmenter (VisGraspSeg), enabling reliable segmentation and dynamic tracking for diverse daily objects. Experimental validations on both self-built code-modulated visual evoked potential (c-VEP) dataset (1344 samples) and two public c-VEP datasets demonstrate that DBFENet achieves the state-of-the-art recognition performance, and the system integrates the DBFENet and proposed vision-guided module and ensures stable multiobject selecting and automatic object grasping in dynamic environments, extending promising applications in healthcare robotics, assistive technology, and industrial automation. The self-built dataset has been made publicly accessible at https://github.com/wtu1020/ ArmBCIsys-Self-built-cVEP-Dataset. Feng Yu 0017, Zhongrui Rao, Neng Chen, Li Liu 0047, Minghua Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | MHC-Segnet: Mamba-Hadamard collaboration segmentation network for multimodal MRI brain tumor
Jiacheng Cao, Liyu Ren, Ao Deng, Feng Yu 0017, Li Liu 0047, Minghua Jiang |
Vis. Comput. | 4 |
| 2024 | DesignGAN: Generation of Hand-Drawn Garment Sketches
Xinrong Hu, Jiwei Huang, Tao Peng 0006, Feng Yu 0017, Jia Chen 0012 |
CGI (1) | 5 |
| 2024 | SCB-LEDN: Lightweight and Efficient Object Detection Network for Student Classroom Behavior
Minghua Jiang, Xingwei Zheng, Mingwei He, Li Liu 0047, Feng Yu 0017 |
CGI (1) | 6 |
| 2024 | MFENet: Multi-scale and Local Frequency Enhancement Network for Skin Lesion Classification
Yuyu Jin, Zhiyong Xiao 0003, Mingwei He, Li Liu 0047, Feng Yu 0017, Minghua Jiang |
CGI (3) | 6 |
| 2024 | A Real-Time Semantic Segmentation Network for Robotic Arm Grasp
Li Liu 0047, Xinlei Zhou, Mingwei He, Feng Yu 0017, Tao Peng 0006, Xinrong Hu, Minghua Jiang |
CGI (3) | 4 |
| 2024 | Smart Clothing System for Arrhythmia Detection Based on Digital Twin Technology
Hanchen Yu, Mingwei He, Feng Yu 0017, Li Liu 0047, Minghua Jiang |
CGI (3) | 3 |
| 2024 | SmPhy: Generating smooth and physically plausible 3D garment animationsabstractDynamic garment simulation plays a crucial role in applications such as virtual try-on and film production. Existing simulation methods face challenges including high computational time, video frame jitter, and limited garment styles. Therefore, we propose SmPhy, a method that takes real videos as input. To alleviate frame jitter in video generation, we employ a temporal perception network for motion smoothing. The temporal physics garment module introduces temporal dependency, utilizing the garment information output from the current frame as input for the next frame, and provides reliable physical constraints to enhance garment deformation effects. Qualitative and quantitative experiments demonstrate that SmPhy reduces time costs and successfully simulates 3D clothing animations closely resembling real-world behaviors. Access links to supporting materials are as follows: https://drive.google.com/file/d/1BIbSI4mT4YgCVFRorszGW9pxbPZo40SH/view?usp=drive_link Jiazhe Miao, Tao Peng 0006, Xinrong Hu, Feng Yu 0017, Minghua Jiang |
ICME | 6 |
| 2024 | Intelligent Wearable System With Motion and Emotion Recognition Based on Digital Twin TechnologyabstractIntelligent wearable systems have been widely used in health monitoring, motion tracking, and engineering safety. However, the single function of current wearable systems cannot satisfy the requirements of complex scenarios, and the wearable systems cannot establish a relationship with the virtual 3D visualization platform. To address these issues, this paper proposes a novel intelligent wearable system with motion and emotion recognition. Multiple sensors are integrated into the system to collect motion and emotion information. In order to achieve accurate classification and recognition of multiple sensor information, we propose a novel human action recognition network called the three-branch spatial-temporal feature extraction network (TB-SFENet), which can obtain more robust features and achieve an accuracy of 97.04% on the UCI-HAR dataset and 92.68% on the UniMiB SHAR dataset. To establish the relationship between the real entity and virtual space, we use digital twin (DT) technology to establish the 3D display DT platform. The platform enables real-time information interaction, such as activity, emotion, location, and monitoring information. Additionally, we establish the TGAM electroencephalogram emotion classification (TEEC) dataset, which contains 120,000 pieces of data, for the proposed system. Experimental results indicate that the proposed system realizes virtual reality information interaction between the personal digital human and actual person based on the intelligent wearable system, which has great potential for applications in intelligent healthcare, virtual reality, and other fields. Feng Yu 0017, Chenyu Yu, Zhangyuan Tian, Jiacheng Cao, Li Liu 0047, Chenghu Du, Minghua Jiang |
IEEE Internet Things J. | 1 |
| 2024 | Human action recognition in immersive virtual reality based on multi-scale spatio-temporal attention networkabstractAbstract Wearable human action recognition (HAR) has practical applications in daily life. However, traditional HAR methods solely focus on identifying user movements, lacking interactivity and user engagement. This paper proposes a novel immersive HAR method called MovPosVR. Virtual reality (VR) technology is employed to create realistic scenes and enhance the user experience. To improve the accuracy of user action recognition in immersive HAR, a multi‐scale spatio‐temporal attention network (MSSTANet) is proposed. The network combines the convolutional residual squeeze and excitation (CRSE) module with the multi‐branch convolution and long short‐term memory (MCLSTM) module to extract spatio‐temporal features and automatically select relevant features from action signals. Additionally, a multi‐head attention with shared linear mechanism (MHASLM) module is designed to facilitate information interaction, further enhancing feature extraction and improving accuracy. The MSSTANet network achieves superior performance, with accuracy rates of 99.33% and 98.83% on the publicly available WISDM and PAMPA2 datasets, respectively, surpassing state‐of‐the‐art networks. Our method showcases the potential to display user actions and position information in a virtual world, enriching user experiences and interactions across diverse application scenarios. Zhiyong Xiao 0003, Xinlei Zhou, Mingwei He, Li Liu 0047, Feng Yu 0017, Minghua Jiang |
Comput. Animat. Virtual Worlds | 6 |
| 2024 | DSANet: A lightweight hybrid network for human action recognition in virtual sportsabstractAbstract Human activity recognition (HAR) has significant potential in virtual sports applications. However, current HAR networks often prioritize high accuracy at the expense of practical application requirements, resulting in networks with large parameter counts and computational complexity. This can pose challenges for real‐time and efficient recognition. This paper proposes a hybrid lightweight DSANet network designed to address the challenges of real‐time performance and algorithmic complexity. The network utilizes a multi‐scale depthwise separable convolutional (Multi‐scale DWCNN) module to extract spatial information and a multi‐layer Gated Recurrent Unit (Multi‐layer GRU) module for temporal feature extraction. It also incorporates an improved channel‐space attention module called RCSFA to enhance feature extraction capability. By leveraging channel, spatial, and temporal information, the network achieves a low number of parameters with high accuracy. Experimental evaluations on UCIHAR, WISDM, and PAMAP2 datasets demonstrate that the network not only reduces parameter counts but also achieves accuracy rates of 97.55%, 98.99%, and 98.67%, respectively, compared to state‐of‐the‐art networks. This research provides valuable insights for the virtual sports field and presents a novel network for real‐time activity recognition deployment in embedded devices. Zhiyong Xiao 0003, Feng Yu 0017, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Minghua Jiang |
Comput. Animat. Virtual Worlds | 2 |
| 2024 | Redundant same sequence point cloud registration
Feng Yu 0017, Zhaoxiang Chen, Jiacheng Cao, Minghua Jiang |
Vis. Comput. | 1 |
| 2024 | Intelligent 3D garment system of the human body based on deep spiking neural networkabstractIntelligent garments, a burgeoning class of wearable devices, have extensive applications in domains such as sports training and medical rehabilitation. Nonetheless, existing research in the smart wearables domain predominantly emphasizes sensor functionality and quantity, often skipping crucial aspects related to user experience and interaction. To address this gap, this study introduces a novel real-time 3D interactive system based on intelligent garments. The system utilizes lightweight sensor modules to collect human motion data and introduces a dual-stream fusion network based on pulsed neural units to classify and recognize human movements, thereby achieving real-time interaction between users and sensors. Additionally, the system in- corporates 3D human visualization functionality, which visualizes sensor data and recognizes human actions as 3D models in realtime, providing accurate and comprehensive visual feedback to help users better understand and analyze the details and features of human motion. This system has significant potential for applications in motion detection, medical monitoring, virtual reality, and other fields. The accurate classification of human actions con- tributes to the development of personalized training plans and injury prevention strategies. This study has substantial implications in the domains of intelligent garments, human motion monitoring, and digital twin visualization. The advancement of this system is expected to propel the progress of wearable technology and foster a deeper comprehension of human motion. Minghua Jiang, Zhangyuan Tian, Chenyu Yu, Yankang Shi, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017 |
Virtual Real. Intell. Hardw. | 8 |
| 2023 | COCCI: Context-Driven Clothing Classification Network
Minghua Jiang, Shuqing Liu, Yankang Shi, Chenghu Du, Guangyu Tang, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017 |
CGI (1) | 9 |
| 2023 | UPDN: Pedestrian Detection Network for Unmanned Aerial Vehicle Perspective
Minghua Jiang, Mengsi Guo, Li Liu 0047, Feng Yu 0017 |
CGI (3) | 5 |
| 2023 | AMDNet: Adaptive Fall Detection Based on Multi-scale Deformable Convolution Network
Minghua Jiang, Keyi Zhang, Yongkang Ma, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017 |
CGI (3) | 7 |
| 2023 | GVPM: Garment Simulation from Video Based on Priori Movements
Jiazhe Miao, Tao Peng 0006, Xinrong Hu, Feng Yu 0017, Minghua Jiang |
CGI (3) | 6 |
| 2023 | AMCNet: Adaptive Matching Constraint for Unsupervised Point Cloud Registration
Feng Yu 0017, Zhuohan Xiao, Zhaoxiang Chen, Li Liu 0047, Minghua Jiang, Xinrong Hu, Tao Peng 0006 |
CGI (1) | 1 |
| 2023 | TSFCloNet: Clothing Classification Algorithm Based on Two-Stream Network StructureabstractIn the fashion field, with the increasing diversity of clothing types and styles, accurate clothing classification becomes very important. However, the complex background and diverse styles of clothing images bring challenges to feature extraction. Classification based on texture features alone may focus too much on details and ignore the overall shape information, thus reducing the accuracy and stability of classification. In order to achieve fast and accurate clothing classification, this paper proposes a two-stream network structure clothing classification algorithm based on shape texture features and multi-feature fusion (TSFCloNet). Its main core is as follows: 1) using the two-stream network structure to extract texture and shape features from the input data set respectively; 2) in the shape feature extraction stream, the clothing shape acquisition module is first used to process the input clothing data set, and the obtained clothing shape data set is input into the ShapeNet feature extraction module to obtain shape feature information; 3) the FFCE (Feature Fusion Channel Enhancement) module is used to fuse the features obtained by the two branches of the structure respectively, and the DSAConv module is used to enhance feature extraction, and the final features are sent to the trained classifier to obtain the clothing style classification results. A large number of experimental results show that the proposed TSFCloNet network achieves higher classification accuracy when dealing with diverse and changeable fashion styles, significantly improving the performance of fashion image classification. Minghua Jiang, Yaxin Zhao, Li Liu 0047, Feng Yu 0017 |
ICPADS | 5 |
| 2023 | GSNet: Generating 3D garment animation via graph skinning networkabstractThe goal of digital dress body animation is to produce the most realistic dress body animation possible. Although a method based on the same topology as the body can produce realistic results, it can only be applied to garments with the same topology as the body. Although the generalization-based approach can be extended to different types of garment templates, it still produces effects far from reality. We propose GSNet, a learning-based model that generates realistic garment animations and applies to garment types that do not match the body topology. We encode garment templates and body motions into latent space and use graph convolution to transfer body motion information to garment templates to drive garment motions. Our model considers temporal dependency and provides reliable physical constraints to make the generated animations more realistic. Qualitative and quantitative experiments show that our approach achieves state-of-the-art 3D garment animation performance. Tao Peng 0006, Jiewen Kuang, Jinxing Liang, Xinrong Hu, Jiazhe Miao, Feng Yu 0017, Minghua Jiang |
Graph. Model. | 8 |
| 2023 | Smart Clothing System With Multiple Sensors Based on Digital Twin TechnologyabstractSmart clothing is widely used for social safety, health monitoring, and sports monitoring. Current research focuses on the use of various materials or sensors to implement smart clothes with different functions, which implies that the functionality of smart clothing depends on the number of sensors used. For existing smart clothing systems, the greatest attention has been given to information processing algorithms and assembly of sensors, and the interaction between users and systems is ignored. To address this gap, this article considers a multifunctional smart clothing system constructed with several sensors. The smart clothing system proposed in this article mainly consists of a hardware module and a software module. Four types of sensors are incorporated into the hardware module to monitor the heart rate, blood oxygen saturation, body temperature, locating information, and activity states; the software module includes the 3-D model based on the user and the feedback system based on digital twin (DT) technology. The DT technology can map the fundamental states of users in terms of the monitoring indices from the hardware module, and give correspondent advice to users. This novel smart clothing system overcomes the lack of an interaction function in existing methods and introduces DT technology into smart wearable devices for the first time. Feng Yu 0017, Minghua Jiang, Zhangyuan Tian, Tao Peng 0006, Xinrong Hu |
IEEE Internet Things J. | 1 |
| 2023 | ClothSeg: semantic segmentation network with feature projection for clothing parsing
Guangyu Tang, Feng Yu 0017, Huiyin Li, Yankang Shi, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Minghua Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | VTON-SCFA: A Virtual Try-On Network Based on the Semantic Constraints and Flow AlignmentabstractAn image-based virtual try-on system transfers an in-shop garment to the corresponding garment region of a reference person, which has huge application potential and commercial value in online clothing shopping. Existing methods have difficulty preserving garment texture and body details because of rough garment alignment and imperfect detail-retention strategies. To address this problem, we propose a virtual try-on network based on semantic constraints and flow alignment. The key idea of the framework is as follows: 1) a global-local semantic predictor (GLSP) is proposed to generate a reasonable target semantic map, which clearly guides the correct alignment of the in-shop garment with the body and the generation of try-on result; and 2) a novel appearance flow-based garment alignment network (AFGAN) is proposed to align the in-shop garment with the body, which is important to preserve maximum garment detail and ensure natural and realistic warping; and 3) we propose a synthesis strategy to integrate the aligned garment and the human body to preserve maximum body detail for generating a realistic result and preventing cross-occlusion and pixel confusion between different body parts. Experiments on the existing benchmark dataset demonstrate that the proposed method achieves the best performance on qualitative and quantitative experiments among the state-of-the-art virtual try-on techniques. Chenghu Du, Feng Yu 0017, Minghua Jiang, Ailing Hua, Tao Peng 0006, Xinrong Hu |
IEEE Trans. Multim. | 2 |
| 2023 | VTNCT: an image-based virtual try-on network by combining feature with pixel transformation
Tao Peng 0006, Feng Yu 0017, Ruhan He, Xinrong Hu, Junping Liu, Minghua Jiang |
Vis. Comput. | 3 |
| 2023 | Three stages of 3D virtual try-on network with appearance flow and shape field
Feng Yu 0017, Minghua Jiang, Ailing Hua, Tao Peng 0006, Xinrong Hu |
Vis. Comput. | 2 |
| 2022 | Multi-Pose Virtual Try-On Via Self-Adaptive Feature FilteringabstractWith the growing trend of virtual try-on, multi-pose tasks attract researchers due to their higher commercial value. Prior methods lack an effective geometric deformation to maintain the original image details resulting in many details loss in the head and garment. To address this problem, we propose a new multi-pose virtual try-on network, which can fit a garment to the corresponding area of a person in arbitrary poses. First, the target pose’s body-semantic distribution is predicted by the target pose point. Second, the in-shop garment and human body are warped based on a human pose to solve the unnatural alignment and the lack of body details by the Deformation Module (DM). Finally, the human body in the given pose and garment is fine generated by the Filtering Synthesis Network (FSN). Compared to state-of-the-art methods with objective experiments on the MPV dataset, the proposed method achieves the best performance in metrics and the rich details in visual results. Chenghu Du, Feng Yu 0017, Minghua Jiang, Tao Peng 0006, Xinrong Hu |
ICASSP | 2 |
| 2022 | Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Captureabstract3D virtual try-on receives widespread attention from scholars due to its great practical and commercial values. In prior methods, the fundamental problems lie in the limitations on texture retention during garment deformation and the lack of feature context capture during depth estimation. To address these problems, we propose a new 3D virtual try-on network via multi-scale characteristic capture (VTON-MC), which can produce an exact 3D model with the generated photo-realistic monocular image. The main processes are as follows: 1) predicting the human semantic-map and aligning the in-shop garment in the human pose using the appearance flow method, 2) synthesizing the human body and the warped garment to gain the image try-on result, and 3) estimating the human double-depth map of the image try-on result to reconstruct desired 3D try-on mesh by designed Depth Estimation Network (DEN). Extensive experiments on existing benchmark datasets demonstrate that VTON-MC outperforms state-of-the-art approaches efficiently. Chenghu Du, Feng Yu 0017, Minghua Jiang, Yaxin Zhao, Tao Peng 0006, Xinrong Hu |
ICASSP | 2 |
| 2022 | High fidelity virtual try-on network via semantic adaptation and distributed componentizationabstractImage-based virtual try-on systems have significant commercial value in online garment shopping. However, prior methods fail to appropriately handle details, so are defective in maintaining the original appearance of organizational items including arms, the neck, and in-shop garments. We propose a novel high fidelity virtual try-on network to generate realistic results. Specifically, a distributed pipeline is used for simultaneous generation of organizational items. First, the in-shop garment is warped using thin plate splines (TPS) to give a coarse shape reference, and then a corresponding target semantic map is generated, which can adaptively respond to the distribution of different items triggered by different garments. Second, organizational items are componentized separately using our novel semantic map-based image adjustment network (SMIAN) to avoid interference between body parts. Finally, all components are integrated to generate the overall result by SMIAN. A priori dual-modal information is incorporated in the tail layers of SMIAN to improve the convergence rate of the network. Experiments demonstrate that the proposed method can retain better details of condition information than current methods. Our method achieves convincing quantitative and qualitative results on existing benchmark datasets. Chenghu Du, Feng Yu 0017, Minghua Jiang, Ailing Hua, Yaxin Zhao, Tao Peng 0006, Xinrong Hu |
Comput. Vis. Media | 2 |
| 2022 | Virtual try-on based on attention U-Net
Xinrong Hu, Jinxing Liang, Feng Yu 0017, Tao Peng 0006 |
Vis. Comput. | 5 |
| 2021 | Cascaded Cross-Domain Fusion of Virtual Try-OnabstractImage-based virtual try-on, aiming to fit new in-shop clothes into a person image, has gained extensive attention in the fields of computer vision and image process community. However, the existing methods are difficult to generate photo-realistic try-on images when large-scale deformations or large occlusions occur. To address this issue, we propose a novel two stage visual try-on network. Specifically, in the first stage, we used a shape matching model to learn the geometric transformation of in-shop clothes. For the second stage, an U-net with cascaded attention mechanism is presented to learn the composition mask which adjust the clothes and rendered persons. The adjusted clothes and the rendered person are combined by the composition mask to get the final try-on result. Experimental results have shown that our method can generate photo-realistic images with no occlusion. Xinrong Hu, Tao Peng 0006, Mingfu Xiong, Feng Yu 0017, Li Li 0094 |
BIBM | 5 |
| 2021 | VTON-HF: High Fidelity Virtual Try-on Network via Semantic AdaptationabstractThe image-based virtual try-on network transfers the target garment item to the corresponding region of the human body. Due to its commercial value in online garment shopping, it has attracted extensive attention from researchers. However, the previous virtual try-on methods are interfered heavily by garments in reference images, so they have defects in maintaining details of human upper limbs, neck, and given garment. Therefore, a novel High Fidelity Virtual Try-on Network via Semantic Adaptation (VTON-HF) is proposed to generate a result with better details. The main processes are as follows: 1) Thin Plate Spline (TPS) warps the target garment coarsely, 2) parsing network generates a target semantic map with the coarse warped garment, 3) our novel Semantic Map-based Image Adjustment Network (SMIAN) generates components separately to avoid interference between image parts with different semantics, 4) SMIAN fuses all components to generate the final result. VTON-HF can retain the maximum amount of detail in the reference garment than previous methods. Our novel architecture generates desired results by fusing separately generated components (garment, upper limb, and neck) and unchanging parts of the reference image. Moreover, our SMIAN incorporates a priori multimodal information in the tail layer, which effectively improves the convergence efficiency of the network. Our method achieves state-of-the-art quantitative results on IS, SSIM, PSNR, and FID using the VITON dataset. (see Fig. 1). Chenghu Du, Feng Yu 0017, Minghua Jiang, Tao Peng 0006, Xinrong Hu |
ICTAI | 2 |
| 2018 | An Image-guided Endoscope System for the Ureter Detection
Enmin Song, Feng Yu 0017, Hong Liu 0005, Youming Wan, Chih-Cheng Hung |
Mob. Networks Appl. | 2 |