Tao Peng 0006

dblp:89/6609-6 · DBLP profile ↗
← Back
89ranked-venue papers
13as first author
86since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 9 first-author · 57 since 2021Computer networks · 11 · 11 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Enhancing GNN-Based Cloth Simulation with Macro-spatial and Local Dynamic Priors
Tao Peng 0006, Chuang Yin, Li Li 0094, Junping Liu, Tongyu Liu
ICIC (6)1
2026 PSR-Diff: Polarization-Guided Diffusion Model for Single Image Specular Highlight Removal
Guobin Zhang, Li Li 0094, Zhaojing Wang, Tao Peng 0006, Xinrong Hu
MMM (2)5
2026 Dapadv: Differentiated adversarial perturbation generation method in problem space for android malware detection
Junwei Tang, Tao Peng 0006
Comput. Secur.3
2026 CPSN: Collision-Overlap and Physics-Based Self-Supervised Neural Cloth Simulation
abstract
ABSTRACT Self‐collision handling remains a fundamental and long‐standing challenge in neural cloth simulation, particularly for loose garments with complex topologies. We propose a self‐supervised neural cloth simulation framework that integrates Gaussian mixture skinning (GMS) with a differentiable collision‐overlap loss to significantly enhance physical plausibility and visual realism. We employ continuous and spatially smooth GMS weights to model vertex‐skeleton coupling, enabling stable deformations under large body motions. To explicitly address cloth self‐collisions, we introduce a differentiable spatial repulsion constraint that suppresses interpenetration and layer‐overlap artifacts. The proposed objective is jointly optimized with physics‐inspired losses, enabling the network to learn consistent cloth dynamics without relying on ground‐truth physical simulations. Experimental results demonstrate improved temporal stability, reduced collision artifacts, and stronger generalization compared to existing self‐supervised methods.
Tao Peng 0006, Xianfang Tang, Li Li 0094, Xinrong Hu
Comput. Animat. Virtual Worlds1
2026 Fusion of microstructural images and constituent properties for elastic property prediction in unidirectional composites: a hybrid ResNet34-MLP approach
Tao Peng 0006, Tongyu Liu, Junping Liu, Xinrong Hu, Li Li 0094
Neural Comput. Appl.1
2026 Shape-texture aware multi-source domain adaptation for industrial anomaly detection
Yaochong Xie, Li Li 0094, Zhaojing Wang, Yaxi Zhou, Tao Peng 0006, Xinrong Hu
Vis. Comput.5
2025 GartransNet: 3D Garments Animation via Transmission Optimized Networks
Tao Peng 0006, Wenjie Yue, Junping Liu, Xinrong Hu, Li Li 0094
CGI (2)1
2025 MIGEdit: Multimodal Interactive Garment Editing
Sicheng Zheng, Bangchao Wang, Jinxing Liang, Li Li 0094, Tao Peng 0006, Junping Liu, Ping Li 0016, Xinrong Hu
CGI (2)5
2025 Accelerating Hierarchical GNN-Based Cloth Simulation via Efficient Message Passing
abstract
To address the efficiency bottlenecks in hierarchical GNN-based cloth simulation, we propose two targeted architectural components that jointly optimize message propagation and topological redundancy. Specifically, we introduce a Redundant Message Reduction strategy that streamlines cloth-body interactions by restricting communication to only necessary stages, and a Proportional Edge Pruning mechanism that adaptively reduces the number of incoming edges per node while retaining sufficient deformation paths. These designs substantially reduce computational overhead with minimal sacrifice of physical fidelity. Extensive experiments across diverse garments and motion sequences demonstrate that our model achieves more than a 40% increase in inference speed and a 22% reduction in training time, while ensuring both numerical and visual fidelity. In addition, qualitative evaluations confirm that our method preserves realistic garment dynamics under challenging poses and unseen clothing structures.
Tao Peng 0006, Chuang Yin, Saishang Zhong, Li Li 0094, Xinrong Hu
CW1
2025 VULDA: Source Code Vulnerability Detection via Local Dependency Context Aggregation on Vulnerability-Aware Code Mapping Graph
Tao Peng 0006, Ling Gui, Junwei Tang, Aoshuang Ye
ICICS (3)1
2025 STCGen: Sketch-based Text-to-Clothing Image Generation with Contour and Style Consistency
abstract
In modern fashion design field, it is a mainstream practice to generate clothing images by combining sketch and text. However, the image quality generated by existing multimodal methods combining sketches and text descriptions is suboptimal, as the clothing in the generated image often lacks contour accuracy and stylistic coherence. In this paper, we present STCGen, an advanced multimodal framework that uses both sketches and text to generate clothing images with improved contours and more consistent style. First, we introduce the sketch prior embedding module, which processes sketches to extract key structural features and ensure the consistency of contours, thereby enhancing image details. Second, we propose a cross space attention mechanism to address the issue of text information loss and ensure stylistic consistency, thereby enhancing overall image coherence. Finally, we propose a network simplification scheme to reduce complexity without compromising the quality of resulting images. Experimental results demonstrate that our method excels in generating high-fidelity clothing images.
Chunxia Xiao, Ruhan He, Jia Chen 0012, Mingfu Xiong, Tao Peng 0006, Xinrong Hu
MMAsia8
2025 CasRPN: Cascade region proposal network for visual tracking
abstract
Trackers based on the region proposal network (RPN) have garnered extensive attention within object tracking. However, current RPN trackers utilizing the ResNet as the feature extraction network only employ its local convolution operation to extract the image features, thereby constraining the model’s comprehension of the global contextual information. To address this issue, this paper employs a Vision Transformer (ViT) to construct global relations across the entire sequence features, enhancing the model’s capability to represent global features. Additionally, traditional single-stage RPN trackers generate candidate boxes through a coarse regression process. Thus, the candidate boxes may only partially cover the object or include excessive background information, diminishing the model’s accuracy in object localization. Consequently, we use a multi-stage RPN to adjust the anchor boxes in a cascade RPN manner. Moreover, traditional multi-stage RPN exists the misalignment problem between the anchor boxes and image features. To further optimize the multi-stage RPN, this paper employs adaptive convolution to align features with anchor boxes. Our tracking method achieves state-of-the-art results on five common tracking benchmark datasets. Specifically, the CasRPN achieves an AUC score of 85.3% on the large-scale Trackingnet dataset.
Jia Chen 0012, Youkang Yuan, Xinrong Hu, Tao Peng 0006
Expert Syst. Appl.5
2025 MelodyTransformer: Improving lyric-to-melody generation by considering melodic features
Ruhan He, Ruixue Liu, Tao Peng 0006, Xinrong Hu
Neurocomputing3
2025 DTDroid: Adversarial Packed Android Malware Detection Based on Traffic and Dynamic Behavioral
abstract
Android has occupied an important share of the operating system of intelligent terminal devices in the Internet of Things (IoT), and the malicious applications of Android have increased rapidly, posing a serious threat to the security of IoT. Machine learning has advanced significantly in the detection of android malware. In order to protect intellectual property, Android developers have begun to use packing techniques to enhance the security of their applications. However, attackers can also pack their malware, which may make feature extraction ineffective and interfere with the prediction results of learning-based classifiers. For this issue, we have designed and implemented a tool by dynamically loading the original DEX using a shell DexClassLoader to generate a batch of packed Android applications. And we have verified that several existing methods fail when faced with packed samples. Therefore, we propose a novel malware detection method called DTDroid that can resist code packing. DTDroid automatically captures network traffic characteristics of target samples based on fuzzy testing and network traffic packet extraction. At the same time, the dynamic behavior characteristics of the target application can be obtained by monitoring the corresponding runtime function calls and system status. The extracted two types of features are contextually spliced and converted into grayscale images, and then detected based on deep learning model. Experimental results show that the detection accuracy of our method reaches 94.22% and 95.14%, respectively, on two kinds of packed datasets, indicating that DTDroid has better robustness for packed samples than the existing methods.
Junwei Tang, Tao Peng 0006, Xiaoyun Yan, Xinrong Hu
IEEE Internet Things J.3
2025 DeFinder: Error-sensitive testing of deep neural networks via vulnerability interpretation
Aoshuang Ye, Benxiao Tang, Jianpeng Ke, Yiru Zhao, Tao Peng 0006
J. Netw. Comput. Appl.6
2025 VULOC: Vulnerability location framework based on assembly code slicing
Xinghang Lv, Jianming Fu, Tao Peng 0006
J. Syst. Softw.3
2025 Enhancing feature interaction for improved generalization in few-shot metal surface defect segmentation
Haijun Yan, Tao Peng 0006, Xinrong Hu, Shuhan Qi
Knowl. Based Syst.3
2025 CST: a melody generation method based on ChatGPT and Structure Transformer
Ruhan He, Ruixue Liu, Tao Peng 0006, Xinrong Hu
Multim. Syst.3
2025 Face photo-sketch portraits transformation via generation pipeline
Mengsi Guo, Mingfu Xiong, Xinrong Hu, Tao Peng 0006
Vis. Comput.5
2025 Enhancing low-frequency stitch code generation for knitted fabrics: an LFSCG-E-Net approach
Jinxing Liang, Kaifang Han, Ruixin Gao, Jiajia Peng, Tao Peng 0006, Xinrong Hu
Vis. Comput.6
2025 TDGar-Ani: temporal motion fusion model and deformation correction network for enhancing garment animation details
Jiazhe Miao, Tao Peng 0006, Xinrong Hu, Li Li 0094
Vis. Comput.2
2025 ViT-BF: vision transformer with border-aware features for visual tracking
Ping Li 0016, Jinxing Liang, Tao Peng 0006, Jia Chen 0012, Li Li 0094, Xinrong Hu, Junping Liu
Vis. Comput.5
2025 Learning monocular face reconstruction from in the wild images using rotation cycle consistency
abstract
With the popularity of the digital human body, monocular three-dimensional (3D) face reconstruction is widely used in fields such as animation and face recognition. Although current methods trained using single-view image sets perform well in monocular 3D face reconstruction tasks, they tend to rely on the constraints of the a priori model or the appearance conditions of the input images, fundamentally because of the inability to propose an effective method to reduce the effects of two-dimensional (2D) ambiguity. To solve this problem, we developed an unsupervised training framework for monocular face 3D reconstruction using rotational cycle consistency. Specifically, to learn more accurate facial information, we first used an autoencoder to factor the input images and applied these factors to generate normalized frontal views. We then proceeded through a differentiable renderer to use rotational consistency to continuously perceive refinement. Our method provided implicit multi-view consistency constraints on the pose and depth information estimation of the input face, and the performance was accurate and robust in the presence of large variations in expression and pose. In the benchmark tests, our method performed more stably and realistically than other methods that used 3D face reconstruction in monocular 2D images.
Xinrong Hu, Kaifan Yang, Ruiqi Luo, Tao Peng 0006, Junping Liu
Virtual Real. Intell. Hardw.4
2025 Deconfounded fashion image captioning with transformer and multimodal retrieval
abstract
Background The annotation of fashion images is a significantly important task in the fashion industry as well as social media and e-commerce. However, owing to the complexity and diversity of fashion images, this task entails multiple challenges, including the lack of fine-grained captions and confounders caused by dataset bias. Specifically, confounders often cause models to learn spurious correlations, thereby reducing their generalization capabilities. Method In this work, we propose the Deconfounded Fashion Image Captioning (DFIC) framework, which first uses multimodal retrieval to enrich the predicted captions of clothing, and then constructs a detailed causal graph using causal inference in the decoder to perform deconfounding. Multimodal retrieval is used to obtain semantic words related to image features, which are input into the decoder as prompt words to enrich sentence descriptions. In the decoder, causal inference is applied to disentangle visual and semantic features while concurrently eliminating visual and language confounding. Results Overall, our method can not only effectively enrich the captions of target images, but also greatly reduce confounders caused by the dataset. To verify the effectiveness of the proposed framework, the model was experimentally verified using the FACAD dataset.
Tao Peng 0006, Weiqiao Yin, Junping Liu, Li Li 0094, Xinrong Hu
Virtual Real. Intell. Hardw.1
2024 DesignGAN: Generation of Hand-Drawn Garment Sketches
Xinrong Hu, Jiwei Huang, Tao Peng 0006, Feng Yu 0017, Jia Chen 0012
CGI (1)4
2024 A Real-Time Semantic Segmentation Network for Robotic Arm Grasp
Li Liu 0047, Xinlei Zhou, Mingwei He, Feng Yu 0017, Tao Peng 0006, Xinrong Hu, Minghua Jiang
CGI (3)5
2024 DS-Seq: Deriving Smooth 3D Human Motion Sequences from Video Time Cues
Tao Peng 0006, Delang Peng, Li Li 0094, Junping Liu, Xinrong Hu
CGI (1)1
2024 GRD: Garment Reconstruction and Draping with Preserved Design Based on 2D Image
Tao Peng 0006, Li Li 0094, Jiazhe Miao, Junping Liu, Xinrong Hu
CGI (2)1
2024 AT-I-FGSM: A novel adversarial CAPTCHA generation method based on gradient adaptive truncation
abstract
Text-based CAPTCHA is widely used in fields such as user identity verification during human-computer interaction in real scenarios. With the development of artificial intelligence, several technologies that automatically bypass CAPTCHAs have emerged, weakening the robustness of CAPTCHAs. In-depth study of adversarial sample technology is needed to further reduce the accuracy of automatic verification code recognition of deep learning models while retaining correct human recognition. We propose a novel method based on gradient adaptive truncation to generate adversarial text-based CAPTCHAs more efficiently. Based on the generated model, our method dynamically adjusts the gradient truncation threshold according to the progress of the perturbation attack method, thereby improving the performance of the sample generation model. On the authoritative dataset, our method is compared with the existing state-of-the-art methods. The results show that our AT-I-FGSM can more effectively reduce the accuracy of automatic recognition models to identify CAPTCHAs and improve the security of CAPTCHAs. At the same time, our method consumes less time in generating CAPTCHAs.
Junwei Tang, Tao Peng 0006, Ruhan He, Xinrong Hu, Changzheng Liu
CSCWD4
2024 Open-Vocabulary RGB-Thermal Semantic Segmentation
Xiaoyun Yan, Zhaojing Wang, Junwei Tang, Yangjun Ou, Xinrong Hu, Tao Peng 0006
ECCV (74)8
2024 SGM: A Dataset for 3D Garment Reconstruction from Single Hand-Drawn Sketch
abstract
High-fidelity garment reconstruction is essential for various applications such as garment design and virtual try-on. While image-based reconstruction methods have made significant progress with deep generative models, generating 3D models from hand-drawn sketches to meet design intentions remains challenging. One of the main obstacles is the limited availability of large-scale 3D garment models accompanied by corresponding sketches. To address this issue, we propose SGM, a comprehensive dataset comprising 656 garment models categorized into short and long sleeves. Each garment model in SGM is accompanied by four types of rendered images and a series of UDF values. Furthermore, we introduce a novel baseline approach for sketch-based garment reconstruction using an end-to-end generative network capable of generating garment models from single hand-drawn sketches. Extensive experimental results highlight the significance and value of our proposed dataset and method. We plan to make SGM publicly available upon publication.
Jia Chen 0012, Jinlong Qin, Saishang Zhong, Xinrong Hu, Tao Peng 0006
ICASSP6
2024 SmPhy: Generating smooth and physically plausible 3D garment animations
abstract
Dynamic garment simulation plays a crucial role in applications such as virtual try-on and film production. Existing simulation methods face challenges including high computational time, video frame jitter, and limited garment styles. Therefore, we propose SmPhy, a method that takes real videos as input. To alleviate frame jitter in video generation, we employ a temporal perception network for motion smoothing. The temporal physics garment module introduces temporal dependency, utilizing the garment information output from the current frame as input for the next frame, and provides reliable physical constraints to enhance garment deformation effects. Qualitative and quantitative experiments demonstrate that SmPhy reduces time costs and successfully simulates 3D clothing animations closely resembling real-world behaviors. Access links to supporting materials are as follows: https://drive.google.com/file/d/1BIbSI4mT4YgCVFRorszGW9pxbPZo40SH/view?usp=drive_link
Jiazhe Miao, Tao Peng 0006, Xinrong Hu, Feng Yu 0017, Minghua Jiang
ICME2
2024 GarTemFormer: Temporal transformer-based for optimizing virtual garment animation
abstract
Virtual garment animation and deformation constitute a pivotal research direction in computer graphics, finding extensive applications in domains such as computer games, animation, and film. Traditional physics-based methods can simulate the physical characteristics of garments, such as elasticity and gravity, to generate realistic deformation effects. However, the computational complexity of such methods hinders real-time animation generation. Data-driven approaches, on the other hand, learn from existing garment deformation data, enabling rapid animation generation. Nevertheless, animations produced using this approach often lack realism, struggling to capture subtle variations in garment behavior. We proposes an approach that balances realism and speed, by considering both spatial and temporal dimensions, we leverage real-world videos to capture human motion and garment deformation, thereby producing more realistic animation effects. We address the complexity of spatiotemporal attention by aligning input features and calculating spatiotemporal attention at each spatial position in a batch-wise manner. For garment deformation, garment segmentation techniques are employed to extract garment templates from videos. Subsequently, leveraging our designed Transformer-based temporal framework, we capture the correlation between garment deformation and human body shape features, as well as frame-level dependencies. Furthermore, we utilize a feature fusion strategy to merge shape and motion features, addressing penetration issues between clothing and the human body through post-processing, thus generating collision-free garment deformation sequences. Qualitative and quantitative experiments demonstrate the superiority of our approach over existing methods, efficiently producing temporally coherent and realistic dynamic garment deformations. • Both spatial and temporal dimensions are considered in terms of human movement. • Feature parameter fusion strategy to integrate human shape and motion features. • The attentional mechanism establishes dependencies between garment frames. • We resolve the garment-body interpenetration issue through post-processing.
Jiazhe Miao, Tao Peng 0006, Xinrong Hu, Li Li 0094
Graph. Model.2
2024 Lightweight Verifiable Privacy-Preserving Data Aggregation for Smart Grids
abstract
As an indispensable part of a smart city, the smart grid has gained widespread attention from industrial and academic communities. How to securely collect users’ real-time energy consumption data to provide services such as big data analytics and demand-response services while ensuring the privacy of individual users is a challenging issue in the smart grid. The privacy-preserving data aggregation (P2DA) suggests a feasible solution. For years, researchers have designed numerous P2DA schemes for securing smart grids. Unfortunately, the majority of them have some security and privacy deficiencies. Other schemes are unsuitable for resource-constrained smart meters due to expensive cryptographic operations. In this work, we design a lightweight verifiable certificate-based P2DA scheme LV-P2DA without pairings for smart grids. We formally prove its security under standard cryptographic assumptions. The performance comparison results illustrate that compared with state-of-the-art solutions, our design achieves at least a 99.43% improvement in computational cost and a 32.96% improvement in communication cost on the smart meter side, respectively.
Duan Guo, Alsharif Abuadbba, Xun Yi, Saru Kumari, Tao Peng 0006
IEEE Internet Things J.7
2024 Android malware detection based on a novel mixed bytecode image combined with attention mechanism
Junwei Tang, Tao Peng 0006, Qiaosen Pi, Ruhan He, Xinrong Hu
J. Inf. Secur. Appl.3
2024 Highlight mask-guided adaptive residual network for single image highlight detection and removal
abstract
Abstract Specular highlights detection and removal is a challenging task. Although various methods exist for removing specular highlights, they often fail to effectively preserve the color and texture details of objects after highlight removal due to the high brightness and nonuniform distribution characteristics of highlights. Furthermore, when processing scenes with complex highlight properties, existing methods frequently encounter performance bottlenecks, which restrict their applicability. Therefore, we introduce a highlight mask‐guided adaptive residual network (HMGARN). HMGARN comprises three main components: detection‐net, adaptive‐removal network (AR‐Net), and reconstruct‐net. Specifically, detection‐net can accurately predict highlight mask from a single RGB image. The predicted highlight mask is then inputted into the AR‐Net, which adaptively guides the model to remove specular highlights and estimate an image without specular highlights. Subsequently, reconstruct‐net is used to progressively refine this result, remove any residual specular highlights, and construct the final high‐quality image without specular highlights. We evaluated our method on the public dataset (SHIQ) and confirmed its superiority through comparative experimental results.
Shuaibin Wang, Li Li 0094, Tao Peng 0006
Comput. Animat. Virtual Worlds4
2024 DSANet: A lightweight hybrid network for human action recognition in virtual sports
abstract
Abstract Human activity recognition (HAR) has significant potential in virtual sports applications. However, current HAR networks often prioritize high accuracy at the expense of practical application requirements, resulting in networks with large parameter counts and computational complexity. This can pose challenges for real‐time and efficient recognition. This paper proposes a hybrid lightweight DSANet network designed to address the challenges of real‐time performance and algorithmic complexity. The network utilizes a multi‐scale depthwise separable convolutional (Multi‐scale DWCNN) module to extract spatial information and a multi‐layer Gated Recurrent Unit (Multi‐layer GRU) module for temporal feature extraction. It also incorporates an improved channel‐space attention module called RCSFA to enhance feature extraction capability. By leveraging channel, spatial, and temporal information, the network achieves a low number of parameters with high accuracy. Experimental evaluations on UCIHAR, WISDM, and PAMAP2 datasets demonstrate that the network not only reduces parameter counts but also achieves accuracy rates of 97.55%, 98.99%, and 98.67%, respectively, compared to state‐of‐the‐art networks. This research provides valuable insights for the virtual sports field and presents a novel network for real‐time activity recognition deployment in embedded devices.
Zhiyong Xiao 0003, Feng Yu 0017, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Minghua Jiang
Comput. Animat. Virtual Worlds4
2024 ANDE: Detect the Anonymity Web Traffic With Comprehensive Model
abstract
The escalating growth of network technology and users poses critical challenges to network security. This paper introduces ANDE, a novel framework designed to enhance the classification accuracy of anonymity networks. ANDE incorporates both raw data features and statistical features extracted from network traffic. Raw data features are transformed into images, enabling recognition and classification using robust image domain models. ANDE combines an enhanced Squeeze-and-Excitation (SE) ResNet with Multilayer Perceptrons (MLP), facilitating concurrent learning and classification of both feature types. Extensive experiments on two publicly available datasets demonstrate the superior performance of ANDE compared to traditional machine learning and deep learning methods. The comprehensive evaluation underscores ANDE’s effectiveness in accurately classifying network traffic within anonymity networks. Additionally, this study empirically validates the efficacy of the SE block in augmenting the classification capabilities of the proposed framework, establishing ANDE as a promising solution for network traffic classification in the realm of network security.
Yunlong Deng, Tao Peng 0006, Bangchao Wang, Gan Wu
IEEE Trans. Netw. Serv. Manag.2
2024 Device Identification Method for Internet of Things Based on Spatial-Temporal Feature Residuals
abstract
In recent years, the Internet of Things (IoT) has penetrated all aspects of our lives through smart cities, health, industries and others that are related to people's livelihood. With the increasing number of IoT devices, more and more personal information is exposed in the network space, which inevitably brings some network security problems. Due to the diversity and heterogeneity of IoT devices, identification of such devices in the complex IoT environments remains a major challenge. Existing deep learning-based device identification methods achieve identification of IoT devices by automatically extracting device traffic features, but usually only single modal features of device traffic are considered, which cannot achieve all-around characterization features of communication traffic and affect the identification results. Therefore, we propose an identification method, termed DMRMTT, that employs a Deep convolutional maxout network and MTT model (Multiple Time-series Transformers) to automatically extract the spatial and temporal features of IoT communication session fingerprints and perform further fusion using the structure of the residual, which makes up for the limitations of the existing methods for studying device traffic. This method can improve the characterization of device traffic behaviour and achieve a more accurate identification of IoT devices. Its efficacy is experimentally validated by using two publicly availbale datasets and compared with existing methods. Results show that our method outperforms other methods in widely used performance metrics and achieves 99.82% identification accuracy, demonstrating its superiority and usefulness in IoT device identification.
Shi Dong 0001, Longhui Shu, Qinyu Xia, Joarder Kamruzzaman, Yuanjun Xia, Tao Peng 0006
IEEE Trans. Serv. Comput.6
2024 Outfit compatibility model using fully connected self-adjusting graph neural network
Li Li 0094, Neng Yu, Tao Peng 0006, Xinrong Hu
Vis. Comput.5
2024 Intelligent 3D garment system of the human body based on deep spiking neural network
abstract
Intelligent garments, a burgeoning class of wearable devices, have extensive applications in domains such as sports training and medical rehabilitation. Nonetheless, existing research in the smart wearables domain predominantly emphasizes sensor functionality and quantity, often skipping crucial aspects related to user experience and interaction. To address this gap, this study introduces a novel real-time 3D interactive system based on intelligent garments. The system utilizes lightweight sensor modules to collect human motion data and introduces a dual-stream fusion network based on pulsed neural units to classify and recognize human movements, thereby achieving real-time interaction between users and sensors. Additionally, the system in- corporates 3D human visualization functionality, which visualizes sensor data and recognizes human actions as 3D models in realtime, providing accurate and comprehensive visual feedback to help users better understand and analyze the details and features of human motion. This system has significant potential for applications in motion detection, medical monitoring, virtual reality, and other fields. The accurate classification of human actions con- tributes to the development of personalized training plans and injury prevention strategies. This study has substantial implications in the domains of intelligent garments, human motion monitoring, and digital twin visualization. The advancement of this system is expected to propel the progress of wearable technology and foster a deeper comprehension of human motion.
Minghua Jiang, Zhangyuan Tian, Chenyu Yu, Yankang Shi, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017
Virtual Real. Intell. Hardw.6
2023 ChatICD: Prompt Learning for Few-shot ICD Coding through ChatGPT
abstract
Automated International Classification of Diseases (ICD) coding involves the automated assignment of diverse disease codes to clinical medical texts. It is considered as a multi-label classification task. Because most ICD codes are rare, the imbalanced distribution and small sample size issue make this task challenging. Inspired by the recent success of ChatGPT and prompt-based fine-tuning, this study proposes a model called ChatICD to address the issue of few-shot ICD coding. First, ChatGPT for data augumentation rephrases the descriptions of ICD codes into multiple samples. Then, ChatICD fine-tunes the pretrained model by generating prompt templates and label mapping words. We conduct an evaluation of ChatICD on benchmark datasets, namely MIMIC-III-50 and MIMIC-III-rare50. On the few-shot ICD coding task of MIMIC-III-rare50, ChatICD achieves macro-F1 and micro-F1 of 35.8% and 38.2% respectively, which is a 5.4% and 5.6% improvement over the current best model.
Junping Liu, Shichen Yang, Tao Peng 0006, Xinrong Hu
BIBM3
2023 MARANet: Multi-scale Adaptive Region Attention Network for Few-Shot Learning
Jia Chen 0012, Xiyang Li, Yangjun Ou, Xinrong Hu, Tao Peng 0006
CGI (1)5
2023 FoldGEN: Multimodal Transformer for Garment Sketch-to-Photo Generation
Jia Chen 0012, Yanfang Wen, Xinrong Hu, Tao Peng 0006
CGI5
2023 COCCI: Context-Driven Clothing Classification Network
Minghua Jiang, Shuqing Liu, Yankang Shi, Chenghu Du, Guangyu Tang, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017
CGI (1)7
2023 AMDNet: Adaptive Fall Detection Based on Multi-scale Deformable Convolution Network
Minghua Jiang, Keyi Zhang, Yongkang Ma, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017
CGI (3)5
2023 Highlight Removal from a Single Image Based on a Prior Knowledge Guided Unsupervised CycleGAN
Yongkang Ma, Li Li 0094, Hao Chen 0140, Tao Peng 0006, Xiong Pan
CGI (1)7
2023 GVPM: Garment Simulation from Video Based on Priori Movements
Jiazhe Miao, Tao Peng 0006, Xinrong Hu, Feng Yu 0017, Minghua Jiang
CGI (3)2
2023 A HRNet-Transformer Network Combining Recurrent-Tokens for Remote Sensing Image Change Detection
Tao Peng 0006, Lingjie Hu, Junping Liu, Xingrong Hu, Ruhan He
CGI (3)1
2023 AMCNet: Adaptive Matching Constraint for Unsupervised Point Cloud Registration
Feng Yu 0017, Zhuohan Xiao, Zhaoxiang Chen, Li Liu 0047, Minghua Jiang, Xinrong Hu, Tao Peng 0006
CGI (1)8
2023 Anomaly Detection of Industrial Products Considering Both Texture and Shape Information
Shaojiang Yuan, Li Li 0094, Neng Yu, Tao Peng 0006, Xinrong Hu, Xiong Pan
CGI (3)4
2023 Monocular 3D Human Pose Estimation Based on Global Temporal-Attentive and Joints-Attention In Video
abstract
Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video, which is widely used in many 3D applications. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local contextual relations of human motion and ignores human joint hierarchies. To address this problem, we propose a Global Temporal-Attentive and Joints-Attention network (GTAJA-Net). This method introduces a Global Attention Feature Integration (GAFI) module and a Motion Tree Fusion Decoder (MTFD) module on the basis of a temporally consistent mesh recovery system (TCMR). A GAFI consisting of a collection of temporal features obtains final temporal features carrying spatial information that enhances temporal correlation and refine the features of the current frame. Meanwhile, MTFD aims at modeling the joint level attention. MTFD considers pose estimation as a top-down hierarchical process similar to SMPL kinematic tree. Though conceptually simple, our GTAJA-Net outperforms the state-of-the-art methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmark datasets. Our code is available at https://github.com/xiangcece/GTAJA-Net.
Ruhan He, Shanshan Xiang, Tao Peng 0006
ICASSP3
2023 Cross-cycle Transformer-based Stitching Method for Low-resolution Borehole Images
abstract
The stitching of borehole images has an important predictive role in safety analysis in the field of geotechnical engineering and intelligent geological exploration. Applying traditional image stitching methods that designed specifically for high-resolution images to low-resolution images will lead to blurred stitching results, stitching seams, fewer matched feature points and difficulties in massive image stitching. To address these problems, we propose an autoencoder-based coarse-to-fine feature extraction network, which can extract image features with high semantic and improves the accuracy of the feature point matching. Besides, we design a cross-cycle Transformer-based image stitching framework, which increase the number of matching feature points by Cross-QuadTree attention and stitch image by affine transformation. Experimental results show that the proposed method can effectively stitch low-resolution geotechnical borehole images with satisfactory visual quality.
Jia Chen 0012, Zhenpeng Fu, Mingfu Xiong, Xinrong Hu, Tao Peng 0006
ICME6
2023 Unsupervised Fashion Style Learning by Solving Fashion Jigsaw Puzzles
abstract
Fashion style learning is the basis for many tasks in fashion AI, such as clothing recommendations, fashion trend analysis and popularity prediction. Most of the existing methods rely on the quality and quantity of the annotations. This paper proposes an efficient two-step unsupervised fashion style learning framework with "Fashion Jigsaw" task and centroid-based density clustering algorithm. First, we design the "Fashion Jigsaw" unsupervised learning task according to the distribution of fashion elements in full-body fashion images. By splitting and recovering fashion images, we pre-train a model that can extract both intra-image and inter-image information. Second, we propose a centroid-based density clustering algorithm and introduce the concept of "centroid" to cluster fashion image features and represent fashion styles. Meanwhile, we keep the noise features to discover the newly sprouted fashion styles. Experiment results demonstrate the effectiveness of our proposed method.
Jia Chen 0012, Haidongqing Yuan, Tao Peng 0006, Xinrong Hu
ICME4
2023 A lightweight method for Android malware classification based on teacher assistant distillation
abstract
In recent years, the growing concern over mobile security and the associated risks posed by mobile malware have prompted an increased focus on utilizing deep learning models for analyzing Android application security. However, the expansion of deep learning model sizes results in an exponential growth of model parameters, demanding significant computing resources for execution. To address this challenge, we propose a lightweight Android malware detection method based on teacher-assistant-student knowledge distillation. Our method enables predicting on local clients, eliminating the need for cloud-base service interactions, and protecting user privacy. We visualize the binary file of the target Android application as an RGB three-channel color image, using ResNeSt50 as the teacher model, and compress it based on knowledge distillation. An assistant model is incorporated to address the issue of insufficient distillation resulting from the significant gap between the teacher and student models. Additionally, we integrate a split-attention mechanism to enhance the ability of the professor model to acquire deep features of malware images. We conduct experiments on Drebin and CICMalDroid 2020 datasets and the results show that the proposed method can ensure that the detection results of student model are more similar to those of the teacher model while reducing model complexity. Our method reduces the number of model parameters by 95% compare to the teacher model while maintaining accuracy. And the accuracy is improved by 0.63% compare to the traditional distillation method.
Junwei Tang, Qiaosen Pi, Ruhan He, Tao Peng 0006, Xinrong Hu
MSN5
2023 Graphormer-Based Contextual Reasoning Network for Small Object Detection
Jia Chen 0012, Xiyang Li, Yangjun Ou, Xinrong Hu, Tao Peng 0006
PRCV (9)5
2023 PTLVD:Program Slicing and Transformer-based Line-level Vulnerability Detection System
abstract
In recent years, deep learning-based software vulnerability detection methods have made significant progress. However, most existing methods focus on detecting vulnerabilities at the function-level or slice-level and cannot pinpoint the exact lines of code that cause the vulnerabilities. Program slicing can extract control and data dependency information from the code to assist deep learning models in detecting vulnerabilities. We propose a novel vulnerability detection model, PTLVD, which generates code gadgets(CGs) by slicing the program based on variables in the code, uses a transformer model for binary classification, and employs our proposed method Integrated Gradients Enhanced with Saliency(IGS) to locate the lines of code that are likely to cause vulnerabilities. IGS enhances the interpretability of the model by integrating the Integrated Gradients and Saliency methods. PTLVD employs an improved method of generating CGs to selectively remove irrelevant code statements, resulting in CGs that contain richer information and enhance the model’s performance. Additionally, during the preprocessing stage, PTLVD removes comments and standardizes code statements onto the same line, which effectively enhances the performance and vulnerability localization capabilities of the model. Experimental results show that, compared to state-of-the-art function-level and slice-level vulnerability detection models, PTLVD improves precision and F1 by 5.25% and 1.79%, respectively. In line-level prediction, Compared to the baseline method, PTLVD not only improved the Top-5 Accuracy by 1.61%, but also successfully reduced the Mean First Ranking by 5.08%.
Tao Peng 0006, Shixu Chen, Junwei Tang, Junping Liu, Xinrong Hu
SCAM1
2023 BovdGFE: buffer overflow vulnerability detection based on graph feature extraction
Xinghang Lv, Tao Peng 0006, Jia Chen 0012, Junping Liu, Xinrong Hu, Ruhan He, Minghua Jiang, Wenli Cao
Appl. Intell.2
2023 GSNet: Generating 3D garment animation via graph skinning network
abstract
The goal of digital dress body animation is to produce the most realistic dress body animation possible. Although a method based on the same topology as the body can produce realistic results, it can only be applied to garments with the same topology as the body. Although the generalization-based approach can be extended to different types of garment templates, it still produces effects far from reality. We propose GSNet, a learning-based model that generates realistic garment animations and applies to garment types that do not match the body topology. We encode garment templates and body motions into latent space and use graph convolution to transfer body motion information to garment templates to drive garment motions. Our model considers temporal dependency and provides reliable physical constraints to make the generated animations more realistic. Qualitative and quantitative experiments show that our approach achieves state-of-the-art 3D garment animation performance.
Tao Peng 0006, Jiewen Kuang, Jinxing Liang, Xinrong Hu, Jiazhe Miao, Feng Yu 0017, Minghua Jiang
Graph. Model.1
2023 Smart Clothing System With Multiple Sensors Based on Digital Twin Technology
abstract
Smart clothing is widely used for social safety, health monitoring, and sports monitoring. Current research focuses on the use of various materials or sensors to implement smart clothes with different functions, which implies that the functionality of smart clothing depends on the number of sensors used. For existing smart clothing systems, the greatest attention has been given to information processing algorithms and assembly of sensors, and the interaction between users and systems is ignored. To address this gap, this article considers a multifunctional smart clothing system constructed with several sensors. The smart clothing system proposed in this article mainly consists of a hardware module and a software module. Four types of sensors are incorporated into the hardware module to monitor the heart rate, blood oxygen saturation, body temperature, locating information, and activity states; the software module includes the 3-D model based on the user and the feedback system based on digital twin (DT) technology. The DT technology can map the fundamental states of users in terms of the monitoring indices from the hardware module, and give correspondent advice to users. This novel smart clothing system overcomes the lack of an interaction function in existing methods and introduces DT technology into smart wearable devices for the first time.
Feng Yu 0017, Minghua Jiang, Zhangyuan Tian, Tao Peng 0006, Xinrong Hu
IEEE Internet Things J.5
2023 ClothSeg: semantic segmentation network with feature projection for clothing parsing
Guangyu Tang, Feng Yu 0017, Huiyin Li, Yankang Shi, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Minghua Jiang
J. Vis. Commun. Image Represent.6
2023 VTON-SCFA: A Virtual Try-On Network Based on the Semantic Constraints and Flow Alignment
abstract
An image-based virtual try-on system transfers an in-shop garment to the corresponding garment region of a reference person, which has huge application potential and commercial value in online clothing shopping. Existing methods have difficulty preserving garment texture and body details because of rough garment alignment and imperfect detail-retention strategies. To address this problem, we propose a virtual try-on network based on semantic constraints and flow alignment. The key idea of the framework is as follows: 1) a global-local semantic predictor (GLSP) is proposed to generate a reasonable target semantic map, which clearly guides the correct alignment of the in-shop garment with the body and the generation of try-on result; and 2) a novel appearance flow-based garment alignment network (AFGAN) is proposed to align the in-shop garment with the body, which is important to preserve maximum garment detail and ensure natural and realistic warping; and 3) we propose a synthesis strategy to integrate the aligned garment and the human body to preserve maximum body detail for generating a realistic result and preventing cross-occlusion and pixel confusion between different body parts. Experiments on the existing benchmark dataset demonstrate that the proposed method achieves the best performance on qualitative and quantitative experiments among the state-of-the-art virtual try-on techniques.
Chenghu Du, Feng Yu 0017, Minghua Jiang, Ailing Hua, Tao Peng 0006, Xinrong Hu
IEEE Trans. Multim.6
2023 VTNCT: an image-based virtual try-on network by combining feature with pixel transformation
Tao Peng 0006, Feng Yu 0017, Ruhan He, Xinrong Hu, Junping Liu, Minghua Jiang
Vis. Comput.2
2023 Three stages of 3D virtual try-on network with appearance flow and shape field
Feng Yu 0017, Minghua Jiang, Ailing Hua, Tao Peng 0006, Xinrong Hu
Vis. Comput.6
2023 Cloth texture preserving image-based 3D virtual try-on
Xinrong Hu, Ruiqi Luo, Junping Liu, Tao Peng 0006
Vis. Comput.6
2022 Few-Shot Detection Based on an Enhanced Prototype for Outdoor Small Forbidden Objects
Jia Chen 0012, Xinzhou Chen, Xinrong Hu, Tao Peng 0006
CGI5
2022 Multi-Pose Virtual Try-On Via Self-Adaptive Feature Filtering
abstract
With the growing trend of virtual try-on, multi-pose tasks attract researchers due to their higher commercial value. Prior methods lack an effective geometric deformation to maintain the original image details resulting in many details loss in the head and garment. To address this problem, we propose a new multi-pose virtual try-on network, which can fit a garment to the corresponding area of a person in arbitrary poses. First, the target pose’s body-semantic distribution is predicted by the target pose point. Second, the in-shop garment and human body are warped based on a human pose to solve the unnatural alignment and the lack of body details by the Deformation Module (DM). Finally, the human body in the given pose and garment is fine generated by the Filtering Synthesis Network (FSN). Compared to state-of-the-art methods with objective experiments on the MPV dataset, the proposed method achieves the best performance in metrics and the rich details in visual results.
Chenghu Du, Feng Yu 0017, Minghua Jiang, Tao Peng 0006, Xinrong Hu
ICASSP5
2022 Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Capture
abstract
3D virtual try-on receives widespread attention from scholars due to its great practical and commercial values. In prior methods, the fundamental problems lie in the limitations on texture retention during garment deformation and the lack of feature context capture during depth estimation. To address these problems, we propose a new 3D virtual try-on network via multi-scale characteristic capture (VTON-MC), which can produce an exact 3D model with the generated photo-realistic monocular image. The main processes are as follows: 1) predicting the human semantic-map and aligning the in-shop garment in the human pose using the appearance flow method, 2) synthesizing the human body and the warped garment to gain the image try-on result, and 3) estimating the human double-depth map of the image try-on result to reconstruct desired 3D try-on mesh by designed Depth Estimation Network (DEN). Extensive experiments on existing benchmark datasets demonstrate that VTON-MC outperforms state-of-the-art approaches efficiently.
Chenghu Du, Feng Yu 0017, Minghua Jiang, Yaxin Zhao, Tao Peng 0006, Xinrong Hu
ICASSP6
2022 An Abnormal Traffic Detection Method for IoT Devices Based on Federated Learning and Depthwise Separable Convolutional Neural Networks
abstract
As a bridge for information interaction between people and things, and things and things, IoT devices bring security issues and data privacy protection issues that have always been the main challenges in the IoT environment. In terms of abnormal traffic detection of IoT devices, data sharing between device data is usually not possible. This makes the deep learning method for model training based on a large amount of data unable to fully exert its strength due to the lack of IoT device attack instances, resulting in the problem of low detection accuracy. To this end, we propose an abnormal traffic detection model for IoT devices, FL-DSCNN (Federated Learning and Depthwise separable convolutional neural networks). First, the mayfly optimization algorithm is used to select the traffic features, and the model training time is reduced by reducing the feature dimension. Then, by introducing the FL framework, the depthwise separable convolutional neural network is used as a local model for collaborative training without sharing private data, avoiding the problem of lack of labeled data due to the “data silos” phenomenon while protecting data privacy. In addition, we experimentally verify the proposed method on the existing public dataset Aposemat IoT-23 dataset and compare and evaluate it with existing methods. The experimental results show that the method can achieve two-class and multi-class detection respectively. The detection accuracy rates of 98.52% and 97.73% prove the progress and superiority of the proposed FL-DSCNN model in the detection of abnormal traffic of IoT devices.
Qinyu Xia, Shi Dong 0001, Tao Peng 0006
IPCCC3
2022 UF-VTON: Toward User-Friendly Virtual Try-On Network
abstract
Image-based virtual try-on aims to transfer a clothes onto a person while preserving both person's and cloth's attributes. However, the existing methods to realize this task require a target clothes, which cannot be obtained in most cases. To address this issue, we propose a novel user-friendly virtual try-on network (UF-VTON), which only requires a person image and an image of another person wearing a target clothes to generate a result of the person wearing the target clothes. Specifically, we adopt a knowledge distillation scheme to construct a new triple dataset for supervised learning, propose a new three-step pipeline (coarse synthesis, clothing alignment, and refinement synthesis) for try-on task, and utilize an end-to-end training strategy to further refine the results. In particular, we design a new synthesis network that includes both CNN blocks and swin-transformer blocks to capture global and local information and generate highly-realistic try-on images. Qualitative and quantitative experiments show that our method achieves the state-of-the-art virtual try-on performance.
Tao Peng 0006, Ruhan He, Xinrong Hu, Junping Liu, Minghua Jiang
ICMR2
2022 PF-VTON: Toward High-Quality Parser-Free Virtual Try-On Network
Tao Peng 0006, Ruhan He, Xinrong Hu, Junping Liu, Minghua Jiang
MMM (1)2
2022 Toward Detail-Oriented Image-Based Virtual Try-On with Arbitrary Poses
Tao Peng 0006, Ruhan He, Xinrong Hu, Junping Liu, Minghua Jiang
MMM (1)2
2022 A Mitmproxy-based Dynamic Vulnerability Detection System For Android Applications
abstract
During the process of pushing patch packets for Android application hotfix, the attacker can hijack and tamper with the dex file due to the lack of adding a digital signature, which leads to code injection with serious consequences. To address the above problems, an dynamic vulnerability detection system based on mitmproxy is primary proposed, which first utilizes mitmproxy to capture all the packets interacted between the client and the server while locating the dex file, then injects the test code into the dex and pushes it to the client for execution using a man-in-the-middle attack, and finally verifies through the log output by the application whether there is a code injection vulnerability. For 1000 applications in the application market, our system successfully detects 34 new unknown applications with dex injection, and the experimental results show that the system is effective in detecting real-world applications with vulnerabilities caused by hotfix.
Xinghang Lv, Tao Peng 0006, Junwei Tang, Ruhan He, Xinrong Hu, Minghua Jiang, Zaihui Deng, Wenli Cao
MSN2
2022 A Mitmproxy-based Dynamic Vulnerability Detection System For Android Applications
abstract
During the process of pushing patch packets for Android application hotfix, the attacker can hijack and tamper with the dex file due to the lack of adding a digital signature, which leads to code injection with serious consequences. To address the above problems, an dynamic vulnerability detection system based on mitmproxy is primary proposed, which first utilizes mitmproxy to capture all the packets interacted between the client and the server while locating the dex file, then injects the test code into the dex and pushes it to the client for execution using a man-in-the-middle attack, and finally verifies through the log output by the application whether there is a code injection vulnerability. For 1000 applications in the application market, our system successfully detects 34 new unknown applications with dex injection, and the experimental results show that the system is effective in detecting real-world applications with vulnerabilities caused by hotfix.
Xinghang Lv, Tao Peng 0006, Junwei Tang, Ruhan He, Xinrong Hu, Minghua Jiang, Zaihui Deng, Wenli Cao
MSN2
2022 Unsupervised Structure Confidence Sampling for Image Inpainting
abstract
Context: Current image inpainting methods show great effects in different applications such as image editing, object removal, art creation and soon, but lack of editability of the inpainting results and convincing unsupervised features.Objective: To improve the existing methods, an optimized framework for image inpainting purpose is proposed based on hierarchical variational auto-encoder (VAE) as well as some optimization strategies.Method: Firstly, the VAE is used to extract the distribution of the features of the masked image in different scales, however, it will cause the distribution offset of extracted features which is unfavorable for image inpainting.Therefore, an optimal strategy that sampling the effective feature and invalid feature separately to avoid the offset of feature distribution of the masked image is integrated into the framework.To further improve the formulation of the proposed framework, the same encoder is used to realize the conversion from two domains to the same domain, which is a benefit to enhance the extraction of effective feature regions.In addition, we also introduce the cycle consistency constraints and GAN constraints into the framework to supervise the inpainting process.Result: Experimental results on the available image dataset demonstrate the effectiveness and superiority of the proposed framework.
Xinrong Hu, Jinxing Liang, Junjie Jin, Junping Liu, Tao Peng 0006, Yuanjun Xia
SEKE6
2022 High fidelity virtual try-on network via semantic adaptation and distributed componentization
abstract
Image-based virtual try-on systems have significant commercial value in online garment shopping. However, prior methods fail to appropriately handle details, so are defective in maintaining the original appearance of organizational items including arms, the neck, and in-shop garments. We propose a novel high fidelity virtual try-on network to generate realistic results. Specifically, a distributed pipeline is used for simultaneous generation of organizational items. First, the in-shop garment is warped using thin plate splines (TPS) to give a coarse shape reference, and then a corresponding target semantic map is generated, which can adaptively respond to the distribution of different items triggered by different garments. Second, organizational items are componentized separately using our novel semantic map-based image adjustment network (SMIAN) to avoid interference between body parts. Finally, all components are integrated to generate the overall result by SMIAN. A priori dual-modal information is incorporated in the tail layers of SMIAN to improve the convergence rate of the network. Experiments demonstrate that the proposed method can retain better details of condition information than current methods. Our method achieves convincing quantitative and qualitative results on existing benchmark datasets.
Chenghu Du, Feng Yu 0017, Minghua Jiang, Ailing Hua, Yaxin Zhao, Tao Peng 0006, Xinrong Hu
Comput. Vis. Media7
2022 Boosting vision transformer for low-resolution borehole image stitching through algebraic multigrid
Jia Chen 0012, Zhenpeng Fu, Xinrong Hu, Tao Peng 0006
Vis. Comput.5
2022 Virtual try-on based on attention U-Net
Xinrong Hu, Jinxing Liang, Feng Yu 0017, Tao Peng 0006
Vis. Comput.6
2021 Subway Driver Behavior Detection Method Based On Multi-features Fusion
abstract
The recognition of subway driver behavior is an important way for early warning of public safety. The current models of behavior recognition focus on action recognition of target objects in large-scene, which are difficult to apply for the subway driver behavior recognition directly because of space-time constraints. RepC3D model is proposed for recognizing subway driver behaviors in the paper. The model fuse the features of C3D model and RepVGG model. Firstly we preprocess the dataset by cutting the subway driver operation video into short videos, then the preprocessed dataset is adopted as the input of RepC3D model and is downsampled with the multiscale convolution layers of the main network VGG, which is used to extract the effective features of the driver's action behavior. Next, as the feature tranning network,RepC3D model identify and classfy the behaviors of the subway driver from the videos. The experimental result shows that the RepC3D model is btteetter than the C3D model and RepVGG model in terms of recognition accuracy, false detection rate, and missed detection rate, the recognition efficiency is also improved. The dataset is available at https://github.com/wtazyy/Datasets.git.
Xinrong Hu, Tao Peng 0006, Junping Liu, Ruhan He
BIBM4
2021 Cascaded Cross-Domain Fusion of Virtual Try-On
abstract
Image-based virtual try-on, aiming to fit new in-shop clothes into a person image, has gained extensive attention in the fields of computer vision and image process community. However, the existing methods are difficult to generate photo-realistic try-on images when large-scale deformations or large occlusions occur. To address this issue, we propose a novel two stage visual try-on network. Specifically, in the first stage, we used a shape matching model to learn the geometric transformation of in-shop clothes. For the second stage, an U-net with cascaded attention mechanism is presented to learn the composition mask which adjust the clothes and rendered persons. The adjusted clothes and the rendered person are combined by the composition mask to get the final try-on result. Experimental results have shown that our method can generate photo-realistic images with no occlusion.
Xinrong Hu, Tao Peng 0006, Mingfu Xiong, Feng Yu 0017, Li Li 0094
BIBM3
2021 DP-VTON: Toward Detail-Preserving Image-Based Virtual Try-on Network
abstract
Image-based virtual try-on systems with the goal of transferring a target clothing item onto the corresponding region of a person have received great attention recently. However, it is still a challenge for the existing methods to generate photo-realistic try-on images while preserving non-target details(Fig. 1). To resolve this issue, we present a novel virtual try-on network, DP-VTON. First, a clothing warping module combines pixel transformation with feature transformation to transform the target clothing. Second, a semantic segmentation prediction module predicts a semantic segmentation map of the person wearing the target clothing. Third, an arm generation module generates arms of the reference image that will be changed after try-on. Finally, the warped clothing, semantic segmentation map, arms image and other non-target details (e.g. face, hair, bottom clothes) are fused together for try-on image synthesis. Extensive experiments demonstrate our system achieves the state-of-the-art virtual try-on performance both qualitatively and quantitatively.1
Tao Peng 0006, Ruhan He, Xinrong Hu, Junping Liu, Minghua Jiang
ICASSP2
2021 VTON-HF: High Fidelity Virtual Try-on Network via Semantic Adaptation
abstract
The image-based virtual try-on network transfers the target garment item to the corresponding region of the human body. Due to its commercial value in online garment shopping, it has attracted extensive attention from researchers. However, the previous virtual try-on methods are interfered heavily by garments in reference images, so they have defects in maintaining details of human upper limbs, neck, and given garment. Therefore, a novel High Fidelity Virtual Try-on Network via Semantic Adaptation (VTON-HF) is proposed to generate a result with better details. The main processes are as follows: 1) Thin Plate Spline (TPS) warps the target garment coarsely, 2) parsing network generates a target semantic map with the coarse warped garment, 3) our novel Semantic Map-based Image Adjustment Network (SMIAN) generates components separately to avoid interference between image parts with different semantics, 4) SMIAN fuses all components to generate the final result. VTON-HF can retain the maximum amount of detail in the reference garment than previous methods. Our novel architecture generates desired results by fusing separately generated components (garment, upper limb, and neck) and unchanging parts of the reference image. Moreover, our SMIAN incorporates a priori multimodal information in the tail layer, which effectively improves the convergence efficiency of the network. Our method achieves state-of-the-art quantitative results on IS, SSIM, PSNR, and FID using the VITON dataset. (see Fig. 1).
Chenghu Du, Feng Yu 0017, Minghua Jiang, Tao Peng 0006, Xinrong Hu
ICTAI6
2021 A Structured Feature Learning Model for Clothing Keypoints Localization
Ruhan He, Yuyi Su, Tao Peng 0006, Jia Chen 0012, Xinrong Hu
MMM (1)3
2021 Wireless Network Abnormal Traffic Detection Method Based on Deep Transfer Reinforcement Learning
abstract
With the continuous development of information technology, the network as the infrastructure of the information age has become an indispensable and vital aspect of our daily lives. With the popularization of 5G technology, the number of handheld devices has increased significantly. Although it has brought great convenience to our production and life, it has also introduced new security risks, making the network more likely to be infiltrated and attacked. Currently, abnormal network traffic detection technology has become a vital part of network security, effectively protecting the network and computer systems from intrusion and maintaining normal operation. In the network abnormal traffic detection experiment based on simulation, most researchers use public and well-known datasets, and different datasets contain different attack samples. When testing on different datasets, the model needs to be retrained, significantly increasing the consumption of computer resources. The paper proposes a wireless network abnormal traffic detection method based on the deep transfer adversarial environment dueling double deep Q-Network (DTAE-Dueling DDQN). First, use the old NSL-KDD dataset to train AE-Dueling DDQN and save the training model weights. Then, use the idea of fine-tuning, transfer the weight of the AE-Dueling DDQN training is completed to the target model, and fine-tune the target model using the newer AWID dataset in the WiFi environment. The experiment compares the current representative deep learning (DL) and deep reinforcement learning (DRL) methods. Experimental results show that our proposed method saves computer resources significantly and achieves good results in all evaluation indicators.
Yuanjun Xia, Shi Dong 0001, Tao Peng 0006
MSN3
2021 A textile fabric classification framework through small motions in videos
Tao Peng 0006, Xianzi Zhou, Junping Liu, Xinrong Hu, Changnian Chen
Multim. Tools Appl.1
2021 Network Abnormal Traffic Detection Model Based on Semi-Supervised Deep Reinforcement Learning
abstract
The rapid development of Internet technology has brought great convenience to our production life, and the ensuing security problems have become increasingly prominent. These problems threaten users’ privacy and pose significant security risks to the normal conduct of many aspects of society, such as politics, economy, culture, and people’s livelihood. The growth of the information transmission rate expands the scope of attacks and provides a more attack environment for intruders. Abnormal detection is an effective security protection technology that can monitor network transmission in real-time, effectively sense external attacks, and provide response decisions for relevant managers. The development of machine learning has also led to the development of abnormal traffic detection technology. The goal has been to use powerful and fast learning algorithms to deal with changing threats and respond in real-time. Most of the current abnormal detection research is based on simulation, using public and well-known datasets. On the one hand, the dataset contains high-dimensional massive data, which traditional machine learning methods cannot be processed. On the other hand, the labeled data scale is far behind the application requirements, and the dataset’s labels are all manually labeled, so the labeling cost is exceptionally high. This paper proposes a semi-supervised Double Deep Q-Network (SSDDQN)-based optimization method for network abnormal traffic detection, mainly based on Double Deep Q-Network (DDQN), a representative of Deep Reinforcement Learning algorithm. In SSDDQN, the current network first adopts the autoencoder to reconstruct the traffic features and then uses a deep neural network as a classifier. The target network first uses the unsupervised learning algorithm K-Means clustering and then uses deep neural network prediction. The experiment uses NSL-KDD and AWID datasets for training and testing and performs a comprehensive comparison with existing machine learning models. The experimental results show that SSDDQN has certain advantages in time complexity and achieved good results in various evaluation metrics.
Shi Dong 0001, Yuanjun Xia, Tao Peng 0006
IEEE Trans. Netw. Serv. Manag.3
2020 HybridGAN: hybrid generative adversarial networks for MR image synthesis
Jia Chen 0012, Mingfu Xiong, Tao Peng 0006, Minghua Jiang, Xiao Qin 0001
Multim. Tools Appl.4
2018 FSLLE: A Fast K Selection Algorithm for Locally Linear Embedding
abstract
Data in a high-dimensional data space may reside in a low-dimensional manifold embedded within the high-dimensional space. Manifold learning discovers intrinsic manifold data structures to facilitate dimensionality reductions. We propose a novel manifold learning technique called fast [Formula: see text] selection for locally linear embedding or FSLLE, which judiciously chooses an appropriate number (i.e., parameter [Formula: see text]) of neighboring points where the local geometric properties are maintained by the locally linear embedding (LLE) criterion. To measure the spatial distribution of a group of neighboring points, FSLLE relies on relative variance and mean difference to form a spatial correlation index characterizing the neighbors’ data distribution. The goal of FSLLE is to quickly identify the optimal value of parameter [Formula: see text], which aims at minimizing the spatial correlation index. FSLLE optimizes parameter [Formula: see text] by making use of the spatial correlation index to discover intrinsic structures of a data point’s neighbors. After implementing FSLLE, we conduct extensive experiments to validate the correctness and evaluate the performance of FSLLE. Our experimental results show that FSLLE outperforms the existing solutions (i.e., LLE and ISOMAP) in manifold learning and dimension reduction. We apply FSLLE to face recognition in which FSLLE achieves higher accuracy than the state-of-the-art face recognition algorithms. FSLLE is superior to the face recognition algorithms, because FSLLE makes a good tradeoff between classification precision and performance.
Jin-Hang Liu, Tao Peng 0006, Kunfang Song, Minghua Jiang, Xinrong Hu, Xiao Qin 0001
Int. J. Comput. Intell. Appl.2
2016 Feedback Control Scheduling in Energy-Efficient and Thermal-Aware Data Centers
abstract
This paper presents a model-predictive control-based scheduling strategy called ThermoRing to reduce cooling costs in data centers. ThermoRing makes use of an online feedback control mechanism to improve thermal management of energy-efficient clusters in a data center. ThermoRing aims at keeping the maximum inlet temperatures of the nodes under a redline temperature limit with little stability errors. Importantly, the ThermoRing approach is capable of dealing with emergency conditions (e.g., node fan shutdown and unexpected rising task arrival rates) by dynamically balancing load among the nodes. ThermoRing incorporates a heat distribution matrix to model the thermal characteristics of a data center housing cluster. ThermoRing is conducive to thermal management in data centers with high-scheduling performance and stability. Using a real-world online bookstore trace, we conduct extensive experiments to compare the performance of ThermoRing with three existing solutions (i.e., C-Oracle, Ad-hoc, and MinHR). The experimental results show that ThermoRing improved the system throughput by more than 10% under regular load conditions and by 40% in emergency cases. ThermoRing also significantly improves the energy efficiency of MinHR, which is a thermal-aware scheduler.
Tao Peng 0006, Xiao Qin 0001, Qiping Hu, Zhijun Fang 0001
IEEE Trans. Syst. Man Cybern. Syst.2