Puming Wang

dblp:175/3338 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0003-1261-8687ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Attention-guided network for infrared unmanned aerial vehicle target detection
Xin Jin 0005, Puming Wang, Shin-Jye Lee, Shaowen Yao 0001, Wangming Lan, Wei Zhou 0011
Eng. Appl. Artif. Intell.4
2026 Adaptive distributed multi-objective collaborative traffic signal control framework based on multi-agent reinforcement learning
Peisong Huang, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao
Future Gener. Comput. Syst.3
2026 CS-DRL: A soft policy update approach for wireless bandwidth allocation using deep reinforcement learning
Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao
Future Gener. Comput. Syst.3
2026 ECMSFNet: Real-Time Infrared Small Target Detection Network With Efficient Convolution and Multiscale Feature Fusion
abstract
With the continuous development of fields such as national defense and military applications, the importance of infrared small target detection (IRSTD) technology based on thermal imaging has become increasingly prominent. However, in practical application scenarios, it remains difficult to effectively extract the features of weak and small targets under low signal-to-noise ratio conditions, while simultaneously suppressing background clutter, preserving target details, and balancing detection accuracy and speed. To address the issues mentioned above, this work proposes a real-time IRSTD network (ECMSFNet) based on efficient convolution and attention-guided multi-scale feature weighting and fusion. During the feature extraction stage, a dual-branch hybrid convolution module (DBHConv) is designed to extract infrared small target features more efficiently. In the feature fusion stage, a three-branch attention-guided module (TBAG) is designed to enhance the input features from both spatial and channel dimensions. By using a three-branch parallel structure to process input features in a differentiated manner, noise is effectively filtered while target detail information is preserved. In addition, to further address the issue of missed detections, a multi-scale feature weighting and fusion module (MSFWF) is designed at the added detection head to adaptively weight the features and optimize the feature propagation path, thereby improving the model’s detection accuracy. Extensive experimental results on multiple datasets demonstrate that the method proposed in this paper outperforms other advanced approaches and achieves a real-time detection speed of 74.63 frames per second. https://github.com/liubiaohua/ECMSFNet.
Xin Jin 0005, Biaohua Liu, Shaowen Yao 0001, Puming Wang
IEEE Internet Things J.5
2025 CoT-NER: A Reasoning Method via Chain of Thought for Chinese Named Entity Recognition
ZhaoKe Long, Shengfa Miao, Puming Wang, Xin Jin 0005, Jing Niu, ShuangFeng Cai
ICIC (24)4
2025 Transferable adversarial attacks for multi-model systems coupling image fusion with classification models
abstract
Abstract Image preprocessing models typically serve as the initial step in advanced visual tasks, aiming to enhance the performance of subsequent tasks. For example, multi-focus image fusion technology significantly improves the performance of downstream semantic classification tasks. However, with the advancement of adversarial attack techniques, these models are facing significant challenges. Previous research has only explored the impact of adversarial attacks on the performance of individual models, lacking an in-depth investigation into the robustness of tasks involving the combination of multiple models. This study aims to delve into the robustness issues of tasks that combine multi-focus image fusion and image classification. To address this challenge, we have designed a new adversarial attack generator specifically for scenarios that combine multi-focus image fusion with image classification. This attack method uses a decision map surrogate model and a binary weight map to precisely add adversarial perturbations to the effective information parts of multi-focus images. It also incorporates attention mechanisms and Grad-CAM technology to optimize the perturbation areas, aiming to disrupt the key features of the fused image to improve the transferability of the attack. Comprehensive experimental results show that this method significantly improves the efficiency of attacks on downstream classification tasks while maintaining the effectiveness of the fusion model.
Xin Jin 0005, Xueshuai Gao, Puming Wang, Shaowen Yao 0001, Wei Zhou 0011
Cybersecur.5
2025 Polyhedral representations with high-frequency for three-dimensional point cloud classification
Xiaoxin Mao, Xue Li 0009, Puming Wang, Xin Jin 0005, Shengfa Miao, Shaowen Yao 0001, Siwang Yang
Eng. Appl. Artif. Intell.3
2025 IDAD: An improved tensor train based distributed DDoS attack detection framework and its application in complex networks
abstract
With the vigorous development of Internet technology, the scale of systems in the network has increased sharply, which provides a great opportunity for potential attacks, especially the Distributed Denial of Service (DDoS) attack. In this case, detecting DDoS attacks is critical to system security. However, current detection methods exhibit limitations, leading to compromises in accuracy and efficiency. To cope with it, three key strategies are implemented in this paper: (i) Using tensors to model large-scale and heterogeneous data in complex networks; (ii) Proposing a denoising algorithm based on the improved and distributed tensor train (IDTT) decomposition, which optimizes the tensor train(TT) decomposition in terms of parallel computation and low-rank estimation; (iii) Combining (i), (ii) and Light Gradient Boosting Machine (LightGBM) classification model, an efficient DDoS attack detection framework is proposed. Datasets CIC-DDoS2019 and NSL-KDD are used to evaluate the framework, and results demonstrate that accuracy can reach 99.19% while having the characteristics of low storage consumption and well speedup ratio.
Qiyuan Fan, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao, Min An
Future Gener. Comput. Syst.3
2025 HPM-GMN: Hierarchical Pooling Multi-level Graph Matching Network
Shengfa Miao, YongKang Mu, Yuling Tian, Yesen Liu, Kuang Li, Puming Wang, Xin Jin 0005, Shaowen Yao 0001
Knowl. Based Syst.9
2025 Mutli-focus image fusion based on guided filter and image matting network
Puchao Zhu, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001
Multim. Tools Appl.3
2025 GDRNet: a channel grouping based time-slice dilated residual network for long-term time-series forecasting
Qingda Bao, Shengfa Miao, Xin Jin 0005, Puming Wang, Shaowen Yao 0001, Da Hu, Ruoshu Wang
J. Supercomput.5
2025 BDIP: An Efficient Big Data-Driven Information Processing Framework and Its Application in DDoS Attack Detection
abstract
With the rapid advancement of 5G communication technology in the era of big data, massive terminal devices connected to the Internet have dramatically increased the scale of network, generating a large amount of high-dimensional and heterogeneous information. This not only enhances the difficulty of information processing in the network, but also poses a severe challenge to data storage and calculation, which has become a big data problem to be solved urgently. To cope with it, this paper proposes an efficient information processing framework and applies it to Distributed Denial of Service (DDoS) attack detection. Overall, three major highlights are made: (i) Tensor is used to represent multi-modal information in large-scale networks; (ii) A novel denoising algorithm based on tensor train(TT) decomposition is proposed, focused on optimizing both computation and correlation; (iii) A big data-driven information processing framework is developed, which includes information preprocessing, denoising and classification. Results in case study indicate that the framework can achieve an accuracy of 99.19%, all while maintaining the great storage advantage, well speedup ratio and strong computing capabilities under the same computational complexity. It can also be generalized to other network data processing scenarios.
Qiyuan Fan, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001, Shengfa Miao, Sizhang Li, Min An
IEEE Trans. Netw. Serv. Manag.3
2025 Crafting imperceptible and transferable adversarial examples: leveraging conditional residual generator and wavelet transforms to deceive deepfake detection
Xin Jin 0005, Puming Wang, Shin-Jye Lee, Shaowen Yao 0001, Wei Zhou 0011
Vis. Comput.4
2024 Enhanced YOLOv7 Model for Aerial Drone Detection in Complex Environments
abstract
With the growing popularity of Unmanned Aerial Vehicles (UAVs) in civilian, commercial, and military applications, the need for robust drone detection systems has become increasingly urgent. However, due to the small size of drones when observed from a distance, most traditional machine learning and two-stage deep learning detection methods currently struggle to capture effective feature information of drones in complex image backgrounds, and often fail to meet the requirements for real-time performance. To address this challenge, we have introduced an improved version of the YOLOv7-Tiny model, known as YOLOv7-ADD, which significantly enhances the detection performance of small target drones at various distances and under complex backgrounds. The model integrates the Scylla-IoU (SIoU) loss function to improve the accuracy of bounding box regression and employs the BiFormer attention mechanism, which dynamically focuses on key features of drones within the detection scene, enhancing the model’s recognition capabilities for drones. Furthermore, the introduction of the Diverse Branch Block (DBB) helps the model capture multi-scale features, optimizing the detection effect for drones of various sizes. Extensive experiments on drones dataset have demonstrated the superior performance of YOLOv7-ADD. It achieved a 2.7% improvement in [email protected] and a 1.1% increase in [email protected]:0.95, while maintaining high FPS detection performance. This provides an efficient drone detection solution for aerial surveillance. The code is available on https://github.com/FuChanglong/ADD.git.
Changlong Fu, Yasu Wu, Xin Jin 0005, Puming Wang
ISPA4
2024 Dif-GAN: A Generative Adversarial Network with Multi-Scale Attention and Diffusion Models for Infrared-Visible Image Fusion
abstract
To obtain fused images with rich information, visible and infrared images are combined. Most current fusion techniques provide decent results. However, they have shortcomings in extracting information of the source images. This limitation prevents the fused images from adequately considering thermal radiation regions and texture details. As a result, the detailed texture information of the source visible image in the final fusion image is much more than the thermal target information of the source infrared image, or vice versa. Since features at a single scale fail to adequately capture the spatial details of complex scenes, a multi-scale attention network is used to extract the deep feature information of source images. For latent variable issues, the Expectation Maximization (EM) technique can yield maximum likelihood estimates. This not only stabilizes the training of the Generative Adversarial Network (GAN) but also aids in addressing the issue of labels lacking in the fusion of visible and infrared images. Although the EM algorithm framework can greatly enhance the training stability of GAN models, the improvement in fusion quality is not large. Therefore, a diffusion model is introduced into the generator to capture the potential joint structure information between infrared and visible images. Massive experiments show that Dif-GAN outperforms the state-of-the-art.
Chengyi Pan, Xiuliang Xi, Xin Jin 0005, Huangqimei Zheng, Puming Wang, Qiang Jiang
ISPA5
2024 Online public opinion time series prediction based on improved N-Beat and multimodal hybrid fusion
abstract
Social media offers a promising way to analyze online public opinion, which has drawn extensive attention from various sectors. In academia, most studies focus on predicting public opinion using unimodal time series methods, paying little attention to multimodal approaches. However, public opinion may be affected by various complex social factors, so it is necessary to explore multimodal elements. Based on the N-Beats model, we propose a novel model, HFN-BeatsConv, which employs a powerful modal alignment strategy, 3D-TCN. Most fuses multimodal data for public opinion prediction. The model employs a component, 3D-TCN, for modal alignment, differs from other research in that it focuses on the time at which the text appears. Subsequently, the model, HFN-BeatsConv, the N-Beats model is enhanced through the utilization of 3D-TCN, which enables the processing of multivariate time series data and the reduction of multimodal time series forecast error. To increase the usability and sustainability of research, this study provides a valuable social media dataset, as a supplementary feature of time series prediction. Through extensive experiments, the proposed method outperforms the existing methods.
Yuling Tian, Shengfa Miao, Shaowen Yao 0001, Puming Wang, Xin Jin 0005
ISPA4
2024 A novel multi-modal incremental tensor decomposition for anomaly detection in large-scale networks
Rongqiao Fan, Qiyuan Fan, Xue Li 0009, Puming Wang, Xin Jin 0005, Shaowen Yao 0001
Inf. Sci.4
2023 A theoretical analysis of continuous firing condition for pulse-coupled neural networks with its applications
Xin Jin 0005, Pingfan Zhang, Youwei He, Puming Wang, Jingyu Hou 0001, Wei Zhou 0011, Shaowen Yao 0001
Eng. Appl. Artif. Intell.5
2022 RM2T2C: Retrospective Multivariate Multistep Transition Tensor Chain Model for User Mobility Pattern Prediction
abstract
With the bloom of intelligent devices over the world, a large scale of user’s trajectory data are collected. How to mine valuable rules from these data and provide services for the industrial community has become an urgent problem. In this article, we propose a multimodal prediction system to infer users’ mobility pattern embedded in heterogeneous data from cyber–physical–social space. According to users’ mobility pattern, the framework can provide smart services for the industrial community. The highlight is the retrospective multivariate multistep transition tensor (${\text{M}^2}{\text{T}^2}$) chain model, which decomposes a large scale of${\text{M}^2}{\text{T}^2}$into a series of small-scale transition tensors (subtransition tensors) with the tensor maximum likelihood estimation method. Then, that one can solve the stationary probability distribution with the small-scale subtransition tensors so as to highly reduce the computation and storage cost. At the same time, the tensor maximum likelihood estimation method avoids the overfitting of${\text{M}^2}{\text{T}^2}$, so the proposed model improves the performance of prediction systems. In the end, several experiments are constructed to evaluate the proposed model.
Puming Wang, Laurence T. Yang, Xue Li 0009, Xiaokang Zhou
IEEE Trans. Ind. Informatics1
2022 TT-TSVD: A Multi-modal Tensor Train Decomposition with Its Application in Convolutional Neural Networks for Smart Healthcare
abstract
Smart healthcare systems are generating a large scale of heterogenous high-dimensional data with complex relationships. It is hard for current methods to analyze such high-dimensional healthcare data. Specifically, the traditional data reduction methods can not keep the correlation among different modalities of data objects, while the latest methods based on tensor singular value decomposition are not effective for data reduction, although they can keep the correlation. This article presents a tensor train-tensor singular value decomposition (TT-TSVD) algorithm for data reduction. Particularly, the presented algorithm balances the correlation-preservation ability of modalities and data reduction ability by combining the advantages of the train structure of the tensor train decomposition and the association relationship between the tensor singular value decomposition retention mode. Extensive experiments are conducted on the convolutional neural network and the results clearly show that the presented algorithm performs effectively for data reduction with a low-loss classification accuracy; what is more, classification accuracy on medical image dataset has been improved a little.
Debin Liu, Laurence T. Yang, Puming Wang, Ruonan Zhao, Qingchen Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 HWOA: an intelligent hybrid whale optimization algorithm for multi-objective task selection strategy in edge cloud computing system
Yan Kang 0003, Xuekun Yang, Bin Pu, Xiaokang Wang 0001, Haining Wang 0006, Puming Wang
World Wide Web7
2021 A fully-automatic image colorization scheme using improved CycleGAN with skip connections
Shanshan Huang 0001, Xin Jin 0005, Jie Li 0023, Shin-Jye Lee, Puming Wang, Shaowen Yao 0001
Multim. Tools Appl.6
2021 The Cyber-Physical-Social Transition Tensor Service Framework
abstract
Cyber-Physical-Social Systems (CPSS) are the extension of Cyber-Physical Systems (CPS) which integrate social characteristics into Cyber-Physical Systems. The emergence of CPSS provides a promising solution for energy and sustainability challenges, at the same time, CPSS make sustainability computing space change from single space to multi-spaces. To cope with the problem, we propose Cyber-Physical-Social Transition Tensor service framework to provide sustainable services for multi-spaces. To the best of our knowledge, this is the first framework that can support multi-space services at one time. The framework solves various stationary probability distributions arising from the multi-relational data from CPSS, then provides services for multi-spaces based on the probability distributions. In order to solve the stationary probability distributions, we develop a novel iterative algorithm, named Intersect Tensor Power Method (ITPM). Furthermore, we prove the convergence of the iterative algorithm, the existence and uniqueness of the stationary probability distribution. This paper's highlight is the Cyber-Physical-Social Transition Tensor (CPST2) model which leverages multivariate transition tensor to seamlessly integrate massive data from CPSS. A case study of multi-space prediction service illustrates our framework, the accuracy reaches 95 percent with three kinds of humans, two hundred and fifty traffic flow slices, and nineteen time slices.
Puming Wang, Laurence T. Yang, Gongwei Qian, Feng Lu 0003
IEEE Trans. Sustain. Comput.1
2020 Data-driven software defined network attack detection : State-of-the-art and perspectives
Puming Wang, Laurence T. Yang, Zhian Ren, Liwei Kuang
Inf. Sci.1
2020 MMDP: A Mobile-IoT Based Multi-Modal Reinforcement Learning Service Framework
abstract
With the development of GPS technology, a new Mobile Internet of Things (M-IoT) is emerging, which perceives the city's rhythm and pulse day and night to collect a large scale of city data. It is urgent to innovate M-IoT service system for these large-scale and heterogeneous data. To cope with the problem, this article proposes a Mobile-IoT based multi-modal reinforcement learning service framework from data perspective, which has three highlights, i) Developing Action-aware High-order Transition Tensor (AHTT) to fuse the heterogeneous data from M-IoTs in a unified form. ii) Developing Multi-modal Markov Decision Process (MMDP) to model the multi-modal reinforcement learning for M-IoT service framework. iii) Developing Tensor Policy Iteration algorithm (TPIA) to solve the optimal tensor policy. Due to using tensor keeps the multi-modal relations of the context information in the process of solving the optimal policy. The proposed M-IoT service system provides more personalized service for taxi drivers. The experiment results shows that most taxi drivers earn more revenue according to the tensor policy.
Puming Wang, Laurence T. Yang, Xue Li 0009, Xiaokang Zhou
IEEE Trans. Serv. Comput.1