VLDB 2026 Research / reviewers in the wild / expert
Peng Gao 0005
dblp:29/5999-5
· DBLP profile ↗
26ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0003-2230-3937ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EnMambaHSI: Enhanced multi-scale spatial-spectral Mamba network for hyperspectral image classification
Peng Gao 0005, Wen-Hua Qin, Fei Wang 0036, Hamido Fujita, Ru-Yue Yuan |
Neurocomputing | 2 |
| 2026 | ZRID-Net: Zero-Reference Real-World Image Dehazing Framework via Deep Self-Decoupling and Reverse Knowledge TransferabstractThis paper investigates one of the most challenging problems in single image dehazing: how to restore haze-free scenes solely from the input observed image without relying on paired or unpaired images and how to extract useful prior information from the observed image to guide the dehazing process. To address these challenges, this paper introduces a novel zero-reference real-world image dehazing method via deep self-decoupling and reverse knowledge transfer (ZRID-Net). Specifically, we first employ a model-driven approach to preliminarily decouple the observed image into coarse-grained components: the haze-free image, transmission map, and atmospheric light. Subsequently, we refine the haze-free image and transmission map separately via a data-driven approach. In addition, we propose a novel reverse knowledge transfer method to exploit latent prior information within hazy images thoroughly for dehazing guidance. This method combines knowledge transfer and contrastive learning to reverse guide the refinement network away from haze characteristics. Finally, a perceptual fusion strategy is employed to obtain haze-free images with high visibility and realism. Extensive experiments demonstrate that the proposed ZRID-Net effectively restores image clarity, enhances structural details, and improves color fidelity across various challenging haze conditions without relying on paired or unpaired supervision. On multiple benchmark datasets, ZRID-Net outperforms existing SOTA approaches in terms of both quantitative metrics and visual quality. The results also confirm its strong generalizability and practical applicability to real-world scenarios. The relevant implementation code can be found at https://github.com/cswangshilong/ZRID-Net. Shilong Wang 0005, Wenqi Ren, Peng Gao 0005, Jiguo Yu, Jianlei Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Towards Patch-Based Noise Compression for Adversarial Attack Against Transformer-Based Visual TrackingabstractIn recent years, with the widespread application of Vision Transformer (ViT) in visual trackers, their robustness has received increasing attention. However, by focusing on global interactions between image patches, ViT reduces sensitivity to local noise, posing new challenges for adversarial attacks. Meanwhile, existing decision-based adversarial attack methods often overlook the differences in noise sensitivity between different patches, further limiting the compression efficiency of adversarial noise, especially in ViT. In visual tracking, existing adversarial attack methods primarily target Siamese network-based trackers, and research on adversarial attacks against Transformer-based trackers, particularly decision-based black-box attacks, is still relatively limited. To implement effective black-box attacks on Transformer-based trackers, this paper innovatively proposes patch-based adversarial noise compression (PANC), a decision-based adversarial attack method. This method effectively compresses adversarial noise patch by patch, significantly improving compression efficiency and attack concealment. PANC also introduces a noise sensitivity matrix that dynamically adds and reduces adversarial noise, optimizing the spatial distribution of noise while decreasing the number of queries. We validated the effectiveness of the proposed PANC attack method on several Transformer-based trackers, including OSTrack, STARK, TransT, and MixformerV2, and three public large-scale benchmark datasets: GOT-10k, TrackingNet, and LaSOT. Experimental results show that compared to the existing state-of-the-art adversarial attack method, the IoU attack, PANC compresses the noise level to 10%, improving the attack effectiveness by 162% with the number of queries of only 45.7%. Furthermore, PANC can serve as an initialization or post-processing optimization strategy for other adversarial attack methods, providing a more flexible and efficient mechanism for adversarial example generation. Our work reveals the vulnerabilities of existing Transformer-based visual trackers and offers new ideas for further improving the efficiency and concealment of adversarial attacks. Peng Gao 0005, Wen-Jia Tang, Fei Wang 0036, Hamido Fujita, Hanan Aljuaid, Ru-Yue Yuan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Learning nested attentional feature fusion network for high performance visual tracking
Peng Gao 0005, Xin-Yue Zhang |
Appl. Intell. | 1 |
| 2025 | Searching a lightweight network architecture for thermal infrared pedestrian tracking
Wen-Jia Tang, Xiao Liu 0004, Peng Gao 0005, Fei Wang 0036, Ru-Yue Yuan |
Appl. Intell. | 3 |
| 2025 | Learning multi-level graph attentional representation for thermal infrared object tracking
Peng Gao 0005, Shi-Min Li, Fei Wang 0036, Hamido Fujita, Hanan Aljuaid, Ru-Yue Yuan |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Towards Greedy Iterative Adversarial Attack With Distortion Maps Against Deep Face RecognitionabstractExisting deep learning-based face recognition models are vulnerable to adversarial attacks due to their inherent network fragility. However, current attack methods generate adversarial examples that often suffer from low visual quality and poor transferability. To address these issues, this paper proposes a novel adversarial attack method, G-FRadv, combining greedy iteration with multi-scale distortion maps to enhance both the attack performance and the visual quality of the adversarial examples. Specifically, G-FRadv first fuses images from different scales to obtain multiple distortion maps. These maps are then partitioned, and the disturbance weight map is coupled with the iteratively sorted gradient information. Finally, the adversarial perturbations generated by different distortion maps are fused and applied to the original image. Experimental results show that the proposed G-FRadv method achieves an average attack success rate 11.38% higher than noise-based methods, and 26.53% higher than makeup-based attack methods, while maintaining better visual quality. Peng Gao 0005, Jiu-Ao Zhu, Wen-Hua Qin |
IEEE Signal Process. Lett. | 1 |
| 2025 | Learning Feature Interaction Alignment Network for Siamese Region Proposal Visual TrackingabstractExisting visual trackers based on Siamese region proposal networks determine the target location and size by comparing features of the target template and search regions. However, due to changes in scale, and appearance, inconsistencies in features may arise, thereby affecting tracking performance. To address this issue, we innovatively introduce a feature interaction alignment network to learn feature interactions between the target template and search regions. This enables the tracker to better align the target-specific features in complex scenarios. Through feature interaction, the tracker can perform feature fusion at a more refined level. Furthermore, to enable existing trackers to learn feature interactions, we improve the loss function to better accommodate the needs of feature interaction alignment. The optimized loss function more effectively balances different types of prediction errors, enhancing the model's adaptability to complex scenarios. Experimental results on four public benchmark datasets show that the proposed feature interaction alignment network improves the accuracy and robustness of the baseline trackers, providing a new direction for improving existing visual tracking methods based on Siamese region proposal networks. Lu-Yao Liu 0003, Peng Gao 0005 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Toward Adaptive Meta-Gradient Adversarial Examples for Visual TrackingabstractIn recent years, visual tracking methods based on convolutional neural networks and transformers have achieved remarkable performance and have been successfully applied in fields such as autonomous driving. However, the numerous security issues exposed by deep learning models have gradually affected the reliable application of visual tracking methods in real-world scenarios. Therefore, how to reveal the security vulnerabilities of existing visual trackers through effective adversarial attacks has become a critical problem that needs to be addressed. To this end, we propose an adaptive meta-gradient adversarial attack (AMGA) method for visual tracking. This method integrates multimodel ensemble and meta-learning strategies, combining momentum mechanisms and Gaussian smoothing, which can significantly enhance the transferability and attack effectiveness of adversarial examples. AMGA randomly selects models from a large model repository, constructs diverse tracking scenarios, and iteratively performs both white- and black-box adversarial attacks in each scenario, optimizing the gradient directions of each model. This paradigm minimizes the gap between white- and black-box adversarial attacks, thus achieving excellent attack performance in black-box scenarios. Extensive experimental results on large-scale datasets, such as OTB2015, LaSOT, and GOT-10 k demonstrate that AMGA significantly improves the attack performance, transferability, and deception of adversarial examples. Wei-Long Tian, Peng Gao 0005, Xiao Liu 0004, Hamido Fujita, Hanan Aljuaid, Maoli Wang |
IEEE Trans. Reliab. | 2 |
| 2024 | Encrypted malicious traffic detection based on natural language processing and deep learning
Xiaodong Zang, Tongliang Wang, Peng Gao 0005, Guowei Zhang 0003 |
Comput. Networks | 5 |
| 2024 | Robust visual tracking with extreme point graph-guided annotation: Approach and experiment
Peng Gao 0005, Xin-Yue Zhang, Xiao-Li Yang, Hamido Fujita, Fei Wang 0036 |
Expert Syst. Appl. | 1 |
| 2024 | Self-supervised BGP-graph reasoning enhanced complex KBQA via SPARQL generation
Yan Yang 0008, Peng Gao 0005, Shangqing Zhao, Yuefeng Chen, Man Lan, Aimin Zhou, Liang He 0001 |
Inf. Process. Manag. | 3 |
| 2024 | In defense and revival of Bayesian filtering for thermal infrared object tracking
Peng Gao 0005, Shi-Min Li, Fei Wang 0036, Ruyue Yuan, Hamido Fujita |
Knowl. Based Syst. | 1 |
| 2023 | Stackelberg-Game-Based Intelligent Offloading Incentive Mechanism for a Multi-UAV-Assisted Mobile-Edge Computing SystemabstractWe study the intelligent offloading problem for a multiple unmanned aerial vehicle (multi-UAV)-assisted mobile-edge computing (MEC) system in an MEC scenario where a natural disaster has damaged the edge server. The study has two steps. First, the task offloading destination is determined by minimizing the total energy consumption of the multi-UAVs in the system. We propose the server selection game-theoretic (SSGT) algorithm and demonstrate its convergence through simulation experiments. Second, we propose an offloading incentive mechanism to price computing resources for a single unmanned aerial vehicle (UAV)-MEC server. Considering the UAV’s power consumption and mobile users’ willingness, we model the interaction between the UAV-MEC server and mobile users as a Stackelberg game. We prove the existence of a Nash equilibrium by theoretical analysis and experimental verification and design the multiround iterative game (MRIG) algorithm based on arithmetic descent to achieve the optimal solution, i.e., the utility tradeoff between the UAV-MEC server and mobile users. Finally, the simulation results show that our proposed scheme can increase the value of overall user satisfaction (SoU) more than other schemes, which proves that the incentive mechanism of resource pricing can supply computing power support for ground mobile users in a UAV-assisted MEC system more effectively. Maoli Wang, Peng Gao 0005, Kunlun Yang |
IEEE Internet Things J. | 3 |
| 2023 | IP traffic behavior characterization via semantic mining
Xiaodong Zang, Maoli Wang, Peng Gao 0005, Guowei Zhang 0003 |
J. Netw. Comput. Appl. | 4 |
| 2023 | Speech-Oriented Sparse Attention Denoising for Voice User Interface Toward Industry 5.0abstractThe adoption of voice user interface (VUI) will promote network automation with enhanced efficiency with reduced simplicity and operating expense in Industry 5.0. Given the noisy environments, speech denoising is indispensable for the VUI in Internet of Things (IoT) or Industrial IoT (IIoT). Despite Transformer's recent success in speech denoising, the adopted full self-attention suffers from quadratic complexity, which challenges the computational power of the IoT/IIoT components. Considering the strong local correlations of speech signals, a speech-oriented sparse attention denoising scheme is developed to keep the meaningful local and global dependencies while mitigating the redundant attentions, resulting in a significant reduction in computational complexity. With the full self-attention as the baseline, experimental results revealed that the proposed scheme achieves a better denoising performance and yields a lower computational cost, indicating the strong potential for various VUI application scenarios in IoT and IIoT toward Industry 5.0. Hongxu Zhu, Qiquan Zhang, Peng Gao 0005, Xinyuan Qian 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Cell tracking with multifeature fusionabstractAbstract Cell tracking is currently a powerful tool in a variety of biomedical research topics. Most cell tracking algorithms follow the tracking by detection paradigm. Detection is critical for subsequent tracking. Unfortunately, very accurate detection is not easy due to many factors like densely populated, low contrast, and possible impurities included. Keeping tracking multiple cells across frames suffers many difficulties, as cells may have similar appearance, they may change their shapes, and nearby cells may interact each other. In this paper, we propose a unified tracking-by-detection framework, where a powerful detector AttentionUnet++, a multimodal extension of the Efficient Convolution Operators algorithm, and an effective data association algorithm are included. Experiments show that the proposed algorithm can outperform many existing cell tracking algorithms. Fei Wang 0036, Shidong Jin, Peng Gao 0005 |
J. Supercomput. | 5 |
| 2023 | Learning Dual-Level Deep Representation for Thermal Infrared TrackingabstractThe feature models used by existing Thermal InfraRed (TIR) tracking methods are usually learned from RGB images due to the lack of a large-scale TIR image training dataset. However, these feature models are less effective in representing TIR objects and they are difficult to effectively distinguish distractors because they do not contain fine-grained discriminative information. To this end, we propose a dual-level feature model containing the TIR-specific discriminative feature and fine-grained correlation feature for robust TIR object tracking. Specifically, to distinguish inter-class TIR objects, we first design an auxiliary multi-classification network to learn the TIR-specific discriminative feature. Then, to recognize intra-class TIR objects, we propose a fine-grained aware module to learn the fine-grained correlation feature. These two kinds of features complement each other and represent TIR objects in the levels of inter-class and intra-class respectively. These two feature models are constructed using a multi-task matching framework and are jointly optimized on the TIR object tracking task. In addition, we develop a large-scale TIR image dataset to train the network for learning TIR-specific feature patterns. To the best of our knowledge, this is the largest TIR tracking training dataset with the richest object class and scenario. To verify the effectiveness of the proposed dual-level feature model, we propose an offline TIR tracker (MMNet) and an online TIR tracker (ECO-MM) based on the feature model and evaluate them on three TIR tracking benchmarks. Extensive experimental results on these benchmarks demonstrate that the proposed algorithms perform favorably against the state-of-the-art methods. Qiao Liu 0001, Di Yuan 0002, Nana Fan, Peng Gao 0005, Xin Li 0034, Zhenyu He 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Encrypted DNS Traffic Analysis for Service Intention InferringabstractService intention refers to what service or which service the server provides. The former includes service classification, service type, or service behavior classification. The latter contains service content classification, such as shopping online or uploading and downloading, etc.. Port-based classification and payload-based classification are two widely used service classification schemes, both of which have many limitations, such as only focusing on server-side scenarios or just designing for non-encrypted requests. In this paper, we propose an encryption-independent approach from a network-side perspective by analyzing the communication behavior of the IPs. Firstly, we identify similar service behavior clusters by employing service influence metrics. Then, we devise a semantic mining mechanism to infer whether they serve a fixed user group or provide interactive service. Finally, we use open-source benchmark datasets, synthetic datasets, and the real Netflow data collected from the China Education Research Network backbone (CERNET) to verify our proposal. Experimental results demonstrate that the accuracy and recall rate of the proposed approach is better than other similar state-of-the-art methods. Besides, our work can also distinguish malicious behavior clusters. Extensive experiments demonstrate that our work is efficient for network management and security monitoring. Xiaodong Zang, Maoli Wang, Peng Gao 0005 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2020 | Learning reinforced attentional representation for end-to-end visual tracking
Peng Gao 0005, Qiquan Zhang, Fei Wang 0036, Liyi Xiao, Hamido Fujita, Yan Zhang 0066 |
Inf. Sci. | 1 |
| 2020 | Siamese attentional keypoint network for high performance visual tracking
Peng Gao 0005, Ruyue Yuan, Fei Wang 0036, Liyi Xiao, Hamido Fujita, Yan Zhang 0066 |
Knowl. Based Syst. | 1 |
| 2019 | Joint Multi-frame Detection and Segmentation for Multi-cell Tracking
Zibin Zhou, Fei Wang 0036, Wenjuan Xi, Huaying Chen, Peng Gao 0005, Chengkang He |
ICIG (2) | 5 |
| 2019 | Learning Cascaded Siamese Networks for High Performance Visual TrackingabstractVisual tracking is one of the most challenging computer vision problems. In order to achieve high performance visual tracking in various negative scenarios, a novel cascaded Siamese network is proposed and developed based on two different deep learning networks: a matching subnetwork and a classification subnetwork. The matching subnetwork is a fully convolutional Siamese network. According to the similarity score between the exemplar image and the candidate image, it aims to search possible object positions and crop scaled candidate patches. The classification subnet-work is designed to further evaluate the cropped candidate patches and determine the optimal tracking results based on the classification score. The matching subnetwork is trained offline and fixed online, while the classification subnetwork performs stochastic gradient descent online to learn more target-specific information. To improve the tracking performance further, an effective classification subnetwork update method based on both similarity and classification scores is utilized for updating the classification subnetwork. Extensive experimental results demonstrate that our proposed approach achieves state-of-the-art performance in recent benchmarks. Peng Gao 0005, Yipeng Ma, Ruyue Yuan, Liyi Xiao, Fei Wang 0036 |
ICIP | 1 |
| 2018 | Efficient Multi-level Correlating for Visual Tracking
Yipeng Ma, Chun Yuan 0003, Peng Gao 0005, Fei Wang 0036 |
ACCV (5) | 3 |
| 2018 | Large Margin Structured Convolution Operator for Thermal Infrared Object TrackingabstractCompared with visible object tracking, thermal infrared (TIR) object tracking can track an arbitrary target in total darkness since it cannot be influenced by illumination variations. However, there are many unwanted attributes that constrain the potentials of TIR tracking, such as the absence of visual color patterns and low resolutions. Recently, structured output support vector machine (SOSVM) and discriminative correlation filter (DCF) have been successfully applied to visible object tracking, respectively. Motivated by these, in this paper, we propose a large margin structured convolution operator (LMSCO) to achieve efficient TIR object tracking. To improve the tracking performance, we employ the spatial regularization and implicit interpolation to obtain continuous deep feature maps, including deep appearance features and deep motion features, of the TIR targets. Finally, a collaborative optimization strategy is exploited to significantly update the operators. Our approach not only inherits the advantage of the strong discriminative capability of SOSVM but also achieves accurate and robust tracking with higher-dimensional features and more dense samples. To the best of our knowledge, we are the first to incorporate the advantages of DCF and SOSVM for TIR object tracking. Comprehensive evaluations on two thermal infrared tracking benchmarks, i.e. VOT-TIR2015 and VOT-TIR2016, clearly demonstrate that our LMSCO tracker achieves impressive results and outperforms most state-of-the-art trackers in terms of accuracy and robustness with sufficient frame rate. Peng Gao 0005, Yipeng Ma, Ke Song 0002, Fei Wang 0036, Liyi Xiao |
ICPR | 1 |
| 2018 | High performance visual tracking with circular and structural operators
Peng Gao 0005, Yipeng Ma, Ke Song 0002, Fei Wang 0036, Liyi Xiao, Yan Zhang 0066 |
Knowl. Based Syst. | 1 |