VLDB 2026 Research / reviewers in the wild / expert
Peng Zhang 0044
dblp:21/1048-44
· DBLP profile ↗
17ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0001-9518-5914ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An LLM-Guided Fuzzing of Proprietary Industrial Communication Protocols with Context Knowledge
Tianci Pan, Huan Qian, Yaowen Zheng, Haining Wang 0001, Peng Zhang 0044, Jiaxing Cheng, Ge Chu, Ke Li 0042, Ming Zhou 0010 |
INFOCOM | 5 |
| 2025 | Dynamic Vulnerability Patching for Heterogeneous Embedded Systems Using Stack Frame ReconstructionabstractExisting dynamic vulnerability patching techniques are not well-suited for embedded devices, especially mission-critical ones such as medical equipment, as they have limited computational power and memory but uninterrupted service requirements. Those devices often lack sufficient idle memory for dynamic patching, and the diverse architectures of embedded systems further complicate the creation of patch triggers that are compatible across various system kernels and hardware platforms. To address these challenges, we propose a hot patching framework called StackPatch that facilitates patch development based on stack frame reconstruction. StackPatch introduces different triggering strategies to update programs stored in memory units. We leverage the exception-handling mechanisms commonly available in embedded processors to enhance StackPatch's adaptability across different processor architectures for control flow redirection. We evaluated StackPatch on embedded devices featuring three major microcontroller (MCU) architectures: ARM, RISC-V, and Xtensa. In the experiments, we used StackPatch to successfully fix 102 publicly disclosed vulnerabilities in real-time operating systems (RTOSes). We applied patching to medical devices, soft programmable logic controllers (PLCs), and network services, with StackPatch consistently completing each vulnerability remediation in less than 260 MCU clock cycles. Ming Zhou 0010, Xupu Hu, Haining Wang 0001, Hui Wen 0001, Limin Sun 0001, Peng Zhang 0044 |
CCS | 7 |
| 2025 | Towards Natural Language-Based Document Image Retrieval: New Dataset and BenchmarkabstractDocument image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic category, e.g., newspapers or receipts. However, these methods struggle to effectively retrieve document images in real-world scenarios where textual queries with fine-grained semantics are usually provided. To bridge this gap, we introduce a new Natural Language-based Document Image Retrieval (NL-DIR) benchmark with corresponding evaluation metrics. In this work, natural language descriptions serve as semantically rich queries for the DIR task. The NL-DIR dataset contains 41K authentic document images, each paired with five high-quality, fine-grained semantic queries generated and evaluated through large language models in conjunction with manual verification. We perform zero-shot and fine-tuning evaluations of existing mainstream contrastive vision-language models and OCR-free visual document understanding (VDU) models. A two-stage retrieval method is further investigated for performance improvement while achieving both time and space efficiency. We hope the proposed NL-DIR benchmark can bring new opportunities and facilitate research for the VDU community. Datasets and codes will be publicly available at huggingface.co/datasets/nianbing/NL-DIR. Xugong Qin, Jun Jie Ou Yang, Peng Zhang 0044, Gangyan Zeng, Hailun Lin |
CVPR | 4 |
| 2025 | CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCRabstractScene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we resort to the Contrastive Language-Image Pre-training (CLIP) models, which have demonstrated the capacity to perceive and understand scene text, making it possible to achieve strictly OCR-free STR. From the perspective of parameter-efficient transfer learning, a lightweight visual position adapter is proposed to provide a positional information complement for the visual encoder. Besides, we introduce a visual context dropout technique to improve the alignment of local visual features. A novel, parameter-free cross-attention mechanism transfers the contrastive relationship between images and text to that between visual tokens and text, producing a rich cross-modal representation, which can be utilized for efficient reranking with a linear classifier. The resulting model, CAYN, which proves that CLIP is Almost all You Need for STR with no more than 0.50M additional parameters required, achieves new state-of-the-art performance on the STR task, with 92.46%/89.49%/85.98% mAP on the SVT/IIIT-STR/TTR datasets. Our findings demonstrate that CLIP can serve as a reliable and efficient solution for OCR-free STR. Xugong Qin, Peng Zhang 0044, Jun Jie Ou Yang, Gangyan Zeng, Wanqian Zhang, Pengwen Dai |
CVPR | 2 |
| 2025 | Malicious DoH Tunnel Traffic Identification Framework Based on DiffFlow-CNNabstractThis study proposes DiffFlow-CNN, a novel framework for identifying malicious DNS over HTTPS (DoH) tunnel traffic, addressing the critical challenge of data imbalance in network security. By transforming network traffic into grayscale images, the framework can leverage context-rich spatial feature extraction to improve the detection accuracy. A diffusion model is employed for data augmentation, generating diverse, highquality malicious traffic samples to mitigate class imbalance. The augmented data is processed by a 2D Convolutional Neural Network (CNN), which effectively classifies traffic into NonDoH, benign DoH, and malicious DoH categories, with further differentiation of malicious types such as Iodine, dns2tcp, and DNSCat2. Experimental results on the CIRA-CIC-DoHBrw-2020 dataset demonstrate that DiffFlow-CNN achieves near-perfect performance, with an accuracy of 99.97%, precision of 99.99%, recall of 99.09%, and F1-score of 99.52% with DiffFlow-CNN. Comparative analysis highlights the superiority of bidirectional flow representations, particularly the Session+MFR method, which leverages early packet information for optimal feature capture. The framework significantly enhances the detection of covert malicious DoH traffic, offering a robust solution for network security management. Weilin Gai, Runqing Zhang, Yunjun Ma, Peng Zhang 0044, Ruoxing Wang |
HPCC | 5 |
| 2025 | LLM-THP: A Large Language Model-Powered Terminal Honeypot Dialogue FrameworkabstractWith the acceleration of digital globalization, cyber threats are showing a trend of complexity and diversification, posing serious security challenges to critical information infrastructure and sensitive data. In this context, the development of efficient and accurate cyber threat detection technologies has become an urgent need to address potential risks and safeguard the security of the digital ecosystem. Terminal Honeypot is a security tool specifically designed to trap and analyze network attacks against end devices. It attracts attackers by simulating real terminal environments, thus collecting attacker behavioral data and attack methods. Development cycle, high resource consumption, and lack of the ability actively adapt to different attacker behaviors. These problems limit the analysis of the depth of the attack and the subsequent collection of attack information. Therefore, in order to adapt to the unknown attacks against Internet devices in the new situation, it is particularly important to design a terminal honeypot that is free from the predefined conditions and can flexibly respond to various attack scenarios. In this paper, a terminal honeypot design method based on LLM (Large Language Model), LLM-THP, is proposed to solve the problem that the existing honeypots are difficult to cope with unknown network threats. Firstly, we study to construct the initial honeypot environment by presetting prompts and design CoT-DPU, Chain-of-Thought Dynamic Prompt Update, which effectively avoids the token overflow problem in multiple attack interactions. For the challenges of attacker interaction, model invocation, and time asynchrony in practical deployment, the LLM-THP framework designs an efficient workflow. Experimental results show that LLM-THP is better than existing mainstream terminal honeypots in terms of honeypot attractiveness, correct response rate, and simulation. Laite Wang, Huan Qian, Weilin Gai, Zhijian Zheng, Peng Zhang 0044, Ruoxing Wang |
HPCC | 6 |
| 2025 | TrapLLM: An LLM-powered Interactive Log-based Honeypot for Real-world Network AttacksabstractWith the continual escalation of cyberattack tactics, zero-day exploits and advanced persistent threats (APTs) characterized by high stealth and dynamic evolution have posed significant challenges to traditional honeypot systems. Existing approaches are limited in service fidelity, interactive intelligence, and awareness of attacker intent, making it difficult to effectively lure advanced attackers or reconstruct threat chains from massive volumes of log data. To address these challenges, this paper introduces large language models (LLMs) as the core driving force to construct an architecture that integrates log-driven data governance with dynamic response generation. Leveraging the semantic understanding and generative capabilities of LLMs, the proposed method enables fine-grained identification of attacker intent, reconstruction of event sequences, and adaptive responses driven by retrieval-augmented generation (RAG), thereby realizing an intelligent closed-loop defense. This innovative integration overcomes the constraints of traditional static rules and low interaction emulation, achieving attack intent analysis and adaptive deception response, and providing an efficient pathway for proactive threat hunting. During a 25day deployment in a real production network environment, the system utilized 25 diversion nodes and 8 types of emulated services to capture over 1.39 million raw attack logs. Analysis revealed multiple attack attempts targeting 9 known Common Vulnerabilities and Exposures (CVE) vulnerabilities, along with a substantial number of high-severity attacks for which specific CVE identifiers could not be determined. Experimental results demonstrate that the proposed approach significantly improves both the accuracy of threat awareness and the timeliness of response, providing reliable support for the evolution of intelligent defense mechanisms. Yunjun Ma, Gangyan Zeng, Peng Zhang 0044, Fuyuan Zhang, Ran Lin, Huan Qian |
TrustCom | 3 |
| 2025 | Encryption Traffic Classification Based on Mining Traffic Context and Transport RelationshipabstractThis paper proposes a novel ETC-MTCTR, which is designed to enable more accurate, versatile and efficient traffic classification in the context of multi-scenario, low-resource encrypted traffic. Through three modules of Datagram Token conversion, pretraining and fine-tuning, the method uses large-scale unlabeled encrypted traffic for pretraining, mining and learning the traffic context and transmission relationship of encrypted traffic classification tasks, so that a small number of labeled data samples can be effectively used in the fine-tuning stage. Significantly improve the performance of the model on specific downstream classification tasks, enhance the accuracy, adaptability and robustness of the model in diverse environments, limited resources and new encryption security protocols, and realize efficient encryption traffic classification in multi-scenario and low-resource background. The results show that ETC-MTCTR achieves the best performance on three tasks: encryption malware classification, VPN encrypted traffic classification, and TLS 1.3 encryption application classification. Its F1 score is improved by 0.22% in the classification task of encrypted malware, 1.4% in the classification task of VPN encrypted traffic App, 4.56% in the classification task of VPN encrypted traffic Service, and 9.89% in the classification task of TLS 1.3 encrypted application, which is significantly better than other comparison methods. Weilin Gai, Runqing Zhang, Peng Zhang 0044 |
WCNC | 6 |
| 2024 | MHPS: Multimodality-Guided Hierarchical Policy Search for Knowledge Graph ReasoningabstractRecently, path inference-based knowledge graph reasoning (KGR) methods have attracted great attention due to their good performance and interpretability. However, as the number of hops increases, the search space grows exponentially, making the reward sparse and the process of reasoning difficult. To alleviate this problem, we propose the Multimodality-guided Hierarchical Policy Search (MHPS) for KGR, which introduces multimodal hierarchical guidance to each layer of policies during policy search. On the one hand, multimodal guidance reserves rich information on different dimensions, providing more opportunities to find better paths. On the other hand, this leads to better interaction between the two agents, resulting in more concise guidance for policy stepping. Experimental results on two public datasets demonstrate that the proposed approach outperforms state-of-the-art methods on multi-hop KGR. Xugong Qin, Peng Zhang 0044, Yongquan He, Xinjian Huang, Ming Zhou 0010, Liehuang Zhu, Qingfeng Tan |
ICASSP | 3 |
| 2024 | Enhancing VPN Traffic Recognition Through CatBoost Feature Extraction and Stacking Ensemble LearningabstractA virtual private network (VPN) often serves as an accessory to conceal the online identities of malicious activities. The identification of VPN tunnels has become a prevalent method for detecting potential security threats or abnormalities. Nevertheless, current deep packet inspection and deep learning approaches encounter challenges such as limited scalability or low accuracy. We introduce a novel approach to address the problems by proposing a supervised protocol-wide flow representation learning approach. Our approach leverages the semantic information inherent in the protocol to generate optimal feature embeddings automatically. Additionally, we propose a stacking ensemble machine algorithm to enhance the accuracy of VPN tunnel identification using the generated feature embeddings. We have implemented a prototype named SA-VPN and conducted a comprehensive evaluation of its effectiveness and efficiency using a significant volume of VPN traffic flows. The results demonstrate that our tool surpasses the performance of current state-of-the-art VPN tunnel identification tools. Ming Zhou 0010, Peng Zhang 0044, Xugong Qin, Xiaoxu Hu |
ICC | 3 |
| 2024 | Improving Multimodal Rumor Detection via Dynamic Graph Modeling
Xiaoxu Hu, Xugong Qin, Peng Zhang 0044, Gangyan Zeng, Runbo Zhao, Xinjian Huang |
ICPR (18) | 4 |
| 2024 | Perception-Enhanced Generative Transformer for Key Information Extraction from Documents
Runbo Zhao, Jun Jie Ou Yang, Xugong Qin, Gangyan Zeng, Xiaoxu Hu, Peng Zhang 0044 |
ICPR (31) | 7 |
| 2024 | Enhancing Feature Selection in IoT Intrusion Detection Using the Ensemble StackingabstractThe Internet of Things (IoT) is increasingly vulnerable to security risks due to new network attacks. Deep learning-based intrusion detection systems (DL-IDS) have emerged as a key solution, but they face challenges like imbalanced datasets and lengthy training times in complex environments. While feature selection algorithms are commonly employed to mitigate these issues, mainstream methods can yield inconsistent results, failing to reflect data characteristics accurately and potentially introducing noise. To address this problem, we propose an ensemble stacking approach to combine multiple feature selection algorithms, thereby minimizing errors from individual approaches. Each feature selection method acts as a base learner to assess feature importance, while logistic regression is a meta-learner to integrate the outputs into a final result. Additionally, we developed a CNN-based intrusion detection model enhanced with BiLSTM and attention mechanisms to improve detection performance. Our approach was tested on the UNSW-NB15 and CIC-IDS2017 datasets, with results indicating a significant improvement in detection performance compared to using a single feature selection method. Zhijian Zheng, Weilin Gai, Peng Zhang 0044, Ming Zhou 0010 |
ISPA | 3 |
| 2024 | Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text RetrievalabstractScene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be available at https://github.com/Gyann-z/FDP. Gangyan Zeng, Yuan Zhang 0013, Dongbao Yang, Peng Zhang 0044, Yiwen Gao 0001, Xugong Qin, Yu Zhou 0015 |
ACM Multimedia | 5 |
| 2024 | SecureNet-AWMI: Safeguarding Network with Optimal Feature Selection AlgorithmabstractDeep learning has emerged as a leading method for detecting network intrusion threats. However, processing large volumes of data increases computational time costs, and noise in the data can reduce detection rates. To address these challenges, feature selection algorithms are essential for balancing time efficiency and detection accuracy. Feature selection algorithms for intrusion detection systems (IDS) face two primary challenges: selecting the most suitable features for the model and managing data imbalances. Traditional methods often rely on manual selection based on feature importance, which can lead to significant computational errors. And they cannot detect attacks with smaller proportions in complex and variable network traffic. We design a secure network intrusion detection framework SecureNet-AWMI to balance attack distribution by augmenting the low-frequency attack samples and reducing the high-frequency attack samples. The core of SecureNet-AWMI is a feature selection component that uses mutual information theory and adjusts weights to account for different types of attacks. To enhance threat detection and classification, we employ an advanced Convolutional Neural Network (CNN) model enhanced with Bidirectional Long Short-Term Memory (BiLSTM) and an attention mechanism. Comparative experiments on three public datasets – CICIDS2017, UNSW-NB15, and NSL-KDD – show that SecureNet-AWMI outperforms current mainstream feature selection and threat classification techniques. Ming Zhou 0010, Zhijian Zheng, Peng Zhang 0044, Sixue Lu, Yamin Xie, Zhongfeng Jin |
TrustCom | 3 |
| 2023 | Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningabstractDue to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection. Despite great progress, these methods show deficiencies in robustness and still suffer from false positives and instance adhesion. Different from existing methods which integrate multiple-granularity features or multiple outputs, we resort to the perspective of representation learning in which auxiliary tasks are utilized to enable the encoder to jointly learn robust features with the main task of per-pixel classification during optimization. For semantic representation learning, we propose global-dense semantic contrast (GDSC), in which a vector is extracted for global semantic representation, then used to perform element-wise contrast with the dense grid features. To learn instance-aware representation, we propose to combine top-down modeling (TDM) with the bottom-up framework to provide implicit instance-level clues for the encoder. With the proposed GDSC and TDM, the encoder network learns stronger representation without introducing any parameters and computations during inference. Equipped with a very light decoder, the detector can achieve more robust real-time scene text detection. Experimental results on four public datasets show that the proposed method can outperform or be comparable to the state-of-the-art on both accuracy and speed. Specifically, the proposed method achieves 87.2% F-measure with 48.2 FPS on Total-Text and 89.6% F-measure with 36.9 FPS on MSRA-TD500 on a single GeForce RTX 2080 Ti GPU. Xugong Qin, Pengyuan Lv, Chengquan Zhang, Yu Zhou 0015, Peng Zhang 0044, Hailun Lin, Weiping Wang 0005 |
ACM Multimedia | 6 |
| 2023 | Vehicle Trajectory Data Mining for Artificial Intelligence and Real-Time Traffic Information ExtractionabstractIt aims to improve the efficiency of information collection and extraction in the current intelligent transportation system, and accurately mine the vehicle trajectory data By using Artificial Intelligence (AI) and Deep Learning methods, the trajectory data generated during vehicle driving are deeply mined and analyzed, and the characteristics of driving behavior of vehicle drivers are modeled and analyzed in detail. Then, a method of mining driving behavior characteristics based on Convolutional Neural Network (CNN) and vehicle trajectory is proposed. Based on the mathematical principle of wavelet packet and Least Square Support Vector Machine (LSSVM), a combined model of trajectory mining is constructed and applied to the short-term prediction of traffic flow. The traffic flow of Binjiang Road and Renmin Road in Guangzhou, Guangdong Province from August 19 to August 21, 2021 is predicted to verify the accuracy of the trajectory mining combined model. The results show that the combination model of data mining has good fitting effect, and the average accuracy is above 0.8. Besides, the effectiveness of the Deep Learning model in driver behavior classification is verified. The accuracy of the classification model is 75.2% for trajectory, and that is 76.8% for driver behavior characteristics. It is of great significance to effectively utilize the knowledge data in Intelligent Transportation System (ITS) and extract valuable information from it, which has certain reference value for the subsequent refined prediction of vehicle behavior. Peng Zhang 0044, Jun Zheng 0008, Hailun Lin, Zhuofeng Zhao, Chao Li 0027 |
IEEE Trans. Intell. Transp. Syst. | 1 |