EDBT 2026 Demo / reviewers in the wild / expert
Haiqiang Fei
dblp:171/1094
· DBLP profile ↗
17ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0002-0431-2605ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 since 2021Security and privacy · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ScenarioFuzz-LLM: Enhancing Diversity in Autonomous Driving Scenario Fuzzing with LLMsabstractAs Autonomous Driving Systems (ADS) are increasingly deployed, ensuring their safety in edge cases becomes critical to preventing catastrophic failures. However, the limited ADS test scenario diversity often hinders the discovery of new defects, especially in complex and rare situations. This paper presents ScenarioFuzz- Llm,a novel method that leverages Large Language Models (LLMs) to enhance the diversity of ADS test scenarios. By incorporating LLMs into a genetic algorithm-based testing framework, ScenarioFuzz- Llmdirects the mutation to address diversity bottlenecks, thereby enabling the exploration of a broader range of edge cases. Our experiments demonstrate that ScenarioFuzz- Llmenhances the number of violation sce-narios by 10.51 % outperforming the state-of-the-art methods, and uncovers 24 unique defects in ADS, three of which are previously undiscovered. These results highlight the superiority of our approach in enhancing ADS testing through more diverse and comprehensive simulation scenarios, ultimately improving the safety of ADS. Shenghao Lin, Fansong Chen, Laile Xi, Kaiyu Xie, Yaowen Zheng, Haiqiang Fei, Yuyan Sun, Hongsong Zhu |
CSCWD | 6 |
| 2025 | FirmEE: Firmware Emulation Enhancement via Automated and Dynamic NVRAM ConfigurationabstractFirmware emulation is a critical method for re-searching embedded systems. However, current approaches to Non- Volatile Random Access Memory (NVRAM) emulation often face challenges such as strong hardware dependency, complex parameter configuration, and the need for extensive manual intervention, which result in low emulation success rates and poor network reachability. Additionally, the lack of transparency during the firmware execution process makes it difficult to track and analyze the causes of emulation failures. To address these challenges, this paper introduces FirmEE, a firmware emulation enhancement system that leverages NVRAM-Sim, which automates the modeling of NVRAM peripherals and simulates the interaction between firmware and NVRAM hardware during parameter requests and assignments. FirmEE dynamically optimizes parameter configurations by constructing an NVRAM value exploration space, utilizing the number of basic blocks executed during firmware startup as reward. This approach facilitates large-scale automated firmware emulation, significantly improving both emulation success rates and network reachability. Moreover, FirmEE provides fine-grained monitoring of the firmware execution process, offering enhanced transparency and deeper insights into system behavior. Experimental results show that FirmEE increases the emulation success rate to 79.41 % and the network reachability rate to 73.09% on a custom dataset comprising 301 firmware images from four mainstream router vendors, significantly outperforming existing methods. Qin Si, Lei Cui 0003, Haiqiang Fei, Hongsong Zhu |
CSCWD | 3 |
| 2025 | HF-Mamba: Improving Multimodal Classification via Hierarchical Fusion Based on Mamba
Yimo Ren, Jinfa Wang, Hong Li 0004, Rongrong Xi, Haiqiang Fei, Hongsong Zhu |
DASFAA (2) | 5 |
| 2025 | Steering Large Language Models for Vulnerability DetectionabstractVulnerability detection remains a critical challenge in the field of security. Many existing approaches extract code representations for vulnerability detection. However, these methods often focus on the overall semantics of the code, neglecting to specifically target vulnerability-related semantics. To address this limitation, we propose a novel LLM steering method designed to steer LLMs to focus on vulnerability concepts, thereby enhancing their performance in vulnerability detection. Specifically, we introduce a vulnerability steering vector that represents the concept of vulnerability in the representation space. This vector is generated using a paired vulnerability-patch function dataset, effectively capturing the essence of vulnerabilities. Experimental results demonstrate that the proposed method significantly improves LLMs' performance and notably outperforms existing SOTA methods in vulnerability detection tasks. Furthermore, we validate the cross-language transferability of the steering vector and explore the explainability of vulnerability detection. Jiayuan Li 0002, Lei Cui 0003, Jie Zhang 0121, Haiqiang Fei, Hongsong Zhu |
ICASSP | 4 |
| 2025 | Lazy-ConSnap: On-Demand Memory Persistence for Efficient Continuous VM Snapshots and Low-Latency RollbackabstractVirtual machine snapshots are critical to service reliability and operational agility in cloud and edge infrastructures. However, under continuous snapshotting, frequent checkpoints impose severe runtime and storage costs, especially for workloads with frequent memory changes. We observe that snapshots frequently store pages never used during online rollback: empirical analysis shows approximately 45 % of pages need not be saved before the next checkpoint. In this paper, we propose Lazy-ConSnap, a VM snapshot system that combines lazy persistence with prediction-based optimization to achieve both storage efficiency and runtime performance. Our approach integrates: (1) a lazy-persistence mechanism using cross-snapshot dirty-page bitmaps to defer saves until pages are re-modified; (2) a history-set prediction algorithm that proactively persists hot pages to reduce costly VM exits; (3) an optimized rollback that reuses memory and loads only modified pages during restoration. Our evaluation with typical workloads over$\mathbf{3 0}$-minute periods shows Lazy-ConSnap achieves up to 6.8 % storage savings (up to$\mathbf{1. 5 G B}$saved), up to$\mathbf{1 4. 3 \%}$rollback speedup, and up to$\mathbf{5. 6 \%}$runtime performance improvement (up to 105s saved) compared to lazy-persistence alone, while maintaining prediction precision above$\mathbf{7 3 \%}$. These gains enable efficient continuous VM snapshots with low-latency recovery for modern cloud environments. Ze Qu, Jiami Lin, Lei Cui 0003, Haiqiang Fei, Hongsong Zhu |
ICPADS | 5 |
| 2025 | VulnTeam: A Team Collaboration Framework for LLM-based Vulnerability DetectionabstractSoftware vulnerability detection is a critical challenge in cyber security. With the rise of deep learning and large language models (LLMs), numerous studies have applied these technologies to vulnerability detection. Existing approaches directly employ prompt engineering, chain-of-thought reasoning, and fine-tuning methods on LLMs, but achieve suboptimal results. To effectively leverage LLMs’ powerful reasoning capabilities for vulnerability detection, we propose VulnTeam, a novel team collaboration framework for LLM vulnerability detection inspired by human expert team collaboration. Specifically, we introduce a dual-stage fine-tuning approach where expert models are first fine-tuned using low-rank adaptation to detect vulnerabilities related to different vulnerability syntactic features, followed by instruction fine-tuning of a leader model responsible for the final decision-making. Ultimately, team members (expert models) and the team leader (leader model) collaborate to detect vulnerabilities. Our experimental evaluation across three LLMs and two datasets demonstrates that VulnTeam significantly enhances LLMs’ vulnerability detection performance (average F1-score improvement of 12.51%). Moreover, VulnTeam-enhanced LLMs substantially outperform previous state-of-the-art (SOTA) vulnerability detection methods (average F1-score improvement of 7.78%). Additionally, we analyze computational costs to validate VulnTeam’s practical applicability. Jiayuan Li 0002, Lei Cui 0003, Wenyan Yu, Haiqiang Fei, Hongsong Zhu |
IJCNN | 4 |
| 2025 | Demystifying Feature Engineering in Malware Analysis of API Call SequencesabstractMachine learning (ML) has been widely used to analyze API call sequences in malware analysis, which typically requires the expertise of domain specialists to extract relevant features from raw data. The extracted features play a critical role in malware analysis. Traditional feature extraction is based on human domain knowledge, while there is a trend of using natural language processing (NLP) for automatic feature extraction. This raises a question: how do we effectively select features for malware analysis based on API call sequences? To answer it, this paper presents a comprehensive study of investigating the impact of feature engineering upon malware classification. We first conducted a comparative performance evaluation under three models, Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and Transformer, with respect to knowledgebased and NLP-based feature engineering methods. We observed that models with knowledge-based feature engineering inputs generally outperform those using NLP-based across all metrics, especially under smaller sample sizes. Then we analyzed a complete set of data features from API call sequences, our analysis reveals that models often focus on features such as handles and virtual addresses, which vary across executions and are difficult for human analysts to interpret. Tianheng Qu, Hongsong Zhu, Limin Sun 0001, Haining Wang 0001, Haiqiang Fei, Zhi Li 0018 |
RAID | 5 |
| 2025 | EHFC: Enhanced Format Clustering via Pre-Trained Traffic Model
Zhen Wang 0043, Laile Xi, Haiqiang Fei, Hong Li 0004, Hongsong Zhu |
WASA (1) | 4 |
| 2025 | When LLMs meet cybersecurity: a systematic literature reviewabstractAbstract The rapid development of large language models (LLMs) has opened new avenues across various fields, including cybersecurity, which faces an evolving threat landscape and demand for innovative technologies. Despite initial explorations into the application of LLMs in cybersecurity, there is a lack of a comprehensive overview of this research area. This paper addresses this gap by providing a systematic literature review, covering the analysis of over 300 works, encompassing 25 LLMs and more than 10 downstream scenarios. Our comprehensive overview addresses three key research questions: the construction of cybersecurity-oriented LLMs, the application of LLMs to various cybersecurity tasks, the challenges and further research in this area. This study aims to shed light on the extensive potential of LLMs in enhancing cybersecurity practices and serve as a valuable resource for applying LLMs in this field. We also maintain and regularly update a list of practical guides on LLMs for cybersecurity at https://github.com/tmylla/Awesome-LLM4Cybersecurity . Jie Zhang 0121, Haoyu Bu, Hui Wen 0001, Yongji Liu, Haiqiang Fei, Rongrong Xi, Hongsong Zhu |
Cybersecur. | 5 |
| 2024 | EasyDetector: Using Linear Probe to Detect the Provenance of Large Language ModelsabstractThe rapid development of large language models (LLMs) has driven significant advancements in various applications. However, the intellectual property of these models often faces risks due to unauthorized reproduction or encapsulation by third parties. In this paper, we propose EasyDetector, a novel approach to detect the provenance of LLMs using linear probes. Our method aims to identify the original source model, even if it has been fine-tuned or encapsulated into another model. Specifically, EasyDetector performs classification on the intermediate layer representations of the new model using linear probes of the original model. Models from the same source exhibit high accuracy, while models from different sources yield low accuracy. Extensive experiments on diverse LLMs demonstrate the effectiveness of EasyDetector in detecting model provenance. The proposed method is lightweight and applicable to various model architectures, holding significant importance for protecting the intellectual property of LLMs. Jie Zhang 0121, Jiayuan Li 0002, Haiqiang Fei, Hongsong Zhu |
TrustCom | 3 |
| 2023 | HEMC: a dynamic behaviour analysis system for malware based on hardware virtualisationabstractSince many malwares disguise themselves by encrypting, obfuscating and recompiling, it is not easy for static analysis methods to recognise new or unknown malwares. This paper proposes a novel dynamic analysis technology based on hardware virtualisation to analyse more malwares with lower computational resources. Firstly, it intercepts the system-call functions to achieve on-demand behaviour analysis by setting special permissions in their physical addresses, which can be dynamically acquired when system-call functions are loaded into memory, as well as only monitoring high-risk functions, which take a small part of the whole functions. Then, this paper utilises copy-on-write technique and incremental image capability to reduce hard drive consumption and hard disk replication time. Finally, this paper proposes a novel approach to capture the return value of system-call functions to deeply analyse the poisoned results of malware samples. Meanwhile, a prototype system, called HEMC, is implemented based on QEMU/KVM . The experiments demonstrate that proposed methods outperform existing methods in efficiency and performance on malware dynamic analysis. Zhenquan Ding, Lei Cui 0003, Haiqiang Fei, Yongji Liu, Zhiyu Hao |
Int. J. Inf. Comput. Secur. | 4 |
| 2021 | VulDetector: Detecting Vulnerabilities Using Weighted Feature Graph ComparisonabstractCode similarity is one promising approach to detect vulnerabilities hidden in software programs. However, due to the complexity and diversity of source code, current methods suffer low accuracy, high false negative and poor performance, especially in analyzing a large program. In this paper, we propose to tackle these problems by presenting VulDetector, a static-analysis tool to detect C/C++ vulnerabilities based on graph comparison at the granularity of function. At the key of VulDetector is a weighted feature graph (WFG) model which characterizes function with a small yet semantically rich graph. It first pinpoints vulnerability-sensitive keywords to slice the control flow graph of a function, thereby reducing the graph size without compromising security-related semantics. Then, each sliced subgraph is characterized using WFG, which provides both syntactic and semantic features in varying degrees of security. As for graph comparison, we take full usage of vulnerability graph and patch graph to improve accuracy. In addition, we propose two optimization methods based on analysis of vulnerabilities. We have implemented VulDetector to automatically detect vulnerabilities in software programs with known vulnerabilities. The experimental results prove the effectiveness and efficiency of VulDetector. Lei Cui 0003, Zhiyu Hao, Haiqiang Fei, Xiao-chun Yun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | CHEAPS2AGA: Bounding Space Usage in Variance-Reduced Stochastic Gradient Descent over Streaming Data and Its Asynchronous Parallel Variants
Yaqiong Peng, Haiqiang Fei, Zhenquan Ding, Zhiyu Hao |
ICA3PP (2) | 2 |
| 2020 | pRnR: A Parallel Record-Replay Framework for Virtual MachinesabstractThe record and replay(RnR) technology of virtual machine(VM) provides the ability to reproduce the past execution of a VM deterministically. It has many promising applications in the cloud environment, including fault tolerance, security analysis, and failure diagnosis. Existing studies in this area pay more effort in optimizing the record method, such as reducing performance penalty and storage costs. However, considering that many practical applications follow the record once, replay many mode, the optimization for the replay is more critical, especially for efficiency. In this paper, we propose pRnR, a novel parallel RnR framework, to support efficient replay. By combining the native RnR framework with an improved continuous snapshots mechanism, pRnR divides the full execution into many independent and complete slices, each of which supports arbitrary replay. In addition, it supports two replay modes to improve replay efficiency, i.e., multi-slice parallel replay and multi-dimension parallel replay. Moreover, we apply our pRnR framework to syscall-based diagnosis to demonstrate its usability. The experimental results show that pRnR is more efficient than existing RnR frameworks. Wei Wang 0428, Lei Cui 0003, Zhiyu Hao, Haiqiang Fei, Chonghua Wang, Yaqiong Peng |
ICCD | 4 |
| 2016 | Piccolo: A Fast and Efficient Rollback System for Virtual Machine ClustersabstractRollback is an effective technique to resume the system execution from a recorded intermediate state upon failures. However, in virtualized environments, rollback of a virtual machine cluster (VMC) produces high network traffic and long service disruption, consequentially imposing significant overhead both on network and applications. In this paper, we propose Piccolo, a fast and efficient rollback system, to restore a VMC from snapshot files over datacenter network. We exploit the similarity among VMC snapshots and leverage multicast to deliver the identical pages across VMs placed on disperse hosts, thereby bypassing transmission of a large number of unnecessary pages. In addition to presenting Piccolo, we detail its implementation, and evaluate it by a set of experiments. The results show that Piccolo could achieve a significant reduction in terms of total sent data, network traffic and rollback latency compared to the existing generic rollback techniques. Lei Cui 0003, Zhiyu Hao, Chonghua Wang, Haiqiang Fei, Zhenquan Ding |
ICPP | 4 |
| 2015 | Lightweight Virtual Machine Checkpoint and Rollback for Long-running Applications
Lei Cui 0003, Zhiyu Hao, Haiqiang Fei, Zhenquan Ding, Bo Li 0005, Peng Liu 0044 |
ICA3PP (3) | 4 |
| 2015 | Traffic Replay in Virtual Network Based on IP-Mapping
Zhiyu Hao, Yongzheng Zhang 0002, Zhenquan Ding, Haiqiang Fei |
ICA3PP (4) | 5 |