Hongsong Zhu

dblp:57/5368 · DBLP profile ↗
← Back
119ranked-venue papers
1as first author
90since 2021 · last 2026
0000-0003-3720-7403ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 39 · 1 first-author · 20 since 2021Security and privacy · 22 · 17 since 2021Artificial intelligence and machine learning · 16 · 14 since 2021Human-computer interaction and ubiquitous computing · 14 · 14 since 2021Software engineering, systems software and programming languages · 11 · 11 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Automated Construction of High-Quality Initial Seed Corpus for Network Protocol Fuzzing
Weicheng Lin, Laile Xi, Yaowen Zheng, Shenghao Lin, Jiaxing Cheng, Zhen Wang 0043, Shizhao Tian, Tianheng Qu, Hongsong Zhu
INFOCOM10
2026 Modubin: A Binary Modularization Approach Based on the Locality of Homologous Functions
Wenyan Yu, Lei Cui 0003, Jiayuan Li 0002, Hong Li 0004, Hongsong Zhu
ICPC7
2026 Bridge: High-Order Taint Vulnerabilities Detection in Linux-Based IoT Firmware
Jiaqian Peng, Puzhuo Liu, Yicheng Zeng, Yongji Liu, Hongsong Zhu
SP7
2026 StruFSM: Byte-level structural modeling for protocol finite state machine inference
Zhen Wang 0043, Yimo Ren, Zhaoteng Yan, Hong Li 0004, Hongsong Zhu
Comput. Networks8
2026 Network Intrusion Detection System Based on Enhanced Dual Gaussian Mixture Variational Autoencoders for Internet of Vehicles
abstract
The Internet of Vehicles (IoV) is highly vulnerable to attacks due to its open communication environment, with new types of attacks continuously emerging. However, existing Network Intrusion Detection Systems (NIDS) often fall short in accuracy, recall rates, false positive rates, and they frequently prove ineffective in detecting new attacks. In this paper, we propose a NIDS for IoV based on Enhanced Dual Gaussian Mixture Variational Autoencoders (EDGMVAE) to detect various attack behaviors, including new attacks. We employ two separate Gaussian Mixture Variational Autoencoders (GMVAE) to conduct unsupervised reconstruction training on normal traffic and known attack traffic. During detection, we obtain the tested traffic’s reconstruction probabilities using these GMVAEs and determine whether an attack has occurred through logical fusion operation. We use two datasets for evaluation, i.e.,the Car Hacking Dataset (CHD) for in-vehicle communication and the UNSW-NB15 Dataset (UBD) for external network communication. The results demonstrate that our method achieves remarkable performance in detecting known attacks, with an accuracy rate of 98.54% on the CHD and 98.59% on the UBD. Moreover, it outperforms other state-of-the-art methods in detecting new attacks. Specifically, on the CHD, the accuracy of the method is 6.66% to 12.83% higher than other methods, while on the UBD, the accuracy is 5.13% to 21.91% higher than other methods.
Shizhao Tian, Yaowen Zheng, Laile Xi, Shenghao Lin, Hongsong Zhu
IEEE Trans. Intell. Transp. Syst.6
2026 SMTFL: Secure Model Training to Untrusted Participants in Federated Learning
abstract
Federated learning is an essential distributed model training technique. However, threats such as gradient inversion attacks and poisoning attacks pose significant risks to both privacy of training data and model correctness. We propose SMTFL, a novel approach for secure model training in federated learning. To safeguard gradients privacy against gradient inversion attacks, clients are dynamically grouped, allowing one client's gradient to be divided to obfuscate the gradients of other clients within the group. This method incorporates checks and balances to reduce the collusion for inferring specific client data. To detect poisoning attacks from malicious clients, we assess the impact of aggregated gradients on the global model's performance, enabling effective identification and exclusion of malicious clients. Each client's gradients are encrypted and stored, with decryption collectively managed by all clients. The detected poisoning gradients are invalidated from the global model through an unlearning method. Compared to related work, SMTFL does not rely on trusted participants, avoids the performance degradation caused with traditional noise-injection, and avoids complex homomorphic encryption during gradient aggregation. SMTFL is evaluated on five datasets, and these results demonstrate its effectiveness in defending against gradient inversion and poisoning attacks. The model accuracy is nearly restored to its pre-attack state when SMTFL is deployed. Furthermore, SMTFL achieves over 95% accuracy in identifying malicious clients while maintaining a false positive rate for honest clients within 5%, which is 6% lower than the latest methods.
Xiaorong Dong, Yimo Ren, Jianhua Wang 0004, Hongsong Zhu, Yongle Chen
IEEE Trans. Mob. Comput.6
2026 PVDetector: Pretrained Vulnerability Detection on Vulnerability-enriched Code Semantic Graph
abstract
Automated vulnerability detection is a critical issue in software security. The advent of Deep Learning (DL) has led to numerous studies employing DL to detect vulnerabilities in software source code. However, existing approaches still perform poorly, particularly with real-world vulnerabilities, due to the difficulty in accurately capturing their properties. To this end, we introduce PVDetector, a DL-based approach that utilizes rich code semantics, incorporates vulnerability knowledge, and leverages pretrained code representations for precise vulnerability detection. At its core, PVDetector employs a new model called Vulnerability-enriched Code Semantic Graph (VCSG), which accurately characterizes functions by distinguishing the semantics of identical variables and more finely capturing control dependencies, data dependencies, and vulnerability relationships. Additionally, we introduce four pretraining tasks specifically designed to learn the semantics of control, data, vulnerability, and variables from the VCSG model. These pretraining tasks significantly enhance PVDetector’s capability to detect vulnerabilities in downstream tasks. Experimental results indicate that PVDetector outperforms SOTAs by 5.0–12.5% in precision, 0.2–9.7% in recall, and 3.0–15.1% in F1-score. Additionally, it supports six programming languages and demonstrates high efficiency (e.g., 10.6 \(\times\) faster than DeepDFA). When applied to seven software products, PVDetector discovered 55 vulnerabilities, including 10 silently patched flaws that had not been previously reported.
Jiayuan Li 0002, Lei Cui 0003, Jie Zhang 0121, Rongrong Xi, Hongsong Zhu
ACM Trans. Softw. Eng. Methodol.6
2025 ScenarioFuzz-LLM: Enhancing Diversity in Autonomous Driving Scenario Fuzzing with LLMs
abstract
As Autonomous Driving Systems (ADS) are increasingly deployed, ensuring their safety in edge cases becomes critical to preventing catastrophic failures. However, the limited ADS test scenario diversity often hinders the discovery of new defects, especially in complex and rare situations. This paper presents ScenarioFuzz- Llm,a novel method that leverages Large Language Models (LLMs) to enhance the diversity of ADS test scenarios. By incorporating LLMs into a genetic algorithm-based testing framework, ScenarioFuzz- Llmdirects the mutation to address diversity bottlenecks, thereby enabling the exploration of a broader range of edge cases. Our experiments demonstrate that ScenarioFuzz- Llmenhances the number of violation sce-narios by 10.51 % outperforming the state-of-the-art methods, and uncovers 24 unique defects in ADS, three of which are previously undiscovered. These results highlight the superiority of our approach in enhancing ADS testing through more diverse and comprehensive simulation scenarios, ultimately improving the safety of ADS.
Shenghao Lin, Fansong Chen, Laile Xi, Kaiyu Xie, Yaowen Zheng, Haiqiang Fei, Yuyan Sun, Hongsong Zhu
CSCWD8
2025 Abnormal Driving Behavior Detection: Deep Reinforcement Learning Based on Expert Guidance
abstract
Abnormal driving behavior is a leading cause of road accidents. Traditional detection methods, relying on classification or unsupervised learning, struggle with accuracy and generalization. Deep Reinforcement Learning (DRL) offers potential but faces challenges such as handling unlabeled and imbalanced data, designing effective reward functions, and ensuring efficient exploration. To address these, we propose an expert-guided DRL framework that integrates a CNN-BiLSTM-Self Attention (CBSA) model and an Isolation Forest (iForest) to guide Proximal Policy Optimization (PPO), enhancing detection accuracy and computational efficiency. Our framework consists of two stages. First, a deep learning model trained on labeled data provides expert guidance. Second, unlabeled data is processed through the pre-trained model and iForest, refining the DRL model via a sparse reward function to detect unknown anomalies while mitigating class imbalance and improving generalization. Experiments on an open-source dataset show our method out-performs baselines, achieving the highest recall (0.70) and Fl-score (0.63). Additionally, attention maps and anomaly heatmaps enhance interpretability, confirming its effectiveness for real-time abnormal driving behavior detection and improved driver safety.
Shenghao Lin, Zhen Wang 0043, Fansong Chen, Yonghe Guo, Yuyan Sun, Hongsong Zhu
CSCWD7
2025 SecRAG: A Graph-Enhanced RAG Framework with Dynamic Prompt for Cybersecurity Applications
abstract
In this paper, we introduce SecRAG, a novel Retrieval-Augmented Generation (RAG) system specifically designed for cybersecurity applications. SecRAG tackles fundamental challenges in context precision and domain-specific terminology through a dual-pronged approach: 1) a data augmentation method optimized for cybersecurity contexts, particularly addressing the RAG system's numerical information sensitivity limitations; and 2) an enhanced dual-level retrieval architecture that integrates graph-based knowledge representation and adaptive text indexing, incorporating a dynamic prompt weighting mechanism based on dual similarity metrics (δ1, δ2) for query relationship and output coherence assessment, to enable comprehensive information discovery. Experimental evaluation on the SecEval benchmark demonstrates that SecRAG achieves stantial improvements over conventional RAG implementations, with a 30% increase in vulnerability node recall rates and reduction in temporal confusion rate to 1.1%. The system has shown particular strengths in specialized areas, achieving overall accuracy rate of 72.02% on the SecEval benchmarks. Our framework effectively addresses the unique challenges of cybersecurity-focused RAG systems, particularly in handling precise numeric attributes and maintaining contextual coherence in multi-turn dialogues.
Yu Qiao 0001, Jie Zhang 0121, Hongsong Zhu
CSCWD6
2025 FirmEE: Firmware Emulation Enhancement via Automated and Dynamic NVRAM Configuration
abstract
Firmware emulation is a critical method for re-searching embedded systems. However, current approaches to Non- Volatile Random Access Memory (NVRAM) emulation often face challenges such as strong hardware dependency, complex parameter configuration, and the need for extensive manual intervention, which result in low emulation success rates and poor network reachability. Additionally, the lack of transparency during the firmware execution process makes it difficult to track and analyze the causes of emulation failures. To address these challenges, this paper introduces FirmEE, a firmware emulation enhancement system that leverages NVRAM-Sim, which automates the modeling of NVRAM peripherals and simulates the interaction between firmware and NVRAM hardware during parameter requests and assignments. FirmEE dynamically optimizes parameter configurations by constructing an NVRAM value exploration space, utilizing the number of basic blocks executed during firmware startup as reward. This approach facilitates large-scale automated firmware emulation, significantly improving both emulation success rates and network reachability. Moreover, FirmEE provides fine-grained monitoring of the firmware execution process, offering enhanced transparency and deeper insights into system behavior. Experimental results show that FirmEE increases the emulation success rate to 79.41 % and the network reachability rate to 73.09% on a custom dataset comprising 301 firmware images from four mainstream router vendors, significantly outperforming existing methods.
Qin Si, Lei Cui 0003, Haiqiang Fei, Hongsong Zhu
CSCWD6
2025 LibRI: A Module Analysis Framework for Identifying Complex Reuse Relationship in Binaries
abstract
With the rapid advancement of collaborative software development, the increased reuse of third-party libraries (TPLs) has introduced new security challenges. Detecting reuse relationships between binary programs and TPLs is vital for software maintenance, vulnerability tracing, and component analysis. However, most existing detection methods are confined to identifying code reuse between binaries, often misclassifying nested and pseudo-propagation reuse as direct reuse. While some methods attempt to analyze these intricate relationships, they depend on pre-collected source code structures or unstable constants, which compromises their generality and accuracy. To tackle these challenges, we introduce LibRI, a framework for identifying complex dependency relationships in C/C++ binaries. LibRI modularizes and matches binaries to analyze module-matching scenarios across multiple TPLs, accurately determining true reuse relationships and identifying original module sources. This facilitates the construction of a detailed reuse relationship graph. Experimental results show that LibRI achieves an accuracy of 0.966 in detecting actual direct reuse, significantly outperforming existing methods. In addition, LibRI is able to build a vulnerability propagation graph of TPLs, identify the propagation paths of TPL vulnerabilities, demonstrating its potential in vulnerability tracking.
Wenyan Yu, Siyuan Li 0014, Mingjiang Huang, Rongrong Xi, Hongsong Zhu
CSCWD6
2025 Microservice Dependency Discovery Based on Spatio-Temporal Network Flow Behavior Modeling
abstract
The microservice architectural style offers significant application scalability and development advantages. Composing monolithic systems into loosely coupled, containerized services reduces deployment and development costs while enhancing the flexibility and resilience of the overall system. However, in large-scale internet applications involving multi-party collaboration and global deployment, the formation of complex microservice dependencies increases the risk of cascading failures. It complicates the process of identifying the source of a fault. Identifying such dependencies in uncontrolled conditions with limited observational data represents a significant challenge. This paper proposes Microservice Dependency Discovery based on Spatio-Temporal Network Flow Behaviour Modelling (Cross-MSDD). This method infers microservice dependencies by modeling spatiotemporal interactions of network traffic. The method employs network flow characteristics to mine frequent behavioral patterns, thereby inferring dependency chains without the necessity for additional tracing tools. The method utilizes network flow characteristics to mine frequent behavior patterns, inferring dependency chains without additional tracing tools, minimizing system interruptions, and protecting request content privacy. To verify the effectiveness of the proposed method, a semi-simulated experimental environment was set up using traffic data from typical microservice applications. The results demonstrate that the method attains an accuracy rate exceeding 96.3% in cross-domain dependency identification, markedly surpassing the performance of existing techniques. This facilitates the practical detection of faults and the mitigation of cascading failure risks, thereby ensuring system stability.
Jinfa Wang, Chunyang Zheng, Hui Wen 0001, Hong Li 0004, Hongsong Zhu
CSCWD6
2025 Research on TTP Data Augmentation Methods Based on the ATT&CK Framework
abstract
As cyber threats escalate, rapid identification and response to attacks are increasingly vital. Cyber Threat Intelli-gence (CTI) is crucial for understanding the threat landscape, and standardized attack frameworks are essential for effective anal-ysis. The MITRE ATT &CK framework has gained widespread adoption for its systematic description of Tactics, Techniques, and Procedures (TTP), aiding security teams in tracking at-tack patterns. However, manual classification of TTP is time-consuming and costly, hindering response efficiency. Although artificial intelligence has advanced automated TTP classification, accuracy still needs improvement due to the scarcity of labeled data, resulting in small and imbalanced datasets. This study introduces a novel TTP data augmentation method to enhance classification accuracy through synthetic data gen-eration. We construct a dataset of 19,716 sentences from the ATT&CK knowledge base and real-world threat reports, Ini-tially, we leverage large language models (LLMs) combined with prompt techniques to generate high-quality synthetic data, followed by semantic filtering and dynamic sampling strategies to further enhance data quality and improve class balance. Experimental results show an average$\mathbf{F}_{1}$score increase of 16.95 % across various classification models, significantly enhancing TTP classification performance.
Xiaodong Xue, Jie Zhang 0121, Tianheng Qu, Rongrong Xi, Hongsong Zhu
CSCWD7
2025 Migratability of Adversarial Node Attacks in Graph Neural Networks: A Comprehensive Study
abstract
Graph Neural Networks (GNNs) are able to discriminate and retain information about the intrinsic properties of nodes, their position in the graph structure, and their relationships to other nodes, thus enabling prediction of node classes. However, GNNs are vulnerable to attacks when performing node categorization tasks. It has been observed that adversarial attacks demonstrate remarkable migration capabilities, suggesting that such attacks are not constrained to a specific model but can achieve comparable attack outcomes on other models as well. This paper presents the inaugural comprehensive examination of the migratability of adversarial attacks within the domain of graph neural networks. In the course of our experiments, we subjected the source graph neural network model to five distinct escape attacks, encompassing a range of dimensions, including the optimization strategy, graph topology, gradient information, reinforcement learning, and agent model. We then migrate the adversarial samples generated by these attacks to the target graph neural network model and evaluate their impact. Our experiments validate the migratability of the adversarial attacks in three different scenarios, i.e., migration across datasets, migration across models, and migration across models and datasets. The experimental results show that the five adversarial attack methods selected in this paper for the node classification task exhibit good migratability in three different migration scenarios. The attacks have a considerable impact on the source graph neural network, and they also demonstrate robust attack effects after migrating to the target graph neural network.
Guoli Zhao, Zhenlu Tan, Kaiyu Xie, Hongsong Zhu
CSCWD7
2025 HF-Mamba: Improving Multimodal Classification via Hierarchical Fusion Based on Mamba
Yimo Ren, Jinfa Wang, Hong Li 0004, Rongrong Xi, Haiqiang Fei, Hongsong Zhu
DASFAA (2)6
2025 Steering Large Language Models for Vulnerability Detection
abstract
Vulnerability detection remains a critical challenge in the field of security. Many existing approaches extract code representations for vulnerability detection. However, these methods often focus on the overall semantics of the code, neglecting to specifically target vulnerability-related semantics. To address this limitation, we propose a novel LLM steering method designed to steer LLMs to focus on vulnerability concepts, thereby enhancing their performance in vulnerability detection. Specifically, we introduce a vulnerability steering vector that represents the concept of vulnerability in the representation space. This vector is generated using a paired vulnerability-patch function dataset, effectively capturing the essence of vulnerabilities. Experimental results demonstrate that the proposed method significantly improves LLMs' performance and notably outperforms existing SOTA methods in vulnerability detection tasks. Furthermore, we validate the cross-language transferability of the steering vector and explore the explainability of vulnerability detection.
Jiayuan Li 0002, Lei Cui 0003, Jie Zhang 0121, Haiqiang Fei, Hongsong Zhu
ICASSP6
2025 Leveraging Fine-Tuned Large Language Models for Device Fingerprint Extraction in IoT Security
Haoyu Bin, Gaosheng Wang, Yimo Ren, Zhi Li 0018, Hongsong Zhu
ICIC (4)6
2025 MalDenoise: Enhancing Robustness of API-Based Malware Detection Against Adversarial Attacks
abstract
API call sequence-based machine learning (ML) models have shown promise in detecting malware with high accuracy. However, the vulnerability of ML models, particularly deep neural networks, to adversarial attacks remains a pressing concern. Unfortunately, existing research efforts have yet to provide a comprehensive defense against various types of adversarial attacks targeted at malware models relying on API call sequences. In this paper, we propose an innovative defense method, named MalDenoise, aiming to fortify the robustness of the API call sequence-based malware detection model against adversarial attacks. It employs a denoising model to filter and denoise the input data for the detection model, eliminating extraneous or irrelevant APIs while accentuating critical components. Furthermore, we have evaluated the performance of MalDenoise against four various adversarial attacks. The results conclusively showcase the superiority of MalDenoise over existing defense methods, attesting to its heightened effectiveness in safeguarding against adversarial threats.
Zuhui Yue, Hongsong Zhu
ICME6
2025 Lazy-ConSnap: On-Demand Memory Persistence for Efficient Continuous VM Snapshots and Low-Latency Rollback
abstract
Virtual machine snapshots are critical to service reliability and operational agility in cloud and edge infrastructures. However, under continuous snapshotting, frequent checkpoints impose severe runtime and storage costs, especially for workloads with frequent memory changes. We observe that snapshots frequently store pages never used during online rollback: empirical analysis shows approximately 45 % of pages need not be saved before the next checkpoint. In this paper, we propose Lazy-ConSnap, a VM snapshot system that combines lazy persistence with prediction-based optimization to achieve both storage efficiency and runtime performance. Our approach integrates: (1) a lazy-persistence mechanism using cross-snapshot dirty-page bitmaps to defer saves until pages are re-modified; (2) a history-set prediction algorithm that proactively persists hot pages to reduce costly VM exits; (3) an optimized rollback that reuses memory and loads only modified pages during restoration. Our evaluation with typical workloads over$\mathbf{3 0}$-minute periods shows Lazy-ConSnap achieves up to 6.8 % storage savings (up to$\mathbf{1. 5 G B}$saved), up to$\mathbf{1 4. 3 \%}$rollback speedup, and up to$\mathbf{5. 6 \%}$runtime performance improvement (up to 105s saved) compared to lazy-persistence alone, while maintaining prediction precision above$\mathbf{7 3 \%}$. These gains enable efficient continuous VM snapshots with low-latency recovery for modern cloud environments.
Ze Qu, Jiami Lin, Lei Cui 0003, Haiqiang Fei, Hongsong Zhu
ICPADS6
2025 TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability Code
abstract
Code reuse in software development frequently facilitates the spread of vulnerabilities, making the scope of affected software in CVE reports imprecise. Traditional methods primarily focus on identifying reused vulnerability code within target software, yet they cannot verify if these vulnerabilities can be triggered in new software contexts. This limitation often results in false positives. In this paper, we introduce TransferFuzz, a novel vulnerability verification framework, to verify whether vulnerabilities propagated through code reuse can be triggered in new software. Innovatively, we collected runtime information during the execution or fuzzing of the basic binary (the vulnerable binary detailed in CVE reports). This process allowed us to extract historical traces, which proved instrumental in guiding the fuzzing process for the target binary (the new binary that reused the vulnerable function). TransferFuzz introduces a unique Key Bytes Guided Mutation strategy and a Nested Simulated Annealing algorithm, which transfers these historical traces to implement trace-guided fuzzing on the target binary, facilitating the accurate and efficient verification of the propagated vulnerability. Our evaluation, conducted on widely recognized datasets, shows that TransferFuzz can quickly validate vulnerabilities previously unverifiable with existing techniques. Its verification speed is 2.5 to 26.2 times faster than existing methods. Moreover, TransferFuzz has proven its effectiveness by expanding the impacted software scope for 15 vulnerabilities listed in CVE reports, increasing the number of affected binaries from 15 to 53. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/TransferFuzz.
Siyuan Li 0014, Yuekang Li, Zuxin Chen, Chaopeng Dong, Yongpan Wang, Hong Li 0004, Yongle Chen, Hongsong Zhu
ICSE8
2025 VulnTeam: A Team Collaboration Framework for LLM-based Vulnerability Detection
abstract
Software vulnerability detection is a critical challenge in cyber security. With the rise of deep learning and large language models (LLMs), numerous studies have applied these technologies to vulnerability detection. Existing approaches directly employ prompt engineering, chain-of-thought reasoning, and fine-tuning methods on LLMs, but achieve suboptimal results. To effectively leverage LLMs’ powerful reasoning capabilities for vulnerability detection, we propose VulnTeam, a novel team collaboration framework for LLM vulnerability detection inspired by human expert team collaboration. Specifically, we introduce a dual-stage fine-tuning approach where expert models are first fine-tuned using low-rank adaptation to detect vulnerabilities related to different vulnerability syntactic features, followed by instruction fine-tuning of a leader model responsible for the final decision-making. Ultimately, team members (expert models) and the team leader (leader model) collaborate to detect vulnerabilities. Our experimental evaluation across three LLMs and two datasets demonstrates that VulnTeam significantly enhances LLMs’ vulnerability detection performance (average F1-score improvement of 12.51%). Moreover, VulnTeam-enhanced LLMs substantially outperform previous state-of-the-art (SOTA) vulnerability detection methods (average F1-score improvement of 7.78%). Additionally, we analyze computational costs to validate VulnTeam’s practical applicability.
Jiayuan Li 0002, Lei Cui 0003, Wenyan Yu, Haiqiang Fei, Hongsong Zhu
IJCNN6
2025 Hybrid Multi-stage Decoding for Few-shot NER with Entity-aware Contrastive Learning
abstract
Few-shot named entity recognition can identify new types of named entities based on a few labeled examples. Previous methods employing token-level or span-level metric learning suffer from the computational burden and a large number of negative sample spans. In this paper, we propose the Hybrid Multistage Decoding for Few-shot NER with Entity-aware Contrastive Learning (MsFNER), which splits the general NER into two stages: entity-span detection and entity classification. There are 3 processes for introducing MsFNER: training, finetuning, and inference. In the training process, we train and get the best entity-span detection model and the entity classification model separately on the source domain using meta-learning, where we create a contrastive learning module to enhance entity representations for entity classification. During finetuning, we finetune the both models on the support dataset of target domain. In the inference process, for the unlabeled query data, we first detect the entity-spans, then the entity-spans are jointly determined by the entity classification model and the KNN. We conduct experiments on the open FewNERD dataset and FewAPTER dataset, the results demonstrate the advance of MsFNER.
Congying Liu, Gaosheng Wang, Xingyuan Wei, Hongsong Zhu
IJCNN5
2025 Lares: LLM-driven Code Slice Semantic Search for Patch Presence Testing
abstract
In modern software ecosystems, 1-day vulnerabilities pose significant security risks due to extensive code reuse. Identifying vulnerable functions in target binaries alone is insufficient; it is also crucial to determine whether these functions have been patched. Existing methods, however, suffer from limited usability and accuracy. They often depend on the compilation process to extract features, requiring substantial manual effort and failing for certain software. Moreover, they cannot reliably differentiate between code changes caused by patches or compilation variations.To overcome these limitations, we propose Lares, a scalable and accurate method for patch presence testing. Lares introduces Code Slice Semantic Search, which directly extracts features from the patch source code and identifies semantically equivalent code slices in the pseudocode of the target binary. By eliminating the need for the compilation process, Lares improves usability, while leveraging large language models (LLMs) for code analysis and SMT solvers for logical reasoning to enhance accuracy. Experimental results show that Lares achieves superior precision, recall, and usability. Furthermore, it is the first work to evaluate patch presence testing across optimization levels, architectures, and compilers. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/Lares.
Siyuan Li 0014, Yaowen Zheng, Hong Li 0004, Jingdong Guo, Chaopeng Dong, Chunpeng Yan, Weijie Wang 0005, Yimo Ren, Limin Sun 0001, Hongsong Zhu
ASE10
2025 Demystifying Feature Engineering in Malware Analysis of API Call Sequences
abstract
Machine learning (ML) has been widely used to analyze API call sequences in malware analysis, which typically requires the expertise of domain specialists to extract relevant features from raw data. The extracted features play a critical role in malware analysis. Traditional feature extraction is based on human domain knowledge, while there is a trend of using natural language processing (NLP) for automatic feature extraction. This raises a question: how do we effectively select features for malware analysis based on API call sequences? To answer it, this paper presents a comprehensive study of investigating the impact of feature engineering upon malware classification. We first conducted a comparative performance evaluation under three models, Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and Transformer, with respect to knowledgebased and NLP-based feature engineering methods. We observed that models with knowledge-based feature engineering inputs generally outperform those using NLP-based across all metrics, especially under smaller sample sizes. Then we analyzed a complete set of data features from API call sequences, our analysis reveals that models often focus on features such as handles and virtual addresses, which vary across executions and are difficult for human analysts to interpret.
Tianheng Qu, Hongsong Zhu, Limin Sun 0001, Haining Wang 0001, Haiqiang Fei, Zhi Li 0018
RAID2
2025 Automated Penetration on Multi-Subnet Environments with Dual-Stage DRL Models
abstract
With the advent of artificial intelligence techniques, the field of Network Attack Defense (NAD) has witnessed a surge in research efforts towards automating penetration testing (PenTest). Our work presents a dual-stage PenTest model aiming at predicting attack paths in network topology and determining payload for vulnerabilities in hosts with deep reinforcement learning models. While constructing training environments, our approach integrates real-world vulnerability environments with virtual network topologies. This allows the model to take into account the process of vulnerability validation with success rate compared to existing work based on fully virtualized targets, while retaining the efficiency of deployment and training provided by virtualization. And we introduce a method that simulate hierarchical network topology with randomized subnets to simulate complex network environments, challenging the agent to adapt and learn effective policies across diverse configurations of the target networks. Our experiments demonstrate the effectiveness of our model in various network sizes. In addition, the results indicate that our approach not only achieves high performance but also maintains stability under the different success rate of vulnerability exploitation, showcasing the robustness and adaptability. Our work contributes to the advancement of automated PenTest by providing a more generalized and efficient solution.
Haoyu Bu, Hui Wen 0001, Hongsong Zhu, Hong Li 0004, Xirui Song, Yimo Ren
SMC3
2025 EHFC: Enhanced Format Clustering via Pre-Trained Traffic Model
Zhen Wang 0043, Laile Xi, Haiqiang Fei, Hong Li 0004, Hongsong Zhu
WASA (1)7
2025 When LLMs meet cybersecurity: a systematic literature review
abstract
Abstract The rapid development of large language models (LLMs) has opened new avenues across various fields, including cybersecurity, which faces an evolving threat landscape and demand for innovative technologies. Despite initial explorations into the application of LLMs in cybersecurity, there is a lack of a comprehensive overview of this research area. This paper addresses this gap by providing a systematic literature review, covering the analysis of over 300 works, encompassing 25 LLMs and more than 10 downstream scenarios. Our comprehensive overview addresses three key research questions: the construction of cybersecurity-oriented LLMs, the application of LLMs to various cybersecurity tasks, the challenges and further research in this area. This study aims to shed light on the extensive potential of LLMs in enhancing cybersecurity practices and serve as a valuable resource for applying LLMs in this field. We also maintain and regularly update a list of practical guides on LLMs for cybersecurity at https://github.com/tmylla/Awesome-LLM4Cybersecurity .
Jie Zhang 0121, Haoyu Bu, Hui Wen 0001, Yongji Liu, Haiqiang Fei, Rongrong Xi, Hongsong Zhu
Cybersecur.9
2025 A Differential Privacy Based Task Offloading Algorithm for Vehicular Edge Computing
abstract
With the advent of Vehicular Ad Hoc Networks (VANETs), Vehicular Edge Computing (VEC) facilitates the execution of vehicular tasks through the Internet. In the VEC architecture, vehicles request task offloading, and a central decision center allocates resources. Effective task offloading algorithms provide optimal and equitable decisions based on objectives such as task latency and system overhead; however, current task offloading algorithms for VEC face challenges in adapting to complex and dynamic road environments. This paper proposes a task-offloading algorithm based on deep reinforcement learning to address the challenges of task offloading in vehicular edge computing. During the task offloading process, the privacy of vehicular task data may be compromised. This study introduces a novel task-offloading algorithm for Vehicular Edge Computing (VEC) that employs differential privacy principles to safeguard the confidentiality of vehicular tasks during the offloading process. The proposed algorithm introduces noise in accordance with the privacy budget during the training process. The study provides a theoretical analysis of privacy, and experimental results based on Attari demonstrate that the proposed differential privacy-based reinforcement learning algorithm exhibits superior convergence compared to existing algorithms. Veins-based simulation experiments on VEC demonstrate that the proposed differential privacy-based task-offloading algorithm can achieve practical offloading while preserving privacy.
Jun Li 0085, Shuqin Zhang, Jinbu Geng, Jizhao Liu, Zenan Wu, Hongsong Zhu
IEEE Internet Things J.6
2025 Automated tactics planning for cyber attack and defense based on large language model agents
Yimo Ren, Jinfa Wang, Hui Wen 0001, Hong Li 0004, Hongsong Zhu
Neural Networks6
2025 Backsolver: Adapting Preceding Execution Paths to Solve Constraints for Concolic Execution
abstract
Concolic execution follows the execution paths of concrete inputs, capable of generating new inputs for unexplored code by solving negated path constraints. However, implicit flows can hinder concolic execution, reducing the code coverage. Implicit flows occur when inputs influence control flow, and the control flow variation affects the values of some variables. During concolic execution, the preceding path selections limit the potential values of these variables. This limitation may result in unsolvable constraints, subsequently restricting the generation of new inputs for unexplored paths. Our insight is that following the same preceding paths is unnecessary, and we can adapt preceding paths to make the latest constraints solvable. We divide states into general states and implicit-flow-solving states (IFSSs). We utilize the general states to perform concolic execution. When solving constraints influenced by implicit flows, we switch to the IFSSs. We use the IFSSs to explore the relevant code region and adapt paths. To mitigate path explosion and construct the relation between inputs and the variables, we merge the IFSSs. State merging does not burden the general states, and we limit the code regions for the IFSSs to minimize the introduced overhead. Finally, we replace the variable symbols in the target constraints with new expressions and attempt to solve the new constraints. We implement our approach in Backsolver and build a test suite to evaluate it. Backsolver successfully identifies all the implicit flows in the test suite and resolves most of them. When evaluated on six real-world binaries, Backsolver resolves the highest number of branches related to implicit flows in total. Besides, Backsolver has the highest code coverage in PlutoSVG and finds a 0-day vulnerability. We reported the vulnerability and obtained a CVE ID.
Yicheng Zeng, Zhanwei Song, Guo Lv, Hongsong Zhu, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.5
2025 RDGV: Reputation-Driven Gradual Verification for Tampering Localization in Cooperative Task Offloading
abstract
In edge computing, offloading complex computation tasks from terminals to nearby edge nodes (ENs) is a critical solution. When an EN cannot complete the tasks independently, it offloads partial tasks to other ENs. These ENs return the results to the original EN, which integrates them before sending the final output to the terminal. This process, called cooperative task offloading, introduces significant security challenges, especially since ENs are typically provided by third parties. Some ENs, driven by self-interest or vulnerability to attack, may provide incorrect results to other ENs (i.e., acting as malicious ENs), ultimately causing the terminal to receive incorrect results. While existing schemes can help terminals detect incorrect results, they fail to locate the malicious ENs, leaving the system in an unreliable state and causing invalid computations based on erroneous intermediate results. We propose a reputation-driven gradual verification scheme (RDGV) to identify and locate malicious ENs. In RDGV, each EN is held accountable for the correctness of its results and faces penalties if the results are found to be incorrect. Successor ENs must verify the intermediate results before utilizing them. That is, gradual verification. An economic incentive rule counters potential attacks from malicious ENs, while reputation, representing EN’s trustworthiness, guides a personalized verification strategy to reduce overall verification overhead. The effectiveness and advantages of RDGV are shown by simulation results and comparison with related work. The findings indicate that honest and continuous service is the optimal strategy for ENs to maintain the credibility of the edge system.
Siyan Zhu, Hongsong Zhu, Yongle Chen
IEEE Trans. Reliab.5
2025 TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated Vulnerability
abstract
Code reuse in software development frequently facilitates the spread of vulnerabilities, leading to imprecise scopes of affected software in CVE reports. Traditional methods focus primarily on detecting reused vulnerability code in target software but lack the ability to confirm whether these vulnerabilities can be triggered in new software contexts. In previous work, we introduced the TransferFuzz framework to address this gap by using historical trace-based fuzzing. However, its effectiveness is constrained by the need for manual intervention and reliance on source code instrumentation. To overcome these limitations, we propose TransferFuzz-Pro, a novel framework that integrates Large Language Model (LLM)-driven code debugging technology. By leveraging LLM for automated, human-like debugging and Proof-of-Concept (PoC) generation, combined with binary-level instrumentation, TransferFuzz-Pro extends verification capabilities to a wider range of targets. Our evaluation shows that TransferFuzz-Pro is significantly faster and can automatically validate vulnerabilities that were previously unverifiable using conventional methods. Notably, it expands the number of affected software instances for 15 CVE-listed vulnerabilities from 15 to 53 and successfully generates PoCs for various Linux distributions. These results demonstrate that TransferFuzz-Pro effectively verifies vulnerabilities introduced by code reuse in target software and automatically generation PoCs.
Siyuan Li 0014, Kaiyu Xie, Yuekang Li, Hong Li 0004, Yimo Ren, Limin Sun 0001, Hongsong Zhu
IEEE Trans. Software Eng.7
2024 Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
abstract
Mining structured knowledge from tweets using named entity recognition (NER) can be beneficial for many downstream applications such as recommendation and intention under standing. With tweet posts tending to be multimodal, multimodal named entity recognition (MNER) has attracted more attention. In this paper, we propose a novel approach, which can dynamically align the image and text sequence and achieve the multi-level cross-modal learning to augment textual word representation for MNER improvement. To be specific, our framework can be split into three main stages: the first stage focuses on intra-modality representation learning to derive the implicit global and local knowledge of each modality, the second evaluates the relevance between the text and its accompanying image and integrates different grained visual information based on the relevance, the third enforces semantic refinement via iterative cross-modal interactions and co-attention. We conduct experiments on two open datasets, and the results and detailed analysis demonstrate the advantage of our model.
Hong Li 0004, Yimo Ren, Jie Liu 0079, Shuaizong Si, Hongsong Zhu, Limin Sun 0001
AAAI6
2024 AS-Fuzzer: An Optimized ADS Fuzzing Method via Scenario Segmentation and Parallel Evolution
abstract
Autonomous Driving Systems (ADS) hold significant potential for enhancing travel convenience. Ensuring the reliability of ADS through efficient and comprehensive simulation testing has garnered substantial attention from researchers. In recent years, various search-based automated test scenario generation methods have been proposed to identify potential ADS defects. However, these methods still face challenges in balancing efficiency with comprehensive testing and suffer from a lack of diversity. To address these challenges, we propose AS-Fuzzer, an optimized ADS fuzzing method based on composite traffic scenario generation. AS- Fuzzer introduces a scenario slicing technique based on traffic road structures, allowing each sliced scenario to evolve independently and in parallel, enhancing interaction rates and balancing efficiency with comprehensive testing. A novel scenario generation method Co-Evolutionary Genetic Algorithms (CEGA), is applied within AS-Fuzzer to improve the diversity of generated scenarios, thereby exploring a wider range of ADS defects. Experimental results demonstrate that the proposed method improves test scenario generation efficiency by 120% compared to the state-of-the-art baseline method. Additionally, in the simulation testing of the proposed method, the interaction rate between the ADS vehicle and non-player characters is 2.79 times that of the baseline method, thereby further enhancing ADS testing efficiency. Furthermore, within the same time frame, the proposed method uncovered 19 types of ADS defects that other baseline methods did not explore, achieving higher ADS defect diversity.
Fansong Chen, Shenghao Lin, Weicheng Lin, Laile Xi, Yongji Liu, Hongsong Zhu
APSEC7
2024 An Efficient Vehicular Intrusion Detection Method Based on Edge Intelligence
abstract
The advancement of the Internet of Vehicles (IoV) has facilitated the integration of intelligent vehicles with Internet connectivity, providing access to a wide range of services that significantly enhance vehicular applications. However, this connectivity also brings about an increased vulnerability to cyber attacks from the internet. Given the limited computing and communication resources available in vehicles, existing intrusion detection methods are ill-suited for vehicular networks. In this paper, we propose a lightweight vehicular intrusion detection method based on Edge Intelligence. The proposed method utilizes edge intelligence to achieve real-time intrusion detection in vehicles. To ensure efficient intrusion detection, we design a lightweight Convolutional Neural Networks (CNN) intrusion detection model and incorporate Auxiliary Classifier Generative Adversarial Networks (ACGAN) for model training. The CNN component will be offloaded to the Edge Cloud to further enhance intrusion detection performance. To address the task offloading optimization problem in edge computing, we introduce a deep reinforcement learning-based task offloading algorithm to allocate the resources of edge cloud for vehicles with limited computing resources. Simulation experiments demonstrate the superiority of proposed vehicular intrusion detection method over existing state-of-the-art methods. The simulation experiments by Veins also show the efficiency of the proposed vehicular intrusion detection.
Jun Li 0085, Shuqin Zhang, Hongsong Zhu, Jizhao Liu
CSCWD3
2024 PG-AID: An Anomaly-based Intrusion Detection Method Using Provenance Graph
abstract
Intrusion detection is a technique used to identify malicious activities that occur in an organization’s information system, and plays a vital role for security of collaborative systems. Provenance graphs, constructed from system-level audit logs, can capture the complex relations between system entities and associate activities across the entire system, thus provide rich contextual information for intrusion detection. As a result, multiple intrusion detection methods leverage provenance graphs to detect stealthy and persistent malicious activities, known as provenancebased intrusion detection systems (PIDS). However, existing PIDS cannot detect malicious activities with fine granularity without prior knowledge of attack patterns. In this paper, we propose PG-AID, an anomaly-based intrusion detection method using provenance graph. PG-AID first converts system-level audit logs into provenance graph data, which are separated into temporal-ordered snapshots. Then to isolate intrusion-related activities, critaical paths in the snapshot are extracted, which are subsequently aggregated to get the snapshot embedding. By modeling the temporal relationships between normal snapshots, PG-AID detects abnormal graphs that exhibits different temporal relations. Finally, critical paths in the abnormal graphs are presented as intrusion indicators. We use DARPA’s (Defense Advanced Research Projects Agency) Transparent Computing (TC) datasets to evaluate PG-AID’s performance. The results show that PG-AID can effectively detect intrusions and provide detailed information about intrusions with low memory utilization.
Lingxiang Meng, Rongrong Xi, Hongsong Zhu
CSCWD4
2024 Fog-Enabled Intrusion Detection Method Integrating Bi-LSTM and Multi-Head Self-Attention for IoT
abstract
The increasing frequency of cyber attacks targeting IoT highlights the crucial need for accurate and real-time intrusion detection methods. Deep learning, renowned for its remarkable pattern recognition and adaptive learning capabilities, emerges as a promising solution. The current deep learning-based intrusion detection methods face two main issues: dependence on cloud computing architecture, making it challenging to meet the real-time requirements of IoT, and a limited focus on the detection of unknown attacks. In this article, we build our deep learning-based intrusion detection models within fog computing architecture to meet the real-time needs of IoT. Initially, we collect traffic data and conduct feature selection in the fog layer. Subsequently, the Bi-LSTM models integrated with Multi-Head Self-Attention mechanism are trained using the feature-selected data in the cloud. To assess our models’ ability to detect unknown attacks, we utilize some known attacks data for model training, and then assess the models’ performance in detecting other attacks. Ultimately, we deploy the models within fog layer to detect intrusions. Our method is evaluated on the Bot-Iot dataset that contains massive IoT traffic. Experimental results reveal that our models exhibit an average accuracy of 99.85% and perform well in detecting unknown attacks.
Shizhao Tian, Zhen Wang 0043, Kaiyu Xie, Fansong Chen, Hongsong Zhu
CSCWD5
2024 Save the Bruised Striver: A Reliable Live Patching Framework for Protecting Real-World PLCs
abstract
Industrial Control Systems (ICS), particularly programmable logic controllers (PLCs) responsible for managing underlying physical infrastructures, often operate for extended periods without interruption. Thus, it is challenging to patch security vulnerabilities of ICS in a timely manner after disclosure because it often necessitates waiting for a rare downtime window. While live patching has been introduced to avoid downtime and maintenance costs, conventional live patching methods are not viable for closed-source PLCs. Without the source code, it is difficult to understand the system behaviors and determine binary patch equivalence. To address these challenges, we present a Reliable Live Patching framework called RLPatch for applying live patches to third-party binary without source code. We design RLPatch to capture real-time conditions and dynamic behaviors of PLCs, which enables DevOps engineers to identify major non-recoverable fault (MNRF) vulnerabilities and generate hot patches. The core of RLPatch is an update agent that inserts breakpoints over the original MNRF code and then directs execution to the patches. To ensure system reliability, we use the unique constraints of PLCs to integrate the update processes with the scan cycle. We leverage RLPatch to patch 20 real vulnerabilities in three widely used Rockwell PLCs. We evaluate RLPatch in a real-world gas pipeline, demonstrating its reliability and effectiveness in practice.
Ming Zhou 0010, Haining Wang 0001, Ke Li 0042, Hongsong Zhu, Limin Sun 0001
EuroSys4
2024 A Relation-Aware Heterogeneous Graph Transformer on Dynamic Fusion for Multimodal Classification Tasks
abstract
Multimodal fusion aims to improve the performance of models for applications by extracting and fusing information in different modalities, including texts, images or others. Recent researches have shown that multimodal fusion is beneficial in many multimedia tasks. In this paper, we study typical multimedia classification tasks in social media posts, including sarcasm detection and sentiment analysis. This paper proposes DMF-RHGT-HPA, including dynamic Fusion multimodal fusion(DMF), a relation-aware heterogeneous graph transformer(RHGT) and hierarchical pooling alignment(HPA). To realize better multimodal fusion, the paper designs it on a heterogeneous graph with dynamic links, without any padding of texts or images. To thoroughly learn the multimodal graph and obtain the representation of nodes, the paper proposes a relation-aware heterogeneous graph transformer to fuse the node-level and edge-level features simultaneously. To get a refined representation of the multimodal graph, the paper designs a hierarchical pooling alignment to gather all nodes’ representations well. Experiments conducted on two primary and public datasets from Twitter and Yelp respectively show the ability of DMF-RHGT-HPA to gain the best performance of sarcasm detection and sentiment analysis, outperforming existing state-of-the-art baselines.
Yimo Ren, Jinfa Wang, Jie Liu 0079, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
ICASSP6
2024 Enhancing Coverage in Stateful Protocol Fuzzing via Value-Based Selection
abstract
The stateful nature inherent in network protocol implementations presents distinctive challenges for testing and verification methods, including Fuzzing. However, not all states hold equal significance. Indiscriminate state selection for fuzzing could lead to intricate path mazes. Similar challenges emerge in the selection for seeds and mutation operators. Therefore, overcoming the efficiency constraints of current fuzzers crucially depends on making precise selections in fuzzing. In this study, we present AcSelector, a novel approach that incorporates the Composite State Model and Mutation Operator Value Table. By offering the most strategic combination of {state, seed, mutation operator}, AcSelector provides systematic guidance for fuzzing, leading to enhanced code coverage. To quantify value of targets, AcSelector employs a principled evaluation strategy. We evaluated AcSelector by fuzzing six network servers from popular open-source projects. Our experimental results demonstrate the effectiveness of AcSelector in increasing code coverage, even under low fuzzing throughput conditions.
Laile Xi, Shenghao Lin, Yuyan Sun, Hongsong Zhu, Limin Sun 0001
ISCC6
2024 Lexicon Graph Adapter Based BERT Model for Chinese Named Entity Recognition
Jie Liu 0079, Yimo Ren, Jinfa Wang, Hongsong Zhu
KSEM (5)5
2024 Symerge: Replacing Calls in Under-Constrained Symbolic Execution and Find Vulnerabilities
Yicheng Zeng, Jiaqian Peng, Jiami Lin, Rongrong Xi, Hongsong Zhu
SecureComm (2)5
2024 EasyDetector: Using Linear Probe to Detect the Provenance of Large Language Models
abstract
The rapid development of large language models (LLMs) has driven significant advancements in various applications. However, the intellectual property of these models often faces risks due to unauthorized reproduction or encapsulation by third parties. In this paper, we propose EasyDetector, a novel approach to detect the provenance of LLMs using linear probes. Our method aims to identify the original source model, even if it has been fine-tuned or encapsulated into another model. Specifically, EasyDetector performs classification on the intermediate layer representations of the new model using linear probes of the original model. Models from the same source exhibit high accuracy, while models from different sources yield low accuracy. Extensive experiments on diverse LLMs demonstrate the effectiveness of EasyDetector in detecting model provenance. The proposed method is lightweight and applicable to various model architectures, holding significant importance for protecting the intellectual property of LLMs.
Jie Zhang 0121, Jiayuan Li 0002, Haiqiang Fei, Hongsong Zhu
TrustCom5
2024 UniTTP: A Unified Framework for Tactics, Techniques, and Procedures Mapping in Cyber Threats
abstract
The increasing complexity of cyber threats necessitates advanced methods for understanding and countering adversarial tactics, techniques, and procedures (TTPs). Despite the support provided by the ATT&CK framework, challenges such as imbalanced sample distribution and technique overlap limit the effective mapping of complex attack patterns. To address these challenges, we propose a framework that integrates language models and advanced artificial intelligence techniques, including hierarchical attention embedding and contrastive learning, to achieve TTP recognition and classification in complex cyber threats. UniTTP consists of five modules: data processing, tactic classification, multi-feature embedding, techniques classification, and LLM post-assistance. By combining these modules, our method not only accurately identifies TTPs within cyber threats, improving the F1 score by 3.23% to 13.66% across different datasets, but also leverages the capabilities of large language models to verify recognition results and deepen the understanding of attack behaviors. This study lays the foundation for robust cyber defense by providing deeper insights into adversary behaviors and enhancing the predictability of complex threats.
Jie Zhang 0121, Hui Wen 0001, Hongsong Zhu
TrustCom4
2024 TM-fuzzer: fuzzing autonomous driving systems through traffic management
Shenghao Lin, Fansong Chen, Laile Xi, Gaosheng Wang, Rongrong Xi, Yuyan Sun, Hongsong Zhu
Autom. Softw. Eng.7
2024 VDTriplet: Vulnerability detection with graph semantics using triplet model
Hao Sun 0028, Lei Cui 0003, Zhenquan Ding, Siyuan Li 0014, Zhiyu Hao, Hongsong Zhu
Comput. Secur.7
2024 KnowCTI: Knowledge-based cyber threat intelligence entity and relation extraction
Gaosheng Wang, Haoyu Bin, Hongsong Zhu
Comput. Secur.6
2024 Cascading Threat Analysis of IoT Devices in Trigger-Action Platforms
abstract
Internet of Things (IoT) platforms have become widely used recently. Facilitated by these IoT platforms, users can easily use programming paradigm to develop customized rules, connect their devices with online services, and realize system automation. However, the attack surface of each device is expanded as the device interactions increase with multiple rules enabled. In this work, we present a framework to analyze the cascading threat based on device interactions in the IFTTT (IF This Then That) platform. We first extract the trigger-action rules from the description text by using an NLP-based method. Then, we create a graph-based model by combining trigger-action rules with three components, to describe the flow information of device interactions. Finally, we propose a graph-searching-based method to discover the paths and starting points of application-level cascading attacks, uncovering the attack surface of devices. We conduct the evaluation on a data set of 305534 applets from the IFTTT platform. The results evidence that cascading attacks exist in IoT deployments but can be captured by our attack surface analysis.
Ke Li 0042, Haining Wang 0001, Ming Zhou 0010, Hongsong Zhu, Limin Sun 0001
IEEE Internet Things J.4
2024 RIETD: A Reputation Incentive Scheme Facilitates Personalized Edge Tampering Detection
abstract
Edge nodes provide service cooperatively in edge computing, where third-party nodes are common. However, they cannot be fully trusted and can intentionally alter service results (e.g., edge tampering). Although some mechanisms can help detect edge tampering, they come with additional detection overhead. It is essential to note that most edge nodes are willing to serve honestly. Therefore, it is reasonable to decrease the detection frequency for those nodes, which helps reduce the overhead. Reputation is the common way to evaluate trustworthiness. In this article, we propose a reputation incentive scheme called RIETD, which evaluates the reputations of edge nodes using their detection results. Moreover, RIETD is loosely coupled with detection mechanisms as an external service. Reputation is the fundamental parameter in RIETD, as it determines an acrlong EN’s appraisal weight on other nodes, personalized detection strategy, and node’s obtained revenues in one service. We demonstrate that RIETD does not significantly reduce the overall detection capability while the overhead is reduced effectively. For instance, when the tampering rate of edge nodes is 10%, and the target detection rate is 90% of a specific detection mechanism, RIETD reduces overhead by approximately 60%. If a full-reputation node is detected to have tampered with the results, its reputation recovery time is similar to the time required for reputation to improve from 0 to 1. Moreover, a node’s expected revenue is lower than that of an honest node, emphasizing the importance of serving honestly and continuously for edge nodes to earn higher revenue.
Fei Lyu 0001, Shuaizong Si, Hongsong Zhu, Limin Sun 0001
IEEE Internet Things J.5
2024 Multi-granularity cross-modal representation learning for named entity recognition on social media
Gaosheng Wang, Hong Li 0004, Jie Liu 0079, Yimo Ren, Hongsong Zhu, Limin Sun 0001
Inf. Process. Manag.6
2024 Few-Shot Malware Classification via Attention-Based Transductive Learning Network
Liting Deng, Chengli Yu, Hui Wen 0001, Mingfeng Xin, Limin Sun 0001, Hongsong Zhu
Mob. Networks Appl.7
2024 FeaShare: Feature Sharing for Computation Correctness in Edge Preprocessing
abstract
Edge preprocessing is a critical service type in edge computing. However, untrusted edges may be malicious to provide incorrect computational results (i.e., edge tampering). Although some studies have considered the correctness of results, they have limitations when applied to edge preprocessing. We present FeaShare, a feature-sharing approach, to verify edge results. The process is integrated into normal service operations. Meanwhile, to overcome feature-based limitations, terminals obtain partial edge results for a set of data by executing a small number of computations. These partial results are leveraged to construct shared features, facilitating the detection of edge tampering even when the tampered portion is not directly related to the features. Subsequently, the shared features are mapped to pseudo-data and added to the terminal's data sequence, preventing features from influencing the results of terminal data. To resist edge attacks, both feature construction and placement are time-dependent and dynamic. FeaShare is not confined to specific edge tasks. We evaluate FeaShare using 3 typical scenes encompassing 5 applications. For instance, the evaluation utilizing the VGG model and CIFAR-10 dataset demonstrates a detection rate of 97%. Terminals perform approximately 10% of the edge's computation operations, and its overhead growth rate is less than 10%.
Haoyu Bin, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
IEEE Trans. Mob. Comput.5
2024 Crafting Binary Protocol Reversing via Deep Learning With Knowledge-Driven Augmentation
abstract
Protocol reverse engineering (PRE) serves as an instrumental tool in various security research, such as protocol fuzzing and intrusion detection. Its primary objective lies in uncovering the format, semantics, and behavior of an unknown protocol without prior information. This paper presents DL-ProS2, a deep learning-based approach for binary protocol reversing, focusing on format segmentation and semantic inference from network traffic. Our approach is underpinned by highlighting the effectiveness of multi-scale features within the network traffic for identifying various types of fields and semantics. Based on this, DL-ProS2 employs a comprehensive end-to-end model that integrates U-Net, siamese network, and BiLSTM-CRF, which enables the effective analysis of unknown protocol traffic to extract the field boundaries and semantics. Meanwhile, to address the issue of limited data diversity and coverage, we implement an innovative knowledge-driven traffic simulation technique. This method harnesses the ChatGPT to extract protocol knowledge from publicly available protocol documents, such as RFCs, as the foundational rules for the simulation. Empirical results substantiate the efficacy of our approach, demonstrating precision rates exceeding 0.95 and recall rates surpassing 0.97 for partially unknown protocol format segmentation and semantic inference. It also retains effectiveness in the inference of completely unknown protocols, with average precision and recall rates of 0.69 and 0.62 for format segmentation, and 0.43 and 0.47 for semantic inference, respectively.
Shouguo Yang, Zhen Wang 0043, Yongji Liu, Hongsong Zhu, Limin Sun 0001
IEEE/ACM Trans. Netw.5
2024 LibAM: An Area Matching Framework for Detecting Third-Party Libraries in Binaries
abstract
Third-party libraries (TPLs) are extensively utilized by developers to expedite the software development process and incorporate external functionalities. Nevertheless, insecure TPL reuse can lead to significant security risks. Existing methods, which involve extracting strings or conducting function matching, are employed to determine the presence of TPL code in the target binary. However, these methods often yield unsatisfactory results due to the recurrence of strings and the presence of numerous similar non-homologous functions. Furthermore, the variation in C/C++ binaries across different optimization options and architectures exacerbates the problem. Additionally, existing approaches struggle to identify specific pieces of reused code in the target binary, complicating the detection of complex reuse relationships and impeding downstream tasks. And, we call this issue the poor interpretability of TPL detection results. In this article, we observe that TPL reuse typically involves not just isolated functions but also areas encompassing several adjacent functions on the Function Call Graph (FCG). We introduce LibAM, a novel Area Matching framework that connects isolated functions into function areas on FCG and detects TPLs by comparing the similarity of these function areas, significantly mitigating the impact of different optimization options and architectures. Furthermore, LibAM is the first approach capable of detecting the exact reuse areas on FCG and offering substantial benefits for downstream tasks. To validate our approach, we compile the first TPL detection dataset for C/C++ binaries across various optimization options and architectures. Experimental results demonstrate that LibAM outperforms all existing TPL detection methods and provides interpretable evidence for TPL detection results by identifying exact reuse areas. We also evaluate LibAM’s scalability on large-scale, real-world binaries in IoT firmware and generate a list of potential vulnerabilities for these devices. Our experiments indicate that the Area Matching framework performs exceptionally well in the TPL detection task and holds promise for other binary similarity analysis tasks. Last but not least, by analyzing the detection results of IoT firmware, we make several interesting findings, for instance, different target binaries always tend to reuse the same code area of TPL. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/LibAM .
Siyuan Li 0014, Yongpan Wang, Chaopeng Dong, Shouguo Yang, Hong Li 0004, Hao Sun 0028, Zhe Lang, Zuxin Chen, Weijie Wang 0005, Hongsong Zhu, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.10
2023 A Privacy-Preserving Online Deep Learning Algorithm Based on Differential Privacy
abstract
Deep Reinforcement Learning (DRL) combines the perceptual capabilities of deep learning with the decision-making capabilities of Reinforcement Learning RL, which can achieve enhanced decision-making. However, the environmental state data contains the privacy of the users. There exists consequently a potential risk of environmental state information being leaked during RL training. Some data desensitization and anonymization technologies are currently being used to protect data privacy. There may still be a risk of privacy disclosure with these desensitization techniques. Meanwhile, policymakers need the environmental state to make decisions, which will cause the disclosure of raw environmental data. To address the privacy issues in DRL, we propose a differential privacy-based online DRL algorithm. The algorithm will add Gaussian noise to the gradients of the deep network according to the privacy budget. More important, we prove tighter bounds for the privacy budget. Furthermore, we train an autocoder to protect the raw environmental state data. In this work, we prove the privacy budget formulation for differential privacy-based online deep RL. Experiments show that the proposed algorithm can improve privacy protection while still having relatively excellent decisionmaking performance.
Jun Li 0085, Fengshi Zhang, Yonghe Guo, Siyuan Li 0014, Guanjun Wu, Dahui Li, Hongsong Zhu
CSCWD7
2023 SynCPFL: Synthetic Distribution Aware Clustered Framework for Personalized Federated Learning
abstract
Federated Learning (FL) is a promising machine learning paradigm for collaborative training on cross-soils in a privacy-protected manner. However, the existence of non-IID data causes problems such as performance degradation and thus becomes one of the key challenges in FL recently. To address this problem, we propose a clustered personalized federated learning method named as SynCPFL. SynCPFL groups clients sharing with the similar data distribution together, thereby facilitating collaboration and producing a better-personalized model for each client. In contrast to existing clustered federated learning methods, SynCPFL does not require multiple rounds of interaction between clients and server, so that the communication overhead is reduced a lot, thereby saving resources of clients. We evaluate SynCPFL on benchmark datasets, the experimental results demonstrate that SynCPFL outperforms existing methods.
Junnan Yin, Yuyan Sun, Lei Cui 0003, Zhengyang Ai, Hongsong Zhu
CSCWD5
2023 MalAder: Decision-Based Black-Box Attack Against API Sequence Based Malware Detectors
abstract
The API call sequence based malware detectors have proven to be promising, especially when incorporated with deep neural networks (DNNs). Several adversarial attack methods are proposed to fool these detectors by introducing undetectable perturbations into normal samples. However, in real-world scenarios, the malware detector provides only the predicted label for a given sample, without exposing its network architecture or output probability, making it challenging for adversarial attacks under the decision-based black-box. Existing work in this area typically relies on random-based methods that suffer high costs and low attack success rates. To address these limitations, we propose a novel decision-based black-box attack against API sequence based malware detectors, called MalAder. Our approach aims to improve the attack success rate as well as query efficiency through a directional perturbation algorithm. First, it utilizes attention-based API ranking to assess the importance of API calls in the context of different API sequences. This assessment guides the insertion position for perturbation. Then, the perturbation is carried out using benign distance perturbing, which gradually shortens the semantic distance from adversarial API sequences to a set of benign samples. Finally, our algorithm iteratively generates adversarial malware samples by performing perturbations. In addition, we have implemented MalAder and evaluated its performance against two classic malware detectors. The results show that MalAder outperforms state-of-the-art decision-based black-box adversarial attacks, proving its effectiveness.
Lei Cui 0003, Hui Wen 0001, Zhi Li 0018, Hongsong Zhu, Zhiyu Hao, Limin Sun 0001
DSN5
2023 CSEDesc: CyberSecurity Event Detection with Event Description
Gaosheng Wang, Shuaizong Si, Hongsong Zhu, Limin Sun 0001
ICANN (3)5
2023 Improving the Modality Representation with multi-view Contrastive Learning for Multimodal Sentiment Analysis
abstract
Modality representation learning is an important problem for multimodal sentiment analysis (MSA), since the highly distinguishable representations can contribute to improving the analysis effect. Previous works of MSA have usually focused on internal fusion strategies for different modalities within one sample, and the external usage of cross reference relations among different samples was given less attention. Recently, the rise of contrastive learning provides powerful clues for us to learn modal representation with stronger discriminative ability. In this study, we explore the approach of representations improvement and devise a three-stages framework with multi-view contrastive learning to refine representations for the specific objectives. Firstly, for each modality, we employ the supervised contrastive learning to pull samples within the same class together while the other samples are pushed apart. Then, a self-supervised contrastive learning is designed for the distilled cross-modal representations after a novel Transformer-based interaction module. At last, we leverage again the supervised contrastive learning to enhance the fused multimodal representation. We conduct extensive experiments on three open datasets, and results show the advance of our model.
Hong Li 0004, Jie Liu 0079, Yimo Ren, Hongsong Zhu, Limin Sun 0001
ICASSP6
2023 CEntRE: A paragraph-level Chinese dataset for Relation Extraction among Enterprises
abstract
Enterprise relation extraction aims to detect pairs of enterprise entities and identify the business relations between them from unstructured or semi-structured text data, and it is crucial for several real-world applications such as risk analysis, rating research and supply chain security. However, previous work mainly focuses on getting attribute information about enterprises like personnel and corporate business, and pays little attention to enterprise relation extraction. To encourage further progress in the research, we introduce the CEntRE, a new dataset constructed from publicly available business news data with careful human annotation and intelligent data processing. Moreover, we propose a joint entity and relation extraction network, which is capable of discovering enterprise entities and extracting business relations between them accurately. The network firstly encodes input sequences with strong semantic augmentation to learn contextual representation for each token, then a conditional random field (CRF) module is used for entity extraction. Subsequently, entity pairs are built and a new encoder based on the entity pairs is applied to get global information for relation extraction. Finally, a biaffine classifier is deployed to classify the relations. Extensive experiments on CEntRE demonstrate the effectiveness of our proposed method compared with other six excellent models, and thus our model can be considered as one strong baseline. The data and code are available at: https://github.com/LiuPeiP-CStMining_Entity_Relations_Among_Enterprises
Hong Li 0004, Yimo Ren, Jie Liu 0079, Fei Lyu 0001, Hongsong Zhu, Limin Sun 0001
IJCNN7
2023 User Recognition of Devices on the Internet based on Heterogeneous Graph Transformer with Partial Labels
abstract
Recognizing the users of devices can easily enable numerous security applications. Due to the lot's kinds of device data and a large number of missing values, it takes work to recognize the users of devices well. The community detection methods based on Graph Neural Networks (GNN) can integrate multi-source data well and cluster devices into communities with the same users. While existing GNN methods face several issues. The methods on homogeneous graphs could not utilize the multi-source data of devices, and most methods on heterogeneous graphs need specific knowledge to design meta paths. Also, the Internet-scale data of devices make it hard to learn the representation thoroughly. Further, most methods need to consider the known partial labels in the early stage of the training process. To improve the performance of user recognition, this paper proposes HGT-PL, namely a Heterogeneous Graph Transformer with Partial Labels, to calculate the representation of devices on the Internet. Then cluster methods are used to realize user recognition. Using graph transformers, HGT-PL deeply learns node features and graph structure on the heterogeneous graph of devices. By Label Encoder, HGT-PL fully utilizes the users of partial devices from preliminary rules with high confidence. Moreover, cluster methods carefully divide and modify the communities with different users. The paper conducts experiments on the web-scale data collected from the Internet. The results show that HGT-PL can recognize users of devices more accurately and effectively, with 0.5121 NMI and 0.3554 ARI, compared with existing GNN methods.
Yimo Ren, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
IJCNN4
2023 UID-Auto-Gen: Extracting Device Fingerprinting from Network Traffic
abstract
The number of Internet device vulnerabilities has been quickly rising in recent years, rendering an explosion of network attacks. Device fingerprinting serves as the primary means for vulnerability awareness and attacker tracking. The current device fingerprinting approach can only achieve model-level identification within the Internet scope or individual-level identification for specific protocols (e.g., SSL) or scenarios (e.g., LAN). However, it is still difficult for these methods to achieve individual-level identification on a global Internet scale. In this paper, we propose a fingerprint extraction approach that is accurate to the individual level of the device by using a combination of clustering, multiple sequence alignment, and based on the geographic location stability of the device. In a continuous 3-month observation for several cities around the world, at least 1.54% of devices can be accurately extracted with unique IDs, with an accuracy rate of 99.30%, which is capable of being used in production environments.
Haoyu Bin, Zhi Li 0018, Rongrong Xi, Hongsong Zhu, Limin Sun 0001
IPCCC6
2023 An Enhanced Vulnerability Detection in Software Using a Heterogeneous Encoding Ensemble
abstract
Detecting vulnerabilities in source code is essential to prevent cybersecurity attacks. Deep learning-based vulnerability detection is an active research topic in software security. However, existing deep learning-based vulnerability detectors (VD) are limited to using either serialization-based or graph-based methods, which do not combine serialized global and structured local information at the same time. As a result, a single method cannot perform well for semantic information that exists in complex source code, leading to low detection accuracy. In this paper, we present EL-VDetect, a stacked ensemble learning approach for vulnerability detection that eliminates these issues. EL-VDetect enhances feature selection techniques to represent the best relevant vulnerability features with the slice code and subgraphs, reducing redundant information of vulnerabilities. Our model combines serialization-based and graph-based neural networks to successfully capture the global and local context information of source code, effectively understands code semantics, and focuses on vulnerable nodes based on the attention mechanism to accurately detect vulnerabilities. To evaluate EL-VDetect's effectiveness, we crawl a real-world dataset from CVEDetails, consisting of functions for eight applications. A comprehensive performance analysis of the real-world dataset shows that EL-VDetect achieves 90.72% accuracy, outperforming baseline deep learning models by 1.75-26.21 %. Our proposed model can better identify vulnerabilities in software than other existing vulnerability detection models.
Hao Sun 0028, Yongji Liu, Zhenquan Ding, Yang Xiao 0011, Zhiyu Hao, Hongsong Zhu
ISCC6
2023 Detecting Vulnerabilities in Linux-Based Embedded Firmware with SSE-Based On-Demand Alias Analysis
abstract
Although the importance of using static taint analysis to detect taint-style vulnerabilities in Linux-based embedded firmware is widely recognized, existing approaches are plagued by following major limitations: (a) Existing works cannot properly handle indirect call on the path from attacker-controlled sources to security-sensitive sinks, resulting in lots of false negatives. (b) They employ heuristics to identify mediate taint source and it is not accurate enough, which leads to high false positives.
Yaowen Zheng, Le Guan, Peng Liu 0005, Hong Li 0004, Hongsong Zhu, Kejiang Ye, Limin Sun 0001
ISSTA7
2023 Software Vulnerability Detection Using an Enhanced Generalization Strategy
Hao Sun 0028, Zhe Bu, Yang Xiao 0011, Chengsheng Zhou, Zhiyu Hao, Hongsong Zhu
SETTA6
2023 DeviceGPT: A Generative Pre-Training Transformer on the Heterogenous Graph for Internet of Things
abstract
Recently, Graph neural networks (GNNs) have been adopted to model a wide range of structured data from academic and industry fields. With the rapid development of Internet technology, there are more and more meaningful applications for Internet devices, including device identification, geolocation and others, whose performance needs improvement. To replicate the several claimed successes of GNNs, this paper proposes DeviceGPT based on a generative pre-training transformer on a heterogeneous graph via self-supervised learning to learn interactions-rich information of devices from its large-scale databases well. The experiments on the dataset constructed from the real world show DeviceGPT could achieve competitive results in multiple Internet applications.
Yimo Ren, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
SIGIR4
2023 HackMentor: Fine-Tuning Large Language Models for Cybersecurity
abstract
The democratization of artificial intelligence has made substantial progress by leveraging open-source large language models (LLMs), enabling researchers across domains to train customized models to meet their specific needs. Given the confidentiality and significance of cybersecurity, obtaining private and localized LLMs is imperative. However, general LLMs are not designed to cater specifically to this field, their general knowledge often falls short when addressing such specialized problems. In this paper, we categorize the domain instructions based on cybersecurity knowledge to guide the construction of high-quality instructions and conversations, ultimately enhancing the specialized capabilities of LLMs. The resulting fine-tuned LLMs, collectively termed HackMentor, are evaluated using WinRate, EloRating, and ZenoEval methods along with other popular LLMs. The experiments demonstrate that the proposed method yields significant performance improvements, surpassing the native LLMs by 10-25% when aligned with cybersecurity prompts. More, HackMentor exhibits comparable conversational quality to ChatGPT, while providing more concise and humanlike responses. This study demonstrates the efficacy of HackMentor in augmenting LLMs for cybersecurity requirements, paving the way for localized LLMs that meet specialized needs without compromising general capabilities.
Jie Zhang 0121, Hui Wen 0001, Liting Deng, Mingfeng Xin, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001
TrustCom7
2023 CL-GAN: A GAN-based continual learning model for generating and detecting AGDs
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001
Comput. Secur.5
2023 Owner name entity recognition in websites based on multiscale features and multimodal co-attention
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001
Expert Syst. Appl.5
2023 Multiview Embedding with Partial Labels to Recognize Users of Devices Based on Unified Transformer
abstract
Recognizing the users of devices (or clusters of devices) who use IP addresses as unique identities on the Internet can easily enable numerous security applications. Fast and accurate user recognition is critical for supervisors to find influenced organizations connected to their networks in light of new security threats. Many users’ information scatters in the multisource data of IP addresses. Up until now, user recognition of devices has had two main problems. On the one hand, existing methods could not fully use multisource data of the IP addresses and wastes the valuable information of labels. On the other hand, only a tiny portion of devices can be tagged with highly confident known users manually, making it an urgent need to infer unknown users of devices. So, the problem of user recognition on devices is to guess the unknown user with multisource data and existing devices with known users. Therefore, this paper proposes a multiview fusion method to deal with multisource data from devices with a small number of manually labelled samples. The paper uses GraphSAGE to obtain an exemplary representation of IP addresses and designs a label encoder to fully use a small number of devices with known users. Then, the paper builds a specific unified transformer to achieve high performance to determine whether two devices have the same user. At the same time, the paper conducts real‐world experiments and finds that the proposed method can achieve 0.9158 accuracy and 0.6131 F1 to find devices with the same users on the constructed dataset in the real world.
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001
Int. J. Intell. Syst.5
2023 Owner name entity recognition in websites based on heterogeneous and dynamic graph transformer
Yimo Ren, Hong Li 0004, Jie Liu 0079, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001
Knowl. Inf. Syst.6
2023 Internet-Scale Fingerprinting the Reusing and Rebranding IoT Devices in the Cyberspace
abstract
Fingerprinting Internet-of-Things(IoT) devices on types and brands is a necessary work for security analysis in the cyberspace. The existing approaches mainly rely on the dominant features of devices which is response to information in order to identify these online devices. However, the web server components reusing and products rebranding are the common phenomenons of these embedded IoT devices. It caused the existing approaches difficult to identify most devices even errors due to the similar responses. In this paper, we present an approach, IoTXray, which improves the work efficiently of information collection about accelerating the relations between reusing/rebranding devices with the corresponding manufacturers. And these relations can generate more accurate and reliable fingerprints than previous approaches. Using the mixed neural networks, IoTXray comprehensively detects the real manufactures of online IoT devices upon three different kinds of data sources. In the experiment, our approach can identify 7,025,854 IoT devices on HTTP-hosts. The identification rate has reached to several times higher than previous approaches. Our approach has especially detected 3,268,953 reusing and 963,653 rebranding devices with their original manufacturers.
Zhaoteng Yan, Zhi Li 0018, Hong Li 0004, Shouguo Yang, Hongsong Zhu, Limin Sun 0001
IEEE Trans. Dependable Secur. Comput.5
2022 Enhancing Chinese Pre-trained Language Model via Heterogeneous Linguistics Graph
abstract
Yanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang, Bowen Yu, Hongsong Zhu, Tingwen Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang 0006, Bowen Yu 0002, Hongsong Zhu, Tingwen Liu
ACL (1)6
2022 Finding Vulnerabilities in Internal-binary of Firmware with Clues
abstract
Embedded devices, represented by Internet of Things devices, bring great convenience to our daily life. Firmware is the core of the embedded device operation. However, vulnerabilities in the firmware can be exploited remotely by hackers through the network. Unfortunately, existing methods are only suitable for finding vulnerabilities in binaries (border-binary) that interact directly with users. When applied to other binaries (internal-binary) that indirectly interact with users, the lack of analysis sources and constraint conditions leads to many false negatives and false positives. In this paper, we propose a new keyword-sensitive data flow analysis approach to address the challenge. Specifically, we leverage crawlers to collect clues related to vulnerability reports from the Internet. Then we use the clues and communication paradigm finders to establish the relationship between different binaries in the firmware sample to form binary dependency graphs. At the same time, based on the functional features, we further dig out the binary relationships that have no Internet clues. Finally, we perform static taint analysis based on binary dependency graphs to determine vulnerabilities. We implemented and evaluated our prototype system FBI. Compared with Karonte, a state-of-the-art tool, FBI found significantly more true positives in Karonte’s data set.
Puzhuo Liu, Dongliang Fang, Shichao Lv, Hongsong Zhu, Limin Sun 0001
ICC6
2022 ProsegDL: Binary Protocol Format Extraction by Deep Learning-based Field Boundary Identification
abstract
Protocol reverse engineering can be applied to various security applications, including fuzzing, malware analysis, and intrusion detection. It aims to acquire an unknown protocol's format, semantic, and behavior specifications, where format extraction is the primary task. One subset of the mainstream research utilizes the network traffic for the reverse analysis. These approaches leverage various algorithms, such as multiple sequence alignment, frequent itemset mining, and information entropy to extract format information from messages. However, they are primarily intended to locate the keyword fields and have limitations in extracting contextual features or dealing with large data sets. This paper presents ProsegDL, a deep learning-based format extraction tool for binary protocol, with a specially designed method of generating training data sets. ProsegDL innovatively leverages image semantic segmentation and siamese network techniques, focusing on extracting the features of fields and identifying field boundaries for fixed format protocols. The tool is evaluated on six popular protocols. The results show that it has at most 13% higher precision, 23% higher recall than the comparison methods when inferring with a small data set, and at most 18% higher precision, 28% higher recall when inferring with a large number of messages.
Jinfa Wang, Shouguo Yang, Yicheng Zeng, Hongsong Zhu, Limin Sun 0001
ICNP6
2022 Multi-features based Semantic Augmentation Networks for Named Entity Recognition in Threat Intelligence
abstract
Extracting cybersecurity entities such as attackers and vulnerabilities from unstructured network texts is an important part of security analysis. However, the sparsity of intelligence data resulted from the higher frequency variations and the randomness of cybersecurity entity names makes it difficult for current methods to perform well in extracting security-related concepts and entities. To this end, we propose a semantic augmentation method which incorporates different linguistic features to enrich the representation of input tokens to detect and classify the cybersecurity names over unstructured text. In particular, we encode and aggregate the constituent feature, morphological feature and part of speech feature for each input token to improve the robustness of the method. More than that, a token gets augmented semantic information from its most similar K words in cybersecurity domain corpus where an attentive module is leveraged to weigh differences of the words, and from contextual clues based on a large-scale general field corpus. We have conducted experiments on the cybersecurity datasets DNRTI and MalwareTextDB, and the results demonstrate the effectiveness of the proposed method.
Hong Li 0004, Zuoguang Wang, Jie Liu 0079, Yimo Ren, Hongsong Zhu
ICPR6
2022 SIFOL: Solving Implicit Flows in Loops for Concolic Execution
abstract
Concolic execution is widely used for binary analysis and is commonly embedded in hybrid fuzzing to find bugs. However, implicit flows in loops can hinder concolic execution and lead to the reduction of code coverage. The implicit flow variables cannot be symbolized and will block the constraint solver from generating new inputs. We propose a new approach to mitigate the problem. We obtain the implicit flow variables by taint analysis in advance and symbolize them during the concolic execution. Then, when the symbols of the variables are in the path constraints and need to be solved, we backtrack to the corresponding loops and perform static symbolic executions in the loops. During the static symbolic executions, we relate the variables with the input symbols by state merging and solve the constraints to generate inputs for new execution paths. We present SIFOL, a hybrid fuzzer based on Driller, and evaluate it on CB-multios. Results show that SIFOL has 5.4% higher code coverage than Driller and finds 5.9% more crashes. Furthermore, after manually adding implicit flows and checks to the target programs, SIFOL only drops 2.6% on coverage and 5.6% on the crash number, while Driller is severely affected (drops 46.1% on coverage and 47.1% on the crash number).
Yicheng Zeng, Jiaqian Peng, Zhanwei Song, Hongsong Zhu, Limin Sun 0001
IPCCC5
2022 Efficient greybox fuzzing of applications in Linux-based IoT devices via enhanced user-mode emulation
abstract
Greybox fuzzing has become one of the most effective vulnerability discovery techniques. However, greybox fuzzing techniques cannot be directly applied to applications in IoT devices. The main reason is that executing these applications highly relies on specific system environments and hardware. To execute the applications in Linux-based IoT devices, most existing fuzzing techniques use full-system emulation for the purpose of maximizing compatibility. However, compared with user-mode emulation, full-system emulation suffersfrom great overhead. Therefore, some previous works, such as Firm-AFL, propose to combine full-system emulation and user-mode emulation to speed up the fuzzing process. Despite the attempts of trying to shift the application towards user-mode emulation, no existing technique supports to execute these applications fully in the user-mode emulation. To address this issue, we propose EQUAFL, which can automatically set up the execution environment to execute embedded applications under user-mode emulation. EQUAFL first executes the application under full-system emulation and observe for the key points where the program may get stuck or even crash during user-mode emulation. With the observed information, EQUAFL can migrate the needed environment for user-mode emulation. Then, EQUAFL uses an enhanced user-mode emulation to replay system calls of network, and resource management behaviors to fulfill the needs of the embedded application during its execution. We evaluate EQUAFL on 70 network applications from different series of IoT devices. The result shows EQUAFL outperforms the state-of-the-arts in fuzzing efficiency (on average, 26 times faster than AFL-QEMU with full-system emulation, 14 times than Firm-AFL). We have also discovered ten vulnerabilities including six CVEs from the tested firmware images.
Yaowen Zheng, Yuekang Li, Cen Zhang, Hongsong Zhu, Yang Liu 0003, Limin Sun 0001
ISSTA4
2022 An Evolutionary Learning Approach Towards the Open Challenge of IoT Device Identification
Jingfei Bian, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
SecureComm4
2022 Detection and Incentive: A Tampering Detection Mechanism for Object Detection in Edge Computing
abstract
The object detection tasks based on edge computing have received great attention. A common concern hasn't been addressed is that edge may be unreliable and uploads the incorrect data to cloud. Existing works focus on the consistency of the transmitted data by edge. However, in cases when the inputs and the outputs are inherently different, the authenticity of data processing has not been addressed. In this paper, we first simply model the tampering detection. Then, bases on the feature insertion and game theory, the tampering detection and economic incentives mechanism (TDEI) is proposed. In tampering detection, terminal negotiates a set of features with cloud and inserts them into the raw data, after the cloud determines whether the results from edge contain the relevant information. The honesty incentives employs game theory to instill the distrust among different edges, preventing them from colluding and thwarting the tampering detection. Meanwhile, the subjectivity of nodes is also considered. TDEI distributes the tampering detection to all edges and realizes the self-detection of edge results. Experimental results based on the KITTI dataset, show that the accuracy of detection is 95% and 80%, when terminal's additional overhead is smaller than 30% for image and 20% for video, respectively. The interference ratios of TDEI to raw data are about 16% for video and 0% for image, respectively. Finally, we discuss the advantage and scalability of TDEI.
Yicheng Zeng, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
SRDS5
2022 Discover the ICS Landmarks Based on Multi-stage Clue Mining
Jie Liu 0079, Jinfa Wang, Hongsong Zhu, Limin Sun 0001
WASA (3)4
2022 Joint Classification of IoT Devices and Relations in the Internet with Network Traffic
abstract
With the rapid growth and popularization of Internet of Things (IoT), more and more devices are deployed in homes, enterprises, cities, etc. The existed methods to classify types and relations of devices are usually two separate tasks. So, it is difficult to quickly provide attributes of devices in the smart network for operators at the same time. At this situation, the paper presents a framework JCIDR for Joint Classification of IoT Device and Relations In the Internet with Network Traffic. By fusing the numerical features and binary image features of traffic, the devices and relations of devices can be recognized simultaneously. The experiment is carried out in a real IoT environment and the accuracy of JCIDR is over 86% with about half time reduction. Therefore, JCIDR could provide operators with a fast, easy, low-cost network device monitoring method without professional equipment or protocols.
Yimo Ren, Hong Li 0004, Shuqin Zhang, Hongsong Zhu, Limin Sun 0001
WCNC6
2021 Discontinuous Named Entity Recognition as Maximal Clique Discovery
abstract
Yucheng Wang, Bowen Yu, Hongsong Zhu, Tingwen Liu, Nan Yu, Limin Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bowen Yu 0002, Hongsong Zhu, Tingwen Liu, Limin Sun 0001
ACL/IJCNLP (1)3
2021 ICS3Fuzzer: A Framework for Discovering Protocol Implementation Bugs in ICS Supervisory Software by Fuzzing
abstract
The supervisory software is widely used in industrial control systems (ICSs) to manage field devices such as PLC controllers. Once compromised, it could be misused to control or manipulate these physical devices maliciously, endangering manufacturing process or even human lives. Therefore, extensive security testing of supervisory software is crucial for the safe operation of ICS. However, fuzzing ICS supervisory software is challenging due to the prevalent use of proprietary protocols. Without the knowledge of the program states and packet formats, it is difficult to enter the deep states for effective fuzzing.
Dongliang Fang, Zhanwei Song, Le Guan, Puzhuo Liu, Anni Peng, Yaowen Zheng, Peng Liu 0005, Hongsong Zhu, Limin Sun 0001
ACSAC9
2021 Asteria: Deep Learning-based AST-Encoding for Cross-platform Binary Code Similarity Detection
abstract
Binary code similarity detection is a fundamental technique for many security applications such as vulnerability search, patch analysis, and malware detection. There is an increasing need to detect similar code for vulnerability search across architectures with the increase of critical vulnerabilities in IoT devices. The variety of IoT hardware architectures and software platforms requires to capture semantic equivalence of code fragments in the similarity detection. However, existing approaches are insufficient in capturing the semantic similarity. We notice that the abstract syntax tree (AST) of a function contains rich semantic information. Inspired by successful applications of natural language processing technologies in sentence semantic understanding, we propose a deep learning-based AST-encoding method, named ASTERIA, to measure the semantic equivalence of functions in different platforms. Our method leverages the Tree-LSTM network to learn the semantic representation of a function from its AST. Then the similarity detection can be conducted efficiently and accurately by measuring the similarity between two representation vectors. We have implemented an open-source prototype of ASTERIA. The Tree-LSTM model is trained on a dataset with 1,022,616 function pairs and evaluated on a dataset with 95,078 function pairs. Evaluation results show that our method outperforms the AST-based tool Diaphora and the-state-of-art method Gemini by large margins with respect to the binary similarity detection. And our method is several orders of magnitude faster than Diaphora and Gemini for the similarity calculation. In the application of vulnerability search, our tool successfully identified 75 vulnerable functions in 5,979 IoT firmware images.
Shouguo Yang, Long Cheng 0005, Yicheng Zeng, Zhe Lang, Hongsong Zhu, Zhiqiang Shi
DSN5
2021 Maximal Clique Based Non-Autoregressive Open Information Extraction
abstract
Open Information Extraction (OpenIE) aims to discover textual facts from a given sentence.In essence, the facts contained in plain text are unordered.However, the popular Ope-nIE systems usually output facts sequentially in the way of predicting the next fact conditioned on the previous decoded ones, which enforce an unnecessary order on the facts and involve the error accumulation between autoregressive steps.To break this bottleneck, we propose MacroIE, a novel non-autoregressive framework for OpenIE.MacroIE firstly constructs a fact graph based on the table filling scheme, in which each node denotes a fact element, and an edge links two nodes that belong to the same fact.Then OpenIE can be reformulated as a non-parametric process of finding maximal cliques from the graph.It directly outputs the final set of facts in one go, thus getting rid of the burden of predicting fact order, as well as the error propagation between facts.Experiments conducted on two benchmark datasets show that our proposed model significantly outperforms current state-of-theart methods, beats the previous systems by as much as 5.7 absolute gain in F1 score.
Bowen Yu 0002, Tingwen Liu, Hongsong Zhu, Limin Sun 0001, Bin Wang 0004
EMNLP (1)4
2021 A Robust IoT Device Identification Method with Unknown Traffic Detection
Xiao Hu 0004, Hong Li 0004, Zhiqiang Shi, Hongsong Zhu, Limin Sun 0001
WASA (1)5
2021 FIUD: A Framework to Identify Users of Devices
Yimo Ren, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
WASA (2)3
2021 Social engineering in cybersecurity: a domain ontology and knowledge graph application examples
abstract
Abstract Social engineering has posed a serious threat to cyberspace security. To protect against social engineering attacks, a fundamental work is to know what constitutes social engineering. This paper first develops a domain ontology of social engineering in cybersecurity and conducts ontology evaluation by its knowledge graph application. The domain ontology defines 11 concepts of core entities that significantly constitute or affect social engineering domain, together with 22 kinds of relations describing how these entities related to each other. It provides a formal and explicit knowledge schema to understand, analyze, reuse and share domain knowledge of social engineering. Furthermore, this paper builds a knowledge graph based on 15 social engineering attack incidents and scenarios. 7 knowledge graph application examples (in 6 analysis patterns) demonstrate that the ontology together with knowledge graph is useful to 1) understand and analyze social engineering attack scenario and incident, 2) find the top ranked social engineering threat elements (e.g. the most exploited human vulnerabilities and most used attack mediums), 3) find potential social engineering threats to victims, 4) find potential targets for social engineering attackers, 5) find potential attack paths from specific attacker to specific target, and 6) analyze the same origin attacks.
Zuoguang Wang, Hongsong Zhu, Limin Sun 0001
Cybersecur.2
2020 TPLinker: Single-stage Joint Extraction of Entities and Relations Through Token Pair Linking
abstract
Extracting entities and relations from unstructured text has attracted increasing attention in recent years but remains challenging, due to the intrinsic difficulty in identifying overlapping relations with shared entities.Prior works show that joint learning can result in a noticeable performance gain.However, they usually involve sequential interrelated steps and suffer from the problem of exposure bias.At training time, they predict with the ground truth conditions while at inference it has to make extraction from scratch.This discrepancy leads to error accumulation.To mitigate the issue, we propose in this paper a one-stage joint extraction model, namely, TPLinker, which is capable of discovering overlapping relations sharing one or both entities while immune from the exposure bias.TPLinker formulates joint extraction as a token pair linking problem and introduces a novel handshaking tagging scheme that aligns the boundary tokens of entity pairs under each relation type.Experiment results show that TPLinker performs significantly better on overlapping and multiple relation extraction, and achieves state-of-the-art performance on two public datasets 1 .
Bowen Yu 0002, Tingwen Liu, Hongsong Zhu, Limin Sun 0001
COLING5
2020 Malware Classification Using Attention-Based Transductive Learning Network
Liting Deng, Hui Wen 0001, Mingfeng Xin, Limin Sun 0001, Hongsong Zhu
SecureComm (2)6
2020 Detecting Internet-Scale NATs for IoT Devices Based on Tri-Net
Zhaoteng Yan, Hui Wen 0001, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001
WASA (1)5
2020 EdgeCC: An Authentication Framework for the Fast Migration of Edge Services Under Mobile Clients
Hongsong Zhu, Hong Li 0004, Limin Sun 0001
WASA (1)3
2019 Remote Fingerprinting on Internet-Wide Printers Based on Neural Network
abstract
Nowadays, a large number of printers are connecting to the Internet. It is undoubtedly true that these online printers are facing severe cyber threats. However, the research/researchers so far cannot answer the current security status of Internet-wide printers. The principal difficulty is to identify the exact brand and model of online printers with high coverage, which is the primary element in vulnerability description. In this work, we design and implement a system called PrinterRadar. Based on Neural Network, PrinterRadar can automatically generate fingerprints from the application layer protocol banners of online printers, and the precision and recall rate of fingerprints can achieve 98% and 97%. By fingerprint matching with the banners which were collected from Censys and Shodan, PrinterRadar found that 508,719 printers were connected to the Internet, covering 92 printer brands and 4,188 printer models. As to the discovered printers, it is about twice the number of those detected by Censys and Shodan.
Zhaoteng Yan, Shichao Lv, Hongsong Zhu, Limin Sun 0001
GLOBECOM4
2019 Side-Channel Information Leakage of Traffic Data in Instant Messaging
abstract
Instant Messaging has been widely applied for both corporate use and personal use in recent years. Major Instant Messaging service providers adopt the Push Technology to ensure the immediacy of message forwarding, which efficiently provides a great convenience for user. However, the immediacy feature causes side-channel information leakage even if some protection measures has been implemented, such as information encryption strategy. In particular, we observe that senders' traffic flows have a strong temporal correlation with those of corresponding recipients, since the messages are forwarded to recipients as soon as they are received by servers. Based on the observation, attackers can infer real-time communications between pairwise users and even the social connections of users. In this paper, we present a methodology framework to validate this side-channel information leakage, which identifies users of real-time communications by matching the pairwise time sequences of traffic flows. We evaluate the method on the collected real-world data. The experimental results show that users' communications can be identified with a high accuracy, and 6 groups of users are inferred to have strong connections based on the data collected from a local area networks.
Ke Li 0042, Hong Li 0004, Hongsong Zhu, Limin Sun 0001, Hui Wen 0001
IPCCC3
2019 An Efficient Greybox Fuzzing Scheme for Linux-based IoT Programs Through Binary Static Analysis
abstract
With the rapid growth of Linux-based IoT devices such as network cameras and routers, the security becomes a concern and many attacks utilize vulnerabilities to compromise the devices. It is crucial for researchers to find vulnerabilities in IoT systems before attackers. Fuzzing is an effective vulnerability discovery technique for traditional desktop programs, but could not be directly applied to Linux-based IoT programs due to the special execution environment requirement. In our paper, we propose an efficient greybox fuzzing scheme for Linux-based IoT programs which consist of two phases: binary static analysis and IoT program greybox fuzzing. The binary static analysis is to help generate useful inputs for efficient fuzzing. The IoT program greybox fuzzing is to reinforce the IoT firmware kernel greybox fuzzer to support IoT programs. We implement a prototype system and the evaluation results indicate that our system could automatically find vulnerabilities in real-world Linux-based IoT programs efficiently.
Yaowen Zheng, Zhanwei Song, Yuyan Sun, Hongsong Zhu, Limin Sun 0001
IPCCC5
2019 Understanding and Securing Device Vulnerabilities through Automated Bug Report Analysis
Xuan Feng 0005, Xiaojing Liao, XiaoFeng Wang 0001, Haining Wang 0001, Qiang Li 0007, Kai Yang 0037, Hongsong Zhu, Limin Sun 0001
USENIX Security Symposium7
2019 FIRM-AFL: High-Throughput Greybox Fuzzing of IoT Firmware via Augmented Process Emulation
Yaowen Zheng, Ali Davanian, Heng Yin 0001, Chengyu Song, Hongsong Zhu, Limin Sun 0001
USENIX Security Symposium5
2019 ONE-Geo: Client-Independent IP Geolocation Based on Owner Name Extraction
Hongsong Zhu, Hai Zhao 0002, Hong Li 0004, Limin Sun 0001
WASA3
2019 Lightweight IoT Malware Visualization Analysis via Two-Bits Networks
Hui Wen 0001, Hongsong Zhu, Limin Sun 0001
WASA5
2019 IoTTracker: An Enhanced Engine for Discovering Internet-of-Thing Devices
abstract
Effectively identifying IoT devices in cyberspace is significant for grasping the security posture of cyberspace. However, there are still some IoT devices without vendor or product keywords in response data that cannot be identified by existing device identification engines. In this paper, we propose a new engine (IoT Tracker)for identifying IoT devices by leveraging the highest similarity of response data between IoT devices of the same vendor or product. Based on the protocol features, IoT Tracker divides application-layer protocols into semistructured data protocols and unstructured data protocols. For each category, IoT Tracker extracts structure structure, style structure or simhash feature. Then, IoTTracker utilizes features extracted from the response data to identify IoT devices. We implement a prototype of our proposed engine and evaluate its effectiveness through real-world experiments. The experimental results show that IoTTracker yield very high accuracy with 95.54 % precision and 93.08 % recall at vendor-level. Compared with existing methods, IoT Tracker adds 40.76% of identifiable devices after de-duplicating the identified dataset.
Xuan Feng 0005, Hongsong Zhu, Limin Sun 0001, Yuchi Zou
WOWMOM4
2019 Towards IP geolocation with intermediate routers based on topology discovery
abstract
IP geolocation determines geographical location by the IP address of Internet hosts. IP geolocation is widely used by target advertising, online fraud detection, cyber-attacks attribution and so on. It has gained much more attentions in these years since more and more physical devices are connected to cyberspace. Most geolocation methods cannot resolve the geolocation accuracy for those devices with few landmarks around. In this paper, we propose a novel geolocation approach that is based on common routers as secondary landmarks (Common Routers-based Geolocation, CRG). We search plenty of common routers by topology discovery among web server landmarks. We use statistical learning to study localized (delay, hop)-distance correlation and locate these common routers. We locate the accurate positions of common routers and convert them as secondary landmarks to help improve the feasibility of our geolocation system in areas that landmarks are sparsely distributed. We manage to improve the geolocation accuracy and decrease the maximum geolocation error compared to one of the state-of-the-art geolocation methods. At the end of this paper, we discuss the reason of the efficiency of our method and our future research.
Hong Li 0004, Qiang Li 0007, Wei Li 0059, Hongsong Zhu, Limin Sun 0001
Cybersecur.5
2018 A graph neural network based efficient firmware information extraction method for IoT devices
abstract
The firmware information for IoT devices includes the manufacturer, the device type, the device model and the firmware version, etc. Identifying firmware information helps build firmware knowledge graph for many security applications, such as homologous analysis and vulnerability detection of firmware. The traditional firmware information identifying method only utilizes the content-based information, lacks the utilization of the structure information of the firmware, and more importantly, it lacks the use of timing information. Lacking of structural information can reduce prediction accuracy, and lacking of timing information will make it difficult to predict the firmware version. In order to address the disadvantages of the existing method, this paper abstracts the directories or files (components) of the firmware into the nodes of the graph and abstracts the relationships between the nodes into the edges of the graph. Timing information such as component creation time and component version are also attached to the node properties to introduce the time sequence features. As a result, the experimental results show that the accuracy of our method is better than that of random forest for the all four tasks (manufacture, device type, device model and firmware version identification). Particularly, and the accuracy rate is greatly improved in the firmware version identification task.
Hong Li 0004, Hui Wen 0001, Hongsong Zhu, Limin Sun 0001
IPCCC4
2018 Mining Human Periodic Behaviors Using Mobility Intention and Relative Entropy
Feng Yi, Libo Yin, Hui Wen 0001, Hongsong Zhu, Limin Sun 0001, Gang Li 0009
PAKDD (1)4
2016 Identification of visible industrial control devices at Internet scale
abstract
Nowadays industrial control devices are crucial for infrastructure-critical systems such as factories, power plants, and water treatment facilities. Devices with IP addresses are visible on the Internet and they connect cyber space and physical world. The first step in protecting devices from attackers is a deep understanding of the devices' characteristics in the cyber space. In this paper, we take a first step in this direction by investigating physical devices running one of the two specific protocols that are widely adopted in industrial control systems. In order to detect these devices in real-time, we propose a two-stage discovery mechanism: first filtering out unqualified hosts from 4 billion remote hosts and then identifying physical devices from qualified candidates. We have conducted a real-world experiment to verify the mechanism and identified dozens of thousands of physical devices from the entire Internet. Results show that our method discovers all devices in 20 hours with 89.5% precision and 79.3% recall.
Xuan Feng 0005, Qiang Li 0007, Qi Han 0001, Hongsong Zhu, Yan Liu 0021, Limin Sun 0001
ICC4
2016 Active Profiling of Physical Devices at Internet Scale
abstract
Nowadays, more and more physical devices embed computing and networking capabilities and are visible on the Internet. These devices include webcams, net-printers, and industrial control equipments, etc. Collecting information about these devices is crucial to preserve cyber-security and facilitate security auditing for system administrators. In this paper, we propose a scalable framework for physical device profiling. It leverages banner grabbing to identify device types and running services, and uses clock skew to determine a device ID. Our framework scales well. We implement a prototype system and use it to profile Webcams and industrial control device. The results show that our system can effectively profile and identify Webcams in real time. We deploy it on the cloud server and use it to detect 4 billion IP addresses to profile 1.2 million Webcams and more than 60 thousand industrial control devices in 20 hours.
Xuan Feng 0005, Qiang Li 0007, Qi Han 0001, Hongsong Zhu, Yan Liu 0021, Limin Sun 0001
ICCCN4
2016 A Lightweight Method for Accelerating Discovery of Taint-Style Vulnerabilities in Embedded Systems
Yaowen Zheng, Zhi Li 0018, Shiran Pan, Hongsong Zhu, Limin Sun 0001
ICICS5
2016 Learning multi-channel correlation filter bank for eye localization
Shiming Ge, Kaixuan Xie, Hongsong Zhu, Shuixian Chen
Neurocomputing5
2015 Vehicle Anomaly Detection Based on Trajectory Data of ANPR System
abstract
This paper proposes a machine-learning technique to detect vehicle anomalies from data captured by automatic number plate recognition (ANPR) system. The proposed anomaly detection technique is specially engineered to exploit both spatial and temporal features of vehicles captured by ANPR system, so as to accurately detect anomaly vehicles. We extensively evaluated the proposed technique using a two- month long dataset collected by a real world ANRP system, which has more than three hundred cameras deployed in a big city of China. The evaluation results show that our technique can effectively detect vehicle anomalies from the huge amount of data collected by the ANPR system. More importantly, our technique significantly outperforms existing schemes especially when the data collected by the ANRP system are noisy due to poor weather condition.
Yuyan Sun, Hongsong Zhu, Limin Sun 0001
GLOBECOM2
2015 Cryptanalysis and improvement of two RFID-OT protocols based on quadratic residues
abstract
The ownership transfer of RFID tag means a tagged product changes control over the supply chain. Recently, Doss et al. proposed two secure RFID tag ownership transfer (RFID-OT) protocols based on quadratic residues. However, we find that they are vulnerable to the desynchronization attack. The attack is probabilistic. As the parameters in the protocols are adopted, the successful probability is 93.75%. We also show that the use of the pseudonym of the tag h(TID) and the new secret key KTIDare not feasible. In order to solve these problems, we propose the improved schemes. Security analysis shows that the new protocols can resist in the desynchronization attack and other attacks. By optimizing the performance of the new protocols, it is more practical and feasible in the large-scale deployment of RFID tags.
Yongming Jin, Hongsong Zhu, Zhiqiang Shi, Xiang Lu 0004, Limin Sun 0001
ICC2
2015 k-Perimeter Coverage Evaluation and Deployment in Wireless Sensor Networks
Changying Li, Jiguo Yu, Hongsong Zhu, Yuyan Sun
WASA4
2015 ROCS: Exploiting FM Radio Data System for Clock Calibration in Sensor Networks
abstract
Clock synchronization is critical for many WSNs due to the need of inter-node coordination and collaborative information processing. Existing protocols based on message passing achieve satisfactory clock synchronization accuracy, however, incur prohibitively high overhead especially in large-scale networks. In this paper, we propose a new clock synchronization approach called ROCS which exploits the radio data system (RDS) from FM radio stations. First, we design a new hardware FM receiver that can extract a periodic pulse from FM broadcasts, referred to as RDS clock. We then conduct a large-scale measurement study of RDS clock in our lab for a period of six days and on a vehicle driving through a metropolitan area of over 40km2. Our results show that RDS clock is highly stable and hence is a viable means to calibrate the clocks of large-scale city-wide sensor networks. To reduce the high power consumption of FM receiver, ROCS adaptively calibrates the native clock via the RDS clock. We implement ROCS in TinyOS on our hardware FM receiver and a TelosB-compatible WSN platform. Our extensive experiments using a 12-node testbed and our driving measurement traces show that ROCS achieves accurate and precise clock synchronization with low power consumption.
Liqun Li, Limin Sun 0001, Guoliang Xing, Wei Huangfu, Ruogu Zhou, Hongsong Zhu
IEEE Trans. Mob. Comput.6
2012 Exploiting ephemeral link correlation for mobile wireless networks
abstract
In wireless mobile networks, energy can be saved by using dynamic transmission scheduling with pre-knowledge about channel conditions. Such pre-knowledge can be obtained via profiling as proposed by several existing systems which assumed that the existence of spatial link correlation makes the measured channel status at one location reusable over a long period of time. Our empirical data, however, tells a different story: spatial link correlation only maintains well within a short duration (from seconds to tens of seconds) while decreases significantly afterwards, a phenomena we call ephemeral link correlation. By leveraging this observation, we design and implement a real-time transmission scheduling system, named PreSeer, on the railway platform for cargo transportation, where the transmission of a sink (to cellular towers) can be scheduled intelligently by utilizing future channel status measured by sinks located in front of it on the same train. We have implemented and evaluated the PreSeer system on the collected data extensively over 7,000-kilometer railway routes during a period of one and half years. Results reveal that PreSeer can help save as much as 40% energy, comparing with three base-line algorithms. More importantly, lessons learned from this major effort provide useful guidelines for transmission scheduling in highly-dynamic mobile environments, where (i) channel measurements cannot be perfectly aligned due to varying vehicle velocities, and (ii) the accuracy of channel measurements is subject to hardware discrepancy as well as environment irregularity.
Wei Liu 0053, Limin Sun 0001, Yunhuai Liu, Hongsong Zhu, Ziguo Zhong, Tian He 0001
SenSys4
2012 An energy-efficient link quality monitoring scheme for wireless networks
abstract
Abstract Link quality is one of the most important factors that affect the performance of wireless networks. In a densely deployed wireless network, continuous link quality monitoring consumes significant amount of energy and bandwidth at each node. In this paper, we propose a sensitivity model and a spatial correlation model that can be used to derive a set of deputy links to monitor, instead of monitoring all of the links in the network. The proposed scheme can improve energy efficiency of the link quality monitoring process. A greedy algorithm is presented to derive the deputy links set based on three different optimization objective functions. Performance of the proposed method is studied extensively and it is shown that the proposed method can save almost 90% energy in typical simulation scenarios than the method of monitoring all links. We also demonstrate that the energy consumption of the greedy deputy set‐based method is upper‐bounded. Copyright © 2010 John Wiley & Sons, Ltd.
Hongsong Zhu, Xinrong Li, Yongjun Xu 0001, Xiaowei Li 0001, Yan Liu 0021
Wirel. Commun. Mob. Comput.1
2011 RestThing: A Restful Web Service Infrastructure for Mash-Up Physical and Web Resources
abstract
In the field of Cyber Physical Systems and Pervasive Computing, physical resources and web resources can be easily handled and seamlessly integrated into our life. However, due to the heterogeneity of devices and tight coupling of individual information systems, the developers cannot easily create their specific applications by combining with physical and web resources. In this paper, we proposed Rest Thing which is a restful web service infrastructure based on REST principles in order to hide the heterogeneity of devices and provide a seamless way to integrate embedded devices with existing web applications. Besides, we implemented a prototyping system, which provided the restful accessible way of the wireless sensors, and built a demo application on the smart phone to collect and merge physical and web resources. Finally, we gave the performance evaluation of the prototyping system.
Weijun Qin, Qiang Li 0007, Limin Sun 0001, Hongsong Zhu, Yan Liu 0021
EUC4
2011 Exploiting FM radio data system for adaptive clock calibration in sensor networks
abstract
Clock synchronization is critical for Wireless Sensor Networks (WSNs) due to the need of inter-node coordination and collaborative information processing. Although many message passing protocols can achieve satisfactory clock synchronization accuracy, they incur prohibitively high overhead when the network scales to more than tens of nodes. An alternative approach is to take advantage of the global time reference induced by existing infrastructures including GPS, timekeeping radio stations, or power grid. However, high power consumption and geographic constraints present them from being widely adopted in WSNs. In this paper, we propose ROCS, a new clock synchronization approach exploiting the Radio Data System (RDS) of FM radios. First, we design a new hardware FM receiver that can extract a periodic pulse from FM broadcasts, referred to as RDS clock. We then conduct a large-scale measurement study of RDS clock in our lab for a period of six days and on a vehicle driving through a metropolitan area of over 40 $km^2$. Our results show that RDS clock is highly stable and hence is a viable means to calibrate the clocks of large-scale city-wide sensor networks. To reduce the high power consumption of FM receiver, ROCS intelligently predicts the time error due to drift, and adaptively calibrates the native clock via the RDS clock. We implement ROCS in TinyOS on our hardware FM receiver and a TelosB-compatible WSN platform. Our extensive experiments using a 12-node testbed and our driving measurement traces show that ROCS achieves accurate and precise clock synchronization with low power consumption.
Liqun Li, Guoliang Xing, Limin Sun 0001, Wei Huangfu, Ruogu Zhou, Hongsong Zhu
MobiSys6
2011 Demo: a sensor network time synchronization protocol based on fm radio data system
abstract
(1) Institute of Software, Chinese Academy of Sciences, China; (2) Graduate University, Chinese Academy of Sciences, China; (3) Department of Computer Science and Engineering, Michigan State University, United States
Liqun Li, Guoliang Xing, Limin Sun 0001, Wei Huangfu, Ruogu Zhou, Hongsong Zhu
MobiSys6
2005 A Self-adaptive Energy-Aware Data Gathering Mechanism for Wireless Sensor Networks
Limin Sun 0001, Ting-Xin Yan, Yanzhong Bi, Hongsong Zhu
ICIC (2)4