Zhi Li 0018

dblp:43/3166-18 · DBLP profile ↗
← Back
37ranked-venue papers
1as first author
28since 2021 · last 2026
0000-0001-7071-2976ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 1 first-author · 7 since 2021Security and privacy · 12 · 11 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NS-FirmID: A Neuro-Symbolic Multi-Agent Framework for Reliable Firmware Version Identification at Internet Scale
Fengshi Zhang, Zhi Li 0018, Shunchao Xu, Yongle Chen, Dongliang Fang, Limin Sun 0001
DSN2
2026 Chronos: Large-Scale Online Firmware Version Detection via Inadvertent Chronological Fingerprints
Fengshi Zhang, Zhi Li 0018, Shunchao Xu, Dongliang Fang, Yongle Chen, Limin Sun 0001
INFOCOM2
2026 NetID-GPT: Adapting large language models for large-scale internet-connected device identification
Zhi Li 0018, Shunchao Xu, Fengshi Zhang, Zhanwei Song, Dongliang Fang, Yongle Chen, Limin Sun 0001
Comput. Networks2
2026 Electromagnetic interference (EMI) backdoor: An EMI-based backdoor attack against computer vision systems
abstract
Recently, computer vision systems, for example, smart traffic surveillance systems, facial recognition systems, etc., have significantly changed our daily life. Even though the neural networks in such systems are known to suffer from backdoor attacks, causing the backdoored models to behave well on benign samples but maliciously on controlled samples (with triggers applied to activate the backdoor), it is generally believed that most of the triggers, when used in physical attacks, are noticeable to victim users and not robust in various settings, such as different angles, distances, lighting conditions, etc. In this paper, we leverage electromagnetic interference (EMI) to produce a specific pattern distortion in images captured by the camera system and utilize the pattern distortion as the backdoor trigger. To avoid the overhead of manually collecting poisoned images, we introduce a simulation sample generation approach, converting clean images to poisoned ones by simulating the distortion caused by EMI against the camera system. Additionally, we propose a contrast loss function to enhance the generalization of backdoor features, improving triggers’ capability to activate the embedded backdoors. We conduct extensive physical experiments using diverse deep neural networks across various camera systems in different practical environments, achieving a 92.54% average backdoor success rate.
Mengjie Sun, Peizhuo Lv, Shengzhi Zhang, Jianshuo Liu, Kai Chen 0012, Hong Li 0004, Zhi Li 0018, Qinhong Jiang, Limin Sun 0001
J. Comput. Secur.7
2026 TLCFI-PLC: Trampoline-Based Lightweight Control Flow Integrity Scheme for Protecting PLC
Kaixiang Liu, Junjiao Liu, Zhiwen Pan, Shichao Lv, Xin Chen 0123, Zhi Li 0018, Yuqi Chen 0001, Limin Sun 0001
IEEE Trans. Inf. Forensics Secur.6
2025 PNetGPT: Proprietary Protocol Network Traffic Generation with Pre-trained Transformer
abstract
Generative pre-trained transformers are exceedingly effective as generative models and classifiers, widely used in natural language processing and computer vision. This work contributes to the exploration of generative pre-trained transformer-based models in the proprietary protocol network traffic. However, building a pre-trained model for proprietary protocol network traffic is non-trivial due to the heterogeneous unknown formats and the extreme scarcity of proprietary protocol network traffic datasets. In this paper, we present PNetGPT, a pre-trained transformer-based model for generating proprietary protocol network traffic. We have constructed the inaugural dataset of 2 real-world proprietary protocols. After training on this dataset, PNetGPT possesses the capacity to generate high-quality proprietary protocol network traffic to support various applications of proprietary protocols, including reverse analysis, protocol fuzzy testing, intrusion detection, etc. We evaluated PNetGPT with two real proprietary protocols and demonstrated state-of-the-art (SOTA) performance in handling heterogeneous unknown formats. The code and datasets are available at: https://github.com/Snail1502/PNetGPT
Zedong Li, Dongliang Fang, Xin Chen 0123, Zhanwei Song, Zhi Li 0018, Shichao Lv, Limin Sun 0001
ICASSP6
2025 Leveraging Fine-Tuned Large Language Models for Device Fingerprint Extraction in IoT Security
Haoyu Bin, Gaosheng Wang, Yimo Ren, Zhi Li 0018, Hongsong Zhu
ICIC (4)5
2025 Exp-Arch: A Novel LLM-Powered Approach for Facilitating Exploit Primitive Assessment in the Linux Kernel
abstract
Transforming Linux kernel exploit primitives into full Privilege Escalation (PE) exploits is a critical, expertiseintensive, and time-consuming challenge, especially with constantly evolving kernel mitigations. While previous research has advanced automated kernel exploit development, these efforts often focused on specialized scenarios rather than providing a generalized, end-to-end framework for diverse primitives. This limitation restricts the exploration of a primitive's true exploit potential. This paper introduces Exp-Arch, a novel approach leveraging Large Language Models (LLMs) for the automated, end-to-end generation of PE exploits from kernel primitives. This process includes comprehensive initial assessment and subsequent exploitation. Exp-Arch's LLM-powered workflow systematically performs an in-depth semantic analysis of the input primitive, devises an intelligent strategic plan for the exploitation route, and then automates the synthesis and iterative closed-loop validation of the final PE exploit code. Exp-Arch offers accurate assessment of a primitive's exploitability and significantly accelerates the exploit development lifecycle. We evaluated ExpArch using various primitives from public Linux kernel 1-day vulnerabilities with commercial LLMs. The results show that Exp-Arch effectively converted 73 % (11 out of 15) of test cases into working kernel exploits, demonstrating its effectiveness in primitive evaluation.
Zuxin Chen, Zhi Li 0018, Zhanwei Song, Zhiqiang Shi, Limin Sun 0001
ICPADS2
2025 Demystifying Feature Engineering in Malware Analysis of API Call Sequences
abstract
Machine learning (ML) has been widely used to analyze API call sequences in malware analysis, which typically requires the expertise of domain specialists to extract relevant features from raw data. The extracted features play a critical role in malware analysis. Traditional feature extraction is based on human domain knowledge, while there is a trend of using natural language processing (NLP) for automatic feature extraction. This raises a question: how do we effectively select features for malware analysis based on API call sequences? To answer it, this paper presents a comprehensive study of investigating the impact of feature engineering upon malware classification. We first conducted a comparative performance evaluation under three models, Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and Transformer, with respect to knowledgebased and NLP-based feature engineering methods. We observed that models with knowledge-based feature engineering inputs generally outperform those using NLP-based across all metrics, especially under smaller sample sizes. Then we analyzed a complete set of data features from API call sequences, our analysis reveals that models often focus on features such as handles and virtual addresses, which vary across executions and are difficult for human analysts to interpret.
Tianheng Qu, Hongsong Zhu, Limin Sun 0001, Haining Wang 0001, Haiqiang Fei, Zhi Li 0018
RAID7
2025 EMFuzz: Use Electromagnetic Fuzzing for Automated Attack Surface Assessment of Actuators
abstract
Actuators are essential components in cyber-physical systems, enabling system modules to perform diverse and complex tasks. Unfortunately, the pursuit of higher functional complexity often correlates with a broader attack surface in actuators. Thus, an efficient automated attack surface assessment is crucial to avoid cyber incidents in critical infrastructures. Limited by enormous parameter spaces, current methods rely on heuristic tests to evaluate interference potential but cannot thoroughly investigate the full spectrum of potential hidden interference. The observation that similar interference trigger configurations lead to the same impact has motivated us to use machine learning algorithms for understanding different impact samples around decision boundaries. By leveraging generalized knowledge of responses against specific attack scenarios, we aim to improve the efficiency of automated attack surface assessment of electromagnetic interference on new targets. To this end, we introduce EMFuzz, an automated mechanism to fuzz hardware to quantify varying adverse effects. We evaluate EMFuzz on 16 new servos within real-world scenarios, where it achieves an 86% accuracy in classifying different attack vectors. With the same test time, EMFuzz uncovers over twice the effective attack configurations of the baseline, greatly improving assessment efficiency. To further validate its efficacy, we apply EMFuzz to assess the attack surface of a new actuator from a robot transfer unit, and it can successfully reveal three distinct adverse effects.
Shiquan Dong, Zhi Li 0018, Jianshuo Liu, Hong Li 0004, Dongliang Fang, Shichao Lv, Haining Wang 0001, Limin Sun 0001
IEEE Trans. Inf. Forensics Secur.2
2025 LLM-Powered Static Binary Taint Analysis
abstract
This article proposes LATTE , the first static binary taint analysis that is powered by a large language model (LLM). LATTE is superior to the state of the art (e.g., Emtaint, Arbiter, Karonte) in three aspects. First, LATTE is fully automated while prior static binary taint analyzers need rely on human expertise to manually customize taint propagation rules and vulnerability inspection rules. Second, LATTE is significantly effective in vulnerability detection, demonstrated by our comprehensive evaluations. For example, LATTE has found 37 new bugs in real-world firmware, which the baselines failed to find. Moreover, 10 of them have been assigned CVE numbers. Lastly, LATTE incurs remarkably low engineering cost, making it a cost-efficient and scalable solution for security researchers and practitioners. We strongly believe that LATTE opens up a new direction to harness the recent advance in LLMs to improve vulnerability analysis for binary programs.
Puzhuo Liu, Chengnian Sun, Yaowen Zheng, Xuan Feng 0005, Zhi Li 0018, Peng Di, Yu Jiang 0001, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.8
2024 MOMR: A Threat in Web Application Due to the Malicious Orchestration of Microservice Requests
abstract
Microservice is an increasingly favored architecture for constructing modern web applications and the fast-paced business requirements facilitate the transmission of microservice traffic among distributed servers. In contrast to traditional architectures, microservice architecture has tight inherent dependencies between microservice units when supporting web application business. Attackers can excavate these dependencies to maliciously orchestrate microservice requests, scheduling microservice traffic to converge on the target link. This attack disrupts link and application quality of service, bringing new potential threats to web applications and cyberspace security. This work analyzes and evaluates the threat due to the malicious orchestration of microservice requests (MOMR) with the initial intention of promoting microservice application security and other information system security based on the microservice architecture. A Cross-Layer Coupling (CLC) model is proposed that aims to describe microservice traffic transmission, which efficiently supports the threat evaluation. A Path-aware Microservice Traffic Scheduling (PMTS) attack method is imposed on the CLC model so that it can construct the MOMR threat accurately. To demonstrate the effectiveness of the proposed method in evaluating the MOMR threat, a comprehensive analysis is performed on a typical microservice application and a semi- physical simulation platform. The result shows the threat causes performance degradation and impacts the network, such as a packet loss rate of up to 79% and an RTT increase of 600% of the target link.
Chunyang Zheng, Jinfa Wang, Shuaizong Si, Zhi Li 0018, Limin Sun 0001
ICC4
2024 LibvDiff: Library Version Difference Guided OSS Version Identification in Binaries
abstract
Open-source software (OSS) has been extensively employed to expedite software development, inevitably exposing downstream software to the peril of potential vulnerabilities. Precisely identifying the version of OSS not only facilitates the detection of vulnerabilities associated with it but also enables timely alerts upon the release of 1-day vulnerabilities. However, current methods for identifying OSS versions rely heavily on version strings or constant features, which may not be present in compiled OSS binaries or may not be representative when only function code changes are made. As a result, these methods are often imprecise in identifying the version of OSS binaries being used.
Chaopeng Dong, Siyuan Li 0014, Shouguo Yang, Yang Xiao 0011, Yongpan Wang, Hong Li 0004, Zhi Li 0018, Limin Sun 0001
ICSE7
2024 Concrete Constraint Guided Symbolic Execution
abstract
Symbolic execution is a popular program analysis technique. It systematically explores all feasible paths of a program but its scalability is largely limited by the path explosion problem, which causes the number of paths proliferates at runtime. A key idea in existing methods to mitigate this problem is to guide the selection of states for path exploration, which primarily relies on the features to represent program states. In this paper, we propose concrete constraint guided symbolic execution, which aims to cover more concrete branches and ultimately improve the overall code coverage during symbolic execution. Our key insight is based on the fact that symbolic execution strives to cover all symbolic branches while concrete branches are neglected, and directing symbolic execution toward uncovered concrete branches has a great potential to improve the overall code coverage. The experimental results demonstrate that our approach can improve the ability of KLEE to both increase code coverage and find more security violations on 10 open-source C programs.
Guowei Yang 0001, Shichao Lv, Zhi Li 0018, Limin Sun 0001
ICSE4
2024 NFCEraser: A Security Threat of NFC Message Modification Caused by Quartz Crystal Oscillator
abstract
Near Field Communication (NFC) has been widely used for rapid data exchange between electronic devices over a very short distance. In this paper, we reveal a new security vulnerability in NFC passive communication channels where transferred data can be modified in real-time. The security threat of data modification posed by this vulnerability is called NFCEraser. Exploiting electromagnetic interference (EMI), NFCEraser injects signals into the crystal oscillator’s electrode and adjusts the amplitude of carrier signals in NFC communication channels. By manipulating the parameters of EMI signals, NFCEraser is able to arbitrarily flip the bits in data payload sent from an NFC peer device, which may cause serious security outcomes. To assess the severity of NFCEraser, we examine six NFC modules under NFC-A/B communication modes and successfully perform reading operations under a variety of data lengths. The experimental results show that NFCEraser can modify data bits in response frames from NFC peer devices with the maximum 89% accuracy, under around 0.21μs latency. Our analysis further shows that NFCEraser can maintain an attack success rate of no less than 85% in environments with typical levels of electromagnetic noise.
Jianshuo Liu, Hong Li 0004, Mengjie Sun, Haining Wang 0001, Hui Wen 0001, Zhi Li 0018, Limin Sun 0001
SP6
2024 Battling against Protocol Fuzzing: Protecting Networked Embedded Devices from Dynamic Fuzzers
abstract
N etworked E mbedded D evices (NEDs) are increasingly targeted by cyberattacks, mainly due to their widespread use in our daily lives. Vulnerabilities in NEDs are the root causes of these cyberattacks. Although deployed NEDs go through thorough code audits, there can still be considerable exploitable vulnerabilities. Existing mitigation measures like code encryption and obfuscation adopted by vendors can resist static analysis on deployed NEDs, but are ineffective against protocol fuzzing. Attackers can easily apply protocol fuzzing to discover vulnerabilities and compromise deployed NEDs. Unfortunately, prior anti-fuzzing techniques are impractical as they significantly slow down NEDs, hampering NED availability. To address this issue, we propose Armor—the first anti-fuzzing technique specifically designed for NEDs. First, we design three adversarial primitives–delay, fake coverage, and forged exception–to break the fundamental mechanisms on which fuzzing relies to effectively find vulnerabilities. Second, based on our observation that inputs from normal users consistent with the protocol specification and certain program paths are rarely executed with normal inputs, we design static and dynamic strategies to decide whether to activate the adversarial primitives. Extensive evaluations show that Armor incurs negligible time overhead and effectively reduces the code coverage (e.g., line coverage by 22%-61%) for fuzzing, significantly outperforming the state of the art.
Puzhuo Liu, Yaowen Zheng, Chengnian Sun, Hong Li 0004, Zhi Li 0018, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.5
2024 Asteria-Pro: Enhancing Deep Learning-based Binary Code Similarity Detection by Incorporating Domain Knowledge
abstract
Widespread code reuse allows vulnerabilities to proliferate among a vast variety of firmware. There is an urgent need to detect these vulnerable codes effectively and efficiently. By measuring code similarities, AI-based binary code similarity detection is applied to detecting vulnerable code at scale. Existing studies have proposed various function features to capture the commonality for similarity detection. Nevertheless, the significant code syntactic variability induced by the diversity of IoT hardware architectures diminishes the accuracy of binary code similarity detection. In our earlier study and the tool Asteria , we adopted a Tree-LSTM network to summarize function semantics as function commonality, and the evaluation result indicates an advanced performance. However, it still has utility concerns due to excessive time costs and inadequate precision while searching for large-scale firmware bugs. To this end, we propose a novel deep learning-enhancement architecture by incorporating domain knowledge-based pre-filtration and re-ranking modules, and we develop a prototype named Asteria-Pro based on Asteria . The pre-filtration module eliminates dissimilar functions, thus reducing the subsequent deep learning-model calculations. The re-ranking module boosts the rankings of vulnerable functions among candidates generated by the deep learning model. Our evaluation indicates that the pre-filtration module cuts the calculation time by 96.9%, and the re-ranking module improves MRR and Recall by 23.71% and 36.4%, respectively. By incorporating these modules, Asteria-Pro outperforms existing state-of-the-art approaches in the bug search task by a significant margin. Furthermore, our evaluation shows that embedding baseline methods with pre-filtration and re-ranking modules significantly improves their precision. We conduct a large-scale real-world firmware bug search, and Asteria-Pro manages to detect 1,482 vulnerable functions with a high precision 91.65%.
Shouguo Yang, Chaopeng Dong, Yang Xiao 0011, Yiran Cheng, Zhiqiang Shi, Zhi Li 0018, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.6
2024 Batch-transformer for scene text image super-resolution
Yaqi Sun, Zhi Li 0018, Kai Yang 0037
Vis. Comput.3
2023 MalAder: Decision-Based Black-Box Attack Against API Sequence Based Malware Detectors
abstract
The API call sequence based malware detectors have proven to be promising, especially when incorporated with deep neural networks (DNNs). Several adversarial attack methods are proposed to fool these detectors by introducing undetectable perturbations into normal samples. However, in real-world scenarios, the malware detector provides only the predicted label for a given sample, without exposing its network architecture or output probability, making it challenging for adversarial attacks under the decision-based black-box. Existing work in this area typically relies on random-based methods that suffer high costs and low attack success rates. To address these limitations, we propose a novel decision-based black-box attack against API sequence based malware detectors, called MalAder. Our approach aims to improve the attack success rate as well as query efficiency through a directional perturbation algorithm. First, it utilizes attention-based API ranking to assess the importance of API calls in the context of different API sequences. This assessment guides the insertion position for perturbation. Then, the perturbation is carried out using benign distance perturbing, which gradually shortens the semantic distance from adversarial API sequences to a set of benign samples. Finally, our algorithm iteratively generates adversarial malware samples by performing perturbations. In addition, we have implemented MalAder and evaluated its performance against two classic malware detectors. The results show that MalAder outperforms state-of-the-art decision-based black-box adversarial attacks, proving its effectiveness.
Lei Cui 0003, Hui Wen 0001, Zhi Li 0018, Hongsong Zhu, Zhiyu Hao, Limin Sun 0001
DSN4
2023 UID-Auto-Gen: Extracting Device Fingerprinting from Network Traffic
abstract
The number of Internet device vulnerabilities has been quickly rising in recent years, rendering an explosion of network attacks. Device fingerprinting serves as the primary means for vulnerability awareness and attacker tracking. The current device fingerprinting approach can only achieve model-level identification within the Internet scope or individual-level identification for specific protocols (e.g., SSL) or scenarios (e.g., LAN). However, it is still difficult for these methods to achieve individual-level identification on a global Internet scale. In this paper, we propose a fingerprint extraction approach that is accurate to the individual level of the device by using a combination of clustering, multiple sequence alignment, and based on the geographic location stability of the device. In a continuous 3-month observation for several cities around the world, at least 1.54% of devices can be accurately extracted with unique IDs, with an accuracy rate of 99.30%, which is capable of being used in production environments.
Haoyu Bin, Zhi Li 0018, Rongrong Xi, Hongsong Zhu, Limin Sun 0001
IPCCC3
2023 Denoising Network of Dynamic Features for Enhanced Malware Classification
abstract
Malware classification based on dynamic feature analysis works by running malware in controlled and isolated environments to observe how it behaves. This technology widely uses the sequence of run-time API calls to classify. Malware often adopts evasion techniques such as obfuscation, encryption, and code injection to obfuscate classification results by introducing noise into the API sequence. The existing methods lack explicit means of filtering noise components in the data, which affects the accuracy of malware detection. To address this issue, we propose DenoMC, a malware classification method with an explicit denoising module. Firstly, we employ dynamic analysis and embedding techniques to encode the API sequence. Then, we introduce a soft thresholding mechanism in the residual network to achieve active filtering of noise components in API sequences. Finally, a BiLSTM model is adopted to enhance the temporal correlation among sequence of API calls and improve classification performance. Experiments conducted on real datasets demonstrate that DenoMC significantly improves malware classification accuracy compared to other state-of-art models. In addition, we validate the effectiveness of each module in DenoMC through extensive ablation studies.
Siyuan Li 0014, Hui Wen 0001, Liting Deng, Zhi Li 0018, Limin Sun 0001
IPCCC6
2023 HackMentor: Fine-Tuning Large Language Models for Cybersecurity
abstract
The democratization of artificial intelligence has made substantial progress by leveraging open-source large language models (LLMs), enabling researchers across domains to train customized models to meet their specific needs. Given the confidentiality and significance of cybersecurity, obtaining private and localized LLMs is imperative. However, general LLMs are not designed to cater specifically to this field, their general knowledge often falls short when addressing such specialized problems. In this paper, we categorize the domain instructions based on cybersecurity knowledge to guide the construction of high-quality instructions and conversations, ultimately enhancing the specialized capabilities of LLMs. The resulting fine-tuned LLMs, collectively termed HackMentor, are evaluated using WinRate, EloRating, and ZenoEval methods along with other popular LLMs. The experiments demonstrate that the proposed method yields significant performance improvements, surpassing the native LLMs by 10-25% when aligned with cybersecurity prompts. More, HackMentor exhibits comparable conversational quality to ChatGPT, while providing more concise and humanlike responses. This study demonstrates the efficacy of HackMentor in augmenting LLMs for cybersecurity requirements, paving the way for localized LLMs that meet specialized needs without compromising general capabilities.
Jie Zhang 0121, Hui Wen 0001, Liting Deng, Mingfeng Xin, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001
TrustCom5
2023 Owner name entity recognition in websites based on heterogeneous and dynamic graph transformer
Yimo Ren, Hong Li 0004, Jie Liu 0079, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001
Knowl. Inf. Syst.5
2023 Spenny: Extensive ICS Protocol Reverse Analysis via Field Guided Symbolic Execution
abstract
Industrial Control System (ICS) protocols have built a tight coupling between ICS components, including industrial software and field controllers such as Programmable Logic Controllers (PLCs). With more ICS components are exposed on the Internet, huge threats are emerging through the exploitation on the inherent defects of ICS protocols. However, the proprietary of ICS protocols makes it extremely hard to build intrusion detection system or perform penetration tests for ICS security reinforcement. In this work, we introduce a symbolic-execution based protocol reverse analysis framework to extract the message format and field type of ICS protocols from real-world PLC firmware. We design new coverage metric and path prioritization strategy to enhance symbolic execution for extensive protocol reverse analysis. Moreover, we propose a field-expression based method on protocol message format inference, along with the analysis on the value ranges of fields which are ignored by previous work. Our evaluation shows that our methods can extract more protocol information during symbolic execution, and achieve high accuracy on protocol reverse analysis compared to Wireshark. Furthermore, we equip the results on private ICS protocols with a black-box fuzzer to test two real-world PLCs. In total, we have found 10 vulnerabilities, including 4 new vulnerabilities.
Zhi Li 0018, Shichao Lv, Limin Sun 0001
IEEE Trans. Dependable Secur. Comput.2
2023 Internet-Scale Fingerprinting the Reusing and Rebranding IoT Devices in the Cyberspace
abstract
Fingerprinting Internet-of-Things(IoT) devices on types and brands is a necessary work for security analysis in the cyberspace. The existing approaches mainly rely on the dominant features of devices which is response to information in order to identify these online devices. However, the web server components reusing and products rebranding are the common phenomenons of these embedded IoT devices. It caused the existing approaches difficult to identify most devices even errors due to the similar responses. In this paper, we present an approach, IoTXray, which improves the work efficiently of information collection about accelerating the relations between reusing/rebranding devices with the corresponding manufacturers. And these relations can generate more accurate and reliable fingerprints than previous approaches. Using the mixed neural networks, IoTXray comprehensively detects the real manufactures of online IoT devices upon three different kinds of data sources. In the experiment, our approach can identify 7,025,854 IoT devices on HTTP-hosts. The identification rate has reached to several times higher than previous approaches. Our approach has especially detected 3,268,953 reusing and 963,653 rebranding devices with their original manufacturers.
Zhaoteng Yan, Zhi Li 0018, Hong Li 0004, Shouguo Yang, Hongsong Zhu, Limin Sun 0001
IEEE Trans. Dependable Secur. Comput.2
2022 Compromised IoT Devices Detection in Smart Home via Semantic Information
abstract
The safety and security of IoT devices in smart home systems is attracting booming attention, due to the cascading threat introduced by the interoperability of IoT devices. It is observed that semantic information of behaviors could be utilized to identify anomalies of IoT devices, and some prior works have attempted to detect single abnormal behavior based on the mined semantic patterns. However, the performance of these methods could be usually affected by some noisy data (e.g., user activities), suffering from false alarms of detection. In this work, we propose a semantic-aware framework of compromised IoT devices detection, which extracts the Mutual Information feature from the semantic information of IoT devices to eliminate interference of the noise. The proposed framework includes three modules: semantic analysis to generate event correlations, feature extraction to construct the feature vector, and detection model to train binary classifiers for detection. To collect real-world data for evaluation, we construct three testbeds of the following scenes: bedroom, living room and kitchen. The performance on the collected dataset shows that our method achieves high accuracy (the average is over 90.0%) on the compromised devices detection.
Ke Li 0042, Zhi Li 0018, Zhimin Gu, Ziying Wang, Limin Sun 0001
ICC2
2022 Characterizing Heterogeneous Internet of Things Devices at Internet Scale Using Semantic Extraction
abstract
Along with the rapid-growth number of Internet of Things (IoT) devices, significant security concerns are raised due to the hidden vulnerabilities among them. Illuminating the characteristics of online devices would shed a light on protecting these potential vulnerable devices. State-of-arts methodologies enumerate devices characteristics as keywords and rules and match them with IoT network data. However, the heterogeneous implementations of IoT devices introduce intricate characteristics features, which impede the large-scale identification. In this work, we close this gap and present a semantic extraction-based approach that can automatically and effectively characterize online devices. We leverage the observation that IoT devices can be identified by analyzing the semantic information of the network packets. Specifically, we first collect the network data of IoT devices and utilize a co-training algorithm to annotate the data. We propose a residual dilate gated convolutional neural network (RDGCNN)-based encoder to extract semantic features from the annotated data. Then, we put forward an entity relationship-based decoder to generate the characteristic triplet (type, brand, and model) of IoT devices by decoding extracted features. We have implemented the prototype of the system and conducted real-world experiments to evaluate the performance. Results show that our approach achieves 92.16% precision and 86.79% recall. In addition, we apply our proposed method to characterize 15 millions IoT devices on the Internet.
Kai Yang 0037, Xiaodong Lin 0001, Zhi Li 0018, Limin Sun 0001
IEEE Internet Things J.4
2022 Understanding Security Risks of Embedded Devices Through Fine-Grained Firmware Fingerprinting
abstract
An increasing number of embedded devices are connecting to the Internet, ranging from cameras, routers to printers, while an adversary can exploit security flaws already known to compromise those devices. Security patches are usually associated with the device firmware, which relies on the device vendors and products. Due to compatibility and release-time issues, many embedded devices are still using outdated firmware with known vulnerabilities or flaws. In this article, we conduct a systematic study on device vulnerabilities by leveraging firmware fingerprints. Specifically, we use a web crawler to gather 9,716 firmware images from official websites of device vendors, and 347,685 security reports scattered across data archives, blogs, and forums. We propose to generate fine-grained fingerprints based on the subtle differences between the filesystems of various firmware images. Furthermore, machine learning algorithms and regex are used to identify device vulnerabilities and corresponding device firmware fingerprints. We perform real-world experiments to validate the performance of the firmware fingerprint, which yields high accuracy of 91% precision and 90% recall. We reveal that 6,898 reports have the firmware and related vulnerability information, and there are more than 10% of firmware vulnerabilities without any patches or solutions for mitigating underlying security risks.
Qiang Li 0007, Dawei Tan, Haining Wang 0001, Zhi Li 0018, Jiqiang Liu
IEEE Trans. Dependable Secur. Comput.5
2020 Detecting Internet-Scale NATs for IoT Devices Based on Tri-Net
Zhaoteng Yan, Hui Wen 0001, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001
WASA (1)4
2018 PANDORA: A Scalable and Efficient Scheme to Extract Version of Binaries in IoT Firmwares
abstract
Open source components are widely used by IoT vendors to develop firmwares in devices. The exposure of vulnerabilities existing in some specific versions of the core components may cause severe security incidents such as the Heartbleed event in 2014 and the Sambacry event in 2016. Extracting the version information from various firmware binaries is significant for evaluating the influence of such incidents and providing emergency response services. To the best of our knowledge, there are still no scalable and efficient extraction methods for binary version information in IoT firmwares. The commonly used method for traditional softwares requires the running up of the firmwares and interaction such as '-version' to obtain the version information. This method is not applicable for IoT devices, as they are built from various platforms which makes it impossible to simulate all of the interested firmwares at large scale. In this paper, we design, implement and evaluate a scalable and efficient binary version extraction framework (termed as PANDORA) for IoT firmwares, which does not rely on the real runtime environment. The main idea of our methodology is to leverage version strings in binaries to get version information. We design a string recover engine (SRE) to recover the missing pieces of those incomplete version strings. We test PANDORA in a dataset containing 2683 IoT binary files. Surprisingly 2267 of them are version-extractable and the recognition rate can reach 84.5%.
Hong Li 0004, Zhi Li 0018, Limin Sun 0001
ICC4
2018 Towards Fine-grained Fingerprinting of Firmware in Online Embedded Devices
abstract
An increasing number of embedded devices are connecting to the Internet at a surprising rate. Those devices usually run firmware and are exposed to the public by device search engines. Firmware in embedded devices comes from different manufacturers and product versions. More importantly, many embedded devices are still using outdated versions of firmware due to compatibility and release-time issues, raising serious security concerns. In this paper, we propose generating fine-grained fingerprints based on the subtle differences between the filesystems of various firmware images. We leverage the natural language processing technique to process the file content and the document object model to obtain the firmware fingerprint. To validate the fingerprints, we have crawled 9,716 firmware images from official websites of device vendors and conducted real-world experiments for performance evaluation. The results show that the recall and precision of the firmware fingerprints exceed 90%. Furthermore, we have deployed the prototype system on Amazon EC2 and collected firmware in online embedded devices across the IPv4 space. Our findings indicate that thousands of devices are still using vulnerable firmware on the Internet.
Qiang Li 0007, Xuan Feng 0005, Haining Wang 0001, Zhi Li 0018, Limin Sun 0001
INFOCOM4
2017 MIAC: A Mobility Intention Auto-Completion Model for Location Prediction
Feng Yi, Zhi Li 0018, Hongtao Wang 0002, Limin Sun 0001
KSEM2
2016 A Lightweight Method for Accelerating Discovery of Taint-Style Vulnerabilities in Embedded Systems
Yaowen Zheng, Zhi Li 0018, Shiran Pan, Hongsong Zhu, Limin Sun 0001
ICICS3
2016 GUIDE: Graphical user interface fingerprints physical devices
abstract
Nowadays, the number of visible physical devices exposed on the Internet is dynamically increasing and they play a crucial role for bridging between the cyber space and the physical world, such as network printer, Webcam, and industrial control devices. Discovering these devices brings about the deep understanding on these devices' characteristics and help secure device security in the cyber space. A device fingerprint is a prerequisite of device discovery in the Internet. However, today's online device search depends on keywords of packet head fields and the keyword collection is done manually. This impedes an accurate and large-scale device discovery, due to high human efforts and inevitable human errors, as well as the difficulty of keeping keywords complete and updated. To address this problem, we propose GUIDE, a framework to automatically generate device fingerprints based on webpages embedded in these devices. In order to demonstrate how GUIDE works, we also develop its prototype system and provide a case study which discover surveillance devices in the cyber space.
Qiang Li 0007, Xuan Feng 0005, Zhi Li 0018, Haining Wang 0001, Limin Sun 0001
ICNP3
2016 Tensor Filter: Collaborative Path Inference from GPS Snippets of Vehicles
Hongtao Wang 0002, Hui Wen 0001, Feng Yi, Zhi Li 0018, Limin Sun 0001
WASA4
2015 A Poisson Distribution Based Topology Control Algorithm for Wireless Sensor Networks Under SINR Model
Kan Yu 0001, Zhi Li 0018, Qiang Li 0007, Jiguo Yu
WASA2
2010 Priority Linear Coding Based Opportunistic Routing for Video Streaming in Ad Hoc Networks
abstract
In this paper, we propose a priority (or progressive) linear coding based opportunistic routing mechanism (OR-PLC) for H.264 video streaming over multi-hop ad hoc networks. OR-PLC assigns different error protection priorities to video packets according to their perceptual importance to mitigate error propagation problem so that the video quality is enhanced in receiver. Furthermore, OR-PLC exploits the broadcast feature of wireless medium to improve the bandwidth utility. Compared with other opportunistic routing schemes, OR-PLC reduces the delay by progressive encoding and decoding. The experiments show that our mechanism outperforms two state-of-the-art schemes, i.e., MORE and MP-RTP. It turns out that OR-PLC delivers more than 3.5 dB PSNR gains in average, while using less bandwidth.
Zhi Li 0018, Limin Sun 0001, Xinyun Zhou, Liqun Li
GLOBECOM1