EDBT 2026 Demo / reviewers in the wild / expert
Hong Li 0004
dblp:93/6234-4
· DBLP profile ↗
83ranked-venue papers
3as first author
60since 2021 · last 2026
0000-0003-1353-7838ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 3 first-author · 13 since 2021Security and privacy · 20 · 18 since 2021Software engineering, systems software and programming languages · 11 · 10 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeechShield: Latency-Efficient and Robust Timbre-Aware Voice Protection Against Speech Synthesis Deepfake Attacks
Jianshuo Liu, Shiquan Dong, Hong Li 0004, Chenghua Gao, Kang G. Shin, Haining Wang 0001, Yimo Ren, Limin Sun 0001 |
DSN | 3 |
| 2026 | Modubin: A Binary Modularization Approach Based on the Locality of Homologous Functions
Wenyan Yu, Lei Cui 0003, Jiayuan Li 0002, Hong Li 0004, Hongsong Zhu |
ICPC | 5 |
| 2026 | StruFSM: Byte-level structural modeling for protocol finite state machine inference
Zhen Wang 0043, Yimo Ren, Zhaoteng Yan, Hong Li 0004, Hongsong Zhu |
Comput. Networks | 7 |
| 2026 | Electromagnetic interference (EMI) backdoor: An EMI-based backdoor attack against computer vision systemsabstractRecently, computer vision systems, for example, smart traffic surveillance systems, facial recognition systems, etc., have significantly changed our daily life. Even though the neural networks in such systems are known to suffer from backdoor attacks, causing the backdoored models to behave well on benign samples but maliciously on controlled samples (with triggers applied to activate the backdoor), it is generally believed that most of the triggers, when used in physical attacks, are noticeable to victim users and not robust in various settings, such as different angles, distances, lighting conditions, etc. In this paper, we leverage electromagnetic interference (EMI) to produce a specific pattern distortion in images captured by the camera system and utilize the pattern distortion as the backdoor trigger. To avoid the overhead of manually collecting poisoned images, we introduce a simulation sample generation approach, converting clean images to poisoned ones by simulating the distortion caused by EMI against the camera system. Additionally, we propose a contrast loss function to enhance the generalization of backdoor features, improving triggers’ capability to activate the embedded backdoors. We conduct extensive physical experiments using diverse deep neural networks across various camera systems in different practical environments, achieving a 92.54% average backdoor success rate. Mengjie Sun, Peizhuo Lv, Shengzhi Zhang, Jianshuo Liu, Kai Chen 0012, Hong Li 0004, Zhi Li 0018, Qinhong Jiang, Limin Sun 0001 |
J. Comput. Secur. | 6 |
| 2026 | FAVDisco: Modeling and Discovering File Access VulnerabilitiesabstractFile access vulnerabilities (FAVs) are one type of security weakness arising from adversary manipulations of file access inputs, posing significant threats to system integrity. Despite their prevalence, FAVs remain underexplored due to limited understanding, complex triggering scenarios, and stealthy and diverse manifestations; these challenges render current detection approaches incomplete and inaccurate. To this end, we conducted an in-depth empirical study across 204 file-related CVEs, uncovering the root cause and trigger mechanisms of FAVs. Based on these findings, we propose an exhaustive accessing model and a specialized threat model that define the adversary and attack surface for FAVs, enabling systematic attribution and analysis of file operations. Furthermore, we propose FAVDisco , a novel framework for discovering FAVs by mutating, triggering, and analyzing file operations. It employs a File Mutator to simulate diverse execution scenarios and an FAV Checker that integrates a model-based adversary controllable checker with pattern-based detection rules to identify FAVs. Implemented on Windows, FAVDisco achieves remarkable performance with 92.1% precision and 83.3% recall on the disclosed FAV detection task, outperforming state-of-the-art methods. Moreover, it uncovers 13 zero-day FAVs in 10 widely used services, with six assigned new CVEs and earning a reward of $29,000 from Microsoft Security Response Center. Beibei Zhao, Wenjie Feng 0001, Qingli Guo, Yingli Sun, Fangming Gu, Xiaorui Gong, Hong Li 0004 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2025 | Microservice Dependency Discovery Based on Spatio-Temporal Network Flow Behavior ModelingabstractThe microservice architectural style offers significant application scalability and development advantages. Composing monolithic systems into loosely coupled, containerized services reduces deployment and development costs while enhancing the flexibility and resilience of the overall system. However, in large-scale internet applications involving multi-party collaboration and global deployment, the formation of complex microservice dependencies increases the risk of cascading failures. It complicates the process of identifying the source of a fault. Identifying such dependencies in uncontrolled conditions with limited observational data represents a significant challenge. This paper proposes Microservice Dependency Discovery based on Spatio-Temporal Network Flow Behaviour Modelling (Cross-MSDD). This method infers microservice dependencies by modeling spatiotemporal interactions of network traffic. The method employs network flow characteristics to mine frequent behavioral patterns, thereby inferring dependency chains without the necessity for additional tracing tools. The method utilizes network flow characteristics to mine frequent behavior patterns, inferring dependency chains without additional tracing tools, minimizing system interruptions, and protecting request content privacy. To verify the effectiveness of the proposed method, a semi-simulated experimental environment was set up using traffic data from typical microservice applications. The results demonstrate that the method attains an accuracy rate exceeding 96.3% in cross-domain dependency identification, markedly surpassing the performance of existing techniques. This facilitates the practical detection of faults and the mitigation of cascading failure risks, thereby ensuring system stability. Jinfa Wang, Chunyang Zheng, Hui Wen 0001, Hong Li 0004, Hongsong Zhu |
CSCWD | 5 |
| 2025 | HF-Mamba: Improving Multimodal Classification via Hierarchical Fusion Based on Mamba
Yimo Ren, Jinfa Wang, Hong Li 0004, Rongrong Xi, Haiqiang Fei, Hongsong Zhu |
DASFAA (2) | 3 |
| 2025 | TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability CodeabstractCode reuse in software development frequently facilitates the spread of vulnerabilities, making the scope of affected software in CVE reports imprecise. Traditional methods primarily focus on identifying reused vulnerability code within target software, yet they cannot verify if these vulnerabilities can be triggered in new software contexts. This limitation often results in false positives. In this paper, we introduce TransferFuzz, a novel vulnerability verification framework, to verify whether vulnerabilities propagated through code reuse can be triggered in new software. Innovatively, we collected runtime information during the execution or fuzzing of the basic binary (the vulnerable binary detailed in CVE reports). This process allowed us to extract historical traces, which proved instrumental in guiding the fuzzing process for the target binary (the new binary that reused the vulnerable function). TransferFuzz introduces a unique Key Bytes Guided Mutation strategy and a Nested Simulated Annealing algorithm, which transfers these historical traces to implement trace-guided fuzzing on the target binary, facilitating the accurate and efficient verification of the propagated vulnerability. Our evaluation, conducted on widely recognized datasets, shows that TransferFuzz can quickly validate vulnerabilities previously unverifiable with existing techniques. Its verification speed is 2.5 to 26.2 times faster than existing methods. Moreover, TransferFuzz has proven its effectiveness by expanding the impacted software scope for 15 vulnerabilities listed in CVE reports, increasing the number of affected binaries from 15 to 53. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/TransferFuzz. Siyuan Li 0014, Yuekang Li, Zuxin Chen, Chaopeng Dong, Yongpan Wang, Hong Li 0004, Yongle Chen, Hongsong Zhu |
ICSE | 6 |
| 2025 | Lares: LLM-driven Code Slice Semantic Search for Patch Presence TestingabstractIn modern software ecosystems, 1-day vulnerabilities pose significant security risks due to extensive code reuse. Identifying vulnerable functions in target binaries alone is insufficient; it is also crucial to determine whether these functions have been patched. Existing methods, however, suffer from limited usability and accuracy. They often depend on the compilation process to extract features, requiring substantial manual effort and failing for certain software. Moreover, they cannot reliably differentiate between code changes caused by patches or compilation variations.To overcome these limitations, we propose Lares, a scalable and accurate method for patch presence testing. Lares introduces Code Slice Semantic Search, which directly extracts features from the patch source code and identifies semantically equivalent code slices in the pseudocode of the target binary. By eliminating the need for the compilation process, Lares improves usability, while leveraging large language models (LLMs) for code analysis and SMT solvers for logical reasoning to enhance accuracy. Experimental results show that Lares achieves superior precision, recall, and usability. Furthermore, it is the first work to evaluate patch presence testing across optimization levels, architectures, and compilers. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/Lares. Siyuan Li 0014, Yaowen Zheng, Hong Li 0004, Jingdong Guo, Chaopeng Dong, Chunpeng Yan, Weijie Wang 0005, Yimo Ren, Limin Sun 0001, Hongsong Zhu |
ASE | 3 |
| 2025 | BinEnhance: An Enhancement Framework Based on External Environment Semantics for Binary Code Search
Yongpan Wang, Hong Li 0004, Xiaojie Zhu, Siyuan Li 0014, Chaopeng Dong, Shouguo Yang, Kangyuan Qin |
NDSS | 2 |
| 2025 | Automated Penetration on Multi-Subnet Environments with Dual-Stage DRL ModelsabstractWith the advent of artificial intelligence techniques, the field of Network Attack Defense (NAD) has witnessed a surge in research efforts towards automating penetration testing (PenTest). Our work presents a dual-stage PenTest model aiming at predicting attack paths in network topology and determining payload for vulnerabilities in hosts with deep reinforcement learning models. While constructing training environments, our approach integrates real-world vulnerability environments with virtual network topologies. This allows the model to take into account the process of vulnerability validation with success rate compared to existing work based on fully virtualized targets, while retaining the efficiency of deployment and training provided by virtualization. And we introduce a method that simulate hierarchical network topology with randomized subnets to simulate complex network environments, challenging the agent to adapt and learn effective policies across diverse configurations of the target networks. Our experiments demonstrate the effectiveness of our model in various network sizes. In addition, the results indicate that our approach not only achieves high performance but also maintains stability under the different success rate of vulnerability exploitation, showcasing the robustness and adaptability. Our work contributes to the advancement of automated PenTest by providing a more generalized and efficient solution. Haoyu Bu, Hui Wen 0001, Hongsong Zhu, Hong Li 0004, Xirui Song, Yimo Ren |
SMC | 4 |
| 2025 | TimeTravel: Real-time Timing Drift Attack on System Time Using Acoustic Waves
Jianshuo Liu, Hong Li 0004, Haining Wang 0001, Mengjie Sun, Hui Wen 0001, Jinfa Wang, Limin Sun 0001 |
USENIX Security Symposium | 2 |
| 2025 | ICSPFuzzer: An Efficient Fuzzing Technique for ICS Protocols
Zhanwei Song, Dongliang Fang, Shunchao Xu, Yaowen Zheng, Hong Li 0004, Shichao Lv, Zhiqiang Shi, Limin Sun 0001 |
WASA (2) | 5 |
| 2025 | EHFC: Enhanced Format Clustering via Pre-Trained Traffic Model
Zhen Wang 0043, Laile Xi, Haiqiang Fei, Hong Li 0004, Hongsong Zhu |
WASA (1) | 6 |
| 2025 | Automated tactics planning for cyber attack and defense based on large language model agents
Yimo Ren, Jinfa Wang, Hui Wen 0001, Hong Li 0004, Hongsong Zhu |
Neural Networks | 5 |
| 2025 | PREXP: Uncovering and Exploiting Security-Sensitive Objects in the Linux KernelabstractSecurity-Sensitive Objects (SSOs) are often critical components in the exploitation of Linux kernel memory corruption vulnerabilities. While existing research has advanced SSOs identification and classification, there remains a significant gap in systematically understanding how these objects can be effectively exploited in real-world security analysis. To address this challenge, we present PREXP, a novel approach to analyzing SSOs exploitability and automating the transformation of Proof-of-Concept (PoC) into exploitable states. Our approach encompasses three key techniques: (1) capability analysis and attribute modeling of vulnerable object (2) extraction and filtering of target SSOs and (3) automatically augmenting PoCs with SSO-specific code to create exploitation capabilities. To evaluate our approach, we tested our prototype on 30 public CVEs, successfully parsing vulnerable object in 22 cases (73.3%) and achieving accurate SSO matches in 18 (60.0%). PREXP outperformed state-of-the-art tools such as SCAVY and AlphaEXP in structure-matching, and enabled the generation of new Control Flow Hijacking Primitives (CFHPs) for 3 previously unexploited vulnerabilities, demonstrating its practical value in real-world exploit development. Zuxin Chen, Yaowen Zheng, Hong Li 0004, Siyuan Li 0014, Weijie Wang 0005, Dongliang Fang, Zhiqiang Shi, Limin Sun 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | EMFuzz: Use Electromagnetic Fuzzing for Automated Attack Surface Assessment of ActuatorsabstractActuators are essential components in cyber-physical systems, enabling system modules to perform diverse and complex tasks. Unfortunately, the pursuit of higher functional complexity often correlates with a broader attack surface in actuators. Thus, an efficient automated attack surface assessment is crucial to avoid cyber incidents in critical infrastructures. Limited by enormous parameter spaces, current methods rely on heuristic tests to evaluate interference potential but cannot thoroughly investigate the full spectrum of potential hidden interference. The observation that similar interference trigger configurations lead to the same impact has motivated us to use machine learning algorithms for understanding different impact samples around decision boundaries. By leveraging generalized knowledge of responses against specific attack scenarios, we aim to improve the efficiency of automated attack surface assessment of electromagnetic interference on new targets. To this end, we introduce EMFuzz, an automated mechanism to fuzz hardware to quantify varying adverse effects. We evaluate EMFuzz on 16 new servos within real-world scenarios, where it achieves an 86% accuracy in classifying different attack vectors. With the same test time, EMFuzz uncovers over twice the effective attack configurations of the baseline, greatly improving assessment efficiency. To further validate its efficacy, we apply EMFuzz to assess the attack surface of a new actuator from a robot transfer unit, and it can successfully reveal three distinct adverse effects. Shiquan Dong, Zhi Li 0018, Jianshuo Liu, Hong Li 0004, Dongliang Fang, Shichao Lv, Haining Wang 0001, Limin Sun 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated VulnerabilityabstractCode reuse in software development frequently facilitates the spread of vulnerabilities, leading to imprecise scopes of affected software in CVE reports. Traditional methods focus primarily on detecting reused vulnerability code in target software but lack the ability to confirm whether these vulnerabilities can be triggered in new software contexts. In previous work, we introduced the TransferFuzz framework to address this gap by using historical trace-based fuzzing. However, its effectiveness is constrained by the need for manual intervention and reliance on source code instrumentation. To overcome these limitations, we propose TransferFuzz-Pro, a novel framework that integrates Large Language Model (LLM)-driven code debugging technology. By leveraging LLM for automated, human-like debugging and Proof-of-Concept (PoC) generation, combined with binary-level instrumentation, TransferFuzz-Pro extends verification capabilities to a wider range of targets. Our evaluation shows that TransferFuzz-Pro is significantly faster and can automatically validate vulnerabilities that were previously unverifiable using conventional methods. Notably, it expands the number of affected software instances for 15 CVE-listed vulnerabilities from 15 to 53 and successfully generates PoCs for various Linux distributions. These results demonstrate that TransferFuzz-Pro effectively verifies vulnerabilities introduced by code reuse in target software and automatically generation PoCs. Siyuan Li 0014, Kaiyu Xie, Yuekang Li, Hong Li 0004, Yimo Ren, Limin Sun 0001, Hongsong Zhu |
IEEE Trans. Software Eng. | 4 |
| 2024 | Hierarchical Aligned Multimodal Learning for NER on Tweet PostsabstractMining structured knowledge from tweets using named entity recognition (NER) can be beneficial for many downstream applications such as recommendation and intention under standing. With tweet posts tending to be multimodal, multimodal named entity recognition (MNER) has attracted more attention. In this paper, we propose a novel approach, which can dynamically align the image and text sequence and achieve the multi-level cross-modal learning to augment textual word representation for MNER improvement. To be specific, our framework can be split into three main stages: the first stage focuses on intra-modality representation learning to derive the implicit global and local knowledge of each modality, the second evaluates the relevance between the text and its accompanying image and integrates different grained visual information based on the relevance, the third enforces semantic refinement via iterative cross-modal interactions and co-attention. We conduct experiments on two open datasets, and the results and detailed analysis demonstrate the advantage of our model. Hong Li 0004, Yimo Ren, Jie Liu 0079, Shuaizong Si, Hongsong Zhu, Limin Sun 0001 |
AAAI | 2 |
| 2024 | Precise and Efficient Third-party Java Libraries Identification Tool for Collaborative SoftwareabstractCollaborative systems frequently depend on various software components, like third-party libraries (TPLs), to execute their functions and expedite the development of the system. The security of an entire collaboration system can be compromised by a TPL that is vulnerable, particularly in an industrial setting. Unfortunately, current TPL detection tools encounter difficulties in precisely identifying version levels and exhibit inefficiency in detecting TPLs on a large scale.To address these challenges, we recommend JHunter, a precise and efficient tool for detecting TPL version details. Our approach involves introducing a novel concept called the attribute class dependency graph (ACDG) as a feature at the package level for TPLs. We then utilise a graph neural network-based method to compare the similarity of ACDGs and identify a list of candidate TPLs. Later, we use more detailed class-level features, such as Control Flow Graphs (CFGs), and constant features to determine version-specific information. We collected 19,095 different versions of TPLs from Maven to build our feature database. Our analysis demonstrates the effectiveness of JHunter on a real-world dataset, achieving F1 scores of 99.34% and 97.28% at the library and version levels, respectively, surpassing previous state-of-the-art (SOTA) results. Hongtu Zhang, Jingdong Guo, Laile Xi, Sidy Tambadou, Fang Zuo, Hong Li 0004 |
CSCWD | 7 |
| 2024 | A Relation-Aware Heterogeneous Graph Transformer on Dynamic Fusion for Multimodal Classification TasksabstractMultimodal fusion aims to improve the performance of models for applications by extracting and fusing information in different modalities, including texts, images or others. Recent researches have shown that multimodal fusion is beneficial in many multimedia tasks. In this paper, we study typical multimedia classification tasks in social media posts, including sarcasm detection and sentiment analysis. This paper proposes DMF-RHGT-HPA, including dynamic Fusion multimodal fusion(DMF), a relation-aware heterogeneous graph transformer(RHGT) and hierarchical pooling alignment(HPA). To realize better multimodal fusion, the paper designs it on a heterogeneous graph with dynamic links, without any padding of texts or images. To thoroughly learn the multimodal graph and obtain the representation of nodes, the paper proposes a relation-aware heterogeneous graph transformer to fuse the node-level and edge-level features simultaneously. To get a refined representation of the multimodal graph, the paper designs a hierarchical pooling alignment to gather all nodes’ representations well. Experiments conducted on two primary and public datasets from Twitter and Yelp respectively show the ability of DMF-RHGT-HPA to gain the best performance of sarcasm detection and sentiment analysis, outperforming existing state-of-the-art baselines. Yimo Ren, Jinfa Wang, Jie Liu 0079, Hong Li 0004, Hongsong Zhu, Limin Sun 0001 |
ICASSP | 5 |
| 2024 | FirmPorter: Porting RTOSes at the Binary Level for Firmware Re-hosting
Mingfeng Xin, Hui Wen 0001, Liting Deng, Hong Li 0004, Qiang Li 0007, Limin Sun 0001 |
ICICS (2) | 4 |
| 2024 | LibvDiff: Library Version Difference Guided OSS Version Identification in BinariesabstractOpen-source software (OSS) has been extensively employed to expedite software development, inevitably exposing downstream software to the peril of potential vulnerabilities. Precisely identifying the version of OSS not only facilitates the detection of vulnerabilities associated with it but also enables timely alerts upon the release of 1-day vulnerabilities. However, current methods for identifying OSS versions rely heavily on version strings or constant features, which may not be present in compiled OSS binaries or may not be representative when only function code changes are made. As a result, these methods are often imprecise in identifying the version of OSS binaries being used. Chaopeng Dong, Siyuan Li 0014, Shouguo Yang, Yang Xiao 0011, Yongpan Wang, Hong Li 0004, Zhi Li 0018, Limin Sun 0001 |
ICSE | 6 |
| 2024 | NFCEraser: A Security Threat of NFC Message Modification Caused by Quartz Crystal OscillatorabstractNear Field Communication (NFC) has been widely used for rapid data exchange between electronic devices over a very short distance. In this paper, we reveal a new security vulnerability in NFC passive communication channels where transferred data can be modified in real-time. The security threat of data modification posed by this vulnerability is called NFCEraser. Exploiting electromagnetic interference (EMI), NFCEraser injects signals into the crystal oscillator’s electrode and adjusts the amplitude of carrier signals in NFC communication channels. By manipulating the parameters of EMI signals, NFCEraser is able to arbitrarily flip the bits in data payload sent from an NFC peer device, which may cause serious security outcomes. To assess the severity of NFCEraser, we examine six NFC modules under NFC-A/B communication modes and successfully perform reading operations under a variety of data lengths. The experimental results show that NFCEraser can modify data bits in response frames from NFC peer devices with the maximum 89% accuracy, under around 0.21μs latency. Our analysis further shows that NFCEraser can maintain an attack success rate of no less than 85% in environments with typical levels of electromagnetic noise. Jianshuo Liu, Hong Li 0004, Mengjie Sun, Haining Wang 0001, Hui Wen 0001, Zhi Li 0018, Limin Sun 0001 |
SP | 2 |
| 2024 | Review of data security within energy blockchain: A comprehensive analysis of storage, management, and utilizationabstractEnergy systems are currently undergoing a transformation towards new paradigms characterized by decarbonization, decentralization, democratization, and digitalization. In this evolving context, energy blockchain, aiming to enhance efficiency, transparency, and security, emerges as an integrated technological solution designed to address the diverse challenges in this field. Data security is essential for the reliable and efficient functioning of energy blockchain. The pressing need to address challenges related to secure data storage, effective data management, and efficient data utilization is increasingly vital. This paper offers a comprehensive survey of academic discourse on energy blockchain data security over the past five years, adopting an all-encompassing perspective that spans data storage, management, and utilization. Our work systematically evaluates and contrasts the strengths and weaknesses of various research methodologies. Additionally, this paper proposes an integrated hierarchical on-chain and off-chain security energy blockchain architecture, specifically designed to meet the complex security requirements of multi-blockchain business environments. Concludingly, this paper identifies key directions for future research, particularly in advancing the integration of storage, management, and utilization of energy blockchain data security. Yunhua He, Zhihao Zhou 0001, Fahui Chong, Bin Wu 0011, Ke Xiao 0001, Hong Li 0004 |
High Confid. Comput. | 7 |
| 2024 | A verifiable and efficient cross-chain calculation model for charging pile reputationabstractTo solve the current situation of low vehicle-to-pile ratio, charging pile(CP) operators incorporate private CPs into the shared charging system. However, the introduction of private CP has brought about the problem of poor service quality. Reputation is a common service evaluation scheme, in which the third-party reputation scheme has the issue of single point of failure; although the blockchain-based reputation scheme solves the single point of failure issue, it also brings the challenges of storage and query efficiency. It is a feasible solution to classify and store information on multiple chains, and at this time, reputation needs to be calculated in a cross-chain mode. Crosschain reputation calculation faces the problems of correctness verification, integrity verification and efficiency. Therefore, this paper proposes a verifiable and efficient cross-chain calculation model for CP reputation. Specially, in this model, we propose a verifiable cross-chain contract calculation scheme that adopts polynomial commitment to solve the problems of polynomial damage and tampering that may be encountered in the crosschain process of outsourced polynomials, so as to ensure the integrity and correctness of polynomial calculations. In addition, the miner selection and incentive mechanism algorithm in this scheme ensures the correctness of extracted information when the outsourced polynomial is calculated on the blockchain. The security analysis and experimental results demonstrate that this scheme is feasible in practice. Yunhua He, Bin Wu 0011, Ke Xiao 0001, Hong Li 0004 |
High Confid. Comput. | 6 |
| 2024 | Multi-granularity cross-modal representation learning for named entity recognition on social media
Gaosheng Wang, Hong Li 0004, Jie Liu 0079, Yimo Ren, Hongsong Zhu, Limin Sun 0001 |
Inf. Process. Manag. | 3 |
| 2024 | Towards Unconstrained Vocabulary Eavesdropping With mmWave Radar Using GANabstractAs acoustic communication systems become increasingly common in our daily life, eavesdropping brings severe security and privacy risks. Current methods of acoustic eavesdropping either provide low resolution due to the use of sub-6 GHz frequencies, work only for limited words based on classification approaches, or cannot work through-wall because of the use of optical sensors. In this article, we presentmilliEar, a mmWave acoustic eavesdropping system that leverages the high-resolution of mmWave FMCW ranging and generative machine learning models to not only extract vibrations but to reconstruct the audio.milliEarcombines speaker vibration estimation with conditional generative adversarial networks to eavesdrop and recover high-quality audios (i.e., with no vocabulary constraints). We implement and evaluatemilliEarusing off-the-shelf mmWave radars deployed in different scenarios and settings. Evaluation results clearly show thatmilliEarcan accurately reconstruct the audio even at different distances, angles, and through the wall with different insulator materials. In addition, our subjective and objective evaluations demonstrate that the reconstructed audio has a strong similarity with the original audio. Pengfei Hu 0001, Wenhao Li 0008, Panneer Selvam Santhalingam, Parth H. Pathak, Hong Li 0004, Huanle Zhang, Xiuzhen Cheng, Prasant Mohapatra |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | FeaShare: Feature Sharing for Computation Correctness in Edge PreprocessingabstractEdge preprocessing is a critical service type in edge computing. However, untrusted edges may be malicious to provide incorrect computational results (i.e., edge tampering). Although some studies have considered the correctness of results, they have limitations when applied to edge preprocessing. We present FeaShare, a feature-sharing approach, to verify edge results. The process is integrated into normal service operations. Meanwhile, to overcome feature-based limitations, terminals obtain partial edge results for a set of data by executing a small number of computations. These partial results are leveraged to construct shared features, facilitating the detection of edge tampering even when the tampered portion is not directly related to the features. Subsequently, the shared features are mapped to pseudo-data and added to the terminal's data sequence, preventing features from influencing the results of terminal data. To resist edge attacks, both feature construction and placement are time-dependent and dynamic. FeaShare is not confined to specific edge tasks. We evaluate FeaShare using 3 typical scenes encompassing 5 applications. For instance, the evaluation utilizing the VGG model and CIFAR-10 dataset demonstrates a detection rate of 97%. Terminals perform approximately 10% of the edge's computation operations, and its overhead growth rate is less than 10%. Haoyu Bin, Hong Li 0004, Hongsong Zhu, Limin Sun 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | LibAM: An Area Matching Framework for Detecting Third-Party Libraries in BinariesabstractThird-party libraries (TPLs) are extensively utilized by developers to expedite the software development process and incorporate external functionalities. Nevertheless, insecure TPL reuse can lead to significant security risks. Existing methods, which involve extracting strings or conducting function matching, are employed to determine the presence of TPL code in the target binary. However, these methods often yield unsatisfactory results due to the recurrence of strings and the presence of numerous similar non-homologous functions. Furthermore, the variation in C/C++ binaries across different optimization options and architectures exacerbates the problem. Additionally, existing approaches struggle to identify specific pieces of reused code in the target binary, complicating the detection of complex reuse relationships and impeding downstream tasks. And, we call this issue the poor interpretability of TPL detection results. In this article, we observe that TPL reuse typically involves not just isolated functions but also areas encompassing several adjacent functions on the Function Call Graph (FCG). We introduce LibAM, a novel Area Matching framework that connects isolated functions into function areas on FCG and detects TPLs by comparing the similarity of these function areas, significantly mitigating the impact of different optimization options and architectures. Furthermore, LibAM is the first approach capable of detecting the exact reuse areas on FCG and offering substantial benefits for downstream tasks. To validate our approach, we compile the first TPL detection dataset for C/C++ binaries across various optimization options and architectures. Experimental results demonstrate that LibAM outperforms all existing TPL detection methods and provides interpretable evidence for TPL detection results by identifying exact reuse areas. We also evaluate LibAM’s scalability on large-scale, real-world binaries in IoT firmware and generate a list of potential vulnerabilities for these devices. Our experiments indicate that the Area Matching framework performs exceptionally well in the TPL detection task and holds promise for other binary similarity analysis tasks. Last but not least, by analyzing the detection results of IoT firmware, we make several interesting findings, for instance, different target binaries always tend to reuse the same code area of TPL. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/LibAM . Siyuan Li 0014, Yongpan Wang, Chaopeng Dong, Shouguo Yang, Hong Li 0004, Hao Sun 0028, Zhe Lang, Zuxin Chen, Weijie Wang 0005, Hongsong Zhu, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2024 | Battling against Protocol Fuzzing: Protecting Networked Embedded Devices from Dynamic FuzzersabstractN etworked E mbedded D evices (NEDs) are increasingly targeted by cyberattacks, mainly due to their widespread use in our daily lives. Vulnerabilities in NEDs are the root causes of these cyberattacks. Although deployed NEDs go through thorough code audits, there can still be considerable exploitable vulnerabilities. Existing mitigation measures like code encryption and obfuscation adopted by vendors can resist static analysis on deployed NEDs, but are ineffective against protocol fuzzing. Attackers can easily apply protocol fuzzing to discover vulnerabilities and compromise deployed NEDs. Unfortunately, prior anti-fuzzing techniques are impractical as they significantly slow down NEDs, hampering NED availability. To address this issue, we propose Armor—the first anti-fuzzing technique specifically designed for NEDs. First, we design three adversarial primitives–delay, fake coverage, and forged exception–to break the fundamental mechanisms on which fuzzing relies to effectively find vulnerabilities. Second, based on our observation that inputs from normal users consistent with the protocol specification and certain program paths are rarely executed with normal inputs, we design static and dynamic strategies to decide whether to activate the adversarial primitives. Extensive evaluations show that Armor incurs negligible time overhead and effectively reduces the code coverage (e.g., line coverage by 22%-61%) for fuzzing, significantly outperforming the state of the art. Puzhuo Liu, Yaowen Zheng, Chengnian Sun, Hong Li 0004, Zhi Li 0018, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Intrusion Detection Based on Sampling and Improved OVA Technique on Imbalanced DataabstractNetwork-based Intrusion Detection(NID) is an effective means to deal with network attacks. NID is able to detect different types of network attacks by analyzing network traffic. However, in the real world, network traffic contains majority and minority class attacks as well as a large number of normal traffic samples. The imbalance in the number of training samples of various types of network traffic makes network intrusion detection very poor. Due to the lack of training samples, traditional NID can’t learn the characteristics of minority class attacks, which leads to the failure of NID to detect minority class attacks. Therefore, in order to solve the problem brought by imbalanced data, we propose a network intrusion detection algorithm based on the sampling and improved One-vs-All(OVA) technique. The dataset is balanced by downsampling the majority class data based on K-means clustering and oversampling the minority class data based on Auxiliary Classifier Generative Adversarial Network(ACGAN), improve classification accuracy through OVA-based model training and testing. We conduct validation experiments on the NSL-KDD dataset, and the experimental results show that the proposed method achieves excellent results in terms of Accuracy, Precision, Recall and F1-score. Compared with existing state-of-the-art methods, the proposed method not only achieves excellent detection performance with low false positive rate, but also addresses the learning problem of imbalanced data more effectively. Yongfei Liu, Hong Li 0004, Wenyuan Zhang 0002, Fei Lyu 0001, Shuaizong Si |
CSCWD | 2 |
| 2023 | Improving the Modality Representation with multi-view Contrastive Learning for Multimodal Sentiment AnalysisabstractModality representation learning is an important problem for multimodal sentiment analysis (MSA), since the highly distinguishable representations can contribute to improving the analysis effect. Previous works of MSA have usually focused on internal fusion strategies for different modalities within one sample, and the external usage of cross reference relations among different samples was given less attention. Recently, the rise of contrastive learning provides powerful clues for us to learn modal representation with stronger discriminative ability. In this study, we explore the approach of representations improvement and devise a three-stages framework with multi-view contrastive learning to refine representations for the specific objectives. Firstly, for each modality, we employ the supervised contrastive learning to pull samples within the same class together while the other samples are pushed apart. Then, a self-supervised contrastive learning is designed for the distilled cross-modal representations after a novel Transformer-based interaction module. At last, we leverage again the supervised contrastive learning to enhance the fused multimodal representation. We conduct extensive experiments on three open datasets, and results show the advance of our model. Hong Li 0004, Jie Liu 0079, Yimo Ren, Hongsong Zhu, Limin Sun 0001 |
ICASSP | 3 |
| 2023 | FlowEmbed: Binary function embedding model based on relational control flow graph and byte sequenceabstractBinary function embedding models are applicable to various downstream tasks within IoT device software systems and have demonstrated advantages in numerous binary analysis tasks, such as vulnerability (homologous) function search and compilation optimization option identification. However, current binary function embedding methods either learn embedding based on code sequence, which lack the program semantics of functions (e.g., control flow, etc.) or based on program structure graphs, which omit global sequential information. As a result, these methods fall short in enabling models to learn the complete semantic of function. In this paper, we introduce FlowEmbed, a novel approach that synergistically integrates control flow and global semantic learning to facilitate exhaustive code comprehension. Initially, FlowEmbed harnesses a distinct relational control flow graph combined with the power of BERT and RGCN models to aptly capture the nuances of control flow semantics. Moreover, by deploying the DPCNN model on a byte sequence constructed from function machine code, FlowEmbed adeptly discerns the inherent global sequential semantics of binary functions. Through rigorous evaluations spanning three IoT-related tasks, FlowEmbed’s efficacy becomes evident, showcasing notable improvements: a 20.6% improvement in compilation optimization option identification, a 1.8% improvement in binary function similarity analysis, and an 11.9% improvement in homologous function search. Collectively, these results underscore FlowEmbed’s superior capability, positioning it as a invaluable asset in a binary analysis application. Yongpan Wang, Chaopeng Dong, Siyuan Li 0014, Renjie Su, Zhanwei Song, Hong Li 0004 |
ICPADS | 7 |
| 2023 | CEntRE: A paragraph-level Chinese dataset for Relation Extraction among EnterprisesabstractEnterprise relation extraction aims to detect pairs of enterprise entities and identify the business relations between them from unstructured or semi-structured text data, and it is crucial for several real-world applications such as risk analysis, rating research and supply chain security. However, previous work mainly focuses on getting attribute information about enterprises like personnel and corporate business, and pays little attention to enterprise relation extraction. To encourage further progress in the research, we introduce the CEntRE, a new dataset constructed from publicly available business news data with careful human annotation and intelligent data processing. Moreover, we propose a joint entity and relation extraction network, which is capable of discovering enterprise entities and extracting business relations between them accurately. The network firstly encodes input sequences with strong semantic augmentation to learn contextual representation for each token, then a conditional random field (CRF) module is used for entity extraction. Subsequently, entity pairs are built and a new encoder based on the entity pairs is applied to get global information for relation extraction. Finally, a biaffine classifier is deployed to classify the relations. Extensive experiments on CEntRE demonstrate the effectiveness of our proposed method compared with other six excellent models, and thus our model can be considered as one strong baseline. The data and code are available at: https://github.com/LiuPeiP-CStMining_Entity_Relations_Among_Enterprises Hong Li 0004, Yimo Ren, Jie Liu 0079, Fei Lyu 0001, Hongsong Zhu, Limin Sun 0001 |
IJCNN | 2 |
| 2023 | User Recognition of Devices on the Internet based on Heterogeneous Graph Transformer with Partial LabelsabstractRecognizing the users of devices can easily enable numerous security applications. Due to the lot's kinds of device data and a large number of missing values, it takes work to recognize the users of devices well. The community detection methods based on Graph Neural Networks (GNN) can integrate multi-source data well and cluster devices into communities with the same users. While existing GNN methods face several issues. The methods on homogeneous graphs could not utilize the multi-source data of devices, and most methods on heterogeneous graphs need specific knowledge to design meta paths. Also, the Internet-scale data of devices make it hard to learn the representation thoroughly. Further, most methods need to consider the known partial labels in the early stage of the training process. To improve the performance of user recognition, this paper proposes HGT-PL, namely a Heterogeneous Graph Transformer with Partial Labels, to calculate the representation of devices on the Internet. Then cluster methods are used to realize user recognition. Using graph transformers, HGT-PL deeply learns node features and graph structure on the heterogeneous graph of devices. By Label Encoder, HGT-PL fully utilizes the users of partial devices from preliminary rules with high confidence. Moreover, cluster methods carefully divide and modify the communities with different users. The paper conducts experiments on the web-scale data collected from the Internet. The results show that HGT-PL can recognize users of devices more accurately and effectively, with 0.5121 NMI and 0.3554 ARI, compared with existing GNN methods. Yimo Ren, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001 |
IJCNN | 3 |
| 2023 | Detecting Vulnerabilities in Linux-Based Embedded Firmware with SSE-Based On-Demand Alias AnalysisabstractAlthough the importance of using static taint analysis to detect taint-style vulnerabilities in Linux-based embedded firmware is widely recognized, existing approaches are plagued by following major limitations: (a) Existing works cannot properly handle indirect call on the path from attacker-controlled sources to security-sensitive sinks, resulting in lots of false negatives. (b) They employ heuristics to identify mediate taint source and it is not accurate enough, which leads to high false positives. Yaowen Zheng, Le Guan, Peng Liu 0005, Hong Li 0004, Hongsong Zhu, Kejiang Ye, Limin Sun 0001 |
ISSTA | 6 |
| 2023 | DeviceGPT: A Generative Pre-Training Transformer on the Heterogenous Graph for Internet of ThingsabstractRecently, Graph neural networks (GNNs) have been adopted to model a wide range of structured data from academic and industry fields. With the rapid development of Internet technology, there are more and more meaningful applications for Internet devices, including device identification, geolocation and others, whose performance needs improvement. To replicate the several claimed successes of GNNs, this paper proposes DeviceGPT based on a generative pre-training transformer on a heterogeneous graph via self-supervised learning to learn interactions-rich information of devices from its large-scale databases well. The experiments on the dataset constructed from the real world show DeviceGPT could achieve competitive results in multiple Internet applications. Yimo Ren, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001 |
SIGIR | 3 |
| 2023 | Enimanal: Augmented cross-architecture IoT malware analysis using graph neural networks
Liting Deng, Hui Wen 0001, Mingfeng Xin, Hong Li 0004, Zhiwen Pan, Limin Sun 0001 |
Comput. Secur. | 4 |
| 2023 | CL-GAN: A GAN-based continual learning model for generating and detecting AGDs
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001 |
Comput. Secur. | 2 |
| 2023 | Owner name entity recognition in websites based on multiscale features and multimodal co-attention
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Multiview Embedding with Partial Labels to Recognize Users of Devices Based on Unified TransformerabstractRecognizing the users of devices (or clusters of devices) who use IP addresses as unique identities on the Internet can easily enable numerous security applications. Fast and accurate user recognition is critical for supervisors to find influenced organizations connected to their networks in light of new security threats. Many users’ information scatters in the multisource data of IP addresses. Up until now, user recognition of devices has had two main problems. On the one hand, existing methods could not fully use multisource data of the IP addresses and wastes the valuable information of labels. On the other hand, only a tiny portion of devices can be tagged with highly confident known users manually, making it an urgent need to infer unknown users of devices. So, the problem of user recognition on devices is to guess the unknown user with multisource data and existing devices with known users. Therefore, this paper proposes a multiview fusion method to deal with multisource data from devices with a small number of manually labelled samples. The paper uses GraphSAGE to obtain an exemplary representation of IP addresses and designs a label encoder to fully use a small number of devices with known users. Then, the paper builds a specific unified transformer to achieve high performance to determine whether two devices have the same user. At the same time, the paper conducts real‐world experiments and finds that the proposed method can achieve 0.9158 accuracy and 0.6131 F1 to find devices with the same users on the constructed dataset in the real world. Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001 |
Int. J. Intell. Syst. | 2 |
| 2023 | Cross-Chain Trusted Service Quality Computing Scheme for Multichain-Model-Based 5G Network Slicing SLAabstractAs a key technology for the development of 5G networks, network slicing is developing rapidly. Although network slicing can realize the flexible division of 5G network resources and quickly customize virtual networks that meet the differentiated needs of customers, it is still difficult to determine the optimal service quality parameters in application scenarios. To solve the problem, this article designs a multichain 5G network slicing service quality computing model to calculate the service quality parameters of the network slicing. The calculated service quality parameters can be used as an adjustment basis in the negotiation of the SLA between the customer and the network operator. However, the traditional method of calculating information across chains will cause frequent information interactions and affect efficiency. Therefore, in this scheme, we deploy a smart contract on each blockchain to calculate the information, which can reduce the frequency of information transmission and improve efficiency. In addition, in order to make the calculation between smart contracts more fluent and the requirements more relevant, this article proposes to coordinate the development of smart contracts through multiple blockchains. Besides, to ensure the cross-chain security calculation, the signature by Cosi protocol and multisigncryption algorithms are used in the transmission of nonprivate information and private information in the cross-chain process, respectively. Security analysis and experimental results prove that the multichain 5G network slicing service quality computing model is feasible and efficient in practice. Yunhua He, Bin Wu 0011, Yigang Yang, Ke Xiao 0001, Hong Li 0004 |
IEEE Internet Things J. | 6 |
| 2023 | DevTag: A Benchmark for Fingerprinting IoT DevicesabstractNowadays, various Internet of Things (IoT) devices, such as routers, webcams, and network printers, have been deployed across the Internet. For security and management purposes, it is important to accurately fingerprint IoT devices. In this work, we build a first benchmark called DevTag (IoT Device Tagging) for fingerprinting IoT devices. Specifically, DevTag supports retrieving packet-level features from IoT devices through two different data collections, passive monitoring, and active probing. For detecting IoT devices, DevTag integrates model-based and rule-based fingerprinting methods. For the model-based detection, we reimplemented five typical deep algorithms to infer IoT device classification models. For the rule-based detection, we generated nearly 41 117 rules in a unified format by analyzing several open-source tools. Furthermore, we conducted a systematic analysis to explore the advantages and limitations of those two methods for detecting IoT devices. Our analysis results reveal that the model-based detection has a significant advantage in distinguishing coarse-grained IoT devices (e.g., device type and vendor), while it is not suitable to detect product information as the label amount is massive. The rule-based detection is capable of extracting fine-grained device information with high precision in a short time. However, rules also suffer several inherent problems, such as multiple matching, conflicting, and overlapping issues. Finally, we implemented and distributed a prototype of DevTag working as the first benchmark for detecting IoT devices in the network community. Shangfeng Wan, Qiang Li 0007, Haining Wang 0001, Hong Li 0004, Limin Sun 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Owner name entity recognition in websites based on heterogeneous and dynamic graph transformer
Yimo Ren, Hong Li 0004, Jie Liu 0079, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001 |
Knowl. Inf. Syst. | 2 |
| 2023 | Model Poisoning Attack on Neural Network Without Reference DataabstractDue to the substantial computational cost of neural network training, adopting third-party models has become increasingly popular. However, recent works demonstrate that third-party models can be poisoned. Nonetheless, most model poisoning attacks require reference data, e.g., training dataset or data belonging to the target label, making them difficult to launch in practice. In this paper, we propose a reference data independent model poisoning attack that can (1) directly search for sensitive features with respect to the target label, (2) quantify the positive and negative effects of the model parameters on sensitive features, and (3) accomplish the training of poisoned model by our parameter selective update strategy. The extensive evaluation on datasets with a few classes and numerous classes show that the attack is (I) effective: the trigger input can be labeled as a deliberate class by the poisoned model with high probability; (II) covert: the performance of the poisoned model is almost indistinguishable from the intact model on non-trigger inputs; and (III) straightforward: an adversary only needs a little background knowledge to launch the attack. Overall, the evaluation results show that our attack achieves 95%, 100%, 81%, 96%, and 96% success rates on Cifar10, Cifar100, ISIC2018, FaceScrub, and ImageNet datasets, respectively. Xianglong Zhang, Huanle Zhang, Hong Li 0004, Dongxiao Yu, Xiuzhen Cheng, Pengfei Hu 0001 |
IEEE Trans. Computers | 4 |
| 2023 | Internet-Scale Fingerprinting the Reusing and Rebranding IoT Devices in the CyberspaceabstractFingerprinting Internet-of-Things(IoT) devices on types and brands is a necessary work for security analysis in the cyberspace. The existing approaches mainly rely on the dominant features of devices which is response to information in order to identify these online devices. However, the web server components reusing and products rebranding are the common phenomenons of these embedded IoT devices. It caused the existing approaches difficult to identify most devices even errors due to the similar responses. In this paper, we present an approach, IoTXray, which improves the work efficiently of information collection about accelerating the relations between reusing/rebranding devices with the corresponding manufacturers. And these relations can generate more accurate and reliable fingerprints than previous approaches. Using the mixed neural networks, IoTXray comprehensively detects the real manufactures of online IoT devices upon three different kinds of data sources. In the experiment, our approach can identify 7,025,854 IoT devices on HTTP-hosts. The identification rate has reached to several times higher than previous approaches. Our approach has especially detected 3,268,953 reusing and 963,653 rebranding devices with their original manufacturers. Zhaoteng Yan, Zhi Li 0018, Hong Li 0004, Shouguo Yang, Hongsong Zhu, Limin Sun 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Towards Practical Binary Code Similarity Detection: Vulnerability Verification via Patch Semantic AnalysisabstractVulnerability is a major threat to software security. It has been proven that binary code similarity detection approaches are efficient to search for recurring vulnerabilities introduced by code sharing in binary software. However, these approaches suffer from high false-positive rates (FPRs) since they usually take the patched functions as vulnerable, and they usually do not work well when binaries are compiled with different compilation settings. To this end, we propose an approach, named Robin , to confirm recurring vulnerabilities by filtering out patched functions. Robin is powered by a lightweight symbolic execution to solve the set of function inputs that can lead to the vulnerability-related code. It then executes the target functions with the same inputs to capture the vulnerable or patched behaviors for patched function filtration. Experimental results show that Robin achieves high accuracy for patch detection across different compilers and compiler optimization levels respectively on 287 real-world vulnerabilities of 10 different software. Based on accurate patch detection, Robin significantly reduces the false-positive rate of state-of-the-art vulnerability detection tools (by 94.3% on average), making them more practical. Robin additionally detects 12 new potentially vulnerable functions. Shouguo Yang, Zhengzi Xu, Yang Xiao 0011, Zhe Lang, Yang Liu 0003, Zhiqiang Shi, Hong Li 0004, Limin Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2022 | Multi-features based Semantic Augmentation Networks for Named Entity Recognition in Threat IntelligenceabstractExtracting cybersecurity entities such as attackers and vulnerabilities from unstructured network texts is an important part of security analysis. However, the sparsity of intelligence data resulted from the higher frequency variations and the randomness of cybersecurity entity names makes it difficult for current methods to perform well in extracting security-related concepts and entities. To this end, we propose a semantic augmentation method which incorporates different linguistic features to enrich the representation of input tokens to detect and classify the cybersecurity names over unstructured text. In particular, we encode and aggregate the constituent feature, morphological feature and part of speech feature for each input token to improve the robustness of the method. More than that, a token gets augmented semantic information from its most similar K words in cybersecurity domain corpus where an attentive module is leveraged to weigh differences of the words, and from contextual clues based on a large-scale general field corpus. We have conducted experiments on the cybersecurity datasets DNRTI and MalwareTextDB, and the results demonstrate the effectiveness of the proposed method. Hong Li 0004, Zuoguang Wang, Jie Liu 0079, Yimo Ren, Hongsong Zhu |
ICPR | 2 |
| 2022 | An Evolutionary Learning Approach Towards the Open Challenge of IoT Device Identification
Jingfei Bian, Hong Li 0004, Hongsong Zhu, Limin Sun 0001 |
SecureComm | 3 |
| 2022 | Detection and Incentive: A Tampering Detection Mechanism for Object Detection in Edge ComputingabstractThe object detection tasks based on edge computing have received great attention. A common concern hasn't been addressed is that edge may be unreliable and uploads the incorrect data to cloud. Existing works focus on the consistency of the transmitted data by edge. However, in cases when the inputs and the outputs are inherently different, the authenticity of data processing has not been addressed. In this paper, we first simply model the tampering detection. Then, bases on the feature insertion and game theory, the tampering detection and economic incentives mechanism (TDEI) is proposed. In tampering detection, terminal negotiates a set of features with cloud and inserts them into the raw data, after the cloud determines whether the results from edge contain the relevant information. The honesty incentives employs game theory to instill the distrust among different edges, preventing them from colluding and thwarting the tampering detection. Meanwhile, the subjectivity of nodes is also considered. TDEI distributes the tampering detection to all edges and realizes the self-detection of edge results. Experimental results based on the KITTI dataset, show that the accuracy of detection is 95% and 80%, when terminal's additional overhead is smaller than 30% for image and 20% for video, respectively. The interference ratios of TDEI to raw data are about 16% for video and 0% for image, respectively. Finally, we discuss the advantage and scalability of TDEI. Yicheng Zeng, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001 |
SRDS | 4 |
| 2022 | Inferring Device Interactions for Attack Path Discovery in Smart Home IoT
Mengjie Sun, Ke Li 0042, Yaowen Zheng, Hong Li 0004, Limin Sun 0001 |
WASA (1) | 5 |
| 2022 | Gradient-Based Adversarial Attacks Against Malware Detection by Instruction Replacement
Jiapeng Zhao, Zhongjin Liu, Xiaoling Zhang 0009, Zhiqiang Shi, Shichao Lv, Hong Li 0004, Limin Sun 0001 |
WASA (1) | 7 |
| 2022 | Joint Classification of IoT Devices and Relations in the Internet with Network TrafficabstractWith the rapid growth and popularization of Internet of Things (IoT), more and more devices are deployed in homes, enterprises, cities, etc. The existed methods to classify types and relations of devices are usually two separate tasks. So, it is difficult to quickly provide attributes of devices in the smart network for operators at the same time. At this situation, the paper presents a framework JCIDR for Joint Classification of IoT Device and Relations In the Internet with Network Traffic. By fusing the numerical features and binary image features of traffic, the devices and relations of devices can be recognized simultaneously. The experiment is carried out in a real IoT environment and the accuracy of JCIDR is over 86% with about half time reduction. Therefore, JCIDR could provide operators with a fast, easy, low-cost network device monitoring method without professional equipment or protocols. Yimo Ren, Hong Li 0004, Shuqin Zhang, Hongsong Zhu, Limin Sun 0001 |
WCNC | 2 |
| 2022 | Reinforcement learning based adversarial malware example generation against black-box detectors
Fangtian Zhong, Pengfei Hu 0001, Hong Li 0004, Xiuzhen Cheng |
Comput. Secur. | 4 |
| 2022 | A Cross-Chain Trusted Reputation Scheme for a Shared Charging Platform Based on BlockchainabstractWith the development of electric vehicles, the shortage of charging piles (CPs) has gradually been exposed. In response to this situation, CP operators have taken private CPs into the shared charging system. Due to the lack of maintenance personnel for private CPs that join shared charging, users often face the problems of damaged CPs and poor service attitudes of CP owners. Reputation solutions based on third-party platforms face a problem of single-point failures and reputation solutions based on blockchain face problems of storage and query efficiency. To improve storage and query efficiency, this article proposes a multichain charging model that stores different types of information on different blockchains. However, it faces the problem of unreliable information called across chains, when calculating reputation across chains. Therefore, this article proposes a cross-chain trusted smart contract ($C_{2}T$smart contract) to ensure the authenticity, real-time, and interchain write mutual exclusion of cross-chain information, making reputation calculation in the multichain charging model more convenient and more accurate. Especially, we propose a data mutual trust mechanism based on Merkle proof to ensure the authenticity of cross-chain information and prevent forged information from participating in calculating reputation. Furthermore, we present a data structure composed of multiple counting Bloom filters (MCBFs) to verify the real time of information and filter out non-real-time information, thereby ensuring the real time of the calculated reputation. In addition, we put forward an algorithm to guarantee the interchain write mutual exclusion by hash mutexes, making the reputation calculation process more accurate and complete. The security analysis and experimental results demonstrate that$C_{2}T$smart contract is feasible in practice. Yunhua He, Bin Wu 0011, Yigang Yang, Ke Xiao 0001, Hong Li 0004 |
IEEE Internet Things J. | 6 |
| 2022 | ShadowPLCs: A Novel Scheme for Remote Detection of Industrial Process Control AttacksabstractIndustrial Control System (ICS) security has become increasingly important as attacks targeting ICSs are more prominent. Although many off-the-shelf industrial network intrusion detection mechanisms have been presented in the past, attackers have always found unique disguisable ways to bypass detections and disrupt actual industrial control processes. To mitigate this deficiency, we present a novel scheme for the detection of industrial process control attacks, calledShadowPLCs. Specifically, the scheme first automatically analyzes the PLC control code, then extracts key parameters of the PLCs including valid register addresses, valid range of values, and control logic rules as a basis for evaluating attacks. The attack behavior is detected in real-time from different perspectives through active communication with PLCs and passive monitoring of the network traffic. We implemented a prototype system with Siemens S7-300 series PLCs as a case study. Our scheme was evaluated using two Siemens S7-300 PLCs deployed on a gas pipeline network platform. Experiments demonstrate that the presented scheme can accurately detect process control attacks in real-time without affecting the normal operations of PLCs. Compared with the other four representative detection models, our scheme has better detection performance with detection accuracy of 97.3 percent. Junjiao Liu, Xiaodong Lin 0001, Xin Chen 0123, Hui Wen 0001, Hong Li 0004, Zhiqiang Shi, Limin Sun 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2021 | A Robust IoT Device Identification Method with Unknown Traffic Detection
Xiao Hu 0004, Hong Li 0004, Zhiqiang Shi, Hongsong Zhu, Limin Sun 0001 |
WASA (1) | 2 |
| 2021 | FIUD: A Framework to Identify Users of Devices
Yimo Ren, Hong Li 0004, Hongsong Zhu, Limin Sun 0001 |
WASA (2) | 2 |
| 2021 | A trusted architecture for EV shared charging based on blockchain technologyabstractWith the development of the Energy Internet and the support of the subsidy policies of various countries, Electric Vehicles(EVs) have ushered in a golden development period. However, the development of EVs needs to solve the problems of insufficient charging piles(CPs) and difficulty in finding CPs. In order to solve the problem of difficult charging of EVs, the concept of shared charging came into being, in which idle CPs or private CPs are shared to meet the charging needs of more people and improve the utilization rate of CPs. However, the shared charging scheme implemented by third-party platforms faces the issue of trust lacking. This paper proposes a blockchain architecture for shared charging, which can use the blockchain to build a trust environment involving private pile owners, charging pile(CP) operators, Electric Vehicle(EV) users, etc.. The blockchain architecture also contains the block structure where pointer was added for quick search, contract content that can automatically execute multi-party contracts to achieve secure computing and reputation-based incentive mechanism to provide high-quality charging services in detail. This architecture establishes the multi-party trust environment for shared charging from three aspects: secure storage, secure computing, and secure incentives. Yunhua He, Bin Wu 0011, Ziye Geng, Ke Xiao 0001, Hong Li 0004 |
High Confid. Comput. | 6 |
| 2020 | VES: A Component Version Extracting System for Large-Scale IoT Firmwares
Xulun Hu, Hong Li 0004, Zhaoteng Yan, Limin Sun 0001 |
WASA (2) | 3 |
| 2020 | EdgeCC: An Authentication Framework for the Fast Migration of Edge Services Under Mobile Clients
Hongsong Zhu, Hong Li 0004, Limin Sun 0001 |
WASA (1) | 4 |
| 2020 | Detecting stealthy attacks on industrial control systems using a permutation entropy-based method
Hong Li 0004, Tom H. Luan, An Yang, Limin Sun 0001, Rui Wang 0079 |
Future Gener. Comput. Syst. | 2 |
| 2019 | An Anonymous Blockchain-Based Logging System for Cloud Computing
Ji-Yao Liu, Yunhua He, Chao Wang 0061, Hong Li 0004, Limin Sun 0001 |
BlockSys | 5 |
| 2019 | SCTM: A Multi-View Detecting Approach Against Industrial Control Systems AttacksabstractOff-the-shelf machine learning based intrusion detection systems (IDS) have proved not suitable for protecting industrial control systems (ICS), as they do not consider cooperative regularities between controllers of control loops, and the serious shortage of attacking training sets. We study the consensus and complementary (2C) features which are widely observed in control loops. Subsequently, a multi-view learning framework is proposed to boost the effectiveness of detecting attacks on ICS by using a large number of unlabeled examples with 2C features. Comprehensive attacks of ICS are designed and implemented on a physical testbed, and the experimental data are collected from the historical sequences and IDS alerts. The experimental results demonstrate that the framework is highly adaptive, and it can rapidly match the dynamics of ICS operating environment. Meanwhile, the effectiveness of the method is discussed when parameters take different values, and it exhibits low false-positive rates but high precision. In addition, the case of error propagation of the framework is analyzed. Ming Zhou 0010, Shichao Lv, Libo Yin, Xin Chen 0123, Hong Li 0004, Limin Sun 0001 |
ICC | 5 |
| 2019 | Side-Channel Information Leakage of Traffic Data in Instant MessagingabstractInstant Messaging has been widely applied for both corporate use and personal use in recent years. Major Instant Messaging service providers adopt the Push Technology to ensure the immediacy of message forwarding, which efficiently provides a great convenience for user. However, the immediacy feature causes side-channel information leakage even if some protection measures has been implemented, such as information encryption strategy. In particular, we observe that senders' traffic flows have a strong temporal correlation with those of corresponding recipients, since the messages are forwarded to recipients as soon as they are received by servers. Based on the observation, attackers can infer real-time communications between pairwise users and even the social connections of users. In this paper, we present a methodology framework to validate this side-channel information leakage, which identifies users of real-time communications by matching the pairwise time sequences of traffic flows. We evaluate the method on the collected real-world data. The experimental results show that users' communications can be identified with a high accuracy, and 6 groups of users are inferred to have strong connections based on the data collected from a local area networks. Ke Li 0042, Hong Li 0004, Hongsong Zhu, Limin Sun 0001, Hui Wen 0001 |
IPCCC | 2 |
| 2019 | ONE-Geo: Client-Independent IP Geolocation Based on Owner Name Extraction
Hongsong Zhu, Hai Zhao 0002, Hong Li 0004, Limin Sun 0001 |
WASA | 5 |
| 2019 | Decentralized Hierarchical Authorized Payment with Online Wallet for Blockchain
Qianwen Wei, Wei Li 0059, Hong Li 0004, Mingsheng Wang |
WASA | 4 |
| 2019 | Towards IP geolocation with intermediate routers based on topology discoveryabstractIP geolocation determines geographical location by the IP address of Internet hosts. IP geolocation is widely used by target advertising, online fraud detection, cyber-attacks attribution and so on. It has gained much more attentions in these years since more and more physical devices are connected to cyberspace. Most geolocation methods cannot resolve the geolocation accuracy for those devices with few landmarks around. In this paper, we propose a novel geolocation approach that is based on common routers as secondary landmarks (Common Routers-based Geolocation, CRG). We search plenty of common routers by topology discovery among web server landmarks. We use statistical learning to study localized (delay, hop)-distance correlation and locate these common routers. We locate the accurate positions of common routers and convert them as secondary landmarks to help improve the feasibility of our geolocation system in areas that landmarks are sparsely distributed. We manage to improve the geolocation accuracy and decrease the maximum geolocation error compared to one of the state-of-the-art geolocation methods. At the end of this paper, we discuss the reason of the efficiency of our method and our future research. Hong Li 0004, Qiang Li 0007, Wei Li 0059, Hongsong Zhu, Limin Sun 0001 |
Cybersecur. | 2 |
| 2019 | Coin Hopping Attack in Blockchain-Based IoTabstractWith dramatic developments of blockchain technology, a number of blockchain-based applications emerge rapidly, among which the incorporation of blockchain into Internet of Things is one of the most valued research direction. Such powerful incorporation is a double-sided sword, i.e., it can benefit both individuals and society but has the vulnerability to coin hopping attack that is a new type of pool mining attack and hard to happen in traditional blockchain networks. In this paper, we theoretically prove the feasibility of coin hopping attack, deeply analyze the conditions of attack implementation, and comprehensively investigate the impacts of coin hopping attack. Moreover, some defense strategies are addressed. To our best knowledge, this paper is the first work targeting coin hopping attack. Saide Zhu, Wei Li 0059, Hong Li 0004, Ling Tian, Guangchun Luo, Zhipeng Cai 0001 |
IEEE Internet Things J. | 3 |
| 2019 | Blockchain for Large-Scale Internet of Things Data Storage and ProtectionabstractWith the dramatically increasing deployment of IoT devices, storing and protecting the large volume of IoT data has become a significant issue. Traditional cloud-based IoT structures impose extremely high computation and storage demands on the cloud servers. Meanwhile, the strong dependencies on the centralized servers bring significant trust issues. To mitigate these problems, we propose a distributed data storage scheme employing blockchain and cetrificateless cryptography. Our scheme eliminates the traditional centralized servers by leveraging the blockchain miners who perform “transaction” verifications and records audit with the help of certificateless cryptography. We present a clear definition of the transactions in a non-cryptocurrency system and illustrate how the transactions are processed. To the best of our knowledge, this is the first work designing a secure and accountable IoT storage system using blockchain. Additionally, we extend our scheme to enable data trading and elaborate how data trading can be efficiently and effectively achieved. Ruinian Li, Tianyi Song, Bo Mei, Hong Li 0004, Xiuzhen Cheng, Limin Sun 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2018 | PANDORA: A Scalable and Efficient Scheme to Extract Version of Binaries in IoT FirmwaresabstractOpen source components are widely used by IoT vendors to develop firmwares in devices. The exposure of vulnerabilities existing in some specific versions of the core components may cause severe security incidents such as the Heartbleed event in 2014 and the Sambacry event in 2016. Extracting the version information from various firmware binaries is significant for evaluating the influence of such incidents and providing emergency response services. To the best of our knowledge, there are still no scalable and efficient extraction methods for binary version information in IoT firmwares. The commonly used method for traditional softwares requires the running up of the firmwares and interaction such as '-version' to obtain the version information. This method is not applicable for IoT devices, as they are built from various platforms which makes it impossible to simulate all of the interested firmwares at large scale. In this paper, we design, implement and evaluate a scalable and efficient binary version extraction framework (termed as PANDORA) for IoT firmwares, which does not rely on the real runtime environment. The main idea of our methodology is to leverage version strings in binaries to get version information. We design a string recover engine (SRE) to recover the missing pieces of those incomplete version strings. We test PANDORA in a dataset containing 2683 IoT binary files. Surprisingly 2267 of them are version-extractable and the recognition rate can reach 84.5%. Hong Li 0004, Zhi Li 0018, Limin Sun 0001 |
ICC | 3 |
| 2018 | A graph neural network based efficient firmware information extraction method for IoT devicesabstractThe firmware information for IoT devices includes the manufacturer, the device type, the device model and the firmware version, etc. Identifying firmware information helps build firmware knowledge graph for many security applications, such as homologous analysis and vulnerability detection of firmware. The traditional firmware information identifying method only utilizes the content-based information, lacks the utilization of the structure information of the firmware, and more importantly, it lacks the use of timing information. Lacking of structural information can reduce prediction accuracy, and lacking of timing information will make it difficult to predict the firmware version. In order to address the disadvantages of the existing method, this paper abstracts the directories or files (components) of the firmware into the nodes of the graph and abstracts the relationships between the nodes into the edges of the graph. Timing information such as component creation time and component version are also attached to the node properties to introduce the time sequence features. As a result, the experimental results show that the accuracy of our method is better than that of random forest for the all four tasks (manufacture, device type, device model and firmware version identification). Particularly, and the accuracy rate is greatly improved in the firmware version identification task. Hong Li 0004, Hui Wen 0001, Hongsong Zhu, Limin Sun 0001 |
IPCCC | 2 |
| 2018 | Robust Network-Based Binary-to-Vector Encoding for Scalable IoT Binary File Retrieval
Hong Li 0004, Zhiqiang Shi, Limin Sun 0001 |
WASA | 2 |
| 2018 | A Secure and Scalable Data Communication Scheme in Smart GridsabstractThe concept of smart grid gained tremendous attention among researchers and utility providers in recent years. How to establish a secure communication among smart meters, utility companies, and the service providers is a challenging issue. In this paper, we present a communication architecture for smart grids and propose a scheme to guarantee the security and privacy of data communications among smart meters, utility companies, and data repositories by employing decentralized attribute based encryption. The architecture is highly scalable, which employs an access control Linear Secret Sharing Scheme (LSSS) matrix to achieve a role‐based access control. The security analysis demonstrated that the scheme ensures security and privacy. The performance analysis shows that the scheme is efficient in terms of computational cost. Chunqiang Hu, Hang Liu 0003, Liran Ma, Yan Huo 0001, Arwa Alrawais, Xiuhua Li 0001, Hong Li 0004, Qingyu Xiong |
Wirel. Commun. Mob. Comput. | 7 |
| 2017 | IHB: A scalable and efficient scheme to identify homologous binaries in IoT firmwaresabstractDue to the extensive code reuse and the widespread use of third-party SDKs, homologous binaries are widely found in IoT firmwares. Once a vulnerability is found in one firmware, other firmwares sharing the similar piece of codes are at high risk. Thus, homologous binary search is of great significance to IoT firmware security analysis. However, there are still no scalable and efficient homologous binary search methods for IoT firmwares. The time complexity of the state-of-the-art method is O(N), and it is not scalable for large-scale IoT firmwares. In this paper, we design, implement, and evaluate a scalable and efficient homologous binary search scheme (termed as IHB) for IoT firmwares with time complexity O(1). The main idea of our methodology is to leverage readable strings in binaries to calculate the similarities between different IoT firmwares. Furthermore, we employ a string filter and the string-based MinHash to achieve both accuracy and efficiency. We test both our scheme and the state-of-the-art methods on a real dataset containing 1024 binary files. The results show that our method is three orders of magnitude more efficient than the existing methods. Meanwhile, our method has a higher true positive rate (92.88%) and a lower false positive rate (2.83%). In the interest of open science, we also make our tools and datasets publicly available to seed future improvements. Hong Li 0004, Zhongjin Liu, Zhiqiang Shi |
IPCCC | 2 |
| 2017 | A Bitcoin Based Incentive Mechanism for Distributed P2P Applications
Yunhua He, Hong Li 0004, Xiuzhen Cheng, Yan Liu 0021, Limin Sun 0001 |
WASA | 2 |
| 2017 | Mobility Intention-Based Relationship Inference from Spatiotemporal Data
Feng Yi, Hong Li 0004, Hongtao Wang 0002, Hui Wen 0001, Limin Sun 0001 |
WASA | 2 |
| 2017 | Mitigating Data Sparsity Using Similarity Reinforcement-Enhanced Collaborative FilteringabstractThe data sparsity problem has attracted significant attention in collaborative filtering-based recommender systems. To alleviate data sparsity, several previous efforts employed hybrid approaches that incorporate auxiliary data sources into recommendation techniques, like content, context, or social relationships. However, due to privacy and security concerns, it is generally difficult to collect such auxiliary information. In this article, we focus on the pure collaborative filtering methods without relying on any auxiliary data source. We propose an improved memory-based collaborative filtering approach enhanced by a novel similarity reinforcement mechanism. It can discover potential similarity relationships between users or items by making better use of known but limited user-item interactions, thus to extract plentiful historical rating information from similar neighbors to make more reliable and accurate rating predictions. This approach integrates user similarity reinforcement and item similarity reinforcement into a comprehensive framework and lets them enhance each other. Comprehensive experiments conducted on several public datasets demonstrate that, in the face of data sparsity, our approach achieves a significant improvement in prediction accuracy when compared with the state-of-the-art memory-based and model-based collaborative filtering algorithms. Weisong Shi, Hong Li 0004 |
ACM Trans. Internet Techn. | 3 |
| 2017 | Why You Go Reveals Who You Know: Disclosing Social Relationship by CooccurrenceabstractThe popularity of location-based services (LBS) and the ubiquity of sensor device have resulted in rich spatiotemporal data. A large number of human behaviors had been recorded including cooccurrence which refers to the phenomenon that two people have been to the same places at the same time. These data enable attackers to infer people’s social relationship based on their cooccurrences and many attack models were proposed. However, current attack models still cannot effectively address the following two challenges: How to distinguish cooccurrences between acquaintances and strangers? What kind of cooccurrence contributes to strong social strength? In this paper, we present a novel social relationship attack model—the Mobility Intention-based Relationship Inference (MIRI) model—which can solve the above two issues. Firstly, we extract mobility intentions and adopt them to characterize cooccurrences. A classification model is trained for attacking social relationship. The experimental results on two real-world datasets demonstrate that the proposed MIRI model can properly differentiate cooccurrences by simultaneously considering spatial and temporal features. The comparison results also indicate that MIRI model significantly outperforms state-of-the-art social relationship attack models. Feng Yi, Hong Li 0004, Hongtao Wang 0002, Limin Sun 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2016 | Side-channel information leakage of encrypted video stream in video surveillance systemsabstractVideo surveillance has been widely adopted to ensure home security in recent years. Most video encoding standards such as H.264 and MPEG-4 compress the temporal redundancy in a video stream using difference coding, which only encodes the residual image between a frame and its reference frame. Difference coding can efficiently compress a video stream, but it causes side-channel information leakage even though the video stream is encrypted, as reported in this paper. Particularly, we observe that the traffic patterns of an encrypted video stream are different when a user conducts different basic activities of daily living, which must be kept private from third parties as obliged by HIPAA regulations. We also observe that by exploiting this side-channel information leakage, attackers can readily infer a user's basic activities of daily living based on only the traffic size data of an encrypted video stream. We validate such an attack using two off-the-shelf cameras, and the results indicate that the user's basic activities of daily living can be recognized with a high accuracy. Hong Li 0004, Yunhua He, Limin Sun 0001, Xiuzhen Cheng, Jiguo Yu |
INFOCOM | 1 |
| 2016 | An Enhanced Structure-Based De-anonymization of Online Social Networks
Hong Li 0004, Cheng Zhang 0018, Yunhua He, Xiuzhen Cheng, Yan Liu 0021, Limin Sun 0001 |
WASA | 1 |
| 2014 | Achieving privacy preservation in WiFi fingerprint-based localizationabstractWiFi fingerprint-based localization is regarded as one of the most promising techniques for indoor localization. The location of a to-be-localized client is estimated by mapping the measured fingerprint (WiFi signal strengths) against a database owned by the localization service provider. A common concern of this approach that has never been addressed in literature is that it may leak the client's location information or disclose the service provider's data privacy. In this paper, we first analyze the privacy issues of WiFi fingerprint-based localization and then propose a Privacy-Preserving WiFi Fingerprint Localization scheme (PriWFL) that can protect both the client's location privacy and the service provider's data privacy. To reduce the computational overhead at the client side, we also present a performance enhancement algorithm by exploiting the indoor mobility prediction. Theoretical performance analysis and experimental study are carried out to validate the effectiveness of PriWFL. Our implementation of PriWFL in a typical Android smartphone and experimental results demonstrate the practicality and efficiency of PriWFL in real-world environments. Hong Li 0004, Limin Sun 0001, Haojin Zhu, Xiang Lu 0004, Xiuzhen Cheng |
INFOCOM | 1 |