Heng Li 0008

dblp:02/3672-8 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-8045-8983ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Do You Really Know What I Am Doing? Backdoor Attacks on Provenance-Based Intrusion Detection Systems
Shaodi Xie, Wei Yuan 0001, Zhu Gong, Heng Li 0008, Tiejun Wu
WWW6
2026 Improving the transferability of targeted adversarial examples by style-agnostic attack
Zimin Mao, Shuijun Yin, Hanwen Zhang 0016, Heng Li 0008, Tiejun Wu, Wei Yuan 0001
Comput. Secur.6
2026 Why Not Diversify Triggers? APK-Specific Backdoor Attack Against Android Malware Detection
abstract
Machine learning-based Android malware detection (AMD) models require abundant data for training robust app classifiers, creating vulnerability to poisoning attacks. Attackers inject poisoned samples into Android app markets, leading to the insertion of a backdoor into the model upon adoption in the training process. Subsequently, attackers can generate evasive malware by embedding a backdoor trigger in malware samples. Currently, research on backdoor attacks towards AMD has just begun to emerge. Existing attacks produce a fixed trigger and apply it to various malware. Once the trigger is discovered by static analysis methods (e.g., software similarity analysis), however, multiple malware carrying this trigger will be simultaneously exposed. To diversify the trigger, we propose anAPK-SpecificBackdoorAttack algorithm (ASBA), which trains a generative adversarial network to generate a specific trigger for every malware sample. Moreover, ASBA manages to make the generated triggers as different as possible, in order to further reduce the likelihood of malware being collectively captured. Extensive experiments have demonstrated that ASBA achieves a 94.6% average attack success rate (ASR) on three datasets, five feature extraction methods and three classification models. Furthermore, compared to state-of-the-art poisoning attack algorithms, ASBA produces more diverse and more effective triggers.
Heng Li 0008, Bang Wu 0002, Cuiying Gao, Wei Yuan 0001, Beihao Xia, Xiapu Luo
IEEE Trans. Dependable Secur. Comput.1
2025 Automated Mass Malware Factory: The Convergence of Piggybacking and Adversarial Example in Android Malicious Software Generation
Heng Li 0008, Bang Wu 0002, Cuiying Gao, Wei Yuan 0001, Xiapu Luo
NDSS1
2025 An Efficient Adversarial Attack on FCG-Based Android Malware Detection Systems
Heng Li 0008, Bang Wu 0002, Wei Yuan 0001, Cuiying Gao, Xinge You, Xiapu Luo
IEEE Trans. Inf. Forensics Secur.1
2024 Efficient Backdoor Attacks for Deep Neural Networks in Real-world Scenarios
abstract
Recent deep neural networks (DNNs) have came to rely on vast amounts of training data, providing an opportunity for malicious attackers to exploit and contaminate the data to carry out backdoor attacks. However, existing backdoor attack methods make unrealistic assumptions, assuming that all training data comes from a single source and that attackers have full access to the training data. In this paper, we introduce a more realistic attack scenario where victims collect data from multiple sources, and attackers cannot access the complete training data. We refer to this scenario as $\textbf{data-constrained backdoor attacks}$. In such cases, previous attack methods suffer from severe efficiency degradation due to the $\textbf{entanglement}$ between benign and poisoning features during the backdoor injection process. To tackle this problem, we introduce three CLIP-based technologies from two distinct streams: $\textit{Clean Feature Suppression}$ and $\textit{Poisoning Feature Augmentation}$. The results demonstrate remarkable improvements, with some settings achieving over $\textbf{100}$% improvement compared to existing attacks in data-constrained scenarios.
Ziqiang Li 0001, Heng Li 0008, Beihao Xia, Yi Wu 0018, Bin Li 0025
ICLR4
2024 A Comprehensive Study of Learning-based Android Malware Detectors under Challenging Environments
abstract
Recent years have witnessed the proliferation of learning-based Android malware detectors. These detectors can be categorized into three types, String-based, Image-based and Graph-based. Most of them have achieved good detection performance under the ideal setting. In reality, however, detectors often face out-of-distribution samples due to the factors such as code obfuscation, concept drift (e.g., software development technique evolution and new malware category emergence), and adversarial examples (AEs). This problem has attracted increasing attention, but there is a lack of comparative studies that evaluate the existing various types of detectors under these challenging environments. In order to fill this gap, we select 12 representative detectors from three types of detectors, and evaluate them in the challenging scenarios involving code obfuscation, concept drift and AEs, respectively. Experimental results reveal that none of the evaluated detectors can maintain their ideal-setting detection performance, and the performance of different types of detectors varies significantly under various challenging environments. We identify several factors contributing to the performance deterioration of detectors, including the limitations of feature extraction methods and learning models. We also analyze the reasons why the detectors of different types show significant performance differences when facing code obfuscation, concept drift and AEs. Finally, we provide practical suggestions from the perspectives of users and researchers, respectively. We hope our work can help understand the detectors of different types, and provide guidance for enhancing their performance and robustness.
Cuiying Gao, Gaozhun Huang, Heng Li 0008, Bang Wu 0003, Yueming Wu 0001, Wei Yuan 0001
ICSE3
2024 Trace-agnostic and Adversarial Training-resilient Website Fingerprinting Defense
abstract
Deep neural network (DNN) based website fingerprinting (WF) attacks can achieve an attack success rate (ASR) of over 90%, seriously threatening the privacy of Tor users. At present, adversarial example (AE) based defenses have demonstrated great potential to defend against WF attacks. However, existing AE-based defenses require knowing a complete traffic trace for adversarial perturbation calculation, which is unrealistic in practice. Moreover, they may become ineffective once adversarial training (AT) is adopted by attackers. To mitigate these two problems, we propose a defense called ALERT. It generates adversarial perturbations without knowing traffic traces, and can effectively resist AT-aided WF attacks. The key idea of ALERT is to produce universal perturbations that vary from user to user. We conduct extensive experiments to evaluate ALERT. In the closed world, ALERT significantly surpasses four representative WF defenses, including the state-of-the-art (SOTA) defense AWA. Specifically, ALERT reduces the ASR of the SOTA DF attack to 12.68% and uses only 20.13% of communication bandwidth. In the open world, ALERT uses only 19.91% of bandwidth, reduces the True Positive Rate (TPR) of the DF attack to 37.46%, obviously outperforming the other defenses.
Litao Qiao, Bang Wu 0002, Heng Li 0008, Cuiying Gao, Wei Yuan 0001, Xiapu Luo
INFOCOM3
2024 Uncovering and Mitigating the Impact of Code Obfuscation on Dataset Annotation with Antivirus Engines
abstract
With the widespread application of machine learning-based Android malware detection methods, building a high-quality dataset has become increasingly important. Existing large-scale datasets are mostly annotated with VirusTotal by aggregating the decisions of antivirus engines, and most of them indiscriminately accept the decisions of all engines. In reality, however, these engines have different capabilities in detecting malware, especially those that have been obfuscated. Previous research has revealed that code obfuscation degrades the detection performance of these engines to varying degrees. This makes us believe that using all engines indiscriminately is unreasonable for dataset annotation. Therefore, in this paper, we first conduct a data-driven evaluation to confirm the negative effects of code obfuscation on engine-based dataset annotation. To gain a deeper understanding of the reasons behind this phenomenon, we evaluate the availability, effectiveness and robustness of every engine under various code obfuscation techniques. Then we categorize the engines and select a set of obfuscation-robust engines. Finally, we conduct comprehensive experiments to verify the effectiveness of the selected engines for dataset annotation. Our experiments show that when 50% obfuscated samples are mixed into the training set, on the classic malware detectors Drebin and Malscan, using our selected engines can effectively improve detection performance by 15.21% and 19.23%, respectively, compared to using all the engines.
Cuiying Gao, Yueming Wu 0001, Heng Li 0008, Wei Yuan 0001, Qidan He, Yang Liu 0003
ISSTA3
2024 Enhancing robustness of person detection: A universal defense filter against adversarial patch attacks
Zimin Mao, Shuiyan Chen, Zhuang Miao, Heng Li 0008, Beihao Xia, Junzhe Cai, Wei Yuan 0001, Xinge You
Comput. Secur.4
2024 Semi-supervised anomaly detection with contamination-resilience and incremental training
Liheng Yuan, Fanghua Ye 0001, Heng Li 0008, Cuiying Gao, Chengqing Yu, Wei Yuan 0001, Xinge You
Eng. Appl. Artif. Intell.3
2023 HARP: Let Object Detector Undergo Hyperplasia to Counter Adversarial Patches
abstract
Adversarial patches can mislead object detectors to produce erroneous predictions. To defend against adversarial patches, one can take two types of protections on the model side, including modifying the detector itself (e.g., adversarial training) or attaching a new model in front of the detector. However, the former often deteriorates clean performance of detectors, and the latter may have high deployment costs caused by too many training parameters. Inspired by the phenomenon of "bone hyperplasia" in human bodies, we present a novel model-side adversarial patch defense, called HARP (Hyperplasia based Adversarial Patch defense). Just as bone hyperplasia can enhance bone strength and skeletal stability, the hyperostosia of detectors can also help to resist adversarial patches. Following this idea, HARP chooses to improve adversarial robustness by "growing" lightweight CNN modules (i.e., hyperplasia modules) on the pre-trained object detectors. We conduct extensive experiments on the PASCAL VOC and COCO datasets to compare HARP with the data-side defense JPEG and the model-side defenses adversarial training, SAC and FNC. Experimental results show that HARP provides excellent defense against adversarial patches while maintaining clean performance, outperforming the compared defense methods. Under PGD-based adaptive attacks, HARP surpasses the recently proposed defense method SAC by 12.5% in mean average precision (mAP) on PASCAL VOC, and 13.2% on COCO dataset. In addition, experiments confirm that the increase in model inference time caused by HARP is almost negligible.
Junzhe Cai, Shuiyan Chen, Heng Li 0008, Beihao Xia, Zimin Mao, Wei Yuan 0001
ACM Multimedia3
2023 Black-box Adversarial Example Attack towards FCG Based Android Malware Detection under Incomplete Feature Information
Heng Li 0008, Zhang Cheng, Bang Wu 0002, Liheng Yuan, Cuiying Gao, Wei Yuan 0001, Xiapu Luo
USENIX Security Symposium1
2023 Detecting Android Malware With Pre-Existing Image Classification Neural Networks
abstract
Android malware detection has attracted increasing attention due to the rapid growth of mobile malware. However, running an in-cloud Android malware detection system usually incurs high hardware and bandwidth costs. This dilemma motivates us to develop a method to repurpose an in-cloud image-classification neural network to detect Android malware. Given an Android app, the proposed method first embeds its features into an image, skillfully perturbs the feature-embedded image, and then feeds the modified image into the in-cloud image classifier. The classifier's outputs are finally mapped into a malware detection result. In addition, two new techniques (perturbation hiding and group mapping) are proposed to reduce the risk of repurposing behavior being recognized and improve detection performance. Experiments show that our perturbations are usually imperceptible to humans, and our method outperforms both traditional machine learning-based detectors and deep learning-based detectors in detection performance.
Shuijun Yin, Heng Li 0008, Minghui Cai, Wei Yuan 0001
IEEE Signal Process. Lett.3
2023 Obfuscation-Resilient Android Malware Analysis Based on Complementary Features
abstract
Existing Android malware detection methods are usually hard to simultaneously resist various obfuscation techniques. Therefore, bytecode-based code obfuscation becomes an effective means to circumvent Android malware analysis. Building obfuscation-resilient Android malware analysis methods is a challenging task, due to the fact that various obfuscation techniques have vastly different effects on code and detection features. To mitigate this problem, we propose combining multiple features that are complementary in combating code obfuscation. Accordingly, we develop an obfuscation-resilient Android malware analysis methodCorDroid, based on two new features: Enhanced Sensitive Function Call Graph (E-SFCG) and Opcode-based Markov transition Matrix (OMM). The first describes sensitive function call relationships, while the second reflects transition probabilities among opcodes. Combining E-SFCG and OMM can well characterize the runtime behavior of Android apps from different perspectives, hence increasing the difficulty of misleading malware analysis through using code obfuscation to affect detection features. To evaluate CorDroid, we generate 74,138 obfuscated samples with 14 different obfuscation techniques, and compare CorDroid with the state-of-the-art detection methods (e.g., MaMaDroid, RevealDroid and APIGraph). In terms of average F1-Score, CorDroid is 29.69% higher than MaMaDroid, 21.80% higher than APIGraph, and 9.71% higher than RevealDroid, respectively. Experiments also validate the complementarity between E-SFCG and OMM, and exhibit the high execution efficiency of CorDroid.
Cuiying Gao, Minghui Cai, Shuijun Yin, Gaozhun Huang, Heng Li 0008, Wei Yuan 0001, Xiapu Luo
IEEE Trans. Inf. Forensics Secur.5
2023 Resisting DNN-Based Website Fingerprinting Attacks Enhanced by Adversarial Training
abstract
Deep neural network (DNN) based website fingerprinting (WF) attacks pose a severe threat to the privacy of Tor users. To overcome this challenge, adversarial perturbation based WF defenses have been recently proposed to fool the classifiers of attackers, through purposefully perturbing the user’s traffic traces. Unfortunately, these defenses significantly deteriorate once the WF attacks are enhanced withadversarial training(AT). AT endows the WF attacks with more powerful website recognition capability, through learning the perturbed traffic traces generated by attackers. To resist the WF attacks enhanced by AT, we develop ablack-boxWF defense, called Acup3. First, Acup3 leveragesmany-to-one website imitationto make the traffic traces associated with different websites look more like each other, increasing the difficulty of website classification. Second, Acup3 generatestrace-agnosticperturbations without accessing traffic traces, making it suitable for practical deployment. Third, Acup3 employsperturbation variationto diversify the traffic traces of different users visiting the same website, making the knowledge learnt from AT less helpful for WF attacks. Therefore, Acup3 is more robust against AT. Experiments demonstrate Acup3 markedly surpasses four representative WF defenses (e.g., Mockingbird and AWA) in defense capability and bandwidth overhead. Facing the state-of-the-art (SOTA) attack Var-CNN enhanced with AT, Acup3 depresses its attack success rate (ASR) from 98% to 24.29% with only 13.95% bandwidth overhead. Compared to the SOTA defense AWA, Acup3 causes a 24.5% larger decrement in ASR of WF attacks, and achieves a more than 100 times faster speed of perturbation generation.
Litao Qiao, Bang Wu 0002, Shuijun Yin, Heng Li 0008, Wei Yuan 0001, Xiapu Luo
IEEE Trans. Inf. Forensics Secur.4
2022 Recent Advances in Concept Drift Adaptation Methods for Deep Learning
abstract
In the ``Big Data'' age, the amount and distribution of data have increased wildly and changed over time in various time-series-based tasks, e.g weather prediction, network intrusion detection. However, deep learning models may become outdated facing variable input data distribution, which is called concept drift. To address this problem, large number of samples are usually required to update deep learning models, which is impractical in many realistic applications. This challenge drives researchers to explore the effective ways to adapt deep learning models to concept drift. In this paper, we first mathematically describe the categories of concept drift including abrupt drift, gradual drift, recurrent drift, incremental drift. We then divide existing studies into two categories (i.e., model parameter updating and model structure updating), and analyze the pros and cons of representative methods in each category. Finally, we evaluate the performance of these methods, and point out the future directions of concept drift adaptation for deep learning.
Liheng Yuan, Heng Li 0008, Beihao Xia, Cuiying Gao, Wei Yuan 0001, Xinge You
IJCAI2
2022 Model scheduling and sample selection for ensemble adversarial example attacks
Zichao Hu, Heng Li 0008, Liheng Yuan, Zhang Cheng, Wei Yuan 0001
Pattern Recognit.2
2021 Robust Android Malware Detection against Adversarial Example Attacks
abstract
Adversarial examples pose severe threats to Android malware detection because they can render the machine learning based detection systems useless. How to effectively detect Android malware under various adversarial example attacks becomes an essential but very challenging issue. Existing adversarial example defense mechanisms usually rely heavily on the instances or the knowledge of adversarial examples, and thus their usability and effectiveness are significantly limited because they often cannot resist the unseen-type adversarial examples. In this paper, we propose a novel robust Android malware detection approach that can resist adversarial examples without requiring their instances or knowledge by jointly investigating malware detection and adversarial example defenses. More precisely, our approach employs a new VAE (variational autoencoder) and an MLP (multi-layer perceptron) to detect malware, and combines their detection outcomes to make the final decision. In particular, we share a feature extraction network between the VAE and the MLP to reduce model complexity and design a new loss function to disentangle the features of different classes, hence improving detection performance. Extensive experiments confirm our model’s advantage in accuracy and robustness. Our method outperforms 11 state-of-the-art robust Android malware detection models when resisting 7 kinds of adversarial example attacks.
Heng Li 0008, Shiyao Zhou, Wei Yuan 0001, Xiapu Luo, Cuiying Gao, Shuiyan Chen
WWW1
2021 Learning features from enhanced function call graphs for Android malware detection
Minghui Cai, Cuiying Gao, Heng Li 0008, Wei Yuan 0001
Neurocomputing4
2021 Black-box attack against handwritten signature verification with region-restricted adversarial perturbations
Heng Li 0008, Hansong Zhang 0002, Wei Yuan 0001
Pattern Recognit.2
2021 A Lightweight On-Device Detection Method for Android Malware
abstract
Android malware poses severe threats to users, hence raising an urgent demand for malware detection. In-cloud Android malware detection often suffers privacy leakage and communication overheads. Therefore, this article focuses on on-device Android malware detection. At present, on-device malware detectors are usually trained on servers and then transplanted to mobile devices (e.g., smartphones). In practice, on-device training is particularly important due to the demand for offline updates. Because mobile devices are limited in resource, however, on-device training is hard to implement, especially for those high-complexity malware detectors. To overcome this challenge, we design a lightweight on-device Android malware detector, based on the recently proposed broad learning method. Our detector mainly uses one-shot computation for model training. Hence it can be fully or incrementally trained directly on mobile devices. As far as detection accuracy is concerned, our detector outperforms the shallow learning-based models, including support vector machine (SVM) and AdaBoost, and approaches the deep learning-based models multilayer perceptron (MLP) and convolutional neural network (CNN). Moreover, our detector is more robust to adversarial examples than the existing detectors, and its robustness can be further improved through on-device model retraining. Finally, its advantages are confirmed by extensive experiments, and its practicality is demonstrated through runtime evaluation on smartphones.
Wei Yuan 0001, Heng Li 0008, Minghui Cai
IEEE Trans. Syst. Man Cybern. Syst.3