VLDB 2026 Research / reviewers in the wild / expert
Guozhu Meng
dblp:134/8681
· DBLP profile ↗
66ranked-venue papers
5as first author
43since 2021 · last 2026
0000-0001-6388-2571ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 27 · 3 first-author · 20 since 2021Software engineering, systems software and programming languages · 23 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
Tong Liu 0027, Zhe Zhao 0007, Guozhu Meng, Kai Chen 0012 |
NDSS | 4 |
| 2026 | SwitchNet: protecting neural networks by structure obfuscation and switch-controlled inferenceabstractAbstract Training deep learning models requires substantial financial and human resources, so once deployed in untrusted environments, these models immediately attract the attention of attackers who seek to steal and misuse them. Traditional model protection methods are ineffective in addressing model accuracy, performance, and proactive defense. To this end, we present an active defensive approach SwitchNet by obfuscating model structure and proposing a switch-controlled mechanism to manage model inference. Specifically, SwitchNet learns the weight distribution of the original model and then constructs confusion layers that are strategically inserted into the original model for structure obfuscation. Each of the model layers is equipped with a switch , which is controlled by a switching policy network. We train this policy network with an adaptive pattern as a “secret key” that can accurately control the switch states, and thereby the model inference process. We conduct a comprehensive theoretical analysis of the perturbation boundary and certify that SwitchNet maintains high robustness under $$\ell _\infty$$ ℓ ∞ perturbations, with certified accuracy exceeding 80% at $$\epsilon = 0.0048$$ ϵ = 0.0048 (CROWN). In addition, we perform extensive experiments on both classical convolutional networks and Vision Transformers. The results show that SwitchNet effectively preserves model accuracy for legitimate users (with only a 0.35% drop), while reducing the accuracy for unauthorized users to near-random guessing. Compared to the state-of-the-art, our approach reduces inference and construction overhead by 20.89% and 12.08%, respectively. Furthermore, SwitchNet proves to be stealthy and resilient against various attacks aimed at detecting or compromising the protection mechanism. Yuling Cai, Guozhu Meng, Yinzhi Cao, Guangdong Bai |
Cybersecur. | 2 |
| 2026 | Revolutionizing electricity theft detection: enhanced accuracy through NILM and multi-source data fusionabstractAbstract Electricity theft detection seeks to thwart the illegal use of electricity, thereby safeguarding the safety and stability of the power system. Traditional methods, which typically rely on aggregated household consumption data to identify theft, often overlook the fact that household consumption is vulnerable to fluctuations in normal user behavior. This results in high false positive and false negative rates. To refine the accuracy, we propose a novel electricity theft detection method based on Non-Intrusive Load Monitoring (NILM) and multi-source data fusion. Our approach employs advanced NILM algorithms to cost-effectively extract individual appliance consumption data from aggregated power signals. We then integrate this data with household aggregate consumption data through a multi-source data fusion architecture. By analyzing the unique consumption patterns of different types of appliances, our approach identifies theft behaviors that cannot be detected by aggregate consumption data alone. Experimental results across three real-world datasets demonstrate that our method significantly outperforms single-source data-based benchmarks, achieving up to a 7.92% gain in F1-score and a 12.6% gain in Precision. Moreover, our method exhibits strong generalization ability across a series of typical machine learning models. Zhiwei Deng, Junsen Feng, Jialing He, Guozhu Meng, Tao Xiang 0001 |
Cybersecur. | 4 |
| 2026 | A survey on physical adversarial attacks against face recognition systems
Mingsi Wang, Jiachen Zhou 0001, Tianlin Li, Guozhu Meng, Kai Chen 0012 |
Neurocomputing | 4 |
| 2025 | AUTHNET: Neural Network with Integrated Authentication LogicabstractModel stealing, i.e., unauthorized access and exfiltration of deep learning models, has emerged as a significant security threat. The misuse and illegal replication of models pose major risks to financial assets and competitive advantage. Traditional protection methods, such as model watermarking, are passive and challenging to enforce, while active defenses often face limitations in terms of efficiency and the security required for widespread deployment. To this end, we propose a native authentication mechanism, called AUTHNET, which integrates authentication logic as part of the model without any additional structures. Our key insight is to reuse redundant neurons with low activation and embed authentication bits in an intermediate layer, called a gate layer. Then, AUTHNET fine-tunes the layers after the gate layer to embed authentication logic so that only inputs with secret key can trigger the correct logic of AUTHNET. It provides the last line of defense, i.e., even being exfiltrated, the model is not usable as the adversary cannot generate valid inputs without the key. We theoretically demonstrate the high sensitivity of AUTHNET to the secret key, which means that precise key provision is essential for achieving good performance of AUTHNET. AUTHNET is compatible with any convolutional neural network, where our extensive evaluations show that AUTHNET successfully achieves the goal in rejecting unauthenticated users (whose average accuracy drops to 22.03%) with a trivial accuracy decrease (1.18% on average) for legitimate users, and is robust against adaptive attacks, providing efficient and lightweight protection. Yuling Cai, Fan Xiang, Guozhu Meng, Yinzhi Cao, Kai Chen 0012 |
ECAI | 3 |
| 2025 | Decictor: Towards Evaluating the Robustness of Decision-Making in Autonomous Driving SystemsabstractAutonomous Driving System (ADS) testing is crucial in ADS development, with the current primary focus being on safety. However, the evaluation of non-safety-critical performance, particularly the ADS's ability to make optimal decisions and produce optimal paths for autonomous vehicles (AVs), is also vital to ensure the intelligence and reduce risks of AVs. Currently, there is little work dedicated to assessing the robustness of ADSs' path-planning decisions (PPDs), i.e., whether an ADS can maintain the optimal PPD after an insignificant change in the environment. The key challenges include the lack of clear oracles for assessing PPD optimality and the difficulty in searching for scenarios that lead to non-optimal PPDs. To fill this gap, in this paper, we focus on evaluating the robustness of ADSs' PPDs and propose the first method, Decictor, for generating nonoptimal decision scenarios (NoDSs), where the ADS does not plan optimal paths for AVs. Decictor comprises three main components: Non-invasive Mutation, Consistency Check, and Feedback. To overcome the oracle challenge, Non-invasive Mutation is devised to implement conservative modifications, ensuring the preservation of the original optimal path in the mutated scenarios. Subsequently, the Consistency Check is applied to determine the presence of nonoptimal PPDs by comparing the driving paths in the original and mutated scenarios. To deal with the challenge of large environment space, we design Feedback metrics that integrate spatial and temporal dimensions of the AV's movement. These metrics are crucial for effectively steering the generation of NoDSs. Therefore, Decictor can generate NoDSs by generating new scenarios and then identifying NoDSs in the new scenarios. We evaluate Decictor on Baidu Apollo, an open-source and production-grade ADS. The experimental results validate the effectiveness of Decictor in detecting non-optimal PPDs of ADSs. It generates 63.9 NoDSs in total, while the best-performing baseline only detects 35.4 NoDSs. Mingfei Cheng, Xiaofei Xie, Yuan Zhou 0005, Junjie Wang 0007, Guozhu Meng, Kairui Yang |
ICSE | 5 |
| 2025 | Measuring and Explaining the Effects of Android App Transformations in Online Malware DetectionabstractIt is well known that antivirus engines are vulnerable to evasion techniques (e.g., obfuscation) that transform malware into its variants.However, it cannot be necessarily attributed to the effectiveness of these evasions, and the limits of engines may also make this unsatisfactory result.In this study, we propose a data-driven approach to measure the effect of app transformations to malware detection, and further explain why the detection result is produced by these engines.First, we develop an interaction model for antivirus engines, illustrating how they respond with different detection results in terms of varying inputs.Six app transformation techniques are implemented in order to generate a large number of Android apps with traceable changes.Then we undertake a onemonth tracking of app detection results from multiple antivirus engines, through which we obtain over 971K detection reports from VirusTotal for 179K apps in total.Last, we conduct a comprehensive analysis of antivirus engines based on these reports from the perspectives of signature-based, static analysis-based, and dynamic analysis-based detection techniques.The results, together with 7 highlighted findings, identify a number of sealed working mechanisms occurring inside antivirus engines and what are the indicators of compromise in apps during malware detection. Guozhu Meng, Zhixiu Guo, Xiaodong Zhang 0014, Haoyu Wang 0001, Kai Chen 0012, Yang Liu 0003 |
Internetware | 1 |
| 2025 | SAP-DIFF: Semantic Adversarial Patch Generation for Black-Box Face Recognition Models via Diffusion ModelsabstractGiven the need to evaluate the robustness of face recognition (FR) models, many efforts have focused on adversarial patch attacks that mislead FR models by introducing localized perturbations. Impersonation attacks are a significant threat because adversarial perturbations allow attackers to disguise themselves as legitimate users. This can lead to severe consequences, including data breaches, system damage, and misuse of resources. However, research on such attacks in FR remains limited. Existing adversarial patch generation methods exhibit limited efficacy in impersonation attacks due to (1) the need for high attacker capabilities, (2) low attack success rates, and (3) excessive query requirements. To address these challenges, we propose a novel method SAP-DIFF that leverages diffusion models to generate adversarial patches via semantic perturbations in the latent space rather than direct pixel manipulation. We introduce an attention disruption mechanism to generate features unrelated to the original face, facilitating the creation of adversarial samples and a directional loss function to guide perturbations toward the target identity's feature space, thereby enhancing attack effectiveness and efficiency. Extensive experiments on popular FR models and datasets demonstrate that our method outperforms state-of-the-art approaches, achieving an average attack success rate improvement of 45.66% (all exceeding 40%), and a reduction in the number of queries by about 40% compared to the SOTA approach. Mingsi Wang, Shuaiyin Yao, Chang Yue, Guozhu Meng |
ICMR | 5 |
| 2025 | Dormant: Defending against Pose-driven Human Image Animation
Jiachen Zhou 0001, Mingsi Wang, Tianlin Li, Guozhu Meng, Kai Chen 0012 |
USENIX Security Symposium | 4 |
| 2025 | Deep learning-based software engineering: progress, challenges, and opportunitiesabstractAbstract Researchers have recently achieved significant advances in deep learning techniques, which in turn has substantially advanced other research disciplines, such as natural language processing, image processing, speech recognition, and software engineering. Various deep learning techniques have been successfully employed to facilitate software engineering tasks, including code generation, software refactoring, and fault localization. Many studies have also been presented in top conferences and journals, demonstrating the applications of deep learning techniques in resolving various software engineering tasks. However, although several surveys have provided overall pictures of the application of deep learning techniques in software engineering, they focus more on learning techniques, that is, what kind of deep learning techniques are employed and how deep models are trained or fine-tuned for software engineering tasks. We still lack surveys explaining the advances of subareas in software engineering driven by deep learning techniques, as well as challenges and opportunities in each subarea. To this end, in this study, we present the first task-oriented survey on deep learning-based software engineering. It covers twelve major software engineering subareas significantly impacted by deep learning techniques. Such subareas spread out through the whole lifecycle of software development and maintenance, including requirements engineering, software development, testing, maintenance, and developer collaboration. As we believe that deep learning may provide an opportunity to revolutionize the whole discipline of software engineering, providing one survey covering as many subareas as possible in software engineering can help future research push forward the frontier of deep learning-based software engineering more systematically. For each of the selected subareas, we highlight the major advances achieved by applying deep learning techniques with pointers to the available datasets in such a subarea. We also discuss the challenges and opportunities concerning each of the surveyed software engineering subareas. Xiangping Chen, Xing Hu 0008, Yuan Huang 0002, He Jiang 0001, Weixing Ji, Yanjie Jiang, Yanyan Jiang 0001, Bo Liu 0094, Hui Liu 0003, Xiaoli Lian, Guozhu Meng, Xin Peng 0001, Hailong Sun 0001, Lin Shi 0006, Bo Wang 0050, Chong Wang 0013, Jifeng Xuan, Xin Xia 0001, Yibiao Yang, Yixin Yang 0006, Li Zhang 0029, Yuming Zhou, Lu Zhang 0023 |
Sci. China Inf. Sci. | 12 |
| 2025 | MalFocus: Locating Malicious Modules in Malware Based on Hybrid Deep LearningabstractIn recent years, binary malware detection has attracted extensive attention from industry and academia. However, most of the existing work only focuses on judging whether a sample is malicious or not, rather than identifying malicious modules in malware. Few studies aiming at locating malicious code work on the function granularity and suffer from inaccuracy. In this paper, we address this problem by locating malicious code at the functional module (FM) granularity, which combines several functions to express the malicious behaviors of malware. We design a tool called MalFocus to automatically divide malware intoFMsand then identify the malicious functional module (MFM) in a multi-model hybrid manner, in which an unsupervised model and an interpretability approach based on a binary classifier are combined, eliminating the workload of labeling malware samples, determining the scope ofMFMsand ranking them according to their maliciousness. The identifiedMFMsare then passed to security analysts for verification, helping to significantly reduce the scope of manual analysis while providing a comprehensive view of the malware attack flow. Additionally, rules derived from the verifiedMFMscan be used to detect variants and new malware families with different functionalities, offering a more general and flexible detection approach. We evaluate MalFocus’s performance on 6764 real-world samples. The results show that MalFocus can correctly identify 95% ofMFMs, outperforming current state-of-the-art work. Weihao Huang, Chaoyang Lin, Lu Xiang, Zhiyu Zhang 0017, Guozhu Meng, Lei Xue 0001, Kai Chen 0012, Zongming Zhang |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | DataElixir: Purifying Poisoned Dataset to Mitigate Backdoor Attacks via Diffusion ModelsabstractDataset sanitization is a widely adopted proactive defense against poisoning-based backdoor attacks, aimed at filtering out and removing poisoned samples from training datasets. However, existing methods have shown limited efficacy in countering the ever-evolving trigger functions, and often leading to considerable degradation of benign accuracy. In this paper, we propose DataElixir, a novel sanitization approach tailored to purify poisoned datasets. We leverage diffusion models to eliminate trigger features and restore benign features, thereby turning the poisoned samples into benign ones. Specifically, with multiple iterations of the forward and reverse process, we extract intermediary images and their predicted labels for each sample in the original dataset. Then, we identify anomalous samples in terms of the presence of label transition of the intermediary images, detect the target label by quantifying distribution discrepancy, select their purified images considering pixel and feature distance, and determine their ground-truth labels by training a benign model. Experiments conducted on 9 popular attacks demonstrates that DataElixir effectively mitigates various complex attacks while exerting minimal impact on benign accuracy, surpassing the performance of baseline defense methods. Jiachen Zhou 0001, Peizhuo Lv, Yibing Lan, Guozhu Meng, Kai Chen 0012, Hualong Ma |
AAAI | 4 |
| 2024 | Demystifying RCE Vulnerabilities in LLM-Integrated AppsabstractLarge Language Models (LLMs) show promise in transforming software development, with a growing interest in integrating them into more intelligent apps. Frameworks like LangChain aid LLM-integrated app development, offering code execution utility/APIs for custom actions. However, these capabilities theoretically introduce Remote Code Execution (RCE) vulnerabilities, enabling remote code execution through prompt injections. No prior research systematically investigates these frameworks' RCE vulnerabilities or their impact on applications and exploitation consequences. Therefore, there is a huge research gap in this field. Tong Liu 0027, Zizhuang Deng, Guozhu Meng, Yuekang Li, Kai Chen 0012 |
CCS | 3 |
| 2024 | Android Malware Family Labeling: Perspectives from the IndustryabstractLabeling and classifying Android malware is important for identifying new threats, triaging security incidents, and demystifying evasion techniques. To automate the malware classification pipeline, state-of-the-art tools such as AVClass and Euphony unify raw labels from commercial antivirus vendors (i.e., VirusTotal) to produce family labels. These tools are widely used for automatic malware classification in both academic research and industry practice. However, they face significant limitations in real-world industrial scenarios with numerous and dynamically changing samples. For example, our industrial practices revealed that VirusTotal's results change over time, leading to temporal inconsistencies in family labeling results that rely on label unification, which can severely impact a company's security posture. Despite this, such issues and challenges remain understudied. In this paper, we present the first systematic measurement study of existing automatic Android malware family labeling systems from various aspects, including label dynamics, consistency, reliability, and etc. Based on a large-scale dataset, we validate that the labeling results of these systems do evolve with time, and such evolution can introduce bias into many previous studies on performance assessments. We also reveal substantial divergence in labeling decisions across different systems when given the same input. Besides, we identify a disclosure priority among families in these systems' labeling processes, which could threaten the industry by allowing malicious actors to exploit these discrepancies. Our findings could benefit both researchers and industry practitioners for further refinement of automatic malware family labeling systems, contributing to their practical applications. Liu Wang 0002, Haoyu Wang 0001, Tao Zhang 0001, Haitao Xu 0002, Guozhu Meng, Peiming Gao, Yi Wang 0013 |
ASE | 5 |
| 2024 | Attribution-guided Adversarial Code Prompt Generation for Code Completion ModelsabstractLarge language models have made significant progress in code completion, which may further remodel future software development. However, these code completion models are found to be highly risky as they may introduce vulnerabilities unintentionally or be induced by a special input, i.e., adversarial code prompt. Prior studies mainly focus on the robustness of these models, but their security has not been fully analyzed. Guozhu Meng, Shangqing Liu, Lu Xiang, Kai Chen 0012, Xiapu Luo, Yang Liu 0003 |
ASE | 2 |
| 2024 | SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-Supervised Learning
Peizhuo Lv, Shenchen Zhu, Shengzhi Zhang, Kai Chen 0012, Ruigang Liang, Chang Yue, Fan Xiang, Yuling Cai, Hualong Ma, Guozhu Meng |
NDSS | 12 |
| 2024 | RTS: A Training-time Backdoor Defense Strategy Based on Weight Residual TendencyabstractThe backdoor attack, caused by malicious samples in the training dataset, has been proven to be a significant threat to the security of deep learning models. The victim model would perform wrongly to the inputs with the trigger. When deployed models in sensitive domains such as facial recognition or autonomous vehicles, the backdoor could lead to potentially disastrous outcomes, like passing the authentication mechanisms and causing traffic accidents. Most of the previous defense works first train the model on the suspicious training dataset and then filter suspicious samples by focusing on the different features in the representation vectors of inputs. The model must be retrained on the filtered training dataset to clear the potential backdoor, bringing in inconvenience and extra computation overhead. This work proposes a training-time sample filtering defense scheme named RTS, with which the model would only be trained once and the potential backdoor could be defended. A pivotal observation is that the target label samples contain richer backdoor-related information compared to other samples, which would result in more adjustments in the model’s weights. Therefore, we filter the poison data with an adaption strategy based on the model’s weight residual tendency in the training process. After all training iterations, we can filter all suspicious samples and obtain a benign model.To demonstrate the effectiveness of the RTS, we conduct experiments on 4 image classification datasets with different model structures at 3 different poison rates (0.01, 0.05, 0.1). Taking the poison rate of 0.05 as an example, the attack success rates on them are decreased to 6.53% (MNIST), 11.88% (CIFAR10), 36.17% (GTSRB), and 12.14% (STL10), while the accuracies of them achieve 96.92% (MNIST), 80.27% (CIFAR10), 83.05% (GTSRB), and 74.9% (STL10). We also find that the effectiveness would not be affected by the target label and suggest the value of σ be in (0.1, 0.4]. About 1% of the training data would be enough for the shadow dataset to decrease the ASR to lower than 20% robustly. These above ablation experiments demonstrate the robustness of RTS. Fan Xiang, Guozhu Meng |
TrustCom | 3 |
| 2024 | Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
Tong Liu 0027, Zhe Zhao 0007, Yinpeng Dong, Guozhu Meng, Kai Chen 0012 |
USENIX Security Symposium | 5 |
| 2024 | Revealing the exploitability of heap overflow through PoC analysisabstractAbstract The exploitable heap layouts are used to determine the exploitability of heap vulnerabilities in general-purpose applications. Prior studies have focused on using fuzzing-based methods to generate more exploitable heap layouts. However, the exploitable heap layout cannot fully demonstrate the exploitability of a vulnerability, as it is uncertain whether the attacker can control the data covered by the overflow. In this paper, we propose the Heap Overflow Exploitability Evaluator (Hoee), a new approach to automatically reveal the exploitability of heap buffer overflow vulnerabilities by evaluating proof-of-concepts (PoCs) generated by fuzzers. Hoee leverages several techniques to collect dynamic information at runtime and recover heap object layouts in a fine-grained manner. The overflow context is carefully analyzed to determine whether the sensitive pointer is corrupted, tainted, or critically used. We evaluate Hoee on 34 real-world CVE vulnerabilities from 16 general-purpose programs. The results demonstrate that Hoee accurately identifies the key factors for developing exploits in vulnerable contexts and correctly recognizes the behavior of overflow. Qintao Shen, Guozhu Meng, Kai Chen 0012 |
Cybersecur. | 2 |
| 2024 | Automated Commit Intelligence by Pre-trainingabstractGitHub commits, which record the code changes with natural language messages for description, play a critical role in software developers’ comprehension of software evolution. Due to their importance in software development, several learning-based works are conducted for GitHub commits, such as commit message generation and security patch identification. However, most existing works focus on customizing specialized neural networks for different tasks. Inspired by the superiority of code pre-trained models, which has confirmed their effectiveness across different downstream tasks, to promote the development of open-source software community, we first collect a large-scale commit benchmark including over 7.99 million commits across 7 programming languages. Based on this benchmark, we present CommitBART, a pre-trained encoder-decoder Transformer model for GitHub commits. The model is pre-trained by three categories (i.e., denoising objectives, cross-modal generation, and contrastive learning) for six pre-training tasks to learn commit fragment representations. Our model is evaluated on one understanding task and three generation tasks for commits. The comprehensive experiments on these tasks demonstrate that CommitBART significantly outperforms previous pre-trained works for code. Further analysis also reveals that each pre-training task enhances the model performance. Shangqing Liu, Yanzhou Li, Xiaofei Xie, Wei Ma 0014, Guozhu Meng, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2023 | Good-looking but Lacking Faithfulness: Understanding Local Explanation Methods through Trend-based TestingabstractWhile enjoying the great achievements brought by deep learning (DL), people are also worried about the decision made by DL models, since the high degree of non-linearity of DL models makes the decision extremely difficult to understand. Consequently, attacks such as adversarial attacks are easy to carry out, but difficult to detect and explain, which has led to a boom in the research on local explanation methods for explaining model decisions. In this paper, we evaluate the faithfulness of explanation methods and find that traditional tests on faithfulness encounter the random dominance problem, i.e., the random selection performs the best, especially for complex data. To further solve this problem, we propose three trend-based faithfulness tests and empirically demonstrate that the new trend tests can better assess faithfulness than traditional tests on image, natural language and security tasks. We implement the assessment system and evaluate ten popular explanation methods. Benefiting from the trend tests, we successfully assess the explanation methods on complex data for the first time, bringing unprecedented discoveries and inspiring future research. Downstream tasks also greatly benefit from the tests. For example, model debugging equipped with faithful explanation methods performs much better for detecting and correcting accuracy and security problems. Jinwen He, Kai Chen 0012, Guozhu Meng, Jiangshan Zhang, Congyi Li |
CCS | 3 |
| 2023 | FMDiv: Functional Module Division on Binary Malware for Accurate Malicious Code LocalizationabstractIn recent years, binary malware detection has attracted extensive attention from industry and academia. However, most of the existing work focuses on determining whether a sample is malicious or not, rather than identifying the malicious essence in malware. Few studies aim at locating malicious code at function granularity and suffer from inaccuracy. In this paper, we solve the problem by dividing malware into Functional Module (FM), which is a better granularity for locating malicious code, as it combines certain functions to express malicious behaviors in malware. We design a tool called FMDiv to automatically unpack and disassemble binary malware and then divide them into FMs based on the function call graph (CG). Meanwhile, one novel feature extraction and embedding method has been adopted to validate the effect of the FM division algorithm and provide one alternative method of characterization for subsequent malicious FM location. We evaluate FMDiv’s performance on 10,440 real-world samples from VIRUSSHARE. The results show that FMDiv can correctly characterize and make FM division of malware, outperforming current state-of-the-art work. Weihao Huang, Chaoyang Lin, Qiucun Yan, Lu Xiang, Zhiyu Zhang 0017, Guozhu Meng, Kai Chen 0012 |
CSCWD | 6 |
| 2023 | ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningabstractLarge-scale pre-trained models such as CodeBERT, GraphCodeBERT have earned widespread attention from both academia and industry. Attributed to the superior ability in code representation, they have been further applied in multiple downstream tasks such as clone detection, code search and code translation. However, it is also observed that these state-of-the-art pre-trained models are susceptible to adversarial attacks. The performance of these pre-trained models drops significantly with simple perturbations such as renaming variable names. This weakness may be inherited by their downstream models and thereby amplified at an unprecedented scale. To this end, we propose an approach namely ContraBERT that aims to improve the robustness of pre-trained models via contrastive learning. Specifically, we design nine kinds of simple and complex data augmentation operators on the programming language (PL) and natural language (NL) data to construct different variants. Furthermore, we continue to train the existing pre-trained models by masked language modeling (MLM) and contrastive pre-training task on the original samples with their augmented variants to enhance the robustness of the model. The extensive ex-periments demonstrate that ContraBERT can effectively improve the robustness of the existing pre-trained models. Further study also confirms that these robustness-enhanced models provide improvements as compared to original models over four popular downstream tasks. Shangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng, Yang Liu 0003 |
ICSE | 4 |
| 2023 | Fairness via Group Contribution MatchingabstractFairness issues in Deep Learning models have recently received increasing attention due to their significant societal impact. Although methods for mitigating unfairness are constantly proposed, little research has been conducted to understand how discrimination and bias develop during the standard training process. In this study, we propose analyzing the contribution of each subgroup (i.e., a group of data with the same sensitive attribute) in the training process to understand the cause of such bias development process. We propose a gradient-based metric to assess training subgroup contribution disparity, showing that unequal contributions from different subgroups are one source of such unfairness. One way to balance the contribution of each subgroup is through oversampling, which ensures that an equal number of samples are drawn from each subgroup during each training iteration. However, we have found that even with a balanced number of samples, the contribution of each group remains unequal, resulting in unfairness under the oversampling strategy. To address the above issues, we propose an easy but effective group contribution matching (GCM) method to match the contribution of each subgroup. Our experiments show that our GCM effectively improves fairness and outperforms other methods significantly. Tianlin Li, Anran Li 0001, Mengnan Du, Aishan Liu, Qing Guo 0005, Guozhu Meng, Yang Liu 0003 |
IJCAI | 7 |
| 2023 | Differential Testing of Cross Deep Learning Framework APIs: Revealing Inconsistencies and Vulnerabilities
Zizhuang Deng, Guozhu Meng, Kai Chen 0012, Tong Liu 0027, Lu Xiang, Chunyang Chen 0001 |
USENIX Security Symposium | 2 |
| 2023 | Aliasing Backdoor Attacks on Pre-trained Models
Cheng'an Wei, Yeonjoon Lee, Kai Chen 0012, Guozhu Meng, Peizhuo Lv |
USENIX Security Symposium | 4 |
| 2023 | SkillSim: voice apps similarity detectionabstractAbstract Virtual personal assistants (VPAs), such as Amazon Alexa and Google Assistant, are software agents designed to perform tasks or provide services to individuals in response to user commands. VPAs extend their functions through third-party voice apps, thereby attracting more users to use VPA-equipped products. Previous studies demonstrate vulnerabilities in the certification, installation, and usage of these third-party voice apps. However, these studies focus on individual apps. To the best of our knowledge, there is no prior research that explores the correlations among voice apps.Voice apps represent a new type of applications that interact with users mainly through a voice user interface instead of a graphical user interface, requiring a distinct approach to analysis. In this study, we present a novel voice app similarity analysis approach to analyze voice apps in the market from a new perspective. Our approach, called SkillSim, detects similarities among voice apps (i.e. skills) based on two dimensions: text similarity and structure similarity. SkillSim measures 30,000 voice apps in the Amazon skill market and reveals that more than 25.9% have at least one other skill with a text similarity greater than 70%. Our analysis identifies several factors that contribute to a high number of similar skills, including the assistant development platforms and their limited templates. Additionally, we observe interesting phenomena, such as developers or platforms creating multiple similar skills with different accounts for purposes such as advertising. Furthermore, we also find that some assistant development platforms develop multiple similar but non-compliant skills, such as requesting user privacy in a non-compliance way, which poses a security risk. Based on the similarity analysis results, we have a deeper understanding of voice apps in the mainstream market. Zhixiu Guo, Ruigang Liang, Guozhu Meng, Kai Chen 0012 |
Cybersecur. | 3 |
| 2023 | Are our clone detectors good enough? An empirical study of code effects by obfuscationabstractAbstract Clone detection has received much attention in many fields such as malicious code detection, vulnerability hunting, and code copyright infringement detection. However, cyber criminals may obfuscate code to impede violation detection. To date, few studies have investigated the robustness of clone detectors, especially in-fashion deep learning-based ones, against obfuscation. Meanwhile, most of these studies only measure the difference between one code snippet and its obfuscation version. However, in reality, the attackers may modify the original code before obfuscating it. Then what we should evaluate is the detection of obfuscated code from cloned code, not the original code. For this, we conduct a comprehensive study evaluating 3 popular deep-learning based clone detectors and 6 commonly used traditional ones. Regarding the data, we collect 6512 clone pairs of five types from the dataset BigCloneBench and obfuscate one program of each pair via 64 strategies of 6 state-of-art commercial obfuscators. We also collect 1424 non-clone pairs to evaluate the false positives. In sum, a benchmark of 524,148 code pairs (either clone or not) are generated, which are passed to clone detectors for evaluation. To automate the evaluation, we develop one uniform evaluation framework, integrating the clone detectors and obfuscators. The results bring us interesting findings on how obfuscation affects the performance of clone detection and what is the difference between traditional and deep learning-based clone detectors. In addition, we conduct manual code reviews to uncover the root cause of the phenomenon and give suggestions to users from different perspectives. Weihao Huang, Guozhu Meng, Chaoyang Lin, Qiucun Yan, Kai Chen 0012, Zhuo Ma 0001 |
Cybersecur. | 2 |
| 2023 | GraphSearchNet: Enhancing GNNs via Capturing Global Dependencies for Semantic Code SearchabstractCode search aims to retrieve accurate code snippets based on a natural language query to improve software productivity and quality. With the massive amount of available programs such as (on GitHub or Stack Overflow), identifying and localizing the precise code is critical for the software developers. In addition, Deep learning has recently been widely applied to different code-related scenarios, e.g., vulnerability detection, source code summarization. However, automated deep code search is still challenging since it requires a high-level semantic mapping between code and natural language queries. Most existing deep learning-based approaches for code search rely on the sequential text i.e., feeding the program and the query as a flat sequence of tokens to learn the program semantics while the structural information is not fully considered. Furthermore, the widely adopted Graph Neural Networks (GNNs) have proved their effectiveness in learning program semantics, however, they also suffer the problem of capturing the global dependencies in the constructed graph, which limits the model learning capacity. To address these challenges, in this paper, we design a novel neural network framework, named GraphSearchNet, to enable an effective and accurate source code search by jointly learning the rich semantics of both source code and natural language queries. Specifically, we propose to construct graphs for the source code and queries with bidirectional GGNN (BiGGNN) to capture the local structural information of the source code and queries. Furthermore, we enhance BiGGNN by utilizing the multi-head attention module to supplement the global dependencies that BiGGNN missed to improve the model learning capacity. The extensive experiments on Java and Python programming language from the public benchmark CodeSearchNet confirm that GraphSearchNet outperforms current state-of-the-art works by a significant margin. Shangqing Liu, Xiaofei Xie, Jing Kai Siow, Lei Ma 0003, Guozhu Meng, Yang Liu 0003 |
IEEE Trans. Software Eng. | 5 |
| 2022 | Understanding Real-world Threats to Deep Learning Models in Android AppsabstractFamous for its superior performance, deep learning (DL) has been popularly used within many applications, which also at the same time attracts various threats to the models. One primary threat is from adversarial attacks. Researchers have intensively studied this threat for several years and proposed dozens of approaches to create adversarial examples (AEs). But most of the approaches are only evaluated on limited models and datasets (e.g., MNIST, CIFAR-10). Thus, the effectiveness of attacking real-world DL models is not quite clear. In this paper, we perform the first systematic study of adversarial attacks on real-world DNN models and provide a real-world model dataset named RWM. Particularly, we design a suite of approaches to adapt current AE generation algorithms to the diverse real-world DL models, including automatically extracting DL models from Android apps, capturing the inputs and outputs of the DL models in apps, generating AEs and validating them by observing the apps' execution. For black-box DL models, we design a semantic-based approach to build suitable datasets and use them for training substitute models when performing transfer-based attacks. After analyzing 245 DL models collected from 62,583 real-world apps, we have a unique opportunity to understand the gap between real-world DL models and contemporary AE generation algorithms. To our surprise, the current AE generation algorithms can only directly attack 6.53% of the models. Benefiting from our approach, the success rate upgrades to 47.35%. Zizhuang Deng, Kai Chen 0012, Guozhu Meng, Xiaodong Zhang 0014 |
CCS | 3 |
| 2022 | Detecting API Missing-Check Bugs Through Complete Cross Checking of Erroneous Returns
Qintao Shen, Guozhu Meng, Kai Chen 0012, Yuqing Zhang 0001 |
Inscrypt | 3 |
| 2022 | LibHunter: An Unsupervised Approach for Third-party Library Detection without Prior KnowledgeabstractThird-party libraries (TPLs) are a significant component of mobile apps. They provide various functionalities, and developers employ them to facilitate app development. TPL detection is a fundamental task in security research, as it can impact other security studies. TPL can act as an assistant to malware detection, privacy leakage detection, etc. Because if a TPL carries malicious code, all apps that integrate the TPL can be considered risky. However, in some studies, TPLs can also act as noise, like app traffic fingerprinting. The TPL and app traffic are mixed during app runtime, making it difficult to fingerprint the app traffic accurately. Unfortunately, all existing TPL detection studies are working with prior knowledge of TPLs, as they need a whitelist or a train on known TPLs. However, new TPLs keep emerging, and it is not feasible for existing works to identify them-especially those who have network behaviors, as they may transfer inappropriate contents in the network. To this end, we propose LibHunter - an approach to identify TPLs without prior knowledge. LibHunter inspects the HTTP(S) traffic, logs the corresponding code execution traces, extracts features from the collected data, and performs a clustering algorithm to obtain TPLs. We apply LibHunter to 3000 apps. Results demonstrate that LibHunter can identify 79 TPLs, and about 60% of them are not detected by all existing works. We perform an analysis to show how important these TPLs are; we also present the visiting graph of these TPLs. Our findings bring light to the research community that existing tools are not accurate when encountering contemporary apps. Huajun Cui, Guozhu Meng, Yuejun Li, Yan Zhang 0014, Jiyan Sun, Dali Zhu, Weiping Wang 0005 |
ISCC | 2 |
| 2022 | Automated Privacy Network Traffic Detection via Self-labeling and LearningabstractWith the increasing popularity of mobile devices, privacy leakage has become more and more serious. The inappropriate behaviors of mobile APPs have brought substantial security risks to the public (e.g., location leakage). Existing solutions detect privacy leakage based on network traffic analysis. However, they can only detect unencrypted traffic, which leads to failures in the face of encrypted traffic. To solve this challenge, we designed an Automated Privacy Traffic Detection system (APTD). APTD can automatically generate self-labeling privacy traffic datasets, learn to identify the encrypted privacy traffic, and accurately assess the risk of privacy leakage. Due to its automation capability, APTD can directly support privacy leakage detection for newly-emerged applications without any system changes. To comprehensively evaluate APTD, we conducted an experiment on 2327 real-world mobile APPs. APTD automatically generated a labeled dataset containing 27343 real-world encrypted traffic traces. Based on the dataset, APTD identifies privacy traffic, and performs a privacy leakage risk assessment of APPs. The results show that APTD achieves 97% accuracy and 99% recall on our dataset and identifies 12 APPs that transmit high-risk privacy data. Yuejun Li, Huajun Cui, Jiyan Sun, Yan Zhang 0014, Guozhu Meng, Weiping Wang 0005 |
ISCC | 6 |
| 2022 | TransRepair: Context-aware Program Repair for Compilation ErrorsabstractAutomatically fixing compilation errors can greatly raise the productivity of software development, by guiding the novice or AI programmers to write and debug code. Recently, learning-based program repair has gained extensive attention and became the state-of-the-art in practice. But it still leaves plenty of space for improvement. In this paper, we propose an end-to-end solution TransRepair to locate the error lines and create the correct substitute for a C program simultaneously. Superior to the counterpart, our approach takes into account the context of erroneous code and diagnostic compilation feedback. Then we devise a Transformer-based neural network to learn the ways of repair from the erroneous code as well as its context and the diagnostic feedback. To increase the effectiveness of TransRepair, we summarize 5 types and 74 fine-grained sub-types of compilations errors from two real-world program datasets and the Internet. Then a program corruption technique is developed to synthesize a large dataset with 1,821,275 erroneous C programs. Through the extensive experiments, we demonstrate that TransRepair outperforms the state-of-the-art in both single repair accuracy and full repair accuracy. Further analysis sheds light on the strengths and weaknesses in the contemporary solutions for future improvement. Shangqing Liu, Guozhu Meng, Xiaofei Xie, Kai Chen 0012, Yang Liu 0003 |
ASE | 4 |
| 2022 | Learning Program Semantics with Code Representations: An Empirical StudyabstractProgram semantics learning is the core and fundamental for various code intelligent tasks e.g., vulnerability detection, clone detection. A considerable amount of existing works propose diverse approaches to learn the program semantics for different tasks and these works have achieved state-of-the-art performance. However, currently, a comprehensive and systematic study on evaluating different program representation techniques across diverse tasks is still missed. From this starting point, in this paper, we conduct an empirical study to evaluate different program representation techniques. Specifically, we categorize current mainstream code representation techniques into four categories i.e., Feature-based, Sequence-based, Tree-based, and Graph-based program representation technique and evaluate its performance on three diverse and popular code intelligent tasks i.e., Code Classification, Vulnerability Detection, and Clone Detection on the public released benchmark. We further design three research questions (RQs) and conduct a comprehensive analysis to investigate the performance. By the extensive experimental results, we conclude that (1) The graph-based representation is superior to the other selected techniques across these tasks. (2) Compared with the node type information used in tree-based and graph-based representations, the node textual information is more critical to learning the program semantics. (3) Different tasks require the task-specific semantics to achieve their highest performance, however combining various program semantics from different dimensions such as control dependency, data dependency can still produce promising results. Jing Kai Siow, Shangqing Liu, Xiaofei Xie, Guozhu Meng, Yang Liu 0003 |
SANER | 4 |
| 2022 | The inconsistency of documentation: a study of online C standard library documentsabstractAbstract The C standard libraries are basic function libraries standardized by the C language. Programmers usually refer to their API documentation provided by third-party websites. Unfortunately, these documents are not necessarily complete or accurate, especially for constraint sentences of API usage, which are called Security Specifications (SSs). SS issues can prevent programmers from following obligatory constraints, which results in API misuse vulnerabilities. Previous work studying SS issues could only find certain types of inaccurate SSs through checking the compliance between API usage and existing SSs. Therefore, we propose a novel approach SSeeker for quickly discovering missing and inaccurate SSs through the inconsistency of semantically similar SSs. More specifically, SSeeker first completes broken sentences and discovers SSs from them by judging their constraint sentiment. Then SSeeker puts semantically similar SSs from different sources into a group, which can be used to discover missing or inaccurate SSs. With the help of SSeeker, we investigated 4 popular online third-party C standard library documents, studied their conformity with the C99 standard, analyzed their APIs and SSs, and discovered 92 prototype issues, 15 web page issues, and 96 SS issues. Ruishi Li, Yunfei Yang 0001, Peiwei Hu, Guozhu Meng |
Cybersecur. | 5 |
| 2022 | Towards Security Threats of Deep Learning Systems: A SurveyabstractDeep learning has gained tremendous success and great popularity in the past few years. However, deep learning systems are suffering several inherent weaknesses, which can threaten the security of learning models. Deep learning’s wide use further magnifies the impact and consequences. To this end, lots of research has been conducted with the purpose of exhaustively identifying intrinsic weaknesses and subsequently proposing feasible mitigation. Yet few are clear about how these weaknesses are incurred and how effective these attack approaches are in assaulting deep learning. In order to unveil the security weaknesses and aid in the development of a robust deep learning system, we undertake an investigation on attacks towards deep learning, and analyze these attacks to conclude some findings in multiple views. In particular, we focus on four types of attacks associated with security threats of deep learning: model extraction attack, model inversion attack, poisoning attack and adversarial attack. For each type of attack, we construct its essential workflow as well as adversary capabilities and attack goals. Pivot metrics are devised for comparing the attack approaches, by which we perform quantitative and qualitative analyses. From the analysis, we have identified significant and indispensable factors in an attack vector, e.g., how to reduce queries to target models, what distance should be used for measuring perturbation. We shed light on 18 findings covering these approaches’ merits and demerits, success probability, deployment complexity and prospects. Moreover, we discuss other potential security weaknesses and possible mitigation which can inspire relevant research in this area. Yingzhe He, Guozhu Meng, Kai Chen 0012, Xingbo Hu, Jinwen He |
IEEE Trans. Software Eng. | 2 |
| 2021 | Why is Your Trojan NOT Responding? A Quantitative Analysis of Failures in Backdoor Attacks of Neural Networks
Xingbo Hu, Yibing Lan, Ruimin Gao, Guozhu Meng, Kai Chen 0012 |
ICA3PP (3) | 4 |
| 2021 | Vall-nut: Principled Anti-Grey box - FuzzingabstractGreybox fuzzing is a widely used technique for software testing that has been adopted by practitioners and researchers to disclose a great number of vulnerabilities in various software. However, adversaries also weaponize greybox fuzzing to mine vulnerabilities for malicious intentions. This poses considerable threats to software systems. To counteract the misuse of greybox fuzzing, we propose VALL-NUT, a novel approach to harden software with properties to combat greybox fuzzing. We dissect the major strategies that facilitate the success of greybox fuzzing, and accordingly propose three types of neutralizing schemesseed queue explosion, seed attenuation, and feedback contamination. We evaluate Vall-nut against the mainstream greybox fuzzers on multiple real-world benchmark programs. The results show that Vall-nut can reduce an average of 34 % code coverage and 76% detected crashes in 24-hour tests. Moreover, we conduct comparisons with two recent studies which show Vall-nut can achieve a superior deduction of detected crashes. Yuekang Li, Guozhu Meng, Jun Xu 0024, Cen Zhang, Hongxu Chen 0001, Xiaofei Xie, Haijun Wang 0002, Yang Liu 0003 |
ISSRE | 2 |
| 2021 | DRMI: A Dataset Reduction Technology based on Mutual Information for Black-box Attacks
Yingzhe He, Guozhu Meng, Kai Chen 0012, Xingbo Hu, Jinwen He |
USENIX Security Symposium | 2 |
| 2021 | Have You been Properly Notified? Automatic Compliance Analysis of Privacy Policy Text with GDPR Article 13abstractWith the rapid development of web and mobile applications, as well as their wide adoption in different domains, more and more personal data is provided, consciously or unconsciously, to different application providers. Privacy policy is an important medium for users to understand what personal information has been collected and used. As data privacy protection is becoming a critical social issue, there are laws and regulations being enacted in different countries and regions, and the most representative one is the EU General Data Protection Regulation (GDPR). It is thus important to detect compliance issues among regulations, e.g., GDPR, with privacy policies, and provide intuitive results for data subjects (i.e., users), data collection party (i.e., service providers) and the regulatory authorities. In this work, we target to solve the problem of compliance analysis between GDPR (Article 13) and privacy policies. We format the task into a combination of a sentence classification step and a rule-based analysis step. We manually curate a corpus of 36,610 labeled sentences from 304 privacy policies, and benchmark our corpus with several standard sentence classifiers. We also conduct a rule-based analysis to detect compliance issues and a user study to evaluate the usability of our approach. The web-based tool AutoCompliance is publicly accessible 1. Shuang Liu 0007, Baiyang Zhao, Renjie Guo, Guozhu Meng, Meishan Zhang |
WWW | 4 |
| 2021 | SEPAL: Towards a Large-scale Analysis of SEAndroid Policy CustomizationabstractNowadays, SEAndroid has been widely deployed in Android devices to enforce security policies and provide flexible mandatory access control (MAC), for the purpose of narrowing down attack surfaces and restricting risky operations. Generally, the original SEAndroid security policy rules are carefully and strictly written and maintained by the Android community. However, in practice, mobile device manufacturers usually have to customize these policy rules and add their own new rules to satisfy their functionality extensions, which breaks the integrity of SEAndroid and causes serious security issues. Still, up to now, it is a challenging task to identify these security issues due to the large and ever-increasing number of policy rules, as well as the complexity of policy semantics. Dongsong Yu, Guangliang Yang 0001, Guozhu Meng, Xiaorui Gong, Xiaobo Xiang, Kai Chen 0012, Wenke Lee, Wenchang Shi |
WWW | 3 |
| 2021 | A Performance-Sensitive Malware Detection System Using Deep Learning on Mobile DevicesabstractCurrently, Android malware detection is mostly performed on server side against the increasing number of malware. Powerful computing resource provides more exhaustive protection for app markets than maintaining detection by a single user. However, apart from the applications (apps) provided by the official market (i.e., Google Play Store), apps from unofficial markets and third-party resources are always causing serious security threats to end-users. Meanwhile, it is a time-consuming task if the app is downloaded first and then uploaded to the server side for detection, because the network transmission has a lot of overhead. In addition, the uploading process also suffers from the security threats of attackers. Consequently, a last line of defense on mobile devices is necessary and much-needed. In this paper, we propose an effective Android malware detection system, MobiTive, leveraging customized deep neural networks to provide a real-time and responsive detection environment on mobile devices. MobiTive is a pre-installed solution rather than an app scanning and monitoring engine using after installation, which is more practical and secure. Although a deep learning-based approach can be maintained on server side efficiently for malware detection, original deep learning models cannot be directly deployed and executed on mobile devices due to various performance limitations, such as computation power, memory size, and energy. Therefore, we evaluate and investigate the following key points: (1) the performance of different feature extraction methods based on source code or binary code; (2) the performance of different feature type selections for deep learning on mobile devices; (3) the detection accuracy of different deep neural networks on mobile devices; (4) the real-time detection performance and accuracy on different mobile devices; (5) the potential based on the evolution trend of mobile devices' specifications; and finally we further propose a practical solution (MobiTive) to detect Android malware on mobile devices. Sen Chen 0001, Xiaofei Xie, Guozhu Meng, Shangwei Lin 0001, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | An empirical assessment of security risks of global Android banking appsabstractMobile banking apps, belonging to the most security-critical app category, render massive and dynamic transactions susceptible to security risks. Given huge potential financial loss caused by vulnerabilities, existing research lacks a comprehensive empirical study on the security risks of global banking apps to provide useful insights and improve the security of banking apps. Sen Chen 0001, Lingling Fan 0003, Guozhu Meng, Ting Su 0001, Minhui Xue 0001, Yinxing Xue, Yang Liu 0003, Lihua Xu |
ICSE | 3 |
| 2020 | A large-scale empirical study on vulnerability distribution within projects and the lessons learnedabstractThe number of vulnerabilities increases rapidly in recent years, due to advances in vulnerability discovery solutions. It enables a thorough analysis on the vulnerability distribution and provides support for correlation analysis and prediction of vulnerabilities. Previous research either focuses on analyzing bugs rather than vulnerabilities, or only studies general vulnerability distribution among projects rather than the distribution within each project. In this paper, we collected a large vulnerability dataset, consisting of all known vulnerabilities associated with five representative open source projects, by utilizing automated crawlers and spending months of manual efforts. We then analyzed the vulnerability distribution within each project over four dimensions, including files, functions, vulnerability types and responsible developers. Based on the results analysis, we presented 12 practical insights on the distribution of vulnerabilities. Finally, we applied such insights on several vulnerability discovery solutions (including static analysis and dynamic fuzzing), and helped them find 10 zero-day vulnerabilities in target projects, showing that our insights are useful. Bingchang Liu, Guozhu Meng, Feng Li 0045, Dandan Sun, Wei Huo 0005, Chao Zhang 0008 |
ICSE | 2 |
| 2020 | A3Ident: A Two-phased Approach to Identify the Leading Authors of Android AppsabstractAuthorship identification is the process of identifying and classifying authors through given codes. Authorship identification can be used in a wide range of software domains, e.g., code authorship disputes, plagiarism detection, exposure of attackers’ identity. Besides the inherent challenges from legacy software development, framework programming and crowdsourcing mode in Android raise the difficulties of authorship identification significantly. More specifically, widespread third party libraries and inherited components (e.g., classes, methods, and variables) dilute the primary code within the entire Android app and blur the boundaries of code written by different authors. However, prior research has not well addressed these challenges.To this end, we design a two-phased approach to attribute the primary code of an Android app to the specific developer. In the first phase, we put forward three types of strategies to identify the relationships between Java packages in an app, which consist of context, semantic and structural relationships. A package aggregation algorithm is developed to cluster all packages that are of high probability written by the same authors. In the second phase, we develop three types of features to capture authors’ coding habits and code stylometry. Based on that, we generate fingerprints for an author from its developed Android apps and employ several machine learning algorithms for authorship classification. We evaluate our approach in three datasets that contain 15,666 apps from 257 distinct developers and achieve a 92.5% accuracy rate on average. Additionally, we test it on 2,900 obfuscated apps and our approach can classify apps with an accuracy rate of 80.4%. Wei Wang 0277, Guozhu Meng, Haoyu Wang 0001, Kai Chen 0012, Weimin Ge, Xiaohong Li 0001 |
ICSME | 2 |
| 2020 | Large-Scale Empirical Studies on Effort-Aware Security Vulnerability Prediction MethodsabstractSecurity vulnerability prediction (SVP) can identify potential vulnerable modules in advance and then help developers to allocate most of the test resources to these modules. To evaluate the performance of different SVP methods, we should take the security audit and code inspection into account and then consider effort-aware performance measures (such as ACC and Popt). However, to the best of our knowledge, the effectiveness of different SVP methods has not been thoroughly investigated in terms of effort-aware performance measures. In this article, we consider 48 different SVP methods, of which 36 are supervised methods and 12 are unsupervised methods. For the supervised methods, we consider 34 software-metric-based methods and two text-mining-based methods. For the software-metric-based methods, in addition to a large number of classification methods, we also consider four state-of-the-art methods (i.e., EALR, OneWay, CBS, and MULTI) proposed in recent effort-aware just-in-time defect prediction studies. For text-mining-based methods, we consider the Bag-of-Word model and the term-frequency-inverse-document-frequency model. For the unsupervised methods, all the modules are ranked in the ascendent order based on a specific metric. Since 12 software metrics are considered when measuring extracted modules, there are 12 different unsupervised methods. To the best of our knowledge, over 40 SVP methods have not been considered in previous SVP studies. In our large-scale empirical studies, we use three real open-source web applications written in PHP as benchmark. These three web applications include 3466 modules and 223 vulnerabilities in total. We evaluate these SVP methods both in the within-project SVP scenario and the cross-project SVP scenario. Empirical results show that two unsupervised methods [i.e., lines of code (LOC) and Halstead's volume (HV)] and four recently proposed state-of-the-art supervised methods (i.e., MULTI, OneWay, CBS, and EALR) can achieve better performance than the other methods in terms of effort-aware performance measures. Then, we analyze the reasons why these six methods can achieve better performance. For example, when using 20% of the entire efforts, we find that these six methods always require more modules to be inspected, especially for unsupervised methods LOC and HV. Finally, from the view of practical vulnerability localization, we find that all the unsupervised methods and the OneWay method have high false alarms before finding the first vulnerable module. This may have an impact on developers' confidence and tolerance, and supervised methods (especially MULTI and text-mining-based methods) are preferred. Xiang Chen 0005, Yingquan Zhao, Zhanqi Cui, Guozhu Meng, Yang Liu 0003 |
IEEE Trans. Reliab. | 4 |
| 2019 | RoLMA: A Practical Adversarial Attack Against Deep Learning-Based LPR Systems
Mingming Zha 0001, Guozhu Meng, Chaoyang Lin, Zhe Zhou 0001, Kai Chen 0012 |
Inscrypt | 2 |
| 2019 | MobiDroid: A Performance-Sensitive Malware Detection System on Mobile PlatformabstractCurrently, Android malware detection is mostly performed on the server side against the increasing number of Android malware. Powerful computing resource gives more exhaustive protection for Android markets than maintaining detection by a single user in many cases. However, apart from the Android apps provided by the official market (i.e., Google Play Store), apps from unofficial markets and third-party resources are always causing a serious security threat to end-users. Meanwhile, it is a time-consuming task if the app is downloaded first and then uploaded to the server side for detection because the network transmission has a lot of overhead. In addition, the uploading process also suffers from the threat of attackers. Consequently, a last line of defense on Android devices is necessary and much-needed. To address these problems, in this paper, we propose an effective Android malware detection system, MobiDroid, leveraging deep learning to provide a real-time secure and fast response environment on Android devices. Although a deep learning-based approach can be maintained on server side efficiently for detecting Android malware, deep learning models cannot be directly deployed and executed on Android devices due to various performance limitations such as computation power, memory size, and energy. Therefore, we evaluate and investigate the different performances with various feature categories, and further provide an effective solution to detect malware on Android devices. The proposed detection system on Android devices in this paper can serve as a starting point for further study of this important area. Sen Chen 0001, Xiaofei Xie, Lei Ma 0003, Guozhu Meng, Yang Liu 0003, Shangwei Lin 0001 |
ICECCS | 5 |
| 2019 | Characterizing Android App Signing IssuesabstractIn the app releasing process, Android requires all apps to be digitally signed with a certificate before distribution. Android uses this certificate to identify the author and ensure the integrity of an app. However, a number of signature issues have been reported recently, threatening the security and privacy of Android apps. In this paper, we present the first large-scale systematic measurement study on issues related to Android app signatures. We first create a taxonomy covering four types of app signing issues (21 anti-patterns in total), including vulnerabilities, potential attacks, release bugs and compatibility issues. Then we developed an automated tool to characterize signature-related issues in over 5 million app items (3 million distinct apks) crawled from Google Play and 24 alternative Android app markets. Our empirical findings suggest that although Google has introduced apk-level signing schemes (V2 and V3) to overcome some of the known security issues, more than 93% of the apps still use only the JAR signing scheme (V1), which poses great security threats. Besides, we also revealed that 7% to 45% of the apps in the 25 studied markets have been found containing at least one signing issue, while a large number of apps have been exposed to security vulnerabilities, attacks and compatibility issues. Among them a considerable number of apps we identified are popular apps with millions of downloads. Finally, our evolution analysis suggested that most of the issues were not mitigated after a considerable amount of time across markets. The results shed light on the emergency for detecting and repairing the app signing issues. Haoyu Wang 0001, Hongxuan Liu, Xusheng Xiao, Guozhu Meng, Yao Guo 0001 |
ASE | 4 |
| 2019 | Securing android applications via edge assistant third-party library detection
Zhushou Tang, Minhui Xue 0001, Guozhu Meng, Chengguo Ying, Yugeng Liu, Jianan He, Haojin Zhu, Yang Liu 0003 |
Comput. Secur. | 3 |
| 2019 | Securing Android App Markets via Modeling and Predicting Malware Spread Between MarketsabstractThe Android ecosystem has recently dominated mobile devices. Android app markets, including official Google Play and other third party markets, are becoming hotbeds, where malware originates and spreads. Android malware has been observed to both propagate within markets and spread between markets. If the spread of Android malware between markets can be predicted, market administrators can take appropriate measures to prevent the outbreak of malware and minimize the damages caused by malware. In this paper, we make the first attempt to protect the Android ecosystem by modeling and predicting the spread of Android malware between markets. To this end, we study the social behaviors that affect the spread of malware, model these spread behaviors with multiple epidemic models, and predict the infection time and order among markets for well-known malware families. To achieve an accurate prediction of malware spread, we model spread behaviors in the following fashion: 1) for a single market, we model the within-market malware growth by considering both the creation and removal of malware; 2) for multiple markets, we determine market relevance by calculating the mutual information among them; and 3) based on the previous two steps, we simulate a susceptible infected model stochastically for spread among markets. The model inference is performed using a publicly available well-labeled dataset AndRadar. To conduct extensive experiments to evaluate our approach, we collected a large number (334,782) of malware samples from 25 Android markets around the world. The experimental results show our approach can depict and simulate the growth of Android malware on a large scale, and predict the infection time and order among markets with 0.89 and 0.66 precision, respectively. Guozhu Meng, Matthew Patrick, Yinxing Xue, Yang Liu 0003, Jie Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | SoProtector: Securing Native C/C++ Libraries for Mobile Applications
Guangquan Xu, Guozhu Meng, James Xi Zheng |
ICA3PP (3) | 3 |
| 2018 | From UI design image to GUI skeleton: a neural machine translator to bootstrap mobile GUI implementationabstractA GUI skeleton is the starting point for implementing a UI design image. To obtain a GUI skeleton from a UI design image, developers have to visually understand UI elements and their spatial layout in the image, and then translate this understanding into proper GUI components and their compositions. Automating this visual understanding and translation would be beneficial for bootstraping mobile GUI implementation, but it is a challenging task due to the diversity of UI designs and the complexity of GUI skeletons to generate. Existing tools are rigid as they depend on heuristically-designed visual understanding and GUI generation rules. In this paper, we present a neural machine translator that combines recent advances in computer vision and machine translation for translating a UI design image into a GUI skeleton. Our translator learns to extract visual features in UI images, encode these features' spatial layouts, and generate GUI skeletons in a unified neural network framework, without requiring manual rule development. For training our translator, we develop an automated GUI exploration method to automatically collect large-scale UI data from real-world applications. We carry out extensive experiments to evaluate the accuracy, generality and usefulness of our approach. Chunyang Chen 0001, Ting Su 0001, Guozhu Meng, Zhenchang Xing, Yang Liu 0003 |
ICSE | 3 |
| 2018 | Large-scale analysis of framework-specific exceptions in Android appsabstractMobile apps have become ubiquitous. For app developers, it is a key priority to ensure their apps' correctness and reliability. However, many apps still suffer from occasional to frequent crashes, weakening their competitive edge. Large-scale, deep analyses of the characteristics of real-world app crashes can provide useful insights to guide developers, or help improve testing and analysis tools. However, such studies do not exist --- this paper fills this gap. Over a four-month long effort, we have collected 16,245 unique exception traces from 2,486 open-source Android apps, and observed that framework-specific exceptions account for the majority of these crashes. We then extensively investigated the 8,243 framework-specific exceptions (which took six person-months): (1) identifying their characteristics (e.g., manifestation locations, common fault categories), (2) evaluating their manifestation via state-of-the-art bug detection techniques, and (3) reviewing their fixes. Besides the insights they provide, these findings motivate and enable follow-up research on mobile apps, such as bug detection, fault localization and patch generation. In addition, to demonstrate the utility of our findings, we have optimized Stoat, a dynamic testing tool, and implemented ExLocator, an exception localization tool, for Android apps. Stoat is able to quickly uncover three previously-unknown, confirmed/fixed crashes in Gmail and Google+; ExLocator is capable of precisely locating the root causes of identified exceptions in real-world apps. Our substantial dataset is made publicly available to share with and benefit the community. Lingling Fan 0003, Ting Su 0001, Sen Chen 0001, Guozhu Meng, Yang Liu 0003, Lihua Xu, Geguang Pu, Zhendong Su 0001 |
ICSE | 4 |
| 2018 | Efficiently manifesting asynchronous programming errors in Android appsabstractAndroid, the #1 mobile app framework, enforces the single-GUI-thread model, in which a single UI thread manages GUI rendering and event dispatching. Due to this model, it is vital to avoid blocking the UI thread for responsiveness. One common practice is to offload long-running tasks into async threads. To achieve this, Android provides various async programming constructs, and leaves evelopers themselves to obey the rules implied by the model. However, as our study reveals, more than 25% apps violate these rules and introduce hard-to-detect, fail-stop errors, which we term as aysnc programming errors (APEs). To this end, this paper introduces APEChecker, a technique to automatically and efficiently manifest APEs. The key idea is to characterize APEs as specific fault patterns, and synergistically combine static analysis and dynamic UI exploration to detect and verify such errors. Among the 40 real-world Android apps, APEChecker unveils and processes 61 APEs, of which 51 are confirmed (83.6% hit rate). Specifically, APEChecker detects 3X more APEs than the state-of-art testing tools (Monkey, Sapienz and Stoat), and reduces testing time from half an hour to a few minutes. On a specific type of APEs, APEChecker confirms 5X more errors than the data race detection tool, EventRacer, with very few false alarms. Lingling Fan 0003, Ting Su 0001, Sen Chen 0001, Guozhu Meng, Yang Liu 0003, Lihua Xu, Geguang Pu |
ASE | 4 |
| 2018 | Are mobile banking apps secure? what can be improved?abstractMobile banking apps, as one of the most contemporary FinTechs, have been widely adopted by banking entities to provide instant financial services. However, our recent work discovered thousands of vulnerabilities in 693 banking apps, which indicates these apps are not as secure as we expected. This motivates us to conduct this study for understanding the current security status of them. First, we take 6 months to track the reporting and patching procedure of these vulnerabilities. Second, we audit 4 state-of the-art vulnerability detection tools on those patched vulnerabilities. Third, we discuss with 7 banking entities via in-person or online meetings and conduct an online survey to gain more feedback from financial app developers. Through this study, we reveal that (1) people may have inconsistent understandings of the vulnerabilities and different criteria for rating severity; (2) state-of-the-art tools are not effective in detecting vulnerabilities that the banking entities most concern; and (3) more efforts should be endeavored in different aspects to secure banking apps. We believe our study can help bridge the existing gaps, and further motivate different parties, including banking entities, researchers and policy makers, to better tackle security issues altogether. Sen Chen 0001, Ting Su 0001, Lingling Fan 0003, Guozhu Meng, Minhui Xue 0001, Yang Liu 0003, Lihua Xu |
ESEC/SIGSOFT FSE | 4 |
| 2018 | DroidEcho: an in-depth dissection of malicious behaviors in Android applicationsabstractA precise representation for attacks can benefit the detection of malware in both accuracy and efficiency. However, it is still far from expectation to describe attacks precisely on the Android platform. In addition, new features on Android, such as communication mechanisms, introduce new challenges and difficulties for attack detection. In this paper, we propose abstract attack models to precisely capture the semantics of various Android attacks, which include the corresponding targets, involved behaviors as well as their execution dependency. Meanwhile, we construct a novel graph-based model called the inter-component communication graph (ICCG) to describe the internal control flows and inter-component communications of applications. The models take into account more communication channel with a maximized preservation of their program logics. With the guidance of the attack models, we propose a static searching approach to detect attacks hidden in ICCG. To reduce false positive rate, we introduce an additional dynamic confirmation step to check whether the detected attacks are false alarms. Experiments show that DroidEcho can detect attacks in both benchmark and real-world applications effectively and efficiently with a precision of 89.5%. Guozhu Meng, Guangdong Bai, Kai Chen 0012, Yang Liu 0003 |
Cybersecur. | 1 |
| 2017 | Mining implicit design templates for actionable code reuseabstractIn this paper, we propose an approach to detecting project-specific recurring designs in code base and abstracting them into design templates as reuse opportunities. The mined templates allow programmers to make further customization for generating new code. The generated code involves the code skeleton of recurring design as well as the semi-implemented code bodies annotated with comments to remind programmers of necessary modification. We implemented our approach as an Eclipse plugin called MICoDe. We evaluated our approach with a reuse simulation experiment and a user study involving 16 participants. The results of our simulation experiment on 10 open source Java projects show that, to create a new similar feature with a design template, (1) on average 69% of the elements in the template can be reused and (2) on average 60% code of the new feature can be adopted from the template. Our user study further shows that, compared to the participants adopting the copy-paste-modify strategy, the ones using MICoDe are more effective to understand a big design picture and more efficient to accomplish the code reuse task. Yun Lin 0001, Guozhu Meng, Yinxing Xue, Zhenchang Xing, Jun Sun 0001, Xin Peng 0001, Yang Liu 0003, Wenyun Zhao, Jin Song Dong 0001 |
ASE | 2 |
| 2017 | Guided, stochastic model-based GUI testing of Android appsabstractMobile apps are ubiquitous, operate in complex environments and are developed under the time-to-market pressure. Ensuring their correctness and reliability thus becomes an important challenge. This paper introduces Stoat, a novel guided approach to perform stochastic model-based testing on Android apps. Stoat operates in two phases: (1) Given an app as input, it uses dynamic analysis enhanced by a weighted UI exploration strategy and static analysis to reverse engineer a stochastic model of the app's GUI interactions; and (2) it adapts Gibbs sampling to iteratively mutate/refine the stochastic model and guides test generation from the mutated models toward achieving high code and model coverage and exhibiting diverse sequences. During testing, system-level events are randomly injected to further enhance the testing effectiveness. Ting Su 0001, Guozhu Meng, Yuting Chen 0001, Geguang Pu, Yang Liu 0003, Zhendong Su 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2017 | Auditing Anti-Malware Tools by Evolving Android Malware and Dynamic Loading TechniqueabstractAlthough a previous paper shows that existing anti-malware tools (AMTs) may have high detection rate, the report is based on existing malware and thus it does not imply that AMTs can effectively deal with future malware. It is desirable to have an alternative way of auditing AMTs. In our previous paper, we use malware samples from android malware collection Genome to summarize a malware meta-model for modularizing the common attack behaviors and evasion techniques in reusable features. We then combine different features with an evolutionary algorithm, in which way we evolve malware for variants. Previous results have shown that the existing AMTs only exhibit detection rate of 20%-30% for 10 000 evolved malware variants. In this paper, based on the modularized attack features, we apply the dynamic code generation and loading techniques to produce malware, so that we can audit the AMTs at runtime. We implement our approach, named Mystique-S, as a service-oriented malware generation system. Mystique-S automatically selects attack features under various user scenarios and delivers the corresponding malicious payloads at runtime. Relying on dynamic code binding (via service) and loading (via reflection) techniques, Mystique-S enables dynamic execution of payloads on user devices at runtime. Experimental results on real-world devices show that existing AMTs are incapable of detecting most of our generated malware. Last, we propose the enhancements for existing AMTs. Yinxing Xue, Guozhu Meng, Yang Liu 0003, Tian Huat Tan, Hongxu Chen 0001, Jun Sun 0001, Jie Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Battery-Aware Mobile Data ServiceabstractSignificant research has been devoted to reduce the energy consumption of mobile devices, but how to increase their energy supply has received far less attention. Moreover, reducing the energy consumption alone does not always extend the device operation time due to a unique battery property - the capacity it delivers hinges critically upon how it is discharged. In this paper, we propose B-MODS, a novel design of battery-aware mobile data service on mobile devices. B-MODS constructs battery-friendly discharge patterns utilizing the recovery effect so as to increase the capacity delivered from batteries while meeting data service requirements. We implement B-MODS as an application layer library on the Android platform. Our experiments with diverse mobile devices under various application scenarios have shown that B-MODS increases the capacity delivery from the battery by up to 49.5 percent, with which an increase in the user-perceived data service utilities of up to 28.6 percent is observed. Liang He 0002, Guozhu Meng, Yu Gu 0001, Cong Liu 0005, Jun Sun 0001, Ting Zhu 0001, Yang Liu 0003, Kang G. Shin |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | Mystique: Evolving Android Malware for Auditing Anti-Malware ToolsabstractIn the arms race of attackers and defenders, the defense is usually more challenging than the attack due to the unpredicted vulnerabilities and newly emerging attacks every day. Currently, most of existing malware detection solutions are individually proposed to address certain types of attacks or certain evasion techniques. Thus, it is desired to conduct a systematic investigation and evaluation of anti-malware solutions and tools based on different attacks and evasion techniques. In this paper, we first propose a meta model for Android malware to capture the common attack features and evasion features in the malware. Based on this model, we develop a framework, MYSTIQUE, to automatically generate malware covering four attack features and two evasion features, by adopting the software product line engineering approach. With the help of MYSTIQUE, we conduct experiments to 1) understand Android malware and the associated attack features as well as evasion techniques; 2) evaluate and compare the 57 off-the-shelf anti-malware tools, 9 academic solutions and 4 App market vetting processes in terms of accuracy in detecting attack features and capability in addressing evasion. Last but not least, we provide a benchmark of Android malware with proper labeling of contained attack and evasion features. Guozhu Meng, Yinxing Xue, Mahinthan Chandramohan, Annamalai Narayanan, Yang Liu 0003, Jie Zhang 0002, Tieming Chen |
AsiaCCS | 1 |
| 2016 | Contextual Weisfeiler-Lehman graph kernel for malware detectionabstractIn this paper, we propose a novel graph kernel specifically to address a challenging problem in the field of cyber-security, namely, malware detection. Previous research has revealed the following: (1) Graph representations of programs are ideally suited for malware detection as they are robust against several attacks, (2) Besides capturing topological neighbourhoods (i.e., structural information) from these graphs it is important to capture the context under which the neighbourhoods are reachable to accurately detect malicious neighbourhoods. We observe that state-of-the-art graph kernels, such as Weisfeiler-Lehman kernel (WLK) capture the structural information well but fail to capture contextual information. To address this, we develop the Contextual Weisfeiler-Lehman kernel (CWLK) which is capable of capturing both these types of information. We show that for the malware detection problem, CWLK is more expressive and hence more accurate than WLK while maintaining comparable efficiency. Through our largescale experiments with more than 50,000 real-world Android apps, we demonstrate that CWLK outperforms two state-of-the-art graph kernels (including WLK) and three malware detection techniques by more than 5.27% and 4.87% F-measure, respectively, while maintaining high efficiency. This high accuracy and efficiency make CWLK suitable for large-scale real-world malware detection. Annamalai Narayanan, Guozhu Meng, Yang Liu 0003, Lihui Chen 0001 |
IJCNN | 2 |
| 2016 | Semantic modelling of Android malware for effective malware comprehension, detection, and classificationabstractMalware has posed a major threat to the Android ecosystem. Existing malware detection tools mainly rely on signature- or feature- based approaches, failing to provide detailed information beyond the mere detection. In this work, we propose a precise semantic model of Android malware based on Deterministic Symbolic Automaton (DSA) for the purpose of malware comprehension, detection and classification. It shows that DSA can capture the common malicious behaviors of a malware family, as well as the malware variants. Based on DSA, we develop an automatic analysis framework, named SMART, which learns DSA by detecting and summarizing semantic clones from malware families, and then extracts semantic features from the learned DSA to classify malware according to the attack patterns. We conduct the experiments in both malware benchmark and 223,170 real-world apps. The results show that SMART builds meaningful semantic models and outperforms both state-of-the-art approaches and anti-virus tools in malware detection. SMART identifies 4583 new malware in real-world apps that are missed by most anti-virus tools. The classification step further identifies new malware variants and unknown families. Guozhu Meng, Yinxing Xue, Zhengzi Xu, Yang Liu 0003, Jie Zhang 0002, Annamalai Narayanan |
ISSTA | 1 |
| 2013 | AUTHSCAN: Automatic Extraction of Web Authentication Protocols from Implementations
Guangdong Bai, Jike Lei, Guozhu Meng, Sai Sathyanarayan Venkatraman, Prateek Saxena, Jun Sun 0001, Yang Liu 0003, Jin Song Dong 0001 |
NDSS | 3 |