Xiao Chen 0002

dblp:05/3054-2 · DBLP profile ↗
← Back
33ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0002-4224-6450ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 19 · 18 since 2021Security and privacy · 8 · 1 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Empirical Study of Vulnerabilities in Python Packages and Their Detection
abstract
ContextIn the rapidly evolving software development landscape, Python stands out for its simplicity, versatility, and extensive ecosystem. Python packages, as units of organization, reusability, and distribution, have become a pressing concern, highlighted by the considerable number of vulnerability reports. As a scripting language, Python often cooperates with other programming languages for performance or interoperability. This also adds complexity to the vulnerabilities inherent to Python packages, and the effectiveness of current vulnerability detection tools remains underexplored in the research community.ObjectivesTo bridge this gap, we present PyVul, the first comprehensive benchmark suite of Pythonpackage vulnerabilities. We use this benchmark to conduct an empirical study that characterizes these vulnerabilities and evaluates the limitations of state-of-the-art detection tools.MethodsWe collect real-world vulnerability reports from GitHub Advisories, Snyk, and Huntr, and curate our benchmark at both the commit level and function level. To improve accuracy, we propose LLM-VDC, a large language model–assisted cleansing method. Based on PyVul, we systematically analyze vulnerabilities and assess the capabilities of both rule-based and machine learning–based detectors.ResultsAfter cleansing, PyVul achieves an accuracy of 100% at the commit level with 1,157 repository snapshots, and 94.0% at the function level with 2,082 vulnerable functions, establishing it as the most precise automatically collected Python vulnerability benchmark. Our empirical analysis reveals that current rule-based vulnerability detectors suffer from mismatches between their assumptions and real-world security scenarios, and limited support for high-order vulnerabilities, cross-language interactions, and Python’s unique language features. On the other hand, ML-based detectors suffer from their inability to reach the necessary context.ConclusionA significant discrepancy exists between the capabilities of existing tools and the demands of effectively identifying real-world security issues in Python packages. PyVul provides a solid foundation for advancing vulnerability research and tool development in this domain.
Haowei Quan, Junjie Wang 0007, Terry Yue Zhuo, Xiao Chen 0002, Xiaoning Du 0001
MSR5
2026 Systematic mapping study to assess security landscape for IoT-based smart farming systems
abstract
• The paper identifies the current trends in key security technologies for IoT-based smart farming through a systematic mapping study. • The paper assesses the technology readiness level (TRL) of each identified key security solution and the characteristics of a security product identified by ISO/IEC 25010. • Mapping of ISO/IEC 25010 security characteristics and Technology Readiness Level Smart farming systems sit at the intersection between three rapidly and independently advancing fields of IoT, Security, and Machine Learning. Its full realisation has tremendous positive impacts on food production; yet agricultural settings come with unique challenges that inhibit the rapid deployment of such state-of-the-art technologies. In this paper, we systematically study the current state of security for IoT-based smart farming research and development landscape and assess the proposed security solutions through the lens of technology readiness levels (TRL) and ISO/IEC 25010 security product evaluation framework. By analysing forty-eight primary studies, we identified the top security technologies under development, the critical security threats being addressed, and the most popularly used machine learning-based security solutions. Furthermore, we found that most of the ISO/IEC 25010 security characteristics considered by the security solutions are currently below TRL 6, indicating that they are well below the deployment readiness levels. Therefore, we recommend several supporting transitional technologies be developed to move the prototype development towards system validation and deployment to avoid the technology “valley of death”, such as farming-specific intrusion detection public datasets and large-scale IoT agriculture testbeds to validate the interoperability and transparency of security solutions at different layers. This systematic mapping study, together with a TRL assessment and ISO 25010 standard mapping, is the first of its kind, intending to provide a standardised comparison of the current state of security technologies for IoT-based smart farms to define a clear roadmap for future research and development. It provides a common terminology for the multidisciplinary stakeholders of smart farming to distinguish between theoretical security concepts and ready-to-deploy solutions, facilitating crucial decisions for investment, deployment, and commercialisation.
Farzana Zahid, Xiao Chen 0002, Shaleeza Sohail, Boyang Li 0005, Melanie Po-Leen Ooi
Comput. Secur.2
2026 DeepDesc: integrating retrieval-augmented generation with large language models for smart contract vulnerability detection
Tao Tan 0002, Xiao Chen 0002
Empir. Softw. Eng.2
2025 BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
abstract
Federated learning (FL) has been widely adopted as a decentralized training paradigm that enables multiple clients to collaboratively learn a shared model without exposing their local data. As concerns over data privacy and regulatory compliance grow, machine unlearning, which aims to remove the influence of specific data from trained models, has become increasingly important in the federated setting to meet legal, ethical, or user-driven demands. However, integrating unlearning into FL introduces new challenges and raises largely unexplored security risks. In particular, adversaries may exploit the unlearning process to compromise the integrity of the global model. In this paper, we present the first backdoor attack in the context of federated unlearning, demonstrating that an adversary can inject backdoors into the global model through seemingly legitimate unlearning requests. Specifically, we propose BadFU, an attack strategy where a malicious client uses both backdoor and camouflage samples to train the global model normally during the federated training process. Once the client requests unlearning of the camouflage samples, the global model transitions into a backdoored state. Extensive experiments under various FL frameworks and unlearning strategies validate the effectiveness of BadFU, revealing a critical vulnerability in current federated unlearning practices and underscoring the urgent need for more secure and robust federated unlearning mechanisms.
Bingguang Lu, Hongsheng Hu, Yuantian Miao, Shaleeza Sohail, Chaoxiang He, Shuo Wang 0012, Xiao Chen 0002
RAID7
2025 A comparative study between android phone and TV apps
Yonghui Liu 0001, Xiao Chen 0002, Yue Liu 0011, Pingfan Kong, Tegawendé F. Bissyandé, Jacques Klein, Xiaoyu Sun 0002, Li Li 0029, Chunyang Chen 0001, John C. Grundy
Autom. Softw. Eng.2
2025 LLM for Mobile: An Initial Roadmap
abstract
When mobile meets LLMs, mobile app users deserve to have more intelligent usage experiences. For this to happen, we argue that there is a strong need to apply LLMs for the mobile ecosystem. We therefore provide a research roadmap for guiding our fellow researchers to achieve that as a whole. In this roadmap, we sum up six directions that we believe are urgently required for research to enable native intelligence in mobile devices. In each direction, we further summarize the current research progress and the gaps that still need to be filled by our fellow researchers.
Daihang Chen, Yonghui Liu 0001, Mingyi Zhou, Yanjie Zhao 0001, Haoyu Wang 0001, Shuai Wang 0011, Xiao Chen 0002, Tegawendé F. Bissyandé, Jacques Klein, Li Li 0029
ACM Trans. Softw. Eng. Methodol.7
2025 Demystifying React Native Android Apps for Static Analysis
abstract
React Native, an open source framework, simplifies cross-platform app development by allowing JavaScript-side code to interact with native-side code. Previous studies disregarded React Native, resulting in insufficient static analysis of React Native app code. This study initiates the investigation of challenges when statically analyzing React Native apps. We propose ReuNify to improve Soot-based static analysis coverage for JavaScript-side and native-side code. ReuNify converts Hermes bytecode to Soot’s intermediate representation. Hermes bytecode, compiled from JavaScript code and integrated into React Native apps, possesses a unique syntax that eludes current JavaScript analyzers. Additionally, we investigate opcode distribution and conduct in-depth analyses of the usage of opcode between popular apps and malware. We also propose a benchmark consisting of 97 control flow-related cases to validate the control flow recovery of the generated intermediate representation. Furthermore, we model the cross-language communication mechanisms of React Native to expand the static analysis coverage for native-side code. Our evaluation demonstrates that ReuNify enables an average increase of 84% in reached nodes within the callgraph and further identifies an average of two additional privacy leaks in taint analysis. In summary, this article demonstrates that ReuNify significantly improves the static analysis for the React Native Android apps.
Yonghui Liu 0001, Xiao Chen 0002, Jordan Samhi, John C. Grundy, Chunyang Chen 0001, Li Li 0029
ACM Trans. Softw. Eng. Methodol.2
2024 Model-less Is the Best Model: Generating Pure Code Implementations to Replace On-Device DL Models
abstract
Recent studies show that on-device deployed deep learning (DL) models, such as those of Tensor Flow Lite (TFLite), can be easily extracted from real-world applications and devices by attackers to generate many kinds of adversarial and other attacks. Although securing deployed on-device DL models has gained increasing attention, no existing methods can fully prevent these attacks. Traditional software protection techniques have been widely explored. If on-device models can be implemented using pure code, such as C++, it will open the possibility of reusing existing robust software protection techniques. However, due to the complexity of DL models, there is no automatic method that can translate DL models to pure code. To fill this gap, we propose a novel method, CustomDLCoder, to automatically extract on-device DL model information and synthesize a customized executable program for a wide range of DL models. CustomDLCoder first parses the DL model, extracts its backend computing codes, configures the extracted codes, and then generates a customized program to implement and deploy the DL model without explicit model representation. The synthesized program hides model information for DL deployment environments since it does not need to retain explicit model representation, preventing many attacks on the DL model. In addition, it improves ML performance because the customized code removes model parsing and preprocessing steps and only retains the data computing process. Our experimental results show that CustomDLCoder improves model security by disabling on-device model sniffing. Compared with the original on-device platform (i.e., TFLite), our method can accelerate model inference by 21.0% and 24.3% on x86-64 and ARM64 platforms, respectively. Most importantly, it can significantly reduce memory consumption by 68.8% and 36.0% on x86-64 and ARM64 platforms, respectively.
Mingyi Zhou, Xiang Gao 0012, John C. Grundy, Chunyang Chen 0001, Xiao Chen 0002, Li Li 0029
ISSTA6
2024 Incremental Context-free Grammar Inference in Black Box Settings
abstract
Black-box context-free grammar inference presents a significant challenge in many practical settings due to limited access to example programs. The state-of-the-art methods, Arvada and Treevada, employ heuristic approaches to generalize grammar rules, initiating from flat parse trees and exploring diverse generalization sequences. We have observed that these approaches suffer from low quality and readability, primarily because they process entire example strings, adding to the complexity and substantially slowing down computations. To overcome these limitations, we propose a novel method that segments example strings into smaller units and incrementally infers the grammar. Our approach, named Kedavra, has demonstrated superior grammar quality (enhanced precision and recall), faster runtime, and improved readability through empirical comparison.
Xiao Chen 0002, Xi Xiao 0001, Xiaoyu Sun 0002, Shaohua Wang 0002, Jitao Han
ASE2
2024 DynaMO: Protecting Mobile DL Models through Coupling Obfuscated DL Operators
abstract
Deploying deep learning (DL) models on mobile applications (Apps) has become ever-more popular. However, existing studies show attackers can easily reverse-engineer mobile DL models in Apps to steal intellectual property or generate effective attacks. A recent approach, Model Obfuscation, has been proposed to defend against such reverse engineering by obfuscating DL model representations, such as weights and computational graphs, without affecting model performance. These existing model obfuscation methods use static methods to obfuscate the model representation, or they use half-dynamic methods but require users to restore the model information through additional input arguments. However, these static methods or half-dynamic methods cannot provide enough protection for on-device DL models. Attackers can use dynamic analysis to mine the sensitive information in the inference codes as the correct model information and intermediate results must be recovered at runtime for static and half-dynamic obfuscation methods. We assess the vulnerability of the existing obfuscation strategies using an instrumentation method and tool, DLModelExplorer, that dynamically extracts correct sensitive model information (i.e., weights, computational graph) at runtime. Experiments show it achieves very high attack performance (e.g., 98.76% of weights extraction rate and 99.89% of obfuscating operator classification rate). To defend against such attacks based on dynamic instrumentation, we propose DynaMO, a Dynamic Model Obfuscation strategy similar to Homomorphic Encryption. The obfuscation and recovery process can be done through simple linear transformation for the weights of randomly coupled eligible operators, which is a fully dynamic obfuscation strategy. Experiments show that our proposed strategy can dramatically improve model security compared with the existing obfuscation strategies, with only negligible overheads for on-device models. Our prototype tool is publicly available at https://github.com/zhoumingyi/DynaMO.
Mingyi Zhou, Xiang Gao 0012, Xiao Chen 0002, Chunyang Chen 0001, John C. Grundy, Li Li 0029
ASE3
2024 How COVID-19 impacts telehealth: an empirical study of telehealth services, users and the use of metaverse
abstract
Since the outbreak of the coronavirus 2019 (COVID-19) pandemic, telehealth services are regarded as a good approach to keep health workers and patients safe while simultaneously managing available resources.In this paper, we discuss the impact that COVID-19 has on telehealth services and on telehealth users' opinion of the service.We collected 245 Android telehealth apps, 144 iOS telehealth apps and 86 telehealth websites, and performed a systematic analysis on this dataset.In this analysis, we conducted a comparison analysis and relevant content analysis of the telehealth apps as well as their security risks.Apart from the mobile platforms, we also inspected the telehealth websites' features, particularly those related to the use of metaverse to improve current telehealth solutions.To further understand people's attitude towards telehealth services, we invited users to participate in a user study aimed at revealing what impact COVID-19 has on users' willingness to adopt telehealth services and revealing the gap between the telehealth service and its users.Our result shows that 27.1% new iOS apps and 27.4% new Android apps were released after the COVID-19 announcement, and a surge of updates were noted within 4 weeks after the COVID-19 announcement.We further found that COVID-19 is frequently mentioned in telehealth app reviews in the second and third quarter of 2020, and the most mentioned aspects related to COVID-19 include family, test result and vaccine.According to our user study, COVID-19 has a significant impact on the selection of telehealth services, especially for female participants, people aged 46-55, and students.The investigation also finds out that the use of metaverse will significantly improves the effectiveness of traditional telehealth solutions.
Lihong Tang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Wei Zhou 0044, Xiaogang Zhu 0001, Yang Xiang 0001
Connect. Sci.3
2024 BOOM: Bottleneck-Aware Opportunistic Multicast Strategy for Cooperative Maritime Sensing
abstract
With the advancements in sensing technologies, maritime sensing has become indispensable in various domains, including logistics, weather forecasting, and marine ranching. However, transmitting large volumes of sensing data faces many challenges in the maritime environment. First, the transmissions purely depend on satellite links often costly and suffer from long propagation latency. On the other hand, traditional unicast transmission results in data duplication, wasting valuable marine communication resources. With the increasing density of sensing devices, the communication distance between maritime sensors has become closer, enabling the deployment of maritime opportunistic networks consisting of device-to-device links. Rather than using unicast transmission over satellite links, employing multicast with opportunistic routing enables simultaneous data transmission to multiple destinations and saves communication resources. Even though the multicast method can avoid redundancy, conducting multicast without considering the maritime characteristics (i.e., the dynamics and the distribution of sensors) may lead to inefficient data delivery. Through real-world experiments, we observe that devices located on the edges of the network have a relatively low receiving rate compared with internal ones and tend to be the bottleneck of the overall multicast progress. Based on this observation, we propose BOOM, a bottleneck-aware opportunistic multicast strategy aiming at reducing multicast latency, taking into account the influence of the bottleneck node and broadcasting rate. Prominently, within maritime scenarios challenged by extreme conditions, such as storms, typhoons, and tsunamis, BOOM’s emphasis encompasses the adaptability of multicast strategies, which necessitates dynamic adjustments in response to equipment failures and shifts in network topology. Through mathematical analysis, we prove the formation of opportunistic multicast is an NP-hard problem and further design a heuristic algorithm based on the convex-hull method to reduce the computational cost in strategy generation. We compare BOOM with four other algorithms using real-world maritime vessel trajectories in various scenarios. The simulation result illustrates that the BOOM achieves a significant reduction in transmission latency, which reduces 36% when sensors are sparsely located in water areas, and the reduction could reach up to 59% when sensors are more dense. Furthermore, in extreme environmental testing conditions, BOOM continues to outperform other algorithms in terms of completion time, with performance improvements of up to 39% and 49% in sparse and dense topology environments, respectively.
Xiao Chen 0002, Chao Zhu 0002, Guanju Shi, Xiang Gao 0013, Yong Cui 0001
IEEE Internet Things J.1
2024 Demystifying the Evolution of Android Malware Variants
abstract
It is important to understand the evolution of Android malware as this facilitates the development of defence techniques by proactively capturing malware features. So far, researchers mainly rely on dendrogram or family-tree analysis for malware's evolutionary development. However, our research finds that these techniques cannot support comprehensive malware evolution modelling, which provides a detailed explanation for why Android malware samples evolve in specific ways. This shortcoming is mainly caused by the coarse-grained clustering and analysis of malware samples. For example, because these works do not divide malware samples of a family into variant sets and explore the evolution principles among those sets, they usually fail to capture new variants that have been empowered by the feature ‘drifting’ in evolution. To address this problem, we propose a fine-grained and in-depth analysis of Android malware. Our experimental work systematically reveals the phylogenetic relationships among the variant sets for a deeper malware evolution analysis. We introduce five metrics: silhouette coefficient, creation date, variant labels, the presentativeness of the variant set formula, and the correctness of the linked edges to evaluate the correctness of our analysis. The results show that our variant clustering achieved a high silhouette value at a small sample distance (0.3), a small standard deviation (three months and 16 days) date based on when the malware samples are lastly modified, a high label consistency (91.4%), a high representativeness (93.1%) of the variant set formula. All the linked variant sets are connected based on our PhyloNet construction rules. We further analyse the coding details of Android malware for each variant set and summarise models of their evolutionary development. In this work, we successfully expose two major models of malware evolution:active evolutionandpassive evolution. We also disclose four technical explanations on the incentives of the two evolution models (two for each model respectively). These findings are valuable for proactive defence against newly emerged malware samples.
Lihong Tang, Xiao Chen 0002, Sheng Wen, Li Li 0029, Marthie Grobler, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.2
2024 Enhancing GUI Exploration Coverage of Android Apps with Deep Link-Integrated Monkey
abstract
Mobile apps are ubiquitous in our daily lives for supporting different tasks such as reading and chatting. Despite the availability of many GUI testing tools, app testers still struggle with low testing code coverage due to tools frequently getting stuck in loops or overlooking activities with concealed entries. This results in a significant amount of testing time being spent on redundant and repetitive exploration of a few GUI pages. To address this, we utilize Android’s deep links, which assist in triggering Android intents to lead users to specific pages and introduce a deep link-enhanced exploration method. This approach, integrated into the testing tool Monkey, gives rise to Delm (Deep Link-enhanced Monkey). Delm oversees the dynamic exploration process, guiding the tool out of meaningless testing loops to unexplored GUI pages. We provide a rigorous activity context mock-up approach for triggering existing Android intents to discover more activities with hidden entrances. We conduct experiments to evaluate Delm’s effectiveness on activity context mock-up, activity coverage, method coverage, and crash detection. The findings reveal that Delm can mock up more complex activity contexts and significantly outperform state-of-the-art baselines with 27.2% activity coverage, 21.13% method coverage, and 23.81% crash detection.
Han Hu 0011, Han Wang 0023, Ruiqi Dong, Xiao Chen 0002, Chunyang Chen 0001
ACM Trans. Softw. Eng. Methodol.4
2023 ModelObfuscator: Obfuscating Model Information to Protect Deployed ML-Based Systems
abstract
More and more edge devices and mobile apps are leveraging deep learning (DL) capabilities. Deploying such models on devices – referred to as on-device models – rather than as remote cloud-hosted services, has gained popularity because it avoids transmitting user’s data off of the device and achieves high response time. However, on-device models can be easily attacked, as they can be accessed by unpacking corresponding apps and the model is fully exposed to attackers. Recent studies show that attackers can easily generate white-box-like attacks for an on-device model or even inverse its training data. To protect on-device models from white-box attacks, we propose a novel technique called model obfuscation. Specifically, model obfuscation hides and obfuscates the key information – structure, parameters and attributes – of models by renaming, parameter encapsulation, neural structure obfuscation, shortcut injection, and extra layer injection. We have developed a prototype tool ModelObfuscator to automatically obfuscate on-device TFLite models. Our experiments show that this proposed approach can dramatically improve model security by significantly increasing the difficulty of parsing models’ inner information, without increasing the latency of DL models. Our proposed on-device model obfuscation has the potential to be a fundamental technique for on-device model deployment. Our prototype tool is publicly available at https://github.com/zhoumingyi/ModelObfuscator.
Mingyi Zhou, Xiang Gao 0012, Jing Wu 0021, John C. Grundy, Xiao Chen 0002, Chunyang Chen 0001, Li Li 0029
ISSTA5
2023 ReuNify: A Step Towards Whole Program Analysis for React Native Android Apps
abstract
React Native is a widely-used open-source frame-work that facilitates the development of cross-platform mobile apps. The framework enables JavaScript code to interact with native-side code, such as Objective-C/Swift for iOS and Java/Kotlin for Android, via a communication mechanism provided by React Native. However, previous research and tools have overlooked this mechanism, resulting in incomplete analysis of React Native app code. To address this limitation, we have developed REUNIFY, a prototype tool that integrates the JavaScript and native-side code of React Native apps into an intermediate language that can be processed by the Soot static analysis framework. By doing so, REUNIFY enables the generation of a comprehensive model of the app's behavior. Our evaluation indicates that, by leveraging REUNIFY, the Soot-based framework can improve its coverage of static analysis for the 1,007 most popular React Native Android apps, augmenting the number of lines of Jimple code by 70%. Additionally, we observed an average increase of 84% in new nodes reached in the callgraph for these apps, after integrating REUNIFY. When REUNIFY is used for taint flow analysis, an average of two additional privacy leaks were identified. Overall, our results demonstrate that REUNIFY significantly enhances the Soot-based framework's capability to analyze React Native Android apps.
Yonghui Liu 0001, Xiao Chen 0002, John C. Grundy, Chunyang Chen 0001, Li Li 0029
ASE2
2023 FloodSFCP: Quality and Latency Balanced Service Function Chain Placement for Remote Sensing in LEO Satellite Network
abstract
Prompted by the significant advancements in image processing technologies and their diverse range of applications, remote sensing satellites are poised for rapid expansion. Nonetheless, offloading the vast amount of remote sensing satellite images to the ground gateway station is inefficient due to the exorbitant costs induced by satellite links, while the limited resources of individual satellites hinder local task processing. With the advancement of the network function virtualization (NFV) technology, a new paradigm for service function chain (SFC) has emerged, which can significantly improve the flexibility and resource utilization of network services and alleviate resource conflicts by dividing large services into smaller ones organized in the form of SFCs. As mega-constellations (e.g., Starlink) developed, the number of low earth orbit (LEO) satellites is increasing. By dividing services into small sub-services and organizing them into SFCs throughout the LEO network, services that cannot be completed by a single satellite can be accomplished through multi-satellite cooperation. However, the quality of the remote sensing service is positively correlated with its latency, and the rapidly changing topology of LEO networks also adds complexity to the SFC placement. Hence, how to select appropriate satellites to place the SFC and modulate service levels, in order to obtain better remote sensing results within an acceptable latency, remains a question. To address these issues, this paper proposes the FloodSFCP, an SFC placement method that aims to increase service quality and decrease latency through offline training and online optimization via deep reinforcement learning, taking into account the variation in LEO network topology. By introducing NoisyNet, Dueling, and N-step learning, we improve the model’s generalization ability and reduce the state space, thus enhancing convergence speed while reducing decision and training time. Experimental results demonstrate that FloodSFCP significantly improves service quality while reducing total decision costs.
Ruoyi Zhang, Chao Zhu 0002, Xiao Chen 0002, Qingyuan Gong, Xinlei Xie, Xiangyuan Bu
SECON3
2023 LazyCow: A Lightweight Crowdsourced Testing Tool for Taming Android Fragmentation
abstract
Android fragmentation refers to the increasing variety of Android devices and operating system versions. Their number make it impossible to test an app on every supported device, resulting in many device compatibility issues and leading to poor user experiences. To mitigate this, a number of works that automatically detect compatibility issues have been proposed. However, current state-of-the-art techniques can only be used to detect specific types of compatibility issues (i.e., compatibility issues caused by API signature evolution), i.e., many other essential categories of compatibility issues are still unknown. For instance, customised OS versions on real devices and semantic OS modifications could result in severe compatibility issues that are difficult to detect statically. In order to address this research gap and facilitate the prospect of taming Android frag- mentation through crowdsourced efforts, we propose LazyCow, a novel, lightweight, crowdsourced testing tool. Our experimental results involving thousands of test cases on real Android devices demonstrate that LazyCow is effective at autonomously identifying and validating API-induced compatibility issues. The source code of both client side and server side are all made publicly available in our artifact package. A demo video of our tool is available at https://www.youtube.com/watch?v=_xzWv_mo5xQ.
Xiaoyu Sun 0002, Xiao Chen 0002, Yonghui Liu 0001, John C. Grundy, Li Li 0029
ESEC/SIGSOFT FSE2
2023 Dynalogue: A Transformer-Based Dialogue System with Dynamic Attention
abstract
Businesses face a range of cyber risks, both external threats and internal vulnerabilities that continue to evolve over time. As cyber attacks continue to increase in complexity and sophistication, more organisations will experience them. For this reason, it is important that organisations seek timely consultancy from cyber professionals so that they can respond to and recover from cyber attacks as quickly as possible. However, huge surges in cyber attacks have long left cyber professionals short of what is required to cover the security needs. This problem is getting worse when an increasing number of people choose to work from home during the pandemic because this situation usually yields extra communication cost.
Rongjunchen Zhang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Surya Nepal, Cécile Paris, Yang Xiang 0001
WWW3
2023 How Does Visualisation Help App Practitioners Analyse Android Apps?
abstract
Behaviour analysis is essential for the security verification of suspicious Android applications, but analysts are usually faced with a huge obstacle when conducting the app behaviour analysis. They are expected to have comprehensive knowledge of different IT fields and a strong awareness of cyber threats. However, training a new security analyst typically requires a significant amount of time and can be extremely costly. Although there are tools available to assist analysts in studying Android behaviour and security, the completion of this task still heavily relies on the experience of the analysts. To address this problem, we recognise visualisation as a promising method and conduct a series of controlled experiments to demonstrate its effectiveness in the context of Android app behaviour and security analysis. We accordingly develop a visualisation tool based on apps’ call graphs (CG) (namedVisualDroid) and conduct an experiment and a follow-up interview. Compared to existing solutions, the results suggest that the CG-based visualisation solution (VisualDroid) can lower the barriers to Android behaviour and security analysis. The user study reveals that the platform includes CG-based visualisation components leads to a statistically significant improvement in Android behaviour analysis and security awareness. More specifically, it improvesAPK Analyzer,JD-GUI,JD-GUI+FlowDroidby 71.4%, 35.7%, and 39.2% in terms of the effectiveness of behaviour analysis. Participants who useVisualDroidalso show improvements in the aspect of security awareness with an increase of 155% againstAPK Analyzer, 96% againstJD-GUI, and 59.3%JD-GUI+FlowDroid.
Lihong Tang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Li Li 0029, Xin Xia 0001, Marthie Grobler, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.3
2023 Demystifying Hidden Sensitive Operations in Android Apps
abstract
Security of Android devices is now paramount, given their wide adoption among consumers. As researchers develop tools for statically or dynamically detecting suspicious apps, malware writers regularly update their attack mechanisms to hide malicious behavior implementation. This poses two problems to current research techniques: static analysis approaches, given their over-approximations, can report an overwhelming number of false alarms, while dynamic approaches will miss those behaviors that are hidden through evasion techniques. We propose in this work a static approach specifically targeted at highlighting hidden sensitive operations (HSOs), mainly sensitive data flows. The prototype version of HiSenDroid has been evaluated on a large-scale dataset of thousands of malware and goodware samples on which it successfully revealed anti-analysis code snippets aiming at evading detection by dynamic analysis. We further experimentally show that, with FlowDroid, some of the hidden sensitive behaviors would eventually lead to private data leaks. Those leaks would have been hard to spot either manually among the large number of false positives reported by the state-of-the-art static analyzers, or by dynamic tools. Overall, by putting the light on hidden sensitive operations, HiSenDroid helps security analysts in validating potentially sensitive data operations, which would be previously unnoticed.
Xiaoyu Sun 0002, Xiao Chen 0002, Li Li 0029, Haipeng Cai, John C. Grundy, Jordan Samhi, Tegawendé F. Bissyandé, Jacques Klein
ACM Trans. Softw. Eng. Methodol.2
2023 Taming Android Fragmentation Through Lightweight Crowdsourced Testing
abstract
Android fragmentation refers to the overwhelming diversity of Android devices and OS versions. These lead to the impossibility of testing an app on every supported device, leaving a number of compatibility bugs scattered in the community and thereby resulting in poor user experiences. To mitigate this, our fellow researchers have designed various works to automatically detect such compatibility issues. However, the current state-of-the-art tools can only be used to detect specific kinds of compatibility issues (i.e., compatibility issues caused by API signature evolution), i.e., many other essential types of compatibility issues are still unrevealed. For example, customized OS versions on real devices and semantic changes of OS could lead to serious compatibility issues, which are non-trivial to be detected statically. To this end, we propose a novel, lightweight, crowdsourced testing approach, to fill this research gap and enable the possibility of taming Android fragmentation through crowdsourced efforts. Specifically, crowdsourced testing is an emerging alternative to conventional mobile testing mechanisms that allow developers to test their products on real devices to pinpoint platform-specific issues. Experimental results on thousands of test cases on real-world Android devices show that is effective in automatically identifying and verifying API-induced compatibility issues. Also, after investigating the user experience through qualitative metrics, users' satisfaction provides strong evidence that is useful and welcome in practice.
Xiaoyu Sun 0002, Xiao Chen 0002, Yonghui Liu 0001, John C. Grundy, Li Li 0029
IEEE Trans. Software Eng.2
2022 Mining Android API Usage to Generate Unit Test Cases for Pinpointing Compatibility Issues
abstract
Despite being one of the largest and most popular projects, the official Android framework has only provided test cases for less than 30% of its APIs. Such a poor test case coverage rate has led to many compatibility issues that can cause apps to crash at runtime on specific Android devices, resulting in poor user experiences for both apps and the Android ecosystem. To mitigate this impact, various approaches have been proposed to automatically detect such compatibility issues. Unfortunately, these approaches have only focused on detecting signature-induced compatibility issues (i.e., a certain API does not exist in certain Android versions), leaving other equally important types of compatibility issues unresolved. In this work, we propose a novel prototype tool, JUnitTestGen, to fill this gap by mining existing Android API usage to generate unit test cases. After locating Android API usage in given real-world Android apps, JUnitTestGen performs inter-procedural backward data-flow analysis to generate a minimal executable code snippet (i.e., test case). Experimental results on thousands of real-world Android apps show that JUnitTestGen is effective in generating valid unit test cases for Android APIs. We show that these generated test cases are indeed helpful for pinpointing compatibility issues, including ones involving semantic code changes.
Xiaoyu Sun 0002, Xiao Chen 0002, Yanjie Zhao 0001, John C. Grundy, Li Li 0029
ASE2
2022 Cross-language Android permission specification
abstract
The Android system manages access to sensitive APIs by permission enforcement. An application (app) must declare proper permissions before invoking specific Android APIs. However, there is no official documentation providing the complete list of permission-protected APIs and the corresponding permissions to date. Researchers have spent significant efforts extracting such API protection mapping from the Android API framework, which leverages static code analysis to determine if specific permissions are required before accessing an API. Nevertheless, none of them has attempted to analyze the protection mapping in the native library (i.e., code written in C and C++), an essential component of the Android framework that handles communication with the lower-level hardware, such as cameras and sensors. While the protection mapping can be utilized to detect various security vulnerabilities in Android apps, such as permission over-privilege, imprecise mapping will lead to false results in detecting such security vulnerabilities. To fill this gap, we thereby propose to construct the protection mapping involved in the native libraries of the Android framework to present a complete and accurate specification of Android API protection. We develop a prototype system, named NatiDroid, to facilitate the cross-language static analysis and compare its performance with two state-of-the-practice tools, termed Axplorer and Arcade. We evaluate NatiDroid on more than 11,000 Android apps, including system apps from custom Android ROMs and third-party apps from the Google Play. Our NatiDroid can identify up to 464 new API-permission mappings, in contrast to the worst-case results derived from both Axplorer and Arcade, where approximately 71% apps have at least one false positive in permission over-privilege. We have disclosed all the potential vulnerabilities detected to the stakeholders.
Xiao Chen 0002, Ruoxi Sun 0001, Minhui Xue 0001, Sheng Wen, M. Ejaz Ahmed, Seyit Ahmet Çamtepe, Yang Xiang 0001
ESEC/SIGSOFT FSE2
2022 Backdoor Attack on Machine Learning Based Android Malware Detectors
abstract
Machine learning (ML) has been widely used for malware detection on different operating systems, including Android. To keep up with malware's evolution, the detection models usually need to be retrained periodically (e.g., every month) based on the data collected in the wild. However, this leads to poisoning attacks, specifically backdoor attacks, which subvert the learning process and create evasion ‘tunnels’ for manipulated malware samples. To date, we have not found any prior research that explored this critical problem in Android malware detectors. Although there are already some similar works in the image classification field, most of those similar ideas cannot be borrowed to solve this problem, because the assumption that the attacker has full control of the training data collection or labelling process is not realistic in real-world malware detection scenarios. In this article, we are motivated to study the backdoor attack against Android malware detectors. The backdoor is created and injected into the model stealthily without access to the training data and activated when an app with the trigger is presented. We demonstrate the proposed attack on four typical malware detectors that have been widely discussed in academia. Our evaluation shows that the proposed backdoor attack achieves up to 99 percent evasion rate over 750 malware samples. Moreover, the above successful attack is realised by a small size of triggers (only four features) and a very low data poisoning rate (0.3 percent).
Xiao Chen 0002, Derui Wang, Sheng Wen, M. Ejaz Ahmed, Seyit Ahmet Çamtepe, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.2
2021 Time-aware User Modeling with Check-in Time Prediction for Next POI Recommendation
abstract
POI (point-of-interest) recommendation as an important type of location-based services has received increasing attention with the rise of location-based social networks. Although significant efforts have been dedicated to learning and recommending users' next POIs based on their historical mobility traces, there still lacks consideration of the discrepancy of users' check-in time preferences and the inherent relationships between POIs and check-in times. To fill this gap, this paper proposes a novel recommendation method which applies multi-task learning over historical user mobility traces known to be sparse. Specifically, we design a cross-graph neural network to obtain time-aware user modeling and control how much information flows across different semantic spaces, which makes up the inadequate representation of existing user modeling methods. In addition, we design a check-in time prediction task to learn users' activities from a time perspective and learn internal patterns between POIs and their check-in times, aiming to reduce the search space to overcome the data sparsity problem. Comprehensive experiments on two real-world public datasets demonstrate that our proposed method outperforms several representative POI recommendation methods with 8.93% to 20.21 % improvement on Recall@1, 5, 10, and 9.25% to 17.56% improvement on Mean Reciprocal Rank.
Xin Wang 0114, Xiao Liu 0004, Li Li 0029, Xiao Chen 0002, Jin Liu 0016, Hao Wu 0010
ICWS4
2021 Characterizing Sensor Leaks in Android Apps
abstract
While extremely valuable to achieve advanced functions, mobile phone sensors can be abused by attackers to implement malicious activities in Android apps, as experimentally demonstrated by many state-of-the-art studies. There is hence a strong need to regulate the usage of mobile sensors so as to keep them from being exploited by malicious attackers. However, despite the fact that various efforts have been put in achieving this, i.e., detecting privacy leaks in Android apps, we have not yet found approaches to automatically detect sensor leaks in Android apps. To fill the gap, we designed and implemented a novel prototype tool, Seeker, that extends the famous FlowDroid tool to detect sensor-based data leaks in Android apps. Seeker conducts sensor-focused static taint analyses directly on the Android apps' bytecode and reports not only sensor-triggered privacy leaks but also the sensor types involved in the leaks. Experimental results using over 40,000 real-world Android apps show that Seeker is effective in detecting sensor leaks in Android apps, and malicious apps are more interested in leaking sensor data than benign apps.
Xiaoyu Sun 0002, Xiao Chen 0002, Kui Liu 0001, Sheng Wen, Li Li 0029, John C. Grundy
ISSRE2
2020 Detecting and Explaining Self-Admitted Technical Debts with Attention-based Neural Networks
abstract
Self-Admitted Technical Debt (SATD) is a sub-type of technical debt. It is introduced to represent such technical debts that are intentionally introduced by developers in the process of software development. While being able to gain short-term benefits, the introduction of SATDs often requires to be paid back later with a higher cost, e.g., introducing bugs to the software or increasing the complexity of the software.
Xin Wang 0114, Jin Liu 0016, Li Li 0029, Xiao Chen 0002, Xiao Liu 0004, Hao Wu 0010
ASE4
2020 SpeedNeuzz: Speed Up Neural Program Approximation with Neighbor Edge Knowledge
abstract
Fuzzing has been a great success in discovering real-world complex programs vulnerabilities. However, fuzzing achieves this effect by blindly generating a large number of test cases, which undoubtedly contains a lot of meaningless mutation inputs. To solve the blindness, machine learning technology is applied to fuzzing in recent work. Some of the machine learning based methods focus on locating and mutating the key bytes in the input, but they do not pay attention to the characteristics in the field of fuzzing when they combine machine learning technology with fuzzing. In this paper, we implement a new fuzzer, called Speed-Neuzz, which uses neural networks to model the branch behaviours of the program based on accurate training data after mitigating the hash collision of AFL. Furthermore, SpeedNeuzz locates and mutates critical bytes in the program input with a gradient-based strategy as well as neighbor edge information. Taking the neighbor edge knowledge into account, we can further reduce the blindness of the mutation based on gradient information so that SpeedNeuzz can generate a large number of quality inputs. Experiments on several real-world programs prove that SpeedNeuzz can achieve higher edge coverage than the state-of-the-art fuzzer NEUZZ under the same time budget.
Xi Xiao 0001, Xiaogang Zhu 0001, Xiao Chen 0002, Sheng Wen, Bin Zhang 0048
TrustCom4
2020 Android HIV: A Study of Repackaging Malware for Evading Machine-Learning Detection
abstract
Machine learning-based solutions have been successfully employed for the automatic detection of malware on Android. However, machine learning models lack robustness to adversarial examples, which are crafted by adding carefully chosen perturbations to the normal inputs. So far, the adversarial examples can only deceive detectors that rely on syntactic features (e.g., requested permissions, API calls,etc.), and the perturbations can only be implemented by simply modifying application’s manifest. While recent Android malware detectors rely more on semantic features from Dalvik bytecode rather than manifest, existing attacking/defending methods are no longer effective. In this paper, we introduce a new attacking method that generates adversarial examples of Android malware and evades being detected by the current models. To this end, we propose a method of applying optimal perturbations onto Android APK that can successfully deceive the machine learning detectors. We develop an automated tool to generate the adversarial examples without human intervention. In contrast to existing works, the adversarial examples crafted by our method can also deceive recent machine learning-based detectors that rely on semantic features such as control-flow-graph. The perturbations can also be implemented directly onto APK’s Dalvik bytecode rather than Android manifest to evade from recent detectors. We demonstrate our attack on two state-of-the-art Android malware detection schemes, MaMaDroid and Drebin. Our results show that the malware detection rates decreased from 96% to 0% in MaMaDroid, and from 97% to 0% in Drebin, with just a small number of codes to be inserted into the APK.
Xiao Chen 0002, Derui Wang, Sheng Wen, Jun Zhang 0010, Surya Nepal, Yang Xiang 0001, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.1
2015 6 million spam tweets: A large ground truth for timely Twitter spam detection
abstract
Twitter has changed the way of communication and getting news for people's daily life in recent years. Meanwhile, due to the popularity of Twitter, it also becomes a main target for spamming activities. In order to stop spammers, Twitter is using Google SafeBrowsing to detect and block spam links. Despite that blacklists can block malicious URLs embedded in tweets, their lagging time hinders the ability to protect users in real-time. Thus, researchers begin to apply different machine learning algorithms to detect Twitter spam. However, there is no comprehensive evaluation on each algorithms' performance for real-time Twitter spam detection due to the lack of large groundtruth. To carry out a thorough evaluation, we collected a large dataset of over 600 million public tweets. We further labelled around 6.5 million spam tweets and extracted 12 light-weight features, which can be used for online detection. In addition, we have conducted a number of experiments on six machine learning algorithms under various conditions to better understand their effectiveness and weakness for timely Twitter spam detection. We will make our labelled dataset for researchers who are interested in validating or extending our work.
Chao Chen 0015, Jun Zhang 0010, Xiao Chen 0002, Yang Xiang 0001, Wanlei Zhou 0001
ICC3
2015 Robust Network Traffic Classification
abstract
As a fundamental tool for network management and security, traffic classification has attracted increasing attention in recent years. A significant challenge to the robustness of classification performance comes from zero-day applications previously unknown in traffic classification systems. In this paper, we propose a new scheme of Robust statistical Traffic Classification (RTC) by combining supervised and unsupervised machine learning techniques to meet this challenge. The proposed RTC scheme has the capability of identifying the traffic of zero-day applications as well as accurately discriminating predefined application classes. In addition, we develop a new method for automating the RTC scheme parameters optimization process. The empirical study on real-world traffic data confirms the effectiveness of the proposed scheme. When zero-day applications are present, the classification performance of the new scheme is significantly better than four state-of-the-art methods: random forest, correlation-based classification, semi-supervised clustering, and one-class SVM.
Jun Zhang 0010, Xiao Chen 0002, Yang Xiang 0001, Wanlei Zhou 0001, Jie Wu 0001
IEEE/ACM Trans. Netw.2
2014 On Addressing the Imbalance Problem: A Correlated KNN Approach for Network Traffic Classification
Di Wu 0050, Xiao Chen 0002, Chao Chen 0015, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001
NSS2