Liu Wang 0002

dblp:10/3099-2 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-3982-4993ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 14 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Doxing-as-a-Service: Demystifying the Chinese Online Doxing Ecosystem
abstract
Doxing refers to the disclosure of personal information without consent, has evolved from sporadic acts of online vigilantism into a structured and commodified practice. In Chinese cyberspace, this shift has produced Doxing-as-a-Service (DaaS), a commercial model in which personal data is retrieved, organized, and traded as on-demand products. This industrialization of privacy violation lowers the barriers to doxing and amplifies its social harms, posing new challenges for building a responsible and safe web. Yet little is known about how DaaS operates or sustains itself, motivating our systematic, data-driven examination of its ecosystem and practices. This paper provides the first systematic study of the Chinese DaaS ecosystem. Analyzing 25,972 messages and 13.22 million subscriber links from 100 major channels on Telegram, we demystify its organization, operations, and user engagement. We find that the DaaS ecosystem operates through a three-tier supply chain linking data providers, service operators, and end users. Operators sustain illicit businesses through six major service categories, persistent advertising, and crypto-based payments, while users interact via specialized group spaces that enable real-time matching, large-scale identity exposure, and community-driven fraud mitigation. Our findings reveal a mature, resilient underground data market operating within mainstream messaging platforms, highlighting new challenges for online privacy and exposing critical vulnerabilities in platform governance and content moderation. This study provides empirical evidence for developing effective regulatory frameworks and accountability mechanisms to mitigate commodified online harms on encrypted messaging services.
Yiran Gao, Pengcheng Xia 0001, Liu Wang 0002, Tianming Liu 0002, Haoyu Wang 0001
WWW3
2025 TacDroid: Detection of Illicit Apps Through Hybrid Analysis of UI-Based Transition Graphs
abstract
Illicit apps have emerged as a thriving underground industry, driven by their substantial profitability. These apps either offer users restricted services (e.g., porn and gambling) or engage in fraudulent activities like scams. Despite the widespread presence of illicit apps, scant attention has been directed towards this issue, with several existing detection methods predominantly relying on static analysis alone. However, given the burgeoning trend wherein an increasing number of mobile apps achieve their core functionality through dynamic resource loading, depending solely on static analysis proves inadequate. To address this challenge, in this paper, we introduce Tac-droid,a novel approach that integrates dynamic analysis for dynamic content retrieval with static analysis to mitigate the limitations inherent in both methods, i.e., the low coverage of dynamic analysis and the low accuracy of static analysis. Specifically, Tacdroid conducts both dynamic and static analyses on an Android app to construct dynamic and static User Interface Transition Graphs (UTGs), respectively. These two UTGs are then correlated to create an intermediate UTG. Subsequently, Tacdroid embeds graph structure and utilizes an enhanced Graph Autoencoder (GAE) model to predict transitions between nodes. Through link prediction, Tacdroid effectively eliminates false positive transition edges stemming from misjudgments in static analysis and supplements false negative transition edges overlooked in the intermediate UTG, thereby generating a comprehensive and accurate UTG. Finally, Tacdroid determines the legitimacy of an app and identifies its category based on the app's UTG. Our evaluation results highlight the outstanding accuracy of Tacdroid in detecting illicit apps. It significantly surpasses the state-of-the-art work, achieving an F1-score of 96.73%. This work represents a notable advancement in the identification and categorization of illicit apps.
Yanchen Lu, Zehua He, Haitao Xu 0002, Zhao Li 0007, Shuai Hao 0001, Liu Wang 0002, Haoyu Wang 0001, Kui Ren 0001
ICSE7
2025 On the (In)Security of Non-resettable Device Identifiers in Custom Android Systems
abstract
User tracking is critical in the mobile ecosystem and relies on device identifiers to build user profiles. Early versions of Android allowed third-party apps to easily access non-resettable identifiers such as serial numbers and IMEI. As privacy concerns grew, Google has tightened identifier access in native Android. In response, stakeholders in custom Android systems introduced covert channels (e.g., system properties and settings) to maintain consistent and stable identifier access across systems and devices, which undoubtedly increases privacy risks. This paper examines the introduction of such channels through system customization and their vulnerability due to poor access control. We present IDRADAR, a scalable and accurate approach for identifying vulnerable properties and settings in custom Android systems. Applying our approach to 1,814 custom ROMs, we identified 8,192 system properties and 3,620 settings that store non-resettable device identifiers. Among these, 3,477 properties and 1,336 settings lack adequate access control and could be exploited by third-party apps to track users without permissions. Further validation on real devices demonstrates the effectiveness of our approach. Compared to state-of-the-art, IDRADAR offers improved scalability and analytical capabilities. Additionally, we investigate the root causes of the access control deficiencies and observe that such vulnerabilities frequently recur across devices from the same OEMs. We have reported our findings to the respective vendors and received positive confirmations. Our work underscores the need for greater scrutiny of covert access to device identifiers and better solutions to safeguard user privacy during system customizations.
Zikan Dong, Liu Wang 0002, Guoai Xu, Haoyu Wang 0001
ASE2
2025 A Big Step Forward? A User-Centric Examination of iOS App Privacy Report and Enhancements
abstract
The prevalent engagement with mobile apps underscores the importance of understanding their data practices. Transparency plays a crucial role in this context, ensuring users to be informed and give consent before any data access occurs. Apple introduced a new feature since iOS 15.2, App Privacy Report, to inform users about detailed insights into apps' data access and sharing. This feature continues Apple's trend of privacy-focused innovations (following Privacy Nutrition Labels), and has been marketed as a big step forward in user privacy. However, its real-world impacts on user privacy and control remain unexamined. We thus proposed an end-to-end study involving systematic assessment of the App Privacy Report's real-world benefits and limitations, LLM-enabled and multi-technique synthesized enhancements, and comprehensive evaluation from both system and user perspectives. Through a structured focus group study with twelve everyday iOS users, we explored their experiences, understanding, and perceptions of the feature, suggesting its limited practical impact resulting from missing important details. We identified two primary user concerns: the clarity of data access purpose and domain description. In response, we proposed enhancements including a purpose inference framework and domain clarification pipeline. We demonstrated the effectiveness and benefits of such enhancements for mobile app users. This work provides practical insights that could help enhance user privacy transparency and discusses areas for future research.
Liu Wang 0002, Dong Wang 0053, Shidong Pan, Haoyu Wang 0001
SP1
2024 MalCertain: Enhancing Deep Neural Network Based Android Malware Detection by Tackling Prediction Uncertainty
abstract
The long-lasting Android malware threat has attracted significant research efforts in malware detection. In particular, by modeling malware detection as a classification problem, machine learning based approaches, especially deep neural network (DNN) based approaches, are increasingly being used for Android malware detection and have achieved significant improvements over other detection approaches such as signature-based approaches. However, as Android malware evolve rapidly and the presence of adversarial samples, DNN models trained on early constructed samples often yield poor decisions when used to detect newly emerging samples. Fundamentally, this phenomenon can be summarized as the uncertainly in the data (noise or randomness) and the weakness in the training process (insufficient training data). Overlooking these uncertainties poses risks in the model predictions. In this paper, we take the first step to estimate the prediction uncertainty of DNN models in malware detection and leverage these estimates to enhance Android malware detection techniques. Specifically, besides training a DNN model to predict malware, we employ several uncertainty estimation methods to train a Correction Model that determines whether a sample is correctly or incorrectly predicted by the DNN model. We then leverage the estimated uncertainty output by the Correction Model to correct the prediction results, improving the accuracy of the DNN model. Experimental results show that our proposed MalCertain effectively improves the accuracy of the underlying DNN models for Android malware detection by around 21% and significantly improves the detection effectiveness of adversarial Android malware samples by up to 94.38%. Our research sheds light on the promising direction that leverages prediction uncertainty to improve prediction-based software engineering tasks.
Guosheng Xu 0001, Liu Wang 0002, Xusheng Xiao, Xiapu Luo, Guoai Xu, Haoyu Wang 0001
ICSE3
2024 Android Malware Family Labeling: Perspectives from the Industry
abstract
Labeling and classifying Android malware is important for identifying new threats, triaging security incidents, and demystifying evasion techniques. To automate the malware classification pipeline, state-of-the-art tools such as AVClass and Euphony unify raw labels from commercial antivirus vendors (i.e., VirusTotal) to produce family labels. These tools are widely used for automatic malware classification in both academic research and industry practice. However, they face significant limitations in real-world industrial scenarios with numerous and dynamically changing samples. For example, our industrial practices revealed that VirusTotal's results change over time, leading to temporal inconsistencies in family labeling results that rely on label unification, which can severely impact a company's security posture. Despite this, such issues and challenges remain understudied. In this paper, we present the first systematic measurement study of existing automatic Android malware family labeling systems from various aspects, including label dynamics, consistency, reliability, and etc. Based on a large-scale dataset, we validate that the labeling results of these systems do evolve with time, and such evolution can introduce bias into many previous studies on performance assessments. We also reveal substantial divergence in labeling decisions across different systems when given the same input. Besides, we identify a disclosure priority among families in these systems' labeling processes, which could threaten the industry by allowing malicious actors to exploit these discrepancies. Our findings could benefit both researchers and industry practitioners for further refinement of automatic malware family labeling systems, contributing to their practical applications.
Liu Wang 0002, Haoyu Wang 0001, Tao Zhang 0001, Haitao Xu 0002, Guozhu Meng, Peiming Gao, Yi Wang 0013
ASE1
2024 Global Prosperity or Local Monopoly? Understanding the Geography of App Popularity
abstract
App stores allow developers to globally distribute their apps to gain more users and attention. In the highly competitive market of app stores, developers need to cater to a large number of users spanning multiple countries. We posit that the characteristics of diverse geographical, linguistic, cultural, societal, and economic environments may impact the adoption of apps. In this paper, we take the first step to characterize popular apps across over 150 countries worldwide, and explore the potential correlations to a number of underlying factors including geography, language as well as cultural, societal, and economic dimensions. Our study is based on a longitudinal (one-year) dataset of daily app popularity from the iOS app stores, covering 154 regions around the world. We reveal that app popularity shows great diversity across the world, while similarities exist among countries that share geographical proximity and linguistic convergence. The differences in app popularity across regions can be further correlated with the cultural model and socioeconomic indices we adopt. On top of the dataset and findings, we implement a prediction task that contributes to app distribution, helping developers choose the right market to distribute and promote their apps. To the best of our knowledge, we are the first to attempt to provide a global understanding of the characteristics of app popularity across the mobile app ecosystem. Our observations can benefit stakeholders in the ecosystem, striving to improve app uptake.
Liu Wang 0002, Conghui Zheng, Haoyu Wang 0001, Xiapu Luo, Gareth Tyson, Yi Wang 0004, Shangguang Wang
MSR1
2023 Re-measuring the Label Dynamics of Online Anti-Malware Engines from Millions of Samples
abstract
VirusTotal is the most widely used online scanning service in both academia and industry. However, it is known that the results returned by antivirus engines are often inconsistent and changing over time. The intrinsic dynamics of VirusTotal labeling have prompted researchers to investigate the characteristics of label dynamics for more effective use. However, they are generally limited in terms of the size and diversity of the datasets used in the measurements. This poses threats to many of their conclusions. In this paper, we perform an extraordinary large-scale study to re-measure the label dynamics of VirusTotal. Our dataset involves all the scan data in VirusTotal over a 14-month period, including over 571 million samples and 847 million reports in total. With this large dataset, we are able to revisit many issues related to the label dynamics of VirusTotal, including the prevalence of label dynamics/silence, the characteristics across file types, the impact of label dynamics on common label aggregation methods, the stabilization patterns of labels, etc. Our measurement reveals some observations that are unknown to the research community and even inconsistent with previous research. We believe that our findings could help researchers advance the understanding of the VirusTotal ecosystem.
Liu Wang 0002, Feng Dong 0008, Haoyu Wang 0001
IMC2
2023 A Deep Dive into the Featured iOS Apps
abstract
Millions of apps in markets have made it difficult for mobile users to find fancy and high quality apps. Mobile app markets have deployed mechanisms to recommend apps to users. Apple usually features apps in the iOS App Store, and mobile users could see the featured apps as soon as they open the App Store. In general, getting apps featured is an achievement all developers strive towards and it is a common belief that getting app featured means that the app is becoming popular. However, the official app recommendation mechanism has not been characterized yet. To fill the void, we present a large-scale and longitudinal study of featured apps on iOS App Store. Specifically, we collaborate with our industry partner to monitor the iOS App Store and collect the daily featured apps in both the US and China, covering a span of over 1.5 years. Based on this comprehensive dataset, we characterize the featured apps from various dimensions and investigate the impact of app recommendation on app popularity. We have revealed a number of observations that are unknown to the community. Most importantly, we observe that although getting featured indeed has a positive effect for most apps, the duration of this effect is short-lived. In addition, there are times when the recommendations are ineffective, and we propose some potential reasons and tips for this. Our study can offer practical implications on app promotion to stakeholders in the mobile app ecosystem.
Liu Wang 0002, Haoyu Wang 0001, Li Li 0029, Yi Wang 0004
Internetware1
2023 Wemint:Tainting Sensitive Data Leaks in WeChat Mini-Programs
abstract
Mini-programs (MiniApps), lightweight versions of full-featured mobile apps that run inside a host app such as WeChat, have become increasingly popular due to their simplified and convenient user experiences. However, MiniApps raise new security and privacy concerns as they can access partially or all of host apps' system resources, including sensitive personal data. While taint detection has been proven effective in addressing this kind of concerns, existing taint detection techniques for mobile apps cannot be directly applied to MiniApps. The main reason is that the key logics of MiniApps are usually written in J avaScript, and its intrinsic characteristics (function-level scope, dynamic types, synchronous programming, and code obfuscation) prevent existing taint detection techniques from precisely propagating the taints. To address this problem, we propose a novel taint detection technique, Wemint, that detects sensitive information leaks in MiniApps. Specifically, Wemint facilitates taint propagation via building a context-based model based on the operational prin-ciple of MiniApps and J avaScript, and addresses asynchronous function calls by modeling their callbacks explicitly in taint rules. In addition, due to the adoption of Abstract Syntax Trees (ASTs) for code representation during taint detection, Wemint exhibits better robustness against the commonly-applied code obfuscation. Our experimental results show that Wemint can effectively detect sensitive information leaks in WeChat MiniApps, as well as trace the path of sensitive data flows. By applying Wemint to over 20K suspicious MiniApps, we found that over 7.5K (36.5 %) of them have sensitive data leaks, and Wemint outperforms the state-of-the-art DoubleX based techniques in detecting these leaks.
Shi Meng, Liu Wang 0002, Shenao Wang 0001, Kailong Wang 0001, Xusheng Xiao, Guangdong Bai, Haoyu Wang 0001
ASE2
2023 DISTDET: A Cost-Effective Distributed Cyber Threat Detection System
Feng Dong 0008, Liu Wang 0002, Xu Nie, Fei Shao, Haoyu Wang 0001, Ding Li 0001, Xiapu Luo, Xusheng Xiao
USENIX Security Symposium2
2022 MalWhiteout: Reducing Label Errors in Android Malware Detection
abstract
Machine learning based Android malware detection has attracted a great deal of research work in recent years. A reliable malware dataset is critical to evaluate the effectiveness of malware detection approaches. Unfortunately, existing malware datasets used in our community are mainly labelled by leveraging existing anti-virus services (i.e., VirusTotal), which are prone to mislabelling. This, however, would lead to the inaccurate evaluation of the malware detection techniques. Removing label noises from Android malware datasets can be quite challenging, especially at a large data scale. To address this problem, we propose an effective approach called MalWhiteout to reduce label errors in Android malware datasets. Specifically, we creatively introduce Confident Learning (CL), an advanced noise estimation approach, to the domain of Android malware detection. To combat false positives introduced by CL, we incorporate the idea of ensemble learning and inter-app relation to achieve a more robust capability in noise detection. We evaluate MalWhiteout on a curated large-scale and reliable benchmark dataset. Experimental results show that MalWhiteout is capable of detecting label noises with over 94% accuracy even at a high noise ratio (i.e., 30%) of the dataset. MalWhiteout outperforms the state-of-the-art approach in terms of both effectiveness (8% to 218% improvement) and efficiency (70 to 249 times faster) across different settings. By reducing label noises, we show that the performance of existing malware detection approaches can be improved.
Liu Wang 0002, Haoyu Wang 0001, Xiapu Luo, Yulei Sui
ASE1
2022 Privacy Analysis of Period Tracking Mobile Apps in the Post-Roe v. Wade Era
abstract
To help people manage their health, period tracking apps have become very popular in recent years. However, the U.S. Supreme Court overturned Roe v. Wade on June 24, 2022. Abortion will be banned in more and more states. Since the health data stored in the period tracking apps can be used to infer whether the user has had or is considering an abortion, mobile users are worrying that these apps may disclose their sensitive information, which can be used to prosecute users. Although period tracking apps have received attention from the research community, no existing work has performed a systematic privacy analysis of these apps, especially in the Post-Roe v. Wade era. To fill the void, this paper presents a comprehensive privacy analysis of popular period tracking apps. We first collect 35 popular period tracking apps from Google Play. Then, we analyze the sensitive user data collected by the period tracking apps using traffic analysis and static analysis. Further we inspect their privacy policies and check the consistency of the privacy policy with the app’s behavior. In addition, we analyze the app reviews to understand the users’ concerns about the period tracking apps. Our study reveals that some period tracking apps have indeed collected sensitive information and have the potential to share the data with third-party authorities. It is urgent for these apps to take action to protect user privacy, and mobile users should pay special attention to this kind of apps they used.
Zikan Dong, Liu Wang 0002, Guoai Xu, Haoyu Wang 0001
ASE2
2022 What did you pack in my app? a systematic analysis of commercial Android packers
abstract
Commercial Android packers have been widely used by developers as a way to protect their apps from being tampered with. However, app packer is usually provided as an online service developed by security vendors, and the packed apps are well protected. It is thus hard to know what exactly is packed in the app, and few existing studies in the community have systematically analyzed the behaviors of commercial app packers. In this paper, we propose PackDiff, a dynamic analysis system to inspect the fine-grained behaviors of commercial packers. By instrumenting the Android system, PackDiff records the runtime behaviors of Android apps (e.g., Linux system call invocations, Java API calls, Binder interactions, etc.), which are further processed to pinpoint the additional sensitive behaviors introduced by packers. By applying PackDiff to roughly 200 apps protected by seven commercial packers, we observe the disappointing facts of existing commercial packers. Most app packers have introduced unnecessary behaviors (e.g., accessing sensitive data), serious performance and compatibility issues, and they can even be abused to create evasive malware and repackaged apps, which contradicts with their design purposes.
Zikan Dong, Hongxuan Liu, Liu Wang 0002, Xiapu Luo, Yao Guo 0001, Guoai Xu, Xusheng Xiao, Haoyu Wang 0001
ESEC/SIGSOFT FSE3
2022 Demystifying "removed reviews" in iOS app store
abstract
The app markets enable users to submit feedback for downloaded apps in the form of star ratings and text reviews, which are meant to be helpful and trustworthy for decision making to both developers and other users. App markets have released strict guidelines/policies for user review submissions. However, there has been growing evidence showing the untrustworthy and poor-quality of app reviews, making the app store review environment a shambles. Therefore, review removal is a common practice, and market maintainers have to remove undesired reviews from the market periodically in a reactive manner. Although some reports and news outlets have mentioned removed reviews, our research community still lacks the comprehensive understanding of the landscape of this kind of reviews. To fill the void, in this paper, we present a large-scale and longitudinal study of removed reviews in iOS App Store. We first collaborate with our industry partner to collect over 30 million removed reviews for 33,665 popular apps over the course of a full year in 2020. This comprehensive dataset enables us to characterize the overall landscape of removed reviews. We next investigate the practical reasons leading to the removal of policy-violating reviews, and summarize several interesting reasons, including fake reviews, offensive reviews, etc. More importantly, most of these mis-behaviors can be reflected on reviews’ basic information including the posters, narrative content, and posting time. It motivates us to design an automated approach to flag the policy-violation reviews, and our experiment result on the labelled benchmark can achieve a good performance (F1=97%). We further make an attempt to apply our approach to the large-scale industry setting, and the result suggests the promising industry usage scenario of our approach. Our approach can act as a gatekeeper to pinpoint policy-violation reviews beforehand, which will be quite effective in improving the maintenance process of app reviews in the industrial setting.
Liu Wang 0002, Haoyu Wang 0001, Xiapu Luo, Tao Zhang 0001, Shangguang Wang, Xuanzhe Liu
ESEC/SIGSOFT FSE1
2021 A Large-Scale Empirical Study of COVID-19 Themed GitHub Repositories
abstract
Ever since the beginning of the outbreak of the COVID-19 pandemic, researchers from interdisciplinary domains have worked together to fight against the crisis. The open source community, plays a vital role in coping with the pandemic which is inherently a collaborative process. Plenty of COVID-19 related datasets, tools, software, deep learning models, are created and shared in research communities with great efforts. However, COVID-19 themed open source projects have not been systematically studied, and we are still unaware how the open source community helps combat COVID-19 in practice. To fill this void, in this paper, we take the first step to study COVID-19 themed repositories in GitHub, one of the most popular collaborative platforms. We have collected over 67K COVID-19 themed GitHub repositories till July 2020. We then characterize them from a number of aspects and classify them into six categories. We further investigate the contribution patterns of the contributors, and development and maintenance patterns of the repositories. This study sheds light on the promising direction of adopting open source technologies and resources to rapidly tackle the worldwide public health emergency in practice, and reveals existing challenges for improvement.
Liu Wang 0002, Guangdong Bai, Haoyu Wang 0001
COMPSAC1
2021 Understanding the Impact of COVID-19 on Github Developers: A Preliminary Study
abstract
The ongoing COVID-19 pandemic has impact almost every aspect of human lives profoundly.This paper investigates the impact of COVID-19 on the activity and contribution of open source software (OSS) developers.Specifically, we make great efforts to harvest the information of all the developers (over 25 million) on GitHub and their contribution activities.With such a large-scale dataset, we perform analysis from four perspectives, including the overall ecosystem level, country level, organization level and developer level, to characterize the impact of COVID-19 on the OSS community.We have revealed a number of interesting observations and trends, which are crucial to understanding the OSS contributors and supporting the collaboration to combat global crisis like COVID-19.
Liu Wang 0002, Guangdong Bai, Weihang Su, Haoyu Wang 0001
SEKE1
2021 A Longitudinal Study of Removed Apps in iOS App Store
abstract
To improve app quality and nip the potential threats in the bud, modern app markets have released strict guidelines along with app vetting process before app publishing. However, there has been growing evidence showing the ineffectiveness of app vetting, making potentially harmful and policy-violation apps sneak into the market from time to time. Therefore, app removal is a common practice, and market maintainers have to remove undesired apps from the market periodically in a reactive manner. Although a number of reports and news media have mentioned removed apps, our research community still lacks the comprehensive understanding of the landscape of this kind of apps. To fill the void, in this paper, we present a large-scale and longitudinal study of removed apps in iOS app store. We first make great efforts to record daily snapshot of iOS app store continuously in a span of 1.5 years. By comparing each two consecutive snapshots, we have collected the information of over 1 million removed apps with their accurate removed date. This comprehensive dataset enables us to characterize the overall landscape of removed apps. We observe that, although most of the removed apps are low-quality apps (e.g., outdated and abandoned), a number of the removed apps are quite popular. We further investigate the practical reasons leading to the removal of such popular apps, and observe several interesting reasons, including ranking fraud, fake description, and content issues, etc. More importantly, most of these mis-behaviors can be reflected on app meta information including app description, app review, and ASO keywords. It motivates us to design an automated approach to flagging the removed apps. Experiment result suggests that, even without accessing to the bytecode of mobile apps, we can identify the removed apps with good performance (F1=83%). Furthermore, we are able to flag the removed apps in advance as long as their inappropriate behaviors appear in their metadata. We believe our approach can work as a whistle blower that pinpoints policy-violation behaviors timely, which will be quite effective in improving the app maintenance process.
Fuqi Lin, Haoyu Wang 0001, Liu Wang 0002, Xuanzhe Liu
WWW3
2021 Beyond the virus: a first look at coronavirus-themed Android malware
Liu Wang 0002, Haoyu Wang 0001, Pengcheng Xia 0001, Yuanchun Li 0003, Lei Wu 0012, Yajin Zhou, Xiapu Luo, Yulei Sui, Yao Guo 0001, Guoai Xu
Empir. Softw. Eng.1
2020 MadDroid: Characterizing and Detecting Devious Ad Contents for Android Apps
abstract
Advertisement drives the economy of the mobile app ecosystem. As a key component in the mobile ad business model, mobile ad content has been overlooked by the research community, which poses a number of threats, e.g., propagating malware and undesirable contents. To understand the practice of these devious ad behaviors, we perform a large-scale study on the app contents harvested through automated app testing. In this work, we first provide a comprehensive categorization of devious ad contents, including five kinds of behaviors belonging to two categories: ad loading content and ad clicking content. Then, we propose MadDroid, a framework for automated detection of devious ad contents. MadDroid leverages an automated app testing framework with a sophisticated ad view exploration strategy for effectively collecting ad-related network traffic and subsequently extracting ad contents. We then integrate dedicated approaches into the framework to identify devious ad contents. We have applied MadDroid to 40,000 Android apps and found that roughly 6% of apps deliver devious ad contents, e.g., distributing malicious apps that cannot be downloaded via traditional app markets. Experiment results indicate that devious ad contents are prevalent, suggesting that our community should invest more effort into the detection and mitigation of devious ads towards building a trustworthy mobile advertising ecosystem.
Tianming Liu 0002, Haoyu Wang 0001, Li Li 0029, Xiapu Luo, Feng Dong 0008, Yao Guo 0001, Liu Wang 0002, Tegawendé F. Bissyandé, Jacques Klein
WWW7