Lingyun Ying

dblp:18/7667 · DBLP profile ↗
← Back
42ranked-venue papers
1as first author
28since 2021 · last 2026
0000-0001-7445-9103ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 32 · 1 first-author · 18 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SPDAgent: Leveraging LLM Agents for Context-Aware Binary Security Patch Detection via Pseudocode Diff Analysis
Fengrui Yang, Xi Xiao 0001, Lingyun Ying, Chao Zhang 0008, Qing Li 0006
DSN4
2026 From Noise to Signal: Precisely Identify Affected Packages of Known Vulnerabilities in npm Ecosystem
Yingyuan Pu, Lingyun Ying, Yacong Gu
NDSS2
2026 Cache Me, Catch You: Cache Related Security Threats in LLM Serving Frameworks
XiangFan Wu, Lingyun Ying, Yacong Gu, Haipeng Qu
NDSS2
2026 Understanding the Status and Strategies of the Code Signing Abuse Ecosystem
Yiming Zhang 0009, Lingyun Ying, Mingming Zhang 0010, Baojun Liu 0002, Hai-Xin Duan, Zi-Quan You
NDSS3
2026 From Obfuscated to Obvious: A Comprehensive JavaScript Deobfuscation Tool for Security Analysis
Dongchao Zhou, Lingyun Ying, Huajun Chai, Dongbin Wang
NDSS2
2026 PIEDChecker: Uncover Permissions-Independent Emulation-Detection Methods in Android System
abstract
For compatibility checks and preventing malicious cheating behaviors in Android systems, It is convenient for benign app developers to utilize emulation-detection technology. However, this technique has been abused by malicious app developers, which causes detection emulation and behavior change, known as anti-emulation behavior, to evade the dynamic analysis performed via the Android emulator. The Android permission mechanism can limit some anti-emulation behaviors, but attackers can still use permission-independent (PI) emulation-detection technology to achieve their goals. In this paper, we propose a static and dynamic combined detection framework namedPIEDCheckerto detect PI emulation-detection apps. This framework can statically identify PI emulation-detection code and dynamically verify PI anti-emulation behaviors.PIEDChecker's performance is validated by 382 manually created test apps and 344 apps with emulation-detection labels. Moreover,PIEDCheckerhas higher accuracy in detecting PI emulation-detection compared with the existing Android malware analysis platforms and academic methods. Moreover, it is found that there are 13,377 apps having PI emulation-detection behaviors within the tested 25,303 apps collected over the past five years. In particular, the detection features of the PI emulation-detection methods are summarized based on the evaluation result.
Weina Niu, Qinsheng Hou, Lingyun Ying, Xiaosong Zhang 0001
IEEE Trans. Dependable Secur. Comput.5
2025 Detecting Malicious Encrypted Traffic with Multimodal Representations
abstract
The rapid advancement of encryption technology enhances network security while enabling hidden attackers to avoid detection. Traditional methods for malicious encrypted traffic detection, which predominantly rely on a single modality such as statistical features or content representations, often fall short of adapting to dynamic network environments. Methods based on graph representations grapple with challenges such as insufficient modeling of the encryption properties and substantial computational resource requirements. Multimodal-based methods seldom consider the graph-based dynamic representation and often overlook the differences in feature spaces. Moreover, these methods are not evaluated for universality across platforms. To solve challenges above, we propose M2D, a multimodal-based framework for malicious encrypted traffic detection suitable for all versions of TLS protocols. M2D extracts (a) heterogeneous graph representation from spatial and temporal features to capture both dynamic patterns and complex interactions between different entities; (b) ciphertext visual representation to enhance content encapsulation; and (c) plaintext representation to explore semantics, then fuses them through the multi-head attention mechanism to emphasize more effective components. Furthermore, we set up an encrypted network traffic dataset generated by sandbox, with session keys embedded for decryption. Experimental results on both public and proposed datasets demonstrate the superior performance of M2D in binary and multi-class classification tasks. Additionally, ablation studies confirm the effectiveness of each component.
Ruijie Zhao 0001, Libo Chen 0001, Lingyun Ying, Zhengguang Han, Zhi Xue
ICC5
2025 Exposing the Hidden Layer: Software Repositories in the Service of Seo Manipulation
abstract
Distinct from traditional malicious packages, this paper uncovers a novel attack vector named “blackhat Search Engine Optimization through REPositories (RepSEO)”. In this approach, attackers carefully craft packages to manipulate search engine results, exploiting the credibility of software repositories to promote illicit websites. Our research presents a systematic analysis of the underground ecosystem of RepSEO, identifying key players such as account providers, advertisers, and publishers. We developed an effective detection tool, applied to a ten-year large-scale dataset of npm, Docker Hub, and NuGet software repositories. This investigation led to the startling discovery of 3,801,682 abusive packages, highlighting the widespread nature of this attack. Our study also delves into the supply chain tactics of these attacks, revealing strategies like the use of self-hosted email services for account registration, redirection methods to obscure landing pages, and rapid deployment techniques by aggressive attackers. Additionally, we explore the profit motives behind these attacks, identifying two primary types of advertisers: survey-based advertisers and malware distribution advertisers. We reported npm, NuGet, and Docker Hub about the RepSEO packages and the related supply chain vulnerabilities of Google, and received their acknowledgments. Software repositories have started removing the abusive packages as of this paper's submission. We also opensource our code and data to facilitate future research.
Mengying Wu, Geng Hong, Wuyuao Mai, Lei Zhang 0096, Yingyuan Pu, Huajun Chai, Lingyun Ying, Hai-Xin Duan, Min Yang 0002
ICSE8
2025 Distilling Benign Knowledge with Fine-Grained AST Fragments for Precise Real-World Web Shell Detection
abstract
Web shell detection has become increasingly crucial with the expansion of cloud computing, where automated malware analysis serves as a foundational approach. A key challenge in malware detection lies in balancing the reduction of false positives with maintaining detection accuracy amid rapid software ecosystem evolution. Existing methods require substantial expert intervention to mitigate false positives and often neglect the resource-intensive measures required to address model degradation caused by software updates. This study introduces ASTBAR, a novel method that extracts fine-grained AST fragments to distill benign behavioral knowledge from webserver software. By leveraging program structure and semantic analysis, ASTBAR generates fragment-level representations of benign samples and employs fragment matching to identify malware. Unlike prior techniques, ASTBAR achieves simultaneous improvements in precision, recall, and adaptability to software evolution. The evaluation results demonstrate that ASTBAR achieves an F1 score of$\mathbf{6 5. 3 5 \%}$, outperforming the state-of-theart methods by$\mathbf{1 0. 3 9 \%}$. In a$\mathbf{1 2}$-month industrial deployment spanning over one million users, ASTBAR maintained a 97.63% recall rat while reducing false positives by 700+ cases daily (equivalent to 30 expert hours).
Mingzhe Gao, Ligeng Chen, Yiling He, Lingyun Ying
IWQoS5
2025 Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild
abstract
Secrets, whether structured like API keys or un-structured like passwords, are essential for securing applications and services. However, the growing use of open -source projects and rapid development cycles has amplified the risk of secret leakage. Current detection tools suffer from high false positive rates and low recall due to simplistic methods like regular expressions and entropy checks, often missing unstructured secrets or mislabeling non-sensitive data. In this paper, we introduce Keysentinel, an advanced automated secret detection tool that addresses these limitations through machine learning, semantic analysis, and prefix matching. To evaluate KEYSENTINEL, we created the first cross-platform benchmark with 11,826 labeled secrets in 1,806,530 files across GitHub, PyPI, and WeChat. We compare Key-sentinelwith six currently available tools. The results show KEYSENTINEL achieves state-of-the-art performance, with precision (91.18%), recall (81.71%), and an F1 score (0.86), surpassing industry-standard tools and significantly reducing false positives. It also outperforms large language models like GPT-4 and o1 in accuracy and cost-effectiveness. Besides, we conduct a large-scale measurement study, analyzing 80,330,098 files from GitHub, PyPI, and WeChat. We found that up to 30% of projects are at risk of secret leaks. Furthermore, we also scan the code base of an IT company to assess real-world secret leakage risks. Our findings underscore the pervasive nature of secret leaks and highlight the urgent need for enhanced secret management practices across platforms.
Zidong Zhang, Lingyun Ying, Huajun Chai, Jiuxin Cao, Hai-Xin Duan
SP3
2025 Dr. Docker: A Large-Scale Security Measurement of Docker Image Ecosystem
abstract
Docker has transformed modern software development, enabling the widespread reuse of containerized applications. Currently, Docker images are primarily distributed through centralized registries, among which Docker Hub is the largest, allowing developers to share and reuse images easily. The threats within these images also spread through the supply chain via dependency relationships, posing risks to anyone using the image and all images built based on it. However, it is unclear to what extent the threats within Docker images are distributed and propagated.
Hequan Shi, Lingyun Ying, Libo Chen 0001, Hai-Xin Duan, Zhi Xue
WWW2
2025 Magnifier: Detecting Network Access via Lightweight Traffic-Based Fingerprints
Wenhao Li 0005, Qiang Wang 0059, Huaifeng Bao, Xiaoyu Zhang 0002, Lingyun Ying, Zhaoxuan Li, Huamin Jin, Shuai Wang 0079
IEEE Trans. Inf. Forensics Secur.5
2024 Toward Understanding the Security of Plugins in Continuous Integration Services
abstract
Mainstream Continuous Integration (CI) platforms have provided the plugin functionality to accelerate the development of CI pipelines. Unfortunately, CI plugins, which are essentially reusable code snippets, also expose new attack surfaces as plugins might be developed by less trusted users. In this paper, we present an in-depth study to understand potential security risks in existing CI plugins. We conduct a comprehensive analysis of plugin implementations on four mainstream CI platforms (GitHub Actions, GitLab CI, CircleCI, and Azure Pipelines), and investigate several weak links in existing plugin distributions and isolation mechanisms. We investigate seven attack vectors that can enable attackers to hijack plugins and distribute malicious code without plugins users being aware, and further exploit hijacked plugins to manipulate the workflow execution. Additionally, we find that plugin dependency (a plugin references other plugins) might further amplify the attack impact of our disclosed attacks. To evaluate the potential impact, we conduct a large-scale measurement on GitHub and GitLab, covering a total of 1,328,912 repositories using the aforementioned CI platforms. Our measurement results show that a large number of repositories and existing plugins, including many widely used ones, are potentially vulnerable to the proposed attacks. We have duly reported the identified vulnerabilities and received positive responses.
Xiaofan Li 0009, Yacong Gu, Chu Qiao, Zhenkai Zhang 0002, Daiping Liu, Lingyun Ying, Hai-Xin Duan, Xing Gao 0001
CCS6
2024 PowerPeeler: A Precise and General Dynamic Deobfuscation Method for PowerShell Scripts
abstract
PowerShell is a powerful and versatile task automation tool. Unfortunately, it is also widely abused by cyber attackers. To bypass malware detection and hinder threat analysis, attackers often employ diverse techniques to obfuscate malicious PowerShell scripts. Existing deobfuscation tools suffer from the limitation of static analysis, which fails to simulate the real deobfuscation process accurately. Accurate, complete, and robust PowerShell script deobfuscation is still a challenging problem.
Huajun Chai, Lingyun Ying, Hai-Xin Duan, Jun Tao 0003
CCS4
2024 MiniCAT: Understanding and Detecting Cross-Page Request Forgery Vulnerabilities in Mini-Programs
abstract
Mini-programs are lightweight apps running in super apps (such as WeChat, Baidu, Alipay, and TikTok), an emerging paradigm in the era of mobile computing. With the growing popularity of mini-programs, there is an increasing concern for their security and privacy. In essence, mini-programs are WebView-based apps. This means that they may be vulnerable to the same security risks associated with web apps. In this work, we discovered a new mini-program vulnerability called MiniCPRF (Cross-Page Request Forgery in Mini-Programs). The exploit of this vulnerability is easy, and the attack consequences are severe, leading to unauthorized operations, such as free shopping, and the exposure of confidential information, such as credit card numbers. The root causes of MiniCPRF can be attributed to multiple design flaws in both mini-programs and their super apps, including the insecure routing mechanism, lack of message integrity check, and plain-text storage. To evaluate the impacts of MiniCPRF, we designed an automated analysis framework called MiniCAT. It can automatically crawl mini-programs, perform static analysis on them, and generate detection reports. In large-scale real-world evaluations with MiniCAT, we identified that 32.0% (13,349/41,726) of analyzable mini-programs are potentially vulnerable to MiniCPRF, including some famous ones with millions of users, such as Sohu and Wenjuanxing. Following the responsible disclosure principle, we have reported verified vulnerable mini-programs to the corresponding vendors and developers, and three real-world cases have been confirmed by CNVD. Additionally, we suggest mitigation strategies to resolve the security issue related to MiniCPRF.
Zidong Zhang, Qinsheng Hou, Lingyun Ying, Wenrui Diao, Yacong Gu, Rui Li 0102, Shanqing Guo, Hai-Xin Duan
CCS3
2024 Enhanced Fast and Reliable Statistical Vulnerability Root Cause Analysis with Sanitizer
abstract
Vulnerability root cause analysis (RCA) is a crucial step following the discovery of vulnerabilities. When faced with a multitude of crashes resulting from fuzzing, effective RCA results can assist developers in swiftly identifying and rectifying the root causes of vulnerabilities. Recently, some methods that rely on statistical behavioral differences to analyze the root causes of vulnerabilities have been introduced. However, they suffer from issues such as high time costs, strong randomness, and imprecise results, rendering them impractical for real-world applications. In this paper, we propose an efficient and accurate statistical analysis-based vulnerability RCA approach named RCLocator. We introduce an enhanced crash information tuple extraction tool based on sanitizer to ensure crash consistency during the mutation process of original files. This approach reduces the time cost of the data augmentation stage and enhances the accuracy of RCA. Furthermore, it provides developers with effective explanations for root cause predicates. We evaluate our approach on RCABench and real-world vulnerabilities. The results indicate that RCLocator significantly outperforms state-of-the-art methods, the probability of obtaining correct root cause analysis results increased by 46.7%, and 9.0 times faster in terms of time.
Zhuo Yan, Haipeng Qu, Lingyun Ying, Q. Chao
ICST3
2024 More Haste, Less Speed: Cache Related Security Threats in Continuous Integration Services
abstract
Continuous Integration (CI) platforms have widely adopted caching to speed up CI task executions by storing and reusing dependent packages. Unfortunately, CI cache also exposes new attack surfaces when cache objects are shared across trust boundaries. In this paper, we systematically investigate potential security threats of CI cache features in seven mainstream CI platforms (CIPs). We find that existing CIPs have isolation issues in their cache sharing and inheritance strategies, potentially raising cache poisoning and data leakage problems. By exploiting these vulnerable mechanisms, we further uncover four attack vectors enabling attackers to stealthily inject malicious code into the cache or steal sensitive data. Even worse, many CIPs provide vulnerable official cache templates that will mistakenly store and expose sensitive data in the cache by default. To understand the potential impact of our disclosed threats, we develop an analysis tool and conduct a large-scale measurement on open-source repositories. Our measurement results show that many popular repositories are potentially affected by these attacks. We also identify 78 repositories that expose their high-value secrets in cache objects and are at risk of secret leakage. We have duly reported identified vulnerabilities to corresponding stakeholders and received positive responses.
Yacong Gu, Lingyun Ying, Huajun Chai, Yingyuan Pu, Hai-Xin Duan, Xing Gao 0001
SP2
2024 Maldet: An Automated Malicious npm Package Detector Based on Behavior Characteristics and Attack Vectors
abstract
With the growing number of software developments and the expansion of functionalities, more and more developers tend to use third-party packages to speed up development and improve efficiency. Due to the openness of npm and the widespread use of Node.js, npm has become the largest open source software ecosystem and therefore a prime target for malicious attackers. Attackers often use various attack vectors to release new malicious packages or destroy existing benign packages, thus posing threats to software that rely on these packages. Therefore, detecting malicious npm packages is of great significance for protecting user security and building a more reliable and secure npm open source software ecosystem.We propose Maldet, an automated malicious npm package detection method. Maldet uses the malicious behavior pattern library we have built as features, training known malicious and benign samples with four different classifiers. In addition, we use a supplementary detector to perform a secondary detect on packages that detected benign, thereby reducing the number of false negatives. The results show that Maldet can detect each package in an average of just a few seconds with 97.12% accuracy, which is superior to other tools, providing high accuracy and fast classification capability.
Haipeng Qu, Lingyun Ying, Linghui Wang
TrustCom3
2024 From Promises to Practice: Evaluating the Private Browsing Modes of Android Browser Apps
abstract
Private browsing is a common feature of web browsers on desktop platforms. This feature protects the privacy of users browsing the Internet and, therefore, is widely welcomed by users. In recent years, with the popularity of smartphones, the private browsing mode has been introduced into mobile browsers. However, its deployment on mobile platforms has not been well evaluated. To bridge the gap, in this work, we systemically studied the private browsing modes of Android browser apps. Specifically, we proposed six private rules for mobile browsers to follow by combining the mobile browsing features with the previous research on private browsing. Furthermore, we designed an automated analysis framework, BroDroid, to detect whether mobile browsers violate these rules. Also, with BroDroid, we evaluated 49 popular browser apps crawled from Google Play. Finally, BroDroid successfully identified 58 violations, some of which come from the promised capabilities of the browser. We reported our discovered issues to the corresponding developers, and four of them (Yandex Browser, Mint Browser, Web Explorer, and Net Fast Web Browser) have acknowledged our findings. Our observation may be the tip of the iceberg, and more efforts should be put into improving the privacy protections of mobile browsers.
Xiaoyin Liu, Wenzhi Li, Qinsheng Hou, Shishuai Yang, Lingyun Ying, Wenrui Diao, Shanqing Guo, Hai-Xin Duan
WWW5
2023 Wolf in Sheep's Clothing: Evaluating Security Risks of the Undelegated Record on DNS Hosting Services
abstract
Leveraging DNS for covert communications is appealing since most networks allow DNS traffic, especially the ones directed toward renowned DNS hosting services. Unfortunately, most DNS hosting services overlook domain ownership verification, enabling miscreants to host undelegated DNS records of a domain they do not own. Consequently, miscreants can conduct covert communication through such undelegated records for whitelisted domains on reputable hosting providers. In this paper, we shed light on the emerging threat posed by undelegated records and demonstrate their exploitation in the wild. To the best of our knowledge, this security risk has not been studied before.
Fenglu Zhang, Baojun Liu 0002, Eihal Alowaisheq, Lingyun Ying, Xiang Li 0108, Zaifeng Zhang, Ying Liu 0024, Hai-Xin Duan, Min Zhang 0054
IMC5
2023 Continuous Intrusion: Characterizing the Security of Continuous Integration Services
abstract
Continuous Integration (CI) is a widely-adopted software development practice for automated code integration. A typical CI workflow involves multiple independent stakeholders, including code hosting platforms (CHPs), CI platforms (CPs), and third party services. While CI can significantly improve development efficiency, unfortunately, it also exposes new attack surfaces. As the code executed by a CI task may come from a less-trusted user, improperly configured CI with weak isolation mechanisms might enable attackers to inject malicious code into victim software by triggering a CI task. Also, one insecure stakeholder can potentially affect the whole process. In this paper, we systematically study potential security threats in CI workflows with multiple stakeholders and major CP components considered. We design and develop an analysis tool, CInspector, to investigate potential vulnerabilities in seven popular CPs, when integrated with three mainstream CHPs. We find that all CPs have the risk of token leakage caused by improper resource sharing and isolation, and many of them utilize over-privileged tokens with improper validity periods. We further reveal four novel attack vectors that allow attackers to escalate their privileges and stealthy inject malicious code by executing a piece of code in a CI task. To understand the potential impact, we conduct a large-scale measurement on the three mainstream CHPs, scrutinizing over 1.69 million repositories. Our quantitative analysis demonstrates that some very popular repositories and large organizations are affected by these attacks. We have duly reported the identified vulnerabilities to CPs and received positive responses.
Yacong Gu, Lingyun Ying, Huajun Chai, Chu Qiao, Hai-Xin Duan, Xing Gao 0001
SP2
2023 Investigating Package Related Security Threats in Software Registries
abstract
Package registries host reusable code assets, allowing developers to share and reuse packages easily, thus accelerating the software development process. Current software registry ecosystems involve multiple independent stakeholders for package management. Unfortunately, abnormal behavior and information inconsistency inevitably exist, enabling adversaries to conduct malicious activities with minimal effort covertly. In this paper, we investigate potential security vulnerabilities in six popular software registry ecosystems. Through a systematic analysis of the official registries, corresponding registry mirrors and registry clients, we identify twelve potential attack vectors, with six of them disclosed for the first time, that can be exploited to distribute malicious code stealthily. Based on these security issues, we build an analysis framework, RScouter, to continuously monitor and uncover vulnerabilities in registry ecosystems. We then utilize RScouter to conduct a measurement study spanning one year over six registries and seventeen popular mirrors, scrutinizing over 4 million packages across 53 million package versions. Our quantitative analysis demonstrates that multiple threats exist in every ecosystem, and some have been exploited by attackers. We have duly reported the identified vulnerabilities to related stakeholders and received positive responses.
Yacong Gu, Lingyun Ying, Yingyuan Pu, Huajun Chai, Xing Gao 0001, Hai-Xin Duan
SP2
2023 RecMaL: Rectify the malware family label via hybrid analysis
Mingzhe Gao, Ligeng Chen, Zhengxuan Liu, Lingyun Ying
Comput. Secur.5
2023 Can We Trust the Phone Vendors? Comprehensive Security Measurements on the Android Firmware Ecosystem
abstract
Android is the most popular smartphone platform with over 85% market share. Its success is built on openness, and phone vendors can utilize the Android source code to make customized products with unique software/hardware features. On the other hand, the fragmentation and customization of Android also bring many security risks that have attracted the attention of researchers. Many efforts were put in to investigate the security of customized Android firmware. However, most of the previous works focus on designing efficient analysis tools or analyzing particular aspects of the firmware. There still lacks a panoramic view of Android firmware ecosystem security and the corresponding understandings based on large-scale firmware datasets. In this work, we made a large-scale comprehensive measurement of the Android firmware ecosystem security. Our study is based on 8,325 firmware images from 153 vendors and 813 Android-related CVEs, which is the largest Android firmware dataset ever used for security measurements. In particular, our study followed a series of research questions, covering vulnerabilities, patches, security updates, and pre-installed apps. To automate the analysis process, we designed a framework,AndScanner+, to complete firmware crawling, firmware parsing, patch analysis, and app analysis. Through massive data analysis and case explorations, several interesting findings are obtained. For example, the patch delay and missing issues are widespread in Android firmware images, say 31.4% and 5.6% of all images, respectively. The latest images of several phones still contain vulnerable pre-installed apps, and even the corresponding vulnerabilities have been publicly disclosed. In addition to data measurements, we also explore the causes behind these security threats through case studies and demonstrate that the discovered security threats can be converted into exploitable vulnerabilities. There are 46 new vulnerabilities found byAndScanner+, 36 of which have been assigned CVE/CNVD IDs. This study provides much new knowledge of the Android firmware ecosystem with a deep understanding of software engineering security practices.
Qinsheng Hou, Wenrui Diao, Chenglin Mao, Lingyun Ying, Xiaofeng Liu 0013, Yuanzhi Li, Shanqing Guo, Meining Nie, Hai-Xin Duan
IEEE Trans. Software Eng.5
2022 DitDetector: Bimodal Learning based on Deceptive Image and Text for Macro Malware Detection
abstract
Macro malware has always been a severe threat to cyber security although the Microsoft Office suite applies the default macro-disabling policy. Among the defense solutions at different stages of the attack chain, document analysis is more targeted through detecting malicious documents with macro malware. It is effective, especially with machine learning methods, but still faces problems handling malware variants, supporting file formats, and attack countermeasures with advanced attack techniques (e.g., Excel 4.0 macro and remote template injection).
Jia Yan 0004, Xiangkun Jia, Lingyun Ying, Purui Su, Zhanyi Wang
ACSAC4
2022 Invoke-Deobfuscation: AST-Based and Semantics-Preserving Deobfuscation for PowerShell Scripts
abstract
In recent years, PowerShell has been widely used in cyber attacks and malicious PowerShell scripts can easily evade the detection of anti-virus software through obfuscation. Existing deobfuscation tools often fail to recover obfuscated scripts correctly due to imprecise obfuscation identification, improper recovery and wrong replacement. In this paper, we propose an AST-based and semantics-preserving deobfuscation approach, Invoke-Deobfuscation. It utilizes recoverable nodes of Abstract Syntax Tree to identify obfuscated pieces precisely, simulates the recovery process through Invoke function and variable tracing, and replaces obfuscated pieces in place to keep the original semantics. We build a large evaluation dataset containing 39,713 wild PowerShell scripts. Compared with the state-of-the-art tools, the experimental results show Invoke-Deobfuscation performs most efficiently. It recovers much more key information than others and significantly reduces samples’ obfuscation score, on average, by 46%. Moreover, 100% of Invoke-Deobfuscation’s results have the same network behavior as the original scripts.
Huajun Chai, Lingyun Ying, Hai-Xin Duan, Daren Zha
DSN2
2022 Large-scale Security Measurements on the Android Firmware Ecosystem
abstract
Android is the most popular smartphone platform with over 85% market share. Its success is built on openness, and phone vendors can utilize the Android source code to make products with unique software/hardware features. On the other hand, the fragmentation and customization of Android also bring many security risks that have attracted the attention of researchers. Many efforts were put in to investigate the security of customized Android firmware. However, most of the previous work focuses on designing efficient analysis tools or analyzing particular aspects of the firmware. There still lacks a panoramic view of Android firmware ecosystem security and the corresponding understandings based on large-scale firmware datasets. In this work, we made a large-scale comprehensive measurement of the Android firmware ecosystem security. Our study is based on 6,261 firmware images from 153 vendors and 602 Android-related CVEs, which is the largest Android firmware dataset ever used for security measurements. In particular, our study followed a series of research questions, covering vulnerabilities, patches, security updates, and pre-installed apps. To automate the analysis process, we designed a framework, AndScanner, to complete ROM crawling, ROM parsing, patch analysis, and app analysis. Through massive data analysis and case explorations, several interesting findings are obtained. For example, the patch delay and missing issues are widespread in Android images, say 24.2% and 6.1% of all images, respectively. The latest images of several phones still contain vulnerable pre-installed apps, and even the corresponding vulnerabilities have been publicly disclosed. In addition to data measurements, we also explore the causes behind these security threats through case studies and demonstrate that the discovered security threats can be converted into exploitable vulnerabilities via 38 newfound vulnerabilities by our framework, 32 of which have been assigned CVE/CNVD numbers. This study provides much new knowledge of the Android firmware ecosystem with deep understanding of software engineering security practices.
Qinsheng Hou, Wenrui Diao, Xiaofeng Liu 0013, Lingyun Ying, Shanqing Guo, Yuanzhi Li, Meining Nie, Hai-Xin Duan
ICSE6
2022 Understanding and Mitigating Label Bias in Malware Classification: An Empirical Study
abstract
Machine learning techniques are promising for malware classification, but there is a neglected problem of label bias in the annotation process which decreases the performance in practice. To understand the label bias problems and existing solutions, we conduct an empirical study based on two Portable Executable (PE) malware sample datasets (i.e., open-sourced BODMAS with 52,793 samples and a new collected MAIN dataset of 153,811 samples), and 67 anti-virus engines in VirusTotal. We first show the two ways of label bias problems, including chaotic naming rules and annotation inconsistency. Then we present the effects of two solutions (i.e., electing one reputable AV engine and aggregating multiple labels based on majority voting) and find they face the problems of feature preference and engine independence. Finally, we propose some recommendations for improvements and get a 7.79% increase in the F1 score (i.e., from 84.83% to 92.62%). The dataset will be open-source for further study.
Jia Yan 0004, Xiangkun Jia, Lingyun Ying, Purui Su
QRS3
2020 NativeX: Native Executioner Freezes Android
abstract
Android is a Linux-based multi-thread open-source operating system that dominates 85% of the worldwide smartphone market share. Though Android has its established management for its framework layer processes, we discovered for the first time that the weak management of native processes is posing tangible threats to Android systems from version 4.2 to 9.0. As a consequence, any third-party application without any permission can freeze the system or force the system to go through a reboot by starving or significantly delaying the critical system services using Android commands in its native processes. We design NativeX to systematically analyze the Android source code to identify the risky Android commands. For each identified risky command, NativeX can automatically generate the PoC (Proof-of-Concept) application, and verify the effectiveness of the generated PoC. We conduct manual vulnerability analysis to reveal two root causes beyond the superficial attack consequences. We further carry out quantitative experiments to demonstrate the attack consequences, including the device temperature surge, the battery degeneration, and the computing performance decrease, based on which, three representative PoC attacks are engineered. Finally, we discuss possible defense approaches to improve the management of Android native processes.
Qinsheng Hou, Lingyun Ying
AsiaCCS3
2017 JGRE: An Analysis of JNI Global Reference Exhaustion Vulnerabilities in Android
abstract
Android system applies a permission-based security model to restrict unauthorized apps from accessing system services, however, this security model cannot constrain authorized apps from sending excessive service requests to exhaust the limited system resource allocated for each system service. As references from native code to a Java object, JNI Global References (JGR) are prone to memory leaks, since they are not automatically garbage collected. Moreover, JGR exhaustion may lead to process abort or even Android system reboot when the victim process could not afford the JGR requests triggered by malicious apps through inter-process communication. In this paper, we perform a systematic study on JGR exhaustion (JGRE) attacks against all system services in Android. Our experimental results show that among the 104 system services in Android 6.0.1, 32 system services have 54 vulnerabilities. Particularly, 22 system services can be successfully attacked without any permission support. After reporting those vulnerabilities to Android security team and getting confirmed, we study the existing ad hoc countermeasures in Android against JGRE attacks. Surprisingly, among the 10 system services that have been protected, 8 system services are still vulnerable to JGRE attacks. Finally, we develop an effective defense mechanism to defeat all identified JGRE attacks by adopting Android's low memory killer (LMK) mechanism.
Yacong Gu, Kun Sun 0001, Purui Su, Qi Li 0002, Yemian Lu, Lingyun Ying, Dengguo Feng
DSN6
2017 A study on a feasible no-root approach on Android
abstract
Root is the administrative privilege on Android, which is however inaccessible on stock Android devices. Due to the desire for privileged functionalities and the reluctance of rooting their devices, Android users seek for no-root approaches, which provide users with part of root privileges without rooting their devices. Existing no-root approaches require users to launch a separate service via Android Debug Bridge (ADB) on an Android device, which would perform user-desired tasks. However, it is unusual for a third-party Android application to work with a separate native service via sockets, and it requires the application developers to have extra knowledge such as Linux programming in application development. In this paper, we propose a feasible no-root approach based on new functionalities added on Android, which creates no separate service but an ADB loopback. To ensure such no-root approach is not misused in a proactive instead of reactive manner, we examine its dark side. We find out that while this approach makes it easy for no-root applications to work, it may lead to a “ permission explosion,” which enables any third-party application to attain shell permissions beyond its granted permissions. The permission explosion can further lead to exploits including privacy leakage, account takeover, application UID abuse, and user input inference. A practical experiment is carried out to evaluate the situation in the real world, which shows that many real-world applications from Google Play and four third-party application markets are indeed vulnerable to these exploits. To mitigate the dark side of the new no-root approach and make it more suitable for users to adopt, we identify the causes of the exploits, and propose a permission-based solution. We also provide suggestions to application developers and application markets on how to prevent these exploits.
Yingjiu Li, Robert H. Deng, Lingyun Ying
J. Comput. Secur.4
2016 Attacks and Defence on Android Free Floating Windows
abstract
Nowadays, the popular Android is so closely involved in people's daily lives that people rely on Android to perform critical operations and trust Android with sensitive information. It is of great importance to guarantee the usability and security of Android which, however, is such a huge system that a potential threat may arise from any part of it. In this paper, we focus on the Free Floating window (FF window) which is a category of windows that can appear freely above any other applications. It can share the screen space with other FF windows, dialogs, and activities. An FF window is flexible in both its appearance and behaviour features. We analyse the behaviour features of FF windows, including the priority in display layer and the capability of processing user-generated events. Three types of attacks via FF windows with delicate design in their appearance and behaviour features are demonstrated, i.e., DoS attack against Android system, GUI hijacking by targeting overlap, and input inference using FF windows as a side channel. To address the threat caused by FF windows, we design a priority framework for FF windows, which protects a sensitive activity/FF window declared by developers from being attacked by any malicious FF windows. A complementary solution is proposed to mitigate the confusion attack from malicious activities. Finally, we provide Android with suggestions on how to manage FF windows.
Lingyun Ying, Yemian Lu, Yacong Gu, Purui Su, Dengguo Feng
AsiaCCS1
2016 Exploiting Android System Services Through Bypassing Service Helpers
Yacong Gu, Lingyun Ying, Yemian Lu, Qi Li 0002, Purui Su
SecureComm3
2015 A Rapid and Scalable Method for Android Application Repackaging Detection
Sibei Jiao, Lingyun Ying, Purui Su, Dengguo Feng
ISPEC3
2015 Xede: Practical Exploit Early Detection
Meining Nie, Purui Su, Qi Li 0002, Zhi Wang 0004, Lingyun Ying, Dengguo Feng
RAID5
2014 Revisiting Node Injection of P2P Botnet
Jia Yan 0004, Lingyun Ying, Yi Yang 0040, Purui Su, Qi Li 0002, Dengguo Feng
NSS2
2014 Automated User Profiling in Location-Based Mobile Messaging Applications
abstract
Location-based messaging applications (LMAs), a kind of messaging applications for mobile devices which enable users to connect with people based on their geographical locations, have recently experienced a huge popularity growth. The killer feature in LMAs that embodies the concept of geo-based instant messaging, named people nearby, allows users at any place to search and communicate with other registered users nearby. In this paper, we discuss a common weakness in LMAs that relates to the abuse of the people nearby function. In this case, rich personal data of registered LMA users can be easily obtained, bringing a chance to perform automated user profiling in LMAs. Specifically, we build an automated and scalable system to construct extended profiles (or we call life profile) of LMA users, which contain not only personal information of LMA users but also the daily activities and social ties inferred from their leaked spatio-temporal privacy. The system is highly adaptable to various applications, requiring no modification of applications or trivial work on protocol reverse engineering. We conduct the evaluation on a large scale for the first time. In our experiment, we succeed to construct life profiles for more than 280,000 users from two popular LMAs. The results of empirical analysis not only validate the existence of the privacy issue in LMAs, but also demonstrate its severity.
Chang Xu 0003, Yi Yang 0040, Lingyun Ying, Purui Su, Dengguo Feng
TrustCom4
2014 Long Term Tracking and Characterization of P2P Botnet
abstract
P2P Botnet is quite robust against various attacks once very effective against centralized network. In this paper, we concentrate on the tracking of P2P botnets, investigate botnet victims which are routable on the Internet, also known as super peers. The super peers are the backbone of the botnet to disseminate its commands and payload updates. Through tracking of three typical live P2P botnets over 6 months and analysis of their network dynamics, we outline a number of descriptive and statistical characterization of super peers, such as geo-location, peer session time and intersession time, in-degree and out-degree distribution, pattern of arrival and departure. In addition, based on the assumption that IP dynamic allocation will not cross the AS (Autonomous System) border, we give out a lower bound estimate of total infected super peers in a conservative manner. We also propose several guidelines on disrupting P2P botnets concerning its various features we have characterized which could be helpful to the security community.
Jia Yan 0004, Lingyun Ying, Yi Yang 0040, Purui Su, Dengguo Feng
TrustCom2
2013 Bind your phone number with caution: automated user profiling through address book matching on smartphone
abstract
Due to the cost-efficient communicating manner and attractive user experience, messenger applications have dominated every smartphone in recent years. Nowadays, Address Book Matching, a new feature that helps people keep in touch with real world contacts, has been loaded in many popular messenger applications, which unfortunately as well brings severe privacy issues to users. In this paper, we propose a novel method to abuse such feature to automatically collect user profiles. This method can be applied to any application equipped with Address Book Matching independent of mobile platforms. We also build a prototype on Android to verify the effectiveness of our method. Moreover, we integrate profiles gathered from different messenger applications and provide insights by performing a consistency and authenticity analysis on user profile fields. As our experiments show, the abuse of Address Book Matching can cause severe user privacy leakage. Finally, we provide some countermeasures for developers to avoid this issue when designing messenger applications.
Lingyun Ying, Sibei Jiao, Purui Su, Dengguo Feng
AsiaCCS2
2013 OSNGuard: Detecting Worms with User Interaction Traces in Online Social Networks
Liang He 0011, Dengguo Feng, Purui Su, Lingyun Ying, Yi Yang 0040, Huafeng Huang, Huipeng Fang
ICICS4
2013 Automatic Polymorphic Exploit Generation for Software Vulnerabilities
Purui Su, Qi Li 0002, Lingyun Ying, Yi Yang 0040, Dengguo Feng
SecureComm4
2010 DepSim: A Dependency-Based Malware Similarity Comparison System
Yi Yang 0040, Lingyun Ying, Rui Wang 0032, Purui Su, Dengguo Feng
Inscrypt2