Shenao Wang 0001

dblp:360/7470 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0003-3818-3343ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 2 first-author · 13 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Seeing is (Not) Believing: The Mirage Card Attack Targeting Online Social Networks
abstract
In the digital era, Online Social Networks (OSNs) have become central to information dissemination, with sharing cards for link previews serving as a key feature.While these cards provide concise snapshots of shared content, their security implications have remained largely overlooked.This paper introduces the Mirage Card Attack, a novel class of attacks that exploits vulnerabilities in sharing card mechanisms across major OSNs.We identify two primary attack vectors: Proxy-Based Redirection and User-Agent-Based Cloaking.These attacks leverage design flaws in Share-SDK implementations and HTML meta tag usage, allowing attackers to bypass existing security measures and present deceptive content to users.Our systematic analysis reveals critical vulnerabilities in current sharing card systems.We demonstrate the feasibility of these attacks through comprehensive evaluations across 8 major OSNs for User-Agent-Based Cloaking and 6 OSNs for Proxy-Based Redirection.Additionally, we analyze 8 widely used card generation tools, uncovering significant security gaps.Our experiments show that some forged cards persist for over 15 days, highlighting the inadequacy of existing detection methods.To evaluate the practical impact of Mirage Card Attacks, we conduct a user study to * Both authors contributed equally to this research.
Wangchenlu Huang, Shenao Wang 0001, Yanjie Zhao 0001, Yuhao Gao, Guosheng Xu 0001, Haoyu Wang 0001
Internetware2
2025 Exploring Typo Squatting Threats in the Hugging Face Ecosystem
abstract
With the rapid advancement of artificial intelligence, pre-trained models (PTMs) have become fundamental building blocks in modern software systems.Model hubs, serving as centralized repositories for these components, have emerged as critical infrastructure in the AI software ecosystem.While prior research has extensively studied typosquatting attacks in traditional package management systems like NPM and PyPI, the security implications of such naming-based vulnerabilities in AI model hubs remain largely unexplored.To fill this gap, we present the first large-scale empirical study on typosquatting threats within the Hugging Face ecosystem, one of the most widely-used open-source model communities.Our research examines three key components: models, datasets, and organizations.Through a systematic analysis of 1,020,755 models (compared against the top 100 most downloaded ones), 219,812 datasets (compared against the top 100 most trending ones), and 127,011 organizations, we discovered concerning patterns of typosquatting that could compromise software supply chain security.Specifically, we identified 1,574 potentially malicious squatting models, with 10.4% exhibiting suspicious and harmful characteristics.Our investigation of datasets revealed 625 cases of typosquatting, where 42.2% showed signs of intentional impersonation based on sampling.Additionally, among the organizations studied, 302 demonstrated squatting patterns that could lead to supply chain attacks, with 4.8% showing explicit malicious intent.These findings highlight the pressing need for better naming conventions and security governance mechanisms in AI model repositories to ensure reliable and secure software development practices.We have reported all identified suspicious resources to Hugging Face for further investigation and potential mitigation measures.
Ningyuan Li 0005, Yanjie Zhao 0001, Shenao Wang 0001, Haoyu Wang 0001
Internetware3
2025 GPT Store Mining and Analysis
abstract
As an important extension of the ChatGPT ecosystem, GPT Store has developed into an active market hosting more than 3 million customized ChatGPTs (GPTs).Despite its large scale, the current academic community still has obvious limitations in its understanding of the ecosystem of this platform.Based on a complete dataset of more than 700,000 GPTs, this paper has achieved a multidimensional analysis of GPT Store.We first systematically examined the platform operation mechanism, covering core elements such as the classification system, interaction mode, and evaluation system.We also comprehensively analyzed the security risks, such as data leakage and jailbreak in GPT Store.Finally, through a user study, this work revealed the behavioral characteristics and experience pain points in real usage scenarios.Based on these findings, we provide operational platform optimization suggestions, including functional improvement, security enhancement, and interaction improvement.This study not only constructs an analytical framework for the GPT Store ecosystem but also provides empirical evidence and optimization directions for its future development.
Dongxun Su, Yanjie Zhao 0001, Xinyi Hou, Shenao Wang 0001, Haoyu Wang 0001
Internetware4
2025 A Characterization Study of Bugs in LLM Agent Workflow Orchestration Frameworks
abstract
Large Language Models (LLMs) have rapidly gained popularity, transforming research and industry. To support their adoption, LLM agent workflow orchestration frameworks (hereinafter referred to as LLM agent frameworks) like LangChain have become essential for building advanced applications. However, their complexity makes bugs inevitable, and these bugs can propagate to downstream applications, causing severe failures or unintended behaviors. In this paper, we first present an abstraction of the structure of mainstream LLM agent frameworks, identifying four key architectural components: data preprocessing, core schema, agent construction, and featured modules. Building on this abstraction, we conduct the first empirical study on LLM agent framework bugs, analyzing 1,026 bug instances extracted from 1,577 real-world bug-related GitHub pull requests (PRs) from three popular LLM agent frameworks: LangChain, LlamaIndex, and Haystack. For each bug, we examine its root cause, symptom, and structural component, providing a systematic taxonomy of nine root causes and six symptom categories. Finally, leveraging the framework structure abstraction and the large-scale empirical study, we perform detailed statistical analysis in terms of the distribution of bugs in different frameworks, the distribution across different framework components, and the relationship between root cause and symptom. The analysis reveals unique challenge patterns compared to traditional software, providing actionable guidance for practitioners on quality assurance.
Ziluo Xue, Yanjie Zhao 0001, Shenao Wang 0001, Kai Chen 0012, Haoyu Wang 0001
ASE3
2025 Demystifying Cookie Sharing Risks in WebView-based Mobile App-in-app Ecosystems
abstract
Mini-programs, an emerging mobile application paradigm within super-apps, offer a seamless and installation-free experience. However, the adoption of the web-view component has disrupted their isolation mechanisms, exposing new attack surfaces and vulnerabilities. In this paper, we introduce a novel vulnerability called Cross Mini-program Cookie Sharing (CMCS), which arises from the shared web-view environment across mini-programs. This vulnerability allows unauthorized data exchange across mini-programs by enabling one mini-program to access cookies set by another within the same web-view context, violating isolation principles. As a preliminary step, we analyzed the web-view mechanisms of four major platforms, including WeChat, AliPay, TikTok, and Baidu, and found that all of them are affected by CMCS vulnerabilities. These findings were responsibly disclosed and acknowledged with two CVEs. Furthermore, we demonstrate the collusion attack enabled by CMCS, where privileged mini-programs exfiltrate sensitive user data via cookies accessible to unprivileged mini-programs. To measure the impact of collusion attacks enabled by CMCS vulnerabilities in the wild, we developed MiCoScan, a static analysis tool that detects mini-programs affected by CMCS vulnerabilities. MiCoScan employs web-view context modeling to identify clusters of mini-programs sharing the same web-view domain and cross-webview data flow analysis to detect sensitive data transmissions to/from web-views. Using MiCoScan, we conducted a large-scale analysis of 351,483 mini-programs, identifying 45,448 clusters sharing web-view domains, 7,965 instances of privileged data transmission, and 9,877 mini-programs vulnerable to collusion attacks. Our findings highlight the widespread prevalence and significant security risks posed by CMCS vulnerabilities, underscoring the urgent need for improved isolation mechanisms in mini-program ecosystems.
Miao Zhang 0011, Shenao Wang 0001, Guilin Zheng, Yanjie Zhao 0001, Haoyu Wang 0001
ASE2
2025 MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid Analysis
abstract
The advent of MiniApps, operating within larger SuperApps, has revolutionized user experiences by offering a wide range of services without the need for individual app downloads. However, this convenience has raised significant privacy concerns, as these MiniApps often require access to sensitive data, potentially leading to privacy violations. Despite existing privacy regulations and platform guidelines, there is a lack of effective mechanisms to safeguard user privacy fully. To address this critical gap, we introduce MiniScope , a novel two-phase hybrid analysis approach, specifically designed for the MiniApp environment. This approach overcomes the limitations of existing static analysis techniques by incorporating UI transition states analysis, cross-package callback control flow resolution, and automated iterative UI exploration. This allows for a comprehensive understanding of MiniApps’ privacy practices, addressing the unique challenges of sub-package loading and event-driven callbacks. Our empirical evaluation of over 120K MiniApps using MiniScope demonstrates its effectiveness in identifying privacy inconsistencies. The results reveal significant issues, with 5.7% of MiniApps over-collecting private data and 33.4% overclaiming data collection. We have responsibly disclosed our findings to 2,282 developers, receiving 44 acknowledgments. These findings emphasize the urgent need for more precise privacy monitoring systems and highlight the responsibility of SuperApp operators to enforce stricter privacy measures.
Shenao Wang 0001, Yuekang Li, Kailong Wang 0001, Yi Liu 0069, Hui Li 0006, Yang Liu 0003, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.1
2025 Large Language Model Supply Chain: A Research Agenda
abstract
The rapid advancement of large language models (LLMs) has revolutionized artificial intelligence, introducing unprecedented capabilities in natural language processing and multimodal content generation. However, the increasing complexity and scale of these models have given rise to a multifaceted supply chain that presents unique challenges across infrastructure, foundation models, and downstream applications. This article provides the first comprehensive research agenda of the LLM supply chain, offering a structured approach to identify critical challenges and opportunities through the dual lenses of software engineering (SE) and security and privacy (S&P). We begin by establishing a clear definition of the LLM supply chain, encompassing its components and dependencies. We then analyze each layer of the supply chain, presenting a vision for robust and secure LLM development, reviewing the current state of practices and technologies, and identifying key challenges and research opportunities. This work aims to bridge the existing research gap in systematically understanding the multifaceted issues within the LLM supply chain, offering valuable insights to guide future efforts in this rapidly evolving domain.
Shenao Wang 0001, Yanjie Zhao 0001, Xinyi Hou, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.1
2025 LLM App Store Analysis: A Vision and Roadmap
abstract
The rapid growth and popularity of large language model (LLM) app stores have created new opportunities and challenges for researchers, developers, users, and app store managers. As the LLM app ecosystem continues to evolve, it is crucial to understand the current landscape and identify potential areas for future research and development. This article presents a forward-looking analysis of LLM app stores, focusing on key aspects such as data mining, security risk identification, development assistance, and market dynamics. Our comprehensive examination extends to the intricate relationships between various stakeholders and the technological advancements driving the ecosystem’s growth. We explore the ethical considerations and potential societal impacts of widespread LLM app adoption, highlighting the need for responsible innovation and governance frameworks. By examining these aspects, we aim to provide a vision for future research directions and highlight the importance of collaboration among stakeholders to address the challenges and opportunities within the LLM app ecosystem. The insights and recommendations provided in this article serve as a foundation for driving innovation, ensuring responsible development, and creating a thriving, user-centric LLM app landscape.
Yanjie Zhao 0001, Xinyi Hou, Shenao Wang 0001, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.3
2024 CanCal: Towards Real-time and Lightweight Ransomware Detection and Response in Industrial Environments
abstract
Ransomware attacks have emerged as one of the most significant cybersecurity threats. Despite numerous methods proposed for detecting and defending against ransomware, existing approaches face two fundamental limitations in large-scale industrial applications: (1) Behavior-based detection engines suffer from the enormous overhead of monitoring all processes and resource constraints for model inference, failing to meet the requirements for real-time detection; (2) Decoy-based detection engines generate an overwhelming number of false positives in large-scale industrial clusters, leading to intolerable disruptions to critical processes and excessive inspection efforts from security analysts. To address these challenges, we propose CanCal, a real-time and lightweight ransomware detection system. Specifically, instead of indiscriminately analyzing all processes, CanCal selectively filters suspicious processes by the monitoring layers and then performs in-depth behavioral analysis to isolate ransomware activities from benign operations, minimizing alert fatigue while ensuring lightweight computational and storage overhead. The experimental results on a large-scale industrial environment (1,761 ransomware, ~ 3 million events, continuous test over 5 months) indicate that CanCal achieves a remarkable 99.65% true positive rate on 555,678 unknown ransomware behavior events, with near-zero false positives. CanCal is as effective as state-of-the-art techniques while enabling rapid inference within 30ms and real-time response within a maximum of 3 seconds. CanCal dramatically reduces average CPU utilization by 91.04% (from 6.7% to 0.6%) and peak CPU utilization by 76.69% (from 26.6% to 6.2%), while avoiding 76.50% (from 3,192 to 750) of the inspection efforts from security analysts. By the time of this writing, CanCal has been integrated into a commercial product and successfully deployed on 3.32 million endpoints for over a year. From March 2023 to April 2024, CanCal successfully detected and thwarted 61 ransomware attacks. A detailed manual forensic analysis of 27 ransomware attacks from March to June 2023 (including 13 n-day exploits and 5 high-risk zero-day attacks) demonstrates the effectiveness of CanCal in combating sophisticated and unknown ransomware threats in real-world scenarios.
Shenao Wang 0001, Feng Dong 0008, Hangfeng Yang, Jingheng Xu, Haoyu Wang 0001
CCS1
2024 GPTZoo: A Large-scale Dataset of GPTs for the Research Community
abstract
The rapid advancements in Large Language Models (LLMs) have revolutionized natural language processing, with GPTs, customized versions of ChatGPT available on the GPT Store, emerging as a prominent technology for specific domains and tasks. To support academic research on GPTs, we introduce GPTZoo, a large-scale dataset comprising 730,420 GPT instances. Each instance includes rich metadata with 21 attributes describing its characteristics, as well as instructions, knowledge files, and third-party services utilized during its development. GPTZoo aims to provide researchers with a comprehensive and readily available resource to study the real-world applications, performance, and potential of GPTs. To facilitate efficient retrieval and analysis of GPTs, we also developed an automated command-line interface (CLI) that supports keyword-based searching of the dataset. To promote open research and innovation, the GPTZoo dataset will undergo continuous updates, and we are granting researchers public access to GPTZoo and its associated tools.
Xinyi Hou, Yanjie Zhao 0001, Shenao Wang 0001, Haoyu Wang 0001
ASE3
2024 Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model Hubs
abstract
The proliferation of pre-trained models (PTMs) and datasets has led to the emergence of centralized model hubs like Hugging Face, which facilitate collaborative development and reuse. However, recent security reports have uncovered vulnerabilities and instances of malicious attacks within these platforms, highlighting growing security concerns. This paper presents the first systematic study of malicious code poisoning attacks on pre-trained model hubs, focusing on the Hugging Face platform. We conduct a comprehensive threat analysis, develop a taxonomy of model formats, and perform root cause analysis of vulnerable formats. While existing tools like Fickling and ModelScan offer some protection, they face limitations in semantic-level analysis and comprehensive threat detection. To address these challenges, we propose MalHug, an end-to-end pipeline tailored for Hugging Face that combines dataset loading script extraction, model deserialization, in-depth taint analysis, and heuristic pattern matching to detect and classify malicious code poisoning attacks in datasets and models. In collaboration with Ant Group, a leading financial technology company, we have implemented and deployed MalHug on a mirrored Hugging Face instance within their infrastructure, where it has been operational for over three months. During this period, MalHug has monitored more than 705K models and 176K datasets, uncovering 91 malicious models and 9 malicious dataset loading scripts. These findings reveal a range of security threats, including reverse shell, browser credential theft, and system reconnaissance. This work not only bridges a critical gap in understanding the security of the PTM supply chain but also provides a practical, industry-tested solution for enhancing the security of pre-trained model hubs.
Shenao Wang 0001, Yanjie Zhao 0001, Xinyi Hou, Kailong Wang 0001, Peiming Gao, Haoyu Wang 0001
ASE2
2024 Towards Robust Detection of Open Source Software Supply Chain Poisoning Attacks in Industry Environments
abstract
The exponential growth of open-source package ecosystems, particularly NPM and PyPI, has led to an alarming increase in software supply chain poisoning attacks. Existing static analysis methods struggle with high false positive rates and are easily thwarted by obfuscation and dynamic code execution techniques. While dynamic analysis approaches offer improvements, they often suffer from capturing non-package behaviors and employing simplistic testing strategies that fail to trigger sophisticated malicious behaviors. To address these challenges, we present OSCAR, a robust dynamic code poisoning detection pipeline for NPM and PyPI ecosystems. OSCAR fully executes packages in a sandbox environment, employs fuzz testing on exported functions and classes, and implements aspect-based behavior monitoring with tailored API hook points. We evaluate OSCAR against six existing tools using a comprehensive benchmark dataset of real-world malicious and benign packages. OSCAR achieves an F1 score of 0.95 in NPM and 0.91 in PyPI, confirming that OSCAR is as effective as the current state-of-the-art technologies. Furthermore, for benign packages exhibiting characteristics typical of malicious packages, OSCAR reduces the false positive rate by an average of 32.06% in NPM (from 34.63% to 2.57%) and 39.87% in PyPI (from 41.10% to 1.23%), compared to other tools, significantly reducing the workload of manual reviews in real-world deployments. In cooperation with Ant Group, a leading financial technology company, we have deployed OSCAR on its NPM and PyPI mirrors since January 2023, identifying 10,404 malicious NPM packages and 1,235 malicious PyPI packages over 18 months. This work not only bridges the gap between academic research and industrial application in code poisoning detection but also provides a robust and practical solution that has been thoroughly tested in a real-world industrial setting.
Shenao Wang 0001, Yanjie Zhao 0001, Peiming Gao, Kailong Wang 0001, Haoyu Wang 0001
ASE3
2023 MalWuKong: Towards Fast, Accurate, and Multilingual Detection of Malicious Code Poisoning in OSS Supply Chains
abstract
In the face of increased threats within software registries and management systems, we address the critical need for effective malicious code detection. In this paper, we propose an innovative approach that integrates source code slicing, inter-procedural analysis, and cross-file inter-procedural analysis, thereby enhancing the detection precision and reducing false positives. This approach has been encapsulated within a multi-analysis-based framework for automatic detection of malicious code in real-world software packages. In its application to major third-party software registries like PyPI and NPM, our framework has proven effective, identifying 130 malicious packages from a total of 169,640 monitored over a continuous period of five weeks. This work advances the current state-of-the-art solution to malicious code detection, demonstrating significant practical impact in strengthening the software supply chain defense.
Ningke Li, Shenao Wang 0001, Mingxi Feng, Kailong Wang 0001, Meizhen Wang, Haoyu Wang 0001
ASE2
2023 Wemint:Tainting Sensitive Data Leaks in WeChat Mini-Programs
abstract
Mini-programs (MiniApps), lightweight versions of full-featured mobile apps that run inside a host app such as WeChat, have become increasingly popular due to their simplified and convenient user experiences. However, MiniApps raise new security and privacy concerns as they can access partially or all of host apps' system resources, including sensitive personal data. While taint detection has been proven effective in addressing this kind of concerns, existing taint detection techniques for mobile apps cannot be directly applied to MiniApps. The main reason is that the key logics of MiniApps are usually written in J avaScript, and its intrinsic characteristics (function-level scope, dynamic types, synchronous programming, and code obfuscation) prevent existing taint detection techniques from precisely propagating the taints. To address this problem, we propose a novel taint detection technique, Wemint, that detects sensitive information leaks in MiniApps. Specifically, Wemint facilitates taint propagation via building a context-based model based on the operational prin-ciple of MiniApps and J avaScript, and addresses asynchronous function calls by modeling their callbacks explicitly in taint rules. In addition, due to the adoption of Abstract Syntax Trees (ASTs) for code representation during taint detection, Wemint exhibits better robustness against the commonly-applied code obfuscation. Our experimental results show that Wemint can effectively detect sensitive information leaks in WeChat MiniApps, as well as trace the path of sensitive data flows. By applying Wemint to over 20K suspicious MiniApps, we found that over 7.5K (36.5 %) of them have sensitive data leaks, and Wemint outperforms the state-of-the-art DoubleX based techniques in detecting these leaks.
Shi Meng, Liu Wang 0002, Shenao Wang 0001, Kailong Wang 0001, Xusheng Xiao, Guangdong Bai, Haoyu Wang 0001
ASE3