Shaofei Li

dblp:251/7559 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0001-6530-5935ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 KnowHow: Automatically Applying High-Level CTI Knowledge for Interpretable and Accurate Provenance Analysis
Yuhan Meng, Shaofei Li, Jiaping Gui, Peng Jiang 0007, Ding Li 0001
NDSS2
2026 Promoguardian: Detecting Promotion Abuse Fraud with Multi-Relation Fused Graph Neural Networks
abstract
As e-commerce platforms develop, fraudulent activities are increasingly emerging, posing significant threats to the security and stability of these platforms. Promotion abuse is one of the fastest-growing types of fraud in recent years and is characterized by users exploiting promotional activities to gain financial benefits from the platform. To investigate this issue, we conduct the first study on promotion abuse fraud in e-commerce platforms MEITUAN. We find that promotion abuse fraud is a group-based fraudulent activity with two types of fraudulent activities: Stocking Up and Cashback Abuse. Unlike traditional fraudulent activities such as fake reviews, promotion abuse fraud typically involves ordinary customers conducting legitimate transactions and these two types of fraudulent activities are often intertwined. To address this issue, we propose leveraging additional information from the spatial and temporal perspectives to detect promotion abuse fraud. In this paper, we introduce PROMOGUARDIAN, a novel multi-relation fused graph neural network that integrates the spatial and temporal information of transaction data into a homogeneous graph to detect promotion abuse fraud. We conduct extensive experiments on real-world data from MEITUAN, and the results demonstrate that our proposed model outperforms state-of-the-art methods in promotion abuse fraud detection, achieving 93.15% precision, detecting 2.1 to 5.0 times more fraudsters, and preventing 1.5 to 8.8 times more financial losses in production environments.
Shaofei Li, Ziqi Zhang 0017, Minyao Hua, Shuli Gao, Zhenkai Liang, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
SP1
2025 Query Provenance Analysis: Efficient and Robust Defense Against Query-Based Black-Box Attacks
abstract
Query-based black-box attacks have emerged as a significant threat to machine learning systems, where adversaries can manipulate the input queries to generate adversarial examples that can cause misclassification of the system. To counter these attacks, researchers have proposed Stateful Defense Models (SDMs) such as BlackLight and PIHA, which can reject queries that are “similar” to historical queries. However, recent studies show that existing approaches are vulnerable to a stronger adaptive attack, Oracle-guided Adaptive Rejection Sampling (OARS). OARS can be easily integrated with existing attack algorithms to evade the SDMs by generating queries with fine-tuned direction and step size of perturbations utilizing the leaked decision boundary from the SDMs. In this paper, we propose a novel approach, Query Provenance Analysis (QPA), for defending against query-based black-box attacks robustly (against both non-adaptive and adaptive attacks) and efficiently (in real-time). Our key insight is that, instead of focusing on individual queries, utilizing features from the query sequence (termed query provenance) can distinguish malicious queries from benign queries more effectively. We construct a query provenance graph to capture the relationship between a new query and prior historical queries, and then design efficient algorithms to detect malicious queries based on the query provenance graphs. We evaluate QPA on four datasets against six query-based attacks and compare QPA with state-of-the-art SDM defenses. The results show that QPA outperforms the baselines regarding defense robustness and efficiency on both non-adaptive and adaptive attacks. Specifically, QPA reduces the Attack Success Rate (ASR) of OARS to 4.08%, which is roughly 20× lower than the baselines. Moreover, QPA achieves higher throughput (up to 7.67×) and lower latency (up to 11.09×) than baselines.
Shaofei Li, Haomin Jia, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
SP1
2025 Not All Exceptions Are Created Equal: Triaging Error Logs in Real-World Enterprises
abstract
Error logs like Java exceptions play a crucial role in diagnosing and resolving errors within the industry. Nonetheless, the extensive logging of Java exceptions may result in exception fatigue in large-scale Java systems at an industrial level, where the frequency of Java exceptions being generated surpasses developers’ ability to manage them effectively. Regrettably, there is a lack of research on the seriousness, prevalence, and solutions to this problem. To close this gap, we first make a comprehensive investigation into the exception fatigue problem within a prominent Internet corporation in China, namely Alibaba, confirming its importance in the industry. Consequently, we introduce a novel solution called ABEL , designed to automatically pinpoint the most relevant exceptions associated with software failures. The key challenge lies in the randomness of exceptions, which prevents existing sequence-based techniques from being effective. To address this challenge, ABEL establishes correlations between Java exceptions and the Key Performance Indicator (KPI) of applications, enabling the identification of exceptions leading to irregularities in KPI. Our evaluation of ABEL across four Java applications and five business KPIs within Alibaba illustrates its capability to pinpoint the primary cause of exception logs with an AC@5 (top-5 accuracy) exceeding 90%, effectively mitigating the exception fatigue problem within Alibaba. Furthermore, it can identify the root-cause exceptions in a real software failure within just 4 minutes, outperforming the manual investigation process by over an hour.
Mengyu Yao, Shaofei Li, Dingyu Yang, Zheshun Wu, Xiaojun Qu, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
ACM Trans. Softw. Eng. Methodol.3
2024 NODLINK: An Online System for Fine-Grained APT Attack Detection and Investigation
Shaofei Li, Feng Dong 0008, Xusheng Xiao, Haoyu Wang 0001, Fei Shao, Jiedong Chen, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
NDSS1
2024 Adonis: Practical and Efficient Control Flow Recovery through OS-level Traces
abstract
Control flow recovery is critical to promise the software quality, especially for large-scale software in production environment. However, the efficiency of most current control flow recovery techniques is compromised due to their runtime overheads along with deployment and development costs. To tackle this problem, we propose a novel solution, Adonis , which harnesses Operating System (OS) -level traces, such as dynamic library calls and system call traces, to efficiently and safely recover control flows in practice. Adonis operates in two steps: It first identifies the call-sites of trace entries, and then it executes a pairwise symbolic execution to recover valid execution paths. This technique has several advantages. First, Adonis does not require the insertion of any probes into existing applications, thereby minimizing runtime cost . Second, given that OS-level traces are hardware-independent, Adonis can be implemented across various hardware configurations without the need for hardware-specific engineering efforts, thus reducing deployment cost . Third, as Adonis is fully automated and does not depend on manually created logs, it circumvents additional development cost . We conducted an evaluation of Adonis on representative desktop applications and real-world IoT applications. Adonis can faithfully recover the control flow with 86.8% recall and 81.7% precision. Compared to the state-of-the-art log-based approach, Adonis can not only cover all the execution paths recovered but also recover 74.9% of statements that cannot be covered. In addition, the runtime cost of Adonis is 18.3× lower than the instrument-based approach; the analysis time and storage cost (indicative of the deployment cost) of Adonis is 50× smaller and 443× smaller than the hardware-based approach, respectively. To facilitate future replication and extension of this work, we have made the code and data publicly available.
Xuanzhe Liu, Chengxu Yang, Ding Li 0001, Shaofei Li, Zhenpeng Chen 0001
ACM Trans. Softw. Eng. Methodol.5
2023 Are we there yet? An Industrial Viewpoint on Provenance-based Endpoint Detection and Response Tools
abstract
Provenance-Based Endpoint Detection and Response (P-EDR) systems are deemed crucial for future Advanced Persistent Threats (APT) defenses. Despite the fact that numerous new techniques to improve P-EDR systems have been proposed in academia, it is still unclear whether the industry will adopt P-EDR systems and what improvements the industry desires for P-EDR systems. To this end, we conduct the first set of systematic studies on the effectiveness and the limitations of P-EDR systems. Our study consists of four components: a one-to-one interview, an online questionnaire study, a survey of the relevant literature, and a systematic measurement study. Our research indicates that all industry experts consider P-EDR systems to be more effective than conventional Endpoint Detection and Response (EDR) systems. However, industry experts are concerned about the operating cost of P-EDR systems. In addition, our research reveals three significant gaps between academia and industry (1) overlooking client-side overhead; (2) imbalancedalarm triage cost and interpretation cost; and (3) excessive server side memory consumption. This paper's findings provide objective data on the effectiveness of P-EDR systems and how much improvements are needed to adopt P-EDR systems in industry.
Feng Dong 0008, Shaofei Li, Peng Jiang 0007, Ding Li 0001, Haoyu Wang 0001, Liangyi Huang, Xusheng Xiao, Jiedong Chen, Xiapu Luo, Yao Guo 0001, Xiangqun Chen
CCS2