Shaohua Li 0002

dblp:83/1926-2 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0001-7556-3615ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 4 first-author · 8 since 2021Computer networks · 4Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Sand: Decoupling Sanitization from Fuzzing for Low Overhead
abstract
Sanitizers provide robust test oracles for various vulnerabilities. Fuzzing on sanitizer-enabled programs has been the best practice to find software bugs. Since sanitizers require heavy program instrumentation to insert run-time checks, sanitizerenabled programs have much higher overhead compared to normally built programs. In this paper, we present SAND, a new fuzzing framework that decouples sanitization from the fuzzing loop. SAND performs fuzzing on a normally built program and only invokes the sanitizerenabled program when input is shown to be interesting. Since most of the generated inputs are not interesting, i.e., not bugtriggering, SAND allows most of the fuzzing time to be spent on the normally built program. We further introduce execution pattern to practically and effectively identify interesting inputs. We implement SAND on top of AFL++ and evaluate it on 20 real-world programs. Our extensive evaluation highlights its effectiveness: in 24 hours, compared to all the baseline fuzzers, SAND significantly discovers more bugs while not missing any.
Ziqiao Kong, Shaohua Li 0002, Heqing Huang 0002, Zhendong Su 0001
ICSE2
2025 Optimizing Input Minimization in Kernel Fuzzing
Hao Sun 0021, Ting Su 0001, Geguang Pu, Shaohua Li 0002
USENIX ATC6
2025 An Empirical Study of Bugs in the rustc Compiler
abstract
Rust is gaining popularity for its well-known memory safety guarantees and high performance, distinguishing it from C/C++ and JVM-based languages. Its compiler, rustc , enforces these guarantees through specialized mechanisms such as trait solving, borrow checking, and specific optimizations. However, Rust’s unique language mechanisms introduce complexity to its compiler, resulting in bugs that are uncommon in traditional compilers. With Rust’s increasing adoption in safety-critical domains, understanding these language mechanisms and their impact on compiler bugs is essential for improving the reliability of both rustc and Rust programs. Such understanding could provide the foundation for developing more effective testing strategies tailored to rustc . Improving the quality of rustc testing is essential for enhancing compiler reliability, which in turn strengthens the safety and correctness of all Rust programs, as compiler bugs can silently propagate into every compiled program. Yet, we still lack a large-scale, detailed, and in-depth study of rustc bugs. To bridge this gap, this work presents a comprehensive and systematic study of rustc bugs, specifically those originating in semantic analysis and intermediate representation (IR) processing, which are stages that implement essential Rust language features such as ownership and lifetimes. Our analysis examines issues and fixes reported between 2022 and 2024, with a manual review of 301 valid issues. We categorize these bugs based on their causes, symptoms, affected compilation stages, and test case characteristics. Additionally, we evaluate existing rustc testing tools to assess their effectiveness and limitations. Our key findings include: (1) rustc bugs primarily arise from Rust’s type system and lifetime model, with frequent errors in the High-Level Intermediate Representation (HIR) and Mid-Level Intermediate Representation (MIR) modules due to complex checkers and optimizations; (2) bug-revealing test cases often involve unstable features, advanced trait usages, lifetime annotations, standard APIs, and specific optimization levels; (3) while both valid and invalid programs can trigger bugs, existing testing tools struggle to detect non-crash errors, underscoring the need for further advancements in rustc testing.
Yang Feng 0003, Yunbo Ni, Shaohua Li 0002, Xizhe Yin, Qingkai Shi, Baowen Xu, Zhendong Su 0001
Proc. ACM Program. Lang.4
2025 Interleaving Large Language Models for Compiler Testing
abstract
Testing compilers with AI models, especially large language models (LLMs), has shown great promise. However, current approaches struggle with two key problems: The generated programs for testing compilers are often too simple, and extensive testing with the LLMs is computationally expensive. In this paper, we propose a novel compiler testing framework that decouples the testing process into two distinct phases: an offline phase and an online phase. In the offline phase, we use LLMs to generate a collection of small but feature-rich code pieces. In the online phase, we reuse these code pieces by strategically combining them to build high-quality and valid test programs, which are then used to test compilers. We implement this idea in a tool, LegoFuzz , for testing C compilers. The results are striking: we found 66 bugs in GCC and LLVM, the most widely used C compilers. Almost half of the bugs are miscompilation bugs, which are serious and hard-to-find bugs that none of the existing LLM-based tools could find. We believe this efficient design opens up new possibilities for using AI models in software testing beyond just C compilers.
Yunbo Ni, Shaohua Li 0002
Proc. ACM Program. Lang.2
2024 UBFuzz: Finding Bugs in Sanitizer Implementations
abstract
In this paper, we propose a testing framework for validating sanitizer implementations in compilers. Our core components are (1) a program generator specifically designed for producing programs containing undefined behavior (UB), and (2) a novel test oracle for sanitizer testing. The program generator employs Shadow Statement Insertion, a general and effective approach for introducing UB into a valid seed program. The generated UB programs are subsequently utilized for differential testing of multiple sanitizer implementations. Nevertheless, discrepant sanitizer reports may stem from either compiler optimization or sanitizer bugs. To accurately determine if a discrepancy is caused by sanitizer bugs, we introduce a new test oracle called crash-site mapping.
Shaohua Li 0002, Zhendong Su 0001
ASPLOS (1)1
2024 Boosting Compiler Testing by Injecting Real-World Code
abstract
We introduce a novel approach for testing optimizing compilers with code from real-world applications The main idea is to construct well-formed programs by fusing multiple code snippets from various realworld projects. The key insight is backed by the fact that the large volume of real-world code exercises rich syntactical and semantic language features, which current engineering-intensive approaches like random program generators are hard to fully support. To construct well-formed programs from real-world code our approach works by (1) extracting real-world code at the granularity of function, (2) injecting function calls into seed programs, and (3) leveraging dynamic execution information to maintain the semantics and build complex data dependencies between injected functions and the seed program. With this idea, our approach complements the existing generators by boosting their expressiveness via fusing real-world code in a semantics-preserving way. We implement our idea in a tool, Creal, to test C compilers. In a nine-month testing period, we have reported 132 bugs to GCC and LLVM, two of the most popular and well-tested C compilers. At the time of writing, 121 of them have been confirmed as unknown bugs, and 101 of them have been fixed. Most of these bugs were miscompilations, and many were recognized as long-latent and critical. Our evaluation results evidently demonstrate the significant advantage of using real-world code to stress-test compilers. We believe this idea will benefit the general compiler testing direction and will be directly applicable to other compilers.
Shaohua Li 0002, Theodoros Theodoridis, Zhendong Su 0001
Proc. ACM Program. Lang.1
2023 Finding Unstable Code via Compiler-Driven Differential Testing
abstract
Unstable code refers to code that has inconsistent or unstable run-time semantics due to undefined behavior (UB) in the program. Compilers exploit UB by assuming that UB never occurs, which allows them to generate efficient but potentially semantically inconsistent binaries. Practitioners have put great research and engineering effort into designing dynamic tools such as sanitizers for frequently occurring UBs. However, it remains a big challenge how to detect UBs that are beyond the reach of current techniques.
Shaohua Li 0002, Zhendong Su 0001
ASPLOS (3)1
2023 Accelerating Fuzzing through Prefix-Guided Execution
abstract
Coverage-guided fuzzing is one of the most effective approaches for discovering software defects and vulnerabilities. It executes all mutated tests from seed inputs to expose coverage-increasing tests. However, executing all mutated tests incurs significant performance penalties---most of the mutated tests are discarded because they do not increase code coverage. Thus, determining if a test increases code coverage without actually executing it is beneficial, but a paradoxical challenge. In this paper, we introduce the notion of prefix-guided execution (PGE) to tackle this challenge. PGE leverages two key observations: (1) Only a tiny fraction of the mutated tests increase coverage, thus requiring full execution; and (2) whether a test increases coverage may be accurately inferred from its partial execution. PGE monitors the execution of a test and applies early termination when the execution prefix indicates that the test is unlikely to increase coverage. To demonstrate the potential of PGE, we implement a prototype on top of AFL++, which we call AFL++-PGE. We evaluate AFL++-PGE on MAGMA, a ground-truth benchmark set that consists of 21 programs from nine popular real-world projects. Our results show that, after 48 hours of fuzzing, AFL++-PGE finds more bugs, discovers bugs faster, and achieves higher coverage. Prefix-guided execution is general and can benefit the AFL-based family of fuzzers.
Shaohua Li 0002, Zhendong Su 0001
Proc. ACM Program. Lang.1
2022 Detecting non-crashing functional bugs in Android apps via deep-state differential analysis
abstract
Non-crashing functional bugs of Android apps can seriously affect user experience. Often buried in rare program paths, such bugs are difficult to detect but lead to severe consequences. Unfortunately, very few automatic functional bug oracles for Android apps exist, and they are all specific to limited types of bugs. In this paper, we introduce a novel technique named deep-state differential analysis, which brings the classical "bugs as deviant behaviors" oracle to Android apps as a generic automatic test oracle. Our oracle utilizes the observations on the execution of automatically generated test inputs that (1) there can be a large number of traces reaching internal app states with similar GUI layouts, and only a small portion of them would reach an erroneous app state, and (2) when performing the same sequence of actions on similar GUI layouts, the outcomes will be limited. Therefore, for each set of test inputs terminating at similar GUI layouts, we manifest comparable app behaviors by appending the same events to these inputs, cluster the manifested behaviors, and identify minorities as possible anomalies. We also calibrate the distribution of these test inputs by a novel input calibration procedure, to ensure the distribution of these test inputs is balanced with rare bug occurrences.
Yanyan Jiang 0001, Ting Su 0001, Shaohua Li 0002, Chang Xu 0001, Jian Lu 0001, Zhendong Su 0001
ESEC/SIGSOFT FSE4
2021 Enabling Cross-Chain Transactions: A Decentralized Cryptocurrency Exchange Protocol
abstract
Inspired by Bitcoin, many different kinds of cryptocurrencies based on blockchain technology have turned up on the market. Due to the special structure of the blockchain, it has been deemed impossible to directly trade between traditional currencies and cryptocurrencies or between different types of cryptocurrencies. Generally, trading between different currencies is conducted through a centralized third-party platform. However, it has the problem of a single point of failure, which is vulnerable to attacks and thus affects the security of the transactions. In this paper, we propose a distributed cryptocurrency trading scheme to solve the problem of centralized exchanges, which can achieve secure trading between different types of cryptocurrencies. Our scheme is implemented with smart contracts on an Ethereum blockchain and deployed on an Ethereum test network. In addition to implementing transactions between individual users, our scheme also allows transactions among multiple users. The experimental result proves that the cost of our scheme is acceptable.
Hangyu Tian, Kaiping Xue, Shaohua Li 0002, Jie Xu 0031, Jianqing Liu, Jun Zhao 0007, David S. L. Wei
IEEE Trans. Inf. Forensics Secur.4
2020 FALCON: A Fourier Transform Based Approach for Fast and Secure Convolutional Neural Network Predictions
abstract
Deep learning as a service has been widely deployed to utilize deep neural network models to provide prediction services. However, this raises privacy concerns since clients need to send sensitive information to servers. In this paper, we focus on the scenario where clients want to classify private images with a convolutional neural network model hosted in the server, while both parties keep their data private. We present FALCON, a fast and secure approach for CNN predictions based on fast Fourier Transform. Our solution enables linear layers of a CNN model to be evaluated simply and efficiently with fully homomorphic encryption. We also introduce the first efficient and privacy-preserving protocol for softmax function, which is an indispensable component in CNNs and has not yet been evaluated in previous work due to its high complexity.
Shaohua Li 0002, Kaiping Xue, Bin Zhu 0010, Chenkai Ding, Xindi Gao, David S. L. Wei
CVPR1
2020 SecGrid: A Secure and Efficient SGX-Enabled Smart Grid System With Rich Functionalities
abstract
Smart grid adopts two-way communication and rich functionalities to gain a positive impact on the sustainability and efficiency of power usage, but on the other hand, also poses serious challenges to customers' privacy. Existing solutions in smart grid usually use cryptographic tools, such as homomorphic encryption, to protect individual privacy, which, however, can only support limited and simple functionalities. Moreover, the resource-constrained smart meters need to perform heavy asymmetric cryptography in these solutions, and thus unnecessarily increases load on smart grid. In this paper, we present a practical and secure SGX-enabled smart grid system, named SecGrid. Our system leverages trusted hardware SGX to ensure that grid utilities can efficiently execute rich functionalities on customers' private data, while guaranteeing their privacy. With our well-devised security protocols in SecGrid, only the smart meters need to perform AES encryption. To validate the superiority of our design, we conduct security analysis and experimentation. Security analysis shows that SecGrid can thwart various attacks from malicious adversaries, and the experimental results show that SecGrid is much faster than the existing privacy-preserving schemes in smart grid.
Shaohua Li 0002, Kaiping Xue, David S. L. Wei, Hao Yue 0001, Nenghai Yu, Peilin Hong
IEEE Trans. Inf. Forensics Secur.1
2019 Healthchain: A Blockchain-Based Privacy Preserving Scheme for Large-Scale Health Data
abstract
With the dramatically increasing deployment of the Internet of Things (IoT), remote monitoring of health data to achieve intelligent healthcare has received great attention recently. However, due to the limited computing power and storage capacity of IoT devices, users' health data are generally stored in a centralized third party, such as the hospital database or cloud, and make users lose control of their health data, which can easily result in privacy leakage and single-point bottleneck. In this paper, we propose Healthchain, a large-scale health data privacy preserving scheme based on blockchain technology, where health data are encrypted to conduct fine-grained access control. Specifically, users can effectively revoke or add authorized doctors by leveraging user transactions for key management. Furthermore, by introducing Healthchain, both IoT data and doctor diagnosis cannot be deleted or tampered with so as to avoid medical disputes. Security analysis and experimental results show that the proposed Healthchain is applicable for smart healthcare system.
Jie Xu 0031, Kaiping Xue, Shaohua Li 0002, Hangyu Tian, Jianan Hong, Peilin Hong, Nenghai Yu
IEEE Internet Things J.3
2019 A Secure and Efficient Access and Handover Authentication Protocol for Internet of Things in Space Information Networks
abstract
Space information network (SIN) makes it possible for any object to be connected to the Internet anywhere, even in the areas with extreme conditions, where a cellular network is not easy to deploy. Access authentication is the key to secure users' access control in SIN, mainly to prevent illegal adversaries from getting access to SIN services. However, the highly complicated communication environment of SIN (e.g., exposed links, higher signal delay, etc.) poses a challenging issue in the design of a secure and efficient authentication scheme. Although some authentication schemes have been proposed for SIN, they are unsuitable for Internet of Things (IoT) in SIN due to the high signaling overhead and insufficient security properties. Therefore, in this paper, we design a provably secure and efficient authentication protocol, along with an efficient handover mechanism, for IoT in SIN. In our design, we introduce a new authentication system model, where the satellites are given the ability to authenticate users to avoid the online involvement of the network control center (NCC) when authenticating users, thereby reducing long authentication delay and avoiding a single point of bottleneck in NCC. Furthermore, the support of batch verification in our design can significantly enhance handover efficiency when a group of users switch to another satellite. Our further analysis shows that our scheme is secure against various attacks and can meet a variety of security requirements. In addition, performance evaluation shows the superiority of our scheme on both delay and handover efficiency compared with existing schemes.
Kaiping Xue, Shaohua Li 0002, David S. L. Wei, Huancheng Zhou, Nenghai Yu
IEEE Internet Things J.3
2019 PPSO: A Privacy-Preserving Service Outsourcing Scheme for Real-Time Pricing Demand Response in Smart Grid
abstract
In power utility service outsourcing, some time-sensitive computations (e.g., dynamic prices prediction) are outsourced to a third-party service provider. This brings in new privacy threats to customers. Although some existing works focus on achieving privacy-preserving temporal and spatial aggregation for one center, they basically cannot be directly applied to the scenario of service outsourcing with multiple centers (e.g., with power utility and service providers). We thus propose a privacy-preserving service outsourcing scheme, called PPSO, for real-time pricing demand response in smart grid with fault tolerance and flexible customers' enrollment and revocation. In our proposed PPSO, power utility can outsource the dynamic pricing prediction to a service provider, while still preserving customers' privacy. Extensive experiment results demonstrate that PPSO has less computation overhead and lower transmission delay compared with existing schemes.
Kaiping Xue, Qingyou Yang, Shaohua Li 0002, David S. L. Wei, Min Peng 0001, Imran Memon, Peilin Hong
IEEE Internet Things J.3
2018 LASA: Lightweight, Auditable and Secure Access Control in ICN with Limitation of Access Times
abstract
Information Centric Networking (ICN), a future network architecture candidate, aims to alleviate the problem of insufficient bandwidth in traditional IP network. In ICN, contents are distributed in the whole network, so access control becomes more intractable. As we know, almost all of existing solutions consider it as a "Yes or No" problem, where a user either has the permission to access the corresponding content or not. However, in many practical situations, a content provider doesn't expect a single authorized user has the ability to access its repertory without times limitation when taking copyright protection into account. In this paper, we propose LASA, a lightweight, auditable and secure solution where legitimate users are limited to access a content provider's data within pre-designate times. In LASA, each content provider sets maximum access times for each legitimate user and edge routers perform authentication and audit based on users' signatures attached to interest packets. Once a legitimate user attempts to exceed his/her limited access times, his/her secret key will be leaked and the dishonest behavior will be detected. Our security analysis shows that LASA can provide signature unforgeability, data confidentiality and other security features. Experiment results show that our scheme LASA brings a little computational cost.
Peixuan He, Yinxin Wan, Qiudong Xia, Shaohua Li 0002, Jianan Hong, Kaiping Xue
ICC4
2018 PPMA: Privacy-Preserving Multisubset Data Aggregation in Smart Grid
abstract
Privacy-preserving data aggregation has been extensively studied in smart grid. However, almost all existing schemes aggregate the total electricity consumption data of the whole user set, which sometimes cannot meet the fine-grained demands from control center in smart grid. In this paper, we propose a privacy-preserving multisubset data aggregation scheme, named PPMA, in smart grid. PPMA can aggregate users' electricity consumption data of different ranges, while guaranteeing the privacy of individual users. Detailed security analysis shows that PPMA can protect individual user's electricity consumption privacy against a strong adversary. In addition, extensive experiments results demonstrate that PPMA has less computation overhead and no more extra communication and storage costs.
Shaohua Li 0002, Kaiping Xue, Qingyou Yang, Peilin Hong
IEEE Trans. Ind. Informatics1
2017 Two-Cloud Secure Database for Numeric-Related SQL Range Queries With Privacy Preserving
abstract
Industries and individuals outsource database to realize convenient and low-cost applications and services. In order to provide sufficient functionality for SQL queries, many secure database schemes have been proposed. However, such schemes are vulnerable to privacy leakage to cloud server. The main reason is that database is hosted and processed in cloud server, which is beyond the control of data owners. For the numerical range query (“>,” “<;,” and so on), those schemes cannot provide sufficient privacy protection against practical challenges, e.g., privacy leakage of statistical properties, access pattern. Furthermore, increased number of queries will inevitably leak more information to the cloud server. In this paper, we propose a two-cloud architecture for secure database, with a series of intersection protocols that provide privacy preservation to various numeric-related range queries. Security analysis shows that privacy of numerical information is strongly protected against cloud providers in our proposed scheme.
Kaiping Xue, Shaohua Li 0002, Jianan Hong, Yingjie Xue, Nenghai Yu, Peilin Hong
IEEE Trans. Inf. Forensics Secur.2