Xiaohai Shi

dblp:246/5359 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 8 since 2021
YearPublicationVenuePosition
2025 RPG: Linux Kernel Fuzzing Guided by Distribution-Specific Runtime Parameter Interfaces
abstract
The Linux distribution kernel differs significantly from the mainline kernel, incorporating additional features and vendor-specific extensions. Among these additions, many runtime parameter interfaces are unique to distribution kernels, which expands the attack surface and increases the risk of potential vulnerabilities. Fuzzing has been used to assess Linux distributions, but existing tools cannot systematically test these distribution-specific interfaces due to two main challenges: (1) generating test cases for these runtime parameter interfaces, and (2) concentrating test resources on the distribution-specific interface code. To address these challenges, we propose RPG, a distribution-specific runtime parameter-guided kernel fuzzer. RPG operates in three phases: First, RPG extracts distribution-specific runtime parameter interfaces. Then, RPG uses LLM and tuning software databases to model each parameter range to generate meaningful interface test cases. Third, RPG utilizes the distribution kernel’s function control flow graph to guide the fuzzer to generate generic test cases that are more closely related to the distribution-specific interface code. We evaluated RPG on four Linux distribution kernels: Ubuntu 22.04, Fedora 42, OpenAnolis 8.8, and OpenAnolis 23.1. RPG detected 22 previously unknown bugs (13 distribution-specific), of which 15 were confirmed and 10 fixed by kernel maintainers. RPG also achieved 20.4% and 21.2% higher branch coverage than Syzkaller and Healer, respectively.
Yuheng Shen, Guoyu Yin, Runzhe Wang, Tao Ma 0006, Xiaohai Shi, Heyuan Shi
ASE7
2025 LLM-assisted Industrial-Scale Differential Testing of Package Incompatibilities in Linux Distributions
abstract
An open source Linux distribution often undergoes version upgrades and migrations, which is prone to incompatibility issues especially when it comes to large-scale software changes. Although differential testing has been widely used in software testing, it is still challenging to apply it for detecting such incompatibilities in the context of industrial settings. In this paper, we report our experience in leveraging LLMs to address the challenges faced by the Linux distribution community. Specifically, we develop an LLM-based differential testing method called Versify to assist maintainers of Linux distributions in locating incompatibilities during version upgrades and migrations. Its trial operation period within the Linux distribution community shows that it uncovered 8,489 instances of differing behavior, of which 644 were prioritized for attention by developers. After deduplication and filtering, 39 unique compatibility reports were identified. Feedback from Linux distributions developers indicates that our reports have provided valuable recommendations for package selection in future OS releases.
Chijin Zhou, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Tao Ma 0006, Zhe Wang 0015, Heyuan Shi
ASE6
2025 DragonRadar: Fuzzing Linux Kernel Deployed in Cloud-Native Environment
abstract
Kata Containers is a secure container runtime with lightweight virtual machines and a customized Linux kernel optimized for cloud-native workloads, which is important for cloud-native systems. Fuzzing is a widely-used technique for detecting kernel vulnerability. However, current kernel fuzzers can't be simply applied to kernels in cloud-native environments because of the discrepancies between test and actual deployment scenarios. This paper introduces DragonRadar, a kernel fuzzing tool adapted for Kata Containers, which aligns the testing environment with cloud deployment realities. We extend to support kernel fuzzing in cloud-native environments by integrating Syzkaller's capability with a lightweight virtual machine manager called Dragonball. The evaluation shows that DragonRadar effectively identifies 25 kernel vulnerabilities in the mainline Linux kernel used in the Kata Containers environment, while maintaining code coverage similar to vanilla Syzkaller. DragonRadar is available at https://github.com/TOBESTONG//DragonRadar.
Heyuan Shi, Weibo Zhang, Runzhe Wang, Xiaohai Shi, Guoyu Yin, Jianzhong Liu, Yuheng Shen
SANER4
2024 Industry Practice of Directed Kernel Fuzzing for Open-source Linux Distribution
abstract
Directed grey-box fuzzing is a widely used automatic testing technique that has helped developers test specific code space in the target program. Although many directed fuzzers are designed to test the Linux kernel, challenges still remain due to the complexity of industrial requirements and deployment environments. In this paper, we collaborate with developers from Alibaba and the OpenAnolis community to conduct an industry practice of directed kernel fuzzing for open-source Linux distribution. We highlight typical challenges in deploying directed kernel fuzzing, including target-related kernel configuration options being disabled, unrelated initial seeds limiting fuzzing startup performance, no support for kernel feature interface fuzzing, independent fuzzer execution limiting fuzzing effectiveness, much manual work to triage and analyze crashes, and hard to integrate into the existing fuzzing framework. We provide solutions to these challenges, which allowed us to discover 11 previously unknown kernel bugs related to cloud-native features, io_uring, and other components in the OpenAnolis Linux distribution.
Heyuan Shi, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Yu Jiang 0001
ASE8
2024 PatchBert: Continuous Stable Patch Identification for Linux Kernel via Pre-trained Model Fine-tuning
abstract
Stable patch identification is crucial in merging patches into stable versions, which helps ensure the stability of the Linux kernel. Although many tools have been proposed to mitigate the manual effort of stable patch identification, challenges still arise because they neglect continuous stable patch tracking and advanced Natural Language Processing (NLP) pre-training techniques. In this paper, in collaboration with developers from the openAnolis Linux operating system distribution community, we present a stable patch identification model called PatchBERT. It utilizes BERT and CodeBERT to capture the semantic patch representation from the commit message and code changes in a patch. We then perform patch classification and output the probability that the patch should be merged into the stable versions. We perform experiments on the dataset used by the previous methods. The experimental results show the superior performance of PatchBERT over state-of-the-art baselines. Additionally, it is common practice to train the model using the latest Linux patches and implement it in a real-world industrial setting. In this exercise, we randomly select 10,000 patches for identification, accurately identifying 8,617 patches and incorrectly identifying 1,383 patches. This practical outcome further confirms the effectiveness and utility of PatchBERT in real-world scenarios.
Heyuan Shi, Runzhe Wang, Yuheng Shen, Yuao Chen, Xiaohai Shi, Yu Jiang 0001
SANER8
2023 KeenTune: Automated Tuning Tool for Cloud Application Performance Testing and Optimization
abstract
The performance testing and optimization of cloud applications is challenging, because manual tuning of cloud computing stacks is tedious and automated tuning tools are rare used for cloud services. To address this issue, we introduce KeenTune, an automated tuning tool designed to optimize application performance and facilitate performance testing. KeenTune is a lightweight and flexible tool that can be deployed with to-be-tuned applications with negligible impact on their performance. Specifically, KeenTune uses a surrogate model that can be implemented with machine learning models to filter out less relevant parameters for efficient tuning. Our empirical evaluation shows that KeenTune significantly enhances the throughput performance of Nginx web servers, resulting in performance improvements of up to 90.43% and 117.23% in certain cases. This study highlights the benefits of using KeenTune for achieving efficient and effective performance testing of cloud applications. The video and source code for KeenTune are provided as supplementary materials.
Qinglong Wang 0003, Runzhe Wang, Xiaohai Shi, Zheng Liu 0022, Tao Ma 0006, Houbing Song, Heyuan Shi
ISSTA4
2023 Adaptive Tracing and Fault Injection based Fault Diagnosis for Open Source Server Software
abstract
The high overhead of tracing, the amount of up-front effort required to select trace points, and the lack of effective data analysis model are the significant barriers to the adoption of intra-component tracing for fault diagnosis today. This paper introduces a novel method for fault diagnosis by combining function level adaptive tracing, fault injection, and graph convolutional network. In order to implement this method, we introduce techniques for (i) selecting function level trace points, (ii) constructing approximate function call trees for programs when using adaptive tracing, and (iii) constructing graph convolutional network with fault injection campaign. We evaluate our method on four widely used open source server software: Redis, Nginx, Httpd, and SQlite. The experimental results show that our method outperforms log-based method, full tracing method, and Gaussian influence method in terms of accuracy, efficiency, and performance impact on the diagnosis target.
Wei Zhang 0248, Bolong Tan, Xiaohai Shi, Jianhui Jiang
QRS4
2022 Industry practice of configuration auto-tuning for cloud applications and services
abstract
Auto-tuning attracts increasing attention in industry practice to optimize the performance of a system with many configurable parameters. It is particularly useful for cloud applications and services since they have complex system hierarchies and intricate knob correlations. However, existing tools and algorithms rarely consider practical problems such as workload pressure control, the support for distributed deployment, and expensive time costs, etc., which are utterly important for enterprise cloud applications and services. In this work, we significantly extend an open source tuning tool – KeenTune to optimize several typical enterprise cloud applications and services. Our practice is in collaboration with enterprise users and tuning tool developers to address the aforementioned problems. Specifically, we highlight five key challenges from our experiences and provide a set of solutions accordingly. Through applying the improved tuning tool to different application scenarios, we achieve 2%-14% improvements for the performance of MySQL, OceanBase, nginx, ingress-nginx, and 5%-70% improvements for the performance of ACK cloud container service.
Runzhe Wang, Qinglong Wang 0003, Heyuan Shi, Yuheng Shen, Zheng Liu 0022, Xiaohai Shi, Yu Jiang 0001
ESEC/SIGSOFT FSE9
2019 Industry practice of coverage-guided enterprise Linux kernel fuzzing
abstract
Coverage-guided kernel fuzzing is a widely-used technique that has helped kernel developers and testers discover numerous vulnerabilities. However, due to the high complexity of application and hardware environment, there is little study on deploying fuzzing to the enterprise-level Linux kernel. In this paper, collaborating with the enterprise developers, we present the industry practice to deploy kernel fuzzing on four different enterprise Linux distributions that are responsible for internal business and external services of the company. We have addressed the following outstanding challenges when deploying a popular kernel fuzzer, syzkaller, to these enterprise Linux distributions: coverage support absence, kernel configuration inconsistency, bugs in shallow paths, and continuous fuzzing complexity. This leads to a vulnerability detection of 41 reproducible bugs which are previous unknown in these enterprise Linux kernel and 6 bugs with CVE IDs in U.S. National Vulnerability Database, including flaws that cause general protection fault, deadlock, and use-after-free.
Heyuan Shi, Runzhe Wang, Xiaohai Shi, Xun Jiao 0002, Houbing Song, Yu Jiang 0001, Jia-Guang Sun 0001
ESEC/SIGSOFT FSE5