Wentai Li

dblp:221/3459 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-7463-9566ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Harmonizing Security and Performance in Microkernel File Servers
Wentai Li, Jinyu Gu 0001, Yubin Xia, Binyu Zang
J. Comput. Sci. Technol.1
2024 Understanding the IO Performance Gap Between OS-Level and VM-Level Containers in High-Density Deployment
abstract
Containers are widely deployed in clouds. There are two common container architectures: operating system-level (OS-level) container and virtual machine-level (VM-level) container. Typical examples are runc and Kata. It is well known that VM- level containers provide better isolation than OS-level containers, but at a higher overhead. Although there are quantitative analyses of the performance gap between these two container architectures, they rarely discuss the performance gap under the constrained resources nrovisioned to containers. Since the high-density deployment of containers is demanding in the cloud, each container is provisioned with limited resources specified by the cgroup mechanism. In this paper, we provide an in-depth analysis of the storage and network (two key aspects) performance differences between runc and Kata under varying resource constraints. We identify configuration implications that are crucial to performance and find that some of them are not exposed by the Kata interfaces. Based on that, we propose a profiling tool to automatically offer configuration suggestions for optimizing container performance. Our evaluation shows that the auto-generated configuration can improve the performance of MySQL by up to 107% in the TPCC benchmark compared with the default Kata setup.
Wentai Li, Kaijun Zhou 0001, Jiacheng Shi 0002, Xingman Chen, Luyuan Wang, Jinyu Gu 0001
ICDCS1
2023 Understanding and Mitigating Twin Function Misuses in Operating System Kernel
abstract
Major operating system kernels expose twin functions, which are groups of internal primitives that have mostly common but slightly diverging semantics, to kernel modules and subsystems. They are created to make the basic primitives work well in various scenarios. Unfortunately, though being expected as solutions, twin functions may turn to problem-makers in practice. As we have observed from over 500 patches applied to upstream Linux and FreeBSD, developers choose an improper one from the twins, leaving the kernel with stability and security bugs as well as error-prone code. In this paper, we aim to understand and mitigate the twin function misuse problem. First, we provide an informative discussion on the misuse-fix patches. We find that violating the constraints from calling context, missing the primitives with better performance, lacking the necessary security enhancements, and breaking the kernel coding style are the four major factors that lead to misuse. We then identify the programming rules from the patches and apply them with a static program analysis tool extended from Coccinelle, including callgraph tainting and type-based function pointer resolving. We have 136 patches accepted by the Linux community and fix 320 new misuses in the upstream Linux kernel.
Jinyu Gu 0001, Jiacheng Shi 0002, Haroran Su, Wentai Li, Binyu Zang, Haibing Guan, Haibo Chen 0001
IEEE Trans. Computers4
2022 EPK: Scalable and Efficient Memory Protection Keys
Jinyu Gu 0001, Wentai Li, Yubin Xia, Haibo Chen 0001
USENIX ATC3
2022 A Hardware-Software Co-design for Efficient Intra-Enclave Isolation
Jinyu Gu 0001, Bojun Zhu, Wentai Li, Yubin Xia, Haibo Chen 0001
USENIX Security Symposium4
2021 Efficiently Recovering Stateful System Components of Multi-server Microkernels
abstract
Microkernel OSes provide OS services through mutually-isolated system servers running in different user processes, which brings stronger fault isolation than monolithic OSes. Nevertheless, considering the fault recovery capability of system servers, most existing microkernel OSes usually do no more than restarting a fault server, which will cause a server to lose all its running states and then may affect all the applications relying on it. In this paper, we present a mechanism named TxIPC that can efficiently recover stateful system servers on microkernel OSes. Since a system server provides the service by inter-process communication (IPC), TxIPC makes it fault resilient by handling each IPC in a transaction-like manner. Specifically, if a fault happens in a server (during one IPC handling procedure), TxIPC aborts all the updates made by the IPC and thus recovers the server from that fault. Evaluations show that TxIPC can enable servers to recover from 99.8% (injected) faults with 3%-45 % performance overhead on application benchmarks, which significantly outperforms existing counterparts.
Wentai Li, Jinyu Gu 0001, Binyu Zang
ICDCS1
2020 Harmonizing Performance and Isolation in Microkernels with Efficient Intra-kernel Isolation and Communication
Jinyu Gu 0001, Xinyue Wu, Wentai Li, Zeyu Mi, Yubin Xia, Haibo Chen 0001
USENIX ATC3
2020 Lookup Table-Based Fast Reliability-Aware Sample Preparation Using Digital Microfluidic Biochips
abstract
Reliability of the prepared fluidic samples is a major concern for automated sample preparation using microfluidic biochips, where induced errors in the resultant concentration values severely affect the assay outcome. However, the existing design automation techniques have not thoroughly considered the reliability model to reduce the induced concentration errors during sample preparation. This article proposes a fast reliability-aware sample preparation (RASP) method for determining the optimized sequence of mixing steps (mixing process) with the enhanced reliability. In RASP, a probabilistic concentration prediction model is proposed for analyzing the reliability of a given mixing process. Based on this probabilistic model, a lookup table construction algorithm along with the table query method is proposed to obtain the optimized mixing process. The simulation results show that for any user-specified target concentration, RASP can effectively determine the optimized mixing process, which generates the droplets with target concentration within the error tolerance of 0.1%. Compared with the state-of-the-art sample preparation algorithm, RASP improves the reliability-related accuracy by 91.4% on average via 2048 testcases.
Lingxuan Shao, Wentai Li, Tsung-Yi Ho, Sudip Roy 0001, Hailong Yao 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 gMig: Efficient GPU Live Migration Optimized by Software Dirty Page for Full Virtualization
abstract
This paper introduces gMig, an open-source and practical GPU live migration solution for full virtualization. By taking advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy combined with the hashing based Software Dirty Page technique to achieve efficient GPU live migration. Particularly, we propose three approaches for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping to adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) One-Shot Pre-Copy, which greatly reduces the rounds of pre-copy of graphics memory. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by 80.0%.
Jiacheng Ma 0001, Yaozu Dong, Wentai Li, Zhengwei Qi, Bingsheng He, Haibing Guan
VEE4
2018 Scalable GPU Virtualization with Dynamic Sharing of Graphics Memory Space
abstract
With increasing GPU-intensive workloads deployed on cloud, cloud service providers are seeking for practical and efficient GPU virtualization solutions. However, the cutting-edge GPU virtualization techniques such as gVirt still suffer from the restriction of scalability, which constrains the number of guest virtual GPU instances. This paper presents gScale, a scalable and practical open source GPU virtualization solution based on gVirt. gScale presents a sharing mechanism which combines partition and sharing together to break the hardware limitation of global graphics memory space. Particularly, we propose two approaches for gScale: (1) the private shadow graphics translation table (GTT) , which enables global graphics memory space sharing among virtual GPUs, (2) ladder mapping and fence memory space pool, which allows CPU access host physical memory space (serving the graphics memory) to bypass global graphics memory space. Furthermore, to mitigate the performance degradation caused by switching private shadow GTT when the number of vGPUs scales up, four other mechanisms are proposed: (1) slot sharing, which improves the performance of vGPU by dividing the high global graphics memory into multiple slots, (2) fine-grained slotting, which provides a flexible virtual graphics memory configuration, (3) predictive GTT copy mechanism, which reduces the performance loss by switching private shadow GTT before context switch, (4) predictive-copy aware scheduling, which maximizes the improvement of predictive GTT copy mechanism in cloud environment. Evaluation shows that gScale scales up to 15 guest virtual GPU instances in Linux or 12 guest virtual GPU instances in Windows, which is 5x and 4x, respectively, that of gVirt. At the same time, gScale incurs a slight but acceptable runtime overhead when hosting multiple virtual GPU instances.
Mochi Xue, Jiacheng Ma 0001, Wentai Li, Yaozu Dong, Zhengwei Qi, Bingsheng He, Haibing Guan
IEEE Trans. Parallel Distributed Syst.3