Chenglai Xiong

dblp:357/2939 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-4822-2916ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 TrustSpace: Trusted memory space against control hijacking attacks
Chenglai Xiong, Guoqi Xie, Xinxin Jiang, Shufeng Chen, Sirong Zhao, Xuejun Yu 0001
J. Syst. Archit.1
2026 DCS3: A Dual-Layer Co-Aware Scheduler With Stealing Balance and Synchronized Priority in Virtualization Environments
abstract
Virtualization environments (e.g., containers and hypervisors) achieve isolation of multiple runtime entities but result in two mutually isolated guest and host layers. Such cross-ayer isolation could cause high latency and low throughput of the system. Previous aware scheduling and double scheduling fail to achieve bidirectional coordination between the guest and host layers. To address this challenge, we develop DCS3, a Dual-layer Co-aware Scheduler that combines stealing balance and synchronized priority. Stealing balancing migrates tasks between virtual CPU (vCPU) queues for load balance based on the workloads of physical CPUs (pCPUs). Synchronized priority dynamically adjusts the thread priorities running on the pCPUs according to the current vCPU workloads. The vCPUs and pC-PUs belong to the guest and host layers, respectively. Compared with aware scheduling, double scheduling, and DCS2 (i.e., DCS3 without synchronized priority), DCS3 has the following obvious advantages: 1) Requests Per Second (RPS) increases by up to 52%, 55%, and 2%, respectively; 2) request latency decreases by up to 72%, 71%, and 20%, respectively.
Chenglai Xiong, Guoqi Xie, Zhongjia Wang, Zhenli He, Shaowen Yao 0001, Jianfeng Tan, Tiwei Bie, Shoumeng Yan
IEEE Trans. Computers1
2026 Paddle Lite on Zephyr: Deploying AI Models in RTOS for Inference Acceleration
abstract
With the rapid development of deep learning techniques in mobile and embedded devices, light-weight inference engines (e.g., Paddle Lite and TensorFlow Lite) are emerged. In some real-time application scenarios, these light-weight inference engines require time acceleration and low memory consumption. Paddle Lite is a well-known open-source inference engine that is fully functional. However, Paddle Lite only supports regular OS (e.g., Linux, Windows, and iOS), making it difficult to achieve time acceleration and low memory consumption for real-time application scenarios during inference. In this brief, we propose the Paddle Lite on Zephyr solution for inference acceleration in RTOS. We first propose a modular compilation method to incorporate the most basic functions of Paddle Lite. To address the system differences between RTOS and Linux, we resolve the system-level and compilation-level issues from modular compilation. We then load the Paddle Lite model into memory as a device when the system starts up. We further design an inference method that skips third-party libraries during inference and thus obtains the same inference results as Linux. We deploy the Paddle Lite on Zephyr and conduct experiments with seven classic Convolutional Neural Network (CNN) models on a single-core CPU. The experiment results show that the average inference time on Zephyr RTOS is reduced by 7%, and the average memory consumption is reduced by 78% compared to Linux. This work has merged an upstream branch of the Paddle Lite.
Guoqi Xie, Wenyan Yan, Chenglai Xiong, Zhenli He, Shaowen Yao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Hypercall-Oriented Abnormal VM Status Detection System: A Non-Intrusive Solution for Both Hypervisor and Guests
abstract
Hypervisor is a VMM (Virtual Machine Monitor) that creates and runs multiple VMs (Virtual Machines) through abstracting resources from a physical machine. Hypercall is a special and crucial call used in virtualized systems as it serves as a main communication channel between VMs and the hypervisor. However, hypercall attacks occur when an attacker manipulates the communication channel, and it could cause abnormal VM status, potentially leading to the abnormal resource allocation of the host OS (Operating System) and crash of VMs. Therefore, the virtualized system should execute abnormal VM status detection to identify potential abnormal behaviors to protect VMs and the host OS; however, existing works are either for reconstructing the hypervisor or hardware isolation, not for the VM status detection for abnormal hypercall.This study develops a hypercall-oriented abnormal VM status detection system called HypercallDetector based on the following three innovations: 1) we implement a hypercall tracing based on eBPF to obtain the hypercall-related running status (including CPU usage, memory usage, network traffic, etc.) of each VM; 2) we implement a window division technology to divide the VM status into multiple status windows of the same size, and appropriate window size with balanced detection precision (95.0%) and latency (within 8.8 ms) obtained by proposing the window regulator; and 3) we implement a CS-H algorithm (Compressing Sensing for Hypercall) to distinguish whether the VM status is abnormal. HypercallDetector shows higher precision and lower latency than its opponent and consumes only 8.6% CPU of single core and 0.3% memory usage when starting 240 VMs.
Fangqi Bi, Guoqi Xie, Zhenli He, Shaowen Yao 0001, Sirong Zhao, Chenglai Xiong, Bo Wan 0008, Yiwen Jiang
IEEE Trans. Computers8
2025 AVL Function Table for LeafHooks Insertion With Obfuscated Control Flow Integrity
abstract
Control flow is the execution order of individual statements, instructions, or function calls within an imperative program. Malicious operation of control flow (e.g., tampering with normal function addresses) leads to severe consequences such as data leakage and system crash. Control Flow Integrity (CFI) is a defense restricting the execution order of program within Control Flow Graph (CFG). IndexHooks is an existing CFI solution designed against forward function calls tampering (including direct and indirect jump). This solution constructs a read-only linear function table that stores function addresses during compilation. Then, IndexHooks checks the table to make program jump to the correct target address during runtime. However, IndexHooks faces limitations in backtracking CFG construction, which can lead to excessive memory usage; the linear structure of the function table is vulnerable to brute force tampering. Addressing the limitations of IndexHooks, this study develops an obfuscated CFI solution called LeafHooks. LeafHooks is implemented during compilation by the LLVM compiler, which performs static analysis and instrumentation on the LLVM Intermediate Representation (IR) code of a program. We make the following three innovations: 1) we propose a speculation-free identification method for indirect function calls by linear traversing and analyzing codes to obtain legal function information (function address); 2) we save this information into a function table in the form of a Balanced Binary Tree (also known as AVL), enhancing the fuzzification of function addresses to defend against brute force; 3) we design a method to simulate control tamper attacks on ARM64 architecture to verify the ability of LeafHooks to protection. LeafHooks shows less overhead than state-of-the-art solutions and reduces 2.9% and 0.55% overhead on average using UnixBench and Phoronix, respectively.
Sirong Zhao, Guoqi Xie, Chenglai Xiong, Kenli Li 0001, Xuejun Yu 0001, Bo Wan 0008, Yiwen Jiang
IEEE Trans. Computers3
2025 Zram Instance Pool Framework for Adaptive Memory Compression in Resource-Sensitive Embedded Operating Systems
abstract
Memory compression can reduce the size of the inactive data in the random access memory (RAM), thereby freeing up unused space and allowing more programs to run; however, current mainstream memory compression frameworks (e.g., Zram and Zswap) and algorithms (e.g., Zstd and Lz4) do not effectively solve the problem of increased CPU utilization, causing they cannot be directly applied to the resource-sensitive embedded operating system, that is, sensitive to both CPU utilization and memory usage. In this study, we develop a Zram instance pool framework called ZramPool for adaptive memory compression. The framework consists of the swap space with multiple Zram instances and the adaptive Zram compression module. Through introducing linear regression analysis, the number of Zram instances can be adaptively adjusted based on the size of the compressed data, allowing Zram instances to work in parallel to match the workload. In ZramPool, we achieve two different requirements of reducing CPU utilization while keeping compression speed and increasing compression speed while keeping CPU utilization. ZramPool is deployed in the embedded Linux OS with a 8GB memory size running on the ARMv8 architecture. For the first requirement, ZramPool can reduce CPU utilization by an average of 11.42% while the compression speed only decreases by an average of 2.4%. For the second requirement, ZramPool can increase compression speed by an average of 11.71% while the CPU utilization only increases by an average of 1.9%.
Yin Deng, Guoqi Xie, Chenglai Xiong, Sirong Zhao, Wei Ren 0002, Kenli Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Holistic WCRT Analysis for Global Fixed-Priority Preemptive Multiprocessor Scheduling
abstract
Many embedded applications demand both resource efficiency and timing guarantee. However, resource sharing naturally complicates the analysis that extracts the worst-case scenario out of contention. Global Fixed-Priority (GFP) preemptive multiprocessor scheduling is one of the mainstream strategies to resolve contention on computational resources. It allows jobs of the same task to be executed on different processors, hence potentially enabling better parallelism and more efficient resource utilization. Unfortunately, its worst-case response time (WCRT) analysis is challenging. Existing approaches divide a high-priority task into three workloads, namely, carry-in workload, body workload, and carry-out workload, trying to optimize them individually. In this work, we propose a holistic WCRT analysis for GFP preemptive multiprocessor scheduling, where a task is no longer divided. Specifically, (i) we establish the tight interference scenario for the task being analyzed to find the most interfering high-priority jobs in any time interval; (ii) we obtain the starting released instant of each high-priority task’s first job to determine the maximum interference from high-priority tasks’ first jobs to the task being analyzed; (iii) we build the worst-case tight interference scenario for the task being analyzed by combining the tight interference scenario and the starting released instants; (iv) we prove that the WCRT of the task being analyzed can be decided by the worst-case tight interference scenario. Evaluation on schedulability shows that our proposed analysis achieves 4.2%-8.6% higher acceptance ratio in randomly generated data sets than the state-of-the-art workload division approaches.
Guoqi Xie, Chenglai Xiong, Renfa Li, Wanli Chang 0001
DAC2