Xiaokang Hu

dblp:205/6979 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ZOC: Elastic and Cost-Efficient Virtual SmartNIC Architecture for Cloud Physical Machines
Naixuan Guan, Xiaokang Hu, Yisheng Xie, Xishi Qiu, Chaojie Liu, Yuchao Cao, Banghao Ying, Dianchen Tian, Yangzeyu Zhang, Hujun Ge, Yibin Shen, Jiesheng Wu
NSDI2
2026 Cacheman: A Comprehensive Last-Level Cache Management System for Multi-tenant Clouds
abstract
Competition for the last-level cache (LLC) is a long-standing issue in multi-tenant cloud environments, often leading to severe performance interference among co-located virtual machines. LLC management in the cloud faces unique challenges, including unpredictable tenant workloads, misaligned performance metrics, and the need to ensure fairness under service level agreements (SLAs). Existing LLC allocation methods fall short in addressing these challenges. We present Cacheman, a comprehensive LLC management system designed from real-world cloud deployment experience. Cacheman introduces a novel gradient-based sharing mechanism for LLC ways, enabling smooth LLC allocation adjustments that simultaneously improve fairness and utilization efficiency. Its real-time allocation algorithm promptly detects and mitigates unfair LLC allocation, adapting to dynamic workloads with second-scale responsiveness. Additionally, Cacheman supports performance consistency for tenants running distributed applications by enforcing negotiated upper bounds on cache usage. Extensive experiments demonstrate that Cacheman effectively achieves its multi-dimensional goals, and long-term production deployment further shows that it significantly reduces SLA violations caused by LLC contention.
Xiaokang Hu, Yuchao Cao, Naixuan Guan, Yifan Wu 0037, Xishi Qiu, Shengdong Dai, Ben Luo, Sanchuan Cheng, Fudong Qiu, Yibin Shen, Jiesheng Wu
PPoPP1
2026 EIDS: A Cloud Intrusion Detection System with High Performance and Maintainability
abstract
Intrusion Detection Systems (IDSes) are widely employed to identify potential attacks in guest virtual machines (VMs). Nonetheless, traditional IDSes fall short of the demands of high-performance clouds. First, monitoring VM events increases the tail latency of guest services. Second, the throughput of traditional IDSes cannot meet high-performance cloud requirements, leading to event loss and reduced detection accuracy. Finally, cloud providers typically run complex IDS tools within the VM. Updating IDS functionality requires modifying guest VMs, which hurts maintainability. To overcome these challenges, this article presents EIDS, a cloud IDS framework with high performance and good maintainability. We observe that the main bottleneck is collecting VM status, and the collected status can be divided into fundamental and supplementary status. EIDS then splits the status collection procedure spatially and temporally. First, we provide a status monitor with a separate architecture that isolates the status collection logic in a microVM, thus minimizing the code in guest VMs and improving maintainability. Second, EIDS introduces a two-phase status collection method to handle multiple events in batches, asynchronously, for high IDS throughput. A tiny tracer, implemented with eBPF, operates inside the user VM to collect fundamental status. The complex status collector runs in an isolated microVM. It utilizes Virtual Machine Introspection (VMI) to gather supplementary status, using the fundamental status to bridge the semantic gap. The status collector batches the collection for multiple events to amortize the fixed overhead of microVM switching and improve event tracing throughput. Finally, to minimize tail latency overhead, a fine-grained and workload-aware scheduler executes IDS logic with small time slices during user VM idle periods. We implemented a prototype of EIDS in Linux-KVM and conducted a comprehensive evaluation. We compared EIDS’s performance with Falco, an open-source IDS widely used by Kubernetes and AWS for runtime security monitoring. The results demonstrate that, compared to Falco, EIDS reduces the 99 th -percentile latency overhead by 97% and achieves a 13.8X improvement in IDS event handling throughput.
Xiaokang Hu, Zhichao Hua 0001, Naixuan Guan, Yibin Shen, Yang Yu 0002, Zeyu Mi, Yubin Xia, Jiesheng Wu
ACM Trans. Comput. Syst.1
2025 Tai Chi: A General High-Efficiency Scheduling Framework for SmartNICs in Hyperscale Clouds
abstract
Cloud service providers increasingly adopt SmartNICs to offload data-plane services (e.g., DPDK and SPDK) and control-plane tasks (such as disk and NIC initialization). Our analysis of production environments reveals that data-plane services statically provision CPUs for peak load, resulting in 67.5% idle CPU cycles during 99% of their runtime in IaaS clouds, leading to wasted CPU resources. On the other hand, control-plane tasks fail to meet critical Service Level Objectives (SLOs), such as virtual machine startup time. Unfortunately, achieving control-plane SLO improvements through co-scheduling with idle data-plane services remains highly challenging, due to the combined effects of intrinsic scheduling latency and the substantial architectural complexity inherent to control-plane ecosystems.
Bang Di, Kaijie Guo, Yibin Shen, Sanchuan Cheng, Fudong Qiu, Xiaokang Hu, Naixuan Guan, Dongdong Huang, Jinhu Li, Yi Wang 0004, Yifang Yang, Yilong Lv, Zhenwei Lu, Jiesheng Wu
SOSP9
2023 Maximizing Throughput for Coexisting Wireless Body Area Networks (WBANs) Based on Optimal Clustering
abstract
The interference of the uplink of coexisting wireless body area networks is studied in this article. When multiple WBANs working in the same channel perform data transmission simultaneously, their uplinks will be interfered with by multiple transmission links of adjacent WBANs, which leads to data packet loss and data transmission failure. Therefore, mitigating interference and maximizing throughput are the main goals of this article. First, the graph coloring algorithm is used to analyze the interference relationship among WBANs and a mathematical model of throughput maximization is formulated. Then, the time slot reallocation algorithm is proposed to achieve node-level optimization but with high complexity. The simulation results show that the proposed scheme can achieve a good packet reception rate and throughput for coexisting WBANs.
Xiaokang Hu, Kunqi Guo, Yuting Qian
IEEE Internet Things J.1
2023 QKPT: Securing Your Private Keys in Cloud With Performance, Scalability and Transparency
abstract
Private key (e.g., RSA key) protection is a significant issue for cloud but existing keyless or keyguard solutions suffer from performance, elasticity or applicability limitations. Recently, represented by Intel KPT, a novel keyguard architecture emerges to combine trusted platform module and crypto accelerator for achieving both security and performance. However, the straight use of KPT for private key protection may not be a good fit in cloud as it incurs challenges on protection capacity, key provisioning latency and transparency. Based on KPT-like hardware, we propose QKPT, a comprehensive key management system to bring your own private keys (BYOPK) into multi-tenant clouds. QKPT introduces a carefully-designed key wrapping layer to overcome these challenges. A small symmetric wrapping key (SWK) is generated for each tenant as the master key to resolve the former two challenges, while a special private key wrapping scheme is adopted to resolve the transparency limitation. Additionally, QKPT incorporates certificate trust to enhance the security of the SWK lifecycle and provides a hardened key server solution without expensive HSM. The evaluation shows that QKPT has a low runtime overhead ($\leq$1.2% for SSL/TLS handshakes) and still greatly outperforms the software baseline (3.5x-17x) owing to the crypto offloading.
Zongpu Zhang, Hubin Zhang, Xiaokang Hu, Jian Li 0021, Weigang Li 0002, Guodong Zhu, Kapil Sood, Brian Will, Haibing Guan
IEEE Trans. Dependable Secur. Comput.4
2022 ES2: Building an Efficient and Responsive Event Path for I/O Virtualization
abstract
Hypervisor intervention in the virtual I/O event path is a main performance bottleneck for I/O virtualization because of the incurred costly VM exits. The shortcomings of prior software solutions against virtual interrupt delivery, a major source of VM exits, promoted the emergence of the hardware-based Posted-Interrupt (PI) technology. PI can provide non-exit interrupt delivery without compromising any virtualization benefit. However, it only acts on the half of the event path, i.e., the interrupt path, while guests I/O requests may also trigger a large amount of VM exits. Additionally, PI may still suffer a severe latency from the vCPU scheduling while delivering interrupts. Aiming at an optimal event path, we propose ES2 to simultaneously improve bidirectional I/O event delivery between guests and their devices. On the basis of PI, ES2 introduces hybrid I/O handling scheme for efficient I/O request delivery and intelligent interrupt redirection for enhanced I/O responsiveness. It does not require any modification to guest OS. We demonstrate that ES2 greatly reduces I/O-related VM exits with the exit handling time (EHT) below 2.5 percent for TCP streams and 0.1 percent for UDP streams, increases guest throughput by 1.9x for Memcached and 1.6x for Nginx, and keeps guest latency at a low level.
Xiaokang Hu, Jian Li 0021, Ruhui Ma, Haibing Guan
IEEE Trans. Cloud Comput.1
2021 A comprehensive test framework for cryptographic accelerators in the cloud
Hubin Zhang, Xiaokang Hu, Jian Li 0021, Haibing Guan
J. Syst. Archit.2
2021 STYX: A Hierarchical Key Management System for Elastic Content Delivery Networks on Public Clouds
abstract
Hosting content delivery networks (CDNs) on clouds has the potential to improve the performance as resources and caches can be placed closer to subscribers. However, avoiding data leakage over an untrusted public cloud is critical, especially for sensitive data such as the SSL private key. The popular Keyless SSL solution allows content owners to retain on-premise custody of SSL private keys on their own key servers, but this solution likely causes performance bottlenecks and impedes the elasticity of CDNs. This paper describes a novel key management system, named STYX, for transmitting trusted data over untrusted channels and storing them on untrusted platforms. STYX accomplishes secure key provisioning for CDN scale-out and the key is securely protected with full revocation rights for CDN scale-in. STYX is implemented as a three-phase hierarchical key management scheme by leveraging Intel Software Guard Extensions (SGX) and QuickAssist Technology (QAT). Furthermore, STYX supports CDN services by integrating Nginx as the SSL termination proxy and the popular Redis/Memcached/Apache as backend caching engines. The performance evaluation shows that STYX significantly outperforms the native HTTPS servers on the CDN node due to QAT acceleration, providing up to a 5× enhancement in throughput and a 50 percent reduction in latency.
Xiaokang Hu, Jian Li 0021, Changzheng Wei, Weigang Li 0002, Haibing Guan
IEEE Trans. Dependable Secur. Comput.1
2020 QWEB: High-Performance Event-Driven Web Architecture With QAT Acceleration
abstract
Hardware accelerators have been a promising solution to reduce the cost of cloud datacenters. This article investigates the acceleration of an important datacenter workload: the web server (or proxy) that faces high computational consumption originated from SSL/TLS processing and HTTP compression. Our study reveals that for the widely-deployed event-driven web architecture, the straight offloading of SSL/TLS or compression tasks suffers from frequent blockings in the offload I/O, leading to the underutilization of both CPU and accelerator resources. To achieve efficient acceleration, we propose QWEB, a comprehensive offload solution based on Intel QuickAssist Technology (QAT). QWEB introduces an asynchronous offload mode for SSL/TLS processing and a pipelining offload mode for HTTP compression, both allowing concurrent offload tasks from a single application process/thread. With these two novel offload modes, the blocking penalty is amortized or even eliminated, and the utilization rate of the parallel computation engines inside the QAT accelerator is greatly increased. The evaluation shows that QWEB provides up to 9x handshake performance with TLS-RSA (2048-bit) over the software baseline. Additionally, the secure data transfer throughput is enhanced by 2x for the SSL/TLS offloading only, 3.5x for the compression offloading only and 5x for the combined offloading.
Jian Li 0021, Xiaokang Hu, David Qian, Changzheng Wei, Gordon McFadden, Brian Will, Weigang Li 0002, Haibing Guan
IEEE Trans. Parallel Distributed Syst.2
2019 QTLS: high-performance TLS asynchronous offload framework with Intel® QuickAssist technology
abstract
Hardware accelerators are a promising solution to optimize the Total Cost of Ownership (TCO) of cloud datacenters. This paper targets the costly Transport Layer Security (TLS) and investigates the TLS acceleration for the widely-deployed event-driven TLS servers or terminators. Our study reveals an important fact: the straight offloading of TLS-involved crypto operations suffers from the frequent long-lasting blockings in the offload I/O, leading to the underutilization of both CPU and accelerator resources.
Xiaokang Hu, Changzheng Wei, Jian Li 0021, Brian Will, Lu Gong, Haibing Guan
PPoPP1
2019 QZFS: QAT Accelerated Compression in File System for Application Agnostic and Cost Efficient Data Storage
Xiaokang Hu, Fuzong Wang, Weigang Li 0002, Jian Li 0021, Haibing Guan
USENIX ATC1
2018 CoINT: Proactive Coordinator for Avoiding Interruptability Holder Preemption Problem in VSMP Environment
abstract
In a Virtual Symmetric Multiprocessing (VSMP) environment, the behavior of hypervisor scheduler can significantly influence a guest's I/O responsiveness. The interrupt remapping mechanism, which can leverage multiple virtual CPUs in the VSMP guest to process I/O events, is known to be an efficient and prevalent solution to improve the I/O performance. However, in this paper we identified a novel challenge called the “Interruptability Holder Preemption” (IHP) problem in interrupt remapping mechanism. The IHP issue presents that a virtual CPU (vCPU) disabling the interruptability of guest's network device is descheduled by the hypervisor scheduler, which can easily invalidate the efficiency of the interrupt remapping mechanism. To solve this problem, we propose CoINT, a gasket coordinator residing in the hypervisor, to substantially enhance the network I/O performance by empowering the hypervisor to be proactively aware of the interruptability information of the guest's network device. ColNT completely eliminates the “Interruptability Holder Preemption” problem and largely reduces I/Ointerrupt processing delay caused by hypervisor scheduler. We implement ColNT in KVM hypervisor and evaluate its efficiency and responsiveness using both macro and microlevel benchmarks. The results show that ColNT can improve the netperf throughput up to 3x compared with native KVM, and up to 1.3x compared with traditional interrupt remapping of hypervisor-Ievel solution, in sacrifice of an negligible and reasonable overhead in the hypervisor.
Xiaokang Hu, Jian Li 0021, Haibing Guan
INFOCOM2
2017 ES2: Aiming at an Optimal Virtual I/O Event Path
abstract
Improving the performance of I/O virtualization is a key issue for cloud and datacenter infrastructures, especially with the rapid increase of network interconnection speeds. Previous efforts have made the performance overhead associated with the virtual I/O data path largely negligible. The remaining bottlenecks mainly lie in the event path: hypervisor interventions trigger costly virtual machine (VM) exits and lead to dramatical performance degradation. Aiming at an optimal virtual I/O event path, we propose ES2, a comprehensive scheme that simultaneously improves bidirectional I/O event delivery between guest VMs and their devices. ES2 can provide efficient I/O request delivery, non-exit interrupt delivery and enhanced I/O responsiveness. Moreover, it does not require any modification to guest operating system (OS) or compromise any virtualization benefit. We demonstrate that ES2 greatly reduces VM exit rate with the time in guest (TIG) for I/O processing above 96% for TCP streams and 99% for UDP streams, increases guest throughput by 1.8× for Memcached and 2× for Apache, and keeps guest latency at a very low level.
Xiaokang Hu, Jian Li 0021, Ruhui Ma, Haibing Guan
ICPP1