Shixiong Qi

dblp:189/8941 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-1367-5544ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 6 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Not A DPU in Name Only! Unleashing RDMA-capable DPUs in Multi-Tenant Serverless Clouds with NADINO
Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan, Diman Zad Tootaghaj, Hardik Soni 0001, Puneet Sharma 0001
EuroSys1
2026 Credit-Guided Congestion Control on Wafer-Scale On-Chip Networks for Molecular Dynamics
abstract
Molecular dynamics (MD) is a cornerstone of scientific computing, but strong scaling often collapses at high parallelism because communication is bursty and highly sensitive to tail latency. MD advances by repeating a fixed timestep loop (one iteration of force computation and state update), and performance is largely determined by how quickly timesteps complete. A key reason is that each timestep contains short, synchronized communication phases, followed by a global dependency before the next timestep. Wafer-scale chips (WSCs) offer cycle-level latency and high on-chip bandwidth, yet their 2D mesh fabrics can still suffer burst-induced queue buildup; existing wavelet scheduling relies on a static stride that either over-injects (triggering credit backpressure) or over-throttles (wasting bandwidth) as conditions evolve.
Shixiong Qi, Zhan Wang 0003, Ning Kang 0007, Fan Yang 0096, Yuanzhe Wang, Guanglei Chen, Guangming Tan, Guojun Yuan
SIGCOMM2
2025 Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
abstract
Serverless computing offers resource efficiency but suffers from a heavyweight data plane. We present Palladium, a DPU-offloaded serverless data plane enabling distributed zero-copy communication. Palladium uses two-sided RDMA and cross-processor shared memory to mitigate limitations of wimpy DPU cores. Its DPU-enabled network engine (DNE) isolates RDMA resources and manages flows across tenants. By converting HTTP/TCP to RDMA at ingress, Palladium reduces protocol overhead on the critical path.
Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan, Diman Zad Tootaghaj, Hardik Soni 0001, Puneet Sharma 0001
SIGCOMM1
2024 SURE: Secure Unikernels Make Serverless Computing Rapid and Efficient
abstract
Current serverless platforms introduce non-trivial overheads when chaining and orchestrating loosely-coupled microservices. Containerized function runtimes are also constrained by insufficient isolation and excessive startup time. This motivates our exploration of a more efficient, secure, and rapid serverless design. We describe SURE, a unikernel-based serverless framework for fast function startup, equipped with a high-performance and secure data plane. SURE's data plane supports distributed zero-copy communication via the seamless interaction between zero-copy protocol stack (Z-stack) and local shared memory processing. To establish a lightweight service mesh, SURE uses library-based sidecars instead of individual userspace sidecars. We leverage Intel's Memory Protection Keys (MPK) as a lightweight capability to ensure safe access to the shared memory data plane. It also isolates the Trusted Computing Base (TCB) components in SURE's function runtime (e.g., library-based sidecar, scheduler, etc) from untrusted user code, while preserving the efficient single-address-space nature of unikernels. In particular, SURE prevents unintended privilege escalation involving MPK with an enhanced TCB. These combined efforts create a more secure and robust data plane while improving throughput up to 79X over Knative, a representative open-source serverless platform.
Federico Parola, Shixiong Qi, Anvaya B. Narappa, K. K. Ramakrishnan, Fulvio Risso
SoCC2
2024 Z-Stack: A High-Performance DPDK-Based Zero-Copy TCP/IP Protocol Stack
abstract
Data centers require high-performance and efficient networking for fast and reliable communication between applications. TCP/IP-based networking still plays a dominant role in data center networking to support a wide range of Layer-4 and Layer-7 applications, such as middleboxes and cloud-based microservices. However, traditional kernel-based TCP/IP stacks face performance challenges due to overheads such as context switching, interrupts, and copying. We present Z-stack, a high-performance userspace TCP/IP stack with a zero-copy design. Utilizing DPDK's Poll Mode Driver, Z-stack bypasses the kernel and moves packets between the NIC and the protocol stack in userspace, eliminating the overhead associated with kernel-based processing. Z-stack em-ploys polling-based packet processing that improves performance under high loads, and eliminates receive livelocks compared to interrupt-driven packet processing. With its zero-copy socket design, Z-stack eliminates copies when moving data between the user application and the protocol stack, which further minimizes latency and improves throughput. In addition, Z-stack seamlessly integrates with shared memory processing within the node, eliminating duplicate protocol processing and serializationldese-rialization overheads for intra-node communication. Z-stack uses F-stack as the starting point which integrates the proven TCP/IP stack from FreeBSD, providing a versatile solution for a variety of cloud use cases and improving performance of data center networking.
Anvaya B. Narappa, Federico Parola, Shixiong Qi, K. K. Ramakrishnan
LANMAN3
2024 SPRIGHT: High-Performance eBPF-Based Event-Driven, Shared-Memory Processing for Serverless Computing
abstract
Serverless computing promises an efficient, low-cost compute capability in cloud environments. However, existing solutions, epitomized by open-source platforms such as Knative, include heavyweight components that undermine this goal of serverless computing. Additionally, such serverless platforms lack dataplane optimizations to achieve efficient, high-performance function chains that facilitate the popular microservices development paradigm. Their use of unnecessarily complex and duplicate capabilities for building function chains severely degrades performance. ‘Cold-start’ latency is another deterrent. We describe, a lightweight, high-performance, responsive serverless framework. exploits shared memory processing and dramatically improves the scalability of the dataplane by avoiding unnecessary protocol processing and serialization-deserialization overheads. extensively leverages event-driven processing with the extended Berkeley Packet Filter (eBPF). We creatively use eBPF’s socket message mechanism to support shared memory processing, with overheads being strictly load-proportional. Compared to constantly-running, polling-based DPDK, achieves the same dataplane performance with 10$\times$less CPU usage under realistic workloads. Additionally, eBPF benefits, by replacing heavyweight serverless components, allowing us to keep functions ‘warm’ with negligible penalty. Our preliminary experimental results show that achieves an order of magnitude improvement in throughput and latency compared to Knative, while substantially reducing CPU usage, and obviates the need for ‘cold-start’.
Shixiong Qi, Leslie Monis, Ziteng Zeng, Ian-Chin Wang, K. K. Ramakrishnan
IEEE/ACM Trans. Netw.1
2023 X-IO: A High-performance Unified I/O Interface using Lock-free Shared Memory Processing
abstract
Cloud-native microservice applications use different communication paradigms to network microservices, including both synchronous and asynchronous I/O for exchanging data. Existing solutions depend on kernel-based networking, incurring significant overheads. The interdependence between microservices for these applications involves considerable communication, including contention between multiple concurrent flows or user sessions. In this paper, we design X-IO, a high-performance unified I/O interface that is built on top of shared memory processing with lock-free producer/consumer rings, eliminating kernel networking overheads and contention. X-IO offers a feature-rich interface. X-IO’s zero-copy interface supports building provides truly zero-copy data transfers between microservices, achieving high performance. X-IO also provides a POSIX-like socket interface using HTTP/REST API to achieve seamless porting of microservices to X-IO, without any change to the application code. X-IO supports concurrent connections for microservices that require distinct user sessions operating in parallel. Our preliminary experimental results show that X-IO’s zero-copy interfaces achieve 2.8x-4.1x performance improvement compared to kernel-based interfaces. Its socket interfaces outperform kernel TCP sockets and achieve performance close to UNIX-domain sockets. The HTTP/REST APIs in X-IO perform 1.4 x-2.3 x better than kernel-based alternatives with concurrent connections.
Shixiong Qi, Han-Sing Tsai, Yu-Sheng Liu, K. K. Ramakrishnan, Jyh-Cheng Chen
NetSoft1
2023 MiddleNet: A Unified, High-Performance NFV and Middlebox Framework With eBPF and DPDK
abstract
Traditional network resident functions (e.g., firewalls, network address translation) and middleboxes (caches, load balancers) have moved from purpose-built appliances to softwarebased components. However, L2/L3 network functions (NFs) are being implemented on Network Function Virtualization (NFV) platforms that extensively exploit kernel-bypass technology. They often use DPDK for zero-copy delivery and high performance. On the other hand, L4/L7 middleboxes, which have a greater emphasis on functionality, take advantage of a full-fledged kernelbased system. L2/L3 NFs and L4/L7 middleboxes continue to be handled by distinct platforms on different nodes. This paper proposes MiddleNet that develops a unified network resident function framework that supports L2/L3 NFs and L4/L7 middleboxes. MiddleNet supports function chains that are essential in both NFV and middlebox environments. MiddleNet uses the Data Plane Development Kit (DPDK) library for zero-copy packet delivery without interrupt-based processing, to enable the ’bumpin-the-wire’ L2/L3 processing performance required of NFV. To support L4/L7 middlebox functionality, MiddleNet utilizes a consolidated, kernel-based protocol stack for processing, avoiding a dedicated protocol stack for each function. MiddleNet fully exploits the event-driven capabilities of the extended Berkeley Packet Filter (eBPF) and seamlessly integrates it with shared memory for high-performance communication in L4/L7 middlebox function chains. The overheads for MiddleNet in L4/L7 are strictly load-proportional, without needing the dedicated CPU cores of DPDK-based approaches. MiddleNet supports flow-dependent packet processing by leveraging Single Root I/O Virtualization (SR-IOV) to dynamically select the packet processing needed (Layers 2 -7). Our experimental results show that MiddleNet achieves high performance in such a unified environment.
Shixiong Qi, Ziteng Zeng, Leslie Monis, K. K. Ramakrishnan
IEEE Trans. Netw. Serv. Manag.1
2022 MiddleNet: A High-Performance, Lightweight, Unified NFV and Middlebox Framework
abstract
Traditional network resident functions (e.g., firewalls, network address translation) and middleboxes (caches, load balancers) have moved from purpose-built appliances to software-based components. However, L2/L3 network functions (NFs) are being implemented on Network Function Virtualization (NFV) platforms that extensively exploit kernel-bypass technology. They often use DPDK for zero-copy delivery and high performance. On the other hand, L4/L7 middleboxes, which usually require full network protocol stack support, take advantage of a full-fledged kernel-based system with a greater emphasis on functionality. Thus, L2/L3 NFs and middleboxes continue to be handled by distinct platforms on different nodes.This paper proposes MiddleNet that seeks to overcome this dichotomy by developing a unified network resident function framework that supports L2/L3 NFs and L4/L7 middleboxes. MiddleNet supports function chains that are essential in both NFV and middlebox environments. MiddleNet uses DPDK for zero-copy packet delivery without interrupt-based processing, to enable the ‘bump-in-the-wire’ L2/L3 processing performance required of NFV. To support L4/L7 middlebox functionality, MiddleNet utilizes a consolidated, kernel-based protocol stack processing, avoiding a dedicated protocol stack for each function. MiddleNet fully exploits the event-driven capabilities provided by the extended Berkeley Packet Filter (eBPF) and seamlessly integrates it with shared memory for high-performance communication in L4/L7 middlebox function chains. The overheads for MiddleNet are strictly load-proportional, without needing the dedicated CPU cores of DPDK-based approaches. MiddleNet supports flow-dependent packet processing by leveraging Single Root I/O Virtualization (SR-IOV) to dynamically select packet processing needed (Layer 2 to Layer 7). Our experimental results show that MiddleNet can achieve high performance in such a unified environment.
Ziteng Zeng, Leslie Monis, Shixiong Qi, K. K. Ramakrishnan
NetSoft3
2022 DEMO: MiddleNet: A High-Performance, Lightweight, Unified NFV & Middlebox Framework
abstract
Softwarized network resident functions have been extensively used to replace purpose-built appliances. However, there is a lack of alternatives for richer network resident functionality with a seamless combination of L2/L3 Network Function Virtualization (NFV) and L4/L7 middleboxes.We propose MiddleNet, a unified L2/L3 NFV and L4/L7 middlebox framework. MiddleNet uses DPDK in L2/L3 NFV to achieve high-performance, zero-copy packet delivery. MiddleNet exploits the event-driven capabilities of extended Berkeley Packet Filter (eBPF) to build up lightweight L4/L7 middleboxes with load-proportional overheads. MiddleNet constructs complex L2/L3 NF and L4/L7 middlebox function chains with low overhead using shared memory communication. With the integration of Single Root I/O Virtualization (SR-IOV), MiddleNet supports dynamically selecting packet processing layers (L2 to L7) based on the flow. In this demo, we show MiddleNet’s operation.
Ziteng Zeng, Leslie Monis, Shixiong Qi, K. K. Ramakrishnan
NetSoft3
2022 L25GC: a low latency 5G core network based on high-performance NFV platforms
abstract
Cellular network control procedures (e.g., mobility, idle-active transition to conserve energy) directly influence data plane behavior, impacting user-experienced delay. Recognizing this control-data plane interdependence, L25GC re-architects the 5G Core (5GC) network, and its processing, to reduce latency of control plane operations and their impact on the data plane. Exploiting shared memory, L25GC eliminates message serialization and HTTP processing overheads, while being 3GPP-standards compliant. We improve data plane processing by factoring the functions to avoid control-data plane interference, and using scalable, flow-level packet classifiers for forwarding-rule lookups. Utilizing buffers at the 5GC, L25GC implements paging, and an intelligent handover scheme avoiding 3GPP's hairpin routing, and data loss caused by limited buffering at 5G base stations, reduces delay and unnecessary message processing. L25GC's integrated failure resiliency transparently recovers from failures of 5GC software network functions and hardware much faster than 3GPP's reattach recovery procedure. L25GC is built based on free5GC, an open-source kernel-based 5GC implementation. L25GC reduces event completion time by ~50% for several control plane events and improves data packet latency (due to improved control plane communication) by ~2×, during paging and handover events, compared to free5GC. L25GC's design is general, although current implementation supports a limited number of user sessions.
Vivek A. Jain, Hao-Tse Chu, Shixiong Qi, Chia-An Lee, Hung-Cheng Chang, Cheng-Ying Hsieh, K. K. Ramakrishnan, Jyh-Cheng Chen
SIGCOMM3
2022 SPRIGHT: extracting the server from serverless computing! high-performance eBPF-based event-driven, shared-memory processing
abstract
Serverless computing promises an efficient, low-cost compute capability in cloud environments. However, existing solutions, epitomized by open-source platforms such as Knative, include heavyweight components that undermine this goal of serverless computing. Additionally, such serverless platforms lack dataplane optimizations to achieve efficient, high-performance function chains that facilitate the popular microservices development paradigm. Their use of unnecessarily complex and duplicate capabilities for building function chains severely degrades performance. 'Cold-start' latency is another deterrent.
Shixiong Qi, Leslie Monis, Ziteng Zeng, Ian-Chin Wang, K. K. Ramakrishnan
SIGCOMM1
2021 Mu: An Efficient, Fair and Responsive Serverless Framework for Resource-Constrained Edge Clouds
abstract
Serverless computing platforms simplify development, deployment, and automated management of modular software functions. However, existing serverless platforms typically assume an over-provisioned cloud, making them a poor fit for Edge Computing environments where resources are scarce. In this paper we propose a redesigned serverless platform that comprehensively tackles the key challenges for serverless functions in a resource constrained Edge Cloud.
Viyom Mittal, Shixiong Qi, Ratnadeep Bhattacharya, Xiaosu Lyu, Sameer G. Kulkarni, Dan Li 0001, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001
SoCC2
2021 Fast Function Instantiation with Alternate Virtualization Approaches
abstract
This paper focuses on the need for emerging domains such as serverless and in-network computing, where applications are often hosted on virtualized compute instances (e.g., containers and unikernels), to have applications startup as quickly as possible. We provide a qualitative and quantitative analysis of containers and unikernels with regard to the startup time. We analyze these in-depth and identify the key components and their impact under scale on the startup latency. We study how startup time scales as we launch multiple instances concurrently. We study the contribution of popular Container Networking Interfaces (CNIs), to the startup time.
Vivek A. Jain, Shixiong Qi, K. K. Ramakrishnan
LANMAN2
2021 Towards a Proactive Lightweight Serverless Edge Cloud for Internet-of-Things Applications
abstract
Edge cloud solutions that bring the cloud closer to the sensors can be very useful to meet the low latency requirements of many Internet-of-Things (IoT) applications. However, IoT traffic can also be intermittent, so running applications constantly can be wasteful. Therefore, having a serverless edge cloud that is responsive and provides low-latency features is a very attractive option for a resource and cost-efficient IoT application environment.In this paper, we discuss the key components needed to support IoT traffic in the serverless edge cloud and identify the critical challenges that make it difficult to directly use existing serverless solutions such as Knative, for IoT applications. These include overhead from heavyweight components for managing the overall system and software adaptors for communication protocol translation used in off-the-shelf serverless platforms that are designed for large-scale centralized clouds. The latency imposed by ‘cold start’ is a further deterrent.To address these challenges we redesign several components of the Knative serverless framework. We use a streamlined protocol adaptor to leverage the MQTT IoT protocol in our serverless framework for IoT event processing. We also create a novel, event-driven proxy based on the extended Berkeley Packet Filter (eBPF), to replace the regular heavyweight Knative queue proxy. Our preliminary experimental results show that the event-driven proxy is a suitable replacement for the queue proxy in an IoT serverless environment and results in lower CPU usage and a higher request throughput.
Ian-Chin Wang, Shixiong Qi, Elizabeth Liri, K. K. Ramakrishnan
NAS2
2021 Assessing Container Network Interface Plugins: Functionality, Performance, and Scalability
abstract
Kubernetes, an open-source container orchestration platform, has been widely adopted by cloud service providers (CSPs) for its advantages in simplifying container deployment, scalability, and scheduling. Networking is one of the central components of Kubernetes, providing connectivity between different Pods (a group of containers) both within the same host and across hosts. To bootstrap Kubernetes networking, the Container Network Interface (CNI) provides a unified interface for the interaction between container runtimes. There are several CNI implementations, available as open-source `CNI plugins'. While they differ in functionality and performance, it is a challenge for a cloud provider to differentiate and choose the appropriate plugin for their environment. In this article, we compare the various open-source CNI plugins available from the community, qualitatively, and through detailed quantitative measurements. With our experimental evaluation, we analyze the overheads and bottlenecks for each CNI plugin, especially because of the interaction with the datapath/iptables as well as the host network stack. Overlay tunnel offload support in the network interface card plays a significant role in achieving the good performance of CNIs that use overlay tunnels for inter-host Pod-to-Pod communication. We also study scalability with an increasing number of Pods, as well as with HTTP workloads, and briefly evaluate Pod startup latency. Our measurement results inform the outline of an ideal CNI environment for Kubernetes.
Shixiong Qi, Sameer G. Kulkarni, K. K. Ramakrishnan
IEEE Trans. Netw. Serv. Manag.1
2020 Understanding Container Network Interface Plugins: Design Considerations and Performance
abstract
Kubernetes, an open-source container orchestration platform, has been widely adopted by cloud service providers (CSPs) for its advantages in simplifying container deployment, scalability and scheduling. Networking is one of the central components of Kubernetes, providing connectivity between different pods (group of containers) both within the same host and across hosts. To bootstrap Kubernetes networking, the Container Network Interface (CNI) provides a unified interface for the interaction between container runtimes. There are several CNI implementations, available as open-source ‘CNI plugins’. While they differ in functionality and performance, it is a challenge for a cloud provider to differentiate and choose the appropriate plugin for their environment. In this paper, we compare the various open source CNI plugins available from the community, qualitatively and through detailed quantitative measurements. With our experimental evaluation, we analyze the overheads and bottlenecks for each CNI plugin, as a result of the network model it implements, interaction with the host network protocol stack and the network policies implemented in iptables rules. The choice of the CNI plugin may also be based on whether intra-host or inter-host communication dominates.
Shixiong Qi, Sameer G. Kulkarni, K. K. Ramakrishnan
LANMAN1
2017 Testudo: A Low Latency and High-Efficient Memory-Centric Network Using Optical Interconnect
abstract
With the continuing-scaling of future multicore processors, the performance requirements on memory access has been put forward much higher. Memory- centric network is deemed as a promising communication paradigm for core-to-memory interconnect in future multicore processors. However, the traditional electrical interconnect has the drawbacks of limited capacity, high communication delay, poor scalability and low energy efficiency, which further limits the performance improvement of the system. To support high-performance communication for memory access, we propose the Testudo architecture, an optically connected memory-centric network (MCN), which utilizes the emerging optical interconnect technology and 3D-stacking memory technology to achieve high bandwidth, low power consumption and high scalability. Testudo is designed based on multiple optical crossbar organized in a torus- like topology. Each optical crossbar is in multiple-write- multiple-read construction, which provides high connectivity for IP cores. By employing an all optical, token-based arbitration scheme with low complexity, the memory access communication is contention-free. Simulation results show that Testudo improves the performance significantly compared to the electrical mesh topology.
Shixiong Qi, Huaxi Gu, Haibo Zhang 0001, Yawen Chen 0001
GLOBECOM1