Ming Liu 0027

dblp:20/2039-27 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-6509-9449ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 19 · 2 first-author · 14 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Understanding and Optimizing Database Pushdown on Disaggregated Storage
abstract
Database pushdown is a widely adopted technique under compute-storage disaggregation. The rising network and I/O speeds, coupled with stagnated compute and memory subsystems of a disaggregated storage architecture in the past decade, render state-of-the-art policy-driven pushdown designs ineffective. This is because the query performance bottleneck has shifted from network and I/O to compute, where computing power at the storage layer becomes scarce.
Yuebin Bai, Ming Liu 0027
ASPLOS (2)4
2026 SG-IOV: Socket-Granular I/O Virtualization for SmartNIC-Based Container Networks
abstract
I/O Virtualization (IOV) is a cornerstone of cloud computing, with container networking as a critical form of IOV in modern cloud paradigms. While container networks serve as feature-rich infrastructure, they incur a high CPU tax yet leave room for efficiency improvement. A natural idea is to offload container networks onto hardware such as SmartNICs via IOV interfaces. However, existing IOV mechanisms, such as SR-IOV, are misaligned with container requirements: limited device scalability versus high container density, packet-layer abstraction versus application-layer processing demands, and coarse-grained virtualization versus fine-grained container workloads.
Chenxingyu Zhao, Jaehong Min, Shengkai Lin, Wei Zhang 0052, Kaiyuan Zhang 0001, Ming Liu 0027, Arvind Krishnamurthy
ASPLOS (2)7
2026 Building A CSFQ-Inspired Transport for Switched CXL Memory Pooling
Zerui Guo, Emily Shriver, Ming Liu 0027
NSDI3
2026 Co-Designing Traffic Control with NVMe-oF for Disaggregated Storage: A Comparative Study of Switched and Switchless SAN Architectures
Chendong Wang, Joontaek Oh, Ming Liu 0027
NSDI3
2026 Understanding and Profiling the Accelerator Chiplet Network Using PingPoint
abstract
Emerging chiplet-based accelerators introduce a new class of intrahost networks—the Accelerator Chiplet Network (ACN)—that links compute chiplets, IO chiplets, and memory modules and increasingly governs application performance. Yet ACN behavior remains largely opaque: existing tools overlook on-package communication and instead attribute overheads to compute or memory subsystems, while ACN-induced latency, bandwidth heterogeneity, and congestion are hard to observe due to proprietary microarchitectures, tight coupling with the execution pipeline, and complex mappings between application activity and hardware.
Junyeol Ryu, Ming Liu 0027, Matthew D. Sinclair
SIGCOMM2
2026 Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QP
abstract
Rack-scale interconnects serve as critical datapaths for emerging communication-intensive systems to scale up. Innovative solutions for this datapath are rising at a rapid pace, especially those based on Ethernet. However, existing hardware-based solutions, such as RDMA, face performance issues, particularly for small-message memory access, and suffer from the inflexibility of hardware-fixed processing. The community is actively pursuing efficient, flexible, and cost-effective rack-scale datapaths.
Chenxingyu Zhao, Jaehong Min, Ming Liu 0027, Arvind Krishnamurthy
SIGCOMM5
2025 Server Chiplet Networking
abstract
Emerging chiplet-based server platforms and the resulting server chiplet networking present a fundamental shift in the (intra-)host network. Unlike conventional monolithic servers, compute chiplets, I/O chiplets, off-chip memory, and peripheral devices communicate through a collection of heterogeneous interconnects and links, formalizing a new server chiplet network substrate that has not been explored before. This paper makes an initial step by characterizing two generations of AMD EPYC chiplet servers, identifying four communication idiosyncrasies, and summarizing the design implications. We outline some future directions under server chiplet networking and discuss how to build next-generation server systems and applications.
Seunghyun An, Joontaek Oh, Ming Liu 0027
HotNets3
2025 RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs
abstract
The emerging microservice/serverless-based cloud programming paradigm and the rising networking speeds leave the RPC stack as the predominant data center tax. Domain-specific hardware acceleration holds the potential to disentangle the overhead and save host CPU cycles. However, state-of-the-art RPC accelerators integrate RPC logic into the CPU or use specialized low-latency interconnects, hardly adopted in commodity servers. To this end, we design and implement RpcNIC, a software-hardware co-designed SmartNIC that enables efficient RPC layer offloading and reconfigurable RPC kernel offloading. RpcNIC connects to the server through the most widely used PCIe interconnect. To grapple with the ramifications of PCIe-induced challenges, RpcNIC introduces three techniques: (a) a target-aware deserializer that effectively batches cross-PCIe writes on the SmartNIC’s SRAM using compacted hardware data structures; (b) a memory-affinity CPU-SmartNIC collaborative serializer, which trades additional host memory copies for slow cross PCIe-transfers; (c) an automatic field update technique that transparently codifies the schema based on dynamic reconfigure RPC kernels to minimize superfluous PCIe traversals. We prototype RpcNIC using the Xilinx U280 FPGA card. On HyperProtoBench, RpcNIC achieves an average of 2.3 × lower RPC layer processing time than a comparable RPC accelerator baseline and demonstrates 2.6 × achievable throughput improvement in the end-to-end cloud workload.
Jie Zhang 0081, Hongjing Huang, Xuzheng Chen, Xiang Li 0205, Jieru Zhao, Ming Liu 0027, Zeke Wang
HPCA6
2025 Building an Elastic Block Storage over EBOFs Using Shadow Views
Ming Liu 0027
NSDI2
2025 Understanding and Profiling NVMe-over-TCP Using ntprof
Yuyuan Kang, Ming Liu 0027
NSDI2
2025 Building Massive MIMO Baseband Processing on a Single-Node Supercomputer
Xincheng Xie, Wentao Hou, Zerui Guo, Ming Liu 0027
NSDI4
2025 White-Boxing RDMA with Packet-Granular Software Control
Chenxingyu Zhao, Jaehong Min, Ming Liu 0027, Arvind Krishnamurthy
NSDI3
2025 Understanding and Profiling CXL.mem Using PathFinder
abstract
CXL.mem and the resulting memory pool are promising and gaining great attention. Unlike local memory, CXL DIMMs stay at the I/O subsystem, whose inferior performance can easily impact the processor pipeline and memory subsystem, yielding performance interference, hardware contention, obscure behaviors, and underutilized communication and computing resources. However, our community lacks a tool to understand and profile the CXL.mem protocol execution end-to-end between CPU and remote DIMM.
Zerui Guo, Yuebin Bai, Mahesh Ketkar, Hugh Wilkinson, Ming Liu 0027
SIGCOMM6
2024 Demystifying Datapath Accelerator Enhanced Off-path SmartNIC
abstract
Network speeds grow quickly in the modern cloud, so SmartNICs are introduced to offload network processing tasks, even application logic. However, typical multicore SmartNICs such as BlueFiled-2 are only capable of processing control-plane tasks with their embedded processors that have limited memory bandwidth and computing power. On the other hand, cloud applications evolve rapidly, such that a limited number of fixed hardware engines in a SmartNIC cannot satisfy the requirements of cloud applications. Therefore, SmartNIC programmers call for a programmable datapath accelerator (DPA) to process network traffic at line rate. However, no existing work has unveiled the performance characteristics of the existing DPA. To this end, we present the first architectural characterization of the latest DPA-enhanced BlueFiled-3 (BF3) SmartNIC. Our evaluation results indicate that BF3's DPA is significantly wimpier than the off-path Arm processor and the host CPU. However, we still identify that DPA has three unique architectural characteristics that unleash the performance potential of DPA. Specifically, we demonstrate how to take advantage of DPA's three architectural characteristics regarding computing, networking, and memory subsystems. Then we propose three important guidelines for programmers to fully unleash the potential of DPA. To demonstrate the effectiveness of our approach, we conduct detailed case studies regarding each guideline. Our case study on key-value aggregation achieves up to$4.3 \times$higher throughput by using our guidelines to optimize memory combinations.
Xuzheng Chen, Jie Zhang 0081, Lingjun Zhu, Yin Zhang 0006, Ming Liu 0027, Zeke Wang
ICNP10
2024 Understanding Routable PCIe Performance for Composable Infrastructures
Wentao Hou, Jie Zhang 0081, Zeke Wang, Ming Liu 0027
NSDI4
2024 eZNS: Elastic Zoned Namespace for Enhanced Performance Isolation and Device Utilization
abstract
Emerging Zoned Namespace (ZNS) SSDs, providing the coarse-grained zone abstraction, hold the potential to significantly enhance the cost efficiency of future storage infrastructure and mitigate performance unpredictability. However, existing ZNS SSDs have a static zoned interface, making them in-adaptable to workload runtime behavior, unscalable to underlying hardware capabilities, and interfering with co-located zones. Applications either under-provision the zone resources yielding unsatisfied throughput, create over-provisioned zones and incur costs, or experience unexpected I/O latencies. We propose eZNS, an elastic-ZNS interface that exposes an adaptive zone with predictable characteristics. eZNS comprises two major components: a zone arbiter that manages zone allocation and active resources on the control plane, and a hierarchical I/O scheduler with read congestion control and write admission control on the data plane. Together, eZNS enables the transparent use of a ZNS SSD and closes the gap between application requirements and zone interface properties. Our evaluations over RocksDB demonstrate that eZNS outperforms a static zoned interface by 17.7% and 80.3% in throughput and tail latency, respectively, at most.
Jaehong Min, Chenxingyu Zhao, Ming Liu 0027, Arvind Krishnamurthy
ACM Trans. Storage3
2023 A Generic Service to Provide In-Network Aggregation for Key-Value Streams
abstract
Key-value stream aggregation is a common operation in distributed systems, which requires intensive computation and network resources. We propose a generic in-network aggregation service for key-value streams, ASK, to accelerate the aggregation operations in diverse distributed applications. ASK is a switch-host co-designed system, where the programmable switch provides a best-effort aggregation service, and the host runs a daemon to interact with applications. ASK makes in-depth optimization tailored to traffic characteristics, hardware restrictions, and network unreliable natures: it vectorizes multiple key-value tuples’ aggregation of one packet in one switch pipeline pass, which improves the per-host’s goodput; it develops a lightweight reliability mechanism for key-value stream’s asynchronous aggregation, which guarantees computation correctness; it designs a hot-key agnostic prioritization for key-skewed workloads, which improves the switch memory utilization. We prototype ASK and use it to support Spark and BytePS. The evaluation shows that ASK could accelerate pure key-value aggregation tasks by up to 155 times and big data jobs by 3-5 times, and be backward compatible with existing INA-empowered distributed training solutions with the same speedup.
Yongchao He, Wenfei Wu, Yanfang Le, Ming Liu 0027, ChonLam Lao
ASPLOS (2)4
2023 Fabric-Centric Computing
abstract
Emerging memory fabrics and the resulting composable infrastructures have fundamentally challenged our conventional wisdom on how to build efficient rack/cluster-scale systems atop. This position paper proposes a new computing paradigm-called Fabric-Centric Computing (FCC)-that views the memory fabric as a first-class citizen to instantiate, orchestrate, and reclaim computations over composable infrastructures. We describe its design principles, report our early experiences, and discuss a new intermediate system stack proposal that harnesses the uniqueness of this cluster interconnect and realizes the vision of FCC.
Ming Liu 0027
HotOS1
2023 LogNIC: A High-Level Performance Model for SmartNICs
abstract
SmartNICs have become an indispensable communication fabric and computing substrate in today’s data centers and enterprise clusters, providing in-network computing capabilities for traversed packets and benefiting a range of applications across the system stack. Building an efficient SmartNIC-assisted solution is generally non-trivial and tedious as it requires programmers to understand the SmartNIC architecture, refactor application logic to match the device’s capabilities and limitations, and correlate an application execution with traffic characteristics. A high-level SmartNIC performance model can decouple the underlying SmartNIC hardware device from its offloaded software implementations and execution contexts, thereby drastically simplifying and facilitating the development process. However, prior architectural models can hardly be applied due to their limited capabilities in dissecting the SmartNIC-offloaded program’s complexity, capturing the nondeterministic overlapping between computation and I/O, and perceiving diverse traffic profiles.
Zerui Guo, Yuebin Bai, Daehyeok Kim, Michael M. Swift, Aditya Akella, Ming Liu 0027
MICRO7
2023 eZNS: An Elastic Zoned Namespace for Commodity ZNS SSDs
Jaehong Min, Chenxingyu Zhao, Ming Liu 0027, Arvind Krishnamurthy
OSDI3
2023 LEED: A Low-Power, Fast Persistent Key-Value Store on SmartNIC JBOFs
abstract
The recent emergence of low-power high-throughput programmable storage platforms-SmartNIC JBOF (just-a-bunch-of-flash)-motivates us to rethink the cluster architecture and system stack for energy-efficient large-scale data-intensive workloads. Unlike conventional systems that use an array of server JBOFs or embedded storage nodes, the introduction of SmartNIC JBOFs has drastically changed the cluster compute, memory, and I/O configurations. Such an extremely imbalanced architecture makes prior system design philosophies and techniques either ineffective or invalid.
Zerui Guo, Chenxingyu Zhao, Yuebin Bai, Michael M. Swift, Ming Liu 0027
SIGCOMM6
2021 Gimbal: enabling multi-tenant storage disaggregation on SmartNIC JBOFs
abstract
Emerging SmartNIC-based disaggregated NVMe storage has become a promising storage infrastructure due to its competitive IO performance and low cost. These SmartNIC JBOFs are shared among multiple co-resident applications, and there is a need for the platform to ensure fairness, QoS, and high utilization. Unfortunately, given the limited computing capability of the SmartNICs and the non-deterministic nature of NVMe drives, it is challenging to provide such support on today's SmartNIC JBOFs.
Jaehong Min, Ming Liu 0027, Tapan Chugh, Chenxingyu Zhao, Andrew Wei, In Hwan Doh, Arvind Krishnamurthy
SIGCOMM2
2021 Automated SmartNIC Offloading Insights for Network Functions
abstract
The gap between CPU and networking speeds has motivated the development of SmartNICs for NF (network functions) offloading. However, offloading performance is predicated upon intricate knowledge about SmartNIC hardware and careful hand-tuning of the ported programs. Today, developers cannot easily reason about the offloading performance or the effectiveness of different porting strategies without resorting to a trial-and-error approach.
Yiming Qiu 0001, Jiarong Xing, Kuo-Feng Hsu, Qiao Kang, Ming Liu 0027, Srinivas Narayana, Ang Chen 0001
SOSP5
2021 Xenic: SmartNIC-Accelerated Distributed Transactions
abstract
High-performance distributed transactions require efficient remote operations on database memory and protocol metadata. The high communication cost of this workload calls for hardware acceleration. Recent research has applied RDMA to this end, leveraging the network controller to manipulate host memory without consuming CPU cycles on the target server. However, the basic read/write RDMA primitives demand trade-offs in data structure and protocol design, limiting their benefits. SmartNICs are a flexible alternative for fast distributed transactions, adding programmable compute cores and on-board memory to the network interface. Applying measured performance characteristics, we design Xenic, a SmartNIC-optimized transaction processing system. Xenic applies an asynchronous, aggregated execution model to maximize network and core efficiency. Xenic's co-designed data store achieves low-overhead remote object accesses. Additionally, Xenic uses flexible, point-to-point communication patterns between SmartNICs to minimize transaction commit latency. We compare Xenic against prior RDMA- and RPC-based transaction systems with the TPC-C, Retwis, and Smallbank benchmarks. Our results for the three benchmarks show 2.42x, 2.07x, and 2.21x throughput improvement, 59%, 42%, and 22% latency reduction, while saving 2.3, 8.1, and 10.1 threads per server.
Henry Schuh, Weihao Liang, Ming Liu 0027, Jacob Nelson 0001, Arvind Krishnamurthy
SOSP3
2020 Clara: Performance Clarity for SmartNIC Offloading
abstract
The gap between CPU and networking speeds has motivated the development of SmartNICs for near-network processing. Recent work has shown that many network functions can benefit from SmartNIC offloading, but identifying the best porting strategy requires hand-tuning and workload-specific optimizations. The developer has no easy way to understand the ported performance beforehand
Yiming Qiu 0001, Qiao Kang, Ming Liu 0027, Ang Chen 0001
HotNets3
2020 Fine-Grained Replicated State Machines for a Cluster Storage System
Ming Liu 0027, Arvind Krishnamurthy, Harsha V. Madhyastha, Rishi Bhardwaj, Chinmay Kamat, Huapeng Yuan, Aditya Jaltade, Roger Liao, Pavan Konka, Anoop Jawahar
NSDI1
2020 Programmable Calendar Queues for High-speed Packet Scheduling
Naveen Kr. Sharma, Chenxingyu Zhao, Ming Liu 0027, Pravein G. Kannan, Changhoon Kim, Arvind Krishnamurthy, Anirudh Sivaraman
NSDI3
2019 Offloading distributed applications onto smartNICs using iPipe
abstract
Emerging Multicore SoC SmartNICs, enclosing rich computing resources (e.g., a multicore processor, onboard DRAM, accelerators, programmable DMA engines), hold the potential to offload generic datacenter server tasks. However, it is unclear how to use a SmartNIC efficiently and maximize the offloading benefits, especially for distributed applications. Towards this end, we characterize four commodity SmartNICs and summarize the offloading performance implications from four perspectives: traffic control, computing capability, onboard memory, and host communication.
Ming Liu 0027, Tianyi Cui, Henry Schuh, Arvind Krishnamurthy, Simon Peter 0001
SIGCOMM1
2019 E3: Energy-Efficient Microservices on SmartNIC-Accelerated Servers
Ming Liu 0027, Simon Peter 0001, Arvind Krishnamurthy, Phitchaya Mangpo Phothilimthana
USENIX ATC1
2018 Approximating Fair Queueing on Reconfigurable Switches
Naveen Kr. Sharma, Ming Liu 0027, Kishore Atreya, Arvind Krishnamurthy
NSDI2
2018 Floem: A Programming System for NIC-Accelerated Network Applications
Phitchaya Mangpo Phothilimthana, Ming Liu 0027, Antoine Kaufmann, Simon Peter 0001, Rastislav Bodík, Thomas E. Anderson
OSDI2
2017 IncBricks: Toward In-Network Computation with an In-Network Cache
abstract
The emergence of programmable network devices and the increasing data traffic of datacenters motivate the idea of in-network computation. By offloading compute operations onto intermediate networking devices (e.g., switches, network accelerators, middleboxes), one can (1) serve network requests on the fly with low latency; (2) reduce datacenter traffic and mitigate network congestion; and (3) save energy by running servers in a low-power mode. However, since (1) existing switch technology doesn't provide general computing capabilities, and (2) commodity datacenter networks are complex (e.g., hierarchical fat-tree topologies, multipath communication), enabling in-network computation inside a datacenter is challenging.
Ming Liu 0027, Jacob Nelson 0001, Luis Ceze, Arvind Krishnamurthy, Kishore Atreya
ASPLOS1