Ming Yan 0009

dblp:51/5332-9 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-1176-0238ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Chariot: Accelerating Distributed Protocols with Data-Path Accelerator in DPUs
Jingqi Feng, Chunpu Huang, Sicheng Liang, Ming Yan 0009, Jie Wu 0003
IWQoS5
2025 WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic Scheduling
abstract
Existing large language model (LLM) serving systems typically batch the compute-bound prefill and I/O-bound decoding phases together.This co-location approach not only leads to significant interference between the two phases but also limits resource allocation and placements.To address these limitations, recent work proposes disaggregating the prefill and decoding phases to enhance performance.However, these works often rely on coarse-grained static scheduling strategies, resulting in imbalanced and insufficient resource utilization.For instance, compute resources for the prefill phase may be overloaded while those for the decoding phase remain idle, resulting in performance bottlenecks.In this paper, we propose WindServe, an efficient phase disaggregated LLM serving system that leverages stream-based, finegrained dynamic scheduling to enhance resource utilization and performance.WindServe features a global scheduler that monitors compute and memory resource usage to dynamically orchestrate cross-phase jobs, effectively reducing queuing delay and KV cache swapping overhead.We also introduce a stall-free rescheduling strategy to saturate the memory resources while minimizing the scheduling overhead from KV cache transfers.Furthermore, we design a stream-based approach to mitigate interference between prefill and decoding jobs.Our evaluation demonstrates that Wind-Serve achieves remarkable stability and SLO attainment under highload scenarios, outperforming state-of-the-art phase-disaggregated LLM serving systems by delivering a 4.28× improvement in TTFT median latency and a 1.5× reduction in TPOT P99 latency.
Jingqi Feng, Rui Zhang 0112, Sicheng Liang, Ming Yan 0009, Jie Wu 0003
ISCA5
2025 Offloading distributed key-value stores with off-path SmartNICs
Shangyi Sun, Rui Zhang 0112, Ming Yan 0009, Jie Wu 0003
Comput. Networks3
2024 Labor: Adaptive Lazy Compaction for Learned Index in LSM-Tree
Chunpu Huang, Lulu Chen, Rui Zhang 0112, Ming Yan 0009, Jie Wu 0003
COCOON (2)5
2024 HR-Tree: A Hybrid PMem-DRAM and Write-Optimized R-Tree for Spatial Data Storage
Rui Zhang 0112, Lulu Chen, Shangyi Sun, Ming Yan 0009, Jie Wu 0003
COCOON (2)5
2024 Revisiting Learned Index with Byte-addressable Persistent Storage
abstract
Byte-addressable Persistent Storage (BPS), such as persistent memory and CXL-enabled SSDs, has become an extension of main memory. This opens up new possibilities for indexes that operate and persist data directly on the memory bus. Recent learned indexes exploit data distribution and have shown great potential for some workloads. Despite some work proposed for integrating learned indexes into BPS, they are mainly based on Intel’s first-generation persistent memory. The current design suffers from the following problems: 1) Excessive storage line accesses due to large node in learned indexes; 2) Inefficient concurrency control due to volatile cache; 3) Write amplification due to mismatch access granularity.
Rui Zhang 0112, Sicheng Liang, Shangyi Sun, Shaonan Ma, Chengying Huan, Lulu Chen, Zhihui Lu 0002, Yang Xu 0010, Ming Yan 0009, Jie Wu 0003
ICPP10
2024 Mitigating Intra-host Network Congestion with SmartNIC
abstract
With the rapid development and wide deployment of high-speed network technologies like RDMA and the relatively stagnant evolution of intra-host resources, intra-host network congestion has become a potential issue that may affect the QoS of network applications. Offloading hotspot data to modern Smart-NICs, enabling hotspot data access completion on the SmartNIC, and reducing intra-host network traffic, is a promising solution to this issue. However, due to the limited SmartNIC resources and the complexity of network application requirements, achieving efficient offload is challenging.We present Magician, an architecture to mitigate intra-host network congestion with SmartNIC. Magician adopts a client-driven data access approach to avoid performance degradation caused by limited SmartNIC resources. Magician also introduces a SmartNIC-oriented hotspot data update strategy that dynamically refreshes hotspot data with minimal overhead. Moreover, we design a server-centric data consistency mechanism to ensure data consistency under concurrent access. We implement Magician within the key-value store. Evaluation of the key-value store with and without Magician suggests that, in the presence of intra-host network congestion, Magician significantly mitigates intra-host network congestion, leading to improved performance of network applications.
Lulu Chen, Chunpu Huang, Rui Zhang 0112, Yiren Zhou, Ming Yan 0009, Jie Wu 0003
IWQoS6
2024 Optimizing Inference Quality with SmartNIC for Recommendation System
abstract
Embedding-based recommendation systems are now widely used to recommend content for users, and have strict requirements on their latency and throughput. However, the latest recommendation models often exceed GPU HBM memory capacity, and the system is often deployed separately on computing nodes for GPU calculating and Parameter Servers for embedding tables’ storage. This architecture leads to a significant amount of network I/O during the inference process and reduces GPU utilization.In this paper, we propose SmartEmb, an inference framework that accelerates the network I/O of embedding table lookups through a specialized control plane of task reordering, prefetching and cache management. We offload these control planes on SmartNIC to avoid contention with the host CPU and gain better performance. We implemented the SmartEmb prototype on BlueField-2 and evaluated its performance. Our evaluation demonstrates that compared to the Nvidia HugeCTR HPS, SmartEmb can improve the quality of service by achieving up to 217% improvement in throughput and reducing latency by up to 190% of overall embedding layer look-ups in inference scenarios.
Ruixin Shi, Ming Yan 0009, Jie Wu 0003
IWQoS2
2023 PFtree: Optimizing Persistent Adaptive Radix Tree for PM Systems on eADR Platform
Rui Zhang 0112, Shangyi Sun, Lulu Chen, Yibo Huang 0005, Ming Yan 0009, Jie Wu 0003
DASFAA (1)6
2022 SKV: A SmartNIC-Offloaded Distributed Key-Value Store
abstract
In data center networks, applications such as dis-tributed key-value stores consume a lot of CPU resources. The performance of the entire system drops significantly under heavy load conditions. In order to improve the performance of key-value stores, many existing studies use RDMA (Remote Direct Memory Access) to reduce the communication overhead. However, RDMA primitives can only offload simple operations to the NIC, such as reading and writing remote memory. With the emergence of new hardware like SmartNICs, we consider whether we can offload more complex operations in distributed key-value stores to SmartNICs to reduce the load on CPU. In this paper we present SKV, a SmartNIC-offloaded distributed key-value store. In order to make full use of the offload ability of the SmartNIC, we make a detailed analysis on the characteristic and architecture of SmartNICs and distributed key-value stores. SKV offloads operations such as data replication to the SmartNIC. We design a new replication mechanism, which enables the server to separate background processing from the interaction with clients in the front. We implement SKV on the Mellanox BlueField SmartNIC. Our evaluations show that SKV improves the overall throughput by 14% and reduces latency by 21 % compared with baseline.
Shangyi Sun, Rui Zhang 0112, Ming Yan 0009, Jie Wu 0003
CLUSTER3
2022 An ultra-low latency and compatible PCIe interconnect for rack-scale communication
abstract
Emerging network-attached resource disaggregation architecture requires ultra-low latency rack-scale communication. However, current hardware offloading (e.g., RDMA) and user-space (e.g., mTCP) communication schemes still rely on heavily layered protocol stacks which requires the translation between PCIe bus and network protocol, or complex connection/memory resource management within RNICs, inevitably bringing latency overhead.
Yibo Huang 0005, Ming Yan 0009, Cunming Liang, Yang Xu 0010, Wenxiong Zou, Yiming Zhang 0018, Rui Zhang 0112, Chunpu Huang, Jie Wu 0003
CoNEXT3
2019 RDMA-driven MongoDB: An approach of RDMA enhanced NoSQL paradigm for large-Scale data processing
Yibo Huang 0005, Zhihui Lu 0002, Ming Yan 0009, Jie Wu 0003, Patrick C. K. Hung, Qifeng Tang
Inf. Sci.4
2017 Computer-aided cirrhosis diagnosis via automatic liver capsule extraction and combined geometry-texture features
abstract
This paper presents a computer-aided system for automatic diagnosis of cirrhosis based on ultrasound images. We first propose a dynamic programming algorithm to automatically extract the liver capsule, and then the continuity and smoothness of capsule serve as an important guideline for image classification. Via the decomposition of the ultrasound image in spatial and gray scales, the density and entropy of suspected nodular areas are used to describe texture features of liver parenchyma. Finally, a trained SVM classifier is applied to classify the samples into normal, mild, moderate and severe clinical stages of the disease. Experiment results show that the proposed method achieves better performance than existing approaches. Moreover, it can be used as an efficient method for early cirrhosis diagnosis in consideration of its high accuracy in distinguishing between normal and abnormal cases.
Xiang Liu 0003, Zhiqin Zhan, Ming Yan 0009, Jia Lin Song, Yan Qiu Chen
ICME3