VLDB 2026 Research / reviewers in the wild / expert
Mingcong Han
dblp:286/1961
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0008-1536-7485ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accurate and Ultra-Fast Launch-Time Validation of Idempotency for GPU KernelsabstractWe discovered that a GPU kernel can have both idempotent and non-idempotent instances depending on the input. These kernels, called conditionally-idempotent, are common in real-world GPU applications—490 out of 547 from six popular applications. This finding reveals a limitation in previous work that statically classifies GPU kernels as idempotent or non-idempotent, potentially compromising the correctness and effectiveness of idempotence-based systems. This paper presents Picker, the first launch-time analysis system for instance-level idempotency validation. Picker accurately validates the idempotency of GPU kernel instances before execution by utilizing launch arguments. Several optimizations are proposed to reduce validation latency to microseconds. Evaluations using representative GPU applications (547 kernels and 18,217 instances) show that Picker accurately identifies idempotent instances with zero false positives and an 18.54% false-negative rate. The launch-time validation completes in under 5 μs for all instances (about 90% under 1 μs). Through integration, Picker reduces checkpoint costs to less than 4% in fault-tolerant systems and decreases preemption latency by 84.2% in scheduling systems. Mingcong Han, Weihang Shen, Rong Chen 0001, Haibo Chen 0001 |
EuroSys | 1 |
| 2026 | DistRS: Disaggregated Reward Service for RLVR with Batch-Level Constraint
Ruidong Zhu, Mingcong Han, Yinmin Zhong, Wencong Xiao, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 2 |
| 2026 | Real-time, Work-conserving GPU Scheduling for Concurrent DNN InferenceabstractMany intelligent applications, such as autonomous driving and virtual reality, require running both latency-critical (real-time) and best-effort deep neural network (DNN) inference tasks to achieve both real-time and work-conserving on the GPU. However, commodity GPUs lack efficient preemptive scheduling support, and existing state-of-the-art approaches either have to monopolize GPU or let real-time tasks to wait for best-effort tasks to complete, resulting in low utilization, high latency, or both. This article presents Reef , the first GPU-accelerated DNN inference serving system that achieves low-latency and work-conserving for concurrent real-time and best-effort tasks. Reef accomplishes this by enabling microsecond-scale kernel preemption and controlled concurrent execution in GPU scheduling. Reef is novel in two ways. First, based on the observation that DNN inference kernels are mostly idempotent, Reef devises a reset-based preemption scheme that launches a real-time kernel on the GPU by proactively killing and restoring best-effort kernels at microsecond-scale. Second, since DNN inference kernels have varied parallelism and predictable latency, Reef proposes a dynamic kernel padding mechanism that dynamically pads the real-time kernel with appropriate best-effort kernels to fully utilize the GPU with negligible overhead. Evaluation using a new DNN inference serving benchmark (DISB) with diverse workloads and a real-world trace on both NVIDIA and AMD GPUs shows that Reef only incurs less than 5% overhead in end-to-end latency for real-time tasks but increases the overall throughput by up to 1.53×, compared to scheduling tasks sequentially. To demonstrate the practical benefits of our approach, we compare Reef with Triton, a widely-adopted production-level serving system. Our evaluation shows that Reef outperforms Triton by 1.12× to 5.20× in end-to-end latency for real-time tasks, while maintaining comparable throughput. Mingcong Han, Rong Chen 0001, Weihang Shen, Hanze Zhang, Haibo Chen 0001 |
ACM Trans. Comput. Syst. | 1 |
| 2025 | XSched: Preemptive Scheduling for Diverse XPUs
Weihang Shen, Mingcong Han, Jialong Liu, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 2 |
| 2025 | PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated SpeculationabstractPhoenixOS (PhOS) is the first OS service that can concurrently checkpoint and restore (C/R) GPU processes—a fundamental capability for critical tasks such as fault tolerance, process migration, and fast startup. While concurrent C/R is well-established on CPUs, it poses unique challenges on GPUs due to their lack of essential features for efficiently tracing concurrent memory reads and writes, such as specific hardware capabilities (e.g., dirty bits) and OS-mediated data paths (e.g., copy-on-write). Xingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao, Rong Chen 0001, Mingcong Han, Jinyu Gu 0001, Haibo Chen 0001 |
SOSP | 6 |
| 2025 | Colocating ML Inference and Training with Fast GPU Memory Handover
Yankui Wang, Mingcong Han, Rong Chen 0001 |
USENIX ATC | 3 |
| 2023 | DArray: A High Performance RDMA-Based Distributed ArrayabstractThis paper presents DArray, a high performance RDMA-based distributed memory system. DArray achieves high performance through three key designs. First, DArray is designed with an object array abstraction, which captures the high-level application semantics and provides a rich set of optimized interfaces with object granularity. Second, DArray adopts distributed cache to absorb remote data accesses. In order to reduce the performance overhead incurred by the cache layer and increase the parallelism, DArray devises a lock-free data access path to the local cache which utilizes reference counters to prevent data races. Finally, based on the observation that most data update operators are associative and commutative, DArray proposes a new "Operate" interface, which enables concurrent data operations on multiple nodes, and extends existing distributed cache coherence protocol to support the new "Operate" semantics. Baorong Ding, Mingcong Han, Rong Chen 0001 |
ICPP | 2 |
| 2022 | Microsecond-scale Preemption for Concurrent GPU-accelerated DNN Inferences
Mingcong Han, Hanze Zhang, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 1 |
| 2021 | ShadowVM: accelerating data plane for data analytics with bare metal CPUs and GPUsabstractWith the development of the big data ecosystem, large-scale data analytics has become more prevalent in the past few years. Apache Spark, etc., provide a flexible approach for scalable processing upon massive data. However, they are not designed for handling computing-intensive workloads due to the restrictions of JVM runtime. In contrast, GPU has been the de facto accelerator for graphics rendering and deep learning in recent years. Nevertheless, the current architecture makes it difficult to take advantage of GPUs and other accelerators in the big data world. Mingcong Han, Shangwei Wu 0001, Chuliang Weng |
PPoPP | 2 |