Hanze Zhang

dblp:330/8813 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0009-0579-5707ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized Coordination
Hanze Zhang, Rong Chen 0001, Xingda Wei, Haibo Chen 0001
Proc. VLDB Endow.1
2026 Real-time, Work-conserving GPU Scheduling for Concurrent DNN Inference
abstract
Many intelligent applications, such as autonomous driving and virtual reality, require running both latency-critical (real-time) and best-effort deep neural network (DNN) inference tasks to achieve both real-time and work-conserving on the GPU. However, commodity GPUs lack efficient preemptive scheduling support, and existing state-of-the-art approaches either have to monopolize GPU or let real-time tasks to wait for best-effort tasks to complete, resulting in low utilization, high latency, or both. This article presents Reef , the first GPU-accelerated DNN inference serving system that achieves low-latency and work-conserving for concurrent real-time and best-effort tasks. Reef accomplishes this by enabling microsecond-scale kernel preemption and controlled concurrent execution in GPU scheduling. Reef is novel in two ways. First, based on the observation that DNN inference kernels are mostly idempotent, Reef devises a reset-based preemption scheme that launches a real-time kernel on the GPU by proactively killing and restoring best-effort kernels at microsecond-scale. Second, since DNN inference kernels have varied parallelism and predictable latency, Reef proposes a dynamic kernel padding mechanism that dynamically pads the real-time kernel with appropriate best-effort kernels to fully utilize the GPU with negligible overhead. Evaluation using a new DNN inference serving benchmark (DISB) with diverse workloads and a real-world trace on both NVIDIA and AMD GPUs shows that Reef only incurs less than 5% overhead in end-to-end latency for real-time tasks but increases the overall throughput by up to 1.53×, compared to scheduling tasks sequentially. To demonstrate the practical benefits of our approach, we compare Reef with Triton, a widely-adopted production-level serving system. Our evaluation shows that Reef outperforms Triton by 1.12× to 5.20× in end-to-end latency for real-time tasks, while maintaining comparable throughput.
Mingcong Han, Rong Chen 0001, Weihang Shen, Hanze Zhang, Haibo Chen 0001
ACM Trans. Comput. Syst.4
2026 Accelerating Million-scale In-network Lock Management using Lock Fission
abstract
Distributed lock services are extensively utilized in distributed systems to serialize concurrent accesses to shared resources. The need for fast and scalable lock services has become more pronounced with decreasing task execution times and expanding dataset scales. However, traditional lock managers, reliant on server CPUs to handle lock requests, experience significant queuing delays in lock grant latency. Advanced network hardware (e.g., programmable switches) presents an avenue to manage locks without queuing delays due to their high packet processing power. Nevertheless, their constrained memory capacity restricts the number of locks they can manage, thereby limiting their efficiency and efficacy in large-scale workloads with millions of locks. This article introduces the concept of lock fission, which enables efficient management of million-scale locks by exploiting both programmable switches and servers. Lock fission decouples lock management into a memory-efficient grant decision process and a latency-insensitive participant maintenance process. This allows the programmable switch to efficiently make grant decisions for numerous locks, while servers asynchronously maintain participants (i.e., holders and waiters). Furthermore, by using the programmable switch for routing, lock fission supports on-demand, fine-grained lock migration, reducing network traffic and lock release delays. Building on this idea, we present FissLock , a fast and scalable in-network lock service for two representative lock management settings: with lock managers on dedicated servers or colocated with applications. Evaluation using various benchmarks and a real-world application shows FissLock ’s efficiency and efficacy. Compared to the state-of-the-art in-network LM, FissLock cuts up to 82.9% (from 44.3%) of median lock grant time in the microbenchmark and improves transaction throughput for TATP and TPC-C by 2.26× and 2.46×.
Hanze Zhang, Rong Chen 0001, Haibo Chen 0001
ACM Trans. Comput. Syst.1
2025 Joint segmentation of retinal layers and fluid lesions in optical coherence tomography with cross-dataset learning
Xiayu Xu, Hualin Wang, Yulei Lu, Hanze Zhang, Tao Tan 0002, Jianqin Lei
Artif. Intell. Medicine4
2024 Fast and Scalable In-network Lock Management Using Lock Fission
Hanze Zhang, Rong Chen 0001, Haibo Chen 0001
OSDI1
2022 Microsecond-scale Preemption for Concurrent GPU-accelerated DNN Inferences
Mingcong Han, Hanze Zhang, Rong Chen 0001, Haibo Chen 0001
OSDI2