VLDB 2026 Research / reviewers in the wild / expert
Zhengding Hu
dblp:359/5899
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0005-8500-6173ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM InferenceabstractLarge-scale Mixture of Experts (MoE) Large Language Models (LLMs) have recently become the frontier open-weight models, achieving remarkable model capability similar to proprietary ones. But their random expert selection mechanism introduces significant data movement overhead that becomes the dominant bottleneck in multi-unit LLM serving systems. To understand the patterns underlying this data movement, we conduct comprehensive data-movement-centric profiling across four state-of-the-art large-scale MoE models released in 2025 (200B-1000B) using over 24,000 requests spanning diverse workloads. We perform systematic analysis from both temporal and spatial perspectives and distill six key insights to guide the design of diverse serving systems. We verify these insights on both future wafer-scale GPU architectures and existing GPU systems. On wafer-scale GPUs, lightweight architectural modifications guided by our insights yield a 6.6$\times$ average speedup across four 200B--1000B models. On existing GPU systems, our insights drive the design of a prefill-aware expert placement algorithm that achieves up to 1.25$\times$ speedup on MoE computation. Our work presents the first comprehensive data-centric analysis of large-scale MoE models together with a concrete design study applying the learned lessons. Our profiling traces are publicly available at \href{https://huggingface.co/datasets/core12345/MoE_expert_selection_trace}{\textcolor{blue}{https://huggingface.co/datasets/core12345/MoE\_expert\_selection\_trace}}. Zhongkai Yu, Yue Guan 0003, Zhengding Hu, Shuyi Pei, Yangwook Kang, Yufei Ding 0001, Po-An Tsai |
ISCA | 5 |
| 2025 | A Fast Sparse Triangular Solve for Structured-grid Problems on Heterogeneous ProcessorsabstractStructured-grid problems are common in scientific computing, particularly in applications like fluid dynamics and electromagnetic simulation. One of the key kernels in solving these problems is Sparse Triangular Solve (SpTRSV), which often becomes a performance bottleneck due to its low computing intensity and inherent internal data dependencies. In structured-grid SpTRSV, the regularity of non-zero distributions and the high parallelism of sparse matrices present opportunities to harness the architectural strengths of modern heterogeneous processors. However, existing SpTRSV algorithms fail to fully exploit these advantages, due to their mismatches in data dependencies, computational order, and memory layouts. In this paper, we introduce a novel SpTRSV algorithm tailored for structured-grids on modern heterogeneous processors. Our approach introduces a two-level blocking strategy to enhance data locality and reduce communication overhead, while a vertical tiling-based pipeline balances parallelism with computational granularity. Additionally, we design hardware-specific adaptive scheduling strategies to accommodate varying degrees of parallelism across distinct architectures. The algorithm has been implemented on two types of heterogeneous processors, NVIDIA GPUs and SW26010-Pro, with hardware-specific optimizations to further improve the performance. Experimental results show that our implementations achieve speedups of more than 1.87 × over state-of-the-art baselines and provide efficient end-to-end solutions with lightweight preprocessing. Zhengding Hu, Yi Zong, Jingwei Sun 0001, Wei Xue 0003, Guangzhong Sun |
ICPP | 1 |
| 2025 | KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsabstractLarge language model (LLM) based agentic workflows have become a popular paradigm for coordinating multiple specialized agents to solve complex tasks. To improve serving efficiency, existing LLM systems employ prefix caching to reuse key-value (KV) tensors corresponding to agents' fixed prompts, thereby avoiding redundant computation across repeated invocations. However, current systems typically evict KV caches using a Least Recently Used (LRU) policy, which fails to anticipate future agent usage and often discards KV caches shortly before their reuse. This leads to frequent cache misses and substantial recomputation or swap- ping overhead. We present KVFlow, a workflow-aware KV cache management framework tailored for agentic workloads. KVFlow abstracts the agent execution schedule as an Agent Step Graph and assigns each agent a steps-to-execution value that estimates its temporal proximity to future activation. These values guide a fine-grained eviction policy at the KV node level, allowing KVFlow to preserve entries likely to be reused and efficiently manage shared prefixes in tree-structured caches. Moreover, KVFlow introduces a fully overlapped KV prefetching mecha- nism, which proactively loads required tensors from CPU to GPU in background threads for agents scheduled in the next step, thereby avoiding cache miss stalls during generation. Compared to SGLang with hierarchical radix cache, KVFlow achieves up to 1.83× speedup for single workflows with large prompts, and up to 2.19× speedup for scenarios with many concurrent workflows. Zaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu, Yue Guan 0003, Wan-Lu Li, Lianhui Qin, Yufei Ding 0001 |
NeurIPS | 4 |
| 2025 | HedraRAG: Co-Optimizing Generation and Retrieval for Heterogeneous RAG Workflows
Zhengding Hu, Vibha Murthy, Zaifeng Pan, Xiaoyi Fang, Yufei Ding 0001 |
SOSP | 1 |
| 2025 | GNNPilot: A Holistic Framework for High-Performance Graph Neural Network Computations on GPUsabstractGraph Neural Networks (GNNs) have emerged as powerful tools for graph-based machine learning tasks, but their performance is often constrained by inefficient sparse operators and limited hardware utilization during multi-operator workflows. This article presents GNNPilot, a holistic optimization framework that addresses these challenges through three key innovations. First, we introduce two packing strategies for gather operators, including neighbor packing for load balancing in sparser graphs, and bin packing with a new sparse format for enhanced data locality in denser graphs. Second, we propose dynamic parallelization methods and a novel row panel-based kernel fusion technique to optimize complex multi-operator GNN models. Third, we develop a lightweight sampling-based auto-tuning mechanism that adapts the framework’s optimization strategies to varying input characteristics. Built upon tensor expression-based intermediate representations, GNNPilot maintains the flexibility to optimize both popular and customized GNN models. Extensive experiments across diverse GNN models and graph datasets demonstrate that GNNPilot achieves substantial speedups over state-of-the-art implementations in both the performance of single operators and the efficiency of end-to-end inference. These results establish GNNPilot as an efficient and adaptive solution for accelerating GNN computations on modern GPU architectures. Zhengding Hu, Jingwei Sun 0001, Guangzhong Sun |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | PckGNN: Optimizing Aggregation Operators with Packing Strategies in Graph Neural NetworksabstractGraph Neural Network (GNN) is one of the most prominent machine learning models. It involves a substantial amount of graph-based aggregation operators, which can be abstracted as sparse matrix computation kernels. Due to the irregularity of the graph adjacency matrix, the aggregation has long been the performance bottleneck of GNN. According to our measurements, existing GNN implementations fall short of achieving optimal performance due to their insufficient consideration of load balancing and data locality. To bridge these performance gaps, we propose PckGNN, which aims to accelerate GNN aggregation operators on GPUs with packing strategies. PckGNN categorizes graph matrices into two types based on different sparsity levels and conducts two packing strategies respectively. For sparser matrices, Neighbor Packing enhances load balancing through a moderate-grained non-zero grouping approach. For denser matrices, Bin Packing exposes more potentials of cache data reuse by bin partitioning, non-zero extracting, format converting and two-level scheduling. Experimental results on SpMM and SDDMM show that PckGNN achieves speedups of 1.46x ∼ 6.14x over existing implementations. When applied GNN inference of three typical models, it achieves speedups of more than 1.29x over the state-of-the-art frameworks. Zhengding Hu, Jingwei Sun 0001, Guangzhong Sun |
IPDPS | 1 |
| 2024 | AG-SpTRSV: An Automatic Framework to Optimize Sparse Triangular Solve on GPUsabstractSparse Triangular Solve (SpTRSV) has long been an essential kernel in the field of scientific computing. Due to its low computational intensity and internal data dependencies, SpTRSV is hard to implement and optimize on graphics processing units (GPUs). Based on our experimental observations, existing implementations on GPUs fail to achieve the optimal performance due to their suboptimal parallelism setups and code implementations plus lack of consideration of the irregular data distribution. Moreover, their algorithm design lacks the adaptability to different input matrices, which may involve substantial manual efforts of algorithm redesigning and parameter tuning for performance consistency. In this work, we propose AG-SpTRSV, an automatic framework to optimize SpTRSV on GPUs, which provides high performance on various matrices while eliminating the costs of manual design. AG-SpTRSV abstracts the procedures of optimizing an SpTRSV kernel as a scheme and constructs a comprehensive optimization space based on it. By defining a unified code template and preparing code variants, AG-SpTRSV enables fine-grained dynamic parallelism and adaptive code optimizations to handle various tasks. Through computation graph transformation and multi-hierarchy heuristic scheduling, AG-SpTRSV generates schemes for task partitioning and mapping, which effectively address the issues of irregular data distribution and internal data dependencies. AG-SpTRSV searches for the best scheme to optimize the target kernel for the specific matrix. A learned lightweight performance model is also introduced to reduce search costs and provide an efficient end-to-end solution. Experimental results with SuiteSparse Matrix Collection on NVIDIA Tesla A100 and RTX 3080 Ti show that AG-SpTRSV outperforms state-of-the-art implementations with geometric average speedups of 2.12x ∼ 3.99x. With the performance model enabled, AG-SpTRSV can provide an efficient end-to-end solution, with preprocessing times ranging from 3.4 to 245 times of the execution time. Zhengding Hu, Jingwei Sun 0001, Guangzhong Sun |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | Toward efficient structured-grid triangular solver on sunway many-core processors
Jianjiang Li, Jiabi Liang, Wei Xue 0003, Zhengding Hu, Jinliang Shi |
J. Supercomput. | 4 |
| 2023 | Rapid simulations of atmospheric data assimilation of hourly-scale phenomena with modern neural networksabstractAtmospheric data assimilation is essential for numerical weather prediction. Ensemble data assimilation connects multiple instances of an atmospheric model through a Kalman filter-based algorithm, which is regarded as a challenging computing task today. In this work, we build a fast, low-cost, and scalable atmospheric data assimilation prototype, DIDA, for the new-generation Sunway supercomputer, including: (1) a framework that enables flexible deployment of components, and manages and optimizes data communication among modules, achieving maximum resource efficiency; (2) an accurate, robust, UNet-based surrogate model for atmospheric dynamic simulation to generate the background ensemble; (3) a batch-LETKF algorithm with high-performance eigenvalue decomposition, which is up to 7.37 times faster than existing numerical libraries while exhibiting almost linear scalability. Experimental evaluations show that our AI-integrated ensemble data assimilation prototype can complete hour-cycle assimilation in minutes, maintain linear scalability, and save an order of magnitude of computing resources, compared with the traditional method. Yiyuan Li, Xiting Ju, Qilong Jia, Yongxiao Zhou, Simeng Qian, Rongfen Lin, Bin Yang 0043, Shupeng Shi, Xin Liu 0081, Jian Tan 0005, Zhengding Hu, Limin Yan, Wei Xue 0003 |
SC | 16 |