VLDB 2026 Research / reviewers in the wild / expert
Zhongming Yu
dblp:302/9573
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Cross-model Fusion-aware Framework for Optimizing (gather-matmul-scatter)s WorkloadabstractModern deep learning models, such as Relation Graph Convolutional Network (RGCN), Sparse Convolutional Networks (SpConv), and Mixture of Experts Networks (MoE), are significantly dependent on the (gather-matmul-scatter) (abbreviated as (g-mm-s) ${ }_{\mathrm{s}}$) workload as their fundamental computational pattern. While existing works have made optimization attempts, several critical challenges remain unsolved, including domain-specific optimization migration, time-consuming exploration, and inefficient dataflow with dynamic inputs.To address these challenges, we introduce Efficient-GMS, a comprehensive framework that enhances ($\mathrm{g}-\mathrm{mm}-\mathrm{s})_{\text {s }}$ workload across diverse input scenarios. Our framework introduces (1) A Fusion-aware framework enabling cross-model optimization migration. We propose a comprehensive dataflow analysis that identifies shared computational patterns across models, enabling the development of four optimized dataflow patterns with vertical and horizontal fusion strategies. (2) Performance model-guided configuration space reduction. We develop a performance model to predict the relative execution efficiency across configurations, thereby reducing the search space and minimizing search time while ensuring optimal configuration selection. (3) Adaptive dataflow selection mechanism. We implement a lightweight heuristic model that dynamically selects optimal dataflow patterns based on the characteristics of the input and the hardware. Experimental results demonstrate that Efficient-GMS achieves significant performance gains, delivering an average end-to-end speedup of $1.46 \times$ in RGCN model, $1.32 \times$ in Sp-Conv-based model, and $1.15 \times$ in MoE model compared to state-of-the-art methods. Yaoxiu Lian, Zhihong Gou, Yibo Han, Zhongming Yu, Sheng Yuan, Zhilin Pei, Xingcheng Zhang, Ningyi Xu, Guohao Dai 0001 |
DAC | 4 |
| 2025 | MAGE: A Multi-Agent Engine for Automated RTL Code GenerationabstractThe automatic generation of RTL code (e.g., Verilog) through natural language instructions has emerged as a promising direction with the advancement of large language models (LLMs). However, producing RTL code that is both syntactically and functionally correct remains a significant challenge. Existing single-LLM-agent approaches face substantial limitations because they must navigate between various programming languages and handle intricate generation, verification, and modification tasks. To address these challenges, this paper introduces MAGE, the first open-source1multi-agent AI system designed for robust and accurate Verilog RTL code generation. We propose a novel high-temperature RTL candidate sampling and debugging system that effectively explores the space of code candidates and significantly improves the quality of the candidates. Furthermore, we design a novel Verilog-state checkpoint checking mechanism that enables early detection of functional errors and delivers precise feedback for targeted fixes, significantly enhancing the functional correctness of the generated RTL code. MAGE achieves a 95.7% rate of syntactic and functional correctness code generation on VerilogEval-Human v2 benchmark, surpassing the state-of-the-art Claude-3.5-sonnet by 23.3%, demonstrating a robust and reliable approach for AIdriven RTL design workflows.1MAGE is open-sourced at https://github.com/stable-lab/ MAGE-A-Multi-Agent-Engine-for-Automated-RTL-Code-Generation Hejia Zhang 0001, Hanxian Huang, Zhongming Yu, Jishen Zhao |
DAC | 4 |
| 2025 | OrcaLoca: An LLM Agent Framework for Software Issue LocalizationabstractRecent developments in Large Language Model (LLM) agents are revolutionizing Autonomous Software Engineering (ASE), enabling automated coding, problem fixes, and feature improvements. However, localization -- precisely identifying software problems by navigating to relevant code sections -- remains a significant challenge. Current approaches often yield suboptimal results due to a lack of effective integration between LLM agents and precise code search mechanisms. This paper introduces OrcaLoca, an LLM agent framework that improves accuracy for software issue localization by integrating priority-based scheduling for LLM-guided action, action decomposition with relevance scoring, and distance-aware context pruning. Experimental results demonstrate that OrcaLoca becomes the new open-source state-of-the-art (SOTA) in function match rate (65.33%) on SWE-bench Lite. It also improves the final resolved rate of an open-source framework by 6.33 percentage points through its patch generation integration. Zhongming Yu, Hejia Zhang 0001, Hanxian Huang, Matrix Yao, Jishen Zhao |
ICML | 1 |
| 2023 | CLAP: Locality Aware and Parallel Triangle Counting with Content Addressable MemoryabstractTriangle counting (TC) is one of the most fundamental graph analysis tools with a wide range of applications. Modern triangle counting algorithms traverse the graph and perform set intersections of neighbor sets to find triangles. However, existing triangle counting approaches suffer from the heavy off-chip memory access and set intersection overhead. Thus, we propose CLAP, the first content addressable memory (CAM) based triangle counting architecture with the software and hardware co-optimizations. To reduce off-chip memory access and the number of set intersections, we propose the first force-based node index reorder method. It simultaneously optimizes both data locality and the computation amount. Compared with random node indices, the reorder method reduces the off-chip memory access and the set intersections by 61% and 64%, respectively, while providing$\mathbf{2.19}\times$end-to-end speedup. To improve the set intersection parallelism, we propose the first CAM-based triangle counting architecture under chip area constraints. We enable the high parallel set intersection by translating it into content search on CAM with full parallelism. Thus, the time complexity of the set intersection reduces from$O(m+n)$or$O(n\log m)$to$O(n)$. Extensive experiments on real-world graphs show that CLAP achieves$\mathbf{39}\times, \mathbf{27}\times$, and$\mathbf{78}\times$speedup over state-of-the-art CPU, GPU, and processing-in-memory baselines, respectively. The software code is available at: https://github.com/thu-nics/CLAP-triangle-counting Tianyu Fu 0004, Chiyue Wei, Zhenhua Zhu 0002, Shang Yang, Zhongming Yu, Guohao Dai 0001, Huazhong Yang, Yu Wang 0002 |
DATE | 5 |
| 2023 | TorchSparse++: Efficient Training and Inference Framework for Sparse Convolution on GPUsabstractSparse convolution plays a pivotal role in emerging workloads, including point cloud processing in AR/VR, autonomous driving, and graph understanding in recommendation systems. Since the computation pattern is sparse and irregular, specialized high-performance kernels are required. Existing GPU libraries offer two dataflow types for sparse convolution. The gather-GEMM-scatter dataflow is easy to implement but not optimal in performance, while the dataflows with overlapped computation and memory access (e.g. implicit GEMM) are highly performant but have very high engineering costs. In this paper, we introduce TorchSparse++, a new GPU library that achieves the best of both worlds. We create a highly efficient Sparse Kernel Generator that generates performant sparse convolution kernels at less than one-tenth of the engineering cost of the current state-of-the-art system. On top of this, we design the Sparse Autotuner, which extends the design space of existing sparse convolution libraries and searches for the best dataflow configurations for training and inference workloads. Consequently, TorchSparse++ achieves 2.9 × , 3.3 × , 2.2 × and 1.7 × measured end-to-end speedup on an NVIDIA A100 GPU over state-of-the-art MinkowskiEngine, SpConv 1.2, TorchSparse and SpConv v2 in inference; and is 1.2-1.3 × faster than SpConv v2 in mixed precision training across seven representative autonomous driving benchmarks. It also seamlessly supports graph convolutions, achieving 2.6-7.6 × faster inference speed compared with state-of-the-art graph deep learning libraries. Our code is publicly released at https://github.com/mit-han-lab/torchsparse. Haotian Tang, Shang Yang, Ke Hong, Zhongming Yu, Xiuyu Li, Guohao Dai 0001, Yu Wang 0002, Song Han 0003 |
MICRO | 5 |
| 2023 | CogDL: A Comprehensive Library for Graph Deep LearningabstractGraph neural networks (GNNs) have attracted tremendous attention from the graph learning community in recent years. It has been widely adopted in various real-world applications from diverse domains, such as social networks and biological graphs. The research and applications of graph deep learning present new challenges, including the sparse nature of graph data, complicated training of GNNs, and non-standard evaluation of graph tasks. To tackle the issues, we present CogDL1, a comprehensive library for graph deep learning that allows researchers and practitioners to conduct experiments, compare methods, and build applications with ease and efficiency. In CogDL, we propose a unified design for the training and evaluation of GNN models for various graph tasks, making it unique among existing graph learning libraries. By utilizing this unified trainer, CogDL can optimize the GNN training loop with several training techniques, such as mixed precision training. Moreover, we develop efficient sparse operators for CogDL, enabling it to become the most competitive graph library for efficiency. Another important CogDL feature is its focus on ease of use with the aim of facilitating open and reproducible research of graph learning. We leverage CogDL to report and maintain benchmark results on fundamental graph tasks, which can be reproduced and directly used by the community. Yukuo Cen, Yan Wang 0120, Yizhen Luo, Zhongming Yu, Xingcheng Yao, Aohan Zeng, Shiguang Guo, Yuxiao Dong, Yang Yang 0009, Peng Zhang 0077, Guohao Dai 0001, Yu Wang 0002, Chang Zhou 0005, Hongxia Yang, Jie Tang 0001 |
WWW | 6 |
| 2023 | Sgap: towards efficient sparse tensor algebra compilation for GPU
Genghan Zhang, Yuetong Zhao, Yanting Tao, Zhongming Yu, Guohao Dai 0001, Sitao Huang, Yuan Wen, Pavlos Petoumenos, Yu Wang 0002 |
CCF Trans. High Perform. Comput. | 4 |
| 2022 | Heuristic adaptability to input dynamics for SpMM on CPUsabstractSparse Matrix-Matrix Multiplication (SpMM) has served as fundamental components in various domains. Many previous studies exploit GPUs for SpMM acceleration because GPUs provide high bandwidth and parallelism. We point out that a static design does not always improve the performance of SpMM on different input data (e.g., >85% performance loss with a single algorithm). In this paper, we consider the challenge of input dynamics from a novel auto-tuning perspective, while following issues remain to be solved: (1) Orthogonal design principles considering sparsity. Orthogonal design principles for such a sparse problem should be extracted to form different algorithms, and further used for performance tuning. (2) Nontrivial implementations in the algorithm space. Combining orthogonal design principles to create new algorithms needs to tackle with new challenges like thread race handling. (3) Heuristic adaptability to input dynamics. The heuristic adaptability is required to dynamically optimize code for input dynamics. Guohao Dai 0001, Guyue Huang, Shang Yang, Zhongming Yu, Yufei Ding 0001, Yuan Xie 0001, Huazhong Yang, Yu Wang 0002 |
DAC | 4 |
| 2022 | Decentralized Time-Delay Control Using Partial Variables With Measurable States for a Class of Interconnected Systems With Time DelaysabstractThis article deals with the problems of stability and control of the interconnected system (IS) with unknown time-varying delays via decentralized time-delay control using partial variables with measurable states. First, the model of the IS with time delays is established, and the relevant control scheme is proposed. The control scheme just needs to control all or partial state variables corresponding to the elements on the main diagonal of the gain matrices, which can reduce the control cost and improve the flexibility of control. In addition, there are no additional restrictions in the process of designing the controller. Second, relevant lemmas are derived. The exponential boundedness and stability analysis of the IS with time delays are presented, respectively, by stability theory, and related results are derived. Meanwhile, the stability domain of the IS is estimated. Besides, the obtained results can also be used for many practical systems, such as the interconnected power system, the multislave teleoperation systems, the brushless dc motor (BLDCM) system, and the chaotic system. Finally, the effectiveness and application of the obtained results are verified by several examples. Zhongming Yu, Xin Dai 0009, Xiaojie Su |
IEEE Trans. Cybern. | 1 |
| 2021 | Exploiting Online Locality and Reduction Parallelism for Sampled Dense Matrix Multiplication on GPUsabstractSampled Dense-Dense Matrix Multiplication (SDDMM) is a core component of many machine learning systems. SDDMM exposes a substantial amount of parallelism that favors throughput-oriented architectures like the GPU. However, accelerating it on GPUs is challenging in two aspects: the poor memory access locality caused by the sparse sampling matrix with the poor parallelism caused by the dot-product reduction of vectors in two dense matrices. To address both challenges, we present PRedS to boost SDDMM efficiency with a suite of Parallel Reduction Scheduling optimizations. PRedS uses Vectorized Coarsen 1-Dimensional Tiling (VCT) to benefit the online locality of loading the dense matrix. PRedS uses Integrated Interleaving Reduction (IIR) to increase thread occupancy in the parallel reduction. PRedS also leverages Warp-Merged Tiling (WMT) to preserve occupancy and parallelism when reducing very long arrays. Enhanced with GPU-intrinsic vectorized memory loading, PRedS achieves a geometric speedup of 29.20× compared to the vendor library. PRedS achieves up to 8.31× speedup over state-of-the-art implementations on the SuiteSparse benchmark. Zhongming Yu, Guohao Dai 0001, Guyue Huang, Yu Wang 0002, Huazhong Yang |
ICCD | 1 |