Yangjie Zhou 0001

dblp:155/7332-1 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-3652-5437ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 CuBridge: An LLM-Based Framework for Understanding and Reconstructing High-Performance Attention Kernels
abstract
Xing Ma, Yangjie Zhou, Wu Sun, Zihan Liu, Jingwen Leng, Yun Lin, Shixuan Sun, Minyi Guo, Jin Song Dong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yangjie Zhou 0001, Wu Sun, Zihan Liu 0002, Jingwen Leng, Yun Lin 0001, Shixuan Sun, Minyi Guo, Jin Song Dong 0001
ACL (1)2
2026 Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
Yukang Chen, Weihao Cui, Han Zhao 0005, Xiaoze Fan, Xusheng Chen, Yangjie Zhou 0001, Shixuan Sun, Bingsheng He, Quan Chen 0002
ASPLOS (2)7
2026 FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
abstract
The scaling of computation throughput continues to outpace improvements in memory bandwidth, making many deep learning workloads memory-bound. Kernel fusion is a key technique to alleviate this problem, but the fusion strategies of existing compilers and frameworks are limited to using local scratchpad memory. When the intermediate results exceed the limited capacity (such as FFN), the fusion fails. Although modern GPUs (like the NVIDIA H100) now incorporate an inter-core connection mechanism known as Distributed Shared Memory (DSM)—providing a larger, high-bandwidth, and low-latency on-chip memory pool—this hardware potential has yet to be exploited by software frameworks. To bridge this gap, we present FlashFuser, the first compiler framework to utilize inter-core connection for kernel fusion on modern GPUs. FlashFuser extends established fusion techniques to the DSM domain through three core contributions. First, we propose a powerful DSM-based communication abstraction that formalizes complex cluster-based data exchange patterns, such as reduce, shuffle and multiply. Second, we introduce a dataflow analyzer that generalizes loop scheduling, resource mapping, and tile selection to the distributed memory hierarchy; it determines the optimal execution order and tile sizes by quantifying data movement across memory levels. Finally, FlashFuser integrates these components into a unified search engine that employs analytical cost modeling and DSM-aware pruning strategies to efficiently discover the optimal execution plan. Our evaluation on an NVIDIA H100 GPU shows that FlashFuser reduces memory access by 58 % and delivers kernel speedups of$3.3 x$against highly-tuned libraries and 4.1x against state-of-the-art compilers, resulting in a$1.24 \times$end-to-end speedup.
Yangjie Zhou 0001, Zihan Liu 0002, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng 0007, Chen Zhang 0001, Anbang Wu, Jingwen Leng
HPCA2
2026 A Full-Stack Framework for GNN Acceleration via Partition-Compiler-Architecture Co-Design
abstract
Graph Neural Networks (GNNs) have achieved remarkable success across domains such as recommendation and scientific computing, yet their practical deployment remains constrained by high execution cost. The diversity of GNN model structures and the sparsity of real-world graphs pose two fundamental challenges for hardware acceleration: supporting heterogeneous operator patterns and achieving high resource utilization under irregular data access. Existing accelerators often address only one aspect, either targeting specific models with hardwired pipelines or applying general architectures with limited efficiency. To address these challenges, we propose SWITCHBLADE, a full-stack framework for GNN acceleration through the coordinated design of partitioning, compilation, and architecture. SWITCHBLADE addresses these challenges through three key components. First, a phase-based intermediate representation unifies diverse GNN models by abstracting computation stages for model-independent code generation. Second, a fine-grained graph partitioner enhances data locality and reduces memory traffic by adapting to graph topology and model semantics. Third, the hardware architecture supports stream-level parallelism and decoupled execution to exploit cross-shard and inter-phase concurrency. Evaluation on representative models and datasets shows that SWITCHBLADE achieves up to 1.85× speedup and 19.03× energy savings over an NVIDIA V100 GPU, while outperforming state-of-the-art GNN accelerators across diverse full-graph workloads, demonstrating both high efficiency and broad model generality.
Yangjie Zhou 0001, Shuwen Lu, Cong Guo 0003, Jingwen Leng, Yufei Ma 0002, Yun Liang 0001, Minyi Guo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Voyager: Input-Adaptive Algebraic Transformations for High-Performance Graph Neural Networks
abstract
Graph neural networks (GNNs) are gaining popularity in diverse application domains and growing in complexity.As a result, it is crucial to achieve high-performance GNN execution.Among various techniques, algebraic transformations, including operator reordering and operator fusion, have been successfully applied to improve the computation and memory access efficiencies of DNN models.However,
Yangjie Zhou 0001, Wenting Shen, Jingwen Leng, Shuwen Lu, Zihan Liu 0002, Weihao Cui, Zhendong Zhang 0004, Wencong Xiao, Baole Ai, Yong Li 0045, Wei Lin 0016, Deze Zeng, Yun Liang 0001, Quan Chen 0001, Ning Liu 0007, Minyi Guo
ASPLOS (3)1
2025 VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
abstract
Vector quantization (VQ), which treats a vector as a compression unit, gains increasing research interests for its potential to accelerate large language models (LLMs). Compared to conventional element-wise quantization methods, VQ algorithms can compress weight and KV cache tensors in LLMs with a greater ratio while maintaining the high model accuracy. However, translating a VQ algorithm’s memory reduction into the actual latency improvement is challenging. We profile and analyze the current approach of integrating VQ into computation kernels and show that its major inefficiency lies in the poor access efficiency of codebooks in VQ algorithms and uncoordinated computation dataflow. Meanwhile, the diversity of VQ algorithms (e.g., different vector sizes and entry counts) and LLMs, computation kernels (e.g matrix-matrix/vector multiplication and attention computation) makes it impractical to manually craft efficient kernel implementations for each specific case. In this work, we design and implement VQ-LLM, an efficient fused VQ kernel generation framework. We first introduce a software abstraction called codebook cache to optimize codebook access efficiency and support the integration of VQ with various computations. The codebook cache adaptively stores different entries across the GPU’s memory hierarchy, including off-chip global memory, on-chip shared memory, and registers. Centered around the codebook cache, we design an efficient computation engine that optimizes memory traffic during computations involving codebooks. This compute engine adopts the codebook-centric dataflow and fusion optimizations. Additionally, we provide adaptive heuristics to tailor parameter selection in our optimizations to diverse VQ configurations. Our optimizations achieve the latency reduction of $\mathbf{6 4. 3 6 \%}$ to $\mathbf{9 9. 1 \%}$ compared to existing open-source implementations. A final comparison with state-of-the-art element-wise quantization methods like AWQ and QoQ shows that our VQ-LLM is practically viable, achieving latencies close or even better latencies to those at equivalent bit-widths, potentially offering greater accuracy.
Zihan Liu 0002, Xinhao Luo, Junxian Guo, Wentao Ni, Yangjie Zhou 0001, Yue Guan 0003, Cong Guo 0003, Weihao Cui, Yu Feng 0007, Minyi Guo, Yuhao Zhu 0001, Minjia Zhang, Jingwen Leng
HPCA5
2025 Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding
abstract
Speculative decoding improves LLM inference by generating and verifying multiple tokens in parallel, but existing systems suffer from suboptimal performance due to a mismatch between dynamic speculation and static runtime assumptions. We present Yggdrasil, a co-designed system that enables latency-optimal speculative decoding through context-aware tree drafting and compiler-friendly execution. Yggdrasil introduces an equal-growth tree structure for static graph compatibility, a latency-aware optimization objective for draft selection, and stage-based scheduling to reduce overhead. Yggdrasil supports unmodified LLMs and achieves up to $3.98\times$ speedup over state-of-the-art baselines across multiple hardware setups.
Yue Guan 0003, Changming Yu, Shihan Fang, Weiming Hu 0005, Zaifeng Pan, Zheng Wang 0075, Zihan Liu 0002, Yangjie Zhou 0001, Yufei Ding 0001, Minyi Guo, Jingwen Leng
NeurIPS8
2025 ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
abstract
Large language model (LLM) decoding suffers from high latency due to fragmented execution across operators and heavy reliance on off-chip memory for data exchange and reduction. This execution model limits opportunities for fusion and incurs significant memory traffic and kernel launch overhead. While modern architectures such as NVIDIA Hopper provide distributed shared memory and low-latency intra-cluster interconnects, they expose only low-level data movement instructions, lacking structured abstractions for collective on-chip communication. To bridge this software-hardware gap, we introduce two cluster-level communication primitives, ClusterReduce and ClusterGather, which abstract common communication patterns and enable structured, high-speed data exchange and reduction between thread blocks within a cluster, allowing intermediate results to be on-chip without involving off-chip memory. Building on these abstractions, we design ClusterFusion, an execution framework that schedules communication and computation jointly to expand operator fusion scope by composing decoding stages such as QKV Projection, Attention, and Output Projection into a single fused kernels. Evaluations on H100 GPUs show that ClusterFusion outperforms state-of-the-art inference frameworks by $1.61\times$ on average in end-to-end latency across different models and configurations.
Xinhao Luo, Zihan Liu 0002, Yangjie Zhou 0001, Shihan Fang, Yu Feng 0007, Chen Zhang 0001, Shixuan Sun, Zhenzhe Zheng 0001, Jingwen Leng, Minyi Guo
NeurIPS3
2025 A Sample-Free Compilation Framework for Efficient Dynamic Tensor Computation
abstract
Dynamic-shape tensor computation poses challenges for shape-specific compilation due to variable input dimensions. Existing compilers rely on shape samples, incurring high tuning costs and performance degradation on unseen inputs. We present Helix, a dynamic tensor compilation framework with sample-free compilation and architecture-guided optimization to achieve both compilation efficiency and shape-general performance. To avoid shape sampling, Helix constructs shape-agnostic compilation by decomposing computations across architectural layers. A bidirectional strategy combines top-down abstraction to align tensor computations with architectural hierarchies, and bottom-up kernel construction to build efficient execution strategies from reusable, architecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through profiling at lower architectural levels, and achieves scalability through architecture-informed modeling at higher levels and runtime. This hierarchical design eliminates shape-specific tuning and enables shape-adaptive execution. Evaluations conducted on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate that Helix reduces compilation time by 174 × over the existing compilers and delivers 2.26 × and 3.29 × execution speedups over vendor libraries and dynamic-shape compilers, respectively.
Yangjie Zhou 0001, Weihao Cui, Zihan Liu 0002, Peng Chen 0035, Mohamed Wahib, Cong Guo 0003, Siyuan Feng 0007, Jintao Meng 0001, Haidong Lan, Jingwen Leng, Yun Lin 0001, Jin Song Dong 0001, Wenxi Zhu, Minwen Deng
SC1
2024 Fractal: Joint Multi-Level Sparse Pattern Tuning of Accuracy and Performance for DNN Pruning
abstract
Model pruning, which eliminates redundant parameters and reduces computational complexity, emerges as a viable strategy for efficient deep neural network (DNN) deployment. Owing to the irregular memory access and computation patterns in the sparse DNN models after pruning, existing arts have suggested various structured sparse patterns to enhance sparse DNN performance. In this work, we propose a unique perspective of understanding existing sparse pattern design as computation-skipping after tiling the tensor computation into multi-level hierarchies. This unified perspective opens up a new design space of multi-level sparse tiling to maximize the sparsity benefits of DNNs, as opposed to the single-level choice in current practices. We present Fractal, an auto-tuning system for sparse patterns that identifies the optimal multi-level sparse tiling pattern. We introduce PatternIR, a novel high-level intermediate representation (IR), to express a diverse range of multi-level sparse patterns. By leveraging insights from prior dense operator optimizations, we translate PatternIR into low-level compiler IRs, facilitating further operator optimization and code generation. Our evaluations demonstrate that Fractal yields substantial speedups of up to on average 3.16× on CUDA Core, 2.52× on TensorCore of GPUs compared to the state-of-art dense baseline under 75% sparsity while upholding minimal accuracy degradation compared to prior sparse operator libraries.
Yue Guan 0003, Changming Yu, Yangjie Zhou 0001, Jingwen Leng, Chao Li 0009, Minyi Guo
ASPLOS (3)3
2023 uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural Networks
abstract
As graph neural networks (GNNs) have achieved great success in many graph learning problems, it is of paramount importance to support their efficient execution. Different graphs and different operators present different patterns during execution. However, there is still a gap in the existing GNN acceleration research to explore adaptive parallelism. We show that existing GNN frameworks rely on handwritten static kernels, which fail to achieve the best performance across different graph operators and input graph structures. In this work, we propose uGrapher, a unified interface that achieves general high performance for different graph operators and datasets. The existing GNN frameworks can easily integrate our design for its simple and unified API. We take a principled approach that decouples a graph operator’s computation and schedule to achieve that. We first build a GNN-specific operator abstraction that incorporates the semantics of graph tensors and graph loops. We explore various schedule strategies based on the abstraction that can balance the well-established trade-off relationship between parallelism, locality, and efficiency. Our evaluation shows that uGrapher can bring up to 29.1× (3.5× on average) performance improvement over the state-of-the-art baselines on two studied NVIDIA GPUs.
Yangjie Zhou 0001, Jingwen Leng, Yaoxu Song, Shuwen Lu, Chao Li 0009, Minyi Guo, Wenting Shen, Yong Li 0045, Wei Lin 0016, Xiangwen Liu
ASPLOS (2)1
2023 AdaptGear: Accelerating GNN Training via Adaptive Subgraph-Level Kernels on GPUs
abstract
Graph neural networks (GNNs) are powerful tools for exploring and learning from graph structures and features. As such, achieving high-performance execution for GNNs becomes crucially important. Prior works have proposed to explore the sparsity (i.e., low density) in the input graph to accelerate GNNs, which uses the full-graph-level or block-level sparsity format. We show that they fail to balance the sparsity benefit and kernel execution efficiency. In this paper, we propose a novel system, referred to as AdaptGear, that addresses the challenge of optimizing GNNs performance by leveraging kernels tailored to the density characteristics at the subgraph level. Meanwhile, we also propose a method that dynamically chooses the optimal set of kernels for a given input graph. Our evaluation shows that AdaptGear can achieve a significant performance improvement, up to 6.49× (1.87× on average), over the state-of-the-art works on two mainstream NVIDIA GPUs across various datasets.
Yangjie Zhou 0001, Yaoxu Song, Jingwen Leng, Zihan Liu 0002, Weihao Cui, Zhendong Zhang 0004, Cong Guo 0003, Quan Chen 0002, Li Li 0012, Minyi Guo
CF1
2023 DistSim: A performance model of large-scale hybrid distributed DNN training
abstract
With the ever-increasing computational demand of DNN training workloads, distributed training has been widely adopted. A combination of data, model and pipeline parallelism strategy, called hybrid parallelism distributed training, is imported to tackle the problem of deploying large-scale models. However, how to evaluate the hybrid strategy and the utilization of each device remains a challenge since existing works either profile on a real large-scale cluster with high time and money costs or only analyze a specific type of parallelism without considering the hybrid parallelism. In this work, we proposed DistSim, an event-based performance model to accurately analyze each device's computation and communication activities with low profiling costs. DistDim breaks down the model into events according to the given distributed strategy, which can be profiled on two nodes. Then DistSim leverages the hierarchy of different parallel strategies to generate the computation and communication event-flow from layer level to model level and finally the activity timeline of each device participating in training. Experiment shows that DistSim can reach <4% errors when predicting distributing training batch time and <5% errors when predicting a single device's activity time in various hybrid strategy settings. We also provide a use-case of DistSim, automatically evaluate and search the best distributed training strategy, and find a hybrid strategy with at most 7.37× throughput improvement.
Guandong Lu, Runzhe Chen, Yakai Wang, Yangjie Zhou 0001, Rui Zhang 0040, Zheng Hu 0002, Yanming Miao, Zhifang Cai, Li Li 0012, Jingwen Leng, Minyi Guo
CF4
2020 Balancing Efficiency and Flexibility for DNN Acceleration via Temporal GPU-Systolic Array Integration
abstract
The research interest in specialized hardware accelerators for deep neural networks (DNN) spikes recently owing to their superior performance and efficiency. However, today’s DNN accelerators primarily focus on accelerating specific "kernels" such as convolution and matrix multiplication, which are vital but only part of an end-to-end DNN-enabled application. Meaningful speedups over the entire application often require supporting computations that are, while massively parallel, ill-suited to DNN accelerators. Integrating a general-purpose processor such as a CPU or a GPU incurs significant data movement overhead and leads to resource under-utilization on the DNN accelerators.We propose Simultaneous Multi-mode Architecture (SMA), a novel architecture design and execution model that offers general-purpose programmability on DNN accelerators in order to accelerate end-to-end applications. The key to SMA is the temporal integration of the systolic execution model with the GPU-like SIMD execution model. The SMA exploits the common components shared between the systolic-array accelerator and the GPU, and provides lightweight reconfiguration capability to switch between the two modes in-situ. The SMA achieves up to 63% performance improvement while consuming 23% less energy than the baseline Volta architecture with TensorCore.
Cong Guo 0003, Yangjie Zhou 0001, Jingwen Leng, Yuhao Zhu 0001, Zidong Du, Quan Chen 0002, Chao Li 0009, Bin Yao 0002, Minyi Guo
DAC2