Zhongkai Yu

dblp:327/2154 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-9893-3498ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM Inference
abstract
Large-scale Mixture of Experts (MoE) Large Language Models (LLMs) have recently become the frontier open-weight models, achieving remarkable model capability similar to proprietary ones. But their random expert selection mechanism introduces significant data movement overhead that becomes the dominant bottleneck in multi-unit LLM serving systems. To understand the patterns underlying this data movement, we conduct comprehensive data-movement-centric profiling across four state-of-the-art large-scale MoE models released in 2025 (200B-1000B) using over 24,000 requests spanning diverse workloads. We perform systematic analysis from both temporal and spatial perspectives and distill six key insights to guide the design of diverse serving systems. We verify these insights on both future wafer-scale GPU architectures and existing GPU systems. On wafer-scale GPUs, lightweight architectural modifications guided by our insights yield a 6.6$\times$ average speedup across four 200B--1000B models. On existing GPU systems, our insights drive the design of a prefill-aware expert placement algorithm that achieves up to 1.25$\times$ speedup on MoE computation. Our work presents the first comprehensive data-centric analysis of large-scale MoE models together with a concrete design study applying the learned lessons. Our profiling traces are publicly available at \href{https://huggingface.co/datasets/core12345/MoE_expert_selection_trace}{\textcolor{blue}{https://huggingface.co/datasets/core12345/MoE\_expert\_selection\_trace}}.
Zhongkai Yu, Yue Guan 0003, Zhengding Hu, Shuyi Pei, Yangwook Kang, Yufei Ding 0001, Po-An Tsai
ISCA1
2026 DomSim: Hardware-Aware Hybrid Fault Simulation With Dominator Tree-Guided Partitioning
abstract
Gate-level fault simulation is a critical step in design for test and functional safety verification of the chip design process, essential to ensuring circuit reliability. As chip complexity grows for mission-critical applications such as autonomous vehicles, medical devices, and military systems, the efficiency of fault simulation increasingly becomes a bottleneck in the chip’s time-to-market. However, existing methods often suffer from computational redundancy, inefficiencies in memory access, or failure to optimize performance for specific CPU hardware platforms. This paper proposes DomSim, a hardware-aware hybrid fault simulation method that combines compiled simulation and event-driven simulation with an optimized computation-to-memory-access ratio. By utilizing circuit information and hierarchical structure provided by dominator trees, DomSim achieves high-quality circuit partitioning, optimizing hardware resource utilization and memory access locality. Furthermore, a parameter adjustment strategy tailored to hardware capabilities and circuit characteristics enables adaptive optimization. Extensive experiments show that DomSim surpasses a commercial tool by 10.29× on average. Further experiments demonstrate that DomSim exhibits good adaptability across different hardware platforms and circuits, highlighting the superiority of our method.
Hui Wang 0152, Zizhen Liu, Jianan Mu, Shengwen Liang, Zhongkai Yu, Zheng Liang 0003, Jiaping Tang, Jing Ye 0001, Xiaowei Li 0001, Bei Yu 0001, Huawei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 KPerfIR: Towards a Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI Workloads
Yue Guan 0003, Yuanwei Fang, Keren Zhou 0001, Corbin Robeck, Manman Ren, Zhongkai Yu, Yufei Ding 0001, Adnan Aziz
OSDI6
2025 Harmonia: A Unified Architecture for Efficient Deep Symbolic Regression
abstract
Symbolic regression (SR), the process of formulating a mathematical expression based on observed data points, is a fundamental task in artificial intelligence but is often hindered by its intense computational demands. Deep-learning-based SR methods (DSR) aim to alleviate these demands by breaking down the SR process into two stages: 1) neural network (NN) inference and 2) Broyden-Fletcher–Goldfarb-Shanno (BFGS) optimization. Although NN accelerators can expedite the NN stage, the performance of the BFGS optimization is compromised due to its poor performance for the variety of transcendental functions. Moreover, the distinct computational characteristics of NN inference and BFGS cause not only low hardware utilization but also significant area waste. To address these issues, we propose Harmonia, a unified architecture with the neural transcendental function unit (NTFU) and the Unified Array for efficient DSR. The NTFU utilizes the radial basis function network (RBFN) as a universal approximator for various transcendental functions, which significantly reduces the heavy transcendental function computation cost. We further propose an efficient training algorithm called random nonlinear optimization (RNO) to obtain a lightweight RBFN without accuracy loss. Moreover, Harmonia supports configurable dataflow which integrates the two computing stages into the Unified Array. Experimental results show that Harmonia achieves hardware utilization of 83.83%, on average. Compared to the GPU baseline, Harmonia achieves$4.8\times $speedup and$47.6\times $energy saving, alongside considerable low area cost.
Tianyun Ma, Yuanbo Wen 0001, Xinkai Song, Pengwei Jin, Husheng Han, Ziyuan Nan, Zhongkai Yu, Shaohui Peng, Yongwei Zhao 0001, Huaping Chen 0001, Zidong Du, Xing Hu 0001, Qi Guo 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2024 Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
abstract
Deploying advanced large language models on edge devices, such as smartphones and robotics, is a growing trend that enhances user data privacy and network connectivity resilience while preserving intelligent capabilities. However, such a task exhibits single-batch computing with incredibly low arithmetic intensity, which poses the significant challenges of huge memory footprint and bandwidth demands on limited edge resources. To address these issues, we introduce Cambricon-LLM, a chiplet-based hybrid architecture with NPU and a dedicated NAND flash chip to enable efficient on-device inference of 70B LLMs. Such a hybrid architecture utilizes both the high computing capability of NPU and the data capacity of the NAND flash chip, with the proposed hardware-tiling strategy that minimizes the data movement overhead between NPU and NAND flash chip. Specifically, the NAND flash chip, enhanced by our innovative in-flash computing and on-die ECC techniques, excels at performing precise lightweight on-die processing. Simultaneously, the NPU collaborates with the flash chip for matrix operations and handles special function computations beyond the flash's on-die processing capabilities. Overall, Cambricon-LLM enables the on-device inference of 70B LLMs at a speed of 3.44 token/s, and 7B LLMs at a speed of 36.34 token/s, which is over 22× to 45× faster than existing flash-offloading technologies, showing the potentiality of deploying powerful LLMs in edge devices.
Zhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai, Ziyuan Nan, Xinkai Song, Yifan Hao 0001, Jie Zhang 0048, Tian Zhi, Yongwei Zhao 0001, Zidong Du, Xing Hu 0001, Qi Guo 0001, Tianshi Chen 0002
MICRO1
2024 Environmental Condition Aware Super-Resolution Acceleration Framework in Server-Client Hierarchies
abstract
In the current landscape, high-resolution (HR) videos have gained immense popularity, promising an elevated viewing experience. Recent research has demonstrated that the video super-resolution (SR) algorithm, empowered by deep neural networks (DNNs), can substantially enhance the quality of HR videos by processing low-resolution (LR) frames. However, the existing DNN models demand significant computational resources, posing challenges for the deployment of SR algorithms on client devices. While numerous accelerators have proposed solutions, their primary focus remains on client-side optimization. In contrast, our research recognizes that the HR video is originally stored in the cloud server and presents an untapped opportunity for achieving both high accuracy and performance improvements. Building on this insight, this article introduces an end-to-end video CODEC-assisted super-resolution (E 2 SR+) algorithm, which tightly integrates the cloud server with the client device to deliver a seamless and real-time video viewing experience. We propose the motion vector search algorithm executed in a cloud server, which can search the motion vectors and residuals for a part of the HR video frames and then pack them as add-ons. We also design an auto-encoder algorithm to down-sample the residuals to save the bitstream cost while guaranteeing the quality of the residuals. Lastly, we propose a reconstruction algorithm performed in the client to quickly reconstruct the corresponding HR frames using the add-ons to skip part of the DNN computations. To implement the E 2 SR+ algorithm, we design corresponding E 2 SR+ architecture in the client, which achieves significant speedup with minimal hardware overhead. Given that the environmental condition varies in the server–client hierarchies, we believe that simply applying E 2 SR+ to all frames is irrational. Accordingly, we offer an environmental condition–aware system to chase the best performance while adapting to the diverse environment. In the system, we design a linear programming (LP) model to simulate the environment and allocate frames to three existing mechanisms. Our experimental results demonstrate that the E 2 SR+ algorithm enhances the peak signal-to-noise ratio by 1.2, 2.5, and 2.3 compared with the state-of-the-art (SOTA) methods EDVR, BasicVSR, and BasicVSR++, respectively. In terms of performance, the E 2 SR+ architecture offers significant improvements over existing SOTA methods. For instance, while BasicVSR++ requires 98 ms on an NVIDIA V100 graphics processing unit (GPU) to generate a 1,280 × 720 HR frame, the E 2 SR+ architecture reduces the execution time to just 39 ms, highlighting the efficiency and effectiveness of our proposed method. Overall, the E 2 SR+ architecture respectively achieves 1.4×, 2.2×, 4.6×, and 442.0× performance improvement compared with ADAS, ISRAcc, the NVIDIA V100 GPU, and a central processing unit. Lastly, the proposed system showcases its superiority and surpasses all the existing mechanisms in terms of execution time when varying environmental conditions.
Zhuoran Song, Zhongkai Yu, Xinkai Song, Yifan Hao 0001, Li Jiang 0002, Naifeng Jing, Xiaoyao Liang
ACM Trans. Archit. Code Optim.2
2022 E2SR: an end-to-end video CODEC assisted system for super resolution acceleration
abstract
Nowadays high-resolution (HR) videos have been a popular choice for a better viewing experience. Recent works have shown that super-resolution (SR) algorithms can provide superior quality HR video by applying the deep neural network (DNN) to each low-resolution (LR) frame. Obviously, such per-frame DNN processing is compute-intensive and hampers the deployment of SR algorithms on mobile devices. Although many accelerators have proposed solutions, they only focus on mobile devices. Differently, we notice that the HR video is originally stored in the cloud server and should be well exploited to gain high accuracy and performance improvement. Based on this observation, this paper proposes an end-to-end video CODEC assisted system (E2SR), which tightly couples the cloud server with the device to deliver a smooth and real-time video viewing experience. We propose the motion vector search algorithm executed in the cloud server, which can search the motion vectors and residuals for part of HR video frames and then pack them as addons. We further propose the reconstruction algorithm executed in the device to fast reconstruct the corresponding HR frames using the addons to skip part of DNN computations. We design the corresponding E2SR architecture to enable the reconstruction algorithm in the device, which achieves significant speedup with minimal hardware overhead. Our experimental results show that the E2SR system achieves 3.4x performance improvement with less than 0.56 PSNR loss compared with the state-of-the-art "EDVR" scheme.
Zhuoran Song, Zhongkai Yu, Naifeng Jing, Xiaoyao Liang
DAC2