Ritchie Zhao

dblp:163/3629 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0003-1656-9165ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Hardware accelerators and domain-specific architectures · 46% Reconfigurable computing and FPGAs · 22% Electronic design automation · 20%
Artificial intelligence
7 papers
Efficient and distributed learning · 60% Deep learning architectures and training · 21% Language models and text generation · 19%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.552025
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025
Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020
Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations · ICLR 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
2.032025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023
Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020
Electronic design automation
high-level synthesis
1.452018
Rosetta: A Realistic High-Level Synthesis Benchmark Suite for Software Programmable FPGAs · FPGA 2018
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017
Machine learning › Efficient and distributed learning › inference efficiency
inference optimization
0.912025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Natural language and speech › Language models and text generation
large language model inference
0.912025
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.912025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Machine learning › Efficient and distributed learning › model compression › parameter compression
mixture-of-experts compression
0.912025
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
KV cache compression
0.912025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic
0.712023
With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.422019
Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs · FPGA 2017
Improving Neural Network Quantization without Retraining using Outlier Channel Splitting · ICML 2019
Machine learning › Deep learning architectures and training
convolutional neural network
0.412019
Building Efficient Deep Neural Networks With Unitary Group Convolutions · CVPR 2019
Machine learning › Efficient and distributed learning › model compression
efficient architecture design
0.412019
Building Efficient Deep Neural Networks With Unitary Group Convolutions · CVPR 2019
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
group convolution
0.412019
Building Efficient Deep Neural Networks With Unitary Group Convolutions · CVPR 2019
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.412019
Improving Neural Network Quantization without Retraining using Outlier Channel Splitting · ICML 2019
Machine learning › Efficient and distributed learning › model compression
quantization
0.412019
Improving Neural Network Quantization without Retraining using Outlier Channel Splitting · ICML 2019
High-performance computing › performance optimization
auto-tuning
0.312017
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
binary neural network accelerator
0.312017
Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs · FPGA 2017
Reconfigurable computing and FPGAs
FPGA accelerator
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Reconfigurable computing and FPGAs
FPGA compilation
0.312017
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Memory systems › memory architecture
multi-bank memory
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Processor architecture and microarchitecture
pipelining
0.312017
Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
processing element array
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Mathematical optimization › online optimization
bandit optimization
0.312017
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.312025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Machine learning › Deep learning architectures and training
transformer
0.312025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Reconfigurable computing and FPGAs
latency-insensitive interface
0.212016
Improving high-level synthesis with decoupled data structure optimization · DAC 2016
Reconfigurable computing and FPGAs
FPGA high-level synthesis
0.212015
Area-efficient pipelining for FPGA-targeted high-level synthesis · DAC 2015
Parallel and multicore computing › task scheduling
pipeline scheduling
0.212015
Area-efficient pipelining for FPGA-targeted high-level synthesis · DAC 2015
Electronic design automation
design space exploration
0.112017
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017

Methods — techniques the papers use, named apart from their topics

quantization · 2.2sparse attention · 1.7dimensionality reduction · 1.7KV cache eviction · 1.7block floating point · 1.3wasserstein barycenter · 0.9residual approximation · 0.9microsoft floating point · 0.9precision gating · 0.4dual-precision activations · 0.4outlier channel splitting · 0.4clipping · 0.4high-level synthesis optimization · 0.3parallel autotuning · 0.3multi-armed bandit · 0.3hazard resolution · 0.3
YearPublicationVenuePosition
2025 RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
abstract
Transformer-based Large Language Models rely critically on the KV cache to efficiently handle extended contexts during the decode phase. Yet, the size of the KV cache grows proportionally with the input length, burdening both memory bandwidth and capacity as decoding progresses. To address this challenge, we present RocketKV, a training-free KV cache compression strategy containing two consecutive stages. In the first stage, it performs coarse-grain permanent KV cache eviction on the input sequence tokens. In the second stage, it adopts a hybrid sparse attention method to conduct fine-grain top-k sparse attention, approximating the attention scores by leveraging both head and sequence dimensionality reductions. We show that RocketKV provides a compression ratio of up to 400×, end-to-end speedup of up to 3.7× as well as peak memory reduction of up to 32.6% in the decode phase on an NVIDIA A100 GPU compared to the full KV cache baseline, while achieving negligible accuracy loss on a variety of long-context tasks. We also propose a variant of RocketKV for multi-turn scenarios, which consistently outperforms other existing methods and achieves accuracy nearly on par with an oracle top-k attention scheme. The source code is available here: https://github.com/NVlabs/RocketKV.
Payman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai, Zhiding Yu, Alexey Tumanov
ICML3
2025 ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
abstract
Mixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for each input token. The sparse structure, while allowing constant time costs, results in space inefficiency: we still need to load all the model parameters during inference. We introduce ResMoE, an innovative MoE approximation framework that utilizes Wasserstein barycenter to extract a common expert (barycenter expert) and approximate the residuals between this barycenter expert and the original ones. ResMoE enhances the space efficiency for inference of large-scale MoE Transformers in a one-shot and data-agnostic manner without retraining while maintaining minimal accuracy loss, thereby paving the way for broader accessibility to large language models. We demonstrate the effectiveness of ResMoE through extensive experiments on Switch Transformer, Mixtral, and DeepSeekMoE models. The results show that ResMoE can reduce the number of parameters in an expert by up to 75% while maintaining comparable performance. The code is available at https://github.com/iDEA-iSAIL-Lab-UIUC/ResMoE, and the supplementary appendix is available at https://famous-blue-raincoat.github.io/mengtingai/files/ResMoE_Appendix.pdf.
Mengting Ai, Tianxin Wei, Yifan Chen 0004, Zhichen Zeng 0001, Ritchie Zhao, Girish Varatkar, Bita Darvish Rouhani, Xianfeng Tang, Hanghang Tong, Jingrui He
KDD (1)5
2023 With Shared Microexponents, A Little Shifting Goes a Long Way
abstract
This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems.
Bita Darvish Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lai Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L'Heureux, Matt Perry, Doug Burger, Eric S. Chung, Zhaoxia Deng, Sam Naghshineh, Jongsoo Park, Maxim Naumov
ISCA2
2020 Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations
Yichi Zhang 0006, Ritchie Zhao, Weizhe Hua, Nayun Xu, G. Edward Suh, Zhiru Zhang
ICLR2
2020 Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point
abstract
In this paper, we explore the limits of Microsoft Floating Point (MSFP), a new class of datatypes developed for production cloud-scale inferencing on custom hardware. Through the co-evolution of hardware design and algorithms, MSFP achieves accuracy comparable to or better than industry standards Bfloat16 and INT8 at 3x and 4x lower cost, respectively. MSFP incurs negligible impact to accuracy (<1%), requires no changes to the model topology, and is integrated with a mature cloud production pipeline. MSFP supports various classes of deep learning models including CNNs, RNNs, and Transformers without modification. Finally, we characterize the accuracy and implementation of MSFP and demonstrate its efficacy on a number of production scenarios, including models that power major online scenarios such as web search, question-answering, and image classification.
Bita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Jeremy Fowers, Kalin Ovtcharov, Anna Vinogradsky, Sarah Massengill, Lita Yang, Ray Bittner, Alessandro Forin, Haishan Zhu, Taesik Na, Prerak Patel, Shuai Che, Lok Chand Koppaka, Subhojit Som, Kaustav Das, Saurabh Tiwary, Steven K. Reinhardt, Sitaram Lanka, Eric S. Chung, Doug Burger
NeurIPS3
2019 Building Efficient Deep Neural Networks With Unitary Group Convolutions
abstract
We propose unitary group convolutions (UGConvs), a building block for CNNs which compose a group convolution with unitary transforms in feature space to learn a richer set of representations than group convolution alone. UGConvs generalize two disparate ideas in CNN architecture, channel shuffling (i.e. ShuffleNet) and block-circulant networks (i.e. CirCNN), and provide unifying insights that lead to a deeper understanding of each technique. We experimentally demonstrate that dense unitary transforms can outperform channel shuffling in DNN accuracy. On the other hand, different dense transforms exhibit comparable accuracy performance. Based on these observations we propose HadaNet, a UGConv network using Hadamard transforms. HadaNets achieve similar accuracy to circulant networks with lower computation complexity, and better accuracy than ShuffleNets with the same number of parameters and floating-point multiplies.
Ritchie Zhao, Jordan Dotzel, Christopher De Sa, Zhiru Zhang
CVPR1
2019 Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
abstract
Quantization can improve the execution latency and energy efficiency of neural networks on both commodity GPUs and specialized accelerators. The majority of existing literature focuses on training quantized DNNs, while this work examines the less-studied topic of quantizing a floating-point model without (re)training. DNN weights and activations follow a bell-shaped distribution post-training, while practical hardware uses a linear quantization grid. This leads to challenges in dealing with outliers in the distribution. Prior work has addressed this by clipping the outliers or using specialized hardware. In this work, we propose outlier channel splitting (OCS), which duplicates channels containing outliers, then halves the channel values. The network remains functionally identical, but affected outliers are moved toward the center of the distribution. OCS requires no additional training and works on commodity hardware. Experimental evaluation on ImageNet classification and language modeling shows that OCS can outperform state-of-the-art clipping techniques with only minor overhead.
Ritchie Zhao, Jordan Dotzel, Christopher De Sa, Zhiru Zhang
ICML1
2018 Rosetta: A Realistic High-Level Synthesis Benchmark Suite for Software Programmable FPGAs
abstract
Modern high-level synthesis (HLS) tools greatly reduce the turn-around time of designing and implementing complex FPGA-based accelerators. They also expose various optimization opportunities, which cannot be easily explored at the register-transfer level. With the increasing adoption of the HLS design methodology and continued advances of synthesis optimization, there is a growing need for realistic benchmarks to (1) facilitate comparisons between tools, (2) evaluate and stress-test new synthesis techniques, and (3) establish meaningful performance baselines to track progress of the HLS technology. While several HLS benchmark suites already exist, they are primarily comprised of small textbook-style function kernels, instead of complete and complex applications. To address this limitation, we introduce Rosetta, a realistic benchmark suite for software programmable FPGAs. Designs in Rosetta are fully-developed applications. They are associated with realistic performance constraints, and optimized with advanced features of modern HLS tools. We believe that Rosetta is not only useful for the HLS research community, but can also serve as a set of design tutorials for non-expert HLS users. In this paper we describe the characteristics of our benchmarks and the optimization techniques applied to them. We further report experimental results on an embedded FPGA device as well as a cloud FPGA platform.
Udit Gupta 0001, Steve Dai, Ritchie Zhao, Nitish Kumar Srivastava, Hanchen Jin, Joseph Featherston, Yi-Hsiang Lai, Gai Liu, Gustavo Angarita Velasquez, Zhiru Zhang
FPGA4
2017 Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis
Steve Dai, Ritchie Zhao, Gai Liu, Shreesha Srinath, Udit Gupta 0001, Christopher Batten, Zhiru Zhang
FPGA2
2017 A Parallel Bandit-Based Approach for Autotuning FPGA Compilation
Chang Xu 0005, Gai Liu, Ritchie Zhao, Guojie Luo, Zhiru Zhang
FPGA3
2017 Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs
Ritchie Zhao, Weinan Song, Tianwei Xing, Jeng-Hau Lin, Mani Srivastava 0001, Rajesh K. Gupta 0001, Zhiru Zhang
FPGA1
2017 Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests
abstract
Modern high-level synthesis (HLS) tools commonly employ pipelining to achieve efficient loop acceleration by overlapping the execution of successive loop iterations. While existing HLS pipelining techniques obtain good performance with low complexity for regular loop nests, they provide inadequate support for effectively synthesizing irregular loop nests. For loop nests with dynamic-bound inner loops, current pipelining techniques require unrolling of the inner loops, which is either very expensive in resource or even inapplicable due to dynamic loop bounds. To address this major limitation, this paper proposes ElasticFlow, a novel architecture capable of dynamically distributing inner loops to an array of processing units (LPUs) in an area-efficient manner. The proposed LPUs can be either specialized to execute an individual inner loop or shared among multiple inner loops to balance the tradeoff between performance and area. A customized banked memory architecture is proposed to coordinate memory accesses among different LPUs to maximize memory bandwidth without significantly increasing memory footprint. We evaluate ElasticFlow using a variety of real-life applications and demonstrate significant performance improvements over a state-of-the-art commercial HLS tool for Xilinx FPGAs.
Gai Liu, Mingxing Tan, Steve Dai, Ritchie Zhao, Zhiru Zhang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 Improving high-level synthesis with decoupled data structure optimization
abstract
Existing high-level synthesis (HLS) tools are mostly effective on algorithm-dominated programs that only use primitive data structures such as fixed size arrays and queues. However, many widely used data structures such as priority queues, heaps, and trees feature complex member methods with data-dependent work and irregular memory access patterns. These methods can be inlined to their call sites, but this does not address the aforementioned issues and may further complicate conventional HLS optimizations, resulting in a low-performance hardware implementation. To overcome this deficiency, we propose a novel HLS architectural template in which complex data structures are decoupled from the algorithm using a latency-insensitive interface. This enables overlapped execution of the algorithm and data structure methods, as well as parallel and out-of-order execution of independent methods on multiple decoupled lanes. Experimental results across a variety of real-life benchmarks show our approach is capable of achieving very promising speedups without causing significant area overhead.
Ritchie Zhao, Gai Liu, Shreesha Srinath, Christopher Batten, Zhiru Zhang
DAC1
2015 Area-efficient pipelining for FPGA-targeted high-level synthesis
abstract
Traditional techniques for pipeline scheduling in high-level synthesis for FPGAs assume an additive delay model where each operation incurs a pre-characterized delay. While a good approximation for some operation types, this fails to consider technology mapping, where a group of logic operations can be mapped to a single look-up table (LUT) and together incur one LUT worth of delay. We propose an exact formulation of the throughput-constrained, mapping-aware pipeline scheduling problem for FPGA-targeted high-level synthesis with area minimization being a primary objective. By taking this cross-layered approach, our technique is able to mitigate the pessimism inherent in static delay estimates and reduce the usage of LUTs and pipeline registers. Experimental results using our method demonstrate improved resource utilization for a number of logic-intensive, real-life benchmarks compared to a state-of-the-art commercial HLS tool for Xilinx FPGAs.
Ritchie Zhao, Mingxing Tan, Steve Dai, Zhiru Zhang
DAC1
2015 ElasticFlow: A Complexity-Effective Approach for Pipelining Irregular Loop Nests
abstract
Modern high-level synthesis (HLS) tools commonly employ pipelining to achieve efficient loop acceleration by overlapping the execution of successive loop iterations. However, existing HLS techniques provide inadequate support for pipelining irregular loop nests that contain dynamic-bound inner loops, where unrolling is either very expensive or not even applicable. To overcome this major limitation, we propose ElasticFlow, a novel architectural synthesis approach capable of dynamically distributing inner loops to an array of loop processing units (LPUs) in a complexity-effective manner. These LPUs can be either specialized to execute an individual loop or shared amongst multiple inner loops for area reduction. We evaluate ElasticFlow using a variety of real-life applications and demonstrate significant performance improvements over a widely used commercial HLS tool for Xilinx FPGAs.
Mingxing Tan, Gai Liu, Ritchie Zhao, Steve Dai, Zhiru Zhang
ICCAD3