VLDB 2026 Research / reviewers in the wild / expert
Ritchie Zhao
dblp:163/3629
· DBLP profile ↗
15ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0003-1656-9165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Hardware accelerators and domain-specific architectures · 46% Reconfigurable computing and FPGAs · 22% Electronic design automation · 20% | |
| Artificial intelligence
7 papers |
Efficient and distributed learning · 60% Deep learning architectures and training · 21% Language models and text generation · 19% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
2.5 | 5 | 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025 Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020 Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations · ICLR 2020 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
2.0 | 3 | 2025 | RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025 With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023 Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020 |
Electronic design automation
high-level synthesis |
1.4 | 5 | 2018 | Rosetta: A Realistic High-Level Synthesis Benchmark Suite for Software Programmable FPGAs · FPGA 2018 Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017 |
Machine learning › Efficient and distributed learning › inference efficiency
inference optimization |
0.9 | 1 | 2025 | RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.9 | 1 | 2025 | RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › parameter compression
mixture-of-experts compression |
0.9 | 1 | 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
KV cache compression |
0.9 | 1 | 2025 | RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic |
0.7 | 1 | 2023 | With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.4 | 2 | 2019 | Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs · FPGA 2017 Improving Neural Network Quantization without Retraining using Outlier Channel Splitting · ICML 2019 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.4 | 1 | 2019 | Building Efficient Deep Neural Networks With Unitary Group Convolutions · CVPR 2019 |
Machine learning › Efficient and distributed learning › model compression
efficient architecture design |
0.4 | 1 | 2019 | Building Efficient Deep Neural Networks With Unitary Group Convolutions · CVPR 2019 |
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
group convolution |
0.4 | 1 | 2019 | Building Efficient Deep Neural Networks With Unitary Group Convolutions · CVPR 2019 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.4 | 1 | 2019 | Improving Neural Network Quantization without Retraining using Outlier Channel Splitting · ICML 2019 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.4 | 1 | 2019 | Improving Neural Network Quantization without Retraining using Outlier Channel Splitting · ICML 2019 |
High-performance computing › performance optimization
auto-tuning |
0.3 | 1 | 2017 | A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
binary neural network accelerator |
0.3 | 1 | 2017 | Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs · FPGA 2017 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.3 | 1 | 2017 | Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Reconfigurable computing and FPGAs
FPGA compilation |
0.3 | 1 | 2017 | A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017 |
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining |
0.3 | 1 | 2017 | Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Memory systems › memory architecture
multi-bank memory |
0.3 | 1 | 2017 | Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Processor architecture and microarchitecture
pipelining |
0.3 | 1 | 2017 | Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017 |
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
processing element array |
0.3 | 1 | 2017 | Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Mathematical optimization › online optimization
bandit optimization |
0.3 | 1 | 2017 | A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.3 | 1 | 2025 | RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2025 | RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025 |
Reconfigurable computing and FPGAs
latency-insensitive interface |
0.2 | 1 | 2016 | Improving high-level synthesis with decoupled data structure optimization · DAC 2016 |
Reconfigurable computing and FPGAs
FPGA high-level synthesis |
0.2 | 1 | 2015 | Area-efficient pipelining for FPGA-targeted high-level synthesis · DAC 2015 |
Parallel and multicore computing › task scheduling
pipeline scheduling |
0.2 | 1 | 2015 | Area-efficient pipelining for FPGA-targeted high-level synthesis · DAC 2015 |
Electronic design automation
design space exploration |
0.1 | 1 | 2017 | A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017 |
Methods — techniques the papers use, named apart from their topics
quantization · 2.2sparse attention · 1.7dimensionality reduction · 1.7KV cache eviction · 1.7block floating point · 1.3wasserstein barycenter · 0.9residual approximation · 0.9microsoft floating point · 0.9precision gating · 0.4dual-precision activations · 0.4outlier channel splitting · 0.4clipping · 0.4high-level synthesis optimization · 0.3parallel autotuning · 0.3multi-armed bandit · 0.3hazard resolution · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache CompressionabstractTransformer-based Large Language Models rely critically on the KV cache to efficiently handle extended contexts during the decode phase. Yet, the size of the KV cache grows proportionally with the input length, burdening both memory bandwidth and capacity as decoding progresses. To address this challenge, we present RocketKV, a training-free KV cache compression strategy containing two consecutive stages. In the first stage, it performs coarse-grain permanent KV cache eviction on the input sequence tokens. In the second stage, it adopts a hybrid sparse attention method to conduct fine-grain top-k sparse attention, approximating the attention scores by leveraging both head and sequence dimensionality reductions. We show that RocketKV provides a compression ratio of up to 400×, end-to-end speedup of up to 3.7× as well as peak memory reduction of up to 32.6% in the decode phase on an NVIDIA A100 GPU compared to the full KV cache baseline, while achieving negligible accuracy loss on a variety of long-context tasks. We also propose a variant of RocketKV for multi-turn scenarios, which consistently outperforms other existing methods and achieves accuracy nearly on par with an oracle top-k attention scheme. The source code is available here: https://github.com/NVlabs/RocketKV. Payman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai, Zhiding Yu, Alexey Tumanov |
ICML | 3 |
| 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual RestorationabstractMixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for each input token. The sparse structure, while allowing constant time costs, results in space inefficiency: we still need to load all the model parameters during inference. We introduce ResMoE, an innovative MoE approximation framework that utilizes Wasserstein barycenter to extract a common expert (barycenter expert) and approximate the residuals between this barycenter expert and the original ones. ResMoE enhances the space efficiency for inference of large-scale MoE Transformers in a one-shot and data-agnostic manner without retraining while maintaining minimal accuracy loss, thereby paving the way for broader accessibility to large language models. We demonstrate the effectiveness of ResMoE through extensive experiments on Switch Transformer, Mixtral, and DeepSeekMoE models. The results show that ResMoE can reduce the number of parameters in an expert by up to 75% while maintaining comparable performance. The code is available at https://github.com/iDEA-iSAIL-Lab-UIUC/ResMoE, and the supplementary appendix is available at https://famous-blue-raincoat.github.io/mengtingai/files/ResMoE_Appendix.pdf. Mengting Ai, Tianxin Wei, Yifan Chen 0004, Zhichen Zeng 0001, Ritchie Zhao, Girish Varatkar, Bita Darvish Rouhani, Xianfeng Tang, Hanghang Tong, Jingrui He |
KDD (1) | 5 |
| 2023 | With Shared Microexponents, A Little Shifting Goes a Long WayabstractThis paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems. Bita Darvish Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lai Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L'Heureux, Matt Perry, Doug Burger, Eric S. Chung, Zhaoxia Deng, Sam Naghshineh, Jongsoo Park, Maxim Naumov |
ISCA | 2 |
| 2020 | Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations
Yichi Zhang 0006, Ritchie Zhao, Weizhe Hua, Nayun Xu, G. Edward Suh, Zhiru Zhang |
ICLR | 2 |
| 2020 | Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointabstractIn this paper, we explore the limits of Microsoft Floating Point (MSFP), a new class of datatypes developed for production cloud-scale inferencing on custom hardware. Through the co-evolution of hardware design and algorithms, MSFP achieves accuracy comparable to or better than industry standards Bfloat16 and INT8 at 3x and 4x lower cost, respectively. MSFP incurs negligible impact to accuracy (<1%), requires no changes to the model topology, and is integrated with a mature cloud production pipeline. MSFP supports various classes of deep learning models including CNNs, RNNs, and Transformers without modification. Finally, we characterize the accuracy and implementation of MSFP and demonstrate its efficacy on a number of production scenarios, including models that power major online scenarios such as web search, question-answering, and image classification. Bita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Jeremy Fowers, Kalin Ovtcharov, Anna Vinogradsky, Sarah Massengill, Lita Yang, Ray Bittner, Alessandro Forin, Haishan Zhu, Taesik Na, Prerak Patel, Shuai Che, Lok Chand Koppaka, Subhojit Som, Kaustav Das, Saurabh Tiwary, Steven K. Reinhardt, Sitaram Lanka, Eric S. Chung, Doug Burger |
NeurIPS | 3 |
| 2019 | Building Efficient Deep Neural Networks With Unitary Group ConvolutionsabstractWe propose unitary group convolutions (UGConvs), a building block for CNNs which compose a group convolution with unitary transforms in feature space to learn a richer set of representations than group convolution alone. UGConvs generalize two disparate ideas in CNN architecture, channel shuffling (i.e. ShuffleNet) and block-circulant networks (i.e. CirCNN), and provide unifying insights that lead to a deeper understanding of each technique. We experimentally demonstrate that dense unitary transforms can outperform channel shuffling in DNN accuracy. On the other hand, different dense transforms exhibit comparable accuracy performance. Based on these observations we propose HadaNet, a UGConv network using Hadamard transforms. HadaNets achieve similar accuracy to circulant networks with lower computation complexity, and better accuracy than ShuffleNets with the same number of parameters and floating-point multiplies. Ritchie Zhao, Jordan Dotzel, Christopher De Sa, Zhiru Zhang |
CVPR | 1 |
| 2019 | Improving Neural Network Quantization without Retraining using Outlier Channel SplittingabstractQuantization can improve the execution latency and energy efficiency of neural networks on both commodity GPUs and specialized accelerators. The majority of existing literature focuses on training quantized DNNs, while this work examines the less-studied topic of quantizing a floating-point model without (re)training. DNN weights and activations follow a bell-shaped distribution post-training, while practical hardware uses a linear quantization grid. This leads to challenges in dealing with outliers in the distribution. Prior work has addressed this by clipping the outliers or using specialized hardware. In this work, we propose outlier channel splitting (OCS), which duplicates channels containing outliers, then halves the channel values. The network remains functionally identical, but affected outliers are moved toward the center of the distribution. OCS requires no additional training and works on commodity hardware. Experimental evaluation on ImageNet classification and language modeling shows that OCS can outperform state-of-the-art clipping techniques with only minor overhead. Ritchie Zhao, Jordan Dotzel, Christopher De Sa, Zhiru Zhang |
ICML | 1 |
| 2018 | Rosetta: A Realistic High-Level Synthesis Benchmark Suite for Software Programmable FPGAsabstractModern high-level synthesis (HLS) tools greatly reduce the turn-around time of designing and implementing complex FPGA-based accelerators. They also expose various optimization opportunities, which cannot be easily explored at the register-transfer level. With the increasing adoption of the HLS design methodology and continued advances of synthesis optimization, there is a growing need for realistic benchmarks to (1) facilitate comparisons between tools, (2) evaluate and stress-test new synthesis techniques, and (3) establish meaningful performance baselines to track progress of the HLS technology. While several HLS benchmark suites already exist, they are primarily comprised of small textbook-style function kernels, instead of complete and complex applications. To address this limitation, we introduce Rosetta, a realistic benchmark suite for software programmable FPGAs. Designs in Rosetta are fully-developed applications. They are associated with realistic performance constraints, and optimized with advanced features of modern HLS tools. We believe that Rosetta is not only useful for the HLS research community, but can also serve as a set of design tutorials for non-expert HLS users. In this paper we describe the characteristics of our benchmarks and the optimization techniques applied to them. We further report experimental results on an embedded FPGA device as well as a cloud FPGA platform. Udit Gupta 0001, Steve Dai, Ritchie Zhao, Nitish Kumar Srivastava, Hanchen Jin, Joseph Featherston, Yi-Hsiang Lai, Gai Liu, Gustavo Angarita Velasquez, Zhiru Zhang |
FPGA | 4 |
| 2017 | Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis
Steve Dai, Ritchie Zhao, Gai Liu, Shreesha Srinath, Udit Gupta 0001, Christopher Batten, Zhiru Zhang |
FPGA | 2 |
| 2017 | A Parallel Bandit-Based Approach for Autotuning FPGA Compilation
Chang Xu 0005, Gai Liu, Ritchie Zhao, Guojie Luo, Zhiru Zhang |
FPGA | 3 |
| 2017 | Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs
Ritchie Zhao, Weinan Song, Tianwei Xing, Jeng-Hau Lin, Mani Srivastava 0001, Rajesh K. Gupta 0001, Zhiru Zhang |
FPGA | 1 |
| 2017 | Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop NestsabstractModern high-level synthesis (HLS) tools commonly employ pipelining to achieve efficient loop acceleration by overlapping the execution of successive loop iterations. While existing HLS pipelining techniques obtain good performance with low complexity for regular loop nests, they provide inadequate support for effectively synthesizing irregular loop nests. For loop nests with dynamic-bound inner loops, current pipelining techniques require unrolling of the inner loops, which is either very expensive in resource or even inapplicable due to dynamic loop bounds. To address this major limitation, this paper proposes ElasticFlow, a novel architecture capable of dynamically distributing inner loops to an array of processing units (LPUs) in an area-efficient manner. The proposed LPUs can be either specialized to execute an individual inner loop or shared among multiple inner loops to balance the tradeoff between performance and area. A customized banked memory architecture is proposed to coordinate memory accesses among different LPUs to maximize memory bandwidth without significantly increasing memory footprint. We evaluate ElasticFlow using a variety of real-life applications and demonstrate significant performance improvements over a state-of-the-art commercial HLS tool for Xilinx FPGAs. Gai Liu, Mingxing Tan, Steve Dai, Ritchie Zhao, Zhiru Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Improving high-level synthesis with decoupled data structure optimizationabstractExisting high-level synthesis (HLS) tools are mostly effective on algorithm-dominated programs that only use primitive data structures such as fixed size arrays and queues. However, many widely used data structures such as priority queues, heaps, and trees feature complex member methods with data-dependent work and irregular memory access patterns. These methods can be inlined to their call sites, but this does not address the aforementioned issues and may further complicate conventional HLS optimizations, resulting in a low-performance hardware implementation. To overcome this deficiency, we propose a novel HLS architectural template in which complex data structures are decoupled from the algorithm using a latency-insensitive interface. This enables overlapped execution of the algorithm and data structure methods, as well as parallel and out-of-order execution of independent methods on multiple decoupled lanes. Experimental results across a variety of real-life benchmarks show our approach is capable of achieving very promising speedups without causing significant area overhead. Ritchie Zhao, Gai Liu, Shreesha Srinath, Christopher Batten, Zhiru Zhang |
DAC | 1 |
| 2015 | Area-efficient pipelining for FPGA-targeted high-level synthesisabstractTraditional techniques for pipeline scheduling in high-level synthesis for FPGAs assume an additive delay model where each operation incurs a pre-characterized delay. While a good approximation for some operation types, this fails to consider technology mapping, where a group of logic operations can be mapped to a single look-up table (LUT) and together incur one LUT worth of delay. We propose an exact formulation of the throughput-constrained, mapping-aware pipeline scheduling problem for FPGA-targeted high-level synthesis with area minimization being a primary objective. By taking this cross-layered approach, our technique is able to mitigate the pessimism inherent in static delay estimates and reduce the usage of LUTs and pipeline registers. Experimental results using our method demonstrate improved resource utilization for a number of logic-intensive, real-life benchmarks compared to a state-of-the-art commercial HLS tool for Xilinx FPGAs. Ritchie Zhao, Mingxing Tan, Steve Dai, Zhiru Zhang |
DAC | 1 |
| 2015 | ElasticFlow: A Complexity-Effective Approach for Pipelining Irregular Loop NestsabstractModern high-level synthesis (HLS) tools commonly employ pipelining to achieve efficient loop acceleration by overlapping the execution of successive loop iterations. However, existing HLS techniques provide inadequate support for pipelining irregular loop nests that contain dynamic-bound inner loops, where unrolling is either very expensive or not even applicable. To overcome this major limitation, we propose ElasticFlow, a novel architectural synthesis approach capable of dynamically distributing inner loops to an array of loop processing units (LPUs) in a complexity-effective manner. These LPUs can be either specialized to execute an individual loop or shared amongst multiple inner loops for area reduction. We evaluate ElasticFlow using a variety of real-life applications and demonstrate significant performance improvements over a widely used commercial HLS tool for Xilinx FPGAs. Mingxing Tan, Gai Liu, Ritchie Zhao, Steve Dai, Zhiru Zhang |
ICCAD | 3 |