VLDB 2026 Research / reviewers in the wild / expert
Keke Zhai
dblp:143/0311
· DBLP profile ↗
7ranked-venue papers
4as first author
2since 2021 · last 2026
0000-0002-5240-8011ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 77% Information retrieval · 12% Indexing and storage engines · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 50% Hardware accelerators and domain-specific architectures · 50% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
1.0 | 1 | 2026 | Request-Only Optimization for Recommendation Systems · SIGIR 2026 |
Recommender systems
large-scale recommendation |
1.0 | 1 | 2026 | Request-Only Optimization for Recommendation Systems · SIGIR 2026 |
Storage systems › key-value storage
embedding table storage |
1.0 | 1 | 2026 | Request-Only Optimization for Recommendation Systems · SIGIR 2026 |
Indexing and storage engines › vector index
approximate nearest neighbor index |
0.3 | 1 | 2026 | SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs · SIGIR 2026 |
Information retrieval › similarity search
nearest neighbor search |
0.3 | 1 | 2026 | SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
model scaling · 3.0multi-task retrieval · 2.0Int8 ANN kernel · 2.0GPU Bloom index · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Request-Only Optimization for Recommendation SystemsabstractRecommendation systems represent one of the largest machine learning applications on the planet -- industry-scale recommendation models are trained with petabytes of data and serve billions of users every day. To utilize the rich user signals in the long user history, these models have been scaled up to unprecedented complexity, up to trillions of floating-point operations (TFLOPs) per example. This scale, coupled with the huge amount of training data, necessitates new storage and training algorithms to efficiently improve the quality of these complex recommendation systems. Lucy Liao, Huihui Cheng, Yanzun Huang, Keke Zhai, Pengchao Wang, Timothy Shi, Xuan Cao, Renqin Cai, Zhaojie Gong, Omkar Vichare, Rui Jian, Leon Gao, Shiyan Deng, Wenlei Xie, Jiaqi Zhai |
SIGIR | 9 |
| 2026 | SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUsabstractServing deep learning based recommendation models (DLRM) at scale is challenging. Existing approaches rely on dedicated ANN indexing and filtering services on CPUs, suffering from non-negligible costs and missing co-design opportunities. Such inefficiency makes them difficult to support complex model architectures, such as learned similarities and multi-task retrieval. In this paper, we present SilverTorch, a model-based serving system that brings all components into one unified model. It unifies model serving by replacing standalone indexing and filtering services with model layers. We propose a model-based GPU Bloom index for feature filtering and a fused Int8 ANN kernel for nearest neighbor search. Through co-design of the ANN search and feature filtering, we reduce GPU memory usage and eliminate computation. Benefiting from this design, we scale up retrieval by introducing an OverArch scoring layer and a multi-task retrieval with a Value Model to aggregate scores. These advancements improve the retrieval accuracy and enable future studies for serving more complex models. Our evaluation on industry-scale datasets shows that SilverTorch achieves up to 23.7× higher throughput compared to the state-of-the-art approaches. We also demonstrate that SilverTorch's solution is 13.35× more cost-efficient than CPU-based solution while improving accuracy via serving more complex models. Bi Xue, Xiaoheng Mao, Xialu Li, Rui Jian, Yanli Zhao, Yanzun Huang, Yijie Deng, Harry Tran, Ryan Chang, Eric Dong, Jiazhou Wang, Keke Zhai, Hongzhang Yin, Pawel Garbacki, Zheng Fang 0009, Yiyi Pan, Min Ni |
SIGIR | 23 |
| 2020 | Batched Small Tensor-Matrix Multiplications on GPUsabstractWe present a fine-tuned library, ZTMM, for batched small tensor-matrix multiplication on GPU architectures. Libraries performing optimized matrix-matrix multiplications involving large matrices are available for many architectures, including a GPU. However, these libraries do not provide optimal performance for applications requiring efficient multiplication of a matrix with a batch of small matrices or tensors. There has been recent interest in developing fine-tuned libraries for batched small matrix-matrix multiplication - these efforts are limited to square matrices. ZTMM supports both square and rectangular matrices. We experimentally demonstrate that our library has significantly higher performance than cuBLAS and Magma libraries. We demonstrate our library's use on a spectral element-based solver called CMT-nek that performs high-fidelity predictive simulations using compressible Navier-Stokes equations. CMT-nek involves three-dimensional tensors, but it is possible to apply the same techniques to higher dimensional tensors. Keke Zhai, Tania Banerjee, Adeesha Wijayasiri, Sanjay Ranka |
HiPC | 1 |
| 2020 | SparsePipe: Parallel Deep Learning for 3D Point CloudsabstractWe propose SparsePipe, an efficient and asynchronous parallelism approach for handling 3D point clouds with multi-GPU training. SparsePipe is built to support 3D sparse data such as point clouds. It achieves this by adopting generalized convolutions with sparse tensor representation to build expressive high-dimensional convolutional neural networks. Compared to dense solutions, the new models can efficiently process irregular point clouds without densely sliding over the entire space, significantly reducing the memory requirements and allowing higher resolutions of the underlying 3D volumes for better performance. SparsePipe exploits intra-batch parallelism that partitions input data into multiple processors and further improves the training throughput with inter-batch pipelining to overlap communication and computing. Besides, it suitably partitions the model when the GPUs are heterogeneous such that the computing is load-balanced with reduced communication overhead. Using experimental results on an eight-GPU platform, we show that SparsePipe can parallelize effectively and obtain better performance on current point cloud benchmarks for both training and inference, compared to its dense solutions. Keke Zhai, Pan He, Tania Banerjee, Anand Rangarajan 0001, Sanjay Ranka |
HiPC | 1 |
| 2020 | Dynamic load balancing for a mesh-based scientific applicationabstractSummary CMT‐nek is a new scientific application for performing high fidelity predictive simulations of particle‐laden, explosively dispersed turbulent flows. CMT‐nek is compute‐intensive and targeted for deployment on exascale platforms. The moving particles are the primary source of load imbalance when the application is executed on parallel processors. In a demonstration problem, all the particles are initially in a closed container until a detonation occurs and the particles move apart. If all processors get an equal share of the fluid domain, then only some of the processors get sections of the domain that are initially laden with particles, leading to disparate loads on the processors. To eliminate load imbalance in different processors and to speed up the makespan, we present different load‐balancing algorithms for CMT‐nek on large‐scale multicore platforms. The load on a processor is determined using different techniques. The performance of the different load‐balancing algorithms is compared, and the associated overheads are analyzed. Evaluations of the application with and without load‐balancing are conducted, and these show that with load‐balancing, simulation time becomes faster by a factor of up to 9.97. The performance was further improved by a factor of up to 1.42 using machine‐learning–based algorithms. Keke Zhai, Tania Banerjee, David Zwick, Jason Hackl, Rahul Koneru, Sanjay Ranka |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Dynamic Load Balancing for Compressible Multiphase TurbulenceabstractCMT-nek is a new scientific application for performing high fidelity predictive simulations of particle laden explosively dispersed turbulent flows. CMT-nek involves detailed simulations, is compute intensive and is targeted to be deployed on exascale platforms. The moving particles are the main source of load imbalance as the application is executed on parallel processors. In a demonstration problem, all the particles are initially in a closed container until a detonation occurs and the particles move apart. If all processors get an equal share of the fluid domain, then only some of the processors get sections of the domain that are initially laden with particles, leading to disparate load on the processors. In order to eliminate load imbalance in different processors and to speedup the makespan, we present different load balancing algorithms for CMT-nek on large scale multicore platforms consisting of hundred of thousands of cores. The detailed process of the load balancing algorithms are presented. The performance of the different load balancing algorithms are compared and the associated overheads are analyzed. Evaluations on the application with and without load balancing are conducted and these show that with load balancing, simulation time becomes faster by a factor of up to 9.97. Keke Zhai, Tania Banerjee, David Zwick, Jason Hackl, Sanjay Ranka |
ICS | 1 |
| 2016 | CP-activated WASD neuronet approach to Asian population prediction with abundant experimental verification
Yunong Zhang, Dongsheng Guo 0001, Ziyi Luo, Keke Zhai, Hongzhou Tan |
Neurocomputing | 4 |