EDBT 2026 Demo / reviewers in the wild / expert
Jiaqi Zhai
dblp:95/9726
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ParetoES: Hardware-Accelerated Sparse Embedding Similarity via Pareto-Optimal Pruning
Jiaqi Zhai, Xuanhua Shi, Wenju Zhao, Chencheng Ye 0001, Shunsen Lv, Zhongtian Long, Bingsheng He, Hai Jin 0001 |
ISCA | 1 |
| 2026 | Request-Only Optimization for Recommendation SystemsabstractRecommendation systems represent one of the largest machine learning applications on the planet -- industry-scale recommendation models are trained with petabytes of data and serve billions of users every day. To utilize the rich user signals in the long user history, these models have been scaled up to unprecedented complexity, up to trillions of floating-point operations (TFLOPs) per example. This scale, coupled with the huge amount of training data, necessitates new storage and training algorithms to efficiently improve the quality of these complex recommendation systems. Lucy Liao, Huihui Cheng, Yanzun Huang, Keke Zhai, Pengchao Wang, Timothy Shi, Xuan Cao, Renqin Cai, Zhaojie Gong, Omkar Vichare, Rui Jian, Leon Gao, Shiyan Deng, Wenlei Xie, Jiaqi Zhai |
SIGIR | 28 |
| 2025 | AccelES: Accelerating Top-K SpMV for Embedding Similarity via Low-bit PruningabstractIn the realm of recommendation systems, achieving real-time performance in embedding similarity tasks is often hindered by the limitations of traditional Top-K sparse matrix-vector multiplication (SpMV) methods, which suffer from high latency due to inefficient memory access patterns. This paper identifies these critical gaps and introduces AccelES, a novel approach that significantly enhances the efficiency of Top-K SpMV. Our method employs a two-stage calculation scheme: the first stage utilizes a compact, low-bit dataset to quickly identify the most relevant entries, while the second stage performs full-precision calculations solely on this pruned subset, thereby minimizing computational overhead. Furthermore, AccelES incorporates innovative matrix representations, Ultra-CSR and Random-CSR, which optimize memory bandwidth utilization. Experimental results demonstrate that AccelES accelerates performance, surpassing state-of-the-art FPGA, GPU, and CPU solutions by factors of 3.4×, 2.5×, and 153.3×, respectively, under controlled conditions. These advancements not only enhance processing speed but also significantly improve real-time performance in recommendation systems, establishing AccelES as a pivotal contribution to the field of Top-K sparse matrix-vector multiplication. Jiaqi Zhai, Xuanhua Shi, Chencheng Ye 0001, Weifang Hu, Bingsheng He, Hai Jin 0001 |
HPCA | 1 |
| 2025 | Retrieval with Learned SimilaritiesabstractRetrieval plays a fundamental role in recommendation systems, search, and natural language processing (NLP) by efficiently finding relevant items from a large corpus given a query. Dot products have been widely used as the similarity function in such tasks, enabled by Maximum Inner Product Search (MIPS) algorithms for efficient retrieval. However, state-of-the-art retrieval algorithms have migrated to learned similarities. These advanced approaches encompass multiple query embeddings, complex neural networks, direct item ID decoding via beam search, and hybrid solutions. Unfortunately, we lack efficient solutions for retrieval in these state-of-the-art setups. Our work addresses this gap by investigating efficient retrieval techniques with expressive learned similarity functions. We establish Mixture-of-Logits (MoL) as a universal approximator of similarity functions, demonstrate that MoL's expressiveness can be realized empirically to achieve superior performance on diverse retrieval scenarios, and propose techniques to retrieve the approximate top-k results using MoL with tight error bounds. Through extensive experimentation, we show that MoL, enhanced by our proposed mutual information-based load balancing loss, sets new state-of-the-art results across heterogeneous scenarios, including sequential retrieval models in recommendation systems and finetuning language models for question answering; and our approximate top-k algorithms outperform baselines by up to 66× in latency while achieving >.99 recall rate compared to exact algorithms. Bailu Ding, Jiaqi Zhai |
WWW | 2 |
| 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsabstractLarge-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge volume of data with thousands of features, most Deep Learning Recommendation Models (DLRMs) in industry fail to scale with compute. Inspired by success achieved by Transformers in language and vision domains, we revisit fundamental design choices in recommendation systems. We reformulate recommendation problems as sequential transduction tasks within a generative modeling framework (``Generative Recommenders''), and propose a new architecture, HSTU, designed for high cardinality, non-stationary streaming recommendation data. HSTU outperforms baselines over synthetic and public datasets by up to 65.8% in NDCG, and is 5.3x to 15.2x faster than FlashAttention2-based Transformers on 8192 length sequences. HSTU-based Generative Recommenders, with 1.5 trillion parameters, improve metrics in online A/B tests by 12.4% and have been deployed on multiple surfaces of a large internet platform with billions of users. More importantly, the model quality of Generative Recommenders empirically scales as a power-law of training compute across three orders of magnitude, up to GPT-3/LLaMa-2 scale, which reduces carbon footprint needed for future model developments, and further paves the way for the first foundation models in recommendations. Jiaqi Zhai, Lucy Liao, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He 0008, Yinghai Lu |
ICML | 1 |
| 2024 | Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash AttentionabstractThe integration of hardware accelerators has significantly advanced the capabilities of modern recommendation systems, enabling the exploration of complex ranking paradigms previously deemed impractical. However, the GPU-based computational costs present substantial challenges. In this paper, we demonstrate our development of an efficiency-driven approach to explore these paradigms, moving beyond traditional reliance on native PyTorch modules. We address the specific challenges posed by ranking models’ dependence on categorical features, which vary in length and complicate GPU utilization. We introduce Jagged Feature Interaction Kernels, a novel method designed to extract fine-grained insights from long categorical features through efficient handling of dynamically sized tensors. We further enhance the performance of attention mechanisms by integrating Jagged tensors with Flash Attention. Our novel Jagged Flash Attention achieves up to 9 × speedup and 22 × memory reduction compared to dense attention. Notably, it also outperforms dense flash attention, with up to 3 × speedup and 53% more memory efficiency. In production models, we observe 10% QPS improvement and 18% memory savings, enabling us to scale our recommendation systems with longer features and more complex architectures. Rengan Xu, Junjie Yang 0005, Yifan Xu 0035, Devashish Shankar, Haoci Zhang, Yuxi Hu 0001, Mingwei Tang, Zehua Zhang 0004, Tunhou Zhang, Dai Li, Gian-Paolo Musumeci, Jiaqi Zhai, Bill Zhu, Hong Yan 0011, Srihari Reddy |
RecSys | 17 |
| 2023 | Revisiting Neural Retrieval on AcceleratorsabstractRetrieval finds a small number of relevant candidates from a large corpus for information retrieval and recommendation applications. A key component of retrieval is to model (user, item) similarity, which is commonly represented as the dot product of two learned embeddings. This formulation permits efficient inference, commonly known as Maximum Inner Product Search (MIPS). Despite its popularity, dot products cannot capture complex user-item interactions, which are multifaceted and likely high rank. We hence examine non-dot-product retrieval settings on accelerators, and propose mixture of logits (MoL), which models (user, item) similarity as an adaptive composition of elementary similarity functions. This new formulation is expressive, capable of modeling high rank (user, item) interactions, and further generalizes to the long tail. When combined with a hierarchical retrieval strategy, h-indexer, we are able to scale up MoL to 100M corpus on a single GPU with latency comparable to MIPS baselines. On public datasets, our approach leads to uplifts of up to 77.3% in hit rate (HR). Experiments on a large recommendation surface at Meta showed strong metric gains and reduced popularity bias, validating the proposed approach's performance and improved generalization. Jiaqi Zhai, Zhaojie Gong, Xiao Sun 0013, Zheng Yan 0007 |
KDD | 1 |
| 2020 | Generating Representative Headlines for News StoriesabstractMillions of news articles are published online every day, which can be overwhelming for readers to follow. Grouping articles that are reporting the same event into news stories is a common way of assisting readers in their news consumption. However, it remains a challenging research problem to efficiently and effectively generate a representative headline for each story. Automatic summarization of a document set has been studied for decades, while few studies have focused on generating representative headlines for a set of articles. Unlike summaries, which aim to capture most information with least redundancy, headlines aim to capture information jointly shared by the story articles in short length and exclude information specific to each individual article. Xiaotao Gu, Yuning Mao, Jiawei Han 0001, You Wu 0001, Cong Yu 0001, Daniel Finnie, Hongkun Yu 0001, Jiaqi Zhai, Nicholas Zukoski |
WWW | 9 |
| 2011 | ATLAS: a probabilistic algorithm for high dimensional similarity searchabstractGiven a set of high dimensional binary vectors and a similarity function (such as Jaccard and Cosine), we study the problem of finding all pairs of vectors whose similarity exceeds a given threshold. The solution to this problem is a key component in many applications with feature-rich objects, such as text, images, music, videos, or social networks. In particular, there are many important emerging applications that require the use of relatively low similarity thresholds. Jiaqi Zhai, Yin Lou, Johannes Gehrke |
SIGMOD Conference | 1 |