EDBT 2026 Demo / reviewers in the wild / expert
Leon Gao
dblp:290/2159
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 100% | |
| Artificial intelligence
3 papers |
Efficient and distributed learning · 68% Deep learning architectures and training · 32% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 54% Storage systems · 46% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
1.6 | 2 | 2026 | Request-Only Optimization for Recommendation Systems · SIGIR 2026 Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022 |
Recommender systems
large-scale recommendation |
1.0 | 1 | 2026 | Request-Only Optimization for Recommendation Systems · SIGIR 2026 |
Storage systems › key-value storage
embedding table storage |
1.0 | 1 | 2026 | Request-Only Optimization for Recommendation Systems · SIGIR 2026 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024 |
Recommender systems
generative recommendation |
0.8 | 1 | 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024 |
Recommender systems
sequential recommendation |
0.8 | 1 | 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.6 | 1 | 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
recommendation model training |
0.6 | 1 | 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022 |
Methods — techniques the papers use, named apart from their topics
model scaling · 3.0software-managed caching · 1.7kernel fusion · 1.7embedding compression · 1.7scaling laws · 1.5generative modeling · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Request-Only Optimization for Recommendation SystemsabstractRecommendation systems represent one of the largest machine learning applications on the planet -- industry-scale recommendation models are trained with petabytes of data and serve billions of users every day. To utilize the rich user signals in the long user history, these models have been scaled up to unprecedented complexity, up to trillions of floating-point operations (TFLOPs) per example. This scale, coupled with the huge amount of training data, necessitates new storage and training algorithms to efficiently improve the quality of these complex recommendation systems. Lucy Liao, Huihui Cheng, Yanzun Huang, Keke Zhai, Pengchao Wang, Timothy Shi, Xuan Cao, Renqin Cai, Zhaojie Gong, Omkar Vichare, Rui Jian, Leon Gao, Shiyan Deng, Wenlei Xie, Jiaqi Zhai |
SIGIR | 18 |
| 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsabstractLarge-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge volume of data with thousands of features, most Deep Learning Recommendation Models (DLRMs) in industry fail to scale with compute. Inspired by success achieved by Transformers in language and vision domains, we revisit fundamental design choices in recommendation systems. We reformulate recommendation problems as sequential transduction tasks within a generative modeling framework (``Generative Recommenders''), and propose a new architecture, HSTU, designed for high cardinality, non-stationary streaming recommendation data. HSTU outperforms baselines over synthetic and public datasets by up to 65.8% in NDCG, and is 5.3x to 15.2x faster than FlashAttention2-based Transformers on 8192 length sequences. HSTU-based Generative Recommenders, with 1.5 trillion parameters, improve metrics in online A/B tests by 12.4% and have been deployed on multiple surfaces of a large internet platform with billions of users. More importantly, the model quality of Generative Recommenders empirically scales as a power-law of training compute across three orders of magnitude, up to GPT-3/LLaMa-2 scale, which reduces carbon footprint needed for future model developments, and further paves the way for the first foundation models in recommendations. Jiaqi Zhai, Lucy Liao, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He 0008, Yinghai Lu |
ICML | 7 |
| 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation modelsabstractDeep learning recommendation models (DLRMs) have been used across many business-critical services at Meta and are the single largest AI application in terms of infrastructure demand in its data-centers. In this paper, we present Neo, a software-hardware co-designed system for high-performance distributed training of large-scale DLRMs. Neo employs a novel 4D parallelism strategy that combines table-wise, row-wise, column-wise, and data parallelism for training massive embedding operators in DLRMs. In addition, Neo enables extremely high-performance and memory-efficient embedding computations using a variety of critical systems optimizations, including hybrid kernel fusion, software-managed caching, and quality-preserving compression. Finally, Neo is paired with ZionEX, a new hardware platform co-designed with Neo's 4D parallelism for optimizing communications for large-scale DLRM training. Our evaluation on 128 GPUs using 16 ZionEX nodes shows that Neo outperforms existing systems by up to 40× for training 12-trillion-parameter DLRM models deployed in production. Dheevatsa Mudigere, Yuchen Hao, Andrew Tulloch, Srinivas Sridharan 0002, Muhammet Mustafa Ozdal, Jade Nie, Jongsoo Park, Jie Amy Yang, Leon Gao, Dmytro Ivchenko, Aarti Basant, Yuxi Hu 0001, Jiyan Yang, Ehsan K. Ardestani, Xiaodong Wang 0020, Rakesh Komuravelli, Ching-Hsiang Chu, Serhat Yilmaz, Jiyuan Qian, Zhuobo Feng, Yinbin Ma, Junjie Yang 0005, Ellie Wen, Chonglin Sun, Whitney Zhao, Dimitry Melts, Krishna Dhulipala, K. R. Kishore, Tyler Graf, Assaf Eisenman, Kiran Kumar Matam, Adi Gangidi, Guoqiang Jerry Chen, Manoj Krishnan, Avinash Nayak, Krishnakumar Nair, Bharath Muthiah, Mahmoud khorashadi, Pallab Bhattacharya, Petr Lapukhov, Maxim Naumov, Ajit Mathews, Lin Qiao, Mikhail Smelyanskiy, Bill Jia, Vijay Rao |
ISCA | 13 |