Zhaojie Gong

dblp:348/9701 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0004-1761-7530ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Recommender systems · 56% Information retrieval · 44%
Artificial intelligence
3 papers
Efficient and distributed learning · 61% Deep learning architectures and training · 39%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.012026
Request-Only Optimization for Recommendation Systems · SIGIR 2026
Recommender systems
large-scale recommendation
1.012026
Request-Only Optimization for Recommendation Systems · SIGIR 2026
Storage systems › key-value storage
embedding table storage
1.012026
Request-Only Optimization for Recommendation Systems · SIGIR 2026
Machine learning › Deep learning architectures and training
transformer
0.812024
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024
Recommender systems
generative recommendation
0.812024
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024
Recommender systems
sequential recommendation
0.812024
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024
Information retrieval › document retrieval › structure-aware retrieval
hierarchical retrieval
0.712023
Revisiting Neural Retrieval on Accelerators · KDD 2023
Information retrieval › similarity search › nearest neighbor search
maximum inner product search
0.712023
Revisiting Neural Retrieval on Accelerators · KDD 2023
Information retrieval › retrieval models
neural retrieval
0.712023
Revisiting Neural Retrieval on Accelerators · KDD 2023

Methods — techniques the papers use, named apart from their topics

model scaling · 3.0scaling laws · 1.5generative modeling · 1.5mixture of logits · 1.3hierarchical indexing · 1.3
YearPublicationVenuePosition
2026 Request-Only Optimization for Recommendation Systems
abstract
Recommendation systems represent one of the largest machine learning applications on the planet -- industry-scale recommendation models are trained with petabytes of data and serve billions of users every day. To utilize the rich user signals in the long user history, these models have been scaled up to unprecedented complexity, up to trillions of floating-point operations (TFLOPs) per example. This scale, coupled with the huge amount of training data, necessitates new storage and training algorithms to efficiently improve the quality of these complex recommendation systems.
Lucy Liao, Huihui Cheng, Yanzun Huang, Keke Zhai, Pengchao Wang, Timothy Shi, Xuan Cao, Renqin Cai, Zhaojie Gong, Omkar Vichare, Rui Jian, Leon Gao, Shiyan Deng, Wenlei Xie, Jiaqi Zhai
SIGIR15
2025 Edge Classification on Imbalanced Multi-relational Graphs
Zhaojie Gong, Yijun Duan, Qiang Ma 0001
ADMA (4)1
2024 Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
abstract
Large-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge volume of data with thousands of features, most Deep Learning Recommendation Models (DLRMs) in industry fail to scale with compute. Inspired by success achieved by Transformers in language and vision domains, we revisit fundamental design choices in recommendation systems. We reformulate recommendation problems as sequential transduction tasks within a generative modeling framework (``Generative Recommenders''), and propose a new architecture, HSTU, designed for high cardinality, non-stationary streaming recommendation data. HSTU outperforms baselines over synthetic and public datasets by up to 65.8% in NDCG, and is 5.3x to 15.2x faster than FlashAttention2-based Transformers on 8192 length sequences. HSTU-based Generative Recommenders, with 1.5 trillion parameters, improve metrics in online A/B tests by 12.4% and have been deployed on multiple surfaces of a large internet platform with billions of users. More importantly, the model quality of Generative Recommenders empirically scales as a power-law of training compute across three orders of magnitude, up to GPT-3/LLaMa-2 scale, which reduces carbon footprint needed for future model developments, and further paves the way for the first foundation models in recommendations.
Jiaqi Zhai, Lucy Liao, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He 0008, Yinghai Lu
ICML8
2023 Revisiting Neural Retrieval on Accelerators
abstract
Retrieval finds a small number of relevant candidates from a large corpus for information retrieval and recommendation applications. A key component of retrieval is to model (user, item) similarity, which is commonly represented as the dot product of two learned embeddings. This formulation permits efficient inference, commonly known as Maximum Inner Product Search (MIPS). Despite its popularity, dot products cannot capture complex user-item interactions, which are multifaceted and likely high rank. We hence examine non-dot-product retrieval settings on accelerators, and propose mixture of logits (MoL), which models (user, item) similarity as an adaptive composition of elementary similarity functions. This new formulation is expressive, capable of modeling high rank (user, item) interactions, and further generalizes to the long tail. When combined with a hierarchical retrieval strategy, h-indexer, we are able to scale up MoL to 100M corpus on a single GPU with latency comparable to MIPS baselines. On public datasets, our approach leads to uplifts of up to 77.3% in hit rate (HR). Experiments on a large recommendation surface at Meta showed strong metric gains and reduced popularity bias, validating the proposed approach's performance and improved generalization.
Jiaqi Zhai, Zhaojie Gong, Xiao Sun 0013, Zheng Yan 0007
KDD2