Yanzun Huang

dblp:244/7049 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0009-1238-0454ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Recommender systems · 77% Information retrieval · 12% Indexing and storage engines · 12%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 50% Hardware accelerators and domain-specific architectures · 50%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.012026
Request-Only Optimization for Recommendation Systems · SIGIR 2026
Recommender systems
large-scale recommendation
1.012026
Request-Only Optimization for Recommendation Systems · SIGIR 2026
Storage systems › key-value storage
embedding table storage
1.012026
Request-Only Optimization for Recommendation Systems · SIGIR 2026
Indexing and storage engines › vector index
approximate nearest neighbor index
0.312026
SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs · SIGIR 2026
Information retrieval › similarity search
nearest neighbor search
0.312026
SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs · SIGIR 2026

Methods — techniques the papers use, named apart from their topics

model scaling · 3.0multi-task retrieval · 2.0Int8 ANN kernel · 2.0GPU Bloom index · 2.0
YearPublicationVenuePosition
2026 Request-Only Optimization for Recommendation Systems
abstract
Recommendation systems represent one of the largest machine learning applications on the planet -- industry-scale recommendation models are trained with petabytes of data and serve billions of users every day. To utilize the rich user signals in the long user history, these models have been scaled up to unprecedented complexity, up to trillions of floating-point operations (TFLOPs) per example. This scale, coupled with the huge amount of training data, necessitates new storage and training algorithms to efficiently improve the quality of these complex recommendation systems.
Lucy Liao, Huihui Cheng, Yanzun Huang, Keke Zhai, Pengchao Wang, Timothy Shi, Xuan Cao, Renqin Cai, Zhaojie Gong, Omkar Vichare, Rui Jian, Leon Gao, Shiyan Deng, Wenlei Xie, Jiaqi Zhai
SIGIR8
2026 SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs
abstract
Serving deep learning based recommendation models (DLRM) at scale is challenging. Existing approaches rely on dedicated ANN indexing and filtering services on CPUs, suffering from non-negligible costs and missing co-design opportunities. Such inefficiency makes them difficult to support complex model architectures, such as learned similarities and multi-task retrieval. In this paper, we present SilverTorch, a model-based serving system that brings all components into one unified model. It unifies model serving by replacing standalone indexing and filtering services with model layers. We propose a model-based GPU Bloom index for feature filtering and a fused Int8 ANN kernel for nearest neighbor search. Through co-design of the ANN search and feature filtering, we reduce GPU memory usage and eliminate computation. Benefiting from this design, we scale up retrieval by introducing an OverArch scoring layer and a multi-task retrieval with a Value Model to aggregate scores. These advancements improve the retrieval accuracy and enable future studies for serving more complex models. Our evaluation on industry-scale datasets shows that SilverTorch achieves up to 23.7× higher throughput compared to the state-of-the-art approaches. We also demonstrate that SilverTorch's solution is 13.35× more cost-efficient than CPU-based solution while improving accuracy via serving more complex models.
Bi Xue, Xiaoheng Mao, Xialu Li, Rui Jian, Yanli Zhao, Yanzun Huang, Yijie Deng, Harry Tran, Ryan Chang, Eric Dong, Jiazhou Wang, Keke Zhai, Hongzhang Yin, Pawel Garbacki, Zheng Fang 0009, Yiyi Pan, Min Ni
SIGIR15
2019 Early Detection of Wheel Spinning: Comparison across Tutors, Models, Features, and Operationalizations
Chuankai Zhang, Yanzun Huang, Dongyang Lu, Weiqi Fang, John C. Stamper, Stephen Fancsali, Kenneth Holstein, Vincent Aleven
EDM2