VLDB 2026 Research / reviewers in the wild / expert
Yuxi Hu 0001
dblp:17/8327-1
· DBLP profile ↗
5ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0005-6608-8441ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIITED: Fine-Grained Embedding Dimension Optimization During Training for Recommender SystemsabstractHuge embedding tables in modern deep learning recommender models (DLRM) require prohibitively large memory during training and inference. This paper proposes FIITED, a system to automatically reduce the memory footprint via FIne-grained In-Training Embedding Dimension pruning. By leveraging the key insight that embedding vectors are not equally important, FIITED adaptively adjusts the dimension of each individual embedding vector during model training, assigning larger dimensions to more important embeddings while adapting to dynamic changes in data. We prioritize embedding dimensions with higher frequencies and gradients as more important. To enable efficient pruning of embeddings and their dimensions during model training, we propose an embedding storage system based on virtually-hashed physically-indexed hash tables. Experiments on two industry models and months of realistic datasets show that FIITED can reduce DLRM embedding size by more than 65’ while preserving model quality, outperforming state-of-the-art in-training embedding pruning methods and increasing the reduction ratio by 1.3× to 1.67×. On public datasets, FIITED can reduce the size of embedding tables by 2.2× to 800× with negligible accuracy drop, achieving 1.05× to 12.5× improvement in reduction ratio compared to baselines while improving model throughput. Qinyi Luo, Penghan Wang, Wei Zhang 0044, Fan Lai 0001, Jiachen Mao, Xiaohan Wei, Wei-Yu Tsai, Yuxi Hu 0001, Xuehai Qian |
IEEE Trans. Computers | 10 |
| 2024 | Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash AttentionabstractThe integration of hardware accelerators has significantly advanced the capabilities of modern recommendation systems, enabling the exploration of complex ranking paradigms previously deemed impractical. However, the GPU-based computational costs present substantial challenges. In this paper, we demonstrate our development of an efficiency-driven approach to explore these paradigms, moving beyond traditional reliance on native PyTorch modules. We address the specific challenges posed by ranking models’ dependence on categorical features, which vary in length and complicate GPU utilization. We introduce Jagged Feature Interaction Kernels, a novel method designed to extract fine-grained insights from long categorical features through efficient handling of dynamically sized tensors. We further enhance the performance of attention mechanisms by integrating Jagged tensors with Flash Attention. Our novel Jagged Flash Attention achieves up to 9 × speedup and 22 × memory reduction compared to dense attention. Notably, it also outperforms dense flash attention, with up to 3 × speedup and 53% more memory efficiency. In production models, we observe 10% QPS improvement and 18% memory savings, enabling us to scale our recommendation systems with longer features and more complex architectures. Rengan Xu, Junjie Yang 0005, Yifan Xu 0035, Devashish Shankar, Haoci Zhang, Yuxi Hu 0001, Mingwei Tang, Zehua Zhang 0004, Tunhou Zhang, Dai Li, Gian-Paolo Musumeci, Jiaqi Zhai, Bill Zhu, Hong Yan 0011, Srihari Reddy |
RecSys | 10 |
| 2023 | AdaEmbed: Adaptive Embedding for Large-Scale Recommendation Models
Fan Lai 0001, Wei Zhang 0044, William Tsai, Xiaohan Wei, Yuxi Hu 0001, Sabin Devkota, Jongsoo Park, Zeliang Chen, Ellie Wen, Paul Rivera, Chun-cheng Jason Chen, Mosharaf Chowdhury |
OSDI | 6 |
| 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation modelsabstractDeep learning recommendation models (DLRMs) have been used across many business-critical services at Meta and are the single largest AI application in terms of infrastructure demand in its data-centers. In this paper, we present Neo, a software-hardware co-designed system for high-performance distributed training of large-scale DLRMs. Neo employs a novel 4D parallelism strategy that combines table-wise, row-wise, column-wise, and data parallelism for training massive embedding operators in DLRMs. In addition, Neo enables extremely high-performance and memory-efficient embedding computations using a variety of critical systems optimizations, including hybrid kernel fusion, software-managed caching, and quality-preserving compression. Finally, Neo is paired with ZionEX, a new hardware platform co-designed with Neo's 4D parallelism for optimizing communications for large-scale DLRM training. Our evaluation on 128 GPUs using 16 ZionEX nodes shows that Neo outperforms existing systems by up to 40× for training 12-trillion-parameter DLRM models deployed in production. Dheevatsa Mudigere, Yuchen Hao, Andrew Tulloch, Srinivas Sridharan 0002, Muhammet Mustafa Ozdal, Jade Nie, Jongsoo Park, Jie Amy Yang, Leon Gao, Dmytro Ivchenko, Aarti Basant, Yuxi Hu 0001, Jiyan Yang, Ehsan K. Ardestani, Xiaodong Wang 0020, Rakesh Komuravelli, Ching-Hsiang Chu, Serhat Yilmaz, Jiyuan Qian, Zhuobo Feng, Yinbin Ma, Junjie Yang 0005, Ellie Wen, Chonglin Sun, Whitney Zhao, Dimitry Melts, Krishna Dhulipala, K. R. Kishore, Tyler Graf, Assaf Eisenman, Kiran Kumar Matam, Adi Gangidi, Guoqiang Jerry Chen, Manoj Krishnan, Avinash Nayak, Krishnakumar Nair, Bharath Muthiah, Mahmoud khorashadi, Pallab Bhattacharya, Petr Lapukhov, Maxim Naumov, Ajit Mathews, Lin Qiao, Mikhail Smelyanskiy, Bill Jia, Vijay Rao |
ISCA | 16 |
| 2012 | On a Triadic Approach to Connect Microstructural Properties to Social Macrostructural PatternsabstractSocial macrostructures, such as structural balance, ranked clusters and transitivity, are of great importance on account of their abilities to reflect the underlying social psychological processes about the formation and evolution of relationships among people. Here we present a detailed study on examining the existence and evolution of social macrostructures in an empirical online social network, and exploring how they can be explained by network micro structural properties, i.e. nodal in degree and out degree and dyadic feature. We establish the micro-macro linkage by analyzing the network triadic patterns. Based on a novel clustering coefficient based network sampling approach, we show that the distribution of observed triad census in our data is low dimensional and can be greatly explained by network dyadic properties. In a time series analysis, we observe that our network exhibits strong tendencies towards balanced, transitive and clustered social macrostructure given the nodal and dyadic characteristics. Our findings supplement the studies on structural properties of online social network by providing more insights on the relation between network macrostructures and the micro-level social processes that result in them. And they form the basis to understand better how online social media systems change the information and communication fabric of our society. Yuxi Hu 0001, Mina Doroud, Shyhtsun Felix Wu |
ASONAM | 1 |