VLDB 2026 Research / reviewers in the wild / expert
Bharath Muthiah
dblp:286/0894
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 100% | |
| Databases, data mining, and information retrieval
3 papers |
Machine learning and data management · 34% Query processing and optimization · 34% Recommender systems · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 87% GPUs and heterogeneous computing · 13% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management › data management for machine learning
embedding table management |
0.8 | 1 | 2024 | OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model · USENIX ATC 2024 |
Query processing and optimization
parallel query processing |
0.8 | 1 | 2024 | OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model · USENIX ATC 2024 |
Machine learning › Efficient and distributed learning
distributed training |
0.6 | 1 | 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022 |
Machine learning › Efficient and distributed learning › model compression › embedding compression
embedding table compression |
0.6 | 1 | 2022 | EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding Table · SC 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.6 | 1 | 2022 | EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding Table · SC 2022 |
Recommender systems › large-scale recommendation
large-scale recommendation model training |
0.6 | 1 | 2022 | EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding Table · SC 2022 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.6 | 1 | 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
recommendation model training |
0.6 | 1 | 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022 |
GPUs and heterogeneous computing › GPU memory management
GPU memory optimization |
0.2 | 1 | 2022 | EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding Table · SC 2022 |
Methods — techniques the papers use, named apart from their topics
tensor-train decomposition · 1.7software-managed caching · 1.7pipeline training · 1.7kernel fusion · 1.7index reordering · 1.7embedding compression · 1.7optimality-guided partitioning · 1.5embedding table parallelization · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model
Zheng Wang 0075, Boyuan Feng, Guyue Huang, Dheevatsa Mudigere, Bharath Muthiah, Ang Li 0006, Yufei Ding 0001 |
USENIX ATC | 6 |
| 2022 | Software-hardware co-design for fast and scalable training of deep learning recommendation modelsabstractDeep learning recommendation models (DLRMs) have been used across many business-critical services at Meta and are the single largest AI application in terms of infrastructure demand in its data-centers. In this paper, we present Neo, a software-hardware co-designed system for high-performance distributed training of large-scale DLRMs. Neo employs a novel 4D parallelism strategy that combines table-wise, row-wise, column-wise, and data parallelism for training massive embedding operators in DLRMs. In addition, Neo enables extremely high-performance and memory-efficient embedding computations using a variety of critical systems optimizations, including hybrid kernel fusion, software-managed caching, and quality-preserving compression. Finally, Neo is paired with ZionEX, a new hardware platform co-designed with Neo's 4D parallelism for optimizing communications for large-scale DLRM training. Our evaluation on 128 GPUs using 16 ZionEX nodes shows that Neo outperforms existing systems by up to 40× for training 12-trillion-parameter DLRM models deployed in production. Dheevatsa Mudigere, Yuchen Hao, Andrew Tulloch, Srinivas Sridharan 0002, Muhammet Mustafa Ozdal, Jade Nie, Jongsoo Park, Jie Amy Yang, Leon Gao, Dmytro Ivchenko, Aarti Basant, Yuxi Hu 0001, Jiyan Yang, Ehsan K. Ardestani, Xiaodong Wang 0020, Rakesh Komuravelli, Ching-Hsiang Chu, Serhat Yilmaz, Jiyuan Qian, Zhuobo Feng, Yinbin Ma, Junjie Yang 0005, Ellie Wen, Chonglin Sun, Whitney Zhao, Dimitry Melts, Krishna Dhulipala, K. R. Kishore, Tyler Graf, Assaf Eisenman, Kiran Kumar Matam, Adi Gangidi, Guoqiang Jerry Chen, Manoj Krishnan, Avinash Nayak, Krishnakumar Nair, Bharath Muthiah, Mahmoud khorashadi, Pallab Bhattacharya, Petr Lapukhov, Maxim Naumov, Ajit Mathews, Lin Qiao, Mikhail Smelyanskiy, Bill Jia, Vijay Rao |
ISCA | 44 |
| 2022 | EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding TableabstractDeep learning Recommendation Models (DLRMs) plays an important role in various application domains. However, existing DLRM training systems require a large number of GPUs due to the memory-intensive embedding tables. To this end, we propose EL-Rec, an efficient computing framework harnessing the Tensor-train (TT) technique to democratize the training of large-scale DLRMs with limited GPU resources. Specifically, EL-Rec optimizes TT decomposition based on key computation primitives of embedding tables and implements a high-performance compressed embedding table which is a drop-in replacement of Pytorch API. EL-Rec introduces an index reordering technique to harvest the performance gains from both local and global information of training inputs. EL-Rec also highlights a pipeline training paradigm to eliminate the communication overhead between the host memory and the training worker. Comprehensive experiments demonstrate that EL-Rec can handle the largest publicly available DLRM dataset with a single GPU and achieves 3× speedup over the state-of-the-art DLRM frameworks. Zheng Wang 0075, Boyuan Feng, Dheevatsa Mudigere, Bharath Muthiah, Yufei Ding 0001 |
SC | 5 |