Pallab Bhattacharya

dblp:178/3027 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0005-7653-744XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
abstract
Characterizing and predicting the training performance of modern machine learning (ML) workloads on compute systems with compute and communication spread between CPUs, GPUs, and network devices is not only the key to optimization and planning but also a complex goal to achieve. The primary challenges include the complexity of synchronization and load balancing between CPUs and GPUs, the variance in input data distribution, and the use of different communication devices and topologies (e.g., NVLink, PCIe, network cards) that connect multiple compute devices, coupled with the desire for flexible training configurations. Built on top of our prior work for single-GPU platforms, we address these challenges and enable multi-GPU performance modeling1by incorporating (1) data-distribution-aware performance models for embedding table lookup, and (2) data movement prediction of communication collectives, into our upgraded performance modeling pipeline equipped with inter-and intra-rank synchronization for ML workloads trained on multi-GPU platforms. Beyond accurately predicting the per-iteration training time of deep learning recommendation models (DLRM) models with random configurations with a geomean error of 5.21% on two multi-GPU platforms, our prediction pipeline generalizes well to other types of ML workloads, such as Transformer-based natural language processing (NLP) models with a geomean error of 3.00%. Moreover, even without actually running ML workloads like DLRMs on the hardware, it is capable of generating insights such as quickly selecting the fastest embedding table sharding configuration (with a success rate of 85%).
Zhongyi Lin, Pallab Bhattacharya, Xizhou Feng, Louis Feng, John D. Owens
IEEE Trans. Parallel Distributed Syst.3
2023 TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory
abstract
The increasing demand for memory in hyperscale applications has led to memory becoming a large portion of the overall datacenter spend. The emergence of coherent interfaces like CXL enables main memory expansion and offers an efficient solution to this problem. In such systems, the main memory can constitute different memory technologies with varied characteristics. In this paper, we characterize memory usage patterns of a wide range of datacenter applications across the server fleet of Meta. We, therefore, demonstrate the opportunities to offload colder pages to slower memory tiers for these applications. Without efficient memory management, however, such systems can significantly degrade performance.
Hasan Al Maruf, Hao Wang 0011, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen 0002, Mosharaf Chowdhury, Shobhit O. Kanaujia, Prakash Chauhan
ASPLOS (3)6
2022 Software-hardware co-design for fast and scalable training of deep learning recommendation models
abstract
Deep learning recommendation models (DLRMs) have been used across many business-critical services at Meta and are the single largest AI application in terms of infrastructure demand in its data-centers. In this paper, we present Neo, a software-hardware co-designed system for high-performance distributed training of large-scale DLRMs. Neo employs a novel 4D parallelism strategy that combines table-wise, row-wise, column-wise, and data parallelism for training massive embedding operators in DLRMs. In addition, Neo enables extremely high-performance and memory-efficient embedding computations using a variety of critical systems optimizations, including hybrid kernel fusion, software-managed caching, and quality-preserving compression. Finally, Neo is paired with ZionEX, a new hardware platform co-designed with Neo's 4D parallelism for optimizing communications for large-scale DLRM training. Our evaluation on 128 GPUs using 16 ZionEX nodes shows that Neo outperforms existing systems by up to 40× for training 12-trillion-parameter DLRM models deployed in production.
Dheevatsa Mudigere, Yuchen Hao, Andrew Tulloch, Srinivas Sridharan 0002, Muhammet Mustafa Ozdal, Jade Nie, Jongsoo Park, Jie Amy Yang, Leon Gao, Dmytro Ivchenko, Aarti Basant, Yuxi Hu 0001, Jiyan Yang, Ehsan K. Ardestani, Xiaodong Wang 0020, Rakesh Komuravelli, Ching-Hsiang Chu, Serhat Yilmaz, Jiyuan Qian, Zhuobo Feng, Yinbin Ma, Junjie Yang 0005, Ellie Wen, Chonglin Sun, Whitney Zhao, Dimitry Melts, Krishna Dhulipala, K. R. Kishore, Tyler Graf, Assaf Eisenman, Kiran Kumar Matam, Adi Gangidi, Guoqiang Jerry Chen, Manoj Krishnan, Avinash Nayak, Krishnakumar Nair, Bharath Muthiah, Mahmoud khorashadi, Pallab Bhattacharya, Petr Lapukhov, Maxim Naumov, Ajit Mathews, Lin Qiao, Mikhail Smelyanskiy, Bill Jia, Vijay Rao
ISCA46
2016 End-to-End Java Security Performance Enhancements for Oracle SPARC Servers
abstract
In this paper we investigate the performance of cryptographic operations, when used in Java applications. We demonstrate the advantage of using built-in hardware accelerator for cryptographic operations on SPARC servers. In particular, we demonstrate the advantage of hardware cryptographic instructions invoked via AES and SHA intrinsics, implemented in the Java Virtual Machine (JVM), over the more traditional Java Native Interface (JNI) calls. For the purpose of our study, we modified the SPECweb2005 benchmark by adding modern banking requirements, and created a new workload which we call the End-to-End Java Security (EEJS) workload. Using the workload, we compare different Java Cryptographic Service Providers (CSPs) and arrive at the conclusion that hardware cryptography has significant performance advantage for Java applications. With the EEJS workload, we also identify several enhancements applicable to the Java Secure Socket Extension (JSSE).
Pallab Bhattacharya, Yao-Min Chen, Shrinivas Joshi, James Cheng
ICPE2
2007 Special Issue on Optoelectronic Devices Based on Quantum Dots
abstract
The twelve articles in this special issue are devoted to optoelectronc device manufacture based on quantum dot technology.
Pallab Bhattacharya, Dieter Bimberg, Yasuhiko Arakawa
Proc. IEEE1