EDBT 2026 Demo / reviewers in the wild / expert
Shubho Sengupta
dblp:147/1220
· DBLP profile ↗
10ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0007-4204-5185ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Speech recognition and synthesis · 34% Reinforcement learning · 28% Efficient and distributed learning · 24% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 24% Performance modeling and evaluation · 23% Hardware reliability and fault tolerance · 22% | |
| Network and information security
1 paper |
Cryptographic protocols and secure computation · 33% Privacy and data protection · 33% Systems and software security · 33% | |
| Human-computer interaction and pervasive computing
1 paper |
Games and playful interaction · 100% |
Topics — the 25 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware reliability and fault tolerance
failure analysis |
0.9 | 1 | 2025 | Revisiting Reliability in Large-Scale Machine Learning Research Clusters · HPCA 2025 |
Performance modeling and evaluation
workload characterization |
0.9 | 1 | 2025 | Revisiting Reliability in Large-Scale Machine Learning Research Clusters · HPCA 2025 |
Privacy and data protection
privacy-preserving machine learning |
0.5 | 1 | 2021 | CrypTen: Secure Multi-Party Computation Meets Machine Learning · NeurIPS 2021 |
Cryptographic protocols and secure computation
secure multiparty computation |
0.5 | 1 | 2021 | CrypTen: Secure Multi-Party Computation Meets Machine Learning · NeurIPS 2021 |
Machine learning › Reinforcement learning › deep reinforcement learning
alphazero |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Games and playful interaction
board games |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2017 | Exploring Sparsity in Recurrent Neural Networks · ICLR (Poster) 2017 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural speech synthesis |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Natural language and speech › Speech recognition and synthesis
prosody prediction |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Machine learning › Efficient and distributed learning › model compression
sparsity |
0.3 | 1 | 2017 | Exploring Sparsity in Recurrent Neural Networks · ICLR (Poster) 2017 |
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Distributed systems
fault tolerance |
0.3 | 1 | 2025 | Revisiting Reliability in Large-Scale Machine Learning Research Clusters · HPCA 2025 |
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
GPUs and heterogeneous computing
deep learning on GPUs |
0.2 | 1 | 2016 | Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2016 | Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016 |
Parallel and multicore computing
graph processing |
0.2 | 1 | 2014 | Navigating the maze of graph analytics frameworks using massive graph datasets · SIGMOD Conference 2014 |
Distributed systems
graph processing systems |
0.2 | 1 | 2014 | Navigating the maze of graph analytics frameworks using massive graph datasets · SIGMOD Conference 2014 |
Machine learning › Efficient and distributed learning
distributed training |
0.1 | 1 | 2021 | CrypTen: Secure Multi-Party Computation Meets Machine Learning · NeurIPS 2021 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Performance modeling and evaluation
bottleneck analysis |
0.1 | 1 | 2014 | Navigating the maze of graph analytics frameworks using massive graph datasets · SIGMOD Conference 2014 |
Methods — techniques the papers use, named apart from their topics
deep neural network · 1.0tensor computation · 1.0automatic differentiation · 1.0mean time to failure estimation · 0.9failure modeling · 0.9monte carlo tree search · 0.8secure multiparty computation · 0.5secure multi-party computation · 0.5persistent kernel · 0.5batch dispatch · 0.5GPU-based inference · 0.5wavenet · 0.3connectionist temporal classification · 0.3recurrent neural network · 0.2system-level analysis · 0.2hand-optimized baseline · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Revisiting Reliability in Large-Scale Machine Learning Research ClustersabstractReliability is a fundamental challenge in operating large-scale machine learning (ML) infrastructures, particularly as the scale of ML models and training clusters continues to grow. Despite decades of research on infrastructure failures, the impact of job failures across different scales remains unclear. This paper presents a view of managing two large, multi-tenant ML clusters, providing quantitative analysis, operational experience, and our own perspective in understanding and addressing reliability concerns at scale. Our analysis reveals that while large jobs are most vulnerable to failures, smaller jobs make up the majority of jobs in the clusters and should be incorporated into optimization objectives. We identify key workload properties, compare them across clusters, and demonstrate essential reliability requirements for pushing the boundaries of ML training at scale.We hereby introduce a taxonomy of failures and key reliability metrics, analyze 11 months of data from two state-of-the-art ML environments with 4 million jobs and over 150 million A100 GPU hours. Building on our data, we fit a failure model to project Mean Time to Failure for various GPU scales. We further propose a method to estimate a related metric, Effective Training Time Ratio, as a function of job parameters, and we use this model to gauge the efficacy of potential software mitigations at scale. Our work provides valuable insights and future research directions for improving the reliability of AI supercomputer clusters, emphasizing the need for flexible, workload-agnostic, and reliability-aware infrastructure, system software, and algorithms. Apostolos Kokolis, Michael Kuchnik, John Hoffman, Adithya Kumar, Parth Malani, Faye Ma, Zach DeVito, Shubho Sengupta, Kalyan Saladi, Carole-Jean Wu |
HPCA | 8 |
| 2024 | Delegated Private Matching For ComputeabstractPrivate matching for compute (PMC) establishes a match between two datasets owned by mutually distrusted parties (C and P) and allows the parties to input more data for the matched records for arbitrary downstream secure computation without rerunning the private matching component. The state-of-the-art PMC protocols only support two parties and assume that both parties can participate in computationally intensive secure computation. We observe that such operational overhead limits the adoption of these protocols to solely powerful entities as small data owners or devices with minimal computing power will not be able to participate. We introduce two protocols to delegate PMC from party P to untrusted cloud servers, called delegates, allowing multiple smaller P parties to provide inputs containing identifiers and associated values. Our Delegated Private Matching for Compute protocols, called DPMC and DsPMC, establish a join between the datasets of party C and multiple delegators P based on multiple identifiers and compute secret shares of associated values for the identifiers that the parties have in common. We introduce a rerandomizable encrypted oblivious pseudorandom function (OPRF) primitive, called EO, which allows two parties to encrypt, mask, and shuffle their data. Note that EO may be of independent interest. Our DsPMC protocol limits the leakages of DPMC by combining our EO scheme and secure three-party shuffling. Finally, our implementation demonstrates the efficiency of our constructions by outperforming related works by approximately 10x for the total protocol execution and by at least 20x for the computation on the delegators. Dimitris Mouris, Daniel Masny, Ni Trieu, Shubho Sengupta, Prasad Buddhavarapu, Benjamin M. Case |
Proc. Priv. Enhancing Technol. | 4 |
| 2022 | Parallel Composition of Weighted Finite-State TransducersabstractFinite-state transducers (FSTs) are frequently used in speech recognition. Transducer composition is an essential operation for combining different sources of information at different granularities. However, composition is also one of the more computationally expensive operations. Due to the heterogeneous structure of FSTs, parallel algorithms for composition are suboptimal in efficiency, generality, or both. We propose an algorithm for parallel composition and implement it on graphics processing units. We benchmark our parallel algorithm on the composition of random graphs and the composition of graphs commonly used in speech recognition. The parallel composition scales better with the size of the input graphs and for large graphs can be as much as 10 to 30 times faster than a sequential CPU algorithm. Shubho Sengupta, Vineel Pratap, Awni Y. Hannun |
ICASSP | 1 |
| 2021 | CrypTen: Secure Multi-Party Computation Meets Machine LearningabstractSecure multi-party computation (MPC) allows parties to perform computations on data while keeping that data private. This capability has great potential for machine-learning applications: it facilitates training of machine-learning models on private data sets owned by different parties, evaluation of one party's private model using another party's private data, etc. Although a range of studies implement machine-learning models via secure MPC, such implementations are not yet mainstream. Adoption of secure MPC is hampered by the absence of flexible software frameworks that `"speak the language" of machine-learning researchers and engineers. To foster adoption of secure MPC in machine learning, we present CrypTen: a software framework that exposes popular secure MPC primitives via abstractions that are common in modern machine-learning frameworks, such as tensor computations, automatic differentiation, and modular neural networks. This paper describes the design of CrypTen and measure its performance on state-of-the-art models for text classification, speech recognition, and image classification. Our benchmarks show that CrypTen's GPU support and high-performance communication between (an arbitrary number of) parties allows it to perform efficient private evaluation of modern machine-learning models under a semi-honest threat model. For example, two parties using CrypTen can securely predict phonemes in speech recordings using Wav2Letter faster than real-time. We hope that CrypTen will spur adoption of secure MPC in the machine-learning community. Brian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta, Mark Ibrahim, Laurens van der Maaten |
NeurIPS | 4 |
| 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZeroabstractThe AlphaGo, AlphaGo Zero, and AlphaZero series of algorithms are remarkable demonstrations of deep reinforcement learning’s capabilities, achieving superhuman performance in the complex game of Go with progressively increasing autonomy. However, many obstacles remain in the understanding of and usability of these promising approaches by the research community. Toward elucidating unresolved mysteries and facilitating future research, we propose ELF OpenGo, an open-source reimplementation of the AlphaZero algorithm. ELF OpenGo is the first open-source Go AI to convincingly demonstrate superhuman performance with a perfect (20:0) record against global top professionals. We apply ELF OpenGo to conduct extensive ablation studies, and to identify and analyze numerous interesting phenomena in both the model training and in the gameplay inference procedures. Our code, models, selfplay datasets, and auxiliary data are publicly available. Yuandong Tian, Jerry Ma, Qucheng Gong, Shubho Sengupta, Zhuoyuan Chen, James Pinkerton, C. Lawrence Zitnick |
ICML | 4 |
| 2017 | Exploring Sparsity in Recurrent Neural Networks
Sharan Narang, Gregory Frederick Diamos, Shubho Sengupta, Erich Elsen |
ICLR (Poster) | 3 |
| 2017 | Deep Voice: Real-time Neural Text-to-SpeechabstractWe present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks: a segmentation model for locating phoneme boundaries, a grapheme-to-phoneme conversion model, a phoneme duration prediction model, a fundamental frequency prediction model, and an audio synthesis model. For the segmentation model, we propose a novel way of performing phoneme boundary detection with deep neural networks using connectionist temporal classification (CTC) loss. For the audio synthesis model, we implement a variant of WaveNet that requires fewer parameters and trains faster than the original. By using a neural network for each component, our system is simpler and more flexible than traditional text-to-speech systems, where each component requires laborious feature engineering and extensive domain expertise. Finally, we show that inference with our system can be performed faster than real time and describe optimized WaveNet inference kernels on both CPU and GPU that achieve up to 400x speedups over existing implementations. Sercan Ö. Arik, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Andrew Gibiansky, Yongguo Kang, John Miller 0001, Andrew Y. Ng, Jonathan Raiman, Shubho Sengupta, Mohammad Shoeybi |
ICML | 11 |
| 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and MandarinabstractWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale. Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu |
ICML | 31 |
| 2016 | Persistent RNNs: Stashing Recurrent Weights On-ChipabstractThis paper introduces a new technique for mapping Deep Recurrent Neural Networks (RNN) efficiently onto GPUs. We show how it is possi- ble to achieve substantially higher computational throughput at low mini-batch sizes than direct implementations of RNNs based on matrix multiplications. The key to our approach is the use of persistent computational kernels that exploit the GPU’s inverted memory hierarchy to reuse network weights over multiple timesteps. Our initial implementation sustains 2.8 TFLOP/s at a mini-batch size of 4 on an NVIDIA TitanX GPU. This provides a 16x reduction in activation memory footprint, enables model training with 12x more parameters on the same hardware, allows us to strongly scale RNN training to 128 GPUs, and allows us to efficiently explore end-to-end speech recognition models with over 100 layers. Gregory Frederick Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates 0002, Erich Elsen, Jesse H. Engel, Awni Y. Hannun, Sanjeev Satheesh |
ICML | 2 |
| 2014 | Navigating the maze of graph analytics frameworks using massive graph datasetsabstractGraph algorithms are becoming increasingly important for analyzing large datasets in many fields. Real-world graph data follows a pattern of sparsity, that is not uniform but highly skewed towards a few items. Implementing graph traversal, statistics and machine learning algorithms on such data in a scalable manner is quite challenging. As a result, several graph analytics frameworks (GraphLab, CombBLAS, Giraph, SociaLite and Galois among others) have been developed, each offering a solution with different programming models and targeted at different users. Unfortunately, the "Ninja performance gap" between optimized code and most of these frameworks is very large (2-30X for most frameworks and up to 560X for Giraph) for common graph algorithms, and moreover varies widely with algorithms. This makes the end-users' choice of graph framework dependent not only on ease of use but also on performance. In this work, we offer a quantitative roadmap for improving the performance of all these frameworks and bridging the "ninja gap". We first present hand-optimized baselines that get performance close to hardware limits and higher than any published performance figure for these graph algorithms. We characterize the performance of both this native implementation as well as popular graph frameworks on a variety of algorithms. This study helps end-users delineate bottlenecks arising from the algorithms themselves vs. programming model abstractions vs. the framework implementations. Further, by analyzing the system-level behavior of these frameworks, we obtain bottlenecks that are agnostic to specific algorithms. We recommend changes to alleviate these bottlenecks (and implement some of them) and reduce the performance gap with respect to native code. These changes will enable end-users to choose frameworks based mostly on ease of use. Nadathur Satish, Narayanan Sundaram, Md. Mostofa Ali Patwary, Jiwon Seo 0002, Jongsoo Park, Muhammad Amber Hassaan, Shubho Sengupta, Zhaoming Yin, Pradeep Dubey |
SIGMOD Conference | 7 |