VLDB 2026 Research / reviewers in the wild / expert
Srinivas Eswar
dblp:224/1565
· DBLP profile ↗
7ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-3418-7796ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Brief Announcement: Optimality Conditions for Parallel Communication-Avoiding Matrix Multiplication with Overlapped CommunicationabstractWhen considering general matrix multiply (GEMM) algorithms for distributed-memory systems, the dominant paradigm is to minimize communication volume. However, to minimize time, one must consider how communication volume interacts with the system characteristics --- namely, communication bandwidth, memory capacity, and the co-scheduling of computation and communication. In this work, we demonstrate that the family of 3D GEMM algorithms --- although reducing overall communication by leveraging extra memory --- fundamentally concentrates communication to the phase of memory filling. This upfront cost hinders the ability to overlap communication with computation, and, consequently, calls for a revised view of the GEMM optimality regions. Our main contribution is derivation of optimality conditions for parallel matrix multiplication that jointly accounts for communication volume and overlap effects, pushing the 3D GEMM optimality region to start with the systems more than 11× larger than previously believed. Mikhail Isaev, Srinivas Eswar, Richard W. Vuduc |
SPAA | 2 |
| 2024 | On Rank Selection for Nonnegative Matrix FactorizationabstractRank selection, i.e. the choice of factorization rank, is the first step in constructing Nonnegative Matrix Factorization (NMF) models. It is a long-standing problem which is not unique to NMF, but arises in most models which attempt to decompose data into its underlying components. Since these models are often used in the unsupervised setting, the rank selection problem is further complicated by the lack of ground truth labels. In this paper, we review and empirically evaluate the most commonly used schemes for NMF rank selection. Srinivas Eswar, Koby Hayashi, Benjamin Cobb, Ramakrishnan Kannan, Grey Ballard, Richard W. Vuduc, Haesun Park |
IEEE Big Data | 1 |
| 2023 | Distributed-Memory Parallel JointNMFabstractJoint Nonnegative Matrix Factorization (JointNMF) is a hybrid method for mining information from datasets that contain both feature and connection information. We propose distributed-memory parallelizations of three algorithms for solving the JointNMF problem based on Alternating Nonnegative Least Squares, Projected Gradient Descent, and Projected Gauss-Newton. We extend well-known communication-avoiding algorithms using a single processor grid case to our coupled case on two processor grids. We demonstrate the scalability of the algorithms on up to 960 cores (40 nodes) with 60% parallel efficiency. The more sophisticated Alternating Nonnegative Least Squares (ANLS) and Gauss-Newton variants outperform the first-order gradient descent method in reducing the objective on large-scale problems. We perform a topic modelling task on a large corpus of academic papers that consists of over 37 million paper abstracts and nearly a billion citation relationships, demonstrating the utility and scalability of the methods. Srinivas Eswar, Benjamin Cobb, Koby Hayashi, Ramakrishnan Kannan, Grey Ballard, Richard W. Vuduc, Haesun Park |
ICS | 1 |
| 2021 | ORCA: Outlier detection and Robust Clustering for Attributed graphs
Srinivas Eswar, Ramakrishnan Kannan, Richard W. Vuduc, Haesun Park |
J. Glob. Optim. | 1 |
| 2021 | PLANC: Parallel Low-rank Approximation with Nonnegativity ConstraintsabstractWe consider the problem of low-rank approximation of massive dense nonnegative tensor data, for example, to discover latent patterns in video and imaging applications. As the size of data sets grows, single workstations are hitting bottlenecks in both computation time and available memory. We propose a distributed-memory parallel computing solution to handle massive data sets, loading the input data across the memories of multiple nodes, and performing efficient and scalable parallel algorithms to compute the low-rank approximation. We present a software package called Parallel Low-rank Approximation with Nonnegativity Constraints, which implements our solution and allows for extension in terms of data (dense or sparse, matrices or tensors of any order), algorithm (e.g., from multiplicative updating techniques to alternating direction method of multipliers), and architecture (we exploit GPUs to accelerate the computation in this work). We describe our parallel distributions and algorithms, which are careful to avoid unnecessary communication and computation, show how to extend the software to include new algorithms and/or constraints, and report efficiency and scalability results for both synthetic and real-world data sets. Srinivas Eswar, Koby Hayashi, Grey Ballard, Ramakrishnan Kannan, Michael A. Matheson, Haesun Park |
ACM Trans. Math. Softw. | 1 |
| 2020 | Distributed-memory parallel symmetric nonnegative matrix factorizationabstractWe develop the first distributed-memory parallel implementation of Symmetric Nonnegative Matrix Factorization (SymNMF), a key data analytics kernel for clustering and dimensionality reduction. Our implementation includes two different algorithms for SymNMF, which give comparable results in terms of time and accuracy. The first algorithm is a parallelization of an existing sequential approach that uses solvers for non symmetric NMF. The second algorithm is a novel approach based on the Gauss-Newton method. It exploits second-order information without incurring large computational and memory costs. We evaluate the scalability of our algorithms on the Summit system at Oak Ridge National Laboratory, scaling up to 128 nodes (4,096 cores) with 70% efficiency. Additionally, we demonstrate our software on an image segmentation task. Srinivas Eswar, Koby Hayashi, Grey Ballard, Ramakrishnan Kannan, Richard W. Vuduc, Haesun Park |
SC | 1 |
| 2019 | A microbenchmark characterization of the Emu chick
Jeffrey Young 0001, Eric R. Hein, Srinivas Eswar, Patrick Lavin, Jiajia Li 0001, E. Jason Riedy, Richard W. Vuduc, Thomas M. Conte |
Parallel Comput. | 3 |