Raghavan Krishnan

dblp:253/0069 · also Krishnan Raghavan · DBLP profile ↗
← Back
5ranked-venue papers in the field
2as first author
2since 2021 · last 2024
0000-0001-9409-2011ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (1 first)Database Systems & Data Management · 1 (1 first)
YearPublicationVenuePosition
2024 Privacy-Preserving Federated Learning for Science: Challenges and Research Directions
abstract
This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific artificial intelligence models, in particular, foundation models (FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy—an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.
Kibaek Kim, Raghavan Krishnan, Olivera Kotevska, Matthieu Dorier, Ravi K. Madduri, Minseok Ryu, Todd S. Munson, Robert B. Ross, Thomas Flynn 0001, Ai Kagawa, Byung-Jun Yoon, Christian Engelmann, Farzad Yousefian
IEEE Big Data2
2024 DISTRI: Development and Integration of Simulation Tools for Resilient Infrastructure
abstract
In contemporary scientific research, data acquisition and analysis platforms have grown increasingly complex, often spanning multiple facilities with diverse internal structures. Efficiently managing the interactions between job scheduling, resource allocation, and networking across these distributed systems requires a robust simulation framework. However, existing simulators fall short in capturing the detailed interactions necessary for comprehensive analysis of large-scale distributed environments. To address this gap, we introduce DISTRI, a versatile framework specifically designed for the development and testing of distributed multi-facility workflows. DISTRI allows for customizable facility configurations and includes built-in support for distributed, resilient scheduling and resource management, alongside detailed network simulation for data communication. Key features of DISTRI encompass inter- and intra-facility resource management, agent-based distributed scheduling, and extensive performance metrics logging for both resource and network management. By providing these essential tools, DISTRI enables thorough analysis and optimization, thereby advancing research in the resilience and efficiency of multi-facility systems.
Imtiaz Mahmud, Pawel Zuk, Cong Wang 0014, Mariam Kiran, Kesheng Wu, Komal Thareja, Raghavan Krishnan, Anirban Mandal, Ewa Deelman
IEEE Big Data7
2019 A Multi-Step Nonlinear Dimension-Reduction Approach with Applications to Big Data
abstract
In this paper, a novel dimension-reduction approach is presented to overcome challenges such as nonlinear relationships, heterogeneity, and noisy dimensions. Initially, the p attributes in the data are first organized into random groups. Next, to systematically remove redundant and noisy dimensions from the data, each group is independently mapped into a low dimensional space via a parametric mapping. The group-wise transformation parameters are estimated using a low-rank approximation of distance covariance. The transformed attributes are reorganized into groups based on the magnitude of their respective eigenvalues. The group-wise organization and reduction process is performed until a user-defined criterion on eigenvalues is satisfied. In addition, novel procedures are introduced to aggregate the transformation parameters when the data is available in batches. Overall performance is demonstrated with extensive simulation analysis on classification by employing 10 data-sets.
Raghavan Krishnan, V. A. Samaranayake, Sarangapani Jagannathan
IEEE Trans. Knowl. Data Eng.1
2018 Distributed Learning of Deep Sparse Neural Networks for High-dimensional Classification
abstract
While analyzing high dimensional data-sets using deep neural network (NN), increased sparsity is desirable but requires careful selection of "sparsity parameters." In this paper, a novel distributed learning methodology is proposed to optimize the NN while addressing this challenge. To address this challenge, the optimal sparsity in the NN is estimated via a two player zero-sum game in the paper. In the proposed game, sparsity parameter is the first player with the aim of increasing sparsity in the NN while NN weights is the second player with the goal of improving its performance in the presence of increased sparsity. To solve the game, additional variables are introduced into the optimization problem such that the output at every layer in the NN depends on this variable instead of the previous layer. Using these additional variables, layer wise cost-functions are derived that are then independently optimized to learn the additional variables, NN weights and the sparsity parameters. To implement the proposed learning procedure in a parallelized and distributed environment, a novel computational algorithm is also proposed. The efficiency of the proposed approach is demonstrated using a total of six data-sets.
Shweta Garg 0003, Raghavan Krishnan, Sarangapani Jagannathan, V. A. Samaranayake
IEEE BigData2
2018 A Minimax Approach for Classification with Big-data
abstract
In this paper, a novel methodology to reduce the generalization errors occurring due to domain shift in big data classification is presented. This reduction is achieved by introducing a suitably selected domain shift to the training data via what is referred to as "distortion model". These distortions are introduced through an affine transformation and additional data-samples are obtained. Next, a deep neural network (NN), referred as "classifier", is used to classify both the original and the additional data samples. By learning from both the original and additional data-samples, the classifier compensates for the domain shift while maintaining its performance on original data. However, as the exact magnitude of the shift one would encounter in real applications is unknown a priori and difficult to predict. The objective is to compensate for the optimal shift that can be introduced by the distortion model without significantly degrading the performance of the model. A two-player zero-sum game is thus designed where the first player is the distortion model with the aim of increasing the domain shift. The classifier then becomes the second player whose aim is to minimize the impact of domain shift. Finally, a direct error-driven learning scheme is utilized to minimize the impact of the classifier while maximizing the domain shift. A comprehensive simulation study is presented where a 12% improvement in the presence of domain shift is demonstrated. The proposed approach is also shown to improve generalization by 6%.
Raghavan Krishnan, Sarangapani Jagannathan, V. A. Samaranayake
IEEE BigData1