Purushotham V. Bangalore

dblp:39/2060 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-1098-9997ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Analysis of I/O Subsystem Utilization on the Theta Supercomputer
abstract
Supercomputing facilities face operational challenges from inefficient I/O patterns that degrade shared storage performance. As HPC workloads diversify, traditional volume-based monitoring fails to distinguish efficient data movement from prolonged, low-bandwidth congestion on shared parallel file systems. This work proposes a hierarchical, application-agnostic methodology using routine Darshan production logs to identify, classify, and prioritize I/O behavioral inefficiencies observable from application-level telemetry at facility scale. Applying this methodology to 14,983 jobs from the ALCF Theta supercomputer, we demonstrate that I/O inefficiency is highly concentrated and repeatable: 87% of high-volume, low-bandwidth I/O volume is attributable to only five users. We map identified inefficiency classes to two distinct diagnostic categories, Application-Limited and Parallelism-Bounded, each with a corresponding intervention strategy requiring minimal user involvement. These results show that routine I/O telemetry can support a practical operational control loop, and the methodology generalizes to any facility deploying Darshan regardless of underlying storage architecture.
Hari Teja Jajula, Purushotham V. Bangalore
HPDC2
2024 Data Security Approaches to Support Open Science on High-Performance Computing Systems
abstract
Research data must be shared with other researchers for further collaboration to support open science. Sharing data with other researchers will save time and money instead of replicating the same research and also enables collaboration. However, open science data are vulnerable to attacks, and attackers can exploit the system, which could lead to harmful consequences from data sharing. This paper explores open science data security challenges and prevention techniques. First, we will review state-of-the-art algorithms and techniques that have been proposed to secure data. Second, we show the difference between generic data security and HPC security and compare the strengths and weaknesses of these approaches. Finally, we suggest some options to address these problems.
Md. Kamal Hossain Chowdhury, Purushotham V. Bangalore
ICCCN2
2024 Understanding GPU Triggering APIs for MPI+X Communication
Patrick G. Bridges, Anthony Skjellum, Evan Drake Suggs, Derek Schafer, Purushotham V. Bangalore
EuroMPI5
2023 Design of a portable implementation of partitioned point-to-point communication primitives
abstract
Abstract The Message Passing Interface (MPI) has been the dominant message passing solution for scientific computing for decades. MPI point‐to‐point communications are highly efficient mechanisms for process‐to‐process communication. However, MPI performance when processes utilize multiple threads is slowed by concurrency protections in the MPI library. MPI's current thread level interface imposes these overheads throughout the library when thread safety is needed. While much work has been done to reduce multithreading overheads in MPI, a solution is needed that reduces the number of messages exchanged in a threaded environment. Partitioned communication is included in the MPI 4.0 standard as an alternative that addresses the challenges of multithreaded communication in MPI today. Partitioned communication reduces overall message volume by creating a buffer‐sharing mechanism between threads such that they can indicate when portions of a communication buffer are available to be sent. Separation of the control and data planes in MPI is enabled by allowing persistent initialization and single occurrence message buffer matching from the indication that the data is ready to be sent. This enables the usage of underlying hardware primitives like triggered operations, where commands (destination, size, etc.) can be set up prior to data buffer readiness and readiness triggered with a simple doorbell/counter later. This approach is useful for future development of MPI operations in environments where traditional networking commands can have performance challenges, like accelerators (GPUs, FPGAs). In this paper, we detail the design and implementation of a layered library (built on top of MPI‐3.1) and an integrated Open MPI solution that supports the new, MPI‐4.0 partitioned communication feature set. The library will enable applications to use currently released MPI implementations and older legacy libraries to provide partitioned communication support while also enabling further exploration of this new communication model in new applications and use cases. We will compare the designs of the library and native Open MPI support, provide performance results and comparisons between the two approaches, and lessons learned from the implementation of partitioned communication in both library and native forms. We find that the native implementation and library have similar performance with a percentage difference under 0.94% in microbenchmarks and performance within 5% for a partitioned communication enabled proxy application.
W. Pepper Marts, Andrew Worley, Prema Soundarajan, Derek Schafer, Matthew G. F. Dosanjh, Ryan E. Grant, Purushotham V. Bangalore, Anthony Skjellum, Sheikh K. Ghafoor
Concurr. Comput. Pract. Exp.7
2021 Understanding the use of message passing interface in exascale proxy applications
abstract
Summary The Exascale Computing Project (ECP) focuses on the development of future exascale‐capable applications. Most ECP applications use the message passing interface (MPI) as their parallel programming model with mini‐apps serving as proxies. This paper explores the explicit usage of MPI in such ECP proxy applications. We empirically analyze 14 proxy applications from the ECP Proxy Apps Suite. We use the MPI profiling interface (PMPI) to collect MPI usage patterns in ECP proxy apps. Our analysis shows that a small subset of features from MPI is commonly used in the proxies of exascale‐capable applications, even when they reference third‐party libraries. This study is intended to provide a better understanding of the use of MPI in current exascale applications. The findings can help focus software investments made for exascale systems in the MPI middleware including optimization, fault‐tolerance, tuning, and hardware‐offload.
Nawrin Sultana, Martin Ruefenacht, Anthony Skjellum, Purushotham V. Bangalore, Ignacio Laguna, Kathryn Mohror
Concurr. Comput. Pract. Exp.4
2021 Implementation and evaluation of MPI 4.0 partitioned communication libraries
Matthew G. F. Dosanjh, Andrew Worley, Derek Schafer, Prema Soundararajan, Sheikh K. Ghafoor, Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant
Parallel Comput.7
2020 Foreword to the Special Issue of the Workshop on Exascale MPI (ExaMPI 2017)
abstract
The aim of the Workshop on Exascale MPI (ExaMPI 2017), held in conjunction with SC17: The International Conference for High Performance Computing, Networking, Storage and Analysis, was to bring together researchers and developers to present and discuss innovative algorithms and concepts in the Message Passing programming model and to create a forum for open and potentially controversial discussions on the future of MPI in the Exascale era. This special issue includes selected papers from this workshop that include innovative algorithms for collective operations, extensions to MPI, including datacentric models, scheduling/routing to avoid network congestion, fault-tolerant communication, interoperability of MPI and PGAS models, and use of MPI in large-scale simulations. The first paper, titled “A survey of MPI usage in the US Exascale Computing Project,” provides an analysis of the survey that was conducted to understand how MPI is currently used and intended to be used by different applications that are part of the Exascale Computing Project (ECP).1 The results of analysis provide specific recommendations for MPI implementors, tool developers, and the MPI Forum. The next three papers focus on the issue of message matching in MPI and provide different options to address message matching at exascale. The paper titled “Tail Queues: A Multi-threaded Matching Architecture” introduces a novel parallel matching architecture and prototype implementation based on MPICH to improve the performance of message matching.2 The paper titled “Communication-Aware Message Matching in MPI” uses a novel message queue architecture that allocates dedicated message queues based on the frequency of communication between various processes to reduce the queue search time and also reduces memory consumption.3 These performance improvements result in a speedup of 5 times on the FDS application. The next paper, titled “Hardware MPI Message Matching: Insights into MPI Matching Behavior to Inform Design,” explores what hardware features are needed to support efficient message matching through the evaluation of message matching characteristics of major MPI implementations.4 The next two papers consider support for fault tolerance. The first paper, titled “EReinit: Scalable and Efficient Fault-Tolerance for Bulk-Synchronous MPI Applications,” describes a global-restart model to improve the recovery time of applications when dealing with faults.5 The second paper, titled “The Unexpected Virtue of Almost: Exploiting MPI Collective Operations to Approximately Coordinate Checkpoints,” describes an uncoordinated checkpointing mechanism that makes use of collective operations already used in an application to force checkpoints.6 The paper titled “Optimizing Point-to-Point Communication between Adaptive MPI Endpoints in Shared Memory” describes an approach to optimize point-to-point communication in a shared memory environment and hence improve MPI multithreading support.7 The paper titled “On the Memory Attribution Problem: A Solution and Case Study Using MPI” describes a solution to capture and analyze memory usage by an MPI application and the MPI library.8 Lastly, the paper titled “Twister2: Design of a Big Data Toolkit” describes an architecture to support different types of data-intensive applications in a unified framework.9 Overall, these nine papers contribute to the knowledge base and advancement of the Message Passing Interface in diverse and useful ways. While illustrating the staying power of MPI after a quarter century, they also point to areas of opportunity for enhancement at Exascale as well as to new potential application areas.
Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant
Concurr. Comput. Pract. Exp.2
2019 Lightweight Fault Tolerance in Pregel-Like Systems
abstract
Pregel-like systems are popular for iterative graph processing thanks to their user-friendly vertex-centric programming model. However, existing Pregel-like systems only adopt a naïve checkpointing approach for fault tolerance, which saves a large amount of data about the state of computation and significantly degrades the failure-free execution performance. Advanced fault tolerance/recovery techniques are left unexplored in the context of Pregel-like systems. This paper proposes a non-invasive lightweight checkpointing (LWCP) scheme which minimizes the data saved to each checkpoint, and additional data required for recovery are generated online from the saved data. This improvement results in 10x speedup in checkpointing, and an integration of it with a recently proposed log-based recovery approach can further speed up recovery when failure occurs. Extensive experiments verified that our proposed LWCP techniques are able to significantly improve the performance of both checkpointing and recovery in a Pregel-like system.
Da Yan 0001, James Cheng, Cheng Long 0001, Purushotham V. Bangalore
ICPP5
2019 Exposition, clarification, and expansion of MPI semantic terms and conventions: is a nonblocking MPI function permitted to block?
abstract
This paper offers a timely study and proposed clarifications, revisions, and enhancements to the Message Passing Interface's (MPI's) Semantic Terms and Conventions. To enhance MPI, a clearer understanding of the meaning of the key terminology has proven essential, and, surprisingly, important concepts remain underspecified, ambiguous and, in some cases, inconsistent and/or conflicting despite 26 years of standardization. This work addresses these concerns comprehensively and usefully informs MPI developers, implementors, those teaching and learning MPI, and power users alike about key aspects of existing conventions, syntax, and semantics. This paper will also be a useful driver for great clarity in current and future standardization and implementation efforts for MPI.
Purushotham V. Bangalore, Rolf Rabenseifner, Daniel J. Holmes, Julien Jaeger, Guillaume Mercier, Claudia Blaas-Schenner, Anthony Skjellum
EuroMPI1
2019 Planning for performance: Enhancing achievable performance for MPI through persistent collective operations
Daniel J. Holmes, Bradley Morgan, Anthony Skjellum, Purushotham V. Bangalore, Srinivas Sridharan 0002
Parallel Comput.4
2018 MPI Derived Datatypes: Performance and Portability Issues
abstract
This paper addresses performance-portability and overall performance issues when derived datatypes are used with four MPI implementations: Open MPI, MPICH, MVAPICH2, and Intel MPI. These comparisons are particularly relevant today since most vendor implementations are now based on Open MPI or MPICH rather than on vendor proprietary code as was more prevalent in the past. Our findings are that, within a single MPI implementation, there are significant differences in performance as a function of it reasonable encodings of derived datatypes as supported by the MPI standard. While this finding may not be surprising, it is important to understand how fundamental vs. arbitrary choices made in early implementation impact the use of derived datatypes to date.
Qingqing Xiong, Purushotham V. Bangalore, Anthony Skjellum, Martin C. Herbordt
EuroMPI2
2017 Architectural implications on the performance and cost of graph analytics systems
abstract
Graph analytics systems have gained significant popularity due to the prevalence of graph data. Many of these systems are designed to run in a shared-nothing architecture whereby a cluster of machines can process a large graph in parallel. In more recent proposals, others have argued that a single-machine system can achieve better performance and/or is more cost-effective. There is however no clear consensus which approach is better. In this paper, we classify existing graph analytics systems into four categories based on the architectural differences, i.e., processing infrastructure (centralized vs distributed), and memory consumption (in-memory vs out-of-core). We select eight open-source systems to cover all categories, and perform a comparative measurement study to compare their performance and cost characteristics across a spectrum of input data, applications, and hardware settings. Our results show that the best performing configuration can depend on the type of applications and input graphs, and there is no dominant winner across all categories. Based on our findings, we summarize the trends in performance and cost, and provide several insights that help to illuminate the performance and resource cost tradeoffs across different graph analytics systems and categories.
Qizhen Zhang 0001, Da Yan 0001, James Cheng, Boon Thau Loo, Purushotham V. Bangalore
SoCC6
2016 Managing I/O Interference in a Shared Burst Buffer System
abstract
In this work, we investigate the problem of inter-application interference in a shared Burst Buffer (BB) system. A BB is a new storage technology for HPC architectures that acts as an intermediate layer between performance-hungry HPC applications and the slow parallel file system. While the BB is meant to alleviate the problem of slow I/O in HPC systems, it is itself prone to performance degradation under interference. We observe that the magnitude of interference effects can reach a level that matters to the HPC system and the jobs that run on it. We investigate I/O scheduling techniques as a mechanism to mitigate BB I/O interference. With our results, we show that scheduling techniques tuned to BBs can control interference and significant performance benefits can be achieved.
Sagar Thapaliya, Purushotham V. Bangalore, Jay F. Lofstead, Kathryn Mohror, Adam Moody
ICPP2
2012 Scheduling and planning job execution of loosely coupled applications
Enis Afgan, Purushotham V. Bangalore, Tibor Skala
J. Supercomput.2
2012 Raising the level of abstraction for developing message passing applications
Ritu Arora, Purushotham V. Bangalore, Marjan Mernik
J. Supercomput.2
2012 Tools and techniques for non-invasive explicit parallelization
Ritu Arora, Purushotham V. Bangalore, Marjan Mernik
J. Supercomput.2
2012 PPModel: a modeling tool for source code maintenance and optimization of parallel programs
Ferosh Jacob, Jeffrey G. Gray, Jeffrey C. Carver, Marjan Mernik, Purushotham V. Bangalore
J. Supercomput.5
2011 Application Information Services for distributed computing environments
Enis Afgan, Purushotham V. Bangalore, Karolj Skala
Future Gener. Comput. Syst.2
2011 A technique for non-invasive application-level checkpointing
Ritu Arora, Purushotham V. Bangalore, Marjan Mernik
J. Supercomput.2
2010 CUDACL: A tool for CUDA and OpenCL programmers
abstract
Graphical Processing Unit (GPU) programming languages are used extensively for general-purpose computations. However, GPU programming languages are at a level of abstraction suitable only for use by expert parallel programmers. This paper presents a new approach through which `C' or Java programmers can access these languages without having to focus on the technical or language-specific details. A prototype of the approach, named CUDACL, is introduced through which a programmer can specify one or more parallel blocks in a file and execute in a GPU. CUDACL also helps the programmer to make CUDA or OpenCL kernel calls inside an existing program. Two scenarios have been successfully implemented to assess the usability and potential of the tool. The tool was created based on a detailed analysis of the CUDA and OpenCL programs. Our evaluation of CUDACL compared to other similar approaches shows the efficiency and effectiveness of CUDACL.
Ferosh Jacob, David Whittaker, Sagar Thapaliya, Purushotham V. Bangalore, Marjan Mernik, Jeffrey G. Gray
HiPC4
2009 GridAtlas - A grid application and resource configuration repository and discovery service
abstract
Although access to grid resources is realized through a standardized interface, independent grid resources are not only managed autonomously but are also accessed as independent entities. Such environment results in configuration differences among individual resources forcing users that access those resources to deal with the variability in resource configurations. This behaviour breaks the concept of interpreting the grid as a unified entity and forces the users to think of the grid in terms of individual resources. Concretely, this variability is expressed through the requirement for the users to explicitly state application installation properties on individual resources during each job submission. This is a tedious, error-prone and unnecessary process that acts as a barrier in the use of the grid. In this paper, a tool named GridAtlas is presented that keeps up with the details of individual resource and application configurations and makes such data easily accessible from a well-known location through Web-service API calls or a Web interface. This paper describes the architecture of the GridAtlas service along with use cases where GridAtlas has been successfully applied and illustrates the benefit of such a service in real grid environments.
Enis Afgan, Purushotham V. Bangalore, Dustin Duncan
CLUSTER2
2009 Cross-domain authorization for federated virtual organizations using the myVocs collaboration environment
abstract
Abstract This paper describes our experiences building and working with the reference implementation of myVocs (my Virtual Organization Collaboration System). myVocs provides a flexible environment for exploring new approaches to security, application development, and access control built from Internet services without a central identity repository. The myVocs framework enables virtual organization (VO) self‐management across unrelated security domains for multiple, unrelated VOs. By leveraging the emerging distributed identity management infrastructure. myVocs provides an accessible, secure collaborative environment using standards for federated identity management and open‐source software developed through the National Science Foundation Middleware Initiative. The Shibboleth software, an early implementation of the Organization for the Advancement of Structured Information Standards Security Assertion Markup Language standard for browser single sign‐on, provides the middleware needed to assert identity and attributes across domains so that access control decisions can be determined at each resource based on local policy. The eduPerson object class for lightweight directory access protocol (LDAP) provides standardized naming, format, and semantics for a global identifier. We have found that a Shibboleth deployment supporting VOs requires the addition of a new VO service component allowing VOs to manage their own membership and control access to their distributed resources. The myVocs system can be integrated with Grid authentication and authorization using GridShib. Copyright © 2008 John Wiley & Sons, Ltd.
Jill B. Gemmill, John-Paul Robinson, Tom Scavo, Purushotham V. Bangalore
Concurr. Comput. Pract. Exp.4
2007 Performance Characterization of BLAST for the Grid
abstract
BLAST is a commonly used bioinformatics application for performing query searches and analysis of biological data. As the amount of search data increases so does job search times. As means of reducing job turnaround times, scientists are resorting to new technologies such as grid computing to obtain needed computational and storage resources. Inherent with advent of new technologies, are additional complexities that arise, forcing scientists to deal with them. Grid computing exemplifies dynamic and transient state of heterogeneous resources that become a major obstacle in realizing user-desired levels of service. Many users do not realize that techniques applied in more traditional cluster environments do not simply transition into grid environment. This paper analyzes resource and application dependencies for BLAST in terms of job parameters that result in performance tradeoffs. We present a set of examples showing performance variability and point out a set of guidelines, which lead to establishing job performance tradeoffs.
Enis Afgan, Purushotham V. Bangalore
BIBE2
2006 Grid-Flow: a Grid-enabled scientific workflow system with a Petri-net-based interface
abstract
Abstract Advances in computer technologies have enabled scientists to explore research issues in their respective domains at scales greater and finer than ever before. The availability of efficient data collection and analysis tools presents researchers with vast opportunities to process heterogeneous data within a distributed environment. To support the opportunities enabled by massive computation, a suitable scientific workflow system is needed to help the users to manage data and programs, and to design reusable procedures of scientific experimental tasks. In this paper, the design and prototype implementation of a scientific workflow infrastructure, called Grid‐Flow, is presented. Grid‐Flow assists researchers in specifying scientific experiments using a Petri‐net‐based interface. The Grid‐Flow infrastructure is designed as a Service Oriented Architecture with multi‐layer component models. The contributions of Grid‐Flow are as follows: (1) a new, lightweight, programmable Grid workflow language, Grid‐Flow Description Language, is provided to describe the workflow process in a Grid environment; (2) a Petri‐net‐based user interface, based on the Generic Modeling Environment, is demonstrated to help the user design the workflow process with a Petri‐net model; and (3) a program integration component of the Grid‐Flow system is presented to integrate all possible programs into the system. Copyright © 2005 John Wiley & Sons, Ltd.
Zhijie Guan, Francisco Hernández, Purushotham V. Bangalore, Jeffrey G. Gray, Anthony Skjellum, Vijay Velusamy
Concurr. Comput. Pract. Exp.3
2006 GAUGE: Grid Automation and Generative Environment
abstract
Abstract The Grid has proven to be a successful paradigm for distributed computing. However, constructing applications that exploit all the benefits that the Grid offers is still not optimal for both inexperienced and experienced users. Recent approaches to solving this problem employ a high‐level abstract layer to ease the construction of applications for different Grid environments. These approaches help facilitate construction of Grid applications, but they are still tied to specific programming languages or platforms. A new approach is presented in this paper that uses concepts of domain‐specific modeling (DSM) to build a high‐level abstract layer. With this DSM‐based abstract layer, the users are able to create Grid applications without knowledge of specific programming languages or being bound to specific Grid platforms. An additional benefit of DSM provides the capability to generate software artifacts for various Grid environments. This paper presents the Grid Automation and Generative Environment (GAUGE). The goal of GAUGE is to automate the generation of Grid applications to allow inexperienced users to exploit the Grid fully. At the same time, GAUGE provides an open framework in which experienced users can build upon and extend to tailor their applications to particular Grid environments or specific platforms. GAUGE employs domain‐specific modeling techniques to accomplish this challenging task. Copyright © 2005 John Wiley & Sons, Ltd.
Francisco Hernández, Purushotham V. Bangalore, Jeffrey G. Gray, Zhijie Guan, Kevin D. Reilly
Concurr. Comput. Pract. Exp.2
2002 Mississippi Computational Web Portal
abstract
Abstract This paper describes design and implementation of an open, extensible object‐oriented framework that allows the integration of new and legacy components into a single user‐friendly Grid computing environment. Thus we extend the researcher's desktop by providing seamless access to remote resources (that is, hardware, software, and data), and thereby simplifying currently difficult to comprehend and changing interfaces and emerging protocols. The user, through the familiar Web browser interface, is able to compose complex computational tasks represented as a collection of middle‐tier objects serving as proxies for services rendered by the back‐end. The proxies through a Grid resource broker use the Grid services, as defined by the Global Grid Forum, to access remote computational resources. The middle‐tier objects are persistent, and therefore once configured, simulation can be reused, shared between users, or undergo transition into operational or educational use. Copyright © 2002 John Wiley & Sons, Ltd.
Tomasz Haupt, Purushotham V. Bangalore, Gregory Henley
Concurr. Comput. Pract. Exp.2
2001 Object-oriented analysis and design of the Message Passing Interface
abstract
Abstract The major contribution of this paper is the application of modern analysis techniques to the important Message Passing Interface standard, work done in order to obtain information useful in designing both application programmer interfaces for object‐oriented languages, and message passing systems. Recognition of ‘Design Patterns’ within MPI is an important discernment of this work. A further contribution is a comparative discussion of the design and evolution of three actual object‐oriented designs for the Message Passing Interface ( MPI‐1SF ) application programmer interface (API), two of which have influenced the standardization of C++ explicit parallel programming with MPI‐2, and which strongly indicate the value of a priori object‐oriented design and analysis of such APIs. Knowledge of design patterns is assumed herein. Discussion provided here includes systems developed at Mississippi State University (MPI++), the University of Notre Dame (OOMPI), and the merger of these systems that results in a standard binding within the MPI‐2 standard. Commentary concerning additional opportunities for further object‐oriented analysis and design of message passing systems and APIs, such as MPI‐2 and MPI/RT, are mentioned in conclusion. Connection of modern software design and engineering principles to high performance computing programming approaches is a new and important further contribution of this work. Copyright © 2001 John Wiley & Sons, Ltd.
Anthony Skjellum, Diane G. Wooley, Michael Wolf, Purushotham V. Bangalore, Andrew Lumsdaine, Jeffrey M. Squyres, Brian C. McCandless
Concurr. Comput. Pract. Exp.5