Sarat Sreepathi

dblp:56/1471 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-4978-9423ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2023 Experiences readying applications for Exascale
abstract
The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.
Nicholas Malaya, O. E. Bronson Messer, Joseph Glenski, Antigoni Georgiadou, Justin Lietz, Kalyana C. Gottiparthi, Marcus S. Day, Jackie Chen, Jon S. Rood, Lucas Esclapez, James B. White III, Gustav R. Jansen, Nicholas Curtis, Stephen Nichols, Jakub Kurzak, Noel Chalmers, Chip Freitag, Paul T. Bauman, Alessandro Fanfarillo, Reuben D. Budiardja, Thomas Papatheodore, Nicholas Frontiere, Damon McDougall, Matthew R. Norman, Sarat Sreepathi, Philip C. Roth, Dmytro Bykov, Noah Wolfe, Paul Mullowney, Markus Eisenbach 0002, Marc T. Henry de Frahan, Wayne Joubert
SC25
2023 The Simple Cloud-Resolving E3SM Atmosphere Model Running on the Frontier Exascale System
Peter M. Caldwell, Luca Bertagna, Conrad Clevenger, Aaron Donahue, James G. Foucar, Oksana Guba, Benjamin R. Hillman, Noel Keen, Jayesh Krishna, Matthew R. Norman, Sarat Sreepathi, Christopher Terai, James B. White III, Andrew G. Salinger, Renata B. McCoy, L. Ruby Leung, David C. Bader, Danqing Wu
SC12
2021 Early Evaluation of Fugaku A64FX Architecture Using Climate Workloads
abstract
The Energy Exascale Earth System Model (E3SM) Project is an ongoing, state-of-the-science Earth system modeling, simulation, and prediction project that targets efficient utilization of U.S. Department of Energy’s (DOE) supercomputers to meet the science needs of the nation and the mission needs of DOE. This work focuses on our early evaluation of the A64FX architecture on Fugaku supercomputer using E3SM benchmarks. We will present results that track hardware trends, facilitate architecture comparison and the specific impact on our workload using an atmospheric model benchmark. We have two variants of the code written in Fortran and C++/Kokkos respectively which were used to collect data on a variety of CPU and GPU platforms. Furthermore, we have conducted a comparative evaluation of the compilers on the A64FX architecture and found GNU to be the best performer for our workload. Our experience so far indicates that Fugaku/A64FX shows promising energy efficiency (performance/Watt) with further performance gains possible through architecture-aware optimization efforts.
Sarat Sreepathi
CLUSTER1
2017 Parallel Multivariate Spatio-Temporal Clustering of Large Ecological Datasets on Hybrid Supercomputers
abstract
A proliferation of data from vast networks of remote sensing platforms (satellites, unmanned aircraft systems (UAS), airborne etc.), observational facilities (meteorological, eddy covariance etc.), state-of-the-art sensors, and simulation models offer unprecedented opportunities for scientific discovery. Unsupervised classification is a widely applied data mining approach to derive insights from such data. However, classification of very large data sets is a complex computational problem that requires efficient numerical algorithms and implementations on high performance computing (HPC) platforms. Additionally, increasing power, space, cooling and efficiency requirements has led to the deployment of hybrid supercomputing platforms with complex architectures and memory hierarchies like the Titan system at Oak Ridge National Laboratory. The advent of such accelerated computing architectures offers new challenges and opportunities for big data analytics in general and specifically, large scale cluster analysis in our case. Although there is an existing body of work on parallel cluster analysis, those approaches do not fully meet the needs imposed by the nature and size of our large data sets. Moreover, they had scaling limitations and were mostly limited to traditional distributed memory computing platforms. We present a parallel Multivariate Spatio-Temporal Clustering (MSTC) technique based on k-means cluster analysis that can target hybrid supercomputers like Titan. We developed a hybrid MPI, CUDA and OpenACC implementation that can utilize both CPU and GPU resources on computational nodes. We describe performance results on Titan that demonstrate the scalability and efficacy of our approach in processing large ecological data sets.
Sarat Sreepathi, Jitendra Kumar 0001, Richard Tran Mills, Forrest M. Hoffman, Vamsi Sripathi, William W. Hargrove
CLUSTER1
2016 Communication Characterization and Optimization of Applications Using Topology-Aware Task Mapping on Large Supercomputers
abstract
On large supercomputers, the job scheduling systems may assign a non-contiguous node allocation for user applications depending on available resources. With parallel applications using MPI (Message Passing Interface), the default process ordering does not take into account the actual physical node layout available to the application. This contributes to non-locality in terms of physical network topology and impacts communication performance of the application. In order to mitigate such performance penalties, this work describes techniques to identify suitable task mapping that takes the layout of the allocated nodes as well as the application's communication behavior into account. During the first phase of this research, we instrumented and collected performance data to characterize communication behavior of critical US DOE (United States - Department of Energy) applications using an augmented version of the mpiP tool. Subsequently, we developed several reordering methods (spectral bisection, neighbor join tree etc.) to combine node layout and application communication data for optimized task placement. We developed a tool called mpiAproxy to facilitate detailed evaluation of the various reordering algorithms without requiring full application executions. This work presents a comprehensive performance evaluation (14,000 experiments) of the various task mapping techniques in lowering communication costs on Titan, the leadership class supercomputer at Oak Ridge National Laboratory.
Sarat Sreepathi, Eduardo F. D'Azevedo, Bobby Philip, Patrick H. Worley
ICPE1
2013 SCORPIO: A scalable two-phase parallel I/O library with application to a large scale subsurface simulator
abstract
Inefficient parallel I/O is known to be a major bottleneck among scientific applications employed on supercomputers as the number of processor cores grows into the thousands. Our prior experience indicated that parallel I/O libraries such as HDF5 that rely on MPI-IO do not scale well beyond 10K processor cores, especially on parallel file systems (like Lustre) with single point of resource contention. Our previous optimization efforts for a massively parallel multi-phase and multi-component subsurface simulator (PFLOTRAN) led to a two-phase I/O approach at the application level where a set of designated processes participate in the I/O process by splitting the I/O operation into a communication phase and a disk I/O phase. The designated I/O processes are created by splitting the MPI global communicator into multiple sub-communicators. The root process in each sub-communicator is responsible for performing the I/O operations for the entire group and then distributing the data to rest of the group. This approach resulted in over 25X speedup in HDF I/O read performance and 3X speedup in write performance for PFLOTRAN at over 100K processor cores on the ORNL Jaguar supercomputer. This research describes the design and development of a general purpose parallel I/O library called Scorpio that incorporates our optimized two-phase I/O approach. The library provides a simplified higher level abstraction to the user, sitting atop existing parallel I/O libraries (such as HDF5) and implements optimized I/O access patterns that can scale on larger number of processors. Performance results with standard benchmark problems and PFLOTRAN indicate that our library is able to maintain the same speedups as before with the added flexibility of being applicable to a wider range of I/O intensive applications.
Sarat Sreepathi, Vamsi Sripathi, Richard Tran Mills, Glenn Hammond, G. (Kumar) Mahinthakumar
HiPC1