Sai Narasimhamurthy

dblp:122/0990 · also Sai B. Narasimhamurthy · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-6448-7239ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 4 since 2021Computer networks · 2 · 2 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2025 Deep learning-based prediction of major page faults in cluster systems
Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy, Aladdin Ayesh
CCF Trans. High Perform. Comput.3
2024 From SSDs Back to HDDs: Optimizing VDO to Support Inline Deduplication and Compression for HDDs as Primary Storage Media
abstract
Deduplication and compression are powerful techniques to reduce the ratio between the quantity of logical data stored and the physical amount of consumed storage. Deduplication can impose significant performance overheads, as duplicate detection for large systems induces random accesses to the backend storage. These random accesses have led to the concern that deduplication for primary storage and HDDs are not compatible. Most inline data reduction solutions are therefore optimized for SSDs and discourage their use for HDDs, even for sequential workloads. In this work, we show that these concerns are valid if and only if the lessons learned from deduplication research are not applied. We have therefore investigated data reduction solutions for primary storage based on the RedHat Virtual Disk Optimizer (VDO) and show that directly applying them can decrease sequential write performance for HDDs by 36×. We then show that slight modifications to VDO plus the integration of a very small SSD area significantly improve performance even beyond the performance without data reduction enabled, making HDDs more cost-efficient for a wide range of mostly sequential cloud workloads than SSDs. Additionally, these VDO optimizations do not require to maintain different code bases for HDDs and SSDs, and we therefore provide the first data reduction solution applicable to both storage media.
Patrick Raaf, André Brinkmann, Eric Borba, Hossein Asadi 0001, Sai Narasimhamurthy, John Bent, Mohamad El-Batal, Reza Salkhordeh
ACM Trans. Storage5
2023 An empirical study of major page faults for failure diagnosis in cluster systems
Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy
J. Supercomput.3
2022 EMOSS '22: Workshop on Emerging Open Storage Systems and Solutions for Data Intensive Computing
abstract
Exciting changes are coming to the fore in the world of Storage and I/O for High Performance and Data Intensive computing. Existing data access and data retrieval methods that suitably support the new generation of data intensive applications in the realm of HPC and AI are being re-assessed. These new applications need to support both evolutionary as well as revolutionary approaches to data access and storage. Open source is enabling the adoption of these new techniques. This workshop looks at cutting edge trends in storage systems and solutions for data intensive HPC and AI applications with specific focus primarily on community driven initiatives and commercial products that may inspire the community.
Sai Narasimhamurthy, Glenn K. Lockwood
HPDC1
2022 NoaSci: A Numerical Object Array Library for I/O of Scientific Applications on Object Storage
abstract
The strong consistency and stateful workflow are seen as the major factors for limiting parallel I/O performance because of the need for locking and state management. While the POSIX-based I/O model dominates modern HPC storage infrastructure, emerging object storage technology can potentially improve I/O performance by eliminating these bottlenecks. Despite a wide deployment on the cloud, its adoption in HPC remains low. We argue one reason is the lack of a suitable programming interface for parallel I/O in scientific applications. In this work, we introduce NoaSci, a Numerical Object Array library for scientific applications. NoaSci supports different data formats (e.g. HDF5, binary), and focuses on supporting nodelocal burst buffers and object stores. We demonstrate for the first time how scientific applications can perform parallel I/O on Seagate’s Motr object store through NoaSci. We evaluate NoaSci’s preliminary performance using the iPIC3D space weather application and position against existing I/O methods.
Steven W. D. Chien, Artur Podobas, Martin Svedin, Andriy Tkachuk, Salem El Sayed, Pawel Andrzej Herman, Ganesan Umanesan, Sai Narasimhamurthy, Stefano Markidis
PDP8
2019 uMMAP-IO: User-Level Memory-Mapped I/O for HPC
abstract
The integration of local storage technologies alongside traditional parallel file systems on HPC clusters, is expected to rise the programming complexity on scientific applications aiming to take advantage of the increased-level of heterogeneity. In this work, we present uMMAP-IO, a user-level memory-mapped I/O implementation that simplifies data management on multi-tier storage subsystems. Compared to the memory-mapped I/O mechanism of the OS, our approach features per-allocation configurable settings (e.g., segment size) and transparently enables access to a diverse range of memory and storage technologies, such as the burst buffer I/O accelerators. Preliminary results indicate that uMMAP-IO provides at least 5-10x better performance on representative workloads in comparison with the standard memory-mapped I/O of the OS, and approximately 20-50% degradation on average compared to using conventional memory allocations without storage support up to 8192 processes.
Sergio Rivas-Gomez, Alessandro Fanfarillo, Sébastien Valat, Christophe Laferriere, Philippe Couvée, Sai Narasimhamurthy, Stefano Markidis
HiPC6
2019 Persistent coarrays: integrating MPI storage windows in coarray fortran
abstract
The inherent integration of novel hardware and software components on HPC is expected to considerably aggravate the Mean Time Between Failures (MTBF) on scientific applications, while simultaneously increase the programming complexity of these clusters. In this work, we present the initial steps towards the integration of transparent resilience support inside Coarray Fortran. In particular, we propose persistent coarrays, an extension of OpenCoarrays that integrates MPI storage windows to leverage its transport layer and seamlessly map coarrays to files on storage. Preliminary results indicate that our approach provides clear benefits on representative workloads, while incurring in minimal source code changes.
Sergio Rivas-Gomez, Alessandro Fanfarillo, Sai Narasimhamurthy, Stefano Markidis
EuroMPI3
2019 SAGE: Percipient Storage for Exascale Data Centric Computing
Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Stefano Markidis, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Dirk Pleiter, Shaun De Witt
Parallel Comput.1
2018 The SAGE project: a storage centric approach for exascale computing: invited paper
abstract
SAGE (Percipient StorAGe for Exascale Data Centric Computing) is a European Commission funded project towards the era of Exascale computing. Its goal is to design and implement a Big Data/Extreme Computing (BDEC) capable infrastructure with associated software stack. The SAGE system follows a storage centric approach as it is capable of storing and processing large data volumes at the Exascale regime.
Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Steven W. D. Chien, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Shaun De Witt, Dirk Pleiter, Stefano Markidis
CF1
2018 MPI windows on storage for HPC applications
Sergio Rivas-Gomez, Roberto Gioiosa, Ivy Bo Peng, Gokcen Kestor, Sai Narasimhamurthy, Erwin Laure, Stefano Markidis
Parallel Comput.5
2016 Improving Collective I/O Performance Using Non-volatile Memory Devices
abstract
Collective I/O is a parallel I/O technique designed to deliver high performance data access to scientific applications running on high-end computing clusters. In collective I/O, write performance is highly dependent upon the storage system response time and limited by the slowest writer. The storage system response time in conjunction with the need for global synchronisation, required during every round of data exchange and write, severely impacts collective I/O performance. Future Exascale systems will have an increasing number of processor cores, while the number of storage servers will remain relatively small. Therefore, the storage system concurrency level will further increase, worsening the global synchronisation problem. Nowadays high performance computing nodes also have access to locally attached solid state drives, effectively providing an additional tier in the storage hierarchy. Unfortunately, this tier is not always fully integrated. In this paper we propose a set of MPI-IO hints extensions that enable users to take advantage of fast, locally attached storage devices to boost collective I/O performance by increasing parallelism and reducing global synchronisation impact in the ROMIO implementation. We demonstrate that by using local storage resources, collective write performance can be greatly improved compared to the case in which only the global parallel file system is used, but can also decrease if the ratio between aggregators and compute nodes is too small.
Giuseppe Congiu, Sai Narasimhamurthy, Tim Süß, André Brinkmann
CLUSTER2
2016 Using Message Logs and Resource Use Data for Cluster Failure Diagnosis
abstract
Failure diagnosis for large compute clusters using only message logs is known to be incomplete. Recent availability of resource use data provides another potentially useful source of data for failure detection and diagnosis. Early work combining message logs and resource use data for failure diagnosis has shown promising results. This paper describes the CRUMEL framework which implements a new approach to combining rationalized message logs and resource use data for failure diagnosis. CRUMEL identifies patterns of errors and resource use and correlates these patterns by time with system failures. Application of CRUMEL to data from the Ranger supercomputer has yielded improved diagnoses over previous research. CRUMEL has: (i) showed that more events correlated with system failures can only be identified by applying different correlation algorithms, (ii) confirmed six groups of errors, (iii) identified Lustre I/O resource use counters which are correlated with occurrence of Lustre faults which are potential flags for online detection of failures, (iv) matched the dates of correlated error events and correlated resource use with the dates of compute node hang-ups and (v) identified two more error groups associated with compute node hang-ups. The pre-processed data will be put on the public domain in September, 2016.
Edward Chuah, Arshad Jhumka, James C. Browne, Nentawe Gurumdimma, Sai Narasimhamurthy, William L. Barth
HiPC5
2015 PIONEER: A Solution to Parallel I/O Workload Characterization and Generation
abstract
The demand for parallel I/O performance continues to grow. However, modelling and generating parallel I/O work-loads are challenging for several reasons including the large number of processes, I/O request dependencies and workload scalability. In this paper, we propose the PIONEER, a complete solution to Parallel I/O workload characterization and gEnERation. The core of PIONEER is a proposed generic workload path, which is essentially an abstract and dense representation of the parallel I/O patterns for all processes in a High Performance Computing (HPC) application. The generic workload path can be built via exploring the inter-processes correlations, I/O dependencies as well as file open session properties. We demonstrate the effectiveness of PIONEER by faithfully generating synthetic workloads for two popular HPC benchmarks and one real HPC application.
Weiping He, David Hung-Chang Du, Sai Narasimhamurthy
CCGRID3
2013 Linking Resource Usage Anomalies with System Failures from Cluster Log Data
abstract
Bursts of abnormally high use of resources are thought to be an indirect cause of failures in large cluster systems, but little work has systematically investigated the role of high resource usage on system failures, largely due to the lack of a comprehensive resource monitoring tool which resolves resource use by job and node. The recently developed TACC_Stats resource use monitor provides the required resource use data. This paper presents the ANCOR diagnostics system that applies TACC_Stats data to identify resource use anomalies and applies log analysis to link resource use anomalies with system failures. Application of ANCOR to first identify multiple sources of resource anomalies on the Ranger supercomputer, then correlate them with failures recorded in the message logs and diagnosing the cause of the failures, has identified four new causes of compute node soft lockups. ANCOR can be adapted to any system that uses a resource use monitor which resolves resource use by job.
Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy, John L. Hammond, James C. Browne, William L. Barth
SRDS3
2006 Coding schemes for integrated transport and storage reliability
abstract
IP storage area networks are based on bulk data transfer over a binary erasure channel and associated bulk data storage in end systems, which are also subjected to erasures. The paper motivates a possibility of an integrated approach to transport and data storage reliability in storage area networks (SAN) which reduces the protocol processing complexities. The integrated transport/storage reliability is sought through the novel Ying-Yang code and the convolution code with new decoding techniques.
Sai Narasimhamurthy, Joseph Y. Hui
IPCCC1
2005 Quanta data storage: an information processing and transportation architecture for storage area networks
abstract
A new architecture for storage area networks (SANs) is proposed for providing efficient information processing and transportation. We deviate from the byte stream oriented transmission control protocol (TCP) transport mechanism to more storage friendly block oriented transport. Data is processed, encrypted, error checked, redundantly encoded, and stored in fixed size blocks called quanta. Each quantum is processed by an effective cross-layer protocol that collapses the protocol stack for security, iWARP and iSCSI functions, transport control, and even redundant arrays of inexpensive disk (RAID) storage. This streamlining produces a highly efficient protocol with fewer memory copies and places most of the computational burden and security safeguard on the client, while the target stores quanta from many clients with minimal processing. We propose a new network RAID storage method using the quantum concept. Also, we unify error control and flow control of the iSCSI and TCP protocols in a manner we believe is more suitable for high data rate and low latency storage applications.
Sai Narasimhamurthy, Prabhanjan C. Gurumohan, S. Sreenivasamurthy, Joseph Y. Hui
IEEE J. Sel. Areas Commun.1
2004 Quanta Data Storage: A Cross Layer Architecture for the Storage Networks
Prabhanjan C. Gurumohan, Sai Narasimhamurthy, Joseph Y. Hui
MSST2