Seetharami R. Seelam

dblp:31/6720 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 8 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use Case
abstract
Scientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services.
Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer
CLOUD8
2022 NVMe Virtualization for Cloud Virtual Machines
abstract
Public clouds are rapidly moving to support Non-Volatile Memory Express (NVMe) based storage to meet the ever-increasing I/O throughput and latency demands of modern workloads. They provide NVMe storage through virtual machines (VMs) where multiple VMs running on a host may share a physical NVMe device. The virtualization method used to share the NVMe capability has important performance, usability and security implications. In this paper, we propose three NVMe storage virtualization methods: PCI device passthrough, virtual block device method, and Storage Performance Development Kit (SPDK) virtual host target method. We evaluate these virtualization methods in terms of performance, scalability, CPU overhead, technology maturity, security, and availability to use one or more of these methods in IBM public cloud.
Lixiang Luo, I-Hsin Chung, Seetharami R. Seelam, Ming-Hung Chen, Yun Joon Soh
ICPE3
2022 Best Practices for HPC Workloads on Public Cloud Platforms: A Guide for Computational Scientists to Use Public Cloud for HPC Workloads
abstract
HPC (high performance computing) applications come with a variety of requirements for computation, communication, and storage; and many of these requirements can be met with commodity technology available in public clouds. In this article, we report on results for several well-known HPC applications on IBM public cloud, and we describe best practices for running such applications on cloud systems in general. Our results show that public clouds are not only ready for HPC workloads, but they can provide performance comparable to, and in some cases better than, current supercomputers.
Robert Walkup, Seetharami R. Seelam, Sophia Wen
ICPE2
2017 Topology-aware GPU scheduling for learning workloads in cloud environments
abstract
Recent advances in hardware, such as systems with multiple GPUs and their availability in the cloud, are enabling deep learning in various domains including health care, autonomous vehicles, and Internet of Things. Multi-GPU systems exhibit complex connectivity among GPUs and between GPUs and CPUs. Workload schedulers must consider hardware topology and workload communication requirements in order to allocate CPU and GPU resources for optimal execution time and improved utilization in shared cloud environments.
Marcelo Amaral, Jorda Polo, David Carrera 0001, Seetharami R. Seelam, Malgorzata Steinder
SC4
2015 Polyglot Application Auto Scaling Service for Platform as a Service Cloud
abstract
Platform as a service (PaaS) is a cloud delivery model that provides software services and solution stacks to enable rapid development, deployment, and operations in many languages and run-times (polyglot). These applications require capabilities to rapidly grow and shrink the underlying resources to satisfy their workload needs. Auto scaling is a service that enables dynamic resource allocation and deal location to match application performance needs and service level agreements. In this paper we present the architecture and implementation of a polyglot auto scaling solution for IBM Blue mix PaaS. Our auto scaling service enables users to describe policies and set thresholds for scaling the applications based on CPU, memory and heap usage for applications developed in different languages (Java, Java Script, Ruby, etc). The auto scaling service consists of a set of monitoring agents, monitoring service, scaling service, and a persistence service. The service is developed with sharedmulti-tenancy model and offered as a managed cloud service. An application attached to the auto scaling service is monitored and its resources will be adjusted based on the auto scaling policies of the user and on the system conditions.
Seetharami R. Seelam, Paolo Dettori, Peter H. Westerink, Ben Bo Yang
IC2E1
2014 Blueprint for Business Middleware as a Managed Cloud Service
abstract
Cloud offers numerous technical middleware services such as databases, caches, messaging systems, and storage but very few business middleware services as first tier managed services. Business middleware such as business process management, business rules, operational decision management, content management and business analytics, if deployed in a cloud environment, is typically only available in a hosted (black-box) model. This is partly due to where cloud is in its evolution, and mostly due to the relatively higher complexity of business middleware vs. technical middleware in the deployment, provisioning, usage, etc. Business middleware consists of multiple functions for business processes design and modeling, execution, optimization, monitoring, and analysis. These functions and their associated complexity have inhibited the wholesale migration of existing business middleware to the cloud. To better understand the complexity in bringing business middleware to the cloud and to develop a systematic cloud enablement approach, we studied the deployment of IBM's Operational Decision Manager (ODM) business middleware product as a managed service (Cloud Decision Service) in IBM's BlueMix cloud platform. Our study indicates that complex middleware must be componentized along functional boundaries, and provide these functions for different business users and developers with cloud experience. In addition, middleware services must leverage other cloud services and they should provide interfaces so that they can be consumed by Java applications as well as by polyglot applications (JavaScript, Ruby, Python, etc). Applications can bind to and use our Cloud Decision Service in a matter of seconds. In contrast, it takes hours to days to setup such a service in the traditional packaged software model. Based on the lessons learned from this experiment we develop a blueprint for enabling high value business middleware as managed cloud services.
Paolo Dettori, David Frank, Seetharami R. Seelam, Pierre Feillet
IC2E3
2011 FAIRIO: An Algorithm for Differentiated I/O Performance
abstract
Providing differentiated service in a consolidated storage environment is a challenging task. To address this problem, we introduce FAIRIO, a cycle-based I/O scheduling algorithm that provides differentiated service to workloads concurrently accessing a consolidated RAID storage system. FAIRIO enforces proportional sharing of I/O service through fair scheduling of disk time. During each cycle of the algorithm, I/O requests are scheduled according to workload weights and disk-time utilization history. Experiments, which were driven by the I/O request streams of real and synthetic I/O benchmarks and run on a modified version of DiskSim, provide evidence of FAIRIO's effectiveness and demonstrate that fair scheduling of disk time is key to achieving differentiated service. In particular, the experimental results show that, for a broad range of workload request types, sizes, and access characteristics, the algorithm provides differentiated storage throughput that is within 10% of being perfectly proportional to workload weights, and, it achieves this with little or no degradation of aggregate throughput. The core design concepts of FAIRIO, including service-time allocation and history-driven compensation, potentially can be used to design I/O scheduling algorithms that provide workloads with differentiated service in storage systems comprised of RAIDs, multiple RAIDs, SANs, and hypervisors for Clouds.
Sarala Arunagiri, Yipkei Kwok, Patricia J. Teller, Ricardo Portillo, Seetharami R. Seelam
SBAC-PAD5
2010 Masking I/O latency using application level I/O caching and prefetching on Blue Gene systems
abstract
In this paper, we present an application-level I/O caching, prefetching, asynchronous system to hide access latency experienced by HPC applications. Our solution of user controllable caching and prefetching system maintains a file-IO cache in the user space of the application, analyzes the I/O access patterns, prefetches requests, and performs write-back of dirty data to storage asynchronously. So each time the application needs the data it does not have to pay the full I/O latency penalty in going to the storage and getting the required data. We have implemented this caching and asynchronous access system on the Blue Gene (BG/L and BG/P) systems. We present experimental results with NAS BT, MADbench, and WRF benchmarks. The results on BG/P system demonstrate that our method hides access latency, enhances application I/O access time by as much as 100%, and improves WRF execution time over 10%.
Seetharami R. Seelam, I-Hsin Chung, John Bauer, Hui-Fang Wen
IPDPS1
2010 Workload performance characterization of DARPA HPCS benchmarks
abstract
Abstract It is critical to understand the workload characteristics and resource usage patterns of available applications to guide the design and development of hardware and software stacks of future machines. In this article, we analyze the workload performance characteristics of three large‐scale DARPA HPCS benchmarks: Hybrid Coordinate Ocean Model, Parallel Ocean Program, and Lattice Boltzemann Magneto‐Hydrodynamics Code while executing on IBM Power5+ processor machines. Our analysis is focused on the CPU/memory performance using Cycles Per Instruction (CPI) model and multiprocess communication performance using MPI traces. For each benchmark, we provide a high‐level performance analysis followed by the hotspot analysis for selected input parameters. Then we present a detailed workload performance characterization using CPI model with data from a unique set of performance counters available on the Power5+ processor system. From communication performance analysis, we describe the sources of load imbalances in the applications and identify the potential impediments to the scalability of the applications under large processor counts. We identify several sources of performance problems that are potential bottlenecks and discuss methods to ameliorate them. We also present a comparative analysis of these benchmarks to summarize the similarities and differences in their performance characteristics. Copyright © 2009 John Wiley & Sons, Ltd.
Seetharami R. Seelam, I-Hsin Chung, Guojing Cong, Hui-Fang Wen, David J. Klepacki
Concurr. Comput. Pract. Exp.1
2009 Tools for scalable performance analysis on Petascale systems
abstract
Tools are becoming increasingly important to efficiently utilize the computing power available in contemporary large scale systems. The drastic increase in the size and the complexity of systems require tools to be scalable while producing meaning full and easily digestible information that may help the user pin-point problems at scale. The goal of this tutorial is to introduce some state-of-the-art performance tools from three different organizations to a diverse audience group. Together these tools provide a broad spectrum of capabilities necessary to analyze the performance of scientific and engineering applications on a variety of large and small scale systems. These tools include: • IBM High Performance Computing Toolkit: The IBM High Performance Computing Toolkit is a suite of performance-related tools and libraries to assist in application tuning. This toolkit is an integrated environment for performance analysis of sequential and parallel applications using the MPI and OpenMP paradigms. Scientists can collect rich performance data from selected parts of an execution, digest the data at a very high level, and plan for improvements within a single unified interface. It provides a common framework for IBM's mid-range server offerings, including pSeries and eSeries servers and Blue Gene systems, on both AIX and Linux. More information cab be found here: http://domino.research.ibm.com/comm/research_projects.nsf/pages/hpct.index.html • Scalable Performance Analysis of Large-Scale Applications (SCALASCA) Toolset: Scalasca is an open-source toolset that can be used to analyze the performance behavior of parallel applications and to identify opportunities for optimization. It has been specifically designed for use on large-scale systems including BlueGene and Cray XT, but is also well-suited for small- and medium-scale HPC platforms. Scalasca supports an incremental performance-analysis procedure that integrates runtime summaries with in-depth studies of concurrent behavior via event tracing, adopting a strategy of successively refined measurement configurations. A distinctive feature is the ability to identify wait states that occur, for example, as a result of unevenly distributed workloads. Especially when trying to scale communication-intensive applications to large processor counts, such wait states can present severe challenges to achieving good performance. Scalasca is developed by the Julich Supercomputing Centre and available under the New BSD open-source license. More information can be found here: http://www.scalasca.org • CEPBA Toolkit: The CEPBA-tools environment is a trace based analysis environment consisting with two major components, Paraver, a browser for traces obtained from a parallel run and Dimemas , a simulator to rebuild the time behavior of a parallel program from a trace. More information can be found here: http://www.bsc.es/plantillaF.php?cat_id=52
I-Hsin Chung, Seetharami R. Seelam, Bernd Mohr, Jesús Labarta
IPDPS2
2009 Towards a framework for automated performance tuning
abstract
As part of the DARPA sponsored high productivity computing systems (HPCS) program, IBM is building petaflop supercomputers that will be fast, power-efficient, and easy to program. In addition to high performance, high productivity to the end user is another prominent goal. The challenge is to develop technologies that bridge the productivity gap - the gap between the hardware complexity and the software limitations. In addition to language, compiler, and runtime research, powerful and user-friendly performance tools are critical in debugging performance problems and tuning for maximum performance. Traditional tools have either focused on specific performance aspects (e.g., communication problems) or provided limited diagnostic capabilities, and using them alone usually do not pinpoint accurately performance problems. Even fewer tools attempt to provide solutions for problems detected. In our study, we develop an open framework that unifies tools, compiler analysis, and expert knowledge to automatically analyze and tune the performance of an application. Preliminary results demonstrated the efficiency of our approach.
Guojing Cong, Seetharami R. Seelam, I-Hsin Chung, Sophia Wen, David J. Klepacki
IPDPS2
2009 Application level I/O caching on Blue Gene/P systems
abstract
In this paper, we present an application level aggressive I/O caching and prefetching system to hide I/O access latency experienced by out-of-core applications. Without the application level prefetching and caching capability, users of I/O intensive applications need to rewrite them with asynchronous I/O calls or restructure their code with MPI-IO calls to efficiently use the large scale system resources. Our proposed solution of user controllable aggressive caching and prefetching system maintains a file-IO cache in the user space of the application, analyzes the I/O access patterns, prefetches requests, and performs write-back of dirty data to storage asynchronously. So each time the application needs the data it does not have to pay the full I/O latency penalty in going to the storage and getting the required data. We have implemented this aggressive caching and asynchronous prefetching on the Blue Gene/P (BGP) system. The preliminary experiment evaluates the caching performance using the WRF benchmark. The results on BGP system demonstrate that our method improves application I/O throughput.
Seetharami R. Seelam, I-Hsin Chung, John Bauer, Hao Yu 0008, Hui-Fang Wen
IPDPS1
2009 Scalability Analysis of Job Scheduling Using Virtual Nodes
Norman Bobroff, Richard Coppinger, Liana L. Fong, Seetharami R. Seelam, Jing Xu 0012
JSSPP4
2008 Workload Performance Characterization of DARPA HPCS Benchmarks
abstract
It is critical to understand the workload characteristics and resource usage patterns of available applications to guide the design and development of hardware and software stacks of the future machines. In this paper, we analyze the workload performance characteristics of three large-scale DARPA HPCS benchmarks: HYCOM, POP, and LBMHD while executing on IBM Power5+ processor machines. Our analysis is focused on CPU/memory performance using cycles per instruction (CPI) model and multiprocess communication performance using MPI traces. For each benchmark, we provide a high level performance analysis followed by the hot-spot analysis of codes for selected input parameters.Then we present a detailed workload performance characterization using CPI model with data from a unique set of performance counters available on the Power5+ processor system. For communication, we describe the sources of load imbalances in the applications and identify the potential impediments to scalability of the applications under large processor counts.We identify several sources of performance problems that are potential bottlenecks and discuss methods to ameliorate them.
Seetharami R. Seelam, I-Hsin Chung, Guojing Cong, Hui-Fang Wen, David J. Klepacki
HPCC1
2008 A framework for automated performance bottleneck detection
abstract
In this paper, we present the architecture design and implementation of a framework for automated performance bottleneck detection. The framework analyzes the time-spent distribution in the application and discovers the performance bottlenecks by using given bottleneck definitions. The user can query the application execution performance to identify performance problems. The design of the framework is flexible and extensible so it can be tailored based on the actual application execution environment and performance tuning requirement. To demonstrate the usefulness of the framework, we apply the framework on a practical DARPA application and show how it helps to identify performance bottlenecks. The framework helps to automate the performance tuning process and improve the user’s productivity.
I-Hsin Chung, Guojing Cong, David J. Klepacki, Simone Sbaraglia, Seetharami R. Seelam, Hui-Fang Wen
IPDPS5
2008 Early experiences in application level I/O tracing on blue gene systems
abstract
On todays massively parallel processing (MPP) supercomputers, it is increasingly important to understand I/O performance of an application both to guide scalable application development and to tune its performance. These two critical steps are often enabled by performance analysis tools to obtain performance data on thousands of processors in an MPP system. To this end, we present the design, implementation, and early experiences of an application level I/O tracing library and the corresponding tool for analyzing and optimizing I/O performance on Blue Gene (BG) MPP systems. This effort was a part of IBM HPC Toolkit for BG systems. To our knowledge, this is the first comprehensive application-level I/O monitoring, playback, and optimizing tool available on BG systems. The preliminary experiments on popular NPB BTIO benchmark show that the tool is much useful on facilitating detailed I/O performance analysis.
Seetharami R. Seelam, I-Hsin Chung, Ding-Yong Hong, Hui-Fang Wen, Hao Yu 0008
IPDPS1
2007 Throttling I/O Streams to Accelerate File-IO Performance
Seetharami R. Seelam, Andre Kerstens, Patricia J. Teller
HPCC1
2007 Modeling the Impact of Checkpoints on Next-Generation Systems
Ron A. Oldfield, Sarala Arunagiri, Patricia J. Teller, Seetharami R. Seelam, Maria Ruiz Varela, Rolf Riesen, Philip C. Roth
MSST4
2007 Virtual I/O scheduler: a scheduler of schedulers for performance virtualization
abstract
Virtualized storage systems are required to service concurrently executing workloads, with potentially diverse data delivery requirements, that are running under multiple operating systems. Although a number of algorithms have been developed for I/O performance virtualization among operating system (OS) instances and their applications, none results in absolute performance virtualization. By absolute performance virtualization we mean that the performance experienced by applications of one operating system does not suffer due to variations in the I/O request stream characteristics of applications of other operating systems. Key requirements of I/O performance virtualization are fairness and performance isolation. In this paper, we present a novel virtual I/O scheduler (VIOS) that provides absolute performance virtualization by being fair in sharing I/O system resources among operating systems and their applications, and provides performance isolation in the face of variations in the characteristics of I/O streams. The VIOS controls the coarse grain allocation of disk time to the different operating system instances and is OS independent; optionally, a set of OS-dependent schedulers may determine the fine-grain interleaving of requests from the corresponding operating systems to the storage system.
Seetharami R. Seelam, Patricia J. Teller
VEE1
2006 Fairness and Performance Isolation: an Analysis of Disk Scheduling Algorithms
abstract
An I/O system using a sharing model provides concurrently executing applications shared access to the underlying I/O resources. Although the existing sharing models of these I/O systems are purported to be fair, none of them results in performance isolation. Failing to provide performance isolation results in unpredictable application performance. Unpredictability in application performance hampers providing quality of service guarantees. In this paper, we present a formal analysis of the fairness properties of various disk scheduling algorithms and an experimental evaluation of their performance isolation properties. We show that none of the existing "fair" scheduling algorithms provides performance isolation
Seetharami R. Seelam, Patricia J. Teller
CLUSTER1
2004 Profiling and Tracing OpenMP Applications with POMP Based Monitoring Libraries
Luiz De Rose, Bernd Mohr, Seetharami R. Seelam
Euro-Par3
2001 Statistical and Dempster-Shafer Techniques in Testing Structural Integrity of Aerospace Structures
Roberto A. Osegueda, Seetharami R. Seelam, Ana C. Holguin, Vladik Kreinovich, Chin-Wang Tao, Hung T. Nguyen 0002
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2