Lavanya Ramakrishnan

dblp:80/1690 · DBLP profile ↗
← Back
47ranked-venue papers
9as first author
6since 2021 · last 2024
0000-0003-1761-4132ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-authorSoftware engineering, systems software and programming languages · 6 · 2 first-authorArtificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Understanding Data Movement Patterns in HPC: A NERSC Case Study
abstract
Scientific experiments are producing unprecedented volumes of data with real-time High Performance Computing (HPC) needs. Understanding and ensuring efficient data movement in these emerging data-intensive workloads is becoming critical for successful workflow execution. The need for end-to-end that integrates compute, network, and storage resources across facilities is resulting in a new integrated infrastructure paradigm. In this paper, we present an extensive analysis of three years of network traffic data from NERSC and identify critical data movement trends while detecting bottlenecks that significantly curtail transfer performance. Our results show that data movement patterns have shifted in the three years, and current infrastructure cannot sufficiently handle competing transfers, leading up to 30% throughput degradation for individual flows. In addition, we provide design recommendations for data movement management in future integrated research infrastructures that aim to reduce data transfer latency, reducing overall time to scientific results.
Anna Giannakou, Damian Hazen, Bjoern Enders, Lavanya Ramakrishnan, Nicholas J. Wright
SC4
2024 SCIPIS: Scalable and concurrent persistent indexing and search in high-end computing systems
Alexandru Iulian Orhean, Anna Giannakou, Lavanya Ramakrishnan, Kyle Chard, Boris Glavic, Ioan Raicu
J. Parallel Distributed Comput.3
2022 SCANNS: Towards Scalable and Concurrent Data Indexing and Searching in High-End Computing System
Alexandru Iulian Orhean, Anna Giannakou, Lavanya Ramakrishnan, Kyle Chard, Ioan Raicu
CCGRID3
2022 Evaluation of a scientific data search infrastructure
abstract
Summary The ability to search over large scientific datasets has become crucial to next‐generation scientific discoveries as data generated from scientific facilities grow dramatically. In previous work, we developed and deployed ScienceSearch, a search infrastructure for scientific data which uses machine learning to automate metadata creation. Our current deployment is deployed atop a container based platform at a HPC center. In this article, we present an evaluation and discuss our experiences with the ScienceSearch infrastructure. Specifically, we present a performance evaluation of ScienceSearch's infrastructure focusing on scalability trends. The obtained results show that ScienceSearch is able to serve up to 130 queries/min with latency under 3 s. We discuss our infrastructure setup and evaluation results to provide our experiences and a perspective on opportunities and challenges of our search infrastructure.
Alexandru Iulian Orhean, Anna Giannakou, Katie Antypas, Ioan Raicu, Lavanya Ramakrishnan
Concurr. Comput. Pract. Exp.5
2021 Assessing data change in scientific datasets
abstract
Summary Scientific datasets are growing rapidly and becoming critical to next‐generation scientific discoveries. The validity of scientific results relies on the quality of data used and data are often subject to change, for example, due to observation additions, quality assessments, or processing software updates. The effects of data change are not well understood and difficult to predict. Datasets are often repeatedly updated and recomputing derived data products quickly becomes time consuming and resource intensive and may in some cases not even be necessary, thus delaying scientific advance. Despite its importance, there is a lack of systematic approaches for best comparing data versions to quantify the changes, and ad‐hoc or manual processes are commonly used. In this article, we propose a novel hierarchical approach for analyzing data changes, including real‐time (online) and offline analyses. We employ a variety of fast‐to‐compute numerical analyses, graphical data change representations, and more resource‐intensive recomputations of a subset of the data product. We illustrate the application of our approach using three scientific diverse use cases, namely, satellite, cosmological, and x‐ray data. The results show that a variety of data change metrics should be employed to enable a comprehensive representation and qualitative evaluation of data changes.
Juliane Mueller 0002, Boris Faybishenko, Deborah A. Agarwal, Chongya Jiang, Youngryel Ryu, Craig Tull, Lavanya Ramakrishnan
Concurr. Comput. Pract. Exp.8
2021 Programming Abstractions for Managing Workflows on Tiered Storage Systems
abstract
Scientific workflows in High Performance Computing ( HPC ) environments are processing large amounts of data. The storage hierarchy on HPC systems is getting deeper, driven by new technologies (NVRAMs, SSDs, etc.) There is a need for new programming abstractions that allow users to seamlessly manage data at the workflow level on multi-tiered storage systems, and provide optimal workflow performance and use of storage resources. In previous work, we introduced a software architecture Managing Data on Tiered Storage for Scientific Workflows (MaDaTS ) that used a Virtual Data Space ( VDS ) abstraction to hide the complexities of the underlying storage system while allowing users to control data management strategies. In this article, we detail the data-centric programming abstractions that allow users to manage a workflow around its data on the storage layer. The programming abstractions simplify data management for scientific workflows on multi-tiered storage systems, without affecting workflow performance or storage capacity. We measure the overheads and effectiveness introduced by the programming abstractions of MaDaTS. Our results show that these abstractions can optimally use the storage capacity in lesser capacity storage tiers, and simplify data management without adding any performance overheads.
Devarshi Ghoshal, Lavanya Ramakrishnan
ACM Trans. Storage2
2020 ASA - The Adaptive Scheduling Architecture
abstract
In High Performance Computing (HPC), resources are controlled by batch systems and may not be available due to long queue waiting times, negatively impacting application deadlines. This is noticeable in low latency scientific workflows where resource planning and timely allocation are key for efficient processing. On the one hand, peak allocations guarantee the fastest possible workflows execution time, at the cost of extended queue waiting times and costly resource usage. On the other hand, dynamic allocations following specific workflow stage requirements optimizes resource usage, though it increases the total workflow makespan. To enable new scheduling strategies and features in workflows, we propose ASA: the Adaptive Scheduling Architecture, a novel scheduling method to reduce perceived queue waiting times as well as to optimize workflows resource usage. Reinforcement learning is used to estimate queue waiting times, and based on these estimates ASA pro-actively submit resource change requests, minimizing total workflow inter-stage waiting times, idle resources, and makespan. Experiments with three scientific workflows at two HPC centers show that ASA combines the best of the two aforementioned approaches, with average queue waiting time and makespan reductions of up to 10% and 2% respectively, with up to 100% prediction accuracy, while obtaining near optimal resource utilization.
Abel Souza, Kristiaan Pelckmans, Devarshi Ghoshal, Lavanya Ramakrishnan, Johan Tordsson
HPDC4
2020 Performance characterization of scientific workflows for the optimal use of Burst Buffers
Christopher S. Daley, Devarshi Ghoshal, Glenn K. Lockwood, Sudip S. Dosanjh, Lavanya Ramakrishnan, Nicholas J. Wright
Future Gener. Comput. Syst.5
2019 Understanding Data Similarity in Large-Scale Scientific Datasets
abstract
Today, scientific experiments and simulations produce massive amounts of heterogeneous data that need to be stored and analyzed. Given that these large datasets are stored in many files, formats and locations, how can scientists find relevant data, duplicates or similarities? In this context, we concentrate on developing algorithms to compare similarity of time series for the purpose of search, classification and clustering. For example, generating accurate patterns from climate related time series is important not only for building models for weather forecasting and climate prediction, but also for modeling and predicting the cycle of carbon, water, and energy. We developed the methodology and ran an exploratory analysis of climatic and ecosystem variables from the FLUXNET2015 dataset. The proposed combination of similarity metrics, nonlinear dimension reduction, clustering methods and validity measures for time series data has never been applied to unlabeled datasets before, and provides a process that can be easily extended to other scientific time series data. The dimensionality reduction step provides a good way to identify the optimum number of clusters, detect outliers and assign initial labels to the time series data. We evaluated multiple similarity metrics, in terms of the internal cluster validity for driver as well as response variables. While the best metric often depends on a number of factor, the Euclidean distance seems to perform well for most variables and also in terms of computational expense.
Payton Linton, William Melodia, Alina Lazar, Deborah A. Agarwal, Ludovico Bianchi, Devarshi Ghoshal, Gilberto Zonta Pastorello, Lavanya Ramakrishnan, Kesheng Wu
IEEE BigData8
2019 Data Jockey: Automatic Data Management for HPC Multi-tiered Storage Systems
abstract
We present the design and implementation of Data Jockey, a data management system for HPC multi-tiered storage systems. As a centralized data management control plane, Data Jockey automates bulk data movement and placement for scientific workflows and integrates into existing HPC storage infrastructures. Data Jockey simplifies data management by eliminating human effort in programming complex data movements, laying datasets across multiple storage tiers when supporting complex workflows, which in turn increases the usability of multitiered storage systems emerging in modern HPC data centers. Specifically, Data Jockey presents a new data management scheme called “goal driven data management” that can automatically infer low-level bulk data movement plans from declarative high-level goal statements that come from the lifetime of iterative runs of scientific workflows. While doing so, Data Jockey aims to minimize data wait times by taking responsibility for datasets that are unused or to be used, and aggressively utilizing the capacity of the upper, higher performant storage tiers. We evaluated a prototype implementation of Data Jockey under a synthetic workload based on a year's worth of Oak Ridge Leadership Computing Facility's (OLCF) operational logs. Our evaluations suggest that Data Jockey leads to higher utilization of the upper storage tiers while minimizing the programming effort of data movement compared to human involved, per-domain adhoc data management scripts.
Woong Shin, Christopher Brumgard, Sudharshan S. Vazhkudai, Devarshi Ghoshal, Sarp Oral, Lavanya Ramakrishnan
IPDPS7
2018 Bringing Data Science to Qualitative Analysis
abstract
Qualitative user research is a human-intensive approach that draws upon ethnographic methods from social sciences to develop insights about work practices to inform software design and development. Recent advances in data science, and in particular, natural language processing (NLP), enables the derivation of machine-generated insights to augment existing techniques. Our work describes our prototype framework based in Jupyter, a software tool that supports interactive data science and scientific computing, that leverages NLP techniques to make sense of transcribed texts from user interviews. This work also serves as a starting point for incorporating data science techniques in the qualitative analyses process.
You-Wei Cheah, Drew Paine, Devarshi Ghoshal, Lavanya Ramakrishnan
eScience4
2018 ScienceSearch: Enabling Search through Automatic Metadata Generation
abstract
Scientific facilities are increasingly generating and handling large amounts of data from experiments and simulations. Next-generation scientific discoveries rely on insights derived from data, especially across domain boundaries. Search capabilities are critical to enable scientists to discover datasets of interest. However, scientific datasets often lack the signals or metadata required for effective searches. Thus, we need formalized methods and systems to automatically annotate scientific datasets from the data and its surrounding context. Additionally, a search infrastructure needs to account for the scale and rate of application data volumes. In this paper, we present ScienceSearch, a system infrastructure that uses machine learning techniques to capture and learn the knowledge, context, and surrounding artifacts from data to generate metadata and enable search. Our current implementation is focused on a dataset from the National Center for Electron Microscopy (NCEM), an electron microscopy facility at Lawrence Berkeley National Laboratory sponsored by the Department of Energy which supports hundreds of users and stores millions of micrographs. In this paper, we describe (a) our search infrastructure and model, (b) methods for generating metadata using machine learning techniques, and (c) optimizations to improve search latency, and deployment on an HPC system. We demonstrate that ScienceSearch is capable of producing valid metadata for NCEM's dataset and providing low-latency good quality search results over a scientific dataset.
Gonzalo Pedro Rodrigo Álvarez, Matthew L. Henderson, Gunther H. Weber, Colin Ophus, Katie Antypas, Lavanya Ramakrishnan
eScience6
2018 Dac-Man: data change management for scientific datasets on HPC systems
Devarshi Ghoshal, Lavanya Ramakrishnan, Deborah A. Agarwal
SC2
2018 Towards understanding HPC users and systems: A NERSC case study
Gonzalo Pedro Rodrigo Álvarez, Per-Olov Östberg, Erik Elmroth, Katie Antypas, Richard A. Gerber, Lavanya Ramakrishnan
J. Parallel Distributed Comput.6
2017 Ten Principles for Creating Usable Software for Science
abstract
The volume and variety of scientific data being generated at experimental facilities requires the seamless interaction of the scientist's knowledge with the large-scale machines and software that is required to process the data. In the last few years, scientific software tools are being developed to address these increasingly complex workflow and data management needs. However, current approaches for designing systems and tools focus on the hardware and software of the machine and do not consider the user. Our experience shows us that user experience research needs to be tightly integrated with the software development life cycle for building sustainable software for science. It has become not just necessary, but critical, to consider the user interaction in the design of the entire system for data-intensive sciences that have complex human interaction with the data, software and systems. The dynamic nature of science projects and the complex roles of personnel in the projects makes it difficult to apply classical user research methodologies from industry. In this paper, we make three specific contributions towards improving the usability and sustainability of scientific software. First, we examine the software life cycle in science environments and identify the differences with commercial software development. Next, we outline ten principles we have developed to guide user engagement and software development and illustrate it with examples from our projects over the last several years. Finally, we provide guidelines to other eScience projects on applying the ten principles in the software development life cycle.
Lavanya Ramakrishnan, Dan Gunter
eScience1
2017 Enabling Workflow-Aware Scheduling on HPC Systems
abstract
Scientific workflows are increasingly common in the workloads of current High Performance Computing (HPC) systems. However, HPC schedulers do not incorporate workflow-specific mechanisms beyond the capacity to declare dependencies between their jobs. Thus, workflows are run as sets of batch jobs with dependencies, which induces long intermediate wait times and, consequently, long workflow turnaround times. Alternatively, to reduce their turnaround time, workflows may be submitted as single pilot jobs that are allocated their maximum required resources for their entire runtime. Pilot jobs achieve shorter turnaround times but reduce the HPC system's utilization because resources may idle during the workflow's execution. We present a workflow-aware scheduling (WoAS) system that enables existing scheduling algorithms to exploit fine-grained information on a workflow's resource requirements and structure without modification. The current implementation of WoAS is integrated into Slurm, a widely used HPC batch scheduler. We evaluate the system using a simulator using real and synthetic workflows and a synthetic baseline workload that captures job patterns observed over three years of workload data from Edison, a large supercomputer hosted at the National Energy Research Scientific Computing Center. Our results show that WoAS reduces workflow turnaround times and improves system utilization without significantly slowing down conventional jobs.
Gonzalo Pedro Rodrigo Álvarez, Erik Elmroth, Per-Olov Östberg, Lavanya Ramakrishnan
HPDC4
2017 MaDaTS: Managing Data on Tiered Storage for Scientific Workflows
abstract
Scientific workflows are increasingly used in High Performance Computing (HPC) environments to manage complex simulation and analyses, often consuming and generating large amounts of data. However, workflow tools have limited support for managing the input, output and intermediate data. The data elements of a workflow are often managed by the user through scripts or other ad-hoc mechanisms. Technology advances for future HPC systems is redefining the memory and storage subsystem by introducing additional tiers to improve the I/O performance of data-intensive applications. These architectural changes introduce additional complexities to managing data for scientific workflows. Thus, we need to manage the scientific workflow data across the tiered storage system on HPC machines. In this paper, we present the design and implementation of MaDaTS (Managing Data on Tiered Storage for Scientific Workflows), a software architecture that manages data for scientific workflows. We introduce Virtual Data Space (VDS), an abstraction of the data in a workflow that hides the complexities of the underlying storage system while allowing users to control data management strategies. We evaluate the data management strategies with real scientific and synthetic workflows, and demonstrate the capabilities of MaDaTS. Our experiments demonstrate the flexibility, performance and scalability gains of MaDaTS as compared to the traditional approach of managing data in scientific workflows.
Devarshi Ghoshal, Lavanya Ramakrishnan
HPDC2
2017 ScSF: A Scheduling Simulation Framework
Gonzalo Pedro Rodrigo Álvarez, Erik Elmroth, Per-Olov Östberg, Lavanya Ramakrishnan
JSSPP4
2017 Web-based visual data exploration for improved radiological source detection
abstract
Summary Radiation detection can provide a reliable means of detecting radiological material. Such capabilities can help to prevent nuclear and/or radiological attacks, but reliable detection in uncontrolled surroundings requires algorithms that account for environmental background radiation. The Berkeley Data Cloud (BDC) facilitates the development of such methods by providing a framework to capture, store, analyze, and share data sets. In the era of big data, both the size and variety of data make it difficult to explore and find data sets of interest and manage the data. Thus, in the context of big data, visualization is critical for checking data consistency and validity, identifying gaps in data coverage, searching for data relevant to an analyst's use cases, and choosing input parameters for analysis. Downloading the data and exploring it on an analyst's desktop using traditional tools are no longer feasible due to the size of the data. This paper describes the design and implementation of a visualization system that addresses the problems associated with data exploration within the context of the BDC. The visualization system is based on a JavaScript front end communicating via REST with a back end web server.
Gunther H. Weber, Mark S. Bandstra, Daniel Chivers, Hamdy H. Elgammal, Valerie C. Hendrix, John Kua, Jonathan S. Maltz, Krishna Muriki, Yeongshnn Ong, Michael J. Quinlan, Lavanya Ramakrishnan, Brian J. Quiter
Concurr. Comput. Pract. Exp.12
2016 Towards Understanding Job Heterogeneity in HPC: A NERSC Case Study
abstract
The high performance computing (HPC) scheduling landscape is changing. Increasingly, there are large scientific computations that include high-throughput, data-intensive, and stream-processing compute models. These jobs increase the workload heterogeneity, which presents challenges for classical tightly coupled MPI job oriented HPC schedulers. Thus, it is important to define new analyses methods to understand the heterogeneity of the workload, and its possible effect on the performance of current systems. In this paper, we present a methodology to assess the job heterogeneity in workloads and scheduling queues. We apply the method on the workloads of three current National Energy Research Scientific Computing Center (NERSC) systems in 2014. Finally, we present the results of such analysis, with an observation that heterogeneity might reduce predictability in the jobs' wait time.
Gonzalo Pedro Rodrigo Álvarez, Per-Olov Östberg, Erik Elmroth, Katie Antypas, Richard A. Gerber, Lavanya Ramakrishnan
CCGrid6
2016 Tigres Workflow Library: Supporting Scientific Pipelines on HPC Systems
abstract
The growth in scientific data volumes has resulted in the need for new tools that enable users to operate on and analyze data on large-scale resources. In the last decade, a number of scientific workflow tools have emerged. These tools often target distributed environments, and often need expert help to compose and execute the workflows. Data-intensive workflows are often ad-hoc, they involve an iterative development process that includes users composing and testing their workflows on desktops, and scaling up to larger systems. In this paper, we present the design and implementation of Tigres, a workflow library that supports the iterative workflow development cycle of data-intensive workflows. Tigres provides an application programming interface to a set of programming templates i.e., sequence, parallel, split, merge, that can be used to compose and execute computational and data pipelines. We discuss the results of our evaluation of scientific and synthetic workflows showing Tigres performs with minimal template overheads (mean of 13 seconds over all experiments). We also discuss various factors (e.g., I/O performance, execution mechansims) that affect the performance of scientific workflows on HPC systems.
Valerie C. Hendrix, James Fox, Devarshi Ghoshal, Lavanya Ramakrishnan
CCGrid4
2016 Considering Time in Designing Large-Scale Systems for Scientific Computing
abstract
High performance computing (HPC) has driven collaborative science discovery for decades. Exascale computing platforms, currently in the design stage, will be deployed around 2022. The next generation of supercomputers is expected to utilize radically different computational paradigms, necessitating fundamental changes in how the community of scientific users will make the most efficient use of these powerful machines. However, there have been few studies of how scientists work with exascale or close-to-exascale HPC systems. Time as a metaphor is so pervasive in the discussions and valuation of computing within the HPC community that it is worthy of close study. We utilize time as a lens to conduct an ethnographic study of scientists interacting with HPC systems. We build upon recent CSCW work to consider temporal rhythms and collective time within the HPC sociotechnical ecosystem and provide considerations for future system design.
Nan-Chen Chen, Sarah S. Poon, Lavanya Ramakrishnan, Cecilia R. Aragon
CSCW3
2016 Cloud computing for data-driven science and engineering
abstract
Cloud computing for data-driven science and engineering* During the past decade, data-driven science and engineering have emerged as a key paradigm for performing scientific research, enabling innovations through new kinds of experiments that were earlier impossible.Today's science has access to advanced instruments like next generation genome sequencers, gigapixel survey telescopes, and networks of sensors that monitor cyber-physical systems, and these are generating datasets that are growing exponentially in complexity and data volume.Big Data, across all dimensions of volume, velocity, variety, and veracity, are offering unique opportunities to enable scientific discovery as well as novel challenges to scientific platforms.Such dynamic, distributed, and data-intensive applications hold the solutions to vital scientific and societal problems of the 21st century.In order to achieve breakthrough in new knowledge, there is a need to develop data-driven system models, perform analytics at large scales, manage data from instruments and analyses, and share and visualize the results with scientific peers and the society at large.To this end, cloud computing offers a computing model for running such data-intensive scientific and engineering applications.Clouds have democratized resource access to underserved disciplines, making it possible to perform nontrivial scientific explorations for just a few hundred dollars.Clouds are particularly cost-effective for Big Data applications due to their co-location of elastic compute resources with data, and their use of commodity hardware, which economizes on costs for non-high performance computing (HPC) workloads.Many contemporary Big Data platforms that have emerged from online enterprises such as Google and Twitter are also optimized for such commodity hardware, as found in their own data centers.Of course, there are costs associated with data transfer and storage, in keeping with the pay-as-you-go model, that may not be well suited for applications requiring frequent transfer of and long-term storage of terabytes of data.Likewise, it is valuable to understand how data-intensive or even HPC applications that have been developed for computing grids, at one end, and applications developed for workstation tools like MATLAB and R, at the other end, can be effectively run on clouds.These are some of the practical realities that are worth exploring on the relevance of clouds for data-driven scientific applications.In this special issue, we have compiled a set of articles that discusses new research, development, and deployment efforts in running eScience and eEngineering workloads on cloud infrastructures and platforms.The open solicitation, which followed the 3rd Workshop on Scientific Cloud Computing (ScienceCloud), invited research and case studies on a variety of topics relevant to data-driven scientific computing on clouds: use of cloud-based technologies to address innovative compute and data-driven scientific problems that are not well served by current HPC clusters and grids, programming platforms for elastic and Big Data applications, performance and cost-effective computing on clouds, and gaps in diverse cloud fabrics and service offerings, among others.In all, the special issue received 28 articles, of which six were selected for publication after multiple rounds of reviews and revisions.The special issue starts with two articles that explore the runtime platform support required for executing Big Data science on clouds.In TomusBlobs: Scalable Data-Intensive Processing on Azure Clouds [1], the authors address the limitations of data storage within IaaS clouds such as Amazon S3 and Microsoft Azure BLOBs that are, while co-located in the data center, not present in the virtual machines (VMs) and need to be accessed over the network.Their distributed storage on VMs is optimized for concurrent access and elastic scaling, even *Corrections added on 5 February 2016, after first online publication: references to the paper "Pilot-abstractions for distributed data-intensive cloud applications" have been removed.
Yogesh L. Simmhan, Lavanya Ramakrishnan, Gabriel Antoniu, Carole A. Goble
Concurr. Comput. Pract. Exp.2
2016 Processing Cassandra Datasets with Hadoop-Streaming Based Approaches
abstract
The progressive transition in the nature of both scientific and industrial datasets has been the driving force behind the development and research interests in the NoSQL model. Loosely structured data poses a challenge to traditional data store systems, and when working with the NoSQL model, these systems are often considered impractical and costly. As the quantity and quality of unstructured data grows, so does the demand for a processing pipeline that is capable of seamlessly combining the NoSQL storage model and a “Big Data” processing platform such as MapReduce. Although MapReduce is the paradigm of choice for data-intensive computing, Java-based frameworks such as Hadoop require users to write MapReduce code in Java while Hadoop Streaming module allows users to define non-Java executables as map and reduce operations. When confronted with legacy C/C++ applications and other non-Java executables, there arises a further need to allow NoSQL data stores access to the features of Hadoop Streaming. We present approaches in solving the challenge of integrating NoSQL data stores with MapReduce under non-Java application scenarios, along with advantages and disadvantages of each approach. We compare Hadoop Streaming alongside our own streaming framework, MARISSA, to show performance implications of coupling NoSQL data stores like Cassandra with MapReduce frameworks that normally rely on file-system based data stores. Our experiments also include Hadoop-C*, which is a setup where a Hadoop cluster is co-located with a Cassandra cluster in order to process data using Hadoop with non-java executables.
Elif Dede, Bedri Sendir, Pinar Kuzlu, J. Weachock, Madhusudhan Govindaraju, Lavanya Ramakrishnan
IEEE Trans. Serv. Comput.6
2015 HPC System Lifetime Story: Workload Characterization and Evolutionary Analyses on NERSC Systems
abstract
High performance computing centers have traditionally served monolithic MPI applications. However, in recent years, many of the large scientific computations have included high throughput and data-intensive jobs. HPC systems have mostly used batch queue schedulers to schedule these workloads on appropriate resources. There is a need to understand future scheduling scenarios that can support the diverse scientific workloads in HPC centers. In this paper, we analyze the workloads on two systems (Hopper, Carver) at the National Energy Research Scientific Computing (NERSC) Center. Specifically, we present a trend analysis towards understanding the evolution of the workload over the lifetime of the two systems.
Gonzalo Pedro Rodrigo Álvarez, Per-Olov Östberg, Erik Elmroth, Katie Antypas, Richard A. Gerber, Lavanya Ramakrishnan
HPDC6
2015 AnalyzeThis: an analysis workflow-aware storage system
abstract
The need for novel data analysis is urgent in the face of a data deluge from modern applications. Traditional approaches to data analysis incur significant data movement costs, moving data back and forth between the storage system and the processor. Emerging Active Flash devices enable processing on the flash, where the data already resides. An array of such Active Flash devices allows us to revisit how analysis workflows interact with storage systems. By seamlessly blending together the flash storage and data analysis, we create an analysis workflow-aware storage system, AnalyzeThis. Our guiding principle is that analysis-awareness be deeply ingrained in each and every layer of the storage, elevating data analyses as first-class citizens, and transforming AnalyzeThis into a potent analytics-aware appliance. We implement the AnalyzeThis storage system atop an emulation platform of the Active Flash array. Our results indicate that AnalyzeThis is viable, expediting workflow execution and minimizing data movement.
Hyogi Sim, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Devesh Tiwari, Ali Anwar 0001, Ali Raza Butt, Lavanya Ramakrishnan
SC7
2015 Performance and energy efficiency of big data applications in cloud environments: A Hadoop case study
Eugen Feller, Lavanya Ramakrishnan, Christine Morin
J. Parallel Distributed Comput.2
2014 Experiences with User-Centered Design for the Tigres Workflow API
abstract
Scientific data volumes have been growing exponentially. This has resulted in the need for new tools that enable users to operate on and analyze data. Cyber infrastructure tools, including workflow tools, that have been developed in the last few years has often fallen short if user needs and suffered from lack of wider adoption. User-centered Design (UCD) process has been used as an effective approach to develop usable software with high adoption rates. However, UCD has largely been applied for user-interfaces and there has been limited work in applying UCD to application program interfaces and cyber infrastructure tools. We use an adapted version of UCD that we refer to as Scientist-Centered Design (SCD) to engage with users in the design and development of Tigres, a workflow application programming interface. Tigres provides a simple set of programming templates (e.g., sequence, parallel, split, merge) that can be can used to compose and execute computational and data transformation pipelines. In this paper, we describe Tigres and discuss our experiences with the use of UCD for the initial development of Tigres. Our experience-to-date is that the UCD process not only resulted in better requirements gathering but also heavily influenced the architecture design and implementation details. User engagement during the development of tools such as Tigres is critical to ensure usability and increase adoption.
Lavanya Ramakrishnan, Sarah S. Poon, Valerie C. Hendrix, Dan Gunter, Gilberto Zonta Pastorello, Deborah A. Agarwal
eScience1
2014 Provisioning, Placement and Pipelining Strategies for Data-Intensive Applications in Cloud Environments
abstract
Clouds are increasingly being used for running data-intensive scientific applications. Data-intensive science applications need performance, scalability and reliability. However, these can be hard to achieve in cloud environments. Intelligent strategies are required to obtain better performance, scalability and reliability on cloud platforms. In this paper, we propose a set of pipelining strategies to effectively utilize provisioned cloud resources. Our experiments on the ExoGENI cloud testbed demonstrates the effectiveness of our approach in increasing performance and reducing failures.
Devarshi Ghoshal, Lavanya Ramakrishnan
IC2E2
2014 Benchmarking MapReduce implementations under different application scenarios
Elif Dede, Zacharia Fadika, Madhusudhan Govindaraju, Lavanya Ramakrishnan
Future Gener. Comput. Syst.4
2014 MARIANE: Using MApReduce in HPC environments
Zacharia Fadika, Elif Dede, Madhusudhan Govindaraju, Lavanya Ramakrishnan
Future Gener. Comput. Syst.4
2014 CAMP: Community Access MODIS Pipeline
Valerie C. Hendrix, Lavanya Ramakrishnan, Youngryel Ryu, Catharine van Ingen, Keith R. Jackson, Deborah A. Agarwal
Future Gener. Comput. Syst.2
2013 On the performance and energy efficiency of Hadoop deployment models
abstract
The exponential growth of scientific and business data has resulted in the evolution of the cloud computing and the MapReduce parallel programming model. Cloud computing emphasizes increased utilization and power savings through consolidation while MapReduce enables large scale data analysis. The Hadoop framework has recently evolved to the standard framework implementing the MapReduce model. In this paper, we evaluate Hadoop performance in both the traditional model of collocated data and compute services as well as consider the impact of separating out the services. The separation of data and compute services provides more flexibility in environments where data locality might not have a considerable impact such as virtualized environments and clusters with advanced networks. In this paper, we also conduct an energy efficiency evaluation of Hadoop on physical and virtual clusters in different configurations. Our extensive evaluation shows that: (1) performance on physical clusters is significantly better than on virtual clusters; (2) performance degradation due to separation of the services depends on the data to compute ratio; (3) application completion progress correlates with the power consumption and power consumption is heavily application specific.
Eugen Feller, Lavanya Ramakrishnan, Christine Morin
IEEE BigData2
2012 Evaluating Hadoop for Data-Intensive Scientific Operations
abstract
Emerging sensor networks, more capable instruments, and ever increasing simulation scales are generating data at a rate that exceeds our ability to effectively manage, curate, analyze, and share it. Data-intensive computing is expected to revolutionize the next-generation software stack. Hadoop, an open source implementation of the MapReduce model provides a way for large data volumes to be seamlessly processed through use of large commodity computers. The inherent parallelization, synchronization and fault-tolerance the model offers, makes it ideal for highly-parallel data-intensive applications. MapReduce and Hadoop have traditionally been used for web data processing and only recently been used for scientific applications. There is a limited understanding on the performance characteristics that scientific data intensive applications can obtain from MapReduce and Hadoop. Thus, it is important to evaluate Hadoop specifically for data-intensive scientific operations -- filter, merge and reorder-- to understand its various design considerations and performance trade-offs. In this paper, we evaluate Hadoop for these data operations in the context of High Performance Computing (HPC) environments to understand the impact of the file system, network and programming modes on performance.
Zacharia Fadika, Madhusudhan Govindaraju, Shane Canon, Lavanya Ramakrishnan
IEEE CLOUD4
2012 MARISSA: MApReduce Implementation for Streaming Science Applications
abstract
MapReduce has since its inception been steadily gaining ground in various scientific disciplines ranging from space exploration to protein folding. The model poses a challenge for a wide range of current and legacy scientific applications for addressing their “Big Data” challenges. For example: MapRe-duce's best known implementation, Apache Hadoop, only offers native support for Java applications. While Hadoop streaming supports applications compiled in a variety of languages such as C, C++, Python and FORTRAN, streaming has shown to be a less efficient MapReduce alternative in terms of performance, and effectiveness. Additionally, Hadoop streaming offers lesser options than its native counterpart, and as such offers less flexibility along with a limited array of features for scientific software. The Hadoop File System (HDFS), a central pillar of Apache Hadoop is not a POSIX compliant file system. In this paper, we present an alternative framework to Hadoop streaming to address the needs of scientific applications: MARISSA (MApReduce Implementation for Streaming Science Applications). We describe MARISSA's design and explain how it expands the scientific applications that can benefit from the MapReduce model. We also compare and explain the performance gains of MARISSA over Hadoop streaming.
Elif Dede, Zacharia Fadika, Jessica Hartog, Madhusudhan Govindaraju, Lavanya Ramakrishnan, Dan Gunter, Shane Canon
eScience5
2011 Adapting MapReduce for HPC environments
abstract
MapReduce is increasingly gaining popularity as a programming model for use in large-scale distributed processing. The model is most widely used when implemented using the Hadoop Distributed File System (HDFS). The use of the HDFS, however, precludes the direct applicability of the model to HPC environments, which use high performance distributed file systems. In such distributed environments, the MapReduce model can rarely make use of full resources, as local disks may not be available for data placement on all the nodes. This work proposes a MapReduce implementation and design choices directly suitable for such HPC environments.
Zacharia Fadika, Elif Dede, Madhusudhan Govindaraju, Lavanya Ramakrishnan
HPDC4
2011 Deadline-sensitive workflow orchestration without explicit resource control
Lavanya Ramakrishnan, Jeffrey S. Chase, Dennis Gannon, Daniel Nurmi, Richard Wolski
J. Parallel Distributed Comput.1
2010 On-demand Overlay Networks for Large Scientific Data Transfers
abstract
Large scale scientific data transfers are central to scientific processes. Data from large experimental facilities have to be moved to local institutions for analysis or often data needs to be moved between local clusters and large supercomputing centers. In this paper, we propose and evaluate a network overlay architecture to enable high-throughput, on-demand, coordinated data transfers over wide-area networks. Our work leverages Phoebus and On-demand Secure Circuits and Advance Reservation System (OSCARS) to provide high performance wide-area network connections. OSCARS enables dynamic provisioning of network paths with guaranteed bandwidth and Phoebus enables the coordination and effective utilization of the OSCARS network paths. Our evaluation shows that this approach leads to improved end-to-end data transfer throughput with minimal overheads. The achieved throughput using our overlay was limited only by the ability of the end hosts to sink the data.
Lavanya Ramakrishnan, Chin Guok, Keith R. Jackson, Ezra Kissel, D. Martin Swany, Deborah A. Agarwal
CCGRID1
2010 WORKEM: Representing and Emulating Distributed Scientific Workflow Execution State
abstract
Scientific workflows have become an integral part of cyberinfrastructure as their computational complexity and data sizes have grown. However, the complexity of the distributed infrastructure makes design of new workflows, determining the right management policies, debugging, testing or reproduction of errors challenging. Today, workflow engines manage the dependencies between tasks of workflows and there are tools available to wrap scientific codes. There is a need for a customizable, isolated and manageable testing container for design, evaluation and deployment of distributed workflows. To build such an environment, we need to be able to model and represent, capture and possibly reuse the execution flows within each task of a workflow that accurately captures the execution behavior. In this paper, we present the design and implementation of WORKEM, an extensible framework that can be used to represent and emulate workflow execution state. We also detail the use of the framework in two specific case studies (a) design and testing of an orchestration system (b) generation of a provenance database. Our evaluation shows that the framework has minimal overheads and can be scaled to run hundreds of workflows in short durations of time and with a high amount of parallelism.
Lavanya Ramakrishnan, Dennis Gannon, Beth Plale
CCGRID1
2010 Defining future platform requirements for e-Science clouds
abstract
Cloud computing has evolved in the commercial space to support highly asynchronous web 2.0 applications. Scientific computing has traditionally been supported by centralized federally funded supercomputing centers and grid resources with a focus on bulk-synchronous compute and data-intensive applications. The scientific computing community has shown increasing interest in exploring cloud computing to serve e-Science applications, with the idea of taking advantage of some of its features such as customizable environments and on-demand resources. Magellan, a recently funded cloud computing project is investigating how cloud computing can serve the needs of mid-range computing and future data-intensive scientific workloads. This paper summarizes the application requirements and business model needed to support the requirements of both existing and emerging science applications, as learned from the early experiences on Magellan and commercial cloud environments. We provide an overview of the capabilities of leading cloud offerings and identify the existent gaps and challenges. Finally, we discuss how the existing cloud software stack may be evolved to better meet e-Science needs, along with the implications for resource providers and middleware developers.
Lavanya Ramakrishnan, Keith R. Jackson, Shane Canon, Shreyas Cholia, John Shalf
SoCC1
2010 Performance Analysis of High Performance Computing Applications on the Amazon Web Services Cloud
abstract
Cloud computing has seen tremendous growth, particularly for commercial web applications. The on-demand, pay-as-you-go model creates a flexible and cost-effective means to access compute resources. For these reasons, the scientific computing community has shown increasing interest in exploring cloud computing. However, the underlying implementation and performance of clouds are very different from those at traditional supercomputing centers. It is therefore critical to evaluate the performance of HPC applications in today's cloud environments to understand the tradeoffs inherent in migrating to the cloud. This work represents the most comprehensive evaluation to date comparing conventional HPC platforms to Amazon EC2, using real applications representative of the workload at a typical supercomputing center. Overall results indicate that EC2 is six times slower than a typical mid-range Linux cluster, and twenty times slower than a modern HPC system. The interconnect on the EC2 cloud platform severely limits performance and causes significant variability.
Keith R. Jackson, Lavanya Ramakrishnan, Krishna Muriki, Shane Canon, Shreyas Cholia, John Shalf, Harvey J. Wasserman, Nicholas J. Wright
CloudCom2
2010 Seeking supernovae in the clouds: a performance study
abstract
Today, our picture of the Universe radically differs from that of just over a decade ago. We now know that the Universe is not only expanding as Hubble discovered in 1929, but that the rate of expansion is accelerating, propelled by mysterious new physics dubbed "Dark Energy." This revolutionary discovery was made by comparing the brightness of nearby Type Ia supernovae (which exploded in the past billion years) to that of much more distant ones (from up to seven billion years ago). The reliability of this comparison hinges upon a very detailed understanding of the physics of the nearby events. As part of its effort to further this understanding, the Nearby Supernova Factory (SNfactory) relies upon a complex pipeline of serial processes that execute various image processing algorithms in parallel on ~10TBs of data.
Keith R. Jackson, Lavanya Ramakrishnan, Karl J. Runge, Rollin C. Thomas
HPDC2
2010 Comparison of resource platform selection approaches for scientific workflows
abstract
Cloud computing is increasingly considered as an additional computational resource platform for scientific workflows. The cloud offers opportunity to scale-out applications from desktops and local cluster resources. Each platform has different properties (e.g., queue wait times in high performance systems, virtual machine startup overhead in clouds) and characteristics (e.g., custom environments in cloud) that makes choosing from these diverse resource platforms for a workflow execution a challenge for scientists. Scientists are often faced with deciding resource platform selection trade-offs with limited information on the actual workflows. While many workflow planning methods have explored resource selection or task scheduling, these methods often require fine-scale characterization of the workflow that is onerous for a scientist. In this paper, we describe our early exploratory work in using blackbox characteristics for a cost-benefit analysis of using different resource platforms. In our blackbox method, we use only limited high-level information on the workflow length, width, and data sizes. The length and width are indicative of the workflow duration and parallelism. We compare the effectiveness of this approach to other resource selection models using two exemplar scientific workflows on desktop, local cluster, HPC center, and cloud platforms. Early results suggest that the blackbox model often makes the same resource selections as a more fine-grained whitebox model. We believe the simplicity of the blackbox model can help inform a scientist on the applicability of a new resource platform, such as cloud resources, even before porting an existing workflow.
Yogesh L. Simmhan, Lavanya Ramakrishnan
HPDC2
2009 VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance
abstract
Today's scientific workflows use distributed heterogeneous resources through diverse grid and cloud interfaces that are often hard to program. In addition, especially for time-sensitive critical applications, predictable quality of service is necessary across these distributed resources. VGrADS' virtual grid execution system (vgES) provides an uniform qualitative resource abstraction over grid and cloud systems. We apply vgES for scheduling a set of deadline sensitive weather forecasting workflows. Specifically, this paper reports on our experiences with (1) virtualized reservations for batchqueue systems, (2) coordinated usage of TeraGrid (batch queue), Amazon EC2 (cloud), our own clusters (batch queue) and Eucalyptus (cloud) resources, and (3) fault tolerance through automated task replication. The combined effect of these techniques was to enable a new workflow planning method to balance performance, reliability and cost considerations. The results point toward improved resource selection and execution management support for a variety of e-Science applications over grids and cloud systems.
Lavanya Ramakrishnan, Charles Koelbel, Yang-Suk Kee, Richard Wolski, Daniel Nurmi, Dennis Gannon, Graziano Obertelli, Asim YarKhan, Anirban Mandal, T. Mark Huang, Kiran Thyagaraja, Dmitrii Zagorodnov
SC1
2008 Performability modeling for scheduling and fault tolerance strategies for scientific workflows
abstract
Scientific applications have diverse characteristics and resource requirements. When combined with the complexity of underlying distributed resources on which they execute (e.g. Grid, cloud computing), these applications can experience significant performance fluctuations as machine reliability varies. Although the performance and reliability of cluster and Grid systems have been studied separately, there has been little analysis of the lost Quality of Service (QoS) experienced with varying availability levels. To enable a dynamic environment that can account for such changes while providing required QoS, next generation tools will need extensible application interfaces that allow users to qualitatively express performance and reliability requirements for the underlying systems. In this paper, we use the concept of performability to capture the degraded performance that might result from varying resource availability. We apply the resulting model to workflow planning and fault tolerance strategies. We present experimental data to validate our model and use simulation results driven by failure data from real HPC systems to demonstrate how the proposed scheme better accounts for resource availability.
Lavanya Ramakrishnan, Daniel A. Reed
HPDC1
2006 Poster reception - Designing a collaborative cyberinfrastructure for event-driven coastal modeling
abstract
The SURA Coastal Ocean Observing & Prediction (SCOOP) program is building cyberinfrastructure (CI) to enable advanced real-time ensemble forecasting of the coastal impacts from storms and hurricanes. This prototype of a reliable, flexible, grid-enabled forecast system integrates real-time distributed data and computer models for the coasts of the southeastern United States. The SCOOP system employs a service-oriented architecture with archive and transport services, metadata catalog, resource management, and portal interfaces. Currently, the SCOOP system uses distributed HPC machines (SCOOP, SURAgrid, others) to meet on-demand requirements. Geospatial web services disseminate the forecast results.We provide the architecture overview and describe the currently deployed system for Hurricane Season 2006 as an example in which a storm advisory automatically initiates a workflow that delivers timely forecasts. The system generates a wind-ensemble and then configures, deploys, and analyzes a variety of water level and wave models across distributed HPC resources to deliver timely forecasts.
Philip Bogden, Gabrielle Allen, Gerry Creager, Sara J. Graves, Rick A. Luettich, Lavanya Ramakrishnan
SC6
2006 Grid allocation and reservation - Toward a doctrine of containment: grid hosting with adaptive resource control
abstract
Grid computing environments need secure resource control and predictable service quality in order to be sustainable. We propose a grid hosting model in which independent, self-contained grid deployments run within isolated containers on shared resource provider sites. Sites and hosted grids interact via an underlying resource control plane to manage a dynamic binding of computational resources to containers. We present a prototype grid hosting system, in which a set of independent Globus grids share a network of cluster sites. Each grid instance runs a coordinator that leases and configures cluster resources for its grid on demand. Experiments demonstrate adaptive provisioning of cluster resources and contrast job-level and container-level resource management in the context of two grid application managers.
Lavanya Ramakrishnan, David Irwin 0001, Laura E. Grit, Aydan R. Yumerefendi, Adriana Iamnitchi, Jeffrey S. Chase
SC1