Narayan Desai

dblp:01/6327 · DBLP profile ↗
← Back
32ranked-venue papers
10as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 9 first-authorApplied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 38% Cloud and datacenter computing · 38% Interconnection networks and networks-on-chip · 12%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › job scheduling
batch scheduling
0.212016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
High-performance computing
supercomputing
0.212016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
Performance modeling and evaluation
benchmarking
0.112016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016

Methods — techniques the papers use, named apart from their topics

scheduling scheme · 0.2comparative study · 0.2benchmarking · 0.2
YearPublicationVenuePosition
2016 Open Issues in Cloud Resource Management
Narayan Desai, Walfredo Cirne
JSSPP1
2016 Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints
abstract
As systems scale toward exascale, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many-such as network bandwidth, I/O bandwidth, or power-have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potential of relaxing network allocation constraints for Blue Gene systems. Our objective is to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is three new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing scheduler on Mira, a 48-rack Blue Gene/Q system at Argonne National Laboratory. Specifically, we use job traces collected from this production system.
Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai
IEEE Trans. Parallel Distributed Syst.7
2015 Improving Batch Scheduling on Blue Gene/Q by Relaxing 5D Torus Network Allocation Constraints
abstract
As systems scale toward exactable, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many -- such as network bandwidth, I/O bandwidth, or power -- have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potentiality of relaxing network allocation constraints for Blue Gene systems. Our objectives to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is two new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing one under different workloads, using job traces collected from the 48-rack Mira, an IBM Blue Gene/Q system at Argonne National Laboratory.
Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai
IPDPS7
2015 A RESTful API for Accessing Microbial Community Data for MG-RAST
abstract
Metagenomic sequencing has produced significant amounts of data in recent years. For example, as of summer 2013, MG-RAST has been used to annotate over 110,000 data sets totaling over 43 Terabases. With metagenomic sequencing finding even wider adoption in the scientific community, the existing web-based analysis tools and infrastructure in MG-RAST provide limited capability for data retrieval and analysis, such as comparative analysis between multiple data sets. Moreover, although the system provides many analysis tools, it is not comprehensive. By opening MG-RAST up via a web services API (application programmers interface) we have greatly expanded access to MG-RAST data, as well as provided a mechanism for the use of third-party analysis tools with MG-RAST data. This RESTful API makes all data and data objects created by the MG-RAST pipeline accessible as JSON objects. As part of the DOE Systems Biology Knowledgebase project (KBase, http://kbase.us) we have implemented a web services API for MG-RAST. This API complements the existing MG-RAST web interface and constitutes the basis of KBase's microbial community capabilities. In addition, the API exposes a comprehensive collection of data to programmers. This API, which uses a RESTful (Representational State Transfer) implementation, is compatible with most programming environments and should be easy to use for end users and third parties. It provides comprehensive access to sequence data, quality control results, annotations, and many other data types. Where feasible, we have used standards to expose data and metadata. Code examples are provided in a number of languages both to show the versatility of the API and to provide a starting point for users. We present an API that exposes the data in MG-RAST for consumption by our users, greatly enhancing the utility of the MG-RAST service.
Andreas Wilke, Jared Bischof, Travis Harrison, Thomas S. Brettin, Mark D'Souza, Wolfgang Gerlach, Hunter Matthews, Tobias Paczian, Jared Wilkening, Elizabeth M. Glass, Narayan Desai, Folker Meyer
PLoS Comput. Biol.11
2014 Workload characterization for MG-RAST metagenomic data analytics service in the cloud
abstract
The cost of DNA sequencing has plummeted in recent years. The consequent data deluge has imposed big burdens for data analysis applications. For example, MG-RAST, a production open-public metagenome annotation service, has experienced increasingly large amount of data submission and has demanded scalable resources for the computational needs. To address this problem, we have developed a scalable platform to port MG-RAST workloads into the cloud, where elastic computing resources can be used on demand. To efficiently utilize such resources, however, one must understand the characteristics of the application workloads. In this paper, we characterize the MG-RAST workloads running in the cloud, from the perspectives of computation, I/O, and data transfer. Insights from this work will help guide application enhancement, service operation, and resource management for MG-RAST and similar big data applications demanding elastic computing resources.
Wei Tang 0001, Jared Bischof, Narayan Desai, Kanak Mahadik, Wolfgang Gerlach, Travis Harrison, Andreas Wilke, Folker Meyer
IEEE BigData3
2014 Data-Aware Resource Scheduling for Multicloud Workflows: A Fine-Grained Simulation Approach
abstract
Cloud infrastructures have seen increasing popularity for addressing the growing computational needs of today's scientific and engineering applications. However, resource management challenges exist in the elastic cloud environment, such as resource provisioning and task allocation, especially when data movement between multiple domains plays an important role. In this work, we study the impact of data-aware resource management and scheduling on scientific workflows in multicloud environments. We develop a workflow simulator based on a network simulation framework for fine-grained simulation for workflow computation and data movement. Using the workload traces from a production metagenomic data analysis service, we evaluate different resource scheduling mechanisms, including proposed data-aware scheduling policies under various resource and bandwidth configurations. The results of this work are expected to answer questions about how to provision computing resources for certain workloads efficiently and how to place tasks across multidomain clouds in order to reduce data movement costs for overall improved system performance.
Wei Tang 0001, Jonathan Jenkins, Folker Meyer, Robert B. Ross, Rajkumar Kettimuthu, Linda Winkler, Xi Yang 0002, Thomas Lehman, Narayan Desai
CloudCom9
2013 A scalable data analysis platform for metagenomics
abstract
With the advent of high-throughput DNA sequencing technology, the analysis and management of the increasing amount of biological sequence data has become a bottleneck for scientific progress. For example, MG-RAST, a metagenome annotation system serving a large scientific community worldwide, has experienced a sustained, exponential growth in data submissions for several years; and this trend is expected to continue. To address the computational challenges posed by this workload, we developed a new data analysis platform, including a data management system (Shock) for biological sequence data and a workflow management system (AWE) supporting scalable, fault-tolerant task and resource management. Shock and AWE can be used to build a scalable and reproducible data analysis infrastructure for upper-level biological data analysis services.
Wei Tang 0001, Jared Wilkening, Narayan Desai, Wolfgang Gerlach, Andreas Wilke, Folker Meyer
IEEE BigData3
2013 Reducing Energy Costs for IBM Blue Gene/P via Power-Aware Job Scheduling
Zhou Zhou 0006, Zhiling Lan, Wei Tang 0001, Narayan Desai
JSSPP4
2013 Poncho: Enabling Smart Administration of Full Private Clouds
Scott Devoid, Narayan Desai, Lorin Hochstein
LISA2
2013 Job scheduling with adjusted runtime estimates on production supercomputers
Wei Tang 0001, Narayan Desai, Daniel Buettner, Zhiling Lan
J. Parallel Distributed Comput.2
2013 Toward balanced and sustainable job scheduling for production supercomputers
Wei Tang 0001, Dongxu Ren, Zhiling Lan, Narayan Desai
Parallel Comput.4
2013 Multi-domain job coscheduling for leadership computing systems
Wei Tang 0001, Narayan Desai, Venkatram Vishwanath, Daniel Buettner, Zhiling Lan
J. Supercomput.2
2011 Evaluating Performance Impacts of Delayed Failure Repairing on Large-Scale Systems
abstract
With the fast improvement in technology, we are now moving toward exascale computing. Many experts predict that exascale computers will have millions of nodes, billions of threads of execution, hundreds of petabytes of inner memory and exabytes of persistent storage. For systems of such a scale, frequent failures are becoming a serious concern. One of the most important reasons is that in a large-scale system it is hard to detect failures. As a result, failure repair may take substantial time. In this paper, we investigate the effect of delayed repairing on two popular types of high-performance computing systems: IBM Blue Gene/P and general cluster. We analyze how delayed failure repairing will affect the performance of jobs when some computing units are at fault but not fixed in time. Our study is based on real workload traces and RAS logs collected from production supercomputing systems. Our Trace-based simulations indicate that fast failure detection and recovery is essential for moving towards petascale and beyond computing.
Zhou Zhou 0006, Wei Tang 0001, Ziming Zheng, Zhiling Lan, Narayan Desai
CLUSTER5
2011 Reducing Fragmentation on Torus-Connected Supercomputers
abstract
Torus-based networks are prevalent on leadership-class petascale systems, providing a good balance between network cost and performance. The major disadvantage of this network architecture is its susceptibility to fragmentation. Many studies have attempted to reduce resource fragmentation in this architecture. Although the approaches suggested can make good allocation decisions reducing fragmentation at job start time, none of them considers a job's wall time, which can cause resource fragmentation when neighboring jobs do not complete closely. In this paper, we propose a wall time-aware job allocation strategy, which adjacently packs jobs that finish around the same time, in order to minimize resource fragmentation caused by job length, discrepancy. Event-driven simulations using real job traces from a production Blue Gene/P system at Argonne National Laboratory demonstrate that our wall time-aware strategy can effectively reduce system fragmentation and improve overall system performance.
Wei Tang 0001, Zhiling Lan, Narayan Desai, Daniel Buettner, Yongen Yu
IPDPS3
2011 Co-analysis of RAS Log and Job Log on Blue Gene/P
abstract
With the growth of system size and complexity, reliability has become of paramount importance for petascale systems. Reliability, Availability, and Serviceability (RAS) logs have been commonly used for failure analysis. However, analysis based on just the RAS logs has proved to be insufficient in understanding failures and system behaviors. To overcome the limitation of this existing methodologies, we analyze the Blue Gene/P RAS logs and the Blue Gene/P job logs in a cooperative manner. From our co-analysis effort, we have identified a dozen important observations about failure characteristics and job interruption characteristics on the Blue Gene/P systems. These observations can significantly facilitate the research in fault resilience of large-scale systems.
Ziming Zheng, Li Yu 0006, Wei Tang 0001, Zhiling Lan, Rinku Gupta, Narayan Desai, Susan Coghlan, Daniel Buettner
IPDPS6
2011 An experience report: porting the MG-RAST rapid metagenomics analysis pipeline to the cloud
abstract
SUMMARY Existing applications in computational biology typically favor a local cluster based integrated computational platform. We present a lessons learned type report for scaling up an existing metagenomics application that outgrew the available local cluster hardware. In our example, removing a number of assumptions linked to tight integration allowed to expand beyond one administrative domain, increase the number and type of machines available for the application, and also improved scaling properties of the application. The assumptions made in designing the computational client make it well suitable for deployment as a virtual machine inside a cloud. This paper discusses the decision process and describes the suitability of deploying various bioinformatics computations to distributed heterogeneous machines. Copyright © 2011 John Wiley & Sons, Ltd.
Andreas Wilke, Jared Wilkening, Elizabeth M. Glass, Narayan Desai, Folker Meyer
Concurr. Comput. Pract. Exp.4
2010 Analyzing and adjusting user runtime estimates to improve job scheduling on the Blue Gene/P
abstract
Backfilling and short-job-first are widely acknowledged enhancements to the simple but popular first-come, first-served job scheduling policy. However, both enhancements depend on user-provided estimates of job runtime, which research has repeatedly shown to be inaccurate. We have investigated the effects of this inaccuracy on backfilling and different queue prioritization policies, determining which part of the scheduling policy is most sensitive. Using these results, we have designed and implemented several estimation-adjusting schemes based on historical data. We have evaluated these schemes using workload traces from the Blue Gene/P system at Argonne National Laboratory. Our experimental results demonstrate that dynamically adjusting job runtime estimates can improve job scheduling performance by up to 20%.
Wei Tang 0001, Narayan Desai, Daniel Buettner, Zhiling Lan
IPDPS2
2009 Fault-aware, utility-based job scheduling on Blue, Gene/P systems
abstract
Job scheduling on large-scale systems is an increasingly complicated affair, with numerous factors influencing scheduling policy. Addressing these concerns results in sophisticated scheduling policies that can be difficult to reason about. In this paper, we present a general utility-based scheduling framework to balance various scheduling requirements and priorities. It enables system owners to customize scheduling policies under different circumstances without changing the scheduling code. We also develop a fault-aware job allocation strategy for Blue Gene/P systems to address the increasing concern of system failures. We demonstrate the effectiveness of these facilities by means of event-driven simulations with real job traces collected from the production Blue Gene/P system at Argonne National Laboratory.
Wei Tang 0001, Zhiling Lan, Narayan Desai, Daniel Buettner
CLUSTER3
2009 Using clouds for metagenomics: A case study
abstract
Cutting-edge sequencing systems produce data at a prodigious rate; and the analysis of these datasets requires significant computing resources. Cloud computing provides a tantalizing possibility for on-demand access to computing resources. However, many open questions remain. We present here a performance assessment of BLAST on real metagenomics data in a cloud setting in order to determine the viability of this approach. BLAST is one of the premier applications in bioinformatics and computational biology and is assumed to consume the vast majority of resources in that area.
Jared Wilkening, Andreas Wilke, Narayan Desai, Folker Meyer
CLUSTER3
2009 Understanding Network Saturation Behavior on Large-Scale Blue Gene/P Systems
abstract
As researchers continue to architect massive-scale systems, it is becoming clear that these systems will utilize a significant amount of shared hardware between processing units. Systems such as the IBM Blue Gene (BG) and Cray XT have started utilizing flat (i.e., scalable) networks, which differ from switched fabrics in that they use a 3D torus or similar topology. This allows the network to grow only linearly with system scale, instead of the super linear growth needed for full fat-tree switched topologies, but at the cost of increased network sharing between processing nodes. While in many cases a full fat-tree is an over estimate of the needed bisectional bandwidth, it is not clear whether the other extreme of a flat topology is sufficient to move data around the network efficiently. In this paper, we study the network behavior of the IBM BG/P using several application communication kernels, and we monitor network congestion behavior based on detailed hardware counters. Our studies scale from small systems to 8 racks (32,768 cores) of BG/P and provide insights into the network communication characteristics of the system.
Pavan Balaji, Harish Naik, Narayan Desai
ICPADS3
2009 Improving Resource Availability by Relaxing Network Allocation Constraints on Blue Gene/P
abstract
High-end computing (HEC) systems have passed the petaflop barrier and continue to move toward the next frontier of {exascale} computing. As companies and research institutes continue to work toward architecting these enormous systems, it is becoming increasingly clear that these systems will utilize a significant amount of shared hardware between processing units, including shared caches, memory management engines, and network infrastructure. While these systems are optimized to use all of the hardware available in a dedicated manner to achieve the best performance, in practice, the shared nature of this hardware makes scheduling applications on it difficult and wasteful. For example, while the IBM Blue Gene/P system has been designed to use a torus network for efficient communication, some of the torus links (especially those connecting different racks) are shared between multiple racks. Thus, a job running on one rack, might preclude another job from running on a second rack in spite of having its compute resources completely idle. In this paper, we assess the relative performance degradation noticed by real applications when such shared network hardware is completely unutilized for some cases. Our measurements on Intrepid, one of the largest Blue Gene/P installations in the world, demonstrate less than 5% degradation for several leadership applications commonly run on the Intrepid system. Further, we demonstrate that the additional scheduling flexibility offered by not sharing such hardware can improve the overall job turnaround time by nearly 40% in some cases.
Narayan Desai, Darius Buntinas, Daniel Buettner, Pavan Balaji
ICPP1
2008 Are nonblocking networks really needed for high-end-computing workloads?
abstract
High-speed interconnects are frequently used to provide scalable communication on increasingly large high-end computing systems. Often, these networks are nonblocking, where there exist independent paths between all pairs of nodes in the system allowing for simultaneous communication with zero network contention. This performance, however, comes at a heavy cost as the number of components needed (and hence cost) increases superlinearly with the number of nodes in the system. In this paper, we study the behavior of real and synthetic supercomputer workloads to understand the impact of the networkpsilas nonblocking capability on overall performance. Starting from a fully nonblocking network, we begin by assessing the worse-case performance degradation caused by removing interstage communication links, resulting in over provisioning and hence potentially blocking in the communication network.We also study the impact of several factors on this behavior, including system workloads, multicore processors, and switch crossbar sizes. Our observations show that a significant reduction in the number of interstage links can be tolerated on all of the workloads analyzed, causing less than 5% overall loss of performance.
Narayan Desai, Pavan Balaji, P. Sadayappan
CLUSTER1
2008 Petascale System Management Experiences
Narayan Desai, Rick Bradshaw, Cory Lueninghoener, Andrew Cherry, Susan Coghlan, William Scullin
LISA1
2007 The computer as software component: A mechanism for developing and testing resource management software
abstract
In this paper, we present an architecture that encapsulates system hardware inside a software component used for job execution and status monitoring. The development of this interface has enabled system simulation, which yields a number of novel benefits, including dramatically improved debug and testing capabilities.
Narayan Desai, Theron Voran, Ewing L. Lusk, Andrew Cherry
CLUSTER1
2006 bcfg2
Narayan Desai
LISA1
2006 Directing Change Using Bcfg2
Narayan Desai, Rick Bradshaw, Cory Lueninghoener
LISA1
2005 A Case Study in Configuration Management Tool Deployment
Narayan Desai, Rick Bradshaw, Scott Matott, Sandra Bittner, Susan Coghlan, Rémy Evard, Cory Lueninghoener, Ti Leggett, John-Paul Navarro, Gene Rackow, Craig Stacey, Tisha Stacey
LISA1
2004 Component-based cluster systems software architecture a case study
abstract
We describe the use of component architecture in an area to which this approach has not been classically applied, the area of cluster system software. By "cluster system software," we mean the collection of programs used in configuring and maintaining individual nodes, together with the software involved in submission, scheduling, monitoring, and termination of jobs. We describe how the component approach maps onto the cluster systems software problem, together with our experiences with the approach in implementing an all-new suite of systems software for a medium-sized cluster with unusually complex systems software requirements.
Narayan Desai, Rick Bradshaw, Ewing L. Lusk, Ralf Butler
CLUSTER1
2003 The ProcessManagement Component of a Scalable Systems Software Environment
abstract
The systems software necessary to operate large-scale parallel computers presents a variety of research and development issues. One approach is to consider systems software as a collection of interacting components, with well-defined published interfaces. The scalable systems software SciDAC project is currently exploring the feasibility of architecting systems software this way. In this paper we present a prototype process manager component for such a system. We describe the component abstractly in terms of its functionality and the interface by which its functionality may be invoked. We propose a precise syntax for this interface and describe one implementation of the process manager component, based on an existing scalable process management system called MPD. We conclude with some experiences using this process manager component in conjunction with other systems software components on a medium-sized Linux cluster.
Ralph M. Butler, Narayan Desai, Andrew Lusk, Ewing L. Lusk
CLUSTER2
2003 BCFG: A Con.guration Management Tool for Heterogeneous Environments
abstract
Since clusters were first introduced, node counts have increased rapidly. Currently, a variety of clusters with more than one thousand nodes are listed on the TOP500 list. In the next three years, clusters with more than four thousand nodes are expected. Cluster management functionality has lagged behind all areas of system software. In order to effectively manage the clusters of today and tomorrow, the basic cluster management software model must change. Current techniques focus on the management of single nodes, as opposed to complete cluster configurations. This approach typically leads to automatic management of compute nodes, while using ad-hoc techniques to manage service nodes. Configuration management is the process where software configurations on clients are installed, updated and verified. To address these issues, we have begun the development of BCFG, a symbolic configuration management tools for heterogeneous clusters. It uses a multi-tiered configuration description.
Narayan Desai, Andrew Lusk, Rick Bradshaw, Rémy Evard
CLUSTER1
2002 Clusters as Large-Scale
abstract
In this paper, we describe the use of a cluster as a generalized facility for development. A development facility is a system used primarily for testing and development activities while being operated reliably for many users. We are in the midst of a project to build and operate a large-scale development facility. We discuss our motivation for using clusters in this way and compare the model with a classic computing facility. We describe our experiences and findings from the first phase of this project. Many of these observations are relevant to the design of standard clusters and to future development facilities.
Rémy Evard, Narayan Desai, John-Paul Navarro, Daniel Nurmi
CLUSTER2
2002 Scalable Cluster Administration - Chiba City I Approach and Lessons Learned
abstract
Systems administrators of large clusters often need to perform the same administrative task hundreds or thousands of times. Administrators have traditionally performed some time-consuming tasks, such as operating system installation, configuration, and maintenance, manually. By combining network services such as DHCP, TFTP, FTP, HTTP, and NFS with remote hardware control and scripted installation, configuration, and maintenance techniques, cluster administrators can automate these administrative tasks. Scalable cluster administration addresses this challenge: What hardware and software design techniques can cluster builders use to automate cluster administration on very large clusters? We describe the approach used in the Mathematics and Computer Science Division of Argonne National Laboratory on Chiba City I, a 314-node Linux cluster; and we analyze the scalability, flexibility, performance and reliability benefits and limitations from that approach.
John-Paul Navarro, Rémy Evard, Daniel Nurmi, Narayan Desai
CLUSTER4