Zhou Zhou 0006

dblp:67/2535-6 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
0since 2021 · last 2018
0000-0002-8637-3654ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 40% High-performance computing · 24% Energy-efficient computing · 21%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › job scheduling
batch scheduling
0.212016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
High-performance computing
supercomputing
0.212016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
Energy-efficient computing
energy-aware scheduling
0.212013
Integrating dynamic pricing of electricity into energy aware scheduling for HPC systems · SC 2013
Cloud and datacenter computing
job scheduling
0.212013
Integrating dynamic pricing of electricity into energy aware scheduling for HPC systems · SC 2013
Performance modeling and evaluation
benchmarking
0.112016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
Energy-efficient computing
power management
0.012013
Integrating dynamic pricing of electricity into energy aware scheduling for HPC systems · SC 2013

Methods — techniques the papers use, named apart from their topics

scheduling scheme · 0.2comparative study · 0.2benchmarking · 0.2trace-based simulation · 0.2
YearPublicationVenuePosition
2018 System-wide trade-off modeling of performance, power, and resilience on petascale systems
Li Yu 0006, Zhou Zhou 0006, Yuping Fan, Michael E. Papka, Zhiling Lan
J. Supercomput.2
2017 Toward Efficient and Flexible Metadata Indexing of Big Data Systems
abstract
In Big Data era, applications are generating orders of magnitude more data in both volume and quantity. While many systems emerge to address such data explosion, the fact that these data's descriptors, i.e., metadata, are also “big” is often overlooked. The conventional approach to address the big metadata issue is to disperse metadata into multiple machines. However, it is extremely difficult to preserve both load-balance and data-locality in this approach. To this end, in this work we propose hierarchical indirection layers for indexing the underlying distributed metadata. By doing this, data locality is achieved efficiently by the indirection while load-balance is preserved. Three key challenges exist in this approach, however: first, how to achieve high resilience; second, how to ensure flexible granularity; third, how to restrain performance overhead. To address above challenges, we design Dindex, a distributed indexing service for metadata. Dindex incorporates a hierarchy of coarse-grained aggregation and horizontal key-coalition. Theoretical analysis shows that the overhead of building Dindex is compensated by only two or three queries. Dindex has been implemented by a lightweight distributed key-value store and integrated to a fully-fledged distributed filesystem. Experiments demonstrated that Dindex accelerated metadata queries by up to 60 percent with a negligible overhead.
Dongfang Zhao 0001, Kan Qiao, Zhou Zhou 0006, Tonglin Li, Zhihan Lyu, Xiaohua Xu 0002
IEEE Trans. Big Data3
2016 Exploring Plan-Based Scheduling for Large-Scale Computing Systems
abstract
As HPC systems scale toward exascale, it becomes critical to manage the underlying resource more effectively. While almost all existing resource management systems schedule jobs in a queuing fashion and have drawbacks of making isolated scheduling decisions that would compromise system performance even with backfilling, plan-based schedulers have the potential to generate better job schedules by producing an execution plan of all waiting jobs but do not receive enough attention. In this paper, we present a novel plan-based scheduling system that utilizes simulated annealing as the optimization engine to support effective resource management on HPC systems. As demonstrated by extensive trace-based simulations with workload traces collected from a wide range of production supercomputers, in comparison with the queue-based scheduling system using FCFS with EASY backfilling, our plan-based scheduling system can reduce the job wait time by 40%, reduce the job response time by 30%, while slightly improving system utilization at the same time. Moreover, our plan-based system is able to run online by solving the scheduling problem at each scheduling iteration within one second, making it practical for production HPC systems.
Xingwu Zhang, Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan
CLUSTER2
2016 Exploiting multi-cores for efficient interchange of large messages in distributed systems
abstract
Summary Conventional data serialization tools assume that objects to be coded are usually small in size so a single CPU core can encode it in a timely manner. In the era of Big Data, however, object gets increasingly complex and larger, which makes data serialization become a new performance bottleneck. This paper describes an approach to parallelize data serialization by leveraging multiple cores. Parallelizing data serialization introduces new questions such as how to split the (sub)objects, how to allocate the available cores, and how to minimize its overhead in practice. In this paper we design a framework for parallelly serializing large objects and analyze the design tradeoffs under different scenarios. To validate the proposed approach, we implemented parallel protocol buffers—the parallel version of Google's Protocol Buffers, a widely‐used data serialization utility. Experimental results confirm the effectiveness of Parallel Protocol Buffers: multiple cores employed in data serialization achieve highly scalable performance and incur negligible overhead. Copyright © 2015 John Wiley & Sons, Ltd.
Dongfang Zhao 0001, Kan Qiao, Zhou Zhou 0006, Tonglin Li, Xiaobing Zhou, Ioan Raicu
Concurr. Comput. Pract. Exp.3
2016 Application power profiling on IBM Blue Gene/Q
Sean Wallace, Zhou Zhou 0006, Venkatram Vishwanath, Susan Coghlan, John R. Tramm, Zhiling Lan, Michael E. Papka
Parallel Comput.2
2016 I/O-aware bandwidth allocation for petascale computing systems
Zhou Zhou 0006, Xu Yang 0009, Dongfang Zhao 0001, Paul M. Rich, Wei Tang 0001, Zhiling Lan
Parallel Comput.1
2016 Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints
abstract
As systems scale toward exascale, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many-such as network bandwidth, I/O bandwidth, or power-have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potential of relaxing network allocation constraints for Blue Gene systems. Our objective is to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is three new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing scheduler on Mira, a 48-rack Blue Gene/Q system at Argonne National Laboratory. Specifically, we use job traces collected from this production system.
Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai
IEEE Trans. Parallel Distributed Syst.1
2015 I/O-Aware Batch Scheduling for Petascale Computing Systems
abstract
In the Big Data era, the gap between the storage performance and an application's I/O requirement is increasing. I/O congestion caused by concurrent storage accesses from multiple applications is inevitable and severely harms the performance. Conventional approaches either focus on optimizing an application's access pattern individually or handle I/O requests on a low-level storage layer without any knowledge from the upper-level applications. In this paper, we present a novel I/O-aware batch scheduling framework to coordinate ongoing I/O requests on petascale computing systems. The motivation behind this innovation is that the batch scheduler has a holistic view of both the system state and jobs' activities and can control the jobs' status on the fly during their execution. We treat a job's I/O requests as periodical subjobs within its lifecycle and transform the I/O congestion issue into a classical scheduling problem. We design two scheduling polices with different scheduling objectives either on user-oriented metrics or system performance. We conduct extensive trace-based simulations using real job traces and I/O traces from a production IBM Blue Gene/Q system. Experimental results demonstrate that our design can improve job performance by more than 30%, as well as increasing system performance.
Zhou Zhou 0006, Xu Yang 0009, Dongfang Zhao 0001, Paul M. Rich, Wei Tang 0001, Zhiling Lan
CLUSTER1
2015 Improving Batch Scheduling on Blue Gene/Q by Relaxing 5D Torus Network Allocation Constraints
abstract
As systems scale toward exactable, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many -- such as network bandwidth, I/O bandwidth, or power -- have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potentiality of relaxing network allocation constraints for Blue Gene systems. Our objectives to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is two new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing one under different workloads, using job traces collected from the 48-rack Mira, an IBM Blue Gene/Q system at Argonne National Laboratory.
Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai
IPDPS1
2015 Quantitative modeling of power performance tradeoffs on extreme scale systems
Li Yu 0006, Zhou Zhou 0006, Sean Wallace, Michael E. Papka, Zhiling Lan
J. Parallel Distributed Comput.2
2014 Balancing job performance with system performance via locality-aware scheduling on torus-connected systems
abstract
Torus-connected network is widely used in modern supercomputers due to its linear per node cost scaling and its competitive overall performance. Job scheduling system plays a critical role for the efficient use of supercomputers. As supercomputers continue growing in size, a fundamental problem arises: how to effectively balance job performance with system performance on torus-connected machines? In this work, we will present a new scheduling design named window-based locality-aware scheduling. Our design contains three novel features. First, rather than one-by-one job scheduling, our design takes a “window” of jobs, i.e. multiple jobs, into consideration for job prioritizing and resource allocation. Second, our design maintains a list of slots to preserve node contiguity information for resource allocation. Finally, we formulate our scheduling decision making into a 0-1 Multiple Knapsack Problem and present two algorithms to solve the problem. A series of trace-based simulations using job logs collected from production supercomputers indicate that this new scheduling design has real potentials and can effectively balance job performance and system performance.
Xu Yang 0009, Zhou Zhou 0006, Wei Tang 0001, Xingwu Zheng, Zhiling Lan
CLUSTER2
2013 Reducing Energy Costs for IBM Blue Gene/P via Power-Aware Job Scheduling
Zhou Zhou 0006, Zhiling Lan, Wei Tang 0001, Narayan Desai
JSSPP1
2013 Integrating dynamic pricing of electricity into energy aware scheduling for HPC systems
abstract
The research literature to date mainly aimed at reducing energy consumption in HPC environments. In this paper we propose a job power aware scheduling mechanism to reduce HPC's electricity bill without degrading the system utilization. The novelty of our job scheduling mechanism is its ability to take the variation of electricity price into consideration as a means to make better decisions of the timing of scheduling jobs with diverse power profiles. We verified the effectiveness of our design by conducting trace-based experiments on an IBM Blue Gene/P and a cluster system as well as a case study on Argonne's 48-rack IBM Blue Gene/Q system. Our preliminary results show that our power aware algorithm can reduce electricity bill of HPC systems as much as 23%.
Xu Yang 0009, Zhou Zhou 0006, Sean Wallace, Zhiling Lan, Wei Tang 0001, Susan Coghlan, Michael E. Papka
SC2
2011 Evaluating Performance Impacts of Delayed Failure Repairing on Large-Scale Systems
abstract
With the fast improvement in technology, we are now moving toward exascale computing. Many experts predict that exascale computers will have millions of nodes, billions of threads of execution, hundreds of petabytes of inner memory and exabytes of persistent storage. For systems of such a scale, frequent failures are becoming a serious concern. One of the most important reasons is that in a large-scale system it is hard to detect failures. As a result, failure repair may take substantial time. In this paper, we investigate the effect of delayed repairing on two popular types of high-performance computing systems: IBM Blue Gene/P and general cluster. We analyze how delayed failure repairing will affect the performance of jobs when some computing units are at fault but not fixed in time. Our study is based on real workload traces and RAS logs collected from production supercomputing systems. Our Trace-based simulations indicate that fast failure detection and recovery is essential for moving towards petascale and beyond computing.
Zhou Zhou 0006, Wei Tang 0001, Ziming Zheng, Zhiling Lan, Narayan Desai
CLUSTER1