EDBT 2026 Demo / reviewers in the wild / expert
Zhiling Lan
dblp:07/2007
· DBLP profile ↗
94ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0002-1047-8724ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 80 · 10 first-author · 13 since 2021Security and privacy · 3Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly SystemsabstractThe Dragonfly network, with its high-radix and low-diameter structure, is a leading interconnect in high-performance computing. A major challenge is workload interference on shared network links. Parallel discrete event simulation (PDES) is commonly used to analyze workload interference. However, high-fidelity PDES is computationally expensive, making it impractical for large-scale or real-time scenarios. Hybrid simulation that incorporates data-driven surrogate models offers a promising alternative, especially for forecasting application runtime, a task complicated by the dynamic behavior of network traffic. We present SMART, a surrogate model that combines graph neural networks (GNNs) and large language models (LLMs) to capture both spatial and temporal patterns from port level router data. SMART outperforms existing statistical and machine learning baselines, enabling accurate runtime prediction and supporting efficient hybrid simulation of Dragonfly networks. Xin Wang 0115, Pietro Lodi Rizzini, Sourav Medya, Zhiling Lan |
AAAI | 4 |
| 2026 | SmartCap: Coordinated CPU-GPU Power Capping for Performance-Assurance Energy EfficiencyabstractPerformance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling methods often rely on exhaustive offline profiling, which becomes impractical due to the large setting space and the high cost of profiling large-scale applications. In this paper, we present OPEN, a framework consists of offline and online phases. The offline phase involves building a performance predictor and constructing an initial dense matrix. In the online phase, OPEN performs lightweight online profiling, and leverages the performance predictor with collaborative filtering to make performance prediction. We evaluate OPEN on multiple heterogeneous systems, including those equipped with A100 and A30 GPUs. Results show that OPEN achieves prediction accuracy up to 98.29\%. This demonstrates that OPEN effectively reduces profiling cost while maintaining high accuracy, making it practical for power-aware performance modeling in modern HPC environments. Overall, OPEN provides a lightweight solution for performance prediction under power constraints, enabling better runtime decisions in power-aware computing environments. Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan |
ICS | 5 |
| 2026 | Beyond Throughput: Performance and Energy Insights of LLM Inference Across AI Accelerators
Giacomo Brunetta, Varuni Sastry 0001, Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan |
IPDPS | 6 |
| 2025 | Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf BenchmarksabstractThe significantly increasing use of artificial intelligence (AI) has led to the availability of specialized AI accelerators, aiming to enhance the performance and energy efficiency of AI workloads. In this paper, we conduct an initial study to evaluate the energy requirements of four AI accelerators: Nvidia A100 GPUs, Intel Habana Gaudi Processing Units (HPUs), Graphcore Bow-Pod64 Intelligence Processing Units (IPUs), and GroqRack Language Processing Units (LPUs) using two popular MLPerf benchmarks: BERT-Large and ResNet50. We report the energy requirements for the two benchmarks to achieve a common MLPerfspecified target accuracy. The benchmarks and AI accelerators were chosen based on the following criteria: publicly available tools or libraries from the vendors to monitor power consumption, publicly available optimized models from the vendors, and access to the AI accelerators. Our experimental results indicate that for ResNet50, Intel Gaudi2 HPUs delivered the highest throughput, the lowest energy consumption for both training and inference, and the highest inference energy efficiency, while Graphcore demonstrated the highest training energy efficiency. For BERTLarge pre-training, Intel Gaudi2 outperformed both Nvidia A100 and Graphcore in terms of time, energy consumption, and training energy efficiency. However, for BERT-Large inference, Nvidia A100 achieved the shortest time and lowest energy consumption; Graphcore exhibited the highest throughput in both pre-training and inference, along with the highest inference energy efficiency. We discuss our observations and findings while exploring the associated tradeoffs. Farah Ferdaus, Xingfu Wu, Valerie Taylor 0001, Zhiling Lan, Sanjif Shanmugavelu, Venkatram Vishwanath, Michael E. Papka |
CCGrid | 4 |
| 2025 | More for Less: Integrating Capability-Predominant and Capacity-Predominant Computing
Michael E. Papka, Zhiling Lan |
JSSPP | 3 |
| 2025 | Directing PDES and Surrogate Models in Loosely Coupled Hybrid Simulations
Kevin A. Brown, Elkin Cruz-Camacho, Kazutomo Yoshii, Xin Wang 0115, Zhiling Lan, Christopher D. Carothers, Robert B. Ross |
SIGSIM-PADS | 5 |
| 2025 | CQSim+: Symbiotic Simulation for Multi-Resource Scheduling in High-Performance Computing
Yash Kurkure, Shambhawi Sharma, Xin Wang 0115, Michael E. Papka, Zhiling Lan |
SIGSIM-PADS | 5 |
| 2025 | Synopsis: MFNetSim: A Multi-Fidelity Network Simulation Framework for Multi-Traffic Modeling of Dragonfly Systems
Xin Wang 0115, Kevin A. Brown, Robert B. Ross, Christopher D. Carothers, Zhiling Lan |
SIGSIM-PADS | 5 |
| 2025 | On the Effectiveness of Unified Memory in Multi-GPU Collective CommunicationabstractModern supercomputers are becoming increasingly dense with accelerators. Industry leaders offer multi-GPU architectures with high interconnection bandwidth between the devices to match the requirements of modern workloads. While those technologies advance, it is up to the programmer to successfully exploit them. Recognizing this burden, multiple abstractions have been built. We focus on the NVIDIA Collective Communication Library (NCCL) and Unified Memory (UM). The former provides MPI-like directives integrated within the GPU runtime, allowing lower latencies and increasing the bandwidth over previous approaches. The latter simplifies the programming paradigm, offering a unified virtual address space. Moreover, it enables memory oversubscription, drastically reducing the efforts towards handling larger problems without completely restructuring the codebase. This work provides the first joint analysis of NCCL and UM from single-node multi-GPU architectures to a production supercomputer. We explore all the available collective communication directives concerning their power requirements and overall throughput. Moreover, we study the effects of various hyperparameters, e.g., message sizes, oversubscription level, and memory advice, on the overall obtainable performance. Our findings showcase how using UM brings negligible increased energy consumption; moreover, in distributed settings, other restricting factors, such as network bottlenecks, surpass the overhead introduced by UM’s page-eviction mechanisms. Riccardo Strina, Ian Di Dio Lavore, Marco D. Santambrogio, Michael E. Papka, Zhiling Lan |
PDP | 5 |
| 2025 | Minimizing Power Waste in Heterogenous Computing via Adaptive Uncore ScalingabstractHigh-performance computing (HPC) systems are essential for scientific discovery and engineering innovation. However, their growing power demands pose significant challenges, particularly as systems scale to the exascale level. Prior uncore frequency tuning studies have primarily focused on conventional HPC workloads running on CPU-only systems. As HPC advances toward heterogeneous computing, integrating diverse GPU workloads on heterogeneous CPU-GPU systems, it becomes imperative to revisit and enhance uncore scaling. Our investigation reveals that uncore frequency scales down only when CPU power approaches its thermal design power (TDP), which is rare in GPU-dominant applications. As a result, modern computing systems experience unnecessary power waste. In this study, we present MAGUS, a user-transparent uncore frequency scaling runtime for heterogeneous computing. MAGUS dynamically adjusts uncore frequencies according to distinct application execution phases, effectively minimizing power waste caused by consistently using maximum uncore frequencies. Our design incorporates several key techniques, including real-time monitoring and prediction of memory accesses, intelligent handling of frequent phase transitions, and leveraging vendor-provided power management features. We evaluate MAGUS with various GPU benchmarks and applications on multiple heterogeneous systems with different CPU and GPU architectures. Experimental results demonstrate that MAGUS achieves up to 27% energy savings compared to the default settings, while maintaining a performance loss of less than 5% and an overhead of under 1%. Seyfal Sultanov, Michael E. Papka, Zhiling Lan |
SC | 4 |
| 2024 | Surrogate Modeling for HPC Application Iteration Times Forecasting with Network FeaturesabstractInterconnect networks are the foundation for modern high performance computing (HPC) systems. Parallel discrete event simulation (PDES), serving as a cornerstone in the study of large-scale networking systems by modeling and simulating the real-world behaviors of HPC facilities, faces escalating computational complexities at an unsustainable scale. The research community is interested in building a surrogate-ready PDES framework where an accurate surrogate model can be used to forecast HPC behaviors and replace computationally expensive PDES phases. In this paper, we focus on forecasting application iteration times, the key indicator of large-scale networking performance, with network features, such as bandwidth-consumed and busy time on routers. We introduce five representative methods, including LAST, Average, ARIMA, LSTM, and the proposed framework LSTM-Feat, to forecast the iteration times of an exemplar application MILC running on a dragonfly system. By incorporating network features, LSTM-Feat can understand dependencies between network features and iteration times, thus facilitating forecasts. The experiments demonstrate the effectiveness of incorporating network features into surrogate models and the potential of surrogate models to accelerate PDES. Xiongxiao Xu, Kevin A. Brown, Tanwi Mallick, Xin Wang 0115, Elkin Cruz-Camacho, Robert B. Ross, Christopher D. Carothers, Zhiling Lan, Kai Shu |
SIGSIM-PADS | 8 |
| 2023 | Interpretable Modeling of Deep Reinforcement Learning Driven SchedulingabstractIn the field of high-performance computing (HPC), there has been recent exploration into the use of deep reinforcement learning for cluster scheduling (DRL scheduling), which has demonstrated promising outcomes. However, a significant challenge arises from the lack of interpretability in deep neural networks (DNN), rendering them as black-box models to system managers. This lack of model interpretability hinders the practical deployment of DRL scheduling. In this work, we present a framework called IRL (Interpretable Reinforcement Learning) to address the issue of interpretability of DRL scheduling. The core idea is to interpret DNN (i.e., the DRL policy) as a decision tree by utilizing imitation learning. Unlike DNN, decision tree models are non-parametric and easily comprehensible to humans. To extract an effective and efficient decision tree, IRL incorporates the Dataset Aggregation (DAgger) algorithm and introduces the notion of critical state to prune the derived decision tree. Through trace-based experiments, we demonstrate that IRL is capable of converting a black-box DNN policy into an interpretable rule-based decision tree while maintaining comparable scheduling performance. Additionally, IRL can contribute to the setting of rewards in DRL scheduling. Boyang Li 0018, Zhiling Lan, Michael E. Papka |
MASCOTS | 2 |
| 2023 | Hybrid PDES Simulation of HPC Networks Using Zombie PacketsabstractHigh-fidelity network simulations provide insights into new realms for high-performance computing (HPC) architectures, although at a high cost. Surrogate models offer a significant reduction in runtime, yet they cannot serve as complete replacements and should be only used when appropriate. Thus the need for hybrid modeling, where high-fidelity simulation and surrogates run side-by-side. We present a surrogate model for HPC networks in which packets bypass the network, and the network state itself is suspended when switching to the surrogate. To bypass the network, every packet is scheduled to arrive at a predicted time in the future estimated from historical data; to suspend the network, all in-flight packets are delivered to their destinations, but they are kept in the system to awaken as zombies when switching back to high-fidelity. Speedup for a hybrid model is relative to the proportion of surrogate to high-fidelity. We obtained a 3 × speedup for a simulation where 70% of virtual time was spent in surrogate mode. When considering the surrogate portion only, the speedup jumps to nearly 20 × on a uniform random network traffic example. The accuracy of the overall simulation increased when the network state was suspended instead of ignored, which demonstrates the need for modeling the network state when transitioning from surrogate back to high-fidelity mode. Elkin Cruz-Camacho, Kevin A. Brown, Xin Wang 0115, Xiongxiao Xu, Kai Shu, Zhiling Lan, Robert B. Ross, Christopher D. Carothers |
SIGSIM-PADS | 6 |
| 2023 | Workload Interference Prevention with Intelligent Routing and Flexible Job Placement on DragonflyabstractDragonfly is an indispensable interconnect topology for exascale HPC systems. To link tens of thousands of compute nodes at a reasonable cost, Dragonfly shares network resources with the entire system such that network bandwidth is not exclusive to any single job. Since HPC systems are usually shared between multiple co-running workloads at the same time, network competition between co-existing workloads is inevitable. This network contention appears as workload interference, where a job’s network communication can be severely delayed by other jobs. Recent studies show that, compared with the deployed adaptive routing algorithms, an intelligent routing solution based on reinforcement learning named Q-adaptive routing can reduce workload interference. In addition to improving routing efficiency, job placement is a simple yet effective method to mitigate workload interference. In this study, we leverage the well-known parallel discrete event simulation toolkit, SST, to investigate workload interference on Dragonfly with three contributions. We first develop an automatic module that serves as the bridge between SST and HPC job scheduler for automatic simulation configuration and automated simulation launching. Next, we propose a flexible job placement strategy that can mitigate workload interference based on workload communication characteristics. Finally, we extensively examine the workload interference under various job placement and routing configurations. Xin Wang 0115, Zhiling Lan |
SIGSIM-PADS | 3 |
| 2023 | Machine Learning for Interconnect Network Traffic Forecasting: Investigation and ExploitationabstractInterconnect networks play a key role in high-performance computing (HPC) systems. Parallel discrete event simulation (PDES) has been a long-standing pillar for studying large-scale networking systems by replicating the real-world behaviors of HPC facilities. However, the simulation requirements and computational complexity of PDES are growing at an intractable rate. An active research topic is to build a surrogate-ready PDES framework where an accurate surrogate model built on machine learning can be used to forecast network traffic for improving PDES. In this paper, we make the first attempt to introduce two representative time series methods, the Autoregressive Integrated Moving Average (ARIMA) and the Adaptive Long Short-Term Memory (ADP-LSTM), to forecast the traffic in interconnect networks, using the Dragonfly system as a representative example. The proposed ADP-LSTM can efficiently adapt to the ever-changing network traffic, facilitating the forecasting capability for intricate network traffic, by incorporating a novel online learning strategy. Our preliminary analysis demonstrates promising results and shows that ADP-LSTM can consistently outperform ARIMA with significantly less time overhead. Xiongxiao Xu, Xin Wang 0115, Elkin Cruz-Camacho, Christopher D. Carothers, Kevin A. Brown, Robert B. Ross, Zhiling Lan, Kai Shu |
SIGSIM-PADS | 7 |
| 2023 | Performance and power modeling and prediction using MuMMI and 10 machine learning methodsabstractSummary Energy‐efficient scientific applications require insight into how high performance computing system features impact the applications' power and performance. This insight can result from the development of performance and power models. In this article, we use the modeling and prediction tool MuMMI (Multiple Metrics Modeling Infrastructure) and 10 machine learning methods to model and predict performance and power consumption and compare their prediction error rates. We use an algorithm‐based fault‐tolerant linear algebra code and a multilevel checkpointing fault‐tolerant heat distribution code to conduct our modeling and prediction study on the Cray XC40 Theta and IBM BG/Q Mira at Argonne National Laboratory and the Intel Haswell cluster Shepard at Sandia National Laboratories. Our experimental results show that the prediction error rates in performance and power using MuMMI are less than 10% for most cases. By utilizing the models for runtime, node power, CPU power, and memory power, we identify the most significant performance counters for potential application optimizations, and we predict theoretical outcomes of the optimizations. Based on two collected datasets, we analyze and compare the prediction accuracy in performance and power consumption using MuMMI and 10 machine learning methods. Xingfu Wu, Valerie Taylor 0001, Zhiling Lan |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | MRSch: Multi-Resource Scheduling for HPCabstractEmerging workloads in high-performance computing (HPC) are embracing significant changes, such as having diverse resource requirements instead of being CPU -centric. This advancement forces cluster schedulers to consider multiple schedulable resources during decision-making. Existing scheduling studies rely on heuristic or optimization methods, which are limited by an inability to adapt to new scenarios for ensuring long-term scheduling performance. We present an intelligent scheduling agent named MRSch for multi-resource scheduling in HPC that leverages direct future prediction (DFP), an advanced multi-objective reinforcement learning algorithm. While DFP demonstrated outstanding performance in a gaming competition, it has not been previously explored in the context of HPC scheduling. Several key techniques are developed in this study to tackle the challenges involved in multi-resource scheduling. These techniques enable MRSch to learn an appropriate scheduling pol-icy automatically and dynamically adapt its policy in response to workload changes via dynamic resource prioritizing. We compare MRSch with existing scheduling methods through extensive trace-base simulations. Our results demonstrate that MRSch improves scheduling performance by up to 48 % compared to the existing scheduling methods. Boyang Li 0018, Yuping Fan, Matthew T. Dearing, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka |
CLUSTER | 4 |
| 2022 | Hybrid Workload Scheduling on HPC SystemsabstractTraditionally, on-demand, rigid, and malleable applications have been scheduled and executed on separate systems. The ever-growing workload demands and rapidly developing HPC infrastructure trigger the interest of converging these applications on a single HPC system. Although allocating the hybrid workloads within one system could potentially improve system efficiency, it is difficult to balance the tradeoff between the responsiveness of on-demand requests, incentive for malleable jobs, and the performance of rigid applications. In this study, we present several scheduling mechanisms to address the issues involved in co-scheduling on-demand, rigid, and malleable jobs on a single HPC system. We extensively evaluate and compare their performance under various configurations and workloads. Our experimental results show that our proposed mechanisms are capable of serving on-demand workloads with minimal delay, offering incentives for declaring malleability, and improving system performance. Yuping Fan, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka |
IPDPS | 2 |
| 2022 | Encoding for Reinforcement Learning Driven Scheduling
Boyang Li 0018, Yuping Fan, Michael E. Papka, Zhiling Lan |
JSSPP | 4 |
| 2022 | Study of Workload Interference with Intelligent Routing on DragonflyabstractDragonfly interconnect is a crucial network technol-ogy for supercomputers. To support exascale systems, network resources are shared such that links and routers are not dedicated to any node pair. While link utilization is increased, workload performance is often offset by network contention. Recently, intelligent routing built on reinforcement learning demonstrates higher network throughput with lower packet latency. However, its effectiveness in reducing workload interference is unknown. In this work, we present extensive network simulations to study multi-workload contention under different routing mechanisms, intelligent routing and adaptive routing, on a large-scale Dragon-fly system. We develop an enhanced network simulation toolkit, along with a suite of workloads with distinctive communication patterns. We also present two metrics to characterize application communication intensity. Our analysis focuses on examining how different workloads interfere with each other under different routing mechanisms by inspecting both application-level and network-level metrics. Several key insights are made from the analysis. Xin Wang 0115, Zhiling Lan |
SC | 3 |
| 2022 | DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance ComputingabstractCluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. The experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%. Yuping Fan, Boyang Li 0018, Dustin Favorite, Naunidh Singh, John T. Childers, Paul M. Rich, William E. Allcock, Michael E. Papka, Zhiling Lan |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2021 | A Dynamic Power Capping Library for HPC ApplicationsabstractAs HPC systems increase in scale and capability, the cost of supplying power to these systems grows significantly. This introduces the urgent need for energy efficient computing through the management of power consumption. The PowerStack initiative defines a holistic power management framework at three levels, i.e., at the cluster level, the job level, and the node level [1] –[3]. It is expected that a cluster will be given a system-wide power budget (aka an allocated power budget), which can be intelligently managed and distributed at the job and node levels. Our work provides a method for intelligently managing node-level power during execution. Zhiling Lan, Xingfu Wu, Valerie Taylor 0001 |
CLUSTER | 2 |
| 2021 | Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly NetworkabstractHigh-radix interconnects such as Dragonfly and its variants rely on adaptive routing to balance network traffic for optimum performance. Ideally, adaptive routing attempts to forward packets between minimal and non-minimal paths with the least congestion. In practice, current adaptive routing algorithms estimate routing path congestion based on local information such as output queue occupancy. Using local information to estimate global path congestion is inevitably inaccurate because a router has no precise knowledge of link states a few hops away. This inaccuracy could lead to interconnect congestion. In this study, we present Q-adaptive routing, a multi-agent reinforcement learning routing scheme for Dragonfly systems. Q-adaptive routing enables routers to learn to route autonomously by leveraging advanced reinforcement learning technology. The proposed Q-adaptive routing is highly scalable thanks to its fully distributed nature without using any shared information between routers. Furthermore, a new two-level Q-table is designed for Q-adaptive to make it computational lightly and saves 50% of router memory usage compared with the previous Q-routing. We implement the proposed Q-adaptive routing in SST/Merlin simulator. Our evaluation results show that Q-adaptive routing achieves up to 10.5% system throughput improvement and 5.2x average packet latency reduction compared with adaptive routing algorithms. Remarkably, Q-adaptive can even outperform the optimal VALn non-minimal routing under the ADV+1 adversarial traffic pattern with up to 3% system throughput improvement and 75% average packet latency reduction. Xin Wang 0115, Zhiling Lan |
HPDC | 3 |
| 2021 | Deep Reinforcement Agent for Scheduling in HPCabstractCluster scheduler is crucial in high-performance computing (HPC). It determines when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a novel, hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. A unique training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. The experiments with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 45%. Yuping Fan, Zhiling Lan, John T. Childers, Paul M. Rich, William E. Allcock, Michael E. Papka |
IPDPS | 2 |
| 2020 | Union: An Automatic Workload Manager for Accelerating Network SimulationabstractWith the rapid growth of the machine learning applications, the workloads of future HPC systems are anticipated to be a mix of scientific simulation, big data analytics, and machine learning applications. Simulation is a great research vehicle to understand the performance implications of co-running scientific applications with big data and machine learning workloads on large-scale systems. In this paper, we present Union, a workload manager that provides an automatic framework to facilitate hybrid workload simulation in CODES. Furthermore, we use Union, along with CODES, to investigate various hybrid workloads composed of traditional simulation applications and emerging learning applications on two dragonfly systems. The experiment results show that both message latency and communication time are important performance metrics to evaluate network interference. Network interference on HPC applications is more reflected by the message latency variation, whereas ML application performance depends more on the communication time. Xin Wang 0115, Misbah Mubarak, Robert B. Ross, Zhiling Lan |
IPDPS | 5 |
| 2019 | Scheduling Beyond CPUs for HPCabstractHigh performance computing (HPC) is undergoing significant changes. The emerging HPC applications comprise both compute- and data-intensive applications. To meet the intense I/O demand from emerging data-intensive applications, burst buffers are deployed in production systems. Existing HPC schedulers are mainly CPU-centric. The extreme heterogeneity of hardware devices, combined with workload changes, forces the schedulers to consider multiple resources (e.g., burst buffers) beyond CPUs, in decision making. In this study, we present a multi-resource scheduling scheme named BBSched that schedules user jobs based on not only their CPU requirements, but also other schedulable resources such as burst buffer. BBSched formulates the scheduling problem into a multi-objective optimization (MOO) problem and rapidly solves the problem using a multi-objective genetic algorithm. The multiple solutions generated by BBSched enables system managers to explore potential tradeoffs among various resources, and therefore obtains better utilization of all the resources. The trace-driven simulations with real system workloads demonstrate that BBSched improves scheduling performance by up to 41% compared to existing methods, indicating that explicitly optimizing multiple resources beyond CPUs is essential for HPC scheduling. Yuping Fan, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka, Brian Austin, David Paul |
HPDC | 2 |
| 2019 | Modeling and Analysis of Application Interference on Dragonfly+abstractDragonfly class of networks are considered as promising interconnects for next-generation supercomputers. While Dragonfly+ networks offer more path diversity than the original Dragonfly design, they are still prone to performance variability due to their hierarchical architecture and resource sharing design. Event-driven network simulators are indispensable tools for navigating complex system design. In this study, we quantitatively evaluate a variety of application communication interactions on a 3,456-node Dragonfly+ system by using the CODES toolkit. This study looks at the impact of communication interference from a user's perspective. Specifically, for a given application submitted by a user, we examine how this application will behave with the existing workload running in the system under different job placement policies. Our simulation study considers hundreds of experiment configurations including four target applications with representative communication patterns under a variety of network traffic conditions. Our study shows that intra-job interference can cause severe performance degradation for communication-intensive applications. Inter-job interference can generally be reduced for applications with one-to-one or one-to-many communication patterns through job isolation. Application with one-to-all communication pattern is resilient to network interference. Xin Wang 0115, Neil McGlohon, Misbah Mubarak, Sudheer Chunduri, Zhiling Lan |
SIGSIM-PADS | 6 |
| 2018 | Trade-Off Study of Localizing Communication and Balancing Network Traffic on a Dragonfly SystemabstractDragonfly networks are being widely adopted in high-performance computing systems. On these networks, however, interference caused by resource sharing can lead to significant network congestion and performance variability. We present a comparative analysis exploring the trade-off between localizing communication and balancing network traffic. We conduct trace-based simulations for applications with different communication patterns, using multiple job placement policies and routing mechanisms. We perform an in-depth performance analysis on representative applications individually and show that different applications have distinct preferences regarding localized communication and balanced network traffic. We further demonstrate the effect of external network interference by introducing background traffic and show that localized communication can help reduce the application performance variation caused by network sharing. Xin Wang 0115, Misbah Mubarak, Xu Yang 0009, Robert B. Ross, Zhiling Lan |
IPDPS | 5 |
| 2018 | System-wide trade-off modeling of performance, power, and resilience on petascale systems
Li Yu 0006, Zhou Zhou 0006, Yuping Fan, Michael E. Papka, Zhiling Lan |
J. Supercomput. | 5 |
| 2017 | Trade-Off Between Prediction Accuracy and Underestimation Rate in Job Runtime EstimatesabstractJob runtime estimates provided by users are widely acknowledged to be overestimated and runtime overestimation can greatly degrade job scheduling performance. Previous studies focus on improving accuracy of job runtime estimates by reducing runtime overestimation, but fail to address the underestimation problem (i.e., the underestimation of job runtimes). Using an underestimated runtime is catastrophic to a job as the job will be killed by the scheduler before completion. We argue that both the improvement of runtime accuracy and the reduction of underestimation rate are equally important. To address this problem, we propose an online runtime adjustment framework called TRIP. TRIP explores the data censoring capability of the Tobit model to improve prediction accuracy while keeping a low underestimation rate of job runtimes. TRIP can be used as a plugin to job scheduler for improving job runtime estimates and hence boosting job scheduling performance. Preliminary results demonstrate that TRIP is capable of achieving high accuracy of 80% and low underestimation rate of 5%. This is significant as compared to other well-known machine learning methods such as SVM, Random Forest, and Last-2 which result in a high underestimation rate (20%-50%). Our experiments further quantify the amount of scheduling performance gain achieved by the use of TRIP. Yuping Fan, Paul M. Rich, William E. Allcock, Michael E. Papka, Zhiling Lan |
CLUSTER | 5 |
| 2017 | Preliminary Interference Study About Job Placement and Routing Algorithms in the Fat-Tree Topology for HPC ApplicationsabstractAmong the high-radix and low-diameter networks, fat-tree topology is commonly used in HPC and datacenter systems. Resource and job management is critically important to mitigate application interference in order to achieve high system performance and utilization. Preliminary studies have shown the effect of job placement on parallel scientific applications performance. In this work we study interference about job placement and routing algorithms on fat-tree system. Applications can be classified into various groups according to the communication patterns. We further combine various job placement policies and routing algorithms and create six different configurations. The system performance is analyzed by performing fine-grained high-fidelity discrete event-driven simulation. Initial experimentation shows that the performance of HPC applications not only is related with its communication pattern, but also relies on the job placement and network routing on fat-tree systems. Peixin Qiao, Xin Wang 0115, Xu Yang 0009, Yuping Fan, Zhiling Lan |
CLUSTER | 5 |
| 2017 | A Preliminary Study of Intra-Application Interference on Dragonfly NetworkabstractDragonfly network is widely used in modern high-performance computing systems. On this network, however, interference caused by network sharing can lead to significant network congestion and degraded performance. In this work, we present a comparative analysis of intra-application interference on applications with nearest neighbor communication, considering various placement strategies. Our results demonstrate that intra-application interference is basically a trade-off between localized communication and balanced network. We further develop alternative job placement policies with the objective to balance the trade-off between localized communication and balanced network. The results demonstrate that our proposed strategies can effectively balance the trade-off and hence reduce intra-application interference on Dragonfly network. Xin Wang 0115, Xu Yang 0009, Misbah Mubarak, Robert B. Ross, Zhiling Lan |
CLUSTER | 5 |
| 2017 | Experience and Practice of Batch Scheduling on Leadership Supercomputers at Argonne
William E. Allcock, Paul M. Rich, Yuping Fan, Zhiling Lan |
JSSPP | 4 |
| 2017 | Topology mapping of irregular parallel applications on torus-connected supercomputers
Jingjin Wu, Xuanxing Xiong, Eduardo Berrocal, Zhiling Lan |
J. Supercomput. | 5 |
| 2017 | Toward General Software Level Silent Data Corruption Detection for Parallel ApplicationsabstractSilent data corruption (SDC) poses a great challenge for high-performance computing (HPC) applications as we move to extreme-scale systems. Mechanisms have been proposed that are able to detect SDC in HPC applications by using the peculiarities of the data (more specifically, its “smoothness” in time and space) to make predictions. However, these data-analytic solutions are still far from fully protecting applications to a level comparable with more expensive solutions such as full replication. In this work, we propose partial replication to overcome this limitation. More specifically, we have observed that not all processes of an MPI application experience the same level of data variability at exactly the same time. Thus, we can smartly choose and replicate only those processes for which the lightweight data-analytic detectors would perform poorly. In addition, we propose a new evaluation method based on the probability that a corruption will pass unnoticed by a particular detector (instead of just reporting overall single-bit precision and recall). In our experiments, we use four applications dealing with different explosions. Our results indicate that our new approach can protect the MPI applications analyzed with 7-70 percent less overhead (depending on the application) than that of full duplication with similar detection recall. Eduardo Berrocal, Leonardo Arturo Bautista-Gomez, Sheng Di, Zhiling Lan, Franck Cappello |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Exploring Plan-Based Scheduling for Large-Scale Computing SystemsabstractAs HPC systems scale toward exascale, it becomes critical to manage the underlying resource more effectively. While almost all existing resource management systems schedule jobs in a queuing fashion and have drawbacks of making isolated scheduling decisions that would compromise system performance even with backfilling, plan-based schedulers have the potential to generate better job schedules by producing an execution plan of all waiting jobs but do not receive enough attention. In this paper, we present a novel plan-based scheduling system that utilizes simulated annealing as the optimization engine to support effective resource management on HPC systems. As demonstrated by extensive trace-based simulations with workload traces collected from a wide range of production supercomputers, in comparison with the queue-based scheduling system using FCFS with EASY backfilling, our plan-based scheduling system can reduce the job wait time by 40%, reduce the job response time by 30%, while slightly improving system utilization at the same time. Moreover, our plan-based system is able to run online by solving the scheduling problem at each scheduling iteration within one second, making it practical for production HPC systems. Xingwu Zhang, Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan |
CLUSTER | 4 |
| 2016 | Exploring Partial Replication to Improve Lightweight Silent Data Corruption Detection for HPC Applications
Eduardo Berrocal, Leonardo Arturo Bautista-Gomez, Sheng Di, Zhiling Lan, Franck Cappello |
Euro-Par | 4 |
| 2016 | Study of Intra- and Interjob Interference on Torus NetworksabstractNetwork contention between concurrently running jobs on HPC systems is a primary cause of performance variability. Optimizing job allocation and avoiding network sharing are hence crucial to alleviate the potential performance degradation. In order to do so effectively, an understanding of the interference among concurrently running jobs, their communication patterns, and contention in the network is required. In this work, we choose three representative HPC applications from the DOE Design Forward Project and conduct detailed simulations on a torus network model to analyze both intra-and interjob interference. By scrutinizing the communication behaviors of these applications, we identify relationships between these behaviors and the possible interference introduced by different job placement policies. Our analyses illuminate a path toward communication pattern awareness in job placement on HPC systems. Xu Yang 0009, John Jenkins, Misbah Mubarak, Xin Wang 0115, Robert B. Ross, Zhiling Lan |
ICPADS | 6 |
| 2016 | A data driven scheduling approach for power management on HPC systemsabstractModern schedulers running on HPC systems traditionally consider the number of resources and the time requested for each job that is to be executed when making scheduling decisions. Until recently this has been sufficient, however as systems get larger, other metrics like power consumption become necessary to ensure system stability. In this paper, we propose a data driven scheduling approach for controlling the power consumption of the entire system under any user defined budget. Here, “data driven” means that our approach actively observes, analyzes, and assesses power behaviors of the system and user jobs to guide scheduling decisions for power management. This design is based on the key observation that HPC jobs have distinct power profiles. Our work contains an empirical analysis of workload power characteristics on a production system, dynamic learner to estimate the job power profile for scheduling, and an online power-aware scheduler for managing the overall system power. Using real workload traces, we demonstrate that our design effectively controls system power consumption while minimizing the impact on system utilization. Sean Wallace, Xu Yang 0009, Venkatram Vishwanath, William E. Allcock, Susan Coghlan, Michael E. Papka, Zhiling Lan |
SC | 7 |
| 2016 | Watch out for the bully!: job interference study on dragonfly networkabstractHigh-radix, low-diameter dragonfly networks will be a common choice in next-generation supercomputers. Preliminary studies show that random job placement with adaptive routing should be the rule of thumb to utilize such networks, since it uniformly distributes traffic and alleviates congestion. Nevertheless, in this work we find that while random job placement coupled with adaptive routing is good at load balancing network traffic, it cannot guarantee the best performance for every job. The performance improvement of communication-intensive applications comes at the expense of performance degradation of less intensive ones. We identify this bully behavior and validate its underlying causes with the help of detailed network simulation and real application traces. We further investigate a hybrid contiguous-noncontiguous job placement policy as an alternative. Initial experimentation shows that hybrid job placement aids in reducing the worst-case performance degradation for less communication-intensive applications while retaining the performance of communication-intensive ones. Xu Yang 0009, John Jenkins, Misbah Mubarak, Robert B. Ross, Zhiling Lan |
SC | 5 |
| 2016 | Application power profiling on IBM Blue Gene/Q
Sean Wallace, Zhou Zhou 0006, Venkatram Vishwanath, Susan Coghlan, John R. Tramm, Zhiling Lan, Michael E. Papka |
Parallel Comput. | 6 |
| 2016 | I/O-aware bandwidth allocation for petascale computing systems
Zhou Zhou 0006, Xu Yang 0009, Dongfang Zhao 0001, Paul M. Rich, Wei Tang 0001, Zhiling Lan |
Parallel Comput. | 7 |
| 2016 | A Scalable, Non-Parametric Method for Detecting Performance Anomaly in Large Scale ComputingabstractAs computer systems continue to grow in scale and complexity, performance problems become common and a major concern for large-scale computing. Performance anomalies caused by application bugs, hardware or software faults, or resource contention can have great impact on system-wide performance and could lead to significant economic losses for service providers. While many detection methods have been presented in the past, the newly emerging challenges are detection scalability and practical use. In this paper, we propose a scalable, non-parametric method for effectively detecting performance anomalies in large-scale systems. The design is generic for anomaly detection in a variety of parallel and distributed systems exhibiting peer-comparable property. It adopts a divide-and-conquer approach to address the scalability challenge and explores the use of non-parametric clustering and two-phase majority voting to improve detection flexibility and accuracy. We derive probabilistic models to quantitatively evaluate our decentralized design. Experiments with a suite of applications on production systems demonstrate that this method outperforms existing methods in terms of detection accuracy with a negligible runtime overhead. Li Yu 0006, Zhiling Lan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation ConstraintsabstractAs systems scale toward exascale, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many-such as network bandwidth, I/O bandwidth, or power-have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potential of relaxing network allocation constraints for Blue Gene systems. Our objective is to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is three new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing scheduler on Mira, a 48-rack Blue Gene/Q system at Argonne National Laboratory. Specifically, we use job traces collected from this production system. Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Comparison of Vendor Supplied Environmental Data Collection MechanismsabstractThe high performance computing landscape is filled with diverse hardware components. A large part of understanding how these components compare to others is by looking at the various environmental aspects of these devices such as power consumption, temperature, etc. Thankfully, vendors of these various pieces of hardware have supported this by providing mechanisms to obtain this data. However, differences not only in the way this data is obtained but also the data which is provided is common between products. In this paper, we take a comprehensive look at the data which is available for the most common pieces of today's HPC landscape, as well as how this data is obtained and how accurate it is. Having surveyed these components, we compare and contrast them noting key differences as well as providing insight into what features future components should have. Sean Wallace, Venkatram Vishwanath, Susan Coghlan, Zhiling Lan, Michael E. Papka |
CLUSTER | 4 |
| 2015 | I/O-Aware Batch Scheduling for Petascale Computing SystemsabstractIn the Big Data era, the gap between the storage performance and an application's I/O requirement is increasing. I/O congestion caused by concurrent storage accesses from multiple applications is inevitable and severely harms the performance. Conventional approaches either focus on optimizing an application's access pattern individually or handle I/O requests on a low-level storage layer without any knowledge from the upper-level applications. In this paper, we present a novel I/O-aware batch scheduling framework to coordinate ongoing I/O requests on petascale computing systems. The motivation behind this innovation is that the batch scheduler has a holistic view of both the system state and jobs' activities and can control the jobs' status on the fly during their execution. We treat a job's I/O requests as periodical subjobs within its lifecycle and transform the I/O congestion issue into a classical scheduling problem. We design two scheduling polices with different scheduling objectives either on user-oriented metrics or system performance. We conduct extensive trace-based simulations using real job traces and I/O traces from a production IBM Blue Gene/Q system. Experimental results demonstrate that our design can improve job performance by more than 30%, as well as increasing system performance. Zhou Zhou 0006, Xu Yang 0009, Dongfang Zhao 0001, Paul M. Rich, Wei Tang 0001, Zhiling Lan |
CLUSTER | 7 |
| 2015 | Lightweight Silent Data Corruption Detection Based on Runtime Data Analysis for HPC ApplicationsabstractNext-generation supercomputers are expected to have more components and, at the same time, consume several times less energy per operation. Consequently, the number of soft errors is expected to increase dramatically in the coming years. In this respect, techniques that leverage certain properties of iterative HPC applications (such as the smoothness of the evolution of a particular dataset) can be used to detect silent errors at the application level. In this paper, we present a pointwise detection model with two phases: one involving the prediction of the next expected value in the time series for each data point, and another determining a range (i.e., normal value interval) surrounding the predicted next-step value. We show that dataset correlation can be used to detect corruptions indirectly and limit the size of the data set to monitor, taking advantage of the underlying physics of the simulation. Our results show that, using our techniques, we can detect a large number of corruptions (i.e., above 90% in some cases) with 84% memory overhead, and 13.75% extra computation time. Eduardo Berrocal, Leonardo Arturo Bautista-Gomez, Sheng Di, Zhiling Lan, Franck Cappello |
HPDC | 4 |
| 2015 | Improving Batch Scheduling on Blue Gene/Q by Relaxing 5D Torus Network Allocation ConstraintsabstractAs systems scale toward exactable, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many -- such as network bandwidth, I/O bandwidth, or power -- have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potentiality of relaxing network allocation constraints for Blue Gene systems. Our objectives to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is two new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing one under different workloads, using job traces collected from the 48-rack Mira, an IBM Blue Gene/Q system at Argonne National Laboratory. Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai |
IPDPS | 3 |
| 2015 | Quantitative modeling of power performance tradeoffs on extreme scale systems
Li Yu 0006, Zhou Zhou 0006, Sean Wallace, Michael E. Papka, Zhiling Lan |
J. Parallel Distributed Comput. | 5 |
| 2015 | Reliability-Aware Speedup Models for Parallel Applications with Coordinated Checkpointing/RestartabstractSpeedup models are powerful analytical tools for evaluating and predicting the performance of parallel applications. Unfortunately, the well-known speedup models like Amdahl’s law and Gustafson’s law do not take reliability into consideration and therefore cannot accurately account for application performance in the presence of failures. In this study, we enhance Amdahl’s law and Gustafson’s law by considering the impact of failures and the effect of coordinated checkpointing/restart. Unlike existing analytical studies relying on Exponential failure distribution alone, in this work we consider both Exponential and Weibull failure distributions in the construction of our reliability-aware speedup models. The derived reliability-aware models are validated through trace-based simulations under a variety of parameter settings. Our trace-based simulations demonstrate these models can effectively quantify failure impact on application speedup. Moreover, we present two case studies to illustrate the use of these reliability-aware speedup models. Ziming Zheng, Li Yu 0006, Zhiling Lan |
IEEE Trans. Computers | 3 |
| 2015 | Hierarchical task mapping for parallel applications on supercomputers
Jingjin Wu, Xuanxing Xiong, Zhiling Lan |
J. Supercomput. | 3 |
| 2014 | Exploring void search for fault detection on extreme scale systemsabstractMean Time Between Failures (MTBF), now calculated in days or hours, is expected to drop to minutes on exascale machines. The advancement of resilience technologies greatly depends on a deeper understanding of faults arising from hardware and software components. This understanding has the potential to help us build better fault tolerance technologies. For instance, it has been proved that combining checkpointing and failure prediction leads to longer checkpoint intervals, which in turn leads to fewer total checkpoints. In this paper we present a new approach for fault detection based on the Void Search (VS) algorithm. VS is used primarily in astrophysics for finding areas of space that have a very low density of galaxies. We evaluate our algorithm using real environmental logs from Mira Blue Gene/Q supercomputer at Argonne National Laboratory. Our experiments show that our approach can detect almost all faults (i.e., sensitivity close to 1) with a low false positive rate (i.e., specificity values above 0.7). We also compare our algorithm with a number of existing detection algorithms, and find that ours outperforms all of them. Eduardo Berrocal, Li Yu 0006, Sean Wallace, Michael E. Papka, Zhiling Lan |
CLUSTER | 5 |
| 2014 | Balancing job performance with system performance via locality-aware scheduling on torus-connected systemsabstractTorus-connected network is widely used in modern supercomputers due to its linear per node cost scaling and its competitive overall performance. Job scheduling system plays a critical role for the efficient use of supercomputers. As supercomputers continue growing in size, a fundamental problem arises: how to effectively balance job performance with system performance on torus-connected machines? In this work, we will present a new scheduling design named window-based locality-aware scheduling. Our design contains three novel features. First, rather than one-by-one job scheduling, our design takes a “window” of jobs, i.e. multiple jobs, into consideration for job prioritizing and resource allocation. Second, our design maintains a list of slots to preserve node contiguity information for resource allocation. Finally, we formulate our scheduling decision making into a 0-1 Multiple Knapsack Problem and present two algorithms to solve the problem. A series of trace-based simulations using job logs collected from production supercomputers indicate that this new scheduling design has real potentials and can effectively balance job performance and system performance. Xu Yang 0009, Zhou Zhou 0006, Wei Tang 0001, Xingwu Zheng, Zhiling Lan |
CLUSTER | 6 |
| 2013 | Application power profiling on IBM Blue Gene/QabstractThe power consumption of state of the art supercomputers, because of their complexity and unpredictable workloads, is extremely difficult to estimate. Accurate and precise results, as are now possible with the latest generation of supercomputers, are therefore a welcome addition to the landscape. Only recently have end users been afforded the ability to access the power consumption of their applications. However, just because it's possible for end users to obtain this data does not mean it's a trivial task. This emergence of new data is therefore not only understudied, but also not fully understood. In this paper, we provide detailed power consumption analysis of microbenchmarks running on Argonne's latest generation of IBM Blue Gene supercomputers, Mira, a Blue Gene/Q system. The analysis is done utilizing our power monitoring library, MonEQ, built on the IBM provided Environmental Monitoring (EMON) API. We describe the importance of sub-second polling of various power domains and the implications they present. To this end, previously well understood applications will now have new facets of potential analysis. Sean Wallace, Venkatram Vishwanath, Susan Coghlan, John R. Tramm, Zhiling Lan, Michael E. Papka |
CLUSTER | 5 |
| 2013 | A Transparent Collective I/O ImplementationabstractAbstract—I/O performance is vital for most HPC applications especially those that generate a vast amount of data with the growth of scale. Many studies have shown that scientific applications tend to issue small and noncontiguous accesses in an interleaving fashion, causing different processes to access overlapping regions. In such scenario, collective I/O is a widely used optimization technique. However, the use of collective I/O deployed in existing MPI implementations is not trivial and sometimes even impossible. Collective I/O is an optimization based on a single collective I/O access. If the data reside in different places (e.g. in different arrays), the application has to maintain a buffer to first combine these data and then perform I/O operations on the buffer rather than the original data pieces. The process is very tedious for application developers. Besides, collective I/O requires the creating of a file view to describe the Yongen Yu, Jingjin Wu, Zhiling Lan, Douglas H. Rudd, Nickolay Y. Gnedin, Andrey V. Kravtsov |
IPDPS | 3 |
| 2013 | Reducing Energy Costs for IBM Blue Gene/P via Power-Aware Job Scheduling
Zhou Zhou 0006, Zhiling Lan, Wei Tang 0001, Narayan Desai |
JSSPP | 2 |
| 2013 | Integrating dynamic pricing of electricity into energy aware scheduling for HPC systemsabstractThe research literature to date mainly aimed at reducing energy consumption in HPC environments. In this paper we propose a job power aware scheduling mechanism to reduce HPC's electricity bill without degrading the system utilization. The novelty of our job scheduling mechanism is its ability to take the variation of electricity price into consideration as a means to make better decisions of the timing of scheduling jobs with diverse power profiles. We verified the effectiveness of our design by conducting trace-based experiments on an IBM Blue Gene/P and a cluster system as well as a case study on Argonne's 48-rack IBM Blue Gene/Q system. Our preliminary results show that our power aware algorithm can reduce electricity bill of HPC systems as much as 23%. Xu Yang 0009, Zhou Zhou 0006, Sean Wallace, Zhiling Lan, Wei Tang 0001, Susan Coghlan, Michael E. Papka |
SC | 4 |
| 2013 | Job scheduling with adjusted runtime estimates on production supercomputers
Wei Tang 0001, Narayan Desai, Daniel Buettner, Zhiling Lan |
J. Parallel Distributed Comput. | 4 |
| 2013 | Toward balanced and sustainable job scheduling for production supercomputers
Wei Tang 0001, Dongxu Ren, Zhiling Lan, Narayan Desai |
Parallel Comput. | 3 |
| 2013 | Multi-domain job coscheduling for leadership computing systems
Wei Tang 0001, Narayan Desai, Venkatram Vishwanath, Daniel Buettner, Zhiling Lan |
J. Supercomput. | 5 |
| 2012 | Filtering log data: Finding the needles in the HaystackabstractLog data is an incredible asset for troubleshooting in large-scale systems. Nevertheless, due to the ever-growing system scale, the volume of such data becomes overwhelming, bringing enormous burdens on both data storage and data analysis. To address this problem, we present a 2-dimensional online filtering mechanism to remove redundant and noisy data via feature selection and instance selection. The objective of this work is two-fold: (i) to significantly reduce data volume without losing important information, and (ii) to effectively promote data analysis. We evaluate this new filtering mechanism by means of real environmental data from the production supercomputers at Oak Ridge National Laboratory and Sandia National Laboratory. Our preliminary results demonstrate that our method can reduce more than 85% disk space, thereby significantly reducing analysis time. Moreover, it also facilitates better failure prediction and diagnosis by more than 20%, as compared to the conventional predictive approach relying on RAS (Reliability, Availability, and Serviceability) events alone. Li Yu 0006, Ziming Zheng, Zhiling Lan, Terry R. Jones, Jim M. Brandt, Ann C. Gentile |
DSN | 3 |
| 2012 | Improving Parallel IO Performance of Cell-based AMR Cosmology ApplicationsabstractTo effectively model various regions with different resolutions, adaptive mesh refinement (AMR) is commonly used in cosmology simulations. There are two well-known numerical approaches towards the implementation of AMR based cosmology simulations: block-based AMR and cell-based AMR. While many studies have been conducted to improve performance and scalability of block-structured AMR applications, little work has been done for cell-based simulations. In this study, we present a parallel IO design for cell-based AMR cosmology applications, in particular, the ART(Adaptive Refinement Tree) code. First, we design a new data format that incorporates a space filling curve to map between spatial and on-disk locations. This indexing not only enables concurrent IO accesses from multiple application processes, but also allows users to extract local regions without significant additional memory, CPU or disk space overheads. Second, we develop a flexible N-M mapping mechanism to harvest the benefits of N-N and N-1 mappings where N is number of application processes and M is a user-tunable parameter for number of files. It not only overcomes the limited bandwidth issue of an N-1 mapping by allowing the creation of multiple files, but also enables users to efficiently restart the application at a variety of computing scales. Third, we develop a user-level library to transparently and automatically aggregate small IO accesses per process to accelerate IO performance. We evaluate this new parallel IO design by means of real cosmology simulations on production HPC system at TACC. Our preliminary results indicate that it can not only provide the functionality required by scientists (e.g., effective extraction of local regions and flexible process-to file mapping), but also significantly improve IO performance. Yongen Yu, Douglas H. Rudd, Zhiling Lan, Nickolay Y. Gnedin, Andrey V. Kravtsov, Jingjin Wu |
IPDPS | 3 |
| 2012 | Hierarchical task mapping of cell-based AMR cosmology simulationsabstractCosmology simulations are highly communication-intensive, thus it is critical to exploit topology-aware task mapping techniques for performance optimization. To exploit the architectural properties of multiprocessor clusters (the performance gap between inter-node and intra-node communication as well as the gap between inter-socket and intra-socket communication), we design and develop a hierarchical task mapping scheme for cell-based AMR (Adaptive Mesh Refinement) cosmology simulations, in particular, the ART application. Our scheme consists of two parts: (1) an inter-node mapping to map application processes onto nodes with the objective of minimizing network traffic among nodes and (2) an intra-node mapping within each node to minimize the maximum size of messages transmitted between CPU sockets. Experiments on production supercomputers with 3D torus and fat-tree topologies show that our scheme can significantly reduce application communication cost by up to 50%. More importantly, our scheme is generic and can be extended to many other applications. Jingjin Wu, Zhiling Lan, Xuanxing Xiong, Nickolay Y. Gnedin, Andrey V. Kravtsov |
SC | 2 |
| 2011 | Performance Emulation of Cell-Based AMR Cosmology SimulationsabstractCosmological simulations are highly complicated, and it is time-consuming to redesign and reimplement the code for improvement. Moreover, it is a risk to implement any idea directly in the code without knowing its effects on performance. In this paper, we design an emulator for cell-based AMR (adaptive mesh refinement) cosmology simulations, in particular, the Adaptive Refinement Tree (ART) application. ART is an advanced "hydro+N-body" simulation tool integrating extensive physics processes for cosmological research. The emulator is designed based on the behaviors of cell-based AMR cosmology simulations, and quantitative performance models are built toward the design of the emulator. Our experiments with realistic cosmology simulations on production supercomputers indicate that the emulator is accurate. Moreover, we evaluate and compare three different load balancing schemes for cell-based cosmology simulations via the emulator. The comparison results provide us useful insight into the performance and scalability of different load balance schemes. Jingjin Wu, Roberto E. González, Zhiling Lan, Nickolay Y. Gnedin, Andrey V. Kravtsov, Douglas H. Rudd, Yongen Yu |
CLUSTER | 3 |
| 2011 | Evaluating Performance Impacts of Delayed Failure Repairing on Large-Scale SystemsabstractWith the fast improvement in technology, we are now moving toward exascale computing. Many experts predict that exascale computers will have millions of nodes, billions of threads of execution, hundreds of petabytes of inner memory and exabytes of persistent storage. For systems of such a scale, frequent failures are becoming a serious concern. One of the most important reasons is that in a large-scale system it is hard to detect failures. As a result, failure repair may take substantial time. In this paper, we investigate the effect of delayed repairing on two popular types of high-performance computing systems: IBM Blue Gene/P and general cluster. We analyze how delayed failure repairing will affect the performance of jobs when some computing units are at fault but not fixed in time. Our study is based on real workload traces and RAS logs collected from production supercomputing systems. Our Trace-based simulations indicate that fast failure detection and recovery is essential for moving towards petascale and beyond computing. Zhou Zhou 0006, Wei Tang 0001, Ziming Zheng, Zhiling Lan, Narayan Desai |
CLUSTER | 4 |
| 2011 | Reducing Fragmentation on Torus-Connected SupercomputersabstractTorus-based networks are prevalent on leadership-class petascale systems, providing a good balance between network cost and performance. The major disadvantage of this network architecture is its susceptibility to fragmentation. Many studies have attempted to reduce resource fragmentation in this architecture. Although the approaches suggested can make good allocation decisions reducing fragmentation at job start time, none of them considers a job's wall time, which can cause resource fragmentation when neighboring jobs do not complete closely. In this paper, we propose a wall time-aware job allocation strategy, which adjacently packs jobs that finish around the same time, in order to minimize resource fragmentation caused by job length, discrepancy. Event-driven simulations using real job traces from a production Blue Gene/P system at Argonne National Laboratory demonstrate that our wall time-aware strategy can effectively reduce system fragmentation and improve overall system performance. Wei Tang 0001, Zhiling Lan, Narayan Desai, Daniel Buettner, Yongen Yu |
IPDPS | 2 |
| 2011 | Co-analysis of RAS Log and Job Log on Blue Gene/PabstractWith the growth of system size and complexity, reliability has become of paramount importance for petascale systems. Reliability, Availability, and Serviceability (RAS) logs have been commonly used for failure analysis. However, analysis based on just the RAS logs has proved to be insufficient in understanding failures and system behaviors. To overcome the limitation of this existing methodologies, we analyze the Blue Gene/P RAS logs and the Blue Gene/P job logs in a cooperative manner. From our co-analysis effort, we have identified a dozen important observations about failure characteristics and job interruption characteristics on the Blue Gene/P systems. These observations can significantly facilitate the research in fault resilience of large-scale systems. Ziming Zheng, Li Yu 0006, Wei Tang 0001, Zhiling Lan, Rinku Gupta, Narayan Desai, Susan Coghlan, Daniel Buettner |
IPDPS | 4 |
| 2011 | FREM: A Fast Restart Mechanism for General Checkpoint/RestartabstractAs failure rate keeps on increasing in large systems, applications running atop restart more frequently than ever. Existing research on checkpoint/restart mainly focuses on optimizing checkpoint operation, without paying much attention to the restart operation. As a result, application restart latency maybe substantial, which greatly threatens system dependability and performance. To attack the restart latency problem, in this paper, we present FREM, a fast restart mechanism for general checkpoint/restart protocols. By dynamically tracking the process data accesses after each checkpoint, FREM masks restart latency by overlapping application recovery with the retrieval of its checkpoint image. We have implemented FREM as a prototype system and tested it under Linux environments. Extensive experiments with real applications demonstrate that it can effectively reduce restart latency by over 50 percent on average, as compared to the conventional restart mechanisms. Zhiling Lan |
IEEE Trans. Computers | 2 |
| 2010 | Analyzing and adjusting user runtime estimates to improve job scheduling on the Blue Gene/PabstractBackfilling and short-job-first are widely acknowledged enhancements to the simple but popular first-come, first-served job scheduling policy. However, both enhancements depend on user-provided estimates of job runtime, which research has repeatedly shown to be inaccurate. We have investigated the effects of this inaccuracy on backfilling and different queue prioritization policies, determining which part of the scheduling policy is most sensitive. Using these results, we have designed and implemented several estimation-adjusting schemes based on historical data. We have evaluated these schemes using workload traces from the Blue Gene/P system at Argonne National Laboratory. Our experimental results demonstrate that dynamically adjusting job runtime estimates can improve job scheduling performance by up to 20%. Wei Tang 0001, Narayan Desai, Daniel Buettner, Zhiling Lan |
IPDPS | 4 |
| 2010 | A study of dynamic meta-learning for failure prediction in large-scale systems
Zhiling Lan, Jiexing Gu, Ziming Zheng, Rajeev Thakur, Susan Coghlan |
J. Parallel Distributed Comput. | 1 |
| 2010 | Toward Automated Anomaly Identification in Large-Scale SystemsabstractWhen a system fails to function properly, health-related data are collected for troubleshooting. However, it is challenging to effectively identify anomalies from the voluminous amount of noisy, high-dimensional data. The traditional manual approach is time-consuming, error-prone, and even worse, not scalable. In this paper, we present an automated mechanism for node-level anomaly identification in large-scale systems. A set of techniques is presented to automatically analyze collected data: data transformation to construct a uniform data format for data analysis, feature extraction to reduce data size, and unsupervised learning to detect the nodes acting differently from others. Moreover, we compare two techniques, principal component analysis (PCA) and independent component analysis (ICA), for feature extraction. We evaluate our prototype implementation by injecting a variety of faults into a production system at NCSA. The results show that our mechanism, in particular, the one using ICA-based feature extraction, can effectively identify faulty nodes with high accuracy and low computation overhead. Zhiling Lan, Ziming Zheng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2009 | Performance under Failures of DAG-based Parallel ComputingabstractAs the scale and complexity of parallel systems continue to grow, failures become more and more an inevitable fact for solving large-scale applications. In this research, we present an analytical study to estimate execution time in the presence of failures of directed acyclic graph (DAG) based scientific applications and provide a guideline for performance optimization. The study is four fold. We first introduce a performance model to predict individual subtask computation time under failures. Next, a layered, iterative approach is adopted to transform a DAG into a layered DAG, which reflects full dependencies among all the subtasks. Then, the expected execution time under failures of the DAG is derived based on stochastic analysis. Unlike existing models, this newly proposed performance model provides both the variance and distribution. It is practical and can be put to real use. Finally, based on the model, performance optimization, weak point identification and enhancement are proposed. Intensive simulations with real system traces are conducted to verify the analytical findings. They show that the newly proposed model and weak point enhancement mechanism work well. Hui Jin 0001, Xian-He Sun, Ziming Zheng, Zhiling Lan |
CCGRID | 4 |
| 2009 | Fault-aware, utility-based job scheduling on Blue, Gene/P systemsabstractJob scheduling on large-scale systems is an increasingly complicated affair, with numerous factors influencing scheduling policy. Addressing these concerns results in sophisticated scheduling policies that can be difficult to reason about. In this paper, we present a general utility-based scheduling framework to balance various scheduling requirements and priorities. It enables system owners to customize scheduling policies under different circumstances without changing the scheduling code. We also develop a fault-aware job allocation strategy for Blue Gene/P systems to address the increasing concern of system failures. We demonstrate the effectiveness of these facilities by means of event-driven simulations with real job traces collected from the production Blue Gene/P system at Argonne National Laboratory. Wei Tang 0001, Zhiling Lan, Narayan Desai, Daniel Buettner |
CLUSTER | 2 |
| 2009 | Reliability-aware scalability models for high performance computingabstractScalability models are powerful analytical tools for evaluating and predicting the performance of parallel applications. Unfortunately, existing scalability models do not quantify failure impact and therefore cannot accurately account for application performance in the presence of failures. In this study, we extend two well-known models, namely Amdahl's law and Gustafson's law, by considering the impact of failures and the effect of fault tolerance techniques on applications. The derived reliability-aware models can be used to predict application scalability in failure-present environments and evaluate fault tolerance techniques. Trace-based simulations via real failure logs demonstrate that the newly developed models provide a better understanding of application performance and scalability in the presence of failures. Ziming Zheng, Zhiling Lan |
CLUSTER | 2 |
| 2009 | System log pre-processing to improve failure predictionabstractLog preprocessing, a process applied on the raw log before applying a predictive method, is of paramount importance to failure prediction and diagnosis. While existing filtering methods have demonstrated good compression rate, they fail to preserve important failure patterns that are crucial for failure analysis. To address the problem, in this paper we present a log preprocessing method. It consists of three integrated steps: (1) event categorization to uniformly classify system events and identify fatal events; (2) event filtering to remove temporal and spatial redundant records, while also preserving necessary failure patterns for failure analysis; (3) causality-related filtering to combine correlated events for filtering through apriori association rule mining. We demonstrate the effectiveness of our preprocessing method by using real failure logs collected from the Cray XT4 at ORNL and the Blue Gene/L system at SDSC. Experiments show that our method can preserve more failure patterns for failure analysis, thereby improving failure prediction by up to 174%. Ziming Zheng, Zhiling Lan, Al Geist |
DSN | 2 |
| 2009 | Fault-Aware Runtime Strategies for High-Performance ComputingabstractAs the scale of parallel systems continues to grow, fault management of these systems is becoming a critical challenge. While existing research mainly focuses on developing or improving fault tolerance techniques, a number of key issues remain open. In this paper, we propose runtime strategies for spare node allocation and job rescheduling in response to failure prediction. These strategies, together with failure predictor and fault tolerance techniques, construct a runtime system called FARS (Fault-Aware Runtime System). In particular, we propose a 0-1 knapsack model and demonstrate its flexibility and effectiveness for reallocating running jobs to avoid failures. Experiments, by means of synthetic data and real traces from production systems, show that FARS has the potential to significantly improve system productivity (i.e., performance and reliability). Zhiling Lan, Prashasta Gujrati, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | A fast restart mechanism for checkpoint/recovery protocols in networked environmentsabstractCheckpoint/recovery has been studied extensively, and various optimization techniques have been presented for its improvement. Regardless of the considerable research efforts, little work has been done on improving its restart latency. The time spent on retrieving and loading the checkpoint image during a recovery is non-trivial, especially in networked environments. With the ever-increasing application memory footprint and system failure rate, it is becoming more of an issue. In this paper, we present a Fast REstart Mechanism called FREM. It allows fast restart of a failed process without requiring the availability of the entire checkpoint image. By dynamically tracking the process data accesses after each checkpoint, FREM masks restart latency by overlapping the computation of the resumed process with the retrieval of its checkpoint image. We have implemented FREM with the BLCR checkpointing tool in Linux systems. Our experiments with the SPEC benchmarks indicate that it can effectively reduce restart latency by 61.96% on average in networked environments. Zhiling Lan |
DSN | 2 |
| 2008 | Dynamic Meta-Learning for Failure Prediction in Large-Scale Systems: A Case StudyabstractDespite great efforts on the design of ultra-reliable components, the increase of system size and complexity has outpaced the improvement of component reliability. As a result, fault management becomes crucial in high performance computing. The advance of fault management relies on effective failure prediction. Despite years of research on failure prediction, it remains an open problem, especially in large-scale systems. In this paper, we address the problem by presenting a dynamic meta-learning prediction engine. It extends our previous work by exploring dynamic training, testing and prediction. Here, the "dynamic" part is from two perspectives: one is to continuously increase the training set during the system operation; and the other is to dynamically modify the rules of failure patterns by tracing prediction accuracy at runtime. Our case study indicates that the proposed predictor is promising by being capable of capturing more than 70% of failures, with the false alarm rate less than 10%. Jiexing Gu, Ziming Zheng, Zhiling Lan, Eva Hocks |
ICPP | 3 |
| 2008 | Enhancing application robustness through adaptive fault toleranceabstractAs the scale of high performance computing (HPC) continues to grow, application fault resilience becomes crucial. To address this problem, we are working on the design of an adaptive fault tolerance system for HPC applications. It aims to enable parallel applications to avoid anticipated failures via preventive migration, and in the case of unforeseeable failures, to minimize their impact through selective checkpointing. Both prior and ongoing work are summarized in this paper. Zhiling Lan, Ziming Zheng, Prashasta Gujrati |
IPDPS | 1 |
| 2008 | Adaptive Fault Management of Parallel Applications for High-Performance ComputingabstractAs the scale of high performance computing (HPC) grows, application fault resilience becomes increasingly important. In this paper, we propose FT-Pro, an adaptive fault management approach that combines the merits of reactive checkpointing and proactive migration. It enables parallel applications to avoid anticipated failures via preventive migration, and in the case of unforeseeable failures, to minimize their impact through selective checkpointing. An adaptation manager is designed for making runtime decision in response to failure prediction. We evaluate FT-Pro through stochastic modeling and case studies with real applications under a wide range of settings. Preliminary results indicate that FT-Pro outperforms periodic checkpointing, in terms of both reducing application completion times and improving resource utilization, by up to 43%. Zhiling Lan |
IEEE Trans. Computers | 1 |
| 2007 | Anomaly localization in large-scale clustersabstractA critical problem facing by managing large-scale clusters is to identify the location of problems in a system in case of unusual events. As the scale of high performance computing (HPC) grows, systems are getting bigger. When a system fails to function properly, health-related data are collected for troubleshooting. However, due to the massive quantities of information obtained from a large number of components, the root causes of anomalies are often buried like needles in a haystack. In this paper, we present a localization method to automatically find out the potential root causes (i.e. a subset of nodes) of the problem from the overwhelming amount of data collected system-wide. System managers can focus on examining these potential locations, thereby significantly reducing human efforts required for anomaly localization. Our method consists of three interrelated steps: (1) feature collection to assemble a feature space for the system; (2) feature extraction to obtain the most significant features for efficient data analysis by applying the principal component analysis (PCA) algorithm; and (3) outlier detection to quickly identify the nodes that are ldquofar awayrdquo from the majority by using the cell-based detection algorithm. Preliminary studies are presented to demonstrate the potential of our method for localizing anomalies in a computing environment where the nodes perform comparable tasks. Ziming Zheng, Zhiling Lan |
CLUSTER | 3 |
| 2007 | A Meta-Learning Failure Predictor for Blue Gene/L SystemsabstractThe demand for more computational power in science and engineering has spurred the design and deployment of ever-growing cluster systems. Even though the individual components used in these systems are highly reliable, the presence of large number of components inevitably increases the failure probability of such systems. Successful prediction of potential failures can greatly enhance various fault tolerance mechanisms used in large clusters, thereby mitigating the adverse impact of failures on system productivity and total cost of ownership. In this paper, we present a three-phase failure predictor to automatically process RAS events and further discover failure patterns for prediction in Blue Gene/L systems. In particular, this paper explores the use of meta- learning to adoptively integrate base methods with the goal to boost prediction accuracy. Experiments with two RAS logs collected from Blue Gene/L systems at ANL and SDSC demonstrate the effectiveness of the proposed failure predictor. Prashasta Gujrati, Zhiling Lan, Rajeev Thakur |
ICPP | 3 |
| 2007 | Fault-Driven Re-Scheduling For Improving System-level Fault ResilienceabstractThe productivity of HPC system is determined not only by their performance, but also by their reliability. The conventional method to limit the impact of failures is checkpointing. However, existing research shows that such a reactive fault tolerance approach can only improve system productivity marginally. Leveraging the recent progress made in the field of failure prediction, we propose fault-driven rescheduling (FARS) to improve system resilience to failures, and investigate the feasibility and effectiveness of utilizing failure prediction to dynamically adjust the placement of active jobs (e.g. running jobs) in response to failure prediction. In particular, a rescheduling algorithm is designed to enable effective job adjustment by evaluating performance impact of potential failures and rescheduling on user jobs. The proposed FARS complements existing research on fault-aware scheduling by allowing user jobs to avoid imminent failures at runtime. We evaluate FARS by using actual workloads and failure events collected from production HPC systems. Our preliminary results show the potential of FARS on improving system resilience to failures. Prashasta Gujrati, Zhiling Lan, Xian-He Sun |
ICPP | 3 |
| 2006 | Evaluating Performance and Scalability of Advanced Accelerator SimulationsabstractAdvanced accelerator simulations have played a prominent role in the design and analysis of modern accelerators. Given that accelerator simulations are computational intensive and various high-end clusters are available for such simulations, it is imperative to study the performance and scalability of accelerator simulations on different production systems. In this paper, we examine the performance and scaling behavior of a DOE SciDAC funded accelerator simulation package called Synergia on three different TOP500 clusters including an IA32 Linux Cluster, an IA64 Linux Cluster, and a SGI Altix 3700 system. The main objective is to understand the impact of different high-end architectures and message layers on the performance of accelerator simulations. Experiments show that IA32 using single-CPU can provide the best performance and scalability, while SGI Altix and IA32 using dual-CPU do not scale past 128 and 256 CPUs respectively. Our analysis also indicates that the existing accelerator simulations have several performance bottlenecks which prevent accelerator simulations to take full advantage of the capabilities provided by teraflop and beyond systems. Zhiling Lan, James F. Amundson, Panagiotis Spentzouris |
CCGRID | 2 |
| 2006 | Exploit Failure Prediction for Adaptive Fault-Tolerance in Cluster ComputingabstractAs the scale of cluster computing grows, it is becoming hard for long-running applications to complete without facing failures on large-scale clusters. To address this issue, checkpointing/restart is widely used to provide the basic fault-tolerant functionality, yet it suffers from high overhead and its reactive characteristic. In this work, we propose FT-Pro, an adaptive fault management mechanism that optimally chooses migration, checkpointing or no action to reduce the application execution time in the presence of failures based on the failure prediction. A cost-based evaluation model is presented for dynamic decision at run-time. Using the actual failure log from a production cluster at NCSA, we demonstrate that even with modest failure prediction accuracy, FT-Pro outperforms the traditional checkpointing/restart strategy by 13%-30% in terms of reducing the application execution time despite failures, which is a significant performance improvement for long-running applications. Zhiling Lan |
CCGRID | 2 |
| 2006 | Poster reception - Improving fault resilience of high performance applicationsabstractFor large-scale systems with hundreds to thousands of nodes, failures are likely to be more frequent as the system reliability decreases exponentially with the increasing count of components. Many parallel applications that span a large number of nodes are designed to run for days or weeks until completion. Hence, application-level fault resilience is of critical importance to the continued scaling of high performance computing (HPC). In this poster, we present and evaluate an adaptive fault resilience framework for HPC applications which adaptively selects an optimal corrective or preventive action based upon failure predictions at runtime. The proposed framework is implemented with a production-level MPI package and assessed with a variety of real-world parallel applications on production HPC systems. The experiment results demonstrate promising performance improvement of FT-Pro against traditional checkpointing/recovery schemes under a wide range of prediction accuracies and application characteristics. Zhiling Lan |
SC | 2 |
| 2006 | DistDLB: Improving cosmology SAMR simulations on distributed computing systems through hierarchical load balancing
Zhiling Lan, Valerie Taylor 0001 |
J. Parallel Distributed Comput. | 1 |
| 2005 | A novel workload migration scheme for heterogeneous distributed computingabstractDynamically partitioning of adaptive applications and migration of excess workload from overloaded processors to underloaded processors during execution are critical techniques needed for distributed computing. Distributed systems differ from traditional parallel systems in that they consist of heterogeneous resources connected with shared networks, thereby preventing existing schemes from benefiting large-scale applications. In particular, the cost entailed by workload migration is significant when the excess workload is transferred across heterogeneous distributed platforms. This paper introduces a novel distributed data migration scheme for large-scale adaptive applications. The major contributions of the paper include: (1) a novel hierarchical data migration scheme is proposed by considering the heterogeneous and dynamic features of distributed computing environments; and (2) a linear programming algorithm is presented to effectively reduce the overhead entailed in migrating excess workload across heterogeneous distributed platforms. Experiment results show that the proposed migration scheme outperforms common-used schemes with respect to reducing the communication cost and the application execution time. Zhiling Lan |
CCGRID | 2 |
| 2003 | Performance Analysis of a Large-Scale Cosmology Application on Three Cluster SystemsabstractA typical cosmological simulation requires a large amount of compute power, which is hard to satisfy with a single machine. Cluster systems provide the opportunity to execute such large-scale applications. In this paper, we investigate and analyze the performance of a large-scale production cosmology application, the ENZO code, on different cluster environments. Three cluster systems, each of them representing a widely-used cluster environment in the area of scientific computing, are used in this work: an IBM SP2 system at SDSC, an IA-64 Linux cluster at NCSA, and a SUN cluster at IIT. The performance is evaluated from three aspects: overall performance, communication characteristics, and load balancing characteristics. The experimental data shows that the cosmology performance on these clusters depends on the system performance and the application characteristics. The application performance on these clusters does not totally match the NPB measurement. Further, it seems that the IA-64 Linux cluster does not scale past 32 CPUs for this application. Zhiling Lan, Prathibha Deshikachar |
CLUSTER | 1 |
| 2003 | Exploring cosmology applications on distributed environments
Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
Future Gener. Comput. Syst. | 1 |
| 2002 | A novel dynamic load balancing scheme for parallel systems
Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
J. Parallel Distributed Comput. | 1 |
| 2001 | Dynamic Load Balancing for Structured Adaptive Mesh Refinement ApplicationsabstractAdaptive Mesh Refinement (AMR) is a type of multiscale algorithm that achieves high resolution in localized regions of dynamic, multidimensional numerical simulations. One of the key issues related to AMR is dynamic load balancing (DLB), which allows large-scale adaptive applications to run efficiently on parallel systems. In this paper we present an efficient DLB scheme for structured AMR (SAMR) applications. Our DLB scheme combines a grid-splitting technique with direct grid movements (e.g., direct movement from an overloaded processor to an underloaded proces sor), for which the objective is to efficiently redistribute workload among all the processors so as to reduce the parallel execution time. The potential benefits of our DLB scheme are examined by incorporating our techniques into a parallel, cosmological application that uses SAMR techniques. Experiments show that by using our scheme, the parallel execution time can be reduced by up to 47% and the quality of load-balancing can be improved by a factor of four. Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
ICPP | 1 |
| 2001 | Dynamic load balancing of SAMR applications on distributed systems
Zhiling Lan, Valerie Taylor 0001, Greg Bryan |
SC | 1 |
| 2000 | Prophesy: An Infrastructure for Analyzing and Modeling the Performance of Parallel and Distributed ApplicationsabstractEfficient execution of applications requires insight into how the system features impact the performance of the application. For distributed systems, the task of gaining this insight is complicated by the complexity of the system features. This insight generally results from significant experimental analysis and possibly the development of performance models. This paper presents the Prophesy project, an infrastructure that aids in gaining this needed insight based upon experience. The core component of Prophesy is a relational database that allows for the recording of performance data, system features and application details. Xingfu Wu, Valerie Taylor 0001, Jonathan Geisler, Zhiling Lan, Rick L. Stevens, Mark Hereld, Ivan R. Judson |
HPDC | 5 |