Xiaojun Ruan

dblp:60/6397 · DBLP profile ↗
← Back
43ranked-venue papers
11as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 5 first-author · 3 since 2021Computer networks · 14 · 4 first-authorSecurity and privacy · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Thermal-aware Energy-efficient Workload Scheduling for Data Centers
abstract
Given the emerging needs of artificial intelligence in both commercial and consumer uses, the rapid increase of energy consumption has become a major concern for data centers. Workload scheduling plays a crucial role in reducing the energy costs associated with large compute-intensive workloads in data centers. However, there has not been much research on reducing the energy costs for GPU-intensive workloads specifically. In this paper, we proposed two thermal-ware scheduling algorithms (GPUThermalAwareMinTemp and GPUThermalAware) for GPU-intensive workloads. Through the integration of machine learning models into a modified version of the GPUCloudSimPlus simulator, we compare the performance of our algorithms with three baseline algorithms by simulating cloud computing environments using a subset of real-world workload traces from Alibaba PAI. Our experimental results show that, compared with the worst-case algorithms, the GPU thermal-aware scheduling algorithms achieved 4.72% lower average host minimum temperatures and saved total energy by 1.17%.
Josh Penaojas, Kousha Salimkhan, Mark Fullton, Henry Locke, Chongye Wang, Xiaojun Ruan, Mahdi Ebrihimi, Xunfei Jiang
CCNC7
2025 Reducing Exposure Bias in Deep Learning Models for GPU Temperature Prediction in Data Centers
abstract
Dynamic thermal-aware load balancing addresses rising data center cooling costs by relying on accurate GPU temperature prediction. Due to the complexity of GPU temperature, deep learning models are widely used, but exposure bias–differences between training and inference environments–often degrades performance in autoregressive settings. While methods like scheduled sampling and attention forcing have been proposed to mitigate exposure bias, their effectiveness for hardware temperature forecasting is not well understood. This study systematically evaluates these mitigation techniques along with patching and modern deep learning architectures for GPU temperature prediction, considering accuracy, robustness, speed, and energy efficiency. Results show that scheduled sampling significantly worsens autoregressive RMSE and exposure bias (p < 0.001) for CNN-BiLSTM and Transformer models, whereas scheduled attention forcing reduces exposure bias but increases RMSE. The classic Transformer achieved the best overall balance with an AR RMSE of 4.34, median exposure bias of 2.4, median inference time of 257 µs, and the lowest training power usage (0.42 kWh), resulting in 22–84% lower carbon emissions. Patch-based models like PatchTST excelled in non-autoregressive settings but only the smallest patch sizes approached the Transformer’s autoregressive robustness at higher computational cost.
Mark Fullton, Xiaojun Ruan, Xunfei Jiang, Mahdi Ebrahimi
ICMLA2
2025 Optimizing parallel I/O performance in NVMe SSDs by Dynamic cache partitioning
Shu Yin 0001, Xiaojun Ruan
Perform. Evaluation3
2024 Dynamic Cache Partitioning for Enhancing Parallel I/O Performance in NVMe SSDs
abstract
Solid State Drive cache, implemented as on-board shared DRAM memory, can significantly enhance 110 performance by caching frequently accessed data. Although SSD caching strategies for single 110 data flows have been extensively explored, studies on cache partitioning to optimize parallel 110 in an SSD are scarce. In this paper, we present a novel dynamic cache partitioning approach designed to improve overall performance of multi-parallel 110 data flows by minimizing per-formance degradation of cache pollution and resource contention. By dynamically adjusting cache partition sizes for each data flow by considering cache sensitivity on performance, our strategy seeks to determine the optimal cache partition sizes to maximize overall 110 throughput. We implemented the strategy in the SSD simulator MQSim and evaluated its performance using various synthetic and real-world workloads. Our experimental results indicate that our dynamic cache partitioning strategy achieves an overall throughput increase of up to 33.22 % compared to shared cache methods and outperforms static cache partitioning strategies by up to 21.19%.
Guancheng Li, Songhui Cao, Shu Yin 0001, Xiaojun Ruan
NAS5
2021 Cached Mapping Table Prefetching for Random Reads in Solid-State Drives
abstract
Data caching strategies and Garbage Collection on SSDs have been extensively explored in the past years. However, the Mapping Table cache performance has not been well studied. Mapping table provides page translation information to Flash Translation Layer (FTL) in order to translate Logical Page Address (LPA) to Physical Page Address (PPA). Missing in mapping table cache causes extra read transactions to flash storage which results in stalls of I/O requests processing in SSDs. Random read requests are affected more than random write requests since write requests can be handled by write cache effectively. In this paper, we analyze the impact of CMT on different random read requests and present a Cached Mapping Table prefetching approach which fetches logical-to-physical page translation information in order to mitigate the stalls in processing random read requests. Our experimental results show an improvement of average request waiting time by up to 13%.
Xiaojun Ruan, Xunfei Jiang, Haiquan Chen 0001
NAS1
2019 Virtual machine allocation and migration based on performance-to-power ratio in energy-efficient clouds
Xiaojun Ruan, Haiquan Chen 0001, Yun Tian 0004, Shu Yin 0001
Future Gener. Comput. Syst.1
2018 DuoFS: A Hybrid Storage System Balancing Energy-Efficiency, Reliability, and Performance
abstract
As the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centers while retaining acceptable systems reliability. We propose a reliable energy-efficient storage system called DuoFS, which aims at balancing the energy efficiency, the reliability and the performance of parallel storage systems by seamlessly integrating one HDD-based file system and one SSD-based file system. At the heart of the DuoFS is a transformative middleware layer that dispatches files to the one of the two independent parallel file systems based on the files' I/O access popularity. By replicating popular files to the SSD-based file system and pushing the HDD-based file system into the low-power mode under light workload conditions, DuoFS can reduce significant energy consumption, avoid major factors that harm the storage systems reliability, and extract SSDs good I/O performance. Experimental results show that the DuoFS system saves up to 40% of energy, achieves up to 50% better I/O performance while only sacrificing less than 15% of the system's reliability.
Shu Yin 0001, Bing Jiao, Xiaomin Zhu 0001, Xiaojun Ruan, Si Chen 0009, Zhuo Tang
PDP4
2017 DuoFS: An Attempt at Energy-Saving and Retaining Reliability of Storage Systems
abstract
As issues of the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centres while retaining acceptable systems reliability. Most energy conservation techniques inevitably have adverse impacts on the parallel disk systems. To address the reliability issues of energy-efficient parallel storage systems, we propose a reliable energy-efficient storage system called DuoFS, which aims at improving both energy efficiency and reliability of parallel storage systems by seamlessly integrating HDDs and SSDs. With the help of the middleware layer, DuoFS can distribute popular data to SSD-based nodes and put HDD-based nodes into the low-power mode under light workload conditions without modification of the parallel systems.
Bing Jiao, Xiaomin Zhu 0001, Xiaojun Ruan, Xiao Qin 0001, Shu Yin 0001
ICDCS3
2016 RESS: A Reliable Energy-Efficient Storage System
abstract
Extracting high I/O performance from parallel file systems is no longer the only goal in modern data centres. As issues of the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centres while retaining acceptable systems reliability. Most energy conservation techniques inevitably have adverse impacts on the parallel disk systems. To address the reliability issues of energy-efficient parallel storage systems, we propose a reliable energy-efficient storage system called RESS, which aims at improving both energy efficiency and reliability of parallel storage systems by seamlessly integrating HDDs and SSDs. At the heart of the RESS is a transformative middleware layer, which reorganizes the I/O workload for the underlying parallel file systems. With the help of the middleware layer, RESS can distribute popular data to SSDs and put HDDs into the low-power mode under light workload conditions without modification of the parallel systems.
Shu Yin 0001, Zhaoyu Xiao, Kenli Li 0001, Jianzhong Huang 0001, Xiaojun Ruan, Xiaomin Zhu 0001, Xiao Qin 0001
ICPADS5
2016 Location-preserved contention-based routing in vehicular ad hoc networks
abstract
Abstract Location privacy protection in vehicular ad hoc networks considers preserving two types of information: the locations and identifications of users. However, existing solutions, which either replace identifications by pseudonyms or hide locations in areas, cannot be directly applied to geographic routing protocols because they degrade network performance. To address this issue, we proposed a location‐preserved contention (LPC) based routing protocol, in which greedy forwarding is achieved using dummy distance to the destination information instead of users’ true locations. Unlike the contention‐based forwarding protocol, the number of duplicated responses in LPC can be reduced by adjusting the parameterα, which is a timer scaling factor. To quantify the efficiency of location privacy protection, an entropy‐based analytical method is proposed. LPC is compared with existing routing and location privacy protection protocols in simulations. Results show that LPC provides 11.7% better network performance and a higher level of location privacy protection than the second best protocol. Copyright © 2014 John Wiley & Sons, Ltd.
Qing Yang 0003, Alvin S. Lim, Xiaojun Ruan, Xiao Qin 0001
Secur. Commun. Networks3
2015 Performance-to-Power Ratio Aware Virtual Machine (VM) Allocation in Energy-Efficient Clouds
abstract
The last decade witnesses a dramatic advance of cloud computing research and techniques. One of the key faced challenges in this field is how to reduce the massive amount of energy consumption in cloud computing data centers. To address this issue, many power-aware virtual machine (VM) allocation and consolidation approaches are proposed to reduce energy consumption efficiently. However, most of those existing efficient cloud solutions save energy cost at a price of the significant performance degradation. In this paper, we present a novel VM allocation algorithm called "PPRGear", which leverages the Performance-to-Power ratios for various host types. By achieving the optimal balance between host utilization and energy consumption, PPRGear is able to guarantee that host computers run at the most power-efficient levels (i.e., the levels with highest Performance-to-Power ratios) so that the energy consumption can be tremendously reduced with little sacrifice of performance. Our extensive experiments with real world traces show that compared with three baseline energy-efficient VM allocation and selection algorithms, PPRGear is able to reduce the energy consumption up to 69.31% for various host computer types with fewer migration and shutdown times and little performance degradation for cloud computing data centers.
Xiaojun Ruan, Haiquan Chen 0001
CLUSTER1
2015 REED: A Reliable Energy-Efficient RAID
abstract
Recent studies indicate that the energy cost and carbon footprint of data centers have become exorbitant. It is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centers. Most energy conservation techniques inevitably have adverse impacts on parallel disk systems. To address the reliability issues of energy-efficient parallel disks, we propose a reliable energy-efficient RAID system called REED, which aims at improving both energy efficiency and reliability of RAID systems by seamlessly integrating HDDs and SSDs. At the heart of REED is a high-performance cache mechanism powered by SSDs, which are serving popular data. Under light workload conditions, REED spins down HDDs into the low-power mode, thereby offering energy conservation. Importantly, during an I/O access turbulence (i.e., I/O load is dynamically and frequently changing), REED is conducive to reducing the number of disk power-state transitions by keeping HDDs in the low-power mode while serving requests with SSDs. We build a model to quantitatively show that REED is capable of improving the reliability of energy-efficient RAIDs. We implement the REED prototype in a real-world RAID-0 system. Our experimental results demonstrate that REED improves the energy-efficiency of conventional RAID-0 by up to 73% while maintaining good reliability.
Shu Yin 0001, Xuewu Li, Kenli Li 0001, Jianzhong Huang 0001, Xiaojun Ruan, Xiaomin Zhu 0001, Wei Cao 0006, Xiao Qin 0001
ICPP5
2014 MINT: A Reliability Modeling Frameworkfor Energy-Efficient Parallel Disk Systems
abstract
The Popular Disk Concentration (PDC) technique and the Massive Array of Idle Disks (MAID) technique are two effective energy conservation schemes for parallel disk systems. The goal of PDC and MAID is to skew I/O load toward a few disks so that other disks can be transitioned to low power states to conserve energy. I/O load skewing techniques like PDC and MAID inherently affect reliability of parallel disks, because disks storing popular data tend to have high failure rates than disks storing cold data. To study reliability impacts of energy-saving techniques on parallel disk systems, we develop a mathematical modeling framework called MINT. We first model the behaviors of parallel disks coupled with power management optimization policies. We make use of data access patterns as input parameters to estimate each disk's utilization and power-state transitions. Then, we derive each disk's reliability in terms of annual failure rate from the disk's utilization, age, operating temperature, and power-state transition frequency. Next, we calculate the reliability of PDC and MAID parallel disk systems in accordance with the annual failure rate of each disk in the systems. Finally, we use real-world trace to validate out MINT model. Validation result shows that the behaviors of PDC and MAID which are modeled by MINT have a similar trend as that in the real-world.
Shu Yin 0001, Xiaojun Ruan, Adam Manzanares, Xiao Qin 0001, Kenli Li 0001
IEEE Trans. Dependable Secur. Comput.2
2012 Multicore-Enabled Smart Storage for Clusters
abstract
We present a multicore-enabled smart storage for clusters in general and MapReduce clusters in particular. The goal of this research is to improve performance of data-intensive parallel applications on clusters by offloading data processing to multicore processors in storage nodes. Compared with traditional storage devices, next-generation disks will have computing capability to reduce computational load of host processors or CPUs. With the advance of processor and memory technologies, smart storage systems are promising devices to perform complex on-disk operations. The proposed smart storage system can avoid moving a huge amount of data back and forth between storage nodes and computing nodes in a cluster. To enhance the performance of data-intensive applications, we have designed a smart storage system called Multicore-enabled Smart Storage (McSD), in which a multicore processor is integrated in storage nodes. We have implemented a programming framework for data-intensive applications running on a computing system coupled with McSD. The programming framework aims at balancing load between computing nodes and multicore-enabled smart storage nodes. To fully utilize multicore processors in smart storage nodes, we have implemented the MapReduce model for McSDs to handle parallel computing on a cluster. A prototype of McSD has been implemented in a cluster connected by Gigabit Ethernet. Experimental results show that McSD can significantly reduce the execution times of three real-world applications - word count, string matching, and matrix multiplication. We demonstrate that the integration of multicore-enabled smart storage with MapReduce clusters is a promising approach to improving overall performance of data-intensive applications on clusters.
Zhiyang Ding, Xunfei Jiang, Shu Yin 0001, Xiao Qin 0001, Kai-Hsiung Chang, Xiaojun Ruan, Mohammed I. Alghamdi, Meikang Qiu
CLUSTER6
2012 Thermal modeling and analysis of storage systems
abstract
Recognizing that power and cooling cost for data centers are increasing, we address in this study the thermal impact of storage systems. In the first phase of this work, we generate the thermal profile of a storage server containing three hard disks. The profiling results show that disks have comparable thermal impacts as processing and networking elements to overall storage node temperature. We develop a thermal model to estimate the outlet temperature of a storage server based on processor and disk utilizations. The thermal model is validated against data acquired by an infrared thermometer as well as build-in temperature sensors on disks. Next, we apply the thermal model to investigate the thermal impact of workload management on storage systems. Our study suggests that disk-aware thermal management techniques have significant impacts on reducing cooling cost of storage systems. We further show that this work can be extended to analysis the cooling cost of data centers with massive storage capacity.
Xunfei Jiang, Mohammed I. Alghamdi, Ji Zhang 0002, Maen M. Al Assaf, Xiaojun Ruan, Tausif Muzaffar, Xiao Qin 0001
IPCCC5
2012 Global workload characterization of a large scale satellite image distribution system
abstract
Online content distribution systems, which store incredibly large amounts of information and provide service to large numbers of users, are becoming increasingly commonplace. To fulfill the wide range of requests sent by different users, these systems must ensure efficient handling of massive amount of data. To achieve this goal, the in-depth analysis and comprehensive understanding of user behaviors are critical. However, analyzing the behaviors of worldwide users with different needs is a very challenging task. This is especially true when historical user behaviors evolve over time or may be affected by unpredictable events. In this paper, we present a number of workload characterization techniques applied to one of the world's largest online satellite image distribution systems operated by the U.S. Geological Survey (USGS) and NASA.
Brian Romoser, Ribel Fares, Peter Janovics, Xiaojun Ruan, Xiao Qin 0001, Ziliang Zong
IPCCC4
2012 Improving write performance by enhancing internal parallelism of Solid State Drives
abstract
Most researches of Solid State Drives (SSDs) architectures rely on Flash Translation Layer (FTL) algorithms and wear-leveling; however, internal parallelism in Solid State Drives has not been well explored. In this research, we proposed a new strategy to improve SSD write performance by enhancing internal parallelism inside SSDs. A SDRAM buffer is added in the design for buffering and scheduling write requests. Because the same logical block numbers may be translated to different physical numbers at different times in FTL, the on-board SDRAM buffer is used to buffer requests at the lower level of FTL. When the buffer is full, same amount of data will be assigned to each storage package in SSDs to enhance internal parallelism. To accurately evaluate performance, we use both synthetic workloads and real-world applications in experiments. We compare the enhanced internal parallelism scheme with the traditional LRU strategy since it is unfair to compare an SSD having buffer with an SSD without a buffer. The simulation results demonstrate that the writing performance of our design is significantly improved compared with the LRU-cache strategy with the same amount of buffer sizes.
Xiaojun Ruan, Ziliang Zong, Mohammed I. Alghamdi, Yun Tian 0004, Xunfei Jiang, Xiao Qin 0001
IPCCC1
2012 ES-MPICH2: A Message Passing Interface with Enhanced Security
abstract
An increasing number of commodity clusters are connected to each other by public networks, which have become a potential threat to security sensitive parallel applications running on the clusters. To address this security issue, we developed a Message Passing Interface (MPI) implementation to preserve confidentiality of messages communicated among nodes of clusters in an unsecured network. We focus on M PI rather than other protocols, because M PI is one of the most popular communication protocols for parallel computing on clusters. Our MPI implementation-called ES-MPICH2-was built based on MPICH2 developed by the Argonne National Laboratory. Like MPICH2, ES-MPICH2 aims at supporting a large variety of computation and communication platforms like commodity clusters and high-speed networks. We integrated encryption and decryption algorithms into the MPICH2 library with the standard MPI interface and; thus, data confidentiality of MPI applications can be readily preserved without a need to change the source codes of the MPI applications. MPI-application programmers can fully configure any confidentiality services in MPICHI2, because a secured configuration file in ES-MPICH2 offers the programmers flexibility in choosing any cryptographic schemes and keys seamlessly incorporated in ES-MPICH2. We used the Sandia Micro Benchmark and Intel MPI Benchmark suites to evaluate and compare the performance of ES-MPICH2 with the original MPICH2 version. Our experiments show that overhead incurred by the confidentiality services in ES-MPICH2 is marginal for small messages. The security overhead in ES-MPICH2 becomes more pronounced with larger messages. Our results also show that security overhead can be significantly reduced in ES-MPICH2 by high-performance clusters. The executable binaries and source code of the ES-MPICH2 implementation are freely available at http:// www.eng.auburn.edu/~xqin/software/es-mpich2/.
Xiaojun Ruan, Qing Yang 0003, Mohammed I. Alghamdi, Shu Yin 0001, Xiao Qin 0001
IEEE Trans. Dependable Secur. Comput.1
2011 Reliability analysis of an energy-aware RAID system
abstract
We develop a mathematical model - MREED - to quantitatively evaluate the failure rate of energy-efficient parallel storage systems. The Power-Aware Redundant Array of Inexpensive Disk (PARAID) aims to reduce energy use of commodity server-class disks without specialized hardware. The goal of PARAID is to skewed striping pattern to adapt to the system load by changing the number of powered disks. By spinning down disks during light workloads, PARAID can reduce power consumption, while still meeting performance demands. We show that MREED can be used to estimate a five-disk PARAID-0 system. We validate the accuracy of MREED using the DiskSim simulator. Our approach shows that MREED can rely on file access pattern to estimate system utilization correctly. Furthermore, even thought PARAID may achieve reasonable reliability, our model shows that PARAID's reliability is affected by data locality.
Shu Yin 0001, Yun Tian 0004, Jiong Xie, Xiao Qin 0001, Mohammed I. Alghamdi, Xiaojun Ruan, Meikang Qiu
IPCCC6
2011 Heat-based dynamic data caching: A load balancing strategy for energy-efficient parallel storage systems with buffer disks
abstract
Performance improvement and energy conservation are two conflicting objectives in large scale parallel storage systems. In this paper, we propose a novel solution to achieve the twin objectives of maximizing performance and minimizing energy consumption of parallel storage systems. Specifically, a buffer-disk based architecture (BUD for short) is designed to conserve energy. A heat-based dynamic data caching strategy is developed to improve performance. The BUD architecture strives to allocate as many requests as possible to buffer disks, thereby keeping a large number of idle data disks in low-power states. This can provide significant opportunities for energy conservation while making buffer disks a potential performance bottleneck. The heat-based data caching strategy aims to achieve good load balancing in buffer disks and alleviate overall performance degradation caused by unbalanced workload. Our experimental results have shown that the proposed BUD framework and dynamic data caching strategy are able to conserve energy by 84.4% for small reads and 78.8% for large reads with slightly degraded response time.
Ziliang Zong, Xiao Qin 0001, Xiaojun Ruan, Mais Nijim
MSST3
2011 EAD and PEBD: Two Energy-Aware Duplication Scheduling Algorithms for Parallel Tasks on Homogeneous Clusters
abstract
High-performance clusters have been widely deployed to solve challenging and rigorous scientific and engineering tasks. On one hand, high performance is certainly an important consideration in designing clusters to run parallel applications. On the other hand, the ever increasing energy cost requires us to effectively conserve energy in clusters. To achieve the goal of optimizing both performance and energy efficiency in clusters, in this paper, we propose two energy-efficient duplication-based scheduling algorithms-Energy-Aware Duplication (EAD) scheduling and Performance-Energy Balanced Duplication (PEBD) scheduling. Existing duplication-based scheduling algorithms replicate all possible tasks to shorten schedule length without reducing energy consumption caused by duplication. Our algorithms, in contrast, strive to balance schedule lengths and energy savings by judiciously replicating predecessors of a task if the duplication can aid in performance without degrading energy efficiency. To illustrate the effectiveness of EAD and PEBD, we compare them with a nonduplication algorithm, a traditional duplication-based algorithm, and the dynamic voltage scaling (DVS) algorithm. Extensive experimental results using both synthetic benchmarks and real-world applications demonstrate that our algorithms can effectively save energy with marginal performance degradation.
Ziliang Zong, Adam Manzanares, Xiaojun Ruan, Xiao Qin 0001
IEEE Trans. Computers3
2011 PRE-BUD: Prefetching for energy-efficient parallel I/O systems with buffer disks
abstract
A critical problem with parallel I/O systems is the fact that disks consume a significant amount of energy. To design economically attractive and environmentally friendly parallel I/O systems, we propose an energy-aware prefetching strategy (PRE-BUD) for parallel I/O systems with disk buffers. We introduce a new architecture that provides significant energy savings for parallel I/O systems using buffer disks while maintaining high performance. There are two buffer disk configurations: (1) adding an extra buffer disk to accommodate prefetched data, and (2) utilizing an existing disk as the buffer disk. PRE-BUD is not only able to reduce the number of power-state transitions, but also to increase the length and number of standby periods. As such, PRE-BUD conserves energy by keeping data disks in the standby state for increased periods of time. Compared with the first prefetching configuration, the second configuration lowers the capacity of the parallel disk system. However, the second configuration is more cost-effective and energy-efficient than the first one. Finally, we quantitatively compare PRE-BUD with both disk configurations against three existing strategies. Empirical results show that PRE-BUD is able to reduce energy dissipation in parallel disk systems by up to 50 percent when compared against a non-energy aware approach. Similarly, our strategy is capable of conserving up to 30 percent energy when compared to the dynamic power management technique.
Adam Manzanares, Xiao Qin 0001, Xiaojun Ruan, Shu Yin 0001
ACM Trans. Storage3
2011 A Message-Scheduling Scheme for Energy Conservation in Multimedia Wireless Systems
abstract
Reducing power consumption of wireless networks has become a major goal in designing modern multimedia wireless systems. In an effort to reduce power consumption, this paper addresses the issue of scheduling real-time messages in multimedia wireless networks subject to both timing and power constraints. A power-consumption model is introduced to calculate power-consumption rates in accordance with message-transmission rates. Next, a new message-scheduling scheme called Power-aware Real-time Message (PARM) is developed to generate message-transmission schedules that minimize power consumption of multimedia wireless-network interfaces and the probability of missing deadlines for real-time messages. With a power-aware scheduling policy in place, the proposed PARM scheme is very energy-efficient. Experimental results based on a wide variety of synthetic workloads and eight real-world applications show that PARM significantly reduces energy dissipation while maintaining low missed rates. PARM reduces power consumption of data transmissions by up to 99.4% (with an average of 86.7%) for synthetic network traffic and saves energy by up to 60.0% (with an average of 34.1%) in the eight real-world applications.
Xiaojun Ruan, Shu Yin 0001, Adam Manzanares, Mohammed I. Alghamdi, Xiao Qin 0001
IEEE Trans. Syst. Man Cybern. Part A1
2010 Location Privacy Protection in Contention Based Forwarding for VANETs
abstract
Compared to traditional wireless network routing protocols, geographic routing provides superior scalability and thus is widely used in vehicular ad hoc networks (VANETs). However, it requires every vehicle to broadcast its location information to its neighboring nodes, and this process will compromise user's location privacy. Existing solutions to this problem can be categorized into two groups: 1) hiding user's location or 2) preserving user's identification information in routing protocols, which drastically reduce network performances. To address this issue, we proposed a dummy-based location privacy protection (DBLPP) routing protocol, in which routing decision is made based upon the dummy distance to the destination (DOD), instead of users' true locations. In this scheme, users' true locations and identification information are preserved, so the user's location privacy is protected. Compared to existing solutions, simulation results show that while DBLPP provides similar network performances as other routing protocols, it achieves a higher level of location privacy protection on vehicles in networks.
Qing Yang 0003, Alvin S. Lim, Xiaojun Ruan, Xiao Qin 0001
GLOBECOM3
2010 Improving Energy Efficiency and Security for Disk Systems
abstract
Improving security and minimizing power consumption are crucial for large-scale data storage systems. Although a handful of studies have been focused on data security and energy efficiency, most of the existing approaches have concentrated on only one of these two metrics. In this paper, we present a new approach to integrating power optimization with security services to enhance the security of energy-efficient large-scale storage systems. In our approach, we make use of the dynamic speed control for power management technique, or DRPM, to conserve energy in secure storage systems. In this study we develop two ways of integrating confidentiality services with the dynamic disk speed control technique. The first strategy - security aggressive in nature - is focused on the improvement of storage system security with less emphasis on energy conservation. The second strategy gives higher priority to energy conservation as opposed to the security optimization. Our experimental results show that the energy-aggressive approach provides better energy savings than the security-aggressive approach. However, the quality of security achieved by the security-aggressive scheme is higher than that of the energy-aggressive approach. Moreover, the empirical results show that energy savings yielded by the two approaches become more pronounced when the data size is increased. The findings illustrate that the response time of the security-aggressive approach is more sensitive to data size than that of the energy-aggressive scheme.
Shu Yin 0001, Mohammed I. Alghamdi, Xiaojun Ruan, Mais Nijim, Ashwin Tamilarasan, Ziliang Zong, Xiao Qin 0001
HPCC3
2010 Energy Efficient Prefetching with Buffer Disks for Cluster File Systems
abstract
Energy efficient computing is becoming increasingly important as the scale of parallel computing systems is expanding. As the processing power of parallel computing systems has been incremented there has been an increased demand for large scale storage systems to store the output of these parallel computing systems. Data centers are growing at an enormous pace and it is important to investigate a means of managing the energy efficiency of large scale parallel storage systems. To address these issues we introduce EEVFS (Energy Efficient Virtual File System), which is able to manage data placement and disk states to help improve the energy efficiency of a parallel disk system. EEVFS places data on the storage disks in an energy efficient layout and attempts to predict when each disk will be idle for a large period of time, facilitating a state transition into the standby state. EEVFS should also maintain relatively high performance, so we have built a load balancing policy into the data partitioning of EEVFS. The implementation architecture and measured results are presented to demonstrate the energy efficiency and performance characteristics of EEVFS.
Adam Manzanares, Xiaojun Ruan, Shu Yin 0001, Jiong Xie, Zhiyang Ding, Yun Tian 0004, James Majors, Xiao Qin 0001
ICPP2
2010 An automatic prefetching and caching system
abstract
Steady improvements in storage capacities and CPU clock speeds intensify the performance bottleneck at the I/O subsystem of modern computers. Caching data can efficiently short circuit costly delays associated with disk accesses. Recent studies have shown that disk I/O performance gains provided by a cache buffer do not scale with cache size. Therefore, new algorithms have to be investigated to better utilize cache buffer space. Predictive prefetching and caching solutions have been shown to improve I/O performance in an efficient and scalable manner in simulation experiments. However, most predictive prefetching algorithms have not yet been implemented in real-world storage systems due to two main limitations: first, the existing prefetching solutions are unable to self regulate based on changing I/O workload; second, excessive number of unneeded blocks are prefetched. Combined, these drawbacks make predictive prefetching and caching a less attractive solution than the simple LRU management. To address these problems, in this paper we propose an automatic prefetching and caching system (or APACS for short), which mitigates all of these shortcomings through three unique techniques, namely: (1) dynamic cache partitioning, (2) prefetch pipelining, and (3) prefetch buffer management. APACS dynamically partitions the buffer cache memory, used for prefetched and cached blocks, by automatically changing buffer/cache sizes in accordance to global I/O performance. The adaptive partitioning scheme implemented in APACS optimizes cache hit ratios, which subsequently accelerates application execution speeds. Experimental results obtained from trace-driven simulations show that APACS outperforms the LRU cache management and existing prefetching algorithms by an average of over 50%.
Joshua Lewis, Mohammed I. Alghamdi, Maen M. Al Assaf, Xiaojun Ruan, Zhiyang Ding, Xiao Qin 0001
IPCCC4
2010 ES-MPICH2: A Message Passing Interface with enhanced security
abstract
In largely distributed clusters, computing nodes are geographically deployed in various computing sites. Information processed in a distributed cluster is shared among a group of distributed processes or users by virtue of messages passing protocols (e.g. message passing interface - MPI) running on the Internet. Because of the open accessible nature of the Internet, data encryption for these large-scale distributed clusters becomes a non-trivial and challenging problem. To address this issue, we enhanced the security of the MPI (Message Passing Interface) protocol by encrypting and decrypting messages sent and received among computing nodes. In this study we focused on MPI rather than other protocols because MPI is one of the most popular communication protocols for cluster computing environments. From among a variety of MPI implementations, we picked MPICH2 developed by the Argonne National Laboratory. The design goal of MPICH2 - a widely used MPI implementation - is to combine portability with high performance. We integrated encryption algorithms into the MPICH2 library so that data confidentiality of MPI applications could be readily preserved without a need to change the source codes of the MPI applications. since we provide a security enhanced MPI-library with the standard MPI interfact, data communications of a conventional MPI program can be secured without converting the program into the corresponding secure version. We used Sandia Micro Benchmark and Intel MPI Benchmarks to evaluate and compared the performance of original MPICH2 and Enhanced Security MPICH2. According to the performance evaluation, ES-MPICH2 provides secured Message Passing Interface by sacrificing reasonable system performance.
Xiaojun Ruan, Qing Yang 0003, Mohammed I. Alghamdi, Shu Yin 0001, Zhiyang Ding, Jiong Xie, Joshua Lewis, Xiao Qin 0001
IPCCC1
2010 Communication-Aware Load Balancing for Parallel Applications on Clusters
abstract
Cluster computing has emerged as a primary and cost-effective platform for running parallel applications, including communication-intensive applications that transfer a large amount of data among the nodes of a cluster via the interconnection network. Conventional load balancers have proven effective in increasing the utilization of CPU, memory, and disk I/O resources in a cluster. However, most of the existing load-balancing schemes ignore network resources, leaving an opportunity to improve the effective bandwidth of networks on clusters running parallel applications. For this reason, we propose a communication-aware load-balancing technique that is capable of improving the performance of communication-intensive applications by increasing the effective utilization of networks in cluster environments. To facilitate the proposed load-balancing scheme, we introduce a behavior model for parallel applications with large requirements of network, CPU, memory, and disk I/O resources. Our load-balancing scheme can make full use of this model to quickly and accurately determine the load induced by a variety of parallel applications. Simulation results generated from a diverse set of both synthetic bulk synchronous and real parallel applications on a cluster show that our scheme significantly improves the performance, in terms of slowdown and turn-around time, over existing schemes by up to 206 percent (with an average of 74 percent) and 235 percent (with an average of 82 percent), respectively.
Xiao Qin 0001, Hong Jiang 0001, Adam Manzanares, Xiaojun Ruan, Shu Yin 0001
IEEE Trans. Computers4
2010 Conserving energy in real-time storage systems with I/O burstiness
abstract
Energy conservation has become a critical problem for real-time embedded storage systems. Although a variety of approaches for reducing energy consumption have been extensively studied, energy conservation for real-time embedded storage systems is still an open problem. In this article, we propose an energy management strategy, I/O Burstiness for Energy Conservation (IBEC), exploiting the burstiness of real-time embedded storage systems applications. Our approach aims at combining the IBEC energy-management strategy with a Linux-based disk block-scheduling mechanism to conserve the energy of storage systems. Extensive experiments are conducted involving a number of synthetic disk traces as well as real-world data-intensive traces. To evaluate the energy efficiency of IBEC, we compare the performance of IBEC against three existing strategies, namely, PA-EDF, DP-EDF, and EDF. Compared with the alternative strategies, IBEC reduces the power consumption of real-time embedded disks system by up to 60%.
Adam Manzanares, Xiaojun Ruan, Shu Yin 0001, Xiao Qin 0001, Adam Roth, Mais Nijim
ACM Trans. Embed. Comput. Syst.2
2009 How reliable are parallel disk systems when energy-saving schemes are involved?
abstract
Many energy conservation techniques have been proposed to achieve high energy efficiency in disk systems. Unfortunately, growing evidence shows that energy-saving schemes in disk drives usually have negative impacts on storage systems. Existing reliability models are inadequate to estimate reliability of parallel disk systems equipped with energy conservation techniques. To solve this problem, we propose a mathematical model - called MINT - to evaluate the reliability of a parallel disk system where energy-saving mechanisms are implemented. In this paper, we focus on modeling the reliability impacts of two well-known energy-saving techniques - the Popular Disk Concentration technique (PDC) and the Massive Array of Idle Disks (MAID). We started this research by investigating how PDC and MAID affect the utilization and power-state transition frequency of each disk in a parallel disk system. We then model the annual failure rate of each disk as a function of the disk's utilization, power state transition frequency as well as operating temperature, because these parameters are key reliability-affecting factors in addition to disk ages. Next, the reliability of a parallel disk system can be derived from the annual failure rate of each disk in the parallel disk system. Finally, we used MINT to study the reliability of a parallel disk system equipped with the PDC and MAID techniques. Experimental results show that PDC is more reliable than MAID when disk workload is low. In contrast, the reliability of MAID is higher than that of PDC under relatively high I/O load.
Shu Yin 0001, Xiaojun Ruan, Adam Manzanares, Xiao Qin 0001
CLUSTER2
2009 HYBUD: An Energy-Efficient Architecture for Hybrid Parallel Disk Systems
abstract
In the past decade parallel disk systems have been highly scalable and able to alleviate the problem of disk I/O bottleneck, thereby being widely used to support a wide range of data-intensive applications. Optimizing energy consumption in parallel disk systems has strong impacts on the cost of backup power-generation and cooling equipment, because a significant fraction of the operation cost of data centres is incurred by energy consumption and cooling. Although flash memory is very energy-efficient compared to disk drives, flash memory is too expensive to use as a major component in large-scale storage systems. In other words, it is not a cost-effective way to make use of large flash memory to build energy-efficient storage systems. To address this problem, in this paper we proposed a hybrid disk architecture or HYBUD that integrates a non-volatile flash memory with buffer disks to build cost-effective and energy-efficient parallel disk systems. While the most popular data sets are cached in flash memory, the second most popular data sets can be stored and retrieved from buffer disks. HYBUD is energy efficient because flash memory coupled with buffer disks can serve a majority of incoming disk requests, thereby keeping a large number of other data disks in the low-power state for longer period times. Furthermore, HYBUD is cost-effective by the virtue of inexpensive buffer disks assisting flash memory to cache a huge amount of popular data. Experimental results demonstratively show that compared with two existing non-hybrid architectures, HYBUD provides significant energy savings for parallel disk systems in a very cost effective way.
Mais Nijim, Adam Manzanares, Xiaojun Ruan, Xiao Qin 0001
ICCCN3
2009 Performance Evaluation of Energy-Efficient Parallel I/O Systems with Write Buffer Disks
abstract
In the past decade, parallel disk systems have been developed to address the problem of I/O performance. A critical challenge with modern parallel I/O systems is that parallel disks consume a significant amount of energy in servers and high performance computers. To conserve energy consumption in parallel I/O systems, one can immediately spin down disks when disk are idle; however, spinning down disks might not be able to produce energy savings due to penalties of spinning operations. Unlike powering up CPUs, spinning down and up disks need physical movements. Therefore, energy savings provided by spinning down operations must offset energy penalties of the disk spinning operations. To substantially reduce the penalties incurred by disk spinning operations, we developed a novel approach to conserving energy of parallel I/O systems with write buffer disks, which are used to accumulate small writes using a log file system. Data sets buffered in the log file system can be transferred to target data disks in a batch way. Thus, buffer disks aim to serve a majority of incoming write requests, attempting to reduce the large number of disk spinning operations by keeping data disks in standby for long period times. Interestingly, the write buffer disks not only can achieve high energy efficiency in parallel I/O systems, but also can shorten response times of write requests. To evaluate the performance and energy efficiency of our parallel I/O systems with buffer disks, we implemented a prototype using a cluster storage system as a testbed. Experimental results show that under light and moderate I/O load, buffer disks can be employed to significantly reduce energy dissipation in parallel I/O systems without adverse impacts on I/O performance.
Xiaojun Ruan, Adam Manzanares, Shu Yin 0001, Ziliang Zong, Xiao Qin 0001
ICPP1
2009 ECOS: An energy-efficient cluster storage system
abstract
Cluster storage systems are essential building blocks for many high-end computing infrastructures. Although energy conservation techniques have been intensively studied in the context of clusters and disk arrays, improving energy efficiency of cluster storage systems remains an open issue. To address this problem, we describe in this paper an approach to implementing an energy-efficient cluster storage system or ECOS for short. ECOS relies on the architecture of cluster storage systems in which each I/O node manages multiple disks - one buffer disk and several data disks. Given an I/O node, the key idea behind ECOS is to redirect disk requests from data disks to the buffer disk. To balance I/O load among I/O nodes, ECOS might redirect requests from one I/O node into the others. Redirecting requests is a driving force of energy saving, and the reason is two-fold. First, ECOS makes an effort to keep buffer disks active while placing data disks into standby in a long time period to conserve energy. Second, ECOS reduces the number of disk spin downs/ups in I/O nodes. The idea of ECOS was implemented in a Linux cluster, where each I/O node contains one buffer disk and two data disks. Experimental results show that ECOS improves the energy efficiency of traditional cluster storage systems where buffer disks are not employed. Adding one extra buffer disk into each I/O node seemingly has negative impact on energy saving. Interestingly, our results indicate that ECOS equipped with extra buffer disks is more energy efficient than the same cluster storage system without the buffer disks. The implication of the experiments is that using existing data disks in I/O nodes to perform as buffer disks can achieve even higher energy efficiency.
Xiaojun Ruan, Shu Yin 0001, Adam Manzanares, Jiong Xie, Zhiyang Ding, James Majors, Xiao Qin 0001
IPCCC1
2009 Improving reliability of energy-efficient parallel storage systems by disk swapping
abstract
The Popular Disk Concentration (PDC) technique and the Massive Array of Idle Disks (MAID) technique are two effective energy saving schemes for parallel disk systems. The goal of PDC and MAID is to skew I/O load towards a few disks so that other disks can be transitioned to low power states to conserve energy. I/O load skewing techniques like PDC and MAID inherently affect reliability of parallel disks because disks storing popular data tend to have high failure rates than disks storing cold data. To achieve good tradeoffs between energy efficiency and disk reliability, we first present a reliability model to quantitatively study the reliability of energy-efficient parallel disk systems equipped with the PDC and MAID schemes. Then, we propose a novel strategy—disk swapping—to improve disk reliability by alternating disks storing hot data with disks holding cold data. We demonstrate that our disk-swapping strategies not only can increase the lifetime of cache disks in MAID-based parallel disk systems, but also can improve reliability of PDC-based parallel disk systems.
Shu Yin 0001, Xiaojun Ruan, Adam Manzanares, Zhiyang Ding, Jiong Xie, James Majors, Xiao Qin 0001
IPCCC2
2009 Can We Improve Energy Efficiency of Secure Disk Systems without Modifying Security Mechanisms?
abstract
Improving energy efficiency of security-aware storage systems is challenging, because security and energy efficiency are often two conflicting goals. The first step toward making the best tradeoffs between high security and energy efficiency is to profile encryption algorithms to decide if storage systems would be able to produce energy savings for security mechanisms. We are focused on encryption algorithms rather than other types of security services, because encryption algorithms are usually computation-intensive. In this study, we used the XySSL libraries and profiled operations of several test problems using Conky - a lightweight system monitor that is highly configurable. Using our profiling techniques we concluded that although 3DES is much slower than AES encryption,it more likely to save energy in security-aware storage systems using 3DES than AES. The CPU is the bottleneck in 3DES, allowing us to take advantage of dynamic power management schemes to conserve energy at the disk level.After profiling several hash functions, we noticed that the CPU is not the bottleneck for any of these functions,indicating that it is difficult to leverage the dynamic power management technique to conserve energy of a single disk where hash functions are implemented for integrity checking.
Xiaojun Ruan, Adam Manzanares, Shu Yin 0001, Mais Nijim, Xiao Qin 0001
NAS1
2009 Energy-Aware Prefetching for Parallel Disk Systems: Algorithms, Models, and Evaluation
abstract
Parallel disk systems consume a significant amount of energy due to the large number of disks. To design economically attractive and environmentally friendly parallel disk systems, in this paper we design and evaluate an energy-aware prefetching strategy for parallel disk systems consisting of a small number of buffer disks and large number of data disks. Using buffer disks to temporarily handle requests for data disks, we can keep data disks in the low-power mode as long as possible. Our prefetching algorithm aims to group many small idle periods in data disks to form large idle periods, which in turn allow data disks to remain in the standby state to save energy. To achieve this goal, we utilize buffer disks to aggressively fetch popular data from regular data disks into buffer disks, thereby putting data disks into the standby state for longer time intervals. A centrepiece in the prefetching mechanism is an energy-saving prediction model, based on which we implement the energy-saving calculation module that is invoked in the prefetching algorithm. We quantitatively compare our energy-aware prefetching mechanism against existing solutions, including the dynamic power management strategy. Experimental results confirm that the buffer-disk-based prefetching can significantly reduce energy consumption in parallel disk systems by up to 50 percent. In addition, we systematically investigate the energy efficiency impact that varying disk power parameters has on our prefetching algorithm.
Adam Manzanares, Xiaojun Ruan, Shu Yin 0001, Mais Nijim, Xiao Qin 0001
NCA2
2009 Dynamic load balancing for I/O-intensive applications on clusters
abstract
Load balancing for clusters has been investigated extensively, mainly focusing on the effective usage of global CPU and memory resources. However, previous CPU- or memory-centric load balancing schemes suffer significant performance drop under I/O-intensive workloads due to the imbalance of I/O load. To solve this problem, we propose two simple yet effective I/O-aware load-balancing schemes for two types of clusters: (1) homogeneous clusters where nodes are identical and (2) heterogeneous clusters, which are comprised of a variety of nodes with different performance characteristics in computing power, memory capacity, and disk speed. In addition to assigning I/O-intensive sequential and parallel jobs to nodes with light I/O loads, the proposed schemes judiciously take into account both CPU and memory load sharing in the system. Therefore, our schemes are able to maintain high performance for a wide spectrum of workloads. We develop analytic models to study mean slowdowns, task arrival, and transfer processes in system levels. Using a set of real I/O-intensive parallel applications and synthetic parallel jobs with various I/O characteristics, we show that our proposed schemes consistently improve the performance over existing non-I/O-aware load-balancing schemes, including CPU- and Memory-aware schemes and a PBS-like batch scheduler for parallel and sequential jobs, for a diverse set of workload conditions. Importantly, this performance improvement becomes much more pronounced when the applications are I/O-intensive. For example, the proposed approaches deliver 23.6--88.0 % performance improvements for I/O-intensive applications such as LU decomposition, Sparse Cholesky, Titan, Parallel text searching, and Data Mining. When I/O load is low or well balanced, the proposed schemes are capable of maintaining the same level of performance as the existing non-I/O-aware schemes.
Xiao Qin 0001, Hong Jiang 0001, Adam Manzanares, Xiaojun Ruan, Shu Yin 0001
ACM Trans. Storage4
2008 Improving reliability and energy efficiency of disk systems via utilization control
abstract
As disk drives become increasingly sophisticated and processing power increases, one of the most critical issues of designing modern disk systems is data reliability. Although numerous energy saving techniques are available for disk systems, most of energy conservation techniques are not effective in reliability critical environments due to their limitation of ignoring the reliability issue. A wide range of factors affect the reliability of disk systems; the most important factors - disk utilization and ages — are the focus of this study. We build a model to quantify the relationship among the disk age, utilization, and failure probabilities. Observing that the reliability of a disk heavily relies on both disk utilization and age, we propose a novel concept of safe utilization zone, where energy of the disk can be conserved without degrading reliability. We investigate an approach to improving both reliability and energy efficiency of disk systems via utilization control, where disk drives are operated in safe utilization zones to minimize the probability of disk failure. In this study, we integrate an existing energy consumption technique that operates the disks at different power modes with our proposed reliability approach. Experimental results show that our approach can significantly improve reliable while achieving high energy efficiency for disk systems.
Kiranmai Bellam, Adam Manzanares, Xiaojun Ruan, Xiao Qin 0001
ISCC3
2008 Improving Security of Real-Time Wireless Networks Through Packet Scheduling [Transactions Letters]
abstract
Modern real-time wireless networks require high security level to assure confidentiality of information stored in packages delivered through wireless links. However, most existing algorithms for scheduling independent packets in real-time wireless networks ignore various security requirements of the packets. Therefore, in this paper we remedy this problem by proposing a novel dynamic security-aware packet-scheduling algorithm, which is capable of achieving high quality of security for realtime packets while making the best effort to guarantee realtime requirements (e.g., deadlines) of those packets. We conduct extensive simulation experiments to evaluate the performance of our algorithm. Experimental results show that compared with two baseline algorithms, the proposed algorithm can substantially improve both quality of security and real-time packet guarantee ratio under a wide range of workload characteristics.
Xiao Qin 0001, Mohammed I. Alghamdi, Mais Nijim, Ziliang Zong, Kiranmai Bellam, Xiaojun Ruan, Adam Manzanares
IEEE Trans. Wirel. Commun.6
2007 Interplay of Security and Reliability using Non-uniform Checkpoints
abstract
Real time applications such as military aircraft flight control systems and online banking are critical with respect to security and reliability. In this paper we presented a way to integrate both by considering confidentiality and integrity services for security and nonuniform checkpoint strategy for reliability. The slack exploitation interacts in subtle ways for security in regards to the placement of checkpoint. The checkpoints are placed in to the task at low frequency in the beginning because the slack available can accommodate a large amount of work at risk and the frequency is increased there after considering the slack available. The security is applied to the data in two ways. First method introduces the security for the entire data at once whereas in the second method the data is divided into n uneven sections and each section is separately secured . That is at the start of the task basic security services are considered depending on the slack available. The security is increased gradually for the rest of the task but if there exist a fault, then at that point the security is maintained at the steady rate because of the limited slack. Compared to the first method the second method can provide up to a 32.3 percent higher security. While compared to the traditional checkpoint strategy, the non-uniform check pointing makes more efficient use of slack while increasing the overall security levels by 34.4 percent for the second method.
Kiranmai Bellam, Raghava K. Vudata, Xiao Qin 0001, Ziliang Zong, Xiaojun Ruan, Mais Nijim
ICCCN5
2007 An Energy-Efficient Scheduling Algorithm Using Dynamic Voltage Scaling for Parallel Applications on Clusters
abstract
In the past decade cluster computing platforms have been widely applied to support a variety of scientific and commercial applications, many of which are parallel in nature. However, scheduling parallel applications on large scale clusters is technically challenging due to significant communication latencies and high energy consumption. As such, shortening schedule length and conserving energy consumption are two major concerns in designing economical and environmentally friendly clusters. In this paper, we propose an energy-efficient scheduling algorithm (TDVAS) using the dynamic voltage scaling technique to provide significant energy savings for clusters. The TDVAS algorithm aims at judiciously leveraging processor idle times to lower processor voltages (i.e., the dynamic voltage scaling technique or DVS), thereby reducing energy consumption experienced by parallel applications running on clusters. Reducing processor voltages, however, can inevitably lead to increased execution times of parallel task. The salient feature of the TDVAS algorithm is to tackle this problem by exploiting tasks precedence constraints. Thus, TDVAS applies the DVS technique to parallel tasks followed by idle processor times to conserve energy consumption without increasing schedule lengths of parallel applications. Experimental results clearly show that the TDVAS algorithm is conducive to reducing energy dissipation in large-scale clusters without adversely affecting system performance.
Xiaojun Ruan, Xiao Qin 0001, Ziliang Zong, Kiranmai Bellam, Mais Nijim
ICCCN1
2007 Energy-Efficient Scheduling for Parallel Applications Running on Heterogeneous Clusters
abstract
High performance clusters have been widely used to provide amazing computing capability for both commercial and scientific applications. However, huge power consumption has prevented the further application of large-scale clusters. Designing energy-efficient scheduling algorithms for parallel applications running on clusters, especially on the high performance heterogeneous clusters, is highly desirable. In this regard, we propose a novel scheduling strategy called energy efficient task duplication schedule (EETDS for short), which can significantly conserve power by judiciously shrinking communication energy cost when allocating parallel tasks to heterogeneous computing nodes. We present the preliminary simulation results for Gaussian and FFT parallel task models to prove the efficiency of our algorithm.
Ziliang Zong, Xiao Qin 0001, Xiaojun Ruan, Kiranmai Bellam, Mais Nijim, Mohammed I. Alghamdi
ICPP3