EDBT 2026 Demo / reviewers in the wild / expert
Tao Li 0006
dblp:75/4601-6
· DBLP profile ↗
144ranked-venue papers
11as first author
9since 2021 · last 2021
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 122 · 9 first-author · 8 since 2021Software engineering, systems software and programming languages · 23 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Security and privacy · 5Artificial intelligence and machine learning · 3Computer networks · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | A Low Cost Weight Obfuscation Scheme for Security Enhancement of ReRAM Based Neural Network AcceleratorsabstractThe resistive random-access memory (ReRAM) based accelerator can execute the large scale neural network (NN) applications in an extremely energy efficient way. However, the non-volatile feature of the ReRAM introduces some security vulnerabilities. The weight parameters of a well-trained NN model deployed on the ReRAM based accelerator are persisted even after the chip is powered off. The adversaries who have the physical access to the accelerator can hence launch the model stealing attack and extract these weights by some micro-probing methods. Run time encryption of the weights is intuitive to protect the NN model but degrades execution performance and device endurance largely. While obfuscation of the weight rows needs to pay the tremendous hardware area overhead in order to achieve the high security. In view of above mentioned problems, in this paper we propose a low cost weight obfuscation scheme to secure the NN model deployed on the ReRAM based accelerators from the model stealing attack. We partition the crossbar into many virtual operation units (VOUs) and perform full permutation on the weights of the VOUs along the column dimension. Without the keys, the attacker cannot perform the correct NN computations even if they have obtained the obfuscated model. Compared with the weight rows based obfuscation, our scheme can achieve the same level of security with less an order of magnitude in the hardware area and power overheads. Tao Li 0006 |
ASP-DAC | 3 |
| 2021 | Fast On-Road Object Detector on ROS-Based Mobile Robot
Gang Wang 0001, Qiudi Song, Tao Li 0006 |
ICA3PP (2) | 3 |
| 2021 | Preface
Xian-He Sun, Dong Li 0001, Wen-Guang Chen, Tao Li 0006, Jiwu Shu, Bo Wu 0002, Jin Xiong, Jinging Xue, Feng Zhang 0007, Jidong Zhai, Zhiia Zhao |
J. Comput. Sci. Technol. | 4 |
| 2021 | Democratic learning: hardware/software co-design for lightweight blockchain-secured on-device machine learningabstractRecently, the trending 5G technology encourages extensive applications of on-device machine learning , which collects user data for model training. This requires cost-effective techniques to preserve the privacy and the security of model training within the resource-constrained environment. Traditional learning methods rely on the trust among the system for privacy and security. However, with the increase of the learning scale, maintaining every edge device’s trustworthiness could be expensive. To cost-effectively establish trust in a trustless environment, this paper proposes democratic learning (DemL), which makes the first step to explore hardware/software co-design for blockchain-secured decentralized on-device learning. By utilizing blockchain’s decentralization and tamper-proofing, our design secures AI learning in a trustless environment. To tackle the extra overhead introduced by blockchain , we propose PoMC (an algorithm and architecture co-design) as a novel blockchain consensus mechanism , which first exploits cross-domain reuse (AI learning and blockchain consensus) in AI learning architecture. Evaluation results show our DemL can protect AI learning from privacy leakage and model pollution, and demonstrated that privacy and security come with trivial hardware overhead and power consumption (2%). We believe that our work will open the door of synergizing blockchain and on-device learning for security and privacy. Mingcong Song, Tao Li 0006, Zhibin Yu 0001, Yuting Dai, Xiaoguang Liu 0001, Gang Wang 0001 |
J. Syst. Archit. | 3 |
| 2021 | LrGAN: A Compact and Energy Efficient PIM-Based Architecture for GAN TrainingabstractAs a powerful unsupervised learning method, Generative Adversarial Network (GAN) plays an essential role in many domains. However, training a GAN imposes four more challenges: (1) intensive communication caused by complex train phases of GAN; (2) much more ineffectual computations caused by peculiar convolutions; (3) more frequent off-chip memory accesses for exchanging intermediate data between the generator and the discriminator; and (4) high energy consumption of unnecessary fine-grained MLC programming. In this article, we propose LrGAN, a PIM-based GAN accelerator, to address the challenges of training GAN. We first propose a zero-free data reshaping scheme for ReRAM-based PIM, which removes the zero-related computations. We then propose a 3D-connected PIM, which can reconfigure connections inside PIM dynamically according to dataflows of propagation and updating. After that, we propose an approximate weight update algorithm to avoid unnecessary fine-grain MLC programming. Finally, we propose LrGAN based on these three techniques, providing different levels of accelerating GAN for programmers. Experiments show that LrGAN achieves 47.2×, 21.42×, and 7.46× speedup over FPGA-based GAN accelerator, GPU platform, and ReRAM-based neural network accelerator respectively. Besides, LrGAN achieves 13.65×, 10.75×, and 1.34× energy saving on average over GPU platform, PRIME, and FPGA-based GAN accelerator, respectively. Haiyu Mao, Jiwu Shu, Mingcong Song, Tao Li 0006 |
IEEE Trans. Computers | 4 |
| 2021 | Toward Efficient Execution of Mainstream Deep Learning Frameworks on Mobile Devices: Architectural ImplicationsabstractIn recent years, continuous growing interests have been seen in bringing artificial intelligence capabilities to mobile devices. However, the related work still faces several issues, such as constrained computation and memory resources, power drain, and thermal limitation. To develop deep learning (DL) algorithms on mobile devices, we need to understand their behaviors. In this article, we explore the architectural behaviors of some mainstream DL frameworks on mobile devices by performing a comprehensive characterization of performance, accuracy, energy efficiency, and thermal behaviors. We experimentally choose four model compression methods to perform on networks and in addition, analyze the related impact on the nodes amount, memory, execution time, model size, inference time, energy consumption, and thermal distribution. With insights into DL-based mobile application characteristics, we hope to guide the design of future smartphone platforms for lower energy consumption. Yuting Dai, Benyong Liu, Tao Li 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Leveraging the Interplay of RAID and SSD for Lifetime Optimization of Flash-Based SSD RAIDabstractFlash-based SSD RAID arrays are increasingly being deployed in data centers. Compared with HDD arrays, SSD arrays drastically enhance I/O performance and density, and reduce power, cooling, and rack space. Nevertheless, SSDs suffer aging issues. Especially, an SSD has limited endurance and needs to be replaced when it reaches to the end of its lifetime. Although prior studies have been conducted to address this disadvantage, effective techniques of RAID/SSD controllers are urgently needed to extend the lifetime of SSD arrays. In this article, we propose a novel RAID architecture, called FreeRAID, to leverage the interplay of RAID and SSD controllers to optimize the lifespan of flash-based SSD arrays. FreeRAID adds a new exploitable phase to the life cycle of flash blocks. In FreeRAID, flash space is separated into normal space and exploitable space, and they are used to serve normal data and approximate data, respectively. We design a dual-space management scheme for RAID controllers to intelligently allocate SSD spaces based on their aging status. Inside an SSD, we propose an adaptive flash translation layer for the SSD controller to maintain the reliability and space efficiency of flash memories. We implemented a prototype of FreeRAID based on an SSD array simulator. Our experiments show that FreeRAID can significantly increase the lifetime by up to 3.07 × compared with conventional SSD-based RAID arrays. Zhaoyan Shen, Chenlin Ma, Zhiping Jia, Tao Li 0006, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | CASpMV: A Customized and Accelerative SpMV Framework for the Sunway TaihuLightabstractThe Sunway TaihuLight, equipped with 10 million cores, is currently the world's third fastest supercomputer. SpMV is one of core algorithms in many high-performance computing applications. This paper implements a fine-grained design for generic parallel SpMV based on the special Sunway architecture and finds three main performance limitations, i.e., storage limitation, load imbalance, and huge overhead of irregular memory accesses. To address these problems, this paper introduces a customized and accelerative framework for SpMV (CASpMV) on the Sunway. The CASpMV customizes an auto-tuning four-way partition scheme for SpMV based on the proposed statistical model, which describes the sparse matrix structure characteristics, to make it better fit in with the computing architecture and memory hierarchy of the Sunway. Moreover, the CASpMV provides an accelerative method and customized optimizations to avoid irregular memory accesses and further improve its performance on the Sunway. Our CASpMV achieves a performance improvement that ranges from 588.05 to 2118.62 percent over the generic parallel SpMV on a CG (which corresponds to an MPI process) of the Sunway on average and has good scalability on multiple CGs. The performance comparisons of the CASpMV with state-of-the-art methods on the Sunway indicate that the sparsity and irregularity of data structures have less impact on CASpMV. Guoqing Xiao 0001, Kenli Li 0001, Yuedan Chen, Wangquan He, Albert Y. Zomaya, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Exploring Highly Dependable and Efficient Datacenter Power System Using Hybrid and Hierarchical Energy BuffersabstractThe massive and irregular load surges challenge datacenter power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial availability issue in modern datacenters which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches. In this paper, we propose Hybrid and Hierarchical Energy Buffering (HHEB), a novel heterogeneous and adaptive scheme that could enable various energy storage devices (ESDs) to be efficiently integrated into existing datacenters for dynamically dealing with power mismatches. Our techniques exploit the diverse characteristics of different ESDs and intelligent load assignment algorithms to improve the dependability and efficiency of datacenter power systems. We evaluate the HHEB design with a prototype. Compared with a homogenous battery energy buffering system, HHEB could improve energy efficiency by 39.7 percent, extend UPS lifetime by 4.7X, promote energy availability by 3.2X, reduce system downtime by 41 percent, and effectively improve the energy availability of various energy buffers in different hierarchies. It allows datacenters to adapt to various power supply anomalies, thereby improving operational efficiency, dependability and availability. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Sustain. Comput. | 4 |
| 2020 | CoExe: An Efficient Co-execution Architecture for Real-Time Neural Network ServicesabstractEnd-to-end latency is sensitive for user-interactive neural network (NN) services on clouds. For periods of high request load, co-locating multiple NN requests has the potential to reduce end-to-end latency. However, current batch-based accelerators lack request-level parallelism support, leaving the queuing time non-optimized. Meanwhile, naively partitioning resources for simultaneous requests suffers from longer execution time as well as lower resource efficiency because different applications utilize separate resources without sharing. To effectively reduce the end-to-end latency for real-time NN requests, we propose CoExe architecture, equipped with a pipeline implementation of a sparsity-driven real-time co-execution model. By leveraging the non-trivial amount of sparse operations during concurrent NNs execution, the end-to-end latency is decreased by up to 12.3× and 2.4× over Eyeriss-like and SCNN at peak workload mode. Besides, we propose row cross (RC) dataflow to reduce data movement cost, and avoid memory duplication. Chubo Liu, Kenli Li 0001, Mingcong Song, Jiechen Zhao 0003, Keqin Li 0001, Tao Li 0006, Zihao Zeng |
DAC | 6 |
| 2020 | BBS: Micro-Architecture Benchmarking Blockchain Systems through Machine Learning and Fuzzy SetabstractDue to the decentralization, irreversibility, and traceability, blockchain has attracted significant attention and has been deployed in many critical industries such as banking and logistics. However, the micro-architecture characteristics of blockchain programs still remain unclear. What's worse, the large number of micro-architecture events make understanding the characteristics extremely difficult. We even lack a systematic approach to identify the important events to focus on. In this paper, we propose a novel benchmarking methodology dubbed BBS to characterize blockchain programs at micro-architecture level. The key is to leverage fuzzy set theory to identify important micro-architecture events after the significance of them is quantified by a machine learning based approach. The important events for single programs are employed to characterize the programs while the common important events for multiple programs form an importance vector which is used to measure the similarity between benchmarks. We leverage BBS to characterize seven and six benchmarks from Blockbench and Caliper, respectively. The results show that BBS can reveal interesting findings. Moreover, by leveraging the importance characterization results, we improve that the transaction throughput of Smallbank from Fabric by 70% while reduce the transaction latency by 55%. In addition, we find that three of seven and two of six benchmarks from Blockbench and Caliper are redundant, respectively. Chao Chen 0022, Zihao Su, Weiguang Chen, Tao Li 0006, Zhibin Yu 0001 |
HPCA | 5 |
| 2020 | QuPAA: Exploiting Parallel and Adaptive Architecture to Scale up Quantum ComputingabstractQuantum computing has been gradually developed from a theory to practice in recent years. The desire for more powerful computational capability keeps promoting the scaling up of quantum computers. For quantum machines with large scales, limited coherence time and bloated instruction bandwidth may obstruct the continuous scaling up. In this paper, we propose a new quantum computing architecture named QuPAA, which leverages multiple layers of arbitrary waveform generator (AWG) units to exploit the parallelism among quantum operations. QuPAA provides a series of mechanisms, including parallelism exploitation, microcode streaming, and a quantum flexible instruction set (QuFIS). These mechanisms assist QuPAA to generate instructions adaptively and reduce the requirement for coherence time and instruction bandwidth. The evaluation results show that, compared to the baseline quantum computer with a single AWG unit, QuPAA reduces 25.9% to 97.4% coherence time requirement, and saves 26.9% to 79.1 % instruction bandwidth. The improvement effectively breaks the limitations and provides strong support for the extension of the scale for quantum computers. Yingxun Fu, Tao Li 0006 |
ICCD | 3 |
| 2020 | Singular Spectrum Analysis for Local Differential Privacy of Classifications in the Smart GridabstractNew privacy implications are induced to individuals and families because of the time-series data classification problem in the Internet of Things such as appliance classifications in the smart grid. To prevent the adversary from inferring the household appliance classification used in the smart grid, a singular spectrum analysis (SSA) has been applied to the local differential privacy (SSA-LDP). First, the Fourier spectrum noise has been added via the geometric sum which has been proved to achieve the Laplace noise distribution. Furthermore, we have proved that the sanitized data through the SSA-LDP is ε -deferentially private for the adversary inference attack. In addition, to achieve a better data utility, a formula has been obtained for the optimal Fourier spectrum noise by decomposing it into the superposition of power spectra of the dominant SSA eigenfilters. Finally, experiments have been performed with a computer-generated data set and a real-world smart-meter data set. Comparisons to other privacy approaches show that the optimized SSA-LDP does achieve a better data utility for a given data privacy. Lu Ou, Zheng Qin 0001, Shaolin Liao, Tao Li 0006, Da-Fang Zhang 0001 |
IEEE Internet Things J. | 4 |
| 2020 | GPU based parallel optimization for real time panoramic video stitchingabstractPanoramic video is a sort of video recorded at the same point of view to record the full scene. With the development of video surveillance and the requirement for 3D converged video surveillance in smart cities, CPU and GPU are required to possess strong processing abilities to make panoramic video. The traditional panoramic products depend on post processing, which results in high power consumption, low stability and unsatisfying performance in real time. In order to solve these problems, we propose a real-time panoramic video stitching framework. The framework we propose mainly consists of three algorithms, L-ORB image feature extraction algorithm, feature point matching algorithm based on LSH and GPU parallel video stitching algorithm based on CUDA. The experiment results show that the algorithm mentioned can improve the performance in the stages of feature extraction of images stitching and matching, the running speed of which is 11.3 times than that of the traditional ORB algorithm and 641 times than that of the traditional SIFT algorithm. Based on analyzing the GPU resources occupancy rate of each resolution image stitching, we further propose a stream parallel strategy to maximize the utilization of GPU resources. Compared with the L-ORB algorithm, the efficiency of this strategy is improved by 1.6–2.5 times, and it can make full use of GPU resources. The performance of the system accomplished in the paper is 29.2 times than that of the former embedded one, while the power dissipation is reduced to 10 W. Chengyao Du, Jingling Yuan, Jiansheng Dong, Lin Li 0001, Mincheng Chen, Tao Li 0006 |
Pattern Recognit. Lett. | 6 |
| 2020 | Temperature-Aware Persistent Data Management for LSM-Tree on 3-D NAND Flash MemoryabstractKey-value (KV) store has been widely deployed in both embedded systems and enterprise systems. Most KV stores today use log structured merge tree (LSM-Tree), as LSM-Tree can eliminate random write operations to the secondary storage and maintain acceptable read performance. LSM-Tree is originally designed for the secondary storage device with hard disk drives. As the emerging storage media, three-dimensional (3-D) flash memory has become the mainstream technology to replace hard disk drives. Different from hard disk drives and the conventional planar flash memory, 3-D flash memory is vulnerable to temperature. High temperature will introduce both charge loss and retention degradation. Since LSM-Tree transfers random write operations to the sequential ones, the access to consecutive physical address in flash memory will cause the temperature issue. This will affect the integrity of data stored in 3-D flash memory. This article presents TLSM, a temperature-aware persistent data management scheme for LSM-Tree-based KV store on 3-D NAND flash memory. TLSM offers both application-level LSM-Tree optimization and firmware-level address management to allocate persistent data to 3-D flash. At the application-level, TLSM presents a novel temperature-aware LSM data structure to reduce the amount of data issued from LSM-Tree to 3-D flash memory. At the firmware-level, TLSM reallocates the data to physical blocks with relatively low temperature. This cross-layer optimization can effectively handle the temperature issue to ensure the data integrity of LSM-Tree in 3-D flash memory. We demonstrate the viability of the proposed scheme using a set of standard benchmarks. Our extensive evaluations show that, TLSM can significantly enhance the data integrity and reduce write amplifications compared to representative schemes. Yi Wang 0003, Jiali Tan, Rui Mao 0001, Tao Li 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | An adaption scheduling based on dynamic weighted random forests for load demand forecasting
Mincheng Chen, Jingling Yuan, Dongling Liu, Tao Li 0006 |
J. Supercomput. | 4 |
| 2020 | aeSpTV: An Adaptive and Efficient Framework for Sparse Tensor-Vector Product Kernel on a High-Performance Computing PlatformabstractMulti-dimensional, large-scale, and sparse data, which can be neatly represented by sparse tensors, are increasingly used in various applications such as data analysis and machine learning. A high-performance sparse tensor-vector product (SpTV), one of the most fundamental operations of processing sparse tensors, is necessary for improving efficiency of related applications. In this article, we propose aeSpTV, an adaptive and efficient SpTV framework on Sunway TaihuLight supercomputer, to solve several challenges of optimizing SpTVon high-performance computing platforms. First, to map SpTV to Sunway architecture and tame expensive memory access latency and parallel writing conflict due to the intrinsic irregularity of SpTV, we introduce an adaptive SpTV parallelization. Second, to co-execute with the parallelization design while still ensuring high efficiency, we design a sparse tensor data structure named CSSoCR. Third, based on the adaptive SpTV parallelization with the novel tensor data structure, we present an autotuner that chooses the most befitting tensor partitioning method for aeSpTV using the variance analysis theory of mathematical statistics to achieve load balance. Fourth, to further leverage the computing power of Sunway, we propose customized optimizations for aeSpTV. Experimental results show that aeSpTV yields good sacalability on both thread-level and process-level parallelism of Sunway. It achieves a maximum GFLOPS of 195.69 on 128 processes. Additionally, it is proved that optimization effects of the partitioning autotuner and optimization techniques are remarkable. Yuedan Chen, Guoqing Xiao 0001, M. Tamer Özsu, Chubo Liu, Albert Y. Zomaya, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | COPA: Highly Cost-Effective Power Back-Up for Green DatacentersabstractTraditional datacenters employ costly diesel generators (DG) and uninterrupted power supplies (UPS) to back up power. However, some or even all racks of a green datacenter can still be powered by renewable energy during grid power outages. This makes the utilization of the DGs and UPSs in green datacenters significantly lower than in traditional datacenters. In this paper, we propose a highly cost-effective power back-up (COPA) approach for green datacenters by leveraging the availability characteristics of renewable energy as well as grid power outages. COPA contributes three new techniques. The first technique, called least UPS capacity planning, determines the least rated power capability and runtime of the UPSs to guarantee the normal operations of a green datacenter during grid power outages. The second technique, named cooperative UPS/renewable power supply, employs UPS and renewable energy at the same time to supply power to each rack when grid power fails. The last one, dubbed renewable-energy-aware dynamic power management, controls the power consumption dynamically based on the available capacity of renewable energy and UPS. We build an experimental cluster consisting of 10 servers, and use four representative benchmarks as well as verified data about the availability characteristics of solar and wind energy to evaluate COPA. The results show that COPA reduces 47 percent and 70 percent of the power back-up cost for a solar energy powered datacenter and a wind energy powered datacenter, respectively. Moreover, COPA guarantees the application's Service Level Agreement (SLA) for at least 20 minutes (over 79 percent outages) and 56 minutes on average while enabling the back-up power to last for at least 2 hours and for 3 hours on average, which cannot be achieved by other under-provisioning power back-up approaches. Junmin Wu, Lieven Eeckhout, Amer Qouneh, Tao Li 0006, Zhibin Yu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | PATCH: Process-Variation-Resilient Space Allocation for Open-Channel SSD with 3D FlashabstractAdvanced three-dimensional (3D) flash memory adopts charge-trap technology that can effectively improve the hit density and reduce the coupling effect. Despite these advantages, 3D charge-trap flash brings a number of new challenges. First, current etching process is unable to manufacture perfect channels with identical feature size. Second, the cell current in 3D charge-trap flash is only 20% compared to planar flash memory, making it difficult to give a reliable sensing margin. These issues are affected by process variation, and they pose threats to the integrity of data stored in 3D charge-trap flash. This paper presents PATCH, a process-variation-resilient space allocation scheme for open-channel SSD with 3D charge-trap flash memory. PATCH is a novel hardware and file system interface that can transparently allocate physical space in the presence of process variation. PATCH utilizes the rich functionalities provided by the system infrastructure of open-channel SSD to reduce the uncorrectable bit errors. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that PATCH can effectively enhance the reliability with negligible extra erase operations in comparison with representative schemes. Yi Wang 0003, Amelie Chi Zhou, Rui Mao 0001, Tao Li 0006 |
DATE | 5 |
| 2019 | Towards Cross-Platform Inference on Edge Devices with Emerging Neuromorphic ArchitectureabstractDeep convolutional neural networks have become the mainstream solution for many artificial intelligence applications. However, they are still rarely deployed on mobile or edge devices due to the cost of a substantial amount of data movement among limited resources. The emerging processing-inmemory neuromorphic architecture offers a promising direction to accelerate the inference process. The key issue becomes how to effectively allocate the processing of inference between computing and storage resources on an edge device.This paper presents Mobile-I, a resource allocation scheme to accelerate the Inference process on Mobile or edge devices. Mobile-I targets at the emerging 3D neuromorphic architecture to reduce the processing latency among computing resources and fully utilize the limited on-chip storage resources. We formulate the target problem as a resource allocation problem and use a software-based solution to offer the cross-platform deployment across multiple mobile or edge devices. We conduct a set of experiments using realistic workloads that are generated from Intel Movidius neural compute stick. Experimental results show that Mobile-I can effectively reduce the processing latency and improve the utilization of computing resources with negligible overhead in comparison with representative schemes. Shangyu Wu, Yi Wang 0003, Amelie Chi Zhou, Rui Mao 0001, Zili Shao, Tao Li 0006 |
DATE | 6 |
| 2019 | Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage ScalingabstractAs the application scope of deep neural networks (DNNs) moves from large-scale data centers to small-scale mobile devices, power wall has become one of the most important obstacles. Voltage scaling is a typical technique enables power saving, but it causes reliability and performance challenges. Therefore, an energy-efficient and reliable scheme for NNs is required to balance above three aspects according to users' requirements for excellent user experience. In this paper, we innovatively propose neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation in NNs and precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Multi-objective optimization and clustering method are combined to find the optimal voltage islands. Finally, we conduct experiment to demonstrate the efficacy of the proposed technique. Jing Wang 0055, Xin Fu 0001, Xingyao Zhang 0002, Lan Gao 0004, Weigong Zhang, Tao Li 0006 |
ICPADS | 8 |
| 2019 | Eager pruning: algorithm and architecture support for fast training of deep neural networksabstractToday's big and fast data and the changing circumstance require fast training of Deep Neural Networks (DNN) in various applications. However, training a DNN with tons of parameters involves intensive computation. Enlightened by the fact that redundancy exists in DNNs and the observation that the ranking of the significance of the weights changes slightly during training, we propose Eager Pruning, which speeds up DNN training by moving pruning to an early stage. Jiaqi Zhang 0002, Mingcong Song, Tao Li 0006 |
ISCA | 4 |
| 2019 | REcache: Efficient Sustainable Energy Management Circuits and Policies for Computing SystemsabstractThe rapidly growing computing systems, such as AI server cluster, IoT devices etc. are facing increasing energy expenditure pressure and the warning of carbon footprint. Designing eco-friendly computing systems which integrated renewable energy sources have attracted considerable attentions recently. Existing schemes either incur green energy efficiency degradation or sacrifice workload performance. This paper proposes REcache (Renewable Energy cache), a sustainable energy management scheme to efficiently utilize green energy for computing systems. Compared to previous proposals, we present a dedicated circuit and energy-aware management policies to coordinate energy harvesting, power management and workload scheduling. We evaluate our scheme through both prototyping and simulation. The experimental results show that the REcache could effectively improve the energy availability 10%, workload performance 5% for different workloads on average. Longjun Liu, Hongbin Sun 0001, Nanning Zheng 0001, Tao Li 0006 |
ISCAS | 5 |
| 2019 | Towards Efficient NVDIMM-based Heterogeneous Storage Hierarchy Management for Big Data WorkloadsabstractIn this paper, we propose a holistic solution to address several important and challenging issues in storage data management in light of emerging NVDIMM-based architecture: namely, new performance modeling, NVDIMM-based migration, and architectural support for NVDIMMs on migration optimization. In particular, a novel NVDIMM-based heterogeneous storage performance model is proposed to effectively address bus contention issues caused by placing NVDIMMs on the memory bus. We also develop an NVDIMM-based lazy migration scheme to effectively minimize adverse effects caused by memory traffic interferences during storage data management processes. Finally, the NVDIMM-based architectural support for migration optimization is proposed to increase channel parallelism in the destination NVDIMMs and bypass buffer caches in the source NVDIMMs, so that the impact of memory traffic can be alleviated. We present detailed evaluation and analysis to quantify how well our techniques can enhance the I/O performances of big workloads via efficient heterogeneous storage hierarchy management. Our experimental results show that overall the proposed techniques yield up to 98% performance improvement over the state-of-the-art techniques. Renhai Chen, Zili Shao, Duo Liu 0002, Zhiyong Feng 0002, Tao Li 0006 |
MICRO | 5 |
| 2019 | An Energy Dynamic Control Algorithm Based on Reinforcement Learning for Data CentersabstractIn recent years, how to use renewable energy to reduce the energy cost of internet data center (IDC) has been an urgent problem to be solved. More and more solutions are beginning to consider machine learning, but many of the existing methods need to take advantage of some future information, which is difficult to obtain in the actual operation process. In this paper, we focus on reducing the energy cost of IDC by controlling the energy flow of renewable energy without any future information. we propose an efficient energy dynamic control algorithm based on the theory of reinforcement learning, which approximates the optimal solution by learning the feedback of historical control decisions. For the purpose of avoiding overestimation, improving the convergence ability of the algorithm, we use the double [Formula: see text]-method to further optimize. The extensive experimental results show that our algorithm can on average save the energy cost by 18.3% and reduce the rate of grid intervention by 26.2% compared with other algorithms, and thus has good application prospects. Yao Xiang, Jingling Yuan, Ruiqi Luo, Xian Zhong, Tao Li 0006 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2019 | A Thermal-Aware Physical Space Reallocation for Open-Channel SSD With 3-D Flash Memoryabstract3-D flash memory faces a number of challenges, including thermal issues and process variation. The high temperature will cause charge loss and lead to the fluctuation of threshold voltage. To address the thermal issue of 3-D flash memory, this paper presents ThermAlloc, a novel thermal-aware physical space allocation strategy for open-channel solid-state drive with 3-D charge trapping flash memory. ThermAlloc permutes the allocation of physical blocks. Consecutively accessed logical blocks are distributed to different physical locations in order to prevent the accumulation of hotspots. The objective is to postpone garbage collection operations and keep the distribution of block temperature well under control. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that ThermAlloc can reduce the peak temperature by 26.94% with 3.15% extra worst-case response time in comparison with the baseline scheme. Yi Wang 0003, Mingxu Zhang, Tao Li 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | A Temperature-Aware Reliability Enhancement Strategy for 3-D Charge-Trap Flash MemoryabstractCompared to the conventional planar flash memory, advanced 3-D flash memory adopts charge-trap technology that can significantly enhance cell density and storage capacity. Despite these advantages, 3-D charge-trap flash memory brings several new challenges. First, charge-trap flash is sensitive to temperature. Recent studies demonstrate that, the high temperature will incur both charge loss and retention degradation. This issue does not happen in 2-D flash memory which adopts floating gate technology. Second, current 3-D charge-trap flash integrates the extra large capacity physical block, and each block contains over 1024 physical pages. The large-capacity block infrastructure will cause extra garbage collection overhead, which makes the thermal issue more complicated. This paper presents TempCure, a temperature-aware reliability enhancement strategy for 3-D charge-trap flash memory. TempCure is a novel hardware and file system interface that can transparently allocate physical space based on the temperature status. TempCure adopts two reliability enhancement strategies, temperature mining and block allotment, to prevent the generation of hotspots and enhance the data integrity of 3-D flash memory. We conduct a set of experiments using standard benchmarks. Experimental results show that TempCure can effectively reduce the peak temperature and block erase counts with negligible timing overhead in comparison with representative schemes. Yi Wang 0003, Jiangfan Huang, Jing Yang 0018, Tao Li 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Performance-Aware Model for Sparse Matrix-Matrix Multiplication on the Sunway TaihuLight SupercomputerabstractGeneral sparse matrix-sparse matrix multiplication (SpGEMM) is one of the fundamental linear operations in a wide variety of scientific applications. To implement efficient SpGEMM for many large-scale applications, this paper proposes scalable and optimized SpGEMM kernels based on COO, CSR, ELL, and CSC formats on the Sunway TaihuLight supercomputer. First, a multi-level parallelism design for SpGEMM is proposed to exploit the parallelism of over 10 millions cores and better control memory based on the special Sunway architecture. Optimization strategies, such as load balance, coalesced DMA transmission, data reuse, vectorized computation, and parallel pipeline processing, are applied to further optimize performance of SpGEMM kernels. Second, we thoroughly analyze the performance of the proposed kernels. Third, a performance-aware model for SpGEMM is proposed to select the most appropriate compressed storage formats for the sparse matrices that can achieve the optimal performance of SpGEMM on the Sunway. The experimental results show the SpGEMM kernels have good scalability and meet the challenge of the high-speed computing of large-scale data sets on the Sunway. In addition, the performance-aware model for SpGEMM achieves an absolute value of relative error rate of 8.31 percent on average when the kernels are executed in one single process and achieves 8.59 percent on average when the kernels are executed in multiple processes. It is proved that the proposed performance-aware model can perform at high accuracy and satisfies the precision of selecting the best formats for SpGEMM on the Sunway TaihuLight supercomputer. Yuedan Chen, Kenli Li 0001, Wangdong Yang, Guoqing Xiao 0001, Xianghui Xie 0001, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | Exploiting Parallelism for CNN Applications on 3D Stacked Processing-In-Memory ArchitectureabstractDeep convolutional neural networks (CNNs) are widely adopted in intelligent systems with unprecedented accuracy but at the cost of a substantial amount of data movement. Although the emerging processing-in-memory (PIM) architecture seeks to minimize data movement by placing memory near processing elements, memory is still the major bottleneck in the entire system. The selection of hyper-parameters in the training of CNN applications requires over hundreds of kilobytes cache capacity for concurrent processing of convolutions. How to jointly explore the computation capability of the PIM architecture and the highly parallel property of neural networks remains a critical issue. This paper presents Para-Net, that exploits Parallelism for deterministic convolutional neural Networks on the PIM architecture. Para- Net achieves data-level parallelism for convolutions by fully utilizing the on-chip processing engine (PE) in PIM. The objective is to capture the characteristics of neural networks and present a hardware-independent design to jointly optimize the scheduling of both intermediate results and computation tasks. We formulate this data allocation problem as a dynamic programming model and obtain an optimal solution. To demonstrate the viability of the proposed Para-Net, we conduct a set of experiments using a variety of realistic CNN applications. The graph abstractions are obtained from deep learning framework Caffe. Experimental results show that Para-Net can significantly reduce processing time and improve cache efficiency compared to representative schemes. Yi Wang 0003, Weixuan 'Vincent' Chen, Jing Yang 0018, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Towards Fast and Lightweight Checkpointing for Mobile Virtualization Using NVRAMabstractCheckpointing is a key enabler of hibernation, live migration and fault-tolerance for virtual machines (VMs) in mobile devices. However, checkpointing a VM is usually heavyweight: the VM's entire memory needs to be dumped to storage, which induces a significant amount of (slow) I/O operations, degrading system performance and user experience. In this paper, we propose FLIC, a fast and lightweight checkpointing machinery for virtualized mobile devices by taking advantages of recent byte-addressable, non-volatile memory (NVRAM). Instead of saving the VM's entire memory to storage, we store its working set pages in NVRAM, avoiding accessing slow flash memory (compared to server-grade SSDs). To further reduce the write activities to flash memory, we propose an energy-efficient data deduplication to eliminate redundant data in VM snapshot and save storage space. Experimental results based on an Exynos 5250 SoC show that our approach can effectively improve the performance of checkpointing in mobile virutalization and save energy. Kan Zhong, Duo Liu 0002, Yunsong Wu, Linbo Long, Weichen Liu 0001, Jinting Ren, Renping Liu 0002, Liang Liang 0002, Zili Shao, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2018 | Exploiting Dynamic Thermal Energy Harvesting for Reusing in Smartphone with Mobile ApplicationsabstractRecently, mobile applications have gradually become performance- and resource- intensive, which results in a massive battery power drain and high surface temperature, and further degrades the user experience. Thus, high power consumption and surface over-heating have been considered as a severe challenge to smartphone design. In this paper, we propose DTEHR, a mobile Dynamic Thermal Energy Harvesting Reusing framework to tackle this challenge. The approach is sustainable in that it generates energy using dynamic Thermoelectric Generators (TEGs). The generated energy not only powers Thermoelectric Coolers (TECs) for cooling down hot-spots, but also recharges micro-supercapacitors (MSCs) for extended smartphone usage. To analyze thermal characteristics and evaluate DTEHR across real-world applications, we build MPPTAT (Multi-comPonent Power and Thermal Analysis Tool), a power and thermal analyzing tool for Android. The result shows that DTEHR reduces the temperature differences between hot areas and cold areas up to 15.4°C (internal) and 7°C (surface). With TEC-based hot-spots cooling, DTEHR reduces the temperature of the surface and internal hot-spots by an average of 8° and 12.8mW respectively. With dynamic TEGs, DTEHR generates 2.7-15mW power, more than hundreds of times of power that TECs need to cool down hot-spots. Thus, extra-generated power can be stored into MSCs to prolong battery life. Yuting Dai, Tao Li 0006, Benyong Liu, Mingcong Song, Huixiang Chen 0001 |
ASPLOS | 2 |
| 2018 | Enabling Efficient Network Service Function Chain Deployment on Heterogeneous Server PlatformabstractNetwork Function Virtualization (NFV) aims to run software-implemented network functions on general hardware such as Commodity Off-the-Shelf (COTS) servers to trade the application-specific performance with generality and re-configurability. Nevertheless, with the wide adoption of general accelerators such as GPU, the researchers seek to boost the performance of software-based network functions while trying to maintain the reusability and programmability in the meantime. The Service Function Chain (SFC) is a key enabler of service flexibility of NFV. The network functions stitch into a chain to provide differentiated services to multi-tenants. However, our characterization results show that existing heterogeneous packet processing frameworks do not handle NFV SFC well since two new overheads, the aggregated processing overheads and co-existence interference overheads, are introduced by SFC.,,,, Motivated by our characterization, we propose NFCompass, a runtime framework that employs SFC re-organization technique and graph-partition based task scheduling technique to conquer the two challenges brought by SFC. By re-organizing the SFC components, the length and complexity of processing paths are reduced and the aggregated overheads are mitigated. By applying the graph-partition based task allocation, better load balance is achieved and the data transfer overheads are considerably reduced. Yang Hu 0001, Tao Li 0006 |
HPCA | 2 |
| 2018 | Towards Efficient Microarchitectural Design for Accelerating Unsupervised GAN-Based Deep LearningabstractRecently, deep learning based approaches have emerged as indispensable tools to perform big data analytics. Normally, deep learning models are first trained with a supervised method and then deployed to execute various tasks. The supervised method involves extensive human efforts to collect and label the large-scale dataset, which becomes impractical in the big data era where raw data is largely un-labeled and uncategorized. Fortunately, the adversarial learning, represented by Generative Adversarial Network (GAN), enjoys a great success on the unsupervised learning. However, the distinct features of GAN, such as massive computing phases and non-traditional convolutions challenge the existing deep learning accelerator designs. In this work, we propose the first holistic solution for accelerating the unsupervised GAN-based Deep Learning. We overcome the above challenges with an algorithm and architecture co-design approach. First, we optimize the training procedure to reduce on-chip memory consumption. We then propose a novel time-multiplexed design to efficiently map the abundant computing phases to our microarchitecture. Moreover, we design high-efficiency dataflows to achieve high data reuse and skip the zero-operand multiplications in the non-traditional convolutions. Compared with traditional deep learning accelerators, our proposed design achieves the best performance (average 4.3X) with the same computing resource. Our design also has an average of 8.3X speedup over CPU and 6.2X energy-efficiency over NVIDIA GPU. Mingcong Song, Jiaqi Zhang 0002, Huixiang Chen 0001, Tao Li 0006 |
HPCA | 4 |
| 2018 | In-Situ AI: Towards Autonomous and Incremental Deep Learning for IoT SystemsabstractRecent years have seen an exploration of data volumes from a myriad of IoT devices, such as various sensors and ubiquitous cameras. The deluge of IoT data creates enormous opportunities for us to explore the physical world, especially with the help of deep learning techniques. Traditionally, the Cloud is the option for deploying deep learning based applications. However, the challenges of Cloud-centric IoT systems are increasing due to significant data movement overhead, escalating energy needs, and privacy issues. Rather than constantly moving a tremendous amount of raw data to the Cloud, it would be beneficial to leverage the emerging powerful IoT devices to perform the inference task. Nevertheless, the statically trained model could not efficiently handle the dynamic data in the real in-situ environments, which leads to low accuracy. Moreover, the big raw IoT data challenges the traditional supervised training method in the Cloud. To tackle the above challenges, we propose In-situ AI, the first Autonomous and Incremental computing framework and architecture for deep learning based IoT applications. We equip deep learning based IoT system with autonomous IoT data diagnosis (minimize data movement), and incremental and unsupervised training method (tackle the big raw IoT data generated in ever-changing in-situ environments). To provide efficient architectural support for this new computing paradigm, we first characterize the two In-situ AI tasks (i.e. inference and diagnosis tasks) on two popular IoT devices (i.e. mobile GPU and FPGA) and explore the design space and tradeoffs. Based on the characterization results, we propose two working modes for the In-situ AI tasks, including Single-running and Co-running modes. Moreover, we craft analytical models for these two modes to guide the best configuration selection. We also develop a novel two-level weight shared In-situ AI architecture to efficiently deploy In-situ tasks to IoT node. Compared with traditional IoT systems, our In-situ AI can reduce data movement by 28-71%, which further yields 1.4X-3.3X speedup on model update and contributes to 30-70% energy saving. Mingcong Song, Kan Zhong, Jiaqi Zhang 0002, Yang Hu 0001, Duo Liu 0002, Weigong Zhang, Jing Wang 0055, Tao Li 0006 |
HPCA | 8 |
| 2018 | Towards Efficient Microarchitecture Design of Simultaneous Localization and Mapping in Augmented Reality EraabstractRecently, augmented reality technologies are debuting to the mainstream markets. Simultaneous Localization and Mapping (SLAM), which serves as the core to drive augmented reality, enables mobile devices to understand "the reality" by recognizing and understanding the surrounding space. However, enabling SLAM on mobile devices still faces challenges due to computation and power limitations. Thus, this paper explores the microarchitecture design space of visual SLAM. First, we conduct a characterization to make the image signal processing stage in the front-end image acquisition stage configurable and explore SLAM's sensitivity to individual stages. Our characterization shows that only demosaicing, gamma compression, and denoising in the image signal processing process have significant influences on SLAM's accuracy. Thus, the image signal processor in SoC in traditional mobile devices can be replaced by an approximation logic to save hardware overhead and energy. Second, we propose a new CGRA architecture, called SL-CGRA, specially tailored to SLAM's workload characteristics. It features a two-level memory design, which includes data-redirection layer support for on-chip memory, and an efficient CGRA memory controller stacked through Through-Via Silicon (TSV) for highly efficient off-chip memory access. Besides, PEs in SL-CGRA support different execution modes to exploit the data-level parallelism and task-level parallelism of SLAM. Our evaluation results show that SL-CGRA achieves good performance and energy efficiency. Huixiang Chen 0001, Yuting Dai, Kan Zhong, Tao Li 0006 |
ICCD | 5 |
| 2018 | Prediction Based Execution on Deep Neural NetworksabstractRecently, deep neural network based approaches have emerged as indispensable tools in many fields, ranging from image and video recognition to natural language processing. However, the large size of such newly developed networks poses both throughput and energy challenges to the underlying processing hardware. This could be the major stumbling block to many promising applications such as self-driving cars and smart cities. Existing work proposes to weed zeros from input neurons to avoid unnecessary DNN computation (zero-valued operand multiplications). However, we observe that many output neurons are still ineffectual even if the zero-removal technique has been applied. These ineffectual output neurons could not pass their values to the subsequent layer, which means all the computations (including zero-valued and non-zero-valued operand multiplications) related to these output neurons are futile and wasteful. Therefore, there is an opportunity to significantly improve the performance and efficiency of DNN execution by predicting the ineffectual output neurons and thus completely avoid the futile computations by skipping over these ineffectual output neurons. To do so, we propose a two-stage, prediction-based DNN execution model without accuracy loss. We also propose a uniform serial processing element (USPE), for both prediction and execution stages to improve the flexibility and minimize the area overhead. To improve the processing throughput, we further present a scale-out design for USPE. Evaluation results over a set of state-of-the-art DNNs show that our proposed design achieves 2.5X speedup and 1.9X energy-efficiency on average over the traditional accelerator. Moreover, by stacking with our design, we can improve Cnvlutin and Stripes by 1.9X and 2.0X on average, respectively. Mingcong Song, Jiechen Zhao 0003, Yang Hu 0001, Jiaqi Zhang 0002, Tao Li 0006 |
ISCA | 5 |
| 2018 | Understanding the Characteristics of Mobile Augmented Reality ApplicationsabstractRecently, augmented reality technologies are debuting to the mainstream markets. Currently, smartphones are still the dominant computing platform for providing mobile augmented reality (MAR) experience to users. Nevertheless, MAR on smartphones faces some key challenges as the rich functionalities increase concerns including battery power drain and thermal dissipation. This paper takes the first step to understand the characteristics of MAR apps from the system and architecture perspectives. We observe that MAR apps exhibit much higher Thread-level Parallelism (TLP) and have a potential to better exploit the big.LITTLE architecture. We further build Multi-comPonent Power and Thermal Analysis Tool (MPPTAT) on Android mobile platform to break down the power consumption and thermal dissipation into hardware components, threads, and phases. Power characterization shows that MAR apps exhibit higher energy consumption compared to other popular mobile non-MAR apps, mainly contributed by the camera. Thread level breakdown shows that the major CPU power consumption is MediaServer, which controls the camera to capture and process multimedia information on smartphones. We also investigate the impact of DVFS, and results indicate that current MAR apps' implementations are mostly CPU intensive and the mobile GPU is not well exploited yet. Moreover, thermal characterization shows that compared to the non-MAR apps, the frequent usage of the camera not only increases the temperature of the camera module but also dissipates the heat to other layers, which directly affect user experience. We believe that our work reveals insights for software and hardware designers to fine-tune augmented reality workloads on mobile platforms. Huixiang Chen 0001, Yuting Dai, Tao Li 0006 |
ISPASS | 5 |
| 2018 | Optimizing RAID/SSD controllers with lifetime extension for flash-based SSD arrayabstractFlash-based SSD RAID arrays are increasingly being deployed in data centers. Compared with HDD arrays, SSD arrays drastically enhance storage density and I/O performance, and reduce power and rack space. Nevertheless, SSDs suffer aging issues. Though prior studies have been conducted to address this disadvantage, effective techniques of RAID/SSD controllers are urgently needed to extend the lifetime of SSD arrays. Zhaoyan Shen, Zili Shao, Tao Li 0006 |
LCTES | 4 |
| 2018 | LerGAN: A Zero-Free, Low Data Movement and PIM-Based GAN ArchitectureabstractAs a powerful unsupervised learning method, Generative Adversarial Network (GAN) plays an important role in many domains such as video prediction and autonomous driving. It is one of the ten breakthrough technologies in 2018 reported in MIT Technology Review. However, training a GAN imposes three more challenges: (1) intensive communication caused by complex train phases of GAN, (2) much more ineffectual computations caused by special convolutions, and (3) more frequent off-chip memory accesses for exchanging inter-mediate data between the generator and the discriminator. In this paper, we propose LerGAN, a PIM-based GAN accelerator to address the challenges of training GAN. We first propose a zero-free data reshaping scheme for ReRAM-based PIM, which removes the zero-related computations. We then propose a 3D-connected PIM, which can reconfigure connections inside PIM dynamically according to dataflows of propagation and updating. Our proposed techniques reduce data movement to a great extent, avoiding I/O to become a bottleneck of training GANs. Finally, we propose LerGAN based on these two techniques, providing different levels of accelerating GAN for programmers. Experiments shows that LerGAN achieves 47.2X, 21.42X and 7.46X speedup over FPGA-based GAN accelerator, GPU platform, and ReRAM-based neural network accelerator respectively. Moreover, LerGAN achieves 9.75X, 7.68X energy saving on average over GPU platform, ReRAM-based neural network accelerator respectively, and has 1.04X energy consuming over FPGA-based GAN accelerator. Haiyu Mao, Mingcong Song, Tao Li 0006, Yuting Dai, Jiwu Shu |
MICRO | 3 |
| 2018 | Start Late or Finish Early: A Distributed Graph Processing System with Redundancy ReductionabstractGraph processing systems are important in the big data domain. However, processing graphs in parallel often introduces redundant computations in existing algorithms and models. Prior work has proposed techniques to optimize redundancies for out-of-core graph systems, rather than distributed graph systems. In this paper, we study various state-of-the-art distributed graph systems and observe root causes for these pervasively existing redundancies. To reduce redundancies without sacrificing parallelism, we further propose SLFE, a distributed graph processing system, designed with the principle of "start late or finish early". SLFE employs a novel preprocessing stage to obtain a graph's topological knowledge with negligible overhead. SLFE's redundancy-aware vertex-centric computation model can then utilize such knowledge to reduce the redundant computations at runtime. SLFE also provides a set of APIs to improve programmability. Our experiments on an 8-machine high-performance cluster show that SLFE outperforms all well-known distributed graph processing systems with the inputs of real-world graphs, yielding up to 75x speedup. Moreover, SLFE outperforms two state-of-the-art shared memory graph systems on a high-end machine with up to 1644x speedup. SLFE's redundancy-reduction schemes are generally applicable to other vertex-centric graph processing systems. Shuang Song 0007, Xu Liu 0001, Qinzhe Wu, Andreas Gerstlauer, Tao Li 0006, Lizy Kurian John |
Proc. VLDB Endow. | 5 |
| 2018 | A Novel ReRAM-Based Processing-in-Memory Architecture for Graph TraversalabstractGraph algorithms such as graph traversal have been gaining ever-increasing importance in the era of big data. However, graph processing on traditional architectures issues many random and irregular memory accesses, leading to a huge number of data movements and the consumption of very large amounts of energy. To minimize the waste of memory bandwidth, we investigate utilizing processing-in-memory (PIM), combined with non-volatile metal-oxide resistive random access memory (ReRAM), to improve both computation and I/O performance. We propose a new ReRAM-based processing-in-memory architecture called RPBFS, in which graph data can be persistently stored and processed in place. We study the problem of graph traversal, and we design an efficient graph traversal algorithm in RPBFS. Benefiting from low data movement overhead and high bank-level parallel computation, RPBFS shows a significant performance improvement compared with both the CPU-based and the GPU-based BFS implementations. On a suite of real-world graphs, our architecture yields a speedup in graph traversal performance of up to 33.8×, and achieves a reduction in energy over conventional systems of up to 142.8×. Zhaoyan Shen, Duo Liu 0002, Zili Shao, H. Howie Huang, Tao Li 0006 |
ACM Trans. Storage | 6 |
| 2018 | A Flattened Metadata Service for Distributed File SystemsabstractKey-Value stores provide scalable metadata service for distributed file systems. However, the metadata's organization itself, which is organized using a directory tree structure, does not fit the key-value access pattern, thereby limiting the performance. To address this issues, we propose a distributed file system with a flattened and fine-grained division metadata service, LocoMeta, to bridge the performance gap between file system metadata and key-value stores. LocoMeta is designed to bridge the gap between file metadata to key-value store with two techniques. First, LocoMeta flattens the directory content and structure, which organizes file and directory index nodes in a flat space while reversely indexing the directory entries. Second, it exploits a fine-grained division method to improve the key-value access performance. Evaluations show that LocoMeta with eight nodes boosts the metadata throughput by five times, which approaches 93 percent throughput of a single-node key-value store, compared to 18 percent in the state-of-the-art IndexFS. Fenlin Liu, Jiwu Shu, Youyou Lu, Tao Li 0006, Yang Hu 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Exploring Customizable Heterogeneous Power Distribution and Management for DatacenterabstractLarge-scale datacenters are facing increasing pressure of capping their carbon emission and power cost. Many leading-edge studies have started to explore server clusters running on multiple power sources. Existing approaches do not sufficiently consider the fine-grained power delivery to satisfy diverse requirements in datacenter, especially in the multi-tenant/colocation datacenter, which may yield low energy utilization. To address the emerging trend and new requirements, this article proposes a novel Datacenter inner Power Switch Network (DiPSN) to improve datacenter power efficiency and user satisfaction. DiPSN is a reconfigurable and easy-to-scale-out power architecture, which enables datacenter to distribute various power sources in a fine-grained manner. Moreover, a tailored machine learning based power source management framework is proposed for DiPSN to dynamically optimize user customized performance metrics and maximize datacenter revenue. Compared with conventional single-switch power distribution system, our DiPSN can be configured to improve solar energy utilization by 39.6 percent, reduce utility power cost by 11.1 percent and improve workload performance by 33.8 percent. Meanwhile, our design can extend battery lifetime by 9.3 percent. This work could provide valuable guidelines for designing heterogeneous power distribution architecture and management methodology in datacenters for improving user-customizable efficiency, sustainability and economy. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Tao Li 0006, Nanning Zheng 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Towards Memory-Efficient Allocation of CNNs on Processing-in-Memory ArchitectureabstractConvolutional neural networks (CNNs) have been successfully applied in artificial intelligent systems to perform sensory processing, sequence learning, and image processing. In contrast to conventional computing-centric applications, CNNs are known to be both computationally and memory intensive. The computational and memory resources of CNN applications are mixed together in the network weights. This incurs a significant amount of data movement, especially for high-dimensional convolutions. The emerging Processing-in-Memory (PIM) alleviates this memory bottleneck by integrating both processing elements and memory into a 3D-stacked architecture. Although this architecture can offer fast near-data processing to reduce data movement, memory is still a limiting factor of the entire system. We observe that an unsolved key challenge is how to efficiently allocate convolutions to 3D-stacked PIM to combine the advantages of both neural and computational processing. This paper presents MemoNet, a memory-efficient data allocation strategy for convolutional neural networks on 3D PIM architecture. MemoNet offers fine-grained parallelism that can fully exploit the computational power of PIM architecture. The objective is to capture the characteristics of neural network applications and perfectly match the underlining hardware resources provided by PIM, resulting in a hardware-independent design to transparently allocate data. We formulate the target problem as a dynamic programming model and present an optimal solution. To demonstrate the viability of the proposed MemoNet, we conduct a set of experiments using a variety of realistic convolutional neural network applications. The extensive evaluations show that, MemoNet can significantly improve the performance and the cache utilization compared to representative schemes. Yi Wang 0003, Weixuan 'Vincent' Chen, Jing Yang 0018, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Towards "Full Containerization" in Containerized Network Function VirtualizationabstractWith exploding traffic stuffing existing network infra-structure, today's telecommunication and cloud service providers resort to Network Function Virtualization (NFV) for greater agility and economics. Pioneer service provider such as AT&T proposes to adopt container in NFV to achieve shorter Virtualized Network Function (VNF) provisioning time and better runtime performance. However, we characterize typical NFV work-loads on the containers and find that the performance is unsatisfactory. We observe that the shared host OS net-work stack is the main bottleneck, where the traffic flow processing involves a large amount of intermediate memory buffers and results in significant last level cache pollution. Existing OS memory allocation policies fail to exploit the locality and data sharing information among buffers. In this paper, we propose NetContainer, a software framework that achieves fine-grained hardware resource management for containerized NFV platform. NetContainer employs a cache access overheads guided page coloring scheme to coordinately address the inter-flow cache access overheads and intra-flow cache access overheads. It maps the memory buffer pages that manifest low cache access overheads (across a flow or among the flows) to the same last level cache partition. NetContainer exploits a footprint theory based method to estimate the cache access overheads and a Min-Cost Max-Flow model to guide the memory buffer mappings. We implement the NetContainer in Linux kernel and extensively evaluate it with real NFV workloads. Exper-imental results show that NetContainer outperforms conventional page coloring-based memory allocator by 48% in terms of successful call rate. Yang Hu 0001, Mingcong Song, Tao Li 0006 |
ASPLOS | 3 |
| 2017 | SmartSwap: High-Performance and User Experience Friendly Swapping in Mobile SystemsabstractWith high-performance mobile processors and large main memory, smartphones are now integrated with more applications and richer functionality than ever. This poses larger memory and storage space demands, however, most mobile systems have limited memory space, which in turn affects user satisfaction. For example, application response time could become longer due to limited memory capacity. Swapping is an effective way to extend memory capacity, but often lead to poor performance in smartphones. Duo Liu 0002, Kan Zhong, Jinting Ren, Tao Li 0006 |
DAC | 5 |
| 2017 | Reducing the "Tax" of Reliability: A Hardware-Aware Method for Agile Data Persistence in Mobile DevicesabstractNowadays, mobile devices are pervasively used by almost everyone. The majority of mobile devices use embedded-Multi Media Cards (eMMC) as storage. However, the crash-proof mechanism of existing I/O stack has not fully exploited the features of eMMC. In some real usage scenarios, the legacy data persistence procedure may dramatically degrade performance of the system. In response to this, this paper exploits the hardware features of eMMC to improve the efficiency of data persistence while preserving the reliability of current mobile systems. We characterize the existing data persistence scheme and observe that the hardware-agnostic design generates excessive non-critical data and adds expensive barriers in data persistence paths. We alleviate these overheads by leveraging eMMC features. Based on evaluations on real systems, our optimizations achieve 5%-31% performance improvement across a wide range of mobile apps. Huixiang Chen 0001, Tao Li 0006 |
DSN | 3 |
| 2017 | Towards Pervasive and User Satisfactory CNN across GPU MicroarchitecturesabstractAccelerating Convolutional Neural Networks (CNNs) on GPUs usually involves two stages: training and inference. Traditionally, this two-stage process is deployed on high-end GPU-equipped servers. Driven by the increase in compute power of desktop and mobile GPUs, there is growing interest in performing inference on various kinds of platforms. In contrast to the requirements of high throughput and accuracy during the training stage, end-users will face diverse requirements related to inference tasks. To address this emerging trend and new requirements, we propose Pervasive CNN (P-CNN), a user satisfaction-aware CNN inference framework. P-CNN is composed of two phases: cross-platform offline compilation and run-time management. Based on users' requirements, offline compilation generates the optimal kernel using architecture-independent techniques, such as adaptive batch size selection and coordinated fine-tuning. The runtime management phase consists of accuracy tuning, execution, and calibration. First, accuracy tuning dynamically identifies the fastest kernels with acceptable accuracy. Next, the run-time kernel scheduler partitions the optimal computing resource for each layer and schedules the GPU thread blocks. If its accuracy is not acceptable to the end-user, the calibration stage selects a slower but more precise kernel to improve the accuracy. Finally, we design a user satisfaction metric for CNNs to evaluate our Pervasive deign. Our evaluation results show P-CNN can provide the best user satisfaction for different inference tasks. Mingcong Song, Yang Hu 0001, Huixiang Chen 0001, Tao Li 0006 |
HPCA | 4 |
| 2017 | A Fast Heuristic Attribute Reduction Algorithm Using SparkabstractEnergy data, which consists of energy consumption statistics and other related data in green data centers, grows dramatically. The energy data has great value, but many attributes within it are redundant and unnecessary. Thus attribute reduction for the energy data has been conceived as a critical step. However, many existing attribute reduction algorithms are often computationally time-consuming. To address these issues, we extend the methodology of rough sets to construct data center energy consumption knowledge representation system. By taking good advantage of in-memory computing, an attribute reduction algorithm for energy data using Spark is proposed. In this algorithm, we use a heuristic formula for measuring the significance of attribute to reduce search space, and an efficient algorithm for simplifying energy consumption decision table, which further improve the computation efficiency. The experimental results show the speed of our algorithm gains up to 0.28X performance improvement over the traditional attribute reduction algorithm using Spark. Mincheng Chen, Jingling Yuan, Lin Li 0001, Dongling Liu, Tao Li 0006 |
ICDCS | 5 |
| 2017 | Complete Tolerance Relation Based Filling Algorithm Using SparkabstractWith the advent of cloud computing, renewable energy is integrated into data center power supply systems increasingly. The power statistics collection may not be available due to the instability of renewable energy, which results in incomplete data. The incomplete energy data will significantly disturb the management of data centers. We further propose a filling algorithm based on complete tolerance class. The algorithm expands the traditional tolerance relation, and fills the missing values of the energy data, which ensures the data integrity. By taking good advantage of in-Memory Computing, We further parallelize and optimize our algorithm using Spark. The experiment results demonstrate that our algorithm outperforms other general filling algorithms in terms of filling accuracy. The proposed algorithm also shows good performance as the missing rate rises up. Jingling Yuan, Yao Xiang, Xian Zhong, Mincheng Chen, Tao Li 0006 |
ICDCS | 5 |
| 2017 | GaaS workload characterization under NUMA architecture for virtualized GPUabstractGraphics-as-a-service (GaaS) is gaining popularity in cloud computing community. There is an emerging trend of running GaaS workload using virtualized GPU in current data center deployment. This paper provides a detailed characterization of GaaS workload under virtualized GPU NUMA environment, and found that: (1) GaaS workloads exhibit different behavior with GPGPU workloads by having more frequent real-time data exchange between CPU and GPU; (2) GaaS workloads have no NUMA overhead, whether considering the influence of remote memory access or the resource contention of CPU uncore. We also test the performance and power tradeoff among the frequency scaling of CPU clock, GPU core clock, and GPU memory clock. Characterization results show that (1) ondemand CPU frequency scaling achieves the best balance between performance and power consumption; (2) GaaS workloads are GPU-computation intensive. GPU memory frequency can be set lower to save energy with little performance sacrifice. Huixiang Chen 0001, Yang Hu 0001, Mingcong Song, Tao Li 0006 |
ISPASS | 5 |
| 2017 | [keynote 2] Paving the way towards nfv: An infrastructure based approachabstractSummary form only give, as follows. Provides an abstract of the keynote presentation and a brief professional biography of the presenter. The complete presentation was not made available for publication as part of the conference proceedings. Network Function Virtualization (NFV) is an initiative driven by the largest service providers (SP) to increase the use of virtualization and integrate intelligence into their network infrastructures. NFV leverages virtualization technology and operates network functions on standard servers to fundamentally decouple the customized and inflexible network hardware. Today, NFV provides a plethora of virtual network functions (VNFs), including gateways, mobile core, deep packet inspection (DPI), security, routing, and traffic management that can be combined to deliver the dynamic customized network service chains. Our exploration on both VM-based and container-based VNFs indicates that the service chaining can pose various challenges to NFV implementation on current commercial off-the-shelf (COTS) server. First, we observe that the NFV packet processing on COTS server exhibits a unique processing pattern - heterogeneous software pipeline, where the NFV traffic flows are processed by a variety of software components sequentially. On modern NUMA-based architecture, the end-to-end performance of NFV traffic flows can be severely affected by placing these heterogeneous software components inappropriately. We develop a thread scheduling mechanism that collaboratively places threads of heterogeneous software pipeline to minimize the end-to-end performance slowdown for NFV traffic flows. In the second part, we characterize the light-weight container-based VNF, which is expected to achieve shorter VNF provisioning time and lower resource overheads. However, we observe that the traffic flow processing in the shared host OS network stack involves a large amount of intermediate memory buffers and results in significant last level cache pollution. We propose NetContainer, a framework that achieves a fine-grained hardware resource management for containerized NFV platform. NetContainer exploits a cache access overhead guided page coloring technique to coordinately manage the inter/intra-flow cache access overheads. Tao Li 0006 |
NAS | 1 |
| 2017 | LocoFS: a loosely-coupled metadata service for distributed file systemsabstractKey-Value stores provide scalable metadata service for distributed file systems. However, the metadata's organization itself, which is organized using a directory tree structure, does not fit the key-value access pattern, thereby limiting the performance. To address this issue, we propose a distributed file system with a loosely-coupled metadata service, LocoFS, to bridge the performance gap between file system metadata and key-value stores. LocoFS is designed to decouple the dependencies between different kinds of metadata with two techniques. First, LocoFS decouples the directory content and structure, which organizes file and directory index nodes in a flat space while reversely indexing the directory entries. Second, it decouples the file metadata to further improve the key-value access performance. Evaluations show that LocoFS with eight nodes boosts the metadata throughput by 5 times, which approaches 93% throughput of a single-node key-value store, compared to 18% in the state-of-the-art IndexFS. Youyou Lu, Jiwu Shu, Yang Hu 0001, Tao Li 0006 |
SC | 5 |
| 2017 | Octopus: an RDMA-enabled Distributed Persistent Memory File System
Youyou Lu, Jiwu Shu, Youmin Chen, Tao Li 0006 |
USENIX ATC | 4 |
| 2017 | Complete tolerance relation based parallel filling for incomplete energy big data
Jingling Yuan, Mincheng Chen, Tao Li 0006 |
Knowl. Based Syst. | 4 |
| 2017 | Achieving Versatile and Simultaneous Cache Optimizations With Nonvolatile SRAMabstractThe efficiency of caches plays a vital role in microprocessors. In this paper, we introduce a novel and flexible cache substrate, which integrates nonvolatile memory devices into the standard SRAM cells. By allowing this nonvolatile SRAM (NV-SRAM) cell to store inconsistent data between SRAM portion and NV portion, we show that the proposed NV2-SRAM cache not only provides enriched functionalities, but also allows simultaneous multiple optimizations. For example, the NV2-SRAM cache can reduce cache misses caused by context-switching and improve the performance by 15%. It can also save up to 67% energy over the SRAM-based cache, outperforming the drowsy cache in terms of both power efficiency and reliability. Moreover, the proposed cache architecture can be used to improve the performance of prefetching by 10%. Comparing with a conventional cache (equipped with a victim buffer) that occupies the same die area, the NV2-SRAM cache gains an 11% performance benefit. To achieve simultaneous optimizations, we propose architecture and OS support to optimize the cache power, performance and reliability concurrently on multicore-based systems. Rui Wang 0014, Dan Jia, Tao Li 0006, Depei Qian 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Oasis: Scaling Out Datacenter Sustainably and EconomicallyabstractAs big data applications proliferate, datacenters today are increasingly looking to adopt a scale-out model. Nevertheless, power capacity has become an important bottleneck that restricts horizontal scaling of servers, especially in datacenters that oversubscribe power infrastructure. When a datacenter hits its ceiling for power provisioning, conventionally the owner has to either build another facility or upgrade existing infrastructure-both approaches add huge cost, require significant time, and can further increase carbon footprint. This paper proposes Oasis, a novel datacenter expansion strategy that enables power-/carbon- constrained servers to scale out economically and sustainably. The basic structure of Oasis, called Oasis Node, naturally supports incremental capacity expansion with near-zero environmental impact since it leverages modular solar panels and distributed battery systems to power newly added servers. To optimize the operation of newly added nodes, we further propose a management framework called Ozone. It allows Oasis to jointly perform power supply switching and server speed scaling to improve efficiency locally and globally. We implement a prototype of Oasis and use it as a research platform for evaluating the design tradeoffs of green scale-out datacenters. With Oasis, a green datacenter could gradually double its capacity with near-oracle performance, extended battery lifetime, and 26 percent cost savings. Chao Li 0009, Yang Hu 0001, Juncheng Gu, Jingling Yuan, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | Managing Battery Aging for High Energy Availability in Green DatacentersabstractEnergy storage devices (ESD), such as UPS batteries, have been repurposed in datacenter as a promising tuning knob for peak power shaving and power cost reducing. However, batteries progressively aging due to irregular usage patterns, which result in less effective capacity and even pose serious threat to server availability. Nevertheless, prior proposals largely ignore the aging issues of battery which may lead to low energy availability for datacenter servers. To fill this critical void, we thoroughly investigate battery aging on a heavily instrumented prototype system over an observation period of ten months. We propose Battery Anti-Aging Treatment Plus (BAAT-P), a novel power delivery architecture included aging management algorithms from the perspective of computing system to hide, reduce, mitigate and plan the battery aging effects for high energy availability in datacenter. Our techniques exploit diverse battery aging mechanisms and dynamic aging management algorithms to provide system-level availability guarantee for datacenter. We evaluate the BAAT-P design with a real prototype. Compared with a battery powered datacenter without aging management policies, the results show that BAAT-P can extend battery lifetime by 72 percent, reduce battery cost by 33 percent and effectively improve energy availability for datacenter servers while maintaining workload performance for the performance critical workloads. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | On the Implication of NTC versus Dark Silicon on Emerging Scale-Out Workloads: The Multi-Core Architecture PerspectiveabstractThe end of Dennard's scaling poses computer systems, especially the datacenters, in front of both power and utilization walls. One possible solution to combat the power and utilization walls is dark silicon where transistors are under-utilized in the chip, but this will result in a diminishing performance. Another solution is Near-Threshold Voltage Computing (NTC) which operates transistors in the near-threshold region and provides much more flexible tradeoffs between power and performance. However, prior efforts largely focus on a specific design option based on the legacy desktop applications, therefore, lacking comprehensive analysis of emerging scale-out applications with multiple design options when dark silicon and/or NTC are/is applied. In this paper, we characterize different perspectives including performance, energy efficiency and reliability in the context of NTC/dark silicon cloud processors running emerging scale-out workloads on various architecture designs. We find NTC is generally an effective way to alleviate the power challenge over scale-out applications compared with dark silicon, it can improve performance by 1.6X, energy efficiency by 50 percent and the reliability problem can be relieved by ECC. Meanwhile, we also observe tiled-OoO architecture improves the performance by 20~370 percent and energy efficiency by 40~600 percent over alternative architecture designs, making it a preferable design paradigm for scale-out workloads. We believe that our observations will provide insights for the design of cloud processors under dark silicon and/or NTC. Jing Wang 0055, Xin Fu 0001, Weigong Zhang, Keni Qiu, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | VarCatcher: A Framework for Tackling Performance Variability of Parallel Workloads on Multi-CoreabstractThe non-deterministic nature of multi-threaded workloads running on multi-core platforms often leads to notable performance variability from run to run. Such variability makes experimental results prone to misinterpretations or misguided claims. To deal with such variability, statistical inference methods are usually used to summarize the experimental results with certain confidence levels by running the experiments or measurements a large number of times. However, such statistical results are often too vague or too simplistic. They are not sufficient to help users understand the causes of such variability, and allow more in-depth analysis on the results or reproduce the results for validation during design space exploration. To allow better analyzability and reproducibility, we propose a framework to tackle such variability, called VarCatcher. The key to VarCatcher is to characterize a parallel execution using Parallel Characteristics Vector (PCV). A clustering-based approach is then used to group runs with similar execution characteristics that can later be used to analyze results in-depth, to customize different evaluation strategies, reproduce the result for variability, to determine the impact of features, or to assist performance diagnosis. We have built a prototype of VarCatcher that includes a user-level toolset for runtime monitoring and measurements using the Intel Processor Trace feature on commodity Intel processors as well as an architecture extension with very low runtime overheads (around 3 and 0.01 percent accordingly). Several case studies confirm that VarCatcher enables several appealing features such as in-depth result analysis, customized evaluation strategies, and reproducibility. Xiaofeng Ji, Shiqiang Yu, Haibo Chen 0001, Tao Li 0006, Pen-Chung Yew, Wenyun Zhao |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | Leveraging Time Prediction and Error Compensation to Enhance the Scalability of Parallel Multi-Core SimulationsabstractDue to synchronization overhead, it is challenging to apply the parallel simulation technique of multi-core processors at larger scales. Although the use of lax synchronization schemes could reduce overhead and balance the load between synchronous points, it introduces timing error and deteriorates simulation accuracy. Through observing the propagation paths of errors, we find that these paths always concentrate on some pivotal events. Based on the observation, we design a delay-calibration mechanism to alleviate errors. We decouple the timing and functional processes of the pivotal events, leveraging prediction technique of delays to connect two categories of the processes. Errors are traced throughout the timing processes of the pivotal events, and are deducted from the predicted delays before the delays are consumed by the functional processes. Therefore, through cleaning the errors at the successive pivot events, the mechanism decreases the simulated time deviations efficiently. Since the prediction and error deduction processes do not have any constraint on synchronizations, our approach largely maintains the scalability of lax synchronization schemes. Furthermore, our proposal is orthogonal to other parallel simulation techniques and can be used in conjunction with them. Experimental results show that error compensation improves the accuracy of lax synchronized simulations by 68 percent and achieves 97.8 percent accuracy when combined with an enhanced lax synchronization. Junmin Wu, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Bridging the Semantic Gaps of GPU Acceleration for Scale-out CNN-based Big Data Processing: Think Big, See SmallabstractConvolutional Neural Networks (CNNs) have substantially advanced the state-of-the-art accuracies of object recognition, which is the core function of a myriad of modern multimedia processing techniques such as image/video processing, speech recognition, and natural language processing. GPU-based accelerators gained increasing attention because a large amount of highly parallel neurons in CNN naturally matches the GPU computation pattern. In this work, we perform comprehensive experiments to investigate the performance bottlenecks and overheads of current GPU acceleration platform for scale-out CNN-based big data processing. Mingcong Song, Yang Hu 0001, Chao Li 0009, Huixiang Chen 0001, Jingling Yuan, Tao Li 0006 |
PACT | 7 |
| 2016 | MCSSim: A memory channel storage simulatorabstractRecently, NVDIMM (Non-Volatile Dual In-line Memory Module) is being widely supported by leading hardware design companies, such as IBM. Nevertheless, existing efforts largely focus on NVDIMM specification and fabrication issues, and the potential performance gains brought by NVDIMM are not fully investigated. In this paper, we present a NVDIMM-based simulator called MCSSim to help study the memory channel storage techniques. MCSSim is a cycle-accurate simulator that is elaborated with the consideration of differences between the memory channel interface and the NAND flash memory features. MCSSim is also implemented with the DRAMSim2 [31] simulator thus enabling the simulation of a variety of hybrid memory systems by combining of DRAM DIMM and NVDIMM. We have done some experiments with MCSSim, and the experimental results show the effectiveness of the proposed simulator. Renhai Chen, Zili Shao, Chia-Lin Yang, Tao Li 0006 |
ASP-DAC | 4 |
| 2016 | Refresh-aware loop scheduling for high performance low power volatile STT-RAMabstractThe highlighted advantages of low leakage power, high storage density and immunity to electronic magnetic radiation make STT-RAM a promising candidate to build cache, SPM or main memory in embedded systems. However, write operations on STT-RAM have considerably longer latency and higher energy consumption than conventional SRAM. To solve this problem, researchers have proposed to relax STT-RAM's non-volatility and to have it work in a fast and low power mode. Under this volatile mode, refresh operations are needed to guarantee data correctness if their lifespan is larger than the retention time. It is observed that this refresh overhead is significant for data in stencil loops with the characteristic of constant read and write dependencies. This paper proposes a loop scheduling technique which can traverse loops in a new direction such that data lifespan can be greatly shortened. Therefore, overall refresh overhead can be efficiently mitigated so as to improve performance and reduce power consumption. The experimental results indicate that access latency and dynamic energy can be improved by 21.4~96.0% and 22.0~95.5% respectively by the proposed scheduling scheme. Keni Qiu, Junpeng Luo, Zhiyao Gong, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Tao Li 0006, Chun Jason Xue |
ICCD | 7 |
| 2016 | An adaptive Non-Uniform Loop Tiling for DMA-based bulk data transfers on many-core processorabstractMesh Network-on-Chip (NoC) is a key fabric to interconnect many cores with desirable scalability, reliability and interoperability. We observe that DMA-based bulk data block transfer exhibits non-negligible NoC latency due to heavy congestions. Loop tiling is an effective way to partition data space for SPM+DMA-based data block transfer. Nevertheless, we observe that the unbalanced NoC latency can degrade the effectiveness of loop tiling in a uniform fashion. In this paper, we propose a NoC-aware Non-Uniform Loop Tiling (NULT) scheme to improve DMA performance. A NULT framework is built on the proposed model to adaptively hide DMA latency into computation time and reduce the overall execution time. The framework first groups cores into different families taking into account their distance-to-data in NoC. Then a heuristic method is presented to solve the near optimal tiling factors for each core family. In this way, different core families are assigned non-uniform tiling sizes. We evaluate the NULT scheme on the NIRGAM platform. Compared to the traditional uniform tiling approach, the proposed NULT technique shows more benefit to overlap memory access time and computation time and thus reduce the overall execution time of a loop nest. Keni Qiu, Yuanhui Ni, Weigong Zhang, Jing Wang 0055, Chun Jason Xue, Tao Li 0006 |
ICCD | 7 |
| 2016 | Exploring Variation-Aware Fault-Tolerant Cache under Near-Threshold ComputingabstractNear threshold voltage computing enables transistor voltage scaling to continue with Moore's Law projection and dramatically improves power and energy efficiency. However, reducing the supply voltage to near-threshold level significantly increases the susceptibility of on-chip caches to process variations, leading to the high error rate. Most existing fault-tolerant schemes significantly sacrifice cache capacity and performance. In this paper, we propose a novel fault-tolerant cache architecture at near-threshold computing, which is suitable for high error rate memories. We first propose a variation-aware skewed-associative cache, and then redirect the faulty blocks to the error-free blocks based on it to explore the fault-tolerance cache design. Unlike previous cache reconfiguration schemes for the fault tolerance, our cache design does not need to sacrifice or disable any fault-free blocks to form a completely functional set. We use all error-free blocks and have the least cache capacity waste. More importantly, since the aging impact could also cause cell failures, our skewed cache takes the aggregated process variation and aging impact into the consideration. Last but not least, our skewed cache design avoids the complex remapping from faulty blocks to the error-free blocks and minimizes the hardware overheads. Our evaluation results show that our variation-aware fault-tolerant cache design exhibits strong capability to tolerate the high error rate, and more excitingly, its effectiveness on reducing the cache miss rate and improving the performance is even more obvious as the supply voltage scales down to the near-threshold region. Jing Wang 0055, Yanjun Liu 0005, Weigong Zhang, Kezhong Lu, Keni Qiu, Xin Fu 0001, Tao Li 0006 |
ICPP | 7 |
| 2016 | HOPE: Enabling Efficient Service Orchestration in Software-Defined Data CentersabstractThe functional scope of today's software-defined data centers (SDDC) has expanded to such an extent that servers face a growing amount of critical background operational tasks like load monitoring, logging, migration, and duplication, etc. These ancillary operations, which we refer to as management operations, often nibble the stringent data center power envelope and exert a tremendous amount of pressure on front-end user tasks. However, existing power capping, peak shaving, and time shifting mechanisms mainly focus on managing data center power demand at the "macro level" -- they do not distinguish ancillary background services from user tasks, and therefore often incur significant performance degradation and energy overhead. Yang Hu 0001, Chao Li 0009, Longjun Liu, Tao Li 0006 |
ICS | 4 |
| 2016 | Towards an Adaptive Multi-Power-Source DatacenterabstractBig data and cloud computing are accelerating the capacity growth of datacenters all over the world. Their energy costs and environmental issues have pushed datacenter operators to explore and integrate alternative energy sources, such as various renewable energy supplies and energy storage devices. Designing datacenters powered by multi-power supplies in the smart grid environment is becoming a promising trend in the next few decades. However, gracefully provisioning various power sources and efficiently manage them in datacenter is a significant challenge. Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Nanning Zheng 0001, Tao Li 0006 |
ICS | 6 |
| 2016 | Scheduling Tasks with Mixed Timing Constraints in GPU-Powered Real-Time SystemsabstractDue to the cost-effective, massive computational power of graphics processing units (GPUs), there is a growing interest of utilizing GPUs in real-time systems. For example GPUs have been applied to automotive systems to enable new advanced and intelligent driver assistance technologies, accelerating the path to self-driving cars. In such systems, GPUs are shared among tasks with mixed timing constraints: real-time (RT) tasks that have to be accomplished before specified deadlines, and non-real-time, best-effort (BE) tasks. In this paper, (1) we propose resource-aware non-uniform slack distribution to enhance the schedulability of RT tasks (the total amount of work of RT tasks whose deadlines can be satisfied on a given amount of resources) in GPU-enabled systems; (2) we propose deadline-aware dynamic GPU partitioning to allow RT and BE tasks to run on a GPU simultaneously, such that BE tasks are not blocked for a long time. Rui Wang 0014, Tao Li 0006, Mingcong Song, Lan Gao 0004, Zhongzhi Luan, Depei Qian 0001 |
ICS | 3 |
| 2016 | Bridging the I/O performance gap for big data workloads: A new NVDIMM-based approachabstractThe long I/O latency posts significant challenges for many data-intensive applications, such as the emerging big data workloads. Recently, the NVDIMM (Non-Volatile Dual In-line Memory Module) technologies provide a promising solution to this problem. By employing non-volatile NAND flash memory as storage media and connecting them via DIMM (Dual Inline Memory Module) slots, the NVDIMM devices are exposed to memory bus so the access latencies due to going through I/O controllers can be significantly mitigated. However, placing NVDIMM on the memory bus introduces new challenges. For instance, by mixing I/O and memory traffic, NVDIMM can cause severe performance degradation on memory-intensive applications. Besides, there exists a speed mismatch between fast memory access and slow flash read/write operations. Moreover, garbage collection (GC) in NAND flash may cause up to several millisecond latency. This paper presents novel, enabling mechanisms that allow NVDIMM to more effectively bridge the I/O performance gap for big data workloads. To address the workload heterogeneity challenge, we develop a scheduling scheme in memory controller to minimize the interference between the native and the I/O-derived memory traffic by exploiting both data access criticality and resource utilization. For NVDIMM controller, several mechanisms are designed to better orchestrate traffic between the memory controller and NAND flash to alleviate the speed discrepancy issue. To mitigate the lengthy GC period, we propose a proactive GC scheme for the NVDIMM controller and flash controller to intelligently synchronize and transfer data involving in forthcoming GC operations. We present detailed evaluation and analysis to quantify how well our techniques fit with the NVDIMM design. Our experimental results show that overall the proposed techniques yield 10%~35% performance improvements over the state-of-the-art baseline schemes. Renhai Chen, Zili Shao, Tao Li 0006 |
MICRO | 3 |
| 2016 | Towards efficient server architecture for virtualized network function deployment: Implications and implementationsabstractRecent years have seen a revolution in network infrastructure brought on by the ever-increasing demands for data volume. One promising proposal to emerge from this revolution is Network Functions Virtualization (NFV), which has been widely adopted by service and cloud providers. The essence of NFV is to run network functions as virtualized workloads on commodity Standard High Volume Servers (SHVS), which is the industry standard. However, our experience using NFV when deployed on modern NUMA-based SHVS paints a frustrating picture. Due to the complexity in the NFV data plane and its service function chain feature, modern NFV deployment on SHVS exhibits a unique processing pattern - heterogeneous software pipeline (HSP), in which the NFV traffic flows must be processed by heterogeneous software components sequentially from the NIC to the end re-ceiver. Since the end-to-end performance of flows is cooperatively determined by the performance of each processing stage, the resource allocation/mapping scheme in NUMA-based SHVS must consider a thread-dependence scheduling to tradeoff the impact of co-located contention and remote packet transmission. In this paper, we develop a thread scheduling mechanism that collaboratively places threads of HSP to minimize the end-to-end performance slowdown for NFV traffic flow. It employs a dynamic programming-based method to search for the optimal thread mapping with negligible overhead. To serve this mechanism, we also develop a performance slowdown estimation model to accurately estimate the performance slowdown at each stage of HSP. We implement our collaborative thread scheduling mechanism on a real system and evaluate it using real workloads. On average, our algorithm outperforms state-of-the-art NUMA-aware and contention-aware scheduling policies by at least 7% on CPU utilization and 23% on traffic throughput with negligible computational overhead (less than 1 second). Yang Hu 0001, Tao Li 0006 |
MICRO | 2 |
| 2016 | Reducing Synchronization Cost for Single-Level Store in Mobile Systems
Yuanchao Xu 0002, Hu Wan 0001, Keni Qiu, Tao Li 0006, Weigong Zhang |
J. Comput. Sci. Technol. | 4 |
| 2016 | Managing Server Clusters on Renewable Energy MixabstractAs climate change has become a global concern and server energy demand continues to soar, many IT companies have started to explore server clusters running on various renewable energy sources. Existing green data center designs often yield suboptimal performance as they only look at a certain specific type of energy source. This article explores data centers powered by hybrid renewable energy systems. We propose GreenWorks, a framework for HPC data centers running on a renewable energy mix. Specifically, GreenWorks features a cross-layer power management scheme tailored to the timing behaviors and capacity constraints of different energy sources. Using realistic workload traces and renewable energy data, we show that GreenWorks could provide a near-optimal workload performance (within 3% difference) on average. It can also reduce the worst-case performance degradation by 43% compared to the state-of-the-art design. Moreover, the performance improvements are based on carbon-neutral operations and are not at the cost of significant efficiency degradation and reduced battery lifecycle. Our technique becomes more efficient when servers become more energy proportional and can effectively handle the ever-increasing depth of renewable power penetration in green data centers. Chao Li 0009, Rui Wang 0014, Depei Qian 0001, Tao Li 0006 |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2016 | WBSP: A Novel Synchronization Mechanism for Architecture Parallel SimulationabstractParallelization is an efficient approach to accelerate multi-core, multi-processor and cluster architecture simulators. Nevertheless, frequent synchronization can significantly hinder the performance of a parallel simulator. A common practice in alleviating synchronization cost is to relax synchronization using lengthened synchronous steps. However, as a side effect, simulation accuracy deteriorates considerably. Through analyzing various factors contributing to the causality error in lax synchronization, we observe that a coherent speed across all nodes is critical to achieve high accuracy. To this end, we propose wall-clock based synchronization (WBSP), a novel mechanism that uses wall-clock time to maintain a coherent running speed across the different nodes by periodically synchronizing simulated clocks with the wall clock within each lax step. Our proposed method only results in a modest precision loss while achieving performance close to lax synchronization. We implement WBSP in a many-core parallel simulator and a cluster parallel simulator. Experimental results show that at a scale of 32-host threads, it improves the performance of the many-core simulator by 4.3χ on average with less than a 5.5 percent accuracy loss compared to the conservative mechanism. On the cluster simulator with 64 nodes, our proposed scheme achieves an 8.3χ speedup compared to the conservative mechanism while yielding only a 1.7 percent accuracy loss. Meanwhile, WBSP outperforms the recent proposed adaptive mechanism on simulations that exhibit heavy traffic. Junmin Wu, Tao Li 0006, Xiufeng Sui |
IEEE Trans. Computers | 3 |
| 2016 | RE-UPS: an adaptive distributed energy storage system for dynamically managing solar energy in green datacenters
Longjun Liu, Hongbin Sun 0001, Chao Li 0009, Yang Hu 0001, Jingmin Xin, Nanning Zheng 0001, Tao Li 0006 |
J. Supercomput. | 7 |
| 2015 | BAAT: Towards Dynamically Managing Battery Aging in Green DatacentersabstractEnergy storage devices (batteries) have shown great promise in eliminating supply/demand power mismatch and reducing energy/power cost in green datacenters. These important components progressively age due to irregular usage patterns, which result in less effective capacity and even pose serious threat to server availability. Nevertheless, prior proposals largely ignore the aging issue of batteries or simply use ad-hoc discharge capping to extend their lifetime. To fill this critical void, we thoroughly investigate battery aging on a heavily instrumented prototype over an observation period of six months. We propose battery anti-aging treatment (BAAT), a novel framework for hiding, reducing, and planning the battery aging effects. We show that BAAT can extend battery lifetime by 69%. It enables datacenters to maximally utilize energy storage resources to enhance availability and boost performance. Moreover, it reduces 26% battery cost and allows datacenters to economically scale in the big data era. Longjun Liu, Chao Li 0009, Hongbin Sun 0001, Yang Hu 0001, Juncheng Gu, Tao Li 0006 |
DSN | 6 |
| 2015 | Understanding the virtualization "Tax" of scale-out pass-through GPUs in GaaS clouds: An empirical studyabstractPass-through techniques enable virtual machines to directly access hardware GPU resources in an exclusive mode, delivering extraordinary graphics performance for client users in GaaS clouds. However, the virtualization overheads of pass-through GPUs may decrease the frame rate of graphics workloads by reducing the occupancy rate of the GPU working queue. In this work, we make the first attempt to characterize pass-through GPUs running in different consolidation scenarios and uncover the root causes of these overheads. Towards this end, we set up state-of-the-art empirical platforms equipped with NVIDIA GRID GPUs and execute graphics intensive workloads running in GaaS clouds. We first demonstrate the existence of virtualization overheads, which can slow down the GPU command generation rate. Compared with a bare-metal system, the performance of pass-through GPUs degrades 9.0% and 21.5% under a single VM and 8-VMs respectively. We analyze the workflow of Windows display driver model and VMEXIT events distribution and identify four factors (i.e. HLT instruction and idle domain, external interrupt delivery, IOMMU, and memory subsystem) that contribute to the performance degradation. Our evaluation results show that: (1) the VM-VMM context switch caused by a HLT instruction and wake-up interrupt injection of an idle domain result in 66. 7% idle time for a single pass-through GPU; (2) the external interrupt delivery and tasklet processing cause additional overheads. When 8 VMs are consolidated, the interrupt delivery processing time and interrupt frequency rise 30.7% and 127.3%, respectively; (3) the existing IOMMU design scales well with pass-through GPUs; and (4) interactions of domain guest's software stacks impact the hardware prefetching mechanism so that it fails to compensate the rapidly growing LLC miss rate when more pass-through GPU VMs are added. To the best of our knowledge, this is the first work that characterizes pass-through GPU virtualization overheads and underlying reasons. This study highlights valuable insights for improving the performance of future virtualized GPU systems. Ming Liu 0006, Tao Li 0006, Neo Jia, Andy Currid, Vladimir Troy |
HPCA | 2 |
| 2015 | Optimization of Resource Allocation and Energy Efficiency in Heterogeneous Cloud Data CentersabstractPerformance and energy efficiency are major concerns in cloud computing data centers. More often, they carry conflicting requirements making optimization a challenge. Further complications arise when heterogeneous hardware and data center management technologies are combined. For example, heterogeneous hardware such as General Purpose Graphics Processing Units (GPGPUs) improve performance at the cost of greater power consumption while virtualization technologies improve resource management and utilization at the cost of degraded performance. In this paper, we focus on exploiting heterogeneity introduced by GPUs to reduce power budget requirements for servers while maintaining performance. To maintain or improve overall server performance at reduced power budget, we propose two enhancements: (a) We borrow power from co-located multithreaded virtual machines (VMs) and reallocate it to GPU VMs. (b) To compensate multi-threaded VMs and re-boost their performance, we propose to borrow virtual computing resources from GPU VMs and reallocate them to CPU VMs. Combining the two techniques minimizes server power budget while maintaining overall server performance. Our results show that server power budget can be reduced by almost 18% at the average cost of 13% performance degradation per virtual machine. In addition, reallocating virtual resources improves the performance of multi-threaded applications by 30% without affecting GPU applications. Combining both techniques reduces server energy consumption by 47 % with minimum performance degradation. Amer Qouneh, Ming Liu 0006, Tao Li 0006 |
ICPP | 3 |
| 2015 | Towards Lightweight and Swift Storage Resource Management in Big Data Cloud EraabstractWorkload IO behavior in modern data centers is fluctuating and unpredictable due to the rapidly adopted, public cloud environment. Nevertheless, existing storage resource management systems, such as VMware SDRS, are incapable of performing real time policy-based storage management due to the high cost of migrating large size virtual disks. Hence, the traditional storage management schemes become ineffective due to the lack of quick response to the frequent IO bursts and the inaccurate storage latency prediction in the light of a highly fluctuating environment. To address the aforementioned issues, we propose LightSRM, which can work properly in a time-variant cloud environment. To mitigate the storage migration cost, we leverage copy-on-write/read snapshots to redirect the IO requests without moving the virtual disk. To support snapshots in storage management, we also build a performance model specifically for snapshots. We employ exponentially weighted moving average with adjustable sliding window to provide quick and accurate performance prediction. Furthermore, we propose a hybrid management scheme, which can dynamically choose either snapshot or migration for fastest performance tuning. We build our prototype in a QEMU/KVM based virtualized environment. Our empirical evaluation results show that snapshot can redirect IO requests in a faster manner than migration can do when the virtual disk size is large. Besides, snapshot method has less disk performance impact on the applications. By employing hybrid snapshot/migration method, LightSRM yields less overall latency, better load balance, and less IO traffic overhead. Ruijin Zhou, Huixiang Chen 0001, Tao Li 0006 |
ICS | 3 |
| 2015 | Towards sustainable in-situ server systems in the big data eraabstractRecent years have seen an explosion of data volumes from a myriad of distributed sources such as ubiquitous cameras and various sensors. The challenges of analyzing these geographically dispersed datasets are increasing due to the significant data movement overhead, time-consuming data aggregation, and escalating energy needs. Rather than constantly move a tremendous amount of raw data to remote warehouse-scale computing systems for processing, it would be beneficial to leverage in-situ server systems (InS) to pre-process data, i.e., bringing computation to where the data is located. Chao Li 0009, Yang Hu 0001, Longjun Liu, Juncheng Gu, Mingcong Song, Xiaoyao Liang, Jingling Yuan, Tao Li 0006 |
ISCA | 8 |
| 2015 | HEB: deploying and managing hybrid energy buffers for improving datacenter efficiency and economyabstractToday, an increasing number of applications and services are being hosted by large-scale data centers. The massive and irregular load surges challenge data center power infrastructures. As a result, power mismatching between supply and demand has emerged as a crucial issue in modern data centers which are either under-provisioned or powered by intermittent power sources. Recent proposals have employed energy storage devices such as the uninterruptible power supply (UPS) systems to address this issue. However, current approaches lack the capacity of efficiently handling the irregular and unpredictable power mismatches. Longjun Liu, Chao Li 0009, Hongbin Sun 0001, Yang Hu 0001, Juncheng Gu, Tao Li 0006, Jingmin Xin, Nanning Zheng 0001 |
ISCA | 6 |
| 2015 | iConn: A Communication Infrastructure for Heterogeneous Computing ArchitecturesabstractRecently, the graphics processing unit (GPU) has made significant progress as a general-purpose parallel processor. The CPU and GPU cooperate together to solve data-parallel and control-intensive real-world applications in an optimized fashion. For example, emerging heterogeneous computing architectures such as Intel Sandy Bridge and AMD Fusion integrate the functionality of the CPU and GPU in a single die. However, the single-die CPU-GPU heterogeneous computing architecture faces the challenge of tight budget of die area. The conventional homogenous interconnect fails to provide satisfactory performance by fully exploiting the given area budget in the heterogeneous processing era. In this article, we aim to implement an interconnect network within an area budget for a CPU-GPU heterogeneous computing architecture. We propose iConn, a 2D mesh-style on-chip heterogeneous communication infrastructure. In iConn, a set of GPU logical units such as the stream processors, the texture units, and the rendering output units form a computing unit (CU). Differing from conventional homogenous router design, iConn adopts nonuniform on-chip routers in order to meet the unique communication demands from each single CPU and CU. The routers can also dynamically allocate their buffers across all virtual channels (VCs) to meet the latency requirements of CPUs and CUs. Moreover, the memory controller scheduling algorithm is modified from traditional load-over-store scheduling in order to prioritize the traffic. Our simulation results show that iConn improves the performance of CPUs by 23.0% and CUs by 9.4%. Zhong-Qi Li, Nilanjan Goswami, Tao Li 0006 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2015 | GPGPU-MiniBench: Accelerating GPGPU Micro-Architecture SimulationabstractGraphics processing units (GPU), due to their massive computational power with up to thousands of concurrent threads and general-purpose GPU (GPGPU) programming models such as CUDA and OpenCL, have opened up new opportunities for speeding up general-purpose parallel applications. Unfortunately, pre-silicon architectural simulation of modern-day GPGPU architectures and workloads is extremely time-consuming. This paper addresses the GPGPU simulation challenge by proposing a framework, called GPGPU-MiniBench, for generating miniature, yet representative GPGPU workloads. GPGPU-MiniBench first summarizes the inherent execution behavior of existing GPGPU workloads in a profile. The central component in the profile is the Divergence Flow Statistics Graph (DFSG), which characterizes the dynamic control flow behavior including loops and branches of a GPGPU kernel. GPGPU-MiniBench generates a synthetic miniature GPGPU kernel that exhibits similar execution characteristics as the original workload, yet its execution time is much shorter thereby dramatically speeding up architectural simulation. Our experimental results show that GPGPU-MiniBench can speed up GPGPU architectural simulation by a factor of 49× on average and up to 589×, with an average IPC error of 4.7 percent across a broad set of GPGPU benchmarks from the CUDA SDK, Rodinia and Parboil benchmark suites. We also demonstrate the usefulness of GPGPU-MiniBench for driving GPU architecture exploration. Zhibin Yu 0001, Lieven Eeckhout, Nilanjan Goswami, Tao Li 0006, Lizy Kurian John, Hai Jin 0001, Cheng-Zhong Xu 0001, Junmin Wu |
IEEE Trans. Computers | 4 |
| 2015 | Aurora: A Cross-Layer Solution for Thermally Resilient Photonic Network-on-ChipabstractWith silicon optical technology moving toward maturity, the use of photonic networks-on-chip (NoCs) for global chip communication is emerging as a promising solution to the communication requirements of future many core processors. It is expected that photonic NoCs will play an important role in alleviating current power, latency, and bandwidth constraints. However, photonic NoCs are sensitive to ambient temperature variations because their basic constituents, ring resonators, are themselves sensitive to those variations. Since ring resonators are basic building blocks for photonic modulators, switches, multiplexers, and demultiplexers, variations of on-chip temperature pose serious challenges to the proper operation of photonic NoCs. Proposed methods that mitigate the effects of temperature at the device level are either difficult to use in CMOS processes or not suitable for large scale implementation. In this paper, we propose Aurora, a thermally resilient photonic NoC architecture design that supports reliable and low bit error rate (BER) on-chip communications in the presence of large temperature variations. Our proposed architecture leverages cross-layer solutions at the device, architecture, and operating system (OS) layers that individually provide considerable improvements and synergistically provide even more significant improvements. To compensate for small temperature variations, our design varies the bias current through ring resonators. For larger temperature variations, we propose architecture-level techniques to reroute messages away from hot regions, and through cooler regions, to their destinations. We also propose a thermal/congestion-aware coscheduling algorithm at the OS level to further lower BER by reorganizing the thermal profile of the chip. Our simulation results show that Aurora provides a robust architectural solution to handle temperature variation effects on future photonic NoCs. For instance, average BER and message error rate are reduced by 96% and 85%, respectively, when the combined thermal optimization scheme [shortest path first+ OS] is applied. From the perspective of power efficiency, Aurora is also superior to conventional photonic NoC architectures by as much as 37%. Zhong-Qi Li, Amer Qouneh, Madhura Joshi, Wangyuan Zhang, Xin Fu 0001, Tao Li 0006 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2014 | Software Transactional Memory for GPU Architectures
Rui Wang 0014, Nilanjan Goswami, Tao Li 0006, Lan Gao 0004, Depei Qian 0001 |
CGO | 4 |
| 2014 | An end-to-end analysis of file system features on sparse virtual disksabstractSoftware Defined Data Center (SDDC) is now an emerging area drawing considerable attention in enterprise computing. Software-Defined Storage (SDS), as a key element to enable the SDDC concept, is considered one of the most disruptive storage technologies in modern times. SDS introduces a variety of novel features and functionalities thereby changing the traditional view of the storage stack. In VMware's ESXi virtualization platform, several sparse virtual disk formats have been implemented to support critical features for SDS such as virtual machine (VM) snapshots, Fault-Tolerance (FT), thin provisioning and linked clones. Each virtual disk format supports unique features that may incur complex interactions with other layers of the storage stack such as guest file systems and storage devices. Ruijin Zhou, Sankaran Sivathanu, Jinpyo Kim, Bing Tsai, Tao Li 0006 |
ICS | 5 |
| 2014 | Optimizing virtual machine consolidation performance on NUMA server architecture for cloud workloadsabstractServer virtualization and workload consolidation enable multiple workloads to share a single physical server, resulting in significant energy savings and utilization improvements. The shift of physical server architectures to NUMA and the increasing popularity of scale-out cloud applications undermine workload consolidation efficiency and result in overall system degradation. In this work, we characterize the consolidation of cloud workloads on NUMA virtualized systems, estimate four different sources of architecture overhead, and explore optimization opportunities beyond the default NUMA-aware hypervisor memory management. Motivated by the observed architectural impact on cloud workload consolidation performance, we propose three optimization techniques incorporating NUMA access overhead into the hypervisor's virtual machine memory allocation and page fault handling routines. Among these, estimation of the memory zone access overhead serves as a foundation for the other two techniques: a NUMA overhead aware buddy allocator and a P2M swap FIFO. Cache hit rate, cycle loss due to cache miss, and IPC serve as indicators to estimate the access cost of each memory node. Our optimized buddy allocator dynamically selects low-overhead memory zones and “proportionally” distributes memory pages across target nodes. The P2M swap FIFO records recently unusedlists for mapping exchanges to rebalance memory access pressure within one domain. Our real system based evaluations show a 41.1% performance improvement when consolidating 16-VMs on a 4-socket server (the proposed allocator contributes 22.8% of the performance gain and the P2M swap FIFO accounts for the rest). Furthermore, our techniques can cooperate well with other methods (i.e. vCPU migration) and scale well when varying VM memory size and the number of sockets in a physical host. Ming Liu 0006, Tao Li 0006 |
ISCA | 2 |
| 2014 | Understanding the Impact of vCPU Scheduling on DVFS-Based Power Management in Virtualized Cloud EnvironmentabstractVirtualized platform has emerged as a prominent environment for cloud computing, especially in today's power-constrained data centers. However, due to a lack of coordination between runtime power management and a virtual CPU (vCPU) scheduler, existing virtualized cloud platform is far from efficient. First, current frequency control mechanism is unable to satisfy the fast-changing vCPU frequency requirement imposed by vCPU scheduler, which we refer to as demand imbalance problem. In addition, newly created vCPUs, if scheduled solely based on fairness, can cause inefficient frequency rise and drop on an unmatched physical core, which we refer to as utilization mismatch problem. In both cases, the system incurs degraded power efficiency and sub-optimal workload performance. In this study we perform a comprehensive analysis on the interplay between vCPU scheduling and processor-centric power control in virtualized cloud environment. Using representative workloads from Cloud Suite and real server deployment, we examine the energy/performance implications of frequency scaling and vCPU scheduling on both single-VM and multi-VM cloud host. We show that existing virtualized platform has the potential to improve energy efficiency and workload performance by 32% and 25%, respectively, if vCPUs are balanced and appropriately scheduled. We also show that dirty page rate, virtual block device processing rate, virtual network packets arrival rate, and network I/O buffer availability are important efficiency indicators for energy-efficient virtualized cloud system design. Ming Liu 0006, Chao Li 0009, Tao Li 0006 |
MASCOTS | 3 |
| 2014 | On Characterization of Performance and Energy Efficiency in Heterogeneous HPC Cloud Data CentersabstractThe relocation of high performance computing systems (HPC) to the cloud poses new challenges for data center architects and IT managers. These challenges are due to heterogeneity injected into data centers by cutting-edge virtualization technologies and hardware accelerators used to support emerging cloud applications and services. Although hardware accelerators like General Purpose Graphics Processing Units (GPGPUs) and virtualization technologies have been well studied and evaluated individually, a detailed analysis of their combined architectures and collective behavior from the data center point of view is lacking. Using real platforms and high performance computing workloads, we study the power performance tradeoffs due to various granularities of heterogeneity across hardware and software layers and expose hidden opportunities for optimizing overall data center efficiency. Our approach is to evaluate server power and performance from a data center point of view as opposed to evaluating hardware accelerators and virtualization technologies themselves. Our results show that performance on cloud is affected by virtualization overhead and fraction of serial code. Moreover, GPU workloads achieve 25% and 30% savings in power and energy consumption when executed on low power platforms, and only 50% of our GPU workloads are more energy efficient than their corresponding CPU implementations. The results also show that it is much more power efficient to collocate GPU virtual machines with non-GPU virtual machines. Amer Qouneh, Nilanjan Goswami, Ruijin Zhou, Tao Li 0006 |
MASCOTS | 4 |
| 2014 | Towards Automated Provisioning and Emergency Handling in Renewable Energy Powered Datacenters
Chao Li 0009, Rui Wang 0014, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006, Depei Qian 0001 |
J. Comput. Sci. Technol. | 8 |
| 2013 | Power-performance co-optimization of throughput core architecture using resistive memoryabstractMassively parallel computing on throughput computers such as GPUs requires myriad memory accesses to register files, on-chip scratchpad, caches, and off-chip DRAM. Unlike CPUs, these processors have a large register file and on-chip scratchpad memory, which consume a significant portion of compute core power (35%-45%). In this paper, we introduce novel throughput architecture by integrating resistive memory (Spin Transfer Torque RAM) inside the compute core, which reduces leakage significantly, but introduces write power overhead and longer write latencies in GPU shared memory and register file accesses. We enhance the compute core by introducing register file organization with differential memory update mechanism to remove update redundancy during write operations. Furthermore, using merged register-write-mechanism and write-back buffer, we coalesce multithreaded GPU register write accesses to save write energy. In addition, we introduce hybrid shared memory design using SRAM and STT-MRAM that provides significant leakage/dynamic power savings without affecting performance. On average, across 23 GPGPU/graphics workloads, our schemes save 46% dynamic power due to register access (83% leakage power saving) with negligible performance degradation. On average, hybrid shared memory provides 10% reduction in dynamic power with maximum 1.6× performance improvement for the current workloads at no additional area overhead. Nilanjan Goswami, Bingyi Cao, Tao Li 0006 |
HPCA | 3 |
| 2013 | Enabling distributed generation powered sustainable high-performance data centerabstractThe necessity for capping carbon emission has significantly restricted the potential of modern data centers. For this matter, both industry and academia are proactively seeking opportunities on cross-layer power management schemes that could open a door for sustainable high-performance computing platform. In this paper we investigate an emerging trend in the IT industry: using promising onsite distributed generation (DG) techniques to provide premium clean energy to the computing load. We develop data center power demand shaping (PDS), a novel technique that allows data centers to utilize onsite green energy efficiently. In contrast to prior design, PDS takes advantage of a so-far unexplored power supply feature, i.e., the load following capabilities of DG systems to avoid the high performance penalty issue incurred during supply tracking. In addition, PDS features two adaptive power management schemes: DGR Boost and UPS Boost. These two workload-aware optimization methods leverage mature computer tuning knobs to achieve attractive data center performance improvement. Using real-world data center traces and industry data of distributed generation systems, we show that our technique can come within 1.2% performance of an ideal oracle, which is roughly a 37% improvement over existing supply tracking based design. Our design could save over 100 metric tons of carbon emissions annually for a 10MW data center. Chao Li 0009, Ruijin Zhou, Tao Li 0006 |
HPCA | 3 |
| 2013 | Exploring high-performance and energy proportional interface for phase change memory systemsabstractPhase change memory is emerging as a promising candidate for building up future energy efficient memory systems. To achieve high-performance and energy proportional design, phase change memory devices need to be reorganized so that (1) the relatively long latency of phase change memory devices should be hidden; (2) unnecessary power waste of phase change memory need to be preserved. Previous studies show that conventional memory ranks could be broken down into multiple smaller ranks for increased concurrency and lower power consumption. Nevertheless, the conventional electrical bus is incapable of supporting a large number of memory chips due to its insufficient load capacity and signal traversing speed. In this paper, we propose a phase change memory system design that leverages the state-of-art photonic links to overcome this issue. Moreover, thanks to the flexibility of photonic links, it is possible to amortize the small-rank penalty (e.g. the rank-to-rank switch overhead) by partitioning the channels either statically or dynamically. Our experimental results show that photonically interconnected phase change memory can increase the system performance (IPC) by up to 19% while saving 35% memory system power. Zhong-Qi Li, Ruijin Zhou, Tao Li 0006 |
HPCA | 3 |
| 2013 | ESPN: A case for energy-star photonic on-chip networkabstractPhotonic Network-on-Chips (NoCs) have recently been proposed due to their inherent low latency and high bandwidth. However, the high static power of the photonic components (e.g. laser source, resonators and waveguides) often results in energy-inefficient architectures. In this paper, we advocate the Energy-Star Photonic Network (ESPN) architecture that optimizes energy utilization via a two-pronged approach: (1) by enabling dynamic resource provisioning, ESPN adapts photonic network resources based on runtime traffic characteristics and (2) by utilizing all-optical adaptive routing, ESPN improves energy efficiency by intelligently exploiting existing network resources without introducing high latency and power hungry auxiliary routing mechanisms. Our evaluation results show that compared to the baseline design, ESPN reduces power and energy consumption under synthetic traffic patterns by 50% and 58% respectively. Zhong-Qi Li, Tao Li 0006 |
ISLPED | 2 |
| 2013 | Chameleon: Adapting throughput server to time-varying green power budget using online learningabstractEco-friendly energy sources (i.e. green power) attract great attention as lowering computer carbon footprint has become a necessity. Existing proposals on managing green energy powered systems show sub-optimal results since they either use rigid load power capping or heavily rely on backup power. We propose Chameleon, a novel adaptive green throughput server. Chameleon comprises of multiple flexible power management policies and leverages learning algorithm to select the optimal operating mode during runtime. The proposed design outperforms the state-of-the-art approach by 13% on performance, improves system MTBF by 42%, and still maintains up to 95% green energy utilization. Chao Li 0009, Rui Wang 0014, Tao Li 0006, Nilanjan Goswami, Depei Qian 0001 |
ISLPED | 4 |
| 2013 | Understanding the implications of virtual machine management on processor microarchitecture designabstractCloud computing has demonstrated tremendous capability in a wide spectrum of online services. Virtualization provides an efficient solution to the utilization of modern multicore processor systems while affording significant flexibility. The growing popularity of virtualized datacenters motivates deeper understanding of the interactions between virtual machine management and the micro-architecture behaviors of the privileged domain. We argue that these behaviors must be factored into the design of processor microarchitecture in virtualized datacenters. In this work, we use performance counters on modern servers to study the micro-architectural execution characteristics of the privileged domain while performing various VM management operations. Our study shows that today's state-of-the-art processor still has room for further optimizations when executing virtualized cloud workloads, particularly in the organization of last level caches and on-chip cache coherence protocol. Specifically, our analysis shows that: shared caches could be partitioned to eliminate interference between the privileged domain and guest domains; the cache coherence protocol could support a high degree of data sharing of the privileged domain; and cache capacity or CPU utilization occupied by the privileged domain could be effectively managed when performing management workflows to achieve high system throughput. Xiufeng Sui, Tao Li 0006, Lixin Zhang 0002 |
ISPASS | 3 |
| 2013 | Wall-clock based synchronization: A parallel simulation technology for cluster systemsabstractA common practice for reducing synchronization overheads in parallel simulation of a large-scale cluster is to relax synchronization with lengthened synchronous steps. However, as a side effect, simulation accuracy degrades considerably. This paper proposes a novel mechanism that keeps the running speeds of different nodes consistent by synchronizing logical clocks with the wall clock periodically within each lax step. Because speed deviations of nodes are the main source of time causality errors, through aligning speeds our mechanism only causes modest precision loss while achieving a close performance to lax synchronization. The experimental results show that it improves the performance by 2 to 11 times relative to the baseline barrier synchronization with a high accuracy (e.g. 99% in most cases). Compared to the recently proposed adaptive mechanism, it also achieves nearly 30% performance improvement. Junmin Wu, Guoliang Chen 0001, Tao Li 0006 |
ISPASS | 4 |
| 2013 | Enabling datacenter servers to scale out economically and sustainablyabstractAs cloud applications proliferate and data-processing demands increase, server resources must grow to unleash the performance of emerging workloads that scale well with large number of compute nodes. Nevertheless, power has become a crucial bottleneck that restricts horizontal scaling (scale out) of server systems, especially in datacenters that employ power over-subscription. When a datacenter hits the maximum capacity of its power provisioning equipment, the owner has to either build another facility or upgrade existing utility power infrastructure -- both approaches add huge capital expenditure, require significant construction lead time, and can further increase the owner's carbon footprint. Chao Li 0009, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006 |
MICRO | 7 |
| 2013 | SolarTune: Real-time scheduling with load tuning for solar energy powered multicore systemsabstractIn this paper, we target at the direct-coupled solar energy powered multicore architectures that provide direct power supply between photovoltaic (PV) generation and the load without the adoption of battery. We present Solar-Tune, a real-time scheduling technique with load tuning for sporadic tasks on solar energy powered multicore systems. The objective is to fully utilize the available solar energy while meeting the deadlines of tasks. To solve the problem, we first perform analysis and formulate the scheduling problem in each duration as an integer linear programming (ILP) model to obtain an optimal schedule. Then we present a heuristic algorithm to dynamically refine the task scheduling based on the predictions of the availability of solar energy. We conduct experiments using real-world meteorological data across different geographic sites. The experimental results show that SolarTune can significantly improve the solar energy utilization ratio and reduce the number of deadline misses compared to the conventional task scheduler. Yi Wang 0003, Renhai Chen, Zili Shao, Tao Li 0006 |
RTCSA | 4 |
| 2013 | Accelerating GPGPU architecture simulationabstractRecently, graphics processing units (GPUs) have opened up new opportunities for speeding up general-purpose parallel applications due to their massive computational power and up to hundreds of thousands of threads enabled by programming models such as CUDA. However, due to the serial nature of existing micro-architecture simulators, these massively parallel architectures and workloads need to be simulated sequentially. As a result, simulating GPGPU architectures with typical benchmarks and input data sets is extremely time-consuming. This paper addresses the GPGPU architecture simulation challenge by generating miniature, yet representative GPGPU kernels. We first summarize the static characteristics of an existing GPGPU kernel in a profile, and analyze its dynamic behavior using the novel concept of the divergence flow statistics graph (DFSG). We subsequently use a GPGPU kernel synthesizing framework to generate a miniature proxy of the original kernel, which can reduce simulation time significantly. The key idea is to reduce the number of simulated instructions by decreasing per-thread iteration counts of loops. Our experimental results show that our approach can accelerate GPGPU architecture simulation by a factor of 88X on average and up to 589X with an average IPC relative error of 5.6%. Zhibin Yu 0001, Lieven Eeckhout, Nilanjan Goswami, Tao Li 0006, Lizy Kurian John, Hai Jin 0001, Cheng-Zhong Xu 0001 |
SIGMETRICS | 4 |
| 2013 | Leveraging phase change memory to achieve efficient virtual machine executionabstractVirtualization technology is being widely adopted by servers and data centers in the cloud computing era to improve resource utilization and energy efficiency. Nevertheless, the heterogeneous memory demands from multiple virtual machines (VM) make it more challenging to design efficient memory systems. Even worse, mission critical VM management activities (e.g. checkpointing) could incur significant runtime overhead due to intensive IO operations. In this paper, we propose to leverage the adaptable and non-volatile features of the emerging phase change memory (PCM) to achieve efficient virtual machine execution. Towards this end, we exploit VM-aware PCM management mechanisms, which 1) smartly tune SLC/MLC page allocation within a single VM and across different VMs and 2) keep critical checkpointing pages in PCM to reduce I/O traffic. Experimental results show that our single VM design (IntraVM) improves performance by 10% and 20% compared to pure SLC- and MLC- based systems. Further incorporating VM-aware resource management schemes (IntraVM+InterVM) increases system performance by 15%. In addition, our design saves 46% of checkpoint/restore duration and reduces 50% of overall IO penalty to the system. Ruijin Zhou, Tao Li 0006 |
VEE | 2 |
| 2013 | Optimizing virtual machine live storage migration in heterogeneous storage environmentabstractVirtual machine (VM) live storage migration techniques significantly increase the mobility and manageability of virtual machines in the era of cloud computing. On the other hand, as solid state drives (SSDs) become increasingly popular in data centers, VM live storage migration will inevitably encounter heterogeneous storage environments. Nevertheless, conventional migration mechanisms do not consider the speed discrepancy and SSD's wear-out issue, which not only causes significant performance degradation but also shortens SSD's lifetime. This paper, for the first time, addresses the efficiency of VM live storage migration in heterogeneous storage environments from a multi-dimensional perspective, i.e., user experience, device wearing, and manageability. We derive a flexible metric (migration cost), which captures various design preference. Based on that, we propose and prototype three new storage migration strategies, namely: 1) Low Redundancy (LR), which generates the least amount of redundant writes; 2) Source-based Low Redundancy (SLR), which keeps the balance between IO performance and write redundancy; and 3) Asynchronous IO Mirroring, which seeks the highest IO performance. The evaluation of our prototyped system shows that our techniques outperform existing live storage migration by a significant margin. Furthermore, by adaptively mixing our proposed schemes, the cost of massive VM live storage migration can be even lower than that of only using the best of individual mechanism. Ruijin Zhou, Fang Liu 0002, Chao Li 0009, Tao Li 0006 |
VEE | 4 |
| 2012 | Integrating nanophotonics in GPU microarchitectureabstractAs high-performance computing device, the GPU has exposed bandwidth and latency bottlenecks in on-chip interconnect and off-chip memory access. To eliminate such bottlenecks, we employ silicon nanophotonics and 3D stacking technologies in GPU microarchitecture. This provides higher communication bandwidth and lower latency signaling mechanisms at reduced power. Furthermore, to insulate the performance of the GPU compute cores from the interconnect bottlenecks we propose a novel interconnect aware thread scheduling scheme to alleviate the traffic congestion. We evaluate a 3D stacked GPU with 2048 SIMD cores having photonic interconnect. The photonic multiple-write-single-read crossbar network with 32B channel bandwidth on average, achieves 96% power reduction. We anticipate that for emerging workloads and microarchitectures the implications of the proposed ideas are far reaching in terms of power and performance. Nilanjan Goswami, Zhong-Qi Li, Ajit Verma, Ramkumar Shankar, Tao Li 0006 |
PACT | 5 |
| 2012 | Aurora: A thermally resilient photonic network-on-chip architectureabstractWith silicon optical technology moving towards maturity, the use of photonic network-on-chip (NoCs) for global chip communication is emerging as a promising solution to communication requirements of future many core processors. It is expected that photonic NoCs will play an important role in alleviating current power, latency, and bandwidth constraints. However, photonic NoCs are sensitive to ambient temperature variations because their basic constituents, ring resonators, are themselves sensitive to those variations. Since ring resonators are basic building blocks for photonic modulators, switches, multiplexers, and demultiplexers, variations of on-chip temperature pose serious challenges to the proper operation of photonic NoCs. Proposed methods that mitigate the effects of temperature at device level are either difficult to use in CMOS processes or not suitable for large scale implementation. In this paper, we propose Aurora, a thermally resilient photonic NoC architecture design that supports reliable and low bit error rate (BER) on-chip communications in the presence of large temperature variations. Our proposed architecture leverages solutions at both device and architecture layers that synergistically provide significant improvements. To compensate for small temperature variations, our design varies the bias current through ring resonators. For larger temperature variations, we propose architecture-level techniques to re-route messages away from hot regions, and through cooler regions, to their destinations, thereby lowering BER. Our simulation results show that Aurora provides a robust architectural solution to handle temperature variation effects on future photonic NoCs. For instance, average BER and message error rate (MER) are reduced by 78% and 30% respectively when the combined device and architectural technique (SPF) is applied. From the perspective of power efficiency, Aurora is also superior to conventional photonic NoC architectures by as much as 33%. Amer Qouneh, Zhong-Qi Li, Madhura Joshi, Wangyuan Zhang, Xin Fu 0001, Tao Li 0006 |
ICCD | 6 |
| 2012 | iSwitch: Coordinating and optimizing renewable energy powered server clustersabstractLarge-scale computing systems such as data centers are facing increasing pressure to cap their carbon footprint. Integrating emerging clean energy solutions into computer system design therefore gains great significance in the green computing era. While some pioneering work on tracking variable power budget show promising energy efficiency, they are not suitable for data centers due to lack of performance guarantee when renewable generation is low and fluctuant. In addition, our characterization of wind power behavior reveals that data centers designed to track the intermittent renewable power incur up to 4X performance loss due to inefficient and redundant load matching activities. As a result, mitigating operational overhead while still maintaining desired energy utilization becomes the most significant challenge in managing server clusters on intermittent renewable energy generation. In this paper we take a first step in digging into the operational overhead of renewable energy powered data center. We propose iSwitch, a lightweight server power management that follows renewable power variation characteristics, leverages existing system infrastructures, and applies supply/load cooperative scheme to mitigate the performance overhead. Comparing with state-of-the-art renewable energy driven system design, iSwitch could mitigate average network traffic by 75%, peak network traffic by 95%, and reduce 80% job waiting time while still maintaining 96% renewable energy utilization. We expect that our work can help computer architects make informed decisions on sustainable and high-performance system design. Chao Li 0009, Amer Qouneh, Tao Li 0006 |
ISCA | 3 |
| 2011 | Helmet: A resistance drift resilient architecture for multi-level cell phase change memory systemabstractPhase change memory (PCM) is emerging as a promising solution for future memory systems and disk caches. As a type of resistive memory, PCM relies on the electrical resistance of Ge2Sb2Te5(GST) to represent stored information. With the adoption of multi-level programming PCM devices, unwanted resistance drift is becoming an increasing reliability concern in future high-density, multi-level cell PCM systems. To address this issue without incurring a significant storage and performance overhead in ECC, conventional design employs a conservative approach, which increases the resistance margin between two adjacent states to combat resistance drift. In this paper, we show that the wider margin adversely impacts the low-power benefit of PCM by incurring up to 2.3X power overhead and causes up to 100X lifetime reduction, thereby exacerbating the wear-out issue. To tolerate resistance drift, we proposed Helmet, a multi-level cell phase change memory architecture that can cost-effectively reduce the readout error rate due to drift. Therefore, we can relax the requirement on margin size, while preserving the readout reliability of the conservative approach, and consequently minimize the power and endurance overhead due to drift. Simulation results show that our techniques are able to decrease the error rate by an average of 87%. Alternatively, for satisfying the same reliability target, our schemes can achieve 28% power savings and a 15X endurance enhancement due to the reduced margin size when compared to the conservative approach. Wangyuan Zhang, Tao Li 0006 |
DSN | 2 |
| 2011 | Mercury: A fast and energy-efficient multi-level cell based Phase Change Memory systemabstractPhase Change Memory (PCM) is one of the most promising technologies among emerging non-volatile memories. PCM stores data in crystalline and amorphous phases of the GST material using large differences in their electrical resistivity. Although it is possible to design a high capacity memory system by storing multiple bits at intermediate levels between the highest and lowest resistance states of PCM, it is difficult to obtain the tight distribution required for accurate reading of the data. Moreover, the required programming latency and energy for a Multiple Level PCM (MLC-PCM) cell is not trivial and can act as a major hurdle in adopting multilevel PCM in a high-density memory architecture design. Furthermore, the effect of process variation (PV) on PCM cell exacerbates the variability in necessary programming current and hence the target resistance spread, leading to the demand for high-latency, multi-iteration-based programming-and-verify write schemes for MLC-PCM. PV-aware control of programming current, programming using staircase down current pulses and programming using increasing reset current pulses are some of the traditional techniques used to achieve optimum programming energy, write latency and accuracy, but they usually target on optimizing only one aspect of the design. In this paper, we address the high-write latency and process variation issues of MLC-PCM by introducing Mercury: A fast and energy efficient multi-level cell based phase change memory architecture. Mercury adapts the programming scheme of a multi-level PCM cell by taking into consideration the initial state of the cell, the target resistance to be programmed and the effect of process variation on the programming current profile of the cell. The proposed techniques act at circuit as well as microarchitecture levels. Simulation results show that Mercury achieves 10% saving in programming latency and 25% saving in programming energy for the PCM memory system compared to that of the traditional methods. Madhura Joshi, Wangyuan Zhang, Tao Li 0006 |
HPCA | 3 |
| 2011 | SolarCore: Solar energy driven multi-core architecture power managementabstractThe global energy crisis and environmental concerns (e.g. global warming) have driven the IT community into the green computing era. Of clean, renewable energy sources, solar power is the most promising. While efforts have been made to improve the performance-per-watt, conventional architecture power management schemes incur significant solar energy loss since they are largely workload-driven and unaware of the supply-side attributes. Existing solar power harvesting techniques improve the energy utilization but increase the environmental burden and capital investment due to the inclusion of large-scale batteries. Moreover, solar power harvesting itself cannot guarantee high performance without appropriate load adaptation. To this end, we propose SolarCore, a solar energy driven, multi-core architecture power management scheme that combines maximal power provisioning control and workload run-time optimization. Using real-world meteorological data across different geographic sites and seasons, we show that SolarCore is capable of achieving the optimal operation condition (e.g. maximal power point) of solar panels autonomously under various environmental conditions with a high green energy utilization of 82% on average. We propose efficient heuristics for allocating the time varying solar power across multiple cores and our algorithm can further improve the workload performance by 10.8% compared with that of round-robin adaptation, and at least 43% compared with that of conventional fixed-power budget control. This paper makes the first step on maximally reducing the carbon footprint of computing systems through the usage of renewable energy sources. We expect that the novel joint optimization techniques proposed in this paper will contribute to building a truly sustainable, high-performance computing environment. Chao Li 0009, Wangyuan Zhang, Chang-Burm Cho, Tao Li 0006 |
HPCA | 4 |
| 2011 | Optimizing throughput/power trade-offs in hardware transactional memory using DVFS and intelligent schedulingabstractPower has emerged as a first-order design constraint in modern processors and has energized microarchitecture researchers to produce a growing number of power optimization proposals. Almost in tandem with the move toward more energy-efficient designs, architects have been increasing the number of processing elements (PEs) on a single chip and promoting the concept of running multithreaded workloads. Nevertheless, software is still lagging behind and is often unable to exploit these additional resources -- giving rise to transactional memory. Transactional memory is a promising programming abstraction that makes it easier for programmers to exploit the resources available in many- core processor systems by removing some of the complexity associated with traditional lock-based programming. This paper proposes new techniques to merge the power and transactional memory domains. Clay Hughes, Tao Li 0006 |
ICS | 2 |
| 2011 | Characterizing and analyzing renewable energy driven data centersabstractAn increasing number of data centers today start to incorporate renewable energy solutions to cap their carbon footprint. However, the impact of renewable energy on large-scale data center design is still not well understood. In this paper, we model and evaluate data centers driven by intermittent renewable energy. Using real-world data center and renewable energy source traces, we show that renewable power utilization and load tuning frequency are two critical metrics for designing sustainable high-performance data centers. Our characterization reveals that load power fluctuation together with the intermittent renewable power supply introduce unnecessary tuning activities, which can increase the management overhead and degrade the performance of renewable energy driven data centers. Chao Li 0009, Amer Qouneh, Tao Li 0006 |
SIGMETRICS | 3 |
| 2010 | Architecting reliable multi-core network-on-chip for small scale processing technologyabstractThe trend towards multi-/many- core design has made network-on-chip (NoC) a crucial component of future microprocessors. With CMOS processing technologies continuously scaling down to the nanometer regime, effects such as process variation (PV) and negative bias temperature instability (NBTI) significantly decrease hardware reliability and lifetime. Therefore, it is imperative for multi-core architects to consider and mitigate these effects in NoCs implemented using small-scale processing technology. This paper reports on a first step to optimize NoC architecture reliability in light of both PV and NBTI effects. We propose novel techniques that can hierarchically alleviate PV and NBTI effects on NoC while leveraging their benign interaction. Our low-level design improves PV and NBTI efficiency of key components (e.g. virtual channel allocation logics, virtual channels) of critical paths of the pipelined router microarchitecture. Our high-level mechanisms leverage NBTI degradation and PV information from multiple routers to intelligently route packets, delivering optimized performance-power-reliability across the NoC substrate. Experimental results show that our intra-router level techniques (i.e. VA_M1 and VC_M2) reduce guardband by 47% while improving network throughput by 24%. Our inter-router optimization scheme (i.e. IR_M3) results in 50% guardband reduction and 19% network latency improvement. Xin Fu 0001, Tao Li 0006, José A. B. Fortes |
DSN | 2 |
| 2009 | Exploring Phase Change Memory and 3D Die-Stacking for Power/Thermal Friendly, Fast and Durable Memory ArchitecturesabstractEmerging three-dimensional (3D) integration technology allows for the direct placement of DRAM on top of a microprocessor, significantly reducing the wire-delay between the two and thereby alleviating memory latency and bandwidth constraints. However, the increase in power density of 3D technology leads to elevated on-chip temperature, which results in an exponential rise in charge leakage of DRAM. Consequently, the refresh frequency of 3D die-stacked DRAM needs to be doubled (or more) to retain data at the expense of additional power overhead. In this work, we investigate using phase-change random access memory (PRAM) as a promising candidate to achieve scalable, low power and thermal friendly memory system architecture in the upcoming 3D-stacking technology era. Using analytical model, circuit- and architectural-level simulations that capture both physical and electrical characteristics of PRAM, we show that the higher temperature of 3D chips is beneficial to PRAM power savings due to its unique, heat-driven programming mechanisms. Moreover, we show that the through silicon vias (TSVs) ubiquitously used in 3D implementations contribute further PRAM power savings due to their substantially lower resistance to the high PRAM programming current. To effectively integrate PRAM into a conventional memory hierarchy, we propose architecture and OS support to address its write latency and reliability disadvantages. We present a hybrid PRAM/DRAM memory architecture and exploit an OS-level paging scheme to improve PRAM write performance and lifetime. Moreover, we leverage the error-correcting capability of strong ECC codes to expand PRAM lifespan and use wear-out aware OS page allocation to minimize ECC performance overhead. Our experimental results show that compared to die-stacked planar DRAM, our design reduces the overall power consumption of the memory system by 54% with 6% performance degradation, consequently alleviating the thermal constraint of 3D chips by up to 4.25degC and achieving a speedup of up to 1.1X. We also show that the lifetime can be improved by a factor of 114X using the proposed endurance optimization schemes. Wangyuan Zhang, Tao Li 0006 |
PACT | 2 |
| 2009 | Soft error vulnerability aware process variation mitigationabstractAs transistor process technology approaches the nanometer scale, process variation significantly affects the design and optimization of high performance microprocessors. Prior studies have shown that chip operating frequency and leakage power can have large variations due to fluctuations in transistor gate length and sub-threshold voltage. In this work, we study the impact of process variation on microarchitecture soft error robustness, an increasing reliability design challenge in the billion-transistor chip era. We explore two techniques that can effectively mitigate the effect of design parameter variation while significantly enhancing microarchitecture soft error reliability. Our first technique is entry-based. It tolerates the deleterious impact of variable latency techniques on soft error reliability by reducing the quantity and residency cycle of vulnerable bits in the microarchitecture structure at a fine granularity. Our second technique is structure-based. It applies body biasing schemes to dynamically adapt transistor sub-threshold voltage (and hence device-level soft error robustness) to the program reliability characteristics at a coarse granularity. We also combine the two techniques which further produces improved results. Compared to existing process variation tolerant schemes, our proposed techniques achieve optimal trade-offs between reliability, performance, and power. To our knowledge, this paper presents the first study on characterizing and optimizing processor microarchitecture resilience to soft errors in light of process variation. Xin Fu 0001, Tao Li 0006, José A. B. Fortes |
HPCA | 2 |
| 2009 | TransMetric: architecture independent workload characterization for transactional memory benchmarksabstractTransactional memory (TM) has emerged as a parallel programming paradigm for multi-core processors yet there is no standardized set of metrics with which to describe their behavior. In this work, we propose a set of transaction-oriented workload characteristics that can accurately capture the behavior of transactional memory programs. We apply principle component analysis and clustering algorithms to analyze the proposed transactional workload characteristics and show that these characteristics are architecturally independent James Poe, Clay Hughes, Tao Li 0006 |
ICS | 3 |
| 2009 | Accurate, scalable and informative design space exploration for large and sophisticated multi-core oriented architecturesabstractAs microprocessors become more complex, early design space exploration plays an essential role in reducing the time to market and post-silicon surprises. The trend toward multi-/many- core processors will result in sophisticated large-scale architecture substrates (e.g. non-uniformly accessed caches interconnected by network-on-chip) that exhibit increasingly complex and heterogeneous behavior. While conventional analytical modeling techniques can be used to efficiently explore the characteristics (e.g. IPC and power) of monolithic architecture design, existing methods lack the ability to accurately and informatively forecast the complex behavior of large and distributed architecture substrates across the design space. This limitation will only be exacerbated with the rapidly increased integration scale (e.g. number of cores per chip). In this paper, we propose novel, multi-scale 2D predictive models which can efficiently reason the characteristics of large and sophisticated multi-core oriented architectures during the design space exploration stage without using detailed cycle-level simulations. Our proposed techniques employ 2D wavelet multiresolution analysis and neural network regression modeling. We extensively evaluate the efficiency of our predictive models in forecasting the complex and heterogeneous characteristics of large and distributed shared cache interconnected by a network on chip in multi-core designs using both multi-programmed and multithreaded workloads. Experimental results show that the models achieve high accuracy while maintaining low complexity and computation overhead. Through case studies, we demonstrate that the proposed techniques can be used to informatively explore and accurately evaluate global, cooperative multi-core resource allocation and thermal-aware designs that cannot be achieved using conventional design exploration methods. Chang-Burm Cho, James Poe, Tao Li 0006, Jingling Yuan |
MASCOTS | 3 |
| 2009 | TransPlant: A parameterized methodology for generating transactional memory workloadsabstractTransactional memory provides a means to bridge the discrepancy between programmer productivity and the difficulty in exploiting thread-level parallelism gains offered by emerging chip multiprocessors. Because the hardware has outpaced the software, there are very few modern multithreaded benchmarks available and even fewer for transactional memory researchers. This hurdle must be overcome for transactional memory research to mature and to gain widespread acceptance. Currently, for performance evaluations, most researchers rely on manually converted lock-based multithreaded workloads or the small group of programs written explicitly for transactional memory. Using converted benchmarks is problematic because they have been tuned so well that they may not be representative of how a programmer will actually use transactional memory. Hand coding stressor benchmarks is unattractive because it is tedious and time consuming. A new parameterized methodology that can automatically generate a program based on the desired high-level program characteristics benefits the transactional memory community. In this work, we propose techniques to generate parameterized transactional memory benchmarks based on a feature set, decoupled from the underlying transactional model. Using principle component analysis, clustering, and raw transactional performance metrics, we show that TransPlant can generate benchmarks with features that lie outside the boundary occupied by these traditional benchmarks. We also show how TransPlant can mimic the behavior of SPLASH-2 and STAMP transactional memory workloads. The program generation methods proposed here will help transactional memory architects select a robust set of programs for quick design evaluations. James Poe, Clay Hughes, Tao Li 0006 |
MASCOTS | 3 |
| 2009 | Characterizing and mitigating the impact of process variations on phase change based memory systemsabstractDynamic Random Access Memory (DRAM) has been used in main memory design for decades. However, DRAM consumes an increasing power budget and faces difficulties in scaling down for small feature size CMOS processing technologies. Compared to conventional DRAM, emerging phase change random access memory (PRAM) demonstrates superior power efficiency and processing scalability as VLSI technologies and integration density continue to advance. Nevertheless, using nano-scale fabrication technologies will unavoidably introduce design parameter variability in the manufacturing stage. In the past, the impact of process variation (PV) on conventional transistor-based storage cells and combinational logic has been studied extensively. However, the implication of PV on non-volatile memory design using emerging phase change techniques has not been well understood. In this paper, we take the first step toward characterizing the effect of process variation on PRAM and explore PV-aware design techniques. We show that process variation increases the PRAM programming power by 96% and degrades PRAM endurance by 50X. Our proposed circuit and two microarchtiecture techniques with system-level support reduce PRAM power by 44%, 59% and 57% and improve PRAM endurance by 27X, 277X and 268X, relative to PV-affected PRAM design. Moreover, we show that the synergy of the proposed cross-layer approaches, which achieve an average 63% power savings and 13050X endurance improvement over the conventional case, provide an attractive design solution to mitigate the deleterious impact of PV for non-volatile memory in the upcoming nano-scale processing technology era. Wangyuan Zhang, Tao Li 0006 |
MICRO | 2 |
| 2008 | Managing multi-core soft-error reliability through utility-driven cross domain optimizationabstractAs semiconductor processing technology continues to scale down, managing reliability becomes an increasingly difficult challenge in high-performance microprocessor design. Transient faults, also known as soft errors, corrupt program data at the circuit level and cause incorrect program execution and system crashes. Future processors will consist of billions of transistors organized as multi-core microarchitectures. Packaging multiple cores (and hence more transistors) onto the same die exposes more devices to soft error strikes. This paper explores utility-function-driven (benefit driven) cross domain optimization for both performance and reliability. We propose the use of utility-based resource management for individual cores while applying utility-based shared cache partitioning across multiple cores. Moreover, we coordinate the optimization of multiple resources based on their cross domain utility information to achieve attractive performance and reliability tradeoffs. Extensive experimental results show that, on average, our utility-driven cross domain optimization reduces the soft error rate of the most vulnerable core in a Chip Multiprocessor (CMP) by up to 35% and improves the CMP’s overall reliability by 22% with less than 3% performance degradation across 15 investigated workloads. Wangyuan Zhang, Tao Li 0006 |
ASAP | 2 |
| 2008 | Archer: A Community Distributed Computing Infrastructure for Computer Architecture Research and Education
Renato J. O. Figueiredo, P. Oscar Boykin, José A. B. Fortes, Tao Li 0006, Jie-Kwon Peir, David Wolinsky, Lizy Kurian John, David R. Kaeli, David J. Lilja, Sally A. McKee, Gokhan Memik, Alain J. Roy, Gary S. Tyson |
CollaborateCom | 4 |
| 2008 | Combined circuit and microarchitecture techniques for effective soft error robustness in SMT processorsabstractAs semiconductor technology scales, reliability is becoming an increasingly crucial challenge in microprocessor design. The rSRAM and voltage scaling are two promising circuit-level radiation hardening techniques to increase soft error robustness of a SRAM-based storage cell. However, applying circuit-level radiation hardening techniques to all on-chip transistors will result in significant overhead in performance and power consumption. In this paper, we propose microarchitecture support that allows cost-effective implementation of radiation hardened key microarchitecture structures (e.g. issue queue and reorder buffer) in SMT processors using soft error robust circuit techniques. Our study shows that the combined circuit and microarchitecture techniques achieve attractive tradeoffs between reliability, performance and power. Xin Fu 0001, Tao Li 0006, José A. B. Fortes |
DSN | 2 |
| 2008 | Optimizing Issue Queue Reliability to Soft Errors on Simultaneous Multithreaded ArchitecturesabstractThe issue queue (IQ) is a key microarchitecture structure for exploiting instruction-level and thread-level parallelism in dynamically scheduled simultaneous multithreaded (SMT) processors. However, exploiting more parallelism yields high susceptibility to transient faults on a conventional IQ. With the rapidly increasing soft error rates, the IQ is likely to be a reliability hot-spot on SMT processors fabricated with advanced technology nodes using smaller and denser transistors with lower threshold voltages and tighter noise margins. In this paper, we explore microarchitecture techniques to optimize IQ reliability to soft error on SMT architectures. We propose to use off-line instruction vulnerability profiling to identify reliability critical instructions. The gathered information is then used to guide reliability-aware instruction scheduling and resource allocation in multithreaded execution environments. We evaluate the efficiency of the proposed schemes across various SMT workload mixes. Extensive simulation results show that, on average, our microarchitecture level soft error mitigation techniques can significantly reduce IQ vulnerability by 42% with 1% performance improvement. To maintain runtime IQ reliability for pre-defined thresholds, we propose dynamic vulnerability management (DVM) mechanisms. Experimental results show that our DVM techniques can effectively achieve desired reliability/performance tradeoffs. Xin Fu 0001, Wangyuan Zhang, Tao Li 0006, José A. B. Fortes |
ICPP | 3 |
| 2008 | Modeling and Analyzing the Effect of Microarchitecture Design Parameters on Microprocessor Soft Error Vulnerability
Chang-Burm Cho, Wangyuan Zhang, Tao Li 0006 |
MASCOTS | 3 |
| 2008 | NBTI tolerant microarchitecture design in the presence of process variationabstractNegative bias temperature instability (NBTI), which reduces the lifetime of PMOS transistors, is becoming a growing reliability concern for sub-micrometer CMOS technologies. Parametric variation introduced by nano-scale device fabrication inaccuracy can exacerbate the PMOS transistor wear-out problem and further reduce the reliable lifetime of microprocessors. In this work, we propose microarchitecture design techniques to combat the combined effect of NBTI and process variation (PV) on the reliability of high-performance microprocessors. Experimental evaluation shows our proposed process variation aware (PV-aware) NBTI tolerant microarchitecture design techniques can considerably improve the lifetime of reliability operation while achieving an attractive trade-off with performance and power. Xin Fu 0001, Tao Li 0006, José A. B. Fortes |
MICRO | 2 |
| 2008 | Microarchitecture soft error vulnerability characterization and mitigation under 3D integration technologyabstractAs semiconductor processing techniques continue to scale down, transient faults, also known as soft errors, are increasingly becoming a reliability threat to high-performance microprocessors fabricated using state-of-the-art CMOS technologies. Emerging 3D chip integration techniques leverage vertically stacked structures to reduce on-chip wire delay and have shown the capability of overcoming interconnect bottlenecks as well as reducing power consumption. While the benefits of 3D die stacking on microprocessor performance and power have been extensively investigated recently, its implication on transient fault susceptibility is largely unknown. In this work, we make the first attempt to characterize microarchitecture soft error vulnerabilities across the stacked chip layers under 3D integration technologies. Using models and simulations that capture soft error physical mechanism and circuit/architecture level impact, our study reveals the opportunities of leveraging 3D integration (e.g. the structure of vertical stacking and the incorporation of heterogeneous process technologies) to achieve enhanced reliability. We showcase that the first characteristic allows outer-layers to shield inter-layers from particle strikes and the second feature enables the deployment of error resilience device techniques (e.g. Silicon-On-Insulator) on vulnerable layers to achieve a reliability target while minimizing manufacturing cost. We further propose a set of microarchitecture techniques which can effectively exploit the reliability benefits offered by 3D technologies. For example, we propose the scheduling of vulnerable in-flight instructions to reliable layers and design robust register files by combing reliability-hardened circuits, program value vulnerability and 3D integration techniques. Experimental results show that these techniques are able to substantially reduce 3D microarchitectures’ soft error rate by up to 88% compared to a planar design. We further evaluate the thermal implication of the proposed techniques and conclude that their impact on chip temperature is negligible. Wangyuan Zhang, Tao Li 0006 |
MICRO | 2 |
| 2008 | ORBIT: Effective Issue Queue Soft-Error Vulnerability Mitigation on Simultaneous Multithreaded Architectures Using Operand Readiness-Based Instruction DispatchabstractWith the advance of semiconductor processing technology, soft errors have become an increasing cause of failures of microprocessors fabricated using smaller and more densely integrated transistors with lower threshold voltages and tighter noise margins. With diminishing performance returns on wider issue superscalar processors, the microprocessor design industry has opted for using simultaneous multithreaded (SMT) architectures in commercial processors to exploit thread-level parallelism (TLP). SMT techniques enhance overall system performance but also introduce greater susceptibility to soft errors - concurrently executing multiple threads exposes many program runtime states to soft-error strikes at any given time. The issue queue (IQ) is a key micro architecture structure to exploit instruction-level and thread-level parallelism. On SMT processors, the IQ buffers a large number of instructions from multiple threads and is more susceptible to soft-error strikes. In this paper, we explore the use of operand-readiness-based instruction dispatch (ORBIT) as an effective mechanism to mitigate IQ soft-error vulnerability on SMT processors. We observe that IQ soft-error vulnerability is largely affected by instructions waiting for their source operands. The overall IQ soft-error vulnerability can be effectively reduced by minimizing the number of waiting instructions and their residency cycles in the IQ. We develop six techniques that aim to improve IQ reliability with negligible performance degradation on SMT processors. Moreover, we extend our techniques with prediction methods that can anticipate the readiness of source operands ahead of time. The ORBIT schemes integrated with reliability-awareness and readiness prediction achieve more attractive reliability/performance trade-offs. The best of the proposed schemes (e.g. Predict_DelayACE) reduces IQ vulnerability by 79% with only 1% throughput IPC and 3% harmonic IPC reduction across all studied workloads. Xin Fu 0001, Tao Li 0006, José A. B. Fortes |
SBAC-PAD | 2 |
| 2008 | Using Analytical Models to Efficiently Explore Hardware Transactional Memory and Multi-Core Co-DesignabstractTransactional memory is emerging as a parallel programming paradigm for multi-core processors. Despite the recent interest in transactional memory, there has been no study to characterize the interaction between hardware transactional memory (HTM) design dimensions and multi-core microarchitecture configuration. In this paper, we investigate the use of analytical modeling techniques to build application-specific performance models for understanding the interaction between HTM and multi-core configurations across large design points and for efficiently exploring the co-design space between the two. A key feature of our modeling technique is the ability to simultaneously capture the individual and combinatorial effects of important HTM design dimensions and core microarchitectural parameters. We show that analytical models can be effective tools for assisting architects in identifying these key effects and the interactions between HTM and multi-core microarchitecture that have a high impact on the performance of transactional memory workloads. The models also enable accurate performance prediction across the joint TM/multi-core design space. By analyzing the regression trees generated from our neural network model building methods, we further reveal heterogeneous interaction between TM workloads, core microarchitectures and TM mechanisms. James Poe, Chang-Burm Cho, Tao Li 0006 |
SBAC-PAD | 3 |
| 2007 | Using Wavelet Domain Workload Execution Characteristics to Improve Accuracy, Scalability and Robustness in Program Phase AnalysisabstractProgram phase analysis has many applications in computer architecture design and optimization. Recently, there has been a growing interest in employing wavelets as a tool for phase analysis. Nevertheless, the examined scope of workload characteristics and the explored benefits due to wavelet-based analysis are quite limited. This work further extends prior research by applying wavelets analysis to abundant types of program execution statistics and quantifying the benefits of wavelet analysis in terms of accuracy, scalability and robustness in phase classification. Experimental results on SPEC CPU 2000 benchmarks show that compared with methods that work in the time domain, wavelet domain phase analysis achieves higher accuracy and exhibits superior scalability and robustness. We examine and contrast the effectiveness of applying wavelets to a wide range of runtime workload execution characteristics. We find that wavelet transform significantly reduces temporal dependence in the sampled workload statistics and therefore simple models which are insufficient in the time domain become quite accurate in the wavelet domain. More attractively, we show that different types of workload execution characteristics in wavelet domain can be assembled together to further improve phase classification accuracy. For long-running, complex and real-world workloads, a scalable phase analysis technique is essential to capture the manifested large-scale program behavior. In this study, we show that such scalability can be achieved by applying wavelet analysis of high dimension sampled workload statistics to alleviate the counter overflow problem which can negatively affect phase classification accuracy. By exploiting the wavelet denoising capability, we show in this paper that phase classification can be performed robustly under program execution variability. To our knowledge, this work presents the first effort on using wavelets to improve scalability and robustness in phase analysis Chang-Burm Cho, Tao Li 0006 |
ISPASS | 2 |
| 2007 | An Analysis of Microarchitecture Vulnerability to Soft Errors on Simultaneous Multithreaded ArchitecturesabstractSemiconductor transient faults (i.e. soft errors) have become an increasingly important threat to microprocessor reliability. Simultaneous multithreaded (SMT) architectures exploit thread-level parallelism to improve overall processor throughput. A great amount of research has been conducted in the past to investigate performance and power issues of SMT architectures. Nevertheless, the effect of multithreaded execution on a microarchitecture's vulnerability to soft error remains largely unexplored. To address this issue, we have developed a microarchitecture level soft error vulnerability analysis framework for SMT architectures. Using a mixed set of SPEC CPU 2000 benchmarks, we quantify the impact of multithreading on a wide range of microarchitecture structures. We examine how the baseline SMT microarchitecture reliability profile varies with workload behavior, the number of threads and fetch policies. Our experimental results show that the overall vulnerability rises in multithreading architectures, while each individual thread shows less vulnerability. By considering both performance and reliability, SMT outperforms superscalar architectures. The SMT reliability and its tradeoff with performance vary across different fetch policies. With a detailed analysis of the experimental results, we point out a set of potential opportunities to reduce SMT microarchitecture vulnerability, which can serve as guidance to exploiting thread-aware reliability optimization techniques in the near future. To our knowledge, this paper presents the first effort to characterize microarchitecture vulnerability to soft error on SMT processors Wangyuan Zhang, Xin Fu 0001, Tao Li 0006, José A. B. Fortes |
ISPASS | 3 |
| 2007 | Informed Microarchitecture Design Space Exploration Using Workload DynamicsabstractProgram runtime characteristics exhibit significant variation. As microprocessor architectures become more complex, their efficiency depends on the capability of adapting with workload dynamics. Moreover, with the approaching billion-transistor microprocessor era, it is not always economical or feasible to design processors with thermal cooling and reliability redundancy capabilities that target an application's worst case scenario. Therefore, analyzing complex workload dynamics early, at the microarchitecture design stage, is crucial to forecast workload runtime behavior across architecture design alternatives and evaluate the efficiency of workload scenario- based architecture optimizations. Existing methods focus exclusively on predicting aggregated workload behavior. In this paper, we propose accurate and efficient techniques and models to reason about workload dynamics across the microarchitecture design space without using detailed cycle- level simulations. Our proposed techniques employ wavelet- based multiresolution decomposition and neural network based non-linear regression modeling. We extensively evaluate the efficiency of our predictive models in forecasting performance, power and reliability domain workload dynamics that the SPEC CPU 2000 benchmarks manifest on high-performance microprocessors with a microarchitecture design space that consists of 9 key parameters. Our results show that the models achieve high accuracy in revealing workload dynamic behavior across a large microarchitecture design space. We also demonstrate that the proposed techniques can be used to efficiently explore workload scenario-driven architecture optimizations. Chang-Burm Cho, Wangyuan Zhang, Tao Li 0006 |
MICRO | 3 |
| 2007 | OS-Aware Branch Prediction: Improving Microprocessor Control Flow Prediction for Operating Systems
Tao Li 0006, Lizy Kurian John, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Juan C. Rubio |
IEEE Trans. Computers | 1 |
| 2006 | Complexity-based program phase analysis and classificationabstractModeling and analysis of program behavior are at the foundation of computer system design and optimization. As computer systems become more adaptive, their efficiency increasingly depends on program dynamic characteristics. Previous studies have revealed that program runtime execution manifests phase behavior. Recently, methods and tools to analyze and classify program phases have also been developed. However, very few studies have been proposed so far to understand and evaluate program phases from their dynamics and complexity perspectives. In this work, we propose new methods, metrics and frameworks which aim to analyze, quantify, and classify the dynamics and complexity of program phases. Our methods use wavelet techniques to represent program phases at multiresolution scales. The cross-correlation coefficients between phase dynamics observed at different scales are then computed as metrics to quantify phase complexity. We propose to apply wavelet-based multiresolution analysis and data clustering to classify program execution into phases that exhibit similar degree of complexity. Experimental results on SPEC CPU 2000 benchmarks show that the proposed schemes classify complexity-based program phases better than currently used approaches. Chang-Burm Cho, Tao Li 0006 |
PACT | 2 |
| 2006 | OS-aware tuning: improving instruction cache energy efficiency on system workloadsabstractLow power has been considered as an important issue in instruction cache (I-cache) designs. Several studies have shown that the I-cache can be tuned to reduce power. These techniques, however, exclusively focus on user-level applications, even though there is evidence that many commercial and emerging workloads often involve heavy use of the operating system (OS). This study goes beyond previous work to explore the opportunities to design energy-efficient I-cache for system workloads. Employing a full-system experimental framework and a wide range of workloads, we characterize user and OS I-cache accesses and motivate OS-aware I-cache tuning to save power. We then present two techniques (OS-aware cache way lookup and OS-aware cache set drowsy mode) to reduce the dynamic and the static power consumption of I-cache. The proposed OS-aware cache way lookup reduces the number of parallel tag comparisons and data array read-outs for cache accesses to save dynamic I-cache power in a given operation mode. The proposed OS-aware cache set drowsy mode puts I-cache regions that are only heavily used by another operation mode to reduce leakage power. The proposed mechanisms require minimal hardware modification and addition. Simulation based experiments show that with no or negligible impact on performance, applying OS-aware tuning techniques yields significant dynamic and static power savings across the experimented applications. To our knowledge, this is the first work to explore cache power optimization by considering the interactions of application-OS-hardware. It is our belief that the proposed techniques can be applied to improve the I-cache energy efficiency on server processors mostly targeting on modern and commercial applications that heavily invoke OS activities Tao Li 0006, Lizy Kurian John |
IPCCC | 1 |
| 2006 | Characterizing Microarchitecture Soft Error Vulnerability Phase BehaviorabstractComputer systems increasingly depend on exploiting program dynamic behavior to optimize performance, power and reliability. Prior studies have shown that program execution exhibits phase behavior in both performance and power domains. Reliabilityoriented program phase behavior, however, remains largely unexplored. As semiconductor transient faults (soft errors) emerge as a critical challenge to reliable system design, characterizing program phase behavior from a reliability perspective is crucial in order to apply dynamic fault-tolerant mechanisms and to optimize performance/reliability trade-offs. In this paper, we compute run-time program vulnerability to soft errors on four microarchitecture structures (i.e. instruction window, reorder buffer, function units and wakeup table) in a high-performance out-of-order execution superscalar processor. Experimental results on the SPEC2000 benchmarks show a considerable amount of time varying behavior in reliability measurements. Our study shows that a single performance metric, such as IPC, cache miss or branch misprediction, is not a good indicator for program vulnerability. The vulnerabilities of the studied microarchitecture structures are then correlated with program code-structure and run-time events to identify vulnerability phase behavior. We observed that both program code-structure and run-time events appear promising in classifying program reliability phase behavior. Overall, performance counter based schemes achieved an average Coefficient of Variation (COV) of 3.5%, 4.5%, 4.3% and 5.7% on the instruction queue, reorder buffer, function units and the wakeup table, while basic block vectors offer COVs of 4.9%, 5.8%, 5.4% and 6% on the four studied microarchitecture structures respectively. We found that in general, tracking performance metrics performs better than tracking control flow in identifying reliability phase behavior of applications. To our knowledge, this paper is the first to characterize program reliability phase behavior at the microarchitecture level. Xin Fu 0001, James Poe, Tao Li 0006, José A. B. Fortes |
MASCOTS | 3 |
| 2005 | Approximate Global Alignment of SequencesabstractWe propose two novel dynamic programming (DP) methods that solve the approximate bounded and unbounded global alignment problems for biological sequences. Our first method solves the bounded alignment problem. It computes the distribution of the edit distance between the remaining suffixes. For a given bound k and approximation p%, it uses this distribution to prune the entries of the DP matrix that will lead to alignments with more than k edit operations with more than p% probability. Our second method addresses the unbounded global alignment problem. For each entry of the distance matrix, it dynamically computes an upper bound to the distance between the unaligned suffixes. This bound, along with the lower bound as computed for the bounded case, is then used to eliminate the entries of the distance matrix. According to our experimental results, our methods are up to three times faster than the competing methods for the bounded alignment and up to two times faster for the unbounded alignment, even with 100% approximation. Our methods use only 17-68% of the space used by the next best competitor. Tamer Kahveci, Venkatakrishnan Ramaswamy, Han Tao, Tao Li 0006 |
BIBE | 4 |
| 2005 | Workload Characterization of Bioinformatics ApplicationsabstractThe exponential growth in the amount of genomic information has spurred growing interest in large scale analysis of genetic data. Bioinformatics applications represent the increasingly important workloads. However, very few results on the behavior of these applications running on the state-of-the-art microprocessor and systems have been published. This paper proposes a suite of widely used bioinformatics applications and studies the execution characteristics of these benchmarks on a representative architecture-the Intel Pentium 4. To understand the impacts and implications of bioinformatics workloads on the microprocessor designs, we contrast the characteristics of bioinformatics workloads and the widely used SPEC 2000 integer benchmarks. Tao Li 0006, Tamer Kahveci, José A. B. Fortes |
MASCOTS | 2 |
| 2005 | Adapting branch-target buffer to improve the target predictability of java codeabstractJava programs are increasing in popularity and prevalence on numerous platforms, including high-performance general-purpose processors. The success of Java technology largely depends on the efficiency in executing the portable Java bytecodes. However, the dynamic characteristics of the Java runtime system present unique performance challenges for several aspects of microarchitecture design. In this work, we focus on the effects of indirect branches on branch-target address prediction performance. Runtime bytecode translation, just-in-time (JIT) compilation, frequent calls to the native interface libraries, and dependence on virtual methods increase the frequency of polymorphic indirect branches. Therefore, accurate target address prediction for indirect branches is very important for Java code.This paper characterizes the indirect branch behavior in Java processing and proposes an adaptive branch-target buffer (BTB) design to enhance the predictability of the targets. Our characterization shows that a traditional BTB will frequently mispredict a few polymorphic indirect branches, significantly deteriorating predictor accuracy in Java processing. Therefore, we propose a rehashable branch-target buffer (R-BTB), which dynamically identifies polymorphic indirect branches and adapts branch-target storage to accommodate multiple targets for a branch.The R-BTB improves the target predictability of indirect branches without sacrificing overall target prediction accuracy. Simulations show that the R-BTB eliminates 61% of the indirect branch mispredictions suffered with a traditional BTB for Java programs running in interpreter mode (46% in JIT mode), which leads to a 57% decrease in overall target address misprediction rate (29% in JIT mode). With an equivalent number of entries, the R-BTB also outperforms the previously proposed target cache scheme for a majority of Java programs by adapting to a greater variety of indirect branch behaviors. Tao Li 0006, Ravi Bhargava, Lizy Kurian John |
ACM Trans. Archit. Code Optim. | 1 |
| 2003 | Routine based OS-aware microprocessor resource adaptation for run-time operating system power savingabstractThe increasingly constrained power budget of today's microprocessor has resulted in a situation where power savings of all components in a system have to be taken into consideration. Operating System (OS) is a major power consumer in many modern applications execution. This paper advocates a routine based OS-aware microprocessor resource adaptation mechanism targeting run-time OS power savings. Simulation results show that compared with the existing sampling-based adaptation schemes, this novel methodology yields more attractive power and performance trade-off on the OS execution. To our knowledge, this paper is the first step to address the power saving issue of the OS itself, an increasingly important area that has been largely overlooked in the previous studies. Tao Li 0006, Lizy Kurian John |
ISLPED | 1 |
| 2003 | Run-time modeling and estimation of operating system power consumptionabstractThe increasing constraints on power consumption in many computing systems point to the need for power modeling and estimation for all components of a system. The Operating System (OS) constitutes a major software component and dissipates a significant portion of total power in many modern application executions. Therefore, modeling OS power is imperative for accurate software power evaluation, as well as power management (e.g. dynamic thermal control and equal energy scheduling) in the light of OS-intensive workloads. This paper characterizes the power behavior of a commercial OS across a wide spectrum of applications to understand OS energy profiles and then proposes various models to cost-effectively estimate its run-time energy dissipation. The proposed models rely on a few simple parameters and have various degrees of complexity and accuracy. Experiments show that compared with cycle-accurate full-system simulation, the model can predict cumulative OS energy to within 1% accuracy for a set of benchmark programs evaluated on a high-end superscalar microprocessor. When applied to track run-time OS energy profiles, the proposed routine level OS power model offers superior accuracy than a simpler, flat OS power model, yielding per-routine estimation error of less than 6%. The most striking observation is the strong correlation between power consumption and the instructions per cycle (IPC) during OS routine executions. Since tools and methodology to measure IPC exist on modern microprocessors, the proposed models can estimate OS power for run-time dynamic thermal and energy management. Tao Li 0006, Lizy Kurian John |
SIGMETRICS | 1 |
| 2002 | Understanding and improving operating system effects in control flow predictionabstractMany modern applications result in a significant operating system (OS) component. The OS component has several implications including affecting the control flow transfer in the execution environment. This paper focuses on understanding the operating system effects on control flow transfer and prediction, and designing architectural support to alleviate the bottlenecks. We characterize the control flow transfer of several emerging applications on a commercial operating system. We find that the exception-driven, intermittent invocation of OS code and the user/OS branch history interference increase the misprediction in both user and kernel code.We propose two simple OS-aware control flow prediction techniques to alleviate the destructive impact of user/OS branch interference. The first one consists of capturing separate branch correlation information for user and kernel code. The second one involves using separate branch prediction tables for user and kernel code. We study the improvement contributed by the OS-aware prediction to various branch predictors ranging from simple Gshare to more elegant Agree, Multi-Hybrid and Bi-Mode predictors. On 32K entries predictors, incorporating OS-aware techniques yields up to 34%, 23%, 27% and 9% prediction accuracy improvement in Gshare, Multi-Hybrid, Agree and Bi-Mode predictors, resulting in up to 8% execution speedup. Tao Li 0006, Lizy Kurian John, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Juan C. Rubio |
ASPLOS | 1 |
| 2002 | Rehashable BTB: An Adaptive Branch Target Buffer to Improve the Target Predictability of Java Code
Tao Li 0006, Ravi Bhargava, Lizy Kurian John |
HiPC | 1 |
| 2002 | Using Complete Machine Simulation for Software Power Estimation: The SoftWatt ApproachabstractPower dissipation has become one of the most critical factors for the continued development of both high-end and low-end computer systems. We present a complete system power simulator, called SoftWatt, that models the CPU, memory hierarchy, and a low-power disk subsystem and quantifies the power behavior of both the application and operating system. This tool, built on top of the SimOS infrastructure, uses validated analytical energy models to identify the power hotspots in the system components, capture relative contributions of the user and kernel code to the system power profile, identify the power-hungry operating system services and characterize the variance in kernel power profile with respect to workload. Our results using Spec JVM98 benchmark suite emphasize the importance of complete system simulation to understand the power impact of architecture and operating system on application execution. Sudhanva Gurumurthi, Anand Sivasubramaniam, Mary Jane Irwin, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Tao Li 0006, Lizy Kurian John |
HPCA | 6 |
| 2001 | Understanding control flow transfer and its predictability in java processingabstractAn in-depth look and understanding of control flow transfer and its predictability can guide architects to adapt control flow prediction hardware in Java processing or finely tune the performance of JVM software on general purpose machines. To our knowledge, this paper provides the first insight of branch behavior on a standard Java Virtual Machine with real workloads. Employing a complete system simulation environment, we profile branch execution characteristics and quantify the performance of a wide range of prediction schemes on both user and kernel code. The impact of different JVM styles (JIT compiler and interpreter) on branch behavior is also studied. We find that: (1) Kernel branches constitute a significant portion of total branch execution in Java processing; (2) Kernel and user code favor different prediction mechanisms; (3) Java processing exercises fairly large number of branch sites and large control flow footprint compared with the execution of benchmarks such as SPECInt95; (4) A major part of the dynamic indirect branches are multiple target (polymorphic) branches. Target addresses of indirect branches, especially those in interpreting mode are highly interleaved and cause high BTB misprediction. 1. Tao Li 0006, Lizy Kurian John |
ISPASS | 1 |
| 2001 | ADir_pNB: A Cost-Effective Way to Implement Full Map Directory-Based Cache Coherence ProtocolsabstractDirectories have been used to maintain cache coherency in shared memory multiprocessors with private caches. The traditional full map directory tracks the exact caching status for each shared memory block and is designed to be efficient and simple. Unfortunately, the inherent directory size explosion makes it unsuitable for large-scale multiprocessors. In this paper, we propose a new directory scheme, dubbed associative full map directory (ADir/sub p/NB) which reduces the directory storage requirement. The proposed ADir/sub p/NB uses one directory entry to maintain the sharing information for a set of exclusively cached memory blocks in a centralized linked list style. By implementing dynamic cache pointer allocation, reclamation, and replacement hints, ADir/sub p/NB can be implemented as "a full map directory with lower directory memory cost". Our analysis indicates that, on a typical architectural paradigm, ADir/sub p/NB reduces memory overhead of a traditional full map directory by up to 70-80 percent. In addition to the low memory overhead, we show that the proposed scheme can be implemented with appropriate protocol modification and hardware addition. Simulation studies indicate that ADir/sub p/NB can achieve a competitive performance with the Dir/sub p/NB. Compared with limited directory schemes, ADir/sub p/NB shows more stable and robust performance results on applications across a spectrum of memory sharing and access patterns due to the elimination of directory overflows. We believe that ADir/sub p/NB can be employed as a design alternative of full map directory for moderately large-scale and fine-grain shared memory multiprocessors. Tao Li 0006, Lizy Kurian John |
IEEE Trans. Computers | 1 |
| 2000 | Using complete system simulation to characterize SPECjvm98 benchmarksabstractComplete system simulation to understand the influence of architecture and operating systems on application execution has been identified to be crucial for systems design. While there have been previous attempts at understanding the architectural impact of Java programs, there has been no prior work investigating the operating system (kernel) activity during their executions. This problem is particularly interesting in the context of Java since it is not only the application that can invoke kernel services, but so does the underlying Java Virtual Machine (JVM) implementation which runs these programs. Further, the JVM style (JIT compiler or interpreter) and the manner in which the different JVM components (such as the garbage collector and class loader) are exercised, can have a significant impact on the kernel activities. Tao Li 0006, Lizy Kurian John, Narayanan Vijaykrishnan, Anand Sivasubramaniam, Jyotsna Sabarinathan, Anupama Murthy |
ICS | 1 |