Xin Fu 0001

dblp:18/2495-1 · DBLP profile ↗
← Back
86ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0002-9458-4769ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 66 · 6 first-author · 24 since 2021Computer networks · 8 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Security and privacy · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Wireless-Aware Energy-Efficient Federated Learning Over Mobile Devices via Algorithm and Hardware Co-Design
abstract
Energy efficiency is essential for federated learning (FL) over mobile devices and its potential prosperous applications. Different from existing communication efficient FL research efforts, which regard communication energy consumption as the bottleneck, we have observed that with ever increasing wireless transmission speed (e.g., Wi-Fi 5 or 5G), the energy consumption of wireless communications for model updates in FL is significantly reduced and sometimes is smaller than that of local on-device training. Motivated by such observations, in this paper, we propose a high-speed wireless communications inspired energy efficient federated learning over mobile devices (EEFL), whose goal is to reduce the overall energy consumption (computing + communication). In particular, we design a novel energy-aware adaptive local update policy for mobile devices by jointly considering FL performance and energy saving of high-speed wireless transmissions. Furthermore, given the device’s local update policy in each FL global round, we advance the dynamic voltage and frequency scaling (DVFS) strategy to minimize local training’s energy consumption by keeping GPU and CPU working at appropriate frequencies without triggering thermal throttling. Extensive experimental results with various learning models, datasets, and wireless transmission environments demonstrate the proposed EEFL’s superiority over the peer designs in terms of energy efficiency.
Rui Chen 0026, Qiyu Wan, Xinyue Zhang 0001, Xiaoqi Qin, Yan-Zhao Hou, Di Wang 0015, Xin Fu 0001, Miao Pan
IEEE Trans. Netw.7
2026 Accelerating Federated Edge Learning via Wireless and Heterogeneity Aware Subnetwork Scheduling
abstract
As a popular distributed learning paradigm, federated learning (FL) over mobile devices fosters numerous applications, while their practical deployment is hindered by participating devices’ computing and communication heterogeneity. Some pioneering research efforts proposed to extract subnetworks from the global model, and assign as large a subnetwork as possible to the device for local training based on its full computing and communications capacity. Although such fixed size subnetwork assignment enables FL training over heterogeneous mobile devices, it is unaware of (i) the dynamic changes of devices’ communication and computing conditions and (ii) FL training progress and its dynamic requirements of local training contributions, both of which may cause very long FL training delay. Motivated by those dynamics, in this paper, we develop a wireless and heterogeneity aware latency efficient FL (WHALE-FL) approach to accelerate FL training through adaptive subnetwork scheduling. Instead of sticking to the fixed size subnetwork, WHALE-FL introduces a novel subnetwork selection utility function to capture device and FL training dynamics, and guides the mobile device to adaptively select the subnetwork size for local training based on (a) its computing and communication capacity, (b) its dynamic computing and/or communication conditions, and (c) FL training status and its corresponding requirements for local training contributions. We provide a theoretical convergence analysis for WHALE-FL with heterogeneous subnetwork assignment, based on which subnetwork structures can be dynamically optimized to reduce the resulting gap to standard full-model FL. Our evaluation shows that, compared with peer designs, WHALE-FL effectively accelerates FL training without sacrificing learning accuracy.
Liang Li 0021, Jiaxiang Geng, Huai-An Su, Xiaoqi Qin, Yan-Zhao Hou, Hao Wang 0022, Xin Fu 0001, Miao Pan
IEEE Trans. Netw.7
2026 Importance Aware Undervolting for Robust Neural Network Training
abstract
Convolutional Neural Network (CNN) is a powerful tool that has been extensively applied to many different applications. However, recent developments at CNN have revealed its vulnerability against adversarial example attacks. By introducing visually undetectable noise to the input image, an adversarial example attack can cause the CNN classifier to make false predictions. Multiple approaches have been proposed to defend against adversarial samples, one of which focuses on injecting noise into CNN during training. However, the existing method cannot generate noise efficiently and introduces extra time and energy overhead. In this paper, we propose an Importance-Aware undervolting training framework to improve CNN robustness. The undervolting technique is employed during training for noise generation at negligible overhead. Meanwhile, we observe that the neuron importance and bit importance in hardware can be leveraged during undervolting CNN training for controllable and flexible noise injection, which improves robustness. We design a position-aware bit mapping method in the memory unit by allocating data bits based on significance. And an importanceaware processing element (PE) mapping is also proposed at the computation unit for noise restriction. Our approach regularizes the noise injected into CNN and serves as an efficient method to defend against adversarial example attacks with significant energy savings. The proposed framework is evaluated through both FPGA implementation and software simulation. The experiment results show that our importance-aware undervolting CNN training achieves 47.8% adversarial accuracy at PGD-10 attack and 47.0% training energy savings.
Xin Fu 0001
IEEE Trans. Sustain. Comput.3
2025 Achieving Lightweight Super-Resolution for Real-Time Computer Graphics
abstract
Image super-resolution (SR) is essential for bridging the gap between modern hardware and real-time computer graphics (CG) applications. It reduces CG workload by allowing low-resolution rendering, with original quality restored later via mathematical operations or machine learning. However, recent learning-based SR methods often rely on complex models, demanding high computational resources and undermining the benefits of reduced rendering workload. Our qualitative and quantitative analysis of the SR process and rendering reveals that readily accessible rendering information can significantly enhance neural network design by serving as additional features. To capitalize on this, we propose CGSR, an optimization framework designed for lightweight real-time super-resolution. CGSR utilizes rendering information to boost both network extensibility and efficiency. It utilizes progressively available rendering information from the pipeline, which arrives earlier than the rendered frame, enabling pre-processing and masking of latency. These features are then integrated into a selected SR network backbone to form a CG-enhanced network. This network is further optimized and refined into a CG-optimized version using neural architecture search (NAS). To improve runtime performance, CGSR also employs rendering-aware hybrid pruning, which dynamically prunes the network based on temporal rendering data. Evaluation results show that CGSR significantly reduces parameter size, multi-add operations, and inference time while maintaining high SR quality across various backbone SR networks.
Yu Wen 0003, Chenhao Xie 0001, Xin Fu 0001
AAAI4
2025 WHALE-FL: Wireless and Heterogeneity Aware Latency Efficient Federated Learning over Mobile Devices via Adaptive Subnetwork Scheduling
abstract
As a popular distributed learning paradigm, federated learning (FL) over mobile devices fosters numerous applications, while their practical deployment is hindered by participating devices' computing and communication heterogeneity. Some pioneering research efforts proposed to extract subnetworks from the global model, and assign as large a subnetwork as possible to the device for local training based on its full computing capacity. Although such fixed size subnetwork assignment enables FL training over heterogeneous mobile devices, it is unaware of (i) the dynamic changes of devices' communication and computing conditions and (ii) FL training progress and its dynamic requirements of local training contributions, both of which may cause very long FL training delay. Motivated by those dynamics, in this paper, we develop a wireless and heterogeneity aware latency efficient FL (WHALE-FL) approach to accelerate FL training through adaptive subnetwork scheduling. Instead of sticking to the fixed size subnetwork, WHALE-FL introduces a novel subnetwork selection utility function to capture device and FL training dynamics, and guides the mobile device to adaptively select the subnetwork size for local training based on (a) its computing and communication capacity, (b) its dynamic computing and/or communication conditions, and (c) FL training status and its corresponding requirements for local training contributions. Our evaluation shows that, compared with peer designs, WHALE-FL effectively accelerates FL training without sacrificing learning accuracy.
Huai-An Su, Jiaxiang Geng, Liang Li 0021, Xiaoqi Qin, Yan-Zhao Hou, Hao Wang 0022, Xin Fu 0001, Miao Pan
AAAI7
2025 MMDFL: Multi-Model-based Decentralized Federated Learning for Resource-Constrained AIoT Systems
abstract
Along with the prosperity of Artificial Intelligence (AI) techniques, more and more Artificial Intelligence of Things (AIoT) applications adopt Federated Learning (FL) to enable collaborative learning without compromising the privacy of devices. Since existing centralized FL methods suffer from the problems of single-point-offailure and communication bottleneck caused by the parameter server, we are witnessing an increasing use of Decentralized Federated Learning (DFL), which is based on Peer-to-Peer (P2P) communication without using a global model. However, DFL still faces three major challenges, i.e., limited computing power and network bandwidth of resource-constrained devices, non-Independent and Identically Distributed (non-IID) device data, and all-neighbor-dependent knowledge aggregation operations, all of which greatly suppress the learning potential of existing DFL methods. To address these problems, this paper presents an efficient DFL framework named MMDFL based on our proposed multi-model-based learning and knowledge aggregation mechanism. Specifically, MMDFL adopts multiple traveler models, which perform local training individually along their traversed devices, accelerating and maximizing knowledge learning and sharing among devices. Moreover, based on our proposed device selection strategy, MMDFL enables each traveler to adaptively explore its next best neighboring device to further enhance the DFL training performance, taking into account issues of data heterogeneity, limited resources and catastrophic forgetting phenomenon. Experimental results from simulation and a real testbed show that, compared with state-of-the-art DFL methods, MMDFL can not only significantly reduce the communication overhead but also achieve better overall classification performance for both IID and non-IID scenarios.
Dengke Yan, Yanxin Yang, Ming Hu 0003, Xin Fu 0001, Mingsong Chen 0001
DAC4
2025 MVMD: A Multi-View Approach for Enhanced Mirror Detection
abstract
In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models. As 3D reconstruction typically relies on multi-view images to capture different perspectives of a scene, detecting and labeling mirrors in multi-view images before reconstruction can effectively address this issue. However, existing methods focus solely on single-image detection, overlooking the rich information provided by multi-view setups. To overcome this limitation, we propose MVMD, a novel Multi-View Mirror Detection method, along with the first database specifically designed for mirror detection in multi-view scenes. The design of MVMD is grounded in the inherent associations between objects seen from different views and those reflected inside and outside of mirrors. These relationships are learned through cross-and self-attention mechanisms. MVMD consists of three key blocks: the Inter-Views Block tracks the shifts of objects within mirrors caused by changes in viewpoint; the Intra-View Block detects object reflections inside mirrors; and the Refinement Block sharpens mirror boundaries and enhances detected details. Experimental results show that our method improves accuracy by up to 2.6% and IoU by up to 11.1%, compared to single-image mirror detection techniques. This substantial improvement makes MVMD particularly effective for computer vision tasks, especially in enhancing the accuracy of 3D reconstruction in mirror-dense environments. Code and data are available at: https://github.com/mvmdwacv25.
Yidan Shen, Yu Wen 0003, Xin Fu 0001
WACV4
2025 Bit-Flip Induced Latency Attacks in Object Detection
abstract
Deep learning and computer vision have experienced significant advancements, particularly in critical applications such as autonomous driving and real-time surveillance, where object detection (OD) plays a pivotal role. Ensuring the accuracy and speed of these systems is paramount to prevent accidents or failures. Recently, latency-based attacks have emerged as a new threat, driven by the essential need for real-time performance in various applications. These attacks target model responsiveness to disrupt system performance without necessarily compromising accuracy. Our preliminary experiments show that introducing just a few bit flips to key parameters in OD models can significantly increase latency, degrading performance. Meanwhile, recent advancements in memory-based attacks, such as Row Hammer [18], demonstrate the ability to conveniently introduce bit flips at desired locations without physical hardware interaction. Based on the observations, we propose a novel attack on OD models that leverages row-hammer to introduce bit-flips via side channels, targeting the non-maximum suppression (NMS) filter and significantly increasing latency. Unlike previous methods that modify input data, our technique ensures efficiency by minimizing bit-flips through critical path exploitation and achieves practical applicability with only a subset of validation data. Experiments across various datasets and models validate our approach, demonstrating latency increases up to 71.6 ms (20.4×) with just 31 bit-flips.
Manojna Sistla, Yu Wen 0003, Aamir Bader Shah, Chenpei Huang, Xuqing Wu 0001, Jiefu Chen, Miao Pan, Xin Fu 0001
WACV9
2025 DAFL: Device-to-Device Transmissions for Delay-Efficient Federated Learning Over Mobile Devices
abstract
Federated learning (FL) over mobile devices is an emerging distributed learning paradigm for numerous delay sensitive applications. In FL, the training delay is composed of the computing and communication delay. Some of the participating mobile devices may have slow local computing or wireless communications, which results in high FL training delay. Intuitively, if fast devices help slow ones, the FL training delay can potentially be reduced. However, helping each other among devices requires frequent transmissions and may cause additional delay. Fortunately, we observe that device-to-device (D2D) transmission, a fast and direct transmission, may be applied between device pairs to mitigate the additional delay from frequent transmissions. Inspired by those observations, we develop the D2D transmission assisted FL (DAFL), a novel FL scheme to improve the training delay over mobile devices. Briefly, we first put the eligible mobile devices into pairs, assigning each pair to one of the four types of relation: 1) similar computing, large communication gap; 2) similar communication, large computing gap; 3) one with faster computing and the other with faster communication; and 4) one with both faster computing and communication. We design the process for each type of device pair to: 1) improve the transmission delay of each pair, by letting the fast device help with the model parameters transmission to the server and 2) improve the computing delay by splitting learning task between paired devices. The emulation results demonstrate that DAFL surpasses existing peer designs in terms of reducing training delay by more than 20%.
Huai-An Su, Pavana Prakash, Rui Chen 0026, Yanmin Gong 0001, Rong Yu 0001, Xin Fu 0001, Miao Pan
IEEE Internet Things J.6
2025 NANI: Energy-efficient Neuron-Aware hardware Noise Injection for adversarial defense using undervolting
Qiyu Wan, Jing Wang 0055, Mingsong Chen 0001, Lu Peng 0001, Xin Fu 0001
J. Syst. Archit.6
2025 Evaluating GPU's Instruction-Level Error Characteristics Under Low Supply Voltages
abstract
Supply voltage underscaling has been an effective approach to improve the energy-efficiency of modern high-performance processors, such as GPUs. However, energy efficiency and reliability are two sides of a trade-off. Undervolting will inevitably undermine reliability, since it reduces chip manufacturers’ voltage guardbands that is designed to ensure correct operations under worst-case scenarios. To achieve optimal energy efficiency while maintaining enough reliability, it is necessary to deeply understand the error characteristics caused by undervolting. Unlike previous works which focus mostly on program level, we perform the first comprehensive instruction-level voltage margin and error characteristics evaluation for GPU architectures. We systematically measure the error probability and patterns of GPU instructions during undervolting. Then, we also analyze the impact of locations (SMs, threads, and bits) and operand data values on the error characteristics. Based on our observations, we reduce the voltage to the minimum safe limit for different instructions which achieves 18.37% energy saving, and we further propose an error detection strategy which reduces the performance and energy overhead by 14.8% with negligible 0.01% degradation for error detection rate.
Jingweijia Tan, Jiashuo Wang, Kaige Yan, Xiaohui Wei 0002, Xin Fu 0001
IEEE Trans. Computers5
2025 AR-Light: Enabling Fast and Lightweight Multi-User Augmented Reality via Semantic Segmentation and Collaborative View Synchronization
abstract
Multi-user Augmented Reality (MuAR) allows multiple users to interact with shared virtual objects, facilitated by exchanging environment information. Current MuAR systems rely on 3D point clouds for real-world analysis, view synchronization, object rendering, and movement tracking. However, the complexity of 3D point clouds leads to significant processing delays, with approximately 80% of overhead in commercial frameworks. This hampers usability and degrades user experience. Our analysis reveals that maintaining the facing side of the real-world scene in a stable environment provides sufficient information for virtual object placement and rendering. To address this, we introduce a lightweight quadtree structure, representing 2D scenes through semantic segmentation and geometry, as an alternative to 3D point clouds. Additionally, we propose a novel correction method to handle potential shifts in virtual object placement during view synchronization among users. Combining all designs, we implement a fast and lightweight MuAR framework namedAR-Lightand test our framework on commercial AR devices. The evaluation results on real-world applications demonstrate that AR-Light can achieve high performance in various real-world scenes while maintaining a comparable virtual object placement accuracy.
Yu Wen 0003, Aamir Bader Shah, Ruizhi Cao, Jiefu Chen, Xuqing Wu 0001, Chenhao Xie 0001, Xin Fu 0001
IEEE Trans. Computers8
2025 Generalized Transitional Markov Chain Monte Carlo Sampling Technique for Bayesian Inversion of Electromagnetic Data
abstract
In the context of Bayesian inversion for scientific and engineering modeling, Markov chain Monte Carlo (MCMC) sampling strategies have become the benchmark due to their flexibility and robustness in dealing with arbitrary posterior probability density functions (PDFs). However, these algorithms have been shown to be inefficient when sampling from high-dimensional posterior distributions or exhibit multimodality and/or strong parameter correlations. In such contexts, transitional MCMC (TMCMC) provides a more efficient alternative. Despite the recent applicability for Bayesian updating and model selection across a variety of disciplines, TMCMC may require a prohibitive number of tempering stages when the prior pdf is significantly different from the target posterior. Furthermore, the need to start with an initial set of samples from the prior distribution may present a challenge when dealing with implicit priors, e.g., based on feasible regions. Finally, TMCMC cannot be used for inverse problems with improper prior PDFs that represent a lack of prior knowledge on all or a subset of parameters. A generalization of TMCMC is proposed that alleviates such challenges and limitations toward providing a more robust and efficient tempering sampling strategy. We present convergence analysis, proving that the distance between the intermediate distributions and the target posterior distribution monotonically decreases as the algorithm proceeds. We also demonstrate the advantages of the proposed generalization through a series of test problems and an engineering application in the oil and gas industry.
Mohammad Khalil, Tommie Catanach, Cosmin Safta, Jiajia Sun, Xuqing Wu 0001, Xin Fu 0001, Jiefu Chen, Yueqin Huang
IEEE Trans. Geosci. Remote. Sens.9
2024 Tuning Quantum Computing Privacy through Quantum Error Correction
abstract
Quantum computing is a promising paradigm for efficiently solving large and high-complexity problems. However, ensuring privacy within this quantum computing necessitates innovative approaches. Existing research has introduced the concept of quantum differential privacy (QDP) to protect data privacy in quantum computing by leveraging quantum noise. Yet, this method faces limitations due to the fixed and uncontrollable nature of the inherent noise, which directly affects the privacy budget of QDP. Addressing this critical gap, our study proposes a novel approach that utilizes quantum error correction (QEC) techniques not only to mitigate quantum computing errors but also to adjust QDP protection levels precisely. By selectively applying QEC to single or multiple qubit gates, we introduce a method to manipulate the quantum noise error rate effectively. Moreover, we derive a new formula for calculating the overall error rate in a quantum circuit and the adjusted privacy budget after QEC operation. Through extensive numerical simulations, we validate the efficacy of utilizing QEC in tuning privacy protection levels within quantum computing.
Keyi Ju, Manojna Sistla, Xinyue Zhang 0001, Aohan Li, Xiaoqi Qin, Xin Fu 0001, Miao Pan
GLOBECOM7
2024 A New Routing Strategy to Improve Success Rates of Quantum Computers
abstract
In the current noisy intermediate-scale quantum (NISQ) Era, Quantum Computing faces significant challenges due to noise, which severely restricts the application of computing complex algorithms. Superconducting quantum chips, one of the pioneer quantum computation technologies, introduce additional noise when moving qubits to adjacent locations for operation on designated two-qubit gates. The current compilers rely on decision models that either count the swap gates or multiply the gate errors when choosing swap paths at the routing stage. Our research has unveiled the overlooked situations for error propagations through the circuit, leading to accumulations that may affect the final output.
Fang Qi, Xin Fu 0001, Xu Yuan 0001, Nian-Feng Tzeng, Lu Peng 0001
ACM Great Lakes Symposium on VLSI2
2024 Safe Offline-to-Online Multi-Agent Decision Transformer: A Safety Conscious Sequence Modeling Approach
abstract
We introduce the Safe Offline-to-Online Multi-Agent Decision Transformer (SO2-MADT), an innovative framework that revolutionizes safety considerations in Multi-agent Reinforcement Learning (MARL) through a novel sequence modeling approach. Leveraging the dynamic capabilities inherent in Decision Transformers, our methodology seamlessly incorporates safety protocols as a cornerstone element, ensuring secure operations throughout both the offline pre-training phase and the adaptive online fine-tuning phase. At the core of our framework lie two pivotal innovations: the Safety-To-Go (STG) token, embedding safety at a macro level, and the Agent Prioritization Module (APM), facilitating explicit credit assignment at a micro level. Through extensive testing against the challenging environments of the StarCraft Multi-Agent Challenge (SMAC) and Multi-agent MuJoCo, our SO2-MADT not only excels in offline pre-training but also demonstrates superior performance during online fine-tuning, without any degradation in performance. The implications of our work provide a pathway for deployment in critical real-world applications where safety is paramount and non-negotiable. The code is available at https://github.com/shahaamirbader/SO2-MADT.
Aamir Bader Shah, Yu Wen 0003, Jiefu Chen, Xuqing Wu 0001, Xin Fu 0001
IROS5
2024 HSAS: Efficient task scheduling for large scale heterogeneous systolic array accelerator cluster
Kaige Yan, Yanshuang Song, Jingweijia Tan, Xiaohui Wei 0002, Xin Fu 0001
Future Gener. Comput. Syst.6
2024 Efficient one-shot Neural Architecture Search with progressive choice freezing evolutionary search
Qiyu Wan, Yu Wen 0003, Mingsong Chen 0001, Jingweijia Tan, Kaige Yan, Xin Fu 0001
Neurocomputing8
2024 Towards High Performance QNNs via Distribution-Based CNOT Gate Reduction
abstract
Quantum Neural Networks (QNNs) are one of the most promising applications that can be implemented on NISQ-era quantum computers. In this study, we observe that QNNs often suffer from gate redundancy, which hugely declines the performance and accuracy of the network. Even state-of-the-art architecture search techniques like QuantumNAS do not completely alleviate this problem. Especially, we find that CNOT gates are major contributors to the execution delay and noise in quantum circuits, and there are many redundant CNOT gates in the QNN post-training. This motivates us to propose a novel distribution-based greedy-search circuit optimization technique that can be employed after the completion of the training process. Our technique significantly reduces the number of CNOT gates in QNNs without affecting the accuracy of the network. With this technique, we have achieved an average of 3× improvement in execution time while reaching a maximum of 12.4× improvement.
Manojna Sistla, Xin Fu 0001
ACM Trans. Archit. Code Optim.3
2024 Enhancing Neural Network Reliability: Insights From Hardware/Software Collaboration With Neuron Vulnerability Quantization
abstract
Ensuring the reliability of deep neural networks (DNNs) is paramount in safety-critical applications. Although introducing supplementary fault-tolerant mechanisms can augment the reliability of DNNs, an efficiency tradeoff may be introduced. This study reveals the inherent fault tolerance of neural networks, where individual neurons exhibit varying degrees of fault tolerance, by thoroughly exploring the structural attributes of DNNs. We thereby develop a hardware/software collaborative method that guarantees the reliability of DNNs while minimizing performance degradation. We introduce the neuron vulnerability factor (NVF) to quantify the susceptibility to soft errors. We propose two efficient methods that leverage the NVF to minimize the negative effects of soft errors on neurons. First, we present a novel computational scheduling scheme. By prioritizing error-prone neurons, the expedited completion of their computations is facilitated to mitigate the risk of neural computing errors that arise from soft errors without sacrificing efficiency. Second, we propose the NVF-guided heterogeneous memory system. We employ variable-strength error-correcting codes and tailor their error-correction mechanisms to the vulnerability profile of specific neurons to ensure a highly targeted approach for error mitigation. Our experimental results demonstrate that the proposed scheme enhances the neural network accuracy by 18% on average, while significantly reducing the fault-tolerance overhead.
Jing Wang 0055, Jinbin Zhu, Xin Fu 0001, Di Zang, Keyao Li, Weigong Zhang
IEEE Trans. Computers3
2023 DAFL: Delay Efficient Federated Learning over Mobile Devices via Device-to-Device Transmissions
abstract
Federated learning (FL) over mobile devices is an emerging distributed learning paradigm for numerous delay sensitive applications. In FL, the training delay is composed of the computing and communication delay. Some of the participating mobile devices may have slow local computing or wireless communications, which results in high FL training delay. Intuitively, if fast devices help slow ones, the FL training delay can potentially be reduced. However, helping each other among devices requires frequent transmissions and may cause additional delay. Fortunately, we observe that Device-to-Device (D2D) transmission, a fast and direct transmission, may be applied between device pairs to mitigate the additional delay from frequent transmissions. Inspired by those observations, we develop the D2D transmission assisted FL (DAFL), a novel FL scheme to improve the training delay over mobile devices. Briefly, we first put the eligible mobile devices into pairs, each pair consisting of a fast and a slow device. Then, we apply D2D transmission between each device pair to: (1) improve the transmission delay of each pair, by letting the fast device help with the model parameters transmission to the server, and (2) improve the computing delay by splitting learning task between paired devices. The emulation results demonstrate that DAFL surpasses existing peer designs in terms of reducing training delay by more than 20%.
Huai-An Su, Pavana Prakash, Rui Chen 0026, Yanmin Gong 0001, Rong Yu 0001, Xin Fu 0001, Miao Pan
GLOBECOM6
2023 Post0-VR: Enabling Universal Realistic Rendering for Modern VR via Exploiting Architectural Similarity and Data Sharing
abstract
To provide users with a fully immersive environment, VR post-processing, which adds numerous realistic effects on the frame after rendering, plays a key role in modern VR systems. Current post-processing is processed separately from normal rendering by the graphics processing unit (GPU). As a result, the GPU needs to first render a high-resolution frame and then add the post-processing effects within a very short time frame. Our in-depth experimental results on commercial VR products demonstrate that the post-processing in VR applications extends the VR frame time by approximately 2X on average. Furthermore, the ever-increasing resolution requirements of modern VR significantly increase the workloads for post-processing in the execution pipeline. This long delay causes VR real-time execution to frequently miss the critical frame-time deadline, thus hurting users’ quality of experience.Based on the analysis of VR post-processing workflow and its common realistic effects, we observe that post-processing shares the same hardware pipeline with normal rendering, and even reuses the intermediate data produced by normal rendering. To fully utilize this hardware-level similarity and capture the data locality, we propose a novel universal realistic rendering architecture for VR, named Post0-VR, which eliminates post-processing by directly merging the common realistic effects into the normal rendering process. Based on our newly proposed VR architecture design, we further propose a dynamic accuracy adjustment method to simplify the normal rendering without hurting users’ perception. The evaluation results on real-world applications demonstrate that Post0-VR can support different types of realistic effects while significantly improving the overall VR rendering performance.
Yu Wen 0003, Chenhao Xie 0001, Shuaiwen Song, Xin Fu 0001
HPCA4
2023 Workie-Talkie: Accelerating Federated Learning by Overlapping Computing and Communications via Contrastive Regularization
abstract
Federated learning (FL) over mobile edge devices is a promising distributed learning paradigm for various mobile applications. However, practical deployment of FL over mobile devices is very challenging because (i) conventional FL incurs huge training latency for mobile edge devices due to interleaved local computing and communications of model updates, (ii) there are heterogeneous training data across mobile edge devices, and (iii) mobile edge devices have hardware heterogeneity in terms of computing and communication capabilities.To address aforementioned challenges, in this paper, we propose a novel "workie-talkie" FL scheme, which can accelerate FL’s training by overlapping local computing and wireless communications via contrastive regularization (FedCR). FedCR can reduce FL’s training latency and almost eliminate straggler issues since it buries/embeds the time consumption of communications into that of local training. To resolve the issue of model staleness and data heterogeneity co-existing, we introduce class-wise contrastive regularization to correct the local training in FedCR. Besides, we jointly exploit contrastive regularization and subnetworks to further extend our FedCR approach to accommodate edge devices with hardware heterogeneity. We deploy FedCR in our FL testbed and conduct extensive experiments. The results show that FedCR outperforms its status quo FL approaches on various datasets and models.
Rui Chen 0026, Qiyu Wan, Pavana Prakash, Lan Zhang 0005, Xu Yuan 0001, Yanmin Gong 0001, Xin Fu 0001, Miao Pan
ICCV7
2023 NAS-SE: Designing A Highly-Efficient In-Situ Neural Architecture Search Engine for Large-Scale Deployment
abstract
The emergence of Neural Architecture Search (NAS) enables an automated neural network development process that potentially replaces manually-enabled machine learning expertise. A state-of-the-art NAS method, namely One-Shot NAS, has been proposed to drastically reduce the lengthy search time for a wide spectrum of conventional NAS methods. Nevertheless, the search cost is still prohibitively expensive for practical large-scale deployment with real-world applications. In this paper, we reveal that the fundamental cause for inefficient deployment of One-Shot NAS in both single-device and large-scale scenarios originates from the massive redundant off-chip weight access during the numerous DNN inference in sequential searching. Inspired by its algorithmic characteristics, we depart from the traditional CMOS-based architecture designs and propose a promising processing-in-memory design alternative to perform in-situ architecture search, which helps fundamentally address the redundancy issue. Moreover, we further discovered two major performance challenges of directly porting the searching process onto the existing PIM-based accelerators: severe pipeline contention and resource under-utilization. By leveraging these insights, we propose the first highly-efficient in-situ One-Shot NAS search engine design, named NAS-SE, for both single-device and large-scale deployment scenarios. NAS-SE is equipped with a two-phased network diversification strategy for eliminating resource contention, and a novel hardware mapping scheme for boosting the resource utilization by an order of magnitude. Our extensive evaluation demonstrates that NAS-SE significantly outperforms the state-of-the-art digital-based customized NAS accelerator (NASA) with an average speedup of 8.8 × and energy-efficiency improvement of 2.05 ×.
Qiyu Wan, Jing Wang 0055, Shuaiwen Song, Xin Fu 0001
MICRO5
2023 EEFL: High-Speed Wireless Communications Inspired Energy Efficient Federated Learning over Mobile Devices
abstract
Energy efficiency is essential for federated learning (FL) over mobile devices and its potential prosperous applications. Different from existing communication efficient FL research efforts, which regard communication energy consumption as the bottleneck, we have observed that with ever increasing wireless transmission speed (e.g., Wi-Fi 5 or 5G), the energy consumption of wireless communications for model updates in FL is significantly reduced and sometimes is smaller than that of local on-device training. Motivated by such observations, in this paper, we propose a high-speed wireless communications inspired energy efficient federated learning over mobile devices (EEFL), whose goal is to reduce the overall energy consumption (computing + communication). In particular, we design a novel energy-aware adaptive local update policy for mobile devices by jointly considering FL performance and energy saving of high-speed wireless transmissions. Furthermore, given the device's local update policy in each FL global round, we advance the dynamic voltage and frequency scaling (DVFS) strategy to minimize local training's energy consumption by keeping GPU and CPU working at appropriate frequencies without triggering thermal throttling. Extensive experimental results with various learning models, datasets, and wireless transmission environments demonstrate the proposed EEFL's superiority over the peer designs in terms of energy efficiency.
Rui Chen 0026, Qiyu Wan, Xinyue Zhang 0001, Xiaoqi Qin, Yan-Zhao Hou, Di Wang 0015, Xin Fu 0001, Miao Pan
MobiSys7
2023 Saca-FI: A microarchitecture-level fault injection framework for reliability analysis of systolic array based CNN accelerator
Jingweijia Tan, Qixiang Wang, Kaige Yan, Xiaohui Wei 0002, Xin Fu 0001
Future Gener. Comput. Syst.5
2023 Efficient Federated Learning for AIoT Applications Using Knowledge Distillation
abstract
As a promising distributed machine learning paradigm, federated learning (FL) trains a central model with decentralized data without compromising user privacy, which makes it widely used by Artificial Intelligence Internet of Things (AIoT) applications. However, the traditional FL suffers from model inaccuracy, since it trains local models only using hard labels of data while useful information of incorrect predictions with small probabilities is ignored. Although various solutions try to tackle the bottleneck of the traditional FL, most of them introduce significant communication overhead, making the deployment of large-scale AIoT devices a great challenge. To address the above problem, this article presents a novel distillation-based FL (DFL) method that enables efficient and accurate FL for AIoT applications. By using knowledge distillation (KD), in each round of FL training, our approach uploads both the soft targets and local model gradients to the cloud server for aggregation, where the aggregation results are then dispatched to AIoT devices for the next round of local training. During the DFL local training, in addition to hard labels, the model predictions approximate soft targets, which can improve model accuracy by leveraging the knowledge of soft targets. To further improve our DFL model performance, we design a dynamic adjustment strategy of loss function weights for tuning the ratio of KD and FL, which can maximize the synergy between soft targets and hard labels. Comprehensive experimental results on well-known benchmarks show that our approach can significantly improve the model accuracy of FL without introducing significant communication overhead.
Tian Liu 0005, Jun Xia 0003, Zhiwei Ling, Xin Fu 0001, Shui Yu 0001, Mingsong Chen 0001
IEEE Internet Things J.4
2023 Accelerating Convolutional Neural Network by Exploiting Sparsity on GPUs
abstract
The convolutional neural network (CNN) is an important deep learning method, which is widely used in many fields. However, it is very time consuming to implement the CNN where convolution usually takes most of the time. There are many zero values in feature maps and filters, which leads to redundant calculations and memory accesses if dense methods are used to compute convolution. Many works recently have made use of sparsity to skip the calculations for zero values to reduce the inference time of the CNN. On the graphics processing unit platform, current works cannot fully exploit the sparsity of the feature map and achieve satisfactory performance. Therefore, we design a new parallel strategy to transform the feature map into a new storage format to avoid the redundant computation of zero values on graphics processing units. Also considering the sparsity in the feature map, we propose a fused storage format to combine the convolution operation with the following pooling operation, to further improve the performance. We carry out experiments with mainstream CNN models and achieve better performance compared with cuDNN and cuSPARSE. For VGG-19, ResNet-50, DenseNet-121, and RegNetX-16GF, 1.97×, 2.23×, 2.74×, and 1.58× speedups respectively are obtained over cuDNN. The speedups over cuSPARSE respectively are 2.10×, 1.83×, 2.35×, and 1.35× when only using the first method.
Weizhi Xu 0001, Yintai Sun, Shengyu Fan, Hui Yu 0010, Xin Fu 0001
ACM Trans. Archit. Code Optim.5
2023 Accelerating Reinforcement Learning-Based CCSL Specification Synthesis Using Curiosity-Driven Exploration
abstract
The Clock Constraint Specification Language (CCSL) has been widely acknowledged as a promising system-level specification for the modeling and analysis of timing behaviors of real-time and embedded systems. However, along with the increasing complexity of modern systems coupled with strict time-to-market constraints, it becomes more and more difficult for requirement engineers to accurately figure out CCSL specifications from natural language-based requirement documents, since they lack both expertise in formal CCSL modeling and design automation tools to support quick and automatic generation of CCSL specifications. To solve the above problem, in this paper we introduce a novel and efficient Reinforcement Learning (RL)-based synthesis approach that can facilitate requirement engineers to quickly figure out their expected CCSL specifications. For a given incomplete CCSL specification, our approach adopts RL-based enumeration to explore all the feasible solutions to fill the holes within CCSL constraints, and leverages curiosity-driven exploration to accelerate the enumeration process. Based on the combination of our proposed curiosity-driven exploration heuristic and deductive reasoning techniques, our approach can not only prune unfruitful enumeration solutions effectively, but also optimize the enumeration process to search for the tightest solution quickly, thus the overall synthesis process can be accelerated dramatically. Comprehensive experimental results demonstrate that our approach significantly outperforms state-of-the-art methods in terms of both synthesis time and synthesis accuracy.
Ming Hu 0003, Min Zhang 0002, Frédéric Mallet, Xin Fu 0001, Mingsong Chen 0001
IEEE Trans. Computers4
2023 Enabling High-Efficient ReRAM-Based CNN Training Via Exploiting Crossbar-Level Insignificant Writing Elimination
abstract
Convolutional neural networks (CNNs) have been widely adopted in many deep learning applications. However, training a deep CNN requests intensive data transfer, which is both time and energy consuming. Using resistive random-access memory (ReRAM) to process data locally in memory is an emerging solution to eliminate the massive data movement. However, training cannot be efficiently supported with current ReRAM-based PIM accelerators because of the frequent and high-cost ReRAM writing operations from the delay, energy, and ReRAM lifetime perspectives. In this paper, we observe that activation induced and weight updating induced writing operations dominate the training energy on ReRAM-based accelerators. We then exploit and leverage a new angle in intermediate data (e.g., activations and errors) sparsity that fits the unique computation pattern in ReRAM crossbars to effectively eliminate the insignificant ReRAM writings, thus, enabling highly efficient CNN training without hurting the training accuracy. The experiment results show our proposed scheme achieves averagely$4.97\times$($19.23\times$) energy saving and$1.38\times$($30.08\times$) speedup compared to the state-of-the-art ReRAM-based accelerator (GPU). Our scheme also achieves$4.55\times$lifetime enhancement compared to the state-of-the-art ReRAM accelerator.
Qiyu Wan, Peixun Ma, Jing Wang 0055, Mingsong Chen 0001, Shuaiwen Song, Xin Fu 0001
IEEE Trans. Computers7
2023 Energy and Reliability-Aware Task Scheduling for Cost Optimization of DVFS-Enabled Cloud Workflows
abstract
Due to the increasing complexity, the execution of workflow applications on cloud typically involves a large number of virtual machines (VMs), which makes the cost as well as energy consumption a great concern. To alleviate this issue, more and more cloud service providers introduce new pricing policies considering Dynamic Voltage and Frequency Scaling (DVFS), where users are charged on the basis of allocated CPU frequencies together with various combinations of VM configurations and prices. However, the customizable CPU frequencies make resource provisioning and scheduling harder to achieve a cost-optimal solution. The things become even worse, since lowering CPU voltages of VMs will increase their chance of suffering soft errors, which results in a high rate of completion time failures of workflow applications. To address the above problem, this paper proposes a novel task scheduling method for the purpose of cost optimization based on the genetic algorithm. By introducing new genetic operators and frequency scaling scheme for DVFS-enabled cloud workflows, our approach can quickly figure out cost-optimal resource provisioning and task scheduling solutions by allocating tasks to appropriate VMs with specific operating frequencies under energy, reliability, makespan and memory constraints. Extensive experiments on various well-known scientific workflow benchmarks validate the effectiveness of the proposed method. Comparing with state-of-the-art methods, our approach can significantly reduce the overall cost and energy consumption without violating the given constraints.
E. Cao, Saira Musa, Mingsong Chen 0001, Tongquan Wei, Xian Wei, Xin Fu 0001, Meikang Qiu
IEEE Trans. Cloud Comput.6
2022 BS-pFL: Enabling Low-Cost Personalized Federated Learning by Exploring Weight Gradient Sparsity
abstract
Recent advancements in Convolution Neural Networks (CNNs) have achieved amazing success in numerous applications. The record-breaking performance of CNNs is usually at the prohibitive training costs, thus all training data are usually processed at the powerful centralized server side, which rises privacy concerns. Federated learning (FL) is a distributed machine learning method over mobile devices to train a global model while keeping decentralized data on devices to preserve the data privacy. However, there are two major limitations to deploy FL on mobile clients. Firstly, on the client side, the limited communication and computation resources on mobile devices cannot well support the full training iterations. Secondly, on the server side, conventional FL only aggregate a common output for all the clients without personalizing the model to each client, which is an important missing feature when clients have heterogeneous data distributions. In this work, we aim to enable low-cost personalized FL by focusing on the weight gradients which are the most important exchanging parameters in FL and meanwhile, dominating the computation and communication cost. We first observe that the client's calculated weight gradients have high sparsity, and the sparse pattern in weight gradients could be predicted via very simple bit-wise operations on a sequence of bits (named bit-stream) instead of conducting expensive high-precision calculations to determine them. Furthermore, a unique pattern is exhibited in each client's uploaded weight gradients according to the distribution of its local training data. Guided by this pattern, each client can get a personalized aggregated model to fit its own data. Hence, we leverage bit-streams to predict weight gradients sparsity for low-cost training on each device, and meanwhile, bit-streams are used to represent the unique sparse pattern of the weight gradient for each client which will guide the model personalization. From our experiments, our proposed framework can improve the computation efficiency by 3.5× on average (up to 4.2×) and reduce the communication cost by 23% on average (up to 41%) while still achieving the state-of-the-art personalized accuracy.
Manojna Sistla, Mingsong Chen 0001, Xin Fu 0001
IJCNN4
2022 Enabling PIM-based AES encryption for online video streaming
Amer Qouneh, Xin Fu 0001
J. Syst. Archit.4
2022 DynamAP: Architectural Support for Dynamic Graph Traversal on the Automata Processor
abstract
Dynamic graph traversals (DGTs) currently are widely used in many important application domains, especially in this big-data era that urgently demands high-performance graph processing and analysis. Unlike static graph traversals, DGTs in real-world application scenarios require not only fast traversal acceleration itself but also, more importantly, a runtime strategy that can effectively accommodate the ever-evolving nature of the graph structure updates followed by a diverse range of graph traversal algorithms . Because of these special features, state-of-the-art designs on conventional compute-centric architectures (e.g., CPU and GPU) struggle to provide sufficient acceleration for DGT processing due to the dominating irregular memory access patterns in graph traversal algorithms and inefficient platform-specific update mechanisms. In this article, we explore the algorithmic features and runtime requirements of real-world DGTs and identify their unique opportunities of acceleration on the recent Micron Automata Processor (AP), an in-situ memory-centric pattern-matching architecture. These features include the natural mapping between traversal algorithms’ path exploration pattern to classic non-deterministic finite automata processing, AP’s architectural and compilation support for DGTs’ evolving traversal operations, and its inherent hardware fitness. However, despite these benefits, enabling highly efficient DGT execution on AP is non-trivial and faces several major challenges. To tackle them, we propose DynamAP , the first AP framework design that enables fast processing for general DGTs. DynamAP is oblivious to periodical traversal algorithm changes and can address the significant overhead caused by frequent graph updates and AP recompilation through our novel hybrid macro designs and associated efficient updating strategies. We evaluate DynamAP against the current DGT designs on a CPU, GPU, and AP with a range of widely adopted DGT algorithms and real-world graphs. For a single update request , our DynamAP achieves an average speedup of 21.3x (up to 39.2x ) over the state-of-the-art implementation on host-AP architecture; an average speedup of 9.2x (up to 14.7x ) and 1.7x (up to 2.8x ) over two highly optimized DGT design frameworks on a 64-GB Intel(R) Xeon CPU and a 32-GB NVIDIA Tesla V100 GPU. DynamAP also maintains high performance and resource utilization for high graph update ratios, and can significantly benefit natural graphs that present a high average vertex degree.
Xingyao Zhang 0002, Donglin Zhuang, Xin Fu 0001, Shuaiwen Song
ACM Trans. Archit. Code Optim.4
2022 PervasiveFL: Pervasive Federated Learning for Heterogeneous IoT Systems
abstract
Federated learning (FL) has been recognized as a promising collaborative on-device machine learning method in the design of Internet of Things (IoT) systems. However, most existing FL methods fail to deal with IoT applications that contain a variety of IoT devices equipped with different types of neural network (NN) models. This is because traditional FL methods assume that local models on devices should have the same architecture as the global model on cloud. To address this problem, we propose a novel framework named PervasiveFL that enables efficient and effective FL among heterogeneous IoT devices. Without modifying original local models, PervasiveFL installs one lightweight NN model named modellet on each device. By using the deep mutual learning (DML) and our entropy-based decision gating (EDG) method, modellets and local models can selectively learn from each other through soft labels using locally captured data. Meanwhile, since modellets are of the same architecture, the learned knowledge by modellets can be shared among devices in a traditional FL manner. In this way, PervasiveFL can be pervasively applied to any heterogeneous IoT system. Comprehensive experimental results on four well-known datasets show that PervasiveFL can not only pervasively enable FL among heterogeneous devices within a large-scale IoT system, but also significantly enhance the inference accuracy of heterogeneous IoT devices with low communication overhead.
Jun Xia 0003, Tian Liu 0005, Zhiwei Ling, Ting Wang 0001, Xin Fu 0001, Mingsong Chen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 η-LSTM: Co-Designing Highly-Efficient Large LSTM Training via Exploiting Memory-Saving and Architectural Design Opportunities
abstract
Recently, the recurrent neural network, or its most popular type—the Long Short Term Memory (LSTM) network— has achieved great success in a broad spectrum of real-world application domains, such as autonomous driving, natural language processing, sentiment analysis, and epidemiology. Due to the complex features of the real-world tasks, current LSTM models become increasingly bigger and more complicated for enhancing the learning ability and prediction accuracy. However, through our in-depth characterization on the state-of-the-art general-purpose deep-learning accelerators, we observe that the LSTM training execution grows inefficient in terms of storage, performance, and energy consumption, under an increasing model size. With further algorithmic and architectural analysis, we identify the root cause for large LSTM training inefficiency: massive intermediate variables. To enable a highly-efficient LSTM training solution for the ever-growing model size, we exploit some unique memory-saving and performance improvement opportunities from the LSTM training procedure, and leverage them to propose the first cross-stack training solution, η-LSTM, for large LSTM models. η-LSTM comprises both software-level and hardware-level innovations that effectively lower the memory footprint upper-bound and excessive data movements during large LSTM training, while also drastically improving training performance and energy efficiency. Experimental results on six real-world large LSTM training benchmarks demonstrate that η-LSTM reduces the required memory footprint by an average of 57.5% (up to 75.8%) and brings down the data movements for weight matrices, activation data, and intermediate variables by 40.9%, 32.9%, and 80.0%, respectively. Furthermore, it outperforms the state-of-the-art GPU implementation for LSTM training by an average of 3.99× (up to 5.73×) on performance and 2.75× (up to 4.25) on energy. We hope this work can shed some light on how to design high logic utilization for future NPUs.
Xingyao Zhang 0002, Haojun Xia, Donglin Zhuang, Xin Fu 0001, Michael B. Taylor, Shuaiwen Song
ISCA5
2021 Shift-BNN: Highly-Efficient Probabilistic Bayesian Neural Network Training via Memory-Friendly Pattern Retrieving
abstract
Bayesian Neural Networks (BNNs) that possess a property of uncertainty estimation have been increasingly adopted in a wide range of safety-critical AI applications which demand reliable and robust decision making, e.g., self-driving, rescue robots, medical image diagnosis. The training procedure of a probabilistic BNN model involves training an ensemble of sampled DNN models, which induces orders of magnitude larger volume of data movement than training a single DNN model. In this paper, we reveal that the root cause for BNN training inefficiency originates from the massive off-chip data transfer by Gaussian Random Variables (GRVs). To tackle this challenge, we propose a novel design that eliminates all the off-chip data transfer by GRVs through the reversed shifting of Linear Feedback Shift Registers (LFSRs) without incurring any training accuracy loss. To efficiently support our LFSR reversion strategy at the hardware level, we explore the design space of the current DNN accelerators and identify the optimal computation mapping scheme to best accommodate our strategy. By leveraging this finding, we design and prototype the first highly efficient BNN training accelerator, named Shift-BNN, that is low-cost and scalable. Extensive evaluation on five representative BNN models demonstrates that Shift-BNN achieves an average of 4.9 × (up to 10.8 ×) boost in energy efficiency and 1.6 × (up to 2.8 ×) speedup over the baseline DNN training accelerator.
Qiyu Wan, Haojun Xia, Xingyao Zhang 0002, Shuaiwen Song, Xin Fu 0001
MICRO6
2021 Enabling Highly Efficient Capsule Networks Processing Through Software-Hardware Co-Design
abstract
As the demand for the image processing increases, the image features become increasingly complicated. Although the Convolutional Neural Network (CNN) have been widely adopted for the imaging processing tasks, it has been found easily misled due to the massive usage of pooling operations. A novel neural network structure called Capsule Networks (CapsNet) is proposed to address the CNN challenge and essentially enhance the learning ability for the image segmentation and object detection. Since the CapsNet contains the high volume of the matrix execution, it has been generally accelerated on modern GPU platforms with the highly optimized deep-learning library. However, the routing procedure of CapsNet introduces the special program and execution features,including massive unshareable intermediate variables and intensive synchronizations, causing inefficient CapsNet execution on modern GPU. To address these challenges, we propose the software-hardware co-designed optimizations, SH-CapsNet, which includes the software-level optimizations namedS-CapsNetand a hybrid computing architecture design namedPIM-CapsNet. In software-level, S-CapsNet reduces the computation and memory accesses by exploiting the computational redundancy and data similarity of the routing procedure. In hardware-level, the PIM-CapsNet leverages the processing-in-memory capability of today's 3D stacked memory to conduct the off-chip in-memory acceleration solution for the routing procedure, while pipelining with the GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet. Evaluation results demonstrate that either our software or hardware optimizations can significantly improve the CapsNet execution efficiency. Together, our co-design can achieve greatly improvement on both performance (3.41 x) and energy savings (68.72 percent) for CapsNet inference, with negligible accuracy loss.
Xingyao Zhang 0002, Xin Fu 0001, Donglin Zhuang, Chenhao Xie 0001, Shuaiwen Song
IEEE Trans. Computers2
2021 A Collaborative and Sustainable Edge-Cloud Architecture for Object Tracking with Convolutional Siamese Networks
abstract
Convolutional Neural Networks (CNNs) are becoming popular in Internet-of-Things (IoT) based object tracking areas, e.g., autonomous driving, commercial surveillance, and intelligent traffic management. However, due to limited processing power of embedded devices and network bandwidth, how to simultaneously guarantee fast object tracking with high accuracy and low energy consumption is still a major challenge, which makes IoT-based vision applications unreliable and unsustainable. To address this problem, this article proposes a collaborative edge-cloud architecture that resorts to cloud for object tracking performance enhancement. By properly offloading computations to cloud and periodically checking tracking status of edge devices through convolutional Siamese networks, our novel edge-cloud architecture enables interactive collaborations between edge devices and cloud servers in order to quickly and accurately rectify tracking errors. Comprehensive experimental results on well-known video object tracking benchmarks show that our architecture can not only significantly improve the performance of object tracking, but also can save the energy consumption of edge devices.
Haifeng Gu, Zishuai Ge, E. Cao, Mingsong Chen 0001, Tongquan Wei, Xin Fu 0001, Shiyan Hu 0001
IEEE Trans. Sustain. Comput.6
2020 Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design
abstract
In recent years, the CNNs have achieved great successes in the image processing tasks, e.g., image recognition and object detection. Unfortunately, traditional CNN's classification is found to be easily misled by increasingly complex image features due to the usage of pooling operations, hence unable to preserve accurate position and pose information of the objects. To address this challenge, a novel neural network structure called Capsule Network has been proposed, which introduces equivariance through capsules to significantly enhance the learning ability for image segmentation and object detection. Due to its requirement of performing a high volume of matrix operations, CapsNets have been generally accelerated on modern GPU platforms that provide highly optimized software library for common deep learning tasks. However, based on our performance characterization on modern GPUs, CapsNets exhibit low efficiency due to the special program and execution features of their routing procedure, including massive unshareable intermediate variables and intensive synchronizations, which are very difficult to optimize at software level. To address these challenges, we propose a hybrid computing architecture design named PIM-CapsNet. It preserves GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet, while pipelining with an off-chip in-memory acceleration solution that effectively tackles routing procedure's inefficiency by leveraging the processing-in-memory capability of today's 3D stacked memory. Using routing procedure's inherent parallellization feature, our design enables hierarchical improvements on CapsNet inference efficiency through minimizing data movement and maximizing parallel processing in memory. Evaluation results demonstrate that our proposed design can achieve substantial improvement on both performance and energy savings for CapsNet inference, with almost zero accuracy loss. The results also suggest good performance scalability in optimizing the routing procedure with increasing network size.
Xingyao Zhang 0002, Shuaiwen Song, Chenhao Xie 0001, Jing Wang 0055, Weigong Zhang, Xin Fu 0001
HPCA6
2020 Fast-BCNN: Massive Neuron Skipping in Bayesian Convolutional Neural Networks
abstract
Bayesian Convolutional Neural Networks (BCNNs) have emerged as a robust form of Convolutional Neural Networks (CNNs) with the capability of uncertainty estimation. A BCNN model is implemented by adding a dropout layer after each convolutional layer in the original CNN. By executing the stochastic inferences many times, BCNNs are able to provide an output distribution that reflects the uncertainty of the final prediction. Repeated inferences in this process lead to much longer execution time, which makes it challenging to apply Bayesian technique to CNNs in real-world applications. In this study, we propose Fast-BCNN, an FPGA-based hardware accelerator design that intelligently skips the redundant computations for two types of neurons during repeated BCNN inferences. Firstly, within a sample inference, we aim to skip the dropped neurons that predetermined by dropout masks. Secondly, by leveraging the information from the first inference and dropout masks, we predict the zero neurons and skip all their corresponding computations during the following sample inferences. Particularly, an optimization algorithm is employed to guarantee the accuracy of zero neuron prediction while achieving the maximal computation reduction. To support our neuron skipping strategy at hardware level, we explore an efficient parallelism for CNN convolution to gracefully skip the corresponding computations for both types of neurons, we then propose a novel PE architecture that accommodates the parallel operation of convolution and prediction with negligible overhead. Experimental results demonstrate that our Fast-BCNN achieves 2.1~8.2× speedup and 44%~84% energy reduction over the baseline CNN accelerator.
Qiyu Wan, Xin Fu 0001
MICRO2
2020 Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage Scaling
abstract
With the platforms of running deep neural networks (DNNs) move from large-scale data centers to handheld devices, power emerge as one of the most significant obstacles. Voltage scaling is a promising technique that enables power saving. Nevertheless, it raises reliability and performance concerns that may undesirably deteriorate NNs accuracy and performance. Consequently, an energy-efficient and reliable scheme is required for NNs to balance the above three aspects with satisfied user experience. To this end, we propose a neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation characteristics in NNs at both inter- and intra-network layers to precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Furthermore, we combine a voltage clustering method and the multi-objective optimization to identify the optimal voltage islands and apply the same voltage to neurons with similar fault tolerance capability. We perform three case studies to demonstrate the efficacy of the proposed techniques.
Jing Wang 0055, Xin Fu 0001, Lan Gao 0004, Weigong Zhang
IEEE Trans. Computers2
2020 Statistical Model Checking-Based Evaluation and Optimization for Cloud Workflow Resource Allocation
abstract
Due to the existence of resource variations, it is very challenging for Cloud workflow resource allocation strategies to guarantee a reliable Quality of Service (QoS). Although dozens of resource allocation heuristics have been developed to improve the QoS of Cloud workflow, it is hard to predict their performance under variations because of the lack of accurate modeling and evaluation methods. So far, there is no comprehensive approach that can quantitatively reason the capability of resource allocation strategies or enable the tuning of parameters to optimize resource allocation solutions under variations. To address the above problems, this paper proposes a novel framework that can evaluate and optimize resource allocation strategies effectively and quantitatively. By using the statistical model checker UPPAAL-SMC and supervised learning approaches, our framework can: i) conduct complex QoS queries on resource allocation instances considering resource variations; ii) make quantitative and qualitative comparisons among resource allocation strategies; iii) enable the tuning of parameters to improve the overall QoS; and iv) support the quick optimization of overall workflow QoS under customer requirements and resource variations. The experimental results demonstrate that our automated framework can support both the Service Level Agreement (SLA) negotiation and workflow resource allocation optimization efficiently.
Mingsong Chen 0001, Saijie Huang, Xin Fu 0001, Xiao Liu 0004, Jifeng He 0001
IEEE Trans. Cloud Comput.3
2020 Toward Customized Hybrid Fuel-Cell and Battery-powered Mobile Device for Individual Users
abstract
Rapidly evolving technologies and applications of mobile devices inevitably increase the power demands on the battery. However, the development of batteries can hardly keep pace with the fast-growing demands, leading to short battery life, which becomes the top complaints from customers. In this article, we investigate a novel energy supply technology, fuel cell (FC), and leverage its advantages of providing long-term energy storage to build a hybrid FC-battery power system. Therefore, mobile device operation time is dramatically extended, and users are no longer bothered by battery recharging. We examine real-world smartphone usage data and find that a naive hybrid power system cannot meet many users’ highly diversified power demands. We thus propose an OS-level power management policy that reduces the device power consumption for each power peak to solve this mismatch. This technique trades the quality-of-service (QoS) for a larger FC ratio in the system and thus much longer device operation time. We further observe that the user’s personality largely determines his/her satisfaction with the QoS degradation and the operation time extension. Thus, applying a hybrid system with fixed configuration (i.e., peak throttling level coupled with corresponding FC/battery ratio) fails to satisfy every user. We then explore customized hybrid system configuration based on each individual user’s personality to deliver the optimal satisfaction for him/her. The experimental results show that our personality-aware hybrid FC-battery solution can achieve 4× longer operation time and 25% higher satisfaction score compared to the common setting for state-of-the-art mobile devices.
Kaige Yan, Jingweijia Tan, Longjun Liu, Xingyao Zhang 0002, Stanko R. Brankovic, Jinghong Chen, Xin Fu 0001
ACM Trans. Embed. Comput. Syst.7
2020 Energy-Efficient GPU L2 Cache Design Using Instruction-Level Data Locality Similarity
abstract
This article presents a novel energy-efficient cache design for massively parallel, throughput-oriented architectures like GPUs. Unlike L1 data cache on modern GPUs, L2 cache shared by all of the streaming multiprocessors is not the primary performance bottleneck, but it does consume a large amount of chip energy. We observe that L2 cache is significantly underutilized by spending 95.6% of the time storing useless data. If such “dead time” on L2 is identified and reduced, L2’s energy efficiency can be drastically improved. Fortunately, we discover that the SIMT programming model of GPUs provides a unique feature among threads: instruction-level data locality similarity, which can be used to accurately predict the data re-reference counts at L2 cache block level. We propose a simple design that leverages this Lo cality S imilarity to build an energy-efficient GPU L2 Cache , named LoSCache . Specifically, LoSCache uses the data locality information from a small group of cooperative thread arrays to dynamically predict the L2-level data re-reference counts of the remaining cooperative thread arrays. After that, specific L2 cache lines can be powered off if they are predicted to be “dead” after certain accesses. Experimental results on a wide range of applications demonstrate that our proposed design can significantly reduce the L2 cache energy by an average of 64% with only 0.5% performance loss. In addition, LoSCache is cost effective, independent of the scheduling policies, and compatible with the state-of-the-art L1 cache designs for additional energy savings.
Jingweijia Tan, Kaige Yan, Shuaiwen Song, Xin Fu 0001
ACM Trans. Design Autom. Electr. Syst.4
2019 LoSCache: Leveraging Locality Similarity to Build Energy-Efficient GPU L2 Cache
abstract
This paper presents a novel energy-efficient cache design for massively parallel, throughput-oriented architectures like GPUs. Unlike L1 data cache on modern GPUs, L2 cache shared by all the streaming multiprocessors is not the primary performance bottleneck but it does consume a large amount of chip energy. We observe that L2 cache is significantly under-utilized by spending 95.6% of the time storing useless data. If such "dead time" on L2 is identified and reduced, L2's energy efficiency can be drastically improved. Fortunately, we discover that the SIMT programming model of GPUs provides a unique feature among threads: instruction-level data locality similarity, which can be used to accurately predict the data re-reference counts at L2 cache block level. We propose a simple design that leverages this Locality Similarity to build an energy-efficient GPU L2 Cache, named LoSCache. Specifically, LoSCache uses the data locality information from a small group of CTAs to dynamically predict the L2-level data re-reference counts of the remaining CTAs. After that, specific L2 cache lines can be powered off if they are predicted to be "dead" after certain accesses. Experimental results on a wide range of applications demonstrate that our proposed design can significantly reduce the L2 cache energy by an average of 64% with only 0.5% performance loss.
Jingweijia Tan, Kaige Yan, Shuaiwen Song, Xin Fu 0001
DATE4
2019 PIM-VR: Erasing Motion Anomalies In Highly-Interactive Virtual Reality World with Customized Memory Cube
abstract
With the revolutionary innovations emerging in the computer graphics domain, virtual reality (VR) has become increasingly popular and mainstream for entertainment, medical simulation and education. In the highly-interactive VR world, the motion-to-photon delay (MPD) which represents the delay from users' head motion to the responded image displayed on their head devices, is the most critical factor for a successful VR experience. Long MPD may cause user to experience significant motion anomalies: judder, lagging and sickness. In order to alleviate this negative effect, asynchronous time warp (ATW) has been proposed by VR vendors to map the rendered stereoscopic frame in the correct position using the latest head-motion information. However, after a careful investigation on the efficiency of the current GPU-accelerated ATW through executing real VR applications on modern VR hardware, we observe that the current ATW design on commercial hardware cannot achieve the ideal MPD and often cause ATW to miss the refresh deadline, resulting in motion anomalies and dropped frame rate. This is caused by two major challenges: inefficient VR execution model and intensive off-chip memory accesses. To tackle these, we propose a preemption-free Processing-In-Memory based ATW design which asynchronously executes ATW within a 3D-stacked memory, without interrupting the rendering tasks on the host GPU. We also identify a redundancy reduction mechanism to further simplify and accelerate the ATW operation. A comprehensive evaluation on our proposed design demonstrates that our PIM-based ATW can achieve the ideal MPD and provide superior user experience. Finally, we provide a design space exploration to showcase different design choices for the PIM-based ATW design.
Chenhao Xie 0001, Xingyao Zhang 0002, Ang Li 0006, Xin Fu 0001, Shuaiwen Song
HPCA4
2019 Reliability Aware Cost Optimization for Memory Constrained Cloud Workflows
E. Cao, Saira Musa, Jianning Zhang, Mingsong Chen 0001, Tongquan Wei, Xin Fu 0001, Meikang Qiu
ICA3PP (2)6
2019 Reliability Enhancement of Neural Networks via Neuron-Level Vulnerability Quantization
Keyao Li, Jing Wang 0055, Xin Fu 0001, Xiufeng Sui, Weigong Zhang
ICA3PP (2)3
2019 Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage Scaling
abstract
As the application scope of deep neural networks (DNNs) moves from large-scale data centers to small-scale mobile devices, power wall has become one of the most important obstacles. Voltage scaling is a typical technique enables power saving, but it causes reliability and performance challenges. Therefore, an energy-efficient and reliable scheme for NNs is required to balance above three aspects according to users' requirements for excellent user experience. In this paper, we innovatively propose neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation in NNs and precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Multi-objective optimization and clustering method are combined to find the optimal voltage islands. Finally, we conduct experiment to demonstrate the efficacy of the proposed technique.
Jing Wang 0055, Xin Fu 0001, Xingyao Zhang 0002, Lan Gao 0004, Weigong Zhang, Tao Li 0006
ICPADS4
2019 OO-VR: NUMA friendly object-oriented VR rendering framework for future NUMA-based multi-GPU systems
abstract
With the strong computation capability, NUMA-based multi-GPU system is a promising candidate to provide sustainable and scalable performance for Virtual Reality (VR) applications and deliver the excellent user experience. However, the entire multi-GPU system is viewed as a single GPU under the single programming model which greatly ignores the data locality among VR rendering tasks during the workload distribution, leading to tremendous remote memory accesses among GPU models (GPMs). The limited inter-GPM link bandwidth (e.g., 64GB/s for NVlink) becomes the major obstacle when executing VR applications in the multi-GPU system. By conducting comprehensive characterizations on different kinds of parallel rendering frameworks, we observe that distributing the rendering object along with its required data per GPM can reduce the inter-GPM memory accesses. However, this object-level rendering still faces two major challenges in NUMA-based multi-GPU system: (1) the large data locality between the left and right views of the same object and the data sharing among different objects and (2) the unbalanced workloads induced by the software-level distribution and composition mechanisms.
Chenhao Xie 0001, Xin Fu 0001, Mingsong Chen 0001, Shuaiwen Song
ISCA2
2019 Improving energy efficiency of mobile devices by characterizing and exploring user behaviors
Kaige Yan, Jingweijia Tan, Xin Fu 0001
J. Syst. Archit.3
2019 Bridging mobile device configuration to the user experience under budget constraint
Kaige Yan, Jingweijia Tan, Xin Fu 0001
Pervasive Mob. Comput.3
2018 Perception-Oriented 3D Rendering Approximation for Modern Graphics Processors
abstract
Anisotropic filtering enabled by modern rasterization-based GPUs provides users with extremely authentic visualization experience, but significantly limits the performance and energy efficiency of 3D rendering process due to its large texture data requirement. To improve 3D rendering efficiency, we build a bridge between anisotropic filtering process and human visual system by analyzing users’ perception on image quality. We discover that anisotropic filtering does not impact user perceived image quality on every pixel. This motives us to approximate the anisotropic filtering process for non-perceivable pixels in order to improve the overall 3D rendering performance without damaging user experience. To achieve this goal, we propose a perceptionoriented runtime approximation model for 3D rendering by leveraging the inner-relationship between anisotropic and isotropic filtering. We also provide a low-cost texture unit design for enabling this approximation. Extensive evaluation on modern 3D games demonstrates that, under a conservative tuning point, our design achieves a significant average speedup of 17% for the overall 3D rendering along with 11% total GPU energy reduction, without visible image quality loss from users’ perception. It also reduces the texture filtering latency by an average of 29%. Additionally, it creates a unique perception-based tuning space for performance-quality tradeoffs on graphics processors.
Chenhao Xie 0001, Xin Fu 0001, Shuaiwen Song
HPCA2
2018 Towards Memory Friendly Long-Short Term Memory Networks (LSTMs) on Mobile GPUs
abstract
Intelligent Personal Assistants (IPAs) with the capability of natural language processing (NLP) are increasingly popular in today's mobile devices. Recurrent neural networks (RNNs), especially one of their forms – Long-Short Term Memory networks (LSTMs), are becoming the core machine learning technique applied in the NLP-based IPAs. With the continuously improved performance of mobile GPUs, local processing has become a promising solution to the large data transmission and privacy issues induced by the cloud-centric computations of IPAs. However, LSTMs exhibit quite inefficient memory access pattern when executed on mobile GPUs due to the redundant data movements and limited off-chip bandwidth. In this study, we aim to explore the memory friendly LSTM on mobile GPUs by hierarchically reducing the off-chip memory accesses. To address the redundant data movements, we propose inter-cell level optimizations that intelligently parallelize the originally sequentially executed LSTM cells (basic units in RNNs, corresponding to neurons in CNNs) to improve the data locality across cells with negligible accuracy loss. To relax the pressure on limited off-chip memory bandwidth, we propose intra-cell level optimizations that dynamically skip the loads and computations of rows in the weight matrices with trivial contribution to the outputs. We also introduce a light-weighted module to the GPUs architecture for the runtime row skipping in weight matrices. Moreover, our techniques are equipped with thresholds which provide a unique tunning space for performance-accuracy trade-offs directly guided by the user preferences. The experimental results show our optimizations achieves substantial improvements on both performance and power with user-imperceptible accuracy loss. And our optimizations exhibit the strong scalability with the increasing input data set. Our user study also shows that our designed system delivers the excellent user experience.
Xingyao Zhang 0002, Chenhao Xie 0001, Jing Wang 0055, Xin Fu 0001
MICRO5
2017 Processing-in-Memory Enabled Graphics Processors for 3D Rendering
abstract
The performance of 3D rendering of Graphics Processing Unit that converts 3D vector stream into 2D frame with 3D image effects significantly impacts users gaming experience on modern computer systems. Due to its high texture throughput requirement, main memory bandwidth becomes a critical obstacle for improving the overall rendering performance. 3D-stacked memory systems such as Hybrid Memory Cube provide opportunities to significantly overcome the memory wall by directly connecting logic controllers to DRAM dies. Although recent works have shown promising improvement in performance by utilizing HMC to accelerate special-purpose applications, a critical challenge of how to effectively leverage its high internal bandwidth and computing capability in GPU for 3D rendering remains unresolved. Based on the observation that texel fetches greatly impact off-chip memory traffic, we propose two architectural designs to enable Processing-In-Memory based GPU for efficient 3D rendering. Additionally, we employ camera angles of pixels to control the performance-quality tradeoff of 3D rendering. Extensive evaluation across several real-world games demonstrates that our design can significantly improve the performance of texture filtering and 3D rendering by an average of 3.97X (up to 6.4X) and 43% (up to 65%) respectively, over the baseline GPU. Meanwhile, our design provides considerable memory traffic and energy reduction without sacrificing rendering quality.
Chenhao Xie 0001, Shuaiwen Song, Jing Wang 0055, Weigong Zhang, Xin Fu 0001
HPCA5
2017 Exploring Energy-Efficient Cache Design in Emerging Mobile Platforms
abstract
Mobile devices are quickly becoming the most widely used processors in consumer devices. Since their major power supply is battery, energy-efficient computing is highly desired. In this article, we focus on energy-efficient cache design in emerging mobile platforms. We observe that more than 40% of L2 cache accesses are OS kernel accesses in interactive smartphone applications. Such frequent kernel accesses cause serious interferences between the user and kernel blocks in the L2 cache, leading to unnecessary block replacements and high L2 cache miss rate. We first propose to statically partition the L2 cache into two separate segments, which can be accessed only by the user code and kernel code, respectively. Meanwhile, the overall size of the two segments is shrunk, which reduces the energy consumption while still maintaining the similar cache miss rate. We then find completely different access behaviors between the two separated kernel and user segments and explore the multi-retention STT-RAM-based user and kernel segments to obtain higher energy savings in this static partition-based cache design. Finally, we propose to dynamically partition the L2 cache into the user and kernel segments to minimize overall cache size. We also integrate the short-retention STT-RAM into this dynamic partition-based cache design for maximal energy savings. The experimental results show that our static technique reduces cache energy consumption by 75% with 2% performance loss, and our dynamic technique further shows strong capability to reduce cache energy consumption by 85% with only 3% performance loss.
Kaige Yan, Lu Peng 0001, Mingsong Chen 0001, Xin Fu 0001
ACM Trans. Design Autom. Electr. Syst.4
2017 Efficient Resource Constrained Scheduling Using Parallel Two-Phase Branch-and-Bound Heuristics
abstract
Branch-and-bound (B&B) approaches are widely investigated in resource constrained scheduling (RCS). However, due to the lack of approaches that can generate a tight schedule at the beginning of the search, B&B approaches usually start with a large initial search space, which makes the following search of an optimal schedule time-consuming. To address this problem, this paper proposes a parallel two-phase B&B approach that can drastically reduce the overall RCS time. This paper makes three major contributions: i) it proposes three partial-search heuristics that can quickly find a tight schedule to compact the initial search space; ii) it presents a two-phase search framework that supports the efficient parallel search of an optimal schedule; iii) it investigates various bound sharing and speculation techniques among collaborative tasks to further improve the parallel search performance at different search phases. The experimental results based on well-established benchmarks demonstrate the efficacy of our proposed approach.
Mingsong Chen 0001, Yongxiang Bao, Xin Fu 0001, Geguang Pu, Tongquan Wei
IEEE Trans. Parallel Distributed Syst.3
2017 On the Implication of NTC versus Dark Silicon on Emerging Scale-Out Workloads: The Multi-Core Architecture Perspective
abstract
The end of Dennard's scaling poses computer systems, especially the datacenters, in front of both power and utilization walls. One possible solution to combat the power and utilization walls is dark silicon where transistors are under-utilized in the chip, but this will result in a diminishing performance. Another solution is Near-Threshold Voltage Computing (NTC) which operates transistors in the near-threshold region and provides much more flexible tradeoffs between power and performance. However, prior efforts largely focus on a specific design option based on the legacy desktop applications, therefore, lacking comprehensive analysis of emerging scale-out applications with multiple design options when dark silicon and/or NTC are/is applied. In this paper, we characterize different perspectives including performance, energy efficiency and reliability in the context of NTC/dark silicon cloud processors running emerging scale-out workloads on various architecture designs. We find NTC is generally an effective way to alleviate the power challenge over scale-out applications compared with dark silicon, it can improve performance by 1.6X, energy efficiency by 50 percent and the reliability problem can be relieved by ECC. Meanwhile, we also observe tiled-OoO architecture improves the performance by 20~370 percent and energy efficiency by 40~600 percent over alternative architecture designs, making it a preferable design paradigm for scale-out workloads. We believe that our observations will provide insights for the design of cloud processors under dark silicon and/or NTC.
Jing Wang 0055, Xin Fu 0001, Weigong Zhang, Keni Qiu, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.2
2017 Interspike-Interval-Based Analog Spike-Time-Dependent Encoder for Neuromorphic Processors
abstract
Von Neumann bottleneck, which refers to the limited throughput between the CPU and memory, has already become a major factor hindering the technical advances of computing systems. In recent years, neuromorphic systems have started to gain the increasing attentions as compact and energy-efficient computing platforms. As one of the most crucial components in the neuromorphic computing systems, neural encoder transforms the stimulus (input signals) into spike trains. In this paper, we adapt the temporal encoding scheme of interspike intervals (ISIs) and present an analog temporal neural encoder with its verification and recovery schemes. The proposed neural encoder allows efficient mapping of signal amplitude information into a spike-time sequence that represents the input data and offers perfect recovery for band-limited stimuli. With the novel iterative structure, the number of spikes increases exponentially with the number of neurons. From the measurements obtained from the fabricated neural encoder chip, our temporal encoder with ISI encoding is proved to be robust and error tolerant.
Chenyuan Zhao, Yang Yi 0002, Xin Fu 0001, Lingjia Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Dolphins First: Dolphin-Aware Communications in Multi-Hop Underwater Cognitive Acoustic Networks
abstract
Acoustic communication is the most versatile and widely used technology for underwater wireless networks. However, the frequencies used by current acoustic modems are heavily overlapped with the cetacean communication frequencies, where the man-made noise of underwater acoustic communications may have harmful or even fatal impact on those lovely marine mammals, e.g., dolphins. To pursue the environmental friendly design for sustainable underwater monitoring and exploration, specifically, to avoid the man-made interference to dolphins, in this paper, we propose a cognitive acoustic transmission scheme, called dolphin-aware data transmission (DAD-Tx), in multi-hop underwater acoustic networks. Different from the collaborative sensing approach and the simplified modeling of dolphins' activities in existing literature, we employ a probabilistic method to capture the stochastic characteristics of dolphins' communications, and mathematically describe the dolphin-aware constraint. Under dolphin-awareness and wireless acoustic transmission constraints, we further formulate the DAD-Tx optimization problem aiming to maximize the end-to-end throughput. Since the formulated problem contains probabilistic constraint and is NP-hard, we leverage Bernstein approximation and develop a three-phase solution procedure with heuristic algorithms for feasible solutions. Simulation results show the effectiveness of the proposed scheme in terms of both network performance and dolphin awareness.
Xuanheng Li, Yi Sun 0009, Yuanxiong Guo, Xin Fu 0001, Miao Pan
IEEE Trans. Wirel. Commun.4
2016 Combating the Reliability Challenge of GPU Register File at Low Supply Voltage
abstract
Supply voltage reduction is an effective approach to significantly reduce GPU energy consumption. As the largest on-chip storage structure, the GPU register file becomes the reliability hotspot that prevents further supply voltage reduction below the safe limit ($V_{min}$) due to process variation effects. This work addresses the reliability challenge of the GPU register file at low supply voltages, which is an essential first step for aggressive supply voltage reduction of the entire GPU chip. To better understand the reliability issues posed by undervolting and its energy-saving potential, we first rigorously model and analyze the process variation impact on the GPU register file at different voltages. By further analyzing the GPU architecture, we make a key observation that the time GPU registers contain useless data (i.e., dead time) is long, providing a unique opportunity to enhance register reliability. We then propose GR-Guard, an architectural solution that leverages long register dead time to enable reliable operations from unreliable register file at low voltages. GR-Guard is both effective and low-cost, and does not affect normal (i.e., non-faulty) register accesses. Experimental results show that for a 28nm baseline GPU under aggressive voltage reduction, GR-Guard can maintain the register file reliability with less than 2\% overall performance degradation, while achieving an average of 31% energy reduction across various applications.
Jingweijia Tan, Shuaiwen Song, Kaige Yan, Xin Fu 0001, Andrés Márquez 0001, Darren J. Kerbyson
PACT4
2016 Exploring Variation-Aware Fault-Tolerant Cache under Near-Threshold Computing
abstract
Near threshold voltage computing enables transistor voltage scaling to continue with Moore's Law projection and dramatically improves power and energy efficiency. However, reducing the supply voltage to near-threshold level significantly increases the susceptibility of on-chip caches to process variations, leading to the high error rate. Most existing fault-tolerant schemes significantly sacrifice cache capacity and performance. In this paper, we propose a novel fault-tolerant cache architecture at near-threshold computing, which is suitable for high error rate memories. We first propose a variation-aware skewed-associative cache, and then redirect the faulty blocks to the error-free blocks based on it to explore the fault-tolerance cache design. Unlike previous cache reconfiguration schemes for the fault tolerance, our cache design does not need to sacrifice or disable any fault-free blocks to form a completely functional set. We use all error-free blocks and have the least cache capacity waste. More importantly, since the aging impact could also cause cell failures, our skewed cache takes the aggregated process variation and aging impact into the consideration. Last but not least, our skewed cache design avoids the complex remapping from faulty blocks to the error-free blocks and minimizes the hardware overheads. Our evaluation results show that our variation-aware fault-tolerant cache design exhibits strong capability to tolerate the high error rate, and more excitingly, its effectiveness on reducing the cache miss rate and improving the performance is even more obvious as the supply voltage scales down to the near-threshold region.
Jing Wang 0055, Yanjun Liu 0005, Weigong Zhang, Kezhong Lu, Keni Qiu, Xin Fu 0001, Tao Li 0006
ICPP6
2016 Redefining QoS and customizing the power management policy to satisfy individual mobile users
abstract
Delivering an excellent use experience to the customers is the top challenge faced by today's mobile device designers and producers. There have been multiple studies on achieving the good trade-offs between QoS and energy to enhance the user experience, however, they generally lack a comprehensive and accurate understanding of QoS, and ignore the fact that each individual user has his/her own preference between QoS and energy. In this study, we overcome these two drawbacks and propose a customized power management policy that dynamically configures the mobile platform to achieve the user-specific optimal QoS and energy trade-offs and hence, satisfy each individual mobile user. We first introduce a novel and comprehensive definition of QoS, and propose the accurate QoS measurement and management methodologies. We then observe that user's personality greatly determines his/her preferences between QoS and energy, and propose an online personality-guided user satisfaction prediction model based on the QoS and energy, guided by the user personality inferred from his/her device usage history. Our validation proves our model can achieve very high prediction accuracy. Finally, we propose our customized power management policy based on the prediction model for individual users. The experiment results show that our technique can improve the user experience by around 36% compared with the state-of-the-art power management policies.
Kaige Yan, Xingyao Zhang 0002, Jingweijia Tan, Xin Fu 0001
MICRO4
2016 Efficient Resource Constrained Scheduling Using Parallel Structure-Aware Pruning Techniques
abstract
Branch-and-bound approaches are promising in pruning fruitless search space during the resource constrained scheduling. However, such approaches only compare the estimated upper and lower bounds of an incomplete schedule to the length of the best feasible schedule at that iteration, which does not fully exploit the potential of the pruning during the search. Aiming to improve the performance of resource constrained scheduling, this paper proposes a parallel structure-aware pruning approach that can traverse the search space significantly faster than state-of-the-art branch-and-bound techniques. This paper makes three major contributions: i) it proposes an efficient pruning technique using the structural scheduling information of the obtained best feasible schedules; ii) it investigates how to perform parallel search to enable efficient multi-directional search and generation of effective fences by tuning the operation enumeration order; and iii) it presents a framework that supports the sharing of minimum upper-bound and fence information among different search tasks to enable efficient parallel structure-aware pruning. The experimental results demonstrate that our parallel pruning approach can drastically reduce the overall resource constrained scheduling time under a wide variety of resource constraints.
Mingsong Chen 0001, Xinqian Zhang, Geguang Pu, Xin Fu 0001, Prabhat Mishra 0001
IEEE Trans. Computers4
2016 Soft error resilience in Big Data kernels through modular analysis
Sui Chen, Greg Bronevetsky, Lu Peng 0001, Bin Li 0008, Xin Fu 0001
J. Supercomput.5
2016 Exploring Soft-Error Robust and Energy-Efficient Register File in GPGPUs using Resistive Memory
abstract
The increasing adoption of graphics processing units (GPUs) for high-performance computing raises the reliability challenge, which is generally ignored in traditional GPUs. GPUs usually support thousands of parallel threads and require a sizable register file. Such large register file is highly susceptible to soft errors and power-hungry. Although ECC has been adopted to register file in modern GPUs, it causes considerable power overhead, which further increases the power stress. Thus, an energy-efficient soft-error protection mechanism is more desirable. Besides its extremely low leakage power consumption, resistive memory (e.g., spin-transfer torque RAM) is also immune to the radiation induced soft errors due to its magnetic field based storage. In this article, we propose to LEverage reSistive memory to enhance the Soft-error robustness and reduce the power consumption (LESS) of registers in the General-Purpose computing on GPUs (GPGPUs). Since resistive memory experiences longer write latency compared to SRAM, we explore the unique characteristics of GPGPU applications to obtain the win-win gains: achieving the near-full soft-error protection for the register file, and meanwhile substantially reducing the energy consumption with negligible performance degradation. Our experimental results show that LESS is able to mitigate the registers soft-error vulnerability by 86% and achieve 61% energy savings with negligible (e.g., 1%) performance degradation.
Jingweijia Tan, Zhi Li 0016, Mingsong Chen 0001, Xin Fu 0001
ACM Trans. Design Autom. Electr. Syst.4
2016 Mitigating the Impact of Hardware Variability for GPGPUs Register File
abstract
As technology keeps scaling down, hardware variability, such as process variations (PV) and negative bias temperature instability (NBTI), emerges as a growing challenge in the modern GPGPUs (general-purpose computing on graphics processing units). PV induces significant delay variations statically, while NBTI dynamically slows down the GPGPUs. Each computing core (i.e., streaming multiprocessor) in GPGPUs supports thousands of simultaneously active threads, and requires a large register file. Such a sizable register file is very sensitive to the hardware variability, and becomes one of the major units in determining the core frequency. In this study, we propose a set of techniques that mitigate both the PV and NBTI impacts on GPGPUs register file. In order to mitigate the susceptibility to PV, we first develop a novel mechanism that classifies registers into fast and slow categories in the highly-banked register architecture to maximize the frequency improvement. We then leverage the unique features in GPGPU applications to effectively tolerate the extra access delay to the slow registers. Moreover, we propose to dynamically balance the utilization across registers to further tolerate the NBTI degradation. Our experimental results show that our proposed techniques optimize GPGPUs performance by 22 percent on average under both PV and NBTI effects.
Jingweijia Tan, Mingsong Chen 0001, Yang Yi 0002, Xin Fu 0001
IEEE Trans. Parallel Distributed Syst.4
2015 POSTER: A Hardware Fingerprint Using GPU Core Frequency Variations
abstract
Hardware primitives provide significant promises to support cryptographic primitives and security mechanisms against various forms of compromises. In this work, we study the intrinsic hardware characteristics of modern graphics processing units (GPUs) due to random manufacturing variations, and exploits the inherent randomness to generate device-specific signatures. In particular, we present a novel GPU-based hardware fingerprint scheme to generate a unique, stable, physically unclonable, unpredictable, and random bit string from the inherent hardware features of a general purpose GPU (GPGPU). The generated fingerprint can be used to implement a physically unclonable function (PUF), and thus to create a trusted computing environment with GPUs as the trust anchor.
Fengjun Li, Xin Fu 0001, Bo Luo
CCS2
2015 Variation-aware evaluation of MPSoC task allocation and scheduling strategies using statistical model checking
Mingsong Chen 0001, Daian Yue, Xiaoke Qin, Xin Fu 0001, Prabhat Mishra 0001
DATE4
2015 Soft-error reliability and power co-optimization for GPGPUS register file using resistive memory
Jingweijia Tan, Zhi Li 0016, Xin Fu 0001
DATE3
2015 Mitigating the Susceptibility of GPGPUs Register File to Process Variations
abstract
As technology keeps scaling down at nano-scale, the increasing process variations (PV) induce significant delay variations and limit the maximum clock frequency in GPGPUs (general-purpose computing on graphics processing units). Each computing core (i.e. streaming multiprocessor) in GPGPUs supports thousands of simultaneously active threads, and requires a large register file. Such a sizeable register file is very sensitive to process variations, and becomes one of the major units in determining the core frequency. In this study, we first develop a novel mechanism that classifies registers into fast and slow categories in the highly-banked register architecture to maximize the frequency improvement. We then leverage the unique features in GPGPU applications to effectively tolerate the extra access delay to the slow registers. Our experimental results show that our proposed techniques are able to significantly optimize GPGPUs performance under process variations.
Jingweijia Tan, Xin Fu 0001
IPDPS2
2015 Characterizing, modeling, and improving the QoE of mobile devices with low battery level
abstract
Mobile users always require an excellent user experience which is the top challenge faced by today's mobile device designers and producers. Mobile devices are battery constrained, thus developing energy-saving techniques to extend the battery life is critical in terms of the user experience. Since the discrepancy between the device energy and battery energy consumption is becoming large when the battery is approaching to depleted, the battery energy-savings mechanisms (instead of the previously explored device energy-saving mechanisms) that target at the low battery level are highly desirable. Besides the battery life, mobile users demand a good system responsiveness, which is also an essential component in the user experience.
Kaige Yan, Xingyao Zhang 0002, Xin Fu 0001
MICRO3
2015 Aurora: A Cross-Layer Solution for Thermally Resilient Photonic Network-on-Chip
abstract
With silicon optical technology moving toward maturity, the use of photonic networks-on-chip (NoCs) for global chip communication is emerging as a promising solution to the communication requirements of future many core processors. It is expected that photonic NoCs will play an important role in alleviating current power, latency, and bandwidth constraints. However, photonic NoCs are sensitive to ambient temperature variations because their basic constituents, ring resonators, are themselves sensitive to those variations. Since ring resonators are basic building blocks for photonic modulators, switches, multiplexers, and demultiplexers, variations of on-chip temperature pose serious challenges to the proper operation of photonic NoCs. Proposed methods that mitigate the effects of temperature at the device level are either difficult to use in CMOS processes or not suitable for large scale implementation. In this paper, we propose Aurora, a thermally resilient photonic NoC architecture design that supports reliable and low bit error rate (BER) on-chip communications in the presence of large temperature variations. Our proposed architecture leverages cross-layer solutions at the device, architecture, and operating system (OS) layers that individually provide considerable improvements and synergistically provide even more significant improvements. To compensate for small temperature variations, our design varies the bias current through ring resonators. For larger temperature variations, we propose architecture-level techniques to reroute messages away from hot regions, and through cooler regions, to their destinations. We also propose a thermal/congestion-aware coscheduling algorithm at the OS level to further lower BER by reorganizing the thermal profile of the chip. Our simulation results show that Aurora provides a robust architectural solution to handle temperature variation effects on future photonic NoCs. For instance, average BER and message error rate are reduced by 96% and 85%, respectively, when the combined thermal optimization scheme [shortest path first+ OS] is applied. From the perspective of power efficiency, Aurora is also superior to conventional photonic NoC architectures by as much as 37%.
Zhong-Qi Li, Amer Qouneh, Madhura Joshi, Wangyuan Zhang, Xin Fu 0001, Tao Li 0006
IEEE Trans. Very Large Scale Integr. Syst.5
2013 Lighting the dark silicon by exploiting heterogeneity on future processors
abstract
As we embrace the deep submicron era, dark silicon caused by the failure of Dennard scaling impedes us from attaining commensurate performance benefit from the increased number of transistors. To alleviate the dark silicon and effectively leverage the advantage of decreased feature size, we consider a set of design paradigms by exploiting heterogeneity in the processor manufacturing. We conduct a thorough investigation on these design patterns from different evaluation perspectives including performance, energy-efficiency, and cost-efficiency. Our observations can provide insightful guidance to the design of future processors in the presence of dark silicon.
Ying Zhang 0016, Lu Peng 0001, Xin Fu 0001
DAC3
2013 Modeling and characterizing GPGPU reliability in the presence of soft errors
Jingweijia Tan, Yang Yi 0002, Fangyang Shen, Xin Fu 0001
Parallel Comput.4
2012 RISE: improving the streaming processors reliability against soft errors in gpgpus
abstract
With hundreds of cores integrated into a single chip, the general-purpose computing on graphic processing units (GPGPUs) provide high computing power to accelerate parallel applications. However, they are prone to manifest high soft-error vulnerability due to the lack of fault detection and tolerance. Especially, streaming processors become the reliability hot-spot in GPGPUs. This paper explores two opportunistic soft-error detection techniques to cost-effectively improve the streaming processors reliability. Observing that the streaming processors are not fully utilized during the branch divergence and pipeline stalls caused by the long latency operations, we propose to Recycle the streaming processors Idle time for Soft-Error detection (RISE) and obtain the good fault coverage with negligible performance degradation. RISE is composed of full-RISE and partial-RISE. Full-RISE selectively triggers the redundancy for a set of warps so that leverages the fully idled streaming processors during the pipeline stall time for the error detection. Partial-RISE performs the redundancy for a number of threads in certain warps using the partially idled streaming processors during the branch divergence. Our experimental results show that RISE shows strong capability in improving the SPs soft-error reliability by 43% with negligible (e.g. 4%) performance loss.
Jingweijia Tan, Xin Fu 0001
PACT2
2012 Aurora: A thermally resilient photonic network-on-chip architecture
abstract
With silicon optical technology moving towards maturity, the use of photonic network-on-chip (NoCs) for global chip communication is emerging as a promising solution to communication requirements of future many core processors. It is expected that photonic NoCs will play an important role in alleviating current power, latency, and bandwidth constraints. However, photonic NoCs are sensitive to ambient temperature variations because their basic constituents, ring resonators, are themselves sensitive to those variations. Since ring resonators are basic building blocks for photonic modulators, switches, multiplexers, and demultiplexers, variations of on-chip temperature pose serious challenges to the proper operation of photonic NoCs. Proposed methods that mitigate the effects of temperature at device level are either difficult to use in CMOS processes or not suitable for large scale implementation. In this paper, we propose Aurora, a thermally resilient photonic NoC architecture design that supports reliable and low bit error rate (BER) on-chip communications in the presence of large temperature variations. Our proposed architecture leverages solutions at both device and architecture layers that synergistically provide significant improvements. To compensate for small temperature variations, our design varies the bias current through ring resonators. For larger temperature variations, we propose architecture-level techniques to re-route messages away from hot regions, and through cooler regions, to their destinations, thereby lowering BER. Our simulation results show that Aurora provides a robust architectural solution to handle temperature variation effects on future photonic NoCs. For instance, average BER and message error rate (MER) are reduced by 78% and 30% respectively when the combined device and architectural technique (SPF) is applied. From the perspective of power efficiency, Aurora is also superior to conventional photonic NoC architectures by as much as 33%.
Amer Qouneh, Zhong-Qi Li, Madhura Joshi, Wangyuan Zhang, Xin Fu 0001, Tao Li 0006
ICCD5
2010 Architecting reliable multi-core network-on-chip for small scale processing technology
abstract
The trend towards multi-/many- core design has made network-on-chip (NoC) a crucial component of future microprocessors. With CMOS processing technologies continuously scaling down to the nanometer regime, effects such as process variation (PV) and negative bias temperature instability (NBTI) significantly decrease hardware reliability and lifetime. Therefore, it is imperative for multi-core architects to consider and mitigate these effects in NoCs implemented using small-scale processing technology. This paper reports on a first step to optimize NoC architecture reliability in light of both PV and NBTI effects. We propose novel techniques that can hierarchically alleviate PV and NBTI effects on NoC while leveraging their benign interaction. Our low-level design improves PV and NBTI efficiency of key components (e.g. virtual channel allocation logics, virtual channels) of critical paths of the pipelined router microarchitecture. Our high-level mechanisms leverage NBTI degradation and PV information from multiple routers to intelligently route packets, delivering optimized performance-power-reliability across the NoC substrate. Experimental results show that our intra-router level techniques (i.e. VA_M1 and VC_M2) reduce guardband by 47% while improving network throughput by 24%. Our inter-router optimization scheme (i.e. IR_M3) results in 50% guardband reduction and 19% network latency improvement.
Xin Fu 0001, Tao Li 0006, José A. B. Fortes
DSN1
2009 Soft error vulnerability aware process variation mitigation
abstract
As transistor process technology approaches the nanometer scale, process variation significantly affects the design and optimization of high performance microprocessors. Prior studies have shown that chip operating frequency and leakage power can have large variations due to fluctuations in transistor gate length and sub-threshold voltage. In this work, we study the impact of process variation on microarchitecture soft error robustness, an increasing reliability design challenge in the billion-transistor chip era. We explore two techniques that can effectively mitigate the effect of design parameter variation while significantly enhancing microarchitecture soft error reliability. Our first technique is entry-based. It tolerates the deleterious impact of variable latency techniques on soft error reliability by reducing the quantity and residency cycle of vulnerable bits in the microarchitecture structure at a fine granularity. Our second technique is structure-based. It applies body biasing schemes to dynamically adapt transistor sub-threshold voltage (and hence device-level soft error robustness) to the program reliability characteristics at a coarse granularity. We also combine the two techniques which further produces improved results. Compared to existing process variation tolerant schemes, our proposed techniques achieve optimal trade-offs between reliability, performance, and power. To our knowledge, this paper presents the first study on characterizing and optimizing processor microarchitecture resilience to soft errors in light of process variation.
Xin Fu 0001, Tao Li 0006, José A. B. Fortes
HPCA1
2008 Combined circuit and microarchitecture techniques for effective soft error robustness in SMT processors
abstract
As semiconductor technology scales, reliability is becoming an increasingly crucial challenge in microprocessor design. The rSRAM and voltage scaling are two promising circuit-level radiation hardening techniques to increase soft error robustness of a SRAM-based storage cell. However, applying circuit-level radiation hardening techniques to all on-chip transistors will result in significant overhead in performance and power consumption. In this paper, we propose microarchitecture support that allows cost-effective implementation of radiation hardened key microarchitecture structures (e.g. issue queue and reorder buffer) in SMT processors using soft error robust circuit techniques. Our study shows that the combined circuit and microarchitecture techniques achieve attractive tradeoffs between reliability, performance and power.
Xin Fu 0001, Tao Li 0006, José A. B. Fortes
DSN1
2008 Optimizing Issue Queue Reliability to Soft Errors on Simultaneous Multithreaded Architectures
abstract
The issue queue (IQ) is a key microarchitecture structure for exploiting instruction-level and thread-level parallelism in dynamically scheduled simultaneous multithreaded (SMT) processors. However, exploiting more parallelism yields high susceptibility to transient faults on a conventional IQ. With the rapidly increasing soft error rates, the IQ is likely to be a reliability hot-spot on SMT processors fabricated with advanced technology nodes using smaller and denser transistors with lower threshold voltages and tighter noise margins. In this paper, we explore microarchitecture techniques to optimize IQ reliability to soft error on SMT architectures. We propose to use off-line instruction vulnerability profiling to identify reliability critical instructions. The gathered information is then used to guide reliability-aware instruction scheduling and resource allocation in multithreaded execution environments. We evaluate the efficiency of the proposed schemes across various SMT workload mixes. Extensive simulation results show that, on average, our microarchitecture level soft error mitigation techniques can significantly reduce IQ vulnerability by 42% with 1% performance improvement. To maintain runtime IQ reliability for pre-defined thresholds, we propose dynamic vulnerability management (DVM) mechanisms. Experimental results show that our DVM techniques can effectively achieve desired reliability/performance tradeoffs.
Xin Fu 0001, Wangyuan Zhang, Tao Li 0006, José A. B. Fortes
ICPP1
2008 NBTI tolerant microarchitecture design in the presence of process variation
abstract
Negative bias temperature instability (NBTI), which reduces the lifetime of PMOS transistors, is becoming a growing reliability concern for sub-micrometer CMOS technologies. Parametric variation introduced by nano-scale device fabrication inaccuracy can exacerbate the PMOS transistor wear-out problem and further reduce the reliable lifetime of microprocessors. In this work, we propose microarchitecture design techniques to combat the combined effect of NBTI and process variation (PV) on the reliability of high-performance microprocessors. Experimental evaluation shows our proposed process variation aware (PV-aware) NBTI tolerant microarchitecture design techniques can considerably improve the lifetime of reliability operation while achieving an attractive trade-off with performance and power.
Xin Fu 0001, Tao Li 0006, José A. B. Fortes
MICRO1
2008 ORBIT: Effective Issue Queue Soft-Error Vulnerability Mitigation on Simultaneous Multithreaded Architectures Using Operand Readiness-Based Instruction Dispatch
abstract
With the advance of semiconductor processing technology, soft errors have become an increasing cause of failures of microprocessors fabricated using smaller and more densely integrated transistors with lower threshold voltages and tighter noise margins. With diminishing performance returns on wider issue superscalar processors, the microprocessor design industry has opted for using simultaneous multithreaded (SMT) architectures in commercial processors to exploit thread-level parallelism (TLP). SMT techniques enhance overall system performance but also introduce greater susceptibility to soft errors - concurrently executing multiple threads exposes many program runtime states to soft-error strikes at any given time. The issue queue (IQ) is a key micro architecture structure to exploit instruction-level and thread-level parallelism. On SMT processors, the IQ buffers a large number of instructions from multiple threads and is more susceptible to soft-error strikes. In this paper, we explore the use of operand-readiness-based instruction dispatch (ORBIT) as an effective mechanism to mitigate IQ soft-error vulnerability on SMT processors. We observe that IQ soft-error vulnerability is largely affected by instructions waiting for their source operands. The overall IQ soft-error vulnerability can be effectively reduced by minimizing the number of waiting instructions and their residency cycles in the IQ. We develop six techniques that aim to improve IQ reliability with negligible performance degradation on SMT processors. Moreover, we extend our techniques with prediction methods that can anticipate the readiness of source operands ahead of time. The ORBIT schemes integrated with reliability-awareness and readiness prediction achieve more attractive reliability/performance trade-offs. The best of the proposed schemes (e.g. Predict_DelayACE) reduces IQ vulnerability by 79% with only 1% throughput IPC and 3% harmonic IPC reduction across all studied workloads.
Xin Fu 0001, Tao Li 0006, José A. B. Fortes
SBAC-PAD1
2007 An Analysis of Microarchitecture Vulnerability to Soft Errors on Simultaneous Multithreaded Architectures
abstract
Semiconductor transient faults (i.e. soft errors) have become an increasingly important threat to microprocessor reliability. Simultaneous multithreaded (SMT) architectures exploit thread-level parallelism to improve overall processor throughput. A great amount of research has been conducted in the past to investigate performance and power issues of SMT architectures. Nevertheless, the effect of multithreaded execution on a microarchitecture's vulnerability to soft error remains largely unexplored. To address this issue, we have developed a microarchitecture level soft error vulnerability analysis framework for SMT architectures. Using a mixed set of SPEC CPU 2000 benchmarks, we quantify the impact of multithreading on a wide range of microarchitecture structures. We examine how the baseline SMT microarchitecture reliability profile varies with workload behavior, the number of threads and fetch policies. Our experimental results show that the overall vulnerability rises in multithreading architectures, while each individual thread shows less vulnerability. By considering both performance and reliability, SMT outperforms superscalar architectures. The SMT reliability and its tradeoff with performance vary across different fetch policies. With a detailed analysis of the experimental results, we point out a set of potential opportunities to reduce SMT microarchitecture vulnerability, which can serve as guidance to exploiting thread-aware reliability optimization techniques in the near future. To our knowledge, this paper presents the first effort to characterize microarchitecture vulnerability to soft error on SMT processors
Wangyuan Zhang, Xin Fu 0001, Tao Li 0006, José A. B. Fortes
ISPASS2
2006 Characterizing Microarchitecture Soft Error Vulnerability Phase Behavior
abstract
Computer systems increasingly depend on exploiting program dynamic behavior to optimize performance, power and reliability. Prior studies have shown that program execution exhibits phase behavior in both performance and power domains. Reliabilityoriented program phase behavior, however, remains largely unexplored. As semiconductor transient faults (soft errors) emerge as a critical challenge to reliable system design, characterizing program phase behavior from a reliability perspective is crucial in order to apply dynamic fault-tolerant mechanisms and to optimize performance/reliability trade-offs. In this paper, we compute run-time program vulnerability to soft errors on four microarchitecture structures (i.e. instruction window, reorder buffer, function units and wakeup table) in a high-performance out-of-order execution superscalar processor. Experimental results on the SPEC2000 benchmarks show a considerable amount of time varying behavior in reliability measurements. Our study shows that a single performance metric, such as IPC, cache miss or branch misprediction, is not a good indicator for program vulnerability. The vulnerabilities of the studied microarchitecture structures are then correlated with program code-structure and run-time events to identify vulnerability phase behavior. We observed that both program code-structure and run-time events appear promising in classifying program reliability phase behavior. Overall, performance counter based schemes achieved an average Coefficient of Variation (COV) of 3.5%, 4.5%, 4.3% and 5.7% on the instruction queue, reorder buffer, function units and the wakeup table, while basic block vectors offer COVs of 4.9%, 5.8%, 5.4% and 6% on the four studied microarchitecture structures respectively. We found that in general, tracking performance metrics performs better than tracking control flow in identifying reliability phase behavior of applications. To our knowledge, this paper is the first to characterize program reliability phase behavior at the microarchitecture level.
Xin Fu 0001, James Poe, Tao Li 0006, José A. B. Fortes
MASCOTS1