Chengmo Yang

dblp:06/4401 · DBLP profile ↗
← Back
74ranked-venue papers
15as first author
11since 2021 · last 2025
0000-0003-0978-1504ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 66 · 15 first-author · 9 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 2 since 2021Security and privacy · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 WSSR: Weight Set Segmentation and Recovery for Fault Resilient Transformers
abstract
The usage of the transformer model architecture has rapidly become a popular choice in fields related to natural language processing, computer vision, and many other applications. These models have rapidly grown in size and complexity, now containing billions of parameters. Due to their high popularity, the integrity of these models becomes a critical concern. Faults such as bit flips in memory could be caused not only by radiation-induced soft errors, electrical disturbances, thermal fluctuations and aging hardware, but also by intentional Bit Flip Attacks (BFAs). While classic fault tolerance solutions such as Error Correcting Codes (ECC) can mitigate such faults to some extent, their application is often constrained by cost, power, or design limitations. To develop a low-cost and effective fault tolerance solution, this work studies the fault resilience of two large-scale transformer models - CLIP-ViT and Google-ViT. Through systematic fault injection, we find that different layers exhibit varying fault sensitivities, and that large-magnitude parameter perturbations correlate strongly with performance degradation. Building upon this insight, we propose WSSR, a lightweight fault detection and recovery method that segments model weights into sets and flags changes in set membership for fault recovery. Our experiments show that WSSR can restore model accuracy to 91.7% and 84.3% of the original for CLIPViT and Google-ViT, respectively, on the ImageNet1K dataset with only 0.87% and 7.37% compute overhead and virtually no storage overhead-making it practical for deployment on resource-constrained systems.
Ntsee Ndingwan, Chengmo Yang
ICCD2
2025 Targeted Fault Injection Attack on Semantic Segmentation Models
abstract
Semantic segmentation, a perception method that labels each pixel in an image with a category or class, is widely used in various domains, such as medical imaging and autonomous driving. The safety-critical nature of these applications imposes strict requirements of the underlying hardware accelerators being secure. Prior studies have shown that hardware accelerators are vulnerable to fault injection attacks, compromising their integrity and reliability. While these fault injection attacks are capable of causing a high accuracy drop, they are difficult to control, as faults affect the computation across random classes. In comparison, this work presents a targeted fault injection attack on black-box segmentation models. It first conducts a vulnerability analysis, demonstrating that faults injected into different parts of the model (e.g., encoder vs decoder) have distinct behaviors in terms of the region within the segmentation map affected and the pixel-level differences caused. Furthermore, this work reveals a linear relationship between the timing and duration of the fault and the region within the segmentation map affected. This translates to a new type of security vulnerability that an adversary can inject faults targeting regions that are more likely to contain critical classes, such as traffic lights or traffic signs, in the context of autonomous driving. The attack is implemented on different segmentation models, including ERFnet, ENet, and FPN with different backbones, to demonstrate its effectiveness.
Jhon Ordoñez, Chengmo Yang
ICCD2
2024 Derailed: Arbitrarily Controlling DNN Outputs with Targeted Fault Injection Attacks
abstract
Hardware accelerators have been widely deployed to improve the efficiency of DNN execution in terms of performance, power, and time predictability. Yet recent studies have shown that DNN accelerators are vulnerable to fault injection attacks, compromising their integrity and reliability. Classic fault injection attacks are capable of causing a high overall accuracy drop. However, one limitation is that they are difficult to control, as faults affect the computation across random classes. In comparison, this paper presents a controlled fault injection attack, capable of derailing arbitrary inputs to a targeted range of classes. Our observation is that the fully connected (FC) layers greatly impact inference results, whereas the computation in the FC layer is typically performed in order. Leveraging this fact, an adversary can perform a controlled fault injection attack even to a black-box DNN model. Specifically, this attack adopts a two-step search process that first identifies the time window during which the FC layer is computed and then pinpoints the targeted classes. This attack is implemented with clock glitching, and the target DNN accelerator is a DPU implemented in the FPGA. The attack is tested on three popular DNN models, namely, ResNet50, InceptionVl, and MobileNetV2. Results show that up to 93 % of inputs are derailed to the attacker-specified classes, demonstrating its effectiveness.
Jhon Ordoñez, Chengmo Yang
DATE2
2024 Enhancing DNN Accelerator Integrity via Selective and Permuted Recomputation
abstract
Hardware accelerators have been widely deployed in many machine learning applications due to their superior performance and energy efficiency. However, these accelerators are vulnerable to fault injection attacks, compromising their integrity and reliability. In particular, recent studies have revealed a targeted attack on black-box DNN models, which, through glitching the execution of the fully connected (FC) layer, is capable of derailing the DNN outputs to arbitrary classes. To defend DNN accelerators against this severe attack, this paper proposes a selective and permuted recomputation scheme. Instead of adopting dual or triple modular redundancy, which incurs high overhead, the proposed scheme selects a subset of critical FC outputs for recomputation. Meanwhile, it permutes the computation of the FC layer to prevent an adversary from pinpointing the exact time of executing the target class. The proposed defense is evaluated on three popular DNN models, namely, ResNet-50, InceptionV3, and MobileNetV3. Results show that under fault injection attacks, it can successfully recover 90--95% of the models' original accuracy, achieved with less than 1.61% runtime overhead and no storage overhead.
Jhon Ordoñez, Chengmo Yang
ICCAD2
2023 NeuroPots: Realtime Proactive Defense against Bit-Flip Attacks in Neural Networks
Qi Liu 0017, Jieming Yin, Wujie Wen, Chengmo Yang, Shi Sha
USENIX Security Symposium4
2021 A Self-Test Framework for Detecting Fault-induced Accuracy Drop in Neural Network Accelerators
abstract
Hardware accelerators built with SRAM or emerging memory devices are essential to the accommodation of the ever-increasing Deep Neural Network (DNN) workloads on resource-constrained devices. After deployment, however, the performance of these accelerators is threatened by the faults in their on-chip and off-chip memories where millions of DNN weights are held. Different types of faults may exist depending on the underlying memory technology, degrading inference accuracy. To tackle this challenge, this paper proposes an online self-test framework that monitors the accuracy of the accelerator with a small set of test images selected from the test dataset. Upon detecting a noticeable level of accuracy drop, the framework uses additional test images to identify the corresponding fault type and predict the severeness of faults by analyzing the change in the ranking of the test images. Experimental results show that our method can quickly detect the fault status of a DNN accelerator and provide accurate fault type and fault severeness information, allowing for subsequent recovery and self-healing process.
Fanruo Meng, Fateme S. Hosseini, Chengmo Yang
ASP-DAC3
2021 Modeling of Threshold Voltage Distribution in 3D NAND Flash Memory
abstract
3D NAND flash memory faces unprecedented complicated interference than planar NAND flash memory, resulting in more concern regarding reliability and performance. Stronger error correction code (ECC) and adaptive reading strategies are proposed to improve the reliability and performance taking a threshold voltage (Vth) distribution model as the backbone. However, the existing modeling methods are challenged to develop such a Vthdistribution model for 3D NAND flash memory. To facilitate it, in this paper, we propose a machine learning-based modeling method. It employs a neural network taking advantage of the existing modeling methods and fully considers multiple interferences and variations in 3D NAND flash memory. Compared with state-of-the-art models, evaluations demonstrate it is more accurate and efficient for predicting Vthdistribution.
Fei Wu 0005, Jian Zhou 0004, Meng Zhang 0014, Chengmo Yang, Zhonghai Lu, Yu Wang 0168, Changsheng Xie 0001
DATE5
2021 Charger-Surfing: Exploiting a Power Line Side-Channel for Smartphone Information Leakage
Patrick Cronin, Xing Gao 0001, Chengmo Yang, Haining Wang 0001
USENIX Security Symposium3
2021 A Compile-Time Framework for Tolerating Read Disturbance in STT-RAM
abstract
Spin-transfer torque magnetic random access memory (STT-RAM) is one of the most promising candidates for next-generation on-chip memories. While STT-RAM offers high density, negligible leakage power, and fast access speed, it also suffers from read-disturbance errors, that is, read operations might accidentally change the value of the accessed memory location. Although these errors could be mitigated by applying restore-after-read operations, the energy overhead would be significant. To reduce such overhead, this article presents an application-level and architecture-independent framework, which selectively inserts restore operations under the guidance of a compiler. This work first introduces a new concept of disturbance chain and then analyzes the vulnerability of each load instruction on a chain to read disturbance errors. This work further proposes a number of compile-time code optimizations to reduce the number of vulnerable loads and hence the associated restore overhead. The proposed compiler optimizations are implemented in LLVM. Experiments in Gem5 show up to 98.6% reduction in the number of restore operations and 48% savings of the energy overhead while maintaining 99.8% coverage of read disturbance errors.
Fateme S. Hosseini, Chengmo Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 DEPS: Exploiting a Dynamic Error Prechecking Scheme to Improve the Read Performance of SSD
abstract
3-D NAND flash memory is gradually being widely used in solid state drives (SSDs), leading to increasing storage capacity. However, the read performance of SSD is sacrificed for decoding operations which are executed to guarantee the data reliability. No matter whether the data have bit errors, they will be sent to error correcting code (ECC) engine to decode, introducing a high read delay of SSD. Error prechecking can help to avoid the redundant decoding operations for the error-free data, but it induces extra checking overhead to the error data. Motivated by this, we carry out comprehensive experiments to analyze the distribution of bit errors in 3-D NAND flash memory. The preliminary experimental results show that there are a large number of pages read without errors in the early lifetime of 3-D NAND flash memory. Based on the observations and analyses, we propose a model to estimate the error-free ratio, and utilize it to design a dynamic error prechecking scheme (DEPS) to bypass the decoding operation for the error-free data in 3-D NAND flash memory and improve the read performance of SSD. Furthermore, by dividing a large page into small subpages, DEPS releases more error-free data, which significantly improves the read performance of SSD. Evaluation results from real-world traces demonstrate that by implementing DEPS, the average read performance of SSD is enhanced by 35%-55% with 3-D MLC NAND flash memory.
Fei Wu 0005, Meng Zhang 0014, Chengmo Yang, Zhonghai Lu, Jiguang Wan 0001, Changsheng Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 Tolerating Defects in Low-Power Neural Network Accelerators Via Retraining-Free Weight Approximation
abstract
Hardware accelerators are essential to the accommodation of ever-increasing Deep Neural Network (DNN) workloads on the resource-constrained embedded devices. While accelerators facilitate fast and energy-efficient DNN operations, their accuracy is threatened by faults in their on-chip and off-chip memories, where millions of DNN weights are held. The use of emerging Non-Volatile Memories (NVM) further exposes DNN accelerators to a non-negligible rate of permanent defects due to immature fabrication, limited endurance, and aging. To tolerate defects in NVM-based DNN accelerators, previous work either requires extra redundancy in hardware or performs defect-aware retraining, imposing significant overhead. In comparison, this paper proposes a set of algorithms that exploit the flexibility in setting the fault-free bits in weight memory to effectively approximate weight values, so as to mitigate defect-induced accuracy drop. These algorithms can be applied as a one-step solution when loading the weights to embedded devices. They only require trivial hardware support and impose negligible run-time overhead. Experiments on popular DNN models show that the proposed techniques successfully boost inference accuracy even in the face of elevated defect rates in the weight memory.
Fateme S. Hosseini, Fanruo Meng, Chengmo Yang, Wujie Wen, Rosario Cammarota
ACM Trans. Embed. Comput. Syst.3
2020 Monitoring the Health of Emerging Neural Network Accelerators with Cost-effective Concurrent Test
abstract
ReRAM-based neural network accelerator is a promising solution to handle the memory-and computation-intensive deep learning workloads. However, it suffers from unique device errors. These errors can accumulate to massive levels during the run time and cause significant accuracy drop. It is crucial to obtain its fault status in real-time before any proper repair mechanism can be applied. However, calibrating such statistical information is non-trivial because of the need of a large number of test patterns, long test time, and high test coverage considering that complex errors may appear in million-to-billion weight parameters. In this paper, we leverage the concept of comer data that can significantly confuse the decision making of neural network model, as well as the training algorithm, to generate only a small set of test patterns that is tuned to be sensitive to different levels of error accumulation and accuracy loss. Experimental results show that our method can quickly and correctly report the fault status of a running accelerator, outperforming existing solutions in both detection efficiency and cost
Qi Liu 0017, Tao Liu 0023, Zihao Liu 0015, Wujie Wen, Chengmo Yang
DAC5
2020 A Crowd-Based Explosive Detection System with Two-Level Feedback Sensor Calibration
abstract
Large, open, public events, such as marathons and festivals, have always presented a unique safety challenge. These sprawling events, which can take up entire city blocks or stretch for many miles, can draw tens to hundreds of thousands of spectators and in some cases have open admission. As it is impracticable to guarantee the subjection of every event-goer to a security screening, we propose a crowd-based explosive detection system that uses a multitude of low-cost ChemFET sensors which are distributed to attendees. As the sensors offer limited accuracy, we further propose a server-based decision-making framework that utilizes a two-level feedback loop between the sensors and the server and explores spatial and temporal locality of the collected data to overcome the inherent low-accuracy of individual sensors. We thoroughly explore two distinct detection schemes, stressing their performance under a myriad of conditions, thus showing that such a crowd-based detection system comprised of low-cost and low-accuracy sensors can deliver high detection accuracy with minimal false positives.
Chengmo Yang, Patrick Cronin, Agamyrat Agambayev, Sule Ozev, A. Enis Çetin, Alex Orailoglu
ICCAD1
2020 Low Overhead Online Data Flow Tracking for Intermittently Powered Non-Volatile FPGAs
abstract
Energy harvesting is an attractive way to power future Internet of Things (IoT) devices since it can eliminate the need for battery or power cables. However, harvested energy is intrinsically unstable. While Field-programmable Gate Array (FPGAs) have been widely adopted in various embedded systems, it is hard to survive unstable power since all the memory components in FPGA are based on volatile Static Random-access Memory (SRAMs). The emerging non-volatile memory-based FPGAs provide promising potentials to keep configuration data on the chip during power outages. Few works have considered implementing efficient runtime intermediate data checkpoint on non-volatile FPGAs. To realize accumulative computation under intermittent power on FPGA, this article proposes a low-cost design framework, Data-Flow-Tracking FPGA (DFT-FPGA), which utilizes binary counters to track intermediate data flow. Instead of keeping all on-chip intermediate data, DFT-FPGA only targets on necessary data that is labeled by off-line analysis and identified by an online tracking system. The evaluation shows that compared with state-of-the-art techniques, DFT-FPGA can realize accumulative computing with less off-line workload and significantly reduce online roll-back time and resource utilization.
Xinyi Zhang 0001, Clay Patterson, Yongpan Liu, Chengmo Yang, Chun Jason Xue, Jingtong Hu
ACM J. Emerg. Technol. Comput. Syst.4
2019 A Fault-Tolerant Neural Network Architecture
abstract
New DNN accelerators based on emerging technologies, such as resistive random access memory (ReRAM), are gaining increasing research attention given their potential of "in-situ" data processing. Unfortunately, device-level physical limitations that are unique to these technologies may cause weight disturbance in memory and thus compromising the performance and stability of DNN accelerators. In this work, we propose a novel fault-tolerant neural network architecture to mitigate the weight disturbance problem without involving expensive retraining. Specifically, we propose a novel collaborative logistic classifier to enhance the DNN stability by redesigning the binary classifiers augmented from both traditional error correction output code (ECOC) and modern DNN training algorithm. We also develop an optimized variable-length "decode-free" scheme to further boost the accuracy under fewer number of classifiers. Experimental results on cutting-edge DNN models and complex datasets show that the proposed fault-tolerant neural network architecture can effectively rectify the accuracy degradation against weight disturbance for DNN accelerators with low cost, thus allowing for its deployment in a variety of mainstream DNNs.
Tao Liu 0023, Wujie Wen, Lei Jiang 0001, Yanzhi Wang 0001, Chengmo Yang, Gang Quan
DAC5
2019 WAS: Wear Aware Superblock Management for Prolonging SSD Lifetime
abstract
Superblocks are widely employed in SSDs for improving performance. However, the standard superblock organization which links blocks with the same block ID across planes into one superblock leads to SSDs' ineluctable lifetime waste due to inter-block wear tolerance variations. This work proposes a wear-aware superblock management, called WAS, which (1) dynamically organizes superblocks according to real-time block wear levels to make strong blocks relieve wear on weak ones, and (2) employs a wear-based garbage collection scheme to reduce inter-block wear gap. Comprehensive experiments are carried out in SSDsim. Results show that WAS greatly prolongs SSD lifetime by 51.3% compared with the state-of-the-art superblock management.
Shunzhuo Wang, Fei Wu 0005, Chengmo Yang, Jiaona Zhou, Changsheng Xie 0001, Jiguang Wan 0001
DAC3
2019 Compiler-Directed and Architecture-Independent Mitigation of Read Disturbance Errors in STT-RAM
abstract
High density, negligible leakage power, and fast read speed have made Spin-Transfer Torque Random Access Memory (STT-RAM) one of the most promising candidates for next generation on-chip memories. However, STT-RAM suffers from read-disturbance errors, that is, read operations might accidentally change the value of the accessed memory location. Although these errors could be mitigated by applying a restore-after-read operation, the energy overhead would be significant. This paper presents an architecture-independent framework to mitigate read disturbance errors while reducing the energy overhead, by selectively inserting restore operations under the guidance of the compiler. For that purpose, the vulnerability of load operations to read disturbance errors is evaluated using a specifically designed fault model; a code transformation technique is developed to reduce the number of vulnerable loads; and, an algorithm is proposed to selectively insert restore operations. The evaluation results show that the proposed technique can effectively reduce up to 97% of restore operations and 66% of the energy overhead while maintaining 99.8% coverage of read disturbance errors.
Fateme S. Hosseini, Chengmo Yang
DATE2
2019 Tuning Track-based NVM Caches for Low-Power IoT Devices
abstract
Track-based non-volatile memories, such as Domain Wall Memory (DWM) and Skyrmion, are promising candidates to be used as CPU caches due to their ultra-high density and low-static power. However, the access latency and energy of these devices are highly affected by the number of shift operation performed. Existing track-based NVM cache designs place logically-adjacent blocks close to each other, resulting in extra shift operations performed to access the data. This paper makes the observation that the block access pattern in cache is typically repetitive with a stride of N. Such pattern motivates us to propose a new cache block placement for track-based NVMs, aiming at reducing the number of shift operations throughout program execution.
Hoda Aghaei Khouzani, Chengmo Yang
ACM Great Lakes Symposium on VLSI2
2019 A Processing-In-Memory Implementation of SHA-3 Using a Voltage-Gated Spin Hall-Effect Driven MTJ-based Crossbar
abstract
Processing-In-Memory (PIM), which implements logic operations within memory cells, opens up a new direction on organizing data and computation. Leveraging resistive or magnetic characteristics of nonvolatile memory (NVM) devices, platforms such as PLiM and ReVAMP have been proposed. This paper presents a PIM implementation of SHA-3, a state-of-the-art secure hash algorithm using a Voltage-Gated Spin Hall-Effect (SHE) Driven magnetic tunnel junction (MTJ) based crossbar, which is able to achieve a complete set of Boolean operations. The work includes the design of the crossbar circuit, the instruction set, and both unpipelined and pipelined implementations of SHA-3. Experimental results show that the proposed SHE MTJ-based implementation is able to achieve 2.16X higher throughput than a state-of-the-art Resistive RAM based SHA-3 implementation. Further throughput improvement can be achieved with multiple message hash (MMH) pipelining.
Chengmo Yang
ACM Great Lakes Symposium on VLSI1
2019 A Scalable and Process Variation Aware NVM-FPGA Placement Algorithm
abstract
As non-volatile memory (NVM) based FPGAs gain increasing popularity, FPGA synthesis tools start to tune the synthesis flow to match NVM characteristics. State-of-the-art NVM FPGA placement algorithms tried to reduce the high reconfiguration cost induced by the costly NVM programming process. However, they are not only limited in scalability but also fail to consider process variation. This paper aims to overcome these limitations. Blocks in the NVM-FPGA are no longer uniform but classified into fast, slow, and dead blocks. Moreover, the proposed placement algorithm reduces computation complexity by not searching the entire design space for an optimal solution with minimum reconfiguration cost, but computing the reconfiguration cost just-in-time. Verilog-to-Routing (VTR)-based implementation confirms its effectiveness in reducing critical path length and speeding up the placement process, while still saving reconfiguration cost by up to 74.2%.
Chengmo Yang
ACM Great Lakes Symposium on VLSI1
2019 Detecting Gas Vapor Leaks through Uncalibrated Sensor Based CPS
abstract
While Volatile Organic Compounds (VOC) and ammonia have a place in our daily lives, their leakage into the environment is harmful to human health. In order to prevent and detect gaseous leaks of harmful VOCs, a cyber-physical system (CPS) comprised of ordinary people or first responders is proposed. This CPS uses small, low-cost sensors coupled to smart phones or mobile devices with the necessary computation and communication capabilities. The efficacy of such a CPS hinges on its ability to address technical challenges stemming from the fact that identically produced sensors may produce different results under the same conditions due to sensor drift, noise, or resolution errors. The proposed system makes use of time-varying signals produced by sensors to detect gas leaks. Sensors sample the gas vapor level in a continuous manner and time-varying sensor data is processed using deep neural networks. One of the neural networks (NN) is an energy efficient Additive Neural Network (AddNet) which can be implemented in host devices. The second NN is the discriminator of a GAN and the third a regular convolutional NN. AddNet produces comparable VOC gas leak detection results to regular convolutional networks while reducing area requirements by two thirds.
Diaa Badawi, Sule Ozev, Jennifer Blain Christen, Chengmo Yang, Alex Orailoglu, A. Enis Çetin
ICASSP4
2019 Covert Data Exfiltration Using Light and Power Channels
abstract
As the Internet of Things (IoT) continues to expand into every facet of our daily lives, security researchers have warned of its myriad security risks. While denial-of-service attacks and privacy violations have been at the forefront of research, covert channel communications remain an important concern. Utilizing a Bluetooth controlled light bulb, we demonstrate three separate covert channels, consisting of current utilization, luminosity and hue. To study the effectiveness of these channels, we implement exfiltration attacks using standard off-the-shelf smart bulbs and RGB LEDs at ranges of up to 160 feet. We analyze the identified channels for throughput, generality and stealthiness, and report transmission speeds of up to 832 bps.
Patrick Cronin, Charles Gouert, Dimitris Mouris, Nektarios Georgios Tsoutsos, Chengmo Yang
ICCD5
2018 Performance analysis on structure of racetrack memory
abstract
Racetrack Memory(RM) has attracted abundant attention of memory researchers recently. RM can achieve ultrahigh storage density, fast access velocity and non-volatility. Former research has demonstrated that RM has potential to serve as on-chip cache or main memory. However, RM has more flexibility and difficulty in design space of main memory because it has more device level design parameters. The layout of macro unit (MU) needs trade-off among area, access performance and energy consumption, and its shift operation introduces extra dimension of design space. In this paper, we explore these design parameters and analyze their relationship in memory design space in both device and system levels. Based on the results, we also propose a hybrid MU structure to further optimize read intensive applications. Experimental results demonstrated the existence of regularity between design parameters and performance features. The optimized layout of racetrack MU is suggested for application areas such as big-data and IoT which need cost-effective and energy-efficient memory respectively. Together with hybrid MU structures, RM can be designed with more flexibility so that specific structures are suitable for specific applications which make “All stack optimization” possible in memory structure level.
Chao Zhang 0007, Qingda Hu, Chengmo Yang, Jiwu Shu
ASP-DAC4
2018 FastGC: accelerate garbage collection via an efficient copyback-based data migration in SSDs
abstract
Copyback is an advanced command contributing to accelerating data migration in garbage collection (GC). Unfortunately, detecting copyback feasibility (whether copyback can be carried out with assurable reliability) against data corruption in the traditional copyback-based GC causes an expensive performance penalty. This paper first explores copyback error characteristics on real NAND flash chips, then proposes a fast garbage collection scheme called FastGC. It utilizes copyback error characteristics to efficiently detect copyback feasibility of data instead of transferring out all valid data for detecting. Experiment results in the SSDsim show the proposed FastGC greatly promotes write response time and read response time by up to 44.2% and 66.3% respectively, compared to the traditional copyback-based GC.
Fei Wu 0005, Jiaona Zhou, Shunzhuo Wang, Yajuan Du, Chengmo Yang, Changsheng Xie 0001
DAC5
2018 A collaborative defense against wear out attacks in non-volatile processors
abstract
While the Internet of Things (IoT) keeps advancing, its full adoption is continually blocked by power delivery problems. One promising solution is Non-Volatile (NV) processors, which harvest energy for themselves and employ a NV memory hierarchy. This allows them to perform computations when power is available, checkpoint and hibernate when power is scarce, and resume their work at a later time. However, utilizing NV memory creates new security vulnerabilities in the form of wear out attacks in the register file. This paper explores the dangers of this design oversight and proposes a mitigation strategy that takes advantage of the unique properties and operating characteristics of NV processors. The proposed defense integrates the power management unit and a two-level register rotation approach, which improves NV processor endurance by 30.1x in attack situations and an average of 7.1x in standard workloads.
Patrick Cronin, Chengmo Yang, Yongpan Liu
DAC2
2018 Architecting data placement in SSDs for efficient secure deletion implementation
abstract
Secure deletion ensures user privacy by permanently removing invalid data from the secondary storage. This process is particularly critical to solid state drives (SSDs) wherein invalid data are generated not only upon deleting a file but also upon updating a file of which the user is not aware. While previous secure deletion schemes are usually applied to all invalid data on the SSD, our observation is that in many cases security is not required for all files on the SSD. This paper proposes an efficient secure deletion scheme targeting only the invalid data of files marked as “secure” by the user. A security-aware data allocation strategy is designed, which separates secure and unsecure data at lower (block) level but mixes them at higher levels of SSD hierarchical organization. Block-level separation minimizes secure deletion cost, while higher-level mixing mitigates the adverse impact of secure deletion on SSD lifetime. A two-level block management scheme is further developed to scatter secure blocks over the SSD for wear leveling. Experiments on real-world benchmarks confirm the advantage of the proposed scheme in reducing secure deletion cost and improving SSD lifetime.
Hoda Aghaei Khouzani, Chen Liu 0013, Chengmo Yang
ICCAD3
2018 Power- and Endurance-Aware Neural Network Training in NVM-Based Platforms
abstract
Neural networks (NNs) have become the go-to tool for solving many real-world recognition and classification tasks with massive and complex data sets. These networks require large data sets for training, which is usually performed on GPUs and CPUs in either a cloud or edge computing setting. No matter where the training is performed, it is subject to tight power/energy and data storage/transfer constraints. While these issues can be mitigated by replacing SRAM/DRAM with nonvolatile memories (NVMs) which offer near-zero leakage power and high scalability, the massive weight updates performed during training shorten NVM endurance and engender high write energy. In this paper, an NVM-friendly NN training approach is proposed. Weight update is redesigned to reduce bit flips in NVM cells. Moreover, two techniques, namely, filter exchange and bitwise rotation, are proposed to respectively balance writes to different weights and to different bits of one weight. The proposed techniques are integrated and evaluated in Caffe. Experimental results show significant power savings and endurance improvements, while maintaining high inference accuracy.
Fanruo Meng, Chengmo Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Reliability-Aware Runtime Adaption Through a Statically Generated Task Schedule
abstract
Device scaling, increasing number of components in a single chip, varying environmental issues, and aging effects have brought severe reliability challenges that impose tight constraints on the operation of a system. To cope with these challenges, this paper proposes a reliability-aware scheduling framework that combines static and dynamic analyses to improve the overall system resiliency to different kinds of faults (i.e., intermittent, transient, and permanent). The static analysis technique employs genetic algorithms to optimize the overall system reliability by considering reliability level (RL) as an intermediate scheduling dimension and creating a task-to-RL mapping. This enables the RL-to-core mapping to be efficiently adapted at runtime according to fault rate variations, while the task-to-RL mapping can still be reused. The dynamic analysis tracks faults appearing in each core and measures the time correlation of those faults to update the RL-to-core mapping. The proposed reliability-aware framework is implemented in a state-of-the-art runtime system, Delaware Adaptive Run-Time System, so as to quantitatively show the advantages of using the overall framework in existing multicore platforms. Experimental results show that the proposed technique delivers up to 30% improvement in application execution time and up to 72% improvement in faults occurring at runtime.
Laura Rozo, Aaron Myles Landwehr, Chengmo Yang
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Leveraging Compiler Optimizations to Reduce Runtime Fault Recovery Overhead
abstract
Smaller feature size, lower supply voltage, and faster clock rates have made modern computer systems more susceptible to faults. Although previous fault tolerance techniques usually target a relatively low fault rate and consider error recovery less critical, with the advent of higher fault rates, recovery overhead is no longer negligible. In this paper, we propose a scheme that leverages and revises a set of compiler optimizations to design, for each application hotspot, a smart recovery plan that identifies the minimal set of instructions to be re-executed in different fault scenarios. Such fault scenario and recovery plan information is efficiently delivered to the processor for runtime fault recovery. The proposed optimizations are implemented in LLVM and GEM5. The results show that the proposed scheme can significantly reduce runtime recovery overhead by 72%.
Fateme S. Hosseini, Pouya Fotouhi, Chengmo Yang, Guang R. Gao
DAC3
2017 Age-aware Logic and Memory Co-Placement for RRAM-FPGAs
abstract
Resistive RAM (RRAM) is a promising non-volatile memory (NVM) device which can replace traditional SRAM as on-chip storage for logic and data in FPGAs. While RRAM outperforms SRAM by offering high scalability, low leakage power, and near-zero power-on delay, RRAM-FPGAs have limited programming cycles, and different writes frequencies of memory and logic blocks make the challenge more severe. To overcome this endurance challenge, we propose an age-aware placement framework for RRAM-FPGAs with uniform reconfigurable logic/memory units. The framework, consisting of a dynamic reconfiguration region allocation algorithm and a logic/memory co-placement algorithm, balances write distributions across the entire FPGA according to logic and memory write frequency differences. The proposed algorithms have been integrated into the VTR synthesis flow. Experiments show that the framework achieves 94.9% write reduction, thus effectively extending RRAM-FPGA programming cycles.
Chengmo Yang, Jingtong Hu
DAC2
2017 Leveraging access port positions to accelerate page table walk in DWM-based main memory
abstract
Domain Wall Memory (DWM) with ultra-high density and comparable read/write latency to DRAM is an attractive replacement for CMOS-based devices. Unlike DRAM, DWM has non-uniform data access latency that is proportional to the number of shift operations. While previous works have demonstrated the feasibility of using DWM as main memory and have proposed different ways to alleviate the impact of shift operations, none of them have addressed the performance-critical metadata accesses, in particular page table accesses. To bridge this gap, this paper aims at accelerating page table walk in DWM-based main memory from two innovative aspects. First of all, we propose a new page table layout and leverage the positions of access ports in DWM to differentiate the state of page table entries. In addition, we propose a technique to pre-align the access ports to the positions to be accessed in the near future, thus hiding shift latency to the maximum extent. Since both address translation and context switching are affected by page table access latency, the proposed technique can effectively improve system performance and user experience.
Hoda Aghaei Khouzani, Pouya Fotouhi, Chengmo Yang, Guang R. Gao
DATE3
2017 CooECC: A Cooperative Error Correction Scheme to Reduce LDPC Decoding Latency in NAND Flash
abstract
The storage capacity of NAND Flash has increased by scaling down to smaller cell size and using multi-level storage technology, but data reliability is degraded by severer retention errors. To ensure data reliability, error correction codes (ECC) are adopted, such as BCH and low-density parity check (LDPC) codes. However, BCH codes are insufficient when raw bit error rates (RBER) caused by retention errors are high. As a result, BCH codes are inevitably replaced with LDPC codes with stronger error correction capability. Traditional LDPC codes are used to independently correct bit errors in the LSB and MSB pages. Unfortunately, decoding latency in such two pages is significantly unbalanced, MSB pages take much higher latency due to higher RBER, leading to suboptimal flash read performance. This paper proposes a cooperative error correction scheme, called CooECC, to reduce LDPC decoding latency of the MSB page in NAND Flash. By exploiting data error characteristics introduced by retention errors, CooECC integrates the decoding result of the LSB page into the initial information of LDPC decoding for the MSB page, making it more accurate. This in turn enables decoding to converge at a higher rate. Simulation results show that for LDPC schemes with information lengths of 2KB and 4KB, the decoding latency can be reduced by up to 87% and 84%, respectively, when RBER is as high as 8.0 × 10^-3.
Meng Zhang 0014, Fei Wu 0005, Yajuan Du, Chengmo Yang, Changsheng Xie 0001, Jiguang Wan 0001
ICCD4
2017 Power-aware and cost-efficient state encoding in non-volatile memory based FPGAs
abstract
Non-volatile memory (NVM)-based FPGAs are expected to replace traditional SRAM-based FPGAs to achieve higher scalability, lower leakage power, and better reliability. In NVM-based FPGAs, dynamic power is the dominant power factor, and flip-flops exhibit the most intensive switching activities. While flip-flops can be implemented with NVM elements such as Magnetic Tunnel Junctions (MTJ), NVM cells suffer from high write energy, making it necessary to reduce dynamic power by minimizing bit flips. Furthermore, flip-flops are used to implement finite state machines, whose power and hardware cost are largely determined by state encoding. In this work, a new state encoding algorithm is proposed to reduce bit flips during state transitions within limited number of flip-flops. The proposed scheme, consisting of a transition graph model, an encoding graph conflict removing algorithm and a hardware efficient encoding algorithm, is able to reduce flip-flops used by 85% and reduce state transition bit flips by 41.1% compared with existing popular encoding solutions.
Abraham Mcllvaine, Chengmo Yang
VLSI-SoC3
2017 Path reuse-aware routing for non-volatile memory based FPGAs
Chengmo Yang
Integr.2
2017 ErasuCrypto: A Light-weight Secure Data Deletion Scheme for Solid State Drives
abstract
Abstract Securely deleting invalid data from secondary storage is critical to protect users’ data privacy against unauthorized accesses. However, secure deletion is very costly for solid state drives (SSDs), which unlike hard disks do not support in-place update. When applied to SSDs, both erasure-based and cryptography-based secure deletion methods inevitably incur large amount of valid data migrations and/or block erasures, which not only introduce extra latency and energy consumption, but also harm SSD lifetime. This paper proposes ErasuCrypto, a light-weight secure deletion framework with low block erasure and data migration overhead. ErasuCrypto integrates both erasurebased and encryption-based data deletion methods and flexibly selects the more cost-effective one to securely delete invalid data. We formulate a deletion cost minimization problem and give a greedy heuristic as the starting point. We further show that the problem can be reduced to a maximum-edge biclique finding problem, which can be effectively solved with existing heuristics. Experiments on real-world benchmarks show that ErasuCrypto can reduce the secure deletion cost of erasurebased scheme by 71% and the cost of cryptographybased scheme by 37%, while guaranteeing 100% security by deleting all the invalid data.
Chen Liu 0013, Hoda Aghaei Khouzani, Chengmo Yang
Proc. Priv. Enhancing Technol.3
2017 Segment and Conflict Aware Page Allocation and Migration in DRAM-PCM Hybrid Main Memory
abstract
Phase change memory (PCM), given its nonvolatility, potential high density, and low standby power, is a promising candidate to be used as main memory in next generation computer systems. However, to hide its shortcomings of limited endurance and slow write performance, state-of-the-art solutions tend to construct a dynamic RAM (DRAM)-PCM hybrid memory and place write-intensive pages in DRAM. While existing optimizations to this hybrid architecture focus on tuning DRAM configurations to reduce the number of write operations to PCM, this paper explores the interactions between DRAM and PCM to improve both the performance and the endurance of a DRAM-PCM hybrid main memory. Specifically, it exploits the flexibility of mapping virtual pages to physical pages, and develops a proactive strategy to allocate pages taking both program segments and DRAM conflict misses into consideration, thus distributing those heavily written pages across different DRAM sets. Meanwhile, a lifetime-aware DRAM replacement algorithm and a conflict-aware page remapping strategy are proposed to further reduce DRAM misses and PCM writes. Experiments confirm that the proposed techniques are able to improve average memory hit time and reduce maximum PCM write counts thus enhancing both performance and lifetime of a DRAM-PCM hybrid main memory.
Hoda Aghaei Khouzani, Fateme S. Hosseini, Chengmo Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 State Asymmetry Driven State Remapping in Phase Change Memory
abstract
Phase change memory (PCM) is one of the most promising candidates to replace DRAM as main memory in deep submicron regime. Regardless of single-level or multiple-level cells, the programming costs to each state exhibit significant asymmetries in latency, energy and endurance. In this paper, we exploit the potential of reducing programming costs in terms of latency, energy, and endurance for PCM through state remapping. First, quantitative programming models are constructed for cost assessments. Then, both dynamic and static remapping schemes are analyzed and compared. The observation that the efficacy of dynamic state remappings is instable motivates us to propose a static remapping technique, which outperforms previous work in cost reduction within much lower implementation overhead. The optimality of the proposed static state remapping is also proved. The evaluation results confirm the efficacy of the proposed state remapping technique in delivering a stable and promising cost reduction in latency, energy, and wear.
Mengying Zhao, Jingtong Hu, Chengmo Yang, Tiantian Liu 0001, Zhiping Jia, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2017 A DWM-Based Stack Architecture Implementation for Energy Harvesting Systems
abstract
Energy harvesting systems tend to use non-volatile processors to conduct computation under intermittent power supplies. While previous implementations of non-volatile processors are based on register architectures, stack architecture, known for its simplicity and small footprint, seems to be a better fit for energy harvesting systems. In this work, Domain Wall Memory (DWM) is used to implement ZPU, the world’s smallest working CPU. Not only does DWM offer ultra-high density and SRAM-comparable access latency, but the sequential access structure of DWM also makes it well suited for a stack whose accesses display high temporal locality. As the performance and energy of DWM are determined by the number of shift operations performed to access the stack, this paper further reduces shift operations through novel data placement and micro-code transformation optimizations. The impact of compiler optimization techniques on the number of shift operations is also investigated so as to select the most effective optimizations for DWM-based stack machine. Experimental studies confirm the effectiveness of the proposed DWM-based stack architectures in improving the performance and energy-efficiency of energy harvesting systems.
Hoda Aghaei Khouzani, Chengmo Yang
ACM Trans. Embed. Comput. Syst.2
2017 Exploiting Multiple Write Modes of Nonvolatile Main Memory in Embedded Systems
abstract
Existing Nonvolatile Memories (NVMs) have many attractive features to be the main memory of embedded systems. These features include low power, high density, and better scalability. Recently, Multilevel Cell (MLC) NVM has gained more and more popularity as it can provide a higher density than the traditional Single-Level Cell (SLC) NVM. However, there are also drawbacks in MLC NVM, namely, limited write endurance and expensive write operation. These two drawbacks have to be overcome before MLC NVM can be practically adopted as the main memory. In MLC Nonvolatile Main Memory (NVMM), two different types of write operations with very diverse data retention times are allowed. The first type maintains data for years but takes a longer time to write and is detrimental to the endurance. The second type maintains data for a short period but takes a shorter time to write. By observing that much of the data written to main memory is temporary and does not need to last long during the execution of a program, in this article, we propose novel task scheduling and write operation selection algorithms to improve MLC NVMM endurance and program efficiency. An Integer Linear Programming (ILP) formulation is first proposed to obtain optimal results. Since ILP takes exponential time to solve, we also propose the Multiwrite Mode-Aware Scheduling (MMAS) algorithm to achieve a near-optimal solution in polynomial time. Additionally, the Dynamical Memory Block Screening (DMS) algorithm is proposed to achieve wear leveling. The experimental results demonstrate that the proposed techniques can greatly improve the lifetime of the MLC NVMM as well as the efficiency of the program.
Mimi Xie, Chengmo Yang, Yiran Chen 0001, Jingtong Hu
ACM Trans. Embed. Comput. Syst.3
2016 A mutual auditing framework to protect IoT against hardware Trojans
abstract
Internet-of-Things (IoT), wherein sensor nodes of different types are used to monitor different objects, are expected to be used in many critical domains. However, hardware Trojans, which are malicious modifications implanted in individual nodes, may utilize the wireless connection facility to leak confidential information or to collude with each other to cause catastrophic failures in the IoT. To defend against these types of network-level threats, our goal is to develop a lightweight framework to monitor communications in the IoT. Instead of relying on a centralized data center to monitor the behavior of all the nodes, we propose to exploit vendor diversity among the nodes to build a distributed framework wherein nodes monitor the trustworthiness of their neighbors. This mutual auditing scheme is able to detect any attempt to leak information or collude with other malicious nodes, thus constructing trustworthy communication channels between untrustworthy nodes.
Chen Liu 0013, Patrick Cronin, Chengmo Yang
ASP-DAC3
2016 Routing path reuse maximization for efficient NV-FPGA reconfiguration
abstract
Non-Volatile memory-based FPGAs (NV-FPGAs) are expected to replace traditional SRAM-based FPGAs to achieve higher scalability and lower power consumption. Yet the slow write performance of NVMs not only challenges FPGA reconfiguration speed and overhead but also constrains the programming cycles of FPGAs. To efficiently configure switch boxes, the majority component of an FPGA, this paper proposes a routing path reuse technique. Technical contributions include a mathematical reconfiguration cost model of routing resources, a reuse-aware routing algorithm, as well as the incorporation of the proposed algorithm into the standard VTR CAD tool. Experiments on standard MCNC benchmarks show that the proposed scheme is able to achieve as much as 40% path reuse rate and reduce as much as 34.0% configuration cost for routing resources.
Patrick Cronin, Chengmo Yang, Jingtong Hu
ASP-DAC3
2016 Towards a Scalable and Write-Free Multi-version Checkpointing Scheme in Solid State Drives
abstract
Flash memory based solid state drives (SSDs) are widely adopted in mobile devices, PCs and data centers, making their reliability critical. Although periodically creating checkpoints is a well-developed technique for traditional hard disk drives, very few work has exploited the unique properties of SSDs to accelerate checkpoint creation and reduce storage cost. More specifically, the remap-on-write property of SSD creates a trail of multiple versions of data, which can be exploited to create multiple checkpoints without engendering extra writes. However, efficiently managing the metadata to support multiple checkpoints is challenging. In this paper, we propose a low storage cost and high performance scheme to support multiple checkpoints in SSDs. Instead of storing a log-based snapshot of the entire Flash Translation Layer (FTL) per checkpoint, we efficiently keep track of the changes to the FTL across multiple checkpoints, thus accelerating the creation, deletion, and activation of checkpoints within minimum storage overhead and minimum impact on regular SSD operations. Experiments on Microsoft real-world benchmarks confirm the advantage of the proposed scheme over a fully snapshot scheme in terms of storage and performance overhead.
Hoda Aghaei Khouzani, Chengmo Yang
DSN2
2016 Fully Exploiting PCM Write Capacity Within Near Zero Cost Through Segment-Based Page Allocation
abstract
Improving the endurance of phase change memory (PCM) is a fundamental issue when PCM technology is considered as an alternative to main memory usage. Existing wear-leveling techniques overcome this challenge through constantly remapping hot virtual pages, thus engendering a fair amount of extra write operations to PCM and imposing considerable performance and energy overhead. Our observation is that it is unnecessary to fully balance the accesses to different physical page frames during the execution of each process. Instead, since endurance is a lifetime factor, the hot virtual pages of different processes can be mapped to different physical pages in the PCM. Leveraging this property, we develop a wear-resistant page allocation algorithm, which exploits the diverse write characteristics of different program segments to improve PCM write endurance within almost no extra remapping cost in terms of energy and performance. The results of experiments conducted based on SPEC benchmarks show that the proposed technique can prolong PCM lifetime by hundreds of times within nearly zero searching and remapping overhead.
Hoda Aghaei Khouzani, Chengmo Yang
ACM J. Emerg. Technol. Comput. Syst.3
2015 Guiding fault-driven adaption in multicore systems through a reliability-aware static task schedule
abstract
Future multicore systems suffer from high and varying fault rates due to device scaling, increasing number of processing notes, varying environmental issues and aging effects. Efficient fault tolerant solutions capable of combining the advantages of static optimization and runtime adaptation are needed. To achieve this goal, we propose a static reliability-aware scheduling technique, aiming to guide runtime adaptation and relieve most of the computational overhead. The proposed static scheduler considers “reliability level” (RL) as an intermediate scheduling dimension and creates a “task-to-RL-to-core” mapping. This enables the “RL-to-core” mapping to be efficiently adapted at runtime according to fault rate variations, while the “task-to-RL” mapping can still be reused. Experimental studies show that by considering fault rates during static scheduling, runtime application execution time can be improved by up to 19% in a non-constant fault rate environment.
Laura A. Rozo Duque, Chengmo Yang
ASP-DAC2
2015 Improving performance and lifetime of DRAM-PCM hybrid main memory through a proactive page allocation strategy
abstract
Phase change memory (PCM), given its non-volatility and low static energy consumption, is a promising candidate to be used as main memory. However, due to its limited endurance and slow write performance, state-of-the-art solutions tend to construct a DRAM-PCM hybrid memory instead of using PCM exclusively. While existing optimizations to this hybrid architecture focus on tuning DRAM configurations to further reduce writes to PCM, we aim at developing a proactive solution. Specifically, we exploit the flexibility of mapping virtual pages to physical pages, and propose a page allocation algorithm that considers both segment information and conflict misses in DRAM to distribute heavily written pages across different DRAM sets. Trace-driven experiments confirm the effectiveness of proposed technique in reducing both DRAM misses and PCM writes, thus simultaneously improving performance and lifetime of DRAM-PCM hybrid main memory.
Hoda Aghaei Khouzani, Chengmo Yang, Jingtong Hu
ASP-DAC2
2015 Checkpoint-aware instruction scheduling for nonvolatile processor with multiple functional units
abstract
Embedded systems powered with harvested energy experience frequent execution interruption due to unstable energy source. Nonvolatile (NV) register based processor is proposed to realize fast resume after power failure. The states in the volatile registers are checkpointed to NV registers. However, frequent checkpointing causes performance degradation and consumes excessive power. In this paper, we propose the checkpoint aware instruction scheduling (CAIS) algorithm to reduce the writes to NV registers. Experiments show that CAIS can improve performance and reduce power consumption.
Mimi Xie, Jingtong Hu, Chengmo Yang, Yiran Chen 0001
ASP-DAC4
2015 Minimizing MLC PCM write energy for free through profiling-based state remapping
abstract
Phase change memory is becoming one of the most promising candidates to replace DRAM as main memory in deep sub-micron regime. Multi-level cell (MLC) PCM outperforms single level cell (SLC) PCM in terms of storage capacity but requires an iterative programming-and-verifying scheme to program cells to different resistance levels. The energy consumed in programming different MLC states varies significantly, thus motivating a state remapping technique to minimize the overall write energy. In this paper, we first compare dynamic and static state remapping strategies in terms of their efficacy in reducing energy, and then propose an effective and low-cost static state remapping algorithm. The experimental studies show 10.6% average (up to 16.9%) reduction in MLC PCM write energy, achieved within negligible hardware and performance overhead. Compared with the most related work, the proposed scheme saves more write energy on average, with near-zero performance, area and energy overhead.
Mengying Zhao, Chengmo Yang, Chun Jason Xue
ASP-DAC3
2015 Improving MPSoC reliability through adapting runtime task schedule based on time-correlated fault behavior
Laura A. Rozo Duque, José Monsalve Diaz, Chengmo Yang
DATE3
2015 Nonvolatile main memory aware garbage collection in high-level language virtual machine
abstract
Non-volatile memories (NVMs) such as Phase Change Memory (PCM) have been considered as promising candidates of next generation main memory for embedded systems due to their attractive features. These features include low power, high density, and better scalability. However, most existing NVMs suffer from two drawbacks, namely, limited write endurance and expensive write operation in terms of both time and energy. These problems are worsen when modern high-level languages employ virtual machine with garbage collector that generates a large amount of extra writes on non-volatile main memory. To tackle this challenge, this paper proposes three techniques: Living Objects Remapping (LORE), Dead Object Stamping (DOS), and Smart Wiping with Maximum Likelihood Estimation (SMILE) to reduce the unnecessary writes when garbage collector handles objects. The experimental results show that the proposed techniques not only significantly reduce the writes during each garbage collection cycle but also greatly improve the performance of virtual machine.
Mimi Xie, Chengmo Yang, Zili Shao, Jingtong Hu
EMSOFT3
2015 Fine-tuning CLB placement to speed up reconfigurations in NVM-based FPGAs
abstract
Non-volatile memories (NVMs) outperform traditional SRAMs in terms of low power consumption, high capacity, near-zero power-on delay, and high error-resistance. Researchers have demonstrated the possibilities of implementing FPGA building blocks with various types of NVMs. However, NVMs also bring several new design challenges to FPGAs: the slow write performance of NVM may degrade FPGA (re)configuration speed, while the limited write endurance of NVM constrains the number of times that the FPGA can be (re)configured. Unfortunately, none of these NVM features are taken into consideration in current FPGA synthesis tools, which have been optimized solely for SRAM-based FPGAs. To tackle this limitation, we propose to make the FPGA placement process aware of the slow and costly NVM writes. Our contributions are three-fold: We first construct mathematical models to characterize reconfiguration costs in NVM-based FPGAs. Second, we identify three types of flexibilities that can be exploited to reduce the reconfiguration cost. Finally, we present three approaches for designers to fine-tune the placement process to balance the reconfiguration cost and traditional timing and routability constraints according to their needs. The proposed algorithms are incorporated in Verilog-to-Routing (VTR) CAD tool. Experiments on standard MCNC benchmark circuits show that our approach eliminates up to 67% NVM writes during the reconfiguration process, thus effectively improving the performance and endurance of NVM-based FPGAs.
Patrick Cronin, Chengmo Yang, Jingtong Hu
FPL3
2015 Secure and Durable (SEDURA): An Integrated Encryption and Wear-leveling Framework for PCM-based Main Memory
abstract
Phase changing memory (PCM) is considered a promising candidate for next-generation main-memory. Despite its advantages of lower power and high density, PCM faces critical security challenges due to its non-volatility: data are still accessible by the attacker even if the device is detached from a power supply. While encryption has been widely adopted as the solution to protect data, it not only creates additional performance and energy overhead during data encryption/decryption, but also hurts PCM lifetime by introducing more writes to PCM cells.
Chen Liu 0013, Chengmo Yang
LCTES2
2015 Non-volatile memories in FPGAs: Exploiting logic similarity to accelerate reconfiguration and increase programming cycles
abstract
Non-volatile memory (NVM) technologies have been known for their advantages of large capacity, low energy consumption, high error-resistance, and near-zero power-on delay. It is expected that they will replace traditional SRAM as FPGA reconfigurable blocks. While NVMs promise FPGAs with more reconfigurable resources, lower power consumption, and higher resilience to power interruptions, they also impose two new design challenges: the slow write performance of NVMs may degrade FPGA reconfiguration speed, while their limited write endurance constrains FPGA programming cycles. To overcome these challenges, we propose a similarity driven approach to reduce reconfiguration cost in NVM-based FPGAs. When synthesizing a new design, its similarity to the design currently on the FPGA is characterized by taking both LUT contents and CLB-level topology into consideration. The reconfiguration cost minimization problem is formulated as a bipartite graph matching problem and solved optimally. Experiments on standard circuit benchmarks show that the proposed algorithms eliminate more than 57.4% of NVM writes during the reconfiguration process, thus effectively improving performance and endurance of NVM-based FPGAs.
Patrick Cronin, Chengmo Yang, Jingtong Hu
VLSI-SoC3
2015 Qualifying non-volatile register files for embedded systems through compiler-directed write minimization and balancing
abstract
Recent research shows non-volatile flip-flops can be attached to register files to hold computation state thus enabling fast recovery upon power failure. However, the endurance limitation of NVM cells challenges their usage for holding register values that are frequently updated during program execution. To extend the lifetime of non-volatile register files, we propose two compiler-directed optimizations. First, through analyzing the register access patterns in frequently executed loops, a minimum set of registers is identified to be periodically written to NVM cells, thus minimizing the total number of writes to the NVM register file. Meanwhile, the register mapping is also adjusted to enable an efficient dynamic register rotation to further balance the writes to different NVM registers. Experimental studies show that the proposed two techniques can significantly extend the lifetime of non-volatile registers, thus qualifying them for various embedded systems.
Chengmo Yang, Maria Ruiz Varela
VLSI-SoC1
2014 Exploiting heterogeneity in MPSoCs to prevent potential trojan propagation across malicious IPs
abstract
Multiprocessor System-on-Chip (MPSoC) platforms face some of the most demanding security concerns, as they process, store, and communicate sensitive information using third-party intellectual property (3PIP) cores. The trend of outsourcing design and fabrication strongly questions the assumption of 3PIP components being trustworthy. While existing research focuses on addressing hardware trojans in individual IPs, this paper improves MPSoC security from another perspective. Specifically, our goal is to prevent trojans in malicious IPs from triggering each other and leading to severe system-wide degradation in security and reliability. We propose to impose trojan isolation constraints during static task scheduling, ensuring that all legal communications on the target MPSoC are between IPs of different types. This in turn enables the runtime system to monitor and detect undesired communication paths, if any. We furthermore pose the security-constrained MPSoC task scheduling as a multi-dimensional optimization problem, and solve it through Integer Linear Programming (ILP), thus minimizing the associated performance, power, and hardware overhead. The results show that trojan isolation can be achieved within one extra vendor and nearly no performance overhead.
Chen Liu 0013, Chengmo Yang
ACM Great Lakes Symposium on VLSI2
2014 Improving multilevel PCM reliability through age-aware reading and writing strategies
abstract
Given its low power consumption and high density, Phase Change Memory (PCM) has been treated as a promising alternative to DRAM for main memory storage. Multilevel Cell (MLC) PCM outperforms regular single level cell (SLC) PCM with even higher information density, yet requires more accurate control for cell reading and writing. More crucially, the resistance of a MLC PCM cell may drift over time, thus introducing high error rate during cell reading if the quantization thresholds are constant. While previous work tries to adjust the quantization method when reading a cell, its accuracy is still limited due to the inter-cell variations in data age. In this paper, we propose various PCM writing and quantization strategies to improve MLC PCM reliability. Cell quantization accuracy is improved by taking into consideration not only the time information but also inter-cell age variations. Moreover, by making the write strategy be aware of time, the inter-level quantization margin can be guaranteed when the cells exhibit large age variations. This time-aware writing scheme is adaptively applied to maximize achievable benefits. The experimental results show that the proposed writing and reading approaches can effectively reduce the quantization error rate in MLC PCM by 95%.
Chen Liu 0013, Chengmo Yang
ICCD2
2014 Leveling to the last mile: Near-zero-cost bit level wear leveling for PCM-based main memory
abstract
Phase change memory (PCM) has demonstrated great potential as an alternative of DRAM to serve as main memory due to its favorable characteristics of non-volatility, scalability and near-zero leakage power. However, the comparatively poor endurance of PCM largely limits its adoption. Wear leveling strategies targeting to even write distributions have been proposed at different granularities and on various memory hierarchies for PCM endurance enhancement. Write operations are distributed across the memory through migrating data from heavily written locations to less burdened ones, which is usually guided by counters recording the number of writes. However, evenly distributing writes at a coarse granularity cannot deliver the best endurance results as write distributions are highly imbalanced even at the bit level. In this work, we propose a near-zero-cost bit-level wear leveling strategy to improve PCM endurance. The proposed technique can be combined with various coarse-grained wear leveling strategies. Experiment results show 102% endurance enhancement on average, which is 34% higher than the most related work, with significantly lower storage, performance and energy overheads.
Mengying Zhao, Liang Shi 0001, Chengmo Yang, Chun Jason Xue
ICCD3
2014 Prolonging PCM lifetime through energy-efficient, segment-aware, and wear-resistant page allocation
abstract
Improving the endurance of Phase change memory (PCM) is a fundamental issue when the technology is considered as an alternative to main memory usage. Existing wear-leveling techniques overcome this challenge through constantly remapping hot virtual pages, engendering a fair amount of extra write operations to PCM and imposing considerable energy overhead. Our observation is that it is unnecessary to fully balance the accesses to different physical pages during the execution of each process. Instead, since endurance is a lifetime factor, the hot virtual pages of different processes can be mapped to different physical pages in the PCM. Leveraging this property, we develop a wear-resistant page allocation algorithm, which exploits the diverse write characteristics of different program segments to improve PCM write endurance within almost no extra remapping cost. Experimental results show that the proposed technique can prolong PCM lifetime by hundreds of times within nearly zero searching and remapping overhead.
Hoda Aghaei Khouzani, Chengmo Yang, Archana Pandurangi
ISLPED3
2014 A Unified Write Buffer Cache Management Scheme for Flash Memory
abstract
NAND flash memory has been widely adopted in embedded systems as secondary storage. However, the further development of flash memory strongly hinges on the tackling of its inherent implausible characteristics, including read-and-write speed asymmetry, inability of in-place updates, and performance-harmful erase operations. While write buffer cache (WBC) has been proposed to enhance the performance of write operations, the development of a unified WBC management scheme that is effective for diverse types of access patterns is still a challenging task. In this paper, a novel WBC management scheme named expectation-based least recently used (ExLRU) is proposed to improve the performance of flash memory through effectively reducing the number of erase operations and write activities. Different from the previous works, ExLRU accurately maintains access history information in the WBC, based on which a novel cost model is constructed to select data with the minimum write cost to write to flash memory. An efficient ExLRU implementation with negligible overhead is developed. Simulation results show that ExLRU outperforms state-of-the-art WBC management schemes under various workloads.
Liang Shi 0001, Jianhua Li 0003, Qing'an Li, Chun Jason Xue, Chengmo Yang, Xuehai Zhou
IEEE Trans. Very Large Scale Integr. Syst.5
2013 Fault detection and recovery efficiency co-optimization through compile-time analysis and runtime adaptation
abstract
The ever scaling-down feature size and noise margin keep elevating hardware failure rates, requiring the incorporation of fault tolerance into computer systems. One fault tolerance scheme that receives a lot of research attention is redundant execution. However, existing solutions are developed under the assumption that the fault rate is low. These techniques either solely focus on fault detection, or sometimes even increase recovery cost to reduce fault detection overhead. The lack of overall efficiency makes them insufficient and inappropriate for embedded systems with tight energy and cost budget. Our study shows that checkpoint frequency and fault rate are two critical parameters determining the overall fault detection and recovery overhead. To co-optimize detection and recovery, we statically construct a mathematical model, capable of taking application and architecture characteristics into consideration and identifying the optimal checkpoint frequency of an application for a given fault rate. Moreover, as the fault rate is infeasible to predict a priori, we furthermore propose a set of heuristics, enabling the system to dynamically monitor the fault rate and adapt the checkpoint frequency accordingly. The efficacy of the static and the adaptive optimizations is evaluated through detailed instructionlevel simulation. The results show that the optimal checkpoint frequency identified by the static model is very close to the actual value (6% deviation) and the run-time adaptation scheme effectively reduces the overhead caused by the unpredictability in fault rate.
Chengmo Yang
CASES2
2013 Boosting efficiency of fault detection and recovery throughapplication-specific comparison and checkpointing
abstract
While the unending technology scaling has brought reliability to the forefront of concerns of semiconductor industry, fault tolerance techniques are still rarely incorporated into existing designs due to their high overhead. One fault tolerance scheme that receives a lot of research attention is duplication and checkpointing. However, most of the techniques in the category employ a blind strategy to compare instruction results, therefore not only generating large overhead in buffering and verifying these values, but also inducing unnecessary rollbacks to recover faults that will never influence subsequent execution. To tackle these issues, we introduce in this paper an approach that identifies the minimum set of instruction results for fault detection and checkpointing. For a given application, the proposed technique first identifies the control and data flow information of each execution hotspot, and then selects only the instruction results that either influence the final program results or are needed during re-execution as the comparison set. Our experimental studies demonstrate that the proposed hotspot-targeting technique is able to reduce nearly 88% of the comparison overhead and mask over 38% of the total injected faults of all the injected faults while at the same time delivering full fault coverage.
Chengmo Yang
LCTES2
2012 Write-activity-aware page table management for PCM-based embedded systems
abstract
Due to its low power consumption and high density, phase change memory (PCM) becomes a promising main-memory alternative to DRAM in embedded systems. PCM, however, has the endurance problem in which the number of rewrites to each cell is quite limited compared with DRAM. Therefore, it is fundamental to eliminate unnecessary writes in PCM-based embedded systems. This paper presents a simple yet effective scheme to solve this problem, through redesigning existing software to exploit write-activity-aware features provided by underlying hardware. Particularly, we target at page table management, a key kernel component residing in the memory management part of the Linux kernel. We present for the first time a write-activity-aware page table management scheme, WAPTM, accomplished through two modifications to the page table initialization and page frame allocation process. The scheme has been implemented in Google Android 2.3 based on ARM architecture and evaluated with real applications on the Android emulator. The experimental results show that the proposed scheme can significantly reduce write activities to page tables in the new kernel compared with the original Android. We hope this work can serve as a first step towards the design of write-activity-aware operating systems via simple and feasible modifications.
Tianzheng Wang 0001, Duo Liu 0002, Zili Shao, Chengmo Yang
ASP-DAC4
2012 Tackling Resource Variations Through Adaptive Multicore Execution Frameworks
abstract
Multicore architectures have been widely adopted to accommodate the rising performance demand in various application domains, ranging from high-end supercomputing to low-end consumer electronics. Yet due to the ever growing integration density and application complexity, such architectures suffer from increased level of core availability variations. At runtime, issues such as device failures, heat buildup, as well as resource competitions and preemptions can make computational resources unavailable, necessitating execution schedules capable of delivering diverse performance levels to match the varying resource allocations. The adaptive execution framework introduced in this paper delivers high-quality schedules capable of predictably reconfiguring execution and gracefully degrading performance in the face of resource unavailability. By adhering to a novel band structure, a set of possible execution schedules are compactly engendered in readiness at compile time, thus delivering predictable responses to runtime resource variations. More importantly, through the exploitation of an extra degree of freedom in the scheduling process, the scheduler can perform task assignments in such a way that adaptivity can be embedded within the preoptimized schedules at almost no cost. The efficacy of the proposed technique is confirmed by incorporating it into a conventional, widely adopted scheduling heuristic and experimentally verifying it in the context of single core degradations.
Chengmo Yang, Alex Orailoglu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 Frugal but flexible multicore topologies in support of resource variation-driven adaptivity
abstract
Given the projected higher variations in the availability of computational resources, adaptive static schedules have been developed to attain high-speed execution reconfiguration with no reliance on any runtime rescheduling decisions. These schedules are able to deliver predictable execution despite the increased levels of device unreliability in future multicore systems. Yet the associated runtime reconfiguration overhead is largely determined by the underlying system topology. Fully connected architectures, although they can effectively hide the overhead in execution migration, become infeasible as the core count grows to hundreds in the near future. We exploit in this paper the high locality associated with adaptive static schedules, and outline a scalable and locally shareable system organization for multicore platforms. With the incorporation of a limited set of neighborhood-centered communication links, threads are allowed to be directly migrated among adjacent cores without physical data movement. At the architecture level, a set of 2-dimensional physical topologies with such a local sharing property embedded is furthermore proposed. The inherent regularity allows these topologies to be adopted as a fixed-silicon multicore platform that can be flexibly redefined according to the parallelism characteristics and resilience needs of each application.
Chengmo Yang, Alex Orailoglu
DATE1
2011 ExLRU: a unified write buffer cache management for flash memory
abstract
NAND flash memory has been widely adopted in embedded systems as secondary storage. Yet the further development of flash memory strongly hinges on the tackling of its inherent implausible characteristics, including read and write speed asymmetry, inability of in-place update, and performance harmful erase operations. While Write Buffer Cache (WBC) has been proposed to enhance the performance of write operations, the development of a unified WBC management scheme that is effective for diverse types of access patterns is still a challenging task. In this paper, a novel WBC management scheme named Expectation-based LRU (ExLRU) is proposed to improve the performance of write operations while at the same time reducing the number of erase operations on flash memory. ExLRU accurately maintains access history information in WBC, based on which a new cost model is constructed to select the data with minimum write cost to be written to flash memory. An efficient ExLRU implementation with negligible hardware overhead is further developed. Simulation results show that ExLRU outperforms state-of-art WBC management schemes under various workloads.
Liang Shi 0001, Jianhua Li 0003, Chun Jason Xue, Chengmo Yang, Xuehai Zhou
EMSOFT4
2011 Migration-aware adaptive MPSoC static schedules with dynamic reconfigurability
Chun Jason Xue, Chengmo Yang, Alex Orailoglu
J. Parallel Distributed Comput.3
2011 Full Fault Resilience and Relaxed Synchronization Requirements at the Cache-Memory Interface
abstract
While multicore platforms promise significant speedup for many current applications, they also suffer from increased reliability problems as a result of ever scaling device size. The projected elevation in fault rate, together with the diverse behavior of fault manifestation, argues for highly efficient solutions of full fault resilience. Traditional duplication and checkpointing strategies typically impose sizable overhead in checkpointing execution results, or in constantly synchronizing two threads for value checking. To reduce such overhead while at the same time delivering full fault resilience, we propose an integrated fault detection and checkpointing framework, wherein the comparison and checkpointing process is performed at the cache-memory interface. By sharing a single cache between two duplicated threads, execution results can be directly verified in the cache before being written back, thus strictly protecting the memory against execution faults. Meanwhile, as unconfirmed data are allowed to be written into the cache, one thread can run well ahead of the other, thus relaxing the straightjacket of the strict execution synchronization model. If a cache block is constantly updated, further synchronization relaxation can be achieved through extending the cache design to duplicate a cache block and skip the comparison of the intermediate values.
Chengmo Yang, Alex Orailoglu
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Fully adaptive multicore architectures through statically-directed dynamic execution reconfigurations
abstract
As a result of ever growing integration density and application complexity, future multicore architectures will suffer from increased levels of core availability variations. Full resource utilization in the face of various levels of resource availability necessitates techniques that compactly engender numerous schedules in readiness at compile time. Such schedules, each of which can make maximum utilization of the available resources, can be adaptively applied at runtime, thus enabling a toleration of up to an arbitrarily large amount of resource variations. Core binding permutations furthermore minimize the performance impact imposed by adaptivity on the pre-reconfiguration schedules while retaining all the concomitant benefits. The efficacy of the proposed technique is confirmed by incorporating it into a conventional, widely adopted scheduling heuristic and experimentally verifying it in the context of multiple core deallocations. This paper thus offers critical improvement over prior state-of-the-art, which targets solely single core failures, a subset of resource variation modes in future nanoscale MPSoCs which are projected to display elevated device failures, heat buildup, resource competition and preemptions.
Chengmo Yang, Alex Orailoglu
VLSI-SoC1
2010 Fine-grained adaptive CMP cache sharing through access history exploitation
abstract
Advances in semiconductor technologies have enabled the integration of multiple processor cores as well as varying sizes of L1 and L2 caches on a single chip. The ever growing complexity and diversity of the associated workloads impose a crucial challenge on the organization and management of the on-chip cache resources. As each core generates a varying amount of accesses to each cache line during execution, sharing a single L2 cache among all the cores can minimize off-chip misses. However, each access to a shared L2 cache imposes significant performance and power overhead, as the tags of all the blocks on a cache line need to be compared in parallel. To efficiently utilize cache resources while saving power, we present in this paper a fine-grained L2 cache management technique with minimum hardware overhead. Each core is allowed to set an ownership bit in an L2 cache block to directly signify the necessity of tag checking, thus reducing the latency and power consumption of each cache access. Joint block ownership approaches provide shareability, thus precluding costly data replication from which private L2 caches typically suffer. Meanwhile, through monitoring line-based access histories, a core that produces a large amount of misses is precluded from replacing blocks belonging to other cores, thus efficiently attaining fine-grained cache partitioning. Experimental results confirm that the proposed technique can effectively reduce the access latency and power consumption of traditional shared L2 caches, accompanied by additionally a slight reduction in the miss rate.
Chengmo Yang, Chun Jason Xue, Alex Orailoglu
VLSI-SoC1
2009 Towards no-cost adaptive MPSoC static schedules through exploitation of logical-to-physical core mapping latitude
abstract
The computing engines of many current applications are powered by MPSoC platforms, which promise significant speedup but induce increased reliability problems as a result of ever growing integration density and chip size. While static MPSoC execution schedules deliver predictable worst-case performance, the absence of dynamic variability unfortunately constrains their usefulness in such an unreliable execution environment. Adaptive static schedules with predictable responses to run-time resource variations have consequently been proposed, yet the extra constraints imposed by adaptivity on task assignment have resulted in schedule length increases. We propose to eradicate the associated performance degradation of such techniques while retaining all the concomitant benefits, by exploiting an inherent degree of freedom in task assignment regarding the logical to physical core mapping. The proposed technique relies on the use of core reordering and rotation through utilizing a graph representation model, which enables a direction translation of inter-core communication paths into order requirements between cores. The algorithmic implementation results confirm that the proposed technique can drastically reduce the schedule length overhead of both pre- and post-reconfiguration schedules.
Chengmo Yang, Alex Orailoglu
DATE1
2009 Processor reliability enhancement through compiler-directed register file peak temperature reduction
abstract
Each semiconductor technology generation brings us closer to the imminent processor architecture heat wall, with all its associated adverse effects on system performance and reliability. Temperature hotspots not only accelerate the physical failure mechanisms such as electromigration and dielectric breakdown, but furthermore make the system more vulnerable to timing-related intermittent failures. Traditional thermal management techniques suffer from considerable performance overhead as the entire processor needs to be stalled or slowed down to preclude heat accumulation. Given the significant temporal and spatial variations of the chip-wide temperature, we propose in this paper a technique that directly targets one of the resources that is most likely to overheat in current processors, namely, the register files. Instead of duplicating or physically distributing the register file, we suggest to attain power density control through exploiting the extant spatial slack associated with register file accesses. Based on application-specific access profiles, a compiler-directed register shuffling strategy is proposed to deterministically construct the logical to physical register mapping in a rotating manner. Simulation results confirm that the proposed technique attains, within a limited hardware budget and negligible performance degradation, effective reduction in peak temperature and hence in the expected fault rates for the entire chip.
Chengmo Yang, Alex Orailoglu
DSN1
2008 A light-weight cache-based fault detection and checkpointing scheme for MPSoCs enabling relaxed execution synchronization
abstract
While technology advances have made MPSoCs a standard architecture for embedded systems, their applicability is increasingly being challenged by dramatic increases in the amount of device failures that may occur during execution. Conventional fault tolerance techniques employ a duplication-and-comparison strategy to detect arbitrary execution faults, as well as a checkpointing-and-rollback strategy to recover from the faulty state. Comparison and checkpointing are performed either at task level, thus imposing a large amount of overhead in verifying and backing up memory pages, or at instruction level, thus necessitating a lock-step execution model which significantly limits the attainable performance. To overcome the shortcomings of both strategies, in this paper we propose a cache-based fault tolerance scheme wherein the comparison and checkpointing process is performed at the cache-memory interface. By allowing two processors that execute duplicated tasks to share a single data cache, the proposed scheme is able to verify execution results before writing them back into memory, thus protecting the memory from being polluted by execution faults. This in turn significantly reduces the checkpointing overhead. Meanwhile, since only the data written into memory are compared, the strict instruction-by-instruction synchronization model used in multithreading processors can be relaxed. The simulation results confirm that the proposed scheme only imposes a performance overhead ranging from 1.4% to 10.4%, while both fault detection and execution checkpointing can be effectively attained.
Chengmo Yang, Alex Orailoglu
CASES1
2007 Light-weight synchronization for inter-processor communication acceleration on embedded MPSoCs
abstract
The advances in semiconductor technologies have placed MPSoCscenter stage as a standard architecture for embedded applications of ever increasing complexity. Efficient utilization of the ample hardware resources requires applications to be decomposed into fine-grained threads, engendering in turn a large amount of interprocessor communications. While fine-grained on-chip interconnects can reduce the data transfer overhead, the traditional synchronization mechanisms, such as spin locks and barriers, still cause significant contention in polling shared variables. To overcome this issue, in this paper we propose a light-weight distributed synchronization mechanism which statically encodes the semantically correct order of accesses to each shared variable. A sharp reduction in the number of code bits is attained through a reference coloring algorithm, which furthermore enables an implementation within negligible hardware overhead. This light-weight synchronization mechanism allows dependent threads to frequently exchange data during execution, in turn enabling the exploration of fine-grained parallelism for applications with complex dependences.
Chengmo Yang, Alex Orailoglu
CASES1
2006 Power-efficient instruction delivery through trace reuse
abstract
As power dissipation inexorably becomes the major bottleneck in system integration and reliability, the front-end instruction delivery path in a traditional out-of-order superscalar processor needs to deliver high application performance in an energy-effective manner. This challenge can be addressed by efficiently reusing the work of fetch and decode performed during preceding loop iterations and resident mostly within the processor itself. As a large percentage of the instructions currently under fetch have previously dispatched copies resident in the Reorder Buffer (ROB), in this paper we develop a mechanism to utilize the ROB as a storage location for previously decoded instructions. Thus instructions can be fed directly from the ROB into the rename and issue stages, enabling the gating off of the fetch and decode logic for large periods of time so as to deliver significant power savings. Power and performance criticality of the ROB requires an efficient reuse identification mechanism; we outline such a cost-efficient Reuse Identification Unit (RIU) which enables effective identification of the matches between the ROB entries and the instructions currently under fetch. Simulation results on both multimedia and SPEC 2000 benchmarks confirm that incorporating the proposed technique on traditional out-of-order superscalar processors results in not only a sight improvement in performance, but also significant savings in the overall system power dissipation, achieved within a limited hardware budget.
Chengmo Yang, Alex Orailoglu
PACT1
2006 Power efficient branch prediction through early identification of branch addresses
abstract
Ever increasing performance requirements have elevated deeply pipelined architectures to a standard even in the embedded processor domain, requiring the incorporation of dynamic branch prediction subsystems to hide the execution latency of control-altering instructions. In this paper a low power early branch identification technique which enables the design of extremely power-efficient branch predictors and BTBs is proposed. Through static extraction of program information regarding the distance to subsequent branches, this technique enables the calculation of the next branch address as soon as the direction of the current branch has been predicted. Early identification of branch addresses enables a complete elimination of the power hungry BTB lookups normally occurring at every execution cycle, as well as a just-in-time wake-up mechanism when accessing "hibernating" entries in complex predictors, switched to power-saving mode to reduce leakage power dissipation. A cost-efficient Branch Identification Unit (BIU) to calculate branch addresses is presented and analyzed in terms of power and timing characteristics. The effectiveness of the proposed BTB access policy and predictor wake-up mechanism is also confirmed by the simulation results of the SPECint 2000 and Media-bench benchmarks.
Chengmo Yang, Alex Orailoglu
CASES1