VLDB 2026 Research / reviewers in the wild / expert
R. Iris Bahar
dblp:68/3828 · also Iris Bahar, Ruth Iris Bahar
· DBLP profile ↗
97ranked-venue papers
18as first author
8since 2021 · last 2026
0000-0001-6927-8527ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 95 · 17 first-author · 7 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing and Optimizing Cache Placement for Secure Memory MetadataabstractTo mitigate well-studied memory vulnerabilities [14], [21], [24], [25], memory devices may implement secure memory [9], [10], [27], [29], [34], an extension to the memory controller logic that guarantees the confidentiality and integrity of data stored in main memory. The memory encryption engine (MEE) is responsible for decrypting and integrity verifying data that comes from the off-chip memory device [10]. To protect outgoing data, the MEE encrypts and protects the integrity of data in 64B-block granularity. To provide confidentiality, the MEE uses counter-mode encryption (CME) [22], [23], [34]. To verify data’s integrity, the MEE maintains a per-block hashed message authentication code (HMAC) [7], [16] in memory. Additionally, Bonsai Merkle Trees (BMT) are used [27] to protect the encryption counter from replay attacks. The root of the BMT is stored on-chip in trusted hardware and serves as a root of trust for authentications. When fetching data from memory, the requisite metadata (i.e., encryption counter, HMAC, and path through the integrity tree) are also fetched to decrypt the data and authenticate its state against the HMAC and trusted root. This ensures that the processor does not perform computation on any corrupted data and that data is private while off-chip. Blake Cragen, R. Iris Bahar, Tamara Silbergleit Lehman |
ISPASS | 3 |
| 2025 | Robot Planning Under Uncertainty for Object Assembly and Troubleshooting Using Human Causal ModelsabstractIn this paper we explore if human mental models of objects, even when flawed, can be integrated with a collaborative robot's decision making framework to allow it to make smarter choices under partial observability for different object-related tasks such as assembly and troubleshooting. We demonstrate how (1) these informative causal models can be extracted from humans through crowdsourcing, (2) object assembly and troubleshooting can be formulated as Partially Observable Markov Decision Processes (POMDPs) and (3) our extracted causal models can be incorporated into those models in the form of approximate priors. Finally, (4) we use systematic experimentation in simulation to demonstrate the success of this approach, with 2 X average improvement in reward observed for object assembly tasks, and 1.4 X average improvement in reward observed for troubleshooting tasks. Semanti Basu, Semir Tatlidil, Tiffany Tran, Serena Saxena, Tom Williams 0001, Steven A. Sloman, R. Iris Bahar |
ICRA | 8 |
| 2024 | A Midsummer Night's Tree: Efficient and High Performance Secure SCMabstractSecure memory is a highly desirable property to prevent memory corruption-based attacks. The emergence of nonvolatile, storage class memory (SCM) devices presents new challenges for secure memory. Metadata for integrity verification, organized in a Bonsai Merkle Tree (BMT), is cached on-chip in volatile caches, and may be lost on a power failure. As a consequence, care is required to ensure that metadata updates are always propagated into SCM. To optimize metadata updates, state-of-the-art approaches propose lazy update crash consistent metadata schemes. However, few consider the implications of their optimizations on on-chip area, which leads to inefficient utilization of scarce on-chip space. In this paper, we propose A Midsummer Night's Tree (AMNT), a novel "tree within a tree" approach to provide crash consistent integrity with low run-time overhead while limiting on-chip area for security metadata. Our approach offloads the potential hardware complexity of our technique to software to keep area overheads low. Our proposed mechanism results in significant improvements (a 41% reduction in execution overhead on average versus the state-of-the-art) for in-memory storage applications while significantly reducing the required on-chip area to implement our protocol. Kidus Workneh, Jac McCarty, Joseph Izraelevitz, Tamara Silbergleit Lehman, R. Iris Bahar |
ASPLOS (3) | 6 |
| 2024 | Invited: Using Causal Information to Enable More Efficient Robot OperationabstractCausal reasoning is a key factor that allows human beings to effortlessly draw parallels between two seemingly different tasks. Given access to a limited number of human-generated causal models, if robots could be taught to perform a similar "transfer", then they could potentially generalize to different domains using limited human intervention. The goal is to effectively generate human-like causal models for an unknown task by learning from a small dataset of human-generated causal models on other related tasks. In this paper we propose two different methods to leverage causal models obtained from humans to generalize to a new task. We ran a user study to obtain causal models on how different light-producing objects function. We then explore one-to-one transfer, where we directly use a causal model generated by a participant to infer the causal model of another object. We also explore many-to-one transfer, where we leverage causal models from different objects and users to infer the causal model of an unknown object. Automated transfer of causal models from humans to unrelated domains has the potential to replicate human-like reasoning in an unknown scenario in the absence of an expert. That chain of reasoning represented through causal models can then be used by autonomous agents to guide several downstream tasks such as object assembly, trouble shooting, and repair. Semanti Basu, Semir Tatlidil, Steven A. Sloman, R. Iris Bahar |
DAC | 5 |
| 2022 | Hardware Acceleration of Nonparametric Belief Propagation for Efficient Robot ManipulationabstractProbabilistic graphical models (PGMs) have been widely used in computer vision, robotics, and statistics. Generative inference algorithms used to solve PGMs, such as belief propagation (BP), involve integration over high-dimensional variables, which becomes computationally infeasible. Drawing inspiration from particle filters, nonparametric belief propagation (NBP) combines the efficiency of Monte-Carlo sampling to capture the belief space of hidden variables while preserving accuracy and algorithm robustness. In this poster presentation, we describe a novel method to accelerate an NBP algorithm in hardware and evaluate our approach on the 6 degree-of-freedom (DoF) articulated object pose estimation problem. Our two major contributions include 1) identifying three independent computation flows in the algorithm to effectively overlap the two main steps of the algorithm: belief update and message update, and 2) creating deeply pipelined processing units that allow for fine-grained parallelism, which helps to balance the workloads on different computation streams. Results from this study demonstrate that our design achieves both improved runtime and energy efficiency. In particular, we achieved 26X energy saving compared to running the algorithm on a Titan Xp GPU, and 10X runtime speedup and 14X energy saving compared to running on the Jetson AGX . We believe that our FPGA implementation can greatly improve particle-based sampling methods for real time applications. Yanqi Liu, Anthony Opipari, Théo Guérin, R. Iris Bahar |
FPGA | 4 |
| 2022 | A Reconfigurable Hardware Library for Robot Scene PerceptionabstractPerceiving the position and orientation of objects (i.e., pose estimation) is a crucial prerequisite for robots acting within their natural environment. We present a hardware acceleration approach to enable real-time and energy efficient articulated pose estimation for robots operating in unstructured environments. Our hardware accelerator implements Nonparametric Belief Propagation (NBP) to infer the belief distribution of articulated object poses. Our approach is on average, 26× more energy efficient than a high-end GPU and 11× faster than an embedded low-power GPU implementation. Moreover, we present a Monte-Carlo Perception Library generated from high-level synthesis to enable reconfigurable hardware designs on FPGA fabrics that are better tuned to user-specified scene, resource, and performance constraints. Yanqi Liu, Anthony Opipari, Odest Chadwicke Jenkins, R. Iris Bahar |
ICCAD | 4 |
| 2022 | HybriDS: Cache-Conscious Concurrent Data Structures for Near-Memory Processing ArchitecturesabstractIn recent years, the ever-increasing impact of memory access bottlenecks has brought forth a renewed interest in near-memory processing (NMP) architectures. In this work, we propose and empirically evaluate hybrid data structures, which are concurrent data structures custom-designed for these new NMP architectures. We focus on cache-optimized data structures, such as skiplists and B+ trees, that are often used as index structures in online transaction processing (OLTP) systems to enable fast key-based lookups. These data structures are hierarchical, where lookups begin at a small number of top-level nodes and diverge to many different node paths as they move down the hierarchy, such that nodes in higher levels benefit more from caching. Our proposed hybrid data structures split traditional hierarchical data structures into a host-managed portion consisting of higher-level nodes and an NMP-managed portion consisting of the remaining lower-level nodes, thus retaining and further enhancing the cache-conscious optimizations of their conventional implementations. Although the idea might seem relatively simple, the splitting of the data structure prompts new synchronization problems, and careful implementation is required to ensure high concurrency and correctness. We provide implementations of a hybrid skiplist and a hybrid B+ tree, and we empirically evaluate them on a cycle-accurate full-system architecture simulator. Our results show that the hybrid data structures have the potential to improve performance by more than 2x compared to state-of-the-art concurrent data structures. Jiwon Choe, Andrew Crotty, Tali Moreshet, Maurice Herlihy, R. Iris Bahar |
SPAA | 5 |
| 2021 | Low Power Shift and Capture through ATPG-Configured Embedded Enable Capture BitsabstractExcessive test power can cause multiple issues at manufacturing as well as during field test. To reduce both shift and capture power during test, we propose a DFT-based approach where we split the scan chains into segments and use extra control bits inserted between the segments to determine whether a particular segment will capture. A significant advantage of this approach is that a standard ATPG tool is capable of automatically generating the appropriate values for the control bits in the test patterns. This is true not only for stuck-at fault test sets, but for Launch-off-Capture (LOC) transition tests as well. It eliminates the need for expensive post processing or modification of the ATPG tool. Up to 37% power reduction can be achieved for a stuck-at test set while up to 35% reduction can be achieved for a transition test set for the circuits studied. Lakshmi Ramakrishnan, Jennifer Dworak, Kundan Nepal, Theodore W. Manikas, R. Iris Bahar |
ITC | 7 |
| 2020 | Hardware Acceleration of Monte-Carlo Sampling for Energy Efficient Robust Robot ManipulationabstractAlgorithms based on Monte-Carlo sampling have been widely adapted in robotics and other areas of engineering due to their performance robustness. However, these sampling-based approaches have high computational requirements, making them unsuitable for real-time applications with tight energy constraints. In this paper, we investigate 6 degree-of-freedom (6DoF) pose estimation for robot manipulation using this method, which uses rendering combined with sequential Monte-Carlo sampling. While potentially very accurate, the significant computational complexity of the algorithm makes it less attractive for mobile robots, where runtime and energy consumption are tightly constrained. To address these challenges, we develop a novel hardware implementation of Monte-Carlo sampling on an FPGA with lower computational complexity and memory usage, while achieving high parallelism and modularization. Our results show 12X-21X improvements in energy efficiency over low-power and high-end GPU implementations, respectively. Moreover, we achieve real time performance without compromising accuracy. Yanqi Liu, Giuseppe Calderoni, R. Iris Bahar |
FPL | 3 |
| 2020 | Hardware Acceleration of Robot Scene Perception AlgorithmsabstractHybrid machine learning algorithms that combine deep learning with probabilistic inference techniques provide highly accurate scene perception for robot manipulation. In particular, a 2-stage approach that combines object detection using convolutional neural networks with Monte-Carlo sampling for pose estimation has been shown to perform particularly well under adversarial scenarios. Unfortunately, this accuracy comes at the cost of high computational complexity, which affects runtime, resource utilization, and energy consumption. This paper describes various challenges in developing complexity-aware techniques for robust robot perception and presents a novel hardware accelerator that addresses these challenge. Experimental results show our design is at least 30% faster and consumes 97% less energy compared to an implementation on a high-end GPU. Compared to a low-power GPU implementation, our design is 95% faster while consuming 96% less energy, demonstrating that accurate, energy-efficient scene perception is possible in real time with targeted hardware acceleration. Yanqi Liu, Can Eren Derman, Giuseppe Calderoni, R. Iris Bahar |
ICCAD | 4 |
| 2019 | IgnoreTM: Opportunistically Ignoring Timing Violations for Energy Savings using HTMabstractEnergy consumption is the dominant factor in many computing systems. Voltage scaling is a widely used technique to lower energy consumption, which exploits supply voltage margins to ensure reliable circuit operation. Aggressive voltage scaling will slow signal propagation; without coherent frequency relaxation, timing violations may be generated. Hardware Transactional Memory (HTM) offers an error recovery mechanism that allows reliable execution and power savings with modest overhead. We propose IgnoreTM, an adaptive error management framework, that tolerates (i.e., opportunistically ignores) timing violations, allowing for more aggressive voltage scaling. Our experimental results show that IgnoreTM allows up to 47% total energy savings with negligible impact on runtime. Dimitra Papagiannopoulou, Sungseob Whang, Tali Moreshet, R. Iris Bahar |
DATE | 4 |
| 2019 | GRIP: Generative Robust Inference and Perception for Semantic Robot Manipulation in Adversarial EnvironmentsabstractRecent advancements have led to a proliferation of machine learning systems used to assist humans in a wide range of tasks. However, we are still far from accurate, reliable, and resource-efficient operations of these systems. For robot perception, convolutional neural networks (CNNs) for object detection and pose estimation are recently coming into widespread use. However, neural networks are known to suffer from overfitting during the training process and are less robust under unforeseen conditions (which makes them especially vulnerable to adversarial scenarios). In this work, we propose Generative Robust Inference and Perception (GRIP) as a two-stage object detection and pose estimation system that aims to combine the relative strengths of discriminative CNNs and generative inference methods to achieve robust estimation. Our results show that a second stage of sample-based generative inference is able to recover from false object detections by CNNs, and produce robust estimations in adversarial conditions. We demonstrate the efficacy of GRIP robustness through comparison with state-of-the-art learning-based pose estimators and pick-and-place manipulation in dark and cluttered environments. Zhiqiang Sui, Zhefan Ye, Yanqi Liu, R. Iris Bahar, Odest Chadwicke Jenkins |
IROS | 6 |
| 2019 | Concurrent Data Structures with Near-Data-Processing: an Architecture-Aware ImplementationabstractRecent advances in memory architectures have provoked renewed interest in near-data-processing (NDP) as way to alleviate the "memory wall" problem. An NDP architecture places logic circuits, such as simple processors, in close proximity to memory. Effective use of NDP architectures requires rethinking data structures and their algorithms. Here, we provide an empirical evaluation of several NDP-aware algorithms for general-purpose concurrent data structures such as linked-lists, skiplists, and FIFO queues. The empirical analysis reveals that the potential benefits of NDP-based concurrent data structures are less than what had been expected in earlier studies. In turn, we introduce lightweight NDP hardware modifications, inspired by initial observations on data access patterns and underlying DRAM activity. Even the minimal changes to hardware significantly improve the performance and energy consumption of NDP-based concurrent data structures, and in many cases, the resulting data structures outperform state-of-the-art concurrent data structures. Jiwon Choe, Amy Huang, Tali Moreshet, Maurice Herlihy, R. Iris Bahar |
SPAA | 5 |
| 2019 | Special Session: Does Approximation Make Testing Harder (or Easier)?abstractMany important application domains, including machine learning, feature intrinsically noise tolerant algorithms. These algorithms process massive, yet noisy and redundant data, by probabilistic and often iterative techniques. As a result, there is a range of valid outputs rather than a single golden value. While this may translate into relaxed constraints for testing and verification of approximate systems, distinguishing actual design bugs from what is being approximated also becomes harder. In this paper, using representative case studies, we pose several challenges for the test and verification community as approximate computing becomes more prevalent as a design of choice in order to achieve performance gains, power or energy savings, improved reliability or reduced software and/or hardware complexity. R. Iris Bahar, Ulya R. Karpuzcu, Sasa Misailovic |
VTS | 1 |
| 2019 | Repurposing FPGAs for Tester Design to Enhance Field-Testing in a 3D Stack
Fanchen Zhang, Kundan Nepal, Jennifer Dworak, Theodore W. Manikas, R. Iris Bahar |
J. Electron. Test. | 7 |
| 2018 | Robust object estimation using generative-discriminative inference for secure robotics applicationsabstractConvolutional neural networks (CNNs) are of increasing widespread use in robotics, especially for object recognition. However, such CNNs still lack several critical properties necessary for robots to properly perceive and function autonomously in uncertain, and potentially adversarial, environments. In this paper, we investigate factors for accurate, reliable, and resource-efficient object and pose recognition suitable for robotic manipulation in adversarial clutter. Our exploration is in the context of a three-stage pipeline of discriminative CNN-based recognition, generative probabilistic estimation, and robot manipulation. This pipeline proposes using a SAmpling Network Density filter, or SAND filter, to recover from potentially erroneous decisions produced by a CNN through generative probabilistic inference. We present experimental results from SAND filter perception for robotic manipulation in tabletop scenes with both benign and adversarial clutter. These experiments vary CNN model complexity for object recognition and evaluate levels of inaccuracy that can be recovered by generative pose inference. This scenario is extended to consider adversarial environmental modifications with varied lighting, occlusions, and surface modifications. Yanqi Liu, Alessandro Costantini, R. Iris Bahar, Zhiqiang Sui, Zhefan Ye, Shiyang Lu, Odest Chadwicke Jenkins |
ICCAD | 3 |
| 2018 | A Sub-Threshold Noise Transient Simulator Based on Integrated Random Telegraph and Thermal Noise ModelingabstractNear-threshold and sub-threshold voltage designs have been identified as possible solutions to overcome the limitations introduced by energy consumption in modern very large scale integration circuits. However, as we approach sub-10 nm transistor technology, aggressive voltage, and gate length scaling will reduce the reliability of logic circuits due to the increasing impact of noise and variability effects. Therefore, designers need new tools to simulate logic circuits in the presence of noise. Time-domain analysis helps understand how transient faults affect a circuit and can guide designers in producing noise-resistant circuitry. However, standard approaches to modeling intrinsic noise sources in the time domain are computationally expensive. Moreover, small noise-driven fluctuations in electron occupation of circuit nodes introduce time-varying biasing point fluctuations, increasing the modeling complexity. To address these challenges, this paper introduces a new approach to modeling thermal noise and random telegraph signal noise directly in the time domain by developing and solving a series of stochastic differential equations. In comparisons to traditional SPICE-based simulations, our approach can provide three orders of magnitude speedup in simulation time without sacrificing accuracy. Moreover, we introduce a novel, iterative threshold-crossing algorithm, aimed at the efficient sampling of rare noise transients. We show that Monte-Carlo simulations based on this approach can detect rare high-amplitude single event transients that would be impossible to uncover with standard transient simulators. Marco Donato, R. Iris Bahar, William R. Patterson, Alexander Zaslavsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Hardware-Software Codesign of Accurate, Multiplier-free Deep Neural NetworksabstractWhile Deep Neural Networks (DNNs) push the state-of-the-art in many machine learning applications, they often require millions of expensive floating-point operations for each input classification. This computation overhead limits the applicability of DNNs to low-power, embedded platforms and incurs high cost in data centers. This motivates recent interests in designing low-power, low-latency DNNs based on fixed-point, ternary, or even binary data precision. While recent works in this area offer promising results, they often lead to large accuracy drops when compared to the floating-point networks. We propose a novel approach to map floating-point based DNNs to 8-bit dynamic fixed-point networks with integer power-of-two weights with no change in network architecture. Our dynamic fixed-point DNNs allow different radix points between layers. During inference, power-of-two weights allow multiplications to be replaced with arithmetic shifts, while the 8-bit fixed-point representation simplifies both the buffer and adder design. In addition, we propose a hardware accelerator design to achieve low-power, low-latency inference with insignificant degradation in accuracy. Using our custom accelerator design with the CIFAR-10 and ImageNet datasets, we show that our method achieves significant power and energy savings while increasing the classification accuracy. Hokchhay Tann, Soheil Hashemi, R. Iris Bahar, Sherief Reda |
DAC | 3 |
| 2017 | Understanding the impact of precision quantization on the accuracy and energy of neural networksabstractDeep neural networks are gaining in popularity as they are used to generate state-of-the-art results for a variety of computer vision and machine learning applications. At the same time, these networks have grown in depth and complexity in order to solve harder problems. Given the limitations in power budgets dedicated to these networks, the importance of low-power, low-memory solutions has been stressed in recent years. While a large number of dedicated hardware using different precisions has recently been proposed, there exists no comprehensive study of different bit precisions and arithmetic in both inputs and network parameters. In this work, we address this issue and perform a study of different bit-precisions in neural networks (from floating-point to fixed-point, powers of two, and binary). In our evaluation, we consider and analyze the effect of precision scaling on both network accuracy and hardware metrics including memory footprint, power and energy consumption, and design area. We also investigate training-time methodologies to compensate for the reduction in accuracy due to limited bit precision and demonstrate that in most cases, precision scaling can deliver significant benefits in design metrics at the cost of very modest decreases in network accuracy. In addition, we propose that a small portion of the benefits achieved when using lower precisions can be forfeited to increase the network size and therefore the accuracy. We evaluate our experiments, using three well-recognized networks and datasets to show its generality. We investigate the trade-offs and highlight the benefits of using lower precisions in terms of energy and memory footprint. Soheil Hashemi, Nicholas Anthony, Hokchhay Tann, R. Iris Bahar, Sherief Reda |
DATE | 4 |
| 2017 | Edge-TM: Exploiting Transactional Memory for Error Tolerance and Energy EfficiencyabstractScaling of semiconductor devices has enabled higher levels of integration and performance improvements at the price of making devices more susceptible to the effects of static and dynamic variability. Adding safety margins (guardbands) on the operating frequency or supply voltage prevents timing errors, but has a negative impact on performance and energy consumption. We propose Edge-TM , an adaptive hardware/software error management policy that ( i ) optimistically scales the voltage beyond the edge of safe operation for better energy savings and ( ii ) works in combination with a Hardware Transactional Memory (HTM)-based error recovery mechanism. The policy applies dynamic voltage scaling (DVS) (while keeping frequency fixed) based on the feedback provided by HTM, which makes it simple and generally applicable. Experiments on an embedded platform show our technique capable of 57% energy improvement compared to using voltage guardbands and an extra 21-24% improvement over existing state-of-the-art error tolerance solutions, at a nominal area and time overhead. Dimitra Papagiannopoulou, Andrea Marongiu, Tali Moreshet, Maurice Herlihy, R. Iris Bahar |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2016 | Thrifty-malloc: A HW/SW codesign for the dynamic management of hardware transactional memory in embedded multicore systemsabstractWe present thrifty-malloc: a transaction-friendly dynamic memory manager for high-end embedded multicore systems. The manager combines modularity, ease-of-use and hardware transactional memory (HTM) compatibility in a light-weight and memory-efficient design. Thrifty-malloc is easy to deploy and configure for non-expert programmers, yet provides good performance with low memory overhead for highly-parallel embedded applications running on massively parallel processor arrays (MPPAs) or many-core architectures. In addition, the transparent mechanisms that increase our manager's resilience to unpredictable dynamic situations incur a low timing overhead in comparison to established techniques. Thomas Carle, Dimitra Papagiannopoulou, Tali Moreshet, Andrea Marongiu, Maurice Herlihy, R. Iris Bahar |
CASES | 6 |
| 2016 | A fast simulator for the analysis of sub-threshold thermal noise transientsabstractThe gate length of CMOS transistors is continuing to shrink down to the sub-10nm region and operating voltages are moving toward near-threshold and even sub-threshold values. With this trend, the number of electrons responsible for the total charge of a CMOS node is greatly reduced. As a consequence, thermal fluctuations that shift a gate from its equilibrium point may no longer have a negligible impact on circuit reliability. Time-domain analysis helps understand how transient faults affect a circuit and can guide designers in producing noise-resistant circuitry. However, modeling thermal noise in the time-domain is computationally very costly. Moreover, small fluctuations in electron occupation introduce time-varying biasing point fluctuations, increasing the modeling complexity. To address these challenges, this paper introduces a new approach to modeling thermal noise directly in the time domain by developing a series of stochastic differential equations (SDE) to model various transient effects in the presence of thermal noise. In comparisons to SPICE-based simulations, our approach can provide 3 orders of magnitude speedup in simulation time, with comparable accuracy. This simulation framework is especially valuable for detecting rare events that could translate into fault-inducing noise transients. While it is computationally infeasible to use SPICE to detect such rare events due to thermal noise, we introduce a new iterative approach that allows detecting 6σ events in a matter of a few hours. Marco Donato, R. Iris Bahar, William R. Patterson, Alexander Zaslavsky |
DAC | 2 |
| 2016 | A low-power dynamic divider for approximate applicationsabstractIn this work, a low-power, low-error divider design is proposed that can achieve significant power and area savings, while introducing insignificant inaccuracies to the output. The design of our divider is highly scalable, offering a wide range of power and inaccuracy trade-offs based on the application requirements. Furthermore, the proposed divider has a lower delay compared to the accurate design, enabling its use on the critical path. We theoretically analyze the error of our design as a function of its configuration, and we thoroughly evaluate the error and power characteristics of our divider in a standalone case and demonstrate that the proposed design can achieve up to 70% in power savings, while introducing an mean average absolute error of only 3.08%. We also implement three image-processing applications in hardware using our divider and conclude that use of the proposed divider will not perceptibly impact their quality-of-service while achieving power benefits of up to 75%. Soheil Hashemi, R. Iris Bahar, Sherief Reda |
DAC | 2 |
| 2016 | Hardware acceleration of feature detection and description algorithms on low-power embedded platformsabstractImage features are broadly used in embedded computer vision applications, from object detection and tracking to motion estimation and 3D reconstruction. Efficient feature extraction and description are crucial due to the real-time requirements of such applications over a constant stream of input data. High-speed computation typically comes at the cost of high power dissipation, yet embedded systems are often highly power constrained, making discovery of power-aware solutions especially critical for these systems. In this paper, we present a power and performance evaluation of three low cost feature detection and description algorithms implemented on various embedded systems (embedded CPUs, GPUs and FPGAs). We show that FPGAs in particular offer attractive solutions for both performance and power and describe several design techniques utilized to accelerate feature extraction and description algorithms on low-cost Zynq SoC FPGAs. Onur Ulusel, Christopher B. Picardo, Christopher B. Harris, Sherief Reda, R. Iris Bahar |
FPL | 5 |
| 2016 | Design of Error-Resilient Logic Gates with Reinforcement Using ImplicationsabstractOperating circuits in the sub-threshold region can save power, but at the cost of higher susceptibility to noise. This paper analyzes various gate-level error-mitigation designs appropriate for sub-threshold circuits. Previous works have proposed a modified version of the Schmitt trigger gate that uses logic implications to reinforce correct functional behavior. However, the increased error resilience requires increased area, delay, and power overhead. To address these shortcomings, we introduce two alternative and less costly approaches to reinforcing correct logic behavior via implications. In addition, to provide more flexibility in implication selection, we consider not just simple implications that reinforce relationships between two signals, but also more complex 3-signal implications within the circuit. Our simulation results demonstrate that these alternative gate structures can outperform the Schmitt trigger version as long as the noise on the reinforcement signals themselves is sufficiently low. Xijing Han, Marco Donato, R. Iris Bahar, Alexander Zaslavsky, William R. Patterson |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | A Simulation Framework for Analyzing Transient Effects Due to Thermal Noise in Sub-Threshold CircuitsabstractNoise analysis in nonlinear logic circuits requires models that take into account time-varying biasing conditions. When considering thermal noise, which moves the circuit away from its equilibrium point, a correct modeling approach has to go beyond the additive white Gaussian noise (AWGN) used in classical noise analysis. Even when accurate models are available, running standard Monte-Carlo simulations that will expose rare soft errors may still be computationally prohibitive. Probabilistic methods are often preferred for estimating the failure rate. However, these approaches may not provide any insight about the dynamic response to noise events. In this paper, we target both problems in the sub-threshold logic application domain. We first provide a time-domain model for fundamental, technology-independent thermal noise in sub-threshold circuits. Then, we use this model to generate noise input files for SPICE transient analysis. The effectiveness of the approach is demonstrated using 7nm FinFET predictive technology models (PTM) for an inverter and a NAND gate. Marco Donato, R. Iris Bahar, William R. Patterson, Alexander Zaslavsky |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | Playing with Fire: Transactional Memory Revisited for Error-Resilient and Energy-Efficient MPSoC ExecutionabstractAs silicon integration technology pushes toward atomic dimensions, errors due to static and dynamic variability are an increasing concern. To avoid such errors, designers often turn to "guardband" restrictions on the operating frequency and voltage. If guardbands are too conservative, they limit performance and waste energy, but less conservative guardbands risk moving the system closer to its Critical Operating Point (COP), a frequency-voltage pair that, if surpassed, causes massive instruction failures. In this paper, we propose a novel scheme that allows to dynamically adjust to an evolving COP and operate at highly reduced margins, while guaranteeing forward progress. Specifically, our scheme dynamically monitors the platform and adaptively adjusts to the COP among multiple cores, using lightweight checkpointing and roll-back mechanisms adopted from Hardware Transactional Memory (HTM) for error recovery. Experiments demonstrate that our technique is particularly effective in saving energy while also offering safe execution guarantees. To the best of our knowledge, this work is the first to describe a full-fledged HTM implementation for error-resilient and energy-efficient MPSoC execution. Dimitra Papagiannopoulou, Andrea Marongiu, Tali Moreshet, Luca Benini, Maurice Herlihy, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 6 |
| 2015 | DRUM: A Dynamic Range Unbiased Multiplier for Approximate ApplicationsabstractMany applications for signal processing, computer vision and machine learning show an inherent tolerance to some computational error. This error resilience can be exploited to trade off accuracy for savings in power consumption and design area. Since multiplication is an essential arithmetic operation for these applications, in this paper we focus specifically on this operation and propose a novel approximate multiplier with a dynamic range selection scheme. We design the multiplier to have an unbiased error distribution, which leads to lower computational errors in real applications because errors cancel each other out, rather than accumulate, as the multiplier is used repeatedly for a computation. Our approximate multiplier design is also scalable, enabling designers to parameterize it depending on their accuracy and power targets. Furthermore, our multiplier benefits from a reduction in propagation delay, which enables its use on the critical path. We theoretically analyze the error of our design as a function of its parameters and evaluate its performance for a number of applications in image processing, and machine classification. We demonstrate that our design can achieve power savings of 54% - 80%, while introducing bounded errors with a Gaussian distribution with near-zero average and standard deviations of 0.45% - 3.61%. We also report power savings of up to 58% when using the proposed design in applications. We show that our design significantly outperforms other approximate multipliers recently proposed in the literature. Soheil Hashemi, R. Iris Bahar, Sherief Reda |
ICCAD | 2 |
| 2015 | Message from the program co-chairsabstractWe would first like to express our many thanks to Resit Sendag, this year's General Chair, as well as the rest of the research community for the opportunity to chair the program for the 10thIEEE International Conference on Networking, Architecture, and Storage (NAS 2015). Jun Wang 0001, R. Iris Bahar |
NAS | 2 |
| 2015 | Repairing a 3-D Die-Stack Using Available Programmable Logicabstract3-D die-stacks hold great promise for increasing system performance, but difficulties in testing dies and assembling a 3-D stack are leading to yield issues and slowing the large scale manufacturing of these devices. In many cases, a single defective die will kill the entire stack. To help mitigate this issue, we explore the possibility of repairing a stack that contains a defective die by utilizing an field programmable gate array (FPGA) that has already been included in the stack for other purposes, such as performance enhancement. Specifically, we propose bypassing the defective portion of a nonprogrammable die by replacing the defective functionality with functionality on the FPGA. In this paper, we discuss what additional logic must be added to an Application-Specific Integrated Circuit (ASIC) die to allow such a bypass to occur. We then show through detailed simulation of a 2.5-D Xilinx FPGA how bypassing of logic can be achieved and throughput maintained even when the two different dies involved operate at different frequencies. Finally, we explore the performance of this technique in a superscalar, out-of-order processor, where different functional units are marked for replacement. Our simulation results show that not only can we salvage a device that would otherwise have to be discarded, but creating multiple copies of the defective partition in the FPGA can allow us to regain performance even when the latency of the units in the FPGA is longer than that of the original defective copy. Kundan Nepal, Soha Alhelaly, Jennifer Dworak, R. Iris Bahar, Theodore W. Manikas, Ping Guikundan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | Energy-Efficient and High-Performance Lock Speculation Hardware for Embedded Multicore SystemsabstractEmbedded systems are becoming increasingly common in everyday life and like their general-purpose counterparts, they have shifted towards shared memory multicore architectures. However, they are much more resource constrained, and as they often run on batteries, energy efficiency becomes critically important. In such systems, achieving high concurrency is a key demand for delivering satisfactory performance at low energy cost. In order to achieve this high concurrency, consistency across the shared memory hierarchy must be accomplished in a cost-effective manner in terms of performance, energy, and implementation complexity. In this article, we propose Embedded-Spec, a hardware solution for supporting transparent lock speculation, without the requirement for special supporting instructions. Using this approach, we evaluate the energy consumption and performance of a suite of benchmarks, exploring a range of contention management and retry policies. We conclude that for resource-constrained platforms, lock speculation can provide real benefits in terms of improved concurrency and energy efficiency, as long as the underlying hardware support is carefully configured. Dimitra Papagiannopoulou, Giuseppe Capodanno, Tali Moreshet, Maurice Herlihy, R. Iris Bahar |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2015 | Introduction to the Special Issue on Reliable, Resilient, and Robust Design of Circuits and SystemsabstractNo abstract available. R. Iris Bahar, Alex K. Jones, Yuan Xie 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2014 | ABACUS: A technique for automated behavioral synthesis of approximate computing circuitsabstractMany classes of applications, especially in the domains of signal and image processing, computer graphics, computer vision, and machine learning, are inherently tolerant to inaccuracies in their underlying computations. This tolerance can be exploited to design approximate circuits that perform within acceptable accuracies but have much lower power consumption and smaller area footprints (and often better run times) than their exact counterparts. In this paper, we propose a new class of automated synthesis methods for generating approximate circuits directly from behavioral-level descriptions. In contrast to previous methods that operate at the Boolean level or use custom modifications, our automated behavioral synthesis method enables a wider range of possible approximations and can operate on arbitrary designs. Our method first creates an abstract synthesis tree (AST) from the input behavioral description, and then applies variant operators to the AST using an iterative stochastic greedy approach to identify the optimal inexact designs in an efficient way. Our method is able to identify the optimal designs that represent the Pareto frontier trade-off between accuracy and power consumption. Our methodology is developed into a tool we call ABACUS, which we integrate with a standard ASIC experimental flow based on industrial tools. We validate our methods on three realistic Verilog-based benchmarks from three different domains - signal processing, computer vision and machine learning. Our tool automatically discovers optimal designs, providing area and power savings of up to 50% while maintaining good accuracy. Kumud Nepal, R. Iris Bahar, Sherief Reda |
DATE | 3 |
| 2014 | Fast Design Exploration for Performance, Power and Accuracy Tradeoffs in FPGA-Based AcceleratorsabstractThe ease-of-use and reconfigurability of FPGAs makes them an attractive platform for accelerating algorithms. However, accelerating becomes a challenging task as the large number of possible design parameters lead to different accelerator variants. In this article, we propose techniques for fast design exploration and multi-objective optimization to quickly identify both algorithmic and hardware parameters that optimize these accelerators. This information is used to run regression analysis and train mathematical models within a nonlinear optimization framework to identify the optimal algorithm and design parameters under various objectives and constraints. To automate and improve the model generation process, we propose the use of L 1 -regularized least squares regression techniques.We implement two real-time image processing accelerators as test cases: one for image deblurring and one for block matching. For these designs, we demonstrate that by sampling only a small fraction of the design space (0.42% and 1.1%), our modeling techniques are accurate within 2%--4% for area and throughput, 8%--9% for power, and 5%--6% for arithmetic accuracy. We show speedups of 340× and 90× in time for the test cases compared to brute-force enumeration. We also identify the optimal set of parameters for a number of scenarios (e.g., minimizing power under arithmetic inaccuracy bounds). Onur Ulusel, Kumud Nepal, R. Iris Bahar, Sherief Reda |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2012 | High Performance Parallel JPEG2000 Streaming Decoder Using GPGPU-CPU Heterogeneous SystemabstractThe JPEG2000 image coding standard provides many superior features compared to JPEG and other compression standards. However, the relatively slow performance of JPEG2000, especially in software implementations, is a critical drawback of the standard. Moreover, as image sizes rapidly grow in size, higher demands on performance for image coding and processing are introduced, making the slow performance of JPEG2000 even further pronounced. While much effort over the past decade has been devoted to accelerating the JPEG2000 encoder, there have been very few studies focusing on improving the performance of the JPEG2000 decoder, despite the fact that the performance of the decoder is just as critical as the encoder. This paper proposes a high-performance JPEG2000 decoder that efficiently exploits the recent improvements of modern parallel programming models and hardware architectures. Specifically, a parallel streaming decoder running on a GPGPU-CPU heterogeneous system is developed to fully exploit the flexibility of the high-performance multi-core CPUs and the massively parallel capability of GPGPUs. In addition, a new task scheduling strategy is developed that exploits the soft-heterogeneity in OpenCL and C/C++ at runtime in order to gain a significant performance boost. Running on a heterogeneous configuration of one Nvidia GTX 480 GPU and one Intel Core i7 CPU, the parallel streaming decoder gains more than 8X speedup in runtime compared to the JasPer JPEG2000 software implementation. Roto Le, Joseph L. Mundy, R. Iris Bahar |
ASAP | 3 |
| 2012 | Fast Multi-Objective Algorithmic Design Co-Exploration for FPGA-based AcceleratorsabstractThe reconfigurability of Field Programmable Gate Arrays (FPGAs) makes them an attractive platform for accelerating algorithms. Accelerating a particular algorithm is a challenging task as the large number of possible algorithmic and hardware design parameters lead to different accelerator variant implementations, each with its own metrics such as performance, area, power, and arithmetic accuracy characteristics. To identify these parameters that optimize the accelerator for certain metrics, we propose techniques for fast design space exploration and non-linear multi-objective optimization (e.g., minimize power under arithmetic inaccuracy bounds). Our methodology samples a small part of the design space and uses measurements from the sampled implementations to train mathematical models for the different metrics. To automate and improve the model generation process, we propose the use of L1-regularized least squares regression techniques. To demonstrate the effectiveness of our approach, we implement a high-throughput real-time accelerator for image debluring. We demonstrate the accuracy (e.g., within 8% for power modeling) of our modeling techniques and their ability to identify the optimal accelerator designs with large speed-ups (340×) in comparison to brute-force enumeration. Kumud Nepal, Onur Ulusel, R. Iris Bahar, Sherief Reda |
FCCM | 3 |
| 2012 | A noise-immune sub-threshold circuit design based on selective use of Schmitt-trigger logicabstractNanoscale circuits operating at sub-threshold voltages are affected by growing impact of random telegraph signal (RTS) and thermal noise. Given the low operational voltages and subsequently lower noise margins, these noise phenomena are capable of changing the value of some of the nodes in the circuit, compromising the reliability of the computation. We propose a method for improving noise-tolerance by selectively applying feed-forward reinforcement to circuits based on use of existing invariant relationships. As reinforcement mechanism, we used a modification of the standard CMOS gates based on the Schmitt trigger circuit. SPICE simulations show our solution offers better noise immunity than both standard CMOS and fully reinforced circuits, with limited area and power overhead. Marco Donato, Fabio Cremona, Warren Jin 0002, R. Iris Bahar, William R. Patterson, Alexander Zaslavsky, Joseph L. Mundy |
ACM Great Lakes Symposium on VLSI | 4 |
| 2012 | NBTI-Aware Data Allocation Strategies for Scratchpad Based Embedded Systems
Cesare Ferri, Dimitra Papagiannopoulou, R. Iris Bahar, Andrea Calimera |
J. Electron. Test. | 3 |
| 2012 | Using implications to choose tests through suspect fault identificationabstractAs circuits continue to scale to smaller feature sizes, wearout and latent defects are expected to cause an increasing number of errors in the field. Online error detection techniques, including logic implication-based checker hardware, are capable of detecting at least some of these errors as they occur. However, recovery may be expensive, and the underlying problem may lead to multiple failures of a core over time. In this article, we will investigate the diagnostic capability of logic implications to identify possible failure locations when an error is detected online. We will then utilize this information to select a highly efficient test set that can be used to effectively test the identified suspect locations in both the failing core and in other identical cores in the system. Jennifer Dworak, Kundan Nepal, Nuno Alves, Yiwen Shi, Nicholas Imbriglia, R. Iris Bahar |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2011 | Dynamic Test Set Selection Using Implication-Based On-Chip DiagnosisabstractWe propose using logic implications as a source of online diagnostic data for on-chip test set selection by taking advantage of their ability to automatically identify a restricted set of faults as the potential cause of an observed error. This information will be used to dynamically choose a test set to detect systematic latent defects or wear out in a multi core system. Nuno Alves, Yiwen Shi, Nicholas Imbriglia, Jennifer Dworak, Kundan Nepal, R. Iris Bahar |
ETS | 6 |
| 2011 | Enhancing online error detection through area-efficient multi-site implicationsabstractWe present a new method to identify multi-site implications that can significantly increase the fault coverage of error-detecting hardware without increasing the area overhead. This method intelligently divides the input space about the functions of internal circuit sites and finds new valuable implications that can share gates in checker logic. Nuno Alves, Yiwen Shi, Jennifer Dworak, R. Iris Bahar, Kundan Nepal |
VTS | 4 |
| 2011 | Test Vector Generation for Post-Silicon Delay Testing Using SAT-Based Decision Problems
Desta Tadesse, R. Iris Bahar, Joel Grodstein |
J. Electron. Test. | 2 |
| 2010 | Improving the testability and reliability of sequential circuits with invariant logicabstractIn this paper, we propose the use of logic implications to enhance online error detection capabilities and to improve the testing efficiency of an integrated circuit. These logic implications are implemented in hardware and help to verify that expected invariant circuit relationships are satisfied during field operation. Thus, any implication violation will indicate the presence of an error due to some faulty circuit behavior. In addition, checking these logic implications in hardware will create additional circuit outputs, which may be useful for compacting $n$-detect test sets. Our results show that logic implications can provide significant error detection and test pattern count reduction with very limited hardware overhead. Nuno Alves, Kundan Nepal, Jennifer Dworak, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 4 |
| 2010 | Numerical queue solution of thermal noise-induced soft errors in subthreshold CMOS devicesabstractPower consumption requirements drive CMOS scaling to ever lower supply voltages, reducing the stability margin with respect to thermal noise and raising the probability for thermally-induced soft errors. Given the long time scale of noise-induced soft errors, conventional Monte Carlo simulations cannot be used to predict error rates and alternative approaches are needed. In this paper, the analysis of thermal fluctuations in a CMOS flip-flop is performed using a 2D queue that maps the available configurations for the flip-flop in terms of electron populations on the two inverters, with the two stable logic states at the opposite corners of the 2D matrix. Trial simulations for model systems show that the thermally-induced logic transitions involve only a limited number of states immediately above and below the main diagonal of the full 2D queue. We present a numerical solution based on variable precision arithmetic for a truncated 2D queue consisting of a variable number of near-diagonal states. It is shown that increasing the width of the near-diagonal queue, an accurate solution for the error rate is asymptotically obtained without the need to consider the full 2D queue. Our approach is used to calculate the mean time to failure of flip-flops built in a 45-nm fully-depleted silicon-on-insulator (FD-SOI) technology modeled in the subthreshold regime, including parasitics. As a predictive tool, the framework can be used to investigate the thermal stability of devices built in future technologies and as a measure of device reliability in VLSI design. Pooya Jannaty, Florian C. Sabou, R. Iris Bahar, Joseph L. Mundy, William R. Patterson, Alexander Zaslavsky |
ACM Great Lakes Symposium on VLSI | 3 |
| 2010 | Energy and Throughput Efficient Transactional Memory for Embedded Multicore Systems
Cesare Ferri, Samantha Wood 0001, Tali Moreshet, R. Iris Bahar, Maurice Herlihy |
HiPEAC | 4 |
| 2010 | Embedded-TM: Energy and complexity-effective hardware transactional memory for embedded multicore systems
Cesare Ferri, Samantha Wood 0001, Tali Moreshet, R. Iris Bahar, Maurice Herlihy |
J. Parallel Distributed Comput. | 4 |
| 2010 | A Cost Effective Approach for Online Error Detection Using Invariant RelationshipsabstractThis paper investigates the use of logic implication checkers for the online detection of errors. A logic implication, or invariant relationship, must hold for all valid input conditions; therefore, any violation of this implication will indicate an error due to an intermittent fault. Techniques are presented to efficiently identify the most useful logic implications to include in checker hardware such that the probability of error detection is maximized while minimizing the additional hardware and delay overhead. Results show that significant error detection is possible-even with only a 10% area overhead-while minimizing impact on delay and power. Nuno Alves, Alison Buben, Kundan Nepal, Jennifer Dworak, R. Iris Bahar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2010 | Temperature-Insensitive Dual- Vth Synthesis for Nanometer CMOS Technologies Under Inverse Temperature DependenceabstractWith the scaling of CMOS technologies, the gap between nominal supply voltage and threshold voltage has decreased significantly. This trend is further amplified in low-power nanometer libraries, which feature cells with identical size and functionality, but different threshold voltages. As a consequence, different cells may have different delay behaviors as the temperature varies within a circuit. For instance, cells with low-threshold devices may experience an increase in delay when temperature increases, whereas cells using high-threshold devices may experience the opposite behavior. The latter effect, also known as inverse temperature dependence (ITD), poses new challenges to circuit designers. Besides making timing analysis more difficult, ITD has important and unforeseeable consequences for power-aware logic synthesis. This paper describes the impact that ITD may have on the design of nanometer circuits. We also provide a threshold voltage assignment algorithm for dual threshold voltage synthesis, which guarantees temperature-insensitive operation of the circuits, together with a significant reduction of both leakage and total power consumption. Experiments performed on a set of standard benchmarks show timing compliance at any operating temperature, and an average leakage reduction around 28% compared to circuits synthesized with a standard synthesis flow that does not take ITD into account. We also apply our proposed synthesis algorithm to a realistic case study consisting of a 32-bit, IEEE-754 floating point unit. Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Detecting errors using multi-cycle invariance informationabstractEnsuring reliable computation at the nanoscale requires mechanisms to detect and correct errors during normal circuit operation. In this paper we propose a method for designing efficient online error detection schemes for circuits based on the identification of invariant relationships in hardware. More specifically, we present a technique that automatically identifies multi-cycle gate-level invariant relationships-where no knowledge of high-level behavioral constraints is required to identify the relationships-and generates the checker logic that verifies these implications. Our results show that cross-cycle implications are particularly useful in discovering difficult-to-detect errors near latch boundaries, and can have a significant impact on boosting error detection rates. Nuno Alves, Kundan Nepal, Jennifer Dworak, R. Iris Bahar |
DATE | 4 |
| 2009 | High-performance, cost-effective heterogeneous 3D FPGA architecturesabstractIn this paper, we propose novel architectural and design techniques for three-dimensional field-programmable gate arrays (3D FPGAs) with Through-Silicon Vias (TSVs). We develop a novel design partitioning methodology that maps the heterogeneous computational resources of an FPGA into a number of die such that the total die area is minimized and the FPGA performance is maximized. Minimizing the total die area leads to direct manufacturing cost savings which is an important incentive to bring 3D technology to the fab and onto the market. An estimation framework is developed to assess the impact of silicon area utilized by 3D interconnect resources while taking into account the large area occupied by TSVs which is crucial to total die area of 3D FPGA. And in order to improve area and performance of 3D FPGA, we design a novel 3D switch box with bypass TSVs. We also analyze the impact of different partitioning strategies on die area and find the optimal number of die that gives the largest reductions in total die area while maximizing the performance. Using a well-developed simulation infrastructure, we show that our methodologies can achieve an average reduction of 27.7% in total die area with a reduced interconnect path delay of about 58%. Roto Le, Sherief Reda, R. Iris Bahar |
FPGA | 3 |
| 2009 | Energy-optimal synchronization primitives for single-chip multi-processorsabstractSynchronization among tasks accounts for a sizable fraction of the energy consumption and execution time of applications running on Multi-Processor Systems-on-Chips platforms. In order to achieve fast and energy-efficient operations, it is therefore essential to implement efficient and power-frugal synchronization primitives. The design of such primitives is complicated by several software and hardware issues, such as: processors running at different speeds, different implementations of the waiting phase upon entering the critical section, and the ratio between static and dynamic power. In this work, we compare a set of classical implementations (i.e., based on busy waiting, or on sleep states) of mutex semaphores, and propose a hybrid (wait/sleep) semaphore in which the sleep state is entered only after a number of busywait cycles. The proposed scheme provides the best overall energy-delay product with respect to previously proposed schemes. Furthermore, we identify an optimal length of the busy-wait cycles, which is empirically shown to depend on the time required to switch from the sleep to the active state. Cesare Ferri, R. Iris Bahar, Mirko Loghi, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | High-performance, cost-effective heterogeneous 3D FPGA architecturesabstractIn this paper, we propose novel architectural and design techniques for three-dimensional field-programmable gate arrays (3D FPGAs) with Through-Silicon Vias (TSVs). We develop a novel design partitioning methodology that maps the heterogeneous computational resources of an FPGA into a number of die such that the total die area is minimized and the FPGA performance is maximized. Minimizing the total die area leads to direct manufacturing cost savings which is an important incentive to bring 3D technology to the fab and onto the market. An estimation framework is developed to assess the impact of silicon area utilized by 3D interconnect resources while taking into account the large area occupied by TSVs which is crucial to total die area of 3D FPGAs. In order to improve area and performance of 3D FPGAs, we design a novel 3D switch box with bypass TSVs. We also analyze the impact of different partitioning strategies on die area and find the optimal number of die that gives the largest reductions in total die area while maximizing the performance. Using a well-developed simulation infrastructure, we show that our methodologies can achieve an average reduction of 27.7% in total die area with a reduced interconnect path delay of about 58%. Roto Le, Sherief Reda, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Compacting test vector sets via strategic use of implicationsabstractAs the complexity of integrated circuits has increased, so has the need for improving testing efficiency. Unfortunately, the types of defects are also becoming more complex, which in turn makes simple approaches for testing inadequate. Using n-detect testing can improve detect coverage; however, this approach can greatly increase the test set size. In this proof-of-concept paper we investigate the use of logic implication checkers, inserted in hardware, as an aid in compacting n-detect test sets. We show that checker hardware with minimal area overhead can reduce test set size by up to 25%. In addition, this implication checker can serve a dual purpose for online error detection. Nuno Alves, Jennifer Dworak, R. Iris Bahar, Kundan Nepal |
ICCAD | 3 |
| 2009 | Reducing the leakage and timing variability of 2D ICcs using 3D ICsabstractThis paper examines the ramifications of using 3D integration technology on the leakage and timing variability of integrated circuits. We develop models that estimate the outcome of mapping a 2D design onto a 3D stack from a process variation perspective. We statistically prove and experimentally demonstrate that 3D integration is a useful technique to combat process variations even if the die/wafers layers involved in 3D stacks are integrated blindly without any parametric tests prior to integration. We further show that if individual die parametric testing information is available, then it is possible to drastically reduce the impact of process variations. We develop fast, near optimal integration strategies based on recursive matching techniques. Our results show that 3D integration can reduce the variability in leakage and timing of planar ICs by around 50% without any testing and by more than 90% with additional test requirements. Sherief Reda, Aung Si, R. Iris Bahar |
ISLPED | 3 |
| 2009 | AutoRex: An automated post-silicon clock tuning toolabstractPost-silicon clock-tuning is a technique used as part of speed-debug efforts to increase the allowable clock frequency of a chip. These days, it is not uncommon for high-end microprocessors to have cores containing a few thousand clock-tuning elements (i.e., variable-delay buffers). Each such buffer can be assigned to one of several possible discrete delay values, as part of the post-silicon speed debugging process. With the proper mix of assignments, many chips that initially could not meet targeted speed requirements, can now run within specification. With thousands of tunable buffers available on chip, the possible combination of assignments to the delay values is quite large. In addition, process variation causes the same design, once fabricated into silicon, to have different critical paths across different chips. Thus a specific buffer-delay assignment that most improves clock frequency for some chips may not be optimal for all chips. In this paper, we propose a tool we call AutoRex, that produces clock-tuning assignments automatically. AutoRex operates by taking data from a volume experiment across multiple process corners and analyzes this data using satisfiability modulo theory (SMT) solvers to create a single ¿recipe¿ for delay buffer assignments such that the clock frequency of the chip is improved as much as possible over the entire sample of chips. Our results show up to a 9% improvement in frequency using AutoRex. Desta Tadesse, Joel Grodstein, R. Iris Bahar |
ITC | 3 |
| 2009 | Introduction to special section: Best of NANOARCH 2008abstractNo abstract available. R. Iris Bahar |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2008 | Temperature-insensitive synthesis using multi-vt librariesabstractTemperature fluctuations can alter the delay in MOS circuits. However, increases in temperature do not always lead to a corresponding increase in circuit delay, specifically when operating at low supply voltages. Instead a temperature inversion effect can be observed on the delay of MOS devices under certain conditions, where the delay actually decreases as temperature increases. Given these non-monotonic effects, guaranteeing timing correctness can no longer be achieved simply by characterizing the design under worst case (i.e., high temperature) conditions. In this paper, we present a synthesis methodology in which multi-Vth design is used to generate temperature-insensitive circuits, while minimizing leakage power dissipation as a side-effect. Our experiments with ISCAS benchmark circuits demonstrate the promise of this approach and show that significant reduction in static power is also possible. Andrea Calimera, Enrico Macii, Massimo Poncino, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 4 |
| 2008 | Energy efficient synchronization techniques for embedded architecturesabstractWe evaluate the energy-efficiency and performance of a number of synchronization mechanisms adapted for embedded devices. We focus on simple hardware accelerators for common software synchronization patterns. We compare the energy efficiency of a range of shared memory benchmarks using both spin-locks and a simple hardware transactional memory. In most cases, transactional memory provides both significantly reduced energy consumption and increased throughput. We also consider applications that employ concurrency patterns based on semaphores, such as pipelines and barriers. We propose and evaluate a novel energy-efficient hardware semaphore construction in which cores spin on local scratchpad memory, reducing the load on the shared bus. Cesare Ferri, Amber Viescas, Tali Moreshet, R. Iris Bahar, Maurice Herlihy |
ACM Great Lakes Symposium on VLSI | 4 |
| 2008 | Reducing leakage power by accounting for temperature inversion dependence in dual-Vt synthesized circuitsabstractThe effects of temperature on delay depend on several parameters, such as cell size, load, supply voltage, and threshold voltage. In particular, variations in Vth can yield a temperature inversion effect causing a decreases of cell delay as temperature increases. This phenomenon, besides affecting timing analysis of a design, has important and unforeseeable consequences on power optimization techniques. In this paper, we focus on the impact of such effects on multi-Vt design; in particular, we show how traditional dual-Vt optimization may yield timing errors in circuits by ignoring temperature effects. Moreover, we present a temperature-aware dual-Vt optimization technique that reduces leakage power and can guarantee that the circuit is timing feasible at the boundary temperatures provided by the technology library. Our experiments show an average 27% leakage reduction with respect to a non temperature-aware design flow. Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino |
ISLPED | 2 |
| 2008 | Using Implications for Online Error DetectionabstractIn this paper, we investigate the use of logic implications for the online detection of intermittent faults and hard-to-detect manufacturing defects. We present techniques to efficiently identify the most powerful circuit implications that can be checked for violations so that the fraction of errors detected can be maximized while minimizing the additional hardware overhead. Importantly, our approach does not require re-synthesis of the targeted logic; the checker logic is added off the critical path and is run in parallel with the regular control logic. Trade-offs can be easily made between additional coverage of errors and additional area overhead. Our results show that significant error detection is possible - even with only a 10% area overhead. Kundan Nepal, Nuno Alves, Jennifer Dworak, R. Iris Bahar |
ITC | 4 |
| 2008 | Fast Measurement of the "Non-Deterministic Zone" in Microprocessor Debug Using Maximum Likelihood EstimationabstractSpeed debug is a critical part of microprocessor diagnosis and debug. During this stage, the test engineer must determine and increase the maximum speed at which the processor can run reliably. One of the difficulties of this stage is that the pass-fail boundary, in practice, is not abrupt, but rather encompasses a non-deterministic region of behavior. Accurately modeling this non-deterministic region is particularly important since it directly influences the amount of time needed for speed debug.Current chip debug efforts often rely on the use of brute force techniques to deduce the shape of the pass-fail boundary region. More specifically, speed debug (i.e., the process of finding and fixing critical paths that prevent a chip from running at a higher frequency) requires a thorough understanding of the width and shape of the pass-fail boundary region where the chip’s behavior is nondeterministic. In this paper, we propose a statistical method based on Maximum Likelihood Estimation (MLE) techniques to infer the underlying shape of the non-deterministic region. Our method was tested on pre-production Intel microprocessors and was successful in modeling the shape of the fuzz-region in substantially feweriterations compared to a brute force approach. Desta Tadesse, R. Iris Bahar, Joel Grodstein |
VTS | 2 |
| 2008 | Introduction to joint ACM JETC/TODAES special issue on new, emerging, and specialized technologiesabstractNo abstract available. R. Iris Bahar, Krishnendu Chakrabarty |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2008 | Parametric yield management for 3D ICs: Models and strategies for improvementabstractThree-Dimensional (3D) Integrated Circuits (ICs) that integrate die with Through-Silicon Vias (TSVs) promise to continue system and functionality scaling beyond the traditional geometric 2D device scaling. 3D integration also improves the performance of ICs by reducing the communication time between different chip components through the use of short TSV-based vertical wires. This reduction is particularly attractive in processors where it is desirable to reduce the access time between the main logic die and the L2 cache or the main memory die. Process variations in 2D ICs lead to a drop in parametric yield (as measured by speed, leakage and sales profits), which forces manufacturers to speed bin their chips and to sell slow chips at reduced prices. In this paper we develop a model to quantify the impact of process variations on the parametric yield of 3D ICs, and then we propose a number of integration strategies that use a graph-theoretic framework to maximize the performance, parametric yield and profits of 3D ICs. Comparing our proposed strategies to current yield-oblivious methods, it is demonstrated that it is possible to increase the number of 3D ICs in the fastest speed bins by almost 2×, while simultaneously reducing the number of slow ICs by 29.4%. This leads to an improvement in performance by up to 6.45% and an increase of about 12.48% in total sales revenue using up-to-date market price models. Cesare Ferri, Sherief Reda, R. Iris Bahar |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2008 | Introduction to joint ACM JETC/TODAES special issue on new, emerging, and specialized technologiesabstractNo abstract available. R. Iris Bahar, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2007 | Interactive presentation: Techniques for designing noise-tolerant multi-level combinational circuitsabstractAs CMOS technology downscales, higher noise levels, wider threshold variation, and low supply voltage will force designers to contend with high rates of soft logical errors and many defective devices. A probabilistic design framework based on Markov random fields (MRF) has been previously proposed to address dynamic fault and noise vulnerability of ultimate digital CMOS circuitry. The idea is to use additional transistors and feedback loops to achieve significant noise immunity and ensure correct logic operations at low VDD. However, the extra reliability achieved in previously published work came at a cost of high transistor counts. In this paper, the authors present techniques to reduce the transistor count of larger multilevel combinational circuits built within the MRF framework by using variable sharing, implied dependence and supergates. Using these techniques the authors show an average reduction of approximately 28% in transistor counts over a range of combinational benchmark circuits built within the MRF framework compared to the best previously published results Kundan Nepal, R. Iris Bahar, Joseph L. Mundy, William R. Patterson, Alexander Zaslavsky |
DATE | 2 |
| 2007 | Accurate timing analysis using SAT and pattern-dependent delay modelsabstractAccurate delay modeling beyond static models is critical to garnering better correlation with post-silicon analysis. Furthermore, post-silicon timing validation requires a pattern-dependent timing model to generate patterns. To address these issues, a timing analysis tool was proposed that integrates a data-dependent delay model into its analysis. The approach solves for the delay by using the concept of circuit unrolling and formulation of timing questions as decision problems for input into a SAT solver. The effectiveness and validity of the proposed methodology is illustrated through experiments on benchmark circuits Desta Tadesse, D. Sheffield, E. Lenge, R. Iris Bahar, Joel Grodstein |
DATE | 4 |
| 2007 | Strategies for improving the parametric yield and profits of 3D ICsabstractThree-Dimensional (3D) Integrated Circuits (ICs) that integrate die with Through-Silicon Vias (TSVs) promise to continue system and functionality scaling beyond the traditional geometric 2D device scaling. 3D integration also improves the performance of ICs by reducing the communication time between different chip components through the use of short TSV-based vertical wires. This reduction is particularly attractive in processors where it is desirable to reduce the access time between the main logic die and the L2 cache or the main memory die. Process variations in 2D ICs lead to a drop in parametric yield (as measured by speed, leakage and sales profits), which forces manufacturers to speed bin their chips and to sell slow chips at reduced prices. In this paper we develop a model to quantify the impact of process variations on the parametric yield of 3D ICs, and then we propose a number of integration strategies that use a graph-theoretic framework to maximize the performance, parametric yield and profits of 3D ICs. Comparing our proposed strategies to current yield-oblivious methods, it is demonstrated that it is possible to increase the number of 3D ICs in the fastest speed bins by almost 2×, while simultaneously reducing the number of slow ICs by 29.4%. This leads to an improvement in performance by up to 6.45% and an increase of about 12.48% in total sales revenue using up-to-date market price models. Cesare Ferri, Sherief Reda, R. Iris Bahar |
ICCAD | 3 |
| 2007 | Designing Nanoscale Logic Circuits Based on Markov Random Fields
Kundan Nepal, R. Iris Bahar, Joseph L. Mundy, William R. Patterson, Alexander Zaslavsky |
J. Electron. Test. | 2 |
| 2006 | A cost-effective implementation of an ECC-protected instruction queue for out-of-order microprocessorsabstractMajor sources of transient errors in microprocessors today include noise and single event upsets. As feature sizes and voltages are reduced to create faster, more efficient, and computationally more powerful processors, these errors will increase significantly. We show that (contrary to conventional wisdom) error correction codes (ECC) can be efficiently utilized to handle these errors as instructions are being processed through the microprocessor pipeline. We will analyze some of the tradeoffs involved in a hardware implementation of ECC for the instruction queue with respect to performance, power, area, and reliability. Specifically, for an environment with high error rates, we show that we can correct all single bit errors with a negligible drop in performance. Our approach can be generalized to other data structures within the microprocessor, including the register file and reorder buffer. Vladimir Stojanovic, R. Iris Bahar, Jennifer Dworak, Richard Weiss 0001 |
DAC | 2 |
| 2006 | Designing MRF based error correcting circuits for memory elementsabstractAs devices are scaled to the nanoscale regime, it is clear that future nanodevices will be plagued by higher soft error rates and reduced noise margins. Traditional implementations of error correcting codes (ECC) can add to the reliability of systems but can be ineffective in highly noisy operating conditions. This paper proposes an implementation of ECC based on the theory of Markov random fields (MRF). The MRF probabilistic model is mapped onto CMOS circuitry, using feedback between transistors to reinforce the correct joint probability of valid logical states. We show that our MRF approach provides superior noise immunity for memory systems that operate under highly noisy conditions Kundan Nepal, R. Iris Bahar, Joseph L. Mundy, William R. Patterson, Alexander Zaslavsky |
DATE | 2 |
| 2006 | Optimizing noise-immune nanoscale circuits using principles of Markov random fieldsabstractAs CMOS devices and operating voltages are scaled down, noise and defective devices will impact the reliability of digital circuits. Probabilistic computing compatible with CMOS offers a possible solution. In this work, we present a new area and power efficient design methodology for the implementation of a probabilistic framework into CMOS technology based on Markov Random Fields (MRF). Using SPICE, we simulate elementary logic components and sample circuits from the MCNC'91 benchmark set and show the area and power benefits compared to older MRF mapping strategies. We also extend our area and power efficient approach to improving the design of a Hamming decoder based on MRF principles. Kundan Nepal, R. Iris Bahar, Joseph L. Mundy, William R. Patterson, Alexander Zaslavsky |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | Trends and Future Directions in Nano Structure Based Computing and FabricationabstractAs silicon CMOS devices are scaled down into the nanoscale regime, new challenges at both the device and system level are arising. While some of these challenges will be overcome in the near future, nanoscale devices will have high manufacturing defect rates and will operate at reduced noise margins, exposing computation to higher soft error rates. Thus, a key challenge for the future will be building fault and defect-tolerant computing systems. Researchers are looking to develop hybrid systems that combine on the same chip CMOS-based circuitry with any number of alternatives, including circuits composed of nanowire or carbon nanotube devices. The big advantage of including these new devices on the same chip is the increased device densities, and potential drop in fabrication costs. On the other hand, integrating very large numbers of devices on a single chip leads to questions of how to manage so many devices with tight constraints on cost, performance, power, and reliability, without having it become a design complexity nightmare. In this paper, we review some key issues and trends arising from nanostructure based computing and fabrication, while providing a few examples of defect-tolerant circuits and architectures currently being proposed as alternatives to "traditional" computing based exclusively on CMOS technology. These include hybrid nanowire/CMOS designs, reconfigurable or redundant architectures, and designs based on probabilistic computing. We end with a discussion on future challenges and direction in nanoscale computing. R. Iris Bahar |
ICCD | 1 |
| 2006 | Energy implications of multiprocessor synchronizationabstractNo abstract available. Tali Moreshet, R. Iris Bahar, Maurice Herlihy |
SPAA | 2 |
| 2006 | Timing analysis for full-custom circuits using symbolic DC formulationsabstractSuccessful analysis of high-speed integrated circuits requires accurate delay computation. A number of delay models have been developed; however, none can claim to be truly robust in the face of large channel-connected regions (CCRs) with input "exclusivity" constraints. A good circuit-level delay model should: 1) consider input exclusivity constraints; 2) handle a wide range of circuit structures; and 3) have a robust underlying framework that can be applied independent of the actual device model. We present a symbolic timing analysis tool that aims to address these three goals. It uses algebraic decision diagrams (ADDs) to estimate delay within a CCR as a function of its inputs while easily handling Boolean input constraints. It starts with a simple linear resistor model for transistors and from there apply various heuristics to improve the delay estimation without altering the symbolic algorithms. It analyzes delay with simple series-parallel reduction when possible and use symbolic matrix techniques to handle more complex circuit structures. The effectiveness of our approach is demonstrated on circuits from industry used in the Alpha 21264 and 21364 instead of the usual International Symposium on Circuits and Systems (ISCAS) or Microelectronics Center of North Carolina (MCNC) benchmarks. Our delay estimates are within 10% of simulation program with integrated circuits emphasis (SPICE) for over 90% of the circuits we simulated. This difference can translate into significant savings in manpower by avoiding the need to verify many unrealizable worst case conditions with other, more costly, simulation techniques Hui-Yuan Song, Kundan Nepal, R. Iris Bahar, Joel Grodstein |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Designing logic circuits for probabilistic computation in the presence of noiseabstractAs Si CMOS devices are scaled down into the nanoscale regime, current computer architecture approaches are reaching their practical limits. Future nano-architectures will confront devices and interconnections with a large number of inherent defects, which motivates the search for new architectural paradigms. In this paper, we examine probabilistic-based design methodologies for nanoscale computer architectures based on Markov random fields (MRF). The MRF approach can express arbitrary logic circuits and the logic operation is achieved by maximizing the probability of correct state configurations in the logic network depending on the interaction of neighboring circuit nodes. The computation proceeds via probabilistic propagation of states through the circuit. Crucially, the MRF logic can be implemented in modified CMOS-based circuitry that trades off circuit area and operation speed for the crucial fault tolerance and noise immunity. This paper builds on the recent demonstration that significant immunity to faulty individual devices or dynamically occurring signal errors can be achieved by the propagation of state probabilities over an MRF network. In particular, we are interested in CMOS-based circuits that work reliably at very low supply voltages (VDD = 0.1–0.2 V), where standard CMOS would fail due to thermal and crosstalk noise, and transistor threshold variation. In this paper, we present results for simulated probabilistic test circuits for elementary logic components and well as small circuits taken from the MCNC91 benchmark suite and we show greatly improved noise immunity operating at very low VDD. The MRF framework extends to all levels of a design, where formally optimum probabilistic computation can be implemented as a natural element of the processing structure. Kundan Nepal, R. Iris Bahar, Joseph L. Mundy, William R. Patterson, Alexander Zaslavsky |
DAC | 2 |
| 2005 | Energy reduction in multiprocessor systems using transactional memoryabstractThe emphasis in microprocessor design has shifted from high performance, to a combination of high performance and low power. Until recently, this trend was mostly true for uniprocessors. In this work we focus on new energy consumption issues unique to multiprocessor systems: synchronization of accesses to shared memory. We investigate and compare different means of providing atomic access to shared memory, including locks and lock-free synchronization (i.e., transactional memory), with respect to energy as well as performance. We show that transactional memory has an advantage in terms of energy consumption over locks, but that this advantage largely depends on the system architecture, the contention level, and the policy of conflict resolution Tali Moreshet, R. Iris Bahar, Maurice Herlihy |
ISLPED | 2 |
| 2005 | Symbolic failure analysis of complex CMOS circuits due to excessive leakage current and charge sharingabstractAs process geometries shrink, leakage currents and charge sharing are becoming increasingly critical problems, especially in full-custom circuit designs. Excessive leakage or charge sharing may cause functional failure at some or all operating conditions. Traditional circuit-analysis techniques may be used to verify if leakage currents are within allowable limits so as not to cause functional failures; however, unless the analysis takes into account specific input constraints for the circuit, the results may be overly pessimistic. Similar limitations exist for charge sharing. In this paper, we approach this verification problem symbolically using algebraic decision diagrams (ADDs). Using ADDs allows us to efficiently analyze leakage and charge sharing within a channel-connected region (CCR) as a function of its inputs. Exclusivity constraints are easily included in the analysis, thus allowing for more accurate (and less pessimistic) results. Our approach is general and can be applied to any arbitrary circuit structure, including a mesh. The effectiveness of our approach is demonstrated on circuits from industry used in the Alpha 21264 and 21364 instead of the usual international symposium on circuits and systems or Microelectronics Center of North Carolina benchmarks. We show that such an analysis can lead to up to a 90% difference in worst-case voltage drop. This difference can translate into significant savings in manpower by avoiding the need to verify many unrealizable worst-case conditions with other, more costly, simulation techniques. R. Iris Bahar, Hui-Yuan Song, Kundan Nepal, Joel Grodstein |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | RESTA: a robust and extendable symbolic timing analysis toolabstractSuccessful timing analysis for high-speed integrated circuits requires accurate delay computation. However, full-custom circuits popular in today's CPU designs make this difficult. A good circuit-level static timing analysis tool should 1) consider both internally or externally specified input constraints; 2) handle a wide range of circuit structures; and 3) have a robust underlying framework that can be applied independent of the actual device model. In this paper, we present RESTA, a Robust and Extendable Symbolic Timing Analysis tool that aims to address these three goals. RESTA estimates the delay for all valid input assignments, while naturally handling input constraints. We start with a simple linear resistor model for transistors and from there apply various heuristics to improve the delay estimation for the circuits without altering the symbolic algorithms. Our worst-case delay estimates are within 10% of SPICE for over 90% of the circuits we simulated. Kundan Nepal, Hui-Yuan Song, R. Iris Bahar, Joel Grodstein |
ACM Great Lakes Symposium on VLSI | 3 |
| 2004 | Reducing Issue Queue Power for Multimedia Applications using a Feedback Control AlgorithmabstractWe propose a dynamic power-aware issue queue in a general-purpose microprocessor for multimedia applications. Resources can be adapted at runtime in accordance with the feedback from a formal closed-loop control system such that power dissipation is reduced while still meeting the real time target for multimedia applications. We partition the issue queue into multiple sets (i.e., FIFOs) such that only instructions at the head of each set are able to issue. We then dynamically reconfigure the issue queue by changing the number and/or size of FIFOs to satisfy specific application's needs. The optimal configuration and timing are predicted by the closed-loop control system. Our results show an average power savings of over 90% in the issue queue with a negligible affect on the performance. Yu Bai 0001, R. Iris Bahar |
ICCD | 2 |
| 2004 | Fetch Halting on Critical Load MissesabstractAs the performance gap between processors and memory systems increases, the CPU spends more time stalled waiting for data from main memory. Critical long latency instructions, such as loads that miss to main memory and floating point arithmetic operations, are primarily responsible for these stalls. We present a technique, Fetch Halting that suspends instruction fetching when the processor is stalled by a critical long latency instruction. This enables us to save power in one of the primary sources of power dissipation, the issue logic. By reducing the occupancy rates in the issue queue and reorder buffer, we save power by disabling a large number of unused queue entries. In order to characterize critical instructions, our approach combines software profiling and hardware monitoring techniques. Statistical profiling information obtained from sample runs is used to identify critical instructions while hardware cache-miss prediction is used to monitor these instructions. We show that, on average, Fetch Halting can reduce issue queue and reorder buffer occupancy rates by 17.2% and 23.4%, respectively, with an average performance loss of only 4.6%. Nikil Mehta, Brian Singer, R. Iris Bahar, Michael Leuchtenburg, Richard Weiss 0001 |
ICCD | 3 |
| 2004 | A low-power in-order/out-of-order issue queueabstractTo better address power concerns, a good design strategy should be flexible enough to dynamically reconfigure available resources according to the application's needs such that extra power is dissipated only when it is really needed. In this work, we focus on power-aware solutions for the issue queue (IQ) in an out-of-order superscalar processor. We propose two schemes that partition the IQ into FIFOs such that only the instructions at the head of each FIFO may request to issue. We then monitor the processor and dynamically vary the number and/or size of FIFOs in accordance with utilization. Experimenting with two different distributions in power dissipation, we show up to 69% reduction in power dissipation in the wakeup and arbitration loop, while constraining performance degradation to be no more than 5%. Yu Bai 0001, R. Iris Bahar |
ACM Trans. Archit. Code Optim. | 2 |
| 2004 | Effects of speculation on performance and issue queue designabstractCurrent trends in microprocessor designs indicate increasing pipeline depth in order to keep up with higher clock frequencies and increased architectural complexity. Speculatively issued instructions are particularly sensitive to increases in pipeline depth. In this brief, we use load hit speculation as an example, and evaluate its cost effectiveness as pipeline depth increases. Our results indicate that as pipeline depth increases, speculation is more essential for performance but can drastically alter the utilization of pipeline resources, particularly the issue queue. We propose an alternative, more cost-effective design that takes into consideration the different issue queue utilization demands without degrading overall processor performance. Tali Moreshet, R. Iris Bahar |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Power-aware issue queue design for speculative instructionsabstractSpeculatively issued instructions may be particularly sensitive to increases in pipeline depth. Our results indicate that as pipeline depth increases, speculation increases the percentage of issue queue instructions that are waiting to be potentially re-issued in case of a mis-speculation. To compensate, issue queues are larger and thus more power hungry. We propose an alternative design called the Dual Issue Queue, that retains pre- and post-issue instructions in separate, smaller queues, saving 18% of issue queue power dissipation without degrading performance. Tali Moreshet, R. Iris Bahar |
DAC | 2 |
| 2003 | A Probabilistic-Based Design Methodology for Nanoscale Computation
R. Iris Bahar, Joseph L. Mundy, Jie Chen 0002 |
ICCAD | 1 |
| 2003 | Symbolic Failure Analysis of Custom Circuits due to Excessive Leakage CurrentabstractAs process geometries shrink, leakage current is becoming an increasingly critical problem, especially in full-custom circuit designs. Excessive leakage may cause functional failure at some or all operating conditions. Traditional circuit analysis techniques may be used to verify if leakage currents are within allowable limits so as not to cause functional failures; however, unless the analysis takes into account specific input constraints for the circuit, the results may be overly pessimistic. We approach this noise analysis problem symbolically using algebraic decision diagrams (ADDs). Using ADDs allows us to efficiently analyze leakage within a channel-connected region (CCR) as a function of its inputs. Exclusivity constraints are easily included in the analysis, thus allowing for more accurate (and less pessimistic) results. Our approach is general and can be applied to any arbitrary circuit structure, including a mesh. The effectiveness of our approach is demonstrated on circuits from industry used in Alpha 21264 and 21364 instead of the usual ISCAS benchmarks. We show that such an analysis can lead to up to a 90% difference in worst case voltage drop. This difference can translate into significant savings in manpower by avoiding the need to verify many unrealizable worst-case conditions with other, more costly, simulation techniques. Hui-Yuan Song, S. Bohidar, R. Iris Bahar, Joel Grodstein |
ICCD | 3 |
| 2001 | Power and energy reduction via pipeline balancingabstractMinimizing power dissipation is an important design requirement for both portable and non-portable systems. In this work, we propose an architectural solution to the power problem that retains performance while reducing power. The technique, known as Pipeline Balancing (PLB), dynamically tunes the resources of a general purpose processor to the needs of the program by monitoring performance within each program. We analyze metrics for triggering PLB, and detail instruction queue design and energy savings based on an extension of the Alpha 21264 processor. Using a detailed simulator, we present component and full chip power and energy savings for single and multi-threaded execution. Results show an issue queue and execution unit power reduction of up to 23% and 13%, respectively, with an average performance loss of 1% to 2%. R. Iris Bahar, Srilatha Manne |
ISCA | 1 |
| 2000 | Power optimization of technology-dependent circuits based on symbolic computation of logic implicationsabstractThis paper presents a novel approach to the problem of optimizing combinational circuits for low power. The method is inspired by the fact that power analysis performed on a technology mapped network gives more realistic estimates than it would at the technology-independent level. After each node's switching activity in the circuit is determined, high-power nodes are eliminated through redundancy addition and removal. To do so, the nodes are sorted according to their switching activity, they are considered one at a time, and learning is used to identify direct and indirect logic implications inside the network. These logic implications are exploited to add gates and connections to the circuit; this may help in eliminating high-power dissipating nodes, thus reducing the total switching activity and power dissipation of the entire circuit. The process is iterative; each iteration starts with a different target node. The end result is a circuit with a decreased switching power. Besides the general optimization algorithm, we propose a new BDD-based method for computing satisfiability and observability implications in a logic network; futhermore, we present heuristic techniques to add and remove redundancy at the technology-dependent level, that is, restructure the logic in selected places without destroying the topology of the mapped circuit. Experimental results show the effectiveness of the proposed technique. On average, power is reduced by 34%, and up to a 64% reduction of power is possible, with a negligible increase in the circuit delay. R. Iris Bahar, Ernest T. Lampe, Enrico Macii |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 1999 | The Non-Critical Buffer: Using Load Latency Tolerance to Improve Data Cache EfficiencyabstractData cache performance is critical to overall processor performance as the latency gap between CPU core and main memory increases. Studies have shown that some loads have latency demands that allow them to be serviced from slower portions of memory, thus allowingmore critical data to be kept in higher levels of the cache. We provide a strategy for identifying this latency-tolerant data at runtime and, using simple heuristics, keep it out of the main cache and place it instead in a small, parallel, associative buffer. Using such a "Non-Critical Buffer" dramatically improves the hit rate for more critical data, and leads to a performance improvement comparable to or better than other traditional cache improvement schemes. IPC improvements of over 4% are seen for some benchmarks. 1. Introduction The performance increase of today's high-end microprocessors is due to many factors, among them is the use of speculative, out-of-order execution with highly accurate branch prediction. Branch p... Brian R. Fisk, R. Iris Bahar |
ICCD | 2 |
| 1998 | Power and performance tradeoffs using various caching strategiesabstractIn this paper, we propose several different data and instruction cache configurations and analyze their power as well as performance implications on the processor. Unlike most existing work in low power microprocessor design, we explore a high performance processor with the latest innovations for performance. Using a detailed, architectural-level simulator, we evaluate full system performance using several different power/performance sensitive cache configurations such as increasing cache size or associativity and including buffers along side L1 caches. We then use the information obtained from the simulator to calculate the energy consumption of the memory hierarchy of the system. As an alternative to simply increasing cache associativity or size to reduce lower-level memory energy consumption (which may have a detrimental effect on on-chip energy consumption), we show that, by using buffers, energy consumption of the memory subsystem may be reduced by as much as 13% for certain data cache configurations and by as much as 23% for certain instruction cache configurations without adversely effecting processor performance or on-chip energy consumption. R. Iris Bahar, Gianluca Albera, Srilatha Manne |
ISLPED | 1 |
| 1997 | Algebraic Decision Diagrams and Their Applications
R. Iris Bahar, Erica A. Frohm, Charles M. Gaona, Gary D. Hachtel, Enrico Macii, Abelardo Pardo, Fabio Somenzi |
Formal Methods Syst. Des. | 1 |
| 1997 | Symbolic timing analysis and resynthesis for low power of combinational circuits containing false pathsabstractThis paper presents applications of algebraic decision diagrams (ADDs) to timing analysis and resynthesis for low power of combinational CMOS circuits. We first propose a symbolic algorithm to perform true delay calculation of a technology mapped network; the procedure we propose, implemented as an extension of the SIS synthesis system, is able to provide more accurate timing information than any other method presented so far; in particular, it is able to compute and store the arrival times of all the gates of the circuit for all possible input vectors, as opposed to the traditional methods which consider only the worst case primary inputs combination. Furthermore, the approach does not require any explicit false path elimination. We then extend our timing analysis tool to the symbolic calculation of required times and slacks, and we use this information to perform resynthesis for low power of the circuit by gate resizing. Our approach takes into account false paths naturally; in fact, it guarantees that resizing of the gates does not increase the true delay of the circuit, even in the presence of false paths. Our experiments have shown that many circuits, originally free of false paths, exhibit a large number of these false paths when optimized for area; therefore, the ability to deal with circuits containing false paths is of primary importance. We present experimental results for ADD-based and static timing analysis-based resynthesis, which clearly show that our tool is superior in the case of circuits containing false paths, but at the same time, it provides competitive results in the case of circuits which are free of false paths. R. Iris Bahar, Gary D. Hachtel, Enrico Macii, Fabio Somenzi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1996 | Symbolic computation of logic implications for technology-dependent low-power synthesisabstractThis paper presents a novel technique for re-synthesizing circuits for low-power dissipation. Power consumption is reduced through redundancy addition and removal by using learning to identify indirect logic implications within a circuit. Such implications are exploited by adding gates and connections to the circuit without altering its overall behavior and thereby enabling us to eliminate other, high power dissipating, nodes. We propose a new BDD-based method for computing indirect implications in a logic network; furthermore, we present heuristic techniques to perform redundancy addition and removal without destroying the topology of the mapped circuit. Experimental results show the effectiveness of the proposed technique in reducing power while keeping within delay and area constraints. R. Iris Bahar, M. Burns, Gary D. Hachtel, Enrico Macii, H. Shin, Fabio Somenzi |
ISLPED | 1 |
| 1995 | Computing the Maximum Power Cycles of a Sequential CircuitabstractThis paper studies the problem of estimating worst case power dissipation in a sequential circuit.We approach this problem by nding the maximum average weight cycles in a weighted directed g r aph.In order to handle practical sized examples, we use symbolic methods, based o n A lgebraic Decision Diagrams (ADDs), for computing the maximum average length cycles as well as the number of gate transitions in the circuit, which is necessary to construct the weighted directed g r aph. Srilatha Manne, Abelardo Pardo, R. Iris Bahar, Gary D. Hachtel, Fabio Somenzi, Enrico Macii, Massimo Poncino |
DAC | 3 |
| 1995 | Boolean techniques for low power driven re-synthesisabstractWe present a boolean technique to reduce power consumption of combinational circuits that have already been optimized for area and delay and then mapped onto a library of gates. In order to achieve a better optimization, we cluster gates by collapsing two or more levels of gates into a single node. When optimizing each cluster, our method extends the algorithms used in ESPRESSO, by adding heuristics that bias the minimization toward lowering the power dissipation in the circuit. The results of our method, on a number of benchmark circuits, show an average of 11% improvement in power savings compared to existing boolean techniques. R. Iris Bahar, Fabio Somenzi |
ICCAD | 1 |
| 1994 | An ADD-based algorithm for shortest path back-tracing of large graphsabstractSymbolic computation techniques play a fundamental role in logic synthesis and formal hardware verification algorithms. Recently, Algebraic Decision Diagrams, i.e., BDDs with a set of constant values different to the set /spl lcub/0,1/spl rcub/, have been used to solve general purpose problems, such as matrix multiplication, shortest path calculation, and solution of linear systems, as well as logic synthesis and formal verification problems, such as timing analysis, probabilistic analysis of finite state machines, and state space decomposition for approximate finite state machine traversal. ADD-based procedures for single-source and all-pairs shortest path weight calculation have appeared to be very effective for the manipulation of large graphs (over 10/sup 27/ vertices and 10/sup 36/ edges). However, for those procedures to be applicable to real problems, for example flow network problems, computing only shortest path weights is not enough; what it is needed is an algorithm that, given the weight of a shortest path between two vertices of a graph, actually determines the sequence of vertices belonging to the shortest path. This paper proposes a symbolic algorithm to execute shortest path back-tracing which exploits the compactness of the ADD data structure to handle large graphs.> R. Iris Bahar, Gary D. Hachtel, Abelardo Pardo, Massimo Poncino, Fabio Somenzi |
Great Lakes Symposium on VLSI | 1 |
| 1994 | A symbolic method to reduce power consumption of circuits containing false paths
R. Iris Bahar, Gary D. Hachtel, Enrico Macii, Fabio Somenzi |
ICCAD | 1 |
| 1993 | Algebraic decision diagrams and their applicationsabstractIn this paper we present theory and experiments on the algebraic decision diagrams (ADDs). These diagrams extend BDD's by allowing values from an arbitrary finite domain to be associated with the terminal nodes. We present a treatment founded in Boolean algebras and discuss algorithms and results in applications like matrix multiplication and shortest path algorithms. Furthermore, we outline possible applications of ADD's to logic synthesis, formal verification, and testing of digital systems. R. Iris Bahar, Erica A. Frohm, Charles M. Gaona, Gary D. Hachtel, Enrico Macii, Abelardo Pardo, Fabio Somenzi |
ICCAD | 1 |