VLDB 2026 Research / reviewers in the wild / expert
Mircea R. Stan
dblp:s/MirceaRStan · also Mircea Stan
· DBLP profile ↗
122ranked-venue papers
15as first author
17since 2021 · last 2026
0000-0003-0577-9976ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 116 · 13 first-author · 15 since 2021Software engineering, systems software and programming languages · 16 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NetLossBench: A Tiered Benchmark for GNN Hardware Trojan Detectors under Partial Netlist ObservationsabstractGraph neural networks for Hardware Trojan detection are typically evaluated on fully observed gate-level netlists, whereas practical settings often provide only partial observations due to reverse engineering limitations, extraction loss, and obfuscation. This makes robustness evaluation inherently challenging. We present NetLossBench, the first tiered robustness benchmark for budgeted stress testing under observation-layer information loss. NetLossBench defines three tiers of loss generation: random loss as a baseline, structured loss guided by topology-aware heuristics to approximate realistic observation loss, and adversarial loss guided by reinforcement learning to approximate worst-case degradation. Results show that random loss can overestimate robustness, while structured and adversarial losses reveal a lower robustness floor. Liangtao Dai, Yimin Gao, Melika Morsali, Mircea R. Stan |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Unlocking High-Performance Low-Power Adiabatic Logic Computing with Modern FinFET Technology Node
Mircea R. Stan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | EAS-CiM 2.0: Event-driven Asynchronous Stream-based Compute-in-Memory Kernels with Scalable PrecisionabstractTraditional edge accelerators struggle to balance performance, accuracy, and power consumption. Our solution leverages Asynchronous Stream Computing (ASC) to implement a fully asynchronous, clockless system where data is encoded in the frequency and duty cycle of streams. This architecture adapts dynamically to available energy resources, adjusting computational precision and speed in response to workload demands. A stimulus-driven model allows the system to process data efficiently when actionable information is present. Analytical and experimental results highlight a tunable trade-off between latency and precision, aligning system performance with real-time requirements. To validate these concepts, we implement a Compute-in-Memory (CiM) based dot product architecture to accelerate Word2Vec Algorithms for Natural Language Processing (NLP) tasks and have exhibited a relative semantic accuracy scaling of 47.81% and a computational efficiency scaling of 43.23x when scaling precision from 12 - 4 bits. Rahul Sreekumar, Melika Morsali, Naomi Solomon, Yimin Gao, Minseong Park, Kyusang Lee, R. K. Krishnamurthy, Mircea R. Stan |
ISCAS | 8 |
| 2025 | Editorial: Renewed Excellence for 2025-2026abstractI am happy and honored to have been reappointed as Editor in Chief (EiC) for the IEEE Transactions on VLSI Systems (TVLSI) for another two-year term. As I continue my efforts to improve the quality of the journal, I am grateful for the renewed trust placed in me by the three IEEE sponsoring societies (CASS, SSCS and CS) and by the VLSI community at large. Contrary to a feared slowdown due to increased difficulties with scaling, the field of Very Large Scale Integration (VLSI) has actually grown at an increasingly fast rate as it provides the hardware backbone for the insatiable AI applications which are taking over the world. The H100/200 GPUs, which are essential for AI training, are the largest “conventional” integrated circuits (IC) with 80 billion transistors, while the wafer-scale WSE2/3, which can provide significant improvements in AI inference, are absolute behemoths with 4 trillion transistors! Mr. Moore can be proud there in heaven for what our industry is able to deliver! Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | A Feedback Self-adaptive Body Biasing-based RF-DC Rectifier for Highly-sensitive RF Energy HarvestingabstractUsing radio frequency (RF) energy to power Internet of Things (IoT) devices over extended distances poses a challenge due to limited input power. As we move farther from the source, the intensity of radio waves diminishes according to an inverse square law. Implementing an internal feedback body biasing connection using flipped-well transistors can dynamically enhance the rectifier’s conductivity and eliminate the need for external circuitry. In this work, we propose a novel feedback self-adaptive body biasing rectifier for RF energy harvesting for the UHF band (902-928 MHz). Based on 22nm fully-depleted silicon-on-insulator (FDSOI) technology, this rectifier achieves the high sensitivity required to harvest weak signals at far distances and a competitive power conversion efficiency (PCE). Furthermore, we introduce a methodology for designing and optimizing a differential impedance matching network (IMN) and rectifier. Based on the simulation results, our proposed 5-stage rectifier exhibits a high sensitivity of -31.4dBm at 1V under a capacitive load, surpassing comparable results achieved by previous works. A good peak PCE of 31.3% under a 5MΩ load at the typical corner makes our rectifier an ideal choice for ultra-low power RF energy harvesting applications. Elisa Pantoja, Yimin Gao, Mircea R. Stan |
ISCAS | 4 |
| 2023 | Hardware Trojans in eNVM Neuromorphic DevicesabstractFast and energy-efficient execution of a DNN on traditional CPU- and GPU-based architectures is challenging due to excessive data movement and inefficient computation. Emerging non-volatile memory (eNVM)-based accelerators that mimic biological neuron computations in the analog domain have shown significant performance improvements. However, the potential security threats in the supply chain of such systems have been largely understudied. This work describes a hardware supply chain attack against analog eNVM neural accelerators by identifying potential Trojan insertion points and proposes a hardware Trojan design that stealthily leaks model parameters while evading detection. Our evaluation shows that such a hardware Trojan can recover over 90% of the synaptic weights. Lingxi Wu, Rahul Sreekumar, Rasool Sharifi, Kevin Skadron, Mircea R. Stan, Ashish Venkat |
DATE | 5 |
| 2023 | FreezeTime: Towards System Emulation through Architectural VirtualizationabstractHigh-end FPGAs enable architecture modeling through emulation with high speed and fidelity. However, the available reconfigurable logic and memory resources limit the size, complexity, and speed of the emulated target designs. The challenge is to map and model large and fast memory hierarchies, such as large caches and mixed main memory, various heterogeneous computation instances, such as CPUs, GPUs, AI/ML processing units and accelerator cores, and communication infrastructure, such as buses and networks. In addition to the spatial dimension, this work uses the temporal dimension, implemented with architectural multiplexing coupled with block-level synchronization, to model a complete system-on-chip architecture. Our approach presents mechanisms to abstract instance plurality while preserving timing in sync. With only a subset of the architecture on the FPGA, we freeze a whole emulated module's activity and state during the additional time intervals necessary for the action on the virtualized modules to elapse. We demonstrate this technique by emulating a hypothetical system consisting of a processor and an SRAM memory too large to map on the FPGA. For this, we modify a LiteX-generated SoC consisting of a VexRISC-V processor and DDR memory, with the memory controller issuing stall signals that freeze the processor, effectively ''hiding'' the memory latency. For Linux boot, we measure significant emulation vs. simulation speedup while matching RTL simulation accuracy. The work is open-sourced. Sergiu Mosanu, Joshua Fixelle, Kevin Skadron, Mircea R. Stan |
FPGA | 4 |
| 2023 | Design Space Exploration of Layer-Wise Mixed-Precision Quantization with Tightly Integrated Edge Inference UnitsabstractLayer-wise mixed-precision quantization (MPQ) has become prevailing for edge inference since it strikes a better balance between accuracy and efficiency compared to the uniform quantization scheme. Existing MPQ strategies either lacked hardware awareness or incurred huge computation costs, which gated their deployment at the edge. In this work, we propose a novel MPQ search algorithm that obtains an optimal scheme by "sampling" layer-wise sensitivity with respect to a newly proposed metric that incorporates both accuracy and proxy of hardware cost. To further efficiently deploy post-training MPQ on edge chips, we propose to tightly integrate the quantized inference units as part of the processor pipeline through micro-architecture and Instruction Set Architecture (ISA) co-design. Evaluation results show that the proposed search algorithm achieves 3% ~ 11% higher inference accuracy with similar hardware cost compared to the state-of-the-art MPQ strategies. In addition, the tightly integrated MPQ units achieve speedup of 15.13x ~ 29.65x compared to a baseline RISC-V processor. Xiaotian Zhao, Yimin Gao, Vaibhav Verma, Ruge Xu, Mircea R. Stan, Xinfei Guo |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | ISVABI: In-Storage Video Analytics Engine with Block InterfaceabstractThe wide use of cameras in the past decade has increased the need to process video data significantly. Due to the large volume of video data, analyzing videos to extract useful information has become a critical challenge. Several prior works have tried to accelerate video analytics workloads by offloading some operations to embedded processors within storage devices. Joshua Fixelle, Pingyi Huo, Mircea R. Stan, Michael P. Mesnier, Narayanan Vijaykrishnan |
LCTES | 4 |
| 2023 | Editorial New Beginnings for IEEE TVLSIabstractIt is with great enthusiasm and honor that I assumed the role of Editor-in-Chief (EiC) for IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) this year. As we embark on this journey, I am grateful for the trust placed in me by the three IEEE sponsoring societies [IEEE Circuits and Systems Society (CASS), IEEE Solid-State Circuits Society (SSCS), and IEEE Computer Society (CS)] and by the VLSI community at large. The field of very large-scale integration (VLSI) design has witnessed remarkable advancements in recent years. From the relentless pursuit of higher transistor densities to the integration of complex 2-D and 3-D systems, the VLSI landscape has evolved rapidly, enabling groundbreaking technologies that have transformed our lives. As we stand at the forefront of this transformative era, it is imperative for us to embrace the paradigm shift that awaits us. Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Editorial Rolling Out the IEEE TVLSI EDICSabstractVLSI Systems research represents a dynamic and expansive domain. Over the years, it has evolved into a creative fusion of theoretical exploration, integrated chip design, performance evaluation, and practical applications related to the wide areas of circuits and systems, computer hardware and solid-state circuits. In acknowledgment of the extensive influence of very large-scale integration (VLSI) systems and the diverse research trends within this field, we have embarked on an Editor’s Information Classification Scheme (EDICS) for IEEE Transactions on Very Large Scale Integration (VLSI) Systems. The primary purpose of these EDICS is to provide a comprehensive description of the focal points within TVLSI. Furthermore, they aid in the judicious allocation of papers to associate editors and reviewers who possess expertise in the specific subject matter of each submission. Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Impact of 3-D Integration on Thermal Performance of RISC-V MemPool Multicore SOCabstractDue to the rise in the number of cores in modern multicore architectures, 3-D integration (i.e., vertical stacking of chips) of system-on-a-chip (SOC) promises better performance due to a drastic reduction in global interconnect lengths and die footprint compared with 2-D counterparts. However, thermal issues are predominant in 3-D-SOCs due to the vertical stacking nature of chips which multiplies the transistor power density by the number of dies within the stack. Also, the reduced lateral heat spreading with aggressive die thinning degrades the ON-chip thermal performances. In this article, we investigate the thermal performance analysis of 3-D-SOC and compare the results with the 2-D-SOC designs for a MemPool multicore SOC with shared L1 scratchpad memory (SPM). Simulation results reveal that the 3-D-SOC using memory-on-logic (MOL) configuration increases the ON-chip maximum temperature by more than 20% compared with the baseline 2-D-SOC and the logic die temperature is relatively higher (3.6%) than the memory die. We also explore the impact of architectural floor-planning effects and 3-D functional partitioning on thermal performance of the MemPool instances in the 3-D-SOC with memory capacity ranging from 1 to 8 MiB and benchmarked the thermal performance with the 2-D-SOC designs. We observe that the junction-to-ambient temperature ($T_{\max }$) increases by 44% and is predominant for the SPM capacity of 8 MiB. Further investigations on various 3-D stacking configurations reveal there is an improvement in thermal performance for MOL over logic-on-memory (LOM) for L1 SPM capacity of 1, 2, and 4 MiB, and LOM over the MOL configuration for L1 SPM capacity of 8 MiB. Sankatali Venkateswarlu, Subrat Mishra, Herman Oprins, Bjorn Vermeersch, Moritz Brunion, Jun-Han Han, Mircea R. Stan, Dwaipayan Biswas, Pieter Weckx, Francky Catthoor |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2022 | PiMulator: a Fast and Flexible Processing-in-Memory Emulation PlatformabstractMotivated by the memory wall problem, researchers propose many new Processing-in-Memory (PiM) architectures to bring computation closer to data. However, evaluating the performance of these emerging architectures involves using a myriad of tools, including circuit simulators, behavioral RTL or software simulation models, hardware approximations, etc. It is challenging to mimic both software and hardware aspects of a PiM architecture using the currently available tools with high performance and fidelity. Until and unless actual products that include PiM become available, the next best thing is to emulate various hardware PiM solutions on FPGA fabric and boards. This paper presents a modular, parameterizable, FPGA synthesizable soft PiM model suitable for prototyping and rapid evaluation of Processing-in-Memory architectures. The PiM model is implemented in System Verilog and allows users to generate any desired memory configuration on the FPGA fabric with complete control over the structure and distribution of the PiM logic units. Moreover, the model is compatible with the LiteX framework, which provides a high degree of usability and compatibility with the FPGA and RISC-V ecosystem. Thus, the framework enables architects to easily prototype, emulate and evaluate a wide range of emerging PiM architectures and designs. We demonstrate strategies to model several pioneering bitwise-PiM architectures and provide detailed benchmark performance results that demonstrate the platform's ability to facilitate design space exploration. We observe an emulation vs. simulation weighted-average speedup of 28× when running a memory benchmark workload. The model can utilize 100% BRAM and only 1% FF and LUT of an Alveo U280 FPGA board. The project is entirely open-source. Sergiu Mosanu, Mohammad Nazmus Sakib, Tommy Tracy II, Ersin Cukurtas, Alif Ahmed, Preslav Ivanov, Samira Manabi Khan, Kevin Skadron, Mircea R. Stan |
DATE | 9 |
| 2022 | Gearbox: a case for supporting accumulation dispatching and hybrid partitioning in PIM-based acceleratorsabstractProcessing-in-memory (PIM) minimizes data movement overheads by placing processing units near each memory segment. Recent PIMs employ processing units with a SIMD architecture. However, kernels with random accesses, such as sparse-matrix-dense-vector (SpMV) and sparse-matrix-sparse-vector (SpMSpV), cannot effectively exploit the parallelism of SIMD units because SIMD's ALUs remain idle until all the operands are collected from local memory segments (memory segment attached to the processing unit) or remote memory segments (other segments of the memory). Marzieh Lenjani, Alif Ahmed, Mircea R. Stan, Kevin Skadron |
ISCA | 3 |
| 2022 | ISKEVA: in-SSD key-value database engine for video analytics applicationsabstractKey-value databases are widely used to store the features or metadata generated from the neural network based video processing platforms. Due to the large volumes of video data, these databases use solid state drives (SSDs) as the primary data storage platform, and user query-based filtering, and retrieval operations on data incur large volume of data movement between the SSD and the host processor. In this paper, we present an in-SSD key-value database which uses the embedded CPU core, and DRAM memory on the SSD to support various queries with predicates and reduce the data movement between SSD and host processor significantly. We augment the SSD flash translation layer with key-value database functions and auxiliary data structures to support the user queries using the embedded core and DRAM memory on SSD. The proposed key-value store prototype on the Cosmos plus OpenSSD board reduces data movement between host processor and SSD by 14.57x, achieves an application-level speedup by 1.16x, and reduced energy consumption by 56% across different types of user queries. Joshua Fixelle, Nagadastagiri Challapalle, Pingyi Huo, Zhaoyan Shen, Zili Shao, Mircea R. Stan, Narayanan Vijaykrishnan |
LCTES | 7 |
| 2022 | Agile-AES: Implementation of configurable AES primitive with agile design approach
Xinfei Guo, Mohamed El-Hadedy 0001, Sergiu Mosanu, Xiangdong Wei, Kevin Skadron, Mircea R. Stan |
Integr. | 6 |
| 2022 | Thermal Performance Analysis of Mempool RISC-V Multicore SoCabstractThe presence of multiple cores in modern multicore architectures makes thermal management and temperature estimation a really challenging task for enhancing reliability and lifespan. Due to the presence of many cores, the core/tile spacing needs to be optimized in order to enhance the thermal coupling between interconnect routing blocks and active tiles. In addition, the tiles activity patterns under partial workload conditions significantly affect the maximum on-chip temperature which results in nonuniform temperature distribution. This is due to poor thermal coupling between neighboring tiles owing to the decrease in spacing between cores. In this article, we investigate the thermal performance analysis of a 256-core (i.e., 64 tiles) Mempool reduced instruction set computer (RISC) V-based architecture considering the impact of inter tiles spacing. Simulation results reveal that lateral heat spreading predominantly affects the thermal performance in multicore architectures under partial workload conditions. We also optimize the thermal performance with different tiles activity pattern. Simulation results reveal that both the maximum on-chip temperature and lateral heat spreading are improved for specific tiles activity patterns. Also the thermal performance analysis considering the “tile-insite effect” reveals that there is little impact on on-chip maximum temperature ($T_{\text {max}}$), but the on-chip thermal gradient ($\Delta T$) and the thermal profile pattern are predominantly affected. Finally, the effect of the secondary heat path toward printed circuit board (PCB) is studied in this work. Sankatali Venkateswarlu, Subrat Mishra, Herman Oprins, Bjorn Vermeersch, Moritz Brunion, Jun-Han Han, Mircea R. Stan, Pieter Weckx, Francky Catthoor |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2020 | FlexAmata: A Universal and Efficient Adaption of Applications to Spatial Automata Processing AcceleratorsabstractPattern matching, especially for complex patterns with many variations, is an important task in many big-data applications and maps well to finite automata. Recently, a variety of research has focused on hardware acceleration of automata processing, especially via spatial architectures that directly map the patterns to massively parallel hardware elements, such as in FPGAs and in-memory solutions. We observed that all existing automata-acceleration architectures are designed based on fixed, 8-bit symbol processing, derived from ASCII processing. However, the alphabet size in pattern-matching applications varies from just a few up to billions of unique symbols. This makes it difficult to provide a universal and efficient mapping of this wide variety of automata applications to existing automata accelerators. Elaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea R. Stan, Kevin Skadron |
ASPLOS | 4 |
| 2020 | Nano-Crossbar based Computing: Lessons Learned and Future DirectionsabstractIn this paper, we first summarize our research activities done through our European Union’s Horizon-2020 project between 2015 and 2019. The project has a goal of developing synthesis and performance optimization techniques for nanocrossbar arrays. For this purpose, different computing models including diode, memristor, FET, and four-terminal switch based models, within different technologies including carbon nanotubes, nanowires, and memristors as well as the CMOS technology have been investigated. Their capabilities to realize logic functions and to tolerate faults have been deeply analyzed. From these experiences, we think that instead of replacing CMOS with a completely new crossbar based technology, developing CMOS compatible crossbar technologies and computing models is a more viable solution to overcome challenges in CMOS miniaturization. At this point, four-terminal switch based arrays, called switching lattices, come forward with their CMOS compatibility feature as well as with their area efficient device and circuit realizations. We have showed that switching lattices can be efficiently implemented using a standard CMOS process to implement logic functions by doing experiments in a 65nm CMOS process. Further in this paper, we make an introduction of realizing memory arrays with switching lattices including ROMs and RAMs. Also we discuss challenges and promises in realizing switching lattices for under 30nm CMOS technologies including FinFET technologies. Mustafa Altun, Ismail Cevik, Ahmet Can Erten, Osman Eksik, Mircea R. Stan, Csaba Andras Moritz |
DATE | 5 |
| 2020 | Grapefruit: An Open-Source, Full-Stack, and Customizable Automata Processing on FPGAsabstractRegular expressions have been widely used in various application domains such as network security, machine learning, and natural language processing. Increasing demand for accelerated regular expressions, or equivalently finite automata, has motivated many efforts in designing FPGA accelerators. However, there is no framework that is publicly available, comprehensive, parameterizable, general, full-stack, and easy-touse, all in one, for design space exploration for a wide range of growing pattern matching applications on FPGAs. In this paper, we present Grapefruit, the first open-source, full-stack, efficient, scalable, and extendable automata processing framework on FPGAs. Grapefruit is equipped with an integrated compiler with many parameters for automata simulation, verification, minimization, transformation, and optimizations. Our modular and standard design allows researchers to add capabilities and explore various features for a target application. Our experimental results show that the hardware generated by Grapefruit performs 9%80% better than prior work that is not fully end-to-end and has 3.4 × higher throughput in a multi-stride solution than a single-stride solution. Reza Rahimi, Elaheh Sadredini, Mircea R. Stan, Kevin Skadron |
FCCM | 3 |
| 2020 | Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical In-Situ AcceleratorsabstractIn-situ approaches process data very close to the memory cells, in the row buffer of each subarray. This minimizes data movement costs and affords parallelism across subarrays. However, current in-situ approaches are limited to only row-wide bitwise (or few-bit) operations applied uniformly across the row buffer. They impose a significant overhead of multiple row activations for emulating 32-bit addition and multiplications using bitwise operations and cannot support operations with data dependencies or based on predicates. Moreover, with current peripheral logic, communication among subarrays is inefficient, and with typical data layouts, bits in a word are not physically adjacent. The key insight of this work is that in-situ, single-word ALUs outperform in-situ, parallel, row-wide, bitwise ALUs by reducing the number of row activations and enabling new operations and optimizations. Our proposed lightweight access and control mechanism, Fulcrum, sequentially feeds data into the single-word ALU and enables operations with data dependencies and operations based on a predicate. For algorithms that require communication among subarrays, we augment the peripheral logic with broadcasting capabilities and a previously-proposed method for low-cost inter-subarray data movement. The sequential processor also enables overlapping of broadcasting and computation, and reuniting bits that are physically adjacent. In order to realize true subarray-level parallelism, we introduce a lightweight column-selection mechanism through shifting one-hot encoded values. This technique enables independent column selection in each subarray. We integrate Fulcrum with Compress Express Link (CXL), a new interconnect standard. Fulcrum with one memory stack delivers on average (up to) 23.4 (76) speedup over a server-class GPU, NVIDIA P100, with three stacks of HBM2 memory, (ii) 70 (228) times speedup per memory stack over the GPU, and (iii) 19 (178.9) times speedup per memory stack over an ideal model of the GPU, which only accounts for the overhead of data movement. Marzieh Lenjani, Patricia Gonzalez-Guerrero, Elaheh Sadredini, Shuangchen Li, Yuan Xie 0001, Ameen Akel, Sean Eilert, Mircea R. Stan, Kevin Skadron |
HPCA | 8 |
| 2020 | Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern MatchingabstractHigh-throughput and concurrent processing of thousands of patterns on each byte of an input stream is critical for many applications with real-time processing needs, such as network intrusion detection, spam filters, virus scanners, and many more. The demand for accelerated pattern matching has motivated several recent in-memory accelerator architectures for automata processing, which is an efficient computation model for pattern matching. Our key observations are: (1) all these architectures are based on 8-bit symbol processing (derived from ASCII), and our analysis on a large set of real-world automata benchmarks reveals that the 8-bit processing dramatically under-utilizes hardware resources, and (2) multi-stride symbol processing, a major source of throughput growth, is not explored in the existing in-memory solutions. This paper presents Impala, a multi-stride in-memory automata processing architecture by leveraging our observations. The key insight of our work is that transforming 8-bit processing to 4-bit processing exponentially reduces hardware resources for state-matching and improves resource utilization. This, in turn, brings the opportunity to have a denser design, and be able to utilize more memory columns to process multiple symbols per cycle with a linear increase in state-matching resources. Impala thus introduces threefold area, throughput, and energy benefits at the expense of increased offline compilation time. Our empirical evaluations on a wide range of automata benchmarks reveal that Impala has on average 2.7× (up to 3.7×) higher throughput per unit area and 1.22× lower power consumption than Cache Automaton, which is the best performing prior work. Elaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea R. Stan, Kevin Skadron |
HPCA | 4 |
| 2020 | Towards on-node Machine Learning for Ultra-low-power Sensors Using Asynchronous Σ Δ StreamsabstractWe propose a novel architecture to enable low-power, complex on-node data processing, for the next generation of sensors for the internet of things (IoT), smartdust, or edge intelligence. Our architecture combines near-analog-memory-computing (NAM) and asynchronous-computing-with-streams (ACS), eliminating the need for ADCs. ACS enables ultra-low power, massive computational resources required to execute on-node complex Machine Learning (ML) algorithms; while NAM addresses the memory-wall that represents a common bottleneck for ML and other complex functions. In ACS an analog value is mapped to an asynchronous stream that can take one of two logic levels ( v h , v l ). This stream-based data representation enables area/power-efficient computing units such as a multiplier implemented as an AND gate yielding savings in power of ∼90% compared to digital approaches. The generation of streams for NAM and ACS in a brute force manner, using analog-to-digital-converters (ADCs) and digital-to-streams-converters, would sky-rocket the power-latency-energy cost making the approach impractical. Our NAM-ACS architecture eliminates expensive conversions, enabling an end-to-end processing on asynchronous streams data-path. We tailor the NAM-ACS architecture for random forest (RaF), an ML algorithm, chosen for its ability to classify using a reduced number of features. Simulations show that our NAM-ACS architecture enables 75% of savings in power compared with a single ADC, obtaining a classification accuracy of 85% using an RaF-inspired algorithm. Patricia Gonzalez-Guerrero, Tommy Tracy II, Xinfei Guo, Rahul Sreekumar, Marzieh Lenjani, Kevin Skadron, Mircea R. Stan |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2020 | Low-Power, Highly Reliable Dynamic Thermal Management by Exploiting Approximate ComputingabstractWith the continuous downscaling of semiconductor processes, the growing power density and thermal issues in multicore processors become more and more challenging, thus reliable dynamic thermal management (DTM) is required to prevent severe challenges in system performance. The accuracy of the thermal profile, delivered to the DTM manager, plays a critical role in the efficiency and reliability of DTM, different sources of noise and variations in deep submicron (DSM) technologies severely affecting the thermal data that can lead to significant degradation of DTM performance. In this article, we propose a novel fault-tolerance scheme exploiting approximate computing to mitigate the DSM effects on DTM efficiency. Approximate computing in hardware design can lead to significant gains in energy efficiency, area, and performance. To exploit this opportunity, there is a need for design abstractions that can systematically incorporate approximation in hardware design which is the main contribution of our work. Our proposed scheme achieves 11.20% lower power consumption, 6.59% smaller area, and 12% reduction in the number of wires, while increasing DTM efficiency by 5.24%. Somayeh Rahimipour, Wameedh Nazar Flayyih, Noor Ain Kamsani, Shaiful J. Hashim, Mircea R. Stan, Fakhrul Z. Rokhani |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | Cross-Layer Resilience: Challenges, Insights, and the Road AheadabstractResilience to errors in the underlying hardware is a key design objective for a large class of computing systems, from embedded systems all the way to the cloud. Sources of hardware errors include radiation, circuit aging, variability induced by manufacturing and operating conditions, manufacturing test escapes, and early-life failures. Many publications have suggested that cross-layer resilience, where multiple error resilience techniques from different layers of the system stack cooperate to achieve cost-effective resilience, is essential for designing cost-effective resilient digital systems. This paper presents a comprehensive overview of cross-layer resilience by addressing fundamental cross-layer resilience questions, by summarizing insights derived from recent advances in cross-layer resilience research, and by discussing future cross-layer resilience challenges. Eric Cheng, Daniel Mueller-Gritschneder, Jacob A. Abraham, Pradip Bose, Alper Buyuktosunoglu, Deming Chen, Hyungmin Cho, Yanjing Li, Uzair Sharif, Kevin Skadron, Mircea R. Stan, Ulf Schlichtmann, Subhasish Mitra |
DAC | 11 |
| 2019 | Flexi-AES: A Highly-Parameterizable Cipher for a Wide Range of Design ConstraintsabstractInterconnected devices communicate efficiently and securely over untrusted networks via security protocols that employ various encryption algorithms, often as hardware modules. State-of-the-art hardware implementations typically focus on optimizing a single metric and are tedious to adapt to a wider set of design constraints. In this work, we develop an open-source, flexible and parameterizable hardware implementation of the Advanced Encryption Standard (AES). We present a feature-rich implementation in Chisel that is simple to employ to any architectures and to fine-tune to specific design requirements. Despite the larger design space, we use 50% fewer lines of code than existing Verilog versions, thus enabling a higher level of development productivity. Sergiu Mosanu, Xinfei Guo, Mohamed El-Hadedy 0001, Lorena Anghel, Mircea R. Stan |
FCCM | 5 |
| 2019 | eAP: A Scalable and Efficient In-Memory Accelerator for Automata ProcessingabstractAccelerating finite automata processing benefits regular-expression workloads and a wide range of other applications that do not map obviously to regular expressions, including pattern mining, bioinformatics, and machine learning. Existing in-memory automata processing accelerators suffer from inefficient routing architectures. They are either incapable of efficiently place-and-route a highly connected automaton or require an excessive amount of hardware resources. Elaheh Sadredini, Reza Rahimi, Vaibhav Verma, Mircea R. Stan, Kevin Skadron |
MICRO | 4 |
| 2019 | Automata Processing in Reconfigurable Architectures: In-the-Cloud Deployment, Cross-Platform Evaluation, and Fast Symbol-Only ReconfigurationabstractWe present a general automata processing framework on FPGAs, which generates an RTL kernel for automata processing together with an AXI and PCIe based I/O circuitry. We implement the framework on both local nodes and cloud platforms (Amazon AWS and Nimbix) with novel features. A full performance comparison of the proposed framework is conducted against state-of-the-art automata processing engines on CPUs, GPUs, and Micron’s Automata Processor using the ANMLZoo benchmark suite and some real-world datasets. Results show that FPGAs enable extremely high-throughput automata processing compared to von Neumann architectures. We also collect the resource utilization and power consumption on the two cloud platforms, and find that the I/O circuitry consumes most of the hardware resources and power. Furthermore, we propose a fast, symbol-only reconfiguration mechanism based on the framework for large pattern sets that cannot fit on a single device and need to be partitioned. The proposed method supports multiple passes of the input stream and reduces the re-compilation cost from hours to seconds. Chunkun Bo, Vinh Dang, Ted Xie, Jack Wadden, Mircea R. Stan, Kevin Skadron |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 46 |
| 2019 | MTTF Enhancement Power-C4 Bump Placement OptimizationabstractConstructing a reliable power delivery network (PDN) is increasingly challenging by further technology scaling. PDN suffers from long-term reliability threats such as electromigration (EM). EM results in permanent failures and directly affects chip lifetime and voltage stability. Loss of limited controlled collapse chip connection (C4) pads to EM makes delivering a stable supply voltage more critical. The C4 bumps failure mechanism depends on current density, on-chip voltage noise, and temperature. In this paper, we develop a statistical simulation framework to analyze the effect of accurate chip temperature on multiple power-bump wearout. Our result shows that using uniform temperature leads, in the pessimistic system, to mean-time-to-failure (MTTF). Meanwhile, the current C4 pad placement optimization algorithms consider the current source or voltage drop as the forces to move the pads that do not lead to maximum MTTF. We propose a new technique to improve the efficiency of optimization algorithms by exploring the temperature as a new virtual force to move the pads. Experimental results show that our algorithm improves the MTTF by 25.43% that will support more off-chip I/O channels in current and near-future technology nodes. Somayeh Rahimipour, Runjie Zhang, Ke Wang 0011, Kevin Skadron, Fakhrul Z. Rokhani, Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | SRAM based opportunistic energy efficiency improvement in dual-supply near-threshold processorsabstractEnergy-efficient microprocessors are essential for a wide range of applications. While near-threshold computing is a promising technique to improve energy efficiency, optimal supply demands from logic core and on-chip memory are conflicting. In this paper, we perform reliability analysis of 6T SRAM and discover imbalanced minimum voltage requirements between read and write operations. We leverage this imbalance property in near-threshold processors equipped with voltage boosting capability by proposing an opportunistic dual-supply switching scheme with a write aggregation buffer. Our results show that proposed technique improves energy efficiency by more than 18% with approximate 8.54% performance speed-up. Yunfei Gu, Dengxue Yan, Vaibhav Verma, Mircea R. Stan, Xuan Zhang 0001 |
DAC | 4 |
| 2018 | Dyhard-DNN: even more DNN acceleration with dynamic hardware reconfigurationabstractDeep Neural Networks (DNNs) have demonstrated their utility across a wide range of input data types, usable across diverse computing substrates, from edge devices to datacenters. This broad utility has resulted in myriad hardware accelerator architectures. However, DNNs exhibit significant heterogeneity in their computational characteristics, e.g., feature and kernel dimensions, and dramatic variances in computational intensity, even between adjacent layers in one DNN. Consequently, accelerators with static hardware parameters run sub-optimally and leave energy-efficiency margins unclaimed. We propose DyHard-DNNs, where accelerator microarchitectural parameters are dynamically reconfigured during DNN execution to significantly improve metrics of interest. We demonstrate the effectiveness of this approach on a configurable SIMD 2D systolic array and show a 15--65% performance improvement (at iso-power) and 25--90% energy improvement (at iso-latency) over the best static configuration in six mainstream DNN workloads. Mateja Putic, Swagath Venkataramani, Schuyler Eldridge, Alper Buyuktosunoglu, Pradip Bose, Mircea R. Stan |
DAC | 6 |
| 2018 | OldSpot: A Pre-RTL Model for Fine-Grained Aging and Lifetime OptimizationabstractModern technologies have been experiencing evergrowing power densities as they scale down, raising temperatures and increasing aging and reliability concerns. High-level simulation is necessary in order to measure aging effects on lifetime. Existing simulation tools are inadequate due to their assumptions about homogeneity and failure tolerance. These limitations reduce their accuracy in modeling heterogeneous systems or systems with shared resources, leading to lifetime overestimation. We propose a new open-source tool called "OldSpot" that relaxes these assumptions and integrates low-level models for aging mechanisms with high-level reliability modeling techniques to enable fine-grained lifetime simulation, and then show how it can be used to improve the lifetime of a multicore system by duplicating functional units rather than adding extra cores. Using this structural duplication, we show area improvement by up to 13% while eliminating performance degradation due to failing resources. Additionally, we use OldSpot to show that lifetime and temperature are not perfectly correlated and that aging must be simulated along with temperature for optimal lifetime. Alec Roelke, Xinfei Guo, Mircea R. Stan |
ICCD | 3 |
| 2018 | Reservoir Computing Based Neural Image FiltersabstractClean images are an important requirement for machine vision systems to recognize visual features correctly. However, the environment, optics, electronics of the physical imaging systems can introduce extreme distortions and noise in the acquired images. In this work, we explore the use of reservoir computing, a dynamical neural network model inspired from biological systems, in creating dynamic image filtering systems that extracts signal from noise using inverse modeling. We discuss the possibility of implementing these networks in hardware close to the sensors. Samiran Ganguly, Yunfei Gu, Yunkun Xie, Mircea R. Stan, Avik W. Ghosh, Nibir K. Dhar |
IECON | 4 |
| 2018 | Tolerating Soft Errors in Processor Cores Using CLEAR (Cross-Layer Exploration for Architecting Resilience)abstractWe present cross-layer exploration for architecting resilience, a first of its kind framework which overcomes a major challenge in the design of digital systems that are resilient to reliability failures: achieve desired resilience targets at minimal costs (energy, power, execution time, and area) by combining resilience techniques across various layers of the system stack (circuit, logic, architecture, software, and algorithm). This is also referred to as cross-layer resilience. In this paper, we focus on radiation-induced soft errors in processor cores. We address both single-event upsets and single-event multiple upsets in terrestrial environments. Our framework automatically and systematically explores the large space of comprehensive resilience techniques and their combinations across various layers of the system stack (586 cross-layer combinations in this paper), derives cost-effective solutions that achieve resilience targets at minimal costs, and provides guidelines for the design of new resilience techniques. Our results demonstrate that a carefully optimized combination of circuit-level hardening, logic-level parity checking, and micro-architectural recovery provides a highly cost-effective soft error resilience solution for general-purpose processor cores. For example, a $50 {\times }$ improvement in silent data corruption (SDC) rate is achieved at only 2.1% energy cost for an out-of-order core (6.1% for an in-order core) with no speed impact. However, (application-aware) selective circuit-level hardening alone, guided by a thorough analysis of the effects of soft errors on application benchmarks, provides a cost-effective soft error resilience solution as well (with ~1% additional energy cost for a $50{\times }$ improvement in SDC rate). Eric Cheng, Shahrzad Mirkhani, Lukasz G. Szafaryn, Chen-Yong Cher, Hyungmin Cho, Kevin Skadron, Mircea R. Stan, Klas Lilja, Jacob A. Abraham, Pradip Bose, Subhasish Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2018 | Controlling the Reliability of SRAM PUFs With Directed NBTI Aging and Recovery
Alec Roelke, Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | REAPR: Reconfigurable engine for automata processingabstractFinite automata have proven their usefulness in high-profile domains ranging from network security to machine learning. While prior work focused on their applicability for purely regular expression workloads such as antivirus and network security rulesets, recent research has shown that automata can optimize the performance for algorithms in other areas such as machine learning and even particle physics. Unfortunately, their emulation on traditional CPU architectures is fundamentally slow and further bottlenecked by memory. In this paper, we present REAPR: Reconfigurable Engine for Automata PRocessing, a flexible framework that synthesizes RTL for automata processing applications as well as I/O to handle data transfer to and from the kernel. We show that even with memory and control flow overheads, FPGAs still enable extremely high-throughput computation of automata workloads compared to other architectures. Ted Xie, Vinh Dang, Jack Wadden, Kevin Skadron, Mircea R. Stan |
FPL | 5 |
| 2017 | Very Low Voltage (VLV) DesignabstractThis paper is a tutorial-style introduction to a special session on: Effective Voltage Scaling in the Late CMOS Era. It covers the fundamental challenges and associated solution strategies in pursuing very low voltage (VLV) designs. We discuss the performance and system reliability constraints that are key impediments to VLV. The associated trade-offs across power, performance and reliability are helpful in inferring the optimal operational voltage-frequency point. This work was performed under the auspices of an ongoing DARPA program (named PERFECT) that is focused on maximizing system-level energy efficiency. Ramon Bertran Monfort, Pradip Bose, David Brooks 0001, Jeff Burns, Alper Buyuktosunoglu, Nandhini Chandramoorthy, Eric Cheng, Martin Cochet, Schuyler Eldridge, Daniel J. Friedman, Hans M. Jacobson, Rajiv V. Joshi, Subhasish Mitra, Robert K. Montoye, Arun Paidimarri, Pritish Parida, Kevin Skadron, Mircea R. Stan, Karthik Swaminathan, Augusto Vega, Swagath Venkataramani, Christos Vezyrtzis, Gu-Yeon Wei, John-David Wellman, Matthew M. Ziegler |
ICCD | 18 |
| 2017 | Cross-Layer Resilience in Low-Voltage Digital Systems: Key InsightsabstractCLEAR (Cross-Layer Exploration for Architecting Resilience) is a first of its kind framework which overcomes a major challenge in the design of digital systems that are resilient to hardware errors: achieve desired resilience targets at low cost (energy, power, execution time, area) by combining resilience techniques across various layers of the system stack (circuit, logic, architecture, software, algorithm). CLEAR automatically and systematically explores the large space of resilience techniques and their combinations, derives cost-effective solutions, provides guidelines for designing new techniques, and offers insights into how to design cost-effective digital systems resilient to hardware errors: 1. circuit-level techniques are crucial; 2. application-level guidance is essential; 3. existing architecture and software techniques are generally expensive or provide too little resilience; 4. some previously published techniques suffer from inaccurate analysis, leading to incorrect conclusions; 5. cost-effective protection from multiple error sources is achieved by combining techniques targeting each specific error source. Eric Cheng, Jacob A. Abraham, Pradip Bose, Alper Buyuktosunoglu, Keith A. Campbell, Deming Chen, Chen-Yong Cher, Hyungmin Cho, Binh Q. Le, Klas Lilja, Shahrzad Mirkhani, Kevin Skadron, Mircea R. Stan, Lukasz G. Szafaryn, Christos Vezyrtzis, Subhasish Mitra |
ICCD | 13 |
| 2017 | Pre-RTL Voltage and Power Optimization for Low-Cost, Thermally Challenged Multicore ChipsabstractThe imminent end of Moore's Law demands increasing complexi-ty to enable continuing improvement in the cost and performance of electronic systems. Complex RTL designs lead to long simulation overheads, making it infeasible to explore large design spac-es. In this work, we present a flow of simulation tools for rapid, high-level, pre-RTL exploration to enable design-space exploration for physically-constrained systems. This flow updates several prior tools and introduces new tools to support optimization across multiple metrics: gem5, a widely-used microarchitecture and memory hierarchy simulator, for performance; McPAT, a generalized, ISA-agnostic power and area modeling tool; HotSpot, a temperature simulator; VoltSpot, a voltage droop simulator; and OldSpot, a planned lifetime and reliability simula-tor. By simulating a workload in gem5 and feeding its results into McPAT and then HotSpot, VoltSpot, and OldSpot, it is possible to explore the effects of workloads on systems with physical con-straints such as power or thermal budgets and lifetime targets without requiring complex RTL design or long RTL simulation. Alec Roelke, Runjie Zhang, Kaushik Mazumdar, Ke Wang 0011, Kevin Skadron, Mircea R. Stan |
ICCD | 6 |
| 2017 | Implications of accelerated self-healing as a key design knob for cross-layer resilience
Xinfei Guo, Mircea R. Stan |
Integr. | 2 |
| 2016 | Work hard, sleep well - Avoid irreversible IC wearout with proactive rejuvenationabstractVarious wearout mechanisms have both a reversible and an irreversible (permanent) part, with some, like BTI and EM having a significant reversible part, while others, like HCI, being mostly irreversible. In this paper we make two contributions. First, we show that the boundary between the reversible and irreversible parts of wearout is not fixed, with the irreversible part becoming at least partially reversible under the right conditions of active accelerated recovery and stress/recovery scheduling. Second, we show that there are certain stress/recovery schedules that can (almost) completely eliminate irreversible wearout, thus allowing significant reductions in necessary design margins. The experiments were done on commercial FPGAs fabricated in a 40nm technology. To fully repair and avoid the irreversible wearout, we propose a biology-inspired sleep-when-getting-tired strategy. The strategy can achieve >60× design margin reduction and ~9% average performance improvement within a 10-year lifetime constraint compared to the no-recovery case. Potential system level implementations (a negative “turbo-boost” like strategy) in multicore and NoC systems are also presented. Xinfei Guo, Mircea R. Stan |
ASP-DAC | 2 |
| 2016 | Clear: cross-layer exploration for architecting resilience combining hardware and software techniques to tolerate soft errors in processor coresabstractWe present a first of its kind framework which overcomes a major challenge in the design of digital systems that are resilient to reliability failures: achieve desired resilience targets at minimal costs (energy, power, execution time, area) by combining resilience techniques across various layers of the system stack (circuit, logic, architecture, software, algorithm). This is also referred to as cross-layer resilience. In this paper, we focus on radiation-induced soft errors in processor cores. We address both single-event upsets (SEUs) and single-event multiple upsets (SEMUs) in terrestrial environments. Our framework automatically and systematically explores the large space of comprehensive resilience techniques and their combinations across various layers of the system stack (798 cross-layer combinations in this paper), derives cost-effective solutions that achieve resilience targets at minimal costs, and provides guidelines for the design of new resilience techniques. We demonstrate the practicality and effectiveness of our framework using two diverse designs: a simple, in-order processor core and a complex, out-of-order processor core. Our results demonstrate that a carefully optimized combination of circuit-level hardening, logic-level parity checking, and micro-architectural recovery provides a highly cost-effective soft error resilience solution for general-purpose processor cores. For example, a 50× improvement in silent data corruption rate is achieved at only 2.1% energy cost for an out-of-order core (6.1% for an in-order core) with no speed impact. However, selective circuit-level hardening alone, guided by a thorough analysis of the effects of soft errors on application benchmarks, provides a cost-effective soft error resilience solution as well (with ~1% additional energy cost for a 50× improvement in silent data corruption rate). Eric Cheng, Shahrzad Mirkhani, Lukasz G. Szafaryn, Chen-Yong Cher, Hyungmin Cho, Kevin Skadron, Mircea R. Stan, Klas Lilja, Jacob A. Abraham, Pradip Bose, Subhasish Mitra |
DAC | 7 |
| 2016 | Generating efficient and high-quality pseudo-random behavior on Automata ProcessorsabstractMicron's Automata Processor (AP) efficiently emulates non-deterministic finite automata and has been shown to provide large speedups over traditional von Neumann execution for massively parallel, rule-based, data-mining and pattern matching applications. We demonstrate the AP's ability to generate high-quality and energy efficient pseudo-random behavior for use in pseudo-random number generation or in chip simulation. By recognizing that transition rules become probabilistic when input characters are randomized, the AP is also capable of simulating Markov chains. Combining hundreds of parallel Markov chains creates high-quality, high-throughput pseudo-random number sequences with greater power efficiency than state-of-the-art CPU and GPU algorithms. This indicates that the AP could potentially accelerate other Markov Chain-based applications such as agent-based simulation. We explore how to achieve throughputs upwards of 40GB/s per AP chip, with power efficiency 6.8x greater than state-of-the-art pseudo-random number generation on GPUs. Jack Wadden, Nathan Brunelle, Ke Wang 0011, Mohamed El-Hadedy 0001, Gabriel Robins, Mircea R. Stan, Kevin Skadron |
ICCD | 6 |
| 2016 | Tolerating the Consequences of Multiple EM-Induced C4 Bump FailuresabstractWith ever-increasing on-chip current density, technology scaling is pushing the electromigration (EM)-induced robustness of silicon chips' controlled collapse chip connection (C4) bump array to its limit. Since the density of C4 bumps is projected to be constant in the future, it is increasingly becoming challenging to guarantee EM-failure free for all power-supply bumps without increasing chip packaging cost or encroaching on bumps sites needed for I/O. In this paper, we develop a statistical simulation framework to analyze the mechanism and consequences of multiple power-bump wearout. Our analysis shows that the penalty of a moderate number of EM-induced power-bump failures is fairly small. A mild increase in on-chip supply voltage noise guardband can tolerate these bump failures and significantly increase a system mean-time-to-failure (MTTF). As a result, the targeted system MTTF can be achieved with significantly reduced power-bump count (e.g., 43% less) and a small extra noise margin (e.g., 0.5% VddIR drop). Runjie Zhang, Brett H. Meyer, Ke Wang 0011, Mircea R. Stan, Kevin Skadron |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | A cross-layer design exploration of charge-recycled power-delivery in many-layer 3d-ICabstract3D-IC technology brings both the opportunities to continue the historical trend of integration-level scaling and the challenges to deliver power reliably and efficiently. Voltage-stacking (V-S), a charge-recycled power delivery scheme that connects the different layers' supply/ground nets into a series stack, provides a scalable solution to the 3D-IC power delivery wall. While prior work has extensively discussed the implementations of V-S at circuit-level, a cross-layer study that examines its system-level implications is missing. In this paper, we start with a circuit implementation of a charge-recycled voltage regulator and build an architecture-level model to study the costs and benefits of utilizing V-S in 3D-IC. Our study shows that by significantly improving the EM-lifetime of C4 and TSV array (e.g., up to 5x) while only marginally increasing the average-case voltage noise (e.g., 0.75% Vdd IR drop), V-S provides a scalable solution for many-layer 3D-IC's power delivery challenge. Runjie Zhang, Kaushik Mazumdar, Brett H. Meyer, Ke Wang 0011, Kevin Skadron, Mircea R. Stan |
DAC | 6 |
| 2015 | Association Rule Mining with the Micron Automata ProcessorabstractAssociation rule mining (ARM) is a widely used data mining technique for discovering sets of frequently associated items in large databases. As datasets grow in size and real-time analysis becomes important, the performance of ARM implementation can impede its applicability. We accelerate ARM by using Micron's Automata Processor (AP), a hardware implementation of non-deterministic finite automata (NFAs), with additional features that significantly expand the APs capabilities beyond those of traditional NFAs. The Apriori algorithm that ARM uses for discovering item sets maps naturally to the massive parallelism of the AP. We implement the multipass pruning strategy used in the Apriori ARM through the APs symbol replacement capability, a form of lightweight reconfigurability. Up to 129X and 49X speedups are achieved by the AP-accelerated Apriori on seven synthetic and real-world datasets, when compared with the Apriori single-core CPU implementation and Eclat, a more efficient ARM algorithm, 6-core multicourse CPU implementation, respectively. The AP-accelerated Apriori solution also outperforms GPU implementations of Eclat especially for large datasets. Technology scaling projections suggest even better speedups from future generations of AP. Ke Wang 0011, Yanjun Qi, Jeffrey J. Fox, Mircea R. Stan, Kevin Skadron |
IPDPS | 4 |
| 2015 | Transient voltage noise in charge-recycled power delivery networks for many-layer 3D-ICabstractAside from the benefits it brings, 3D-IC technology inevitably exacerbates the difficulty of power delivery with volumetrically increasing power consumption. Recent work managed to “recycle” current within the 3D stack by linking the different layers' supply/ground nets into a series connection. This charge-recycled (also known as voltage-stacked, or V-S) scheme provides a scalable solution for 3D-IC's power delivery because it supports an arbitrary number of layers with a constant off-chip current demand. Although prior work has studied the circuit implementation of a V-S power delivery network (PDN) and its current-reduction benefits, a whole-system evaluation of V-S PDNs' transient voltage noise and a noise comparison between the V-S PDN and the traditional PDN are missing. In this paper, we build a system-level model to examine voltage-stacked 3D-ICs' transient noise and explore the impact of different PDN design parameters and workload behaviors. Our results show that compared with the traditional PDN scheme, V-S provides stronger isolation for cross-layer noise interference, which in turn grants higher performance benefits for run-time noise mitigation techniques, such as dynamic margin adaptation. We observe that, compared with traditional PDNs, V-S PDNs provide up to 60% lower transient noise in the worst-case scenario. Furthermore, we show that V-S PDNs significantly reduce the packaging cost, because their noise is almost insensitive to the package impedance (e.g., a 300% impedance increase only raises worst-case noise by less than 0.3% Vdd). Runjie Zhang, Kaushik Mazumdar, Brett H. Meyer, Ke Wang 0011, Kevin Skadron, Mircea R. Stan |
ISLPED | 6 |
| 2014 | Walking pads: Fast power-supply pad-placement optimizationabstractWe propose a novel C4 pad placement optimization framework for 2D power delivery grids: Walking Pads (WP). WP optimizes pad locations by moving pads according to the “virtual forces” exerted on them by other pads and current sources in the system. WP algorithms achieve the same IR drop as state-of-the-art techniques, but are up to 634X faster. We further propose an analytical model relating pad count and IR drop for determining the optimal pad count for a given IR drop budget. Ke Wang 0011, Brett H. Meyer, Runjie Zhang, Kevin Skadron, Mircea R. Stan |
ASP-DAC | 5 |
| 2014 | Modeling and Experimental Demonstration of Accelerated Self-Healing TechniquesabstractIn this paper we postulate that future electronics systems will use sleep time as an active recovery period essential for their overall performance. Our hypothesis is that by explicitly controlling the ratio of sleep vs. active and sleep conditions (e.g. higher temperatures, negative voltages), we can deeply rejuvenate electronic systems periodically to improve their metrics. We perform a series of stress and recovery experiments using commercial FPGAs to demonstrate several cases where we bring stressed chips back to within 90% of their original margin by actively rejuvenating for only 1/4 of the stress time. We validate our experiments against extracted models and present potential applications to multi-core systems. Xinfei Guo, Wayne P. Burleson, Mircea R. Stan |
DAC | 3 |
| 2014 | Computing with Hybrid CMOS/STO CircuitsabstractRecent research in spin torque nano-oscillators (STNO) have opened the possibility of using electron spin to generate sustained microwave oscillations. Furthermore, the experimental verification of synchronization of STNOs could allow communication and computation with nanoscaled oscillators. In this paper, we propose a hybrid MOSFET/STNO array which can be used for pattern recognition applications. First, we show that an array of electrically coupled STNOs obey the dynamics of Kuramoto's weakly coupled oscillators [1]. This behavior allows us to use the STNO array to implement the oscillatory neurocomputer proposed by Hoppensteadt et. al. [2]. Mehdi Kabir, Mircea R. Stan |
DAC | 2 |
| 2014 | Walking Pads: Managing C4 Placement for Transient Voltage Noise MinimizationabstractTransient voltage noise, including resistive and reactive noise, causes timing errors at runtime. We introduce a heuristic framework---Walking Pads---to minimize transient voltage violations by optimizing power supply pad placement. We show that the steady-state optimal design point differs from the transient optimum, and further noise reduction can be achieved with transient optimization. Our methodology significantly reduces voltage violations by balancing the average transient voltage noise of the four branches at each pad site. When we optimize pad placement using a representative stressmark, voltage violations are reduced 46-80% across 11 Parsec benchmarks with respect to the results from IR-drop-optimized pad placement. We also show that the allocation of on-chip decoupling capacitance significantly influences the optimal locations of pads. Ke Wang 0011, Brett H. Meyer, Runjie Zhang, Mircea R. Stan, Kevin Skadron |
DAC | 4 |
| 2014 | Architecture implications of pads as a scarce resourceabstractDue to non-ideal technology scaling, delivering a stable supply voltage is increasingly challenging. Furthermore, competition for limited chip interface resources (i.e., C4 pads) between power supply and I/O, and the loss of such resources to electromigration, means that constructing a power delivery network (PDN) that satisfies noise margins without compromising performance is and will remain a critical problem for architects and circuit designers alike. Simple guardbanding will no longer work, as the consequent performance penalty will grow with technology scaling. In this paper, we develop a pre-RTL PDN model, VoltSpot, for the purpose of studying the performance and noise tradeoffs among power supply and I/O pad allocation, the effectiveness of noise mitigation techniques, and the consequent implications of electromigration-induced PDN pad failure. Our simulations demonstrate that, despite their integral role in the PDN, power/ground pads can be aggressively reduced (by conversion into I/O pads) to their electromigration limit with minimal performance impact from extra voltage noise - provided the system implements a suitable noise-mitigation strategy. The key observation is that even though reducing power/ground pads significantly increases the number of voltage emergencies, the average noise amplitude increase is small. Overall, we can triple I/O bandwidth while maintaining target lifetimes and incurring only 1.5% slowdown. Runjie Zhang, Ke Wang 0011, Brett H. Meyer, Mircea R. Stan, Kevin Skadron |
ISCA | 4 |
| 2014 | A multi-output on-chip switched-capacitor DC-DC converter for near- and sub-threshold power modesabstractThis paper presents a novel multi-output on-chip switched-capacitor (SC) DC-DC converter simultaneously providing two output voltages (2VDD/3 and VDD/3, in addition to the available full VDD) that enables the use of three different power modes for optimizing power/performance trade-offs: super-threshold (full VDD), near-threshold (2VDD/3) and sub-threshold (VDD/3). Unlike previously proposed SC converters, the multiple conversion ratios of the novel converter are achieved without changing the topology of the circuit, thus fewer components being needed to support multiple voltages. Above 84% and 70% efficiencies are obtained in simulation over a range of load currents from 0.4 mA to 5 mA, and from 0.4 mA to 1 mA, with conversion ratios of 2/3 and 1/3, respectively. The paper also presents a novel optimization method to save area by using unequally sized flying capacitors. The target application is for systems with discrete power modes; also for systems that use dithering to emulate a continuous range of voltages such as the Panoptic Dynamic Voltage Scaling (PDVS). Yingbo Zhao, Yintang Yang, Kaushik Mazumdar, Xinfei Guo, Mircea R. Stan |
ISCAS | 5 |
| 2013 | Architectural implications of spatial thermal filtering
Karthik Sankaranarayanan, Brett H. Meyer, Wei Huang 0004, Robert J. Ribando, Hossein Haj-Hariri, Mircea R. Stan, Kevin Skadron |
Integr. | 6 |
| 2013 | Modeling Power Consumption of NAND Flash Memories Using FlashPowerabstractFlash is the most popular solid-state memory technology used today. A range of consumer electronics products, such as cell-phones and music players, use flash memory for storage and flash memory is increasingly displacing hard disk drives as the primary storage device in laptops, desktops, and servers. There is a rich microarchitectural design space for flash memory, and there are several architectural options for incorporating flash into the memory hierarchy. Exploring this design space requires detailed insights into the power characteristics of flash memory. In this paper, we present FlashPower, a detailed power model for the two most popular variants of NAND flash, namely, the single-level cell (SLC) and 2-bit Multi-Level Cell (MLC) based flash memory chips. FlashPower is built on top of CACTI, a widely used tool in the architecture community for studying various memory organizations. FlashPower takes several parameters like the device technology, microarchitectural layout, bias voltages and workload parameters as input to estimate the power consumption of a flash chip during its various operating modes. We validate FlashPower against chip power measurements from several different manufacturers and show that our results are comparable to the actual chip measurements. We illustrate the versatility of the tool in a design space exploration of power optimal flash memory array configurations. Vidyabhushan Mohan, Trevor Bunker, Laura M. Grupp, Sudhanva Gurumurthi, Mircea R. Stan, Steven Swanson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2012 | Breaking the power delivery wall using voltage stackingabstractWe propose the use of voltage stacking for addressing some of the power delivery issues for many-core processors. To demonstrate the effectiveness of our method we first design a proxy for a many-core stacked processor in the form of a regular structure using multiple ring oscillators where we can control the voltage, frequency and switching activity for individual rings. For intermediate voltage rail regulation, we propose a push pull-based switched capacitor regulator designed specifically for balancing the stacked loads. Detailed Spice simulation results for the prototype model show a 4× reduction in supply current when using 4 layers of voltage stacking. We further validate our method by designing a voltage-stacked structure using two PIC cores. Kaushik Mazumdar, Mircea R. Stan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2012 | A new taxonomy for reconfigurable prefix addersabstractWhile previous taxonomies for prefix adders have focused on the design space of such adders with fixed topologies (in terms of fanout, radix, logic depth, wiring tracks), our work considers the design space of reconfigurable prefix adders (with applications in fault-tolerant adder design) by introducing several new degrees of freedom in the design space. Fault tolerance in general requires redundancy, so we start with a redundant structure which is a superset of the entire family of prefix adders with fanout-of-2 and then prune the redundant structure accordingly for the defect-free initial state or for when defects occur. In addition to the traditional prefix adders (Kogge-Stone, Han-Carlson, Brent-Kung) our taxonomy proposes several new variations that are equivalent in performance and complexity to the traditional structures, yet can mask different sets of faults. Stevo D. Bailey, Mircea R. Stan |
ISCAS | 2 |
| 2012 | Self-assembled multiferroic magnetic QCA structures for low power systemsabstractThe discovery of multiferroic materials has lead to a great interest in creating logic circuits which exploit both the magnetic and electrical properties of these materials. In this work we focus on self-assembled array structures of Magnetic Quantum Cellular Automata (MQCA) composed of multiferroic nanopillars which can be configured and clocked using solely electric fields. Furthermore, due to the switching nature of these nanopillars, the arrays can be reconfigured to implement read/write memory or multiple logic circuits in a similar fashion to FPGAs. Finally, we develop a SPICE model for the multiferroic nanopillars and demonstrate the functionality of the array. Mircea R. Stan, Mehdi Kabir, Jiwei Lu, Stuart A. Wolf |
ISCAS | 1 |
| 2012 | ArchFP: Rapid prototyping of pre-RTL floorplans
Gregory G. Faust, Runjie Zhang, Kevin Skadron, Mircea R. Stan, Brett H. Meyer |
VLSI-SoC | 4 |
| 2012 | Tracking On-Chip Age Using Distributed, Embedded SensorsabstractRecent works show bias temperature instability (BTI) is a detrimental hard-aging mechanism in CMOS circuit design. Negative BTI (NBTI) alone degrades circuit speed upwards of 20% over a 10 year life-span. Having the ability to track the actual aging process provides one method to reduce large design margins that are otherwise required to offset circuit aging. This work extends previous research by contributing a sensing scheme that employs on-chip sensors capable of accurately tracking NBTI pMOS current degradations across process, temperature, and varying activity factors. Results show that a 7600$\mu{\hbox {m}}^{2}$sensing area achieves an overall system accuracy of 90% at a voltage threshold precision of 2 mV. We thoroughly describe the sensor design and the underlying statistics used to determine overall accuracy and precision. Furthermore, a novel sensor distribution method is presented that uses an existing scan-chain methodology to mask the overhead of adding the on-chip sensors. Stuart N. Wooters, Adam C. Cabe, Zhenyu Qi 0001, Jiajing Wang, Randy W. Mann, Benton H. Calhoun, Mircea R. Stan, Travis N. Blalock |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2011 | Experimental demonstration of standby power reduction using voltage stacking in an 8Kb embedded FDSOI SRAMabstractVoltage stacking has been proposed as an efficient solution for power delivery in high performance processors, for 3D ICs, for pin-limited ICs, and for implicit sleep mode (standby) DC/DC conversion. In this paper we demonstrate voltage stacking for an 8Kb embedded SRAM in 180nm fully-depleted SOI (FDSOI) which leads to 88.6% reduction in standby power, including overhead. The SRAM is formed of two 4Kb subarrays which are powered in parallel during active mode, and stacked in series during standby. The SRAM uses no explicit decoupling or regulating and achieves active-to-sleep and sleep-to-active transitions of less than 10ns and a breakeven time of 20ns. Adam C. Cabe, Mircea R. Stan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | RAMA: a self-assembled multiferroic magnetic QCA for low power systemsabstractRecently, with the discovery of multiferroic materials, there has been a great interest in creating logic devices which exploit both magnetic and electric properties of these materials. This paper proposes a reconfigurable array of magnetic automata (RAMA) made of multiferroic nanopillars which can be operated using electric fields. Furthermore, due to the switching nature of these nanopillars, the array can be reconfigured to implement multiple logic circuits in a similar fashion to FPGAs. The paper proposes a compact model of the multiferroic switching mechanism which can be used to describe the behavior of the nanopillars in a circuit simulator. In addition, results from micromagnetic simulations of the MQCA bits indicate that it can operate with energy consumptions that are magnitudes lower than conventional CMOS technologies. Finally, the paper discusses the reliability of the nanopillar switching and suggests ways to optimize the error rates. Mehdi Kabir, Mircea R. Stan, Stuart A. Wolf, Ryan B. Comes, Jiwei Lu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Relaxing non-volatility for fast and energy-efficient STT-RAM cachesabstractSpin-Transfer Torque RAM (STT-RAM) is an emerging non-volatile memory technology that is a potential universal memory that could replace SRAM in processor caches. This paper presents a novel approach for redesigning STT-RAM memory cells to reduce the high dynamic energy and slow write latencies. We lower the retention time by reducing the planar area of the cell, thereby reducing the write current, which we then use with CACTI to design caches and memories. We simulate quad-core processor designs using a combination of SRAM- and STT-RAM-based caches. Since ultra-low retention STT-RAM may lose data, we also provide a preliminary evaluation for a simple, DRAM-style refresh policy. We found that a pure STT-RAM cache hierarchy provides the best energy efficiency, though a hybrid design of SRAM-based L1 caches with reduced-retention STT-RAM L2 and L3 caches eliminates performance loss while still reducing the energy-delay product by more than 70%. Clinton Wills Smullen IV, Vidyabhushan Mohan, Anurag Nigam, Sudhanva Gurumurthi, Mircea R. Stan |
HPCA | 5 |
| 2011 | The STeTSiMS STT-RAM simulation and modeling systemabstractThere is growing interest in emerging non-volatile memory technologies such as Phase-Change Memory, Memristors, and Spin-Transfer Torque RAM (STT-RAM). STT-RAM, in particular, is experiencing rapid development that can be difficult for memory systems researchers to take advantage of. What is needed are techniques that enable designers to explore the potential of recent STT-RAM designs and adjust the performance without needing a detailed understanding of the physics. In this paper, we present the STeTSiMS STT-RAM Simulation and Modeling System to assist memory systems researchers. After providing background on the operation of STT-RAM magnetic tunnel junctions (MTJs), we demonstrate how to fit three different published MTJ models to our model and normalize their characteristics with respect to common metrics. The high-speed switching behavior of the designs is evaluated using macromagnetic simulations. We have also added a first-order model for STT-RAM memory arrays to the CACTI memory modeling tool, which we then use to evaluate the performance, energy consumption, and area for: (i) a high-performance cache, (ii) a high-capacity cache, and (iii) a high-density memory. Clinton Wills Smullen IV, Anurag Nigam, Sudhanva Gurumurthi, Mircea R. Stan |
ICCAD | 4 |
| 2011 | Delivering on the promise of universal memory for spin-transfer torque RAM (STT-RAM)
Anurag Nigam, Clinton Wills Smullen IV, Vidyabhushan Mohan, Eugene Chen, Sudhanva Gurumurthi, Mircea R. Stan |
ISLPED | 6 |
| 2010 | Stacking SRAM banks for ultra low power standby mode operationabstractOn-chip SRAM caches have come to dominate the total chip area and leakage power consumed in state-of-the-art microprocessor designs. Such large memories are necessary to attain high performance, however it is critical to minimize the idle currents drawn while these SRAM banks are inactive. This work proposes a novel voltage reduction technique to reduce SRAM leakage power during the standby mode. The design employs an implicit voltage reduction method that "stacks" SRAM banks in series while these blocks are inactive. No explicit DC/DC converters are required to achieve the reduced voltages, which leads to large area reductions over techniques requiring on-chip regulation circuits. This stacking technique reduces the voltage on each block close to the absolute data retention voltage (DRV) of each cell, and achieves a maximum leakage power reduction of 93% from the active power mode. Simulation results show the stability of the scheme around corners, process variations, and on-chip noise. Adam C. Cabe, Zhenyu Qi 0001, Mircea R. Stan |
DAC | 3 |
| 2010 | SRAM-based NBTI/PBTI sensor system designabstractNBTI has been a major aging mechanism for advanced CMOS technology and PBTI is also looming as a big concern. This work first proposes a compact on-chip sensor design that tracks both NBTI and PBTI for both logic and SRAM circuits. Embedded in an SRAM array the sensor takes the form of a 6T SRAM cell and is at least 30x smaller than previous designs. Extensively reusing the SRAM peripheral circuitry minimizes control logic overhead. Sensing overhead is further amortized as the sensors can be both reconfigured and recycled as functional SRAM cells, potentially increasing SRAM yield when other bit cells fail due to initial process variation or long time aging effects. The paper also proposes a variation-aware sensor system design methodology by quantifying and leveraging the tradeoff between the size and number of sensors and the system sensing precision. Design examples show that a system of 500 sensors can achieve 4mV precision with 98.8% confidence, and a system of 1K sensors designed for 1M SRAM bit cells achieves 2000x area overhead reduction compared to a worst-case based approach. Zhenyu Qi 0001, Jiajing Wang, Adam C. Cabe, Stuart N. Wooters, Travis N. Blalock, Benton H. Calhoun, Mircea R. Stan |
DAC | 7 |
| 2010 | FlashPower: A detailed power model for NAND flash memoryabstractFlash memory is widely used in consumer electronics products, such as cell-phones and music players, and is increasingly displacing hard disk drives as the primary storage device in laptops, desktops, and even servers. There is a rich microarchitectural design space for flash memory and there are several architectural options for incorporating flash into the memory hierarchy. Exploring this design space requires detailed insights into the power characteristics of flash memory. In this paper, we present FlashPower, a detailed analytical power model for Single-Level Cell (SLC) based NAND flash memory, which is used in high-performance flash products. We have integrated FlashPower with CACTI 5.3, which is widely used in the architecture community for studying memory organizations. FlashPower takes as input device technology and microarchitectural parameters to estimate the power consumed by a flash chip during its various operating modes. We have validated FlashPower against published chip power measurements and show that they are comparable. Vidyabhushan Mohan, Sudhanva Gurumurthi, Mircea R. Stan |
DATE | 3 |
| 2010 | How I Learned to Stop Worrying and Love Flash Endurance
Vidyabhushan Mohan, Taniya Siddiqua, Sudhanva Gurumurthi, Mircea R. Stan |
HotStorage | 4 |
| 2010 | Temperature-to-power mappingabstractAccurate power maps are useful for power model validation, process variation characterization, leakage estimation, and power optimization, but are hard to measure directly. Deriving power maps from measured thermal maps is the inverse problem of the power-to-temperature mapping, extensively studied through thermal simulation. Until recently this inverse heat conduction problem has received little attention in the microarchitecture research community. This paper first identifies the source of difficulties for the problem. The inverse mapping is then performed by applying constraints from microarchitecture-level observations. The inherent large sensitivity of the resultant power map is minimized through thermal map-filtering and constrained least-squares optimization. Choices of filter parameters and optimization constraints are investigated and their effects are evaluated. Furthermore, the paper highlights the differences between the grid and block modeling in the inverse mapping which were often ignored by previous schemes. The proposed methods reduce the mapping error by more than 10× compared to unoptimized solutions. To our best knowledge this is the first work to quantitatively evaluate and minimize the noise effect in the temperature to power mapping problem at the microarchitecture level for both grid and block mode, and for the steady and transient case. Zhenyu Qi 0001, Brett H. Meyer, Wei Huang 0004, Robert J. Ribando, Kevin Skadron, Mircea R. Stan |
ICCD | 6 |
| 2010 | The Promise of Nanomagnetics and Spintronics for Future Logic and Universal MemoryabstractThis paper is both a review of some recent developments in the utilization of magnetism for applications to logic and memory and a description of some new innovations in nanomagnetics and spintronics. Nanomagnetics is primarily based on the magnetic interactions, while spintronics is primarily concerned with devices that utilize spin polarized currents. With the end of complementary metal-oxide-semiconductor (CMOS) in sight, nanomagnetics can provide a new paradigm for information process using the principles of magnetic quantum cellular automata (MQCA). This paper will review and describe these principles and then introduce a new nonlithographic method of producing reconfigurable arrays of MQCAs and/or storage bits that can be configured electrically. Furthermore, this paper will provide a brief description of magnetoresistive random access memory (MRAM), the first mainstream spintronic nonvolatile random access memory and project how far its successor spin transfer torque random access memory (STT-RAM) can go to provide a truly universal memory that can in principle replace most, if not all, semiconductor memories in the near future. For completeness, a description of an all-metal logic architecture based on magnetoresistive structures (transpinnor) will be described as well as some approaches to logic using magnetic tunnel junctions (MTJs). Stuart A. Wolf, Jiwei Lu, Mircea R. Stan, Eugene Chen, Daryl M. Treger |
Proc. IEEE | 3 |
| 2009 | Graphene Devices, Interconnect and Circuits - Challenges and OpportunitiesabstractGraphene has recently emerged as a serious contender for the post silicon era. Graphene nanoribbon (GNR) devices have similar performance characteristics to carbon nanotube (CNT) ones. However, lithographic patterning methods applied to graphene can avoid the degree of chirality control and alignment issues typical of CNTs, and GNR devices and GNR interconnect can in principle be seamlessly obtained by patterning single graphene sheets, thus leading to monolithically device-interconnect structures. Electrically doped GNR devices in series and in parallel can be used for creating complex GNR FET digital circuits. There are also several important challenges facing the graphene ldquobrave new world,rdquo but many of the difficulties hopefully will have tractable solutions. This paper examines the topic of GNR FET circuit design from a bottom-up theoretical perspective, starting with GNR device and interconnect modeling and simulation, while trying to reconcile theory with some recent experimental results. Mircea R. Stan, Dincer Unluer, Avik W. Ghosh, Frank Tseng |
ISCAS | 1 |
| 2009 | Differentiating the roles of IR measurement and simulation for power and temperature-aware designabstractIn temperature-aware design, the presence or absence of a heatsink fundamentally changes the thermal behavior with important design implications. In recent years, chip-level infrared (IR) thermal imaging has been gaining popularity in studying thermal phenomena and thermal management, as well as reverse-engineering chip power consumption. Unfortunately, IR thermal imaging needs a peculiar cooling solution, which removes the heatsink and applies an IR-transparent liquid flow over the exposed bare die to carry away the dissipated heat. Because this cooling solution is drastically different from a normal thermal package, its thermal characteristics need to be closely examined. In this paper, we characterize the differences between two cooling configurations-forced air flow over a copper heatsink (AIR-SINK) and laminar oil flow over bare silicon (OIL-SILICON). For the comparison, we modify the HotSpot thermal model by adding the IR-transparent oil flow and the secondary heat transfer path through the package pins, hence modeling what the IR camera actually sees at runtime. We show that OIL-SILICON and AIR-SINK are significantly different in both transient and steady-state thermal responses. OIL-SILICON has a much slower short-term transient response, which makes dynamic thermal management less efficient. In addition, for OIL-SILICON, the direction of oil flow plays an important role by changing hot spot location, thus impacting hot spot identification and thermal sensor placement. These results imply that the power- and temperature-aware design process cannot just rely on IR measurements. Simulation and IR measurement are both needed and are complementary techniques. Wei Huang 0004, Kevin Skadron, Sudhanva Gurumurthi, Robert J. Ribando, Mircea R. Stan |
ISPASS | 5 |
| 2009 | Sensitivity-Based Optimization of Disk ArchitectureabstractStorage plays a pivotal role in the performance of many applications. Many applications, especially those that run on servers, are I/O intensive and therefore require high performance storage systems. These high-end storage systems consume a large amount of power, the bulk of which is due to the disk drives. Optimizing disk architectures is a design time as well as a run time issue and requires balancing between performance and power. There are different figures of merit, such as performance and energy, and a large space of design and runtime "knobs" that can be used to optimize disk drive behavior. Given such a large space, it is desirable to have a systematic methodology to optimally set these knobs to satisfy our figures of merit as efficiently as possible. In this paper we present the sensitivity-based optimization methodology for disk architectures (SODA), which leverages results previously obtained in digital circuit design optimization scenarios. Using detailed models of the electro-mechanical behavior of disk drives and a suite of realistic workloads, we show how SODA can aid in design and runtime optimization of disk drive architectures. Sriram Sankar, Yan Zhang 0028, Sudhanva Gurumurthi, Mircea R. Stan |
IEEE Trans. Computers | 4 |
| 2008 | Many-core design from a thermal perspectiveabstractAir cooling limits have been a major design challenge in recent years for integrated circuits. Multi-core exacerbates thermal challenges because power scales with the number of cores, but also creates new opportunities for temperature-aware design, because multi-core designs offer more design parameters than single-core designs. This paper investigates the relationship between core size and on-chip hot spot temperature and shows that with the same power density, smaller cores are cooler than larger cores due to a spatial low-pass filtering effect of temperature. This phenomenon suggests that designs exploiting low-pass filtering can dissipate more power within the same cooling budget than contemporary designs. Wei Huang 0004, Mircea R. Stan, Karthik Sankaranarayanan, Robert J. Ribando, Kevin Skadron |
DAC | 2 |
| 2008 | NBTI resilient circuits using adaptive body biasingabstractReliability has become a practical concern in today's VLSI design with advanced technologies. In-situ sensors have been proposed for reliability monitoring to provide advance warnings before system errors occur. This paper presents a reliability monitor design for NBTI (Negative Bias Temperature Instability). NBTI is recognized as very critical as it leads to short device lifetime. The proposed reliability monitor not only tracks the NBTI effect but also mitigates the degradation by forward biasing the PMOS. A worst case scenario static stress experiment demonstrates two orders of magnitude improvement in system lifetime using PTM 65nm technology. A ring oscillator example shows how frequency degradation can be compensated. Deployment of the proposed NBTI monitor is also discussed and two compatible strategies are provided to incorporate these monitors efficiently: the first focuses on low area overhead while the second features low power. Zhenyu Qi 0001, Mircea R. Stan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | Intra-disk Parallelism: An Idea Whose Time Has ComeabstractServer storage systems use a large number of disks to achieve high performance, thereby consuming a significant amount of power. In this paper, we propose to significantly reduce the power consumed by such storage systems via intra-disk parallelism, wherein disk drives can exploit parallelism in the I/O request stream. Intra-disk parallelism can facilitate replacing a large disk array with a smaller one, using the minimum number of disk drives needed to satisfy the capacity requirements. We show that the design space of intra-disk parallelism is large and present a taxonomy to formulate specific implementations within this space. Using a set of commercial workloads, we perform a limit study to identify the key performance bottlenecks that arise when we replace a storage array that is tuned to provide high performance with a single high-capacity disk drive. We show that it is possible to match, and even surpass, the performance of a storage array for these workloads by using a single disk drive of sufficient capacity that exploits intra-disk parallelism, while significantly reducing the power consumed by the storage system. We evaluate the performance and power consumption of disk arrays composed of intra-disk parallel drives, and discuss engineering and cost issues related to the implementation and deployment of such disk drives. Sriram Sankar, Sudhanva Gurumurthi, Mircea R. Stan |
ISCA | 3 |
| 2008 | Sensitivity Based Power Management of Enterprise Storage Systems
Sriram Sankar, Sudhanva Gurumurthi, Mircea R. Stan |
MASCOTS | 3 |
| 2008 | Accurate, Pre-RTL Temperature-Aware Design Using a Parameterized, Geometric Thermal ModelabstractPreventing silicon chips from negative, even disastrous thermal hazards has become increasingly challenging these days; considering thermal effects early in the design cycle is thus required. To achieve this, an accurate yet fast temperature model together with an early-stage, thermally optimized, design flow are needed. In this paper, we present an improved block-based compact thermal model (HotSpot 4.0) that automatically achieves good accuracy even under extreme conditions. The model has been extensively validated with detailed finite-element thermal simulation tools. We also show that properly modeling package components and applying the right boundary conditions are crucial to making full-chip thermal models like HotSpot accurately resemble what happens in the real world. Ignoring or over-simplifying package components can lead to inaccurate temperature estimations and potential thermal hazards that are costly to fix in later designs stages. Such a full-chip and package thermal model can then be incorporated into a thermally optimized design flow where it acts as an efficient communication medium among computer architects, circuit designers and package designers in early microprocessor design stages, to achieve early and accurate design decisions and also faster design convergence. For example, the temperature-leakage interaction can be readily analyzed within such a design flow to predict potential thermal hazards such as thermal runaway. Wei Huang 0004, Karthik Sankaranarayanan, Kevin Skadron, Robert J. Ribando, Mircea R. Stan |
IEEE Trans. Computers | 5 |
| 2007 | SODA: Sensitivity Based Optimization of Disk ArchitectureabstractStorage plays a pivotal role in the performance of many applications. Optimizing disk architectures is a design-time as well as a run-time issue and requires balancing between performance, power and capacity. The design space is large and there are many "knobs" that can be used to optimize disk drive behavior. Here we present a sensitivity-based optimization for disk architectures (SODA) which leverages results from digital circuit design. Using detailed models of the electro-mechanical behavior of disk drives and a suite of realistic workloads, we show how SODA can aid in design and runtime optimization. Yan Zhang 0028, Sudhanva Gurumurthi, Mircea R. Stan |
DAC | 3 |
| 2007 | Temperature-aware circuit design using adaptive body biasingabstractDue to continuously increasing active power dissipation die temperatures exhibit significant spatial and temporal variability. Targeting the worst-case temperature is not optimal as performance will be lost for non-worst cases. This paper proposes a temperature-adaptive body bias technique that can dynamically recover this lost performance. Two applications are described that demonstrate the effectiveness of proposed technique. Yan Zhang 0028, Mircea R. Stan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Structured and tuned array generation (STAG) for high-performance random logicabstractRegularly structured design techniques can combat complexity on a variety of fronts. We present the Structured and Tuned Array Generation (STAG) design methodology, which provides a complete design solution from logic to layout for regularly structured circuits. The STAG circuit tuning constraints are a key component of the methodology. The tuning contraints first guide a SPICE-level tuner to a violation free region in the design space. Secondly, the tuning methodology provides flexibility for targeting a variety of design contraints and objectives. Design examples illustrate STAG's ability for fast turnaround time as well as for high performance and timing critical random logic. Matthew M. Ziegler, Gary S. Ditlow, Stephen V. Kosonocky, Zhenyu Qi 0001, Mircea R. Stan |
ACM Great Lakes Symposium on VLSI | 5 |
| 2007 | Designing CMOS/molecular memories while considering device parameter variationsabstractIn recent years, many advances have been made in the development of molecular scale devices. Experimental data shows that these devices have potential for use in both memory and logic. This article describes the challenges faced in building crossbar array-based molecular memory and develops a methodology to optimize molecular scale architectures based on experimental device data taken at room temperature. In particular, issues in reading and writing such as memory using CMOS are discussed, and a solution is introduced for easily reading device conductivity states (typically characterized by very small currents). Additionally, a metric is derived to determine the voltages for writing to the crossbar array. The proposed memory design is also simulated with consideration to device parameter variations. Thus, the results presented here shed light on important design choices to be made at multiple abstraction levels, from devices to architectures. Simulation results, incorporating experimental device data, are presented using Cadence Spectre. Garrett S. Rose, Yuxing Yao, James M. Tour, Adam C. Cabe, Nadine Gergel-Hackett, Nabanita Majumdar, John C. Bean, Lloyd R. Harriott, Mircea R. Stan |
ACM J. Emerg. Technol. Comput. Syst. | 9 |
| 2007 | Interconnect Lifetime Prediction for Reliability-Aware SystemsabstractThermal effects are becoming a limiting factor in high-performance circuit design due to the strong temperature dependence of leakage power, circuit performance, IC package cost, and reliability. While many interconnect reliability models assume a constant temperature, this paper analyzes the effects of temporal and spatial thermal gradients on interconnect lifetime in terms of electromigration, and presents a physics-based dynamic reliability model which returns reliability equivalent temperature and current density that can be used in traditional reliability analysis tools. The model is verified with numerical simulations and reveals that blindly using the maximum temperature leads to too pessimistic lifetime estimation. Therefore, the proposed model not only increases the accuracy of reliability estimates, but also enables designers to reclaim design margin in reliability-aware design. In addition, the model is useful for improving the performance of temperature-aware runtime management by modeling system lifetime as a resource to be consumed at a stress-dependent rate Zhijian Lu, Wei Huang 0004, Mircea R. Stan, Kevin Skadron, John C. Lach |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2006 | Procrastinating voltage scheduling with discrete frequency setsabstractThis paper presents an efficient method to find the optimal intra-task voltage/frequency scheduling for single tasks in practical real-time systems using statistical workload information. Our method is analytic in nature and proved to be optimal. Simulation results verify our theoretical analysis and show significant energy savings over previous methods. In addition, in contrast to the previous techniques in which all available frequencies are used in a schedule, we find that, by carefully selecting a subset of a small number of frequencies, one can still design a reasonably good schedule while avoiding unnecessary transition overheads. Zhijian Lu, Yan Zhang 0028, Mircea R. Stan, John C. Lach, Kevin Skadron |
DATE | 3 |
| 2006 | A programmable majority logic array using molecular scale electronicsabstractWe present a computer architecture design that utilizes molecular electronic devices fabricated from self-assembled monolayers (SAM) of molecules. This architecture has been designed for a process being developed at the University of Virginia where molecules are assembled via vapor phase deposition. As this process lends itself nicely to developing multiple layers of devices on a single substrate, the circuit and architectural designs presented here exist in three dimensions. Through this work we show how molecular electronics naturally allows for 3D integration at the nanoscale.The design consists of two types of molecular devices: switches with memory and resonant tunneling diodes (RTD). These switching molecules are patterned between layers of metal to form a uniform crossbar array that can behave similar to a programmable logic array (PLA). In this design, the crossbar arrays drive rows of circuits, referred to as Goto pairs, consisting of two stacked molecular RTDs. As described in previous work, a Goto pair with only one resistor driving its input functions as a latch whereas the circuit acts as a majority gate when driven by multiple resistors. Since the crossbar arrays can be programmed such that the various Goto pairs are driven by different numbers of inputs, the overall circuit has the ability to implement logic as networks of latches and majority gates. Using crossbar arrays in this way is somewhat different from other approaches in that logic is not simply implemented in the array but is a product of both the array and the Goto pair functionality. Garrett S. Rose, Mircea R. Stan |
FPGA | 2 |
| 2006 | Design approaches for hybrid CMOS/molecular memory based on experimental device dataabstractIn recent years many advances have been made in the development of molecular scale devices. Experimental data shows that these devices have potential for use in both memory and logic. This paper describes the challenges faced in building crossbar array based molecular memory, and develops a methodology to optimize molecular scale architectures based on experimental device data taken at room temperature. In particular, we discuss reading and writing such memory using CMOS and compiling a solution for easily reading device conductivity states (typically characterized by very small currents). Additionally, a metric is derived to determine the voltages for writing to the crossbar array. Simulation results, incorporating experimental device data, are presented using Cadence Spectre. Garrett S. Rose, Adam C. Cabe, Nadine Gergel-Hackett, Nabanita Majumdar, Mircea R. Stan, John C. Bean, Lloyd R. Harriott, Yuxing Yao, James M. Tour |
ACM Great Lakes Symposium on VLSI | 5 |
| 2006 | HotSpot: A Compact Thermal Modeling Methodology for Early-Stage VLSI DesignabstractThis paper presents HotSpot-a modeling methodology for developing compact thermal models based on the popular stacked-layer packaging scheme in modern very large-scale integration systems. In addition to modeling silicon and packaging layers, HotSpot includes a high-level on-chip interconnect self-heating power and thermal model such that the thermal impacts on interconnects can also be considered during early design stages. The HotSpot compact thermal modeling approach is especially well suited for preregister transfer level (RTL) and presynthesis thermal analysis and is able to provide detailed static and transient temperature information across the die and the package, as it is also computationally efficient. Wei Huang 0004, Shougata Ghosh, Sivakumar Velusamy, Karthik Sankaranarayanan, Kevin Skadron, Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2005 | Optimal procrastinating voltage scheduling for hard real-time systemsabstractThis paper presents an optimal procrastinating voltage scheduling (OP-DVS) for hard real-time systems using stochastic workload information. Algorithms are presented for both single-task and multi-task workloads. Offline calculations provide real-time guarantees for worst-case execution, and online scheduling reclaims slack time and schedules tasks accordingly. The OP-DVS algorithm is provably optimal in terms of energy minimization with no deadline misses. Simulation results show up to 30% energy savings for single-task workloads and 74% for multi-task workloads compared to using a constant worst-case execution voltage. The complexity of the algorithm for multi-task workloads is linear to the number of tasks involved. Yan Zhang 0028, Zhijian Lu, John C. Lach, Kevin Skadron, Mircea R. Stan |
DAC | 5 |
| 2005 | Monitoring Temperature in FPGA based SoCsabstractFPGA logic densities continue to increase at a tremendous rate. This has had the undesired consequence of increased power density, which manifests itself as higher on-die temperatures and local hotspots. Sophisticated packaging techniques have become essential to maintain the health of the chip. In addition to static techniques to reduce the temperature, dynamic thermal management techniques are essential. Such techniques rely on accurate on-chip temperature information. In this paper, we present the design of a system that monitors the temperatures at various locations on the FPGA. This system is composed of a controller interfacing to an array of temperature sensors that are implemented on the FPGA fabric. Such a system can be used to implement dynamic thermal management techniques. We cross validate the sensor readings with values obtained from HotSpot, a pre-RTL architectural level thermal modeling tool. Sivakumar Velusamy, Wei Huang 0004, John C. Lach, Mircea R. Stan, Kevin Skadron |
ICCD | 4 |
| 2005 | The need for a full-chip and package thermal model for thermally optimized IC designsabstractModeling and analyzing detailed die temperature with a full-chip thermal model at early design stages is important to discover and avoid potential thermal hazards. However, omitting important aspects of package details in a thermal model can result in significant temperature estimation errors. In this paper, we discuss the applications of an existing compact thermal model that models both die and package temperature details. As an example, a thermally selfconsistent leakage power calculation of a POWER4-like microprocessor design is presented. We then demonstrate the importance of including detailed package information in the thermal model by several examples considering the impact of thermal interface material (TIM), which glues the die to the heat spreader. The fact that detailed package information is needed to build an accurate compact thermal model implies a design flow, in which the chip- and package-level compact thermal model acts as a convenient medium for more productive collaborations among circuit designers, computer architects and package designers, leading to early and efficient evaluations of different design tradeoffs for an optimal design from a thermal point of view. Categories and Subject Descriptors: Wei Huang 0004, Eric Humenay, Kevin Skadron, Mircea R. Stan |
ISLPED | 4 |
| 2004 | System level leakage reduction considering the interdependence of temperature and leakageabstractThe high leakage devices in nanometer technologies as well as the low activity rates in system-on-a-chip (SOC) contribute to the growing significance of leakage power at the system level. We first present system-level leakage-power modeling and characteristics and discuss ways to reduce leakage for caches. Considering the interdependence between leakage power and temperature, we then discuss thermal runaway and dynamic power and thermal management (DPTM) to reduce power and prevent thermal violations. We show that a thermal-independent leakage model may hide actual failures of DPTM. Finally, we present voltage scaling considering DPTM for different packaging options. We show that the optimal Vdd for the best throughput may be smaller than the largest Vdd allowed by the given packaging platform, and that advanced cooling techniques can improve throughput significantly. Lei He 0001, Weiping Liao, Mircea R. Stan |
DAC | 3 |
| 2004 | Compact thermal modeling for temperature-aware designabstractThermal design in sub-100nm technologies is one of the major challenges to the CAD community. In this paper, we first introduce the idea of temperature-aware design. We then propose a compact thermal model which can be integrated with modern CAD tools to achieve a temperature-aware design methodology. Finally, we use the compact thermal model in a case study of microprocessor design to show the importance of using temperature as a guideline for the design. Results from our thermal model show that a temperature-aware design approach can provide more accurate estimations, and therefore better decisions and faster design convergence. Wei Huang 0004, Mircea R. Stan, Kevin Skadron, Karthik Sankaranarayanan, Shougata Ghosh, Sivakumar Velusamy |
DAC | 2 |
| 2004 | State-Preserving vs. Non-State-Preserving Leakage Control in CachesabstractThis paper compares the effectiveness of state-preserving and non-state-preserving techniques for leakage control in caches by comparing drowsy cache and gated-V/sub ss/ for data caches using 70nm technology parameters. To perform the comparison, we introduce "HotLeakage", a new architectural model for subthreshold and gate leakage that explicitly models the effects of temperature, voltage, and parameter variations, and has the ability to recalculate leakage currents dynamically as temperature and voltage change at runtime due to operating conditions, DVS techniques, etc. By comparing drowsy-cache and gated-V/sub ss/ at different L2 latencies and different gate oxide thickness values, we are able to identify a range of operating parameters at which gated-V/sub ss/ is more energy efficient than drowsy-cache, even though gated-V/sub ss/ does not preserve data in cache lines that have been deactivated. We are also able to show potential further benefits of gated-V/sub ss/ if an effective dynamic adaptation technique can be found. These results debunk a fairly widespread belief that state-preserving techniques are inherently superior to non-state-preserving techniques. Yingmin Li, Dharmesh Parikh, Yan Zhang 0028, Karthik Sankaranarayanan, Mircea R. Stan, Kevin Skadron |
DATE | 5 |
| 2004 | A Unified Design Space for Regular Parallel Prefix AddersabstractWe consider sparsity, fanout, and radix as three dimensions in the design space of regular parallel prefix adders and present a unified formalism to describe such structures. Matthew M. Ziegler, Mircea R. Stan |
DATE | 2 |
| 2004 | Interconnect lifetime prediction under dynamic stress for reliability-aware designabstractThermal effects are becoming a limiting factor in high-performance circuit design due to the strong temperature-dependence of leakage power, circuit performance, IC package cost and reliability. While many interconnect reliability models assume a constant temperature, this paper presents a physics-based model for estimating interconnect lifetime for any time-varying temperature/current profile. This model is verified with numerical solutions. With this model, we show that designers may be more aggressive with the temperature profiles that are allowed on a chip. In fact, our model reveals that when the temperature magnitude variation is small, average temperature (instead of worst-case temperature) can be used to accurately predict interconnect lifetime, allowing for significant design margin reclamation in reliability-aware design. Even when the variation of temperature magnitude is large, our model shows that using the maximum temperature is still too conservative for interconnect lifetime prediction. Therefore, our model not only increases the accuracy of reliability estimates, but also enables designers to consider more aggressive designs. This model is similarly useful for temperature-aware dynamic runtime management. Zhijian Lu, Wei Huang 0004, John C. Lach, Mircea R. Stan, Kevin Skadron |
ICCAD | 4 |
| 2004 | Temperature-aware microarchitecture: Modeling and implementationabstractWith cooling costs rising exponentially, designing cooling solutions for worst-case power dissipation is prohibitively expensive. Chips that can autonomously modify their execution and power-dissipation characteristics permit the use of lower-cost cooling solutions while still guaranteeing safe temperature regulation. Evaluating techniques for thisdynamic thermal management(DTM), however, requires a thermal model that is practical for architectural studies.This paper describesHotSpot, an accurate yet fast and practical model based on an equivalent circuit of thermal resistances and capacitances that correspond to microarchitecture blocks and essential aspects of the thermal package. Validation was performed using finite-element simulation. The paper also introduces several effective methods for DTM: "temperature-tracking" frequency scaling, "migrating computation" to spare hardware units, and a "hybrid" policy that combines fetch gating with dynamic voltage scaling. The latter two achieve their performance advantage by exploiting instruction-level parallelism, showing the importance of microarchitecture research in helping control the growth of cooling costs.Modeling temperature at the microarchitecture level also shows that power metrics are poor predictors of temperature, that sensor imprecision has a substantial impact on the performance of DTM, and that the inclusion of lateral resistances for thermal diffusion is important for accuracy. Kevin Skadron, Mircea R. Stan, Karthik Sankaranarayanan, Wei Huang 0004, Sivakumar Velusamy, David Tarjan |
ACM Trans. Archit. Code Optim. | 2 |
| 2004 | Power-Aware Branch Prediction: Characterization and DesignabstractThis uses Wattch and the SPEC 2000 integer and floating-point benchmarks to explore the role of branch predictor organization in power/energy/performance trade offs for processor design. Even though the direction predictor by itself represents less than 1 percent of the processor's total power dissipation, prediction accuracy is nevertheless a powerful lever on processor behavior and program execution time. A thorough study of branch predictor organizations shows that, as a general rule, to reduce overall energy consumption in the processor, it is worthwhile to spend more power in the branch predictor if this results in more accurate predictions that improve running time. This not only improves performance, but can also improve the energy-delay product by up to 20 percent. Three techniques, however, can reduce power dissipation without harming accuracy. Banking reduces the portion of the branch predictor that is active at any one time. A new on-chip structure, the prediction probe detector (PPD), uses predecode bits to entirely eliminate unnecessary predictor and branch target buffer (BTB) accesses. Despite the extra power that must be spent accessing it, the PPD reduces local predictor power and energy dissipation by about 31 percent and overall processor power and energy dissipation by 3 percent. These savings can be further improved by using profiling to annotate branches, identifying those that are highly biased and do not require static prediction. Finally, we explore the effectiveness of a previously proposed technique, pipeline gating, and find that, even with adaptive control based on recent predictor accuracy, pipeline gating yields little or no energy savings. Dharmesh Parikh, Kevin Skadron, Yan Zhang 0028, Mircea R. Stan |
IEEE Trans. Computers | 4 |
| 2004 | Large-signal two-terminal device model for nanoelectronic circuit analysisabstractAs the nanoelectronics field reaches the maturity needed for circuit-level integration, modeling approaches are needed that can capture nonclassical behaviors in a compact manner. This paper proposes a universal device model (UDM) for two-terminal devices that addresses the challenge of correctly balancing accuracy, complexity, and flexibility. The UDM qualitatively captures fundamental classical and quantum phenomena and enables nanoelectronic circuit design and simulation. We discuss the motivation behind this modeling approach as well as the underlying details of the model. Furthermore, we present a circuit example of the model in action. Garrett S. Rose, Matthew M. Ziegler, Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | Reducing Multimedia Decode Power using Feedback ControlabstractDespite recent advances, battery life continues to be a limiting factor in mobile multimedia systems. Significant energy savings can be achieved by adapting systems at runtime to match the execution requirements of different tasks. We introduce an online dynamic voltage/frequency scaling (DVS) feedback technique that reduces voltage and frequency to match the playback rate. A PI controller adjusts the decoder's speed to keep constant the occupancy of the buffer between the decoder and the display, effectively matching the average decode rate to the display rate without the need for any off-line profiling. MPEG simulation results show that this technique reduces decoder power consumption while providing strong real-time guarantees. Zhijian Lu, John C. Lach, Mircea R. Stan, Kevin Skadron |
ICCD | 3 |
| 2003 | Temperature-Aware MicroarchitectureabstractWith power density and hence cooling costs rising exponentially, processor packaging can no longer be designed for the worst case, and there is an urgent need for runtime processor-level techniques that can regulate operating temperature when the package's capacity is exceeded. Evaluating such techniques, however, requires a thermal model that is practical for architectural studies.This paper describes HotSpot, an accurate yet fast model based on an equivalent circuit of thermal resistances and capacitances that correspond to microarchitecture blocks and essential aspects of the thermal package. Validation was performed using finite-element simulation. The paper also introduces several effective methods for dynamic thermal management (DTM): "temperature-tracking" frequency scaling, localized toggling, and migrating computation to spare hardware units. Modeling temperature at the microarchitecture level also shows that power metrics are poor predictors of temperature, and that sensor imprecision has a substantial impact on the performance of DTM. Kevin Skadron, Mircea R. Stan, Wei Huang 0004, Sivakumar Velusamy, Karthik Sankaranarayanan, David Tarjan |
ISCA | 2 |
| 2003 | Molecular electronics: from devices and interconnect to circuits and architectureabstractAs the dominating CMOS technology is fast approaching a "brick wall," new opportunities arise for competing solutions. Nanoelectronics has achieved several breakthroughs lately and promises to overcome many of the limitations intrinsic to current semiconductor approaches. Most of the results in this area reported until now focus on devices and interconnect; this work goes several steps further and presents issues related to circuits and architecture. Based on proposed nanoscale interconnect and device structures, we explore the design space available to the nanoelectronic circuit designer and system architect. Mircea R. Stan, Paul D. Franzon, Seth Copen Goldstein, John C. Lach, Matthew M. Ziegler |
Proc. IEEE | 1 |
| 2002 | Control-theoretic dynamic frequency and voltage scaling for multimedia workloadsabstractThis paper describes a formal feedback-control algorithm for dynamic voltage/frequency scaling (DVS) in a portable multimedia system to save power while maintaining a desired playback rate. Our algorithm is similar in complexity to the previously-proposed change-point detection algorithm [19] but does a better job of maintaining stable throughput and is not dependent on the assumption of an exponential distribution of the frame decoding rate. For approximately the same energy savings as reported by [19], our controller is able to keep the average frame delay within 10% of the target more than 90% of the time, whereas the change-point detection algorithm kept the average frame delay with 10% of the target only 70% or less of the time executing the same workload. Zhijian Lu, Jason Hein, Marty Humphrey, Mircea R. Stan, John C. Lach, Kevin Skadron |
CASES | 4 |
| 2002 | The Selective Pull-Up (SP) Noise Immunity Scheme for Dynamic CircuitsabstractSummary form only given. Noise is an important consideration in the design of integrated circuits. Increased immunity to noise, however, typically comes at the expense of increased delay. So, it is very important to have an adequate noise immunity with a minimum penalty in performance. "Global" noise immunity schemes can be used when the noise is approximately the same on all nodes in the circuit; but when a few nodes are noisier then others much better results can be obtained by selective noise immunity schemes. The selective pull-up (SP) technique for dynamic circuits is a method for improving the noise immunity of inputs selectively, so that the least penalty in delay is paid for inputs that intrinsically have higher noise immunity. Mircea R. Stan, Avishek Panigrahi |
DATE | 1 |
| 2002 | Power Issues Related to Branch PredictionabstractThis paper explores the role of branch predictor organization in power/energy/performance tradeoffs for processor design. We find that as a general rule, to reduce overall energy consumption in the processor it is worthwhile to spend more power in the branch predictor if this results in more accurate predictions that improve running time. Two techniques, however, provide substantial reductions in power dissipation without harming accuracy. Banking reduces the portion of the branch predictor that is active at any one time. And a new on-chip structure, the prediction probe detector (PPD), can use pre-decode bits to entirely eliminate unnecessary predictor and branch target buffer (BTB) accesses. Despite the extra power that must be spent accessing the PPD, it reduces local predictor power and energy dissipation by about 45% and overall processor power and energy dissipation by 5-6%. Dharmesh Parikh, Kevin Skadron, Yan Zhang 0028, Marco Barcella, Mircea R. Stan |
HPCA | 5 |
| 2002 | Control-Theoretic Techniques and Thermal-RC Modeling for Accurate and Localized Dynamic Thermal ManagementabstractThis paper proposes the use of formal feedback control theory as a way to implement adaptive techniques in the processor architecture. Dynamic thermal management (DTM) is used as a test vehicle, and variations of a PID controller (Proportional-Integral-Differential) are developed and tested for adaptive control of fetch "toggling." To accurately test the DTM mechanism being proposed, this paper also develops a thermal model based on lumped thermal resistances and thermal capacitances. This model is computationally efficient and tracks temperature at the granularity of individual functional blocks within the processor. Because localized heating occurs much faster than chip-wide heating, some parts of the processor are more likely, to be "hot spots" than others. Experiments using Wattch and the SPEC2000 benchmarks show that the thermal trigger threshold can be set within 0.2/spl deg/ of the maximum temperature and yet never enter thermal emergency. This cuts the performance loss of DTM by 65% compared to the previously described fetch toggling technique that uses a response of fixed magnitude. Kevin Skadron, Tarek F. Abdelzaher, Mircea R. Stan |
HPCA | 3 |
| 2002 | A Case for CMOS/nano co-designabstractThe challenge of extending Moore's Law past the physical and economic barriers of present semiconductor technologies calls for novel nanoelectronic solutions. Circuits composed of mixed silicon semiconductors and nanoelectronics can provide a means for gradually switching technology paradigms. We suggest a design methodology to accompany this concept. Furthermore, we explore design tradeoffs for a nanoscale crossbar technology that supports CMOS/nano co-design. Matthew M. Ziegler, Mircea R. Stan |
ICCAD | 2 |
| 2002 | Circuit-level techniques to control gate leakage for sub-100nm CMOSabstractAlthough still negligible for state-of-the-art CMOS, gate leakage will become significant in the future for sub-100nm technologies, due to the scaling of oxide thickness. We propose several circuit techniques to control gate leakage based on the fact that PMOS transistors with SiO2 gate oxide have an order of magnitude smaller gate leakage than NMOS transistors in the same technology. First, we compare n-type domino with p-type domino circuits in terms of performance, leakage and switching power, and explore the different tradeoffs between performance and power. Second, we compare n-type with p-type gating for MTCMOS to control the leakage during sleep. The proposed circuits are simulated for a predictive 70nm CMOS technology with 10Å gate oxide thickness and 1.2V supply voltage. Fatih Hamzaoglu, Mircea R. Stan |
ISLPED | 2 |
| 2002 | Odd/even bus invert with two-phase transfer for buses with couplingabstractThe coupling capacitances between on-chip bus lines become dominant in deep-submicron technologies. Coding to reduce the switching activity of the individual lines was enough to reduce power on buses in older technologies, but new coding techniques that reduce the coupling activity between lines are needed for deep-submicron buses. One such coding technique uses the simple observation that coupling capacitances are always charged and discharged by activity on neighboring bus lines, where one line has an odd number and the other has an even number (if bus lines are numbered "in-order"). We thus propose to reduce the coupling activity by independently controlling the odd and even bus lines with two separate lines, the Odd Invert, and Even Invert line, respectively. We obtain significant reductions in power simply by comparing the coupling activity for the four possible cases of the Odd and Even Invert lines (00, 01, 10, 11), and then choosing the value with the smallest coupling activity to transmit on the bus. Even after encoding, the coupling activity for a pair of bus lines is still strongly dependent on the data. In particular the toggling sequences 01→10 and 10→01 result in 4 times more coupling energy dissipation than other coupling events. We thus propose a targeted Two-Phase transfer in order to reduce total power only on the pairs of lines that carry such toggling events. Yan Zhang 0028, John C. Lach, Kevin Skadron, Mircea R. Stan |
ISLPED | 4 |
| 2002 | Analysis of dual-VT SRAM cells with full-swing single-ended bit line sensing for on-chip cacheabstractThis paper compares different high-V/sub T/ and dual-V/sub T/ design choices for a large on-chip cache with single-ended sensing in a 0.13 /spl mu/m technology generation. The analysis shows that the best design is the one using a dual-V/sub T/ cell, with minimum channel length pass transistors, and low-V/sub T/ peripheral circuits. This dual-V/sub T/ circuit provides 20% performance gain with only 1.3/spl times/ larger active leakage power, and 2.4% larger cell area compared to the best design using high-V/sub T/ cells with nonminimum channel length pass transistors. Fatih Hamzaoglu, Yibin Ye, Ali Keshavarzi, Kevin Zhang 0001, Siva G. Narendra, Shekhar Borkar, Mircea R. Stan, Vivek De |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2001 | Low-power CMOS with subvolt supply voltagesabstractWe first present a circuit taxonomy along the space and time dimensions, which is useful for classifying generic low-power techniques, followed by an analysis of optimal power supply and threshold voltages and transistor sizing for minimizing the energy-delay product of a class of complementary metal-oxide-semiconductor (CMOS) digital circuits. Mircea R. Stan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1999 | Challenges in clockgating for a low power ASIC methodologyabstractGating the clock is an important technique used in low power design to disable unused modules of a circuit, Gating can save power by both preventing unnecessary activity in the logic modules as well as by eliminating power dissipation in the clock distribution network.There is an inherent pitfall though in implementing gating groups for hierarchical gated clock distribution because the groups are typically developed at the logic level with no information of the physical layout of the clocktree.Depending on the distribution of underlying sinks, maintaining gating groups can cause a wiring overhead that is potentially greater than the savings due to reduced switching.We look at modtjications of zero- skew tree algorithms to consider both the physical and logical aspects of hierarchical gating.The algorithms are applied to data taken from a low power ASK design.The best gated clocktree is created using both physical and logical information. David Garrett, Mircea R. Stan, Alvar Dean |
ISLPED | 2 |
| 1998 | Low power architecture of the soft-output Viterbi algorithmabstractAn important technique for reducing pow er consumption in VLSI systems is strength reduction, the substitution of a less-costly operation such as a shift, for a more-costly operation such a multiplication. Using a logarithmic number represen tation provides sev eral opportunities for strength reductions; in particular, m ultiplicationis performed as the fixed-point addition of logarithms, and extracting a square root is implemented via a shift. These reductions occur transparently at the hardware level; consequently relativ ely little algorithmic modification is required, and they are readily applicable to adaptive filtering. For performing Givens rotations in the QR decomposition recursiv e least squares adaptive filter, logarithmic arithmetic is shown to compare favorably to other strength reduction techniques, such as CORDIC arithmetic, in terms of switched capacitance and numerical accuracy. David Garrett, Mircea R. Stan |
ISLPED | 2 |
| 1998 | Low threshold CMOS circuits with low standby currentabstractMulti-Voltage CMOS (MVCMOS) is a design methodology for very low power supply voltages that uses low-threshold transistors in series with the supply rails. The control voltages on the gating transistors need to be outside of the Vdd - Vss range (hence the name MVCMOS) in order to reduce the standby current, but the resulting circuits operate at lower supply voltages and have a lower area overhead than the previously proposed Multi-Threshold CMOS (MTCMOS). Mircea R. Stan |
ISLPED | 1 |
| 1998 | Long and Fast Up/Down CountersabstractThis paper presents recent advances in the design of constant-time up/down counters in the general context of fast counter design. An overview of existing techniques for the design of long and fast counters reveals several methods closely related to the design of fast adders, as well as some techniques that are only valid for counter design. The main idea behind the novel up/down counters is to recognize that the only extra difficulty with an up/down (vs. up-only or down-only) counter is when the counter changes direction from counting up to counting down (and vice-versa). For dealing with this difficulty, the new design uses a "shadow" register for storing the previous counter state. When counting only up or only down, the counter functions like a standard up-only or down-only constant time counter, but, when it changes direction instead of trying to compute the new value (which typically requires carry propagation), it simply uses the contents of the shadow register which contains the exact desired previous value. An alternative approach for restoring the previous state in constant time is to store the carry bits in a Carry/Borrow register. Mircea R. Stan, Alexandre F. Tenca, Milos D. Ercegovac |
IEEE Trans. Computers | 1 |
| 1997 | Synchronous Up/Down Counter with Clock Period Independent of Counter SizeabstractThe theory and practice of up-only or down-only prescaled (or constant time) counters is well understood both in industry and in the academia. Such counters are obtained by partitioning the counter into sub-blocks in order to be able to anticipate the CARRY propagation inside each block (similar to a carry-select adder). When properly designed, prescaled counters have a clock period independent of counter size. Until now it was not known whether it is possible to design a constant time up/down binary counter. The paper presents the theory behind building a synchronous up/down counter of arbitrary length and with period independent of counter size. The main idea behind the novel up/down counter is to recognize that the only extra difficulty with an up/down (vs. up-only or down-only) constant time counter is when the counter changes "direction" from counting up to counting down and vice-versa. For dealing with this difficulty the new design uses a "shadow" register inside each sub-block with the purpose of always storing the previous block value. When counting only up or only down the counter function like a standard up-only or down-only constant time counter but when it changes direction, instead of trying to compute the new value (which typically requires carry propagation), it simply uses the contents of the shadow register which contains the exact desired previous value. A 54-bit up/down counter running at 40 MHz was implemented in an Atmel AT6000 FPGA and similar up/down counters can be implemented in any technology. Mircea R. Stan |
IEEE Symposium on Computer Arithmetic | 1 |
| 1997 | Power reduction techniques for a spread spectrum based correlatorabstractThis paper presents the design of a low power spread spec-trum correlator. We look at two major approaches and eval-uate the best alternative for power reduction. We first consider a shift register FIFO implementation and look at reducing the switching activity for the arithmetic operations with a change in the addition algorithm. The correlation calculation can be modified to include storage of the previ-ous result so that arithmetic circuits need only compute the difference between the present and next value. A binary adder tree with bypass can then reduce power by shutting off unnecessary computations. We then look at minimizing the power for sample storage by limiting the amount of data moved per cycle. This can be achieved by using a register file FIFO implementation. Interestingly, the two power min-imization techniques, bypass adder tree and register file FIFO implementation, were found to be strongly non-orthogonal, with the final effect that the register file changes the data statistics in such a way that it cancels the savings for the adder tree with bypass. The final solution of a register file with standard adder tree was found to have the lowest power dissipation. Using Bus-Invert for encod-ing the data as it enters the FIFO further reduces the power consumption due to the global bus of the register file. David Garrett, Mircea R. Stan |
ISLPED | 2 |
| 1997 | Low-power encodings for global communication in CMOS VLSIabstractTechnology trends and especially portable applications are adding a third dimension (power) to the previously two-dimensional (speed, area) VLSI design space. A large portion of power dissipation in high performance CMOS VLSI is due to the inherent difficulties in global communication at high rates and we propose several approaches to address the problem. These techniques can be generalized at different levels in the design process. Global communication typically involves driving large capacitive loads which inherently require significant power. However, by carefully choosing the data representation, or encoding, of these signals, the average and peak power dissipation can be minimized. Redundancy can be added in space (number of bus lines), time (number of cycles) and voltage (number of distinct amplitude levels). The proposed codes can be used on a class of terminated off-chip board-level buses with level signaling, or on tristate on-chip buses with level or transition signaling. Mircea R. Stan, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1996 | Two dimensional codes for low powerabstractCoding was previously proposed for reducing power consumption in CMOS. The original formulations use extra redundancy in space (number of bus lines) for reducing the bus transition activity (and consequently the dynamic power and simultaneous switching noise). This paper proposes several new coding techniques for low power. First it looks at codes in which redundancy in time is used for reduced bus activity. Two-dimensional codes with redundancy in both time and space can then be developed for extra power reduction. Interestingly, these two-dimensional codes can be unrolled in either space or time in order to obtain new one-dimensional codes in the other dimension. More powerful codes using Run-Length Limited (RLL), phase-modulation and amplitude-modulation techniques are finally proposed. Mircea R. Stan, Wayne P. Burleson |
ISLPED | 1 |
| 1995 | Coding a terminated bus for low powerabstractCoding was proposed as a general method of decreasing power dissipation for the I/O. Lower power dissipation can be obtained by using extra bus liner for coding the data. This paper presents an application of the general theory of limited-weight codes for a class of parallel terminated buses with pull-up terminators (e.g. Rambus). Power dissipation on such a bus-line is larger for a logical 1 and it follows that patterns with few 1s should be chosen. A perfect k/2-limited weight code equivalent to the previously proposed Bus-Invert method and a novel non-perfect 3-limited weight code are described. Both codes can be algorithmically generated and practical issues related to their implementation on the Rambus are discussed. Mircea R. Stan, Wayne P. Burleson |
Great Lakes Symposium on VLSI | 1 |
| 1995 | Bus-invert coding for low-power I/OabstractTechnology trends and especially portable applications drive the quest for low-power VLSI design. Solutions that involve algorithmic, structural or physical transformations are sought. The focus is on developing low-power circuits without affecting too much the performance (area, latency, period). For CMOS circuits most power is dissipated as dynamic power for charging and discharging node capacitances. This is why many promising results in low-power design are obtained by minimizing the number of transitions inside the CMOS circuit. While it is generally accepted that because of the large capacitances involved much of the power dissipated by an IC is at the I/O little has been specifically done for decreasing the I/O power dissipation. We propose the bus-invert method of coding the I/O which lowers the bus activity and thus decreases the I/O peak power dissipation by 50% and the I/O average power dissipation by up to 25%. The method is general but applies best for dealing with buses. This is fortunate because buses are indeed most likely to have very large capacitances associated with them and consequently dissipate a lot of power.> Mircea R. Stan, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |