Engin Ipek

dblp:i/EnginIpek · DBLP profile ↗
← Back
37ranked-venue papers
8as first author
2since 2021 · last 2025
0000-0003-2809-5809ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 8 first-author · 2 since 2021Software engineering, systems software and programming languages · 13 · 4 first-author · 1 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
25 papers
Memory systems · 36% Storage systems · 15% Hardware accelerators and domain-specific architectures · 13%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 30 heaviest of 68, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
0.962018
Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018
AC-DIMM: associative computing with STT-MRAM · ISCA 2013
Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computing · ISCA 2010
Storage systems › computational storage
in-storage computing
0.912025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025
Storage systems
key-value storage
0.912025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025
Hardware accelerators and domain-specific architectures › accelerator integration
near-storage accelerator
0.912025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025
Memory systems
DRAM
0.752019
Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019
Architecting phase change memory as a scalable dram alternative · ISCA 2009
Self-Optimizing Memory Controllers: A Reinforcement Learning Approach · ISCA 2008
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.632018
Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018
AC-DIMM: associative computing with STT-MRAM · ISCA 2013
Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computing · ISCA 2010
Memory systems
processing-in-memory
0.532018
Memristive Boltzmann machine: A hardware accelerator for combinatorial optimization and deep learning · HPCA 2016
AC-DIMM: associative computing with STT-MRAM · ISCA 2013
Enabling Scientific Computing on Memristive Accelerators · ISCA 2018
Memory systems › in-memory computing
in-situ computing accelerator
0.512021
An Analog Preconditioner for Solving Linear Systems · HPCA 2021
High-performance computing › numerical linear algebra
preconditioner
0.512021
An Analog Preconditioner for Solving Linear Systems · HPCA 2021
High-performance computing
scientific computing
0.512021
An Analog Preconditioner for Solving Linear Systems · HPCA 2021
Memory systems › non-volatile memory
resistive memory
0.532016
Memristive Boltzmann machine: A hardware accelerator for combinatorial optimization and deep learning · HPCA 2016
A resistive TCAM accelerator for data-intensive computing · MICRO 2011
Dynamically replicated memory: building reliable systems from nanoscale resistive memories · ASPLOS 2010
Hardware reliability and fault tolerance
error correction
0.422019
Making Memristive Neural Network Accelerators Reliable · HPCA 2018
Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019
Hardware reliability and fault tolerance
memory reliability
0.422018
Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018
Dynamically replicated memory: building reliable systems from nanoscale resistive memories · ASPLOS 2010
Hardware accelerators and domain-specific architectures › machine learning accelerator › inference accelerator
neural network inference accelerator
0.412020
Commutative Data Reordering: A New Technique to Reduce Data Movement Energy on Sparse Inference Workloads · ISCA 2020
Memory systems
memory controller
0.432013
A programmable memory controller for the DDRx interfacing standards · ACM Trans. Comput. Syst. 2013
PARDIS: A programmable memory controller for the DDRx interfacing standards · ISCA 2012
Self-Optimizing Memory Controllers: A Reinforcement Learning Approach · ISCA 2008
Energy-efficient computing › energy-efficient communication
interconnect energy
0.422015
More is less: improving the energy efficiency of data movement via opportunistic use of sparse codes · MICRO 2015
DESC: energy-efficient data exchange using synchronized counters · MICRO 2013
Energy-efficient computing
power management
0.432017
Voltage Regulator Efficiency Aware Power Management · ASPLOS 2017
A programmable memory controller for the DDRx interfacing standards · ACM Trans. Comput. Syst. 2013
PARDIS: A programmable memory controller for the DDRx interfacing standards · ISCA 2012
Memory systems › DRAM
DRAM refresh
0.412019
Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019
Storage systems › storage reliability › durability
retention failure
0.412019
Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019
Hardware reliability and fault tolerance
arithmetic codes
0.312018
Making Memristive Neural Network Accelerators Reliable · HPCA 2018
Hardware reliability and fault tolerance
error-correcting codes for memory
0.312018
Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › in-memory computing accelerator
memristive crossbar accelerator
0.312018
Enabling Scientific Computing on Memristive Accelerators · ISCA 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.312018
Making Memristive Neural Network Accelerators Reliable · HPCA 2018
Storage systems › storage reliability
scrubbing
0.312018
Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018
High-performance computing
sparse linear solver
0.312018
Enabling Scientific Computing on Memristive Accelerators · ISCA 2018
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation
0.312018
Enabling Scientific Computing on Memristive Accelerators · ISCA 2018
Memory systems
content-addressable memory
0.322013
AC-DIMM: associative computing with STT-MRAM · ISCA 2013
A resistive TCAM accelerator for data-intensive computing · MICRO 2011
Processor architecture and microarchitecture
chip multiprocessor
0.342010
Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computing · ISCA 2010
Coordinated management of multiple interacting resources in chip multiprocessors: A machine learning approach · MICRO 2008
Core fusion: accommodating software diversity in chip multiprocessors · ISCA 2007
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.312017
Voltage Regulator Efficiency Aware Power Management · ASPLOS 2017
Memory systems › processing-in-memory
near-data processing
0.312025
ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025

Methods — techniques the papers use, named apart from their topics

domain decomposition · 0.5bit-slicing · 0.5traveling salesman problem · 0.4capacitated vehicle routing problem · 0.4approximation method · 0.4error-correcting codes · 0.4operation scheduling · 0.3fixed-point emulation of floating point · 0.3data-aware encoding · 0.3arithmetic coding · 0.3processing-in-memory · 0.2boltzmann machine · 0.2RRAM · 0.2
YearPublicationVenuePosition
2025 ANVIL: An In-Storage Accelerator for Name-Value Data Stores
abstract
Name-value pairs (NVPs) are a widely-used abstraction to organize data in millions of applications.At a high level, an NVP associates a name (e.g., array index, key, hash) with each value in a collection of data.Specific NVP data store formats can vary widely, ranging from simple arrays/dictionaries and lookup tables to key-value stores and data mining workloads.Despite their importance, existing optimizations for NVPs are limited to only a single data store format, as the broad definition of NVPs allows for significant heterogeneity in encoding and implementation.We propose ANVIL, the first end-to-end system that allows programmers to broadly accelerate most formats of NVPs.With a conventional solid-state drive (SSD), large-scale NVP lookups can saturate both external and internal SSD bandwidth, as every NVP in the data store needs to be sent back to the host CPU to check for a matching name.ANVIL makes use of in-storage processing to avoid reading out any data for names that do not match, by performing name match checks directly inside the SSD's NAND flash chips.We demonstrate that ANVIL can substantially reduce disk I/O, reduce metadata overheads, and provide speedups of 4.0×, 25×, and 14.6% over a conventional SSD, for three different NVP workloads (database transactions, analytics, and graph processing).
Ryan Wong 0001, Nikita Kim, Aniket Das, Kevin Higgs, Engin Ipek, Sapan Agarwal, Saugata Ghose, Ben Feinberg
ISCA5
2021 An Analog Preconditioner for Solving Linear Systems
abstract
Over the past decade as Moore's Law has slowed, the need for new forms of computation that can provide sustainable performance improvements has risen. A new method, called in situ computing, has shown great potential to accelerate matrix vector multiplication (MVM), an important kernel for a diverse range of applications from neural networks to scientific computing. Existing in situ accelerators for scientific computing, however, have a significant limitation: these accelerators provide no acceleration for preconditioning-a key bottleneck in linear solvers and in scientific computing workflows. This paper enables in situ acceleration for state-of-the-art linear solvers by demonstrating how to use a new in situ matrix inversion accelerator for analog preconditioning. As existing techniques that enable high precision and scalability for in situ MVM are inapplicable to in situ matrix inversion, new techniques to compensate for circuit non-idealities are proposed. Additionally, a new approach to bit slicing that enables splitting operands across multiple devices without external digital logic is proposed. For scalability, this paper demonstrates how in situ matrix inversion kernels can work in tandem with existing domain decomposition techniques to accelerate the solutions of arbitrarily large linear systems. The analog kernel can be directly integrated into existing preconditioning workflows, leveraging several well-optimized numerical linear algebra tools to improve the behavior of the circuit. The result is an analog preconditioner that is more effective (up to 50% fewer iterations) than the widely used incomplete LU factorization preconditioner, ILU(0), while also reducing the energy and execution time of each approximate solve operation by 1025x and 105x respectively.
Ben Feinberg, Ryan Wong 0001, T. Patrick Xiao, Christopher H. Bennett, Jacob N. Rohan, Erik G. Boman, Matthew J. Marinella, Sapan Agarwal, Engin Ipek
HPCA9
2020 Commutative Data Reordering: A New Technique to Reduce Data Movement Energy on Sparse Inference Workloads
abstract
Data movement is a significant and growing consumer of energy in modern systems, from specialized low-power accelerators to GPUs with power budgets in the hundreds of Watts. Given the importance of the problem, prior work has proposed designing interconnects on which the energy cost of transmitting a 0 is significantly lower than that of transmitting a 1. With such an interconnect, data movement energy is reduced by encoding the transmitted data such that the number of 1s is minimized. Although promising, these data encoding proposals do not take full advantage of application level semantics. As an example of a neglected optimization opportunity, consider the case of a dot product computation as part of a neural network inference task. The order in which the neural network weights are fetched and processed does not affect correctness, and can be optimized to further reduce data movement energy.This paper presents commutative data reordering (CDR), a hardware-software approach that leverages the commutative property in linear algebra to strategically select the order in which weight matrix coefficients are fetched from memory. To find a low-energy transmission order, weight ordering is modeled as an instance of one of two well-studied problems, the Traveling Salesman Problem and the Capacitated Vehicle Routing Problem. This reduction makes it possible to leverage the vast body of work on efficient approximation methods to find a good transmission order. CDR exploits the indirection inherent to sparse matrix formats such that no additional metadata is required to specify the selected order. The hardware modifications required to support CDR are minimal, and incur an area penalty of less than 0.01% when implemented on top of a mobile-class GPU. When applied to 7 neural network inference tasks running on a GPU-based system, CDR respectively reduces average DRAM IO energy by 53.1% and 22.2% over the data bus invert encoding scheme used by LPDDR4, and the recently proposed Base + XOR encoding. These savings are attained with no changes to the mobile system software and no runtime performance penalty.
Ben Feinberg, Benjamin C. Heyman, Darya Mikhailenko, Ryan Wong 0001, An C. Ho, Engin Ipek
ISCA6
2019 Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data
abstract
DRAM refresh is responsible for significant performance and energy overheads in a wide range of computer systems, from mobile platforms to datacenters [1] . With the growing demand for DRAM capacity and the worsening retention time characteristics of deeply scaled DRAM, refresh is expected to become an even more pronounced problem in future technology generations [2] . This paper examines content aware refresh, a new technique that reduces the refresh frequency by exploiting the unidirectional nature of DRAM retention errors: assuming that a logical 1 and 0 respectively are represented by the presence and absence of charge, 1-to-0 failures are much more likely than 0-to-1 failures. As a result, in a DRAM system that uses a block error correcting code (ECC) to protect memory, blocks with fewer 1s can attain a specified reliability target (i.e., mean time to failure) with a refresh rate lower than that which is required for a block with all 1s. Leveraging this key insight, and without compromising memory reliability, the proposed content aware refresh mechanism refreshes memory blocks with fewer 1s less frequently. To keep the overhead of tracking multiple refresh rates manageable, refresh groups-groups of DRAM rows refreshed together-are dynamically arranged into one of a predefined number of refresh bins and refreshed at the rate determined by the ECC block with the greatest number of 1s in that bin. By tailoring the refresh rate to the actual content of a memory block rather than assuming a worst case data pattern, content aware refresh respectively outperforms DRAM systems that employ RAS-only Refresh, all-bank Auto Refresh, and per-bank Auto Refresh mechanisms by 12, 8, and 13 percent. It also reduces DRAM system energy by 15, 13, and 16 percent as compared to these systems.
Mahdi Nazm Bojnordi, Engin Ipek
IEEE Trans. Computers4
2018 Making Memristive Neural Network Accelerators Reliable
abstract
Deep neural networks (DNNs) have attracted substantial interest in recent years due to their superior performance on many classification and regression tasks as compared to other supervised learning models. DNNs often require a large amount of data movement, resulting in performance and energy overheads. One promising way to address this problem is to design an accelerator based on in-situ analog computing that leverages the fundamental electrical properties of memristive circuits to perform matrix-vector multiplication. Recent work on analog neural network accelerators has shown great potential in improving both the system performance and the energy efficiency. However, detecting and correcting the errors that occur during in-memory analog computation remains largely unexplored. The same electrical properties that provide the performance and energy improvements make these systems especially susceptible to errors, which can severely hurt the accuracy of the neural network accelerators. This paper examines a new error correction scheme for analog neural network accelerators based on arithmetic codes. The proposed scheme encodes the data through multiplication by an integer, which preserves addition operations through the distributive property. Error detection and correction are performed through a modulus operation and a correction table lookup. This basic scheme is further improved by data-aware encoding to exploit the state dependence of the errors, and by knowledge of how critical each portion of the computation is to overall system accuracy. By leveraging the observation that a physical row that contains fewer 1s is less susceptible to an error, the proposed scheme increases the effective error correction capability with less than 4.5% area and less than 4.7% energy overheads. When applied to a memristive DNN accelerator performing inference on the MNIST and ILSVRC-2012 datasets, the proposed technique reduces the respective misclassification rates by 1.5x and 1.1x.
Ben Feinberg, Engin Ipek
HPCA3
2018 Enabling Scientific Computing on Memristive Accelerators
abstract
Linear algebra is ubiquitous across virtually every field of science and engineering, from climate modeling to macroeconomics. This ubiquity makes linear algebra a prime candidate for hardware acceleration, which can improve both the run time and the energy efficiency of a wide range of scientific applications. Recent work on memristive hardware accelerators shows significant potential to speed up matrix-vector multiplication (MVM), a critical linear algebra kernel at the heart of neural network inference tasks. Regrettably, the proposed hardware is constrained to a narrow range of workloads: although the eight-to 16-bit computations afforded by memristive MVM accelerators are acceptable for machine learning, they are insufficient for scientific computing where high-precision floating point is the norm. This paper presents the first proposal to enable scientific computing on memristive crossbars. Three techniques are explored — reducing overheads by exploiting exponent range locality, early termination of fixed-point computation, and static operation scheduling — that together enable a fixed-point memristive accelerator to perform high-precision floating point without the exorbitant cost of naïve floating-point emulation on fixed-point hardware. A heterogeneous collection of crossbars with varying sizes is proposed to efficiently handle sparse matrices, and an algorithm for mapping the dense subblocks of a sparse matrix to an appropriate set of crossbars is investigated. The accelerator can be combined with existing GPU-based systems to handle datasets that cannot be efficiently handled by the memristive accelerator alone. The proposed optimizations permit the memristive MVM concept to be applied to a wide range of problem domains, respectively improving the execution time and energy dissipation of sparse linear solvers by 10.3x and 10.9x over a purely GPU-based system.
Ben Feinberg, Uday Kumar Reddy Vengalam, Nathan Whitehair, Engin Ipek
ISCA5
2018 Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories
abstract
DRAM density scaling has become increasingly difficult due to challenges in maintaining a sufficiently high storage capacitance and a sufficiently low leakage current at nanoscale feature sizes. Non-volatile memories (NVMs) have drawn significant attention as potential DRAM replacements because they represent information using resistance rather than electrical charge. Spin-torque transfer magnetoresistive RAM (STT-MRAM) is one of the most promising NVM technologies due to its relatively low write energy, high speed, and high endurance. However, STT-MRAM suffers from its own scaling problems. As the size of the storage element decreases with technology scaling, STT-MRAM retention error rates are expected to increase, which will require multi-bit error-correcting code (ECC) and periodic scrubbing. We introduce the Sanitizer architecture, which mitigates the performance and energy overheads of ECC and scrubbing in future STT-MRAM based main memories. To reduce the scrubbing rate, a coarse-grained, multi-bit ECC mechanism with a 12.5 percent storage overhead is used. To avoid fetching multiple blocks from memory and performing costly ECC checks on every read, the memory regions that will likely be accessed in the near future are predicted and proactively scrubbed. Compared to a conventional STT-MRAM system, Sanitizer improves performance by 1.22× and reduces end-to-end system energy by 22 percent.
Mahdi Nazm Bojnordi, Qing Guo 0004, Engin Ipek
IEEE Trans. Computers4
2017 Voltage Regulator Efficiency Aware Power Management
abstract
Conventional off-chip voltage regulators are typically bulky and slow, and are inefficient at exploiting system and workload variability using Dynamic Voltage and Frequency Scaling (DVFS). On-die integration of voltage regulators has the potential to increase the energy efficiency of computer systems by enabling power control at a fine granularity in both space and time. The energy conversion efficiency of on-chip regulators, however, is typically much lower than off-chip regulators, which results in significant energy losses. Fine-grained power control and high voltage regulator efficiency are difficult to achieve simultaneously, with either emerging on-chip or conventional off-chip regulators.
Victor W. Lee, Engin Ipek
ASPLOS3
2016 Memristive Boltzmann machine: A hardware accelerator for combinatorial optimization and deep learning
abstract
The Boltzmann machine is a massively parallel computational model capable of solving a broad class of combinatorial optimization problems. In recent years, it has been successfully applied to training deep machine learning models on massive datasets. High performance implementations of the Boltzmann machine using GPUs, MPI-based HPC clusters, and FPGAs have been proposed in the literature. Regrettably, the required all-to-all communication among the processing units limits the performance of these efforts. This paper examines a new class of hardware accelerators for large-scale combinatorial optimization and deep learning based on memristive Boltzmann machines. A massively parallel, memory-centric hardware accelerator is proposed based on recently developed resistive RAM (RRAM) technology. The proposed accelerator exploits the electrical properties of RRAm to realize in situ, fine-grained parallel computation within memory arrays, thereby eliminating the need for exchanging data between the memory cells and the computational units. Two classical optimization problems, graph partitioning and boolean satisfiability, and a deep belief network application are mapped onto the proposed hardware. As compared to a multicore system, the proposed accelerator achieves 57x higher performance and 25x lower energy with virtually no loss in the quality of the solution to the optimization problems. The memristive accelerator is also compared against an RRAM based processing-in-memory (PIM) system, with respective performance and energy improvements of 6.89x and 5.2x.
Mahdi Nazm Bojnordi, Engin Ipek
HPCA2
2016 Reducing data movement energy via online data clustering and encoding
abstract
Modern computer systems expend significant amounts of energy on transmitting data over long and highly capacitive interconnects. A promising way of reducing the data movement energy is to design the interconnect such that the transmission of 0s is considerably cheaper than that of 1s. Given such an interconnect with asymmetric transmission costs, data movement energy can be reduced by encoding the transmitted data such that the number of 1s in each transmitted codeword is minimized. This paper presents a new data encoding technique based on online data clustering that exploits this opportunity. The transmitted data blocks are dynamically clustered based on the similarities between their binary representations. Each data block is expressed as the bitwise XOR between one of multiple cluster centers and a residual with a small number of 1s. The data movement energy is minimized by sending the residual along with an identifier that specifies which cluster center to use in decoding the transmitted data. At runtime, the proposed approach continually updates the cluster centers based on the observed data to adapt to phase changes. The proposed technique is compared to three previously proposed energy-efficient data encoding techniques on a set of 14 applications. The results indicate respective energy savings of 5%, 9%, and 12% in DDR4, LPDDR3, and last level cache subsystems as compared to the best existing baseline encoding technique.
Engin Ipek
MICRO2
2016 Back to the Future: Current-Mode Processor in the Era of Deeply Scaled CMOS
abstract
This paper explores the use of MOS current-mode logic (MCML) as a fast and low noise alternative to static CMOS circuits in microprocessors, thereby improving the performance, energy efficiency, and signal integrity of future computer systems. The power and ground noise generated by an MCML circuit is typically 10-100× smaller than the noise generated by a static CMOS circuit. Unlike static CMOS, whose dominant dynamic power is proportional to the frequency, MCML circuits dissipate a constant power independent of clock frequency. Although these traits make MCML highly energy efficient when operating at high speeds, the constant static power of MCML poses a challenge for a microarchitecture that operates at the modest clock rate and with a low activity factor. To address this challenge, a single-core microarchitecture for MCML is explored that exploits the C-slow retiming technique, and operates at a high frequency with low complexity to save energy. This design principle contrasts with the contemporary multicore design paradigm for static CMOS that relies on a large number of gates operating in parallel at the modest speeds. The proposed architecture generates 10-40× lower power and ground noise, and operates within 13% of the performance (i.e., 1/ExecutionTime) of a conventional, eight-core static CMOS processor while exhibiting 1.6× lower energy and 9% less area. Moreover, the operation of an MCML processor is robust under both systematic and random variations in transistor threshold voltage and effective channel length.
Yanwei Song, Mahdi Nazm Bojnordi, Alexander E. Shapiro, Eby G. Friedman, Engin Ipek
IEEE Trans. Very Large Scale Integr. Syst.6
2016 Reducing Switching Latency and Energy in STT-MRAM Caches With Field-Assisted Writing
abstract
A field-assisted spin-torque transfer magnetoresistive RAM (STT-MRAM) cache is presented for the use in high-performance energy-efficient microprocessors. Adding field assistance reduces the switching latency by a factor of 4. An array model is developed to evaluate the switching energy for different field currents and array sizes. Several STT-MRAM-based cells demonstrate a 55% energy reduction as compared with an SRAM cache subsystem. As compared with STT-MRAM caches with subbank buffering and differential writes, a field-assisted STT-MRAM cache improves the system performance by 28%, with a 6.7% increase in energy.
Ravi Patel 0001, Qing Guo 0004, Engin Ipek, Eby G. Friedman
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Architecting a MOS current mode logic (MCML) processor for fast, low noise and energy-efficient computing in the near-threshold regime
abstract
Near-threshold computing (NTC) is an effective technique for improving the energy efficiency of a CMOS microprocessor, but suffers from a significant performance loss and an increased sensitivity to voltage noise. MOS current-mode logic (MCML), a differential logic family, maintains a low voltage swing and a constant current, making it inherently fast and low-noise. These traits make MCML a natural selection to implement an NTC processor; however, MCML suffers from a high static power regardless of the clock frequency or the level of switching activity, which would result in an inordinate energy consumption in a large scale IC. To address this challenge, this paper explores a single-core microarchitecture for MCML that takes advantage of C-slow retiming technique, and runs at a high frequency with low complexity to save energy. This design principle is opposite to the contemporary multicore design paradigm for static CMOS that relies on a large number of gates running in parallel at modest speeds. When compared to an eight-core static CMOS processor operating in the near-threshold regime, the proposed processor exhibits 3x higher performance, 2x lower energy, and 10 x lower voltage noise, while maintaining a similar level of power dissipation.
Yanwei Song, Mahdi Nazm Bojnordi, Alexander E. Shapiro, Engin Ipek, Eby G. Friedman
ICCD5
2015 Energy-efficient data movement with sparse transition encoding
abstract
Data movement over long on-chip interconnects is a major contributor to system energy. This paper presents novel signaling and encoding techniques that toget her improve the energy efficiency of data communication between the processor cores and the last level cache. The proposed techniques make the interconnect energy proportional to the number of ones in the transferred data block (i.e., the block's hamming weight), regardless of the previous state of the interconnect. The hamming weight of each cache block is kept low through a sparse data encoding approach to minimize interconnect energy. Simulation results show that the proposed communication scheme reduces the overall L2 cache energy by 30% on a set of eleven parallel applications with a 0.5% average performance degradation.
Yanwei Song, Mahdi Nazm Bojnordi, Engin Ipek
ICCD3
2015 Enabling energy efficient Hybrid Memory Cube systems with erasure codes
abstract
The Hybrid Memory Cube (HMC) is a promising alternative to DDRx memory due to its potential to achieve significantly higher bandwidth. However, the high static power of an HMC device compromises power efficiency when the device is lightly utilized. Activating a sleeping HMC takes over 2µs, which makes it challenging to manage HMC power without a substantial degradation in system performance. We introduce a new technique that alleviates the long wakeup penalty of an HMC by employing erasure codes. Inaccessible data stored in a sleeping HMC module can be reconstructed by decoding related data retrieved from other active HMCs, rather than waiting for the sleeping HMC module to become active. This approach makes it possible to tolerate the latency penalty incurred when switching an HMC between active and sleep modes, thereby enabling a power-capped HMC system. Simulations show that the proposed architecture outperforms a current HMC-based multicore system by 6.2×, and reduces the system energy by 5.3× under the same power budget as the multicore baseline.
Yanwei Song, Mahdi Nazm Bojnordi, Engin Ipek
ISLPED4
2015 More is less: improving the energy efficiency of data movement via opportunistic use of sparse codes
abstract
Data movement over long and highly capacitive interconnects is responsible for a large fraction of the energy consumed in nanometer ICs. DDRx, the most broadly adopted family of DRAM interfaces, contributes significantly to the overall system energy in a wide range of computer systems. To reduce the energy cost of data transfers, DDR4 adopts a pseudo open-drain IO circuit that consumes power only when transmitting or receiving a 0, which makes the IO energy proportional to the number of 0s transferred over the data bus. A data bus invert (DBI) coding technique is therefore supported by the DDR4 standard to encode each byte using a small number of 0s. Although sparse coding techniques that are more advanced than DBI can reduce the IO power further, the relatively high bandwidth overhead of these codes has heretofore prevented their application to the DDRx bus.
Yanwei Song, Engin Ipek
MICRO2
2014 Field driven STT-MRAM cell for reduced switching latency and energy
abstract
A field driven approach to STT-MRAM switching is proposed as a method for reducing the switching latency of an MTJ in high performance caches. An MRAM array model is presented to characterize the switching energy and maximum achievable reduction in energy using the field driven approach. The switching latency per bit is reduced by more than a factor of ten. The resultant switching energy per bit is reduced by 82% as compared to a standard STT-MRAM.
Ravi Patel 0001, Engin Ipek, Eby G. Friedman
ISCAS2
2013 AC-DIMM: associative computing with STT-MRAM
abstract
With technology scaling, on-chip power dissipation and off-chip memory bandwidth have become significant performance bottlenecks in virtually all computer systems, from mobile devices to supercomputers. An effective way of improving performance in the face of bandwidth and power limitations is to rely on associative memory systems. Recent work on a PCM-based, associative TCAM accelerator shows that associative search capability can reduce both off-chip bandwidth demand and overall system energy. Unfortunately, previously proposed resistive TCAM accelerators have limited flexibility: only a restricted (albeit important) class of applications can benefit from a TCAM accelerator, and the implementation is confined to resistive memory technologies with a high dynamic range (RHigh/RLow), such as PCM.
Qing Guo 0004, Ravi Patel 0001, Engin Ipek, Eby G. Friedman
ISCA4
2013 DESC: energy-efficient data exchange using synchronized counters
abstract
Increasing cache sizes in modern microprocessors require long wires to connect cache arrays to processor cores. As a result, the last-level cache (LLC) has become a major contributor to processor energy, necessitating techniques to increase the energy efficiency of data exchange over LLC interconnects.
Mahdi Nazm Bojnordi, Engin Ipek
MICRO2
2013 A programmable memory controller for the DDRx interfacing standards
abstract
Modern memory controllers employ sophisticated address mapping, command scheduling, and power management optimizations to alleviate the adverse effects of DRAM timing and resource constraints on system performance. A promising way of improving the versatility and efficiency of these controllers is to make them programmable—a proven technique that has seen wide use in other control tasks, ranging from DMA scheduling to NAND Flash and directory control. Unfortunately, the stringent latency and throughput requirements of modern DDRx devices have rendered such programmability largely impractical, confining DDRx controllers to fixed-function hardware. This article presents the instruction set architecture (ISA) and hardware implementation of PARDIS, a programmable memory controller that can meet the performance requirements of a high-speed DDRx interface. The proposed controller is evaluated by mapping previously proposed DRAM scheduling, address mapping, refresh scheduling, and power management algorithms onto PARDIS. Simulation results show that the average performance of PARDIS comes within 8% of fixed-function hardware for each of these techniques; moreover, by enabling application-specific optimizations, PARDIS improves system performance by 6 to 17% and reduces DRAM energy by 9 to 22% over four existing memory controllers.
Mahdi Nazm Bojnordi, Engin Ipek
ACM Trans. Comput. Syst.2
2012 Overcoming single-thread performance hurdles in the core fusion reconfigurable multicore architecture
abstract
Though the prime target of multicore architectures is parallel and multithreaded workloads (which favors maximum core count), executing sequential code fast continues to remain critical (which benefits from maximum core size). This poses a difficult design trade-off. Core Fusion is a recently-proposed reconfigurable multicore architecture that attempts to circumvent this compromise by "fusing" groups of fundamentally independent cores into larger, more aggressive processors dynamically as needed. In this way, it accommodates highly parallel, partially parallel, multiprogrammed, and sequential codes with ease.
Janani Mukundan, Saugata Ghose, Robert Karmazin, Engin Ipek, José F. Martínez
ICS4
2012 PARDIS: A programmable memory controller for the DDRx interfacing standards
abstract
Modern memory controllers employ sophisticated address mapping, command scheduling, and power management optimizations to alleviate the adverse effects of DRAM timing and resource constraints on system performance. A promising way of improving the versatility and efficiency of these controllers is to make them programmable - a proven technique that has seen wide use in other control tasks ranging from DMA scheduling to NAND Flash and directory control. Unfortunately, the stringent latency and throughput requirements of modern DDRx devices have rendered such programmability largely impractical, confining DDRx controllers to fixed-function hardware. This paper presents the instruction set architecture (ISA) and hardware implementation of PARDIS, a programmable memory controller that can meet the performance requirements of a high-speed DDRx interface. The proposed controller is evaluated by mapping previously proposed DRAM scheduling, address mapping, refresh scheduling, and power management algorithms onto PARDIS. Simulation results show that the average performance of PARDIS comes within 8% of fixed-function hardware for each of these techniques; moreover, by enabling application-specific optimizations, PARDIS improves system performance by 6-17% and reduces DRAM energy by 9-22% over four existing memory controllers.
Mahdi Nazm Bojnordi, Engin Ipek
ISCA2
2011 A resistive TCAM accelerator for data-intensive computing
abstract
Power dissipation and off-chip bandwidth restrictions are critical challenges that limit microprocessor performance. Ternary content addressable memories (TCAM) hold the potential to address both problems in the context of a wide range of data-intensive workloads that benefit from associative search capability. Power dissipation is reduced by eliminating instruction processing and data movement overheads present in a purely RAM based system. Bandwidth demand is lowered by processing data directly on the TCAM chip, thereby decreasing off-chip traffic. Unfortunately, CMOS-based TCAM implementations are severely power- and area-limited, which restricts the capacity of commercial products to a few megabytes, and confines their use to niche networking applications.
Qing Guo 0004, Engin Ipek
MICRO4
2010 Dynamically replicated memory: building reliable systems from nanoscale resistive memories
abstract
DRAM is facing severe scalability challenges in sub-45nm tech- nology nodes due to precise charge placement and sensing hur- dles in deep-submicron geometries. Resistive memories, such as phase-change memory (PCM), already scale well beyond DRAM and are a promising DRAM replacement. Unfortunately, PCM is write-limited, and current approaches to managing writes must de- commission pages of PCM when the first bit fails.
Engin Ipek, Jeremy Condit, Ed Nightingale, Doug Burger, Thomas Moscibroda
ASPLOS1
2010 Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computing
abstract
As CMOS scales beyond the 45nm technology node, leakage concerns are starting to limit microprocessor performance growth. To keep dynamic power constant across process generations, traditional MOSFET scaling theory prescribes reducing supply and threshold voltages in proportion to device dimensions, a practice that induces an exponential increase in subthreshold leakage. As a result, leakage power has become comparable to dynamic power in current-generation processes, and will soon exceed it in magnitude if voltages are scaled down any further. Beyond this inflection point, multicore processors will not be able to afford keeping more than a small fraction of all cores active at any given moment. Multicore scaling will soon hit a power wall. This paper presents resistive computation, a new technique that aims at avoiding the power wall by migrating most of the functionality of a modern microprocessor from CMOS to spin-torque transfer magnetoresistive RAM (STT-MRAM)---a CMOS-compatible, leakage-resistant, non-volatile resistive memory technology. By implementing much of the on-chip storage and combinational logic using leakage-resistant, scalable RAM blocks and lookup tables, and by carefully re-architecting the pipeline, an STT-MRAM based implementation of an eight-core Sun Niagara-like CMT processor reduces chip-wide power dissipation by 1.7× and leakage power by 2.1× at the 32nm technology node, while maintaining 93% of the system throughput of a CMOS-based design.
Engin Ipek, Tolga Soyata
ISCA2
2009 Architecting phase change memory as a scalable dram alternative
abstract
Memory scaling is in jeopardy as charge storage and sensing mechanisms become less reliable for prevalent memory technologies, such as DRAM. In contrast, phase change memory (PCM) storage relies on scalable current and thermal mechanisms. To exploit PCM's scalability as a DRAM alternative, PCM must be architected to address relatively long latencies, high energy writes, and finite endurance.We propose, crafted from a fundamental understanding of PCM technology parameters, area-neutral architectural enhancements that address these limitations and make PCM competitive with DRAM. A baseline PCM system is 1.6x slower and requires 2.2x more energy than a DRAM system. Buffer reorganizations reduce this delay and energy gap to 1.2x and 1.0x, using narrow rows to mitigate write energy and multiple rows to improve locality and write coalescing. Partial writes enhance memory endurance, providing 5.6 years of lifetime. Process scaling will further reduce PCM energy costs and improve endurance.
Benjamin C. Lee, Engin Ipek, Onur Mutlu, Doug Burger
ISCA2
2009 Better I/O through byte-addressable, persistent memory
abstract
Modern computer systems have been built around the assumption that persistent storage is accessed via a slow, block-based interface. However, new byte-addressable, persistent memory technologies such as phase change memory (PCM) offer fast, fine-grained access to persistent storage.
Jeremy Condit, Ed Nightingale, Christopher Frost 0001, Engin Ipek, Benjamin C. Lee, Doug Burger, Derrick Coetzee
SOSP4
2008 Self-Optimizing Memory Controllers: A Reinforcement Learning Approach
abstract
Efficiently utilizing off-chip DRAM bandwidth is a critical issuein designing cost-effective, high-performance chip multiprocessors(CMPs). Conventional memory controllers deliver relativelylow performance in part because they often employ fixed,rigid access scheduling policies designed for average-case applicationbehavior. As a result, they cannot learn and optimizethe long-term performance impact of their scheduling decisions,and cannot adapt their scheduling policies to dynamic workloadbehavior.We propose a new, self-optimizing memory controller designthat operates using the principles of reinforcement learning (RL)to overcome these limitations. Our RL-based memory controllerobserves the system state and estimates the long-term performanceimpact of each action it can take. In this way, the controllerlearns to optimize its scheduling policy on the fly to maximizelong-term performance. Our results show that an RL-basedmemory controller improves the performance of a set of parallelapplications run on a 4-core CMP by 19% on average (upto 33%), and it improves DRAM bandwidth utilization by 22%compared to a state-of-the-art controller.
Engin Ipek, Onur Mutlu, José F. Martínez, Rich Caruana
ISCA1
2008 Coordinated management of multiple interacting resources in chip multiprocessors: A machine learning approach
abstract
Efficient sharing of system resources is critical to obtaining high utilization and enforcing system-level performance objectives on chip multiprocessors (CMPs). Although several proposals that address the management of a single microarchitectural resource have been published in the literature, coordinated management of multiple interacting resources on CMPs remains an open problem.
Ramazan Bitirgen, Engin Ipek, José F. Martínez
MICRO2
2008 Efficient architectural design space exploration via predictive modeling
abstract
Efficiently exploring exponential-size architectural design spaces with many interacting parameters remains an open problem: the sheer number of experiments required renders detailed simulation intractable. We attack this via an automated approach that builds accurate predictive models. We simulate sampled points, using results to teach our models the function describing relationships among design parameters. The models can be queried and are very fast, enabling efficient design tradeoff discovery. We validate our approach via two uniprocessor sensitivity studies, predicting IPC with only 1--2% error. In an experimental study using the approach, training on 1% of a 250-K-point CMP design space allows our models to predict performance with only 4--5% error. Our predictive modeling combines well with techniques that reduce the time taken by each simulation experiment, achieving net time savings of three-four orders of magnitude.
Engin Ipek, Sally A. McKee, Rich Caruana, Bronis R. de Supinski, Martin Schulz 0001
ACM Trans. Archit. Code Optim.1
2007 Utilizing Dynamically Coupled Cores to Form a Resilient Chip Multiprocessor
abstract
Aggressive CMOS scaling will make future chip multiprocessors (CMPs) increasingly susceptible to transient faults, hard errors, manufacturing defects, and process variations. Existing fault-tolerant CMP proposals that implement dual modular redundancy (DMR) do so by statically binding pairs of adjacent cores via dedicated communication channels and buffers. This can result in unnecessary power and performance losses in cases where one core is defective (in which case the entire DMR pair must be disabled), or when cores exhibit different frequency/leakage characteristics due to process variations (in which case the pair runs at the speed of the slowest core). Static DMR also hinders power density/thermal management, as DMR pairs running code with similar power/thermal characteristics are necessarily placed next to each other on the die. We present dynamic core coupling (DCC), an architectural technique that allows arbitrary CMP cores to verify each other's execution while requiring no static core binding at design time or dedicated communication hardware. Our evaluation shows that the performance overhead of DCC over a CMP without fault tolerance is 3% on SPEC2000 benchmarks, and is within 5% for a set of scalable parallel scientific and data mining applications with up to eight threads (16 processors). Our results also show that DCC has the potential to significantly outperform existing static DMR schemes.
Christopher LaFrieda, Engin Ipek, José F. Martínez, Rajit Manohar
DSN2
2007 A Reconfigurable Chip Multiprocessor Architecture to Accommodate Software Diversity
abstract
We present core fusion, a reconfigurable chip multiprocessor (CMP) architecture where groups of fundamentally independent cores can dynamically morph into a larger CPU, or they can be used as distinct processing elements, as needed at run time by applications. Core fusion gracefully accommodates software diversity and incremental parallelization in CMPs. It provides a single execution model across all configurations, requires no additional programming effort or specialized compiler support, maintains ISA compatibility, and leverages mature micro-architecture technology.
Engin Ipek, Meyrem Kirman, Nevin Kirman, José F. Martínez
IPDPS1
2007 Core fusion: accommodating software diversity in chip multiprocessors
abstract
This paper presents core fusion, a reconfigurable chip multiprocessor(CMP) architecture where groups of fundamentally independent cores can dynamically morph into a larger CPU, or they can be used as distinct processing elements, as needed at run time by applications. Core fusion gracefully accommodates software diversity and incremental parallelization in CMPs. It provides a single execution model across all configurations, requires no additional programming effort or specialized compiler support, maintains ISA compatibility, and leverages mature micro-architecture technology.
Engin Ipek, Meyrem Kirman, Nevin Kirman, José F. Martínez
ISCA1
2007 Predicting parallel application performance via machine learning approaches
abstract
Abstract Consistently growing architectural complexity and machine scales make the creation of accurate performance models for large‐scale applications increasingly challenging. Traditional analytic models are difficult and time consuming to construct, and are often unable to capture full system and application complexity. To address these challenges, we automatically build models based on execution samples. We use multilayer neural networks, because they can represent arbitrary functions and handle noisy inputs robustly. In this paper we focus on two well‐known parallel applications whose variations in execution times are not well understood: SMG 2000, a semicoarsening multigrid solver, and HPL, an open‐source implementation of LINPACK. We sparsely sample performance data on two radically different platforms across large, multidimensional parameter spaces and show that our models based on these data can predict performance within 2% to 7% of actual application runtimes. Copyright © 2007 John Wiley & Sons, Ltd.
Engin Ipek, Sally A. McKee, Bronis R. de Supinski, Martin Schulz 0001, Rich Caruana
Concurr. Comput. Pract. Exp.2
2006 Efficiently exploring architectural design spaces via predictive modeling
abstract
Architects use cycle-by-cycle simulation to evaluate design choices and understand tradeoffs and interactions among design parameters. Efficiently exploring exponential-size design spaces with many interacting parameters remains an open problem: the sheer number of experiments renders detailed simulation intractable. We attack this problem via an automated approach that builds accurate, confident predictive design-space models. We simulate sampled points, using the results to teach our models the function describing relationships among design parameters. The models produce highly accurate performance estimates for other points in the space, can be queried to predict performance impacts of architectural changes, and are very fast compared to simulation, enabling efficient discovery of tradeoffs among parameters in different regions. We validate our approach via sensitivity studies on memory hierarchy and CPU design spaces: our models generally predict IPC with only 1-2% error and reduce required simulation by two orders of magnitude. We also show the efficacy of our technique for exploring chip multiprocessor (CMP) design spaces: when trained on a 1% sample drawn from a CMP design space with 250K points and up to 55x performance swings among different system configurations, our models predict performance with only 4-5% error on average. Our approach combines with techniques to reduce time per simulation, achieving net time savings of three-four orders of magnitude.
Engin Ipek, Sally A. McKee, Rich Caruana, Bronis R. de Supinski, Martin Schulz 0001
ASPLOS1
2006 Dynamic program phase detection in distributed shared-memory multiprocessors
abstract
We present a novel hardware mechanism for dynamic program phase detection in distributed shared-memory (DSM) multiprocessors. We show that successful hardware mechanisms for phase detection in uniprocessors do not necessarily work well in DSM systems, since they lack the ability to incorporate the parallel application's global execution information and memory access behavior based on data distribution. We then propose a hardware extension to a well-known uniprocessor mechanism that significantly improves phase detection in the context of DSM multiprocessors. The resulting mechanism is modest in size and complexity, and is transparent to the parallel application.
Engin Ipek, José F. Martínez, Bronis R. de Supinski, Sally A. McKee, Martin Schulz 0001
IPDPS1
2005 An Approach to Performance Prediction for Parallel Applications
Engin Ipek, Bronis R. de Supinski, Martin Schulz 0001, Sally A. McKee
Euro-Par1