EDBT 2026 Demo / reviewers in the wild / expert
Engin Ipek
dblp:i/EnginIpek
· DBLP profile ↗
37ranked-venue papers
8as first author
2since 2021 · last 2025
0000-0003-2809-5809ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 8 first-author · 2 since 2021Software engineering, systems software and programming languages · 13 · 4 first-author · 1 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
25 papers |
Memory systems · 36% Storage systems · 15% Hardware accelerators and domain-specific architectures · 13% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 30 heaviest of 68, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
non-volatile memory |
0.9 | 6 | 2018 | Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018 AC-DIMM: associative computing with STT-MRAM · ISCA 2013 Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computing · ISCA 2010 |
Storage systems › computational storage
in-storage computing |
0.9 | 1 | 2025 | ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025 |
Storage systems
key-value storage |
0.9 | 1 | 2025 | ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025 |
Hardware accelerators and domain-specific architectures › accelerator integration
near-storage accelerator |
0.9 | 1 | 2025 | ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025 |
Memory systems
DRAM |
0.7 | 5 | 2019 | Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019 Architecting phase change memory as a scalable dram alternative · ISCA 2009 Self-Optimizing Memory Controllers: A Reinforcement Learning Approach · ISCA 2008 |
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM |
0.6 | 3 | 2018 | Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018 AC-DIMM: associative computing with STT-MRAM · ISCA 2013 Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computing · ISCA 2010 |
Memory systems
processing-in-memory |
0.5 | 3 | 2018 | Memristive Boltzmann machine: A hardware accelerator for combinatorial optimization and deep learning · HPCA 2016 AC-DIMM: associative computing with STT-MRAM · ISCA 2013 Enabling Scientific Computing on Memristive Accelerators · ISCA 2018 |
Memory systems › in-memory computing
in-situ computing accelerator |
0.5 | 1 | 2021 | An Analog Preconditioner for Solving Linear Systems · HPCA 2021 |
High-performance computing › numerical linear algebra
preconditioner |
0.5 | 1 | 2021 | An Analog Preconditioner for Solving Linear Systems · HPCA 2021 |
High-performance computing
scientific computing |
0.5 | 1 | 2021 | An Analog Preconditioner for Solving Linear Systems · HPCA 2021 |
Memory systems › non-volatile memory
resistive memory |
0.5 | 3 | 2016 | Memristive Boltzmann machine: A hardware accelerator for combinatorial optimization and deep learning · HPCA 2016 A resistive TCAM accelerator for data-intensive computing · MICRO 2011 Dynamically replicated memory: building reliable systems from nanoscale resistive memories · ASPLOS 2010 |
Hardware reliability and fault tolerance
error correction |
0.4 | 2 | 2019 | Making Memristive Neural Network Accelerators Reliable · HPCA 2018 Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019 |
Hardware reliability and fault tolerance
memory reliability |
0.4 | 2 | 2018 | Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018 Dynamically replicated memory: building reliable systems from nanoscale resistive memories · ASPLOS 2010 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › inference accelerator
neural network inference accelerator |
0.4 | 1 | 2020 | Commutative Data Reordering: A New Technique to Reduce Data Movement Energy on Sparse Inference Workloads · ISCA 2020 |
Memory systems
memory controller |
0.4 | 3 | 2013 | A programmable memory controller for the DDRx interfacing standards · ACM Trans. Comput. Syst. 2013 PARDIS: A programmable memory controller for the DDRx interfacing standards · ISCA 2012 Self-Optimizing Memory Controllers: A Reinforcement Learning Approach · ISCA 2008 |
Energy-efficient computing › energy-efficient communication
interconnect energy |
0.4 | 2 | 2015 | More is less: improving the energy efficiency of data movement via opportunistic use of sparse codes · MICRO 2015 DESC: energy-efficient data exchange using synchronized counters · MICRO 2013 |
Energy-efficient computing
power management |
0.4 | 3 | 2017 | Voltage Regulator Efficiency Aware Power Management · ASPLOS 2017 A programmable memory controller for the DDRx interfacing standards · ACM Trans. Comput. Syst. 2013 PARDIS: A programmable memory controller for the DDRx interfacing standards · ISCA 2012 |
Memory systems › DRAM
DRAM refresh |
0.4 | 1 | 2019 | Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019 |
Storage systems › storage reliability › durability
retention failure |
0.4 | 1 | 2019 | Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data · IEEE Trans. Computers 2019 |
Hardware reliability and fault tolerance
arithmetic codes |
0.3 | 1 | 2018 | Making Memristive Neural Network Accelerators Reliable · HPCA 2018 |
Hardware reliability and fault tolerance
error-correcting codes for memory |
0.3 | 1 | 2018 | Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › in-memory computing accelerator
memristive crossbar accelerator |
0.3 | 1 | 2018 | Enabling Scientific Computing on Memristive Accelerators · ISCA 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.3 | 1 | 2018 | Making Memristive Neural Network Accelerators Reliable · HPCA 2018 |
Storage systems › storage reliability
scrubbing |
0.3 | 1 | 2018 | Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main Memories · IEEE Trans. Computers 2018 |
High-performance computing
sparse linear solver |
0.3 | 1 | 2018 | Enabling Scientific Computing on Memristive Accelerators · ISCA 2018 |
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation |
0.3 | 1 | 2018 | Enabling Scientific Computing on Memristive Accelerators · ISCA 2018 |
Memory systems
content-addressable memory |
0.3 | 2 | 2013 | AC-DIMM: associative computing with STT-MRAM · ISCA 2013 A resistive TCAM accelerator for data-intensive computing · MICRO 2011 |
Processor architecture and microarchitecture
chip multiprocessor |
0.3 | 4 | 2010 | Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computing · ISCA 2010 Coordinated management of multiple interacting resources in chip multiprocessors: A machine learning approach · MICRO 2008 Core fusion: accommodating software diversity in chip multiprocessors · ISCA 2007 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.3 | 1 | 2017 | Voltage Regulator Efficiency Aware Power Management · ASPLOS 2017 |
Memory systems › processing-in-memory
near-data processing |
0.3 | 1 | 2025 | ANVIL: An In-Storage Accelerator for Name-Value Data Stores · ISCA 2025 |
Methods — techniques the papers use, named apart from their topics
domain decomposition · 0.5bit-slicing · 0.5traveling salesman problem · 0.4capacitated vehicle routing problem · 0.4approximation method · 0.4error-correcting codes · 0.4operation scheduling · 0.3fixed-point emulation of floating point · 0.3data-aware encoding · 0.3arithmetic coding · 0.3processing-in-memory · 0.2boltzmann machine · 0.2RRAM · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ANVIL: An In-Storage Accelerator for Name-Value Data StoresabstractName-value pairs (NVPs) are a widely-used abstraction to organize data in millions of applications.At a high level, an NVP associates a name (e.g., array index, key, hash) with each value in a collection of data.Specific NVP data store formats can vary widely, ranging from simple arrays/dictionaries and lookup tables to key-value stores and data mining workloads.Despite their importance, existing optimizations for NVPs are limited to only a single data store format, as the broad definition of NVPs allows for significant heterogeneity in encoding and implementation.We propose ANVIL, the first end-to-end system that allows programmers to broadly accelerate most formats of NVPs.With a conventional solid-state drive (SSD), large-scale NVP lookups can saturate both external and internal SSD bandwidth, as every NVP in the data store needs to be sent back to the host CPU to check for a matching name.ANVIL makes use of in-storage processing to avoid reading out any data for names that do not match, by performing name match checks directly inside the SSD's NAND flash chips.We demonstrate that ANVIL can substantially reduce disk I/O, reduce metadata overheads, and provide speedups of 4.0×, 25×, and 14.6% over a conventional SSD, for three different NVP workloads (database transactions, analytics, and graph processing). Ryan Wong 0001, Nikita Kim, Aniket Das, Kevin Higgs, Engin Ipek, Sapan Agarwal, Saugata Ghose, Ben Feinberg |
ISCA | 5 |
| 2021 | An Analog Preconditioner for Solving Linear SystemsabstractOver the past decade as Moore's Law has slowed, the need for new forms of computation that can provide sustainable performance improvements has risen. A new method, called in situ computing, has shown great potential to accelerate matrix vector multiplication (MVM), an important kernel for a diverse range of applications from neural networks to scientific computing. Existing in situ accelerators for scientific computing, however, have a significant limitation: these accelerators provide no acceleration for preconditioning-a key bottleneck in linear solvers and in scientific computing workflows. This paper enables in situ acceleration for state-of-the-art linear solvers by demonstrating how to use a new in situ matrix inversion accelerator for analog preconditioning. As existing techniques that enable high precision and scalability for in situ MVM are inapplicable to in situ matrix inversion, new techniques to compensate for circuit non-idealities are proposed. Additionally, a new approach to bit slicing that enables splitting operands across multiple devices without external digital logic is proposed. For scalability, this paper demonstrates how in situ matrix inversion kernels can work in tandem with existing domain decomposition techniques to accelerate the solutions of arbitrarily large linear systems. The analog kernel can be directly integrated into existing preconditioning workflows, leveraging several well-optimized numerical linear algebra tools to improve the behavior of the circuit. The result is an analog preconditioner that is more effective (up to 50% fewer iterations) than the widely used incomplete LU factorization preconditioner, ILU(0), while also reducing the energy and execution time of each approximate solve operation by 1025x and 105x respectively. Ben Feinberg, Ryan Wong 0001, T. Patrick Xiao, Christopher H. Bennett, Jacob N. Rohan, Erik G. Boman, Matthew J. Marinella, Sapan Agarwal, Engin Ipek |
HPCA | 9 |
| 2020 | Commutative Data Reordering: A New Technique to Reduce Data Movement Energy on Sparse Inference WorkloadsabstractData movement is a significant and growing consumer of energy in modern systems, from specialized low-power accelerators to GPUs with power budgets in the hundreds of Watts. Given the importance of the problem, prior work has proposed designing interconnects on which the energy cost of transmitting a 0 is significantly lower than that of transmitting a 1. With such an interconnect, data movement energy is reduced by encoding the transmitted data such that the number of 1s is minimized. Although promising, these data encoding proposals do not take full advantage of application level semantics. As an example of a neglected optimization opportunity, consider the case of a dot product computation as part of a neural network inference task. The order in which the neural network weights are fetched and processed does not affect correctness, and can be optimized to further reduce data movement energy.This paper presents commutative data reordering (CDR), a hardware-software approach that leverages the commutative property in linear algebra to strategically select the order in which weight matrix coefficients are fetched from memory. To find a low-energy transmission order, weight ordering is modeled as an instance of one of two well-studied problems, the Traveling Salesman Problem and the Capacitated Vehicle Routing Problem. This reduction makes it possible to leverage the vast body of work on efficient approximation methods to find a good transmission order. CDR exploits the indirection inherent to sparse matrix formats such that no additional metadata is required to specify the selected order. The hardware modifications required to support CDR are minimal, and incur an area penalty of less than 0.01% when implemented on top of a mobile-class GPU. When applied to 7 neural network inference tasks running on a GPU-based system, CDR respectively reduces average DRAM IO energy by 53.1% and 22.2% over the data bus invert encoding scheme used by LPDDR4, and the recently proposed Base + XOR encoding. These savings are attained with no changes to the mobile system software and no runtime performance penalty. Ben Feinberg, Benjamin C. Heyman, Darya Mikhailenko, Ryan Wong 0001, An C. Ho, Engin Ipek |
ISCA | 6 |
| 2019 | Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable DataabstractDRAM refresh is responsible for significant performance and energy overheads in a wide range of computer systems, from mobile platforms to datacenters [1] . With the growing demand for DRAM capacity and the worsening retention time characteristics of deeply scaled DRAM, refresh is expected to become an even more pronounced problem in future technology generations [2] . This paper examines content aware refresh, a new technique that reduces the refresh frequency by exploiting the unidirectional nature of DRAM retention errors: assuming that a logical 1 and 0 respectively are represented by the presence and absence of charge, 1-to-0 failures are much more likely than 0-to-1 failures. As a result, in a DRAM system that uses a block error correcting code (ECC) to protect memory, blocks with fewer 1s can attain a specified reliability target (i.e., mean time to failure) with a refresh rate lower than that which is required for a block with all 1s. Leveraging this key insight, and without compromising memory reliability, the proposed content aware refresh mechanism refreshes memory blocks with fewer 1s less frequently. To keep the overhead of tracking multiple refresh rates manageable, refresh groups-groups of DRAM rows refreshed together-are dynamically arranged into one of a predefined number of refresh bins and refreshed at the rate determined by the ECC block with the greatest number of 1s in that bin. By tailoring the refresh rate to the actual content of a memory block rather than assuming a worst case data pattern, content aware refresh respectively outperforms DRAM systems that employ RAS-only Refresh, all-bank Auto Refresh, and per-bank Auto Refresh mechanisms by 12, 8, and 13 percent. It also reduces DRAM system energy by 15, 13, and 16 percent as compared to these systems. Mahdi Nazm Bojnordi, Engin Ipek |
IEEE Trans. Computers | 4 |
| 2018 | Making Memristive Neural Network Accelerators ReliableabstractDeep neural networks (DNNs) have attracted substantial interest in recent years due to their superior performance on many classification and regression tasks as compared to other supervised learning models. DNNs often require a large amount of data movement, resulting in performance and energy overheads. One promising way to address this problem is to design an accelerator based on in-situ analog computing that leverages the fundamental electrical properties of memristive circuits to perform matrix-vector multiplication. Recent work on analog neural network accelerators has shown great potential in improving both the system performance and the energy efficiency. However, detecting and correcting the errors that occur during in-memory analog computation remains largely unexplored. The same electrical properties that provide the performance and energy improvements make these systems especially susceptible to errors, which can severely hurt the accuracy of the neural network accelerators. This paper examines a new error correction scheme for analog neural network accelerators based on arithmetic codes. The proposed scheme encodes the data through multiplication by an integer, which preserves addition operations through the distributive property. Error detection and correction are performed through a modulus operation and a correction table lookup. This basic scheme is further improved by data-aware encoding to exploit the state dependence of the errors, and by knowledge of how critical each portion of the computation is to overall system accuracy. By leveraging the observation that a physical row that contains fewer 1s is less susceptible to an error, the proposed scheme increases the effective error correction capability with less than 4.5% area and less than 4.7% energy overheads. When applied to a memristive DNN accelerator performing inference on the MNIST and ILSVRC-2012 datasets, the proposed technique reduces the respective misclassification rates by 1.5x and 1.1x. Ben Feinberg, Engin Ipek |
HPCA | 3 |
| 2018 | Enabling Scientific Computing on Memristive AcceleratorsabstractLinear algebra is ubiquitous across virtually every field of science and engineering, from climate modeling to macroeconomics. This ubiquity makes linear algebra a prime candidate for hardware acceleration, which can improve both the run time and the energy efficiency of a wide range of scientific applications. Recent work on memristive hardware accelerators shows significant potential to speed up matrix-vector multiplication (MVM), a critical linear algebra kernel at the heart of neural network inference tasks. Regrettably, the proposed hardware is constrained to a narrow range of workloads: although the eight-to 16-bit computations afforded by memristive MVM accelerators are acceptable for machine learning, they are insufficient for scientific computing where high-precision floating point is the norm. This paper presents the first proposal to enable scientific computing on memristive crossbars. Three techniques are explored — reducing overheads by exploiting exponent range locality, early termination of fixed-point computation, and static operation scheduling — that together enable a fixed-point memristive accelerator to perform high-precision floating point without the exorbitant cost of naïve floating-point emulation on fixed-point hardware. A heterogeneous collection of crossbars with varying sizes is proposed to efficiently handle sparse matrices, and an algorithm for mapping the dense subblocks of a sparse matrix to an appropriate set of crossbars is investigated. The accelerator can be combined with existing GPU-based systems to handle datasets that cannot be efficiently handled by the memristive accelerator alone. The proposed optimizations permit the memristive MVM concept to be applied to a wide range of problem domains, respectively improving the execution time and energy dissipation of sparse linear solvers by 10.3x and 10.9x over a purely GPU-based system. Ben Feinberg, Uday Kumar Reddy Vengalam, Nathan Whitehair, Engin Ipek |
ISCA | 5 |
| 2018 | Sanitizer: Mitigating the Impact of Expensive ECC Checks on STT-MRAM Based Main MemoriesabstractDRAM density scaling has become increasingly difficult due to challenges in maintaining a sufficiently high storage capacitance and a sufficiently low leakage current at nanoscale feature sizes. Non-volatile memories (NVMs) have drawn significant attention as potential DRAM replacements because they represent information using resistance rather than electrical charge. Spin-torque transfer magnetoresistive RAM (STT-MRAM) is one of the most promising NVM technologies due to its relatively low write energy, high speed, and high endurance. However, STT-MRAM suffers from its own scaling problems. As the size of the storage element decreases with technology scaling, STT-MRAM retention error rates are expected to increase, which will require multi-bit error-correcting code (ECC) and periodic scrubbing. We introduce the Sanitizer architecture, which mitigates the performance and energy overheads of ECC and scrubbing in future STT-MRAM based main memories. To reduce the scrubbing rate, a coarse-grained, multi-bit ECC mechanism with a 12.5 percent storage overhead is used. To avoid fetching multiple blocks from memory and performing costly ECC checks on every read, the memory regions that will likely be accessed in the near future are predicted and proactively scrubbed. Compared to a conventional STT-MRAM system, Sanitizer improves performance by 1.22× and reduces end-to-end system energy by 22 percent. Mahdi Nazm Bojnordi, Qing Guo 0004, Engin Ipek |
IEEE Trans. Computers | 4 |
| 2017 | Voltage Regulator Efficiency Aware Power ManagementabstractConventional off-chip voltage regulators are typically bulky and slow, and are inefficient at exploiting system and workload variability using Dynamic Voltage and Frequency Scaling (DVFS). On-die integration of voltage regulators has the potential to increase the energy efficiency of computer systems by enabling power control at a fine granularity in both space and time. The energy conversion efficiency of on-chip regulators, however, is typically much lower than off-chip regulators, which results in significant energy losses. Fine-grained power control and high voltage regulator efficiency are difficult to achieve simultaneously, with either emerging on-chip or conventional off-chip regulators. Victor W. Lee, Engin Ipek |
ASPLOS | 3 |
| 2016 | Memristive Boltzmann machine: A hardware accelerator for combinatorial optimization and deep learningabstractThe Boltzmann machine is a massively parallel computational model capable of solving a broad class of combinatorial optimization problems. In recent years, it has been successfully applied to training deep machine learning models on massive datasets. High performance implementations of the Boltzmann machine using GPUs, MPI-based HPC clusters, and FPGAs have been proposed in the literature. Regrettably, the required all-to-all communication among the processing units limits the performance of these efforts. This paper examines a new class of hardware accelerators for large-scale combinatorial optimization and deep learning based on memristive Boltzmann machines. A massively parallel, memory-centric hardware accelerator is proposed based on recently developed resistive RAM (RRAM) technology. The proposed accelerator exploits the electrical properties of RRAm to realize in situ, fine-grained parallel computation within memory arrays, thereby eliminating the need for exchanging data between the memory cells and the computational units. Two classical optimization problems, graph partitioning and boolean satisfiability, and a deep belief network application are mapped onto the proposed hardware. As compared to a multicore system, the proposed accelerator achieves 57x higher performance and 25x lower energy with virtually no loss in the quality of the solution to the optimization problems. The memristive accelerator is also compared against an RRAM based processing-in-memory (PIM) system, with respective performance and energy improvements of 6.89x and 5.2x. Mahdi Nazm Bojnordi, Engin Ipek |
HPCA | 2 |
| 2016 | Reducing data movement energy via online data clustering and encodingabstractModern computer systems expend significant amounts of energy on transmitting data over long and highly capacitive interconnects. A promising way of reducing the data movement energy is to design the interconnect such that the transmission of 0s is considerably cheaper than that of 1s. Given such an interconnect with asymmetric transmission costs, data movement energy can be reduced by encoding the transmitted data such that the number of 1s in each transmitted codeword is minimized. This paper presents a new data encoding technique based on online data clustering that exploits this opportunity. The transmitted data blocks are dynamically clustered based on the similarities between their binary representations. Each data block is expressed as the bitwise XOR between one of multiple cluster centers and a residual with a small number of 1s. The data movement energy is minimized by sending the residual along with an identifier that specifies which cluster center to use in decoding the transmitted data. At runtime, the proposed approach continually updates the cluster centers based on the observed data to adapt to phase changes. The proposed technique is compared to three previously proposed energy-efficient data encoding techniques on a set of 14 applications. The results indicate respective energy savings of 5%, 9%, and 12% in DDR4, LPDDR3, and last level cache subsystems as compared to the best existing baseline encoding technique. Engin Ipek |
MICRO | 2 |
| 2016 | Back to the Future: Current-Mode Processor in the Era of Deeply Scaled CMOSabstractThis paper explores the use of MOS current-mode logic (MCML) as a fast and low noise alternative to static CMOS circuits in microprocessors, thereby improving the performance, energy efficiency, and signal integrity of future computer systems. The power and ground noise generated by an MCML circuit is typically 10-100× smaller than the noise generated by a static CMOS circuit. Unlike static CMOS, whose dominant dynamic power is proportional to the frequency, MCML circuits dissipate a constant power independent of clock frequency. Although these traits make MCML highly energy efficient when operating at high speeds, the constant static power of MCML poses a challenge for a microarchitecture that operates at the modest clock rate and with a low activity factor. To address this challenge, a single-core microarchitecture for MCML is explored that exploits the C-slow retiming technique, and operates at a high frequency with low complexity to save energy. This design principle contrasts with the contemporary multicore design paradigm for static CMOS that relies on a large number of gates operating in parallel at the modest speeds. The proposed architecture generates 10-40× lower power and ground noise, and operates within 13% of the performance (i.e., 1/ExecutionTime) of a conventional, eight-core static CMOS processor while exhibiting 1.6× lower energy and 9% less area. Moreover, the operation of an MCML processor is robust under both systematic and random variations in transistor threshold voltage and effective channel length. Yanwei Song, Mahdi Nazm Bojnordi, Alexander E. Shapiro, Eby G. Friedman, Engin Ipek |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2016 | Reducing Switching Latency and Energy in STT-MRAM Caches With Field-Assisted WritingabstractA field-assisted spin-torque transfer magnetoresistive RAM (STT-MRAM) cache is presented for the use in high-performance energy-efficient microprocessors. Adding field assistance reduces the switching latency by a factor of 4. An array model is developed to evaluate the switching energy for different field currents and array sizes. Several STT-MRAM-based cells demonstrate a 55% energy reduction as compared with an SRAM cache subsystem. As compared with STT-MRAM caches with subbank buffering and differential writes, a field-assisted STT-MRAM cache improves the system performance by 28%, with a 6.7% increase in energy. Ravi Patel 0001, Qing Guo 0004, Engin Ipek, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Architecting a MOS current mode logic (MCML) processor for fast, low noise and energy-efficient computing in the near-threshold regimeabstractNear-threshold computing (NTC) is an effective technique for improving the energy efficiency of a CMOS microprocessor, but suffers from a significant performance loss and an increased sensitivity to voltage noise. MOS current-mode logic (MCML), a differential logic family, maintains a low voltage swing and a constant current, making it inherently fast and low-noise. These traits make MCML a natural selection to implement an NTC processor; however, MCML suffers from a high static power regardless of the clock frequency or the level of switching activity, which would result in an inordinate energy consumption in a large scale IC. To address this challenge, this paper explores a single-core microarchitecture for MCML that takes advantage of C-slow retiming technique, and runs at a high frequency with low complexity to save energy. This design principle is opposite to the contemporary multicore design paradigm for static CMOS that relies on a large number of gates running in parallel at modest speeds. When compared to an eight-core static CMOS processor operating in the near-threshold regime, the proposed processor exhibits 3x higher performance, 2x lower energy, and 10 x lower voltage noise, while maintaining a similar level of power dissipation. Yanwei Song, Mahdi Nazm Bojnordi, Alexander E. Shapiro, Engin Ipek, Eby G. Friedman |
ICCD | 5 |
| 2015 | Energy-efficient data movement with sparse transition encodingabstractData movement over long on-chip interconnects is a major contributor to system energy. This paper presents novel signaling and encoding techniques that toget her improve the energy efficiency of data communication between the processor cores and the last level cache. The proposed techniques make the interconnect energy proportional to the number of ones in the transferred data block (i.e., the block's hamming weight), regardless of the previous state of the interconnect. The hamming weight of each cache block is kept low through a sparse data encoding approach to minimize interconnect energy. Simulation results show that the proposed communication scheme reduces the overall L2 cache energy by 30% on a set of eleven parallel applications with a 0.5% average performance degradation. Yanwei Song, Mahdi Nazm Bojnordi, Engin Ipek |
ICCD | 3 |
| 2015 | Enabling energy efficient Hybrid Memory Cube systems with erasure codesabstractThe Hybrid Memory Cube (HMC) is a promising alternative to DDRx memory due to its potential to achieve significantly higher bandwidth. However, the high static power of an HMC device compromises power efficiency when the device is lightly utilized. Activating a sleeping HMC takes over 2µs, which makes it challenging to manage HMC power without a substantial degradation in system performance. We introduce a new technique that alleviates the long wakeup penalty of an HMC by employing erasure codes. Inaccessible data stored in a sleeping HMC module can be reconstructed by decoding related data retrieved from other active HMCs, rather than waiting for the sleeping HMC module to become active. This approach makes it possible to tolerate the latency penalty incurred when switching an HMC between active and sleep modes, thereby enabling a power-capped HMC system. Simulations show that the proposed architecture outperforms a current HMC-based multicore system by 6.2×, and reduces the system energy by 5.3× under the same power budget as the multicore baseline. Yanwei Song, Mahdi Nazm Bojnordi, Engin Ipek |
ISLPED | 4 |
| 2015 | More is less: improving the energy efficiency of data movement via opportunistic use of sparse codesabstractData movement over long and highly capacitive interconnects is responsible for a large fraction of the energy consumed in nanometer ICs. DDRx, the most broadly adopted family of DRAM interfaces, contributes significantly to the overall system energy in a wide range of computer systems. To reduce the energy cost of data transfers, DDR4 adopts a pseudo open-drain IO circuit that consumes power only when transmitting or receiving a 0, which makes the IO energy proportional to the number of 0s transferred over the data bus. A data bus invert (DBI) coding technique is therefore supported by the DDR4 standard to encode each byte using a small number of 0s. Although sparse coding techniques that are more advanced than DBI can reduce the IO power further, the relatively high bandwidth overhead of these codes has heretofore prevented their application to the DDRx bus. Yanwei Song, Engin Ipek |
MICRO | 2 |
| 2014 | Field driven STT-MRAM cell for reduced switching latency and energyabstractA field driven approach to STT-MRAM switching is proposed as a method for reducing the switching latency of an MTJ in high performance caches. An MRAM array model is presented to characterize the switching energy and maximum achievable reduction in energy using the field driven approach. The switching latency per bit is reduced by more than a factor of ten. The resultant switching energy per bit is reduced by 82% as compared to a standard STT-MRAM. Ravi Patel 0001, Engin Ipek, Eby G. Friedman |
ISCAS | 2 |
| 2013 | AC-DIMM: associative computing with STT-MRAMabstractWith technology scaling, on-chip power dissipation and off-chip memory bandwidth have become significant performance bottlenecks in virtually all computer systems, from mobile devices to supercomputers. An effective way of improving performance in the face of bandwidth and power limitations is to rely on associative memory systems. Recent work on a PCM-based, associative TCAM accelerator shows that associative search capability can reduce both off-chip bandwidth demand and overall system energy. Unfortunately, previously proposed resistive TCAM accelerators have limited flexibility: only a restricted (albeit important) class of applications can benefit from a TCAM accelerator, and the implementation is confined to resistive memory technologies with a high dynamic range (RHigh/RLow), such as PCM. Qing Guo 0004, Ravi Patel 0001, Engin Ipek, Eby G. Friedman |
ISCA | 4 |
| 2013 | DESC: energy-efficient data exchange using synchronized countersabstractIncreasing cache sizes in modern microprocessors require long wires to connect cache arrays to processor cores. As a result, the last-level cache (LLC) has become a major contributor to processor energy, necessitating techniques to increase the energy efficiency of data exchange over LLC interconnects. Mahdi Nazm Bojnordi, Engin Ipek |
MICRO | 2 |
| 2013 | A programmable memory controller for the DDRx interfacing standardsabstractModern memory controllers employ sophisticated address mapping, command scheduling, and power management optimizations to alleviate the adverse effects of DRAM timing and resource constraints on system performance. A promising way of improving the versatility and efficiency of these controllers is to make them programmable—a proven technique that has seen wide use in other control tasks, ranging from DMA scheduling to NAND Flash and directory control. Unfortunately, the stringent latency and throughput requirements of modern DDRx devices have rendered such programmability largely impractical, confining DDRx controllers to fixed-function hardware. This article presents the instruction set architecture (ISA) and hardware implementation of PARDIS, a programmable memory controller that can meet the performance requirements of a high-speed DDRx interface. The proposed controller is evaluated by mapping previously proposed DRAM scheduling, address mapping, refresh scheduling, and power management algorithms onto PARDIS. Simulation results show that the average performance of PARDIS comes within 8% of fixed-function hardware for each of these techniques; moreover, by enabling application-specific optimizations, PARDIS improves system performance by 6 to 17% and reduces DRAM energy by 9 to 22% over four existing memory controllers. Mahdi Nazm Bojnordi, Engin Ipek |
ACM Trans. Comput. Syst. | 2 |
| 2012 | Overcoming single-thread performance hurdles in the core fusion reconfigurable multicore architectureabstractThough the prime target of multicore architectures is parallel and multithreaded workloads (which favors maximum core count), executing sequential code fast continues to remain critical (which benefits from maximum core size). This poses a difficult design trade-off. Core Fusion is a recently-proposed reconfigurable multicore architecture that attempts to circumvent this compromise by "fusing" groups of fundamentally independent cores into larger, more aggressive processors dynamically as needed. In this way, it accommodates highly parallel, partially parallel, multiprogrammed, and sequential codes with ease. Janani Mukundan, Saugata Ghose, Robert Karmazin, Engin Ipek, José F. Martínez |
ICS | 4 |
| 2012 | PARDIS: A programmable memory controller for the DDRx interfacing standardsabstractModern memory controllers employ sophisticated address mapping, command scheduling, and power management optimizations to alleviate the adverse effects of DRAM timing and resource constraints on system performance. A promising way of improving the versatility and efficiency of these controllers is to make them programmable - a proven technique that has seen wide use in other control tasks ranging from DMA scheduling to NAND Flash and directory control. Unfortunately, the stringent latency and throughput requirements of modern DDRx devices have rendered such programmability largely impractical, confining DDRx controllers to fixed-function hardware. This paper presents the instruction set architecture (ISA) and hardware implementation of PARDIS, a programmable memory controller that can meet the performance requirements of a high-speed DDRx interface. The proposed controller is evaluated by mapping previously proposed DRAM scheduling, address mapping, refresh scheduling, and power management algorithms onto PARDIS. Simulation results show that the average performance of PARDIS comes within 8% of fixed-function hardware for each of these techniques; moreover, by enabling application-specific optimizations, PARDIS improves system performance by 6-17% and reduces DRAM energy by 9-22% over four existing memory controllers. Mahdi Nazm Bojnordi, Engin Ipek |
ISCA | 2 |
| 2011 | A resistive TCAM accelerator for data-intensive computingabstractPower dissipation and off-chip bandwidth restrictions are critical challenges that limit microprocessor performance. Ternary content addressable memories (TCAM) hold the potential to address both problems in the context of a wide range of data-intensive workloads that benefit from associative search capability. Power dissipation is reduced by eliminating instruction processing and data movement overheads present in a purely RAM based system. Bandwidth demand is lowered by processing data directly on the TCAM chip, thereby decreasing off-chip traffic. Unfortunately, CMOS-based TCAM implementations are severely power- and area-limited, which restricts the capacity of commercial products to a few megabytes, and confines their use to niche networking applications. Qing Guo 0004, Engin Ipek |
MICRO | 4 |
| 2010 | Dynamically replicated memory: building reliable systems from nanoscale resistive memoriesabstractDRAM is facing severe scalability challenges in sub-45nm tech- nology nodes due to precise charge placement and sensing hur- dles in deep-submicron geometries. Resistive memories, such as phase-change memory (PCM), already scale well beyond DRAM and are a promising DRAM replacement. Unfortunately, PCM is write-limited, and current approaches to managing writes must de- commission pages of PCM when the first bit fails. Engin Ipek, Jeremy Condit, Ed Nightingale, Doug Burger, Thomas Moscibroda |
ASPLOS | 1 |
| 2010 | Resistive computation: avoiding the power wall with low-leakage, STT-MRAM based computingabstractAs CMOS scales beyond the 45nm technology node, leakage concerns are starting to limit microprocessor performance growth. To keep dynamic power constant across process generations, traditional MOSFET scaling theory prescribes reducing supply and threshold voltages in proportion to device dimensions, a practice that induces an exponential increase in subthreshold leakage. As a result, leakage power has become comparable to dynamic power in current-generation processes, and will soon exceed it in magnitude if voltages are scaled down any further. Beyond this inflection point, multicore processors will not be able to afford keeping more than a small fraction of all cores active at any given moment. Multicore scaling will soon hit a power wall. This paper presents resistive computation, a new technique that aims at avoiding the power wall by migrating most of the functionality of a modern microprocessor from CMOS to spin-torque transfer magnetoresistive RAM (STT-MRAM)---a CMOS-compatible, leakage-resistant, non-volatile resistive memory technology. By implementing much of the on-chip storage and combinational logic using leakage-resistant, scalable RAM blocks and lookup tables, and by carefully re-architecting the pipeline, an STT-MRAM based implementation of an eight-core Sun Niagara-like CMT processor reduces chip-wide power dissipation by 1.7× and leakage power by 2.1× at the 32nm technology node, while maintaining 93% of the system throughput of a CMOS-based design. Engin Ipek, Tolga Soyata |
ISCA | 2 |
| 2009 | Architecting phase change memory as a scalable dram alternativeabstractMemory scaling is in jeopardy as charge storage and sensing mechanisms become less reliable for prevalent memory technologies, such as DRAM. In contrast, phase change memory (PCM) storage relies on scalable current and thermal mechanisms. To exploit PCM's scalability as a DRAM alternative, PCM must be architected to address relatively long latencies, high energy writes, and finite endurance.We propose, crafted from a fundamental understanding of PCM technology parameters, area-neutral architectural enhancements that address these limitations and make PCM competitive with DRAM. A baseline PCM system is 1.6x slower and requires 2.2x more energy than a DRAM system. Buffer reorganizations reduce this delay and energy gap to 1.2x and 1.0x, using narrow rows to mitigate write energy and multiple rows to improve locality and write coalescing. Partial writes enhance memory endurance, providing 5.6 years of lifetime. Process scaling will further reduce PCM energy costs and improve endurance. Benjamin C. Lee, Engin Ipek, Onur Mutlu, Doug Burger |
ISCA | 2 |
| 2009 | Better I/O through byte-addressable, persistent memoryabstractModern computer systems have been built around the assumption that persistent storage is accessed via a slow, block-based interface. However, new byte-addressable, persistent memory technologies such as phase change memory (PCM) offer fast, fine-grained access to persistent storage. Jeremy Condit, Ed Nightingale, Christopher Frost 0001, Engin Ipek, Benjamin C. Lee, Doug Burger, Derrick Coetzee |
SOSP | 4 |
| 2008 | Self-Optimizing Memory Controllers: A Reinforcement Learning ApproachabstractEfficiently utilizing off-chip DRAM bandwidth is a critical issuein designing cost-effective, high-performance chip multiprocessors(CMPs). Conventional memory controllers deliver relativelylow performance in part because they often employ fixed,rigid access scheduling policies designed for average-case applicationbehavior. As a result, they cannot learn and optimizethe long-term performance impact of their scheduling decisions,and cannot adapt their scheduling policies to dynamic workloadbehavior.We propose a new, self-optimizing memory controller designthat operates using the principles of reinforcement learning (RL)to overcome these limitations. Our RL-based memory controllerobserves the system state and estimates the long-term performanceimpact of each action it can take. In this way, the controllerlearns to optimize its scheduling policy on the fly to maximizelong-term performance. Our results show that an RL-basedmemory controller improves the performance of a set of parallelapplications run on a 4-core CMP by 19% on average (upto 33%), and it improves DRAM bandwidth utilization by 22%compared to a state-of-the-art controller. Engin Ipek, Onur Mutlu, José F. Martínez, Rich Caruana |
ISCA | 1 |
| 2008 | Coordinated management of multiple interacting resources in chip multiprocessors: A machine learning approachabstractEfficient sharing of system resources is critical to obtaining high utilization and enforcing system-level performance objectives on chip multiprocessors (CMPs). Although several proposals that address the management of a single microarchitectural resource have been published in the literature, coordinated management of multiple interacting resources on CMPs remains an open problem. Ramazan Bitirgen, Engin Ipek, José F. Martínez |
MICRO | 2 |
| 2008 | Efficient architectural design space exploration via predictive modelingabstractEfficiently exploring exponential-size architectural design spaces with many interacting parameters remains an open problem: the sheer number of experiments required renders detailed simulation intractable. We attack this via an automated approach that builds accurate predictive models. We simulate sampled points, using results to teach our models the function describing relationships among design parameters. The models can be queried and are very fast, enabling efficient design tradeoff discovery. We validate our approach via two uniprocessor sensitivity studies, predicting IPC with only 1--2% error. In an experimental study using the approach, training on 1% of a 250-K-point CMP design space allows our models to predict performance with only 4--5% error. Our predictive modeling combines well with techniques that reduce the time taken by each simulation experiment, achieving net time savings of three-four orders of magnitude. Engin Ipek, Sally A. McKee, Rich Caruana, Bronis R. de Supinski, Martin Schulz 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2007 | Utilizing Dynamically Coupled Cores to Form a Resilient Chip MultiprocessorabstractAggressive CMOS scaling will make future chip multiprocessors (CMPs) increasingly susceptible to transient faults, hard errors, manufacturing defects, and process variations. Existing fault-tolerant CMP proposals that implement dual modular redundancy (DMR) do so by statically binding pairs of adjacent cores via dedicated communication channels and buffers. This can result in unnecessary power and performance losses in cases where one core is defective (in which case the entire DMR pair must be disabled), or when cores exhibit different frequency/leakage characteristics due to process variations (in which case the pair runs at the speed of the slowest core). Static DMR also hinders power density/thermal management, as DMR pairs running code with similar power/thermal characteristics are necessarily placed next to each other on the die. We present dynamic core coupling (DCC), an architectural technique that allows arbitrary CMP cores to verify each other's execution while requiring no static core binding at design time or dedicated communication hardware. Our evaluation shows that the performance overhead of DCC over a CMP without fault tolerance is 3% on SPEC2000 benchmarks, and is within 5% for a set of scalable parallel scientific and data mining applications with up to eight threads (16 processors). Our results also show that DCC has the potential to significantly outperform existing static DMR schemes. Christopher LaFrieda, Engin Ipek, José F. Martínez, Rajit Manohar |
DSN | 2 |
| 2007 | A Reconfigurable Chip Multiprocessor Architecture to Accommodate Software DiversityabstractWe present core fusion, a reconfigurable chip multiprocessor (CMP) architecture where groups of fundamentally independent cores can dynamically morph into a larger CPU, or they can be used as distinct processing elements, as needed at run time by applications. Core fusion gracefully accommodates software diversity and incremental parallelization in CMPs. It provides a single execution model across all configurations, requires no additional programming effort or specialized compiler support, maintains ISA compatibility, and leverages mature micro-architecture technology. Engin Ipek, Meyrem Kirman, Nevin Kirman, José F. Martínez |
IPDPS | 1 |
| 2007 | Core fusion: accommodating software diversity in chip multiprocessorsabstractThis paper presents core fusion, a reconfigurable chip multiprocessor(CMP) architecture where groups of fundamentally independent cores can dynamically morph into a larger CPU, or they can be used as distinct processing elements, as needed at run time by applications. Core fusion gracefully accommodates software diversity and incremental parallelization in CMPs. It provides a single execution model across all configurations, requires no additional programming effort or specialized compiler support, maintains ISA compatibility, and leverages mature micro-architecture technology. Engin Ipek, Meyrem Kirman, Nevin Kirman, José F. Martínez |
ISCA | 1 |
| 2007 | Predicting parallel application performance via machine learning approachesabstractAbstract Consistently growing architectural complexity and machine scales make the creation of accurate performance models for large‐scale applications increasingly challenging. Traditional analytic models are difficult and time consuming to construct, and are often unable to capture full system and application complexity. To address these challenges, we automatically build models based on execution samples. We use multilayer neural networks, because they can represent arbitrary functions and handle noisy inputs robustly. In this paper we focus on two well‐known parallel applications whose variations in execution times are not well understood: SMG 2000, a semicoarsening multigrid solver, and HPL, an open‐source implementation of LINPACK. We sparsely sample performance data on two radically different platforms across large, multidimensional parameter spaces and show that our models based on these data can predict performance within 2% to 7% of actual application runtimes. Copyright © 2007 John Wiley & Sons, Ltd. Engin Ipek, Sally A. McKee, Bronis R. de Supinski, Martin Schulz 0001, Rich Caruana |
Concurr. Comput. Pract. Exp. | 2 |
| 2006 | Efficiently exploring architectural design spaces via predictive modelingabstractArchitects use cycle-by-cycle simulation to evaluate design choices and understand tradeoffs and interactions among design parameters. Efficiently exploring exponential-size design spaces with many interacting parameters remains an open problem: the sheer number of experiments renders detailed simulation intractable. We attack this problem via an automated approach that builds accurate, confident predictive design-space models. We simulate sampled points, using the results to teach our models the function describing relationships among design parameters. The models produce highly accurate performance estimates for other points in the space, can be queried to predict performance impacts of architectural changes, and are very fast compared to simulation, enabling efficient discovery of tradeoffs among parameters in different regions. We validate our approach via sensitivity studies on memory hierarchy and CPU design spaces: our models generally predict IPC with only 1-2% error and reduce required simulation by two orders of magnitude. We also show the efficacy of our technique for exploring chip multiprocessor (CMP) design spaces: when trained on a 1% sample drawn from a CMP design space with 250K points and up to 55x performance swings among different system configurations, our models predict performance with only 4-5% error on average. Our approach combines with techniques to reduce time per simulation, achieving net time savings of three-four orders of magnitude. Engin Ipek, Sally A. McKee, Rich Caruana, Bronis R. de Supinski, Martin Schulz 0001 |
ASPLOS | 1 |
| 2006 | Dynamic program phase detection in distributed shared-memory multiprocessorsabstractWe present a novel hardware mechanism for dynamic program phase detection in distributed shared-memory (DSM) multiprocessors. We show that successful hardware mechanisms for phase detection in uniprocessors do not necessarily work well in DSM systems, since they lack the ability to incorporate the parallel application's global execution information and memory access behavior based on data distribution. We then propose a hardware extension to a well-known uniprocessor mechanism that significantly improves phase detection in the context of DSM multiprocessors. The resulting mechanism is modest in size and complexity, and is transparent to the parallel application. Engin Ipek, José F. Martínez, Bronis R. de Supinski, Sally A. McKee, Martin Schulz 0001 |
IPDPS | 1 |
| 2005 | An Approach to Performance Prediction for Parallel Applications
Engin Ipek, Bronis R. de Supinski, Martin Schulz 0001, Sally A. McKee |
Euro-Par | 1 |