VLDB 2026 Research / reviewers in the wild / expert
Ulya R. Karpuzcu
dblp:89/5932
· DBLP profile ↗
38ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-9238-4256ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 7 · 4 since 2021Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid Ising FPGA-COBI Architecture with Hardware-Based Problem DecompositionabstractMany combinatorial optimization problems map naturally to Ising Hamiltonians, $H({\text{s}}) = - \sum\nolimits_{i,j} {{J_{ij}}} {s_i}{s_j} - \sum\nolimits_i {{h_i}} {s_i}$ , and CMOS ring-oscillator Ising machines solve them in microseconds at milliwatts [1] , [2] . Their key limitation is capacity : the number of spins one solver core can process in a single solve. Because each hardware spin represents one binary Ising variable, capacity directly sets the largest problem solvable in one shot. Our 28 nm five-core COBI chip solves a 45-spin all-to-all subproblem per core in 77.5 µ s, so larger instances require iterative decomposition. This shifts the bottleneck from analog solving to digital orchestration: a CPU-based decomposer needs ∼321 µ s/iter over PCIe, 4× the core solve time, leaving the solver idle 84.9% of the time. We instead co-locate an FPGA decomposer with the chip and derive sizing laws for the required parallelism, achieving 1.93× geomean speedup and > 40× energy reduction vs. an optimized C++ baseline. Ruihong Yin, Chaohui Li, Ahmet Efe, Abhimanyu Kumar, Ziqing Zeng, Ulya R. Karpuzcu, Sachin S. Sapatnekar, Chris H. Kim |
FCCM | 7 |
| 2026 | MIBID: Model Based Fault Diagnosis on Ising MachinesabstractModel-Based Diagnosis (MBD) identifies faulty components in complex systems by reasoning over a model of expected behavior and observations. Computing minimal-cardinality diagnoses—those involving the smallest number of faulty components—is NP-hard and becomes challenging for large systems due to the combinatorial growth of possible fault combinations. SAT-based formulations provide a compact representation of diagnostic constraints, allowing the diagnosis problem to be expressed as a combinatorial optimization problem. Since Ising Machines are well suited for solving such problems, this representation offers a natural pathway for mapping MBD to Ising-based computation. In this work, we present MIBID, a framework that maps SAT-based MBD formulations to an Ising model for computing minimal-cardinality diagnoses. The framework also incorporates hardware-aware pre-processing and decomposition to adapt the formulation to the capabilities of an Ising Machine. Experimental results on a manufactured Ising Machine show competitive performance with SAT-based methods for single minimal diagnoses. For multiple-diagnosis tasks, MIBID enumerates up to \(34\%\) and \(43.59\%\) more diagnoses than state-of-the-art under the weak and strong fault models, respectively, thereby providing broader coverage of plausible fault explanations. Nafisa Sadaf Prova, Ahmet Efe, Abhimanyu Kumar, Chris H. Kim, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ACM Great Lakes Symposium on VLSI | 6 |
| 2026 | SATIC: An Optimizing Ising Compiler for SAT(isfiability)
Ahmet Efe, M. Hüsrev Cilasun, Abhimanyu Kumar, Nafisa Sadaf Prova, Ziqing Zeng, Tahmida Islam, Ruihong Yin, Chaohui Li, Peter Kreye, Chris H. Kim, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ISCA | 12 |
| 2025 | The Case for Secure Miniservers Beyond the EdgeabstractBeyond edge devicescan function off the power grid and without batteries, making them suitable for deployment in hard-to-reach environments. As the energy budget is extremely tight, energy-hungry long-distance communication required for offloading computation or reporting results to a server becomes a significant limitation. Based on the observation that the energy required for communication decreases with shorter distances, this paper makes a case for the deployment ofsecure beyond edge miniservers. These are strategically positioned, lightweight local servers designed to support beyond edge devices without compromising the privacy of sensitive information. We demonstrate that even for relatively small scale representative computations – which are more likely to fit into the tight power budget of a beyond edge device for local processing – deploying a beyond edge miniserver can lead to higher performance. To this end, we consider representative deployment scenarios of practical importance, including but not limited to agricultural systems or building structures, where beyond edge miniservers enable highly energy-efficient real-time data processing. Salonik Resch, M. Hüsrev Cilasun, Zamshed I. Chowdhury, Masoud Zabihi, Yang Lv 0003, Jianping Wang 0006, Sachin S. Sapatnekar, Ismail Akturk, Ulya R. Karpuzcu |
IEEE Trans. Computers | 9 |
| 2024 | On Gate Flip Errors in Computing-In-MemoryabstractComputing-in-memory (CIM) architectures that perform logic gate operations directly within memory arrays, in-situ, are particularly effective in addressing memory-induced performance bottlenecks. When paired with nonvolatile memory, energy efficiency in performing bulk bitwise logic operations can reach unprecedented levels. However, unlocking this potential is not possible if functional correctness is compromised. In this paper we present a CIM-specific class of functional errors termed gate flips, where parametric variations make a logic gate behave as another. Through detailed functional and electrical characterization we demonstrate that gate flips stem from a significant subclass of write errors. Accordingly, we introduce an abstract model to enable efficient functional reliability assessment and to guide design decisions in forming universal CIM gate libraries. We also evaluate the impact on the end accuracy of computation using representative benchmarks. Zamshed I. Chowdhury, M. Hüsrev Cilasun, Salonik Resch, Masoud Zabihi, Yang Lv 0003, Brandon Zink, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
DATE | 9 |
| 2024 | On Error Correction for Nonvolatile Processing-In-MemoryabstractProcessing in memory (PiM) represents a promising computing paradigm to enhance performance of numerous dataintensive applications. Variants performing computing directly in emerging nonvolatile memories can deliver very high energy efficiency. PiM architectures directly inherit the vulnerabilities of the underlying memory substrates, but they also are subject to errors due to the computation in place. Numerous well-established error correcting codes (ECC) for memory exist, and are also considered in the PiM context, however, they typically ignore errors that occur throughout computation. In this paper we revisit the error correction design space for nonvolatile PiM, considering both storage/memory and computation-induced errors, surveying several self-checking and homomorphic approaches. We propose several solutions and analyze their complex performance-area-coverage trade-off, using three representative nonvolatile PiM technologies. All of these solutions guarantee single error correction for both, bulk bitwise computations and ordinary memory/storage errors. M. Hüsrev Cilasun, Salonik Resch, Zamshed I. Chowdhury, Masoud Zabihi, Yang Lv 0003, Brandon Zink, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ISCA | 9 |
| 2023 | On Endurance of Processing in (Nonvolatile) MemoryabstractProcessing-in-Memory (PIM) architectures have gained popularity due to their ability to alleviate the memory wall by performing large numbers of operations within the memory itself. On top of this, nonvolatile memory (NVM) technologies offer highly energy-efficient operations, rendering processing in NVM especially promising. Unfortunately, a major drawback is that NVM has limited endurance. Even when used for standard memory, nonvolatile technologies face limited lifetimes, which is exacerbated by imbalanced usage of memory cells. PIM significantly increases the number of operations the memory is required to perform, making the problem much worse. In this work, we quantitatively analyze the impact of PIM applications on endurance considering representative memory technologies. Our findings indicate that limited endurance can easily block the performance and energy efficiency potential of PIM architectures. Even the best known technologies of today can fall short of meeting practical lifetime expectations. This highlights the importance of research efforts to improve endurance especially at the device technology level. Our study represents the first step in characterizing the very demanding endurance needs of PIM applications to derive a detailed technology level design specification. Salonik Resch, M. Hüsrev Cilasun, Zamshed I. Chowdhury, Masoud Zabihi, Zhengyang Zhao 0001, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ISCA | 8 |
| 2022 | GeNVoM: Read Mapping Near Non-Volatile MemoryabstractDNA sequencing is the physical/biochemical process of identifying the location of the four bases (Adenine, Guanine, Cytosine, Thymine) in a DNA strand. As semiconductor technology revolutionized computing, modern DNA sequencing technology (termed Next Generation Sequencing, NGS) revolutionized genomic research. As a result, modern NGS platforms can sequence hundreds of millions of short DNA fragments in parallel. The sequenced DNA fragments, representing the output of NGS platforms, are termed reads. Besides genomic variations, NGS imperfections induce noise in reads. Mapping each read to (the most similar portion of) a reference genome of the same species, i.e., read mapping, is a common critical first step in a diverse set of emerging bioinformatics applications. Mapping represents a search-heavy memory-intensive similarity matching problem, therefore, can greatly benefit from near-memory processing. Intuition suggests using fast associative search enabled by Ternary Content Addressable Memory (TCAM) by construction. However, the excessive energy consumption and lack of support for similarity matching (under NGS and genomic variation induced noise) renders direct application of TCAM infeasible, irrespective of volatility, where only non-volatile TCAM can accommodate the large memory footprint in an area-efficient way. This paper introduces GeNVoM, a scalable, energy-efficient and high-throughput solution. Instead of optimizing an algorithm developed for general-purpose computers or GPUs, GeNVoM rethinks the algorithm and non-volatile TCAM-based accelerator design together from the ground up. Thereby GeNVoM can improve the throughput by up to 3.67×; the energy consumption, by up to 1.36×, when compared to an ASIC baseline, which represents one of the highest-throughput implementations known. S. Karen Khatamifard, Zamshed I. Chowdhury, Nakul Pande, Meisam Razaviyayn, Chris H. Kim, Ulya R. Karpuzcu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | Energy-efficient and Reliable Inference in Nonvolatile Memory under Extreme Operating ConditionsabstractBeyond-edge devices can operate outside the reach of the power grid and without batteries. Such devices can be deployed in large numbers in regions that are difficult to access. Using machine learning, these devices can solve complex problems and relay valuable information back to a host. Many such devices deployed in low Earth orbit can even be used as nanosatellites. Due to the harsh and unpredictable nature of the environment, these devices must be highly energy-efficient, be capable of operating intermittently over a wide temperature range, and be tolerant of radiation. Here, we propose a non-volatile processing-in-memory architecture that is extremely energy-efficient, supports minimal overhead checkpointing for intermittent computing, can operate in a wide range of temperatures, and has a natural resilience to radiation. Salonik Resch, S. Karen Khatamifard, Zamshed I. Chowdhury, Masoud Zabihi, Zhengyang Zhao 0001, M. Hüsrev Cilasun, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2021 | CAMeleon: Reconfigurable B(T)CAM in Computational RAMabstractEmbedded/edge computing comes with a very stringent hardware resource (area) budget and a need for extreme energy efficiency. This motivates repurposing, i.e., reconfiguring hardware resources on demand, where the overhead of reconfiguration itself is subject to the very same tight budgets in area and energy efficiency. Numerous applications running on resource constrained environments such as wearable devices and Internet-of-Things incorporate CAM (Content Addressable Memory) as a key computational building block. In this paper we present CAMeleon -- a novel energy-efficient compute substrate which can seamlessly be reconfigured to perform CAM operations in addition to logic and memory functions. CAMeleon has a similar level of latency to conventional CAM designs based on SRAM and emerging memory technologies (such as STT-MTJ, ReRAM and PCM), however, performs CAM operations more energy-efficiently, consumes less area, and can support traditional logic and memory functions beyond CAM operations on demand thanks to its reconfigurability. Zamshed I. Chowdhury, Salonik Resch, M. Hüsrev Cilasun, Zhengyang Zhao 0001, Masoud Zabihi, Sachin S. Sapatnekar, Jianping Wang 0006, Ulya R. Karpuzcu |
ACM Great Lakes Symposium on VLSI | 8 |
| 2021 | Spiking Neural Networks in Spintronic Computational RAMabstractSpiking Neural Networks (SNNs) represent a biologically inspired computation model capable of emulating neural computation in human brain and brain-like structures. The main promise is very low energy consumption. Classic Von Neumann architecture based SNN accelerators in hardware, however, often fall short of addressing demanding computation and data transfer requirements efficiently at scale. In this article, we propose a promising alternative to overcome scalability limitations, based on a network of in-memory SNN accelerators, which can reduce the energy consumption by up to 150.25= when compared to a representative ASIC solution. The significant reduction in energy comes from two key aspects of the hardware design to minimize data communication overheads: (1) each node represents an in-memory SNN accelerator based on a spintronic Computational RAM array, and (2) a novel, De Bruijn graph based architecture establishes the SNN array connectivity. M. Hüsrev Cilasun, Salonik Resch, Zamshed I. Chowdhury, Erin Olson, Masoud Zabihi, Zhengyang Zhao 0001, Thomas Peterson, Keshab K. Parhi, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ACM Trans. Archit. Code Optim. | 11 |
| 2020 | CRAFFT: High Resolution FFT Accelerator In Spintronic Computational RAMabstractHigh resolution Fast Fourier Transform (FFT) is important for various applications while increased memory access and parallelism requirement limits the traditional hardware. In this work, we explore acceleration opportunities for high resolution FFTs in spintronic computational RAM (CRAM) which supports true in-memory processing semantics. We experiment with Spin-Torque-Transfer (STT) and Spin-Hall-Effect (SHE) based CRAMs in implementing CRAFFT, a high resolution FFT accelerator in memory. For one million point fixed-point FFT, we demonstrate that CRAFFT can provide up to 2.57× speedup and 673× energy reduction. We also provide a proof-of-concept extension to floating-point FFT. M. Hüsrev Cilasun, Salonik Resch, Zamshed I. Chowdhury, Erin Olson, Masoud Zabihi, Zhengyang Zhao 0001, Thomas Peterson, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
DAC | 10 |
| 2020 | ACR: Amnesic Checkpointing and RecoveryabstractSystematic checkpointing of the machine state makes restart of execution from a safe state possible upon detection of an error. The time and energy overhead of checkpointing, however, grows with the frequency of checkpointing. Considering the growth of expected error rates, amortizing this overhead becomes especially challenging, as checkpointing frequency tends to increase with increasing error rates. Based on the observation that due to imbalanced technology scaling, recomputing a data value can be more energy efficient than retrieving (i.e., loading) a stored copy, this paper explores how recomputation of data values (which otherwise would be read from a checkpoint from memory or secondary storage) can reduce the machine state to be checkpointed, and thereby, the checkpointing overhead. Even in a relatively small scale system, recomputation-based checkpointing can reduce the storage overhead by up to 23.91%; time overhead, by 11.92%; and energy overhead, by 12.53%, respectively. Ismail Akturk, Ulya R. Karpuzcu |
HPCA | 2 |
| 2020 | MOUSE: Inference In Non-volatile Memory for Energy Harvesting ApplicationsabstractThere is increasing demand to bring machine learning capabilities to low power devices. By integrating the computational power of machine learning with the deployment capabilities of low power devices, a number of new applications become possible. In some applications, such devices will not even have a battery, and must rely solely on energy harvesting techniques. This puts extreme constraints on the hardware, which must be energy efficient and capable of tolerating interruptions due to power outages. Here, we propose an in-memory machine learning accelerator utilizing non-volatile spintronic memory. The combination of processing-in-memory and non-volatility provides a key advantage in that progress is effectively saved after every operation. This enables instant shut down and restart capabilities with minimal overhead. Additionally, the operations are highly energy efficient leading to low power consumption. Salonik Resch, S. Karen Khatamifard, Zamshed I. Chowdhury, Masoud Zabihi, Zhengyang Zhao 0001, M. Hüsrev Cilasun, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
MICRO | 9 |
| 2020 | Dual-precision fixed-point arithmetic for low-power ray-triangle intersections
Krishna Rajan, Soheil Hashemi, Ulya R. Karpuzcu, Michael C. Doggett, Sherief Reda |
Comput. Graph. | 3 |
| 2020 | PIMBALL: Binary Neural Networks in Spintronic MemoryabstractNeural networks span a wide range of applications of industrial and commercial significance. Binary neural networks (BNN) are particularly effective in trading accuracy for performance, energy efficiency, or hardware/software complexity. Here, we introduce a spintronic, re-configurable in-memory BNN accelerator, PIMBALL: P rocessing I n M emory B NN A cce L(L) erator, which allows for massively parallel and energy efficient computation. PIMBALL is capable of being used as a standard spintronic memory (STT-MRAM) array and a computational substrate simultaneously. We evaluate PIMBALL using multiple image classifiers and a genomics kernel. Our simulation results show that PIMBALL is more energy efficient than alternative CPU-, GPU-, and FPGA-based implementations while delivering higher throughput. Salonik Resch, S. Karen Khatamifard, Zamshed I. Chowdhury, Masoud Zabihi, Zhengyang Zhao 0001, Jianping Wang 0006, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ACM Trans. Archit. Code Optim. | 8 |
| 2019 | True In-memory Computing with the CRAM: From Technology to ApplicationsabstractNo abstract available. Masoud Zabihi, Zhengyang Zhao 0001, Zamshed I. Chowdhury, Salonik Resch, Mahendra DC, Thomas Peterson, Ulya R. Karpuzcu, Jianping Wang 0006, Sachin S. Sapatnekar |
ACM Great Lakes Symposium on VLSI | 7 |
| 2019 | POWERT Channels: A Novel Class of Covert CommunicationExploiting Power Management VulnerabilitiesabstractTo be able to meet demanding application performance requirements within a tight power budget, runtime power management must track hardware activity at a very fine granularity in both space and time. This gives rise to sophisticated power management algorithms, which need the underlying system to be both highly observable (to be able to sense changes in instantaneous power demand timely) and controllable (to be able to react to changes in instantaneous power demand timely). The end goal is allocating the power budget, which itself represents a very critical shared resource, in a fair way among active tasks of execution. Fundamentally, if not carefully managed, any system-wide shared resource can give rise to covert communication. Power budget does not represent an exception, particularly as systems are becoming more and more observable and controllable. In this paper, we demonstrate how power management vulnerabilities can enable covert communication over a previously unexplored, novel class of covert channels which we will refer to as POWERT channels. We also provide a comprehensive characterization of the POWERT channel capacity under various sharing and activity scenarios. Our analysis based on experiments on representative commercial systems reveal a peak channel capacity of 121.6 bits per second (bps). S. Karen Khatamifard, Amitabh Das, Selçuk Köse, Ulya R. Karpuzcu |
HPCA | 5 |
| 2019 | Special Session: Does Approximation Make Testing Harder (or Easier)?abstractMany important application domains, including machine learning, feature intrinsically noise tolerant algorithms. These algorithms process massive, yet noisy and redundant data, by probabilistic and often iterative techniques. As a result, there is a range of valid outputs rather than a single golden value. While this may translate into relaxed constraints for testing and verification of approximate systems, distinguishing actual design bugs from what is being approximated also becomes harder. In this paper, using representative case studies, we pose several challenges for the test and verification community as approximate computing becomes more prevalent as a design of choice in order to achieve performance gains, power or energy savings, improved reliability or reduced software and/or hardware complexity. R. Iris Bahar, Ulya R. Karpuzcu, Sasa Misailovic |
VTS | 2 |
| 2019 | In-Memory Processing on the Spintronic CRAM: From Hardware Design to Application MappingabstractThe Computational Random Access Memory (CRAM) is a platform that makes a small modification to a standard spintronics-based memory array to organically enable logic operations within the array. CRAM provides a true in-memory computational platform that can perform computations within the memory array, as against other methods that send computational tasks to a separate processor module or a near-memory module at the periphery of the memory array. This paper describes how the CRAM structure can be built and utilized, accounting for considerations at the device, gate, and functional levels. Techniques for constructing fundamental gates are first overviewed, accounting for electrical and noise margin considerations. Next, these logic operations are composed to schedule operations in the array that implement basic arithmetic operations such as addition and multiplication. These methods are then demonstrated on 2D convolution with multibit data, and a binary neural inference engine. The performance of the CRAM is analyzed on near-term and longer-term spintronic device technologies. Significant improvements in energy and execution time for the CRAM-based implementation over a near-memory processing system are demonstrated, and can be attributed to the ability of CRAM to overcome the memory access bottleneck, and to provide high levels of parallelism to the computation. Masoud Zabihi, Zamshed I. Chowdhury, Zhengyang Zhao 0001, Ulya R. Karpuzcu, Jianping Wang 0006, Sachin S. Sapatnekar |
IEEE Trans. Computers | 4 |
| 2019 | Exploiting Algorithmic Noise Tolerance for Scalable On-Chip Voltage RegulationabstractWith the advent of on-chip digital low-dropout (DLDO) regulators, distributed on-chip voltage regulation has become increasingly promising. Environmental and operating conditions have been demonstrated to degrade DLDO performance, which directly affects execution accuracy. The area overhead (OH) needed to compensate aging-induced voltage noise degradation can be significant. Accordingly, in this paper, the algorithmic noise tolerance of certain processor components is exploited as an area-quality control knob to trade the program output quality for area OH. Furthermore, efficient and lightweight techniques utilizing a unidirectional shift register and reduced clock pulsewidth triggering are proposed to realize a novel aging-aware (AA) DLDO to achieve a better area and quality tradeoff. Owing to the large number and distributed nature of voltage regulators, with the proposed design, both the number of regulators utilized in the system and the size of each local regulator are scalable to satisfy the needs of different applications and processor components with varying algorithmic noise tolerance. It is demonstrated through simulation of an IBM POWER8 like processor that the proposed AA design can achieve up to, respectively, 43.2% and 3x transient and steady-state performance improvement. Additionally, more than 10% area OH saving can be achieved over a 5-year period. S. Karen Khatamifard, Ulya R. Karpuzcu, Selçuk Köse |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Mitigation of NBTI induced performance degradation in on-chip digital LDOsabstractOn-chip digital low-dropout voltage regulators (LDOs) have recently gained impetus and drawn significant attention for integration within both mobile devices and micro-processors. Although the benefits of easy integration and fast response speed surpass analog LDOs and other voltage regulator types, NBTI induced performance degradation is typically overlooked. The conventional bi-directional shift register based controller can even exacerbate the degradation, which has been demonstrated theoretically and through practical applications. In this paper, a novel uni-directional shift register is proposed to evenly distribute the electrical stress and mitigate the NBTI effects under arbitrary load conditions with nearly no extra power and area overhead. The benefits of the proposed design as well as reliability aware design considerations are explored and highlighted through simulation of an IBM POWER8 like processor under several benchmark applications. It is demonstrated that the proposed NBTI-aware design can achieve up to 43.2% performance improvement as compared to a conventional one. S. Karen Khatamifard, Ulya R. Karpuzcu, Selçuk Köse |
DATE | 3 |
| 2017 | AMNESIAC: Amnesic Automatic ComputerabstractDue to imbalances in technology scaling, the energy consumption of data storage and communication by far exceeds the energy consumption of actual data production, i.e., computation. As a consequence, recomputing data can become more energy efficient than storing and retrieving precomputed data. At the same time, recomputation can relax the pressure on the memory hierarchy and the communication bandwidth. This study hence assesses the energy efficiency prospects of trading computation for communication. We introduce an illustrative proof-of-concept design, identify practical limitations, and provide design guidelines. Ismail Akturk, Ulya R. Karpuzcu |
ASPLOS | 2 |
| 2017 | ThermoGater: Thermally-Aware On-Chip Voltage RegulationabstractTailoring the operating voltage to fine-grain temporal changes in the power and performance needs of the workload can effectively enhance power efficiency. Therefore, power-limited computing platforms of today widely deploy integrated (i.e., on-chip) voltage regulation which enables fast fine-grain voltage control. Voltage regulators convert and distribute power from an external energy source to the processor. Unfortunately, power conversion loss is inevitable and projected integrated regulator designs are unlikely to eliminate this loss even asymptotically. Reconfigurable power delivery by selective shut-down, i.e., gating, of distributed on-chip regulators in response to spatio-temporal changes in power demand can sustain operation at the minimum conversion loss. However, even the minimum conversion loss is sizable, and as conversion loss gets dissipated as heat, on-chip regulators can easily cause thermal emergencies due to their small footprint. S. Karen Khatamifard, Weize Yu, Selçuk Köse, Ulya R. Karpuzcu |
ISCA | 5 |
| 2017 | Efficiency, Stability, and Reliability Implications of Unbalanced Current Sharing Among Distributed On-Chip Voltage RegulatorsabstractPower delivery networks with distributed on-chip voltage regulators (VRs) serve as an effective way for fast localized voltage regulation within modern microprocessors. Without careful consideration of the interactions among the distributed VRs and the power grid, unbalanced current sharing (CS) among those regulators may, however, lead to efficiency degradations, stability, and reliability issues, and even malfunctions of the regulators. This paper is a first attempt to investigate the efficiency, stability, and reliability implications of unbalanced CS among distributed on-chip VRs. Benefits of balanced CS are demonstrated with concrete examples, showing the necessity of an appropriate current balancing scheme. An adaptive reference voltage control method and the corresponding control algorithms specifically for distributed on-chip VRs are proposed to balance the CS among regulators at different locations. The proposed techniques successfully balance the CS among distributed VRs and can be applied to different regulator types. Simulation results based on practical microprocessor setups confirm the efficiency, stability, and reliability implications. S. Karen Khatamifard, Orhun Aras Uzun, Ulya R. Karpuzcu, Selçuk Köse |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | VARIUS-TC: A modular architecture-level model of parametric variation for thin-channel switchesabstractUnder aggressive miniaturization, unconventional digital switches rapidly come to light, which introduce new sources of variation in design parameters, and hence challenge the manufacturing process further. As a result, performance and power of manufactured hardware becomes greatly unpredictable. Characterizing variation-incurred unpredictability at early stages of the design necessitates dependable architecture-level models of variation, which distill device- and circuit-level details to accurately evaluate system-level implications. In this paper, we introduce a modular architecture-level model of parametric variation to address this challenge. As a case study, we refine our discussion to a representative class of emerging thin-channel switches, FinFETs. S. Karen Khatamifard, Michael Resch 0002, Nam Sung Kim, Ulya R. Karpuzcu |
ICCD | 4 |
| 2016 | Snatch: Opportunistically reassigning power allocation between processor and memory in 3D stacksabstractThe pin count largely determines the cost of a chip package, which is often comparable to the cost of a die. In 3D processor-memory designs, power and ground (P/G) pins can account for the majority of the pins. This is because packages include separate pins for the disjoint processor and memory power delivery networks (PDNs). Supporting separate PDNs and P/G pins for processor and memory is inefficient, as each set has to be provisioned for the worst-case power delivery requirements. In this paper, we propose to reduce the number of P/G pins of both processor and memory in a 3D design, and dynamically and opportunistically divert some power between the two PDNs on demand. To perform the power transfer, we use a small bidirectional on-chip voltage regulator that connects the two PDNs. Our concept, called Snatch, is effective. It allows the computer to execute code sections with high processor or memory power requirements without having to throttle performance. We evaluate Snatch with simulations of an 8-core multicore stacked with two memory dies. In a set of compute-intensive codes, the processor snatches memory power for 30% of the time on average, speeding-up the codes by up to 23% over advanced turbo-boosting; in memory-intensive codes, the memory snatches processor power. Alternatively, Snatch can reduce the package cost by about 30%. Dimitrios Skarlatos 0002, Renji Thomas, Aditya Agrawal, Shibin Qin, Robert C. N. Pilawa-Podgurski, Ulya R. Karpuzcu, Radu Teodorescu, Nam Sung Kim, Josep Torrellas |
MICRO | 6 |
| 2016 | Accuracy Bugs: A New Class of Concurrency Bugs to Exploit Algorithmic Noise ToleranceabstractParallel programming introduces notoriously difficult bugs, usually referred to as concurrency bugs. This article investigates the potential for deviating from the conventional wisdom of writing concurrency bug--free, parallel programs. It explores the benefit of accepting buggy but approximately correct parallel programs by leveraging the inherent tolerance of emerging parallel applications to inaccuracy in computations. Under algorithmic noise tolerance, a new class of concurrency bugs, accuracy bugs, degrade the accuracy of computation (often at acceptable levels) rather than causing catastrophic termination. This study demonstrates how embracing accuracy bugs affects the application output quality and performance and analyzes the impact on execution semantics. Ismail Akturk, Riad Akram, Mohammad Majharul Islam, Abdullah Muzahid, Ulya R. Karpuzcu |
ACM Trans. Archit. Code Optim. | 5 |
| 2016 | System-Level Power Analysis of a Multicore Multipower Domain Processor With ON-Chip Voltage RegulatorsabstractIn this paper, we study two different ON-chip power delivery schemes, namely, fully integrated voltage regulator (FIVR) and low-dropout regulator (LDO), and analyze their effect on total system power under process variation, assuming a realistic dynamic voltage-frequency scaling (DVFS) system. The impact of different task scheduling algorithms on the overall system power was also analyzed. We find that in a hypothetical 256-core processor, under a per-core DVFS assumption, the FIVR-based power delivery consumes 20% less power than the LDO-based one for a 50% throughput. However, as the number of cores in the processor reduces, the difference in power consumption between the FIVR-based and LDO-based power delivery schemes becomes smaller. For example, in the case of a 16-core processor with per-core DVFS capability, FIVR-based design was found to consume about the same power as the LDO-based design. Ayan Paul, Sang Phill Park, Dinesh Somasekhar, Young Moon Kim, Nitin Borkar, Ulya R. Karpuzcu, Chris H. Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2015 | Comparison of single-ISA heterogeneous versus wide dynamic range processors for mobile applicationsabstractMobile computing devices demand processors to offer a wide range of performance/power trade-offs so that they can provide much needed high performance or low power consumption depending on a given operating requirement. While dynamic voltage/frequency scaling (DVFS) has been the most powerful technique to provide such trade-offs, few processor vendors have the capability to provide a sufficient DVFS range requiring joint optimization of devices and circuits. Facing such a challenge, two promising approaches are proposed: scaling the amount of processor resources such as on-chip memory and execution units, i.e., dynamic resource scaling (DRS) and switching between big out-of-order (OoO) and little in-order cores in a single-ISA heterogeneous processor such as ARM's big. LITTLE. In this paper, we compare a single-ISA heterogeneous processor with a wide dynamic range (WDR) processor augmented with DRS in terms of (1) device-, circuit-, architecture-level implications, (2) design challenges, (3) area, (4) performance, and (5) energy efficiency. We evaluate a big. LITTLE processor (as a representative of single-ISA heterogeneous processor) based on Cortex-A15/A7 and a WDR processor based on Cortex-A15 running various mobile and SPEC2006 benchmarks. Our experiments demonstrate that the WDR processor combined with DRS can deliver energy efficiency close to the single-ISA heterogeneous processor, depending on the power overhead of circuit implementation to provide the wide DVFS range. Hamid Reza Ghasemi, Ulya R. Karpuzcu, Nam Sung Kim |
ICCD | 2 |
| 2014 | Accordion: Toward soft Near-Threshold Voltage ComputingabstractWhile more cores can find place in the unit chip area every technology generation, excessive growth in power density prevents simultaneous utilization of all. Due to the lower operating voltage, Near-Threshold Voltage Computing (NTC) promises to fit more cores in a given power envelope. Yet NTC prospects for energy efficiency disappear without mitigating (i) the performance degradation due to the lower operating frequency; (ii) the intensified vulnerability to parametric variation. To compensate for the first barrier, we need to raise the degree of parallelism - the number of cores engaged in computation. NTC-prompted power savings dominate the power cost of increasing the core count. Hence, limited parallelism in the application domain constitutes the critical barrier to engaging more cores in computation. To avoid the second barrier, the system should tolerate variation-induced errors. Unfortunately, engaging more cores in computation exacerbates vulnerability to variation further. To overcome NTC barriers, we introduce Accordion, a novel, light-weight framework, which exploits weak scaling along with inherent fault tolerance of emerging R(ecognition), M(ining), S(ynthesis) applications. The key observation is that the problem size not only dictates the number of cores engaged in computation, but also the application output quality. Consequently, Accordion designates the problem size as the main knob to trade off the degree of parallelism (i.e. the number of cores engaged in computation), with the degree of vulnerability to variation (i.e. the corruption in application output quality due to variation-induced errors). Parametric variation renders ample reliability differences between the cores. Since RMS applications can tolerate faults emanating from data-intensive program phases as opposed to control, variation-afflicted Accordion hardware executes fault-tolerant data-intensive phases on error-prone cores, and reserves reliable cores for control. Ulya R. Karpuzcu, Ismail Akturk, Nam Sung Kim |
HPCA | 1 |
| 2014 | Low-Cost Per-Core Voltage Domain Support for Power-Constrained High-Performance ProcessorsabstractPer-core voltage domains can improve performance under a power constraint. Most commercial processors, however, only have a single voltage domain for all processor cores. This is because splitting the single voltage domain into per-core voltage domains and powering them with multiple off-chip voltage regulators (VRs) incur a high cost for the platform and package designs. Although using on-chip switching VRs can be an alternative solution, integrating high-quality inductors for VRs with cores has been a technical challenge. In this paper, we propose a cost-effective power delivery technique to support per-core voltage domains. Our technique is based on the observations that: 1) core-to-core (C2C) voltage variations are relatively small for most execution intervals when the voltages/frequencies are optimized to maximize performance under a power constraint and 2) per-core power-gating devices augmented with feedback control circuitry can serve as low-cost VRs that can provide high efficiency in situations like 1). Our experimental results show that processors using our technique can achieve power efficiency as high as those using the per-core on-chip switching VRs at a much lower cost. Abhishek A. Sinkar, Hamid Reza Ghasemi, Michael J. Schulte, Ulya R. Karpuzcu, Nam Sung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | EnergySmart: Toward energy-efficient manycores for Near-Threshold ComputingabstractWhile Near-Threshold Voltage Computing (NTC) is a promising approach to push back the manycore power wall, it suffers from a high sensitivity to parameter variations. One possible way to cope with variations is to use multiple on-chip voltage (Vdd) domains. However, this paper finds that such an approach is energy inefficient. Consequently, for NTC, we propose a manycore organization that has a single Vdddomain and relies on multiple frequency domains to tackle variation. We call it EnergySmart. For this approach to be competitive, it has to be paired with effective core assignment strategies and also support fine-grain (i.e., short-interval) DVFS. This paper shows that, at NTC, a simple chip with a single Vdddomain can deliver a higher performance per watt than one with multiple Vdddomains. Ulya R. Karpuzcu, Abhishek A. Sinkar, Nam Sung Kim, Josep Torrellas |
HPCA | 1 |
| 2012 | VARIUS-NTV: A microarchitectural model to capture the increased sensitivity of manycores to process variations at near-threshold voltagesabstractNear-Threshold Computing (NTC), where the supply voltage is only slightly higher than the threshold voltage of transistors, is a promising approach to attain energy-efficient computing. Unfortunately, compared to the conventional Super-Threshold Computing (STC), NTC is more sensitive to process variations, which results in higher power consumption and lower frequencies than would otherwise be possible, and potentially a non-negligible fault rate. To help address variations at NTC at the architecture level, this paper presents the first microarchitectural model of process variations for NTC. The model, called VARIUS-NTV, extends the existing VARIUS variation model. Its key aspects include: (i) adopting a gate-delay model and an SRAM cell type that are tailored to NTC, (ii) modeling SRAM failure modes emerging at NTC, and (iii) accounting for the impact of leakage in SRAM models. We evaluate a simulated 11nm, 288-core tiled manycore at both NTC and STC. The results show higher frequency and power variations within the NTC chip. For example, the maximum difference in on-chip tile frequency is ≈2.3× at STC and ≈3.7× at NTC. We also validate our model against an experimental chip. Ulya R. Karpuzcu, Krishna B. Kolluru, Nam Sung Kim, Josep Torrellas |
DSN | 1 |
| 2010 | LeadOut: Composing low-overhead frequency-enhancing techniques for single-thread performance in configurable multicoresabstractDespite the ubiquity of multicores, it is as important as ever to deliver high single-thread performance. An appealing way to accomplish this is by shutting down the idle cores in the chip and running the busy, performance-critical core(s) at higher-than-nominal frequencies. To enable such frequencies, two low-overhead approaches either boost voltage beyond nominal values, or pair cores in leader-checker configurations and let them run beyond safe frequency margins. We observe that, in a large multicore with varying numbers of busy cores, individual application of either of these two techniques is suboptimal. Each alone is often unable to bring the multicore all the way to its power or temperature envelopes due to limitations in supply voltage or error rate. Moreover, we show that the two techniques are complementary, and can be synergistically combined to unlock much higher levels of single-thread performance. Finally, we demonstrate a dynamic controller that optimizes the two techniques. Our data shows that, given a 16-core multi-core where half of the cores are already busy, an additional, performance-critical thread now attains 34% higher performance than before, while consuming 220% more power. Brian Greskamp, Ulya R. Karpuzcu, Josep Torrellas |
HPCA | 2 |
| 2009 | Blueshift: Designing processors for timing speculation from the ground upabstractSeveral recent processor designs have proposed to enhance performance by increasing the clock frequency to the point where timing faults occur, and by adding error-correcting support to guarantee correctness. However, such timing speculation (TS) proposals are limited in that they assume traditional design methodologies that are suboptimal under TS. In this paper, we present a new approach where the processor itself is designed from the ground up for TS. The idea is to identify and optimize the most frequently-exercised critical paths in the design, at the expense of the majority of the static critical paths, which are allowed to suffer timing errors. Our approach and design optimization algorithm are called BlueShift. We also introduce two techniques that, when applied under BlueShift, improve processor performance: on-demand selective biasing (OSB) and path constraint tuning (PCT). Our evaluation with modules from the OpenSPARC T1 processor shows that, compared to conventional TS, BlueShift with OSB speeds up applications by an average of 8% while increasing the processor power by an average of 12%. Moreover, compared to a high-performance TS design, BlueShift with PCT speeds up applications by an average of 6% with an average processor power overhead of 23% . providing a way to speed up logic modules that is orthogonal to voltage scaling. Brian Greskamp, Ulya R. Karpuzcu, Jeffrey J. Cook, Josep Torrellas, Deming Chen, Craig B. Zilles |
HPCA | 3 |
| 2009 | Accurate microarchitecture-level fault modeling for studying hardware faultsabstractDecreasing hardware reliability is expected to impede the exploitation of increasing integration projected by Moore's Law. There is much ongoing research on efficient fault tolerance mechanisms across all levels of the system stack, from the device level to the system level. High-level fault tolerance solutions, such as at the microarchitecture and system levels, are commonly evaluated using statistical fault injections with microarchitecture-level fault models. Since hardware faults actually manifest at a much lower level, it is unclear if such high level fault models are acceptably accurate. On the other hand, lower level models, such as at the gate level, may be more accurate, but their increased simulation times make it hard to track the system-level propagation of faults. Thus, an evaluation of high-level reliability solutions entails the classical tradeoff between speed and accuracy. This paper seeks to quantify and alleviate this tradeoff. We make the following contributions: (1) We introduce SWAT-Sim, a novel fault injection infrastructure that uses hierarchical simulation to study the system-level manifestations of permanent (and transient) gate-level faults. For our experiments, SWAT-Sim incurs a small average performance overhead of under 3x, for the components we simulate, when compared to pure microarchitectural simulations. (2) We study system-level manifestations of faults injected under different microarchitecture-level and gate-level fault models and identify the reasons for the inability of microarchitecture-level faults to model gate-level faults in general. (3) Based on our analysis, we derive two probabilistic microarchitecture-level fault models to mimic gate-level stuck-at and delay faults. Our results show that these models are, in general, inaccurate as they do not capture the complex manifestation of gate-level faults. The inaccuracies in existing models and the lack of more accurate microarchitecture-level models motivate using infrastructures similar to SWAT-Sim to faithfully model the microarchitecture-level effects of gate-level faults. Man-Lap Li, Pradeep Ramachandran, Ulya R. Karpuzcu, Siva Kumar Sastry Hari, Sarita V. Adve |
HPCA | 3 |
| 2009 | The BubbleWrap many-core: popping cores for sequential accelerationabstractMany-core scaling now faces a power wall. The gap between the number of cores that fit on a die and the number that can operate simultaneously under the power budget is rapidly increasing with technology scaling. In future designs, many of the cores may have to be dormant at any given time to meet the power budget. To push back the many-core power wall, this paper proposes Dynamic Voltage Scaling for Aging Management (DVSAM) --- a new scheme for managing processor aging to attain higher performance or lower power consumption. In addition, this paper introduces the BubbleWrap many-core, a novel architecture that makes extensive use of DVSAM. BubbleWrap identifies the most power-efficient set of cores in a variation-affected chip --- the largest set that can be simultaneously powered-on --- and designates them as Throughput cores dedicated to parallel-section execution. The rest of the cores are designated as Expendable and are dedicated to accelerating sequential sections. BubbleWrap attains maximum sequential acceleration by sacrificing Expendable cores one at a time, running them at elevated supply voltage for a significantly shorter service life each, until they completely wear-out and are discarded --- figuratively, as if popping bubbles in bubble wrap that protects Throughput cores. In simulated 32-core chips, BubbleWrap provides substantial improvements over a plain chip. For example, on average, one design runs fully-sequential applications at a 16% higher frequency, and fully-parallel ones with a 30% higher throughput. Ulya R. Karpuzcu, Brian Greskamp, Josep Torrellas |
MICRO | 1 |