EDBT 2026 Demo / reviewers in the wild / expert
Koushik Chakraborty
dblp:92/4010
· DBLP profile ↗
59ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-0228-2737ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 14 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probabilistic Verification for Modular Network-on-Chip Systems
Nick Waddoups, Jonah Boe, Arnd Hartmanns, Prabal Basu, Sanghamitra Roy, Koushik Chakraborty, Zhen Zhang 0006 |
VMCAI | 6 |
| 2026 | LOgIQ: Log-Domain Optimization for Hardware-Efficient Inference and Quantization of TransformersabstractQuantization enables efficient large language model (LLM) inference by reducing memory and computation cost, but most existing methods sacrifice hardware simplicity through floating-point fallback, dequantization logic, or metadata routing. These overheads disrupt datapath regularity and limit gains in timing and area, especially in multiply-accumulate (MAC)-heavy transformer layers. We propose LogiMAC, a hardware-aligned quantization scheme that encodes weights in the log domain using shift-based arithmetic. LogiMAC embeds all decoding logic within the MAC unit, eliminates runtime scaling, and supports sublevel refinement without multipliers or look-up tables. The resulting datapath supports both shift directions with optional gating for unused paths. We evaluate LogiMAC across seven modern LLMs (ranging from 2B to 8B parameters) and compare it to strong baselines including SmoothQuant, GOBO, and SpinQuant. Our results show that while functional accuracy remains competitive, LogiMAC reduces area by$1.8\times $and dynamic power by$5.9\times $(in power-gated mode) compared to standard integer baselines. Timing analysis reveals over 50% of LogiMAC paths complete within just12%of the normalized delay seen in other schemes. Validated over 34 million generated tokens, our findings demonstrate that accurate low-bit quantization can coexist with extreme hardware efficiency. Tanzeel-ur-Rehman Khan, Sanghamitra Roy, Koushik Chakraborty |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | STRIVE: Empowering a Low Power Tensor Processing Unit with Fault Detection and Error ResilienceabstractRapid growth in Deep Neural Network (DNN) workloads has increased the energy footprint of the Artificial Intelligence (AI) computing realm. For optimum energy efficiency, we propose operating a DNN hardware in the Low-Power Computing (LPC) region. However, operating at LPC causes increased delay sensitivity to Process Variation (PV). Delay faults are an intriguing consequence of PV. In this article, we demonstrate the vulnerability of DNNs to delay variations, substantially lowering the prediction accuracy. To overcome delay faults, we present STRIVE—a post-fabrication fault detection and reactive error reduction technique. We also introduce a time-borrow correction technique to ensure error-free DNN computation. Noel Daniel Gundi, Sanghamitra Roy, Koushik Chakraborty |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | STRIVE: Enabling Choke Point Detection and Timing Error Resilience in a Low-Power Tensor Processing UnitabstractRapid growth in Deep Neural Network (DNN) workloads has increased the energy footprint of the Artificial Intelligence (AI) computing realm. For optimum energy efficiency, we propose operating a DNN hardware in the Low-Power Computing (LPC) region. However, operating at LPC causes increased delay sensitivity to Process Variation (PV). Delay faults are an intriguing consequence of PV. In this paper, we demonstrate the vulnerability of DNNs to delay variations, substantially lowering the prediction accuracy. To overcome delay faults, we present STRIVE—a post-fabrication fault detection and reactive error reduction technique. We also introduce a time-borrow correction technique to ensure error-free DNN computation. Noel Daniel Gundi, Zinnia Muntaha Mowri, Andrew Chamberlin, Sanghamitra Roy, Koushik Chakraborty |
DAC | 5 |
| 2021 | UPTPU: Improving Energy Efficiency of a Tensor Processing Unit through Underutilization Based Power-GatingabstractThe AI boom is bringing a plethora of domain-specific architectures for Neural Network computations. Google’s Tensor Processing Unit (TPU), a Deep Neural Network (DNN) accelerator, has replaced the CPUs/GPUs in its data centers, claiming more than 15 × rate of inference. However, the unprecedented growth in DNN workloads with the widespread use of AI services projects an increasing energy consumption of TPU based data centers. In this work, we parametrize the extreme hardware underutilization in TPU systolic array and propose UPTPU: an intelligent, dataflow adaptive power-gating paradigm to provide a staggering 3.5 × – 6.5× energy efficiency to TPU for different input batch sizes. Pramesh Pandey, Noel Daniel Gundi, Koushik Chakraborty, Sanghamitra Roy |
DAC | 3 |
| 2021 | Probabilistic Verification for Reliability of a Two-by-Two Network-on-Chip System
Riley Roberts, Arnd Hartmanns, Prabal Basu, Sanghamitra Roy, Koushik Chakraborty, Zhen Zhang 0006 |
FMICS | 6 |
| 2021 | EFFORT: A Comprehensive Technique to Tackle Timing Violations and Improve Energy Efficiency of Near-Threshold Tensor Processing UnitsabstractModern deep neural network (DNN) applications demand a remarkable processing throughput usually unmet by traditional Von Neumann architectures. Consequently, hardware accelerators, comprising a sea of multiplier-and-accumulate (MAC) units, have recently gained prominence in accelerating DNN inference engine. For example, tensor processing units (TPUs) account for a lion’s share of Google’s datacenter inference operations. The proliferation of real-time DNN predictions is accompanied by a tremendous energy budget. In quest of trimming the energy footprint of DNN accelerators, we propose Energy eFFicient and errOr Resilient TPU (EFFORT)—an energy optimized, yet high-performance TPU architecture, operating at the near-threshold computing (NTC) region. EFFORT promotes a better-than-worst case design by operating the NTC TPU at a substantially high frequency while keeping the voltage at the NTC nominal value. In order to tackle the timing errors due to such aggressive operation, we employ an opportunistic error mitigation strategy. In addition, we implement anin situclock gating architecture, drastically reducing the MACs’ dynamic power consumption. Compared to a cutting-edge error mitigation technique for TPUs, EFFORT enables up to$2.5\times $better performance at NTC with only 4% average accuracy drop across six out of eight DNN benchmarks. Noel Daniel Gundi, Tahmoures Shabanian, Prabal Basu, Pramesh Pandey, Sanghamitra Roy, Koushik Chakraborty |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | EFFORT: Enhancing Energy Efficiency and Error Resilience of a Near-Threshold Tensor Processing UnitabstractModern deep neural network (DNN) applications demand a remarkable processing throughput usually unmet by traditional Von Neumann architectures. Consequently, hardware accelerators, comprising a sea of multiplier and accumulate (MAC) units, have recently gained prominence in accelerating DNN inference engine. For example, Tensor Processing Units (TPU) account for a lion's share of Google's datacenter inference operations. The proliferation of real-time DNN predictions is accompanied with a tremendous energy budget. In quest of trimming the energy footprint of DNN accelerators, we propose EFFORT-an energy optimized, yet high performance TPU architecture, operating at the Near-Threshold Computing (NTC) region. EFFORT promotes a better-than-worst-case design by operating the NTC TPU at a substantially high frequency while keeping the voltage at the NTC nominal value. In order to tackle the timing errors due to such aggressive operation, we employ an opportunistic error mitigation strategy. Additionally, we implement an in-situ clock gating architecture, drastically reducing the MACs' dynamic power consumption. Compared to a cutting-edge error mitigation technique for TPUs, EFFORT enables up to 2.5× better performance at NTC with only 2% average accuracy drop across 3 out of 4 DNN datasets. Noel Daniel Gundi, Tahmoures Shabanian, Prabal Basu, Pramesh Pandey, Sanghamitra Roy, Koushik Chakraborty, Zhen Zhang 0006 |
ASP-DAC | 6 |
| 2020 | GreenTPU: Predictive Design Paradigm for Improving Timing Error Resilience of a Near-Threshold Tensor Processing UnitabstractThe emergence of hardware accelerators has brought about several orders of magnitude improvement in the speed of the deep neural-network (DNN) inference. Among such DNN accelerators, the Google tensor processing unit (TPU) has transpired to be the best-in-class, offering more than 15× speedup over the contemporary GPUs. However, the rapid growth in several DNN workloads conspires to escalate the energy consumptions of the TPU-based data-centers. In order to restrict the energy consumption of TPUs, we propose GreenTPU-a low-power near-threshold (NTC) TPU design paradigm. To ensure a high inference accuracy at a low-voltage operation, GreenTPU identifies the patterns in the error-causing activation sequences in the systolic array, and prevents further timing errors from similar patterns by intermittently boosting the operating voltage of the specific multiplier-and-accumulator units in the TPU. Compared to a cutting-edge timing error mitigation technique for TPUs, GreenTPU enables 2× to 3× higher performance (TOPS) in an NTC TPU, with a minimal loss in the prediction accuracy. Pramesh Pandey, Prabal Basu, Koushik Chakraborty, Sanghamitra Roy |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Exploring Warp Criticality in Near-Threshold GPGPU Applications Using a Dynamic Choke Point AnalysisabstractGeneral-purpose graphics processing units (GPGPUs), due to their enormous parallelism, have found ubiquitous applications in parallel computing. However, their peak power rating has also increased over the years. As a consequence, near-threshold computing (NTC) has come to the rescue. However, a severe device-level delay variability arising from process variation (PV) can significantly diminish the NTC system performance. In this article, we examine choke points-a unique device-level characteristic of PV at NTC-that can exacerbate the delays of the GPGPU parallel warps. In order to improve the NTC GPU performance, we propose a family of holistic circuit-architectural solutions, referred to as choke-point-aware warp speculator (CPAWS). CPAWS identifies the choke point-induced critical warps in GPGPU applications and improves their execution latencies. Compared to a state-of-the-art warp scheduling policy, our best scheme improves the performance and energy efficiency of an NTC GPU by ~39% and ~31%, respectively. Sourav Sanyal, Prabal Basu, Aatreyi Bal, Sanghamitra Roy, Koushik Chakraborty |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | GreenTPU: Improving Timing Error Resilience of a Near-Threshold Tensor Processing UnitabstractThe emergence of hardware accelerators has brought about several orders of magnitude improvement in the speed of the deep neural-network (DNN) inference. Among such DNN accelerators, Google Tensor Processing Unit (TPU) has transpired to be the best-in-class, offering more than 15× speedup over the contemporary GPUs. However, the rapid growth in several DNN workloads conspires to escalate the energy consumptions of the TPU-based data-centers. In order to restrict the energy consumption of TPUs, we propose Green TPU---a low-power near-threshold (NTC) TPU design paradigm. To ensure a high inference accuracy at a low-voltage operation, GreenTPU identifies the patterns in the error-causing activation sequences in the systolic array, and prevents further timing errors from the same sequence by intermittently boosting the operating voltage of the specific multiplier-and-accumulator units in the TPU. Compared to a cutting-edge timing error mitigation technique for TPUs, GreenTPU enables 2X--3X higher performance in an NTC TPU, with a minimal loss in the prediction accuracy. Pramesh Pandey, Prabal Basu, Koushik Chakraborty, Sanghamitra Roy |
DAC | 3 |
| 2019 | Predicting Critical Warps in Near-Threshold GPGPU Applications using a Dynamic Choke Point AnalysisabstractGeneral purpose graphics processing units (GP-GPU) can significantly improve the power consumption at the NTC operating region. However, process variation (PV) can drastically reduce its performance. In this paper, we examine choke points-a unique device-level characteristic of PV at NTC-that can exacerbate the warp criticality problem. We show that the modern warp schedulers cannot tackle the choke point induced critical warps in an NTC GPU. We propose Warp Latency Booster, a circuit-architectural solution to dynamically predict the critical warps and accelerate them in their respective execution units. Our best scheme achieves an average improvement of ~32% and ~41% in performance, and ~21% and ~19% in energy-efficiency, respectively, over two state-of-the-art warp schedulers. Sourav Sanyal, Prabal Basu, Aatreyi Bal, Sanghamitra Roy, Koushik Chakraborty |
DATE | 5 |
| 2019 | Probabilistic Verification for Reliable Network-on-Chip System Design
Arnd Hartmanns, Prabal Basu, Rajesh J. S., Koushik Chakraborty, Sanghamitra Roy, Zhen Zhang 0006 |
FMICS | 5 |
| 2018 | Trident: A comprehensive timing error resilient technique against choke points at NTC
Aatreyi Bal, Sanghamitra Roy, Koushik Chakraborty |
DATE | 3 |
| 2018 | Reliability and Uniformity Enhancement in 8T-SRAM based PUFs operating at NTCabstractSRAM-based PUFs (SPUFs) have emerged as promising security primitives for low-power devices. However, operating 8T-SPUFs at Near-Threshold Computing (NTC) realm is plagued by exacerbated process variation (PV) sensitivity which thwarts their reliable operation. In this paper, we demonstrate the massive degradation in the reliability and uniformity characteristics of 8T-SPUF. By exploiting the opportunities bestowed by schematic asymmetry of 8T-SPUF cells, we propose biasing and sizing based design strategies. Our techniques achieve an immense improvement of more than 55% in the percentage of unreliable cells and improves the proximity to ideal uniformity by 82%, over a baseline NTC 8T-SPUF with no enhancement. Pramesh Pandey, Asmita Pal, Koushik Chakraborty, Sanghamitra Roy |
ISLPED | 3 |
| 2018 | ACE-GPU: Tackling Choke Point Induced Performance Bottlenecks in a Near-Threshold Computing GPUabstractThe proliferation of multicore devices with a strict thermal budget has aided to the research in Near-Threshold Computing (NTC). However, the operation of a Graphics Processing Unit (GPU) at the NTC region has still remained recondite. In this work, we explore an important reliability predicament of NTC, called choke points, that severely throttles the performance of GPUs. Employing a cross-layer methodology, we demonstrate the potency of choke points in inducing timing errors in a GPU, operating at the NTC region. We propose a holistic circuit-architectural solution, that promotes an energy-efficient NTC-GPU design paradigm by gracefully tackling the choke point induced timing errors. Our proposed scheme offers 3.18x and 88.5% improvements in NTC-GPU performance and energy delay product, respectively, over a state-of-the-art timing error mitigation technique, with marginal area and power overheads. Tahmoures Shabanian, Aatreyi Bal, Prabal Basu, Koushik Chakraborty, Sanghamitra Roy |
ISLPED | 4 |
| 2018 | Trident: Comprehensive Choke Error Mitigation in NTC Systems
Aatreyi Bal, Sanghamitra Roy, Koushik Chakraborty |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Dynamic Choke Sensing for Timing Error Resilience in NTC SystemsabstractProcess variation (PV) is a conspicuous predicament for submicrometer VLSI circuits. In this paper, we illustrate “choke points” as a vital consequence of PV in the near-threshold computing domain. Choke points are PV affected sensitized logic gates with increased delay deviation. They dominate the choice of critical paths postfabrication. To mitigate the timing errors induced thereby, we propose dynamic choke sensing (DCS). This technique senses the timing error causing opcode sequences, and uses the knowledge to prevent similar sequences from causing errors in the future. We propose two variants of our scheme. Our techniques provide ~55% improvement in performance and ~73% improvement in energy efficiency as compared with popular timing error mitigation scheme, Razor, with minimal area and power overheads. Aatreyi Bal, Shamik Saha, Sanghamitra Roy, Koushik Chakraborty |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Revamping timing error resilience to tackle choke points at NTC systemsabstractProcess variation is a conspicuous predicament for sub-micron VLSI circuits. In this paper, we illustrate “choke points” as a vital consequence of process variation in the Near Threshold Computing (NTC) domain. Choke points are process variation affected sensitized logic gates with increased delay deviation. They dominate the choice of critical paths postfabrication. To mitigate the timing errors induced thereby, we propose Dynamic Choke Sensing (DCS). This technique senses the timing error causing opcode sequences, and uses the knowledge to prevent similar sequences from causing errors in future. Our scheme provides 25%-160% improvement in performance and 50%-90% improvement in energy efficiency as compared to contemporary timing error mitigation schemes, with minimal area and power overheads. Aatreyi Bal, Shamik Saha, Sanghamitra Roy, Koushik Chakraborty |
DATE | 4 |
| 2017 | SSAGA: SMs Synthesized for Asymmetric GPGPU ApplicationsabstractThe emergence of GPGPU applications, bolstered by flexible GPU programming platforms, has created a tremendous challenge in maintaining high energy efficiency in modern GPUs. In this article, we demonstrate that customizing a Streaming Multiprocessor (SM) of a GPU at a lower frequency is significantly more energy efficient compared to employing DVFS on an SM designed for a high-frequency operation. Using a system-level CAD technique, we proposeSSAGA—Streaming Multiprocessors Synthesized for Asymmetric GPGPU Applications—an energy-efficient GPU design paradigm. SSAGA creates architecturally identical SM cores, customized for different voltage-frequency domains. Our rigorous cross-layer methodology demonstrates an average of 20% improvement in energy efficiency over a spatially multitasking GPU across a range of GPGPU applications. Shamik Saha, Prabal Basu, Chidhambaranathan Rajamanikkam, Aatreyi Bal, Koushik Chakraborty, Sanghamitra Roy |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2017 | IcoNoClast: Tackling Voltage Noise in the NoC Power Supply Through Flow-Control and Routing AlgorithmsabstractPower supply noise (PSN) is a growing concern in modern multiprocessor system-on-chips (MPSoCs). The advent of new architectures, such as the network-on-chip (NoC), the standard for on-chip communication in MPSoCs, has given rise to new challenges in maintaining reliable and energy-efficient operation. The growing NoC power footprint, increase in the transistor current, and high switching speed of the logic devices exacerbate the peak PSN in the NoC power delivery network (PDN). Hence, preserving power supply integrity in the NoC PDN is critical. In this paper, we propose IcoNoClast, a collection of a novel flow-control protocol (PAF) and an adaptive routing algorithm (PSN-aware routing), to mitigate the PSN in NoCs. Our best scheme achieves~15% and~12% improvements in the regional peak PSN and energy efficiency across a range of PARSEC benchmarks, with a 4.1% performance overhead and marginal area and power footprints. Prabal Basu, Rajesh J. S., Koushik Chakraborty, Sanghamitra Roy |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | SwiftGPU: fostering energy efficiency in a near-threshold GPU through a tactical performance boostabstractIn this paper, we investigate the challenges of preserving energy-efficiency in a Near-Threshold Computing (NTC) GPU. Two key factors can significantly undermine the efficacy of GPUs at NTC: (a) elongated delays at NTC make the GPU applications severely sensitive toMulti-cycle Latency Datapaths (MLDs) within the GPU pipeline; and (b) process variation (PV) at NTC induces a substantial performance variance. To address these emerging challenges, we propose SwiftGPU---an energyefficient GPU design paradigm at NTC. SwiftGPU dynamically adjusts the degree of parallelization, and the speed of the MLDs within each stream core of the GPU. The proposed scheme achieves an average of~15% improvement in energy-efficiency over an ideal PV-free GPU, operating at the Super-Threshold regime. SwiftGPU incurs marginal area, wire-length and power overheads of 0.65%, 2.6% and 3.7%, respectively. Prabal Basu, Shamik Saha, Koushik Chakraborty, Sanghamitra Roy |
DAC | 4 |
| 2016 | Catching the flu: emerging threats from a third party power management unitabstractPower management units (PMU) have come into the spotlight with energy efficiency becoming a first order constraint in MPSoC designs. To cater to the exponential rise in power events, and to meet the demands of tight power and energy budgets, PMUs are evolving to more complex and intelligent designs. In an era defined by energy efficient computing, a malicious circuit embedded in a third party PMU can adversely affect the operation of the entire MPSoC. Rajesh J. S., Chidhambaranathan Rajamanikkam, Koushik Chakraborty, Sanghamitra Roy |
DAC | 3 |
| 2016 | Synergistic timing speculation for multi-threaded programsabstractIn this paper, we address the problem of timing speculation for multi-threaded workloads executing on a multi-core processor. Our approach is based on a new observation --- heterogeneity in path sensitization delays across different threads in multi-threaded programs. Leveraging this heterogeneity, we propose Synergistic Timing Speculation (SynTS) to jointly optimize the energy and execution time of multithreaded applications. In particular, SynTS uses a sampling based online error probability estimation technique, coupled with a polynomial time algorithm, to optimally determine the voltage, frequency and the amount of timing speculation for each thread. Our experimental evaluations, based on detailed cross-layer simulations, demonstrate that SynTS reduces energy delay product by up to 21%, compared to existing timing speculation schemes. Atif Yasin, Jeff Zhang 0001, Siddharth Garg, Sanghamitra Roy, Koushik Chakraborty |
DAC | 6 |
| 2016 | PRADA: Combating voltage noise in the NoC power supply through flow-control and routing algorithms
Prabal Basu, Rajesh J. S., Koushik Chakraborty, Sanghamitra Roy |
DATE | 3 |
| 2016 | BoostNoC: power efficient network-on-chip architecture for near threshold computingabstractWhile near threshold design space provides a promising approach towards energy-efficient computing, it is plagued by sub-optimal performance. Application characteristics and hardware non-idealities of conventional architectures (optimized for the nominal voltage) prevent us from fully leveraging the potential of NTC systems. Further, the popular approach of increasing the computational core count to compensate for the performance loss severely burdens the on-chip communication fabric with an increased communication demand. In this work, we quantitatively analyze the performance bottleneck created by a conventional NoC architecture in many-core NTC systems. To reclaim the performance lost due to a sub-optimal NoC, we propose BoostNoC — a power efficient, multi-layered network-on-chip architecture. BoostNoC improves the system performance by nearly 2× over a conventional NTC system. Further, we improve the energy efficiency by 1.4× with the use of drowsy routers. Chidhambaranathan Rajamanikkam, Rajesh J. S., Koushik Chakraborty, Sanghamitra Roy |
ICCAD | 3 |
| 2015 | Opportunistic turbo execution in NTC: exploiting the paradigm shift in performance bottlenecksabstractIn this paper, we investigate an intriguing shifting trend in performance bottlenecks for Near-Threshold Computing (NTC) processors. Our study demonstrates that the traditional memory latency bottleneck is largely superseded by the bottlenecks of Long Latency Datapaths (LLDs) within a processor core. To exploit this paradigm shift, we propose Opportunistic Turbo Execution (OTE). OTE dynamically boosts the performance of LLDs, by several factors, improving both performance and energy efficiency in an NTC core. Using a comprehensive circuit-architectural analysis, we demonstrate a 42.2% improvement in energy efficiency over a recently proposed technique, across a range of benchmarks. Dieudonne Manzi, Sanghamitra Roy, Koushik Chakraborty |
DAC | 4 |
| 2015 | Tackling voltage emergencies in NoC through timing error resilienceabstractAggressive technology scaling exacerbates the problem of voltage emergencies in emerging MPSoC systems. Network-on-Chips, the de-facto standard for connecting on-chip components in forthcoming devices play a central role in providing robust and reliable communication. In this work, we propose DrNoC (droop resilient network-on-chip)-two microarchitectural techniques to mitigate voltage emergency-induced timing errors in NoCs and preserve error-free communication throughout the network. DrNoC employs frequency downscaling and a pipeline error-recovery mechanism to reclaim corrupted flits in the router. Compared to the recently proposed NSFTR fault-tolerant technique, DrNoC offers a 27% improvement in energy-delay efficiency. Rajesh J. S., Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
ISLPED | 3 |
| 2015 | Runtime Detection of a Bandwidth Denial Attack from a Rogue Network-on-ChipabstractIn this paper, we propose a covert threat model for MPSoCs designed using 3rd party Network-on-Chips (NoC). We illustrate that a malicious NoC can disrupt the availability of on-chip resources, thereby causing large performance bottlenecks for the software running on the MPSoC platform. We then propose a runtime latency auditor that enables an MPSoC integrator to monitor the trustworthiness of the deployed NoC throughout the chip lifetime. For the proposed technique, our comprehensive cross-layer analysis indicates modest overheads of 12.73% in area, 9.844% in power and 5.4% in terms of network latency. Rajesh J. S., Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
NOCS | 3 |
| 2015 | DARP-MP: Dynamically Adaptable Resilient Pipeline Design in Multicore ProcessorsabstractIn this article, we demonstrate that the sensitized path delays in various microprocessor pipe stages exhibit intriguing temporal and spatial variations during the execution of real-world applications. To effectively exploit these delay variations, we propose dynamically adaptable resilient pipeline (DARP)—a series of runtime techniques to boost power-performance efficiency and fault tolerance in a pipelined microprocessor. DARP employs early error prediction to avoid a major portion of the timing errors. We combine DARP with the state-of-art topologically homogeneous and power-performance heterogeneous (THPH) architecture to build up a new frontier for the energy efficiency of multicore processors (DARP-MP). Using a rigorous circuit-architectural infrastructure, we demonstrate that DARP substantially improves the multicore processor performance (9.4--20%) and energy efficiency (10--28.6%) compared to state-of-the-art techniques. The energy-efficiency improvements of DARP-MP are 42% and 49.9% compared against the original THPH and another state-of-art multicore power management scheme, respectively. Sanghamitra Roy, Koushik Chakraborty |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2015 | Wearout Resilience in NoCs Through an Aging Aware Adaptive Routing AlgorithmabstractContinuous technology scaling has made aging mechanisms, such as negative bias temperature instability and electromigration primary concerns in network-on-chip (NoC) designs. In this paper, we extensively analyze the effects of these aging mechanisms on NoC routers and links. We observe a critical need of a robust aging-aware routing algorithm that not only reduces power-performance overheads caused due to aging degradation, but also minimizes the stress experienced by heavily utilized routers and links. To solve this problem, we propose an aging-aware adaptive routing algorithm and a router microarchitecture that routes the packets along the paths, which are both least congested and experience minimum aging degradation. After an extensive experimental analysis using real workloads, we observe 13% and 12.17% average overhead reduction in network latency and energy-delay product per flit, a 10.4% improvement in performance, and a 60% improvement in mean time to failure using our aging-aware routing algorithm. Dean Michael Ancajas, Kshitij Bhardwaj, Koushik Chakraborty, Sanghamitra Roy |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Fort-NoCs: Mitigating the Threat of a Compromised NoCabstractIn this paper, we uncover a novel and imminent threat to an emerging computing paradigm: MPSoCs built with 3rd party IP NoCs. We demonstrate that a compromised NoC (C-NoC) can enable a range of security attacks with an accomplice software component. To counteract these threats, we propose Fort-NoCs, a series of techniques that work together to provide protection from a C-NoC in an MPSoC. Fort-NoCs's foolproof protection disables covert backdoor activation, and reduces the chance of a successful side-channel attack by "clouding" the information obtained by an attacker. Compared to recently proposed techniques, Fort-NoCs offers a substantially better protection with lower overheads. Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
DAC | 2 |
| 2014 | DARP: Dynamically Adaptable Resilient Pipeline design in microprocessorsabstractIn this paper, we demonstrate that the sensitized path delays in various microprocessor pipe stages exhibit intriguing temporal and spatial variations during the execution of real world applications. To effectively exploit these delay variations, we propose Dynamically Adaptable Resilient Pipeline (DARP)-a series of runtime techniques to boost power performance efficiency and fault tolerance in a pipelined microprocessor. DARP employs early error prediction to avoid a major portion of the timing errors. Using a rigorous circuit-architectural infrastructure, we demonstrate substantial improvements in the performance (9.4-20%) and energy efficiency (6.4-27.9%), compared to state-of-the-art techniques. Sanghamitra Roy, Koushik Chakraborty |
DATE | 3 |
| 2014 | Dark Silicon Aware Multicore Systems: Employing Design Automation With Architectural InsightabstractThe emergence of dark silicon-a fundamental design constraint absent in past generations-brings intriguing challenges and opportunities to microprocessor design. In this brief, we demonstrate the challenges of comparing competing design styles across different technology generations in a dark silicon era. We provide a new metric to guide a dark silicon aware (DSA) system design and propose a stochastic optimization algorithm for DSA multicore system design. Our technique shows 11%-58% and 5.7-5.8 times improvement in energy efficiency for forthcoming technology generations with two multicore design styles: cores with various voltage-frequency domains and cores with heterogeneous microarchitectures. Jason M. Allred, Sanghamitra Roy, Koushik Chakraborty |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Exploring High-Throughput Computing Paradigm for Global RoutingabstractWith aggressive technology scaling, the complexity of the global routing problem is poised to grow rapidly. Solving such a large computational problem demands a high-throughput hardware platform such as modern graphics processing units (GPUs). In this paper, we explore a hybrid GPU-CPU high-throughput computing environment as a scalable alternative to the traditional CPU-based router. We introduce net-level concurrency (NLC), which is a novel parallel model for router algorithms and aims to exploit concurrency at the level of individual nets. To efficiently uncover NLC, we design a scheduler to create groups of nets that can be routed in parallel. At its core, our scheduler employs a novel algorithm to dynamically analyze data dependencies between multiple nets. We believe such an algorithm can lay the foundation for uncovering data-level parallelism in routing, which is a necessary requirement for employing high-throughput hardware. Detailed simulation results show an average of 4× speedup over NTHU-Route 2.0 with negligible loss in solution quality. To the best of our knowledge, this is the first work on utilizing GPUs for global routing. Yiding Han, Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | DMR3D: dynamic memory relocation in 3D multicore systemsabstractThree-dimensional Multicore Systems present unique opportunities for proximity driven data placement in the memory banks. Coupled with distributed memory controllers, a design trend seen in recent systems, we propose a Dynamic Memory Relocator for 3D Multicores (DMR3D) to dynamically migrate physical pages among different memory controllers. Our proposed technique avoids long interconnect delays, and increases the use of vertical interconnect, thereby substantially reducing memory access latency and communication energy. Our techniques show 30% and 25% average performance and communication energy improvement on real world applications. Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
DAC | 2 |
| 2013 | HCI-tolerant NoC router microarchitectureabstractThe trend towards massive parallel computing has necessitated the need for an On-Chip communication framework that can scale well with the increasing number of cores. At the same time, technology scaling has made transistors susceptible to a multitude of reliability issues (NBTI, HCI, TDDB). In this work, we propose an HCI-Tolerant microarchitecture for an NoC Router by manipulating the switching activity around the circuit. We find that most of the switching activity (the primary cause of HCI degradation) are only concentrated in a few parts of the circuit, severely degrading some portions more than others. Our techniques increase the lifetime of an NoC router by balancing this switching activity. Compared to an NoC without any reliability techniques, our best schemes improve the switching activity distribution, clock cycle degradation, system performance and energy delay product per flit by 19%, 26%, 11% and 17%, respectively, on an average. Dean Michael Ancajas, James McCabe Nickerson, Koushik Chakraborty, Sanghamitra Roy |
DAC | 3 |
| 2013 | Efficiently tolerating timing violations in pipelined microprocessorsabstractEarly prediction of an upcoming timing violation presents a tremendous opportunity to mask the performance overhead of tolerating these faults. In this paper, we explore several techniques for optimizing instruction scheduling in an Out-of-Order pipeline, exploiting this new perspective in robust system design. Compared to recently proposed stall based techniques for tolerating predictable timing violations, we demonstrate a massive reduction in performance overhead, while supporting correct execution in faulty environments (64--97% across different benchmarks). Koushik Chakraborty, Brennan Cozzens, Sanghamitra Roy, Dean Michael Ancajas |
DAC | 1 |
| 2013 | Proactive aging management in heterogeneous NoCs through a criticality-driven routing approachabstractThe emergence of power efficient heterogeneous NoCs presents an intriguing challenge in NoC reliability, particularly due to aging degradation. To effectively tackle this challenge, this work presents a dynamic routing algorithm that exploits the architecture level criticality of network packets while routing. Our proposed framework uses a Wearout Monitoring System (to track NBTI effect) and architecture-level criticality information to create a routing policy that restricts aging degradation with minimal impact on system level performance. Compared to the state-of-the-art BRAR (Buffered-Router Aware Routing), our best scheme achieves 38%, 53% and 29% improvements on network latency, system performance and Energy Delay Product per Flit (EDPPF) overheads, respectively. Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
DATE | 2 |
| 2013 | Long term sustainability of differentially reliable systems in the dark silicon eraabstractAs transistor miniaturization continues, providing robustness and computational correctness comes with rising power, performance, and area overhead costs. However, the diversity of software error tolerance is increasing as modern society embraces ubiquitous computing. This diversity can be exploited by differentially reliable (DR) multicore systems. The rising level of dark silicon-the portion of a chip that must remain inactive due to power budget constraints-makes such DR systems even more attractive when compared to homogeneous designs because power efficiency is improved with the increased flexibility of dynamically selecting appropriate cores for a given software workload. However, ensuring the long-term sustainability of these DR systems is a profound challenge. Asymmetric utilization of cores, differential aging degradation, and manufacturing process variation alter the relative reliability of DR system components, degrading and even eliminating the energy efficiency advantage. In this paper, we propose a feedback control based thread-to-core mapping framework to ensure longterm sustainability and extend the energy efficiency of a DR system. Over a ten-year lifespan, we analyze our approach on two DR design techniques and respectively demonstrate 14.4-16.3% and 26.1-31.0% in sustained energy-efficiency benefits, surpassing the recently proposed race-to-idle approach. Jason M. Allred, Sanghamitra Roy, Koushik Chakraborty |
ICCD | 3 |
| 2013 | A global router on GPU architectureabstractIn the modern VLSI design flow, global router is often utilized to provide fast and accurate congestion analysis for upstream processes to improve the design routability. Global routing parallelization is a good candidate to speedup its runtime performance while delivering very competitive solution quality. In this paper, we first study the cause of insufficient exploitable concurrency of the existing net level concurrency model, which has become a major bottleneck for parallelizing the emerging design problems. Then, we mitigate this limitation with a novel fine grain parallel model, with which a GPU based multi-agent global router is designed. Our experimental results indicate that the parallel model can effectively support the GPU based global router, and deliver stable solutions. The runtime comparison with NCTUgr2 has shown that upto 3.9× speedup is achieved by the GPU based router. Yiding Han, Koushik Chakraborty, Sanghamitra Roy |
ICCD | 2 |
| 2013 | Architecturally Homogeneous Power-Performance Heterogeneous Multicore SystemsabstractDynamic voltage and frequency scaling (DVFS), a widely adopted technique to ensure safe thermal characteristics while delivering superior energy efficiency, is rapidly becoming inefficient with technology scaling due to two critical factors: 1) inability to scale the supply voltage due to reliability concerns and 2) dynamic adaptations through DVFS cannot alter underlying power hungry circuit characteristics, designed for the nominal frequency. In this paper, we show that DVFS scaled circuits substantially lag in energy efficiency, by 22%-86%, compared to ground up designs for target frequency levels. We propose architecturally homogeneous power-performance heterogeneous multicore systems, a fundamentally alternate means to design energy efficient multicore systems. Using a system level computer-aided design (CAD) approach, we seamlessly integrate architecturally identical cores, designed for different voltage-frequency domains. We use a combination of standard cell library based CAD flow and full system architectural simulation to demonstrate 11%-22% improvement in energy efficiency using our design paradigm. Koushik Chakraborty, Sanghamitra Roy |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Towards graceful aging degradation in NoCs through an adaptive routing algorithmabstractContinuous technology scaling has made aging mechanisms such as Negative Bias Temperature Instability (NBTI) and electromigration primary concerns in Network-on-Chip (NoC) designs. In this paper, we model the effects of these aging mechanisms on NoC components such as routers and links using a novel reliability metric called Traffic Threshold per Epoch (TTpE). We observe a critical need of a robust aging-aware routing algorithm that not only reduces power-performance overheads caused due to aging degradation but also minimizes the stress experienced by heavily utilized routers and links. To solve this problem, we propose an aging-aware adaptive routing algorithm and a router microarchitecture that routes the packets along the paths which are both least congested and experience minimum aging stress. After an extensive experimental analysis using real workloads, we observe a 13%, 12.7% average overhead reduction in network latency and Energy-Delay-Product-Per-Flit (EDPPF) and a 10.4% improvement in performance using our aging-aware routing algorithm. Kshitij Bhardwaj, Koushik Chakraborty, Sanghamitra Roy |
DAC | 2 |
| 2012 | Predicting timing violations through instruction-level path sensitization analysisabstractIn this paper, we present a novel technique for early prediction of timing violations in high-performance pipelined microprocessors. We show that a static instruction in a microprocessor, identified by its Program Counter (PC), is an excellent predictor of an upcoming timing violation. Our analysis combines architectural data collected from real program execution with gate level logic analysis. Exploiting this PC based timing violation predictability, we propose a robust system design that predicts and tolerates timing violations seamlessly in a pipelined microprocessor. Under two different faulty environments, we show 20.9-89.8% and 14.6-80.6% average performance improvements in real programs over other state-of-the-art techniques, respectively. Sanghamitra Roy, Koushik Chakraborty |
DAC | 2 |
| 2012 | An MILP-based aging-aware routing algorithm for NoCsabstractNetwork-on-Chip (NoC) architectures have emerged as a better replacement of the traditional bus-based communication in the many-core era. However, continuous technology scaling has made aging mechanisms such as Negative Bias Temperature Instability (NBTI) and electromigration primary concerns in NoC design. In this paper1, we propose a novel system-level aging model to model the effects of asymmetric aging in NoCs. We observe a critical need of a holistic aging analysis, which when combined with power-performance optimization, poses a multi-objective design challenge. To solve this problem, we propose a Mixed Integer Linear Programming (MILP)-based aging-aware routing algorithm that optimizes the various design constraints using a multi-objective formulation. After an extensive experimental analysis using real workloads, we observe a 62.7%, 46% average overhead reduction in network latency and Energy-Delay-Product-Per-Flit (EDPPF) and a 41% improvement in Instructions Per Cycle (IPC) using our aging-aware routing algorithm. Kshitij Bhardwaj, Koushik Chakraborty, Sanghamitra Roy |
DATE | 2 |
| 2012 | DOC: Fast and accurate congestion analysis for global routingabstractThis work presents a fast and accurate congestion analysis tool at the global routing stage. It focuses on capturing the difficult-to-solve congestion in global routing designs. The proposed framework identifies the routing congestion using a novel Orthogonal Congestion Correlation (OCC) factor, which identifies the hard-to-route hot-spots. A key contribution of this work is a fast global router to minimize congestion caused by long nets and accurately reveal the distribution of hard-to-route spots due to high density short nets. The global router uses a dynamic representation of net to allow fast topology transformation. The proposed framework can evaluate the routability of a placement solution, and be utilized to aid the placer for a congestion-aware design. Yiding Han, Koushik Chakraborty, Sanghamitra Roy |
ICCD | 2 |
| 2012 | Mitigating NBTI in the physical register file through stress predictionabstractDegradation of transistor parameter values due to Negative Bias Temperature Instability (NBTI) has emerged as a major reliability problem in current and future transistor generations. NBTI Aging of SRAM cell leads to a lower noise margin, thereby increasing the failure rate. The physical register file, which consists of an array of SRAM cells, can suffer from data loss, leading to system failure. In this paper, we explore a novel approach by investigating NBTI stress and mitigation at the instruction granularity. While a wide range of NBTI stress exists in different registers, the stress induced by specific instructions is highly predictable. Using such a prediction mechanism, we propose an NBTI tolerant power efficient physical register file design. Our approach improves the noise margin in a register file by 20%, 32%, and 125% for the 45nm, 32nm, and 22nm technology nodes, respectively. Overall, we observe 14.8% power saving and a 19.8% area penalty in the register file. Saurabh Kothawade, Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
ICCD | 3 |
| 2012 | Designing for dark silicon: a methodological perspective on energy efficient systemsabstractThe emergence of dark silicon - a fundamental design constraint absent in the past generations - brings intriguing challenges and opportunities in microprocessor design. To gracefully embrace dark silicon, design methodologies must adapt themselves to identify progressive systems that can effectively exploit the growing dark silicon. We demonstrate that relying on traditional design metrics may lead to sub-optimal design choices with the rise of the dark silicon area. We provide a new metric to guide a dark silicon aware system design and propose a stochastic optimization algorithm for dark silicon aware multicore system design. Our design approach shows 7-23% benefit in upcoming technology generations. Jason M. Allred, Sanghamitra Roy, Koushik Chakraborty |
ISLPED | 3 |
| 2012 | Supporting Overcommitted Virtual Machines through Hardware Spin DetectionabstractMultiprocessor operating systems (OSs) pose several unique and conflicting challenges to System Virtual Machines (System VMs). For example, most existing system VMs resort to gang scheduling a guest OS's virtual processors (VCPUs) to avoid OS synchronization overhead. However, gang scheduling is infeasible for some application domains, and is inflexible in other domains. In an overcommitted environment, an individual guest OS has more VCPUs than available physical processors (PCPUs), precluding the use of gang scheduling. In such an environment, we demonstrate a more than two-fold increase in application runtime when transparently virtualizing a chip-multiprocessor's cores. To combat this problem, we propose a hardware technique to detect when a VCPU is wasting CPU cycles, and preempt that VCPU to run a different, more productive VCPU. Our technique can dramatically reduce cycles wasted on OS synchronization, without requiring any semantic information from the software. We then present a server consolidation case study to demonstrate the potential of more flexible scheduling policies enabled by our technique. We propose one such policy that logically partitions the CMP cores between guest VMs. This policy increases throughput by 10-25 percent for consolidated server workloads due to improved cache locality and core utilization. Koushik Chakraborty, Philip M. Wells, Gurindar S. Sohi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | Stack Aware Threshold Voltage Assignment in 3-D Multicore DesignsabstractDue to the inherent nature of heat flow in 3-D integrated circuits, stacked dies exhibit a wide range of thermal characteristics. The temperature of dies progressively increases with increasing distance from the heat sink. This heterogeneous temperature profile coupled with the strong dependence of leakage on temperature and process variation plays havoc in achieving system level energy efficiency in such systems, complicating the task of power provisioning in 3-D multicores. In this paper, we address this power provisioning challenge in 3-D ICs by advocating a novel stack aware microprocessor design paradigm, where the circuit designers are aware of the intended placement of a die in a 3-D stack. We present a concrete application of this paradigm through a stack aware threshold voltage (Vt) assignment algorithm for a 3-D multicore system, where we specifically account for: 1) the change in the role of leakage power; 2) expected operating frequency; and 3) dependency of PV induced leakage variation andVtlevels. Our stack aware scheme tunesVtassignment based on the vertical placement of the die in a 3-D stack. Detailed simulation based experiments with our proposed algorithm show 4%-19% improvement in energy efficiency for a typical multicore system organized as 3-D stacked dies. Koushik Chakraborty, Sanghamitra Roy |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Topologically homogeneous power-performance heterogeneous multicore systemsabstractDynamic Voltage and Frequency Scaling (DVFS), a widely adopted technique to ensure safe thermal characteristics while delivering superior energy efficiency, is rapidly becoming inefficient with technology scaling due to two critical factors: (a) inability to scale the supply voltage due to reliability concerns; and (b) dynamic adaptations through DVFS cannot alter underlying power hungry circuit characteristics, designed for the nominal frequency. In this paper, we show that DVFS scaled circuits substantially lag in energy efficiency, by 22-86%, compared to ground up designs for target frequency levels. We propose Topologically Homogeneous Power-Performance Heterogeneous multicore systems (THPH), a fundamentally alternate means to design energy efficient multicore systems. Using a system level CAD approach, we seamlessly integrate architecturally identical cores, designed for different voltage-frequency (VF) domains. We use a combination of standard cell library based CAD flow and full system architectural simulation to demonstrate 11-22% improvement in energy efficiency using our design paradigm. Koushik Chakraborty, Sanghamitra Roy |
DATE | 1 |
| 2011 | Exploring high throughput computing paradigm for global routingabstractWith aggressive technology scaling, the complexity of the global routing problem is poised to rapidly grow. Solving such a large computational problem demands a high throughput hardware platform such as modern Graphics Processing Units (GPU). In this work, we explore a hybrid GPU-CPU high-throughput computing environment as a scalable alternative to the traditional CPU-based router. We introduce Net Level Concurrency (NLC): a novel parallel model for router algorithms that aims to exploit concurrency at the level of individual nets. To efficiently uncover NLC, we design a Scheduler to create groups of nets that can be routed in parallel. At its core, our Scheduler employs a novel algorithm to dynamically analyze data dependencies between multiple nets. We believe such an algorithm can lay the foundation for uncovering data-level parallelism in routing: a necessary requirement for employing high throughput hardware. Detailed simulation results show an average of 4X speedup over NTHU-Route 2.0 with negligible loss in solution quality. To the best of our knowledge, this is the first work on utilizing GPUs for global routing. Yiding Han, Dean Michael Ancajas, Koushik Chakraborty, Sanghamitra Roy |
ICCAD | 3 |
| 2011 | Design and Implementation of a Throughput-Optimized GPU Floorplanning AlgorithmabstractIn this article, we propose a novel floorplanning algorithm for GPUs. Floorplanning is an inherently sequential algorithm, far from the typical programs suitable for Single-Instruction Multiple-Thread (SIMT)-style concurrency in a GPU. We propose a fundamentally different approach of exploring the floorplan solution space, where we evaluate concurrent moves on a given floorplan. We illustrate several performance optimization techniques for this algorithm in GPUs. To improve the solution quality, we present a comprehensive exploration of the design space, including various techniques to adapt the annealing approach in a GPU. Compared to the sequential algorithm, our techniques achieve 6--188X speedup for a range of MCNC and GSRC benchmarks, while delivering comparable or better solution quality. Yiding Han, Koushik Chakraborty, Sanghamitra Roy, Vilasita Kuntamukkala |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2010 | Microarchitecture aware gate sizing: A framework for circuit-architecture co-optimizationabstractModern high performance microprocessors experience substantially lower utilization in many of their structural components. To recover energy efficiency from lower utilization, system architects resort to dynamic voltage frequency scaling (DVFS). In this paper, we demonstrate that dynamic adaptations using DVFS are markedly energy inefficient than techniques that design circuits ground up for lower performance. We propose a novel microarchitecture aware gate sizing and threshold voltage assignment algorithm to mitigate this current limitation. Our technique is the first of its kind that exploits architectural slack in gate sizing, and leverages on-chip redundancy and slack. We evaluate this circuit-architectural co-optimization framework in a superscalar processor by combining standard cell based gate sizing flows with state-of-the-art architectural simulation. Our results show 17-46% improvement in the datapath energy efficiency over traditional circuit designs incorporating DVFS schemes. Sanghamitra Roy, Koushik Chakraborty |
ICCD | 2 |
| 2009 | Mixed-mode multicore reliabilityabstractFuture processors are expected to observe increasing rates of hardware faults. Using Dual-Modular Redundancy (DMR), two cores of a multicore can be loosely coupled to redundantly execute a single software thread, providing very high coverage from many difference sources of faults. This reliability, however, comes at a high price in terms of per-thread IPC and overall system throughput. We make the observation that a user may want to run both applications requiring high reliability, such as financial software, and more fault tolerant applications requiring high performance, such as media or web software, on the same machine at the same time. Yet a traditional DMR system must fully operate in redundant mode whenever any application requires high reliability. This paper proposes a Mixed-Mode Multicore (MMM), which enables most applications, including the system software, to run with high reliability in DMR mode, while applications that need high performance can avoid the penalty of DMR. Though conceptually simple, two key challenges arise: 1) care must be taken to protect reliable applications from any faults occurring to applications running in high performance mode, and 2) the desire to execute additional independent software threads for a performance application complicates the scheduling of computation to cores. After solving these issues, an MMM is shown to improve overall system performance, compared to a traditional DMR system, by approximately 2X when one reliable and one performance application are concurrently executing. Philip M. Wells, Koushik Chakraborty, Gurindar S. Sohi |
ASPLOS | 2 |
| 2008 | Adapting to intermittent faults in multicore systemsabstractFuture multicore processors will be more susceptible to a variety of hardware failures. In particular, intermittent faults, caused in part by manufacturing, thermal, and voltage variations, can cause bursts of frequent faults that last from several cycles to several seconds or more. Due to practical limitations of circuit techniques, cost-effective reliability will likely require the ability to temporarily suspend execution on a core during periods of intermittent faults. Philip M. Wells, Koushik Chakraborty, Gurindar S. Sohi |
ASPLOS | 2 |
| 2007 | Adapting to Intermittent Faults in Future Multicore Systems
Philip M. Wells, Koushik Chakraborty, Gurindar S. Sohi |
PACT | 2 |
| 2006 | Hardware support for spin management in overcommitted virtual machinesabstractMultiprocessor operating systems (OSs) pose several unique and conflicting challenges to System Virtual Machines (System VMs). For example, most existing system VMs resort to gang scheduling a guest OS's virtual processors (VCPUs) to avoid OS synchronization overhead. However, gang scheduling is infeasible for some application domains, and is inflexible in other domains.In an overcommitted environment, an individual guest OS has more VCPUs than available physical processors (PCPUs), precluding the use of gang scheduling. In such an environment, we demonstrate a more than two-fold increase in runtime when transparently virtualizing a chip-multiprocessor's cores. To combat this problem, we propose a hardware technique to detect several cases when a VCPU is not performing useful work, and suggest preempting that VCPU to run a different, more productive VCPU. Our technique can dramatically reduce cycles wasted on OS synchronization, without requiring any semantic information from the software.We then present a case study, typical of server consolidation, to demonstrate the potential of more flexible scheduling policies enabled by our technique. We propose one such policy that logically partitions the CMP cores between guest VMs. This policy increases throughput by 10-25% for consolidated server workloads due to improved cache locality and core utilization, and substantially improves performance isolation in private caches. Philip M. Wells, Koushik Chakraborty, Gurindar S. Sohi |
PACT | 2 |
| 2006 | Computation spreading: employing hardware migration to specialize CMP cores on-the-flyabstractIn canonical parallel processing, the operating system (OS) assigns a processing core to a single thread from a multithreaded server application. Since different threads from the same application often carry out similar computation, albeit at different times, we observe extensive code reuse among different processors, causing redundancy (e.g., in our server workloads, 45-65% of all instruction blocks are accessed by all processors). Moreover, largely independent fragments of computation compete for the same private resources causing destructive interference. Together, this redundancy and interference lead to poor utilization of private microarchitecture resources such as caches and branch predictors.We present Computation Spreading (CSP), which employs hardware migration to distribute a thread's dissimilar fragments of computation across the multiple processing cores of a chip multiprocessor (CMP), while grouping similar computation fragments from different threads together. This paper focuses on a specific example of CSP for OS intensive server applications: separating application level (user) computation from the OS calls it makes.When performing CSP, each core becomes temporally specialized to execute certain computation fragments, and the same core is repeatedly used for such fragments. We examine two specific thread assignment policies for CSP, and show that these policies, across four server workloads, are able to reduce instruction misses in private L2 caches by 27-58%, private L2 load misses by 0-19%, and branch mispredictions by 9-25%. Koushik Chakraborty, Philip M. Wells, Gurindar S. Sohi |
ASPLOS | 1 |