EDBT 2026 Demo / reviewers in the wild / expert
Shounak Chakraborty 0001
dblp:156/8565-1
· DBLP profile ↗
21ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0003-1679-6210ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 9 first-author · 17 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VLIM: Verified Loop Interchange for Optimised Matrix MultiplicationabstractLoop optimisations are essential for achieving high performance in modern computing, particularly for memory-intensive operations. However, while unverified optimisers achieve impressive speedups, their manual application is error-prone and challenging to verify, making them risky in high-assurance computing platforms. This paper introduces VLIM, a novel rewrite algebra, to overcome these difficulties, enabling the development and automatic verification of loop transformations within the Capla programming language, a formally defined front-end for the Compcert verified compiler. Our framework allows compiler developers to define rewrite rules, with correctness proofs automatically derived through rewrite composition, ensuring semantic preservation during optimisation. We demonstrate the effectiveness of our approach, VLIM, by implementing a loop interchange optimisation and evaluating its impact on matrix multiplication performance. Empirical analyses show significant performance improvements: for a 1000 × 1000 matrix, loop interchange using VLIM reduced runtime by 36.6% and 74.6% when compiled with Compcert and Clang, respectively. This work advances the state-of-the-art in verified compilation, offering a promising direction for developing high-performance, formally verified software. Oliver Turner, Shounak Chakraborty 0001 |
DATE | 2 |
| 2025 | PRECIOUS: Approximate Real-Time Computing in MLC-MRAM Based Heterogeneous CMPsabstractEnhancing quality of service (QoS) in approximate-computing (AC) based real-time systems, without violating power limits is becoming increasingly challenging due to contradictory constraints, i.e., power consumption and time criticality, as multicore computing platforms are becoming heterogeneous. To fulfill these constraints and optimise system QoS, AC tasks should be judiciously mapped on such platforms. However, prior approaches rarely considered the problem of AC task deployment on heterogeneous platforms. Moreover, the majority of prior approaches typically neglect the runtime architectural phenomena, which can be accounted for along with the approximation tolerance of the applications to enhance the QoS. We presentPRECIOUS, a novel hybrid offline-online approach that firstschedules AC real-timetasks on aheterogeneous multicorewith an objective to maximise QoS and determines the appropriate cluster for each task constrained by a system-wide power limit, deadline, and task-dependency. At runtime,PRECIOUSintroduces novel architectural techniques for the AC tasks, where tasks are executed on a heterogeneous platform equipped withmultilevel-cell (MLC)-MRAMbased last-level cache to improve energy efficiency and performance by prudentially leveraging storage density of MLC-MRAM while ameliorating associated high write latency and write energy. Our novel block management for the MLC-MRAM cache further improves performance of the system, which we exploit opportunistically to enhance system QoS, and turn off processor cores during the dynamically generated slacks.PRECIOUS-Offlineachieves up to 76% QoS for a specific task-set, surpassing prior art, whereasPRECIOUS-Onlineenhances QoS by 9.0% by reducing cache miss-rate by 19% on a 64-core heterogeneous system without incurring any energy overhead over a conventional MRAM based cache design. Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier |
IEEE Trans. Computers | 2 |
| 2025 | HotReRAM: A Performance-Power-Thermal Simulation Framework for ReRAM-Based CachesabstractThis article proposes a comprehensive thermal modeling and simulation framework, HotReRAM, for resistive RAM (ReRAM)-based caches that is verified against a memristor circuit-level model. The simulation is driven by power traces based on cache accesses for detailed temperature modeling over time. HotReRAM models power at a fine-grain level and generates temperature traces for different cache regions together with detailed analyses of thermal stability, retention time and write latency. Combining HotReRAM with gem5, a full-system simulator, and NVSim, a power simulator, for ReRAM enables temporal and spatial modeling of crucial ReRAM characteristics. This integration allows designers and architects to analyze various cache characteristics within a single cache bank and address thermal-induced issues when designing ReRAM caches. Our simulation results for an 8-MiB ReRAM cache show that the spatial thermal variance can be as high as 7 K for a single cache bank, whereas the temporal thermal variance is more than 40 K. Such temperature variances impact retention time with a standard deviation of 3.9–10.2 for a set of benchmark applications, where the write latency can increase by up to 14.5%. Shounak Chakraborty 0001, Thanasin Bunnam, Jedsada Arunruerk, Sukarn Agarwal, Shengqi Yu, Rishad A. Shafik, Magnus Själander |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | MAFin: Maximizing Accuracy in FinFET based Approximated Real-Time ComputingabstractWe propose MAFin that exploits the unique temperature effect inversion (TEI) property of a FinFET based multicore platform, where processing speed increases with temperature, in the context of approximate real-time computing. In approximate real-time computing platforms, the execution of each task can be divided into two parts: (i) the mandatory part, execution of which provides a result of acceptable quality, followed by (ii) the optional part, that can be executed partially or fully to refine the initially obtained result in order to increase the result-accuracy (QoS) without violating deadlines. With an objective to maximize the QoS for a FinFET based multicore system, MAFin, our proposed real-time scheduler first derives a task-to-core allocation, while respecting system-wide constraints and prepares a schedule. During execution, MAFin further increases the achieved QoS, while balancing the performance and temperature on-the-fly by incorporating a prudential temperature cognizant frequency management mechanism and guarantees imposed constraints. Specifically, MAFin exploits the TEI property of FinFET based processors, where processor-speed is enhanced at the increased temperature, to reduce the execution time of the individual tasks. This reduced execution-time is then traded off either to enhance QoS by executing more from the tasks' optional parts or to improve energy efficiency by turning off the core. While surpassing prior art, MAFin achieves 70% QoS, which is further enhanced by 8.3% in online, with a maximum EDP gain of up to 12%, based on benchmark based evaluation on a 4-core based system. Shounak Chakraborty 0001, Sangeet Saha, Magnus Själander, Klaus D. McDonald-Maier |
DAC | 1 |
| 2024 | TEEMO: Temperature Aware Energy Efficient Multi-Retention STT-RAM Cache ArchitectureabstractThe potential benefits of high density, non-volatility, and reduced leakage-power consumption make STT-RAM a credible successor to SRAM in caches. However, STT-RAMs experience higher write energy and latency, curtailing their potential for commercial implementation. Relaxing the retention time of STT-RAM can overcome these downsides by reducing both write latency and energy. But, significant reduction in retention time can result in premature expiry of blocks requiring frequent refreshes or write backs, which can increase the cache miss-rate and impact performance.Our proposed technique, TEEMO, divides an STT-RAM-based last level cache (LLC) set-wise into two parts with different retention times. Fetched cache blocks are placed in the corresponding sets based on the requests type (i.e., read or write). By dynamically tracking recent access intensities, blocks are prudentially managed such that write-intensive blocks are directed to LLC sets with lower retention time, whereas LLC sets with higher retention time handle read-intensive blocks. Moreover, maintaining uniform temperature across an LLC bank is crucial as temperature directly affects the performance, retention time, and lifetime of the STT-RAM cells. By employing a write-counter-based dynamic block allocation, TEEMO balances write accesses across the cache sets to maintain a uniform power density, and hence, the temperature, across the LLC bank.Our evaluation shows that TEEMO reduces the energy-delay product with up to 44.8% and improves performance by 12.5%, on average, over a baseline SRAM based ISO-area LLC. TEEMO reduces the spatial thermal variance from 8.1 °C to 4.2 °C, and reduces chip failure rate by more than 95% over prior art. Sukarn Agarwal, Shounak Chakraborty 0001, Magnus Själander |
IPDPS | 2 |
| 2024 | ARCTIC: Approximate Real-Time Computing in a Cache-Conscious Multicore EnvironmentabstractImproving result-accuracy in approximate computing (AC) based time-critical systems, without violating power constraints of the underlying circuitry, is gradually becoming challenging with the rapid progress in technology scaling. The execution span of each AC real-time tasks can be split into a couple of parts: (i) the mandatory part, execution of which offers a result of acceptable quality, followed by (ii) the optional part, which can be executed partially or completely to refine the initially obtained result in order to increase the result-accuracy, while respecting the time-constraint. In this article, we introduce a novel hybrid offline-online scheduling strategy, for AC real-time tasks. The goal of real-time scheduler of is to maximise the results-accuracy (QoS) of the task-set with opportunistic shedding of the optional part, while respecting system-wide constraints. During execution, retains exclusive copy of the private cache blocks only in the local caches in a multi-core system and no copies of these blocks are maintained at the other caches, and improves performance (i.e., reduces execution-time) by accumulating more live blocks on-chip. Combining offline scheduling with the online cache optimization improves both QoS and energy efficiency. While surpassing prior arts, our proposed strategy reduces the task-rejection-rate by up to 25%, whereas enhances QoS by 10%, with an average energy-delay-product gain of up to 9.1%, on an 8-core system. Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | TREAFET: Temperature-Aware Real-Time Task Scheduling for FinFET based MulticoresabstractThe recent shift in the VLSI industry from conventional MOSFET to FinFET for designing contemporary chip-multiprocessor (CMP) has noticeably improved hardware platforms’ computing capabilities, but at the cost of several thermal issues. Unlike the conventional MOSFET, FinFET devices experience a significant increase in circuit speed at a higher temperature, called temperature effect inversion (TEI), but higher temperature can also curtail the circuit lifetime due to self-heating effects (SHEs). These fundamental thermal properties of FinFET introduced a new challenge for scheduling time-critical tasks on FinFET-based multicores that how to exploit TEI towards improving performance while combating SHEs. In this work,TREAFET, a temperature-aware real-time scheduler, attempts to exploit the TEI feature of FinFET-based multicores in a time-critical computing paradigm. At first, the overall progress of individual tasks is monitored, tasks are allocated to the cores, and finally, a schedule is prepared. By considering the thermal profiles of the individual tasks and the current thermal status of the cores, hot tasks are assigned to the cold cores and vice-versa. Finally, the performance and temperature are balanced on-the-fly by incorporating a prudential voltage scaling towards exploiting TEI while guaranteeing the deadline and thermal safety. Moreover,TREAFETstimulates the average runtime frequency by employing an opportunistic energy-adaptive voltage spiking mechanism, in which energy saving during memory stalls at the cores is traded off during the time slice having the spiked voltage. Simulation results claimTREAFETmaintains a safe and stable thermal status (peak temperature below 80 °C) and improves frequency up to 17% over the assigned value, which ensures legitimate time-critical performance for a variety of workloads while surpassing a state-of-the-art technique. The stimulated frequency inTREAFETalso finishes the tasks early, thus providing opportunities to save energy by power gating the cores, and achieves a 24% energy delay product (EDP) gain on average. Shounak Chakraborty 0001, Yanshul Sharma, Sanjay Moulik |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | Architecting Selective Refresh based Multi-Retention Cache for Heterogeneous System (ARMOUR)abstractThe increasing use of chiplets, and the demand for high-performance yet low-power systems, will result in heterogeneous systems that combine both CPUs and accelerators (e.g., general-purpose GPUs). Chiplet based designs also enable the inclusion of emerging memory technologies, since such technologies can reside on a separate chiplet without requiring complex integration in existing high-performance process technologies. One such emerging memory technology is spin-transfer torque (STT) memory, which has the potential to replace SRAM as the last-level cache (LLC). STT-RAM has the advantage of high density, non-volatility, and reduced leakage power, but suffers from a higher write latency and energy, as compared to SRAM. However, by relaxing the retention time, the write latency and energy can be reduced at the cost of the STT-RAM becoming more volatile. The retention time and write latency/energy can be traded against each other by creating an LLC with multiple retention zones. With a multi-retention LLC, the challenge is to direct the memory accesses to the most advantageous zone, to optimize for overall performance and energy efficiency. We propose ARMOUR, a mechanism for efficient management of memory accesses to a multi-retention LLC, where based on the initial requester (CPU or GPU) the cache blocks are allocated in the high (CPU) or low (GPU) retention zone. Furthermore, blocks that are about to expire are either refreshed (CPU) or written back (GPU). In addition, ARMOUR evicts CPU blocks with an estimated short lifetime, which further improves cache performance by reducing cache pollution. Our evaluation shows that ARMOUR improves average performance by 28.9% compared to a baseline STT-RAM based LLC and reduces the energy-delay product (EDP) by 74.5% compared to an iso-area SRAM LLC. Sukarn Agarwal, Shounak Chakraborty 0001, Magnus Själander |
DAC | 2 |
| 2023 | DELICIOUS: Deadline-Aware Approximate Computing in Cache-Conscious MulticoreabstractEnhancing result-accuracy in approximate computing (AC) based real-time systems, without violating power constraints of the underlying hardware, is a challenging problem. Execution of such AC real-time applications can be split into two parts: (i)the mandatory part, execution of which provides a result of acceptable quality, followed by (ii)the optional part, that can be executed partially or fully to refine the initially obtained result in order to increase the result-accuracy, without violating the time-constraint. This article introducesDELICIOUS, a novel hybrid offline-onlinescheduling strategyfor AC real-time dependent tasks. By employing an efficientheuristic algorithm,DELICIOUSfirst generates a schedule for a task-set with an objective to maximize the results-accuracy, while respecting system-wide constraints. During execution,DELICIOUSthen introduces aprudential cache resizingthat reduces temperature of the adjacent cores, by generating thermal buffers at the turned off cache ways.DELICIOUSfurther trades off this thermal benefits by enhancing the processing speed of the cores for a stipulated duration, calledV/F Spiking, without violating the power budget of the core, to shorten the execution length of the tasks. This reduced runtime is exploited either to enhance result-accuracy by dynamically adjusting the optional part, or to reduce temperature by enabling sleep mode at the cores. While surpassing the prior art,DELICIOUSoffers 80% result-accuracy with its scheduling strategy, which is further enhanced by 8.3% in online, while reducing runtime peak temperature by 5.8°C on average, as shown by benchmark based evaluation on a 4-core based multicore. Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Rahul Gangopadhyay, Magnus Själander, Klaus D. McDonald-Maier |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | STIFF: thermally safe temperature effect inversion aware FinFET based multi-coreabstractFinFET, a non-planar device, has become the prevalent choice for chip-multiprocessor (CMP) designs due to its lower leakage and improved scalability as compared to planar CMOS devices. FinFETs are fundamentally different from conventional CMOS circuits in terms of circuit-delay vs. temperature, i.e., circuit-delay decreases in FinFET at higher temperature even in the super threshold supply-voltage regime. Such characteristic of FinFET is known as temperature effect inversion (TEI). But, a drastic increase in channel temperature may lead to an increase in leakage consumption and may accelerate the circuit aging process due to the self-heating effect (SHE). This paper introduces STIFF, which balances the upsides of TEI against the potential hazardous SHE in a FinFET based CMP. Basically, STIFF exploits online performance statistics to determine the thermal intensity of cores and local caches, and scales the supply-voltage prudentially to maintain a stable core-frequency and local-cache performance on-the-fly by exploiting TEI, while reducing the SHE. Our simulation results show that, STIFF is able to maintain a stable frequency of 3.7GHz of the cores with a small standard deviation of 0.23, while maintaining a safe temperature during execution, and it outperforms a state-of-the-art DVFS technique for the FinFET based cores. STIFF also maintains a stable access time at the local L1 caches, while ensuring thermal safety by introducing a cache access cognizant scaling of the supply voltage of the individual L1 cache-banks without any noticeable performance-loss. Shounak Chakraborty 0001, Vassos Soteriou, Magnus Själander |
CF | 1 |
| 2022 | RESTORE: Real-Time Task Scheduling on a Temperature Aware FinFET based MulticoreabstractIn this work, we propose RESTORE that exploits the unique thermal feature of FinFET based multicore platforms, where processing speed increases with temperature, in the context of time-criticality to meet other design constraints of real-time systems. RESTORE is a temperature aware real-time scheduler for FinFET based multicore system that first derives a task-to-core allocation, and prepares a schedule. Next, it balances the performance and temperature on the fly by incorporating a prudential temperature cognizant voltage/frequency scaling while guaranteeing task deadlines. Simulation results claim, RESTORE is able to maintain a safe and stable thermal status (peak temperature below 80°C), hence the frequency (3.7 GHz on an average), that ensures legitimate time-critical performance for a variety of workloads while surpassing state-of-the-arts. Yanshul Sharma, Sanjay Moulik, Shounak Chakraborty 0001 |
DATE | 3 |
| 2022 | ACCURATE: Accuracy Maximization for Real-Time Multicore Systems With Energy-Efficient Way-Sharing CachesabstractImproving result accuracy in approximate computing (AC)-based real-time applications without violating deadlines has recently become an active research domain. Execution time of AC real-time tasks can individually be separated into: execution of the mandatory part to obtain a result of acceptable quality, followed by a partial/complete execution of the optional part to improve the result accuracy of the initial result within a given deadline. However, obtaining higher result accuracy at the cost of enhanced execution time may lead to deadline violation, along with higher energy usage. We present ACCURATE, a novel hybrid offline–online approximate real-time scheduling approach that first schedules AC-based tasks on multicore with an objective to maximize result accuracy and determines operational processing speeds for each task constrained by system-wide power limit, deadline, and task dependency. At runtime, by employing a way-sharing technique (WH_LLC) at the last level cache (LLC), ACCURATE improves performance, which is further leveraged, to enhance result accuracy by executing more from the optional part and to improve the energy efficiency of the cache by turning off a controlled number of cache ways. ACCURATE also exploits the slacks either to improve the result accuracy of the tasks or to enhance the energy efficiency of the underlying system, or both. ACCURATE achieves 85% QoS with 36% average reduction in cache leakage consumption with a 24% average gain in energy-delay product (EDP) for a 4-core-based chip multiprocessor (CMP) with 6.4% average improvement in performance. Sangeet Saha, Shounak Chakraborty 0001, Xiaojun Zhai, Shoaib Ehsan, Klaus D. McDonald-Maier |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | ETA-HP: an energy and temperature-aware real-time scheduler for heterogeneous platforms
Yanshul Sharma, Shounak Chakraborty 0001, Sanjay Moulik |
J. Supercomput. | 2 |
| 2021 | ABACa: Access Based Allocation on Set Wise Multi-Retention in STT-RAM Last Level CacheabstractExhibition of potential advantages of high density, non-volatility, and low static power consumption makes STTRAM a credible successor to SRAM in caches. However, higher write energy and latency of the STT-RAM limit its potential towards commercial usage. Relaxation of STT-RAM’s retention time can be a viable solution to alleviate these obstacles by reducing both write time and energy. However, significant reduction in retention time might lead to premature expiry of the blocks requiring frequent refreshes or write-backs, which can incorporate unnecessary stalls along with the increased miss-rate.This paper proposes ABACa, an approach that logically bifurcates a cache set-wise for two different retention times where cache blocks are segregated upon their arrival and placed in the corresponding set, accordingly. In particular, if a block’s arrival is triggered by a read miss, the block is placed into a set with a higher retention time, called as read-set. On the other hand, the block is placed into a write-set having a lower retention time, if the block’s arrival is caused by a write miss. Our empirical analysis shows that, ABACa achieves a significant improvement of 40.75% in miss-rate and 61.35% EDP (Energy Delay Product) gain compared to baseline multi-retention STT-RAM-based and SRAM-based last level caches, respectively. Sukarn Agarwal, Shounak Chakraborty 0001 |
ASAP | 2 |
| 2021 | SEAMERS: A Semi-partitioned Energy-Aware scheduler for heterogeneous MulticorEReal-time Systems
Sanjay Moulik, Zinea Das, Rajesh Devaraj, Shounak Chakraborty 0001 |
J. Syst. Archit. | 4 |
| 2021 | WaFFLe: Gated Cache-Ways with Per-Core Fine-Grained DVFS for Reduced On-Chip Temperature and Leakage ConsumptionabstractManaging thermal imbalance in contemporary chip multi-processors (CMPs) is crucial in assuring functional correctness of modern mobile as well as server systems. Localized regions with high activity, e.g., register files, ALUs, FPUs, and so on, experience higher temperatures than the average across the chip and are commonly referred to as hotspots. Hotspots affect functional correctness of the underlying circuitry and a noticeable increase in leakage power, which in turn generates heat in a self-reinforced cycle. Techniques that reduce the severity of or completely eliminate hotspots can maintain functional correctness along with improving performance of CMPs. Conventional dynamic thermal management targets the cores to reduce hotspots but often ignores caches, which are known for their high leakage power consumption. This article presents WaFFLe , an approach that targets the leakage power of the last-level cache (LLC) and hotspots occurring at the cores. WaFFLe turns off LLC-ways to reduce leakage power and to generate on-chip thermal buffers. In addition, fine-grained DVFS is applied during long LLC miss induced stalls to reduce core temperature. Our results show that WaFFLe reduces peak and average temperature of a 16-core based homogeneous tiled CMP with up to 8.4 ֯ C and 6.2 ֯ C, respectively, with an average performance degradation of only 2.5 %. We also show that WaFFLe outperforms a state-of-the-art cache-based technique and a greedy DVFS policy. Shounak Chakraborty 0001, Magnus Själander |
ACM Trans. Archit. Code Optim. | 1 |
| 2021 | Prepare: Power-Aware Approximate Real-time Task Scheduling for Energy-Adaptive QoS MaximizationabstractAchieving high result-accuracy in approximate computing (AC) based real-time applications without violating power constraints of the underlying hardware is a challenging problem. Execution of such AC real-time tasks can be divided into the execution of the mandatory part to obtain a result of acceptable quality, followed by a partial/complete execution of the optional part to improve accuracy of the initially obtained result within the given time-limit. However, enhancing result-accuracy at the cost of increased execution length might lead to deadline violations with higher energy usage. We propose Prepare , a novel hybrid offline-online approximate real-time task-scheduling approach, that first schedules AC-based tasks and determines operational processing speeds for each individual task constrained by system-wide power limit, deadline, and task-dependency. At runtime, by employing fine-grained DVFS, the energy-adaptive processing speed governing mechanism of Prepare reduces processing speed during each last level cache miss induced stall and scales up the processing speed once the stall finishes to a higher value than the predetermined one. To ensure on-chip thermal safety, this higher processing speed is maintained only for a short time-span after each stall, however, this reduces execution times of the individual task and generates slacks. Prepare exploits the slacks either to enhance result-accuracy of the tasks, or to improve thermal and energy efficiency of the underlying hardware, or both. With a 70 - 80% workload, Prepare offers 75% result-accuracy with its constrained scheduling, which is enhanced by 5.3% for our benchmark based evaluation of the online energy-adaptive mechanism on a 4-core based homogeneous chip multi-processor, while meeting the deadline constraint. Overall, while maintaining runtime thermal safety, Prepare reduces peak temperature by up to 8.6 °C for our baseline system. Our empirical evaluation shows that constrained scheduling of Prepare outperforms a state-of-the-art scheduling policy, whereas our runtime energy-adaptive mechanism surpasses two current DVFS based thermal management techniques. Shounak Chakraborty 0001, Sangeet Saha, Magnus Själander, Klaus D. McDonald-Maier |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2019 | Exploring the Role of Large Centralised Caches in Thermal Efficient Chip DesignabstractIn the era of short channel length, Dynamic Thermal Management (DTM) has become a challenging task for the architects and designers engineering modern Chip Multi-Processors (CMPs). Ever-increasing demand of processing power along with the developed integration technology produces CMPs with high power density, which in turn increases effective chip temperature. This increased temperature leads to increase in the reliability issues for the chip-circuitry with significant increment in leakage power consumption. Recent DTM techniques apply DVFS or Task Migration to reduce temperature at the cores, the hottest on-chip components, but often ignore the on-chip hot caches. To commensurate the high data demand of these cores, most of the modern CMPs are equipped with large multi-level on-chip caches, out of which on-chip Last Level Caches (LLCs) occupy the largest on-chip area. These LLCs are accounted for their significantly high leakage power consumption that can also potentially generate on-chip hotspots at the LLCs similar to the cores. As power consumption constructs the backbone of heat dissipation, hence, this work dynamically shrinks cache size while maintaining performance constraint to reduce LLC leakage, primarily. These turned-off cache portions further work as on-chip thermal buffers for reducing average and peak temperature of the CMP without affecting the computation. Simulation results claim that, at a minimal penalty on the performance, proposed cache-based thermal management having 8MB centralised multi-banked shared LLC gives around 5°C reduction in peak and average chip temperature, which are comparable with a Greedy DVFS policy. Shounak Chakraborty 0001, Hemangee K. Kapoor |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2018 | Utility Aware Snoozy Caches for Energy Efficient Chip Multi-ProcessorsabstractHeavy leakage power consumption of on-chip last level caches (LLCs) has become the primary obstacle for architecting chip multi-processors (CMPs) in recent times. As leakage power has a direct relationship with the supply voltage, hence, periodic access profile based dynamic voltage scaling (DVS) in the LLC banks can be a promising option towards reducing this heavy cache leakage. A plethora of prior attempts have reduced this by anticipating working set size (WSS) of the applications and eventually putting some portions of the cache banks in low power mode. This proposed work aims to reduce leakage by putting a whole LLC bank into a low power (snoozy) mode through exploiting DVS at cache banks having minimal usages. Additionally, the resulting performance impacts of the low power snoozy mode are alleviated further by putting some snoozy banks in active mode on-demand. Experimental evaluations using full system simulation on a multi-banked 2MB 8-way set associative L2 cache show 10% more leakage savings on an average over a prior drowsy technique. Ashwini A. Kulkarni, Shounak Chakraborty 0001, Shrinivas P. Mahajan, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Analysing the Role of Last Level Caches in Controlling Chip TemperatureabstractDynamic Thermal Management (DTM) has become a major concern for the chip-designers, as it becomes a challenging task in recent power densed high performance Chip Multi-Processors (CMPs), due to integration of more on-chip components to meet ever increasing demand of processing power. The increased chip temperature incorporates severe circuit errors along with significant increment in leakage power consumption. Traditional DTM techniques apply DVFS or task migration to reduce core temperature, as cores are considered as the hottest on-chip components. Additionally, to commensurate high data demand of these high performance cores, large on-chip Last Level Caches (LLCs) are attached, which are the principal contributors to the on-chip leakage power consumption and occupy the largest on-chip area. As power consumption reduction plays the pivotal role in temperature reduction, hence, this work dynamically shrinks the cache size not only to reduce leakage power consumption, but also, to create on-chip thermal buffers for reducing average chip temperature by exploiting the heat transfer physics. Cache resizing decisions are taken based upon the generated cache hotspots and/or the access patterns, during process execution. Simulation results of the proposed thermal management method are compared with an existing DVFS based method (at cores) and a prior drowsy cache based technique to show its effectiveness. Shounak Chakraborty 0001, Hemangee K. Kapoor |
IEEE Trans. Sustain. Comput. | 1 |
| 2016 | Static energy reduction by performance linked dynamic cache resizingabstractThe increased power density with short channel effect in modern transistors significantly increases the leakage energy consumptions of on-chip Last Level Caches (LLCs) in recent Chip Multi-Processors(CMPs). Performance linked dynamic shrinking in the LLC size is a promising option for reducing cache leakage. Prior works attempt to reduce the cache leakage by predicting Working Set Size(WSS) of the applications and by putting some cache portions in low power mode. This paper aims to reduce leakage energy by using a combination of cache bank shutdown and way shutdown. The banks with minimal usages are candidates for shutdown. In banks with average usages, some ways are turned off to save leakage. To mitigate the impact of smaller set-size, we apply dynamic associativity management technique. Experimental evaluation using full system simulation on a 4MB 8-way set associative L2 cache gives 70% average savings in static energy with 35% average savings in EDP. In case application's cache demand increases we can turn-on some ways to maintain performance. Shounak Chakraborty 0001, Hemangee K. Kapoor |
VLSI-SoC | 1 |