Sukarn Agarwal

dblp:191/6879 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0003-1292-3235ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 9 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 PRECIOUS: Approximate Real-Time Computing in MLC-MRAM Based Heterogeneous CMPs
abstract
Enhancing quality of service (QoS) in approximate-computing (AC) based real-time systems, without violating power limits is becoming increasingly challenging due to contradictory constraints, i.e., power consumption and time criticality, as multicore computing platforms are becoming heterogeneous. To fulfill these constraints and optimise system QoS, AC tasks should be judiciously mapped on such platforms. However, prior approaches rarely considered the problem of AC task deployment on heterogeneous platforms. Moreover, the majority of prior approaches typically neglect the runtime architectural phenomena, which can be accounted for along with the approximation tolerance of the applications to enhance the QoS. We presentPRECIOUS, a novel hybrid offline-online approach that firstschedules AC real-timetasks on aheterogeneous multicorewith an objective to maximise QoS and determines the appropriate cluster for each task constrained by a system-wide power limit, deadline, and task-dependency. At runtime,PRECIOUSintroduces novel architectural techniques for the AC tasks, where tasks are executed on a heterogeneous platform equipped withmultilevel-cell (MLC)-MRAMbased last-level cache to improve energy efficiency and performance by prudentially leveraging storage density of MLC-MRAM while ameliorating associated high write latency and write energy. Our novel block management for the MLC-MRAM cache further improves performance of the system, which we exploit opportunistically to enhance system QoS, and turn off processor cores during the dynamically generated slacks.PRECIOUS-Offlineachieves up to 76% QoS for a specific task-set, surpassing prior art, whereasPRECIOUS-Onlineenhances QoS by 9.0% by reducing cache miss-rate by 19% on a 64-core heterogeneous system without incurring any energy overhead over a conventional MRAM based cache design.
Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier
IEEE Trans. Computers3
2025 HotReRAM: A Performance-Power-Thermal Simulation Framework for ReRAM-Based Caches
abstract
This article proposes a comprehensive thermal modeling and simulation framework, HotReRAM, for resistive RAM (ReRAM)-based caches that is verified against a memristor circuit-level model. The simulation is driven by power traces based on cache accesses for detailed temperature modeling over time. HotReRAM models power at a fine-grain level and generates temperature traces for different cache regions together with detailed analyses of thermal stability, retention time and write latency. Combining HotReRAM with gem5, a full-system simulator, and NVSim, a power simulator, for ReRAM enables temporal and spatial modeling of crucial ReRAM characteristics. This integration allows designers and architects to analyze various cache characteristics within a single cache bank and address thermal-induced issues when designing ReRAM caches. Our simulation results for an 8-MiB ReRAM cache show that the spatial thermal variance can be as high as 7 K for a single cache bank, whereas the temporal thermal variance is more than 40 K. Such temperature variances impact retention time with a standard deviation of 3.9–10.2 for a set of benchmark applications, where the write latency can increase by up to 14.5%.
Shounak Chakraborty 0001, Thanasin Bunnam, Jedsada Arunruerk, Sukarn Agarwal, Shengqi Yu, Rishad A. Shafik, Magnus Själander
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 TEEMO: Temperature Aware Energy Efficient Multi-Retention STT-RAM Cache Architecture
abstract
The potential benefits of high density, non-volatility, and reduced leakage-power consumption make STT-RAM a credible successor to SRAM in caches. However, STT-RAMs experience higher write energy and latency, curtailing their potential for commercial implementation. Relaxing the retention time of STT-RAM can overcome these downsides by reducing both write latency and energy. But, significant reduction in retention time can result in premature expiry of blocks requiring frequent refreshes or write backs, which can increase the cache miss-rate and impact performance.Our proposed technique, TEEMO, divides an STT-RAM-based last level cache (LLC) set-wise into two parts with different retention times. Fetched cache blocks are placed in the corresponding sets based on the requests type (i.e., read or write). By dynamically tracking recent access intensities, blocks are prudentially managed such that write-intensive blocks are directed to LLC sets with lower retention time, whereas LLC sets with higher retention time handle read-intensive blocks. Moreover, maintaining uniform temperature across an LLC bank is crucial as temperature directly affects the performance, retention time, and lifetime of the STT-RAM cells. By employing a write-counter-based dynamic block allocation, TEEMO balances write accesses across the cache sets to maintain a uniform power density, and hence, the temperature, across the LLC bank.Our evaluation shows that TEEMO reduces the energy-delay product with up to 44.8% and improves performance by 12.5%, on average, over a baseline SRAM based ISO-area LLC. TEEMO reduces the spatial thermal variance from 8.1 °C to 4.2 °C, and reduces chip failure rate by more than 95% over prior art.
Sukarn Agarwal, Shounak Chakraborty 0001, Magnus Själander
IPDPS1
2024 ARCTIC: Approximate Real-Time Computing in a Cache-Conscious Multicore Environment
abstract
Improving result-accuracy in approximate computing (AC) based time-critical systems, without violating power constraints of the underlying circuitry, is gradually becoming challenging with the rapid progress in technology scaling. The execution span of each AC real-time tasks can be split into a couple of parts: (i) the mandatory part, execution of which offers a result of acceptable quality, followed by (ii) the optional part, which can be executed partially or completely to refine the initially obtained result in order to increase the result-accuracy, while respecting the time-constraint. In this article, we introduce a novel hybrid offline-online scheduling strategy, for AC real-time tasks. The goal of real-time scheduler of is to maximise the results-accuracy (QoS) of the task-set with opportunistic shedding of the optional part, while respecting system-wide constraints. During execution, retains exclusive copy of the private cache blocks only in the local caches in a multi-core system and no copies of these blocks are maintained at the other caches, and improves performance (i.e., reduces execution-time) by accumulating more live blocks on-chip. Combining offline scheduling with the online cache optimization improves both QoS and energy efficiency. While surpassing prior arts, our proposed strategy reduces the task-rejection-rate by up to 25%, whereas enhances QoS by 10%, with an average energy-delay-product gain of up to 9.1%, on an 8-core system.
Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Architecting Selective Refresh based Multi-Retention Cache for Heterogeneous System (ARMOUR)
abstract
The increasing use of chiplets, and the demand for high-performance yet low-power systems, will result in heterogeneous systems that combine both CPUs and accelerators (e.g., general-purpose GPUs). Chiplet based designs also enable the inclusion of emerging memory technologies, since such technologies can reside on a separate chiplet without requiring complex integration in existing high-performance process technologies. One such emerging memory technology is spin-transfer torque (STT) memory, which has the potential to replace SRAM as the last-level cache (LLC). STT-RAM has the advantage of high density, non-volatility, and reduced leakage power, but suffers from a higher write latency and energy, as compared to SRAM. However, by relaxing the retention time, the write latency and energy can be reduced at the cost of the STT-RAM becoming more volatile. The retention time and write latency/energy can be traded against each other by creating an LLC with multiple retention zones. With a multi-retention LLC, the challenge is to direct the memory accesses to the most advantageous zone, to optimize for overall performance and energy efficiency. We propose ARMOUR, a mechanism for efficient management of memory accesses to a multi-retention LLC, where based on the initial requester (CPU or GPU) the cache blocks are allocated in the high (CPU) or low (GPU) retention zone. Furthermore, blocks that are about to expire are either refreshed (CPU) or written back (GPU). In addition, ARMOUR evicts CPU blocks with an estimated short lifetime, which further improves cache performance by reducing cache pollution. Our evaluation shows that ARMOUR improves average performance by 28.9% compared to a baseline STT-RAM based LLC and reduces the energy-delay product (EDP) by 74.5% compared to an iso-area SRAM LLC.
Sukarn Agarwal, Shounak Chakraborty 0001, Magnus Själander
DAC1
2023 Compound Memory Models
abstract
Today's mobile, desktop, and server processors are heterogeneous, consisting not only of CPUs but also GPUs and other accelerators. Such heterogeneous processors are starting to expose a shared memory interface across these devices.Given that each of these individual devices typically supports a distinct instruction set architecture and a distinct memory consistency model, it is not clear what the memory consistency model of the heterogeneous machine should be. In this paper, we answer this question by formalizing "compound" memory models: we present a compositional operational model describing the resulting model when devices with distinct consistency models are fused together. We instantiate our model with the compound x86TSO/PTX model -- a CPU enforcing x86TSO and a GPU enforcing the PTX model. A key result is that the x86TSO/PTX compound model retains compiler mappings from the language-based (scoped) C memory model. This means that threads mapped to the x86TSO device can continue to use the already proven C-to-x86TSO compiler mapping, and the same for PTX.
Andres Goens, Soham Chakraborty 0001, Susmit Sarkar, Sukarn Agarwal, Nicolai Oswald, Vijay Nagarajan
Proc. ACM Program. Lang.4
2023 DELICIOUS: Deadline-Aware Approximate Computing in Cache-Conscious Multicore
abstract
Enhancing result-accuracy in approximate computing (AC) based real-time systems, without violating power constraints of the underlying hardware, is a challenging problem. Execution of such AC real-time applications can be split into two parts: (i)the mandatory part, execution of which provides a result of acceptable quality, followed by (ii)the optional part, that can be executed partially or fully to refine the initially obtained result in order to increase the result-accuracy, without violating the time-constraint. This article introducesDELICIOUS, a novel hybrid offline-onlinescheduling strategyfor AC real-time dependent tasks. By employing an efficientheuristic algorithm,DELICIOUSfirst generates a schedule for a task-set with an objective to maximize the results-accuracy, while respecting system-wide constraints. During execution,DELICIOUSthen introduces aprudential cache resizingthat reduces temperature of the adjacent cores, by generating thermal buffers at the turned off cache ways.DELICIOUSfurther trades off this thermal benefits by enhancing the processing speed of the cores for a stipulated duration, calledV/F Spiking, without violating the power budget of the core, to shorten the execution length of the tasks. This reduced runtime is exploited either to enhance result-accuracy by dynamically adjusting the optional part, or to reduce temperature by enabling sleep mode at the cores. While surpassing the prior art,DELICIOUSoffers 80% result-accuracy with its scheduling strategy, which is further enhanced by 8.3% in online, while reducing runtime peak temperature by 5.8°C on average, as shown by benchmark based evaluation on a 4-core based multicore.
Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Rahul Gangopadhyay, Magnus Själander, Klaus D. McDonald-Maier
IEEE Trans. Parallel Distributed Syst.3
2021 ABACa: Access Based Allocation on Set Wise Multi-Retention in STT-RAM Last Level Cache
abstract
Exhibition of potential advantages of high density, non-volatility, and low static power consumption makes STTRAM a credible successor to SRAM in caches. However, higher write energy and latency of the STT-RAM limit its potential towards commercial usage. Relaxation of STT-RAM’s retention time can be a viable solution to alleviate these obstacles by reducing both write time and energy. However, significant reduction in retention time might lead to premature expiry of the blocks requiring frequent refreshes or write-backs, which can incorporate unnecessary stalls along with the increased miss-rate.This paper proposes ABACa, an approach that logically bifurcates a cache set-wise for two different retention times where cache blocks are segregated upon their arrival and placed in the corresponding set, accordingly. In particular, if a block’s arrival is triggered by a read miss, the block is placed into a set with a higher retention time, called as read-set. On the other hand, the block is placed into a write-set having a lower retention time, if the block’s arrival is caused by a write miss. Our empirical analysis shows that, ABACa achieves a significant improvement of 40.75% in miss-rate and 61.35% EDP (Energy Delay Product) gain compared to baseline multi-retention STT-RAM-based and SRAM-based last level caches, respectively.
Sukarn Agarwal, Shounak Chakraborty 0001
ASAP1
2021 Improving the Performance of Hybrid Caches Using Partitioned Victim Caching
abstract
Non-Volatile Memory technologies are coming as a viable option on account of the high density and low-leakage power over the conventional SRAM counterpart. However, the increased write latency reduces their chances as a substitute for SRAM. To attenuate this problem, a hybrid STT-RAM-SRAM architecture is proposed where with large STT-RAM ways, the small SRAM ways are incorporated for handling the write operations. However, the performance gain obtained from such an architecture is not as much as expected on account of the larger miss rate caused by smaller SRAM partition. This, in turn, may limit the amount of cache capacity. This article attempts to reduce the miss penalty and improve the average memory access time by retaining the victims evicted from the hybrid cache in a smaller, fully associative SRAM structure called the victim cache. The victim cache is accessed on a miss in the primary hybrid cache. Hits in the victim cache require an exchange of the block between the main hybrid cache and the victim cache. In such cases, to effectively place the required block in the appropriate region of the main hybrid cache, we propose an access-based block placement technique. Besides, to manage the runtime load and the uneven evictions of the SRAM partition, we also present a dynamic region-based victim cache partitioning method to hold the victims dedicated to each region. Experimental evaluation on a full system simulator shows significant improvement in the performance and execution time along with a reduction in the overall miss rate. The proposed policy also increases the endurance of Hybrid Cache Architectures (HCA) by reducing writes in the STT partition.
Sukarn Agarwal, Hemangee K. Kapoor
ACM Trans. Embed. Comput. Syst.1
2020 DidaSel: dirty data based selection of VC for effective utilization of NVM buffers in on-chip interconnects
abstract
In a multi-core system, communication across cores is managed by an on-chip interconnect called Network-on-Chip (NoC). The utilization of NoC results in limitations such as high communication delay and high network power consumption. The buffers of the NoC router consume a considerable amount of leakage power. This paper attempts to reduce leakage power consumption by using Non-Volatile Memory technology-based buffers. NVM technology has the advantage of higher density and low leakage but suffers from costly write operation, and weaker write endurance. These characteristics impact on the total network power consumption, network latency, and lifetime of the router as a whole.
Khushboo Rani, Sukarn Agarwal, Hemangee K. Kapoor
ISLPED2
2020 Reuse Distance-based Victim Cache for Effective Utilisation of Hybrid Main Memory System
abstract
Hybrid main memories comprising DRAM and Non-volatile memories (NVM) are projected as potential replacements of the traditional DRAM-based memories. However, traditional cache management policies designed for improving the hit rate lack awareness of the comparative latency of read-write for NVM blocks where the write latency is more than the read latency. Therefore, developing cache management techniques that reduce costly write-backs of the NVM blocks, yet maintain a fair hit rate in the cache, is of paramount importance. We propose two techniques based on the use of a small victim cache associated with the last-level cache that helps in retaining on the chip critical DRAM and NVM blocks. Victim cache being a scarce resource, we intend to keep only performance-critical blocks in the victim cache by exploiting the idea of reuse distance. The first technique, Victim Cache Replacement Policy, works on the replacement policy of the victim cache by preferential eviction of DRAM blocks over NVM blocks. However, the second technique, Prioritized Partitioning of victim cache, logically partitions the victim cache, giving a smaller share to the DRAM blocks and a relatively larger share to the NVM blocks. Experimental evaluation on full-system simulator shows significant improvement in system performance and reduction in the number of write-backs to the NVM partition of the main memory compared to the baseline and existing technique. Additionally, NVM reads and DRAM miss rate are also improved, leading to further performance enhancement.
Arijit Nath, Sukarn Agarwal, Hemangee K. Kapoor
ACM Trans. Design Autom. Electr. Syst.2
2019 Enhancing the Lifetime of Non-Volatile Caches by Exploiting Module-Wise Write Restriction
abstract
The emerging Non-Volatile Memory (NVM) technologies offer a good combination of high density and near-zero leakage power, becoming the strongest candidate in the memory hierarchy including caches. However, the weak write endurance of these memories creates a bottleneck towards their employment in the cache hierarchy. This weak endurance shows its effects due to the write variations introduced by the applications and the existing cache management policies. Such variations result in early breakdown of the NVM cells reducing the effective lifetime of the NVM memory component. This paper proposes a technique to mitigate intra-set write variation, i.e. write variations occurring within the cache set. Our policy divides the cache logically into multiple equal-sized modules. During execution, the writes are distributed uniformly across different ways of the different modules within the set. Experimental results using full system simulation show that the proposed technique reduces the intra-set write variation significantly over the baseline and the existing techniques.
Sukarn Agarwal, Hemangee K. Kapoor
ACM Great Lakes Symposium on VLSI1
2019 Towards Optimizing Refresh Energy in embedded-DRAM Caches using Private Blocks
abstract
In recent years, the increased working set size of applications craves for more memory demand in terms of large size Last Level Caches (LLC). To fulfill this, embedded DRAM (eDRAM) caches have been considered as one of the best alternatives over conventional SRAM caches. eDRAM has a property of low leakage and provides more capacity in the same area footprint of SRAM. However, its retention period consumes significant refresh energy in the periodic refresh. In this paper, we present an approach to minimize the total energy spent on refreshes by considering the presence of private blocks in the LLC. Our approach restricts refreshing of those blocks that are loaded exclusively from the main memory on an LLC miss. Experimental result using full system simulation show 55% reduction in the total number of refreshes compared to baseline policy; and 62% reduction in total power consumption over SRAM.
Sheel Sindhu Manohar, Sukarn Agarwal, Hemangee K. Kapoor
ACM Great Lakes Symposium on VLSI2
2019 Improving the Lifetime of Non-Volatile Cache by Write Restriction
abstract
The attractive features such as low static power and high density exhibited by the Non-Volatile Memory (NVM) technologies makes them a promising candidate in the memory hierarchy, including caches. However, the limited write endurance with the write variations governed by the access patterns and the applied replacement policies reduce the chance of NVMs as a successor of SRAM. These write variations are of concern as they not only breakdown the NVM cells but also reduce the effective lifetime. This paper proposes efficient techniques to mitigate the intra-set write variation to improve the lifetime of the NVM cache. Our first two techniques partition the cache into windows of equal size and distribute the writes uniformly across the cache set by employing the window as write-restricted or read-only. The selection of the window in these techniques is by rotation or with the help of counters. In our third technique, different cache ways are employed as a write-restricted over the period of execution to distribute the writes uniformly. Experimental results using full system simulation show the significant reduction in intra-set write variation along with improvement in the cache lifetime.
Sukarn Agarwal, Hemangee K. Kapoor
IEEE Trans. Computers1
2018 Reuse-Distance-Aware Write-Intensity Prediction of Dataless Entries for Energy-Efficient Hybrid Caches
Sukarn Agarwal, Hemangee K. Kapoor
IEEE Trans. Very Large Scale Integr. Syst.1
2017 Targeting inter set write variation to improve the lifetime of non-volatile cache using fellow sets
abstract
High density and low static power exhibited by nonvolatile technologies (NVM) have made them popular candidates in the memory hierarchy, including caches. Writes within a cache set are governed by the access pattern as well as replacement policies, leading to a large write variation. This variation is of concern as it leads to early breakdown of the NVM cells due to large writes thus reducing the effective lifetime. This paper presents a technique to improve the lifetime of non-volatile caches by reducing the inter-set write variation. Our policy partitions the cache sets into groups called fellow groups. Every set has two logical parts: Normal and Reserved. Sets within a fellow group can use the reserved parts from their fellow sets to distribute the writes uniformly. Experimental results using full system simulation show that the proposed technique shows significant reduction in inter-set write variation over the baseline and existing technique.
Sukarn Agarwal, Hemangee K. Kapoor
VLSI-SoC1
2016 Restricting writes for energy-efficient hybrid cache in multi-core architectures
abstract
Emerging non-volatile memory technology Spin Transfer Torque Random Access Memory (STT-RAM) is a good candidate for the Last Level Cache (LLC) on account of high density, good scalability and low power consumption. However, expensive write operation reduces their chances as a replacement of SRAM. To handle these expensive write operations, an STT-RAM/SRAM hybrid cache architecture is proposed that reduces the number of writes and energy consumption of the STT-RAM region in the LLC by considering the existence of private blocks. Our approach allocates dataless entries for such kind of blocks when they are loaded in the LLC on a miss. We make changes in the conventional MESI protocol by adding new states to deal with the dataless entries. Experimental results using full system simulator shows 73% savings in write operations and 20% energy savings compared to an existing policy.
Sukarn Agarwal, Hemangee K. Kapoor
VLSI-SoC1