VLDB 2026 Research / reviewers in the wild / expert
Hemangee K. Kapoor
dblp:75/2282 · also Hemangee Kalpesh Kapoor
· DBLP profile ↗
59ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0002-9376-7686ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 52 · 3 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Theory of computation · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpALEn: Sparsity Aware Load Balancing Inference Engine for Neural NetworkabstractIn the field of computer science, Convolutional Neural Network (CNN) algorithms are a crucial tool that contributes to the advancement of Computer Vision. CNNs are composed of an enormous number of multiplication and addition operations performed on the input data to calculate the probability and predict the output. In multiplication, if any of the operands is zero-valued, then it is irrelevant in that particular output, and hence, these computations can be omitted to avoid unnecessary computation. In this study, we propose an architecture that performs the computations through parallel Processing Elements(PEs) and is also capable of skipping the ineffectual zero-valued computation to improve PE utilisation. Our proposed work, SpALEn, adopts the channel-first dataflow and is designed to perform the inference function with zero-skipping for enhanced performance. Moreover, due to the skipped computation, load imbalance occurs as the computation workload varies among the PEs. A dynamic logic is designed to mitigate this and ensure the hardware resources are utilised thoroughly. SpALEn achieves a speedup of 12 \(\times\) in comparison to a dense architecture. Imlijungla Longchar, Hemangee K. Kapoor |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2025 | WEnSIBR: Compression using dynamic bases supported with encoding and zone wise wear leveling for NVMs
Swati Upadhyay, Arijit Nath, Hemangee K. Kapoor |
J. Syst. Archit. | 3 |
| 2024 | Opportunistic Migration for Hybrid Memories While Mitigating Aging EffectsabstractHybrid memory systems composed of Non-volatile memory (NVM) and DRAM to exploit the high density of NVM and low access latency of DRAM. Phase Change Memory (PCM), a type of non-volatile memory, is a viable choice for main memory. High write latency and high voltage requirements for PCM lead to Biased Temperature Instability (BTI) aging and performance degradation. De-stressing the memory circuit at regular intervals controls BTI aging. Memory performance can be enhanced by migrating the highest write count memory pages across memory units. Migration and de-stress halt the service of regular requests and affect the performance of the system. Therefore, it is crucial to control migration and de-stress to enhance hybrid memory performance while mitigating BTI aging. We propose DOPMig, a de-stress-aware page migration technique. The policy migrates write-intensive pages to DRAM at regular intervals but opportunistically parallel to the de-stress operation. This method of background migration helps to reduce the migration overheads and improves performance. DOPMig achieves a performance gain of 22%, and improves memory service rate by 15%, and increases DRAM access by 24%. N. S. Aswathy, Hemangee K. Kapoor |
ICCD | 2 |
| 2024 | DynaCache: A Checkpoint Aware Reconfigurable Cache for Intermittently Powered Computing SystemsabstractEnergy harvesting devices are rapidly evolving to rival battery-backed technologies. Batteries have a shorter life- time and need maintenance compared to capacitors. Moreover, the usage of batteries comes with undeniable environmental costs. Distinct challenges have emerged such as performance enhancement, crash consistency of data, energy management, etc. To tackle these challenges we propose DynaCache: A checkpoint-aware reconfigurable cache for Intermittent powered computing systems. It is an intermittent computing architecture consisting of a reconfigurable L1 cache that dynamically trans- forms its write policy to strike a favorable trade-off between per- formance, data consistency, and forward progress of a program. Our proposal achieves a speedup in performance of 1.85x and 1.5x compared to non-volatile write-back cache and volatile write- through cache respectively. It also achieves a 99% reduction in dirty data compared to a volatile write-back cache guaranteeing a near absolute data consistency in a scenario of a power failure. Rishabh Mahanta, Hemangee K. Kapoor |
VLSI-SoC | 2 |
| 2024 | Migration-aware slot-based memory request scheduler to guarantee QoS in DRAM-PCM hybrid memories
N. S. Aswathy, Hemangee K. Kapoor |
J. Syst. Archit. | 2 |
| 2024 | AmLuCEP: Amalgamating LUT-based Compression and Adaptive Encoding Assisted Block Placement To Improve Lifetime of PCM-based Main MemoriesabstractWith the rising demands for high capacity memory and poor scalability of the existing DRAM-based main memories, the emerging Non-volatile memories captures higher attention due to their high density and low leakage power consumption. However, the possible consideration of such memories as alternatives of DRAM is largely hindered by their intrinsic drawbacks like high write latency, high write energy and low write endurance. In this article, we propose an integrated solution by combining the effect of compression and encoding assisted block placement to improve lifetime of NVMs. We have developed a compression technique called LUT_Comp by exploiting the word-level redundancy present in the words of the incoming cache blocks to NVM. LUT_Comp remains effective in reducing bit-flips in NVMs by offering a balance in compression ratio and coverage (Cov). Additionally, we also propose an encoding based block placement policy that places the compressed blocks in the appropriate half within the memory. The integrated approach of compression and block placement termed AmLuCEP offers a uniform bit-flips distribution while further reducing bit-flips in NVM. Experimental results show that AmLuCEP reduces bit-flips by 54%, 42%, 37%, 21%; energy consumption by 41%, 28%, 24%, 13%; and improves lifetime by 57%, 40%, 38%, 21% over baseline and the existing techniques READ [ 38 ], COEF [ 39 ] and SELEC [ 13 ], respectively. Arijit Nath, Hemangee K. Kapoor |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | CCGRID 2023: A Holistic Approach to Inclusion and Belongingabstract“CCGRID will act with responsibility as its primary consideration; with equity, diversity, and inclusion as its central goals.” from the CCGRID 2023 web site [1] Beth Plale, Preeti Malakar, Meenakshi D'Souza, Hemangee K. Kapoor, Yogesh L. Simmhan, Ilkay Altintas, S. Manohar 0001 |
CCGrid | 4 |
| 2023 | Look before you leap: An Access-based Prudent Page Migration for Hybrid MemoriesabstractHybrid memory composed of DRAM and PCM exploits benefit of both types of memory. The random page placement in such memories may cause write-intensive pages to be placed in PCM partition, which may adversely affect the memory performance due to the higher write latency of PCM. Migration of write-intensive pages to DRAM helps in improving memory service time. Existing techniques migrate pages having write access count greater than a predefined threshold. These techniques do not examine the access pattern once the choice to migrate the page has been made. This might lead to unnecessary migrations because the page may have been hot before the decision, but the number of access may have dropped after migration. To accurately identify the hot page, we propose an access-based prudent page migration method which uses an eDRAM buffer to migrate hot pages from PCM to DRAM. In this paper, we present a look-before-you-leap migration technique where after a page is identified as a hot page, makes a thoughtful decision regarding whether to migrate or not to migrate it. N. S. Aswathy, Hemangee K. Kapoor |
VLSI-SoC | 3 |
| 2023 | ADaMaT: Towards an Adaptive Dataflow for Maximising Throughput in Neural Network InferenceabstractWith the development of research in hardware for Convolutional Neural Network(CNNs) Algorithms, it becomes crucial to examine the different aspects of hardware design. CNNs are mainly used in computer vision applications, and translating these algorithms into hardware calls for adopting appropriate dataflow to improve the utilisation of hardware resources resulting in higher throughput. In particular, the inference task at each neuron position can be assigned to a compute unit in the hardware accelerator, and several such neuron positions can be completed in parallel. We observe that adopting a static dataflow for an architecture can result in the under-utilisation of resources because of the different dimensions of the data in the network. The motivation of this paper is built upon the need for adaptive dataflow for the design to improve the multiply-and-accumulate (MAC) utilisation in CNNs. We propose a method, ADaMaT, which adapts the dataflow at runtime by appropriately assigning tasks to the MAC units depending on the dimensions of the layers instead of a pre-determined assignment. The adaptive assignment tries to maximise the MAC utilisation and improve the throughput. We have performed a comparative analysis among different static dataflows and our proposed ADaMaT dataflow. Imlijungla Longchar, Hemangee K. Kapoor |
VLSI-SoC | 2 |
| 2023 | ALAMNI: Adaptive LookAside Memory Based Near-Memory Inference Engine for Eliminating Multiplications in Real-TimeabstractThe CNN algorithm seeks high performance and energy efficiency in real-time inference. The costly off-chip memory accesses put additional burdens on CNN's execution. Towards avoiding off-chip accesses, we propose ALAMNI, a novel near-memory architecture that expedites the CNNs in the logic layer of the Hybrid Memory Cube. We exploit intra- and inter-vault parallelism to accelerate the highly parallel CNN operations. The proposed ALAMNI replaces costly multiplications of CNNs with lookaside memory (LAM) based searches. The proposed ALAMNI policy is effective on unseen data as it discards the data pre-profiling overhead by an adaptive LAM update policy. The ALAMNI controller keeps the most frequent triplets of weight (W), activation (A), and multiplication result (M),, in the LAM to eliminate redundant computations. As an optimization, we incorporate a bitmasking concept to raise the hit rate of LAMs and further amortize computations. We also present a study on the relation between the amount of bitmasking and the loss of classification accuracy of the popular ConvNets. We keep the bitmasking as a reconfigurable feature of the ALAMNI units to achieve desired classification accuracy. Experimental results show substantial improvement in the system's performance and energy efficiency compared to the baseline and state-of-the-art. Palash Das 0001, Shashank Sharma 0002, Hemangee K. Kapoor |
IEEE Trans. Computers | 3 |
| 2023 | CAPMIG: Coherence-Aware Block Placement and Migration in Multiretention STT-RAM CachesabstractIn recent years, the increased working set size of applications craves more memory demand in terms of large-sized last-level caches (LLCs). To fulfill this, one of the promising technology is STTRAM. However, high write energy and write latency make it challenging to adopt it on a wide scale. Multiretention STTRAM caches have been considered an improvisation over standard STTRAM caches by reducing their retention time which reduces the write latency. Here, we have to negotiate with a refresh operation by applying various refresh management techniques. However, its retention period consumes significant refresh energy in the periodic refresh. In this article, we take help from the coherence protocol and decide the best retention type for each block. A block loaded on a write access is likely to get more writes in the future and is therefore loaded in the lowest retention time region. Similarly, instruction blocks are loaded in the highest retention time region. This helps in reducing the number of refreshes incurred by the blocks. During runtime, the blocks may change their access patterns, requiring a change in their retention region. This article also proposes a migration policy to relocate the blocks to appropriate regions during runtime. Identification of zero data value blocks and not refreshing them is an additional augmentation to our proposal. The experimental result using full system simulation shows a good reduction in the number of refreshes and energy consumption over the baseline designs. Sheel Sindhu Manohar, Hemangee K. Kapoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | A Predictable QoS-aware Memory Request Scheduler for Soft Real-time SystemsabstractA memory controller manages the flow of data to and from attached memory devices. The order in which a set of contending memory requests from different tasks are serviced significantly influences the rate of progress and completion times of these tasks. This in turn may affect the Quality-of-Service (QoS) delivered by these tasks. In this article, we focus towards the design of a QoS-aware memory controller targeted towards soft real-time systems. The proposed memory controller tries to generate an urgency-based schedule for the contending memory requests based on the allowable response time latencies associated with each request. The objective is to improve task-level response time predictability while maximizing acquired QoS. Exhaustive experiments carried out using real memory traces and standard simulation tools exhibit the practical efficacy of the proposed memory controller design. N. S. Aswathy, Arnab Sarkar 0001, Hemangee K. Kapoor |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2023 | Adaptive distribution of control messages for improving bandwidth utilization in multiple NoC
Sonal Yadav, Vijay Laxmi, Hemangee K. Kapoor, Manoj Singh Gaur, Amit Kumar 0046 |
J. Supercomput. | 3 |
| 2022 | Hydra: A near hybrid memory accelerator for CNN inferenceabstractConvolutional neural network (CNN) accelerators often suffer from limited off-chip memory bandwidth and on-chip capacity constraints. One solution to this problem is near-memory or in-memory processing. Non-volatile memory, such as phase-change memory (PCM), has emerged as a promising DRAM alternative. It is also used in combination with DRAM, forming a hybrid memory. Though near-memory processing (NMP) has been used to accelerate the CNN inference, the feasibility/efficacy of NMP remained unexplored for a hybrid main memory system. Additionally, PCMs are also known to have low write endurance, and therefore, the tremendous amount of writes generated by the accelerators can drastically hamper the longevity of the PCM memory. In this work, we propose Hydra, a near hybrid memory accelerator integrated close to the DRAM to execute inference. The PCM banks store the models that are only read by the memory controller during the inference. For entire forward propagation (inference), the intermediate writes from Hydra are entirely performed to the DRAM, eliminating PCM-writes to enhance PCM lifetime. Unlike the other in-DRAM processing-based works, Hydra does not eliminate any multiplication operations by using binary or ternary neural networks, making it more suitable for the requirement of high accuracy. We also exploit inter- and intra-chip (DRAM chip) parallelism to improve the system's performance. On average, Hydra achieves around 20x performance improvements over the in-DRAM processing-based state-of-the-art works while accelerating the CNN inference. Palash Das 0001, Ajay Joshi, Hemangee K. Kapoor |
DATE | 3 |
| 2022 | SRS-Mig: Selection and Run-time Scheduling of page Migration for improved response time in hybrid PCM-DRAM memoriesabstractHybrid memory systems with a combination of DRAM and Non-Volatile Memory (NVM) types can make use of scalability and performance of both NVM and DRAM. Random placement of pages in Phase Change Memory (PCM) with more write accesses incurs higher write latencies. So, migrating write intensive pages from PCM to DRAM helps to reduce execution time and memory response time for applications. Existing techniques mainly focus on selecting the page migration candidate and migrate it immediately when it becomes eligible. This direct migration approach can hamper the response time of regular memory accesses. So, in our paper, we identify migration candidates and in addition, schedule when they can be migrated to DRAM. To realize this, we have used Selection and Run-time Scheduling of page Migration (SRS-Mig), a frame-based scheduling approach for migrations and read/write requests. SRS-Mig reduces migration overhead and guarantees future accesses to migrated pages to yield an improved execution time and memory response time for the applications. Experimental evaluation shows 30% improvement in execution time; 26% improvement memory response time, and considerable energy savings with the existing baseline techniques. N. S. Aswathy, Sreesiddesh Bhavanasi, Arnab Sarkar 0001, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | CoSeP: Compression and Content-based Selection Procedure to Improve Lifetime of Encrypted Non-Volatile Main MemoriesabstractIn this paper, we propose a technique called CoSeP that combines the effect of compression and the content of the compressed blocks to reduce bit-flips in the encrypted PCM-based main memories. The blocks are compressed using the technique (out of FPC, BDI, and COMF) that offers minimum block size when the sizes of the two smallest compressed blocks are non-similar. However, for compressed blocks of similar sizes, the block is compressed using the technique that encounters minimum bit-flips, which reduces bit-flips further. Experimental results show that our technique gives a substantial reduction in bit-flips and improvements in lifetime compared to baseline and state-of-the-art techniques. Arijit Nath, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | Exploiting successive identical words and differences with dynamic bases for effective compression in Non-Volatile MemoriesabstractEmerging Non-volatile memories are considered as potential candidates for replacing traditional DRAM in main memory. However, downsides like long write latency, high write energy, and low write endurance make their direct adoption in the memory hierarchy challenging. Approaches that reduce the number of bits written are beneficial to overcome such drawbacks. Swati Upadhyay, Arijit Nath, Hemangee K. Kapoor |
ISLPED | 3 |
| 2022 | ZaLoBI: Zero avoiding Load Balanced Inference acceleratorabstractConvolutional neural networks are prevalent machine learning tools used in computer vision. Their ubiquitous use and high compute requirement have given rise to the design and development of accelerators for the same. Among several approaches to improve the performance of these accelerators, exploiting data sparsity has become very popular. Along similar lines, this paper proposes a design that skips the computation of zero-valued data operands and achieves better speedup. The savings in zero-valued computations also results in energy savings. The proposed accelerator exploits two levels of data parallelism to distribute work across multiple processing elements (PEs). The random distribution of zero values results in certain PEs getting idle due to the skipping of computations, thus creating load imbalance in the system. To address this issue, we extend our contribution in performing load balancing by dynamically scheduling tasks to the idle PEs. Our zero avoiding load-balanced accelerator (ZaLoBI) achieves around 76% and 5.57% speedup over the respective baselines and also outperforms the state-of-the-art works while saving energy. Imlijungla Longchar, Palash Das 0001, Hemangee K. Kapoor |
VLSI-SoC | 3 |
| 2022 | SWEL-COFAE : Wear Leveling and Adaptive Encoding Assisted Compression of Frequent Words in Non-Volatile Main MemoriesabstractEmerging Non-Volatile memories such as Phase Change Memory (PCM) and Resistive RAM are projected as potential replacements of the traditional DRAM-based main memories. However, limited write endurance and high write energy limit their chances of adoption as a mainstream main memory standard. Therefore, developing solutions that enhance the lifetime of these memories while offering a decent system performance has a great impact in building future large capacity and energy-efficient main memories. In this paper, we propose a word-level compression scheme called COMF to reduce bitflips in PCMs by removing the most repeated words from the cache blocks before writing into memory. COMF is augmented with an adaptive granularity based encoding technique to form COFAE, which reduces the bitflips to a further extent. We also propose SWEL-COFAE, which is an intra-line stride-based wear leveling technique to improve lifetime by balancing the bitflip pressure within the cells of the memory lines. Experimental results show that the proposed technique improves lifetime by 103% and reduces bitflips and energy by 60% and 59% respectively over baseline Arijit Nath, Hemangee K. Kapoor |
IEEE Trans. Computers | 2 |
| 2022 | CORIDOR: Using COherence and TempoRal LocalIty to Mitigate Read Disurbance ErrOR in STT-RAM CachesabstractIn the deep sub-micron region, “spin-transfer torque RAM” (STT-RAM ) suffers from “read-disturbance error” (RDE) , whereby a read operation disturbs the stored data. Mitigation of RDE requires restore operations, which imposes latency and energy penalties. Hence, RDE presents a crucial threat to the scaling of STT-RAM. In this paper, we offer three techniques to reduce the restore overhead. First, we avoid the restore operations for those reads, where the block will get updated at a higher level cache in the near future. Second, we identify read-intensive blocks using a lightweight mechanism and then migrate these blocks to a small SRAM buffer. On a future read to these blocks, the restore operation is avoided. Third, for data blocks having zero value, a write operation is avoided, and only a flag is set. Based on this flag, both read and restore operations to this block are avoided. We combine these three techniques to design our final policy, named CORIDOR. Compared to a baseline policy, which performs restore operation after each read, CORIDOR achieves a 31.6% reduction in total energy and brings the relative CPI (cycle-per-instruction) to 0.64×. By contrast, an ideal RDE-free STT-RAM saves 42.7% energy and brings the relative CPI to 0.62×. Thus, our CORIDOR policy achieves nearly the same performance as an ideal RDE-free STT-RAM cache. Also, it reaches three-fourths of the energy-saving achieved by the ideal RDE-free cache. We also compare CORIDOR with four previous techniques and show that CORIDOR provides higher restore energy savings than these techniques. Sheel Sindhu Manohar, Sparsh Mittal, Hemangee K. Kapoor |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | Pop-Crypt: Identification and Management of Popular Words for Enhancing Lifetime of EnCrypted Nonvolatile Main MemoriesabstractEmerging nonvolatile memories (NVMs) are considered as potential replacements of the traditional DRAM-based main memories. However, the nonvolatility feature of the NVMs may lead to the stealing of sensitive data stored in NVMs due to their prolonged data retention. Memory encryption turns out to be a viable option to provide data security. However, the existing encryption techniques, on account of the diffusion property, increase the number of bit-flips in the NVM cells, thus leading to their early wear out. Therefore, security and lifetime issues of the NVMs are difficult to go hand in hand. In this article, we identify words that are repeated in several memory blocks and term them as popular words. The proposal is to avoid the encryption of the popular words by maintaining them in a reference table. In our proposal, Pop-Crypt, every block to be written to the memory gets partially encrypted in which the popular words are replaced with pointers to the reference table, and other words get encrypted. The partially encrypted blocks (PEBs) reduce the number of bit-flips in PCM, thereby improving its lifetime significantly. Experimental results show that Pop-Crypt considerably improves lifetime, energy consumption, and system performance over baseline and a state-of-the-art technique. Arijit Nath, Hemangee K. Kapoor |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | A Soft Real-time Memory Request Scheduler for Phase Change Memory SystemsabstractPhase Change Memory (PCM) has emerged as a viable alternative to traditional DRAM memories especially in real-time embedded systems, due to their higher density and lower leakage power dissipation. However, PCM comes with its own drawbacks. Although, the performances of DRAM and PCM are comparable for memory reads, PCM is about three times slower in terms of write latency, and suffers from significantly lower write endurance. The high write latency of PCM may be detrimental to delivered QoS and may lead to deadline misses in real-time systems. To circumvent the problem, this paper proposes a novel memory scheduling scheme which employs separate write request buffer in order to prioritize reads over writes. The read requests are scheduled using an urgency based scheduler where urgency depends on allowable response times of tasks. The write requests are serviced when there are no pending reads using a similar urgency based scheduler as used for read requests. Experimental evaluation using standard benchmarks reveal that the proposed scheme is able to achieve better normalized QoS compared to existing scheduling techniques for PCM and comparable access latencies with respect to DRAM. N. S. Aswathy, Hemangee K. Kapoor, Arnab Sarkar 0001 |
RTCSA | 2 |
| 2021 | CLU: A Near-Memory Accelerator Exploiting the Parallelism in Convolutional Neural NetworksabstractConvolutional/Deep Neural Networks (CNNs/DNNs) are rapidly growing workloads for the emerging AI-based systems. The gap between the processing speed and the memory-access latency in multi-core systems affects the performance and energy efficiency of the CNN/DNN tasks. This article aims to alleviate this gap by providing a simple and yet efficient near-memory accelerator-based system that expedites the CNN inference. Towards this goal, we first design an efficient parallel algorithm to accelerate CNN/DNN tasks. The data is partitioned across the multiple memory channels (vaults) to assist in the execution of the parallel algorithm. Second, we design a hardware unit, namely the convolutional logic unit (CLU), which implements the parallel algorithm. To optimize the inference, the CLU is designed, and it works in three phases for layer-wise processing of data. Last, to harness the benefits of near-memory processing (NMP), we integrate homogeneous CLUs on the logic layer of the 3D memory, specifically the Hybrid Memory Cube (HMC). The combined effect of these results in a high-performing and energy-efficient system for CNNs/DNNs. The proposed system achieves a substantial gain in the performance and energy reduction compared to multi-core CPU- and GPU-based systems with a minimal area overhead of 2.37%. Palash Das 0001, Hemangee K. Kapoor |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2021 | TARTS: A Temperature-Aware Real-Time Deadline-Partitioned Fair Scheduler
Sanjay Moulik, Arnab Sarkar 0001, Hemangee K. Kapoor |
J. Syst. Archit. | 3 |
| 2021 | nZESPA: A Near-3D-Memory Zero Skipping Parallel Accelerator for CNNsabstractConvolutional neural networks (CNNs) are one of the most popular machine learning tools for computer vision. The ubiquitous use in several applications with its high computation-cost has made it lucrative for optimization through accelerated architecture. State-of-the-art has either exploited the parallelism of CNNs, or eliminated computations through sparsity or used near-memory processing (NMP) to accelerate the CNNs. We introduce NMP-fully sparse architecture, which acquires all three capabilities. The proposed architecture is parallel and hence processes the independent CNN tasks concurrently. To exploit the sparsity, the proposed system employs a dataflow, namely, Near-3D-Memory Zero Skipping Parallel dataflow or nZESPA dataflow. This dataflow maintains the compressed-sparse encoding of data that skips all ineffectual zero-valued computations of CNNs. We design a custom accelerator which employs the nZESPA dataflow. The grids of nZESPA modules are integrated into the logic layer of the hybrid memory cube. This integration saves a significant amount of off-chip communications while implementing the concept of NMP. We compare the proposed architecture with three other architectures which either do not exploit sparsity (NMP-dense) or do not employ NMP (traditional-fully sparse) or do not include both (traditional-dense). The proposed system outperforms the baselines in terms of performance and energy consumption while executing CNN inference. Palash Das 0001, Hemangee K. Kapoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Investigating Frequency Scaling, Nonvolatile, and Hybrid Memory Technologies for On-Chip Routers to Support the Era of Dark SiliconabstractIn the era of dark silicon, several components on the chip [i.e., cores, memory, and network on chip (NoC)] need to be powered-off or run in low-power mode. This is mainly due to the increased leakage power consumption at smaller technology nodes. Other than the power consumed by cores and caches, power and performance of the interconnects is a significant factor as the communication network consumes a considerable share of the power budget. In particular, the buffers used at every port of the NoC router consume considerable dynamic as well as static power. To support dark silicon and save energy, a popular approach is to power off the routers and wake them up when needed. However, this affects the packet latency, and we need to observe the traffic through the nodes to decide turning the routers ON-OFF. In this article, we propose to keep the routers always powered ON to maintain constant connectivity and investigate various approaches. One proposal is to frequency scale the routers connected to powered OFF nodes, and the other proposals are to use a combination of SRAM and nonvolatile spin-transfer torque random access memory-based VCs in the routers. By managing which VCs to be active at a given time, we achieve energy savings. The proposals are evaluated by varying the percentage of dark nodes on the chip. The experimental results show that all proposals yield significant energy savings while maintaining connectivity. Khushboo Rani, Hemangee K. Kapoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Improving the Performance of Hybrid Caches Using Partitioned Victim CachingabstractNon-Volatile Memory technologies are coming as a viable option on account of the high density and low-leakage power over the conventional SRAM counterpart. However, the increased write latency reduces their chances as a substitute for SRAM. To attenuate this problem, a hybrid STT-RAM-SRAM architecture is proposed where with large STT-RAM ways, the small SRAM ways are incorporated for handling the write operations. However, the performance gain obtained from such an architecture is not as much as expected on account of the larger miss rate caused by smaller SRAM partition. This, in turn, may limit the amount of cache capacity. This article attempts to reduce the miss penalty and improve the average memory access time by retaining the victims evicted from the hybrid cache in a smaller, fully associative SRAM structure called the victim cache. The victim cache is accessed on a miss in the primary hybrid cache. Hits in the victim cache require an exchange of the block between the main hybrid cache and the victim cache. In such cases, to effectively place the required block in the appropriate region of the main hybrid cache, we propose an access-based block placement technique. Besides, to manage the runtime load and the uneven evictions of the SRAM partition, we also present a dynamic region-based victim cache partitioning method to hold the victims dedicated to each region. Experimental evaluation on a full system simulator shows significant improvement in the performance and execution time along with a reduction in the overall miss rate. The proposed policy also increases the endurance of Hybrid Cache Architectures (HCA) by reducing writes in the STT partition. Sukarn Agarwal, Hemangee K. Kapoor |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2020 | ZENCO: Zero-bytes based ENCOding for Non-Volatile Buffers in On-Chip InterconnectsabstractWith multiple cores integrated on the same die, communication across cores is managed by on-chip interconnect called Network-on-Chip (NoC). Power and performance of these interconnect become a significant factor as the communication network has limitations of high network power consumption and delay. The buffers used in the NoC router consume a considerable amount of dynamic as well as static power. This paper attempts to reduce static power consumption by using Non-Volatile Memory technology based STT-RAM buffers. STT-RAM technology has the advantage of higher density and low leakage, but suffer from costly write operation, and weaker write endurance. These characteristics on whole impacts on the total network power consumption, network latency, and lifetime of the router. In this paper, we propose a compression technique at Network Interface, which is based on zero bytes present in data packets. We also propose a compression with the wear-leveling technique, which reduces the write variation in VCs to reduce uneven writes across the buffers.Experimental evaluation on full system simulator shows that proposed policy obtains 0.37 compression ratio and 63% reduction of total network flit. All these results in a significant decrease in total network power. The policies also show remarkable improvement in the lifetime with wear-leveling compression. Khushboo Rani, Hemangee K. Kapoor |
DAC | 2 |
| 2020 | Dimming Hybrid Caches to Assist in Temperature Control of Chip MultiProcessorsabstractThe continuous rise of on-chip components like cores and caches has brought enormous computing capabilities at the cost of high leakage power and temperature. A recent study has shown a substantial spatial temperature variance in modern large on-chip caches. This high temperature elevates the cooling cost and becomes responsible for the thermal breakdown of the chip. One solution to reduce the leakage is the use of non-volatile memory (NVM) like STT-RAM. Other includes incorporating the concept of dark silicon. In this paper, we amalgamate the idea of using STT-RAM in the last level cache (LLC) and the dark silicon approach to shut down certain cache ways to leverage the benefits from both. We address the downsides like higher access latencies of STT-RAM by the use of hybrid cache (SRAM + STT-RAM) and weak endurance of the STT-RAM by wear leveling. We propose a system to handle three different temperature thresholds (high, medium, and low) by appropriately selecting the type of cache ways to be powered off. The proposed system delivers up to 5.38 K reduction in temperature compared to the baseline, 93% reduction in leakage power with an EDP gain up to 92%. Chirag Joshi, Palash Das 0001, Ashwini A. Kulkarni, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | WELCOMF: wear leveling assisted compression using frequent words in non-volatile main memoriesabstractEmerging Non-Volatile memories such as Phase Change Memory (PCM) and Resistive RAM are projected as potential replacements of the traditional DRAM-based main memories. However, limited write endurance and high write energy limit their chances of adoption as a mainstream main memory standard. Arijit Nath, Hemangee K. Kapoor |
ISLPED | 2 |
| 2020 | DidaSel: dirty data based selection of VC for effective utilization of NVM buffers in on-chip interconnectsabstractIn a multi-core system, communication across cores is managed by an on-chip interconnect called Network-on-Chip (NoC). The utilization of NoC results in limitations such as high communication delay and high network power consumption. The buffers of the NoC router consume a considerable amount of leakage power. This paper attempts to reduce leakage power consumption by using Non-Volatile Memory technology-based buffers. NVM technology has the advantage of higher density and low leakage but suffers from costly write operation, and weaker write endurance. These characteristics impact on the total network power consumption, network latency, and lifetime of the router as a whole. Khushboo Rani, Sukarn Agarwal, Hemangee K. Kapoor |
ISLPED | 3 |
| 2020 | Reuse Distance-based Victim Cache for Effective Utilisation of Hybrid Main Memory SystemabstractHybrid main memories comprising DRAM and Non-volatile memories (NVM) are projected as potential replacements of the traditional DRAM-based memories. However, traditional cache management policies designed for improving the hit rate lack awareness of the comparative latency of read-write for NVM blocks where the write latency is more than the read latency. Therefore, developing cache management techniques that reduce costly write-backs of the NVM blocks, yet maintain a fair hit rate in the cache, is of paramount importance. We propose two techniques based on the use of a small victim cache associated with the last-level cache that helps in retaining on the chip critical DRAM and NVM blocks. Victim cache being a scarce resource, we intend to keep only performance-critical blocks in the victim cache by exploiting the idea of reuse distance. The first technique, Victim Cache Replacement Policy, works on the replacement policy of the victim cache by preferential eviction of DRAM blocks over NVM blocks. However, the second technique, Prioritized Partitioning of victim cache, logically partitions the victim cache, giving a smaller share to the DRAM blocks and a relatively larger share to the NVM blocks. Experimental evaluation on full-system simulator shows significant improvement in system performance and reduction in the number of write-backs to the NVM partition of the main memory compared to the baseline and existing technique. Additionally, NVM reads and DRAM miss rate are also improved, leading to further performance enhancement. Arijit Nath, Sukarn Agarwal, Hemangee K. Kapoor |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | Improving Static Power Efficiency via Placement of Network Demultiplexer over Control Plane of Router in Multi-NoCsabstractNetwork Demultiplexer (Net-Demux) is an essential hardware unit in multiple NoCs for traffic distribution between the NoC networks. This paper proposes a novel idea of the placement of Net-Demux at the control plane of switch allocator of the router to improve static power and energy efficiency as compared to conventional data plane placement at the Network Interface (NI). Sonal Yadav, Vijay Laxmi, Manoj Singh Gaur, Hemangee K. Kapoor |
DAC | 4 |
| 2019 | Enhancing the Lifetime of Non-Volatile Caches by Exploiting Module-Wise Write RestrictionabstractThe emerging Non-Volatile Memory (NVM) technologies offer a good combination of high density and near-zero leakage power, becoming the strongest candidate in the memory hierarchy including caches. However, the weak write endurance of these memories creates a bottleneck towards their employment in the cache hierarchy. This weak endurance shows its effects due to the write variations introduced by the applications and the existing cache management policies. Such variations result in early breakdown of the NVM cells reducing the effective lifetime of the NVM memory component. This paper proposes a technique to mitigate intra-set write variation, i.e. write variations occurring within the cache set. Our policy divides the cache logically into multiple equal-sized modules. During execution, the writes are distributed uniformly across different ways of the different modules within the set. Experimental results using full system simulation show that the proposed technique reduces the intra-set write variation significantly over the baseline and the existing techniques. Sukarn Agarwal, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | Towards Optimizing Refresh Energy in embedded-DRAM Caches using Private BlocksabstractIn recent years, the increased working set size of applications craves for more memory demand in terms of large size Last Level Caches (LLC). To fulfill this, embedded DRAM (eDRAM) caches have been considered as one of the best alternatives over conventional SRAM caches. eDRAM has a property of low leakage and provides more capacity in the same area footprint of SRAM. However, its retention period consumes significant refresh energy in the periodic refresh. In this paper, we present an approach to minimize the total energy spent on refreshes by considering the presence of private blocks in the LLC. Our approach restricts refreshing of those blocks that are loaded exclusively from the main memory on an LLC miss. Experimental result using full system simulation show 55% reduction in the total number of refreshes compared to baseline policy; and 62% reduction in total power consumption over SRAM. Sheel Sindhu Manohar, Sukarn Agarwal, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Cost effective routing techniques in 2D mesh NoC using on-chip transmission lines
Dipika Deb, John Jose, Shirshendu Das, Hemangee K. Kapoor |
J. Parallel Distributed Comput. | 4 |
| 2019 | Dynamic reconfiguration of embedded-DRAM caches employing zero data detection based refresh optimisation
Sheel Sindhu Manohar, Hemangee K. Kapoor |
J. Syst. Archit. | 2 |
| 2019 | Improving the Lifetime of Non-Volatile Cache by Write RestrictionabstractThe attractive features such as low static power and high density exhibited by the Non-Volatile Memory (NVM) technologies makes them a promising candidate in the memory hierarchy, including caches. However, the limited write endurance with the write variations governed by the access patterns and the applied replacement policies reduce the chance of NVMs as a successor of SRAM. These write variations are of concern as they not only breakdown the NVM cells but also reduce the effective lifetime. This paper proposes efficient techniques to mitigate the intra-set write variation to improve the lifetime of the NVM cache. Our first two techniques partition the cache into windows of equal size and distribute the writes uniformly across the cache set by employing the window as write-restricted or read-only. The selection of the window in these techniques is by rotation or with the help of counters. In our third technique, different cache ways are employed as a write-restricted over the period of execution to distribute the writes uniformly. Experimental results using full system simulation show the significant reduction in intra-set write variation along with improvement in the cache lifetime. Sukarn Agarwal, Hemangee K. Kapoor |
IEEE Trans. Computers | 2 |
| 2019 | Exploring the Role of Large Centralised Caches in Thermal Efficient Chip DesignabstractIn the era of short channel length, Dynamic Thermal Management (DTM) has become a challenging task for the architects and designers engineering modern Chip Multi-Processors (CMPs). Ever-increasing demand of processing power along with the developed integration technology produces CMPs with high power density, which in turn increases effective chip temperature. This increased temperature leads to increase in the reliability issues for the chip-circuitry with significant increment in leakage power consumption. Recent DTM techniques apply DVFS or Task Migration to reduce temperature at the cores, the hottest on-chip components, but often ignore the on-chip hot caches. To commensurate the high data demand of these cores, most of the modern CMPs are equipped with large multi-level on-chip caches, out of which on-chip Last Level Caches (LLCs) occupy the largest on-chip area. These LLCs are accounted for their significantly high leakage power consumption that can also potentially generate on-chip hotspots at the LLCs similar to the cores. As power consumption constructs the backbone of heat dissipation, hence, this work dynamically shrinks cache size while maintaining performance constraint to reduce LLC leakage, primarily. These turned-off cache portions further work as on-chip thermal buffers for reducing average and peak temperature of the CMP without affecting the computation. Simulation results claim that, at a minimal penalty on the performance, proposed cache-based thermal management having 8MB centralised multi-banked shared LLC gives around 5°C reduction in peak and average chip temperature, which are comparable with a Greedy DVFS policy. Shounak Chakraborty 0001, Hemangee K. Kapoor |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2019 | Write Variation Aware Buffer Assignment for Improved Lifetime of Non-Volatile Buffers in On-Chip InterconnectsabstractWith multiple cores integrated on the same die, communication across cores is managed by on-chip interconnect called network-on-chip (NoC). Power and performance of these interconnect is a significant factor as the communication network consumes a considerable share of the power budget. In particular, the buffers used at every port of the NoC router consume considerable dynamic as well as static power. This paper attempts to reduce static power consumption by using non-volatile memory technology-based spin-transfer torque random access memory (STT-RAM) buffers. STT-RAM technology has the advantage of high density and low leakage but suffers from weaker write endurance. This impacts the lifetime of the router as a whole. The buffers in a router are allocated to virtual networks (VNets) and in-turn to virtual channels (VCs) within each VNet. To reduce uneven writes across the buffers, we propose policies to reduce intra-VNet write variation and inter-VNet write variation. The former performs write variation aware VC allocation in each VNet, and the latter does write variation aware buffer assignments to each VNet. Experimental evaluation on full system simulator shows that proposed policies reduce write variation to almost 0% and improve lifetime by 3.3 and 19.9 times for intra-VNet and inter-VNet, respectively. We also get significant gains in the energy delay product. Khushboo Rani, Hemangee K. Kapoor |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Towards Near-Data Processing of Compare Operations in 3D-Stacked MemoryabstractThe gap between the processing speed and memory access speed of the modern multi-core systems has become a bottleneck for the emerging data-intensive workloads. In this scenario, it has become a smarter idea to move some amount of computation closer to the data, thus stimulating the concept of near-data processing (NDP). Compare or scanning, the core operations of many applications, typically in a database, can leverage the benefits of NDP. We propose near-data compare unit (NDCU), a less-invasive hardware, that can be integrated with the existing ecosystem of the hybrid memory cube (HMC). While integrating NDCU, we have designed two full-system architectures, one is lighter NDP with no parallelism (NNP) and the second is NDP with vault level parallelism (NVLP). While the first architecture is more power and area efficient, the second one is very fast with negligible overheads. With the motive of carrying out scan operation, we have specifically implemented 'compare-n-hit', 'compare-n-count' and 'compare-n-max' operations on both row-store and column-store databases and found significant improvements over conventional CPU-based system. We get around 2.3x and 37x performance improvement in NNP and NVLP architectures respectively. In both the designs, we reduce the energy consumption by around 8x on an average. Palash Das 0001, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Utility Aware Snoozy Caches for Energy Efficient Chip Multi-ProcessorsabstractHeavy leakage power consumption of on-chip last level caches (LLCs) has become the primary obstacle for architecting chip multi-processors (CMPs) in recent times. As leakage power has a direct relationship with the supply voltage, hence, periodic access profile based dynamic voltage scaling (DVS) in the LLC banks can be a promising option towards reducing this heavy cache leakage. A plethora of prior attempts have reduced this by anticipating working set size (WSS) of the applications and eventually putting some portions of the cache banks in low power mode. This proposed work aims to reduce leakage by putting a whole LLC bank into a low power (snoozy) mode through exploiting DVS at cache banks having minimal usages. Additionally, the resulting performance impacts of the low power snoozy mode are alleviated further by putting some snoozy banks in active mode on-demand. Experimental evaluations using full system simulation on a multi-banked 2MB 8-way set associative L2 cache show 10% more leakage savings on an average over a prior drowsy technique. Ashwini A. Kulkarni, Shounak Chakraborty 0001, Shrinivas P. Mahajan, Hemangee K. Kapoor |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | Analysing the Role of Last Level Caches in Controlling Chip TemperatureabstractDynamic Thermal Management (DTM) has become a major concern for the chip-designers, as it becomes a challenging task in recent power densed high performance Chip Multi-Processors (CMPs), due to integration of more on-chip components to meet ever increasing demand of processing power. The increased chip temperature incorporates severe circuit errors along with significant increment in leakage power consumption. Traditional DTM techniques apply DVFS or task migration to reduce core temperature, as cores are considered as the hottest on-chip components. Additionally, to commensurate high data demand of these high performance cores, large on-chip Last Level Caches (LLCs) are attached, which are the principal contributors to the on-chip leakage power consumption and occupy the largest on-chip area. As power consumption reduction plays the pivotal role in temperature reduction, hence, this work dynamically shrinks the cache size not only to reduce leakage power consumption, but also, to create on-chip thermal buffers for reducing average chip temperature by exploiting the heat transfer physics. Cache resizing decisions are taken based upon the generated cache hotspots and/or the access patterns, during process execution. Simulation results of the proposed thermal management method are compared with an existing DVFS based method (at cores) and a prior drowsy cache based technique to show its effectiveness. Shounak Chakraborty 0001, Hemangee K. Kapoor |
IEEE Trans. Sustain. Comput. | 2 |
| 2018 | Reuse-Distance-Aware Write-Intensity Prediction of Dataless Entries for Energy-Efficient Hybrid Caches
Sukarn Agarwal, Hemangee K. Kapoor |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Targeting inter set write variation to improve the lifetime of non-volatile cache using fellow setsabstractHigh density and low static power exhibited by nonvolatile technologies (NVM) have made them popular candidates in the memory hierarchy, including caches. Writes within a cache set are governed by the access pattern as well as replacement policies, leading to a large write variation. This variation is of concern as it leads to early breakdown of the NVM cells due to large writes thus reducing the effective lifetime. This paper presents a technique to improve the lifetime of non-volatile caches by reducing the inter-set write variation. Our policy partitions the cache sets into groups called fellow groups. Every set has two logical parts: Normal and Reserved. Sets within a fellow group can use the reserved parts from their fellow sets to distribute the writes uniformly. Experimental results using full system simulation show that the proposed technique shows significant reduction in inter-set write variation over the baseline and existing technique. Sukarn Agarwal, Hemangee K. Kapoor |
VLSI-SoC | 2 |
| 2017 | Dynamic Associativity Management in Tiled CMPs by Runtime Adaptation of Fellow SetsabstractThe non-uniform distribution of memory accesses among the cache sets results in some sets being used heavily while certain others remaining underutilized. Dynamic associativity management (DAM) is a technique to allow the heavily used sets to distribute their load among the lightly used sets thus improving the overall utilization of the cache. CMP-SVR is a previously proposed DAM based technique, where each set is divided into two sections: normal storage (NT) and reserve storage (RT). Some number of ways (25 to 50 percent) from each set are reserved for RT and the remaining ways belong to NT. The sets are divided into groups called fellow-groups and a set can use the reserve-ways of its fellow sets to increase its associativity during execution. Though CMP-SVR improves performance the formation of its fellow-groups is static: once created it never changes. It has been observed that some fellow-groups have more number of heavily used sets than the other fellow-groups. As a result the cache loads are not uniformly distributed among the fellow-groups. Also the behavior of sets changes dynamically: a lightly used set may become heavily used after a number of execution cycles. This paper studies the behavior of each set in detail and proposes a DAM based technique which improves the performance compared to other DAM based techniques. The proposed technique called FS-DAM dynamically creates fellow-groups based on the current set loads ensuring that the heavily used sets are evenly distributed among all the fellow-groups. Such distribution increases the utilization of the cache and hence improves performance. Full system simulation shows an average of 6.62 and 16.74 percent improvements, in FS-DAM as compared to CMP-SVR, in terms of CPI (Cycles Per Instruction) and MPKI (Miss Per Thousand Instructions) respectively. Comparing with Z-Cache the improvements are 6.21 percent (CPI) and 14.65 percent (MPKI). The proposed policy also shows better performance over V-Way and SBC. Shirshendu Das, Hemangee K. Kapoor |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Restricting writes for energy-efficient hybrid cache in multi-core architecturesabstractEmerging non-volatile memory technology Spin Transfer Torque Random Access Memory (STT-RAM) is a good candidate for the Last Level Cache (LLC) on account of high density, good scalability and low power consumption. However, expensive write operation reduces their chances as a replacement of SRAM. To handle these expensive write operations, an STT-RAM/SRAM hybrid cache architecture is proposed that reduces the number of writes and energy consumption of the STT-RAM region in the LLC by considering the existence of private blocks. Our approach allocates dataless entries for such kind of blocks when they are loaded in the LLC on a miss. We make changes in the conventional MESI protocol by adding new states to deal with the dataless entries. Experimental results using full system simulator shows 73% savings in write operations and 20% energy savings compared to an existing policy. Sukarn Agarwal, Hemangee K. Kapoor |
VLSI-SoC | 2 |
| 2016 | Static energy reduction by performance linked dynamic cache resizingabstractThe increased power density with short channel effect in modern transistors significantly increases the leakage energy consumptions of on-chip Last Level Caches (LLCs) in recent Chip Multi-Processors(CMPs). Performance linked dynamic shrinking in the LLC size is a promising option for reducing cache leakage. Prior works attempt to reduce the cache leakage by predicting Working Set Size(WSS) of the applications and by putting some cache portions in low power mode. This paper aims to reduce leakage energy by using a combination of cache bank shutdown and way shutdown. The banks with minimal usages are candidates for shutdown. In banks with average usages, some ways are turned off to save leakage. To mitigate the impact of smaller set-size, we apply dynamic associativity management technique. Experimental evaluation using full system simulation on a 4MB 8-way set associative L2 cache gives 70% average savings in static energy with 35% average savings in EDP. In case application's cache demand increases we can turn-on some ways to maintain performance. Shounak Chakraborty 0001, Hemangee K. Kapoor |
VLSI-SoC | 2 |
| 2016 | A Framework for Block Placement, Migration, and Fast Searching in Tiled-DNUCA ArchitectureabstractMulticore processors have proliferated several domains ranging from small-scale embedded systems to large data centers, making tiled CMPs (TCMPs) the essential next-generation scalable architecture. NUCA architectures help in managing the capacity and access time for such larger cache designs. It divides the last-level cache (LLC) into multiple banks connected through an on-chip network. Static NUCA (SNUCA) has a fixed address mapping policy, whereas dynamic NUCA (DNUCA) allows blocks to relocate nearer to the processing cores at runtime. To allow this, DNUCA divides the banks into multiple banksets and a block can be placed in any bank within a particular bankset. The entire bankset may need to be searched to access a block. Optimal bankset searching mechanisms are essential for getting the benefits from DNUCA. This article proposes a DNUCA-based TCMP architecture called TLD-NUCA. It reduces the LLC access time of TCMP and also allows a heavily loaded bank to distribute its load among the underused banks. Instead of other DNUCA designs, TLD-NUCA considers larger banksets. Such relaxations result in more uniform load distribution than existing DNUCA-based TCMP (T-DNUCA). Considering larger banksets improves the utilization factor, but T-DNUCA cannot implement it because of its expensive searching mechanism. TLD-NUCA uses a centralized directory, called TLD, to search a block from all the banks. Also, the proposed block placement policy reduces the instances when the central TLD needs to be contacted. It does not require the expensive simultaneous search as needed by T-DNUCA. Better cache utilization and a reduction in LLC access time improve the miss rate as well as the average memory access time (AMAT). Improving the miss rate and AMAT results in improvements in cycles per instructions (CPI). Experimental analysis found that TLD-NUCA improves performance by 6.5% as compared to T-DNUCA. The improvement is 13% as compared to the SNUCA-based TCMP design. Shirshendu Das, Hemangee K. Kapoor |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2013 | Towards a Better Cache Utilization Using Controlled Cache PartitioningabstractMany multi-core processors nowadays employ a shared Last Level Cache (LLC). Partitioning LLC becomes more important as LLC is shared among the cores. Past research has demonstrated that the traditional least recently used (LRU) based partitioning cum replacement policy has adverse effects on parameters like instruction per cycle (IPC), miss rate and speedup. This leads to poor performance in an environment when multiple cores compete for one global LLC. Applications, enjoying locality of reference are purely benefited by LRU, however LRU fails for the applications showing working set size (WSS) large than the LLC size. In this work, we propose a scheme which allows cores to steal/donate their lines upto a threshold and give them a chance to adjust their partition when there is a miss. Instead of maintaining strict target partitioning, we introduce a flexible threshold window. Our evaluation with multiprogrammed workloads shows significant performance improvement. Prateek D. Halwe, Shirshendu Das, Hemangee K. Kapoor |
DASC | 3 |
| 2013 | A formal framework for interfacing mixed-timing systems
Shirshendu Das, Parasara Sridhar Duggirala, Hemangee K. Kapoor |
Integr. | 3 |
| 2013 | Formal Approach for DVS-Based Power Management for Multiple Server System in Presence of Server Failure and RepairabstractThe paper presents a DVS-based power management policy for multiprocessor systems. The aim is to optimize power consumption by keeping the job loss probability as a system-wide constraint. Optimal values for service rate are computed using an ideal setting where speed can change continuously. As real processors have discrete speed levels, we switch between two nearby speeds to achieve the optimal rate. We develop a formal model of such a system using the probabilistic model checker PRISM and prove properties satisfied by the system. We demonstrated the applicability of the policy on multiple servers and under both kinds of deadlines: DBS and DES. For a constraint value of 25%, the DBS model achieved power savings of 29.46% in theoretical, 8.75% in actual, and 7.23% in leakage power. The DES model achieved power savings of 30% in theoretical, 11.9% in actual, and 8.7% in leakage power. For robustness, a repair facility was used which can have repairmen varying from one to the total number of servers. Lalit Chandnani, Hemangee K. Kapoor |
IEEE Trans. Ind. Informatics | 2 |
| 2013 | Design and formal verification of a hierarchical cache coherence protocol for NoC based multiprocessors
Hemangee K. Kapoor, Praveen Kanakala, Malti Verma, Shirshendu Das |
J. Supercomput. | 1 |
| 2009 | A Process Algebraic View of Latency-Insensitive SystemsabstractLatency-insensitive (LI) systems are those which can function correctly in spite of delays along its connecting wires. This delay is assumed to be a multiple of the clock period. The paper presents a single-clock process algebraic model for such systems. It gives the definitions for LI computational blocks and LI connectors. Important properties for these are shown to be satisfied. Composition of such modules can be done by the parallel composition operator of the process algebra. Conditions are given to check for liveness and deadlock freedom of LI systems. Comparison of latency equivalence between streams of events can be done using the model and this leads to a method of proving latency-equivalent modules. The paper is a step toward high-level specification and verification of such systems. The work can be extended to address more complex interconnections by modeling the underlying finite-state machines. Hemangee K. Kapoor |
IEEE Trans. Computers | 1 |
| 2007 | Controllable Delay-Insensitive Processes
Mark B. Josephs, Hemangee K. Kapoor |
Fundam. Informaticae | 2 |
| 2006 | Formal Modelling and Verification of an Asynchronous DLX PipelineabstractA five stage pipeline of an asynchronous DLX processor is modelled and its control flow is verified. The model is built using an asynchronous pipeline of latches separated by processing logic. We model two versions of this pipeline: one using latch controllers with four-phase semi-decoupled and another using fully-decoupled protocol. All the processing units are modelled as processes in the PROMELA language of the Spin tool. The model is verified in Spin by means of assertions, LTL properties and progress labels. A useful observation obtained from the study is that: although the semi-decoupled protocol has the potential to hold a data item in every latch, in the presence of processing logic, at most alternate stages can execute at a given time. Its implication being, in the case of control and data hazards no pipeline stalls are necessary, in the case of fully decoupled version, all stages could execute valid instructions at the same time. All the models were verified to be free from deadlock Hemangee K. Kapoor |
SEFM | 1 |
| 2006 | Verification and Implementation of Delay-Insensitive Processes in Restrictive Environments
Hemangee K. Kapoor, Mark B. Josephs, Dennis P. Furey |
Fundam. Informaticae | 1 |
| 2004 | Decomposing specifications with concurrent outputs to resolve state coding conflicts in asynchronous logic synthesisabstractSynthesis of asynchronous logic using the tool Petrify requires a state graph with a complete state coding. It is common for specifications to exhibit concurrent outputs, but Petrify is sometimes unable to resolve the state coding conflicts that arise as a result, and hence cannot synthesise a circuit. A pair of decomposition heuristics (expressed in the language of Delay-Insensitive Sequential Processes) are given that helps one to obtain a synthesisable specification. The second heuristic has been successfully applied to a set of nine benchmarks to obtain significant reductions both in area and in synthesis time, compared with synthesis performed on the original specifications. Hemangee K. Kapoor, Mark B. Josephs |
DAC | 1 |
| 2004 | Modelling and verification of delay-insensitive circuits using CCS and the Concurrency Workbench
Hemangee K. Kapoor, Mark B. Josephs |
Inf. Process. Lett. | 1 |