VLDB 2026 Research / reviewers in the wild / expert
Shirshendu Das
dblp:131/1057
· DBLP profile ↗
18ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeSTAR: Hardware Trojans and its mitigation strategy in NoC routers
Josna Philomina, Rekha K. James, Shirshendu Das, Palash Das 0001, Daleesha M. Viswanathan |
Integr. | 3 |
| 2026 | Exploiting virtual channel allocation policies in STT-RAM buffers of NoC routers through hardware Trojan
Josna Philomina, Rekha K. James, Palash Das 0001, Shirshendu Das, Daleesha M. Viswanathan |
J. Syst. Archit. | 4 |
| 2025 | SmartDeCoup: Decoupling the STT-RAM LLC for even write distribution and lifetime improvement
Prabuddha Sinha, Krishna Prathik B. V., Shirshendu Das, T. Venkata Kalyan |
J. Syst. Archit. | 3 |
| 2024 | $\mathcal{F}lush+early\mathcal{R}\text{ELOAD}$: Covert Channels Attack on Shared LLC Using MSHR MergingabstractModern multiprocessors include multiple cores, all of them share a large Last Level Cache (LLC). Because of the shared nature of LLC, different timing channel attacks can be built by exploiting LLC behaviors. Covert Channel and Side Channel are two well-known attacks used to steal sensitive information from a secure application. While several countermeasures have already been proposed to prevent these attacks, the possibility of discovering new variants cannot be overlooked. In this paper, we propose a covert channel attack designed to circumvent the state-of-the-art countermeasure for preventing the Flush+Reload attack. Experimental results indicate that the proposed attack renders the current state-of-the-art countermeasure futile and ineffective. Aditya S. Gangwar, Prathamesh Nitin Tanksale, Shirshendu Das, Sudeepta Mishra |
DATE | 3 |
| 2024 | RSPP: Restricted Static Pseudo-Partitioning for Mitigation of Cross-Core Covert Channel AttacksabstractCache timing channel attacks exploit the inherent properties of cache memories: hit and miss time along with the shared nature of the cache to leak secret information. The side channel and covert channel are the two well-known cache timing channel attacks. In this article, we propose Restricted Static Pseudo-Partitioning (RSPP), an effective partition-based mitigation mechanism that restricts the cache access of only the adversaries involved in the attack. It has an insignificant impact of only 1% in performance, as the benign processes have access to the full cache and restrictions are limited only to the suspicious processes and cache sets. It can be implemented with a maximum storage overhead of 1.45% of the total Last-Level Cache (LLC) size. This article presents three variations of the proposed attack mitigation mechanism: RSPP, simplified-RSPP (S-RSPP) and corewise-RSPP (C-RSPP) with different hardware overheads. A full system simulator is used for evaluating the performance impact of RSPP. A detailed experimental analysis with different LLC and attack parameters is also discussed. RSPP is also compared with the existing defense mechanisms effective against cross-core covert channel attacks. Jaspinder Kaur, Shirshendu Das |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | TPPD: Targeted Pseudo Partitioning based Defence for cross-core covert channel attacks
Jaspinder Kaur, Shirshendu Das |
J. Syst. Archit. | 2 |
| 2022 | WinDRAM: Weak rows as in-DRAM cacheabstractSummary The primary factor responsible for increasing the refresh rate is the presence of weak rows in a DRAM. They have a shorter retention time and lose charge faster than regular rows. Recently, a technique known as in‐DRAM cache was introduced in which some DRAM rows act as a separate module. The in‐DRAM cache can be used for a variety of purposes in DRAM. We present WinDRAM, an in‐DRAM cache comprised of all the DRAM's weak rows. The most recently accessed rows are copied into the in‐DRAM cache so that when the row is accessed again, both rows (original and copy) can be activated at the same time. Such simultaneous activation reduces activation time and, as a result, DRAM access latency. Dual‐row activation is the term for this concept. Because weak rows are part of the in‐DRAM cache and are frequently accessed, WinDRAM does not perform a periodic refresh on them. Existing techniques based on in‐DRAM cache do not design the in‐DRAM cache using weak rows. WinDRAM proposes a novel idea by designing the in‐DRAM cache using weak rows. Because the weak rows do not need to be refreshed, the refresh interval of the remaining rows can be increased, resulting in a refresh rate reduction of 80% to 90%. The speedup is 15% to 25% faster than standard DRAM and 12.77% faster than previous work for high memory‐intensive workloads. Overall energy consumption is also reduced by 10% to 15%. Sudershan Kumar, Prabuddha Sinha, Shirshendu Das |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Process variation aware DRAM-Cache resizing
Bindu Agarwalla, Shirshendu Das, Nilkanta Sahu |
J. Syst. Archit. | 2 |
| 2021 | Exploiting Secrets by Leveraging Dynamic Cache Partitioning of Last Level CacheabstractDynamic cache partitioning for shared Last Level Caches (LLC) is deployed in most modern multicore systems to achieve process isolation and fairness among the applications and avoid security threats. Since LLC has visibility of all cache blocks requested by several applications running on a multicore system, a malicious application can potentially threaten the system that can leverage the dynamic partitioning schemes applied to the LLCs by creating a timing-based covert channel attack. We call it as Cache Partitioned Covert Channel (CPCC) attack. The malicious applications may contain a trojan and a spy and use the underlying shared memory to create the attack. Through this attack, secret pieces of information like encryption keys or any secret information can be transmitted between the intended parties. We have observed that CPCC can target single or multiple cache sets to achieve a higher transmission rate with a maximum error rate of 5% only. The paper also addresses a few defense strategies that can avoid such cache partitioning based covert channel attacks. Anurag Agarwal, Jaspinder Kaur, Shirshendu Das |
DATE | 3 |
| 2021 | A Fairness Conscious Cache Replacement Policy for Last Level CacheabstractMulticore systems with shared Last Level Cache (LLC) possess a bigger challenge in allocating the LLC space among multiple applications running in the system. Since all applications use the shared LLC, interference caused by them may evict important blocks of other applications that result in premature eviction and may also lead to thrashing. Replacement policies applied locally to a set distributes the sets in a dynamic way among the applications. However, previous work on replacement techniques focused on the re-reference aspect of a block or application behavior to improve the overall system performance. The paper proposes a novel cache replacement technique Application Aware Re-reference Interval Prediction (AARIP) that considers application behavior, re-reference interval, and premature block eviction for replacing a cache block. Experimental evaluation on a four-core system shows that AARIP achieves an overall performance improvement of 7.28%, throughput by 4.9%, and improves overall system fairness by 7.85%, as compared to the traditional SRRIP replacement policy. Kousik Kumar Dutta, Prathamesh Nitin Tanksale, Shirshendu Das |
DATE | 3 |
| 2021 | Efficient Cache Resizing policy for DRAM-based LLCs in ChipMultiprocessors
Bindu Agarwalla, Shirshendu Das, Nilkanta Sahu |
J. Syst. Archit. | 2 |
| 2021 | Towards Enhanced System Efficiency while Mitigating Row HammerabstractIn recent years, DRAM-based main memories have become susceptible to the Row Hammer (RH) problem, which causes bits to flip in a row without accessing them directly. Frequent activation of a row, called an aggressor row , causes its adjacent rows’ ( victim ) bits to flip. The state-of-the-art solution is to refresh the victim rows explicitly to prevent bit flipping. There have been several proposals made to detect RH attacks. These include both probabilistic as well as deterministic counter-based methods. The technique of handling RH attacks, however, remains the same. In this work, we propose an efficient technique for handling the RH problem. We show that the mechanism is agnostic of the detection mechanism. Our RH handling technique omits the necessity of refreshing the victim rows. Instead, we use a small non-volatile Spin-Transfer Torque Magnetic Random Access Memory (STTRAM) that ensures no unnecessary refreshes of the victim rows on the DRAM device and thus allowing more time for normal applications in the same DRAM device. Our model relies on the migration of the aggressor rows. This accounts for removing blocking of the DRAM operations due to the refreshing of victim rows incurred in the previous solution. After extensive evaluation, we found that, compared to the conventional RH mitigation techniques, our model minimizes the blocking time of the memory that is imposed due to explicit refreshing by an average of 80.72% in the worst-case scenario and provides energy savings of about 15.82% on average, across different types of RH-based workloads. A lookup table is necessary to pinpoint the location of a particular row, which, when combined with the STTMRAM, limits the storage overhead to 0.39% of a 2 GB DRAM. Our proposed model prevents repeated refreshing of the same victim rows in different refreshing windows on the DRAM device and leads to an efficient RH handling technique. Kaustav Goswami 0002, Dip Sankar Banerjee, Shirshendu Das |
ACM Trans. Archit. Code Optim. | 3 |
| 2019 | Cost effective routing techniques in 2D mesh NoC using on-chip transmission lines
Dipika Deb, John Jose, Shirshendu Das, Hemangee K. Kapoor |
J. Parallel Distributed Comput. | 3 |
| 2017 | Dynamic Associativity Management in Tiled CMPs by Runtime Adaptation of Fellow SetsabstractThe non-uniform distribution of memory accesses among the cache sets results in some sets being used heavily while certain others remaining underutilized. Dynamic associativity management (DAM) is a technique to allow the heavily used sets to distribute their load among the lightly used sets thus improving the overall utilization of the cache. CMP-SVR is a previously proposed DAM based technique, where each set is divided into two sections: normal storage (NT) and reserve storage (RT). Some number of ways (25 to 50 percent) from each set are reserved for RT and the remaining ways belong to NT. The sets are divided into groups called fellow-groups and a set can use the reserve-ways of its fellow sets to increase its associativity during execution. Though CMP-SVR improves performance the formation of its fellow-groups is static: once created it never changes. It has been observed that some fellow-groups have more number of heavily used sets than the other fellow-groups. As a result the cache loads are not uniformly distributed among the fellow-groups. Also the behavior of sets changes dynamically: a lightly used set may become heavily used after a number of execution cycles. This paper studies the behavior of each set in detail and proposes a DAM based technique which improves the performance compared to other DAM based techniques. The proposed technique called FS-DAM dynamically creates fellow-groups based on the current set loads ensuring that the heavily used sets are evenly distributed among all the fellow-groups. Such distribution increases the utilization of the cache and hence improves performance. Full system simulation shows an average of 6.62 and 16.74 percent improvements, in FS-DAM as compared to CMP-SVR, in terms of CPI (Cycles Per Instruction) and MPKI (Miss Per Thousand Instructions) respectively. Comparing with Z-Cache the improvements are 6.21 percent (CPI) and 14.65 percent (MPKI). The proposed policy also shows better performance over V-Way and SBC. Shirshendu Das, Hemangee K. Kapoor |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | A Framework for Block Placement, Migration, and Fast Searching in Tiled-DNUCA ArchitectureabstractMulticore processors have proliferated several domains ranging from small-scale embedded systems to large data centers, making tiled CMPs (TCMPs) the essential next-generation scalable architecture. NUCA architectures help in managing the capacity and access time for such larger cache designs. It divides the last-level cache (LLC) into multiple banks connected through an on-chip network. Static NUCA (SNUCA) has a fixed address mapping policy, whereas dynamic NUCA (DNUCA) allows blocks to relocate nearer to the processing cores at runtime. To allow this, DNUCA divides the banks into multiple banksets and a block can be placed in any bank within a particular bankset. The entire bankset may need to be searched to access a block. Optimal bankset searching mechanisms are essential for getting the benefits from DNUCA. This article proposes a DNUCA-based TCMP architecture called TLD-NUCA. It reduces the LLC access time of TCMP and also allows a heavily loaded bank to distribute its load among the underused banks. Instead of other DNUCA designs, TLD-NUCA considers larger banksets. Such relaxations result in more uniform load distribution than existing DNUCA-based TCMP (T-DNUCA). Considering larger banksets improves the utilization factor, but T-DNUCA cannot implement it because of its expensive searching mechanism. TLD-NUCA uses a centralized directory, called TLD, to search a block from all the banks. Also, the proposed block placement policy reduces the instances when the central TLD needs to be contacted. It does not require the expensive simultaneous search as needed by T-DNUCA. Better cache utilization and a reduction in LLC access time improve the miss rate as well as the average memory access time (AMAT). Improving the miss rate and AMAT results in improvements in cycles per instructions (CPI). Experimental analysis found that TLD-NUCA improves performance by 6.5% as compared to T-DNUCA. The improvement is 13% as compared to the SNUCA-based TCMP design. Shirshendu Das, Hemangee K. Kapoor |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2013 | Towards a Better Cache Utilization Using Controlled Cache PartitioningabstractMany multi-core processors nowadays employ a shared Last Level Cache (LLC). Partitioning LLC becomes more important as LLC is shared among the cores. Past research has demonstrated that the traditional least recently used (LRU) based partitioning cum replacement policy has adverse effects on parameters like instruction per cycle (IPC), miss rate and speedup. This leads to poor performance in an environment when multiple cores compete for one global LLC. Applications, enjoying locality of reference are purely benefited by LRU, however LRU fails for the applications showing working set size (WSS) large than the LLC size. In this work, we propose a scheme which allows cores to steal/donate their lines upto a threshold and give them a chance to adjust their partition when there is a miss. Instead of maintaining strict target partitioning, we introduce a flexible threshold window. Our evaluation with multiprogrammed workloads shows significant performance improvement. Prateek D. Halwe, Shirshendu Das, Hemangee K. Kapoor |
DASC | 2 |
| 2013 | A formal framework for interfacing mixed-timing systems
Shirshendu Das, Parasara Sridhar Duggirala, Hemangee K. Kapoor |
Integr. | 1 |
| 2013 | Design and formal verification of a hierarchical cache coherence protocol for NoC based multiprocessors
Hemangee K. Kapoor, Praveen Kanakala, Malti Verma, Shirshendu Das |
J. Supercomput. | 4 |