EDBT 2026 Demo / reviewers in the wild / expert
T. Venkata Kalyan
dblp:19/574 · also Venkata Kalyan Tavva
· DBLP profile ↗
12ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-6596-2292ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Strata: Proactive Page Placement in Hybrid Memory SystemsabstractHybrid memory systems that integrate high-bandwidth memory (HBM) with off-chip DRAM employ page migration mechanisms to balance performance and capacity. The state-of-the-art technique Bumblebee introduces a hotness-based page allocation mechanism in which new pages are placed in HBM only if they remain present in an HBM hot-page queue and if free HBM capacity is available; otherwise, pages are allocated in off-chip DRAM. While this approach can potentially reduce unnecessary migrations and HBM pollution, it remains conservative with respect to initial page placement. Our proposal, Strata , proactively assigns newly allocated pages to HBM, rather than deferring HBM placement until sufficient hotness history is observed. This increases the fraction of early accesses served by HBM, reducing memory access latency and improving overall IPC by \(27.18\%\) over Bumblebee. Upasna, T. Venkata Kalyan |
CF | 2 |
| 2025 | SmartDeCoup: Decoupling the STT-RAM LLC for even write distribution and lifetime improvement
Prabuddha Sinha, Krishna Prathik B. V., Shirshendu Das, T. Venkata Kalyan |
J. Syst. Archit. | 4 |
| 2025 | Selective Subarray Isolation for Mitigating RowHammer AttackabstractRowHammer is a severe circuit-level vulnerability in DRAM-based main memories that allows attackers to flip the bits stored in DRAM rows by repeatedly accessing the nearby rows. Due to density scaling, newer generation DRAM chips are found to be increasingly more vulnerable to RowHammer attacks, motivating researchers from both academia and industry to come up with new RowHammer attack patterns and mitigation strategies that can be widely adopted. However, the question remains whether the mitigation strategies available now can secure DRAM-based memory in the future. We propose three approaches to mitigate RowHammer attacks by exploiting subarray isolation. A subarray is a collection of DRAM rows in a DRAM bank where each subarray operates independently. In the first approach, known as Subarray Isolation (SI), data from different domains are allocated to separate subarrays in DRAM. The SI strategy naively allocates subarrays to domains, greatly hampering the bank-level parallelism in memory accesses, leading to a significant performance loss. The second approach, namely, Selective Subarray Isolation (SSI), improves this aspect. With the SSI strategy, we allocate only confidential data from different domains to separate subarrays. The non-confidential data of the domains will share the subarrays as in the conventional case. Our evaluations show that the SSI strategy performs better compared to state-of-the-art mitigation strategies when the amount of confidential data is less. To further improve performance, we propose the third approach, namely Finer Selective Subarray Isolation (FSSI), which allocates separate partitions protected with guard rows within a subarray to confidential data from different domains. Our evaluations show that, of the three approaches, the FSSI strategy performs the best. Compared to baseline without any RowHammer protection, the FSSI strategy experiences an average performance drop of 0.89% for 50% of confidential data, but for 10% and 20% of confidential data, it shows an improvement of 1.43% and 1.28%, respectively. We also observe that the FSSI strategy is the most energy efficient among the state-of-the-art RowHammer mitigation techniques. Note that all our proposed strategies do not incur hardware overhead for performing RowHammer mitigation. Praseetha M, Madhu Mutyam, T. Venkata Kalyan |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Cache Line Pinning for Mitigating Row Hammer AttackabstractRowHammer attack is a serious security threat to DRAM-based memory that causes bit flips in nearby rows when a DRAM row is accessed frequently. Many mitigation strategies are proposed against the RowHammer attack, and a few of the mitigation strategies are adopted and implemented by the hardware vendors. But even the latest generations of DRAM-based memory with in-DRAM mitigation are found vulnerable to the RowHammer attack. Praseetha M, Madhu Mutyam, T. Venkata Kalyan |
ICPP | 3 |
| 2022 | Techniques to Improve Write and Retention Reliability of STT-MRAM Memory SubsystemabstractSpin transfer torque magneto-resistive random-access memory (STT-MRAM) has many advantages, such as scalability, persistence, practically infinite endurance, and fast access speed, that make it a promising and emerging technology for memory. However, this technology has multiple reliability issues, such as read and write reliability, higher write power, and long write latency, etc. At elevated temperatures, these issues exacerbate further. As the temperature increases massively in the latest compute nodes, we need to study and understand the effect of temperature on STT-MRAM memory writes and reliability. In this article, we propose the temperature-aware memory controller (MC) and device architecture techniques specific to STT-MRAM technology, which can improve write reliability, retention reliability, and memory power without sacrificing the performance. Our simulation results show that the proposed techniques cumulatively improve the write bit error rate (BER) on an average by$603\times $, increase retention reliability by 65%, along with 27% power reduction and 5.8% improved system performance over the baseline STT-MRAM-based memory subsystem. Saravanan Sethuraman, T. Venkata Kalyan, M. B. Srinivas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory SubsystemabstractSpin-transfer torque magneto-resistive random-access memory (STT-MRAM) is an exciting new emerging technology, being considered as a strong candidate to fill the gaps in the existing memory hierarchy between DRAM and the secondary memory. STT-MRAM has adequate endurance. However, unresolved write switching and read reliability issues still exist at the functional operating temperature corners. One biggest challenge is that the read bit error rate (RBER) is not at an acceptable level for system reliability across the wide operating temperature range. We present an STT-MRAM memory subsystem that is fully compatible with existing DDR-based DIMM designs and evaluate read disturb and read sense bit-error rate (BER) under various operating temperature conditions. We propose temperature aware adaptive techniques for reliable reads at the rank level. The proposed temperature adaptation technique improves overall reliability of the DDR4 STT-MRAM-based memory subsystem with an optimal read current considering an acceptable 64-byte cacheline BER. Our full system simulations show 1000× order of improvements toward a cell raw read disturb BER along with 5% reduction in memory power and less than 1% impact on overall system performance. Saravanan Sethuraman, T. Venkata Kalyan, Karthick Rajamani, Chitra K. Subramanian, Kyu-Hyoun Kim, Hillery C. Hunter, M. B. Srinivas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | A Scalable and Energy-Efficient Concurrent Binary Search Tree With FatnodesabstractIn the recent past, devising algorithms for concurrent data structures has been driven by the need for scalability. Further, there is an increased traction across the industry towards power efficient concurrent data structure designs. In this context, we introduce a scalable and energy-efficient concurrent binary search tree with fatnodes (namely, FatCBST), and present algorithms to perform basic operations on it. Unlike a single node with one value, a fatnode consists of a set of values. FatCBST minimizes structural changes while performing update operations on the tree. In addition, fatnodes help to exploit the spatial locality in the cache hierarchy and also reduce the height of the tree. FatCBST allows multiple threads to perform update operations on an existing fatnode simultaneously. Experimental results show that for low contention workloads as well as large set sizes, FatCBST scales well and also provides high performance-per-watt values as compared to the state-of-the-art implementations. For high contention workloads with small set sizes, FatCBST suffers from contention. Praveen Alapati, T. Venkata Kalyan, Madhu Mutyam |
IEEE Trans. Sustain. Comput. | 2 |
| 2014 | Data remapping for an energy efficient burst chop in DRAM memory systemsabstractIn modern day systems, main memory contributes significantly to the overall power consumption. One of the features provided by JEDEC DDR3 standard onwards is Burst Chop (BC) through which the Burst Length of the data access commands (CAS) can be configured. This work aims to improve the energy efficiency of the DRAM memory by exploiting the existing BC features for half writes (writes in which either the first half or second half of the cache block is dirty). We propose to change the mapping of words of a cache block to the DRAM devices in order to reduce the number of devices involved in half writes. With our new mapping, we achieve average memory power savings of 3.27% with negligible impact on performance. Sudharsan Jagathrakshakan, T. Venkata Kalyan, Madhu Mutyam |
PACT | 2 |
| 2014 | Scattered refresh: An alternative refresh mechanism to reduce refresh cycle timeabstractWith realization of high density DRAM devices, the amount of time spent in refreshing a DRAM bank is increasing. This reduces the availability of the bank to the requests from the processing cores, leading to degradation in performance. In this work we target to reduce the refresh cycle time of the DRAM device by scattering the rows in a refresh operation to different subarrays and leveraging the available parallelism in their access. Considering 8Gb devices, we show that Scattered Refresh achieves up to 10.2% of overall system performance improvement. Scattered Refresh, being orthogonal to the existing refresh handling techniques, can be employed along with any of them, boosting their effectiveness further. T. Venkata Kalyan, Ravi Kasha, Madhu Mutyam |
ASP-DAC | 1 |
| 2014 | SFFMap: Set-First Fill mapping for an energy efficient pipelined data cacheabstractConventionally, consecutively addressed blocks are mapped onto different sets in cache. In this work, we propose a new block address mapping, Set-First Fill (SFFMap), for pipelined L1 data caches wherein consecutively addressed data blocks are mapped onto the same set. This increases the inter-block spatial locality within the cache set. In order to exploit SFFMap, we propose to store and if possible, access the most recently used set in the cache's pipeline registers. Further, selective access (SSA) and selective update (SSU) techniques are proposed for set-buffer to increase the effectiveness of SFFMap. Our experimental evaluation for in-order and out-of-order processors with an 8-way set-associative data cache shows that SFFMap, together with SSA and SSU, achieves around 27% reduction in dynamic energy and 4-5% performance improvement. The proposed techniques need minor modifications to the existing hardware, making it an adoptable design. Pritam Majumder, T. Venkata Kalyan, Madhu Mutyam |
ICCD | 2 |
| 2014 | EFGR: An Enhanced Fine Granularity Refresh Feature for High-Performance DDR4 DRAM DevicesabstractHigh-density DRAM devices spend significant time refreshing the DRAM cells, leading to performance drop. The JEDEC DDR4 standard provides a Fine Granularity Refresh (FGR) feature to tackle refresh. Motivated by the observation that in FGR mode, only a few banks are involved, we propose an Enhanced FGR (EFGR) feature that introduces three optimizations to the basic FGR feature and exposes the bank-level parallelism within the rank even during the refresh. The first optimization decouples the nonrefreshing banks. The second and third optimizations determine the maximum number of nonrefreshing banks that can be active during refresh and selectively precharge the banks before refresh, respectively. Our simulation results show that the EFGR feature is able to recover almost 56.6% of the performance loss incurred due to refresh operations. T. Venkata Kalyan, Ravi Kasha, Madhu Mutyam |
ACM Trans. Archit. Code Optim. | 1 |
| 2008 | Word-interleaved cache: an energy efficient data cache architectureabstractWe propose a novel energy-efficient data cache architecture, namely, word-interleaved (WI) cache. In theWI cache, a cache block is distributed uniformly among the different cache ways and each line of a cache way holds some words of the block. This distribution provides an opportunity to activate/deactivate the cache ways based on the requested address's offset, thus minimizing the overall cache access energy. For a 4-way set associative cache of size 16KB and blocksize 32B, the proposed technique accomplishes dynamic energy savings of 54.2% without considering fast hits and 62.3% when fast hits are considered, with small performance degradation and negligible area overhead. T. Venkata Kalyan, Madhu Mutyam |
ISLPED | 1 |