EDBT 2026 Demo / reviewers in the wild / expert
Madhu Mutyam
dblp:m/MadhuMutyam
· DBLP profile ↗
47ranked-venue papers
10as first author
4since 2021 · last 2025
0000-0003-1638-4195ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorTheory of computation · 5 · 4 first-authorArtificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Selective Subarray Isolation for Mitigating RowHammer AttackabstractRowHammer is a severe circuit-level vulnerability in DRAM-based main memories that allows attackers to flip the bits stored in DRAM rows by repeatedly accessing the nearby rows. Due to density scaling, newer generation DRAM chips are found to be increasingly more vulnerable to RowHammer attacks, motivating researchers from both academia and industry to come up with new RowHammer attack patterns and mitigation strategies that can be widely adopted. However, the question remains whether the mitigation strategies available now can secure DRAM-based memory in the future. We propose three approaches to mitigate RowHammer attacks by exploiting subarray isolation. A subarray is a collection of DRAM rows in a DRAM bank where each subarray operates independently. In the first approach, known as Subarray Isolation (SI), data from different domains are allocated to separate subarrays in DRAM. The SI strategy naively allocates subarrays to domains, greatly hampering the bank-level parallelism in memory accesses, leading to a significant performance loss. The second approach, namely, Selective Subarray Isolation (SSI), improves this aspect. With the SSI strategy, we allocate only confidential data from different domains to separate subarrays. The non-confidential data of the domains will share the subarrays as in the conventional case. Our evaluations show that the SSI strategy performs better compared to state-of-the-art mitigation strategies when the amount of confidential data is less. To further improve performance, we propose the third approach, namely Finer Selective Subarray Isolation (FSSI), which allocates separate partitions protected with guard rows within a subarray to confidential data from different domains. Our evaluations show that, of the three approaches, the FSSI strategy performs the best. Compared to baseline without any RowHammer protection, the FSSI strategy experiences an average performance drop of 0.89% for 50% of confidential data, but for 10% and 20% of confidential data, it shows an improvement of 1.43% and 1.28%, respectively. We also observe that the FSSI strategy is the most energy efficient among the state-of-the-art RowHammer mitigation techniques. Note that all our proposed strategies do not incur hardware overhead for performing RowHammer mitigation. Praseetha M, Madhu Mutyam, T. Venkata Kalyan |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | Cache Line Pinning for Mitigating Row Hammer AttackabstractRowHammer attack is a serious security threat to DRAM-based memory that causes bit flips in nearby rows when a DRAM row is accessed frequently. Many mitigation strategies are proposed against the RowHammer attack, and a few of the mitigation strategies are adopted and implemented by the hardware vendors. But even the latest generations of DRAM-based memory with in-DRAM mitigation are found vulnerable to the RowHammer attack. Praseetha M, Madhu Mutyam, T. Venkata Kalyan |
ICPP | 2 |
| 2024 | Selective Memory Compression for GPU Memory Oversubscription ManagementabstractUnified virtual memory for discrete GPUs helps increase programmer productivity by abstracting the presence of different memories. However, fault-driven migrations and remote fault handling slow down the applications considerably. During memory oversubscription, the overheads increase drastically due to increased page thrashing and page evictions. We propose a GPU memory compression system, Selective Memory Compression (SMC), that selectively compresses read-only pages to increase the effective memory size while avoiding costly page remappings due to page overflows. We also propose a line packing scheme for compressed pages, Split Linearly Compressed Pages (SLCP), that minimizes unused space, gives better compressibility, and reduces extra memory accesses performed to fetch data in a compressed memory system. We show that under 125% and 150% oversubscription, SMC combined with SLCP gives 53% and 60% performance improvement, respectively, over a baseline that uses the state-of-the-art eviction policy. Abdun Nihaal, Madhu Mutyam |
ICPP | 2 |
| 2023 | Formal Modeling and Verification of Security Properties of a Ransomware-Resistant SSDabstractSolid-state drives (SSDs) are fast emerging as the primary choice for data storage in diverse domains. However, data protection against ransomware attacks on such storage devices is a crucial challenge. A recent research work proposed an In-SSD ransomware protection mechanism to improve SSD data security. The SSD flash translation layer (FTL), which consists of an address mapping, garbage collection, and wear leveling components, is inherently the most complex part of an SSD controller. The In-SSD ransomware protection mechanism adds a new component to the FTL, making the FTL design more complex. The new component interacts with almost all other units of FTL for recovery from a ransomware attack. Such a design would naturally require rigorous verification, which is the focus of this article. In this work, we discuss the derivation of a set of security properties related to protection from ransomware that covers the security requirements and the design; we next prove such properties using symbolic model checking. Shivani Tripathy, Debiprasanna Sahoo, Manoranjan Satpathy, Madhu Mutyam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Router Buffer Caching for Managing Shared Cache Blocks in Tiled Multi-Core ProcessorsabstractMultiple cores in a tiled multi-core processor are connected using a network-on-chip mechanism. All these cores share the last-level cache (LLC). For large-sized LLCs, generally, non-uniform cache architecture design is considered, where the LLC is split into multiple slices. Accessing highly shared cache blocks from an LLC slice by several cores simultaneously results in congestion at the LLC, which in turn increases the access latency. To deal with this issue, we propose a congestion management technique in the LLC that equips the NoC router with small storage to keep a copy of heavily shared cache blocks. To identify highly shared cache blocks, we also propose a prediction classifier in the LLC controller. We implement our technique in Sniper, an architectural simulator for multi-core systems, and evaluate its effectiveness by running a set of parallel benchmarks. Our experimental results show that the proposed technique is effective in reducing the LLC access time. Joe Augustine, Kanakagiri Raghavendra, John Jose, Madhu Mutyam |
ICCD | 4 |
| 2020 | Fuzzy fairness controller for NVMe SSDsabstractModern NVMe SSDs are widely deployed in diverse domains due to characteristics like high performance, robustness, and energy efficiency. It has been observed that the impact of interference among the concurrently running workloads on their overall response time differs significantly in these devices, which leads to unfairness. Workload intensity is a dominant factor influencing the interference. Prior works use a threshold value to characterize a workload as high-intensity or low-intensity; this type of characterization has drawbacks due to lack of information about the degree of low- or high-intensity. Shivani Tripathy, Debiprasanna Sahoo, Manoranjan Satpathy, Madhu Mutyam |
ICS | 4 |
| 2020 | Optimization of Intercache Traffic Entanglement in Tagless Caches With Tiling OpportunitiesabstractSo-called “tagless” caches have become common as a means to deal with the vast L4 last-level caches (LLCs) enabled by increasing device density, emerging memory technologies, and advanced integration capabilities (e.g., 3-D). Tagless schemes often result in intercache entanglement between tagless cache (L4) and the cache (L3) stewarding its metadata. We explore new cache organization policies that mitigate overheads stemming from the intercache-level replacement entanglement. We incorporate support for explicit tiling shapes that can better match software access patterns to improve the spatial and temporal locality of large block allocations in many essential computational kernels. To address entanglement overheads and pathologies, we propose new replacement policies and energy-friendly mechanisms for tagless LLCs, such as restricted block caching (RBC) and victim tag buffer caching (VBC) to incorporate L4 eviction costs into L3 replacement decisions efficiently. We evaluate our schemes on a range of linear algebra kernels that are software tiled. RBC and VBC demonstrate a reduction in memory traffic of 83/4.4/67% and 69/35.5/76% for 8/32/64 MB L4s, respectively. Besides, RBC and VBC provide speedups of 16/0.3/0.6% and 15.7/1.8/0.8%, respectively, for systems with 8/32/64 MB L4, over a tagless cache with an LRU policy in the L3. We also show that matching the shape of the hardware allocation for each tagless region superblocks to the access order of the software tile improves latency by 13.4% over the baseline tagless cache with reductions in memory traffic of 51% over linear superblocks. S. R. Swamy Saranam Chongala, Sumitha George, Hariram Thirucherai Govindarajan, Jagadish Kotra, Madhu Mutyam, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | A Scalable and Energy-Efficient Concurrent Binary Search Tree With FatnodesabstractIn the recent past, devising algorithms for concurrent data structures has been driven by the need for scalability. Further, there is an increased traction across the industry towards power efficient concurrent data structure designs. In this context, we introduce a scalable and energy-efficient concurrent binary search tree with fatnodes (namely, FatCBST), and present algorithms to perform basic operations on it. Unlike a single node with one value, a fatnode consists of a set of values. FatCBST minimizes structural changes while performing update operations on the tree. In addition, fatnodes help to exploit the spatial locality in the cache hierarchy and also reduce the height of the tree. FatCBST allows multiple threads to perform update operations on an existing fatnode simultaneously. Experimental results show that for low contention workloads as well as large set sizes, FatCBST scales well and also provides high performance-per-watt values as compared to the state-of-the-art implementations. For high contention workloads with small set sizes, FatCBST suffers from contention. Praveen Alapati, T. Venkata Kalyan, Madhu Mutyam |
IEEE Trans. Sustain. Comput. | 3 |
| 2019 | POSTER: Variable Sized Cache-Block CompactionabstractData blocks compressed to different sizes can be stored together inside a single cache-block to increase space utilization. However, the lack of a common size offset makes it challenging to locate individual blocks without additional tag overhead. We propose Variable Sized Cache-Block Compaction (VSCC) that allows us to store variable sized compressed blocks together and locate them inside a cache-block by using their compression encodings - available inside the tag metadata. We introduce a novel read/write scheme and a new BDI compression encoding, which reduce the necessary operations by 50%. Experimental results reveal that VSCC outperforms state-of-the-art techniques from the performance and energy point of view while keeping the storage overheads within acceptable limits. Sayantan Ray, Madhu Mutyam |
PACT | 2 |
| 2019 | Post-Model Validation of Victim DRAM CachesabstractFormal modeling and analysis of a victim DRAM cache has already been discussed in the existing literature. These works use interacting state machines to model states and transitions of a victim DRAM cache. In this work, we address model-code conformance between a formal model of the victim DRAM cache and a simulator obtained from it. Our work focuses on a two-step approach to validate a DRAM cache implementation. In the first step, we use a technique based on it Feedback-Directed Random Testing to reverse engineer the state models from the execution traces. This process helps us to match the implementation with the state machines associated with the formal model and to verify some of the state-based properties. In the second step, we instrument the implementation with monitors and validate the liveness and safety properties (during run-time) which have been proved earlier in the formal model. Debiprasanna Sahoo, Shivani Tripathy, Manoranjan Satpathy, Madhu Mutyam |
ICCD | 4 |
| 2019 | Formal Modeling and Verification of a Victim DRAM CacheabstractThe emerging Die-stacking technology enables DRAM to be used as a cache to break the “Memory Wall” problem. Recent studies have proposed to use DRAM as a victim cache in both CPU and GPU memory hierarchies to improve performance. DRAM caches are large in size and, hence, when realized as a victim cache, non-inclusive design is preferred. This non-inclusive design adds significant differences to the conventional DRAM cache design in terms of its probe, fill, and writeback policies. Design and verification of a victim DRAM cache can be much more complex than that of a conventional DRAM cache. Hence, without rigorous modeling and formal verification, ensuring the correctness of such a system can be difficult. The major focus of this work is to show how formal modeling is applied to design and verify a victim DRAM cache. In this approach, we identify the agents in the victim DRAM cache design and model them in terms of interacting state machines. We derive a set of properties from the specifications of a victim cache and encode them using Linear Temporal Logic. The properties are then proven using symbolic and bounded model checking. Finally, we discuss how these properties are related to the dataflow paths in a victim DRAM cache. Debiprasanna Sahoo, Swaraj Sha, Manoranjan Satpathy, Madhu Mutyam, S. Ramesh 0002, Partha S. Roop |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2018 | CAMO: A novel cache management organization for GPGPUsabstractGPGPUs are now commonly used as co-processors of CPUs for the computation of data parallel and throughputintensive algorithms. However, memory available in GPGPUs is limited for many applications of interest; there is a continuous demand for increased memory of such applications. Several techniques like multi-steaming or pinned memory are frequently employed to mitigate these issues to some extent. However, these techniques either suffer from latency overhead or increase programming complexity. GPUdmm uses GPU DRAM as a cache of CPU; key problems in this design are inefficient memory access data-path and tag access overhead. In this context, we present CAMO, a novel cache memory organization for GPGPUs which addresses the limitations of pinned memory technique and GPUdmm. First, it uses GPU DRAM as a victim cache of LLC that improves the performance by delivering data faster to the SMs. Second, it uses ATCache, a CPU based DRAM cache tag management technique. ATCache reduces the number of DRAM cache accesses. We implement CAMO within the GPGPU-Sim framework and show that its average performance - when compared with pinned memory - increases by a factor of 1.87x and the peak performance growth being 4.67x. In addition, CAMO outperforms GPUdmm on an average by a factor of 15.9% and maximum speedup by a factor of 80%. Debiprasanna Sahoo, Swaraj Sha, Manoranjan Satpathy, Madhu Mutyam, Laxmi N. Bhuyan |
ASP-DAC | 4 |
| 2018 | Formal Modeling and Verification of Controllers for a Family of DRAM CachesabstractDie-stacking technology enables the use of a high density DRAM as a cache. Major processor vendors have recently started using these stacked DRAM modules as the last level cache of their products. These stacked DRAM modules provide high bandwidth with relatively low latency compared to the off-package DRAM modules. Recent studies on DRAM caches propose several variants to optimize performance and power of the systems. However, none of the existing works discuss its design and verification aspect. DRAM cache controller (DCC) design is significantly complex in comparison to a conventional DRAM-based main memory controller. This is because it involves controlling both the timing aspect of DRAM system as well as the functional aspect of cache. Therefore, without rigorous modeling and verification of such designs, it would be difficult to ensure correctness. In the current research, we focus on the design and verification issues of DCC. We select a common variant of DRAM cache and build a formal model of its controller in terms of interacting state machines; we term the common variant as the baseline and its model as the base model. We then verify safety, liveness, and timing properties of this variant using model checking. Next, we demonstrate how the formal models and the associated properties of other variants of DCCs can be derived from the base model in a systematic way. Analyzing the individual DRAM cache variations, we observe that most of the variants exhibit product-line characteristics. Debiprasanna Sahoo, Swaraj Sha, Manoranjan Satpathy, Madhu Mutyam, S. Ramesh 0002, Partha S. Roop |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Concurrent Treaps
Praveen Alapati, Swamy Saranam, Madhu Mutyam |
ICA3PP | 3 |
| 2017 | RCTP: Region Correlated Temporal PrefetcherabstractHardware prefetcher is an essential component of modern processors that helps in boosting system performance by fetching the data before processor demands for the same. Hardware prefetching techniques have been proposed to exploit various kinds of access patterns. However, there are applications that are highly irregular in nature that evolved in the past decade, and have massive memory footprint. Temporal prefetching techniques are effective in predicting the future addresses of these irregular applications. Prior works on temporal prefetching use large data structures to store temporal patterns and future memory accesses are predicted using these patterns. However, these techniques predict future accesses only when there is a pattern that is already trained for a given cache line address. To address this issue, we propose Region Correlated Temporal Prefetcher (RCTP). Our technique correlates temporal patterns of memory regions and predicts future accesses for the region whose patterns are yet to be populated. Thus, RCTP helps in predicting cache line addresses whose first access is yet to happen unlike the traditional temporal prefetchers. We evaluate RCTP on SPEC CPU 2006, CRONO, and PBBS benchmark suites. RCTP outperforms the state-of-the-art temporal prefetcher named ISB by 26%, and a recent delta prefetcher called VLDP by 6%. This improvement comes with a hardware overhead of 1KB over ISB; however, RCTP does not require off-chip storage, unlike other temporal prefetchers. Dennis Antony Varkey, Biswabandan Panda, Madhu Mutyam |
ICCD | 3 |
| 2017 | MBZip: Multiblock Data CompressionabstractCompression techniques at the last-level cache and the DRAM play an important role in improving system performance by increasing their effective capacities. A compressed block in DRAM also reduces the transfer time over the memory bus to the caches, reducing the latency of a LLC cache miss. Usually, compression is achieved by exploiting data patterns present within a block. But applications can exhibit data locality that spread across multiple consecutive data blocks. We observe that there is significant opportunity available for compressing multiple consecutive data blocks into one single block, both at the LLC and DRAM. Our studies using 21 SPEC CPU applications show that, at the LLC, around 25% (on average) of the cache blocks can be compressed into one single cache block when grouped together in groups of 2 to 8 blocks. In DRAM, more than 30% of the columns residing in a single DRAM page can be compressed into one DRAM column, when grouped together in groups of 2 to 6. Motivated by these observations, we propose a mechanism, namely, MBZip, that compresses multiple data blocks into one single block (called a zipped block), both at the LLC and DRAM. At the cache, MBZip includes a simple tag structure to index into these zipped cache blocks and the indexing does not incur any redirectional delay. At the DRAM, MBZip does not need any changes to the address computation logic and works seamlessly with the conventional/existing logic. MBZip is a synergistic mechanism that coordinates these zipped blocks at the LLC and DRAM. Further, we also explore silent writes at the DRAM and show that certain writes need not access the memory when blocks are zipped. MBZip improves the system performance by 21.9%, with a maximum of 90.3% on a 4-core system. Kanakagiri Raghavendra, Biswabandan Panda, Madhu Mutyam |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | PBC: Prefetched Blocks CompactionabstractCache compression improves the performance of a multi-core system by being able to store more cache blocks in a compressed format. Compression is achieved by exploiting data patterns present within a block. For a given cache space, compression increases the effective cache capacity. However, this increase is limited by the number of tags that can be accommodated at the cache. Prefetching is another technique that improves system performance by fetching the cache blocks ahead of time into the cache and hiding the off-chip latency. Commonly used hardware prefetchers, such as stream and stride, fetch multiple contiguous blocks into the cache. In this paper we propose prefetched blocks compaction (PBC) wherein we exploit the data patterns present across these prefetched blocks. PBC compacts the prefetched blocks into a single block with a single tag, effectively increasing the cache capacity. We also modify the cache organization to access these multiple cache blocks residing in a single block without any need for extra tag look-ups. PBC improves the system performance by 11.1 percent with a maximum of 43.4 percent on a four-core system. Kanakagiri Raghavendra, Biswabandan Panda, Madhu Mutyam |
IEEE Trans. Computers | 3 |
| 2014 | Data remapping for an energy efficient burst chop in DRAM memory systemsabstractIn modern day systems, main memory contributes significantly to the overall power consumption. One of the features provided by JEDEC DDR3 standard onwards is Burst Chop (BC) through which the Burst Length of the data access commands (CAS) can be configured. This work aims to improve the energy efficiency of the DRAM memory by exploiting the existing BC features for half writes (writes in which either the first half or second half of the cache block is dirty). We propose to change the mapping of words of a cache block to the DRAM devices in order to reduce the number of devices involved in half writes. With our new mapping, we achieve average memory power savings of 3.27% with negligible impact on performance. Sudharsan Jagathrakshakan, T. Venkata Kalyan, Madhu Mutyam |
PACT | 3 |
| 2014 | Scattered refresh: An alternative refresh mechanism to reduce refresh cycle timeabstractWith realization of high density DRAM devices, the amount of time spent in refreshing a DRAM bank is increasing. This reduces the availability of the bank to the requests from the processing cores, leading to degradation in performance. In this work we target to reduce the refresh cycle time of the DRAM device by scattering the rows in a refresh operation to different subarrays and leveraging the available parallelism in their access. Considering 8Gb devices, we show that Scattered Refresh achieves up to 10.2% of overall system performance improvement. Scattered Refresh, being orthogonal to the existing refresh handling techniques, can be employed along with any of them, boosting their effectiveness further. T. Venkata Kalyan, Ravi Kasha, Madhu Mutyam |
ASP-DAC | 3 |
| 2014 | Minimally buffered single-cycle deflection routerabstractWith the drift from computation centric designs to communication centric designs in the Chip Multi Processor (CMP) era, the interconnect fabric is gaining more importance. An efficient NoC in terms of power, area and average flit latency has a huge impact on the overall performance of a CMP. In the current work, we propose MinBSD - a minimally buffered, single cycle, deflection router. It incorporates different operations (Injection, Ejection, Preemption, Re-injection) in a single module to handle the traffic effectively and ensures smooth flow of flits through router pipeline. It performs overlapped execution of independent operations. These factors not only make MinBSD to operate in a single cycle but also to reduce the critical path latency resulting in a faster interconnect network. Experimental results show that MinBSD reduces the average flit latency on real work loads, reduces die area and power consumption when compared to the existing state-of-the-art minimally buffered deflection routers. Gnaneswara Rao Jonna, John Jose, Rachana Radhakrishnan, Madhu Mutyam |
DATE | 4 |
| 2014 | SFFMap: Set-First Fill mapping for an energy efficient pipelined data cacheabstractConventionally, consecutively addressed blocks are mapped onto different sets in cache. In this work, we propose a new block address mapping, Set-First Fill (SFFMap), for pipelined L1 data caches wherein consecutively addressed data blocks are mapped onto the same set. This increases the inter-block spatial locality within the cache set. In order to exploit SFFMap, we propose to store and if possible, access the most recently used set in the cache's pipeline registers. Further, selective access (SSA) and selective update (SSU) techniques are proposed for set-buffer to increase the effectiveness of SFFMap. Our experimental evaluation for in-order and out-of-order processors with an 8-way set-associative data cache shows that SFFMap, together with SSA and SSU, achieves around 27% reduction in dynamic energy and 4-5% performance improvement. The proposed techniques need minor modifications to the existing hardware, making it an adoptable design. Pritam Majumder, T. Venkata Kalyan, Madhu Mutyam |
ICCD | 3 |
| 2014 | Using packet information for efficient communication in NoCsabstractMultithreaded workloads like SPLASH2 and PARSEC generate heavy network traffic. This traffic consists of different packets injected by various nodes at various points of time. A packet consists of three essential components viz., source, destination and data. In this work, we study the opportunity to increase the efficiency of the underlying network. We come up with novel methods to share various components of the packets present in a router at any time. We provide an analysis of the performance gains over contemporary optimization techniques. We conduct experiments on a 64 node setup as well as a 512 node setup and record the energy consumption, IPC gains, network latency and throughputs. We show that our technique outperforms the contemporary Hamiltonian routing by 8.43% and VCTM routing by 7.7% on an average, in terms of IPC speedup. Prasanna Venkatesh Rengasamy, Madhu Mutyam |
NOCS | 2 |
| 2014 | EFGR: An Enhanced Fine Granularity Refresh Feature for High-Performance DDR4 DRAM DevicesabstractHigh-density DRAM devices spend significant time refreshing the DRAM cells, leading to performance drop. The JEDEC DDR4 standard provides a Fine Granularity Refresh (FGR) feature to tackle refresh. Motivated by the observation that in FGR mode, only a few banks are involved, we propose an Enhanced FGR (EFGR) feature that introduces three optimizations to the basic FGR feature and exposes the bank-level parallelism within the rank even during the refresh. The first optimization decouples the nonrefreshing banks. The second and third optimizations determine the maximum number of nonrefreshing banks that can be active during refresh and selectively precharge the banks before refresh, respectively. Our simulation results show that the EFGR feature is able to recover almost 56.6% of the performance loss incurred due to refresh operations. T. Venkata Kalyan, Ravi Kasha, Madhu Mutyam |
ACM Trans. Archit. Code Optim. | 3 |
| 2014 | Implementation and Analysis of History-Based Output Channel Selection Strategies for Adaptive Routers in Mesh NoCsabstractThe efficiency and effectiveness of an adaptive router in an NoC-based multicore system is evaluated by the performance it achieves under varying inter-core communication traffic. A well-designed selection strategy plays an important role in an adaptive router to act upon dynamic traffic variations. The effectiveness of a selection strategy depends on what metric is used to represent congestion, how precisely this metric captures the actual congestion, and how much cost is involved in capturing the congestion on a real-time scale. Congestion is formed over a period of time due to cumulative and chain reaction effects. We propose novel history-based selection strategies that could be used with any adaptive, deadlock-free, minimal routing in mesh NoCs. Buffer occupancy time and rate of flit flow across reachable ports of neighboring routers in the recent past are captured, propagated, and maintained in a cost-effective way to compute the selection metric. Experimental results on real and synthetic workloads show that our proposed selection strategies significantly outperform state-of-the-art techniques. John Jose, Madhu Mutyam |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2013 | DeBAR: deflection based adaptive router with minimal bufferingabstractEnergy efficiency of the underlying communication framework plays a major role in the performance of multicore systems. NoCs with buffer-less routing are gaining popularity due to simplicity in the router design, low power consumption, and load balancing capacity. With minimal number of buffers, deflection routers evenly distribute the traffic across links. In this paper, we propose an adaptive deflection router, DeBAR, that uses a minimal set of central buffers to accommodate a fraction of mis-routed flits. DeBAR incorporates a hybrid flit ejection mechanism that gives the effect of dual ejection with a single ejection port, an innovative adaptive routing algorithm, and a selective flit buffering based on flit marking. Our proposed router design reduces the average flit latency and the deflection rate, and improves the throughput with respect to the existing minimally buffered deflection routers without any change in the critical path. John Jose, Bhawna Nayak, Damarla Kranthi Kumar, Madhu Mutyam |
DATE | 4 |
| 2013 | SLIDER: Smart Late Injection DEflection Router for mesh NoCsabstractNetwork-on-Chip (NoC) provides a scalable communication interface for processing cores in large multicore systems. An efficient NoC router should not only minimize the average packet latency of the network but also have minimum pipeline latency, area, and power. Area and power overheads are affecting the scalability and popularity of traditional input buffered routers. In this context minimally buffered deflection routers are emerging as a cost effective alternative. We propose SLIDER, Smart Late Injection DEflection Router, that uses side buffers for accommodating a fraction of deflected flits. The main contributions of this work are smart late injection and selective flit preemption. In SLIDER the injection stage is kept at the end of the router pipeline. This reduces the contention in the arbitration stage, eliminates unwanted intra-router movement of flits and effectively utilizes the idle output channels. We parallelize independent operations in the router pipeline and reduce the pipeline latency by 25%. Experimental results on synthetic and real workloads show that SLIDER reduces average flit latency, channel wastage, and deflection rate, and increases throughput in the network when compared to the state-of-the-art minimally buffered deflection routers. Bhawna Nayak, John Jose, Madhu Mutyam |
ICCD | 3 |
| 2012 | SkipCache: miss-rate aware cache managementabstractIt is common for computers to have multi-level caches. This piece of work revolves around one question: Are all levels needed by all applications during all phases of their execution?, especially in the multi programmed scenario where giving the entire cache to one application and depriving the other might actually increase the performance. Kanakagiri Raghavendra, Tripti S. Warrier, Madhu Mutyam |
PACT | 3 |
| 2012 | TRACKER: A low overhead adaptive NoC router with load balancing selection strategyabstractThe effectiveness of an adaptive router in a Network on Chip (NoC) is evaluated by the selection metric it uses and its impact on overall performance. In this paper, we propose a flit flow history based load balancing selection strategy that can be used in any adaptive routers for output port selection. Using this selection strategy, we propose an adaptive router TRACKER, that keeps track of flow of flits through all its ports and updates this tracked information to its neighbors in a cost effective manner. Routers make use of these flit flow estimates to compute a novel selection metric for output port selection of incoming flits. TRACKER outperforms the baseline adaptive router architectures using odd-even routing model with conventional selection metrics like count of free virtual channels, count of fluid buffers and buffer occupancy time at reachable downstream neighbors. John Jose, K. V. Mahathi, J. Shiva Shankar, Madhu Mutyam |
ICCAD | 4 |
| 2012 | Fibonacci Codes for Crosstalk AvoidanceabstractPropagation delay across long on-chip buses is significant when adjacent wires are transitioning in opposite direction (i.e., crosstalk transitions) as compared to transitioning in the same direction. By exploiting Fibonacci number system, we propose a family of Fibonacci coding techniques for crosstalk avoidance, relate them to some of the existing crosstalk avoidance techniques, and show how the encoding logic of one technique can be modified to generate codewords of the other technique. Madhu Mutyam |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Prevention flow-control for low latency torus networks-on-chipabstractThe challenge for on-chip networks is to provide low latency communication in a very low power budget. To reduce the latency and keep the simplicity of a mesh network, torus network is proposed. As torus networks have inherent circular dependency, additional effort is needed to prevent deadlock, even if deadlock free routing algorithms are used. Arpit Joshi, Madhu Mutyam |
NOCS | 2 |
| 2011 | Timing variation-aware scheduling and resource binding in high-level synthesisabstractDue to technological scaling, process variations have increased significantly, resulting in large variations in the delay of the functional units. Hence, the worst-case approach is becoming increasingly pessimistic in meeting a certain performance yield. The problem therefore is to increase the performance as much as possible while maintaining the desired yield. In this work, we introduce an integer linear programming (ILP) formulation for scheduling and resource binding in high-level synthesis (HLS) which tries to mitigate the effect of timing variations. In the presence of delay variations of resources, as chained resources can give a better latency and performance yield trade-off, instead of considering them independently, we consider external chaining of resources, that is, two or more resources are connected by external wiring, and exploit operation chaining. Without violating the yield constraints, the proposed ILP formulation chains two consecutive operations and binds these chained operations to chained resources for minimizing the overall latency of the schedule. Our ILP formulation also makes sure that two consecutive operations can be chained over multiple clock cycles so that it becomes possible to access the data in the middle of the chained operations at the start of the clock steps over which the operations are chained. By solving our ILP formulation using ILOG CPLEX, we show that our mechanism achieves lesser latency in most cases, compared to the no-chaining case. Significant performance improvement is achieved even for the 100% yield case, which has never been demonstrated in any published work, to the best of our knowledge. Kartikey Mittal, Arpit Joshi, Madhu Mutyam |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2009 | Process-Variation-Aware Adaptive Cache Architecture and ManagementabstractFabricating circuits that employ ever-smaller transistors leads to dramatic variations in critical process parameters. This in turn results in large variations in execution/access latencies of different hardware components. This situation is even more severe for memory components due to minimum-sized transistors used in their design. Current design methodologies that are tuned for the worst case scenarios are becoming increasingly pessimistic from the performance angle, and thus, may not be a viable option at all for future designs. This paper makes two contributions targeting on-chip data caches. First, it presents an adaptive cache management policy based on nonuniform cache access. Second, it proposes a latency compensation approach that employs several circuit-level techniques to change the access latency of select cache lines based on the criticalities of the load instructions that access them. Our experiments reveal that both these techniques can recover significant amount of the lost performance due to worst case designs. Madhu Mutyam, Feng Wang 0004, Krishnan Ramakrishnan, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Yuan Xie 0001, Mary Jane Irwin |
IEEE Trans. Computers | 1 |
| 2009 | Selective shielding technique to eliminate crosstalk transitionsabstractWith CMOS process technology scaling to deep submicron level, propagation delay across long on-chip buses is becoming one of the main performance limiting factors in high-performance designs. Propagation delay is very significant when adjacent wires are transitioning in opposite direction as compared to transitioning in the same direction. As opposite transitions on adjacent wires (called as crosstalk transitions ) have significant impact on propagation delay, several bus encoding techniques have been proposed in literature to eliminate such transitions. We propose selective shielding technique to eliminate crosstalk transitions. We show that the selective shielding technique requires ⌈3 n /2⌉ wires to encode a n -bit bus. SPICE simulations by considering 90nm technology nodes reveal that, for uniformly distributed random data, our technique achieves nearly 39% (21%) delay savings over 10 mm -length uncoded 32-bit bus for pipelined (nonpipelined) data transmission at the cost of nearly 7% energy overhead. Madhu Mutyam |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2008 | Block remap with turnoff: A variation-tolerant cache design techniqueabstractWith reducing feature size, the effects of process variations are becoming more and more predominant. Memory components such as on-chip caches are more susceptible to such variations because of high density and small sized transistors present in them. Process variations can result in high access latency and leakage energy dissipation. This may lead to a functionally correct chip being rejected, resulting in reduced chip yield. In this paper, by considering a process variation affected on-chip data cache, we first analyze performance loss due to worst-case design techniques such as accessing the entire cache with the worst-case access latency or turning off the process variation affected cache blocks, and show that the worst-case design techniques result in significant performance loss and/or high leakage energy. Then by exploiting the fact that not all applications require full associativity at set-level, we propose a variation-tolerant design technique, namely, block remap with turnoff (BRT), to minimize performance loss and leakage energy consumption. In BRT technique we selectively turnoff few blocks after rearranging them in such a way that all sets get almost equal number of process variation affected blocks. By turning off process variation affected blocks of a set, leakage energy can be minimized and the set can be accessed with low latency at the cost of reduced set associativity. We validate our technique by running SPEC2000 CPU benchmark-suite on Simplescalar simulator and show that our technique significantly reduces the performance loss and leakage energy consumption due to process variations. Mohammed Abid Hussain, Madhu Mutyam |
ASP-DAC | 2 |
| 2008 | Process Variation Aware Issue Queue DesignabstractIn sub-90 nm process technology it becomes harder to control the fabrication process, which in turn causes variations between the design-time parameters and the fabricated parameters. Variations in the critical process parameters can result in significant fluctuations in the switching speed and leakage power consumption of different transistors in the same chip. In this paper, we study the impact of process variation on issue queues. Due to process variation, issue queues can take variable access latency. In order to work with nonuniform access latency issue queues, by exploiting ready operands of instructions at dispatch time, we propose a process variation aware issue queue design. Experimental results reveal that, for a 64-entry issue queue with half of the entries affected by process variation, our technique recovers most of the lost performance due to process variation and incurs a performance penalty of less than 2% with respect to the performance of issue queues without process variation. Kanakagiri Raghavendra, Madhu Mutyam |
DATE | 2 |
| 2008 | Power management of variation aware chip multiprocessorsabstractFaced with the challenge of finding ways to use an ever-growing transistor budget, microarchitects have begun to move towards the chip multiprocessors (CMPs) as an attractive solution. CMPs have become a common way of reducing chip complexity and power consumption while maintaining high performance. Multiple cores are replicated on a single chip, resulting in a potential linear scaling of performance. Cores are becoming sufficiently small with technology scaling. As technology continues to scale, inter-die and intra-die variations in process parameters can result in significant impact on performance and power consumption, leading to asymmetry among the cores that were designed to be symmetric. Adaptive voltage scaling can be used to bring all cores to the same performance level leaving only core-to-core power variations. The goal of our work is to find the optimal frequency that balances performance with power against asymmetry. We also demonstrate that traditional task scheduling techniques need to be revisited to mitigate the effects of process variations. Abu Saad Papa, Madhu Mutyam |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | Word-interleaved cache: an energy efficient data cache architectureabstractWe propose a novel energy-efficient data cache architecture, namely, word-interleaved (WI) cache. In theWI cache, a cache block is distributed uniformly among the different cache ways and each line of a cache way holds some words of the block. This distribution provides an opportunity to activate/deactivate the cache ways based on the requested address's offset, thus minimizing the overall cache access energy. For a 4-way set associative cache of size 16KB and blocksize 32B, the proposed technique accomplishes dynamic energy savings of 54.2% without considering fast hits and 62.3% when fast hits are considered, with small performance degradation and negligible area overhead. T. Venkata Kalyan, Madhu Mutyam |
ISLPED | 2 |
| 2007 | Working with process variation aware caches
Madhu Mutyam, Narayanan Vijaykrishnan |
DATE | 1 |
| 2007 | Selective shielding: a crosstalk-free bus encoding techniqueabstractWith CMOS process technology scaling to deep submicron level, propagation delay across long on-chip buses is becoming one of the main performance limiting factors in high-performance designs. Propagation delay is very significant when adjacent wires are transitioning in opposite direction (i.e., crosstalk transitions) as compared to transitioning in the same direction. As crosstalk transitions have significant impact on propagation delay, several bus encoding techniques have been proposed in literature to eliminate such transitions. In this work, we propose a technique, namely, selective shielding, to eliminate crosstalk transitions. Compared to the conventional shielding technique, our technique significantly reduces the number of extra wires. We give a lower bound on the number of wires required to encode n-bit data using the selective shielding technique. We show that our technique achieves better energy savings and requires less area as compared to the other techniques. Madhu Mutyam |
ICCAD | 1 |
| 2006 | Delay and peak power minimization for on-chip buses using temporal redundancyabstractIn this paper, we propose a novel temporal redundancy based encoding technique for delay and peak power minimization. The proposed encoding scheme is tested with the SPEC2000 CINT benchmarks for 90nm and 65nm technologies. The experimental results show that our approach is very effective in reducing the peak power. From the delay perspective, our approach reduces the delay by at least 11% (4%) in the address (data) buses compared to the data transmission without encoding. K. Najeeb, V. Kamakoti 0001, Madhu Mutyam |
ACM Great Lakes Symposium on VLSI | 4 |
| 2006 | Compiler-directed thermal management for VLIW functional unitsabstractAs processors, memories, and other components of today's embedded systems are pushed to higher performance in more enclosed spaces, processor thermal management is quickly becoming a limiting design factor. While previous proposals mostly approached this thermal management problem from circuit and architecture angles, software can also play an important role in identifying and eliminating thermal hotspots as it is the main factor that shapes the order and frequency of accesses to different hardware components in the chip. This is particularly true for compiler-scheduled Very Long Instruction Word (VLIW) datapath.In this paper, we focus on a compiler-based approach to make the thermal profile more balanced in the integer functional units of VLIW architectures. For balanced thermal behavior and peak temperature minimization, we propose techniques based on load balancing across the integer functional units with or without rotation of functional unit usage. As leakage power is exponentially dependent on temperature and temperature is dependent on total power (i.e., switching and leakage), in our techniques, we also consider leakage power optimization by IPC tuning (instructions issued per cycle). By taking a code that is already scheduled for maximum performance as input, our scheduling strategies modify this performance-oriented schedule for balanced thermal behavior with negligible performance degradation. We simulate our scheduling strategies using a framework that consists of the Trimaran infrastructure, a power model, and the HotSpot. Our experimental results using several benchmark programs reveal that the peak temperature can be reduced through compiler scheduling. Madhu Mutyam, Feihui Li, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin |
LCTES | 1 |
| 2005 | On Characterizing Recursively Enumerable Languages by Insertion Grammars
Madhu Mutyam, Kamala Krithivasan, A. Siddhartha Reddy |
Fundam. Informaticae | 1 |
| 2005 | Rewriting P systems: improved hierarchies
Madhu Mutyam |
Theor. Comput. Sci. | 1 |
| 2003 | Array-rewriting P systems
Rodica Ceterchi, Madhu Mutyam, Gheorghe Paun, K. G. Subramanian 0001 |
Nat. Comput. | 2 |
| 2002 | Generalized normal form for rewriting P systems
Madhu Mutyam, Kamala Krithivasan |
Acta Informatica | 1 |
| 2002 | Contextual P Systems
Kamala Krithivasan, Madhu Mutyam |
Fundam. Informaticae | 2 |
| 2001 | P Systems with Membrane Creation: Universality and Efficiency
Madhu Mutyam, Kamala Krithivasan |
MCU | 1 |