VLDB 2026 Research / reviewers in the wild / expert
Manu Perumkunnil Komalan
dblp:144/4600 · also Manu Komalan, Manu Perumkunnil
· DBLP profile ↗
21ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-0029-6548ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 4 first-author · 13 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A System Level Performance Evaluation for Superconducting Digital SystemsabstractSuperconducting Digital (SCD) technology offers significant potential for enhancing the performance of next generation large scale compute workloads. By leveraging advanced lithography and a 300 mm platform, SCD devices can reduce energy consumption and boost computational power. This paper presents a cross-layer modeling approach to evaluate the system-level performance benefits of SCD architectures for Large Language Model (LLM) training and inference. Our findings, based on experimental data and Pulse Conserving Logic (PCL) design principles, demonstrate substantial performance gain in both training and inference. We are, thus, able to convincingly show that the SCD technology can address memory and interconnect limitations of present day solutions for next-generation compute systems. Joyjit Kundu, Debjyoti Bhattacharjee, Nathan Josephsen, Ankit Pokhrel, Udara De Silva, Wenzhe Guo, Steven Van Winckel, Steven Brebels, Quentin Herr, Anna Herr, Manu Perumkunnil Komalan |
DATE | 11 |
| 2025 | ARC: Application-Level Refinement and Cache Mapping for Performance Optimization on the EdgeabstractRecent advances in applications that are highly dependent on efficient cache utilization, in addition to the rapid growth of Edge computing systems deployed with emerging processors, generate a complex paradigm across the hardware and software continuum. In this work, we propose ARC, a novel systematic exploration methodology for application-level refinement and cache configuration mapping over emerging architectures for performance optimization. More specifically, our solution relies on workload partitioning and source code slicing mechanisms aiming to boost co-exploration of cache configuration parameters. Our proposed methodology is evaluated on a real-life IoT biomedical use case deployed over GEM5 RISC-V simulated system, showing that i) the co-impact of source code refinement and effective cache configuration leads to 61.1% execution time optimization, ii) the effective application organization and refinement leads to reduced hardware complexity. Last, we provide guidelines for application cache-friendly source code organization for performance optimization. Manolis Katsaragakis, Christos P. Lamprakos, Peter Kourzanov, Manu Perumkunnil Komalan, Lazaros Papadopoulos, Francky Catthoor, Dimitrios Soudris |
ISCAS | 4 |
| 2024 | Hier-3D: A Methodology for Physical Hierarchy Exploration of 3-D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. Complemented with an automated floorplanning solution, the flow allows for system-level physical and architectural exploration of 3D designs. As a result, we significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 -2.2× compared with 2D, where all metrics are improved simultaneously, including up to 20% power savings. Nesara Eranna Bethur, Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Design Technology co-optimization of 1D-1VCMA to improve read performance for SCM applicationsabstract1-diode 1-Voltage controlled magnetic anisotropy (1D-1VCMA) can be an option for Storage Class Memory (SCM) to bridge the latency gap between DRAM and flash memory. It has low sneak current, high non-linearity and low IR drop. This paper presents the Design Technology Co-optimization (DTCO) study of 1D-1VCMA stack to improve the performance and energy. Thanks to precessional switching of VCMA, the write operation is very fast, but the read determines overall latency as read before write is needed to ensure reliable write operations. The read performance of 1D-1VCMA is penalized due to high VCMA MTJ resistance, hence impacting the overall performance. To improve the read performance, this paper explores two solutions: 1) reducing the VCMA RA product, and 2) improving the read circuit. These solutions improve the read performance by 36% and 260%, respectively. Mohit Gupta 0004, Manu Perumkunnil Komalan, Dwaipayan Biswas, Saeideh Alinezhad Chamazcoti, Gouri Sankar Kar, Arnaud Furnémont, Julien Ryckaert |
ISCAS | 2 |
| 2023 | Impact of interconnects enhancement on SRAM design beyond 5nm technology nodeabstractThis paper presents an extensive study of 6T-SRAM based on FinFET for advanced technology nodes beyond 5nm. We deduce that parasitic resistance becomes the main bottleneck for SRAM design at these nodes. SRAM's writing margin and read speed are impacted due to the increased Bit-Line (BL) and Word-Line (WL) resistance. This work primarily explores two possible solutions to improve the parasitic resistance at advanced process technology nodes: 1) strapping of BL and WL to higher metal, and 2) adopting the resistance optimized BEOL. Strapping BL and WL to higher metal layer improves the Write Trip Point (WTP) by ~100mV and the critical path delay by 24% at the cost of 50% higher energy. Resistance optimized BEOL can improve WTP by ~50mV more and delay by 25% more, at the cost of increased energy consumption (8%). Mohit Gupta 0004, Pieter Weckx, Manu Perumkunnil Komalan, Julien Ryckaert |
ISCAS | 3 |
| 2023 | AMPeD: An Analytical Model for Performance in Distributed Training of TransformersabstractTransformers are a class of machine learning models that have piqued high interest recently due to a multitude of reasons. They can process multiple modalities efficiently and have excellent scalability. Despite these obvious advantages, training these large models is very time-consuming. Hence, there have been efforts to speed up the training process using efficient distributed implementations. Many different types of parallelism have been identified that can be employed standalone or in combination. However, naively combining different parallelization schemes can incur significant communication overheads, thereby potentially defeating the purpose of distributed training. Thus, it becomes vital to predict the right mapping of different parallelisms to the underlying system architecture. In this work, we propose AMPeD, an analytical model for performance in distributed training of transformers. It exposes all the transformer model parameters, potential parallelism choices (along with their mapping onto the system), the accelerator as well as system architecture specffications as tunable knobs, thereby enabling hardware-software co-design. With the help of 3 case studies, we show that the combinations of parallelisms predicted to be efficient by AMPeD conform with the results from the state-of-the-art literature. Using AMPeD, we also show that future distributed systems consisting of optical communication substrates can train large models up to 4× faster as compared to the current state-of the-art systems without modifying the peak computational power of the accelerators. Finally, we validate AMPeD with in-house experiments on real systems and via published literature. The max. observed error is limited to 12%. The model is available here: https://github.com/CSA-infra/AMPeD Diksha Moolchandani, Joyjit Kundu, Frederik Ruelens, Peter Vrancx, Timon Evenblij, Manu Perumkunnil Komalan |
ISPASS | 6 |
| 2023 | Beyond RSS: Towards Intelligent Dynamic Memory Management (Work in Progress)abstractThe main goal of dynamic memory allocators is to minimize memory fragmentation. Fragmentation stems from the interaction between workload behavior and allocator policy. There are, however, no works systematically capturing said interaction. We view this gap as responsible for the absence of a standardized, quantitative fragmentation metric, the lack of workload dynamic memory behavior characterization techniques, and the absence of a standardized benchmark suite targeting dynamic memory allocation. Such shortcomings are profoundly asymmetric to the operation’s ubiquity. Christos P. Lamprakos, Sotirios Xydis, Peter Kourzanov, Manu Perumkunnil Komalan, Francky Catthoor, Dimitrios Soudris |
MPLR | 4 |
| 2023 | Exploring Pareto-Optimal Hybrid Main Memory Configurations Using Different Emerging MemoriesabstractMain memory system design and corresponding technology requirements have become increasingly challenging for data-dominated high-performance applications. To address the leakage and scalability issues of the conventional DRAM-based memory, new memory technologies with ultra-low leakage and potential for high scalability have been explored extensively over the last decade. However, none of them are mature enough to serve as a drop-in replacement for DRAM. In this paper, we propose a hybrid main memory system solution for utilizing new memory technologies with specific features, based on the target application characteristics and system configurations. To this end, we examine two new memories, 1S-1VCMA and IGZO-based DRAM, along with conventional DRAM in the context of hybrid main memory solutions for high-capacity and low-power Pareto-optimizations, respectively. To better evaluate the power and performance, we consider the page-fault modeling in our evaluations. The results of the simulation show that different combinations of memory technologies in the hybrid memory system, different memory capacities, and different storage systems could provide a promising solution in the system regarding the characteristics of running applications and the requirements of the system. Saeideh Alinezhad Chamazcoti, Mohit Gupta 0004, Hyungrock Oh, Timon Evenblij, Francky Catthoor, Manu Perumkunnil Komalan, Gouri Sankar Kar, Arnaud Furnémont |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Hier-3D: A Hierarchical Physical Design Methodology for Face-to-Face-Bonded 3D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. We significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 to 2.2 × compared with 2D, where all metrics are improved simultaneously, including up to power savings. Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
ISLPED | 6 |
| 2022 | Analyzing the Electromigration Challenges of Computation in Resistive MemoriesabstractPerforming the computation in memory (CiM) based on the resistive non-volatile memories can significantly improve the energy efficiency and performance of data-intensive and deep learning applications. Activating multiple rows of the memories at the same time is required in Multiply and Accumulation (MAC) operation of neural networks. This simultaneous activation, however, increases the current density of the shared interconnect, which exacerbates the Electromigration (EM) risk. This paper analyzes the EM phenomenon in CiM-oriented MAC paradigms based on emerging non-volatile resistive memories including Spin Transfer Torque Magnetic RAM (STT-MRAM), Redox-based RAM (ReRAM), and Phase Change Memory (PCM). We show how EM is exacerbated compared to normal memory architectures. For EM analysis in CiM, we modify the existing EM models, and consider different interconnect and array dimensions. We also propose the EM-aware row activation pattern as effective means to mitigate the EM degradations in the analog MAC paradigms. Mahta Mayahinia, Mehdi Baradaran Tahoori, Manu Perumkunnil Komalan, Kris Croes, Francky Catthoor |
ITC | 3 |
| 2022 | Time-Dependent Electromigration Modeling for Workload-Aware Design-Space Exploration in STT-MRAMabstractElectromigration (EM) has been known as a reliability threatening factor for back-end-of-the-line interconnects. Spin-transfer torque magnetic RAM (STT-MRAM) is an emerging nonvolatile memory that has gained a lot of attention in recent years. However, relatively large operational current magnitude is a challenge for this technology, and hence, EM can be a potential reliability concern, even for the signal lines of this memory. A workload-aware EM modeling needs to capture time-dependent current density in the memory signal lines and to be able to predict the effect of the EM phenomenon on the interconnect for its entire lifetime. In this work, we present methods to effectively model the workload-dependent EM-induced meantime to failure (MTTF) in typical STT-MRAM arrays under a variety of realistic workloads. This allows performing the design-space exploration to co-optimize reliability and other design metrics. Mahta Mayahinia, Mehdi Baradaran Tahoori, Manu Perumkunnil Komalan, Houman Zahedmanesh, Kris Croes, Tommaso Marinelli, José Ignacio Gómez, Timon Evenblij, Gouri Sankar Kar, Francky Catthoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Microarchitectural Exploration of STT-MRAM Last-level Cache Parameters for Energy-efficient DevicesabstractAs the technology scaling advances, limitations of traditional memories in terms of density and energy become more evident. Modern caches occupy a large part of a CPU physical size and high static leakage poses a limit to the overall efficiency of the systems, including IoT/edge devices. Several alternatives to CMOS SRAM memories have been studied during the past few decades, some of which already represent a viable replacement for different levels of the cache hierarchy. One of the most promising technologies is the spin-transfer torque magnetic RAM (STT-MRAM), due to its small basic cell design, almost absent static current and non-volatility as an added value. However, nothing comes for free, and designers will have to deal with other limitations, such as the higher latencies and dynamic energy consumption for write operations compared to reads. The goal of this work is to explore several microarchitectural parameters that may overcome some of those drawbacks when using STT-MRAM as last-level cache (LLC) in embedded devices. Such parameters include: number of cache banks, number of miss status handling registers (MSHRs) and write buffer entries, presence of hardware prefetchers. We show that an effective tuning of those parameters may virtually remove any performance loss while saving more than 60% of the LLC energy on average. The analysis is then extended comparing the energy results from calibrated technology models with data obtained with freely available tools, highlighting the importance of using accurate models for architectural exploration. Tommaso Marinelli, José Ignacio Gómez, Christian Tenllado, Manu Perumkunnil Komalan, Mohit Gupta 0004, Francky Catthoor |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | Memory Hierarchy Calibration Based on Real Hardware In-order Cores for Accurate SimulationabstractComputer system simulators are major tools used by architecture researchers. Two key elements play a role in the credibility of simulator results: (1) the simulator's accuracy, and (2) the quality of the baseline architecture. Some simulators, such as gem5, already provide highly accurate parameterized models. However, finding the right values for all these parameters to faithfully model a real architecture is still a problem. In this paper, we calibrate the memory hierarchy of an in-order core gem5 simulation to accurately model a real mobile Arm SoC. We execute small programs, which we design to stress specific parts of the memory system, to deduce key parameter values for the model. We compare the execution of SPEC CPU2006 benchmarks on the real hardware with the gem5 simulation. Our results show that our calibration reduces the average and worst-case IPC error by 36 % and 50%, respectively, when compared with a gem5 simulation configured with the default parameters. Quentin Huppert, Timon Evenblij, Manu Perumkunnil Komalan, Francky Catthoor, Lionel Torres, David Novo |
DATE | 3 |
| 2021 | Power, Performance, Area and Cost Analysis of Memory-on-Logic Face-to-Face Bonded 3D Processor DesignsabstractIn this paper, we present a power, performance, area and cost (PPAC) analysis for large-scale 3D processor designs based on wafer-to-wafer bonding. From the evaluation of our cost model, we investigate a typically disregarded opportunity in 3D that is area savings due to buffer savings and better routability, offering unexpected cost savings. We explore the viability of this factor with the feedback of a state-of-the-art 3D memory-on-logic implementation flow. We show how this affects the PPAC of full-chip GDS implementations of a large-scale manycore processor design. Experiments show that our memory-on-logic 3D implementation offers 7% silicon area savings, resulting in 53.5% footprint reduction. We also obtain a 40% power-performance-cost improvement compared with 2D counterparts Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Dragomir Milojevic, Francky Catthoor, Manu Perumkunnil Komalan, Sung Kyu Lim |
ISLPED | 7 |
| 2021 | High-Performance Logic-on-Memory Monolithic 3-D IC Designs for Arm Cortex-A ProcessorsabstractMonolithic 3-D IC (M3-D) is a promising solution to improve the performance and energy-efficiency of modern processors. But, designers are faced with challenges in design tools and methodologies, especially for power and thermal verifications. We developed a new physical design flow that optimally places and routes cache modules in one tier and logic gates in the other. Our tool also builds high-quality clock and power delivery networks targeting logic-on-memory M3-D designs. Finally, we developed a sign-off analysis tool flow to evaluate power, performance, area (PPA), thermal, and voltage-drop quality for given M3-D designs. Using our complete register transfer level (RTL)-to-Graphic Design System (GDS) tool flow, we designed commercial quality 2-D and M3-D implementation of Arm Cortex-A7 and Cortex-A53 processors in a commercial 28-nm technology. Experimental results show that our 3-D processors offer 20% (A7) and 21% (A53) performance gain, compared with their 2-D commercial counterparts. The voltage-drop degradation of our 3-D Cortex-A7 and Cortex-A53 processors is less than 3% of the supply voltage, while temperature increase is 10.71 °C and 13.04 °C, respectively. Lingjun Zhu, Lennart Bamberg, Sai Pentapati, Kyungwook Chang, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Brian Cline, Saurabh Sinha 0001, Alberto García Ortiz, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2019 | Process, Circuit and System Co-optimization of Wafer Level Co-Integrated FinFET with Vertical Nanosheet Selector for STT-MRAM ApplicationsabstractWe present for the first time a co-integrated FinFET with vertical nanosheet transistor (VFET) process on a 300 mm silicon wafer for STT-MRAM applications and its related avenues with a holistic design-technology-co-optimization (DTCO) and power-performance-area-cost (PPAC) approach. The STT-MRAM bitcell and a 2 Mbit macro have been optimized and designed to address the viability of the co-integration process and advantages of vertical channel transistors for STT-MRAM selectors. An architectural system simulator GEM5 has been also employed with Polybench workloads to assess energy saving at system-level. In order to enable this co-integration, four extra masks are required, which costs below 10% in embedded chips. A 36% area reduction can be achieved for the STT-MRAM bitcell implemented with VFET selectors. With a UVLT flavor, the STT-MRAM bitcell comprising of 3-nanosheet could deliver the same performance of the 4-fin LVT FinFET selector. A 2 Mbit STT-MRAM macro designed with VFET selector can offer a 17% and a 21% reduction for read access latency and energy per operation respectively, and a 10% for write energy per operation. A 7% energy saving for the STT-MRAM L2 cache using VFET selector has been observed at the system level with Polybench workloads. Trong Huynh Bao, Anabela Veloso, Sushil Sakhare, Philippe Matagne, Julien Ryckaert, Manu Perumkunnil Komalan, Davide Crotti, Farrukh Yasin, Alessio Spessot, Arnaud Furnémont, Gouri Sankar Kar, Anda Mocuta |
DAC | 6 |
| 2019 | A Comparative Analysis on the Impact of Bank Contention in STT-MRAM and SRAM Based LLCsabstractSpin Transfer Torque Magnetic RAM (STT-MRAM) is being extensively considered as a promising replacement for Last Level Caches (LLC), due to its high density, low leakage and non-volatility. However, writes to STT-MRAM are energy intensive and have a high latency. While the high dynamic energy consumption during writes can be compensated by the low static energy consumption, the high latency results in performance degradation. This work shows that in contrast to SRAM-based LLCs, the performance degradation for STT-MRAM is primarily due to bank contention, when trying to satisfy a read request while the bank is being written. We holistically explore the effects of cache banking and cache contention on energy and performance in the LLC of mobile multicore systems, with in-order cores or with out-of-order cores. The detail of the analysis is enabled by highly accurate cache models, based on a 28nm SRAM industry compiler, and an in-house developed STT-MRAM compiler, which generates full STT-MRAM macro designs with silicon-validated MTJ stack and complete parasitic extraction at the 28nm node. Our results show that there is a clear difference in the energy-performance optimal banking configuration between STT-MRAM caches and SRAM caches. These low contention STT-MRAM cache designs with the optimal number of banks save at least 60% cache energy while losing at most single digit percentages in system performance compared to SRAM cache designs. This show an increased potential of using STT-MRAM as a replacement for SRAM in an LLC. Timon Evenblij, Christian Tenllado, Manu Perumkunnil Komalan, Francky Catthoor, Sushil Sakhare, Peter Debacker, Gouri Sankar Kar, Arnaud Furnémont, Nicolas Bueno, José Ignacio Gómez |
ICCD | 3 |
| 2018 | Main memory organization trade-offs with DRAM and STT-MRAM options based on gem5-NVMain simulation frameworksabstractCurrent main memory organizations in embedded and mobile application systems are DRAM dominated. The ever-increasing gap between today's processor and memory speeds makes the DRAM subsystem design a major aspect of computer system design. However, the limitations to DRAM scaling and other challenges like refresh provide undesired trade-offs between performance, energy and area to be made by architecture designers. Several emerging NVM options are being explored to at least partly remedy this but today it is very hard to assess the viability of these proposals because the simulations are not fully based on realistic assumptions on the NVM memory technologies and on the system architecture level. In this paper, we propose to use realistic, calibrated STT-MRAM models and a well calibrated cross-layer simulation and exploration framework, named SEAT, to better consider technologies aspects and architecture constraints. We will focus on general purpose/mobile SoC multi-core architectures. We will highlight results for a number of relevant benchmarks, representatives of numerous applications based on actual system architecture. The most energy efficient STT-MRAM based main memory proposal provides an average energy consumption reduction of 27% at the cost of 2x the area and the least energy efficient STT-MRAM based main memory proposal provides an average energy consumption reduction of 8% at the around the same area or lesser when compared to DRAM. Manu Perumkunnil Komalan, Hyungrock Oh, Matthias Hartmann, Sushil Sakhare, Christian Tenllado, José Ignacio Gómez, Gouri Sankar Kar, Arnaud Furnémont, Francky Catthoor, Sophiane Senni, David Novo, Abdoulaye Gamatié, Lionel Torres |
DATE | 1 |
| 2017 | Cross-layer design and analysis of a low power, high density STT-MRAM for embedded systemsabstractSTT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has attracted considerable attention of late since it is the most promising logic compatible nonvolatile memory that is suitable for advanced logic nodes (N28 and beyond) in terms of endurance, speed and power. Embedded STT-MRAM has thus been proposed as a candidate for emerging low standby-power connectivity systems such IoT (Internet-of-Things) and wearables. We utilize the high performance CoFeB based perpendicular MTJ (pMTJ) device to realize a low power and highly dense STT-MRAM array for such systems. This study is carried out on the TSMC 28nm technology node and includes a complete cross-layer design and analysis framework ranging from device modeling to circuit design, layout and system implementation. The process variations and temperature (PT) impact on the MTJ for the STT-MRAM design (and correspondingly the total energy consumption and performance of the system) is also analyzed. We report a ∼85% reduction in the energy consumption compared to the baseline SRAM based system for near negligible performance penalty (<5%). Manu Perumkunnil Komalan, Sushil Sakhare, Trong Huynh Bao, Siddharth Rao, Christian Tenllado, José Ignacio Gómez, Gouri Sankar Kar, Arnaud Furnémont, Francky Catthoor |
ISCAS | 1 |
| 2015 | System level exploration of a STT-MRAM based level 1 data-cache
Manu Perumkunnil Komalan, Christian Tenllado, José Ignacio Gómez, Francisco Tirado, Francky Catthoor |
DATE | 1 |
| 2014 | Feasibility exploration of NVM based I-cache through MSHR enhancementsabstractSRAM based memory systems are plagued by a number of problems like sub-threshold leakage and susceptibility to read/write failure with dynamic voltage scaling schemes or low supply voltage. Non-Volatile Memory (NVM) technologies are being explored extensively nowadays to replace the conventional SRAM memories even for level 1 (L1) caches. These NVMs like Spin Torque Transfer RAM (STT-MRAM), Resistive-RAM (ReRAM) and Phase Change RAM (PRAM) are less hindered by leakage problems with technology scaling and consume lesser area. However, simple replacement of SRAM by NVMs is not a viable option due to their write related issues. The main focus of this paper is the exploration of write delay and write energy issues in a NVM based L1 Instruction cache (I-cache) for an ARM like single core system. We propose a NVM I-cache and extend its MSHR (Miss Status Handling Register) functionality to address the NVMs write related issues. According to our simulations, appropriate tuning of selective architecture parameters can reduce the performance penalty introduced by the NVM (∼45%) to extremely tolerable levels (∼1%) and show energy gains up to 35%. Furthermore, on configuring our modified NVM based system to occupy area comparable to the original SRAM-based configuration, it outperforms the SRAM baseline and leads to even more energy savings. Manu Perumkunnil Komalan, José Ignacio Gómez, Christian Tenllado, Praveen Raghavan, Matthias Hartmann, Francky Catthoor |
DATE | 1 |