Preeti Ranjan Panda

dblp:46/1929 · DBLP profile ↗
← Back
80ranked-venue papers
20as first author
21since 2021 · last 2025
0000-0002-2508-7531ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 76 · 20 first-author · 21 since 2021Software engineering, systems software and programming languages · 15 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Cryo-CACTI: Cryogenic-Aware CACTI for Cache Modeling Down to 10K in Advanced 7nm FinFETs
abstract
Cryogenic circuits are currently employed in fields such as quantum computing, particle detectors, magnetic resonance imaging, and space applications. While cryogenic circuits are being researched, there is limited work on designing cryogenic caches at temperatures below 77K. Moreover, there is no tool to estimate the delay, power, and area of cryogenic caches at advanced technology nodes. Our research focuses on the development of cryogenic caches tailored for the 7nm technology node, operating at 10K. However, a key challenge is the lack of cryogenic measurement data, especially in recent technologies. Consequently, through conducting our own FinFET transistor measurements, we calibrate cryogenic transistor models at 10K. With the 7nm cryogenic transistor data, we modelCryo-CACTIfor cryogenic caches (due to cache’s vital role in improving performance and their considerable share in area and power of the processor). Using Cryo-CACTI, our evaluation reveals considerable improvements in the energy efficiency (up to 99%) of cryogenic caches of larger sizes compared to the caches at room temperature (300K). Additionally, we explore alternative cache configurations at circuit-level to optimize cryogenic operation. Furthermore, we use Cryo-CACTI to explore the performance/energy consumption of cryogenic caches while simulating workloads such as SPEC CPU2017 and machine learning via neural networks.Cryo-CACTI is available for download athttps://github.com/marg-tools/Cryo-CACTI
Divya Praneetha Ravipati, Victor M. van Santen, Shivendra Singh Parihar, Yogesh Singh Chauhan, Preeti Ranjan Panda, Hussam Amrouch
IEEE Trans. Computers5
2025 FLASH: Deadline-Aware Flexible LLC Arbitration and Scheduling for Hardware Accelerators
abstract
Integrating domain-specific hardware accelerators on modern systems on chips (SoCs) has enabled complex applications, such as vision, natural language processing, autonomous driving, and augmented reality, on small form factors. This leads to challenges in the integration of accelerators, with high memory bandwidth requirements and strict deadlines, on the system’s memory hierarchy. The system-level shared cache, or last-level cache (LLC), is a critical resource shared by multi-core processors, GPUs, and hardware accelerators in modern heterogeneous SoCs. It significantly reduces the bottleneck at the off-chip memory and delivers high performance. With the integration of accelerators on the LLC gaining momentum, the on-chip shared cache management becomes vital. If not managed intelligently, the interference between cache requests from the cores and the accelerators can significantly deteriorate their performance. Given the architectural differences between DRAM and cache systems, the off-chip memory management strategies explored by previous works cannot be extended to the LLC. We propose a deadline-aware flexible LLC arbitration and scheduling framework, FLASH , to dynamically partition the LLC bandwidth between the accelerators and multi-core processors to meet the deadline given for the accelerator while minimizing the impact on the performance of the cores. FLASH arbitrates between the requests from the cores and the accelerators and schedules the requests depending on the accelerator’s progress and its chances of meeting the deadline. We evaluate FLASH across different workloads and hardware accelerator configurations to show that it not only achieves significantly better performance for the cores than other static scheduling policies but also significantly reduces the deadline miss rates of the accelerator.
Ayushi Agarwal, Pulkit Goel, P. J. Joseph, Prokash Ghosh, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.6
2025 SHARP: SHARing-Aware Cache Writeback byPass
abstract
In modern multi-processor systems-on-chips (MPSoCs), writebacks from the private caches to the shared cache can introduce significant performance bottlenecks, especially because multiple threads from different co-executing programs contend for the shared cache resources. Intelligent cache bypass decisions for writebacks help mitigate such contention and enhance the utilization of the shared cache. Most prior cache bypass strategies account for contention for shared cache capacity by focusing primarily on data reuse, with only recent research beginning to consider bandwidth contention also in dynamic bypass decisions. However, data sharing, a crucial characteristic of modern multithreaded workloads, remains largely overlooked by state-of-the-art cache bypass decisions. Bypassing highly shared cache lines can increase the volume of main memory accesses, potentially resulting in performance bottlenecks. We introduce SHARP, a novel cache bypass policy that incorporates three key factors: data sharing, contention, and data reuse, into its dynamic bypass decisions for cache writebacks. In addition to prioritizing the caching of data with high reuse, we prioritize the caching of data shared across multiple threads to enhance cache utilization. We dynamically modulate our bypass decisions, employing aggressive bypass for writebacks when shared cache contention is high, while employing conservative bypass when contention is low. Experiments across a diverse set of PARSEC workloads demonstrate that SHARP improves overall system throughput by 12% and 8% compared to the no-bypass baseline and the state-of-the-art bypass baseline, respectively. SHARP also reduces the overall cache energy consumption by 14% over the no-bypass baseline.
Dinesh Joshi, Aritra Bagchi, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.3
2025 FARRE: Fairness Aware Request Response Arbitration in Shared Caches
abstract
Contention in shared caches caused by concurrently executing applications can lead to overall performance degradation in multiprocessor systems-on-chip (MPSoCs). To address this issue, various shared cache arbitration techniques have been proposed to manage cache bandwidth contention. These techniques focus on enhancing overall system performance; however, this optimization often comes at the expense of system fairness, leading to some applications experiencing disproportionate slowdowns, or, in the worst case, starvation. Therefore, an effective shared cache bandwidth management policy is needed to optimize performance while ensuring fairness across applications. We propose FARRE , a novel fairness aware request-response arbitration technique for shared caches. FARRE is designed to optimize performance while attempting to maintain a user-defined fairness threshold. We evaluate its effectiveness through extensive simulations including comparisons against state-of-the-art arbitration schemes. The results show that FARRE is able to maintain or exceed the input fairness thresholds, and improves system performance over standard fair scheduling policies such as round-robin; the performance improvement is 14% for lower fairness thresholds such as 0.5, and could even gain 5% performance for aggressive thresholds such as 0.9. Additionally, compared to the best performance optimization techniques, FARRE achieves 81% higher fairness.
Garima Modi, Priyanka Singla 0001, Neetu Jindal, Ayan Mandal, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.5
2024 APPAMM: Memory Management for IPsec Application on Heterogeneous SoCs
abstract
To keep up with the growing computational demands of current-day applications, SoC design has shifted towards heterogeneous architectures with CPUs and domain-specific accelerators. These accelerators demand high on-chip and off-chip memory bandwidth and require efficient management of shared system resources. We characterize Internet Protocol Security (IPsec), a high-throughput application, by collecting the memory traces of this application running on the accelerators and CPU cores of NXP LX2160A SoC. We use this characterization to design a simulation infrastructure for simulating IPsec on possible domain-specific architectural extensions and perform a design-space exploration across various general-purpose memory management policies. We propose APPAMM, an application-specific predictive packet-aware memory management policy using the knowledge of IPsec to improve performance for next-generation SoCs. Using our approach to manage memory for different input packet streams, we report improvements of up to 22x in the packet drop rate and peak throughput.
Ayushi Agarwal, Radhika Dharwadkar, Isaar Ahmad, P. J. Joseph, Prokash Ghosh, Preeti Ranjan Panda
VLSI-SoC8
2024 CAPE: Criticality-Aware Performance and Energy Optimization Policy for NCFET-Based Caches
abstract
Caches are crucial yet power-hungry components in present-day computing systems. With the Negative Capacitance Fin Field-Effect Transistor (NCFET) gaining significant attention due to its internal voltage amplification, allowing for better operation at lower voltages (stronger ON-current and reduced leakage current), the introduction of NCFET technology in caches can reduce power consumption without loss in performance. Apart from the benefits offered by the technology, we leverage the unique characteristics offered by NCFETs and propose a dynamic voltage scaling based criticality-aware performance and energy optimization policy (CAPE) for on-chip caches. We present the first work towards optimizing energy in NCFET-based caches with minimal impact on performance. Compared to operating at a nominal voltage of 0.7 V, CAPE shows improvement in Last-Level Cache (LLC) energy savings by up to 19.2%, while the baseline policies devised for traditional CMOS- (/FinFET-) based caches are ineffective in improving NCFET-based LLC energy savings. Compared to the considered baseline policies, our CAPE policy also demonstrates better LLC energy-delay product (EDP) and throughput savings.
Divya Praneetha Ravipati, Ramanuj Goel, Victor M. van Santen, Hussam Amrouch, Preeti Ranjan Panda
IEEE Trans. Computers5
2024 NOVELLA: Nonvolatile Last-Level Cache Bypass for Optimizing Off-Chip Memory Energy
abstract
Contemporary multiprocessor systems-on-chips (MPSoCs) continue to confront energy-related challenges, primarily originating from off-chip data movements. Nonvolatile memories (NVMs) emerge as a promising solution with their high-storage density and low leakage, yet they suffer from slow and expensive write operations. Writebacks from higher-level caches and responses from off-chip memory create significant contention at the shared nonvolatile last-level cache (LLC), affecting system performance with increased queuing for critical reads. Previous research primarily addresses the performance issues by trying to mitigate contention through the bypassing of NVM writes. Nevertheless, off-chip memory energy, one of the most critical components of system energy, remains unaddressed by state-of-the-art bypass policies. While certain energy components, such as leakage and refresh, depend on system performance, performance-optimizing bypass policies may not ensure energy efficiency. Aggressive bypass decisions aimed only at performance enhancement could degrade cache reuse, potentially outweighing reductions in leakage and refresh energies with the increase in off-chip dynamic energy. While both performance and off-chip memory energy are influenced by both cache contention and reuse, the tradeoffs for achieving optimal performance versus optimal energy are different. We introduce nonvolatile last-level cache bypass for optimizing off-chip memory energy (NOVELLA), a novel bypass policy for the nonvolatile LLC, to optimize off-chip memory energy by exploiting tradeoffs between cache contention and reuse, achieving a balance across different components of the energy. Compared to a naïve no-bypass baseline, while state-of-the-art reuse-aware bypass solutions reduce off-chip memory energy consumption by up to 8%, and a contention- and reuse-aware bypass baseline by 12%, NOVELLA achieves significant energy savings of 21% across diverse SPEC workloads.
Aritra Bagchi, Ohm Rishabh, Preeti Ranjan Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 3D-TemPo: Optimizing 3-D DRAM Performance Under Temperature and Power Constraints
abstract
3-D DRAM provides a significant performance boost resulting from substantial memory bandwidth. However, the stacked memory architecture exhibits high power density, causing thermal hotspots. Further, systems under power constraints require careful planning for intelligent allocation of the available power to their various components. A straightforward dynamic power management policy of allocating more power to potentially high memory activity 3-D DRAM ranks so as to maximize system performance causes a rise in the temperature of such ranks, making them susceptible to thermal stalls and shutdown by dynamic thermal management (DTM) strategies. A rise in rank temperature, in turn, increases the leakage power of memory ranks, affecting power budgeting decisions. Thus, a coordinated strategy for power budgeting and thermal management is needed. We propose an adjacency-aware dynamic power budgeting technique, 3D-TemPo, which dynamically performs a reward-based power allocation to memory ranks, in order to maximize 3-D DRAM performance under power and thermal constraints, and is sensitive to strong thermal correlations between vertically adjacent ranks. We evaluate 3D-TemPo using SPEC CPU2017 and PARSEC 2.1 benchmark suites and observe speedups of$1\times $to$17.94\times $compared to baseline strategies.
Shailja Pandey, Sayam Sethi, Preeti Ranjan Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 COBRRA: COntention-aware cache Bypass with Request-Response Arbitration
abstract
In modern multi-processor systems-on-chip (MPSoCs), requests from different processor cores, accelerators, and their responses from the lower-level memory contend for the shared cache bandwidth, making it a critical performance bottleneck. Prior research on shared cache management has considered requests from cores but has ignored crucial contributions from their responses. Prior cache bypass techniques focused on data reuse and neglected the system-level implications of shared cache contention. We propose COBRRA, a novel shared cache controller policy that mitigates the contention by aggressively bypassing selected responses from the lower-level memory and scheduling the remaining requests and responses to the cache efficiently. COBRRA is able to improve the average performance of a set of 15 SPEC workloads by 49% and 33% compared to the no-bypass baseline and the best-performing state-of-the-art bypass solution, respectively. Furthermore, COBRRA reduces the overall cache energy consumption by 38% and 31% compared to the no-bypass baseline and the most energy-efficient state-of-the-art bypass solution, respectively.
Aritra Bagchi, Dinesh Joshi, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.3
2024 NeuroTAP: Thermal and Memory Access Pattern-Aware Data Mapping on 3D DRAM for Maximizing DNN Performance
abstract
Deep neural networks (DNNs) have been widely adopted, owing to break-through performance and high accuracy. DNNs exhibit varying memory behavior involving specific and recognizable memory access patterns and access intensity, depending on the selected data reuse in different layers. Such applications have high memory bandwidth demands due to aggressive computations, performing several billion-floating-point-operations-per-second (BFLOPs). 3D DRAMs, providing very high memory access bandwidth, are extensively employed to break the memory wall , bridging the gap between compute and memory while running DNNs. However, the vertical integration in 3D DRAM introduces serious thermal issues, resulting from high power density and close proximity of memory cells, and requires dynamic thermal management (DTM). To unleash the true potential of 3D DRAM and exploit the enormous bandwidth under thermal constraints, there is a need to intelligently map the DNN application’s data across memory channels, pseudo-channels, and banks, minimizing the effective memory latency and reducing the thermal-induced application slowdown. The specific memory access patterns exhibited by a DNN layer execution are crucial to determine a favorable data mapping method for 3D DRAM dies that potentially causes minimal thermal impact and also maximizes DRAM bandwidth utilization. In this work, we propose an application-aware and thermal-sensitive data mapping that intelligently assigns portions of the 3D DRAM to DNN layers, leveraging the knowledge about layer’s memory access patterns and minimizing DTM-induced performance overheads. Additionally, we also deploy a DRAM low-power states based DTM mechanism to keep the 3D DRAM within safe thermal limits. Using our proposal, we observe a performance improvement of 1% to 61%, and memory energy savings of 1% to 55% for popular DNNs over state-of-the-art DTM strategies while running DNN inference.
Shailja Pandey, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.2
2024 POEM: Performance Optimization and Endurance Management for Non-volatile Caches
abstract
Non-volatile memories (NVMs), with their high storage density and ultra-low leakage power, offer promising potential for redesigning the memory hierarchy in next-generation Multi-Processor Systems-on-Chip (MPSoCs). However, the adoption of NVMs in cache designs introduces challenges such as NVM write overheads and limited NVM endurance. The shared NVM cache in an MPSoC experiences requests from different processor cores and responses from the off-chip memory when the requested data is not present in the cache. Besides, upon evictions of dirty data from higher-level caches, the shared NVM cache experiences another source of write operations, known as writebacks . These sources of write operations—writebacks and responses—further exacerbate the contention for the shared bandwidth of the NVM cache and create significant performance bottlenecks. Uncontrolled write operations can also affect the endurance of the NVM cache, posing a threat to cache lifetime and system reliability. Existing strategies often address either performance or cache endurance individually, leaving a gap for a holistic solution. This study introduces the Performance Optimization and Endurance Management (POEM) methodology, a novel approach that aggressively bypasses cache writebacks and responses to alleviate the NVM cache contention. Contrary to the existing bypass policies that do not pay adequate attention to the shared NVM cache contention and focus too much on cache data reuse, POEM’s aggressive bypass significantly improves the overall system performance, even at the expense of data reuse. POEM also employs effective wear leveling to enhance the NVM cache endurance by careful redistribution of write operations across different cache lines. Across diverse workloads, POEM yields an average speedup of 34% over a naïve baseline and 28.8% over a state-of-the-art NVM cache bypass technique while enhancing the cache endurance by 15% over the baseline. POEM also explores diverse design choices by exploiting a key policy parameter that assigns varying priorities to the two system-level objectives.
Aritra Bagchi, Dharamjeet, Ohm Rishabh, Manan Suri, Preeti Ranjan Panda
ACM Trans. Design Autom. Electr. Syst.5
2024 NeuroCool: Dynamic Thermal Management of 3D DRAM for Deep Neural Networks through Customized Prefetching
abstract
Deep neural network (DNN) implementations are typically characterized by huge datasets and concurrent computation, resulting in a demand for high memory bandwidth due to intensive data movement between processors and off-chip memory. Performing DNN inference on general-purpose cores/edge is gaining attraction to enhance user experience and reduce latency. The mismatch in the CPU and conventional DRAM speed leads to under-utilization of the compute capabilities, causing increased inference time. 3D DRAM is a promising solution to effectively fulfill the bandwidth requirement of high-throughput DNNs. However, due to high power density in stacked architectures, 3D DRAMs need dynamic thermal management (DTM), resulting in performance overhead due to memory-induced CPU throttling. We study the thermal impact of DNN applications running on a 3D DRAM system, and make a case for a memory temperature-aware customized prefetch mechanism to reduce DTM overheads and significantly improve performance. In our proposed NeuroCool DTM policy, we intelligently place either DRAM ranks or tiers in low power state, using the DNN layer characteristics and access rate. We establish the generalization of our approach through training and test datasets comprising diverse data points from widely used DNN applications. Experimental results on popular DNNs show that NeuroCool results in a average performance gain of 44% (as high as 52%) and memory energy improvement of 43% (as high as 69%) over general-purpose DTM policies.
Shailja Pandey, Lokesh Siddhu, Preeti Ranjan Panda
ACM Trans. Design Autom. Electr. Syst.3
2023 Education Abstract: Thermal Challenges and Mitigation in 3D DRAM
Preeti Ranjan Panda, Shailja Pandey
CODES+ISSS1
2023 CABARRE: Request Response Arbitration for Shared Cache Management
abstract
Modern multi-processor systems-on-chip (MPSoCs) are characterized by caches shared by multiple cores. These shared caches receive requests issued by the processor cores. Requests that are subject to cache misses may result in the generation of responses . These responses are received from the lower level of the memory hierarchy and written to the cache. The outstanding requests and responses contend for the shared cache bandwidth. To mitigate the impact of the cache bandwidth contention on the overall system performance, an efficient request and response arbitration policy is needed. Research on shared cache management has neglected the additional cache contention caused by responses, which are written to the cache. We propose CABARRE , a novel request and response arbitration policy at shared caches, so as to improve the overall system performance. CABARRE shows a performance improvement of 23% on average across a set of SPEC workloads compared to straightforward adaptations of state-of-the-art solutions.
Garima Modi, Aritra Bagchi, Neetu Jindal, Ayan Mandal, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.5
2023 Dynamic Thermal Management of 3D Memory through Rotating Low Power States and Partial Channel Closure
abstract
Modern high-performance and high-bandwidth three-dimensional (3D) memories are characterized by frequent heating. Prior art suggests turning off hot channels and migrating data to the background DDR memory, incurring significant performance and energy overheads. We propose three Dynamic Thermal Management (DTM) approaches for 3D memories, reducing these overheads. The first approach, Rotating-channel Low-power-state-based DTM (RL-DTM) , minimizes the energy overheads by avoiding data migration. RL-DTM places 3D memory channels into low power states instead of turning them off. Since data accesses are disallowed during low power state, RL-DTM balances each channel’s low-power-state duration. The second approach, Masked rotating-channel Low-power-state-based DTM (ML-DTM) , is a fine-grained policy that minimizes the energy-delay product (EDP) and improves the performance of RL-DTM by considering the channel access rate. The third strategy, Partial channel closure and ML-DTM , minimizes performance overheads of existing channel-level turn-off-based policies by closing a channel only partially and integrating ML-DTM, reducing the number of channels being turned off. We evaluate the proposed DTM policies using various mixes of SPEC benchmarks and multi-threaded workloads and observe them to significantly improve performance, energy, and EDP over state-of-the-art approaches for different 3D memory architectures.
Lokesh Siddhu, Aritra Bagchi, Rajesh Kedia, Isaar Ahmad, Shailja Pandey, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.6
2023 Performance and Energy Studies on NC-FinFET Cache-Based Systems With FN-McPAT
abstract
To understand performance and energy tradeoffs in CPU–memory systems at lower geometries and new technologies, there is a need to update the processor and cache models used by instruction-level simulators. We improve the existing McPAT tool to support the 14-nm FinFET commercial technology, while respecting McPAT’s overall modeling methodology. We also include the results from the BOOM CPU core, synthesized with FinFET technology, into the McPAT tool to model the core components. For the first time, we extend McPAT to support the negative capacitance fin field-effect transistor (NC-FinFET), an emerging transistor technology with subthreshold swing (SS) below 60 mV/decade and unique leakage characteristics. Experiments using our FN-McPAT tool indicate that the NC-FinFET-based system is more energy-efficient relative to the FinFET-based system for memory-intensive workloads and vice versa for the compute-intensive workloads while operating at the highest voltage and frequency. In addition, we analyze the performance and energy consumption of last-level caches (LLCs) operating at various voltages and report novel insights into the energy consumption behavior for the NC-FinFET-based LLC. FN-McPAT is available for download athttps://github.com/marg-tools/FN-McPAT.
Divya Praneetha Ravipati, Victor M. van Santen, Sami Salamin, Hussam Amrouch, Preeti Ranjan Panda
IEEE Trans. Very Large Scale Integr. Syst.5
2022 CoreMemDTM: Integrated Processor Core and 3D Memory Dynamic Thermal Management for Improved Performance
abstract
The growing performance of processors and 3D memories has resulted in higher power densities and temperatures. Dynamic thermal management (DTM) policies for processor cores and memory have received significant research attention, but existing solutions address processors and 3D memories independently, which causes overcompensation, and there is a need to coordinate the DTM of the two subsystems. Further, existing CPU DTM policies slow down heated cores significantly, increasing the overall execution time and performance overheads. We propose CoreMemDTM, a technique for integrating processor core and 3D memory DTM policies that attempts to minimize performance overheads. We suggest employing DTM depending on the thermal margin since safe temperature thresholds might differ for the two subsystems. We propose a stall-balanced core DVFS policy for core thermal management that enables distributed cooling, decreasing overheads. We evaluate CoreMemDTM using ten different SPEC CPU2017 workloads across various safe temperature thresholds and observe average execution time and energy improvements of 14% and 36% compared to state-of-the-art DTM policies.
Lokesh Siddhu, Rajesh Kedia, Preeti Ranjan Panda
DATE3
2022 CoMeT: An Integrated Interval Thermal Simulation Toolchain for 2D, 2.5D, and 3D Processor-Memory Systems
abstract
Processing cores and the accompanying main memory working in tandem enable modern processors. Dissipating heat produced from computation remains a significant problem for processors. Therefore, the thermal management of processors continues to be an active subject of research. Most thermal management research is performed using simulations, given the challenges in measuring temperatures in real processors. Fast yet accurate interval thermal simulation toolchains remain the research tool of choice to study thermal management in processors at the system level. However, the existing toolchains focus on the thermal management of cores in the processors, since they exhibit much higher power densities than memory. The memory bandwidth limitations associated with 2D processors lead to high-density 2.5D and 3D packaging technology: 2.5D packaging technology places cores and memory on the same package; 3D packaging technology takes it further by stacking layers of memory on the top of cores themselves. These new packagings significantly increase the power density of the processors, making them prone to overheating. Therefore, mitigating thermal issues in high-density processors (packaged with stacked memory) becomes even more pressing. However, given the lack of thermal modeling for memories in existing interval thermal simulation toolchains, they are unsuitable for studying thermal management for high-density processors. To address this issue, we present the first integrated Core and Memory interval Thermal (CoMeT) simulation toolchain.CoMeTcomprehensively supports thermal simulation of high- and low-density processors corresponding to four different core-memory (integration) configurations—off-chip DDR memory, off-chip 3D memory, 2.5D, and 3D.CoMeTsupports several novel features that facilitate overlying system research.CoMeTadds only an additional ~5% simulation-time overhead compared to an equivalent state-of-the-art core-only toolchain. The source code ofCoMeThas been made open for public use under theMITlicense.
Lokesh Siddhu, Rajesh Kedia, Shailja Pandey, Martin Rapp, Anuj Pathania, Jörg Henkel, Preeti Ranjan Panda
ACM Trans. Archit. Code Optim.7
2022 NeuroMap: Efficient Task Mapping of Deep Neural Networks for Dynamic Thermal Management in High-Bandwidth Memory
abstract
High-bandwidth memory (HBM) offers breakthrough memory bandwidth through its vertically stacked memory architecture and through-silicon via (TSV)-based fast interconnect. However, the stacked architecture leads to high-power density causing thermal issues when running modern memory-hungry workloads such as deep neural networks (DNNs). Prior works on dynamic thermal management (DTM) of 3-D DRAM do not consider the physical structure of HBM and often lead to heavy DTM-induced performance penalty. We propose an application-aware efficient task mapping and migration-based DTM policy that maps DNN instances to cores through exploiting the channel layout of HBM and leveraging the significant temperature gradient across DRAM dies while making thermal decisions. We utilize the variation in the memory access behavior of DNN layers and attempt to minimize stalling due to thermal hotspots in the HBM stack. We also use application-aware dynamic voltage and frequency scaling (DVFS) and DRAM low-power states to further improve performance. Experimental results on workloads comprising seven popular DNNs show that NeuroMap results in an average execution time and memory energy reduction of 39% and 40%, respectively, over state-of-the-art DTM mechanisms.
Shailja Pandey, Preeti Ranjan Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 FN-CACTI: Advanced CACTI for FinFET and NC-FinFET Technologies
abstract
Cache memories are an indispensable component of many processor-based systems and contribute significantly to the overall area, power consumption, and delay. This leads to an important role played by modeling tools for estimating the area, power consumption, and access time of cache memories. However, existing modeling tools such as CACTI and its various extensions have been primarily designed using data from various projections. For the first time, we propose an entire flow for obtaining/calibrating the transistor characteristics from a commercial technology and use these characteristics within CACTI. We also improve the modeling approach to make them more fine-grained and follow recent manufacturing trends suitable for FinFET technology. Further, for the first time, we extend CACTI to support negative capacitance fin field effect transistor (NC-FinFET), an emerging technology depicting negative capacitance whose current and capacitive characteristics are very different compared to those of the FinFET. We use the proposed tool (FN-CACTI) to identify NC-FinFET-based caches to be significantly more energy-efficient than corresponding FinFET-based caches. We also study an application of FN-CACTI to determine optimal voltages corresponding to the lowest energy consumption for NC-FinFET and FinFET-based caches of various sizes.
Divya Praneetha Ravipati, Rajesh Kedia, Victor M. van Santen, Jörg Henkel, Preeti Ranjan Panda, Hussam Amrouch
IEEE Trans. Very Large Scale Integr. Syst.5
2021 Leakage-Aware Dynamic Thermal Management of 3D Memories
abstract
3D memory systems offer several advantages in terms of area, bandwidth, and energy efficiency. However, thermal issues arising out of higher power densities have limited their widespread use. While prior works have looked at reducing dynamic power through reduced memory accesses, in these memories, both leakage and dynamic power consumption are comparable. Furthermore, as the temperature rises, the leakage power increases, creating a thermal-leakage loop. We study the impact of leakage power on 3D memory temperature and propose turning OFF specific memory channels to meet thermal constraints. Data is migrated to a 2D memory before closing a 3D channel. We introduce an analytical model to assess the 2D memory delay and use the model to guide data migration decisions. The above strategy is referred to asFastCooland provides an improvement of 22%, 19%, and 32% on average (up to 57%, 72%, and 82%) in performance, memory energy, and energy-delay product (EDP), respectively, on different workloads consisting of SPEC CPU2006 benchmarks. We further propose a thermal management strategy namedEnergy-Efficient FastCool (EEFC), which improves upon FastCool by selecting the channels to be closed by considering temperature, leakage, access rate, and position of various 3D memory channels at runtime. Our experiments demonstrate that EEFC leads to an additional improvement of up to 30%, 30%, and 51% in performance, memory energy, and EDP compared to FastCool. Finally, we analyze the effects of process variations on the efficiency of the proposed FC and EEFC strategies. Variation in the manufacturing process causes changes in the leakage power and temperature profile. Since EEFC considers both while selecting channels for closure, it is more resilient to process variations and achieves a lower application execution time and memory energy compared to FastCool.
Lokesh Siddhu, Rajesh Kedia, Preeti Ranjan Panda
ACM Trans. Design Autom. Electr. Syst.3
2020 Enhancing Network-on-Chip Performance by Reusing Trace Buffers
abstract
Ensuring the functional correctness of networks-on-chip (NoCs) can be particularly challenging, and communication-centric debug methodologies have been widely used by engineers to validate NoC functionality during post-silicon validation. Design-for-debug structures, such as trace buffers and monitors, are usually inserted in such systems-on-chip to enhance signal visibility. However, this debug hardware becomes underutilized once the chip goes into production. While the size and organization of the router buffers directly impact network throughput, these buffers also dominate the on-chip router area. We propose a scheme augmented virtual channel (AugVC) to reuse trace buffers to augment router buffers, with the objective of improving the overall network performance. The experimental results for a 64-node mesh network show that our proposed approach can reduce latency by up to 38.25% for transpose traffic compared to a baseline design with reduced buffer sizes. We also propose an extension, output port directed virtual channel (ODVC), that uses a modified virtual channel assignment strategy, on the basis of the designated output port of a network packet. This strategy reduces the average packet latency and area of the router by 45% and 32.4%, respectively.
Neetu Jindal, Shubhani Gupta, Divya Praneetha Ravipati, Preeti Ranjan Panda, Smruti R. Sarangi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 REAL: REquest Arbitration in Last Level Caches
abstract
Shared last level caches (LLC) of multicore systems-on-chip are subject to a significant amount of contention over a limited bandwidth, resulting in major performance bottlenecks that make the issue a first-order concern in modern multiprocessor systems-on-chip. Even though shared cache space partitioning has been extensively studied in the past, the problem of cache bandwidth partitioning has not received sufficient attention. We demonstrate the occurrence of such contention and the resulting impact on the overall system performance. To address the issue, we perform detailed simulations to study the impact of different parameters and propose a novel cache bandwidth partitioning technique, called REAL , that arbitrates among cache access requests originating from different processor cores. It monitors the LLC access patterns to dynamically assign a priority value to each core. Experimental results on different mixes of benchmarks show up to 2.13× overall system speedup over baseline policies, with minimal impact on energy.
Sakshi Tiwari, Shreshth Tuli, Isaar Ahmad, Ayushi Agarwal, Preeti Ranjan Panda, Sreenivas Subramoney
ACM Trans. Embed. Comput. Syst.5
2019 DHOOM: Reusing Design-for-Debug Hardware for Online Monitoring
abstract
Runtime verification employs dedicated hardware or software monitors to check whether program properties hold at runtime. However, these monitors often incur high area and performance overheads depending on whether they are implemented in hardware or software. In this work, we propose DHOOM, an architectural framework for runtime monitoring of program assertions, which exploits the combination of a reconfigurable fabric present alongside a processor core with the vestigial on-chip Design-for-Debug hardware. This combination of hardware features allows DHOOM to minimize the overall performance overhead of runtime verification, even when subject to a given area constraint. We present an algorithm for dynamically selecting an effective subset of assertion monitors that can be accommodated in the available programmable fabric, while instrumenting the remaining assertions in software. We show that our proposed strategy, while respecting area constraints, reduces the performance overhead of runtime verification by up to 32% when compared with a baseline of software-only monitors.
Neetu Jindal, Sandeep Chandran, Preeti Ranjan Panda, Sanjiva Prasad, Abhay Mitra, Kunal Singhal, Shikhar Tuli
DAC3
2019 FastCool: Leakage Aware Dynamic Thermal Management of 3D Memories
abstract
3D memory systems offer several advantages in terms of area, bandwidth, and energy efficiency. However, thermal issues arising out of higher power densities have limited their widespread use. While prior works have looked at reducing dynamic power through reduced memory accesses, in these memories, both leakage and dynamic power consumption are comparable. Furthermore, as the temperature rises the leakage power increases, creating a thermal-leakage loop. We study the impact of leakage power on 3D memory temperature and propose turning OFF hot channels to meet thermal constraints. Data is migrated to a 2D memory before closing a 3D channel. We introduce an analytical model to assess the 2D memory delay and use the model to guide data migration decisions. Our experiments show that the proposed optimization improves performance by 27% on an average (up to 66%) over state-of-the-art strategies.
Lokesh Siddhu, Preeti Ranjan Panda
DATE2
2019 Alleria: An Advanced Memory Access Profiling Framework
abstract
Application analysis and simulation tools are used extensively by embedded system designers to improve existing optimization techniques or develop new ones. We propose the Alleria framework to make it easier for designers to comprehensively collect critical information such as virtual and physical memory addresses, accessed values, and thread schedules about one or more target applications. Such profilers often incur substantial performance overheads that are orders of magnitude larger than native execution time. We discuss how that overhead can be significantly reduced using a novel profiling mechanism called adaptive profiling. We develop a heuristic-based adaptive profiling mechanism and evaluate its performance using single-threaded and multi-threaded applications. The proposed technique can improve profiling throughput by up to 145% and by 37% on an average, enabling Alleria to be used to comprehensively profile applications with a throughput of over 3 million instructions per second.
Hadi Brais, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.2
2019 PredictNcool: Leakage Aware Thermal Management for 3D Memories Using a Lightweight Temperature Predictor
abstract
Recent research on mitigating thermal problems in 3D memories has covered reactive strategies that reduce memory power consumption, and thereby, performance, when the memory temperature reaches the maximum operating limit. Such techniques could benefit from temperature prediction and avoid unnecessary invocations and state transitions of the thermal management strategy. We develop an accurate steady state temperature predictor for thermal management of 3D memories. We utilize the symmetries in the floorplan, along with other design insights, to reduce the predictor’s model parameters, making it lightweight and suitable for runtime thermal management. Using the temperature prediction, we introduce PredictNcool , a proactive thermal management strategy to reduce application runtime and memory energy. We compare PredictNcool with two recent thermal management strategies and our experiments show that the proposed optimization results in performance improvements of 28% and 5%, and memory subsystem energy reductions of 38% and 12% (on average).
Lokesh Siddhu, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.2
2018 Reusing Trace Buffers as Victim Caches
Neetu Jindal, Preeti Ranjan Panda, Smruti R. Sarangi
IEEE Trans. Very Large Scale Integr. Syst.2
2017 A coordinated multi-agent reinforcement learning approach to multi-level cache co-partitioning
abstract
The widening gap between the processor and memory performance has led to the inclusion of multiple levels of caches in the modern multi-core systems. Processors with simultaneous multithreading (SMT) support multiple hardware threads on the same physical core, which results in shared private caches. Any inefficiency in the cache hierarchy can negatively impact the system performance and motivates the need to perform a co-optimization of multiple cache levels by trading off individual application throughput for better system throughput and energy-delay-product (EDP). We propose a novel coordinated multiagent reinforcement learning technique for performing Dynamic Cache Co-partitioning, called Machine Learned Caches (MLC). MLC has low implementation overhead and does not require any special hardware data profilers. We have validated our proposal with 15 8-core workloads created using Spec2006 benchmarks and found it to be an effective co-partitioning technique. MLC exhibited system throughput and EDP improvements of up to 14% (gmean:9.35%) and 19.2% (gmean: 13.5%) respectively. We believe this is the first attempt at addressing the problem of multi-level cache co-partitioning.
Rahul Jain 0004, Preeti Ranjan Panda, Sreenivas Subramoney
DATE2
2017 Reusing trace buffers to enhance cache performance
abstract
With the increasing complexity of modern Systems-on-Chip, the possibility of functional errors escaping design verification is growing. Post-silicon validation targets the discovery of these errors in early hardware prototypes. Due to limited visibility and observability, dedicated design-for-debug (DFD) hardware such as trace buffers are inserted to aid post-silicon validation. In spite of its benefit, such hardware incurs area overheads, which impose size limitations. However, the overhead could be overcome if the area dedicated to DFD could be reused in-field. In this work, we present a novel method for reusing an existing trace buffer as a victim cache of a processor to enhance performance. The trace buffer storage space is reused for the victim cache, with a small additional controller logic. Experimental results on several benchmarks and trace buffer sizes show that the proposed approach can enhance the average performance by up to 8.3% over a baseline architecture. We also propose a strategy for dynamic power management of the structure, to enable saving energy with negligible impact on performance.
Neetu Jindal, Preeti Ranjan Panda, Smruti R. Sarangi
DATE2
2017 Cooperative Multi-Agent Reinforcement Learning-Based Co-optimization of Cores, Caches, and On-chip Network
abstract
Modern multi-core systems provide huge computational capabilities, which can be used to run multiple processes concurrently. To achieve the best possible performance within limited power budgets, the various system resources need to be allocated effectively. Any mismatch between runtime resource requirement and allocation leads to a sub-optimal energy-delay product (EDP). Different optimization techniques exist for addressing the problem of mismatch between the dynamic requirement and runtime allocation of the system resources. Choosing between multiple optimizations at runtime is complex due to the non-additive effects, making the scenario suitable for the application of machine learning techniques. We present a novel method, Machine Learned Machines (MLM), by using online reinforcement learning (RL) to perform dynamic partitioning of the last level cache (LLC), along with dynamic voltage and frequency scaling (DVFS) of the core and uncore (interconnection network and LLC). We have proposed and evaluated three different MLM co-optimization techniques based on independent and cooperative multi-agent learners. We show that the co-optimization results in a much lower system EDP than any of the techniques applied individually. We explore various RL models targeted toward optimization of different system metrics and study their effects on a system EDP, system throughput (STP), and Fairness. The various proposed techniques have been extensively evaluated with a mix of 20 workloads on a 4-core system using Spec2006 benchmarks. We have further evaluated our cooperative MLM techniques on a 16-core system. The results show an average of 20.5% and 19.1% system EDP improvement on a 4-core and 16-core system, respectively, with limited degradation of STP and Fairness.
Rahul Jain 0004, Preeti Ranjan Panda, Sreenivas Subramoney
ACM Trans. Archit. Code Optim.2
2017 Managing Trace Summaries to Minimize Stalls During Postsilicon Validation
abstract
On-chip trace buffers are increasingly being used for at-speed debug during postsilicon validation. The limited size of these buffers results in their frequent overflowing. In scenarios when such overflowing is not desirable, the chip is stalled, and the state data recorded in these buffers are transferred off-chip. Such frequent stalling significantly impedes efficient debugging. We propose a novel scheme to minimize the number of such stalls using a portion of the trace buffer to also store summaries of trace messages. We describe an overlapped trace buffer architecture that uses a reduced number of ports to capture tapered summaries where both detailed and summary versions of traces are stored simultaneously. We propose a simple hardware structure to generate two kinds of trace summaries-spatial and temporal-as specified by the validation engineer. We introduce a storage specification language that allows the validation engineer to unambiguously specify the information to be captured in these summaries to the debug hardware. We demonstrate that our proposal significantly reduces the number of stalls for off-chip transfer of captured traces in four bug scenarios that are representative of different classes of bugs encountered during postsilicon validation.
Sandeep Chandran, Preeti Ranjan Panda, Smruti R. Sarangi, Ayan Bhattacharyya, Deepak Chauhan, Sharad Kumar
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Extending trace history through tapered summaries in post-silicon validation
abstract
On-chip trace buffers are increasingly being used for at-speed debug during post-silicon validation. However, the activity history captured by these buffers is small due to their limited size. We propose a novel scheme that extends the captured trace history (by upto 162%) by using a portion of the trace buffer to also store summaries of trace messages. We describe an Overlapped trace buffer architecture that uses a reduced number of ports to capture tapered summaries where both detailed and summary versions of traces are stored simultaneously. We demonstrate the usefulness of the proposed methodology for debugging various classes of bugs encountered during post-silicon validation.
Sandeep Chandran, Preeti Ranjan Panda, Deepak Chauhan, Sharad Kumar, Smruti R. Sarangi
ASP-DAC2
2016 Machine Learned Machines: Adaptive co-optimization of caches, cores, and On-chip Network
Rahul Jain 0004, Preeti Ranjan Panda, Sreenivas Subramoney
DATE2
2016 Data Flow Transformation for Energy-Efficient Implementation of Givens Rotation-Based QRD
abstract
QR decomposition (QRD), a matrix decomposition algorithm widely used in embedded application domain, can be realized in a large number of valid processing sequences that differ significantly in the number of memory accesses and computations, and hence the overall implementation energy. With modern low-power embedded processors evolving toward register files with wide memory interfaces and vector functional units (FUs), data flow in these algorithms needs to be carefully devised to efficiently utilize the costly wide memory accesses and the vector FUs. In this article, we present an energy-efficient data flow transformation strategy for the Givens rotation--based QRD.
Namita Sharma 0001, Preeti Ranjan Panda, Francky Catthoor, Min Li 0001, Prashant Agrawal
ACM Trans. Embed. Comput. Syst.2
2016 Integrated Exploration Methodology for Data Interleaving and Data-to-Memory Mapping on SIMD Architectures
abstract
This work presents a methodology for efficient exploration of data interleaving and data-to-memory mapping options for Single Instruction Multiple Data (SIMD) platform architectures. The system architecture consists of a reconfigurable clustered scratch-pad memory and a SIMD functional unit, which performs the same operation on multiple input data in parallel. The memory accesses contribute substantially to the overall energy consumption of an embedded system executing a data intensive task. The scope of this work is the reduction of the overall energy consumption by increasing the utilization of the functional units and decreasing the number of memory accesses. The presented methodology is tested using a number of benchmark applications with holes in their access scheme. Potential gains are calculated based on the energy models, both for the processing and the memory part of the system. The reduction in energy consumption after efficient interleaving and mapping of data is between 40% and 80% for the complete system and the studied benchmarks.
Iasonas Filippopoulos, Namita Sharma 0001, Francky Catthoor, Per Gunnar Kjeldsberg, Preeti Ranjan Panda
ACM Trans. Embed. Comput. Syst.5
2016 Partitioning and Data Mapping in Reconfigurable Cache and Scratchpad Memory-Based Architectures
abstract
Scratchpad memory (SPM) is considered a useful component in the memory hierarchy, solely or along with caches, for meeting the power and energy constraints as performance ceases to be the sole criteria for processor design. Although the efficiency of SPM is well known, its use has been restricted owing to difficulties in programmability. Real applications usually have regions that are amenable to exploitation by either SPM or cache and hence can benefit if the two are used in conjunction. Dynamically adjusting the local memory resources to suit application demand can significantly improve the efficiency of the overall system. In this article, we propose a compiler technique to map application data objects to the SPM-cache and also partition the local memory between the SPM and cache depending on the dynamic requirement of the application. First, we introduce a novel graph-based structure to tackle data allocation in an application. Second, we use this to present a data allocation heuristic to map program objects for a fixed-size SPM-cache hybrid system that targets whole program optimization. We finally extend this formulation to adapt the SPM and cache sizes, as well as the data allocation as per the requirement of different application regions. We study the applicability of the technique on various workloads targeted at both SPM-only and hardware reconfigurable memory systems, observing an average of 18% energy-delay improvement over state-of-the-art techniques.
Prasenjit Chakraborty, Preeti Ranjan Panda, Sandeep Sen
ACM Trans. Design Autom. Electr. Syst.2
2016 Area-Aware Cache Update Trackers for Postsilicon Validation
abstract
The internal state of the complex modern processors often needs to be dumped out frequently during postsilicon validation. Since the caches hold most of the state, the volume of data dumped and the transfer time are dominated by the large caches present in the architecture. The limited bandwidth to transfer data present in these large caches off-chip results in stalling the processor for long durations when dumping the cache contents off-chip. To alleviate this, we propose to transfer only those cache lines that were updated since the previous dump. Since maintaining a bit-vector with a separate bit to track the status of individual cache lines is expensive, we propose two methods: 1) where a bit tracks multiple cache lines and 2) an Interval Table which stores only the starting and ending addresses of continuous runs of updated cache lines. Both methods require significantly lesser space compared with a bit-vector, and allow the designer to choose the amount of space to allocate for this design-for-debug feature. The impact of reducing storage space is that some nonupdated cache lines are dumped too. We attempt to minimize such overheads. We propose a scheme to share such cache update tracking hardware (or Update Trackers) across multiple caches in case of physically distributed caches so that they are replicated fewer times, thereby limiting the area overhead. We show that the proposed Update Trackers occupy less than 1% of cache area for both the shared and distributed caches.
Sandeep Chandran, Smruti R. Sarangi, Preeti Ranjan Panda
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Array Interleaving - An Energy-Efficient Data Layout Transformation
abstract
Optimizations related to memory accesses and data storage make a significant difference to the performance and energy of a wide range of data-intensive applications. These techniques need to evolve with modern architectures supporting wide memory accesses. We investigate array interleaving , a data layout transformation technique that achieves energy efficiency by combining the storage of data elements from multiple arrays in contiguous locations, in an attempt to exploit spatial locality. The transformation reduces the number of memory accesses by loading the right set of data into vector registers, thereby minimizing redundant memory fetches. We perform a global analysis of array accesses, and account for possibly different array behavior in different loop nests that might ultimately lead to changes in data layout decisions for the same array across program regions. Our technique relies on detailed estimates of the savings due to interleaving, and also the cost of performing the actual data layout modifications. We also account for the vector register widths and the possibility of choosing the appropriate granularity for interleaving. Experiments on several benchmarks show a 6--34% reduction in memory energy due to the strategy.
Namita Sharma 0001, Preeti Ranjan Panda, Francky Catthoor, Praveen Raghavan, Tom Vander Aa
ACM Trans. Design Autom. Electr. Syst.2
2014 Array scalarization in high level synthesis
abstract
Parallelism across loop iterations present in behavioral specifications can typically be exposed and optimized using well known techniques such as Loop Unrolling. However, since behavioral arrays are usually mapped to memories (SRAM) during synthesis, performance bottlenecks arise due to memory port constraints. We study array scalarization, the transformation of an array into a group of scalar variables. We propose a technique for selectively scalarizing arrays for improving the performance of synthesized designs by taking into consideration the latency benefits as well as the area overhead caused by using discrete registers for storing array elements instead of denser SRAM. Our experiments on several benchmark examples indicate promising speedups of more than 10x for several designs due to scalarization.
Preeti Ranjan Panda, Namita Sharma 0001, Arun Kumar Pilania, Gummidipudi Krishnaiah, Sreenivas Subramoney, Ashok Jagannathan
ASP-DAC1
2014 Energy optimization in Android applications through wakelock placement
abstract
Energy efficiency is a critical factor in mobile systems, and a significant body of recent research efforts has focused on reducing the energy dissipation in mobile hardware and applications. The Android OS Power Manager provides programming interface routines called wakelocks for controlling the activation state of devices on a mobile system. An appropriate placement of wakelock acquire and release functions in the application can make a significant difference to the energy consumption. In this paper, we propose a data flow analysis based strategy for determining the placement of wakelock statements corresponding to the uses of devices in an application. Our experimental evaluation on a set of Android applications show significant (up to 32%) energy savings with the proposed optimization strategy.
Faisal Alam, Preeti Ranjan Panda, Nikhil Tripathi, Namita Sharma 0001, Sanjiv Narayan
DATE2
2014 Energy efficient data flow transformation for Givens Rotation based QR Decomposition
abstract
QR Decomposition (QRD) is a typical matrix decomposition algorithm that shares many common features with other algorithms such as LU and Cholesky decomposition. The principle can be realized in a large number of valid processing sequences that differ significantly in the number of memory accesses and computations, and hence, the overall implementation energy. With modern low power embedded processors evolving towards register files with wide memory interfaces and vector functional units (FUs), the data flow in matrix decomposition algorithms needs to be carefully devised to achieve energy efficient implementation. In this paper, we present an efficient data flow transformation strategy for the Givens Rotation based QRD that optimizes data memory accesses. We also explore different possible implementations for QRD of multiple matrices using the SIMD feature of the processor. With the proposed data flow transformation, a reduction of up to 36% is achieved in the overall energy over conventional QRD sequences.
Namita Sharma 0001, Preeti Ranjan Panda, Min Li 0001, Prashant Agrawal, Francky Catthoor
DATE2
2014 High level energy modeling of controller logic in data caches
abstract
In modern embedded processor caches, a significant amount of energy dissipation occurs in the controller logic part of the cache. Previous power/energy modeling tools have focused on the core memory part of the cache. We propose energy models for two of these modules -- Write Buffer and Replacement logic. Since this hardware is generally synthesized by designers, our power models are also based on empirical data. We found a linear dependence of the per-access write buffer energy on the write buffer depth and write width. We validated our models on several different benchmark examples, using different technology nodes. Our models generate energy estimates that are within 4.2% of those measured by detailed power simulations, making the models valuable mechanisms for rapid energy estimates during architecture exploration.
Preeti Ranjan Panda, Srikanth Chandrasekaran, Namita Sharma 0001, Sarath Kumar Kandalam, Nagaraj N.
ACM Great Lakes Symposium on VLSI1
2014 Shared-port register file architecture for low-energy VLIW processors
abstract
We propose a reduced-port Register File (RF) architecture for reducing RF energy in a VLIW processor. With port reduction, RF ports need to be shared among Function Units (FUs), which may lead to access conflicts, and thus, reduced performance. Our solution includes (i) a carefully designed RF-FU interconnection network that permits port sharing with minimum conflicts and without any delay/energy overheads, and (ii) a novel scheduling and binding algorithm that reduces the performance penalty. With our solution, we observed as much as 83% RF energy savings with no more than a 10% loss in performance for a set of Mediabench and Mibench benchmarks.
Neeraj Goel, Preeti Ranjan Panda
ACM Trans. Archit. Code Optim.3
2013 SPM-Sieve: A framework for assisting data partitioning in scratch pad memory based systems
abstract
Modern system architectures sometimes include scratch pad memories (SPM) in their memory hierarchy to take advantage of their simpler design, in an attempt to meet the system area, performance, and power budget. These systems employing SPM can be broadly categorized as: (a) cacheless systems with only SPM, (b) hybrid systems with both cache and SPM, and (c) reconfigurable systems with the provision to reconfigure local memory as either cache, SPM, or a combination of the two. However SPM based systems have needed larger efforts spent on their programming, mainly due to allocating data and orchestrating data transfers explicitly by soft-ware. Tight product development cycles require faster development and porting of diverse applications to multiple SPM based architectures. In this paper we present SPM-Sieve, a profile-based tool and framework targeted for SPM based architectures that generates partitioning decisions of the first level memory in the system hierarchy, and suggests object mapping amongst the memory partitions without resorting to detailed simulation of all configurations. This is done by natively executing an application and using minimal target architecture specification, which not only provides early information influencing data organization in the application, but also provides a foundation for other more sophisticated algorithms to produce optimized allocations. We demonstrate the utility and generality of SPM-Sieve by evaluating it on a large number of SPEC2000 benchmarks targeted for a 128KB first level memory. We evaluate its effectiveness by performing simulation studies comparing the partition suggested by the tool against varying partition sizes, and observe that its suggestions are very competitive for SPM based architectures with and without caches.
Prasenjit Chakraborty, Preeti Ranjan Panda
CASES2
2013 Space sensitive cache dumping for post-silicon validation
abstract
The internal state of complex modern processors often needs to be dumped out frequently during post-silicon validation. Since the last level cache (considered L2 in this paper) holds most of the state, the volume of data dumped and the transfer time are dominated by the L2 cache. The limited bandwidth to transfer data off-chip coupled with the large size of L2 cache results in stalling the processor for long durations when dumping the cache contents off-chip. To alleviate this, we propose to transfer only those cache lines that were updated since the previous dump. Since maintaining a bit-vector with a separate bit to track the status of individual cache lines is expensive, we propose 2 methods: (i) where a bit tracks multiple cache lines and (ii) an Interval Table which stores only the starting and ending addresses of continuous runs of updated cache lines. Both methods require significantly lesser space compared to a bit-vector, and allow the designer to choose the amount of space to allocate for this design-for-debug (DFD) feature. The impact of reducing storage space is that some non-updated cache lines are dumped too. We attempt to minimize such overheads. Further, the Interval Table is independent of the cache size which makes it ideal for large caches. Through experimentation, we also determine the break-even point below which a t-lines/bit bit-vector is beneficial compared to an Interval Table.
Sandeep Chandran, Smruti R. Sarangi, Preeti Ranjan Panda
DATE3
2013 Data memory optimization in LTE downlink
abstract
Optimizations related to memory accesses and data storage make a significant difference to the performance and energy of a wide range of data-intensive applications. Such strategies need to evolve with modern SoC and processor architectures, which lead to new optimization opportunities. In this paper, we focus on data memory optimization for LTE downlink receiver as this is a data- and computation-intensive part of the LTE application with tight energy and latency constraints. We study the data dependencies globally and conclude that by providing data samples from the antennas in interleaved form at the FFT input, we can achieve 7-15% reduction in memory access energy over an optimized implementation without any performance overhead.
Namita Sharma 0001, Tom Vander Aa, Prashant Agrawal, Praveen Raghavan, Preeti Ranjan Panda, Francky Catthoor
ICASSP5
2012 Integrating software caches with scratch pad memory
abstract
Software cache refers to cache functionality emulated in software on a compiler-controlled Scratch Pad Memory (SPM). Such structures are useful when standard SPM allocation strategies cannot be used due to hard-to-analyze memory reference patterns in the source code. SPM data allocation strategies generally rely on compile-time inference of spatial and temporal reuse, with the general flow being the copying of a block/tile of array data into the SPM, followed by its processing, and finally, copying back. However, when array index functions are complicated due to conditionals, complex expressions, and dependence on run-time data, the SPM compiler has to rely on expensive DMA for individual words, leading to poor performance. Software caches (SWC) can play a crucial role in improving performance under such circumstances -- their access times are longer than those for direct SPM access, but they retain the advantages (present in hardware caches) of exploiting spatial and temporal locality discovered at run-time. We present the first automated compiler data allocation strategy that considers the presence of a software cache in SPM space, and makes decisions on which arrays should be accessed through it, at which times. Arrays could be accessed differently in different parts of a program, and our algorithm analyzes such uses and considers the possibility of selectively accessing an array through the SWC only when it is efficient, based on a cost model of the overheads involved in SPM/SWC transitions. We implemented our technique in an LLVM based framework and experimented with several applications on a Cell based machine. Our technique results in up to 82% overall performance improvement over a conventional SPM mapping algorithm and up to 27% over a typical SWC-enhanced implementation.
Prasenjit Chakraborty, Preeti Ranjan Panda
CASES2
2012 Efficient on-line algorithm for maintaining k-cover of sparse bit-strings
abstract
We consider the on-line problem of representing a sparse bit string by a set of k intervals, where k is much smaller than the length of the string. The goal is to minimize the total length of these intervals under the condition that each 1-bit must be in one of these intervals. We give an efficient greedy algorithm which takes time O(log k) per update (an update involves converting a 0-bit to a 1-bit), which is independent of the size of the entire string. We prove that this greedy algorithm is 2-competitive. We use a natural linear programming relaxation for this problem, and analyze the algorithm by finding a dual feasible solution whose value matches the cost of the greedy algorithm.
Amit Kumar 0001, Preeti Ranjan Panda, Smruti R. Sarangi
FSTTCS2
2011 A SysML Profile for Development and Early Validation of TLM 2.0 Models
Preeti Ranjan Panda
ECMFA3
2011 A UML based framework for efficient validation of TLM 2 models
Preeti Ranjan Panda
FDL3
2011 Compressing Cache State for Postsilicon Processor Debug
abstract
During postsilicon processor debugging, we need to frequently capture and dump out the internal state of the processor. Since internal state constitutes all memory elements, the bulk of which is composed of cache, the problem is essentially that of transferring cache contents off-chip, to a logic analyser. In order to reduce the transfer time and save expensive logic analyser memory, we propose to compress the cache contents on their way out. We present a hardware compression engine for cache data using a Cache-Aware Compression strategy that exploits knowledge of the cache fields and their behavior to achieve an effective compression. Experimental results indicate that the technique results in 7-31 percent better compression than one that treats the data as just one long bit stream. We also describe and evaluate a parallel compression architecture that uses multiple compression engines, resulting in a 54 percent reduction in transfer time.
Preeti Ranjan Panda, M. Balakrishnan, Anant Vishnoi
IEEE Trans. Computers1
2010 Enhancing post-silicon processor debug with Incremental Cache state Dumping
abstract
During post-silicon validation/debug of processors, it is common to alternate between two phases: processor execution and state dump. The state dump, where the entire processor state is dumped off-chip to a logic analyzer for further processing, is a major bottleneck. We present a technique for improving debug efficiency by reducing the volume of cache data dumped off-chip, while still capturing the complete state. The reduction is achieved by introducing hardware mechanisms to transmit only the portion of the cache that was updated since the last dump. We propose two design alternatives based on whether or not the processor is permitted to continue execution during the dump: Blocking Incremental Cache Dumping (BICD) and Non-blocking Incremental Cache Dumping (NICD). We observe a 64% reduction in overall cache lines dumped and the dump time reduces to an average of 16.8% and 0.0002% for BICD and NICD respectively.
Preeti Ranjan Panda, Anant Vishnoi, M. Balakrishnan
VLSI-SoC1
2009 Online cache state dumping for processor debug
abstract
Post-silicon processor debugging is frequently carried out in a loop consisting of several iterations of the following two key steps: (i) processor execution for some duration, followed by (ii) dumping out of the processor's internal state into an external logic analyzer for further offline processing. Internal state of the processor is dominated by the L2 cache. During the process of dumping the cache content, the processor's execution is halted so that the state can be faithfully reproduced offline. In order to reduce the duration for which the processor is halted, and indirectly reduce debug time, we propose two Online Cache Dumping strategies, Retransmit Non-dumped Line (RNL) and Dump History Table (DHT), with the objective of transferring the cache contents while the processor is executing, and yet maintaining fidelity of the dumped data. For typical experimental debug scenarios, we observe that the effective dump times are reduced to between 0.01% and 3.5% of the original times. We also employ compression to reduce the cache content transfer time and logic analyzer space. Our experiments indicate an average compression ratio of 59.2%.
Anant Vishnoi, Preeti Ranjan Panda, M. Balakrishnan
DAC2
2009 A generic platform for estimation of multi-threaded program performance on heterogeneous multiprocessors
abstract
This paper deals with a methodology for software estimation to enable design space exploration of heterogeneous multiprocessor systems. Starting from fork-join representation of application specification along with high level description of multiprocessor target architecture and mapping of application components onto architecture resource elements, it estimates the performance of application on target multiprocessor architecture. The methodology proposed includes the effect of basic compiler optimizations, integrates light weight memory simulation and instruction mapping for complex instruction to improve the accuracy of software estimation. To estimate performance degradation due to contention for shared resources like memory and bus, synthetic access traces coupled with interval analysis technique is employed. The methodology has been validated on a real heterogeneous platform. Results show that using estimation it is possible to predict performance with average errors of around 11%.
Aryabartta Sahu, M. Balakrishnan, Preeti Ranjan Panda
DATE3
2009 Cache aware compression for processor debug support
abstract
During post-silicon processor debugging, we need to frequently capture and dump out the internal state of the processor. Since internal state constitutes all memory elements, the bulk of which is composed of cache, the problem is essentially that of transferring cache contents off-chip, to a logic analyzer. In order to reduce the transfer time and save expensive logic analyzer memory, we propose to compress the cache contents on their way out. We present a hardware compression engine for cache data using a Cache Aware Compression strategy that exploits knowledge of the cache fields and their behavior to achieve an effective compression. Experimental results indicate that the technique results in 7-31% better compression than one that treats the data as just one long bit stream. We also describe and evaluate a parallel compression architecture that uses multiple compression engines, resulting in a 54% reduction in transfer time.
Anant Vishnoi, Preeti Ranjan Panda, M. Balakrishnan
DATE2
2008 REWIRED - Register Write Inhibition by Resource Dedication
abstract
We propose REWIRED (register write inhibition by resource dedication), a technique for reducing power during high level synthesis (HLS) by selectively inhibiting the storage of function unit (FU) output data into registers. Registers are generally inferred in HLS when data produced in one clock cycle is used in a later cycle. However, when it can be established that the input registers to an FU are not changing values during a certain period, the outputs during this period can be directly read off the FU output pins without needing to store them in registers. When the life-times of such data are short, it may be possible to completely eliminate the register storage operation, thereby reducing power. We present a genetic algorithm formulation and a heuristic for maximizing the number of register stores that can be inhibited in a scheduled data flow graph (DFG) during behavioral synthesis.
Pushkar Tripathi, Rohan Jain, Srikanth Kurra, Preeti Ranjan Panda
ASP-DAC4
2008 Texture filter memory: a power-efficient and scalable texture memory architecture for mobile graphics processors
abstract
With increasing interest in sophisticated graphics capabilities in mobile systems, energy consumption of graphics hardware is becoming a major design concern in addition to the traditional performance enhancement criteria. Among the different steps in the graphics processing pipeline, we have observed that memory accesses during texture mapping - a highly memory intensive phase - contribute 30–40% of the energy consumed in typical embedded graphics processors. This makes the texture mapping subsystem an attractive candidate for energy optimization. We argue that a standard cache hierarchy, commonly used by researchers and commercial graphics processors for texture mapping, is wasteful of energy, and propose the Texture Filter Memory, an energy efficient architecture that exploits locality and the relatively high degree of predictability in texture memory access patterns. Our architecture consumes 75% lesser energy for texturing in a fixed function pipeline, incurring no performance overhead and a small area overhead over conventional texture mapping hardware.
B. V. N. Silpa, Anjul Patney, Tushar Krishna, Preeti Ranjan Panda, G. S. Visweswaran
ICCAD4
2007 The impact of loop unrolling on controller delay in high level synthesis
abstract
Loop unrolling is a well-known compiler optimization that can lead to significant performance improvements. When used in high level synthesis (HLS) unrolling can affect the controller complexity and delay. We study the effect of the loop unrolling factor on the delay of controllers generated during HLS. We propose a technique to predict controller delay as a function of the loop unrolling factor, and use this prediction with other search space pruning methods to automatically determine the optimal loop unrolling factor that results in a controller whose delay fits into a specified time budget, without an exhaustive exploration. Experimental results indicate delay predictions that are close to measured delays, yet significantly faster than exhaustive synthesis
Srikanth Kurra, Neeraj Kumar Singh 0004, Preeti Ranjan Panda
DATE3
2007 An Efficient Pipelined VLSI Architecture for Lifting-Based 2D-Discrete Wavelet Transform
abstract
The discrete wavelet transform (DWT) forms the core of the JPEG2000 image compression algorithm. JPEG2000 standard defines an irreversible DWT by a lifting scheme of factorized coefficients using (9, 7) Daubechies coefficients. The paper proposed optimizations on lifting based 1D-DWT data flow graph resulting in a new pipelining scheme, which is more power-efficient than existing approaches. The authors have shown that the constant multipliers defined by JPEG2000 standard, have latency of about 1.6 times the general adder latency. Hence, while finding the critical path, the basic assumption of multiplier latency being much greater than the adder latency does not hold for DWT in JPEG2000. How 75% multiplications can be reduced at the scaling step by this 2D-DWT implementation was also shown. This multiplication reduction not only saves power but also results in area saving by eliminating three multipliers in the hardware
Rahul Jain 0004, Preeti Ranjan Panda
ISCAS2
2006 Abridged addressing: a low power memory addressing strategy
abstract
The memory subsystem is known to comprise a significant fraction of the power dissipation in embedded systems. The memory addressing strategy, which determines the sequence of addresses appearing on the memory address bus as well as the switching activity in the addressing logic, has a major impact on the memory subsystem power dissipation. We present a novel addressing strategy, Abridged Addressing, that helps reduce system power dissipation by substantially reducing both the address bus switching as well the addressing logic power. The strategy, which relies on minimizing register accesses in the addressing logic, helps overcome some of the limitations of existing approaches: the address bus switching is low; there is very little area, performance, and power overhead; and the addressing hardware is simpler, making the technique suitable for both on-chip and off-chip memory, as well as single-port and multi-port memories
Preeti Ranjan Panda
ASP-DAC1
2006 Rapid estimation of control delay from high-level specifications
abstract
We address the problem of estimating controller delay from high-level specifications during behavioral synthesis. Typically, the critical path of a synthesised behavioral design goes through both the datapath and the control logic; yet most scheduling algorithms account only for datapath and ignore control delay, leading to timing uncertainties in the resulting designs. We present an estimation technique for computing a fast, robust, scalable, and reasonably accurate approximation of the control delay from behavioral specifications. The delay estimate is formulated in terms of the properties of the input specification and other inputs to the synthesis process such as resource constraints.
Gagan Raj Gupta 0001, Madhur Gupta, Preeti Ranjan Panda
DAC3
2005 Evaluation of Bus Based Interconnect Mechanisms in Clustered VLIW Architectures
abstract
With new sophisticated compiler technology, it is possible to schedule distant instructions efficiently. As a consequence, the amount of exploitable instruction level parallelism (ILP) in applications has gone up considerably. However, monolithic register file VLIW architectures present scalability problems due to a centralized register file which is far slower than the functional units (FU). Clustered VLIW architectures, with a subset of FU connected to any RF are the solution to this scalability problem. Recent studies with a wide variety of inter-cluster interconnection mechanisms have presented substantial gains in performance (number of cycles) over the most studied RF-to-RF type interconnections. However, these studies have compared only one or two design points in the RF-to-RF interconnects design space. In this paper, we extend the previous reported work. We consider both multi-cycle and pipelined buses. To obtain realistic bit latencies, we synthesized the various architectures and found out post layout clock periods. The results demonstrate that while there is very little variation in interconnect area, all the bus based architectures are heavily performance constrained. Also, neither multi-cycle nor pipelined buses or increasing the number of buses itself is able to achieve performance comparable to point-to-point type interconnects.
Anup Gangwar, M. Balakrishnan, Preeti Ranjan Panda
DATE3
2003 Memory allocation and mapping in high-level synthesis - an integrated approach
abstract
With the increasing design complexity and performance requirement, data arrays in behavioral specification are usually mapped to memories in behavioral synthesis. This paper describes a new algorithm that overcomes two limitations of the previous works on the problem of memory-allocation and array-mapping to memories. Specifically, its key features are a tight link to the scheduling effect, which was totally or partially ignored by the existing memory synthesis systems, and supporting nonuniform access speeds among the ports of memories, which greatly diversify the possible (practical) memory configurations. Experimental data on a set of benchmark filter designs are provided to show the effectiveness of the proposed exploration strategy in finding globally best memory configurations.
Jaewon Seo, Preeti Ranjan Panda
IEEE Trans. Very Large Scale Integr. Syst.3
2002 An integrated algorithm for memory allocation and assignment in high-level synthesis
abstract
With the increasing design complexity and performance requirement, data arrays in behavioral specification are usually mapped to fast on-chip memories in behavioral synthesis. This paper describes a new algorithm that overcomes two limitations of the previous works on the problem of memory-allocation and array-mapping to memories. Specifically, its key features are (1) a tight link to the scheduling effect, which was totally or partially ignored by the existing memory synthesis systems, and supporting (2) non-uniform access speeds among the ports of memories, which greatly diversify the possible (practical) memory configurations. Experimental data on a set of benchmark filter designs are provided to show the effectiveness of the proposed exploration strategy in finding globally best memory configurations.
Jaewon Seo, Preeti Ranjan Panda
DAC3
2002 Memory Architectures for Embedded Systems-On-Chip
Preeti Ranjan Panda, Nikil Dutt
HiPC1
2002 An energy-conscious algorithm for memory port allocation
abstract
Multiport memories are extensively used in modern system designs because of the performance advantages they offer. The increased memory access throughput could lead to significantly faster schedules in behavioral synthesis. However, they also have an associated area and energy penalty. We describe a technique for mapping data accesses to multiport memories during behavioral synthesis that results in significantly better energy characteristics than an unoptimized multiport design. The technique consists of an initial colouring of the array access nodes in the data flow graph based on spatial locality, followed by attempts to consecutively access memory locations with the same colour on the same port. Our experiments on several applications indicate a significant reduction in address bus switching activity, leading to an overall energy reduction over an unoptimized design, while still maintaining a performance advantage over a single-port solution.
Preeti Ranjan Panda, Lakshmikantam Chitturi
ICCAD1
2001 Data and memory optimization techniques for embedded systems
abstract
We present a survey of the state-of-the-art techniques used in performing data and memory-related optimizations in embedded systems. The optimizations are targeted directly or indirectly at the memory subsystem, and impact one or more out of three important cost metrics: area, performance, and power dissipation of the resulting implementation. We first examine architecture-independent optimizations in the form of code transoformations. We next cover a broad spectrum of optimization techniques that address memory architectures at varying levels of granularity, ranging from register files to on-chip memory, data caches, and dynamic memory (DRAM). We end with memory addressing related issues.
Preeti Ranjan Panda, Francky Catthoor, Nikil Dutt, Koen Danckaert, Erik Brockmeyer, Chidamber Kulkarni, Arnout Vandecappelle, Per Gunnar Kjeldsberg
ACM Trans. Design Autom. Electr. Syst.1
2000 On-chip vs. off-chip memory: the data partitioning problem in embedded processor-based systems
abstract
Efficient utilization of on-chip memory space is extremely important in modern embedded system applications based on processor cores. In addition to a data cache that interfaces with slower off-chip memory, a fast on-chip SRAM, called Scratch-Pad memory, is often used in several applications, so that critical data can be stored there with a guaranteed fast access time. We present a technique for efficiently exploiting on-chip Scratch-Pad memory by partitioning the application's scalar and arrayed variables into off-chip DRAM and on-chip Scratch-Pad SRAM, with the goal of minimizing the total execution time of embedded applications. We also present extensions of our proposed memory assignment strategy to handle context switching between multiple programs, as well as a generalized memory hierarchy. Our experiments on code kernels from typical applications show that our technique results in significant performance improvements.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
ACM Trans. Design Autom. Electr. Syst.1
1999 Memory bank customization and assignment in behavioral synthesis
abstract
With increasing design complexity and chip area, on-chip memory has become an important component whose integration needs to be addressed during system design. Modern embedded DRAM technology allows for large amounts of on-chip memory space. However, in order to utilize the available memory intelligently, the memory has to be appropriately customized for the specific application. We address the topic of incorporating the application-specific customization of memory bank configuration into behavioral synthesis. The strategy involves a partitioning of behavioral arrays into memory banks based on a cost function that estimates the performance implications. For a given candidate partition, we present a heuristic for determining the access sequence that minimizes page misses in a bank while respecting data dependences. The output of the exploration is a graph displaying the variation of delay and memory area with the bank configuration. Our experiments on several memory-intensive examples confirm that the exploration results can provide critical feedback to the designer about the optimal memory configuration for a given application.
Preeti Ranjan Panda
ICCAD1
1999 Augmenting Loop Tiling with Data Alignment for Improved Cache Performance
abstract
Loop blocking (tiling) is a well-known compiler optimization that helps improve cache performance by dividing the loop iteration space into smaller blocks (tiles); reuse of array elements within each tile is maximized by ensuring that the working set for the tile fits into the data cache. Padding is a data alignment technique that involves the insertion of dummy elements into a data structure for improving cache performance. In this work, we present DAT, a technique that augments loop tiling with data alignment, achieving improved efficiency (by ensuring that the cache is never under-utilized) as well as improved flexibility (by eliminating self-interference cache conflicts independent of the tile size). This results in a more stable and better cache performance than existing approaches, in addition to maximizing cache utilization, eliminating self-interference, and minimizing cross-interference conflicts. Further, while all previous efforts are targeted at programs characterized by the reuse of a single array, we also address the issue of minimizing conflict misses when several tiled arrays are involved. To validate our technique, we ran extensive experiments using both simulations as well as actual measurements on SUN Sparc5 and Sparc10 workstations. The results on benchmarks exhibiting varying memory access patterns demonstrate the effectiveness of our technique through consistently high hit ratios and improved performance across varying problem sizes.
Preeti Ranjan Panda, Hiroshi Nakamura, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Computers1
1999 Local memory exploration and optimization in embedded systems
abstract
Embedded processor-based systems allow for the tailoring of the on-chip memory architecture based on application specific requirements. We present an analytical strategy for exploring the on-chip memory architecture for a given application, based on a memory performance estimation scheme. The analytical technique has the important advantage of enabling a fast evaluation of candidate memory architectures in the early stages of system design. Many digital signal-processing applications involve array accesses and loop nests that can benefit from such an exploration. Our experiments demonstrate that our estimations closely follow the actual simulated performance at significantly reduced run times.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1999 Low-power memory mapping through reducing address bus activity
abstract
Arrays in behavioral specifications that are too large to fit into on-chip registers are usually mapped to off-chip memories during behavioral synthesis. We address the problem of system power reduction through transition count minimization on the memory address bus when these arrays are accessed from memory. We exploit regularity and spatial locality in the memory accesses and determine the mapping of behavioral array references to physical memory locations to minimize address bus transitions. We describe array mapping strategies for two important memory configurations: all behavioral arrays mapped to a single off-chip memory and arrays mapped into multiple memory modules drawn from a library. For the single memory configuration, we describe a heuristic for selecting a memory mapping scheme to achieve low power for each behavioral array. For mapping into a library of multiple memory modules, we formulate the problem as three logical-to-physical memory mapping subtasks and present experiments demonstrating the transition count reductions based on our approach. Our experiments on several image processing benchmarks show power savings of up to 63% through reduced transition activity on the memory address bus in the single memory case. We also observe a further transition count reduction by a factor of 1.5-6.7 over a straightforward mapping scheme in the multiple memories configuration.
Preeti Ranjan Panda, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.1
1998 Data Cache Sizing for Embedded Processor Applications
abstract
We present a technique for determining the best data cache size required for a given memory-intensive application. A careful memory and cache line assignment strategy based on the analysis of the array access patterns effects a significant reduction in the required data cache size, with no negative impact on the performance, thereby freeing vital on-chip silicon area for other hardware resources. Experiments on several benchmark kernels performed on LSI Logic's CW4001 embedded processor simulator confirm the soundness of our cache sizing and memory assignment strategy and the accuracy of our analytical predictions.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
DATE1
1998 Incorporating DRAM access modes into high-level synthesis
abstract
Memory-intensive behaviors often contain large arrays that are synthesized into off-chip memories. With the increasing gap between on-chip and off-chip memory access delays, it is imperative to exploit the efficient access mode features of modern-day memories (e.g., page-mode DRAM's) in order to alleviate the memory bandwidth bottleneck. Although recent research efforts in high-level synthesis (HLS) have addressed the issue of memory-based synthesis, current techniques are unable to exploit efficiently the special access modes of these off-chip memories, resulting in significantly inferior performance using these memory library parts. Our work addresses this issue by (a) modeling realistic off-chip memory access modes for HLS, (b) presenting algorithms to infer applicability of HLS with these memory access modes, and (c) transforming input behavior to provide further memory access optimizations during HLS. We demonstrate the utility of our approach using a suite of memory-intensive benchmarks with a realistic DRAM library module. Experimental results show a significant performance improvement (more than 40%) as a result of our optimization techniques.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1997 Exploiting off-chip memory access modes in high-level synthesis
abstract
Memory-intensive behaviors often contain large arrays that are synthesized into off-chip memories. With the increasing gap between on-chip and off-chip memory access delays, it is imperative to exploit the efficient access mode features of modern-day memories (e.g. page-mode DRAMs) in order to alleviate the memory bandwidth bottleneck. Our work addresses this issue by: (a) modeling realistic off-chip memory access modes for High-level Synthesis (HLS), (b) presenting algorithms to infer applicability of HLS with these memory access modes, and (c) transforming input behavior to provide further memory access optimizations during HLS. We demonstrate the utility of our approach using a suite of memory-intensive benchmarks with a realistic DRAM library module. Experimental results show a significant performance improvement (more than 40%) as a result of our optimization techniques.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
ICCAD1
1997 A Data Alignment Technique for Improving Cache Performance
abstract
We address the problem of improving the data cache performance of numerical applications-specifically, those with blocked (or tiled) loops. We present DAT, a data alignment technique utilizing array-padding, to improve program performance through minimizing cache conflict misses. We describe algorithms for selecting tile sizes for maximizing data cache utilization, and computing pad sizes for eliminating self-interference conflicts in the chosen tile. We also present a generalization of the technique to handle applications with several tiled arrays. Our experimental results comparing our technique with previous published approaches on machines with different cache configurations show consistently good performance on several benchmark programs, for a variety of problem sizes.
Preeti Ranjan Panda, Hiroshi Nakamura, Nikil Dutt, Alexandru Nicolau
ICCD1
1997 Memory data organization for improved cache performance in embedded processor applications
abstract
Code generation for embedded processors opens up the possibility for several performance optimization techniques that have been ignored by traditional compilers due to compilation time constraints. We present techniques that take into account the parameters of the data caches for organizing scalar and array variables declared in embedded code into memory, with the objective of improving data cache performance. We present techniques for clustering variables to minimize compulsory cache misses, and for solving the memory assignment problem to minimize conflict cache misses. Our experiments with benchmark code kernels from DSP and other domains on the CW4001 embedded processor from LSI Logic indicate significant improvements in data cache performance by the application of our memory organization technique.
Preeti Ranjan Panda, Nikil Dutt, Alexandru Nicolau
ACM Trans. Design Autom. Electr. Syst.1
1996 Low-power mapping of behavioral arrays to multiple memories
abstract
Large data arrays in behavioral specifications are usually mapped to off-chip memories during system synthesis. We address the problem of system power reduction through transition count minimization on the address bus during memory accesses, when mapping behavioral arrays to multiple memory modules drawn from a library. We formulate the problem as three logical-to-physical memory mapping subtask, provide algorithms for each subtask, and present experiments that demonstrate the transition count reductions based on our approach. Our experiments show a transition count reduction by a factor of 1.5-6.7 over a straightforward mapping scheme.
Preeti Ranjan Panda, Nikil Dutt
ISLPED1
1991 A Flexible Scheme for State Assignment Based on Characteristics of the FSM
abstract
The authors present a novel scheme for state assignment based on the premise that there is a strong correlation between the FSM (finite state machine) characteristics and the state assignment technique to be used. Based on the nature of the FSM, one of four state assignment schemes is selected by the system. The selection of one of these techniques is done automatically by the system. An option is provided to further optimize the generated solution using simulated annealing. Results on MCNC benchmarks indicate that the flexible methodology for state assignment leads to area and delay values that are better in most cases than those obtained by using existing state assignment schemes.>
Biswadip Mitra, Preeti Ranjan Panda, Parimal Pal Chaudhuri
ICCAD2