EDBT 2026 Demo / reviewers in the wild / expert
Narasinga Rao Miniskar
dblp:82/7182
· DBLP profile ↗
22ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0001-8259-8891ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 10 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PowerMappeR: Power-Optimized Mapping of SNNs onto ReRAM Crossbars coupled via Packet-Switched NoCs
Devin Pohl, Kazi Asifuzzaman, Aaron R. Young, Narasinga Rao Miniskar, Jeffrey S. Vetter |
IPDPS | 4 |
| 2026 | Accl++ : A high-productivity programming language for performance and code portability on heterogeneous systemsabstractThis work describes the Accl++ programming language for heterogeneous computing. Accl++ is embedded in the C++ language and implemented as a C++ library. The language allows for describing device code and execution libraries and includes primitives for runtime compilation (RTC), thereby enabling code portability across disparate devices. Here, we demonstrate Accl++’s capability by coding different benchmarks from different domains. This work also describes an analysis of the overheads introduced by the Accl++ RTC support as well as the performance of the generated code for the Accl++ kernels. Accl++ improves heterogeneous code portability without incurring high levels of overhead. In two different heterogeneous systems—one composed of two 32-core AMD EPYC 7513 CPUs and two NVIDIA A100 GPUs and the other composed of two 12-core AMD EPYC 7272 CPUs and two AMD MI100 GPUs—Accl++ enables the execution of the same binary application with observed RTC overheads in the range of 3%–10% of the total kernel execution time, resulting in performance levels similar to those of native CUDA/HIP/OpenCL code. Marc González 0001, Pedro Valero-Lara, Mohammad Alaul Haque Monil, Seyong Lee, Beau Johnston, Aaron R. Young, Narasinga Rao Miniskar, Keita Teranishi, Jeffrey S. Vetter |
Future Gener. Comput. Syst. | 7 |
| 2025 | POSTER: IRISX: A Dynamic Trade-off System for Performance Portability on Multi-Accelerator PlatformsabstractContemporary high performance computing, cloud, and embedded systems are equipped with various combinations of heterogeneous processors. This has led to the development of portable abstractions to target these heterogeneous systems. However, these abstractions fail to deliver performance portability as they cannot easily modify key parameters such as the amount of concurrency, the representation of the computation graph, the kernel implementation, and data transfers for varying problem sizes and hardware architectures. To address this wide parameter space, this work presents IRISX, a framework that dynamically finds the performance portable solution based on application characteristics and underlying hardware. By changing the representation of the computation at the kernel level, task level, and graph level, IRISX finds the appropriate set of heterogeneous processors that provides the best performance while ensuring portability to different heterogeneous systems. Sanil Rao, Mohammad Alaul Haque Monil, Het Mankad, Narasinga Rao Miniskar, Keita Teranishi, Jeffrey S. Vetter, Franz Franchetti |
PACT | 4 |
| 2025 | Mapping Spiking Neural Networks to Heterogeneous Crossbar Architectures using Integer Linear ProgrammingabstractAdvances in novel hardware devices and architectures allow Spiking Neural Network (SNN) evaluation using ultra-low power, mixed-signal, memristor crossbar arrays. As individual network sizes quickly scale beyond the dimensional capabilities of single crossbars, networks must be mapped onto multiple crossbars. Crossbar sizes within modern Memristor Crossbar Architectures (MCAs) are determined predominately not by device technology but by network topology; more, smaller crossbars consume less area thanks to the high structural sparsity found in larger, brain-inspired SNNs. Motivated by continuing increases in SNN sparsity due to improvements in training methods, we propose utilizing heterogeneous crossbar sizes to further reduce area consumption. This approach was previously unachievable as prior compiler studies only explored solutions targeting homogeneous MCAs. Our work improves on the state-of-the-art by providing Integer Linear Programming (ILP) formulations supporting arbitrarily heterogeneous architectures. By modeling axonal interactions between neurons, our methods produce better mappings while removing inhibitive a priori knowledge requirements. We first show a 16.7-27.6% reduction in area consumption for square-crossbar homogeneous architectures. Then, we demonstrate 66.9-72.7% further reduction when using a reasonable configuration of heterogeneous crossbar dimensions. Next, we present a new optimization formulation capable of minimizing the number of inter-crossbar routes. When applied to solutions already near-optimal in area, an 11.9-26.4% routing reduction is observed without impacting area consumption. Finally, we present a profile-guided optimization capable of minimizing the number of runtime spikes between crossbars. Compared to the best-area-then-route optimized solutions, we observe a further 0.5-14.8% inter-crossbar spike reduction while requiring 1–3 orders of magnitude less solver time. Devin Pohl, Aaron R. Young, Kazi Asifuzzaman, Narasinga Rao Miniskar, Jeffrey S. Vetter |
DATE | 4 |
| 2025 | ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors
Kazi Asifuzzaman, Aaron R. Young, Prasanna Date, Shruti R. Kulkarni, Narasinga Rao Miniskar, Matthew J. Marinella, Jeffrey S. Vetter |
Euro-Par (2) | 5 |
| 2025 | IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous ComputingabstractIn the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems. Narasinga Rao Miniskar, Aaron R. Young, Mohammad Alaul Haque Monil, Kazi Asifuzzaman, Beau Johnston, Keita Teranishi, Jeffrey S. Vetter |
ICPP | 1 |
| 2024 | CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous SystemsabstractPerformance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems. Norihisa Fujita, Beau Johnston, Narasinga Rao Miniskar, Ryohei Kobayashi 0001, Mohammad Alaul Haque Monil, Keita Teranishi, Seyong Lee, Jeffrey S. Vetter, Taisuke Boku |
e-Science | 3 |
| 2023 | A 3D Implementation of Convolutional Neural Network for Fast InferenceabstractLow latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology. Narasinga Rao Miniskar, Pruek Vanna-Iampikul, Aaron R. Young, Sung Kyu Lim, Frank Liu 0001, Jieun Yoo, Corrinne Mills, Farah Fahim, Jeffrey S. Vetter |
ISCAS | 1 |
| 2022 | Ultra Low Latency Machine Learning for Scientific Edge ApplicationsabstractIn this paper, we present an FPGA design of an extremely low latency scientific machine learning application at the edge. Real-time prediction of errant high-energy particle beams at scientific facilities such as Spallation Neutron Source (SNS) is crucial to avoid damages to the equipment. Machine learning techniques are becoming increasingly effective to detect subtle signatures of the errant beams in the noisy sensor signals. However, to minimize potential damage done by errant beam, real-time errant beam detection has to be completed with extremely low latency, usually less than 1 microsecond. By stream processing the input features and employing out-of-order execution of decision nodes among the decision trees, we demonstrate that our highly efficient FPGA implementation can achieve 60 nanoseconds of computing latency for complex random forest models with 10,000 input features. Narasinga Rao Miniskar, Aaron R. Young, Frank Liu 0001, Willem Blokland, Anthony M. Cabrera, Jeffrey S. Vetter |
FPL | 1 |
| 2022 | IRIS-BLAS: Towards a Performance Portable and Heterogeneous BLAS LibraryabstractThis paper presents IRIS-BLAS, a novel heterogeneous and performance portable BLAS library. IRIS-BLAS is built on top of the IRIS runtime and multiple vendor and open-source BLAS libraries. It can transparently use all the architectures/devices available in a heterogeneous system, using the appropriate BLAS library based on the task mapping at run time. Thus, IRIS-BLAS is portable across a broad spectrum of architectures and BLAS libraries, alleviating the worry of application developers about modifying the application source code. Even though the emphasis is on portability, IRIS-BLAS provides competitive or even better performance than other state-of-the-art references. Moreover, IRIS-BLAS offers new features such as efficiently using extremely heterogeneous systems composed of multiple GPUs from different hardware vendors. Narasinga Rao Miniskar, Mohammad Alaul Haque Monil, Pedro Valero-Lara, Frank Liu 0001, Jeffrey S. Vetter |
HIPC | 1 |
| 2021 | A Memory Efficient Lock-Free Circular QueueabstractHardware queues are import in many applications, such as data transfer, synchronization of concurrent modules with the need of mutual exclusion constructs. State of the art bounded (of a fixed size) lock free circular queues are implemented either by read/write atomic operations, or barrier conditions, or by separating dequeue and enqueue operations. However, these queues always require an unused element at all the times to safeguard the front and rear pointers of the queue, so as to avoid data race conditions, which leads to the waste of memory. The waste of memory is especially disadvantageous in applications such as I/O data transfer, and image transfer between processing filters, when large element size is needed, We propose a lock- free solution of the bounded circular queue through read/write atomic operations, but without the need of an extra element in the queue. The proposed solution is implemented and verified in both Verilog and 'C' languages. We also demonstrate its effectiveness by comparing its area and delay metrics with the implementations of other existing designs of queue. Narasinga Rao Miniskar, Frank Liu 0001, Jeffrey S. Vetter |
ISCAS | 1 |
| 2020 | Deffe: a data-efficient framework for performance characterization in domain-specific computingabstractAs the computer architecture community moves toward the end of traditional device scaling, domain-specific architectures are becoming more pervasive. Given the number of diverse workloads and emerging heterogeneous architectures, exploration of this design space is a constrained optimization problem in a high-dimensional parameter space. In this respect, predicting workload performance both accurately and efficiently is a critical task for this exploration. In this paper, we present Deffe: a framework to estimate workload performance across varying architectural configurations. Deffe uses machine learning to improve the performance of this design space exploration. By casting the work of performance prediction itself as transfer learning tasks, the modelling component of Deffe can leverage the learned knowledge on one workload and "transfer" it to a new workload. Our extensive experimental results on a contemporary architecture toolchain (RISC-V and GEM5) and infrastructure show that the method can achieve superior testing accuracy with an effective reduction of 32-80× in terms of the amount of required training data. The overall run-time can be reduced from 400 hours to 5 hours when executed over 24 CPU cores. The infrastructure component of Deffe is based on scalable and easy-to-use open-source software components. Frank Liu 0001, Narasinga Rao Miniskar, Dwaipayan Chakraborty, Jeffrey S. Vetter |
CF | 2 |
| 2019 | Low Complex & High Accuracy Computation Approximations to Enable On-Device RNN ApplicationsabstractRecurrent Neural Networks (RNN) have demonstrated excellent results for various Automatic Speech Recognition (ASR) and Natural Language Processing (NLP) tasks. However, executing RNNs requires huge memory and computations which makes it difficult to achieve real time performance on low power devices like smartphones. Hence, currently ASR and NLP applications such as voice assistants are using cloud based solutions. In this paper, to enable on-device inference, we propose efficient approximations for weights of FC layers and activation functions to reduce the computational complexity. The proposed approximations eliminate multiplications, divisions and exponential operations by replacing them with simple arithmetic operations (shifts, additions) to significantly reduce the computation requirements without any perceivable loss of functional accuracy. The approximations also reduce the memory size and bandwidth requirements. We also present a lightweight VLIW based DSP architecture with these approximations to enable on-device inference. The approximations have been tested on the proposed DSP with various RNN applications like EESEN, LRCN and S2VT. The results with approximations show - accuracies similar to that of float (32-bit) reference, ~ 8x-12× performance gains, ~ 2x-4x gains in memory requirement and bandwidth. Moreover, the activation approximation results show better average and peak errors compared to the State of the Art. Sirish Kumar Pasupuleti, Raj Narayana Gadde, Vasanthakumar Rajagopal, Ashok Vishnoi, N. Chandra Sekhar, R. Chandra Kumar, Narasinga Rao Miniskar |
ISCAS | 7 |
| 2018 | An Intelligent Bandwidth Manager for CNN Applications on Embedded DevicesabstractAdapting complex Convolution Neural Network (CNN) applications on embedded processors is a challenge due to the massive memory bandwidth and computational requirements. In particular, the CNN memory bandwidth requirement poses a huge challenge for the processors with Scratch Pad Memory (SPM), usually of limited size. In this paper, we present an Intelligent Bandwidth Manager (IBWM) to efficiently handle the CNN bandwidth for SPM based processors. The proposed IBWM is a two fold approach which includes Intelligent SPM Manager (ISM) to optimize the number of accesses to SDRAM by analysing the data patterns, and Feature Map Compression (FMC) to further reduce the bandwidth by exploiting the feature map data sparsity. The IBWM is independent of any processor architecture and can be adopted in any processor with SPM. The proposed IBWM is experimented with ResNet-50 [1] and AlexNet [2] networks on a Samsung Reconfigurable Processor (SRP) [3] for various SPM sizes. The SDRAM bandwidth results show, 2x improvement compared to MIT Eyeriss [4] for AlexNet, and 4x-8x improvement compared to primitive bandwidth management techniques for AlexNet and ResNet-50. The proposed method achieves the bandwidth closer to the minimum possible bandwidth. Sirish Kumar Pasupuleti, Aishwarya Rajaram, Narasinga Rao Miniskar, Raj Narayana Gadde, Deepanshu Yadvandu, Vasanthakumar Rajagopal, Ashok Vishnoi, Chandra Kumar Ramasamy |
ICIP | 3 |
| 2018 | Accurate and Efficient Fixed Point Inference for Deep Neural NetworksabstractDeploying DNNs on embedded devices is a challenge because of their high memory and computational requirements. Performing DNN inference in lesser bit-width fixed point arithmetic is seen as a crucial step in realizing DNNs on embedded devices. State-of-the-art methods achieve floating point accuracy using re-training and complex activation normalization methods. In this paper we propose an accurate and efficient end-to-end DNN inference on 16-bit fixed point arithmetic. We prove that floating point accuracy can be achieved with a simple quantization method of using powers of 2 as scale factors coupled with our optimal bit-width estimation algorithm without using re-training. Additionally, it leads to efficient activation normalization using only arithmetic shifts. We show that the combination of our quantization method and activation normalization maximizes SIMD throughput resulting in 2× to 6× gain in execution time compared to floating point inference. Experimental results demonstrate that our method generalizes to different networks giving same or better accuracy compared to floating point for classification, regression and recurrent networks. Vasanthakumar Rajagopal, Chandra Kumar Ramasamy, Ashok Vishnoi, Raj Narayana Gadde, Narasinga Rao Miniskar, Sirish Kumar Pasupuleti |
ICIP | 5 |
| 2018 | Optimal SDRAM Buffer Allocator for Efficient Reuse of Layer IO in CNNs Inference FrameworkabstractDeep Learning based applications are becoming increasingly ubiquitous. The new generation smart phones are adapting lot of applications built on deep learning technology. However, adapting complex Deep Neural Network (DNN) applications on embedded processors is a huge challenge not only due to huge computational requirement, but also due to the massive SDRAM memory requirements of network layer IO buffers. Hence, an efficient reuse of layer IO buffers is required. However, it is challenging to reuse layer IO buffers because of complex network topology and large number of layers in the network. In this paper, we present an optimal SDRAM buffer allocator to minimize the overall SDRAM memory requirement of layer IO buffers, which works for any complex networks. The proposed SDRAM buffer allocator is integrated with ICNN (Inference only Convolutional Neural Networks) framework which is an extension of Caffe framework. Our framework with optimal SDRAM buffer allocator is experimented with popular AlexNet, GoogLeNet, ResNet-50 and Inception-ResNet-v2 CNNs. The results show 2× to 30× reduction in SDRAM footprint when compared with Caffe/TensorFlow frameworks and 26% to 47% reduction when compared with MXNet framework. Narasinga Rao Miniskar, Sirish Kumar Pasupuleti, Vasanthakumar Rajagopal, Ashok Vishnoi, Chandra Kumar Ramasamy, Raj Narayana Gadde |
ISCAS | 1 |
| 2017 | A novel method to regenerate an optimal CNN by exploiting redundancy patterns in the networkabstractDeploying Convolution Neural Networks (CNN) based computer vision applications on low-power embedded devices is challenging due to massive computation and memory bandwidth requirements. Research is on-going on faster algorithms, network pruning, and model compression techniques to produce light-weight networks. In this paper, we propose a novel method which exploits a redundancy pattern in the network to regenerate an efficient and functionally identical CNN for a given network. We identify the pattern based on the layer parameters (kernel size and stride) and data flow analysis among the layers to avoid the redundant processing and memory requirements while maintaining identical accuracy. Our proposed method augments the state-of-the-art pruning and model compression techniques to achieve further performance boost-up. The proposed method is experimented with the Caffe [1] framework for ResNet-50 [2] inference on Samsung smartphone with an octa-core ARM Cortex-A53 processor. The results show an improvement of 4x in performance and memory at layer level, ~22% performance improvement and 6% memory reduction at network level. Sirish Kumar Pasupuleti, Narasinga Rao Miniskar, Vasanthakumar Rajagopal, Raj Narayana Gadde |
ICIP | 2 |
| 2017 | Fast cycle-accurate compile based simulator for reconfigurable processorabstractReconfigurable Processor (RP) provides great flexibility of hardware re-configurability through software solution for high-performance computing. RP is used as a DSP in Samsung DTV and Camera to run the Audio, Video codecs and image quality enhancement algorithms. RP runs in two modes: VLIW (Very Large Instruction Word) and CGRA (Coarse Grain Reconfigurable Array). To minimize the time-to-market of products, application developers of the RP require a fast and cycle-accurate profiling-enabled simulator to verify the functionality of applications and to optimize the hot-spots (performance critical codes). Further, fast simulation is necessary to verify the functionalities for certification of various standards of multimedia (DivX, Dolby, etc.). The state-of-the-art RP simulators are very slow running at 2 MIPS (Million instructions per second) in CGRA mode and 4 MIPS in VLIW mode, running on x86 host processor at 3.4 GHz. We propose FastSim, a fast compile based cycle-accurate simulator for RP, which yields simulation speed of 900 MIPS in CGRA mode and 1600 MIPS in VLIW mode with > 99.5% core cycle accuracy. Thus, FastSim enables ~400x faster simulation when compared to existing RP simulators. FastSim in VLIW mode is 2x faster when compared to the state-of-the-art functional simulators, and 16x faster when compared to cycle-accurate simulators, of the VLIW and RISC processors in the industry. FastSim speed is also comparable to application running time on native x86. The faster simulation speed is achieved with the use of an innovative maximal static analysis in both VLIW and CGRA modes. Narasinga Rao Miniskar, Raj Narayana Gadde, Young-chul Rams Cho, Sukjin Kim |
ISCAS | 1 |
| 2016 | Intra mode power saving methodology for CGRA-based reconfigurable processor architecturesabstractReconfigurable processors (RP) such as Samsung Reconfigurable Processor (SRP) are best suited for wide range of embedded DSP application domains such as image, audio, video and vision processing. It provides great flexibility of hardware reconfigurability through software solution which provides high-performance computing with low-energy and fast time-to-market. Power consumption of the processor architecture is one of the important considerations for mobile DSP solutions. However, the performance critical code sections of applications mapped to RPs may not utilize all minicores (group of functional units, local register files and their connections) due to low ILP (Instruction Level Parallelism) even after exploring software pipelining by the compiler. The unused minicores can consume considerable amount of power (including power leakage) from configuration memory banks, functional units, local register files, and also from central register files of RP architectures. There are approaches in RP to power-gate the minicores when we switch from VLIW (Very Large Instruction Word) mode of RP operation to CGRA (Coarse Grained Reconfigurable Array) mode (inter-mode power-gating exploration). In this paper, we propose a novel programmer directive based power-gating technique for RP which explores unused resources during VLIW and CGRA mode of execution (intra-mode power-gating). Our approach has shown up to 33% power savings in the SRP CGRA mode and up to 56% power savings in SRP VLIW mode, with additional power-gating circuits that contribute to <;1% increase in die area. Narasinga Rao Miniskar, Rahul R. Patil, Raj Narayana Gadde, Young-chul Rams Cho, Sukjin Kim, Shi Hwa Lee |
ISCAS | 1 |
| 2014 | Retargetable automatic generation of compound instructions for CGRA based reconfigurable processor applicationsabstractReconfigurable processors such as SRP (Samsung Reconfigurable Processors) have become increasingly important, which enables just enough flexibility of accepting software solutions and providing application specific hardware configurability for faster time-to-market, lower development cost and higher performance while maintaining lower energy consumption and area. The reconfigurable processor compilation framework supports wide range of architectures through architecture description template for different domains of applications such as image processing, multimedia, video, and graphics. These architectures support several domain specific compound instructions (also called as intrinsics), which are computationally efficient when compared to the set of general instructions in the processor. Application developers have to use these intrinsics in their programs according to the architecture, which can result very inefficient usage, tedious and more error-prone. Moreover, the intrinsics provided by the architecture need constant reference to the intrinsics file during development. In this paper, we propose a retargetable novel methodology for the automatic generation of compound instructions for a given architecture and application source code at compile time. Our approach is able to consider ~75% of total intrinsics in the architectures with the success rate of > 90% in identifying the intrinsics in the benchmarks such as AVC, OpenGL Full Engine and OpenGL Vector benchmarks. Narasinga Rao Miniskar, Soma Kohli, Haewoo Park, Donghoon Yoo |
CASES | 1 |
| 2012 | Function inlining and loop unrolling for loop acceleration in reconfigurable processorsabstractThe next generation SoCs for consumer electronics need software solutions for faster time-to-market, lower development cost and higher performance while maintaining lower energy consumption and area. As a result, reconfigurable processors (RPs) have become increasingly important, which enables just enough exibility of accepting software solutions and providing application-specific hardware reconfigurability. Samsung Electronics has developed a reconfigurable processor called Samsung Reconfigurable Processor (SRP), which is the basis of our work. Though, the SRP is a powerful processor, it requires a smart and intelligent compiler to compile the application software while exploring its reconfigurable architecture. The existing compiler for the SRP does not support functional inlining and loop unrolling, and no study has yet been done on these optimizations for the RPs. In this paper, we study the impact of these optimizations on the performance of applications for the SRP processor and we also show how these optimizations are supported in the SRP compiler. We analyze the performance improvement due to these optimizations on various benchmarks namely Sobel Edge filter, JPEG decoder, and Luma Deblocking filter of the H.264 standard. Our experimental results have shown about 83% gain on performance with the functional inlining optimization and the loop unrolling optimization when compared to the original code for Sobel filter and JPEG encoder, and 11% gain on performance for Luma Deblock filter. Narasinga Rao Miniskar, Pankaj Shailendra Gode, Soma Kohli, Donghoon Yoo |
CASES | 1 |
| 2010 | PinComm: Characterizing Intra-application Communication for the Many-Core EraabstractAs the number of cores in both embedded Multi-Processor Systems-on-Chip and general purpose processors keeps rising, on-chip communication becomes more and more important. In order to write efficient programs for these architectures it is therefore necessary to have a good idea of the communication behavior of an application. We present a communication profiler that extracts this behavior from compiled, sequential or parallel C/C++ programs, and constructs a dynamic data-flow graph at the level of major functional blocks. In contrast to existing methods of measuring inter-program communication, our tool automatically generates the program's data-flow graph and is less demanding for the developer. It can also be used to view differences between program phases (such as different video frames), which allows both input- and phase-specific optimizations to be made. We will also describe briefly how this information can subsequently be used to guide the effort of parallelizing the application, to co-design the software, memory hierarchy and communication hardware, and to provide new sources of communication-related runtime optimizations. Wim Heirman, Dirk Stroobandt, Narasinga Rao Miniskar, Roel Wuyts, Francky Catthoor |
ICPADS | 3 |