M. Balakrishnan

dblp:24/4827 · DBLP profile ↗
← Back
53ranked-venue papers
10as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3Computer networks · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorTheory of computation · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 EXPRESS: A Framework for Execution Time Prediction of Concurrent CNNs on Xilinx DPU Accelerator
abstract
Deep learning Processor Unit (DPU) is a highly configurable CNN accelerator that supports a variety of CNNs and can be implemented with multiple instances on the same FPGA. Many applications deploy concurrent execution of different CNNs and in such a setting, an execution time predictor can help “optimize” the DPU configurations to meet the performance requirements of different tasks. We characterize CNN execution on DPUs and reduce the variability in execution time due to interference from the operating system. Subsequently, we propose a machine learning-based framework (EXPRESS) to predict the execution time of any given CNN on a DPU configuration, considering CNN, DPU, and bus characteristics. We improvise EXPRESS to support heterogeneous CNNs in EXPRESS-2.0 by making features independent of the number of CNNs. Our entire experimentation is based on data from a real FPGA board for 16 standard CNNs. Our frameworks, EXPRESS and EXPRESS-2.0, significantly outperform state-of-the-art by achieving an average execution time prediction error of 2.2% and 0.7%, respectively. We illustrate the effectiveness of this low prediction error for design space exploration, which is very useful for embedded system application developers.
Shikha Goel, Rajesh Kedia, Rijurekha Sen, M. Balakrishnan
ACM Trans. Embed. Comput. Syst.4
2022 EXPRESS: CNN EXecution Time PREdiction for DPU DeSign Space Exploration
abstract
Deep learning Processor Units (DPUs) from Xilinx are design-time configurable CNN accelerators for FPGAs. We propose EXPRESS, which predicts the execution time of any given CNN on a DPU. EXPRESS incorporates the effect of bus connections into prediction. As a DPU is invoked by a host CPU to process a CNN layer by layer, EXPRESS considers the CPU and the DPU execution time for predicting the end-to-end processing time. EXPRESS has an average prediction error of 2.2% and significantly outperforms state-of-the-art.
Shikha Goel, Rajesh Kedia, Rijurekha Sen, M. Balakrishnan
FPT4
2022 Beacon Placement and Signal Strength Estimation to Improve Localization Coverage and Accuracy
abstract
Estimating the location of a mobile device using Bluetooth signal strength observations is an active research area due to its promising application in location-based services. Field deployment of localization systems is facing many challenges. This includes optimal placement of sensors to improve the localization coverage and modeling the uncertainty in wireless signals given the complexity of signal propagation in the presence of heterogeneous factors like operating environment and obstacles. In this paper, we first present a heuristics-based technique for the placement of beacons assuming map data is available. Our beacon placement algorithm results in a significant improvement in configuration time and coverage while using a smaller number of beacons. Further, we propose a Gaussian-process-based novel hybrid kernel function to generate a robust likelihood model for signal strength estimation. Our experiments over 83 randomly sampled data points demonstrate a superior ranging and localization performance of our approach.
Vikas Upadhyay, Kashi Roy, M. Balakrishnan
IPIN3
2021 EnergyNN: Energy Estimation for Neural Network Inference Tasks on DPU
abstract
Convolutional Neural Networks (CNNs) are increasingly becoming popular in embedded and energy limited mobile applications. Hardware designers have proposed various accelerators to speed up the execution of CNNs on embedded platforms. Deep Learning Processor Unit (DPU) is one such generic CNN accelerator for Xilinx platforms that can execute any CNN on one or more DPUs configured on an FPGA. In a period of rapid growth in CNN algorithms and the availability of multiple configurations of CNN accelerators (like DPU), the design space is expanding fast. These design points show significant trade-off in execution time, energy consumption and application performance measured in terms of accuracy. To be able to perform this trade-off, we propose a methodology for energy estimation of a CNN running on a DPU. We build an energy model using characteristics of few CNNs and use this model for energy prediction of other unseen CNNs. We evaluate our approach using 16 different standard and popular CNNs with an average prediction error of 9.9%. Energy estimation can be useful in various scheduling applications where one can choose from multiple CNNs based on its energy consumption. We demonstrate the utility of our approach in a drone that is deployed for detecting objects on the ground.
Shikha Goel, M. Balakrishnan, Rijurekha Sen
FPL2
2021 VmAP: A Fair Metric for Video Object Detection
abstract
Video object detection is the task of detecting objects in a sequence of frames, typically, with a significant overlap in content among consecutive frames. Mean Average Precision (mAP) was originally proposed for evaluating object detection techniques in independent frames, but has been used for evaluating video based object detectors as well. This is undesirable since the average precision over all frames masks the biases that a certain object detector might have against certain types of objects depending on the number of frames for which the object is present in a video sequence. In this paper we show several disadvantages of mAP as a metric for evaluating video based object detection. Specifically, we show that: (a) some object detectors could be severely biased against some specific kind of objects, such as small, blurred, or low contrast objects, and such differences may not reflect in mAP based evaluation, (b) operating a video based object detector at the best frame based precision/recall value (high F1 score) may lead to many false positives without a significant increase in the number of objects detected. (c) mAP does not take into account that tracking can be potentially used to recover missed detections in the temporal neighborhood while this can be account for while evaluating detectors. As an alternate, we suggest a novel evaluation metric (VmAP) which takes the focus away from evaluating detections on every frame. Unlike mAP, VmAP rewards a high recall of different object views throughout the video. We form sets of bounding boxes having similar views of an object in a temporal neighborhood and use a set-level recall for evaluation. We show that VmAP is able to address all the challenges with the mAP listed above. Our experiments demonstrate hidden biases in object detectors, shows upto 99% reduction in false positives while maintaining similar object recall and shows a 9% improvement in correlation with post-tracking performance.
Anupam Sobti, Vaibhav Mavi, M. Balakrishnan, Chetan Arora 0001
ACM Multimedia3
2021 Performance-Energy Trade-off in Modern CMPs
abstract
Chip multiprocessors (CMPs) are ubiquitous in all computing systems ranging from high-end servers to mobile devices. In these systems, energy consumption is a critical design constraint as it constitutes the most significant operating cost for computing clouds. Analogous to this, longer battery life continues to be an essential user concern in mobile devices. To optimize on power consumption, modern processors are designed with Dynamic Voltage and Frequency Scaling (DVFS) support at the individual core as well as the uncore level. This allows fine-grained control of performance and energy. For an n core processor with m core and uncore frequency choices, the total DVFS configuration space is now m (n+1) (with the uncore accounting for the + 1). In addition to that, in CMPs, the performance-energy trade-off due to core/uncore frequency scaling concerning a single application cannot be determined independently as cores share critical resources like the last level cache (LLC) and the memory. Thus, unlike the uni-processor environment, the energy consumption of an application running on a CMP depends not only on its characteristics but also on those of its co-runners (applications running on other cores). The key objective of our work is to select a suitable core and uncore frequency that minimizes power consumption while limiting application performance degradation within certain pre-defined limits (can be termed as QoS requirements). The key contribution of our work is a learning-based model that is able to capture the interference due to shared cache, bus bandwidth, and memory bandwidth between applications running on multiple cores and predict near-optimal frequencies for core and uncore.
Solomon Abera, M. Balakrishnan
ACM Trans. Archit. Code Optim.2
2019 GRanDE: Graphical Representation and Design Space Exploration of Embedded Systems
abstract
Tasks executing computer vision and machine learning algorithms are becoming popular on embedded platforms. A key characteristic of such tasks is the presence of modes providing different levels of application performance in terms of metrics like accuracy. The system designer has the flexibility to select an appropriate mode for executing such tasks. Secondly, the designer also has the traditional flexibility of choosing suitable components to build the execution platform. Thirdly, the system performance might vary with various external factors (known as context), and during the initial stages of system design, the designer might have the flexibility to support only a subset of the possible contexts. This three-fold flexibility in the hands of the designer has not been explored simultaneously in prior works and raises the complexity of designing embedded systems many-fold. In this paper, we address the design of such systems through a novel framework named GRanDE (Graphical Representation and Design Space Exploration). GRanDE consists of a comprehensive graphical representation to capture the three aspects of the design space discussed earlier. Further, we transform this representation into Constraint Logic Programming (CLP) constructs, which could be used to interactively explore and prune the design space. We demonstrate the applicability of the proposed framework on an embedded system named MAVI having ~1.3 million design points. The generated CLP program could prune up to 99.74% of the design space of MAVI.
Rajesh Kedia, M. Balakrishnan, Kolin Paul
DSD2
2019 Multi-sensor Energy Efficient Obstacle Detection
abstract
With the improvement in technology, both the cost and the power requirement of cameras, as well as other sensors have come down significantly. It has allowed these sensors to be integrated into portable as well as wearable systems. Such systems are usually operated in a hands-free and always-on manner where they need to function continuously in a variety of scenarios. In such situations, relying on a single sensor or a fixed sensor combination can be detrimental to both performance as well as energy requirements. Consider the case of an obstacle detection task. Here using an RGB camera helps in recognizing the obstacle type but takes much more energy than an ultrasonic sensor. Infrared cameras can perform better than RGB camera at night but consume twice the energy. Therefore, an efficient system must use a combination of sensors, with an adaptive control that ensures the use of the sensors appropriate to the context. In this adaptation, one needs to consider both performance and energy and their trade-off. In this paper, we explore the strengths of different sensors as well their trade-off for developing a deep neural network based wearable device. We choose a specific case study in the context of a mobility assistance device for the visually impaired. The device detects obstacles in the path of a visually impaired person and is required to operate both at day and night with minimal energy to increase the usage time on a single charge. The device employs multiple sensors: ultrasonic sensor, RGB Camera, and NIR Camera along with a deep neural network accelerator for speeding up computation. We show that by adaptively choosing the appropriate sensor for the context, we can achieve up to 90% reduction in energy while maintaining comparable performance to a single sensor system.
Anupam Sobti, M. Balakrishnan, Chetan Arora 0001
DSD2
2019 Equivalence Checking and Compaction of n-input Majority Terms Using Implicants of Majority
Rajeswari Devadoss, Kolin Paul, M. Balakrishnan
J. Electron. Test.3
2018 Object Detection in Real-Time Systems: Going Beyond Precision
abstract
Applications like autonomous driving, industrial robotics, surveillance, and wearable assistive technology rely on object detectors as an integral part of the system. Thus, an increase in performance of object detectors directly affects the quality of such systems. In the recent years, convolutional neural networks (CNNs) and its variants emerged as the state of art in object detection, where performance is usually measured either in terms of mean average precision (mAP) or number of frames processed per second (fps). Many applications which use object detectors are resource constrained in practice. Even though it is clear from the published results, that a frame-level analysis of the system in terms of mAP or fps proves the superiority of one algorithm over the other, we observe that such metrics do not necessarily apply to real time applications with resource constraints. A slower algorithm even though highly accurate may need to drop frames to maintain the necessary frame rate and lose on the accuracy. We propose a closer look at the metrics used for performance in real-time applications, and suggest some new evaluation criterion. Our comparison of state of the art detectors on these metrics has also thrown some surprises in terms of conventional wisdom, which we present in this paper. Our framework is available at https://www.github.com/anupamsobti/object-detectionreal-time-systems.
Anupam Sobti, Chetan Arora 0001, M. Balakrishnan
WACV3
2014 LightSim: A leakage aware ultrafast temperature simulator
abstract
In this paper, we propose the design of an ultra-fast temperature simulator (LightSim) that can perform both steady state and transient thermal analysis, and also take the effect of leakage power into account. We use a novel Hankel transform based technique to derive a transient version of the Green's function for a chip, which takes into account the feedback loop between temperature and leakage. Subsequently, we calculate the temperature map of a chip by convolving the derived Green's function with the power map. Our simulator is at least 3500 times faster than HotSpot, and at least 2.3 times faster than competing research prototypes [4, 12]. The total error is limited to 0.18 °C.
Smruti R. Sarangi, Gayathri Ananthanarayanan, M. Balakrishnan
ASP-DAC3
2014 High Level Design Approach to Accelerate De Novo Genome Assembly Using FPGAs
abstract
Many scientific applications take a very long time to execute on general purpose processors. Speedups can be obtained by using specialized hardware in conjunction with the processors. FPGA based accelerators are known to be effective for reducing the execution time of many scientific applications. Since FPGAs are configurable, they can be customized to implement a variety of processing elements as accelerators. The process of mapping algorithm to architecture is complex, as the design space is large. System simulation is usually employed to carry out the exploration, in spite of the fact that simulation models take significantly large amount of time to execute. High level design space exploration helps in taking the required decisions to arrive at an optimal design. In this paper we describe design space exploration carried out for accelerating de novo genome assembly using FPGAs. Three models at various levels of abstraction were used. We discuss how the simulation time of these models influence the choice of design parameters at different levels of abstraction. We illustrate this process by using the high level models to evaluate Hard Embedded Blocks (HEBs) in FPGAs for accelerating the de novo genome assembly application.
B. Sharat Chandra Varma 0001, Kolin Paul, M. Balakrishnan
DSD3
2014 Mapping Tasks to a Dynamically Reconfigurable Coarse Grained Array
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) have become popular in recent times as the increased transistor densities have enabled greater integration of increasingly complex “compute cores”. These devices pack massive compute power and can be effectively used to build efficient solutions for applications which have a significant degree of parallelism. In many cases, these CGRAs are also partially reconfigurable. Clearly to make effective use of these highly “parallel compute platforms”, a good mapping flow is required to map the parallelism that is present in a target application.
Mansureh S. Moghaddam, Kolin Paul, M. Balakrishnan
FCCM3
2014 Edutactile - A Tool for Rapid Generation of Accurate Guideline-Compliant Tactile Graphics for Science and Mathematics
Mrinal Mech, Kunal Kwatra, Supriya Das, Piyush Chanana, Rohan Paul, M. Balakrishnan
ICCHP (2)6
2014 Super edge magic graceful graphs
G. Marimuthu, M. Balakrishnan
Inf. Sci.2
2013 A path-guided audio based indoor navigation system for persons with visual impairment
abstract
Independent path-based mobility in an unfamiliar indoor environment is a common problem faced by visually impaired community. We present the design of an infra-red based active wayfinding system for the visually impaired. Our proposed system: downloads the floor plan of the building, locates and tracks the user inside the building, finds the shortest path and provides step-by-step direction to the destination using voice messages. The audio instructions include active guidance for impending turns in the path of travel, distance of each section between turns, obstacle warning instructions and position correction messages when the user gets lost. Results from a needs finding study with visually impaired individuals formed the design of the system. We then deployed the system in a building and field tested it with users using a standardized before-and-after study. The comparison of the results demonstrated that the system is usable and useful.
Dhruv Jain, Akhil Jain, Rohan Paul, Akhila Komarika, M. Balakrishnan
ASSETS5
2013 FAssem: FPGA Based Acceleration of De Novo Genome Assembly
abstract
Next generation sequencing technologies produce large amounts of data at very low cost. They produce short reads of DNA fragments. These fragments have many overlaps, lots of repeats and may also include sequencing errors. The assembly process involves merging these sequences to form the original sequences. In recent years many software programs have been developed for this purpose. All of them take significant amount of time to execute. Velvet is a commonly used de novo assembly program. We propose a method to reduce the overall time for assembly by using pre-processing of the short read data on FPGAs and processing its output using Velvet. We show significant speed-ups with slight or no compromise on the quality of the assembled output.
B. Sharat Chandra Varma 0001, Kolin Paul, M. Balakrishnan, Dominique Lavenier
FCCM3
2012 E-super vertex magic labelings of graphs
G. Marimuthu, M. Balakrishnan
Discret. Appl. Math.2
2012 System-Level Design Space Exploration Methodology for Energy-Efficient Sensor Node Configurations: An Experimental Validation
abstract
For sensor nodes deployed at small distances, computation energy along with radio energy determines the battery life. We have proposed a system-level design space exploration methodology in [1] for searching an energy-efficient error-correcting code (ECC). This methodology takes into account the computation and the radio energy in an integrated manner. In this paper, we validate this methodology by deploying the Imote2 nodes and measuring energy values under different operating modes, e.g., with and without ECC. In this process, we propose a validation framework and node energy model. Experimental results validate the methodology and show that with ECC we can save up to 14% transmitter energy under a certain set of conditions. The main contribution of this paper is that it establishes experimentally that our methodology is effective in exploration of various node configurations and finding an energy-efficient solution.
Sonali Chouhan, M. Balakrishnan, Ranjan Bose
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2012 Measures and Countermeasures for Null Frequency Jamming of On-Demand Routing Protocols in Wireless Ad Hoc Networks
abstract
Distributed network protocols operate similar to periodic state machines, utilizing internal states and timers for network coordination, which creates opportunities for carefully engineered radio jamming to target the protocol operating periods and disrupt network communications. Such periodic attacks targeting specific protocol period/frequency of operation is referred to as Null Frequency Jamming (NFJ). Our hypothesis is that NFJ is a pervasive phenomenon in dynamic systems, including wireless ad-hoc networks. This paper aims to test the hypothesis by investigating NFJ targeted at the on-demand routing protocols for ad-hoc networks. Our mathematical analysis and simulation results show substantial degradation in end-to-end network throughput at certain null periods/frequencies, where the jamming periodicity self-synchronizes with the route-recovery cycle. We also study an effective countermeasure, randomized route-recovery periods, for eliminating the presence of predictable null frequencies and mitigating the impact of NFJ. Our analytical model and simulation results validate the effectiveness of randomized route recovery with appropriately chosen randomization ranges.
M. Balakrishnan, Hong Huang 0003, Rafael Asorey-Cacheda, Satyajayant Misra, Sandeep Pawar, Yousef Jaradat
IEEE Trans. Wirel. Commun.1
2011 Architecture and tools for programmable QCA
abstract
Quantum-dot Cellular Automata (QCA) is a nano-scale compute fabric being explored by the VLSI research community as the difficulties in shrinking CMOS transistors mount. The paradigm promises high device densities and power-efficiency, and has unique properties that make it an interesting candidate for programmable devices. In this work, we propose a specialized architecture for programmable devices using QCA, present design rules for circuit design using the architecture and introduce a simulation engine tuned to efficiently simulate QCA circuits designed for this architecture.
Rajeswari Devadoss, Kolin Paul, M. Balakrishnan
FPT3
2011 p-QCA: A Tiled Programmable Fabric Architecture Using Molecular Quantum-Dot Cellular Automata
abstract
Quantum-dot cellular automata is an interesting computation fabric with many never-seen-before properties. However, no programmable fabric scheme has utilized all these properties effectively. We propose an architecture for a programmable device using molecular QCA which exploits all the specialities of the fabric. The architecture taps the flexibility provided by the clocking system of molecular QCA to build a simple tile-based programmable device with the 3-input Majority gate as the fundamental logic element. Observing how a QCA structure can behave as either an interconnect or a logic gate depending on clocking, the proposed architecture merges routing and logic elements, thus drastically changing how programmable fabrics have been designed.
Rajeswari Devadoss, Kolin Paul, M. Balakrishnan
ACM J. Emerg. Technol. Comput. Syst.3
2011 Compressing Cache State for Postsilicon Processor Debug
abstract
During postsilicon processor debugging, we need to frequently capture and dump out the internal state of the processor. Since internal state constitutes all memory elements, the bulk of which is composed of cache, the problem is essentially that of transferring cache contents off-chip, to a logic analyser. In order to reduce the transfer time and save expensive logic analyser memory, we propose to compress the cache contents on their way out. We present a hardware compression engine for cache data using a Cache-Aware Compression strategy that exploits knowledge of the cache fields and their behavior to achieve an effective compression. Experimental results indicate that the technique results in 7-31 percent better compression than one that treats the data as just one long bit stream. We also describe and evaluate a parallel compression architecture that uses multiple compression engines, resulting in a 54 percent reduction in transfer time.
Preeti Ranjan Panda, M. Balakrishnan, Anant Vishnoi
IEEE Trans. Computers2
2010 A tiled programmable fabric using QCA
abstract
Quantum-dot Cellular Automata is an interesting computation fabric with many never-seen-before properties. However, no programmable fabric scheme has utilized all these properties effectively. We propose an architecture for a programmable device using QCA which exploits all the specialities of the fabric. The architecture taps the flexibility provided by the clocking system of QCA to build a simple tile based programmable device with the 3-input Majority gate as the fundamental logic element. Observing how a QCA structure can behave as either an interconnect or a logic gate depending on clocking, the proposed architecture merges routing and logic elements, thus drastically changing how programmable fabrics have been designed.
Rajeswari Devadoss, Kolin Paul, M. Balakrishnan
FPT3
2010 Enhancing post-silicon processor debug with Incremental Cache state Dumping
abstract
During post-silicon validation/debug of processors, it is common to alternate between two phases: processor execution and state dump. The state dump, where the entire processor state is dumped off-chip to a logic analyzer for further processing, is a major bottleneck. We present a technique for improving debug efficiency by reducing the volume of cache data dumped off-chip, while still capturing the complete state. The reduction is achieved by introducing hardware mechanisms to transmit only the portion of the cache that was updated since the last dump. We propose two design alternatives based on whether or not the processor is permitted to continue execution during the dump: Blocking Incremental Cache Dumping (BICD) and Non-blocking Incremental Cache Dumping (NICD). We observe a 64% reduction in overall cache lines dumped and the dump time reduces to an average of 16.8% and 0.0002% for BICD and NICD respectively.
Preeti Ranjan Panda, Anant Vishnoi, M. Balakrishnan
VLSI-SoC3
2009 Online cache state dumping for processor debug
abstract
Post-silicon processor debugging is frequently carried out in a loop consisting of several iterations of the following two key steps: (i) processor execution for some duration, followed by (ii) dumping out of the processor's internal state into an external logic analyzer for further offline processing. Internal state of the processor is dominated by the L2 cache. During the process of dumping the cache content, the processor's execution is halted so that the state can be faithfully reproduced offline. In order to reduce the duration for which the processor is halted, and indirectly reduce debug time, we propose two Online Cache Dumping strategies, Retransmit Non-dumped Line (RNL) and Dump History Table (DHT), with the objective of transferring the cache contents while the processor is executing, and yet maintaining fidelity of the dumped data. For typical experimental debug scenarios, we observe that the effective dump times are reduced to between 0.01% and 3.5% of the original times. We also employ compression to reduce the cache content transfer time and logic analyzer space. Our experiments indicate an average compression ratio of 59.2%.
Anant Vishnoi, Preeti Ranjan Panda, M. Balakrishnan
DAC3
2009 A generic platform for estimation of multi-threaded program performance on heterogeneous multiprocessors
abstract
This paper deals with a methodology for software estimation to enable design space exploration of heterogeneous multiprocessor systems. Starting from fork-join representation of application specification along with high level description of multiprocessor target architecture and mapping of application components onto architecture resource elements, it estimates the performance of application on target multiprocessor architecture. The methodology proposed includes the effect of basic compiler optimizations, integrates light weight memory simulation and instruction mapping for complex instruction to improve the accuracy of software estimation. To estimate performance degradation due to contention for shared resources like memory and bus, synthetic access traces coupled with interval analysis technique is employed. The methodology has been validated on a real heterogeneous platform. Results show that using estimation it is possible to predict performance with average errors of around 11%.
Aryabartta Sahu, M. Balakrishnan, Preeti Ranjan Panda
DATE2
2009 Cache aware compression for processor debug support
abstract
During post-silicon processor debugging, we need to frequently capture and dump out the internal state of the processor. Since internal state constitutes all memory elements, the bulk of which is composed of cache, the problem is essentially that of transferring cache contents off-chip, to a logic analyzer. In order to reduce the transfer time and save expensive logic analyzer memory, we propose to compress the cache contents on their way out. We present a hardware compression engine for cache data using a Cache Aware Compression strategy that exploits knowledge of the cache fields and their behavior to achieve an effective compression. Experimental results indicate that the technique results in 7-31% better compression than one that treats the data as just one long bit stream. We also describe and evaluate a parallel compression architecture that uses multiple compression engines, resulting in a 54% reduction in transfer time.
Anant Vishnoi, Preeti Ranjan Panda, M. Balakrishnan
DATE3
2009 An experimental validation of system level design space exploration methodology for energy efficient sensor nodes
abstract
For closely deployed sensor nodes, computation energy along with radio energy determines the battery life. We have proposed a system level design space exploration methodology [1] for searching an energy efficient error correcting code (ECC). This methodology takes into account the computation and the radio energy in an integrated manner. In this paper we validate this methodology by deploying Imote2 nodes and measuring energy values under different operating modes, e.g., with and without ECC. Experimental results validate the methodology and show that with ECC we can save upto 13% transmitter energy for certain set of conditions. This paper validates experimentally that our methodology is effective in exploration of various node configurations and finding an energy efficient solution.
Sonali Chouhan, M. Balakrishnan, Ranjan Bose
ISLPED2
2009 A Framework for Energy-Consumption-Based Design Space Exploration for Wireless Sensor Nodes
abstract
In this paper, we first establish that, in wireless sensor networks, operating over ldquosmallrdquo distances, both computation energy and radio energy influence the battery life. In such a scenario, to evaluate the utility of error-correcting codes (ECCs) from an energy perspective, one has to consider the energy consumed in encoding-decoding and transmitting additional ldquoredundantrdquo bits vis-a-vis the energy saved due to coding gain. This paper presents a framework for evaluating various ECCs based on a comprehensive energy model of a sensor node. The framework supports exploration of sensor node design space with application- and deployment-related parameters, like distance, bit error rate, path loss exponent, as well as the modulation scheme and ECC parameters. The exploration results show that, as compared to the uncoded-data transmission, the energy-optimal ECC saves 15%-60% node energy for the given parameters.
Sonali Chouhan, Ranjan Bose, M. Balakrishnan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Integrated energy analysis of error correcting codes and modulation for energy efficient wireless sensor nodes
abstract
Optimizing energy consumption is a key objective in designing wireless sensor nodes. It has been shown earlier that the node energy is strongly influenced by the modulation and the error correcting code (ECC) used. The utility of using ECC from an energy perspective is determined by the energy saving due to the ECC coding gain vis-a-vis the energy overhead of "redundant" bits and of energy saving. Furthermore, the node energy varies with the change in errorcorrecting capability and code word length of a particular ECC as well as the modulation constellation size. The ECC coding gain is influenced by the constellation size. In this paper, the node energy variations with ECC and modulation parameters are analyzed for an energy optimal node design for the nodes operating in the additive white Gaussian noise channel. Based on this analysis, we compute the per information bit node energy and this is used to select an "optimal" ECC and modulation scheme pair. Our results show that the energy optimal ECC-modulation pair selected for some specific operating conditions could save as much as 50% energy. In nutshell, our work is targeted towards reducing the search space and finding an energy optimal ECC modulation pair for the given environment and application.
Sonali Chouhan, Ranjan Bose, M. Balakrishnan
IEEE Trans. Wirel. Commun.3
2008 A framework for energy consumption based design space exploration for wireless sensor nodes
abstract
In wireless sensor networks due to the small transmission distances involved, the computation energy along with the radio energy determines the battery life. Energy consumption of error control codes (ECCs) is a complex function of the energy consumption in computing encoding-decoding, transmitting "redundant" bits and energy saved by coding gain. This paper presents a methodology, which integrates computation and radio energy for searching an energy optimal ECC. Based on this methodology, a design space exploration framework and the energy model of sensor node have been developed. Exploration results show that the energy optimal ECC saves 15-58% node energy for given parameters.
Sonali Chouhan, M. Balakrishnan, Ranjan Bose
ISLPED2
2007 A Behavioral Synthesis Approach for Distributed Memory FPGA Architectures
abstract
This paper presents an approach for efficiently mapping loops and array intensive applications onto FPGA architectures with distributed RAMs, multipliers and logic. We perform a data dependency based, two level partitioning of the application's iteration space under target FPGA architectural constraints, to achieve better performance. It is shown that, this approach can result in a super-linear speedup; linear speedup due to concurrent computation on multiple compute elements and additional speedup due to improvement in the clock frequency (up to 30%). The clock period reduction is made possible because computation and accesses are now localized, i.e. the compute elements interact only with memories which are close by.
Ashutosh Pal, M. Balakrishnan
FPL2
2007 Impact of intercluster communication mechanisms on ILP in clustered VLIW architectures
abstract
VLIW processors have started gaining acceptance in the embedded systems domain. However, monolithic register file VLIW processors with a large number of functional units are not viable. This is because of the need for a large number of ports to support FU requirements, which makes them expensive and extremely slow. A simple solution is to break the register file into a number of smaller register files with a subset of FUs connected to it. These architectures are termed clustered VLIW processors.
Anup Gangwar, M. Balakrishnan
ACM Trans. Design Autom. Electr. Syst.2
2006 New approach to architectural synthesis: incorporating QoS constraint
abstract
Embedded applications like video decoding, video streaming and those in the network domain, typically have a Quality of Service (QoS) requirement which needs to be met. Apart from being a design constraint, it can also be considered as a flexibility that the design does not have to work under worst case data condition. For example, in the case of video decoding with variable decoding time for individual frames, it may be adequate that only a fraction of frames (say 90%) needs to be decoded. In this work, we propose a novel method of exploiting this flexibility for efficient partitioning and mapping in the architectural synthesis of the application at hand. We translate QoS specification of the overall application to the time constraint on the individual components constituting the application and use this knowledge in an optimum synthesis technique based on Mixed Integer Linear Programming (MILP). We study this in the context of MPEG2 Decoder and show that the approach can be used to obtain optimal time of execution as well as energy reduction while meeting the QoS requirements.
Harsh Dhand, Basant Kumar Dwivedi, M. Balakrishnan
EMSOFT3
2005 Evaluation of Bus Based Interconnect Mechanisms in Clustered VLIW Architectures
abstract
With new sophisticated compiler technology, it is possible to schedule distant instructions efficiently. As a consequence, the amount of exploitable instruction level parallelism (ILP) in applications has gone up considerably. However, monolithic register file VLIW architectures present scalability problems due to a centralized register file which is far slower than the functional units (FU). Clustered VLIW architectures, with a subset of FU connected to any RF are the solution to this scalability problem. Recent studies with a wide variety of inter-cluster interconnection mechanisms have presented substantial gains in performance (number of cycles) over the most studied RF-to-RF type interconnections. However, these studies have compared only one or two design points in the RF-to-RF interconnects design space. In this paper, we extend the previous reported work. We consider both multi-cycle and pipelined buses. To obtain realistic bit latencies, we synthesized the various architectures and found out post layout clock periods. The results demonstrate that while there is very little variation in interconnect area, all the bus based architectures are heavily performance constrained. Also, neither multi-cycle nor pipelined buses or increasing the number of buses itself is able to achieve performance comparable to point-to-point type interconnects.
Anup Gangwar, M. Balakrishnan, Preeti Ranjan Panda
DATE2
2005 SMPS: an FPGA-based prototyping environment for multiprocessor embedded systems (abstract only)
abstract
Streaming media applications represent an important class of applications for embedded systems. Recent advances in design-space exploration of architectures for such applications have pointed towards the suitability of Multiprocessor System on Chip (SoC) solutions. Multiprocessor SoCs not only offer higher performance, but can also lead to solutions which are cheaper cost wise. A typical synthesis methodology for such architectures would require a validation stage at the end of final system integration. The wide availability of cheap and large FPGA devices, advances in automatic synthesis from VHDL/Verilog and abundance of high performance computing platforms enables the design of a generic validation system for such Multiprocessor SoCs.In this paper we present the design and implementation of Srijan Multiprocessor Prototyping System (SMPS). SMPS is a system for rapid prototyping and validation of single chip application specific multiprocessor systems. The individual computing elements are RISC processors, coprocessors which lie in the processor pipeline, and ASICs which connect directly to system bus. The system is a tightly coupled multiprocessor with shared memory and shared address space. A Real-time Operating System (RTOS) provides task scheduling and access to shared resources. The system is presented as a parameterized VHDL based on the open source Sparc~V8 compliant LEON processor and a homegrown light-weight RTOS, RtKer-MP. The entire VHDL is configurable using a GUI, has support for cache coherency, choice of arbitration policy and easy integration of custom processing engines. RtKer-MP allows for a pluggable scheduler, dynamic and static scheduling policies, static and dynamic task migrations domains and variable interruption frequencies for separate processors. The pluggable scheduler interface allows for quick exploration of various scheduling policies for a feedback to the estimation systems.
Ankit Mathur, Mayank Agarwal, Soumyadeb Mitra, Anup Gangwar, M. Balakrishnan, Subhashis Banerjee
FPGA5
2004 An efficient technique for exploring register file size in ASIP design
abstract
Performance estimation is a crucial operation which drives the design space exploration in an application-specific instruction set processor synthesis. With the increase in the level of integration, the design space has considerably expanded, which makes the simulation-driven techniques inadequate due to their slow speed. Alternatively, there are approaches which estimate performance by scheduling the application on the available processor resources without generating code. These are much more effective in exploring a large design space. The technique reported in this paper presents a scheduler-based approach for register file-size exploration. The performance is estimated by computing the number of spills for a particular register file size. The concept of register reuse chains is used for local register allocation, while live variable analysis done across the blocks is used to estimate global register needs. The technique is fast, accurate, retargetable, and does not require code generation. We have generated and validated execution-time estimates for selected benchmarks for a RISC (ARM7TDMI) and a VLIW (TM-1000) processor. Further, we have shown that this approach is much faster than simulator-based techniques.
Manoj Kumar Jain, M. Balakrishnan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Exploring Storage Organization in ASIP Synthesis
abstract
Performance estimation which drives the design space exploration is usually done by simulation. With increasing dimensions of the design space, simulator based approaches become too time consuming. In the domain of application specific instruction set processors (ASIP), this problem can be solved by scheduler based approaches, which are much faster. However, existing scheduler based approaches do not help in exploring storage organization. We present a scheduler based technique for exploring register file size, number of register windows and cache configurations in an integrated manner. Performances for different register file sizes are estimated by predicting the number of memory spills and its delay. The technique employed does not require explicit register assignment. The number of context switches leading to spills is estimated for evaluating the time penalty due to a limited number of register windows and cache simulator is used for estimating cache performance. The proposed technique has been validated for several benchmarks over a range of processors by comparing our estimates to the results obtained from standard simulation tools. The processors include ARM7TDMI, LEON and Trimedia (TM-1000).
Manoj Kumar Jain, M. Balakrishnan
DSD2
2002 An efficient technique for exploring register file size in ASIP synthesis
abstract
Performance estimation is a crucial operation which drives the design space exploration in Application Speci c Instruction Set Processors (ASIP) synthesis. The usual approach to estimate performance is to do simulation. With increasing dimensions of the design space, simulator based approaches become too time consuming. This problem can be solved by scheduler based approaches, which are much faster. However existing scheduler based approaches do not help in exploring storage organization. This paper presents a scheduler based technique for exploring register le size in ASIP synthesis.
Manoj Kumar Jain, M. Balakrishnan
CASES2
2001 Analysis of the influence of register file size on energyconsumption, code size, and execution time
abstract
Interest in low-power embedded systems has increased considerably in the past few years. To produce low-power code and to allow an estimation of power consumption of software running on embedded systems, a power model was developed based on physical measurement using an evaluation board and integrated into a compiler and profiler. The compiler uses the power information to choose instruction sequences consuming less power, whereas the profiler gives information about the total power consumed during execution of the generated program. The used compiler is parameterized such that, e.g., the register file size may be changed. The resulting code is evaluated with respect to code size, performance, and power consumption for different register file sizes. The extracted information is especially useful during application analysis and architecture space exploration in application-specific integrated processor (ASIP) design. Our analysis gives the designer the ability to estimate the desirable register file size for an ASIP design. The size of the register file should be considered as a design parameter since it has a strong impact on the energy consumption of embedded systems.
Lars Wehmeyer, Manoj Kumar Jain, Stefan Steinke, Peter Marwedel, M. Balakrishnan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2000 Speeding up power estimation of embedded software
abstract
Power is increasingly becoming a design constraint for embedded systems. A processor is responsible for energy consumption on account of the software component of the embedded system. The power estimation of this component is a major concern due to the rising complexities of processors and the slow estimation tools. This work attempts to estimate the energy dissipation of the PR1900 processor based on instruction set model with improved accuracy. The model is integrated in a simulation framework and validated. Over 200 times speedup has been obtained with average 1.4% loss in accuracy over gate level estimation. Analysis of the energy dissipated by the instruction vis a vis the processor architecture has been carried out and a substantial reduction in the measurement effort to build the processor energy model has been achieved.
Akshaye Sama, J. F. M. Theeuwen, M. Balakrishnan
ISLPED3
2000 Allocation of FIFO structures in RTL data paths
abstract
Along with functional units, storage and interconnects contribute significantly to data path costs. This paper addresses the issue of reducing the costs of storage and interconnect. In a post-datapath synthesis phase, one or more queues can be allocated and variables bound to it, with the goal of reducing storage and interconnect costs. Further, in contrast to earlier work, we support “irregular” cdfgs and multicycle functional units for queue synthesis. Initial results on HLS benchmark examples have been encouraging, and show the potential of using queue synthesis to reduce datapath cost. A novel feature of our work is the formulation of the problem for a variety of FIFO structures with their own “queueing” criteria.
M. Balakrishnan, Heman Khanna
ACM Trans. Design Autom. Electr. Syst.1
1999 Hardware/Software Partitioning Between Microprocessor and Reconfigurable Hardware
abstract
No abstract available.
Sanjiv Kapoor, M. Balakrishnan
FPGA3
1998 Direct mapping of RTL structures onto LUT-based FPGA's
abstract
The problem of mapping synthesized RTL structures onto look-up table (LUT)-based field programmable gate arrays (FPGAs) is addressed in this paper. The key distinctive feature of this work is a novel approach to perform the mapping by utilizing the iterative nature of the data path components. The approach exploits the regularity of data path components by slicing the components and mapping slices of one or more connected components together. This is in contrast to other FPGA mapping techniques which start from Boolean networks. Both cost optimal and delay optimal mappings are supported. The objective in cost optimal mapping is to cover a given data path network with minimum number of CLBs. Similarly in delay optimal mapping, the objective is to reduce the number of CLB levels in the critical combinational logic paths. Implementation of these mapping techniques with LUT based FPGAs as target technology results in a significant reduction in cost (CLB count) and critical path delays (CLB levels).
A. R. Naseer, M. Balakrishnan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1995 Buffer constraints in a variable-rate packetized video system
abstract
This paper provides an analysis of the buffers in a variable rate video encoding system and provides necessary conditions that need to be met by the channel and the encoder buffer control mechanism in order to ensure that the decoder and encoder buffers do not underflow or overflow. It is also shown that this can be achieved through the monitoring of just the encoder buffer. The paper also provides an algorithm that allows encoders to change their rate, by dynamically changing their logical buffer sizes, and lists the conditions that need to be met prior to the change in channel rate.
M. Balakrishnan
ICIP1
1989 Integrated Scheduling and Binding: A Synthesis Approach for Design Space Exploration
abstract
Synthesis of digital systems, involves a number of tasks ranging from scheduling to generating interconnections. The interrelationship between these tasks implies that good designs can only be generated by considering the overall impact of a design decision. The approach presented in this paper provides a framework for integrating scheduling decisions with binding decisions. The methodology supports allocation of a wider mix of operator modules and covers the design space more effectively. The process itself can be described as incremental synthesis and is thus well-suited for applications involving partial pre-synthesized structures.
M. Balakrishnan, Peter Marwedel
DAC1
1988 A Semantic Approach for Modular Synthesis of VLSI Systems
M. Balakrishnan, S. Sutarwala, Arun K. Majumdar, Dilip K. Banerji, James G. Linders
Inf. Process. Lett.1
1988 Synthesis of decentralised controllers from high level description
abstract
The present approaches to automated digital synthesis realise the control using a centralised controller. This paper presents a technique for synthesizing decentralised controllers based on interconnections between control and data path. The algorithm is based on analyzing the control flow and suitably partitioning it to realise decentralised controllers. The controllers, thus synthesized, preserve the control structure in the source description and have simple inter-controller communication. The algorithm is illustrated with an example and the controllers are mapped to pla like structures.
M. Balakrishnan, Arun K. Majumdar, Dilip K. Banerji, James G. Linders
Microprocess. Microprogramming1
1988 Allocation of multiport memories in data path synthesis
abstract
An algorithm to synthesize registers using multiport memories during data-path synthesis is presented. The proposed approach considers not only the access requirements of registers but also their interconnection to operators in order to minimize required interconnections. The same approach can be applied to select the optimum number of buses in a multibus architecture. The method is illustrated with an example.>
M. Balakrishnan, Arun K. Majumdar, Dilip K. Banerji, James G. Linders, Jayanti C. Majithia
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1987 An efficient retargetable microprogram generating system
M. Balakrishnan, Pramod Chandra P. Bhatt, Bharat B. Madan
Microprocessing and Microprogramming1
1986 A survey of microprogramming languages
M. Balakrishnan, Bharat B. Madan, Pramod Chandra P. Bhatt
Microprocessing and Microprogramming1
1982 A multi-channel microprogrammed FFT processor
abstract
The complexity of implementing a multi-channel real-time FFT processor is examined with emphasis on memory specifications and processor performance, Design alternatives are considered and finally the design of an optimal processor is presented. The designed system is microprogrammed and optimised for pipeline processing. A reduction of a factor of four in the mass storage speed has been achieved at the expense of a small size high-speed memory.
M. Balakrishnan, A. V. S. M. Rao, Rajendar Bahl
ICASSP1