EDBT 2026 Demo / reviewers in the wild / expert
Bharadwaj S. Amrutur
dblp:66/4689 · also Bharadwaj Amrutur
· DBLP profile ↗
45ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0007-5062-6174ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gradient-Weighted Feature Back-Projection: A Fast Alternative to Feature Distillation in 3D Gaussian SplattingabstractWe propose a training-free method for feature field rendering in 3D Gaussian Splatting, enabling fast and scalable embedding of high-dimensional features into 3D scenes. Unlike training-based feature distillation methods, which are computationally expensive and often yield feature embeddings that poorly reflect the rendered semantics, our approach back-projects 2D features onto pre-trained 3D Gaussians using influence weights derived from the rendering equation. This projection produces a queryable 3D feature field, validated on tasks including 2D and 3D segmentation, affordance transfer, and identity encoding, spanning queries using language, pixel, and synthetic embeddings. These capabilities, in turn, enable downstream applications in augmented and virtual reality, interactive scene editing, and robotics. Across different tasks, our method achieves performance comparable to or better than training-based approaches, while significantly reducing computational cost. The project page is at https://jojijoseph.github.io/3dgs-backprojection. Joji Joseph, Bharadwaj S. Amrutur, Shalabh Bhatnagar |
SIGGRAPH Asia | 2 |
| 2023 | A Co-Simulation Framework for Communication and Control in Autonomous Multi-Robot SystemsabstractMulti-Robot Systems (MRS) are transforming diverse domains like logistics, cargo management, and agriculture. However, ensuring that the behavior is correct under various network conditions within such complex environments is challenging, and meeting the desired automation goals is difficult. We propose the CORNET 2.0 co-simulation framework to jointly and accurately simulate multi-agent robotic systems within physical environments and the communication network models within such environments. Our modular framework allows diverse robot and network models to seamlessly integrate to simulate the robot's autonomy, physical space, and network features, such as latency, throughput, and loss intrinsic to the network topology and communication technology. A key novelty of CORNET 2.0 is its accurate synchronizing of mobility and time, which ensures that the physical location of a robot at a point in time, and the network properties and packets that flow from that location, are aligned. This is vital to model and validate MRS coordination algorithms that rely on network interactions. We provide a detailed evaluation of CORNET 2.0 in modeling real-world MRS use cases, such as leader-follower and warehouse environments, that help highlights the benefits. Srikrishna Acharya, Mukunda Bharatheesha, Yogesh L. Simmhan, Bharadwaj S. Amrutur |
IROS | 4 |
| 2021 | Spatial Reasoning from Natural Language Instructions for Robot ManipulationabstractRobots that can manipulate objects in unstructured environments and collaborate with humans can benefit immensely by understanding natural language. We propose a pipelined architecture of two stages to perform spatial reasoning on the text input. All the objects in the scene are first localized, and then the instruction for the robot in natural language and the localized co-ordinates are mapped to the start and end co-ordinates corresponding to the locations where the robot must pick up and place the object respectively. We show that representing the localized objects by quantizing their positions to a binary grid is preferable to representing them as a list of 2D co-ordinates. We also show that attention improves generalization and can overcome biases in the dataset. The proposed method is used to pick-and-place playing cards using a robot arm. Sagar Venkatesh Gubbi, Anirban Biswas, Raviteja Upadrashta, Vikram Srinivasan, Partha Talukdar, Bharadwaj S. Amrutur |
ICRA | 6 |
| 2021 | Translating Natural Language Instructions to Computer Programs for Robot ManipulationabstractIt is highly desirable for robots that work alongside humans to be able to understand instructions in natural language. Existing language conditioned imitation learning models directly predict the actuator commands from the image observation and the instruction text. Rather than directly predicting actuator commands, we propose translating the natural language instruction to a Python function which queries the scene by accessing the output of the object detector and controls the robot to perform the specified task. This enables the use of non-differentiable modules such as a constraint solver when computing commands to the robot. Moreover, the labels in this setup are significantly more informative computer programs that capture the intent of the expert rather than teleoperated demonstrations. We show that the proposed method performs better than training a neural network to directly predict the robot actions. Sagar Venkatesh Gubbi, Raviteja Upadrashta, Bharadwaj S. Amrutur |
IROS | 3 |
| 2021 | Exploiting Application Tolerance for Functional SafetyabstractAs the use of safety critical systems is becoming more prevalent, there is a need to reduce the implementation overhead required to provide safety. The conventional design of such systems does not consider application behaviours, thereby resulting in a pessimistic design where the safety provided is often not required during large periods of the application execution. In this paper, we analyse the different phases of an application during its overall execution life cycle, together with the embedded threads to perform specific operations, and propose a new methodology for protection of the safety critical application threads. We show the benefits of this method and the ability to build lower cost systems which are functionally safe using the flexibility which is embedded inside the application itself. Two new application based protection schemes, based on altering the application execution parameters (e.g. control loop frequency) and redundant execution of selective threads, are proposed. For these experiments, we have used commercial off the shelf components without any hardware functional safety features and implemented safety measures by augmenting the application software. Experiments on Electric Vehicle Traction (EVT) and On-Board Charger (OBC) applications indicate overall MIPS savings between 70% to 95%. These results indicate that a careful design of the application can itself be the first step to protect the integrated circuits which drive them. Prasanth Viswanathan Pillai, Rubin A. Parekhji, Bharadwaj S. Amrutur |
ITC | 3 |
| 2020 | Multi-Instance Aware Localization for End-to-End Imitation LearningabstractExisting architectures for imitation learning using image-to-action policy networks perform poorly when presented with an input image containing multiple instances of the object of interest, especially when the number of expert demonstrations available for training are limited. We show that end-to-end policy networks can be trained in a sample efficient manner by (a) appending the feature map output of the vision layers with an embedding that can indicate instance preference or take advantage of an implicit preference present in the expert demonstrations, and (b) employing an autoregressive action generator network for the control layers. The proposed architecture for localization has improved accuracy and sample efficiency and can generalize to the presence of more instances of objects than seen during training. When used for end-to-end imitation learning to perform reach, push, and pick-and-place tasks on a real robot, training is achieved with as few as 15 expert demonstrations. Sagar Venkatesh Gubbi, Raviteja Upadrashta, Shishir Kolathaya, Bharadwaj S. Amrutur |
IROS | 4 |
| 2020 | Teaching Robots Novel Objects by Pointing at ThemabstractRobots that must operate in novel environments and collaborate with humans must be capable of acquiring new knowledge from human experts during operation. We propose teaching a robot novel objects it has not encountered before by pointing a hand at the new object of interest. An end-to-end neural network is used to attend to the novel object of interest indicated by the pointing hand and then to localize the object in new scenes. In order to attend to the novel object indicated by the pointing hand, we propose a spatial attention modulation mechanism that learns to focus on the highlighted object while ignoring the other objects in the scene. We show that a robot arm can manipulate novel objects that are highlighted by pointing a hand at them. We also evaluate the performance of the proposed architecture on a synthetic dataset constructed using emojis and on a real-world dataset of common objects. Sagar Venkatesh Gubbi, Raviteja Upadrashta, Shishir Kolathaya, Bharadwaj S. Amrutur |
RO-MAN | 4 |
| 2020 | Learning Stable Manoeuvres in Quadruped Robots from Expert DemonstrationsabstractWith the research into development of quadruped robots picking up pace, learning based techniques are being explored for developing locomotion controllers for such robots. A key problem is to generate leg trajectories for continuously varying target linear and angular velocities, in a stable manner. In this paper, we propose a two pronged approach to address this problem. First, multiple simpler policies are trained to generate trajectories for a discrete set of target velocities and turning radius. These policies are then augmented using a higher level neural network for handling the transition between the learned trajectories. Specifically, we develop a neural network based filter that takes in target velocity, radius and transforms them into new commands that enable smooth transitions to the new trajectory. This transformation is achieved by learning from expert demonstrations. An application of this is the transformation of a novice user's input into an expert user's input, thereby ensuring stable manoeuvres regardless of the user's experience. Training our proposed architecture requires much less expert demonstrations compared to standard neural network architectures. Finally, we demonstrate experimentally these results in the in-house quadruped Stoch 2. Sashank Tirumala, Sagar Venkatesh Gubbi, Kartik Paigwar, Aditya Sagi, Ashish Joglekar, Shalabh Bhatnagar, Ashitava Ghosal, Bharadwaj S. Amrutur, Shishir Kolathaya |
RO-MAN | 8 |
| 2019 | Realizing Learned Quadruped Locomotion Behaviors through Kinematic Motion PrimitivesabstractHumans and animals are believed to use a very minimal set of trajectories to perform a wide variety of tasks including walking. Our main objective in this paper is two fold 1) Obtain an effective tool to realize these basic motion patterns for quadrupedal walking, called the kinematic motion primitives (kMPs), via trajectories learned from deep reinforcement learning (D-RL) and 2) Realize a set of behaviors, namely trot, walk, gallop and bound from these kinematic motion primitives in our custom four legged robot, called the “Stoch”. D-RL is a data driven approach, which has been shown to be very effective for realizing all kinds of robust locomotion behaviors, both in simulation and in experiment. On the other hand, kMPs are known to capture the underlying structure of walking and yield a set of derived behaviors. We first generate walking gaits from D-RL, which uses policy gradient based approaches. We then analyze the resulting walking by using principal component analysis. We observe that the kMPs extracted from PCA followed a similar pattern irrespective of the type of gaits generated. Leveraging on this underlying structure, we then realize walking in Stoch by a straightforward reconstruction of joint trajectories from kMPs. This type of methodology improves the transferability of these gaits to real hardware, lowers the computational overhead on-board, and also avoids multiple training iterations by generating a set of derived behaviors from a single learned gait. Abhik Singla, Shounak Bhattacharya, Dhaivat Dholakiya, Shalabh Bhatnagar, Ashitava Ghosal, Bharadwaj S. Amrutur, Shishir Kolathaya |
ICRA | 6 |
| 2019 | One-Shot Object Localization Using Learnt Visual Cues via Siamese NetworksabstractA robot that can operate in novel and unstructured environments must be capable of recognizing new, previously unseen, objects. In this work, a visual cue is used to specify a novel object of interest which must be localized in new environments. An end-to-end neural network equipped with a Siamese network is used to learn the cue, infer the object of interest, and then to localize it in new environments. We show that a simulated robot can pick-and-place novel objects pointed to by a laser pointer. We also evaluate the performance of the proposed approach on a dataset derived from the Omniglot handwritten character dataset and on a small dataset of toys. Sagar Venkatesh Gubbi, Bharadwaj S. Amrutur |
IROS | 2 |
| 2019 | Learning Active Spine Behaviors for Dynamic and Efficient Locomotion in Quadruped RobotsabstractIn this work, we provide a simulation framework to perform systematic studies on the effects of spinal joint compliance and actuation on bounding performance of a 16-DOF quadruped spined robot Stoch 2. Fast quadrupedal locomotion with active spine is an extremely hard problem, and involves a complex coordination between the various degrees of freedom. Therefore, past attempts at addressing this problem have not seen much success. Deep-Reinforcement Learning seems to be a promising approach, after its recent success in a variety of robot platforms, and the goal of this paper is to use this approach to realize the aforementioned behaviors. With this learning framework, the robot reached a bounding speed of 2.1m /s with a maximum Froude number of 2. Simulation results also show that use of active spine, indeed, increased the stride length, improved the cost of transport, and also reduced the natural frequency to more realistic values. Shounak Bhattacharya, Abhik Singla, Abhimanyu, Dhaivat Dholakiya, Shalabh Bhatnagar, Bharadwaj S. Amrutur, Ashitava Ghosal, Shishir Kolathaya |
RO-MAN | 6 |
| 2019 | Trajectory based Deep Policy Search for Quadrupedal WalkingabstractIn this paper, we explore a specific form of deep reinforcement learning (D-RL) technique for quadrupedal walking—trajectory based policy search via deep policy networks. Existing approaches determine optimal policies for each time step, whereas we propose to determine an optimal policy for each walking step. We justify our approach based on the fact that animals including humans use “low” dimensional trajectories at the joint level to realize walking. We will construct these trajectories by using Bézier polynomials, with the coefficients being determined by a parameterized policy. In order to maintain smoothness of the trajectories during step transitions, hybrid invariance conditions are also applied. The action is computed at the beginning of every step, and a linear PD control law is applied to track at the individual joints. After each step, reward is computed, which is then used to update the new policy parameters for the next step. After learning an optimal policy, i.e., an optimal walking gait for each step, we then successfully play them in a custom built quadruped robot, Stoch 2, thereby validating our approach. Shishir Kolathaya, Ashitava Ghosal, Bharadwaj S. Amrutur, Ashish Joglekar, Suhan Shetty, Dhaivat Dholakiya, Abhimanyu, Aditya Sagi, Shounak Bhattacharya, Abhik Singla, Shalabh Bhatnagar |
RO-MAN | 3 |
| 2017 | Safety analysis for integrated circuits in the context of hybrid systemsabstractMany real-life systems have integrated circuits interacting with physical systems in safety critical applications. These systems are called hybrid systems. The safety analysis of integrated circuits used in such systems is typically done in isolation of the end application and associated physical system, and hence results in the need to take recourse to conservative design techniques utilizing costly redundancy. We are gradually moving away from the paradigm of independently designing the digital and physical parts of hybrid systems towards simultaneous considerations for both. These systems have an acceptable tolerance determined by the application due to the inertial nature of the physical system, error tolerance capability in closed loop applications, built-in hardware and software functionality, etc. In this paper, we perform a comparative study of integrated circuit safety analysis as practiced today and system level application specific safety analysis that incorporates a physical system. We propose an improved method based upon the divide and conquer approach for such co-analysis to address practical limitations associated with adopting system level analysis techniques during integrated circuit design. Experimental results for a representative motor control system indicate that the application has an error tolerance of 92–160 cycles of closed loop operation for worst case errors and a control value error tolerance in the range of 5–7% at different operating conditions. Incorporation of application tolerance results in up to 4.3X reduction in the number of hardware elements which need to be protected. Prasanth Viswanathan Pillai, Rubin A. Parekhji, Bharadwaj S. Amrutur |
ITC | 3 |
| 2016 | A Variation-Tolerant Replica-Based Reference-Generation Technique for Single-Ended Sensing in Wide Voltage-Range SRAMsabstractThe most promising SRAM cells capable of operating over a wide range of supply voltages contain single-ended read ports. These systems require an external reference voltage that suitably scales to enable error-free operation of the memory, as the supply voltage is scaled. This paper presents a replica-based reference-generation technique for wide voltage range SRAMs. The proposed approach tracks the memory over the large range of supply voltages, and is tunable to extend functionality down to subthreshold voltages. In addition, a tunable delay-based timing-generation scheme is employed to enable memory functionality, in the presence of increased variation at subthreshold voltages. Configuration bits are set using a random-sampling-based Built-in Self-Test algorithm that significantly speeds up the tuning process. A 4-kb array, using the conventional 8T cell, implemented in the UMC 130-nm process, is demonstrated to function from 1.2 V down to 310 mV (at 1.3 MHz and 6.45 pJ/access). The memory consumes 0.115 pJ/bit/access at the energy optimum point of 400 mV. Viveka Konandur Rajanna, Bharadwaj S. Amrutur |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Improved Methods for Accurate Safety Analysis of Real-Life SystemsabstractIntegrated circuits are being used in different applications which are not always known at the time of specification and design creation. Safety standards specify that certain design processes be followed to guarantee safety of the applications in which these circuits are being used. As a result, the design phase is followed (often mandated) by an evaluation phase, wherein the safety worthiness of the circuit must be ascertained. In this paper, we perform a detailed study of such an evaluation as practised in the industry, understand the limitations, and propose techniques to improve the existing methodology. The improvements proposed are: (i) Capturing workload diversity as input constraints (values and sequence). (ii) Modelling application specific performance tolerance. (iii) Illustrating how physical system can be included into this analysis using a suitable representation. (iv) Budgeting of tolerance across various interacting modules to reduce computational complexity of safety analysis. Experimental results to illustrate suitability of the proposed methods are presented using a set of ITC benchmark circuits and two representative industrial circuits. Prasanth Viswanathan Pillai, Rubin A. Parekhji, Bharadwaj S. Amrutur |
ATS | 3 |
| 2015 | Performance analysis of wireless devices for a campus-wide IoT networkabstractTo select an appropriate technology for the deployment of an Internet-of-Things (IoT) network inside the Indian Institute of Science (IISc) campus, we first compare available wireless technologies based on their data sheets. After selecting two of the best available sub-GHz devices, we characterize them by performing controlled lab experiments. Next we test these sub-GHz modules in different real world environments such as open ground, straight road, moderately and densely wooded area, inside a concrete building and on building roof-tops. We then compare their performances for characterization of the wireless channels in different environments. In the end, we propose a sensor and network plan towards monitoring water resources inside the IISc campus. Nihesh Rathod, Pratik Jain, Renu Subramanian, Siddhesh Yawalkar, Mallikarjun Sunkenapally, Bharadwaj S. Amrutur, Rajesh Sundaresan |
WiOpt | 6 |
| 2015 | All Digital Energy Sensing for Minimum Energy TrackingabstractMinimizing energy consumption is of utmost importance in an energy starved system with relaxed performance requirements. This brief presents a digital energy sensing method that requires neither a constant voltage reference nor a time reference. An energy minimizing loop uses this to find the minimum energy point and sets the supply voltage between 0.2 and 0.5 V. Energy savings up to 1275% over existing minimum energy tracking techniques in the literature is achieved. Sagar Venkatesh Gubbi, Bharadwaj S. Amrutur |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A Digital Frequency Multiplication Technique for Energy Efficient TransmittersabstractA logic gate-based digital frequency multiplication technique for low-power frequency synthesis is presented. The proposed digital edge combining approach offers broadband operation with low-power and low-area advantages and is a promising candidate for low-power frequency synthesis in deep submicrometer CMOS technologies. Chip prototype of the proposed frequency multiplication-based 2.4-GHz binary frequency-shift-keying (BFSK)/amplitude shift keying (ASK) transmitter (TX) was fabricated in 0.13-μm CMOS technology. The TX achieves maximum data rates of 3 and 20 Mb/s for BFSK and ASK modulations, respectively, consuming a 14-mA current from 1.3 V supply voltage. The corresponding energy efficiencies of the TX are 3.6 nJ/bit for BFSK and 0.91 nJ/bit for ASK modulations. R. R. Manikandan, Abhishek Kumar 0007, Bharadwaj S. Amrutur |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Toward a Scalable Working Set Size Estimation Method and Its Application for Chip MultiprocessorsabstractIt is essential to accurately estimate the working set size (WSS) of an application for various optimizations such as to partition cache among virtual machines or reduce leakage power dissipated in an over-allocated cache by switching it off. However, the state-of-the-art heuristics such as average memory access latency (AMAL) or cache miss ratio (CMR) are poorly correlated to the WSS of an application due to 1) over-sized caches and 2) their dispersed nature. Past studies focus on estimating WSS of an application executing on a uniprocessor platform. Estimating the same for a chip multiprocessor (CMP) with a large dispersed cache is challenging due to the presence of concurrently executing threads/processes. Hence, we propose a scalable, highly accurate method to estimate WSS of an application. We call this method “tagged WSS (TWSS)” estimation method. We demonstrate the use of TWSS to switch-off the over-allocated cache ways in Static and Dynamic NonUniform Cache Architectures (SNUCA, DNUCA) on a tiled CMP. In our implementation of adaptable way SNUCA and DNUCA caches, decision of altering associativity is taken by each L2 controller. Hence,this approach scales better with the number of cores present on a CMP. It gives overall (geometric mean) 26% and 19% higher energy-delay product savings compared to AMAL and CMR heuristics on SNUCA, respectively. Aparna Mandke Dani, Bharadwaj S. Amrutur, Y. N. Srikant |
IEEE Trans. Computers | 2 |
| 2014 | Skip Decision and Reference Frame Selection for Low-Complexity H.264/AVC Surveillance Video CodingabstractH.264/advanced video coding surveillance video encoders use the Skip mode specified by the standard to reduce bandwidth. They also use multiple frames as reference for motion-compensated prediction. In this paper, we propose two techniques to reduce the bandwidth and computational cost of static camera surveillance video encoders without affecting detection and recognition performance. A spatial sampler is proposed to sample pixels that are segmented using a Gaussian mixture model. Modified weight updates are derived for the parameters of the mixture model to reduce floating point computations. A storage pattern of the parameters in memory is also modified to improve cache performance. Skip selection is performed using the segmentation results of the sampled pixels. The second contribution is a low computational cost algorithm to choose the reference frames. The proposed reference frame selection algorithm reduces the cost of coding uncovered background regions. We also study the number of reference frames required to achieve good coding efficiency. Distortion over foreground pixels is measured to quantify the performance of the proposed techniques. Experimental results show bit rate savings of up to 94.5% over methods proposed in literature on video surveillance data sets. The proposed techniques also provide up to 74.5% reduction in compression complexity without increasing the distortion over the foreground regions in the video sequence. Pushkar Gorur, Bharadwaj S. Amrutur |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Time-Based All-Digital Technique for Analog Built-in Self-TestabstractA scheme for built-in self-test of analog signals with minimal area overhead for measuring on-chip voltages in an all-digital manner is presented. The method is well suited for a distributed architecture, where the routing of analog signals over long paths is minimized. A clock is routed serially to the sampling heads placed at the nodes of analog test voltages. This sampling head present at each test node, which consists of a pair of delay cells and a pair of flip-flops, locally converts the test voltage to a skew between a pair of subsampled signals, thus giving rise to as many subsampled signal pairs as the number of nodes. To measure a certain analog voltage, the corresponding subsampled signal pair is fed to a delay measurement unit to measure the skew between this pair. The concept is validated by designing a test chip in a UMC 130-nm CMOS process. Sub-millivolt accuracy for static signals is demonstrated for a measurement time of a few seconds, and an effective number of bits of 5.29 is demonstrated for low-bandwidth signals in the absence of sample-and-hold circuitry. Rajath Vasudevamurthy, Pratap Kumar Das, Bharadwaj S. Amrutur |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | An 8-to-1 bit 1-MS/s SAR ADC With VGA and Integrated Data Compression for Neural RecordingabstractLow power consumption per channel and data rate minimization are two key challenges which need to be addressed in future generations of neural recording systems (NRS). Power consumption can be reduced by avoiding unnecessary processing whereas data rate is greatly decreased by sending spike time-stamps along with spike features as opposed to raw digitized data. Dynamic range in NRS can vary with time due to change in electrode-neuron distance or background noise, which demands adaptability. An analog-to-digital converter (ADC) is one of the most important blocks in a NRS. This paper presents an 8-bit SAR ADC in 0.13- μm CMOS technology along with input and reference buffer. A novel energy efficient digital-to-analog converter switching scheme is proposed, which consumes 37% less energy than the present state-of-the-art. The use of a ping-pong input sampling scheme is emphasized for multichannel input to alleviate the bandwidth requirement of the input buffer. To reduce the data rate, the A/D process is only enabled through the in-built background noise rejection logic to ensure that the noise is not processed. The ADC resolution can be adjusted from 8 to 1 bit in 1-bit step based on the input dynamic range. The ADC consumes 8.8 μW from 1 V supply at 1 MS/s speed. It achieves effective number of bits of 7.7 bits and FoM of 42.3 fJ/conversion-step. Vikram Chaturvedi, Tejasvi Anand, Bharadwaj S. Amrutur |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Efficient cache exploration method for a tiled chip multiprocessorabstractPast studies use deterministic models to evaluate optimal cache configuration or to explore its design space. However, with the increasing number of components present on a chip multiprocessor (CMP), deterministic approaches do not scale well. Hence, we apply probabilistic genetic algorithms (GA) to determine a near-optimal cache configuration for a sixteen tiled CMP. We propose and implement a faster trace based approach to estimate fitness of a chromosome. It shows up-to 218x simulation speedup over the cycle-accurate architectural simulation. Our methodology can be applied to solve other cache optimization problems such as design space exploration of cache and its partitioning among applications/virtual machines. Aparna Mandke Dani, Y. N. Srikant, Bharadwaj S. Amrutur |
HiPC | 3 |
| 2012 | TCP: Thread Contention Predictor for Parallel ProgramsabstractWith proliferation of chip multicores (CMPs) on desktops and embedded platforms, multi-threaded programs have become ubiquitous. Existence of multiple threads may cause resource contention, such as, in on-chip shared cache and interconnects, depending upon how they access resources. Hence, we propose a tool - Thread Contention Predictor (TCP) to help quantify the number of threads sharing data and their sharing pattern. We demonstrate its use to predict a more profitable shared, last level on-chip cache (LLC) access policy on CMPs. Our cache configuration predictor is 2.2 times faster compared to the cycle-accurate simulations. We also demonstrate its use for identifying hot data structures in a program which may cause performance degradation due to false data sharing. We fix layout of such data structures and show up-to 10% and 18% improvement in execution time and energy-delay product (EDP), respectively. Aparna Mandke Dani, Bharadwaj S. Amrutur, Y. N. Srikant, Chiranjib Bhattacharyya |
PDP | 2 |
| 2012 | Dynamic Supply and Threshold Voltage Scaling for CMOS Digital Circuits Using In-Situ Power MonitorabstractA generalized power tracking algorithm that minimizes power consumption of digital circuits by dynamic control of supply voltage and the body bias is proposed. A direct power monitoring scheme is proposed that does not need any replica and hence can sense total power consumed by load circuit across process, voltage, and temperature corners. Design details and performance of power monitor and tracking algorithm are examined by a simulation framework developed using UMC 90-nm CMOS triple well process. The proposed algorithm with direct power monitor achieves a power savings of 42.2% for activity of 0.02 and 22.4% for activity of 0.04. Experimental results from test chip fabricated in AMS 350 nm process shows power savings of 46.3% and 65% for load circuit operating in super threshold and near sub-threshold region, respectively. Measured resolution of power monitor is around 0.25 mV and it has a power overhead of 2.2% of die power. Issues with loop convergence and design tradeoff for power monitor are also discussed in this paper. Nandish Mehta, Bharadwaj S. Amrutur |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Speeded up Gaussian Mixture Model algorithm for background subtractionabstractAdaptive Gaussian Mixture Models (GMM) have been one of the most popular and successful approaches to perform foreground segmentation on multimodal background scenes. However, the good accuracy of the GMM algorithm comes at a high computational cost. An improved GMM technique was proposed by Zivkovic to reduce computational cost by minimizing the number of modes adaptively. In this paper, we propose a modification to his adaptive GMM algorithm that further reduces execution time by replacing expensive floating point computations with low cost integer operations. To maintain accuracy, we derive a heuristic that computes periodic floating point updates for the GMM weight parameter using the value of an integer counter. Experiments show speedups in the range of 1.33 - 1.44 on standard video datasets where a large fraction of pixels are multimodal. Pushkar Gorur, Bharadwaj S. Amrutur |
AVSS | 2 |
| 2011 | A power scalable receiver front-end at 2.4 GHzabstractWe propose a Low Noise Amplifier (LNA) architecture for power scalable receiver front end (FE) for Zigbee. The motivation for power scalable receiver is to enable minimum power operation while meeting the run-time performance needed. We use simple models to find empirical relations between the available signal and interference levels to come up with required Noise Figure (NF) and 3rd order Intermodulation Product (IIP3) numbers. The architecture has two independent digital knobs to control the NF and IIP3. Acceptable input match while using adaptation has been achieved by using an Active Inductor configuration for the source degeneration inductor of the LNA. The low IF receiver front end (LNA with I and Q mixers) was fabricated in 130 nm RFCMOS process and tested. Kaushik Ghosal, S. A. Kannan, Bharadwaj S. Amrutur |
ISCAS | 3 |
| 2011 | A mostly-digital analog scan-out chain for low bandwidth voltage measurement for analog IP testabstractA method of precise measurement of on-chip analog voltages in a mostly-digital manner, with minimal overhead, is presented. A pair of clock signals is routed to the node of an analog voltage. This analog voltage controls the delay between this pair of clock signals, which is then measured in an all-digital manner using the technique of sub-sampling. This sub-sampling technique, having measurement time and accuracy trade-off, is well suited for low bandwidth signals. This concept is validated by designing delay cells, using current starved inverters in UMC 130nm CMOS process. Sub-mV accuracy is demonstrated for a measurement time of few seconds. Rajath Vasudevamurthy, Pratap Kumar Das, Bharadwaj S. Amrutur |
ISCAS | 3 |
| 2011 | Memory Bandwidth and Power Reduction Using Lossy Reference Frame Compression in Video EncodingabstractLarge external memory bandwidth requirement leads to increased system power dissipation and cost in video coding application. Majority of the external memory traffic in video encoder is due to reference data accesses. We describe a lossy reference frame compression technique that can be used in video coding with minimal impact on quality while significantly reducing power and bandwidth requirement. The low cost transformless compression technique uses lossy reference for motion estimation to reduce memory traffic, and lossless reference for motion compensation (MC) to avoid drift. Thus, it is compatible with all existing video standards. We calculate the quantization error bound and show that by storing quantization error separately, bandwidth overhead due to MC can be reduced significantly. The technique meets key requirements specific to the video encode application. 24-39% reduction in peak bandwidth and 23-31% reduction in total average power consumption are observed for IBBP sequences. Ajit Gupte, Bharadwaj S. Amrutur, Mahesh Mehendale, Ajit V. Rao, Madhukar Budagavi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | 0.84 ps Resolution Clock Skew Measurement via SubsamplingabstractAn all-digital on-chip clock skew measurement system via subsampling is presented. The clock nodes are subsampled with a near-frequency asynchronous sampling clock to result in beat signals which are themselves skewed in the same proportion but on a larger time scale. The beat signals are then suitably masked to extract only the skews of the rising edges of the clock signals. We propose a histogram of the arithmetic difference of the beat signals which decouples the relationship of clock jitter to the minimum measurable skew, and allows skews arbitrarily close to zero to be measured with a precision limited largely by measurement time, unlike the conventional XOR based histogram approach. We also analytically show that the proposed approach leads to an unbiased estimate of skew. The measured results from a 65 nm delay measurement front-end indicate that for an input skew range of ±1 fan-out-of-4 (FO4) delay, ±3σ resolution of 0.84 ps can be obtained with an integral error of 0.65 ps. We also experimentally demonstrate that a frequency modulation on a sampling clock maintains precision, indicating the robustness of the technique to jitter. We also show how FM modulation helps in restoring precision in case of rationally related clocks. Bharadwaj S. Amrutur, Pratap Kumar Das, Rajath Vasudevamurthy |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Adaptive Keeper Design for Dynamic Logic Circuits Using Rate Sensing TechniqueabstractThe increasing variability in device leakage has made the design of keepers for wide OR structures a challenging task. The conventional feedback keepers (CONV) can no longer improve the performance of wide dynamic gates for the future technologies. In this paper, we propose an adaptive keeper technique called rate sensing keeper (RSK) that enables faster switching and tracks the variation across different process corners. It can switch upto 1.9× faster (for 20 legs) than CONV and can scale upto 32 legs as against 20 legs for CONV in a 130-nm 1.2-V process. The delay tracking is within 8% across the different process corners. We demonstrate the circuit operation of RSK using a 32 × 8 register file implemented in an industrial 130-nm 1.2-V CMOS process. The performance of individual dynamic logic gates are also evaluated on chip for various keeper techniques. We show that the RSK technique gives superior performance compared to the other alternatives such as Conditional Keeper (CKP) and current mirror-based keeper (LCR). Rakesh Gnana David Jeyasingh, Navakanta Bhat, Bharadwaj S. Amrutur |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | In-situ power monitoring scheme and its application in dynamic voltage and threshold scaling for digital CMOS integrated circuitsabstractAn in-situ power monitoring technique for Dynamic Voltage and Threshold scaling (DVTS) systems is proposed which measures total power consumed by load circuit using sleep transistor acting as power sensor. Design details of power monitor are examined using simulation framework in UMC 90nm CMOS process. Experimental results of test chip fabricated in AMS 0.35µm CMOS process are presented. The test chip has variable activity between 0.05 and 0.5 and has PMOS VTH control through nWell contact. Maximum resolution obtained from power monitor is 0.25mV. Overhead of power monitor in terms of its power consumption is 0.244 mW (2.2% of total power of load circuit). Lastly, power monitor is used to demonstrate closed loop DVTS system. DVTS algorithm shows 46.3% power savings using in-situ power monitor. Nandish Mehta, Gururaj V. Naik, Bharadwaj S. Amrutur |
ISLPED | 3 |
| 2010 | False Error Vulnerability Study of On-line Soft Error Detection Mechanisms
M. Kiran Kumar Reddy, Bharadwaj S. Amrutur, Rubin A. Parekhji |
J. Electron. Test. | 2 |
| 2010 | Voltage and Temperature Aware Statistical Leakage Analysis Framework Using Artificial Neural NetworksabstractArtificial neural networks (ANNs) have shown great promise in modeling circuit parameters for computer aided design applications. Leakage currents, which depend on process parameters, supply voltage and temperature can be modeled accurately with ANNs. However, the complex nature of the ANN model, with the standard sigmoidal activation functions, does not allow analytical expressions for its mean and variance. We propose the use of a new activation function that allows us to derive an analytical expression for the mean and a semi-analytical expression for the variance of the ANN-based leakage model. To the best of our knowledge this is the first result in this direction. Our neural network model also includes the voltage and temperature as input parameters, thereby enabling voltage and temperature aware statistical leakage analysis (SLA). All existing SLA frameworks are closely tied to the exponential polynomial leakage model and hence fail to work with sophisticated ANN models. In this paper, we also set up an SLA framework that can efficiently work with these ANN models. Results show that the cumulative distribution function of leakage current of ISCAS'85 circuits can be predicted accurately with the error in mean and standard deviation, compared to Monte Carlo-based simulations, being less than 1% and 2% respectively across a range of voltage and temperature values. Janakiraman Viraraghavan, Bharadwaj S. Amrutur |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | An adaptive, feature-based low power motion estimation algorithmabstractMotion Estimation is one of the most power hungry operations in video coding. While optimal search (eg. full search) methods give best quality, non optimal methods are often used in order to reduce cost and power. Various algorithms have been used in practice that trade off quality vs. complexity. Global elimination is an algorithm based on pixel averaging to reduce complexity of motion search while keeping performance close to that of full search. We propose an adaptive version of the global elimination algorithm that extracts individual macro-block features using Hadamard transform to optimize the search. Performance achieved is close to the full search method and global elimination. Operational complexity and hence power is reduced by 30% to 45% compared to global elimination method. Ajit Gupte, Bharadwaj S. Amrutur |
ICME | 2 |
| 2008 | False Error Study of On-line Soft Error Detection MechanismsabstractWith technology scaling, vulnerability to soft errors in random logic is increasing. There is a need for on-line error detection and protection for logic gates even at sea level. The error checker is the key element for an on-line detection mechanism. We compare three different checkers for error detection from the point of view of area, power and false error detection rates. We find that the double sampling checker (used in Razor), is the simplest and most area and power efficient, but suffers from very high false detection rates of 1.15 times the actual error rates. We also find that the alternate approaches of triple sampling and integrate and sample method (I&S) can be designed to have zero false detection rates, but at an increased area, power and implementation complexity. The triple sampling method has about 1.74 times the area and twice the power as compared to the Double Sampling method and also needs a complex clock generation scheme. The I&S method needs about 16% more power with 0.58 times the area as double sampling, but comes with more stringent implementation constraints as it requires detection of small voltage swings. M. Kiran Kumar Reddy, Bharadwaj S. Amrutur, Rubin A. Parekhji |
IOLTS | 2 |
| 2008 | Quadrature generation techniques for frequency multiplication based oscillatorsabstractFrequency multiplication (FM) can be used to design low power frequency synthesizers. This is achieved by running the VCO at a much reduced frequency, while employing a power efficient frequency multiplier, and also thereby eliminating the first few dividers. Quadrature signals can be generated by frequency- multiplying low frequency I/Q signals, however this also multiplies the quadrature error of these signals. Another way is generating additional edges from the low-frequency oscillator (LFO) and develop a quadrature FM. This makes the I-Q precision heavily dependent on process mismatches in the ring oscillator. In this paper we examine the use of fewer edges from LFO and a single stage polyphase filter to generate approximate quadrature signals, which is then followed by an injection-locked quadrature VCO to generate high- precision I/Q signals. Simulation comparisons with the existing approach shows that the proposed method offers very good phase accuracy of 0.5deg with only a modest increase in power dissipation for 2.4 GHz IEEE 802.15.4 standard using UMC 0.13 mum RFCMOS technology. Jagdish Nayayan Pandey, Bharadwaj S. Amrutur, Sudhir S. Kudva |
ISCAS | 2 |
| 2008 | Power reduction in on-chip interconnection network by serializationabstractWe explore the use of serialization in on-chip buses for reducing interconnect energy. Serialization reduces wire density and hence the coupling capacitance between adjacent data bits. This enables higher data rates, thus making it possible to send multiple data bits on a single wire within a single clock cycle. Energy reduction is brought about as a result of the decreased coupling capacitance, however this is offset by increased size and number of repeaters to obtain higher speed. A critical delay exists above which serialization is more energy efficient. We find this critical delay for a 2:1 serialization by solving an optimization problem formulated as minimization of power with serialization factor (1 or 2), area, bandwidth and frequency as constraints, and having repeater size, number of repeaters, and wire dimensions as design variables. We find that for delays above 40% of minimum delay for the wire, double pumping is more energy efficient across a range of technology nodes and supply voltages and matches well with a simple analytical derivation. Arvind Madan, Bharadwaj S. Amrutur |
ISLPED | 2 |
| 2008 | Optimal power and noise allocation for analog and digital sections of a low power radio receiverabstractWe determine the optimal allocation of power between the analog and digital sections of an RF receiver, while meeting the BER constraint. Unlike conventional RF receiver designs, we treat the SNR at the output of the analog front end (SNRAD) as a design parameter rather than a specification to arrive at this optimal allocation. We first determine the relationship of the SNRAD to the resolution and operating frequency of the digital section. We then use power models for the analog and digital sections to solve the power minimization problem. As an example, we consider a 802.15.4 compliant low-IF receiver operating at 2.4 GHz in 0.13μm technology with 1.2 V power supply. We find that the overall receiver power is minimized by having the analog front end provide an SNR of 1.3dB and the ADC and the digital section operate at 1-bit resolution with 18MHz sampling frequency while achieving a power dissipation of 7mW. Kannan A. Sankaragomathi, Manodipan Sahoo, Satyam Dwivedi, Bharadwaj S. Amrutur, Navakanta Bhat |
ISLPED | 4 |
| 2007 | Dynamic Cache Placement with Two-level Mapping to Reduce Conflict Misses
Kaushik Rajan, R. Govindarajan, Bharadwaj S. Amrutur |
PACT | 3 |
| 2007 | INTACTE: an interconnect area, delay, and energy estimation tool for microarchitectural explorationsabstractPrior work on modeling interconnects has focused on optimizing the wire and repeater design for trading off energy and delay, and is largely based on low level circuit parameters. Hence these models are hard to use directly to make high level microarchitectural trade-offs in the initial exploration phase of a design. In this paper, we propose INTACTE, a tool that can be used by architects toget reasonably accurate interconnect area, delay, and power estimates based on a few architecture level parameters for the interconnect such as length, width (in number of bits), frequency, and latency for a specified technology and voltage. Rahul Nagpal, Arvind Madan, Bharadwaj S. Amrutur, Y. N. Srikant |
CASES | 3 |
| 2007 | Area efficient loop filter design for charge pump phase locked loopabstractIn this paper, two new dual-path based area efficient loop filter circuits are proposed for Charge Pump Phase Locked Loop(CPPLL). The proposed circuits were designed in 0.25μ CSM analog process with 1.8V supply. The proposed circuits achieved up to 85% savings in capacitor area. Simulations showed good match of the new circuits with the conventional circuit. The proposed circuits are particularly useful in applications that demand low die area. R. G. Raghavendra, Bharadwaj S. Amrutur |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | A scalable low power issue queue for large instruction window processorsabstractLarge instruction windows and issue queues are key to exploiting greater instruction level parallelism in out-of-order superscalar processors. However, the cycle time and energy consumption of conventional large monolithic issue queues are high. Previous efforts to reduce cycle time segment the issue queue and pipeline wakeup. Unfortunately, this results in significant IPC loss. Other proposals which address energy efficiency issues by avoiding only the unnecessary tag-comparisons do not reduce broadcasts. These schemes also increase the issue latency.To address both these issues comprehensively, we propose the Scalable Lowpower Issue Queue (SLIQ). SLIQ augments a pipelined issue queue with direct indexing to mitigate the problem of delayed wakeups while reducing the cycle time. Also, the SLIQ design naturally leads to significant energy savings by reducing both the number of tag broadcasts and comparisons required.A 2 segment SLIQ incurs an average IPC loss of 0.2% over the entire SPEC CPU2000 suite, while achieving a 25.2% reduction in issue latency when compared to a monolithic 128-entry issue queue for an 8-wide superscalar processor. An 8 segment SLIQ improves scalability by reducing the issue latency by 38.3% while incurring an IPC loss of only 2.3%. Further, the 8 segment SLIQ significantly reduces the energy consumption and energy-delay product by 48.3% and 67.4% respectively on average. Rajesh Vivekanandham, Bharadwaj S. Amrutur, R. Govindarajan |
ICS | 2 |
| 2006 | Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regionsabstractCMPs enable simultaneous execution of multiple applications on the same platforms that share cache resources. Diversity in the cache access patterns of these simultaneously executing applications can potentially trigger inter-application interference, leading to cache pollution. Whereas a large cache can ameliorate this problem, the issues of larger power consumption with increasing cache size, amplified at sub-100nm technologies, makes this solution prohibitive. In this paper, in order to address the issues relating to power-aware performance of caches, we propose a caching structure that addresses the following: 1) Definition of application-specific cache partitions as an aggregation of caching units (molecules). The parameters of each molecule namely size, associativity and line size are chosen so that the power consumed by it and access time are optimal for the given technology. 2) Application-specific resizing of cache partitions with variable and adaptive associativity per cache line, way size and variable line size. 3) A replacement policy that is transparent to the partition in terms of size, heterogeneity in associativity and line size. Through simulation studies we establish the superiority of molecular cache (caches built as aggregations of molecules) that offers a 29% power advantage over that of an equivalently performing traditional cache Keshavan Varadarajan, S. K. Nandy 0001, Vishal Sharda, Bharadwaj S. Amrutur, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell |
MICRO | 4 |
| 1992 | A projective geometry architecture for scientific computationabstractA large fraction of scientific and engineering computations involve sparse matrices. While dense matrix computations can be parallelized relatively easily, sparse matrices with arbitrary or irregular structure pose a real challenge to designers of highly parallel machines. A recent paper by N.K. Karmarkar (1991) proposed a new parallel architecture for sparse matrix computations based on finite projective geometries. Mathematical structure of these geometries plays an important role in defining the interconnections between the processors and memories in this architecture, and also aids in efficiently solving several difficult problems (such as load balancing, data-routing, memory-access conflicts, etc.) that are encountered in the design of parallel systems. The authors discuss some of the key issues in the system design of such a machine, and show how exploiting the structure of the geometry results in an efficient hardware implementation of the machine. They also present circuit designs and simulation results for key elements of the system: a 200 MHz pipelined memory; a pipelined multiplier based on an adder unit with a delay of 2 ns; and a 500 Mbit/s CMOS input/output buffer.> Bharadwaj S. Amrutur, Rajeev Joshi, Narendra K. Karmarkar |
ASAP | 1 |