Vijay Raghunathan

dblp:80/6713 · DBLP profile ↗
← Back
75ranked-venue papers
11as first author
10since 2021 · last 2026
0000-0003-4713-5386ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 56 · 9 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 2 since 2021Computer networks · 15 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 2 · 1 since 2021
YearPublicationVenuePosition
2026 COSMOS: Designing Energy-Efficient Context-Aware Multimodal Cognitive Systems
Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED5
2026 InterAxNN: Reconfigurable and Approximate in-Memory Processing Accelerator for Ultra-Low-Power Binary Neural Network Inference in Intermittently Powered Systems
abstract
In this work, we propose InterAxNN , an energy-aware approximate hardware architecture to perform vector-matrix multiplications in the binary precision regime for energy-constrained intermittently powered systems (IPS). In contrast to existing XNOR multiply-and-accumulate (MAC) operations implemented widely for binary neural networks (BNNs), we design a novel reconfigurable XNOR-MAC and AND-MAC memory macro to perform approximate binary precision operations, targeted for systems with extreme energy constraints. The proposed macro design is integrated with the ability to modify the MAC mode during run-time depending on instantaneous energy and power transients. We utilize the unique attributes of ferroelectric transistors (FeFETs) to implement the proposed ultra-low power BNN engine performing in-memory computing for artificial intelligence (AI) workloads. Subsequently, we leverage the quality configurable compute-in-memory-based hardware accelerator to implement InterAxNN based on a TI MSP430-based microcontroller. We evaluate the proposed InterAxNN concerning two baselines: (a) standard von Neumann computing architecture-based-microcontroller platform (MCU), and (b) MCU with a state-of-the-art low energy accelerator (MCU+LEA), and observe significant performance and energy benefits. Experimental results performed using a TI MSP430FR5379 IPS system show 448×–581× uplift in forward progress for 2%–8% accuracy loss for MNIST, 4%–5% accuracy loss for EMNIST, and 1%–4% reduction in accuracy for QMNIST, respectively, using MLP on a MCU+LEA platform with Unified NVM architecture. The AND-MAC mode in InterAxNN results in 91×–127× amount of additional forward progress over XNOR-MAC for 1%–2%, 4%–19%, and 1%–5% higher quality degradation for MNIST, EMNIST, and QMNIST, respectively.
Arnab Raha, Sandeep Krishna Thirumala, Sumeet Kumar Gupta, Vijay Raghunathan
ACM Trans. Design Autom. Electr. Syst.4
2025 GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
abstract
Graph Neural Networks (GNNs) are crucial for learning and reasoning over graph-structured data, with applications in network analysis, recommendation systems, and speech analytics. Deploying them on edge devices, such as client PCs and laptops, enables real-time processing, enhances privacy, and reduces cloud dependency. For instance, GNNs can augment Retrieval-Augmented Generation (RAG) for Large Language Models (LLMs) and enable event-based vision tasks. However, irregular memory access, sparse graphs, and dynamic structures lead to high latency and energy consumption on resource-constrained devices. Modern edge processors combine CPUs, GPUs, and NPUs, where NPUs excel at data-parallel tasks but face challenges with irregular GNN computations. To address these gaps, we present GraNNite, the first hardware-aware framework tailored to optimize GNN deployment on commercial-off-the-shelf (COTS) state-of-the-art (SOTA) DNN accelerators using a systematic three-step methodology: (1) enabling GNN execution on NPUs, (2) optimizing performance, and (3) trading accuracy for further performance and energy efficiency gains. Towards that end, the first category includes techniques such as GraphSplit for workload distribution and StaGr for static graph aggregation, while GrAd and NodePad handle real-time updates for dynamic graphs. Next, performance improvement is acquired through techniques such as EffOp for control-heavy operations and GraSp for sparsity exploitation. For Graph Convolution layers, PreG, SymG, and CacheG reduce redundancy and memory transfers. The final class of techniques deals with quality vs efficiency tradeoffs – QuantGr applies INT8 quantization to lower memory usage and computation time, while GrAx1, GrAx2, and GrAx3 optimize graph attention, broadcast-add, and sample-and-aggregate (SAGE)-max aggregation for higher throughput with minimal quality loss. Experimental evaluations on Intel® Core™ Ultra Series 1 and 2 AI PCs demonstrate that GraNNite achieves speedups of 2.6× to 7.6× over default NPU mappings, with energy efficiency improvements up to 8.6× compared to CPUs and GPUs. Across various GNN models, GraNNite delivers up to 10.8× and 6.7× higher performance than CPUs and GPUs, respectively. Our code implementation is available at this link.
Arghadip Das, Shamik Kundu, Arnab Raha, Soumendu Kumar Ghosh, Deepak Mathaikutty, Vijay Raghunathan
IJCNN6
2025 Demo Abstract: ECO: Low Power Context-Aware Multimodal AI on NPUs
abstract
We present ECO, the first system enabling efficient multimodal AI deployment on commercial Neural Processing Units (NPUs) through context-aware sensor and compute optimizations. ECO introduces runtime-tunable, NPU-architecture-aware knobs—approximate interpolation, quantization, and model scaling—that adapt to system conditions such as energy availability and sensor reliability. Deployed on an Intel Core Ultra Series 2 NPU with RGB and LiDAR inputs for a semantic segmentation application, ECO achieves up to 4.9× performance and 11.3× energy-efficiency improvement over CPU. Compared to systems lacking runtime context adaptability, ECO preserves higher segmentation quality (48.1 mean IoU in %, referred to as IoU hereafter) vs. 37.9 IoU under energy constraints and restores accuracy from 30.6 IoU to 40.0 IoU in sensor failure scenarios. The demo video and the ECO codebase are available at https://github.com/arghadippurdue/ECO%5FDemo.
Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED5
2025 SecuPilot: A Security Coprocessor-Integrated Platform for Autonomous UAV Security
abstract
This article introduces SecuPilot, a Security Coprocessor-integrated platform designed to enhance the resilience and operational security of autonomous Unmanned Aerial Vehicles (UAVs) in increasingly adversarial environments. Recognizing the critical role UAVs play in diverse applications, SecuPilot builds upon established architectures by incorporating a dedicated security module that performs a series of comprehensive preflight checks to establish a robust root of trust. Once airborne, the platform continuously monitors the UAV’s operational state during runtime to infer and maintain its security, ensuring that any anomalies are promptly identified and addressed without disrupting mission-critical functions. We develop a custom Hardware-in-the-Loop (HITL) simulation framework replicating realistic operational scenarios and adversarial conditions to validate the system’s performance and effectiveness. This rigorous evaluation demonstrates that SecuPilot can successfully mitigate potential threats while preserving the essential performance characteristics of the UAV. The results underscore the viability of a scalable, hardware-centric approach to UAV security, paving the way for safer autonomous systems capable of operating in complex, threat-prone environments.
Yatharth Agarwal, Vijay Raghunathan
ACM Trans. Embed. Comput. Syst.2
2024 Toward Energy-Efficient Collaborative Inference Using Multisystem Approximations
abstract
Cooperative inference applications have seen considerable potential with distributed deep neural networks (DDNNs). One use for DDNNs is the classification of 3-D objects from a set of 2-D images or views. This approach is also known as multiview convolutional neural networks (MVCNNs). However, due to the intensive computational demands, substantial communication overhead, high-inference delay, and energy limits, it is difficult to deploy MVCNN on resource-constrained edge devices. This article proposes for the first time the concept of distributed approximate systems (DRAX), which employs a multidevice approach to approximate computing and uses synergistic approximations of various edge computing systems to enable energy-efficient collaborative DDNN inference.DRAXperforms a significance-aware approximation of multiple nodes and prunes the large design space using the nonuniform contribution of various perspectives/views to the final inference to achieve optimal quality-energy tradeoff. In addition, we also propose a novel remaining energy-aware heuristic, which dynamically chooses the approximation degree based on the user-provided quality bounds and further increases the system lifetime. The experimental results obtained from a prototype of a 12-view 3-D object classification system implemented on an Intel Stratix IV FPGA development board demonstrate substantial energy savings ($2.6 \times$to$8\times$) for minimal (<1%) application-level quality loss.
Arghadip Das, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
IEEE Internet Things J.4
2024 PArtNNer: Platform-Agnostic Adaptive Edge-Cloud DNN Partitioning for Minimizing End-to-End Latency
abstract
The last decade has seen the emergence of Deep Neural Networks (DNNs) as the de facto algorithm for various computer vision applications. In intelligent edge devices, sensor data streams acquired by the device are processed by a DNN application running on either the edge device itself or in the cloud. However, “edge-only” and “cloud-only” execution of State-of-the-Art DNNs may not meet an application’s latency requirements due to the limited compute, memory, and energy resources in edge devices, dynamically varying bandwidth of edge-cloud connectivity networks, and temporal variations in the computational load of cloud servers. This work investigates distributed (partitioned) inference across edge devices (mobile/end device) and cloud servers to minimize end-to-end DNN inference latency. We study the impact of temporally varying operating conditions and the underlying compute and communication architecture on the decision of whether to run the inference solely on the edge, entirely in the cloud, or by partitioning the DNN model execution among the two. Leveraging the insights gained from this study and the wide variation in the capabilities of various edge platforms that run DNN inference, we propose PArtNNer , a platform-agnostic adaptive DNN partitioning algorithm that finds the optimal partitioning point in DNNs to minimize inference latency. PArtNNer can adapt to dynamic variations in communication bandwidth and cloud server load without requiring pre-characterization of underlying platforms. Experimental results for six image classification and object detection DNNs on a set of five commercial off-the-shelf compute platforms and three communication standards indicate that PArtNNer results in 10.2× and 3.2× (on average) and up to 21.1× and 6.7× improvements in end-to-end inference latency compared to execution of the DNN entirely on the edge device or entirely on a cloud server, respectively. Compared to pre-characterization-based partitioning approaches, PArtNNer converges to the optimal partitioning point 17.6× faster.
Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan, Anand Raghunathan
ACM Trans. Embed. Comput. Syst.3
2023 Energy-Efficient Approximate Edge Inference Systems
abstract
The rapid proliferation of the Internet of Things and the dramatic resurgence of artificial intelligence based application workloads have led to immense interest in performing inference on energy-constrained edge devices. Approximate computing (a design paradigm that trades off a small degradation in application quality for disproportionate energy savings) is a promising technique to enable energy-efficient inference at the edge. This article introduces the concept of an approximate edge inference system ( AxIS ) and proposes a systematic methodology to perform joint approximations between different subsystems in a deep neural network (DNN)-based edge inference system, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use a smart camera system that executes various DNN-based image classification and object detection applications to illustrate how the sensor, memory, compute, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: (a) Cam Edge , where the DNN is executed locally on the edge device, and (b) Cam Cloud , where the edge device sends the captured image to a remote cloud server that executes the DNN. We have prototyped such an approximate inference system using an Intel Stratix IV GX-based Terasic TR4-230 FPGA development board. Experimental results obtained using six large DNNs and four compact DNNs running image classification applications demonstrate significant energy savings (≈ 1.6× -4.7× for large DNNs and ≈ 1.5× -3.6× for small DNNs), for minimal (<1%) loss in application-level quality. Furthermore, results using four object detection DNNs exhibit energy savings of ≈ 1.5× -5.2× for similar quality loss. Compared to approximating a single subsystem in isolation, AxIS achieves 1.05× -3.25× gains in energy savings for image classification and 1.35× -4.2× gains for object detection on average, for minimal (<1%) application-level quality loss.
Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ACM Trans. Embed. Comput. Syst.3
2022 Exploring the Design of Energy-Efficient Intermittently Powered Systems Using Reconfigurable Ferroelectric Transistors
abstract
In this article, we explore the design of energy-efficient intermittently powered systems (IPSs) using reconfigurable-ferroelectric transistors (R-FEFETs). Utilizing the dynamic tunability between volatile and nonvolatile modes of operation in R-FEFETs, we design nonvolatile flip-flops (NVFFs) and memory (NVM) suitable for IPS. We present two variants of R-FEFET-based NVFFs (RNVFFs): 1) with automatic backup and 2) with need-based backup. While the former offers high backup energy efficiency, the latter offers low normal operation energy. We also present an IPS-specific R-FEFET-based NVM (3T-R) with high energy efficiency compared with FEFET-based 2T NVM. Leveraging these nonvolatile circuits, we map the microcontroller unit (MCU) core registers of an IPS to RNVFFs and its on-chip memory to 3T-R. Subsequently, we analyze system-level implications of improving NVFFs and NVM individually by using R-FEFETs compared with existing FEFET-based designs. Our system-level simulations demonstrate that although we improve the register energy by 55%–67%, the total memory and system-level energy savings obtained from just improving the NVFFs (registers) in the microcontroller core are only 0.60%–5.78% and 0.31%–3.18%, respectively. However, improving the NVM by using 3T-R results in a much larger total memory and system-level energy savings in the range of 37%–40% and 20%–22%, respectively, in the context of a state-of-the-art IPS.
Sandeep Krishna Thirumala, Arnab Raha, Sumeet Kumar Gupta, Vijay Raghunathan
IEEE Trans. Very Large Scale Integr. Syst.4
2021 Special Session: Approximate TinyML Systems: Full System Approximations for Extreme Energy-Efficiency in Intelligent Edge Devices
abstract
Approximate computing (AxC) has advanced from being an emerging design paradigm to becoming one of the most popular and effective methods of energy optimization for applications in the domains of computer vision, image/video processing, data mining, analytics, and search. The simultaneous rise of artificial intelligence (AI) has provided an additional thrust to the adoption of various AxC techniques in intelligent edge platforms where energy-efficiency is not only desirable but necessary. In spite of the big rise in interest for AxC, the adoption of approximate hardware has mostly been limited to only one component of the system (usually the processing subsystem) which often contributes only a fraction of the overall system-level power. A full system approach to AxC enables us to extend approximations to other subsystems, such as the memory, sensor, and communications subsystems. This paper presents the foundational concepts of an approximate TinyML system that applies approximations synergistically to multiple subsystems in an edge inference device. These approximations are applied intelligently to significantly reduce energy while incurring a negligible loss in application-level quality. We demonstrate multiple versions of an approximate smart camera system that can execute state-of-the-art deep neural networks (DNNs) while consuming only a fraction of the total energy in a typical system.
Arnab Raha, Soumendu Kumar Ghosh, Debabrata Mohapatra, Deepak Mathaikutty, Raymond Sung, Cormac Brick, Vijay Raghunathan
ICCD7
2020 Communication-efficient View-Pooling for Distributed Multi-View Neural Networks
abstract
Multi-view object detection or the problem of detecting an object using multiple viewpoints, is an important problem in computer vision with varied applications such as distributed smart cameras and collaborative drone swarms. Multi-view object detection algorithms based on deep neural networks (DNNs) achieve high accuracy by view pooling, or aggregating features corresponding to the different views. However, when these algorithms are realized on networks of edge devices, the communication cost incurred by view pooling often dominates the overall latency and energy consumption.In this paper, we propose techniques for communication-efficient view pooling that can be used to improve the efficiency of distributed multi-view object detection and apply them to state-of-the-art multi-view DNNs. First, we propose significance-aware feature selection, which identifies and communicates only those features from each view that are likely to impact the pooled result (and hence, the final output of the DNN). Second, we propose multi-resolution view pooling, which divides views into dominant and non-dominant views, and down-scales the features from non-dominant views using an additional network layer before communicating them for pooling. The dominant and non-dominant views are pooled separately and the results are jointly used to derive the final classification. We implement and evaluate the proposed pooling schemes using a model test-bed of twelve Raspberry Pi 3b+ devices and show that they achieve 9X - 36X reduction in data communicated and 1.8X reduction in inference latency, with no degradation in accuracy.
Manik Singhal, Vijay Raghunathan, Anand Raghunathan
DATE2
2020 IPS-CiM: Enhancing Energy Efficiency of Intermittently-Powered Systems with Compute-in-Memory
abstract
Intermittently Powered Systems (IPS) have an ability to sustain computation progress across multiple power cycles in the presence of unreliable and sporadic harvested energy. However, with the emergence of data-intensive applications to be processed on energy-constrained IPS, it becomes challenging to handle large amounts of data with standard IPS architectures due to the von-Neumann bottleneck. To address this issue, we propose a compute-in-memory (CiM) engine which alleviates the memory-processor bottleneck and enhances energy-efficiency for transient computing workloads in IPS. We present a ferroelectric transistor (FEFET) based memory architecture which supports (a) nonvolatile memory (NVM) storage, (b) standard Boolean and arithmetic operations, (c) cyclic redundancy check for error detection and (d) edge-sensing for wireless sensory networks. Using the proposed CiM engine as a unified NVM, we construct an integrated IPS-CiM architecture based on the TI MSP430 microcontroller system with supply capacitances in the range of 10 nF -1 μF. We evaluate the proposed design with two baselines: hybrid SRAM+NVM and unified NVM architectures, both of which perform standard out-of-memory computing. We observe that for 1μF supply capacitance, IPS-CiM results in energy and performance benefits in the range of 35X-450X and 32X-400X, respectively over conventional microcontroller-based systems.
Sandeep Krishna Thirumala, Arnab Raha, Vijay Raghunathan, Sumeet Kumar Gupta
ICCD3
2020 Approximate inference systems (AxIS): end-to-end approximations for energy-efficient inference at the edge
abstract
The rapid proliferation of the Internet-of-Things (IoT) and the dramatic resurgence of artificial intelligence (AI) based application workloads has led to immense interest in performing inference on energy-constrained edge devices. Approximate computing (a design paradigm that yields large energy savings at the cost of a small degradation in application quality) is a promising technique to enable energy-efficient inference at the edge. This paper introduces the concept of an approximate inference system (AxIS) and proposes a systematic methodology to perform joint approximations across different subsystems in a deep neural network-based inference system, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use a smart camera system that executes various convolutional neural network (CNN) based image recognition applications to illustrate how the sensor, memory, compute, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: (a) Camedge, where the CNN executes locally on the edge device, and (b) Camcloud, where the edge device sends the captured image to a remote cloud server that executes the CNN. We have prototyped such an approximate inference system using an Altera Stratix IV GX-based Terasic TR4-230 FPGA development board. Experimental results obtained using six CNNs demonstrate significant energy savings (around 1.7× for Camedge and 3.5× for Camcloud) for minimal (< 1%) loss in application quality. Compared to approximating a single subsystem in isolation, AxIS achieves additional energy benefits of 1.6×--1.7× (Camedge) and 1.4×--3.4× (Camcloud) on average for minimal application-level quality loss.
Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED3
2020 Approximate Memory Compression
abstract
Memory subsystems are a major energy bottleneck in computing platforms due to frequent transfers between processors and off-chip memory. We propose approximate memory compression, a technique that leverages the intrinsic resilience of emerging workloads such as machine learning and data analytics to reduce off-chip memory traffic, thereby improving energy and performance. We realize approximate memory compression by enhancing the memory controller to be aware of approximate memory regions-regions in memory that contain approximation-resilient data-and to transparently compress (decompress) the data written to (read from) these regions. To provide control over approximations, each approximate memory region is associated with an error constraint such as the maximum error that may be introduced in each data element. The quality-aware memory controller subjects memory transactions to a compression scheme that introduces approximations, thereby reducing memory traffic, while adhering to the specified error constraint for each approximate memory region. A software interface is provided to allow programmers to identify data structures (DSs) that are resilient to approximations. A runtime quality control framework automatically determines the error constraints for the identified DSs such that a given target application-level quality is maintained. We evaluate our proposal by applying it to three different main memory technologies in the context of a general-purpose computing system-DDR3 DRAM, LPDDR3 DRAM, and spin-transfer torque magnetic RAM (STT-MRAM). To demonstrate the feasibility of the proposed concepts, we also implement a hardware prototype using the Intel UniPHY-DDR3 memory controller and Nios-II processor, a Hynix DDR3 DRAM module, and a Stratix-IV field-programmable gate array (FPGA) development board. Across a wide range of machine learning benchmarks, approximate memory compression obtains significant benefits in main memory energy (1.18× for DDR3 DRAM, 1.52× for LPDDR3 DRAM, and 2.0× for STT-MRAM) and a simultaneous improvement in execution time (5.2% for DDR3 DRAM, 5.4% for LPDDR3 DRAM, and 9.3% for STT-MRAM) with nearly identical application output quality.
Ashish Ranjan 0001, Arnab Raha, Vijay Raghunathan, Anand Raghunathan
IEEE Trans. Very Large Scale Integr. Syst.3
2018 SYNCVIBE: Fast and Secure Device Pairing through Physical Vibration on Commodity Smartphones
abstract
The emergence of the Internet of Things (IoT) and pervasive computing challenges in securely and conveniently connecting devices with limited user interfaces. In particular, discovering and bootstrapping a wireless connection (e.g., Wi-Fi and Bluetooth Low Energy) between two devices that share no prior knowledge, commonly known as pairing, often requires users to go through cumbersome tasks of manually discovering the target device and entering a long passkey. When the devices do not have a proper user interface to enter a passkey, the security of pairing is often given up, leaving the communication vulnerable to a number of attacks. To alleviate this challenge, we propose a usable and secure out-of-band (OOB) communication method called SyncVibe, leveraging the inherent nature of close-proximity transmission of mechanical vibration. SyncVibe utilizes a vibration motor and an accelerometer, that are already ubiquitously available or easy to embed in mobile and wearable devices, to transmit and receive pairing information. By simply keeping two devices in direct contact, the user can bootstrap a secure, high-bandwidth wireless connection without manual pairing procedures. The proposed method maximizes accuracy and effective data throughput with a vibration clock recovery technique, which inserts a minimal amount of extra bit patterns to assure synchronization between the transmitter and the receiver. In addition, SyncVibe can automatically adjust its detection thresholds in response to various vibration noises and transmission media. Our implementation of SyncVibe demonstrates high-accuracy transmission, proving itself as a suitable OOB communication channel for short data transmission for secure device pairing.
Kyuin Lee, Vijay Raghunathan, Anand Raghunathan, Younghyun Kim 0001
ICCD2
2018 Dual Mode Ferroelectric Transistor based Non-Volatile Flip-Flops for Intermittently-Powered Systems
abstract
In this work, we propose dual mode ferroelectric transistors (D-FEFETs) that exhibit dynamic tuning of operation between volatile and non-volatile modes with the help of a control signal. We utilize the unique features of D-FEFET to design two variants of non-volatile flip-flops (NVFFs). In both designs, D-FEFETs are operated in the volatile mode for normal operations and in the non-volatile mode to backup the state of the flip-flop during a power outage. The first design comprises of a truly embedded non-volatile element (D-FEFET) which enables a fully automatic backup operation. In the second design, we introduce need-based backup, which lowers energy during normal operation at the cost of area with respect to the first design. Compared to a previously proposed FEFET based NVFF, the first design achieves 19% area reduction along with 96% lower backup energy and 9% lower restore energy, but at 14%-35% larger operation energy. The second design shows 11% lower area, 21% lower backup energy, 16% decrease in backup delay and similar operation energy but with a penalty of 17% and 19% in the restore energy and delay, respectively. System-level analysis of the proposed NVFFs in context of a state-of-the-art intermittently-powered system using real benchmarks yielded 5%-33% energy savings.
Sandeep Krishna Thirumala, Arnab Raha, Hrishikesh Jayakumar, Kaisheng Ma, Narayanan Vijaykrishnan, Vijay Raghunathan, Sumeet Kumar Gupta
ISLPED6
2018 D-PUF: An Intrinsically Reconfigurable DRAM PUF for Device Authentication and Random Number Generation
abstract
Physically Unclonable Functions (PUFs) have proved to be an effective and low-cost measure against counterfeiting by providing device authentication and secure key storage services. Memory-based PUF implementations are an attractive option due to the ubiquitous nature of memory in electronic devices and the requirement of minimal (or no) additional circuitry. Dynamic Random Access Memory-- (DRAM) based PUFs are particularly advantageous due to their large address space and multiple controllable parameters during response generation. However, prior works on DRAM PUFs use a static response-generation mechanism making them vulnerable to security attacks. Further, they result in slow device authentication, are not applicable to commercial off-the-shelf devices, or require DRAM power cycling prior to authentication. In this article, we propose D-PUF, an intrinsically reconfigurable DRAM PUF based on the idea of DRAM refresh pausing. A key feature of the proposed DRAM PUF is reconfigurability , that is, by varying the DRAM refresh-pause interval, the challenge-response behavior of the PUF can be altered, making it robust to various attacks. The article is broadly divided into two parts. In the first part, we demonstrate the use of D-PUF in performing device authentication through a secure, low-overhead methodology. In the second part, we show the generation of true random numbers using D-PUF. The design is implemented and validated using an Altera Stratix IV GX FPGA-based Terasic TR4-230 development board and several off-the-shelf 1GB DDR3 DRAM modules. Our experimental results demonstrate a 4.3×-6.4× reduction in authentication time compared to prior work. Using controlled temperature and accelerated aging tests, we also demonstrate the robustness of our authentication mechanism to temperature variations and aging effects. Finally, the ability of the design to generate random numbers is verified using the NIST Statistical Test Suite.
Soubhagya Sutar, Arnab Raha, Devadatta M. Kulkarni, Rajeev Shorey, Jeffrey D. Tew, Vijay Raghunathan
ACM Trans. Embed. Comput. Syst.6
2018 Designing Energy-Efficient Intermittently Powered Systems Using Spin-Hall-Effect-Based Nonvolatile SRAM
abstract
Intermittently powered systems represent a new class of batteryless devices that operate solely on energy harvested from their environment. Due to the unreliable nature of ambient energy sources, these devices experience frequent intervals of power loss, leading to sudden reboots. Tolerating such power supply disruptions require the ability to rapidly checkpoint/save system state when power loss is imminent and restore it at the start of the next power cycle to continue computations in a seamless manner. A typical microcontroller used in these systems consists of a fast nonvolatile SRAM and a nonvolatile Flash storage. Prior work has shown how emerging nonvolatile memory technologies such as STT-MRAM can improve the energy efficiency of these systems, either by using STT-MRAM as a drop-in replacement for Flash (henceforth referred to as the SRAM+STT-MRAM memory configuration) or using STT-MRAM as unified memory (henceforth referred to as the unified STT-MRAM memory configuration). However, both these configurations have significant drawbacks. Using the SRAM+STT-MRAM configuration leads to high checkpointing overhead due to the inefficient write operations of STT-MRAM whereas using the unified STT-MRAM configuration is inefficient due to executing every program instruction directly from STT-MRAM. This paper proposes a novel Spin Hall Effect-based nonvolatile-SRAM (SNVRAM) bit-cell that combines the nonvolatility of spin devices with the speed and energy efficiency of conventional 6T SRAM cells. We explore the use of the proposed SNVRAM to replace the SRAM in a transiently powered system to mitigate the drawbacks of the aforementioned memory configurations. Simulation results using a set of evaluation benchmarks demonstrate that the SNVRAM+STT-MRAM configuration leads to significant memory energy benefits of $2.6\times $ and $2.8\times $ on average, compared to the SRAM+STT-MRAM and unified STT-MRAM memory configurations, respectively.
Arnab Raha, Akhilesh Jaiswal 0001, Syed Shakib Sarwar, Hrishikesh Jayakumar, Vijay Raghunathan, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2018 Approximating Beyond the Processor: Exploring Full-System Energy-Accuracy Tradeoffs in a Smart Camera System
abstract
The intrinsic error resilience exhibited by emerging application domains enables new avenues for energy optimization of computing systems, namely, the introduction of a small amount of approximations during system operation in exchange for substantial energy savings. Prior work in the area of approximate computing has focused on individual subsystems of a computing system, for example, the computational subsystem or the memory subsystem. Since they focus only on individual subsystems, these techniques are unable to exploit the large energy-saving opportunities that stem from adopting a full-system perspective and approximating multiple subsystems of a computing platform simultaneously in a coordinated manner. This paper proposes a systematic methodology to perform joint approximations across different subsystems, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use the example of a smart camera system that executes various computer vision and image processing applications to illustrate how the sensing, memory, processing, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: 1) a compute-intensive smart camera system, AxSYScomp, where the error-resilient application executes locally within the camera and produces the final application output, and 2) a communication-intensive smart camera system, AxSYScomp, that sends the captured image to a remote cloud server, where the error-resilient application is executed and the final output is generated. We have implemented such an approximate smart camera system using an Altera Stratix IV GX FPGA development board, a Terasic TRDB-D5M 5-Megapixel camera module, a Terasic RFS WiFi module, and a 1-GB DDR3 dynamic random access memory small outline dual in-line memory module (SODIMM). Experimental results obtained using six application benchmarks demonstrate significant energy savings (around 7.5x for AxSYScomp and 4x on average for AxSYScomp) for minimal (<;1%) loss in application quality. Compared to approximating a single subsystem, the proposed full-system approximation methodology achieves additional energy benefits of 3.5x-5.5x (in the case of AxSYScomp) and 1.8x-3.7x (in the case of AxSYScomm) on average for a minimal (<;1%) application-level quality loss.
Arnab Raha, Vijay Raghunathan
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Towards Full-System Energy-Accuracy Tradeoffs: A Case Study of An Approximate Smart Camera System
abstract
The intrinsic error resilience exhibited by emerging application domains enables a new dimension for energy optimization of computing systems, namely the introduction of a controlled amount of approximations during system operation in exchange for substantial energy savings. Prior work in the area of approximate computing has focused on individual subsystems of a computing system, e.g., the computational subsystem or the memory subsystem. Since they focus only on individual subsystems, these techniques are unable to exploit the large energy-saving opportunities that stem from adopting a full-system perspective and approximating multiple subsystems of a computing platform simultaneously in a coordinated manner. This paper proposes a systematic methodology to perform joint approximations across different subsystems, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use the example of a smart camera system that executes various computer vision and image processing applications to illustrate how the sensing, memory, and processing subsystems can all be approximated synergistically. We have implemented such an approximate smart camera system using an Altera Stratix IV GX FPGA development board, a Terasic TRDB-D5M 5 Megapixel camera module, and a 1GB DDR3 SODIMM module. Experimental results obtained using six application benchmarks demonstrate significant energy savings (around 7.5X on average) for minimal (< 1%) loss in application quality. Compared to approximating a single subsystem, the proposed full-system approximation methodology achieves additional energy benefits of 3.5X - 5.5X on average for minimal (< 1%) quality loss.
Arnab Raha, Vijay Raghunathan
DAC2
2017 Spike timing dependent plasticity based enhanced self-learning for efficient pattern recognition in spiking neural networks
abstract
Spike Timing Dependent Plasticity (STDP), wherein synaptic weights are modified based on the temporal correlation between a pair of pre- and post-synaptic (post-neuronal) spikes, is widely used to implement unsupervised learning in Spiking Neural Networks (SNNs). In general, STDP-based learning models disregard the information embedded in post-neuronal spiking frequency. We observe that updating the synaptic weights at the instants of every post-neuronal spike while ignoring the spiking frequency could potentially cause them to learn overlapping representations of multiple input patterns sharing common features. We present STDP-based enhanced plasticity mechanisms that account for the spiking frequency to achieve efficient synaptic learning. First, we utilize low-pass filtered neuronal membrane potential to obtain an estimate of the spiking frequency. We perform STDP-driven weight updates in the event of a post-spike if the filtered potential exceeds a definite threshold. This ensures that plasticity is effected on the dominantly firing neuron that indicates a strong bias in learning the input pattern. Synaptic updates are restrained in the case of sporadic neuronal spiking activity, which implies a weak correlation with the input pattern. This enhances the quality of features encoded by the synapses, resulting in an improvement of 5.8% in the classification accuracy of an SNN of 100 neurons trained for digit recognition. Our simulations further show that the enhanced scheme provides a reduction of 2 χ in the number of weight updates, which leads to improved energy efficiency in event-driven SNN implementations. Second, we explore a neuronal spike-count based enhanced plasticity mechanism. The synapses are modified at the instant of a post-spike if the neuron had fired a certain number of spikes since the preceding update instant. This scheme performs delayed updates at suitable neuronal spiking instants to learn improved synaptic representations. Using this technique, the classification accuracy increased by 4% with 5.2× reduction in the number of weight updates.
Gopalakrishnan Srinivasan, Sourjya Roy, Vijay Raghunathan, Kaushik Roy 0001
IJCNN3
2017 AXSERBUS: A quality-configurable approximate serial bus for energy-efficient sensing
abstract
Mobile, wearable, and implantable devices integrate an increasing number and variety of sensors such as microphones, image sensors, and accelerometers. These devices spend substantial amounts of time reading the sensors within them, thereby incurring significant energy dissipation over off-chip serial interconnects. This paper proposes AXSERBUS, a quality-configurable approximate serial bus that exploits the locality of sensory data and the error resiliency of sensing applications to reduce energy dissipation. AXSERBUS significantly reduces signal transitions by encoding the differentials of sensory data in three encoding modes, depending on the magnitude of the differentials: very small differentials are zeroed out, incurring no energy dissipation; intermediate differentials are encoded using special low-transition count patterns; and for high differentials, the absolute value (not the differential) of the data is transmitted. Compared to previous schemes, the proposed multi-level encoding results in more data being encoded as low-energy patterns. In addition, in the intermediate differential encoding mode, the differentials are encoded in an approximate manner, and the approximation bounds are proportional to the magnitude of the differentials. Since small differentials are more frequent than large differentials in sensory data, the proposed encoding scheme also minimizes quality degradation. We demonstrate that AXSERBUS achieves improved energy vs. quality tradeoffs compared to previous schemes. In the context of an optical character recognition (OCR) application, AXSERBUS achieves 79.4% reduction in dynamic power dissipation, while maintaining accuracy above 95%.
Younghyun Kim 0001, Setareh Behroozi, Vijay Raghunathan, Anand Raghunathan
ISLPED3
2017 Approximate memory compression for energy-efficiency
abstract
Memory subsystems are a major energy bottleneck in computing platforms due to frequent transfers between processors and off-chip memory. We propose approximate memory compression, a technique that leverages the intrinsic resilience of emerging workloads such as machine learning and data analytics to reduce off-chip memory traffic and energy. To realize approximate memory compression, we enhance the memory controller to be aware of memory regions that contain approximation-resilient data, and to transparently compress/decompress the data written to/read from these regions. To provide control over approximations, the quality-aware memory controller conforms to a specified error constraint for each approximate memory region. We design a software interface that programmers can use to identify data structures that are resilient to approximations. We also propose a runtime quality control framework that automatically determines the error constraints for the identified data structures such that a given target application-level quality is maintained. We evaluate our proposal by implementing a hardware prototype using the Intel UniPHY-DDR3 memory controller and NIOS-II processor, a Hynix DDR3 DRAM module, and a Stratix-IV FPGA development board. Across a suite of 8 machine learning benchmarks, approximate memory compression obtains a 1.28× benefit in DRAM energy and a simultaneous 11.5% improvement in execution time for a small (<; 1.5%) loss in output quality.
Ashish Ranjan 0001, Arnab Raha, Vijay Raghunathan, Anand Raghunathan
ISLPED3
2017 Quality Configurable Approximate DRAM
abstract
Approximate computing is an emerging design paradigm that leverages the inherent error tolerance present in many applications to improve their power consumption and performance. Due to the forgiving nature of these error-resilient applications, precise input data is not always necessary for them to produce outputs of acceptable quality. This makes the memory subsystem (i.e., the place where data is stored), a suitable component for introducing approximations in return for substantial energy savings. Towards this end, this paper proposes a systematic methodology for constructing a quality configurable approximate DRAM system. Our design is based upon an extensive experimental characterization of memory errors as a function of the DRAM refresh-rate. Leveraging the insights gathered from this characterization, we propose four novel strategies for partitioning the DRAM in a system into a number of quality bins based on the frequency, location, and nature of bit errors in each of the physical pages, while also taking into account the property of variable retention time exhibited by DRAM cells. During data allocation, critical data is placed in the highest quality bin (that contains only accurate pages) and approximate data is allocated to bins sorted in descending order of quality, with the refresh rate serving as the quality control knob. We validate our proposed scheme on several error-resilient applications implemented using an Altera Stratix IV GX FPGA based Terasic TR4-230 development board containing a 1GB DDR3 DRAM module. Experimental results demonstrate a significant improvement in the energy-quality trade-off compared to previous work and show a reduction in DRAM refresh power of up to 73 percent on average with minimal loss in output quality.
Arnab Raha, Soubhagya Sutar, Hrishikesh Jayakumar, Vijay Raghunathan
IEEE Trans. Computers4
2017 Energy-Aware Memory Mapping for Hybrid FRAM-SRAM MCUs in Intermittently-Powered IoT Devices
abstract
Forecasts project that by 2020, there will be around 50 billion devices connected to the Internet of Things (IoT), most of which will operate untethered and unplugged. While environmental energy harvesting is a promising solution to power these IoT edge devices, it introduces new complexities due to the unreliable nature of ambient energy sources. In the presence of an unreliable power supply, frequent checkpointing of the system state becomes imperative, and recent research has proposed the concept of in-situ checkpointing by using ferroelectric RAM (FRAM), an emerging non-volatile memory technology, as unified memory in these systems. Even though an entirely FRAM-based solution provides reliability, it is energy inefficient compared to SRAM due to the higher access latency of FRAM. On the other hand, an entirely SRAM-based solution is highly energy efficient but is unreliable in the face of power loss. This paper advocates an intermediate approach in hybrid FRAM-SRAM microcontrollers that involves judicious memory mapping of program sections to retain the reliability benefits provided by FRAM while performing almost as efficiently as an SRAM-based system. We propose an energy-aware memory mapping technique that maps different program sections to the hybrid FRAM-SRAM microcontroller such that energy consumption is minimized without sacrificing reliability. Our technique consists of eM-map , which performs a one-time characterization to find the optimal memory map for the functions that constitute a program and energy-align , a novel hardware-software technique that aligns the system’s powered-on time intervals to function execution boundaries, which results in further improvements in energy efficiency and performance. Experimental results obtained using the MSP430FR5739 microcontroller demonstrate a significant performance improvement of up to 2x and energy reduction of up to 20% over a state-of-the-art FRAM-based solution. Finally, we present a case study that shows the implementation of our techniques in the context of a real IoT application.
Hrishikesh Jayakumar, Arnab Raha, Jacob R. Stevens, Vijay Raghunathan
ACM Trans. Embed. Comput. Syst.4
2017 qLUT: Input-Aware Quantized Table Lookup for Energy-Efficient Approximate Accelerators
abstract
Approximate computing has emerged as a popular design paradigm for optimizing the performance and energy consumption of error-resilient applications in domains such as machine learning, graphics, data analytics, etc . Numerous techniques for approximate computing have been proposed at different layers of the system stack, from circuits to architecture to software. In this work, we propose a new technique, called quantized table lookup , for approximating the meta-functions used in the core computational kernels of error-resilient applications. In contrast to prior work that directly approximates the functionality of the meta-functions, the proposed technique instead approximates the input data to the meta-functions by reducing/quantizing them to a much smaller set of values that we call quantized inputs . The small number of quantized inputs enables us to completely replace the energy-intensive arithmetic units in the meta-function with small and energy-efficient lookup tables (called quantized lookup tables or q LUT) that contain precomputed output values corresponding to the quantized inputs. The proposed approximation technique is not only highly generic, but also inherently quality-configurable and input-aware. Quality-configurability and input-awareness are achieved by modulating the size of the q LUT as well as selecting the values of the quantized inputs judiciously based on the statistics of the original input data. To evaluate the proposed technique, we have implemented the dominant meta-functions of nine error-resilient application benchmarks as quantized table lookup based hardware accelerators using 45nm technology. Experimental results demonstrate average energy savings of 46% at the application-level for minimal (<1%) loss in output quality.
Arnab Raha, Vijay Raghunathan
ACM Trans. Embed. Comput. Syst.2
2017 Energy-Efficient Reduce-and-Rank Using Input-Adaptive Approximations
abstract
Approximate computing is an emerging design paradigm that exploits the intrinsic ability of applications to produce acceptable outputs even when their computations are executed approximately. In this paper, we explore approximate computing for a key computation pattern, reduce-andrank (RnR), which is prevalent in a wide range of workloads, including video processing, recognition, search, and data mining. An RnR kernel performs a reduction operation (e.g., distance computation, dot product, and L1-norm) between an input vector and each of a set of reference vectors, and ranks the reduction outputs to select the top reference vectors for the current input. We propose three complementary approximation strategies for the RnR computation pattern. The first is interleaved reductionand-ranking, wherein the vector reductions are decomposed into multiple partial reductions and interleaved with the rank computation. Leveraging this transformation, we propose the use of intermediate reduction results and ranks to identify future computations that are likely to have a low impact on the output, and can, hence, be approximated. The second strategy, inputsimilarity-based approximation, exploits the spatial or temporal correlation of inputs (e.g., pixels of an image or frames of a video) to identify computations that are amenable to approximation. The third strategy, reference vector reordering, rearranges the order in which the reference vectors are processed such that vectors that are relatively more critical in evaluating the correct output, are processed at the beginning of RnR operation. The number of these critical reference vectors is usually small, which renders a substantial portion of the total computation to be amenable to approximation. These strategies address a key challenge in approximate computing-identification of which computations to approximate-and may be used to drive any approximation mechanism, such as computation skipping or precision scaling to realize performance and energy improvements. A second key challenge in approximate computing is that the extent to which computations can be approximated varies significantly from application to application, and across inputs for even a single application. Hence, input-adaptive approximation, or the ability to automatically modulate the degree of approximation based on the nature of each individual input, is essential for obtaining optimal energy savings. In addition, to enable quality configurability in RnR kernels, we propose a kernel-level quality metric that correlates well to application-level quality, and identify key parameters that can be used to tune the proposed approximation strategies dynamically. We develop a runtime framework that modulates the identified parameters during the execution of RnR kernels to minimize their energy while meeting a given target quality. To evaluate the proposed concepts, we designed quality-configurable hardware implementations of six RnR-based applications from the recognition, mining, search, and video processing application domains in 45-nm technology. Our experiments demonstrate a 1.13×-3.18× reduction in energy consumption with virtually no loss in output quality (<;0.5%) at the application level. The energy benefits further improve up to 3.43× and 3.9× when the quality constraints are relaxed to 2.5% and 5%, respectively.
Arnab Raha, Swagath Venkataramani, Vijay Raghunathan, Anand Raghunathan
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Energy-efficient system design for IoT devices
abstract
It is projected that, within the coming decade, there will be more than 50 billion smart objects connected to the Internet of Things (IoT). These smart objects, which connect the physical world with the world of computing infrastructure, are expected to pervade all aspects of our daily lives and revolutionize a number of application domains such as healthcare, energy conservation, transportation, etc. In this paper, we present an overview of the challenges involved in designing energy-efficient IoT edge devices and describe recent research that has proposed promising solutions to address these challenges. First, we outline the challenges involved in efficiently supplying power to an IoT device. Next, we discuss the role of emerging memory technologies in making IoT devices energy-efficient. Finally, we discuss the potential impact that approximate computing can have in increasing the energy-efficiency of wearables and other compute-intensive IoT devices.
Hrishikesh Jayakumar, Arnab Raha, Younghyun Kim 0001, Soubhagya Sutar, Woo Suk Lee, Vijay Raghunathan
ASP-DAC6
2016 D-PUF: an intrinsically reconfigurable DRAM PUF for device authentication in embedded systems
abstract
Physically Unclonable Functions (PUFs) have proved to be effective and low-cost measure against counterfeiting by providing device authentication and secure key storage services. Memory based PUF implementations are an attractive option due to the ubiquitous nature of memory in electronic devices and the requirement of minimal (or no) additional circuitry. DRAM based PUFs are particularly advantageous due to their large address space and multiple controllable parameters during response generation. However, prior works on DRAM PUFs use a static response generation mechanism making them vulnerable to security attacks. Further, they result in very slow device authentication, are not applicable to off-the-shelf DRAM modules, or require DRAM power cycling prior to authentication.
Soubhagya Sutar, Arnab Raha, Vijay Raghunathan
CASES3
2016 TeleProbe: Zero-power Contactless Probing for Implantable Medical Devices
abstract
The lack of post-deployment visibility into system behavior is one of the major challenges in ensuring the reliable operation of implantable medical devices (IMDs). While wireless connectivity is becoming common in IMDs for monitoring device status, conventional wireless links incur significant energy overheads for data acquisition, processing, and active radio transmission. While low-power transceivers have been introduced to reduce the energy consumed by the radio itself, the energy consumed by the microcontroller for processing data and controlling the radio has often been overlooked. As a result, in IMDs that have a stringent energy constraint, prolonged signal monitoring over a wireless channel is infeasible due to this prohibitively high power consumption.
Woo Suk Lee, Younghyun Kim 0001, Vijay Raghunathan
ISLPED3
2016 Sleep-Mode Voltage Scaling: Enabling SRAM Data Retention at Ultra-Low Power in Embedded Microcontrollers
Hrishikesh Jayakumar, Arnab Raha, Vijay Raghunathan
ACM Trans. Embed. Comput. Syst.3
2016 CO-GPS: Energy Efficient GPS Sensing with Cloud Offloading
abstract
Location is a fundamental service for mobile computing. Typical GPS receivers, although widely available for navigation purposes, may consume too much energy to be useful for many applications. Observing that in many sensing scenarios, the location information can be post-processed when the data is uploaded to a server, we design a cloud-offloaded GPS (CO-GPS) solution that allows a sensing device to aggressively duty-cycle its GPS receiver and log just enough raw GPS signal for post-processing. Leveraging publicly available information such as GNSS satellite ephemeris and an Earth elevation database, a cloud service can derive good quality GPS locations from a few milliseconds of raw data. Using our design of a portable sensing device platform called CLEON, we evaluate the accuracy and efficiency of the solution. Compared to more than 30 seconds of heavy signal processing on standalone GPS receivers, we can achieve three orders of magnitude lower energy consumption per location tagging.
Jie Liu 0001, Bodhi Priyantha, Ted Hart, Yuzhe Jin, Woo Suk Lee, Vijay Raghunathan, Heitor S. Ramos, Qiang Wang 0001
IEEE Trans. Mob. Comput.6
2016 Input-Based Dynamic Reconfiguration of Approximate Arithmetic Units for Video Encoding
abstract
The field of approximate computing has received significant attention from the research community in the past few years, especially in the context of various signal processing applications. Image and video compression algorithms, such as JPEG, MPEG, and so on, are particularly attractive candidates for approximate computing, since they are tolerant of computing imprecision due to human imperceptibility, which can be exploited to realize highly power-efficient implementations of these algorithms. However, existing approximate architectures typically fix the level of hardware approximation statically and are not adaptive to input data. For example, if a fixed approximate hardware configuration is used for an MPEG encoder (i.e., a fixed level of approximation), the output quality varies greatly for different input videos. This paper addresses this issue by proposing a reconfigurable approximate architecture for MPEG encoders that optimizes power consumption with the goal of maintaining a particular Peak Signal-to-Noise Ratio (PSNR) threshold for any video. Toward this end, we design reconfigurable adder/subtractor blocks (RABs), which have the ability to modulate their degree of approximation, and subsequently integrate these blocks in the motion estimation and discrete cosine transform modules of the MPEG encoder. We propose two heuristics for automatically tuning the approximation degree of the RABs in these two modules during runtime based on the characteristics of each individual video. Experimental results show that our approach of dynamically adjusting the degree of hardware approximation based on the input video respects the given quality bound (PSNR degradation of 1%-10%) across different videos while achieving a power saving up to 38% over a conventional nonapproximated MPEG encoder architecture. Note that although the proposed reconfigurable approximate architecture is presented for the specific case of an MPEG encoder, it can be easily extended to other DSP applications.
Arnab Raha, Hrishikesh Jayakumar, Vijay Raghunathan
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Quality-aware data allocation in approximate DRAM?
abstract
Approximate computing is an emerging design paradigm that leverages the inherent error tolerance present in many applications to optimize their power consumption and performance. Due to the forgiving nature of these error-resilient applications, highly precise input data is not always necessary for them to produce outputs of acceptable quality. This makes memory, the place where data is stored, a suitable component for introducing errors or approximations in return for considerable energy savings. Towards this end, this paper proposes, for the first time, a systematic way for constructing a quality-aware approximate DRAM system. Our design is based upon an extensive experimental characterization of memory errors as a function of the DRAM refresh rate. Leveraging the insights gathered from this characterization, we propose four novel strategies for partitioning the DRAM into a number of quality bins based on the frequency, location, and nature of bit errors in each of the physical pages. During allocation, critical data is placed in the highest quality bin containing only accurate pages and approximate data is allocated to bins sorted in descending order of quality. We validate our proposed scheme on several error-resilient applications implemented using an Altera Stratix IV GX FPGA based Terasic TR4-230 development board containing a 1GB DDR3 DRAM module. Experimental results demonstrate a significant improvement in the energy-quality trade-off compared to previous work and show a reduction in DRAM refresh power of up to 73% with minimal loss in output quality.
Arnab Raha, Hrishikesh Jayakumar, Soubhagya Sutar, Vijay Raghunathan
CASES4
2015 Vibration-based secure side channel for medical devices
abstract
Implantable and wearable medical devices are used for monitoring, diagnosis, and treatment of an ever-increasing range of medical conditions, leading to an improved quality of life for patients. The addition of wireless connectivity to medical devices has enabled post-deployment tuning of therapy and access to device data virtually anytime and anywhere but, at the same time, has led to the emergence of security attacks as a critical concern. While cryptography and secure communication protocols may be used to address most known attacks, the lack of a viable secure connection establishment and key exchange mechanism is a fundamental challenge that needs to be addressed. We propose a vibration-based secure side channel between an external device (medical programmer or smartphone) and a medical device. Vibration is an intrinsically short-range, user-perceptible channel that is suitable for realizing physically secure communication at low energy and size/weight overheads. We identify and address key challenges associated with the vibration channel, and propose a vibration-based wakeup and key exchange scheme, named SecureVibe, that is resistant to battery drain attacks. We analyze the risk of acoustic eavesdropping attacks and propose an acoustic masking countermeasure. We demonstrate and evaluate vibration-based wakeup and key exchange between a smartphone and a prototype medical device in the context of a realistic human body model.
Younghyun Kim 0001, Woo Suk Lee, Vijay Raghunathan, Niraj K. Jha, Anand Raghunathan
DAC3
2015 Quality configurable reduce-and-rank for energy efficient approximate computing
Arnab Raha, Swagath Venkataramani, Vijay Raghunathan, Anand Raghunathan
DATE3
2015 Message from the program chairs
abstract
It is our great pleasure to welcome you to the 2015 ACM/IEEE International Symposium on Low Power Electronics and Design - ISLPED'15, in the “eternal city” of Rome, Italy. This year's symposium continues its two decade long tradition of being the premier forum for presentation of research results and industrial experience reports on leading-edge issues in low power design. ISLPED has always been unique in the sense that it brings together researchers and practitioners interested in various aspects of low power design at a single venue and provides them an opportunity to share their perspectives with each other.
Ruchir Puri, Vijay Raghunathan
ISLPED2
2015 VIDalizer: An energy efficient video streamer
abstract
Recent years have witnessed a significant rise in the number, duration and variety of video contents, which contribute to the bulk of internet traffic. With increase in smartphone and tablet users, watching videos on mobile devices has become one of its most popular use cases. These devices live on limited battery energy which is still a major bottleneck and a source of user dissatisfaction during video playback. In this paper we introduce an intermediate framework called VIDalizer for power efficient video delivery to smartphones and tablets. This almost transparent to the user, battery aware framework takes away some of the video processing overhead from the device and intelligently tunes its parameters customized for the mobile device while delivering the video using a novel transport protocol. Our preliminary results show that this framework can significantly reduce energy consumption up to 45%–55% of a mobile device without compromising user experience.
Arnab Raha, Subrata Mitra, Vijay Raghunathan, Sanjay G. Rao
WCNC3
2015 QuickRecall: A HW/SW Approach for Computing across Power Cycles in Transiently Powered Computers
abstract
Transiently Powered Computers (TPCs) are a new class of batteryless embedded systems that depend solely on energy harvested from external sources for performing computations. Enabling long-running computations on TPCs is a major challenge due to the highly intermittent nature of the power supply (often bursts of < 100ms), resulting in frequent system reboots. Prior work seeks to address this issue by frequently checkpointing system state in flash memory, preserving it across power cycles. However, this involves a substantial overhead due to the high erase/write times of flash memory. This article proposes the use of Ferroelectric RAM (FRAM), an emerging nonvolatile memory technology that combines the benefits of SRAM and flash, to seamlessly enable long-running computations in TPCs. We propose a lightweight, in-situ checkpointing technique for TPCs using FRAM that consumes only 30 nJ while decreasing the time taken for saving and restoring a checkpoint to only 21.06μ s , which is over two orders of magnitude lower than the corresponding overhead using flash. We have implemented and evaluated our technique, Q uick R ecall , using the TI MSP430FR5739 FRAM-enabled microcontroller. Experimental results show that our highly-efficient checkpointing translate to significant speedup (1.25x - 8.4x) in program execution time and reduction (∼3x) in application-level energy consumption.
Hrishikesh Jayakumar, Arnab Raha, Woo Suk Lee, Vijay Raghunathan
ACM J. Emerg. Technol. Comput. Syst.4
2015 iTCP: an intelligent TCP with neural network based end-to-end congestion control for ad-hoc multi-hop wireless mesh networks
A. B. M. Alim Al Islam, Vijay Raghunathan
Wirel. Networks2
2015 SymCo: Symbiotic Coexistence of Single-hop and Multi-hop Transmissions in Next-generation Wireless Mesh Networks
A. B. M. Alim Al Islam, Vijay Raghunathan
Wirel. Networks2
2015 SiAc: simultaneous activation of heterogeneous radios in high data rate multi-hop wireless networks
A. B. M. Alim Al Islam, Vijay Raghunathan
Wirel. Networks2
2014 Powering the internet of things
abstract
Various industry forecasts project that, by 2020, there will be around 50 billion devices connected to the Internet of Things (IoT), helping to engineer new solutions to societal-scale problems such as healthcare, energy conservation, transportation, etc. Most of these devices will be wireless due to the expense, inconvenience, or in some cases, the sheer infeasibility of wiring them. Further, many of them will have stringent size constraints. With no cord for power and limited space for a battery, powering these devices (to achieve several months to possibly years of unattended operation) becomes a daunting challenge. This paper highlights some promising directions for addressing this challenge, focusing on three main building blocks: (a) the design of ultra-low power hardware platforms that integrate computing, sensing, storage, and wireless connectivity in a tiny form factor, (b) the development of intelligent system-level power management techniques, and (c) the use of environmental energy harvesting to make IoT devices self-powered, thus decreasing -- in some cases, even eliminating -- their dependence on batteries. We discuss these building blocks in detail and illustrate case-studies of systems that use them judiciously, including the QUBE wireless embedded platform, which exploits the characteristics of emerging non-volatile memory technologies to seamlessly and efficiently enable long-running computations in systems that experience frequent power loss (i.e., intermittently powered systems).
Hrishikesh Jayakumar, Kangwoo Lee, Woo Suk Lee, Arnab Raha, Younghyun Kim 0001, Vijay Raghunathan
ISLPED6
2012 Modeling, design and cross-layer optimization of polysilicon solar cell based micro-scale energy harvesting systems
abstract
This paper presents modeling, design, and cross-layer optimization of polysilicon solar cell based micro-scale energy harvesting systems. The proposed design methodology is suitable for energy harvesting systems employed in low-cost applications, such as wireless sensor networks. Our approach is unique in achieving maximum output power by cross-layer optimization of poly-silicon solar cells (thickness, grain boundaries and cell configuration) and power converter circuits. Simulation results indicate that optimizing the solar cell along with the power converter improves the system output power by 16% compared to a baseline approach of optimizing the two components separately.
Elif S. Mungan, Chao Lu 0005, Vijay Raghunathan, Kaushik Roy 0001
ISLPED3
2012 Multi-armed Bandit Congestion Control in Multi-hop Infrastructure Wireless Mesh Networks
abstract
Congestion control in multi-hop infrastructure wireless mesh networks is both an important and a unique problem. It is unique because it has two prominent causes of failed transmissions which are difficult to tease apart - lossy nature of wireless medium and high extent of congestion around gateways in the network. The concurrent presence of these two causes limits applicability of already available congestion control mechanisms, proposed for wireless networks. Prior mechanisms mainly focus on the former cause, ignoring the latter one. Therefore, we address this issue to design an end-to-end congestion control mechanism for infrastructure wireless mesh networks in this paper. We formulate the congestion control problem and map that to the restless multi-armed bandit problem, a well-known decision problem in the literature. Then, we propose three myopic policies to achieve a near-optimal solution for the mapped problem since no optimal solution is known to this problem. We perform comparative evaluation through ns-2 simulation and a real testbed experiment with a wireline TCP variant and a wireless TCP protocol. The evaluation reveals that our proposed mechanism can achieve up to 52% increased network throughput and 34% decreased average energy consumption per transmitted bit in comparison to the other end-to-end congestion control variants.
A. B. M. Alim Al Islam, S. M. Iftekharul Alam, Vijay Raghunathan, Saurabh Bagchi
MASCOTS3
2012 A Cross-Layer Analytical Model to Estimate the Capacity of a WiMAX Network
abstract
Specialized physical and medium access control layer operations of WiMAX introduce a challenging problem to formulate a precise cross-layer analytical model for its capacity estimation. Although the cross-layer formulation is yet to be attempted in the literature, it can effectively serve as a fast and cost-effective tool to facilitate future deployment planning, efficient network maintenance, etc., and thus can expedite to cope up with the recent blast in WiMAX deployment. Therefore, we propose a novel formulation of cross-layer analytical model for WiMAX capacity estimation in this paper. Our formulation addresses intricate interrelationships between different layers in a protocol stack. Besides, the formulation separately addresses both reliable and unreliable transport layer transmissions. We verify our formulated model using ns-2 simulation, which reveals that the model can estimate the capacity of a WiMAX network with as low as 1% average error having a standard deviation of only 1%. Consequently, we find that the proposed model can achieve up to a 97% decrease in the average error compared to other available analytical models in the literature. Finally, we present different applications of the proposed model by utilizing it in network planning and protocol overhead analysis.
A. B. M. Alim Al Islam, Vijay Raghunathan
MASCOTS2
2011 Stage number optimization for switched capacitor power converters in micro-scale energy harvesting
abstract
Micro-scale energy harvesting has become an increasingly viable and promising option for powering ultra-low power systems. A power converter is a key component in micro-scale energy harvesting systems. Various design parameters of the power converter, most notably the number of stages in a multi-stage power converter, play a crucial role in determining the amount of electrical power that can be extracted from a micro-scale energy transducer such as a miniature solar cell. Existing stage number optimization techniques for switched capacitor power converters, when used for energy harvesting systems, result in a substantial degradation in the amount of harvested electrical power. To address this problem, this paper proposes a new stage number optimization technique for switched capacitor power converters that maximizes the net harvested power in micro-scale energy harvesting systems. The proposed technique is based on a new figure-of-merit that is well suited for energy-harvesting systems. We have validated the proposed technique through circuit simulations using IBM 65nm technology. Our simulation results demonstrate that the proposed stage number optimization technique results in an increase of 60%-290% in net harvested power, compared to existing stage number optimization techniques.
Chao Lu 0005, Sang Phill Park, Vijay Raghunathan, Kaushik Roy 0001
DATE3
2011 Backpacking: Deployment of Heterogeneous Radios in High Data Rate Sensor Networks
abstract
The early success of wireless sensor networks has led to a new generation of increasingly sophisticated sensor network applications, such as HP's CeNSE. These applications demand high network throughput that easily exceeds the capability of low-power 802.15.4 radios that are most commonly used in today's sensor nodes. To address this issue, this paper investigates an energy-efficient approach to supplementing an 802.15.4 based sensor network with high bandwidth, high power, longer range radios such as 802.11. Exploiting a key observation that the high bandwidth radio achieves low energy consumption per transmitted bit of data due to its inherent transmission efficiency, we propose a hybrid network architecture that utilizes an optimal density of dual-radio (802.15.4 and 802.11) nodes to augment a sensor network having only 802.15.4 radios. We present a cross-layer mathematical model to calculate this optimal density, which strikes a balance between the low energy per bit of the high-bandwidth radio and the low sleep power of 802.15.4 radio. Experimental results obtained using a wireless testbed reveal that our architecture improves the average energy per bit, the time elapsed before half of the nodes drain their battery, and the end-to-end delay by 62%, 106%, and 73% respectively, compared to a network that uses only 802.15.4 radios.
A. B. M. Alim Al Islam, Mohammad Sajjad Hossain, Vijay Raghunathan, Y. Charlie Hu
ICCCN3
2011 μSETL: A set based programming abstraction for wireless sensor networks
Mohammad Sajjad Hossain, A. B. M. Alim Al Islam, Milind Kulkarni 0001, Vijay Raghunathan
IPSN4
2011 Aveksha: a hardware-software approach for non-intrusive tracing and profiling of wireless embedded systems
abstract
It is important to get an idea of the events occurring in an embedded wireless node when it is deployed in the field, away from the convenience of an interactive debugger. Such visibility can be useful for post-deployment testing, replay-based debugging, and for performance and energy profiling of various software components. Prior software-based solutions to address this problem have incurred high execution overhead and intrusiveness. The intrusiveness changes the intrinsic timing behavior of the application, thereby reducing the fidelity of the collected profile. Prior hardware-based solutions have involved the use of dedicated ASICs or other tightly coupled changes to the embedded node's processor, which significantly limits their applicability.
Matthew Tan Creti, Mohammad Sajjad Hossain, Saurabh Bagchi, Vijay Raghunathan
SenSys4
2011 End-to-end congestion control in wireless mesh networks using a neural network
abstract
Maintaining the performance of reliable transport protocols, such as TCP, over wireless mesh networks is a challenging problem due to the unique characteristics of wireless mesh networks such as the lossy nature of the communication medium, absence of a base station, similarity in traffic pattern experienced by neighboring mesh nodes, etc. One of the reasons for the poor performance of conventional TCP variants over wireless mesh networks is that the congestion control mechanisms in conventional TCP variants do not explicitly account for these unique characteristics. To address this problem, this paper proposes a novel neural network based congestion control technique for reliable data transfer over wireless mesh networks. We analyze the proposed congestion control technique in detail and incorporate it into TCP to create a variant that we name intelligent TCP or iTCP. We evaluate the performance of iTCP using ns-2 simulations. Our results demonstrate that our proposed congestion control technique exhibits a significant improvement in total network throughput and average energy consumption per bit compared to congestion control techniques used in other variants of TCP.
A. B. M. Alim Al Islam, Vijay Raghunathan
WCNC2
2010 Micro-scale energy harvesting: a system design perspective
abstract
Harvesting electrical power from environmental energy sources is an attractive and increasingly feasible option for several micro-scale electronic systems such as biomedical implants and wireless sensor nodes that need to operate autonomously for long periods of time (months to years). However, designing highly efficient micro-scale energy harvesting systems requires an in-depth understanding of various design considerations and tradeoffs. This paper provides an overview of the area of micro-scale energy harvesting and discusses the various challenges and considerations involved from a system-design perspective.
Chao Lu 0005, Vijay Raghunathan, Kaushik Roy 0001
ASP-DAC2
2010 Efficient power conversion for ultra low voltage micro scale energy transducers
abstract
Energy harvesting has emerged as a feasible and attractive option to improve battery lifetime in micro-scale electronic systems such as biomedical implants and wireless sensor nodes. A key challenge in designing micro-scale energy harvesting systems is that miniature energy transducers (e.g., photovoltaic cells, thermo-electric generators, and fuel cells) output very low voltages (0-0.4V). Therefore, a fully on-chip power converter (usually based on a charge pump) is used to boost the output voltage of the energy transducer and transfer charge into an energy buffer for storage. However, the charge transfer capability of widely used linear charge pump based power converters degrades when used with ultra-low voltage energy transducers. This paper presents the design of a new tree topology charge pump that has a reduced charge sharing time, leading to an improved charge transfer capability. The proposed design has been implemented using 65nm technology and circuit simulations demonstrate that the proposed design results in an increase of up to 30% in harvested power compared to existing linear charge pumps.
Chao Lu 0005, Sang Phill Park, Vijay Raghunathan, Kaushik Roy 0001
DATE3
2010 AEGIS: A Lightweight Firewall for Wireless Sensor Networks
Mohammad Sajjad Hossain, Vijay Raghunathan
DCOSS2
2010 AEGIS: a rule based framework for traffic gatekeeping in wireless sensor networks
abstract
Traffic gatekeeping, typically implemented using firewalls, is an essential functionality in today's networked computing systems such as desktop PCs. Traffic gatekeeping provides protection against a variety of over-the-network security attacks and can also be used to implement various communication resource management policies. With the development of technologies such as IPv6 and 6LoWPAN that pave the way for Internet-connected embedded systems and sensor networks, traffic gatekeeping will soon become a must-have for these systems as well. Towards this, we present Aegis, a lightweight, rule-based traffic gatekeeping framework for sensor networks. We present an overview of the Aegis architecture, its implementation, and two usage scenarios.
Mohammad Sajjad Hossain, Vijay Raghunathan
IPSN2
2010 Maximum power point considerations in micro-scale solar energy harvesting systems
abstract
Maximum power point (MPP) tracking is a technique to maximize the amount of power harvested from energy transducers such as solar cells. MPP tracking presents new design challenges when used in the context of micro-scale energy harvesting systems, where the area dedicated to solar cells is small (in the range of sub or a few cm2) and hence, the power output is in the range of a few mW. This paper provides an overview of several low-overhead MPP tracking approaches that are attractive for micro-scale solar energy harvesting. These include: design-time component matching method, fractional open-circuit voltage or fractional short-circuit current method, and variants of the generic hill-climbing approach. We also illustrate using a simple case study, how MPP from a full-system perspective may differ from the MPP of the photovoltaic module itself in micro-scale harvesting systems.
Chao Lu 0005, Vijay Raghunathan, Kaushik Roy 0001
ISCAS2
2010 Analysis and design of ultra low power thermoelectric energy harvesting systems
abstract
Thermal energy harvesting using micro-scale thermoelectric generators is a promising approach to alleviate the power supply challenge in ultra low power systems. In thermal energy harvesting systems, energy is extracted from the transducer using an interface circuitry, which plays a key role in the determining the energy extraction efficiency. This paper presents techniques for the systematic modeling, analysis, and design of interface circuitry used in micro-scale thermoelectric energy harvesting systems. We characterize the electrical behavior of a micro-scale thermoelectric transducer connected to a step-up charge pump based power converter and model the relationship between the transducer output voltage and the charge pump switching frequency. We model various power loss components inside the interface circuitry and present an analytical design methodology that estimates optimal parameter values for the interface circuitry. These parameter values lead to maximum net output power being delivered to the energy buffer. We have implemented various interface circuitries using IBM 65nm technology to verify our proposed models and methodology. Circuit simulation results show that the proposed methodology accurately estimates the maximum power point voltage of the system with an error of 3%.
Chao Lu 0005, Sang Phill Park, Vijay Raghunathan, Kaushik Roy 0001
ISLPED3
2009 Green at the micro-scale: towards self-powered embedded systems
abstract
The field of green computing, which has gained tremendous importance and popularity over the past few years, has mainly focused on large computing systems such as server farms. This talk makes the case that it is equally important for green computing to be applied at the micro-scale, targeting the billions of embedded devices found everywhere around us. A promising start in this direction is to power embedded systems using environmental (renewable) energy sources such as solar, wind, vibration, etc., which raises the possibility of self-sustaining embedded systems. This tutorial highlights the challenges involved in architecting such self-sustaining embedded systems and discusses various hardware and software techniques for their efficient design and operation.
Vijay Raghunathan
ISLPED1
2008 HERMES: A Software Architecture for Visibility and Control in Wireless Sensor Network Deployments
abstract
Designing reliable software for sensor networks is challenging because application developers have little visibility into, and understanding of the post-deployment behavior of code executing on resource constrained nodes in remote and ill-reproducible environments. To address this problem, this paper presents HERMES, a lightweight framework and prototype tool that provides fine-grained visibility and control of a sensor node's software at run-time. HERMES's architecture is based on the notion of interposition, which enables it to provide these properties in a minimally intrusive manner, without requiring any modification to software applications being observed and controlled. HERMES provides a general, extensible, and easy-to-use framework for specifying which software components to observe and control as well as when and how this observation and control is done. We have implemented and tested a fully functional prototype of HERMES for the SOS sensor operating system. Our performance evaluation, using real sensor nodes as well as cycle-accurate simulation, shows that HERMES successfully achieves its objective of providing fine-grained and dynamic visibility and control without incurring significant resource overheads. We demonstrate the utility and flexibility of HERMES by using our prototype to design, implement, and evaluate three case-studies: debugging and testing deployed sensor network applications, performing transparent software updates in sensor nodes, and implementing network traffic shaping and resource policing.
Nupur Kothari, Kiran Nagaraja, Vijay Raghunathan, Florin Sultan, Srimat T. Chakradhar
IPSN3
2006 Harvesting aware power management for sensor networks
abstract
Energy harvesting offers a promising alternative to solve the sustainability limitations arising from battery size constraints in sensor networks. Several considerations in using an environmental energy source are fundamentally different from using batteries. Rather than a limit on the total energy, harvesting transducers impose a limit on the instantaneous power available. Further, environmental energy availability is often highly variable and a deterministic metric such as residual battery capacity is not available to characterize the energy source. The different nodes in a sensor network may also have different energy harvesting opportunities. Since the same end-user performance may be achieved using different workload allocations at multiple nodes, it is important to adapt the workload allocation to the spatio-temporal energy availability profile in order to enable energy-neutral operation of the network. This paper describes power management techniques for such energy harvesting sensor networks. Platform design considerations as well as power scaling techniques at the node-level and network-level are described.
Aman Kansal, Jason Hsu, Mani Srivastava 0001, Vijay Raghunathan
DAC4
2006 Adaptive duty cycling for energy harvesting systems
abstract
Harvesting energy from the environment is feasible in many applications to ameliorate the energy limitations in sensor networks. In this paper, we present an adaptive duty cycling algorithm that allows energy harvesting sensor nodes to autonomously adjust their duty cycle according to the energy availability in the environment. The algorithm has three objectives, namely (a) achieving energy neutral operation, i.e., energy consumption should not be more than the energy provided by the environment, (b) maximizing the system performance based on an application utility model subject to the above energy-neutrality constraint, and (c) adapting to the dynamics of the energy source at run-time. We present a model that enables harvesting sensor nodes to predict future energy opportunities based on historical data. We also derive an upper bound on the maximum achievable performance assuming perfect knowledge about the future behavior of the energy source. Our methods are evaluated using data gathered from a prototype solar energy harvesting platform and we show that our algorithm can utilize up to 58% more environmental energy compared to the case when harvesting-aware power management is not used.
Jason Hsu, Sadaf Zahedi, Aman Kansal, Mani Srivastava 0001, Vijay Raghunathan
ISLPED5
2006 Design and power management of energy harvesting embedded systems
abstract
Harvesting energy from the environment is a desirable and increasingly important capability in several emerging applications of embedded systems such as sensor networks, biomedical implants, etc. While energy harvesting has the potential to enable near-perpetual system operation, designing an efficient energy harvesting system that actually realizes this potential requires an in-depth understanding of several complex tradeoffs. These tradeoffs arise due to the interaction of numerous factors such as the characteristics of the harvesting transducers, chemistry and capacity of the batteries used (if any), power supply requirements and power management features of the embedded system, application behavior, etc. This paper surveys the various issues and tradeoffs involved in designing and operating energy harvesting embedded systems. System design techniques are described that target high conversion and storage efficiency by extracting the most energy from the environment and making it maximally available for consumption. Harvesting aware power management techniques are also described, which reconcile the very different spatio-temporal characteristics of energy availability and energy usage within a system and across a network.
Vijay Raghunathan, Pai H. Chou
ISLPED1
2005 Design considerations for solar energy harvesting wireless embedded systems
abstract
Sustainable operation of battery powered wireless embedded systems (such as sensor nodes) is a key challenge, and considerable research effort has been devoted to energy optimization of such systems. Environmental energy harvesting, in particular solar based, has emerged as a viable technique to supplement battery supplies. However, designing an efficient solar harvesting system to realize the potential benefits of energy harvesting requires an in-depth understanding of several factors. For example, solar energy supply is highly time varying and may not always be sufficient to power the embedded system. Harvesting components, such as solar panels, and energy storage elements, such as batteries or ultracapacitors, have different voltage-current characteristics, which must be matched to each other as well as the energy requirements of the system to maximize harvesting efficiency. Further, battery non-idealities, such as self-discharge and round trip efficiency, directly affect energy usage and storage decisions. The ability of the system to modulate its power consumption by selectively deactivating its sub-components also impacts the overall power management architecture. This paper describes key issues and tradeoffs which arise in the design of solar energy harvesting, wireless embedded systems and presents the design, implementation, and performance evaluation of Heliomote, our prototype that addresses several of these issues. Experimental results demonstrate that Heliomote, which behaves as a plug-in to the Berkeley/Crossbow motes and autonomously manages energy harvesting and storage, enables near-perpetual, harvesting aware operation of the sensor node.
Vijay Raghunathan, Aman Kansal, Jason Hsu, Jonathan Friedman, Mani Srivastava 0001
IPSN1
2005 Heliomote: enabling long-lived sensor networks through solar energy harvesting
abstract
No abstract available.
Kris Lin, Jennifer Yu, Jason Hsu, Sadaf Zahedi, Jonathan Friedman, Aman Kansal, Vijay Raghunathan, Mani Srivastava 0001
SenSys8
2005 Energy-aware wireless systems with adaptive power-fidelity tradeoffs
abstract
Wireless networked embedded systems, such as multimedia terminals, sensor nodes, etc., present a rich domain for making energy/performance/quality tradeoffs based on application needs, network conditions, etc. Energy awareness in these systems is the ability to perform tradeoffs between available battery energy and application quality requirements. In this paper, we show how operating system directed dynamic voltage scaling and dynamic power management can provide for such a capability. We propose a real-time scheduling algorithm that uses runtime feedback about application behavior to provide adaptive power-fidelity tradeoffs. We demonstrate our approach in the context of a static priority-based preemptive task scheduler. Simulation results show that the proposed algorithm results in significant energy savings compared to state-of-the-art dynamic voltage scaling schemes with minimal loss in system fidelity. We have implemented our scheduling algorithm into the eCos real-time operating system running on an Intel XScale-based variable voltage platform. Experimental results obtained using this platform confirm the effectiveness of our technique
Vijay Raghunathan, Cristiano Pereira, Mani Srivastava 0001, Rajesh K. Gupta 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2004 Joint end-to-end scheduling, power control and rate control in multi-hop wireless networks
abstract
This paper addresses the problem of joint scheduling, power control and rate control while maximizing end-to-end data rates in multi-hop wireless networks. Using a "physical layer" network model that explicitly takes into account interference due to spatial spectrum reuse, we formulate the throughput maximization problem as a mixed integer linear programming problem (MILP). While a MILP based approach yields an optimal solution, it does not scale well to large networks. To address this issue, we also present a computationally efficient water-filling based heuristic. Simulation results, obtained using our heuristic, highlight several capacity related tradeoffs that arise in wireless ad-hoc networks. Prior work only provides either asymptotic results on ad-hoc network capacity or, at best, techniques for computing loose upper bounds for throughput in specific instances of networks.
Gautam Kulkarni, Vijay Raghunathan, Mani Srivastava 0001
GLOBECOM2
2004 Experience with a low power wireless mobile computing platform
abstract
A detailed power analysis of a multi-radio mobile platform highlights the complex tradeoffs between the computation, storage, and communication subsystems. The particular mobile device, which does not include an LCD or other on-board display, can be used as a source for audio or video media files, or a source/sink for secure data transfers. A version of the device has been augmented with fine-grained power monitoring capability and used to obtain detailed measurements of power dissipation in the various subsystems. Analysis of these measurements sheds light on the power consumption characteristics of different applications, thereby providing hints to system designers about potential areas for optimization. Specifically, this work contrasts the power efficiency of the various wireless technologies supported by the system.
Vijay Raghunathan, Trevor Pering, Roy Want, Alex Nguyen 0005, Peter Jensen
ISLPED1
2004 Energy efficient wireless packet scheduling and fair queuing
abstract
As embedded systems are being networked, often wirelessly, an increasingly larger share of their total energy budget is due to the communication. This necessitates the development of power management techniques that address communication subsystems, such as radios, as opposed to computation subsystems, such as embedded processors, to which most of the research effort thus far has been devoted. In this paper, we present techniques for energy efficient packet scheduling and fair queuing in wireless communication systems. Our techniques are based on an extensive slack management approach that dynamically adapts the output rate of the system in accordance with the input packet arrival rate. We use a recently proposed radio power management technique, dynamic modulation scaling (DMS), as a control knob to enable energy-latency trade-offs during wireless packet transmission. We first analyze a single input stream scenario, and describe a rate adaptation technique that results in significantly lower energy consumption (reductions of up to 10 ×), while still bounding the resulting packet delays. By appropriately setting the various parameters of our algorithm, the system can be made to traverse the energy-latency-fidelity trade-off space. We extend our techniques to a multiple input stream scenario, and present E 2 WFQ , an energy efficient version of the weighted fair queuing (WFQ) algorithm for fair packet scheduling. Simulation results show that large energy savings can be obtained through the use of E 2 WFQ , with only a small, bounded increase in worst case packet latency. Further, our results demonstrate that E 2 WFQ does not adversely affect the throughput allocation (and hence, fairness) of WFQ.
Vijay Raghunathan, Saurabh Ganeriwal, Mani Srivastava 0001, Curt Schurgers
ACM Trans. Embed. Comput. Syst.1
2003 A survey of techniques for energy efficient on-chip communication
abstract
Interconnects have been shown to be a dominant source of energy consumption in modern day System-on-Chip (SoC) designs. With a large (and growing) number of electronic systems being designed with battery considerations in mind, minimizing the energy consumed in on-chip interconnects becomes crucial. Further, the use of nanometer technologies is making it increasingly important to consider reliability issues during the design of SoC communication architectures. Continued supply voltage scaling has led to decreased noise margins, making interconnects more susceptible to noise sources such as crosstalk, power supply noise, radiation induced defects, etc. The resulting transient faults cause the interconnect to behave as an unreliable transport medium for data signals. Therefore, fault tolerant communication mechanisms, such as Automatic Repeat Request (ARQ), Forward Error Correction (FEC), etc., which have been widely used in the networking community, are likely to percolate to the SoC domain.This paper presents a survey of techniques for energy efficient on-chip communication. Techniques operating at different levels of the communication design hierarchy are described, including circuit-level techniques, such as low voltage signaling, architecture-level techniques, such as communication architecture selection and bus isolation, system-level techniques, such as communication based power management and dynamic voltage scaling for interconnects, and network-level techniques, such as error resilient encoding for packetized on-chip communication. Emerging technologies, such as Code Division Multiple Access (CDMA) based buses, and wireless interconnects are also surveyed.
Vijay Raghunathan, Mani Srivastava 0001, Rajesh K. Gupta 0001
DAC1
2003 Energy efficiency and fairness tradeoffs in multi-resource, multi-tasking embedded systems
abstract
This paper presents techniques for optimizing the energy efficiency of multi-resource, multi-tasking embedded systems. Low power design of individual system resources, such as embedded processors, has been extensively studied in the past. However, system-level techniques, such as those presented in this paper, which exploit the synergy between various system resources, achieve levels of energy efficiency that cannot be obtained by considering individual resources independently. We demonstrate that, in multi-resource embedded systems that concurrently execute multiple applications, there exists a tradeoff between resource management efficiency and resource allocation fairness. By solving the multi-resource energy optimization problem in the context of an embedded sensor system, we show that our techniques enable the system designer to traverse this efficiency-fairness tradeoff space.
Sung I. Park, Vijay Raghunathan, Mani Srivastava 0001
ISLPED2
2003 Power management for energy-aware communication systems
abstract
System-level power management has become a key technique to render modern wireless communication devices economically viable. Despite their relatively large impact on the system energy consumption, power management for radios has been limited to shutdown-based schemes, while processors have benefited from superior techniques based on dynamic voltage scaling (DVS). However, similar scaling approaches that trade-off energy versus performance are also available for radios. To utilize these in radio power management, existing packet scheduling policies have to be thoroughly rethought to make them energy-aware, essentially opening a whole new set of challenges the same way the introduction of DVS did to CPU task scheduling. We use one specific scaling technique, dynamic modulation scaling (DMS), as a vehicle to outline these challenges, and to introduce the intricacies caused by the nonpreemptive nature of packet scheduling and the time-varying wireless channel.
Curt Schurgers, Vijay Raghunathan, Mani Srivastava 0001
ACM Trans. Embed. Comput. Syst.2
2002 E2WFQ: an energy efficient fair scheduling policy for wireless systems
abstract
As embedded systems are being networked, often wirelessly, an increasingly larger share of their total energy budget is due to the communication. This necessitates the development of power management techniques that address communication subsystems, such as radios, as opposed to computation subsystems, such as embedded processors, to which most of the research effort thus far has been devoted. In this paper, we present E2WFQ, an energy efficient version of the Weighted Fair Queuing (WFQ) algorithm for packet scheduling in communication systems. We employ a recently proposed radio power management technique, Dynamic Modulation Scaling (DMS), as a control knob to enable energy-latency tradeoffs during wireless packet scheduling. The use of E2WFQ results in an energy aware packet scheduler, which exploits the statistics of the input arrival pattern as well as the variability in packet lengths. Simulation results show that large savings in energy consumption can be obtained through the use of our scheduling scheme, compared to conventional WFQ, with only a small, bounded increase in worst case packet latency.
Vijay Raghunathan, Saurabh Ganeriwal, Curt Schurgers, Mani Srivastava 0001
ISLPED1
2001 Transient Power Management Through High Level Synthesis
abstract
The use of nanometer technologies is making it increasingly important to consider transient characteristics of a circuit's power dissipation (e.g., peak power, and power gradient or differential) in addition to its average power consumption. Current transient power analysis and reduction approaches are mostly at the transistor- and logic-levels. We argue that, as was the case with average power minimization, architectural solutions to transient power problems can complement and significantly extend the scope of lower-level techniques. In this work, we present a high-level synthesis approach to transient power management. We demonstrate how high-level synthesis can impact the cycle-by-cycle peak power and peak power differential for the synthesized implementation. Further, we demonstrate that it is necessary to consider transient power metrics judiciously in order to minimize or avoid area and performance overheads. In order to alleviate the limits on parallelism imposed by peak power constraints, we propose a novel technique based on the selective insertion of data monitor operations in the behavioral description. We present enhanced scheduling algorithms that can accept constraints on transient power characteristics (in addition to the conventional resource and performance constraints). Experimental results on several example designs obtained using a state-of-the-art commercial design flow and technology library indicate that high-level synthesis with transient power management results in significant benefits-peak power reductions of up to 32% (average of 25%), and peak power differential reductions of up to 59% (average of 42%)-with minimal performance overheads.
Vijay Raghunathan, Srivaths Ravi 0001, Anand Raghunathan, Ganesh Lakshminarayana
ICCAD1
2001 Adaptive Power-Fidelity in Energy-Aware Wireless Embedded System
abstract
Energy aware system operation, and not just low power hardware, is an important requirement for wireless embedded systems. These systems, such as wireless multimedia terminals or wireless sensor nodes, combine (soft) real-time constraints on computation and communication with requirements of long battery lifetime. In this paper, we present an OS-directed dynamic power management technique for such systems that goes beyond conventional techniques to provide an adaptive power vs. fidelity trade-off. The ability of wireless systems to adapt to changing fidelity in the form of data losses and errors is used to tradeoff against energy consumption. We also exploit system workload variation to proactively manage energy resources by predicting processing requirements. The supply voltage, and clock frequency are set according to predicted computation requirements of a specific task instance, and an adaptive feedback control machanism is used to keep system fidelity (deadline misses) within specifications. We present the theoretical framework underlying our approach in the context of both a static priority-based preemptive task scheduler as well as a dynamic priority based one, and present simulation-based performance analysis that shows that our technique provides large energy savings (up to 76%) with little loss in fidelity (<4%). Further, we describe the implementation of our technique in the eCos real-time operating system (RTOS) running on a StrongARM processor to illustrate the issues involved in enhancing RTOSs for energy awareness.
Vijay Raghunathan, Papleologos Spanos, Mani Srivastava 0001
RTSS1
2000 Integrating variable-latency components into high-level synthesis
abstract
Components used as building blocks (e.g., functional units) in conventional HLS techniques are assumed to have fixed latency values. Variable-latency units exhibit the property that the number of cycles taken to compute their outputs varies depending on the input values. While variable-latency units offer potential for performance improvement, we demonstrate that realization of this potential requires that HLS be adapted suitably (sub-optimal use of variable-latency units can lead to performance degradation, or unnecessarily high area overheads). Our techniques to incorporate variable-latency units into HLS ensure that the performance improvement is maximized, while minimizing area overheads or satisfying resource constraints. These techniques are not restricted to specific HLS tools/algorithms, and can be plugged in to any generic HLS system. Since area overheads may still be incurred due to the use of variable-latency units, we present a novel technique, based on the concept of reduced variable-latency units, to further reduce area overheads. Reduced variable-latency units only implement the low-latency case behavior of complete variable-latency units. We demonstrate that the use of reduced variable-latency units significantly reduces area overheads, and sometimes results in improvements in performance while simultaneously reducing the area of the register transfer level implementation. Experimental results show that the proposed variable-latency-unit-based synthesis techniques achieve a performance improvement of up to 1.6/spl times/ (average of 1.4/spl times/) over a state-of-the-art HLS tool, with minimal area overheads (average of 5.3%). The use of reduced variable-latency units leads to a performance improvement of up to 1.6/spl times/ (average of 1.3/spl times/), with a simultaneous area reduction of up to 17.9% (10.6% on the average).
Vijay Raghunathan, Srivaths Ravi 0001, Ganesh Lakshminarayana
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1