VLDB 2026 Research / reviewers in the wild / expert
Arnab Raha
dblp:69/10426
· DBLP profile ↗
51ranked-venue papers
13as first author
26since 2021 · last 2026
0000-0002-8848-1069ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 11 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COSMOS: Designing Energy-Efficient Context-Aware Multimodal Cognitive Systems
Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan |
ISLPED | 4 |
| 2026 | InterAxNN: Reconfigurable and Approximate in-Memory Processing Accelerator for Ultra-Low-Power Binary Neural Network Inference in Intermittently Powered SystemsabstractIn this work, we propose InterAxNN , an energy-aware approximate hardware architecture to perform vector-matrix multiplications in the binary precision regime for energy-constrained intermittently powered systems (IPS). In contrast to existing XNOR multiply-and-accumulate (MAC) operations implemented widely for binary neural networks (BNNs), we design a novel reconfigurable XNOR-MAC and AND-MAC memory macro to perform approximate binary precision operations, targeted for systems with extreme energy constraints. The proposed macro design is integrated with the ability to modify the MAC mode during run-time depending on instantaneous energy and power transients. We utilize the unique attributes of ferroelectric transistors (FeFETs) to implement the proposed ultra-low power BNN engine performing in-memory computing for artificial intelligence (AI) workloads. Subsequently, we leverage the quality configurable compute-in-memory-based hardware accelerator to implement InterAxNN based on a TI MSP430-based microcontroller. We evaluate the proposed InterAxNN concerning two baselines: (a) standard von Neumann computing architecture-based-microcontroller platform (MCU), and (b) MCU with a state-of-the-art low energy accelerator (MCU+LEA), and observe significant performance and energy benefits. Experimental results performed using a TI MSP430FR5379 IPS system show 448×–581× uplift in forward progress for 2%–8% accuracy loss for MNIST, 4%–5% accuracy loss for EMNIST, and 1%–4% reduction in accuracy for QMNIST, respectively, using MLP on a MCU+LEA platform with Unified NVM architecture. The AND-MAC mode in InterAxNN results in 91×–127× amount of additional forward progress over XNOR-MAC for 1%–2%, 4%–19%, and 1%–5% higher quality degradation for MNIST, EMNIST, and QMNIST, respectively. Arnab Raha, Sandeep Krishna Thirumala, Sumeet Kumar Gupta, Vijay Raghunathan |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2026 | Algorithm-Hardware Co-Design of Digital Compute-in-Memory Architecture Supporting Flexible and Temporal N:M SparsityabstractStructured pruning with fixed N:M sparsity ratios in large language models (LLMs) significantly constrains model expressivity, often leading to suboptimal accuracy. While supporting multiple N:M configurations can enhance representational flexibility, such a capability typically introduces substantial hardware complexity and overhead. To overcome these limitations, we first present FLOW, a flexible, layer-wise, outlier-density-aware N:M sparsity selection framework. FLOW adaptively determines the optimal N and M values per layer within a specified range by jointly considering the magnitude and distribution of outliers, thereby improving sparsity allocation and model fidelity. Extending this idea, we introduce FLOW++, which generalizes flexible N:M sparsity to the temporal domain for reasoning LLMs. FLOW++ enables temporal pruning by dynamically adapting sparsity patterns across reasoning thought types. To support efficient deployment of models with such dynamically varying sparsity patterns, we propose FlexCiM, a flexible, low-overhead, digital compute-in-memory (DCiM) architecture. FlexCiM partitions the DCiM macro into smaller submacros, which are adaptively aggregated and disaggregated through distribution and merging mechanisms for different values of N and M. We conduct experiments across different LLM families and state-space models and conclusively demonstrate that the proposed algorithm-hardware co-design framework achieves up to 36% higher accuracy,$1.75\times $faster inference, and$1.5\times $lower energy consumption compared to existing alternatives, establishing an effective balance between flexibility and hardware efficiency in sparse LLM inference. Code is available at:https://github.com/FLOW-open-project/FLOW Akshat Ramachandran, Souvik Kundu 0009, Arnab Raha, Shamik Kundu, Deepak Mathaikutty, Tushar Krishna |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Machine Learning-Driven STL Generation for Enhancing Functional Safety of E/E SystemsabstractThe increasing complexity of safety-critical hardware systems demands advanced methods for ensuring functional safety (FuSa). Traditional techniques like ATPG and BIST are intrusive, requiring additional hardware and disrupting operations, making them unsuitable for in-field testing. To address this, for the first time, we propose a machine learning (ML)-driven automated Self-Test Library (STL) generation for seamless in-field testing during idle periods, ensuring uninterrupted fault detection and high system performance. Utilizing reinforcement learning, the STL generates design-specific test patterns, achieving up to $57.57 \%$ improvement in fault coverage and up to $85 \%$ efficiency compared to existing pattern-based testing, enhancing FuSa in mission-critical applications. Sanjay Das, Swastik Bhattacharya, Anand Menon, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
DAC | 8 |
| 2025 | Enhancing AMS Circuit Reliability: An Anomaly Dataset for Functional Safety Research in Automotive SoCs
Sanjay Das, Anand Menon, Omar Abiola Abioye, Afreen Fatimah Khazi-Syed, Jonathan Edward Lee, Ayush Arunachalam, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
ACM Great Lakes Symposium on VLSI | 11 |
| 2025 | GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing UnitsabstractGraph Neural Networks (GNNs) are crucial for learning and reasoning over graph-structured data, with applications in network analysis, recommendation systems, and speech analytics. Deploying them on edge devices, such as client PCs and laptops, enables real-time processing, enhances privacy, and reduces cloud dependency. For instance, GNNs can augment Retrieval-Augmented Generation (RAG) for Large Language Models (LLMs) and enable event-based vision tasks. However, irregular memory access, sparse graphs, and dynamic structures lead to high latency and energy consumption on resource-constrained devices. Modern edge processors combine CPUs, GPUs, and NPUs, where NPUs excel at data-parallel tasks but face challenges with irregular GNN computations. To address these gaps, we present GraNNite, the first hardware-aware framework tailored to optimize GNN deployment on commercial-off-the-shelf (COTS) state-of-the-art (SOTA) DNN accelerators using a systematic three-step methodology: (1) enabling GNN execution on NPUs, (2) optimizing performance, and (3) trading accuracy for further performance and energy efficiency gains. Towards that end, the first category includes techniques such as GraphSplit for workload distribution and StaGr for static graph aggregation, while GrAd and NodePad handle real-time updates for dynamic graphs. Next, performance improvement is acquired through techniques such as EffOp for control-heavy operations and GraSp for sparsity exploitation. For Graph Convolution layers, PreG, SymG, and CacheG reduce redundancy and memory transfers. The final class of techniques deals with quality vs efficiency tradeoffs – QuantGr applies INT8 quantization to lower memory usage and computation time, while GrAx1, GrAx2, and GrAx3 optimize graph attention, broadcast-add, and sample-and-aggregate (SAGE)-max aggregation for higher throughput with minimal quality loss. Experimental evaluations on Intel® Core™ Ultra Series 1 and 2 AI PCs demonstrate that GraNNite achieves speedups of 2.6× to 7.6× over default NPU mappings, with energy efficiency improvements up to 8.6× compared to CPUs and GPUs. Across various GNN models, GraNNite delivers up to 10.8× and 6.7× higher performance than CPUs and GPUs, respectively. Our code implementation is available at this link. Arghadip Das, Shamik Kundu, Arnab Raha, Soumendu Kumar Ghosh, Deepak Mathaikutty, Vijay Raghunathan |
IJCNN | 3 |
| 2025 | Demo Abstract: ECO: Low Power Context-Aware Multimodal AI on NPUsabstractWe present ECO, the first system enabling efficient multimodal AI deployment on commercial Neural Processing Units (NPUs) through context-aware sensor and compute optimizations. ECO introduces runtime-tunable, NPU-architecture-aware knobs—approximate interpolation, quantization, and model scaling—that adapt to system conditions such as energy availability and sensor reliability. Deployed on an Intel Core Ultra Series 2 NPU with RGB and LiDAR inputs for a semantic segmentation application, ECO achieves up to 4.9× performance and 11.3× energy-efficiency improvement over CPU. Compared to systems lacking runtime context adaptability, ECO preserves higher segmentation quality (48.1 mean IoU in %, referred to as IoU hereafter) vs. 37.9 IoU under energy constraints and restores accuracy from 30.6 IoU to 40.0 IoU in sensor failure scenarios. The demo video and the ECO codebase are available at https://github.com/arghadippurdue/ECO%5FDemo. Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan |
ISLPED | 4 |
| 2025 | Demo Abstract: A Low-Power Real-Time Hardware Accelerator for Edge Detection Using Stochastic ComputingabstractWe present a low-power, stochastic computing-based method for real-time video edge detection. Traditional Sobel-based pipelines are often resource-intensive and consume substantial power. In this work, we simplify the Sobel operator within a stochastic computing framework to achieve significant reductions in hardware complexity and energy consumption. We implement the proposed design on a Basys 3 FPGA interfaced with an OV7670 camera, demonstrating real-time performance. Experimental results show up to 17% energy savings, 84% reduction in LUT utilization, and a 68% decrease in RAM storage and a substantial reduction in hardware footprint compared to a traditional implementation. Demo video link: https://github.com/arghadippurdue/StoBelDemo. Priyajit Ghosh, Rajarshi Mukherjee, Auro Anand Saha, Sutirtha Naha, Arghadip Das, Arnab Raha, Mrinal K. Naskar |
ISLPED | 6 |
| 2025 | Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory AcceleratorabstractLarge language model (LLM) pruning with fixed N:M structured sparsity significantly limits the expressivity of the sparse model, yielding sub-optimal performance. On the contrary, support for more than one N:M pattern to provide sparse representational freedom yields a costly overhead in the hardware. To mitigate these challenges for LLMs, we first present a flexible layer-wise outlier-density-aware N:M sparsity (FLOW) selection method. FLOW enables the identification of optimal layer-wise N and M values (from a given range) by simultaneously accounting for the presence and distribution of outliers, allowing a higher degree of representational freedom. To deploy the sparse models with such N:M flexibility, we then present a flexible low overhead, digital computein-memory architecture (FlexCiM). FlexCiM enables support for diverse sparsity patterns by partitioning a digital CiM (DCiM) macro into smaller sub-macros which are adaptively aggregated and disaggregated through distribution and merging mechanisms for different values of N and M. Extensive experiments on both transformer-based and recurrence-based state space foundation models (SSMs) demonstrate FLOW to outperform existing alternatives with an accuracy improvement of up to 36%, while FlexCiM delivers up to 1.75× lower inference latency and 1.5× lower energy consumption compared to existing sparse accelerators. Code is available at: https://github.com/FLOW-open-project/FLOW Akshat Ramachandran, Souvik Kundu 0009, Arnab Raha, Shamik Kundu, Deepak K. Mathaikutty, Tushar Krishna |
ISLPED | 3 |
| 2025 | OpenAssert: Towards Secure Assertion Generation using Large Language ModelsabstractAssertions are critical components used in hardware verification, ensuring robust functionality, fortifying design security, and providing essential verification features. Traditional hardware assertion methods are not automated, complicate security audits, and require effort, causing prolonged development cycles. Recent studies have highlighted the potential of commercial Large Language Models (LLMs) to generate security-focused assertions by leveraging textual data from design specifications. However, reliance on proprietary models like GPT-4 severely jeopardizes IP privacy and data confidentiality, undermining transparency and accountability in data handling practices. In this paper, we address secure hardware assertion generation by proposing a practical approach to significantly enhance the feasibility of open-source LLMs. Our proposed method, OpenAssert, involves fine-tuning existing models to be utilized locally at the user’s end without compromising confidentiality. Additionally, we employ Retrieval Augmentation Generation to refine these models, mitigating hallucinations and security-related errors. OpenAssert demonstrates improvements, achieving up to a 44% increase in rouge-1 score, a 49% improvement in cosine similarity, and a 43.4% reduction in word error rate for security-critical designs compared to open-source models. Anand Menon, Samit Shahnawaz Miftah, Amisha Srivastava, Shamik Kundu, Shovik Kundu, Arnab Raha, Suvadeep Banerjee, Deepak Mathaikutty, Kanad Basu |
VTS | 6 |
| 2024 | Graph Learning-based Fault Criticality Analysis for Enhancing Functional Safety of E/E SystemsabstractThe increasing complexity of Electrical and Electronic (E/E) systems underscores the need for protective measures to ensure functional safety (FuSa) in high-assurance environments. This entails the identification and fortification of vulnerable nodes to enhance system reliability during mission-critical scenarios. Traditionally, the assessment of E/E system reliability has relied on fault injection (FI) techniques and simulations. However, FI faces challenges in coping with escalating design complexity, including resource demands and timing overheads. Furthermore, it falls short in identifying critical components that may lead to functional failures. To address these challenges, we propose a Machine Learning (ML)-based framework for predicting critical nodes in hardware designs. The process begins with constructing a graph from the design netlist, forming the foundation for training a Graph Convolutional Network (GCN). The GCN model utilizes graph node attributes, node labels, and edge connections to learn and predict critical nodes in the circuit. The model furnishes up to 93.7% accuracy in identifying vulnerable circuit nodes during evaluation on diverse designs such as Synchronous Dynamic Random Access Memory (SDRAM) controller, OpenRISC 1200 (OR1200) modules. Furthermore, we incorporate an explainability analysis to interpret individual node predictions. This analysis discerns the critical design factors influencing fault criticality in the design. Moreover, to the best of our knowledge, we, for the first time, perform a regression analysis to generate node criticality scores, quantifying the degrees of criticality, that can enable prioritizing resources towards critical nodes. Sanjay Das, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
DAC | 6 |
| 2024 | SwiSS: Switchable Single-Sided Sparsity-based DNN AcceleratorsabstractDeep Neural Networks (DNNs) exhibit sparsity in both activation and weight tensors, but certain layers have higher weight sparsity, while others have higher activation sparsity. This challenges the conventional approach of fixing sparsity acceleration to either weights or activations alone. Conversely, harnessing both-sided sparsity necessitates complex design logic for identifying participating non-zero weights and activation pairs during a multiply-accumulate operation, leading to a significant impact on energy efficiency and area overhead in the edge accelerator. In this paper, we, for the first time, exploit the unbalanced sparsity in DNNs to propose the concept of dynamically Switchable Single-sided Sparsity, SwiSS, to improve energy efficiency in edge DNN accelerators. Through a novel self-adaptive dynamic sparsity selection algorithm, SwiSS can determine whether to enable one-sided weight or one-sided activation sparsity for a sparsity-enabled DNN accelerator. This capability allows SwiSS to dynamically exploit both sides of sparsity while maximizing the associated power and area benefits in the accelerator. Evaluation on state-of-the-art network-dataset configurations conducted on FlexNN [12] accelerator architecture demonstrates that SwiSS yields up to 30.76% and 8.29% improvements in power and area overheads, respectively (which translates to 1.42X and 1.08X improvement in TOPS/W and TOPS/mm2, respectively), compared to a combined two-sided sparsity scenario, with a negligible drop in sparsity acceleration. Shamik Kundu, Soumendu Kumar Ghosh, Arnab Raha, Deepak Mathaikutty |
ISLPED | 3 |
| 2024 | Toward Energy-Efficient Collaborative Inference Using Multisystem ApproximationsabstractCooperative inference applications have seen considerable potential with distributed deep neural networks (DDNNs). One use for DDNNs is the classification of 3-D objects from a set of 2-D images or views. This approach is also known as multiview convolutional neural networks (MVCNNs). However, due to the intensive computational demands, substantial communication overhead, high-inference delay, and energy limits, it is difficult to deploy MVCNN on resource-constrained edge devices. This article proposes for the first time the concept of distributed approximate systems (DRAX), which employs a multidevice approach to approximate computing and uses synergistic approximations of various edge computing systems to enable energy-efficient collaborative DDNN inference.DRAXperforms a significance-aware approximation of multiple nodes and prunes the large design space using the nonuniform contribution of various perspectives/views to the final inference to achieve optimal quality-energy tradeoff. In addition, we also propose a novel remaining energy-aware heuristic, which dynamically chooses the approximation degree based on the user-provided quality bounds and further increases the system lifetime. The experimental results obtained from a prototype of a 12-view 3-D object classification system implemented on an Intel Stratix IV FPGA development board demonstrate substantial energy savings ($2.6 \times$to$8\times$) for minimal (<1%) application-level quality loss. Arghadip Das, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan |
IEEE Internet Things J. | 3 |
| 2024 | DiagNNose: Toward Error Localization in Deep Learning Hardware-Based on VTA-TVM StackabstractLow-level hardware faults manifested in a Deep learning (DL) accelerator usher in graceless degradation of high-level classification accuracy, which can eventuate to catastrophic circumstances. This violates the crucial Functional Safety (FuSa) of the DL accelerator, maintaining which is imperative in high-assurance applications. Conventional techniques for error localization incur high-test efforts, without regards to the unique challenges posed by DL systems. In this direction, we propose DiagNNose, a two-tier machine learning-based error localization framework for on-line fault management in DL accelerators. We develop a novel diagnostic pattern selection algorithm to obtain a minimal subset of functional test patterns, that are executed in the accelerator in mission mode. By extracting and analyzing dataflow-based features from the intermediate computations of the general matrix multiply (GEMM) core, a lightweight multilayer perceptron accomplishes bit-level error localization in 8-bit, 16-bit, and 32-bit datapath units with high fidelity. We have limited ourselves to a single accelerator design, i.e., the versatile tensor accelerator (VTA) architecture to evaluate our proposed DiagNNose framework. On executing state-of-the-art deep neural networks trained on ImageNet; error localization using only 30 diagnostic functional test patterns demonstrate up to 98.4% diagnosability, thereby demonstrating an improvement of 54.63% over a random test pattern set, with as low as 4.95% overhead in the DL accelerator in mission mode. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | PArtNNer: Platform-Agnostic Adaptive Edge-Cloud DNN Partitioning for Minimizing End-to-End LatencyabstractThe last decade has seen the emergence of Deep Neural Networks (DNNs) as the de facto algorithm for various computer vision applications. In intelligent edge devices, sensor data streams acquired by the device are processed by a DNN application running on either the edge device itself or in the cloud. However, “edge-only” and “cloud-only” execution of State-of-the-Art DNNs may not meet an application’s latency requirements due to the limited compute, memory, and energy resources in edge devices, dynamically varying bandwidth of edge-cloud connectivity networks, and temporal variations in the computational load of cloud servers. This work investigates distributed (partitioned) inference across edge devices (mobile/end device) and cloud servers to minimize end-to-end DNN inference latency. We study the impact of temporally varying operating conditions and the underlying compute and communication architecture on the decision of whether to run the inference solely on the edge, entirely in the cloud, or by partitioning the DNN model execution among the two. Leveraging the insights gained from this study and the wide variation in the capabilities of various edge platforms that run DNN inference, we propose PArtNNer , a platform-agnostic adaptive DNN partitioning algorithm that finds the optimal partitioning point in DNNs to minimize inference latency. PArtNNer can adapt to dynamic variations in communication bandwidth and cloud server load without requiring pre-characterization of underlying platforms. Experimental results for six image classification and object detection DNNs on a set of five commercial off-the-shelf compute platforms and three communication standards indicate that PArtNNer results in 10.2× and 3.2× (on average) and up to 21.1× and 6.7× improvements in end-to-end inference latency compared to execution of the DNN entirely on the edge device or entirely on a cloud server, respectively. Compared to pre-characterization-based partitioning approaches, PArtNNer converges to the optimal partitioning point 17.6× faster. Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan, Anand Raghunathan |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Enhanced ML-Based Approach for Functional Safety Improvement in Automotive AMS CircuitsabstractThe extensive adoption of safety-critical applications in high-assurance environments, such as the automotive domain, has laid emphasis on safeguarding the reliability and Functional Safety (FuSa) of the Electrical and/or Electronic (E/E) components constituting such systems. Most modern automotive Systems-on-Chips (SoCs) comprise Analog and Mixed Signal (AMS) circuits, which are more susceptible to faults than their digital equivalents. However, their attributes of operating in the continuous signal region can be leveraged to perform early anomaly detection, which could facilitate the subversion of the eventual hardware failure state, thereby improving the FuSa of the system. To this end, we had proposed a novel unsupervised learning-based early anomaly detection framework catered to automotive AMS circuits (in ITC 2022). However, existing approaches to AMS FuSa violation detection are limited by pre-specified feature inputs, and lack rationale for identifying signals to be monitored to perform anomaly detection. To address these issues as well as further augment our original solution, in this paper, we propose a novel anomaly detection strategy that involves: (1) a genetic algorithm-based feature selection approach, (2) a novel signal selection algorithm that ascertains the best intermediate circuit signal, for furnishing enhanced anomaly detection accuracy, while reducing the associated detection latency, and (3) an explainable AI (XAI)-based framework that boosts user interpretability and transparency of the anomaly detection framework. This XAI approach, in turn, can be provided as feedback to the designer during circuit design and validation. The proposed approach is evaluated using case studies of two representative AMS circuits, which are prevalent in automotive SoCs. Our experimental analyses demonstrate that the proposed approach furnishes up to 100% detection accuracy and 2.3× reduction in detection time compared to our existing framework, in addition to providing insights by improving transparency of the anomaly detection framework, thereby exhibiting the efficacy of our solution. Ayush Arunachalam, Sanjay Das, Monikka Rajan, Xiankun Jin, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu |
ITC | 7 |
| 2023 | HIPEDAP: Energy-Efficient Hardware Accelerators for Hidden Periodicity DetectionabstractHidden periodicity detection (HPD) forms the basis of various emerging and complex applications such as detecting tandem repeats in DNA, absence seizure detection in EEG signals,etc.. The solutions to the period estimation problem were not satisfactorily accurate until Ramanujan sums (RS) were used to explore the periodic decomposition of signals. Its use in hidden periodicity detection was streamlined to form Ramanujan Filter Bank (RFB), but its usage in the applications proved to be computationally expensive. This paper proposes HIPEDAP, an efficient set of hardware accelerators for hidden periodicity detection applications using Ramanujan Filter Bank. HIPEDAPis developed by proposing several incrementally efficient microarchitectures, from Arch-A to E targeting improvements in different aspects of the design such as area, power, and performance. Further, the inherent error resilience exhibited by HPD applications enables us to propose an approximate architecture Arch-F, that synergistically applies multiple approximation techniques such as approximate adder, multiplier, and precision scaling on top of Arch-E, resulting in significant performance and energy benefits. Experimental results obtained after synthesizing the microarchitectures on 45 nm technology demonstrate that the optimized Arch-E design is able to achieve 4.7X, 8.2X, and 1.7X improvements in terms of area, frequency of execution, and power, respectively. Further, Arch-F demonstrate additional power savings of 14.6% on average (max 32.2%) over Arch-E for almost no loss in application-level quality. Finally, across a suite of practical applications, HIPEDAPexhibited a speed-up in the range of$5.1 \;{\times }\; 10^{2}$X to$3.2 \;{\times }\; 10^{4}$compared to its software implementations. Arghadip Das, Chandrachur Majumder, Debaprasad De, Arnab Raha, Mrinal K. Naskar |
IEEE Trans. Computers | 4 |
| 2023 | Trouble-Shooting at GAN Point: Improving Functional Safety in Deep Learning AcceleratorsabstractThe proliferation of Deep Neural Networks (DNNs) in real-time mission critical applications has promoted the implementation of custom-built DNN inference accelerators. These accelerators require a considerable amount of on-chip memory to store millions of trained DNN parameters for executing inference at the edge. Drastic technology scaling in recent years have made these memory circuits highly vulnerable to faults due to various reasons like aging, latent defects, single event upsets, etc. Such faults are highly detrimental to the classification accuracy of the DNN accelerator, leading to the crucial Functional Safety (FuSa) violation. This can eventuate to catastrophic circumstances, when used in mission-critical applications. In order to detect such violations in mission mode, we propose to generate a set of functional test patterns by leveraging the concept of Generative Adversarial Networks (GANs), that are independent of the DNN model and the accelerator characteristics. Our experimental results demonstrate that, the generated test patterns significantly improve FuSa violation detection coverage by up to 130.28%, compared to existing techniques. To the best of our knowledge, this is the first work that generates GAN-based test patterns in order to perform FuSa violation detection in mission-critical DNN accelerators. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Computers | 3 |
| 2023 | A Novel Low-Power Compression Scheme for Systolic Array-Based Deep Learning AcceleratorsabstractThe proliferation of deep learning algorithms has catalyzed their utilization to solve a multitude of real-world problems. Algorithms such as deep neural networks (DNNs) are compute- and power-intensive, thereby accentuating the development of hardware platforms like DNN inference accelerators. However, inference execution of large DNNs in resource-constrained environments induces energy bottlenecks in these accelerators. Since large DNNs consist of hundreds of millions of trained parameters, accessing them from the accelerator memory incurs substantial energy. To address this challenge, we propose HardCompress, which, to the best of our knowledge, is the first low-power solution that uses traditional compression strategies pertaining to commercial DNN accelerators in resource-constrained IoT edge devices. The three-step approach involves hardware-based post-quantization trimming of weights, followed by their dictionary-based compression and subsequent decompression by a low-power hardware engine during inference in the accelerator. We evaluate the proposed solution on lightweight networks trained on the MNIST dataset, the compact model trained on the CIFAR-10 dataset, and large DNNs trained on the ImageNet dataset. Performance of HardCompress at different quantization levels has been analyzed. Furthermore, to quantify the effectiveness of the proposed solution, an energy framework that contrasts the DRAM energies of the original and HardCompressed models has been developed. Finally, a fault injection framework which compares the fault resilience of the original model with its HardCompressed counterpart is also proposed. Our results exhibit that HardCompress, without any performance degradation in large DNNs, furnishes a maximum compression of 99.27%, equivalent to$137\times $reduction in memory footprint and 0.07 J for 8-bit quantization in the systolic array-based DNN accelerator. Furthermore, our proposed low-power decompression engine incurs an area overhead of only 0.02%; thus, enabling HardCompress’ utilization in resource-constrained environments. Ayush Arunachalam, Shamik Kundu, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Energy-Efficient Approximate Edge Inference SystemsabstractThe rapid proliferation of the Internet of Things and the dramatic resurgence of artificial intelligence based application workloads have led to immense interest in performing inference on energy-constrained edge devices. Approximate computing (a design paradigm that trades off a small degradation in application quality for disproportionate energy savings) is a promising technique to enable energy-efficient inference at the edge. This article introduces the concept of an approximate edge inference system ( AxIS ) and proposes a systematic methodology to perform joint approximations between different subsystems in a deep neural network (DNN)-based edge inference system, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use a smart camera system that executes various DNN-based image classification and object detection applications to illustrate how the sensor, memory, compute, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: (a) Cam Edge , where the DNN is executed locally on the edge device, and (b) Cam Cloud , where the edge device sends the captured image to a remote cloud server that executes the DNN. We have prototyped such an approximate inference system using an Intel Stratix IV GX-based Terasic TR4-230 FPGA development board. Experimental results obtained using six large DNNs and four compact DNNs running image classification applications demonstrate significant energy savings (≈ 1.6× -4.7× for large DNNs and ≈ 1.5× -3.6× for small DNNs), for minimal (<1%) loss in application-level quality. Furthermore, results using four object detection DNNs exhibit energy savings of ≈ 1.5× -5.2× for similar quality loss. Compared to approximating a single subsystem in isolation, AxIS achieves 1.05× -3.25× gains in energy savings for image classification and 1.35× -4.2× gains for object detection on average, for minimal (<1%) application-level quality loss. Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Unsupervised Learning-based Early Anomaly Detection in AMS Circuits of Automotive SoCsabstractWith the proliferation of safety-critical applications in the automotive domain, it is imperative to guarantee the functional safety of circuits and components constituting automotive systems, e.g., the electrical and/or electronic subsystems in automotive vehicles. Analog and Mixed-Signal (AMS) circuits, prevalent in such systems, are more susceptible to faults than their digital counterparts, due to advanced manufacturing nodes, parametric perturbations, environmental stress, etc. However, their continuous signal characteristics provide an opportunity for early anomaly detection, which in turn, facilitates the deployment of safety mechanisms to prevent eventual system failure. Towards this end, we propose a novel unsupervised machine learning-based framework to perform early anomaly detection in AMS circuits. Our approach involves anomaly injection in various circuit locations and individual components to develop a training dataset encompassing a wide range of possible anomalous scenarios, feature extraction from observation signals, and clustering algorithms to facilitate anomaly detection. To this end, we propose a novel centroid selection technique for the unsupervised learning algorithms, which is tailored for detecting anomalies in AMS circuits. This approach furnishes high fidelity anomaly detection by identifying the ideal cluster centers corresponding to anomalous and non-anomalous signals. Furthermore, time series-based analysis is proposed to improve and expedite the anomaly detection performance. We evaluated our solution using a case study of two AMS circuits commonly present in automotive systems-on-chips. Our experimental results exhibit that the proposed approach furnishes up to 100% accuracy. Additionally, the time series-based technique reduces the anomaly detection latency by 5×, thereby demonstrating the efficacy of our solution. Ayush Arunachalam, Athulya Kizhakkayil, Shamik Kundu, Arnab Raha, Suvadeep Banerjee, Robert Jin, Kanad Basu |
ITC | 4 |
| 2022 | RIBoNN: Designing Robust In-Memory Binary Neural Network AcceleratorsabstractRRAM crossbar-based accelerators show promise to execute compute intensive Deep Learning applications at the edge. For highly energy-constrained systems, Binary Neural Networks (BNNs) have gained momentum in recent times as the reduced precision alleviates the costs associated with storage, compute and communication. However, faults manifested in a unit bitcell of a RRAM crossbar-based accelerator may lead to drastic degradation in accuracy of the BNN, resulting in unintended system behavior. In this paper, we propose RIBoNN, a robust RRAM-based in-memory BNN accelerator, that consists of a 2T2R differential bitcell as the basic element of the crossbar. By leveraging the inherent characteristics of the proposed bitcell, RIBoNN is capable of achieving in-situ fault tolerance, thus circumventing the need to stall the deployed application for detection or diagnosis at the edge. RIBoNN, when evaluated on image-based datasets yields up to 96.57 % improvement in BNN classification accuracy, at a fault rate of 5 %; thereby demonstrating significant fault-tolerance over the state-of-the-art XNOR-RRAM BNN accelerator. Even though RiBoNN furnishes a negligible energy overhead of 2.62% over XNOR-RRAM, our proposed accelerator significantly reduces the inference latency by performing 24.4 % faster MAC operations with identical area footprint, while providing immense fault tolerance at the edge. Shamik Kundu, Akul Malhotra, Arnab Raha, Sumeet Kumar Gupta, Kanad Basu |
ITC | 3 |
| 2022 | Special Session: Effective In-field Testing of Deep Neural Network Hardware AcceleratorsabstractOngoing research to obtain high performance Deep Neural Network (DNN) executions have led to the development of customized purpose-built deep learning inference accelerators. DNN accelerators are susceptible to faults, due to high-energy particles, process variations, temperature and structural deformities manifesting as latent defects. These faults can introduce misclassification, thereby jeopardizing the Functional Safety (FuSa) of the accelerator in mission mode, which can eventuate to disastrous consequences, including loss of human lives. In this paper, we explore the impact of such faults on the FuSa of a DNN accelerator by varying the network parameters, position and characteristics of the injected fault across multiple exhaustive datasets. Furthermore, we analyze the efficiency of a software-based self test scheme to detect FuSa violations in the accelerator in mission mode, that employs functional test patterns, akin to instances in the application dataset. The test patterns, selected from the dataset of the DNN, furnish up to 100% coverage with cardinality as low as 0.1% of the entire test dataset. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Kanad Basu |
VTS | 3 |
| 2022 | Exploring the Design of Energy-Efficient Intermittently Powered Systems Using Reconfigurable Ferroelectric TransistorsabstractIn this article, we explore the design of energy-efficient intermittently powered systems (IPSs) using reconfigurable-ferroelectric transistors (R-FEFETs). Utilizing the dynamic tunability between volatile and nonvolatile modes of operation in R-FEFETs, we design nonvolatile flip-flops (NVFFs) and memory (NVM) suitable for IPS. We present two variants of R-FEFET-based NVFFs (RNVFFs): 1) with automatic backup and 2) with need-based backup. While the former offers high backup energy efficiency, the latter offers low normal operation energy. We also present an IPS-specific R-FEFET-based NVM (3T-R) with high energy efficiency compared with FEFET-based 2T NVM. Leveraging these nonvolatile circuits, we map the microcontroller unit (MCU) core registers of an IPS to RNVFFs and its on-chip memory to 3T-R. Subsequently, we analyze system-level implications of improving NVFFs and NVM individually by using R-FEFETs compared with existing FEFET-based designs. Our system-level simulations demonstrate that although we improve the register energy by 55%–67%, the total memory and system-level energy savings obtained from just improving the NVFFs (registers) in the microcontroller core are only 0.60%–5.78% and 0.31%–3.18%, respectively. However, improving the NVM by using 3T-R results in a much larger total memory and system-level energy savings in the range of 37%–40% and 20%–22%, respectively, in the context of a state-of-the-art IPS. Sandeep Krishna Thirumala, Arnab Raha, Sumeet Kumar Gupta, Vijay Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Special Session: Approximate TinyML Systems: Full System Approximations for Extreme Energy-Efficiency in Intelligent Edge DevicesabstractApproximate computing (AxC) has advanced from being an emerging design paradigm to becoming one of the most popular and effective methods of energy optimization for applications in the domains of computer vision, image/video processing, data mining, analytics, and search. The simultaneous rise of artificial intelligence (AI) has provided an additional thrust to the adoption of various AxC techniques in intelligent edge platforms where energy-efficiency is not only desirable but necessary. In spite of the big rise in interest for AxC, the adoption of approximate hardware has mostly been limited to only one component of the system (usually the processing subsystem) which often contributes only a fraction of the overall system-level power. A full system approach to AxC enables us to extend approximations to other subsystems, such as the memory, sensor, and communications subsystems. This paper presents the foundational concepts of an approximate TinyML system that applies approximations synergistically to multiple subsystems in an edge inference device. These approximations are applied intelligently to significantly reduce energy while incurring a negligible loss in application-level quality. We demonstrate multiple versions of an approximate smart camera system that can execute state-of-the-art deep neural networks (DNNs) while consuming only a fraction of the total energy in a typical system. Arnab Raha, Soumendu Kumar Ghosh, Debabrata Mohapatra, Deepak Mathaikutty, Raymond Sung, Cormac Brick, Vijay Raghunathan |
ICCD | 1 |
| 2021 | Toward Functional Safety of Systolic Array-Based Deep Learning Hardware AcceleratorsabstractHigh accuracy and ever-increasing computing power have made deep neural networks (DNNs) the algorithm of choice for various machine learning, computer vision, and image processing applications across the computing spectrum. To this end, Google developed the tensor processing unit (TPU) to accelerate the computationally intensive matrix multiplication operation of a DNN on its systolic array architecture. Faults manifested in the datapath of such a systolic array due to latent manufacturing defects or single-event effects may lead to functional safety (FuSa) violation. Although DNNs are known to resist minor perturbations with their inherent fault-tolerant characteristics, we show that the classification accuracy of the model plummets from 97.4% to 7.75% with a minimal fault rate of 0.0003% in the accelerator, implying catastrophic circumstances when deployed across mission-critical systems. Hence, to ensure FuSa of such accelerators, this article provides an extensive FuSa assessment of the accelerator exposed to faults in the datapath, by varying the network parameters, position, and characteristics of the induced error across multiple exhaustive data sets. Furthermore, we propose two novel strategies to obtain a diminutive set of functional test patterns to detect FuSa violation in a DNN accelerator. Our experimental results demonstrate that the obtained test sets can achieve an average of 92.63% (in some cases, up to 100%) fault coverage with cardinality as low as 0.1% of the entire test data set. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | IPS-CiM: Enhancing Energy Efficiency of Intermittently-Powered Systems with Compute-in-MemoryabstractIntermittently Powered Systems (IPS) have an ability to sustain computation progress across multiple power cycles in the presence of unreliable and sporadic harvested energy. However, with the emergence of data-intensive applications to be processed on energy-constrained IPS, it becomes challenging to handle large amounts of data with standard IPS architectures due to the von-Neumann bottleneck. To address this issue, we propose a compute-in-memory (CiM) engine which alleviates the memory-processor bottleneck and enhances energy-efficiency for transient computing workloads in IPS. We present a ferroelectric transistor (FEFET) based memory architecture which supports (a) nonvolatile memory (NVM) storage, (b) standard Boolean and arithmetic operations, (c) cyclic redundancy check for error detection and (d) edge-sensing for wireless sensory networks. Using the proposed CiM engine as a unified NVM, we construct an integrated IPS-CiM architecture based on the TI MSP430 microcontroller system with supply capacitances in the range of 10 nF -1 μF. We evaluate the proposed design with two baselines: hybrid SRAM+NVM and unified NVM architectures, both of which perform standard out-of-memory computing. We observe that for 1μF supply capacitance, IPS-CiM results in energy and performance benefits in the range of 35X-450X and 32X-400X, respectively over conventional microcontroller-based systems. Sandeep Krishna Thirumala, Arnab Raha, Vijay Raghunathan, Sumeet Kumar Gupta |
ICCD | 2 |
| 2020 | Approximate inference systems (AxIS): end-to-end approximations for energy-efficient inference at the edgeabstractThe rapid proliferation of the Internet-of-Things (IoT) and the dramatic resurgence of artificial intelligence (AI) based application workloads has led to immense interest in performing inference on energy-constrained edge devices. Approximate computing (a design paradigm that yields large energy savings at the cost of a small degradation in application quality) is a promising technique to enable energy-efficient inference at the edge. This paper introduces the concept of an approximate inference system (AxIS) and proposes a systematic methodology to perform joint approximations across different subsystems in a deep neural network-based inference system, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use a smart camera system that executes various convolutional neural network (CNN) based image recognition applications to illustrate how the sensor, memory, compute, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: (a) Camedge, where the CNN executes locally on the edge device, and (b) Camcloud, where the edge device sends the captured image to a remote cloud server that executes the CNN. We have prototyped such an approximate inference system using an Altera Stratix IV GX-based Terasic TR4-230 FPGA development board. Experimental results obtained using six CNNs demonstrate significant energy savings (around 1.7× for Camedge and 3.5× for Camcloud) for minimal (< 1%) loss in application quality. Compared to approximating a single subsystem in isolation, AxIS achieves additional energy benefits of 1.6×--1.7× (Camedge) and 1.4×--3.4× (Camcloud) on average for minimal application-level quality loss. Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan |
ISLPED | 2 |
| 2020 | Approximate Memory CompressionabstractMemory subsystems are a major energy bottleneck in computing platforms due to frequent transfers between processors and off-chip memory. We propose approximate memory compression, a technique that leverages the intrinsic resilience of emerging workloads such as machine learning and data analytics to reduce off-chip memory traffic, thereby improving energy and performance. We realize approximate memory compression by enhancing the memory controller to be aware of approximate memory regions-regions in memory that contain approximation-resilient data-and to transparently compress (decompress) the data written to (read from) these regions. To provide control over approximations, each approximate memory region is associated with an error constraint such as the maximum error that may be introduced in each data element. The quality-aware memory controller subjects memory transactions to a compression scheme that introduces approximations, thereby reducing memory traffic, while adhering to the specified error constraint for each approximate memory region. A software interface is provided to allow programmers to identify data structures (DSs) that are resilient to approximations. A runtime quality control framework automatically determines the error constraints for the identified DSs such that a given target application-level quality is maintained. We evaluate our proposal by applying it to three different main memory technologies in the context of a general-purpose computing system-DDR3 DRAM, LPDDR3 DRAM, and spin-transfer torque magnetic RAM (STT-MRAM). To demonstrate the feasibility of the proposed concepts, we also implement a hardware prototype using the Intel UniPHY-DDR3 memory controller and Nios-II processor, a Hynix DDR3 DRAM module, and a Stratix-IV field-programmable gate array (FPGA) development board. Across a wide range of machine learning benchmarks, approximate memory compression obtains significant benefits in main memory energy (1.18× for DDR3 DRAM, 1.52× for LPDDR3 DRAM, and 2.0× for STT-MRAM) and a simultaneous improvement in execution time (5.2% for DDR3 DRAM, 5.4% for LPDDR3 DRAM, and 9.3% for STT-MRAM) with nearly identical application output quality. Ashish Ranjan 0001, Arnab Raha, Vijay Raghunathan, Anand Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Dual Mode Ferroelectric Transistor based Non-Volatile Flip-Flops for Intermittently-Powered SystemsabstractIn this work, we propose dual mode ferroelectric transistors (D-FEFETs) that exhibit dynamic tuning of operation between volatile and non-volatile modes with the help of a control signal. We utilize the unique features of D-FEFET to design two variants of non-volatile flip-flops (NVFFs). In both designs, D-FEFETs are operated in the volatile mode for normal operations and in the non-volatile mode to backup the state of the flip-flop during a power outage. The first design comprises of a truly embedded non-volatile element (D-FEFET) which enables a fully automatic backup operation. In the second design, we introduce need-based backup, which lowers energy during normal operation at the cost of area with respect to the first design. Compared to a previously proposed FEFET based NVFF, the first design achieves 19% area reduction along with 96% lower backup energy and 9% lower restore energy, but at 14%-35% larger operation energy. The second design shows 11% lower area, 21% lower backup energy, 16% decrease in backup delay and similar operation energy but with a penalty of 17% and 19% in the restore energy and delay, respectively. System-level analysis of the proposed NVFFs in context of a state-of-the-art intermittently-powered system using real benchmarks yielded 5%-33% energy savings. Sandeep Krishna Thirumala, Arnab Raha, Hrishikesh Jayakumar, Kaisheng Ma, Narayanan Vijaykrishnan, Vijay Raghunathan, Sumeet Kumar Gupta |
ISLPED | 2 |
| 2018 | D-PUF: An Intrinsically Reconfigurable DRAM PUF for Device Authentication and Random Number GenerationabstractPhysically Unclonable Functions (PUFs) have proved to be an effective and low-cost measure against counterfeiting by providing device authentication and secure key storage services. Memory-based PUF implementations are an attractive option due to the ubiquitous nature of memory in electronic devices and the requirement of minimal (or no) additional circuitry. Dynamic Random Access Memory-- (DRAM) based PUFs are particularly advantageous due to their large address space and multiple controllable parameters during response generation. However, prior works on DRAM PUFs use a static response-generation mechanism making them vulnerable to security attacks. Further, they result in slow device authentication, are not applicable to commercial off-the-shelf devices, or require DRAM power cycling prior to authentication. In this article, we propose D-PUF, an intrinsically reconfigurable DRAM PUF based on the idea of DRAM refresh pausing. A key feature of the proposed DRAM PUF is reconfigurability , that is, by varying the DRAM refresh-pause interval, the challenge-response behavior of the PUF can be altered, making it robust to various attacks. The article is broadly divided into two parts. In the first part, we demonstrate the use of D-PUF in performing device authentication through a secure, low-overhead methodology. In the second part, we show the generation of true random numbers using D-PUF. The design is implemented and validated using an Altera Stratix IV GX FPGA-based Terasic TR4-230 development board and several off-the-shelf 1GB DDR3 DRAM modules. Our experimental results demonstrate a 4.3×-6.4× reduction in authentication time compared to prior work. Using controlled temperature and accelerated aging tests, we also demonstrate the robustness of our authentication mechanism to temperature variations and aging effects. Finally, the ability of the design to generate random numbers is verified using the NIST Statistical Test Suite. Soubhagya Sutar, Arnab Raha, Devadatta M. Kulkarni, Rajeev Shorey, Jeffrey D. Tew, Vijay Raghunathan |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Designing Energy-Efficient Intermittently Powered Systems Using Spin-Hall-Effect-Based Nonvolatile SRAMabstractIntermittently powered systems represent a new class of batteryless devices that operate solely on energy harvested from their environment. Due to the unreliable nature of ambient energy sources, these devices experience frequent intervals of power loss, leading to sudden reboots. Tolerating such power supply disruptions require the ability to rapidly checkpoint/save system state when power loss is imminent and restore it at the start of the next power cycle to continue computations in a seamless manner. A typical microcontroller used in these systems consists of a fast nonvolatile SRAM and a nonvolatile Flash storage. Prior work has shown how emerging nonvolatile memory technologies such as STT-MRAM can improve the energy efficiency of these systems, either by using STT-MRAM as a drop-in replacement for Flash (henceforth referred to as the SRAM+STT-MRAM memory configuration) or using STT-MRAM as unified memory (henceforth referred to as the unified STT-MRAM memory configuration). However, both these configurations have significant drawbacks. Using the SRAM+STT-MRAM configuration leads to high checkpointing overhead due to the inefficient write operations of STT-MRAM whereas using the unified STT-MRAM configuration is inefficient due to executing every program instruction directly from STT-MRAM. This paper proposes a novel Spin Hall Effect-based nonvolatile-SRAM (SNVRAM) bit-cell that combines the nonvolatility of spin devices with the speed and energy efficiency of conventional 6T SRAM cells. We explore the use of the proposed SNVRAM to replace the SRAM in a transiently powered system to mitigate the drawbacks of the aforementioned memory configurations. Simulation results using a set of evaluation benchmarks demonstrate that the SNVRAM+STT-MRAM configuration leads to significant memory energy benefits of $2.6\times $ and $2.8\times $ on average, compared to the SRAM+STT-MRAM and unified STT-MRAM memory configurations, respectively. Arnab Raha, Akhilesh Jaiswal 0001, Syed Shakib Sarwar, Hrishikesh Jayakumar, Vijay Raghunathan, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Approximating Beyond the Processor: Exploring Full-System Energy-Accuracy Tradeoffs in a Smart Camera SystemabstractThe intrinsic error resilience exhibited by emerging application domains enables new avenues for energy optimization of computing systems, namely, the introduction of a small amount of approximations during system operation in exchange for substantial energy savings. Prior work in the area of approximate computing has focused on individual subsystems of a computing system, for example, the computational subsystem or the memory subsystem. Since they focus only on individual subsystems, these techniques are unable to exploit the large energy-saving opportunities that stem from adopting a full-system perspective and approximating multiple subsystems of a computing platform simultaneously in a coordinated manner. This paper proposes a systematic methodology to perform joint approximations across different subsystems, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use the example of a smart camera system that executes various computer vision and image processing applications to illustrate how the sensing, memory, processing, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: 1) a compute-intensive smart camera system, AxSYScomp, where the error-resilient application executes locally within the camera and produces the final application output, and 2) a communication-intensive smart camera system, AxSYScomp, that sends the captured image to a remote cloud server, where the error-resilient application is executed and the final output is generated. We have implemented such an approximate smart camera system using an Altera Stratix IV GX FPGA development board, a Terasic TRDB-D5M 5-Megapixel camera module, a Terasic RFS WiFi module, and a 1-GB DDR3 dynamic random access memory small outline dual in-line memory module (SODIMM). Experimental results obtained using six application benchmarks demonstrate significant energy savings (around 7.5x for AxSYScomp and 4x on average for AxSYScomp) for minimal (<;1%) loss in application quality. Compared to approximating a single subsystem, the proposed full-system approximation methodology achieves additional energy benefits of 3.5x-5.5x (in the case of AxSYScomp) and 1.8x-3.7x (in the case of AxSYScomm) on average for a minimal (<;1%) application-level quality loss. Arnab Raha, Vijay Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Towards Full-System Energy-Accuracy Tradeoffs: A Case Study of An Approximate Smart Camera SystemabstractThe intrinsic error resilience exhibited by emerging application domains enables a new dimension for energy optimization of computing systems, namely the introduction of a controlled amount of approximations during system operation in exchange for substantial energy savings. Prior work in the area of approximate computing has focused on individual subsystems of a computing system, e.g., the computational subsystem or the memory subsystem. Since they focus only on individual subsystems, these techniques are unable to exploit the large energy-saving opportunities that stem from adopting a full-system perspective and approximating multiple subsystems of a computing platform simultaneously in a coordinated manner. This paper proposes a systematic methodology to perform joint approximations across different subsystems, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use the example of a smart camera system that executes various computer vision and image processing applications to illustrate how the sensing, memory, and processing subsystems can all be approximated synergistically. We have implemented such an approximate smart camera system using an Altera Stratix IV GX FPGA development board, a Terasic TRDB-D5M 5 Megapixel camera module, and a 1GB DDR3 SODIMM module. Experimental results obtained using six application benchmarks demonstrate significant energy savings (around 7.5X on average) for minimal (< 1%) loss in application quality. Compared to approximating a single subsystem, the proposed full-system approximation methodology achieves additional energy benefits of 3.5X - 5.5X on average for minimal (< 1%) quality loss. Arnab Raha, Vijay Raghunathan |
DAC | 1 |
| 2017 | Approximate memory compression for energy-efficiencyabstractMemory subsystems are a major energy bottleneck in computing platforms due to frequent transfers between processors and off-chip memory. We propose approximate memory compression, a technique that leverages the intrinsic resilience of emerging workloads such as machine learning and data analytics to reduce off-chip memory traffic and energy. To realize approximate memory compression, we enhance the memory controller to be aware of memory regions that contain approximation-resilient data, and to transparently compress/decompress the data written to/read from these regions. To provide control over approximations, the quality-aware memory controller conforms to a specified error constraint for each approximate memory region. We design a software interface that programmers can use to identify data structures that are resilient to approximations. We also propose a runtime quality control framework that automatically determines the error constraints for the identified data structures such that a given target application-level quality is maintained. We evaluate our proposal by implementing a hardware prototype using the Intel UniPHY-DDR3 memory controller and NIOS-II processor, a Hynix DDR3 DRAM module, and a Stratix-IV FPGA development board. Across a suite of 8 machine learning benchmarks, approximate memory compression obtains a 1.28× benefit in DRAM energy and a simultaneous 11.5% improvement in execution time for a small (<; 1.5%) loss in output quality. Ashish Ranjan 0001, Arnab Raha, Vijay Raghunathan, Anand Raghunathan |
ISLPED | 2 |
| 2017 | Quality Configurable Approximate DRAMabstractApproximate computing is an emerging design paradigm that leverages the inherent error tolerance present in many applications to improve their power consumption and performance. Due to the forgiving nature of these error-resilient applications, precise input data is not always necessary for them to produce outputs of acceptable quality. This makes the memory subsystem (i.e., the place where data is stored), a suitable component for introducing approximations in return for substantial energy savings. Towards this end, this paper proposes a systematic methodology for constructing a quality configurable approximate DRAM system. Our design is based upon an extensive experimental characterization of memory errors as a function of the DRAM refresh-rate. Leveraging the insights gathered from this characterization, we propose four novel strategies for partitioning the DRAM in a system into a number of quality bins based on the frequency, location, and nature of bit errors in each of the physical pages, while also taking into account the property of variable retention time exhibited by DRAM cells. During data allocation, critical data is placed in the highest quality bin (that contains only accurate pages) and approximate data is allocated to bins sorted in descending order of quality, with the refresh rate serving as the quality control knob. We validate our proposed scheme on several error-resilient applications implemented using an Altera Stratix IV GX FPGA based Terasic TR4-230 development board containing a 1GB DDR3 DRAM module. Experimental results demonstrate a significant improvement in the energy-quality trade-off compared to previous work and show a reduction in DRAM refresh power of up to 73 percent on average with minimal loss in output quality. Arnab Raha, Soubhagya Sutar, Hrishikesh Jayakumar, Vijay Raghunathan |
IEEE Trans. Computers | 1 |
| 2017 | Energy-Aware Memory Mapping for Hybrid FRAM-SRAM MCUs in Intermittently-Powered IoT DevicesabstractForecasts project that by 2020, there will be around 50 billion devices connected to the Internet of Things (IoT), most of which will operate untethered and unplugged. While environmental energy harvesting is a promising solution to power these IoT edge devices, it introduces new complexities due to the unreliable nature of ambient energy sources. In the presence of an unreliable power supply, frequent checkpointing of the system state becomes imperative, and recent research has proposed the concept of in-situ checkpointing by using ferroelectric RAM (FRAM), an emerging non-volatile memory technology, as unified memory in these systems. Even though an entirely FRAM-based solution provides reliability, it is energy inefficient compared to SRAM due to the higher access latency of FRAM. On the other hand, an entirely SRAM-based solution is highly energy efficient but is unreliable in the face of power loss. This paper advocates an intermediate approach in hybrid FRAM-SRAM microcontrollers that involves judicious memory mapping of program sections to retain the reliability benefits provided by FRAM while performing almost as efficiently as an SRAM-based system. We propose an energy-aware memory mapping technique that maps different program sections to the hybrid FRAM-SRAM microcontroller such that energy consumption is minimized without sacrificing reliability. Our technique consists of eM-map , which performs a one-time characterization to find the optimal memory map for the functions that constitute a program and energy-align , a novel hardware-software technique that aligns the system’s powered-on time intervals to function execution boundaries, which results in further improvements in energy efficiency and performance. Experimental results obtained using the MSP430FR5739 microcontroller demonstrate a significant performance improvement of up to 2x and energy reduction of up to 20% over a state-of-the-art FRAM-based solution. Finally, we present a case study that shows the implementation of our techniques in the context of a real IoT application. Hrishikesh Jayakumar, Arnab Raha, Jacob R. Stevens, Vijay Raghunathan |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | qLUT: Input-Aware Quantized Table Lookup for Energy-Efficient Approximate AcceleratorsabstractApproximate computing has emerged as a popular design paradigm for optimizing the performance and energy consumption of error-resilient applications in domains such as machine learning, graphics, data analytics, etc . Numerous techniques for approximate computing have been proposed at different layers of the system stack, from circuits to architecture to software. In this work, we propose a new technique, called quantized table lookup , for approximating the meta-functions used in the core computational kernels of error-resilient applications. In contrast to prior work that directly approximates the functionality of the meta-functions, the proposed technique instead approximates the input data to the meta-functions by reducing/quantizing them to a much smaller set of values that we call quantized inputs . The small number of quantized inputs enables us to completely replace the energy-intensive arithmetic units in the meta-function with small and energy-efficient lookup tables (called quantized lookup tables or q LUT) that contain precomputed output values corresponding to the quantized inputs. The proposed approximation technique is not only highly generic, but also inherently quality-configurable and input-aware. Quality-configurability and input-awareness are achieved by modulating the size of the q LUT as well as selecting the values of the quantized inputs judiciously based on the statistics of the original input data. To evaluate the proposed technique, we have implemented the dominant meta-functions of nine error-resilient application benchmarks as quantized table lookup based hardware accelerators using 45nm technology. Experimental results demonstrate average energy savings of 46% at the application-level for minimal (<1%) loss in output quality. Arnab Raha, Vijay Raghunathan |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2017 | Energy-Efficient Reduce-and-Rank Using Input-Adaptive ApproximationsabstractApproximate computing is an emerging design paradigm that exploits the intrinsic ability of applications to produce acceptable outputs even when their computations are executed approximately. In this paper, we explore approximate computing for a key computation pattern, reduce-andrank (RnR), which is prevalent in a wide range of workloads, including video processing, recognition, search, and data mining. An RnR kernel performs a reduction operation (e.g., distance computation, dot product, and L1-norm) between an input vector and each of a set of reference vectors, and ranks the reduction outputs to select the top reference vectors for the current input. We propose three complementary approximation strategies for the RnR computation pattern. The first is interleaved reductionand-ranking, wherein the vector reductions are decomposed into multiple partial reductions and interleaved with the rank computation. Leveraging this transformation, we propose the use of intermediate reduction results and ranks to identify future computations that are likely to have a low impact on the output, and can, hence, be approximated. The second strategy, inputsimilarity-based approximation, exploits the spatial or temporal correlation of inputs (e.g., pixels of an image or frames of a video) to identify computations that are amenable to approximation. The third strategy, reference vector reordering, rearranges the order in which the reference vectors are processed such that vectors that are relatively more critical in evaluating the correct output, are processed at the beginning of RnR operation. The number of these critical reference vectors is usually small, which renders a substantial portion of the total computation to be amenable to approximation. These strategies address a key challenge in approximate computing-identification of which computations to approximate-and may be used to drive any approximation mechanism, such as computation skipping or precision scaling to realize performance and energy improvements. A second key challenge in approximate computing is that the extent to which computations can be approximated varies significantly from application to application, and across inputs for even a single application. Hence, input-adaptive approximation, or the ability to automatically modulate the degree of approximation based on the nature of each individual input, is essential for obtaining optimal energy savings. In addition, to enable quality configurability in RnR kernels, we propose a kernel-level quality metric that correlates well to application-level quality, and identify key parameters that can be used to tune the proposed approximation strategies dynamically. We develop a runtime framework that modulates the identified parameters during the execution of RnR kernels to minimize their energy while meeting a given target quality. To evaluate the proposed concepts, we designed quality-configurable hardware implementations of six RnR-based applications from the recognition, mining, search, and video processing application domains in 45-nm technology. Our experiments demonstrate a 1.13×-3.18× reduction in energy consumption with virtually no loss in output quality (<;0.5%) at the application level. The energy benefits further improve up to 3.43× and 3.9× when the quality constraints are relaxed to 2.5% and 5%, respectively. Arnab Raha, Swagath Venkataramani, Vijay Raghunathan, Anand Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Energy-efficient system design for IoT devicesabstractIt is projected that, within the coming decade, there will be more than 50 billion smart objects connected to the Internet of Things (IoT). These smart objects, which connect the physical world with the world of computing infrastructure, are expected to pervade all aspects of our daily lives and revolutionize a number of application domains such as healthcare, energy conservation, transportation, etc. In this paper, we present an overview of the challenges involved in designing energy-efficient IoT edge devices and describe recent research that has proposed promising solutions to address these challenges. First, we outline the challenges involved in efficiently supplying power to an IoT device. Next, we discuss the role of emerging memory technologies in making IoT devices energy-efficient. Finally, we discuss the potential impact that approximate computing can have in increasing the energy-efficiency of wearables and other compute-intensive IoT devices. Hrishikesh Jayakumar, Arnab Raha, Younghyun Kim 0001, Soubhagya Sutar, Woo Suk Lee, Vijay Raghunathan |
ASP-DAC | 2 |
| 2016 | D-PUF: an intrinsically reconfigurable DRAM PUF for device authentication in embedded systemsabstractPhysically Unclonable Functions (PUFs) have proved to be effective and low-cost measure against counterfeiting by providing device authentication and secure key storage services. Memory based PUF implementations are an attractive option due to the ubiquitous nature of memory in electronic devices and the requirement of minimal (or no) additional circuitry. DRAM based PUFs are particularly advantageous due to their large address space and multiple controllable parameters during response generation. However, prior works on DRAM PUFs use a static response generation mechanism making them vulnerable to security attacks. Further, they result in very slow device authentication, are not applicable to off-the-shelf DRAM modules, or require DRAM power cycling prior to authentication. Soubhagya Sutar, Arnab Raha, Vijay Raghunathan |
CASES | 2 |
| 2016 | Sleep-Mode Voltage Scaling: Enabling SRAM Data Retention at Ultra-Low Power in Embedded Microcontrollers
Hrishikesh Jayakumar, Arnab Raha, Vijay Raghunathan |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Input-Based Dynamic Reconfiguration of Approximate Arithmetic Units for Video EncodingabstractThe field of approximate computing has received significant attention from the research community in the past few years, especially in the context of various signal processing applications. Image and video compression algorithms, such as JPEG, MPEG, and so on, are particularly attractive candidates for approximate computing, since they are tolerant of computing imprecision due to human imperceptibility, which can be exploited to realize highly power-efficient implementations of these algorithms. However, existing approximate architectures typically fix the level of hardware approximation statically and are not adaptive to input data. For example, if a fixed approximate hardware configuration is used for an MPEG encoder (i.e., a fixed level of approximation), the output quality varies greatly for different input videos. This paper addresses this issue by proposing a reconfigurable approximate architecture for MPEG encoders that optimizes power consumption with the goal of maintaining a particular Peak Signal-to-Noise Ratio (PSNR) threshold for any video. Toward this end, we design reconfigurable adder/subtractor blocks (RABs), which have the ability to modulate their degree of approximation, and subsequently integrate these blocks in the motion estimation and discrete cosine transform modules of the MPEG encoder. We propose two heuristics for automatically tuning the approximation degree of the RABs in these two modules during runtime based on the characteristics of each individual video. Experimental results show that our approach of dynamically adjusting the degree of hardware approximation based on the input video respects the given quality bound (PSNR degradation of 1%-10%) across different videos while achieving a power saving up to 38% over a conventional nonapproximated MPEG encoder architecture. Note that although the proposed reconfigurable approximate architecture is presented for the specific case of an MPEG encoder, it can be easily extended to other DSP applications. Arnab Raha, Hrishikesh Jayakumar, Vijay Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Quality-aware data allocation in approximate DRAM?abstractApproximate computing is an emerging design paradigm that leverages the inherent error tolerance present in many applications to optimize their power consumption and performance. Due to the forgiving nature of these error-resilient applications, highly precise input data is not always necessary for them to produce outputs of acceptable quality. This makes memory, the place where data is stored, a suitable component for introducing errors or approximations in return for considerable energy savings. Towards this end, this paper proposes, for the first time, a systematic way for constructing a quality-aware approximate DRAM system. Our design is based upon an extensive experimental characterization of memory errors as a function of the DRAM refresh rate. Leveraging the insights gathered from this characterization, we propose four novel strategies for partitioning the DRAM into a number of quality bins based on the frequency, location, and nature of bit errors in each of the physical pages. During allocation, critical data is placed in the highest quality bin containing only accurate pages and approximate data is allocated to bins sorted in descending order of quality. We validate our proposed scheme on several error-resilient applications implemented using an Altera Stratix IV GX FPGA based Terasic TR4-230 development board containing a 1GB DDR3 DRAM module. Experimental results demonstrate a significant improvement in the energy-quality trade-off compared to previous work and show a reduction in DRAM refresh power of up to 73% with minimal loss in output quality. Arnab Raha, Hrishikesh Jayakumar, Soubhagya Sutar, Vijay Raghunathan |
CASES | 1 |
| 2015 | Quality configurable reduce-and-rank for energy efficient approximate computing
Arnab Raha, Swagath Venkataramani, Vijay Raghunathan, Anand Raghunathan |
DATE | 1 |
| 2015 | VIDalizer: An energy efficient video streamerabstractRecent years have witnessed a significant rise in the number, duration and variety of video contents, which contribute to the bulk of internet traffic. With increase in smartphone and tablet users, watching videos on mobile devices has become one of its most popular use cases. These devices live on limited battery energy which is still a major bottleneck and a source of user dissatisfaction during video playback. In this paper we introduce an intermediate framework called VIDalizer for power efficient video delivery to smartphones and tablets. This almost transparent to the user, battery aware framework takes away some of the video processing overhead from the device and intelligently tunes its parameters customized for the mobile device while delivering the video using a novel transport protocol. Our preliminary results show that this framework can significantly reduce energy consumption up to 45%–55% of a mobile device without compromising user experience. Arnab Raha, Subrata Mitra, Vijay Raghunathan, Sanjay G. Rao |
WCNC | 1 |
| 2015 | QuickRecall: A HW/SW Approach for Computing across Power Cycles in Transiently Powered ComputersabstractTransiently Powered Computers (TPCs) are a new class of batteryless embedded systems that depend solely on energy harvested from external sources for performing computations. Enabling long-running computations on TPCs is a major challenge due to the highly intermittent nature of the power supply (often bursts of < 100ms), resulting in frequent system reboots. Prior work seeks to address this issue by frequently checkpointing system state in flash memory, preserving it across power cycles. However, this involves a substantial overhead due to the high erase/write times of flash memory. This article proposes the use of Ferroelectric RAM (FRAM), an emerging nonvolatile memory technology that combines the benefits of SRAM and flash, to seamlessly enable long-running computations in TPCs. We propose a lightweight, in-situ checkpointing technique for TPCs using FRAM that consumes only 30 nJ while decreasing the time taken for saving and restoring a checkpoint to only 21.06μ s , which is over two orders of magnitude lower than the corresponding overhead using flash. We have implemented and evaluated our technique, Q uick R ecall , using the TI MSP430FR5739 FRAM-enabled microcontroller. Experimental results show that our highly-efficient checkpointing translate to significant speedup (1.25x - 8.4x) in program execution time and reduction (∼3x) in application-level energy consumption. Hrishikesh Jayakumar, Arnab Raha, Woo Suk Lee, Vijay Raghunathan |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2014 | ASLAN: Synthesis of approximate sequential circuitsabstractMany applications produce acceptable results when their underlying computations are executed in an approximate manner. For such applications, approximate circuits enable hardware implementations that exhibit improved efficiency for a given quality. Previous efforts have largely focused on the design of approximate combinational logic blocks such as adders and multipliers. In practice, however, designers are concerned with the quality of outputs generated by a sequential circuit after several cycles of computation, rather than an embedded combinational block. We propose ASLAN (Automatic methodology for Sequential Logic ApproximatioN), the first effort towards the synthesis of approximate sequential circuits. Given a sequential circuit and an output quality constraint, ASLAN creates an approximate version of the circuit that consumes lower energy, while meeting the specified quality bound. The key challenges in approximating sequential circuits are (i) to model how errors due to approximations are generated, re-circulate through the combinational logic over multiple cycles of operation, and eventually impact quality of the final output, and (ii) to select the most beneficial approximations, i.e., those that result in higher energy savings for smaller impact on quality. ASLAN addresses the first challenge by constructing a virtual Sequential Quality Constraint Circuit (SQCC) and utilizing formal verification techniques to ensure that the selected approximations meet the quality constraint. To address the second challenge, ASLAN identifies combinational blocks in the sequential circuit that are amenable to approximation, generates local quality-energy trade-off curves for them, and uses a gradient-descent approach to iteratively approximate the entire sequential circuit. We used ASLAN to automatically synthesize approximate versions of ten sequential benchmarks, resulting in energy reductions of 1.20X-2.44X for tight quality constraints, and 1.32X-4.42X for moderate quality constraints. We present case studies of using the approximate circuits generated by ASLAN in two popular applications - MPEG Encoding and K-Means Clustering - obtaining 1.32X energy savings with 0.5% PSNR degradation, and 1.26X energy savings with 0.8% increase in mean cluster radius, respectively. Ashish Ranjan 0001, Arnab Raha, Swagath Venkataramani, Kaushik Roy 0001, Anand Raghunathan |
DATE | 2 |
| 2014 | Powering the internet of thingsabstractVarious industry forecasts project that, by 2020, there will be around 50 billion devices connected to the Internet of Things (IoT), helping to engineer new solutions to societal-scale problems such as healthcare, energy conservation, transportation, etc. Most of these devices will be wireless due to the expense, inconvenience, or in some cases, the sheer infeasibility of wiring them. Further, many of them will have stringent size constraints. With no cord for power and limited space for a battery, powering these devices (to achieve several months to possibly years of unattended operation) becomes a daunting challenge. This paper highlights some promising directions for addressing this challenge, focusing on three main building blocks: (a) the design of ultra-low power hardware platforms that integrate computing, sensing, storage, and wireless connectivity in a tiny form factor, (b) the development of intelligent system-level power management techniques, and (c) the use of environmental energy harvesting to make IoT devices self-powered, thus decreasing -- in some cases, even eliminating -- their dependence on batteries. We discuss these building blocks in detail and illustrate case-studies of systems that use them judiciously, including the QUBE wireless embedded platform, which exploits the characteristics of emerging non-volatile memory technologies to seamlessly and efficiently enable long-running computations in systems that experience frequent power loss (i.e., intermittently powered systems). Hrishikesh Jayakumar, Kangwoo Lee, Woo Suk Lee, Arnab Raha, Younghyun Kim 0001, Vijay Raghunathan |
ISLPED | 4 |
| 2012 | An optimal sensor deployment scheme to ensure multi level coverage and connectivity in wireless sensor networksabstractThis paper introduces an optimal deployment algorithm of sensors in a given region to provide desired coverage and connectivity for a wireless sensor network. Our paper utilizes two separate procedures for covering different regions of a symmetrical rectangular area. The proposed method divides the given area of interest into two distinct sub-regions termed as the central and edge regions. In each region, a unique scheme is used to determine the number and location of sensors required to monitor and completely cover the region keeping the connectivity and coverage ranges of the sensors, their hardware specification and the dimensions of the region concerned as the constraints. Our scheme reduces the overhead in determining the position of sensors for deployment by following a coverage and connectivity algorithm of lesser complexity rather than those present in related schemes. Finally, we compare our deployment scheme with the interpolation scheme of [2] in regions of different dimensions and different coverage and connectivity levels with different sensing ranges of the sensors to show our cost efficiency over the latter. Arnab Raha, Shovan Maity, Mrinal K. Naskar, Omar Alfandi, Dieter Hogrefe |
IWCMC | 1 |
| 2012 | Fuzzy Logic Election of Node for Routing in WSNsabstractSensor nodes of Wireless Sensor Networks (WSNs) are resource constraints in energy, memory, processing and communication bandwidth. Since they are operated by battery, their life span is limited. Specially, energy conservation is very important issue in the WSN, because it directly affects the life of the node as well as the entire network. Here, we develop a new way of electing a node among many trustworthy nodes for routing processes. This method consumes the energies of network nodes based on Fuzzy logic applied on their residual energy, trust level and distance from the Base Station. The proposed method elects one indispensible node for participating in routing among many worthy nodes. Hence, this method of election of node for routing in WSN sees the conservation of nodes energies go by very smooth and justifying, thereby increasing the life of the WSN. Shaik Sahil Babu, Arnab Raha, Mrinal K. Naskar, Omar Alfandi, Dieter Hogrefe |
TrustCom | 2 |