VLDB 2026 Research / reviewers in the wild / expert
Anup Das 0001
dblp:95/9639 · also Anup K. Das 0001
· DBLP profile ↗
85ranked-venue papers
29as first author
43since 2021 · last 2026
0000-0002-5673-2636ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 67 · 25 first-author · 32 since 2021Software engineering, systems software and programming languages · 19 · 11 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robertha: Eigenspectrum Regularized Attention for Robust Natural Language UnderstandingabstractWe study asymmetric vulnerability to embedding corruption in encoder-based language models, where uniform perturbations dis-proportionately degrade low-magnitude embeddings compared to high-magnitude ones. Since critical words concentrate in low-norm space, this asymmetry causes catastrophic degradation of grammatical and semantic structure even under moderate corruption. Existing robustness approaches either sacrifice clean performance or fail to generalize to higher corruption levels. To address this problem, we propose Robertha, an attention mechanism built on Modern Hopfield Networks, in which semantic patterns act as stable states (attractors) that pull corrupted embeddings toward correct representations. We introduce iterative refinement for differential recovery: heavily corrupted embeddings require multiple convergence steps, while lightly corrupted embeddings converge quickly. To strengthen this mechanism, we introduce Eigenspectrum Regularization (ESR), which enforces low-rank key structures by controlling eigenvalue entropy, creating strong, well separated attractors with wide recovery basins. Across 13 GLUE and SuperGLUE tasks, Robertha significantly outperforms existing robustness methods while maintaining competitive clean performance. Andreia Podasca, Anup Das 0001 |
ACL (1) | 2 |
| 2026 | Sparse Compressed Quantized Dataflow Architecture (QUDA) for Neuromorphic ComputingabstractCompressed Sparse Row (CSR) encoding stores only non-zero elements of a sparse matrix along with indexing information for efficient MVM. We propose QUDA, a heterogeneous 2D dataflow architecture that operates directly on CSR-encoded, quantized weights to improve memory efficiency and utilization. It combines flow-through systolic array (FTSA) columns for partial accumulation with a final output-stationary systolic array (OSSA) column for accumulation and spike generation. Our results show significant improvements in resource utilization and power consumption compared to state-of-the-art designs. Shadi Matinizadeh, Anup Das 0001 |
FCCM | 2 |
| 2025 | Exploring Dendritic Computation in Bio-Inspired Architectures for Dynamic ProgrammingabstractDynamic programming is a classical optimization technique that systematically decomposes a complex problem into simpler sub-problems to find an optimal solution. We explore the use of bio-inspired architectures to find the shortest path between two nodes in a graph using dynamic programming. We leverage dendritic computations, which are linear and nonlinear mechanisms in neuronal dendrites that allow to implement different computational primitives. We exploit two key mechanisms: 1) a dendrite acts as a delay line to propagate an excitatory post-synaptic potential to the soma, and 2) a feedback mechanism from the soma into the dendrites to control this delay. Our key ideas are the following. First, we model each node on a graph as a leaky integrate-and-fire (LIF) neuron, supporting the two dendritic mechanisms. We use a countdown counter to implement forward propagation of a delayed synaptic potential and eligibility trace-based feedback to update the delay by incorporating the cost of edges in a graph. Next, we formulate dynamic programming in terms of the time to the first spike in neurons. We breakdown the shortest path problem into sub-problems of finding the earliest firing times of neurons, and iteratively building the final solution from these smaller sub-problems by tracing backward. We implement this approach for several real-world graphs and show its scalability. We also show early prototyping on a Virtex UltraScale FPGA. Anup Das 0001 |
DATE | 1 |
| 2025 | Online Learning for Dynamic Structural Characterization in Electron Energy Loss SpectroscopyabstractIn-situ Electron Energy Loss Spectroscopy (EELS) is a crucial technique for determining the elemental composition of materials through EELS Spectrum Images (EELS-SI). While recent innovations have made it possible for EELS-SI data acquisition at rates of 400 frames per second with near-zero read noise, the challenge lies in processing this massive stream of real-time data to capture nanoscale dynamic changes. This task demands advanced machine learning methods capable of identifying subtle and complex features in EELS spectra. Furthermore, the EELS data acquired in difficult experimental conditions often suffer from a low signal-to-noise ratio (SNR), leading to unreliable classification and limiting their utility. In response to this critical need, we introduce a spiking neural network (SNN)-based Variational Autoencoder (VAE) that embeds spectral data into a latent space, facilitating precise prediction of structural changes. VAEs are designed to learn efficient low-dimensional representations while capturing the inherent variability in the data, making them highly effective for processing multidimensional data. Additionally, SNNs, which use biological neurons, offer unmatched scalability and energy efficiency by processing information through binary spikes, making them ideal for high-throughput data. We validate our framework using MXene annealing data, achieving denoised spectrum images with an SNR of 28.3dB. For the first time, we present a fully online learning solution for dynamic structural tracking, implemented directly in hardware, eliminating the traditional bottleneck of offline training. Our method achieves reliable, real-time, on-device characterization of high-speed EELS data when evaluated on an FPGA platform. Joint experiments with the SNN-VAE model on both spiking autoencoder hardware and a softwaretrained hybrid configuration of hardware spiking encoders demonstrated latency reductions of 25.2x, 93.7x, and 1.04x, 4.5x in energy savings, respectively, compared to baseline. M. Lakshmi Varshika, Jonathan Hollenbach, Nicolas Bohm Agostini, Ankur Limaye, Antonino Tumeo, Anup Das 0001 |
DATE | 6 |
| 2025 | Hierarchical Model-Based Approach for Concurrent Testing of Neuromorphic ArchitectureabstractNeuromorphic architectures that implement spiking neural networks provide a biologically inspired and energy-efficient approach to processing information. These systems use spike trains, where the timing and frequency of spikes drive computation, offering unique advantages in dynamic and event-driven tasks. This paper develops a concurrent testing methodology for neuromorphic architectures, emphasizing Error Detection and Isolation (EDI) through a hierarchical model-based redundancy framework. Our approach uses a software-based monitoring system that compares the discrepancies between the observed and predicted behavior of hardware-mapped neurons at both the system and the neuron levels. We identify key statistical properties of spike trains that are critical for error detection and develop computationally efficient machine learning models to forecast these properties. By combining real-time observations with predictions of neuron behavior, our EDI methodology ensures robust fault detection and isolation. Experimental evaluations using an open source neuromorphic processor design executing benchmark datasets, MNIST, FashionMNIST, and SVHN, demonstrate the effectiveness. We observe high fault coverage with reduced computational overhead, making the EDI scheme suitable for real-time use in neuromorphic systems. Abhishek Kumar Mishra 0002, Anup Das 0001, Nagarajan Kandasamy |
DSN | 3 |
| 2025 | Optimizing Memory Latency and Bandwidth of Spiking Neural Network Accelerators on FPGA via Sparse HashingabstractSpiking Neural Networks (SNNs) exploit the natural temporal sparsity by processing discrete events. Introducing additional weight sparsity can further improve their energy efficiency when implemented in hardware. We propose SPARSH, a novel memory organization technique that uses sparse hashing to allow efficient storage and fast access while minimizing memory over-provisioning. Through SPARSH, we make the following four key contributions. First, we introduce an efficient hash function to evenly distribute nonzero weights to buckets that are optimized for an FPGA memory system and placed in a compact memory space to achieve storage efficiency. Second, we use a controller that integrates a Bloom filter to detect sparsity and skip memory accesses for weights and activations that are zero, and for neurons that are in refractory state. This improves memory bandwidth utilization. Third, we propose a scheduler that improves memory latency by prioritizing accesses that hit in the same bucket over other accesses. Finally, we propose an algorithm to select the design parameters of SPARSH based on the sparsity of a target SNN. We implement SPARSH for a recent SNN accelerator on a Virtex UltraScale and evaluate using seven SNNs models. We show that SPARSH reduces memory over-provisioning by 3.2×, access latency by 77%, and bandwidth utilization by 68% with a marginal increase in resource utilization. Shadi Matinizadeh, M. Lakshmi Varshika, Anup Das 0001 |
ICCAD | 3 |
| 2025 | Invited Paper: Analyzing the Robustness of Neuromorphic Computing in the Presence of Variability in Non-Volatile MemoryabstractAnalog neuromorphic architectures using nonvolatile memory (NVM)-based crossbar arrays are a promising approach to accelerate deep learning workloads by efficiently performing the core matrix-vector multiplications (MVMs). A crossbar array leverages the current accumulation property of NVM cells to enable parallel MVM computation naturally through basic circuit laws (Ohm’s Law and Kirchhoff’s Current Law). Programming a crossbar for MVM involves mapping the synaptic weights of a deep learning model (i.e., the matrix elements) to conductance values of its NVM cells. Unfortunately, NVMs suffer from variability, including stochastic device-to-device (DDV) and cycle-to-cycle (CCV) variations in conductance-versus-pulse characteristics, which lead to imprecise weight mapping and reduced model performance.We present a novel mathematical framework to analyze the robustness of neuromorphic computing in the presence of variability in NVM. Specifically, we formulate this robustness analysis as finding the minimum perturbation in the feature space (hidden layers) that is sufficient to change the estimated output label, thereby resulting in incorrect classification. We model this as a quadratic programming problem with linear constraints, where we minimize the L2 norm squared of the perturbation subject to crossing the decision boundary to the nearest class. This results in a convex optimization problem with a quadratic objective function and linear inequality constraints. We solve the problem using Lagrangian method with Karush-Kuhn-Tucker (KKT) conditions, finding the minimum perturbation as an orthogonal projection onto the closest decision boundary. The closed-form solution provides both the direction (difference of weight vectors) and magnitude (scaled by confidence difference and inverse squared norm of weight difference) of the minimum perturbation. We present an extensive analysis and evaluation of the mathematical framework on a deep neural network trained on MNIST and Fashion-MNIST datasets. We also present an interpretation of the minimum perturbation, outlining the potential for algorithm-hardware co-design. Andreia Podasca, Anup Das 0001 |
ICCAD | 2 |
| 2025 | Neuromorphic Architectures for Scientific Computing: a Structural Characterization Case StudyabstractNeuromorphic computing offers a promising paradigm for energy-efficient edge processing in scientific applications, such as the real-time analysis of Electron Energy Loss Spectroscopy (EELS) data from Transmission Electron Microscopes (TEMs). Current methods, primarily based on Spiking Variational Autoencoders (S-VAE), are constrained by high computational overhead. To address this, we propose an energy-efficient Spiking Hopfield Network (S-Hopfield) for online encoding and decoding of structural dynamics. Our approach leverages the inherent associative memory of Hopfield networks to robustly denoise and reconstruct spectral images, outperforming an S-VAE model in both image quality metrics and hardware efficiency. Quantitatively, the S-Hopfield network achieved a Mean Squared Error (MSE) of 0.54, a 28% improvement over the S-VAE’s MSE of 0.75. On a Xilinx Virtex-7 FPGA, the S-Hopfield’s core inference engine consumed a mere 0.25 W, representing a 51% reduction in power compared to the S-VAE’s 0.51 W. These results demonstrate that the S-Hopfield network provides a superior, low-power solution for real-time spectral analysis at the edge, paving the way for autonomous experimental control in material science. M. Lakshmi Varshika, Jonathan Hollenbach, Nicolas Bohm Agostini, Ankur Limaye, Marco Minutoli, Vito Giovanni Castellana, Joseph B. Manzano, Anup Das 0001, Mitra Taheri, Antonino Tumeo |
ICCAD | 8 |
| 2025 | A Framework for Automatic Synthesis of Neuromorphic Architectures with Heterogeneous Integration of CMOS and MemristorsabstractA hybrid CMOS-memristor design can significantly enhance the energy efficiency of neuromorphic systems, particularly those implementing spiking neural networks (SNNs). In such a hybrid design, neurons are implemented using CMOS transistors, while synaptic weights are implemented using memristive devices such as resistive RAM (RRAM). We propose a framework for automatic synthesis of such designs at the SPICE level starting from an SNN model defined in a high-level language such as Python. Given the ubiquity of PyTorch in the machine learning community and for demonstration purposes, the frontend of the proposed framework is integrated with a torch-based SNN simulator for model specification and training. Its backend is integrated with a SPICE simulator, e.g., Synopsys HSPICE. We built an open-source application programming interface (API) to compile an SNN model down to its hybrid implementation as a crossbar-based or layer-based microarchitecture, which can subsequently be simulated to verify the design for a wide range of learning tasks and datasets. We show the capability of this framework to perform circuit-oriented design space exploration. Sarah Johari, Arghavan Mohammadhassani, Anup Das 0001 |
ISCAS | 3 |
| 2025 | A Digital Neuromorphic Architecture for Unsupervised Shortest Path Computation on Real-World GraphsabstractGraphs are popular tools for analyzing interconnected data entities. We propose SENTIENCE, a novel approach to computing the shortest path in a graph in an unsupervised manner drawing inspiration from the hippocampus. SENTIENCE uses laterally-connected neurons to represent nodes and synapses to represent edges. The strength of a synaptic connection is encoded as the axonal delay. When a neuron (source) is excited, a wavefront of neural activity is created that propagates through the graph via the connected nodes. SENTIENCE uses the Eligibility Propagation (E-Prop) algorithm to learn the sequence of wave movement through the shortest path in a graph, which can be obtained by identifying earliest firing (eligible) neighbors by tracing back in time from when the wavefront reaches a destination. We introduce a lightweight hardware design for the axonal plasticity and the E-Prop learning mechanism of SENTIENCE. We propose a tile-based architecture to address the scalability of SENTIENCE for real-world graphs with irregular data dependency. We evaluate SENTIENCE on a Versal VPK 180 FPGA and show that SENTIENCE consumes on average 10× less resources and 12× less power compared to a state-of-the-art. Arghavan Mohammadhassani, Shadi Matinizadeh, M. Lakshmi Varshika, Anup Das 0001 |
ISCAS | 4 |
| 2025 | Improving Scalability of NoC-Based Neuromorphic Hardware with Compressed AER (C-AER) ProtocolabstractDigital neuromorphic architectures use the address event representation protocol (AER), which encodes a spike into the address of it’s originating neuron. We show that, this baseline AER protocol is inefficient in network-on-chip (NoC)-based many-core designs, increasing the congestion, delay and energy consumption on the NoC. We introduce CAERPRO, a novel spike encoding protocol for many-core neuromorphic hardware. The key idea of CAERPRO is to organize the output spikes of a core (in space and time) as a single binary spike string and compress it into fewer bits using a static compression technique. We implement CAERPRO for a state-of-the-art NoC-based neuromorphic hardware and synthesize it using Synopsys Design compiler in 32nm technology. Our results using 4 workloads show that, compared to baseline AER protocol, CAERPRO reduces energy consumption on the NoC by an average 14.8× and latency of spike events by an average 9.6×. Furthermore, we show that CAERPRO compression and decompression hardware consume only 2.97% of the total power of a neuromorphic core at 1Mhz clock frequency and introduce a marginal area overhead of 0.26%. Arghavan Mohammadhassani, Sarah Johari, Krupa Tishbi, Anup Das 0001 |
ISCAS | 4 |
| 2025 | Mapping and scheduling spiking neural networks on segmented ladder bus architectures
Phu Khanh Huynh, Francky Catthoor, Anup Das 0001 |
J. Syst. Archit. | 3 |
| 2025 | Design Flow for Scheduling Spiking Deep Convolutional Neural Networks on Heterogeneous Neuromorphic System-on-chipabstractNeuromorphic systems-on-chip (NSoCs) integrate CPU cores and neuromorphic hardware accelerators on the same chip. These platforms can execute spiking deep convolutional neural networks (SDCNNs) with a low energy footprint. Modern NSoCs are heterogeneous in terms of their computing, communication, and storage resources. This makes scheduling SDCNN operations a combinatorial problem of exploring an exponentially large state space in determining mapping, ordering, and timing of operations to achieve a target hardware performance, e.g., throughput. We propose a systematic design flow to schedule SDCNNs on an NSoC. Our scheduler, called SMART (SDCNN MApping, OrdeRing, and Timing), branches the combinatorial optimization problem into computationally relaxed sub-problems that generate fast solutions without significantly compromising the solution quality. SMART improves performance by efficiently incorporating the heterogeneity in computing, communication, and storage resources. SMART operates in four steps. First, it creates a self-timed execution schedule to map operations to compute resources, maximizing throughput. Second, it uses an optimization strategy to distribute activation and synaptic weights to storage resources, minimizing data communication-related overhead. Third, it constructs an inter-processor communication (IPC) graph with a transaction order for its communication actors. This transaction order is created using a transaction partial order algorithm, which minimizes contention on the shared communication resources. Finally, it schedules this IPC graph to hardware by overlapping communication with the computation, and leveraging operation, pipeline, and batch parallelism. We evaluate SMART using 10 representative image, object, and language-based SDCNNs. Results show that SMART increases throughput by an average 23%, compared to a state-of-the-art scheduler. SMART is implemented entirely in software as a compiler extension. It does not require any change in a neuromorphic hardware or its interface to CPUs. It improves throughput with only a marginal increase in the compilation time. SMART is released under the open-source MIT licensing at https://github.com/drexel-DISCO/SMART to foster future research. Anup Das 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | Clustering and Allocation of Spiking Neural Networks on Crossbar-Based Neuromorphic ArchitectureabstractNeuromorphic hardware, designed to mimic the neural structure of the human brain, offers an energy-efficient platform for implementing machine-learning models in the form of Spiking Neural Networks (SNNs). Achieving efficient SNN execution on this hardware requires careful consideration of various objectives, such as optimizing utilization of individual neuromorphic cores and minimizing inter-core communication. Unlike previous approaches that overlooked the architecture of the neuromorphic core when clustering the SNN into smaller networks, our approach uses architecture-aware algorithms to ensure that the resulting clusters can be effectively mapped to the core. We base our approach on a crossbar architecture for each neuromorphic core. We start with a basic architecture where neurons can only be mapped to the columns of the crossbar. Our technique partitions the SNN into clusters of neurons and synapses, ensuring that each cluster fits within the crossbar's confines, and when multiple clusters are allocated to a single crossbar, we maximize resource utilization by efficiently reusing crossbar resources. We then expand this technique to accommodate an enhanced architecture that allows neurons to be mapped not only to the crossbar's columns but also to its rows, with the aim of further optimizing utilization. To evaluate the performance of these techniques, assuming a multi-core neuromorphic architecture, we assess factors such as the number of crossbars used and the average crossbar utilization. Our evaluation includes both synthetically generated SNNs and spiking versions of well-known machine-learning models: LeNet, AlexNet, DenseNet, and ResNet. We also investigate how the structure of the SNN impacts solution quality and discuss approaches to improve it. Ilknur Mustafazade, Nagarajan Kandasamy, Anup Das 0001 |
CF | 3 |
| 2024 | An Open-Source And Extensible Framework for Fast Prototyping and Benchmarking of Spiking Neural Network HardwareabstractSpiking neural networks (SNNs) are bioplausible machine learning models that use discrete spikes to encode, compute, and transmit information. Combined with event-driven low-power hardware, SNNs can improve the energy efficiency of learning tasks. Although there have been several efforts to build SNN hardware, there is no uniform framework to verify and benchmark these designs in terms of key hardware performance metrics such as inference accuracy, area, power consumption, and throughput. We propose PRONTO, an open-source and extensible framework to verify SNN hardware for different learning tasks and datasets. Given the ubiquity of PyTorch in the machine learning community and for demonstration purposes, the frontend of PRONTO is integrated with a torch-based SNN simulator for model specification and training. Its backend is integrated with an open-source quantized SNN hardware. PRONTO interfaces with a torch code to generate input stimuli which are then driven to SNN hardware through a configurable SystemVerilog testbench, verifying the design across various SNN-specific configurations. PRONTO utilizes a dataflow-based approach to validate SNN models that are segmented and run on a mix of software and hardware platforms. We describe PRONTO and evaluate it using six datasets spanning image, audio, and text classification. We present benchmark results for various input settings. PRONTO is available under an open-source licensing to provide a platform to evaluate all current and future SNN hardware designs. We believe PRONTO will substantially reduce the design verification effort, thus facilitating fast design prototyping. Shadi Matinizadeh, Anup Das 0001 |
FPL | 2 |
| 2024 | Sparsity Aware Learning in Feedback-Driven Differential Recurrent Neural Networks
Ankita Paul, Anup Das 0001 |
ICANN (4) | 2 |
| 2024 | Learning in Recurrent Spiking Neural Networks with Sparse Full-FORCE Training
Ankita Paul, Anup Das 0001 |
ICANN (10) | 2 |
| 2024 | Neuromorphic Computing for Graph AnalyticsabstractFinding the single-source shortest paths (SSSP) in a weighted graph is a fundamental problem in computer science with many practical applications, including network flow and social network analysis. Unfortunately, fast and efficient SSSP computations still remain a major bottleneck for shared memory systems, even with multiple processing cores. In this work, our aim is to address SSSP and other graph analytics using neuromorphic computing, which is an emerging computing paradigm inspired by the computations in the mammalian brain. In our architecture, nodes of a graph are modeled as integrate-and-fire (IF) neurons, while edges are modeled as synaptic connections. Unlike a conventional neuromorphic system that uses electrical synapses where the synaptic strength of a connection is represented using a weight, we propose a novel approach where the synaptic strength is represented using the speed of propagation of an action potential (spike). Therefore, edge capacities are modeled as axonal delays in our architecture and are modulated by the precise timing of spikes (plasticity). When a spike is injected into a neuron representing the source node, it induces a spike to its neighbors (connected neurons) after a delay, where the delay is proportional to the capacity of the corresponding edge. These neurons in turn induce spikes to their respective neighbors, creating a spiking wavefront that propagates through the architecture. Anup Das 0001 |
ICCAD | 1 |
| 2024 | Data Driven Learning of Aperiodic Nonlinear Dynamic Systems Using Spike Based Reservoirs-in-ReservoirabstractMimicking an aperiodic nonlinear dynamic system is challenging as it is difficult to represent it using closed-form equations. A feedback-driven spike-based recurrent spiking neural network is a powerful computational model that can mimic such dynamical systems. We propose reservoirs-in-reservoir (R-i-R), a novel architecture to mimic the frequent pattern changes in space and time of an aperiodic nonlinear dynamic system. Here, a large reservoir is built by connecting multiple small reservoirs to a common output. These small reservoirs are individually specialized to mimic a portion of the input dynamic. The internal recurrent connections of each reservoir and its readout are trained using a recursive least squares (RLS)-based full first-order and reduced control error (full-FORCE) algorithm. To make the entire R-i-R architecture adaptable to the change in periodicity of an input, we implement a new cost function that incorporates a unique forgetting factor to control the fading and wind-up of the covariance matrix of each reservoir during training. We evaluate R-i-R using seven aperiodic nonlinear dynamic systems. We show that R-i-R with rate encoding reduces the error rate by an average 59% with 1.8X reduction in network size compared to state-of-the-art. To improve energy efficiency, we implement a time-to-first-spike encoding and show an average reduction of 39. 5% in the number of spikes. Ankita Paul, Nagarajan Kandasamy, Kapil R. Dandekar, Anup Das 0001 |
IJCNN | 4 |
| 2024 | Wafer2Spike: Spiking Neural Network for Wafer Map Pattern ClassificationabstractIn integrated circuit design, the analysis of wafer map patterns is critical to improve yield and detect manufacturing issues. We develop Wafer2Spike, an architecture for wafer map pattern classification using a spiking neural network (SNN), and demonstrate that a well-trained SNN achieves superior performance compared to deep neural network-based solutions. Wafer2Spike achieves an average classification accuracy of 98% on the WM-811k wafer benchmark dataset. It is also superior to existing approaches for classifying defect patterns that are underrepresented in the original dataset. Wafer2Spike achieves this improved precision with great computational efficiency. Abhishek Kumar Mishra 0002, Anush Niranjan Lingamoorthy, Anup Das 0001, Nagarajan Kandasamy |
ITC | 4 |
| 2024 | Model-Based Approach Towards Correctness Checking of Neuromorphic Computing SystemsabstractNeuromorphic hardware that emulates the neural structure of the human brain can implement machine learning models in an extremely energy-efficient manner. It is especially suitable for executing spiking neural networks (SNNs) which comprise spiking neurons interconnected via synapses. The underlying computation is based on spike trains in which the location and frequency of spikes that occur within the network guide the execution. This paper develops a fault detection and isolation (FDI) methodology to monitor the correctness of a neuromorphic program’s execution using model-based redundancy in which a software-based monitor compares discrepancies between the behavior of neurons mapped to hardware and that predicted by a corresponding mathematical model. We identify properties of spike trains generated by neurons that can be used for fault detection and build machine learning models to forecast these properties. Predictions from these models, which describe the nominal behavior of neurons, when combined with real-time observations, form the basis for FDI. Experiments using CARLSim, a high-fidelity SNN simulator, show that the proposed approach achieves high fault coverage using models that can operate with low computational overhead in real time. Abhishek Kumar Mishra 0002, Anup Das 0001, Nagarajan Kandasamy |
PRDC | 2 |
| 2024 | Efficient Built-In Self-Test Strategy for Neuromorphic Hardware Based On Alarm PlacementabstractNeuromorphic hardware that mimics the neural structure of the brain can efficiently run spiking neural networks (SNNs) to perform various machine learning tasks. Built-in self-test capability (BIST) is critical to ensure the reliability and functionality of these systems when deployed in mission-critical applications such as autonomous vehicles. We introduce an online BIST strategy for neuromorphic hardware that aims to maximize fault coverage while reducing the testing time needed to detect and isolate faulty components. Our key contribution is the adaptation of the alarm placement problem for this purpose. The topology of a typical SNN allows for multiple fault propagation paths through the network, and we take advantage of this property to test (or place alarms on) only the minimum number of neurons necessary to isolate any faulty neuron in the SNN. The alarms themselves detect erroneous behavior by measuring the dissimilarity between the observed and expected spike trains generated by a neuron based on the applied test pattern. The efficacy of the BIST approach is evaluated using multiple SNN models in single- and multiple-fault scenarios. We also evaluate different spike dissimilarity metrics in terms of their fault-detection effectiveness. Ilknur Mustafazade, Anup Das 0001, Nagarajan Kandasamy |
PRDC | 2 |
| 2024 | WaferCap: Open Classification of Wafer Map Patterns using Deep Capsule NetworkabstractIn integrated circuit design, analysis of wafer map patterns is critical to enhance yield and detect manufacturing issues. With the emergence of novel wafer map patterns, there is increasing need for robust artificial intelligence models that can both accurately classify seen patterns and while also detecting ones not seen during training, a capability known as open world classification. We develop a novel solution to this problem: WaferCap, a Deep Capsule Network designed for wafer map pattern classification and equipped with a rejection mechanism. When evaluated using the WM-811k dataset, WaferCap significantly surpasses existing methods, achieving 99% accuracy for fully seen patterns while demonstrating robust performance in open-world settings by effectively detecting unseen wafer map patterns. Abhishek Kumar Mishra 0002, Mohammad Ershad Shaik, Anush Niranjan Lingamoorthy, Anup Das 0001, Nagarajan Kandasamy, Nur A. Touba |
VTS | 5 |
| 2023 | Hardware-Software Co-Design for On-Chip Learning in AI SystemsabstractSpike-based convolutional neural networks (CNNs) are empowered with on-chip learning in their convolution layers, enabling the layer to learn to detect features by combining those extracted in the previous layer. We propose ECHELON, a generalized design template for a tile-based neuromorphic hardware with on-chip learning capabilities. Each tile in ECHELON consists of a neural processing units (NPU) to implement convolution and dense layers of a CNN model, an on-chip learning unit (OLU) to facilitate spike-timing dependent plasticity (STDP) in the convolution layer, and a special function unit (SFU) to implement other CNN functions such as pooling, concatenation, and residual computation. These tile resources are interconnected using a shared bus, which is segmented and configured via the software to facilitate parallel communication inside the tile. Tiles are themselves interconnected using a classical Network-on-Chip (NoC) interconnect. We propose a system software to map CNN models to ECHELON, maximizing the performance. We integrate the hardware design and software optimization within a co-design loop to obtain the hardware and software architectures for a target CNN, satisfying both performance and resource constraints. In this preliminary work, we show the implementation of a tile on a FPGA and some early evaluations. Using 8 STDP-enabled CNN models, we show the potential of our co-design methodology to optimize hardware resources. M. Lakshmi Varshika, Abhishek Kumar Mishra 0002, Nagarajan Kandasamy, Anup Das 0001 |
ASP-DAC | 4 |
| 2023 | Preserving Privacy of Neuromorphic Hardware From PCIe Congestion Side-Channel AttackabstractNeuromorphic systems are equipped with software-managed scratchpad to cache intermediate results and synaptic weights of a machine learning model. PCIe (Peripheral Component Interconnect Express) is the de facto protocol to interface between scratchpad and main memory. Congestion happens when PCIe traffic overwhelms the PCIe link capacity. This introduces transmission delay, which not only impacts model performance but also leaks sensitive information about a user (the victim).In this paper, we show that inefficient data placement in scratchpad using state-of-the-art compilers may trigger significant data movement over PCIe. An attacker can measure the PCIe congestion to indirectly infer the victim’s model. Therefore, the delay from PCIe congestion can be exploited as a side-channel.We propose a compiler extension to intelligently manage scratchpad in order to improve model privacy. First, we formulate a design metric to assess the vulnerability of a model to PCIe congestion side-channel attack. Next, we propose an optimization strategy integrated within the compiler to identify contents that should be retained inside scratchpad to minimize this design metric. Finally, we propose a Hill Climbing heuristic to allocate model operations to neuromorphic tiles and improve privacy by efficiently utilizing their on-chip scratchpad capacity.We evaluate our privacy-preserving model execution (PrivacyX) to mitigate PCIe congestion side-channel attack using one attack scenario and 16 image, object, and language-based machine learning models. We show that PrivacyX significantly reduces the vulnerability of a model to PCIe congestion side-channel attack compared to baseline compilers. We also show that PrivacyX, which is managed entirely in software, is complementary to several hardware-based privacy preserving solutions. Anup Das 0001 |
COMPSAC | 1 |
| 2023 | Online Performance Monitoring of Neuromorphic Computing SystemsabstractNeuromorphic computation is based on spike trains in which the location and frequency of spikes occurring within the network guide the execution. This paper develops a frame-work to monitor the correctness of a neuromorphic program’s execution using model-based redundancy in which a software-based monitor compares discrepancies between the behavior of neurons mapped to hardware and that predicted by a corresponding mathematical model in real time. Our approach reduces the hardware overhead needed to support the monitoring infrastructure and minimizes intrusion on the executing application. Fault-injection experiments utilizing CARLSim, a high-fidelity SNN simulator, show that the framework achieves high fault coverage using parsimonious models which can operate with low computational overhead in real time. Abhishek Kumar Mishra 0002, Anup Das 0001, Nagarajan Kandasamy |
ETS | 2 |
| 2023 | ACM TECS Special Issue on Embedded System Security TutorialsabstractNo abstract available. Aviral Shrivastava, Jian-Jia Chen, Akash Kumar 0001, Anup Das 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | A design methodology for fault-tolerant computing using astrocyte neural networksabstractWe propose a design methodology to facilitate fault tolerance of deep learning models. First, we implement a many-core fault-tolerant neuromorphic hardware design, where neuron and synapse circuitries in each neuromorphic core are enclosed with astrocyte circuitries, the star-shaped glial cells of the brain that facilitate self-repair by restoring the spike firing frequency of a failed neuron using a closed-loop retrograde feedback signal. Next, we introduce astrocytes in a deep learning model to achieve the required degree of tolerance to hardware faults. Finally, we use a system software to partition the astrocyte-enabled model into clusters and implement them on the proposed fault-tolerant neuromorphic design. We evaluate this design methodology using seven deep learning inference models and show that it is both area- and power-efficient. Murat Isik, Ankita Paul, M. Lakshmi Varshika, Anup Das 0001 |
CF | 4 |
| 2022 | Design of Many-Core Big Little µBrains for Energy-Efficient Embedded Neuromorphic ComputingabstractAs spiking-based deep learning inference applications are increasing in embedded systems, these systems tend to integrate neuromorphic accelerators such as µBrain to improve energy efficiency. We propose a µBrain-based scalable many-core neuromorphic hardware design to accelerate the computations of spiking deep convolutional neural networks (SDCNNs). To increase energy efficiency, cores are designed to be heterogeneous in terms of their neuron and synapse capacity (i.e., big vs. little cores), and they are interconnected using a parallel segmented bus interconnect, which leads to lower latency and energy compared to a traditional mesh-based Network-on-Chip (NoC). We propose a system software framework called SentryOS to map SDCNN inference applications to the proposed design. SentryOS consists of a compiler and a run-time manager. The compiler compiles an SDCNN application into sub-networks by exploiting the internal architecture of big and little µBrain cores. The run-time manager schedules these sub-networks onto cores and pipeline their execution to improve throughput. We evaluate the proposed big little many-core neuromorphic design and the system software framework with five commonly-used SDCNN inference applications and show that the proposed solution reduces energy (between 37% and 98%), reduces latency (between 9% and 25%), and increases application throughput (between 20% and 36%). We also show that SentryOS can be easily extended for other spiking neuromorphic accelerators such as Loihi and DYNAPs. M. Lakshmi Varshika, Adarsha Balaji, Federico Corradi, Anup Das 0001, Jan Stuijt, Francky Catthoor |
DATE | 4 |
| 2022 | Multiscale Voxel Based Decoding For Enhanced Natural Image Reconstruction From Brain ActivityabstractReconstructing perceived images from human brain activity monitored by functional magnetic resonance imaging (fMRI) is hard, especially for natural images. Existing methods often result in blurry and unintelligible reconstructions with low fidelity. In this study, we present a novel approach for enhanced image reconstruction, in which existing methods for object decoding and image reconstruction are merged together. This is achieved by conditioning the reconstructed image to its decoded image category using a class-conditional generative adversarial network and neural style transfer. The results indicate that our approach improves the semantic similarity of the reconstructed images and can be used as a general framework for enhanced image reconstruction. Mali Halac, Murat Isik, Hasan Ayaz, Anup Das 0001 |
IJCNN | 4 |
| 2022 | CARLsim 6: An Open Source Library for Large-Scale, Biologically Detailed Spiking Neural Network SimulationabstractMature simulation systems for Spiking Neural Networks (SNNs) become more relevant than ever for understanding the brain and supporting neuromorphic computing. The CARL-sim SNN platform is one of the first Open Source simulation systems that utilized CUDA GPUs to address the tremendous parallel processing demands of natural brains. It has evolved over almost a decade in numerous scientific research projects requiring efficient biologically plausible modeling at scale. With its sixth major release, CARLsim 6 respects this legacy by supporting the latest versions of operating systems, development tool chains, multi-core computers, and of course GPUs. It runs on a range of platforms; from Notebooks up to the NVIDIA DGX-A100 supercomputer, and is used in biologically plausible simulations of the hippocampus and neocortex. The latest version has added flexibility for incorporating long-term and short-term synaptic plasticity. Neuromodulation is an important property of neurobiology that can lead to rapid few shot learning, network rewiring, and neural activity modulation. Because of this, CARLsim 6 now supports four multiple neuromodulators for simulating neural excitability and synaptic plasticity. Lars Niedermeier, Kexin Chen 0002, Jinwei Xing, Anup Das 0001, Jeffrey Kopsick, Eric Scott, Nate Sutton, Killian Weber, Nikil Dutt, Jeffrey L. Krichmar |
IJCNN | 4 |
| 2022 | Learning in Feedback-driven Recurrent Spiking Neural Networks using full-FORCE TrainingabstractFeedback-driven recurrent spiking neural networks (RSNNs) are powerful computational models that can mimic dynamic systems. However, the presence of a feedback loop from the readout to the recurrent layer de-stabilizes the learning mechanism and prevents it from converging. Here, we propose a supervised training procedure for RSNNs, where a second network is introduced only during the training, to provide hint for the target dynamics. The proposed training procedure consists of generating targets for both recurrent and readout layers (i.e., for a full RSNN system). It uses the recursive least square-based First-Order and Reduced Control Error (FORCE) algorithm to fit the activity of each layer to its target. The proposed full-FORCE training procedure reduces the amount of modifications needed to keep the error between the output and target close to zero. These modifications control the feedback loop, which causes the training to converge. We demonstrate the improved performance and noise robustness of the proposed full-FORCE training procedure to model 8 dynamical systems using RSNNs with leaky integrate and fire (LIF) neurons and rate coding. For energy-efficient hardware implementation, an alternative time-to-first-spike (TTFS) coding is implemented for the full-FORCE training procedure. Compared to rate coding, full-FO RCE with TTFS coding generates fewer spikes and facilitates faster convergence to the target dynamics. Ankita Paul, Stefan Wagner 0021, Anup Das 0001 |
IJCNN | 3 |
| 2022 | Real-Time Scheduling of Machine Learning Operations on Heterogeneous Neuromorphic SoCabstractNeuromorphic Systems-on-Chip (NSoCs) are becoming heterogeneous by integrating general-purpose processors (GPPs) and neural processing units (NPUs) on the same SoC. For embedded systems, an NSoC may need to execute user applications built using a variety of machine learning models. We propose a real-time scheduler, called PRISM, which can schedule machine learning models on a heterogeneous NSoC either individually or concurrently to improve their system performance. PRISM consists of the following four key steps. First, it constructs an interprocessor communication (IPC) graph of a machine learning model from a mapping and a self-timed schedule. Second, it creates a transaction order for the communication actors and embeds this order into the IPC graph. Third, it schedules the graph on an NSoC by overlapping communication with the computation. Finally, it uses a Hill Climbing heuristic to explore the design space of mapping operations on GPPs and NPUs to improve the performance. Unlike existing schedulers which use only the NPUs of an NSoC, PRISM improves performance by enabling batch, pipeline, and operation parallelism via exploiting a platform's heterogeneity. For use-cases with concurrent applications, PRISM uses a heuristic resource sharing strategy and a non-preemptive scheduling to reduce the expected wait time before concurrent operations can be scheduled on contending resources. Our extensive evaluations with 20 machine learning workloads show that PRISM significantly improves the performance per watt for both individual applications and use-cases when compared to state-of-the-art schedulers. Anup Das 0001 |
MEMOCODE | 1 |
| 2022 | Design-Technology Co-Optimization for NVM-Based Neuromorphic Processing ElementsabstractAn emerging use case of machine learning (ML) is to train a model on a high-performance system and deploy the trained model on energy-constrained embedded systems. Neuromorphic hardware platforms, which operate on principles of the biological brain, can significantly lower the energy overhead of an ML inference task, making these platforms an attractive solution for embedded ML systems. We present a design-technology tradeoff analysis to implement such inference tasks on the processing elements (PEs) of a non-volatile memory (NVM)-based neuromorphic hardware. Through detailed circuit-level simulations at scaled process technology nodes, we show the negative impact of technology scaling on the information-processing latency, which impacts the quality of service of an embedded ML system. At a finer granularity, the latency inside a PE depends on (1) the delay introduced by parasitic components on its current paths, and (2) the varying delay to sense different resistance states of its NVM cells. Based on these two observations, we make the following three contributions. First, on the technology front, we propose an optimization scheme where the NVM resistance state that takes the longest time to sense is set on current paths having the least delay, and vice versa, reducing the average PE latency, which improves the quality of service. Second, on the architecture front, we introduce isolation transistors within each PE to partition it into regions that can be individually power-gated, reducing both latency and energy. Finally, on the system-software front, we propose a mechanism to leverage the proposed technological and architectural enhancements when implementing an ML inference task on neuromorphic PEs of the hardware. Evaluations with a recent neuromorphic hardware architecture show that our proposed design-technology co-optimization approach improves both performance and energy efficiency of ML inference tasks without incurring high cost-per-bit. Shihao Song, Adarsha Balaji, Anup Das 0001, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic HardwareabstractSpiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches. Shihao Song, Harry Chong, Adarsha Balaji, Anup Das 0001, James A. Shackleford, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | Endurance-Aware Mapping of Spiking Neural Networks to Neuromorphic HardwareabstractNeuromorphic computing systems are embracing memristors to implement high density and low power synaptic storage as crossbar arrays in hardware. These systems are energy efficient in executing Spiking Neural Networks (SNNs). We observe that long bitlines and wordlines in a memristive crossbar are a major source of parasitic voltage drops, which create current asymmetry. Through circuit simulations, we show the significant endurance variation that results from this asymmetry. Therefore, if the critical memristors (ones with lower endurance) are overutilized, they may lead to a reduction of the crossbar's lifetime. We propose eSpine, a novel technique to improve lifetime by incorporating the endurance variation within each crossbar in mapping machine learning workloads, ensuring that synapses with higher activation are always implemented on memristors with higher endurance, and vice versa. eSpine works in two steps. First, it uses the Kernighan-Lin Graph Partitioning algorithm to partition a workload into clusters of neurons and synapses, where each cluster can fit in a crossbar. Second, it uses an instance of Particle Swarm Optimization (PSO) to map clusters to tiles, where the placement of synapses of a cluster to memristors of a crossbar is performed by analyzing their activation within the workload. We evaluate eSpine for a state-of-the-art neuromorphic hardware model with phase-change memory (PCM)-based memristors. Using 10 SNN workloads, we demonstrate a significant improvement in the effective lifetime. Twisha Titirsha, Shihao Song, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Nagarajan Kandasamy, Francky Catthoor |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Improving Inference Lifetime of Neuromorphic Systems via Intelligent Synapse MappingabstractNon-Volatile Memories (NVMs) such as Resistive RAM (RRAM) are used In neuromorphic systems to Implement high-density and low-power analog synaptic weights. Unfortunately, an RRAM cell can switch its state after reading its content a certain number of times. Such behavior challenges the integrity and program-onee-read-many-times philosophy of implementing machine learning inference on neuromorphic systems, impacting the Quality-of-Serviee (QoS). Elevated temperatures and frequent usage can significantly shorten the number of times an RRAM cell can be reliably read before it becomes absolutely necessary to reprogram. We propose an architectural solution to extend the read endurance of RRAM-based neuromorphic systems. We make two key contributions. First, we formulate the read endurance of an RRAM cell as a function of the programmed synaptic weight and its activation within a machine learning workload. Second, we propose an intelligent workload mapping strategy incorporating the endurance formulation to place the synapses of a machine learning model onto the RRAM cells of the hardware. The objective is to extend the inference lifetime, defined as the number of times the model can be used to generate output (inference) before the trained weights need to be reprogrammed on the RRAM cells of the system. We evaluate our architectural solution with machine learning workloads on a eyele-aeeurate simulator of an RRAM-based neuromorphic system. Our results demonstrate a significant increase in inference lifetime with only a minimal performance impact. Shihao Song, Twisha Titirsha, Anup Das 0001 |
ASAP | 3 |
| 2021 | Aging-Aware Request Scheduling for Non-Volatile Main MemoryabstractModern computing systems are embracing non-volatile memory (NVM) to implement high-capacity and low-cost main memory. Elevated operating voltages of NVM accelerate the aging of CMOS transistors in the peripheral circuitry of each memory bank. Aggressive device scaling increases power density and temperature, which further accelerates aging, challenging the reliable operation of NVM-based main memory. We propose HEBE, an architectural technique to mitigate the circuit aging-related problems of NVM-based main memory. HEBE is built on three contributions. First, we propose a new analytical model that can dynamically track the aging in the peripheral circuitry of each memory bank based on the bank's utilization. Second, we develop an intelligent memory request scheduler that exploits this aging model at run time to de-stress the peripheral circuitry of a memory bank only when its aging exceeds a critical threshold. Third, we introduce an isolation transistor to decouple parts of a peripheral circuit operating at different voltages, allowing the decoupled logic blocks to undergo long-latency de-stress operations independently and off the critical path of memory read and write accesses, improving performance. We evaluate HEBE with workloads from the SPEC CPU2017 Benchmark suite. Our results show that HEBE significantly improves both performance and lifetime of NVM-based main memory. Shihao Song, Anup Das 0001, Onur Mutlu, Nagarajan Kandasamy |
ASP-DAC | 2 |
| 2021 | On the role of system software in energy management of neuromorphic computingabstractNeuromorphic computing systems such as DYNAPs and Loihi have recently been introduced to the computing community to improve performance and energy efficiency of machine learning programs, especially those that are implemented using Spiking Neural Network (SNN). The role of a system software for neuromorphic systems is to cluster a large machine learning model (e.g., with many neurons and synapses) and map these clusters to the computing resources of the hardware. In this work, we formulate the energy consumption of a neuromorphic hardware, considering the power consumed by neurons and synapses, and the energy consumed in communicating spikes on the interconnect. Based on such formulation, we first evaluate the role of a system software in managing the energy consumption of neuromorphic systems. Next, we formulate a simple heuristic-based mapping approach to place the neurons and synapses onto the computing resources to reduce energy consumption. We evaluate our approach with 10 machine learning applications and demonstrate that the proposed mapping approach leads to a significant reduction of energy consumption of neuromorphic computing systems. Twisha Titirsha, Shihao Song, Adarsha Balaji, Anup Das 0001 |
CF | 4 |
| 2021 | Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning AcceleratorsabstractThe growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion in availability of domain-specific accelerators, which struggle to support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a tool to quickly and automatically transition from algorithm definition to hardware implementation and explore the design space along a variety of SWaP (size, weight and Power) metrics. The software defined architectures (SODA) synthesizer implements a modular compiler-based infrastructure for the end-to-end generation of machine learning accelerators, from high-level frameworks to hardware description language. Neuromorphic computing, mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders-of-magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lacks the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper, we discuss the design of such an integrated generator, leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect in the SODA frontend that allows expressing spiking neural network concepts (e.g., spiking sequences, transformation, and manipulation) and we discuss how to enable the mapping of spiking neurons to the related specialized hardware (which could be generated through middle-end and backend layers of the SODA Synthesizer). We then discuss the opportunities for further integration offered by the hardware compilation infrastructure, providing a path towards the generation of complex hybrid artificial intelligence systems. Serena Curzel, Nicolas Bohm Agostini, Shihao Song, Ismet Dagli, Ankur Limaye, Cheng Tan 0002, Marco Minutoli, Vito Giovanni Castellana, Vinay Amatya, Joseph B. Manzano, Anup Das 0001, Fabrizio Ferrandi, Antonino Tumeo |
ICCAD | 11 |
| 2021 | A Design Flow for Mapping Spiking Neural Networks to Many-Core Neuromorphic HardwareabstractThe design of many-core neuromorphic hardware is becoming increasingly complex as these systems are now expected to execute large machine-learning models. A predictable design flow is needed to guarantee real-time performance such as latency and throughput without significantly increasing the buffer requirement of computing cores. Synchronous Data Flow Graphs (SDFGs) have been previously used for predictable mapping of streaming applications to multiprocessor systems. We propose an SDFG-based design flow to map spiking neural networks (SNNs) to many-core neuromorphic hardware with the objective of exploring the tradeoff between throughput and buffer-size requirements. The proposed design flow integrates an iterative partitioning approach based on Kernighan-Lin graph partitioning heuristic to create SNN clusters such that each cluster can be mapped to a core of the hardware. The partitioning approach minimizes inter-cluster spike communication, which improves latency on the shared interconnect of the hardware. Next, the design flow maps clusters to cores using Particle Swarm Optimization (PSO), an evolutionary algorithm, while exploring the design space of throughput and buffer size. Pareto-optimal mappings are retained from the design flow, allowing system designers to select a Pareto mapping that satisfies throughput and buffer-size requirements of the design. We evaluated the developed design flow using five large-scale convolutional neural network (CNN) models. Results demonstrate 63% higher maximum throughput and 10% lower buffer-size requirement compared to state-of-the-art dataflow-based mapping solutions. Shihao Song, M. Lakshmi Varshika, Anup Das 0001, Nagarajan Kandasamy |
ICCAD | 3 |
| 2021 | Special Session: Reliability Analysis for AI/ML HardwareabstractArtificial intelligence (AI) and Machine Learning (ML) are becoming pervasive in today's applications, such as autonomous vehicles, healthcare, aerospace, cybersecurity, and many critical applications. Ensuring the reliability and robustness of the underlying AI/ML hardware becomes our paramount importance. In this paper, we explore and evaluate the reliability of different AI/ML hardware. The first section outlines the reliability issues in a commercial systolic array-based ML accelerator in the presence of faults engendering from device-level non-idealities in the DRAM. Next, we quantified the impact of circuit-level faults in the MSB and LSB logic cones of the Multiply and Accumulate (MAC) block of the AI accelerator on the AI/ML accuracy. Finally, we present two key reliability issues- circuit aging and endurance in emerging neuromorphic hardware platforms and present our system-level approach to mitigate them. Shamik Kundu, Kanad Basu, Mehdi Sadi, Twisha Titirsha, Shihao Song, Anup Das 0001, Ujjwal Guin |
VTS | 6 |
| 2021 | Dynamic Reliability Management in Neuromorphic ComputingabstractNeuromorphic computing systems execute machine learning tasks designed with spiking neural networks. These systems are embracing non-volatile memory to implement high-density and low-energy synaptic storage. Elevated voltages and currents needed to operate non-volatile memories cause aging of CMOS-based transistors in each neuron and synapse circuit in the hardware, drifting the transistor’s parameters from their nominal values. If these circuits are used continuously for too long, the parameter drifts cannot be reversed, resulting in permanent degradation of circuit performance over time, eventually leading to hardware faults. Aggressive device scaling increases power density and temperature, which further accelerates the aging, challenging the reliable operation of neuromorphic systems. Existing reliability-oriented techniques periodically de-stress all neuron and synapse circuits in the hardware at fixed intervals, assuming worst-case operating conditions, without actually tracking their aging at run-time. To de-stress these circuits, normal operation must be interrupted, which introduces latency in spike generation and propagation, impacting the inter-spike interval and hence, performance (e.g., accuracy). We observe that in contrast to long-term aging, which permanently damages the hardware, short-term aging in scaled CMOS transistors is mostly due to bias temperature instability. The latter is heavily workload-dependent and, more importantly, partially reversible. We propose a new architectural technique to mitigate the aging-related reliability problems in neuromorphic systems by designing an intelligent run-time manager (NCRTM), which dynamically de-stresses neuron and synapse circuits in response to the short-term aging in their CMOS transistors during the execution of machine learning workloads, with the objective of meeting a reliability target. NCRTM de-stresses these circuits only when it is absolutely necessary to do so, otherwise reducing the performance impact by scheduling de-stress operations off the critical path. We evaluate NCRTM with state-of-the-art machine learning workloads on a neuromorphic hardware. Our results demonstrate that NCRTM significantly improves the reliability of neuromorphic hardware, with marginal impact on performance. Shihao Song, Jui Hanamshet, Adarsha Balaji, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Nagarajan Kandasamy, Francky Catthoor |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2020 | PyCARL: A PyNN Interface for Hardware-Software Co-Simulation of Spiking Neural NetworkabstractWe present PyCARL, a PyNN-based common Python programming interface for hardware-software cosimulation of spiking neural network (SNN). Through PyCARL, we make the following two key contributions. First, we provide an interface of PyNN to CARLsim, a computationally- efficient, GPU-accelerated and biophysically-detailed SNN simulator. PyCARL facilitates joint development of machine learning models and code sharing between CARLsim and PyNN users, promoting an integrated and larger neuromorphic community. Second, we integrate cycle-accurate models of state-of-the-art neuromorphic hardware such as TrueNorth, Loihi, and DynapSE in PyCARL, to accurately model hardware latencies, which delay spikes between communicating neurons, degrading performance of machine learning models. PyCARL allows users to analyze and optimize the performance difference between software-based simulation and hardware-oriented simulation. We show that system designers can also use PyCARL to perform design-space exploration early in the product development stage, facilitating faster time-to-market of neuromorphic products. Adarsha Balaji, Prathyusha Adiraju, Hirak J. Kashyap, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Francky Catthoor |
IJCNN | 4 |
| 2020 | Exploiting inter- and intra-memory asymmetries for data mapping in hybrid tiered-memoriesabstractModern computing systems are embracing hybrid memory comprising of DRAM and non-volatile memory (NVM) to combine the best properties of both memory technologies, achieving low latency, high reliability, and high density. A prominent characteristic of DRAM-NVM hybrid memory is that it has NVM access latency much higher than DRAM access latency. We call this inter-memory asymmetry. We observe that parasitic components on a long bitline are a major source of high latency in both DRAM and NVM, and a significant factor contributing to high-voltage operations in NVM, which impact their reliability. We propose an architectural change, where each long bitline in DRAM and NVM is split into two segments by an isolation transistor. One segment can be accessed with lower latency and operating voltage than the other. By introducing tiers, we enable non-uniform accesses within each memory type (which we call intra-memory asymmetry), leading to performance and reliability trade-offs in DRAM-NVM hybrid memory. Shihao Song, Anup Das 0001, Nagarajan Kandasamy |
ISMM | 2 |
| 2020 | Improving phase change memory performance with data content aware accessabstractPhase change memory (PCM) is a scalable non-volatile memory technology that has low access latency (like DRAM) and high capacity (like Flash). Writing to PCM incurs significantly higher latency and energy penalties compared to reading its content. A prominent characteristic of PCM’s write operation is that its latency and energy are sensitive to the data to be written as well as the content that is overwritten. We observe that overwriting unknown memory content can incur significantly higher latency and energy compared to overwriting known all-zeros or all-ones content. This is because all-zeros or all-ones content is overwritten by programming the PCM cells only in one direction, i.e., using either SET or RESET operations, not both. Shihao Song, Anup Das 0001, Onur Mutlu, Nagarajan Kandasamy |
ISMM | 2 |
| 2020 | Compiling Spiking Neural Networks to Neuromorphic HardwareabstractMachine learning applications that are implemented with spike-based computation model, e.g., Spiking Neural Network (SNN), have a great potential to lower the energy consumption when executed on a neuromorphic hardware. How- ever, compiling and mapping an SNN to the hardware is challenging, especially when compute and storage resources of the hardware (viz. crossbars) need to be shared among the neurons and synapses of the SNN. We propose an approach to analyze and compile SNNs on resource-constrained neuromorphic hardware, providing guarantees on key performance metrics such as execution time and throughput. Our approach makes the following three key contributions. First, we propose a greedy technique to partition an SNN into clusters of neurons and synapses such that each cluster can fit on to the resources of a crossbar. Second, we exploit the rich semantics and expressiveness of Synchronous Dataflow Graphs (SDFGs) to represent a clustered SNN and analyze its performance using Max-Plus Algebra, considering the available compute and storage capacities, buffer sizes, and communication bandwidth. Third, we propose a self-timed execution-based fast technique to compile and admit SNN-based applications to a neuromorphic hardware at run-time, adapting dynamically to the available resources on the hard- ware. We evaluate our approach with standard SNN-based applications and demonstrate a significant performance improvement compared to current practices. Shihao Song, Adarsha Balaji, Anup Das 0001, Nagarajan Kandasamy, James A. Shackleford |
LCTES | 3 |
| 2020 | Mapping Spiking Neural Networks to Neuromorphic HardwareabstractNeuromorphic hardware implements biological neurons and synapses to execute a spiking neural network (SNN)-based machine learning. We present SpiNeMap, a design methodology to map SNNs to crossbar-based neuromorphic hardware, minimizing spike latency and energy consumption. SpiNeMap operates in two steps: SpiNeCluster and SpiNePlacer. SpiNeCluster is a heuristic-based clustering technique to partition an SNN into clusters of synapses, where intracluster local synapses are mapped within crossbars of the hardware and intercluster global synapses are mapped to the shared interconnect. SpiNeCluster minimizes the number of spikes on global synapses, which reduces spike congestion and improves application performance. SpiNePlacer then finds the best placement of local and global synapses on the hardware using a metaheuristic-based approach to minimize energy consumption and spike latency. We evaluate SpiNeMap using synthetic and realistic SNNs on a state-of-the-art neuromorphic hardware. We show that SpiNeMap reduces average energy consumption by 45% and spike latency by 21%, compared to the best-performing SNN mapping technique. Adarsha Balaji, Francky Catthoor, Anup Das 0001, Yuefeng Wu, Khanh Huynh, Francesco Dell'Anna, Giacomo Indiveri, Jeffrey L. Krichmar, Nikil Dutt, Siebren Schaafsma |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Design Methodology for Embedded Approximate Artificial Neural NetworksabstractArtificial neural networks (ANNs) have demonstrated significant promise while implementing recognition and classification applications. The implementation of pre-trained ANNs on embedded systems requires representation of data and design parameters in low-precision fixed-point formats; which often requires retraining of the network. For such implementations, the multiply-accumulate operation is the main reason for resultant high resource and energy requirements. To address these challenges, we present Rox-ANN, a design methodology for implementing ANNs using processing elements (PEs) designed with low-precision fixed-point numbers and high performance and reduced-area approximate multipliers on FPGAs. The trained design parameters of the ANN are analyzed and clustered to optimize the total number of approximate multipliers required in the design. With our methodology, we achieve insignificant loss in application accuracy. We evaluated the design using a LeNet based implementation of the MNIST digit recognition application. The results show a 65.6%, 55.1% and 18.9% reduction in area, energy consumption and latency for a PE using 8-bit precision weights and activations and approximate arithmetic units, when compared to 16-bit full precision, accurate arithmetic PEs. Adarsha Balaji, Salim Ullah, Anup Das 0001, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Exploration of Segmented Bus As Scalable Global Interconnect for Neuromorphic ComputingabstractSpiking Neural Networks (SNNs) are efficient computation models for spatio-temporal pattern recognition on resource and power constrained platforms. Dedicated SNN hardware, also called neuromorphic hardware, can further reduce the energy consumption of these platforms. A neuromorphic hardware consists of crossbars, which are arrangements of input and output neurons with fully-connected synapses. Time-multiplexed interconnects are used to communicate spikes between crossbars. When a SNN model is mapped on multiple crossbars, the time-multiplexed interconnect increases spike latency and energy consumption, and disorders spike arrivals at output neurons, which reduces application accuracy. In this paper, we propose segmented bus interconnect for global synapses in a neuromorphic architecture. The objective is to reduce power consumption and enable parallel processing compared to traditional time-multiplexed interconnects. The fundamental idea for the segmented bus is to partition a single bus into several segments, with the segmentation switches controlled by software. We evaluate the scalability of segmented bus using synthetic applications. Our results show that segmented bus reduces the latency and energy consumption of the global synapse network significantly with respect to state-of-the-art techniques. Adarsha Balaji, Yuefeng Wu, Anup Das 0001, Francky Catthoor, Siebren Schaafsma |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Enabling and Exploiting Partition-Level Parallelism (PALP) in Phase Change MemoriesabstractPhase-change memory (PCM) devices have multiple banks to serve memory requests in parallel . Unfortunately, if two requests go to the same bank , they have to be served one after another , leading to lower system performance . We observe that a modern PCM bank is implemented as a collection of partitions that operate mostly independently while sharing a few global peripheral structures, which include the sense amplifiers (to read) and the write drivers (to write). Based on this observation, we propose PALP , a new mechanism that enables partition-level parallelism within each PCM bank, and exploits such parallelism by using the memory controller’s access scheduling decisions. PALP consists of three new contributions. First , we introduce new PCM commands to enable parallelism in a bank’s partitions in order to resolve the read-write bank conflicts, with no changes needed to PCM logic or its interface. Second , we propose simple circuit modifications that introduce a new operating mode for the write drivers, in addition to their default mode of serving write requests. When configured in this new mode, the write drivers can resolve the read-read bank conflicts, working jointly with the sense amplifiers. Finally , we propose a new access scheduling mechanism in PCM that improves performance by prioritizing those requests that exploit partition-level parallelism over other requests, including the long outstanding ones. While doing so, the memory controller also guarantees starvation-freedom and the PCM’s running-average-power-limit (RAPL). We evaluate PALP with workloads from the MiBench and SPEC CPU2017 Benchmark suites. Our results show that PALP reduces average PCM access latency by 23%, and improves average system performance by 28% compared to the state-of-the-art approaches. Shihao Song, Anup Das 0001, Onur Mutlu, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | VRL-DRAM: improving DRAM performance via variable refresh latencyabstractA DRAM chip requires periodic refresh operations to prevent data loss due to charge leakage in DRAM cells. Refresh operations incur significant performance overhead as a DRAM bank/rank becomes unavailable to service access requests while being refreshed. In this work, our goal is to reduce the performance overhead of DRAM refresh by reducing the latency of a refresh operation. We observe that a significant number of DRAM cells can retain their data for longer than the worst-case refresh period of 64ms. Such cells do not always need to be fully refreshed; a low-latency partial refresh is sufficient for them. Anup Das 0001, Hasan Hassan, Onur Mutlu |
DAC | 1 |
| 2018 | Mapping of local and global synapses on spiking neuromorphic hardwareabstractSpiking Neural Networks (SNNs) are widely deployed to solve complex pattern recognition, function approximation and image classification tasks. With the growing size and complexity of these networks, hardware implementation becomes challenging because scaling up the size of a single array (crossbar) of fully connected neurons is no longer feasible due to strict energy budget. Modern neromorphic hardware integrates small-sized crossbars with time-multiplexed interconnects. Partitioning SNNs becomes essential in order to map them on neuromorphic hardware with the major aim to reduce the global communication latency and energy overhead. To achieve this goal, we propose our instantiation of particle swarm optimization, which partitions SNNs into local synapses (mapped on crossbars) and global synapses (mapped on time-multiplexed interconnects), with the objective of reducing spike communication on the interconnect. This improves latency, power consumption as well as application performance by reducing inter-spike interval distortion and spike disorders. Our framework is implemented in Python, interfacing CARLsim, a GPU-accelerated application-level spiking neural network simulator with an extended version of Noxim, for simulating time-multiplexed interconnects. Experiments are conducted with realistic and synthetic SNN-based applications with different computation models, topologies and spike coding schemes. Using power numbers from in-house neuromorphic chips, we demonstrate significant reductions in energy consumption and spike latency over PACMAN, the widely-used partitioning technique for SNNs on SpiNNaker. Anup Das 0001, Yuefeng Wu, Khanh Huynh, Francesco Dell'Anna, Francky Catthoor, Siebren Schaafsma |
DATE | 1 |
| 2018 | Dataflow-Based Mapping of Spiking Neural Networks on Neuromorphic HardwareabstractSpiking Neural Networks (SNNs) are powerful computation engines for pattern recognition and image classification applications. Apart from application performance such as recognition and classification accuracy, system performance such as throughput becomes important when executing these applications on a hardware. We propose a systematic design-flow to map SNN-based applications on a crossbar-based neuromorphic hardware, guaranteeing application as well as system performance. Synchronous Dataflow Graphs (SDFGs) are used to model these applications with extended semantics to represent neural network topologies. Self-timed scheduling is then used to analyze throughput, incorporating hardware constraints such as synaptic memory, communication and I/O bandwidth of crossbars. Our design-flow integrates CARLsim, a GPU-accelerated application-level SNN simulator with SDF3, a tool for mapping SDFG on hardware. We conducted experiments with realistic and synthetic SNNs on representative neuromorphic hardware, demonstrating throughput-resource trade-offs for a given application performance. For throughput-constrained applications, we show average 20% reduction of hardware usage with 19% reduction in energy consumption. For throughput-scalable applications, we show an average 53% higher throughput compared to a state-of-the-art approach. Anup Das 0001, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Unsupervised heart-rate estimation in wearables with Liquid states and a probabilistic readout
Anup Das 0001, Paruthi Pradhapan, Willemijn Groenendaal, Prathyusha Adiraju, Raj Thilak Rajan, Francky Catthoor, Siebren Schaafsma, Jeffrey L. Krichmar, Nikil Dutt, Chris Van Hoof |
Neural Networks | 1 |
| 2017 | Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUsabstractModern mobile and embedded devices are required to be increasingly energy-efficient while running more sophisticated tasks, causing the CPU design to become more complex and employ more energy-saving techniques. This has created a greater need for fast and accurate power estimation frameworks for both run-time CPU energy management and design-space exploration. We present a statistically rigorous and novel methodology for building accurate run-time power models using performance monitoring counters (PMCs) for mobile and embedded devices, and demonstrate how our models make more efficient use of limited training data and better adapt to unseen scenarios by uniquely considering stability. Our robust model formulation reduces multicollinearity, allows separation of static and dynamic power, and allows a 100× reduction in experiment time while sacrificing only 0.6% accuracy. We present a statistically detailed evaluation of our model, highlighting and addressing the problem of heteroscedasticity in power modeling. We present software implementing our methodology and build power models for ARM Cortex-A7 and Cortex-A15 CPUs, with 3.8% and 2.8% average error, respectively. We model the behavior of the nonideal CPU voltage regulator under dynamic CPU activity to improve modeling accuracy by up to 5.5% in situations where the voltage cannot be measured. To address the lack of research utilizing PMC data from real mobile devices, we also present our data acquisition method and experimental platform software. We support this paper with online resources including software tools, documentation, raw data and further results. Matthew J. Walker, Stephan Diestelhorst, Andreas Hansson 0001, Anup Das 0001, Sheng Yang 0003, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | The slowdown or race-to-idle question: Workload-aware energy optimization of SMT multicore platforms under process variation
Anup Das 0001, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 1 |
| 2016 | Workload Change Point Detection for Runtime Thermal Management of Embedded SystemsabstractApplications executed on multicore embedded systems interact with system software [such as the operating system (OS)] and hardware, leading to widely varying thermal profiles which accelerate some aging mechanisms, reducing the lifetime reliability. Effectively managing the temperature therefore requires: 1) autonomous detection of changes in application workload and 2) appropriate selection of control levers to manage thermal profiles of these workloads. In this paper, we propose a technique for workload change detection using density ratio-based statistical divergence between overlapping sliding windows of CPU performance statistics. This is integrated in a runtime approach for thermal management, which uses reinforcement learning to select workload-specific thermal control levers by sampling on-board thermal sensors. Identified control levers override the OSs native thread allocation decision and scale hardware voltage-frequency to improve average temperature, peak temperature, and thermal cycling. The proposed approach is validated through its implementation as a hierarchical runtime manager for Linux, with heuristic-based thread affinity selected from the upper hierarchy to reduce thermal cycling and learningbased voltage-frequency selected from the lower hierarchy to reduce average and peak temperatures. Experiments conducted with mobile, embedded, and high performance applications on ARM-based embedded systems demonstrate that the proposed approach increases workload change detection accuracy by an average 3.4×, reducing the average temperature by 4 °C-25 °C, peak temperature by 6 °C-24 °C, and thermal cycling by 7%-35% over state-of-the-art approaches. Anup Das 0001, Geoff V. Merrett, Mirco Tribastone, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Graceful Performance Modulation for Power-Neutral Transient Computing SystemsabstractTransient computing systems do not have energy storage, and operate directly from energy harvesting. These systems are often faced with the inherent challenge of low-current or transient power supply. In this paper, we propose “power-neutral” operation, a new paradigm for such systems, whereby the instantaneous power consumption of the system must match the instantaneous harvested power. Power neutrality is achieved using a control algorithm for dynamic frequency scaling, modulating system performance gracefully in response to the incoming power. Detailed system model is used to determine design parameters for selecting the system voltage thresholds where the operating frequency will be raised or lowered, or the system will be hibernated. The proposed control algorithm for power-neutral operation is experimentally validated using a microcontroller incorporating voltage threshold-based interrupts for frequency scaling. The microcontroller is powered directly from real energy harvesters; results demonstrate that a power-neutral system sustains operation for 4%-88% longer with up to 21% speedup in application execution. Domenico Balsamo, Anup Das 0001, Alex S. Weddell, Davide Brunelli, Bashir M. Al-Hashimi, Geoff V. Merrett, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Hibernus++: A Self-Calibrating and Adaptive System for Transiently-Powered Embedded DevicesabstractEnergy harvesters are being used to power autonomous systems, but their output power is variable and intermittent. To sustain computation, these systems integrate batteries or supercapacitors to smooth out rapid changes in harvester output. Energy storage devices require time for charging and increase the size, mass, and cost of systems. The field of transient computing moves away from this approach, by powering the system directly from the harvester output. To prevent an application from having to restart computation after a power outage, approaches such as Hibernus allow these systems to hibernate when supply failure is imminent. When the supply reaches the operating threshold, the last saved state is restored and the operation is continued from the point it was interrupted. This paper proposes Hibernus++ to intelligently adapt the hibernate and restore thresholds in response to source dynamics and system load properties. Specifically, capabilities are built into the system to autonomously characterize the hardware platform and its performance during hibernation in order to set the hibernation threshold at a point which minimizes wasted energy and maximizes computation time. Similarly, the system auto-calibrates the restore threshold depending on the balance of energy supply and consumption in order to maximize computation time. Hibernus++ is validated both theoretically and experimentally on microcontroller hardware using both synthesized and real energy harvesters. Results show that Hibernus++ provides an average 16% reduction in energy consumption and an improvement of 17% in application execution time over state-of-the-art approaches. Domenico Balsamo, Alex S. Weddell, Anup Das 0001, Alberto Rodriguez Arreola, Davide Brunelli, Bashir M. Al-Hashimi, Geoff V. Merrett, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Learning Transfer-Based Adaptive Energy Minimization in Embedded SystemsabstractEmbedded systems execute applications with varying performance requirements. These applications exercise the hardware differently depending on the computation task, generating varying workloads with time. Energy minimization with such workload and performance variations within (intra) and across (inter) applications is particularly challenging. To address this challenge, we propose an online approach, capable of minimizing energy through adaptation to these variations. At the core of this approach is a reinforcement learning algorithm that suitably selects the appropriate voltage/frequency scaling (VFS) based on workload predictions to meet the applications' performance requirements. The adaptation is then facilitated and expedited through learning transfer, which uses the interaction between the application, runtime, and hardware layers to adjust the VFS. The proposed approach is implemented as a power governor in Linux and extensively validated on an ARM Cortex-A8 running different benchmark applications. We show that with intra- and inter-application variations, our proposed approach can effectively minimize energy consumption by up to 33% compared to the existing approaches. Scaling the approach to multicore systems, we also demonstrate that it can minimize energy by up to 18% with 2× reduction in the learning time when compared with an existing approach. Rishad A. Shafik, Sheng Yang 0003, Anup Das 0001, Luis Alfonso Maeda-Nunez, Geoff V. Merrett, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Adaptive and Hierarchical Runtime Manager for Energy-Aware Thermal Management of Embedded SystemsabstractModern embedded systems execute applications, which interact with the operating system and hardware differently depending on the type of workload. These cross-layer interactions result in wide variations of the chip-wide thermal profile. In this article, a reinforcement learning-based runtime manager is proposed that guarantees application-specific performance requirements and controls the POSIX thread allocation and voltage/frequency scaling for energy-efficient thermal management. This controls three thermal aspects: peak temperature, average temperature, and thermal cycling. Contrary to existing learning-based runtime approaches that optimize energy and temperature individually, the proposed runtime manager is the first approach to combine the two objectives, simultaneously addressing all three thermal aspects. However, determining thread allocation and core frequencies to optimize energy and temperature is an NP-hard problem. This leads to exponential growth in the learning table (significant memory overhead) and a corresponding increase in the exploration time to learn the most appropriate thread allocation and core frequency for a particular application workload. To confine the learning space and to minimize the learning cost, the proposed runtime manager is implemented in a two-stage hierarchy: a heuristic-based thread allocation at a longer time interval to improve thermal cycling, followed by a learning-based hardware frequency selection at a much finer interval to improve average temperature, peak temperature, and energy consumption. This enables finer control on temperature in an energy-efficient manner while simultaneously addressing scalability, which is a crucial aspect for multi-/many-core embedded systems. The proposed hierarchical runtime manager is implemented for Linux running on nVidia’s Tegra SoC, featuring four ARM Cortex-A15 cores. Experiments conducted with a range of embedded and cpu-intensive applications demonstrate that the proposed runtime manager not only reduces energy consumption by an average 15% with respect to Linux but also improves all the thermal aspects—average temperature by 14°C, peak temperature by 16°C, and thermal cycling by 54%. Anup Das 0001, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Reliability and Energy-Aware Mapping and Scheduling of Multimedia Applications on Multiprocessor SystemsabstractLifetime reliability is an emerging concern in multiprocessor systems as escalating power density and hence temperature variation continues to accelerate wear-out leading to a growing prominence of device defects. In this paper, we propose a system-level approach that involves performance-aware mapping of multimedia applications on a multiprocessor system to jointly minimize energy consumption and temperature related wear-out. Fundamental to this approach is a simplified temperature model that incorporates not only the transient and the steady-state behavior (temporal effect), but also the temperature dependency on the surrounding cores (spatial effect). This model is validated against the temperature obtained using theHotSpottool with transient and steady-state simulations, and is shown to be accurate within 5.5°C, leading to an MTTF estimation accuracy of an average 21 percent with respect to the state-of-the-art approaches. The proposed temperature model is integrated in a gradient-based fast heuristic that controls the voltage and frequency of the cores to limit the average and peak temperature leading to a longer lifetime, simultaneously minimizing the energy consumption. Lifetime computation considers task remapping, which is a common feature available in modern multiprocessor systems. A linear programming approach is then proposed to distribute the cores of a multiprocessor system among concurrent applications to maximize the lifetime. Experiments conducted with a set of synthetic and real-life applications represented as synchronous data flow graphs demonstrate that the proposed approach minimizes energy consumption by an average 24 percent with 47 percent increase in lifetime. For concurrent applications, the proposed lifetime-aware core distribution results in an average 10 percent improvement in lifetime as compared to performance-based core distribution. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Workload uncertainty characterization and adaptive frequency scaling for energy minimization of embedded systems
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 1 |
| 2015 | Hardware-software interaction for run-time power optimization: A case study of embedded Linux on multicore smartphonesabstractApplications running on smartphones interact with the hardware and the system software differently, resulting in widely varying power consumption and hence thermal profiles. Typically, these smartphone platforms expose some hardware power control features to users, controlled through software governors such as cpufreq for dynamic voltage-frequency scaling (DVFS) and cpuquiet for dynamic core selection (DCS). Operating systems on these platforms manage these governors conservatively, independent of application's performance requirement. To address this, we propose an alternative approach, which uses reinforcement learning to explore the trade-off between power saving opportunities using DVFS and DCS and application's performance at run-time. The objective is to reduce power consumption, taking into consideration dynamic power, leakage power, and the inter-dependency between temperature and power. The reinforcement learning-based control is validated as a case-study on ARM A15-based nvidia's tegra smartphone through its implementation as a run-time manager (RTM). This RTM interfaces with different hardware performance counters and the embedded Linux Operating System through (1) the cpuquiet API to select cores at run-time; and (2) the cpufreq API to scale the frequency of active cores. Experiments with mobile and high performance applications demonstrate that the proposed approach achieves an average 22% (7-40%) power reduction compared to existing techniques. Anup Das 0001, Matthew J. Walker, Andreas Hansson 0001, Bashir M. Al-Hashimi, Geoff V. Merrett |
ISLPED | 1 |
| 2015 | Execution Trace-Driven Energy-Reliability Optimization for Multimedia MPSoCsabstractMultiprocessor systems-on-chip (MPSoCs) are becoming a popular design choice in current and future technology nodes to accommodate the heterogeneous computing demand of a multitude of applications enabled on these platform. Streaming multimedia and other communication-centric applications constitute a significant fraction of the application space of these devices. The mapping of an application on an MPSoC is an NP-hard problem. This has attracted researchers to solve this problem both as stand-alone (best-effort) and in conjunction with other optimization objectives, such as energy and reliability. Most existing studies on energy-reliability joint optimization are static—that is, design time based. These techniques fail to capture runtime variability such as resource unavailability and dynamism associated with application behaviors, which are typical of multimedia applications. The few studies that consider dynamic mapping of applications do not consider throughput degradation, which directly impacts user satisfaction. This article proposes a runtime technique to analyze the execution trace of an application modeled as Synchronous Data Flow Graphs (SDFGs) to determine its mapping on a multiprocessor system with heterogeneous processing units for different fault scenarios. Further, communication energy is minimized for each of these mappings while satisfying the throughput constraint. Experiments conducted with synthetic and real SDFGs demonstrate that the proposed technique achieves significant improvement with respect to the state-of-the-art approaches in terms of throughput and storage overhead with less than 20% energy overhead. Anup Das 0001, Amit Kumar Singh 0002, Akash Kumar 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2015 | Autonomous Soft-Error Tolerance of FPGA Configuration BitsabstractField-programmable gate arrays (FPGAs) are increasingly susceptible to radiation-induced single event upsets (SEUs). These upsets are predominant in a space environment; however, with increasing use of static RAM (SRAM) in modern FPGAs, these SEUs are gaining prominence even in a terrestrial environment. SEUs can flip SRAM bits of FPGA, potentially altering the functionality of the implemented design. This has motivated FPGA designers to investigate techniques to protect the FPGA configuration bits against such inadvertent bit flips (soft error). Traditionally, triple modular redundancy (TMR) is used to protect the FPGA bit flips. Increasing design complexity and limited battery life motivate for alternative approaches for soft-error tolerance. In this article, we propose a technique to improve autonomous fault-masking capabilities of a design by maximizing the number of zeros or ones in lookup tables (LUTs). The technique analyzes critical configuration bits and utilizes spare resources (XOR gates and carry chains) of FPGAs to selectively manipulate the logic implemented in LUTs using two operations: LUT restructuring and LUT decomposition. We implemented the proposed approach for Xilinx Virtex-6 FPGAs and validated the same with a wide set of designs from the MCNC, IWLS 2005, and ITC99 benchmark suites. Results demonstrate that the proposed logic restructuring maximizes logic 0 (or 1) of LUTs by an average of 20%, achieving 80% fault masking with no area overhead. The fault rate of the entire design is reduced by 60% on average as compared to the existing techniques. Furthermore, the logic decomposition algorithm provides incremental fault-tolerance capabilities and achieves an additional 5% fault masking with an average 7% increase in slice usage. The complete methodology is implemented into a tool for Xilinx FPGA and is made available online for the benefit of the research community. The algorithms are lightweight, and the whole design flow (including Xilinx Place and Route) was completed in 75 minutes for the largest benchmark in the set. Anup Das 0001, Shyamsundar Venkataraman, Akash Kumar 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2014 | Reinforcement Learning-Based Inter- and Intra-Application Thermal Optimization for Lifetime Improvement of Multicore SystemsabstractThe thermal profile of multicore systems vary both within an application's execution (intra) and also when the system switches from one application to another (inter). In this paper, we propose an adaptive thermal management approach to improve the lifetime reliability of multicore systems by considering both inter- and intra-application thermal variations. Fundamental to this approach is a reinforcement learning algorithm, which learns the relationship between the mapping of threads to cores, the frequency of a core and its temperature (sampled from on-board thermal sensors). Action is provided by overriding the operating system's mapping decisions using affinity masks and dynamically changing CPU frequency using in-kernel governors. Lifetime improvement is achieved by controlling not only the peak and average temperatures but also thermal cycling, which is an emerging wear-out concern in modern systems. The proposed approach is validated experimentally using an Intel quad-core platform executing a diverse set of multimedia benchmarks. Results demonstrate that the proposed approach minimizes average temperature, peak temperature and thermal cycling, improving the mean-time-to-failure (MTTF) by an average of 2x for intra-application and 3x for inter-application scenarios when compared to existing thermal management techniques. Furthermore, the dynamic and static energy consumption are also reduced by an average 10% and 11% respectively. Anup Das 0001, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi, Akash Kumar 0001, Bharadwaj Veeravalli |
DAC | 1 |
| 2014 | Temperature aware energy-reliability trade-offs for mapping of throughput-constrained applications on multimedia MPSoCsabstractThis paper proposes a design-time (offline) analysis technique to determine application task mapping and scheduling on a multiprocessor system and the voltage and frequency levels of all cores (offline DVFS) that minimize application computation and communication energy, simultaneously minimizing processor aging. The proposed technique incorporates (1) the effect of the voltage and frequency on the temperature of a core; (2) the effect of neighboring cores' voltage and frequency on the temperature (spatial effect); (3) pipelined execution and cyclic dependencies among tasks; and (4) the communication energy component which often constitutes a significant fraction of the total energy for multimedia applications. The temperature model proposed here can be easily integrated in the design space exploration for multiprocessor systems. Experiments conducted with MPEG-4 decoder on a real system demonstrate that the temperature using the proposed model is within 5% of the actual temperature clearly demonstrating its accuracy. Further, the overall optimization technique achieves 40% savings in energy consumption with 6% increase in system lifetime. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
DATE | 1 |
| 2014 | Combined DVFS and mapping exploration for lifetime and soft-error susceptibility improvement in MPSoCsabstractEnergy and reliability optimization are two of the most critical objectives for the synthesis of multiprocessor systems-on-chip (MPSoCs). Task mapping has shown significant promise as a low cost solution in achieving these objectives as standalone or in tandem as well. This paper proposes a multi-objective design space exploration to determine the mapping of tasks of an application on a multiprocessor system and voltage/frequency level of each tasks (exploiting the DVFS capabilities of modern processors) such that the reliability of the platform is improved while fulfilling the energy budget and the performance constraint set by system designers. In this respect, the reliability of a given MPSoC platform incorporates not only the impact of voltage and frequency on the aging of the processors (wear-out effect) but also on the susceptibility to soft-errors - a joint consideration missing in all existing works in this domain. Further, the proposed exploration also incorporates soft-error tolerance by selective replication of tasks, making the proposed approach an interesting blend of reactive and proactive fault-tolerance. The combined objective of minimizing core aging together with the susceptibility to transient faults under a given performance/energy budget is solved by using a multi-objective genetic algorithm exploiting tasks' mapping, DVFS and selective replication as tuning knobs. Experiments conducted with reallife and synthetic application graphs clearly demonstrate the advantage of the proposed approach. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Cristiana Bolchini, Antonio Miele |
DATE | 1 |
| 2014 | Criticality-aware scrubbing mechanism for SRAM-based FPGAsabstractScrubbing has been considered as an effective mechanism to provide fault-tolerance in Static-RAM (SRAM)-based Field Programmable Gate Arrays (FPGAs). However, the current scrubbing techniques execute without considering the criticality and timing of the user tasks implemented in the FPGA. They often do not execute the scrubbing process in the right instant, which minimizes the probability of each task being executed without transient faults. Moreover, these current solutions are not adapted to the tasks' fault-tolerance requirements, since they may not properly protect the most critical tasks in the system. However, if they do it, they waste resources with the less critical tasks. In this paper, a new scrubbing mechanism is proposed. This new approach adapts the scrubbing mechanism to the tasks' execution, by a proper scheduling and according to their criticality. A proposed heuristic finds a feasible scrubbing schedule for each hardware task. Firstly, the minimum scrubbing periods are computed according to the criticality of each implemented hardware task. Secondly, a proper scrubbing schedule following the EDL (Earliest Deadline as Late as possible) algorithm is found, maximizing the reliability of the system. The experimental results show up to 79% improvements on the system reliability, achieved without wasting scrubbing resources. Shyamsundar Venkataraman, Anup Das 0001, Akash Kumar 0001 |
FPL | 3 |
| 2014 | A bit-interleaved embedded hamming scheme to correct single-bit and multi-bit upsets for SRAM-based FPGAsabstractSingle Event Upsets (SEUs) inadvertently change the configuration bits of Static-RAM (SRAM)-based Field Programmable Gate Arrays (FPGAs), leading to erroneous output until the error has been corrected. Scrubbing using an Error Correction Code (ECC) such as hamming is a popular method to correct such faults. However, current works either require a large external memory to store the ECCs or can at most correct only one error in a frame. This paper proposes a novel bit-interleaved embedded hamming scheme along with scrubbing, to correct single (SBUs) and multi-bit upsets (MBUs) in SRAM-based FPGAs. This scheme does not require an external memory to store the ECCs, as they are embedded within the configuration memory itself. Experiments conducted on various benchmarks show that the proposed scheme can handle multiple errors per frame very well, with an embedding efficiency of over 99.3%. Shyamsundar Venkataraman, Anup Das 0001, Akash Kumar 0001 |
FPL | 3 |
| 2014 | Communication and migration energy aware task mapping for reliable multiprocessor systems
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
Future Gener. Comput. Syst. | 1 |
| 2014 | Energy-aware task mapping and scheduling for reliable embedded computing systemsabstractTask mapping and scheduling are critical in minimizing energy consumption while satisfying the performance requirement of applications enabled on heterogeneous multiprocessor systems. An area of growing concern for modern multiprocessor systems is the increase in the failure probability of one or more component processors. This is especially critical for applications where performance degradation (e.g., throughput) directly impacts the quality of service requirement. This article proposes a design-time (offline) multi-criterion optimization technique for application mapping on embedded multiprocessor systems to minimize energy consumption for all processor fault-scenarios. A scheduling technique is then proposed based on self-timed execution to minimize the schedule storage and construction overhead at runtime. Experiments conducted with synthetic and real applications from streaming and nonstreaming domains on heterogeneous MPSoCs demonstrate that the proposed technique minimizes energy consumption by 22% and design space exploration time by 100x, while satisfying the throughput requirement for all processor fault-scenarios. For scalable throughput applications, the proposed technique achieves 30% better throughput per unit energy, compared to the existing techniques. Additionally, the self-timed execution-based scheduling technique minimizes schedule construction time by 95% and storage overhead by 92%. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Aging-aware hardware-software task partitioning for reliable reconfigurable multiprocessor systemsabstractHomogeneous multiprocessor systems with reconfigurable area (also known as Reconfigurable Multiprocessor Systems) are emerging as a popular design choice in current and future technology nodes to meet the heterogeneous computing demand of a multitude of applications enabled on these platforms. Application specific mapping decisions on such a platform involve partitioning a given application into software tasks (executed on one or more of the general purpose processors, GPPs) and the hardware tasks (realized as dedicated hardware on the reconfigurable area) to optimize and/or satisfy design constraints such as reliability, performance and design cost. Improving the reliability considering transient faults by increasing the number of checkpoints negatively impacts the reliability considering permanent faults. This trade-off is ignored in all prior studies on task mapping and scheduling. This paper proposes an optimization technique to decide the optimal number of checkpoints for the software tasks which minimizes aging of the GPPs while maximizing the transient fault-tolerance of the overall platform (GPPs and the reconfigurable area) and satisfying design cost and performance. Experiments conducted with synthetic and real-life application task graphs (cyclic and acyclic) demonstrate that the proposed technique minimizes aging and improves the platform lifetime by an average 60% as compared to the existing transient fault-aware techniques. Further, a gradient-based heuristic is proposed to minimize the design space exploration time by upto 500× with less than 5% deviation from optimal solution. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
CASES | 1 |
| 2013 | Energy optimization by exploiting execution slacks in streaming applications on multiprocessor systemsabstractDynamic voltage and frequency scaling (DVFS) offers great potential for optimizing the energy efficiency of Multiprocessor Systems-on-Chip (MPSoCs). The conventional approaches for processor voltage and frequency adjustment are not suitable for streaming multimedia applications due to the cyclic nature of dependencies in the executing tasks which can potentially violate the throughput constraints. In this paper, we propose a methodology that applies DVFS for such cyclic dependent tasks. The methodology involves an off-line analysis that assumes worst-case execution times of tasks to identify the executions that can be slowed down and an on-line analysis to utilize the slacks arising from tasks that finish their execution before the worst-case execution times. Thus, the methodology minimizes energy consumption during both off-line and on-line analysis while satisfying the throughput constraints. Experiments based on models of real-life streaming multimedia applications show that the proposed methodology reduces the overall energy consumption by 43% when compared to existing approaches. Amit Kumar Singh 0002, Anup Das 0001, Akash Kumar 0001 |
DAC | 2 |
| 2013 | Reliability-driven task mapping for lifetime extension of networks-on-chip based multiprocessor systemsabstractShrinking transistor geometries, aggressive voltage scaling and higher operating frequencies have negatively impacted the lifetime reliability of embedded multi-core systems. In this paper, a convex optimization-based task-mapping technique is proposed to extend the lifetime of a multiprocessor systems-on-chip (MPSoCs). The proposed technique generates mappings for every application enabled on the platform with variable number of cores. Based on these results, a novel 3D-optimization technique is developed to distribute the cores of an MPSoC among multiple applications enabled simultaneously. Additionally, reliability of the underlying network-on-chip links is also addressed by incorporating aging of links in the objective function. Our formulations are developed for directed acyclic graphs (DAGs) and synchronous dataflow graphs (SDFGs), making our approach applicable for streaming as well as non-streaming applications. Experiments conducted with synthetic and real-life application graphs demonstrate that the proposed approach extends the lifetime of an MPSoC by more than 30% when applications are enabled individually as well as in tandem. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
DATE | 1 |
| 2013 | Communication and migration energy aware design space exploration for multicore systems with intermittent faultsabstractShrinking transistor geometries, aggressive voltage scaling and higher operating frequencies have negatively impacted the dependability of embedded multicore systems. Most existing research works on fault-tolerance have focused on transient and permanent faults of cores. Intermittent faults are a separate class of defects resulting from on-chip temperature, pressure and voltage variations and lasting for a few cycles to several seconds or more. Operations of cores impacted by intermittent faults are suspended during these cycles but come back alive when conditions become favorable. This paper proposes a technique to model the availability of multiprocessor systems-on-chip (MPSoCs) with intermittent and reparable device defects. This model is based on Markov chain with stochastic fault distribution and can be applied even for permanent faults. Based on this model, a design space pruning technique is proposed to select a set of task mappings (with variable resource usage), which minimizes the task communication energy while satisfying the MPSoC availability constraint. Moreover, task migration overhead is also minimized, which is an important consideration for frequently occurring intermittent and temperature related faults, where prolonged system downtime during task re-mapping is not desired. Experiments conducted with real-life and synthetic application task graphs demonstrate that the proposed technique minimizes communication energy by 30% and reduces migration overhead by 50% as compared to the existing approaches. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
DATE | 1 |
| 2013 | RAPIDITAS: RAPId Design-Space-Exploration Incorporating Trace-Based Analysis and SimulationabstractSimulation-based Design Space Exploration (DSE) to evaluate all possible mappings for a given application and Multiprocessor-System-on-Chip (MPSoC) platform is computationally costly for large problems. Even using efficient exploration methodologies to evaluate the mappings cannot overcome the evaluation time bottleneck. This paper presents a novel DSE methodology that analyzes the execution trace to prune the vast design space. Simulations are employed only on the pruned design points (mappings), hence reducing the number of simulations. The methodology performs iterative exploration and provides premier mappings requiring different number of processors, which can be used at run-time subject to desired performance and available platform processors. We evaluate our methodology by using models of real-life multimedia applications and demonstrate that the DSE time is reduced by 72% while generating high quality mappings. Amit Kumar Singh 0002, Anup Das 0001, Akash Kumar 0001 |
DSD | 2 |
| 2013 | Improving autonomous soft-error tolerance of FPGA through LUT configuration bit manipulationabstractSoft-errors in LUT configuration bits of FPGAs can alter the functionality of an implemented design, rendering it useless, unless re-programmed. This paper proposes a technique to improve autonomous fault-masking capabilities of a design by maximizing the number of zeros or ones in LUTs. The technique utilizes spare resources (XOR gates and carry chain) of FPGA devices to selectively manipulate LUT contents using two operations - LUT restructuring and LUT decomposition. Experiments conducted with a wide set of benchmarks from MCNC, IWLS 2005 and ITC99 benchmark suite on Xilinx Virtex 6 FPGA board demonstrate that the proposed methodology maximizes logic 0/1 of LUTs by an average 20% achieving 80% fault-masking with no area overhead. The fault-rate of the entire design is reduced by 60% on average as compared to the existing techniques. Further, an additional 5% fault-masking can be achieved with a 7% increase in slice usage. Anup Das 0001, Shyamsundar Venkataraman, Akash Kumar 0001 |
FPL | 1 |
| 2012 | Minimizing Power Consumption of Spatial Division Based Networks-on-Chip Using Multi-path and Frequency ReductionabstractWith an increasing number of processing elements being integrated on a single die, networks-on-chip (NoCs) are emerging as a significant contributor to overall chip power consumption. While some solutions have been proposed to reduce this power consumption, none of them can be applied to spatial division multiplexing (SDM)-based NoCs. In this paper, we introduce a method to minimize the power consumption of an SDM-based NoC by frequency minimization, while still satisfying the bandwidth requirements. The problem is integrated with the connection-routing problem which is modeled as a mixed-integer quadratic constrained problem (MIQCP). However, solving this MIQCP formulation directly using existing solvers is infeasible for large use-cases. We propose a two-step approach by first computing the minimum feasible frequency for the entire network taking bandwidth of all connections into consideration. This first step reduces the frequency-minimization-routing MIQCP problem into a routing-only mixed-integer linear programming (MILP) problem. In the second step, this MILP problem is solved using a standard ILP solver. Two other techniques are proposed to solve the routing and frequency minimization problem. Experiments are performed with synthetic examples and a case-study with JPEG decoder to evaluate the performance and results of the three methods. MILP-based approach achieves up to 55% power reduction as compared to the other methods albeit at the cost of higher execution time. Sheng Hao Wang, Anup Das 0001, Akash Kumar 0001, Henk Corporaal |
DSD | 2 |
| 2012 | Energy-Aware Communication and Remapping of Tasks for Reliable Multimedia Multiprocessor SystemsabstractShrinking transistor geometries, aggressive voltage scaling and higher operating frequencies have negatively impacted the dependability of embedded multiprocessor systems-on-chip (MPSoCs). Fault-tolerance and energy efficiency are the two most desired features of modern-day MPSoCs. For most of the multimedia applications, task communication energy constitutes more than 40% of the overall application energy. In this paper, an integer linear programming (ILP) based approach is proposed to reduce the communication energy and fault-tolerant migration overhead of throughput-constrained multimedia applications modeled using synchronous data flow graphs (SDFGs). The ILP is solved at compile-time for all fault-scenarios to generate task-core mappings satisfying an application throughput requirement. These mappings are stored in a table which is looked up at run-time as and when faults occur. Experiments conducted with real and synthetic applications demonstrate that the proposed technique reduces communication energy by an average 40% and migration overhead by 33% as compared to the existing fault-tolerant techniques. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli |
ICPADS | 1 |
| 2012 | Fault-aware task re-mapping for throughput constrained multimedia applications on NoC-based MPSoCsabstractShrinking transistor geometry and aggressive voltage scaling are leading to growing concerns on the reliability of multiprocessor systems. Majority of streaming multimedia applications are characterized by fixed throughput requirements; violation of which directly impacts user experience. None of the prior research considers joint treatment of throughput and task-migration overhead, both of which are essential for fault-tolerance of throughput-constrained multimedia multiprocessor systems. In this paper, we propose to remap tasks from faulty processors with the objective of minimizing the migration overhead while satisfying throughput constraints. The proposed technique is based on extensive design-time analysis of different fault scenarios to determine optimal mappings from the throughput-migration overhead Pareto space. These mappings are stored in a table and are looked-up at run-time to migrate tasks as and when faults occur. Applications are modeled using Synchronous Data Flow graphs (SDFG) to consider cyclic dependencies of tasks, typically found in multimedia systems. Experiments performed with synthetic and real application graphs demonstrate that the migration overhead can be reduced by 26% on average while still meeting throughput constraints. Moreover, by selecting an appropriate initial processor-task mapping, migration overhead can be further reduced by 15% on average. Anup Das 0001, Akash Kumar 0001 |
RSP | 1 |
| 2012 | A design flow for partially reconfigurable heterogeneous multi-processor platformsabstractModern multiprocessor systems-on-chip (MPSoCs) are expected to handle multi-application use cases. As the number and complexity of these applications scale, resource allocation to meet the application throughput requirement is becoming quite a challenge. In this paper, a complete design flow is proposed for partially reconfigurable heterogeneous MPSoC platforms. The proposed flow determines the minimum resources required to map and guarantee the throughput of applications in all use-cases. Further, a suitable mapping for each application is chosen so that energy consumption is minimized. Experiments conducted with a set of synthetic benchmarks and real-life applications clearly demonstrate the advantage of our approach over homogeneous or fully reconfigurable designs. The proposed design flow achieves more than 50% energy savings when the number of configurations is not optimized. With configuration-optimization, our flow results in 75% reduction in the number of configurations with 5% reduction in energy. Jiashu Li, Anup Das 0001, Akash Kumar 0001 |
RSP | 2 |
| 2010 | A study on performance benefits of core morphing in an asymmetric multicore processorabstractMulticore architectures are designed so as to provide an acceptable level of performance per unit power for the majority of applications. Consequently, we must occasionally expect applications that could have benefited from a more powerful core in terms of either lower execution time and/or lower energy consumed. Fusing some of the resources of two (or more) cores to configure a more powerful core for such instances is a natural approach to deal with those few applications that have very high performance demands. However, a recent study has shown that fusing homogeneous cores is unlikely to benefit applications. In this paper we study the potential performance benefits of core morphing in a heterogeneous multicore processor that can be reconfigured at runtime. We consider as an example a dual core processor with one of the two cores being designed to target integer intensive applications while the other is better suited to floating-point intensive applications. These two cores can be fused into a single powerful core when an application that can benefit from such fusion is executing. We first discuss the design principles of the two individual cores so that the majority of the benchmarks that we consider execute in a satisfactory way. We then show that a small subset of the considered applications can greatly benefit from core morphing even in the case where two applications that could have been executed in parallel on the two cores are run, for some percentage of time, on the single morphed core. Our results indicate that a performance gain of up to 100% is achievable at a small hardware overhead of less than 1%. Anup Das 0001, Rance Rodrigues, Israel Koren, Sandip Kundu |
ICCD | 1 |