EDBT 2026 Demo / reviewers in the wild / expert
Adarsha Balaji
dblp:235/8858
· DBLP profile ↗
11ranked-venue papers
5as first author
6since 2021 · last 2023
0000-0002-3535-8788ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Application - Hardware Co-Optimization of Crossbar-Based Neuromorphic SystemsabstractSpiking Neural Networks (SNNs) executed on neuro-morphic hardware (NmC) have shown great potential to perform a class of learning and inference tasks with low latency and high energy efficiency. However, there are an ever increasing number of design hyperparameters related to the learning algorithms, neuron and synaptic models, neural network architectures and neuromorphic hardware that are required to design the SNN model. These hyperparameters determine the accuracy and en-ergy efficiency of the trained model and often require application, hardware, and software expertise, and are often determined by trial and error, which is a difficult, ad-hoc, and time consuming process. Therefore, there is a need for an automated hyperpa-rameter tuning framework that can search a set of software and hardware hyperparameters to find an SNN model with optimal application accuracy and energy efficiency, when inferred on a NmC. To this end, we propose a framework to co-optimize the accuracy and energy consumption of an SNN model executed on an cross-bar based NmC. The proposed framework integrates three key components: (1) SNNTorch, to train and validate SNN model, (2) SNN-Neurosim, a custom energy estimator for cross-bar based NmC and (3) DeepHyper, a scalable hyperparameter tuning approach to explore the hyperparameter search space of the SNN model and neuromorphic hardware. We evaluate the framework using SNN models trained for scientific applications to (1) parameterize the planetary boundary layer (PBL) using data from the Weather Research Forecast (WRF) model, (2) detect Bragg diffraction peaks for use in X-ray based diffraction microscopy and (3) the reconstruction of high-fidelity amplitude and phase for ptychographic imaging. Adarsha Balaji, Prasanna Balaprakash |
ICMLA | 1 |
| 2022 | Design of Many-Core Big Little µBrains for Energy-Efficient Embedded Neuromorphic ComputingabstractAs spiking-based deep learning inference applications are increasing in embedded systems, these systems tend to integrate neuromorphic accelerators such as µBrain to improve energy efficiency. We propose a µBrain-based scalable many-core neuromorphic hardware design to accelerate the computations of spiking deep convolutional neural networks (SDCNNs). To increase energy efficiency, cores are designed to be heterogeneous in terms of their neuron and synapse capacity (i.e., big vs. little cores), and they are interconnected using a parallel segmented bus interconnect, which leads to lower latency and energy compared to a traditional mesh-based Network-on-Chip (NoC). We propose a system software framework called SentryOS to map SDCNN inference applications to the proposed design. SentryOS consists of a compiler and a run-time manager. The compiler compiles an SDCNN application into sub-networks by exploiting the internal architecture of big and little µBrain cores. The run-time manager schedules these sub-networks onto cores and pipeline their execution to improve throughput. We evaluate the proposed big little many-core neuromorphic design and the system software framework with five commonly-used SDCNN inference applications and show that the proposed solution reduces energy (between 37% and 98%), reduces latency (between 9% and 25%), and increases application throughput (between 20% and 36%). We also show that SentryOS can be easily extended for other spiking neuromorphic accelerators such as Loihi and DYNAPs. M. Lakshmi Varshika, Adarsha Balaji, Federico Corradi, Anup Das 0001, Jan Stuijt, Francky Catthoor |
DATE | 2 |
| 2022 | Design-Technology Co-Optimization for NVM-Based Neuromorphic Processing ElementsabstractAn emerging use case of machine learning (ML) is to train a model on a high-performance system and deploy the trained model on energy-constrained embedded systems. Neuromorphic hardware platforms, which operate on principles of the biological brain, can significantly lower the energy overhead of an ML inference task, making these platforms an attractive solution for embedded ML systems. We present a design-technology tradeoff analysis to implement such inference tasks on the processing elements (PEs) of a non-volatile memory (NVM)-based neuromorphic hardware. Through detailed circuit-level simulations at scaled process technology nodes, we show the negative impact of technology scaling on the information-processing latency, which impacts the quality of service of an embedded ML system. At a finer granularity, the latency inside a PE depends on (1) the delay introduced by parasitic components on its current paths, and (2) the varying delay to sense different resistance states of its NVM cells. Based on these two observations, we make the following three contributions. First, on the technology front, we propose an optimization scheme where the NVM resistance state that takes the longest time to sense is set on current paths having the least delay, and vice versa, reducing the average PE latency, which improves the quality of service. Second, on the architecture front, we introduce isolation transistors within each PE to partition it into regions that can be individually power-gated, reducing both latency and energy. Finally, on the system-software front, we propose a mechanism to leverage the proposed technological and architectural enhancements when implementing an ML inference task on neuromorphic PEs of the hardware. Evaluations with a recent neuromorphic hardware architecture show that our proposed design-technology co-optimization approach improves both performance and energy efficiency of ML inference tasks without incurring high cost-per-bit. Shihao Song, Adarsha Balaji, Anup Das 0001, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic HardwareabstractSpiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches. Shihao Song, Harry Chong, Adarsha Balaji, Anup Das 0001, James A. Shackleford, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | On the role of system software in energy management of neuromorphic computingabstractNeuromorphic computing systems such as DYNAPs and Loihi have recently been introduced to the computing community to improve performance and energy efficiency of machine learning programs, especially those that are implemented using Spiking Neural Network (SNN). The role of a system software for neuromorphic systems is to cluster a large machine learning model (e.g., with many neurons and synapses) and map these clusters to the computing resources of the hardware. In this work, we formulate the energy consumption of a neuromorphic hardware, considering the power consumed by neurons and synapses, and the energy consumed in communicating spikes on the interconnect. Based on such formulation, we first evaluate the role of a system software in managing the energy consumption of neuromorphic systems. Next, we formulate a simple heuristic-based mapping approach to place the neurons and synapses onto the computing resources to reduce energy consumption. We evaluate our approach with 10 machine learning applications and demonstrate that the proposed mapping approach leads to a significant reduction of energy consumption of neuromorphic computing systems. Twisha Titirsha, Shihao Song, Adarsha Balaji, Anup Das 0001 |
CF | 3 |
| 2021 | Dynamic Reliability Management in Neuromorphic ComputingabstractNeuromorphic computing systems execute machine learning tasks designed with spiking neural networks. These systems are embracing non-volatile memory to implement high-density and low-energy synaptic storage. Elevated voltages and currents needed to operate non-volatile memories cause aging of CMOS-based transistors in each neuron and synapse circuit in the hardware, drifting the transistor’s parameters from their nominal values. If these circuits are used continuously for too long, the parameter drifts cannot be reversed, resulting in permanent degradation of circuit performance over time, eventually leading to hardware faults. Aggressive device scaling increases power density and temperature, which further accelerates the aging, challenging the reliable operation of neuromorphic systems. Existing reliability-oriented techniques periodically de-stress all neuron and synapse circuits in the hardware at fixed intervals, assuming worst-case operating conditions, without actually tracking their aging at run-time. To de-stress these circuits, normal operation must be interrupted, which introduces latency in spike generation and propagation, impacting the inter-spike interval and hence, performance (e.g., accuracy). We observe that in contrast to long-term aging, which permanently damages the hardware, short-term aging in scaled CMOS transistors is mostly due to bias temperature instability. The latter is heavily workload-dependent and, more importantly, partially reversible. We propose a new architectural technique to mitigate the aging-related reliability problems in neuromorphic systems by designing an intelligent run-time manager (NCRTM), which dynamically de-stresses neuron and synapse circuits in response to the short-term aging in their CMOS transistors during the execution of machine learning workloads, with the objective of meeting a reliability target. NCRTM de-stresses these circuits only when it is absolutely necessary to do so, otherwise reducing the performance impact by scheduling de-stress operations off the critical path. We evaluate NCRTM with state-of-the-art machine learning workloads on a neuromorphic hardware. Our results demonstrate that NCRTM significantly improves the reliability of neuromorphic hardware, with marginal impact on performance. Shihao Song, Jui Hanamshet, Adarsha Balaji, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Nagarajan Kandasamy, Francky Catthoor |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | PyCARL: A PyNN Interface for Hardware-Software Co-Simulation of Spiking Neural NetworkabstractWe present PyCARL, a PyNN-based common Python programming interface for hardware-software cosimulation of spiking neural network (SNN). Through PyCARL, we make the following two key contributions. First, we provide an interface of PyNN to CARLsim, a computationally- efficient, GPU-accelerated and biophysically-detailed SNN simulator. PyCARL facilitates joint development of machine learning models and code sharing between CARLsim and PyNN users, promoting an integrated and larger neuromorphic community. Second, we integrate cycle-accurate models of state-of-the-art neuromorphic hardware such as TrueNorth, Loihi, and DynapSE in PyCARL, to accurately model hardware latencies, which delay spikes between communicating neurons, degrading performance of machine learning models. PyCARL allows users to analyze and optimize the performance difference between software-based simulation and hardware-oriented simulation. We show that system designers can also use PyCARL to perform design-space exploration early in the product development stage, facilitating faster time-to-market of neuromorphic products. Adarsha Balaji, Prathyusha Adiraju, Hirak J. Kashyap, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Francky Catthoor |
IJCNN | 1 |
| 2020 | Compiling Spiking Neural Networks to Neuromorphic HardwareabstractMachine learning applications that are implemented with spike-based computation model, e.g., Spiking Neural Network (SNN), have a great potential to lower the energy consumption when executed on a neuromorphic hardware. How- ever, compiling and mapping an SNN to the hardware is challenging, especially when compute and storage resources of the hardware (viz. crossbars) need to be shared among the neurons and synapses of the SNN. We propose an approach to analyze and compile SNNs on resource-constrained neuromorphic hardware, providing guarantees on key performance metrics such as execution time and throughput. Our approach makes the following three key contributions. First, we propose a greedy technique to partition an SNN into clusters of neurons and synapses such that each cluster can fit on to the resources of a crossbar. Second, we exploit the rich semantics and expressiveness of Synchronous Dataflow Graphs (SDFGs) to represent a clustered SNN and analyze its performance using Max-Plus Algebra, considering the available compute and storage capacities, buffer sizes, and communication bandwidth. Third, we propose a self-timed execution-based fast technique to compile and admit SNN-based applications to a neuromorphic hardware at run-time, adapting dynamically to the available resources on the hard- ware. We evaluate our approach with standard SNN-based applications and demonstrate a significant performance improvement compared to current practices. Shihao Song, Adarsha Balaji, Anup Das 0001, Nagarajan Kandasamy, James A. Shackleford |
LCTES | 2 |
| 2020 | Mapping Spiking Neural Networks to Neuromorphic HardwareabstractNeuromorphic hardware implements biological neurons and synapses to execute a spiking neural network (SNN)-based machine learning. We present SpiNeMap, a design methodology to map SNNs to crossbar-based neuromorphic hardware, minimizing spike latency and energy consumption. SpiNeMap operates in two steps: SpiNeCluster and SpiNePlacer. SpiNeCluster is a heuristic-based clustering technique to partition an SNN into clusters of synapses, where intracluster local synapses are mapped within crossbars of the hardware and intercluster global synapses are mapped to the shared interconnect. SpiNeCluster minimizes the number of spikes on global synapses, which reduces spike congestion and improves application performance. SpiNePlacer then finds the best placement of local and global synapses on the hardware using a metaheuristic-based approach to minimize energy consumption and spike latency. We evaluate SpiNeMap using synthetic and realistic SNNs on a state-of-the-art neuromorphic hardware. We show that SpiNeMap reduces average energy consumption by 45% and spike latency by 21%, compared to the best-performing SNN mapping technique. Adarsha Balaji, Francky Catthoor, Anup Das 0001, Yuefeng Wu, Khanh Huynh, Francesco Dell'Anna, Giacomo Indiveri, Jeffrey L. Krichmar, Nikil Dutt, Siebren Schaafsma |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Design Methodology for Embedded Approximate Artificial Neural NetworksabstractArtificial neural networks (ANNs) have demonstrated significant promise while implementing recognition and classification applications. The implementation of pre-trained ANNs on embedded systems requires representation of data and design parameters in low-precision fixed-point formats; which often requires retraining of the network. For such implementations, the multiply-accumulate operation is the main reason for resultant high resource and energy requirements. To address these challenges, we present Rox-ANN, a design methodology for implementing ANNs using processing elements (PEs) designed with low-precision fixed-point numbers and high performance and reduced-area approximate multipliers on FPGAs. The trained design parameters of the ANN are analyzed and clustered to optimize the total number of approximate multipliers required in the design. With our methodology, we achieve insignificant loss in application accuracy. We evaluated the design using a LeNet based implementation of the MNIST digit recognition application. The results show a 65.6%, 55.1% and 18.9% reduction in area, energy consumption and latency for a PE using 8-bit precision weights and activations and approximate arithmetic units, when compared to 16-bit full precision, accurate arithmetic PEs. Adarsha Balaji, Salim Ullah, Anup Das 0001, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | Exploration of Segmented Bus As Scalable Global Interconnect for Neuromorphic ComputingabstractSpiking Neural Networks (SNNs) are efficient computation models for spatio-temporal pattern recognition on resource and power constrained platforms. Dedicated SNN hardware, also called neuromorphic hardware, can further reduce the energy consumption of these platforms. A neuromorphic hardware consists of crossbars, which are arrangements of input and output neurons with fully-connected synapses. Time-multiplexed interconnects are used to communicate spikes between crossbars. When a SNN model is mapped on multiple crossbars, the time-multiplexed interconnect increases spike latency and energy consumption, and disorders spike arrivals at output neurons, which reduces application accuracy. In this paper, we propose segmented bus interconnect for global synapses in a neuromorphic architecture. The objective is to reduce power consumption and enable parallel processing compared to traditional time-multiplexed interconnects. The fundamental idea for the segmented bus is to partition a single bus into several segments, with the segmentation switches controlled by software. We evaluate the scalability of segmented bus using synthetic applications. Our results show that segmented bus reduces the latency and energy consumption of the global synapse network significantly with respect to state-of-the-art techniques. Adarsha Balaji, Yuefeng Wu, Anup Das 0001, Francky Catthoor, Siebren Schaafsma |
ACM Great Lakes Symposium on VLSI | 1 |