EDBT 2026 Demo / reviewers in the wild / expert
M. Lakshmi Varshika
dblp:300/8341 · also Mirtinti Lakshmi Varshika
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-4828-1142ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Online Learning for Dynamic Structural Characterization in Electron Energy Loss SpectroscopyabstractIn-situ Electron Energy Loss Spectroscopy (EELS) is a crucial technique for determining the elemental composition of materials through EELS Spectrum Images (EELS-SI). While recent innovations have made it possible for EELS-SI data acquisition at rates of 400 frames per second with near-zero read noise, the challenge lies in processing this massive stream of real-time data to capture nanoscale dynamic changes. This task demands advanced machine learning methods capable of identifying subtle and complex features in EELS spectra. Furthermore, the EELS data acquired in difficult experimental conditions often suffer from a low signal-to-noise ratio (SNR), leading to unreliable classification and limiting their utility. In response to this critical need, we introduce a spiking neural network (SNN)-based Variational Autoencoder (VAE) that embeds spectral data into a latent space, facilitating precise prediction of structural changes. VAEs are designed to learn efficient low-dimensional representations while capturing the inherent variability in the data, making them highly effective for processing multidimensional data. Additionally, SNNs, which use biological neurons, offer unmatched scalability and energy efficiency by processing information through binary spikes, making them ideal for high-throughput data. We validate our framework using MXene annealing data, achieving denoised spectrum images with an SNR of 28.3dB. For the first time, we present a fully online learning solution for dynamic structural tracking, implemented directly in hardware, eliminating the traditional bottleneck of offline training. Our method achieves reliable, real-time, on-device characterization of high-speed EELS data when evaluated on an FPGA platform. Joint experiments with the SNN-VAE model on both spiking autoencoder hardware and a softwaretrained hybrid configuration of hardware spiking encoders demonstrated latency reductions of 25.2x, 93.7x, and 1.04x, 4.5x in energy savings, respectively, compared to baseline. M. Lakshmi Varshika, Jonathan Hollenbach, Nicolas Bohm Agostini, Ankur Limaye, Antonino Tumeo, Anup Das 0001 |
DATE | 1 |
| 2025 | Optimizing Memory Latency and Bandwidth of Spiking Neural Network Accelerators on FPGA via Sparse HashingabstractSpiking Neural Networks (SNNs) exploit the natural temporal sparsity by processing discrete events. Introducing additional weight sparsity can further improve their energy efficiency when implemented in hardware. We propose SPARSH, a novel memory organization technique that uses sparse hashing to allow efficient storage and fast access while minimizing memory over-provisioning. Through SPARSH, we make the following four key contributions. First, we introduce an efficient hash function to evenly distribute nonzero weights to buckets that are optimized for an FPGA memory system and placed in a compact memory space to achieve storage efficiency. Second, we use a controller that integrates a Bloom filter to detect sparsity and skip memory accesses for weights and activations that are zero, and for neurons that are in refractory state. This improves memory bandwidth utilization. Third, we propose a scheduler that improves memory latency by prioritizing accesses that hit in the same bucket over other accesses. Finally, we propose an algorithm to select the design parameters of SPARSH based on the sparsity of a target SNN. We implement SPARSH for a recent SNN accelerator on a Virtex UltraScale and evaluate using seven SNNs models. We show that SPARSH reduces memory over-provisioning by 3.2×, access latency by 77%, and bandwidth utilization by 68% with a marginal increase in resource utilization. Shadi Matinizadeh, M. Lakshmi Varshika, Anup Das 0001 |
ICCAD | 2 |
| 2025 | Neuromorphic Architectures for Scientific Computing: a Structural Characterization Case StudyabstractNeuromorphic computing offers a promising paradigm for energy-efficient edge processing in scientific applications, such as the real-time analysis of Electron Energy Loss Spectroscopy (EELS) data from Transmission Electron Microscopes (TEMs). Current methods, primarily based on Spiking Variational Autoencoders (S-VAE), are constrained by high computational overhead. To address this, we propose an energy-efficient Spiking Hopfield Network (S-Hopfield) for online encoding and decoding of structural dynamics. Our approach leverages the inherent associative memory of Hopfield networks to robustly denoise and reconstruct spectral images, outperforming an S-VAE model in both image quality metrics and hardware efficiency. Quantitatively, the S-Hopfield network achieved a Mean Squared Error (MSE) of 0.54, a 28% improvement over the S-VAE’s MSE of 0.75. On a Xilinx Virtex-7 FPGA, the S-Hopfield’s core inference engine consumed a mere 0.25 W, representing a 51% reduction in power compared to the S-VAE’s 0.51 W. These results demonstrate that the S-Hopfield network provides a superior, low-power solution for real-time spectral analysis at the edge, paving the way for autonomous experimental control in material science. M. Lakshmi Varshika, Jonathan Hollenbach, Nicolas Bohm Agostini, Ankur Limaye, Marco Minutoli, Vito Giovanni Castellana, Joseph B. Manzano, Anup Das 0001, Mitra Taheri, Antonino Tumeo |
ICCAD | 1 |
| 2025 | A Digital Neuromorphic Architecture for Unsupervised Shortest Path Computation on Real-World GraphsabstractGraphs are popular tools for analyzing interconnected data entities. We propose SENTIENCE, a novel approach to computing the shortest path in a graph in an unsupervised manner drawing inspiration from the hippocampus. SENTIENCE uses laterally-connected neurons to represent nodes and synapses to represent edges. The strength of a synaptic connection is encoded as the axonal delay. When a neuron (source) is excited, a wavefront of neural activity is created that propagates through the graph via the connected nodes. SENTIENCE uses the Eligibility Propagation (E-Prop) algorithm to learn the sequence of wave movement through the shortest path in a graph, which can be obtained by identifying earliest firing (eligible) neighbors by tracing back in time from when the wavefront reaches a destination. We introduce a lightweight hardware design for the axonal plasticity and the E-Prop learning mechanism of SENTIENCE. We propose a tile-based architecture to address the scalability of SENTIENCE for real-world graphs with irregular data dependency. We evaluate SENTIENCE on a Versal VPK 180 FPGA and show that SENTIENCE consumes on average 10× less resources and 12× less power compared to a state-of-the-art. Arghavan Mohammadhassani, Shadi Matinizadeh, M. Lakshmi Varshika, Anup Das 0001 |
ISCAS | 3 |
| 2023 | Hardware-Software Co-Design for On-Chip Learning in AI SystemsabstractSpike-based convolutional neural networks (CNNs) are empowered with on-chip learning in their convolution layers, enabling the layer to learn to detect features by combining those extracted in the previous layer. We propose ECHELON, a generalized design template for a tile-based neuromorphic hardware with on-chip learning capabilities. Each tile in ECHELON consists of a neural processing units (NPU) to implement convolution and dense layers of a CNN model, an on-chip learning unit (OLU) to facilitate spike-timing dependent plasticity (STDP) in the convolution layer, and a special function unit (SFU) to implement other CNN functions such as pooling, concatenation, and residual computation. These tile resources are interconnected using a shared bus, which is segmented and configured via the software to facilitate parallel communication inside the tile. Tiles are themselves interconnected using a classical Network-on-Chip (NoC) interconnect. We propose a system software to map CNN models to ECHELON, maximizing the performance. We integrate the hardware design and software optimization within a co-design loop to obtain the hardware and software architectures for a target CNN, satisfying both performance and resource constraints. In this preliminary work, we show the implementation of a tile on a FPGA and some early evaluations. Using 8 STDP-enabled CNN models, we show the potential of our co-design methodology to optimize hardware resources. M. Lakshmi Varshika, Abhishek Kumar Mishra 0002, Nagarajan Kandasamy, Anup Das 0001 |
ASP-DAC | 1 |
| 2022 | A design methodology for fault-tolerant computing using astrocyte neural networksabstractWe propose a design methodology to facilitate fault tolerance of deep learning models. First, we implement a many-core fault-tolerant neuromorphic hardware design, where neuron and synapse circuitries in each neuromorphic core are enclosed with astrocyte circuitries, the star-shaped glial cells of the brain that facilitate self-repair by restoring the spike firing frequency of a failed neuron using a closed-loop retrograde feedback signal. Next, we introduce astrocytes in a deep learning model to achieve the required degree of tolerance to hardware faults. Finally, we use a system software to partition the astrocyte-enabled model into clusters and implement them on the proposed fault-tolerant neuromorphic design. We evaluate this design methodology using seven deep learning inference models and show that it is both area- and power-efficient. Murat Isik, Ankita Paul, M. Lakshmi Varshika, Anup Das 0001 |
CF | 3 |
| 2022 | Design of Many-Core Big Little µBrains for Energy-Efficient Embedded Neuromorphic ComputingabstractAs spiking-based deep learning inference applications are increasing in embedded systems, these systems tend to integrate neuromorphic accelerators such as µBrain to improve energy efficiency. We propose a µBrain-based scalable many-core neuromorphic hardware design to accelerate the computations of spiking deep convolutional neural networks (SDCNNs). To increase energy efficiency, cores are designed to be heterogeneous in terms of their neuron and synapse capacity (i.e., big vs. little cores), and they are interconnected using a parallel segmented bus interconnect, which leads to lower latency and energy compared to a traditional mesh-based Network-on-Chip (NoC). We propose a system software framework called SentryOS to map SDCNN inference applications to the proposed design. SentryOS consists of a compiler and a run-time manager. The compiler compiles an SDCNN application into sub-networks by exploiting the internal architecture of big and little µBrain cores. The run-time manager schedules these sub-networks onto cores and pipeline their execution to improve throughput. We evaluate the proposed big little many-core neuromorphic design and the system software framework with five commonly-used SDCNN inference applications and show that the proposed solution reduces energy (between 37% and 98%), reduces latency (between 9% and 25%), and increases application throughput (between 20% and 36%). We also show that SentryOS can be easily extended for other spiking neuromorphic accelerators such as Loihi and DYNAPs. M. Lakshmi Varshika, Adarsha Balaji, Federico Corradi, Anup Das 0001, Jan Stuijt, Francky Catthoor |
DATE | 1 |
| 2021 | A Design Flow for Mapping Spiking Neural Networks to Many-Core Neuromorphic HardwareabstractThe design of many-core neuromorphic hardware is becoming increasingly complex as these systems are now expected to execute large machine-learning models. A predictable design flow is needed to guarantee real-time performance such as latency and throughput without significantly increasing the buffer requirement of computing cores. Synchronous Data Flow Graphs (SDFGs) have been previously used for predictable mapping of streaming applications to multiprocessor systems. We propose an SDFG-based design flow to map spiking neural networks (SNNs) to many-core neuromorphic hardware with the objective of exploring the tradeoff between throughput and buffer-size requirements. The proposed design flow integrates an iterative partitioning approach based on Kernighan-Lin graph partitioning heuristic to create SNN clusters such that each cluster can be mapped to a core of the hardware. The partitioning approach minimizes inter-cluster spike communication, which improves latency on the shared interconnect of the hardware. Next, the design flow maps clusters to cores using Particle Swarm Optimization (PSO), an evolutionary algorithm, while exploring the design space of throughput and buffer size. Pareto-optimal mappings are retained from the design flow, allowing system designers to select a Pareto mapping that satisfies throughput and buffer-size requirements of the design. We evaluated the developed design flow using five large-scale convolutional neural network (CNN) models. Results demonstrate 63% higher maximum throughput and 10% lower buffer-size requirement compared to state-of-the-art dataflow-based mapping solutions. Shihao Song, M. Lakshmi Varshika, Anup Das 0001, Nagarajan Kandasamy |
ICCAD | 2 |