Gourav Datta

dblp:250/9607 · DBLP profile ↗
← Back
24ranked-venue papers
13as first author
23since 2021 · last 2026
0000-0002-5380-2619ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Shallow Enough? A Cross-Architecture Study of Ultra-Low-Depth Neural Networks for Edge Inference
abstract
Edge deployment imposes strict latency, memory, and energy constraints that scale directly with network depth, yet the question of which architectural paradigm offers the best accuracy-efficiency tradeoff at ultra-low depth (four to six layers) remains open. We present the first systematic comparison of convolutional, pure transformer, and hybrid CNN-transformer models under fixed shallow depth budgets. We develop an analytical framework characterizing the representational cost of local convolutional versus global attention-based feature mixing as a function of depth, and validate it with experiments on ImageNet-1K across architectures spanning MobileNetV2, DeiT, Swin, MobileViT, and EfficientFormer. Beyond FLOPs and accuracy, we report real-world latency and energy measurements on an NVIDIA Jetson Nano and examine deployment feasibility on Cortex-M class microcontrollers. Our experiments reveal how accuracy, latency, and memory scale across all three architecture families as depth decreases, providing practitioners with direct guidance on which model class to choose for a given depth and hardware budget.
Chengwei Zhou, Haotian Yu, Shoma Yukawa, Deniz Najafi, Shaahin Angizi, Gourav Datta
ACM Great Lakes Symposium on VLSI6
2025 Analog vs. Digital In-Sensor Computing: A Tale of Two Paradigms
Gourav Datta, Chengwei Zhou
ACM Great Lakes Symposium on VLSI1
2025 Dynamic SpikFormer: Low-Latency & Energy-Efficient Spiking Neural Networks with Dynamic Time Steps for Vision Transformers
abstract
Spiking Neural Networks (SNNs) have emerged as a popular spatio-temporal computing paradigm for complex vision tasks. Recently proposed SNN training algorithms have significantly reduced the number of time steps (down to 1) for improved latency and energy efficiency, however, they target only convolutional neural networks (CNN). These algorithms, when applied to the recently spotlighted vision transformers (ViT), either require a large number of time steps or fail to converge. Based on the analysis of the histograms of the ANN and SNN activation maps, we hypothesize that each ViT block has a different sensitivity to the number of time steps. We propose a novel training framework that dynamically allocates the number of time steps to each ViT module depending on a trainable score assigned to each timestep. In particular, we generate a scalar binary time step mask that filters spikes emitted by each neuron in a leaky-integrate-and-fire (LIF) layer. The resulting SNNs have high activation sparsity and require only accumulate operations (AC), except for the input embedding layer, in contrast to expensive multiply-and-accumulates (MAC) needed in traditional ViTs. This yields significant improvements in energy efficiency. We evaluate our training framework and resulting SNNs on image recognition tasks including CIFAR10, CIFAR100, and ImageNet with different ViT architectures. We obtain a test accuracy of 95.97% with 4.97 time steps with direct encoding on CIFAR10.
Gourav Datta, Zeyu Liu 0003, Anni Li, Peter A. Beerel
ICASSP1
2025 Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics
abstract
Vision Transformers (ViTs) have emerged as a powerful architecture for computer vision tasks due to their ability to model long-range dependencies and global contextual relationships. However, their substantial compute and memory demands hinder efficient deployment in scenarios with strict energy and bandwidth limitations. In this work, we propose Opto-ViT, the first near-sensor, region-aware ViT accelerator leveraging silicon photonics (SiPh) for real-time and energy-efficient vision processing. Opto-ViT features a hybrid electronic-photonic architecture, where the optical core handles compute-intensive matrix multiplications using Vertical-Cavity Surface-Emitting Lasers (VCSELs) and Microring Resonators (MRs), while nonlinear functions and normalization are executed electronically. To reduce redundant computation and patch processing, we introduce a lightweight Mask Generation Network (MGNet) that identifies regions of interest in the current frame and prunes irrelevant patches before ViT encoding. We further co-optimize the ViT backbone using quantization-aware training and matrix decomposition tailored for photonic constraints. Experiments across device fabrication, circuit and architecture co-design, to classification, detection, and video tasks demonstrate that Opto-ViT achieves 100.4 KFPS/W with up to 84% energy savings with less than 1.6% accuracy loss, while enabling scalable and efficient ViT deployment at the edge.
Mehrdad Morsali, Chengwei Zhou, Deniz Najafi, Sreetama Sarkar, Pietro Mercati, Navid Khoshavi, Peter A. Beerel, Mahdi Nikdast, Gourav Datta, Shaahin Angizi
ICCAD9
2025 GAN-BiLSTM-HDC: A Hybrid Framework for Robust and Hardware-Efficient Malware Detection
abstract
Hyperdimensional Computing (HDC) has emerged as a hardware-efficient paradigm for embedded malware detection, offering strong parallelism and low complexity. However, the accuracy and robustness of HDC classifiers remain highly dependent on the diversity and quality of training data, leaving them vulnerable to novel threats. To address this challenge, we introduce a generative adversarial network (GAN)-assisted augmentation framework for the Microprocessor without Interlocked Pipelined Stages-32 (MIPS32) malware generation. The GAN is trained on real-world MIPS32 malware binaries to produce previously unseen instruction sequences. The synthetic code stacks are filtered using a custom MIPS32 assembler for syntactic validation and a Bidirectional Long Short-Term Memory (BiLSTM)-based semantic critic to ensure logical coherence. Only validated samples are retained to expand the training set for the HDC classifier, thereby strengthening generalization and resilience against novel malware variants. Our preliminary results show an average generator loss ($\mathbf{G}$) of 2.68 over 200 epochs and a discriminator loss (D) converging to 0.63, indicating that the GAN is learning to generate realistic and diverse outputs. This hybrid GAN-BiLSTM-HDC framework shows strong potential for enhancing classification accuracy, resilience, and efficiency in resource-constrained, real-time malware detection systems.
Emilien Meyer, Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Lida Kouhalvandi, Gourav Datta, Sercan Aygün, M. Hassan Najafi
ICCD5
2025 MaskVD: Region Masking for Efficient Video Object Detection
abstract
Video tasks are compute-heavy and thus pose a challenge when deploying in real-time applications, particularly for tasks that require state-of-the-art Vision Trans-formers (ViTs). Several research efforts have tried to address this challenge by leveraging the fact that large portions of the video undergo very little change across frames leading to redundant computations in frame-based video processing. In particular, some works leverage pixel or semantic differences across frames, however, this yields limited latency benefits with significantly increased memory overhead. This paper, in contrast, presents a strategy for masking regions in video frames that leverages the semantic information in images and the temporal correlation between frames to significantly reduce FLOPs and latency with little to no penalty in performance over base-line models. In particular, we demonstrate that by lever-aging extracted features from previous frames, ViT back-bones directly benefit from region masking, skipping up to 80% of input regions, improving FLOPs and latency by 3.14x and 1.5x. We improve memory and latency over the state-of-the-art (SOTA) by 2.3x and 1.14x, while maintaining similar detection performance. Additionally, our approach demonstrates promising results on convolutional neural networks (CNNs) and provides latency improvements over the SOTA up to 1.3x using specialized computational kernels. Code is available at: https://github.com/sreetamasarkar/MaskVD
Sreetama Sarkar, Gourav Datta, Souvik Kundu 0002, Chirayata Bhattacharyya, Peter A. Beerel
WACV2
2024 Toward High-Accuracy, Programmable Extreme-Edge Intelligence for Neuromorphic Vision Sensors utilizing Magnetic Domain Wall Motion-based MTJ
abstract
The desire to empower resource-limited edge devices with computer vision (CV) must overcome the high energy consumption of collecting and processing vast sensory data. To address the challenge, this work proposes an energy-efficient non-von-Neumann in-pixel processing solution for neuromorphic vision sensors employing emerging (X) magnetic domain wall magnetic tunnel junction (MDWMTJ) for the first time, in conjunction with CMOS-based neuromorphic pixels. Our hybrid CMOS+X approach performs in-situ massively parallel asynchronous analog convolution, exhibiting low power consumption and high accuracy across various CV applications by leveraging the non-volatility and programmability of the MDWMTJ. Moreover, our developed device-circuit-algorithm co-design framework captures device constraints (low tunnel-magnetoresistance, low dynamic range) and circuit constraints (non-linearity, process variation, area consideration) based on monte-carlo simulations and device parameters utilizing GF22nm FD-SOI technology. Our experimental results suggest we can achieve an average of 45.3% reduction in backend-processor energy, maintaining similar front-end energy compared to the state-of-the-art and high accuracy of 79.17% and 95.99% on the DVS-CIFAR10 and IBM DVS128-Gesture datasets, respectively.
Md. Abdullah-Al Kaiser, Gourav Datta, Peter A. Beerel, Akhilesh Jaiswal 0001
DAC2
2024 Training Ultra-Low-Latency Spiking Neural Networks from Scratch
abstract
Spiking Neural networks (SNN) have emerged as an attractive spatio-temporal computing paradigm for a wide range of low-power vision tasks. However, state-of-the-art (SOTA) SNN models either incur multiple time steps which hinder their deployment in real-time use cases or increase the training complexity significantly. To mitigate this concern, we present a training framework (from scratch) for SNNs with ultra-low (down to 1) time steps that leverages the Hoyer regularizer. We calculate the threshold for each BANN layer as the Hoyer extremum of a clipped version of its activation map. The clipping value is determined through training using gradient descent with our Hoyer regularizer. We evaluate the efficacy of our training framework on large-scale vision tasks, including traditional and event-based image recognition and object detection. Our experiments demonstrate up to 34× increase in compute efficiency with a marginal accuracy/mAP drop compared to non-spiking networks. Finally, we implement our framework in the Lava-DL library, thereby enabling the deployment of our SNN models in the Loihi neuromorphic chip.
Gourav Datta, Zeyu Liu 0003, Peter A. Beerel
ICASSP1
2024 Can we get the best of both Binary Neural Networks and Spiking Neural Networks for Efficient Computer Vision?
abstract
Binary Neural networks (BNN) have emerged as an attractive computing paradigm for a wide range of low-power vision tasks. However, state-of-the-art (SOTA) BNNs do not yield any sparsity, and induce a significant number of non-binary operations. On the other hand, activation sparsity can be provided by spiking neural networks (SNN), that too have gained significant traction in recent times. Thanks to this sparsity, SNNs when implemented on neuromorphic hardware, have the potential to be significantly more power-efficient compared to traditional artifical neural networks (ANN). However, SNNs incur multiple time steps to achieve close to SOTA accuracy. Ironically, this increases latency and energy---costs that SNNs were proposed to reduce---and presents itself as a major hurdle in realizing SNNs’ theoretical gains in practice. This raises an intriguing question: *Can we obtain SNN-like sparsity and BNN-like accuracy and enjoy the energy-efficiency benefits of both?* To answer this question, in this paper, we present a training framework for sparse binary activation neural networks (BANN) using a novel variant of the Hoyer regularizer. We estimate the threshold of each BANN layer as the Hoyer extremum of a clipped version of its activation map, where the clipping value is trained using gradient descent with our Hoyer regularizer. This approach shifts the activation values away from the threshold, thereby mitigating the effect of noise that can otherwise degrade the BANN accuracy. Our approach outperforms existing BNNs, SNNs, and adder neural networks (that also avoid energy-expensive multiplication operations similar to BNNs and SNNs) in terms of the accuracy-FLOPs trade-off for complex image recognition tasks. Downstream experiments on object detection further demonstrate the efficacy of our approach. Lastly, we demonstrate the portability of our approach to SNNs with multiple time steps. Codes are publicly available [here](https://github.com/godatta/Ultra-Low-Latency-SNN).
Gourav Datta, Zeyu Liu 0003, Peter A. Beerel
ICLR1
2024 LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units
abstract
Transformer models have demonstrated high accuracy in numerous applications but have high complexity and lack sequential processing capability making them ill-suited for many streaming applications at the edge where devices are heavily resource-constrained. Thus motivated, many researchers have proposed reformulating the transformer models as RNN modules which modify the self-attention computation with explicit states. However, these approaches often incur significant performance degradation. The ultimate goal is to develop a model that has the following properties: parallel training, streaming and low-cost inference, and state-of-the-art (SOTA) performance. In this paper, we propose a new direction to achieve this goal. We show how architectural modifications to a fully-sequential recurrent model can help push its performance toward Transformer models while retaining its sequential processing capability. Specifically, inspired by the recent success of Legendre Memory Units (LMU) in sequence learning tasks, we propose LMUFormer, which augments the LMU with convolutional patch embedding and convolutional channel mixer. Moreover, we present a spiking version of this architecture, which introduces the benefit of states within the patch embedding and channel mixer modules while simultaneously reducing the computing complexity. We evaluated our architectures on multiple sequence datasets. Of particular note is our performance on the Speech Commands V2 dataset (35 classes). In comparison to SOTA transformer-based models within the ANN domain, our LMUFormer demonstrates comparable performance while necessitating a remarkable $70\times$ reduction in parameters and a substantial $140\times$ decrement in FLOPs. Furthermore, when benchmarked against extant low-complexity SNN variants, our model establishes a new SOTA with an accuracy of 96.12\%. Additionally, owing to our model's proficiency in real-time data processing, we are able to achieve a 32.03\% reduction in sequence length, all while incurring an inconsequential decline in performance.
Zeyu Liu 0003, Gourav Datta, Anni Li, Peter A. Beerel
ICLR2
2024 FixPix: Fixing Bad Pixels using Deep Learning
Sreetama Sarkar, Xinan Ye, Gourav Datta, Peter A. Beerel
ICPR (3)3
2023 Technology-Circuit-Algorithm Tri-Design for Processing-in-Pixel-in-Memory (P2M)
abstract
The massive amounts of data generated by camera sensors motivate data processing inside pixel arrays, i.e., at the extreme-edge. Several critical developments have fueled recent interest in the processing-in-pixel-in-memory paradigm for a wide range of visual machine intelligence tasks, including (1) advances in 3D integration technology to enable complex processing inside each pixel in a 3D integrated manner while maintaining pixel density, (2) analog processing circuit techniques for massively parallel low-energy in-pixel computations, and (3) algorithmic techniques to mitigate non-idealities associated with analog processing through hardware-aware training schemes. This article presents a comprehensive technology-circuit-algorithm landscape that connects technology capabilities, circuit design strategies, and algorithmic optimizations to power, performance, area, bandwidth reduction, and application-level accuracy metrics. We present our results using a comprehensive co-design framework incorporating hardware and algorithmic optimizations for various complex real-life visual intelligence tasks mapped onto our P2M paradigm.
Md. Abdullah-Al Kaiser, Gourav Datta, Sreetama Sarkar, Souvik Kundu 0002, Zihan Yin, Manas Garg, Ajey P. Jacob, Peter A. Beerel, Akhilesh Jaiswal 0001
ACM Great Lakes Symposium on VLSI2
2023 In-Sensor & Neuromorphic Computing Are all You Need for Energy Efficient Computer Vision
abstract
Due to the high activation sparsity and use of accumulates (AC) instead of expensive multiply-and-accumulates (MAC), neuromorphic spiking neural networks (SNNs) have emerged as a promising low-power alternative to traditional DNNs for several computer vision (CV) applications. However, most existing SNNs require multiple time steps for acceptable inference accuracy, hindering real-time deployment and increasing spiking activity and, consequently, energy consumption. Recent works proposed direct encoding that directly feeds the analog pixel values in the first layer of the SNN in order to significantly reduce the number of time steps. Although the overhead for the first layer MACs with direct encoding is negligible for deep SNNs and the CV processing is efficient using SNNs, the data transfer between the image sensors and the downstream processing costs significant bandwidth and may dominate the total energy. To mitigate this concern, we propose an in-sensor computing hardware-software co-design framework for SNNs targeting image recognition tasks. Our approach reduces the bandwidth between sensing and processing by 12−96× and the resulting total energy by 2.32× compared to traditional CV processing, with a 3.8% reduction in accuracy on ImageNet.
Gourav Datta, Zeyu Liu 0003, Md. Abdullah-Al Kaiser, Souvik Kundu 0002, Joe Mathai, Zihan Yin, Ajey P. Jacob, Akhilesh Jaiswal 0001, Peter A. Beerel
ICASSP1
2023 ViTA: A Vision Transformer Inference Accelerator for Edge Applications
abstract
Vision Transformer models, such as ViT, Swin Transformer, and Transformer-in-Transformer, have recently gained significant traction in computer vision tasks due to their ability to capture the global relation between features which leads to superior performance. However, they are compute-heavy and difficult to deploy in resource-constrained edge devices. Existing hardware accelerators, including those for the closely-related BERT transformer models, do not target highly resource-constrained environments. In this paper, we address this gap and propose ViTA - a configurable hardware accelerator for inference of vision transformer models, targeting resource-constrained edge computing devices and avoiding repeated off-chip memory accesses. We employ a head-level pipeline and inter-layer MLP optimizations, and can support several commonly used vision transformer models with changes solely in our control logic. We achieve nearly 90% hardware utilization efficiency on most vision transformer models, report a power of 0.88W when synthesised with a clock of 150 MHz, and get reasonable frame rates - all of which makes ViTA suitable for edge applications.
Shashank Nag, Gourav Datta, Souvik Kundu 0002, Nitin Chandrachoodan, Peter A. Beerel
ISCAS2
2023 Bridging the Gap Between Spiking Neural Networks & LSTMs for Latency & Energy Efficiency
abstract
Spiking Neural Networks (SNNs) have emerged as an attractive spatio-temporal computing paradigm for complex vision tasks. However, most existing works yield models that require many time steps and do not leverage the inherent temporal dynamics of spiking neural networks, even for sequential tasks. Motivated by this observation, we propose an optimized spiking long short-term memory networks (LSTM) training framework that involves a novel ANN-to-SNN conversion framework, followed by SNN fine-tuning via backpropagation through time (BPTT). In particular, we propose novel activation functions in the source LSTM architecture and convert a judiciously selected subset of them to leaky-integrate-and-fire (LIF) activations with optimal bias shifts. Moreover, we propose a pipelined parallel processing scheme that hides the SNN time steps, significantly improving system latency, especially for long sequences. The resulting SNNs have high activation sparsity and require only accumulate operations (AC), in contrast to expensive multiply-and-accumulates (MAC) needed for ANNs, except for the input layer when using direct encoding, yielding significant improvements in energy efficiency. We evaluate our framework on sequential learning tasks including temporal MNIST, Google Speech Commands (GSC), and UCI Smartphone datasets on different LSTM architectures. We obtain test accuracy of 94.75 % with only 2 time steps on the GSC dataset with$\sim 4.1\times$lower energy than an iso-architecture standard LSTM.
Gourav Datta, Haoqin Deng, Robert Aviles, Zeyu Liu 0003, Peter A. Beerel
ISLPED1
2023 Self-Attentive Pooling for Efficient Deep Learning
abstract
Efficient custom pooling techniques that can aggressively trim the dimensions of a feature map for resource-constrained computer vision applications have recently gained significant traction. However, prior pooling works extract only the local context of the activation maps, limiting their effectiveness. In contrast, we propose a novel non-local self-attentive pooling method that can be used as a drop-in replacement to the standard pooling layers, such as max/average pooling or strided convolution. The proposed self-attention module uses patch embedding, multihead self-attention, and spatial-channel restoration, followed by sigmoid activation and exponential soft-max. This self-attention mechanism efficiently aggregates dependencies between non-local activation patches during downsampling. Extensive experiments on standard object classification and detection tasks with various convolutional neural network (CNN) architectures demonstrate the superiority of our proposed mechanism over the state-of-the-art (SOTA) pooling techniques. In particular, we surpass the test accuracy of existing pooling techniques on different variants of MobileNet-V2 on ImageNet by an average of ~1.2%. With the aggressive down-sampling of the activation maps in the initial layers (providing up to 22x reduction in memory consumption), our approach achieves 1.43% higher test accuracy compared to SOTA techniques with iso-memory footprints. This enables the deployment of our models in memory-constrained devices, such as micro-controllers (without losing significant accuracy), because the initial activation maps consume a significant amount of on-chip memory for high-resolution images required for complex vision tasks. Our pooling method also leverages channel pruning to further reduce memory footprints. Codes are available at https://github.com/CFun/Non-Local-Pooling.
Gourav Datta, Souvik Kundu 0002, Peter A. Beerel
WACV2
2023 Enabling ISPless Low-Power Computer Vision
abstract
Current computer vision (CV) systems use an image signal processing (ISP) unit to convert the high resolution raw images captured by image sensors to visually pleasing RGB images. Typically, CV models are trained on these RGB images and have yielded state-of-the-art (SOTA) performance on a wide range of complex vision tasks, such as object detection. In addition, in order to deploy these models on resource-constrained low-power devices, recent works have proposed in-sensor and in-pixel computing approaches that try to partly/fully bypass the ISP and yield significant bandwidth reduction between the image sensor and the CV processing unit by downsampling the activation maps in the initial convolutional neural network (CNN) layers. However, direct inference on the raw images degrades the test accuracy due to the difference in covariance of the raw images captured by the image sensors compared to the ISP-processed images used for training. Moreover, it is difficult to train deep CV models on raw images, because most (if not all) large-scale open-source datasets consist of RGB images. To mitigate this concern, we propose to invert the ISP pipeline, which can convert the RGB images of any dataset to its raw counterparts, and enable model training on raw images. We release the raw version of the COCO dataset, a large-scale benchmark for generic high-level vision tasks. For ISP-less CV systems, training on these raw images result in a ∼7.1% increase in test accuracy on the visual wake works (VWW) dataset compared to relying on training with traditional ISP-processed RGB datasets. To further improve the accuracy of ISP-less CV models and to increase the energy and bandwidth benefits obtained by in-sensor/in-pixel computing, we propose an energy-efficient form of analog in-pixel demosaicing that may be coupled with in-pixel CNN computations. When evaluated on raw images captured by real sensors from the PASCALRAW dataset, our approach results in a 8.1% increase in mAP. Lastly, we demonstrate a further 20.5% increase in mAP by using a novel application of few-shot learning with thirty shots each for the novel PASCALRAW dataset, constituting 3 classes. Codes are available at https://github.com/godatta/ISP-less-CV.
Gourav Datta, Zeyu Liu 0003, Zihan Yin, Linyu Sun, Akhilesh Jaiswal 0001, Peter A. Beerel
WACV1
2022 Can Deep Neural Networks be Converted to Ultra Low-Latency Spiking Neural Networks?
abstract
Spiking neural networks (SNNs), that operate via binary spikes distributed over time, have emerged as a promising energy efficient ML paradigm for resource-constrained devices. However, the current state-of-the-art (SOTA) SNNs require multiple time steps for acceptable inference accuracy, increasing spiking activity and, consequently, energy consumption. SOTA training strategies for SNNs involve conversion from a non-spiking deep neural network (DNN). In this paper, we determine that SOTA conversion strategies cannot yield ultra low latency because they incorrectly assume that the DNN and SNN pre-activation values are uniformly distributed. We propose a new training algorithm that accurately captures these distributions, minimizing the error between the DNN and converted SNN. The resulting SNNs have ultra low latency and high activation sparsity, yielding significant improvements in compute efficiency. In particular, we evaluate our framework on image recognition tasks from CIFAR-10 and CIFAR-100 datasets on several VGG and ResNet architectures. We obtain top-1 accuracy of 64.19% with only 2 time steps on the CIFAR-100 dataset with ∼159.2× lower compute energy compared to an iso-architecture standard DNN. Compared to other SOTA SNN models, our models perform inference 2.5-8× faster (i.e., with fewer time steps).
Gourav Datta, Peter A. Beerel
DATE1
2022 Asd-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers
abstract
Multimodal active speaker detection (ASD) methods assign a speaking/not-speaking label per individual in a video clip. ASD is critical for applications such as natural human-computer interaction, speaker diarization, and video reframing. Recent work has shown the success of transformers in multimodal settings, thus we propose a novel framework that leverages modern transformer and concatenation mechanisms to efficiently capture the interaction between audio and video modalities for ASD. We achieve mAP similar to state-of-the-art (93.0% vs 93.5%) on the AVA-ActiveSpeaker dataset. Further, our model has ~3× smaller size (15.23MB vs 49.82MB), reduced FLOPs count (11.8 vs 14.3), and lower training time (15h vs 38h). To verify our model is making predictions from the right visual cues, we computed saliency maps over input images. We found that in addition to mouth regions, the nose, cheek, and area under the eye were helpful in identifying active speakers. Our ablation study reveals that the mouth region alone achieved lower mAP (91.9% vs 93.0%) compared to full face region, supporting our hypothesis that facial expressions in addition to mouth region are useful for ASD.
Gourav Datta, Tyler Etchart, Vivek Yadav, Varsha Hedau, Pradeep Natarajan, Shih-Fu Chang
ICASSP1
2022 P2M-DeTrack: Processing-in-Pixel-in-Memory for Energy-efficient and Real-Time Multi-Object Detection and Tracking
abstract
Today’s high resolution, high frame rate cameras in autonomous vehicles generate a large volume of data that needs to be transferred and processed by a downstream processor or machine learning (ML) accelerator to enable intelligent computing tasks, such as multi-object detection and tracking. The massive amount of data transfer incurs significant energy, latency, and bandwidth bottlenecks, which hinders real-time processing. To mitigate this problem, we propose an algorithm-hardware co-design framework called Processing-in-Pixel-in-Memory-based object Detection and Tracking (P2M-DeTrack). P2M-DeTrack is based on a custom faster R-CNN-based model that is distributed partly inside the pixel array (front-end) and partly in a separate FPGA/ASIC (back-end). The proposed front-end in-pixel processing down-samples the input feature maps significantly with judiciously optimized strided convolution and pooling. Compared to a conventional baseline design that transfers frames of RGB pixels to the back-end, the resulting P2M-DeTrack designs reduce the data bandwidth between sensor and back-end by up to 24×. The designs also reduce the sensor and total energy (obtained from in-house circuit simulations at Globalfoundries 22nm technology node) per frame by 5.7× and 1.14×, respectively. Lastly, they reduce the sensing and total frame latency by an estimated 1.7× and 3×, respectively. We evaluate our approach on the multi-object object detection (tracking) task of the large-scale BDD100K dataset and observe only a 0.5% reduction in the mean average precision (0.8% reduction in the identification F1 score) compared to the state-of-the-art.
Gourav Datta, Souvik Kundu 0002, Zihan Yin, Joe Mathai, Zeyu Liu 0003, Mulin Tian, Shunlin Lu, Ravi Teja Lakkireddy, Andrew G. Schmidt, Wael Abd-Almageed, Ajey P. Jacob, Akhilesh Jaiswal 0001, Peter A. Beerel
VLSI-SoC1
2021 Training Energy-Efficient Deep Spiking Neural Networks with Single-Spike Hybrid Input Encoding
abstract
Spiking Neural Networks (SNNs) have emerged as an attractive alternative to traditional deep learning frameworks, since they provide higher computational efficiency in event driven neuromorphic hardware. However, the state-of-the-art (SOTA) SNNs suffer from high inference latency, resulting from inefficient input encoding and training techniques. The most widely used input coding schemes, such as Poisson based rate-coding, do not leverage the temporal learning capabilities of SNNs. This paper presents a training framework for low-latency energy-efficient SNNs that uses a hybrid encoding scheme at the input layer in which the analog pixel values of an image are directly applied during the first timestep and a novel variant of spike temporal coding is used during subsequent timesteps. In particular, neurons in every hidden layer are restricted to fire at most once per image which increases activation sparsity. To train these hybrid-encoded SNNs, we propose a variant of the gradient descent based spike timing dependent backpropagation (STDB) mechanism using a novel cross entropy loss function based on both the output neurons' spike time and membrane potential. The resulting SNNs have reduced latency and high activation sparsity, yielding significant improvements in computational efficiency. In particular, we evaluate our proposed training scheme on image classification tasks from CIFAR-10 and CIFAR-100 datasets on several VGG architectures. We achieve top-l accuracy of 66.46% with 5 timesteps on the CIFAR-100 dataset with ~125x less compute energy than an equivalent standard ANN. Additionally, our proposed SNN performs 5–300 x faster inference compared to other state-of-the-art rate or temporally coded SNN models.
Gourav Datta, Souvik Kundu 0002, Peter A. Beerel
IJCNN1
2021 Spike-Thrift: Towards Energy-Efficient Deep Spiking Neural Networks by Limiting Spiking Activity via Attention-Guided Compression
abstract
The increasing demand for on-chip edge intelligence has motivated the exploration of algorithmic techniques and specialized hardware to reduce the computation energy of current machine learning models. In particular, deep spiking neural networks (SNNs) have gained interest because their event-driven hardware implementations can consume very low energy. However, minimizing average spiking activity and thus energy consumption while preserving accuracy in deep SNNs remains a significant challenge and opportunity. This paper proposes a novel two-step SNN compression technique to reduce their spiking activity while maintaining accuracy that involves compressing specifically-designed artificial neural networks (ANNs) that are then converted into the target SNNs. Our approach uses an ultra-high ANN compression technique that is guided by the attention-maps of an uncompressed meta-model. We then evaluate the firing threshold of each ANN layer and start with the trained ANN weights to perform a sparse-learning-based supervised SNN training to minimize the number of time steps required while retaining compression. To evaluate the merits of the proposed approach, we performed experiments with variants of VGG and ResNet, on both CIFAR-10 and CIFAR-100, and VGG16 on Tiny-ImageNet. SNN models generated through the proposed technique yield state-of-the-art compression ratios of up to 33.4× with no significant drop in accuracy compared to baseline unpruned counterparts. As opposed to the existing SNN pruning methods we achieve up to 8.3× better compression with no drop in accuracy. Moreover, compressed SNN models generated by our methods can have up to 12.2× better compute energy-efficiency compared to ANNs that have a similar number of parameters.
Souvik Kundu 0002, Gourav Datta, Massoud Pedram, Peter A. Beerel
WACV2
2021 Metastability in Superconducting Single Flux Quantum (SFQ) Logic
abstract
Superconducting digital electronics, especially Single Flux Quantum (SFQ), has emerged as a promising beyond-CMOS technology with Josephson junctions (JJ) as the active device. It has the potential to meet the booming demands of lower power consumption and higher operation speeds in the electronics industry and future exascale supercomputing systems. Despite these promises, scaling SFQ circuits remains a serious challenge that motivates the support of multiple SFQ clock domains. Towards this end, this paper analyzes the impact of setup time violations and metastability in SFQ circuits comparing the derived analytical models to their CMOS counterparts. It also proposes new techniques to reduce the average latency in metastability-tolerant SFQ synchronizers, and evaluates their effects on the layout and critical margin of the design. It further extends the proposed model to estimate the Mean Time Between Failure (MTBF) of flip-flop-based synchronizers and shows that their MTBF with the current feature sizes is unaffected by noise, similar to CMOS. Finally, it curve fits this model to simulations using the state-of-the-art SFQ5ee process and shows that a two-flop SFQ synchronizer with a clock frequency of 25 GHz has an estimated MTBF of ~106years.
Gourav Datta, Yunkun Lin, Bo Zhang 0098, Peter A. Beerel
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 Modeling and Characterization of Metastability in Single Flux Quantum (SFQ) Synchronizers
abstract
Despite the promises of low-power and high-frequency of single-flux quantum (SFQ) technology, scaling these circuits remains a serious challenge that motivates the support of multiple SFQ clock domains. Towards this end, this paper analyzes the impact of setup time violations and metastability in SFQ circuits comparing the derived analytical models to their CMOS counterparts. It then extends this model to estimate the Mean Time Between Failure (MTBF) of flip-flop-based synchronizers and curve fits this model to simulations in the state-of-the-art SFQ5ee process. Interestingly, we find a two-flop SFQ synchronizer has an estimated MTBF of ~106years.
Gourav Datta, Peter A. Beerel
ISCAS1