EDBT 2026 Demo / reviewers in the wild / expert
Priyesh Shukla
dblp:260/0640
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-9868-8944ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MOSAIC: Collaborative Compute-in-Memory µArrays for Flexible and Scalable Deep LearningabstractCompute-in-Memory (CiM) architectures, particularly those leveraging SRAM-based arrays, present significant opportunities for accelerating deep learning by mitigating data movement bottlenecks in traditional von Neumann systems. SRAM-based CiM offers notable advantages, including speed, seamless CMOS integration, and compatibility with existing System-on-Chip (SoC) designs. However, challenges persist, primarily stemming from analog-domain computations that necessitate analog-to-digital (A/D) and digital-to-analog (D/A) conversions, leading to reduced accuracy, increased power overhead, and rigid operational constraints. To overcome these limitations, we propose MOSAIC, a novel CiM architecture designed around three foundational principles: (1) co-designing deep learning operators for CiM such as employing multiplication-free computations and frequency-domain processing to minimize/eliminate DAC/ADC overheads; (2) leveraging a memory-immersed digitization approach that utilizes parasitic bit-lines as capacitive DACs within CiM arrays, thereby significantly reducing peripheral complexity and enhancing scalability; and (3) orchestrating inference over a network of compact CiM µArrays by dynamically interconnecting them to provide flexibility, efficiency, and minimized computational overhead for varied inference workload characteristics. Collectively with these fundamental innovations, MOSAIC addresses critical bottlenecks in accuracy, scalability, and flexibility, unlocking CiM’s full potential for efficient deep learning in embedded systems. Amit Ranjan Trivedi, Shamma Nasrin, Priyesh Shukla, Nastaran Darabi, Divake Kumar, Dinithi Jayasuriya, Nethmi Jayasinghe |
VTS | 3 |
| 2024 | Navigating the Unknown: Uncertainty-Aware Compute-in-Memory Autonomy of Edge RoboticsabstractThis paper addresses the challenging problem of energy-efficient and uncertainty-aware pose estimation in insect-scale drones, which is crucial for tasks such as surveillance in constricted spaces and for enabling non-intrusive spatial intelligence in smart homes. Since tiny drones operate in highly dynamic environments, where factors like lighting and human movement impact their predictive accuracy, it is crucial to deploy uncertainty-aware prediction algorithms that can account for environmental variations and express not only the prediction but also confidence in the prediction. We address both of these challenges with Compute-in-Memory (CIM) which has become a pivotal technology for deep learning acceleration at the edge. While traditional CIM techniques are promising for energy-efficient deep learning, to bring in the robustness of uncertainty-aware predictions at the edge, we introduce a suite of novel techniques: First, we discuss CIM-based acceleration of Bayesian filtering methods uniquely by leveraging the Gaussian-like switching current of CMOS inverters along with co-design of kernel functions to operate with extreme parallelism and with extreme energy efficiency. Secondly, we discuss the CIM-based acceleration of variational inference of deep learning models through probabilistic processing while unfolding iterative computations of the method with a compute reuse strategy to significantly minimize the workload. Overall, our co-design methodologies demonstrate the potential of CIM to improve the processing efficiency of uncertainty-aware algorithms by orders of magnitude, thereby enabling edge robotics to access the robustness of sophisticated prediction frameworks within their extremely stringent area/power resources. Nastaran Darabi, Priyesh Shukla, Dinithi Jayasuriya, Divake Kumar, Alex C. Stutts, Amit Ranjan Trivedi |
DATE | 2 |
| 2023 | Robust Monocular Localization of Drones by Adapting Domain Maps to Depth Prediction InaccuraciesabstractWe present a novel monocular localization framework by jointly training deep learning-based depth prediction and Bayesian filtering-based pose reasoning. The proposed cross-modal framework significantly outperforms deep learning-only predictions with respect to model scalability and tolerance to environmental variations. Specifically, we show little-to-no degradation of pose accuracy even with extremely poor depth estimates from a lightweight depth predictor. Our framework also maintains high pose accuracy in extreme lighting variations compared to standard deep learning, even without explicit domain adaptation. By openly representing the map and intermediate feature maps (such as depth estimates), our framework also allows for faster updates and reusing intermediate predictions for other tasks, such as obstacle avoidance, resulting in much higher resource efficiency. Priyesh Shukla, Sureshkumar S., Alex C. Stutts, Sathya Ravi, Theja Tulabandhula, Amit Ranjan Trivedi |
ICASSP | 1 |
| 2023 | MC-CIM: Compute-in-Memory With Monte-Carlo Dropouts for Bayesian Edge IntelligenceabstractWe propose MC-CIM, a compute-in-memory (CIM) framework for robust, yet low power, Bayesian edge intelligence. Deep neural networks (DNN) with deterministic weights cannot express their prediction uncertainties, thereby pose critical risks for applications where the consequences of mispredictions are fatal such as surgical robotics. To address this limitation, Bayesian inference of a DNN has gained attention. Using Bayesian inference, not only the prediction itself, but the prediction confidence can also be extracted for planning risk-aware actions. However, Bayesian inference of a DNN is computationally expensive, ill-suited for real-time and/or edge deployment. An approximation to Bayesian DNN using Monte Carlo Dropout (MC-Dropout) has shown high robustness along with low computational complexity. Enhancing the computational efficiency of the method, we discuss a novel CIM module that can perform in-memory probabilistic dropout in addition to in-memory weight-input scalar product to support the method. We also propose a compute-reuse reformulation of MC-Dropout where each successive instance can utilize the product-sum computations from the previous iteration. Even more, we discuss how the random instances can be optimally ordered to minimize the overall MC-Dropout workload by exploiting combinatorial optimization methods. Application of the proposed CIM-based MC-Dropout execution is discussed for MNIST character recognition and visual odometry (VO) of autonomous drones. The framework reliably gives prediction confidence amidst non-idealities imposed by MC-CIM to a good extent. Proposed MC-CIM with$16\times 31$SRAM array, 0.85 V supply, 16nm low-standby power (LSTP) technology consumes 32 pJ for 30 MC-Dropout instances of probabilistic inference in its most optimal computing and peripheral configuration, saving$\sim 34$% energy compared to typical execution. Priyesh Shukla, Shamma Nasrin, Nastaran Darabi, Wilfred Gomes, Amit Ranjan Trivedi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Non van-Neumann Anomaly Detection in Multi-Channel Time-Series using Charge Trap Transistor CrossbarsabstractMany of the existing techniques to detect anomalies in multi-channel time-series lack flexibility and incur significant processing-overheads; therefore, real-time, flexible anomaly detection in resource-constrained edge devices is still an open problem. Addressing the above challenges, we present an ultra-low-power non von-Neumann framework for statistical modeling based anomaly detection in multi-channel time-series. To disruptively minimize the power dissipation for anomaly detection and enhance scalability to complex time-series statistics, we pursue a co-designed approach where the anomaly statistics are modelled using an unconventional Harmonic-Mean of Gaussian-like (HMG) functions. We show that multivariate HMG functions can be implemented simply by exploiting the short-circuit current of multi-input inverters. A non-von Neumann crossbar of charge trap inverters implements a mixture of HMG function and stores model weights. On Yahoo real time-series dataset, our anomaly detection approach achieves an f1-score higher than 0.85 even in the presence of significant process variation and consumes 181fJ/sensor sample for a three-channel time-series. Compared to baseline digital and analog approaches for anomaly detection, our framework is $\sim 40 \times$ and $\sim 6 \times$ more energy efficient, respectively, while being more scalable to time-series dimension. Ahish Shylendra, Priyesh Shukla, Amit Ranjan Trivedi |
ISCAS | 2 |
| 2022 | Ultralow-Power Localization of Insect-Scale Drones: Interplay of Probabilistic Filtering and Compute-in-MemoryabstractWe propose a novel compute-in-memory (CIM)-based ultralow-power framework for probabilistic localization of insect-scale drones. Localization is a critical subroutine for path planning and rotor control in drones, where a drone is required to continuously estimate its pose (position and orientation) in flying space. The conventional probabilistic localization approaches rely on the 3-D Gaussian mixture model (GMM)-based representation of a 3-D map. A GMM model with hundreds of mixture functions is typically needed to adequately learn and represent the intricacies of the map. Meanwhile, localization using complex GMM map models is computationally intensive. Since insect-scale drones operate under extremely limited area/power budget, continuous localization using GMM models entails much higher operating energy, thereby limiting flying duration and/or size of the drone due to a larger battery. Addressing the computational challenges of localization in an insect-scale drone using a CIM approach, we propose a novel framework of 3-D map representation using a harmonic mean of the “Gaussian-like” mixture (HMGM) model. We show thatshort-circuit currentof a multiinput floating-gate CMOS-based inverter follows the harmonic mean of a Gaussian-like function. Therefore, the likelihood function useful for drone localization can be efficiently implemented by connecting many multiinput inverters in parallel, each programmed with the parameters of the 3-D map model represented as HMGM. When the depth measurements are projected to the input of the implementation, the summed current of the inverters emulates the likelihood of the measurement. We have characterized our approach on an RGB-D scenes dataset. The proposed localization framework is$\sim 25\times $energy-efficient than the traditional, 8-bit digital GMM-based processor paving the way for tiny autonomous drones. Priyesh Shukla, Ankith Muralidhar, Nick Iliev, Theja Tulabandhula, Sawyer B. Fuller, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Compute-in-Memory Upside Down: A Learning Operator Co-Design Perspective for ScalabilityabstractThis paper discusses the potential of model-hardware co-design to simplify the implementation complexity of compute-in-SRAM deep learning considerably. Although compute-in-SRAM has emerged as a promising approach to improve the energy efficiency of DNN processing, current implementations suffer due to complex and excessive mixed-signal peripherals, such as the need for parallel digital-to-analog converters (DACs) at each input port. Comparatively, our approach inherently obviates complex peripherals by co-designing learning operators to SRAM's operational constraints. For example, our co-designed implementation is DAC-free even for multibit precision DNN processing. Additionally, we also discuss the interaction of our compute-in-SRAM operator with Bayesian inference of DNNs. We show a synergistic interaction of Bayesian inference with our framework, where Bayesian methods allow achieving similar accuracy with much smaller network size. Although each iteration of sample-based Bayesian inference is computationally expensive, the cost is minimized by our compute-in-SRAM approach. Meanwhile, by reducing the network size, Bayesian methods reduce the footprint cost of compute-in-SRAM implementation, which is a crucial concern for the method. We characterize this interaction for deep learning-based pose (position and orientation) estimation for a drone. Shamma Nasrin, Priyesh Shukla, Shruthi Jaisimha, Amit Ranjan Trivedi |
DATE | 2 |
| 2020 | MC2RAM: Markov Chain Monte Carlo Sampling in SRAM for Fast Bayesian InferenceabstractThis work discusses the implementation of Marko Chain Monte Carlo (MCMC) sampling from an arbitrary Gaussian mixture model (GMM) within SRAM. We show a novel architecture of SRAM by embedding it with random number generators (RNGs), digital-to-analog converters (DACs), and analog-to-digital converters (ADCs) so that SRAM arrays can be used for high performance Metropolis-Hastings (MH) algorithm-based MCMC sampling. Most of the expensive computations are performed within the SRAM and can be parallelized for high speed sampling. Our iterative compute flow minimizes data movement during sampling. We characterize power-performance trade-off of our design by simulating on 45 nm CMOS technology. For a two-dimensional, two mixture GMM, the implementation consumes ~ 91μW power per sampling iteration and produces 500 samples in 2000 clock cycles on an average at 1 GHz clock frequency. Our study highlights interesting insights on how low-level hardware non-idealities can affect high-level sampling characteristics, and recommends ways to optimally operate SRAM within area/power constraints for high performance sampling. Priyesh Shukla, Ahish Shylendra, Theja Tulabandhula, Amit Ranjan Trivedi |
ISCAS | 1 |
| 2020 | Low Power Unsupervised Anomaly Detection by Nonparametric Modeling of Sensor StatisticsabstractThis article presents anomaly detection by examining sensor stream statistics (AEGIS), a novel mixed-signal framework for real-time AEGIS. AEGIS utilizes kernel density estimation (KDE)-based nonparametric density estimation to generate a real-time statistical model of the sensor data stream. The likelihood estimate of the sensor data point can be obtained based on the generated statistical model to detect outliers. We present CMOS Gilbert Gaussian cell-based design to realize Gaussian kernels for KDE. For outlier detection, the decision boundary is defined in terms of kernel standard deviation (σKernel) and likelihood threshold (PThres). We adopt a sliding window to update the detection model in real time. We use time-series data set provided from Yahoo to benchmark the performance of AEGIS. A f1-score higher than 0.87 is achieved by optimizing parameters such as length of the sliding window and decision thresholds which are programmable in AEGIS. Discussed architecture is designed using 45-nm technology node and our approach on average consumes ~75-μW power at a sampling rate of 2 MHz while using ten recent inlier samples for density estimation. Ahish Shylendra, Priyesh Shukla, Saibal Mukhopadhyay, Swarup Bhunia, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |