EDBT 2026 Demo / reviewers in the wild / expert
B. Sharat Chandra Varma 0001
dblp:127/1294 · also Bogaraju Sharatchandra Varma, Sharatchandra Varma Bogaraju
· DBLP profile ↗
7ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-8082-8822ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | F2Opt: Novel Fine-Tuning and Folding Algorithms for FPGA-Based DNN AcceleratorsabstractFPGAs, with their parallelism, low power consumption, and reconfigurability, offer an ideal solution for accelerating quantized deep neural networks (qDNNs) on resource-constrained edge devices. They enable enhanced latency, reduced energy consumption, and improved computational efficiency. However, existing frameworks to accelerate qDNN inference on FPGA face challenges in fine-tuning the deployed models on the FPGAs, as expensive re-synthesis and re-mapping of the accelerator is required. Additionally, the folding algorithm for DNN compute engines in these frameworks introduces substantial padding overheads to align with memory widths. This paper introduces a novel evolutionary algorithm-based hardware-in-the-loop (EvoHIL) framework. EvoHIL uses hardware-level weight bit-flip operations to activate neurons and improve accelerator accuracy without requiring DNN re-training and rebuilding the entire accelerator. We also propose a novel algorithmic optimization (Aopt) for folding across DNN accelerator layers. Aopt optimally aligns folding factors with memory width, eliminating excessive padding overheads and improving throughput while reducing memory and resource utilization. Experimental results demonstrate the effectiveness of these solutions. EvoHIL optimization enhanced the accuracy of a binarized convolutional neural network (BCNN) accelerator to nearly 86%. Aopt delivered significant improvements, including up to 96.77 % padding overhead reduction, 33.2 % increased throughput, and 25 % reduced runtime compared to FINN. Muhammad Shakeel Akram, B. Sharat Chandra Varma 0001, Vincent Meyers, Mehdi Baradaran Tahoori, Dewar Finlay |
FPL | 2 |
| 2025 | Toward TinyDPFL systems for real-time cardiac healthcare: Trends, challenges, and system-level perspectives on AI algorithms, hardware, and edge intelligenceabstractDespite rapid advances in medical technology, cardiac diseases remain the leading cause of global mortality, with arrhythmias that pose significant diagnostic and treatment challenges. This survey presents a comprehensive review of 176 state-of-the-art contributions in machine learning (ML), federated learning (FL), TinyML, and hardware acceleration for efficient, real-time, and privacy-preserving cardiac diagnosis and care. Explores both software and hardware advancements, including differential privacy (DP), quantized neural networks, and FPGA (Field Programmable Gate Array)-based implementations optimized for edge devices and wearable devices. Key challenges, such as latency, energy constraints, adversarial robustness, and personalization, are systematically examined. The survey synthesizes solutions across algorithmic innovations, secure and adaptive FL frameworks, and neuromorphic and sparse architectures, especially FPGA-based solutions, for resource-aware inference and training. Informed by original research, it highlights emerging directions: AI-driven data mining, DP for quantized models, continual learning (CL) on the edge, FPGA-accelerators including quantized DNN, SNN, and Sparse architectures, tuneable/reconfigurable FPGA-based TinyDPFL, Multimodal heterogeneous FL, real-time adversarial detection via model watermarking. This work offers a unified system-level perspective bridging ML algorithms and edge AI hardware, guiding the development of scalable, adaptive, and trustworthy cardiac healthcare systems. Beyond surveying existing literature, it proposes forward-looking design principles to advance intelligent, secure, and practical digital cardiology. Muhammad Shakeel Akram, B. Sharat Chandra Varma 0001, Aqib Javed, Jim Harkin, Dewar Finlay |
J. Syst. Archit. | 2 |
| 2024 | An Energy-Efficient Artefact Detection Accelerator on FPGAs for Hyper-Spectral Satellite ImageryabstractHyper-Spectral Imaging (HSI) is a crucial technique used to analyse remote sensing data acquired from Earth observation satellites. The rich spatial and spectral information obtained through HSI allows for better characterisation and exploration of the Earth's surface over traditional techniques like RGB and Multi-Spectral imaging on the downlinked image data at ground stations. In some cases, these images do not contain meaningful information due to the presence of clouds or other artefacts, limiting their usefulness. Transmission of such artefact HSI images leads to wasteful use of already scarce energy and time costs required for communication. While detecting such artefacts prior to transmitting the HSI image is desirable, the computational complexity of these algorithms and the limited power budget on satellites (especially CubeSats) are key constraints. This paper presents an unsupervised learning-based convolutional autoencoder (CAE) model for artefact identification of acquired HSI images at the satellite and a deployment architecture on AMD's Zynq Ultrascale FPGAs. The model is trained and tested on widely used HSI image datasets: Indian Pines, Salinas Valley, the University of Pavia and the Kennedy Space Center. For deployment, the model is quantised to 8-bit precision, fine-tuned using the Vitis-AI framework and integrated as a subordinate accelerator using AMD's Deep-Learning Processing Units (DPU) instance on the Zynq device. Our tests show that the model can process each spectral band in an HSI image in 4 ms, 2.6× better than INT8 inference on Nvidia's Jetson platform & 1.27× better than SOTA artefact detectors. Our model also achieves an fl-score of 92.8 % and FPR of 0 % across the dataset, while consuming 21.52 mJ per HSI image, 3.6 ×better than INT8 Jetson inference & 7.5 × better than SOTA artefact detectors, making it a viable architecture for deployment in CubeSats. Cornell Castelino, Shashwat Khandelwal, Shanker Shreejith, B. Sharat Chandra Varma 0001 |
DSD | 4 |
| 2022 | Low-Latency In Situ Image Analytics With FPGA-Based Quantized Convolutional Neural NetworkabstractReal-time in situ image analytics impose stringent latency requirements on intelligent neural network inference operations. While conventional software-based implementations on the graphic processing unit (GPU)-accelerated platforms are flexible and have achieved very high inference throughput, they are not suitable for latency-sensitive applications where real-time feedback is needed. Here, we demonstrate that high-performance reconfigurable computing platforms based on field-programmable gate array (FPGA) processing can successfully bridge the gap between low-level hardware processing and high-level intelligent image analytics algorithm deployment within a unified system. The proposed design performs inference operations on a stream of individual images as they are produced and has a deeply pipelined hardware design that allows all layers of a quantized convolutional neural network (QCNN) to compute concurrently with partial image inputs. Using the case of label-free classification of human peripheral blood mononuclear cell (PBMC) subtypes as a proof-of-concept illustration, our system achieves an ultralow classification latency of 34.2 [Formula: see text] with over 95% end-to-end accuracy by using a QCNN, while the cells are imaged at throughput exceeding 29 200 cells/s. Our QCNN design is modular and is readily adaptable to other QCNNs with different latency and resource requirements. Maolin Wang 0002, Kelvin C. M. Lee, Bob M. F. Chung, B. Sharat Chandra Varma 0001, Ho-Cheung Ng, Justin S. J. Wong, Ho Cheung Shum, Kevin K. Tsia, Hayden Kwok-Hay So |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | Real-time object detection and classification for high-speed asymmetric-detection time-stretch optical microscopy on FPGAabstractA real-time object detection and classification system using FPGA developed for high-speed asymmetric time-stretched optical microscopy (ATOM) framework is presented. Due to the massive amount of data generated by optical frontend, storing the raw data for offline post-processing is slow and impractical for the targeted single cell analysis applications. The proposed FPGA solution eliminates the need to transfer and persist the entire raw data by processing low-level signals and forming high-level images in real-time. Objects of interest are detected and segmented from the image stream and a classifier subsequently performs high-level analysis on the segmented images. When compared with existing software-based post-processing workflow, this FPGA-based approach will improve both the number of objects captured per experiment and the overall end-to-end object classification performance. The system also allows co-optimization between optical system, low-level signal processing and image analytic in a unified environment that enables new scientific discoveries previously unachievable. Maolin Wang 0002, Ho-Cheung Ng, Bob M. F. Chung, B. Sharat Chandra Varma 0001, Manish Kumar Jaiswal, Kevin K. Tsia, Ho Cheung Shum, Hayden Kwok-Hay So |
FPT | 4 |
| 2014 | High Level Design Approach to Accelerate De Novo Genome Assembly Using FPGAsabstractMany scientific applications take a very long time to execute on general purpose processors. Speedups can be obtained by using specialized hardware in conjunction with the processors. FPGA based accelerators are known to be effective for reducing the execution time of many scientific applications. Since FPGAs are configurable, they can be customized to implement a variety of processing elements as accelerators. The process of mapping algorithm to architecture is complex, as the design space is large. System simulation is usually employed to carry out the exploration, in spite of the fact that simulation models take significantly large amount of time to execute. High level design space exploration helps in taking the required decisions to arrive at an optimal design. In this paper we describe design space exploration carried out for accelerating de novo genome assembly using FPGAs. Three models at various levels of abstraction were used. We discuss how the simulation time of these models influence the choice of design parameters at different levels of abstraction. We illustrate this process by using the high level models to evaluate Hard Embedded Blocks (HEBs) in FPGAs for accelerating the de novo genome assembly application. B. Sharat Chandra Varma 0001, Kolin Paul, M. Balakrishnan |
DSD | 1 |
| 2013 | FAssem: FPGA Based Acceleration of De Novo Genome AssemblyabstractNext generation sequencing technologies produce large amounts of data at very low cost. They produce short reads of DNA fragments. These fragments have many overlaps, lots of repeats and may also include sequencing errors. The assembly process involves merging these sequences to form the original sequences. In recent years many software programs have been developed for this purpose. All of them take significant amount of time to execute. Velvet is a commonly used de novo assembly program. We propose a method to reduce the overall time for assembly by using pre-processing of the short read data on FPGAs and processing its output using Velvet. We show significant speed-ups with slight or no compromise on the quality of the assembled output. B. Sharat Chandra Varma 0001, Kolin Paul, M. Balakrishnan, Dominique Lavenier |
FCCM | 1 |