Sidharth Maheshwari

dblp:152/5202 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-9665-5698ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Dynamic Tsetlin Machine Accelerators for On-Chip Training Using FPGAs
abstract
The increased demand for data privacy and security in machine learning (ML) applications has put impetus on effective edge training on Internet-of-Things (IoT) nodes. Edge training aims to leverage speed, energy efficiency and adaptability within the resource constraints of the nodes. Deploying and training Deep Neural Networks (DNNs)-based models at the edge, although accurate, posit significant challenges from the back-propagation algorithm’s complexity, bit precision trade-offs, and heterogeneity of DNN layers. This paper presents a Dynamic Tsetlin Machine (DTM) training accelerator as an alternative to DNN implementations. DTM utilizes logic-based on-chip inference with finite-state automata-driven learning within the same Field Programmable Gate Array (FPGA) package. Underpinned on the Vanilla and Coalesced Tsetlin Machine algorithms, the dynamic aspect of the accelerator design allows for a run-time reconfiguration targeting different datasets, model architectures, and model sizes without resynthesis. This makes the DTM suitable for targeting multivariate sensor-based edge tasks. Compared to DNNs, DTM trains with fewer multiply-accumulates, devoid of derivative computation. It is a data-centric ML algorithm that learns by aligning Tsetlin automata with input data to form logical propositions enabling efficient Look-up-Table (LUT) mapping and frugal Block RAM usage in FPGA training implementations. The proposed accelerator offers 2.54x more Giga operations per second per Watt (GOP/s per W) and uses 6x less power than the next-best comparable design.
Gang Mao, Tousif Rahman, Sidharth Maheshwari, Bob Pattison, Rishad A. Shafik, Alexandre Yakovlev
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 MATADOR: Automated System-on-Chip Tsetlin Machine Design Generation for Edge Applications
abstract
System-on-Chip Field-Programmable Gate Arrays (SoC-FPGAs) offer significant throughput gains for machine learning (ML) edge inference applications via the design of co-processor accelerator systems. However, the design effort for training and translating ML models into SoC-FPGA solutions can be substantial and requires specialist knowledge aware trade-offs between model performance, power consumption, latency and resource utilization. Contrary to other ML algorithms, Tsetlin Machine (TM) performs classification by forming logic proposition between boolean actions from the Tsetlin Automata (the learning elements) and boolean input features. A trained TM model, usually, exhibits high sparsity and considerable overlapping of these logic propositions both within and among the classes. The model, thus, can be translated to RTL-level design using a miniscule number of AND and NOT gates. This paper presents MATADOR, an automated boolean-to-silicon tool with GUI interface capable of implementing optimized accelerator design of the TM model onto SoC-FPGA for inference at the edge. It offers automation of the full development pipeline: model training, system level design generation, design verification and deployment. It makes use of the logic sharing that ensues from propositional overlap and creates a compact design by effectively utilizing the TM model's sparsity. MATADOR accelerator designs are shown to be up to 13.4x faster, up to 7x more resource frugal and up to 2x more power efficient when compared to the state-of-the-art Quantized and Binary Deep Neural Network implementations.
Tousif Rahman, Gang Mao, Sidharth Maheshwari, Rishad A. Shafik, Alexandre Yakovlev
DATE3
2023 REDRESS: Generating Compressed Models for Edge Inference Using Tsetlin Machines
abstract
Inference at-the-edge using embedded machine learning models is associated with challenging trade-offs between resource metrics, such as energy and memory footprint, and the performance metrics, such as computation time and accuracy. In this work, we go beyond the conventional Neural Network based approaches to explore Tsetlin Machine (TM), an emerging machine learning algorithm, that uses learning automata to create propositional logic for classification. We use algorithm-hardware co-design to propose a novel methodology for training and inference of TM. The methodology, called REDRESS, comprises independent TM training and inference techniques to reduce the memory footprint of the resulting automata to target low and ultra-low power applications. The array of Tsetlin Automata (TA) holds learned information in the binary form as bits: {0,1}, called excludes and includes, respectively. REDRESS proposes a lossless TA compression method, called the include-encoding, that stores only the information associated with includes to achieve over 99% compression. This is enabled by a novel computationally minimal training procedure, called the Tsetlin Automata Re-profiling, to improve the accuracy and increase the sparsity of TA to reduce the number of includes, hence, the memory footprint. Finally, REDRESS includes an inherently bit-parallel inference algorithm that operates on the optimally trained TA in the compressed domain, that does not require decompression during runtime, to obtain high speedups when compared with the state-of-the-art Binary Neural Network (BNN) models. In this work, we demonstrate that using REDRESS approach, TM outperforms BNN models on all design metrics for five benchmark datasets viz. MNIST, CIFAR2, KWS6, Fashion-MNIST and Kuzushiji-MNIST. When implemented on an STM32F746G-DISCO microcontroller, REDRESS obtained speedups and energy savings ranging 5-5700× compared with different BNN models.
Sidharth Maheshwari, Tousif Rahman, Rishad A. Shafik, Alexandre Yakovlev, Ashur Rafiev, Lei Jiao 0001, Ole-Christoffer Granmo
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 An FPGA Based Energy-Efficient Read Mapper With Parallel Filtering and In-Situ Verification
abstract
In the assembly pipeline of Whole Genome Sequencing (WGS), read mapping is a widely used method to re-assemble the genome. It employs approximate string matching and dynamic programming-based algorithms on a large volume of data and associated structures, making it a computationally intensive process. Currently, the state-of-the-art data centers for genome sequencing incur substantial setup and energy costs for maintaining hardware, data storage and cooling systems. To enable low-cost genomics, we propose an energy-efficient architectural methodology for read mapping using a single system-on-chip (SoC) platform. The proposed methodology is based on the q-gram lemma and designed using a novel architecture for filtering and verification. The filtering algorithm is designed using a parallel sorted q-gram lemma based method for the first time, and it is complemented by an in-situ verification routine using parallel Myers bit-vector algorithm. We have implemented our design on the Zynq Ultrascale+ XCZU9EG MPSoC platform. It is then extensively validated using real genomic data to demonstrate up to 7.8× energy reduction and up to 13.3× less resource utilization when compared with the state-of-the-art software and hardware approaches.
Venkateshwarlu Y. Gudur, Sidharth Maheshwari, Amit Acharyya, Rishad A. Shafik
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 PLEDGER: Embedded Whole Genome Read Mapping using Algorithm-HW Co-design and Memory-aware Implementation
abstract
With over 6000 known genetic disorders, genomics is a key driver to transform the current generation of healthcare from reactive to personalized, predictive, preventive and participatory (P4) form. High throughput sequencing technologies produce large volumes of genomic data, making genome reassembly and analysis computationally expensive in terms of performance and energy. In this paper, we propose an algorithm-hardware co-design driven acceleration approach for enabling translational genomics. Core to our approach is a Pyopencl based tooL for gEnomic workloaDs tarGeting Embedded platforms (PLEDGER). PLEDGER is a scalable, portable and energy-efficient solution to genomics targeting low-cost embedded platforms. It is a read mapping tool to reassemble genome, which is a crucial prerequisite to genomics. Using bit-vectors and variable level optimisations, we propose a low-memory footprint, dynamic programming based filtration and verification kernel capable of accelerated parallel heterogeneous executions. We demonstrate, for the first time, mapping of real reads to whole human genome on a memory-restricted embedded platform using novel memory-aware preprocessed data structures. We compare the performance and accuracy of PLEDGER with state-of-the-art RazerS3, Hobbes3, CORAL and REPUTE on two systems: 1) Intel i7-8750H CPU + Nvidia GTX 1050 Ti, 2) Odroid N2 with 6 cores: 4xCortex-A73 + 2xCortex-A53 and Mali GPU. PLEDGER demonstrates persistent energy and accuracy advantages compared to state-of-the-art read mappers producing up to 11× speedups and 5.9× energy savings compared to state-of-the-art hardware resources.
Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Venkateshwarlu Y. Gudur, Amit Acharyya
DATE1
2021 CORAL: Verification-Aware OpenCL Based Read Mapper for Heterogeneous Systems
abstract
Genomics has the potential to transform medicine from reactive to a personalized, predictive, preventive, and participatory (P4) form. Being a Big Data application with continuously increasing rate of data production, the computational costs of genomics have become a daunting challenge. Most modern computing systems are heterogeneous consisting of various combinations of computing resources, such as CPUs, GPUs, and FPGAs. They require platform-specific software and languages to program making their simultaneous operation challenging. Existing read mappers and analysis tools in the whole genome sequencing (WGS) pipeline do not scale for such heterogeneity. Additionally, the computational cost of mapping reads is high due to expensive dynamic programming based verification, where optimized implementations are already available. Thus, improvement in filtration techniques is needed to reduce verification overhead. To address the aforementioned limitations with regards to the mapping element of the WGS pipeline, we propose a Cross-platfOrm Read mApper using opencL (CORAL). CORAL is capable of executing on heterogeneous devices/platforms, simultaneously. It can reduce computational time by suitably distributing the workload without any additional programming effort. We showcase this on a quadcore Intel CPU along with two Nvidia GTX 590 GPUs, distributing the workload judiciously to achieve up to 2× speedup compared to when, only, the CPUs are used. To reduce the verification overhead, CORAL dynamically adapts k-mer length during filtration. We demonstrate competitive timings in comparison with other mappers using real and simulated reads. CORAL is available at: https://github.com/nclaes/CORAL.
Sidharth Maheshwari, Venkateshwarlu Y. Gudur, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 REPUTE: An OpenCL based Read Mapping Tool for Embedded Genomics
abstract
Genomics is transforming medicine from reactive to personalized, predictive, preventive and participatory (P4). The massive amount of data produced by genomics is a major challenge as it requires extensive computational capabilities, consuming large amounts of energy. A crucial prerequisite for computational genomics is genome assembly but the existing mapping tools used are predominantly software based, optimized for homogeneous high-performance systems. In this paper, we propose an OpenCL based REad maPper for heterogeneoUs sysTEms (REPUTE), which can use diverse and parallel compute and storage devices effectively. Core to this tool are dynamic programming based filtration and verification kernel to map the reads on multiple devices, concurrently. We show hardware/ software co-design and implementations of REPUTE across different platforms, and compare it with state-of-the-art mappers. We demonstrate the performance of mappers on two systems: 1) Intel CPU + 2×Nvidia GPUs; 2) HiKey970 embedded SoC with ARM Cortex-A73/A53 cores. The results show that REPUTE outperforms other read mappers in most cases producing up to 13× speedup with better or comparable accuracy. We also demonstrate that the embedded implementation can achieve up to 27× energy savings, enabling low-cost genomics.
Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya
DATE1
2020 Accelerated Filtering and in situ Verification for Energy-Optimized Genome Read Mapping
abstract
Whole genome sequencing (WGS) includes sequencing and assembly pipelines to extract biological genomes for new advances in healthcare, agriculture and environmental research. It produces small random sections of the genome, called reads, and then re-assembled by mapping those reads to a reference genome. This process called read mapping produces a large volume of data, which are disparately processed by compute- and memory-intensive filtering and verification algorithms. As such, the problem of energy-frugal read mapping has remained an open challenge. In this paper, we propose an accelerated read mapping methodology with combined filtering and verification, implemented on an FPGA platform. Core to our methodology is an algorithm based on q-gram lemma for filtration with Myers bit-vector for verification in tandem. Through in situ verification, the proposed implementation optimizes resource utilization between filtration and verification and introduces parallel pipelines in computation and storage processes. Our experimental analysis shows that this methodology gives up to 8.7× energy efficiency when implemented on the Zynq Ultrascale+ FPGA platform, compared with the state-of-the-art software and hardware approaches.
Venkateshwarlu Y. Gudur, Sidharth Maheshwari, Rishad A. Shafik, Amit Acharyya
ISCAS2
2016 Vector FPGA acceleration of 1-D DWT computations using sparse matrix skeletons
abstract
We can exploit application-specific sparse structure and distribution of non-zero coefficients in Discrete Wavelet Transform (DWT) matrices to significantly improve the performance of 1-D DWT mapped to FPGA-based soft vector processors. We reformulate DWT computations specifically in terms of sparse matrix operations, where the transformation matrices have a repeating block with a fixed non-zero pattern, which we refer to as a skeleton. We exploit this property to transform the original DWT matrix into a Modified-Matrix-Form to expose abundant soft vector parallelism in the dot products. The resulting form can also be readily compiled into low-level DMA routines for boosting memory throughput. We autogenerate vector routines and memory access sequences tailored for parametric combinations of DWT filter sizes, and decomposition levels as required by the application domain. When compared to embedded ARMv7 32b CPU implementations using optimized OpenBLAS routines, soft vector implementation on the Xilinx Zedboard and Altera DE2/DE4 platforms demonstrate speedups of 12-103×.
Sidharth Maheshwari, Gourav Modi, Siddhartha 0001, Nachiket Kapre
FPL1
2014 Multi-directional error correction schemes for SRAM-based FPGAs
abstract
Readback scrubbing is considered as an effective mechanism to correct errors in Static-RAM (SRAM)-based Field Programmable Gate Arrays (FPGAs). However, current solutions have a low error correction percentage per unit area overhead. This paper proposes two new error detection/correction mechanisms that combine frame readback scrubbing with error correction codes (ECCs) that are applied in multiple directions, to achieve a high error correction percentage per unit area overhead. Experiments conducted show that the proposed schemes have an excellent error correction percentage (over 99%), especially for multi-bit upsets, while using up to 59.37% lesser area overhead compared with other state-of-the-art.
Shyamsundar Venkataraman, Sidharth Maheshwari, Akash Kumar 0001
FPL3
2013 Accurate and reliable 3-lead to 12-lead ECG reconstruction methodology for remote health monitoring applications
abstract
Standard 12-lead (S12) system and Mason-Likar 12-lead (ML12) system despite of being most acceptable systems for clinical usage are not the preferred lead systems for remote monitoring (RM) applications. Usually RM applications involve wireless transmission of signals and a 2-3 lead system is preferred for bandwidth and storage limitations and data transmission time. Generally, ECG compression techniques are applied for the same, however, compression ratio (CR) depends on the number of channels and decreases with the increase in number of channels. Thus, it facilitates the usage of a 2-3 lead system. However, a reduced lead (RL) system with 2-3 leads may be inadequate for the information desired by the cardiologists who are accustomed to S12 or ML12 system pertaining to its decades old usage. In this paper, we attempt to provide solution to both technical and non-technical limitations of RM applications. We reconstruct S12 and ML12 systems from Reduced 3-lead (R3L) system comprising of basis leads I, II, V2using personalized or patient-specific transformation. Two separate investigations have been carried out for S12 and ML12 with their corresponding R3L systems comprising of their respective basis leads. PhysioNet PTBDB and INCARTDB after wavelet based preprocessing were used in this investigation. R2statistics, correlation (rx) and regression (bx) coefficients were used to evaluate reconstructed signal against the original signal and the mean values obtained were 96.53%, 0.982 and 0.968 (S12) and 96.53%, 0.982 and 0.968 (ML12) respectively. R3L system reduces number of leads and electrodes from 12 and 10 to 3 and 5 respectively, lowers bandwidth and storage requirements, data transmission time and increases CR. The study shows that basis leads obtained from S12 outperforms the basis leads of ML12 for reconstruction of precordial leads.
Sidharth Maheshwari, Amit Acharyya, Pachamuthu Rajalakshmi, Paolo Emilio Puddu, Michele Schiariti
Healthcom1