EDBT 2026 Demo / reviewers in the wild / expert
Gurpreet S. Kalsi
dblp:241/0915 · also Gurpreet Singh Kalsi
· DBLP profile ↗
10ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0000-7018-7585ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ace-of-Spades: Accelerating Spatially Sparse Convolution for 3D Scene UnderstandingabstractSemantic understanding of 3D scenes is fundamental to many applications like robotics, autonomous driving, AR/VR. State-of-the-art methods for different 3D scene understanding tasks use 3D convolutional neural networks (CNNs) operating on point clouds. Convolution on spatially sparse data like point cloud involves irregular data accesses and compute patterns leading to poor utilization and energy efficiency in CPU/GPU implementations. The existing CNN accelerators designed for weight/activation sparsity cannot be efficiently repurposed for 3D spatially sparse CNNs given the fundamental differences in locating non-zero operands and granularity of work-dispatches. To address the dataflow challenges due to spatial sparsity and the need for specialized microarchitecture for spatially sparse convolution we present Ace-of-Spade s (AoS), an algorithm-dataflow-architecture co-designed system. AoS enables the data reuse among spatially proximate points using a locality-aware metadata structure along with a surface orientation aware point cloud reordering algorithm. AoS uses a novel technique for spatial sparsity aware selection of optimal data tiles by modelling the sparsity induced variations in the point cloud with a near-zero latency overheads. To accelerate computation on spatially sparse data, we propose a novel hardware accelerator Ss p nna with a front-end to convert varying number of operations per point into a stream of dense work dispatches to the backend compute engine. The compute engine further exploits weight and input feature data reuse through dynamic systolic grouping and multicast interconnects. The Ss p nna core together with the 64 KB of L1 memory requires 0.31 mm 2 of area in 10nm process at 1 GHz. Overall, AoS achieves speedup/energy savings of 19.9x / 49.9x and 2.2x / 7.1x over the state-of-the-art CPU and GPU implementations respectively. Om Ji Omer, Prashant Laddha, Gurpreet S. Kalsi, Kamlesh R. Pillai, Anirud Thyagharajan, Ahimanyu Kulkarni, Anbang Yao, Yurong Chen 0001, Sreenivas Subramoney |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | ApHMM: Accelerating Profile Hidden Markov Models for Fast and Energy-efficient Genome AnalysisabstractProfile hidden Markov models (pHMMs) are widely employed in various bioinformatics applications to identify similarities between biological sequences, such as DNA or protein sequences. In pHMMs, sequences are represented as graph structures, where states and edges capture modifications (i.e., insertions, deletions, and substitutions) by assigning probabilities to them. These probabilities are subsequently used to compute the similarity score between a sequence and a pHMM graph. The Baum-Welch algorithm, a prevalent and highly accurate method, utilizes these probabilities to optimize and compute similarity scores. Accurate computation of these probabilities is essential for the correct identification of sequence similarities. However, the Baum-Welch algorithm is computationally intensive, and existing solutions offer either software-only or hardware-only approaches with fixed pHMM designs. When we analyze state-of-the-art works, we identify an urgent need for a flexible, high-performance, and energy-efficient hardware-software co-design to address the major inefficiencies in the Baum-Welch algorithm for pHMMs. We introduce ApHMM , the first flexible acceleration framework designed to significantly reduce both computational and energy overheads associated with the Baum-Welch algorithm for pHMMs. ApHMM employs hardware-software co-design to tackle the major inefficiencies in the Baum-Welch algorithm by (1) designing flexible hardware to accommodate various pHMM designs, (2) exploiting predictable data dependency patterns through on-chip memory with memoization techniques, (3) rapidly filtering out unnecessary computations using a hardware-based filter, and (4) minimizing redundant computations. ApHMM achieves substantial speedups of 15.55×–260.03×, 1.83×–5.34×, and 27.97× when compared to CPU, GPU, and FPGA implementations of the Baum-Welch algorithm, respectively. ApHMM outperforms state-of-the-art CPU implementations in three key bioinformatics applications: (1) error correction, (2) protein family search, and (3) multiple sequence alignment, by 1.29×–59.94×, 1.03×–1.75×, and 1.03×–1.95×, respectively, while improving their energy efficiency by 64.24×–115.46×, 1.75×, and 1.96×. Can Firtina, Kamlesh R. Pillai, Gurpreet S. Kalsi, Bharathwaj Suresh, Damla Senol Cali, Jeremie S. Kim, Taha Shahroodi, Meryem Banu Cavlak, Joël Lindegger, Mohammed Alser, Juan Gómez-Luna, Sreenivas Subramoney, Onur Mutlu |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | SeGraM: a universal hardware accelerator for genomic sequence-to-graph and sequence-to-sequence mappingabstractA critical step of genome sequence analysis is the mapping of sequenced DNA fragments (i.e., reads) collected from an individual to a known linear reference genome sequence (i.e., sequence-to-sequence mapping). Recent works replace the linear reference sequence with a graph-based representation of the reference genome, which captures the genetic variations and diversity across many individuals in a population. Mapping reads to the graph-based reference genome (i.e., sequence-to-graph mapping) results in notable quality improvements in genome analysis. Unfortunately, while sequence-to-sequence mapping is well studied with many available tools and accelerators, sequence-to-graph mapping is a more difficult computational problem, with a much smaller number of practical software tools currently available. Damla Senol Cali, Konstantinos Kanellopoulos, Joël Lindegger, Zülal Bingöl, Gurpreet S. Kalsi, Ziyi Zuo, Can Firtina, Meryem Banu Cavlak, Jeremie S. Kim, Nika Mansouri-Ghiasi, Gagandeep Singh 0002, Juan Gómez-Luna, Nour Almadhoun, Mohammed Alser, Sreenivas Subramoney, Can Alkan, Saugata Ghose, Onur Mutlu |
ISCA | 5 |
| 2022 | Compute-In-Memory Using 6T SRAM for a Wide Variety of WorkloadsabstractThis paper presents a split wordline 6T SRAM based compute-in-memory subarray with variable multi-bit precision for input operands and outputs. The split wordlines of the 6T cell enable sign segregation, thus allowing arbitrary sign/magnitude multiply and accumulate (MAC) operations. The arbitrary signmagnitude MAC operation extends the usage of the MAC array for DSP workloads as well. On top of that, the split wordline 6T cell is more resilient to write-disturb, which is a major concern for conventional 6T cell based compute-in-memory operation. The proposed system was designed and implemented using 65nm UMC technology and the energy efficiency and throughput are found to be 80.1 TOPS/W and 496 GOPS respectively, for maximum precision of inputs(4 bits)/outputs(5 bits). The maximum throughput and energy efficiency of the design are found to be 780 GOPS and 94.8 TOPS/W respectively. Hand-written digit recognition application mapped on the proposed system showed a maximum accuracy degradation of 0.2% as compared to that obtained from software with the same input/output quantization. Pramod Kumar Bharti, Kamlesh R. Pillai, Sagar Varma Sayyaparaju, Gurpreet S. Kalsi, Joycee Mekie, Sreenivas Subramoney |
ISCAS | 5 |
| 2021 | ONT-X: An FPGA Approach to Real-time Portable Genomic AnalysisabstractOxford Nanopore Technologies (ONT) MinION is a pocket-sized portable DNA sequencer for on-the-field DNA sequencing. The ONT data analysis pipeline requires substantial amounts of computational resources for the basecalling and read alignment step in the pipeline. We study the compute and memory requirement of the basecalling step and present the detailed microarchitecture for accelerating it on FPGA. We also make the case that an embedded FPGA-based SoC with partial reconfiguration (PR) capability is an ideal substrate for accelerating ONT data analysis pipeline, due to the diversity in these application kernels and the evolving nature of the algorithms. Focusing on two widely used ONT data analysis pipelines, we present two modes of PR operation to cater to either latency or throughput requirements of these two pipelines and demonstrate how to amortize the partial reconfiguration cost, overlap the kernels, minimize the data transfer cost and achieve 111× higher throughput and 6× reduction in latency using these two PR modes. C. N. Ramachandra, Anirban Nag, Rajeev Balasubramonian, Gurpreet S. Kalsi, Kamlesh R. Pillai, Sreenivas Subramoney |
FCCM | 4 |
| 2020 | PSB-RNN: A Processing-in-Memory Systolic Array Architecture using Block Circulant Matrices for Recurrent Neural NetworksabstractRecurrent Neural Networks (RNNs) are widely used in Natural Language Processing (NLP) applications as they inherently capture contextual information across spatial and temporal dimensions. Compared to other classes of neural networks, RNNs have more weight parameters as they primarily consist of fully connected layers. Recently, several techniques such as weight pruning, zero-skipping, and block circulant compression have been introduced to reduce the storage and access requirements of RNN weight parameters. In this work, we present a ReRAM crossbar based processing-in-memory (PIM) architecture with systolic dataflow incorporating block circulant compression for RNNs. The block circulant compression decomposes the operations in a fully connected layer into a series of Fourier transforms and point-wise operations resulting in reduced space and computational complexity. We formulate the Fourier transform and point-wise operations into in-situ multiply-and-accumulate (MAC) operations mapped to ReRAM crossbars for high energy efficiency and throughput. We also incorporate systolic dataflow for communication within the crossbar arrays, in contrast to broadcast and multicast communications, to further improve energy efficiency. The proposed architecture achieves average improvements in compute efficiency of 44× and 17× over a custom FPGA architecture and conventional crossbar based architecture implementations, respectively. Nagadastagiri Challapalle, Sahithi Rampalli, Makesh Chandran, Gurpreet S. Kalsi, Sreenivas Subramoney, Jack Sampson, Narayanan Vijaykrishnan |
DATE | 4 |
| 2020 | Descriptor Scoring for Feature Selection in Real-Time Visual SlamabstractMany emerging applications of Visual SLAM running on resource constrained hardware platforms impose very aggressive pose accuracy requirements and highly demanding latency constraints. To achieve the required pose accuracy under constrained compute budget, real-time SLAM implementations have to work with few but highly repeatable and invariant features. While many state-of-the-art techniques, proposed for selecting good features to track, do address some of these concerns, they are computationally complex and therefore, not suitable for power, latency and cost sensitive edge devices. On the other hand, simpler feature selection methods based on detector (corner) score, lack in identifying features with required invariance and trackability. We present a notion of feature descriptor score as a measure of invariance under distortions. We further propose feature selection method based on descriptor score requiring very minimal compute and demonstrate its performance with binary descriptors on an EKF based visual inertial odometry (VIO). Compared to detector score based methods, our method provides an improvement up to 10% in ATE (Absolute Trajectory Error) score on EuroC dataset. Prashant Laddha, Om Ji Omer, Gurpreet S. Kalsi, Dipan Mandal, Sreenivas Subramoney |
ICIP | 3 |
| 2020 | GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence AnalysisabstractGenome sequence analysis has enabled significant advancements in medical and scientific areas such as personalized medicine, outbreak tracing, and the understanding of evolution. To perform genome sequencing, devices extract small random fragments of an organism's DNA sequence (known as reads). The first step of genome sequence analysis is a computational process known as read mapping. In read mapping, each fragment is matched to its potential location in the reference genome with the goal of identifying the original location of each read in the genome. Unfortunately, rapid genome sequencing is currently bottlenecked by the computational power and memory bandwidth limitations of existing systems, as many of the steps in genome sequence analysis must process a large amount of data. A major contributor to this bottleneck is approximate string matching (ASM), which is used at multiple points during the mapping process. ASM enables read mapping to account for sequencing errors and genetic variations in the reads. We propose GenASM, the first ASM acceleration framework for genome sequence analysis. GenASM performs bitvectorbased ASM, which can efficiently accelerate multiple steps of genome sequence analysis. We modify the underlying ASM algorithm (Bitap) to significantly increase its parallelism and reduce its memory footprint. Using this modified algorithm, we design the first hardware accelerator for Bitap. Our hardware accelerator consists of specialized systolic-array-based compute units and on-chip SRAMs that are designed to match the rate of computation with memory capacity and bandwidth, resulting in an efficient design whose performance scales linearly as we increase the number of compute units working in parallel. We demonstrate that GenASM provides significant performance and power benefits for three different use cases in genome sequence analysis. First, GenASM accelerates read alignment for both long reads and short reads. For long reads, GenASM outperforms state-of-the-art software and hardware accelerators by 116× and 3.9×, respectively, while reducing power consumption by 37× and 2.7×. For short reads, GenASM outperforms state-of-the-art software and hardware accelerators by 111× and 1.9×. Second, GenASM accelerates pre-alignment filtering for short reads, with 3.7× the performance of a state-of-the-art pre-alignment filter, while reducing power consumption by 1.7× and significantly improving the filtering accuracy. Third, GenASM accelerates edit distance calculation, with 22-12501× and 9.3-400× speedups over the state-of-the-art software library and FPGA-based accelerator, respectively, while reducing power consumption by 548-582× and 67×. We conclude that GenASM is a flexible, high-performance, and low-power framework, and we briefly discuss four other use cases that can benefit from GenASM. Damla Senol Cali, Gurpreet S. Kalsi, Zülal Bingöl, Can Firtina, Lavanya Subramanian, Jeremie S. Kim, Rachata Ausavarungnirun, Mohammed Alser, Juan Gómez-Luna, Amirali Boroumand, Anant Nori, Allison Scibisz, Sreenivas Subramoney, Can Alkan, Saugata Ghose, Onur Mutlu |
MICRO | 2 |
| 2020 | Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network AccelerationabstractThis paper presents a Look-Up Table (LUT) based Processing-In-Memory (PIM) technique with the potential for running Neural Network inference tasks. We implement a bitline computing free technique to avoid frequent bitline accesses to the cache sub-arrays and thereby considerably reducing the memory access energy overhead. LUT in conjunction with the compute engines enables sub-array level parallelism while executing complex operations through data lookup which otherwise requires multiple cycles. Sub-array level parallelism and systolic input data flow ensure data movement to be confined to the SRAM slice.Our proposed LUT based PIM methodology exploits substantial parallelism using look-up tables, which does not alter the memory structure/organization, that is, preserving the bit-cell and peripherals of the existing SRAM monolithic arrays. Our solution achieves 1.72× higher performance and 3.14x lower energy as compared to a state-of-the-art processing-in-cache solution. Sub-array level design modifications to incorporate LUT along with the compute engines will increase the overall cache area by 5.6%. We achieve 3.97x speedup w.r.t neural network systolic accelerator with a similar area. The re-configurable nature of the compute engines enables various neural network operations and thereby supporting sequential networks (RNNs) and transformer models. Our quantitative analysis demonstrates 101×, 3× faster execution and 91×, 11× energy efficient than CPU and GPU respectively while running the transformer model, BERT-Base. Akshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Rangachar Srinivasa, Makesh Chandran, Kamlesh R. Pillai, Om Ji Omer, Narayanan Vijaykrishnan, Sreenivas Subramoney |
MICRO | 2 |
| 2019 | Visual Inertial Odometry At the Edge: A Hardware-Software Co-design Approach for Ultra-low Latency and PowerabstractVisual Inertial Odometry (VIO) is used for estimating pose and trajectory of a system and is a foundational requirement in many emerging applications like AR/VR, autonomous navigation in cars, drones and robots. In this paper, we analyze key compute bottlenecks in VIO and present a highly optimized VIO accelerator based on a hardware-software codesign approach. We detail a set of novel micro-architectural techniques that optimize compute, data movement, bandwidth and dynamic power to make it possible to deliver high quality of VIO at ultra-low latency and power required for budget constrained edge devices. By offloading the computation of the critical linear algebra algorithms from the CPU, the accelerator enables high sample rate IMU usage in VIO processing while acceleration of image processing pipe increases precision, robustness and reduces IMU induced drift in final pose estimate. The proposed accelerator requires a small silicon footprint (1.3 mm2in a 28nm process at 600 MHz), utilizes a modest on-chip shared SRAM (560KB) and achieves 10x speedup over a software-only implementation in terms of image sample-based pose update latency while consuming just 2.2 mW power. In a FPGA implementation, using the EuRoC VIO dataset (VGA 30fps images and 100Hz IMU) the accelerator design achieves pose estimation accuracy (loop closure error) comparable to a software based VIO implementation. Dipan Mandal, Srivatsava Jandhyala, Om Ji Omer, Gurpreet S. Kalsi, Biji George, Gopi Neela, Santhosh Kumar Rethinagiri, Sreenivas Subramoney, Lance Hacking, Jim Radford, Eagle Jones, Belliappa Kuttanna, Hong Wang 0003 |
DATE | 4 |