EDBT 2026 Demo / reviewers in the wild / expert
Kamlesh R. Pillai
dblp:278/5959
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ace-of-Spades: Accelerating Spatially Sparse Convolution for 3D Scene UnderstandingabstractSemantic understanding of 3D scenes is fundamental to many applications like robotics, autonomous driving, AR/VR. State-of-the-art methods for different 3D scene understanding tasks use 3D convolutional neural networks (CNNs) operating on point clouds. Convolution on spatially sparse data like point cloud involves irregular data accesses and compute patterns leading to poor utilization and energy efficiency in CPU/GPU implementations. The existing CNN accelerators designed for weight/activation sparsity cannot be efficiently repurposed for 3D spatially sparse CNNs given the fundamental differences in locating non-zero operands and granularity of work-dispatches. To address the dataflow challenges due to spatial sparsity and the need for specialized microarchitecture for spatially sparse convolution we present Ace-of-Spade s (AoS), an algorithm-dataflow-architecture co-designed system. AoS enables the data reuse among spatially proximate points using a locality-aware metadata structure along with a surface orientation aware point cloud reordering algorithm. AoS uses a novel technique for spatial sparsity aware selection of optimal data tiles by modelling the sparsity induced variations in the point cloud with a near-zero latency overheads. To accelerate computation on spatially sparse data, we propose a novel hardware accelerator Ss p nna with a front-end to convert varying number of operations per point into a stream of dense work dispatches to the backend compute engine. The compute engine further exploits weight and input feature data reuse through dynamic systolic grouping and multicast interconnects. The Ss p nna core together with the 64 KB of L1 memory requires 0.31 mm 2 of area in 10nm process at 1 GHz. Overall, AoS achieves speedup/energy savings of 19.9x / 49.9x and 2.2x / 7.1x over the state-of-the-art CPU and GPU implementations respectively. Om Ji Omer, Prashant Laddha, Gurpreet S. Kalsi, Kamlesh R. Pillai, Anirud Thyagharajan, Ahimanyu Kulkarni, Anbang Yao, Yurong Chen 0001, Sreenivas Subramoney |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2024 | ApHMM: Accelerating Profile Hidden Markov Models for Fast and Energy-efficient Genome AnalysisabstractProfile hidden Markov models (pHMMs) are widely employed in various bioinformatics applications to identify similarities between biological sequences, such as DNA or protein sequences. In pHMMs, sequences are represented as graph structures, where states and edges capture modifications (i.e., insertions, deletions, and substitutions) by assigning probabilities to them. These probabilities are subsequently used to compute the similarity score between a sequence and a pHMM graph. The Baum-Welch algorithm, a prevalent and highly accurate method, utilizes these probabilities to optimize and compute similarity scores. Accurate computation of these probabilities is essential for the correct identification of sequence similarities. However, the Baum-Welch algorithm is computationally intensive, and existing solutions offer either software-only or hardware-only approaches with fixed pHMM designs. When we analyze state-of-the-art works, we identify an urgent need for a flexible, high-performance, and energy-efficient hardware-software co-design to address the major inefficiencies in the Baum-Welch algorithm for pHMMs. We introduce ApHMM , the first flexible acceleration framework designed to significantly reduce both computational and energy overheads associated with the Baum-Welch algorithm for pHMMs. ApHMM employs hardware-software co-design to tackle the major inefficiencies in the Baum-Welch algorithm by (1) designing flexible hardware to accommodate various pHMM designs, (2) exploiting predictable data dependency patterns through on-chip memory with memoization techniques, (3) rapidly filtering out unnecessary computations using a hardware-based filter, and (4) minimizing redundant computations. ApHMM achieves substantial speedups of 15.55×–260.03×, 1.83×–5.34×, and 27.97× when compared to CPU, GPU, and FPGA implementations of the Baum-Welch algorithm, respectively. ApHMM outperforms state-of-the-art CPU implementations in three key bioinformatics applications: (1) error correction, (2) protein family search, and (3) multiple sequence alignment, by 1.29×–59.94×, 1.03×–1.75×, and 1.03×–1.95×, respectively, while improving their energy efficiency by 64.24×–115.46×, 1.75×, and 1.96×. Can Firtina, Kamlesh R. Pillai, Gurpreet S. Kalsi, Bharathwaj Suresh, Damla Senol Cali, Jeremie S. Kim, Taha Shahroodi, Meryem Banu Cavlak, Joël Lindegger, Mohammed Alser, Juan Gómez-Luna, Sreenivas Subramoney, Onur Mutlu |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | Compute-In-Memory Using 6T SRAM for a Wide Variety of WorkloadsabstractThis paper presents a split wordline 6T SRAM based compute-in-memory subarray with variable multi-bit precision for input operands and outputs. The split wordlines of the 6T cell enable sign segregation, thus allowing arbitrary sign/magnitude multiply and accumulate (MAC) operations. The arbitrary signmagnitude MAC operation extends the usage of the MAC array for DSP workloads as well. On top of that, the split wordline 6T cell is more resilient to write-disturb, which is a major concern for conventional 6T cell based compute-in-memory operation. The proposed system was designed and implemented using 65nm UMC technology and the energy efficiency and throughput are found to be 80.1 TOPS/W and 496 GOPS respectively, for maximum precision of inputs(4 bits)/outputs(5 bits). The maximum throughput and energy efficiency of the design are found to be 780 GOPS and 94.8 TOPS/W respectively. Hand-written digit recognition application mapped on the proposed system showed a maximum accuracy degradation of 0.2% as compared to that obtained from software with the same input/output quantization. Pramod Kumar Bharti, Kamlesh R. Pillai, Sagar Varma Sayyaparaju, Gurpreet S. Kalsi, Joycee Mekie, Sreenivas Subramoney |
ISCAS | 3 |
| 2021 | ONT-X: An FPGA Approach to Real-time Portable Genomic AnalysisabstractOxford Nanopore Technologies (ONT) MinION is a pocket-sized portable DNA sequencer for on-the-field DNA sequencing. The ONT data analysis pipeline requires substantial amounts of computational resources for the basecalling and read alignment step in the pipeline. We study the compute and memory requirement of the basecalling step and present the detailed microarchitecture for accelerating it on FPGA. We also make the case that an embedded FPGA-based SoC with partial reconfiguration (PR) capability is an ideal substrate for accelerating ONT data analysis pipeline, due to the diversity in these application kernels and the evolving nature of the algorithms. Focusing on two widely used ONT data analysis pipelines, we present two modes of PR operation to cater to either latency or throughput requirements of these two pipelines and demonstrate how to amortize the partial reconfiguration cost, overlap the kernels, minimize the data transfer cost and achieve 111× higher throughput and 6× reduction in latency using these two PR modes. C. N. Ramachandra, Anirban Nag, Rajeev Balasubramonian, Gurpreet S. Kalsi, Kamlesh R. Pillai, Sreenivas Subramoney |
FCCM | 5 |
| 2020 | Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network AccelerationabstractThis paper presents a Look-Up Table (LUT) based Processing-In-Memory (PIM) technique with the potential for running Neural Network inference tasks. We implement a bitline computing free technique to avoid frequent bitline accesses to the cache sub-arrays and thereby considerably reducing the memory access energy overhead. LUT in conjunction with the compute engines enables sub-array level parallelism while executing complex operations through data lookup which otherwise requires multiple cycles. Sub-array level parallelism and systolic input data flow ensure data movement to be confined to the SRAM slice.Our proposed LUT based PIM methodology exploits substantial parallelism using look-up tables, which does not alter the memory structure/organization, that is, preserving the bit-cell and peripherals of the existing SRAM monolithic arrays. Our solution achieves 1.72× higher performance and 3.14x lower energy as compared to a state-of-the-art processing-in-cache solution. Sub-array level design modifications to incorporate LUT along with the compute engines will increase the overall cache area by 5.6%. We achieve 3.97x speedup w.r.t neural network systolic accelerator with a similar area. The re-configurable nature of the compute engines enables various neural network operations and thereby supporting sequential networks (RNNs) and transformer models. Our quantitative analysis demonstrates 101×, 3× faster execution and 91×, 11× energy efficient than CPU and GPU respectively while running the transformer model, BERT-Base. Akshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Rangachar Srinivasa, Makesh Chandran, Kamlesh R. Pillai, Om Ji Omer, Narayanan Vijaykrishnan, Sreenivas Subramoney |
MICRO | 5 |