EDBT 2026 Demo / reviewers in the wild / expert
Mayank Kabra
dblp:29/2110
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0000-2304-3124ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesabstractNear-data processing (NDP) mitigates the data movement bottleneck in modern computing systems by performing computation close to where the data resides. Solid-state drives (SSDs) are well suited for NDP because they: (1) store large application datasets that exceed main memory capacity, and (2) contain multiple heterogeneous computation resources, e.g., general-purpose embedded cores in the SSD controller, DRAM chips, and NAND flash chips, which enable three NDP paradigms: in-storage processing (ISP), processing using DRAM in the SSD (PuD-SSD), and in-flash processing (IFP). These resources offer massive internal parallelism and enable in-place computation, which reduces data movement across the memory hierarchy. A large body of prior SSD-based NDP techniques operate in isolation, mapping computations to only one or two NDP paradigms (i.e., ISP, PuD-SSD, or IFP) within the SSD. These techniques (1) are tailored to specific workloads or kernels, (2) do not offload computations across all three NDP paradigms in the SSD and thus fail to exploit the full computational potential of an SSD, and (3) lack programmer-transparency, often requiring significant manual effort to identify offloadable code regions and map them to the SSD computation resources, which limits their general applicability and ease of deployment. While several prior works propose techniques to partition computation between the host and near-memory accelerators, adapting these techniques to SSDs offers limited benefits because they (1) ignore the heterogeneity of the SSD computation resources, and (2) make offloading decisions based on limited factors such as bandwidth utilization, data movement cost, or memory intensity, while ignoring key factors such as resource utilization. We propose Conduit, a general-purpose, programmertransparent NDP framework for SSDs that accelerates a broad range of workloads by leveraging available SSD computation resources. Conduit operates in two stages. At compile time, Conduit executes a custom compiler (e.g., LLVM) pass that (i) vectorizes suitable application code segments into single-instruction multiple-data (SIMD) operations that align with the SSD's page layout, and (ii) embeds metadata (e.g., operation type, operand sizes) into the vectorized instructions to guide runtime offloading decisions. At runtime, within the SSD, Conduit performs instruction-granularity offloading by evaluating six key application and system features (e.g., operation type, computation resource utilization, data dependence delay), and uses a cost function to select the most suitable SSD computation resource to execute each vectorized instruction. We evaluate Conduit and two prior NDP offloading techniques using an in-house event-driven SSD simulator on six data-intensive applications (e.g., large language model inference and training, encryption). Conduit outperforms the best-performing prior offloading policy by$1.8 \times$and reduces energy consumption by 46 %, with small latency and storage overheads, and no additional hardware cost. Rakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta, Rahul Bera, Nika Mansouri-Ghiasi, Nanditha Rao, Qingcai Jiang, Andreas Kosmas Kakolyris, Yu Liang 0004, Mohammad Sadrosadati, Onur Mutlu |
HPCA | 3 |
| 2026 | Long integer NTT execution on UPMEM-PIM for 128-bit secure fully homomorphic encryptionabstractFully Homomorphic Encryption (FHE) enables secure computations on encrypted data, hence becoming an appealing technology for privacy-preserving data processing. A core kernel in many cryptographic and FHE workloads is the Number Theoretic Transform (NTT). While NTT involves frequent non-contiguous data accesses, limiting overall performance, processing–in–memory (PIM) has the potential to address this limitation. PIM, performing computations close to the data, reduces the need for extensive data transfers between memory and compute units. However, the performance of current PIM solutions is limited by inherent factors related to the integration of processing capabilities within memory modules. In this article we analyze the performance trade-offs of NTT kernel designs along with optimized modular multiplication algorithms on PIM systems based on UPMEM hardware. Our results include significant performance improvements of up to 2.9 × over state–of–the–art approaches on UPMEM-PIM, while preserving, for the first time in the literature, 128-bit security at high precision. Tathagata Barik, Priyam Mehta, Zaira Pindado, Harshita Gupta, Mayank Kabra, Mohammad Sadrosadati, Onur Mutlu, Antonio J. Peña |
Future Gener. Comput. Syst. | 5 |
| 2025 | CIPHERMATCH: Accelerating Homomorphic Encryption-Based String Matching via Memory-Efficient Data Packing and In-Flash ProcessingabstractHomomorphic encryption (HE) allows secure computation on encrypted data without revealing the original data, providing significant benefits for privacy-sensitive applications. Many cloud computing applications (e.g., DNA read mapping, biometric matching, web search) use exact string matching as a key operation. However, prior string matching algorithms that use homomorphic encryption are limited by high computational latency caused by the use of complex operations and data movement bottlenecks due to the large encrypted data size. In this work, we provide an efficient algorithm-hardware codesign to accelerate HE-based secure exact string matching. We propose CIPHERMATCH, which (i) reduces the increase in memory footprint after encryption using an optimized software-based data packing scheme, (ii) eliminates the use of costly homomorphic operations (e.g., multiplication and rotation), and (iii) reduces data movement by designing a new in-flash processing (IFP) architecture. Mayank Kabra, Rakesh Nadig, Harshita Gupta, Rahul Bera, Manos Frouzakis, Vamanan Arulchelvan, Yu Liang 0004, Haiyu Mao, Mohammad Sadrosadati, Onur Mutlu |
ASPLOS (2) | 1 |
| 2025 | Proteus: Achieving High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible ArithmeticabstractProcessing-using-DRAM (PUD) is a paradigm where the analog operational properties of DRAM are used to perform bulk logic operations.While PUD promises high throughput at low energy and area cost, we uncover three limitations of existing PUD approaches that lead to significant inefficiencies: (i) static data representation, i.e., two's complement with fixed bit-precision, leading to unnecessary computation over useless (i.e., inconsequential) data; (ii) support for only throughput-oriented execution, where the high latency of Geraldo F. Oliveira, Mayank Kabra, Kangqi Chen, A. Giray Yaglikçi, Melina Soysal, Mohammad Sadrosadati, Joaquín Olivares 0001, Saugata Ghose, Juan Gómez-Luna, Onur Mutlu |
ICS | 2 |
| 2023 | GCells: A Graph-Search Approach to Design Custom Cells for Computational SubsystemsabstractStandard cell design is a challenge considering its impact on the overall synthesis of the design. The standard cells are generally offered by the manufacturers and the cells are updated to match the advancement in the fabrication facility. However matching of the system designed by the non-manufacturing team with the standard cells may not always present the best results. The exploration of standard cell design is very limited and not much is disclosed due to intellectual protection (IPs). This paper proposes a graph-search algorithm to extract the popular set of connected nodes, where the graph represents the computational subsystem and the nodes reflect individual gates. The proposed graph-search approach was applied on MAC designs of different bit-widths to extract top four ranked custom cells of four inputs and further characterized to incorporate them in the CMOS implemented standard cell library. The extracted top four cells were optimized for transistor widths to three different targets including critical path delay, product-of-power-and-delay (PDP), and equal rising and falling resistances separately. The augmented library with extracted custom cells was applied for synthesizing MAC design of four different bit-widths using three different set of library cells. The custom cells in the range of 38.23% to 52.80% were mapped in the synthesized versions of MAC designs for all the three optimized versions of cells, which validates the approach and emphasizes the need to explore new cells for acquiring performance and power-efficient subsystem designs. The standard cell library incorporated with custom cells offered maximum performance improvement of 35.3% and power savings of 56% over the standard cell library when synthesized for MAC designs of different bit-widths. The custom library synthesized MAC designs when adopted as hardware accelerator units for LeNet, AlexNet, and VGG-16 network showcased a performance gain of 23% to 33.80%, over the MAC designs synthesized by the original standard cell library. The graph-search approach is a step towards automating the custom cells for the design under synthesis. The approach has the potential to realize complex functions with customized cells, applicable to modern day SoC design. Mayank Kabra, Shreyas V. S, Prashanth H. C., Kedar Deshpande, Madhav Rao |
DSD | 1 |
| 2023 | Design and Evaluation of Finite Field Multipliers Using Fast XNOR CellsabstractThe current polynomial multiplication is built on conventional CMOS cells, and no major changes are explored in the standard cell library to improve the performance. Hence state-of-the-art (SOTA) finite field multipliers of operand sizes ranging from 93 to 409 bits were designed and evaluated by adopting faster XNOR cells. The hardware metrics in the form of gates usage, and propagation delay were compared. The SOTA multipliers of different approaches including Conventional Algorithm~(CA), Karatsuba Algorithm~(KA), Overlap free Karatsuba Algorithm~(OKA), and Overlap-free based multiplication strategy(OBS) were designed and synthesized through ASIC flow using 45~nm GPDK library files. The fast XNOR cell adopted SOTA multipliers improved the compute delay in the range of 8.24% to 33.45%, 8% to 37.05%, 4.63% to 18.36%, and 1.01% to 38.73% for OKA, OBS, CA, and KA respectively. All the design files are made freely available for further usage to research and designers' community. Nitin D. Patwari, Anjul Srivastav, Mayank Kabra, Prashanth Jonna, Madhav Rao |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Understanding classifier errors by examining influential neighborsabstractModern supervised learning algorithms can learn very accurate and complex discriminating functions. But when these classifiers fail, this complexity can also be a drawback because there is no easy, intuitive way to diagnose why they are failing and remedy the problem. This important question has received little attention. To address this problem, we propose a novel method to analyze and understand a classifier's errors. Our method centers around a measure of how much influence a training example has on the classifier's prediction for a test example. To understand why a classifier is mispredicting the label of a given test example, the user can find and review the most influential training examples that caused this misprediction, allowing them to focus their attention on relevant areas of the data space. This will aid the user in determining if and how the training data is inconsistently labeled or lacking in diversity, or if the feature representation is insufficient. As computing the influence of each training example is computationally impractical, we propose a novel distance metric to approximate influence for boosting classifiers that is fast enough to be used interactively. We also show several novel use paradigms of our distance metric. Through experiments, we show that it can be used to find incorrectly or inconsistently labeled training examples, to find specific areas of the data space that need more training data, and to gain insight into which features are missing from the current representation. Mayank Kabra, Alice Robie, Kristin Branson |
CVPR | 1 |
| 2007 | Learning the structure of manifolds using random projectionsabstractWe present a simple variant of the k-d tree which automatically adapts to intrinsic low dimensional structure in data. Yoav Freund, Sanjoy Dasgupta, Mayank Kabra, Nakul Verma |
NIPS | 3 |