Sumukh Pinge

dblp:361/6313 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0009-3186-610XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective Communication
abstract
Collective communication operations in AI and HPC workloads generate heavy network traffic. Offloading these operations to network switches reduces latency, but performing arithmetic and replication at line rate is difficult, especially as port counts and link speeds grow. Existing in-network approaches rely on accumulation buffers that not only limit throughput but also require complex state management to handle stragglers and congestion. We present CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs). Each CE combines spatial and temporal parallelism to perform reductions without accumulation buffers. CEs compose into k-ary n-tree topologies, scaling to 32- and 256-port switches while preserving non-blocking throughput. Source routing and flit-level synchronization keep per-switch state minimal. Our FPGA implementation shows that CCSwitch's quaternary-tree reduction fabric uses up to 23% fewer LUTs and 12–30% fewer flip-flops than a comparable Clos-based design at equal throughput. Enabling the full feature set—source routing, replication, and time-multiplexed VCs—uses 1.4–1.8× more LUTs than the circuit-switched baseline, well below the 3–5× overhead typical of packet-switched NoC routers, while supporting concurrent collectives on shared links.
Sumukh Pinge, Hardik Soni 0001, Bob Lantz, Khaled Diab 0001, Lianjie Cao, Tajana Rosing, Puneet Sharma 0001
SIGCOMM1
2026 Proxima: Near-Storage Acceleration for Graph-Based Approximate Nearest Neighbor Search in 3D NAND
abstract
Approximate nearest neighbor search (ANNS) plays an indispensable role in a wide variety of applications, including recommendation systems, information retrieval, and semantic search. Among the cutting-edge ANNS algorithms, graph-based approaches provide superior accuracy and scalability on massive datasets. However, the best-performing graph-based ANNS solutions incur tens of hundreds of memory footprints as well as costly distance computation, thus hindering their efficient deployment at scale. The 3D NAND flash is emerging as a promising device for data-intensive applications due to its high density and nonvolatility. In this work, we present the near-storage processing (NSP)-based ANNS solution Proxima to accelerate graph-based ANNS with algorithm-hardware co-design in 3D NAND flash. Proxima significantly reduces the complexity of graph search by leveraging the distance approximation and early termination. On top of the algorithmic enhancement, we implement the Proxima search algorithm in 3D NAND flash using the heterogeneous integration technique. To maximize 3D NAND’s bandwidth utilization, we present a customized dataflow and optimized data allocation scheme. Our evaluation results show that, compared to graph ANNS on CPU and GPU, Proxima achieves a magnitude improvement in throughput or energy efficiency. Proxima yields 7× to 13× speedup over existing ASIC designs. Furthermore, Proxima achieves a good balance between accuracy, efficiency, and storage density compared to previous NSP-based accelerators.
Po-Kai Hsu, Jaeyoung Kang 0001, Minxuan Zhou, Sumukh Pinge, Shimeng Yu, Tajana Rosing
IEEE Trans. Computers6
2026 HyperMetric: Efficient Hyperdimensional Computing With Metric Learning for Robust Edge Intelligence
abstract
Hyperdimensional computing (HDC) is emerging as an efficient and robust computing paradigm that has strong resilience to various types of errors. The error robustness nature of HDC makes it a good match for error-prone memory systems. However, the mechanisms behind HDCs robustness are not fully understood. In this work, we propose HyperMetric, a framework to train highly robust and hardware-friendly HDC models. We found that HDC’s error resilience is driven by Hamming distance margin between hypervectors. Based on this, we propose HyperMetric training that is based on metric learning in order to optimize for high robustness. The experiments show that HyperMetric trained HDC models deliver up to 17W larger Hamming distance margin and up to 14.3 We accelerate HyperMetric trained models using ReRAM. As compared to state-of-the-art HDC algorithms OnlineHD and HyDREA, HyperMetric ReRAM accelerator is > 20% more accurate for computing-in-memory (CIM) errors and > 10% more accurate for bit errors even in the face of variations. Furthermore, HyperMetric hardware is 35% more accurate in comparison with existing tinyHD and GENERIC accelerators in the face of 3× ReRAM resistance variance, and 20% more accurate with BER of up to 20% due to voltage scaling while keeping a good balance between area, power, and processing laten
Sean Fuhrman, Keming Fan, Sumukh Pinge, Wei-Chen Chen, Tajana Rosing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) Flash
abstract
The rapid expansion of mass spectrometry (MS) data, now exceeding hundreds of terabytes, poses significant challenges for efficient, large-scale library search — a critical component for drug discovery. Traditional processors struggle to handle this data volume efficiently, making in-storage computing (ISP) a promising alternative. This work introduces an ISP architecture leveraging a 3D Ferroelectric NAND (FeNAND) structure, providing significantly higher density, faster speeds, and lower voltage requirements compared to traditional NAND flash. Despite its superior density, the NAND structure has not been widely utilized in ISP applications due to limited throughput associated with row-by-row reads from serially connected cells. To overcome these limitations, we integrate hyperdimensional computing (HDC), a brain-inspired paradigm that enables highly parallel processing with simple operations and strong error tolerance. By combining HDC with the proposed dual-bound approximate matching (D-BAM) distance metric, tailored to the FeNAND structure, we parallelize vector computations to enable efficient MS spectral library search, achieving 43× speedup and 21× higher energy efficiency over state-of-the-art 3D NAND methods, while maintaining comparable accuracy.
Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H. Pantha, Po-Kai Hsu, Zihan Xia 0002, Flavio Ponzina, Winston Chern, Taeyoung Song, Priyankka Gundlapudi Ravikumar, Mengkun Tian, Lance Fernandes, Hari Jayasankar, Chinsung Park, Amrit Garlapati, Kijoon Kim, Jongho Woo, Suhwan Lim, Wanki Kim, Daewon Ha, Duygu Kuzum, Shimeng Yu, Tajana Rosing, Mingu Kang
ICCAD1
2025 PATHE: A Privacy-Preserving Database Pattern Search Platform with Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) enables secure computation on encrypted data without decryption, allowing a great opportunity for privacy-preserving computation. Many companies maintain extensive, high-quality databases to deliver services, making preserving data privacy during the database pattern searches crucial. With FHE, the server can take encrypted queries from clients and search through the reference database on the server without decryption, thus guaranteeing data security for all parties. While FHE provides a promising solution to data privacy, it has severe drawbacks of explosive memory requirements and excessive latency, which amplify the computational and memory inefficiencies for database search applications.To address these, we propose PATHE that exploits FHE and hyperdimensional computing (HDC), which provides high parallelism, excellent robustness to errors, for high-performance privacy-preserving database search. On the software side, we propose an FHE-friendly PATHE algorithm that leverages efficient FHE-HDC search and a scheme-switching-based argmax to support database search and maintain comparable accuracy to the state-of-the-art. On the hardware side, PATHE proposes an efficient and scalable FHE accelerator system using Compute Express Link (CXL) for large-scale FHE database search, along with a novel, storage-aware dataflow designed to optimize memory and storage transfers for large database workloads. We evaluate PATHE on the large-scale encrypted database of protein mass spectra, PATHE achieves 2.1× speedup and 1.7× better energy efficiency compared to the baseline system.
Xuan Wang 0040, Minxuan Zhou, Gabrielle De Micheli, Yujin Nam, Sumukh Pinge, Augusto Vega, Tajana Rosing
ICCAD5
2025 HPVM-HDC: A Heterogeneous Programming System for Accelerating Hyperdimensional Computing
abstract
Hyperdimensional Computing (HDC), a technique inspired by cognitive models of computation, has been proposed as an efficient and robust alternative basis for machine learning.HDC programs are often manually written in low-level and target specific languages targeting CPUs, GPUs, and FPGAs-these codes cannot be easily retargeted onto HDC-specific accelerators.No previous programming system enables productive development of HDC programs and generates efficient code for several hardware targets.We propose a heterogeneous programming system for HDC: a novel programming language, HDC++, for writing applications using a unified programming model, including HDC-specific primitives to improve programmability, and a heterogeneous compiler, HPVM-HDC, that provides an intermediate representation for compiling HDC programs to many hardware targets.We implement two tuning optimizations, automatic binarization and reduction perforation, that exploit the error resilient nature of HDC.Our evaluation shows that HPVM-HDC generates performance-competitive code for CPUs and GPUs, achieving a geomean speed-up of 1.17x over optimized baseline CUDA implementations with a geomean * Equally contributing authors.
Russel Arbore, Xavier Routh, Abdul Rafae Noor, Akash Kothari, Haichao Yang, Sumukh Pinge, Minxuan Zhou, Tajana Rosing, Vikram S. Adve
ISCA7
2025 SmartMS: Efficient Hierarchical Database Search for Mass Spectrometry via Processing-in-Memory
abstract
The acceleration of Mass Spectrometry (MS) library search is crucial for advancing scientific and pharmaceutical research. Recent methodologies leverage Hyperdimensional Computing (HDC) to encode reference and query spectra as high-dimensional vectors, enabling highly parallel similarity computations. In this context, Processing-In-Memory (PIM) has emerged as a promising solution, offering orders of magnitude improvements in computational speed compared to GPU-based approaches when handling large-scale libraries. However, bruteforce search methods remain computationally intensive, exacerbating the high energy demands associated with MS library search operations in data centers. In this work, we propose SmartMS, a novel tool that leverages HDC to construct a multi-level database structure, reducing search complexity from linear to logarithmic while maintaining compatibility with PIM-based accelerators. SmartMS improves identification accuracy by 3% while delivering a 33× improvement in speed and a 58× energy reduction, with a negligible increase in memory requirements of 0.5% when compared to the current state of the art.
Flavio Ponzina, Sumukh Pinge, Abhijay Deevi, Yilin Ge, Mingu Kang, Tajana Rosing
ISLPED2
2024 Efficient Open Modification Spectral Library Searching in High-Dimensional Space with Multi-Level-Cell Memory
abstract
Open Modification Search (OMS) is a promising algorithm for mass spectrometry analysis that enables the discovery of modified peptides. However, OMS encounters challenges as it exponentially extends the search scope. Existing OMS accelerators either have limited parallelism or struggle to scale effectively with growing data volumes. In this work, we introduce an OMS accelerator utilizing multi-level-cell (MLC) RRAM memory to enhance storage capacity by 3x. Through in-memory computing, we achieve up to 77x faster data processing with two to three orders of magnitude better energy efficiency. Testing was done on a fabricated MLC RRAM chip. We leverage hyperdimensional computing to tolerate up to 10% memory errors while delivering massive parallelism in hardware.
Keming Fan, Wei-Chen Chen, Sumukh Pinge, H.-S. Philip Wong, Tajana Rosing
DAC3
2024 SpectraFlux: Harnessing the Flow of Multi-FPGA in Mass Spectrometry Clustering
abstract
The identification and quantification of proteins through mass spectrometry (MS) are foundational to proteomics, offering insights into biological systems and disease states. However, current clustering tools struggle to process large-scale datasets. We propose SpectraFlux, a multiple FPGA-based architecture for accelerated mass spectrum clustering that outperforms existing CPU, GPU, and FPGA designs. It employs heterogeneous clustering kernels for adaptive bucket size management and optimizes memory usage by distinguishing between on-chip and high-bandwidth memory (HBM) storage solutions. SpectraFlux is built upon the TAPA-CS framework, which automatically compiles and partitions a large dataflow design across multiple chips with RDMA-based inter-FPGA communication. Our solution shows a 2.7X speed up on a quad-FPGA platform compared to a single FPGA. Additionally, we introduce a refined cost model for frame-based inter-FPGA communication to better accommodate the variable data rates inherent in proteomic data processing. This reduces the inter-FPGA data movement by up to 73%. Finally, SpectraFlux achieves speedups of up to 11X and 17X over SOTA FPGA and GPU accelerators, respectively.
Neha Prakriya, Sumukh Pinge, Jason Cong, Tajana Rosing
DAC3
2024 SpecHD: Hyperdimensional Computing Framework for FPGA-Based Mass Spectrometry Clustering
abstract
Mass spectrometry-based proteomics is a key enabler for personalized healthcare, providing a deep dive into the complex protein compositions of biological systems. This technology has vast applications in biotechnology and biomedicine but faces significant computational bottlenecks. Current methodologies often require multiple hours or even days to process extensive datasets, particularly in the domain of spectral clustering. To tackle these inefficiencies, we introduce SpecHD, a hyperdimensional computing (HDC) framework supplemented by an FPGA-accelerated architecture with integrated near-storage preprocessing. Utilizing streamlined binary operations in an HDC environment, SpecHD capitalizes on the low-latency and parallel capabilities of FPGAs. This approach markedly improves clustering speed and efficiency, serving as a catalyst for real-time, high-throughput data analysis in future healthcare applications. Our evaluations demonstrate that SpecHD not only maintains but often surpasses existing clustering quality metrics while drastically cutting computational time. Specifically, it can cluster a large-scale human proteome dataset-comprising 25 million MS/MS spectra and 131 GB of MS data-in just 5 minutes. With energy efficiency exceeding 31x and a speedup factor that spans a range of 6x to 54x over existing state-of-the-art solutions, SpecHD emerges as a promising solution for the rapid analysis of mass spectrometry data with great implications for personalized healthcare.
Sumukh Pinge, Jaeyoung Kang 0001, Niema Moshiri, Wout Bittremieux, Tajana Rosing
DATE1
2023 HyperMetric: Robust Hyperdimensional Computing on Error-prone Memories using Metric Learning
abstract
Hyperdimensional computing (HDC) is emerging as an efficient and robust computing paradigm that has strong resilience to various types of errors. The robustness of HDC makes it a good match for error-prone memory systems. In this work, we propose HyperMetric, a framework to develop highly robust and hardware-friendly HDC models. First, we propose HyperMetric training which is based on metric learning to optimize for high robustness. The experiments show that HyperMetric-trained HDC models deliver up to 17× larger distance margin and 14.3% accuracy gain. Compared to state-of-the-art HDC algorithms OnlineHD [1] and HyDREA [2], HyperMetric ReRAM accelerator is > 20% more accurate for computing-in-memory (CIM) errors and > 10% more accurate for bit errors even in the face of variations. Furthermore, HyperMetric hardware is 35% more accurate in comparison with state of the art tinyHD [3] and GENERIC [4] accelerators in the face of 3× ReRAM resistance variance, and 20% more accurate with BER of up to 20% due to voltage scaling while keeping a good balance between area, power, and processing latency.
Viji Swaminathan, Sumukh Pinge, Sean Fuhrman, Tajana Rosing
ICCD3