Sudipta Mondal

dblp:218/1179 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2024 Hardware Acceleration of Inference on Dynamic GNNs
abstract
Dynamic graph neural networks (DGNNs) play a crucial role in applications that require inferencing on graph-structured data, where the connectivity and features of the graph evolve over time. The proposed platform integrates graph neural network (GNN) and recurrent neural network (RNN) components of DGNNs, providing a unified platform that captures spatial and temporal information. Novel contributions include optimized cache reuse, a novel caching policy, and efficient GNN-RNN pipelining. Average energy efficiency gains of 8393X, 183x, and 87X -- 10X, and inference speedups of 1796X, 77X, and 21x -- 2.4X, over Intel Xeon Gold CPU, NVIDIA V100 GPU, and prior approaches, respectively, are demonstrated across multiple graph datasets and multiple DGNNs.
Sudipta Mondal, Sachin S. Sapatnekar
ISLPED1
2023 A Multicore GNN Training Accelerator
abstract
Graph neural networks (GNN) are vital for analytics on real-world problems with graph models. This work develops a multicore GNN training accelerator and develops multicore-specific optimizations for superior performance. It uses enhanced multicore-specific dynamic caching to circumvent the costs of irregular DRAM access patterns of graph-structured data. A novel feature vector segmentation approach is used to maximize on-chip data reuse with high on-chip computation per memory access, reducing data access latency, using a machine learning model for optimal performance. The work presents a major advance over prior FPGA/ASIC GNN accelerators by handling significantly larger datasets (with up to 8.6M vertices) on a variety of GNN models. On average, training speedup of 17× and energy efficiency improvement of 322× is achieved over DGL on a GPU; a speedup of 14× with 268× lower energy is shown over GPU-based GNNAdvisor; and 11× and 24× speedups are obtained over ASIC-based Rubik and FPGA-based GraphACT.
Sudipta Mondal, Ramprasath Srinivasa Gopalakrishnan, Ziqing Zeng, Kishor Kunal, Sachin S. Sapatnekar
ISLPED1
2023 A Unified Engine for Accelerating GNN Weighting/Aggregation Operations, With Efficient Load Balancing and Graph-Specific Caching
abstract
Graph neural networks (GNNs) analysis engines are vital for real-world problems that use large graph models. Challenges for a GNN hardware platform include the ability to 1) host a variety of GNNs; 2) handle high sparsity in input vertex feature vectors and the graph adjacency matrix and the accompanying random memory access patterns; and 3) maintain load-balanced computation in the face of uneven workloads, induced by high sparsity and power-law vertex degree distributions. This article proposes GNNIE, an accelerator designed to run a broad range of GNNs. It tackles workload imbalance by 1) splitting vertex feature operands into blocks; 2) reordering and redistributing computations; and 3) using a novel flexible MAC architecture. It adopts a graph-specific, degree-aware caching policy that is well suited to real-world graph characteristics. The policy enhances on-chip data reuse and avoids random memory access to DRAM. GNNIE achieves average speedups of$7197\times $over a CPU and$17.81\times $over a GPU over multiple datasets on graph attention networks (GATs), graph convolutional networks (GCNs), GraphSAGE, GINConv, and DiffPool. Compared to prior approaches, GNNIE achieves an average speedup of$5\times $over HyGCN (which cannot implement GATs) for GCN, GraphSAGE, and GINConv. GNNIE achieves an average speedup of$1.3\times $over AWB-GCN (which runs only GCNs), despite using$3.4\times $fewer processing units.
Sudipta Mondal, Susmita Dey Manasi, Kishor Kunal, Ramprasath Srinivasa Gopalakrishnan, Ziqing Zeng, Sachin S. Sapatnekar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 GNNIE: GNN inference engine with load-balancing and graph-specific caching
abstract
Graph neural networks (GNN) inferencing involves weighting vertex feature vectors, followed by aggregating weighted vectors over a vertex neighborhood. High and variable sparsity in the input vertex feature vectors, and high sparsity and power-law degree distributions in the adjacency matrix, can lead to (a) unbalanced loads and (b) inefficient random memory accesses. GNNIE ensures load-balancing by splitting features into blocks, proposing a flexible MAC architecture, and employing load (re)distribution. GNNIE's novel caching scheme bypasses the high costs of random DRAM accesses. GNNIE shows high speedups over CPUs/GPUs; it is faster and runs a broader range of GNNs than existing accelerators.
Sudipta Mondal, Susmita Dey Manasi, Kishor Kunal, Ramprasath Srinivasa Gopalakrishnan, Sachin S. Sapatnekar
DAC1
2018 Pre-assembly testing of interconnects in embedded multi-die interconnect bridge (EMIB) dies
abstract
The embedded multi-die interconnect bridge (EMIB) is an advanced packaging technology for 2.5D integration. This paper presents a bridge test architecture based on the proposed IEEE Std. P1838. The proposed test method enables access to interconnects at a pre-assembly stage by pairing the interconnects using metal shorts and probing on coarse-pitch C4 bumps. It can efficiently detect resistive-open and resistive-short defects in the bridge interconnects and micro-bumps. Simulation results are presented to evaluate the range of defects that can be detected by the proposed method.
Sudipta Mondal, Krishnendu Chakrabarty
DATE1