Susmita Dey Manasi

dblp:185/5739 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-9358-6255ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2025 Performance Analysis of CNN Inference/Training with Convolution and Non-Convolution Operations on ASIC Accelerators
abstract
Today’s performance analysis frameworks for deep learning accelerators suffer from two significant limitations. First, although modern convolutional neural networks (CNNs) consist of many types of layers other than convolution, especially during training, these frameworks largely focus on convolution layers only. Second, these frameworks are generally targeted towards inference and lack support for training operations. This work proposes a novel open-source performance analysis framework, SimDIT, for general ASIC-based systolic hardware accelerator platforms. The modeling effort of SimDIT comprehensively covers convolution and non-convolution operations of both CNN inference and training on a highly parameterizable hardware substrate. SimDIT is integrated with a backend silicon implementation flow and provides detailed end-to-end performance statistics (i.e., data access cost, cycle counts, energy, and power) for executing CNN inference and training workloads. SimDIT-enabled performance analysis reveals that on a 64×64 processing array, non-convolution operations constitute 59.5% of total runtime for ResNet-50 training workload. In addition, by optimally distributing available off-chip DRAM bandwidth and on-chip SRAM resources, SimDIT achieves 18× performance improvement over a generic static resource allocation for ResNet-50 inference.
Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew B. Kahng, Sean Kinzer, Susmita Dey Manasi, Sachin S. Sapatnekar, Zhiang Wang
ACM Trans. Design Autom. Electr. Syst.5
2024 An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning Accelerators
abstract
Parameterizable machine learning (ML) accelerators are the product of recent breakthroughs in ML. To fully enable their design space exploration (DSE), we propose a physical-design-driven, learning-based prediction framework for hardware-accelerated deep neural network (DNN) and non-DNN ML algorithms. It adopts a unified approach that combines power, performance, and area (PPA) analysis with frontend performance simulation, thereby achieving a realistic estimation of both backend PPA and system metrics such as runtime and energy. In addition, our framework includes a fully automated DSE technique, which optimizes backend and system metrics through an automated search of architectural and backend parameters. Experimental studies show that our approach consistently predicts backend PPA and system metrics with an average 7% or less prediction error for the ASIC implementation of two deep learning accelerator platforms, VTA and VeriGOOD-ML, in both a commercial 12 nm process and a research-oriented 45 nm process.
Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Sayak Kundu, Rohan Mahapatra, Susmita Dey Manasi, Sachin S. Sapatnekar, Zhiang Wang, Ziqing Zeng
ACM Trans. Design Autom. Electr. Syst.8
2023 Reusing GEMM Hardware for Efficient Execution of Depthwise Separable Convolution on ASIC-Based DNN Accelerators
abstract
Deep learning (DL) accelerators are optimized for standard convolution. However, lightweight convolutional neural networks (CNNs) use depthwise convolution (DwC) in key layers, and the structural difference between DwC and standard convolution leads to significant performance bottleneck in executing lightweight CNNs on such platforms. This work reuses the fast general matrix-vector multiplication (GEMM) core of DL accelerators by mapping DwC to channel-wise parallel matrix-vector multiplications. An analytical framework is developed to guide pre-RTL hardware choices, and new hardware modules and software support are developed for end-to-end evaluation of the solution. This GEMM-based DwC execution strategy offers substantial performance gains for lightweight CNNs: 7× speedup and 1.8× lower off-chip communication for MobileNet-v1 over a conventional DL accelerator, and 74× speedup over a CPU, and even 1.4× speedup over a power-hungry GPU.
Susmita Dey Manasi, Suvadeep Banerjee, Abhijit Davare, Anton A. Sorokin, Steven M. Burns, Desmond Kirkpatrick, Sachin S. Sapatnekar
ASP-DAC1
2023 A Unified Engine for Accelerating GNN Weighting/Aggregation Operations, With Efficient Load Balancing and Graph-Specific Caching
abstract
Graph neural networks (GNNs) analysis engines are vital for real-world problems that use large graph models. Challenges for a GNN hardware platform include the ability to 1) host a variety of GNNs; 2) handle high sparsity in input vertex feature vectors and the graph adjacency matrix and the accompanying random memory access patterns; and 3) maintain load-balanced computation in the face of uneven workloads, induced by high sparsity and power-law vertex degree distributions. This article proposes GNNIE, an accelerator designed to run a broad range of GNNs. It tackles workload imbalance by 1) splitting vertex feature operands into blocks; 2) reordering and redistributing computations; and 3) using a novel flexible MAC architecture. It adopts a graph-specific, degree-aware caching policy that is well suited to real-world graph characteristics. The policy enhances on-chip data reuse and avoids random memory access to DRAM. GNNIE achieves average speedups of$7197\times $over a CPU and$17.81\times $over a GPU over multiple datasets on graph attention networks (GATs), graph convolutional networks (GCNs), GraphSAGE, GINConv, and DiffPool. Compared to prior approaches, GNNIE achieves an average speedup of$5\times $over HyGCN (which cannot implement GATs) for GCN, GraphSAGE, and GINConv. GNNIE achieves an average speedup of$1.3\times $over AWB-GCN (which runs only GCNs), despite using$3.4\times $fewer processing units.
Sudipta Mondal, Susmita Dey Manasi, Kishor Kunal, Ramprasath Srinivasa Gopalakrishnan, Ziqing Zeng, Sachin S. Sapatnekar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 GNNIE: GNN inference engine with load-balancing and graph-specific caching
abstract
Graph neural networks (GNN) inferencing involves weighting vertex feature vectors, followed by aggregating weighted vectors over a vertex neighborhood. High and variable sparsity in the input vertex feature vectors, and high sparsity and power-law degree distributions in the adjacency matrix, can lead to (a) unbalanced loads and (b) inefficient random memory accesses. GNNIE ensures load-balancing by splitting features into blocks, proposing a flexible MAC architecture, and employing load (re)distribution. GNNIE's novel caching scheme bypasses the high costs of random DRAM accesses. GNNIE shows high speedups over CPUs/GPUs; it is faster and runs a broader range of GNNs than existing accelerators.
Sudipta Mondal, Susmita Dey Manasi, Kishor Kunal, Ramprasath Srinivasa Gopalakrishnan, Sachin S. Sapatnekar
DAC2
2021 Fast and Efficient Constraint Evaluation of Analog Layout Using Machine Learning Models
abstract
Placement algorithms for analog circuits explore numerous layout configurations in their iterative search. To steer these engines towards layouts that meet the electrical constraints on the design, this work develops a fast feasibility predictor to guide the layout engine. The flow first discerns rough bounds on layout parasitics and prunes the feature space. Next, a Latin hypercube sampling technique is used to sample the reduced search space, and the labeled samples are classified by a linear support vector machine (SVM). If necessary, a denser sample set is used for the SVM, or if the constraints are found to be nonlinear, a multilayer perceptron (MLP) is employed. The resulting machine learning model demonstrated to rapidly evaluate candidate placements in a placer, and is used to build layouts for several analog blocks.
Tonmoy Dhar, Jitesh Poojary, Kishor Kunal, Meghna Madhusudan, Arvind K. Sharma, Susmita Dey Manasi, Jiang Hu 0001, Ramesh Harjani, Sachin S. Sapatnekar
ASP-DAC7
2021 DeepOpt: Optimized Scheduling of CNN Workloads for ASIC-based Systolic Deep Learning Accelerators
abstract
Scheduling computations in each layer of a convolutional neural network on a deep learning (DL) accelerator involves a large number of choices, each of which involves a different set of memory reuse and memory access patterns. Since memory transactions are the primary bottleneck in DL acceleration, these choices can strongly impact the energy and throughput of the accelerator. This work proposes an optimization framework, DeepOpt, for general ASIC-based systolic hardware accelerators for layer-specific and hardware-specific scheduling strategy for each layer of a CNN to optimize energy and latency. Optimal hardware allocation significantly reduces execution cost as compared to generic static hardware resource allocation, e.g., improvements of up to 50x in the energy-delay product for VGG-16 and 41x for GoogleNet-v1.
Susmita Dey Manasi, Sachin S. Sapatnekar
ASP-DAC1
2021 VeriGOOD-ML: An Open-Source Flow for Automated ML Hardware Synthesis
abstract
This paper introduces VeriGOOD-ML, an automated methodology for generating Verilog with no human in the loop, starting from a high-level description of a machine learning (ML) algorithm in a standard format such as ONNX. The Verilog RTL is then translated through a back-end design flow to GDSII, driven by a design planning approach that is well tailored to the macro-intensive nature of ML platforms. VeriGOOD-ML uses three approaches to build ML hardware: the TABLA platform uses a dataflow architecture that is well suited to non-DNN ML algorithms; the GeneSys platform, with a systolic array and a SIMD array, is optimized for implementing DNNs; and the Axiline approach synthesizes small ML algorithms by hardcoding the structure of the algorithm into hardware, thus trading off flexibility for performance and power. The overall approach explores the design space of platform configurations and Pareto-optimal-PPA back-end implementations to yield designs that represent different tradeoffs at the algorithmic level between area, power, performance, and execution time. The overall methodology, from architecture to back-end design to hardware implementation, is described in this paper, and the results of VeriGOOD-ML are demonstrated on a set of ML benchmarks.
Hadi Esmaeilzadeh, Soroush Ghodrati, Jie Gu 0003, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Rohan Mahapatra, Susmita Dey Manasi, Edwin Mascarenhas, Sachin S. Sapatnekar, Ravi Varadarajan, Zhiang Wang, Hanyang Xu 0002, Brahmendra Reddy Yatham, Ziqing Zeng
ICCAD9
2021 SeFAct2: Selective Feature Activation for Energy-Efficient CNNs Using Optimized Thresholds
abstract
This work presents a framework for dynamic energy reduction in hardware accelerators for convolutional neural networks (CNNs). The key idea is based on the early prediction of the features that may be important, with the deactivation of computations related to unimportant features and static bitwidth reduction. The former is applied in late layers of the CNN, while the latter is more effective in the early layers. The procedure includes a methodology for automated threshold tuning to detect feature activation. For various state-of-the-art neural networks, the results show that energy savings of up to about 30% are achievable, after accounting for all implementation overheads, with a small loss in the accuracy.
Farhana Sharmin Snigdha, Susmita Dey Manasi, Jiang Hu 0001, Sachin S. Sapatnekar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 NeuPart: Using Analytical Models to Drive Energy-Efficient Partitioning of CNN Computations on Cloud-Connected Mobile Clients
abstract
Data processing on convolutional neural networks (CNNs) places a heavy burden on energy-constrained mobile platforms. This article optimizes energy on a mobile client by partitioning CNN computations between in situ processing on the client and offloaded computations in the cloud. A new analytical CNN energy model is formulated, capturing all major components of the in situ computation, for ASIC-based deep learning accelerators. The model is benchmarked against measured silicon data. The analytical framework is used to determine the optimal energy partition point between the client and the cloud at runtime. On standard CNN topologies, partitioned computation is demonstrated to provide significant energy savings on the client over a fully cloud-based computation or fully in situ computation. For example, at 80 Mbps effective bit rate and 0.78 W transmission power, the optimal partition for AlexNet [SqueezeNet] saves up to 52.4% [73.4%] energy over a fully cloud-based computation and 27.3% [28.8%] energy over a fully in situ computation.
Susmita Dey Manasi, Farhana Sharmin Snigdha, Sachin S. Sapatnekar
IEEE Trans. Very Large Scale Integr. Syst.1
2019 SeFAct: selective feature activation and early classification for CNNs
abstract
This work presents a dynamic energy reduction approach for hardware accelerators for convolutional neural networks (CNN). Two methods are used: (1) an adaptive data-dependent scheme to selectively activate a subset of all neurons, by narrowing down the possible activated classes (2) static bitwidth reduction. The former is applied in late layers of the CNN, while the latter is more effective in early layers. Even accounting for the implementation overheads, the results show 20%--25% energy savings with 5--10% accuracy loss.
Farhana Sharmin Snigdha, Ibrahim Ahmed 0002, Susmita Dey Manasi, Meghna G. Mankalale, Jiang Hu 0001, Sachin S. Sapatnekar
ASP-DAC3
2016 Gate/source-overlapped heterojunction Tunnel FET-based LAMSTAR neural network and its Application to EEG Signal Classification
abstract
This paper explores reduced complexity physical implementation of self-organizing-map (SOM) and LAMSTAR (Large Scale Memory Storage and Retrieval) neural network. Unique Gaussian IDS-VGScharacteristic of emerging gate/source-overlapped heterojunction Tunnel FET (SO-HTFET) is utilized to simplify the complexity of a SOM. For a given pattern, SO-HTFET-based SOM performs associative processing between the applied pattern feature and the stored neuron states. SO-HTFET reduces the SOM computing cell to just a single transistor. This is remarkable considering that a conventional digital SOM cell will require more than 100 transistors. IDS-VGS variance of SO-HTFET is modulated by varying its drain-to-source voltage (VDS). This enables dynamic adaptation of distance measures in SO-HTFET-based SOM. Various SOM-modules are combined in a LAMSTAR network with link weights to facilitate deep learning and integration of various features of the applied pattern in a decision making process. Electroencephalogram (EEG) classification is studied using SO-HTFET-based LAMSTAR. SO-HTFET enables a higher number of hidden neurons in LAMSTAR by reducing the complexity of SOM and thereby, improves classification accuracy than a conventional design. EEG classification accuracy is specifically evaluated for fixed neuron and dynamic neuron approaches. The optimal variance of SO-HTFET IDS-VGSis extracted for these approaches.
Susmita Dey Manasi, Amit Ranjan Trivedi
IJCNN1