Sumit K. Mandal

dblp:229/4302 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-9294-1603ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 7 first-author · 19 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Accurate Analytical Modeling for NoCs with Hybrid Arbitration under High Traffic Injection
abstract
Analytical performance modeling of Networks-on-Chip (NoC) are important for fast design space exploration and quick pre-silicon evaluation. Existing NoC performance analysis techniques assume certain micro-architectural details (e.g., a particular arbitration technique) to be homogeneous across the entire NoC. However, emerging NoC architectures may have hybrid arbitration across the NoC to ensure high throughput. Moreover, existing analytical models estimating performance of NoCs with finite buffers fail to analyze the performance of the NoC accurately under high traffic injection which occur in several modern-day server as well as client applications. In this work, we propose a performance analysis technique for NoCs with hybrid arbitration under high traffic injection. We propose a novel transformation to accurately compute the waiting time of the queues under hybrid arbitration. We also develop a technique to compute the effective arrival statistics to the queues when the desired injection rate is high. Thorough experimental evaluation with a wide range of injection rates at the queues of an industrial NoC show that our proposed analytical model incurs only 7% error on average and 3 orders of speed-up with respect to cycle-accurate simulation under high traffic injection.
Rahul Tripathy, Mohammad Majharul Islam, Riad Akram, Raid Ayoub, Sumit K. Mandal
DATE5
2026 FARE: A Fine-grained Pipelined Reconfigurable FlashAttention Kernel
abstract
Transformer algorithm-based large machine learning (ML) models are widely used in various application domains [3]. We observe that the attention operation of a transformer algorithm consumes a significant amount of execution time when executed on a GPU, affecting real-time latency for inference. While other major operations in the transformer algorithm (e.g., feed-forward operation) consist of weight stationary matrix multiplications, attention consists of non-linear operations and a sequence of matrix multiplications. This sequence of operations in attention results in memory access bottlenecks in GPUs, directly affecting the latency. While there exists an implementation of an attention kernel to minimize memory access bottlenecks -- FlashAttention [1] -- it lacks reconfigurability. Moreover, GPUs provide limited knobs for customizing the attention datapath. There exist application-specific accelerators for ML models with transformer algorithms, but they too lack reconfigurability [2]. FPGAs, in contrast, enable application-specific dataflows, giving us full control over the design. Therefore, in this work, we propose a fine-grained pipelined reconfigurable flash-attention kernel targeting FPGA platforms. The fine-grained pipelining allows us to execute different stages of the attention mechanism in parallel. It carefully constructs different stages of a flash-attention kernel to reduce resource usage as well as improve the performance. As ML models scale and energy demands increase, this makes FPGA-based attention the sustainable choice. Experimental evaluation on Xilinx Zedboard (Zynq 7000) with different parameters of the attention kernel shows that our proposed design improves the performance by up to 2.9× while reducing the resource usage by up to 16.50× over state-of-the-art [4].
Kaushikkumar S. Rathva, Aakarsh Alam, Srini Srinivasan, Sumit K. Mandal
FPGA4
2026 Demo Abstract: NETRA - Energy-efficient Navigation at Edge towards Real-time Assistance for Visually Impaired Individuals
Meenakshi Atkade, Sumit K. Mandal
ISLPED2
2026 How Many Shots Does It Take? A Noise-Aware Quantum Resource Allocation Framework
abstract
Any algorithm execution on quantum computers requires several repeated and costly executions (known as shots) to obtain reliable results. In this work, we propose a closed-form accurate analytical expression to determine optimal number of shots required for reliable execution of any algorithm on a quantum computer. We also present a theoretically grounded technique to distribute fixed shot budget across different partitions in a quantum circuit minimizing the total error. Our proposed analytical model helps to reduce the shots associated with reliable execution of quantum algorithms by about 58\% compared to current practice, in turn reducing the energy consumption by upto 62\%. Furthermore, our proposed optimal shot allocation technique across different partitions reduces total error by up to 73\% compared to conventional approaches.
Prateek P. Kulkarni, Sumit K. Mandal
ISLPED2
2026 FlexNoC: Fast and Flexible Analysis for NoCs with Arbitrary Topologies and Hybrid Arbitration
abstract
Performance analysis of Network-on-Chips(NoC) plays a crucial role in design space exploration of SoCs, but traditional cycle-accurate NoC simulation often limits the ability to explore a large design space efficiently due to their notoriously slow execution. There exist several lightweight performance analysis techniques to reduce the design-space exploration time for NoCs, but all of them lack flexibility. In this work, we present FlexNoC - an end-to-end fast and flexible NoC performance analysis framework based on analytical modeling grounded on queuing theory. FlexNoC considers NoCs with irregular topologies and hybrid arbitration which no existing NoC performance analysis framework considers. We establish a Domain-Specific Language (DSL) to describe an NoC with any topology. Specifically, we extend the DOT language using ANTLR-based grammar to support custom NoC primitives such as injectors, queues, servers, arbiters, sinks, and splits. The DSL enables user-defined network components and their interconnections. The queuing theory based analytical model which is the backbone of the framework incorporates hybrid arbitration, along with finite buffers to accurately capture complex interactions between queues present in any given NoC. FlexNoC is accurate in NoC performance estimation and three orders of magnitude faster than cycle accurate NoC simulation - offering an efficient platform for rapid design space exploration and early-stage NoC performance optimization. Moreover, we demonstrate that FlexNoC, through rapid design space exploration, unlocks new insights regarding arbitration techniques at router ports.
Anuparna Ganguly, Rahul Tripathy, Tyson Loveless, Mohammad Majharul Islam, Sumit K. Mandal
ISPASS5
2026 Characterizing and Accelerating Spacecraft Onboard Workloads on RISC-V Platform
abstract
This paper studies 28 production spacecraft workloads from ISRO (Indian Space Research Organisation) missions to find out which computations actually limit performance on resource-constrained space processors. Spacecraft onboard processors face strict power and thermal constraints while performing critical real-time calculations for navigation, control, and data processing. Using a Basic Computational Element (BCE) breakdown, we find that software trigonometric calls (up to 36% of execution time) and small matrix operations (13%–41%) are the main bottlenecks across navigation, attitude control, guidance, and remote-sensing workloads. To check whether these findings translate into useful hardware, we build two domainspecific accelerators—a CORDIC (COordinate Rotation DIgital Computer) trigonometric unit and a 3 × 3 systolic matrix accelerator—integrated into a RISC-V processor through custom instruction set extensions (RVTrig, RVMatrix) and evaluated using cycle-approximate gem5 simulation. Across 18 representative workloads, the combined accelerators give a 3.13× geometric mean speedup and 37.2% energy reduction, reaching up to 9.29× on the most accelerator-friendly workload. Section VI introduces a timestep-based roofline model that tracks algorithmic progress rather than raw operations, so that an accelerated system no longer looks slower under the standard FLOPs-based view.
Boul Chandra Garai, Sumit K. Mandal, K. R. Yogesh Prasad, R. Govindarajan
IEEE Trans. Computers2
2025 Interconnect Performance Estimation for ML Accelerators via Lightweight Analytical Model
abstract
Machine learning (ML) algorithms are being used in real time in various applications now-a-days. Since state-of-theart high performing ML algorithms are computationally intensive, there exist accelerators which reduce latency and/or improve energy-efficiency of computer systems executing ML algorithms. Performance estimation frameworks have been constructed to perform extensive design space exploration while designing these ML data flow accelerators. While the performance estimation of compute elements of data flow accelerators consists of high-level models, the performance estimation of communication elements of the accelerators still relies on cycle accurate simulations. However, cycle accurate simulations of the interconnect (communication elements) are notoriously slow and it increases the execution time of the performance estimation frameworks, hindering design space exploration. Existing analytical model-based interconnect performance estimation techniques are not applicable for the interconnects of data flow accelerators since they exhibit specific communication patterns. Therefore, we develop an end-to-end framework to estimate interconnect performance for ML data flow accelerators. Extensive experimental evaluations on different ML algorithms show that our proposed framework estimates the interconnect performance with less than 5% error.
Rahul Tripathy, Sumit K. Mandal
ISPASS2
2025 HISIM: Analytical Performance Modeling and Design Space Exploration of 2.5D/3D Integration for AI Computing
abstract
Monolithic designs face significant fabrication cost and data movement challenges, especially when executing complex and diverse AI models. Advanced 2.5D/3D packaging promises high bandwidth and connection density to overcome these challenges, yet it also introduces new electro-thermal constraints. This article develops a suite of analytical performance models to enable efficient benchmarking of a 2.5D/3D heterogeneous system for energy-efficient AI computing. These models encompass various performance metrics related to computing units, network-on-chip (NoC), and network-on-package (NoP). The results are summarized into a new tool, HISIM, which is$10^{4} \times $–$10^{6} \times $faster than state-of-the-art AI benchmark tools. Furthermore, HISIM integrates rapid thermal simulation for the 2.5D/3D system, helping shed light on both the potential and limitations of 2.5D/3D heterogeneous integration (HI) on representative AI algorithms. The code of HISIM is available athttps://github.com/mec-UMN/HISIM.
Zhenyu Wang 0016, Pragnya Sudershan Nalla, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Jae-sun Seo, Vidya A. Chhabria, Jeff Zhang 0001, Chaitali Chakrabarti, Ümit Y. Ogras, Yu Cao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Exploiting 2.5D/3D Heterogeneous Integration for AI Computing
abstract
The evolution of AI algorithms has not only revolutionized many application domains, but also posed tremendous challenges on the hardware platform. Advanced packaging technology today, such as 2.5D and 3D interconnection, provides a promising solution to meet the ever-increasing demands of bandwidth, data movement, and system scale in AI computing. This work presents HISIM, a modeling and benchmarking tool for chiplet-based heterogeneous integration. HISIM emphasizes the hierarchical interconnection that connects various chiplets through network-on-package. It further integrates technology roadmap, power/latency prediction, and thermal analysis together to support electro-thermal co-design. Leveraging HISIM with in-memory computing chiplets, we explore the advantages and limitations of 2.5D and 3D heterogenous integration on representative AI algorithms, such as DNNs, transformers, and graph neural networks.
Zhenyu Wang 0016, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Yaotian Liu, Jae-sun Seo, Chaitali Chakrabarti, Ümit Y. Ogras, Vidya A. Chhabria, Jeff Zhang 0001, Yu Cao 0001
ASPDAC4
2024 DISHA: Low-Energy Sparse Transformer at Edge for Outdoor Navigation for the Visually Impaired Individuals
abstract
Assistive technology for visually impaired individuals is extremely useful to make them independent of another human being in performing day-to-day chores and instill confidence in them. One of the important aspects of assistive technology is outdoor navigation for visually impaired people. While there exist several techniques for outdoor navigation in the literature, they are mainly limited to obstacle detection. However, navigating a visually impaired person through the sidewalk (while the person is walking outside) is important too. Moreover, the assistive technology should ensure low-energy operation to extend the battery life of the device. Therefore, in this work, we propose an end-to-end technology deployed on an edge device to assist visually impaired people. Specifically, we propose a novel pruning technique for transformer algorithm which detects sidewalk. The pruning technique ensures low latency of execution and low energy consumption when the pruned transformer algorithm is deployed on the edge device. Extensive experimental evaluation shows that our proposed technology provides up to 32.49% improvement in accuracy and 1.4 hours of extension in battery life with respect to a baseline technique.
Praveen Nagil, Sumit K. Mandal
ISLPED2
2023 A Lightweight Congestion Control Technique for NoCs with Deflection Routing
abstract
Network-on-Chip (NoC) congestion builds up during heavy traffic load and leads to wasted link bandwidth, crippling the system performance. We propose a lightweight machine learning-based technique that helps predict congestion in the net-work by collecting features related to traffic at each destination and labelling it using a novel time reversal approach. The labelled data is used to design a low overhead and an explainable decision tree model used at runtime congestion control. Experimental evaluations with synthetic and real traffic on industrial$\boldsymbol{6\times 6}$NoC show that the proposed approach increases fairness and memory read bandwidth by up to 114% with respect to existing congestion control technique while incurring less than 0.01% of overhead.
Shruti Yadav Narayana, Sumit K. Mandal, Raid Ayoub, Michael Kishinevsky, Ümit Y. Ogras
DATE2
2023 Achieving Datacenter-scale Performance through Chiplet-based Manycore Architectures
abstract
Chiplet-based 2.5D systems that integrate multiple smaller chips on a single die are gaining popularity for executing both compute-and data-intensive applications. While smaller chips (chiplets) reduce fabrication costs, they also provide less functionality. Hence, manufacturing several smaller chiplets and combining them into a single system enables the functionality of a larger monolithic chip without prohibitive fabrication costs. The chiplets are connected through the network-on-interposer (NoP). Designing a high-performance and energy-efficient NoP architecture is essential as it enables large-scale chiplet integration. This paper highlights the challenges and existing solutions for designing suitable NoP architectures targeted for 2.5D systems catered to datacenter-scale applications. We also highlight the future research challenges stemming from the current state-of-the-art to make the NoP-based 2.5D systems widely applicable.
Sumit K. Mandal, Janardhan Rao Doppa, Ümit Y. Ogras, Partha Pratim Pande
DATE2
2023 Fast Performance Analysis for NoCs With Weighted Round-Robin Arbitration and Finite Buffers
abstract
Weighted round-robin (WRR) arbitration provides global fairness in networks-on-chip (NoCs) as opposed to the commonly used round-robin and priority-based arbitration techniques. However, the large number of weights explodes the design space and exacerbates performance (latency-throughput) tuning. Therefore, fast and accurate performance analysis techniques for NoCs are crucial for accelerating design space exploration and accurate pre-silicon evaluation. This article presents the first comprehensive performance analysis technique for NoCs with WRR arbitration and finite buffers. It can handle bursty traffic and is scalable to large NoC sizes. The proposed technique first estimates the probability that a queue is full and uses this result to compute the modified service time and queuing delay. Thorough experimental evaluations with synthetic traffic and real applications show that the proposed analytical model is always more than 10% accurate compared to cycle-accurate simulations. Moreover, the proposed performance analysis technique is five orders of magnitude faster than cycle-accurate simulations for a$16\times16$mesh NoC.
Sumit K. Mandal, Shruti Yadav Narayana, Raid Ayoub, Michael Kishinevsky, Ahmed Abousamra, Ümit Y. Ogras
IEEE Trans. Very Large Scale Integr. Syst.1
2022 Big-Little Chiplets for In-Memory Acceleration of DNNs: A Scalable Heterogeneous Architecture
abstract
Monolithic in-memory computing (IMC) architectures face significant yield and fabrication cost challenges as the complexity of DNNs increases. Chiplet-based IMCs that integrate multiple dies with advanced 2.5D/3D packaging offers a low-cost and scalable solution. They enable heterogeneous architectures where the chiplets and their associated interconnection can be tailored to the non-uniform algorithmic structures to maximize IMC utilization and reduce energy consumption. This paper proposes a heterogeneous IMC architecture with big-little chiplets and a hybrid network-on-package (NoP) to optimize the utilization, interconnect bandwidth, and energy efficiency. For a given DNN, we develop a custom methodology to map the model onto the big-little architecture such that the early layers in the DNN are mapped to the little chiplets with higher NoP bandwidth and the subsequent layers are mapped to the big chiplets with lower NoP bandwidth. Furthermore, we achieve a scalable solution by incorporating a DRAM into each chiplet to support a wide range of DNNs beyond the area limit. Compared to a homogeneous chiplet-based IMC architecture, the proposed big-little architecture achieves up to 329× improvement in the energy-delay-area product (EDAP) and up to 2× higher IMC utilization. Experimental evaluation of the proposed big-little chiplet-based RRAM IMC architecture for ResNet-50 on ImageNet shows 259×, 139×, and 48× improvement in energy-efficiency at lower area compared to Nvidia V100 GPU, Nvidia T4 GPU, and SIMBA architecture, respectively.
A. Alper Goksoy, Sumit K. Mandal, Zhenyu Wang 0016, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001
ICCAD3
2022 Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous Processor
abstract
RF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design.
Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue
ISCAS31
2022 Impact of On-chip Interconnect on In-memory Acceleration of Deep Neural Networks
abstract
With the widespread use of Deep Neural Networks (DNNs), machine learning algorithms have evolved in two diverse directions—one with ever-increasing connection density for better accuracy and the other with more compact sizing for energy efficiency. The increase in connection density increases on-chip data movement, which makes efficient on-chip communication a critical function of the DNN accelerator. The contribution of this work is threefold. First, we illustrate that the point-to-point (P2P)-based interconnect is incapable of handling a high volume of on-chip data movement for DNNs. Second, we evaluate P2P and network-on-chip (NoC) interconnect (with a regular topology such as a mesh) for SRAM- and ReRAM-based in-memory computing (IMC) architectures for a range of DNNs. This analysis shows the necessity for the optimal interconnect choice for an IMC DNN accelerator. Finally, we perform an experimental evaluation for different DNNs to empirically obtain the performance of the IMC architecture with both NoC-tree and NoC-mesh. We conclude that, at the tile level, NoC-tree is appropriate for compact DNNs employed at the edge, and NoC-mesh is necessary to accelerate DNNs with high connection density. Furthermore, we propose a technique to determine the optimal choice of interconnect for any given DNN. In this technique, we use analytical models of NoC to evaluate end-to-end communication latency of any given DNN. We demonstrate that the interconnect optimization in the IMC architecture results in up to 6 × improvement in energy-delay-area product for VGG-19 inference compared to the state-of-the-art ReRAM-based IMC architectures.
Sumit K. Mandal, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001
ACM J. Emerg. Technol. Comput. Syst.2
2022 SWAP: A Server-Scale Communication-Aware Chiplet-Based Manycore PIM Accelerator
abstract
Processing-in-memory (PIM) is a promising technique to accelerate deep learning (DL) workloads. Emerging DL workloads (e.g., ResNet with 152 layers) consist of millions of parameters, which increase the area and fabrication cost of monolithic PIM accelerators. The fabrication cost challenge can be addressed by 2.5-D systems integrating multiple PIM chiplets connected through a network-on-package (NoP). However, server-scale scenarios simultaneously execute multiple compute-heavy DL workloads, leading to significant interchiplet data volume. State-of-the-art NoP architectures proposed in the literature do not consider the nature of DL workloads. In this article, we propose a novel server scale 2.5-D manycore architecture called SWAP that accounts for the traffic characteristics of DL applications. Comprehensive experimental evaluations with different system sizes as well as diverse emerging DL workloads demonstrate that SWAP achieves significant performance and energy consumption improvements with much lower fabrication cost than state-of-the-art NoP topologies.
Sumit K. Mandal, Janardhan Rao Doppa, Ümit Y. Ogras, Partha Pratim Pande
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Theoretical Analysis and Evaluation of NoCs with Weighted Round-Robin Arbitration
abstract
Fast and accurate performance analysis techniques are essential in early design space exploration and pre-silicon evaluations, including software eco-system development. In particular, on-chip communication continues to play an increasingly important role as the many-core processors scale up. This paper presents the first performance analysis technique that targets networks-on-chip (NoCs) that employ weighted round-robin (WRR) arbitration. Besides fairness, WRR arbitration provides flexibility in allocating bandwidth proportionally to the importance of the traffic classes, unlike basic round-robin and priority-based arbitration. The proposed approach first estimates the effective service time of the packets in the queue due to WRR arbitration. Then, it uses the effective service time to compute the average waiting time of the packets. Next, we incorporate a decomposition technique to extend the analytical model to handle NoC of any size. The proposed approach achieves less than 5% error while executing real applications and 10% error under challenging synthetic traffic with different burstiness levels.
Sumit K. Mandal, Jie Tong, Raid Ayoub, Michael Kishinevsky, Ahmed Abousamra, Ümit Y. Ogras
ICCAD1
2021 SIAM: Chiplet-based Scalable In-Memory Acceleration with Mesh for Deep Neural Networks
abstract
In-memory computing (IMC) on a monolithic chip for deep learning faces dramatic challenges on area, yield, and on-chip interconnection cost due to the ever-increasing model sizes. 2.5D integration or chiplet-based architectures interconnect multiple small chips (i.e., chiplets) to form a large computing system, presenting a feasible solution beyond a monolithic IMC architecture to accelerate large deep learning models. This paper presents a new benchmarking simulator, SIAM, to evaluate the performance of chiplet-based IMC architectures and explore the potential of such a paradigm shift in IMC architecture design. SIAM integrates device, circuit, architecture, network-on-chip (NoC), network-on-package (NoP), and DRAM access models to realize an end-to-end system. SIAM is scalable in its support of a wide range of deep neural networks (DNNs), customizable to various network structures and configurations, and capable of efficient design space exploration. We demonstrate the flexibility, scalability, and simulation speed of SIAM by benchmarking different state-of-the-art DNNs with CIFAR-10, CIFAR-100, and ImageNet datasets. We further calibrate the simulation results with a published silicon result, SIMBA. The chiplet-based IMC architecture obtained through SIAM shows 130 and 72 improvement in energy-efficiency for ResNet-50 on the ImageNet dataset compared to Nvidia V100 and T4 GPUs.
Sumit K. Mandal, Manvitha Pannala, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001
ACM Trans. Embed. Comput. Syst.2
2021 FLASH: Fast Neural Architecture Search with Hardware Optimization
abstract
Neural architecture search (NAS) is a promising technique to design efficient and high-performance deep neural networks (DNNs). As the performance requirements of ML applications grow continuously, the hardware accelerators start playing a central role in DNN design. This trend makes NAS even more complicated and time-consuming for most real applications. This paper proposes FLASH, a very fast NAS methodology that co-optimizes the DNN accuracy and performance on a real hardware platform. As the main theoretical contribution, we first propose the NN-Degree, an analytical metric to quantify the topological characteristics of DNNs with skip connections (e.g., DenseNets, ResNets, Wide-ResNets, and MobileNets). The newly proposed NN-Degree allows us to do training-free NAS within one second and build an accuracy predictor by training as few as 25 samples out of a vast search space with more than 63 billion configurations. Second, by performing inference on the target hardware, we fine-tune and validate our analytical models to estimate the latency, area, and energy consumption of various DNN architectures while executing standard ML datasets. Third, we construct a hierarchical algorithm based on simplicial homology global optimization (SHGO) to optimize the model-architecture co-design process, while considering the area, latency, and energy consumption of the target hardware. We demonstrate that, compared to the state-of-the-art NAS approaches, our proposed hierarchical SHGO-based algorithm enables more than four orders of magnitude speedup (specifically, the execution time of the proposed algorithm is about 0.1 seconds). Finally, our experimental evaluations show that FLASH is easily transferable to different hardware architectures, thus enabling us to do NAS on a Raspberry Pi-3B processor in less than 3 seconds.
Guihong Li, Sumit K. Mandal, Ümit Y. Ogras, Radu Marculescu
ACM Trans. Embed. Comput. Syst.2
2021 Front-End Architecture Design for Low-Complexity 3-D Ultrasound Imaging Based on Synthetic Aperture Sequential Beamforming
abstract
The 3-D ultrasound imaging provides distinct advantages over its 2-D counterpart leading to a more accurate analysis of tumors and cysts. However, the front end of a 3-D system must receive and process data at prodigious rates, making it impractical for power-constrained portable systems. Synthetic aperture sequential beamforming (SASB) is an ultrasound beamforming technique that splits the computation into two stages, such that the computation in Stage 1 can be completed in the power-constrained front end while the remaining computation can be done elsewhere. In this article, we present several algorithmic and architectural techniques to enable efficient computation of Stage 1 processing without compromising imaging quality. Specifically, we present algorithmic techniques that reduce the computational complexity in Stage 1 by 17× through a systematic reduction in the number of apodization coefficients. We propose a 3-D die stacked architecture where the signals received by 961 active transducers are digitized, routed by a network-onchip, and processed in parallel. This architecture does not require the explicit storage of incoming data samples. We synthesize the architecture using TSMC 28-nm technology node. The front-end power consumption is around 1.5 W, making it suitable for portable applications.
Jian Zhou 0012, Sumit K. Mandal, Brendan L. West, Siyuan Wei, Ümit Y. Ogras, Oliver Kripfgans, J. Brian Fowlkes, Thomas F. Wenisch, Chaitali Chakrabarti
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Online Adaptive Learning for Runtime Resource Management of Heterogeneous SoCs
abstract
Dynamic resource management has become one of the major areas of research in modern computer and communication system design due to lower power consumption and higher performance demands. The number of integrated cores, level of heterogeneity and amount of control knobs increase steadily. As a result, the system complexity is increasing faster than our ability to optimize and dynamically manage the resources. Moreover, offline approaches are sub-optimal due to workload variations and large volume of new applications unknown at design time. This paper first reviews recent online learning techniques for predicting system performance, power, and temperature. Then, we describe the use of predictive models for online control using two modern approaches: imitation learning (IL) and an explicit nonlinear model predictive control (NMPC). Evaluations on a commercial mobile platform with 16 benchmarks show that the IL approach successfully adapts the control policy to unknown applications. The explicit NMPC provides 25% energy savings compared to a state-of-the-art algorithm for multi-variable power management of modern GPU sub-systems.
Sumit K. Mandal, Ümit Y. Ogras, Janardhan Rao Doppa, Raid Ayoub, Michael Kishinevsky, Partha Pratim Pande
DAC1
2020 Performance Analysis of Priority-Aware NoCs with Deflection Routing under Traffic Congestion
abstract
Priority-aware networks-on-chip (NoCs) are used in industry to achieve predictable latency under different workload conditions. These NoCs incorporate deflection routing to minimize queuing resources within routers and achieve low latency during low traffic load. However, deflected packets can exacerbate congestion during high traffic load since they consume the NoC bandwidth. State-of-the-art analytical models for priority-aware NoCs ignore deflected traffic despite its significant latency impact during congestion. This paper proposes a novel analytical approach to estimate end-to-end latency of priority-aware NoCs with deflection routing under bursty and heavy traffic scenarios. Experimental evaluations show that the proposed technique outperforms alternative approaches and estimates the average latency for real applications with less than 8% error compared to cycle-accurate simulations.
Sumit K. Mandal, Anish Krishnakumar, Raid Ayoub, Michael Kishinevsky, Ümit Y. Ogras
ICCAD1
2020 Runtime Task Scheduling Using Imitation Learning for Heterogeneous Many-Core Systems
abstract
Domain-specific systems-on-chip, a class of heterogeneous many-core systems, is recognized as a key approach to narrow down the performance and energy-efficiency gap between custom hardware accelerators and programmable processors. Reaching the full potential of these architectures depends critically on optimally scheduling the applications to available resources at runtime. Existing optimization-based techniques cannot achieve this objective at runtime due to the combinatorial nature of the task scheduling problem. As the main theoretical contribution, this article poses scheduling as a classification problem and proposes a hierarchical imitation learning (IL)-based scheduler that learns from an Oracle to maximize the performance of multiple domain-specific applications. Extensive evaluations with six streaming applications from wireless communications and radar domains show that the proposed IL-based scheduler approximates an offline Oracle policy with more than 99% accuracy for performance- and energy-based optimization objectives. Furthermore, it achieves almost identical performance to the Oracle with a low runtime overhead and successfully adapts to new applications, many-core system configurations, and runtime variations in application characteristics.
Anish Krishnakumar, Samet E. Arda, A. Alper Goksoy, Sumit K. Mandal, Ümit Y. Ogras, Anderson Luiz Sartor, Radu Marculescu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 An Energy-aware Online Learning Framework for Resource Management in Heterogeneous Platforms
abstract
Mobile platforms must satisfy the contradictory requirements of fast response time and minimum energy consumption as a function of dynamically changing applications. To address this need, systems-on-chip (SoC) that are at the heart of these devices provide a variety of control knobs, such as the number of active cores and their voltage/frequency levels. Controlling these knobs optimally at runtime is challenging for two reasons. First, the large configuration space prohibits exhaustive solutions. Second, control policies designed offline are at best sub-optimal, since many potential new applications are unknown at design-time. We address these challenges by proposing an online imitation learning approach. Our key idea is to construct an offline policy and adapt it online to new applications to optimize a given metric (e.g., energy). The proposed methodology leverages the supervision enabled by power-performance models learned at runtime. We demonstrate its effectiveness on a commercial mobile platform with 16 diverse benchmarks. Our approach successfully adapts the control policy to an unknown application after executing less than 25% of its instructions.
Sumit K. Mandal, Ganapati Bhat, Janardhan Rao Doppa, Partha Pratim Pande, Ümit Y. Ogras
ACM Trans. Design Autom. Electr. Syst.1
2019 Analytical Performance Models for NoCs with Multiple Priority Traffic Classes
abstract
Networks-on-chip (NoCs) have become the standard for interconnect solutions in industrial designs ranging from client CPUs to many-core chip-multiprocessors. Since NoCs play a vital role in system performance and power consumption, pre-silicon evaluation environments include cycle-accurate NoC simulators. Long simulations increase the execution time of evaluation frameworks, which are already notoriously slow, and prohibit design-space exploration. Existing analytical NoC models, which assume fair arbitration, cannot replace these simulations since industrial NoCs typically employ priority schedulers and multiple priority classes. To address this limitation, we propose a systematic approach to construct priority-aware analytical performance models using micro-architecture specifications and input traffic. Our approach decomposes the given NoC into individual queues with modified service time to enable accurate and scalable latency computations. Specifically, we introduce novel transformations along with an algorithm that iteratively applies these transformations to decompose the queuing system. Experimental evaluations using real architectures and applications show high accuracy of 97% and up to 2.5× speedup in full-system simulation.
Sumit K. Mandal, Raid Ayoub, Michael Kishinevsky, Ümit Y. Ogras
ACM Trans. Embed. Comput. Syst.1
2019 Dynamic Resource Management of Heterogeneous Mobile Platforms via Imitation Learning
abstract
The complexity of heterogeneous mobile platforms is growing at a rate faster than our ability to manage them optimally at runtime. For example, state-of-the-art systems-on-chip (SoCs) enable controlling the type (Big/Little), number, and frequency of active cores. Managing these platforms becomes challenging with the increase in the type, number, and supported frequency levels of the cores. However, existing solutions used in mobile platforms still rely on simple heuristics based on the utilization of cores. This paper presents a novel and practical imitation learning (IL) framework for dynamically controlling the type (Big/Little), number, and the frequencies of active cores in heterogeneous mobile processors. We present efficient approaches for constructing an Oracle policy to optimize different objective functions, such as energy and performance per Watt (PPW). The Oracle policies enable us to design low-overhead power management policies that achieve near-optimal performance matching the Oracle. Experiments on a commercial platform with 19 benchmarks show on an average 101% PPW improvement compared to the default interactive governor.
Sumit K. Mandal, Ganapati Bhat, Chetan Arvind Patil, Janardhan Rao Doppa, Partha Pratim Pande, Ümit Y. Ogras
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Online learning for adaptive optimization of heterogeneous SoCs
abstract
Energy efficiency and performance of heterogeneous multiprocessor systems-on-chip (SoC) depend critically on utilizing a diverse set of processing elements and managing their power states dynamically. Dynamic resource management techniques typically rely on power consumption and performance models to assess the impact of dynamic decisions. Despite the importance of these decisions, many existing approaches rely on fixed power and performance models learned offline. This paper presents an online learning framework to construct adaptive analytical models. We illustrate this framework for modeling GPU frame processing time, GPU power consumption and SoC power-temperature dynamics. Experiments on Intel Atom E3826, Qualcomm Snapdragon 810, and Samsung Exynos 5422 SoCs demonstrate that the proposed approach achieves less than 6% error under dynamically varying workloads.
Ganapati Bhat, Sumit K. Mandal, Ujjwal Gupta, Ümit Y. Ogras
ICCAD2