Bonita Bhaskaran

dblp:66/754 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Revisiting Microelectronics Resilience and Reliability in the Era of AI
abstract
The proliferation of resource-intensive large language model (LLM) workloads has driven interest in consistency and reliability of their performance on mobile SOCs. Prior work on latency prediction for Deep Neural Network (DNN) execution has primarily examined training predictors for CNNs for each chip under test and requires re-training or fine-tuning to adapt to new hardware or account for variations in performance between chips. To fill this gap, we present a latency prediction framework for LLMs on mobile GPUs that allows easy adaptation to performance variation in new devices and variability-aware estimation of LLM workload latency without requiring extensive access to low-level compilation features such as Instruction-Set Architecture (ISA) information. Our approach has been tested on multiple mobile devices, achieving sub-15% error in latency prediction for multiple LLMs.
Arjun Chaudhuri, Bonita Bhaskaran
VTS2
2022 Observation Point Insertion Using Deep Learning
abstract
Silent Data Corruption (SDC) is one of the critical problems in the field of testing, where errors or corruption do not manifest externally. As a result, there is increased focus on improving the outgoing quality of dies by striving for better correlation between structural and functional patterns to achieve a low DPPM. This is very important for NVIDIA's chips due to the various markets we target; for example, automotive and data center markets have stringent in-field testing requirements. One aspect of these efforts is to also target better testability while incurring lower test cost. Since structural testing is faster than functional tests, it is important to make these structural test patterns as effective as possible and free of test escapes. However, with the rising cell count in today's digital circuits, it is becoming increasingly difficult to sensitize faults and propagate the fault effects to scan-flops or primary outputs. Hence, methods to insert observation points to facilitate the detection of hard-to-detect (HtD) faults are being increasingly explored. In this work, we propose an Observation Point Insertion (OPI) scheme using deep learning with the motivation of achieving - 1) better quality test points than commercial EDA tools leading to a potential lower pattern count 2) faster turnaround time to generate the test points. In order to achieve better pattern compaction than commercial EDA tools, we employ Graph Convolutional Networks (GCNs) to learn the topology of logic circuits along with the features that influence its testability. The graph structures are subsequently used to train two GCN-type deep learning models - the first model predicts signal probabilities at different nets and the second model uses these signal probabilities along with other features to predict the reduction in test-pattern count when OPs are inserted at different locations in the design. The features we consider include structural features like gate type, gate logic, reconvergent-fanouts and testability features like SCOAP. Our simulation results indicate that the proposed machine learning models can predict the probabilistic testability metrics with reasonable accuracy and can identify observation points that reduce pattern count.
Bonita Bhaskaran, Sanmitra Banerjee, Kaushik Narayanun, Shao-Chun Hung, Seyed Nima Mozaffari, Tung-Che Liang
ICCAD1
2022 On-Die Noise Measurement During Automatic Test Equipment (ATE) Testing and In-System-Test (IST)
abstract
It is realized that having a method/apparatus that accurately measures voltage noise is imperative for ATE and SLT testing to 1) reliably sign-off on production patterns; 2) effectively optimize low power settings of scan architecture; 3) screen for defects during ATE testing and apply structural patterns on SLT at desired Voltage/Frequency points; This requires optimized noise profiles and hence localized noise monitors for appropriate tuning. In addition, with NVIDIA’s chips foraying into the automotive space, functional safety has gained utmost priority, and the noise profile during In-System Test (IST) helps catch reliability and aging-related defects in the field that show up after stress and degradation. This paper proposes an enhancement to the in-system Noise Measurement macro (NMEAS) to record voltage noise during the application of structural DFT patterns, such as in ATE and SLT testing, which was not possible in the conventional noise measurement methods. The introduced technique utilizes a continuous free-running fast clock that feeds functional frequency to NMEAS during test which allows it to measure the voltage noise of the chip during both shift and capture phases. Also, a novel enable generation logic and a counter are introduced that allow for more precise characterization of the measured voltage noise data.
Seyed Nima Mozaffari, Bonita Bhaskaran, Shantanu Sarangi, Suhas M. Satheesh, Kuo Lin Fu, Nithin Valentine, P. Manikandan, Mahmut Yilmaz
VTS2
2019 An Efficient Supervised Learning Method to Predict Power Supply Noise During At-speed Test
abstract
The Power Distribution Network (PDN) is designed for worst-case power-hungry functional use-cases. Most often Design for Test (DFT) scenarios are not accounted for, while optimizing the PDN design. Automatic Test Pattern Generation (ATPG) tools typically follow a greedy algorithm to achieve maximum fault coverage with short test times. This causes Power Supply Noise (PSN) during scan testing to be much higher than functional mode since switching activity is higher by an order of magnitude. Understanding the noise characteristics through exhaustive pattern simulation is extremely machine and memory intensive and requires unsustainably long runtimes. Hence, we aggressively limit switching factors to conservative estimates and rely on post-silicon noise characterization to optimize test vectors. In this work, we propose a novel method to predict simultaneous switching noise using fast Deep Neural Networks (DNNs) such as Fully Connected Network, Convolutional Neural Network, and Natural Language Processing. Our approach, that is based on pre-silicon ATPG vectors, is significantly faster than conventional estimation methods and can potentially reduce the test time.
Seyed Nima Mozaffari, Bonita Bhaskaran, Kaushik Narayanun, Ayub Abdollahian, Vinod Pagalone, Shantanu Sarangi, Jonathon E. Colburn
ITC2
2019 A Novel Graph Coloring Based Solution for Low-Power Scan Shift
abstract
During scan shift, high simultaneous toggling of sequential logic on a System-on-Chip (SoC) can result in increased Power Supply Noise (PSN). The problem gets exacerbated when the switching logic is present in neighboring blocks on the SoC that share the same power rails. To solve this voltage noise problem, we propose a new graph coloring algorithm that assigns staggered shift-clocks to the SoC blocks such that (i) no two neighboring blocks use the same shift-clock (to reduce local hotspots), and (ii) the number of scan cells toggling per shift clock is equalized (to reduce global noise). The new algorithm takes into account the total number of scan flops per block, and the assignment of stagger clocks is done such that the total number of scan flops that toggle per staggered shift-clock is balanced at the power rail-level. Using silicon data from NVIDIA's recently taped-out chips, we show that the stagger assignment using our new algorithm results in at 70% PSN reduction compared to conventional scan shift and around 21% PSN reduction compared to the previously proposed stagger assignment solutions.
Saurabh Gupta 0005, Bonita Bhaskaran, Shantanu Sarangi, Ayub Abdollahian, Jennifer Dworak
VTS2
2019 Special Session: In-System-Test (IST) Architecture for NVIDIA Drive-AGX Platforms
abstract
Safety is one of the crucial features of autonomous drive platforms, and semiconductor chips used in these architectures must guarantee functional safety aspects mandated by ISO 26262 standard. To monitor the failures due to field defects, in-system-structural-tests are automatically run during key-on and/or key-off. Upon detection of any permanent defects by the in-system-test (IST) architecture, Drive platform responds to achieve the fail-safe state of the system. In this paper, we present the IST architecture that helps with achieving highest functional safety levels on the NVIDIA Drive platform.
Pavan Kumar Datla Jagannadha, Mahmut Yilmaz, Milind Sonawane, Sailendra Chadalavada, Shantanu Sarangi, Bonita Bhaskaran, Shashank Bajpai, Venkat Abilash Reddy Nerallapally, Jayesh Pandey, Sam Jiang
VTS6
2017 At-speed capture global noise reduction & low-power memory test architecture
abstract
Traditionally, DFT patterns exacerbate dynamic power consumption in large ASICs. At-speed scan and memory tests are sensitive to voltage droop and peak current because the power grid is designed for functional power viruses (maximum workload applications) whose power consumption is much lower than DFT patterns. Our goal in this work is to ensure that the quality of test is not compromised while power is constrained to be within sign-off power budgets. We present an IEEE 1500-compliant Global Low Power Capture (GLPC) architecture with minimized interconnects between sub-blocks. For memory tests, we also present an extension of the architecture, Low Power MBIST (LP-MBIST) which shuts down the toggling of logic flops. Experimental results for both architectures show appreciable dynamic power reduction on recently taped out 16nm ASIC chips.
Bonita Bhaskaran, Sailendra Chadalavada, Shantanu Sarangi, Nithin Valentine, Venkat Abilash Reddy Nerallapally, Ayub Abdollahian
VTS1
2016 Advanced test methodology for complex SoCs
abstract
This paper presents the latest test methodology for NVIDIA's multi-billion transistor Mobile System on Chip (SoC) and Graphics Processing Unit (GPU). The paper describes the innovations that enhance the SoC plug-n-play scheme in terms of DFT. It also demonstrates how the architecture enables ultra-low pin count testing together with test data reuse and efficient test scheduling to improve the test quality while lowering the test cost. We present a scalable scan interface methodology coupled with core isolation and advanced clocking design while keeping the overall power budget for test within the limits of SoC Thermal Design Power (TDP). Silicon results are shared to demonstrate the effectiveness of this architecture.
Pavan Kumar Datla Jagannadha, Mahmut Yilmaz, Milind Sonawane, Sailendra Chadalavada, Shantanu Sarangi, Bonita Bhaskaran, Ayub Abdollahian
ITC6
2016 Test method and scheme for low-power validation in modern SOC integrated circuits
abstract
Test Mode power can be 5X higher than functional power in GPUs, while the power grid is designed only for worst-case functional toggle. The large simultaneous switching noise induced on the power rails during at-speed capture testing is constrained by means of hardware solution. To determine the best low power mode for ATPG, we propose novel techniques to: estimate global peak current (di), determine local droop trend and validate and further optimize chosen power settings with exhaustive post-silicon power mode tuning. During Power Optimization (PO) phase, the measured clock frequency (fclk) and Vdroop are analyzed on every pattern and test coverage and pattern count are optimized for the production pattern set. We share correlation results and Power Supply Noise (PSN) distribution for the production pattern set on recent 28-nm GPUs.
Bonita Bhaskaran, Amit Sanghani, Kaushik Narayanun, Ayub Abdollahian, Amit Laknaur
VTS1
2016 A programmable method for low-power scan shift in SoC integrated circuits
abstract
We present a programmable method for shift-clock stagger assignment to reduce power supply noise during system-on-chip (SoC) testing. An SoC design is typically composed of several blocks and two neighboring blocks that share the same power rails should not be toggled at the same time during shift. Therefore, the proposed programmable method does not assign the same stagger value to neighboring blocks. The positions of all blocks are first analyzed and the shared boundary length between blocks is then calculated. Based on the position relationships between the blocks, a mathematical model is presented to derive optimal result for small-to-medium sized problems. For larger designs, a heuristic algorithm is proposed and evaluated. We present assignment results as well as power-analysis results and silicon data for industry designs to highlight the effectiveness of the proposed method.
Ran Wang 0002, Bonita Bhaskaran, Karthikeyan Natarajan, Ayub Abdollahian, Kaushik Narayanun, Krishnendu Chakrabarty, Amit Sanghani
VTS2
2007 DFT Techniques and Automation for Asynchronous NULL Conventional Logic Circuits
abstract
Conventional automatic test pattern generation (ATPG) algorithms fail when applied to asynchronous NULL convention logic (NCL) circuits due to the absence of a global clock and presence of more state-holding elements, leading to poor fault coverage. This paper presents a design-for-test (DFT) approach aimed at making asynchronous NCL designs testable using conventional ATPG programs. We propose an automatic DFT insertion flow (ADIF) methodology that performs scan and test point insertion on NCL designs to improve test coverage, using a custom ATPG library. Experimental results show significant increase in fault coverage for NCL cyclic and acyclic pipelined designs.
Venkat Satagopan, Bonita Bhaskaran, Waleed K. Al-Assadi, Scott C. Smith, Sindhu Kakarla
IEEE Trans. Very Large Scale Integr. Syst.2