Shantanu Sarangi

dblp:35/10050 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 4 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Next-Gen Scalable In-System-Test Architecture for Nvidia Automotive Platform
Sailendra Chadalavada, Milind Sonawane, Saranyan Sarangan, Alex Hsu, Pavan Javvaji, Shantanu Sarangi
VTS6
2024 A Scalable & Cost Efficient Next-Gen Scan Architecture: Streaming Scan Test via NVIDIA MATHS
abstract
Streaming scan test architectures can greatly optimize the test data delivery to large industrial designs. This paper discusses what happens when such architectures are combined with nearly unlimited data bandwidth provided by NVIDIA MATHS (Mechanism to Access Test-Data over High-Speed Link). There are multiple techniques for efficient use of scan bandwidth and its impact on the overall test cost and test quality. We have also architected various debug techniques for silicon bring-up. This scan architecture was designed for highest throughput to test multiple dies in parallel with lowest test power and best diagnosability.
Kunal Jain Mangilal, Mahmut Yilmaz, Vishal Agarwal, Shantanu Sarangi, Kaushik Narayanun
ITC4
2022 On-Die Noise Measurement During Automatic Test Equipment (ATE) Testing and In-System-Test (IST)
abstract
It is realized that having a method/apparatus that accurately measures voltage noise is imperative for ATE and SLT testing to 1) reliably sign-off on production patterns; 2) effectively optimize low power settings of scan architecture; 3) screen for defects during ATE testing and apply structural patterns on SLT at desired Voltage/Frequency points; This requires optimized noise profiles and hence localized noise monitors for appropriate tuning. In addition, with NVIDIA’s chips foraying into the automotive space, functional safety has gained utmost priority, and the noise profile during In-System Test (IST) helps catch reliability and aging-related defects in the field that show up after stress and degradation. This paper proposes an enhancement to the in-system Noise Measurement macro (NMEAS) to record voltage noise during the application of structural DFT patterns, such as in ATE and SLT testing, which was not possible in the conventional noise measurement methods. The introduced technique utilizes a continuous free-running fast clock that feeds functional frequency to NMEAS during test which allows it to measure the voltage noise of the chip during both shift and capture phases. Also, a novel enable generation logic and a counter are introduced that allow for more precise characterization of the measured voltage noise data.
Seyed Nima Mozaffari, Bonita Bhaskaran, Shantanu Sarangi, Suhas M. Satheesh, Kuo Lin Fu, Nithin Valentine, P. Manikandan, Mahmut Yilmaz
VTS3
2022 NVIDIA MATHS: Mechanism to Access Test-Data over High-Speed Links
abstract
MATHS (Mechanism to Access Test-Data over High-Speed Link) provides a high-throughput PCIe based system to structurally test system-on-chips (SOCs) at wafer and system-level. The system removes the need for expensive test equipment by eliminating the input/output pin (IO) requirements and memory per IO needs. It simplifies the ATE architecture and design to enable smaller form factors and reduce capital costs of ownership. MATHS enables eliminating the assembly test-insertion by directly testing the SOCs on system level platforms, further reducing the costs. Since the mechanism is based on PCIe standards, it is highly portable across all platforms including ATE, system-level test, board, and in-field testing.
Mahmut Yilmaz, Pavan Kumar Datla Jagannadha, Kaushik Narayanun, Shantanu Sarangi, Francisco Da Silva, Joe Sarmiento, Smbat Tonoyan, Ashwin Chintaluri, Animesh Khare, Milind Sonawane, Anitha Kalva, Alex Hsu, Jayesh Pandey
VTS4
2019 An Efficient Supervised Learning Method to Predict Power Supply Noise During At-speed Test
abstract
The Power Distribution Network (PDN) is designed for worst-case power-hungry functional use-cases. Most often Design for Test (DFT) scenarios are not accounted for, while optimizing the PDN design. Automatic Test Pattern Generation (ATPG) tools typically follow a greedy algorithm to achieve maximum fault coverage with short test times. This causes Power Supply Noise (PSN) during scan testing to be much higher than functional mode since switching activity is higher by an order of magnitude. Understanding the noise characteristics through exhaustive pattern simulation is extremely machine and memory intensive and requires unsustainably long runtimes. Hence, we aggressively limit switching factors to conservative estimates and rely on post-silicon noise characterization to optimize test vectors. In this work, we propose a novel method to predict simultaneous switching noise using fast Deep Neural Networks (DNNs) such as Fully Connected Network, Convolutional Neural Network, and Natural Language Processing. Our approach, that is based on pre-silicon ATPG vectors, is significantly faster than conventional estimation methods and can potentially reduce the test time.
Seyed Nima Mozaffari, Bonita Bhaskaran, Kaushik Narayanun, Ayub Abdollahian, Vinod Pagalone, Shantanu Sarangi, Jonathon E. Colburn
ITC6
2019 A Novel Graph Coloring Based Solution for Low-Power Scan Shift
abstract
During scan shift, high simultaneous toggling of sequential logic on a System-on-Chip (SoC) can result in increased Power Supply Noise (PSN). The problem gets exacerbated when the switching logic is present in neighboring blocks on the SoC that share the same power rails. To solve this voltage noise problem, we propose a new graph coloring algorithm that assigns staggered shift-clocks to the SoC blocks such that (i) no two neighboring blocks use the same shift-clock (to reduce local hotspots), and (ii) the number of scan cells toggling per shift clock is equalized (to reduce global noise). The new algorithm takes into account the total number of scan flops per block, and the assignment of stagger clocks is done such that the total number of scan flops that toggle per staggered shift-clock is balanced at the power rail-level. Using silicon data from NVIDIA's recently taped-out chips, we show that the stagger assignment using our new algorithm results in at 70% PSN reduction compared to conventional scan shift and around 21% PSN reduction compared to the previously proposed stagger assignment solutions.
Saurabh Gupta 0005, Bonita Bhaskaran, Shantanu Sarangi, Ayub Abdollahian, Jennifer Dworak
VTS3
2019 Special Session: In-System-Test (IST) Architecture for NVIDIA Drive-AGX Platforms
abstract
Safety is one of the crucial features of autonomous drive platforms, and semiconductor chips used in these architectures must guarantee functional safety aspects mandated by ISO 26262 standard. To monitor the failures due to field defects, in-system-structural-tests are automatically run during key-on and/or key-off. Upon detection of any permanent defects by the in-system-test (IST) architecture, Drive platform responds to achieve the fail-safe state of the system. In this paper, we present the IST architecture that helps with achieving highest functional safety levels on the NVIDIA Drive platform.
Pavan Kumar Datla Jagannadha, Mahmut Yilmaz, Milind Sonawane, Sailendra Chadalavada, Shantanu Sarangi, Bonita Bhaskaran, Shashank Bajpai, Venkat Abilash Reddy Nerallapally, Jayesh Pandey, Sam Jiang
VTS5
2019 Hybrid Performance Modeling for Optimization of In-System-Structural-Test (ISST) Latency
abstract
In-System-Test (IST) is one of the most advanced feature of autonomous drive platforms to monitor the semiconductor chip failures due to field defects. This is achieved by application of structural-tests (Memory BIST and/or Logic BIST) during key-on and/or key-off functional events. The time duration for application of these structural tests, detection of defects, drive platform reaction, and leading it to a fail-safe state must be short enough to avoid hazards caused by defects in safety modules (SMs). The optimization of diagnostic test interval (test application and result analysis) is important to reduce the overall in-system-structural test latency. In this paper, we present the hybrid performance modeling methodology adopted for IST architecture to optimize the design implementation with lowest possible and acceptable diagnostic test latency for key-on and/or key-off application of structural-test in Drive platform.
Milind Sonawane, Venkat Abilash Reddy Nerallapally, Alex Hsu, Shantanu Sarangi
VTS4
2017 At-speed capture global noise reduction & low-power memory test architecture
abstract
Traditionally, DFT patterns exacerbate dynamic power consumption in large ASICs. At-speed scan and memory tests are sensitive to voltage droop and peak current because the power grid is designed for functional power viruses (maximum workload applications) whose power consumption is much lower than DFT patterns. Our goal in this work is to ensure that the quality of test is not compromised while power is constrained to be within sign-off power budgets. We present an IEEE 1500-compliant Global Low Power Capture (GLPC) architecture with minimized interconnects between sub-blocks. For memory tests, we also present an extension of the architecture, Low Power MBIST (LP-MBIST) which shuts down the toggling of logic flops. Experimental results for both architectures show appreciable dynamic power reduction on recently taped out 16nm ASIC chips.
Bonita Bhaskaran, Sailendra Chadalavada, Shantanu Sarangi, Nithin Valentine, Venkat Abilash Reddy Nerallapally, Ayub Abdollahian
VTS3
2016 Advanced test methodology for complex SoCs
abstract
This paper presents the latest test methodology for NVIDIA's multi-billion transistor Mobile System on Chip (SoC) and Graphics Processing Unit (GPU). The paper describes the innovations that enhance the SoC plug-n-play scheme in terms of DFT. It also demonstrates how the architecture enables ultra-low pin count testing together with test data reuse and efficient test scheduling to improve the test quality while lowering the test cost. We present a scalable scan interface methodology coupled with core isolation and advanced clocking design while keeping the overall power budget for test within the limits of SoC Thermal Design Power (TDP). Silicon results are shared to demonstrate the effectiveness of this architecture.
Pavan Kumar Datla Jagannadha, Mahmut Yilmaz, Milind Sonawane, Sailendra Chadalavada, Shantanu Sarangi, Bonita Bhaskaran, Ayub Abdollahian
ITC5
2016 Flexible scan interface architecture for complex SoCs
abstract
Non-standardized scan interface within and across system-on-chips (SoCs) limits test-data reuse for intellectual properties (IPs). To overcome this limitation, we present a flexible and dynamic scan interface architecture that enables reuse of test-data for a given IP across SoCs with different scan pin configurations. The dynamic nature of this architecture also enables variable shift frequencies across different IPs in a given SoC. The architecture decouples the scan pin requirements from the design cycle of the IPs. It also uses bidirectional scan pins to further reduce test cost by using as few as two pins.
Milind Sonawane, Sailendra Chadalavada, Shantanu Sarangi, Amit Sanghani, Mahmut Yilmaz, Pavan Kumar Datla Jagannadha, Jonathon E. Colburn
VTS3
2016 Dynamic docking architecture for concurrent testing and peak power reduction
abstract
Interdependence of the clocking architecture across IPs and overall peak power consumption is a major bottleneck that prevents concurrent yet independent testing of an IP at a higher clock frequency. We use a dynamic clocking architecture that eliminates these dependencies and reduces peak shift power by using clock phase staggering at a granular level during system-on-chip (SoC) testing. A SoC design is typically composed of several Intellectual Property (IPs), some of which may be replicated. Generating a full set of test patterns targeting all IPs at the same time is computationally intensive and may be constrained by project schedule. Using this architecture, production test patterns are generated independently at the IP level and applied concurrently at the SoC level without exceeding the power budget of the chip during test. We present various aspects of the clocking architecture design along with simulation and silicon results to highlight the effectiveness of this architecture.
Milind Sonawane, Pavan Kumar Datla Jagannadha, Sailendra Chadalavada, Shantanu Sarangi, Mahmut Yilmaz, Amit Sanghani, Karthikeyan Natarajan, Jonathon E. Colburn, Anubhav Sinha
VTS4
2011 A clock-gating based capture power droop reduction methodology for at-speed scan testing
abstract
Excessive power dissipation caused by large amount of switching activities has been a major issue in scan-based testing. For large designs, the excessive switching activities during launch cycle can cause severe power droop, which cannot be recovered before capture cycle, rendering the at-speed scan testing more susceptible to the power droop. In this paper, we present a methodology to avoid power droop during scan capture without compromising at-speed test coverage. It is based on the use of a low area overhead hardware controller to control the clock gates. The methodology is ATPG (Automatic Test Pattern Generation)-independent, hence pattern generation time is not affected and pattern manipulation is not required. The effectiveness of this technique is demonstrated on several industrial designs.
Amit Sanghani, Shantanu Sarangi
DATE3