Kaushik Narayanun

dblp:180/6657 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 since 2021
YearPublicationVenuePosition
2024 A Scalable & Cost Efficient Next-Gen Scan Architecture: Streaming Scan Test via NVIDIA MATHS
abstract
Streaming scan test architectures can greatly optimize the test data delivery to large industrial designs. This paper discusses what happens when such architectures are combined with nearly unlimited data bandwidth provided by NVIDIA MATHS (Mechanism to Access Test-Data over High-Speed Link). There are multiple techniques for efficient use of scan bandwidth and its impact on the overall test cost and test quality. We have also architected various debug techniques for silicon bring-up. This scan architecture was designed for highest throughput to test multiple dies in parallel with lowest test power and best diagnosability.
Kunal Jain Mangilal, Mahmut Yilmaz, Vishal Agarwal, Shantanu Sarangi, Kaushik Narayanun
ITC5
2022 Observation Point Insertion Using Deep Learning
abstract
Silent Data Corruption (SDC) is one of the critical problems in the field of testing, where errors or corruption do not manifest externally. As a result, there is increased focus on improving the outgoing quality of dies by striving for better correlation between structural and functional patterns to achieve a low DPPM. This is very important for NVIDIA's chips due to the various markets we target; for example, automotive and data center markets have stringent in-field testing requirements. One aspect of these efforts is to also target better testability while incurring lower test cost. Since structural testing is faster than functional tests, it is important to make these structural test patterns as effective as possible and free of test escapes. However, with the rising cell count in today's digital circuits, it is becoming increasingly difficult to sensitize faults and propagate the fault effects to scan-flops or primary outputs. Hence, methods to insert observation points to facilitate the detection of hard-to-detect (HtD) faults are being increasingly explored. In this work, we propose an Observation Point Insertion (OPI) scheme using deep learning with the motivation of achieving - 1) better quality test points than commercial EDA tools leading to a potential lower pattern count 2) faster turnaround time to generate the test points. In order to achieve better pattern compaction than commercial EDA tools, we employ Graph Convolutional Networks (GCNs) to learn the topology of logic circuits along with the features that influence its testability. The graph structures are subsequently used to train two GCN-type deep learning models - the first model predicts signal probabilities at different nets and the second model uses these signal probabilities along with other features to predict the reduction in test-pattern count when OPs are inserted at different locations in the design. The features we consider include structural features like gate type, gate logic, reconvergent-fanouts and testability features like SCOAP. Our simulation results indicate that the proposed machine learning models can predict the probabilistic testability metrics with reasonable accuracy and can identify observation points that reduce pattern count.
Bonita Bhaskaran, Sanmitra Banerjee, Kaushik Narayanun, Shao-Chun Hung, Seyed Nima Mozaffari, Tung-Che Liang
ICCAD3
2022 NVIDIA MATHS: Mechanism to Access Test-Data over High-Speed Links
abstract
MATHS (Mechanism to Access Test-Data over High-Speed Link) provides a high-throughput PCIe based system to structurally test system-on-chips (SOCs) at wafer and system-level. The system removes the need for expensive test equipment by eliminating the input/output pin (IO) requirements and memory per IO needs. It simplifies the ATE architecture and design to enable smaller form factors and reduce capital costs of ownership. MATHS enables eliminating the assembly test-insertion by directly testing the SOCs on system level platforms, further reducing the costs. Since the mechanism is based on PCIe standards, it is highly portable across all platforms including ATE, system-level test, board, and in-field testing.
Mahmut Yilmaz, Pavan Kumar Datla Jagannadha, Kaushik Narayanun, Shantanu Sarangi, Francisco Da Silva, Joe Sarmiento, Smbat Tonoyan, Ashwin Chintaluri, Animesh Khare, Milind Sonawane, Anitha Kalva, Alex Hsu, Jayesh Pandey
VTS3
2019 An Efficient Supervised Learning Method to Predict Power Supply Noise During At-speed Test
abstract
The Power Distribution Network (PDN) is designed for worst-case power-hungry functional use-cases. Most often Design for Test (DFT) scenarios are not accounted for, while optimizing the PDN design. Automatic Test Pattern Generation (ATPG) tools typically follow a greedy algorithm to achieve maximum fault coverage with short test times. This causes Power Supply Noise (PSN) during scan testing to be much higher than functional mode since switching activity is higher by an order of magnitude. Understanding the noise characteristics through exhaustive pattern simulation is extremely machine and memory intensive and requires unsustainably long runtimes. Hence, we aggressively limit switching factors to conservative estimates and rely on post-silicon noise characterization to optimize test vectors. In this work, we propose a novel method to predict simultaneous switching noise using fast Deep Neural Networks (DNNs) such as Fully Connected Network, Convolutional Neural Network, and Natural Language Processing. Our approach, that is based on pre-silicon ATPG vectors, is significantly faster than conventional estimation methods and can potentially reduce the test time.
Seyed Nima Mozaffari, Bonita Bhaskaran, Kaushik Narayanun, Ayub Abdollahian, Vinod Pagalone, Shantanu Sarangi, Jonathon E. Colburn
ITC3
2017 Test-cost optimization in a scan-compression architecture using support-vector regression
abstract
Scan compression is widely used in high-volume testing of complex integrated circuits. With an increase in design complexity, the increased density of unknown (X) values from output responses reduces compression efficiency. In order to effectively block X values and maximize the effectiveness of test compression, a scan-compression architecture has recently been proposed, in which deterministic test patterns can be loaded into selected scan cells by controlling the initial state of the pseudo-random pattern generator (PRPG). A careful selection of the PRPG length is however essential to reduce test cost. We propose an optimization method based on support-vector regression to determine the PRPG length for test-cost reduction in a given scan-compression architecture. A correlation-based feature selection methodology is also proposed to reduce the amount of data needed for the accurate selection of the PRPG length. Experimental results on industrial designs highlight the effectiveness of the proposed method.
Jonathon E. Colburn, Vinod Pagalone, Kaushik Narayanun, Krishnendu Chakrabarty
VTS4
2016 Test method and scheme for low-power validation in modern SOC integrated circuits
abstract
Test Mode power can be 5X higher than functional power in GPUs, while the power grid is designed only for worst-case functional toggle. The large simultaneous switching noise induced on the power rails during at-speed capture testing is constrained by means of hardware solution. To determine the best low power mode for ATPG, we propose novel techniques to: estimate global peak current (di), determine local droop trend and validate and further optimize chosen power settings with exhaustive post-silicon power mode tuning. During Power Optimization (PO) phase, the measured clock frequency (fclk) and Vdroop are analyzed on every pattern and test coverage and pattern count are optimized for the production pattern set. We share correlation results and Power Supply Noise (PSN) distribution for the production pattern set on recent 28-nm GPUs.
Bonita Bhaskaran, Amit Sanghani, Kaushik Narayanun, Ayub Abdollahian, Amit Laknaur
VTS3
2016 A programmable method for low-power scan shift in SoC integrated circuits
abstract
We present a programmable method for shift-clock stagger assignment to reduce power supply noise during system-on-chip (SoC) testing. An SoC design is typically composed of several blocks and two neighboring blocks that share the same power rails should not be toggled at the same time during shift. Therefore, the proposed programmable method does not assign the same stagger value to neighboring blocks. The positions of all blocks are first analyzed and the shared boundary length between blocks is then calculated. Based on the position relationships between the blocks, a mathematical model is presented to derive optimal result for small-to-medium sized problems. For larger designs, a heuristic algorithm is proposed and evaluated. We present assignment results as well as power-analysis results and silicon data for industry designs to highlight the effectiveness of the proposed method.
Ran Wang 0002, Bonita Bhaskaran, Karthikeyan Natarajan, Ayub Abdollahian, Kaushik Narayanun, Krishnendu Chakrabarty, Amit Sanghani
VTS5