EDBT 2026 Demo / reviewers in the wild / expert
Karthikeyan Natarajan
dblp:21/9680
· DBLP profile ↗
9ranked-venue papers
1as first author
4since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Power Aware TestabstractPower handling during test is an important requirement that needs to be considered during chip design, silicon bring-up, and in-system testing. In this tutorial, we will start by reviewing the importance of power and listing the different problems faced with poor power intent. We will then give an overview of different power aspects related to test, from RTL implementation to in-system validation, and how each step can impact the overall performance. Next, we will introduce different design-for-test (DFT) techniques that help improve power planning to produce optimized quality of results (QoR). Finally, we will present data sets on how each of the listed techniques implemented on real designs produces desired results. Likith Kumar Manchukonda, Karthikeyan Natarajan, Manish Arora |
ETS | 2 |
| 2022 | Comprehensive Power-Aware ATPG Methodology for Complex Low-Power DesignsabstractTest-power is an important factor that needs to be managed during Automatic Test Pattern Generation (ATPG), silicon bring-up, and in-system test. The current low-power ATPG techniques involve functional clock-gating approaches during capture as well as software and hardware based adjacent-fill techniques during shift, capture for limiting sequential cell toggling activity to reduce the overall power on Automatic Test Equipment (ATE). However, the current approach over-relies on sequential switching activity without considering associated combinational switching activity which results in an inconsistency between the power estimations during ATPG and actual test power on the ATE. Power simulation sign-off data provides more accurate modeling of design's switching-activity and power characteristics. In this paper, we present a new technique that addresses the mismatch between power estimations during ATPG and actual power on ATE by using power-simulation sign-off data during ATPG which allows generating patterns with a capture power profile that more accurately matches the power consumption during ATE test. This technique results in reduced IR-drop and dramatically improves the yield. In this paper, we also present the silicon data from MediaTek designs showing that lower peak power, lower IR drop, and lower Vmin was achieved using the new comprehensive Power-Aware ATPG. Khader S. Abdel-Hafez, Michael Dsouza, Likith Kumar Manchukonda, Elddie Tsai, Karthikeyan Natarajan, Ting-Pu Tai, Wenhao Hsueh, Smith Lai |
ITC | 5 |
| 2022 | High Speed IO Access for Test forms the foundation for Silicon Lifecycle ManagementabstractWith the rapid growth in semiconductor complexity, higher expectations for SoC performance and longevity, there is a need to continuously monitor the silicon throughout its life cycle to maximize performance and identify defects before they impact system operations. SCAN Vectors are reused for Silicon Lifecycle Management (SLM) along with sensors which are embedded into the silicon device for monitoring and detecting failures before they impact system functions. The volume of such data generally tends to be huge which requires a robust network and access mechanism. With the industry moving to using the High-Speed IO for In-Field Test Access, this provides us an opportunity to use the same access mechanism for SLM. The High-Speed IO access mechanism provides plenty of bandwidth as the native protocol of these interfaces are used. In this paper we will provide an overview of the High-Speed IO access solution which enables the SCAN Vectors for In-Field SLM applications. We also explore HSAT architectures for SLM structures like SCAN, JTAG, sensors and monitors. Brendan Tully, Karthikeyan Natarajan |
ITC | 3 |
| 2022 | Novel Technique for Manufacturing & In-system Testing of Large Scale SoC using Functional Protocol Based High-Speed I/OabstractTest time continues to increase due to the growing design size and complexity of modern large-scale System-on-Chip (SoCs). An increasing adoption of chiplet-based design method further exacerbates this problem as it reduces the number of pins available for test pattern application. Additionally, the SoCs deployed in safety-critical applications require in-system testing with very short test times. In this paper, we introduce a novel solution that addresses these challenges by using the existing high-speed functional interfaces of an SoC, such as PCIe or USB, to perform both scan test and in-system test with easy accessibility. This solution uses the native protocol of these scalable high-speed interfaces to deliver packetized test data to the device-under-test (DUT) at significantly faster speeds than can be achieved using General Purpose Input/Outputs (GPIOs), reducing test time. This paper provides the details of the solution and its practical implementation along with silicon data on an Amazon Machine Learning (ML) SoC. Brendan Tully, Abhijeet Samudra, Ajay Nagarandal, Karthikeyan Natarajan, Rahul Singhal |
VTS | 5 |
| 2019 | High Performance Graph Convolutional Networks with Applications in Testability AnalysisabstractApplications of deep learning to electronic design automation (EDA) have recently begun to emerge, although they have mainly been limited to processing of regular structured data such as images. However, many EDA problems require processing irregular structures, and it can be non-trivial to manually extract important features in such cases. In this paper, a high performance graph convolutional network (GCN) model is proposed for the purpose of processing irregular graph representations of logic circuits. A GCN classifier is firstly trained to predict observation point candidates in a netlist. The GCN classifier is then used as part of an iterative process to propose observation point insertion based on the classification results. Experimental results show the proposed GCN model has superior accuracy to classical machine learning models on difficult-to-observation nodes prediction. Compared with commercial testability analysis tools, the proposed observation point insertion flow achieves similar fault coverage with an 11% reduction in observation points and a 6% reduction in test pattern count. Yuzhe Ma, Haoxing Ren, Brucek Khailany, Harbinder Sikka, Lijuan Luo, Karthikeyan Natarajan, Bei Yu 0001 |
DAC | 6 |
| 2018 | Lossless Parallel Implementation of a Turbo Decoder on GPUabstractTurbo decoders use the recursive BCJR algorithm which is computationally intensive and hard to parallelise. The branch metric and extrinsic log-likelihood ratio computations are easily parallelisable, but the forward and backward metric computation is not parallelisable without compromising bit error rate. This paper proposes a lossless parallelisation technique for Turbo decoders on Graphics Processing Units (GPU). The recursive forward and backward metric computation is formulated as prefix (scan) matrix multiplication problem which is computed on the GPU using parallel prefix sum computation technique. Overall, this method achieves a throughput of 73 Mbps for a 3GPP LTE compliant turbo decoder without any BER loss and latency as low as 61 μs. Karthikeyan Natarajan, Nitin Chandrachoodan |
HiPC | 1 |
| 2016 | Dynamic docking architecture for concurrent testing and peak power reductionabstractInterdependence of the clocking architecture across IPs and overall peak power consumption is a major bottleneck that prevents concurrent yet independent testing of an IP at a higher clock frequency. We use a dynamic clocking architecture that eliminates these dependencies and reduces peak shift power by using clock phase staggering at a granular level during system-on-chip (SoC) testing. A SoC design is typically composed of several Intellectual Property (IPs), some of which may be replicated. Generating a full set of test patterns targeting all IPs at the same time is computationally intensive and may be constrained by project schedule. Using this architecture, production test patterns are generated independently at the IP level and applied concurrently at the SoC level without exceeding the power budget of the chip during test. We present various aspects of the clocking architecture design along with simulation and silicon results to highlight the effectiveness of this architecture. Milind Sonawane, Pavan Kumar Datla Jagannadha, Sailendra Chadalavada, Shantanu Sarangi, Mahmut Yilmaz, Amit Sanghani, Karthikeyan Natarajan, Jonathon E. Colburn, Anubhav Sinha |
VTS | 7 |
| 2016 | A programmable method for low-power scan shift in SoC integrated circuitsabstractWe present a programmable method for shift-clock stagger assignment to reduce power supply noise during system-on-chip (SoC) testing. An SoC design is typically composed of several blocks and two neighboring blocks that share the same power rails should not be toggled at the same time during shift. Therefore, the proposed programmable method does not assign the same stagger value to neighboring blocks. The positions of all blocks are first analyzed and the shared boundary length between blocks is then calculated. Based on the position relationships between the blocks, a mathematical model is presented to derive optimal result for small-to-medium sized problems. For larger designs, a heuristic algorithm is proposed and evaluated. We present assignment results as well as power-analysis results and silicon data for industry designs to highlight the effectiveness of the proposed method. Ran Wang 0002, Bonita Bhaskaran, Karthikeyan Natarajan, Ayub Abdollahian, Kaushik Narayanun, Krishnendu Chakrabarty, Amit Sanghani |
VTS | 3 |
| 2011 | Design and implementation of a time-division multiplexing scan architecture using serializer and deserializer in GPU chipsabstractWe present the design and implementation details of a time-division demultiplexing/multiplexing based scan architecture using serializer/deserializer. This is one of the key DFT features implemented on NVIDIA's Fermi family GPU (Graphic Processing Unit) chips. We provide a comprehensive description on the architecture and specifications. We also depict a compact serializer/deserializer module design, test timing consideration, design rule and test pattern verification. Finally, we show silicon data collected from Fermi GPUs. Amit Sanghani, Karthikeyan Natarajan |
VTS | 3 |