Liyang Lai

dblp:96/1318 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0003-1041-8980ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 7 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A testability-driven technology mapping method for optimized test point insertion
Xiaoze Lin, Liyang Lai, Biwei Xie, Huawei Li 0001
Integr.2
2025 An Efficient Parallel Fault Simulator for Functional Patterns on Multi-Core Systems
abstract
Fault simulation targeting functional patterns emerges as an essential mechanism within functional safety, crucial for validating the effectiveness of safety mechanisms. The acceleration of fault simulation for functional patterns is imperative for boosting the efficiency and adaptability of functional safety verification, presenting a significant yet unresolved challenge. In the paper, we propose an efficient fault simulator for functional patterns, utilizing three techniques including fault filtering, fault grouping, and CPU-based parallelism. The integration of these three techniques, tailored to the characteristics of functional patterns, reduces the runtime of fault simulation from different perspectives. The experimental results show that on a 48-core system, an average 79x speedup can be achieved by our parallel fault simulator against a commercial tool.
Xiaoze Lin, Liyang Lai, Huawei Li 0001, Biwei Xie
DATE2
2024 Sideway Scan, Solving the Achilles' Heel of Scan-based Diagnosis
abstract
the IC industry has seen tremendous value in diagnosis-driven yield analysis for speeding up physical failure analysis and yield ramp-up. However poor diagnostic resolution of chain failures in scan-based tests is known to be the Achilles’ heel in this flow. To address the problem, we propose sideway scan, a hardware-based solution for chain diagnosis. Scan chains are divided into multiple groups by clock domains or layout constraints. By introducing cyclic sideway transmission paths across scan chains within a group, the architecture guarantees that the snapshot of a broken chain can be transmitted to one or more good chains and the broken chain can capture value from good chains. Compared with state-of-the-art schemes such as reversible scan and two-dimensional scan, it is the first time a solution can simultaneously accomplish the objectives of manageable hardware overhead, high resolution in volume diagnosis, and flexibility in adaptive diagnosis. Simulation-based experiments validate the effectiveness of sideway scan in diagnosing chain failures caused by single and multiple faults.
Liyang Lai, Zefan Lin, Qitao Wang
ITC-Asia1
2024 Parallel Static Learning Toward Heterogeneous Computing Architectures
abstract
Static learning is a learning algorithm for finding additional implicit implications between gates in a netlist. In automatic test pattern generation (ATPG) the learned implications help recognize conflicts and redundancies early, and thus greatly improve the performance of ATPG. Though ATPG can further benefit from multiple runs of incremental or dynamic learning, it is only feasible when the learning process is fast enough. In this article, we study the performance optimization of static learning through parallelization on heterogeneous computing architectures, which includes multicore microprocessors (CPUs), and graphics processing units (GPUs). We discuss the advantages and limitations of each of these architectures. With their specific features in mind, we propose two different parallelization strategies that are tailored to multicore CPUs and GPUs. Speedup and performance scalability of the two proposed parallel algorithms are analyzed. It is demonstrated that for million-gate designs, close to linear performance gain is achieved on multicore CPUs, and up to$260\times $speedup over a commercial tool can be obtained on a single graphic card.
Xiaoze Lin, Liyang Lai, Huawei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 GPU-Based Concurrent Static Learning
abstract
Static learning is a learning algorithm for retrieving implicit logical relationships between nodes in a netlist. The learning results play an important role in improving automatic test pattern generation (ATPG), such as increasing fault coverage and reducing pattern count. In this work, we study accelerating static learning on graphics processing units (GPUs). By tailoring to the architectural features of GPUs, an algorithm of concurrent static learning is proposed. Multiple learning jobs are carried out simultaneously or concurrently in the same netlist. Moreover, the forward and backward implications of these concurrent jobs are processed as a whole, which leads to better utilization of the computing resources on GPUs. Experiments show that the algorithm can achieve up to 253x speedup against a single-threaded commercial tool and is about 1.8 times better than existing GPU-based solutions.
Huaxiao Liang, Xiaoze Lin, Liyang Lai, Naixing Wang
ITC3
2023 Adaptive Multidimensional Parallel Fault Simulation Framework on Heterogeneous System
abstract
Fault simulation is a critical component of the automatic test pattern generation (ATPG) tool, which is widely used in chip development. The CPU–GPU heterogeneous system can accelerate fault simulation. However, existing work faces the following challenges: 1) Path Divergence: The simulation path of different faults is not uniform, which leads to low parallel efficiency of different GPU threads; 2) Unbalanced Workload: The load of different computing units is not balanced, leading to serious differences in the execution time of each part; and 3) Poor Scalability: When the circuit scale increases, the GPU memory is limited and the simulation has strong structural dependence, which makes the simulation difficult. In this work, we propose an adaptive multidimensional parallel fault simulation framework based on the CPU–GPU heterogeneous system. We adaptively select different simulation approaches according to different circuit scales. In detail, we use the fanout-free region (FFR) grouping method to solve the problem of path divergence. We also use a combination of static and dynamic load balancing to tradeoff data handling and the execution time of each computing unit. We limit the queue length used in the GPU to improve the scalability of the simulation. To further accelerate, we propose the 4-D parallel architecture on multiple GPUs. Extensive experimental results show that our fault simulator based on 8 GPU is$105.7\times $faster than the commercial tool on average. For tens of millions of gate-level circuits, our fault simulator based on one GPU is up to$25.9\times $faster than the CPU single-threaded simulator.
Jingbo Hu, Guohao Dai 0001, Liuzheng Wang, Liyang Lai, Huazhong Yang, Yu Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 Scalable Parallel Static Learning
abstract
Static learning is a learning algorithm for finding additional implicit implications between gates in a netlist. In automatic test pattern generation (ATPG) the learned implications help recognize conflicts and redundancies early, and thus greatly improve the performance of ATPG. Though ATPG can further benefit from multiple runs of incremental or dynamic learning, it is only feasible when the learning process is fast enough. In the paper, we study speeding up static learning through parallelization on heterogeneous computing platform, which includes multi-core microprocessors (CPUs), and graphics processing units (GPUs). We discuss the advantages and limitations in each of these architectures. With their specific features in mind, we propose two different parallelization strategies that are tailored to multi-core CPUs and GPUs. Speedup and performance scalability of the two proposed parallel algorithms are analyzed. As far as we know, this is the first time that parallel static learning is studied in the literature.
Xiaoze Lin, Liyang Lai, Huawei Li 0001
ITC-Asia2
2020 GPU-based Hybrid Parallel Logic Simulation for Scan Patterns
abstract
GPGPU, general-purpose computing on graphics processing units, has been witnessed growing from a niche to a mainstream computing paradigm in the last decade. It is widely deployed in machine learning, artificial intelligence, and many scientific applications. In this article, we study gate-level logic simulation for scan test patterns, one of the key algorithmic components for test generation, fault grading and design rule check. We are exploring if GPGPU can deliver scalable performance speedup as promised by its massive parallelism. We discuss the limitations of the state of art work. With GPUs' architectural features in mind, a novel algorithm of hybrid race-tolerant parallel logic simulation is proposed to unleash its immense power of parallelization. Its unique integration of oblivious simulation and event-driven simulation leads to a scalable realization of parallel logic simulation for scan patterns. For the first time in the literature, it is demonstrated that a modest GPU can handle an industrial design of over twenty million gates with excellent performance scalability.
Liyang Lai, Qiting Zhang, Kun-Han Tsai, Wu-Tung Cheng
ITC-Asia1
2020 GPGPU-Based ATPG System: Myth or Reality?
abstract
General-purpose computing on graphics processing units (GPGPUs) is a programming model that uses graphics cards to perform computations traditionally done by CPU. It began to become practical with the advent of programmable shaders and floating-point support on GPU in around 2001. The spread of GPGPU has been accelerated with introduction of CUDA from NVIDIA in 2006 and later OpenCL in 2009. Nowadays GPGPU is widely deployed in various applications, such as data mining, artificial intelligence, and many scientific computations. GPGPU seemingly promises immense parallelism with massive concurrent cores, and thus much shorter run times. This is true for algorithms that bear intrinsic data and task parallelism, such as image and video processing. For an ATPG system where some algorithms are sequential in nature, the speedup is not easy to achieve in the real world. Flaws in setting up speedup evaluation can lead to false promises. Will GPGPU-based ATPG system become a reality? Or it is just a myth. In this paper, we try to provide an answer by surveying state-of-the-art works and by analyzing practical aspects of today's industrial designs.
Liyang Lai, Kun-Han Tsai, Huawei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 Diagnosis and Layout Aware (DLA) Scan Chain Stitching
abstract
Without appropriate stitching of scan chains, even with good diagnosis algorithm and diagnostic pattern generation, the chain diagnostic resolution may still be bad. In this paper, we propose a novel pattern-independent diagnosis and layout aware (DLA) scan chain stitching method: 1) the resolution is improved by increasing and properly distributing the sensitive scan cells, which can capture useful diagnostic information under both single- and multiple-fault situations; and 2) the scan cell layout placement is taken into account to reduce routing overhead and hence preserve the chip performance. Experiments using two different techniques to diagnose ISCAS'89/ITC'99 benchmark circuits with/without embedded scan compaction show the effectiveness of the proposed method in improving the diagnostic resolution. Impacts on chip performance, embedded scan compaction, transition fault coverage, and test power dissipation are negligible. The proposed method is also successfully applied to an industry circuit manufactured with 20-nm technology. The silicon results show 7× average resolution improvement comparing to without using the DLA scan chain stitching.
Jing Ye 0001, Yu Huang 0005, Yu Hu 0001, Wu-Tung Cheng, Ruifeng Guo, Liyang Lai, Ting-Pu Tai, Xiaowei Li 0001, Wei-pin Changchien, Daw-Ming Lee, Ji-Jan Chen, Sandeep C. Eruvathi, Kartik K. Kumara, Charles C. C. Liu, Sam Pan
IEEE Trans. Very Large Scale Integr. Syst.6
2013 Diagnosis and Layout Aware (DLA) scan chain stitching
abstract
Without appropriate stitching of scan chains, even with good diagnosis algorithm and diagnostic pattern generation, it may still result in bad scan chain diagnostic resolution. To improve the diagnostic resolution, we propose a novel Diagnosis and Layout Aware (DLA) scan chain stitching method, which is pattern independent and supports embedded scan compaction. It is based on three ideas: (1) increasing the number of sensitive scan cells, which can capture useful diagnostic information; (2) properly distributing the sensitive scan cells along the scan chains to enhance the overall resolution; (3) stitching scan cells based on their placement at layout to preserve the chip performance. Experiments on ISCAS'89/ITC'99 benchmark circuits and a real industry circuit based on 20nm technology with silicon results show that, the proposed DLA scan chain stitching method effectively improves the resolution, with negligible impact on chip performance, embedded scan compaction, transition fault coverage, and test power dissipation. The silicon results even show 7X average resolution improvement comparing to without using the proposed method.
Jing Ye 0001, Yu Huang 0005, Yu Hu 0001, Wu-Tung Cheng, Ruifeng Guo, Liyang Lai, Ting-Pu Tai, Xiaowei Li 0001, Wei-pin Changchien, Daw-Ming Lee, Ji-Jan Chen, Sandeep C. Eruvathi, Kartik K. Kumara, Charles C. C. Liu, Sam Pan
ITC6
2009 STDF Memory Fail Datalog Standard
abstract
Yield learning in modern technologies requires fail data logging from the scan and memory structural tests to gain insight into the failing location inside a chip.Currently there is no standard format to store the fail data in an efficient way. A group of more than 20companies from ATE, EDA, Semiconductor and Yield Management companies has been working to enhance the Standard Fail Data log Format (STDF) V4 to enable efficient fail data log for scan and memory fails. This paper describes the proposed memory fail datalog format.
Ajay Khoche, Jay Katz, Sauro Landini, Kochen Liao, Neetu Agrawal, Glenn Plowman, Songlin Zuo, Liyang Lai, John Rowe, Thomas Zanon
VTS8
2008 Detection and Diagnosis of Static Scan Cell Internal Defect
abstract
In this paper, we study the impact, detection and diagnosis of the defect inside a scan cell, which is called scan cell internal defect. We first use SPICE simulation to understand how a scan cell internal defect impacts the operation of a single scan cell. To study the detectability and diagnosability of a scan cell internal defect in a production test environment, we inject scan cell internal defects into a scan-based industrial design and perform fault simulation by using production scan test patterns. Next, we evaluate how effective an existing scan chain diagnosis technique based on traditional fault models can diagnose scan cell internal defect. We finally propose a new diagnosis algorithm to improve scan cell internal defect diagnostic resolution using scan cell internal fault model. Experimental results show the effectiveness of the proposed scan cell internal fault diagnosis technique.
Ruifeng Guo, Liyang Lai, Yu Huang 0005, Wu-Tung Cheng
ITC2
2007 Programmable Scan-Based Logic Built-In Self Test
abstract
This paper presents a programmable approach for performing scan-based logic built-in self test. This approach combines the techniques of reseeding and weighted random patterns testing. Reseeding is used to encode the bias cube and weighted patterns are used to fine tune the weight set. Experimental results show fault coverage comparable to ATPG can be achieved. Most importantly, the scheme fits well in the system test environment and high fault coverage can be obtained with a small number of reconfigurations on the BIST controller.
Liyang Lai, Wu-Tung Cheng, Thomas Rinderknecht
ATS1
2006 Signature Based Diagnosis for Logic BIST
abstract
This paper presents a new approach for performing logic BIST diagnosis exclusively using MISR signatures. Unlike conventional logic BIST diagnosis approaches which require either huge test time or complicated logic BIST design and ATE flow, signature based diagnosis does not require dynamically changing MISR operations for each failing device. Our experimental data shows that signature based diagnosis can achieve similar diagnosis resolution with manageable diagnosis run time while eliminating most of the complexity associated with the traditional approach to logic BIST diagnostics
Wu-Tung Cheng, Thomas Rinderknecht, Liyang Lai, Chris Hill
ITC4
2005 Hardware Ef.cient LBISTWith Complementary Weights
abstract
In this paper, a novel logic BIST (built-in self test) scheme with complementary weights is proposed. The BIST implementation combines random patterns with complementary-weight weighted patterns. A heuristic algorithm based on deterministic test set is developed to compute weight set with complementary weights. Hardware similar to bit-flipping is used to produce complementary weights. For random resistant ISCAS circuits, complete fault coverage can be achieved with very low hardware overhead. Experiments show that two complementary weights are sufficient for weighted random pattern testing and it presents a novel direction for exploiting weighted patterns.
Liyang Lai, Janak H. Patel, Thomas Rinderknecht, Wu-Tung Cheng
ICCD1
2004 Logic BIST with Scan Chain Segmentation
abstract
This work presents a novel BIST (built-in self test) scheme with scan chain segmentation. In the scheme, a combination of pseudo random patterns and single-weight patterns have been applied to CUT (circuit under test). Scan chain is partitioned into multiple segments delimited by inverters. When a single weighted pattern is applied to a segmented scan chain, successive segments receive bit patterns with complementary weights. Several segment configurations may be required to achieve full fault coverage. In this scheme the control logic is inside the scan path and built-in self test can be implemented without compromising timing performance of CUT. Experiments show that our scheme can obtain very good fault coverage. Hardware implementation is simple and straightforward.
Liyang Lai, Janak H. Patel, Thomas Rinderknecht, Wu-Tung Cheng
ITC1
2004 Logic BIST Using Constrained Scan Cells
abstract
This paper presents a novel scan cell based control point insertion technique which eliminates timing degradation of conventional control points in built-in self test (BIST) applications. In this approach, control points are encoded into scan chains. Observation points are applied to enhance fault coverage. At each phase, a set of control points are activated to detect a set of target faults. Compared to conventional test point insertion, scan cell based control points improve controllability of the core logic without compromising timing performance of circuit under test (CUT). Experimental results show that close to stuck-at fault coverage by automatic test pattern generation (ATPG) can be achieved by our BIST method.
Liyang Lai, Thomas Rinderknecht, Wu-Tung Cheng, Janak H. Patel
VTS1