EDBT 2026 Demo / reviewers in the wild / expert
Yu Huang 0005
dblp:39/6301-5
· DBLP profile ↗
104ranked-venue papers
33as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 96 · 32 first-author · 27 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scan Chain Reordering for Improving Test Coverage with Compression
Hairui Cai, Zezhong Wang 0006, Yu Huang 0005, Naixing Wang, Zhouxing Su, Zhipeng Lv |
VTS | 3 |
| 2026 | C2C: Cell-to-Cell Controllability Evaluation for Partial Scan Selection
Hairui Cai, Liuzheng Wang, Lingxiang Liao, Yu Huang 0005, Zhouxing Su, Zhipeng Lv |
VTS | 7 |
| 2026 | Coverage-Aware Scan-Chain Reordering Under Iso-Power Constraints for Programmable Low-Power LBIST
Yumei Hu, Hairui Cai, Xiangheng Xie, Zhipeng Lv, Zhouxing Su, Yu Huang 0005, Zezhong Wang 0006 |
VTS | 7 |
| 2025 | ATPG-Based Weighted Scan Chain Control for Programmable Low-Power LBISTabstractLogic built-in self-test (LBIST) suffers from excessive power consumption due to high toggling rates caused by pseudo-random patterns. This paper presents a programmable low-power LBIST scheme that leverages scan chain weighting based on ATPG-guided fault analysis. By analyzing the distribution of specified bits across ATPG-generated test cubes, each scan chain is assigned a weight indicating its relative contribution to fault detection. Chains are then grouped into seven activation levels, each mapped to a distinct toggle probability to balance power and test coverage. A configurable control circuit based on shift and hold registers generates the required low-power signals. Experimental results on industrial-scale designs demonstrate that the proposed method achieves significantly higher fault coverage under identical power constraints compared to a commercial LBIST solution. Yumei Hu, Hairui Cai, Xiaohui Xue, Yu Huang 0005, Zhipeng Lv, Zhouxing Su, Zezhong Wang 0006 |
ICCD | 5 |
| 2025 | FGNN2: A Powerful Pretraining Framework for Learning the Logic Functionality of CircuitsabstractLearning feasible representation from raw gate-level circuits is essential for incorporating machine learning techniques in logic synthesis, physical design, or verification. Existing structure-based learning methods tend to concentrate mainly on the graph topology, often neglecting logic functionality. This oversight frequently results in a failure to capture the underlying semantics, thereby limiting their overall applicability. To address the concern, we propose a novel circuit representation learning framework, FGNN2, that utilizes a contrastive scheme to effectively extract generic functionality knowledge. We construct a comprehensive pretraining dataset through a customized circuit augmentation scheme. We have also developed a novel contrastive loss function to capture the relative functional distance between different circuits, and to generate representations that are invariant to the input order. In addition, we employed a customized graph neural network (GNN) architecture to better align with the above framework. Comprehensive experiments on the multiple complex real-world designs demonstrate that our proposed solution significantly outperforms the state-of-the-art circuit representation learning flows. Ziyi Wang 0010, Zhuolun He, Guangliang Zhang, Qiang Xu 0001, Tsung-Yi Ho, Yu Huang 0005, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | HDLdebugger: Streamlining HDL debugging with Large Language ModelsabstractIn the domain of chip design, hardware description languages (HDLs) play a pivotal role. However, due to the inherent complexity of HDLs and the scarcity of high-quality debugging resources, HDL bug fixing remains a challenging and time-consuming task, even for seasoned engineers. Consequently, there is a pressing need to develop automated HDL code debugging models, which can alleviate the burden on hardware engineers. Despite the strong capabilities of large language models (LLMs) in generating, completing, and debugging software code, their utilization in the specialized field of HDL debugging has been limited and, to date, has not yielded satisfactory results. In this paper, we propose an LLM-assisted HDL debugging framework, namely HDLdebugger, which consists of HDL debugging data generation via a reverse engineering approach, a search engine for retrieval-augmented generation, and a retrieval-augmented LLM fine-tuning approach. Through the integration of these components, HDLdebugger can automate and streamline HDL debugging for chip design. Our comprehensive experiments, conducted on an HDL code dataset sourced from Industry, reveal that HDLdebugger outperforms 13 cutting-edge LLM baselines, displaying exceptional effectiveness in HDL code debugging. Xufeng Yao, Haoyang Li 0002, Tsz Ho Chan, Wenyi Xiao, Mingxuan Yuan, Yu Huang 0005, Lei Chen 0002, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2024 | BetterV: Controlled Verilog Generation with Discriminative GuidanceabstractDue to the growing complexity of modern Integrated Circuits (ICs), there is a need for automated circuit design methods. Recent years have seen increasing research in hardware design language generation to facilitate the design process. In this work, we propose a Verilog generation framework, BetterV, which fine-tunes large language models (LLMs) on processed domain-specific datasets and incorporates generative discriminators for guidance on particular design demands. Verilog modules are collected, filtered, and processed from the internet to form a clean and abundant dataset. Instruct-tuning methods are specially designed to fine-tune the LLMs to understand knowledge about Verilog. Furthermore, data are augmented to enrich the training set and are also used to train a generative discriminator on particular downstream tasks, providing guidance for the LLMs to optimize Verilog implementation. BetterV has the ability to generate syntactically and functionally correct Verilog, outperforming GPT-4 on the VerilogEval benchmark. With the help of task-specific generative discriminators, BetterV achieves remarkable improvements on various electronic design automation (EDA) downstream tasks, including netlist node reduction for synthesis and verification runtime reduction with Boolean Satisfiability (SAT) solving. Zehua Pei, Hui-Ling Zhen, Mingxuan Yuan, Yu Huang 0005, Bei Yu 0001 |
ICML | 4 |
| 2024 | Diagnosis of Defects on Global SignalsabstractGlobal control signals account for over 10% of the circuit area and are challenging to diagnose when defects occur. This paper introduces a global signal diagnosis technique that can be easily integrated into the existing chain failure diagnosis flow in volume diagnosis. In addition to clock and scan enable signals, set and reset signals are also taken into account during diagnosis. The layout topology of global nets is utilized to improve the accuracy and resolution of the diagnosis result. The effectiveness of this method is validated through several physical failure analysis results. Weiming Zhang 0005, Xiaotian Ding, Yu Huang 0005 |
ITC | 6 |
| 2024 | Large circuit models: opportunities and challengesabstractAbstract Within the electronic design automation (EDA) domain, artificial intelligence (AI)-driven solutions have emerged as formidable tools, yet they typically augment rather than redefine existing methodologies. These solutions often repurpose deep learning models from other domains, such as vision, text, and graph analytics, applying them to circuit design without tailoring to the unique complexities of electronic circuits. Such an “AI4EDA” approach falls short of achieving a holistic design synthesis and understanding, overlooking the intricate interplay of electrical, logical, and physical facets of circuit data. This study argues for a paradigm shift from AI4EDA towards AI-rooted EDA from the ground up, integrating AI at the core of the design process. Pivotal to this vision is the development of a multimodal circuit representation learning technique, poised to provide a comprehensive understanding by harmonizing and extracting insights from varied data sources, such as functional specifications, register-transfer level (RTL) designs, circuit netlists, and physical layouts. We champion the creation of large circuit models (LCMs) that are inherently multimodal, crafted to decode and express the rich semantics and structures of circuit data, thus fostering more resilient, efficient, and inventive design methodologies. Embracing this AI-rooted philosophy, we foresee a trajectory that transcends the current innovation plateau in EDA, igniting a profound “shift-left” in electronic design methodology. The envisioned advancements herald not just an evolution of existing EDA tools but a revolution, giving rise to novel instruments of design-tools that promise to radically enhance design productivity and inaugurate a new epoch where the optimization of circuit performance, power, and area (PPA) is achieved not incrementally, but through leaps that redefine the benchmarks of electronic systems’ capabilities. Zhufei Chu, Wenji Fang, Tsung-Yi Ho, Ru Huang 0001, Yu Huang 0005, Sadaf Khan, Yun Liang 0001, Yibo Lin, Guojie Luo, Hongyang Pan, Zhengyuan Shi, Guangyu Sun 0003, Dimitrios Tsaras, Runsheng Wang, Ziyi Wang 0010, Xinming Wei, Zhiyao Xie, Qiang Xu 0001, Chenhao Xue, Junchi Yan, Bei Yu 0001, Mingxuan Yuan, Evangeline F. Y. Young, Xuan Zeng 0001, Haoyi Zhang, Zuodong Zhang, Hui-Ling Zhen, Binwu Zhu, Keren Zhu 0001, Sunan Zou |
Sci. China Inf. Sci. | 7 |
| 2024 | Erratum to: Large circuit models: opportunities and challenges
Zhufei Chu, Wenji Fang, Tsung-Yi Ho, Ru Huang 0001, Yu Huang 0005, Sadaf Khan, Yun Liang 0001, Yibo Lin, Guojie Luo, Hongyang Pan, Zhengyuan Shi, Guangyu Sun 0003, Dimitrios Tsaras, Runsheng Wang, Ziyi Wang 0010, Xinming Wei, Zhiyao Xie, Qiang Xu 0001, Chenhao Xue, Junchi Yan, Bei Yu 0001, Mingxuan Yuan, Evangeline F. Y. Young, Xuan Zeng 0001, Haoyi Zhang, Zuodong Zhang, Hui-Ling Zhen, Binwu Zhu, Keren Zhu 0001, Sunan Zou |
Sci. China Inf. Sci. | 7 |
| 2023 | Improve Volume Physical-Aware Diagnosis via Active Pattern SamplingabstractVolume diagnosis is an essential step in diagnosis driven yield analysis. Generally, diagnosis run time increases when more failing patterns are collected in a fail log. In order to improve diagnosis throughput, pattern sampling, which only uses a subset of the failing patterns, is a common practice in volume diagnosis. Compared to the results achieved by using all the failing patterns, traditional pattern sampling has a negative effect on the diagnosis accuracy and resolution. In this paper, we propose a layout-aware active pattern sampling method which improves the quality of diagnosis in terms of accuracy and resolution over the traditional pattern sampling method. Meanwhile, it achieves higher diagnosis throughput compared with the non-sampling methodology. Diagnostic results on four industrial designs show that the average suspects per symptom are reduced by 10.77% ~35.86 % compared to the traditional pattern sampling method. Besides, our algorithm has an advantage of identifying more actual defective locations compared with the traditional sampling or non-sampling method. Jiaxing Gao, Yu Huang 0005, Xiaotian Ding, Weiming Zhang 0005 |
ATS | 4 |
| 2023 | Fault Simulation Acceleration Based on ARM Multi-core CPU ArchitectureabstractFault simulation plays an important role in ATPG and fault diagnosis of integrated circuits. However, with the increasing complexity of chips, the simulation of tens of millions faults of VLSI needs a lot of time and computing resources. To improve the simulation efficiency, many methods and technologies have emerged, such as GPU acceleration and distributed computing. However, there are some challenges and limitations to these approaches, such as high cost, high energy consumption, and programming complexity. In contrast, the acceleration of fault simulation based on ARM multi-core CPU has the characteristics of strong multi-threaded parallel computing capability, low cost, low energy consumption, and easy implementation. Therefore, a new method for accelerating fault simulation based on ARM multi-core CPU is proposed in this paper, which adopts an enhanced parallel simulation method. In this method, each node has its own memory space and processor, and the CPU can work at full load. This paper also provides solutions to NUMA affinity, cross-node access and other problems. This paper will elaborate on these methods and demonstrate their effectiveness and accuracy in fault simulation. Shi-Jie Ye, Yun-Ju Liu, Liuzheng Wang, Hui-Ling Zhen, Weiming Zhang 0005, Yu Huang 0005 |
ATS | 6 |
| 2023 | TOFU: A Two-Step Floorplan Refinement Framework for Whitespace ReductionabstractFloorplanning, as an early step in physical design, will greatly affect the PPA of the later stages. To achieve better performance while main-taining relatively the same chip size, the utilization of the generated floorplan needs to be high and constraints related to design rules, routability, power should be honored. In this paper, we propose a two-step framework, called TOFU, for floorplan whitespace reduction with fixed-outline and soft/pre- placed/hard modules modeled. Whitespace is first reduced by iteratively refining the locations of modules. Then the modules near whitespace will be changed into rectilinear shapes to further improve the utilization. To ensure the legality and quality of the intermediate floorplan during the refinement process, a constraint graph-based legalizer with a novel constraint graph construction method is proposed. Experimental results show that the whitespace of the initial floorplans generated by Corblivar [1] can be reduced by about 70% on average and up to 90% in several cases. Moreover, the resulting wirelength is also 3% shorter due to a higher utilization. Shixiong Kai, Chak-Wa Pui, Shougao Jiang, Bin Wang 0034, Yu Huang 0005, Jianye Hao |
DATE | 6 |
| 2023 | EffiSyn: Efficient Logic Synthesis with Dynamic Scoring and PruningabstractLogic synthesis tools synthesize circuit structures to optimize specific targets given reasonable constraints and runtime using a set of well-defined operators. The efficiency of these operators is critical to achieving better runtime and optimization convergence. However, most synthesis operators are designed heuristically with fixed and sub-optimal traversal orders of nodes, cuts, and candidate subgraphs that are independent of circuit structures and functionalities. This leads to redundant computation and loss of optimization opportunities. Due to the spatial structure similarity, sub-circuits already synthesized in the same circuit contain meaningful information to guide more efficient synthesis for unvisited sub-circuits. Historical evaluation and gain features can learn from conflicts and be utilized to predict synthesis gain and prune invalid sub-circuits. Instead of using high-weight feature extraction and models, we utilize efficient prediction models with lightweight structural and functional features to reduce overhead. We thus propose a generalizable and dynamic scoring and pruning framework EffiSyn to accelerate logic synthesis operators while maintaining synthesis effectiveness. For example, we further improve the highly optimized operator drw by scoring and pruning invalid cuts and precomputed subgraphs. Extensive experiments on 20 public and industrial circuits validate that EffiSyn can accelerate drw by about 35% with negligible effectiveness loss or even with effectiveness improvement. Experiments over diverse circuits and synthesis sequences also validate the generalization of the proposed framework. Xing Li 0023, Lei Chen 0002, Jiantang Zhang, Shuang Wen 0007, Weihua Sheng, Yu Huang 0005, Mingxuan Yuan |
ICCAD | 6 |
| 2023 | A Unified Framework for Layout Pattern Analysis With Deep Causal EstimationabstractThe decrease of feature size and the growing complexity of the fabrication process lead to more failures in manufacturing semiconductor devices. Therefore, identifying the root cause layout patterns of failures becomes increasingly crucial for yield improvement. In this article, a novel layout-aware diagnosis-based layout pattern analysis framework is proposed to identify the root cause efficiently. At the first stage of the framework, an encoder network trained using contrastive learning is used to extract representations of layout snippets that are invariant to trivial transformations, including shift, rotation, and mirroring, which are then clustered to form layout patterns. At the second stage, we model the causal relationship between any potential root cause layout patterns and the systematic defects by a structural causal model, which is then used to estimate the average causal effect (ACE) of candidate layout patterns on the systematic defect to identify the true root cause. Experimental results on real industrial cases demonstrate that our framework outperforms a commercial tool with higher accuracies and around$\times 8.4$speedup on average. Ran Chen 0001, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Bei Yu 0001, Pengyun Li, Yu Huang 0005, Jianye Hao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | Old School, New Primitive: Toward Scalable PUF-Based Authenticated Encryption Scheme in IoTabstractThe Internet of Things (IoT) facilitates the information exchange between people and smart devices. It needs cryptographic measures to secure its communications and interconnected objects. However, cyber–physical attacks pose a great challenge to the protection of secret keys inside. Physically unclonable function (PUF) is a promising hardware primitive with unclonable structures providing tamper evidence for a device. Moreover, a PUF instance has a unique set of randomized challenge–response pairs. Although it can be integrated into a security scheme to replace long-term keys, designing a dedicated PUF-based cryptographic algorithm that supports peer-to-peer communication remains a challenging field to explore. In this article, we propose SPEAR, a scalable PUF-based authenticated encryption (AE) scheme that uses no cryptographic primitives other than PUF and hash functions. SPEAR can be deployed on peer IoT devices that have performed a handshake protocol to obtain shared credentials. Its security under the chosen ciphertext attack is formally proved using the game-playing technique, and it is still secure when attackers physically extract the credentials. In addition, we give a variant,$x$SPEAR, to involve associated data and avoid the nonce reuse problem. Compared to other PUF-based ciphers, it performs better in terms of storage overhead and PUF evaluation times. SPEAR first realizes scalable AE based on PUF and can be a practical solution for IoT. Dawu Gu, Yu Huang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Accelerate SAT-based ATPG via Preprocessing and New Conflict Management HeuristicsabstractDue to the continuous advancement of semicon-ductor technologies, there are more defects than ever widely distributed in manufactured chips. In order to meet the high product quality and low defective-parts-per-million (DPPM) goals, Boolean Satisfiability (SAT) technique has been shown to be a robust alternative to conventional APTG techniques, especially for hard-to-detect faults. However, the SAT-based ATPG still confronts two challenges. The first one is to reduce extra computational overhead of SAT modeling, i.e. to transform a circuit testing problem to a Conjunctive Normal Form (CNF) which is the foundation of modern SAT solvers. The second one lies in the SAT solver's efficiency which is brought by the loss of structural information during CNF transformation. In this work, we propose a new SAT-based ATPG approach to address the two challenges mentioned above: (1) To reduce CNF transformation overhead, we utilize a simulation-driven pre-processing for narrowing down the fault propagation and activation logic cones, leading to an improvement in CNF transformation and reduction in runtime. (2) To further improve the solving efficiency, We propose new ranking-based heuristics to build more effective conflict database, enabling the direct solving for small scale instance and a looking-head method for large scale ones. Extensive experimental results on industrial circuits demonstrate that on average the proposed approach could cover 89.67% of the faults failed by a commercial ATPG tool with a comparable runtime. Junhua Huang, Hui-Ling Zhen, Naixing Wang, Mingxuan Yuan, Yu Huang 0005, Jiping Tao |
ASP-DAC | 6 |
| 2022 | Functionality matters in netlist representation learningabstractLearning feasible representation from raw gate-level netlists is essential for incorporating machine learning techniques in logic synthesis, physical design, or verification. Existing message-passing-based graph learning methodologies focus merely on graph topology while overlooking gate functionality, which often fails to capture underlying semantic, thus limiting their generalizability. To address the concern, we propose a novel netlist representation learning framework that utilizes a contrastive scheme to acquire generic functional knowledge from netlists effectively. We also propose a customized graph neural network (GNN) architecture that learns a set of independent aggregators to better cooperate with the above framework. Comprehensive experiments on multiple complex real-world designs demonstrate that our proposed solution significantly outperforms state-of-the-art netlist feature learning flows. Ziyi Wang 0010, Zhuolun He, Guangliang Zhang, Qiang Xu 0001, Tsung-Yi Ho, Bei Yu 0001, Yu Huang 0005 |
DAC | 8 |
| 2022 | LHNN: lattice hypergraph neural network for VLSI congestion predictionabstractPrecise congestion prediction from a placement solution plays a crucial role in circuit placement. This work proposes the lattice hypergraph (LH-graph), a novel graph formulation for circuits, which preserves netlist data during the whole learning process, and enables the congestion information propagated geometrically and topologically. Based on the formulation, we further developed a heterogeneous graph neural network architecture LHNN, jointing the routing demand regression to support the congestion spot classification. LHNN constantly achieves more than 35% improvements compared with U-nets and Pix2Pix on the F1 score. We expect our work shall highlight essential procedures using machine learning for congestion prediction. Bowen Wang 0017, Guibao Shen, Dong Li 0016, Jianye Hao, Wulong Liu, Yu Huang 0005, Hongzhong Wu, Yibo Lin, Guangyong Chen, Pheng-Ann Heng |
DAC | 6 |
| 2022 | RCANet: Root Cause Analysis via Latent Variable Interaction Modeling for Yield ImprovementabstractIdentifying root causes of systematic defects is a crucial step in yield enhancement process of integrated circuit (IC) manufacturing. With increasing complexity of fabrication processes and decreasing sizes of pattern features, more systematic defects occur at advanced technology nodes, and traditional methods are unfeasible to directly identify failure causes, due to expensive time and labor costs. Root cause analysis (RCA) technology is thus studied to automatically identify common root causes in a short time. In this paper, we develop RCANet, an end-to-end unsupervised learning-based RCA framework, which analyses diagnosis reports of failing dies within a wafer and identifies both layout-aware and cell-internal root causes efficiently. Experimental results on designs with different technologies demonstrate that RCANet outperforms both a commercial tool and the state-of-the-art method. Xiaopeng Zhang 0009, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Evangeline F. Y. Young, Pengyun Li, Yu Huang 0005, Jianye Hao |
ITC | 8 |
| 2022 | Neural Fault Analysis for SAT-based ATPGabstractContinued advances in process technology have led to a relentless increase in the design complexity of integrated circuits (ICs). In order to meet the increasing demand of low defective-parts-per-million (DPPM) and high product quality of the complex circuit designs, Boolean Satisfactory (SAT) has worked as a robust alternative to conventional APTG techniques. In SAT-based ATPG, logic cones related to the target faults are transformed to Boolean formulas, and standard SAT solving procedures are then used for solving these formulas. Recently, artificial intelligence (AI) techniques have shown great potential in speeding-up SAT solvers. However, the high diversity of the structural characteristics within the logic cones of target faults limits the AI techniques being used for SAT-based ATPG. To meet this challenge, this paper proposes a neural fault analysis technology that is made up of a multi-stage learning model and the testability classifier to highly increase the SAT-based ATPG solving efficiency. The multi-stage learning model is composed of a generative model with a topology structure discriminator and a conflict structure discriminator. It is trained for high-quality data synthesis. Then the testability classifier is trained for adaptive heuristic selection and effective initialization in SAT-based ATPG. Experimental results on both open-source and industrial circuits demonstrate that the neural fault analysis can reduce the SAT solving time by 34.79% and reduce the runtime of SAT-based ATPG by 7.43% on average. It is also shown that the proposed neural fault analysis can cover 9.14% of the faults failed by the conventional SAT-based ATPG framework with a comparable runtime. Junhua Huang, Hui-Ling Zhen, Naixing Wang, Mingxuan Yuan, Yu Huang 0005 |
ITC | 6 |
| 2022 | DeepTPI: Test Point Insertion with Deep Reinforcement LearningabstractTest point insertion (TPI) is a widely used technique for testability enhancement, especially for logic built-in self-test (LBIST) due to its relatively low fault coverage. In this paper, we propose a novel TPI approach based on deep reinforcement learning (DRL), named DeepTpi. Unlike previous learning-based solutions that formulate the TPI task as a supervised-learning problem, we train a novel DRL agent, instantiated as the combination of a graph neural network (GNN) and a Deep Q-Learning network (DQN), to maximize the test coverage improvement. Specifically, we model circuits as directed graphs and design a graph-based value network to estimate the action values for inserting different test points. The policy of the DRL agent is defined as selecting the action with the maximum value. Moreover, we apply the general node embeddings from a pretrained model to enhance node features, and propose a dedicated testability-aware attention mechanism for the value network. Experimental results on circuits with various scales show that DeepTPI significantly improves test coverage compared to the commercial DFT tool. The code of this work is available at https://github.com/cure-lab/DeepTPI. Zhengyuan Shi, Min Li 0019, Sadaf Khan, Liuzheng Wang, Naixing Wang, Yu Huang 0005, Qiang Xu 0001 |
ITC | 6 |
| 2022 | Compression-Aware ATPGabstractThe remarkable growth of the circuit size and complexity is primarily due to the advances of VLSI design and manufacturing technologies. The on-chip linear sequential test compression has become the de facto industrial mainstream DFT methodology in reducing the overall cost of testing large chips. In this paper, we propose a novel and efficient compression-aware ATPG method to significantly boost the performance of ATPG and reduce pattern count. The proposed approach first analyzes the intrinsic dependency of the equations determined by a linear sequential decompressor to build Maximal Linear Independent Group (MLIG). Next it computes the implied values given ATPG generated test cubes based on MLIG. The implied values can significantly improve the test compaction and reduce the ATPG run time due to the reduction of value conflicts between ATPG and compression. The proposed approach does not affect test coverage, requires no extra hardware support, and can be applied to any linear sequential compression scheme. Experimental results on several industrial designs demonstrate that on average the proposed approach can reduce the number of ATPG patterns by 7.57% and the ATPG run time by 28.4%. Zezhong Wang 0006, Naixing Wang, Yu Huang 0005 |
ITC | 5 |
| 2021 | Enhancements of Model and Method in Lithography Hotspot IdentificationabstractThe manufacturing of integrated circuits (ICs) has been continuously improved through the advancement of fabrication technology nodes. However the lithography hotspots (HSs) caused by optical diffraction problems seriously affect the yield and reliability of ICs. Although lithography simulation can accurately capture the HSs through physically simulating the lithography process, it requires a lot of computing resources, which usually takes > 100 CPU ·$h$/mm2[1]. Due to the image recognition nature, the state-of-the-art HS identification algorithms based on deep learning have obvious advantages in reducing run time comparing to the traditional algorithms. However, its accuracy still needs to be enhanced since there are many false alarms of non-hotspots (NHSs) and escapes of the real HSs, which makes it difficult to be a signoff technique. In this paper, we propose two enhancements in HS identification. First, a hybrid deep learning model is proposed in lithography HS identification, which includes a CNN model to combine physical features. Second, an ensemble learning method is proposed based on multiple submodels. The proposed enhanced model and method can achieve high HS identification accuracy on the benchmarks 1–4 of the ICCAD 2012 dataset with recall> 98.8%. In addition, it can achieve even 100% recall on the benchmark 1 and benchmark 3 while maintaining the precision at a high level with 53.6% and 87.1%, respectively. Moreover, for the first time it can achieve not only 100% recall on benchmark 5, but also high precision of 61.8%, which is much higher than any published deep learning methods for HSs identification, as far as we know. The proposed model and methodology can be applied in industrial IC designs due to its effectiveness and efficiency. Xuanyu Huang, Yu Huang 0005 |
DATE | 3 |
| 2021 | A Unified Framework for Layout Pattern Analysis with Deep Causal EstimationabstractThe decrease of feature size and the growing complexity of the fabrication process lead to more failures in manufacturing semiconductor devices. Therefore, identifying the root cause layout patterns of failures becomes increasingly crucial for yield improvement. In this paper, a novel layout-aware diagnosis-based layout pattern analysis framework is proposed to identify the root cause efficiently. At the first stage of the framework, an encoder network trained using contrastive learning is used to extract representations of layout snippets that are invariant to trivial transformations including shift, rotation, and mirroring, which are then clustered to form layout patterns. At the second stage, we model the causal relationship between any potential root cause layout patterns and the systematic defects by a structural causal model, which is then used to estimate the Average Causal Effect (ACE) of candidate layout patterns on the systematic defect to identify the true root cause. Experimental results on real industrial cases demonstrate that our framework outperforms a commercial tool with higher accuracies and around x8.4 speedup on average. Ran Chen 0001, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Bei Yu 0001, Pengyun Li, Yu Huang 0005, Jianye Hao |
ICCAD | 8 |
| 2021 | Diagnosis and Yield LearningabstractWith the advanced technology used in the semiconductor manufacture process, systematic defects related to layout patterns and litho process are the major cause of yield issues. In this industry session, we invite three experts from three EDA companies to share their experiences of diagnosis and yield learning. Yu Huang 0005, Wu-Tung Cheng, Ruifeng Guo, Sameer Chillarige |
ITC-Asia | 1 |
| 2021 | Automotive Test and ReliabilityabstractAs automobiles become increasingly computerized, the safety requirement of hundreds of ICs used in a car is growing rapidly. Testing of such automotive ICs is more stringently than testing other ICs because the quality, reliability and functional safety of such ICs are extremely important as they are directly related to human lives. Many standards such as ISO 26262 are used to guide for high quality and long-term reliability driven by functional safety requirements. In this industry session, we invite 3 experts from IC design and EDA companies. They share their experiences with many case studies, and focus on different aspects of automotive testing, online monitoring, reliability and in-system functional safety. Yu Huang 0005, David Francis, Yervant Zorian, Nilanjan Mukherjee 0001 |
ITC-Asia | 1 |
| 2021 | Adaptive NN-based Root Cause Analysis in Volume Diagnosis for Yield ImprovementabstractRoot Cause Analysis (RCA) is a critical technology for yield improvement in integrated circuit manufacture. Traditional RCA prefers unsupervised algorithms such as Expectation Maximization based on Bayesian models. However, these methods are severely limited by the weak predictive capability of statistical models and can’t effectively transfer the yield learning experience from old designs and processes to the new ones. Motivated by recent advancements of deep learning, in this paper we propose a Neural-Network-based adaptive framework for RCA in yield improvement. The proposed framework consists of an inference module and a self-adaptive module. The former receives volume diagnosis reports and predicts the root cause distributions. The latter is able to adapt the inference module to new designs and processes based on a few of targeted samples without any manual adjustment. Experimental results show that a relatively large improvement on accuracy is achieved by the proposed framework on simulated diagnosis data. Furthermore, the transferring capability of the self-adaptive module is also validated by the results. Ruosheng Xu, Shangling Jui, Zhihao Ding, Pengyun Li, Yu Huang 0005 |
ITC | 8 |
| 2021 | Testability-Aware Low Power Controller Design with Evolutionary LearningabstractXORNet-based low power controller is a popular technique to reduce circuit transitions in scan-based testing. However, existing solutions construct the XORNet evenly for scan chain control, and it may result in sub-optimal solutions without any design guidance. In this paper, we propose a novel testability-aware low power controller with evolutionary learning. The XORNet generated from the proposed genetic algorithm (GA) enables adaptive control for scan chains according to their usages, thereby significantly improving XORNet encoding capacity, reducing the number of failure cases with ATPG and decreasing test data volume. Experimental results indicate that under the same control bits, our GA-guided XORNet design can improve the fault coverage by up to 2.11%. The proposed GA-guided XORNets also allows reducing the number of control bits, and the total testing time decreases by 20.78% on average and up to 47.09% compared to the existing design without sacrificing test coverage. Min Li 0019, Zhengyuan Shi, Zezhong Wang 0006, Yu Huang 0005, Qiang Xu 0001 |
ITC | 5 |
| 2021 | Special Session - Test for AI Chips: from DFT to On-line TestingabstractThis special session focuses on test for artificial intelligence (AI) chips, with important issues from design for test (DFT) to on-line testing. The first talk discusses different DFT implementations and their tradeoffs as well as test access and configuration infrastructure for AI chips with many cores. The second talk discusses low-cost on-line fault detection and hardware salvaging techniques for neural network processors. The last talk gives case studies for testing industrial AI SOC chips, with an emphasis on the automatic test pattern generation (ATPG) methodology. Huawei Li 0001, Xiaowei Li 0001, Yu Huang 0005, Ying Wang 0001, Gary Guo |
VTS | 3 |
| 2020 | Efficient Prognostication of Pattern Count with Different Input Compression RatiosabstractA novel method to efficiently and accurately prognosticate the pattern count at different input compression ratios with the Embedded Deterministic Test (EDT) compression technology is proposed. With this method the total ATPG run time can be significantly reduced compared to the currently used trial-and-error method. Fong-Jyun Tsai, Chong-Siao Ye, Yu Huang 0005, Kuen-Jong Lee, Wu-Tung Cheng, Sudhakar M. Reddy, Mark Kassab, Janusz Rajski |
ETS | 3 |
| 2020 | Estimation of Test Data Volume for Scan Architectures with Different Numbers of Input ChannelsabstractOver the past two decades, test data compression has become a de facto technology used in large industrial designs to reduce the overall test cost. During DFT planning, it is very important to understand the impact of using different numbers of input/output channels on test coverage, test cycles, and test data volume. In this paper, an efficient method to estimate the test data volume with different input channel counts using the Embedded Deterministic Test (EDT) compression technology is proposed. The results can then be used to quickly determine the scan configuration that results in the least or near least test data volume. With this method, the total ATPG run time can be reduced by a factor of more than 10X compared to the currently used trial-and-error method. Fong-Jyun Tsai, Chong-Siao Ye, Yu Huang 0005, Kuen-Jong Lee, Wu-Tung Cheng, Sudhakar M. Reddy, Mark Kassab, Janusz Rajski, Shi-Xuan Zheng |
ITC-Asia | 3 |
| 2020 | Fast Bring-Up of an AI SoC through IEEE 1687 Integrating Embedded TAPs and IEEE 1500 InterfacesabstractComplex application specific SoC are being developed for hardware support artificial intelligence (AI) applications. Such a complex SoCs are integrating a large number of on-chip and off-chip memories, numerous cores and interfaces including in our case a hierarchy of embedded TAPs, as well as security measures and Design-For-Test (DFT) structures. In this case-study paper, we demonstrate using IEEE 1687-2014 (IJTAG) to integrate all these different components into a single, unifying methodology. From this, we derive the benefits of workflow efficiency and fast silicon bring-up. For example, we can report that silicon bring-up of the DFT of the entire SoC was completed in about 4 days, and other bring-up aspects of the Soc were also completed in very little time. Haiying Ma, Ligang Lu, Haitao Qian, Fanjin Meng, Rahul Singhal, Martin Keim, Yu Huang 0005 |
ITC | 9 |
| 2020 | Prediction of Test Pattern Count and Test Data Volume for Scan Architectures under Different Input Channel ConfigurationsabstractAs the complexity of industrial integrated circuits continue to increase rapidly, test data compression has now become a de facto technology for large designs to reduce the overall test cost. During the design for test (DFT) planning, it is critical to understand the impact of using different numbers of input/output test channels on test coverage, test cycles, and test data volume. In this paper, two approaches to predict the test pattern counts and test data volumes with different input channel counts are presented, one with the compression tool able to generate channel-scaling patterns and the other without this capability. The results can be used to determine the scan test configuration that results in the smallest or near smallest test data volume. Experiments on industrial circuits show that the average error rates of pattern count prediction for most circuits are less than 10% for both approaches. The error rates of the predicted smallest data volumes are all less than 3.5%. The total ATPG run time can be reduced by a factor of more than 10X compared to the currently used trial-and-error approach. Fong-Jyun Tsai, Chong-Siao Ye, Kuen-Jong Lee, Shi-Xuan Zheng, Yu Huang 0005, Wu-Tung Cheng, Sudhakar M. Reddy, Mark Kassab, Janusz Rajski, Chen Wang 0014, Justyna Zawada |
ITC | 5 |
| 2020 | Effective Design of Layout-Friendly EDT DecompressorabstractThis paper proposes an innovative design methodology for layout-friendly decompressor used in EDT compression architecture. A segmented decompressor architecture is proposed, in which each segment drives a subset of scan chains. The EDT input channel injectors are carefully selected to maximize the encoding capacity for all scan chains. Experimental results with several large industrial designs demonstrate that using the proposed technology, the routing congestion introduced by EDT decompressor is reduced significantly with negligible impact on test coverage and improved pattern count. Yu Huang 0005, Janusz Rajski, Mark Kassab, Nilanjan Mukherjee 0001, Jeffrey Mayer |
VTS | 1 |
| 2020 | Diagnosis of Intermittent Scan Chain Faults Through a Multistage Neural Network Reasoning ProcessabstractDiagnosis of intermittent scan chain failures still remains a hard problem. In this article, we demonstrate that the use of artificial neural networks (ANNs) can lead to significantly higher accuracy. The key of this method is a multistage process incorporating ANNs with gradually refined focuses. During this process, the final fault suspect is elected through multiple rounds of ANN inference, instead of just one round. At each stage, identification of a proper Affine Group, used as the “candidate set of scan cells for the next round of ANN inference,” will influence the final diagnostic accuracy. Thus, we propose a validation-based learning procedure for Affine Group derivation to further boost the final diagnostic accuracy. The experimental results on benchmark circuits have shown that this method is, on the average, 17.46% more accurate than a state-of-the-art commercial tool for intermittent stuck-at-0 faults. Mason Chern, Shih-Wei Lee, Shi-Yu Huang, Yu Huang 0005, Gaurav Veda, Kun-Han Tsai, Wu-Tung Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Low Cost Hypercompression of Test DataabstractThis article presents a next-generation test data compression scheme. It builds on the isometric compression paradigm, but makes it more flexible and elevates encoding efficiency to values unachievable through state-of-the-art sequential compression schemes. Furthermore, its programmable selection of full-toggle scan chains ensures high test coverage and virtually eliminates compression aborts. The presented approach follows from a fundamental observation that among test cube care bits, only a very few have a status of necessary assignments (their locations cannot be changed), whereas the remaining ones have alternative sites. These test cubes are used to form circular test templates which synergistically control a decompressor and guide back ATPG to find assignments yielding highly compressible test patterns. A redesigned low-silicon-area decompressor is also capable of reducing switching rates in scan chains with a new test power control scheme. The experimental results obtained for large industrial designs and other benchmark circuits confirm the superiority of the proposed scheme over existing techniques and are reported herein. Yu Huang 0005, Sylwester Milewski, Janusz Rajski, Jerzy Tyszer, Chen Wang 0014 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Improving scan chain diagnostic accuracy using multi-stage artificial neural networksabstractDiagnosis of intermittent scan chain failures remains a hard problem. We demonstrate that Artificial Neural Networks (ANNs) can be used to achieve significantly higher accuracy. The key is to take on domain knowledge and use a multi-stage process incorporating ANNs with gradually refined focuses. Experimental results on benchmark circuits show that this method is, on average, 20% more accurate than a state-of-the-art commercial tool for intermittent stuck-at faults, and improves the hit rate from 25.3% to 73.9% for some test-case. Mason Chern, Shih-Wei Lee, Shi-Yu Huang, Yu Huang 0005, Gaurav Veda, Kun-Han Tsai, Wu-Tung Cheng |
ASP-DAC | 4 |
| 2019 | Deep Learning Based Test Compression AnalyzerabstractWith the increase in design complexity and test data volume, compressed tests together with on-chip test decompression hardware such as Embedded Deterministic Test (EDTTM) are widely used in industry in order to reduce test cost. One of the challenges of such Design-for-Test (DFT) technology is to determine a set of optimal parameters such as the number of scan chains, scan channels, power budget, etc. such that it can reach the highest test coverage with a minimum amount of test data volume whilst satisfying various other constraints. To achieve the optimal compression configuration quickly, in this work deep learning technology based on Tensorflow is explored to estimate the test coverage and the data volume for a design when employing EDT under a given set of circuit parameters. Based on the estimated data, the optimal test architecture is also predicted, yielding a more efficient approach compared to the currently used trial-and-error methods. To demonstrate the advantages of our deep learning approach over the currently used utility, we present experimental data for eight industrial designs. Cheng-Hung Wu, Yu Huang 0005, Kuen-Jong Lee, Wu-Tung Cheng, Gaurav Veda, Sudhakar M. Reddy, Chun-Cheng Hu, Chong-Siao Ye |
ATS | 2 |
| 2019 | Non-Adaptive Pattern Reordering to Improve Scan Chain Diagnostic ResolutionabstractThe following topics are dealt with: integrated circuit testing; fault diagnosis; integrated circuit design; logic testing; logic design; system-on-chip; design for testability; learning (artificial intelligence); field programmable gate arrays; IEEE standards. Yu Huang 0005, Jakub Janicki, Szczepan Urban |
ETS | 1 |
| 2019 | A Case Study of Testing Strategy for AI SoCabstractRecent advances in artificial intelligence (AI) are becoming a driving force behind the technological revolution and industrial transformation leading to economic and social development. Application specific AI SoCs are being developed at different companies to accelerate processing of the data-intensive AI computations. There are many new challenges in designing and implementing Design-For-Test (DFT) logic for AI SoCs. In this paper, we share our experiences with DFT implementation for our AI SoC. To achieve lower power and higher bandwidth for AI SoC, we use high speed Serdes PHY with lower threshold voltage, which uses many SoC pins. Therefore, it has a negative impact on DFT and ATPG due to lack of reusable IOs that can be used as scan test channels. In this paper, we present our solution and tradeoffs made to optimize DFT silicon area overhead, test cost, test coverage, pre-silicon verification run time with ready-to-use silicon bring-up methodologies. Haiying Ma, Quan Jing, Yu Huang 0005, Rahul Singhal, Fanjin Meng |
ITC-Asia | 5 |
| 2019 | Innovative Practices on DFT for AI ChipsabstractHardware acceleration for Artificial Intelligence (AI) is now a very competitive and rapidly evolving market. As a result, fast time-to-market is a leading concern for this segment. To speed up time-to-market and ensure quality, new design-for-test (DFT) architectures, new DFT methodologies and technologies are emerging. AI chips are typically very big with many identical and non- identical cores, distributed memories, high-speed IOs, which makes testing of such a gigantic SoC a very challenging task. We found it is very important for us to understand these new challenges from the point of views of DFT engineers. In this innovative practice session, we invited three DFT experts from three AI chip companies to share their experiences of DFT on AI chips. Iris Ma, Hui King Lau, Joseph A. Reynick, Yu Huang 0005 |
VTS | 4 |
| 2018 | Industrial Case Studies of SoC Test Scheduling Optimization by Selecting Appropriate EDT ArchitecturesabstractModern large system-on-chip (SoC) designs typically have hundreds of cores. Each core requires a certain number of input/output test channels. At the chip level, however, the total number of test pins is limited such that all core-level test channels cannot be accessed at the same time. Therefore, hierarchical pattern retargeting is required for SoC test. Test scheduling algorithms can be applied to reduce the total test time. In this paper, we add the EDT architectures as one additional dimension of parameter into the prior test scheduling algorithm. Experimental results based on real case studies show that with the proposed flow, the test time can be further reduced up to approximately 24%. Guoliang Li 0004, Henry Zhao, Qinfu Yang, Yu Huang 0005 |
ITC-Asia | 5 |
| 2018 | Hypercompression of Test PatternsabstractThe paper presents a novel test data compression scheme. This low-silicon-area solution builds on the isometric compression paradigm, but makes it more flexible, elevates encoding efficiency to values unachievable through any conventional type of sequential compression, and ensures high test coverage due to programmable selection of full toggle scan chains. The presented approach follows from a fundamental observation that only a few specified positions in test cubes are necessary to detect faults, while the remaining ones have alternative sites. Such test cubes are used to form circular test templates which synergistically control a decompressor and guide ATPG to find assignments yielding highly compressible test cubes. A redesigned decompressor is also capable of reducing switching rates in scan chains with a new test power control scheme. Experimental results obtained for large industrial designs confirm superiority of the proposed scheme over state-of-the-art techniques and are reported herein. Yu Huang 0005, Sylwester Milewski, Janusz Rajski, Jerzy Tyszer, Chen Wang 0014 |
ITC | 1 |
| 2018 | Special session on machine learning for test and diagnosisabstractThe special session focuses on using Machine Learning (ML) techniques on different applications in test and diagnosis. The first contribution discusses how to close the gap between working silicon and a working system by using ML. The second presentation then talks an alternative ML view and its various applications such as functional verification, Fmax prediction, and production yield optimization. The last presentation discusses using supervised ML on volume diagnosis to further improve the accuracy of identifying root causes. Krishnendu Chakrabarty, Li-C. Wang, Gaurav Veda, Yu Huang 0005 |
VTS | 4 |
| 2018 | Innovative practices on machine learning for emerging applicationsabstractThe IP session focuses on using Machine Learning (ML) techniques on several emerging applications. The first contribution discusses hotspot detection by using ML. The second presentation then talks a data-driven health monitoring solution. The last contribution discusses using ML to emulate hardware Trojans. Kareem Madkour, Zhaobo Zhang, Alfred L. Crouch, Peter L. Levin, Eve Hunter, Yu Huang 0005 |
VTS | 6 |
| 2017 | Scan Chain Diagnosis Based on Unsupervised Machine LearningabstractScan-based testing has proven to be a cost-effective method for achieving good test coverage in digital circuits. It was reported in prior papers that about 30% to 50% of all failing die were due to defects that cause scan chains to fail [1][2]. Therefore, scan chain failure diagnosis is very important to improve yield. The previously proposed methods of chain diagnosis were primarily based on either deterministic fault models and simulation algorithms or probabilistic analysis. To handle hard-to-model defect behaviors more robustly, in this paper, we propose a new scan chain diagnosis algorithm based on unsupervised machine learning. Its application on "scannable memory designs" (SMD) is demonstrated to illustrate the effectiveness of the proposed algorithm. Yu Huang 0005, Brady Benware, Randy Klingenberg, Huaxing Tang, Jayant Dsouza, Wu-Tung Cheng |
ATS | 1 |
| 2015 | Advancements in diagnosis driven yield analysis (DDYA): A survey of state-of-the-art scan diagnosis and yield analysis technologiesabstractIn this paper, we surveyed the recent advancements in DDYA, which includes scan-based diagnosis technologies and diagnosis driven yields analysis. Multiple industrial cases studies are given to illustrate the values of the DDYA flow. By using the advanced DDYA, diagnosis can be more accurate, fast and informative. More importantly it can help improve yield by finding the systematic defect, identifying the root causes, correlating diagnosis results with DFM and picking highly possible die and suspect for PFA to validate all the findings. Yu Huang 0005, Wu-Tung Cheng |
ETS | 1 |
| 2015 | Scan Test Bandwidth Management for Ultralarge-Scale System-on-Chip ArchitecturesabstractThis paper presents several techniques employed to resolve problems surfacing when applying scan bandwidth management to large industrial multicore system-on-chip (SoC) designs with embedded test data compression. These designs pose significant challenges to the channel management scheme, flow, and tools. This paper introduces several test logic architectures that facilitate preemptive test scheduling for SoC circuits with embedded deterministic test-based test data compression. The same solutions allow efficient handling of physical constraints in realistic applications. Finally, state-of-the-art SoC test scheduling algorithms are rearchitected accordingly by making provisions for: 1) setting up time-effective test configurations; 2) optimization of SoC pin partitions; 3) allocation of core-level channels based on scan data volume; and 4) more flexible core-wise usage of automatic test equipment channel resources. A detailed case study is illustrated herein with a variety of experiments allowing one to learn how to tradeoff different architectures and test-related factors. Wu-Tung Cheng, Grady Giles, Yu Huang 0005, Jakub Janicki, Mark Kassab, Grzegorz Mrugalski, Nilanjan Mukherjee 0001, Janusz Rajski, Jerzy Tyszer |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Diagnosis and Layout Aware (DLA) Scan Chain StitchingabstractWithout appropriate stitching of scan chains, even with good diagnosis algorithm and diagnostic pattern generation, the chain diagnostic resolution may still be bad. In this paper, we propose a novel pattern-independent diagnosis and layout aware (DLA) scan chain stitching method: 1) the resolution is improved by increasing and properly distributing the sensitive scan cells, which can capture useful diagnostic information under both single- and multiple-fault situations; and 2) the scan cell layout placement is taken into account to reduce routing overhead and hence preserve the chip performance. Experiments using two different techniques to diagnose ISCAS'89/ITC'99 benchmark circuits with/without embedded scan compaction show the effectiveness of the proposed method in improving the diagnostic resolution. Impacts on chip performance, embedded scan compaction, transition fault coverage, and test power dissipation are negligible. The proposed method is also successfully applied to an industry circuit manufactured with 20-nm technology. The silicon results show 7× average resolution improvement comparing to without using the DLA scan chain stitching. Jing Ye 0001, Yu Huang 0005, Yu Hu 0001, Wu-Tung Cheng, Ruifeng Guo, Liyang Lai, Ting-Pu Tai, Xiaowei Li 0001, Wei-pin Changchien, Daw-Ming Lee, Ji-Jan Chen, Sandeep C. Eruvathi, Kartik K. Kumara, Charles C. C. Liu, Sam Pan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Diagnose Failures Caused by Multiple Locations at a TimeabstractFault diagnosis plays an important role in physical failure analysis and yield learning process. With tens of billions of transistors being integrated in one chip, multiple faults may exist. With multiple faults, fault masking and reinforcing effects may appear. They may cause the conventional single-fault-based diagnosis methods such as the single location at a time (SLAT) to be invalid. The popular SLAT approach fails if there are not enough SLAT patterns that can be explained by a single stuck-at fault. Moreover, a real silicon defect may behave as different fault models (DM) under different failing patterns, which may invalidate the SLAT approach that uses a single-fault model across all failing patterns. In this paper, we introduce the concept of fault element to support multiple fault models, and use a fault-element graph (FEG) to consider fault masking and reinforcing effects among multiple faults. Based on the FEGs of all failing patterns, the most likely fault locations and their fault elements are iteratively identified. Meanwhile, the FEGs are iteratively pruned to keep track of the remaining multiple fault effects until all the fault locations are identified and all the FEGs are reduced to null. Experiments demonstrate that the proposed diagnosis method can identify the locations of multiple faults even under DM with high diagnostic accuracy and resolution. Jing Ye 0001, Yu Hu 0001, Xiaowei Li 0001, Wu-Tung Cheng, Yu Huang 0005, Huaxing Tang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2013 | EDT bandwidth management - Practical scenarios for large SoC designsabstractThe paper discusses practical issues involved in applying scan bandwidth management to large industrial system-on-chip (SoC) designs deploying embedded test data compression. These designs pose significant challenges to the channel bandwidth management methodology itself, flow, and tools. The paper introduces several test logic architectures that facilitate preemptive test scheduling for SoC circuits with EDT-based test data compression. Moreover, some recently proposed SoC test scheduling algorithms are refined accordingly by making provision for (1) setting up test configurations minimizing test time, (2) optimization of SoC pin allocation based on scan data volume, and (3) handling physical constraints in realistic applications. Detailed presentation of a case study is illustrated with a variety of experiments that allow one to learn how to tradeoff different architectures and test scheduling. Jakub Janicki, Jerzy Tyszer, Wu-Tung Cheng, Yu Huang 0005, Mark Kassab, Nilanjan Mukherjee 0001, Janusz Rajski, Grady Giles |
ITC | 4 |
| 2013 | Diagnosis and Layout Aware (DLA) scan chain stitchingabstractWithout appropriate stitching of scan chains, even with good diagnosis algorithm and diagnostic pattern generation, it may still result in bad scan chain diagnostic resolution. To improve the diagnostic resolution, we propose a novel Diagnosis and Layout Aware (DLA) scan chain stitching method, which is pattern independent and supports embedded scan compaction. It is based on three ideas: (1) increasing the number of sensitive scan cells, which can capture useful diagnostic information; (2) properly distributing the sensitive scan cells along the scan chains to enhance the overall resolution; (3) stitching scan cells based on their placement at layout to preserve the chip performance. Experiments on ISCAS'89/ITC'99 benchmark circuits and a real industry circuit based on 20nm technology with silicon results show that, the proposed DLA scan chain stitching method effectively improves the resolution, with negligible impact on chip performance, embedded scan compaction, transition fault coverage, and test power dissipation. The silicon results even show 7X average resolution improvement comparing to without using the proposed method. Jing Ye 0001, Yu Huang 0005, Yu Hu 0001, Wu-Tung Cheng, Ruifeng Guo, Liyang Lai, Ting-Pu Tai, Xiaowei Li 0001, Wei-pin Changchien, Daw-Ming Lee, Ji-Jan Chen, Sandeep C. Eruvathi, Kartik K. Kumara, Charles C. C. Liu, Sam Pan |
ITC | 2 |
| 2013 | Distributed dynamic partitioning based diagnosis of scan chainabstractDiagnosis memory footprint for large designs is growing as design sizes grow such that the diagnosis throughput for given computational resources becomes a bottleneck in volume diagnosis. In this paper, we propose a scan chain diagnosis flow based on dynamic design partitioning and distributed diagnosis architecture that can improve the diagnosis throughput over one order of magnitude. Yu Huang 0005, Xiaoxin Fan, Huaxing Tang, Wu-Tung Cheng, Brady Benware, Sudhakar M. Reddy |
VTS | 1 |
| 2012 | A Hybrid Flow for Memory Failure Bitmap ClassificationabstractFailure bitmaps of manufactured memory arrays may contain the information of some systematic defects and have hence been used to monitor the process and to improve the memory yield. It is important to have an accurate flow to classify the memory failure bitmap signatures. The memory bitmap signature classification can be either dictionary based or machine learning based. This paper introduces a hybrid flow that can combine dictionary based and machine learning based methods. The proposed method can enhance the accuracy of signature classification, and more importantly, it has the capability of learning new memory bitmap signatures unseen before. Yu Huang 0005, Wu-Tung Cheng, Chris Schuermyer, Eric Faehn, Ruth Farrugia |
Asian Test Symposium | 2 |
| 2012 | Improved volume diagnosis throughput using dynamic design partitioningabstractA method based on dynamic design partition is presented to increase the throughput of volume diagnosis by increasing the number of failing dies diagnosed within a given time T using given constrained computational resources C. Recently we proposed a static design partitioning method to reduce the diagnosis memory footprint for large designs [1] to achieve this objective. The method in [1] is applied once for each design without using the information of test patterns and failure files, and then diagnosis is performed on an appropriate block(s) of the design partition for a failure file. Even though the memory footprint of diagnosis is reduced the diagnosis quality is impacted to unacceptable levels for some types of defects such as bridges. In this paper, we propose a new failure dependent design partitioning method to improve volume diagnosis throughput with a minimal impact on diagnosis quality. For each failure file, the proposed method first determines the small partition needed to diagnose this failure, and then performs the diagnosis on this partition instead of the complete design. Since the partition is far smaller, both the run time and the memory usage of diagnosis can be significantly reduced better than when earlier proposed static partition is used. Extensive experiments were conducted on several large industrial designs to validate the proposed method. It has been observed that the typical partition size for various defects is less than 3% of the size of the original design. Also diagnosis runs much faster (>;2X) on the partition. Combining these two factors, the throughput of volume diagnosis can be improved by an order of magnitude. Xiaoxin Fan, Huaxing Tang, Yu Huang 0005, Wu-Tung Cheng, Sudhakar M. Reddy, Brady Benware |
ITC | 3 |
| 2012 | Addressing Visual Consistency in Video Retargeting: A Refined Homogeneous ApproachabstractFor the video retargeting problem which adjusts video content into a smaller display device, it is not clear how to balance the three conflicting design objectives: 1) visual interestingness preservation; 2) temporal retargeting consistency; and 3) nondeformation. To understand their perceptual importance, we first identify that the latter two play a dominating role in making the retargeting results appealing. Then a statistical study on human response to the targeting scale is carried out, suggesting that the global preservation of contents pursued by most existing approaches is not necessary. Based on the newly prioritized objectives and the statistical findings, we design a video retargeting system which, as a refined homogeneous approach, addresses the temporal consistency issue holistically and is still capable of preserving high degree of visual interestingness. In particular, we propose a volume retargeting cost metric to jointly consider the retargeting objectives and formulate video retargeting as an optimization problem in graph representation. A dynamic programming solution is then given. In addition, we introduce a nonlinear fusion based attention model to measure the visual interestingness distribution. The experiment results from both image rendering and subjective tests indicate that our proposed attention modeling and video retargeting system outperform their conventional methods, respectively. Zheng Yuan 0001, Taoran Lu, Yu Huang 0005, Dapeng Oliver Wu, Hong Heather Yu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Virtual Circuit Model for Low Power Scan Testing in Linear Decompressor-Based Compression EnvironmentabstractLarge test data volume and high test power consumption are two major concerns for the industry when testing large integrated circuits. Linear decompress or-based compression (LDC) is efficient in reducing test data volume, while X-filling during ATPG can efficiently reduce test power with low overhead. However, traditional X-filling methods cannot be reused in the LDC environment. In this paper, we propose a virtual circuit model to make the linear de-compressor transparent to the external testing. As a result, existing X-filling methods can be reused to reduce test power. Experimental results on benchmark circuits demonstrate the efficiency of the proposed approach. Jia Li 0022, Yu Huang 0005 |
Asian Test Symposium | 4 |
| 2011 | Virtual ads insertion in street building views for augmented realityabstractThis paper is proposing a new framework of virtual ads insertion for Augmented Reality (AR) in a specific real scene, building facades of the street view videos. In this framework, specific region detection, camera orientation estimation, visual tracking, motion filtering and virtual-real blending are discussed. For the building façade, vanishing points are fully identified to detect a rectangular planar structure. Meanwhile, dynamic registration and color harmonization are designed to improve visual acuity of virtual ads insertion. Experimental results are given to show the good performance of the proposed framework. Yu Huang 0005, Hong Heather Yu |
ICIP | 1 |
| 2011 | Video summarization with semantic concept preservationabstractA compelling video summarization should allow viewers to understand the summary content and recover the original plot correctly. To this end, we materialize the abstract elements that are cognitively informative for viewers as concepts. They implicitly convey the semantic structure and are instantiated by semantically redundant instances. Then we analyze that a good summary should i) keep various concepts complete and balanced so as to give viewers comparable cognitive clues from a complete perspective ii) pursue the most saliency so that the rendered summary is attractive to human perception. We then formulate video summarization as an integer programming problem and give a ranking based solution. We also propose a novel method to discover the latent concepts by spectral clustering of bag-of-words features. Experiment results on human evaluation scores demonstrate that our summarization approach performs well in terms of the informativeness, enjoyability and scalibility. Zheng Yuan 0001, Taoran Lu, Dapeng Oliver Wu, Yu Huang 0005, Hong Heather Yu |
MUM | 4 |
| 2011 | Prediction of compression bound and optimization of compression architecture for linear decompression-based schemesabstractOn-chip linear decompression-based schemes have been widely adopted by industrial circuits nowadays to effectively reduce the ever increasing test data volume and test time. Though they can easily achieve relatively high compression ratio, there is a bound of effective compression ratio for these compression schemes. Prior work tried to address this problem by trying different compression architectures to identify this compression bound. However, they can not predict this compression bound efficiently. In this paper, we will first analyze the correlation between the effective compression ratio and the compression architecture, thus to predict that compression bound efficiently. In addition, this paper will also propose how to design the compression architecture for target effective compression ratio with one-pass calculation, which was usually done by a time-consuming try-and-error process as well in the current DFT flow. Experimental results show the accuracy of the prediction and the effectiveness of the compression architecture design. Jia Li 0022, Yu Huang 0005 |
VTS | 2 |
| 2010 | Emulating and diagnosing IR-drop by using dynamic SDFabstractThe Standard Delay Format (SDF) information is very important in timing-aware simulation of VLSI designs. However, conventionally, SDF is only design-dependent, but pattern-independent, which is called static SDF in this paper. Static SDF ignores all dynamic pattern dependent parameters, such as IR drop and crosstalk. In this paper, we propose a novel pattern-dependent SDF (called dynamic SDF) generation technique, and apply it to take IR-drop effects into consideration. With the proposed IR-drop-aware SDF generation technique, we improve the accuracy of simulation, and perform diagnosis on the failed patterns to pin point the pattern-dependent IR-drop defects in our design. Experimental results demonstrate the efficiency of this method when used for transition delay fault pattern application and diagnosis. Yu Huang 0005, Ruifeng Guo, Wu-Tung Cheng, Mark Tehranipoor |
ASP-DAC | 2 |
| 2010 | Enhance Profiling-Based Scan Chain Diagnosis by Pattern MaskingabstractIn prior work, profiling-based chain diagnosis methodologies were developed. This paper discusses the challenges associated with profiling-based chain diagnosis in the production test environment with limited failure buffer capacity on tester. We propose the following two pattern masking application flows to enhance diagnosis resolution in this scenario: (1) generic pattern masking and (2) adaptive pattern masking. Experimental results illustrate that pattern masking will enable profiling-based scan chain diagnosis in volume diagnosis environment. Wu-Tung Cheng, Yu Huang 0005 |
Asian Test Symposium | 2 |
| 2010 | Full-circuit SPICE simulation based validation of dynamic delay estimationabstractPower supply noise may have big impacts on the design performance in the latest technologies. Accurately mapping the IR-drop effect to real delay is a challenging task, which will directly impact the accuracy of IR-drop related performance evaluation, test, and diagnosis. In this paper, we first present our previous work on setting up an IR2Delay database for addressing this issue. We then propose a flow to validate this database by comparing it with full-circuit SPICE simulation results. In this flow, mixed-signal simulation is used, which can reuse the existing digital testbench as stimuli while maintaining the accuracy of SPICE simulation. Yu Huang 0005, Pinki Mallick, Wu-Tung Cheng, Mark Tehranipoor |
ETS | 2 |
| 2010 | Video retargeting with nonlinear spatial-temporal saliency fusionabstractVideo retargeting (resolution adaptation) is a challenging problem for its highly subjective nature. In this paper, a nonlinear saliency fusing approach, that considers human perceptual characteristics for automatic video retargeting, is being proposed. First, we incorporate features from phase spectrum of quaternion Fourier Transform (PQFT) in spatial domain and global motion residual based on matched feature points by the Kanade-Lucas-Tomasi (KLT) tracker in temporal domain. In addition, under a cropping-and-scaling retargeting framework, we propose content-aware information loss metrics and a hierarchical search to find optimal cropping window parameters. Results show the success of our approach on detecting saliency regions and retargeting on images and videos. Taoran Lu, Zheng Yuan 0001, Yu Huang 0005, Dapeng Oliver Wu, Hong Heather Yu |
ICIP | 3 |
| 2010 | Video retargeting: A visual-friendly dynamic programming approachabstractVideo retargeting is the task of fitting standard-sized video into arbitrary screen. A compelling retargeting attempts to preserve most visual information of original video as well as deliver a temporally consistent retargeted view. To handle long video sequences, we perform the task on a shot/subshot basis. For each frame, a crop pane is determined to optimally select a region of interest as the retargeted frame in two stages: i.e. minimizing visual information loss (intra-frame consideration) to yield source and destination crop pane parameters at boundary frames and minimizing visual information loss accumulation under the visual inertness (inter-frame consideration) constraints to search for a smooth transition of crop pane across interior frames. The second minimization process is remodeled as the shortest-path problem in graph theory and the parametric transition of crop panes is solved by dynamic programming. Experiments demonstrate our approach preserves salient regions of original video whilst offering eye-friendly visual consistency. Zheng Yuan 0001, Taoran Lu, Yu Huang 0005, Dapeng Oliver Wu, Hong Heather Yu |
ICIP | 3 |
| 2010 | Case study of scan chain diagnosis and PFA on a low yield waferabstractIn this poster, we share our industrial experiences on running chain diagnosis and PFA (Physical Failure Analysis) on a wafer that suffered from low yield. In addition, case study on PFA will be illustrated. Yu Huang 0005, Brady Benware, Wu-Tung Cheng, Ting-Pu Tai, Feng-Ming Kuo, Yuan-Shih Chen |
ITC | 1 |
| 2010 | Test cycle power optimization for scan-based designsabstractExtraordinary power consumption during the scan test may inadvertently cause a functional good die to fail. This paper proposes a peak power reduction algorithm for the scan test which considers both the shift cycles and capture cycles simultaneously to limit the peak power of all test cycles during the test generation. In addition, the analysis also recommends the types of circuit structures that are more suitable to add test logic for maximum power reduction with the minimum test cost. The proposed methodology is highly efficient and can be applied to large industrial designs. Kun-Han Tsai, Yu Huang 0005, Wu-Tung Cheng, Ting-Pu Tai, Augusli Kifli |
ITC | 2 |
| 2010 | On Reducing Scan Shift Activity at RTLabstractPower dissipation in digital circuits during scan-based test is generally much higher than that during functional operation. Unfortunately, this increased test power can create hot spots that may damage the silicon, the bonding wires, and even the package. It can also cause intensive erosion of conductors-severely decreasing the reliability of a device. Finally, excessive test power may also result in extra yield loss. To address these issues, this paper first presents a detailed investigation of a benchmark circuit's switching activity during different modes of operation. Specifically, the average number of transitions in the combinational logic of a benchmark circuit during scan shift is found to be approximately 2.5 times more than the average number of transitions during the circuit's normal functional operation. A DFT-based approach for reducing circuit switching activity during scan shift is proposed. Instead of inserting additional logic at the gate level that may introduce additional delay on critical paths, the proposed method modifies the design at the register transfer level (RTL) and uses the synthesis tools to automatically deal with timing analysis and optimization. Our experiments show that significant power reduction can be achieved with very low overhead. Elif Alpaslan, Yu Huang 0005, Xijiang Lin, Wu-Tung Cheng, Jennifer Dworak |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Scan Chain Diagnosis by Adaptive Signal Profiling with Manufacturing ATPG PatternsabstractIn the past, software based scan chain defect diagnosis can be roughly classified into two categories (1) model-based algorithms, and (2) data-driven algorithms. In this paper we first analyze the advantages and disadvantages of each category of the chain diagnosis algorithms. Next, an adaptive signal profiling algorithm that can use manufacturing ATPG scan patterns is proposed for scan chain diagnosis. Finally, several case studies and their PFA results are presented to validate the accuracy and effectiveness of the proposed algorithm. Yu Huang 0005, Wu-Tung Cheng, Ruifeng Guo, Ting-Pu Tai, Feng-Ming Kuo, Yuan-Shih Chen |
Asian Test Symposium | 1 |
| 2009 | On Improving Diagnostic Test Generation for Scan Chain FailuresabstractIn this paper, we present test generation procedures to improve scan chain failure diagnosis. The proposed test generation procedures improve diagnostic resolution by using multi-cycle scan test patterns. A diagnostic test generation flow to speed up diagnosis is proposed to address the issue of long run times of test generation and large number of test patterns for the cases where the range of suspected cells is large. Experimental results on several industrial designs show the effectiveness of the proposed procedures in improving diagnostic resolution, reducing run times of test generation and also reducing the number of test patterns. Ruifeng Guo, Wu-Tung Cheng, Sudhakar M. Reddy, Yu Huang 0005 |
Asian Test Symposium | 5 |
| 2008 | Observation Point Oriented Deterministic Diagnosis Pattern Generation (DDPG) for Chain DiagnosisabstractScan is a widely used Design-for-Testability technique to improve test and diagnosis quality. Many defects may cause scan chains to fail. In this paper, an observation point oriented Deterministic Diagnostic Pattern Generation (DDPG) method was proposed for compound defects, which tolerates the system defects during scan chain diagnosis. Instead of sensitizing multiple paths proposed in our prior work, the proposed new DDPG method directly targets as many observation points as possible to observe the loading error occurred on the targeted scan cell. Experimental results on ISCASpsila89 benchmark circuits show that the proposed DDPG method improves the effectiveness and efficiency of diagnosing compound defects, compared to our prior research. Yu Hu 0001, Yu Huang 0005, Jing Ye 0001, Xiaowei Li 0001 |
ATS | 3 |
| 2008 | Diagnose Multiple Stuck-at Scan Chain FaultsabstractPrior effect-cause based chain diagnosis algorithms suffer from accuracy and performance problems when multiple stuck-at faults exist on the same scan chain. In this paper, we propose new chain diagnosis algorithms based on dominant fault pair to enhance diagnosis accuracy and efficiency. Several heuristic techniques are proposed, which include (1) double candidate range calculation, (2) dynamic learning and (3) two- dimensional space linear search. The experimental results illustrate the effectiveness and efficiency of the proposed chain diagnosis algorithms. Yu Huang 0005, Wu-Tung Cheng, Ruifeng Guo |
ETS | 1 |
| 2008 | Detection and Diagnosis of Static Scan Cell Internal DefectabstractIn this paper, we study the impact, detection and diagnosis of the defect inside a scan cell, which is called scan cell internal defect. We first use SPICE simulation to understand how a scan cell internal defect impacts the operation of a single scan cell. To study the detectability and diagnosability of a scan cell internal defect in a production test environment, we inject scan cell internal defects into a scan-based industrial design and perform fault simulation by using production scan test patterns. Next, we evaluate how effective an existing scan chain diagnosis technique based on traditional fault models can diagnose scan cell internal defect. We finally propose a new diagnosis algorithm to improve scan cell internal defect diagnostic resolution using scan cell internal fault model. Experimental results show the effectiveness of the proposed scan cell internal fault diagnosis technique. Ruifeng Guo, Liyang Lai, Yu Huang 0005, Wu-Tung Cheng |
ITC | 3 |
| 2008 | Deterministic Diagnostic Pattern Generation (DDPG) for Compound DefectsabstractScan chain failure diagnosis has become an important means for silicon debug and yield improvement. Although plenty of prior work discussed how to perform scan chain diagnosis, most of the previously proposed techniques made an assumption that the system logic is fault-free, which could be an impractical assumption leading to incorrect diagnostic results. In this paper, we propose a scan chain deterministic diagnostic pattern generation (DDPG) method that can tolerate the faults in the system logic without degradation of chain diagnostic resolution and precision. The entire flow includes three steps. In the first step, patterns are created to propagate the state of a targeted scan cell to as many reliable observation points as possible. In the second step, the load error probability of each targeted scan cell is calculated based on the hamming distances between the observed responses and the expected good or faulty responses. In the last step, a suspect profile is plotted, which can be used to identify the suspect scan cell(s) based on ranking scores. Experimental results show that the diagnostic resolution and precision are not degraded even with dozens of faults injected into the system logic. Yu Hu 0001, Huawei Li 0001, Xiaowei Li 0001, Jing Ye 0001, Yu Huang 0005 |
ITC | 6 |
| 2008 | Reducing Scan Shift Power at RTLabstractPower consumption during scan-based test becomes a concern in nanometer technologies. Previous test power reduction techniques that insert additional logic in gate-level circuits may result in timing violations. In this paper, we show that the problem can be solved at the RTL instead so that the timing and area constraints will be handled automatically by synthesis tools. Using a signal probabilistic approach proposed previously, we identify power-sensitive scan cells at the prototyping gate level, and we map these cells to their corresponding signal/variable bits at the RT-level. Additional RTL code is added to freeze these power- sensitive bits in order to reduce scan shift power consumption. Experimental results on ITC99 benchmarks show that on average more than 22% power reduction can be achieved when we only freeze the top 1% of power-sensitive bits at RTL. The flow is more practical in terms of timing closure than doing the same at the gate-level. Elif Alpaslan, Yu Huang 0005, Xijiang Lin, Wu-Tung Cheng, Jennifer Dworak |
VTS | 2 |
| 2008 | Scan Shift Power Reduction by Freezing Power Sensitive Scan Cells
Xijiang Lin, Yu Huang 0005 |
J. Electron. Test. | 2 |
| 2007 | Fault Dictionary Based Scan Chain Failure DiagnosisabstractIn this paper, we present a fault dictionary based scan chain failure diagnosis technique. We first describe a technique to create small dictionaries for scan chain faults by storing differential signatures. Based on the differential signatures stored in a fault dictionary, we can quickly identify single stuck-at fault or timing fault in a faulty chain. We further develop a novel technique to diagnose some multiple stuck-at faults in a single scan chain. Comparing with fault simulation based diagnosis technique, the proposed fault dictionary based diagnosis technique is up to 130 times faster with same level of diagnosis accuracy and resolution. Ruifeng Guo, Yu Huang 0005, Wu-Tung Cheng |
ATS | 2 |
| 2007 | Programmable Logic BIST for At-speed TestabstractIn this paper, we propose a novel programmable logic BIST controller that can facilitate at-speed test for the design with multiple clock domains and multiple clock frequencies. Moreover, a static analysis method is also proposed to optimize the BIST test pattern allocation for testing the timing faults in different intra/inter clock domains when the maximum number of applied BIST test patterns is specified. Experimental results show the effectiveness of the proposed method on achieving higher test coverage than the method with test patterns evenly distributed among different test sessions. Yu Huang 0005, Xijiang Lin |
ATS | 1 |
| 2007 | A RTL Testability Analyzer Based on Logical Virtual PrototypingabstractIn this paper, we propose a novel RTL testability analyzer based on logical virtual prototyping. A very fast synthesis engine is utilized to create a gate level hierarchical netlist with generic gates, which we call a logical virtual prototype, in this paper. Subsequently, ATPG and testability analysis are performed on the logical virtual prototype to provide RTL designers with a wealth of information that would allow them: (1) Accurately estimate / predict test coverage. (2) Generate patterns that can be used for gate-level design. (3) Identify hard-to-test design blocks in RTL, which can then be redesigned to improve coverage. Yu Huang 0005, Nilanjan Mukherjee 0001, Wu-Tung Cheng, Greg Aldrich |
ATS | 1 |
| 2007 | Effect of IR-Drop on Path Delay Testing Using Statistical AnalysisabstractIR-drop has become a major source of delay defects in deep sub-micron VLSI designs. In this work, we analyze the effect of IR-drop in path-delay test and how to obtain more accurate delay information of critical paths. For possible regions with IR-drop, we perform timing analysis on these nodes such that a certain amount of voltage drop can be associated with extra delays on victim nodes. Power analysis is conducted to determine the occurrence probability of a certain voltage drop. These probability values are used to weigh the extra delays caused by IR-drop of all victim nodes, which are then accumulated along each path. Experimental results show that such a process can effectively take the small delays caused by IR-drop into consideration and can have a significant impact on the identification and analysis of critical paths. Yu Huang 0005 |
ATS | 3 |
| 2007 | Scan Diagnosis and Its Successful Industrial ApplicationsabstractAs technologies move below 130nm, the IC industry has seen a significant change in defect type encountered. Feature-related defects are becoming more prevalent than particle-driven defects in nanometer designs. The process and design variances require checks for the design-for-manufacturing (DFM) issues in order to achieve a high yield. Scan diagnosis targeted for the nanometer designs can provide quick, accurate and reliable failure information from the production environment. The ranked, fault classified and physically linked scan diagnosis results can, in turn, provide the guides for DFM checks. High volume diagnosis provides data to yield management system for statistical analysis. This presentation briefly explains the technology behind the scene, discusses scan diagnosis applications and shows the results. Wu-Tung Cheng, Yu Huang 0005, Martin Keim, Randy Klingenberg |
ATS | 3 |
| 2007 | Dynamic learning based scan chain diagnosisabstractScan chain defect diagnosis is important to silicon debug and yield enhancement. Traditional simulation-based chain diagnosis algorithms may take long run time if a large number of simulations are required. In this paper, a novel dynamic learning based scan chain diagnosis is proposed to speedup the diagnosis run time. Experimental results illustrate that by using the proposed dynamic learning techniques, the diagnosis run time can be reduced about 10X on average Yu Huang 0005 |
DATE | 1 |
| 2007 | A complete test set to diagnose scan chain failuresabstractIn this paper, we present a test generation algorithm to improve scan chain failure diagnosis resolution. The proposed test generation algorithm creates a complete test set that guarantees each defective scan cell has unique failing behavior. This algorithm handles stuck-at fault and timing fault models. Problems and solutions that may happen in practical usage are discussed. We further extend the test generation algorithm to handle multiple failing scan chains and designs with embedded scan compression logic. Experimental results show the effectiveness of the proposed diagnostic test generation algorithm. Ruifeng Guo, Yu Huang 0005, Wu-Tung Cheng |
ITC | 2 |
| 2007 | Diagnose compound scan chain and system logic defectsabstractScan based diagnosis can be of great help to guide physical failure analysis, which is critical for the success of silicon debug and yield ramp up. In practice, diagnosis becomes more difficult if scan chain defects and system logic defects co-exist on one die, which are called compound defects in this paper. We first describe the challenges in diagnosing this type of compound defects. A novel diagnosis flow is proposed to diagnose the compound defects on scan chain and system logic. The diagnosis methodology was successfully applied in industrial designs. Yu Huang 0005, Wu-Tung Cheng, Ruifeng Guo, Will Hsu, Yuan-Shih Chen, Albert Mann |
ITC | 1 |
| 2007 | Effects of Embedded Decompression and Compaction Architectures on Side-Channel Attack ResistanceabstractAttack resistance has been a critical concern for security-related applications. Various side-channel attacks can he launched to retrieve security information such as encryption key. Prior work does not consider the presence of embedded compression architectures and their impacts on security. This paper analyzes the complexity of side-channel attacking on designs with embedded decompression and compaction circuit. A possible attacking strategy was first presented and performs analysis on the complexity of attacking circuits with EDT architecture. The probabilistic analysis was then extended to the more general compaction schemes using distance coding. The authors have shown that successful attacking of designs with embedded decompressor and compactor is extremely difficult. The complexity is much higher than the results shown in prior work assuming no decompression and compaction. It indicates that the use of embedded compression architectures can achieve higher security level. Yu Huang 0005 |
VTS | 2 |
| 2006 | Diagnosis of defects on scan enable and clock treesabstractIn this paper, we propose a different software based method that does not require any extra effort manipulating test parameters. First, we assume the defects are on scan chains and use previously published chain diagnosis algorithms presented in Y. Huang et al. (2005) to identify the suspect scan cells. Secondly, if there is at least one faulty chain that is modeled with stuck-at-X fault, we attempt to diagnose with stuck-at-0 fault model at scan enable. As we mentioned earlier that the shift operation is incorrect when the scan cell value is obtained from the system logic for each shift cycle, it is very likely we see both stuck-at-1 and stuck-at-0 at scan cells. So stuck-at-X fault model at scan cells is a sign of the stuck-at-0 fault model for scan enable defects Yu Huang 0005, Keith Gallie |
DATE | 1 |
| 2006 | Diagnosis with Limited Failure InformationabstractThis paper discusses the challenges associated with diagnosing chain integrity and system logic failures in the production test environment with limited failure information. The following three methods were proposed to enhance diagnosis resolution in this scenario: (1) static pattern re-ordering (2) dynamic pattern re-ordering (3) per-pin based diagnosis. Experimental results illustrate that per-pin based failure logging and diagnosis algorithms enable scan chain diagnosis in volume production environment. Successful application of per-pin based chain diagnosis is demonstrated with an industrial case Yu Huang 0005, Wu-Tung Cheng, Nagesh Tamarapalli, Janusz Rajski, Randy Klingenberg, Will Hsu, Yuan-Shih Chen |
ITC | 1 |
| 2005 | Using fault model relaxation to diagnose real scan chain defectsabstractSoftware-based scan chain fault diagnosis is typically composed of two steps. First, scan chain flush patterns are used to identify faulty chains and fault models. This is followed by chain diagnosis using scan patterns in the second step. In this paper, we target chain diagnosis on one special category of chain faults: intermittent scan chain faults. It is showed that these faults may not be modeled correctly in the first step. Hence, a novel diagnosis methodology based on scan chain fault model relaxation is proposed. Yu Huang 0005, Wu-Tung Cheng, Greg Crowell |
ASP-DAC | 1 |
| 2005 | Off-shore outsource DFT vs. build off-shore branch officesabstractOutsourcing is the result of maturation in an industry: as an industry matures, skills are automated and inevitably become commoditized. When commoditization occurs, the work goes to those who best compete with the lowest common denominators of trained labor, low-cost resources, and affordable capital. Outsourcing seems to be an attractive means for cost containment for semiconductor industry as both market and profit margins have tightened. Corporations should focus on core and outsource context. Core means the activity that directly affects the competitive advantage of an organization, which differentiates a company from its competitors. All other activities are context. The good news is that one company's context may be another company's core. Yu Huang 0005 |
ITC | 1 |
| 2005 | Compressed pattern diagnosis for scan chain failuresabstractIn scan based designs, 10%-30% defects are in scan chains. Hence scan chain fault diagnosis becomes an important process for silicon debug and yield ramp up. With embedded compression techniques getting popular, chain diagnosis on devices with the embedded compression techniques becomes a challenge. In this paper, we provide a general methodology that can be applied for performing chain diagnosis in the context of any embedded compression techniques with any existing chain diagnosis algorithms. The proposed methodology enables seamless reuse of the existing chain diagnosis infrastructure with compressed test data. Experimental results show that with compressed patterns, the chain diagnosis resolution can be enhanced up to one order of magnitude with only 25% of failure cycles collected from ATE, compared to the diagnosis results with uncompressed patterns. Yu Huang 0005, Wu-Tung Cheng, Janusz Rajski |
ITC | 1 |
| 2004 | Compactor Independent Direct DiagnosisabstractIn scan test environment, designs with embedded compression techniques can achieve dramatic reduction in test data volume and test application time. However, performing fault diagnosis with the reduced test data becomes a challenge. In this paper, we provide a general methodology based on circuit transformation technique that can be applied for performing fault diagnosis in the context of any compression technique. The proposed methodology enables seamless reuse of the existing standard ATPG based diagnosis infrastructure with compressed test data. Experimental results indicate that the diagnostic resolution of devices with embedded compression is comparable with that of devices without embedded compression. Wu-Tung Cheng, Kun-Han Tsai, Yu Huang 0005, Nagesh Tamarapalli, Janusz Rajski |
Asian Test Symposium | 3 |
| 2004 | Intermittent Scan Chain Fault Diagnosis Based on Signal Probability AnalysisabstractA new algorithm to diagnose intermittent scan chain fault in scan-based designs is proposed in this paper. An intermittent scan chain fault sometimes is triggered and sometimes is not triggered during scan chain shifting, which makes it very difficult to locate the fault sites. In this paper, we provide answers to three questions: (1) Why intermittent scan chain faults happen? (2) Why diagnosis of this type of faults is necessary? (3) How to diagnose this type of faults? The experimental results presented demonstrate that the proposed diagnosis algorithm is effective for large industrial designs with multiple intermittent scan chain faults. Yu Huang 0005, Wu-Tung Cheng, Cheng-Ju Hsieh, Huan-Yung Tseng, Alou Huang, Yu-Ting Hung |
DATE | 1 |
| 2003 | Efficient Diagnosis for Multiple Intermittent Scan Chain Hold-Time FaultsabstractWhen VLSI design and process enter the stage of ultra deep submicron (UDSM), process variations, signal integrity (SI) and design integrity (DI) issues can no longer be ignored. These factors introduce some new problems in VLSI design, test and diagnosis, which increase lime-to-market, time-to-volume and cost for silicon debug. Intermittent scan chain hold-time fault is one of such problems we encountered in practice. The fault sites have to be located to speedup silicon debug and improve yield. Recent study of the problem proposed a statistical algorithm to diagnose the faulty scan chains if only one fault per chain. Based on the previous work, in this paper, an efficient diagnosis algorithm is proposed to diagnose faulty scan chains with multiple faults per chain. The presented experimental results on industrial designs show that the proposed algorithm achieves good diagnosis resolution in reasonable time. Yu Huang 0005, Wu-Tung Cheng, Cheng-Ju Hsieh, Huan-Yung Tseng, Alou Huang, Yu-Ting Hung |
Asian Test Symposium | 1 |
| 2003 | Using embedded infrastructure IP for SOC post-silicon verificationabstractThis paper presents a method to embed an FPGA core in a SOC as an infrastructure IP that can exploit transaction-based verification methodology to verify and debug the first silicon. The primary objective for post-silicon verification is to reduce the time taken for validating the first silicon. Additionally, verifying silicon at chip-level is expected to speedup the silicon debugging and thus reduce the time to market. Experimental results presented in this paper demonstrate that the proposed method can be implemented with small overhead. Yu Huang 0005, Wu-Tung Cheng |
DAC | 1 |
| 2003 | Statistical Diagnosis for Intermittent Scan Chain Hold-Time FaultabstractIntermittent scan chain hold-time fault is discussed in this paper and a method to diagnose the faulty site in a scan chain is proposed as well. Unlike the previous scan chain diagnosis methods that targeted permanent faults only, the proposed method targets both permanent faults and intermittent faults. Three ideas are presented in this paper. First an enhanced upper bound on the location of candidate faulty scan cells is obtained. Second a new method to determine a lower bound is proposed. Finally a statistical diagnosis algorithm is proposed to calculate the probabilities of the bounded set of candidate faulty scan cells. The proposed algorithm is shown to be efficient and effective for large industrial designs with multiple faulty scan chains. 1 Yu Huang 0005, Wu-Tung Cheng, Sudhakar M. Reddy, Cheng-Ju Hsieh, Yu-Ting Hung |
ITC | 1 |
| 2003 | SOC Test Scheduling Using Simulated AnnealingabstractWe propose an SOC test scheduling method based on simulated annealing. In our method, the test scheduling is formulated as a two-dimensional bin packing problem (rectangle packing) and a data structure called a sequence pair is used to represent the placement of the rectangles. Simulated annealing is used to find the optimal test schedule by altering an initial sequence pair and changing the width of the core wrapper. We also propose a method of wrapper design for cores without internal scan chains. Experiments are conducted on ITC'02 benchmarks, showing that overall the proposed method provides better solutions compared to earlier methods. Sudhakar M. Reddy, Irith Pomeranz, Yu Huang 0005 |
VTS | 4 |
| 2002 | Core - Clustering Based SOC Test Scheduling OptimizationabstractIn this paper, a method is presented to schedule tests for core-based SoCs to achieve optimal test completion time for the SoC design by simultaneously determining optimal core clustering, core cluster wrapper width, and pin mapping. For the first time the above mentioned techniques are applied concurrently to solve the SoC test scheduling problem. A heuristic algorithm implementing these techniques to determine an optimal solution is proposed. Yu Huang 0005, Sudhakar M. Reddy, Wu-Tung Cheng |
Asian Test Symposium | 1 |
| 2002 | Optimal Core Wrapper Width Selection and SOC Test Scheduling Based on 3-D Bin Packing AlgorithmabstractThis paper presents a method to consider a given SOC with pin and peak power constraints, and simultaneously (1) determine an optimal wrapper width for each core, (2) allocate SOC pins to cores and (3) schedule core tests to minimize the test completion time. For the first time the stated problem is formulated as a restricted 3 dimensional bin-packing problem and a heuristic to determine an optimal solution is proposed. Yu Huang 0005, Sudhakar M. Reddy, Wu-Tung Cheng, Paul Reuter, Nilanjan Mukherjee 0001, Chien-Chung Tsai, Omer Samman, Yahya Zaidan |
ITC | 1 |
| 2002 | On Concurrent Test of Core-Based SOC Design
Yu Huang 0005, Wu-Tung Cheng, Chien-Chung Tsai, Nilanjan Mukherjee 0001, Omer Samman, Yahya Zaidan, Sudhakar M. Reddy |
J. Electron. Test. | 1 |
| 2002 | Synthesis of Scan Chains for Netlist Descriptions at RT-Level
Yu Huang 0005, Chien-Chung Tsai, Nilanjan Mukherjee 0001, Omer Samman, Wu-Tung Cheng, Sudhakar M. Reddy |
J. Electron. Test. | 1 |
| 2001 | Resource Allocation and Test Scheduling for Concurrent Test of Core-Based SoC DabstractA method to solve the resource allocation and test scheduling problems together in order to achieve concurrent test for core-based system-on-chip (SOC) designs is presented in this paper. The primary objective for concurrent SOC test is to reduce test application time. The methodology used in this paper is not limited to any specific test access mechanism (TAM). Additionally, it can also be applied for test budgeting during the design phase to obtain a tradeoff between test application time and SOC pins needed. In this paper, the above problem is formulated as a well-known 2-dimensional bin-packing problem. A best fit heuristic algorithm is employed to obtain satisfactory results. Yu Huang 0005, Wu-Tung Cheng, Chien-Chung Tsai, Nilanjan Mukherjee 0001, Omer Samman, Yahya Zaidan, Sudhakar M. Reddy |
Asian Test Symposium | 1 |
| 2001 | On RTL scan designabstractThis paper presents a methodology to insert scan paths in a functional Register Transfer Level (RTL) specification of a design that can exploit existing functional paths between sequential elements in the original circuit for establishing scan chains. The primary objective for RTL scan insertion is to reduce the time taken for DFT, and thus reduce the time to market. Additionally, building scan chains at the functional RT-Level is expected to reduce the total area overhead introduced by full scan without compromising the fault coverage achieved. In addition, it often eliminates the delay associated with the additional multiplexer as a part of a conventional scan-cell in high performance designs. Experimental results presented in this paper demonstrate that the proposed method achieves the above objectives while also achieving higher fault coverages for most of the benchmark circuits considered. Yu Huang 0005, Chien-Chung Tsai, Nilanjan Mukherjee 0001, Omer Samman, Dan Devries, Wu-Tung Cheng, Sudhakar M. Reddy |
ITC | 1 |
| 2000 | Improving the Proportion of At-Speed Tests in Scan BISTabstractA method to select the lengths of functional sequences in a BIST scheme for scan designs is proposed in this paper. A functional sequence is a sequence of primary input vectors applied when the circuit operates as a sequential circuit, without using scan. These sequences can be applied at-speed, i.e., at the normal circuit clock speed. The objectives set for choosing the lengths of the functional sequences are to increase the number of vectors applied at-speed, and to reduce the number of settings of functional sequence lengths, without compromising the fault coverage achieved. The experimental results presented demonstrate that compared to earlier methods, the proposed method achieves the above objectives while also achieving higher fault coverages for most of the benchmark circuits considered. Yu Huang 0005, Irith Pomeranz, Sudhakar M. Reddy, Janusz Rajski |
ICCAD | 1 |