VLDB 2026 Research / reviewers in the wild / expert
Huaguo Liang
dblp:30/6037
· DBLP profile ↗
103ranked-venue papers
10as first author
63since 2021 · last 2026
0000-0002-0307-7236ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 95 · 8 first-author · 59 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorSecurity and privacy · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LTTL: A Low-Overhead and Triple-Node-Upset-Tolerant Latch Design for Aerospace ApplicationsabstractAs the feature size of the CMOS technology keeps scaling down, the charge sharing caused by radiation is becoming more and more prominent, and the occurrence possibility of the triple-node upset (TNU) increases significantly. In this paper, we propose a low-overhead and TNU-tolerant latch (LTTL) that leverages three parallel storage cells and an output-level error interceptive module to achieve complete TNU tolerance while minimizing design overhead. The optimized structure eliminates redundant devices and employs a high-speed D-to-Q path, significantly reducing delay-area-power product (DAPP). Even any three nodes of the latch are flipped at the same time, the output of the latch can retain the original value. Simulation results not only confirm the TNU tolerance of the proposed latch but also demonstrate that the latch can provide a 57% reduction in delay, 20% reduction in area, and 62% reduction in DAPP on average compared to state-of-the-art TNU-tolerant latches. Zikang Ma, Zhongyu Gao, Qianhui Liu, Yi Man, Huaguo Liang, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 6 |
| 2026 | A Phase Delay Comparison-Based TRNG Controlled by TDC
Peiyang Kang, Ying Zhang 0118, Yingchun Lu, Huaguo Liang |
ISCAS | 6 |
| 2026 | Dynamic Meta-Learning with Attentional Prototypes for Few-Shot Wafer Defect Recognition
Tianming Ni, Meifang Yu, Huaguo Liang, Senling Wang, Muyang Cheng, Mu Nie |
J. Electron. Test. | 4 |
| 2026 | RO-like ring-based TRNG with adaptive mode switching for enhanced entropy Harvesting
Jinlin Chen, Huaguo Liang, Yingchun Lu |
Integr. | 2 |
| 2026 | Design of a dynamic obfuscation-based strong PUF resistant to modeling attacks and mutual authentication protocol
Yingchun Lu, Huaguo Liang, Zhengfeng Huang, Jinlin Chen, Xiumin Xu |
Integr. | 4 |
| 2026 | FTPUF:Feedback structure of TERO PUF for high reliability
Yingchun Lu, Xinkai Wu, Jinlin Chen, Huaguo Liang, Zhengfeng Huang, Xiumin Xu |
Integr. | 4 |
| 2026 | Testing method for marginal defects based on dynamic critical resistance
Zhiwei Shao, Huaguo Liang, Shichao Bai, Yingchun Lu, Zhengfeng Huang |
Integr. | 2 |
| 2026 | Ring Oscillator-Based PreBond TSV Testing Method With Classification and Grading of DefectsabstractThe immaturity of manufacturing processes often leads to a high incidence of defects in through-silicon vias (TSVs). Prebond TSV testing is essential for optimizing the yield of chip-based integrated circuits. However, current testing methods are limited by their incomprehensive fault coverage and difficulty detecting subtle defects. Furthermore, these methods exhibit significant performance variability due to changes in process angle, power supply voltage, and temperature (PVT). To overcome these limitations, this paper introduces an innovative Ring Oscillator (RO)-based prebond test method specifically designed for TSVs, with a robust system for defect classification and grading. By sampling each node of the RO oscillating ring, the proposed method enhances the resolution of the Time-to-Digital Converter, thereby improving the defect detection capability. Additionally, a weak current source, constructed utilizing the unique properties of MOS transistors, enables the precise detection of open faults, resistive open defects with Ropen ≥ 1.5 K, and leakage defects with Rleak ≤ 10 G. To further mitigate the impact of PVT variations on test results, the paper integrates advanced machine learning techniques for defect classification and grading, providing valuable insights for fault bin classification and fault diagnosis. This innovative approach contributes significantly to the advancement of 3D IC reliability assessment. Xianrui Dou, Huaguo Liang, Zhengfeng Huang, Yingchun Lu, Jun Liu 0070 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | Lightweight High-Throughput Portable Multi-Mode Reconfigurable Integrated PUF-TRNGabstractPrivacy-Preserving Mutual Authentication (PPMA) protocols utilize Physical Unclonable Function (PUF) and True Random Number Generator (TRNG) as security primitives to protect privacy. To ensure the security of Internet of Things (IoT) nodes in untrusted environments, PPMA keys and encrypted data must reside on the same chip. The concept of integrating PUF and TRNG on a single device has thus emerged as a new security paradigm. This paper proposes a novel lightweight, portable, multi-mode reconfigurable integrated PUF-TRNG architecture resistant to machine learning attacks. Through co-design, the architecture achieves the integration and switching among three modes: Arbiter Physical Unclonable Function (APUF), Ring Oscillator Physical Unclonable Function (RO PUF), and TRNG. The APUF mode leverages the RO PUF mode for assistance, endowing it with machine learning resistance, where the highest prediction rates using Logistic Regression (LR), Support Vector Machine (SVM), CMA (Comparative Model Analysis), and Deep Neural Network (DNN) algorithms are only around 60%. Additionally, a lightweight authentication protocol is proposed to further enhance resistance against machine learning attacks. In TRNG mode, the architecture has two outputs, each capable of generating random numbers at 800 Mbps, resulting in a total throughput of 1600 Mbps. The generated random numbers have successfully passed various tests, including NIST SP800-22, NIST SP800-90B, AIS-31, and TESTU01. In the NIST SP800-22 test, the pass rates for both outputs of the Artix-7 and Kintex-7 FPGAs are approximately 99%. Jinlin Chen, Mingjing Qiu, Peiyang Kang, Zhengfeng Huang, Yingchun Lu, Huaguo Liang, Yaohua Xu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2026 | A Parallel Feedback Obfuscation Strong PUF Against Machine-Learning Modeling Attacks and Lightweight Authentication ProtocolabstractArbiter physical unclonable function (APUF) is a hardware security primitive that generates security keys by utilizing unavoidable process variations during chip manufacturing. However, the structure based on linear additive function makes it vulnerable to machine learning (ML) attacks. This paper proposes a parallel feedback obfuscation PUF (PFO PUF) design, which uses intermediate arbitration signals of the lower-layer APUF to generate the hidden challenge of the upper-layer APUF, enhancing the overall nonlinearity of the structure. The obfuscation module makes weight judgment for intermediate arbitration signals of upper-layer and lower-layer APUFs, which obfuscates the real response of PUF. We further design a variant of PFO PUF called reconfigured challenge obfuscation PFO PUF (RPFO PUF) and propose its lightweight device authentication protocol. RPFO PUF enhances the resistance of the original PFO PUF against reverse engineering (RE) and improves its Strict Avalanche Criterion (SAC) characteristic by reordering the challenges and incorporating weak PUF responses. The proposed PFO PUF and RPFO PUF were comprehensively evaluated via Python-based simulations and FPGA measurements. In Python simulations, both designs show strong resistance to state-of-the-art ML attacks, with logistic regression (LR), support vector machine (SVM), and covariance matrix adaptation evolution strategies (CMA-ES) yielding near 50% prediction accuracies under various PUF configurations. Although deep neural network (DNN) achieves up to 69.52% prediction accuracy on the PFO PUF, it drops to ∼50% on the RPFO PUF. FPGA results further confirm this, with the (32, 11)-RPFO PUF achieving a maximum prediction accuracy of only 51.47% across all four ML attacks. Moreover, both designs incur low hardware overheads, requiring just 743 and 2145 gate equivalents (GEs), respectively. Zhengfeng Huang, Yankun Lin, Yingchun Lu, Huaguo Liang, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | Multi-channel TRNG based on Scalable Cascaded Full Feedback Ring OscillatorabstractTrue random number generator (TRNG) is a key component in ensuring hardware security, and with the development of technologies such as high-speed communications, there is a higher demand for its generation rate. In this work, an ultra-high throughput rate TRNG based on a scalable cascaded full feedback ring oscillator (CFFRO) as the entropy source circuit is presented and implemented on Xilinx Artix-7, Kintex-7, and zynq UltraScale+ FPGAs devices. Unlike previous works, the proposed CFFRO is designed to be constructed as multiple parallel internal ROs, which, in turn, are sequentially cascaded and coupled to each other to disrupt the frequency spectrum of each ring oscillator and enhance the output uncertainty. Each internal ring oscillator in CFFRO can be used as an output for random numbers, creating multi-channel TRNG with parallel outputs and single-channel TRNG with multi-bit serial outputs. Measurements of the sequences extracted by both random number output schemes of TRNG show good randomness in the NIST SP800-22 suits and high entropy values in both the NIST SP 800-90B and AIS-31 testing suits, and the Dieharder suite verified the robustness under voltage and temperature variations. Moreover, due to the good extensibility of CFFRO, TRNGs with 2–8 channel counts are implemented in this work. At a sampling frequency of 400 MHz, the random sequences generated by 2–8 channel TRNGs can pass the tests. Peiyang Kang, Yaohua Xu, Huaguo Liang, Zhengfeng Huang, Yingchun Lu |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2026 | Reusing LBIST as a Physically Unclonable Function: A Low-Overhead PUF Capturing Circuit Transient ResponsesabstractPhysical unclonable functions (PUFs) are becoming a crucial solution for addressing hardware security challenges in integrated circuits. However, their widespread adoption is often hindered by the significant overhead of dedicated PUF circuits. This article introduces LBIST-PUF, a novel intrinsic PUF that minimizes overhead by ingeniously reusing the existing logic built-in self-test (LBIST) infrastructure. The design generates unique chip fingerprints through capturing transient responses of the circuit under test (CUT) with a high-frequency configurable clock. To ensure robustness, we incorporate a compensation circuit and a signature correction algorithm, enhancing reliability against environmental variations. Experimental evaluation confirms that the LBIST-PUF achieves near-ideal performance: reliability close to 100%, uniqueness of 50.01%, and a pass rate exceeding 95% on the National Institute of Standards and Technology (NIST) statistical tests. These results underscore the potential of our design as a secure, low-cost authentication solution for Internet of Things (IoT) applications. Zhiwei Shao, Huaguo Liang, Shichao Bai, Hao Lv 0008, Zhengfeng Huang, Yingchun Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | TNURML: Triple-Node-Upset-Recovery Magnetic Latch Design with Non-Volatility for Aerospace ApplicationsabstractAs semiconductor technology advances, radiative-particle-induced soft errors and power consumption are becoming major concerns for digital circuits in aerospace applications. Radiation hardening by design and magnetic tunnel junctions (MTJs) are widely employed to address these concerns. In this paper, a novel latch, called TNURML, that can completely recover from triple-node upsets (TNUs), is proposed. The embedded MTJs provide non-volatility and are compatible with traditional CMOS processes. The TNURML employs a TNU-recovery module as well as a pair of MTJs for backup and recovery operations. Extensive simulations demonstrate the excellent TNU-recovery capability and non-volatility of the TNURML latch at the cost of slightly increased area overhead. The proposed TNURML latch reduces 43.63% of delay and 48.23% of power on average when compared to the state-of-the-art latches. Aibin Yan, Zhiyuan Pei, Cuiyun Jiang, Huaguo Liang, Xiaoqing Wen, Patrick Girard 0001 |
ISCAS | 6 |
| 2025 | A High-Precision Pre-Bond TSV Delay-Fault Detection Technique Using Digitally Controlled Delay LinesabstractTSV fabrication in 3D ICs remains immature, causing resistive open-circuit and leakage faults. Pre-bond testing of TSVs improves 3D IC yield and performance. This paper proposes a fully digital pre-bond TSV delay-fault detection scheme using a dual-stage DCDL and a Bang-Bang Phase Detector, a delayed resolution of about 1.52 ps was achieved. 45 nm CMOS HSPICE simulations verify detection of Ropen≥ 0.2 kΩ and Rleak≤ 3.3 MΩ, significantly enhancing accuracy. Zhongyuan Liang, Huaguo Liang, Xianrui Dou |
ITC-Asia | 2 |
| 2025 | High throughput true random number generator based on dynamically superimposed hybrid entropy sources
Yingchun Lu, Changlong Cao, Yang Li 0010, Huaguo Liang, Lixiang Ma |
Integr. | 4 |
| 2025 | Low test cost adaptive testing method for high yield IC products
Yuqi Pan, Huaguo Liang, Zhengfeng Huang, Maoxiang Yi, Yingchun Lu |
Integr. | 2 |
| 2025 | Lightweight high-throughput true random number generator based on state switchable ring oscillator
Shehui Wu, Huaguo Liang, Hao Lv 0008, Maoxiang Yi, Yingchun Lu |
Integr. | 2 |
| 2025 | Semi-supervised lithography hotspot detection based on feature fusion and residual attention
Xinzhong Xiao, Wenxin Huang, Ruijun Ma 0002, Fuxin Tang, Pan Qi, Huaguo Liang |
Integr. | 8 |
| 2025 | ESegNet-ILT: An end-to-end mask optimization method in VLSI design flow based on enhanced SegNet
Yong Xue, Yu Zhang 0162, Ruijun Ma 0002, Huaguo Liang, Zhengfeng Huang |
Integr. | 7 |
| 2025 | Improve SAC in PUFs: Metric, Analysis, Algorithm, and ApplicationabstractPhysical Unclonable Functions (PUFs) are crucial for lightweight authentication in the Internet of Things (IoT), but existing PUFs often have poor statistical properties and are vulnerable to machine learning attacks. Designs implementing the Strict Avalanche Criterion (SAC) lack sufficient theoretical foundation. This paper introduces quantitative metrics to evaluate the SAC performance of strong PUFs and conducts rigorous analysis on Arbiter PUFs (APUFs) and their classic variants, addressing imprecision in existing methods. Based on these metrics, we developed an algorithm to optimize the SAC performance of strong PUFs by adjusting the challenge sequences. This optimization improved the SAC performance of the 2-XOR APUF by 59% without additional resource consumption, making it comparable to the 4-XOR APUF; We also provided mathematical proof for the optimal solution. Furthermore, we propose the SAC Optimized Shuffled XOR Arbiter PUF (SOS XOR APUF), which improves SAC performance by 84% compared to the 3-XOR APUF with the same entropy source. It addresses the inherent defect of poor statistical properties when two adjacent bits in the challenge flip, achieving a theoretical response flip probability of 0.5. The SOS XOR APUF resists existing machine learning attacks---including Logistic Regression (LR), Covariance Matrix Adaptation Evolution Strategy (CMA-ES), Artificial Neural Network (ANN), and Deep Neural Network (DNN)---with prediction accuracy below 55%. Finally, we designed and verified an authentication protocol based on this PUF in the ProVerif environment, achieving mutual authentication between IoT devices and servers, preventing secret information from being stolen, and enhancing the security of the PUF structure. Zhengfeng Huang, Fansheng Zeng, Jingchang Bian, Huaguo Liang, Yingchun Lu, Tianming Ni |
IEEE Internet Things J. | 5 |
| 2025 | SegNet-OPC: A Mask Optimization Framework in VLSI Design Flow Based on Semantic Segmentation Network
Pan Qi, Fuxin Tang, Huaguo Liang, Zhengfeng Huang |
J. Comput. Sci. Technol. | 4 |
| 2025 | Ultra-High Efficiency TRNG IP Based on Mesh Topology of Coupled-XORabstractThe true random number generator is capable of generating completely random and unpredictable sequences, and plays a crucial role in various fields such as cryptography, encryption communication, and random algorithms. To meet the demand for high-throughput true random number generators in modern high-speed systems, a lightweight TRNG design is proposed. It utilizes a mesh topology of coupled-XOR as entropy source and generates a highly compact and high throughput true random number generator by coupling oscillators in the network. The generated random sequences have successfully passed the NIST SP 800-22, TESTU01, NIST SP 800-90B, and AIS-31 tests. It achieved ultra-high throughput of 2.1Gbps and 2.4Gbps on the Xilinx Artix-7 and Kintex-7 series development boards, respectively, achieving efficient utilization of hardware resources. Compared with other works, this design has significant advantages in terms of resource utilization and throughput. Yingchun Lu, Enpu Xu, Jinlin Chen, Huaguo Liang, Zhengfeng Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Fabdb: a low-latency fault-tolerant architecture based on dynamic bypass for network-on-chip
Zhuoxuan Ji, Jianhua Li 0003, Huaguo Liang |
J. Supercomput. | 4 |
| 2025 | Pulse-Based Prebond TSV TestingabstractDue to the immaturity of the manufacturing process, numerous faults often occur in through-silicon vias (TSVs). Prebond TSV testing is crucial in enhancing the performance and yield of chiplet-based integrated chips. However, most existing test methods suffer from the test resolution and hard-to-detect weak faults. A novel prebond TSV test method based on the pulse is proposed to improve the test circuit. By introducing pMOS as a driver in pulse detection, TSV leakage faults can be directly tested, thus improving the resolution of leakage faults’ detection. In addition, the range of test pulsewidth to digital code conversion is effectively improved by the ring oscillator (RO) for coarse detection and pulse shrinking for fine detection, avoiding the problem of large overheads that would be brought about by solely increasing the pulse shrinking chain. The results validated by HSPICE simulation show that it can detect open faults, resistive open faults with$R_{\text {open}} \gt $$0.9~{\mathrm {K}} {\mathrm {\Omega }}$, leakage faults with$R_{\text {leak}} \lt $$30~{\mathrm {G}} {\mathrm {\Omega }}$, and compound faults consisting of resistive open faults and leakage faults. Xianrui Dou, Huaguo Liang, Zhengfeng Huang, Yingchun Lu, Maoxiang Yi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | BF PUF: A Modeling Attack-Resistant Strong PUF Based on Bent FunctionsabstractStrong physical unclonable functions (PUFs) are promising circuits for lightweight Internet of Things (IoT) authentication and security. However, existing strong PUFs exhibit very low cryptographic nonlinearity (NL), making them vulnerable to machine learning (ML) modeling and cryptanalytic attack. To address this issue, we propose the Bent function PUF (BF PUF) based on Maiorana-McFarland (M-M) constructed Bent functions, which obfuscates the responses of the strong PUF to enhance resistance against modeling attacks. The core idea is to employ the M-M construction method for Bent functions to ensure maximum cryptographic NL to resist modeling attacks. A Feistel network is configured using weak PUF responses as keys to achieve device-specific and unpredictable mappings of input challenges while meeting the requirements of the M-M Bent function construction. A Python-based model of the BF PUF was developed, and simulation results indicate that the cryptographic NL of the proposed BF PUF outperformsk-xorarbiter PUFs (APUFs) (${k} =2$, 4, 6). The proposed BF PUF was also implemented and evaluated on the FPGA hardware platform. The experimental results show that under modeling attacks using four ML algorithms—logistic regression (LR), artificial neural networks (ANNs), deep neural networks (DNNs), and covariance matrix adaptation evolution strategies (CMA-ES)—the best prediction accuracy under these four modeling attack algorithms is 52.60%. The reliability under temperature fluctuations ranging from$- 10~^{\circ }$C to$80~^{\circ }$C is between 84.20% and 99.78%. Zhengfeng Huang, Fansheng Zeng, Yanqiao Chi, Yankun Lin, Yingchun Lu, Huaguo Liang, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | A TSV Misalignment-Based Repair Architecture in 3-D ChipsabstractAs a critical component of 3-D integrated circuits (3D-ICs), the quality of through-silicon vias (TSVs) significantly impacts the yield and reliability of 3D-ICs, especially the clustered faults during manufacturing. In this article, a repair architecture based on TSV misalignment is proposed. This architecture achieves a higher repair rate by physically connecting the signal not to its closest TSV but only to the TSVs far away from each other. Experimental results show that the average repair rate of the proposed architecture increases by 13.42% compared to the existing repair architectures of the same type for clustered faults. Compared to the router-based architecture, the proposed architecture has a similar average repair rate with less than 0.15% difference in fewer than eight clustered faults, reducing the delay and MUX area overhead by 70.27% and 54.17%, respectively. Huaguo Liang, Jiahui Xiao, Xianrui Dou, Tianming Ni, Yingchun Lu, Zhengfeng Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | A Block-Group-Based Redundant TSV Architecture for Clustered FaultsabstractThree-dimensional integrated circuits (3-D ICs) based on through-silicon vias (TSVs) have tremendous advantages, such as high performance, low power consumption, and heterogeneous integration. However, TSV faults significantly diminish the yield of 3-D ICs. To address this issue, this article proposes a block-group-based redundant TSV (RTSV) architecture. In this architecture, TSVs are organized into multiple blocks, with four blocks constituting a basic unit for the repair of faulty TSVs (FTSV). Within each basic unit, four TSVs situated at the same position across four blocks form a group that shares an RTSV. This methodology effectively improves the repair rate for clustered TSV faults through cross-block and cross-group connections among TSVs. In addition, we propose a TSV repair algorithm that dynamically updates the difficulty values (DUDVs) of FTSVs. This approach uses quadtrees with FTSVs as the root node to compute the repair difficulty values, selecting the most difficult TSV for repair in each iteration. The difficulty values are continuously updated throughout the repair process to further improve the repair rate. Experimental results show that the proposed method achieves repair rates of 98.7% and 99.9% for random and clustered TSV faults, respectively. Jun Liu 0070, Tianhao Du, Mulin Ye, Xi Wu 0003, Tianming Ni, Huaguo Liang |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | RHT_NoC: A Reconfigurable Hybrid Topology Architecture for Chiplet-Based Multicore SystemabstractChiplet-based system-on-chip (SoC) architectures, leveraging 2.5-D/3-D integration technologies, provide scalable solutions for a wide range of applications. Achieving high performance and cost-effectiveness in these systems relies heavily on optimizing die-to-die interconnect topologies and designs, which are essential for seamless interchiplet communication. This article introduces a reconfigurable hybrid topology (RHT) architecture designed for chiplet-based multicore systems. RHT achieves high performance and energy efficiency by dynamically reconfiguring the network topology to traffic variations, adaptively selecting transport subnets, and optimizing link bandwidth allocation, thereby minimizing congestion and maximizing packet throughput. Furthermore, RHT leverages global traffic information to dynamically combine Torus loops, maximizing opportunities for rapid packet transmission delivery while guaranteeing minimal hop counts. Moreover, RHT accelerates packet transmission via bufferless combined loops, extending the continuous sleeping periods of routers, improves power gating efficiency, and significantly reduces static power consumption. Simulation results indicate that the Mesh-DyRing achieves over a 40% reduction in network latency and more than a 20% decrease in power consumption overhead compared to the baseline design. When compared to WiNoC, an advanced hybrid wired-wireless topology design, the Mesh-DyRing-PG configuration reduces power consumption by 56.2% while maintaining equivalent average network latency. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | IRCA-TRNG: A Lightweight Dual-Ring Chaotic TRNG With Perturbation Refresh for High ThroughputabstractAs a core component in the field of information security, the true random number generator (TRNG) produces high-entropy random numbers by extracting unpredictable noise from the physical environment, exhibiting nonreproducibility and resistance to prediction. To address the challenges posed by interference in high-speed systems, maintaining stable throughput and entropy sources for TRNG, this article proposes an optimized TRNG that utilizes chaotic interference to refresh the cellular automata (IRCA) iterative algorithm. The IRCA-TRNG utilizes a self-timed ring oscillator (STR) and jitter to perturb the operation of chaotic cellular automata (CA) cells, achieving a high-throughput TRNG. The generated random sequences have successfully passed NIST SP800-22, TESTU01, NIST SP800-90B, and AIS-31 tests. A throughput of 1040 Mb/s has been achieved on Xilinx Artix-7 and PYNQ-K2 series development boards. Compared with the state-of-the-art works, the proposed TRNG demonstrates significant advantages in resource utilization and performance quality factors. Peiyang Kang, Deqin Shi, Yaohua Xu, Yunlai Zhu, Zhengfeng Huang, Huaguo Liang, Yingchun Lu, Aibin Yan, Ying Zhang 0118 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | MTCX: Ultra-high Throughput TRNG Based on Mesh topology of Coupled-XORabstractThe true random number generator (TRNG) is capable of generating completely random and unpredictable sequences. TRNG is an important hardware security primitive, it plays a critical role in identity verification, prevention of replay attacks and key exchange. To meet the demand for high-throughput true random number generators in modern high-speed systems, a lightweight design is proposed. It utilizes a mesh topological of coupled-XOR as entropy source and generates a highly compact and high throughput true random number generator by coupling oscillators in the network. The generated random sequences have successfully passed the NIST SP800-22 and AIS-31 tests. It achieved ultra-high throughput of 2.1Gbps and 2.4Gbps on the Xilinx Artix-7 FPGA and Kintex-7 FPGA, respectively, achieving efficient utilization of hardware resources. Compared with other architectures, this design has significant advantages in terms of throughput and resource utilization. Yingchun Lu, Enpu Xu, Huaguo Liang, Cuiyun Jiang, Lixiang Ma |
ATS | 3 |
| 2024 | A RO-Integrated-LFSR-Based Nonlinear Strong PUF with Intrinsic Modeling Attacks ResilienceabstractPhysical Unclonable Functions (PUF) are important hardware security primitives used for generating keys and identity authentication, with wide applications in the Internet of Things security. However, the security of strong PUF faces serious threats from modeling attacks, especially in the case of Arbiter PUF and their variants that involve additive linear integration of entropy sources. This paper proposes a Ring-Oscillator-Integrated-Linear-Feedback-Shift-Register-based PUF (ROinLFSR PUF) that achieves immunity to modeling attacks by highly nonlinearly integrating independent responses from weak RO PUFs using a configurable LFSR. To increase the efficiency of entropy extraction in hardware resources, dual entropy sources extraction is performed on the period and duty cycle of RO. Python simulation and FPGA experimental results demonstrate that the proposed PUF has intrinsic resilience against modeling attacks. And the proposed PUF achieves good results in reliability, uniqueness, uniformity, and randomness. Jingchang Bian, Zhengfeng Huang, Yankun Lin, Huaguo Liang, Aibin Yan |
ITC-Asia | 6 |
| 2024 | PFO PUF: A Lightweight Parallel Feed Obfuscation PUF Resistant to Machine Learning AttacksabstractArbiter Physically Unclonable Functions (APUFs) are hardware security primitives that leverage manufacturing process variation to generate security keys. They can produce exponential challenge-response pairs (CRPs) with minimal hardware overhead. However, the symmetric nature of linear additive functions makes them vulnerable to modeling attacks rooted in machine learning. To address this issue, this paper introduces a novel design called Parallel Feed Obfuscation PUF (PFO PUF). In this approach, the intermediate decision signals from the lower APUF are used as a concealed challenge for the upper APUF, enhancing the overall nonlinearity of the dual-APUF. Additionally, obfuscation modules are employed to determine the weights of the intermediate decision signals from both the upper and lower APUFs, protecting the actual response. Experimental results demonstrate that the proposed PFO PUF effectively withstands four advanced machine learning attack algorithms, including Logistic Regression (LR), Support Vector Machine (SVM), Deep Feedforward Neural Network (DFNN), and Efficient CANDECOMP/PARAFAC Tensor Regression Network (ECPTRN). The prediction accuracy of these four algorithms is consistently below 66.30%. Compared with other enhanced structures based on APUF, PFO-PUF only uses 493 LUTs and has lower resource overhead. Zhengfeng Huang, Yankun Lin, Fansheng Zeng, Jingchang Bian, Huaguo Liang, Yingchun Lu, Xiaoqing Wen, Tianming Ni |
ITC-Asia | 6 |
| 2024 | A Quadruple-Node Upsets Hardened Latch Design Based on Cross-Coupled Elements
Zhengfeng Huang, Zishuai Li, Huaguo Liang, Tianming Ni, Aibin Yan |
J. Electron. Test. | 4 |
| 2024 | Wafer-level Adaptive Testing Based on Dual-Predictor Collaborative Decision
Yuqi Pan, Huaguo Liang, Jinxing Qu, Zhengfeng Huang, Maoxiang Yi, Yingchun Lu |
J. Electron. Test. | 2 |
| 2024 | A High-Performance Quadruple-Node-Upset-Tolerant Latch Design and an Algorithm for Tolerance Verification of Hardened Latches
Xuewei Qin, Ruijun Ma 0002, Chaoming Liu, Huaguo Liang |
J. Electron. Test. | 7 |
| 2024 | A self-training end-to-end mask optimization framework based on semantic segmentation network
Fuxin Tang, Pan Qi, Huaguo Liang, Zhengfeng Huang |
Integr. | 5 |
| 2024 | Design of novel low cost triple-node-upset self-recoverable hardened latch
Ruijun Ma 0002, Zhengfeng Huang, Huaguo Liang, Haojie Sun, Chaoming Liu |
Integr. | 5 |
| 2024 | F-Bypass: A Low-Power Network-on-Chip Design Utilizing Bypass to Improve Network ConnectivityabstractWith the development of transistor feature size to nanometer level, static power consumption has gradually become the main factor affecting the overall power consumption of network-on-chip (NoC). Power gating is an effective technology to reduce static power consumption, but it also brings new challenges, such as BET violation, wake-up latency and network connectivity. Therefore, a power gating method is needed to improve NoC performance and reduce static power consumption. This article proposes a low-power bypass method, namely Forwarding bypass (F-Bypass). First, F-Bypass adds bypass paths between all input and output ports and the network interface (NI) and connects the pop-up port and injection port in NI through the bypass path. When the router is powered off, F-Bypass performs wake-up-free packet transmission, which reduces the break-even time (BET) violation and cumulative wake-up latency while ensuring network connectivity. Secondly, this article adds the modified VC state table to NI so that the power-off router can perform normal traffic control. Finally, a new wake-up criterion is proposed, which can effectively avoid the frequent wake-up of power-off routers, and the detailed hardware implementation of F-Bypass is provided. The simulation results under integrated traffic load show that compared with the traditional scheme, the delay of F-Bypass is reduced by 2.2%, the throughput is increased by 13.1%, and the total static power consumption is reduced by 75.2%. Key performance indicators are superior to other solutions, and the increased area cost is moderate. Shuaijie Yuan, Jianhua Li 0003, Huaguo Liang |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2024 | Traffic-oriented reconfigurable NoC with augmented inter-port buffer sharingabstractAs the number of cores in a multicore system increases, the communication pressure on the interconnection network also increases. The network-on-chip (NoC) architecture is expected to take on the ever-expanding communication demands triggered by the ever-increasing number of cores. The communication behavior of the NoC architecture exhibits significant spatial-temporal variation, posing a considerable challenge for NoC reconfiguration. In this paper, we propose a traffic-oriented reconfigurable NoC with augmented inter-port buffer sharing to adapt to the varying traffic flows with a high flexibility. First, a modified input port is introduced to support buffer sharing between adjacent ports. Specifically, the modified input port can be dynamically reconfigured to react to on-demand traffic. Second, it is ascertained that a centralized output-oriented buffer management works well with the reconfigurable input ports. Finally, this reconfiguration method can be implemented with a low overhead hardware design without imposing a great burden on the system implementation. The experimental results show that compared to other proposals, the proposed NoC architecture can greatly reduce the packet latency and improve the saturation throughput, without incurring significant area and power overhead. Chenglong Sun, Huaguo Liang |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2024 | Design Guidelines and Feedback Structure of Ring Oscillator PUF for Performance ImprovementabstractThe physical unclonable function (PUF) is a hardware security primitive that is used to generate secret keys or identity authentication for chips using random manufacturing process variation (MPV). The PUF based on ring oscillator (RO PUF) has been extensively studied in recent years because of its high robustness and ease of design. Although the performance has been optimized in previous studies, several uniqueness, reliability, and theoretical foundation concerns still remain. This article presents a transistor-level parameters-based quantitative theoretical model, which clearly reveals several design guidelines for improving the reliability of RO PUF. Furthermore, a PUF based on a feedback ring oscillator (RO) structure is proposed, which combined the RO topology and the drafting effect of XOR gates to enhance the uniqueness and reliability. The correctness of the theoretical model was verified by the SPICE simulation experiment result. And in the FPGA experiment result, the uniqueness and the reliability of feedback RO PUF using the same hardware resources on the same chip was superior to that of RO PUF. The theoretical research method of RO PUF used can be widely applied to other PUFs using ring topology and feedback RO PUF is a great substitute for RO PUF. Zhengfeng Huang, Jingchang Bian, Yankun Lin, Huaguo Liang, Tianming Ni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | DBU-PG: energy-efficient noc design using dual-buffering power gating
Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang |
J. Supercomput. | 6 |
| 2024 | A Reliability-Aware Splitting Duty-Cycle Physical Unclonable Function Based on Trade-off Process, Voltage, and Temperature VariationsabstractThe physical unclonable function (PUF) is a hardware security primitive that can be used to prevent malicious attacks aimed at obtaining device information at the hardware level. The ring oscillator (RO) PUF has attracted considerable research attention. To improve the reliability of the RO PUF under voltage and temperature changes, the response of the duty-cycle (DC) PUF was obtained by comparing the duty cycle of the RO rather than the period. However, this method reduces the effective utilization of process variations, which limits its implementation in mature advanced manufacturing processes. In this study, a splitting duty-cycle (SDC) PUF was proposed to balance the effective extraction of process variations and robustness under voltage and temperature changes. The sensibility formula between the performance of SDC PUF and process, voltage, and temperature was established through a circuit model and statistical methodology, and the comprehensive characteristics of SDC PUF were analyzed theoretically. Next, 16 SDC PUFs with 128-bit responses were implemented and measured on a Xilinx Virtex-7 device. The experimental results revealed that the average native reliability of SDC PUF was 98.97%, and the reliability was 97.32% under various voltage and temperature conditions. This result revealed advantages over the DC PUF implemented in the same device. The uniqueness of the SDC PUF was 50.42%, and it passed the NIST SP 800-22 randomness and autocorrelation function tests. Jingchang Bian, Zhengfeng Huang, Huaguo Liang |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | A Self-Biased Current Reference Source-Based Pre-Bond TSV Test SolutionabstractIn the realm of 3-D integrated circuits (ICs), the presence of faults within through-silicon vias (TSVs) poses a significant threat, potentially undermining their yield and reliability. Thus, a meticulous TSV testing procedure is imperative to guarantee the optimal performance of 3-D ICs. This article introduces a novel pre-bond TSV testing approach. The method employs a self-biased current reference source (SCRS) to charge and discharge TSVs. This approach not only detects open faults, leakage faults, and compound faults but also distinguishes between these fault types. Additionally, the SCRS allows for a reduction in current magnitude, effectively prolonging the charge and discharge duration of TSVs, thereby enhancing fault detection capabilities. The implementation of a SCRS bolsters the resilience of the proposed method against variations in process, voltage, and temperature (PVT). Experimental results substantiate the superiority of this technique in terms of fault detection efficacy, and robustness in the face of PVT fluctuations when compared to alternative approaches. Jun Liu 0070, Songren Cheng, Xi Wu 0003, Huaguo Liang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Low Overhead and High Stability Radiation-Hardened Latch for Double/Triple Node Upsets
Zhengfeng Huang, Hao Wang 0169, Dongxing Ma, Huaguo Liang, Aibin Yan |
J. Electron. Test. | 4 |
| 2023 | Dynamic detection of wireless interface faults and fault-tolerant routing algorithm in WiNoC
Chang Qian, Wu Zhou 0007, Qi Wang 0027, Huaguo Liang |
Integr. | 6 |
| 2023 | Design of approximate Booth multipliers based on error compensation
Yongxia Sheng, Huaguo Liang, Bao Fang, Cuiyun Jiang, Zhengfeng Huang, Maoxiang Yi, Yingchun Lu |
Integr. | 2 |
| 2023 | Improving power and performance of on-chip network through virtual channel sharing and power gating
Wu Zhou 0007, Huaguo Liang |
Integr. | 4 |
| 2023 | LQNTL: Low-overhead quadruple-node-upset self-recovery latch based on triple-mode redundancy
Ruijun Ma 0002, Huaguo Liang, Zhengfeng Huang, Chaoming Liu |
Integr. | 4 |
| 2023 | URMP: using reconfigurable multicast path for NoC-based deep neural network accelerators
Chenglong Sun, Qi Wang 0027, Huaguo Liang |
J. Supercomput. | 5 |
| 2023 | Low-overhead TRNG based on MUX for cryptographic protection using multiphase sampling
Huaguo Liang, Yingchun Lu |
J. Supercomput. | 2 |
| 2023 | High-efficiency TRNG Design Based on Multi-bit Dual-ring OscillatorabstractUnpredictable true random numbers are required in security technology fields such as information encryption, key generation, mask generation for anti-side-channel analysis, algorithm initialization, and so on. At present, the true random number generator (TRNG) is not enough to provide fast random bits by low-speed bits generation. Therefore, it is necessary to design a faster TRNG. This work presents an ultra-compact TRNG with high throughput based on a novel extendable dual-ring oscillator (DRO). Owing to multiple bits output per cycle in DRO can be used to obtain the original random sequence, the proposed DRO achieves a maximum resource utilization to build a more efficient TRNG, compared with the conventional TRNG system based on ring oscillator (RO), which only has a single output and needs to build multiple groups of ring oscillators. TRNG based on the 2-bit DRO and its 8-bit derivative structure has been verified on Xilinx Artix-7 and Kintex-7 FPGA under the automatic layout and routing and has achieved a throughput of 550 Mbps and 1,100 Mbps, respectively. Moreover, in terms of throughput performance over operating frequency, hardware consumption, and entropy, the proposed scheme has obvious advantages. Finally, the generated sequences show good randomness in the test of NIST SP800-22 and Dieharder test suite and pass the entropy estimation test kit NIST SP800-90B and AIS-31. Yingchun Lu, Huaguo Liang, Maoxiang Yi, Zhengfeng Huang, Yuanming Ma |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2023 | RMC_NoC: A Reliable On-Chip Network Architecture With Reconfigurable Multifunctional ChannelabstractAs chip fabrication has advanced to the nano level, the increased link density has heightened the risk of failures. The potential performance drawbacks resulting from these link failures have become a critical challenge in the design of reliable network-on-chip (NoC) systems. Fault-tolerant routing algorithms have proven to be effective strategies for handling this issue by diverting packets away from failed links to prevent congestion. However, these algorithms often result in excessive packet diversion, especially in the presence of a higher failure rate, which can significantly constrain the network’s behavior. This article introduces a novel NoC design with reconfigurable multifunctional channels (RMC_NoC). This design dynamically adapts the channel functions in response to network conditions to ensure that packets from failed links follow their original paths. In addition, it presents a channel buffer bubble flow control mechanism that can resolve congestion by redistributing congested traffic within the channel buffer. The evaluation results demonstrate that our approach ensures superior network communication even in the presence of permanent link failures, with minimal area overhead and power consumption. Moreover, our system exhibits lower latency and higher throughput compared to state-of-the-art fault-tolerant methods across various link failure rates. Notably, even at a severe failure rate of 30%, RMC_NoC exhibits only a 16.3% increase in latency compared to an ideal failure-free environment (Baseline) while still maintaining system communication capabilities to a considerable extent. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Energy-Efficient Multiple Network-on-Chip Architecture With Bandwidth ExpansionabstractAs technology feature sizes diminish to the nanometer regime, the leakage power crisis has become a major challenge in network-on-chip (NoC) design. Power gating (PG) is used to mitigate growing leakage power as an effective static power-saving technique. Applying PG in a multiple NoC (Multi-NoC) rather than a traditional NoC is a promising solution. However, limited by the channel width of the subnets, the increase in packet length will bring a severe serialization issue and performance loss. Previous Multi-NoC schemes have to wake up more subnets to minimize the performance loss, which also sacrifices their energy efficiency. In this article, we introduce an architecture, namely, BandExp, which allows subnets to expand their bandwidth by utilizing the idle physical links of other subnets. More bandwidth helps subnets mitigate the serialization issue and reduce the performance loss. Meanwhile, other subnets gain longer sleep cycles and thus save more energy. Evaluation results indicate that compared to the state-of-the-art Catnap, the proposed architecture reduces the average packet latency and execution time of different benchmarks by 19.3% and 3.2%, respectively. Also, the net static energy of the network is reduced by 23.2% on average, while the incurred area overhead is only 1.3%. Wu Zhou 0007, Zhengfeng Huang, Huaguo Liang, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | A Low Power-Consumption Triple-Node-Upset-Tolerant Latch Design
Yingchun Lu, Guangzhen Hu, Hao Wang 0169, Huaguo Liang, Maoxiang Yi, Zhengfeng Huang |
J. Electron. Test. | 6 |
| 2022 | A reconfigurable test method based on LFSR for 3D stacking integrated circuits
Chen Tian 0011, Jianyong Lu, Liu Jun, Huaguo Liang, Yingchun Lu, Maoxiang Yi |
Integr. | 4 |
| 2022 | Architecting a congestion pre-avoidance and load-balanced wireless network-on-chip
Chenglong Sun, Huaguo Liang |
J. Parallel Distributed Comput. | 3 |
| 2021 | A N: 1 Single-Channel TDMA Fault-Tolerant Technique for TSVs in 3D-ICsabstractAs the number of 3D-IC stacks increases, defects of through silicon via (TSV) in manufacturing and bonding process seriously affect the yield and reliability of the chip. Comparing to discarding these defective ones, some faulttolerant architectures are proposed, however, these existing schemes have great hardware overhead. In the paper, an N:1 single-channel time division multiple access (TDMA) faulttolerant technique using redundant TSV to tolerate TSV defect is proposed. Data is grouped and transmitted through TSV by TDMA mechanism, which reduces the number of TSV. The number of groups depends on bandwidth and hardware. The N: 1 single-channel TDMA structure is designed to use a single TSV to accomplish the time-sharing transmission of a group of signals. Each signal TSV is equipped with an additional TSV to improve the fault-tolerant coverage. The functions are verified on 40 nm Xilinx virtex-6 FPGA. The simulation results of Design Compiler based on 45 nm PTM show that the fault coverage rate can be increased to 100%, and the area overhead is reduced by 60.8% compared with the existing methods. Huaguo Liang, Danqing Li, Tianming Ni, Zhengfeng Huang, Cuiyun Jiang |
ITC-Asia | 1 |
| 2021 | Neural Network-based Online Fault Diagnosis in Wireless-NoC Systems
Qi Wang 0027, Yingchun Lu, Huaguo Liang, Dakai Zhu 0001 |
J. Electron. Test. | 4 |
| 2021 | Approximate multipliers based on a novel unbiased approximate 4-2 compressor
Bao Fang, Huaguo Liang, Dawen Xu 0002, Maoxiang Yi, Yongxia Sheng, Cuiyun Jiang, Zhengfeng Huang, Yingchun Lu |
Integr. | 2 |
| 2021 | A Cost-Effective TSV Repair Architecture for Clustered Faults in 3-D ICabstractDue to the winding level of the thinned wafers and the surface roughness of silicon dies, the through-silicon vias (TSVs) defect tend to be clustered, reducing the yield of 3-D integrated circuit significantly. To tackle this fault clustering problem, the existing TSV repair methods adopt the TSV redundancy idea, which brings a major cost to 3-D integration. In this brief, a honeycomb-TDMA TSV design is proposed to mitigate the impact of multiple clustered faults without the need of redundant TSVs (RTSVs), thereby decreasing the area overhead and enhances the yield. The yield of the honeycomb-TDMA architecture can achieve 91.38%-99.67% for different benchmark circuits from IWLS 2005, which has the highest yield. Furthermore, our design achieves total additional hardware (timing delay overhead) reduction by 83.70%-86.85% (46.01%-55.96%), 66.89%-73.25% (29.41%-38.49%), 68.02%-74.20% (41.40%-52.20%), 60.60%-68.18% (18.09%-33.18%), and 75.86%-80.52% (3.05%-20.91%), respectively, compared with router-based, ring-based, group-based, cellular-based, and honeycomb-based methods. Therefore, the proposed architecture is the best choice in terms of yields, hardware overhead, and timing delay. Tianming Ni, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | High-Throughput Portable True Random Number Generator Based on Jitter-Latch StructureabstractUnder the requirement of highly reliable encryption, the design of true random number generators (TRNGs) based on field-programmable gate arrays (FPGAs) is receiving increased attention. Although TRNGs based on ring oscillators (ROs) and phase-locked loops (PLLs) have the advantages of small resource overhead and high throughput, there are problems such as instability of randomness and poor portability. To improve the randomness, portability, and throughput of a random number generator, we design a TRNG whose randomness is generated by the oscillation of self-timed rings (STRs) and accurately extracted by a jitter-latch structure. The portability of the structure is verified by electronic design automation (EDA) tools. Under the condition of 0°C-80°C ambient temperature and 1.0 ± 0.1 V output voltage, the proposed structure is tested many times on Xilinx Spartan-6 and Virtex-6 FPGAs with an automatic routing mode. Theoretical analysis shows that this method can effectively improve the coverage of jitter and reduce the migration phenomenon. Experimental results show excellent performance in randomness, robustness, and portability, and the throughput reaches 100 Mbps. Xinyu Wang 0027, Huaguo Liang, Maoxiang Yi, Zhengfeng Huang, Haochen Qi, Yingchun Lu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Pure Digital Scalable Mixed Entropy Separation Structure for Physical Unclonable Function and True Random Number GeneratorabstractThis study presents a pure digital scalable mixed entropy separation structure for the physical unclonable function (PUF) and true random number generator (TRNG), which is implemented on Xilinx field-programmable gate arrays (FPGAs). The mixed entropy separation structure in this study is to solve the problems of unstable output in the existing PUF structure and the poor scalability of the TRNG and PUF design. The proposed design has the following innovations: 1) concept of sensitive entropy and the corresponding processing method are proposed for the first time; 2) design does not need to modify the internal design of the entropy source, and it is suitable for most entropy source arrays; and 3) adjustable PUF bit width and TRNG throughput enhance the scalability of the architecture under different requirements. The structure is simulated and validated on two Xilinx FPGAs and tested under nominal working conditions. The results show that the 512-bit PUF entropy source after treatment is 92.7% more stable than the 1024-bit entropy source before treatment, the random number after treatment passes all types of randomness tests, and the minimum entropy is more than 0.8. Under various conditions within the ranges of 0 °C–80 °C and 0.8–1.2 V, the TRNG output remains stable after treatment; the maximum intra-Hamming distance of the treated PUF is 5.1028%, and the average intra-Hamming distance is 2.7065%. Yingchun Lu, Xinyu Wang 0027, Maoxiang Yi, Zhengfeng Huang, Huaguo Liang |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2021 | Reliability Evaluation and Analysis of FPGA-Based Neural Network Acceleration SystemabstractPrior works typically conducted the fault analysis of neural network accelerator computing arrays with simulation and focused on the prediction accuracy loss of the neural network models. There is still a lack of systematic fault analysis of the neural network acceleration system that considers both the accuracy degradation and system exceptions, such as system stall and running overtime. To that end, we implemented a representative neural network accelerator and corresponding fault injection modules on a Xilinx ARM-FPGA platform and evaluated the reliability of the system under different fault injection rates when a series of typical neural network models are deployed on the neural network acceleration system. The entire fault injection and reliability evaluation system is open-sourced on GitHub. With comprehensive experiments on the system, we identify the system exceptions based on the various abnormal behaviors of the FPGA-based neural network acceleration system and analyze the underlying reasons. Particularly, we find that the probability of the system exceptions dominates the reliability of the system. The faults also incur accuracy degradation of the neural network models, but the influence depends on the applications of the models and can vary greatly. In addition, we also evaluated the use of conventional triple modular redundancy (TMR) and demonstrated the challenge of TMR with both experiments and analytical models, which may shed light on the reliability design of the FPGA-based neural network acceleration system. Dawen Xu 0002, Ziyang Zhu, Cheng Liu 0008, Ying Wang 0001, Lei Zhang 0008, Huaguo Liang, Huawei Li 0001, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2020 | Multi-task Scheduling for PIM-based Heterogeneous Computing SystemabstractProcessing-in-Memory (PIM) or Near-Data Processing has been recognized as the most potential solution to resolve the ever-aggravating memory wall especially as the thrive of memory-intensive scale-out workloads such as graph computing and data analytics. However, when the future computing system becomes more and more likely to adopt PIM architectures as a type of the storage and processing component, there is a lack of literature and research work on the general scheduling framework with the emerging heterogeneous system except for some ad-hoc task partitioning methods with specialized PIM designs. This work is the first to propose a formalized model to quantitatively describe the multi-task scheduling problem in PIM+CPU platform without loss of generality, and also an optimized task mapping-and-scheduling algorithm to boost the hardware utility for these novel heterogeneous systems. The proposed scheduling framework is fully aware of the data access bandwidth and processing capability distinction between the CPU and PIM devices, and also the implications of task mapping on the bandwidth contention, data communication intensity and hardware utility for the concurrent workloads. Experimental results show that, compared to the traditional scheduling algorithm for heterogeneous system, the proposed method is able to improve the system performance by over 10% and the energy efficiency by almost 10% for multi-core scale-out applications. Dawen Xu 0002, Cheng Chu, Cheng Liu 0008, Ying Wang 0001, Xianzhong Zhou, Lei Zhang 0008, Huaguo Liang, Huawei Li 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | A Hybrid Computing Architecture for Fault-tolerant Deep Learning AcceleratorsabstractRegular 2D computing array is widely utilized for the processing of the major neural network operations in many deep learning accelerators (DLAs). Hardware failures on the array can lead to considerable computing errors and prediction accuracy loss. Prior works proposed to add homogeneous redundant PEs to each row or column of the regular computing array to mitigate faulty PEs, but they may fail to recover the computing array from faults when the number of faulty PEs in a row or column exceeds the number of redundant PEs in the corresponding row or column. The problem gets worse when the faults are not evenly distributed across the computing array. To address the problem, we propose a hybrid computing architecture (HCA) for fault-tolerant DLAs. Instead of adding homogeneous redundant PEs to the regular computing array of DLAs, it has a dot-production processing unit (DPPU) to recompute the operations that are mapped to the faulty PEs concurrently without performance penalty under moderate fault injection. Even under high fault injection, HCA can be degraded smoothly and remains functional. In addition, DPPU exploits the parallelism within each operation and processes the network operations sequentially, so it can tolerate faulty PEs in arbitrary locations and ensures steady performance under distinct fault distributions. According to our experiments, HCA shows significantly higher reliability and performance under various fault injection with comparable chip area penalty compared to the conventional redundancy approaches. Dawen Xu 0002, Cheng Chu, Cheng Liu 0008, Ying Wang 0001, Lei Zhang 0008, Huaguo Liang, Kwang-Ting Cheng |
ICCD | 7 |
| 2020 | Dual-Interlocked-Storage-Cell-Based Double-Node-Upset Self-Recoverable Flip-Flop Design for Safety-Critical ApplicationsabstractThis paper presents a novel dual-interlocked storage-cell (DICE)-based double-node-upset (DNU) self-recoverable, namely DURI-FF, in the nano-scale CMOS technology. The master latch of the DURI-FF cell consists of three transmission gates (TGs) and three interlocked DICEs with three common nodes. The common nodes are connected to TGs for value initialization. The slave latch of the DURI-FF cell comprises six TGs, six inverters and three interlocked DICEs. The outputs of the inverters respectively feed the internal nodes of the slave latch. The interlocked DICEs make the master latch and the slave latch DNU self-recoverable. Simulation results validate the DNU self-recoverability of the proposed DURI-FF cell. Moreover, compared with the state-of-the-art hardened flip-flop cells, the proposed DURI-FF cell achieves roughly 43% delay reduction at the cost of moderate silicon area and power dissipation. Aibin Yan, Zhelong Xu, Jie Cui 0004, Zuobin Ying, Zhengfeng Huang, Huaguo Liang, Patrick Girard 0001, Xiaoqing Wen |
ISCAS | 6 |
| 2020 | LCHR-TSV: Novel Low Cost and Highly Repairable Honeycomb-Based TSV Redundancy Architecture for Clustered FaultsabstractDue to the winding level of the thinned wafers and the surface roughness of silicon dies, the quality of through-silicon vias (TSVs) varies during the fabrication and bonding process. If one TSV exhibits a defect during its manufacturing process, the probability of multiple defects occurring in the TSVs neighboring the faulty TSV increases, i.e., the TSV defects tend to be clustered, which significantly reduces the yield of 3-D integrated circuit. To resolve the clustered TSV faults, router-based, ring-based, group-based, and cellular-based redundant TSV (RTSV) architectures were proposed. However, the repair rate is low and the hardware overhead as well as delay overhead is high. In this article, we propose a honeycomb-based RTSV architecture to utilize the area and delay more efficiently as well as to maintain high yield. The simulation results show that the proposed architecture has a 99.84% repair rate for uniform faults and an 81.42% repair rate for highly clustered faults. The proposed design achieves a 51.66% reduction of hardware overhead compared with the router-based design and a 20.69%, 46.93%, 34.17%, and 11.15% reduction of total delay compared with ring-based, router-based, group-based, and cellular-based methods, respectively. Tianming Ni, Huaguo Liang, Aibin Yan, Zhengfeng Huang, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Architecture of Cobweb-Based Redundant TSV for Clustered FaultsabstractIn this brief, a cobweb-based redundant through-silicon-via (TSV) design is proposed with efficient hardware as well as high repair rate to repair clustered faulty TSVs (FTSVs). The experimental simulation results demonstrate that for highly clustered faults, the repair rate of the proposed RTSV method is 48.59% and 1.75% higher than that of the ring-based and router-based RTSV methods, respectively. Furthermore, the proposed design can achieve 63.93% and 16.34% hardware reductions compared with the router-based and the ring-based design, respectively. Tianming Ni, Dongsheng Liu 0001, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | A Novel Triple-Node-Upset-Tolerant CMOS Latch Design using Single-Node-Upset-Resilient CellsabstractNano-scale CMOS circuits are vulnerable to single-event triple-node-upsets (SETUs). This paper proposes the design of a novel CMOS latch to tolerate any SETU using single-node-upset-resilient cells converged at a highly reliable node. The latch makes use of three single-node-upset-resilient cells, each of which mainly consists of triple mutually feeding back 2-input C-elements. These cells have a common converged output node feeding back to the output of the latch, making the latch capable of tolerating any SETU. Simulation results not only confirm the SETU tolerance capability but also show a significant area-power-delay-product reduction of 96.81% for the proposed latch compared with the only existing SETU hardened latch. Zhiyuan Song, Aibin Yan, Jie Cui 0004, Xiaoqing Wen, Chaoping Lai, Zhengfeng Huang, Huaguo Liang |
ITC-Asia | 9 |
| 2019 | Novel Application of Deep Learning for Adaptive Testing Based on Long Short-Term MemoryabstractAdaptive testing is a promising approach that practically ensures cost reduction and reliability for test strategy. In adaptive testing, the test content or pass/fail limits are not fixed as in conventional test, but depend on other test results of the currently or historically tested data. Based on recent progress in machine learning, a new Long Short-Term Memory (LSTM) which is more advanced than simple Recurrent Neuron Network (RNN) is proposed for defect screening. The simulation results have been compared with the other deep learning and traditional methods when patterns are increased and decreased. The comparisons show that the proposed RNN-based LSTM method has achieved remarkable improvements, i.e. 4.3% accuracy improvement and 2.32s time reduction during the test process. Tai Song, Huaguo Liang, Zhengfeng Huang, Maoxiang Yi, Xiangsheng Fang, Aibin Yan |
VTS | 2 |
| 2019 | DVFS Based Error Avoidance Strategy in Wireless Network-on-Chip
Qi Wang 0027, Lizhu Hu, Huaguo Liang |
J. Electron. Test. | 4 |
| 2019 | CPCA: An efficient wireless routing algorithm in WiNoC for cross path congestion awareness
Jianhua Li 0003, Chenglong Sun, Huaguo Liang, Gaoming Du |
Integr. | 5 |
| 2019 | A Pulse Shrinking-Based Test Solution for Prebond Through Silicon via in 3-D ICsabstractSince the physical defects such as resistive open and leakage in through silicon vias (TSVs) caused by immature manufacturing techniques tend to undermine the reliability and yield of 3-D integrated circuits, it is very important to test the TSV as early as possible in the fabrication process. There are some shortcomings in the existing prebond TSV test techniques, such as incomprehensive fault coverage, large area overhead, and additional test time. To overcome these problems, a noninvasive solution for prebond TSV test based on pulse shrinking is proposed in this paper. This method makes use of the fact that defects in TSV lead to variation in the propagation delay-the rise and fall times are first transformed into pulse width, and the pulse shrinking technique is used to digitize the pulse width into a digital code which is then compared with an expected value for a fault-free TSV. Experiments on defect detection are carried out using HSPICE simulations with realistic models for 45-nm CMOS technology. The results show that the proposed method performs better than the existing methods in terms of fault coverage, area overhead, and test time. Maoxiang Yi, Jingchang Bian, Tianming Ni, Cuiyun Jiang, Huaguo Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2018 | A Dictionary-Based Test Data Compression Method Using Tri-State CodingabstractWith the rapid development of integrated circuit manufacturing processes, the degree of integration of system-on-chip(SoC) has increased dramatically. In this paper, a dictionary-based test data compression method with tri-state coding is proposed to reduce the increasing test data volume. Firstly the partial input reduction is used to preprocess the test set, and then the tri-state coding is presented to mark the index so that it can encode anywhere in the test set. Experimental results on ISCAS'89 benchmark circuits show that the average compression rate of the proposed scheme reaches 73.92% by increasing the hardware area overhead slightly. Chenxin Lin, Huaguo Liang, Fuji Ren |
ATS | 3 |
| 2018 | A Hybrid DMR Latch to Tolerate MNU Using TDICE and WDICEabstractWith technology scaling, nanoscale CMOS becomes more sensitive to Multiple Node Upsets (MNUs). This paper presents a Multiple Node Upsets Tolerant Hardened Latch based on hybrid Double Modular Redundancy. The proposed latch consists of two elementary cells derived from DICE: one cell is referred to as TDICE cell with four additional NMOS transistors in the feedback lines, the other cell is referred to as WDICE cell with two additional NMOS transistors and two additional PMOS transistors in the feedback lines. Additional transistors in the feedback line of DICE cell improves the resilience to multiple-node upset. Extensive simulation results show the proposed latch can tolerate the DNU with the probability of 100%, and tolerate the TNU with the probability of 95.70%. Also the proposed latch can make a good tradeoff among area, delay, power and robustness. Zhengfeng Huang, Zian Su, Huaguo Liang, Huijie Yao, Tianming Ni |
ATS | 4 |
| 2018 | A Low-Cost High-Efficiency True Random Number Generator on FPGAsabstractTrue random number generator (TRNG), essential component in cryptographic equipment, which can generate unpredictable and irreproducible key string has an important effect on information encryption. In this work, a novel low-cost, high-efficiency true random number generator based on the ring oscillator is implemented on FPGAs. Forming a tapped delay line by utilizing the fast carry logic on FPGA, we have improved the efficiency of the entropy extraction from the jitter of single transition event rather than multiple jitter accumulation like most RO-based TRNGs. In order to achieve low cost and high throughput, the delay of the ring oscillator has been optimized by deeply studying LUT structure and routing resources. The proposed architecture has been validated on Xilinx Virtex-6 FPGA, which obtains a high throughput of about 100 Mbps while occupying just 25 slices and provides robustness across a wide range of temperature (0 °C ~ 80 °C), voltage (0.9 V ~ 1.1 V) and process variation (multiple chips). And the generated random bitstreams have passed all tests in the NIST statistical test suite. Gaoliang Ma, Huaguo Liang, Zhengfeng Huang, Maoxiang Yi, Xiumin Xu |
ATS | 2 |
| 2018 | An All-Digital and Jitter-Quantizing True Random Number Generator in SRAM-Based FPGAsabstractThis paper describes a novel all-digital true rand-om number generator (TRNG) in SRAM-based field programable gate arrays (FPGAs), which utilizes vernier technique to high precisely quantize random edge jitter caused by thermal noise in order for on-die entropy extraction. The TRNG is implemented in three ML605 platforms and experimental result shows that the TRNG presents a high quality of randomness (passing all NIST random tests with high p-values), a high throughput of 127 Mbps, and a good tolerance to bias phenolmenon induced by process, voltage, and temperature (PVT) variations. Xiumin Xu, Huaguo Liang, Gaoliang Ma, Zhengfeng Huang, Maoxiang Yi, Tianming Ni, Yingchun Lu |
ATS | 2 |
| 2018 | A High Reliability FPGA Chip Identification Generator Based on PDLsabstractPhysical Unclonable Functions (PUFs) promise cheap, efficient, and secure identification and authentication of devices, especially in FPGAs, which have been widely used. Various PUF implementation techniques have been proposed to translate chip-specific variations into a unique chip ID. It is difficult to guarantee the stability of the chip ID generation due to the complex operating environment. To solve this problem, in this paper, the Programmable Delay Lines (PDLs) was utilized to configure the ring oscillator to improve the stability of ID gener-ation. Compared with the original RO PUF, the proposed structure does not add extra overhead, but instead saves resources due to the compact layout. Experimental results demonstrate that the chip ID generated by our configurable ring oscillator (RO) PUFs is random (passing the NIST randomness test), and multiple measurements under a wide range of operating environments show that the proposed PUF is highly reliable (the bit flip rate is reduced from approximately 1.0% to 0 at nominal temperature and voltage conditions). Huaguo Liang, Zhengfeng Huang, Maoxiang Yi, Xiumin Xu |
ATS | 2 |
| 2018 | Novel low cost and DNU online self-recoverable RHBD latch design for nanoscale CMOSabstractThis paper presents a novel low cost and double node upset (DNU) online self-recoverable latch design using radiation hardening by design (RHBD) technology. The latch mainly consists of 8 interlocked input-split inverters. Since all internal nodes are interlocked, if any of the possible node pairs occurs a DNU, the latch can restore back. Simulation results have demonstrated the DNU online self-recoverability and also demonstrated that the proposed latch design saves 71.68% transmission delay, 72.92% power dissipation and 93.69% comprehensive delay-power-area product (DPAP) on average, compared with the up-to-date DNU online self-recoverable latch designs. Aibin Yan, Chaoping Lai, Yinlei Zhang, Chunming Liu, Zhile Chen, Jie Cui 0004, Huaguo Liang |
ISCAS | 9 |
| 2018 | MTTF-Aware Reliability Task Scheduling for PIM-Based Heterogeneous Computing SystemabstractProcessing-in-Memory (PIM) has been recognized as the most feasible solution to resolve the ever-aggravating memory wall especially as the boom of memory-intensive scale-out workloads such as graph computing and data analytics. However, when the future computing system becomes more and more likely to adopt PIM architectures as a type of the storage and processing unit, existing aging-award task scheduling algorithms for heterogeneous systems do not consider memory interference in PIM+CPU system, deducing an inaccurate task runtime and temperature which will over-estimate MTTF. We proposed a quantitatively formalized model for the aging reliability of PIM+CPU heterogeneous system and MTTF-ALG (a MTTF-based task scheduling algorithm) to balance the MTTF of whole system. Experimental results show that, compared to the traditional scheduling algorithm for heterogeneous system, the proposed method is able to reduce MTTF variation over 60.2% on average and the runtime by 15.3% on average for PIM+CPU system. Desong Pang, Dawen Xu 0002, Ying Wang 0001, Huaguo Liang |
ITC-Asia | 4 |
| 2018 | An improved communication scheme for non-HOL-blocking wireless NoC
Kun Xing, Zhengfeng Huang, Huaguo Liang |
Integr. | 5 |
| 2017 | HLDTL: High-performance, low-cost, and double node upset tolerant latch designabstractThis paper presents a high-performance, low-cost, and double node upset (DNU) tolerant latch design. The latch mainly constructs from a 3-input Muller C-element at the output stage and a single node upset resilient cell for keeping data, and the cell mainly consists of triple mutual feedback 2-input Muller C-elements, thus the latch is DNU tolerant. Using fewer CMOS transistors, clock gating technique, and high-speed transmission path, the latch also performs with lower cost penalties. Simulation results have demonstrated the DNU tolerability and a ~97.78% area-power-delay product saving for the latch design on average compared with the DNU tolerant latch designs. Aibin Yan, Zhengfeng Huang, Maoxiang Yi, Jie Cui 0004, Huaguo Liang |
VTS | 5 |
| 2017 | Double-Node-Upset-Resilient Latch Design for Nanoscale CMOS TechnologyabstractThis brief presents a double-node-upset-resilient latch (DNURL) design in 22-nm CMOS technology. The latch comprises three interlocked single-node-upset-resilient cells and each of the cells mainly consists of three mutually feeding back Muller C-elements. Simulation results demonstrate the double-node upset resilience and a 73.0% delay-power-area product saving on average compared with the up-to-date DNURL designs. Aibin Yan, Zhengfeng Huang, Maoxiang Yi, Xiumin Xu, Huaguo Liang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2016 | Novel Low Cost and Double Node Upset Tolerant Latch Design for Nanoscale CMOS TechnologyabstractThis paper presents a novel low cost and double node upset tolerant latch design in 22nm CMOS technology. The latch mainly comprises a single node upset resilient cell which feeds back to a 3-input Muller C-element at output stage. Simulation results demonstrate the double node upset tolerance and an 81.2% area-power-delay product saving for the latch design on average. Aibin Yan, Zhengfeng Huang, Xiangsheng Fang, Huaguo Liang |
ATS | 5 |
| 2015 | MTTF-Aware Reliability Task Scheduling for Heterogeneous Multicore System
Huaguo Liang, Yangyang Dai, Maoxiang Yi, Dawen Xu 0002, Zhengfeng Huang |
ICA3PP (2) | 1 |
| 2015 | Pulse shrinkage based pre-bond through silicon vias test in 3D ICabstractDefects in TSV not only lead to variation in the propagation delay but also in the transition delay of the net connected to the TSV. A non-invasive approach for pre-bond TSV test based on pulse shrinkage is proposed to detect resistive open and leakage fault. TSVs are used as capacitive loads of their driving gates, then the pulse visiting the cyclic shrinkage cells will be shrunk until it vanishes completely. The shrinkage amount is digitized into a digital code to compare with an expected value of fault free. Experiments on fault detection are presented through HSPICE simulations using realistic models for a 45 nm CMOS technology. The results show the effectiveness in the detection of resistive open defects 0.2kΩ above and equivalent leakage resistance less than 40MΩ. The estimated design for testability area cost of our method is negligible for realistic dies. Chang Hao, Huaguo Liang |
VTS | 2 |
| 2015 | A High Performance SEU Tolerant Latch
Zhengfeng Huang, Huaguo Liang, Sybille Hellebrand |
J. Electron. Test. | 2 |
| 2014 | Design of a Radiation Hardened Latch for Low-Power CircuitsabstractAs technology node entered the era of nanotechnology, a latch is much more susceptible to soft errors caused by energetic particles in space radiation environment. In order to enhance the Single Event Upset (SEU) -tolerance capability of a latch, this paper presents an interlocking soft error hardened latch (ISEHL) which is suitable for low-power circuits. The proposed latch is based on three C-elements which are errors tolerable, and the logic state of each C-element is determined by the output state of two other C-elements, which constitute an interlocking soft error hardened latch. The simulation results show that the proposed ISEHL latch can not only be applied to clock-gating circuits but also perform with 41% power as well as 95% Power Delay Product (PDP) saving as comparing with the FERST latch which performs an equivalent superior SEU-tolerance ability. Huaguo Liang, Zhengfeng Huang, Aibin Yan |
ATS | 1 |
| 2013 | A dynamic self-adaptive correction method for error resilient applicationabstractThe aggressive scaling down technology has posed transistor aging to be a new challenging to the reliability of circuits. Transistor aging could cause the gradual degradation of circuit performance and eventually lead to timing error. In this paper, a dynamic self-adaptive method is proposed to protect the circuit from the influence of transistor aging. This makes use of aging detection sensors and self-adaptive clock scaling cell. Aging sensors would automatically wake up the clock scaling cell to shift the clock phase of circuits when an error occurs. Then the timing error would be masked by a second sampling with the shifted clock. The method is simulated by Hspice using 65nm technology. The evaluation results show that this method is effective to error resilient with no impact on normal function of circuits, and it improves the MTTF by 1.16 times with 22.73% circuit overheads on average when the phase difference is 20% clock cycle. Luming Yan, Huaguo Liang, Zhengfeng Huang |
DATE | 2 |
| 2010 | A Novel x -ploiting Strategy for Improving Performance of Test Data CompressionabstractA precomputed core test set contains a large number of don't cares (x's) that can be effectively exploited to improve test data compression (TDC). Extending pattern run-length coding, we present a novel strategy that propagates thex's of a reference pattern to a new reference pattern in such a way that the reference pattern is xor-ed with the pattern to be encoded. Thex-propagating strategy can increase the probability of a reference pattern being coding-compatible with the pattern to be encoded, and its validity can be established by filling somex's of the already encoded patterns in backtracing way. How our strategy is used for TDC is demonstrated. Experimental results for large ISCAS89 benchmarks show that, compared to the recently proposed schemes, our technique can effectively improve compression and simplify on-chip decoder, and work better when used for core-unified TDC. Maoxiang Yi, Huaguo Liang, Lei Zhang 0008, Wenfa Zhan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | A Test Vector Compression/Decompression Scheme Based on Logic Operation between Adjacent Bits (LOBAB) CodingabstractA new test vector compression/decompression scheme, namely a scheme of logic operation between adjacent bits (LOBAB) is presented, which is based on bitwise logic operation between itself and its previous bit. It turns all kinds of series including continuous series, such as a series of all 0s and all 1s, and reversal series, such as a series of 01 and 10, into series of all 0s by logic operation between adjacent bits. On one hand, the two kinds of series, continuous series and reversal series, are both taken into account, which decreases the number of division to the original test data. On the other hand, all series are turned into series of all 0s, which eases the process of encoding and decoding. Compared with other already known schemes this scheme has some characteristics, such as high compression ratio, easy control and implementation. The performance of the algorithm is mathematically analyzed and its merits are experimentally confirmed on the larger examples of the ISCAS89 benchmark circuits. Huaguo Liang, Wenfa Zhan, Cuiyun Jiang |
PRDC | 1 |
| 2009 | Impact of Hazards on Pattern Selection for Small Delay DefectsabstractHazards ubiquitously exist in combinational circuits, and then should be taken into account for delay testing. This paper analyzes the impact of hazards on small-delay defect (SDD) detection, and presents a new test pattern selection method considering hazards. The concept of arrival time window is introduced and the concept of output deviation is redefined to accurately reflect the pattern capability on SDD detection. A new signal transition probability calculation method is presented to calculate output deviation more practical than that without considering hazards. Patterns from an N-detect test set for transition faults are then selected according to their output deviations. Experimental results show that, for the same pattern count, the patterns selected by the proposed method excite more long paths, and are capable of detecting more small delay defects at the early stage of delay testing compared to the method without considering hazards. Huawei Li 0001, Yinghua Min, Xiaowei Li 0001, Huaguo Liang |
PRDC | 5 |
| 2008 | A New Radiation Hardened by Design Latch for Ultra-Deep-Sub-Micron TechnologiesabstractSoft errors induced by cosmic radiation have become an urgent issue for ultra-deep-sub-micron (UDSM) technologies. In this paper, we propose a new radiation hardened by design latch (RHBDL). RHBDL can improve robustness by masking the soft errors induced by SEU and SET. We evaluate the propagation delay, power dissipation and power delay product of RHBDL using SPICE simulations. Compared with existing reported solutions such as TMR-latch, RHBDL is less SEU-sensitive, more area efficient, delay and power efficient. Zhengfeng Huang, Huaguo Liang |
IOLTS | 2 |
| 2007 | Block Marking and Updating Coding in Test Data Compression for SoCabstractA novel test data compression coding scheme, block marking and updating coding, is proposed in this paper. Test data in the test set was divided into successive fixed-length vectors, called blocks, and then they were marked according to their compatibility compared with a reference vector. An operation that is similar to difference and a technique that strategically fills in don 't-care bits are combined to increase the probabilities of compatibility or inverse compatibility. It effectively compresses test data and its decompression structure is very simple. Experimental results of ISCAS-89 benchmark circuits show that the scheme is very effective. Lei Zhang 0008, Huaguo Liang, Wenfa Zhan, Cuiyun Jiang |
ATS | 2 |
| 2007 | A Novel Collaborative Scheme of Test Data Compression Based on Fixed-Plus-variable-Length CodingabstractA novel collaborative scheme of test data compression based on fixed-plus-variable-length (FPVL) coding is presented, with which the test data can be compressed efficiently. In this scheme, code words are divided into fixed-length head and variable-length tail. In order to obtain further compression, the highest bit of the tail is reduced from the code words, because all of the highest bits in the tail section of the tail are the same as 1. A special shift counter is also used, which further eases the control circuit. Experimental results of the Mintest fault sets of part ofISCAS-89 benchmark circuits show that the proposed scheme is obviously better than traditional coding methods in the compression ratio and the implementation of decompression, such as Golomb, FDR, VIHC, v9C coding. Wenfa Zhan, Huaguo Liang, Zhengfeng Huang |
CSCWD | 2 |
| 2007 | Test data compression scheme based on variable-to-fixed-plus-variable-length coding
Wenfa Zhan, Huaguo Liang, Zhengfeng Huang |
J. Syst. Archit. | 2 |
| 2005 | A BIST Scheme Based on Selecting State Generation of Folding CountersabstractIn this paper, a BIST scheme based on selecting state generation of folding counters is presented. LFSR is used to encode the seeds of the folding counters, where folding distances (or indexes) are stored to control deterministic test patterns generation, so that the generated test set is completely equal to the original test set. This scheme solves compression of the deterministic test set and overcomes overlapping and redundancy of test patterns produced by the different seeds. Experimental results prove that it not only achieves higher test data compression ratio, but also efficiently reduces test application time, and that the average test application time is only four percent of that of the same type scheme. Huaguo Liang, Maoxiang Yi, Xiangsheng Fang, Cuiyun Jiang |
Asian Test Symposium | 1 |
| 2003 | Sharing BIST with Multiple Cores for System-on-a-ChipabstractA novel architecture based on mixed mode BIST for sharing among multiple logic cores on an system-on-a-chip is presented. In the architecture a single-polynomial LFSR with maximum degree in the multiple cores can be selected to generate pseudo-random patterns to cover the easy to detect faults for the all cores. For the remaining faults of the each core deterministic test patterns can be compressed by a two-dimensional compression scheme, where the LFSR encodes the seeds of a folding counter as the seeds of the LFSR so as to reduce amount of test data storage, and all of the cores under test can use the unique LFSR to decompress the encoded seeds. Experimental results indicate that the proposed scheme can achieve a significant amount of compression for test data storage, and the simple and flexible architecture can be directly embedded on chip for systems-on-a-chip test. Huaguo Liang, Cuiyun Jiang |
Asian Test Symposium | 1 |
| 2002 | Two-Dimensional Test Data Compression for Scan-Based Deterministic BIST
Huaguo Liang, Sybille Hellebrand, Hans-Joachim Wunderlich |
J. Electron. Test. | 1 |
| 2002 | A Mixed-Mode BIST Scheme Based on Folding Compression
Huaguo Liang, Sybille Hellebrand, Hans-Joachim Wunderlich |
J. Comput. Sci. Technol. | 1 |
| 2001 | Two-dimensional test data compression for scan-based deterministic BISTabstractA novel architecture for scan-based mixed mode BIST is presented. To reduce the storage requirements for the deterministic patterns it relies on a two-dimensional compression scheme, which combines the advantages of known vertical and horizontal compression techniques. To reduce both the number of patterns to be stored and the number of bits to be stored for each pattern, deterministic test cubes are encoded as seeds of an LFSR (horizontal compression), and the seeds are again compressed into seeds of a folding counter sequence (vertical compression). The proposed BIST architecture is fully compatible with standard scan design, simple and flexible, so that sharing between several logic cores is possible. Experimental results show that the proposed scheme requires less test data storage than previously published approaches providing the same flexibility and scan compatibility. Huaguo Liang, Sybille Hellebrand, Hans-Joachim Wunderlich |
ITC | 1 |
| 2001 | A Mixed Mode BIST Scheme Based on Reseeding of Folding Counters
Sybille Hellebrand, Huaguo Liang, Hans-Joachim Wunderlich |
J. Electron. Test. | 2 |
| 2000 | A mixed mode BIST scheme based on reseeding of folding countersabstractIn this paper a new scheme for deterministic and mixed mode scan-based BIST is presented. It relies on a new type of test pattern generator which resembles a programmable Johnson counter and is called folding counter. Both the theoretical background and practical algorithms are presented to characterize a set of deterministic test cubes by a reasonably small number of seeds for a folding counter. Combined with classical approaches for test width compression and with pseudorandom pattern generation these new techniques provide an efficient and flexible solution for scan-based BIST. Experimental results show that the proposed scheme outperforms previously published approaches based on the reseeding of LFSRs or Johnson counters. Sybille Hellebrand, Hans-Joachim Wunderlich, Huaguo Liang |
ITC | 3 |