VLDB 2026 Research / reviewers in the wild / expert
Tianming Ni
dblp:193/1937
· DBLP profile ↗
63ranked-venue papers
11as first author
55since 2021 · last 2026
0000-0001-6272-8660ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 10 first-author · 48 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Meta-Learning with Attentional Prototypes for Few-Shot Wafer Defect Recognition
Tianming Ni, Meifang Yu, Huaguo Liang, Senling Wang, Muyang Cheng, Mu Nie |
J. Electron. Test. | 1 |
| 2026 | A Configurable Delay Transient-Effect Ring Oscillator PUF against modeling attacks
Tianming Ni, Mu Nie, Senling Wang, Jingchang Bian |
Integr. | 1 |
| 2026 | Adversarial noise-perturbed feature fusion for deep multi-view clustering with joint optimization
Yong Wang 0008, Lihao Yang, Yazhou Ren 0001, Yourui Huang, Guifu Lu, Tianming Ni |
Pattern Recognit. | 7 |
| 2026 | Automated Co-Optimization Framework of Feature Selection and Ensemble Learning for Wafer Yield PredictionabstractEfficient and automated wafer yield prediction is central to cost control and process optimization in intelligent semiconductor manufacturing. As the initial testing stage in wafer inspection, Wafer Acceptance Testing (WAT) data contain critical process-related information. However, its high dimensionality, redundancy, and nonlinear inter-dependencies pose significant challenges to conventional yield prediction models, such as high computational overhead and limited generalization capability. Moreover, existing approaches often lack an automated framework capable of jointly addressing feature redundancy and model complexity. This paper proposes a collaborative and automated prediction framework that integrates feature selection and machine learning. First, an enhanced Binary Zebra Population Optimization Algorithm (BZPOA) is introduced, which incorporates redesigned exploration and development mechanisms to automatically identify key feature subsets from high-dimensional parameters, substantially reducing data redundancy and computational dimensionality. Second, a Bayesian hyperparameter-optimized XGBoost model is constructed, utilizing the Tree-structured Parzen Estimator (TPE) to achieve deep co-optimization of model parameters and feature space, thereby overcoming the inefficiency and overfitting issues commonly associated with manual parameter tuning. Experiments on a real-world dataset demonstrate that the proposed framework achieves average Recall, Precision, and F1-scores of 0.872, 0.917, and 0.902, respectively. Compared with the full-feature baseline, the BZPOA-selected feature subset improves predictive performance by 5%–10%, attains an AUC of 0.904, and significantly reduces per-wafer prediction time. Cross-factory transfer experiments further confirm the robustness of the proposed system. Tianming Ni, Muyang Cheng, Jingchang Bian, Senling Wang, Xiaoqing Wen, Mu Nie |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Unifying Modality and Scale: Visual Mamba for Feature Fusion in RGB-X Crowd CountingabstractCrowd counting has long been a crucial topic in the domains of computer vision and video surveillance. In particular, with the widespread use of thermal or depth cameras, RGB-X crowd counting has emerged as a prominent research focus. Although depth or infrared images provide complementary information, the core challenge remains in effectively unifying heterogeneous cross-modality and cross-scale information to form a comprehensive representation of crowd distributions in complex scenes. To address this problem, we propose a novel Mamba-based framework, termed UMS-VMamba-CC for multi-modal (RGB-X) crowd counting. Specifically, we design the cross-modality disentanglement fusion visual Mamba (CMDF-VMamba) that uses self-supervised learning to decompose modality-invariant and modality-specific features in spatial-frequency domains, followed by multi-modal feature aggregation via the gating mechanism. For cross-scale fusion, we design the cross-scale pyramid fusion visual Mamba (CSPF-VMamba), which adopts the bi-directional pyramid structure to incorporate low-level features into high-level representations during downsampling, while upsampling high-level features and aggregating them with low-level features through state space contextual modeling. Comprehensive experiments on multiple mainstream datasets demonstrate that the UMS-VMamba-CC framework achieves competitive performance for RGB-X crowd counting. Yaocong Hu, Mengbo Jia, Pindeng Wang, Wenbo Zhu 0002, Huanjie Tao, Tianming Ni, Teng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | A Parallel Feedback Obfuscation Strong PUF Against Machine-Learning Modeling Attacks and Lightweight Authentication ProtocolabstractArbiter physical unclonable function (APUF) is a hardware security primitive that generates security keys by utilizing unavoidable process variations during chip manufacturing. However, the structure based on linear additive function makes it vulnerable to machine learning (ML) attacks. This paper proposes a parallel feedback obfuscation PUF (PFO PUF) design, which uses intermediate arbitration signals of the lower-layer APUF to generate the hidden challenge of the upper-layer APUF, enhancing the overall nonlinearity of the structure. The obfuscation module makes weight judgment for intermediate arbitration signals of upper-layer and lower-layer APUFs, which obfuscates the real response of PUF. We further design a variant of PFO PUF called reconfigured challenge obfuscation PFO PUF (RPFO PUF) and propose its lightweight device authentication protocol. RPFO PUF enhances the resistance of the original PFO PUF against reverse engineering (RE) and improves its Strict Avalanche Criterion (SAC) characteristic by reordering the challenges and incorporating weak PUF responses. The proposed PFO PUF and RPFO PUF were comprehensively evaluated via Python-based simulations and FPGA measurements. In Python simulations, both designs show strong resistance to state-of-the-art ML attacks, with logistic regression (LR), support vector machine (SVM), and covariance matrix adaptation evolution strategies (CMA-ES) yielding near 50% prediction accuracies under various PUF configurations. Although deep neural network (DNN) achieves up to 69.52% prediction accuracy on the PFO PUF, it drops to ∼50% on the RPFO PUF. FPGA results further confirm this, with the (32, 11)-RPFO PUF achieving a maximum prediction accuracy of only 51.47% across all four ML attacks. Moreover, both designs incur low hardware overheads, requiring just 743 and 2145 gate equivalents (GEs), respectively. Zhengfeng Huang, Yankun Lin, Yingchun Lu, Huaguo Liang, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2026 | EMFFTrans: Efficient Multi-Scale Feature Fusion Transformer for Road Scene Semantic Segmentation
Yaocong Hu, Pindeng Wang, Mengbo Jia, Jinwen Hong, Guoyang Wan, Huicheng Yang, Bingyou Liu, Tianming Ni, Teng Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 11 |
| 2026 | A Novel Approach to Reducing Testing Costs and Minimizing Defect Escapes Using Dynamic Neighborhood Range and Shapley ValuesabstractWafer acceptance testing (WAT) is a process that is used to assess the quality and reliability of manufactured wafers. This technique for the early detection and screening of chips allows for improvements in their reliability and performance during the manufacture of semiconductor devices. The automatic test equipment (ATE) used for processing millions of wafers is susceptible to a number of issues, including the absence of data values, the presence of redundant parameters, and categorical imbalance. These issues increase the cost of data processing and impede an investigation into the relationship between WAT and feature diagnostics. In this study, we propose a method with a low test escape rate based on a multi-objective optimization algorithm to reduce the cost of testing and minimize the number of defective dice that go undetected. The proposed method retains outliers, dynamically selects the range of the neighborhood to reduce the cost of testing, and uses Shapley values to analyze a WAT dataset to determine the importance of features of the data. The multi-objective optimization algorithm ranks features by their importance and applies an adaptive method to eliminate features with a low overall correlation, thereby reducing the risk that defective dice are undetected. Tianming Ni, Wangsheng Rui, Cheng Zhuo, Yu Li 0007, Xiaoqing Wen, Mu Nie |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2026 | Functional Fault Impact Probability Prediction using Spatio-Temporal Graph Convolutional NetworkabstractLogic-level defects that escape manufacturing tests pose reliability risks in modern systems and require functional testing to identify their activation and propagation behaviors. However, effective functional testing is limited by the high cost of long-cycle fault simulations. To address this challenge, we propose a Spatio-Temporal Graph Convolutional Network framework to efficiently and accurately predict the Fault Impact Probability on the circuit's function cross-over multiple function cycles, enabling rapid quantitative assessment of functionally possible faults. Our method represents gate-level netlists as spatio-temporal graphs, capturing both structural connectivity and short-range signal-propagation dynamics. With dedicated spatial and temporal encoders, the proposed ST-GCN enables accurate prediction of multi-cycle circuit-level FIP. Experiments on ISCAS’89 benchmarks show that the approach reduces fault-simulation cost by over an order of magnitude while maintaining high accuracy (mean absolute error as low as 0.024 for 5-cycle predictions). The framework supports both testability-metric-based and simulation-based feature construction, enabling a tunable balance between efficiency and accuracy. A case study on test point selection further demonstrates that using predicted FIPs to guide observation-point placement improves the detectability of multi-cycle, hard-to-detect circuit-level faults. Overall, this work provides a scalable solution for circuit-level multi-cycle fault-impact assessment and can be readily integrated into functional test generation and other Electronic Design Automation workflows. Shaoqi Wei, Senling Wang, Hiroshi Kai, Yoshinobu Higami, Ruijun Ma 0002, Tianming Ni, Xiaoqing Wen, Hiroshi Takahashi |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2026 | RMC PUF: A Highly Reliable PUF Architecture Based on Recursive Markov Chain ObfuscationabstractPhysical unclonable functions (PUFs) are critical hardware security primitives that extract entropy from intrinsic manufacturing process variations (MPVs) to generate unique cryptographic responses. They have demonstrated extensive application prospects in lightweight encryption and authentication for Internet of Things (IoT) devices. However, the classic arbiter PUF (APUF) is inherently vulnerable to machine learning (ML) modeling attacks due to the linear mathematical model of its delay circuits. To address this challenge, this article proposes a novel postprocessing obfuscation architecture named recursive Markov chain PUF (RMC PUF). The proposed architecture couples an APUF core with Markov obfuscation modules. By exploiting the intermediate stage responses of the APUF as an entropy source, the system constructs a 1-D recursive stochastic state transition mechanism. This approach recursively transforms and obfuscates the raw entropy based on Markov Chain theory, achieving a PUF system with stability and robust resistance against ML attacks. Through this efficient recursive obfuscation strategy, the design achieves precise utilization of the APUF’s intrinsic entropy. The work is validated through both numerical simulation and FPGA prototyping. Experimental results show that the prototype maintains a high reliability of 98.9%–99.9% within a voltage range of 0.8–1.2 V, and a reliability of 96.1%–99.7% within a temperature range from$- 10~{^{\circ }}$C to$80~{^{\circ }}$C. Comparative analysis indicates that the proposed RMC PUF significantly outperforms existing structures in security, limiting the prediction accuracy of four ML attack models [logistic regression (LR), artificial neural networks (ANNs), deep neural networks (DNNs), and covariance matrix adaptive evolutionary strategy (CMA-ES)] to no more than 54.72%. Zhengfeng Huang, Xuxiang Sun, Yingchun Lu, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2026 | NVLIM: MTJ and CMOS-Based Nonvolatile Latch Design With Protection Against Triple-Node-Upsets for Robust ComputingabstractSoft errors and power dissipation emerge as critical challenges in developing high-reliability and cost-sensitive embedded systems. To address these issues, the magnetic tunnel junction (MTJ) is considered a promising solution due to its nonvolatility and its compatibility with traditional CMOS manufacturing processes. In this work, we propose a novel nonvolatile (NV) latch consisting of inverters and MTJs, namely, NVLIM, which provides nonvolatility and robust partial tolerance against triple-node-upsets (TNUs) at low cost. NVLIM integrates a TNU-tolerant block based on CMOS with a backup-restore block using MTJs. Simulation results incorporating process, voltage, and temperature (PVT) variations, bias temperature instability (BTI) impact, and Monte Carlo simulations demonstrate the balanced performance in terms of nonvolatility, robust partial TNU tolerance, and comprehensive overhead of the proposed latch. Aibin Yan, Litao Wang, Zhengfeng Huang, Qingyang Zhang 0001, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Software-Defined Secure Island for Testing Chiplet SystemsabstractChiplet systems that stack heterogeneous dies via 2.5D/3D integration use JTAG-based test access ports (TAPs) to validate inter-die links and enable in-field diagnosis. However, these TAPs create a shared attack surface: penetrating a single die can potentially expose control of the entire stack. Current countermeasures involve embedding a complete cryptographic engine in each chiplet, which increases the area and locks the protocol at tape-out, leaving it susceptible to future unknown attacks. This paper proposes a Software-Defined Secure Island (SDSI) architecture that decouples security policies from hardwired logic while satisfying non-functional requirements. Each chiplet instantiates an ultra-lightweight Secure Island Controller (SIC) macro that performs only two XORs and two additions per handshake. Meanwhile, a centralized secure island runs heavyweight cryptography in firmware using the multi-round SASL-JTAG+ protocol. SDSI enables scalable, adaptable test access protection by decoupling security from hardware. An FPGA implementation shows that the SIC macro occupies 11% of the area of an AES- 128 core and 37% of a SHA256 core, yet it supports 256 - to 512-bit keys with negligible growth. A security analysis demonstrates immunity to replay attacks because the authentication data is refreshed with each session. All future upgrades, such as longer keys, stronger hashes, and additional rounds, are delivered via firmware, providing scalable, field-upgradable protection for heterogeneous chiplet systems. Hisashi Okamoto, Senling Wang, Hiroshi Kai, Hiroyuki Yotsuyanagi, Yoshinobu Higami, Tianming Ni, Tai Song, Hiroshi Takahashi, Xiaoqing Wen |
ATS | 6 |
| 2025 | H3Match: A Hybrid Heterogeneous Hypergraph Matching Method for Subcircuit IdentificationabstractSubcircuit matching is widely applied in logic synthesis, design verification, hardware security, etc. Previous works employ redundant circuit representations, coupled with timeconsuming enumerative search methods. Subsequent works use a hybrid “approximate filtering - exact verification” framework, but the numerous false negatives predicted by the graph neural network (GNN) based filtering lead to severe matching failure. In this paper, an improved hybrid method named H3Match is proposed to achieve a better tradeoff between runtime, accuracy, and false negative rate. First, we model the circuits as hypergraphs to fully capture the topology and construct diverse heterogeneous hyperedge features to facilitate the learning of circuit topologies. Second, to reduce the false negatives, we reformulate the subgraph matching problem as matching directed acyclic graphs (DAGs) with embedded circular structure information and develop a directed GNN-based approximate matching approach to identify potential matching subcircuits. Finally, we propose a general mixed integer nonlinear programming (MINLP) formulation for exact verification, with convergency speed accelerated by extracting the initial solution from the results in approximate matching. Experimental results show that our approximate method outperforms state-of-the-art (SOTA) methods by 4.31% in accuracy while achieving virtually zero false negatives. Our exact verification is on average $4.16 \times$ faster than SOTA exact methods. Overall, the end-to-end flow achieves a $7.08 \times$ speedup compared to existing approaches. Qingsong Peng, Tianming Ni, Tinghuan Chen, Qi Sun 0002, Cheng Zhuo |
DAC | 4 |
| 2025 | Efficient Modulated State Space Model for Mixed-Type Wafer Defect Pattern RecognitionabstractAccurate and efficient wafer defect detection is crucial in semiconductor manufacturing to maintain product quality and optimize yield. Traditional methods struggle with the complexity and diversity of modern wafer defect patterns. While deep learning approaches are effective, they are often resource-intensive, posing challenges for real-time deployment in industrial settings. To solve these problems, we propose an Efficient Modulated State Space Model (EM-SSM) for mixed-type wafer defect recognition, optimized with knowledge distillation to balance accuracy and efficiency. Our framework captures size-dependent relationships and improves defect-specific feature representation to recognize complex defects precisely. Specifically, we introduce an efficient directional modulation mechanism to refine spatial recognition of defect patterns. To further improve inference efficiency, we propose a deep-to-shallow distillation method that transfers knowledge from deeper networks to lighter networks, reducing inference time without compromising classification accuracy. Experimental results on the MixedWM38 wafer dataset with 38 defect types show that our model achieves 99.0% accuracy, outperforming traditional methods in both accuracy and efficiency. Our model offers a scalable solution for modern semiconductor defect detection. Mu Nie, Shidong Zhu, Aibin Yan, Cheng Zhuo, Xiaoqing Wen, Tianming Ni |
DATE | 6 |
| 2025 | NVSRLO: A FeFET-Based Non-Volatile and SEU-Recoverable Latch Design with Optimized OverheadabstractThis paper presents a FeFET-based non-volatile and single-event upset (SEU) recoverable latch, namely NVSRLO, which does not require any extra control signals. Simulation results show that the proposed latch provides non-volatility and SEU-recovery with optimized overhead. Compared with existing non-volatile latches, NVSRLO significantly reduces delay, power, and delay-power-area product at the cost of area. Aibin Yan, Wangjin Jiang, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
DATE | 5 |
| 2025 | A lightweight general PUF framework for resisting machine learning attacks
Tianming Ni, Zhengfeng Huang, Aibin Yan, Senling Wang, Xiaoqing Wen, Mu Nie, Jingchang Bian |
Integr. | 1 |
| 2025 | Improve SAC in PUFs: Metric, Analysis, Algorithm, and ApplicationabstractPhysical Unclonable Functions (PUFs) are crucial for lightweight authentication in the Internet of Things (IoT), but existing PUFs often have poor statistical properties and are vulnerable to machine learning attacks. Designs implementing the Strict Avalanche Criterion (SAC) lack sufficient theoretical foundation. This paper introduces quantitative metrics to evaluate the SAC performance of strong PUFs and conducts rigorous analysis on Arbiter PUFs (APUFs) and their classic variants, addressing imprecision in existing methods. Based on these metrics, we developed an algorithm to optimize the SAC performance of strong PUFs by adjusting the challenge sequences. This optimization improved the SAC performance of the 2-XOR APUF by 59% without additional resource consumption, making it comparable to the 4-XOR APUF; We also provided mathematical proof for the optimal solution. Furthermore, we propose the SAC Optimized Shuffled XOR Arbiter PUF (SOS XOR APUF), which improves SAC performance by 84% compared to the 3-XOR APUF with the same entropy source. It addresses the inherent defect of poor statistical properties when two adjacent bits in the challenge flip, achieving a theoretical response flip probability of 0.5. The SOS XOR APUF resists existing machine learning attacks---including Logistic Regression (LR), Covariance Matrix Adaptation Evolution Strategy (CMA-ES), Artificial Neural Network (ANN), and Deep Neural Network (DNN)---with prediction accuracy below 55%. Finally, we designed and verified an authentication protocol based on this PUF in the ProVerif environment, achieving mutual authentication between IoT devices and servers, preventing secret information from being stolen, and enhancing the security of the PUF structure. Zhengfeng Huang, Fansheng Zeng, Jingchang Bian, Huaguo Liang, Yingchun Lu, Tianming Ni |
IEEE Internet Things J. | 7 |
| 2025 | Cost Efficient Flip-Flop Designs With Multiple-Node Upset-Tolerance and Algorithm-Based VerificationsabstractThis article presents radiation-hardened flip-flop (FF) designs capable of tolerating soft errors, e.g., single-node upsets (SNUs), double-node upsets (DNUs) and multiple-node upsets (MNUs). First, a 2-input FF and a 3-input FF are proposed as the baseline FFs that not only, respectively, tolerate SNUs and DNUs but also exhibit cost efficiency in terms of delay, power, and area. Through adding two stages of c-elements, a 4-input FF and a 5-input FF are proposed as the baseline FFs as well. Utilizing the structural characteristics of these FFs, an$N-1$input FF and an N input FF are proposed as the extended FFs capable of tolerating more node upsets. Moreover, a highly efficient algorithm for verifying MNU-tolerance of these FFs is proposed. Algorithm and HSPICE-tool-based verification results both demonstrate the MNU-tolerance for the proposed FFs with more inputs. Aibin Yan, Zhengfeng Huang, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | BF PUF: A Modeling Attack-Resistant Strong PUF Based on Bent FunctionsabstractStrong physical unclonable functions (PUFs) are promising circuits for lightweight Internet of Things (IoT) authentication and security. However, existing strong PUFs exhibit very low cryptographic nonlinearity (NL), making them vulnerable to machine learning (ML) modeling and cryptanalytic attack. To address this issue, we propose the Bent function PUF (BF PUF) based on Maiorana-McFarland (M-M) constructed Bent functions, which obfuscates the responses of the strong PUF to enhance resistance against modeling attacks. The core idea is to employ the M-M construction method for Bent functions to ensure maximum cryptographic NL to resist modeling attacks. A Feistel network is configured using weak PUF responses as keys to achieve device-specific and unpredictable mappings of input challenges while meeting the requirements of the M-M Bent function construction. A Python-based model of the BF PUF was developed, and simulation results indicate that the cryptographic NL of the proposed BF PUF outperformsk-xorarbiter PUFs (APUFs) (${k} =2$, 4, 6). The proposed BF PUF was also implemented and evaluated on the FPGA hardware platform. The experimental results show that under modeling attacks using four ML algorithms—logistic regression (LR), artificial neural networks (ANNs), deep neural networks (DNNs), and covariance matrix adaptation evolution strategies (CMA-ES)—the best prediction accuracy under these four modeling attack algorithms is 52.60%. The reliability under temperature fluctuations ranging from$- 10~^{\circ }$C to$80~^{\circ }$C is between 84.20% and 99.78%. Zhengfeng Huang, Fansheng Zeng, Yanqiao Chi, Yankun Lin, Yingchun Lu, Huaguo Liang, Jingchang Bian, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2025 | A TSV Misalignment-Based Repair Architecture in 3-D ChipsabstractAs a critical component of 3-D integrated circuits (3D-ICs), the quality of through-silicon vias (TSVs) significantly impacts the yield and reliability of 3D-ICs, especially the clustered faults during manufacturing. In this article, a repair architecture based on TSV misalignment is proposed. This architecture achieves a higher repair rate by physically connecting the signal not to its closest TSV but only to the TSVs far away from each other. Experimental results show that the average repair rate of the proposed architecture increases by 13.42% compared to the existing repair architectures of the same type for clustered faults. Compared to the router-based architecture, the proposed architecture has a similar average repair rate with less than 0.15% difference in fewer than eight clustered faults, reducing the delay and MUX area overhead by 70.27% and 54.17%, respectively. Huaguo Liang, Jiahui Xiao, Xianrui Dou, Tianming Ni, Yingchun Lu, Zhengfeng Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A Block-Group-Based Redundant TSV Architecture for Clustered FaultsabstractThree-dimensional integrated circuits (3-D ICs) based on through-silicon vias (TSVs) have tremendous advantages, such as high performance, low power consumption, and heterogeneous integration. However, TSV faults significantly diminish the yield of 3-D ICs. To address this issue, this article proposes a block-group-based redundant TSV (RTSV) architecture. In this architecture, TSVs are organized into multiple blocks, with four blocks constituting a basic unit for the repair of faulty TSVs (FTSV). Within each basic unit, four TSVs situated at the same position across four blocks form a group that shares an RTSV. This methodology effectively improves the repair rate for clustered TSV faults through cross-block and cross-group connections among TSVs. In addition, we propose a TSV repair algorithm that dynamically updates the difficulty values (DUDVs) of FTSVs. This approach uses quadtrees with FTSVs as the root node to compute the repair difficulty values, selecting the most difficult TSV for repair in each iteration. The difficulty values are continuously updated throughout the repair process to further improve the repair rate. Experimental results show that the proposed method achieves repair rates of 98.7% and 99.9% for random and clustered TSV faults, respectively. Jun Liu 0070, Tianhao Du, Mulin Ye, Xi Wu 0003, Tianming Ni, Huaguo Liang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | A Response-Nonlinearized DEMUX-TDC PUF for Resistance Against Modeling Attacks and Secure Authentication ProtocolsabstractAs a critical hardware security primitive for the authentication within the Internet of Things (IoT), the physical unclonable function (PUF) represents an innovative security design paradigm for integrated circuits. However, the linear challenge-response mapping of the arbiter PUF (APUF) and variants render these structures more susceptible to modeling attacks due to their delayed linear structure. In this article, we propose a nonlinearized demultiplexer time-to-digital converter (DEMUX-TDC) PUF. This PUF uses a quantized delay difference technique to alter the traditional response generation mechanism, demonstrating robust resistance against modeling attacks. First, the proposed PUF employs a segmented APUF variant structure configured in both front and back segmented modes to generate a source of delay difference entropy. Additionally, the scheme incorporates a multilayer differential tapped TDC circuit to quantize the delay differences into digital codes, followed by a linear feedback shift register (LFSR) to obfuscate the final output response. We further propose a highly secure mutual authentication protocol based on reconfigurable PUF, by leveraging the characteristics of front and back segments of the PUF’s challenge. Evaluation of the proposed scheme on the implementation of Xilinx Virtex-7 and Spartan-6 field-programmable gate array (FPGA) demonstrates that the uniqueness and uniformity can reach ideal value, in the condition of prediction accuracy across six modeling attacks remaining around 50%. Tianming Ni, Mu Nie, Aibin Yan, Senling Wang, Xiaoqing Wen, Jingchang Bian |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Cost-Optimized Double-Node-Upset-Recovery Latch Designs With Aging Mitigation and Algorithm-Based Verification for Long-Term Robustness EnhancementabstractWith the continuous advancement of CMOS technologies, soft errors, such as single-node upset (SNU) and double-node upset (DNU), caused by radiation in nanoscale integrated circuits, are becoming increasingly prominent. Meanwhile, transistor aging mitigation is indispensable for long-term robustness enhancement. First, to reduce the impact of radiation on circuits, we propose a novel DNU-recovery latch with low cost, namely, DURLC, only consisting of four dual-input C-elements (CEs) and four clock-gated input-split inverters for the storage of values. Second, we propose a DNU-recovery latch with moderate cost, namely, DURMC, based on seven CEs and four inverters, for convenience to optimize the latch to alleviate aging. The proposed DNU-recovery latch with mitigated aging is called DURMA. The latch employs a high-speed path to reduce delay without sacrificing performance when mitigating aging issues. Finally, we propose an algorithm-based verification method to validate the DNU recovery of the proposed latches. The simulation results show that, compared with the state-of-the-art robust latches, the proposed latches have the advantages of DNU recovery with moderate and even low cost, and meanwhile, aging is effectively mitigated for the DURMA latch. Aibin Yan, Changli Hu, Na Bai, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Design of Nonvolatile and Multinode-Upset Recoverable Latches Based on Magnetic Tunnel Junction and CMOSabstractSpintronic devices, such as magnetic tunnel junctions (MTJs), are promising for space applications due to their radiation hardness and nonvolatility. However, as semiconductor technology advances, CMOS peripheral circuits are becoming vulnerable to double node upset (DNU) as well as triple node upset (TNU). This brief proposes two nonvolatile and robust latch designs primarily composed of MTJs and C-elements (CEs). Both designs offer nonvolatility and self-recovery from multiple-node upsets. Simulation results demonstrate that the proposed latches provide nonvolatility and complete protection against multiple-node upsets with balanced overhead. Aibin Yan, Yongkang Xu, Na Bai, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Reconfigurable Fault-Tolerant Link With Bandwidth Expansion for 2.5-D Chiplet-Based SystemsabstractThe 2.5-D chiplet-based systems offer a promising path toward higher performance and integration density, but the reliability of interchiplet links, particularly vertical links (VLs), poses a significant challenge. This article proposes reconfigurable bidirectional link (ReBL), a novel architecture providing robust fault tolerance and enhanced bandwidth for these critical interconnects. ReBL features three key innovations: 1) a robust fault-tolerant mechanism leveraging ReBLs that dynamically adapts to and mitigates permanent link failures; 2) a dynamic bandwidth expansion technique that significantly enhances the system efficiency by utilizing idle links to optimize resource allocation and throughput; and 3) a virtual channel (VC) allocation strategy that guarantees deadlock-free operations through strategic channel partitioning and assignment. Evaluations using synthetic traffic and PARSEC benchmarks demonstrate ReBL’s significant advantages under high fault rates. Compared with the state-of-the-art reliable and deadlock-free routing (ReD) approach, ReBL achieves an average reduction of 21.3% in packet latency and 4.9% in application execution time across the evaluated benchmarks. These benefits are achieved with only a 6.06% area overhead over baseline. Wu Zhou 0007, Le Luo 0002, Fulong Chen 0002, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | IDLD: Interlocked Dual-Circle Latch Design with Low Cost and Triple-Node-Upset-Recovery for Aerospace ApplicationsabstractModern powerful CMOS chips are usually highly integrated and implemented with aggressively shrunk technology nodes. In radiation environment, under charge-sharing mechanism, one particle striking can simultaneously impact multiple nodes causing double-node-upsets (DNUs) and triple-node-upsets (TNUs). In this paper, we propose an Interlocked Dual-circle Latch Design, namely IDLD, with low cost and TNU recovery for aerospace applications. IDLD consists of four transmission gates and twelve 2-input C-elements (CEs) implemented in 22nm CMOS process. Simulation results demonstrate the complete TNU recovery as well as cost-effectiveness for the proposed IDLD latch. Aibin Yan, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 6 |
| 2024 | PFO PUF: A Lightweight Parallel Feed Obfuscation PUF Resistant to Machine Learning AttacksabstractArbiter Physically Unclonable Functions (APUFs) are hardware security primitives that leverage manufacturing process variation to generate security keys. They can produce exponential challenge-response pairs (CRPs) with minimal hardware overhead. However, the symmetric nature of linear additive functions makes them vulnerable to modeling attacks rooted in machine learning. To address this issue, this paper introduces a novel design called Parallel Feed Obfuscation PUF (PFO PUF). In this approach, the intermediate decision signals from the lower APUF are used as a concealed challenge for the upper APUF, enhancing the overall nonlinearity of the dual-APUF. Additionally, obfuscation modules are employed to determine the weights of the intermediate decision signals from both the upper and lower APUFs, protecting the actual response. Experimental results demonstrate that the proposed PFO PUF effectively withstands four advanced machine learning attack algorithms, including Logistic Regression (LR), Support Vector Machine (SVM), Deep Feedforward Neural Network (DFNN), and Efficient CANDECOMP/PARAFAC Tensor Regression Network (ECPTRN). The prediction accuracy of these four algorithms is consistently below 66.30%. Compared with other enhanced structures based on APUF, PFO-PUF only uses 493 LUTs and has lower resource overhead. Zhengfeng Huang, Yankun Lin, Fansheng Zeng, Jingchang Bian, Huaguo Liang, Yingchun Lu, Xiaoqing Wen, Tianming Ni |
ITC-Asia | 10 |
| 2024 | Test Point Selection for Multi-Cycle Logic BIST using Multivariate Temporal-Spatial GCNsabstractThis paper proposes a novel Test Point Insertion (TPI) strategy to enhance the testability for multi-cycle Built-In Self-Test (BIST) for logic circuits. The approach leverages Multivariate Temporal-Spatial Graph Convolutional Neural Networks (MTS-GCN) and Reinforcement Learning to identify optimal Test Points (TPs). The proposed TPI method treats the testability information of a logic circuit as time-series data and employs Multivariate Time-Series Graph Neural Networks (MTGNN) to capture the relationship between the circuit's structural (spatial information) attributes and the temporal variability of signal line testability across capture cycles. A subsequent Multi-Layer Perceptron (MLP) computes the metric for each signal line to pinpoint potential TPs based on the extracted temporal-spatial features. Experimental evaluation based on benchmark circuits confirms the efficacy of the proposed model, which is trained with Deep Q-Networks (DQN), in improving the fault detection for multi-cycle logic BIST. Senling Wang, Shaoqi Wei, Hisashi Okamoto, Tatusya Nishikawa, Hiroshi Kai, Yoshinobu Higami, Hiroyuki Yotsuyanagi, Ruijun Ma 0002, Tianming Ni, Hiroshi Takahashi, Xiaoqing Wen |
ITC-Asia | 9 |
| 2024 | A Quadruple-Node Upsets Hardened Latch Design Based on Cross-Coupled Elements
Zhengfeng Huang, Zishuai Li, Huaguo Liang, Tianming Ni, Aibin Yan |
J. Electron. Test. | 5 |
| 2024 | Design Guidelines and Feedback Structure of Ring Oscillator PUF for Performance ImprovementabstractThe physical unclonable function (PUF) is a hardware security primitive that is used to generate secret keys or identity authentication for chips using random manufacturing process variation (MPV). The PUF based on ring oscillator (RO PUF) has been extensively studied in recent years because of its high robustness and ease of design. Although the performance has been optimized in previous studies, several uniqueness, reliability, and theoretical foundation concerns still remain. This article presents a transistor-level parameters-based quantitative theoretical model, which clearly reveals several design guidelines for improving the reliability of RO PUF. Furthermore, a PUF based on a feedback ring oscillator (RO) structure is proposed, which combined the RO topology and the drafting effect of XOR gates to enhance the uniqueness and reliability. The correctness of the theoretical model was verified by the SPICE simulation experiment result. And in the FPGA experiment result, the uniqueness and the reliability of feedback RO PUF using the same hardware resources on the same chip was superior to that of RO PUF. The theoretical research method of RO PUF used can be widely applied to other PUFs using ring topology and feedback RO PUF is a great substitute for RO PUF. Zhengfeng Huang, Jingchang Bian, Yankun Lin, Huaguo Liang, Tianming Ni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | MURLAV: A Multiple-Node-Upset Recovery Latch and Algorithm-Based Verification MethodabstractIn advanced CMOS technologies, integrated circuits are sensitive to multiple-node-upsets (MNUs) induced in harsh radiation environments. The existing verification of the reliability of latches highly relies on electronic design automation (EDA) tools considering complex error-injection scenarios. In this paper, we propose a novel latch, namely MURLAV, protected against quadruple node-upsets (QNUs) induced in harsh radiation environments, as well as an algorithmic error-recovery verification method. The latch provides complete recovery from all QNUs with a formed redundant structure. The algorithm can simplify the verification process and demonstrate the QNU recovery for the proposed MURLAV latch. Simulation results demonstrate that the proposed latch can recover from any QNU and that it has lower area and delay overhead. Compared with existing latches of the same type, the proposed MURLAV latch achieves an overhead reduction of 34% in silicon area and 15% in delay on average at the cost of moderate power consumption. Aibin Yan, Zhongyu Gao, Zhengfeng Huang, Tianming Ni, Jie Cui 0004, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Introduction to the Special Issue on Design for Testability and Reliability of Security-aware HardwareabstractThe research on design for testability and reliability of security-aware hardware has been important in both academia and industry. With ever-growing globalization, commercial hardware design, manufacturing, transportation, and supply now involve many different countries, resulting in aggravated vulnerability from hardware design to manufacturing. Hardware with malicious purposes implanted from the third-party manufacturing process may control the operation of a circuit and tamper its functions, causing serious security issues. However, hardware includes not only devices and circuits but also systems. An important fact is that testability, reliability, and security technologies come from different design layers, but the impact evaluation is conducted at the system level. In other words, the testability, reliability, and security design of different layers can be carried out in a holistic manner to achieve optimization for the whole system. In addition, the testability, reliability, and security design technologies of each design layer can be collaboratively conducted to achieve better performance. The testability, reliability, and security tradeoff has garnered attention from academia and industry, particularly in the Post-Moore Era, due to the complexities and opportunities arising from new architectures and technologies. Tianming Ni, Xiaoqing Wen, Hussam Amrouch, Cheng Zhuo, Peilin Song |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2024 | Nonvolatile Latch Designs With Node-Upset Tolerance and Recovery Using Magnetic Tunnel Junctions and CMOSabstractAs semiconductor technologies scale down, radiative-particle-induced soft errors and static power consumption are becoming major concerns for digital circuits. Magnetic-tunnel-junctions (MTJs) are widely used to address these concerns. MTJs are nonvolatile (NV) and compatible with traditional CMOS processes. In this article, we first propose a double-node-upset (DNU) tolerant and NV latch, i.e., M-TPDICE-V2, providing high reliability. In addition, we further propose an advanced latch, namely, M-8C, that is able to completely recover from single-node upsets (SNUs) and DNUs. M-8C uses a DNU recovery module and a backup and restore module based on a pair of MTJs. Furthermore, we propose a universal backup and restore module suitable for any latch providing nonvolatility. We simulate the proposed latches using the Synopsys HSPICE tool with a 45-nm CMOS process model. Simulation results confirm the superior capabilities of our proposed M-TPDICE-V2 and M-8C latches. M-TPDICE-V2 exhibits strong SNU and DNU tolerance and nonvolatility, while the M-8C latch provides complete DNU recovery capabilities. Aibin Yan, Litao Wang, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Enhancing Defect Diagnosis and Localization in Wafer Map Testing Through Weakly Supervised LearningabstractDefect diagnosis and localization in wafer maps are crucial tasks in semiconductor manufacturing. Existing deep learning methods often require pixel-level annotations, making them impractical for large-scale deployment. In this paper, we propose a novel weakly supervised learning approach to achieving high-precision defect identification and effective localization with only image-level labels. By leveraging the information of defect types and locations, we introduce a weighted fusion of activation maps, called Class Activation Map (CAM), to highlight classspecific regions. We further enhance defect localization accuracy and completeness by employing optimized region growing operations to eliminate noise in defect regions. Moreover, we present an optimized inference method that provides meaningful visual explanations for defect recognition. Experimental results on real-world wafer map images demonstrate the effectiveness of our approach in accurately segmenting defect patterns with no pixel-level annotations. By training the model solely on wafer map image classification labels, our proposed model significantly improves defect recognition, facilitating efficient defect analysis in semiconductor manufacturing. The proposed weakly supervised learning approach offers a practical solution for defect diagnosis and localization, with the potential of widespread adoption in the semiconductor industry. Mu Nie, Wankou Yang, Senling Wang, Xiaoqing Wen, Tianming Ni |
ATS | 6 |
| 2023 | Advanced DICE Based Triple-Node-Upset Recovery Latch with Optimized Overhead for Space ApplicationsabstractWith the rapid advancement of CMOS technologies, integrated circuits are becoming more prone to soft errors, e.g., triple-node upsets (TNUs). In this paper, to effectively tolerate TNUs, an input-split C-element-based DICEs (IC-DICEs) based TNU-recovery latch is proposed. The latch employs three interlocked IC-DICEs to allow recovering from any TNU. Simulations demonstrate the TNU recovery of the latch, and also demonstrate that the proposed latch can reduce delay by 87.21%, area by 27.04%, and delay-area-power product (DAPP) by 87.44% on average, compared to the alternative latches. Aibin Yan, Xuehua Li, Zhongyu Gao, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen |
ATS | 5 |
| 2023 | Two Highly Reliable and High-Speed SRAM Cells for Safety-Critical Applications
Aibin Yan, Yang Chang, Jing Xiang, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 7 |
| 2023 | A Low Area and Low Delay Latch Design with Complete Double-Node-Upset-Recovery for Aerospace Applications
Aibin Yan, Shaojie Wei, Jinjun Zhang, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 6 |
| 2023 | Design of A Highly Reliable and Low-Power SRAM With Double-Node Upset Recovery for Safety-critical ApplicationsabstractFor high-speed operations, low power consumption and small silicon area, transistors are being scaled aggressively. Meanwhile, circuit reliability is facing greater challenges in advanced technologies. In this paper, a highly reliable and low-power SRAM with double-node-upset (DNU) recovery, namely HRLP16T, is proposed for safety-critical fields. HRLP16T can recover from single-node-upset (SNU) at all the sensitive nodes, and it has eight node pairs recoverable from DNUs. Simulation results demonstrate its advantages in terms of delay and power consumption over typical existing SRAM cell designs. Aibin Yan, Jing Xiang, Zhengfeng Huang, Tianming Ni, Jie Cui 0004, Patrick Girard 0001, Xiaoqing Wen |
ITC-Asia | 4 |
| 2023 | A Low Overhead and Double-Node-Upset Self-Recoverable LatchabstractWith the rapid advancement of semiconductor technologies, integrated circuits, especially storage elements (e.g., latches) have become increasingly vulnerable to soft errors. In order to effectively tolerate double-node-upsets (DNUs) caused by radiation and reduce the power and area of latches, this paper proposes a DNU self-recoverable latch with low overhead in terms of power and area. The proposed latch mainly comprises seven 2-input C-elements and two inverters to achieve DNU self-recovery. Simulation results show that the proposed latch can recover from all possible DNUs and that it can reduce delay by 45.7%, power by 29.1%, area by 65.9%, and area-power-delay-product by 87.4%, on average, compared to typical existing DNU self-recoverable latches. Aibin Yan, Tianming Ni, Jie Cui 0004, Zhengfeng Huang, Patrick Girard 0001, Xiaoqing Wen |
ITC-Asia | 3 |
| 2023 | LDAVPM: A Latch Design and Algorithm-Based Verification Protected Against Multiple-Node-Upsets in Harsh Radiation EnvironmentsabstractIn deep nano-scale and high-integration CMOS technologies, storage circuits have become increasingly sensitive to charge-sharing-induced multiple-node-upsets (MNUs) that include double, triple, and quadruple node-upsets. Currently, verifications for error recovery of existing latches highly rely on EDA tools with complex error-injection combinations. In this article, a latch design protected against MNUs in the harsh radiation as well as an algorithm-based verification process is proposed. Due to the constructed redundant feedback loops, the latch can completely recover from any MNU. Algorithm-based verification and simulations both demonstrate the MNU recovery of the proposed latch. Simulation results demonstrate the low area overhead of the proposed latch compared with the only one existing of the same type. Aibin Yan, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Design of True Random Number Generator Based on Multi-Ring Convergence Oscillator Using Short Pulse Enhanced RandomnessabstractThe entropy source structure with embedded XOR gates in a ring oscillator (RO) as a true random number generator (TRNG) can improve the speed of accumulating jitter in the oscillator. However, the XOR gate has a certain response time to the input change, and when the input changes too fast, the XOR gate will output short pulses. In this paper, we propose a TRNG design based on a multi-ring convergence oscillator (MRCO) making use of the characteristics of short pulses. We study the output of the XOR gate when facing different inputs. By modeling the time of a fibonacci ring oscillator (FIRO) as an example, we find that the loss of short pulses in an inverter chain is the reason for making the FIRO enter into periodic oscillation. This phenomenon suppresses the accumulation of jitter and occurs periodically in existing structures. Our proposed structure uses independent sub-rings to accumulate jitter, allowing the main-ring to quickly generate short pulses to provide analog randomness. The proposed TRNG design is implemented in Xilinx Virtex-6 FPGA. The experimental results show that it has the highest ratio of throughput rate to hardware resources. The generated random sequence pass both NIST SP800-22 test and NIST SP800-90B test. Tianming Ni, Qingsong Peng, Jingchang Bian, Zhengfeng Huang, Aibin Yan, Senling Wang, Xiaoqing Wen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | A Flexible and High-Performance Lattice-Based Post-Quantum Crypto Secure CoprocessorabstractProgress of quantum computing technology seriously threaten the industrial information security based on traditional public-key cryptosystem. Thus, the cryptosystem with anti-quantum attack characteristics is gradually becoming a significant research in the security field. In this article, a flexible and high-performance secure coprocessor is designed for security in industrial processes, which can execute the post-quantum cryptographic algorithm Saber efficiently. Custom instruction set and arithmetic accelerators are proposed to effectively optimize the flexibility of system architecture, and improve the performance of calculation. The hardware implementation results show that the maximum operating frequency of the coprocessor can reach 345 MHz. Compared with related state-of-the-art works, it achieves the highest operating frequency on the same Xilinx UltraScale+ FPGA platform, performing the encryption and decryption operations within 13.5 and 15.4μs, respectively. Meanwhile, this article achieves 1.7/3.1/5.9× area-time product improvements in look-up table flip-flop block memory storage with good flexibility. Dongsheng Liu 0001, Xiang Li 0220, Xingjie Liu, Jiahao Lu 0002, Xuecheng Zou, Ang Hu, Tianming Ni |
IEEE Trans. Ind. Informatics | 10 |
| 2023 | Worst-case Power Integrity Prediction Using Convolutional Neural NetworkabstractPower integrity analysis is an essential step in power distribution network (PDN) sign-off to ensure the performance and reliability of chips. However, with the growing PDN size and increasing scenarios to be validated, it becomes very time- and resource-consuming to conduct full-stack PDN simulation to check the power integrity for different test vectors. Recently, various works have proposed machine learning–based methods for PDN power integrity prediction, many of which still suffer from large training overhead, inefficiency, or non-scalability. Thus, this article proposed an efficient and scalable framework for the worst-case power integrity prediction, which can handle general tasks including dynamic noise prediction and bump current prediction. The framework first reduces the spatial and temporal redundancy in the PDN and input current vector and then employs efficient feature extraction as well as a novel convolutional neural network architecture to predict the worst-case power integrity. Experimental results show that the proposed framework consistently outperforms the commercial tool and the state-of-the-art machine learning method with only 0.63–1.02% mean relative error and 25–69× speedup for noise prediction and 0.22–1.06% mean relative error and 24–64× speedup for bump current prediction. Yufei Chen 0007, Yucheng Wang 0005, Tianming Ni, Zhiguo Shi 0001, Xunzhao Yin, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2022 | SCLCRL: Shuttling C-elements based Low-Cost and Robust Latch Design Protected against Triple Node Upsets in Harsh Radiation EnvironmentsabstractAs the CMOS technology is continuously scaling down, nano-scale integrated circuits are becoming susceptible to harsh-radiation induced soft errors, such as double-node upsets (DNUs) and triple-node upsets (TNUs). This paper presents a shuttle C-elements based low-cost and robust latch (namely SCLCRL) that can recover from any TNU in harsh radiation environments. The latch comprises seven primary storage nodes and seven secondary storage nodes. Each pair of primary nodes feeds a secondary node through one C-element (CE) and each pair of secondary nodes feeds a primary node through another CE, forming redundant feedback loops to robustly retain values. Simulation results validate all key TNUs' recoverability features of the proposed latch. Simulation results also demonstrate that the proposed SCLCRL latch can approximately save 29% silicon area and 47% D-Q delay on average at the cost of moderate power, compared with the state-of-the-art TNU-recoverable reference latches of the same-type. Aibin Yan, Shiwei Huang, Zijie Zhai, Xiangyu Cheng, Jie Cui 0004, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
DATE | 7 |
| 2022 | A Highly Robust, Low Delay and DNU-Recovery Latch Design for Nanoscale CMOS TechnologyabstractWith the advancement of semiconductor technologies, nano-scale CMOS circuits have become more vulnerable to soft errors, such as single-node-upsets (SNUs) and double-node-upsets (DNUs). In order to effectively tolerate DNUs caused by radiation and reduce the delay and area consumption of latches, this paper proposes a DNU resilient latch in the nanoscale CMOS technology. The latch mainly comprises four input-split inverters and four 2-input C-elements. Since all internal nodes are interlocked, the latch can recover from all possible DNUs. Simulation results show that, compared with the state-of-the-art DNU self-recovery latch designs, the proposed latch can save 64.51% transmission delay and 56.88% delay-area-power-product (DAPP) on average, respectively. Aibin Yan, Shaojie Wei, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 6 |
| 2022 | Cost-Optimized and Robust Latch Hardened against Quadruple Node Upsets for Nanoscale CMOSabstractWith the aggressive reduction of CMOS transistor feature sizes, the soft error rate of nano-scale integrated circuits increases exponentially. In this paper, we propose a novel cost-optimized and robust latch, namely CRLHQ, hardened against quadruple-node-upsets (QNUs) for nanoscale CMOS technologies. The latch mainly comprises a 5×5 matrix based on interlocked source-drain cross-coupled inverters to robustly store logic values. Owing to the redundant constructed feedback loops, the latch can recover from all possible QNUs. Simulation results demonstrate all key QNUs' recovery of the proposed CRLHQ latch. Simulation results also show that the proposed latch can approximately reduce the D-Q delay by 44.3%, the silicon area by 7.3% and the delay-area-power product (DAPP) by 14.2%, compared with the state-of-the-art same-type reference latches that can recover from any QNU. Aibin Yan, Shukai Song, Jixiang Zhang 0007, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
ITC-Asia | 6 |
| 2022 | Valid test pattern identification for VLSI adaptive test
Tai Song, Tianming Ni, Zhengfeng Huang, Jinlei Wan |
Integr. | 2 |
| 2022 | A double-node-upset completely tolerant CMOS latch design with extremely low cost for high-performance applications
Aibin Yan, Kuikui Qian, Tai Song, Zhengfeng Huang, Tianming Ni, Xiaoqing Wen |
Integr. | 5 |
| 2022 | Fortune: A New Fault-Tolerance TSV Configuration in Router-Based Redundancy StructureabstractIn three-dimensional integrated circuits (3D-ICs), through silicon via (TSV) is a critical technique in providing vertical connections. However, yield is one of the key obstacles to adopt the TSV-based 3D-ICs technology in the industry. Various fault-tolerance structures using redundant TSVs to repair faulty functional TSVs have been proposed in the literature for yield and reliability enhancement. However, the TSV repair paths under delay constraint cannot always be generated due to the lack of appropriate repair algorithms. In this article, we propose an effective TSV repair strategy for the router-based TSV redundancy architecture, taking into account the delay overhead. First, we prove that the router-based fault-tolerance structure configuration (RFSC) with the delay constraint is equivalent to the length-bounded multicommodity flow (LBMCF) problem. Then, an integer linear programming (ILP) formulation with acceptable scalability is presented to solve the LBMCF problem. The experimental results demonstrate that, compared with state-of-the-art fault-tolerance designs, the proposed ILP model can provide higher yield and lower delay overhead. Qi Xu 0004, Hao Geng, Tianming Ni, Song Chen 0001, Bei Yu 0001, Yi Kang, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | A 4NU-Recoverable and HIS-Insensitive Latch Design for Highly Robust Computing in Harsh Radiation EnvironmentsabstractThis paper proposes a 4-node-upset (4NU) recoverable and high-impedance-state (HIS) insensitive latch design, namely QRHIL, for highly robust computing in harsh radiation environments. The latch mainly comprises a 5×5 looped C-element matrix to store values and provide complete 4NU recovery. Owing to the multiple-level error-interception of the 5×5 C-element matrix, the latch can recover from all possible 4NUs; thus, the latch is insensitive to HIS. Simulation results demonstrate the 4NU-recovery of the proposed latch. The results also show that the latch can approximately save 46% D-Q delay and 46% CLK-Q delay owing to the use of a high-speed D-Q path and clock-gating, compared with the state-of-the-art 3NU-recoverable latch (TNURL) that is not 4NU-recoverable. Aibin Yan, Aoran Cao, Zhengzheng Fan, Zhelong Xu, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
ACM Great Lakes Symposium on VLSI | 5 |
| 2021 | Kelvin Bridge Structure Based TSV Test for Weak FaultsabstractDue to the immaturity of manufacturing process, TSV is vulnerable to a variety of defects, which brings new testing challenges. Most of the existing test methods are suffer from the test resolution and difficult to detect weak faults. Borrowing the wisdom of Kelvin Bridge, a non-invasive test method is proposed to detect resistive open fault and leakage fault. By adjusting the resistances on the bridge arm to make them change in equal proportion, the adverse effects of contact resistance and parasitic resistance on the wire can be eliminated. HSPICE simulation using 45 nm CMOS technology show that it can successfully detect resistive open fault above 0.1 $\Omega$ and leakage fault below 10 $ M\Omega$. The effectiveness of the test scheme is further proved by process-voltage-temperature (PVT) analysis. Chang Hao, Zhengfeng Huang, Tianming Ni |
ITC-Asia | 3 |
| 2021 | A N: 1 Single-Channel TDMA Fault-Tolerant Technique for TSVs in 3D-ICsabstractAs the number of 3D-IC stacks increases, defects of through silicon via (TSV) in manufacturing and bonding process seriously affect the yield and reliability of the chip. Comparing to discarding these defective ones, some faulttolerant architectures are proposed, however, these existing schemes have great hardware overhead. In the paper, an N:1 single-channel time division multiple access (TDMA) faulttolerant technique using redundant TSV to tolerate TSV defect is proposed. Data is grouped and transmitted through TSV by TDMA mechanism, which reduces the number of TSV. The number of groups depends on bandwidth and hardware. The N: 1 single-channel TDMA structure is designed to use a single TSV to accomplish the time-sharing transmission of a group of signals. Each signal TSV is equipped with an additional TSV to improve the fault-tolerant coverage. The functions are verified on 40 nm Xilinx virtex-6 FPGA. The simulation results of Design Compiler based on 45 nm PTM show that the fault coverage rate can be increased to 100%, and the area overhead is reduced by 60.8% compared with the existing methods. Huaguo Liang, Danqing Li, Tianming Ni, Zhengfeng Huang, Cuiyun Jiang |
ITC-Asia | 4 |
| 2021 | Parallel DICE Cells and Dual-Level CEs based 3-Node-Upset Tolerant Latch Design for Highly Robust ComputingabstractWith the rapid advancement of design and manufacturing technologies of nano-scale CMOS circuits, latches are becoming increasingly sensitive to multiple-node-upsets caused by harsh radiation effects. In this paper, a Parallel Dual-interlocked-storage-cells (DICEs) and Dual-level C-elements (CEs) based 3-node-upset (3NU)-Tolerant Latch, namely PDDCTL, design for highly robust computing, is proposed. The latch comprises five transmission gates, two DICEs and three CEs. Due to the use of two single-node-upset self-recoverable DICEs and three error-interceptive CEs, the latch can provide complete 3NU-tolerance with low cost. Simulation results not only confirm the 3NU-tolerance of the proposed latch but also demonstrate that the delay-power-area product of the PDDCTL latch is reduced by 68.82% on average compared with the state-of-the-art 3NU hardened latch designs. Aibin Yan, Zijie Zhai, Lele Wang 0011, Jixiang Zhang 0007, Ningning Cui, Tianming Ni, Xiaoqing Wen |
ITC-Asia | 6 |
| 2021 | Design of Radiation Hardened Latch and Flip-Flop with Cost-Effectiveness for Low-Orbit Aerospace Applications
Aibin Yan, Aoran Cao, Zhelong Xu, Jie Cui 0004, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
J. Electron. Test. | 5 |
| 2021 | A Cost-Effective TSV Repair Architecture for Clustered Faults in 3-D ICabstractDue to the winding level of the thinned wafers and the surface roughness of silicon dies, the through-silicon vias (TSVs) defect tend to be clustered, reducing the yield of 3-D integrated circuit significantly. To tackle this fault clustering problem, the existing TSV repair methods adopt the TSV redundancy idea, which brings a major cost to 3-D integration. In this brief, a honeycomb-TDMA TSV design is proposed to mitigate the impact of multiple clustered faults without the need of redundant TSVs (RTSVs), thereby decreasing the area overhead and enhances the yield. The yield of the honeycomb-TDMA architecture can achieve 91.38%-99.67% for different benchmark circuits from IWLS 2005, which has the highest yield. Furthermore, our design achieves total additional hardware (timing delay overhead) reduction by 83.70%-86.85% (46.01%-55.96%), 66.89%-73.25% (29.41%-38.49%), 68.02%-74.20% (41.40%-52.20%), 60.60%-68.18% (18.09%-33.18%), and 75.86%-80.52% (3.05%-20.91%), respectively, compared with router-based, ring-based, group-based, cellular-based, and honeycomb-based methods. Therefore, the proposed architecture is the best choice in terms of yields, hardware overhead, and timing delay. Tianming Ni, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | A Sextuple Cross-Coupled SRAM Cell Protected against Double-Node UpsetsabstractIn this paper, we propose a sextuple cross-coupled SRAM cell, namely SCCS18T, protected against double-node upsets. Since the proposed SCCS18T cell forms a large feedback loop for value retention and error interception, the cell can provide self-recoverability from any single-node upsets (SNUs) and partial double-node upsets (DNUs). Moreover, the proposed cell has optimized operation speed due to the use of six access transistors. Simulation results show that the SCCS18T cell can save approximately 65% read access time at the cost of 49% power dissipation and 50% silicon area on average, compared with typical hardened SRAM cells. Aibin Yan, Jun Zhou 0016, Jie Cui 0004, Tianming Ni, Xiaoqing Wen, Patrick Girard 0001 |
ATS | 5 |
| 2020 | Design of a Highly Reliable SRAM Cell with Advanced Self-Recoverability from Soft ErrorsabstractIn this paper, a highly reliable SRAM cell, namely SESRS cell, is proposed. Since the cell has a special feedback mechanism among its internal nodes and has more access transistors compared to a standard SRAM cell, the SESRS cell provides the following advantages: (1) it can self-recover from single node upsets (SNUs) and double-node upsets (DNUs); (2) it can reduce power consumption by 49.78% and silicon area by 7.92%, compared with the only existing SRAM cell which can self-recover from all possible DNUs. Simulation results validate the robustness of the proposed SESRS cell. Moreover, compared with the state-of-the-art hardened SRAM cells, the proposed SESRS cell can reduce read access time by 61.93% on average. Zhengda Dou, Aibin Yan, Jun Zhou 0016, Yuanjie Hu, Tianming Ni, Jie Cui 0004, Patrick Girard 0001, Xiaoqing Wen |
ITC-Asia | 6 |
| 2020 | Information Assurance Through Redundant Design: A Novel TNU Error-Resilient Latch for Harsh Radiation EnvironmentabstractIn nano-scale CMOS technologies, storage cells such as latches are becoming increasingly sensitive to triple-node-upset (TNU) errors caused by harsh radiation effects. In the context of information assurance through redundant design, this article proposes a novel low-cost and TNU on-line self-recoverable latch design which is robust against harsh radiation effects. The latch mainly consists of a series of mutually interlocked 3-input Muller C-elements (CEs) that forms a circular structure. The output of any CE in the latch respectively feeds back to one input of some specified downstream CEs, making the latch completely self-recoverable from any possible TNU, i.e., the latch is completely TNU-resilient. Simulation results demonstrate the complete TNU-resiliency of the proposed latch. In addition, due to the use of fewer transistors and a high-speed path, the proposed latch reduces the delay-power-area product by approximately 91 percent compared with the state-of-the-art TNU hardened latch (TNUHL), which cannot provide a complete TNU-resiliency. Aibin Yan, Yuanjie Hu, Jie Cui 0004, Zhengfeng Huang, Tianming Ni, Patrick Girard 0001, Xiaoqing Wen |
IEEE Trans. Computers | 6 |
| 2020 | LCHR-TSV: Novel Low Cost and Highly Repairable Honeycomb-Based TSV Redundancy Architecture for Clustered FaultsabstractDue to the winding level of the thinned wafers and the surface roughness of silicon dies, the quality of through-silicon vias (TSVs) varies during the fabrication and bonding process. If one TSV exhibits a defect during its manufacturing process, the probability of multiple defects occurring in the TSVs neighboring the faulty TSV increases, i.e., the TSV defects tend to be clustered, which significantly reduces the yield of 3-D integrated circuit. To resolve the clustered TSV faults, router-based, ring-based, group-based, and cellular-based redundant TSV (RTSV) architectures were proposed. However, the repair rate is low and the hardware overhead as well as delay overhead is high. In this article, we propose a honeycomb-based RTSV architecture to utilize the area and delay more efficiently as well as to maintain high yield. The simulation results show that the proposed architecture has a 99.84% repair rate for uniform faults and an 81.42% repair rate for highly clustered faults. The proposed design achieves a 51.66% reduction of hardware overhead compared with the router-based design and a 20.69%, 46.93%, 34.17%, and 11.15% reduction of total delay compared with ring-based, router-based, group-based, and cellular-based methods, respectively. Tianming Ni, Huaguo Liang, Aibin Yan, Zhengfeng Huang, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Architecture of Cobweb-Based Redundant TSV for Clustered FaultsabstractIn this brief, a cobweb-based redundant through-silicon-via (TSV) design is proposed with efficient hardware as well as high repair rate to repair clustered faulty TSVs (FTSVs). The experimental simulation results demonstrate that for highly clustered faults, the repair rate of the proposed RTSV method is 48.59% and 1.75% higher than that of the ring-based and router-based RTSV methods, respectively. Furthermore, the proposed design can achieve 63.93% and 16.34% hardware reductions compared with the router-based and the ring-based design, respectively. Tianming Ni, Dongsheng Liu 0001, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | A Pulse Shrinking-Based Test Solution for Prebond Through Silicon via in 3-D ICsabstractSince the physical defects such as resistive open and leakage in through silicon vias (TSVs) caused by immature manufacturing techniques tend to undermine the reliability and yield of 3-D integrated circuits, it is very important to test the TSV as early as possible in the fabrication process. There are some shortcomings in the existing prebond TSV test techniques, such as incomprehensive fault coverage, large area overhead, and additional test time. To overcome these problems, a noninvasive solution for prebond TSV test based on pulse shrinking is proposed in this paper. This method makes use of the fact that defects in TSV lead to variation in the propagation delay-the rise and fall times are first transformed into pulse width, and the pulse shrinking technique is used to digitize the pulse width into a digital code which is then compared with an expected value for a fault-free TSV. Experiments on defect detection are carried out using HSPICE simulations with realistic models for 45-nm CMOS technology. The results show that the proposed method performs better than the existing methods in terms of fault coverage, area overhead, and test time. Maoxiang Yi, Jingchang Bian, Tianming Ni, Cuiyun Jiang, Huaguo Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | A Hybrid DMR Latch to Tolerate MNU Using TDICE and WDICEabstractWith technology scaling, nanoscale CMOS becomes more sensitive to Multiple Node Upsets (MNUs). This paper presents a Multiple Node Upsets Tolerant Hardened Latch based on hybrid Double Modular Redundancy. The proposed latch consists of two elementary cells derived from DICE: one cell is referred to as TDICE cell with four additional NMOS transistors in the feedback lines, the other cell is referred to as WDICE cell with two additional NMOS transistors and two additional PMOS transistors in the feedback lines. Additional transistors in the feedback line of DICE cell improves the resilience to multiple-node upset. Extensive simulation results show the proposed latch can tolerate the DNU with the probability of 100%, and tolerate the TNU with the probability of 95.70%. Also the proposed latch can make a good tradeoff among area, delay, power and robustness. Zhengfeng Huang, Zian Su, Huaguo Liang, Huijie Yao, Tianming Ni |
ATS | 6 |
| 2018 | An All-Digital and Jitter-Quantizing True Random Number Generator in SRAM-Based FPGAsabstractThis paper describes a novel all-digital true rand-om number generator (TRNG) in SRAM-based field programable gate arrays (FPGAs), which utilizes vernier technique to high precisely quantize random edge jitter caused by thermal noise in order for on-die entropy extraction. The TRNG is implemented in three ML605 platforms and experimental result shows that the TRNG presents a high quality of randomness (passing all NIST random tests with high p-values), a high throughput of 127 Mbps, and a good tolerance to bias phenolmenon induced by process, voltage, and temperature (PVT) variations. Xiumin Xu, Huaguo Liang, Gaoliang Ma, Zhengfeng Huang, Maoxiang Yi, Tianming Ni, Yingchun Lu |
ATS | 7 |