Wei-Kai Liu

dblp:32/1743 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0003-4075-3499ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 7 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Runtime Fault Localization in Deep Neural Network Accelerators
abstract
Systolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Although fault detection and repair techniques have been proposed to enhance the robustness of systolic arrays, fault localization remains an open problem. We propose a fault tolerance framework including run-time based fault detection and fault localization, both leveraging functional data to generate checksums on-the-fly. This approach enables error detection and localization during normal operation, avoiding the need for dedicated test patterns or additional downtime. Experimental evaluation shows that the proposed fault localization architecture incurs an area overhead less than 2% for a 256× 256 systolic array. In simulations, the proposed method achieves 100% fault detection and localization in a 256× 256 systolic array.
Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty
ACM Trans. Design Autom. Electr. Syst.1
2025 Identifying System-on-Chip Security Assets with Structure-Based Analysis
abstract
In the evolving field of hardware design, ensuring the security of System-on-Chips (SoCs) has become increasingly vital. As SoCs grow in complexity, integrating components from various sources, the identification and protection of security assets are crucial to prevent vulnerabilities. Traditional methods of identifying these assets are manual and time-intensive. To address this challenge, automated tools for security asset identification are essential, enabling faster and more accurate detection of critical assets early in the design process. In this paper, we propose a framework for the automated identification of security assets within SoCs. By transforming register-transfer level (RTL) code into graphs and leveraging deep neural networks (DNNs) to classify assets based on their structural patterns, our approach can effectively differentiate between security and non-security assets. Experimental results show that the proposed method achieves high classification accuracy, with the model reaching up to 99% accuracy in identifying security assets, significantly reducing the need for manual intervention.
Wei-Kai Liu, Benjamin Tan 0001, Krishnendu Chakrabarty
DAC1
2025 Patchability-Driven Design Exploration for System-on-Chip Patching Architectures
abstract
As System-on-Chip (SoC) designs become increasingly complex, ensuring comprehensive verification has become more challenging, leading to overlooked hardware bugs that can be found in the field. Addressing hardware bugs post-deployment is difficult, as they typically cannot be easily fixed like software bugs. To tackle this issue, hardware-based patching mechanisms have emerged as a potential solution for providing in-field fixes. However, the lack of a standardized method to evaluate the ”patchability” of different designs complicates the integration of patching infrastructure into SoCs. In this article, we propose a fully parameterized Patch Support Block (PSB) architecture that can be tailored for various hardware designs, enabling post-deployment patching. We introduce a novel patchability score formulation that provides a quantifiable metric for evaluating the effectiveness of patching designs. Our approach considers both the observability and controllability of the patching hardware and provides a framework for system integrators to maximize patchability while managing resource constraints. Through experimentation with multiple design configurations, we demonstrate how our methodology can enhance patchability in hardware systems and provide security-related fixes for SoCs in real-world scenarios.
Wei-Kai Liu, Benjamin Tan 0001, Krishnendu Chakrabarty
ACM Trans. Design Autom. Electr. Syst.1
2024 Theoretical Patchability Quantification for IP-Level Hardware Patching Designs
abstract
As the complexity of System-on-Chip (SoC) designs continues to increase, ensuring thorough verification becomes a significant challenge for system integrators. The complexity of verification can result in undetected bugs. Unlike software or firmware bugs, hardware bugs are hard to fix after deployment and they require additional logic, i.e., patching logic integrated with the design in advance in order to patch. However, the absence of a standardized metric for defining “patchability” leaves system integrators relying on their understanding of each IP and security requirements to engineer ad hoc patching designs. In this paper, we propose a theoretical patchability quantification method to analyze designs at the Register Transfer Level (RTL) with provided patching options. Our quantification defines patchability as a combination of observability and controllability so that we can analyze and compare the patchability of IP variations. This quantification is a systematic approach to estimate each patching architecture’s ability to patch at run-time and complements existing patching works. In experiments, we compare several design options of the same patching architecture and discuss their differences in terms of theoretical patchability and how many potential weaknesses can be mitigated.
Wei-Kai Liu, Benjamin Tan 0001, Jason M. Fung, Krishnendu Chakrabarty
ASPDAC1
2024 Effective Runtime Fault Detection for DNN Accelerators
abstract
Systolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing of matrix multiplication, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Even though many algorithm-based fault tolerance (ABFT) algorithms have been proposed to detect and correct errors in matrix multiplication, these ABFT methods cannot detect many errors originating from the accelerator hardware. We propose a run-time based fault detection technique leveraging functional data to generate checksums on-the-fly, avoiding the requirement for test patterns. Experimental evaluation shows that the proposed fault detection architecture can achieve 100% test coverage while incurring an area overhead of less than 2% for a 256 × 256 systolic array.
Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty
ATS1
2023 Hardware-Supported Patching of Security Bugs in Hardware IP Blocks
abstract
To satisfy various design requirements and application needs, designers integrate multiple intellectual property blocks (IPs) to produce a system on chip (SoC). For improved survivability, designers should be able to patch the SoC to mitigate potential security issues arising from hardware IPs; for increased flexibility, we propose adding programmable hardware-based support for monitoring and bug mitigation. However, it is a challenge to decide how much additional cost a designer should expend up front to deal with unknown, future issues. We propose an approach that guides designers toward maximizing the benefits of adding “patchability” to various IPs in the system, given a target resource overhead. We frame the design problem as an integer quadratic program and show that our approach achieves superior patchability compared to the naïve and baseline approaches for a given cost limit. Experimental results show that when we set a cost limit of 2% field-programmable gate array adaptive logic module usage, our solution can generate a viable patching infrastructure with six patching blocks offering patches for seven different services in our case study.
Wei-Kai Liu, Benjamin Tan 0001, Jason M. Fung, Ramesh Karri, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Don't CWEAT It: Toward CWE Analysis Techniques in Early Stages of Hardware Design
abstract
To help prevent hardware security vulnerabilities from propagating to later design stages where fixes are costly, it is crucial to identify security concerns as early as possible, such as in RTL designs. In this work, we investigate the practical implications and feasibility of producing a set of security-specific scanners that operate on Verilog source files. The scanners indicate parts of code that might contain one of a set of MITRE's common weakness enumerations (CWEs). We explore the CWE database to characterize the scope and attributes of the CWEs and identify those that are amenable to static analysis. We prototype scanners and evaluate them on 11 open source designs - 4 system-on-chips (SoC) and 7 processor cores - and explore the nature of identified weaknesses. Our analysis reported 53 potential weaknesses in the OpenPiton SoC used in [email protected], 11 of which we confirmed as security concerns.
Baleegh Ahmad, Wei-Kai Liu, Luca Collini, Hammond A. Pearce, Jason M. Fung, Jonathan Valamehr, Mohammad Bidmeshki, Piotr Sapiecha, Krishnendu Chakrabarty, Ramesh Karri, Benjamin Tan 0001
ICCAD2
2021 Time-Division Multiplexing Based System-Level FPGA Routing
abstract
Multi-FPGA system prototyping has become popular for modern VLSI logic verification, but such a system realization is often limited by its number of inter-FPGA connections. As a result, time-division multiplexing (TDM) is employed to accommodate more inter-FPGA signals than the connections in a multi-FPGA system. However, the inter-FPGA signal delay induced by TDM becomes significant due to time-multiplexing. Researchers have shown that TDM ratios (signal time-multiplexing ratios) significantly affect the performance of a multi-FPGA system and inter-FPGA routing highly influences the quality of this system. This paper presents a framework to minimize the system clock period for a system-level FPGA while considering the inter-FPGA routing topology and the timing criticality of nets. Our framework consists of two stages: (1) a distributed profiling scheme to generate the desired net-ordering and then alleviate the routing congestion, and (2) a net-/edge-based refinement to assign TDM ratios efficiently with a strict decrease in the ratios. Based on the 2019 CAD contest at ICCAD benchmarks and the contest evaluation metric with both quality and efficiency, experimental results show that our framework achieves the best overall score among all the participating teams and published works.
Wei-Kai Liu, Ming-Hung Chen, Chen-Chia Chang, Yao-Wen Chang
ICCAD1
1997 A Parallel Algorithm for Constructing a Labeled Tree
abstract
A tree T is labeled when the n vertices are distinguished from one another by names such as v/sub 1/, v/sub 2/...v/sub n/. Two labeled trees are considered to be distinct if they have different vertex labels even though they might be isomorphic. According to Cayley's tree formula, there are n/sup n-2/ labeled trees on n vertices. Prufer used a simple way to prove this formula and demonstrated that there exists a mapping between a labeled tree and a number sequence. From his proof, we can find a naive sequential algorithm which transfers a labeled tree to a number sequence and vice versa. However, it is hard to parallelize. In this paper, we shall propose an O(log n) time parallel algorithm for constructing a labeled tree by using O(n) processors and O(n log n) space on the EREW PRAM computational model.
Yue-Li Wang, Hon-Chan Chen, Wei-Kai Liu
IEEE Trans. Parallel Distributed Syst.3