Liuzheng Wang

dblp:124/2724 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021
YearPublicationVenuePosition
2026 C2C: Cell-to-Cell Controllability Evaluation for Partial Scan Selection
Hairui Cai, Liuzheng Wang, Lingxiang Liao, Yu Huang 0005, Zhouxing Su, Zhipeng Lv
VTS3
2023 Fault Simulation Acceleration Based on ARM Multi-core CPU Architecture
abstract
Fault simulation plays an important role in ATPG and fault diagnosis of integrated circuits. However, with the increasing complexity of chips, the simulation of tens of millions faults of VLSI needs a lot of time and computing resources. To improve the simulation efficiency, many methods and technologies have emerged, such as GPU acceleration and distributed computing. However, there are some challenges and limitations to these approaches, such as high cost, high energy consumption, and programming complexity. In contrast, the acceleration of fault simulation based on ARM multi-core CPU has the characteristics of strong multi-threaded parallel computing capability, low cost, low energy consumption, and easy implementation. Therefore, a new method for accelerating fault simulation based on ARM multi-core CPU is proposed in this paper, which adopts an enhanced parallel simulation method. In this method, each node has its own memory space and processor, and the CPU can work at full load. This paper also provides solutions to NUMA affinity, cross-node access and other problems. This paper will elaborate on these methods and demonstrate their effectiveness and accuracy in fault simulation.
Shi-Jie Ye, Yun-Ju Liu, Liuzheng Wang, Hui-Ling Zhen, Weiming Zhang 0005, Yu Huang 0005
ATS3
2023 Adaptive Multidimensional Parallel Fault Simulation Framework on Heterogeneous System
abstract
Fault simulation is a critical component of the automatic test pattern generation (ATPG) tool, which is widely used in chip development. The CPU–GPU heterogeneous system can accelerate fault simulation. However, existing work faces the following challenges: 1) Path Divergence: The simulation path of different faults is not uniform, which leads to low parallel efficiency of different GPU threads; 2) Unbalanced Workload: The load of different computing units is not balanced, leading to serious differences in the execution time of each part; and 3) Poor Scalability: When the circuit scale increases, the GPU memory is limited and the simulation has strong structural dependence, which makes the simulation difficult. In this work, we propose an adaptive multidimensional parallel fault simulation framework based on the CPU–GPU heterogeneous system. We adaptively select different simulation approaches according to different circuit scales. In detail, we use the fanout-free region (FFR) grouping method to solve the problem of path divergence. We also use a combination of static and dynamic load balancing to tradeoff data handling and the execution time of each computing unit. We limit the queue length used in the GPU to improve the scalability of the simulation. To further accelerate, we propose the 4-D parallel architecture on multiple GPUs. Extensive experimental results show that our fault simulator based on 8 GPU is$105.7\times $faster than the commercial tool on average. For tens of millions of gate-level circuits, our fault simulator based on one GPU is up to$25.9\times $faster than the CPU single-threaded simulator.
Jingbo Hu, Guohao Dai 0001, Liuzheng Wang, Liyang Lai, Huazhong Yang, Yu Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 DeepTPI: Test Point Insertion with Deep Reinforcement Learning
abstract
Test point insertion (TPI) is a widely used technique for testability enhancement, especially for logic built-in self-test (LBIST) due to its relatively low fault coverage. In this paper, we propose a novel TPI approach based on deep reinforcement learning (DRL), named DeepTpi. Unlike previous learning-based solutions that formulate the TPI task as a supervised-learning problem, we train a novel DRL agent, instantiated as the combination of a graph neural network (GNN) and a Deep Q-Learning network (DQN), to maximize the test coverage improvement. Specifically, we model circuits as directed graphs and design a graph-based value network to estimate the action values for inserting different test points. The policy of the DRL agent is defined as selecting the action with the maximum value. Moreover, we apply the general node embeddings from a pretrained model to enhance node features, and propose a dedicated testability-aware attention mechanism for the value network. Experimental results on circuits with various scales show that DeepTPI significantly improves test coverage compared to the commercial DFT tool. The code of this work is available at https://github.com/cure-lab/DeepTPI.
Zhengyuan Shi, Min Li 0019, Sadaf Khan, Liuzheng Wang, Naixing Wang, Yu Huang 0005, Qiang Xu 0001
ITC4
2010 Achieving page-mapping FTL performance at block-mapping FTL cost by hiding address translation
abstract
Flash Translation Layer (FTL) is one of the most important components of SSD, whose main purpose is to perform logical to physical address translation in a way that is suitable to the unique physical characteristics of the Flash memory technology. The pure page-mapping FTL scheme, arguably the best FTL scheme due to its ability to map any logical page number (LPN) to any physical page number (PPN) to minimize erase operations, cannot be practically deployed since it consumes a prohibitively large RAM (SRAM or DRAM) space to store the page-mapping table for an SSD of moderate to large size. Alternatives to the pure page-mapping FTL, such as block-mapping FTLs, hybrid FTLs (e.g., FAST) and the latest demand-based page-mapping FTLs (e.g., DFTL), require significantly less RAM space but suffer from a few performance issues. Block-mapping FTLs perform poorly with higher erasure counts, particularly under random write workloads. Hybrid FTL schemes incur costly merge operations that hurt performance and increase the erasure counts. Performances of demand-based FTLs heavily depend on workload characteristics such as access locality, read/write ratio and request arrival interval time. This paper proposes a new FTL scheme, called HAT, to achieve the performance of a pure page-mapping FTL at the RAM cost of a block-mapping FTL while consuming lower energy, by hiding the address translation (HAT). The basic idea behind our scheme is to create a separate access path to read/write the address mapping information to significantly Hide the Address-Translation latency by incorporating a low energy-consuming solid-state memory device that stores the entire page mapping table. We implement an SSD simulator, SSDsim, to validate our HAT design and evaluate its performance. The extensive trace-driven simulation results show that the performance of HAT is within 0.8% of the pure page-mapping FTL, while consuming about 50% of the energy.
Yang Hu 0007, Hong Jiang 0001, Dan Feng 0001, Lei Tian 0001, Shu Ping Zhang, Jingning Liu, Wei Tong 0001, Liuzheng Wang
MSST9