Zhiteng Chao

dblp:282/9344 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
15since 2021 · last 2026
0009-0006-2926-7499ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 9 first-author · 15 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 AssertMiner: Module-Level Spec Generation and Assertion Mining using Static Analysis Guided LLMs
abstract
Assertion-based verification (ABV) is a key approach to checking whether a logic design complies with its architectural specifications. Existing assertion generation methods based on design specifications typically produce only top-level assertions, overlooking verification needs on the implementation details in the modules at the micro-architectural level, where design errors occur more frequently. To address this limitation, we present AssertMiner, a module-level assertion generation framework that leverages static information generated from abstract syntax tree (AST) to assist LLMs in mining assertions. Specifically, it performs AST-based structural extraction to derive the module call graph, I/O table, and dataflow graph, guiding the LLM to generate module-level specifications and mine module-level assertions. Our evaluation demonstrates that AssertMiner outperforms existing methods such as AssertLLM and Spec2Assertion in generating high-quality assertions for modules. When integrated with these methods, AssertMiner can enhance the structural coverage and significantly improve the error detection capability, enabling a more comprehensive and efficient verification process.
Hongqin Lyu, Yonghao Wang, Zhiteng Chao, Huawei Li 0001
ASP-DAC4
2026 Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToT
abstract
Large language models (LLMs) hold promise for automating integrated circuit (IC) engineering using register transfer level (RTL) hardware description languages (HDLs) like Verilog. However, challenges remain in ensuring the quality of Verilog generation. Complex designs often fail in a single generation due to the lack of targeted decoupling strategies, and evaluating the correctness of decoupled sub-tasks remains difficult. While the chain-of-thought (CoT) method is commonly used to improve LLM reasoning, it has been largely ineffective in automating IC design workflows, requiring manual intervention. The key issue is controlling CoT reasoning direction and step granularity, which do not align with expert RTL design knowledge. This paper introduces VeriBToT, a specialized LLM reasoning paradigm for automated Verilog generation. By integrating Top-down and design-for-verification (DFV) approaches, VeriBToT achieves self-decoupling and self-verification of intermediate steps, constructing a Backtrack Tree of Thought with formal operators. Compared to traditional CoT paradigms, our approach enhances Verilog generation while optimizing token costs through flexible modularity, hierarchy, and reusability.
Zhiteng Chao, Yonghao Wang, Tenghui Hua, Husheng Han, Tianmeng Yang, Jianan Mu, Bei Yu 0001, Rui Zhang 0040, Jing Ye 0001, Huawei Li 0001
DATE1
2026 CoverAssert: Iterative LLM Assertion Generation Driven by Functional Coverage via Syntax-Semantic Representations
abstract
LLMs can generate SystemVerilog assertions (SVAs) from natural language specs, but single-pass outputs often lack functional coverage due to limited IC design understanding. We propose CoverAssert, an iterative framework that clusters semantic and AST-based structural features of assertions, maps them to specifications, and uses functional coverage feedback to guide LLMs in prioritizing uncovered points. Experiments on four open-source designs show that integrating CoverAssert with AssertLLM and Spec2Assertion improves average improvements of 9.57% in branch coverage, 9.64% in statement coverage, and 15.69% in toggle coverage.
Yonghao Wang, Yang Yin, Hongqin Lyu, Zhiteng Chao, Mingyu Shi, Wenchao Ding 0007, Yunlin Du, Jing Ye 0001, Huawei Li 0001
DATE5
2026 Iterative LLM-Based Assertion Generation Using Syntax-Semantic Representations for Functional Coverage-Guided Verification
Yonghao Wang, Yang Yin, Hongqin Lyu, Zhiteng Chao, Wenchao Ding 0007, Jing Ye 0001, Huawei Li 0001
ETS5
2026 AssertMiner-pro: Enhanced module-level spec generation and assertion mining with LLM guided by top-down hierarchical strategies
Yonghao Wang, Hongqin Lyu, Boling Chen, Mingyu Shi, Zhiteng Chao, Huawei Li 0001
Integr.6
2025 AssertGen: Enhancement of LLM-aided Assertion Generation through Cross-Layer Signal Bridging
abstract
Assertion-based verification (ABV) serves as a crucial technique for ensuring that register-transfer level (RTL) designs adhere to their specifications. While Large Language Model (LLM) aided assertion generation approaches have recently achieved remarkable progress, existing methods are still unable to effectively identify the relationship from the behavioral interactions of signals across different layers, which leads to the insufficiency of the generated assertions. To address this issue, we propose AssertGen, an assertion generation framework that automatically generates SystemVerilog assertions (SVA). AssertGen first extracts verification objectives from specifications using a chain-of-thought (CoT) reasoning strategy, then bridges corresponding signals between these objectives and the RTL code to construct a cross-layer signal chain, and finally generates SVAs based on the LLM. Experimental results demonstrate that AssertGen outperforms the existing state-of-the-art methods across several key metrics, such as pass rate of formal property verification (FPV), cone of influence (COI), proof core and mutation testing coverage.
Hongqin Lyu, Yonghao Wang, Yunlin Du, Mingyu Shi, Zhiteng Chao, Wenxing Li, Huawei Li 0001
ATS5
2025 PastATPG: A Hybrid ATPG Framework for Better Test Compaction with Partial Assignment SAT
abstract
In automatic test pattern generation (ATPG), SAT-based methods are typically used to complement structural approaches, especially for addressing hard-to-detect faults. However, as the size and complexity of circuits grow, SAT-based ATPG faces challenges like pattern inflation and excessive runtime, limiting its overall performance. The key problem lies in the fact that current mainstream SAT solvers perform complete assignments for all primary inputs of the fault’s transitive fanin cone without considering the detection of other faults, making test compaction extremely difficult and time consuming. In this paper, a novel SAT solver PA-MiniSat is proposed, which is capable of generating partial assignments for solving variables and significantly reduces the number of specified bits in test cubes. As an extension of MiniSat, it employs a full-literal watching technique and a circuit-adapted heuristic branching strategy, achieving overall improved performance in ATPG. Based on PA-MiniSat, a hybrid ATPG framework PastATPG is proposed for better test compaction, which tightly integrates structural algorithms with the SAT solver into the unified test compaction flow. Experimental results demonstrate that our method outperforms other SAT solvers in pattern compaction and, in some cases, even surpasses commercial ATPG tools in terms of speed. The code is available at https://github.com/sklp-eda-lab/PastATPG.
Zhiteng Chao, Xindi Zhang 0001, Jianan Mu, Zizhen Liu, Shengwen Liang, Shaowei Cai 0001, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001
DAC1
2025 MOSS: Multi-Modal Representation Learning on Sequential Circuits
abstract
Deep learning has significantly advanced Electronic Design Automation (EDA), with circuit representation learning emerging as a key area for modeling the relationship between a circuit’s structure and functionality. Existing methods primarily use either Large Language Models (LLMs) for Register Transfer Level (RTL) code analysis or Graph Neural Networks (GNNs) for netlist modeling. While LLMs excel at high-level functional understanding, they struggle with detailed netlist behavior. GNNs, however, face challenges when scaling to larger sequential circuits due to long-range information dependencies and insufficient functional supervision, leading to decreased accuracy and limited generalization. To address these challenges, we propose MOSS, a multimodal framework that integrates GNNs with LLMs for sequential circuit modeling. By enhancing D-type Flip-Flop (DFF) node features with embeddings from fine-tuned LLMs on RTL code, we focus the GNN on critical anchor points, reducing reliance on long-range dependencies. The LLM also provides global circuit embeddings, offering efficient supervision for functionality-related tasks. Additionally, MOSS introduces an adaptive aggregation method and a two-phase propagation mechanism in the GNN to better model signal propagation and sequential feedback within the circuit. Experimental results demonstrate that MOSS significantly improves the accuracy of functionality and performance predictions for sequential circuits compared to existing methods, particularly in larger circuits where previous models struggle. Specifically, MOSS achieves a $\mathbf{9 5. 2 \%}$ accuracy in arrival time prediction.
Jianan Mu, Tianmeng Yang, Silin Liu, Yihan Wen, Hui Wang 0152, Zhiteng Chao, Husheng Han, Zizhen Liu, Shengwen Liang, Jing Ye 0001, Bei Yu 0001, Xiaowei Li 0001, Huawei Li 0001
DAC12
2025 TESLA: Testability Enhancement for Shift-Left Automation via Multi-LLM Collaboration
abstract
The "Shift-Left" Design-for-Test (DFT) paradigm has gained significant attention in recent years, enabling early-stage testability enhancement at the Register Transfer Level (RTL) to optimize Power-Performance-Area-Testability (PPAT) trade-offs and accelerate Time-to-Market (TTM). However, existing methods struggle to perform quantitative testability analysis at the RTL stage, particularly in Partial Scan Selection (PSS) and Test Point Insertion (TPI), due to the lack of structured netlist representations and cross-stage optimization. To address this challenge, we propose TESLA, a multi-LLM collaboration framework that autonomously performs PSS and TPI at the RTL stage. TESLA leverages the semantic understanding capabilities of Large Language Models (LLMs) to analyze RTL Verilog code and optimize testability without requiring synthesis. Two key data augmentation strategies are introduced for efficient Instruction Tuning: (1) back-annotating heuristic PSS results from the synthesized netlist to RTL, and (2) utilizing advanced LLMs guided by DFT knowledge to generate synthetic RTL TPI training data. Furthermore, we integrate Direct Preference Optimization (DPO) to refine LLM decision-making, incorporating real feedback from commercial EDA tools to align optimization objectives with practical testability metrics. The experimental results demonstrate that our proposed approach achieves better test coverage compared to other RTL stage PSS and TPI combination schemes on the majority of circuits in the RTLLM benchmark, while also reducing the number of patterns for a significant portion of the circuits. On the larger, hierarchical OpenCores benchmark, our approach surpasses the solution combining heuristic PSS and commercial DFT tool’s TPI, achieving improvements on the same two test metrics.
Zhiteng Chao, Rengang Zhang, Hongqin Lyu, Wenxing Li, Zizhen Liu, Jianan Mu, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001
ITC1
2025 HighTPI: A Hierarchical Graph Based Intelligent Method for Test Point Insertion
abstract
As integrated circuits grow in complexity, test point insertion (TPI) has become vital for enhancing testability and improving reliability in design for test (DFT). Recent studies have shown the effectiveness of deep learning-based TPI using graph neural networks (GNNs) in improving test quality. However, the high cost of collecting training data, incomplete capture of the intrinsic characteristics of circuits, and the vast search space in large circuits hinder the performance of existing intelligent approaches. This paper introduces HighTPI, a two-stage learning approach for TPI to effectively reduce the number of test patterns, which leverages hierarchical graph representation by constructing a hypergraph based on hypernodes in fanout-free regions (FFRs). HighTPI better captures multi-fanout reconvergence information while lowering the cost of obtaining ground-truth labels due to the smaller scale of the FFR-based hypergraph. Two specialized GNNs are designed in stage I to select candidate insertion points for observation and control points, respectively. This integration of expert knowledge through supervised learning helps guide the reinforcement learning process in stage II, mitigating the challenges of sparse rewards and a large decision space. The experimental results demonstrate that HighTPI outperforms other TPI methods in terms of the trade-off between pattern reduction and fault coverage enhancement.
Zhiteng Chao, Hongqin Lyu, Minjun Wang, Wenxing Li, Zizhen Liu, Jianan Mu, Shengwen Liang, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001
VTS1
2025 A fast test compaction method using dedicated Pure MaxSAT solver embedded in DFT flow
Zhiteng Chao, Xindi Zhang 0001, Junying Huang, Zizhen Liu, Jing Ye 0001, Shaowei Cai 0001, Huawei Li 0001, Xiaowei Li 0001
Integr.1
2025 Memory-Efficient and Adaptive Heterogeneous Framework for Gate-Level Fault Simulation
abstract
Gate-level fault simulation is essential for automatic test pattern generation (ATPG). The traditional event-driven simulation is time-consuming due to the large number of faults. While parallel fault simulation with GPGPUs shows promise, it faces reduced parallel efficiency on large circuits. This is mainly due to the increased space required to store fault values, limiting the number of faults that can be processed in parallel and preventing full utilization of the GPU’s capabilities. In this study, we propose a memory-efficient fault machine implementation FM gpu based on a circular vector, which is tailored for GPU fault simulation with some sacrifices of time efficiency and a variable length limit. We also propose a fully adaptive parallel fault simulation framework based on the CPU-GPU heterogeneous system, which includes two stages on the GPU and performs CPU simulation at the same time. All parameters related to GPU memory optimization and workload balancing in the framework can be adjusted adaptively. The experimental results demonstrate that our method achieves better memory efficiency and speedup compared to the previous GPU fault simulation methods, a maximum speedup of 137.48× compared to the baseline open-source simulator with 32 threads, and a maximum speedup of 2.52× compared to a 32-thread commercial tool.
Zhiteng Chao, Junying Huang, Wenjie Li 0004, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001
ACM Trans. Design Autom. Electr. Syst.1
2024 A Fast Test Compaction Method for Commercial DFT Flow Using Dedicated Pure-MaxSAT Solver
abstract
Minimizing the testing cost is crucial in the context of the design for test (DFT) flow. In our observation, the test patterns generated by commercial ATPG tools in test compression mode still contain redundancy. To tackle this obstacle, we propose a post-flow static test compaction method that utilizes a partial fault dictionary instead of a full fault dictionary, and leverages a dedicated Pure-MaxSAT solver to re-compact the test patterns generated by commercial ATPG tools. We also observe that commercial ATPG tools offer a more comprehensive selection of candidate patterns for compaction in the “n-detect” mode, leading to superior compaction efficacy. In experiments on ISCAS89, ITC99, and open-source RISC-V CPU benchmarks, our method achieves an average reduction of 21.58% and a maximum of 29.93% in test cycles evaluated by commercial tools while maintaining fault coverage. Furthermore, our approach demonstrates improved performance compared with existing methods.
Zhiteng Chao, Xindi Zhang 0001, Junying Huang, Jing Ye 0001, Shaowei Cai 0001, Huawei Li 0001, Xiaowei Li 0001
ASPDAC1
2024 A Static Test Compaction Method Based on GCN Assisted Fault Gate Classification
abstract
Static test compaction aims to reduce the number of generated test patterns after automatic test pattern generation (ATPG) to enable one pattern to detect more faults. However, existing traditional algorithms require the establishment and maintenance of a fault detection profile to obtain essential faults, identified as those detectable exclusively by a single pattern (denoted as 1-D), incurring substantial computational overhead. We propose a novel fault gates classification approach based on graph convolutional network (GCN). By categorizing fault gates into with and without hard-to-detect faults, we selectively construct a partial fault detection profile only for the fault gates with hard-to-detect faults. Partial fault detection profile effectively reduces the time spent on establishing and maintaining it, as well as the time cost of obtaining essential faults. The experiment shows that our improved method can increase pattern reduction efficiency while accelerating, and its impact on fault coverage can be ignored. Compared with the original algorithm, the maximum acceleration ratio is 6.44×, and the number of patterns is reduced by up to 20.68%.
Zhiteng Chao, Qinluan Dai, Zizhen Liu, Wenxing Li, Hongqin Lyu, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001
ITC-Asia1
2023 A Distributed ATPG System Combining Test Compaction Based on Pure MaxSAT
abstract
As the target of test synthesis is to obtain highly compacted test patterns with acceptable fault coverage, automatic test pattern generation (ATPG) plays an important role in the design for test (DFT) process. Distributed ATPG systems have been designed to harness the parallelism of computer architectures to accelerate this process. However, due to delayed communication among distributed nodes, the redundancy of certain computations and substantial pattern expansion issues may arise. To tackle this problem, this paper proposes a test compaction module based on Pure MaxSAT to re-compact the patterns generated by the distributed ATPG system, significantly reducing the number of patterns without loss of fault coverage. Techniques such as partial fault-dropping, two-stage compaction, and building a fault dictionary with machine word all-fill are integrated to reduce the cost of test compaction and internal communication overhead within the distributed framework. Experimental results indicate that the number of patterns generated by the distributed ATPG system integrated with the test compaction module is greatly reduced within an acceptable time overhead.
Zhiteng Chao, Senlin Wang, Pengyu Tian, Shuwen Yuan, Huawei Li 0001, Jing Ye 0001, Xiaowei Li 0001
ATS1
2020 Optimization Space Exploration of Hardware Design for CRYSTALS-KYBER
abstract
Public key cryptography is important in the global communication digital infrastructure. However, the emergence of quantum computer and Shor algorithm has greatly threatened the security of public key cryptography. The CRYSTALS-KYBER, as a lattice-based KEM algorithm, passed three rounds of a global solicitation for post-quantum cryptography algorithms held by the National Institute of Standards and Technology (NIST). This paper explores the implementation and optimization space of hardware design according to CRYSTALS-KYBER algorithm. We analyze its software code and try different strategies to optimize the hardware implementation, and conduct comparative analysis in terms of area and speed. The experimental results show that the performance can be greatly improved by moderately optimizing the loops. In comparison with optimal results of the work [12], our optimizations improve the performance by up to 74.6% for encapsulation algorithm and 54.4% for decapsulation algorithm.
Zhiteng Chao, Jing Ye 0001, Wen Wang 0007, Yuan Cao 0003, Xiaowei Li 0001, Huawei Li 0001
ATS2