Jianwang Zhai

dblp:291/6350 · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
34since 2021 · last 2026
0000-0002-1581-3536ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 4 first-author · 33 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Timing-Aware Optimization of Die-Level Routing and TDM Assignment for Multi-FPGA Systems
abstract
The escalating scale and complexity of modern circuits demand multi-FPGA emulation platforms that incorporate multi-die architectures. However, most existing routers remain FPGA-level, optimizing wire-length or total Time-Division Multiplexing (TDM) ratios while disregarding die-level load imbalance and path-level slack. They result in suboptimal performance and timing violations. In this paper, we propose a timing-aware co-optimization framework for die-level routing and TDM assignment, explicitly linking physical constraints to critical path timing slack. The proposed flow features a timing-aware load-balanced die-level router with timing path compression and a timing graph-based TDM assignment. Experiments on industrial designs show that the proposed method improves the worst-path slack by 98% over the existing methods.
Haoyuan Li 0004, Chunyan Pei, Jianwang Zhai, Wenjian Yu
ASP-DAC4
2026 HLS-Timer: Fine-Grained Path-Level Timing Estimation for High-Level Synthesis
abstract
Electronic Design Automation (EDA) requires early stage timing guidance to maximize optimization potential. Accurate timing estimation is essential in the High-Level Synthesis (HLS) stage. However, current HLS tools often produce inaccurate timing predictions, resulting in unmet performance targets. While post-synthesis EDA tool chains can provide precise timing analysis, their exhaustive methodologies are prohibitively time-consuming and computationally expensive. Recent machine learning approaches have demonstrated promising results in predicting design-level timing metrics in HLS designs, such as Worst Negative Slack (WNS) and Critical Path (CP) delay. Nevertheless, fine-grained, path-level timing estimation remains an unresolved challenge. In this work, we present HLS-Timer, the first path-level timing estimator for HLS. The proposed framework employs a graph-based representation of the HLS design, integrating local structural features with global contextual information to model timing paths and provide accurate, finegrained delay predictions. Experimental results demonstrate that on previously unseen designs HLS-Timer achieves exceptional accuracy in path-level delay estimation (Pearson $\mathbf{R}=\mathbf{0. 9 4}, \mathbf{R}^{\mathbf{2}} =0.93$, MAPE $=18.96 \%$), highlighting its strong generalization capability. Furthermore, it surpasses state-of-the-art baselines in design-level timing prediction, reducing MAPE to $9.97 \%$ for WNS and $6.79 \%$ for CP delays.
Zibo Hu, Zhe Lin 0001, Renjing Hou, Xingyu Qin, Jianwang Zhai
ASP-DAC5
2026 MAEDA: An LLM-Powered Multi-Agent Evaluation Framework for EDA Tool Documentation QA
abstract
Large Language Models (LLMs) have shown remarkable capability in knowledge-intensive scenarios, such as electronic design automation (EDA) tool documentation question answering (QA), due to their ability to process and generate contextually rich, domain-specific information. Evaluating LLM outputs is paramount, as it directly impacts their accuracy, effectiveness, and trustworthiness in practical applications. In this paper, we introduce MAEDA, a novel LLM-powered multi-agent evaluation framework that utilizes multiple fine-tuned LLM agents working collaboratively to assess common error types encountered in EDA tool documentation QA. Specifically, we design customized point-to-point alignment and chain-of-thought (CoT) reasoning strategies tailored to specific agents, enhancing both fine-tuning and inference capabilities. Experimental results demonstrate that MAEDA outperforms state-of-the-art (SOTA) general-purpose and cross-domain evaluation frameworks in accurately identifying error types specific to this domain. Our benchmark is publicly available at https://github.com/Rayzzz14/MAEDA-DATE26/.
Yuan Pu 0001, Hairuo Han, Yuntao Nie, Jiajun Qin, Yuhan Qin, Tairu Qiu, Zhuolun He, Jianwang Zhai, Bei Yu 0001
DATE9
2026 DynaOpt: A Heterogeneous Logic Optimization Framework with Dynamic Sequence Generation
abstract
Heterogeneous logic optimization improves circuit quality by partitioning a design and leveraging the best Directed Acyclic Graph (DAG) representation for each region. However, existing frameworks are limited by their reliance on applying fixed, pre-defined optimization scripts to these partitions. This approach fails to adapt to the specific structure of each partition or its impact on global circuit metrics. This paper introduces DynaOpt, a framework that overcomes this limitation by dynamically generating tailored optimization sequences. After partitioning the circuit with a timing and structure-aware algorithm and selecting the optimal DAG for each partition, DynaOpt discovers a bespoke optimization sequence for each sub-circuit. The key to this process is our novel, globally-aware fitness function, which guides a Genetic Algorithm (GA) by efficiently approximating the impact of local changes on the final circuit quality. Experiments demonstrate that DynaOpt achieves a significant improvement in Quality of Results (QoR) over the state-of-the-art (SOTA) framework. This validates the effectiveness of generating custom optimization sequences and addresses the fundamental limitations of relying on pre-defined sequences.
Xingyu Qin, Guande Dong, Jianwang Zhai
DATE3
2026 AutoShrink: Adaptive Search Space Shrinkage for Large-Scale Pareto Optimization of HLS Designs
abstract
High-level synthesis (HLS) streamlines accelerator customization by delivering a high-level hardware programming paradigm enriched with a variety of optimization directives. However, the quality of HLS designs is largely determined by the selection of directives in navigating trade-offs among multiple design metrics, a non-trivial process that can significantly prolong design turnaround time. Design space exploration (DSE) serves as a promising solution to this problem, but existing studies on DSE suffer from a lack of efficiency or generalization capability in large-scale application scenarios. To address this problem, this paper proposes AutoShrink, a DSE engine that automatically and adaptively shrinks the large search space of an HLS design to gradually retain only high-quality solutions. AutoShrink incorporates: (1) a comprehensive design space pruning strategy that integrates domain knowledge and consolidates the joint effect of directives; and (2) an importance-guided Pareto optimization algorithm that dynamically tracks the importance ranking of the applied directives and leverages this ranking to effectively steer the search toward Pareto-optimal solutions. Experimental results demonstrate that AutoShrink efficiently achieves a close approximation of the Pareto frontier across diverse benchmarks with design spaces scaling up to 1016, which attains an average deviation of only 8.1%, outperforming three generic optimization methods and three state-of-the-art customized approaches by 5.73× and 4.47×, respectively.
Yingxin Zeng, Binghao Cheng, Jianwang Zhai, Zhe Lin 0001
DATE3
2026 Structural Timing-Aware Circuit Partitioning with Feasibility Constraints for Multi-Chiplet Design
abstract
Chiplet-based heterogeneous integration has become a scalable paradigm for modern VLSI systems, where early-stage chiplet partitioning critically affects timing quality after physical design. However, conventional hypergraph partitioning primarily optimizes connectivity, while accurate timing analysis is unavailable or too expensive to obtain at early stages. To address this issue, we present a timing-aware chiplet partitioning framework that integrates structural timing modeling into hypergraph construction while preserving logic module binding and physical feasibility constraints. By embedding timing sensitivity into hyperedge weights, the proposed method guides partitioning to better preserve timing-critical connections during early design stages. A constraint-preserving hypergraph construction strategy is further introduced to maintain implementation consistency across the physical design flow. Experimental results show that the proposed framework significantly improves post-placement timing metrics, and the benefit becomes more pronounced as the chiplet count increases.
Kanglin Tian, Jianwang Zhai, Xiuli Fu
ACM Great Lakes Symposium on VLSI4
2026 Etch-Explorer: A Robust Bayesian Optimization Framework for Stringent Constrained Plasma Etching
abstract
Plasma etching is a critical process in semiconductor manufacturing, yet discovering optimal recipes is hindered by the constraint collapse challenge, in which the “golden” intervals that simultaneously satisfy multiple stringent interval constraints are extremely sparse in the high-dimensional design space. Traditional optimization methods often struggle with this sparsity and the complex physical coupling of plasma reactions. To address these limitations, we propose Etch-Explorer, a robust Bayesian Optimization (BO)-based framework that features three synergistic innovations. The framework first employs a Heterogeneous Active Sampling (HAS) strategy to capture space skeletons and physical boundaries, effectively mitigating initial search blindness in sparse regions. Subsequently, a Joint-Constraint-Aware Acquisition Function (JCAF) is leveraged to provide risk-aware navigation by explicitly modeling the joint satisfaction probabilities across multiple interval objectives. To support precise decision-making, a Deep Residual Process Emulator (ResSAN-DTS) is integrated to capture deep non-linear physical couplings with high fidelity. Experimental results demonstrate that Etch-Explorer significantly outperforms state-of-the-art methods in search efficiency and success rate, successfully locating optimal recipes within stringent constraints while substantially reducing wafer costs.
Jianwang Zhai
ACM Great Lakes Symposium on VLSI4
2026 Performance Pragma-Based Design Space Pruning and Exploration for High-Level Synthesis
Donghao Guo, Zhe Lin 0001, Jianwang Zhai
ISCAS3
2026 CHASE: A CHiplet Architecture Simulation and Exploration Framework with Decoupled Multi-Fidelity Optimization
abstract
Chiplet-based architecture is a promising emerging technology with benefits in cost, reusability, and performance. However, designing a complicated system to fulfill the comprehensive design metrics is challenging, and designers frequently suffer from tedious evaluation iterations. We propose the CHASE framework, i.e., a CHiplet-based Architecture Simulation and Exploration framework, which jointly considers both performance metrics and manufacturing. In the framework, simulation component ChipletSIM offers holistic modeling of chiplet-based architectures, integrating critical performance metrics (e.g., latency, power) and manufacturing metrics (e.g., yield, cost) across design stages. The exploration component, ChipletDSE, adopts a decoupled multi-fidelity exploration strategy to boost design exploration efficiency and reduce resource consumption. Our framework substantially improves the probability of attaining optimal designs in the early design phase via a comprehensive simulation process and an efficient exploration approach. Compared to previous methods, the experimental results demonstrate the effectiveness of the CHASE framework in comprehensive simulation and efficient exploration.
Shixin Chen, Jianwang Zhai, Bei Yu 0001
ISPD3
2025 The Survey of 2.5D Integrated Architecture: An EDA perspective
abstract
Enhancing performance while reducing costs is the fundamental design philosophy of integrated circuits (ICs). With advancements in packaging technology, interposer-based chiplet architecture has emerged as a promising solution. Chiplet integration, often referred to as 2.5D IC, offers significant benefits, including cost-effectiveness, reusability, and improved performance. However, realizing these advantages heavily relies on effective electronic design automation (EDA) processes. EDA plays a crucial role in optimizing architecture design, partitioning, combination, physical design, reliability analysis, etc. Currently, optimizing the automation methodologies for chiplet architecture is a popular focus; therefore, we propose a survey to summarize current methods and discuss future directions. This paper will review the research literature on design automation methods for chiplet-based architectures, highlighting current challenges and exploring opportunities in 2.5D IC from an EDA perspective. We expect this survey will provide valuable insights for the future development of EDA tools chiplet-based integrated architectures.
Shixin Chen, Zichao Ling, Jianwang Zhai, Bei Yu 0001
ASP-DAC4
2025 PIRLLS: Pretraining with Imitation and RL Finetuning for Logic Synthesis
abstract
As a key step in digital integrated circuit (IC) design, logic synthesis involves various logic optimization algorithms, where the quality of results (QoR) depends heavily on the optimization sequence used. Exploring the optimization space is challenging as the number of potential optimal permutations grows exponentially. Traditional methods rely on manual adjustments by experts, but are difficult to deal with complex and different circuits, leading to significant optimality gaps. Many automatic methods have been introduced, but still face problems of low generalization and low efficiency.
Guande Dong, Jianwang Zhai, Hongtao Cheng, Chuan Shi 0001
ASP-DAC2
2025 FTAFP: A Feedthrough-Aware Floorplanner for Hierarchical Design of Large-Scale SoCs
abstract
Floorplanning is a critical step in the physical design of digital integrated circuits (ICs). As circuit complexity grows, the hierarchical design paradigm of large-scale systems on chips (SoCs) is gradually emerging, introducing new optimization challenges, particularly with feedthrough. Feedthrough is a through-module connection, yet it would require additional buffers and ports inside the module for data transmission. Excessive feedthroughs will inevitably hinder the routability within reusable modules, causing congestion and timing problems. However, few works have addressed the challenges of feedthrough modeling and optimization.
Kanglin Tian, Jianwang Zhai, Shixiong Kai, Bei Yu 0001
ASP-DAC3
2025 IRGNN: A Graph-based Framework Integrating Numerical Solution and Point Cloud for Static IR Drop Prediction
abstract
With the continued scaling of integrated circuits (ICs), IR drop analysis for on-chip power grids (PGs) is crucial but increasingly computationally demanding. Traditional numerical methods deliver high accuracy but are prohibitively time-intensive, while various machine learning (ML) methods have been introduced to alleviate these computational burdens. However, most CNN-based methods ignore the fine structure and topological information of PGs, and face interpretability or scalability issues. In this work, we propose a novel graphbased framework, IRGNN, leveraging the PG topology with the integration of numerical solutions and point clouds. Our framework applies a numerical solver, AMG-PCG, to generate rough numerical solutions as a reliable interpretability foundation for ML. Then, to capture PG topology, we regard nodes of PG as point clouds and extract point cloud features, and we introduce a novel graph structure, IRGraph. Furthermore, a novel graph-based model IRGNN is designed, incorporating a designed neighbor distance attention (NDA) layer for distanceaware PG features aggregation and graph transformer (GT) layer to capture global information. It should be noted that our framework can analyze the IR drop of each node in PG, which CNN-based methods cannot do. Experimental evaluations demonstrate that our framework achieves significantly higher accuracy than previous CNN-based approaches and numerical solvers while substantially reducing computation time.
Yueyue Xi, Jianwang Zhai, Jingyu Jia, Jiawei Liu 0006, Chuan Shi 0001
DAC3
2025 HeteroSVD: Efficient SVD Accelerator on Versal ACAP with Algorithm-Hardware Co-Design
abstract
Singular value decomposition (SVD) is a matrix factorization technique widely used in signal processing and recommendation systems, etc. In general, the time complexity of SVD algorithms is cubic to the problem size, making SVD algorithms difficult to meet stringent performance requirements in real-time. However, existing FPGA and GPU solutions fall short of jointly optimizing latency, throughput, and power consumption. To settle this issue, this paper proposes HeteroSVD, a heterogeneous reconfigurable accelerator for SVD computation on the Versal ACAP platform. HeteroSVD introduces a system-level SVD decomposition mechanism and proposes an algorithm-hardware co-design method to optimize SVD ordering jointly and AI engine (AIE)-centric dataflow and placement with Versal. Furthermore, in order to improve the quality of results (QoR) and facilitate micro-architecture selection, we introduce an automatic optimization framework that performs accurate performance modeling and fast design space exploration. Experiment results demonstrate that HeteroSVD reduces the latency by $1.98 \times$ over existing FPGA accelerators and outperforms GPU solutions with an improvement of up to $7.22 \times$ in latency, $1.77 \times$ in throughput, and $13.18 \times$ in energy efficiency.
Xinya Luan, Zhe Lin 0001, Jianwang Zhai
DAC4
2025 VSpGEMM: Exploiting Versal ACAP for High-Performance SpGEMM Acceleration
abstract
Sparse general matrix-matrix multiplication (SpGEMM) serves as a fundamental operation in real-world applications such as deep learning. Different from general matrix multiplication, matrices in SpGEMM are highly sparse and therefore require a compact representation. This places an additional burden on data preprocessing and exchanging and also causes irregular memory access patterns, which can in turn lead to communication and computation bottlenecks. To break these bottlenecks, we present VSpGEMM, a hardware accelerator for SpGEMM that is tailored and optimized on Versal ACAP. Firstly, a new storage format called BCSX is proposed in VSpGEMM, which offers a unified and block-wise compression strategy to deal with both row-major and columnmajor representation of non-zero data, enabling fixed-pattern memory accesses and effective data preloading. Secondly, a multi-level tiling mechanism is introduced to decompose the holistic SpGEMM into multiple computation granularities that fit into the AI Engines (AIEs) on Versal in a hierarchical manner, enhancing data reuse. Thirdly, a hybrid partitioning scheme is presented to orchestrate both the AIEs and programmable logic (PL) for intermediate product merging, which together resolve the issues of high memory utilization and communication demand. Experimental results demonstrate a $2.65 \times$ speedup over state-of-the-art (SOTA) GEMM design on Versal and an average $33.62 \times$ improvement in energy efficiency compared to cuSPARSE on RTX 4090 GPU, showing the efficacy of VSpGEMM.
Zhe Lin 0001, Xinya Luan, Jianwang Zhai
DAC4
2025 Truly Pre-Routing Timing Prediction via Considering Power Delivery Network
abstract
Fast and accurate pre-routing timing prediction is essential in the chip design flow. However, existing machine learning (ML)assisted pre-routing timing methods often overlook the impact of power delivery networks (PDNs), which contribute to IR drop and routing congestion. This limitation can make these methods less practical for realworld circuit design flows. To address this, we propose two specialized encoders-an IR drop-aware encoder and a routing congestion-aware encoder-that effectively capture PDN effects through multimodal fusion of netlist, layout, and PDN data. To mitigate the challenges of imbalanced multimodal fusion, we further develop a Pareto optimization approach to ensure balanced utilization of all modalities, enhancing timing prediction accuracy. Comprehensive experiments on large-scale open-source designs using TSMC’s 16 nm technology node validate the superiority of our model over state-of-the-art pre-routing timing prediction methods.
Yuyang Ye 0001, Mingwei He, Lizheng Ren, Jianwang Zhai, Tinghuan Chen, Jun Yang 0006, Longxing Shi
DAC4
2025 IR-Fusion: A Fusion Framework for Static IR Drop Analysis Combining Numerical Solution and Machine Learning
abstract
IR Drop analysis for on-chip power grids (PGs) is vital but computationally challenging due to the rapid growth in the integrated circuit (IC) scale. Traditional numerical methods employed by current EDA software are accurate but extremely time-consuming. To achieve rapid analysis of IR drop, various machine learning (ML) methods have been introduced to address the inefficiency of numerical methods. However, the issue of interpretability or scalability has been limiting practical applications. In this work, we propose IR-Fusion, which aims to combine numerical methods with ML to achieve the trade-off and complementarity between accuracy and efficiency in static IR drop analysis. Specifically, the numerical method is used to obtain rough solutions and ML models are utilized to improve accuracy further. In our framework, an efficient numerical solver, AMG-PCG, is applied to get rough numerical solutions. Then, based on the numerical solution, the fusion of hierarchical numerical-structural information representing the multilayer structure of the PG is employed, and an Inception Attention U-Net model is designed to capture details and interaction of features at different scales. To cope with the limitations and diversity of PG designs, an augmented curriculum learning strategy is applied to the training phase. Evaluation of IR-Fusion shows that its accuracy is significantly better than previous ML-based methods while requiring considerably less iteration on solver to achieve the same accuracy compared with numerical methods.
Jianwang Zhai, Jingyu Jia, Jiawei Liu 0006, Bei Yu 0001, Chuan Shi 0001
DATE2
2025 WideGate: Beyond Directed Acyclic Graph Learning in Subcircuit Boundary Prediction
abstract
Subcircuit boundary prediction is an important application of machine learning in logical analysis, effectively supporting tasks such as functional verification and logic optimization. Existing methods often convert circuits into and-inverter graphs and then use directed acyclic graph neural networks to perform this task. However, two key characteristics of subcircuit boundary prediction do not align with the fundamental assumptions of directed acyclic graph (DAG) learning, which limits the model's expressiveness and generalization capabilities. To break these assumptions, we propose WideGate, which includes a receptive field generation module that extends beyond the fanin cone and fanout cone, as well as an adaptive aggregation module that focuses on boundaries. Extensive experiments show that WideGate significantly outperforms existing methods in terms of prediction accuracy and training efficiency for sub circuit boundary prediction. The code is available at https://github.com/BUPT-GAMMA/WideGate.
Jiawei Liu 0006, Zhiyan Liu, Jianwang Zhai, Zhengyuan Shi, Qiang Xu 0001, Bei Yu 0001, Chuan Shi 0001
DATE4
2025 AuxiliarySRAM: Exploring Elastic On-Chip Memory in 2.5D Chiplet Systems Design
Zichao Ling, Yixin Xuan, Jianwang Zhai
ACM Great Lakes Symposium on VLSI5
2025 MILS: Modality Interaction Driven Learning for Logic Synthesis
Jiawei Liu 0006, Jianwang Zhai, Chuan Shi 0001
ACM Great Lakes Symposium on VLSI3
2025 Transferable Parasitic Estimation via Graph Contrastive Learning and Label Rebalancing in AMS Circuits
abstract
Graph representation learning on Analog-Mixed Signal (AMS) circuits is crucial for various downstream tasks, e.g., parasitic estimation. However, the scarcity of design data, the unbalanced distribution of labels, and the inherent diversity of circuit implementations pose significant challenges to learning robust and transferable circuit representations. To address these limitations, we propose CircuitGCL, a novel graph contrastive learning framework that integrates representation scattering and label rebalancing to enhance transferability across heterogeneous circuit graphs. CircuitGCL employs a self-supervised strategy to learn topology-invariant node embeddings through hyperspherical representation scattering, eliminating dependency on large-scale data. Simultaneously, balanced mean squared error (BMSE) and balanced softmax cross-entropy (BSCE) losses are introduced to mitigate label distribution disparities between circuits, enabling robust and transferable parasitic estimation. Evaluated on parasitic capacitance estimation (edge-level task) and ground capacitance classification (node-level task) across TSMC 28nm AMS designs, CircuitGCL outperforms all state-of-the-art (SOTA) methods, with the R2improvement of 33.64% ~ 44.20% for edge regression and F1-score gain of 0.9× ~ 2.1× for node classification. Our code is available at https://github.com/ShenShan123/CircuitGCL.
Shan Shen, Shenglu Hua, Jiawei Liu 0006, Jianwang Zhai, Chuan Shi 0001, Wenjian Yu
ICCAD5
2025 HieRFP: A Hierarchical Recognition and Floorplanning Framework for Reusable Modules
abstract
As the beginning stage of physical design, floor-planning is critical to the quality of chip design. However, with the widespread use of hierarchical modular design and reusable modules, floorplanning for large-scale systems-on-chip (SoC) has become more complex. Reusable modules must be efficiently identified and utilized to address this complexity while considering their symmetry. In this work, we propose HieRFP, a framework for identifying reusable modules and performing floorplanning considering symmetry. Firstly, a hierarchical clustering method is proposed to automatically recognize the symmetry of reusable modules based on netlist and module outlines. Then, the hierarchical representation combined with the corner stitching compliant-based method is used to enable floorplanning that considers symmetry and overcomes the challenge of rectilinear modules. Experimental results demonstrate the superiority of HieRFP, achieving better optimization in terms of wirelength and area compared to existing floorplanners, effectively exploiting the symmetry of reusable modules.
Kanglin Tian, Jianwang Zhai
ISCAS3
2025 DrlGoFPGA: FPGA Global Placement Considering Input-Output Buffer Based on Deep Reinforcement Learning and Gradient Optimization
abstract
The placement of the input-output buffer (IOBUF) can impact the performance and power consumption of the FPGA. The existing global placement (GP) methods lack consideration for IOBUF, resulting in a decrease in placement and routing quality. To address this issue, we propose a GP framework, DrlGoFPGA, which combines IOBUF placement based on deep reinforcement learning (DRL) with other instances placement based on gradient optimization (GO). A policy network structure with multi-action sampling is designed to accelerate the running speed of DRL, and a parallelizable reward function is designed to optimize each IOBUF placement action and avoid sparse reward problems. Then, an IOBUF line-network relationship (ILNR) graph creation method is designed to improve the agent’s ability to explore optimal solutions, and the graph features of ILNR by capturing them through a graph neural network embedded in the convolutional neural network. Finally, an IOBUF placement legalization method is designed to ensure that the IOBUF position meets the FPGA architecture. The experimental results show that compared with the state-of-the-art placement tools based on GO, DrlGoFPGA can improve GP speed by 13.2%-7×, half-perimeter wirelength by 0.2%-2.6%, and wirelength by 0.2%-1.5% and the IOBUF placement model has good generalization.
Jianwang Zhai, Liuyu Xiang, Zixi Huang, Zhaofeng He 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Towards Automated RISC-V Microarchitecture Design with Reinforcement Learning
abstract
Microarchitecture determines the implementation of a microprocessor. Designing a microarchitecture to achieve better performance, power, and area (PPA) trade-off has been increasingly difficult. Previous data-driven methodologies hold inappropriate assumptions and lack more tightly coupling with expert knowledge. This paper proposes a novel reinforcement learning-based (RL) solution that addresses these limitations. With the integration of microarchitecture scaling graph, PPA preference space embedding, and proposed lightweight environment in RL, experiments using commercial electronic design automation (EDA) tools show that our method achieves an average PPA trade-off improvement of 16.03% than previous state-of-the-art approaches with 4.07× higher efficiency. The solution qualities outperform human implementations by at most 2.03× in the PPA trade-off.
Jianwang Zhai, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
AAAI2
2024 Fast Estimation for Electromigration Nucleation Time Based on Random Activation Energy Model
abstract
Electromigration (EM) has attracted significant interest in recent years, because the current density of on-chip power delivery networks (PDNs) is always increasing. However, the EM phenomenon is affected by the randomness of the annealing process during nanofabrication, which requires more reliable statistical models for EM analysis. In this work, we propose a fast estimation method for EM nucleation time based on the random activation energy model. Experiments demonstrate that our method can accurately and efficiently analyze the nucleation time distribution under random processes, and achieve 39.1% improvement in estimation speed compared with the previous work.
Jingyu Jia, Jianwang Zhai
DATE2
2024 PGAU: Static IR Drop Analysis for Power Grid using Attention U-Net Architecture and Label Distribution Smoothing
abstract
As feature sizes shrink, the on-chip power grid (PG) faces serious power integrity issues, and static IR drop analysis becomes critical for PG design and optimization. Many machine learning (ML) based methods have been proposed to address the inefficiencies of traditional numerical methods. However, many previous works have ignored the problems of feature confusion and imbalance IR drop distribution. In this work, we propose novel feature augmentation and selection methods to solve the feature confusion problem and use the label distribution smoothing (LDS) technique to handle unbalanced labels. Importantly, we design a static IR drop analysis model for PG using the Attention U-Net architecture (PGAU). Furthermore, two real-world datasets are used for evaluation. Experiments show that our model outperforms baselines, with a 2.6% improvement in the correlation coefficient (CC) and a 22.2% reduction in the mean absolute error (MAE). Moreover, our model is highly transferable and performs better against never-before-seen designs.
Jiawei Liu 0006, Jianwang Zhai, Jingyu Jia, Chuan Shi 0001
ACM Great Lakes Symposium on VLSI3
2024 PolarGate: Breaking the Functionality Representation Bottleneck of And-Inverter Graph Neural Network
abstract
Understanding the functionality of Boolean networks is crucial for processes such as functional equivalence checking, logic synthesis and malicious logic identification. With the proliferation of deep learning in electronic design automation (EDA), graph neural networks (GNNs) are widely used for embedding the and-inverter graphs (AIGs), a standard form of Boolean networks, into vectorized representation. A key challenge in the use of GNN for Boolean representation is that although GNNs can well encapsulate the structural properties of AIGs, they usually fail to fully capture the functionality of Boolean logic. Moreover, most GNNs designed for AIGs (also called AIGNNs) either rely on a large amount of training data or require complex supervisory tasks, making it difficult to maintain high training efficiency and prediction accuracy. In this work, for the first time, we focus on breaking the bottleneck of AIGNNs by augmenting their capability of functional representation, providing an efficient solution called PolarGate, which naturally aligns the message passing process with the logical functionality of AIGs. Specifically, we map the behavior of the logic gate into an ambipolar state space, customize differentiable logical operators, and design a functionality-aware message passing strategy. Experimental results on two logically related tasks (i.e., signal probability prediction and truth-table distance prediction) show that PolarGate outperforms the state-of-the-art GNN-based methods for Boolean representation, with an improvement of 62.1% (40.6%) in learning capability and 79.5% (85.6%) in efficiency on two tasks. The code is avaliable at https://github.com/BUPT-GAMMA/PolarGate.
Jiawei Liu 0006, Jianwang Zhai, Zhe Lin 0001, Bei Yu 0001, Chuan Shi 0001
ICCAD2
2024 Effective Resource Model and Cost Scheme for Maze Routing in 3D Global Routing
abstract
Routing is an essential step in the design closure of integrated circuits (IC) and has become the runtime bottleneck in the physical design flow of very large-scale integrated (VLSI) circuits. The maze routing approach, which is mostly used in the rip-up and reroute (RRR) stage of global routing, closely affects routing efficiency and quality. In this paper, we proposed an effective resource model and dynamic congestion sensitivity adjusting method for multi-level 3D maze routing, to improve the efficiency and quality of solutions for 3D global routing. The proposed method takes into account the resource distribution in the maze routing planning process, can optimize the composition of the solution space, and can flexibly adjust the congestion sensitivity. The experimental results show that after integrating the proposed model in a high-performance 3D global router, better routing quality can be achieved in shorter runtime.
Jianwang Zhai, Zhongdong Qi
ISCAS2
2024 BOOM-Explorer: RISC-V BOOM Microarchitecture Design Space Exploration
abstract
Microarchitecture parameters tuning is critical in the microprocessor design cycle. It is a non-trivial design space exploration (DSE) problem due to the large solution space, cycle-accurate simulators’ modeling inaccuracy, and high simulation runtime for performance evaluations. Previous methods require massive expert efforts to construct interpretable equations or high computing resource demands to train black-box prediction models. This article follows the black-box methods due to better solution qualities than analytical methods in general. We summarize two learned lessons and propose BOOM-Explorer accordingly. First, embedding microarchitecture domain knowledge in the DSE improves the solution quality. Second, BOOM-Explorer makes the microarchitecture DSE for register-transfer-level designs within the limited time budget feasible. We enhance BOOM-Explorer with the diversity-guidance, further improving the algorithm performance. Experimental results with RISC-V Berkeley-Out-of-Order Machine under 7-nm technology show that our proposed methodology achieves an average of 18.75% higher Pareto hypervolume, 35.47% less average distance to reference set, and 65.38% less overall running time compared to previous approaches.
Qi Sun 0002, Jianwang Zhai, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
ACM Trans. Design Autom. Electr. Syst.3
2023 Microarchitecture Power Modeling via Artificial Neural Network and Transfer Learning
abstract
Accurate and robust power models are highly demanded to explore better CPU designs. However, previous learning-based power models ignore the discrepancies in data distribution among different CPU designs, making it difficult to use data from the historical configuration to aid modeling for new target configuration. In this paper, we investigate the transferability of power models and propose a microarchitecture power modeling method based on transfer learning (TL). A novel TL method for artificial neural network (ANN)-based power models is proposed, where cross-domain mixup generates more auxiliary samples close to the target configuration to fill in the distribution discrepancy and domain-adversarial training extracts domain-invariant features to complete the target model construction. Experiments show that our method greatly improves the model transferability and can effectively utilize the knowledge of the existing CPU configuration to facilitate target power model construction.
Jianwang Zhai, Yici Cai, Bei Yu 0001
ASP-DAC1
2023 McPAT-Calib: A RISC-V BOOM Microarchitecture Power Modeling Framework
abstract
Power efficiency has become a nonneglected issue of modern CPUs. Therefore, accurate and robust power models are highly demanded in academia and industry. However, it is hard for existing power models to balance modeling speed, generality, and accuracy well. This article introduces McPAT-Calib, a microarchitecture power modeling framework, which combines McPAT with machine learning (ML) calibration and active learning (AL) sampling. McPAT-Calib can quickly and accurately estimate the power of different benchmarks executed on different CPU configurations, and provide an effective evaluation tool for the early design stage. First, McPAT-7nm is introduced to support the preliminary analytical power modeling for the 7-nm technology node. Then, a wide range of modeling features are identified, and automatic feature selection and advanced nonlinear regression are used to calibrate the McPAT-7nm modeling results, greatly improving the accuracy. Moreover, a novel AL approach termed power greedy sampling (PowerGS) embedded with domain knowledge is leveraged to reduce the modeling cost effectively. We use up to 15 configurations of the RISC-V Berkeley out-of-order machine (BOOM) along with 80 benchmarks, targeting 7-nm technology, to extensively evaluate McPAT-Calib. Compared with state-of-the-art (SOTA) microarchitecture power models, McPAT-Calib can reduce the mean absolute percentage error (MAPE) under different cross-validation (CV) strategies by 3.64%–6.14% (absolute reduction). Meanwhile, PowerGS is superior to the existing AL approaches, which can significantly reduce the demand for labeled samples to speed up model construction. The effectiveness of the overall modeling and estimation flow with AL sampling has also been verified.
Jianwang Zhai, Binwu Zhu, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Microarchitecture Design Space Exploration via Pareto-Driven Active Learning
abstract
Microarchitecture design is a key stage of processor development involving various core design metrics, e.g., performance, power consumption, etc. However, due to the high complexity and huge design space of microarchitecture, it becomes challenging to get better designs quickly. In this article, we propose a microarchitecture design space exploration (DSE) approach via Pareto-driven active learning (AL). First, a more accurate dynamic tree ensemble model is used to guide the exploration and can give the importance of each design parameter. Then, a Pareto-driven AL approach is proposed that prioritizes the exploration of designs with larger hypervolume contributions in the predicted Pareto fronts and allows the acceptance of poor solutions to handle model inaccuracies. Finally, a parallel strategy is utilized to speed up the exploration. The experimental results on the 7-nm RISC-V Berkeley out-of-order machine (BOOM) show that our method can find diversified designs converging to real Pareto fronts more efficiently, achieving better exploration quality and efficiency than previous work.
Jianwang Zhai, Yici Cai
IEEE Trans. Very Large Scale Integr. Syst.1
2021 BOOM-Explorer: RISC-V BOOM Microarchitecture Design Space Exploration Framework
abstract
The microarchitecture design of a processor has been increasingly difficult due to the large design space and time-consuming verification flow. Previously, researchers rely on prior knowledge and cycle-accurate simulators to analyze the performance of different microarchitecture designs but lack sufficient discussions on methodologies to strike a good balance between power and performance. This work proposes an automatic framework to explore microarchitecture designs of the RISC-V Berkeley Out-of-Order Machine (BOOM), termed as BOOM-Explorer, achieving a good trade-off on power and performance. Firstly, the framework utilizes an advanced microarchitecture-aware active learning (MicroAL) algorithm to generate a diverse and representative initial design set. Secondly, a Gaussian process model with deep kernel learning functions (DKL-GP) is built to characterize the design space. Thirdly, correlated multi-objective Bayesian optimization is leveraged to explore Pareto-optimal designs. Experimental results show that BOOM-Explorer can search for designs that dominate previous arts and designs developed by senior engineers in terms of power and performance within a much shorter time.
Qi Sun 0002, Jianwang Zhai, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
ICCAD3
2021 McPAT-Calib: A Microarchitecture Power Modeling Framework for Modern CPUs
abstract
Energy efficiency has become the core issue of modern CPUs, and it is difficult for existing power models to balance speed, generality, and accuracy. This paper introduces McPAT-Calib, a microarchitecture power modeling framework, which combines McPAT with machine learning (ML) calibration methods. McPAT-Calib can quickly and accurately estimate the power of different benchmarks running on different CPU configurations, and provide an effective evaluation tool for the design of modern CPUs. First, McPAT-7nm is introduced to support the analytical power modeling for the 7nm technology node. Then, a wide range of modeling features are identified, and automatic feature selection and advanced regression methods are used to calibrate the McPAT-7nm modeling results, which greatly improves the generality and accuracy. Moreover, a sampling algorithm based on active learning (AL) is leveraged to effectively reduce the labeling cost. We use up to 15 configurations of 7nm RISC-V Berkeley Out-of-Order Machine (BOOM) along with 80 benchmarks to extensively evaluate the proposed framework. Compared with state-of-the-art microarchitecture power models, McPAT-Calib can reduce the mean absolute percentage error (MAPE) of shuffle-split cross-validation by 5.95%. More importantly, the MAPE is reduced by 6.14% and 3.64% for the evaluations of unknown CPU configurations and benchmarks, respectively. The AL sampling algorithm can reduce the demand of labeled samples by 50 %, while the accuracy loss is only 0.44 %.
Jianwang Zhai, Binwu Zhu, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
ICCAD1