VLDB 2026 Research / reviewers in the wild / expert
Yuan Pu 0001
dblp:221/5439-1
· DBLP profile ↗
27ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0002-1322-5642ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 8 first-author · 25 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAEDA: An LLM-Powered Multi-Agent Evaluation Framework for EDA Tool Documentation QAabstractLarge Language Models (LLMs) have shown remarkable capability in knowledge-intensive scenarios, such as electronic design automation (EDA) tool documentation question answering (QA), due to their ability to process and generate contextually rich, domain-specific information. Evaluating LLM outputs is paramount, as it directly impacts their accuracy, effectiveness, and trustworthiness in practical applications. In this paper, we introduce MAEDA, a novel LLM-powered multi-agent evaluation framework that utilizes multiple fine-tuned LLM agents working collaboratively to assess common error types encountered in EDA tool documentation QA. Specifically, we design customized point-to-point alignment and chain-of-thought (CoT) reasoning strategies tailored to specific agents, enhancing both fine-tuning and inference capabilities. Experimental results demonstrate that MAEDA outperforms state-of-the-art (SOTA) general-purpose and cross-domain evaluation frameworks in accurately identifying error types specific to this domain. Our benchmark is publicly available at https://github.com/Rayzzz14/MAEDA-DATE26/. Yuan Pu 0001, Hairuo Han, Yuntao Nie, Jiajun Qin, Yuhan Qin, Tairu Qiu, Zhuolun He, Jianwang Zhai, Bei Yu 0001 |
DATE | 2 |
| 2026 | Smart-PCLib: A LLM-based Multi-Agent Framework for Automated PCB Component Library Generation
Zhaohai Di, Jindong Tu, Yuan Pu 0001, Jiawei Liu 0006, Chong Tong, Tsung-Yi Ho, Bei Yu 0001, Tinghuan Chen |
DATE | 4 |
| 2026 | RATuner: Retrieval-Augmented VLSI Flow Design Parameter Tuning Framework
Peng Xu 0052, Ziyang Yu 0001, Yuan Pu 0001, Xinyun Zhang 0001, Donger Luo, Hao Geng, Tsung-Yi Ho, Bei Yu 0001 |
DATE | 3 |
| 2026 | IncreMacro-3D: Incremental Macro Placement for Face-to-Face Stacked Memory-on-Logic 3D ICsabstractFace-to-face stacked 3D ICs, such as memory-on-logic (MoL) architectures, have emerged as a promising solution to overcome the limitations of traditional 2D integration by offering enhanced performance, power efficiency, and density. Given the increasing design complexity of modern system-on-chips (SoCs), achieving high-quality macro placement is critical, as it plays a decisive role in determining the final performance, power, and area (PPA) metrics. However, existing RTL-to-GDS 3D physical design flows for MoL 3D ICs rely heavily on manual macro placement, which becomes increasingly challenging and time-consuming for modern SoCs with a vast number of macros. In this paper, we introduce an innovative macro placement algorithm, IncreMacro-3D, which employs graph neural network-based macro repartitioning and 3D macro position refinement, thereby facilitating subsequent steps in 3D physical design flow. The experimental results on several benchmark circuits demonstrate that the proposed approach can reduce the routed wirelength, worst negative slack (WNS), total negative slack (TNS), and total power consumption by 6.1%, 44.2%, 62.8%, and 0.6% compared to state-of-the-art analytical placer for MoL 3D ICs. Lancheng Zou, Sing Sen Ye, Yuan Pu 0001, Jiaxi Jiang, Siting Liu 0002, Yuxuan Zhao 0001, Bei Yu 0001 |
DATE | 4 |
| 2026 | Invited: Infusing EDA Knowledge into LLM Systems: An Information-Source PerspectiveabstractLarge language models have shown remarkable potential for electronic design automation (EDA), yet building effective LLM systems for EDA remains challenging due to complex tool-specific terminology and documentation. This paper surveys knowledge injection techniques that infuse domain expertise into LLM systems for EDA. We examine three complementary approaches: finetuning, which encodes EDA knowledge into model parameters through training on domain corpora and synthetic data; retrieval-augmented generation (RAG), which dynamically retrieves from external knowledge bases; and multi-agent flow, which decomposes complex tasks across specialized agents and leverages environment feedback for iterative refinement. As a case study, we present a graph-based RAG approach that addresses global queries requiring cross-chunk reasoning. The method trains document-customized embeddings via contrastive learning on knowledge graphs, detects semantically related entities using HDBSCAN clustering, and generates textual summaries integrated through hybrid retrieval. Experiments on OpenROAD documentation demonstrate significant improvements in answering global queries while maintaining local query performance. These findings highlight that domain customization is essential for effective knowledge injection, and graph-based techniques are particularly promising as they inherently encode domain knowledge through entity extraction and relationship modeling. Yuhan Qin, Yuan Pu 0001, Tairu Qiu, Zhuolun He, Bei Yu 0001 |
ISPD | 2 |
| 2026 | RegPlace: Regularity-Aware Placement for Full-System DNN Accelerator DesignsabstractThe rise of deep neural network accelerators demands physical design tools that recognize spatial regularity patterns. Traditional placers, unaware of the regularity of spatial arrays, produce suboptimal solutions. This work proposes RegPlace, a regularity-aware placement algorithm for full-system DNN accelerators that automatically identifies processing elements using graph convolutional networks and employs a variance-based soft regularity loss to guide optimization. Compared to state-of-the-art methods, our approach achieves up to 6% wirelength reduction while maintaining comparable runtime, with post-placement metrics further confirming its effectiveness. Jiaxi Jiang, Yuan Pu 0001, Yuxuan Zhao 0001, Peiyu Liao, Zuodong Zhang, Yibo Lin, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | PAPlace: Performance-Driven Differentiable Analog PlacementabstractAnalog circuit placement is crucial for optimal performance, but achieving a decent layout demands expertise and time. Recent advances in machine learning techniques have shown promising results in modeling analog layout performance. PAPlace further extends these methods and integrates them into the core analog placement engine, allowing direct optimization of the post-layout performance effectively. Our approach proposes a differentiable prediction model that combines layout and wiring information into a non-linear analog placement engine. We then incorporate the differentiable performance model into a gradient-descent-based global placement engine. A multi-objective optimization method is further proposed to find the common gradient descent direction for different metrics. The experimental results on benchmarks under the TSMC 40nm technology node demonstrate the superiority of the proposed framework compared with the cutting-edge works, with up to 2163.00μ V , 73.95dB, 62.25MHz, 57.84dB improvement in Offset Voltage, CMRR, BandWidth, DC Gain metrics. Peng Xu 0052, Yuan Pu 0001, Keren Zhu 0001, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | Late Breaking Results: Hybrid Logic Optimization with Predictive Self-SupervisionabstractHybrid optimization is an emerging approach in logic synthesis, focusing on applying diverse optimization methods to different parts of a logic circuit. This paper analyzes the relationship between each vertex and its corresponding optimization method. We extract a subgraph centered on each vertex and quantify the logic optimization results of these subgraphs as vertex features. Based on these features, we propose a circuit partitioning method to cluster the logic circuit, enabling the final optimized circuit to be constructed by merging clusters optimized with their respective methods. Additionally, we introduce a self-supervised prediction model to efficiently obtain vertex features. The experimental results targeting LUT mapping demonstrate that our method achieves improvements of $8.48 \%$ in area and 9.81% in delay compared to the state-of-the-art. Rongliang Fu, Zhengyuan Shi, Yuan Pu 0001, Junying Huang, Qiang Xu 0001, Tsung-Yi Ho |
DAC | 5 |
| 2025 | DSPlacer: DSP Placement for FPGA-based CNN AcceleratorabstractDeploying convolutional neural networks (CNNs) on hardware platforms like Field Programmable Gate Arrays (FPGAs) has garnered significant attention due to their inherent flexibility and parallelism. Achieving optimal timing closure remains a critical challenge, as placement directly impacts clock frequency and throughput. Existing approaches often face scalability issues with large designs or fail to formalize placement rules into automated algorithms. In this paper, we propose DSPlacer, a novel DSP placement framework designed for diverse CNN accelerator architectures in the context of FPGA design. The proposed approach iteratively optimizes the placement of datapath DSPs to enhance timing performance. To achieve this, DSPlacer integrates several advanced techniques, including graph convolutional network-based datapath DSP identification, DSP graph construction, min-cost-flow DSP assignment, and integer linear programming (ILP)-based cascade constraint legalization. These techniques collectively address two key requirements for datapath DSP placement: (1) cascading datapath DSPs to achieve a compact layout, and (2) preserving direct datapath information between the processing system and programmable logic. The framework has been evaluated on multiple academic benchmarks and compared against AMD Xilinx Vivado 2020.2 and AMF-Placer 2.0. Experimental results demonstrate that DSPlacer improves Worst Negative Slack (WNS) by 32% and 65%, respectively, highlighting its efficacy and superiority. Baohui Xie, Xinrui Zhu, Yuan Pu 0001, Tongkai Wu, Xiaofeng Zou, Bei Yu 0001, Tinghuan Chen |
DAC | 4 |
| 2025 | MM-GRADE: A Multi-Modal EDA Tool Documentation QA Framework Leveraging Retrieval Augmented GenerationabstractThe complexity of EDA tools necessitates the development of advanced documentation query answering systems to enhance user efficiency and reduce the associated learning curve. Recent innovations in the use of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) for EDA tool documentation have demonstrated significant progress; however, these approaches typically lack the multi-modal capabilities required to effectively handle visual data, such as circuit layout and GUI screenshots provided through user input. To address the concern, we introduce a multi-modal RAG system that incorporates two domain-customized modules: a multi-modal retriever model finetuned by the customized bilevel hard negative mining (BHNM) strategy, and a vision large language model (VLLM) finetuned using a tailored extract-score-answer pipeline. Moreover, we have manually curated ORD-MMBench, a multi-modal QA benchmark comprising 120 high-quality question-document-answer triplets based on OpenROAD documentation. Experimental results demonstrate that our customized RAG framework outperforms state-of-the-art multi-modal RAG flows and models on ORD-MMBench. Yuan Pu 0001, Zhuolun He, Shutong Lin, Jiajun Qin, Xinyun Zhang 0001, Hairuo Han, Haisheng Zheng, Cheng Zhuo, Qi Sun 0002, David Z. Pan, Bei Yu 0001 |
ICCAD | 1 |
| 2025 | GPU Acceleration for Versatile Buffer InsertionabstractWith the advancement of circuit design complexity and technology nodes, buffer insertion has become pivotal in mitigating timing violations, significantly impacting the physical design development cycle and highlighting the necessity for acceleration methodologies. In this paper, we present BIGX, a GPU-accelerated algorithmic framework for buffer insertion. BIGX is versatile and can be adapted to implement different dynamic programming (DP) based buffering algorithms for repairing various types of timing violations. In particular, we introduce MCDP, a dedicated DP-based buffering algorithm for repairing maximum capacitance violations, and propose a parallel version of Van Ginneken’s algorithm for setup violations, both algorithms are incorporated and implemented in BIGX. Furthermore, to overcome the runtime limitations of DP-based buffering algorithms, BIGX adopts a distributed Branch Merge algorithm based on bucket sorting, which fully leverages the hierarchical memory architecture of modern GPUs to achieve substantial speedups while preserving solution quality. Experimental results on industrial benchmarks demonstrate that, with the integration of MCDP, BIGX repairs 96.6% of maximum capacitance violations. Compared to OpenROAD, BIGX with MCDP repairs 2.54x more maximum capacitance violations and delivers a 3.37x speedup. Additionally, BIGX accelerates the Van Ginneken’s algorithm by 11.68x while maintaining comparable solution quality to its CPU-based counterpart. Yuan Pu 0001, Yuhao Ji, Siying Yu, Zuodong Zhang, Zizheng Guo 0001, Zhuolun He, Yibo Lin, David Z. Pan, Bei Yu 0001 |
ICCAD | 1 |
| 2025 | Circuit Representation Learning with Masked Gate Modeling and Verilog-AIG AlignmentabstractUnderstanding the structure and function of circuits is crucial for electronic design automation (EDA).
Circuits can be formulated as And-Inverter graphs (AIGs), enabling efficient implementation of representation learning through graph neural networks (GNNs).
Masked modeling paradigms have been proven effective in graph representation learning.
However, masking augmentation to original circuits will destroy their logical equivalence, which is unsuitable for circuit representation learning.
Moreover, existing masked modeling paradigms often prioritize structural information at the expense of abstract information such as circuit function.
To address these limitations, we introduce MGVGA, a novel constrained masked modeling paradigm incorporating masked gate modeling (MGM) and Verilog-AIG alignment (VGA).
Specifically, MGM preserves logical equivalence by masking gates in the latent space rather than in the original circuits, subsequently reconstructing the attributes of these masked gates.
Meanwhile, large language models (LLMs) have demonstrated an excellent understanding of the Verilog code functionality.
Building upon this capability, VGA performs masking operations on original circuits and reconstructs masked gates under the constraints of equivalent Verilog codes, enabling GNNs to learn circuit functions from LLMs.
We evaluate MGVGA on various logic synthesis tasks for EDA and show the superior performance of MGVGA compared to previous state-of-the-art methods.
Our code is available at https://github.com/wuhy68/MGVGA. Haoyuan Wu, Haisheng Zheng, Yuan Pu 0001, Bei Yu 0001 |
ICLR | 3 |
| 2025 | HeLO: A Heterogeneous Logic Optimization Framework by Hierarchical Clustering and Graph LearningabstractModern very large-scale integration (VLSI) designs usually consist of modules with various topological structures and functionalities. To better optimize such large and heterogeneous logic networks, it is essential to identify the structural and functional characteristics of its modules, and represent them with appropriate DAG types (such as AIG, MIG, XAG, etc.) for logic optimization. This paper proposes HeLO, a hetero-DAG logic optimization framework empowered by hierarchical clustering and graph learning. HeLO leverages a hierarchical clustering algorithm, which splits the original Boolean network into sub-circuits by considering both topological and functional characteristics. A novel graph neural network model is customized to generate the topological-functional embedding (used for distance calculation in hierarchical clustering) and predict the best-fit DAG type of each sub-circuit. Experimental results demonstrate that HeLO outperforms LSOracle, the SOTA heterogeneous logic optimization framework, in terms of node-depth product (for technology-independent logic optimization) and delay-area product (for technology mapping) by 8.7% and 6.9%, respectively. Yuan Pu 0001, Fangzhou Liu 0005, Zhuolun He, Keren Zhu 0001, Rongliang Fu, Ziyi Wang 0010, Tsung-Yi Ho, Bei Yu 0001 |
ISPD | 1 |
| 2025 | IncreMacro: Incremental Macro Placement RefinementabstractThis article proposes$\textsf {IncreMacro}$, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this,$\textsf {IncreMacro}$utilizes several methods, including kd-tree-based macro diagnosis, gradient-based macro shifting, constraint-graph-based LP for macro legalization, and diffusion-based cell migration. By employing these techniques iteratively,$\textsf {IncreMacro}$meets two critical solution requirements of macro placement: 1) pushing macros toward the chip boundary and 2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into$\textsf {AutoDMP}$and$\textsf {DREAMPlace}~4.0$, and is evaluated on seven RISC-V benchmark circuits and four TILOS macro placement circuit designs at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by$\textsf {AutoDMP}~(\textsf {DREAMPlace}~4.0$), our approach reduces routed wirelength by 15.1% (14.9%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 99.9 (82.6%) and 99.9% (81.3%), and reduces the total power consumption by 4.4% (4.3%). Meanwhile, compared with$\textsf {IncreMacro}$[1], our approach augmented with the cell migration algorithm improves the routed WNS and TNS by 24.7% and 23.1%, and remains the average routed wirelength and total power consumption almost unchanged. Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Jiajun Qin, Haisheng Zheng, Yibo Lin, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QAabstractRetrieval augmented generation (RAG) improves the accuracy and dependability of generative AI models by integrating factual information from external databases. This technique is widely used in tasks involving document-grounded question answering (QA). While these RAG systems are extensively pretrained on general-purpose documents, they face considerable limitations when applied to specialized, knowledgeintensive fields such as electronic design automation (EDA). This paper addresses such issue by proposing a customized RAG framework along with three domain-specific techniques for EDA tool documentation QA, including a contrastive learning scheme for text embedding model fine-tuning, a reranker distilled from proprietary LLM, and a generative LLM fine-tuned with highquality domain corpus. To further unleash the extraordinary language capacity of LLMs in the domain of EDA-tool documentation QA, we propose to train LLMs as the reranker model with our customized two-stage traning scheme, which consists of the point-wise instruction tuning stage and the pairwise learn-to-rank (LTR) stage. Finally, we have developed and released a documentation QA evaluation benchmark, ORD-QA, for OpenROAD, an advanced RTL-to-GDSII design platform. Experimental results demonstrate that our proposed RAG flow and techniques have achieved superior performance on ORD-QA as well as on a commercial tool, compared with state-of-thearts. Furthermore, compared with the SOTA reranker models, our LLM reranker prominently improves the document retrieval accuracy and thus leads to better QA quality. The ORD-QA benchmark and the training dataset for our customized RAG flow are open-source at https://github.com/lesliepy99/RAG-EDA. Yuan Pu 0001, Zhuolun He, Tairu Qiu, Haoyuan Wu, Qi Sun 0002, Cheng Zhuo, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | PRO-TIME: Prerouting Optimization-Aware Timing Prediction via Multimodal LearningabstractFast and accurate pre-routing timing prediction is crucial in the very-large-scale integration (VLSI) design flow. Existing machine learning (ML)-assisted pre-routing timing evaluators neglect the impact of timing optimization, which may render their approaches impractical in real circuit design flows. To address the challenges posed by timing optimization, we propose PRO-TIME, a pre-routing optimization-aware timing prediction framework that is driven by multimodal learning. Specifically, we propose a novel endpoint embedding framework that integrates both netlist and layout information. A customized graph neural network (GNN) model is used for extracting endpoint-wise netlist information, which is motivated by the delay propagation process. Meanwhile, we apply the U-net model with a masking strategy to extract endpoint-wise layout information. Furthermore, we propose an adaptive layout mask adjustment scheme to boost performance by leveraging the layout information more effectively. Comprehensive experiments on large-scale RISC-V designs with advanced 7-nm technology node demonstrate the superiority of our model compared to the state-of-the-art pre-routing timing evaluators. Ziyi Wang 0010, Siting Liu 0002, Yuan Pu 0001, Song Chen 0001, Tsung-Yi Ho, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | ParSGCN: Bridging the Gap Between Emulation Partitioning and SchedulingabstractEfficient functional verification is crucial in the very-large-scale integration (VLSI) design flow. Existing processor-based emulation systems suffer from low efficiency due to the gap between partitioning and scheduling during compilation. To address the above concern, we propose ParSGCN, a scheduling-friendly emulation compilation flow that considers the objective of scheduling during partitioning. To incorporate the hard-to-perceive look-ahead information about scheduling, we embed it into a net cut probability distribution, which is easier to utilize. We estimate this probability distribution using a tailored variant of graph convolutional network (GCN) that is trained through a customized loss function and a large dataset of real-world compilation solutions. Additionally, we have developed a set of novel techniques to guide the emulation partitioning process using the estimated probability distribution. The proposed method is integrated into an industrial emulator and evaluated on large-scale designs with up to over 100 million cells. Comprehensive experimental results demonstrate the effectiveness of ParSGCN, showcasing an average improvement of 16.38%, 26.04%, and 19.52% in the best, worst, and median solution quality, respectively, based on 50 runs. Ziyi Wang 0010, Wenqian Zhao 0002, Yuan Pu 0001, Lei Chen 0031, Wilson W. K. Thong, Weihua Sheng, Tsung-Yi Ho, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Large Language Models for EDA: Future or Mirage?abstractIn this article, we explore the burgeoning intersection of large language models (LLMs) and electronic design automation (EDA). We critically assess whether LLMs represent a transformative future for EDA or merely a fleeting mirage. By organizing existing research into four critical domains of EDA—code generation, verification and debugging, knowledge representation and retrieval, and optimization/modeling—we provide a comprehensive overview of the current state-of-the-art. The survey concludes with a 5-level roadmap to guide the progressive integration and advancement of LLMs in EDA. Ultimately, this article aims to provide a comprehensive, evidence-based perspective on the role of LLMs in shaping the future of EDA. Zhuolun He, Yuan Pu 0001, Haoyuan Wu, Tairu Qiu, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | EasyMRC: Efficient Mask Rule Checking via Representative Edge SamplingabstractThe photolithography process is getting more sophisticated with technology node scaling down and VLSI designs becoming complex. As photomask patterns get finer, mask rule checks (MRCs) are inevitable to avoid discrepancies in the layout and to ensure manufacturability. This paper introduces an efficient mask rule checking approach that utilizes a representative edge sampling scheme. The representative edge sampling scheme selects a subset of edges and points of each polygon that capture its contour, meanwhile greatly reducing the number of edges involved in actual checking. Experimental results demonstrate that the proposed approach achieves significant speedup compared with the state-of-the-art academic tool. Zhuolun He, Yuan Pu 0001, Wenjian Yu, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | NeuroSelect: Learning to Select Clauses in SAT SolversabstractModern SAT solvers depend on conflict-driven clause learning to avoid recurring conflicts. Deleting less valuable learned clauses is a crucial component of modern SAT solvers to ensure efficiency. However, a single clause deletion policy cannot guarantee optimal performance on all SAT instances. This paper introduces a new clause deletion metric to diversify existing clause deletion policies. Then, we propose to use machine learning to evaluate and select clause deletion policies adaptively based on the input instance. We show that our method can reduce the runtime of the state-of-the-art SAT solver Kissat by 5.8% on large industry benchmarks. Hongduo Liu, Peng Xu 0052, Yuan Pu 0001, Lihao Yin, Hui-Ling Zhen, Mingxuan Yuan, Tsung-Yi Ho, Bei Yu 0001 |
DAC | 3 |
| 2024 | Lesyn: Placement-aware Logic Resynthesis for Non-Integer Multiple-Cell-Height DesignsabstractNon-integer multiple cell height (NIMCH) standard-cell libraries offer promising co-optimization for power, performance and area in advanced technology nodes. However, such non-uniform design introduces new layout constraints where any sub-region can only accommodate gates of the same cell height due to manufacturability concerns. The existing physical design flow for NIMCH circuits, which handles the layout constraint by clustering and relocating gates according to their cell heights, often leads to substantial gate displacement that harms circuit performance. To alleviate the above issue, this paper proposes a row-based logic resynthesis procedure that explicitly adjusts cell heights after initial placement without changing cell positions. Experiment results demonstrate that compared with the conventional NIMCH physical design flow, our proposed approach can reduce the maximal delay by 26.1%. Yuan Pu 0001, Fangzhou Liu 0005, Yu Zhang 0189, Zhuolun He, Yibo Lin, Kai-Yuan Chao, Bei Yu 0001 |
DAC | 1 |
| 2024 | Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QAabstractRetrieval augmented generation (RAG) enhances the accuracy and reliability of generative AI models by sourcing factual information from external databases, which is extensively employed in document-grounded question-answering (QA) tasks. Off-the-shelf RAG flows are well pretrained on general-purpose documents, yet they encounter significant challenges when being applied to knowledge-intensive vertical domains, such as electronic design automation (EDA). This paper addresses such issue by proposing a customized RAG framework along with three domain-specific techniques for EDA tool documentation QA, including a contrastive learning scheme for text embedding model fine-tuning, a reranker distilled from proprietary LLM, and a generative LLM fine-tuned with high-quality domain corpus. Furthermore, we have developed and released a documentation QA evaluation benchmark, ORD-QA, for OpenROAD, an advanced RTL-to-GDSII design platform. Experimental results demonstrate that our proposed RAG flow and techniques have achieved superior performance on ORD-QA as well as on a commercial tool, compared with state-of-the-arts. The ORD-QA benchmark and the training dataset for our customized RAG flow are open-source at https://github.com/lesliepy99/RAG-EDA. Yuan Pu 0001, Zhuolun He, Tairu Qiu, Haoyuan Wu, Bei Yu 0001 |
ICCAD | 1 |
| 2024 | IncreMacro: Incremental Macro Placement RefinementabstractThis paper proposes IncreMacro, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this, IncreMacro utilizes several methods including kd-tree-based macro diagnosis, gradient-based macro shifting and constraint-graph-based LP for macro legalization. By employing these techniques iteratively, IncreMacro meets two critical solution requirements of macro placement: (1) pushing macros to the chip boundary; and (2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into DREAMPlace and AutoDMP, and is evaluated on several RISC-V benchmark circuits at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by DREAMPlace (AutoDMP), IncreMacro reduces routed wirelength by 6.5% (16.8%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 59.9% (99.6%) and 63.9% (99.9%), and reduces the total power consumption by 3.3% (4.9%). Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Haisheng Zheng, Yibo Lin, Bei Yu 0001 |
ISPD | 1 |
| 2024 | Multi-Electrostatics Based Placement for Non-Integer Multiple-Height CellsabstractA circuit design incorporating non-integer multi-height (NIMH) cells, such as a combination of 8-track and 12-track cells, offers increased flexibility in optimizing area, timing, and power simultaneously. The conventional approach for placing NIMH cells involves using commercial tools to generate an initial global placement, followed by a legalization process that divides the block area into row regions with specific heights and relocates cells to rows of matching height. However, such placement flow often causes significant disruptions in the initial placement results, resulting in inferior wirelength. To address this issue, we propose a novel multi-electrostatics-based global placement algorithm that utilizes the NIMH-aware clustering method to dynamically generate rows. This algorithm directly tackles the global placement problem with NIMH cells. Specifically, we utilize an augmented Lagrangian formulation along with a preconditioning technique to achieve high-quality solutions with fast and robust numerical convergence. Experimental results on the OpenCores benchmarks demonstrate that our algorithm achieves about 12% improvements on HPWL with 23.5X speed up on average, outperforming state-of-the-art approaches. Furthermore, our placement solutions demonstrate a substantial improvement in WNS and TNS by 22% and 49% respectively. These results affirm the efficiency and effectiveness of our proposed algorithm in solving row-based placement problems for NIMH cells. Yu Zhang 0189, Yuan Pu 0001, Fangzhou Liu 0005, Peiyu Liao, Kai-Yuan Chao, Keren Zhu 0001, Yibo Lin, Bei Yu 0001 |
ISPD | 2 |
| 2023 | Restructure-Tolerant Timing Prediction via Multimodal FusionabstractFast and accurate pre-routing timing prediction is crucial in the very-large-scale integration (VLSI) design flow. Existing machine learning (ML)-assisted pre-routing timing evaluators neglect the impact of timing optimization, which may render their approaches impractical in real circuit design flows. To model the impact of timing optimization, we propose an endpoint embedding framework that integrates netlist-layout information via multimodal fusion. An end-to-end flow is further developed for pre-routing restructure-tolerant prediction on global timing metrics. Comprehensive experiments on large-scale RISC-V designs with advanced 7-nm technology node demonstrate the superiority of our model compared to the SOTA pre-routing timing evaluators. Ziyi Wang 0010, Siting Liu 0002, Yuan Pu 0001, Song Chen 0001, Tsung-Yi Ho, Bei Yu 0001 |
DAC | 3 |
| 2023 | FastGR: Global Routing on CPU-GPU with Heterogeneous Task Graph Scheduler (Extended Abstract)abstractRunning time is a key metric across the standard physical design flow stages. However, with the rapid growth in design sizes, routing runtime has become the runtime bottleneck in the physical design flow. To improve the effectiveness of the modern global router, we propose a global routing framework with GPU-accelerated routing algorithms and a heterogeneous task graph scheduler, called FastGR. Its runtime-oriented version FastGRL achieves 2.489× speedup compared with the state-of-the-art global router. Furthermore, the GPU-accelerated L-shape pattern routing used in FastGRL can contribute to 9.324× speedup over the sequential algorithm on CPU. Its quality-oriented version FastGRH offers further quality improvement over FastGRL with similar acceleration. Siting Liu 0002, Yuan Pu 0001, Peiyu Liao, Hongzhong Wu, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001 |
IJCAI | 2 |
| 2023 | FastGR: Global Routing on CPU-GPU With Heterogeneous Task Graph SchedulerabstractRunning time is a key metric across the standard physical design flow stages. However, with the rapid growth in design sizes, routing runtime has become the runtime bottleneck in the physical design flow. As a result, speeding routing becomes a critical and pressing task for IC design automation. Aside from the running time, we need to evaluate the quality of the global routing solution since a poor global routing engine degrades the solution performance after the entire routing stage. This work takes both of them into consideration. We propose a global routing framework with GPU-accelerated routing algorithms and a heterogeneous task graph scheduler, called FastGR, to accelerate the procedure of the modern global router and improve its effectiveness. Its runtime-oriented version$\text {FastGR}^{\text {L}}$achieves$2.489\times $speedup compared with the state-of-the-art global router. Furthermore, the GPU-accelerated L-shape pattern routing algorithm used in$\text {FastGR}^{\text {L}}$can contribute to$9.324\times $speedup over the sequential algorithm on CPU. Its quality-oriented version$\text {FastGR}^{\text {H}}$offers a 27.855% improvement of the number of shorts over the runtime-oriented version and still gets$1.970\times $faster than the most advanced global router. Siting Liu 0002, Yuan Pu 0001, Peiyu Liao, Hongzhong Wu, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |