EDBT 2026 Demo / reviewers in the wild / expert
Zhuolun He
dblp:194/3907
· DBLP profile ↗
44ranked-venue papers
10as first author
39since 2021 · last 2026
0009-0009-4909-6588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 10 first-author · 34 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CombRewriter: Enabling Combinational Logic Simplification in MLIR-Based Hardware CompilerabstractModern Hardware Description Languages (HDLs) play a pivotal role in enabling swift and adaptable hardware development. A hardware compiler translates high-level designer intents into a concrete hardware implementation, the quality of which directly determines ultimate circuit performance. However, current hardware compilers may overlook opportunities for combinational logic simplification, leading to RTL code that contains redundant logic and degrades the Quality of Results (QoR) of the synthesized netlist. This paper presents CombRewriter, a novel approach that incorporates compilation-level optimization techniques into combinational logic simplification. Experimental results demonstrate that the proposed method effectively reduces netlist area. Haisheng Zheng, Zhuolun He, Yuzhe Ma, Bei Yu 0001 |
ASP-DAC | 2 |
| 2026 | MAEDA: An LLM-Powered Multi-Agent Evaluation Framework for EDA Tool Documentation QAabstractLarge Language Models (LLMs) have shown remarkable capability in knowledge-intensive scenarios, such as electronic design automation (EDA) tool documentation question answering (QA), due to their ability to process and generate contextually rich, domain-specific information. Evaluating LLM outputs is paramount, as it directly impacts their accuracy, effectiveness, and trustworthiness in practical applications. In this paper, we introduce MAEDA, a novel LLM-powered multi-agent evaluation framework that utilizes multiple fine-tuned LLM agents working collaboratively to assess common error types encountered in EDA tool documentation QA. Specifically, we design customized point-to-point alignment and chain-of-thought (CoT) reasoning strategies tailored to specific agents, enhancing both fine-tuning and inference capabilities. Experimental results demonstrate that MAEDA outperforms state-of-the-art (SOTA) general-purpose and cross-domain evaluation frameworks in accurately identifying error types specific to this domain. Our benchmark is publicly available at https://github.com/Rayzzz14/MAEDA-DATE26/. Yuan Pu 0001, Hairuo Han, Yuntao Nie, Jiajun Qin, Yuhan Qin, Tairu Qiu, Zhuolun He, Jianwang Zhai, Bei Yu 0001 |
DATE | 8 |
| 2026 | Invited: Infusing EDA Knowledge into LLM Systems: An Information-Source PerspectiveabstractLarge language models have shown remarkable potential for electronic design automation (EDA), yet building effective LLM systems for EDA remains challenging due to complex tool-specific terminology and documentation. This paper surveys knowledge injection techniques that infuse domain expertise into LLM systems for EDA. We examine three complementary approaches: finetuning, which encodes EDA knowledge into model parameters through training on domain corpora and synthetic data; retrieval-augmented generation (RAG), which dynamically retrieves from external knowledge bases; and multi-agent flow, which decomposes complex tasks across specialized agents and leverages environment feedback for iterative refinement. As a case study, we present a graph-based RAG approach that addresses global queries requiring cross-chunk reasoning. The method trains document-customized embeddings via contrastive learning on knowledge graphs, detects semantically related entities using HDBSCAN clustering, and generates textual summaries integrated through hybrid retrieval. Experiments on OpenROAD documentation demonstrate significant improvements in answering global queries while maintaining local query performance. These findings highlight that domain customization is essential for effective knowledge injection, and graph-based techniques are particularly promising as they inherently encode domain knowledge through entity extraction and relationship modeling. Yuhan Qin, Yuan Pu 0001, Tairu Qiu, Zhuolun He, Bei Yu 0001 |
ISPD | 4 |
| 2026 | Routing-aware Legal Hybrid Bonding Terminal Assignment for 3D Face-to-Face Stacked ICsabstractFace-to-face (F2F) stacked three-dimensional (3D) IC is a promising alternative for scaling beyond Moore’s Law. In F2F 3D ICs, dies are connected through bonding terminals whose positions can significantly impact routing performance. Further, there exists resource competition among all the 3D nets due to the constrained bonding terminal number. In advanced technology nodes, traditional bonding terminal planning may also introduce legality challenges of bonding terminals, as the metal pitches can be much smaller than the sizes of bonding terminals. Previous works attempt to insert bonding terminals automatically using existing 2D commercial P&R tools and then consider interdie connection legality, but they fail to take the legality and routing performance into account simultaneously. In this article, we provide a novel bonding terminal assignment formulation for effective routing-aware bonding terminal planning. We explore the generalized assignment formulation and provide the routability guidance in our hybrid bonding terminal assignment problem. Our framework, BTAssign , offers a strict legality guarantee and an iterative solution. We provide two versions of the BTAssign framework, BTAssign-WL [ 1 ] and BTAssign-R, which BTAssign-R extends BTAssign-WL [ 1 ] by considering routability. The experiments are conducted on 18 open source designs with various 3D net densities and the most advanced bonding scale. The results reveal that all the testing cases with different partitioning and placement strategies could gain benefits from our BTAssign framework. Siting Liu 0002, Jieya Zhou, Jiaxi Jiang, Zhuolun He, Ziyi Wang 0010, Yibo Lin, Bei Yu 0001, Martin D. F. Wong |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | Efficient OpAmp Adaptation for Zoom Attention to Golden ContextsabstractLarge language models (LLMs) have shown significant promise in question-answering (QA) tasks, particularly in retrieval-augmented generation (RAG) scenarios and long-context applications. However, their performance is hindered by noisy reference documents, which often distract from essential information. Despite fine-tuning efforts, Transformer-based architectures struggle to prioritize relevant content. This is evidenced by their tendency to allocate disproportionate attention to irrelevant or later-positioned documents. Recent work proposes the differential attention mechanism to address this issue, but this mechanism is limited by an unsuitable common-mode rejection ratio (CMRR) and high computational costs. Inspired by the operational amplifier (OpAmp), we propose the OpAmp adaptation to address these challenges, which is implemented with adapters efficiently. By integrating the adapter into pre-trained Transformer blocks, our approach enhances focus on the golden context without costly training from scratch. Empirical evaluations on noisy-context benchmarks reveal that our Qwen2.5-OpAmp-72B model, trained with our OpAmp adaptation, surpasses the performance of state-of-the-art LLMs, including DeepSeek-V3 and GPT-4o.Our code is available at https://github.com/wuhy68/OpampAdapter. Haoyuan Wu, Rui Ming, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
ACL (1) | 4 |
| 2025 | Swift or Exact? Boosting Efficient Microarchitecture DSE via Multi-fidelity Partial Order PredictionabstractA significant challenge in microarchitecture design space exploration (DSE) lies in the time-intensive synthesis and simulation process, making rapid design exploration infeasible. While the simulation tools offer reports on performance, power, and area (PPA) in the different stages, the PPA reports at early stages may fail to reflect the true relative qualities for various designs, i.e., with low fidelities. To address these limitations, we propose a novel multi-fidelity optimization algorithm tailored for multi-stage optimization problems. The proposed method employs a non-linear Gaussian process model to effectively fuse data from different stages with different fidelities, minimizing the need for expensive high-fidelity data while maximizing accuracy. Furthermore, a logical regression function and a multi-objective partial order relation are introduced to evaluate the reliability of low-fidelity data, mitigating their potential inaccuracies. Experiments demonstrate that our proposed multi-fidelity optimization algorithm can approximate the Pareto front of the direct design space in a shorter time with better performance. Hao Geng, Zhuolun He, Qi Sun 0002, Cheng Zhuo |
DAC | 3 |
| 2025 | ChatLS: Multimodal Retrieval-Augmented Generation and Chain-of-Thought for Logic Synthesis Script CustomizationabstractLarge Language Models (LLMs) have demonstrated significant potential in automating the Electronic Design Automation (EDA) process through effective integration with EDA tools. This paper targets the customization of logic synthesis scripts, which is crucial for accommodating the unique characteristics of each design in the EDA workflow. The proposed framework, called ChatLS, integrates multimodal retrieval-augmented generation (RAG) and chain-of-thought (CoT) reasoning, enabling LLMs to collaboratively analyze design features and precisely customize synthesis scripts. Experimental results demonstrate that ChatLS has achieved superior performance in customizing synthesis scripts with a commercial logic synthesis tool. Haisheng Zheng, Haoyuan Wu, Zhuolun He |
DAC | 3 |
| 2025 | iRw: An Intelligent RewritingabstractThis paper proposes a novel machine learning-driven rewriting algorithm to optimize And-Inverter Graphs (AIGs) for refining combinational logic prior to technology mapping. The algorithm, called i Rw, iteratively extracts subcircuits in AIGs and replaces them with more streamlined implementations. These subcircuits are identified using an original extraction algorithm, while the compact implementations are produced through rewriting techniques guided by a machine learning model. This approach efficiently enables the generation of logically equivalent subcircuits with minimal overhead. Experiments on benchmark circuits indicate that the proposed methodology outperforms state-of-the-art AIG rewriting techniques in both quality and runtime. Haisheng Zheng, Haoyuan Wu, Zhuolun He, Yuzhe Ma, Bei Yu 0001 |
DATE | 3 |
| 2025 | MM-GRADE: A Multi-Modal EDA Tool Documentation QA Framework Leveraging Retrieval Augmented GenerationabstractThe complexity of EDA tools necessitates the development of advanced documentation query answering systems to enhance user efficiency and reduce the associated learning curve. Recent innovations in the use of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) for EDA tool documentation have demonstrated significant progress; however, these approaches typically lack the multi-modal capabilities required to effectively handle visual data, such as circuit layout and GUI screenshots provided through user input. To address the concern, we introduce a multi-modal RAG system that incorporates two domain-customized modules: a multi-modal retriever model finetuned by the customized bilevel hard negative mining (BHNM) strategy, and a vision large language model (VLLM) finetuned using a tailored extract-score-answer pipeline. Moreover, we have manually curated ORD-MMBench, a multi-modal QA benchmark comprising 120 high-quality question-document-answer triplets based on OpenROAD documentation. Experimental results demonstrate that our customized RAG framework outperforms state-of-the-art multi-modal RAG flows and models on ORD-MMBench. Yuan Pu 0001, Zhuolun He, Shutong Lin, Jiajun Qin, Xinyun Zhang 0001, Hairuo Han, Haisheng Zheng, Cheng Zhuo, Qi Sun 0002, David Z. Pan, Bei Yu 0001 |
ICCAD | 2 |
| 2025 | GPU Acceleration for Versatile Buffer InsertionabstractWith the advancement of circuit design complexity and technology nodes, buffer insertion has become pivotal in mitigating timing violations, significantly impacting the physical design development cycle and highlighting the necessity for acceleration methodologies. In this paper, we present BIGX, a GPU-accelerated algorithmic framework for buffer insertion. BIGX is versatile and can be adapted to implement different dynamic programming (DP) based buffering algorithms for repairing various types of timing violations. In particular, we introduce MCDP, a dedicated DP-based buffering algorithm for repairing maximum capacitance violations, and propose a parallel version of Van Ginneken’s algorithm for setup violations, both algorithms are incorporated and implemented in BIGX. Furthermore, to overcome the runtime limitations of DP-based buffering algorithms, BIGX adopts a distributed Branch Merge algorithm based on bucket sorting, which fully leverages the hierarchical memory architecture of modern GPUs to achieve substantial speedups while preserving solution quality. Experimental results on industrial benchmarks demonstrate that, with the integration of MCDP, BIGX repairs 96.6% of maximum capacitance violations. Compared to OpenROAD, BIGX with MCDP repairs 2.54x more maximum capacitance violations and delivers a 3.37x speedup. Additionally, BIGX accelerates the Van Ginneken’s algorithm by 11.68x while maintaining comparable solution quality to its CPU-based counterpart. Yuan Pu 0001, Yuhao Ji, Siying Yu, Zuodong Zhang, Zizheng Guo 0001, Zhuolun He, Yibo Lin, David Z. Pan, Bei Yu 0001 |
ICCAD | 6 |
| 2025 | HeLO: A Heterogeneous Logic Optimization Framework by Hierarchical Clustering and Graph LearningabstractModern very large-scale integration (VLSI) designs usually consist of modules with various topological structures and functionalities. To better optimize such large and heterogeneous logic networks, it is essential to identify the structural and functional characteristics of its modules, and represent them with appropriate DAG types (such as AIG, MIG, XAG, etc.) for logic optimization. This paper proposes HeLO, a hetero-DAG logic optimization framework empowered by hierarchical clustering and graph learning. HeLO leverages a hierarchical clustering algorithm, which splits the original Boolean network into sub-circuits by considering both topological and functional characteristics. A novel graph neural network model is customized to generate the topological-functional embedding (used for distance calculation in hierarchical clustering) and predict the best-fit DAG type of each sub-circuit. Experimental results demonstrate that HeLO outperforms LSOracle, the SOTA heterogeneous logic optimization framework, in terms of node-depth product (for technology-independent logic optimization) and delay-area product (for technology mapping) by 8.7% and 6.9%, respectively. Yuan Pu 0001, Fangzhou Liu 0005, Zhuolun He, Keren Zhu 0001, Rongliang Fu, Ziyi Wang 0010, Tsung-Yi Ho, Bei Yu 0001 |
ISPD | 3 |
| 2025 | Divergent Thoughts toward One Goal: LLM-based Multi-Agent Collaboration System for Electronic Design AutomationabstractHaoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Haoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
NAACL (Long Papers) | 3 |
| 2025 | On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language UnderstandingabstractLarge language models (LLMs) achieve remarkable performance in code generation tasks.
However, a significant performance disparity persists between popular programming languages (e.g., Python, C++) and others.
To address this capability gap, we leverage the code translation task to train LLMs, thereby facilitating the transfer of coding proficiency across diverse programming languages.
Moreover, we introduce OORL for training, a novel reinforcement learning (RL) framework that integrates on-policy and off-policy strategies.
Within OORL, on-policy RL is applied during code translation, guided by a rule-based reward signal derived from unit tests.
Complementing this coarse-grained rule-based reward, we propose Group Equivalent Preference Optimization (GEPO), a novel preference optimization method.
Specifically, GEPO trains the LLM using intermediate representations (IRs) groups.
LLMs can be guided to discern IRs equivalent to the source code from inequivalent ones, while also utilizing signals about the mutual equivalence between IRs within the group.
This process allows LLMs to capture nuanced aspects of code functionality.
By employing OORL for training with code translation tasks, LLMs improve their recognition of code functionality and their understanding of the relationships between code implemented in different languages.
Extensive experiments demonstrate that our OORL for LLMs training with code translation tasks achieves significant performance improvements on code benchmarks across multiple programming languages. Haoyuan Wu, Rui Ming, Jilong Gao, Hangyu Zhao, Xueyi Chen, Yikai Yang, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
NeurIPS | 8 |
| 2025 | IncreMacro: Incremental Macro Placement RefinementabstractThis article proposes$\textsf {IncreMacro}$, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this,$\textsf {IncreMacro}$utilizes several methods, including kd-tree-based macro diagnosis, gradient-based macro shifting, constraint-graph-based LP for macro legalization, and diffusion-based cell migration. By employing these techniques iteratively,$\textsf {IncreMacro}$meets two critical solution requirements of macro placement: 1) pushing macros toward the chip boundary and 2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into$\textsf {AutoDMP}$and$\textsf {DREAMPlace}~4.0$, and is evaluated on seven RISC-V benchmark circuits and four TILOS macro placement circuit designs at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by$\textsf {AutoDMP}~(\textsf {DREAMPlace}~4.0$), our approach reduces routed wirelength by 15.1% (14.9%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 99.9 (82.6%) and 99.9% (81.3%), and reduces the total power consumption by 4.4% (4.3%). Meanwhile, compared with$\textsf {IncreMacro}$[1], our approach augmented with the cell migration algorithm improves the routed WNS and TNS by 24.7% and 23.1%, and remains the average routed wirelength and total power consumption almost unchanged. Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Jiajun Qin, Haisheng Zheng, Yibo Lin, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QAabstractRetrieval augmented generation (RAG) improves the accuracy and dependability of generative AI models by integrating factual information from external databases. This technique is widely used in tasks involving document-grounded question answering (QA). While these RAG systems are extensively pretrained on general-purpose documents, they face considerable limitations when applied to specialized, knowledgeintensive fields such as electronic design automation (EDA). This paper addresses such issue by proposing a customized RAG framework along with three domain-specific techniques for EDA tool documentation QA, including a contrastive learning scheme for text embedding model fine-tuning, a reranker distilled from proprietary LLM, and a generative LLM fine-tuned with highquality domain corpus. To further unleash the extraordinary language capacity of LLMs in the domain of EDA-tool documentation QA, we propose to train LLMs as the reranker model with our customized two-stage traning scheme, which consists of the point-wise instruction tuning stage and the pairwise learn-to-rank (LTR) stage. Finally, we have developed and released a documentation QA evaluation benchmark, ORD-QA, for OpenROAD, an advanced RTL-to-GDSII design platform. Experimental results demonstrate that our proposed RAG flow and techniques have achieved superior performance on ORD-QA as well as on a commercial tool, compared with state-of-thearts. Furthermore, compared with the SOTA reranker models, our LLM reranker prominently improves the document retrieval accuracy and thus leads to better QA quality. The ORD-QA benchmark and the training dataset for our customized RAG flow are open-source at https://github.com/lesliepy99/RAG-EDA. Yuan Pu 0001, Zhuolun He, Tairu Qiu, Haoyuan Wu, Qi Sun 0002, Cheng Zhuo, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | FGNN2: A Powerful Pretraining Framework for Learning the Logic Functionality of CircuitsabstractLearning feasible representation from raw gate-level circuits is essential for incorporating machine learning techniques in logic synthesis, physical design, or verification. Existing structure-based learning methods tend to concentrate mainly on the graph topology, often neglecting logic functionality. This oversight frequently results in a failure to capture the underlying semantics, thereby limiting their overall applicability. To address the concern, we propose a novel circuit representation learning framework, FGNN2, that utilizes a contrastive scheme to effectively extract generic functionality knowledge. We construct a comprehensive pretraining dataset through a customized circuit augmentation scheme. We have also developed a novel contrastive loss function to capture the relative functional distance between different circuits, and to generate representations that are invariant to the input order. In addition, we employed a customized graph neural network (GNN) architecture to better align with the above framework. Comprehensive experiments on the multiple complex real-world designs demonstrate that our proposed solution significantly outperforms the state-of-the-art circuit representation learning flows. Ziyi Wang 0010, Zhuolun He, Guangliang Zhang, Qiang Xu 0001, Tsung-Yi Ho, Yu Huang 0005, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Large Language Models for EDA: Future or Mirage?abstractIn this article, we explore the burgeoning intersection of large language models (LLMs) and electronic design automation (EDA). We critically assess whether LLMs represent a transformative future for EDA or merely a fleeting mirage. By organizing existing research into four critical domains of EDA—code generation, verification and debugging, knowledge representation and retrieval, and optimization/modeling—we provide a comprehensive overview of the current state-of-the-art. The survey concludes with a 5-level roadmap to guide the progressive integration and advancement of LLMs in EDA. Ultimately, this article aims to provide a comprehensive, evidence-based perspective on the role of LLMs in shaping the future of EDA. Zhuolun He, Yuan Pu 0001, Haoyuan Wu, Tairu Qiu, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | EasyMRC: Efficient Mask Rule Checking via Representative Edge SamplingabstractThe photolithography process is getting more sophisticated with technology node scaling down and VLSI designs becoming complex. As photomask patterns get finer, mask rule checks (MRCs) are inevitable to avoid discrepancies in the layout and to ensure manufacturability. This paper introduces an efficient mask rule checking approach that utilizes a representative edge sampling scheme. The representative edge sampling scheme selects a subset of edges and points of each polygon that capture its contour, meanwhile greatly reducing the number of edges involved in actual checking. Experimental results demonstrate that the proposed approach achieves significant speedup compared with the state-of-the-art academic tool. Zhuolun He, Yuan Pu 0001, Wenjian Yu, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | LSTP: A Logic Synthesis Timing PredictorabstractThe ever-growing complexity of modern VLSI circuits brings about a substantial increase in the design cycle. As for logic synthesis, how to efficiently obtain physical characteristics of a design for subsequent design space exploration emerges as a critical issue. In this paper, we propose ${\mathsf{LSTP}}$, an ML-based logic synthesis predictor, which can rapidly predict the post-synthesis timing of a broad range of circuit designs. Specifically, we explicitly take optimization sequences into consideration so that we can comprehend the synergy between optimization passes and their effects on netlists. Experimental results demonstrate that we outperform state-of-the-art remarkably. Haisheng Zheng, Zhuolun He, Fangzhou Liu 0005, Zehua Pei, Bei Yu 0001 |
ASPDAC | 2 |
| 2024 | PDRC: Package Design Rule Checking via GPU-Accelerated Geometric Intersection Algorithms for Non-Manhattan GeometryabstractWith the emergence of chiplet technology, the scale of IC packaging design has been steadily increasing, making the utilization of traditional design rule checking (DRC) methods more time-consuming. In this paper, we propose PDRC, a package-level design rule checker for non-manhattan geometry with GPU acceleration. PDRC employs hierarchical interval lists within an iterative parallel sweepline framework to implement the geometric intersection algorithm, thereby finishing design rule checking tasks. Experimental results have demonstrated 30 - 50 times speedup achieved by PDRC compared with two CPU-based checkers. Jiaxi Jiang, Lancheng Zou, Wenqian Zhao 0002, Zhuolun He, Tinghuan Chen, Bei Yu 0001 |
DAC | 4 |
| 2024 | Lesyn: Placement-aware Logic Resynthesis for Non-Integer Multiple-Cell-Height DesignsabstractNon-integer multiple cell height (NIMCH) standard-cell libraries offer promising co-optimization for power, performance and area in advanced technology nodes. However, such non-uniform design introduces new layout constraints where any sub-region can only accommodate gates of the same cell height due to manufacturability concerns. The existing physical design flow for NIMCH circuits, which handles the layout constraint by clustering and relocating gates according to their cell heights, often leads to substantial gate displacement that harms circuit performance. To alleviate the above issue, this paper proposes a row-based logic resynthesis procedure that explicitly adjusts cell heights after initial placement without changing cell positions. Experiment results demonstrate that compared with the conventional NIMCH physical design flow, our proposed approach can reduce the maximal delay by 26.1%. Yuan Pu 0001, Fangzhou Liu 0005, Yu Zhang 0189, Zhuolun He, Yibo Lin, Kai-Yuan Chao, Bei Yu 0001 |
DAC | 4 |
| 2024 | CBTune: Contextual Bandit Tuning for Logic SynthesisabstractLogic synthesis pre-optimization involves applying a sequence of transformations called synthesis flow to reduce the circuit's Boolean logic graph, like AIG. However, the challenge lies in selecting and arranging these transformations due to the exponentially expanding solution space. In this work, we propose CBTune, a novel online learning framework that utilizes a contextual bandit algorithm to explore the solution space and generate synthesis flows efficiently. We develop the Syn-LinUCB algorithm as the agent, which incorporates circuit characteristics and leverages long-term payoffs to guide decision-making, thus ef-fectively preventing getting trapped in local optima. Experimental results show that our framework achieves the optimal synthesis flow with a lower time cost, substantially reducing the number of AIG nodes and 6-LUTs compared to SOTA approaches. Fangzhou Liu 0005, Zehua Pei, Ziyang Yu 0001, Haisheng Zheng, Zhuolun He, Tinghuan Chen, Bei Yu 0001 |
DATE | 5 |
| 2024 | Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General TasksabstractLarge language models (LLMs) have demonstrated considerable proficiency in general natural language processing (NLP) tasks.Instruction tuning, a successful paradigm, enhances the ability of LLMs to follow natural language instructions and exhibit robust generalization across general tasks.However, these models often encounter performance limitations across multiple tasks due to constrained model capacity.Expanding this capacity during the instruction tuning phase poses significant challenges.To address this issue, we introduce parameter-efficient sparsity crafting (PESC), which crafts dense models into sparse models using the mixture-of-experts (MoE) architecture.PESC integrates adapters into the MoE layers of sparse models, differentiating experts without altering the individual weights within these layers.This method significantly reduces computational costs and GPU memory requirements, facilitating model capacity expansion through a minimal parameter increase when guaranteeing the quality of approximation in function space compared to original sparse upcycling.Our empirical evaluation demonstrates the effectiveness of the PESC method.Using PESC during instruction tuning, our best sparse model outperforms other sparse and dense models and exhibits superior general capabilities compared to GPT-3.5.Our code is available at https://github.com/wuhy68/ Parameter-Efficient-MoE. Haoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu 0001 |
EMNLP | 3 |
| 2024 | Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QAabstractRetrieval augmented generation (RAG) enhances the accuracy and reliability of generative AI models by sourcing factual information from external databases, which is extensively employed in document-grounded question-answering (QA) tasks. Off-the-shelf RAG flows are well pretrained on general-purpose documents, yet they encounter significant challenges when being applied to knowledge-intensive vertical domains, such as electronic design automation (EDA). This paper addresses such issue by proposing a customized RAG framework along with three domain-specific techniques for EDA tool documentation QA, including a contrastive learning scheme for text embedding model fine-tuning, a reranker distilled from proprietary LLM, and a generative LLM fine-tuned with high-quality domain corpus. Furthermore, we have developed and released a documentation QA evaluation benchmark, ORD-QA, for OpenROAD, an advanced RTL-to-GDSII design platform. Experimental results demonstrate that our proposed RAG flow and techniques have achieved superior performance on ORD-QA as well as on a commercial tool, compared with state-of-the-arts. The ORD-QA benchmark and the training dataset for our customized RAG flow are open-source at https://github.com/lesliepy99/RAG-EDA. Yuan Pu 0001, Zhuolun He, Tairu Qiu, Haoyuan Wu, Bei Yu 0001 |
ICCAD | 2 |
| 2024 | Routing-aware Legal Hybrid Bonding Terminal Assignment for 3D Face-to-Face Stacked ICsabstractFace-to-face (F2F) stacked 3D IC is a promising alternative for scaling beyond Moore's Law. In F2F 3D ICs, dies are connected through bonding terminals whose positions can significantly impact routing performance. Further, there exists resource competition among all the 3D nets due to the constrained bonding terminal number. In advanced technology nodes, such 3D integration may also introduce legality challenges of bonding terminals, as the metal pitches can be much smaller than the sizes of bonding terminals. Previous works attempt to insert bonding terminals automatically using existing 2D commercial P&R tools and then consider inter-die connection legality, but they fail to take the legality and routing performance into account simultaneously. In this paper, we explore the formulation of the generalized assignment in the hybrid bonding terminal assignment problem. Our framework, BTAssign, offers a strict legality guarantee and an iterative solution. The experiments are conducted on 18 open-source designs with various 3D net densities and the most advanced bonding scale. The results reveal that BTAssign can achieve improvements in routed wirelength under all testing conditions from 1.0% to 5.0% with a tolerable runtime overhead. Siting Liu 0002, Jiaxi Jiang, Zhuolun He, Ziyi Wang 0010, Yibo Lin, Bei Yu 0001, Martin D. F. Wong |
ISPD | 3 |
| 2024 | Large Language Models for EDA: Future or Mirage?abstractIn this paper, we explore the burgeoning intersection of Large Language Models (LLMs) and Electronic Design Automation (EDA). We critically assess whether LLMs represent a transformative future for EDA or merely a fleeting mirage. By analyzing current advancements, challenges, and potential applications, we dissect how LLMs can revolutionize EDA processes like design, verification, and optimization. Furthermore, we contemplate the ethical implications and feasibility of integrating these models into EDA workflows. Ultimately, this paper aims to provide a comprehensive, evidence-based perspective on the role of LLMs in shaping the future of EDA. Zhuolun He, Bei Yu 0001 |
ISPD | 1 |
| 2024 | IncreMacro: Incremental Macro Placement RefinementabstractThis paper proposes IncreMacro, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this, IncreMacro utilizes several methods including kd-tree-based macro diagnosis, gradient-based macro shifting and constraint-graph-based LP for macro legalization. By employing these techniques iteratively, IncreMacro meets two critical solution requirements of macro placement: (1) pushing macros to the chip boundary; and (2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into DREAMPlace and AutoDMP, and is evaluated on several RISC-V benchmark circuits at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by DREAMPlace (AutoDMP), IncreMacro reduces routed wirelength by 6.5% (16.8%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 59.9% (99.6%) and 63.9% (99.9%), and reduces the total power consumption by 3.3% (4.9%). Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Haisheng Zheng, Yibo Lin, Bei Yu 0001 |
ISPD | 3 |
| 2024 | ChatEDA: A Large Language Model Powered Autonomous Agent for EDAabstractThe integration of a complex set of Electronic Design Automation (EDA) tools to enhance interoperability is a critical concern for circuit designers. Recent advancements in large language models (LLMs) have showcased their exceptional capabilities in natural language processing and comprehension, offering a novel approach to interfacing with EDA tools. This research paper introduces ChatEDA, an autonomous agent for EDA empowered by a large language model, AutoMage, complemented by EDA tools serving as executors. ChatEDA streamlines the design flow from the Register-Transfer Level (RTL) to the Graphic Data System Version II (GDSII) by effectively managing task decomposition, script generation, and task execution. Through comprehensive experimental evaluations, ChatEDA has demonstrated its proficiency in handling diverse requirements, and our fine-tuned AutoMage model has exhibited superior performance compared to GPT-4 and other similar LLMs. Haoyuan Wu, Zhuolun He, Xinyun Zhang 0001, Xufeng Yao, Su Zheng, Haisheng Zheng, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel ExecutionabstractDeep learning frameworks optimize the computation graphs and intra-operator computations to boost the inference performance on GPUs, while inter-operator parallelism is usually ignored. In this paper, a unified framework, AutoGraph, is proposed to obtain highly optimized computation graphs in favor of parallel executions of GPU kernels. A novel dynamic programming algorithm, combined with backtracking search, is adopted to explore the optimal graph optimization solution, with the fast performance estimation from the mixed critical path cost. Accurate runtime information based on GPU Multi-Stream launched with CUDA Graph is utilized to determine the convergence of the optimization. Experimental results demonstrate that our method achieves up to 3.47x speedup over existing graph optimization methods. Moreover, AutoGraph outperforms state-of-the-art parallel kernel launch frameworks by up to 1.26x. Yuxuan Zhao 0001, Qi Sun 0002, Zhuolun He, Bei Yu 0001 |
AAAI | 3 |
| 2023 | OpenDRC: An Efficient Open-Source Design Rule Checking Engine with Hierarchical GPU AccelerationabstractDesign rule checking (DRC) is an essential procedure in physical verification, yet few open-source DRC tools are accessible in academia. To fill in the gap, we present OpenDRC, an open-source DRC engine that aims for extremely high efficiency. OpenDRC maintains hierarchical layouts with layer-wise bounding volume hierarchies and performs adaptive row-based partition to identify independent regions for check pruning and/or parallel processing. For common design rules, OpenDRC provides a sequential mode that runs cell-level sweeplines, and a parallel mode that launches edge-based GPU check kernels. Experiments demonstrate that OpenDRC outperforms state-of-the-art multi-threading and GPU-accelerated design rule checkers. The source code is available at https://github.com/opendrc/opendrc. Zhuolun He, Yihang Zuo, Jiaxi Jiang, Haisheng Zheng, Yuzhe Ma, Bei Yu 0001 |
DAC | 1 |
| 2023 | Efficient Design Rule Checking with GPU AccelerationabstractDesign rule checking (DRC) is an essential part of the chip design flow, which ensures that manufacturing requirements are conformed to avoid a chip failure. With the rapid increase of design scales, DRC has been suffering from runtime overhead. To overcome this challenge, we propose to accelerate DRC algorithms by harnessing the power of graphics processing units (GPUs). Specifically, we first explore an efficient data transfer approach for geometry information of a layout. Then we investigate GPU-based scanline algorithms to accommodate both intra-polygon checking and intre-polygon checking based on the characteristics of the design rules. Experimental results show that the proposed GPU-accelerated method can substantially outperform a multi-threaded DRC algorithm using CPU. Compared with the baseline with 24 threads, we can achieve an average speedup of 36 × and 201 × for spacing rule checks and enclosing rule checks on a metal layer, respectively. Zhuolun He, Yuzhe Ma, Bei Yu 0001 |
DATE | 3 |
| 2023 | Invited Paper: Heterogeneous Acceleration for Design Rule CheckingabstractThe advances of heterogeneous CPU-GPU computing platforms have marked their great potential for algorithm acceleration. Yet, how to orchestrate such hybrid devices remains a concern for programmers and researchers. To obtain a desired performance gain, different strategies, such as parallel computing, heterogeneous scheduling, and data movement minimization, should be carefully considered and effectively combined. In this talk, we first review efforts for efficient design rule checking in the literature. Then, we introduce the parallel sweepline paradigm for design rule checking, and demonstrate how to accelerate design rule checking with such paradigm on heterogeneous CPU-GPU platforms. Zhuolun He, Bei Yu 0001 |
ICCAD | 1 |
| 2023 | AlphaSyn: Logic Synthesis Optimization with Efficient Monte Carlo Tree SearchabstractRecent years have seen rising research in logic synthesis recipe generation to improve the Quality-of-Result (QoR). However, existing approaches typically have low efficiency and are stuck at local optima. In this work, we propose a logic synthesis optimization framework, AlphaSyn, that incorporates a domain-specific Monte Carlo tree search (MCTS) algorithm. AlphaSyn enables exploration across the entire search space while optimizing sampling points utilization. We further develop a synthesis-specific upper confidence bound for trees (SynUCT) algorithm for the selection phase and a well-designed learning strategy to enhance the stability of the MCTS algorithm. The AlphaSyn algorithm is fully parallelized for efficiency with asynchronous MCTS exploration and significance-base resource allocation. For standard-cell technology mapping on the ASAP 7nm library among other tasks, experimental results show that AlphaSyn outperforms SOTA FlowTune with an average 8.74% area reduction and$\boldsymbol{1.24}\times$runtime speedup. Zehua Pei, Fangzhou Liu 0005, Zhuolun He, Guojin Chen, Haisheng Zheng, Keren Zhu 0001, Bei Yu 0001 |
ICCAD | 3 |
| 2023 | Efficient Super-Resolution System With Block-Wise Hybridization and Quantized Winograd on FPGAabstractSuper-resolution (SR) techniques aim to restore a high-resolution (HR) image from low-resolution (LR) images, which are often used to assist the enhancement of image/video quality under the rapid development of HR and high-frame-rate media. Recently, neural network (NN)-based methods perform much better image reconstruction quality than classical approaches. However, the unacceptable computation complexity as well as the huge memory footprints of NNs limit the throughputs and scalability of these SR systems. In this work, we analyze several key issues in the design of NN-based SR systems first. Then, we propose a three-level systematic optimization methodology for SR systems to reduce computation overhead and keep image quality. At the algorithm level, we introduce image blocking to SR tasks and develop a block-wise SR algorithm based on the hybrid of NN and interpolation with a consistent image block evaluation metric. The configurable hybrid parameters help the SR algorithm to achieve a flexible tradeoff between the computation overhead and image quality. At the operator level, we focus on the transpose convolution operators commonly used for upsampling in SR NNs. We propose an efficient Winograd-based transposed convolution acceleration method. Through the efficient subconvolutions conversion and the Winograd specialization, this methods enables unified Winograd transformations and simplified data access patterns. At the data level, we propose a novel quantization method for Winograd-aware SR NNs to get better-quantized accuracy. Comprehensive evaluations demonstrate the effectiveness of these optimizations. Our SR system reduces a large number of multiplications with great scalability and supports 4K@120 fps and 8K@30 fps outputs with acceptable image quality degradation. Bizhao Shi, Jiaxi Zhang 0001, Zhuolun He, Xuechao Wei, Sicheng Li 0001, Guojie Luo, Hongzhong Zheng, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Efficient Arithmetic Block Identification With Graph Learning and Network-FlowabstractArithmetic block identification in gate-level netlists plays an essential role for various purposes, including malicious logic detection, functional verification, or macro-block optimization. However, current methods usually suffer from either low performance or poor scalability. To address the issue, we come up with a novel framework based on graph learning and network flow analysis, that extracts desired logic components from a complete circuit netlist. We design a novel asynchronous bidirectional graph neural network (ABGNN) dedicated to representation learning on directed acyclic graphs. In addition, we develop a convex cost network-flow-based datapath extraction approach to match the predicted block inputs with predicted block outputs. Experimental results on open-source RISC-V CPU designs demonstrate that our proposed solution significantly outperforms several state-of-the-art arithmetic block identification flows. Ziyi Wang 0010, Zhuolun He, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Functionality matters in netlist representation learningabstractLearning feasible representation from raw gate-level netlists is essential for incorporating machine learning techniques in logic synthesis, physical design, or verification. Existing message-passing-based graph learning methodologies focus merely on graph topology while overlooking gate functionality, which often fails to capture underlying semantic, thus limiting their generalizability. To address the concern, we propose a novel netlist representation learning framework that utilizes a contrastive scheme to acquire generic functional knowledge from netlists effectively. We also propose a customized graph neural network (GNN) architecture that learns a set of independent aggregators to better cooperate with the above framework. Comprehensive experiments on multiple complex real-world designs demonstrate that our proposed solution significantly outperforms state-of-the-art netlist feature learning flows. Ziyi Wang 0010, Zhuolun He, Guangliang Zhang, Qiang Xu 0001, Tsung-Yi Ho, Bei Yu 0001, Yu Huang 0005 |
DAC | 3 |
| 2022 | X-Check: CPU-Accelerated Design Rule Checking via Parallel Sweepline AlgorithmsabstractDesign rule checking (DRC) is essential in physical verification to ensure high yield and reliability for VLSI circuit designs. To achieve reasonable design cycle time, acceleration for computationally intensive DRC tasks has been demanded to accommodate the ever-growing complexity of modern VLSI circuits. In this paper, we propose X-Check, a GPU-accelerated design rule checker. X-Check integrates novel parallel sweepline algorithms, which are both efficient in practice and with nontrivial theoretical guarantees. Experimental results have demonstrated significant speedup achieved by X-Check compared with a multi-threaded CPU checker. Zhuolun He, Yuzhe Ma, Bei Yu 0001 |
ICCAD | 1 |
| 2021 | Physical Synthesis for Advanced Neural Network ProcessorsabstractThe remarkable breakthroughs in deep learning have led to a dramatic thirst for computational resources to tackle interesting real-world problems. Various neural network processors have been proposed for the purpose, yet, far fewer discussions have been made on the physical synthesis for such specialized processors, especially in advanced technology nodes. In this paper, we review several physical synthesis techniques for advanced neural network processors. We especially argue that datapath design is an essential methodology in the above procedures due to the organized computational graph of neural networks. As a case study, we investigate a wafer-scale deep learning accelerator placement problem in detail. Zhuolun He, Peiyu Liao, Siting Liu 0002, Yuzhe Ma, Yibo Lin, Bei Yu 0001 |
ASP-DAC | 1 |
| 2021 | Graph Learning-Based Arithmetic Block IdentificationabstractArithmetic block identification in gate-level netlist is an essential procedure for malicious logic detection, functional verification, or macro-block optimization. We argue that existing methods suffer either scalability or performance issues. To address the problem, we propose a graph learning-based solution that promises to extract desired logic components from a complete design netlist. We further design a novel asynchronous bidirectional graph neural network (ABGNN) dedicated to representation learning on directed acyclic graphs. Experimental results on open-source RISC-V CPU designs demonstrate that our proposed solution significantly outperforms several state-of-the-art arithmetic block identification flows. Zhuolun He, Ziyi Wang 0010, Bei Yu 0001 |
ICCAD | 1 |
| 2020 | Learn to Floorplan through Acquisition of Effective Local Search HeuristicsabstractAutomatic heuristic design through reinforcement learning opens a promising direction for solving computationally difficult problems. Unlike most previous works that aimed at solution construction, we explore the possibility of acquiring local search heuristics through massive search experiments. To illustrate the applicability, an agent is trained to perform a walk in the search space by selecting a candidate neighbor solution at each step. Specifically, we target the floorplanning problem, where a neighbor solution is generated through perturbing the sequence pair encoding of a floorplan. Experimental results demonstrate the efficacy of the acquired heuristics as well as the potential of automatic heuristic design. Zhuolun He, Yuzhe Ma, Peiyu Liao, Ngai Wong 0001, Bei Yu 0001, Martin D. F. Wong |
ICCD | 1 |
| 2020 | Understanding Graphs in EDA: From Shallow to Deep LearningabstractAs the scale of integrated circuits keeps increasing, it is witnessed that there is a surge in the research of electronic design automation (EDA) to make the technology node scaling happen. Graph is of great significance in the technology evolution since it is one of the most natural ways of abstraction to many fundamental objects in EDA problems like netlist and layout, and hence many EDA problems are essentially graph problems. Traditional approaches for solving these problems are mostly based on analytical solutions or heuristic algorithms, which require substantial efforts in designing and tuning. With the emergence of the learning techniques, dealing with graph problems with machine learning or deep learning has become a potential way to further improve the quality of solutions. In this paper, we discuss a set of key techniques for conducting machine learning on graphs. Particularly, a few challenges in applying graph learning to EDA applications are highlighted. Furthermore, two case studies are presented to demonstrate the potential of graph learning on EDA applications. Yuzhe Ma, Zhuolun He, Wei Li 0159, Bei Yu 0001 |
ISPD | 2 |
| 2020 | Deep Model Compression and Inference Speedup of Sum-Product Networks on Tensor TrainsabstractSum-product networks (SPNs) constitute an emerging class of neural networks with clear probabilistic semantics and superior inference speed over other graphical models. This brief reveals an important connection between SPNs and tensor trains (TTs), leading to a new canonical form which we call tensor SPNs (tSPNs). Specifically, we demonstrate the intimate relationship between a valid SPN and a TT. For the first time, through mapping an SPN onto a tSPN and employing specially customized optimization techniques, we demonstrate improvements up to a factor of 100 on both model compression and inference speedup for various data sets with negligible loss in accuracy. Ching Yun Ko, Cong Chen 0003, Zhuolun He, Kim Batselier, Ngai Wong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | FPGA-Based Real-Time Super-Resolution System for Ultra High Definition VideosabstractThe market benefits from a barrage of Ultra High Definition (Ultra-HD) displays, yet most extant cameras are barely equipped with Full-HD video capturing. In order to upgrade existing videos without extra storage costs, we propose an FPGA-based super-resolution system that enables real-time Ultra-HD upscaling in high quality. Our super-resolution system crops each frame into blocks, measures their total variation values, and dispatches them accordingly to a neural network or an interpolation module for upscaling. This approach balances the FPGA resource utilization, the attainable frame rate, and the image quality. Evaluations demonstrate that the proposed system achieves superior performance in both throughput and reconstruction quality, comparing to current approaches. Zhuolun He, Hanxian Huang, Ming Jiang 0001, Yuanchao Bai, Guojie Luo |
FCCM | 1 |
| 2017 | FPGA Acceleration for Computational Glass-Free Displays
Zhuolun He, Guojie Luo |
FPGA | 1 |