Tinghuan Chen

dblp:175/6569 · DBLP profile ↗
← Back
61ranked-venue papers
7as first author
57since 2021 · last 2026
0000-0002-9195-6619ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 54 · 6 first-author · 52 since 2021Software engineering, systems software and programming languages · 10 · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FD-MAGRPO: Functionality-Driven Multi-Agent Group Relative Policy Optimization for Analog-LDO Sizing
abstract
This paper introduces the Functionality-Driven Multi-Agent Group Relative Policy Optimization (FD-MAGRPO) algorithm, which is designed to enhance exploration efficiency in reinforcement learning (RL) for analog integrated circuit sizing. Our proposed method integrates two key innovations: (1) a critic-free multi-agent optimization framework based on Group Relative Policy Optimization (GRPO), that eliminates the critic network and achieves stable and efficient policy updates; and (2) a functionality-driven grouping strategy, that enables agents to coordinate exploration by functional roles instead of circuit blocks, thereby improving credit assignment and cooperation. Experimental results on practical low-dropout regulator (LDO) circuits with 65–179 design parameters show that the proposed method achieves rapid convergence with only 800–3000 simulations, yielding a 4.8×–13.0× speedup over state-of-the-art methods. Mathematical analysis and empirical studies validate that the combination of critic-free optimization and functionality-based grouping leads to higher exploration efficiency and faster convergence. The proposed method enables the discovery of higher circuit performances that are inaccessible to conventional approaches, establishing FD-MAGRPO as a robust and efficient solution for complex analog-LDO sizing tasks.
Haoning Jiang, Zhuoli Ouyang, Tinghuan Chen, Junmin Jiang
AAAI5
2026 KCLNet: Electrically Equivalence-Oriented Graph Representation Learning for Analog Circuits
abstract
Digital circuit representation learning has made remarkable progress in electronic design automation, effectively supporting critical tasks such as testability analysis and logic reasoning. However, representation learning for analog circuits remains challenging due to their continuous electrical characteristics compared to the discrete states of digital circuits. This paper presents a direct current (DC) electrically equivalent-oriented analog representation learning framework, named KCLNet. We will open-source the dataset and code upon publication. It comprises an asynchronous graph neural network structure with electrically-simulated message passing and a representation learning method inspired by Kirchhoff's Current Law (KCL). This method maintains the orderliness of the circuit embedding space by enforcing the equality of the sum of outgoing and incoming current embeddings at each node, which significantly enhances the generalization ability of circuit embeddings. KCLNet offers a novel and effective solution for analog circuit representation learning with electrical constraints preserved. Experimental results demonstrate that our method achieves significant performance in a variety of downstream tasks, e.g., analog circuit classification, subcircuit detection, and circuit edit distance prediction.
Peng Xu 0052, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
AAAI3
2026 Synergistic Bayesian Optimization and Reinforcement Learning with Bidirectional Interaction for Efficient VLSI Constraint Tuning
Jiayi Tu 0001, Jindong Tu, Meng Zhang 0010, Tinghuan Chen
ASP-DAC6
2026 IP-Matcher: An Efficient One-to-Many Matching Framework for Analog Circuit Design and Reusing
abstract
The design efficiency of analog circuits is generally lower than that of digital circuits, presenting a significant bottleneck in the current integrated circuit industry. One promising method to accelerate design processes is the modular design philosophy adapted from digital methodologies. However, there is a lack of an efficient framework for reusing mature analog circuit topologies and the corresponding layout designs. To achieve a rapid design iteration while utilizing specialized expertise in design, we propose IP-Matcher, an efficient IP-based analog circuit matching and reusing framework. The framework consists of three components: Analog Graph Converter, Analog IP Manager, and IP-based Matcher, which collaborate to enhance both matching accuracy and speed, thereby improving analog IP reusability. We leverage the unique characteristics of analog circuits to significantly prune the matching space, overcoming the limitations of traditional circuit matching strategies. Experimental results show that our work not only outperforms the state-of-the-art method by 32% in accuracy but also achieves a 16× speedup.
Shixin Chen, Peng Xu 0052, Tinghuan Chen, Bei Yu 0001
DATE4
2026 Smart-PCLib: A LLM-based Multi-Agent Framework for Automated PCB Component Library Generation
Zhaohai Di, Jindong Tu, Yuan Pu 0001, Jiawei Liu 0006, Chong Tong, Tsung-Yi Ho, Bei Yu 0001, Tinghuan Chen
DATE9
2026 PCB-Migrator: Automated PCB PnR Migration
abstract
Despite the availability of numerous frameworks and tools for automated PCB placement and routing, the industry still relies heavily on expert designers to ensure layout reliability and performance. However, when design requirements change, such as adjustments to board dimensions or the addition of new obstacles, experts must often recreate similar layouts from scratch, leading to substantial inefficiencies in both time and resources. To address this challenge, we introduce PCB-Migrator, an automated framework for PCB layout migration. Our approach leverages an offset constraint graph to capture positional relationships among components in the referenced design and effectively map them onto the new PCB. Additionally, PCB-Migrator builds routing path graphs to extract routing characteristics from the reference layout and applies graph matching to guide the routing process on the new board. Experimental results demonstrate that PCB-Migrator outperforms existing baselines, achieving faster runtimes while preserving the key design characteristics and performance of the referenced PCB.
Yaohui Han, Beichen Li 0003, Rongliang Fu, Qunsong Ye, Bei Yu 0001, Tsung-Yi Ho, Tinghuan Chen
DATE9
2026 EDA Flow Matters: Stage-Aware Parameter Optimization of Tool Chain
abstract
Optimizing Electronic Design Automation (EDA) tool parameters with only dozens of affordable evaluations represents one of the most challenging problems in today’s EDA flow management, where each experiment costs hours to days yet directly impacts final PPA outcomes. While Bayesian Optimization (BO) naturally fits such sample-constrained scenarios, it models the entire EDA flow as a monolithic formulation, blindly ignoring the sequential structure that each stage in the EDA flow affects the next. In this work, we propose a stage-aware optimization framework that fundamentally rethinks EDA parameter tuning. The proposed stage-aware Gaussian process explicitly models cascading relationships between EDA stages through interconnected GP layers, extracting abundant information from each expensive evaluation. To better meet realistic needs, we further introduce Expected Hypervolume Improvement (EHVI)-Efficiency, a time-aware acquisition function that exploits evaluation runtime estimation and EDA tools’ checkpoint reuse to balance design metrics’ expected improvement against EDA flow’s computational cost. Experiments and ablation studies on 6 designs across 3 process nodes demonstrate the effectiveness of our proposed method.
Xinheng Li, Donger Luo, Peng Xu 0052, Ziyang Yu 0001, Qi Sun 0002, Tinghuan Chen, Bei Yu 0001, Hao Geng
DATE6
2026 A 8-19-GHz Current-Reused Wideband LNA With Dual Transformer Feedback
Zushuai Xie, Haojie Gong, Tinghuan Chen, Zhikuang Cai, Zixuan Wang 0022
ISCAS4
2026 A Compact Wide-Band RF Energy Harvesting Front-End Based on an On-Chip IMN with 51.78% Peak PCE
Zushuai Xie, Tinghuan Chen, Zhikuang Cai, Zixuan Wang 0022
ISCAS4
2026 Attention-Based EDA Tool Parameter Explorer: From Hybrid Parameters to Multi-QoR Metrics
Donger Luo, Qi Sun 0002, Peng Xu 0052, Su Zheng, Qi Xu 0004, Tinghuan Chen, Bei Yu 0001, Hao Geng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 Ckt2Vec: Efficient Electrical Encoding for Analog Circuit Representations in Vector Space
abstract
Representation learning for analog circuits is challenging due to the continuous electrical characteristics of devices, compared to the discrete states of digital circuits. While graph neural networks (GNNs) show promise in analog circuit tasks, existing methods neglect the intrinsic electrical properties governing device-specific behaviors. Traditional device feature encoding methods present limitations: one-hot encoding is space-consuming and fails to effectively characterize inter-device similarities, while text encoding introduces erroneous estimation. We propose Ckt2Vec, a novel framework that integrates electrical characteristics into analog circuit representation learning. By encoding frequency-domain embeddings of current-voltage (I-V) curves via a spectral extractor, Ckt2Vec compresses nonlinear device-specific behaviors into low-dimensional embeddings while preserving physical fidelity. A graph-based contrastive learning approach further generates hierarchical circuit representations, capturing both block- and system-level interactions. Evaluated on three downstream tasks, including circuit classification, subcircuit detection, and circuit edit distance prediction, Ckt2Vec outperforms traditional one-hot and text-based encoding methods with less space consumption and better capability in capturing analog behavior.
Peng Xu 0052, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 A Survey of the First TinyML@ICCAD Contest for Ventricular Arrhythmia Detection by Artificial Intelligence on Low-power Microprocessor
abstract
Artificial intelligence has achieved remarkable success in various real-world applications. However, the challenge lies in its implementation on hardware platforms with constrained resources and low power while maintaining real-time capabilities. Edge artificial intelligence, in particular, stands as a pivotal field for the practical deployment of AI. The 41st IEEE/ACM International Conference on Computer-Aided Design introduced the inaugural TinyML Design Contest in 2022. The contest entailed a rigorous, multi-month research and development competition, focusing on the creation of real-time detection algorithms for life-threatening ventricular arrhythmia. These algorithms were required to be deployable on the low-power microprocessor NUCLEO-L432KC. Open to multi-person teams worldwide, the contest garnered 150 teams participation teams from 50+ organizations, with 41 teams successfully completing the challenge. Our SEUer team secured the second place. This article provides a detailed exposition of the contest, offering insights into its structure and objectives. Furthermore, it analyzes and discusses the methods developed by some of the entries as well as representative results. Finally, the article concludes with directions for future improvements.
Meng Zhang 0010, Tinghuan Chen, Jun Yang 0006
ACM Trans. Embed. Comput. Syst.4
2026 PAPlace: Performance-Driven Differentiable Analog Placement
abstract
Analog circuit placement is crucial for optimal performance, but achieving a decent layout demands expertise and time. Recent advances in machine learning techniques have shown promising results in modeling analog layout performance. PAPlace further extends these methods and integrates them into the core analog placement engine, allowing direct optimization of the post-layout performance effectively. Our approach proposes a differentiable prediction model that combines layout and wiring information into a non-linear analog placement engine. We then incorporate the differentiable performance model into a gradient-descent-based global placement engine. A multi-objective optimization method is further proposed to find the common gradient descent direction for different metrics. The experimental results on benchmarks under the TSMC 40nm technology node demonstrate the superiority of the proposed framework compared with the cutting-edge works, with up to 2163.00μ V , 73.95dB, 62.25MHz, 57.84dB improvement in Offset Voltage, CMRR, BandWidth, DC Gain metrics.
Peng Xu 0052, Yuan Pu 0001, Keren Zhu 0001, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.4
2025 SMART-GPO: Gate-Level Sensitivity Measurement with Accurate Estimation for Glitch Power Optimization
abstract
Dynamic power consumption is a significant concern in modern integrated circuits. This issue is primarily caused by signal toggling, including unwanted toggles known as glitches. With the number of operations increasing in circuits, glitches can lead to significant additional dynamic power. This paper presents SMART-GPO, a novel framework that efficiently and accurately estimates and reduces glitch power. Our approach samples cycles for accurate glitch estimation, followed by gate-sizing and Vth assignment to optimize glitch power based on sensitivity measurements. We validated SMART-GPO on the Berkeley Out-of-Order Machine (BOOM) and Rocket SoCs with TSMC N28 technology. It achieves a mean absolute percentage error (MAPE) of 2% on glitch power estimation when running power analysis on only 1% simulation cycles. The optimization results demonstrate that our framework reduces glitch power by more than 9%, which outperforms previous approaches substantially.
Yikang Ouyang, Yuchao Wu, Dongsheng Zuo, Subhendu Roy, Tinghuan Chen, Zhiyao Xie, Yuzhe Ma
ASP-DAC5
2025 H3Match: A Hybrid Heterogeneous Hypergraph Matching Method for Subcircuit Identification
abstract
Subcircuit matching is widely applied in logic synthesis, design verification, hardware security, etc. Previous works employ redundant circuit representations, coupled with timeconsuming enumerative search methods. Subsequent works use a hybrid “approximate filtering - exact verification” framework, but the numerous false negatives predicted by the graph neural network (GNN) based filtering lead to severe matching failure. In this paper, an improved hybrid method named H3Match is proposed to achieve a better tradeoff between runtime, accuracy, and false negative rate. First, we model the circuits as hypergraphs to fully capture the topology and construct diverse heterogeneous hyperedge features to facilitate the learning of circuit topologies. Second, to reduce the false negatives, we reformulate the subgraph matching problem as matching directed acyclic graphs (DAGs) with embedded circular structure information and develop a directed GNN-based approximate matching approach to identify potential matching subcircuits. Finally, we propose a general mixed integer nonlinear programming (MINLP) formulation for exact verification, with convergency speed accelerated by extracting the initial solution from the results in approximate matching. Experimental results show that our approximate method outperforms state-of-the-art (SOTA) methods by 4.31% in accuracy while achieving virtually zero false negatives. Our exact verification is on average $4.16 \times$ faster than SOTA exact methods. Overall, the end-to-end flow achieves a $7.08 \times$ speedup compared to existing approaches.
Qingsong Peng, Tianming Ni, Tinghuan Chen, Qi Sun 0002, Cheng Zhuo
DAC5
2025 LMM-IR: Large-Scale Netlist-Aware Multimodal Framework for Static IR-Drop Prediction
abstract
Static IR drop analysis is a fundamental and critical task in the field of chip design. Nevertheless, this process can be quite time-consuming, potentially requiring several hours. Moreover, addressing IR drop violations frequently demands iterative analysis, thereby causing the computational burden. Therefore, fast and accurate IR drop prediction is vital for reducing the overall time invested in chip design. In this paper, we firstly propose a novel multimodal approach that efficiently processes SPICE files through large-scale netlist transformer (LNT). Our key innovation is representing and processing netlist topology as 3D point cloud representations, enabling efficient handling of netlist with up to hundreds of thousands to millions nodes. All types of data, including netlist files and image data, are encoded into latent space as features and fed into the model for static voltage drop prediction. This enables the integration of data from multiple modalities for complementary predictions. Experimental results demonstrate that our proposed algorithm can achieve the best F1 score and the lowest MAE among the winning teams of the ICCAD 2023 contest and the state-of-theart algorithms.
Zhen Wang 0030, Hongquan He, Qi Xu 0004, Tinghuan Chen, Hao Geng
DAC5
2025 Efficient Continuous Logic Optimization with Diffusion Model
abstract
The logic synthesis optimization flow is crucial to the quality of results (QoR), which applies a sequence of transformations to a design. Recently, there has been a growing focus on the automatic optimization of synthesis flows to improve QoR, utilizing techniques such as Bayesian optimization and reinforcement learning, which may fall short in efficiency due to the exponentially large search space. In contrast, continuous optimization offers notable efficiency advantages by leveraging the explicit gradient. However, despite its potential, several significant concerns remain to be addressed. On one hand, it is essential to obtain a reliable gradient. On the other hand, a major challenge arises from the fact that searching within a continuous space can yield solutions that deviate from feasible ones. In this paper, we propose an efficient approach to optimize synthesis sequences within a continuous latent space. Specifically, the gradient information is derived from a QoR surrogate model, while the discrepancies between solutions and feasible transformations are minimized by a diffusion model. Experimental results on extensive benchmarks demonstrate that the proposed method not only achieves lower area and delay but also improves efficiency by 5 X to 130 X, compared with previous methods.
Yikang Ouyang, Jiadong Zhu, Tinghuan Chen, Yuzhe Ma
DAC4
2025 DSPlacer: DSP Placement for FPGA-based CNN Accelerator
abstract
Deploying convolutional neural networks (CNNs) on hardware platforms like Field Programmable Gate Arrays (FPGAs) has garnered significant attention due to their inherent flexibility and parallelism. Achieving optimal timing closure remains a critical challenge, as placement directly impacts clock frequency and throughput. Existing approaches often face scalability issues with large designs or fail to formalize placement rules into automated algorithms. In this paper, we propose DSPlacer, a novel DSP placement framework designed for diverse CNN accelerator architectures in the context of FPGA design. The proposed approach iteratively optimizes the placement of datapath DSPs to enhance timing performance. To achieve this, DSPlacer integrates several advanced techniques, including graph convolutional network-based datapath DSP identification, DSP graph construction, min-cost-flow DSP assignment, and integer linear programming (ILP)-based cascade constraint legalization. These techniques collectively address two key requirements for datapath DSP placement: (1) cascading datapath DSPs to achieve a compact layout, and (2) preserving direct datapath information between the processing system and programmable logic. The framework has been evaluated on multiple academic benchmarks and compared against AMD Xilinx Vivado 2020.2 and AMF-Placer 2.0. Experimental results demonstrate that DSPlacer improves Worst Negative Slack (WNS) by 32% and 65%, respectively, highlighting its efficacy and superiority.
Baohui Xie, Xinrui Zhu, Yuan Pu 0001, Tongkai Wu, Xiaofeng Zou, Bei Yu 0001, Tinghuan Chen
DAC8
2025 Rank-based Multi-objective Approximate Logic Synthesis via Monte Carlo Tree Search
abstract
Approximate Logic Synthesis (ALS) is an automated technique designed for error-tolerant applications, optimizing delay, area, and power under specified error constraints. However, existing methods typically focus on either delay reduction or area minimization, often leading to local optima in multi-objective optimization. This paper proposes a rankbased multi-objective ALS framework using Monte Carlo Tree Search (MCTS). It develops non-dominated circuit ranking, to guide MCTS in exploring local approximate changes (LACs) across the entire circuit and generate approximate circuit sets with great optimization potential. Additionally, a Rank-Transformer model is introduced to predict pathdomain ranks, enhancing the application of high-quality LACs within circuit paths. Experimental results show that our framework achieves faster and more efficient optimization in delay and area simultaneously compared to state-of-the-art methods.
Yuyang Ye 0001, Xiangfei Hu, Peng Xu 0052, Yu Gong 0002, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
DAC6
2025 Truly Pre-Routing Timing Prediction via Considering Power Delivery Network
abstract
Fast and accurate pre-routing timing prediction is essential in the chip design flow. However, existing machine learning (ML)assisted pre-routing timing methods often overlook the impact of power delivery networks (PDNs), which contribute to IR drop and routing congestion. This limitation can make these methods less practical for realworld circuit design flows. To address this, we propose two specialized encoders-an IR drop-aware encoder and a routing congestion-aware encoder-that effectively capture PDN effects through multimodal fusion of netlist, layout, and PDN data. To mitigate the challenges of imbalanced multimodal fusion, we further develop a Pareto optimization approach to ensure balanced utilization of all modalities, enhancing timing prediction accuracy. Comprehensive experiments on large-scale open-source designs using TSMC’s 16 nm technology node validate the superiority of our model over state-of-the-art pre-routing timing prediction methods.
Yuyang Ye 0001, Mingwei He, Lizheng Ren, Jianwang Zhai, Tinghuan Chen, Jun Yang 0006, Longxing Shi
DAC5
2025 Timing-Driven Approximate Logic Synthesis Based on Double-Chase Grey Wolf Optimizer
abstract
With the shrinking technology nodes, timing optimization becomes increasingly challenging. Approximate logic synthesis (ALS) can perform local approximate changes (LACs) on circuits to optimize timing with the cost of slight inaccuracy. However, existing ALS methods that focus solely on critical path depth reduction or area minimization are not optimal in timing optimization. This paper proposes an effective timing-driven ALS framework, where we employ a double-chase grey wolf optimizer to explore and apply LACs, simultaneously bringing excellent critical path shortening and area reduction under error constraints. Subsequently, it utilizes post-optimization under area constraints to convert area reduction into further timing improvement, thus achieving maximum critical path delay reduction. According to experiments on open-source circuits with 28nm technology, compared to the SOTA method, our framework can generate approximate circuits with greater critical path delay reduction under different error and area constraints.
Xiangfei Hu, Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001
DATE3
2025 RSizing: Robust Bayesian Optimization for Analog Circuit Sizing Under Process Variations
abstract
The increasing complexity of CMOS technology and circuit designs has intensified the need for robust analog design automation tools that can handle process variations effectively. This paper presents RSizing, a novel approach for analog circuit sizing that optimizes performance while ensuring robustness against process variations. Our method employs a three-phase strategy: First, it identifies promising design regions through nominal condition optimization to prune the design space efficiently. Second, it performs variation-aware optimization using heteroscedastic Gaussian processes (HGP) to model circuit performance under process variations, capturing the non-uniform nature of process-induced fluctuations across the design space. The HGP models are combined with an efficient acquisition function based on Thompson sampling to guide the exploration of robust designs using limited Monte Carlo simulations. Finally, it refines the solutions through additional targeted Monte Carlo simulations and model calibration. Experimental results on three benchmark circuits demonstrate that RSizing achieves superior performance compared to existing methods, consistently meeting yield requirements while optimizing multiple performance metrics with significantly reduced computational cost.
Jindong Tu, Peng Xu 0052, Zushuai Xie, Bei Yu 0001, Tinghuan Chen
ICCAD8
2025 Hierarchical Behavioral Learning-Based Dynamic Electromigration Analysis for Signal Networks
Jindong Tu, Tinghuan Chen, Qi Sun 0002, Cheng Zhuo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 IncreMacro: Incremental Macro Placement Refinement
abstract
This article proposes$\textsf {IncreMacro}$, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this,$\textsf {IncreMacro}$utilizes several methods, including kd-tree-based macro diagnosis, gradient-based macro shifting, constraint-graph-based LP for macro legalization, and diffusion-based cell migration. By employing these techniques iteratively,$\textsf {IncreMacro}$meets two critical solution requirements of macro placement: 1) pushing macros toward the chip boundary and 2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into$\textsf {AutoDMP}$and$\textsf {DREAMPlace}~4.0$, and is evaluated on seven RISC-V benchmark circuits and four TILOS macro placement circuit designs at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by$\textsf {AutoDMP}~(\textsf {DREAMPlace}~4.0$), our approach reduces routed wirelength by 15.1% (14.9%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 99.9 (82.6%) and 99.9% (81.3%), and reduces the total power consumption by 4.4% (4.3%). Meanwhile, compared with$\textsf {IncreMacro}$[1], our approach augmented with the cell migration algorithm improves the routed WNS and TNS by 24.7% and 23.1%, and remains the average routed wirelength and total power consumption almost unchanged.
Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Jiajun Qin, Haisheng Zheng, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 SMART: Graph Learning-Boosted Subcircuit Matching for Large-Scale Analog Circuits
abstract
Subcircuit matching in a large-scale analog circuit is a fundamental problem in VLSI computer-aided design (CAD). Existing approaches suffer from a poor scalability issue for a large-scale analog circuit. In this article, we propose a graph learning-boosted subcircuit matching framework for large-scale analog circuits named SMART, consisting of two stages. In the first stage, we customize hypergraph neural networks to map circuit topology for embedding space. Then, coarse subcircuit recognition is directly performed in the embedding space by geometric relations between the query circuit and all candidate subcircuits within the target circuit. In the second stage, a radial matching method, including device attribute matching, connection relationship matching and uniqueness-based matching, is customized to perform fine matching and obtain matches between interconnections and devices in the query circuit and candidate subcircuits. Experimental results show our SMART can outperform state-of-the-art search-based method VF3 and learning-based method NeuroMatch, and achieve the fastest speed. Specifically, using our framework for subcircuit matching can achieve up to$135\times $speedup with slight accuracy loss, and up to$7\times $speedup while maintaining 100% accuracy.
Jindong Tu, Pengjia Li, Peng Xu 0052, Qianru Zhang, Sanping Wan, Yongsheng Sun, Bei Yu 0001, Tinghuan Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2025 PARoute2: Enhanced Analog Routing via Performance-Drive Guidance Generation
abstract
Analog routing is crucial for performance optimization in analog circuit design, but conventionally takes significant development time and requires design expertise. Recent research has attempted to use machine learning (ML) to generate guidance to preserve circuit performance after analog routing. These methods face challenges such as expensive data acquisition and biased guidance. This article presents AnalogFold, a new paradigm of analog routing that leverages ML to provide performance-oriented routing guidance. Our approach learns performance-driven routing guidance and uses it to help automatic routers for performance-driven routing optimization. We propose to use a 3DGNN that incorporates cost-aware distance to make accurate predictions on post-layout performance. A pool-assisted potential relaxation process derives the effective routing guidance. The experimental results on multiple benchmarks under the TSMC 40 nm technology node demonstrate the superiority of the proposed framework compared to the cutting-edge works.
Peng Xu 0052, Jindong Tu, Guojin Chen, Keren Zhu 0001, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 Learning-Driven Physically Aware Large-Scale Circuit Gate Sizing
abstract
Gate sizing plays an important role in timing optimization after physical design. Existing machine learning-based gate sizing works cannot optimize timing on multiple timing paths simultaneously and neglect the physical constraint on layouts. They cause suboptimal sizing solutions and low-efficiency issues when compared with commercial gate sizing tools. In this work, we propose a learning-driven physically aware gate sizing framework to optimize timing performance on large-scale circuits efficiently. In our gradient descent optimization-based work, for obtaining accurate gradients, a multimodal gate sizing-aware timing model is achieved via learning timing information on multiple timing paths and physical information on multiple-scaled layouts jointly. Then, gradient generation based on the sizing-oriented estimator and adaptive back-propagation are developed to update gate sizes. Our results demonstrate that our work achieves higher-timing performance improvements in a faster way compared with the commercial gate sizing tool.
Yuyang Ye 0001, Peng Xu 0052, Lizheng Ren, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Algorithm-Hardware Co-design for Accelerating Depthwise Separable CNNs
abstract
Depthwise separable convolution (DSC) is a popular method for constructing lightweight neural networks. However, the pointwise convolution (PWC) has a much larger number of parameters than the depthwise convolution (DWC), causing the imbalanced parameter ratio of PWC to DWC. In this article, we propose an efficient and hardware-efficiency convolution (Shared Kernel sliding on channel Convolution, SKC) to replace the redundant PWC in DSC for a balanced parameter ratio, where SKC customizes the sharing kernel in the channel dimension to reduce the number of parameters, and the local connection in the channel dimension reduces the computation. Furthermore, the proposed SKC is suitable for Winograd acceleration, and the large kernel decomposition method is introduced to facilitate its use. We implement the first Winograd-based FPGA hardware accelerator for DSCNets. The shared 1D and 2D Winograd convolution computing engine is proposed to compute the proposed DSC consisting of DWC and SKC efficiently. An alternating loading and reusing storage approach is developed to efficiently load SKC input feature maps. Experimental results show our DSC-based accelerator can achieve 20× higher power efficiency at the cost of a small loss of accuracy by algorithm-hardware co-design compared with traditional accelerators.
RenGang Li, Tinghuan Chen, Meng Zhang 0010, Henk Corporaal
ACM Trans. Design Autom. Electr. Syst.4
2025 Rank-DSE: Neural Pareto Comparator of Microarchitecture Design Space Exploration
abstract
The complexity of microarchitecture design has surged due to the expanding design space and time-intensive verification processes. Existing regression-based machine learning methods struggle with inaccurate estimations because of limited training samples. To address these challenges, we propose Rank-DSE, a novel framework for microarchitecture design space exploration (DSE) that leverages a Neural Pareto Comparator (NPC) to directly model the comparative relationships between different architecture designs. Rank-DSE bypasses the inaccuracies of absolute PPA (performance, power, area) predictions by focusing on relative comparisons. The NPC computes the probability of one architecture dominating another and employs semi-supervised learning to reduce the reliance on labeled data. Additionally, a reinforcement-learning-based sampling scheme with an updating baseline Pareto set accelerates the exploration process. Experimental results on the ICCAD 2021 benchmark demonstrate that Rank-DSE achieves superior search quality and cost-efficiency compared to state-of-the-art methods. Specifically, Rank-DSE improves hypervolume by up to 7% while reducing exploration cost by 53.09% compared to cutting-edge approaches. These results highlight the advantages of Rank-DSE in terms of efficiency and effectiveness for microarchitecture DSE.
Peng Xu 0052, Su Zheng, Mingzi Wang, Ziyang Yu 0001, Shixin Chen, Tinghuan Chen, Keren Zhu 0001, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.6
2024 PDRC: Package Design Rule Checking via GPU-Accelerated Geometric Intersection Algorithms for Non-Manhattan Geometry
abstract
With the emergence of chiplet technology, the scale of IC packaging design has been steadily increasing, making the utilization of traditional design rule checking (DRC) methods more time-consuming. In this paper, we propose PDRC, a package-level design rule checker for non-manhattan geometry with GPU acceleration. PDRC employs hierarchical interval lists within an iterative parallel sweepline framework to implement the geometric intersection algorithm, thereby finishing design rule checking tasks. Experimental results have demonstrated 30 - 50 times speedup achieved by PDRC compared with two CPU-based checkers.
Jiaxi Jiang, Lancheng Zou, Wenqian Zhao 0002, Zhuolun He, Tinghuan Chen, Bei Yu 0001
DAC5
2024 WinoGen: A Highly Configurable Winograd Convolution IP Generator for Efficient CNN Acceleration on FPGA
abstract
The convolution neural network (CNN) has been widely adopted in computer vision tasks. In the FPGA-based CNN accelerator design, Winograd convolution can effectively improve computation performance and save hardware resources. However, building efficient and highly compatible IP for arbitrary Winograd convolution on FPGA remains underexplored. To address this issue, we propose a novel and efficient reformulation of Winograd convolution, named Structured Direct Winograd Convolution (SDW). We further develop WinoGen, a Chisel-based highly configurable Winograd convolution IP generator. Given arbitrary input/output tile size and kernel size, it can generate optimized high-performance IP automatically. Meanwhile, our generated IP can be compatible with multiple kernel sizes and tile sizes. Experimental results show that the IP generated by WinoGen achieves DSP efficiency up to 3.80 GOPS/DSP and energy efficiency up to 652.77 GOPS/W while showing 2.45× and 3.10× improvements when processing a same CNN model compared with state-of-the-arts.
Pengjia Li, Shixin Chen, Beichen Li 0003, Chong Tong, Jianlei Yang 0001, Tinghuan Chen, Bei Yu 0001
DAC8
2024 Performance-driven Analog Routing via Heterogeneous 3DGNN and Potential Relaxation
abstract
Analog routing is crucial for performance optimization in analog circuit design, but conventionally takes significant development time and requires design expertise. Recent research has attempted to use machine learning (ML) to generate guidance to preserve circuit performance after analog routing. These methods face challenges such as expensive data acquisition and biased guidance. This paper presents AnalogFold, a new paradigm of analog routing that leverages ML to provide performance-oriented routing guidance. Our approach learns performance-driven routing guidance and uses it to help automatic routers for performance-driven routing optimization. We propose to use a 3DGNN that incorporates cost-aware distance to make accurate predictions on post-layout performance. A pool-assisted potential relaxation process derives the effective routing guidance. The experimental results on multiple benchmarks under the TSMC 40nm technology node demonstrate the superiority of the proposed framework compared to the cutting-edge works.
Peng Xu 0052, Guojin Chen, Keren Zhu 0001, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
DAC4
2024 CBTune: Contextual Bandit Tuning for Logic Synthesis
abstract
Logic synthesis pre-optimization involves applying a sequence of transformations called synthesis flow to reduce the circuit's Boolean logic graph, like AIG. However, the challenge lies in selecting and arranging these transformations due to the exponentially expanding solution space. In this work, we propose CBTune, a novel online learning framework that utilizes a contextual bandit algorithm to explore the solution space and generate synthesis flows efficiently. We develop the Syn-LinUCB algorithm as the agent, which incorporates circuit characteristics and leverages long-term payoffs to guide decision-making, thus ef-fectively preventing getting trapped in local optima. Experimental results show that our framework achieves the optimal synthesis flow with a lower time cost, substantially reducing the number of AIG nodes and 6-LUTs compared to SOTA approaches.
Fangzhou Liu 0005, Zehua Pei, Ziyang Yu 0001, Haisheng Zheng, Zhuolun He, Tinghuan Chen, Bei Yu 0001
DATE6
2024 Attention-Based EDA Tool Parameter Explorer: From Hybrid Parameters to Multi-QoR metrics
abstract
Improving the outcomes of very-large-scale integration design without altering the underlying design enablement, such as process, device, interconnect, and IPs, is critical for integrated circuit (IC) designers. Parameter tuning for electronic design automation (EDA) tools is an emerging technology for improving the final design Quality-of-Result (QoR). However, many complex heuristics have been accreted upon previous complex heuristics integrated into tools, resulting in a vast number of tunable parameters. Even worse, these parameters include both continuous and discrete ones, making the parameter tuning process laborious and challenging. In this paper, we propose an attention-based EDA tool parameter explorer. A self-attention mechanism is developed to navigate the parameter importance. A hybrid space Gaussian process model is leveraged to optimize continuous and discrete parameters jointly, capturing their complex interactions. In addition, considering multiple QoR metrics and the large amount of time required to invoke EDA tools, a customized acquisition function based on expected hypervolume improvement (EHVI) is proposed to enable multi-objective optimization and parallel evaluation. Experimental results on a set of IWLS2005 benchmarks demonstrate the effectiveness and efficiency of our method.
Donger Luo, Qi Sun 0002, Qi Xu 0004, Tinghuan Chen, Hao Geng
DATE4
2024 A Graph-Learning-Driven Prediction Method for Combined Electromigration and Thermomigration Stress on Multi-Segment Interconnects
abstract
As technology advances, the temperature gradient in the interconnects becomes more significant, which causes serious thermomigration. Simulating the coupling effects of thermomigration (TM) and electromigration (EM) on large-scale circuits is very time-consuming caused by a substantial increase in computational complexity. Recently, some researchers utilized graph learning-based methods to predict EM stress in medium-scale cases. Unfortunately, these works overlooked the effects of TM. To predict the EM - TM stress of large-scale interconnects accurately and efficiently, we propose a framework based on Graph Attention Networks (GATs) with a customized alternating aggregation method for collecting information in junctions and branches of interconnects jointly. The experimental results show that our work achieves an average relative error of less than 1 % compared to the commercial software COMSOL for inter-connects consisting of fewer than 200 segments. Furthermore, our method also achieves 9037 x speedup in predicting the OpenROAD test circuit with a maximum segment number reaching 10807.
Yunfan Zuo, Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Longxing Shi
DATE4
2024 RISCSparse: Point Cloud Inference Engine on RISC-V Processor
abstract
Machine learning on point clouds is increasingly accessible at the edge, notably in applications such as autonomous driving. However, the sparse and irregular nature of point clouds presents significant latency challenges on general-purpose hardware. RISC-V, with its evolving ecosystem, offers a promising platform for embedding intelligence at the edge due to its full-stack scalability. This paper focuses on the advanced point cloud operation known as submanifold convolution (SC), deploying submanifold sparse convolutional networks (SSCNs) on a RISC-V System-on-Chip (SoC) designed within the Chipyard framework. We address three critical bottlenecks of SSCNs- Rule Map Construction (Mapping), Gather-MatMul-Scatter (GMS), and uncombined operation - to meet the real-time inference requirement for the on-chip implementation. By leveraging the RISC-V Vector extension and Gemmini, an open-source full-stack DNN accelerator generator, we vectorize the Mapping process, offload GEMM-related operations to the Gemmini Systolic Array, and cooperatively use the Systolic Array and vector processing units to reduce the memory footprint. Our evaluations show that the RISC-V-based SSCNs implementation achieves an average of 11.73× and 13.1× overall speedups with a small workload compared to TorchSparse on Edge-CPU, a state-of-the-art point cloud inference engine, for 3D segmentation and detection tasks, respectively. When contrasting with TorchSparse on an Edge-GPU, our implementation still delivers a notable improvement, with average speedups of 1.63× for 3D segmentation and 1.07× for detection tasks.
Shangran Lin, Xinrui Zhu, Baohui Xie, Tinghuan Chen, Cheng Zhuo, Qi Sun 0002, Bei Yu 0001
ICCAD4
2024 IncreMacro: Incremental Macro Placement Refinement
abstract
This paper proposes IncreMacro, a novel approach for macro placement refinement in the context of integrated circuit (IC) design. The suggested approach iteratively and incrementally optimizes the placement of macros in order to enhance IC layout routability and timing performance. To achieve this, IncreMacro utilizes several methods including kd-tree-based macro diagnosis, gradient-based macro shifting and constraint-graph-based LP for macro legalization. By employing these techniques iteratively, IncreMacro meets two critical solution requirements of macro placement: (1) pushing macros to the chip boundary; and (2) preserving the original macro relative positional relationship. The proposed approach has been incorporated into DREAMPlace and AutoDMP, and is evaluated on several RISC-V benchmark circuits at the 7-nm technology node. Experimental results show that, compared with the macro placement solution provided by DREAMPlace (AutoDMP), IncreMacro reduces routed wirelength by 6.5% (16.8%), improves the routed worst negative slack (WNS) and total negative slack (TNS) by 59.9% (99.6%) and 63.9% (99.9%), and reduces the total power consumption by 3.3% (4.9%).
Yuan Pu 0001, Tinghuan Chen, Zhuolun He, Haisheng Zheng, Yibo Lin, Bei Yu 0001
ISPD2
2024 An analysis of TinyML@ICCAD for implementing AI on low-power microprocessor
Meng Zhang 0010, Tinghuan Chen, Jun Yang 0006
Sci. China Inf. Sci.5
2024 Timing-Driven Technology Mapping Approximation Based on Reinforcement Learning
abstract
As the transistor technology nodes shrink into the nano-scales, timing guardbands caused by aging effects and process variations continue to increase. Approximate computing can eliminate aging-and-variation-induced timing guardbands without sacrificing the design performance. It can apply local approximate changes (LACs) automatically in circuits to reduce critical path delay. However, efficiently achieving timing optimization under error distance constraints is still tricky. This work proposes an automated timing-driven technology mapping approximation framework based on reinforcement learning (RL). The framework uses path-weighted graph neural networks (PGNNs) to embed RL states and timing path-aware LAC candidates to construct RL action spaces. It can efficiently eliminate timing guardbands induced by aging and variation. Our proposed circuit-agnostic framework operates on the gate-level netlists. According to experiments on the open-source circuits using TSMC 28nm and 16nm technology under aging and variation conditions, our framework can achieve an average 24.78% critical path delay reduction under 5 different error distance constraints and 4.83× speedup, compared with a state-of-the-art method.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Fast Constraints Tuning via Transfer Learning and Multiobjective Optimization
abstract
As the complexity of very-large-scale integration (VLSI) increases, empirically determining the design constraints necessary to achieve the optimal performance, power, and area (PPA) within the electronic design automation (EDA) workflow becomes more challenging. Design space exploration is capable of effectively and automatically identifying the design constraints required to attain the optimal PPA in VLSI designs. However, the absence of prior knowledge can lead to less efficient explorations. This paper proposes a novel fast constraint tuning framework via transfer learning and multi-objective Bayesian optimization (MOBO) to find the optimal design constraints. Firstly, we introduce transfer learning into multi-objective Bayesian optimization by Gaussian Copula and transform the PPA data into residual observations. We propose to transfer the prior information of the implemented technologies to the advanced technology to optimize the parameter design space under the advanced technology. Secondly, we propose Gaussian process regression with an auto-encoder-based deep kernel as a surrogate model in MOBO. The auto-encoder-based deep kernel can extract more input features to make the surrogate model more precise. We employ the batch uncertainty-aware search acquisition function to improve exploration efficiency. Using this surrogate model and this acquisition function in MOBO can reduce the amount that EDA tools need to run. The average EDA tools running times of the proposed model is 204, and the average ADRS is 0.0373. Compared to state-of-the-art approaches, experiments on a CPU design reveal that a higher-quality Pareto frontier can be provided with a shorter running time.
Meng Zhang 0010, Yifan Niu, Zewei Chen, Yajun Ha, Tinghuan Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2024 Wages: The Worst Transistor Aging Analysis for Large-scale Analog Integrated Circuits via Domain Generalization
abstract
Transistor aging leads to the deterioration of analog circuit performance over time. The worst aging degradation is used to evaluate the circuit reliability. It is extremely expensive to obtain it since several circuit stimuli need to be simulated. The worst degradation collection cost reduction brings an inaccurate training dataset when a machine learning (ML) model is used to fast perform the estimation. Motivated by the fact that there are many similar subcircuits in large-scale analog circuits, in this article we propose Wages to train an ML model on an inaccurate dataset for the worst aging degradation estimation via a domain generalization technique. A sampling-based method on the feature space of the transistor and its neighborhood subcircuit is developed to replace inaccurate labels. A consistent estimation for the worst degradation is enforced to update model parameters. Label updating and model updating are performed alternately to train an ML model on the inaccurate dataset. Experimental results on the very advanced 5 nm technology node show that our Wages can significantly reduce the label collection cost with a negligible estimation error for the worst aging degradations compared to the traditional methods.
Tinghuan Chen, Hao Geng, Qi Sun 0002, Sanping Wan, Yongsheng Sun, Huatao Yu, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.1
2024 DeepOTF: Learning Equations-constrained Prediction for Electromagnetic Behavior
abstract
High-quality passive devices are becoming increasingly important for the development of mobile devices and telecommunications, but obtaining such devices through simulation and analysis of electromagnetic (EM) behavior is time-consuming. To address this challenge, artificial neural network (ANN) models have emerged as an effective tool for modeling EM behavior, with NeuroTF being a representative example. However, these models are limited by the specific form of the transfer function, leading to discontinuity issues and high sensitivities. Moreover, previous methods have overlooked the physical relationship between distributed parameters, resulting in unacceptable numeric errors in the conversion results. To overcome these limitations, we propose two different neural network architectures: DeepOTF and ComplexTF. DeepOTF is a data-driven deep operator network for automatically learning feasible transfer functions for different geometric parameters. ComplexTF utilizes complex-valued neural networks to fit feasible transfer functions for different geometric parameters in the complex domain while maintaining causality and passivity. Our approach also employs an Equations-constraint Learning scheme to ensure the strict consistency of predictions and a dynamic weighting strategy to balance optimization objectives. The experimental results demonstrate that our framework shows superior performance than baseline methods, achieving up to 1,700× higher accuracy.
Peng Xu 0052, Tinghuan Chen, Guojin Chen, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.3
2023 Mixed-Type Wafer Failure Pattern Recognition
abstract
The ongoing evolution in process fabrication enables us to step below the 5nm technology node. Although foundries can pattern and etch smaller but more complex circuits on silicon wafers, a multitude of challenges persist. For example, defects on the surface of wafers are inevitable during manufacturing. To increase the yield rate and reduce time-to-market, it is vital to recognize these failures and identify the failure mechanisms of these defects. Recently, applying machine learning-powered methods to combat single defect pattern classification has made significant progress. However, as the processes become increasingly complicated, various single-type defect patterns may emerge and be coupled on a wafer and thus shape a mixed-type pattern. In this paper, we will survey the recent pace of progress on advanced methodologies for wafer failure pattern recognition, especially for mixed-type one. We sincerely hope this literature review can highlight the future directions and promote the advancement of the wafer failure pattern recognition.
Hao Geng, Qi Sun 0002, Tinghuan Chen, Qi Xu 0004, Tsung-Yi Ho, Bei Yu 0001
ASP-DAC3
2023 Graph-Learning-Driven Path-Based Timing Analysis Results Predictor from Graph-Based Timing Analysis
abstract
With diminishing margins in advanced technology nodes, the performance of static timing analysis (STA) is a serious concern, including accuracy and runtime. The STA can generally be divided into graph-based analysis (GBA) and path-based analysis (PBA). For GBA, the timing results are always pessimistic, leading to overdesign during design optimization. For PBA, the timing pessimism is reduced via propagating real path-specific slews with the cost of severe runtime overheads relative to GBA. In this work, we present a fast and accurate predictor of post-layout PBA timing results from inexpensive GBA based on deep edge-featured graph attention network, namely deep EdgeGAT. Compared with the conventional machine and graph learning methods, deep EdgeGAT can learn global timing path information. Experimental results demonstrate that our predictor has the potential to substantially predict PBA timing results accurately and reduce timing pessimism of GBA with maximum error reaching 6.81 ps, and our work achieves an average 24.80× speedup faster than PBA using the commercial STA tool.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
ASP-DAC2
2023 Fast and Accurate Wire Timing Estimation Based on Graph Learning
abstract
Accurate wire timing estimation has become a bottleneck in timing optimization since it needs a long turn-around time using a sign-off timer. The gate timing can be calculated accurately using lookup tables in cell libraries. In comparison, the accuracy and efficiency of wire timing calculation for complex RC nets are extremely hard to trade-off. The limited number of wire paths opens a door for the graph learning method in wire timing estimation. In this work, we present a fast and accurate wire timing estimator based on a novel graph learning architecture, namely GNNTrans. It can generate wire path representations by aggregating local structure information and global relationships of whole RC nets, which cannot be collected with traditional graph learning work efficiently. Experimental results on both tree-like and non-tree nets demonstrate improved accuracy, with the max error of wire delay being lower than 5 ps. In addition, our estimator can predict the timing of over 200K nets in less than 100 secs. The fast and accurate work can be integrated into incremental timing optimization for routed designs.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
DATE2
2023 TRouter: Thermal-Driven PCB Routing via Nonlocal Crisscross Attention Networks
abstract
In this article, we propose TRouter, a thermal-driven printed circuit board (PCB) routing framework via a machine-learning model. The model is designed to capture the long-range spatial information from the PCB layout and predict thermal distribution. The information contains pads, vias, components and wire segments. A gradient in each grid cell obtained from the backpropagation is integrated into a full-board routing algorithm to guide thermal-aware wire detour and via punching. To achieve a significant speedup, we construct a conflict graph according to whether overlapping among convex hulls of nets. A greedy-based method is adopted to remove nonroot nodes from all nodes. Then, a task graph is constructed to improve the parallelism. We conduct experiments on open-source benchmarks to illustrate our TRouter can achieve significant speedup and lower-temperature designs, compared with a state-of-the-art PCB routing algorithm.
Tinghuan Chen, Silu Xiong, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 PTPT: Physical Design Tool Parameter Tuning via Multi-Objective Bayesian Optimization
abstract
Physical design flow through associated electronic design automation (EDA) tools plays an imperative role in the advanced integrated circuit design. Mostly, the parameters fed into physical design tools are mainly manually picked based on the domain knowledge of the experts. Nevertheless, owing to the ever-shrinking scaling down of technology nodes and the complexity of the design space spanned by combinations of the parameters, even coupled with the time-consuming simulation process, such manual explorations for parameter configurations of physical design tools have become extremely laborious. There exist a few works in the field of design flow parameter tuning. However, very limited prior arts explore the complex correlations among multiple quality-of-result (QoR) metrics of interest (e.g., delay, power, and area) and explicitly optimize these goals simultaneously. To overcome these weaknesses and seek effective parameter settings of physical design tools, in this article, we propose a multi-objective Bayesian optimization (BO) framework with a multi-task Gaussian model as the surrogate model. An information gain-based acquisition function is adopted to sequentially choose candidates for tool simulation to efficiently approximate the Pareto-optimal parameter configurations. The experimental results on three industrial benchmarks under the 7-nm technology node demonstrate the superiority of the proposed framework compared to the cutting-edge works.
Hao Geng, Tinghuan Chen, Yuzhe Ma, Binwu Zhu, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Aging-Aware Critical Path Selection via Graph Attention Networks
abstract
In advanced technology nodes, aging effects like negative and positive bias temperature instability (NBTI and PBTI) become increasingly significant, making timing closure and optimization more challenging. Unfortunately, conventional critical path (CP) selection tools used in reliability-aware design flow cannot accurately identify CPs under different aging conditions. To address this issue, we propose an aging-aware CP selection flow comprising two parts: 1) critical cell detection and 2) path criticality (PC) computation. We employ graph-attention (GAT) networks to predict the critical cells in the aged circuits, and a PC computation algorithm that takes into account circuit-level and transistor-level parameters to generate PC rank lists. Our experimental results demonstrate that our GAT model outperforms classical machine learning models in detecting critical cells. Additionally, compared with the commercial tool, our aging-aware flow achieves an average accuracy of 99.52%, 98.69%, and 97.20% for top-10%, top-5%, and top-1% path sets, respectively, in five industrial designs subjected to different aging conditions and workloads.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Techniques for CAD Tool Parameter Auto-tuning in Physical Synthesis: A Survey (Invited Paper)
abstract
As the technology node of integrated circuits rapidly goes beyond 5nm, synthesis-centric modern very large-scale integration (VLSI) design flow is facing ever-increasing design complexity and suffering the pressure of time-to-market. During the past decades, synthesis tools have become progressively sophisticated and offer countless tunable parameters that can significantly influence design quality. Nevertheless, owing to the time-consuming tool evaluation plus a limitation to one possible parameter combination per synthesis run, manually searching for optimal configurations of numerous parameters proves to be elusive. What's worse, tiny perturbations to these parameters can result in very large variations in the Quality-of-Results (QoR). Therefore, automatic tool parameter tuning to reduce human cost and tool evaluation cost is in demand. Machine-learning techniques provide chances to enable the auto-tuning process of tool parameters. In this paper, we will survey the recent pace of progress on advanced parameter auto-tuning flows of physical synthesis tools. We sincerely expect this survey can enlighten the future development of parameter auto-tuning methodologies.
Hao Geng, Tinghuan Chen, Qi Sun 0002, Bei Yu 0001
ASP-DAC2
2022 A fast parameter tuning framework via transfer learning and multi-objective bayesian optimization
abstract
Design space exploration (DSE) can automatically and effectively determine design parameters to achieve the optimal performance, power and area (PPA) in very large-scale integration (VLSI) design. The lack of prior knowledge causes low efficient exploration. In this paper, a fast parameter tuning framework via transfer learning and multi-objective Bayesian optimization is proposed to quickly find the optimal design parameters. Gaussian Copula is utilized to establish the correlation of the implemented technology. The prior knowledge is integrated into multi-objective Bayesian optimization through transforming the PPA data to residual observation. The uncertainty-aware search acquisition function is employed to explore design space efficiently. Experiments on a CPU design show that this framework can achieve a higher quality of Pareto frontier with less design flow running than state-of-the-art methodologies.
Tinghuan Chen, Jiaxin Huang 0010, Meng Zhang 0010
DAC2
2022 Deep H-GCN: Fast Analog IC Aging-Induced Degradation Estimation
abstract
With continued scaling, the transistor aging induced by hot carrier injection (HCI) and bias temperature instability (BTI) causes an increasing failure of nanometer-scale integrated circuits (ICs). Compared to digital ICs, analog ICs are more susceptible to aging effects. The industrial large-scale analog ICs bring grand challenges in the efficiency of aging verification. In this article, we propose a heterogeneous graph convolutional network (H-GCN) to fast estimate aging-induced transistor degradation in analog ICs. To characterize the multityped devices and connection pins, a heterogeneous directed multigraph is adopted to efficiently represent the topology of analog ICs. A latent space mapping method is used to transform the feature vector of all typed devices into a unified latent space. We further extend the proposed H-GCN to be a deep version via initial residual connections and identity mappings. The extended deep H-GCN can extract information from multihop devices without an oversmoothing issue. A probability-based neighborhood sampling method on the bipartite graph is adopted to ease the model training on large-scale graphs and achieve good scalability. Experiments on very advanced 5-nm industrial benchmarks show that, compared to traditional graph learning methods and static aging reliability simulations by an industrial design-for-reliability (DFR) tool, the proposed deep H-GCN can achieve more accurate estimations of aging-induced transistor degradation. Compared to the dynamic and static aging reliability simulations, our extended deep H-GCN, on average, can achieve$241\times $and$39\times $speedup, respectively.
Tinghuan Chen, Qi Sun 0002, Canhui Zhan, Changze Liu, Huatao Yu, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 An Efficient Sharing Grouped Convolution via Bayesian Learning
abstract
Compared with traditional convolutions, grouped convolutional neural networks are promising for both model performance and network parameters. However, existing models with the grouped convolution still have parameter redundancy. In this article, concerning the grouped convolution, we propose a sharing grouped convolution structure to reduce parameters. To efficiently eliminate parameter redundancy and improve model performance, we propose a Bayesian sharing framework to transfer the vanilla grouped convolution to be the sharing structure. Intragroup correlation and intergroup importance are introduced into the prior of the parameters. We handle the Maximum Type II likelihood estimation problem of the intragroup correlation and intergroup importance by a group LASSO-type algorithm. The prior mean of the sharing kernels is iteratively updated. Extensive experiments are conducted to demonstrate that on different grouped convolutional neural networks, the proposed sharing grouped convolution structure with the Bayesian sharing framework can reduce parameters and improve prediction accuracy. The proposed sharing framework can reduce parameters up to 64.17%. For ResNeXt-50 with the sharing grouped convolution on ImageNet dataset, network parameters can be reduced by 96.875% in all grouped convolutional layers, and accuracies are improved to 78.86% and 94.54% for top-1 and top-5, respectively.
Tinghuan Chen, Qi Sun 0002, Meng Zhang 0010, Hao Geng, Qianru Zhang, Bei Yu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Correlated Multi-objective Multi-fidelity Optimization for HLS Directives Design
abstract
High-level synthesis (HLS) tools have gained great attention in recent years because it emancipates engineers from the complicated and heavy hardware description language writing and facilitates the implementations of modern applications (e.g., deep learning models) on Field-programmable Gate Array (FPGA) , by using high-level languages and HLS directives. However, finding good HLS directives is challenging, due to the time-consuming design processes, the balances among different design objectives, and the diverse fidelities (accuracies of data) of the performance values between the consecutive FPGA design stages. To find good HLS directives, a novel automatic optimization algorithm is proposed to explore the Pareto designs of the multiple objectives while making full use of the data with different fidelities from different FPGA design stages. Firstly, a non-linear Gaussian process (GP) is proposed to model the relationships among the different FPGA design stages. Secondly, for the first time, the GP model is enhanced as correlated GP (CGP) by considering the correlations between the multiple design objectives, to find better Pareto designs. Furthermore, we extend our model to be a deep version deep CGP (DCGP) by using the deep neural network to improve the kernel functions in Gaussian process models, to improve the characterization capability of the models, and learn better feature representations. We test our design method on some public benchmarks (including general matrix multiplication and sparse matrix-vector multiplication) and deep learning-based object detection model iSmart2 on FPGA. Experimental results show that our methods outperform the baselines significantly and facilitate the deep learning designs on FPGA.
Qi Sun 0002, Tinghuan Chen, Siting Liu 0002, Jianli Chen, Hao Yu 0001, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.2
2021 Analog IC Aging-induced Degradation Estimation via Heterogeneous Graph Convolutional Networks
abstract
With continued scaling, transistor aging induced by Hot Carrier Injection and Bias Temperature Instability causes a gradual failure of nanometer-scale integrated circuits (ICs). In this paper, to characterize the multi-typed devices and connection ports, a heterogeneous directed multigraph is adopted to efficiently represent analog IC post-layout netlists. We investigate a heterogeneous graph convolutional network (H-GCN) to fast and accurately estimate aging-induced transistor degradation. In the proposed H-GCN, an embedding generation algorithm with a latent space mapping method is developed to aggregate information from the node itself and its multi-typed neighboring nodes through multi-typed edges. Since our proposed H-GCN is independent of dynamic stress conditions, it can replace static aging analysis. We conduct experiments on very advanced 5nm industrial designs. Compared to traditional machine learning and graph learning methods, our proposed H-GCN can achieve more accurate estimations of aging-induced transistor degradation. Compared to an industrial reliability tool, our proposed H-GCN can achieve 24.623x speedup on average.
Tinghuan Chen, Qi Sun 0002, Canhui Zhan, Changze Liu, Huatao Yu, Bei Yu 0001
ASP-DAC1
2021 Correlated Multi-objective Multi-fidelity Optimization for HLS Directives Design
abstract
High-level synthesis (HLS) tools have gained great attention in recent years because it emancipates engineers from the complicated and heavy hardware description language writing, by using high-level languages and HLS directives. However, previous works seem powerless, due to the time-consuming design processes, the contradictions among design objectives, and the accuracy difference between the three stages (fidelities). To find good HLS directives, in this paper, a novel correlated multi-objective non-linear optimization algorithm is proposed to explore the Pareto solutions while making full use of data from different fidelities. A non-linear Gaussian process is proposed to model relationships among the analysis reports from different fidelities for the same objective. For the first time, correlated multivariate Gaussian process models are introduced into this domain to characterize the complex relationships of multiple objectives in each design fidelity. A tree-based method is proposed to erase invalid solutions and obviously non-optimal solutions. Experimental results show that our non-linear and pioneering correlated models can approximate the Pareto-frontier of the directive design space in a shorter time with much better performance and good stability, compared with the state-of-the-art.
Qi Sun 0002, Tinghuan Chen, Siting Liu 0002, Jin Miao, Jianli Chen, Hao Yu 0001, Bei Yu 0001
DATE2
2021 Fast and Efficient DNN Deployment via Deep Gaussian Transfer Learning
abstract
Deep neural networks (DNNs) have been widely used recently while their hardware deployment optimizations are very time-consuming and the historical deployment knowledge is not utilized efficiently. In this paper, to accelerate the optimization process and find better deployment configurations, we propose a novel transfer learning method based on deep Gaussian processes (DGPs). Firstly, a deep Gaussian process (DGP) model is built on the historical data to learn empirical knowledge. Secondly, to transfer knowledge to a new task, a tuning set is sampled for the new task under the guidance of the DGP model. Then DGP is tuned according to the tuning set via maximum-a-posteriori (MAP) estimation to accommodate for the new task and finally used to guide the deployments of the task. The experiments show that our method achieves the best inference latencies of convolutions while accelerating the optimization process significantly, compared with previous arts.
Qi Sun 0002, Tinghuan Chen, Hao Geng, Xinyun Zhang 0001, Bei Yu 0001
ICCV3
2021 Leveraging Spatial Correlation for Sensor Drift Calibration in Smart Building
abstract
Sensor drift is an intractable obstacle to practical temperature measurement in smart building. In this article, we propose a sensor spatial correlation model. Given prior knowledge, maximum a posteriori (MAP) estimation is performed to calibrate drifts. MAP is formulated as a nonconvex problem with three hyper-parameters. An alternating-based method is proposed to solve this nonconvex formulation. Cross-validation, Gibbs expectation-maximization (EM) and variational Bayesian EM (VB-EM) are further exploited to determine hyper-parameters. Experimental results on widely used benchmarks from the simulator EnergyPlus demonstrate that compared with state-of-the-art methods, the proposed framework can achieve a robust drift calibration and a better tradeoff between accuracy and runtime. On average, compared with state-of-the-art, the proposed framework can achieve about 3× accuracy improvement. In order to attain the same drift calibration accuracy with VB-EM, Gibbs EM needs 10 000 samples, which will incur a 30× runtime overhead.
Tinghuan Chen, Bingqing Lin, Hao Geng, Shiyan Hu 0001, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Sensor Drift Calibration via Spatial Correlation Model in Smart Building
abstract
Sensor drift is an intractable obstacle to practical temperature measurement in smart building. In this paper, we propose a sensor spatial correlation model. Given prior knowledge, Maximum-aposteriori (MAP) estimation is performed to calibrate drifts. MAP is formulated as a non-convex problem with three hyper-parameters. An alternating-based method is proposed to solve this non-convex formulation. Cross-validation and Expectation-maximum with Gibbs sampling are further to determine hyper-parameters. Experimental results show that on benchmarks from simulator EnergyPlus, compared with state-of-the-art method, the proposed framework can achieve a robust drift calibration and a better trade-off between accuracy and runtime.
Tinghuan Chen, Bingqing Lin, Hao Geng, Bei Yu 0001
DAC1
2019 Power-Driven DNN Dataflow Optimization on FPGA
abstract
Deep neural networks (DNNs) have been proven to achieve unprecedented success on modern artificial intelligence (AI) tasks, which have also greatly motivated the rapid developments of novel DNN models and hardware accelerators. Many challenges still remain towards the design of power efficient DNN accelerator due to the intrinsically intensive data computation and transmission in DNN algorithms. However, most existing efforts in the domain have taken latency as the sole optimization objective, which may often result in sub-optimality in power consumption. In this paper, we propose a framework to optimize the power efficiency of DNN dataflow on FPGA while maximally minimizing the impact on latency. We first propose power and latency models that are built upon different dataflow configurations. Then a power-driven dataflow formulation is proposed, which enables a hierarchical exploration strategy on the dataflow configurations, leading to efficient power consumption at limited latency loss. Experimental results have demonstrated the effectiveness of our proposed models and exploration strategies, where power improvement has shown up to 31% with latency degradation of no worse than 6.5%.
Qi Sun 0002, Tinghuan Chen, Jin Miao, Bei Yu 0001
ICCAD2
2019 Recent advances in convolutional neural network acceleration
Qianru Zhang, Meng Zhang 0010, Tinghuan Chen, Zhifei Sun, Yuzhe Ma, Bei Yu 0001
Neurocomputing3
2018 Electricity Theft Detection Using Generative Models
abstract
Advanced metering infrastructure (AMI) plays an important role in smart grid. On one hand, AMI makes the smart grid more vulnerable to cyber attacks. On the other hand, large amount of available usage data helps detect energy thefts using machine learning methods. In this paper, we focus on energy theft that results in customer usage pattern change in utility database. To overcome the imbalance problem between normal and anomaly behavior data, we propose an anomaly detection framework called semi-supervised generative Gaussian mixture model, which can be controlled with detection indicator thresholds to adjust the intensity of detection. Human knowledge is successfully introduced into the model using detection indicators. We analyze it with various machine learning based methods including one-class SVM and autoencoder, and show that our framework has the most effective performance validated by simulation that is based on real-world energy consumption data.
Qianru Zhang, Meng Zhang 0010, Tinghuan Chen, Jinan Fan
ICTAI3