Yuyang Ye 0001

dblp:194/4226-1 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-0726-0468ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 8 first-author · 22 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 DCTDSE: A Bimodal Design Space Exploration Flow via Discrete-Continuous Transformation
abstract
The conservation core accelerator presents a promising avenue for improving the computational efficiency of specific applications. However, its design space is ultra-high-dimensional, significantly increasing the exploratory effort required to identify the optimal design across performance, power, and area metrics. Furthermore, the discrete nature of the microarchitecture space renders conventional search methods ineffective. To tackle these challenges, we propose the Discrete Continuous Transformation to speedup Design Space Exploration, namely DCTDSE. It can operate in either offline or online mode. In offline mode, it transforms the original discrete design space into a continuous space, builds predictive models, performs parallel gradient-based optimization, and maps the results back to the discrete domain. In the online mode, DCTDSE refines the models by iteratively resampling previously found solutions, thereby enhancing exploration quality while maintaining moderate runtime overhead. Experimental results indicate that DCTDSE achieves a 3.9× to 40× speedup over benchmark methods in offline mode. In online mode, it provides a 2.5× speedup, with a 21% reduction in exploration quality relative to the most accurate comparison method.
Shuaibo Huang, Liangji Wu, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization
abstract
Electronic design automation (EDA) tools are critical in the very large scale integration (VLSI) flow. To address the challenges posed by the extensive search space and intricate feature interactions, statistical and machine-learning methods have been employed. These methods aim to model tool parameters and treat the tuning process as a regression task. However, these regression-based methods suffer from inaccurate estimations owing to limited training samples. To address this issue, we propose a ranking-based tool parameter tuning framework, called RankTuner, which directly learns the dominant relationship between parameters. RankTuner utilizes a pairwise Gaussian process to estimate the probability and uncertainty of the dominance relationship. Our approach also integrates a Duel-Thompson sampling method to balance exploration and exploitation in parameter selections. A dimensionality reduction scheme with random embedding and trust region techniques is incorporated to enable parallel searches. Experimental results demonstrate the superiority of RankTuner compared to the cutting-edge tool parameter tuning methods.
Peng Xu 0052, Su Zheng, Yuyang Ye 0001, Hao Geng, Tsung-Yi Ho, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 StatCHAR: Statistical Timing Characterization Framework via Heterogeneous Graph Attention Network and Active Learning With Parasitic RC Reduction
abstract
Statistical timing characterization for standard cell library poses significant challenges to accuracy and runtime cost. Prior analytical and learning-based methods neglect the profound influence induced by the layout-dependent parasitic resistor and capacitor (RC) network in cell netlist as well as the timing correlation between the topological structures of cells and process, voltage, and temperature (PVT) corners for model training, resulting in tremendous simulation effort and poor accuracy. In this work, a Statistical timing Characterization framework via Heterogeneous graph attention network and Active learning with parasitic RC Reduction (StatCHAR) is proposed, where the transistors and parasitic RC in cell are represented as heterogeneous nodes for graph learning and redundant RC nodes are removed to alleviate node imbalance issue and improve accuracy. The significant training data are selected from the full characterization set with active learning strategy to achieve the optimal balance between simulation overhead for the training set and prediction precision for the remaining test set. The proposed framework was validated with typical standard cells under multiple PVT corners with TSMC 22nm process, which achieves an excellent prediction with a relative Root Mean Square Error (rRMSE) of only 2.43% with only 11.9% of total characterization data for training, demonstrating an accuracy improvement of 2.7$\times \sim 12.1\times $for statical timing analysis on benchmark circuits compared to competitive learning based methods and a characterization runtime reduction by 7.5$\times $.
Peng Cao 0002, Zeyuan Deng, Yuhan Dong, Yuyang Ye 0001, Jun Yang 0006
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 Rank-based Multi-objective Approximate Logic Synthesis via Monte Carlo Tree Search
abstract
Approximate Logic Synthesis (ALS) is an automated technique designed for error-tolerant applications, optimizing delay, area, and power under specified error constraints. However, existing methods typically focus on either delay reduction or area minimization, often leading to local optima in multi-objective optimization. This paper proposes a rankbased multi-objective ALS framework using Monte Carlo Tree Search (MCTS). It develops non-dominated circuit ranking, to guide MCTS in exploring local approximate changes (LACs) across the entire circuit and generate approximate circuit sets with great optimization potential. Additionally, a Rank-Transformer model is introduced to predict pathdomain ranks, enhancing the application of high-quality LACs within circuit paths. Experimental results show that our framework achieves faster and more efficient optimization in delay and area simultaneously compared to state-of-the-art methods.
Yuyang Ye 0001, Xiangfei Hu, Peng Xu 0052, Yu Gong 0002, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
DAC1
2025 Truly Pre-Routing Timing Prediction via Considering Power Delivery Network
abstract
Fast and accurate pre-routing timing prediction is essential in the chip design flow. However, existing machine learning (ML)assisted pre-routing timing methods often overlook the impact of power delivery networks (PDNs), which contribute to IR drop and routing congestion. This limitation can make these methods less practical for realworld circuit design flows. To address this, we propose two specialized encoders-an IR drop-aware encoder and a routing congestion-aware encoder-that effectively capture PDN effects through multimodal fusion of netlist, layout, and PDN data. To mitigate the challenges of imbalanced multimodal fusion, we further develop a Pareto optimization approach to ensure balanced utilization of all modalities, enhancing timing prediction accuracy. Comprehensive experiments on large-scale open-source designs using TSMC’s 16 nm technology node validate the superiority of our model over state-of-the-art pre-routing timing prediction methods.
Yuyang Ye 0001, Mingwei He, Lizheng Ren, Jianwang Zhai, Tinghuan Chen, Jun Yang 0006, Longxing Shi
DAC1
2025 Timing-Driven Approximate Logic Synthesis Based on Double-Chase Grey Wolf Optimizer
abstract
With the shrinking technology nodes, timing optimization becomes increasingly challenging. Approximate logic synthesis (ALS) can perform local approximate changes (LACs) on circuits to optimize timing with the cost of slight inaccuracy. However, existing ALS methods that focus solely on critical path depth reduction or area minimization are not optimal in timing optimization. This paper proposes an effective timing-driven ALS framework, where we employ a double-chase grey wolf optimizer to explore and apply LACs, simultaneously bringing excellent critical path shortening and area reduction under error constraints. Subsequently, it utilizes post-optimization under area constraints to convert area reduction into further timing improvement, thus achieving maximum critical path delay reduction. According to experiments on open-source circuits with 28nm technology, compared to the SOTA method, our framework can generate approximate circuits with greater critical path delay reduction under different error and area constraints.
Xiangfei Hu, Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001
DATE2
2025 GraphCAD: Leveraging Graph Neural Networks for Accuracy Prediction Handling Crosstalk-affected Delays
abstract
As chip fabrication technology advances, the capacitive effects between wires have become increasingly pronounced, making crosstalk-induced incremental delay a serious issue. Traditional static timing analysis involves complex and iterative calculations through timing windows, requiring precise alignment of aggressor and victim nets, along with delay and slew estimations, which significantly increase runtime and licensing costs. In our work, we develop a Graph Neural Network framework to predict crosstalk-affected delays, focusing on the impacts of the coupling effect and overlapping nets. Moreover, we employ a curriculum learning strategy that gradually integrates aggressors with victims, improving model convergence through progressively complex scenarios. Experimental results show that our framework precisely predicts crosstalk-affected delays, matching commercial tools' performance with a fivefold speedup.
Fangzhou Liu 0005, Guannan Guo, Yuyang Ye 0001, Ziyi Wang 0010, Wenjie Fu 0003, Weihua Sheng, Bei Yu 0001
ISPD3
2025 ARS-Flow 2.0: An enhanced design space exploration flow for accelerator-rich system based on active learning
Shuaibo Huang, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
Integr.2
2025 An Optimization-Aware Prerouting Timing Prediction Framework Based on Multimodal Learning
abstract
Accurate and efficient prerouting timing estimation is particularly crucial during placement to alleviate time-consuming design iterations. Machine-learning (ML)-based methods have been introduced recently to predict the post-routing timing results at placement stage, but most of them neglect the impact of timing optimization during physical design, suffering from accuracy loss due to inconsistent circuit netlist. In this work, an optimization-aware prerouting timing prediction framework based on multimodal learning is proposed to calibrate the timing changes between placement and routing stages, where the local netlist and layout information are extracted by graph neural network (GNN) and convolutional neural network (CNN), respectively, while the global information along the path is further extracted by Transformer network. Based on the predicted post-routing timing results by the proposed framework, timing optimization guidance is generated to enhance traditional design flow with better physical implementation quality. Experimental results demonstrate that for the OpenCores benchmark circuits under TSMC 22nm process, the proposed framework achieves significant correlation and accuracy improvement with an average of 0.9219 in terms of R2 score and 2.12% of mean absolute percentage error (MAPE) as well as an average runtime acceleration of$645\times $compared with traditional design flow on testing designs. With the timing optimization guidance, significant worst negative slack (WNS) and total negative slack (TNS) improvement are achieved compared with traditional flow after placement and routing, respectively, without noticeable area, power, wire length, and the number of design rule check (DRC) violations increase.
Peng Cao 0002, Yusen Qin, Guoqing He, Zhanhua Zhang, Yuyang Ye 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 Learning-Driven Physically Aware Large-Scale Circuit Gate Sizing
abstract
Gate sizing plays an important role in timing optimization after physical design. Existing machine learning-based gate sizing works cannot optimize timing on multiple timing paths simultaneously and neglect the physical constraint on layouts. They cause suboptimal sizing solutions and low-efficiency issues when compared with commercial gate sizing tools. In this work, we propose a learning-driven physically aware gate sizing framework to optimize timing performance on large-scale circuits efficiently. In our gradient descent optimization-based work, for obtaining accurate gradients, a multimodal gate sizing-aware timing model is achieved via learning timing information on multiple timing paths and physical information on multiple-scaled layouts jointly. Then, gradient generation based on the sizing-oriented estimator and adaptive back-propagation are developed to update gate sizes. Our results demonstrate that our work achieves higher-timing performance improvements in a faster way compared with the commercial gate sizing tool.
Yuyang Ye 0001, Peng Xu 0052, Lizheng Ren, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 A Hybrid Domain and Pipelined Analog Computing Chain for MVM Computation
abstract
In this article, a stream-architecture and pipelined hybrid computing chain is presented to process matrix-vector multiplication (MVM). In each stage of the computing chain, a primary multiply-accumulate (MAC) stage consisting of charge, time, and digital domain processing units makes signed or unsigned$8\times 1\times 8$bit MAC operations and MSB quantization. Based on the stream architecture, the length of the computing chain can be configured to fit different MVM applications. In the charge-domain MAC unit, a double-plate sampling and weighted capacitor array with writing yield and efficiency enhanced 7T bitcell and three-step weighting scheme is implemented. To utilize the speed and resolution advantages of time-domain computing, a high linearity voltage-to-time converter (VTC) followed by a dynamic tristate delay chain is proposed to transfer and store MAC values from the charge domain in the time domain. To realize fast analog readout, a folding type and distributed time-to-digital converter (TDC) is proposed. To fully eliminate the offset and variation in the distributed TDC, a specific residue readout timing and back-end calibration scheme are applied. In the digital domain, a double-input and double-clock dynamic D flip-flop is built to realize partial sum transmission and accumulation in a single cycle with low energy and area consumption. Post-simulation results show that this computing chain can achieve 20.89–40.72-TOPS/W energy efficiency and 4.498-TOPS/mm2 throughput.
Tianzhu Xiong, Yuyang Ye 0001, Xin Si, Jun Yang 0006
IEEE Trans. Very Large Scale Integr. Syst.2
2024 Heterogeneous Graph Attention Network Based Statistical Timing Library Characterization with Parasitic RC Reduction
abstract
Statistical timing characterization for standard cell library poses significant challenge to accuracy and runtime cost. Prior analytical and machine learning-based methods neglect the profound influence induced by layout-dependent parasitic resistor and capacitor (RC) network in cell netlist as well as the timing correlation between topological structures of cells and process, voltage, and temperature (PVT) corners, resulting in tremendous simulation effort and/or poor accuracy. In this work, an accurate and efficient statistical cell timing library characterization framework is proposed based on heterogeneous graph attention network (HGAT) assisted with parasitic RC reduction approach, where the transistors and parasitic RC in cell are represented as heterogeneous nodes for graph learning and redundant RC nodes are removed to alleviate node imbalance issue and improve prediction accuracy. The proposed framework was validated with TSMC 22nm standard cells under multiple PVT corners to predict the standard deviation of cell delay with the error of 2.67% on average for all validated cells in terms of relative Root Mean Squared Error (rRMSE) with $3 \times $ characterization runtime speedup, achieving $2.7 \sim 6.9 \times $ accuracy improvement compared with prior works. The predicted statistical timing libraries were further validated with ISCAS’89 benchmark circuits for statistical static timing analysis (SSTA), where the critical path delay at $3 \sigma$ percentile point is reported with the average mismatch of $1.34 ps$ compared with foundry-provided library, showing $10.7 \sim 14.5 \times $ better accuracy than the competitive approaches.
Yuyang Ye 0001, Guoqing He, Peng Cao 0002
ASPDAC2
2024 An Optimization-aware Pre-Routing Timing Prediction Framework Based on Heterogeneous Graph Learning
abstract
Accurate and efficient pre-routing timing estimation is particularly crucial in timing-driven placement, as design iterations caused by timing divergence are time-consuming. However, existing machine learning prediction models overlook the impact of timing optimization techniques during routing stage, such as adjusting gate sizes or swapping threshold voltage types to fix routing-induced timing violations. In this work, an optimization-aware pre-routing timing prediction framework based on heterogeneous graph learning is proposed to calibrate the timing changes introduced by wire parasitic and optimization techniques. The path embedding generated by the proposed framework fuses learned local information from graph neural network and global information from transformer network to perform accurate endpoint arrival time prediction. Experimental results demonstrate that the proposed framework achieves an average accuracy improvement of 0.10 in terms of R2score on testing designs and brings average runtime acceleration of three orders of magnitude compared with the design flow.
Guoqing He, Yuyang Ye 0001, Peng Cao 0002
ASPDAC3
2024 ARS-Flow: A Design Space Exploration Flow for Accelerator-rich System based on Active Learning
abstract
Surrogate model-based design space exploration (DSE) is the mainstream method to search for optimal microarchitecture designs. However, it is hard to build accurate models for accelerator-rich systems within limited samples due to its high dimensional characteristic. Moreover, it is easy to fall into local optimal or difficult to converge. To solve these two problems, we propose a DSE flow based on active learning, namely ARS-Flow. It is featured with Pareto-region-oriented stochastic resampling method (PRSRS) and multiobjective genetic algorithm with self-adaptive hyperparameter control (SAMOGA). Taking the gem5-SALAM system for illustration, the proposed method can build more accurate models and find better microarchitecture designs with acceptable runtime costs.
Shuaibo Huang, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
ASPDAC2
2024 A Graph-Learning-Driven Prediction Method for Combined Electromigration and Thermomigration Stress on Multi-Segment Interconnects
abstract
As technology advances, the temperature gradient in the interconnects becomes more significant, which causes serious thermomigration. Simulating the coupling effects of thermomigration (TM) and electromigration (EM) on large-scale circuits is very time-consuming caused by a substantial increase in computational complexity. Recently, some researchers utilized graph learning-based methods to predict EM stress in medium-scale cases. Unfortunately, these works overlooked the effects of TM. To predict the EM - TM stress of large-scale interconnects accurately and efficiently, we propose a framework based on Graph Attention Networks (GATs) with a customized alternating aggregation method for collecting information in junctions and branches of interconnects jointly. The experimental results show that our work achieves an average relative error of less than 1 % compared to the commercial software COMSOL for inter-connects consisting of fewer than 200 segments. Furthermore, our method also achieves 9037 x speedup in predicting the OpenROAD test circuit with a maximum segment number reaching 10807.
Yunfan Zuo, Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Longxing Shi
DATE2
2024 RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization
abstract
Electronic Design Automation (EDA) tools are critical in the Very Large Scale Integration (VLSI) flow. To address the challenges posed by the extensive search space and intricate feature interactions, statistical and machine-learning methods have been employed. These methods aim to model tool parameters and treat the tuning process as a regression task. However, these regression-based methods suffer from inaccurate estimations owing to limited training samples. To address this issue, we propose a ranking-based tool parameter tuning framework, called RankTuner, which directly learns the dominant relationship between parameters. RankTuner utilizes a pairwise Gaussian process to estimate the probability and uncertainty of the dominance relationship. Our approach also integrates a Duel-Thompson sampling method to balance exploration and exploitation in parameter selections. A dimensionality reduction scheme with random embedding and trust region techniques is incorporated to enable parallel searches. Experimental results demonstrate the superiority of RankTuner compared to the cutting-edge tool parameter tuning methods.
Peng Xu 0052, Su Zheng, Yuyang Ye 0001, Hao Geng, Tsung-Yi Ho, Bei Yu 0001
ICCAD3
2024 Timing-Driven Technology Mapping Approximation Based on Reinforcement Learning
abstract
As the transistor technology nodes shrink into the nano-scales, timing guardbands caused by aging effects and process variations continue to increase. Approximate computing can eliminate aging-and-variation-induced timing guardbands without sacrificing the design performance. It can apply local approximate changes (LACs) automatically in circuits to reduce critical path delay. However, efficiently achieving timing optimization under error distance constraints is still tricky. This work proposes an automated timing-driven technology mapping approximation framework based on reinforcement learning (RL). The framework uses path-weighted graph neural networks (PGNNs) to embed RL states and timing path-aware LAC candidates to construct RL action spaces. It can efficiently eliminate timing guardbands induced by aging and variation. Our proposed circuit-agnostic framework operates on the gate-level netlists. According to experiments on the open-source circuits using TSMC 28nm and 16nm technology under aging and variation conditions, our framework can achieve an average 24.78% critical path delay reduction under 5 different error distance constraints and 4.83× speedup, compared with a state-of-the-art method.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Graph-Learning-Driven Path-Based Timing Analysis Results Predictor from Graph-Based Timing Analysis
abstract
With diminishing margins in advanced technology nodes, the performance of static timing analysis (STA) is a serious concern, including accuracy and runtime. The STA can generally be divided into graph-based analysis (GBA) and path-based analysis (PBA). For GBA, the timing results are always pessimistic, leading to overdesign during design optimization. For PBA, the timing pessimism is reduced via propagating real path-specific slews with the cost of severe runtime overheads relative to GBA. In this work, we present a fast and accurate predictor of post-layout PBA timing results from inexpensive GBA based on deep edge-featured graph attention network, namely deep EdgeGAT. Compared with the conventional machine and graph learning methods, deep EdgeGAT can learn global timing path information. Experimental results demonstrate that our predictor has the potential to substantially predict PBA timing results accurately and reduce timing pessimism of GBA with maximum error reaching 6.81 ps, and our work achieves an average 24.80× speedup faster than PBA using the commercial STA tool.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
ASP-DAC1
2023 Fast and Accurate Wire Timing Estimation Based on Graph Learning
abstract
Accurate wire timing estimation has become a bottleneck in timing optimization since it needs a long turn-around time using a sign-off timer. The gate timing can be calculated accurately using lookup tables in cell libraries. In comparison, the accuracy and efficiency of wire timing calculation for complex RC nets are extremely hard to trade-off. The limited number of wire paths opens a door for the graph learning method in wire timing estimation. In this work, we present a fast and accurate wire timing estimator based on a novel graph learning architecture, namely GNNTrans. It can generate wire path representations by aggregating local structure information and global relationships of whole RC nets, which cannot be collected with traditional graph learning work efficiently. Experimental results on both tree-like and non-tree nets demonstrate improved accuracy, with the max error of wire delay being lower than 5 ps. In addition, our estimator can predict the timing of over 200K nets in less than 100 secs. The fast and accurate work can be integrated into incremental timing optimization for routed designs.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
DATE1
2023 FPGNN-ATPG: An Efficient Fault Parallel Automatic Test Pattern Generator
abstract
The advanced multi-core technology enables parallel computing to speed up Automatic Test Pattern Generation (ATPG). The main challenge is to solve an increasing number of hard-to-solve faults effectively. In this paper, we develop an efficient parallel computing system for the ATPG program, i.e., FPGNN-ATPG, which is consisted of two parts: graph-neural-networks-based (GNN-based) fault classification and fault-driven deterministic test pattern generator (DTPG). The end-to-end GNN-based classifier can predict fault types with superior accuracy compared with classical machine learning methods. And the fault-driven DTPG can solve different types of faults in parallel without runtime overhead. According to the experimental results on an 8-core machine, our FPGNN-ATPG framework obtains an average of 7.56X speedup while reducing 14.13% pattern count ratio with full 100% fault coverage for 10 industrial instances.
Yuyang Ye 0001, Zonghui Wang, Zun Xue, Hao Yan 0002
ACM Great Lakes Symposium on VLSI1
2023 Optimized matrix ordering of sparse linear solver using a few-shot model for circuit simulation
Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
Integr.2
2023 Aging-Aware Critical Path Selection via Graph Attention Networks
abstract
In advanced technology nodes, aging effects like negative and positive bias temperature instability (NBTI and PBTI) become increasingly significant, making timing closure and optimization more challenging. Unfortunately, conventional critical path (CP) selection tools used in reliability-aware design flow cannot accurately identify CPs under different aging conditions. To address this issue, we propose an aging-aware CP selection flow comprising two parts: 1) critical cell detection and 2) path criticality (PC) computation. We employ graph-attention (GAT) networks to predict the critical cells in the aged circuits, and a PC computation algorithm that takes into account circuit-level and transistor-level parameters to generate PC rank lists. Our experimental results demonstrate that our GAT model outperforms classical machine learning models in detecting critical cells. Additionally, compared with the commercial tool, our aging-aware flow achieves an average accuracy of 99.52%, 98.69%, and 97.20% for top-10%, top-5%, and top-1% path sets, respectively, in five industrial designs subjected to different aging conditions and workloads.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1