Yao Wang 0002

dblp:72/628-2 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0005-3652-0424ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 3D Chiplet Partitioning and Floorplanning Interaction with Vertical Bonding Consideration
abstract
The emerging technologies of 3D chiplet integration offer a promising path to increase functional density and communication efficiency beyond the limitations of traditional 2D designs. Among various methods, hybrid bonding enables high-density vertical connections with reduced parasitics and latency. However, few existing approaches explicitly support this vertical interconnect scenario, and most treat partitioning and floorplanning as separate stages. To fully exploit the benefits of vertical integration, partitioning and floorplanning need to be considered jointly to balance bond demand and bond supply. In this paper, we propose a unified framework for 3D chiplet partitioning and floorplanning under fine-pitch bonding technologies. Experiments on industry-standard benchmarks demonstrate a 10%∼40% reduction in HPWL and over 60% decrease in inter-die vertical connection overflow, confirming the effectiveness and superiority of our approach over current methods.
Mengen Chen, Yao Wang 0002, Yang Guo 0003
DATE4
2025 SI-Aware Wire Timing Prediction at Pre-Routing Stage with Multi-Corner Consideration
abstract
Timing closure is a critical but effort-taking task in VLSI designs. Early design stages have relatively ample room for changes that can fix timing problems in a proactive manner. However, accurate timing prediction is very challenging at early stages due to the absence of information determined by later stages in the design flow. At the pre-routing stage, it is generally believed that the prediction of wire delay is more complicated than that of gate delay, since the former is highly dependent on the routing information and PVT conditions. Addressing that, in this work, the prediction model is studied and the importance of multiple features, including Signal-Integrity (SI) related ones, is explored, with the purpose of boosting the turnaround time of physical design and reducing the performance penalty caused by worst-case scenario assumptions. Experimental results show that the proposed timing predictor has achieved a correlation of over 0.98 with the sign-off timing results in SI-mode under multi-corner scenarios.
Renjun Zhao, Yao Wang 0002, Chang Liu 0019, Yang Guo 0003
ASP-DAC4
2024 A survey of compute nodes with 100 TFLOPS and beyond for supercomputers
Junsheng Chang, Kai Lu 0001, Yang Guo 0003, Yongwen Wang, Libo Huang 0002, Yao Wang 0002, Biwei Zhang
CCF Trans. High Perform. Comput.8
2024 Hierarchical Mapping of Large-Scale Spiking Convolutional Neural Networks Onto Resource-Constrained Neuromorphic Processor
abstract
Neuromorphic processors have been designed as non-von Neumann systems for energy-efficient spiking neural network (SNN) execution. Spiking convolutional neural networks (SCNNs), combining the advantage of convolutional neural network (CNN) and SNN, have been widely applied to vision tasks. However, as the scale of SCNNs increases, executing large-scale SCNNs on resource-constrained neuromorphic processor faces many challenges, including massive synapse pruning caused by resource competition, execution performance degradation, etc. Addressing these problems, we propose an efficient approach to map large-scale SCNNs onto resource-constrined neuromorphic processor. The approach consists of three steps: splitting, partitioning, and mapping. We explore three acyclic splitting strategies to divide large-scale SCNNs into subnetworks without cyclic dependency. Axon sharing is the guiding principle to partition subnetworks into multiple clusters. To obtain an optimal cluster-to-core mapping scheme, we use Non-dominated Sorting Genetic Algorithm to collaboratively optimize two metrics. We evaluate our approach with eight realistic SCNN applications. The results show that compared with existing state-of-the-art methods, our approach significantly reduces the synapse pruning and accuracy loss, and increases the execution performance.
Xun Xiao, Yao Wang 0002, Junbo Tie, Lei Wang 0011, Weixia Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 A General Layout Pattern Clustering Using Geometric Matching-based Clip Relocation and Lower-bound Aided Optimization
abstract
With the continuous shrinking of feature size, detection of lithography hotspots has been raised as one of the major concerns in Design-for-Manufacturability (DFM) of semiconductor processing. Hotspot detection, along with other DFM measures, trades off turnaround time for the yield of IC manufacturing, and thus a simplified but wide-ranging pattern definition is a key to the problem. Layout pattern clustering methods, which group geometrically similar layout clips into clusters, have been vastly proposed to identify layout patterns efficiently. To minimize the clustering number for subsequent DFM processing, in this article, we propose a geometric-matching-based clip relocation technique to increase the opportunity of pattern clustering. Particularly, we formulate the lower bound of the clustering number as a maximum-clique problem, and we have also proved that the clustering problem can be solved by the result of the maximum-clique very efficiently. Compared with the experimental results of the state-of-the-art approaches on ICCAD 2016 Contest benchmarks, the proposed method can achieve the optimal solutions for all benchmarks with very competitive runtime. To evaluate the scalability, the ICCAD 2016 Contest benchmarks are extended and evaluated. Moreover, experimental results on the extended benchmarks demonstrate that our method can reduce the cluster number by 16.59% on average, while the runtime is 74.11% faster on large-scale benchmarks compared with previous works.
Yao Wang 0002, Zhiyong Fu, Yang Guo 0003
ACM Trans. Design Autom. Electr. Syst.2
2023 A Soft-Error Mitigation Approach Using Pulse Quenching Enhancement at Detailed Placement for Combinational Circuits
abstract
As technology continuously shrinks, radiation-induced soft errors have become a great threat to the circuit reliability. Among all the causes, the Single-Event Transient (SET) effect is the dominating one for the radiation-induced soft errors. SET-induced soft errors can be mitigated by multiple methods. In terms of area and power overhead, blocking SET propagation is considered to be the most efficient way for soft error reduction. It is found that the SET pulse width can be shrunk by a pulse quenching effect, which can be utilized to mitigate soft errors without introducing any area and power overhead. In this article, we present an effective detailed placer to exploit the pulse quenching effect for soft error reduction in combinational circuits. In our method, the quenching effect enhancement is globally optimized while the cell displacement is minimized. The experimental results demonstrate that our method reduces the soft error vulnerability of the circuits by 29.53% versus 18.38% of the state-of-the-art solution. Meanwhile, our method has a minimal effect on the displacement and half-perimeter wire length (HPWL) compared with the previous solutions, which means a minimum timing influence to the original design.
Yao Wang 0002, Chang Liu 0019, Qiang Wu 0015, Juan Luo, Yang Guo 0003
ACM Trans. Design Autom. Electr. Syst.2
2023 Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISA
abstract
In recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming.
Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001
IEEE Trans. Parallel Distributed Syst.4
2022 Accurate timing prediction at placement stage with look-ahead RC network
abstract
Timing closure is a critical but effort-taking task in VLSI designs. In placement stage, a fast and accurate net delay estimator is highly desirable to guide the timing optimization prior to routing, and thus reduce the timing pessimism and shorten the design turn-around time. To handle the timing uncertainty at the placement stage, we propose a fast net delay timing predictor based on machine learning, which extract the fully timing features using a look-ahead RC network. Experimental results show that the proposed timing predictor has achieved average correlation over 0.99 with the post-routing sign-off timing results obtained in Synopsys PrimeTime.
Zhiyong Fu, Yao Wang 0002, Chang Liu 0019, Yang Guo 0003
DAC3
2022 Unicorn: a multicore neuromorphic processor with flexible fan-in and unconstrained fan-out for neurons
abstract
Neuromorphic processor is popular due to its high energy efficiency for spatio-temporal applications. However, when running the spiking neural network (SNN) topologies with the ever-growing scale, existing neuromorphic architectures face challenges due to their restrictions on neuron fan-in and fan-out. This paper proposes Unicorn, a multicore neuromorphic processor with a spike train sliding multicasting mechanism (STSM) and neuron merging mechanism (NMM) to support unconstrained fan-out and flexible fan-in of neurons. Unicorn supports 36K neurons and 45M synapses and thus supports a variety of neuromorphic applications. The peak performance and energy efficiency of Unicorn reach 36TSOPS and 424GSOPS/W respectively. Experimental results show that Unicorn can achieve 2×-5.5× energy reduction over the state-of-the-art neuromorphic processor when running an SNN with a relatively large fan-out and fan-in.
Lei Wang 0011, Yao Wang 0002, LingHui Peng, Xun Xiao, Weixia Xu 0001
DAC3
2022 Dynamic Vision Sensor Based Gesture Recognition Using Liquid State Machine
Xun Xiao, Lei Wang 0011, Lianhua Qu, Shasha Guo 0001, Yao Wang 0002, Ziyang Kang
ICANN (3)6
2020 Maximum Clique Based Method for Optimal Solution of Pattern Classification
abstract
As the transistor feature size continuously shrinks, design for manufacturability (DFM) has become a crucial concern. Layout pattern classification, which groups geometrically similar layout clips into clusters, has been widely utilized in a variety of DFM applications, such as hotspot library generation, hierarchical data storage, and systematic yield optimization. In this paper, we have proposed a maximum clique based method to obtain the lower bound of the clustering count and have proven that the lower bound is the theoretical optimal solution. To solve the clustering problem, we formulate it as a Set-Covering Problem (SCP) and utilize the result of the maximum clique to help the SCP quickly converge. Compared with the experimental results of the state-of-the-art approaches on ICCAD 2016 Contest benchmarks, our proposed method can achieve optimal solutions for all benchmarks with an approximate minimum run-time.
Zhiyong Fu, Yao Wang 0002, Yang Guo 0003
ICCD4
2020 Lithography Hotspot Detection with FFT-based Feature Extraction and Imbalanced Learning Rate
abstract
With the increasing gap between transistor feature size and lithography manufacturing capability, the detection of lithography hotspots becomes a key stage of physical verification flow to enhance manufacturing yield. Although machine learning approaches are distinguished for their high detection efficiency, they still suffer from problems such as large-scale layout and class imbalance. In this article, we develop a hotspot detection model based on machine learning with high performance. In the proposed model, we first apply an Fast Fourier Transform--based feature extraction method that can compress large-scale layout to a multi-dimensional representation with much smaller size while preserving the discriminative layout pattern information to improve the detection efficiency. Second, addressing the class imbalance problem, we propose a new technique called imbalanced learning rate and embed it into the convolutional neural network model to further reduce false alarms without accuracy decay. Compared with the results of current state-of-the-art approaches on ICCAD 2012 Contest benchmarks, our proposed model can achieve better solutions in many evaluation metrics, including the official metrics.
Shizhe Zhou, Rui Li 0019, Yao Wang 0002, Yang Guo 0003
ACM Trans. Design Autom. Electr. Syst.5
2017 A Mixed-Size Monolithic 3D Placer with 2D Layout Inheritance
abstract
Monolithic 3D IC is a high integration density emerging technology in the age of both "More Moore" and "More-than-Moore". In this paper, we propose a novel method of generating mixed-size 3D placement based on transforming a 2D placement result. Experimental results indicate that, when compared with the input 2D placement, the 3D placer can reduce the wirelength by 57%, and provide a four-layer 3D chip footprint of about one quarter of the 2D counterpart. Moreover, our placer can preserve the layout information from the 2D placement input, which means that the 2D placement quality can be inherited in the 3D placement results. Compared with an analytical wirelength-driven placer, our placer achieves 34% benefit on 2D layout inheritance and 12% benefit on runtime with acceptable (4%) wirelength cost.
Yao Wang 0002, Yang Guo 0003, Sorin Cotofana
ACM Great Lakes Symposium on VLSI2
2016 Ripple 2.0: Improved Movement of Cells in Routability-Driven Placement
abstract
Routability is one of the most important problems in high-performance circuit designs. From the viewpoint of placement design, two major factors cause routing congestion: (i) interconnections between cells and (ii) connections on macro blockages. In this article, we present a routability-driven placer, Ripple 2.0, which emphasizes both kinds of routing congestion. Several techniques will be presented, including (i) cell inflation with routing path consideration, (ii) congested cluster optimization, (iii) routability-driven cell spreading, and (iv) simultaneous routing and placement for routability refinement. With the official evaluation protocol, Ripple 2.0 outperforms other published academic routability-driven placers. Compared with top results in the ICCAD 2012 contest, Ripple 2.0 achieves a better detailed routing solution obtained by a commercial router.
Yao Wang 0002, Yang Guo 0003, Evangeline F. Y. Young
ACM Trans. Design Autom. Electr. Syst.2
2014 Analysis of the impact of spatial and temporal variations on the stability of SRAM arrays and the mitigation technique using independent-gate devices
Yao Wang 0002, Sorin Cotofana
J. Parallel Distributed Comput.1
2013 Lifetime reliability assessment with aging information from low-level sensors
abstract
Aggressive technology scaling has led Integrated Circuits (ICs) suffer from ever-increasing wearout effects. As a consequence, Dynamic Reliability Management (DRM) becomes an essential approach to assure IC's lifetime reliability. Accurate and efficient reliability modeling from low-level aging sensor measurements is critical to DRM systems. This work presents a Time-Sharing Sensing (TSS) method for $V_{th}$-sensor based DRM to assess the dynamic NBTI-induced degradation experienced by the circuit under monitoring. SPICE simulation results suggest that the proposed TSS method can accurately capture the circuit reliability status under random stress conditions.
Yao Wang 0002, Sorin Cotofana
ACM Great Lakes Symposium on VLSI1
2013 A direct measurement scheme of amalgamated aging effects with novel on-chip sensor
abstract
Aggressive technology scaling has led to a significant reduction of device reliability. As a consequence Integrated Circuits (ICs) reliability became a major issue and Dynamic Reliability Management (DRM) schemes have been proposed to assure ICs' lifetime reliability. Though, up to date, various aging sensors have been proposed, few of them can provide real quantitative aging measurements. In view of this, we propose a direct measuring scheme by using the drain current as aging indicator. We designed a novel on-chip aging sensor able to detect the amalgamated aging effects of ICs caused by joint failure mechanisms. This is achieved by detecting the peak power supply current (Ipp) degradation from the device and/or circuit, which is a signature of the total drain current. Unlike the existing aging sensors which indirectly estimate the aging status of a device, the proposed sensor allows for direct aging assessment for single device and/or circuit blocks. Simulation results using the TSMC 65nm technology indicate that the proposed sensor can operate at 1GHz. Accelerated test simulation in Cadence for a set of ISCAS85 benchmark circuits indicates that the drain current exhibits a similar aging rate as the threshold voltage for the entire circuit lifetime, but with a better sensitivity towards the End-of-Life (EOL), which demonstrates the validity and practical relevance of the proposed aging monitoring framework.
Nicoleta Cucu Laurenciu, Yao Wang 0002, Sorin Cotofana
VLSI-SoC2