Rongjian Liang

dblp:257/8611 · DBLP profile ↗
← Back
34ranked-venue papers
12as first author
32since 2021 · last 2026
0000-0001-8626-2359ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 12 first-author · 32 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 C3PO: Commercial-Quality Global Placement via Coherent, Concurrent Timing, Routability, and Wirelength Optimization
abstract
Despite achieving orders-of-magnitude runtime speedup, GPU-accelerated placers (GPU-Placers) still have extremely limited industrial adoption, largely due to the wide gaps in Power, Performance, and Area (PPA) metrics compared to those well-established CPU-centric commercial Physical Design (PD) tools. To overcome this issue, we introduce C3PO, the first commercial-quality, differentiable, multi-objective global placer that performs concurrent timing, routability, and wirelength optimization in a coherent manner with custom CUDA kernels. Particularly, we propose a convex-based framework that dynamically computes objective weights at each placement iteration by solving a quadratic problem, eliminating the need of manual parameter tuning. In the experiments, we rigorously validate C3PO with an industry-leading commercial PD tool and demonstrate that on 8 designs from TILOS [1] and IWLS [2] in ASAP 7nm [3], C3PO consistently outperforms the commercial tool by up to 16.7% in routed wirelength and 19.6% in switching power with complete full-flow validation.
Yi-Chen Lu, Hao-Hsiang Hsiao, Rongjian Liang, Haoxing Ren
ASP-DAC3
2026 LegoMap: Optimization for High-Throughput Transformer Computing on AI Engine-Based FPGAs
Hailiang Hu, Haodong Chang, Donghao Fang, Zhenrui Wang, Wuxi Li, Rongjian Liang, Bo Yuan 0001, Jiang Hu 0001
FCCM6
2026 MapTune: Versatile ASIC Technology Mapping via Reinforcement Learning Guided Library Tuning
abstract
Technology mapping involves mapping logical circuits to a library of standard cells. Traditionally, a full technology library is used, leading to a large search space and potential runtime overhead. Motivated by randomly sampled technology mapping case studies, we propose MapTune to address this challenge by utilizing reinforcement learning to make design-specific cell selection choices. By learning from the environment and guided by the reward, MapTune refines the cell selection process, resulting in a reduced search space and potentially improved mapping quality. The effectiveness of MapTune is evaluated on a wide range of benchmarks, different technology libraries, and various technology mappers. The empirical results demonstrate that MapTune achieves higher mapping accuracy and reduces delay/area across various circuit designs, technology libraries, and mappers. The article also discusses the Pareto-Optimal exploration and confirms the perpetual delay-area tradeoff. Conducted on benchmark suites ISCAS 85/89, ITC/ISCAS 99, VTR8.0, and EPFL benchmarks, the post-technology mapping and post-sizing quality-of-results (QoR) have been significantly improved, with average Area-Delay Product (ADP) improvement of 16.56% among all different exploration settings in MapTune. The improvements consistently remained for four different technologies (7 nm, 45 nm, 130 nm, and 180 nm) with various mappers including both state-of-the-art open-source and commercial synthesis tools.
Mingju Liu, Daniel Robinson, Johannes Maximilian Kühn, Rongjian Liang, Haoxing Ren, Cunxi Yu
ACM Trans. Design Autom. Electr. Syst.5
2025 DCO-3D: Differentiable Congestion Optimization in 3D ICs
abstract
State-of-the-art 3D IC flows fail to consider 3D congestion during earlier stages, leading to excessive use of end-of-flow ECO resources for routability correction that severely degrades full-chip Power, Performance, and Area metrics. We present DCO-3D, a Machine Learning-based routability-aware 3D PD flow that performs early post-route congestion prediction using Siamese Networks and resolves the predicted hotspots using a fully differentiable 3D cell spreading with Graph Neural Network. On 6 industrial designs in a commercial 3nm node, DCO-3D improves Pin-3D, the known best Pin-3D flow, by up to 47.2% in overflow, 86.2% in TNS and 5.1% in power at signoff.
Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Anthony Agnesina, Rongjian Liang, Yuan-Hsiang Lu, Haoxing Ren, Sung Kyu Lim
DAC5
2025 Reinforcement Learning-Driven Window Selection for Enhanced Window-Based Rip-up and Reroute in Chip Detailed Routing
abstract
With increasingly complex design rules and pin density in advanced technology nodes, achieving a violation-free layout has become more challenging, also making rip-up and reroute (RUR) the most runtime-intensive component of detailed routing. We propose a novel reinforcement learning (RL)based approach to enhance the window-based RUR process. Our method features a dynamic window generation strategy that adjusts window size and position based on the distribution of design rule violations (DRV), enabling efficient targeting of congested areas. By leveraging the predictive capabilities of RL, our approach aims to minimize DRVs and achieve high-quality routing results. Experimental results demonstrate that our method outperforms the state-of-theart detailed routers, TritonRoute, achieving a DRV-free solution, averagely improving wirelength by 0.07%, via count by 2.42%, and consuming almost the same average runtime.
Yu-Chan Keng, Yu-Chun Pai, Wen-Hao Liu 0001, Haoxing Ren, Danny Liu, Rongjian Liang, Mark Ho, Anthony Agnesina, Yih-Lang Li
DAC6
2025 INSTA: An Ultra-Fast, Differentiable, Statistical Static Timing Analysis Engine for Industrial Physical Design Applications
abstract
Prior GPU-accelerated Static Timing Analysis (GPU-STA) works all struggle to find industrial adoption, primarily because they aim to build standalone timing engines that can never emulate the proprietary delay models used in commercial tools. In this paper, we adopt a different philosophy by presenting INSTA, the first-ever differentiable, statistical GPU-STA engine that achieves unprecedented accuracy and scalability by a one-time initialization from any reference tool, bringing two transformative capabilities to Physical Design (PD): (1) rapid, high-fidelity timing analysis for incremental netlist update, and (2) gradient-based truly-global timing optimization at scale. Notably, INSTA demonstrates a near-perfect 0.999 correlation with an industryleading signoff tool on a 15 -million-pin design in a commercial 3 nm node with runtime under 0.1 seconds. Experimental results showcase INSTA’s capability through three PD applications: (1) serving as a fast evaluator in an industrial gate sizing flow, achieving $\mathbf{2 5 x}$ faster incremental update_timing runtime with almost no accuracy loss; (2) INSTA-Size, a gradient-based gate sizer that achieves up to $\mathbf{1 5 \%}$ better Total Negative Slack (TNS) than the reference signoff engine by sizing $68 \%$ fewer amount of cells; and (3) INSTA-Place, a differentiable timingdriven global placer that outperforms the state-of-the-art net-weighting placer by up to 16% in Half-Perimeter Wirelegnth (HPWL) and 59.4% in TNS on the ICCAD’15 benchmark [15].
Yi-Chen Lu, Zhizheng Guo, Kishor Kunal, Rongjian Liang, Haoxing Ren
DAC4
2025 Hybrid Exact and Heuristic Efficient Transistor Network Optimization for Multi-Output Logic
abstract
With the approaching post-Moore era, it is becoming increasingly impractical to decrease the transistor size in digital VLSI for better performance. To address this issue, one approach is to optimize the digital circuit at the transistor level to reduce the transistor count. Although previous works have explored ways to conduct transistor network optimization, most of these efforts have focused on single-output networks or applied heuristics only, limiting their scope or optimization quality. In this paper, we propose an exact transistor network optimization algorithm that supports multi-output logic and is formulated as a SAT problem. Our approach maintains a high optimization level by employing the exact algorithm, while also incorporating a hybrid process that uses a heuristic algorithm to predict the solution range as a guidance for better efficiency. Experimental results show that the proposed algorithm has a 5.32% better optimization level given 54% less runtime compared with the state-of-the-art work.
Lang Feng 0001, Rongjian Liang, Hongxin Kong
DATE2
2025 Invited Paper: 2025 ICCAD CAD Contest Problem C: Incremental Placement Optimization Beyond Detailed Placement: Simultaneous Gate Sizing, Buffering, and Cell Relocation
abstract
Late-stage placement optimization is where real PPA trade-offs surface, and where conventional heuristic passes tend to get trapped in small, local neighborhoods. We frame an invited "Problem C" contest that treats this stage as a global, multi-operator search over gate sizing, buffer/inverter-pair insertion, and legal cell relocation, with strict reproducibility and legality. Our core belief grounded in production experience is that GPU batching and differentiable guidance expand the tractable search space: you can score and steer thousands of coordinated moves per iteration, not just a handful, and do so under tight runtime budgets. Submissions must produce a replayable ECO changelist and a final legal DEF; a standardized evaluation flow computes timing, power, and wirelength and combines them with displacement and runtime into the contest score. The specification is designed to encourage pragmatic use of gradient signals and tensorized batching without mandating any single method, enabling participants to leverage novel GPU tools to deliver industrially deployable PPA gains.
Yi-Chen Lu, Rongjian Liang, Wen-Hao Liu 0001, Haoxing Ren
ICCAD2
2025 LLM4Verilog: Building Large-Scale, High-Quality Data Infrastructure for Verilog Code Generation via Community Efforts
abstract
Despite recent advancements in code generation with large language models (LLMs), generating hardware code such as Verilog remains a significant challenge due to the scarcity of large-scale, high-quality datasets in the hardware domain. Existing approaches, including scraping open-source repositories and relying on manually curated datasets, often suffer from limited diversity, quality, and scalability. To address these limitations, we introduce LLM4Verilog, an exploratory, collaborative initiative aimed at constructing a large-scale, high-quality, open-source Verilog dataset. Our initiative integrates a community-driven data collection pipeline with a two-stage data filtering technique to ensure high dataset quality. The first stage removes duplicates and low-quality samples, resulting in a large-scale dataset called LLM4Verilog-complete. The second stage applies an LLM-driven quality scoring method, VeriScore, to perform fine-grained filtering and produce a high-quality, ready-to-use dataset called LLM4Verilog-filtered. We evaluate the effectiveness of these datasets through fine-tuning three different LLMs on our dataset, achieving 6.6%~11.2% and 5.3%~13.2% improvements in pass@1 scores on VerilogEval-human and VerilogEval-Machine, respectively, compared to models fine-tuned with prior state-of-the-art datasets. Notably, these improvements are achieved without relying on complex fine-tuning or data augmentation techniques, highlighting our dataset’s strong potential as a foundational resource for enhancing LLMs’ Verilog code generation capabilities. For more information about our initiative and resulting dataset, please refer to https://nvlabs.github.io/LLM4HWDesign/.
Zhongzhi Yu, Chaojian Li, Yongan Zhang, Nathaniel Ross Pinckney, Wenfei Zhou, Rongjian Liang, Haoxing Ren, Yingyan (Celine) Lin
ICCAD7
2025 GOALPlace: Begin with the End in Mind
abstract
Co-optimizing placement with congestion is integral to achieving high-quality designs. This paper presents GOALPlace, a learning-based approach to improving placement congestion by controlling cell density. It efficiently learns from an EDA tool's post-route optimized results and uses an empirical Bayes technique to adapt the target to a specific placer's solutions, effectively beginning with the end in mind. Our method enhances correlation with the tool's router and timing-opt engine, while solving placement globally without expensive incremental congestion estimation and mitigation methods. A statistical analysis with hierarchical netlist clustering establishes the importance of density and the potential for an adequate cell density target across placements. Our experiments show that our method, when integrated into an academic GPU-accelerated global placer, consistently produces macro and standard cell placements that match or exceed the quality of commercial tools. Our empirical Bayes methodology also shows a substantial quality improvement over leading academic mixed-size placers, achieving up to 10× fewer design rule check (DRC) violations, a 5% decrease in wirelength, and a 30% and 60% reduction in worst and total negative slack (WNS/TNS).
Anthony Agnesina, Rongjian Liang, Geraldo Pradipta, Anand Rajaram, Haoxing Ren
ISPD2
2025 Invited: ISPD 2025 Performance-Driven Large Scale Global Routing Contest
abstract
Global routing is a critical aspect of VLSI design, significantly impacting timing, power consumption, and routability. The ISPD2024 contest focused on addressing the scalability challenges of global routing by leveraging GPU and machine learning techniques. Building on this foundation, the ISPD2025 contest introduces several important updates to better reflect real-world routing challenges. These updates include the provision of industry-standard input files for more precise modeling and integration with OpenROAD for accurate performance assessment. Collectively, these updates aim to bring the contest closer to practical routing scenarios, fostering the development of scalable and efficient solutions for large-scale chip designs.
Rongjian Liang, Anthony Agnesina, Wen-Hao Liu 0001, Matt Liberty, Hsin-Tzu Chang, Haoxing Ren
ISPD1
2025 LEGO-Size: LLM-Enhanced GPU-Optimized Signoff-Accurate Differentiable VLSI Gate Sizing in Advanced Nodes
abstract
On-Chip Variation (OCV)-aware and Path-Based Analysis (PBA) accurate timing optimization achieved by gate sizing (including Vth-assignment) remains a pivotal step in modern signoff. However, in advanced nodes (e.g., 3nm), commercial tools often yield suboptimal results due to the intricate design demands and the vast choices of library cells that require substantial runtime and computational resources for exploration. To address these challenges, we introduce LEGO-Size, a generative framework that harnesses the power of Large Language Models (LLMs) and GPU-accelerated differentiable techniques for efficient gate sizing. LEGO-Size introduces three key innovations. First, it considers timing paths as sequences of tokenized library cells, casting gate sizing prediction as a language modeling task and solving it with self-supervised learning and supervised fine-tuning. Second, it employs a Graph Transformer (GT) with a linear-complexity attention mechanism for netlist encoding, enabling LLMs to make sizing decisions from a global perspective. Third, it integrates a differentiable Static Timing Analysis (STA) engine to refine LLM-predicted gate size probabilities by directly optimizing Total Negative Slack (TNS) through gradient descent. Experimental results on 5 unseen million-gate industrial designs in a commercial 3nm node show that LEGO-Size achieves up to 125x speed up with 37% TNS improvement over an industry-leading commercial signoff tool with minimal power and area overhead.
Yi-Chen Lu, Kishor Kunal, Geraldo Pradipta, Rongjian Liang, Ravikishore Gandikota, Haoxing Ren
ISPD4
2025 Crane: Inter-Layer Scheduling Framework for DNN Inference and Training Co-Support on Tiled Architecture
abstract
Tiled architectures have emerged as a compelling platform for scaling deep neural network (DNN) execution, offering both compute density and communication efficiency.To harness their full potential, effective inter-layer scheduling is crucial for managing operation order, memory behavior, and compute resource coordination.However, current schedulers often fall short due to three persistent issues: incomplete treatment of core design factors, limited flexibility in handling diverse workload structures, and reliance on heuristic search algorithms with poor convergence.In this work, we trace these limitations to the absence of a unified and expressive scheduling representation.We introduce Crane, a framework that addresses these gaps through a hierarchical tableformat abstraction capable of encoding rich scheduling semantics.Crane supports both inference and training workloads, and reformulates scheduling as a mathematically structured optimization problem, enabling more complete and efficient exploration of the scheduling space.Evaluations show that Crane reduces energydelay product by up to 21.01× and improves scheduling speed by at least 2.82× over state-of-the-art baselines.
Yu Gong 0003, Lingyi Huang, Haodong Chang, Rongjian Liang, Cheng Yang 0013, Zhexiang Tang, Jiang Hu 0001, Bo Yuan 0001
MICRO4
2025 DiMO-CNN: Deep Learning Toolkit-Accelerated Analytical Modeling and Optimization of CNN Hardware and Dataflow
abstract
The growing complexity of CNNs demands both hardware acceleration design and dataflow mapping solutions. The large co-design solution space presents a huge challenge. We introduce an analytical model for assessing CNN hardware design and dataflow solutions, using a matrix-based approach. Our co-optimization method, combining nonlinear programming and parallel local search, excels in addressing the power-performance-area tradeoff. The average relative error of our analytical model compared with Timeloop is as small as 1%. Compared to state-of-the-art methods, our co-optimization achieves solutions with average$3.14\times $shorter inference latency,$\mathbf {68.2\%}$less power consumption, and$\mathbf {74\%}$less area on all testcases. It also provides a$200\times $speedup of optimization runtime.
Jianfeng Song, Rongjian Liang, Bo Yuan 0001, Jiang Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 DGR: Differentiable Global Router
abstract
Modern VLSI design flows necessitate fast and high-quality global routers. In this paper, we introduce DGR, a differentiable global router capable of concurrent optimization for hundreds of thousands of nets 1. Our innovation lies in the development of a routing Directed Acyclic Graph (DAG) forest to represent the 2D pattern routing space for all nets, enabling coordinated selection of Steiner trees and 2-pin routing paths from a global perspective. For efficient search within the DAG forest, we relax the discrete search space to be continuous and develop a differentiable solver accelerated by deep learning toolkits on GPUs. Experimental results demonstrate that DGR substantially mitigates routing overflow while concurrently reducing total wirelengths from 0.95% to 4.08% and via numbers from 1.28% to 2.54% in congested testcases compared to state-of-the-art academic global routers. Additionally, DGR exhibits favorable scalability in both runtime and memory with respect to the number of nets.
Wei Li 0159, Rongjian Liang, Anthony Agnesina, Chia-Tung Ho, Anand Rajaram, Haoxing Ren
DAC2
2024 DiMO-Sparse: Differentiable Modeling and Optimization of Sparse CNN Dataflow and Hardware Architecture
abstract
Many real-world CNNs exhibit sparsity, a characteristic that has primarily been utilized in manual design processes and has received little attention in existing automatic optimization techniques. To the best of our knowledge, this paper presents the first systematic investigation of automatic dataflow and hardware optimization for sparse CNN computation. A differentiable PPA (Power Performance Area) model incorporating stochastic modeling of sparse CNN workloads is developed to enable fast nonlinear optimization solving and massively parallel local search-based discretization. Experimental results on public domain testcases demonstrate the efficacy of the proposed approach, achieving an average of 5× and 10× better PPA than the previous work for two different sparsity patterns.
Jianfeng Song, Rongjian Liang, Yu Gong 0003, Bo Yuan 0001, Jiang Hu 0001
DATE2
2024 2024 ICCAD CAD Contest Problem C: Scalable Logic Gate Sizing Using ML Techniques and GPU Acceleration
abstract
Logic gate sizing plays a vital role in timing optimization, especially as Moore's Law slows, shifting greater responsibility to EDA tools to enhance power, performance, and area (PPA), as these gains are no longer achieved solely through scaling and process advancements. There is an increasing need to push the limits of logic gate sizing to extract every possible improvement in PPA. With recent breakthroughs in machine learning (ML) and the computational power of GPUs, there is significant potential to elevate logic gate sizing algorithms to new heights. This contest aims to advance logic gate sizing and push the boundaries of PPA improvement through innovative EDA tools that leverage machine learning and GPU acceleration. As part of the contest, an infrastructure has been developed to enable ML and GPU-accelerated logic gate sizing algorithms, including the release of benchmarks in both standard EDA and ML-friendly formats, along with examples of incorporating "ML inside" EDA tools through Python APIs. The contest leverages the open-source EDA tool OpenROAD and ML-friendly data representation format, CircuitOps, to lower barriers to entry by providing accessible formats and tools, allowing participants to build on existing software without redundancy. With over 25 teams actively participating, the contest highlights growing interest and potential to push the boundaries of timing optimization.
Bing-Yue Wu, Rongjian Liang, Geraldo Pradipta, Anthony Agnesina, Haoxing Ren, Vidya A. Chhabria
ICCAD2
2024 Invited Paper: LLM4HWDesign Contest: Constructing a Comprehensive Dataset for LLM-Assisted Hardware Code Generation with Community Efforts
abstract
Large Language Models (LLMs) show promise in streamlining hardware design, particularly in hardware code generation. However, the development of LLMs for this domain is severely hindered by the scarcity of large-scale, high-quality, and publicly accessible hardware code datasets. This shortage limits the effective fine-tuning of LLMs, impeding their ability to acquire hardware domain knowledge and generate practical designs. To address this challenge, we have organized the first-of-its-kind LLM4HWDesign contest, a community-driven initiative aimed at constructing a large-scale, high-quality dataset for hardware code generation. The contest adopts a two-phase approach, focusing on expanding the scale and quality of an existing hardware code generation dataset, respectively. By harnessing the collective efforts of the hardware design community, the LLM4HWDesign contest seeks to establish a critical resource for advancing LLM-assisted hardware design workflows. The primary goal of this initiative is to deliver a comprehensive dataset compiled from participants' submissions. We hope the released dataset will significantly advance the field of LLM-assisted hardware design and provide substantial benefits to the broader hardware community.
Zhongzhi Yu, Chaojian Li, Yongan Zhang, Nathaniel Ross Pinckney, Wenfei Zhou, Rongjian Liang, Haoxing Ren, Yingyan (Celine) Lin
ICCAD8
2024 GPU/ML-Enhanced Large Scale Global Routing Contest
abstract
Modern VLSI design flows demand scalable global routing techniques applicable across diverse design stages. In response, the ISPD 2024 contest pioneers the first GPU/ML-enhanced global routing competition, selecting advancements in GPU-accelerated computing platforms and machine learning techniques to address scalability challenges. Large-scale benchmarks, containing up to 50 million cells, offer test cases to assess global routers' runtime and memory scalability. The contest provides simplified input/output formats and performance metrics, framing global routing challenges as mathematical optimization problems and encouraging diverse participation. Two sets of evaluation metrics are introduced: the primary one concentrates on global routing applications to guide post-placement optimization and detailed routing, focusing on congestion resolution and runtime scalability. Special honor is given based on the second set of metrics, placing additional emphasis on runtime efficiency and aiming at guiding early-stage planning.
Rongjian Liang, Anthony Agnesina, Wen-Hao Liu 0001, Haoxing Ren
ISPD1
2024 MedPart: A Multi-Level Evolutionary Differentiable Hypergraph Partitioner
abstract
State-of-the-art hypergraph partitioners, such as hMETIS, usually adopt a multi-level paradigm for efficiency and scalability. However, they are prone to getting trapped in local minima due to their reliance on refinement heuristics and overlooking global structural information during coarsening. SpecPart, the most advanced academic hypergraph partitioning refinement method, improves partitioning by leveraging spectral information. Still, its success depends heavily on the quality of initial input solutions. This work introduces MedPart, a multi-level evolutionary differentiable hypergraph partitioner. MedPart follows the multi-level paradigm but addresses its limitations by using fast spectral coarsening and introducing a novel evolutionary differentiable algorithm to optimize each coarsening level. Moreover, by analogy between hypergraph partitioning and deep graph learning, our evolutionary differentiable algorithm can be accelerated with deep graph learning toolkits on GPUs. Experiments on public benchmarks consistently show MedPart outperforming hMETIS and achieving up to a 30% improvement in cut size for some benchmarks compared to the best-published solutions, including those from SpecPart---moreover, MedPart's runtime scales linearly with the number of hyperedges.
Rongjian Liang, Anthony Agnesina, Haoxing Ren
ISPD1
2024 OpenROAD and CircuitOps: Infrastructure for ML EDA Research and Education
abstract
Traditional electronic design automation (EDA) techniques struggle to fulfill the stringent efficiency and quick turnaround demands of complex integrated systems. Machine learning (ML) strategies for EDA (“ML EDA”) are pivotal in transforming EDA to address these challenges. However, they encounter significant obstacles due to inadequate infrastructure, ranging from datasets to software interfaces. This paper demonstrates a software infrastructure for ML EDA built on two key technologies: (i) OpenROAD’s Python APIs, and (ii) NVIDIA’s CircuitOps, an EDA data representation format tailored for ML, facilitating ML EDA applications. The paper illustrates three ML EDA examples that utilize the established OpenROAD and CircuitOps infrastructure.
Vidya A. Chhabria, Wenjing Jiang, Andrew B. Kahng, Rongjian Liang, Haoxing Ren, Sachin S. Sapatnekar, Bing-Yue Wu
VTS4
2023 BufFormer: A Generative ML Framework for Scalable Buffering
abstract
Buffering is a prevalent interconnect optimization technique to help timing closure and is often performed after placement. A common buffering approach is to construct a Steiner tree and then buffers are inserted on the tree based on Ginneken-Lillis style algorithm. Such an approach is difficult to scale with large nets. Our work attempts to solve this problem with a generative machine-learning (ML) approach without Steiner tree construction. Our approach can extract and reuse knowledge from high quality samples and therefore has significantly improved scalability. A generative ML framework, BufFormer, is proposed to construct abstract tree topology while simultaneously determining buffer sizes & locations. A baseline method, FLUTE-based Steiner tree construction followed by Ginneken-Lillis style buffer insertion, is implemented to generate training samples. After training, BufFormer can produce solutions for unseen nets highly comparable to baseline results with a correlation coefficient 0.977 in terms of buffer area and 0.934 for driver-sink delays. On average, BufFormer-generated tree achieves similar delays with slightly larger buffer area. And up to 160X speedup can be achieved for large nets when running on a GPU over the baseline on a single CPU thread.
Rongjian Liang, Siddhartha Nath, Anand Rajaram, Jiang Hu 0001, Haoxing Ren
ASP-DAC1
2023 Late Breaking Results: Test Selection For RTL Coverage By Unsupervised Learning From Fast Functional Simulation
abstract
Functional coverage closure is an important but RTL simulation intensive aspect of constrained random verification. To reduce these computational demands, we propose test selection for functional coverage via machine learning (ML) based anomaly detection in the structural coverage space of fast functional simulators. We achieve promising results on two units from a state-of-the-art production GPU design. With our approach, an up to 85% RTL simulation runtime reduction can be achieved when compared to baseline constrained random test selection while achieving the same RTL functional coverage.
Rongjian Liang, Nathaniel Ross Pinckney, Yuji Chai, Haoxin Ren, Brucek Khailany
DAC1
2023 Invited Paper: CircuitOps: An ML Infrastructure Enabling Generative AI for VLSI Circuit Optimization
abstract
An innovative ML infrastructure named CircuitOps is developed to streamline dataset generation and model inference for various generative AI (GAI)-based circuit optimization tasks. Addressing the challenges of the absence of a shared Intermediate Representation (IR), steep EDA learning curves, and AI-unfriendly data structures, we propose solutions that empower efficient data handling. Our contributions encompass the following: (1) labeled property graphs (LPGs) as IR for flexible netlist representation and efficient parallel processing; (2) tools-agnostic IR generation from standard EDA files; (3) customizable dataset generation facilitated through AI-friendly LPGs; (4) gRPC-based inference deployment. Compared with using Tcl interfaces of EDA design tools, CircuitOps achieves a significant 99× dataset generation speedup and 75K nets per second transfer throughput, validating its effectiveness in optimizing GAI tasks.
Rongjian Liang, Anthony Agnesina, Geraldo Pradipta, Vidya A. Chhabria, Haoxing Ren
ICCAD1
2022 Mapping Large Scale Finite Element Computing on to Wafer-Scale Engines
abstract
The finite element method has wide applications and often presents a computing challenge due to huge problem sizes and slow convergence rate. A leading-edge computing acceleration approach is to leverage wafer-scale engine, which contains more than 800K processing elements. The effectiveness of this approach heavily depends on how to map a finite element computing task onto such enormous hardware space. A mapping method is introduced to partition an object space into computing kernels, which are further placed onto processing elements. This method achieves the best overall result in terms of computing accuracy and communication cost among all the ISPD 2021 contest participants.
Yishuang Lin, Rongjian Liang, Hailiang Hu, Jiang Hu 0001
ASP-DAC2
2022 Deep Learning Toolkit-Accelerated Analytical Co-Optimization of CNN Hardware and Dataflow
abstract
The continuous growth of CNN complexity not only intensifies the need for hardware acceleration but also presents a huge challenge. That is, the solution space for CNN hardware design and dataflow mapping becomes enormously large besides the fact that it is discrete and lacks a well behaved structure. Most previous works either are stochastic metaheuristics, such as genetic algorithm, which are typically very slow for solving large problems, or rely on expensive sampling, e.g., Gumbel Softmax-based differentiable optimization and Bayesian optimization. We propose an analytical model for evaluating power and performance of CNN hardware design and dataflow solutions. Based on this model, we introduce a co-optimization method consisting of nonlinear programming and parallel local search. A key innovation in this model is its matrix form, which enables the use of deep learning toolkit for highly efficient computations of power/performance values and gradients in the optimization. In handling power-performance tradeoff, our method can lead to better solutions than minimizing a weighted sum of power and latency. The average relative error of our model compared with Timeloop is as small as 1%. Compared to state-of-the-art methods, our approach achieves solutions with up to 1.7 × shorter inference latency, 37.5% less power consumption, and 3 × less area on ResNet 18. Moreover, it provides a 6.2 × speedup of optimization runtime.
Rongjian Liang, Jianfeng Song, Bo Yuan 0001, Jiang Hu 0001
ICCAD1
2022 A Stochastic Approach to Handle Non-Determinism in Deep Learning-Based Design Rule Violation Predictions
abstract
Deep learning is a promising approach to early DRV (Design Rule Violation) prediction. However, non-deterministic parallel routing hampers model training and degrades prediction accuracy. In this work, we propose a stochastic approach, called LGC-Net, to solve this problem. In this approach, we develop new techniques of Gaussian random field layer and focal likelihood loss function to seamlessly integrate Log Gaussian Cox process with deep learning. This approach provides not only statistical regression results but also classification ones with different thresholds without retraining. Experimental results with noisy training data on industrial designs demonstrate that LGC-Net achieves significantly better accuracy of DRV density prediction than prior arts.
Rongjian Liang, Hua Xiang 0001, Jinwook Jung, Jiang Hu 0001, Gi-Joon Nam
ICCAD1
2022 Design Rule Violation Prediction at Sub-10-nm Process Nodes Using Customized Convolutional Networks
abstract
As the semiconductor process technology advances into sub-10-nm regime, cell pin accessibility, which is a complex joint effect from the pin shape and nearby blockages, becomes a main cause for design rule violations (DRVs). Therefore, a machine-learning model for DRV prediction needs to consider both very high-resolution pin shape patterns and low-resolution layout information as input features. A new convolutional neural network technique, J-Net, is introduced for the prediction with mixed resolution features. This is a customized architecture that is flexible for handling various input and output resolution requirements. It can be applied at placement stage without using global routing information. This technique is evaluated on 12 industrial designs at a 7-nm technology node. The results show that the J-Net-based binary classifier can improve the true positive rate by 37%, 40%, and 7%, respectively, compared to extensions of three recent works, with similar false positive rates.
Rongjian Liang, Hua Xiang 0001, Diwesh Pandey, Lakshmi N. Reddy, Shyam Ramji, Gi-Joon Nam, Jiang Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Preplacement Net Length and Timing Estimation by Customized Graph Neural Network
abstract
Net length is a key proxy metric for optimizing timing and power across various stages of a standard digital design flow. However, the bulk of net length information is not available until cell placement, and hence, it is a significant challenge to explicitly consider net length optimization in design stages prior to placement, such as logic synthesis. In addition, the absence of net length information makes accurate preplacement timing estimation extremely difficult. Poor predictability on the timing not only affects timing optimizations but also hampers the accurate evaluation of synthesis solutions. This work addresses these challenges by a preplacement prediction flow with estimators on both net length and timing. We propose a graph attention network (GAT) method with customization, called Net2, to estimate individual net length before cell placement. Its accuracy-oriented version Net2a achieves about 15% better accuracy than several previous works in identifying both long nets and long critical paths. Its fast version Net2f is more than$1000\times $faster than placement while still outperforms previous works and other neural network techniques in terms of various accuracy metrics. Based on net size estimations, we propose the first machine learning-based preplacement timing estimator. Compared with the preplacement timing report from commercial tools, it improves the correlation coefficient in arc delays by 0.08, and reduces the mean absolute error in slack, worst negative slack, and total negative slack estimations by more than 50%.
Zhiyao Xie, Rongjian Liang, Jiang Hu 0001, Chen-Chia Chang, Jingyu Pan, Yiran Chen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Net2: A Graph Attention Network Method Customized for Pre-Placement Net Length Estimation
abstract
Net length is a key proxy metric for optimizing timing and power across various stages of a standard digital design flow. However, the bulk of net length information is not available until cell placement, and hence it is a significant challenge to explicitly consider net length optimization in design stages prior to placement, such as logic synthesis. This work addresses this challenge by proposing a graph attention network method with customization, called Net2, to estimate individual net length before cell placement. Its accuracy-oriented version Net2a achieves about 15% better accuracy than several previous works in identifying both long nets and long critical paths. Its fast version Net2f is more than 1000x faster than placement while still outperforms previous works and other neural network techniques in terms of various accuracy metrics.
Zhiyao Xie, Rongjian Liang, Jiang Hu 0001, Yixiao Duan, Yiran Chen 0001
ASP-DAC2
2021 Automatic Routability Predictor Development Using Neural Architecture Search
abstract
The rise of machine learning technology inspires a boom of its applications in electronic design automation (EDA) and helps improve the degree of automation in chip designs. However, manually crafted machine learning models require extensive human expertise and tremendous engineering efforts. In this work, we leverage neural architecture search (NAS) to automate the development of high-quality neural architectures for routability prediction, which can help to guide cell placement toward routable solutions. Our search method supports various operations and highly flexible connections, leading to architectures significantly different from all previous human-crafted models. Experimental results on a large dataset demonstrate that our automatically generated neural architectures clearly outperform multiple representative manually crafted solutions. Compared to the best case of manually crafted models, NAS-generated models achieve 5.85% higher Kendall's$T$in predicting the number of nets with DRC violations and 2.12% better area under ROC curve (ROC-AUC) in DRC hotspot detection. Moreover, compared with human-crafted models, which easily take weeks to develop, our efficient NAS approach finishes the whole automatic search process with only 0.3 days.
Chen-Chia Chang, Jingyu Pan, Tunhou Zhang, Zhiyao Xie, Jiang Hu 0001, Weiyi Qi, Chung-Wei Lin, Rongjian Liang, Joydeep Mitra, Elias Fallon, Yiran Chen 0001
ICCAD8
2021 FlowTuner: A Multi-Stage EDA Flow Tuner Exploiting Parameter Knowledge Transfer
abstract
EDA tools provide a large spectrum of parameters to help designers achieve the maximized PPA of designs. The corresponding enormous solution space, however, hinders designers from navigating towards optimal solutions. In this paper, we propose a multi-stage automatic flow tuning tool, named FlowTuner, for efficient and effective parameter tuning of VLSI design flow. It utilizes both exploitation using transferred parameter knowledge from archival design data and exploration via a multi-stage cooperative co-evolutionary framework. Furthermore, novel flow jump-start and early-stop techniques are developed to reduce the overall runtime for tuning. Experiments on a set of IWLS 2005 benchmark circuits through a commercial tool flow demonstrate that FlowTuner produces considerably better design outcomes in 50 % shorter turnaround time compared to the state-of-the-art flow tuning techniques.
Rongjian Liang, Jinwook Jung, Hua Xiang 0001, Lakshmi N. Reddy, Alexey Lvov, Jiang Hu 0001, Gi-Joon Nam
ICCAD1
2020 Routing-Free Crosstalk Prediction
abstract
Interconnect spacing is getting increasingly smaller in advanced technology nodes, which adversely increases the capacitive coupling of adjacent interconnect wires. It makes crosstalk a significant contributor to signal integrity and timing, and it is now imperative to prevent crosstalk-induced noise and delay issues in the earlier stages of VLSI design flow. Nonetheless, since the crosstalk effect depends primarily on the switching of neighboring nets, accurate crosstalk evaluation is only viable at the late stages of design flow with routing information available, e.g., after detailed routing. There have also been previous efforts in early-stage crosstalk prediction, but they mostly rely on time-expensive trial routing. In this work, we propose a machine learning-based routing-free crosstalk prediction framework. Given a placement, we identify routing and net topology-related features, along with electrical and logical features, which affect crosstalk-induced noise and delay. We then employ machine learning techniques to train the crosstalk prediction models, which can be used to identify crosstalk-critical nets in placement stages. Experimental results demonstrate that the proposed method can instantly classify more than 70% of crosstalk-critical nets after placement with a false-positive rate of less than 2%.
Rongjian Liang, Zhiyao Xie, Jinwook Jung, Vishnavi Chauha, Yiran Chen 0001, Jiang Hu 0001, Hua Xiang 0001, Gi-Joon Nam
ICCAD1
2020 DRC Hotspot Prediction at Sub-10nm Process Nodes Using Customized Convolutional Network
abstract
As the semiconductor process technology advances into sub-10nm regime, cell pin accessibility, which is a complex joint effect from the pin shape and nearby blockages, becomes a main cause for DRC violations. Therefore, a machine learning model for DRC hotspot prediction needs to consider both very high-resolution pin shape patterns and low-resolution layout information as input features. A new convolutional neural network technique, J-Net, is introduced for the prediction with mixed resolution features. This is a customized architecture that is flexible for handling various input and output resolution requirements. It can be applied at placement stage without using global routing information. This technique is evaluated on 12 industrial designs at 7nm technology node. The results show that it can improve true positive rate by 37%, 40% and 14% respectively, compared to three recent works, with similar false positive rates.
Rongjian Liang, Hua Xiang 0001, Diwesh Pandey, Lakshmi N. Reddy, Shyam Ramji, Gi-Joon Nam, Jiang Hu 0001
ISPD1