VLDB 2026 Research / reviewers in the wild / expert
Haoxing Ren
dblp:58/2578 · also Haoxing Mark Ren
· DBLP profile ↗
107ranked-venue papers
15as first author
78since 2021 · last 2026
0000-0003-1028-3860ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 102 · 15 first-author · 74 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | C3PO: Commercial-Quality Global Placement via Coherent, Concurrent Timing, Routability, and Wirelength OptimizationabstractDespite achieving orders-of-magnitude runtime speedup, GPU-accelerated placers (GPU-Placers) still have extremely limited industrial adoption, largely due to the wide gaps in Power, Performance, and Area (PPA) metrics compared to those well-established CPU-centric commercial Physical Design (PD) tools. To overcome this issue, we introduce C3PO, the first commercial-quality, differentiable, multi-objective global placer that performs concurrent timing, routability, and wirelength optimization in a coherent manner with custom CUDA kernels. Particularly, we propose a convex-based framework that dynamically computes objective weights at each placement iteration by solving a quadratic problem, eliminating the need of manual parameter tuning. In the experiments, we rigorously validate C3PO with an industry-leading commercial PD tool and demonstrate that on 8 designs from TILOS [1] and IWLS [2] in ASAP 7nm [3], C3PO consistently outperforms the commercial tool by up to 16.7% in routed wirelength and 19.6% in switching power with complete full-flow validation. Yi-Chen Lu, Hao-Hsiang Hsiao, Rongjian Liang, Haoxing Ren |
ASP-DAC | 5 |
| 2026 | Differentiable Tier Assignment for Timing and Congestion-Aware Routing in 3D ICsabstractState-of-the-art (SOTA) 3D physical design (PD) flows extend commercial 2D place-and-route (P&R) tools to enable signoff-quality 3D IC implementation through double metal stacking and inter-die metal layer sharing. While metal layer sharing introduces additional routing resources, the substantially higher manufacturing cost of face-to-face (F2F) inter-die vias compared to intra-die vias necessitates 3D-aware routing strategies to manage routability-cost trade-offs. To address this, we propose differentiable routing guidance for 3D ICs (DRG-3D), a GPUaccelerated differentiable optimization framework that provides routing guidance for 3D ICs. DRG-3D formulates a fully differentiable objective that simultaneously optimizes key 3D design metrics: routing congestion, wirelength, via cost, and F 2 F -via cost, which enables efficient and scalable gradient-based optimization over large-scale netlists. Experimental results show that DRG-3D outperforms the SOTA Pin-3D flow, achieving up to 8.37% reduction in routing overflow, 23.99% reduction in total negative slack (TNS), and 18.05% reduction in post-route timing violations. Yuan-Hsiang Lu, Hao-Hsiang Hsiao, Yi-Chen Lu, Haoxing Ren, Sung Kyu Lim |
ASP-DAC | 4 |
| 2026 | Code, Not Canvas: Multi-Agent Layout Generation Beyond Vision ModelsabstractRule-constrained chip layout generation is crucial in semiconductor manufacturing industry, providing significant resouces for technology developement and data-driven methodologies. Generative AI has become the mainstream solution for layout generation backboned with GAN, ViT or Diffussion models. These methods leverages two phase flow including squish topology generation and DRC-aware geometry filling. However, vision model-based approach is sub-optimal and lacks controllability of generated layouts. To address this, we reformulate layout generation as a coding problem which offers the user max controllability of layout generation and bridges the gap between vision-based geometry generation and LLM’s coding capability. Specifically, we deliver a multi-agent framework to write Python code to create diverse GDSII layouts following given design rule constraints. We demonstrate superior performance for layout generation over SOTA vision-based models, GPT-4o and Cursor in terms of DRC-clean pattern diversity. Haoxing Ren |
ASP-DAC | 2 |
| 2026 | Invited: Polymath: Self-Improving Hierarchical Workflow for Multi-Domain Problem SolvingabstractLarge language models (LLMs) excel at solving complex tasks by executing agentic workflows composed of detailed instructions and structured operations. However, building agents for diverse applications by manually embedding foundation models into agentic systems such as Chain-of-Thought, Self-Reflection, and ReACT through text interfaces limits scalability and efficiency. Recently, researchers have explored automating workflow generation using code-based representations, but most methods depend on labeled data, limiting their applicability to real-world, dynamic hardware design problems. We introduce Polymath, a self-improving agent with a dynamic hierarchical workflow that combines task flow graphs with code-represented workflows to address these challenges. Polymath employs an experience-driven optimization framework that integrates multi-level graph optimization using surrogate scores from historical evaluations with a self-reflection-guided evolutionary algorithm for workflow refinement, enabling unsupervised self-improvement without labeled data. Experiments show that Polymath outperforms a leading commercial agentic system by 16.23% pass@1 and 11.47% pass@3 on hardware benchmarks, and achieves an average 8.1% improvement over state-of-the-art baselines on coding, math, and multi-turn QA tasks. Chia-Tung Ho, Abhishek B. Akkur, Haoxing Ren |
ISPD | 5 |
| 2026 | GrandPlan: Differentiable, Simultaneous Top-Level Floorplanning and Partition-Level Cell Placement for Large-Scale IP-CoresabstractTop-level floorplanning is a critical step in industrial physical design, where the die is partitioned into exactly abutted regions with carefully allocated areas to enable efficient hierarchical place-and-route and achieve desired power, performance, and area (PPA) trade-offs. In current practice, however, floorplanning remains largely manual and sub-optimal, as designers rely on RTL hierarchy with limited physical guidance; commercial tools cannot feasibly perform flat optimization at IP-core scale. As a result, late-stage routability-driven partition resizing often triggers cascading boundary changes, disrupting neighboring partitions and significantly increasing turnaround time and engineering cost. To address this challenge, we present GrandPlan, a GPU-accelerated, differentiable, end-to-end framework that co-optimizes top-level floorplanning and partition-level cell placement within a single automated loop. Leveraging custom CUDA kernels, GrandPlan generates clean, rectilinear partition boundaries while concurrently placing macros and standard cells. The framework consists of three tightly coupled stages: (1) flat IP-core placement with differentiable grouping objectives, (2) boundary refinement via simulated annealing under area and routability constraints, and (3) routability-aware fence-region placement. Experiments on eight large-scale industrial IP-cores (up to 25M cells) show that GrandPlan reduces total wirelength by up to 14% and cross-partition (feedthrough) wirelength by 27% on average compared to human-expert-crafted baselines, with an average runtime of only 1.2 hours. Zhili Xiong, Yi-Chen Lu, David Z. Pan, Haoxing Ren |
ISPD | 4 |
| 2026 | LiDAR 3.0: Photonics-Aware Planning-Guided Automated Electrical Routing for Large-Scale Active Photonic Integrated CircuitsabstractThe rising demand for AI training and inference, as well as scientific computing, combined with stringent latency and energy budgets, is driving the adoption of integrated photonics for computing, sensing, and communications. As active photonic integrated circuits (PICs) scale in device count and functional heterogeneity, physical implementation by manual scripting and ad-hoc edits is no longer tenable. This creates an immediate need for an electronic–photonic design automation (EPDA) stack in which physical design automation is a core capability. However, there is currently no end-to-end fully automated routing flow that coordinates photonic waveguides and on-chip metal interconnect. Critically, available digital VLSI and analog/custom routers are not directly applicable to PIC metal routing due to a lack of customization to handle constraints induced by photonic devices and waveguides. We present, to our knowledge, the first end-to-end routing framework LiDAR 3.0 for large-scale active PICs that addresses waveguides and metal wires within a unified flow. We introduce a physically-aware global planner that generates congestion- and crossing-aware routing guides while explicitly accounting for the region of photonic components and waveguides. We further propose a sequence-consistent track assignment and a soft guidance-assisted detailed routing to speed up the routing process with significantly optimized routability and via usage. Evaluated on various large PIC designs, our router delivers fast, high-quality active PIC routing solutions with fewer vias, lower congestion, and competitive runtime relative to manual and existing VLSI router baselines; on average it reduce via count by ~99%, user-specified design rule violation by ~98%, and runtime by 17x, establishing a practical foundation for EPDA at system scale. Hongjian Zhou, Nicholas Gangi, Meng Zhang 0023, Haoxing Ren, Rena Huang, Jiaqi Gu 0002 |
ISPD | 6 |
| 2026 | FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification
Lily Jiaxin Wan, Chia-Tung Ho, Yunsheng Bai, Cunxi Yu, Deming Chen, Haoxing Ren |
VTS | 6 |
| 2026 | LiDAR 2.0: Hierarchical Curvy Waveguide Detailed Routing for Large-Scale Photonic Integrated CircuitsabstractDriven by innovations in photonic computing and interconnects, photonic integrated circuit (PIC) designs advance and grow in complexity. Traditional manual physical design processes have become increasingly cumbersome. Available PIC layout tools are mostly schematic-driven, which has not alleviated the burden of manual waveguide planning and layout drawing. Previous research in PIC automated routing is largely adapted from electronic design, focusing on high-level planning and overlooking photonic-specific constraints such as curvy waveguides, bending, and port alignment. As a result, they fail to scale and cannot generate DRV-free layouts, highlighting the need for dedicated electronic-photonic design automation tools to streamline PIC physical design. In this work, we present LiDAR, the first automated PIC detailed router for large-scale designs. It features a grid-based, curvy-aware A* engine with adaptive crossing insertion, congestion-aware net ordering, and insertion-loss optimization. To enable routing in more compact and complex designs, we further extend our router to hierarchical routing as LiDAR 2.0. It introduces redundant-bend elimination, crossing space preservation, and routing order refinement for improved conflict resilience. We also develop and open-source a YAML-based PIC intermediate representation and diverse benchmarks, including TeMPO, GWOR, and Bennes, which feature hierarchical structures and high crossing densities. Evaluations across various benchmarks show that LiDAR 2.0 produces nearly DRV-free layouts, achieving up to 16% lower insertion loss and 7.69× speedup over prior methods on spacious cases, and 9% lower insertion loss with 6.95× speedup over LiDAR 1.0 on compact cases. Our codes are open-sourced at link. Hongjian Zhou, Ziang Yin, Nicholas Gangi, Z. Rena Huang, Haoxing Ren, Joaquin Matres, Jiaqi Gu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | MapTune: Versatile ASIC Technology Mapping via Reinforcement Learning Guided Library TuningabstractTechnology mapping involves mapping logical circuits to a library of standard cells. Traditionally, a full technology library is used, leading to a large search space and potential runtime overhead. Motivated by randomly sampled technology mapping case studies, we propose MapTune to address this challenge by utilizing reinforcement learning to make design-specific cell selection choices. By learning from the environment and guided by the reward, MapTune refines the cell selection process, resulting in a reduced search space and potentially improved mapping quality. The effectiveness of MapTune is evaluated on a wide range of benchmarks, different technology libraries, and various technology mappers. The empirical results demonstrate that MapTune achieves higher mapping accuracy and reduces delay/area across various circuit designs, technology libraries, and mappers. The article also discusses the Pareto-Optimal exploration and confirms the perpetual delay-area tradeoff. Conducted on benchmark suites ISCAS 85/89, ITC/ISCAS 99, VTR8.0, and EPFL benchmarks, the post-technology mapping and post-sizing quality-of-results (QoR) have been significantly improved, with average Area-Delay Product (ADP) improvement of 16.56% among all different exploration settings in MapTune. The improvements consistently remained for four different technologies (7 nm, 45 nm, 130 nm, and 180 nm) with various mappers including both state-of-the-art open-source and commercial synthesis tools. Mingju Liu, Daniel Robinson, Johannes Maximilian Kühn, Rongjian Liang, Haoxing Ren, Cunxi Yu |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | Intelligent OPC Engineer Assistant for Semiconductor ManufacturingabstractAdvancements in chip design and manufacturing have enabled the processing of complex tasks such as deep learning and natural language processing, paving the way for the development of artificial general intelligence (AGI). AI, on the other hand, can be leveraged to innovate and streamline semiconductor technology from planning and implementation to manufacturing. In this paper, we present Intelligent OPC Engineer Assistant, an AI/LLM-powered methodology designed to solve the core manufacturing-aware optimization problem known as Optical Proximity Correction (OPC). The methodology involves a reinforcement learning-based OPC recipe search and a customized multi-modal agent system for recipe summarization. Experiments demonstrate that our methodology can efficiently build OPC recipes on various chip designs with specially handled design topologies, a task that typically requires the full-time effort of OPC engineers with years of experience. Guojin Chen, Bei Yu 0001, Haoxing Ren |
AAAI | 4 |
| 2025 | VerilogCoder: Autonomous Verilog Coding Agents with Graph-based Planning and Abstract Syntax Tree (AST)-based Waveform Tracing ToolabstractDue to the growing complexity of modern Integrated Circuits (ICs), automating hardware design can prevent a significant amount of human error from the engineering process and result in less errors. Verilog is a popular hardware description language for designing and modeling digital systems; thus, Verilog generation is one of the emerging areas of research to facilitate the design process. In this work, we propose VerilogCoder, a system of multiple Artificial Intelligence (AI) agents for Verilog code generation, to autonomously write Verilog code and fix syntax and functional errors using collaborative Verilog tools (i.e., syntax checker, simulator, and waveform tracer). Firstly, we propose a task planner that utilizes a novel Task and Circuit Relation Graph retrieval method to construct a holistic plan based on module descriptions. To debug and fix functional errors, we develop a novel and efficient abstract syntax tree (AST)-based waveform tracing tool, which is integrated within the autonomous Verilog completion flow. The proposed methodology successfully generates 94.2% syntactically and functionally correct Verilog code, surpassing the state-of-the-art methods by 33.9% on the VerilogEval-Human v2 benchmark. Chia-Tung Ho, Haoxing Ren, Brucek Khailany |
AAAI | 2 |
| 2025 | ChipAlign: Instruction Alignment in Large Language Models for Chip Design via Geodesic InterpolationabstractRecent advancements in large language models (LLMs) have expanded their application across various domains, including chip design, where domain-adapted chip models like ChipNeMo have emerged. However, these models often struggle with instruction alignment, a crucial capability for LLMs that involves following explicit human directives. This limitation impedes the practical application of chip LLMs, including serving as assistant chatbots for hardware design engineers. In this work, we introduce ChipAlign, a novel approach that utilizes a training-free model merging strategy, combining the strengths of a general instruction-aligned LLM with a chip-specific LLM. By considering the underlying manifold in the weight space, ChipAlign employs geodesic interpolation to effectively fuse the weights of input LLMs, producing a merged model that inherits strong instruction alignment and chip expertise from the respective instruction and chip LLMs. Our results demonstrate that ChipAlign significantly enhances instruction-following capabilities of existing chip LLMs, achieving up to a 26.6% improvement on the IFEval benchmark, while maintaining comparable expertise in the chip domain. This improvement in instruction alignment also translates to notable gains in instruction-involved QA tasks, delivering performance enhancements of 3.9% on the OpenROAD QA benchmark and 8.25% on production-level chip QA benchmarks, surpassing state-of-the-art baselines. Chenhui Deng, Yunsheng Bai, Haoxing Ren |
DAC | 3 |
| 2025 | GEM: GPU-Accelerated Emulator-Inspired RTL SimulationabstractIn this paper, we present a GPU-accelerated RTL simulator addressing critical challenges in high-speed circuit verification. Traditional CPU-based RTL simulators struggle with scalability and performance, and while FPGA-based emulators offer acceleration, they are costly and less accessible. Previous GPU-based attempts have failed to speed up RTL simulation due to the heterogeneous nature of circuit partitions, which conflicts with the SIMT (Single Instruction, Multiple Thread) paradigm of GPUs. Inspired by the design of emulators, our approach introduces a novel virtual Very Long Instruction Word (VLIW) architecture, designed for efficient CUDA execution. We also design a flow that maps circuit logic to the architecture in a process analogous to the FPGA CAD flow. This architecture mitigates issues of irregular memory access and thread divergence, unlocking GPU potential for RTL simulation. Our solution achieves up to $64 \times$ speed-up over the best CPU simulators, democratizing high-speed RTL simulation with accessible hardware and establishing a new frontier for GPUaccelerated circuit verification. Zizheng Guo 0001, Yanqing Zhang 0002, Runsheng Wang, Yibo Lin, Haoxing Ren |
DAC | 5 |
| 2025 | DCO-3D: Differentiable Congestion Optimization in 3D ICsabstractState-of-the-art 3D IC flows fail to consider 3D congestion during earlier stages, leading to excessive use of end-of-flow ECO resources for routability correction that severely degrades full-chip Power, Performance, and Area metrics. We present DCO-3D, a Machine Learning-based routability-aware 3D PD flow that performs early post-route congestion prediction using Siamese Networks and resolves the predicted hotspots using a fully differentiable 3D cell spreading with Graph Neural Network. On 6 industrial designs in a commercial 3nm node, DCO-3D improves Pin-3D, the known best Pin-3D flow, by up to 47.2% in overflow, 86.2% in TNS and 5.1% in power at signoff. Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Anthony Agnesina, Rongjian Liang, Yuan-Hsiang Lu, Haoxing Ren, Sung Kyu Lim |
DAC | 7 |
| 2025 | Reinforcement Learning-Driven Window Selection for Enhanced Window-Based Rip-up and Reroute in Chip Detailed RoutingabstractWith increasingly complex design rules and pin density in advanced technology nodes, achieving a violation-free layout has become more challenging, also making rip-up and reroute (RUR) the most runtime-intensive component of detailed routing. We propose a novel reinforcement learning (RL)based approach to enhance the window-based RUR process. Our method features a dynamic window generation strategy that adjusts window size and position based on the distribution of design rule violations (DRV), enabling efficient targeting of congested areas. By leveraging the predictive capabilities of RL, our approach aims to minimize DRVs and achieve high-quality routing results. Experimental results demonstrate that our method outperforms the state-of-theart detailed routers, TritonRoute, achieving a DRV-free solution, averagely improving wirelength by 0.07%, via count by 2.42%, and consuming almost the same average runtime. Yu-Chan Keng, Yu-Chun Pai, Wen-Hao Liu 0001, Haoxing Ren, Danny Liu, Rongjian Liang, Mark Ho, Anthony Agnesina, Yih-Lang Li |
DAC | 4 |
| 2025 | INSTA: An Ultra-Fast, Differentiable, Statistical Static Timing Analysis Engine for Industrial Physical Design ApplicationsabstractPrior GPU-accelerated Static Timing Analysis (GPU-STA) works all struggle to find industrial adoption, primarily because they aim to build standalone timing engines that can never emulate the proprietary delay models used in commercial tools. In this paper, we adopt a different philosophy by presenting INSTA, the first-ever differentiable, statistical GPU-STA engine that achieves unprecedented accuracy and scalability by a one-time initialization from any reference tool, bringing two transformative capabilities to Physical Design (PD): (1) rapid, high-fidelity timing analysis for incremental netlist update, and (2) gradient-based truly-global timing optimization at scale. Notably, INSTA demonstrates a near-perfect 0.999 correlation with an industryleading signoff tool on a 15 -million-pin design in a commercial 3 nm node with runtime under 0.1 seconds. Experimental results showcase INSTA’s capability through three PD applications: (1) serving as a fast evaluator in an industrial gate sizing flow, achieving $\mathbf{2 5 x}$ faster incremental update_timing runtime with almost no accuracy loss; (2) INSTA-Size, a gradient-based gate sizer that achieves up to $\mathbf{1 5 \%}$ better Total Negative Slack (TNS) than the reference signoff engine by sizing $68 \%$ fewer amount of cells; and (3) INSTA-Place, a differentiable timingdriven global placer that outperforms the state-of-the-art net-weighting placer by up to 16% in Half-Perimeter Wirelegnth (HPWL) and 59.4% in TNS on the ICCAD’15 benchmark [15]. Yi-Chen Lu, Zhizheng Guo, Kishor Kunal, Rongjian Liang, Haoxing Ren |
DAC | 5 |
| 2025 | BOSON-1: Understanding and Enabling Physically-Robust Photonic Inverse Design with Adaptive Variation-Aware Subspace OptimizationabstractNanophotonic device design aims to optimize pho-tonic structures to meet specific requirements across various applications. Inverse design has unlocked non-intuitive, high-dimensional design spaces, enabling the discovery of compact, high-performance device topologies beyond traditional heuristic or analytic methods. The adjoint method, which calculates analytical gradients for all design variables using just two electromagnetic simulations, enables efficient navigation of this complex space. However, many inverse-designed structures, while numerically plausible, are difficult to fabricate and highly sensitive to physical variations, limiting their practical use. The discrete material distributions with numerous local-optimal structures also pose significant optimization challenges, often causing gradient-based methods to converge on suboptimal designs. In this work, we formulate inverse design as a fabrication-restricted, discrete, prob-abilistic optimization problem and introduce BOSON−1, an end-to-end, adaptive, variation-aware subspace optimization framework to address the challenges of manufacturability, robustness, and optimizability. We explicitly consider the fabrication process and differentiably optimize the design in the fabricable subspace. To overcome optimization difficulty, we propose dense target-enhanced gradient flows to mitigate misleading local optima and introduce a conditional subspace optimization strategy to create high-dimensional tunnels to escape local optima. Furthermore, we significantly reduce the prohibitive runtime associated with optimizing across exponential variation samples through an adaptive sampling-based robust optimization method, ensuring both efficiency and variation robustness. On three representative photonic device benchmarks, our proposed inverse design methodology BOSON−1delivers fabricable structures and achieves the best convergence and performance under realistic variations, outperforming prior arts with 74.3% post-fabrication performance. Pingchuan Ma 0012, Zhengqi Gao, Amir Begovic, Meng Zhang 0023, Haoxing Ren, Z. Rena Huang, Duane S. Boning, Jiaqi Gu 0002 |
DATE | 6 |
| 2025 | FVEval: Understanding Language Model Capabilities in Formal Verification of Digital HardwareabstractThe remarkable reasoning and code generation capabilities of large language models (LLMs) have spurred significant interest in applying LLMs to enable task automation in digital chip design. In particular, recent work has investigated early ideas of applying these models to formal verification (FV), an approach to verifying hardware implementations that can provide strong guarantees of confidence but demands significant amounts of human effort. While the value of LLM-driven automation is evident, our understanding of model performance, however, has been hindered by the lack of holistic evaluation. In response, we present FVEval, the first comprehensive benchmark and evaluation framework for characterizing LLM performance in tasks pertaining to FV. The benchmark consists of three sub-tasks that measure LLM capabilities at different levels-from the generation of SystemVerilog assertions (SVA) given natural language descriptions to reasoning about the design RTL and suggesting assertions directly without additional human input. As test instances, we present both collections of expert-written verification collateral and methodologies to scalably generate synthetic examples aligned with industrial FV workflows. A wide range of existing LLMs, both proprietary and open-source, are evaluated against FVEval, based on which we investigate where today's LLMs stand and how we might further enable their application toward improving productivity in digital FV. Our benchmark and evaluation code is available at https://github.com/NVlabs/FVEval. Ghaith Bany Hamad, Syed Suhaib, Haoxing Ren |
DATE | 5 |
| 2025 | ChipVQA: Benchmarking Visual Language Models for Chip DesignabstractLarge-language models (LLMs) have exhibited great potential to assist chip designs and analysis. Recent research and efforts are mainly focusing on text-based tasks including general QA, debugging, design tool scripting, and so on. However, chip design and implementation workflow usually require a visual understanding of diagrams, flow charts, graphs, schematics, waveforms, etc, which demands the development of multimodality foundation models. In this paper, we propose ChipVQA, a benchmark designed to evaluate the capability of visual language models for chip design. ChipVQA includes 142 carefully designed and collected VQA questions covering five chip design disciplines: Digital Design, Analog Design, Architecture, Physical Design and Semiconductor Manufacturing. Unlike existing VQA benchmarks, ChipVQA questions are carefully designed by chip design experts and require indepth domain knowledge and reasoning to solve. We conduct comprehensive evaluations on both open-source and proprietary multimodal models that are greatly challenged by the benchmark suit. ChipVQA is available at https://github.com/phdyang007/chipvqa. Qijing Huang 0001, Nathaniel Ross Pinckney, Walker J. Turner, Wenfei Zhou, Yanqing Zhang 0002, Chia-Tung Ho, Chen-Chia Chang, Haoxing Ren |
DATE | 9 |
| 2025 | SimPart: A Simple Yet Effective Replication-Aided Partitioning Algorithm for Logic Simulation on GPU
Yi-Hua Chung, Shui Jiang, Wan-Luan Lee, Yanqing Zhang 0002, Haoxing Ren, Tsung-Yi Ho, Tsung-Wei Huang |
Euro-Par (3) | 5 |
| 2025 | BUFFALO: PPA-Configurable, LLM-based Buffer Tree Generation via Group Relative Policy OptimizationabstractBuffer insertion is a critical netlist optimization technique in Physical Design (PD) that balances trade-offs between Power, Performance, and Area (PPA) metrics. Traditional buffering methods rely heavily on local heuristics, which do not scale and often result in globally sub-optimal solutions. Prior Machine Learning (ML) techniques such as BufFormer attempted to alleviate this limitation but remain prohibitively time-consuming (and sub-optimal) due to their incremental nature. In this paper, we introduce BUFFALO, a generative buffer insertion framework that, for the first time in PD, formulates buffer tree generation as a sequence-to-sequence task solved by Large Language Models (LLMs). Particularly, given a design, BUFFALO performs single-shot generation of buffer trees for all fanout-violating nets and INSTA-selected timing critical nets. Furthermore, Group Relative Policy Optimization (GRPO), a Reinforcement Learning (RL) technique, is employed to refine predicted solutions in a PPA-configurable manner. Experimental results on 9 full-chip designs in a 7nm node demonstrate that BUFFALO outperforms an industry-leading commercial PD tool by 71% in Total Negative Slack (TNS), 67.69% in Worst Negative Slack (WNS), and 83x in runtime without incurring additional power consumption. Hao-Hsiang Hsiao, Yi-Chen Lu, Sung Kyu Lim, Haoxing Ren |
ICCAD | 4 |
| 2025 | Invited Paper: LLM-Enhanced GPU-Optimized Physical Design at ScaleabstractModern Physical Design (PD) flows face a dual challenge: proprietary, heterogeneous design data and the rapid evolution of process nodes, both of which block models from transferring to new chips. To overcome these hurdles, we demonstrate a unified, data-driven framework that distills critical netlist optimization moves, including gate sizing, buffer insertion, and cell relocation, into "optimization primitives" learned by Large Language Models (LLMs). Particularly, we develop a high-quality, synthetic optimization data generation pipeline with commercial tools at scale, while using a GPU-accelerated differentiable Static Timing Analysis (STA) engine to create fast feedback loop, enabling end-to-end gradient propagation to guide model learning. By training on both real and synthetic data across multiple technology generations, our approach captures fundamental PD optimization patterns that transfer seamlessly to unseen designs, overcoming the constraints of fragmented design representations and proprietary data in industrial PD flows. Yi-Chen Lu, Hao-Hsiang Hsiao, Haoxing Ren |
ICCAD | 3 |
| 2025 | Invited Paper: 2025 ICCAD CAD Contest Problem C: Incremental Placement Optimization Beyond Detailed Placement: Simultaneous Gate Sizing, Buffering, and Cell RelocationabstractLate-stage placement optimization is where real PPA trade-offs surface, and where conventional heuristic passes tend to get trapped in small, local neighborhoods. We frame an invited "Problem C" contest that treats this stage as a global, multi-operator search over gate sizing, buffer/inverter-pair insertion, and legal cell relocation, with strict reproducibility and legality. Our core belief grounded in production experience is that GPU batching and differentiable guidance expand the tractable search space: you can score and steer thousands of coordinated moves per iteration, not just a handful, and do so under tight runtime budgets. Submissions must produce a replayable ECO changelist and a final legal DEF; a standardized evaluation flow computes timing, power, and wirelength and combines them with displacement and runtime into the contest score. The specification is designed to encourage pragmatic use of gradient signals and tensorized batching without mandating any single method, enabling participants to leverage novel GPU tools to deliver industrially deployable PPA gains. Yi-Chen Lu, Rongjian Liang, Wen-Hao Liu 0001, Haoxing Ren |
ICCAD | 4 |
| 2025 | Leveraging GPU for Better Detailed Placement QualityabstractIn the physical design flow, detailed placement is critical for wirelength optimization and routability enhancement. While many CPU-based approaches emphasize wirelength reduction, GPU-based approaches primarily focus on accelerating detailed placement without sacrificing quality. However, leveraging GPU parallelism to further improve placement quality remains largely underexplored. As existing optimization steps have approached the practical limits of wirelength optimization, achieving further improvements has become increasingly challenging. In this work, we propose a GPU-based detailed placement flow featuring Simultaneous Row and Order Assignment (SROA) step, a novel step that integrates dynamic programming-based row assignment with heuristic-based cell position adjustment. SROA enables efficient exploration of a significantly larger solution space on GPUs. Experimental results show that our approach preserves routability while reducing detailed routing wirelength and via count by 1.35% and 1.03%, respectively, compared to the ABCDPlace solutions. Chen-Han Lu, Wen-Hao Liu 0001, Haoxing Ren, Ting-Chi Wang |
ICCAD | 3 |
| 2025 | LLM4Verilog: Building Large-Scale, High-Quality Data Infrastructure for Verilog Code Generation via Community EffortsabstractDespite recent advancements in code generation with large language models (LLMs), generating hardware code such as Verilog remains a significant challenge due to the scarcity of large-scale, high-quality datasets in the hardware domain. Existing approaches, including scraping open-source repositories and relying on manually curated datasets, often suffer from limited diversity, quality, and scalability. To address these limitations, we introduce LLM4Verilog, an exploratory, collaborative initiative aimed at constructing a large-scale, high-quality, open-source Verilog dataset. Our initiative integrates a community-driven data collection pipeline with a two-stage data filtering technique to ensure high dataset quality. The first stage removes duplicates and low-quality samples, resulting in a large-scale dataset called LLM4Verilog-complete. The second stage applies an LLM-driven quality scoring method, VeriScore, to perform fine-grained filtering and produce a high-quality, ready-to-use dataset called LLM4Verilog-filtered. We evaluate the effectiveness of these datasets through fine-tuning three different LLMs on our dataset, achieving 6.6%~11.2% and 5.3%~13.2% improvements in pass@1 scores on VerilogEval-human and VerilogEval-Machine, respectively, compared to models fine-tuned with prior state-of-the-art datasets. Notably, these improvements are achieved without relying on complex fine-tuning or data augmentation techniques, highlighting our dataset’s strong potential as a foundational resource for enhancing LLMs’ Verilog code generation capabilities. For more information about our initiative and resulting dataset, please refer to https://nvlabs.github.io/LLM4HWDesign/. Zhongzhi Yu, Chaojian Li, Yongan Zhang, Nathaniel Ross Pinckney, Wenfei Zhou, Rongjian Liang, Haoxing Ren, Yingyan (Celine) Lin |
ICCAD | 9 |
| 2025 | Apollo: Automated Routing-Informed Placement for Large-Scale Photonic Integrated CircuitsabstractAs technology advances, photonic integrated circuits (PICs) are rapidly scaling in size and complexity, with modern designs integrating thousands of components to meet the demands of artificial intelligence (AI), high-performance computing, and chip-to-chip optical interconnects. However, the analog custom layout nature of photonics, the curvy waveguide structures, and single-layer routing resources impose stringent physical constraints, such as minimum bend radii and waveguide crossing penalties, which make manual layout the de facto standard. This manual process takes weeks to complete and is error-prone, which is fundamentally unscalable for large-scale PIC systems. Existing automation solutions have adopted force-directed placement on small benchmarks with tens of components, with limited routability and scalability. To fill this fundamental gap in the electronic-photonic design automation (EPDA) toolchain, we present Apollo, the first GPU-accelerated, routing-informed placement framework tailored for large-scale PICs. Apollo features an asymmetric bending-aware wirelength function with explicit modeling of waveguide routing congestion and crossings to preserve enough routing spacing for routability maximization. Meanwhile, conditional projection is employed to gradually enforce a variety of user-defined layout constraints, including alignment, spacing, etc. This constrained optimization is accelerated and stabilized by a custom blockwise adaptive Nesterov-accelerated optimizer, ensuring stable and high-quality convergence. To catalyze research in PIC layout automation, we also develop and open-source large-scale PIC benchmarks derived from real-world photonic tensor core designs. Compared to existing methods, Apollo can generate high-quality layouts for large-scale PICs with an average routing success rate of 94.79% across all benchmarks within minutes. By tightly coupling placement with physical-aware routing, Apollo establishes a new paradigm for automated PIC design—bringing intelligent, scalable layout synthesis to the forefront of next-generation EPDA. Our code is open-sourced at link*. Hongjian Zhou, Nicholas Gangi, Z. Rena Huang, Haoxing Ren, Jiaqi Gu 0002 |
ICCAD | 5 |
| 2025 | CraftRTL: High-quality Synthetic Data Generation for Verilog Code Models with Correct-by-Construction Non-Textual Representations and Targeted Code RepairabstractDespite the significant progress made in code generation with large language models, challenges persist, especially with hardware description languages such as Verilog. This paper first presents an analysis of fine-tuned LLMs on Verilog coding, with synthetic data from prior methods. We identify two main issues: difficulties in handling non-textual representations (Karnaugh maps, state-transition diagrams and waveforms) and significant variability during training with models randomly making ''minor'' mistakes. To address these limitations, we enhance data curation by creating correct-by-construction data targeting non-textual representations. Additionally, we introduce an automated framework that generates error reports from various model checkpoints and injects these errors into open-source code to create targeted code repair data. Our fine-tuned Starcoder2-15B outperforms prior state-of-the-art results by 3.8\%, 10.9\%, 6.6\% for pass@1 on VerilogEval-Machine, VerilogEval-Human, and RTLLM. Yunda Tsai, Wenfei Zhou, Haoxing Ren |
ICLR | 4 |
| 2025 | GOALPlace: Begin with the End in MindabstractCo-optimizing placement with congestion is integral to achieving high-quality designs. This paper presents GOALPlace, a learning-based approach to improving placement congestion by controlling cell density. It efficiently learns from an EDA tool's post-route optimized results and uses an empirical Bayes technique to adapt the target to a specific placer's solutions, effectively beginning with the end in mind. Our method enhances correlation with the tool's router and timing-opt engine, while solving placement globally without expensive incremental congestion estimation and mitigation methods. A statistical analysis with hierarchical netlist clustering establishes the importance of density and the potential for an adequate cell density target across placements. Our experiments show that our method, when integrated into an academic GPU-accelerated global placer, consistently produces macro and standard cell placements that match or exceed the quality of commercial tools. Our empirical Bayes methodology also shows a substantial quality improvement over leading academic mixed-size placers, achieving up to 10× fewer design rule check (DRC) violations, a 5% decrease in wirelength, and a 30% and 60% reduction in worst and total negative slack (WNS/TNS). Anthony Agnesina, Rongjian Liang, Geraldo Pradipta, Anand Rajaram, Haoxing Ren |
ISPD | 5 |
| 2025 | DRC-Coder: Automated DRC Checker Code Generation Using LLM Autonomous AgentabstractIn the advanced technology nodes, the integrated design rule checker (DRC) is often utilized in place and route tools for fast optimization loops for power-performance-area. Implementing integrated DRC checkers to meet the standard of commercial DRC tools demands extensive human expertise to interpret foundry specifications, analyze layouts, and debug code iteratively. However, this labor-intensive process, requiring to be repeated by every update of technology nodes, prolongs the turnaround time of designing circuits. In this paper, we present DRC-Coder, a multi-agent framework with vision capabilities for automated DRC code generation. By incorporating vision language models and large language models (LLM), DRC-Coder can effectively process textual, visual, and layout information to perform rule interpretation and coding by two specialized LLMs. We also design an auto-evaluation function for LLMs to enable DRC code debugging. Experimental results show that targeting on a sub-3nm technology node for a state-of-the-art standard cell layout tool, DRC-Coder achieves perfect F1 score 1.000 in generating DRC codes for meeting the standard of a commercial DRC tool, highly outperforming standard prompting techniques (F1=0.631). DRC-Coder can generate code for each design rule within four minutes on average, which significantly accelerates technology advancement and reduces engineering costs. Chen-Chia Chang, Chia-Tung Ho, Yiran Chen 0001, Haoxing Ren |
ISPD | 5 |
| 2025 | Invited: ISPD 2025 Performance-Driven Large Scale Global Routing ContestabstractGlobal routing is a critical aspect of VLSI design, significantly impacting timing, power consumption, and routability. The ISPD2024 contest focused on addressing the scalability challenges of global routing by leveraging GPU and machine learning techniques. Building on this foundation, the ISPD2025 contest introduces several important updates to better reflect real-world routing challenges. These updates include the provision of industry-standard input files for more precise modeling and integration with OpenROAD for accurate performance assessment. Collectively, these updates aim to bring the contest closer to practical routing scenarios, fostering the development of scalable and efficient solutions for large-scale chip designs. Rongjian Liang, Anthony Agnesina, Wen-Hao Liu 0001, Matt Liberty, Hsin-Tzu Chang, Haoxing Ren |
ISPD | 6 |
| 2025 | LEGO-Size: LLM-Enhanced GPU-Optimized Signoff-Accurate Differentiable VLSI Gate Sizing in Advanced NodesabstractOn-Chip Variation (OCV)-aware and Path-Based Analysis (PBA) accurate timing optimization achieved by gate sizing (including Vth-assignment) remains a pivotal step in modern signoff. However, in advanced nodes (e.g., 3nm), commercial tools often yield suboptimal results due to the intricate design demands and the vast choices of library cells that require substantial runtime and computational resources for exploration. To address these challenges, we introduce LEGO-Size, a generative framework that harnesses the power of Large Language Models (LLMs) and GPU-accelerated differentiable techniques for efficient gate sizing. LEGO-Size introduces three key innovations. First, it considers timing paths as sequences of tokenized library cells, casting gate sizing prediction as a language modeling task and solving it with self-supervised learning and supervised fine-tuning. Second, it employs a Graph Transformer (GT) with a linear-complexity attention mechanism for netlist encoding, enabling LLMs to make sizing decisions from a global perspective. Third, it integrates a differentiable Static Timing Analysis (STA) engine to refine LLM-predicted gate size probabilities by directly optimizing Total Negative Slack (TNS) through gradient descent. Experimental results on 5 unseen million-gate industrial designs in a commercial 3nm node show that LEGO-Size achieves up to 125x speed up with 37% TNS improvement over an industry-leading commercial signoff tool with minimal power and area overhead. Yi-Chen Lu, Kishor Kunal, Geraldo Pradipta, Rongjian Liang, Ravikishore Gandikota, Haoxing Ren |
ISPD | 6 |
| 2025 | GPU-Accelerated Inverse Lithography Towards High Quality Curvy Mask GenerationabstractInverse Lithography Technology (ILT) has emerged as a promising solution for photo mask design and optimization. Relying on multi-beam mask writers, ILT enables the creation of free-form curvilinear mask shapes that enhance printed wafer image quality and process window. However, a major challenge in implementing curvilinear ILT for large-scale production is mask rule checking, an area currently under development by foundries and EDA vendors. Although recent research has incorporated mask complexity into the optimization process, much of it focuses on reducing e-beam shots, which does not align with the goals of curvilinear ILT. In this paper, we introduce a GPU-accelerated ILT algorithm that improves not only contour quality and process window but also the precision of curvilinear mask shapes. Our experiments on open benchmarks demonstrate a significant advantage of our algorithm over leading academic ILT engines. Source code will be available at https://github.com/phdyang007/curvyILT. Haoxing Ren |
ISPD | 2 |
| 2025 | Cypress: VLSI-Inspired PCB Placement with GPU AccelerationabstractThe scale of printed circuit board (PCB) designs has increased significantly, with modern commercial designs featuring more than 10,000 components. However, the placement process heavily relies on manual efforts that take weeks to complete, highlighting the need for automated PCB placement methods. The challenges of PCB placement arise from its flexible design space and limited routing resources. Existing automated PCB placement tools have achieved limited success in quality and scalability. In contrast, very large-scale integration (VLSI) placement methods have proven to be scalable for designs with millions of cells and delivering high-quality results. Therefore, we propose Cypress, a scalable, GPU-accelerated PCB placement method inspired by VLSI. It incorporates tailored cost functions, constraint handling, and optimized techniques adapted for PCB layouts. In addition, there is an increasing demand for realistic and open-source benchmarks to (1) enable meaningful comparisons between tools and (2) establish performance baselines to track progress in PCB placement technology. To address this gap, we present a PCB benchmark suite synthesized from real commercial designs. We evaluate our method against state-of-the-art commercial and academic PCB placement tools with the benchmark suite. Our approach demonstrates a 1-5.9X higher routability on the proposed benchmarks. For fully routed designs, Cypress achieves 1-19.7X shorter routed track lengths. With GPU acceleration, Cypress delivers up to 492.3X speedup in run time. Finally, we demonstrate scalability to real commercial designs, a capability unmatched by existing tools. Niansong Zhang, Anthony Agnesina, Noor Shbat, Yuval Leader, Zhiru Zhang, Haoxing Ren |
ISPD | 6 |
| 2025 | Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code GenerationabstractThe application of large language models (LLMs) to digital hardware code generation is an emerging field, with most LLMs primarily trained on natural language and software code. Hardware code like Verilog constitutes a small portion of training data, and few hardware benchmarks exist. The open-source VerilogEval benchmark, released in November 2023, provided a consistent evaluation framework for LLMs on code completion tasks. Since then, both commercial and open models have seen significant development. In this work, we evaluate new commercial and open models since VerilogEval’s original release—including GPT-4o, GPT-4 Turbo, Llama3.1 (8B/70B/405B), Llama3 70B, Mistral Large, DeepSeek Coder (33B and 6.7B), CodeGemma 7B, and RTL-Coder—against an improved VerilogEval benchmark suite. We find measurable improvements in state-of-the-art models: GPT-4o achieves a 63% pass rate on specification-to-RTL tasks. The recently released and open Llama3.1 405B achieves a 58% pass rate, almost matching GPT-4o, while the smaller domain-specific RTL-Coder 6.7B models achieve an impressive 34% pass rate. Additionally, we enhance VerilogEval’s infrastructure by automatically classifying failures, introducing in-context learning support, and extending the tasks to specification-to-RTL translation. We find that prompt engineering remains crucial for achieving good pass rates and varies widely with model and task. A benchmark infrastructure that allows for prompt engineering and failure analysis is essential for continued model development and deployment. Nathaniel Ross Pinckney, Christopher Batten, Haoxing Ren, Brucek Khailany |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | DGR: Differentiable Global RouterabstractModern VLSI design flows necessitate fast and high-quality global routers. In this paper, we introduce DGR, a differentiable global router capable of concurrent optimization for hundreds of thousands of nets 1. Our innovation lies in the development of a routing Directed Acyclic Graph (DAG) forest to represent the 2D pattern routing space for all nets, enabling coordinated selection of Steiner trees and 2-pin routing paths from a global perspective. For efficient search within the DAG forest, we relax the discrete search space to be continuous and develop a differentiable solver accelerated by deep learning toolkits on GPUs. Experimental results demonstrate that DGR substantially mitigates routing overflow while concurrently reducing total wirelengths from 0.95% to 4.08% and via numbers from 1.28% to 2.54% in congested testcases compared to state-of-the-art academic global routers. Additionally, DGR exhibits favorable scalability in both runtime and memory with respect to the number of nets. Wei Li 0159, Rongjian Liang, Anthony Agnesina, Chia-Tung Ho, Anand Rajaram, Haoxing Ren |
DAC | 7 |
| 2024 | RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelabstractThis paper presents RTLFixer, a novel framework enabling automatic syntax errors fixing for Verilog code with Large Language Models (LLMs). Despite LLM's promising capabilities, our analysis indicates that approximately 55% of errors in LLM-generated Verilog are syntax-related, leading to compilation failures. To tackle this issue, we introduce a novel debugging framework that employs Retrieval-Augmented Generation (RAG) and ReAct prompting, enabling LLMs to act as autonomous agents in interactively debugging the code with feedback. This framework demonstrates exceptional proficiency in resolving syntax errors, successfully correcting about 98.5% of compilation errors in our debugging dataset, comprising 212 erroneous implementations derived from the VerilogEval benchmark. Our method leads to 32.3% and 10.1% increase in pass@1 success rates in the VerilogEval-Machine and VerilogEval-Human benchmarks, respectively. The source code and benchmark are available at https://github.com/NVlabs/RTLFixer. Yunda Tsai, Haoxing Ren |
DAC | 3 |
| 2024 | BoolGebra: Attributed Graph-Learning for Boolean Algebraic ManipulationabstractLogic optimization is an essential stage in the design automation flow for digital systems as the performance of the system at logic level can have significant impacts on the final chip area, timing closure, and the power efficiency of the system. Logic optimization is a technology-independent circuit optimization at the logic level conducted on multi-level technology-independent representations such as And-Inverter-Graphs (AIGs) [1] and Majority-Inverter-Graphs (MIGs) [2] of the digital logic. Existing state-of-the-art (SOTA) Directed-Acyclic-Graphs (DAGs) aware Boolean optimization algorithms, such as structural rewriting (rw) [1], resubstitution (rs) [3], and refactoring (rf) [1] in ABC [4], are conducted on the AIG data structure with a graph-level single optimization concept, i.e., all nodes in the graph have one same fixed optimization opportunity, while overlooking other potential optimization opportunities. [5] proposes orchestrated logic optimization, which is a fine-grained node-level logic optimization method incorporating multiple optimization techniques within a single AIG traversal. However, the enlarged search space pose a significant challenge in searching optimal solutions without domain knowledge. Anthony Agnesina, Yanqing Zhang 0002, Haoxing Ren, Cunxi Yu |
DATE | 4 |
| 2024 | GL0AM: GPU Logic Simulation Using 0-Delay and Re-simulation Acceleration MethodabstractIn this work, we present GL0AM, a novel GPU accelerated logic simulator that performs delay annotated gate-level simulation, supporting a wide range of sequential components, including SRAMs, and simulation scenarios. We propose a methodology to perform simulation in 2 phases in order to increase the parallelism exposed to the GPU, where the first phase performs 0-delay cycle simulation that records sequential gate (flops, latches, and clock-gates) waveform results, and the second phase performs re-simulation to attain the remaining combinational gate waveforms. We propose to use netlist graph partitioning to minimize synchronization overheads and increase parallelism during the 0-delay simulation phase. GL0AM achieves simulation execution time speedup of 19-537X on an NVIDIA H100 GPU when compared to a commercial simulator across a diverse set of benchmarks. GL0AM provides a 2-5X speedup, iso- GPU platform and benchmark, over a recent state-of-the-art GPU-accelerated logic simulator.1 Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
ICCAD | 2 |
| 2024 | Differentiable Edge-based OPCabstractOptical proximity correction (OPC) is crucial for pushing the boundaries of semiconductor manufacturing and enabling the continued scaling of integrated circuits. While pixel-based OPC, termed as inverse lithography technology (ILT), has gained research interest due to its flexibility and precision. Its complexity and intricate features can lead to challenges in mask writing, increased defects, and higher costs, hence hindering widespread industrial adoption. In this paper, we propose DiffOPC, a differentiable OPC framework that enjoys the virtue of both edge-based OPC and ILT. By employing a mask rule-aware gradient-based optimization approach, DiffOPC efficiently guides mask edge segment movement during mask optimization, minimizing wafer error by propagating true gradients from the cost function back to the mask edges. Our approach achieves lower edge placement error while reducing manufacturing cost by half compared to state-of-the-art OPC techniques, bridging the gap between the high accuracy of pixel-based OPC and the practicality required for industrial adoption, thus offering a promising solution for advanced semiconductor manufacturing. Guojin Chen, Haoxing Ren, Bei Yu 0001, David Z. Pan |
ICCAD | 3 |
| 2024 | 2024 ICCAD CAD Contest Problem C: Scalable Logic Gate Sizing Using ML Techniques and GPU AccelerationabstractLogic gate sizing plays a vital role in timing optimization, especially as Moore's Law slows, shifting greater responsibility to EDA tools to enhance power, performance, and area (PPA), as these gains are no longer achieved solely through scaling and process advancements. There is an increasing need to push the limits of logic gate sizing to extract every possible improvement in PPA. With recent breakthroughs in machine learning (ML) and the computational power of GPUs, there is significant potential to elevate logic gate sizing algorithms to new heights. This contest aims to advance logic gate sizing and push the boundaries of PPA improvement through innovative EDA tools that leverage machine learning and GPU acceleration. As part of the contest, an infrastructure has been developed to enable ML and GPU-accelerated logic gate sizing algorithms, including the release of benchmarks in both standard EDA and ML-friendly formats, along with examples of incorporating "ML inside" EDA tools through Python APIs. The contest leverages the open-source EDA tool OpenROAD and ML-friendly data representation format, CircuitOps, to lower barriers to entry by providing accessible formats and tools, allowing participants to build on existing software without redundancy. With over 25 teams actively participating, the contest highlights growing interest and potential to push the boundaries of timing optimization. Bing-Yue Wu, Rongjian Liang, Geraldo Pradipta, Anthony Agnesina, Haoxing Ren, Vidya A. Chhabria |
ICCAD | 5 |
| 2024 | Invited Paper: LLM4HWDesign Contest: Constructing a Comprehensive Dataset for LLM-Assisted Hardware Code Generation with Community EffortsabstractLarge Language Models (LLMs) show promise in streamlining hardware design, particularly in hardware code generation. However, the development of LLMs for this domain is severely hindered by the scarcity of large-scale, high-quality, and publicly accessible hardware code datasets. This shortage limits the effective fine-tuning of LLMs, impeding their ability to acquire hardware domain knowledge and generate practical designs. To address this challenge, we have organized the first-of-its-kind LLM4HWDesign contest, a community-driven initiative aimed at constructing a large-scale, high-quality dataset for hardware code generation. The contest adopts a two-phase approach, focusing on expanding the scale and quality of an existing hardware code generation dataset, respectively. By harnessing the collective efforts of the hardware design community, the LLM4HWDesign contest seeks to establish a critical resource for advancing LLM-assisted hardware design workflows. The primary goal of this initiative is to deliver a comprehensive dataset compiled from participants' submissions. We hope the released dataset will significantly advance the field of LLM-assisted hardware design and provide substantial benefits to the broader hardware community. Zhongzhi Yu, Chaojian Li, Yongan Zhang, Nathaniel Ross Pinckney, Wenfei Zhou, Rongjian Liang, Haoxing Ren, Yingyan (Celine) Lin |
ICCAD | 9 |
| 2024 | ILILT: Implicit Learning of Inverse Lithography TechnologiesabstractLithography, transferring chip design masks to the silicon wafer, is the most important phase in modern semiconductor manufacturing flow. Due to the limitations of lithography systems, Extensive design optimizations are required to tackle the design and silicon mismatch. Inverse lithography technology (ILT) is one of the promising solutions to perform pre-fabrication optimization, termed mask optimization. Because of mask optimization problems’ constrained non-convexity, numerical ILT solvers rely heavily on good initialization to avoid getting stuck on sub-optimal solutions. Machine learning (ML) techniques are hence proposed to generate mask initialization for ILT solvers with one-shot inference, targeting faster and better convergence during ILT. This paper addresses the question of whether ML models can directly generate high-quality optimized masks without engaging ILT solvers in the loop. We propose an implicit learning ILT framework: ILILT, which leverages the implicit layer learning method and lithography-conditioned inputs to ground the model. Trained to understand the ILT optimization procedure, ILILT can outperform the state-of-the-art machine learning solutions, significantly improving efficiency and quality. Haoxing Ren |
ICML | 2 |
| 2024 | Novel Transformer Model Based Clustering Method for Standard Cell Design AutomationabstractStandard cells are essential components of modern digital circuit designs. With process technologies advancing beyond 5nm, more routability issues have arisen due to the decreasing number of routing tracks (RTs), increasing number and complexity of design rules, and strict patterning rules. The standard cell design automation framework is able to automatically design standard cell layouts, but it is struggling to resolve the severe routability issues in advanced nodes. As a result, a better and more efficient standard cell design automation method that can not only resolve the routability issue but also scale to hundreds of transistors to shorten the development time of standard cell libraries is highly needed and essential. Chia-Tung Ho, Ajay Chandna, David Guan, Alvin Ho, Haoxing Ren |
ISPD | 7 |
| 2024 | GPU/ML-Enhanced Large Scale Global Routing ContestabstractModern VLSI design flows demand scalable global routing techniques applicable across diverse design stages. In response, the ISPD 2024 contest pioneers the first GPU/ML-enhanced global routing competition, selecting advancements in GPU-accelerated computing platforms and machine learning techniques to address scalability challenges. Large-scale benchmarks, containing up to 50 million cells, offer test cases to assess global routers' runtime and memory scalability. The contest provides simplified input/output formats and performance metrics, framing global routing challenges as mathematical optimization problems and encouraging diverse participation. Two sets of evaluation metrics are introduced: the primary one concentrates on global routing applications to guide post-placement optimization and detailed routing, focusing on congestion resolution and runtime scalability. Special honor is given based on the second set of metrics, placing additional emphasis on runtime efficiency and aiming at guiding early-stage planning. Rongjian Liang, Anthony Agnesina, Wen-Hao Liu 0001, Haoxing Ren |
ISPD | 4 |
| 2024 | MedPart: A Multi-Level Evolutionary Differentiable Hypergraph PartitionerabstractState-of-the-art hypergraph partitioners, such as hMETIS, usually adopt a multi-level paradigm for efficiency and scalability. However, they are prone to getting trapped in local minima due to their reliance on refinement heuristics and overlooking global structural information during coarsening. SpecPart, the most advanced academic hypergraph partitioning refinement method, improves partitioning by leveraging spectral information. Still, its success depends heavily on the quality of initial input solutions. This work introduces MedPart, a multi-level evolutionary differentiable hypergraph partitioner. MedPart follows the multi-level paradigm but addresses its limitations by using fast spectral coarsening and introducing a novel evolutionary differentiable algorithm to optimize each coarsening level. Moreover, by analogy between hypergraph partitioning and deep graph learning, our evolutionary differentiable algorithm can be accelerated with deep graph learning toolkits on GPUs. Experiments on public benchmarks consistently show MedPart outperforming hMETIS and achieving up to a 30% improvement in cut size for some benchmarks compared to the best-published solutions, including those from SpecPart---moreover, MedPart's runtime scales linearly with the number of hyperedges. Rongjian Liang, Anthony Agnesina, Haoxing Ren |
ISPD | 3 |
| 2024 | Challenges for Automating PCB LayoutabstractPrinted circuit board (PCB) design is typically semi-automated or fully manual. However, in recent years, the scale of PCB designs has rapidly enlarged, such that the engineering effort of manual design has increased dramatically. Therefore, the criticality of automation emerges. PCB houses are looking for productivity improvement that is contributed by automation. In this talk, the speaker will give a short tutorial about how a PCB design is done today and then indicate the challenges and opportunities for PCB design automation. Wen-Hao Liu 0001, Anthony Agnesina, Haoxing Ren |
ISPD | 3 |
| 2024 | OpenROAD and CircuitOps: Infrastructure for ML EDA Research and EducationabstractTraditional electronic design automation (EDA) techniques struggle to fulfill the stringent efficiency and quick turnaround demands of complex integrated systems. Machine learning (ML) strategies for EDA (“ML EDA”) are pivotal in transforming EDA to address these challenges. However, they encounter significant obstacles due to inadequate infrastructure, ranging from datasets to software interfaces. This paper demonstrates a software infrastructure for ML EDA built on two key technologies: (i) OpenROAD’s Python APIs, and (ii) NVIDIA’s CircuitOps, an EDA data representation format tailored for ML, facilitating ML EDA applications. The paper illustrates three ML EDA examples that utilize the established OpenROAD and CircuitOps infrastructure. Vidya A. Chhabria, Wenjing Jiang, Andrew B. Kahng, Rongjian Liang, Haoxing Ren, Sachin S. Sapatnekar, Bing-Yue Wu |
VTS | 5 |
| 2024 | Domain-Adapted LLMs for VLSI Design and Verification: A Case Study on Formal VerificationabstractLarge language models (LLMs) present unprecedented opportunities in task automation for industrial chip design and verification that can yield significant improvements in engineering productivity. Instead of deploying off-the-shelf LLMs, we present our methodology for adapting a language model to the domain of VLSI design, and we show that our domain-adapted model, ChipNeMo, achieves improved performance against models of similar size on benchmarks concerning chip design and electronic design automation (EDA). We finally present a case study on the prospective of applying LLMs to hardware formal verification. Our results indicate that the largest and most capable models, such as GPT-4, are able to generate syntactically correct SVA implementations, yet there exists room for improvement in ensuring precise reflection of user intent given as high-level natural language descriptions of formal properties. Ghaith Bany Hamad, Syed Suhaib, Haoxing Ren |
VTS | 5 |
| 2024 | DAG-Aware Synthesis OrchestrationabstractModern logic synthesis techniques use multi-level technology-independent representations like And-Inverter-Graphs (AIGs) for digital logic. This involves structural rewriting, resubstitution, and refactoring based on directed-acyclic-graph (DAGs) traversal. Existing DAG-aware logic synthesis algorithms are designed to perform one specific optimization during a single DAG traversal. However, we empirically identify and demonstrate that these algorithms are limited in quality-of-results due to the solely considered optimization operation in the design concept. This work proposes Synthesis Orchestration, which is a fine-grained node-level optimization implying multiple optimizations during the single traversal of the graph. Our experimental results are comprehensively conducted on all 104 designs collected from ISCAS’85/89/99, VTR, and EPFL benchmark suites. The orchestration algorithms consistently outperform existing optimizations, rewriting, resubstitution, refactoring, leading to an average of 4% more node reduction with reasonable runtime cost for the single optimization. Moreover, we evaluate the orchestration algorithm in the sequential optimization, and as a plug-in algorithm in resyn and resyn3 flows in ABC, which demonstrate consistent logic minimization improvements (1%, 4.7% and 11.5% more node reduction on average). Finally, we integrate the orchestration into OpenROAD for end-to-end performance evaluations. Our results demonstrate the advantages of the orchestration optimization techniques, even after technology mapping and post-routing in the design flow. Mingju Liu, Haoxing Ren, Alan Mishchenko, Cunxi Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | GAN-Place: Advancing Open Source Placers to Commercial-quality Using Generative Adversarial Networks and Transfer LearningabstractRecently, GPU-accelerated placers such as DREAMPlace and Xplace have demonstrated their superiority over traditional CPU-reliant placers by achieving orders of magnitude speed up in placement runtime. However, due to their limited focus in placement objectives (e.g., wirelength and density), the placement quality achieved by DREAMPlace or Xplace is not comparable to that of commercial tools. In this article, to bridge the gap between open source and commercial placers, we present a novel placement optimization framework named GAN-Place that employs generative adversarial learning to transfer the placement quality of the industry-leading commercial placer, Synopsys ICC2, to existing open source GPU-accelerated placers (DREAMPlace and Xplace). Without the knowledge of the underlying proprietary algorithms or constraints used by the commercial tools, our framework facilitates transfer learning to directly enhance the open source placers by optimizing the proposed differentiable loss that denotes the “similarity” between DREAMPlace- or Xplace-generated placements and those in commercial databases. Experimental results on seven industrial designs not only show that our GAN-Place immediately improves the Power, Performance, and Area metrics at the placement stage but also demonstrates that these improvements last firmly to the post-route stage, where we observe improvements by up to 8.3% in wirelength, 7.4% in power, and 37.6% in Total Negative Slack on a commercial CPU benchmark. Yi-Chen Lu, Haoxing Ren, Hao-Hsiang Hsiao, Sung Kyu Lim |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | BufFormer: A Generative ML Framework for Scalable BufferingabstractBuffering is a prevalent interconnect optimization technique to help timing closure and is often performed after placement. A common buffering approach is to construct a Steiner tree and then buffers are inserted on the tree based on Ginneken-Lillis style algorithm. Such an approach is difficult to scale with large nets. Our work attempts to solve this problem with a generative machine-learning (ML) approach without Steiner tree construction. Our approach can extract and reuse knowledge from high quality samples and therefore has significantly improved scalability. A generative ML framework, BufFormer, is proposed to construct abstract tree topology while simultaneously determining buffer sizes & locations. A baseline method, FLUTE-based Steiner tree construction followed by Ginneken-Lillis style buffer insertion, is implemented to generate training samples. After training, BufFormer can produce solutions for unseen nets highly comparable to baseline results with a correlation coefficient 0.977 in terms of buffer area and 0.934 for driver-sink delays. On average, BufFormer-generated tree achieves similar delays with slightly larger buffer area. And up to 160X speedup can be achieved for large nets when running on a GPU over the baseline on a single CPU thread. Rongjian Liang, Siddhartha Nath, Anand Rajaram, Jiang Hu 0001, Haoxing Ren |
ASP-DAC | 5 |
| 2023 | Enabling Scalable AI Computational Lithography with Physics-Inspired ModelsabstractComputational lithography is a critical research area for the continued scaling of semiconductor manufacturing process technology by enhancing silicon printability via numerical computing methods. Today's solutions for these problems are primarily CPU-based and require many thousands of CPUs running for days to tape out a modern chip. We seek AI/GPU-assisted solutions for the two problems, aiming at improving both runtime and quality. Prior academic research has proposed using machine learning for lithography modeling and mask optimization, typically represented as image-to-image mapping problems, where convolution layer backboned UNets and ResNets are applied. However, due to the lack of domain knowledge integrated into the framework designs, these solutions have been limited by their application scenarios or performance. Our method aims to tackle the limitations of such previous CNN-based solutions by introducing lithography bias into the neural network design, yielding a much more efficient model design and significant performance improvements. Haoxing Ren |
ASP-DAC | 2 |
| 2023 | GenFuzz: GPU-accelerated Hardware Fuzzing using Genetic Algorithm with Multiple InputsabstractHardware fuzzing has emerged as a promising automatic verification technique to efficiently discover and verify hardware vulnerabilities. However, hardware fuzzing can be extremely time-consuming due to compute-intensive iterative simulations. While recent research has explored several approaches to accelerate hardware fuzzing, nearly all of them are limited to single-input fuzzing using one thread of a CPU-based simulator. As a result, we propose Gen-Fuzz, a GPU-accelerated hardware fuzzer using a genetic algorithm with multiple inputs. Measuring experimental results on a real industrial design, we show that GenFuzz running on a single A6000 GPU and eight CPU cores achieves 80× runtime speed-up when compared to state-of-the-art hardware fuzzers. Dian-Lun Lin, Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany, Shih-Hsin Wang, Tsung-Wei Huang |
DAC | 3 |
| 2023 | Invited Paper: CircuitOps: An ML Infrastructure Enabling Generative AI for VLSI Circuit OptimizationabstractAn innovative ML infrastructure named CircuitOps is developed to streamline dataset generation and model inference for various generative AI (GAI)-based circuit optimization tasks. Addressing the challenges of the absence of a shared Intermediate Representation (IR), steep EDA learning curves, and AI-unfriendly data structures, we propose solutions that empower efficient data handling. Our contributions encompass the following: (1) labeled property graphs (LPGs) as IR for flexible netlist representation and efficient parallel processing; (2) tools-agnostic IR generation from standard EDA files; (3) customizable dataset generation facilitated through AI-friendly LPGs; (4) gRPC-based inference deployment. Compared with using Tcl interfaces of EDA design tools, CircuitOps achieves a significant 99× dataset generation speedup and 75K nets per second transfer throughput, validating its effectiveness in optimizing GAI tasks. Rongjian Liang, Anthony Agnesina, Geraldo Pradipta, Vidya A. Chhabria, Haoxing Ren |
ICCAD | 5 |
| 2023 | Invited Paper: VerilogEval: Evaluating Large Language Models for Verilog Code GenerationabstractThe increasing popularity of large language models (LLMs) has paved the way for their application in diverse domains. This paper proposes a benchmarking framework tailored specifically for evaluating LLM performance in the context of Verilog code generation for hardware design and verification. We present a comprehensive evaluation dataset consisting of 156 problems from the Verilog instructional website HDLBits. The evaluation set consists of a diverse set of Verilog code generation tasks, ranging from simple combinational circuits to complex finite state machines. The Verilog code completions can be automatically tested for functional correctness by comparing the transient simulation outputs of the generated design with a golden solution. We also demonstrate that the Verilog code generation capability of pretrained language models could be improved with supervised fine-tuning by bootstrapping with LLM generated synthetic problem-code pairs. Nathaniel Ross Pinckney, Brucek Khailany, Haoxing Ren |
ICCAD | 4 |
| 2023 | An Adversarial Active Sampling-Based Data Augmentation Framework for AI-Assisted Lithography ModelingabstractLithography modeling is a crucial problem in chip design to ensure the manufacturability of chip design masks. It requires rigorous simulations of optical and chemical models that are computationally expensive. Recent developments in machine learning have provided alternative solutions in replacing time-consuming lithography simulations with deep neural networks. However, considerable accuracy drop still impede its industrial adoption. Most importantly, the quality and quantity of the training dataset directly affects the model performance. To tackle this problem, we propose a Litho-Aware Data Augmentation (LADA) framework to resolve the limited data dilemma and improve the machine learning model performance. First, we pretrain the neural networks for lithography modeling and a gradient-friendly StyleGAN2 generator. We then perform adversarial active sampling to generate informative and synthetic in-distribution mask designs. These synthetic mask images augment the original limited training dataset used to finetune the lithography model for improved performance. Experimental results demonstrate that LADA can successfully improve the model robustness and exploit the neural network capacity by narrowing the performance gap between the training and testing data instances. Brucek Khailany, Haoxing Ren |
ICCAD | 4 |
| 2023 | Reinforcement Learning Guided Detailed Routing for Custom CircuitsabstractDetailed routing is the most tedious and complex procedure in design automation and has become a determining factor in layout automation in advanced manufacturing nodes. Despite continuing advances in custom integrated circuit (IC) routing research, industrial custom layout flows remain heavily manual due to the high complexity of the custom IC design problem. Besides conventional design objectives such as wirelength minimization, custom detailed routing must also accommodate additional constraints (e.g., path-matching) across the analog/mixed-signal (AMS) and digital domains, making an already challenging procedure even more so. This paper presents a novel detailed routing framework for custom circuits that leverages deep reinforcement learning to optimize routing patterns while considering custom routing constraints and industrial design rules. Comprehensive post-layout analyses based on industrial designs demonstrate the effectiveness of our framework in dealing with the specified constraints and producing sign-off-quality routing solutions. Hao Chen 0059, Kai-Chieh Hsu, Walker J. Turner, Po-Hsuan Wei, Keren Zhu 0001, David Z. Pan, Haoxing Ren |
ISPD | 7 |
| 2023 | AutoDMP: Automated DREAMPlace-based Macro PlacementabstractMacro placement is a critical very large-scale integration (VLSI) physical design problem that significantly impacts the design power-performance-area (PPA) metrics. This paper proposes AutoDMP, a methodology that leverages DREAMPlace, a GPU-accelerated placer, to place macros and standard cells concurrently in conjunction with automated parameter tuning using a multi-objective hyperparameter optimization technique. As a result, we can generate high-quality predictable solutions, improving the macro placement quality of academic benchmarks compared to baseline results generated from academic and commercial tools. AutoDMP is also computationally efficient, optimizing a design with 2.7 million cells and 320 macros in 3 hours on a single NVIDIA DGX Station A100. This work demonstrates the promise and potential of combining GPU-accelerated algorithms and ML techniques for VLSI design automation. Anthony Agnesina, Puranjay Rajvanshi, Geraldo Pradipta, Austin Jiao, Ben Keller, Brucek Khailany, Haoxing Ren |
ISPD | 8 |
| 2023 | NVCell 2: Routability-Driven Standard Cell Layout in Advanced Nodes with Lattice Graph Routability ModelabstractStandard cells are essential components of modern digital circuit designs. With process technologies advancing beyond the 5nm node, more routability issues have arisen due to the decreasing number of routing tracks, increasing number and complexity of design rules, and strict patterning rules. Automatic standard cell synthesis tools are struggling to design cells with severe routability issues. In this paper, we propose a routability-driven standard cell synthesis framework using a novel pin density aware congestion metric, lattice graph routability modelling approach, and dynamic external pin allocation methodology to generate routability optimized layouts. On a benchmark of 94 complex and hard-to-route standard cells, NVCell 2 improves the number of routable and LVS/DRC clean cell layouts by 84.0% and 87.2%, respectively. NVCell 2 can generate 98.9% of cells LVS/DRC clean, with 13.9% of the cells having smaller area, compared to an industrial standard cell library with over 1000 standard cells. Chia-Tung Ho, Alvin Ho, Matthew Fojtik, Shang Wei, Brucek Khailany, Haoxing Ren |
ISPD | 8 |
| 2023 | DREAM-GAN: Advancing DREAMPlace towards Commercial-Quality using Generative Adversarial LearningabstractDREAMPlace is a renowned open-source placer that provides GPU-acceleratable infrastructure for placements of Very-Large-Scale-Integration (VLSI) circuits. However, due to its limited focus on wirelength and density, existing placement solutions of DREAMPlace are not applicable to industrial design flows. To improve DREAMPlace towards commercial-quality without knowing the black-boxed algorithms of the tools, in this paper, we present DREAM-GAN, a placement optimization framework that advances DREAMPlace using generative adversarial learning. At each placement iteration, aside from optimizing the wirelength and density objectives of the vanilla DREAMPlace, DREAM-GAN computes and optimizes a differentiable loss that denotes the similarity score between the underlying placement and the tool-generated placements in commercial databases. Experimental results on 5 commercial and OpenCore designs using an industrial design flow implemented by Synopsys ICC2 not only demonstrate that DREAM-GAN significantly improves the vanilla DREAMPlace at the placement stage across each benchmark, but also show that the improvements last firmly to the post-route stage, where we observe improvements by up to 8.3% in wirelength and 7.4% in total power. Yi-Chen Lu, Haoxing Ren, Hao-Hsiang Hsiao, Sung Kyu Lim |
ISPD | 2 |
| 2023 | Introduction to the Special Issue on Machine Learning for CAD/EDAabstractNo abstract available. Yibo Lin, Avi Ziv, Haoxing Ren |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | Generative self-supervised learning for gate sizing: invitedabstractSelf-supervised learning has shown great promise in leveraging large amounts of unlabeled data to achieve higher accuracy than supervised learning methods in many domains. Generative self-supervised learning can generate new data based on the trained data distribution. In this paper, we evaluate the effectiveness of generative self-supervised learning on combinational gate sizing in VLSI designs. We propose a novel use of Transformers for gate sizing when trained on a large dataset generate from a commercial EDA tool. We demonstrate that our trained model can achieve 93% accuracy, 1440X speedup and fast design convergence when compared to a leading commercial EDA tool. Siddhartha Nath, Geraldo Pradipta, Corey Hu, Brucek Khailany, Haoxing Ren |
DAC | 6 |
| 2022 | Generic lithography modeling with dual-band optics-inspired neural networksabstractLithography simulation is a critical step in VLSI design and optimization for manufacturability. Existing solutions for highly accurate lithography simulation with rigorous models are computationally expensive and slow, even when equipped with various approximation techniques. Recently, machine learning has provided alternative solutions for lithography simulation tasks such as coarse-grained edge placement error regression and complete contour prediction. However, the impact of these learning-based methods has been limited due to restrictive usage scenarios or low simulation accuracy. To tackle these concerns, we introduce an dual-band optics-inspired neural network design that considers the optical physics underlying lithography. To the best of our knowledge, our approach yields the first published via/metal layer contour simulation at 1nm2/pixel resolution with any tile size. Compared to previous machine learning based solutions, we demonstrate that our framework can be trained much faster and offers a significant improvement on efficiency and image quality with 20× smaller model size. We also achieve 85× simulation speedup over traditional lithography simulator with ~ 1% accuracy loss. Zongyi Li, Kumara Sastry, Saumyadip Mukhopadhyay, Mark Kilgard, Anima Anandkumar, Brucek Khailany, Haoxing Ren |
DAC | 9 |
| 2022 | GATSPI: GPU accelerated gate-level simulation for power improvementabstractIn this paper, we present GATSPI, a novel GPU accelerated logic gate simulator that enables ultra-fast power estimation for industry-sized ASIC designs with millions of gates. GATSPI is written in PyTorch with custom CUDA kernels for ease of coding and maintainability. It achieves simulation kernel speedup of up to 1668X on a single-GPU system and up to 7412X on a multiple-GPU system when compared to a commercial gate-level simulator running on a single CPU core. GATSPI supports a range of simple to complex cell types from an industry standard cell library and SDF conditional delay statements without requiring prior calibration runs and produces industry-standard SAIF files from delay-aware gate-level simulation. Finally, we deploy GATSPI in a glitch-optimization flow, achieving a 1.4% power saving with a 449X speedup in turnaround time compared to a similar flow using a commercial simulator. Yanqing Zhang 0002, Haoxing Ren, Akshay Sridharan, Brucek Khailany |
DAC | 2 |
| 2022 | Routability-Aware Placement for Advanced FinFET Mixed-Signal Circuits using Satisfiability Modulo TheoriesabstractDue to the increasingly complex design rules and geo-metric layout constraints within advanced FinFET nodes, automated placement of full-custom analog/mixed-signal (AMS) designs has become increasingly challenging. Compared with traditional planar nodes, AMS circuit layout is dramatically different for FinFET technologies due to strict design rules and grid-based restrictions for both placement and routing. This limits previous analog placement approaches in effectively handling all of the new constraints while adhering to the new layout style. Additionally, limited work has demonstrated effective routability modeling, which is crucial for successful routing. This paper presents a robust analog placement framework using satisfiability modulo theories (SMT) for efficient constraint handling and routability modeling. Experimental results based on industrial designs show the effectiveness of the proposed framework in optimizing placement metrics while satisfying the specified constraints. Hao Chen 0059, Walker J. Turner, David Z. Pan, Haoxing Ren |
DATE | 4 |
| 2022 | TAG: Learning Circuit Spatial Embedding from LayoutsabstractAnalog and mixed-signal (AMS) circuit designs still rely on human design expertise. Machine learning has been assisting circuit design automation by replacing human experience with artificial intelligence. This paper presents TAG, a new paradigm of learning the circuit representation from layouts leveraging Text, self Attention and Graph. The embedding network model learns spatial information without manual labeling. We introduce text embedding and a self-attention mechanism to AMS circuit learning. Experimental results demonstrate the ability to predict layout distances between instances with industrial FinFET technology benchmarks. The effectiveness of the circuit representation is verified by showing the transferability to three other learning tasks with limited data in the case studies: layout matching prediction, wirelength estimation, and net parasitic capacitance prediction. Keren Zhu 0001, Hao Chen 0059, Walker J. Turner, George F. Kokai, Po-Hsuan Wei, David Z. Pan, Haoxing Ren |
ICCAD | 7 |
| 2022 | TransSizer: A Novel Transformer-Based Fast Gate SizerabstractGate sizing is a fundamental netlist optimization move and researchers have used supervised learning-based models in gate sizers. Recently, Reinforcement Learning (RL) has been tried for sizing gates (and other EDA optimization problems) but are very runtime-intensive. In this work, we explore a novel Transformer-based gate sizer, TransSizer, to directly generate optimized gate sizes given a placed and unoptimized netlist. TransSizer is trained on datasets obtained from real tapeout-quality industrial designs in a foundry 5nm technology node. Our results indicate that TransSizer achieves 97% accuracy in predicting optimized gate sizes at the postroute optimization stage. Furthermore, TransSizer has a speedup of ~1400× while delivering similar timing, power and area metrics when compared to a leading-edge commercial tool for sizing-only optimization. Siddhartha Nath, Geraldo Pradipta, Corey Hu, Brucek Khailany, Haoxing Ren |
ICCAD | 6 |
| 2022 | Why are Graph Neural Networks Effective for EDA Problems?: (Invited Paper)abstractIn this paper, we discuss the source of effectiveness of Graph Neural Networks (GNNs) in EDA, particularly in the VLSI design automation domain. We argue that the effectiveness comes from the fact that GNNs implicitly embed the prior knowledge and inductive biases associated with given VLSI tasks, which is one of the three approaches to make a learning algorithm physics-informed. These inductive biases are different to those common used in GNNs designed for other structured data, such as social networks and citation networks. We will illustrate this principle with several recent GNN examples in the VLSI domain, including predictive tasks such as switching activity prediction, timing prediction, parasitics prediction, layout symmetry prediction, as well as optimization tasks such as gate sizing and macro and cell transistor placement. We will also discuss the challenges of applications of GNN and the opportunity of applying self-supervised learning techniques with GNN for VLSI optimization. Haoxing Ren, Siddhartha Nath, Yanqing Zhang 0002, Hao Chen 0059 |
ICCAD | 1 |
| 2022 | From RTL to CUDA: A GPU Acceleration Flow for RTL Simulation with Batch StimulusabstractHigh-throughput RTL simulation is critical for verifying today’s highly complex SoCs. Recent research has explored accelerating RTL simulation by leveraging event-driven approaches or partitioning heuristics to speed up simulation on a single stimulus. To further accelerate throughput performance, industry-quality functional verification signoff must explore running multiple stimulus (i.e., batch stimulus) simultaneously, either with directed tests or random inputs. In this paper, we propose RTLFlow, a GPU-accelerated RTL simulation flow with batch stimulus. RTLflow first transpiles RTL into CUDA kernels that each simulates a partition of the RTL simultaneously across multiple stimulus. It also leverages CUDA Graph and pipeline scheduling for efficient runtime execution. Measuring experimental results on a large industrial design (NVDLA) with 65536 stimulus, we show that RTLflow running on a single A6000 GPU can achieve a 40 × runtime speed-up when compared to an 80-thread multi-core CPU baseline. Dian-Lun Lin, Haoxing Ren, Yanqing Zhang 0002, Brucek Khailany, Tsung-Wei Huang |
ICPP | 2 |
| 2022 | AutoCRAFT: Layout Automation for Custom Circuits in Advanced FinFET TechnologiesabstractDespite continuous efforts in layout automation for full-custom circuits, including analog/mixed-signal (AMS) designs, automated layout tools have not yet been widely adopted in current industrial full-custom design flows due to the high circuit complexity and sensitivity to layout parasitics. Nevertheless, the strict design rules and grid-based restrictions in nanometer-scale FinFET nodes limit the degree of freedom in full-custom layout design and thus reduce the gap between automation tools and human experts. This paper presents AutoCRAFT, an automatic layout generator targeting region-based layouts for advanced FinFET-based full-custom circuits. AutoCRAFT uses specialized place-and-route (P&R) algorithms to handle various design constraints while adhering to typical FinFET layout styles. Verified by comprehensive post-layout analyses, AutoCRAFT has achieved promising preliminary results in generating sign-off quality layouts for industrial benchmarks. Hao Chen 0059, Walker J. Turner, Sanquan Song, Keren Zhu 0001, George F. Kokai, Brian Zimmer, C. Thomas Gray, Brucek Khailany, David Z. Pan, Haoxing Ren |
ISPD | 10 |
| 2022 | Embracing Machine Learning in EDAabstractThe application of machine learning (ML) in EDA is a hot research trend. To use ML in EDA, it is nature to think from the ML method point of view, i.e. supervised learning, reinforcement learning and unsupervised learning. Based on this point of view, we can roughly classify the ML applications in EDA into three categories: prediction, optimization, and generation. The prediction category applies supervised learning methods to predict design quality of result (QoR) metrics. There are two kinds of QoR metrics that benefit from the prediction. One kind of metrics are those that can be determined at the current design stage but calculating them consumes a lot of computing resources. For ex-ample, [11] [12] leverage ML to predict circuit power consumption without expensive simulations. The other kind of metrics are those that depend on future design stages. For example, [8] predicts post layout parasitics from schematic of analog circuits. The optimization category applies Bayesian Optimization (BO)and reinforcement learning (RL) to directly optimize EDA problems.BO treats the optimization objective as a blackbox function and tries to find optimal solutions by iteratively sampling the solution space. For example, [5] proposes to use BO with graph embedding and neural network-based surrogate model to size analog circuits. RL treats the optimization objective as the reward from an environment, and trains agents to maximize the reward. [7] proposes to use RL to optimize macro placement, and [9] proposes to use RL to optimize parallel prefix circuit structures. The generation category applies generative models such as generative adversarial networks (GANs) to directly generate solutions to EDA problems. Generative models can learn from previous optimized data distribution and generate solutions for a new problem instance without going through iterative processes like BO or RL. For example, [10] builds a conditional GAN model that learns to generate optical proximity correction (OPC) layout from the original mask. Haoxing Ren |
ISPD | 1 |
| 2021 | Standard Cell Routing with Reinforcement Learning and Genetic Algorithm in Advanced Technology NodesabstractStandard cell layout in advanced technology nodes are done manually in the industry today. Automating standard cell layout process, in particular the routing step, are challenging because of the constraints of enormous design rules. In this paper we propose a machine learning based approach that applies genetic algorithm to create initial routing candidates and uses reinforcement learning (RL) to fix the design rule violations incrementally. A design rule checker feedbacks the violations to the RL agent and the agent learns how to fix them based on the data. This approach is also applicable to future technology nodes with unseen design rules. We demonstrate the effectiveness of this approach on a number of standard cells. We have shown that it can route a cell which is deemed unroutable manually, reducing the cell size by 11%. Haoxing Ren, Matthew Fojtik |
ASP-DAC | 1 |
| 2021 | Invited- NVCell: Standard Cell Layout in Advanced Technology Nodes with Reinforcement LearningabstractHigh quality standard cell layout automation in advanced technology nodes is still challenging in the industry today because of complex design rules. In this paper we introduce an automatic standard cell layout generator called NVCell that can generate layouts with equal or smaller area for over 90% of single row cells in an industry standard cell library on an advanced technology node. NVCell leverages reinforcement learning (RL) to fix design rule violations during routing and to generate efficient placements. Haoxing Ren, Matthew Fojtik |
DAC | 1 |
| 2021 | MAVIREC: ML-Aided Vectored IR-Drop Estimation and ClassificationabstractVectored IR drop analysis is a critical step in chip signoff that checks the power integrity of an on-chip power delivery network. Due to the prohibitive runtimes of dynamic IR drop analysis, the large number of test patterns must be whittled down to a small subset of worst-case IR vectors. Unlike the traditional slow heuristic method that select a few vectors with incomplete coverage, MAVIREC uses machine learning techniques -- 3D convolutions and regression-like layers -- for accurately recommending a larger subset of test patterns that exercise worst-case scenarios. In under 30 minutes, MAVIREC profiles 100K-cycle vectors and provides better coverage than a state-of-the-art industrial flow. Further, MAVIREC's IR drop predictor shows 10x speedup with under 4mV RMSE relative to an industrial flow. Vidya A. Chhabria, Yanqing Zhang 0002, Haoxing Ren, Ben Keller, Brucek Khailany, Sachin S. Sapatnekar |
DATE | 3 |
| 2021 | Parasitic-Aware Analog Circuit Sizing with Graph Neural Networks and Bayesian OptimizationabstractLayout parasitics significantly impact the performance of analog integrated circuits, leading to discrepancies between schematic and post-layout performance and requiring several iterations to achieve design convergence. Prior work has accounted for parasitic effects during the initial design phase but relies on automated layout generation for estimating parasitics. In this work, we leverage recent developments in parasitic prediction using graph neural networks to eliminate the need for in-the-loop layout generation. We propose an improved surrogate performance model using parasitic graph embeddings from the pre-trained parasitic prediction network. We further leverage dropout as an efficient prediction of uncertainty for Bayesian optimization to automate transistor sizing. Experimental results demonstrate the proposed surrogate model has 20% better R2 prediction score and improves optimization convergence by 3.7 times and 2.1 times compared to conventional Gaussian process regression and neural network based Bayesian linear regression, respectively. Furthermore, the inclusion of parasitic prediction in the optimization loop could guarantee satisfaction of all design constraints, while schematic-only optimization fail numerous constraints if verified with parasitic estimations. Walker J. Turner, George F. Kokai, Brucek Khailany, David Z. Pan, Haoxing Ren |
DATE | 6 |
| 2021 | 2021 ICCAD CAD Contest Problem C: GPU Accelerated Logic RewritingabstractLogic rewriting is an important optimization function that can improve Quality of Results (QoR) in modern VLSI circuits. This optimization function usually has a greedy approach and involves steps such as graph traversal, cut computation and ranking, and functional matching. For logic rewriting to be effective in improving the QoR, there should be many local rewriting iterations which can be very slow for industrial level benchmark circuits. One effective solution to speed up the logic rewriting operation is to upload its time consuming steps to Graphics Processing Units (GPUs) to benefit from massively parallel computations that is available there. In this regard, the present contest problem studies the possibility of using GPUs in accelerating a classical logic rewriting function. State-of-the-art large-scale open-source benchmark circuits as well as industrial-level designs will be used to test the GPU accelerated logic rewriting function. Ghasem Pasandi, Sreedhar Pratty, Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
ICCAD | 5 |
| 2021 | Optimizing VLSI Implementation with Reinforcement Learning - ICCAD Special Session PaperabstractReinforcement learning (RL) has gained attention recently as an optimization algorithm for chip design. This method treats many chip design problems as Markov decision problems (MDPs), where design optimization objectives are converted into rewards given by the environment and design variables are converted into actions provided to the environment. Some recent examples include applications of RL to macro placement and standard cell layout routing. We believe RL can be applied to nearly all aspects of VLSI implementation flows, since many VLSI implementation problems are often NP-complete and state-of-art algorithms cannot be guaranteed to be optimal. With enough training data, it is possible to achieve better results with RL. In this paper we review recent advances in applying RL to VLSI implementation problems such as cell layout, synthesis, placement, routing and parameter tuning. We discuss the challenges of applying RL to VLSI implementation flows and propose future research directions for overcoming these challenges. Haoxing Ren, Saad Godil, Brucek Khailany, Robert Kirby 0001, Haiguang Liao, Siddhartha Nath, Jonathan Raiman, Rajarshi Roy 0003 |
ICCAD | 1 |
| 2021 | DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI PlacementabstractPlacement for very large-scale integrated (VLSI) circuits is one of the most important steps for design closure. We propose a novel GPU-accelerated placement framework DREAMPlace, by casting the analytical placement problem equivalently to training a neural network. Implemented on top of a widely adopted deep learning toolkit PyTorch, with customized key kernels for wirelength and density computations, DREAMPlace can achieve around 40× speedup in global placement without quality degradation compared to the state-of-the-art multithreaded placer RePlAce. We believe this work shall open up new directions for revisiting classical EDA problems with advancements in AI hardware and software. Yibo Lin, Zixuan Jiang, Jiaqi Gu 0002, Wuxi Li, Shounak Dhar, Haoxing Ren, Brucek Khailany, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | FIST: A Feature-Importance Sampling and Tree-Based Method for Automatic Design Flow Parameter TuningabstractDesign flow parameters are of utmost importance to chip design quality and require a painfully long time to evaluate their effects. In reality, flow parameter tuning is usually performed manually based on designers' experience in an ad hoc manner. In this work, we introduce a machine learning-based automatic parameter tuning methodology that aims to find the best design quality with a limited number of trials. Instead of merely plugging in machine learning engines, we develop clustering and approximate sampling techniques for improving tuning efficiency. The feature extraction in this method can reuse knowledge from prior designs. Furthermore, we leverage a state-of-the-art XGBoost model and propose a novel dynamic tree technique to overcome overfitting. Experimental results on benchmark circuits show that our approach achieves 25% improvement in design quality or 37% reduction in sampling cost compared to random forest method, which is the kernel of a highly cited previous work. Our approach is further validated on two industrial designs. By sampling less than 0.02% of possible parameter sets, it reduces area by 1.83% and 1.43% compared to the best solutions hand-tuned by experienced designers. Zhiyao Xie, Guanqi Fang, Yu-Hung Huang, Haoxing Ren, Yanqing Zhang 0002, Brucek Khailany, Shao-Yun Fang, Jiang Hu 0001, Yiran Chen 0001, Erick Carvajal Barboza |
ASP-DAC | 4 |
| 2020 | PowerNet: Transferable Dynamic IR Drop Estimation via Maximum Convolutional Neural NetworkabstractIR drop is a fundamental constraint required by almost all chip designs. However, its evaluation usually takes a long time that hinders mitigation techniques for fixing its violations. In this work, we develop a fast dynamic IR drop estimation technique, named PowerNet, based on a convolutional neural network (CNN). It can handle both vector-based and vectorless IR analyses. Moreover, the proposed CNN model is general and transferable to different designs. This is in contrast to most existing machine learning (ML) approaches, where a model is applicable only to a specific design. Experimental results show that PowerNet outperforms the latest ML method by 9% in accuracy for the challenging case of vectorless IR drop and achieves a 30× speedup compared to an accurate IR drop commercial tool. Further, a mitigation tool guided by PowerNet reduces IR drop hotspots by 26% and 31% on two industrial designs, respectively, with very limited modification on their power grids. Zhiyao Xie, Haoxing Ren, Brucek Khailany, Ye Sheng, Santosh Santosh, Jiang Hu 0001, Yiran Chen 0001 |
ASP-DAC | 2 |
| 2020 | ParaGraph: Layout Parasitics and Device Parameter Prediction using Graph Neural NetworksabstractLayout-dependent parasitics and device parameters significantly impact integrated circuit performance and are often the cause of slow convergences between schematic and layout designs. Circuit designers typically estimate parasitics from past experience, resulting in variability between designers and the potential for inaccuracies. In this paper, we present ParaGraph: a graph neural network model to predict net parasitics and device parameters by converting circuit schematics into graphs and leveraging key modeling techniques based on GraphSage, Relation GCN and Graph Attention Networks. Furthermore, the use of ensemble modeling increases model accuracy over a large range of prediction values. Trained on a large dataset of industrial circuits, the model achieves an average prediction R2of 0.772 (110% better than XGBoost) and reduces average simulation errors from over 100% with designer's estimation to less than 10%. Haoxing Ren, George F. Kokai, Walker J. Turner, Ting-Sheng Ku |
DAC | 1 |
| 2020 | GRANNITE: Graph Neural Network Inference for Transferable Power EstimationabstractThis paper introduces GRANNITE, a GPU-accelerated novel graph neural network (GNN) model for fast, accurate, and transferable vector-based average power estimation. During training, GRANNITE learns how to propagate average toggle rates through combinational logic: a netlist is represented as a graph, register states and unit inputs from RTL simulation are used as features, and combinational gate toggle rates are used as labels. A trained GNN model can then infer average toggle rates on a new workload of interest or new netlists from RTL simulation results in a few seconds. Compared to traditional power analysis using gate-level simulations, GRANNITE achieves >18.7X speedup with an error of only <; 5.5% across a diverse set of benchmark circuits. Compared to a GPU-accelerated conventional probabilistic switching activity estimation approach, GRANNITE achieves much better accuracy (on average 25.9% lower error) at similar runtimes. Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
DAC | 2 |
| 2020 | Opportunities for RTL and Gate Level Simulation using GPUs (Invited Talk)abstractThis paper summarizes the opportunities in accelerating simulation on parallel processing hardware platforms such as GPUs. First, we give a summary of prior art. Then, we propose the idea that coding frameworks usually used for popular machine learning (ML) topics, such as PyTorch/DGL.ai, can also be used for exploring simulation purposes. We demo a crude oblivious two-value cycle gate-level simulator using the higher level ML framework APIs that exhibits >20X speedup, despite its simplistic construction. Next, we summarize recent advances in GPU features that may provide additional opportunities to further state-of-the-art results. Finally, we conclude and touch upon some potential areas for furthering research into the topic of GPU accelerated simulation. Yanqing Zhang 0002, Haoxing Ren, Brucek Khailany |
ICCAD | 2 |
| 2020 | Problem C: GPU Accelerated Logic Re-simulation : (Invited Talk)abstractLogic "re"-simulation can be defined as gate level simulation where the input waveforms at every primary input and pseudo-primary input (such as register/RAM outputs) are known. Such waveforms could come from the unit's RTL simulation trace or Automatic Test Pattern Generation (ATPG) vectors. This type of simulation is useful in doing functional verification on gate level netlists and power analysis, since we can take the known trace on all primary and pseudo-primary inputs, re-simulate the trace using propagation of signals through timing-aware gate-level combinational logic, and verify that results at the primary and pseudo-primary outputs match the reference RTL simulation results. However, gate level simulation is usually much slower than RTL simulation. Thus, there is motivation for faster solutions. In this contest, we ask contestants to use Graphic Processing Units (GPUs) to speedup the re-simulation task. Yanqing Zhang 0002, Haoxing Ren, Ben Keller, Brucek Khailany |
ICCAD | 2 |
| 2020 | ABCDPlace: Accelerated Batch-Based Concurrent Detailed Placement on Multithreaded CPUs and GPUsabstractPlacement is an important step in modern verylarge-scale integrated (VLSI) designs. Detailed placement is a placement refining procedure intensively called throughout the design flow, thus its efficiency has a vital impact on design closure. However, since most detailed placement techniques are inherently greedy and sequential, they are generally difficult to parallelize. In this article, we present a concurrent detailed placement framework, ABCDPlace, exploiting multithreading and graphic processing unit (GPU) acceleration. We propose batch-based concurrent algorithms for widely adopted sequential detailed placement techniques, such as independent set matching, global swap, and local reordering. The experimental results demonstrate that ABCDPlace can achieve 2× -5× faster runtime than sequential implementations with multithreaded CPU and over 10× with GPU on ISPD 2005 contest benchmarks without quality degradation. On larger industrial benchmarks, we show more than 16× speedup with GPU over the state-of-the-art sequential detailed placer. ABCDPlace finishes the detailed placement of a 10-million-cell industrial design in 1 min. Yibo Lin, Wuxi Li, Jiaqi Gu 0002, Haoxing Ren, Brucek Khailany, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI PlacementabstractPlacement for very-large-scale integrated (VLSI) circuits is one of the most important steps for design closure. This paper proposes a novel GPU-accelerated placement framework DREAMPlace, by casting the analytical placement problem equivalently to training a neural network. Implemented on top of a widely-adopted deep learning toolkit PyTorch, with customized key kernels for wirelength and density computations, DREAMPlace can achieve over 30× speedup in global placement without quality degradation compared to the state-of-the-art multi-threaded placer RePlAce. We believe this work shall open up new directions for revisiting classical EDA problems with advancement in AI hardware and software. Yibo Lin, Shounak Dhar, Wuxi Li, Haoxing Ren, Brucek Khailany, David Z. Pan |
DAC | 4 |
| 2019 | High Performance Graph Convolutional Networks with Applications in Testability AnalysisabstractApplications of deep learning to electronic design automation (EDA) have recently begun to emerge, although they have mainly been limited to processing of regular structured data such as images. However, many EDA problems require processing irregular structures, and it can be non-trivial to manually extract important features in such cases. In this paper, a high performance graph convolutional network (GCN) model is proposed for the purpose of processing irregular graph representations of logic circuits. A GCN classifier is firstly trained to predict observation point candidates in a netlist. The GCN classifier is then used as part of an iterative process to propose observation point insertion based on the classification results. Experimental results show the proposed GCN model has superior accuracy to classical machine learning models on difficult-to-observation nodes prediction. Compared with commercial testability analysis tools, the proposed observation point insertion flow achieves similar fault coverage with an 11% reduction in observation points and a 6% reduction in test pattern count. Yuzhe Ma, Haoxing Ren, Brucek Khailany, Harbinder Sikka, Lijuan Luo, Karthikeyan Natarajan, Bei Yu 0001 |
DAC | 2 |
| 2019 | PRIMAL: Power Inference using Machine LearningabstractThis paper introduces PRIMAL, a novel learning-based framework that enables fast and accurate power estimation for ASIC designs. PRIMAL trains machine learning (ML) models with design verification testbenches for characterizing the power of reusable circuit building blocks. The trained models can then be used to generate detailed power profiles of the same blocks under different workloads. We evaluate the performance of several established ML models on this task, including ridge regression, gradient tree boosting, multi-layer perceptron, and convolutional neural network (CNN). For average power estimation, ML-based techniques can achieve an average error of less than 1% across a diverse set of realistic benchmarks, outperforming a commercial RTL power estimation tool in both accuracy and speed (15x faster). For cycle-by-cycle power estimation, PRIMAL is on average 50x faster than a commercial gate-level power analysis tool, with an average error less than 5%. In particular, our CNN-based method achieves a 35x speed-up and an error of 5.2% for cycle-by-cycle power estimation of a RISC-V processor core. Furthermore, our case study on a NoC router shows that PRIMAL can achieve a small estimation error of 4.5% using cycle-approximate traces from SystemC simulation. Haoxing Ren, Yanqing Zhang 0002, Ben Keller, Brucek Khailany, Zhiru Zhang |
DAC | 2 |
| 2019 | Routability-Driven Macro Placement with Embedded CNN-Based Prediction ModelabstractWith the dramatic shrink of feature size and the advance of semiconductor technology nodes, numerous and complicated design rules need to be followed, and a chip design can only be taped-out after passing design rule check (DRC). The high design complexity seriously deteriorates design routability, which can be measured by the number of DRC violations after the detailed routing stage. In addition, a modern large-scaled design typically consists of many huge macros due to the wide use of intellectual properties (IPs). Empirically, the placement of these macros greatly determines routability, while there exists no effective cost metric to directly evaluate a macro placement because of the extremely high complexity and unpredictability of cell placement and routing. In this paper, we propose the first work of routability-driven macro placement with deep learning. A convolutional neural network (CNN)-based routability prediction model is proposed and embedded into a macro placer such that a good macro placement with minimized DRC violations can be derived through a simulated annealing (SA) optimization process. Experimental results show the accuracy of the predictor and the effectiveness of the macro placer. Yu-Hung Huang, Zhiyao Xie, Guanqi Fang, Tao-Chun Yu, Haoxing Ren, Shao-Yun Fang, Yiran Chen 0001, Jiang Hu 0001 |
DATE | 5 |
| 2019 | Toward Intelligent Physical Design: Deep Learning and GPU AccelerationabstractDeep learning (DL) has achieved tremendous success in computer vision, natural language processing and gaming. Would DL help push physical design toward a more intelligent paradigm to meet the post-Moore era design automation challenges? We will discuss several supervised deep learning models and our works applying these models in physical design and related EDA domains. In addition to supervised learning, we will also illustrate that reinforcement learning and unsupervised learning with deep learning can be applied to design automation tasks. Furthermore, we will show the potential of GPU accelerated physical design with an example of placement engines implemented with deep learning frameworks such as TensorFlow and PyTorch. With deep learning and GPU accelerated algorithms, we could make physical design more intelligent. Haoxing Ren |
ISPD | 1 |
| 2018 | RouteNet: routability prediction for mixed-size designs using convolutional neural networkabstractEarly routability prediction helps designers and tools perform preventive measures so that design rule violations can be avoided in a proactive manner. However, it is a huge challenge to have a predictor that is both accurate and fast. In this work, we study how to leverage convolutional neural network to address this challenge. The proposed method, called RouteNet, can either evaluate the overall routability of cell placement solutions without global routing or predict the locations of DRC (Design Rule Checking) hotspots. In both cases, large macros in mixed-size designs are taken into consideration. Experiments on benchmark circuits show that RouteNet can forecast overall routability with accuracy similar to that of global router while using substantially less runtime. For DRC hotspot prediction, RouteNet improves accuracy by 50% compared to global routing. It also significantly outperforms other machine learning approaches such as support vector machine and logistic regression. Zhiyao Xie, Yu-Hung Huang, Guanqi Fang, Haoxing Ren, Shao-Yun Fang, Yiran Chen 0001, Jiang Hu 0001 |
ICCAD | 4 |
| 2013 | Intuitive ECO synthesis for high performance circuitsabstractIn the IC industry, chip design cycles are becoming more compressed, while designs themselves are growing in complexity. These trends necessitate efficient methods to handle late-stage engineering change orders (ECOs) to the functional specification, often in response to errors discovered after much of the implementation is finished. Past ECO synthesis algorithms have typically treated ECOs as functional errors and applied error diagnosis techniques to solve them. However, error diagnosis methods are primarily geared towards finding a single change, and moreover, tend to be computationally complex. In this paper, we propose a unique methodology that can systematically incorporate human intuition into the ECO process. Our methodology involves finding a set of directly substitutable points known as functional correspondences between the original implementation and the new specification by using name-preserving synthesis and user hints, to diminish the size of the ECO problem. On average, our approach can reduce the size of logic changes by 94% from those reported in current literature. We then incorporate our logic ECO changes into an incremental physical synthesis flow to demonstrate its usability in an industrial setting. Our ECO synthesis methodology is evaluated on high-performance industrial designs. Results indicate that post-ECO worst negative slack (WNS) improved 14% and total negative slack (TNS) improved 46% over pre-ECO. Haoxing Ren, Ruchir Puri, Lakshmi N. Reddy, Smita Krishnaswamy, Cindy Washburn, Joel Earl, Joachim Keinert |
DATE | 1 |
| 2013 | LatchPlanner: latch placement algorithm for datapath-oriented high-performance VLSI designsabstractIn this paper, we present a novel algorithm for latch placement, LatchPlanner which enables a placement engine to deliver high quality placement for datapath-oriented design. Datapath-oriented VLSI designs are in general hand-crafted by human at high cost, as understanding and capturing datapath structure is critical for the performance. The conventional placement algorithms by itself cannot exploit the underlying datapath due to lack of logic structure recognition and inaccurate/approximated wirelength estimation. LatchPlanner addresses such drawbacks by placing and fixing latches in the datapath context, a key element in datapath structure. By taking placed/fixed latches as constraints, a placer can find a more datapath-friendly placement effectively, which results in higher-quality hardware. LatchPlanner begins latch clustering/sizing/ordering to prepare the following steps, a) global latch placement based on linear programming to place latch clusters, and b) local latch placement based on network flow optimization to place latches within each cluster. Experimental results on eighteen industrial benchmarks show that LatchPlanner improves total wirelength by 32%, total negative slack by 25%, and area by 3% without CPU overhead over a commercial placement engine, and delivers near semi-custom-quality solutions. Minsik Cho, Hua Xiang 0001, Haoxing Ren, Matthew M. Ziegler, Ruchir Puri |
ICCAD | 3 |
| 2013 | Network flow based datapath bit slicingabstractIn deep sub-micro designs, more functions are integrated into one chip, and datapath has become a critical part of the design. Typical datapath consists an array of bit slices. The inherent high degree regularity of datapaths is especially attractive to the placement and routing to achieve regular layout with high density and high performance. However, the current design methodology may generate inferior datapath designs because the datapath regularity cannot be well understood by the traditional design tools. In previous works, several techniques are proposed to preserve/re-identify datapath structures. However, they either restrict the datapath optimization or have little tolerance on bit slice difference. Hua Xiang 0001, Minsik Cho, Haoxing Ren, Matthew M. Ziegler, Ruchir Puri |
ISPD | 3 |
| 2010 | History-based VLSI legalization using network flowabstractIn VLSI placement, legalization is an essential step where the overlaps between gates/macros must be removed. In this paper, we introduce a history-based legalization algorithm with min-cost network flow optimization. We find a legal solution with the minimum deviation from a given placement to fully honor/preserve the initial placement, by solving a gate-centric network flow formulation in an iterative manner. In order to realize a flow into gate movements, we develop efficient techniques which solve an approximated Subset-sum problem. Over the iterations, we factor into our formulation the history which captures a set of likely-to-fail gate movements. Such a history-based scheme enables our algorithm to intelligently legalize highly complex designs. Experimental results on over 740 real cases show that our approach is significantly superior to the existing algorithms in terms of failure rate (no failure) as well as quality of results (55% less max-deviation). Minsik Cho, Haoxing Ren, Hua Xiang 0001, Ruchir Puri |
DAC | 2 |
| 2010 | Logical and physical restructuring of fan-in treesabstractA symmetric-function fan-in tree (SFFT) is a fanout-free cone of logic that computes a symmetric function, so that all of the leaf nets in its support set are commutative. Such trees are frequently found in designs, especially when the design originated as two-level logic.These trees are usually created during logic synthesis, when there is no knowledge of the locations of the tree root or of the source gates of the leaf nets. Because of this, large SFFTs present a challenge to placement algorithms. The result is that the tree placements are generally far from optimal, leading to wiring congestion, excess buffering, and timing problems. Restructuring such trees can produce a more placeable and wire-efficient design.In this paper, we propose algorithms to identify and to restructure SFFTs during physical design. The key feature of an SFFT is that it can be implemented with various structures of a uniform set of gates with commutative inputs, i.e. AND, OR, or XOR. Drawing on the flexibility of SFFT logic structures, the proposed tree restructuring algorithm uses existing placement information to rebuild the SFFTs with reduced tree wire lengths. The experimental results demonstrate the efficiency and effectiveness of the algorithms. Hua Xiang 0001, Haoxing Ren, Louise Trevillyan, Lakshmi N. Reddy, Ruchir Puri, Minsik Cho |
ISPD | 2 |
| 2009 | DeltaSyn: An efficient logic difference optimizer for ECO synthesisabstractDuring the IC design process, functional specifications are often modified late in the design cycle, after placement and routing are completed. However, designers are left either to manually process such modifications by hand or to restart the design process from scratch---a very costly option. In order to address this issue, we present DeltaSyn, a method for generating a highly optimized logic difference between a modified high-level specification and an implemented design. DeltaSyn has the ability to locate boundaries in implemented logic within which changes can be confined. DeltaSyn demarcates the boundary in two phases. The first phase employs fast functional and structural analysis techniques to identify equivalent signals forming the input-side boundary of the changes. The second phase locates the outputside boundary of the changes through a novel dynamic algorithm that detects matching logic downstream from the changes required by the ECO. Experiments on industrial designs show that together these techniques successfully implement ECOs while preserving an average of 97% of the existing logic. Unlike previous approaches, the use of bitparallel logic simulation and fast SAT solvers enables high performance and scalability. DeltaSyn can process and verify a typical ECO for a design of around 10K gates in about 200 seconds or less. Smita Krishnaswamy, Haoxing Ren, Nilesh Modi, Ruchir Puri |
ICCAD | 2 |
| 2009 | Low cost test point insertion without using extra registers for high performance designabstractThis paper presents a new approach to improve random test coverage during physical synthesis for high performance design. This new approach performs test point insertion (TPI) to improve testability of random resistant nets. Conventional test point insertion approaches add extra registers either as control points or observation points to improve controllability or observability. However, adding extra registers has many disadvantages in high performance design. It might degrade the design performance in term of power and timing. It might also end up with even worse testability. The new approach does not add any registers; instead it only uses existing signals as test points. The test points are selected from logic paths with the most timing slack and physical placement proximity. This saves silicon area, reduces power consumption, and minimizes the design changes. The new approach also automates the test insertion process by integrating it in the physical synthesis flow. The integration helps reduce design closure turn around time significantly. Production results show nearly zero performance degradation from this approach and better testability improvement as compared to a manual test point insertion approach. Furthermore, this paper proposes a novel idea of exploiting unreachable states for test point insertion. Preliminary results on this idea are also given. Haoxing Ren, Mary P. Kusko, Victor N. Kravets, Rona Yaari |
ITC | 1 |
| 2007 | Hippocrates: First-Do-No-Harm Detailed PlacementabstractPhysical synthesis optimizations and engineering change orders typically change the locations of cells, resize cells or add more cells to the design after global placement. Unfortunately, those changes usually lead to wirelength increases; thus another pass of optimizations to further improve wirelength, timing and routing congestion characteristics is required. Simple wirelength-driven detailed placement techniques could be useful in this scenario. While such techniques can help to reduce wirelength, ones without careful timing constraint considerations might degrade the timing characteristics (worst negative slack, total negative slack, etc) and/or introduce more electrical violations (exceeding maximum output load constraints and maximum input slew constraints). In this paper, we propose a new detailed placement paradigm, which use a set of pin-based timing and electrical constraints in detailed placement to prevent it from degrading timing or violating electrical constraints while reducing wire-length, thus dubbed as Hippocrates: FIRST-DO-NO-HARM optimizations. Our experimental results show great promises. By honoring these constraints, our detailed placement technique not only reduces total wirelength (TWL), but also significantly improves timing, achieving 37% better total negative slack (TNS). Haoxing Ren, David Z. Pan, Charles J. Alpert, Gi-Joon Nam, Paul G. Villarrubia |
ASP-DAC | 1 |
| 2007 | RQL: Global Placement via Relaxed Quadratic Spreading and LinearizationabstractThis paper describes a simple and effective quadratic placement algorithm called RQL. We show that a good quadratic placement, followed by local wirelength-driven spreading can produce excellent results on large-scale industrial ASIC designs. As opposed to the current top performing academic placers [4, 7, 11], RQL does not embed a linearization technique within the solver. Instead, it only requires a simpler, pure quadratic objective function in the spirit of [8, 10, 23]. Experimental results show that RQL outperforms all available academic placers on the ISPD-2005 placement contest benchmarks. In particular, RQL obtains an average wire-length improvement of 2.8%, 3.2%, 5.4%, 8.5%, and 14.6% versus mPL6 [5], NTUPlace3 [7], Kraftwerk [20], APlace2.0 [11], and Capo10.2 [18], respectively. In addition, RQL is three, seven, and ten times faster than mpL6, Capo10.2, and APlace2.0, respectively. On the ISPD-2006 placement contest benchmarks, on average, RQL obtains the best scaled wirelength among all available academic placers. Natarajan Viswanathan, Gi-Joon Nam, Charles J. Alpert, Paul G. Villarrubia, Haoxing Ren, Chris C. N. Chu |
DAC | 5 |
| 2007 | Techniques for Fast Physical SynthesisabstractThe traditional purpose of physical synthesis is to perform timing closure , i.e., to create a placed design that meets its timing specifications while also satisfying electrical, routability, and signal integrity constraints. In modern design flows, physical synthesis tools hardly ever achieve this goal in their first iteration. The design team must iterate by studying the output of the physical synthesis run, then potentially massage the input, e.g., by changing the floorplan, timing assertions, pin locations, logic structures, etc., in order to hopefully achieve a better solution for the next iteration. The complexity of physical synthesis means that systems can take days to run on designs with multimillions of placeable objects, which severely hurts design productivity. This paper discusses some newer techniques that have been deployed within IBM's physical synthesis tool called PDS that significantly improves throughput. In particular, we focus on some of the biggest contributors to runtime, placement, legalization, buffering, and electric correction, and present techniques that generate significant turnaround time improvements Charles J. Alpert, Shrirang K. Karandikar, Zhuo Li 0001, Gi-Joon Nam, Stephen T. Quay, Haoxing Ren, Cliff C. N. Sze, Paul G. Villarrubia, Mehmet Can Yildiz |
Proc. IEEE | 6 |
| 2007 | Diffusion-Based Placement Migration With Application on LegalizationabstractPlacement migration is the movement of cells within an existing placement to address a variety of postplacement design-closure issues, such as timing, routing congestion, signal integrity, and heat distribution. To fix a design problem, one would like to perturb the design as little as possible while preserving the integrity of the original placement. This paper presents a new diffusion-based placement method based on a discrete approximation to the closed-form solution of the continuous diffusion equation. It has the advantage of smooth spreading, which helps preserve neighborhood characteristics of the original placement. Applying this technique to placement legalization demonstrates significant improvements in wire length and timing compared with other commonly used techniques. Haoxing Ren, David Z. Pan, Charles J. Alpert, Paul G. Villarrubia, Gi-Joon Nam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Diffusion-based placement migrationabstractPlacement migration is the movement of cells within an existing placement to address a variety of post-placement design closure issues, such as timing, routing congestion, signal integrity, and heat distribution. To fix a design problem, one would like to perturb the design as little as possible while preserving the integrity of the original placement. This work presents a new diffusion-based placement method based on a discrete approximation to a closedform solution of the continuous diffusion equation. It has the advantage of smooth spreading, which helps preserve neighborhood characteristics of the original placement. Applying this technique to placement legalization demonstrates significant improvements in wire length and timing compared to other commonly used techniques. Haoxing Ren, David Z. Pan, Charles J. Alpert, Paul G. Villarrubia |
DAC | 1 |
| 2005 | Computational geometry based placement migrationabstractPlacement migration is a critical step to address a variety of post-placement design closure issues, such as timing, routing congestion, signal integrity, and heat distribution. To fix a design problem, one would like to perturb the design as little as possible while preserving the integrity of the original placement. This work presents a novel computational geometry based placement migration method, and a new stability metric to more accurately measure the "similarity" between two placements. It has two stages, a bin-based spreading at coarse scale and a Delaunay triangulation based spreading at finer grain. It has clear advantage over conventional legalization algorithms such that the neighborhood characteristics of the original placement are preserved. Thus, the placement migration is much more stable, which is important to maintain. Applying this technique to placement legalization demonstrates significant improvements in wire length and stability compared to other popular legalization algorithms. Tao Luo 0002, Haoxing Ren, Charles J. Alpert, David Z. Pan |
ICCAD | 2 |
| 2005 | Sensitivity guided net weighting for placement-driven synthesisabstractNet weighting is a key technique in timing-driven placement (TDP), which plays a crucial role for deep submicron very large scale integration of physical synthesis and timing closure. A popular way to assign net weight is based on its slack, such that the worst negative slack (WNS) of the entire circuit may be minimized. While WNS is an important optimization metric, another figure of merit (FOM), defined as the total slack difference compared to a certain slack threshold for all timing end points, is of equal importance to measure the overall timing closure result for highly complex modern application specific integrated circuits and microprocessor designs. Moreover, to optimally assign net weight for timing closure, the effect of net weighting on timing should be carefully studied. In this paper, we perform a comprehensive analysis of the wirelength, slack, and FOM sensitivities to the net weight, and propose a new net weighting scheme based on those sensitivities. Such sensitivity analysis implicitly takes potential physical synthesis effect into consideration. The experiments on a set of industrial circuits show promising results for both stand-alone TDP and physical synthesis afterwards. Haoxing Ren, David Z. Pan, David S. Kung 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | True crosstalk aware incremental placement with noise mapabstractCrosstalk noise has become an important issue as technology scales down for timing and signal integrity closure. Existing works to fix crosstalk noise are mostly done at the routing or post routing stage, which may be too late. Since placement determines the overall routing congestion, which correlates with the coupling capacitance, which in turn correlates with the crosstalk noise, placement shall be a good level to do early noise mitigation. The only existing work for the crosstalk aware placement (to our best knowledge) is by Lou and Chen (2004), which uses the coupling capacitance map to guide placement. However, crosstalk is determined not only by the coupling capacitance, but also by many other factors, such as the driver resistance of the victim net and the coupling location (near source vs near sink coupling) (Cong et al., 2001). We introduce a concept of noise map which takes those factors into account. Guided by this accurate noise map explicitly, we propose an incremental placement technique to mitigate noise without disturbing the global placement order. Our incremental placement has two key steps, namely noise aware cell inflation and local refinement. Experimental results on industrial circuits show that our approach is able to reduce the number of top noise nets by 25% and improve the timing (300ps on the worst slack), with no wire length penalty or CPU overhead. Our incremental approach is also able to maintain the placement stability. Haoxing Ren, David Z. Pan, Paul G. Villarrubia |
ICCAD | 1 |
| 2004 | Sensitivity guided net weighting for placement driven synthesisabstractNet weighting is a key technique in large scale timing driven placement, which plays a crucial role for deep submicron physical synthesis and timing closure. A popular way to assign net weight is based on the slack of the nets, trying to minimize the worst negative slack (WNS) for the entire circuit. While WNS is an important optimization metric, another figure of merit (FOM), defined as the total slack difference compared to a certain slack threshold for all timing end points, is of equivalent importance to measure the overall timing closure result for highly complex modern ASIC and microprocessor designs. In this paper, we perform a comprehensive analysis of the slack and FOM sensitivities to the net weight, and propose a new net weighting scheme based on the slack and FOM sensitivities. Such sensitivity analysis implicitly takes potential physical synthesis effect into consideration. Experiment results on a set of industrial circuits are promising for both stand-alone timing driven placement and physical synthesis afterwards. Haoxing Ren, David Z. Pan, David S. Kung 0001 |
ISPD | 1 |