EDBT 2026 Demo / reviewers in the wild / expert
Zhiang Wang
dblp:309/4641
· DBLP profile ↗
27ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0002-6669-9702ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 27 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PigMap3: A Physically Aware Incremental Mapping Framework with On-the-fly Post-Layout Critical Path Tracking
Hongyang Pan, Cunqing Lan, Zhiang Wang, Xuan Zeng 0001, Fan Yang 0001, Keren Zhu 0001 |
ASP-DAC | 3 |
| 2026 | HOLMES: Hierarchical Optimization with poLygonal ModEling for Large-Scale AMS Placement
Yujie Yan, Jiahua Liu, Zecheng Xu, Linxi Qiu, Yumao Wu, Zhiang Wang, Changhao Yan, Zhaori Bi, Keren Zhu 0001 |
ISCAS | 7 |
| 2026 | LSMC Meets GPU Acceleration: Scalable and High-Quality Multi-Row Detailed Placement
Andrew B. Kahng, Jason Liang, Zhiang Wang |
ISCAS | 3 |
| 2026 | PhySeqForm: A Data-Driven, Physical Synthesis Sequence Former
Cunqing Lan, Zijian Jiang, Hongyang Pan, Zhiang Wang, Keren Zhu 0001 |
ISCAS | 4 |
| 2026 | RC-Scaled Timing-Driven Routing: Bridging Targeted Timing Optimization and Massively
Parallel Global Routing, Zecheng Xu, Boxiang Song, Zhiang Wang, Fan Yang 0001, Keren Zhu 0001, Xuan Zeng 0001 |
ISCAS | 5 |
| 2026 | Layout-Aware Standard Cell Synthesis via Reparameterization Multi-Task Bayesian Optimization
Zhouyang Wu, Ruiyu Lyu, Keren Zhu 0001, Zhiang Wang, Zhaori Bi, Changhao Yan, Xuan Zeng 0001 |
ISCAS | 4 |
| 2026 | Invited: Toward Sustainable and Transparent Benchmarking for Academic Physical Design ResearchabstractThis paper presents RosettaStone 2.0, an open benchmark translation and evaluation framework built on OpenROAD-Research [1]. RosettaStone 2.0 provides complete RTL-to-GDS reference flows for both conventional 2D designs and Pin-3D-style face-to-face (F2F) hybrid-bonded 3D designs, enabling rigorous apples-to-apples comparison across planar and three-dimensional implementation settings. The framework is integrated within OpenROAD-flowscripts (ORFS)-Research [2]; it incorporates continuous integration (CI)-based regression testing and provides a standardized evaluation pipeline based on the METRICS2.1 convention, with structured logs and reports generated by ORFS-Research. To support transparent and reproducible research, RosettaStone 2.0 further provides a community-facing leaderboard, which is governed by verified pull requests and enforced through Developer Certificate of Origin (DCO) compliance. Andrew B. Kahng, Zhiang Wang, Zhiyu Zheng |
ISPD | 3 |
| 2026 | Invited: Post-Placement Buffering and Sizing ContestabstractThe ISPD 2026 Contest [22] challenges participants to develop post-detailed placement buffering and sizing tools that optimize timing and fix electrical rule check (ERC) violations under real-world constraints. Unlike prior contests, this contest emphasizes practical physical design challenges including fixed macros and I/Os, power delivery network (PDN) blockages, soft placement blockages, and fixed routing resources. The contest provides eight public benchmarks and four hidden benchmarks, with a range from 15K to 1.4M instances, in the ASAP7 7nm technology node [4] with multi-threshold voltage cell libraries. Evaluation is performed using the open-source OpenROAD infrastructure, with scoring based on timing (total negative slack), power (dynamic and leakage) and penalties for ERC violations, displacement, routing congestion and runtime. This paper describes the contest problem formulation, benchmarks, evaluation methodology, a review of related contests and a two-year roadmap for continuation in the ISPD 2027 Contest. Andrew B. Kahng, Seokhyeong Kang, Sayak Kundu, Yiting Liu 0002, Davit Markarian, Seonghyeon Park, Zhiang Wang |
ISPD | 7 |
| 2026 | An Updated Assessment of Reinforcement Learning for Macro PlacementabstractWe provide an improved assessment of Google Brain’s deep reinforcement learning approach to macro placement [29] and its updated Circuit Training (CT) implementation in GitHub [53]. A stronger simulated annealing (SA) baseline leverages the “go-with-the-winners” metaheuristic [3] and a multi-threading implementation. We develop and release new public benchmarks in sub-10nm technology: LEF/DEF for Google’s 7nm TSMC Ariane protobuf and scaled variants, as well as testcases implemented in the open-source ASAP7 7nm research enablement. We evaluate from-scratch training and fine-tuning results for the latest “AlphaChip” release of Circuit Training, alongside multiple alternative macro placers. We also study the recently-published pre-training guidance in [53]. A commercial place-and-route tool is used to provide “true reward” post-route power, performance and area metrics. All data, evaluation flows and related scripts are publicly available in theMacroPlacementGitHub repository [63]. Our study affords insights into reproducibility and reporting in the research literature, and points out still-missing confirmations (e.g., of CT’s scalability and pre-training methodology) that remain open questions for the research community. Chung-Kuan Cheng, Andrew B. Kahng, Sayak Kundu, Yucheng Wang 0016, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | PigMap2: A Physical Information-Guided Technology Mapping Framework
Cunqing Lan, Hongyang Pan, Zhiang Wang, Xuan Zeng 0001, Fan Yang 0001, Keren Zhu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | ChipletPart: Cost-Aware Partitioning for 2.5D SystemsabstractIndustry adoption of chiplets has been growing as chiplets are a cost-effective option for making large, high-performance systems. Consequently, partitioning large systems into chiplets is increasingly important. In this work, we introduce ChipletPart —a cost-driven 2.5D system partitioner that addresses the unique constraints of chiplet systems, including complex objective functions, limited reach of inter-chiplet I/O transceivers, and the assignment of heterogeneous manufacturing technologies to different chiplets. ChipletPart integrates a sophisticated chiplet cost model with a genetic algorithm (GA)-based technology assignment and partitioning methodology, along with a simulated annealing (SA)-based chiplet floorplanner. Our results show that ChipletPart : (i) reduces chiplet cost by up to 58% (20% geometric mean) compared to state-of-the-art min-cut partitioners, which often yield floorplan-infeasible solutions; (ii) generates partitions with up to 47% (6% geometric mean) lower cost compared to the prior work Floorplet ; (iii) reduces chiplet cost up to 48% (30% geometric mean) compared to Chipletizer , while consistently producing I/O-feasible chiplet solutions across all testcases; and (iv) for the testcases we study, heterogeneous integration reduces cost by up to 43% (15% geometric mean) compared to homogeneous implementations. Additionally, we explore Bayesian optimization (BO) for finding low cost and floorplan-feasible chiplet solutions with technology assignments. On some testcases, our BO framework achieves better system cost (up to 5.3% improvement) with higher runtime overhead (up to 4×) compared to our GA-based framework. We also present case studies that show how changes in packaging and inter-chiplet signaling technologies can affect partitioning solutions. Finally, ChipletPart , the underlying chiplet cost model, and our chiplet testcase generator are available as open-source tools for the community. Alexander Graening, Puneet Gupta 0001, Andrew B. Kahng, Bodhisatta Pramanik, Zhiang Wang |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | Use Cases and Deployment of ML in IC Physical DesignabstractML for IC physical design must be deployed in order to have business impacts. However, deployment in production must navigate many practical considerations, including choice of targets, skillsets and infrastructure, expectations and resources, data, and "MLOps". Furthermore, usage of ML is not the same as IC design practice and capability. In this invited paper, we give perspectives on basic strategies for selecting applications and pursuing deployment for ML in IC physical design. Example aspects include checklists for data and ML models, evaluation of model performance and progress on the path to deployment, the shifting landscape of MLOps, and challenges of "LLM-ability". Amur Ghose, Andrew B. Kahng, Sayak Kundu, Yiting Liu 0002, Bodhisatta Pramanik, Zhiang Wang, Dooseok Yoon |
ASP-DAC | 6 |
| 2025 | Invited: IEEE DATC RDF-2025: Enabling an EDA Research EcosystemabstractOver the past year, IEEE CEDA DATC has continued to improve the DATC Robust Design Flow (RDF) while continuing to expand initiatives that advance open infrastructures and culture changes, serving the global community of EDA researchers and users. This invited paper focuses on three highlights: (1) establishment of an accessible, "contrib-like" GitHub resource that provides a more accessible environment for OpenROAD- and OpenROAD-flow-scripts-based research works; (2) the first-ever permission mechanism and benchmarking results for a commercial EDA P&R tool, published with permissions developed with the tool vendor (Siemens EDA); and (3) efforts that support a nascent "ML EDA Commons". The paper also provides brief reviews of the past year’s RDF developments and roadmap updates. Vidya A. Chhabria, Amur Ghose, Vikram Gopalakrishnan, Andrew B. Kahng, Sayak Kundu, Yiting Liu 0002, Zhiang Wang, Bing-Yue Wu |
ICCAD | 7 |
| 2025 | DG-RePlAce: A Dataflow-Driven GPU-Accelerated Analytical Global Placement Framework for Machine Learning AcceleratorsabstractGlobal placement (GP) is a fundamental step in VLSI physical design. The wide use of 2-D processing element (PE) arrays in machine learning accelerators poses new challenges of scalability and quality of results (QoR) for state-of-the-art academic global placers. In this work, we develop DG-RePlAce, a new and fast GPU-accelerated GP framework built on top of the OpenROAD infrastructure, which exploits the inherent dataflow and datapath structures of machine learning accelerators. Experimental results with a variety of machine learning accelerators using a commercial 12-nm enablement show that, compared with RePlAce (DREAMPlace), our approach achieves an average reduction in routed wirelength by$10\%~(7\%)$and total negative slack (TNS) by$31\%~(34\%)$, with faster GP and on-par total runtimes relative to DREAMPlace. Empirical studies on the TILOS MacroPlacement Benchmarks further demonstrate that post-route improvements over RePlAce and DREAMPlace may reach beyond the motivating application to machine learning accelerators. Andrew B. Kahng, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Performance Analysis of CNN Inference/Training with Convolution and Non-Convolution Operations on ASIC AcceleratorsabstractToday’s performance analysis frameworks for deep learning accelerators suffer from two significant limitations. First, although modern convolutional neural networks (CNNs) consist of many types of layers other than convolution, especially during training, these frameworks largely focus on convolution layers only. Second, these frameworks are generally targeted towards inference and lack support for training operations. This work proposes a novel open-source performance analysis framework, SimDIT, for general ASIC-based systolic hardware accelerator platforms. The modeling effort of SimDIT comprehensively covers convolution and non-convolution operations of both CNN inference and training on a highly parameterizable hardware substrate. SimDIT is integrated with a backend silicon implementation flow and provides detailed end-to-end performance statistics (i.e., data access cost, cycle counts, energy, and power) for executing CNN inference and training workloads. SimDIT-enabled performance analysis reveals that on a 64×64 processing array, non-convolution operations constitute 59.5% of total runtime for ResNet-50 training workload. In addition, by optimally distributing available off-chip DRAM bandwidth and on-chip SRAM resources, SimDIT achieves 18× performance improvement over a generic static resource allocation for ResNet-50 inference. Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew B. Kahng, Sean Kinzer, Susmita Dey Manasi, Sachin S. Sapatnekar, Zhiang Wang |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2024 | Strengthening the Foundations for IC Physical Design and ML EDA ResearchabstractOver the past year, IEEE CEDA DATC has continued to improve the DATC Robust Design Flow (RDF) while also advancing open infrastructure for research, including machine learning for electronic design automation (ML EDA). The 2024 RDF release includes new standalone and integrated global placement and macro placement engines, as well as a CCS-based delay calculator. Advances in baselines and benchmarks include the addition of new benchmarks for macro placement and logic gate sizing, as well as further efforts to establish calibrations of both optimizations and analyses to aid assessments of research progress in EDA. Additional efforts to promote open and reproducible research include refined proxy research enablements and enhanced ML EDA infrastructure through the development and use of new formats, the release of datasets, and the development of Python APIs in OpenROAD. Vidya A. Chhabria, Vikram Gopalakrishnan, Andrew B. Kahng, Sayak Kundu, Zhiang Wang, Bing-Yue Wu, Dooseok Yoon |
ICCAD | 5 |
| 2024 | Physically Aware Synthesis Revisited: Guiding Technology Mapping with Primitive Logic Gate PlacementabstractA typical VLSI design flow is divided into separated front-end logic synthesis and back-end physical design (PD) stages, which often require costly iterations between these stages to achieve design closure. Existing approaches face significant challenges, notably in utilizing feedback from physical metrics to better adapt and refine synthesis operations, and in establishing a unified and comprehensive metric. This paper introduces a new Primitive logic gate placement guided technology MAPping (PigMAP) framework to address these challenges. With approximating technology-independent spatial information, we develop a novel wirelength (WL) driven mapping algorithm to produce PD-friendly netlists. PigMAP is equipped with two schemes: a performance mode that focuses on optimizing the critical path WL to achieve high performance, and a power mode that aims to minimize the total WL, resulting in balanced power and performance outcomes. We evaluate our framework using the EPFL benchmark suites with ASAP7 technology, using the OpenROAD tool for place-and-route. Compared with OpenROAD flow scripts, performance mode reduces delay by 14% while increasing power consumption by only 6%. Meanwhile, power mode achieves a 3% improvement in delay and a 9% reduction in power consumption. Hongyang Pan, Cunqing Lan, Yiting Liu 0002, Zhiang Wang, Li Shang 0001, Xuan Zeng 0001, Fan Yang 0001, Keren Zhu 0001 |
ICCAD | 4 |
| 2024 | K-SpecPart: Supervised Embedding Algorithms and Cut Overlay for Improved Hypergraph PartitioningabstractState-of-the-art hypergraph partitioners follow the multilevel paradigm that constructs multiple levels of progressively coarser hypergraphs that are used to drive cut refinement on each level of the hierarchy. Multilevel partitioners are subject to two limitations: 1) hypergraph coarsening processes rely on local neighborhood structure without fully considering the global structure of the hypergraph and 2) refinement heuristics risk entrapment in local minima. In this article, we describe K-SpecPart, a supervised spectral framework for multiway partitioning that directly tackles these two limitations. K-SpecPart relies on the computation of generalized eigenvectors and supervised dimensionality reduction techniques to generate vertex embeddings. These are computational primitives that are not only fast, but embeddings also capture global structural properties of the hypergraph that are not explicitly considered by existing partitioners. K-SpecPart then converts the vertex embeddings into multiple partitioning solutions. Unlike multilevel partitioners that only consider the best solution, K-SpecPart introduces the idea of “ensembling” multiple solutions via a cut-overlay clustering technique that often enables the use of computationally demanding partitioning methods such as integer linear programming (ILP). Using the output of a standard partitioner as a supervision hint, K-SpecPart effectively combines the strengths of established multilevel partitioning techniques with the benefits of spectral graph theory and other combinatorial algorithms. K-SpecPart significantly extends ideas and algorithms that first appeared in our previous work on the bipartitioner SpecPart (Bustany et al., ICCAD 2022). Our experiments demonstrate the effectiveness of K-SpecPart. For bipartitioning, K-SpecPart produces solutions with up to ~15% cutsize improvement over SpecPart. For multiway partitioning, K-SpecPart produces solutions with up to ~20% cutsize improvement for smaller$K$, and maintains ~2% improvement even when$K$is increased to 128, over leading partitioners hMETIS and KaHyPar. Ismail Bustany, Andrew B. Kahng, Ioannis Koutis, Bodhisatta Pramanik, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Hier-RTLMP: A Hierarchical Automatic Macro Placer for Large-Scale Complex IP BlocksabstractIn a typical RTL to GDSII flow, floorplanning or macro placement is a critical step in achieving decent quality of results (QoR). Moreover, in today’s physical synthesis flows (e.g., Synopsys Fusion Compiler or Cadence Genus iSpatial), a floorplan.def with macro and IO pin placements is typically needed as an input to the front-end physical synthesis. Recently, with the increasing complexity of IP blocks, and in particular with auto-generated RTL for machine learning (ML) accelerators, the number of macros in a single RTL block can easily run into the several hundreds. This makes the task of generating an automatic floorplan (.def) with IO pin and macro placements for front-end physical synthesis even more critical and challenging. The so-called peripheral approach of forcing macros to the periphery of the layout is no longer viable when the ratio of the sum of the macro perimeters to the floorplan perimeter is large, since this increases the required stacking depth of macros. In this article, we develop a novel multilevel physical planning approach that exploits the hierarchy and dataflow inherent in the design RTL, and describe its realization in a new hierarchical macro placer, Hier-RTLMP. Hier-RTLMP borrows from traditional approaches used in manual system-on-chip (SoC) floorplanning to create an automatic macro placement for use with large IP blocks containing very large numbers of macros. Empirical studies demonstrate substantial improvements over the previous RTL-MP macro placement approach (Kahng et al., 2022), and promising post-route improvements relative to a leading commercial place-and-route tool. Andrew B. Kahng, Ravi Varadarajan, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning AcceleratorsabstractParameterizable machine learning (ML) accelerators are the product of recent breakthroughs in ML. To fully enable their design space exploration (DSE), we propose a physical-design-driven, learning-based prediction framework for hardware-accelerated deep neural network (DNN) and non-DNN ML algorithms. It adopts a unified approach that combines power, performance, and area (PPA) analysis with frontend performance simulation, thereby achieving a realistic estimation of both backend PPA and system metrics such as runtime and energy. In addition, our framework includes a fully automated DSE technique, which optimizes backend and system metrics through an automated search of architectural and backend parameters. Experimental studies show that our approach consistently predicts backend PPA and system metrics with an average 7% or less prediction error for the ASIC implementation of two deep learning accelerator platforms, VTA and VeriGOOD-ML, in both a commercial 12 nm process and a research-oriented 45 nm process. Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Sayak Kundu, Rohan Mahapatra, Susmita Dey Manasi, Sachin S. Sapatnekar, Zhiang Wang, Ziqing Zeng |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2023 | An Open-Source Constraints-Driven General Partitioning Multi-Tool for VLSI Physical DesignabstractWith the increasing complexity of IC products, large-scale designs must be efficiently partitioned into multiple blocks, tiles, or devices for concurrent backend place-and-route (P&R) implementation. State-of-the-art partitioners focus on balanced min-cut without considering constraints such as timing or heterogeneity of resource types. They are thus increasingly unsuitable for current physical design requirements. We introduce TritonPart, the first open-source, constraints-driven partitioning tool for VLSI physical design. TritonPart employs efficient algorithms to handle constraints, including multi-dimensional balance, embedding, and timing constraints. Our experimental work affirms its benefits. For standard min-cut partitioning, TritonPart outperforms hMETIS [17], with improvements of up to ~20% on some benchmarks. For embedding-aware partitioning, TritonPart effectively leverages the embeddings generated by SpecPart [4] and improves upon it by ~2%. For timing-aware partitioning, TritonPart significantly reduces the number of cuts on timing-critical paths and prevents timing-noncritical paths from becoming critical (~21X, ~119X reduction relative to hMETIS and KaHyPar [31], respectively). Ismail Bustany, Grigor Gasparyan, Andrew B. Kahng, Ioannis Koutis, Bodhisatta Pramanik, Zhiang Wang |
ICCAD | 6 |
| 2023 | Invited Paper: IEEE CEDA DATC Emerging Foundations in IC Physical Design and MLCAD ResearchabstractRecent activities of the IEEE CEDA DATC strengthen the DATC Robust Design Flow (RDF) and broadly support research on machine learning for CAD/EDA (MLCAD). The RDF-2023 version of the RDF adds standalone and integrated netlist partitioners, a detailed placement optimizer, dynamic power analysis, and enablement of new directions (design-technology co-optimization and 3D layout). Advancement of benchmarking practices and strong baselines has continued - e.g., the MacroPlacement effort introduced in RDF-2022 now has new benchmarks, integration of the AutoDMP macro placer, and baseline solutions generated by Simulated Annealing and human experts. Other DATC efforts have focused on proxies and other elements of MLCAD research enablement. These include real and synthetic benchmarks tailored for IR drop analysis, a calibration methodology for research PDKs, and artificial netlist generation for data augmentation and design space coverage of netlists used in model training. We conclude with directions for future DATC efforts. Jinwook Jung, Andrew B. Kahng, Sayak Kundu, Zhiang Wang, Dooseok Yoon |
ICCAD | 4 |
| 2023 | Assessment of Reinforcement Learning for Macro PlacementabstractWe provide open, transparent implementation and assessment of Google Brain's deep reinforcement learning approach to macro placement (Nature) and its Circuit Training (CT) implementation in GitHub. We implement in open-source key "blackbox" elements of CT, and clarify discrepancies between CT and Nature. New testcases on open enablements are developed and released. We assess CT alongside multiple alternative macro placers, with all evaluation flows and related scripts public in GitHub. Our experiments also encompass academic mixed-size placement benchmarks, as well as ablation and stability studies. We comment on the impact of Nature and CT, as well as directions for future research. Chung-Kuan Cheng, Andrew B. Kahng, Sayak Kundu, Yucheng Wang 0016, Zhiang Wang |
ISPD | 5 |
| 2022 | SpecPart: A Supervised Spectral Framework for Hypergraph Partitioning Solution ImprovementabstractState-of-the-art hypergraph partitioners follow the multilevel paradigm that constructs multiple levels of progressively coarser hypergraphs that are used to drive cut refinements on each level of the hierarchy. Multilevel partitioners are subject to two limitations: (i) Hypergraph coarsening processes rely on local neighborhood structure without fully considering the global structure of the hypergraph. (ii) Refinement heuristics can stagnate on local minima. In this paper, we describe SpecPart, the first supervised spectral framework that directly tackles these two limitations. SpecPart solves a generalized eigenvalue problem that captures the balanced partitioning objective and global hypergraph structure in a low-dimensional vertex embedding while leveraging initial high-quality solutions from multilevel partitioners as hints. SpecPart further constructs a family of trees from the vertex embedding and partitions them with a tree-sweeping algorithm. Then, a novel overlay of multiple tree-based partitioning solutions, followed by lifting to a coarsened hypergraph, where an ILP partitioning instance is solved to alleviate local stagnation. We have validated SpecPart on multiple sets of benchmarks. Experimental results show that for some benchmarks, our SpecPart can substantially improve the cutsize by more than 50% with respect to the best published solutions obtained with leading partitioners hMETIS and KaHyPar. Ismail Bustany, Andrew B. Kahng, Ioannis Koutis, Bodhisatta Pramanik, Zhiang Wang |
ICCAD | 5 |
| 2022 | IEEE CEDA DATC: Expanding Research Foundations for IC Physical Design and ML-Enabled EDAabstractThis paper describes new elements in the RDF-2022 release of the DATC Robust Design Flow, along with other activities of the IEEE CEDA DATC. The RosettaStone initiated with RDF-2021 has been augmented to include 35 benchmarks and four open-source technologies (ASAP7, NanGate45 and SkyWater130HS/HD), plus timing-sensible versions created using path-cutting. The Hier-RTLMP macro placer is now part of DATC RDF, enabling macro placement for large modern designs with hundreds of macros. To establish a clear baseline for macro placers, new open-source benchmark suites on open PDKs, with corresponding flows for fully reproducible results, are provided. METRICS2.1 infrastructure in OpenROAD and OpenROAD-flow-scripts now uses native JSON metrics reporting, which is more robust and general than the previous Python script-based method. Calibrations on open enablements have also seen notable updates in the RDF. Finally, we also describe an approach to establishing a generic, cloud-native large-scale design of experiments for ML-enabled EDA. Our paper closes with future research directions related to DATC's efforts. Jinwook Jung, Andrew B. Kahng, Ravi Varadarajan, Zhiang Wang |
ICCAD | 4 |
| 2022 | RTL-MP: Toward Practical, Human-Quality Chip Planning and Macro PlacementabstractIn a typical RTL-to-GDSII flow, floorplanning plays an essential role in achieving decent quality of results (QoR). A good floorplan typically requires interaction between the frontend designer, who is responsible for the functionality of the RTL, and the backend physical design engineer. The increasing complexity of macro-dominated designs (especially machine learning accelerators with autogenerated RTL) has made the floorplanning task even more challenging and time-consuming. In this paper, we propose RTL-MP, a novel macro placer which utilizes RTL information and tries to "mimic" the interaction between the frontend RTL designer and the backend physical design engineer to produce human-quality floorplans. By exploiting the logical hierarchy and processing logical modules based on connection signatures, RTL-MP can capture the dataflow inherent in the RTL and use the dataflow information to guide macro placement. We also apply autotuning to optimize hyperparameter settings based on input designs. We have built RTL-MP based on OpenROAD infrastructure and applied RTL-MP to a set of industrial designs. RTL-MP outperforms state-of-the-art commercial macro placers and achieves QoR similar to that of handcrafted floorplans. Andrew B. Kahng, Ravi Varadarajan, Zhiang Wang |
ISPD | 3 |
| 2021 | VeriGOOD-ML: An Open-Source Flow for Automated ML Hardware SynthesisabstractThis paper introduces VeriGOOD-ML, an automated methodology for generating Verilog with no human in the loop, starting from a high-level description of a machine learning (ML) algorithm in a standard format such as ONNX. The Verilog RTL is then translated through a back-end design flow to GDSII, driven by a design planning approach that is well tailored to the macro-intensive nature of ML platforms. VeriGOOD-ML uses three approaches to build ML hardware: the TABLA platform uses a dataflow architecture that is well suited to non-DNN ML algorithms; the GeneSys platform, with a systolic array and a SIMD array, is optimized for implementing DNNs; and the Axiline approach synthesizes small ML algorithms by hardcoding the structure of the algorithm into hardware, thus trading off flexibility for performance and power. The overall approach explores the design space of platform configurations and Pareto-optimal-PPA back-end implementations to yield designs that represent different tradeoffs at the algorithmic level between area, power, performance, and execution time. The overall methodology, from architecture to back-end design to hardware implementation, is described in this paper, and the results of VeriGOOD-ML are demonstrated on a set of ML benchmarks. Hadi Esmaeilzadeh, Soroush Ghodrati, Jie Gu 0003, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Rohan Mahapatra, Susmita Dey Manasi, Edwin Mascarenhas, Sachin S. Sapatnekar, Ravi Varadarajan, Zhiang Wang, Hanyang Xu 0002, Brahmendra Reddy Yatham, Ziqing Zeng |
ICCAD | 13 |