EDBT 2026 Demo / reviewers in the wild / expert
Anthony Agnesina
dblp:221/2864
· DBLP profile ↗
30ranked-venue papers
11as first author
23since 2021 · last 2025
0000-0003-0393-0230ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 11 first-author · 23 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DCO-3D: Differentiable Congestion Optimization in 3D ICsabstractState-of-the-art 3D IC flows fail to consider 3D congestion during earlier stages, leading to excessive use of end-of-flow ECO resources for routability correction that severely degrades full-chip Power, Performance, and Area metrics. We present DCO-3D, a Machine Learning-based routability-aware 3D PD flow that performs early post-route congestion prediction using Siamese Networks and resolves the predicted hotspots using a fully differentiable 3D cell spreading with Graph Neural Network. On 6 industrial designs in a commercial 3nm node, DCO-3D improves Pin-3D, the known best Pin-3D flow, by up to 47.2% in overflow, 86.2% in TNS and 5.1% in power at signoff. Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Anthony Agnesina, Rongjian Liang, Yuan-Hsiang Lu, Haoxing Ren, Sung Kyu Lim |
DAC | 4 |
| 2025 | Reinforcement Learning-Driven Window Selection for Enhanced Window-Based Rip-up and Reroute in Chip Detailed RoutingabstractWith increasingly complex design rules and pin density in advanced technology nodes, achieving a violation-free layout has become more challenging, also making rip-up and reroute (RUR) the most runtime-intensive component of detailed routing. We propose a novel reinforcement learning (RL)based approach to enhance the window-based RUR process. Our method features a dynamic window generation strategy that adjusts window size and position based on the distribution of design rule violations (DRV), enabling efficient targeting of congested areas. By leveraging the predictive capabilities of RL, our approach aims to minimize DRVs and achieve high-quality routing results. Experimental results demonstrate that our method outperforms the state-of-theart detailed routers, TritonRoute, achieving a DRV-free solution, averagely improving wirelength by 0.07%, via count by 2.42%, and consuming almost the same average runtime. Yu-Chan Keng, Yu-Chun Pai, Wen-Hao Liu 0001, Haoxing Ren, Danny Liu, Rongjian Liang, Mark Ho, Anthony Agnesina, Yih-Lang Li |
DAC | 8 |
| 2025 | GOALPlace: Begin with the End in MindabstractCo-optimizing placement with congestion is integral to achieving high-quality designs. This paper presents GOALPlace, a learning-based approach to improving placement congestion by controlling cell density. It efficiently learns from an EDA tool's post-route optimized results and uses an empirical Bayes technique to adapt the target to a specific placer's solutions, effectively beginning with the end in mind. Our method enhances correlation with the tool's router and timing-opt engine, while solving placement globally without expensive incremental congestion estimation and mitigation methods. A statistical analysis with hierarchical netlist clustering establishes the importance of density and the potential for an adequate cell density target across placements. Our experiments show that our method, when integrated into an academic GPU-accelerated global placer, consistently produces macro and standard cell placements that match or exceed the quality of commercial tools. Our empirical Bayes methodology also shows a substantial quality improvement over leading academic mixed-size placers, achieving up to 10× fewer design rule check (DRC) violations, a 5% decrease in wirelength, and a 30% and 60% reduction in worst and total negative slack (WNS/TNS). Anthony Agnesina, Rongjian Liang, Geraldo Pradipta, Anand Rajaram, Haoxing Ren |
ISPD | 1 |
| 2025 | Invited: ISPD 2025 Performance-Driven Large Scale Global Routing ContestabstractGlobal routing is a critical aspect of VLSI design, significantly impacting timing, power consumption, and routability. The ISPD2024 contest focused on addressing the scalability challenges of global routing by leveraging GPU and machine learning techniques. Building on this foundation, the ISPD2025 contest introduces several important updates to better reflect real-world routing challenges. These updates include the provision of industry-standard input files for more precise modeling and integration with OpenROAD for accurate performance assessment. Collectively, these updates aim to bring the contest closer to practical routing scenarios, fostering the development of scalable and efficient solutions for large-scale chip designs. Rongjian Liang, Anthony Agnesina, Wen-Hao Liu 0001, Matt Liberty, Hsin-Tzu Chang, Haoxing Ren |
ISPD | 2 |
| 2025 | Cypress: VLSI-Inspired PCB Placement with GPU AccelerationabstractThe scale of printed circuit board (PCB) designs has increased significantly, with modern commercial designs featuring more than 10,000 components. However, the placement process heavily relies on manual efforts that take weeks to complete, highlighting the need for automated PCB placement methods. The challenges of PCB placement arise from its flexible design space and limited routing resources. Existing automated PCB placement tools have achieved limited success in quality and scalability. In contrast, very large-scale integration (VLSI) placement methods have proven to be scalable for designs with millions of cells and delivering high-quality results. Therefore, we propose Cypress, a scalable, GPU-accelerated PCB placement method inspired by VLSI. It incorporates tailored cost functions, constraint handling, and optimized techniques adapted for PCB layouts. In addition, there is an increasing demand for realistic and open-source benchmarks to (1) enable meaningful comparisons between tools and (2) establish performance baselines to track progress in PCB placement technology. To address this gap, we present a PCB benchmark suite synthesized from real commercial designs. We evaluate our method against state-of-the-art commercial and academic PCB placement tools with the benchmark suite. Our approach demonstrates a 1-5.9X higher routability on the proposed benchmarks. For fully routed designs, Cypress achieves 1-19.7X shorter routed track lengths. With GPU acceleration, Cypress delivers up to 492.3X speedup in run time. Finally, we demonstrate scalability to real commercial designs, a capability unmatched by existing tools. Niansong Zhang, Anthony Agnesina, Noor Shbat, Yuval Leader, Zhiru Zhang, Haoxing Ren |
ISPD | 2 |
| 2024 | DGR: Differentiable Global RouterabstractModern VLSI design flows necessitate fast and high-quality global routers. In this paper, we introduce DGR, a differentiable global router capable of concurrent optimization for hundreds of thousands of nets 1. Our innovation lies in the development of a routing Directed Acyclic Graph (DAG) forest to represent the 2D pattern routing space for all nets, enabling coordinated selection of Steiner trees and 2-pin routing paths from a global perspective. For efficient search within the DAG forest, we relax the discrete search space to be continuous and develop a differentiable solver accelerated by deep learning toolkits on GPUs. Experimental results demonstrate that DGR substantially mitigates routing overflow while concurrently reducing total wirelengths from 0.95% to 4.08% and via numbers from 1.28% to 2.54% in congested testcases compared to state-of-the-art academic global routers. Additionally, DGR exhibits favorable scalability in both runtime and memory with respect to the number of nets. Wei Li 0159, Rongjian Liang, Anthony Agnesina, Chia-Tung Ho, Anand Rajaram, Haoxing Ren |
DAC | 3 |
| 2024 | BoolGebra: Attributed Graph-Learning for Boolean Algebraic ManipulationabstractLogic optimization is an essential stage in the design automation flow for digital systems as the performance of the system at logic level can have significant impacts on the final chip area, timing closure, and the power efficiency of the system. Logic optimization is a technology-independent circuit optimization at the logic level conducted on multi-level technology-independent representations such as And-Inverter-Graphs (AIGs) [1] and Majority-Inverter-Graphs (MIGs) [2] of the digital logic. Existing state-of-the-art (SOTA) Directed-Acyclic-Graphs (DAGs) aware Boolean optimization algorithms, such as structural rewriting (rw) [1], resubstitution (rs) [3], and refactoring (rf) [1] in ABC [4], are conducted on the AIG data structure with a graph-level single optimization concept, i.e., all nodes in the graph have one same fixed optimization opportunity, while overlooking other potential optimization opportunities. [5] proposes orchestrated logic optimization, which is a fine-grained node-level logic optimization method incorporating multiple optimization techniques within a single AIG traversal. However, the enlarged search space pose a significant challenge in searching optimal solutions without domain knowledge. Anthony Agnesina, Yanqing Zhang 0002, Haoxing Ren, Cunxi Yu |
DATE | 2 |
| 2024 | 2024 ICCAD CAD Contest Problem C: Scalable Logic Gate Sizing Using ML Techniques and GPU AccelerationabstractLogic gate sizing plays a vital role in timing optimization, especially as Moore's Law slows, shifting greater responsibility to EDA tools to enhance power, performance, and area (PPA), as these gains are no longer achieved solely through scaling and process advancements. There is an increasing need to push the limits of logic gate sizing to extract every possible improvement in PPA. With recent breakthroughs in machine learning (ML) and the computational power of GPUs, there is significant potential to elevate logic gate sizing algorithms to new heights. This contest aims to advance logic gate sizing and push the boundaries of PPA improvement through innovative EDA tools that leverage machine learning and GPU acceleration. As part of the contest, an infrastructure has been developed to enable ML and GPU-accelerated logic gate sizing algorithms, including the release of benchmarks in both standard EDA and ML-friendly formats, along with examples of incorporating "ML inside" EDA tools through Python APIs. The contest leverages the open-source EDA tool OpenROAD and ML-friendly data representation format, CircuitOps, to lower barriers to entry by providing accessible formats and tools, allowing participants to build on existing software without redundancy. With over 25 teams actively participating, the contest highlights growing interest and potential to push the boundaries of timing optimization. Bing-Yue Wu, Rongjian Liang, Geraldo Pradipta, Anthony Agnesina, Haoxing Ren, Vidya A. Chhabria |
ICCAD | 4 |
| 2024 | GPU/ML-Enhanced Large Scale Global Routing ContestabstractModern VLSI design flows demand scalable global routing techniques applicable across diverse design stages. In response, the ISPD 2024 contest pioneers the first GPU/ML-enhanced global routing competition, selecting advancements in GPU-accelerated computing platforms and machine learning techniques to address scalability challenges. Large-scale benchmarks, containing up to 50 million cells, offer test cases to assess global routers' runtime and memory scalability. The contest provides simplified input/output formats and performance metrics, framing global routing challenges as mathematical optimization problems and encouraging diverse participation. Two sets of evaluation metrics are introduced: the primary one concentrates on global routing applications to guide post-placement optimization and detailed routing, focusing on congestion resolution and runtime scalability. Special honor is given based on the second set of metrics, placing additional emphasis on runtime efficiency and aiming at guiding early-stage planning. Rongjian Liang, Anthony Agnesina, Wen-Hao Liu 0001, Haoxing Ren |
ISPD | 2 |
| 2024 | MedPart: A Multi-Level Evolutionary Differentiable Hypergraph PartitionerabstractState-of-the-art hypergraph partitioners, such as hMETIS, usually adopt a multi-level paradigm for efficiency and scalability. However, they are prone to getting trapped in local minima due to their reliance on refinement heuristics and overlooking global structural information during coarsening. SpecPart, the most advanced academic hypergraph partitioning refinement method, improves partitioning by leveraging spectral information. Still, its success depends heavily on the quality of initial input solutions. This work introduces MedPart, a multi-level evolutionary differentiable hypergraph partitioner. MedPart follows the multi-level paradigm but addresses its limitations by using fast spectral coarsening and introducing a novel evolutionary differentiable algorithm to optimize each coarsening level. Moreover, by analogy between hypergraph partitioning and deep graph learning, our evolutionary differentiable algorithm can be accelerated with deep graph learning toolkits on GPUs. Experiments on public benchmarks consistently show MedPart outperforming hMETIS and achieving up to a 30% improvement in cut size for some benchmarks compared to the best-published solutions, including those from SpecPart---moreover, MedPart's runtime scales linearly with the number of hyperedges. Rongjian Liang, Anthony Agnesina, Haoxing Ren |
ISPD | 2 |
| 2024 | Challenges for Automating PCB LayoutabstractPrinted circuit board (PCB) design is typically semi-automated or fully manual. However, in recent years, the scale of PCB designs has rapidly enlarged, such that the engineering effort of manual design has increased dramatically. Therefore, the criticality of automation emerges. PCB houses are looking for productivity improvement that is contributed by automation. In this talk, the speaker will give a short tutorial about how a PCB design is done today and then indicate the challenges and opportunities for PCB design automation. Wen-Hao Liu 0001, Anthony Agnesina, Haoxing Ren |
ISPD | 2 |
| 2024 | Hier-3D: A Methodology for Physical Hierarchy Exploration of 3-D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. Complemented with an automated floorplanning solution, the flow allows for system-level physical and architectural exploration of 3D designs. As a result, we significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 -2.2× compared with 2D, where all metrics are improved simultaneously, including up to 20% power savings. Nesara Eranna Bethur, Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | On Legalization of Die Bonding Bumps and Pads for 3-D ICsabstractAs state-of-the-art 3D IC Place-and-Route flows were designed with older technology nodes and aggressive bonding pitch assumptions, they introduce an unacceptable number of 3D via overlap violations during routing in real-world scenarios. Specifically, when dealing with higher via pitch to wire size ratios using more advanced technology nodes than they were designed for, these flows struggle to comply with width and spacing rules. In this paper, we propose a novel 3D via legalization stage and a subsequent refinement stage during routing to address this issue. Two independent via legalization methods are introduced: a force-based algorithm and a bipartite-matching algorithm with Bayesian optimization. Our two legalization methods, along with the refinement stage, are compatible with various process nodes, bonding technologies, and partitioning styles. By implementing the modified 3D routing with the proposed legalizers, we successfully eliminate all 3D via overlap violations while minimizing the impact on performance, power, or area. Sai Pentapati, Anthony Agnesina, Moritz Brunion, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Invited Paper: CircuitOps: An ML Infrastructure Enabling Generative AI for VLSI Circuit OptimizationabstractAn innovative ML infrastructure named CircuitOps is developed to streamline dataset generation and model inference for various generative AI (GAI)-based circuit optimization tasks. Addressing the challenges of the absence of a shared Intermediate Representation (IR), steep EDA learning curves, and AI-unfriendly data structures, we propose solutions that empower efficient data handling. Our contributions encompass the following: (1) labeled property graphs (LPGs) as IR for flexible netlist representation and efficient parallel processing; (2) tools-agnostic IR generation from standard EDA files; (3) customizable dataset generation facilitated through AI-friendly LPGs; (4) gRPC-based inference deployment. Compared with using Tcl interfaces of EDA design tools, CircuitOps achieves a significant 99× dataset generation speedup and 75K nets per second transfer throughput, validating its effectiveness in optimizing GAI tasks. Rongjian Liang, Anthony Agnesina, Geraldo Pradipta, Vidya A. Chhabria, Haoxing Ren |
ICCAD | 2 |
| 2023 | AutoDMP: Automated DREAMPlace-based Macro PlacementabstractMacro placement is a critical very large-scale integration (VLSI) physical design problem that significantly impacts the design power-performance-area (PPA) metrics. This paper proposes AutoDMP, a methodology that leverages DREAMPlace, a GPU-accelerated placer, to place macros and standard cells concurrently in conjunction with automated parameter tuning using a multi-objective hyperparameter optimization technique. As a result, we can generate high-quality predictable solutions, improving the macro placement quality of academic benchmarks compared to baseline results generated from academic and commercial tools. AutoDMP is also computationally efficient, optimizing a design with 2.7 million cells and 320 macros in 3 hours on a single NVIDIA DGX Station A100. This work demonstrates the promise and potential of combining GPU-accelerated algorithms and ML techniques for VLSI design automation. Anthony Agnesina, Puranjay Rajvanshi, Geraldo Pradipta, Austin Jiao, Ben Keller, Brucek Khailany, Haoxing Ren |
ISPD | 1 |
| 2023 | On Legalization of Die Bonding Bumps and Pads for 3D ICsabstractState-of-the-art 3D IC Place-and-Route flows were designed with older technology nodes and aggressive bonding pitch assumptions. As a result, these flows fail to honor the width and spacing rules for the 3D vias with realistic pitch values. We propose a critical new 3D via legalization stage during routing to reduce such violations. A force-based solver and bipartite-matching algorithm with Bayesian optimization are presented as viable legalizers and are compatible with various process nodes, bonding technologies, and partitioning types. With the modified 3D routing, we reduce the 3D via violations by more than 10× with zero impact on performance, power, or area. Sai Pentapati, Anthony Agnesina, Moritz Brunion, Sung Kyu Lim |
ISPD | 2 |
| 2023 | Parameter Optimization of VLSI Placement Through Deep Reinforcement LearningabstractCritical to achieving power–performance–area goals, a human engineer typically spends a considerable amount of time tuning the multiple settings of a commercial placer. This article proposes a deep reinforcement learning (RL) framework to optimize the placement parameters of a commercial electronic design automation (EDA) tool. We build an autonomous agent that learns to tune parameters without human intervention and domain knowledge, trained solely by RL from self-search. To generalize to unseen netlists, we use a mixture of handcrafted features from graph topology theory and graph embeddings generated using unsupervised graph neural networks. Our RL algorithms are chosen to overcome the sparsity of data and latency of placement runs. As a result, our trained RL agent achieves up to 11% and 2.5% wire length improvements on unseen netlists compared with a human engineer and a state-of-the-art tool auto-tuner in just one placement iteration ($20\times $and$50\times $fewer iterations). In addition, the success of the RL agent is measured using a statistical test with theoretical guarantees and an optimized sample size. Anthony Agnesina, Kyungwook Chang, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | A PPA Study of Reinforced Placement Parameter Autotuning: Pseudo-3D vs. True-3D PlacersabstractD Place and Route (P&R) flows either involve true-3D placement algorithms or use commercial 2D tools to transform a 2D design into a 3D design. Irrespective of the nature of the placers, several placement parameters in these tools affect the quality of the final 3D designs. Different parameter settings work well with different circuits, and it is impossible to manually tune them for a particular circuit. Automated approaches involving reinforcement learning have been shown to adapt and learn the parameter settings and create trained models. However, their effectiveness depends on the input dataset quality. Using a set of 10 netlists and 10–21 handpicked placement parameters in P&R flows involving pseudo-3D or true-3D placement, the dataset quality is analyzed. The datasets are the design metrics obtained through different P&R stages, such as placement optimization, clock tree synthesis, or 3D partitioning and global routing. The training runtime and the quality of the final design metrics are compared. On a pseudo-3D flow, the training takes around 126–290 hours, whereas, on a true-3D placer-based flow, it takes around 305–410 hours. It is observed that the datasets obtained from different stages lead to drastically different final design results. With the RL-based training processes, the quality of results in 3D designs improves by up to 23.7% compared to their corresponding untrained P&R flows. Gauthaman Murali, Anthony Agnesina, Sung Kyu Lim |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2022 | MemPool-3D: Boosting Performance and Efficiency of Shared-L1 Memory Many-Core Clusters with 3D IntegrationabstractThree-dimensional integrated circuits promise power, performance, and footprint gains compared to their 2D counter-parts, thanks to drastic reductions in the interconnects' length through their smaller form factor. We can leverage the potential of 3D integration by enhancing MemPool, an open-source many-core design with 256 cores and a shared pool of L1 scratchpad memory connected with a low-latency interconnect. MemPool's baseline 2D design is severely limited by routing congestion and wire propagation delay, making the design ideal for 3D integration. In architectural terms, we increase MemPool's scratchpad memory capacity beyond the sweet spot for 2D designs, improving performance in a common digital signal processing kernel. We propose a 3D MemPool design that leverages a smart partitioning of the memory resources across two layers to balance the size and utilization of the stacked dies. In this paper, we explore the architectural and the technology parameter spaces by analyzing the power, performance, area, and energy efficiency of MemPool instances in 2D and 3D with 1 MiB, 2 MiB, 4 MiB, and 8 MiB of scratchpad memory in a commercial 28 nm technology node. We observe a performance gain of 9.1% when running a matrix multiplication on MemPool-3D with 4 MiB of scratchpad memory compared to the MemPool 2D counterpart. In terms of energy efficiency, we can implement the MemPool-3D instance with 4 MiB of L1 memory on an energy budget 15 % smaller than its 2D counterpart, and 3.7 % smaller than the MemPool-2D instance with a fourth of the L1 scratchpad memory capacity. Matheus A. Cavalcante, Anthony Agnesina, Samuel Riedel, Moritz Brunion, Alberto García Ortiz, Dragomir Milojevic, Francky Catthoor, Sung Kyu Lim, Luca Benini |
DATE | 2 |
| 2022 | Hier-3D: A Hierarchical Physical Design Methodology for Face-to-Face-Bonded 3D ICsabstractHierarchical very-large-scale integration (VLSI) flows are an understudied yet critical approach to achieving design closure at giga-scale complexity and gigahertz frequency targets. This paper proposes a novel hierarchical physical design flow enabling the building of high-density and commercial-quality two-tier face-to-face-bonded hierarchical 3D ICs. We significantly reduce the associated manufacturing cost compared to existing 3D implementation flows and, for the first time, achieve cost competitiveness against the 2D reference in large modern designs. Experimental results on complex industrial and open manycore processors demonstrate in two advanced nodes that the proposed flow provides major power, performance, and area/cost (PPAC) improvements of 1.2 to 2.2 × compared with 2D, where all metrics are improved simultaneously, including up to power savings. Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Matheus A. Cavalcante, Samuel Riedel, Luca Benini, Sung Kyu Lim |
ISLPED | 1 |
| 2022 | ART-3D: Analytical 3D Placement with Reinforced Parameter Tuning for Monolithic 3D ICsabstractIn this paper, we show that true 3D placement approaches, enhanced with reinforcement learning, can offer further PPA improvements over pseudo-3D approaches. To accomplish this goal, we integrate an academic true 3D placement engine into a commercial-grade 3D physical design flow, creating ART-3D flow (Analytical 3D Placement with Reinforced Parameter Tuning-based 3D flow). We use a reinforcement learning (RL) framework to find optimized placement parameter settings of the true 3D placement engine for a given netlist and perform high-quality 3D placement. We then use an efficient 3D optimization and routing engine based on a commercial place and route (P&R) tool to maintain or improve the benefits reaped from true 3D placement till design signoff. We evaluate our 3D flow by designing several gate-only and processor benchmarks on a commercial 28nm technology node. Our proposed 3D flow involving true 3D placement offers the best PPA results compared to existing 3D P&R flows and reduces power consumption by up to 31%, improves effective frequency by up to 25%, and therefore reduces power-delay product by up to 43% compared with commercial 2D IC design flow. These improvements predominantly come from RL-based parameter tuning, as it improves the performance of the 3D placer by up to 12%. Gauthaman Murali, Sandra Maria Shaji, Anthony Agnesina, Guojie Luo, Sung Kyu Lim |
ISPD | 3 |
| 2022 | A Clock Tree Prediction and Optimization Framework Using Generative Adversarial LearningabstractModern physical design flows highly depend on design space exploration to find the commercial tools’ clock tree synthesis (CTS) parameters that lead to optimized clock trees. However, such exploration is often time-consuming and computationally inefficient. In this article, we overcome this drawback by proposing a novel framework named GAN-CTS, which utilizes conditional generative adversarial network (GAN) to predict and optimize CTS outcomes. Our framework is built upon three sequential learning stages. First, to precisely characterize distinct designs, we leverage transfer learning to extract netlist features directly from placement images. Second, we perform regression learning using various methods to predict the target CTS outcomes and demonstrate that the proposed multitask learning approach achieves better accuracy than the meta-modeling method adopted by previous works. To fully benefit from the predictions made by our framework, we further quantitatively interpret the importance of each CTS input parameter subject to various design objectives through attribution-based learning. Finally, generative adversarial learning is leveraged to optimize the target clock metrics with the guidance provided by the pretrained regression model. To substantiate the generality of our framework, we perform validations on four unseen netlists that are not utilized in the training process. The experimental results conducted on real-world designs demonstrate that our framework: 1) achieves an average prediction error of 3%; 2) improves the commercial tool’s auto-generated clock tree by 20.7% in clock power, 21.5% in clock wirelength, 36.1% in the worst skew; and 3) reaches an F1-score of 0.93 in the classification task of determining successful and failed CTS runs. Yi-Chen Lu, Jeehyun Lee 0002, Anthony Agnesina, Kambiz Samadi, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Power, Performance, Area and Cost Analysis of Memory-on-Logic Face-to-Face Bonded 3D Processor DesignsabstractIn this paper, we present a power, performance, area and cost (PPAC) analysis for large-scale 3D processor designs based on wafer-to-wafer bonding. From the evaluation of our cost model, we investigate a typically disregarded opportunity in 3D that is area savings due to buffer savings and better routability, offering unexpected cost savings. We explore the viability of this factor with the feedback of a state-of-the-art 3D memory-on-logic implementation flow. We show how this affects the PPAC of full-chip GDS implementations of a large-scale manycore processor design. Experiments show that our memory-on-logic 3D implementation offers 7% silicon area savings, resulting in 53.5% footprint reduction. We also obtain a 40% power-performance-cost improvement compared with 2D counterparts Anthony Agnesina, Moritz Brunion, Alberto García Ortiz, Dragomir Milojevic, Francky Catthoor, Manu Perumkunnil Komalan, Sung Kyu Lim |
ISLPED | 1 |
| 2020 | VLSI Placement Parameter Optimization using Deep Reinforcement LearningabstractThe quality of placement is essential in the physical design flow. To achieve PPA goals, a human engineer typically spends a considerable amount of time tuning the multiple settings of a commercial placer (e.g. maximum density, congestion effort, etc.). This paper proposes a deep reinforcement learning (RL) framework to optimize the placement parameters of a commercial EDA tool. We build an autonomous agent that learns to tune parameters optimally without human intervention and domain knowledge, trained solely by RL from self-search. To generalize to unseen netlists, we use a mixture of handcrafted features from graph topology theory along with graph embeddings generated using unsupervised Graph Neural Networks. Our RL algorithms are chosen to overcome the sparsity of data and latency of placement runs. Our trained RL agent achieves up to 11% and 2.5% wirelength improvements on unseen netlists compared with a human engineer and a state-of-the-art tool auto-tuner, in just one placement iteration (20× and 50× less iterations). Anthony Agnesina, Kyungwook Chang, Sung Kyu Lim |
ICCAD | 1 |
| 2020 | A Fault-Tolerant and High-Speed Memory Controller Targeting 3D Flash Memory Cubes for Space ApplicationsabstractIn this work, we develop a new 3D flash memory cube architecture that integrates multiple flash dies and their logic controller in a unique and optimized fashion for space applications. In our Loaf-of-Bread (LOB) configuration, flash dies are standing up and bonded laterally instead of the conventional pancake-style vertical die stacking. Our LOB allows the flash dies to be bonded without the use of through-silicon-vias and micro-bumps. Instead, we insert a redistribution layer in between two adjacent dies to bring the IO signals to the bottom, where a logic controller die collects the IOs from all flash dies and coordinates the communication with an off-cube host processor. Thus, our LOB configuration allows the users to form flash memory cubes using off-the-shelf 2D flash dies and complete the integration at a packaging house instead of a fab. A key element in our LOB flash cube is the logic controller architecture that supports fault-tolerant and energy-efficient operation of the cube. We develop the controller architecture and validate the functionality using C++ emulation and FPGA prototyping. Compared with a state-of-the-art space-grade flash memory module, our system shows a 20X bandwidth improvement in a smaller form factor along with a 25X better ratio of density per volume. Anthony Agnesina, Da Eun Shim, James Yamaguchi, Christian Krutzik, John Carson, Dan Nakamura, Sung Kyu Lim |
ICCD | 1 |
| 2020 | Improving FPGA-Based Logic Emulation Systems through Machine LearningabstractWe present a machine learning (ML) framework to improve the use of computing resources in the FPGA compilation step of a commercial FPGA-based logic emulation flow. Our ML models enable highly accurate predictability of the final place and route design qualities, runtime, and optimal mapping parameters. We identify key compilation features that may require aggressive compilation efforts using our ML models. Experiments based on our large-scale database from an industry’s emulation system show that our ML models help reduce the total number of jobs required for a given netlist by 33%. Moreover, our job scheduling algorithm based on our ML model reduces the overall time to completion of concurrent compilation runs by 24%. In addition, we propose a new method to compute “recommendations” from our ML model to perform re-partitioning of difficult partitions. Tested on a large-scale industry system on chip design, our recommendation flow provides additional 15% compile time savings for the entire system on chip. To exploit our ML model inside the time-critical multi-FPGA partitioning step, we implement it in an optimized multi-threaded representation. Anthony Agnesina, Sung Kyu Lim, Etienne Lepercq, Jose Escobedo Del Cid |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2020 | A COTS-Based Novel 3-D DRAM Memory Cube Architecture for Space ApplicationsabstractThe first mainstream products in three-dimensional integrated circuit (3-D IC) design are memory devices where multiple memory tiers are horizontally integrated to offer manifold improvements when compared with their 2-D counterparts. Unfortunately, none of these existing 3-D memory cubes are ready for harsh space environments. This article introduces a new memory cube architecture for space, based on the vertical integration of multiple commercial-off-the-shelf, 3-D stacked, dynamic random-access memory (DRAM) memory devices with a custom radiation-hardened-by-design controller. Our solution offers high memory capacity, increased bandwidth, fault tolerance, and improved size-weight-and-power characteristics needed for space missions. Validation and functional evaluation of the application-specific integrated circuit (ASIC) controller will be conducted prior to tape-out on a custom FPGA-based emulator platform integrating the 3-D stack. The selected test methodology ensures high-quality register transfer level (RTL) as well as allows to subject the cube structure to radiation testing. The proposed design concept allows for flexibility in the choice of the DRAM die in the case of technology road-map changes or unsatisfactory radiation results. Anthony Agnesina, James Yamaguchi, Christian Krutzik, John Carson, Jean Yang-Scharlotta, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Reducing Compilation Effort in Commercial FPGA Emulation Systems Using Machine LearningabstractThis paper presents a machine learning (ML) framework to improve the use of computing resources in the FPGA compilation step of a commercial FPGA-based logic emulation flow. Our ML models enable highly accurate predictability of the final P&R design qualities, runtime, and optimal mapping parameters. We identify key compilation features that may require aggressive compilation efforts using our ML models. Experiments based on our large-scale database from an industry's emulation system show that our ML models help reduce the total number of jobs required for a given netlist by 33%. Moreover, our job scheduling algorithm based on our ML model reduces the overall time to completion of concurrent compilation runs by 24%. In addition, we propose a new method to compute “recommendations” from our ML model, in order to perform repartitioning of difficult partitions. Tested on a large-scale industry SoC design, our recommendation flow provides additional 15% compile time savings for the entire SoC. Anthony Agnesina, Etienne Lepercq, Jose Escobedo Del Cid, Sung Kyu Lim |
ICCAD | 1 |
| 2019 | GAN-CTS: A Generative Adversarial Framework for Clock Tree Prediction and OptimizationabstractIn this paper, we propose a complete framework named GAN-CTS which utilizes conditional generative adversarial network (GAN) and reinforcement learning to predict and optimize clock tree synthesis (CTS) outcomes. To precisely characterize different netlists, we leverage transfer learning to extract design features directly from placement images. Based on the proposed framework, we further quantitatively interpret the importance of each CTS input parameter subject to various design objectives. Finally, to prove the generality of our framework, we conduct experiments on the unseen netlists which are not utilized in the training process. Experimental results performed on industrial designs demonstrate that our framework (1) achieves an average prediction error of 3%, (2) improves the commercial tool's auto-generated clock tree by 51.5% in clock power, 18.5% in clock wirelength, 5.3% in the maximum skew, and (3) reaches an F1-score of 0.952 in the classification task of determining successful and failed CTS processes. Yi-Chen Lu, Jeehyun Lee 0002, Anthony Agnesina, Kambiz Samadi, Sung Kyu Lim |
ICCAD | 3 |
| 2018 | A novel 3D DRAM memory cube architecture for space applicationsabstractThe first mainstream products in 3D IC design are memory devices where multiple memory tiers are horizontally integrated to offer manifold improvements compared with their 2D counterparts. Unfortunately, none of these existing 3D memory cubes are ready for harsh space environments. This paper presents a new memory cube architecture for space, based on vertical integration of Commercial-Off-The-Shelf (COTS), 3D stacked, DRAM memory devices with a custom Radiation-Hardened-By-Design (RHBD) controller offering high memory capacity, robust reliability and low latency. Validation and evaluation of the ASIC controller will be conducted prior to tape-out on a custom FPGA-based emulator platform integrating the 3D-stack. Anthony Agnesina, Amanvir Sidana, James Yamaguchi, Christian Krutzik, John Carson, Jean Yang-Scharlotta, Sung Kyu Lim |
DAC | 1 |