Keren Zhu 0001

dblp:139/1776 · DBLP profile ↗
← Back
56ranked-venue papers
6as first author
48since 2021 · last 2026
0000-0003-2698-141XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 53 · 6 first-author · 46 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 PigMap3: A Physically Aware Incremental Mapping Framework with On-the-fly Post-Layout Critical Path Tracking
Hongyang Pan, Cunqing Lan, Zhiang Wang, Xuan Zeng 0001, Fan Yang 0001, Keren Zhu 0001
ASP-DAC6
2026 Submodular Maximization-inspired Adaptive Routing Bend Space Planning
abstract
Routing can greatly impact tape-out chip performance by determining the physical layout of metal wire segments. As designs grow in complexity and size, modern routing frameworks struggle to manage limited routing resources among numerous nets efficiently. In this paper, we introduce a novel adaptive routing bend space planning framework, ARSP, that adaptively adjusts the routing bend space for each net based on the availability of routing resources throughout the routing flow. ARSP is built on a well-defined submodular maximization problem and uses an efficient approximation algorithm to ensure sub-optimal performance. Integrating ARSP with state-of-the-art routing flows shows an average improvement of 6.34% and 5.11% in reducing shorts and spacing violations, respectively. Additionally, our adaptive planning framework outperforms all static routing space planning strategies in both effectiveness and efficiency, showing the necessity of adaptive planning.
Siting Liu 0002, Peng Xu 0052, Peiyu Liao, Keren Zhu 0001, Yibo Lin, Bei Yu 0001
DATE4
2026 A Physically-aware Framework for Joint MBFF Synthesis with OPTICS-based Debanking
abstract
MBFF banking is a standard technique for clock power reduction in modern IC design, yet two systemic flaws in prior works limit its practical gains: geometric abstractions that ignore placement congestion, and open-loop workflows where late legalization failures nullify power savings. We propose a self-correcting framework that co-optimizes banking, placement, and debanking via three innovations: (1) Mahalanobis-distance clustering for placement-feasible MBFF formation; (2) a legalization-driven feedback loop with cost-aware debanking to recover unplaceable MBFFs; and (3) an OPTICS-based debanking that splits problematic MBFFs at highest-cost boundaries. Evaluated on ICCAD 2024 CAD Contest Problem B and large-scale benchmarks, our framework outperforms the 1st, 2nd, and 3rd place winners by 2.3%, 9.3%, and 4.0% in final weighted score, respectively.
Benchao Zhu, Yang Liu 0376, Jianli Chen, Keren Zhu 0001
ACM Great Lakes Symposium on VLSI5
2026 HOLMES: Hierarchical Optimization with poLygonal ModEling for Large-Scale AMS Placement
Yujie Yan, Jiahua Liu, Zecheng Xu, Linxi Qiu, Yumao Wu, Zhiang Wang, Changhao Yan, Zhaori Bi, Keren Zhu 0001
ISCAS10
2026 PhySeqForm: A Data-Driven, Physical Synthesis Sequence Former
Cunqing Lan, Zijian Jiang, Hongyang Pan, Zhiang Wang, Keren Zhu 0001
ISCAS5
2026 RC-Scaled Timing-Driven Routing: Bridging Targeted Timing Optimization and Massively
Parallel Global Routing, Zecheng Xu, Boxiang Song, Zhiang Wang, Fan Yang 0001, Keren Zhu 0001, Xuan Zeng 0001
ISCAS7
2026 Layout-Aware Standard Cell Synthesis via Reparameterization Multi-Task Bayesian Optimization
Zhouyang Wu, Ruiyu Lyu, Keren Zhu 0001, Zhiang Wang, Zhaori Bi, Changhao Yan, Xuan Zeng 0001
ISCAS3
2026 PigMap2: A Physical Information-Guided Technology Mapping Framework
Cunqing Lan, Hongyang Pan, Zhiang Wang, Xuan Zeng 0001, Fan Yang 0001, Keren Zhu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 DAMIL-DCIM+: Automated Dataflow-Aware Layout Synthesis for Digital CIM With Self-Assembled Bitcell Units and MILP-Based Optimization
abstract
Digital computing-in-memory (DCIM) systems integrate complex digital logic with parasitic-sensitive bitcell arrays, presenting unique physical design challenges. Conventional design strategies often fall short in these systems due to irregular dataflow patterns and excessive interconnect lengths, which degrade performance and increase parasitic effects. As a result, current DCIM implementations frequently rely on manual layout, which is both time-consuming and a major bottleneck in the design cycle. While existing DCIM layout synthesis frameworks attempt to automate this process using template-based placement methods inspired by manual design, their rigid constraints can lead to inefficient area utilization and increased core sizes. To address these limitations, we propose DAMIL-DCIM+, a novel placement framework that combines the structural clarity of template-based methods with the flexibility of optimization-based techniques. Specifically, DAMIL-DCIM+ employs a global dataflow-aware floorplan to guide placement and leverages MILP-based detailed placement to optimize wirelength and preserve dataflow regularity. Inspired by self-assembling design principles, this approach enables scalable and structured integration of parasitic-sensitive components. The hybrid methodology of DAMIL-DCIM+ reduces total wirelength, lowers parasitic effects, and enhances performance while maintaining design regularity. Experimental results on a 28nm DCIM circuit demonstrate that DAMIL-DCIM+ improves operating frequency by 25.2% and reduces power consumption by 19.6% compared to Cadence Innovus, without increasing core area.
Xinglong Yan, Zecheng Xu, Keren Zhu 0001, Shuo Li 0008, Fan Yang 0001, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 PAPlace: Performance-Driven Differentiable Analog Placement
abstract
Analog circuit placement is crucial for optimal performance, but achieving a decent layout demands expertise and time. Recent advances in machine learning techniques have shown promising results in modeling analog layout performance. PAPlace further extends these methods and integrates them into the core analog placement engine, allowing direct optimization of the post-layout performance effectively. Our approach proposes a differentiable prediction model that combines layout and wiring information into a non-linear analog placement engine. We then incorporate the differentiable performance model into a gradient-descent-based global placement engine. A multi-objective optimization method is further proposed to find the common gradient descent direction for different metrics. The experimental results on benchmarks under the TSMC 40nm technology node demonstrate the superiority of the proposed framework compared with the cutting-edge works, with up to 2163.00μ V , 73.95dB, 62.25MHz, 57.84dB improvement in Offset Voltage, CMRR, BandWidth, DC Gain metrics.
Peng Xu 0052, Yuan Pu 0001, Keren Zhu 0001, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.3
2025 MARIO: A Superadditive Multi-Algorithm Interworking Optimization Framework for Analog Circuit Sizing
abstract
Numeric optimization methods are widely utilized to tackle complex analog circuit sizing problems, where the challenges include expensive simulations, non-linearity, and high parameter dimensionality. However, the diverse characteristics exhibited by different circuits result in varied optimization landscapes, making it difficult to identify a single algorithm that consistently outperforms others across all problems. In this paper, we introduce a multi-algorithm interworking optimization framework, which achieves optimization superadditivity based on a pool of member algorithms and a powerful algorithm-interworking protocol. We propose a computing resource reallocation method, which employs multitask Gaussian process regression and portfolio optimization techniques, leading to flexible and prudent online adaption of member algorithms. To efficiently utilize the computing resources for local exploitation, an evaluation data broadcast strategy enables cooperativeness across member algorithms. Besides, algorithms with different modeling overheads are integrated time-adaptively via an asynchronous parallelization mechanism. Comparative experiments against state-of-the-art algorithmcombining tools and optimization algorithms demonstrate the superiority of the proposed optimization framework.
Wangzhen Li, Ruiyu Lyu, Changhao Yan, Keren Zhu 0001, Zhaori Bi, Dian Zhou, Xuan Zeng 0001
DAC5
2025 Decoupling Analog Circuit Representation from Technology for Behavior-Centric Optimization
abstract
Analog IC design is mainly manual and implemented at the device level. A major reason is circuit behavior-extraction. Unlike its digital counterpart, analog IC design is strongly coupled with technology nodes and is difficult to represent by an abstract behavioral model. The lack of accurate and efficient analog modeling has become a bottleneck in analog design automation. This paper proposes a behavior-centric optimization framework for analog circuits that represents circuit behavior using transistor electrical properties instead of sizes, improving model generalization and reducing optimization complexity. To characterize the process, we propose a method for mapping transistor electrical properties to sizes. Moreover, we developed a radial basis functions-based Kolmogorov-Arnold network (RBF-KAN) to accurately approximate circuit nonlinear behavior with limited simulations. Compared to blackbox modeling, our approach enables constructing surrogate models via KAN under a set specification with just a few hundred simulations. Experiments on the testing suite showed our framework achieved a $1.76 \times$ to $2.64 \times$ improvement in large signal figure of merit (FOM) and $1.73 \times$ to $2.48 \times$ in small signal FOM over state-of-the-art methods, while also enabling $3.5 \times$ to $6.2 \times$ acceleration in design porting.
Jintao Li 0002, Haochang Zhi, Jiang Xiao 0002, Keren Zhu 0001, Yun Li 0002
DAC4
2025 ELMap: Area-Driven LUT Mapping with $k$-LUT Network Exact Synthesis
abstract
Mapping to$k$-input lookup tables ($k$-LUTs) is a critical process in field-programmable gate array (FPGA) synthesis. However, the structure of the subject graph can introduce structural bias, which refers to the dependency of mapping results on the inherent graph structure, often leading to suboptimal results. To address this, we present ELMap, an area-driven LUT mapping framework. It incorporates structural choice during the collapsing phase. This enables dynamic decomposition, maximizing local-to-global optimization transfer. To ensure seamless integration between the optimization and mapping processes, ELMap leverages exact$k$-LUT synthesis to generate area-optimal sub-LUT networks. Experiments on the EPFL benchmark suite demonstrate that ELMap significantly outperforms state-of-the-art methods. Specifically, in 6-LUT mapping, ELMap reduces the average LUT area by 8.5% and improves the area-depth-product (ADP) by 5.8%. In 4-LUT remapping, it reduces the average LUT area by 17.6% and improves the ADP by 2.4%.
Hongyang Pan, Keren Zhu 0001, Fan Yang 0001, Zhufei Chu, Xuan Zeng 0001
DATE2
2025 DAMIL-DCIM: A Digital CIM Layout Synthesis Framework with Dataflow-Aware Floorplan and MILP-Based Detailed Placement
abstract
Digital computing-in-memory (DCIM) systems integrate complex digital logic with parasitic-sensitive bitcell arrays. Conventional physical design strategies degrade DCIM performance due to a lack of dataflow regularity and excessive wirelength. As a result, current DCIM design often relies on manual layout, which is time-consuming and a bottleneck in the design cycle. Existing layout synthesis frameworks for DCIM often mimic the manual approach and employ a template-based method for DCIM placement. However, overly constrained templates lead to an excessive core area, resulting in high costs in practice. In this work, we introduce DAMIL-DCIM, a novel placement framework that bridges template-based techniques with optimization-based placement methods. DAMIL-DCIM utilizes a global dataflow-aware floorplan inspired by template methods and further optimizes the layout using MILP(Mixed Integer Linear Programming)-based detailed placement. The combination of global floorplanning and placement optimization reduces total wire length while maintaining dataflow regularity, resulting in lower parasitic and enhanced performance. Experimental results show, on a practical 28nm DCIM circuit, our approach improves frequency by 25.2% and reduces power consumption by 19.6% compared to Cadence Innovus, while maintaining the same core area.
Fan Yang 0001, Keren Zhu 0001, Xuan Zeng 0001
DATE4
2025 LCTMwalk: GPU-Accelerated Transient Thermal Simulation for Liquid-Cooled 2.5D/3D ICs via Random Walks on Circuit Networks of Modified Compact Thermal Models
abstract
Thermal issues are critical in 2.5D/3D IC design, and liquid cooling provides an effective solution for heat dissipation. Widely used compact thermal models (CTMs) convert chips into circuit networks for fast thermal simulations. However, current matrix-solving acceleration methods for CTM-derived circuits are inadequate for high-speed iterative transient thermal analysis of large-scale liquid-cooled 2.5D/3D ICs during design optimization. In contrast, the random walk method can provide fast solutions for local nodes in large-scale circuit networks, but it is not applicable to the circuit networks of the CTMs with liquid cooling. In this paper, we propose LCTMwalk, a novel GPU-accelerated random walk method for transient thermal analysis of liquid-cooled 2.5D/3D ICs. To enable random walks on the liquid-cooled CTM-derived circuit network, we replace the voltage-controlled current source model with the diode model. Additionally, we improve the transient analysis by using a time-backward random walk with time-domain path reuse, accelerating the solution of temperature at local circuit nodes. Experimental results show LCTMwalk can solve million-scale cases in only 500 ms, and achieves a 14-22× speedup compared to the state-of-the-art alternating direction implicit (ADI) method with GPU. Besides, LCTMwalk exhibits good generalizability and can be applied to various 2.5D/3D IC structures with high accuracy (error<1 K compared to 3D-ICE).
Zhixuan Dong, Yonghan Luo, Changhao Yan, Zhaori Bi, Keren Zhu 0001, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
ICCAD6
2025 NSTherm: An Error-Bounded Network-Stochastic Fusion Thermal Simulator for Geometry-Adaptable Chiplets via Diffeomorphic Mapping and Neural-Guided Variance Reduction
abstract
For highly integrated, thermally constrained chiplets, the design process requires iterative shape optimization, making rapid thermal simulation across varying geometries critically important. Existing deterministic approaches, such as COMSOL and HotSpot require solving large-scale linear systems, incurring expensive computational costs. Stochastic methods suffer from slow convergence, demanding excessive resources for high-precision results. Current neural network (NN)-based methods necessitate retraining upon geometry modifications, limiting adaptability. Meanwhile, neural networks suffer from the absence of provable error bounds, introducing three fundamental risks in practical deployment. We enable the fast solution of heat equations for varying geometries and propose a novel solver that integrates operator learning with stochastic methods. By employing diffeomorphic mapping, our approach addresses the challenge of operator networks in handling shape variations. Furthermore, the network’s predictions guide the stochastic method for variance reduction, which extremely accelerates the traditional stochastic method, while the stochastic results provide error guarantees and corrections for the neural network’s outputs. Extensive experiments show that we achieve a speedup of 10.69-23.04× over commercial field solver COMSOL and a speedup of 5.20-11.87× over the traditional stochastic methods.
Zhixuan Dong, Yonghan Luo, Changhao Yan, Keren Zhu 0001, Zhaori Bi, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
ICCAD6
2025 BAGNet: A Boundary-Aware Graph Neural Network for SRAM Yield Analysis in Post-LayoutSimulation
abstract
Yield analysis has grown in significance with the increasing integration of SRAM arrays. The post-layout simulation introduces strong inter-column correlations in SRAM caused by parasitic parameters, thereby complicating yield analysis. However, most existing methods only consider the pre-layout simulation of SRAM circuits. In this paper, we present BAGNet: a boundary-aware Graph Neural Network (GNN) for SRAM yield analysis in post-layout simulation. We introduce a GNN module that learns the graph representations of SRAM arrays while generating feature vectors. We then construct an accurate surrogate model by the Multilayer Perceptron (MLP) to provide predictions for circuit performances. Given that delineating failure boundaries is vital for yield estimation, we propose an innovative nonlinear mapping strategy and an adaptive iterative strategy integrated with BAGNet, thus endowing our model with boundary-aware capability. After the model is built, we employ the importance sampling (IS) method on our surrogate model to deliver efficient and accurate yield estimation without time-consuming circuit simulations. Experimental results demonstrate that BAGNet outperforms the state-of-the-art method with 1.823.52x speedup, without losing accuracy.
Haoyang Sang, Changhao Yan, Zhaori Bi, Keren Zhu 0001, Xuan Zeng 0001
ICCAD4
2025 Seeing Through Designs: Attention-Based Knowledge Transfer for Preference-Guided Microarchitecture Search
abstract
Modern processor microarchitectures face increasing complexity, leading to larger search spaces and lengthy design-to-silicon validation flows. While reusing design knowledge across architectures offers potential efficiency gains, the common practice remains specific-architecture search due to inherent discrepancies in power, performance, and area (PPA) metrics between designs. We propose an attention-based microarchitecture search framework for effective cross-architecture knowledge transfer. Our approach propose a cross-attention network to capture interdependencies between microarchitectural topology and design tool configurations, enabling knowledge adaptation across architectures with minimal fine-tuning. Additionally, we complement it with an uncertainty-guided optimization strategy that efficiently navigates search based on specific user preferences. Experimental results demonstrate our approach outperforms previous methods with 68.16% higher hypervolume indicators and 3.85× speed-up of time in reaching the same hypervolume. Furthermore, our approach successfully discovers design points that meet user-specified PPA targets that state-of-the-art (SOTA) methods failed to identify. Our code is publicly available at https://github.com/MarsH3107/ICAN, enabling broader adoption and encouraging further research in transferable processor design optimization.
Zhaori Bi, Ming Zhu 0016, Qiwei Zhan, Keren Zhu 0001, Fan Yang 0001, Changhao Yan, Dian Zhou, Xuan Zeng 0001
ICCAD6
2025 HeLO: A Heterogeneous Logic Optimization Framework by Hierarchical Clustering and Graph Learning
abstract
Modern very large-scale integration (VLSI) designs usually consist of modules with various topological structures and functionalities. To better optimize such large and heterogeneous logic networks, it is essential to identify the structural and functional characteristics of its modules, and represent them with appropriate DAG types (such as AIG, MIG, XAG, etc.) for logic optimization. This paper proposes HeLO, a hetero-DAG logic optimization framework empowered by hierarchical clustering and graph learning. HeLO leverages a hierarchical clustering algorithm, which splits the original Boolean network into sub-circuits by considering both topological and functional characteristics. A novel graph neural network model is customized to generate the topological-functional embedding (used for distance calculation in hierarchical clustering) and predict the best-fit DAG type of each sub-circuit. Experimental results demonstrate that HeLO outperforms LSOracle, the SOTA heterogeneous logic optimization framework, in terms of node-depth product (for technology-independent logic optimization) and delay-area product (for technology mapping) by 8.7% and 6.9%, respectively.
Yuan Pu 0001, Fangzhou Liu 0005, Zhuolun He, Keren Zhu 0001, Rongliang Fu, Ziyi Wang 0010, Tsung-Yi Ho, Bei Yu 0001
ISPD4
2025 PARoute2: Enhanced Analog Routing via Performance-Drive Guidance Generation
abstract
Analog routing is crucial for performance optimization in analog circuit design, but conventionally takes significant development time and requires design expertise. Recent research has attempted to use machine learning (ML) to generate guidance to preserve circuit performance after analog routing. These methods face challenges such as expensive data acquisition and biased guidance. This article presents AnalogFold, a new paradigm of analog routing that leverages ML to provide performance-oriented routing guidance. Our approach learns performance-driven routing guidance and uses it to help automatic routers for performance-driven routing optimization. We propose to use a 3DGNN that incorporates cost-aware distance to make accurate predictions on post-layout performance. A pool-assisted potential relaxation process derives the effective routing guidance. The experimental results on multiple benchmarks under the TSMC 40 nm technology node demonstrate the superiority of the proposed framework compared to the cutting-edge works.
Peng Xu 0052, Jindong Tu, Guojin Chen, Keren Zhu 0001, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Rethinking Logic Rewriting: Technology-Aware Subgraph Matching with Exact Synthesis
abstract
Logic synthesis is crucial in digital design automation, significantly enhancing performance, reducing area, and lowering power consumption through technology-independent optimization followed by technology mapping. Logic rewriting, a key strategy for optimization, iteratively replaces portions of logic circuits with more compact implementations. Despite historical advancements, challenges remain in subgraph selection, technology-dependent metrics, and performance-runtime trade-offs. This article presents a novel Te chnology- a ware logic R e W riting ( TeaRW ) framework to address these challenges. TeaRW incorporates a technology-aware rewriting algorithm that evaluates post-mapping netlist metrics during the technology-independent optimization phase. It employs four distinct subgraph rewriting techniques to maximize the effectiveness of local optimization. For efficiency, TeaRW utilizes an optimized logic representation database derived from exact synthesis, enabling cost-effective replacements. Experimental results on real-world benchmarks show improvements over the ABC tool, including an average Area-Delay-Product (ADP) improvement of 8.18% in delay-oriented optimization and 0.28% in area-oriented optimization when compared to state-of-the-art optimization scripts.
Hongyang Pan, Keren Zhu 0001, Fan Yang 0001, Xuan Zeng 0001, Yun Shao 0008, Zhufei Chu
ACM Trans. Design Autom. Electr. Syst.2
2025 Rank-DSE: Neural Pareto Comparator of Microarchitecture Design Space Exploration
abstract
The complexity of microarchitecture design has surged due to the expanding design space and time-intensive verification processes. Existing regression-based machine learning methods struggle with inaccurate estimations because of limited training samples. To address these challenges, we propose Rank-DSE, a novel framework for microarchitecture design space exploration (DSE) that leverages a Neural Pareto Comparator (NPC) to directly model the comparative relationships between different architecture designs. Rank-DSE bypasses the inaccuracies of absolute PPA (performance, power, area) predictions by focusing on relative comparisons. The NPC computes the probability of one architecture dominating another and employs semi-supervised learning to reduce the reliance on labeled data. Additionally, a reinforcement-learning-based sampling scheme with an updating baseline Pareto set accelerates the exploration process. Experimental results on the ICCAD 2021 benchmark demonstrate that Rank-DSE achieves superior search quality and cost-efficiency compared to state-of-the-art methods. Specifically, Rank-DSE improves hypervolume by up to 7% while reducing exploration cost by 53.09% compared to cutting-edge approaches. These results highlight the advantages of Rank-DSE in terms of efficiency and effectiveness for microarchitecture DSE.
Peng Xu 0052, Su Zheng, Mingzi Wang, Ziyang Yu 0001, Shixin Chen, Tinghuan Chen, Keren Zhu 0001, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.7
2024 ISOP-Yield: Yield-Aware Stack-Up Optimization for Advanced Package using Machine Learning
abstract
High-speed cross-chip interconnects and packaging are critical for the overall performance of modern heterogeneous integrated computing systems. Recent studies have developed automatic stack-up design optimization methods for high-density interconnect (HDI) printed circuit board (PCB). However, few have considered the impact of manufacturing variation and the resulting yield issue in high-volume manufacturing (HVM). In this paper, we propose a novel framework for automatic stack-up design, optimizing the interconnect performance with a given yield requirement. The proposed framework utilizes the smooth and gradient-available machine learning surrogate model, employing a first-order Taylor expansion to approximate the output performance distribution. Experimental results demonstrate that our method effectively boosts the yield rate compared to the existing stack-up optimization framework. In addition, the proposed yield-aware algorithm shows an average of 49.96% efficiency improvement in yield-aware figure of merits compared to the state-of-the-art input noise-aware Bayesian optimization algorithm for high yield targets.
Hyunsu Chae, Keren Zhu 0001, Bhyrav Mutnury, Zixuan Jiang, Daniel De Araujo, Douglas Wallace, Douglas Winterberg, Adam R. Klivans, David Z. Pan
ASPDAC2
2024 A Study on Exploring and Exploiting the High-dimensional Design Space for Analog Circuit Design Automation : (Invited Paper)
abstract
The escalated intricacy of analog circuits, compounded by the high-dimensional nature of the design space, introduces complexities in optimizing circuit performance. Since the evaluation cost, often through circuit simulation, is resource-intensive and time-consuming, it is crucial to obtain a feasible design with a decent Figure of Merit (FOM) value within a limited simulation budget. In this study, we conduct an in-depth review and analysis of cutting-edge exploration and exploitation techniques developed to address the intricacies encountered in analog circuit design automation. Moreover, to enable algorithmic comparisons and advance the state of the field, we provide benchmarks encompassing analog circuit netlists with high-dimensional design variables, which empower researchers to rigorously assess and refine their optimization algorithms, leading to enhanced efficacy and novel developments.
Ruiyu Lyu, Aidong Zhao, Zhaori Bi, Keren Zhu 0001, Fan Yang 0001, Changhao Yan, Dian Zhou, Xuan Zeng 0001
ASPDAC5
2024 Performance-Driven Analog Layout Automation: Current Status and Future Directions (Invited Paper)
abstract
Optimizing circuit performance presents a pivotal challenge in the realm of automatic analog physical design. The intricacy of analog performance arises from its sensitivity to layout implementation, frequently lacking a viable approach for direct optimization. This talk initiates with a comprehensive overview of the present challenges and the techniques currently in use. The emphasis will be laid on the recent advancements in employing black-box optimization for enhancing analog performance. Subsequently, we will delve into a detailed case study and analysis of post-layout performance distribution for a typical analog circuit. This study will showcase various layout implementations generated by the open-source analog layout generator, MAGICAL. Future directions will be discussed based on the case study.
Peng Xu 0052, Jintao Li 0002, Tsung-Yi Ho, Bei Yu 0001, Keren Zhu 0001
ASPDAC5
2024 Performance-driven Analog Routing via Heterogeneous 3DGNN and Potential Relaxation
abstract
Analog routing is crucial for performance optimization in analog circuit design, but conventionally takes significant development time and requires design expertise. Recent research has attempted to use machine learning (ML) to generate guidance to preserve circuit performance after analog routing. These methods face challenges such as expensive data acquisition and biased guidance. This paper presents AnalogFold, a new paradigm of analog routing that leverages ML to provide performance-oriented routing guidance. Our approach learns performance-driven routing guidance and uses it to help automatic routers for performance-driven routing optimization. We propose to use a 3DGNN that incorporates cost-aware distance to make accurate predictions on post-layout performance. A pool-assisted potential relaxation process derives the effective routing guidance. The experimental results on multiple benchmarks under the TSMC 40nm technology node demonstrate the superiority of the proposed framework compared to the cutting-edge works.
Peng Xu 0052, Guojin Chen, Keren Zhu 0001, Tinghuan Chen, Tsung-Yi Ho, Bei Yu 0001
DAC3
2024 A Data-Driven Analog Circuit Synthesizer with Automatic Topology Selection and Sizing
abstract
Despite significant recent advancements in analog design automation, analog front-end design remains a challenge characterized by its heavy reliance on human designer expertise together with extensive trial-and-error simulations. In this paper, we present a novel data-driven analog circuit synthesizer with automatic topology selection and sizing. We propose a modular approach to build a comprehensive, parameterized circuit topology library. Instead of starting from an exhaustive dataset, which is often not available or too expensive to build, we build an adaptive topology dataset, which can later be enhanced with synthetic data generated using variational autoencoders (VAE), a generative machine learning technique. This integration bolsters our methodology's predictive capabilities, minimizing the risk of inadvertent oversight of viable topologies. To ensure accuracy and robustness, the predicted topology is re-sized for verification and further performance optimization. Our experiments, which involve over 360 OPAMP topologies and over 540K data points demonstrate our framework's capability to identify optimal topology and its sizing within minutes, achieving design quality comparable to that of experienced designers.
Souradip Poddar, Ahmet Faruk Budak, Linran Zhao, Chen-Hao Hsu, Supriyo Maji, Keren Zhu 0001, Yaoyao Jia, David Z. Pan
DATE6
2024 AnalogGym: An Open and Practical Testing Suite for Analog Circuit Synthesis
abstract
Recent advances in machine learning (ML) for automating analog circuit synthesis have been significant, yet challenges remain. A critical gap is the lack of a standardized evaluation framework, compounded by various process design kits (PDKs), simulation tools, and a limited variety of circuit topologies. These factors hinder direct comparisons and the validation of algorithms. To address these shortcomings, we introduced AnalogGym, an open-source testing suite designed to provide fair and comprehensive evaluations. AnalogGym includes 30 circuit topologies in five categories: sensing front ends, voltage references, low dropout regulators, amplifiers, and phase-locked loops. It supports several technology nodes for academic and commercial applications and is compatible with commercial simulators such as Cadence Spectre, Synopsys HSPICE, and the open-source simulator Ngspice. AnalogGym standardizes the assessment of ML algorithms in analog circuit synthesis and promotes reproducibility with its open datasets and detailed benchmark specifications. AnalogGym's user-friendly design allows researchers to easily adapt it for robust, transparent comparisons of state-of-the-art methods, while also exposing them to real-world industrial design challenges, enhancing the practical relevance of their work. Additionally, we have conducted a comprehensive comparison study of various analog sizing methods on AnalogGym, highlighting the capabilities and advantages of different approaches. AnalogGym is available in the GitHub repository1. The documentations are also available at2.
Jintao Li 0002, Haochang Zhi, Ruiyu Lyu, Wangzhen Li, Zhaori Bi, Keren Zhu 0001, Yanhan Zeng, Weiwei Shan, Changhao Yan, Fan Yang 0001, Yun Li 0002, Xuan Zeng 0001
ICCAD6
2024 Revisiting sensitivity-based analog sizing with derivative-aware Bayesian optimization and error-suppressed adjoint analysis
abstract
Current state-of-the-art (SOTA) analog circuit sizing methods predominantly rely on derivative-free algorithms. However, these methods struggle with sample efficiency due to the lack of derivative information, acting as a bottleneck for further advancements. In contrast, classic sensitivity analysis computes partial derivatives of circuit performance with respect to design parameters, enabling efficient first-order optimization. Yet, sensitivity-driven analog sizing has seen limited use due to: 1) accumulated numerical errors from nonlinear devices, and 2) the complex, non-convex nature of circuit optimization problems, which makes local search methods like gradient descent ineffective for global optimization. To address these challenges, this paper equips SOTA analog sizing algorithms with derivative awareness and proposes DarBO, a Derivative-aware Bayesian Optimization method. DarBO uses derivatives from error-suppressed adjoint sensitivity analysis to improve Gaussian process posteriors in local optimization, enhancing convergence with fewer circuit simulations. For global exploration, DarBO adapts a derivative-aware Gaussian mixture model (d-GMM) for region partitioning and a gradient-driven Monte Carlo tree search (d-MCTS) for subregion selection. By bridging classic sensitivity-driven analog sizing with SOTA Bayesian optimization algorithms, DarBO offers an efficient and robust solution for analog circuit sizing. Experimental results show that DarBO achieves up to 5.0 × acceleration in terms of the number of circuit simulations compared to existing first-order and derivative-free optimization methods.
Ruiyu Lyu, Aidong Zhao, Keren Zhu 0001, Zhaori Bi, Changhao Yan, Fan Yang 0001, Dian Zhou, Xuan Zeng 0001
ICCAD4
2024 Physically Aware Synthesis Revisited: Guiding Technology Mapping with Primitive Logic Gate Placement
abstract
A typical VLSI design flow is divided into separated front-end logic synthesis and back-end physical design (PD) stages, which often require costly iterations between these stages to achieve design closure. Existing approaches face significant challenges, notably in utilizing feedback from physical metrics to better adapt and refine synthesis operations, and in establishing a unified and comprehensive metric. This paper introduces a new Primitive logic gate placement guided technology MAPping (PigMAP) framework to address these challenges. With approximating technology-independent spatial information, we develop a novel wirelength (WL) driven mapping algorithm to produce PD-friendly netlists. PigMAP is equipped with two schemes: a performance mode that focuses on optimizing the critical path WL to achieve high performance, and a power mode that aims to minimize the total WL, resulting in balanced power and performance outcomes. We evaluate our framework using the EPFL benchmark suites with ASAP7 technology, using the OpenROAD tool for place-and-route. Compared with OpenROAD flow scripts, performance mode reduces delay by 14% while increasing power consumption by only 6%. Meanwhile, power mode achieves a 3% improvement in delay and a 9% reduction in power consumption.
Hongyang Pan, Cunqing Lan, Yiting Liu 0002, Zhiang Wang, Li Shang 0001, Xuan Zeng 0001, Fan Yang 0001, Keren Zhu 0001
ICCAD8
2024 Multi-Electrostatics Based Placement for Non-Integer Multiple-Height Cells
abstract
A circuit design incorporating non-integer multi-height (NIMH) cells, such as a combination of 8-track and 12-track cells, offers increased flexibility in optimizing area, timing, and power simultaneously. The conventional approach for placing NIMH cells involves using commercial tools to generate an initial global placement, followed by a legalization process that divides the block area into row regions with specific heights and relocates cells to rows of matching height. However, such placement flow often causes significant disruptions in the initial placement results, resulting in inferior wirelength. To address this issue, we propose a novel multi-electrostatics-based global placement algorithm that utilizes the NIMH-aware clustering method to dynamically generate rows. This algorithm directly tackles the global placement problem with NIMH cells. Specifically, we utilize an augmented Lagrangian formulation along with a preconditioning technique to achieve high-quality solutions with fast and robust numerical convergence. Experimental results on the OpenCores benchmarks demonstrate that our algorithm achieves about 12% improvements on HPWL with 23.5X speed up on average, outperforming state-of-the-art approaches. Furthermore, our placement solutions demonstrate a substantial improvement in WNS and TNS by 22% and 49% respectively. These results affirm the efficiency and effectiveness of our proposed algorithm in solving row-based placement problems for NIMH cells.
Yu Zhang 0189, Yuan Pu 0001, Fangzhou Liu 0005, Peiyu Liao, Kai-Yuan Chao, Keren Zhu 0001, Yibo Lin, Bei Yu 0001
ISPD6
2024 Large circuit models: opportunities and challenges
abstract
Abstract Within the electronic design automation (EDA) domain, artificial intelligence (AI)-driven solutions have emerged as formidable tools, yet they typically augment rather than redefine existing methodologies. These solutions often repurpose deep learning models from other domains, such as vision, text, and graph analytics, applying them to circuit design without tailoring to the unique complexities of electronic circuits. Such an “AI4EDA” approach falls short of achieving a holistic design synthesis and understanding, overlooking the intricate interplay of electrical, logical, and physical facets of circuit data. This study argues for a paradigm shift from AI4EDA towards AI-rooted EDA from the ground up, integrating AI at the core of the design process. Pivotal to this vision is the development of a multimodal circuit representation learning technique, poised to provide a comprehensive understanding by harmonizing and extracting insights from varied data sources, such as functional specifications, register-transfer level (RTL) designs, circuit netlists, and physical layouts. We champion the creation of large circuit models (LCMs) that are inherently multimodal, crafted to decode and express the rich semantics and structures of circuit data, thus fostering more resilient, efficient, and inventive design methodologies. Embracing this AI-rooted philosophy, we foresee a trajectory that transcends the current innovation plateau in EDA, igniting a profound “shift-left” in electronic design methodology. The envisioned advancements herald not just an evolution of existing EDA tools but a revolution, giving rise to novel instruments of design-tools that promise to radically enhance design productivity and inaugurate a new epoch where the optimization of circuit performance, power, and area (PPA) is achieved not incrementally, but through leaps that redefine the benchmarks of electronic systems’ capabilities.
Zhufei Chu, Wenji Fang, Tsung-Yi Ho, Ru Huang 0001, Yu Huang 0005, Sadaf Khan, Yun Liang 0001, Yibo Lin, Guojie Luo, Hongyang Pan, Zhengyuan Shi, Guangyu Sun 0003, Dimitrios Tsaras, Runsheng Wang, Ziyi Wang 0010, Xinming Wei, Zhiyao Xie, Qiang Xu 0001, Chenhao Xue, Junchi Yan, Bei Yu 0001, Mingxuan Yuan, Evangeline F. Y. Young, Xuan Zeng 0001, Haoyi Zhang, Zuodong Zhang, Hui-Ling Zhen, Binwu Zhu, Keren Zhu 0001, Sunan Zou
Sci. China Inf. Sci.39
2024 Erratum to: Large circuit models: opportunities and challenges
Zhufei Chu, Wenji Fang, Tsung-Yi Ho, Ru Huang 0001, Yu Huang 0005, Sadaf Khan, Yun Liang 0001, Yibo Lin, Guojie Luo, Hongyang Pan, Zhengyuan Shi, Guangyu Sun 0003, Dimitrios Tsaras, Runsheng Wang, Ziyi Wang 0010, Xinming Wei, Zhiyao Xie, Qiang Xu 0001, Chenhao Xue, Junchi Yan, Bei Yu 0001, Mingxuan Yuan, Evangeline F. Y. Young, Xuan Zeng 0001, Haoyi Zhang, Zuodong Zhang, Hui-Ling Zhen, Binwu Zhu, Keren Zhu 0001, Sunan Zou
Sci. China Inf. Sci.39
2024 ISOP+: Machine Learning-Assisted Inverse Stack-Up Optimization for Advanced Package Design
abstract
The future of computing requires heterogeneous integration, including the recent adoption of chiplet methodology, where high-speed cross-chip interconnects and packaging are critical for the overall system performance. As an example of advanced packaging, a high-density interconnect (HDI) printed circuit board (PCB) has been widely used in complex electronics ranging from cell phones to computing servers. A modern HDI PCB may have over 20 layers, each with its unique material properties and geometrical dimensions, i.e., stack-up, to meet various design constraints and performance requirements. Stack-up design is usually done manually in the industry, where experienced designers may devote many hours adjusting the physical dimensions and materials in order to meet the desired specifications. This process, however, is time-consuming, tedious, and suboptimal, largely depending on the designer’s expertise. In this article, we propose to automate the stack-up design with a new framework, ISOP+, using machine learning (ML) for inverse stack-up optimization for advanced package design with adaptive weight adjustment and multilevel optimization. Given a target design specification, ISOP+ automatically searches for ideal stack-up design parameters while optimizing performance. A novel ML-assisted hyperparameter optimization method is developed to make the search efficient and reliable. Experimental results demonstrate that ISOP+ is better in figure-of-merit (FoM) than conventional simulated annealing and Bayesian optimization algorithms, with all our design targets met with a shorter runtime. We also compare our fully automated ISOP+ with expert designers in the industry and achieve very promising results, with orders of magnitude reduction of turn-around time.
Hyunsu Chae, Keren Zhu 0001, Bhyrav Mutnury, Douglas Wallace, Douglas Winterberg, Daniel De Araujo, Jay Reddy, Adam R. Klivans, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 ISOP: Machine Learning-Assisted Inverse Stack-Up Optimization for Advanced Package Design
abstract
Future computing calls for heterogeneous integration, e.g., the recent adoption of the chiplet methodology. However, high-speed cross-chip interconnects and packaging shall be critical for the overall system performance. As an example of advanced packaging, a high-density interconnect (HDI) printed circuit board (PCB) has been widely used in complex electronics from cell phones to computing servers. A modern HDI PCB may have over 20 layers, each with its unique material properties and geometrical dimensions, i.e., stack-up, to meet various design constraints and performance optimizations. However, stack-up design is usually done manually in the industry, where experienced designers may devote many hours to adjusting the physical dimensions and materials to meet the desired specifications. This process, however, is time-consuming, tedious, and sub-optimal, largely depending on the designer's expertise. In this paper, we propose to automate the stack-up design with a new framework, ISOP, using machine learning for inverse stack-up optimization for advanced package design. Given a target design specification, ISOP automatically searches for ideal stack-up design parameters while optimizing performance. We develop a novel machine learning-assisted hyper-parameter optimization method to make the search efficient and reliable. Experimental results demonstrate that ISOP is better in figure-of-merit (FoM) than conventional simulated annealing and Bayesian optimization algorithms, with all our design targets met with a shorter runtime. We also compare our fully-automated ISOP with expert designers in the industry and achieve very promising results, with orders of magnitude reduction of turn-around time.
Hyunsu Chae, Bhyrav Mutnury, Keren Zhu 0001, Douglas Wallace, Douglas Winterberg, Daniel De Araujo, Jay Reddy, Adam R. Klivans, David Z. Pan
DATE3
2023 Practical Layout-Aware Analog/Mixed-Signal Design Automation with Bayesian Neural Networks
abstract
The high simulation cost has been a bottleneck of practical analog/mixed-signal design automation. Many learning-based algorithms require thousands of simulated data points, which is impractical for expensive to simulate circuits. We propose a learning-based algorithm that can be trained using a small amount of data and, therefore, scalable to tasks with expensive simulations. Our efficient algorithm solves the post-layout performance optimization problem where simulations are known to be expensive. Our comprehensive study also solves the schematic-level sizing problem. For efficient optimization, we utilize Bayesian Neural Networks as a regression model to approximate circuit performance. For layout-aware optimization, we handle the problem as a multi-fidelity optimization problem and improve efficiency by exploiting the correlations from cheaper evaluations. We present three test cases to demonstrate the efficiency of our algorithms. Our tests prove that the proposed approach is more efficient than conventional baselines and state-of -the-art algorithms.
Ahmet Faruk Budak, Keren Zhu 0001, David Z. Pan
ICCAD2
2023 AlphaSyn: Logic Synthesis Optimization with Efficient Monte Carlo Tree Search
abstract
Recent years have seen rising research in logic synthesis recipe generation to improve the Quality-of-Result (QoR). However, existing approaches typically have low efficiency and are stuck at local optima. In this work, we propose a logic synthesis optimization framework, AlphaSyn, that incorporates a domain-specific Monte Carlo tree search (MCTS) algorithm. AlphaSyn enables exploration across the entire search space while optimizing sampling points utilization. We further develop a synthesis-specific upper confidence bound for trees (SynUCT) algorithm for the selection phase and a well-designed learning strategy to enhance the stability of the MCTS algorithm. The AlphaSyn algorithm is fully parallelized for efficiency with asynchronous MCTS exploration and significance-base resource allocation. For standard-cell technology mapping on the ASAP 7nm library among other tasks, experimental results show that AlphaSyn outperforms SOTA FlowTune with an average 8.74% area reduction and$\boldsymbol{1.24}\times$runtime speedup.
Zehua Pei, Fangzhou Liu 0005, Zhuolun He, Guojin Chen, Haisheng Zheng, Keren Zhu 0001, Bei Yu 0001
ICCAD6
2023 Reinforcement Learning Guided Detailed Routing for Custom Circuits
abstract
Detailed routing is the most tedious and complex procedure in design automation and has become a determining factor in layout automation in advanced manufacturing nodes. Despite continuing advances in custom integrated circuit (IC) routing research, industrial custom layout flows remain heavily manual due to the high complexity of the custom IC design problem. Besides conventional design objectives such as wirelength minimization, custom detailed routing must also accommodate additional constraints (e.g., path-matching) across the analog/mixed-signal (AMS) and digital domains, making an already challenging procedure even more so. This paper presents a novel detailed routing framework for custom circuits that leverages deep reinforcement learning to optimize routing patterns while considering custom routing constraints and industrial design rules. Comprehensive post-layout analyses based on industrial designs demonstrate the effectiveness of our framework in dealing with the specified constraints and producing sign-off-quality routing solutions.
Hao Chen 0059, Kai-Chieh Hsu, Walker J. Turner, Po-Hsuan Wei, Keren Zhu 0001, David Z. Pan, Haoxing Ren
ISPD5
2023 Joint Optimization of Sizing and Layout for AMS Designs: Challenges and Opportunities
abstract
Recent advances in analog device sizing algorithms show promising results on the automatic schematic design. However, the majority of the sizing algorithms are based on schematic-level simulations and layout-agnostic. The physical layout implementation brings extra parasitics to the analog circuits, leading to discrepancies between schematic and post-layout performance. This performance gap raises questions about the effectiveness of automatic analog device sizing tools. Prior work has leveraged procedural layout generation to account for layout-induced parasitics in the sizing process. However, the need for layout templates makes such methodology limited in application. In this paper, we propose to bridge automatic analog sizing with post-layout performance using state-of-the-art optimization-based analog layout generators. A quantitative study is conducted to measure the impact of layout awareness in state-of-the-art device sizing algorithms. Furthermore, we present our perspectives on the future directions in layout-aware analog circuit schematic design.
Ahmet Faruk Budak, Keren Zhu 0001, Hao Chen 0059, Souradip Poddar, Linran Zhao, Yaoyao Jia, David Z. Pan
ISPD2
2023 Hierarchical Analog and Mixed-Signal Circuit Placement Considering System Signal Flow
abstract
Placement is a critical step in layout automation for analog and mixed-signal (AMS)-integrated circuits (ICs). It determines the proximity of devices and influences the wiring topology, significantly impacting post-routing parasitics and coupling capacitance. Existing analog placement techniques mainly focus on geometric constraints in analog building blocks. However, there yet lacks an effective way to consider the system-level signal flow for sensitive AMS circuits. Leveraging prior knowledge from schematics, we propose considering the critical signal paths in automatic AMS placement. A multilevel analog layout automation flow is further developed to reduce manual efforts in synthesizing hierarchical AMS circuits. Experimental results demonstrate the efficiency and effectiveness of our proposed framework with a 22.8% reduction in routed wirelength compared to state-of-the-art AMS placer and a 10-dB improvement in the signal-to-noise-and-distortion ratio (SNDR) for an ADC.
Keren Zhu 0001, Hao Chen 0059, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Automating Analog Constraint Extraction: From Heuristics to Learning: (Invited Paper)
abstract
Analog layout synthesis has recently received much attention to mitigate the increasing cost of manual layout efforts. To achieve the desired performance and design specifications, generating layout constraints is critical in fully automated netlist-to-GDSII analog layout flow. However, there is a big gap between automatic constraint extraction and constraint management in analog layout synthesis. This paper introduces the existing constraint types for analog layout synthesis and points out the recent research trends in automating analog constraint extraction. Specifically, the paper reviews the conventional graph heuristic methods such as graph similarity and the recent machine learning approach leveraging graph neural networks. It also discusses challenges and research opportunities.
Keren Zhu 0001, Hao Chen 0059, David Z. Pan
ASP-DAC1
2022 Generative-Adversarial-Network-Guided Well-Aware Placement for Analog Circuits
abstract
Generating wells for transistors is an essential challenge in analog circuit layout synthesis. While it is closely related to analog placement, very little research has explicitly considered well generation within the placement process. In this work, we propose a new analytical well-aware analog placer. It uses a generative adversarial network (GAN) for generating wells and guides the placement process. A global placement algorithm spreads the modules given the GAN guidance and optimizes for area and wirelength. Well-aware legalization techniques then legalize the global placement results and produce the final placement solutions. By allowing well sharing between transistors and explicitly considering wells in placement, the proposed framework achieves more than 74% improvement in the area and more than 26% reduction in half-perimeter wirelength over existing placement methodologies in experimental results.
Keren Zhu 0001, Hao Chen 0059, Xiyuan Tang, Wei Shi 0011, Nan Sun 0001, David Z. Pan
ASP-DAC1
2022 Reinforcement Learning for Electronic Design Automation: Case Studies and Perspectives: (Invited Paper)
abstract
Reinforcement learning (RL) algorithms have recently seen rapid advancement and adoption in the field of electronic design automation (EDA) in both academia and industry. In this paper, we first give an overview of RL and its applications in EDA. In particular, we discuss three case studies: chip macro placement, analog transistor sizing, and logic synthesis. In collaboration with Google Brain, we develop a hybrid RL and analytical mixed -size placer and achieve better results with less training time on public and proprietary benchmarks. Working with Intel, we develop an RL-inspired optimizer for analog circuit sizing, combining the strengths of deep neural networks and reinforcement learning to achieve state-of-the-art black-box optimization results. We also apply RL to the popular logic synthesis framework ABC and obtain promising results. Through these case studies, we discuss the advantages, disadvantages, opportunities, and challenges of RL in EDA.
Ahmet Faruk Budak, Zixuan Jiang, Keren Zhu 0001, Azalia Mirhoseini, Anna Goldie, David Z. Pan
ASP-DAC3
2022 TAG: Learning Circuit Spatial Embedding from Layouts
abstract
Analog and mixed-signal (AMS) circuit designs still rely on human design expertise. Machine learning has been assisting circuit design automation by replacing human experience with artificial intelligence. This paper presents TAG, a new paradigm of learning the circuit representation from layouts leveraging Text, self Attention and Graph. The embedding network model learns spatial information without manual labeling. We introduce text embedding and a self-attention mechanism to AMS circuit learning. Experimental results demonstrate the ability to predict layout distances between instances with industrial FinFET technology benchmarks. The effectiveness of the circuit representation is verified by showing the transferability to three other learning tasks with limited data in the case studies: layout matching prediction, wirelength estimation, and net parasitic capacitance prediction.
Keren Zhu 0001, Hao Chen 0059, Walker J. Turner, George F. Kokai, Po-Hsuan Wei, David Z. Pan, Haoxing Ren
ICCAD1
2022 Fuse and Mix: MACAM-Enabled Analog Activation for Energy-Efficient Neural Acceleration
abstract
Analog computing has been recognized as a promising low-power alternative to digital counterparts for neural network acceleration. However, conventional analog computing is mainly in a mixed-signal manner. Tedious analog/digital (A/D) conversion cost significantly limits the overall system's energy efficiency. In this work, we devise an efficient analog activation unit with magnetic tunnel junction (MTJ)-based analog content-addressable memory (MACAM), simultaneously realizing nonlinear activation and A/D conversion in a fused fashion. To compensate for the nascent and therefore currently limited representation capability of MACAM, we propose to mix our analog activation unit with digital activation dataflow. A fully differential framework, SuperMixer, is developed to search for an optimized activation workload assignment, adaptive to various activation energy constraints. The effectiveness of our proposed methods is evaluated on a silicon photonic accelerator. Compared to standard activation implementation, our mixed activation system with the searched assignment can achieve competitive accuracy with >60% energy saving on A/D conversion and activation.
Hanqing Zhu, Keren Zhu 0001, Jiaqi Gu 0002, Harrison Jin, Ray T. Chen, Jean Anne C. Incorvia, David Z. Pan
ICCAD2
2022 AutoCRAFT: Layout Automation for Custom Circuits in Advanced FinFET Technologies
abstract
Despite continuous efforts in layout automation for full-custom circuits, including analog/mixed-signal (AMS) designs, automated layout tools have not yet been widely adopted in current industrial full-custom design flows due to the high circuit complexity and sensitivity to layout parasitics. Nevertheless, the strict design rules and grid-based restrictions in nanometer-scale FinFET nodes limit the degree of freedom in full-custom layout design and thus reduce the gap between automation tools and human experts. This paper presents AutoCRAFT, an automatic layout generator targeting region-based layouts for advanced FinFET-based full-custom circuits. AutoCRAFT uses specialized place-and-route (P&R) algorithms to handle various design constraints while adhering to typical FinFET layout styles. Verified by comprehensive post-layout analyses, AutoCRAFT has achieved promising preliminary results in generating sign-off quality layouts for industrial benchmarks.
Hao Chen 0059, Walker J. Turner, Sanquan Song, Keren Zhu 0001, George F. Kokai, Brian Zimmer, C. Thomas Gray, Brucek Khailany, David Z. Pan, Haoxing Ren
ISPD4
2021 Universal Symmetry Constraint Extraction for Analog and Mixed-Signal Circuits with Graph Neural Networks
abstract
Recent research trends in analog layout synthesis aim for a fully automated netlist-to-GDSII design flow with minimum human efforts. Due to the sensitiveness of analog circuit layouts, symmetry matching between critical building blocks and devices can significantly impact the overall circuit performance. Therefore, providing accurate symmetry constraints for automated layout synthesis tools is crucial to achieving high-quality layouts. This paper presents a novel graph-learning-based framework leveraging unsupervised learning to recognize circuit matching structures by making the most of numerous unlabeled circuits. The proposed framework supports both system-level and device-level symmetry constraints extraction for various large-scale analog/mixed-signal systems. Experimental results show that our framework outperforms state-of-the-art symmetry constraint detection algorithms with remarkable accuracy and runtime improvement.
Hao Chen 0059, Keren Zhu 0001, Xiyuan Tang, Nan Sun 0001, David Z. Pan
DAC2
2021 OpenSAR: An Open Source Automated End-to-end SAR ADC Compiler
abstract
Despite recent developments in automated analog sizing and analog layout generation, there is doubt whether analog design automation techniques could scale to system-level designs. On the other hand, analog designs are considered major roadblocks for open source hardware with limited available design automation tools. In this work, we present OpenSAR, the first open source automated end-to-end successive approximation register (SAR) analog-to-digital converter (ADC) compiler. OpenSAR only requires system performance specifications as the minimal input and outputs DRC and LVS clean layouts. Compared with prior work, we leverage automated placement and routing to generate analog building blocks, removing the need to design layout templates or libraries. We optimize the redundant non-binary capacitor digital-to-analog converter (CDAC) array design for yield considerations with a template-based layout generator that interleaves capacitor rows and columns to reduce process gradient mismatch. Post layout simulations demonstrate that the generated prototype designs achieve state-of-the-art resolution, speed, and energy efficiency.
Xiyuan Tang, Keren Zhu 0001, Hao Chen 0059, Nan Sun 0001, David Z. Pan
ICCAD3
2020 S3DET: Detecting System Symmetry Constraints for Analog Circuits with Graph Similarity
abstract
Symmetry and matching between critical building blocks have a significant impact on analog system performance. However, there is limited research on generating system level symmetry constraints. In this paper, we propose a novel method of detecting system symmetry constraints for analog circuits with graph similarity. Leveraging spectral graph analysis and graph centrality, the proposed algorithm can be applied to circuits and systems of large scale and different architectures. To the best of our knowledge, this is the first work in detecting system level symmetry constraints for analog and mixed-signal (AMS) circuits. Experimental results show that the proposed method can achieve high accuracy of 88.3% with low false alarm rate of less than 1.1% in largescale AMS designs.
Wuxi Li, Keren Zhu 0001, Biying Xu, Yibo Lin, Linxiao Shen, Xiyuan Tang, Nan Sun 0001, David Z. Pan
ASP-DAC3
2020 Closing the Design Loop: Bayesian Optimization Assisted Hierarchical Analog Layout Synthesis
abstract
Existing analog layout synthesis tools provide little guarantee to post layout performance and have limited capabilities of handling system-level designs. In this paper, we present a closed-loop hierarchical analog layout synthesizer, capable of handling system designs. To ensure system performance, the building block layout implementations are optimized efficiently, utilizing post layout simulations with multi-objective Bayesian optimization. To the best of our knowledge, this is the first work demonstrating success in automated layout synthesis on generic analog system designs. Experimental results show our synthesized continuous-time ΔΣ modulator (CTDSM) achieves post layout performance of 65.9dB in signal to noise and distortion ratio (SNDR), compared with 67.8dB in the schematic design.
Keren Zhu 0001, Xiyuan Tang, Biying Xu, Wei Shi 0011, Nan Sun 0001, David Z. Pan
DAC2
2020 Towards Decrypting the Art of Analog Layout: Placement Quality Prediction via Transfer Learning
abstract
Despite tremendous efforts in analog layout automation, little adoption has been demonstrated in practical design flows. Traditional analog layout synthesis tools use various heuristic constraints to prune the design space to ensure post layout performance. However, these approaches provide limited guarantee and poor generalizability due to a lack of model mapping layout properties to circuit performance. In this paper, we attempt to shorten the gap in post layout performance modeling for analog circuits with a quantitative statistical approach. We leverage a state-of-the-art automatic analog layout tool and industry-level simulator to generate labeled training data in an automated manner. We propose a 3D convolutional neural network (CNN) model to predict the relative placement quality using well-crafted placement features. To achieve data-efficiency for practical usage, we further propose a transfer learning scheme that greatly reduces the amount of data needed. Our model would enable early pruning and efficient design explorations for practical layout design flows. Experimental results demonstrate the effectiveness and generalizability of our method across different operational transconductance amplifier (OTA) designs.
Keren Zhu 0001, Jiaqi Gu 0002, Linxiao Shen, Xiyuan Tang, Nan Sun 0001, David Z. Pan
DATE2
2020 An Efficient Training Framework for Reversible Neural Architectures
Zixuan Jiang, Keren Zhu 0001, Jiaqi Gu 0002, David Z. Pan
ECCV (27)2
2020 Effective Analog/Mixed-Signal Circuit Placement Considering System Signal Flow
abstract
Placement is among the most critical steps in analog/mixed-signal (AMS) circuit layout synthesis. It implicitly determines the wiring topology and therefore has considerable impacts on post-layout parasitics and coupling. Existing analog placement techniques are mainly focusing on geometric constraints in analog building blocks. However, there yet lacks an effective way to consider the systemlevel signal flow for sensitive AMS circuits. Leveraging prior knowledge from schematics, we propose to consider the critical signal paths in automatic AMS placement and present an efficient framework. Experimental results demonstrate our proposed framework's efficiency and effectiveness with a 22.8% reduction in routed wire-length compared to state-of-the-art AMS placer and 10 dB improvement in the signal-to-noise-and-distortion ratio (SNDR) for an ADC.
Keren Zhu 0001, Hao Chen 0059, Xiyuan Tang, Nan Sun 0001, David Z. Pan
ICCAD1
2020 Toward Silicon-Proven Detailed Routing for Analog and Mixed-Signal Circuits
abstract
Detailed routing is an intricate and tedious procedure in design automation and has become a crucial step for advanced node enablement. Compared with its advances in digital design, detailed routing for analog/mixed-signal (AMS) integrated circuits (ICs) is still heavily manual. In AMS designs, the sensitive net coupling issues and analog-specific constraints make detailed routing even more challenging. This work presents a novel and efficient detailed routing framework for automated AMS layout synthesis considering industrial design rules as well as analog-specific geometric and electrical constraints. Experimental results demonstrate the efficiency and effectiveness of our approach in optimizing circuit performance while satisfying the specified constraints. Post-layout simulations further prove that our detailed routing results can achieve sign-off quality.
Hao Chen 0059, Keren Zhu 0001, Xiyuan Tang, Nan Sun 0001, David Z. Pan
ICCAD2
2019 MAGICAL: Toward Fully Automated Analog IC Layout Leveraging Human and Machine Intelligence: Invited Paper
abstract
Despite tremendous advancement of digital IC design automation tools over the last few decades, analog IC layout is still heavily manual which is very tedious and error-prone. This paper will first review the history, challenges, and current status of analog IC layout automation. Then, we will present MAGICAL, a human-intelligence inspired, fully-automated analog IC layout system currently being developed under the DARPA IDEA program. It starts from an unannotated netlist, performs automatic layout constraint extraction and device generation, then performs placement and post-placement optimization, followed by routing to obtain the final GDSII layout. Various analytical, heuristic, and machine learning algorithms will be discussed. MAGICAL has obtained promising preliminary results. We will conclude the paper with further discussions on challenges and future directions for fully-automated analog IC layout.
Biying Xu, Keren Zhu 0001, Yibo Lin, Shaolan Li, Xiyuan Tang, Nan Sun 0001, David Z. Pan
ICCAD2
2019 GeniusRoute: A New Analog Routing Paradigm Using Generative Neural Network Guidance
abstract
Due to sensitive layout-dependent effects and varied performance metrics, analog routing automation for performance-driven layout synthesis is difficult to generalize. Existing research has proposed a number of heuristic layout constraints targeting specific performance metrics. However, previous frameworks fail to automatically combine routing with human intelligence. This paper proposes a novel, fully automated, analog routing paradigm that leverages machine learning to provide routing guidance, mimicking the sophisticated manual layout approaches. Experiments show that the proposed methodology obtains significant improvements over existing techniques and achieves competitive performance to manual layouts while being capable of generalizing to circuits of different functionality.
Keren Zhu 0001, Yibo Lin, Biying Xu, Shaolan Li, Xiyuan Tang, Nan Sun 0001, David Z. Pan
ICCAD1