Su Zheng

dblp:88/5153 · DBLP profile ↗
← Back
32ranked-venue papers
15as first author
30since 2021 · last 2026
0000-0003-1159-1611ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 10 first-author · 27 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CausalTuner: Will Causality Help High-Dimensional EDA Tool Parameter Tuning
abstract
Electronic Design Automation (EDA) tools are central to Very Large Scale Integration (VLSI) design, where numerous parameters govern the Quality-of-Result (QoR) metrics, including performance, power, and area. The high dimensionality of the parameter space, coupled with complex interactions, makes manual tuning inefficient and hinders the scalability of automated methods. Existing methods typically treat parameters as flat vectors, neglecting the EDA flow’s hierarchical causal structure, where early-stage decisions constrain later downstream stages. To address this, we propose CausalTuner, a causality-aware design space exploration framework for efficient parameter tuning. It employs a hybrid causal attention mechanism to capture stage-wise parameter interactions and embeds them into deep kernel Gaussian processes for accurate and generalizable surrogate modeling. The causal exploration strategies enhance sampling efficiency. Experiments show that CausalTuner outperforms state-of-the-art methods in both final QoR and efficiency.
Ziyang Yu 0001, Peng Xu 0052, Su Zheng, Hao Geng, Bei Yu 0001, Martin D. F. Wong
ASP-DAC3
2026 Scalable Second-Order Optimizer for Full-Chip Inverse Lithography Techniques
abstract
Full-chip inverse lithography techniques (ILT) represent an advanced methodology for next-generation mask optimization, enhancing sub-wavelength patterning but often facing prohibitive computational costs. State-of-the-art methods rely on iterative first-order optimizers, which exhibit slow convergence, often requiring hundreds of iterations. This inefficiency is compounded by the high overhead of repeated Fast Fourier Transform (FFT) operations and inter-GPU communication per iteration. To overcome this fundamental bottleneck, we propose a scalable second-order optimizer for full-chip ILT. Our approach leverages second-order curvature information via the Hessian matrix to achieve dramatically faster convergence and superior pattern fidelity compared to conventional first-order methods. Crucially, we address the prohibitive cost of exact Hessian computation by employing Hutchinson’s method to efficiently approximate the Hessian diagonal. Combined with exponential moving average (EMA) and gradient modulation techniques, our optimizer achieves significant performance gains. Experimental results demonstrate substantial improvements in both runtime efficiency (reduced iterations) and solution quality (enhanced pattern fidelity) compared to existing first-order ILT methods, paving the way for practical full-chip ILT.
Su Zheng, Ziyang Yu 0001, Bei Yu 0001, Martin D. F. Wong
DATE1
2026 Planning methods for accounting and taxation in new energy enterprises in the context of energy transition
Su Zheng
Serv. Oriented Comput. Appl.1
2026 HyDAS: Hybrid Domain Deformed Attention for Selective Hotspot Detection
abstract
Technology node scaling is challenged in many aspects, including pitch reduction, patterning flexibility, and lithography process variability during manufacturing. Without exception, layout hotspot detection, one of the critical steps to achieving design closure, also requires upgrading the associated techniques. With the rapid development of deep learning techniques, the detector exploiting convolutional neural network (CNN) is superior to ones based on pattern matching and classical machine learning algorithms. However, due to the local nature of CNN, the traditional CNN-based detector fails to model the relationship between the patterns in a large-sized layout, resulting in ignoring the impact of light propagation and some optical effects during photolithography. Even worse, another challenge arises from the fact that engineers cannot fully trust the results of learning model-based detectors, especially when handling some complicated layout patterns in practice. This makes it very difficult to deploy the detectors. Observing the facts, we propose a vision transformer (ViT) model-based layout hotspot detector with a deformed attention mechanism, where the training paradigm is inspired by the large pre-trained foundation model (e.g., OpenAI’s GPT-n series) and fine-tuning. Considering the light diffraction during photolithography, the hybrid domain (i.e., spatial and spectral domain) layout inputs via multi-channel are leveraged. Besides, our proposed detector integrates a selective option where the model can choose to do prediction or send to engineers based on the misclassification risk level. Experimental results on the ICCAD2012 metal layer benchmarks and ICCAD2020 via layer benchmarks demonstrate the effectiveness and efficiency of our approach. We have made the ICCAD2020 dataset publicly available to support further research in hotspot detection, enable benchmarking across different process nodes and layout types, and facilitate reproducibility in the field. The dataset is accessible at https: //github.com/shadowior/ICCAD2020.
Qi Sun 0002, Su Zheng, Xinyun Zhang 0001, Bei Yu 0001, Hao Geng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 Attention-Based EDA Tool Parameter Explorer: From Hybrid Parameters to Multi-QoR Metrics
Donger Luo, Qi Sun 0002, Peng Xu 0052, Su Zheng, Qi Xu 0004, Tinghuan Chen, Bei Yu 0001, Hao Geng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization
abstract
Electronic design automation (EDA) tools are critical in the very large scale integration (VLSI) flow. To address the challenges posed by the extensive search space and intricate feature interactions, statistical and machine-learning methods have been employed. These methods aim to model tool parameters and treat the tuning process as a regression task. However, these regression-based methods suffer from inaccurate estimations owing to limited training samples. To address this issue, we propose a ranking-based tool parameter tuning framework, called RankTuner, which directly learns the dominant relationship between parameters. RankTuner utilizes a pairwise Gaussian process to estimate the probability and uncertainty of the dominance relationship. Our approach also integrates a Duel-Thompson sampling method to balance exploration and exploitation in parameter selections. A dimensionality reduction scheme with random embedding and trust region techniques is incorporated to enable parallel searches. Experimental results demonstrate the superiority of RankTuner compared to the cutting-edge tool parameter tuning methods.
Peng Xu 0052, Su Zheng, Yuyang Ye 0001, Hao Geng, Tsung-Yi Ho, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Curvilinear Optical Proximity Correction via Cardinal Spline
abstract
This paper presents a novel curvilinear optical proximity correction (OPC) framework. The proposed approach involves representing mask patterns with control points, which are interconnected through cardinal splines. Mask optimization is achieved by iteratively adjusting these control points, guided by lithography simulation. To ensure compliance with mask rule checking (MRC) criteria, we develop comprehensive methods for checking width, space, area, and curvature. Additionally, to match the performance of inverse lithography techniques (ILT), we design algorithms to fit ILT results and resolve MRC violations. Extensive experiments demonstrate the effectiveness of our methodology, highlighting its potential as a viable OPC/ILT alternative.
Su Zheng, Xiaoxiao Liang, Ziyang Yu 0001, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
DAC1
2025 RuleLearner: OPC Rule Extraction From Inverse Lithography Technique Engine
abstract
Model-based optical proximity correction (OPC) with subresolution assist feature (SRAF) generation is a critical standard practice for compensating lithography distortions in the fabrication of integrated circuits at advanced technology nodes. Typical model-based OPC and SRAF algorithms involve the selection of user-controlled rule parameters. Conventionally, these rules are heuristically determined and applied globally throughout the correction regions, which can be time consuming and require expert knowledge of the tool. Additionally, the correlations of rule parameters to the objectives are highly nonlinear. All these factors make designing a high-performance OPC engine for complex metal designs a nontrivial task. This article proposes RuleLearner, a comprehensive mask optimization system designed for SRAF generation and model-based OPC in real industrial scenarios. The proposed framework learns from the guidance of an information-augmented inverse lithography technique engine, which, although expressive for complex designs, is expensive to generate refined masks for a whole set of design clips. Considering the nonlinearity and the tradeoff between local and global performance, the extracted rule value distributions are further optimized with customized natural gradients. The sophisticated SRAF generation, the edge segmentation and movements are then guided by the rule parameter. Experimental results show that RuleLearner can be applied across different complex design patterns and achieve the best lithographic performance and computational efficiency.
Ziyang Yu 0001, Su Zheng, Wenqian Zhao 0002, Xiaoxiao Liang, Guojin Chen, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Streamlining Computational Lithography With Efficient Pattern Database
abstract
In the pursuit of advancing computational lithography, this paper introduces a novel pattern database framework designed to support related tasks. The proposed framework is built upon three core components: an unsupervised metric learning method for robust pattern embedding, a vector database for swift pattern retrieval, and an efficient algorithm dedicated to pattern clustering. These elements synergize to significantly enhance the efficiency and effectiveness of various computational lithography methods. In downstream tasks, our framework provides accurate lithography hotspot detection through pattern retrieval, streamlines inverse lithography technique (ILT) by leveraging solution reusing, and facilitates the exploration of ILT & source parameters based on the pattern clustering results. Collectively, these advancements culminate in a comprehensive improvement in computational lithography, offering a scalable solution for the ever-evolving demands of this field.
Su Zheng, Wenqian Zhao 0002, Shuyuan Sun, Fan Yang 0001, Bei Yu 0001, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Lay-Net: Grafting Netlist Knowledge on Layout-Based Congestion Prediction
abstract
Congestion modeling is crucial for enhancing the routability of VLSI placement solutions. The underutilization of netlist information constrains the efficacy of existing layout-based congestion modeling techniques. We devise a novel approach that grafts netlist-based message passing into a layout-based model, thereby achieving a better knowledge fusion between layout and netlist to improve congestion prediction performance. The innovative heterogeneous message-passing paradigm more effectively incorporates routing demand into the model by considering connections between cells, overlaps of nets, and interactions between cells and nets. Leveraging multi-scale features, the proposed model effectively captures connection information across various ranges, addressing the issue of inadequate global information present in existing models. Using contrastive learning and mini-Gnet techniques allows the model to learn and represent features more effectively, boosting its capabilities and achieving superior performance. Extensive experiments demonstrate a notable performance enhancement of the proposed model compared to existing methods.Our code is available at: https://github.com/lanchengzou/congPred.
Lancheng Zou, Su Zheng, Peng Xu 0052, Siting Liu 0002, Bei Yu 0001, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Rank-DSE: Neural Pareto Comparator of Microarchitecture Design Space Exploration
abstract
The complexity of microarchitecture design has surged due to the expanding design space and time-intensive verification processes. Existing regression-based machine learning methods struggle with inaccurate estimations because of limited training samples. To address these challenges, we propose Rank-DSE, a novel framework for microarchitecture design space exploration (DSE) that leverages a Neural Pareto Comparator (NPC) to directly model the comparative relationships between different architecture designs. Rank-DSE bypasses the inaccuracies of absolute PPA (performance, power, area) predictions by focusing on relative comparisons. The NPC computes the probability of one architecture dominating another and employs semi-supervised learning to reduce the reliance on labeled data. Additionally, a reinforcement-learning-based sampling scheme with an updating baseline Pareto set accelerates the exploration process. Experimental results on the ICCAD 2021 benchmark demonstrate that Rank-DSE achieves superior search quality and cost-efficiency compared to state-of-the-art methods. Specifically, Rank-DSE improves hypervolume by up to 7% while reducing exploration cost by 53.09% compared to cutting-edge approaches. These results highlight the advantages of Rank-DSE in terms of efficiency and effectiveness for microarchitecture DSE.
Peng Xu 0052, Su Zheng, Mingzi Wang, Ziyang Yu 0001, Shixin Chen, Tinghuan Chen, Keren Zhu 0001, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.2
2025 Bridging Hotspot Detection and Mask Optimization via Domain-Crossing Masked Layout Modeling
abstract
With the rapid development of semiconductors, the size of transistors is continuously scaling down. The shrinking circuit size poses great challenges to optical proximity correction (OPC) and hotspot detection (HSD). Recent advancements in OPC and HSD commonly employ deep neural networks, achieving impressive performance within a limited runtime. Based on these achievements, we observe that deep-learning-based models of both HSD and OPC require knowledge of layout structure information. Furthermore, these two tasks are closely related to the lithography process during chip manufacturing. Observing such strong relationships, we propose that integrating OPC and HSD into a unified deep learning model will contribute to the performance of both tasks. To bridge the relationship between OPC and HSD, we first pre-train a layout understanding model built on the mask modeling technique, which effectively captures the layout geometric information, and then the pre-trained model can be easily fine-tuned on HSD and OPC with limited data. To fully pre-train the layout understanding model (LUM), we create a large layout dataset using layout generation techniques, solving the data-hungry issues. Experimental results show that the fine-tuned LUM model achieves remarkable performance on both OPC and HSD tasks.
Binwu Zhu, Su Zheng, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
ACM Trans. Design Autom. Electr. Syst.2
2024 SoC-Tuner: An Importance-guided Exploration Framework for DNN-targeting SoC Design
abstract
Designing a system-on-chip (SoC) for deep neural network (DNN) acceleration requires balancing multiple metrics such as latency, power, and area. However, most existing methods ignore the interactions among different SoC components and rely on inaccurate and error-prone evaluation tools, leading to inferior SoC design. In this paper, we present SoC-Tuner, a DNN-targeting exploration framework to find the Pareto optimal set of SoC configurations efficiently. Our framework constructs a thorough SoC design space of all components and divides the exploration into three phases. We propose an importance-based analysis to prune the design space, a sampling algorithm to select the most representative initialization points, and an information-guided multi-objective optimization method to balance multiple design metrics of SoC design. We validate our framework with the actual very-large-scale-integration (VLSI) flow on various DNN benchmarks and show that it outperforms previous methods. To the best of our knowledge, this is the first work to construct an exploration framework of SoCs for DNN acceleration.
Shixin Chen, Su Zheng, Wenqian Zhao 0002, Bei Yu 0001
ASPDAC2
2024 Fracturing-aware Curvilinear ILT via Circular E-beam Mask Writer
abstract
Inverse lithography technology (ILT) plays a crucial role in optical proximity correction, tending to generate curvilinear masks for optimal process windows. Traditional curvilinear mask manufacturing involves fracturing into rectangles, requiring expensive mask write times. A novel E-beam mask writer that writes variable radius circles per shot significantly reduces the shot count for curvilinear masks. To exploit this mask writer's benefits, we present two methods to generate circular fracturing-aware masks. The first one converts pixel-based masks from existing ILT methods into circle-based masks using predefined rules. The second one integrates circular constraints into the ILT process, generating circle-based masks directly via optimization. Extensive experimental results validate both approaches' effectiveness.
Xinyun Zhang 0001, Su Zheng, Guojin Chen, Binwu Zhu, Hong Xu 0001, Bei Yu 0001
DAC2
2024 EMOGen: Enhancing Mask Optimization via Pattern Generation
abstract
Layout pattern generation via deep generative models is a promising methodology for building practical large-scale pattern libraries. However, although improving optical proximity correction (OPC) is a major target of existing pattern generation methods, they are not explicitly trained for OPC and integrated into OPC methods. In this paper, we propose EMOGen to enable the co-evolution of layout pattern generation and learning-based OPC methods. With the novel co-evolution methodology, we achieve up to 39% enhancement in OPC and 34% improvement in pattern legalization.
Su Zheng, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
DAC1
2024 RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization
abstract
Electronic Design Automation (EDA) tools are critical in the Very Large Scale Integration (VLSI) flow. To address the challenges posed by the extensive search space and intricate feature interactions, statistical and machine-learning methods have been employed. These methods aim to model tool parameters and treat the tuning process as a regression task. However, these regression-based methods suffer from inaccurate estimations owing to limited training samples. To address this issue, we propose a ranking-based tool parameter tuning framework, called RankTuner, which directly learns the dominant relationship between parameters. RankTuner utilizes a pairwise Gaussian process to estimate the probability and uncertainty of the dominance relationship. Our approach also integrates a Duel-Thompson sampling method to balance exploration and exploitation in parameter selections. A dimensionality reduction scheme with random embedding and trust region techniques is incorporated to enable parallel searches. Experimental results demonstrate the superiority of RankTuner compared to the cutting-edge tool parameter tuning methods.
Peng Xu 0052, Su Zheng, Yuyang Ye 0001, Hao Geng, Tsung-Yi Ho, Bei Yu 0001
ICCAD2
2024 Improving Neural ODE Training with Temporal Adaptive Batch Normalization
abstract
Neural ordinary differential equations (Neural ODEs) is a family of continuous-depth neural networks where the evolution of hidden states is governed by learnable temporal derivatives. We identify a significant limitation in applying traditional Batch Normalization (BN) to Neural ODEs, due to a fundamental mismatch --- BN was initially designed for discrete neural networks with no temporal dimension, whereas Neural ODEs operate continuously over time. To bridge this gap, we introduce temporal adaptive Batch Normalization (TA-BN), a novel technique that acts as the continuous-time analog to traditional BN. Our empirical findings reveal that TA-BN enables the stacking of more layers within Neural ODEs, enhancing their performance. Moreover, when confined to a model architecture consisting of a single Neural ODE followed by a linear layer, TA-BN achieves 91.1\% test accuracy on CIFAR-10 with 2.2 million parameters, making it the first \texttt{unmixed} Neural ODE architecture to approach MobileNetV2-level parameter efficiency. Extensive numerical experiments on image classification and physical system modeling substantiate the superiority of TA-BN compared to baseline methods.
Su Zheng, Zhengqi Gao, Fan-Keng Sun, Duane S. Boning, Bei Yu 0001, Martin D. F. Wong
NeurIPS1
2024 Floorplet: Performance-Aware Floorplan Framework for Chiplet Integration
abstract
A chiplet is an integrated circuit (IC) that encompasses a well-defined subset of an overall systems functionality. In contrast to traditional monolithic system-on-chips (SoCs), chipletbased architecture can reduce costs and increase reusability, representing a promising avenue for continuing Moore’s Law. Despite the advantages of multi-chiplet architectures, floorplan design in a chiplet-based architecture has received limited attention. Conflicts between cost and performance necessitate a trade-off in chiplet floorplan design since additional latency introduced by advanced packaging can decrease performance. Consequently, balancing performance, cost, area, and reliability is of paramount importance. To address this challenge, we propose Floorplet (Floorplan chiplet), a framework comprising simulation tools for performance reporting and comprehensive models for cost and reliability optimization. Our framework employs the open-source Gem5 simulator to establish the relationship between performance and floorplan for the first time, guiding the floorplan optimization of multi-chiplet architecture. The experimental results show that our method decreases inter-chiplet communication costs by 24.81%.
Shixin Chen, Shanyi Li, Zhen Zhuang, Su Zheng, Zheng Liang 0003, Tsung-Yi Ho, Bei Yu 0001, Alberto L. Sangiovanni-Vincentelli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
abstract
The integration of a complex set of Electronic Design Automation (EDA) tools to enhance interoperability is a critical concern for circuit designers. Recent advancements in large language models (LLMs) have showcased their exceptional capabilities in natural language processing and comprehension, offering a novel approach to interfacing with EDA tools. This research paper introduces ChatEDA, an autonomous agent for EDA empowered by a large language model, AutoMage, complemented by EDA tools serving as executors. ChatEDA streamlines the design flow from the Register-Transfer Level (RTL) to the Graphic Data System Version II (GDSII) by effectively managing task decomposition, script generation, and task execution. Through comprehensive experimental evaluations, ChatEDA has demonstrated its proficiency in handling diverse requirements, and our fine-tuned AutoMage model has exhibited superior performance compared to GPT-4 and other similar LLMs.
Haoyuan Wu, Zhuolun He, Xinyun Zhang 0001, Xufeng Yao, Su Zheng, Haisheng Zheng, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 L2O-ILT: Learning to Optimize Inverse Lithography Techniques
abstract
Inverse lithography technique (ILT) is one of the most widely used resolution enhancement techniques (RETs) to compensate for the diffraction effect in the lithography process. However, ILT suffers from runtime overhead issues with the shrinking size of technology nodes. In this article, our proposed L2O-ILT framework unrolls the iterative ILT optimization algorithm into a learnable neural network with high interpretability, which can generate a high-quality initial mask for fast refinement. Experimental results demonstrate that our method achieves better performance on both mask printability and runtime than the previous methods.
Binwu Zhu, Su Zheng, Ziyang Yu 0001, Guojin Chen, Yuzhe Ma, Fan Yang 0001, Bei Yu 0001, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 HierCGRA: A Novel Framework for Large-scale CGRA with Hierarchical Modeling and Automated Design Space Exploration
abstract
Coarse-grained reconfigurable arrays (CGRAs) are promising design choices in computation-intensive domains, since they can strike a balance between energy efficiency and flexibility. A typical CGRA comprises processing elements (PEs) that can execute operations in applications and interconnections between them. Nevertheless, most CGRAs suffer from the ineffectiveness of supporting flexible architecture design and solving large-scale mapping problems. To address these challenges, we introduce HierCGRA, a novel framework that integrates hierarchical CGRA modeling, Chisel-based Verilog generation, LLVM-based data flow graph (DFG) generation, DFG mapping, and design space exploration (DSE). With the graph homomorphism (GH) mapping algorithm, HierCGRA achieves a faster mapping speed and higher PE utilization rate compared with the existing state-of-the-art CGRA frameworks. The proposed hierarchical mapping strategy achieves 41× speedup on average compared with the ILP mapping algorithm in CGRA-ME. Furthermore, the automated DSE based on Bayesian optimization achieves a significant performance improvement by the heterogeneity of PEs and interconnections. With these features, HierCGRA enables the agile development for large-scale CGRA and accelerates the process of finding a better CGRA architecture.
Sichao Chen, Su Zheng, Guowei Zhu, Jingyuan Li 0003, Yazhou Yan, Yuan Dai, Wenbo Yin, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.3
2023 Mitigating Distribution Shift for Congestion Optimization in Global Placement
abstract
The placement and routing (PnR) flow plays a critical role in physical design. Poor routing congestion is a possible problem causing severe routing detours, which can lead to deteriorated timing performance or even routing failure. Deep-learning-based congestion prediction model is designed to guide the global placement process in previous work. However, the distribution shift problem in this method limits its performance. In this paper, we mitigate the distribution shift problem with a look-ahead mechanism inspired by optical flow prediction and an invariant feature space learning technique. With the proposed method, we can achieve better congestion prediction performance and less-congested placement results.
Su Zheng, Lancheng Zou, Siting Liu 0002, Yibo Lin, Bei Yu 0001, Martin D. F. Wong
DAC1
2023 Lay-Net: Grafting Netlist Knowledge on Layout-Based Congestion Prediction
abstract
Congestion modeling is a key point for improving the routability of VLSI placement solutions. The underuti-lization of netlist information limits the performance of ex-isting layout-based congestion modeling methods. Combining the knowledge from netlist and layout, we graft netlist-based message passing on a layout-based model to achieve better congestion prediction performance. The novel heterogeneous message-passing paradigm better embeds the routing demand into the model by considering both connections between cells and overlaps of nets. With the help of multi-scale features, the proposed model can effectively capture connection information across different ranges, overcoming the problem of insufficient global information in existing models. Based on the advancements, the proposed model achieves significant improvement compared with existing methods.
Su Zheng, Lancheng Zou, Peng Xu 0052, Siting Liu 0002, Bei Yu 0001, Martin D. F. Wong
ICCAD1
2023 LithoBench: Benchmarking AI Computational Lithography for Semiconductor Manufacturing
abstract
Computational lithography provides algorithmic and mathematical support for resolution enhancement in optical lithography, which is the critical step in semiconductor manufacturing. The time-consuming lithography simulation and mask optimization processes limit the practical application of inverse lithography technology (ILT), a promising solution to the challenges of advanced-node lithography. Although various machine learning methods for ILT have shown promise for reducing the computational burden, this field is in lack of a dataset that can train the models thoroughly and evaluate the performance comprehensively. To boost the development of AI-driven computational lithography, we present the LithoBench dataset, a collection of circuit layout tiles for deep-learning-based lithography simulation and mask optimization. LithoBench consists of more than 120k tiles that are cropped from real circuit designs or synthesized according to the layout topologies of famous ILT testcases. The ground truths are generated by a famous lithography model in academia and an advanced ILT method. Based on the data, we provide a framework to design and evaluate deep neural networks (DNNs) with the data. The framework is used to benchmark state-of-the-art models on lithography simulation and mask optimization. We hope LithoBench can promote the research and development of computational lithography. LithoBench is available at https://anonymous.4open.science/r/lithobench-APPL.
Su Zheng, Binwu Zhu, Bei Yu 0001, Martin D. F. Wong
NeurIPS1
2023 Boosting VLSI Design Flow Parameter Tuning with Random Embedding and Multi-objective Trust-region Bayesian Optimization
abstract
Modern very large-scale integration (VLSI) design requires the implementation of integrated circuits using electronic design automation (EDA) tools. Due to the complexity of EDA algorithms, there are numerous tool parameters that have imperative impacts on the chip design quality. Manual selection of parameter values is excessively laborious and constrained by experts’ experience. Due to the high complexity and lack of parallelization, most existing parameter tuning methods cannot make sufficient exploration in a large search space. In this article, we boost the efficiency and performance of parameter tuning with random embedding and multi-objective trust-region Bayesian optimization. Random embedding can effectively cut down the number of variables in the search process and thus reduce the runtime of Bayesian optimization. Multi-objective trust-region Bayesian optimization allows the algorithm to explore diverse solutions with excellent parallelism. Due to the ability to do more exploration in limited runtime, the proposed framework can achieve better performance than existing methods in our experiments.
Su Zheng, Hao Geng, Bei Yu 0001, Martin D. F. Wong
ACM Trans. Design Autom. Electr. Syst.1
2022 GRAEBO: FPGA General Routing Architecture Exploration via Bayesian Optimization
abstract
Modern FPGAs utilize complex routing architectures to optimize the area, critical path delay, and power consumption. General Routing Block (GRB) models the routing resources of modern FPGAs, enabling the design of better routing architectures than previous academic FPGAs based on the CB-SB model. However, the design space of the GRB model is too large to be explored manually. In this paper, we propose GRAEBO, a design space exploration (DSE) algorithm for FPGA routing architectures based on Bayesian optimization, which can optimize and accelerate the DSE by balancing exploration and exploitation. Moreover, we design pruning rules to further improve the DSE efficiency, which can serve as a multi-fidelity acceleration method. GRAEBO obtains better area, delay, and area-delay product than a 142-channel baseline CB-SB architecture, with improvements of 8%, 19%, and 26%, respectively. Compared to the GRB architecture found by the simulated annealing algorithm, GRAEBO achieves 9% smaller area, 5% shorter delay, and 13% better area-delay product on the VTR benchmarks.
Su Zheng, Jiadong Qian, Hao Zhou 0008, Lingli Wang
FPL1
2022 Low Error-Rate Approximate Multiplier Design for DNNs with Hardware-Driven Co-Optimization
abstract
In this paper, two approximate 3 × 3 multipliers are proposed and the synthesis results of the ASAP-7nm process library justify that they can reduce the area by 31.38% and 36.17%, and the power consumption by 36.73% and 35.66% compared with the exact multiplier, respectively. They can be aggregated with a 2 × 2 multiplier to produce an 8 × 8 multiplier with low error-rate based on the distribution of DNN weights. We propose a hardware-driven software co-optimization method to improve the DNN accuracy by retraining. Based on the proposed two approximate 3-bit multipliers, three approximate 8-bit multipliers with low error-rate are designed for DNNs. Compared with the exact 8-bit unsigned multiplier, our design can achieve a significant advantage over other approximate multipliers on the public dataset.
Jide Zhang, Su Zheng, Zhen Li 0059, Lingli Wang
ISCAS3
2022 HEAM: High-Efficiency Approximate Multiplier optimization for Deep Neural Networks
abstract
We propose an optimization method for the automatic design of approximate multipliers, which minimizes the average error according to the operand distributions. Our multiplier achieves up to 50.24% higher accuracy than the best reproduced approximate multiplier in DNNs, with 15.76% smaller area, 25.05% less power consumption, and 3.50% shorter delay. Compared with an exact multiplier, our multiplier reduces the area, power consumption, and delay by 44.94%, 47.63%, and 16.78%, respectively, with negligible accuracy losses. The tested DNN accelerator modules with our multiplier obtain up to 18.70% smaller area and 9.99% less power consumption than the original modules.
Su Zheng, Zhen Li 0059, Jingbo Gao, Jide Zhang, Lingli Wang
ISCAS1
2022 Adaptable Approximate Multiplier Design Based on Input Distribution and Polarity
abstract
Approximate computing is an efficient approach to reduce the design complexity for error-resilient applications. Multipliers are key arithmetic units in many applications, such as deep neural networks (DNNs) and digital signal processing (DSP) systems. In this article, an open-source adaptable approximate multiplier design driven by input distribution and polarity is proposed to generate optimized approximate multipliers to trade off between the application-level performance and the hardware cost. The proposed method minimizes the average square of the absolute error of an approximate multiplier according to the probability distributions of operands extracted from the target application with consideration of input polarity, achieving low hardware cost and negligible application-level performance loss. The proposed method can generate unsigned multipliers (or signed multipliers) based on the Braun multiplier (or Baugh–Wooley multiplier). To demonstrate the effectiveness of the method, three different-scale quantized DNNs, including LeNet, AlexNet, and VGG16 with 8$\times $8 unsigned multiplication and an adaptive least mean square (LMS)-based finite impulse response (FIR) filter with 16$\times $16 fixed-point signed multiplication, are evaluated. In the DNN training process, a noise training technique is adopted to reduce the accuracy loss due to the approximation. When compared to the state-of-the-art approximate multipliers, the generated multipliers can achieve up to 26.4% and 27.1% product of power, delay, and area gains with negligible application-level performance loss in VGG16 and FIR applications, respectively.
Zhen Li 0059, Su Zheng, Jide Zhang, Jingbo Gao, Jun Tao 0001, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.2
2021 FastCGRA: A Modeling, Evaluation, and Exploration Platform for Large-Scale Coarse-Grained Reconfigurable Arrays
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) provide sufficient flexibility in domain-specific applications with high hardware efficiency, which make CGRAs suitable for fast-evolving fields such as neural network acceleration and edge computing. To meet the requirement of the fast evolution, we propose FastCGRA, the modeling, mapping, and exploration platform for large-scale CGRAs. FastCGRA supports hierarchical architecture description and automatic switch module generation. Connectivity-aware packing and graph partition algorithms are designed to reduce the complexity of placement and routing. The graph homomorphism placement algorithm in FastCGRA enables efficient placement on large-scale CGRAs. The packing and placement algorithms cooperate with a negotiation-based routing algorithm to form an integral mapping procedure. FastCGRA can support the modeling and mapping of large-scale CGRAs with significantly higher placement and routing efficiency than existing platforms. The automatic switch module generation method can reduce the complexity of CGRA interconnection design. With these features, FastCGRA can boost the exploration of large-scale CGRAs.
Su Zheng, Kaisen Zhang, Yaoguang Tian, Wenbo Yin, Lingli Wang, Xuegong Zhou
FPT1
2019 Targeted Black-Box Adversarial Attack Method for Image Classification Models
abstract
Deep neural networks (DNNs) are widely applied to image classification tasks. Due to the fact that these models are usually vulnerable, subtle perturbations of pixels may lead to classification errors, which poses a serious threat to the success of DNN applications. Moreover, perturbations of pixels can also corrupt other pattern recognition models such as Naive Bayes (NB), Decision Tree (DT) and Random Forest (RF). In this paper, a general method is proposed to carry out targeted black-box attacks for image classification models. The proposed method can achieve targeted fool rates (TFRs) of 0.873 and 0.781 on CIFAR-10 dataset with and without the access to the training set of the target model respectively. For cross-model attacks, the proposed method can still achieve a TFR of 0.630 on CIFAR- 10. Furthermore, the proposed method is able to mount attacks for up to 100 classes on CIFAR-100 dataset with a TFR of 0.721, successfully handling 99 cases for each class. In our experiments, the proposed method shows higher performance and higher reliability than other black-box attack methods, with 0.123 greater maximum TFR and 0.602 greater minimum TFR than previous methods UPSET and ANGRI on CIFAR-10 in attacks trained on a single model.
Su Zheng, Lingli Wang
IJCNN1
2007 Surprising creativity: a cognitive framework for interactive exhibits designed for children
abstract
Interactive exhibits in museums are providing exciting and dynamic learning experiences with significant potential to stimulate children's creativity. However, current sophisticated interfaces designed to deliver easily accessible information are not teaching the fundamental skills necessarily to foster genuine creative outcomes. The aim of our research is to promote a design methodology that fosters children's creativity, helping them to gain the formative skills necessary to nurture the process of creative learning. There needs to be more encouragement to motivate children's curiosity and the promotion of observational skills that can help them realise the creative possibilities to be derived from everyday experiences. This paper describes the development of the Creative Surprise Model (CSM): a cognitive framework that informs a methodology to support interactive design practitioners. It identifies the motivational link between surprise emotion and the generation of creativity. We demonstrate how it is applied by describing a real life design task.
Su Zheng, Adrian Bromage, Martin Adam, Stephen A. R. Scrivener
Creativity & Cognition1