EDBT 2026 Demo / reviewers in the wild / expert
Wei W. Xing
dblp:284/1640
· DBLP profile ↗
46ranked-venue papers
7as first author
42since 2021 · last 2026
0000-0002-3177-8478ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 6 first-author · 35 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM InferenceabstractWe introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where experts use models of varying capacities based on task complexity. The framework features three key innovations: (1) sparse expert activation through unified gating networks, (2) specialized expert trees that leverage domain-specific expertise while optimizing the quality-efficiency trade-off, and (3) collaborative debate mechanisms for conflicting solutions. Additionally, MoT includes a shared baseline tree with early stopping—activated experts perform lightweight validation and terminate early when confidence is high. Experiments across five benchmarks (GSM8K, MATH, AIME 2024, MMLU, HotpotQA) show that MoT achieves 2-7 percentage point accuracy improvements while reducing LLM calls by 37-40% compared to existing multi-path methods. Yangbo Wei, Zhen Huang 0007, Shaoqiang Lu, Junhong Qian, Dongge Qin, Ting-Jung Lin, Wei W. Xing, Lei He 0001 |
AAAI | 7 |
| 2026 | VFlow: Discovering Optimal Agentic Workflows for Verilog Generation
Yangbo Wei, Zhen Huang 0007, Lei He 0001, Ting-Jung Lin, Wei W. Xing |
ASP-DAC | 6 |
| 2026 | ConvGA: Convolution Network-Guided Genetic Algorithm for Optimal PDN Decoupling DesignabstractPower supply noise has emerged as a critical bottleneck in modern integrated circuit design, where increasing current densities and higher operating frequencies pose significant challenges to system reliability. While decoupling capacitors (decaps) serve as the primary solution for suppressing power delivery network (PDN) noise, determining their optimal values and placement remains computationally prohibitive using traditional methods. This aritcle introduces ConvGA, a novel framework that seamlessly integrates convolutional neural networks (CNN) with genetic algorithms to revolutionize PDN decap optimization. At the heart of ConvGA is a specialized CNN architecture trained on comprehensive boundary element method (BEM) simulations, enabling ultra-fast impedance prediction for arbitrary PCB configurations. Our CNN achieves remarkable accuracy while reducing impedance computation time from hours to mere milliseconds—a 500× speedup over conventional BEM calculations. This acceleration enables the genetic algorithm to efficiently explore vast design spaces through adaptive population control and dynamic constraint mechanisms, systematically minimizing both the number of required capacitors and the deviation from target impedance. Extensive experiments on industrial-scale PDNs demonstrate that ConvGA achieves a 15× reduction in optimization time while requiring 30% fewer capacitors compared to state-of-the-art methods, consistently producing high-quality solutions across diverse PDN configurations. Yuchuan Lin, Ning Xu 0006, Wei W. Xing, Yuanqing Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | ASAP: Accelerating Corner-Based Timing Analysis With Bayesian Active Self-Attention Neural ProcessabstractWith the advancement of modern nanoscale technology nodes, Static Timing Analysis (STA) has become an indispensable technique for ensuring circuit reliability and performance across diverse process conditions. However, traditional STA methods scale poorly to the explosion of process corners in the nanoscale fabrication technology. Despite some seminal works in using AI to accelerate such processes, they either lack reliability or stability. To this end, we introduce ASAP, a novel approach addressing this challenge by combining both the latest deep learning methods and the classical Bayesian models to deliver scalable and accurate predictions with a self-calibration strategy to ensure reliability. Technically, the ASAP novelly integrates self-attention to help identify and prioritize crucial features under various input conditions and employs Neural Process to make confidence-based predictions for the final timing results. Furthermore, ASAP is equipped with Active Learning for self-refinement and self-correction. Experimental evaluations on benchmark circuits demonstrate that our method surpasses state-of-the-art work in STA accuracy by 18% in terms of prediction accuracy. Longze Wang, Wei W. Xing, Zhelong Wang, Christos P. Sotiriou, Nikolaos Sketopoulos, Ning Xu 0006, Yuanqing Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | NeuralMesh: Neural Network For FEM Mesh Generation in 2.5D/3D Chiplet Thermal SimulationabstractAdvanced integrated circuit (IC) systems increasingly utilize chiplet-based packaging with complex $2.5 \mathrm{D} / 3 \mathrm{D}$ structures and dense Through-Silicon Via (TSV) arrays. While the Finite Element Method (FEM) provides high-fidelity thermal simulation for these systems, its computational efficiency degrades significantly when generating and optimizing meshes for intricate geometries. To address these performance limitations while preserving simulation accuracy, we present NeuralMesh, a novel framework that accelerates thermal analysis of chiplet-based ICs. Our approach integrates deep learning and geometric analysis to optimize mesh generation without the need for iterative refinement steps. NeuralMesh first employs an enhanced segmentation model to predict thermal distributions based on geometric, material, and power parameters. These predictions, combined with key geometric features, guide the optimization of an initial coarse FEM mesh. By eliminating traditional iterative mesh refinement, our framework achieves up to $45.00 \times$ mesh generation speedup while maintaining thermal accuracy within 0.8% of commercial COMSOL simulations. It reduces the number of mesh elements in unimportant areas, which represents a speed improvement of the subsequent thermal simulation. This advancement enables rapid yet precise thermal analysis essential for modern IC package design. Pengju Chen, Dan Niu, Dekang Zhang, Depeng Xie, Zhou Jin 0001, Wei W. Xing, Lei He 0001 |
DAC | 7 |
| 2025 | Self-Attention to Operator Learning-based 3D-IC Thermal SimulationabstractThermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDESolving based methods, while accurate, are too slow for iterative design. Machine learning approaches like FNO provide faster alternatives but suffer from high-frequency information loss and high-fidelity data dependency. We introduce Self-Attention UNet Fourier Neural Operator (SAU-FNO), a novel framework combining self-attention and U-Net with FNO to capture longrange dependencies and model local high-frequency features effectively. Transfer learning is employed to fine-tune low-fidelity data, minimizing the need for extensive high-fidelity datasets and speeding up training. Experiments demonstrate that SAUFNO achieves state-of-the-art thermal prediction accuracy and provides an $842 \times$ speedup over traditional FEM methods, making it an efficient tool for advanced 3D IC thermal simulations. Zhen Huang 0007, Wenkai Yang, Muxi Tang, Depeng Xie, Ting-Jung Lin, Yu Zhang 0086, Wei W. Xing, Lei He 0001 |
DAC | 8 |
| 2025 | Accuracy Is Not Always We Need: Precision-Aware Bayesian Yield OptimizationabstractIntegrated circuit yield optimization plays a vital role in ensuring reliable semiconductor manufacturing, directly impacting both product quality and production costs. Current approaches to yield optimization face two fundamental challenges that limit their practical effectiveness. First, yield estimation requires intensive computational resources. Second, traditional black-box optimization methods inefficiently allocate these resources across design candidates. Most existing approaches compound these issues by performing detailed yield estimations uniformly across all candidates, regardless of their potential quality. To address these limitations, we introduce a novel precision-aware yield optimization framework that intelligently adapts computational resource allocation based on each design candidate’s predicted performance. Our approach moves beyond simple simulation counting by incorporating a Figure of Merit (FoM) as a continuous quality metric. By combining a Continuous AutoRegression model to characterize the relationship between true yield and precision levels with a sophisticated multi-fidelity acquisition strategy, our framework achieves optimal resource distribution. Experimental validation on four industry-standard benchmark circuits demonstrates that our method converges with fewer than 1,000 simulations, reducing simulation costs by over $10 \times$ while achieving better final designs and robustness than state-of-the-art high-fidelity approaches. Jing Kou, Zidong Chen, Haiyan Qin, Wang Kang 0001, Wei W. Xing |
DAC | 6 |
| 2025 | Multi-Agent Yield Analysis For Circuit DesignabstractSemiconductor yield estimation presents a critical challenge in modern manufacturing, directly impacting production costs and market competitiveness. Traditional estimation methods, particularly Monte Carlo simulation, while reliable, become computationally prohibitive for complex modern circuits. Contemporary approaches, including importance sampling and machine learning techniques, face fundamental limitations in consistency across circuit topologies and practical validation. This work introduces YieldAgent, a novel Large Language Model (LLM)-powered framework that revolutionizes yield estimation through dynamic integration of multiple analytical strategies. YieldAgent employs a three-layer agent architecture to analyze circuit characteristics and historical data, optimizing estimation methods while balancing computational efficiency and precision. The framework incorporates Retrieval-Augmented Generation for domain knowledge integration and Tree-structured Parzen Estimators for dynamic hyperparameter optimization. Experimental validation across 12nm and 40nm technology nodes demonstrates that YieldAgent reduces computational overhead by up to $2.9 \times$ while maintaining or exceeding state-of-the-art accuracy. The system’s ability to adapt across different circuit topologies and technology nodes establishes a new paradigm for scalable, intelligent yield estimation in electronic design automation. Haiyan Qin, Jing Kou, Wang Kang 0001, Wei W. Xing |
DAC | 5 |
| 2025 | FUSIS: Fusing Surrogate Models and Importance Sampling for Efficient Yield EstimationabstractAs process nodes continue to shrink, yield estimation has become increasingly critical in modern circuit design. Traditional approaches face significant challenges: surrogate-based methods often struggle with robustness and accuracy, whereas importance sampling (IS)-based methods suffer from high simulation costs. To address these challenges simultaneously, we propose FUSIS, a unified framework that combines the strengths of surrogate-based and IS-based approaches. Unlike conventional surrogate-based methods that directly replace SPICE simulations for performance predictions, FUSIS employs a Deep Kernel support vector machine (SVM) as an approximation of the indicator function, which is further utilized to construct a quasi-optimal proposal distribution for IS to accelerate convergence. To further mitigate yield estimation bias caused by surrogate inaccuracies, we introduce a novel correction factor to adjust the IS-based yield estimation. Experiments conducted on SRAM and analog circuits demonstrate that FUSIS significantly improves accuracy by up to 24.84% (8.67% on average) while achieving up to 29.54× (10.30× on average) speedup in efficiency compared to seven state-of-the-art methods. Wei W. Xing |
DATE | 2 |
| 2025 | LaRED: Efficient IR Drop Predictor with Layout-Preserving Rebuilder-Encoder-Decoder ArchitectureabstractIn the realm of integrated circuit verification, IR drop analysis plays a crucial role. Recent advancements in machine learning (ML) significantly enhance its efficiency, yet many current approaches fail to fully leverage the input structure of feature maps and the transmission mechanism of Power Delivery Network (PDN) layouts. To bridge these gaps, we introduce Layout-Preserving Rebuilder-Encoder-Decoder Architecture Predictor (LaRED), which employs a novel Rebuilder-Encoder-Decoder (RED) architecture and utilizes an innovative downsampling approach and upsampling framework to optimize its perception of instances and the transmission of features. LaRED captures information from various regions with asymmetric topological structure while preserving and transferring layout characteristics through deformable convolution, hybrid downsampling, cascaded upsampling, and attentional feature fusion. The rebuilder rebuilds raw input, whereas the encoder ensures comprehensive feature transmission across all instances. The decoder then facilitates seamless transfer of feature information across layers. This approach enables LaRED to integrate chip features of varying topologies and scales, enhancing its representational power. Compared to the current State-Of-The-Art (SOTA), MAUnet, LaRED achieves accuracy improvements of 34.6% to 42.6% in benchmark tests, establishing it as the new standard in static IR drop analysis for integrated circuit design with ML techniques. The code is available at https://github.com/Todi85/LaRED. Chengxuan Yu, Yanshuang Teng, Wenhao Dai, Yongjiang Li, Wei W. Xing, Dan Niu, Zhou Jin 0001 |
DATE | 5 |
| 2025 | DIVE: Dynamic Information-Guided Variable Expansion for Deeper Analog Circuit OptimizationabstractIn analog circuit design, transistor sizing remains a critical challenge due to high-dimensional parameter spaces and expensive simulations. While Bayesian optimization shows promise, existing methods struggle with the "curse of dimensionality." Inspired by expert designers’ workflow of focusing on key parameters first before gradually optimizing secondary parameters, we introduce DIVE (Dynamic Information-guided Variable Expansion), a framework that reformulates parameter optimization as an information efficiency maximization problem. Unlike approaches that explore all parameters simultaneously, DIVE progressively includes design variables with the highest information content. Our framework introduces three innovations: (1) constraint-aware weighted mutual information analysis that evaluates parameters’ contributions to objectives and constraints; (2) an adaptive variable inclusion mechanism that determines when to expand the optimization space; and (3) a mutual information-guided kernel learning strategy for Gaussian process models. Evaluations across multiple analog circuits demonstrate that DIVE achieves a 1.61×-21.11× reduction in required simulations while delivering up to 2.69× spec improvements over the state-of-the-art methods. By progressing from simple to complex parameter spaces based on information theory, DIVE establishes a new paradigm for reaching deeper optima by mimicking experienced designers’ design philosophy in circuit optimization. Our code is available1. Zhuohua Liu, Weilun Xie, Yuanqi Hu, Wei W. Xing |
ICCAD | 6 |
| 2025 | ASTRA: Automatic Sizing of Transistors with Reasoning AgentsabstractAdvancing technology nodes have significantly increased the complexity of transistor sizing in analog circuit design. Although artificial intelligence (AI) techniques show potential, their lack of integrated domain expertise often leads to slow convergence in practical applications. We propose ASTRA (Automatic Sizing of Transistors with Reasoning Agents), a novel optimization framework that implements the Model Context Protocol (MCP) to create structured reasoning pathways between Large Language Models (LLMs), domain knowledge bases, and Bayesian Optimization (BO). ASTRA introduces a two-stage process: first, MCP-guided design initialization that leverages Retrieval-Augmented Generation (RAG) to quickly identify feasible regions using gm/ID methodology; and second, BO-based optimization focused on critical transistors, identified through LLM reasoning with data-driven validation. A key innovation of ASTRA is its ability to seamlessly integrate with and enhance virtually any existing transistor sizing algorithm at minimal additional cost. Unlike purely data-driven or black-box LLM approaches, ASTRA maintains traceable decision processes that can be verified and refined. Evaluated on three real-world analog circuits, ASTRA enhances multiple classical optimization methods, achieving up to 4.35× fewer simulation iterations and 2.36× performance improvements, demonstrating its effectiveness as a general open-source framework for advancing analog circuit sizing.1 Wei W. Xing, Baowen Ou, Zhuohua Liu, Yuanqi Hu |
ICCAD | 1 |
| 2025 | SetupKit: Efficient Multi-Corner Setup/Hold Time Characterization Using Bias-Enhanced Interpolation and Active LearningabstractAccurate setup/hold time characterization is crucial for modern chip timing closure, but its reliance on potentially millions of SPICE simulations across diverse process-voltage-temperature (PVT) corners creates a major bottleneck, often lasting weeks or months. Existing methods suffer from slow search convergence and inefficient exploration, especially in the multi-corner setting. We introduce SetupKit, a novel framework designed to break this bottleneck using statistical intelligence, circuit analysis and active learning (AL). SetupKit integrates three key innovations: BEIRA, a bias-enhanced interpolation search derived from statistical error modeling to accelerate convergence by overcoming stagnation issues, initial search interval estimation by circuit analysis and AL strategy using Gaussian Process. This AL component intelligently learns PVT-timing correlations, actively guiding the expensive simulations to the most informative corners, thus minimizing redundancy in multi-corner characterization. Evaluated on industrial 22nm standard cells across 16 PVT corners, SetupKit demonstrates a significant 2.4× overall CPU time reduction (from 720 to 290 days on a single core) compared to standard practices, drastically cutting characterization time. SetupKit offers a principled, learning-based approach to library characterization, addressing a critical EDA challenge and paving the way for more intelligent simulation management. Junzhuo Zhou, Haoxuan Xia, Yuxin Yan, Chengyu Zhu, Ting-Jung Lin, Wei W. Xing, Lei He 0001 |
ICCAD | 7 |
| 2025 | OpenYield: An Open-Source SRAM Yield Analysis and Optimization Benchmark SuiteabstractStatic Random-Access Memory (SRAM) yield analysis is essential for semiconductor innovation, yet research progress faces a critical challenge: the large gap between simplified academic models and the complexities observed in practice. The lack of open, higher-fidelity benchmarks has hindered reproducibility and transferability, as promising academic techniques often fail to carry over to more realistic settings. We present OpenYield, an open-source ecosystem that aims to narrow this gap through three contributions: (i) An SRAM circuit generator that explicitly incorporates second-order effects (interconnect/line parasitics, inter-cell leakage coupling, and peripheralcircuit variations) that are commonly omitted in academic studies. (ii) A standardized evaluation platform with a simple interface and baseline yield-analysis implementations to enable fair comparisons and reproducible research on these higherfidelity circuits. (iii) An optimization platform for transistor-level sizing under these models, supporting reproducible studies of robustness/efficiency trade-offs. OpenYield aims to foster more reproducible and transferable progress in SRAM-yield research. The framework is publicly available at OpenYield:URL. Shan Shen, Xingyang Li, Zhuohua Liu, Junhao Ma, Yiheng Wu, Yuquan Sun, Wei W. Xing |
ICCD | 8 |
| 2025 | LVFGen: Efficient Liberty Variation Format (LVF) Generation Using Variational Analysis and Active LearningabstractAs transistor dimensions shrink, process variations significantly impact circuit performance, signifying the need for accurate statistical circuit analysis. In digital circuit timing analysis, the Liberty Variation Format (LVF) has emerged as an industrial leading representation of timing distributions in cell libraries at 22 nm and below. However, LVF characterization relies on the Monte Carlo (MC) method, which requires excessive SPICE simulations of cells with process variations. Similar challenges also exist for uncertainty propagation and quantification in chip manufacturing and the broader scientific communities. To resolve this foundational challenge, this paper presents LVFGen, a novel method that reduces the simulation costs of MC while generate high-accuracy LVF library. LVFGen utilizes an active learning strategy based on variational analysis to identify process variation samples that impact timing distributions more significantly. Compared to the state-of-the-art Quasi-MC method, LVFGen demonstrates an overall 2.27× speedup in LVF library generation within an accuracy level of 5k-sample MC and a 4.06× speedup within a 100k-sample MC accuracy. Junzhuo Zhou, Haoxuan Xia, Wei W. Xing, Ting-Jung Lin, Lei He 0001 |
ISPD | 3 |
| 2025 | ML-PTA: A Two-Stage ML-Enhanced Framework for Accelerating Nonlinear DC Circuit Simulation With Pseudo-Transient AnalysisabstractDirect current (DC) analysis lies at the heart of integrated circuit design in seeking DC operating points. Although pseudo-transient analysis (PTA) methods have been widely used in DC analysis in both industry and academia, their initial parameters and stepping strategy require expert knowledge and labor tuning to deliver efficient performance, which hinders their further applications. In this paper, we leverage the latest advancements in machine learning to deploy PTA with more efficient setups for different problems. More specifically, active learning, which automatically draws knowledge from other circuits, is used to provide suitable initial parameters for PTA solver, and then calibrate on-the-fly to further accelerate the simulation process using TD3-based reinforcement learning (RL). To expedite model convergence, we introduce dual agents and a public sampling buffer in our RL method to enhance sample utilization. To further improve the learning efficiency of the RL agent, we incorporate imitation learning to improve reward function and introduce supervised learning to provide a better dual-agent rotation strategy. We make the proposed algorithm a general out-of-the-box SPICE-like solver and assess it on a variety of circuits, demonstrating up to 3.10× reduction in NR iterations for the initial stage and 285.71× for the RL stage. Zhou Jin 0001, Wenhao Li 0017, Haojie Pei, Xiaru Zha, Yichao Dong, Xiang Jin, Dan Niu, Wei W. Xing |
IEEE Trans. Computers | 9 |
| 2025 | Carbon Nanotube Interconnect Optimizations With Bayesian Neural Network and Bayesian OptimizationabstractAs Cu interconnects near their physical limits with continued technology scaling, carbon nanotube (CNT) interconnects have emerged as a promising alternative due to their excellent conductivity. However, fabrication immaturity introduces significant process variations, causing discrepancies between ideal and actual performance. This article presents a novel approach to optimize CNT interconnects considering process variations. We first develop a parameterized CNT interconnect model that accounts for process variations. Using this model, a Bayesian Neural Network (BNN) is proposed to predict performance distributions by leveraging its inherent uncertainty. We then introduce a Bayesian optimization framework that uses the BNN’s posterior to jointly optimize interconnect parameters and buffer insertion, targeting Area-Delay Product (ADPIn this work, we use area-delay product ziegler2001optimal as the performance metric, though other metrics can also be applied within our framework.) with process variations. Experimental results demonstrate the effectiveness of our approach. The proposed BNN model achieves over 95% prediction accuracy for interconnects performance distributions. Compared to existing methods, our method achieves an average ADP improvement of 20.3% over the state-of-the-art methods and 13% over the standard Monte Carlo method. Compared with Monte Carlo method, our method also achieves an average 8.8x acceleration. Moreover, the optimized CNT interconnects show an average improvement of 82.8% in ADP and 68.7% in delay compared to Cu interconnects. This work offers an effective method for optimizing CNT interconnects under process variations and highlights their potential as a viable alternative to Cu interconnects in future integrated circuits. Zhelong Wang, Wei W. Xing, Ning Xu 0006, Yuanqing Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | ModelGen: Automating Semiconductor Parameter Extraction with Large Language Model AgentsabstractDevice models require large numbers of parameters to characterize complex physical effects. Although the latest advancements in machine learning and automated tools have drastically improved efficiency over the classic methods, they still demand a considerable amount of human intervention in the loop to gain accuracy. This drastically limits further automation. Inspired by the success of Multimodal Large Language Models (MLLMs) in addressing tasks across diverse fields, we propose ModelGen, the first in-depth study to leverage MLLMs with RAG (Retrieval-Augmented Generation) to significantly reduce human effort in parameter extraction for compact model. Our contributions include (1) Automated Agentic Workflow Construction that learns to build and refine extraction workflows through iterative optimization, (2) MLLM Judge, a visual scoring mechanism that evaluates fitting quality using actual device characteristic plots rather than simple numerical metrics, and (3) Model-specific RAG for providing relevant domain knowledge during the extraction process. Experimental results demonstrate that ModelGen achieves a 26.8%–33.1% improvement in pass@1,3,5 compared to base LLM methods. The system completes complex model extractions for BSIMs and ASM-HEMT in hours (up to 168× faster) rather than days or weeks, making parameter extraction more accessible to non-experts while maintaining professional engineer-level accuracy. Yangbo Wei, Zhanfei Chen, Jinlong Yan, Ting-Jung Lin, Zhen Huang 0007, Wei W. Xing, Lei He 0001 |
ACM Trans. Design Autom. Electr. Syst. | 9 |
| 2024 | Multi-Resolution Active Learning of Fourier Neural OperatorsabstractFourier Neural Operator (FNO) is a popular operator learning framework. It not only achieves the state-of-the-art performance in many tasks, but also is efficient in training and prediction. However, collecting training data for the FNO can be a costly bottleneck in practice, because it often demands expensive physical simulations. To overcome this problem, we propose Multi-Resolution Active Learning of FNO (MRA-FNO), which can dynamically select the input functions and resolutions to lower the data cost as much as possible while optimizing the learning efficiency. Specifically, we propose a probabilistic multi-resolution FNO and use ensemble Monte-Carlo to develop an effective posterior inference algorithm. To conduct active learning, we maximize a utility-cost ratio as the acquisition function to acquire new examples and resolutions at each step. We use moment matching and the matrix determinant lemma to enable tractable, efficient utility computation. Furthermore, we develop a cost annealing framework to avoid over-penalizing high-resolution queries at the early stage. The over-penalization is severe when the cost difference is significant between the resolutions, which renders active learning often stuck at low-resolution queries and inferior performance. Our method overcomes this problem and applies to general multi-fidelity active learning and optimization problems. We have shown the advantage of our method in several benchmark operator learning tasks. The code is available at https://github.com/shib0li/MRA-FNO. Xin Yu 0003, Wei W. Xing, Robert M. Kirby, Akil Narayan 0001, Shandian Zhe |
AISTATS | 3 |
| 2024 | Equation Discovery with Bayesian Spike-and-Slab Priors and Efficient KernelsabstractDiscovering governing equations from data is important to many scientific and engineering applications. Despite promising successes, existing methods are still challenged by data sparsity and noise issues, both of which are ubiquitous in practice. Moreover, state-of-the-art methods lack uncertainty quantification and/or are costly in training. To overcome these limitations, we propose a novel equation discovery method based on Kernel learning and BAyesian Spike-and-Slab priors (KBASS). We use kernel regression to estimate the target function, which is flexible, expressive, and more robust to data sparsity and noises. We combine it with a Bayesian spike-and-slab prior — an ideal Bayesian sparse distribution — for effective operator selection and uncertainty quantification. We develop an expectation-propagation expectation-maximization (EP-EM) algorithm for efficient posterior inference and function estimation. To overcome the computational challenge of kernel regression, we place the function values on a mesh and induce a Kronecker product construction, and we use tensor algebra to enable efficient computation and optimization. We show the advantages of KBASS on a list of benchmark ODE and PDE discovery tasks. The code is available at \url{https://github.com/long-da/KBASS}. Da Long, Wei W. Xing, Aditi S. Krishnapriyan, Robert M. Kirby, Shandian Zhe, Michael W. Mahoney |
AISTATS | 2 |
| 2024 | CIS: Conditional Importance Sampling for Yield Optimization of Analog and SRAM CircuitsabstractYield optimization is one of the central challenges in submicrometer integrated circuit manufacture. Classic yield optimization methods rely on importance sampling (IS) to provide efficient and robust yield estimation for each individual design. Despite its success, such an approach is still computationally expensive due to the large number of calculations for many different designs. To resolve this challenge, we propose conditional importance sampling (CIS) that can approximate the optimal proposal distribution for any given design by leveraging the power of the modern deep-learning-based sampling method, conditional normalizing flow. More importantly, CIS generalizes well to unseen design and thus can deliver effective yield optimization with a small number of expensive simulations. To conduct yield optimization efficiently with consideration of creditable uncertainty, we propose a novel Important Sampling Bayesian optimization (ISBO) using a deep-warped gradient-boosting regression (GBR). The proposed method is extensively evaluated against five state-of-the-art baselines; the results show that the proposed method delivers superior performance: a speedup of $1.10 \times-l0.46\times$ ($4.45 \times$ on average) with even higher yield designs, an improvement of $1.1 \times-10 \times$ ($4.44 \times$ on average) in consideration of the Optimality-Cost Ratio, and most importantly, excellent robustness and consistency in all our extensive experiments on analog and SRAM circuits. Wei W. Xing |
ASPDAC | 2 |
| 2024 | BoCNT: A Bayesian Optimization Framework for Global CNT Interconnect OptimizationabstractAs the prevailing copper interconnect technology advances to its fundamental physical limit, interconnect delay due to ever-increasing wire resistivity shows a significant impact on circuit performance. Bundled single-wall carbon nanotubes (SWCNTs) interconnects have emerged as a promising candidate technique to replace copper interconnects thanks to their superior conductivity and immunity to electromigration. To deliver satisfying performance within low power consumption, the CNT interconnect timing is optimized by adjusting either the interconnect geometry, e.g., CNT diameter and nanotube pitch, or the buffer insertion. These two operations are normally optimized separately, which leads to an inferior design that is not global optimum. To resolve this problem, we first propose a model that parameterizes SWCNT interconnects. We then leverage the known Bayesian optimization to optimize SWCNT global interconnect and buffer insertion simultaneously to promote interconnect performance. The proposed method is assessed based on a set of interconnect benchmarks at 22nm technology node. Compared to the state-of-the-art methods, the Bayesian co-optimization technique can reduce more than 17% power delay product (PDP). Additionally, we evaluate the SWCNT interconnect performance at 32nm, 22nm and 16nm technology nodes. Compared to the SOTA method with the same wire dimension, SWCNT interconnects optimized by the proposed method can further reduce delay and PDP relative to copper by 35% and 45% on average, which highlights the promising prospect of the SWCNT interconnect technology and the effectiveness of our proposed technique. Ning Xu 0006, Wei W. Xing, Yuanqing Cheng |
ASPDAC | 3 |
| 2024 | Unleashing the Potential of AQFP Logic Placement via Entanglement Entropy and ProjectionabstractAdiabatic quantum-flux-parametron (AQFP) logic, known for its energy efficiency, has emerged as a prominent superconductor-based logic family, surpassing traditional rapid single flux quantum (RSFQ) logic. In AQFP circuits, each cell operates on AC power, serving as both a power supply and clock signal to drive data flow across clock phases. However, signal attenuation with increasing wirelength may result in more potential data errors. To address this, rows of buffers are inserted as repeaters to ensure data synchronization and avoid wirelength violations. However, these inserted buffer rows in the AQFP placement significantly amplifies power consumption and circuit delay. To address these challenges, in this paper, we propose an innovative and analytical method for the placement of AQFP. The proposed method aims at minimizing the need for additional buffers. The framework incorporates two key features: (1) entanglement entropy for topology initialization and (2) projection for placement and buffering. These features offer advantages such as avoiding intensive computations, including fix-order Lagrangian optimization in large-scale scenarios, while significantly reducing the required number of buffer rows. The experimental results validate the efficiency of the proposed framework, demonstrating an average reduction of 81% in the required number of buffers and acceleration of 1.88x in the processing time compared with the state-of-the-art method. Yinuo Bai 0002, Enxin Yi, Wei W. Xing, Bei Yu 0001, Zhou Jin 0001 |
DAC | 3 |
| 2024 | MAUnet: Multiscale Attention U-Net for Effective IR Drop PredictionabstractThe efficient analysis of power grids is a crucial yet computationally challenging task in integrated circuit (IC) design, given the shrinking power supply voltage of ultra deep-submicron VLSI design. Different from the conventional modified nodal analysis technique, this paper introduces MAUnet, an innovative machine-learning model that redefines state-of-the-art full-chip static IR drop prediction. MAUnet ingeniously integrates multi-scale convolutional blocks, attention mechanisms, and U-Net architecture to optimize prediction accuracy. The multi-scale convolutional blocks significantly enhance feature extraction from image-based data, while the attention mechanism precisely identifies hotspot regions. The U-Net architecture, on the other hand, enables scalable image-to-image prediction applicable to circuits of any size. Uniquely, MAUnet also incorporates a pioneering fusion method that synergies both power grids and image-based data. Additionally, we introduce a low-rank approximation transfer learning technique to extend MAUnet's applicability to unseen test cases. Benchmark tests validate MAUnet's superior performance, achieving an average error of less than 6% relative to the average IR drop on three benchmarks. The performance enhancements offered by our proposed method are substantial, outperforming the current state-of-the-art method, IREDGe, by considerable margins of 29%, 65%, and 68% in three canonical benchmarks. Transfer learning is validated to enable model to achieve effective improvement on real circuit test cases. Compared to commercial tools, which often require hours to deliver results, the proposed method provides orders of magnitude speed-up with negligible error in practice. Yuanqing Cheng, Yage Lin, Kelin Peng, Shunchuan Yang, Zhou Jin 0001, Wei W. Xing |
DAC | 7 |
| 2024 | KATO: Knowledge Alignment And Transfer for Transistor Sizing Of Different Design and TechnologyabstractAutomatic transistor sizing in circuit design continues to be a formidable challenge. Despite that Bayesian optimization (BO) has achieved significant success, it is circuit-specific, limiting the accumulation and transfer of design knowledge for broader applications. This paper proposes (1) efficient automatic kernel construction, (2) the first transfer learning across different circuits and technology nodes for BO, and (3) a selective transfer learning scheme to ensure only useful knowledge is utilized. These three novel components are integrated into BO with Multi-objective Acquisition Ensemble (MACE) to form Knowledge Alignment and Transfer Optimization (KATO) to deliver state-of-the-art performance: up to 2x simulation reduction and 1.2x design improvement over the baselines. Wei W. Xing, Weijian Fan, Zhuohua Liu, Yuanqi Hu |
DAC | 1 |
| 2024 | Every Failure Is A Lesson: Utilizing All Failure Samples To Deliver Tuning-Free Efficient Yield EvaluationabstractYield estimation and optimization have become increasingly important for circuit design as technology nodes scale down. Simple yet well-established minimal norm importance sampling (MNIS) still serves as an industrial standard due to its robustness and reliability. In this study, we generalize the classic MNIS and propose Every Failure Is A Lesson (EFIAL) to utilize every failure sample (instead of one in MNIS) to construct the proposal distribution. EFIAL is completely tuning-free and the update computation complexity is only O(M) (M is the number of failure samples) by utilizing the blessing of dimensionality. The idea of EFIAL is then extended to the state-of-the-art (SOTA) pre-sampling method, onion sampling, to significantly boost efficiency, by up to 9.08x (4.68x on average). Extensive evaluations against SOTA yield estimation methods reveal that EFIAL achieves a speedup of up to 13.54x (5.16x on average) and an accuracy improvement of up to 24.91%. Wei W. Xing, Weijian Fan, Lei He 0001 |
DAC | 1 |
| 2024 | LVF2: A Statistical Timing Model based on Gaussian Mixture for Yield Estimation and Speed BinningabstractAs transistor size continues to scale down, process variation has become an essential factor determining semiconductor yield and economic return. The Liberty Variation Format (LVF) is the current industrial standard that expresses statistical timing behaviors based on single Gaussian model. However, it loses accuracy when the timing distribution is non-Gaussian due to growing process variations. This paper proposes a novel LVF2 distribution model that combines two weighted skewed-normal (SN) distributions, which better captures the multi-Gaussian timing distribution while maintaining backward compatibility with LVF. Experiments using TSMC 22nm standard cells show that, compared to LVF, LVF2 reduces binning error by 7.74X in delay and 9.56X in transition time, and reduces 3σ-yield error by 4.79X and 7.18X in delay and transition time, respectively. The error reduction for path delay is diminished due to Central Limit Theorem (CLT). But it is still 2X for a typical circuit path with 8 Fanout-of-4 (FO4) inverter delays. Junzhuo Zhou, Haoxuan Xia, Leilei Jin, Xiao Shi 0001, Wei W. Xing, Ting-Jung Lin, Lei He 0001 |
DAC | 7 |
| 2024 | Beyond the Yield Barrier: Variational Importance Sampling Yield AnalysisabstractOptimal mean shift vector (OMSV)-based importance sampling methods have long been prevalent in yield estimation and optimization as an industry standard. However, most OMSV-based methods are designed heuristically without a rigorous understanding of their limitations. To this end, we propose VIS, the first variational analysis framework for yield problems, enabling a systematic refinement for OMSV. For instance, VIS reveals that the classic OMSV is suboptimal, and the optimal/true OMSV should always stay beyond the failure boundary, which enables a free improvement for all OMSV-based methods immediately. Using VIS, we show a progressive refinement for the classic OMSV including incorporation of full covariance in closed form, adjusting for asymmetric failure distributions, and capturing multiple failure regions, each of which contributes to a progressive improvement of more than 2×. Inheriting the simplicity of OMSV, the proposed method retains simplicity and robustness yet achieves up to 29.03× speedup over the state-of-the-art (SOTA) methods. We also demonstrate how the SOTA yield optimization, ASAIS, can immediately benefit from our True OMSV, delivering a 1.20× and 1.27× improvement in performance and efficiency, respectively, without additional computational overhead. Lei He 0001, Wei W. Xing |
ICCAD | 3 |
| 2024 | Pseudo Adjoint Optimization: Harnessing the Solution Curve for SPICE AccelerationabstractPseudo transient analysis (PTA) has been a promising solution for direct current (DC) analysis of transistor-level circuit simulation. Despite its popularity, PTA requires meticulous hyperparameter tuning for optimal performance. In this paper, we propose pseudo adjoint optimization, Soda-PTA, which models the PTA solution curve (which is used to measure convergence) using a neural ordinary differential equation (Neural ODE) and deriving explicit gradients of the Newton-Raphson (NR) iteration w.r.t. the PTA hyperparameters through the classic adjoint method, enabling effective optimization of the PTA hyperparameters. To generalize Soda-PTA for unseen circuits, we further introduce a graph convolution network to transfer optimal PTA hyperparameters from the other circuits to the target one. Soda-PTA is implemented in an out-of-the-box SPICE simulator. Through extensive experiments, Soda-PTA demonstrates superior acceleration performance: an average speedup of 1.53x over the state-of-the-art BoA-PTA while ensuring superior convergence and up to 22.12x speedup compared to the native PTA solver. Jiatai Sun, Xiaru Zha, Chao Wang 0120, Dan Niu, Wei W. Xing, Zhou Jin 0001 |
ICCAD | 6 |
| 2024 | An End-to-End In-Memory Computing System Based on a 40-nm eFlash-Based IMC SoC: Circuits, Toolchains, and Systems Co-Design FrameworkabstractDespite its promising potential for Artificial Intelligence (AI) applications, current In-Memory Computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance including (a) an 8-bit hardware-friendly Quantization-Aware Training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data, (b) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips, and (c) an efficient mapping strategy based on the Integer Linear Programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40nm eFlash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 hours for voice recognition, a 21.53% improvement for the Perceptual Evaluation of Speech Quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection. Tianshuo Bai, Wanru Mao, Guangyao Wang, Aifei Zhang, Shihang Fu, Shuaikai Liu, Jianchao Hu, Xitong Yang, Biao Pan, Wei W. Xing, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2024 | Multicorner Timing Analysis Acceleration for Iterative Physical Design of ICsabstractWe propose a multi-corner multi-stage timing analysis prediction framework using a generalized linear model with latent features. We then further improve such methods using kernel trick extension, transfer learning with knowledge from previous designs, and multi-output feature engineering to deliver state-of-the-art (SOTA) prediction accuracy with very limited training data. Most importantly, our method is equipped with a Bayesian decision strategy to deliver reliable predictions with accuracy close to 100%, pushing the frontier of the machine-learning-based STA for practical implementation in the industry environment, where reliability is highly desired. Experimental results show that the accuracy of our proposed method outperforms the SOTA competitors by up to 4x and can improve prediction accuracy to 100% with little extra STA executions. Wei W. Xing, Longze Wang, Zhelong Wang, Zhaoyu Shi, Ning Xu 0006, Yuanqing Cheng, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | High-Dimensional Yield Estimation Using Shrinkage Deep Features and Maximization of Integral Entropy ReductionabstractDespite the fast advances in high-sigma yield analysis with the help of machine learning techniques in the past decade, one of the main challenges, the curse of "dimensionality", which is inevitable when dealing with modern large-scale circuits, remains unsolved. To resolve this challenge, we propose an absolute shrinkage deep kernel learning, ASDK, which automatically identifies the dominant process variation parameters in a nonlinear-correlated deep kernel and acts as a surrogate model to emulate the expensive SPICE simulation. To further improve the yield estimation efficiency, we propose a novel maximization of approximated entropy reduction for an efficient model update, which is also enhanced with parallel batch sampling for parallel computing, making it ready for practical deployment. Experiments on SRAM column circuits demonstrate the superiority of ASDK over the state-of-the-art (SOTA) approaches in terms of accuracy and efficiency with up to 11.1x speedup over SOTA methods. Guohao Dai 0002, Wei W. Xing |
ASP-DAC | 3 |
| 2023 | Seeking the Yield Barrier: High-Dimensional SRAM Evaluation Through Optimal ManifoldabstractBeing able to efficiently obtain an accurate estimate of the failure probability of SRAM components has become a central issue as model circuits shrink their scale to submicrometer with advanced technology nodes. In this work, we revisit the classic norm minimization method. We then generalize it with infinite components and derive the novel optimal manifold concept, which bridges the surrogate-based and importance sampling (IS) yield estimation methods. We then derive a sub-optimal manifold, optimal hypersphere, which leads to an efficient sampling method being aware of the failure boundary called onion sampling. Finally, we use a neural coupling flow (which learns from samples like a surrogate model) as the IS proposal distribution. These combinations give rise to a novel yield estimation method, named Optimal Manifold Important Sampling (OPTIMIS), which keeps the advantages of the surrogate and IS methods to deliver state-of-the-art performance with robustness and consistency, with up to 3.5x in efficiency and 3x in accuracy over the best of SOTA methods in High-dimensional SRAM evaluation. Guohao Dai 0002, Wei W. Xing |
DAC | 3 |
| 2023 | TOTAL: Multi-Corners Timing Optimization Based on Transfer and Active LearningabstractIn modern advanced integrated circuit design, a design normally needs to be progressively optimized until the static timing analysis (STA) of full process corners meets the timing constraints. To improve efficiency, using machine learning to predict the path timings directly in order to reduce the extensive time-consuming SPICE simulations has become a promising technique to approach fast design closure. However, current methods lack both flexibility and reliability to be used in a practical industrial environment. To resolve these challenges, we propose TOTAL, which is constructed using a generalized linear model with latent features to effectively capture knowledge transferred from previous designs and delivers state-of-the-art (SOTA) prediction accuracy that is up to 6.6x improvement over the competitors in terms of mean absolute error (MAE). Most importantly, TOTAL is equipped with a Bayesian decision strategy to actively update uncertain predictions and deliver reliable predictions with accuracy close to 100%, pushing the frontier of the machine-learning-based STA for practical implementation. Wei W. Xing, Rongqi Lu, Zhelong Wang, Ning Xu 0006, Yuanqing Cheng, Weisheng Zhao 0001 |
DAC | 1 |
| 2023 | OPT: Optimal Proposal Transfer for Efficient Yield Optimization for Analog and SRAM CircuitsabstractYield optimization is one of the central challenges in submicrometer integrated circuit manufacture. However, yield optimization is computationally expensive due to intensive yield estimation and intractable optimization processes. In this work, we first reinvent the state-of-the-art all sensitivity adversarial importance sampling (ASAIS) yield optimization from a Laplace approximation perspective, which also reveals its limitations and suggests improvements. We then generalize it with infinite components and discover the key ingredient in yield optimization to be an effective proposal distribution transfer (OPT) procedure, which is captured using conditional normalizing flow (CNF). To deliver a reliable yield optimization pipeline that accounts for the uncertainty due to the lack of data, we propose sequential ensemble, the first empirical uncertainty estimation that enables tractable Bayesian yield optimization without introducing an extra surrogate for the first time. We conduct extensive experiments against five state-of-the-art baselines and show that the proposed method delivers superior performance: a speedup of 1.01x-11.94x (5.57x on average) with higher yield designs, and most importantly, excellent robustness and consistency in all our experiments on analog and SRAM circuits. Guohao Dai 0002, Yuanqing Cheng, Wang Kang 0001, Wei W. Xing |
ICCAD | 5 |
| 2023 | BoA-PTA: A Bayesian Optimization Accelerated PTA Solver for SPICE SimulationabstractOne of the greatest challenges in integrated circuit design is the repeated executions of computationally expensive SPICE simulations, particularly when highly complex chip testing/verification is involved. Recently, pseudo-transient analysis (PTA) has shown to be one of the most promising continuation SPICE solvers. However, the PTA efficiency is highly influenced by the inserted pseudo-parameters. In this work, we proposed BoA-PTA, a Bayesian optimization accelerated PTA that can substantially accelerate simulations and improve convergence performance without introducing extra errors. Furthermore, our method does not require any pre-computation data or offline training. The acceleration framework can either speed up ongoing, repeated simulations (e.g., Monte-Carlo simulations) immediately or improve new simulations of completely different circuits. BoA-PTA is equipped with cutting-edge machine learning techniques, such as deep learning, Gaussian process, Bayesian optimization, non-stationary monotonic transformation, and variational inference via reparameterization. We assess BoA-PTA in 43 benchmark circuits and real industrial circuits against other SOTA methods and demonstrate an average of 1.5x (maximum 3.5x) for the benchmark circuits and up to 250x speedup for the industrial circuit designs over the original CEPTA without sacrificing any accuracy. Wei W. Xing, Xiang Jin, Tian Feng 0002, Dan Niu, Weisheng Zhao 0001, Zhou Jin 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | Physics Informed Deep Kernel LearningabstractDeep kernel learning is a promising combination of deep neural networks and nonparametric function estimation. However, as a data driven approach, the performance of deep kernel learning can still be restricted by scarce or insufficient data, especially in extrapolation tasks. To address these limitations, we propose Physics Informed Deep Kernel Learning (PI-DKL) that exploits physics knowledge represented by differential equations with latent sources. Specifically, we use the posterior function sample of the Gaussian process as the surrogate for the solution of the differential equation, and construct a generative component to integrate the equation in a principled Bayesian hybrid framework. For efficient and effective inference, we marginalize out the latent variables in the joint probability and derive a collapsed model evidence lower bound (ELBO), based on which we develop a stochastic model estimation algorithm. Our ELBO can be viewed as a nice, interpretable posterior regularization objective. On synthetic datasets and real-world applications, we show the advantage of our approach in both prediction accuracy and uncertainty quantification. The code is available at https://github.com/GregDobby/PIDKL. Zheng Wang 0042, Wei W. Xing, Robert M. Kirby, Shandian Zhe |
AISTATS | 2 |
| 2022 | Accelerating nonlinear DC circuit simulation with reinforcement learningabstractDC analysis is the foundation for nonlinear electronic circuit simulation. Pseudo transient analysis (PTA) methods have gained great success among various continuation algorithms. However, PTA tends to be computationally intensive without careful tuning of parameters and proper stepping strategies. In this paper, we harness the latest advancing in machine learning to resolve these challenges simultaneously. Particularly, an active learning is leveraged to provide a fine initial solver environment, in which a TD3-based Reinforcement Learning (RL) is implemented to accelerate the simulation on the fly. The RL agent is strengthen with dual agents, priority sampling, and cooperative learning to enhance its robustness and convergence. The proposed algorithms are implemented in an out-of-the-box SPICElike simulator, which demonstrated a significant speedup: up to 3.1X for the initial stage and 234X for the RL stage. Zhou Jin 0001, Haojie Pei, Yichao Dong, Xiang Jin, Wei W. Xing, Dan Niu |
DAC | 6 |
| 2022 | Efficient bayesian yield analysis and optimization with active learningabstractYield optimization for circuit design is computationally intensive due to the expensive yield estimation based on Monte Carlo methods and the difficult optimization process. In this work, a uniform framework to solve these problems simultaneously is proposed. Firstly, a novel efficient Bayesian yield analysis framework, BYA, is proposed by deriving a Bayesian estimation for the yield and introducing active learning based on reductions of integral entropy. A tractable convolutional entropy infill technique is then proposed to efficiently solve the entropy reduction problem. Lastly, we extend BYA for yield optimization by transforming knowledge across the design space and variational space. Experimental results based on SRAM and adder circuits show that BYA is 410x faster (in terms of the number of simulations) than standard MC and averagely 10x (up to 10000x) more accurate than the state-of-the-art method for yield estimation, and is about 5x faster than the SOTA yield optimization methods. Xiang Jin, Linxu Shi, Wang Kang 0001, Wei W. Xing |
DAC | 5 |
| 2022 | E-LMC: Extended Linear Model of Coregionalization for Spatial Field PredictionabstractPhysical simulations based on partial differential equations typically generate spatial fields results, which are utilized to calculate specific properties of a system for engineering design and optimization. Due to the intensive computational burden of the simulations, a surrogate model mapping the low-dimensional inputs to the spatial fields are commonly built based on a relatively small dataset. To resolve the challenge of predicting the whole spatial field, the popular linear model of coregional-ization (LMC) can disentangle complicated correlations within the high-dimensional spatial field outputs and deliver accurate predictions. However, LMC fails if the spatial field cannot be well approximated by a linear combination of base functions with latent processes. In this paper, we present the Extended Linear Model of Coregionalization (E-LMC) by introducing an invertible neural network to linearize the highly complex and nonlinear spatial fields so that the LMC can easily generalize to nonlinear problems while preserving the traceability and scalability. Several real-world applications demonstrate that E-LMC can exploit spatial correlations effectively, showing a maximum improvement of about 40% over the original LMC and outperforming the other state-of-the-art spatial field models. Shihong Wang, Yichen Meng, Wei W. Xing |
IJCNN | 4 |
| 2022 | GAR: Generalized Autoregression for Multi-Fidelity FusionabstractIn many scientific research and engineering applications, where repeated simulations of complex systems are conducted, a surrogate is commonly adopted to quickly estimate the whole system. To reduce the expensive cost of generating training examples, it has become a promising approach to combine the results of low-fidelity (fast but inaccurate) and high-fidelity (slow but accurate) simulations. Despite the fast developments of multi-fidelity fusion techniques, most existing methods require particular data structures and do not scale well to high-dimensional output. To resolve these issues, we generalize the classic autoregression (AR), which is wildly used due to its simplicity, robustness, accuracy, and tractability, and propose generalized autoregression (GAR) using tensor formulation and latent features. GAR can deal with arbitrary dimensional outputs and arbitrary multifidelity data structure to satisfy the demand of multi-fidelity fusion for complex problems; it admits a fully tractable likelihood and posterior requiring no approximate inference and scales well to high-dimensional problems. Furthermore, we prove the autokrigeability theorem based on GAR in the multi-fidelity case and develop CIGAR, a simplified GAR with the same predictive mean accuracy but requires significantly less computation. In experiments of canonical PDEs and scientific computational examples, the proposed method consistently outperforms the SOTA methods with a large margin (up to 6x improvement in RMSE) with only a few high-fidelity training samples. Wei W. Xing |
NeurIPS | 3 |
| 2021 | Multi-Fidelity High-Order Gaussian Processes for Physical SimulationabstractThe key task of physical simulation is to solve partial differential equations (PDEs) on discretized domains, which is known to be costly. In particular, high-fidelity solutions are much more expensive than low-fidelity ones. To reduce the cost, we consider novel Gaussian process (GP) models that leverage simulation examples of different fidelities to predict high-dimensional PDE solution outputs. Existing GP methods are either not scalable to high-dimensional outputs or lack effective strategies to integrate multi-fidelity examples. To address these issues, we propose Multi-Fidelity High-Order Gaussian Process (MFHoGP) that can capture complex correlations both between the outputs and between the fidelities to enhance solution estimation, and scale to large numbers of outputs. Based on a novel nonlinear coregionalization model, MFHoGP propagates bases throughout fidelities to fuse information, and places a deep matrix GP prior over the basis weights to capture the (nonlinear) relationships across the fidelities. To improve inference efficiency and quality, we use bases decomposition to largely reduce the model parameters, and layer-wise matrix Gaussian posteriors to capture the posterior dependency and to simplify the computation. Our stochastic variational learning algorithm successfully handles millions of outputs without extra sparse approximations. We show the advantages of our method in several typical applications. Zheng Wang 0042, Wei W. Xing, Robert M. Kirby, Shandian Zhe |
AISTATS | 2 |
| 2020 | Infinite ShapeOdds: Nonparametric Bayesian Models for Shape RepresentationsabstractLearning compact representations for shapes (binary images) is important for many applications. Although neural network models are very powerful, they usually involve many parameters, require substantial tuning efforts and easily overfit small datasets, which are common in shape-related applications. The state-of-the-art approach, ShapeOdds, as a latent Gaussian model, can effectively prevent overfitting and is more robust. Nonetheless, it relies on a linear projection assumption and is incapable of capturing intrinsic nonlinear shape variations, hence may leading to inferior representations and structure discovery. To address these issues, we propose Infinite ShapeOdds (InfShapeOdds), a Bayesian nonparametric shape model, which is flexible enough to capture complex shape variations and discover hidden cluster structures, while still avoiding overfitting. Specifically, we use matrix Gaussian priors, nonlinear feature mappings and the kernel trick to generalize ShapeOdds to a shape-variate Gaussian process model, which can grasp various nonlinear correlations among the pixels within and across (different) shapes. To further discover the hidden structures in data, we place a Dirichlet process mixture (DPM) prior over the representations to jointly infer the cluster number and memberships. Finally, we exploit the Kronecker-product structure in our model to develop an efficient, truncated variational expectation-maximization algorithm for model estimation. On synthetic and real-world data, we show the advantage of our method in both representation learning and latent structure discovery. Wei W. Xing, Shireen Y. Elhabian, Robert M. Kirby, Ross T. Whitaker, Shandian Zhe |
AAAI | 1 |
| 2020 | Scalable Gaussian Process Regression NetworksabstractGaussian process regression networks (GPRN) are powerful Bayesian models for multi-output regression, but their inference is intractable. To address this issue, existing methods use a fully factorized structure (or a mixture of such structures) over all the outputs and latent functions for posterior approximation, which, however, can miss the strong posterior dependencies among the latent variables and hurt the inference quality. In addition, the updates of the variational parameters are inefficient and can be prohibitively expensive for a large number of outputs. To overcome these limitations, we propose a scalable variational inference algorithm for GPRN, which not only captures the abundant posterior dependencies but also is much more efficient for massive outputs. We tensorize the output space and introduce tensor/matrix-normal variational posteriors to capture the posterior correlations and to reduce the parameters. We jointly optimize all the parameters and exploit the inherent Kronecker product structure in the variational model evidence lower bound to accelerate the computation. We demonstrate the advantages of our method in several real-world applications. Wei W. Xing, Robert M. Kirby, Shandian Zhe |
IJCAI | 2 |
| 2020 | Multi-Fidelity Bayesian Optimization via Deep Neural NetworksabstractBayesian optimization (BO) is a popular framework for optimizing black-box functions. In many applications, the objective function can be evaluated at multiple fidelities to enable a trade-off between the cost and accuracy. To reduce the optimization cost, many multi-fidelity BO methods have been proposed. Despite their success, these methods either ignore or over-simplify the strong, complex correlations across the fidelities. While the acquisition function is therefore easy and convenient to calculate, these methods can be inefficient in estimating the objective function. To address this issue, we propose Deep Neural Network Multi-Fidelity Bayesian Optimization (DNN-MFBO) that can flexibly capture all kinds of complicated relationships between the fidelities to improve the objective function estimation and hence the optimization performance. We use sequential, fidelity-wise Gauss-Hermite quadrature and moment-matching to compute a mutual information-based acquisition function in a tractable and highly efficient way. We show the advantages of our method in both synthetic benchmark datasets and real-world applications in engineering design. Wei W. Xing, Robert M. Kirby, Shandian Zhe |
NeurIPS | 2 |
| 2019 | Scalable High-Order Gaussian Process RegressionabstractWhile most Gaussian processes (GP) work focus on learning single-output functions, many applications, such as physical simulations and gene expressions prediction, require estimations of functions with many outputs. The number of outputs can be much larger than or comparable to the size of training samples. Existing multi-output GP models either are limited to low-dimensional outputs and restricted kernel choices, or assume oversimplified low-rank structures within the outputs. To address these issues, we propose HOGPR, a High-Order Gaussian Process Regression model, which can flexibly capture complex correlations among the outputs and scale up to a large number of outputs. Specifically, we tensorize the high-dimensional outputs, introducing latent coordinate features to index each tensor element (i.e., output) and to capture their correlations. We then generalize a multilinear model to a hybrid of a GP and latent GP model. The model is endowed with a Kronecker product structure over the inputs and the latent features. Using the Kronecker product properties and tensor algebra, we are able to perform exact inference over millions of outputs. We show the advantage of the proposed model on several real-world applications. Shandian Zhe, Wei W. Xing, Robert M. Kirby |
AISTATS | 2 |