Zhengqi Gao

dblp:256/9403 · DBLP profile ↗
← Back
27ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-1515-4198ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 10 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SPIPE: Differentiable SPICE-Level Co-Simulation Program for Integrated Photonics and Electronics
abstract
Heterogeneous photonic-electronic systems, such as co-packaged optics and photonic-electronic artificial intelligence (AI) accelerators, are rapidly gaining traction but also pose significant design challenges due to distinct design methodologies. Digital and analog electronics are typically described using hardware description languages and SPICE, respectively, whereas photonic devices and systems are represented using permittivity tensors on the Yee grid and the Scattering matrix formulation. This disparity necessitates an end-to-end photonic-electronic cosimulation tool to streamline co-design. Most preliminary cosimulation approaches rely on translating photonic compact models into Verilog-A or SPICE models to simulate everything there, which not only introduces the additional complexity of model conversion but also has potential numerical stability problems. Additionally, another critical functionality missing from the current implementation is enabling gradient calculation in these co-simulators, which will be crucial for end-to-end gradient-based electronic-photonic system optimization. To address these challenges, we introduce SPIPE, a differentiable SPICE-level co-simulation framework for integrated photonic-electronic systems. SPIPE is the first co-simulator to overcome model conversion issues and to provide differentiability. Numerical experiments on several circuits confirm the accuracy of SPIPE when compared to analytical solutions and real-world experimental data. Furthermore, in cases where existing simulators are applicable, SPIPE achieves a runtime reduction of 2!+85!A compared to an industry-standard simulator. SPIPE features an integrated simulation interface with a low usage barrier, opening avenues for more accessible and effective photonic-electronic co-design. SPIPE is open sourced: https://github.com/zhengqigao/spipe.
Zhengqi Gao, Jiaqi Gu 0002, Luca Daniel, Ronald A. Rohrer, Duane S. Boning
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 BOSON-1: Understanding and Enabling Physically-Robust Photonic Inverse Design with Adaptive Variation-Aware Subspace Optimization
abstract
Nanophotonic device design aims to optimize pho-tonic structures to meet specific requirements across various applications. Inverse design has unlocked non-intuitive, high-dimensional design spaces, enabling the discovery of compact, high-performance device topologies beyond traditional heuristic or analytic methods. The adjoint method, which calculates analytical gradients for all design variables using just two electromagnetic simulations, enables efficient navigation of this complex space. However, many inverse-designed structures, while numerically plausible, are difficult to fabricate and highly sensitive to physical variations, limiting their practical use. The discrete material distributions with numerous local-optimal structures also pose significant optimization challenges, often causing gradient-based methods to converge on suboptimal designs. In this work, we formulate inverse design as a fabrication-restricted, discrete, prob-abilistic optimization problem and introduce BOSON−1, an end-to-end, adaptive, variation-aware subspace optimization framework to address the challenges of manufacturability, robustness, and optimizability. We explicitly consider the fabrication process and differentiably optimize the design in the fabricable subspace. To overcome optimization difficulty, we propose dense target-enhanced gradient flows to mitigate misleading local optima and introduce a conditional subspace optimization strategy to create high-dimensional tunnels to escape local optima. Furthermore, we significantly reduce the prohibitive runtime associated with optimizing across exponential variation samples through an adaptive sampling-based robust optimization method, ensuring both efficiency and variation robustness. On three representative photonic device benchmarks, our proposed inverse design methodology BOSON−1delivers fabricable structures and achieves the best convergence and performance under realistic variations, outperforming prior arts with 74.3% post-fabrication performance.
Pingchuan Ma 0012, Zhengqi Gao, Amir Begovic, Meng Zhang 0023, Haoxing Ren, Z. Rena Huang, Duane S. Boning, Jiaqi Gu 0002
DATE2
2025 MAPS: Multi-Fidelity AI-Augmented Photonic Simulation and Inverse Design Infrastructure
abstract
Inverse design has emerged as a transformative approach for photonic device optimization, enabling exploration of high-dimensional, non-intuitive design spaces to create ultra-compact, high-performance devices, advancing photonic inte-grated circuits (PICs) in computing and interconnects. However, practical challenges, such as suboptimal device performance compared to manual designs, limited manufacturability, high sensitivity to variations, computational inefficiency, and lack of interpretability, have hindered its adoption in commercial hardware. Recent advancements in AI-assisted photonic simulation and design offer transformative potential, accelerating simulations and design generation by orders of magnitude over traditional numerical methods. Despite these breakthroughs, the lack of an open-source, standardized infrastructure and evaluation bench-mark limits accessibility and cross-disciplinary collaboration. To address this, we introduce MAPS, a multi-fidelity AI-augmented photonic simulation and inverse design infrastructure, designed to bridge this gap. MAP S features three synergistic components: 1 MAPS-Data: A dataset acquisition framework for generating multi-fidelity, richly labeled device designs using intelligent sampling strategies, providing high-quality data for AI-for-optics research. 2 MAPS-Train: A flexible AI-for-photonics training framework, offering hierarchical data loading pipeline, customizable model construction, support for data- and physics-driven losses, and comprehensive evaluation metrics. 3 MAPS-InvDes: An advanced adjoint method-based inverse design toolkit that abstracts complex physics but exposes flexible optimization steps, integrates pre-trained AI models, and incorporates fabrication-aware variation models, for real-world applicability. This infrastructure MAPS provides a unified, open-source platform for developing, benchmarking, and advancing AI-assisted photonic design workflows, accelerating innovation in photonic hardware optimization and scientific machine learning.
Pingchuan Ma 0012, Zhengqi Gao, Meng Zhang 0023, Mark Ren, Z. Rena Huang, Duane S. Boning, Jiaqi Gu 0002
DATE2
2025 REG: Rectified Gradient Guidance for Conditional Diffusion Models
abstract
Guidance techniques are simple yet effective for improving conditional generation in diffusion models. Albeit their empirical success, the practical implementation of guidance diverges significantly from its theoretical motivation. In this paper, we reconcile this discrepancy by replacing the scaled marginal distribution target, which we prove theoretically invalid, with a valid scaled joint distribution objective. Additionally, we show that the established guidance implementations are approximations to the intractable optimal solution under no future foresight constraint. Building on these theoretical insights, we propose rectified gradient guidance (REG), a versatile enhancement designed to boost the performance of existing guidance methods. Experiments on 1D and 2D demonstrate that REG provides a better approximation to the optimal solution than prior guidance techniques, validating the proposed theoretical framework. Extensive experiments on class-conditional ImageNet and text-to-image generation tasks show that incorporating REG consistently improves FID and Inception/CLIP scores across various settings compared to its absence.
Zhengqi Gao, Kaiwen Zha, Zihui Xue, Duane S. Boning
ICML1
2025 KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
abstract
Multi-agent large language model (LLM) systems are increasingly adopted for complex language processing tasks that require communication and coordination among agents. However, these systems often suffer substantial overhead from repeated reprocessing of overlapping contexts across agents. In typical pipelines, once an agent receives a message from its predecessor, the full context-including prior turns-must be reprocessed from scratch, leading to inefficient processing. While key-value (KV) caching is an effective solution for avoiding redundant computation in single-agent settings where prefixes remain unchanged, it cannot be directly reused in multi-agent scenarios due to diverging prefixes introduced by agent-specific context extensions. We identify that the core challenge lies in the offset variance of KV-caches across agents. To address this, we propose **KVCOMM**, a training-free framework that enables efficient prefilling in multi-agent inference by reusing KV-caches and aligning cache offsets of overlapping contexts under diverse prefix contexts. KVCOMM estimates and adjusts KV-caches for shared content by referencing a pool of cached examples—termed *anchors*—that store observed cache deviations under varying prefixes. The anchor pool is maintained and updated online, allowing dynamic adaptation to distinct user requests and context structures. KVCOMM achieves over 70% reuse rate across diverse multi- agent workloads, including retrieval-augmented generation, math reasoning, and collaborative coding tasks, all without quality degradation. Particularly, when each fully-connected agent receives 1K input tokens with 512 prefix tokens and 512 output tokens under a five-agent setting, KVCOMM achieves up to 7.8× speedup compared to the standard prefill pipeline, reducing TTFT from ∼430ms to ∼55ms. Code is available at https://github.com/FastMAS/KVCOMM.
Hancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang, Yuzhe Fu, Ming-Yu Chung, Yueqian Lin, Danyang Zhuo, Yiran Chen 0001
NeurIPS2
2025 RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
abstract
Reinforcement learning (RL) has recently emerged as a compelling approach for enhancing the reasoning capabilities of large language models (LLMs), where an LLM generator serves as a policy guided by a verifier (reward model). However, current RL post-training methods for LLMs typically use verifiers that are fixed (rule-based or frozen pretrained) or trained discriminatively via supervised fine-tuning (SFT). Such designs are susceptible to reward hacking and generalize poorly beyond their training distributions. To overcome these limitations, we propose Tango, a novel framework that uses RL to concurrently train both an LLM generator and a verifier in an interleaved manner. A central innovation of Tango is its generative, process-level LLM verifier, which is trained via RL and co-evolves with the generator. Importantly, the verifier is trained solely based on outcome-level verification correctness rewards without requiring explicit process-level annotations. This generative RL-trained verifier exhibits improved robustness and superior generalization compared to deterministic or SFT-trained verifiers, fostering effective mutual reinforcement with the generator. Extensive experiments demonstrate that both components of Tango achieve state-of-the-art results among 7B/8B-scale models: the generator attains best-in-class performance across five competition-level math benchmarks and four challenging out-of-domain reasoning tasks, while the verifier leads on the ProcessBench dataset. Remarkably, both components exhibit particularly substantial improvements on the most difficult mathematical reasoning problems.
Kaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong, Duane S. Boning, Dina Katabi
NeurIPS2
2024 NOFIS: Normalizing Flow for Rare Circuit Failure Analysis
abstract
Accurate estimation of rare failure occurrence probability is crucial for ensuring the proper and reliable functioning of integrated circuits (ICs). Conventional Monte Carlo methods are inefficient, demanding an exorbitant number of samples to achieve reliable estimates. Inspired by the exact sampling capabilities of normalizing flows, we revisit this problem and propose normalizing flow assisted importance sampling, termed NOFIS. NOFIS first learns a sequence of proposal distributions associated with predefined nested subset events by minimizing KL divergence losses. Next, it estimates the rare event probability by utilizing importance sampling in conjunction with the last proposal. The efficacy of our NOFIS method is substantiated through comprehensive qualitative visualizations, affirming the optimality of the learned proposal distribution, as well as 10 quantitative experiments, which highlight NOFIS's superior accuracy over baseline approaches.
Zhengqi Gao, Dinghuai Zhang, Luca Daniel, Duane S. Boning
DAC1
2024 KirchhoffNet: A Scalable Ultra Fast Analog Neural Network
abstract
In this paper, we leverage a foundational principle of analog electronic circuitry, Kirchhoff's current and voltage laws, to introduce a distinctive class of neural network models termed KirchhoffNet. Essentially, KirchhoffNet is an analog circuit that can function as a neural network, utilizing its initial node voltages as the neural network input and the node voltages at a specific time point as the output. The evolution of node voltages within the specified time is dictated by learnable parameters on the edges connecting nodes. We demonstrate that KirchhoffNet is governed by a set of ordinary differential equations (ODEs), and notably, even in the absence of traditional layers (such as convolution layers), it attains state-of-the-art performances across diverse and complex machine learning tasks. Most importantly, KirchhoffNet can be potentially implemented as a low-power analog integrated circuit, leading to an appealing property --- irrespective of the number of parameters within a KirchhoffNet, its on-chip forward calculation can always be completed within a short time. This characteristic makes KirchhoffNet a promising and fundamental paradigm for implementing large-scale neural networks, opening a new avenue in analog neural networks for AI. Our source code and model checkpoints are publicly available: https://github.com/zhengqigao/kirchhoffnet.
Zhengqi Gao, Fan-Keng Sun, Ronald A. Rohrer, Duane S. Boning
ICCAD1
2024 Improving Neural ODE Training with Temporal Adaptive Batch Normalization
abstract
Neural ordinary differential equations (Neural ODEs) is a family of continuous-depth neural networks where the evolution of hidden states is governed by learnable temporal derivatives. We identify a significant limitation in applying traditional Batch Normalization (BN) to Neural ODEs, due to a fundamental mismatch --- BN was initially designed for discrete neural networks with no temporal dimension, whereas Neural ODEs operate continuously over time. To bridge this gap, we introduce temporal adaptive Batch Normalization (TA-BN), a novel technique that acts as the continuous-time analog to traditional BN. Our empirical findings reveal that TA-BN enables the stacking of more layers within Neural ODEs, enhancing their performance. Moreover, when confined to a model architecture consisting of a single Neural ODE followed by a linear layer, TA-BN achieves 91.1\% test accuracy on CIFAR-10 with 2.2 million parameters, making it the first \texttt{unmixed} Neural ODE architecture to approach MobileNetV2-level parameter efficiency. Extensive numerical experiments on image classification and physical system modeling substantiate the superiority of TA-BN compared to baseline methods.
Su Zheng, Zhengqi Gao, Fan-Keng Sun, Duane S. Boning, Bei Yu 0001, Martin D. F. Wong
NeurIPS2
2023 The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation
Zihui Xue, Zhengqi Gao, Sucheng Ren, Hang Zhao 0021
ICLR2
2023 Nominality Score Conditioned Time Series Anomaly Detection by Point/Sequential Reconstruction
abstract
Time series anomaly detection is challenging due to the complexity and variety of patterns that can occur. One major difficulty arises from modeling time-dependent relationships to find contextual anomalies while maintaining detection accuracy for point anomalies. In this paper, we propose a framework for unsupervised time series anomaly detection that utilizes point-based and sequence-based reconstruction models. The point-based model attempts to quantify point anomalies, and the sequence-based model attempts to quantify both point and contextual anomalies. Under the formulation that the observed time point is a two-stage deviated value from a nominal time point, we introduce a nominality score calculated from the ratio of a combined value of the reconstruction errors. We derive an induced anomaly score by further integrating the nominality score and anomaly score, then theoretically prove the superiority of the induced anomaly score over the original anomaly score under certain conditions. Extensive studies conducted on several public datasets show that the proposed framework outperforms most state-of-the-art baselines for time series anomaly detection.
Chih-Yu Lai, Fan-Keng Sun, Zhengqi Gao, Jeffrey H. Lang, Duane S. Boning
NeurIPS3
2023 Correlated Bayesian Model Fusion: Efficient High-Dimensional Performance Modeling of Analog/RF Integrated Circuits Over Multiple Corners
abstract
Efficient high-dimensional performance modeling of analog/RF circuits over multiple corners is an important-yet-challenging task. In this article, we propose a novel performance modeling approach for analog/RF circuits, referred to as correlated Bayesian model fusion (C-BMF). The key idea is to encode the correlation information for both model template and coefficient magnitude among different corners by using a unified prior distribution. Next, the prior distribution is combined with a few simulation samples via Bayesian inference to efficiently determine the unknown model coefficients. Two circuit examples designed in a commercial 40-nm CMOS process demonstrate that C-BMF achieves about$2\times $cost reduction over the traditional state-of-the-art modeling technique without surrendering any accuracy.
Zhengqi Gao, Fa Wang, Jun Tao 0001, Yangfeng Su, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Circuit Theory of Time Domain Adjoint Sensitivity
abstract
It was originally stated that convolution operations were required to implement adjoint sensitivity in the time domain. In this article, we revisit time-domain adjoint sensitivity with a circuit theoretic approach and an efficient solution is clearly stated in terms of device level. Key is the linearization of the energy storage elements (e.g., capacitance and inductance) and nonlinear memoryless elements (e.g., MOS, BJT DC characteristics) at each time step. Due to the finite precision of computation, numerical errors that accumulate across timesteps can arise in nonlinear elements. A methodology to suppress that error is introduced. Numerical results demonstrate that the proposed method achieves accuracy while significantly reducing computational runtime.
Danyal Ahsanullah, Zhengqi Gao, Ronald A. Rohrer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Unleashing the Power of Graph Spectral Sparsification for Power Grid Analysis via Incomplete Cholesky Factorization
abstract
Graph spectral sparsification-based preconditioning technique has shown promising results for power grid analysis. However, the conventional methods converge slowly for high-accuracy requirement. In this work, we propose an efficient approach to address this issue. Instead of using the Cholesky factorization, we employ the incomplete Cholesky factorization to factorize the spectral sparsifier. We also propose a concept of graph spectral pattern, which can further reduce the preconditioned conjugate gradient (PCG) iterations using less number of nonzeros. Experiments show that under 10−6 relative tolerance, our proposed preconditioning technique achieves$1.17\times $speedup compared to AMGPCG in average; compared to the conventional spectral sparsification-based preconditioning techniques, our proposed approach achieves up to$8.53\times $speedup of the factorization,$8.74\times $speedup of the PCG iteration, and$5.6\times $speedup of the total time. Moreover, the speedup of the total time continues to enlarge for higher-accuracy requirement, e.g., 10−12. Finally, but not least, our method is compatible with existing graph spectral sparsification algorithms for power grid analysis.
Chunqiao Li, Chengtao An, Zhengqi Gao, Fan Yang 0001, Yangfeng Su, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Co-advise: Cross Inductive Bias Distillation
abstract
The inductive bias of vision transformers is more relaxed that cannot work well with insufficient data. Knowledge distillation is thus introduced to assist the training of transformers. Unlike previous works, where merely heavy convolution-based teachers are provided, in this paper, we delve into the influence of models inductive biases in knowledge distillation (e.g., convolution and involution). Our key observation is that the teacher accuracy is not the dominant reason for the student accuracy, but the teacher inductive bias is more important. We demonstrate that lightweight teachers with different architectural inductive biases can be used to co-advise the student transformer with outstanding performances. The rationale behind is that models designed with different inductive biases tend to focus on diverse patterns, and teachers with different inductive biases attain various knowledge despite being trained on the same dataset. The diverse knowledge provides a more precise and comprehensive description of the data and compounds and boosts the performance of the student during distillation. Furthermore, we propose a token inductive bias alignment to align the inductive bias of the token with its target teacher model. With only lightweight teachers provided and using this cross inductive bias distillation method, our vision transformers (termed as CiT) outperform all previous vision transformers (ViT) of the same architecture on ImageNet. Moreover, our small size model CiT-SAK further achieves 82.7% Top-1 accuracy on ImageNet without modifying the attention module of the ViT. Code is available at https://github.com/OliverRensu/co-advise.
Sucheng Ren, Zhengqi Gao, Tianyu Hua, Zihui Xue, Yonglong Tian, Shengfeng He, Hang Zhao 0021
CVPR2
2022 A Simple Data Mixing Prior for Improving Self-Supervised Learning
abstract
Data mixing (e.g., Mixup, Cutmix, ResizeMix) is an essential component for advancing recognition models. In this paper, we focus on studying its effectiveness in the self-supervised setting. By noticing the mixed images that share the same source images are intrinsically related to each other, we hereby propose SDMP, short for Simple Data Mixing Prior, to capture this straightforward yet essential prior, and position such mixed images as additional positive pairs to facilitate self-supervised representation learning. Our experiments verify that the proposed SDMP enables data mixing to help a set of self-supervised learning frameworks (e.g., MoCo) achieve better accuracy and out-of-distribution robustness. More notably, our SDMP is the first method that successfully leverages data mixing to improve (rather than hurt) the performance of Vision Transformers in the self-supervised setting. Code is publicly available at https://github.com/OliverRensu/SDMP.
Sucheng Ren, Zhengqi Gao, Shengfeng He, Alan L. Yuille, Yuyin Zhou, Cihang Xie
CVPR3
2022 Learning from Multiple Annotator Noisy Labels via Sample-Wise Label Fusion
Zhengqi Gao, Fan-Keng Sun, Mingran Yang, Sucheng Ren, Zikai Xiong, Marc Engeler, Antonio Burazer, Linda Wildling, Luca Daniel, Duane S. Boning
ECCV (24)1
2022 NeurOLight: A Physics-Agnostic Neural Operator Enabling Parametric Photonic Device Simulation
abstract
Optical computing has become emerging technology in next-generation efficient artificial intelligence (AI) due to its ultra-high speed and efficiency. Electromagnetic field simulation is critical to the design, optimization, and validation of photonic devices and circuits.However, costly numerical simulation significantly hinders the scalability and turn-around time in the photonic circuit design loop. Recently, physics-informed neural networks were proposed to predict the optical field solution of a single instance of a partial differential equation (PDE) with predefined parameters. Their complicated PDE formulation and lack of efficient parametrization mechanism limit their flexibility and generalization in practical simulation scenarios. In this work, for the first time, a physics-agnostic neural operator-based framework, dubbed NeurOLight, is proposed to learn a family of frequency-domain Maxwell PDEs for ultra-fast parametric photonic device simulation. Specifically, we discretize different devices into a unified domain, represent parametric PDEs with a compact wave prior, and encode the incident light via masked source modeling. We design our model to have parameter-efficient cross-shaped NeurOLight blocks and adopt superposition-based augmentation for data-efficient learning. With those synergistic approaches, NeurOLight demonstrates 2-orders-of-magnitude faster simulation speed than numerical solvers and outperforms prior NN-based models by ~54% lower prediction error using ~44% fewer parameters.
Jiaqi Gu 0002, Zhengqi Gao, Chenghao Feng, Hanqing Zhu, Ray T. Chen, Duane S. Boning, David Z. Pan
NeurIPS2
2022 Efficient Non-Monte-Carlo Yield Estimation
abstract
Parametric yield estimation is a critical component in the Integrated Circuit design flow. We propose an efficient non-Monte-Carlo yield estimation method. Key is the use of sensitivity information efficiently obtained with the nominal circuit response. Based on Taylor expansion, the circuit performance can be approximated with a multivariate Gaussian distribution. Combining this with the circuit performance specifications, the yield can be estimated efficiently by repeatedly sampling from the obtained Gaussian distribution. Also proposed is an efficient method to identify impactful factors leading to yield loss (e.g., the most or least sensitive process variables; the tightest performance specifications) based on a multidimensional Venn diagram. Circuit examples demonstrate that the proposed method can well estimate the yield while significantly reducing the number of circuit simulations. Moreover, the proposed yield analysis method can provide useful hints for yield enhancement.
Zhengqi Gao, Ronald A. Rohrer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Fast Statistical Analysis of Rare Failure Events With Truncated Normal Distribution in High-Dimensional Variation Space
abstract
In this article, to accurately estimate the rare failure rates for large-scale circuits (e.g., SRAM) where process variations are modeled as truncated normal distributions in high-dimensional space, we propose a novel truncated scaled-sigma sampling (T-SSS) method. Similar to scaled-sigma sampling (SSS), T-SSS distorts the truncated normal distributions by a scaling factor, resulting in an analytical model for failure rate estimation. By drawing random samples from the distorted distribution and estimating a sequence of scaled failure rates, we can solve all unknown model coefficients and predict the original failure rate by extrapolation. The accuracy of T-SSS is further assessed by estimating its confidence interval (CI) based on resampling. Our numerical results demonstrate that the proposed T-SSS method can achieve superior accuracy over the state-of-the-art method without increasing the computational cost.
Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Bayesian Inference on Introduced General Region: An Efficient Parametric Yield Estimation Method for Integrated Circuits
abstract
In this paper, we propose an efficient parametric yield estimation method based on Bayesian Inference. By observing that nowadays analog and mixed-signal circuit is designed via a multi-stage flow, and that the circuit performance correlation of early stage and late stage is naturally symmetrical, we introduce a general region to capture the common features of the early and late stage. Meanwhile, two private regions are also incorporated to represent the unique features of these two stages respectively. Afterwards, we introduce classifiers one for each region to explicitly encode the correlation information. Next, we set up a graphical model, and consequently adopt Bayesian Inference to calculate the model parameters. Finally, based on the obtained optimal model parameters, we can accurately and efficiently estimate the parametric yield with a simple sampling method. Our numerical experiments demonstrate that compared to the state-of-the-art algorithms, our proposed method can better estimate the yield while significantly reducing the number of circuit simulations.
Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001
ASP-DAC1
2021 Multimodal Knowledge Expansion
abstract
The popularity of multimodal sensors and the accessibility of the Internet have brought us a massive amount of unlabeled multimodal data. Since existing datasets and well-trained models are primarily unimodal, the modality gap between a unimodal network and unlabeled multi-modal data poses an interesting problem: how to transfer a pre-trained unimodal network to perform the same task with extra unlabeled multimodal data? In this work, we propose multimodal knowledge expansion (MKE), a knowledge distillation-based framework to effectively utilize multimodal data without requiring labels. Opposite to traditional knowledge distillation, where the student is designed to be lightweight and inferior to the teacher, we observe that a multimodal student model consistently rectifies pseudo labels and generalizes better than its teacher. Extensive experiments on four tasks and different modalities verify this finding. Furthermore, we connect the mechanism of MKE to semi-supervised learning and offer both empirical and theoretical explanations to understand the expansion capability of a multimodal student.1
Zihui Xue, Sucheng Ren, Zhengqi Gao, Hang Zhao 0021
ICCV3
2020 Exploring a Bayesian Optimization Framework Compatible with Digital Standard Flow for Soft-Error-Tolerant Circuit
abstract
Soft error is a major reliability concern in advanced technology nodes. Although mitigating Soft Error Rate (SER) will inevitably sacrifice area and power, few studies paid attention to optimization methods to explore trade-offs between area, power and SER. This paper proposes an optimization framework based on Bayesian approach for soft-error-tolerant circuit design. It comprises two steps:1) data preprocessing and 2) Bayesian optimization. In the preprocessing step, a strategy incorporating k-means algorithm and a novel sequencing algorithm is used to cluster Flip-Flops (FFs) with similar SER in order to reduce the dimensionality for the subsequent step. Bayesian Neural Network (BNN) is the applied surrogate model for acquiring the posterior distribution of three design metrics, while the Lower confidence bound (LCB) functions are employed as acquisition functions to select the next point based on BNN when optimizing. Finally, the non-dominated sorting genetic algorithm (NSGA-II) is used to search the Pareto Optimal Front (POF) solutions of three LCB functions. Experimental results demonstrate the proposed framework has a 1.4x improvement in accuracy and a 70% reduction in SER with acceptable increases in power and area.
Yan Li 0084, Xiaoyoung Zeng, Zhengqi Gao, Liyu Lin, Jun Tao 0001, Jun Han 0003, Xu Cheng 0002, Mehdi Baradaran Tahoori, Xiaoyang Zeng
DAC3
2020 Multi-Corner Parametric Yield Estimation via Bayesian Inference on Bernoulli Distribution with Conjugate Prior
abstract
To efficiently estimate parametric yields over multiple process, voltage, temperature corners for binary output circuits, we propose a novel Bayesian Inference method based on Bernoulli distribution with conjugate prior in this paper. The key idea is to adopt a product of Beta distributions as the conjugate prior for the yields and encode circuit performance correlations among different corners into this prior. Next, the hyper-parameters are optimized by using multi-start Quasi-Newton method, and the yields over different corners are estimated via maximum-a-posteriori. Two circuit examples demonstrate that the proposed method achieves up to 3.0× cost reduction over the state-of-the-art methods without surrendering any accuracy.
Jiahe Shi, Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001
ISCAS2
2020 Efficient Rare Failure Analysis Over Multiple Corners via Correlated Bayesian Inference
abstract
In this article, we propose an efficient correlated Bayesian inference (CBI) method to estimate the system-level failure rates for large-scale circuit systems over multiple process corners. The key idea is to encode the correlations of circuit performances among the different corners into the prior distributions of several carefully defined failure events. The hyper-parameters of these distributions can be learned from a few simulation samples via Bayesian inference and, next, the system-level failure rates over different corners can be simultaneously estimated by taking into account these prior distributions. An iteratively constrained inference method is further developed to guarantee the numerical stability of the proposed method and legalize all estimated failure rates. The numerical experiments demonstrate that compared to the state-of-the-art algorithm, the proposed method can achieve around 10× runtime reduction without surrendering any accuracy.
Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Efficient Parametric Yield Estimation Over Multiple Process Corners via Bayesian Inference Based on Bernoulli Distribution
abstract
Parametric yield estimation over multiple process corners plays an important role in robust circuit design. In this article, we propose a novel Bayesian inference method based on Bernoulli distribution (BI-BD) to efficiently estimate the multicorner yields for binary output circuit. The key idea is to encode the circuit performance correlation among different corners as our prior knowledge. Consequently, after combining a few simulation samples, the yield estimation over all corners can be calibrated via Bayesian inference based on iterative reweighted least squares (IRLS) and expectation maximization (EM). A circuit example demonstrates that the proposed BI-BD method can achieve up to 2.0 × cost reduction over the conventional Monte Carlo method without surrendering any accuracy.
Zhengqi Gao, Jun Tao 0001, Dian Zhou, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Efficient Performance Trade-off Modeling for Analog Circuit based on Bayesian Neural Network
abstract
In this paper, we propose an efficient performance trade-off modeling method for analog circuit based on Bayesian Neural Network (BNN). First, we use a single BNN to simultaneously model multiple performances of interest (PoIs) of an analog circuit. This BNN model can be trained by using a novel automatic differential variational inference (ADVI) method with affordable computational cost. Next, the performance trade-off model can be extracted by embedding BNN into Bayesian optimization framework combined with a modified multi-objective evolutionary method. Since the correlations among different PoIs are implicitly encoded in the BNN model, the proposed method can capture the performance trade-off model efficiently and accurately. The numerical experiments demonstrate that compared to the state-of-the-art algorithms, the proposed method can achieve up to 2× runtime reduction without surrendering any accuracy.
Zhengqi Gao, Jun Tao 0001, Fan Yang 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001
ICCAD1