Sheng-Guo Wang

dblp:15/8083 · also Shengguo Wang · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 1 first-author · 8 since 2021Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 LCTMwalk: GPU-Accelerated Transient Thermal Simulation for Liquid-Cooled 2.5D/3D ICs via Random Walks on Circuit Networks of Modified Compact Thermal Models
abstract
Thermal issues are critical in 2.5D/3D IC design, and liquid cooling provides an effective solution for heat dissipation. Widely used compact thermal models (CTMs) convert chips into circuit networks for fast thermal simulations. However, current matrix-solving acceleration methods for CTM-derived circuits are inadequate for high-speed iterative transient thermal analysis of large-scale liquid-cooled 2.5D/3D ICs during design optimization. In contrast, the random walk method can provide fast solutions for local nodes in large-scale circuit networks, but it is not applicable to the circuit networks of the CTMs with liquid cooling. In this paper, we propose LCTMwalk, a novel GPU-accelerated random walk method for transient thermal analysis of liquid-cooled 2.5D/3D ICs. To enable random walks on the liquid-cooled CTM-derived circuit network, we replace the voltage-controlled current source model with the diode model. Additionally, we improve the transient analysis by using a time-backward random walk with time-domain path reuse, accelerating the solution of temperature at local circuit nodes. Experimental results show LCTMwalk can solve million-scale cases in only 500 ms, and achieves a 14-22× speedup compared to the state-of-the-art alternating direction implicit (ADI) method with GPU. Besides, LCTMwalk exhibits good generalizability and can be applied to various 2.5D/3D IC structures with high accuracy (error<1 K compared to 3D-ICE).
Zhixuan Dong, Yonghan Luo, Changhao Yan, Zhaori Bi, Keren Zhu 0001, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
ICCAD7
2025 NSTherm: An Error-Bounded Network-Stochastic Fusion Thermal Simulator for Geometry-Adaptable Chiplets via Diffeomorphic Mapping and Neural-Guided Variance Reduction
abstract
For highly integrated, thermally constrained chiplets, the design process requires iterative shape optimization, making rapid thermal simulation across varying geometries critically important. Existing deterministic approaches, such as COMSOL and HotSpot require solving large-scale linear systems, incurring expensive computational costs. Stochastic methods suffer from slow convergence, demanding excessive resources for high-precision results. Current neural network (NN)-based methods necessitate retraining upon geometry modifications, limiting adaptability. Meanwhile, neural networks suffer from the absence of provable error bounds, introducing three fundamental risks in practical deployment. We enable the fast solution of heat equations for varying geometries and propose a novel solver that integrates operator learning with stochastic methods. By employing diffeomorphic mapping, our approach addresses the challenge of operator networks in handling shape variations. Furthermore, the network’s predictions guide the stochastic method for variance reduction, which extremely accelerates the traditional stochastic method, while the stochastic results provide error guarantees and corrections for the neural network’s outputs. Extensive experiments show that we achieve a speedup of 10.69-23.04× over commercial field solver COMSOL and a speedup of 5.20-11.87× over the traditional stochastic methods.
Zhixuan Dong, Yonghan Luo, Changhao Yan, Keren Zhu 0001, Zhaori Bi, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
ICCAD8
2025 pPIRW: An Efficient and Accurate Precalculation Path Integral Random Walk Solver for Steady-State Thermal Simulation With Robin Boundary Conditions
abstract
With the rapid increase of the transistor number in VLSI, rapidly rising power density and temperatures make heat dissipation a major challenge in IC design and manufacturing. However, conventional deterministic thermal analysis methods have difficulties in obtaining local temperature solutions efficiently, and the existing stochastic method is inaccurate when dealing with thermal analysis problems involving Robin boundary conditions (BCs). In this article, a highly parallelized path integral random walk (PIRW) solver is innovatively proposed for steady-state thermal analysis with mixed BCs, especially Robin BCs. The rigorous calculation of the local time and the Feynman-Kac functional$\hat {e}_{c}(t)$are adopted to accurately handle Neumann and Robin BCs for the first time. Furthermore, based on the PIRW, we propose an accurate and microsecond-level precalculation PIRW (pPIRW) predictor, which precalculates time-consuming random walks, obtains temperatures by simple vector multiplication, and therefore is suitable for proactive thermal management. The pPIRW essentially calculates a partial inverse of large-scale matrices constructed from the finite difference-based compact thermal models (CTMs). Experimental results show that compared with 3D-ICE, the PIRW solver maintains high accuracy with a negligible error within$0.5~^{\circ }$C, achieves$136\times $–$209\times $speedup and$8.53\times $–$11.1\times $storage space reduction with all three kinds of BCs, and decreases to$1\times $–$1.53\times $speedup for lacking the absorbing Dirichlet boundary. The pPIRW further has speed improvement of 2.5e$4\times $–6.7e$6\times $and memory reduction of$36.6\times $–$42.5\times $over PIRW without loss of accuracy. Integrated within a thermal management strategy, the pPIRW predictor can eliminate all thermal conflicts while maintaining the highest working frequency. Meanwhile, pPIRW achieves$29.2\times $speedup and$634\times $memory reduction over the CTM during the offline precalculation stage.
Zhixuan Dong, Longlong Yang, Cuiyang Ding, Changhao Yan, Zhaori Bi, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 APPLE-DSE: Asynchronous Parallel Pareto Set Learning for Microarchitecture Design Space Exploration
abstract
The synthesizable and parameterizable RISC-V microarchitecture, combined with multiobjective optimization-based design space exploration (DSE), facilitates agile adaptation to various microprocessor designs for customized applications. However, to enhance design quality, DSE must consider both architecture parameters and EDA tool parameters, resulting in exponentially increased optimization complexity with the dimensionality of parameters. Exhaustively exploring the whole design space is impossible. Additionally, due to the time-consuming nature of microprocessor simulation, minimizing the number of simulations is imperative. Addressing these challenges, we propose asynchronous parallel Pareto set learning for microarchitecture DSE (APPLE-DSE). APPLE-DSE utilizes the Pareto set learning (PSL) technique to obtain an approximate Pareto front with a “light-weight” evaluation. PSL captures the structural characteristics of the Pareto set (PS) guided by the surrogate models, enabling it to explore any tradeoff area in the approximate PS. Employing the probabilistic reparameterization (PR) technique, APPLE-DSE adapts PSL to handle discrete variables. Furthermore, APPLE-DSE incorporates a simulation time-aware asynchronous parallel scheduling strategy to further enhance optimization efficiency. Experimental results show that APPLE-DSE achieves a maximum improvement of 16.81% in hypervolume within the same time budget and a$127.73\times $speedup in algorithm run time per iteration compared to state-of-the-art methods.
Tianning Gao, Zhaori Bi, Changhao Yan, Fan Yang 0001, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 EVDMARL: Efficient Value Decomposition-based Multi-Agent Reinforcement Learning with Domain-Randomization for Complex Analog Circuit Design Migration
abstract
Automated analog circuit design migration significantly alleviates the burden on designers in circuit sizing under various operating conditions. Conventional methods model the migration problem as black-box optimization, requiring excessive iterations of costly simulations to converge. Reinforcement learning exhibits significant promise in transfer learning, as it enables the generation of circuits that fulfill specifications efficiently. The paper proposes a novel value decomposition-based multi-agent reinforcement learning framework, aiming to model complex analog circuits and eliminate the need for manually defined specifications of sub-circuits for new operating conditions. Additionally, it incorporates domain randomization techniques to efficiently generate circuits that meet unforeseen scenarios with minimal simulations. Experiment demonstrates that our algorithm can efficiently generate circuits meeting specifications under new operating conditions in few number of steps, outperforming state-of-the-art methods.
Handa Sun, Zhaori Bi, Wenning Jiang, Ye Lu 0005, Changhao Yan, Fan Yang 0001, Wenchuang Hu, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
DAC8
2024 A novel biometric authentication scheme with privacy protection based on SVM and ZKP
abstract
Biometric authentication is a very convenient and user-friendly method. The popularity of this method requires strong privacy-preserving technology to prevent the disclosure of template information. Most of the existing privacy protection technologies rely on classic encryption techniques, such as homomorphic encryption, which incur huge system overhead and cannot be popularized. To address these issues, we propose a novel biometric authentication scheme with privacy protection based on support vector machine and zero knowledge proof (BioAu–SVM+ZKP). BioAu–SVM+ZKP allows users to authenticate themselves to different service providers without disclosing any biometric template information. The evidence is generated through the zero-knowledge proof utilizing polynomial commitments. Our approach for generating a unique and repeatable biometric identifier from the user’s fingerprint image leverages the multi-classification property of SVM. Notably, our scheme not only reduces the communication overhead but also provides the privacy protection features. Besides, the communication overhead of BioAu–SVM+ZKP is constant. We have simulated the authentication scheme on the common dataset NIST, analyzed the performance and proved the security.
Chunjie Guo, Lin You, Gengran Hu, Sheng-Guo Wang, Chengtang Cao
Comput. Secur.5
2024 Secure and Efficient Biometric-Based Anonymous Authentication Scheme for Mobile-Edge Computing
abstract
Currently, biometric-based authentication schemes have been widely deployed in the mobile edge computing environment to ensure the authenticity of the edge nodes’ identities. While some solutions for mobile edge computing adopting a tripartite architecture employ anonymous authentication to protect the edge nodes’ sensitive biometric data and transaction information, these solutions come with their own set of challenges. Specifically, the service provider is unable to hold the edge nodes accountable for any violation committed during the transactions. Moreover, the edge nodes are unable to revoke their identity information stored in the registration center. To address these issues, in this work, we introduce a secure and efficient biometric-based anonymous authentication scheme for mobile edge computing. Our approach utilizes non-interactive zero-knowledge (NIZK) arguments and linear encryption to achieve anonymity and traceability of the edge nodes’ identities, while also using accumulators to implement the revocability of the edge nodes’ identities, ensuring that the edge nodes with revoked identities cannot pass the verification of the service provider. Additionally, we design an efficient batch verification protocol to handle anonymous authentication messages for large-scale edge nodes. The feasibility of the proposed scheme could be proven positively by both of the experimental results and the security analysis. The experimental results show that our approach requires up to 418 ms to create an anonymous identity and 16.2 ms to verify it. The time consumption for batch verification of the anonymous identities of the 1,200 edge nodes is only 54.4 ms.
Lin You, Gengran Hu, Sheng-Guo Wang
IEEE Internet Things J.4
2024 ROI-HIT: Region of Interest-Driven High-Dimensional Microarchitecture Design Space Exploration
abstract
Exploring the design space of RISC-V processors faces significant challenges due to the vastness of the high-dimensional design space and the associated expensive simulation costs. This work proposes a region of interest (ROI)-driven method, which focuses on the promising ROIs to reduce the over-exploration on the huge design space and improve the optimization efficiency. A tree structure based on self-organizing map (SOM) networks is proposed to partition the design space into ROIs. To reduce the high dimensionality of design space, a variable selection technique based on a sensitivity matrix is developed to prune unimportant design parameters and efficiently hit the optimum inside the ROIs. Moreover, an asynchronous parallel strategy is employed to further save the time taken by simulations. Experimental results demonstrate the superiority of our proposed method, achieving improvements of up to 43.82% in performance, 33.20% in power consumption, and 11.41% in area compared to state-of-the-art methods.
Tianning Gao, Aidong Zhao, Zhaori Bi, Changhao Yan, Fan Yang 0001, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 An Efficient Yield Estimation Method for Layouts of High Dimensional and High Sigma SRAM Arrays
abstract
This paper firstly focuses on yield estimation problem on post-layout-simulation of high dimensional SRAM arrays. Post-layout-simulation is much more credible than pre-simulation. However, it introduces strong relationship among SRAM columns. The Multi-Fidelity Gaussian Process model between the small and the large SRAM arrays near Optimal Shift Vector (OSV) is built. An iterative strategy is proposed and Multi-Modal method is applied to obtain more prior knowledge of the small SRAM arrays and further accelerate convergence. Experimental results show that the proposed method can gain 5-7x speedup with less relative errors than the state-of-the-art method for 384D cases.
Changhao Yan, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
DATE3
2021 A Novel and Unified Full-Chip CMP Model Aware Dummy Fill Insertion Framework With SQP-Based Optimization Method
abstract
Dummy filling is widely applied to significantly improve the planarity of topographic patterns for the chemical mechanical polishing process in VLSI manufactures. The main challenge of dummy filling is balancing multiple objectives, such as fill amounts, planarity, parasitic capacitance, etc. An obvious drawback of traditional rule-based dummy filling methods is pattern densities, instead of post-chemical mechanical polishing (CMP) topographies, being included in optimization objectives. Although the quality of post-CMP topography strongly depends on pattern features of layouts, especially the density uniformity, however, experimental results show that chip surface variations are not exactly the same as density variations. In this article, a unified dummy fill insertion optimization framework is proposed, integrated with the multiple starting points-sequential quadratic programming (MSP-SQP) optimization solver, where all objectives are considered without approximation. Inside this framework, a full-chip CMP simulator is first integrated to evaluate the planarity of the chip surface. By selecting the initial points smartly with heuristic prior knowledge, the proposed method can be effectively accelerated. The effectiveness of the proposed algorithm is verified with the average 25.8% improvement of quality compared with rule-based methods.
Junzhe Cai, Changhao Yan, Yudong Tao, Yibo Lin, Sheng-Guo Wang, David Z. Pan, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 An Efficient and Robust Yield Optimization Method for High-dimensional SRAM Circuits
abstract
Due to time-consuming SPICE simulations and extremely low failure rates, yield optimization for large static random access memory (SRAM) circuits is still a challenging problem. In this paper, a novel robust yield optimization problem is firstly proposed for SRAM circuits, where robust means considering design and process parameter variations simultaneously. Both a multi-fidelity Gaussian process regression model, which utilizes the strong nonlinear relationship between small and large SRAM columns, and a Bayesian optimization framework are applied to guide the sampling of the expensive large SRAM circuits. A multimodal problem is formulated to find all peaks and valleys on the small SRAM circuits. Such precomputational knowledge can accelerate the convergence of the proposed multi-fidelity and Bayesian optimization framework. Experimental results show that robust yield is essential to yield optimization, for traditional optimal design will degenerate with 4-5 orders of magnitude of yields, if design variations considered, and it doesn't coincide with the new optimum under the robust yield. The proposed method can gain a 3~4× speedup compared to the state-of-the-art method without loss of accuracy.
Tianchen Gu, Changhao Yan, Xiulong Wu, Fan Yang 0001, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
DAC6
2018 A general graph based pessimism reduction framework for design optimization of timing closure
abstract
In this paper, we develop a general pessimism reduction framework for design optimization of timing closure. Although the modified graph based timing analysis (mGBA) slack model can be readily formulated into a quadratic programming problem with constraints, the realistic difficulty is the size of the problem. A critical path selection scheme, a uniform sampling method with the sparse characteristics of the optimal solution, and a stochastic conjugate gradient method are proposed to accelerate the optimization solver. This modified GBA is embedded into design optimization of timing closure. Experimental results show that the proposed solver can achieve 13.82x speedup than gradient descent method with similar accuracy. With mGBA, the optimization of timing closure can achieve a better performance on area, leakage power, buffer counts.
Fulin Peng, Changhao Yan, Chunyang Feng, Jianquan Zheng, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
DAC5
2018 An efficient Bayesian yield estimation method for high dimensional and high sigma SRAM circuits
abstract
With increasing dimension of variation space and computational intensive circuit simulation, accurate and fast yield estimation of realistic SRAM chip remains a significant and complicated challenge. In this paper, du Experiment results show that the proposed method has an almost constant time complexity as the dimension increases, and gains 6x speedup over the state-of-the-art method in the 485D cases.
Jinyuan Zhai, Changhao Yan, Sheng-Guo Wang, Dian Zhou
DAC3
2018 An Efficient Non-Gaussian Sampling Method for High Sigma SRAM Yield Analysis
abstract
Yield 1 analysis of SRAM is a challenging issue, because the failure rates of SRAM cells are extremely small. In this article, an efficient non-Gaussian sampling method of cross entropy optimization is proposed for estimating the high sigma SRAM yield. Instead of sampling with the Gaussian distribution in existing methods, a non-Gaussian distribution, i.e., a joint one-dimensional generalized Pareto distribution and ( n -1)-dimensional Gaussian distribution, is taken as the function family of practical distribution, which is proved to be more suitable to fit the ideal distribution in the view of extreme failure event. To minimize the cross entropy between practical and ideal distributions, a sequential quadratic programing solver with multiple starting points strategy is applied for calculating the optimal parameters of practical distributions. Experimental results show that the proposed non-Gaussian sampling is a 2.2--4.1× speedup over the Gaussian sampling, on the whole, it is about a 1.6--2.3× speedup over state-of-the-art methods with low- and high-dimensional cases without loss of accuracy
Jinyuan Zhai, Changhao Yan, Sheng-Guo Wang, Dian Zhou, Hai Zhou 0001, Xuan Zeng 0001
ACM Trans. Design Autom. Electr. Syst.3
2017 Optimization and Quality Estimation of Circuit Design via Random Region Covering Method
abstract
Random region covering is a global optimization technique that explores the landscape by introducing multiple random starting points to initiate the local optimization solvers. This study applies the random region covering technique to circuit design automation and proposes a theory to explain why this technique is efficient at searching for the global optimum. In addition to analyzing the efficiency of the random region covering algorithm, the theory gives a probability-based estimation of the goodness of the optimization result. To enhance the efficiency of the random region covering technique, this work evaluates the boundary of top performance regions and proposes a modified random region covering method that only performs the global optimization on the top design region. The results from a large number of mathematical experiments verify the proposed methodology. The optimized designs of a class-E power amplifier and a wide load range operational amplifier outperform both manual designs and other state-of-the-art optimization techniques.
Zhaori Bi, Dian Zhou, Sheng-Guo Wang, Xuan Zeng 0001
ACM Trans. Design Autom. Electr. Syst.3
2016 A novel unified dummy fill insertion framework with SQP-based optimization method
abstract
Dummy fill insertion is widely applied to significantly improve the planarity of topographic patterns for chemical mechanical polishing process in VLSI manufacture. However, these dummies will lead to additional parasitic capacitance and deteriorate the circuit performance. The main challenge of dummy filling algorithms is how to balance multiple objectives, such as fill amount, density variation, parasitic capacitance, etc. which is the aim of ICCAD 2014 DFM contest. Traditional dummy fill insertion methods are no longer applicable because they generate large amount of fills or take unaffordable time. In this paper, we propose a unified dummy fill insertion optimization framework based on multi-starting points and sequential quadratic programming optimization solver, where all objectives are considered simultaneously without approximation. Selecting the initial points smartly with prior knowledge, the proposed method can be effectively accelerated. Even without any prior knowledge, it can also reach high fill quality by random initial points with high scalability. The proposed algorithm is verified by ICCAD 2014 DFM contest benchmark, which shows better quality of dummy filling over the state-of-the-art algorithms.
Yudong Tao, Changhao Yan, Yibo Lin, Sheng-Guo Wang, David Z. Pan, Xuan Zeng 0001
ICCAD4
2015 Rapid estimation of the probability of SRAM failure via adaptive multi-level sliding-window statistical method
Changhao Yan, Xuan Zeng 0001, Sheng-Guo Wang
Integr.4
2014 New Efficient Regression Method for Local AADT Estimation via SCAD Variable Selection
abstract
This paper focuses on the estimation and variable selection for the local annual average daily traffic (AADT). The variable selection procedure by smoothly clipped absolute deviation penalty is proposed. It can simultaneously select significant variables and estimate unknown regression coefficients in one step. The estimation algorithm and the tuning parameters selection are presented. The data from Mecklenburg County, North Carolina, USA, in 2007 are used for demonstration with our proposed variable selection procedures. The results show that this penalized regression technology improves the local AADT estimation along with satellite information, and it outperforms some other benchmark models.
Bingduo Yang, Sheng-Guo Wang, Yuanlu Bao
IEEE Trans. Intell. Transp. Syst.2
2012 A Key Sharing Fuzzy Vault Scheme
Lin You, Mengsheng Fan, Sheng-Guo Wang, Fenghai Li
ICICS4
2011 A new method for multiparameter robust stability distribution analysis of linear analog circuits
abstract
A correlation-first bisection method is proposed for analyzing the robust stability distribution of linear analog circuits in the multi-parameter space. This new method first transfers the complex multi-parameter robust stability problem into nonlinear inequalities by the Routh criterion, and then solves them by interval arithmetic and new bisection strategy. The axis with strong relationship to the functions dominating the stability is bisected. Furthermore, the Monte Carlo method is adopted for the uncertain subdomains to increase the convergence speed of bisection methods as the cube number increases. The proposed method has no error in both stable and unstable areas, and high efficiency to determine the complex boundaries between the stable and unstable areas. Numerical results validate this new method.
Changhao Yan, Sheng-Guo Wang, Xuan Zeng 0001
ICCAD2
2010 Modeling of Distributed RLC Interconnect and Transmission Line via Closed Forms and Recursive Algorithms
abstract
This paper presents the closed forms of the state-space models and the recursive algorithms of the transfer function models for fast and accurate modeling of the distributedRLCinterconnect and transmission lines, which may be evenly or unevenly distributed. Considered models include the distributedRLCinterconnect lines with or without external source and load connection. The effective closed forms and recursive algorithms do not involve any matrix inverse, LU matrix factorization, or matrix multiplication, thus reducing the computation complexity dramatically. Especially, the computation complexity of the closed forms for any evenly or unevenly distributedRLCinterconnect line circuits is onlyO(1) orO(m), respectively, in sense of the scalar multiplication times, wherem¿Nof the system order. The features of new recursive algorithms are two recursive s-polynomials and the low computation complexity. Examples illustrate the new methods in both time and frequency domains. Comparing with the PSpice, the new methods can dramatically reduce the runtime of the time responses and the Bode plots by 25% - 98.5% in the examples. The results can be applied to theRLCinterconnect analysis and model reduction as a key to new approach.
Sheng-Guo Wang
IEEE Trans. Very Large Scale Integr. Syst.1