Zhenya Zhou

dblp:167/1447 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 8 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MISP-Net: Significantly Reducing Transient Backward Steppings via Novel Multi-step Irregular Sequence Prediction
abstract
In the post-layout simulation for large-scale integrated circuits, Transient Analysis (TA), determining the time-domain response over a specified time interval, is essential and time-consuming. Especially, a mass of backward steppings and low simulation efficiency occur without proper settings of Newton-Raphson (NR) initial solution and accurate Local Truncation Error (LTE) estimation. In this work, a novel multi-step irregular sequence prediction model (MISP-Net) is proposed to predict multiple NR initial solutions and precise LTE estimations by just one inference step. This model is constructed by an Irregular Multiple Timesteps Prediction Module (IMTP) and a Irregular Multi-step Solution Prediction Module (IMSP). In IMSP, to improve the irregular prediction performance, a Dual-branch Irregular Feature Pyramid (DIFP) equipped with lightweight Multi-Channel Irregular Time Attention (MITA) are designed. We assess the proposed MISP-Net in the real large-scale industrial circuits on a commercial SPICE simulator. Compared with the commercial SPICE and the SOTA ISPT-Net model, significant backward stepping reductions are achieved: up to 78.57% for NR nonconvergence case and 76.62% for LTE overlimit case, respectively. And the prediction time for NR initial solution in our model is remarkably reduced by up to 5.58× compared to the SOTA ISPT-Net model.
Yichao Dong, Dan Niu, Chao Wang 0120, Zhenya Zhou, Zhou Jin 0001, Changyin Sun 0001
DATE4
2026 Efficient Parallel ILU Factorization and Forward/Backward Substitution with Application to Large-Scale Nonlinear Circuit Simulation
abstract
Efficient techniques are proposed for parallel incomplete LU (ILU) factorization and forward/backward substitution, for the sparse matrices with the same sparsity pattern. These parallel algorithms are then used as a preconditioner for the generalized minimal residual (GMRES) algorithm to obtain an ILU-GMRES solver for large-scale circuit simulation. The novelty of the parallel ILU and substitution algorithms includes a subtree-based task scheduling scheme, a nested dissection-based approach for generating task queues, and the task packing and reverse-order execution techniques for forward/backward substitution. Experiments on 43 matrices dumped from circuit simulation show that the 8-thread parallel ILU with threshold (ILUT) factorization and forward/backward substitution with the proposed techniques achieve 4.2 \(\times\) and 3.3 \(\times\) parallel speedups on average, respectively. The proposed parallel ILUT-GMRES solver runs 5.3 \(\times\) , on average, faster than PARDISO on these benchmarks. When integrated into Ngspice, it enables up to 2.3 \(\times\) and 1.7 \(\times\) faster execution of a step of Newton-Raphson iteration than the commercial parallel HSPICE, for the DC analysis and time integration stages, respectively.
Jiawen Cheng, Shan Shen, Zhenya Zhou, Wenjian Yu
ACM Trans. Design Autom. Electr. Syst.6
2025 A Geometry-Material Aware Point Cloud Transformer for Large-scale Unstructured Thermal Analysis in 2.5D ICs
abstract
Thermal management in large-scale unstructured 2.5D ICs faces the challenges due to the integration of complex geometries and heterogeneous materials. Existing deep learning (DL) methods urgently require a memory-efficient and high-fidelity unstructured representation method for multiscale complex ICs to simultaneously model macroscopic components and microscopic structure. Moreover, it further needs to achieve multiscale geometric thermal feature capture and thermal distribution difference adaptation among heterogeneous materials. Combining a multiscale unstructured point-cloud representation, this paper introduces Therm-PCT, a geometry-material aware point-cloud transformer framework to achieve high-accuracy thermal and its gradient prediction. Therm-PCT incorporates three key modules: adaptive multipath-coupled diffusion (AMD), a wavelet-based fine-grained recovery (WFR), and a thermal-aware Mixture-of-Material-Experts (TA-MoME) adapter. AMD adaptively learns heat diffusion path interaction with serialization-gate-based attention. Furthermore, the WFR module recovers fine-grained thermal gradients through high-frequency wavelet domain enhancement, and the TA-MoME adapter adapts to heterogeneous material by dynamically routing material-specific experts. Experiments demonstrate that the Thermal-PCT’s accuracy performance metric improvements are substantial, outperforming the newly proposed method FSA-Heat, by considerable margins of 78.03%, 84.00%, 67.61%, and 78.25% in 80 K-scale point clouds. It also achieves a 147× speed-up compared to the commercial software COMSOL. Additionally, Therm-PCT shows the potential of zero-shot generalization up to 0.4 M-scale points (5.7× than training scale) and robust performance on unseen geometric shapes.
Dekang Zhang, Dan Niu, Yichao Cao, Yichao Dong, Zhenya Zhou, Zhou Jin 0001
ICCAD5
2024 MSH: A Multi-Stage HiZ-Aware Homotopy Framework for Nonlinear DC Analysis
abstract
Nonlinear DC analysis is one of the most important tasks in transistor-level circuit simulation. Homotopy gains great success to eliminate non-convergence problem occurred in the Newton-Raphson (NR) based methods. However, nonlinear circuits with DC-path available high impedance (HiZ) nodes may fail to converge with homotopy methods due to sufficiently large resistance compared to homotopy insertions, leading to an insufficiently close enough initial-guess. In this paper, we propose a HiZ-aware homotopy framework, MSH, enabling multi-stage continuation for HiZ nodes and others separately to enhance simulation convergence. In addition, a brand-new homotopy function with limited current gain variation for MOS transistors is utilized to ensure smoother solution curve and better efficiency. Moreover, we trace the solution curve with arclength by considering homotopy parameters as unknown variables to better ensure convergence. The effectiveness of our proposed homotopy framework is demon-strated on large-scale industrial-level circuits.
Zhou Jin 0001, Tian Feng 0002, Dan Niu, Zhenya Zhou, Cheng Zhuo
DATE5
2024 ISPT-Net: A Noval Transient Backward-Stepping Reduction Policy by Irregular Sequential Prediction Transformer
abstract
In the post-layout simulation for large-scale integrated circuits, transient analysis (TA), determining the time-domain response over a specified time interval, is essential and important. However, it tends to be computationally intensive and quite time-consuming without proper settings of NR initial solution and accurate LTE estimation for determining the next transient timestep, which will lead to a mass of backward-steppings. In this paper, an irregular sequential prediction transformer named ISPT-Net is proposed to predict accurately transient solution as NR initial solution and further obtain precise LTE estimation for setting next timestep. The ISPT-Net is strengthened with timestep positional encoding module (TPE), frequency- and timestep-sensitive muti-head self-attention module (FT-MSA) to enhance irregular sequence feature extraction and prediction accuracy. We assess ISPT-Net in the real large-scale industrial circuits on a commercial SPICE simulator, and achieve a remarkable backward stepping reduction: up to 14.43X for NR nonconvergence case and 4.46X for LTE overlimit case while guaranteeing higher solution accuracy.
Yichao Dong, Dan Niu, Zhou Jin 0001, Chuan Zhang 0001, Changyin Sun 0001, Zhenya Zhou
DATE6
2024 Normal and Shear Compliance Estimation for Inclined Fractures Using Full-Waveform Sonic Log Data
abstract
In this work, we propose a phase delay approach to estimate the normal and shear fracture compliances utilizing refracted P- and S-waves from full-waveform sonic (FWS) log data generated by a monopole transmitter in a fluid-filled borehole. We derive analytical plane-wave expressions of fracture-induced P- and S-wave phase time delays for inclined compliant fractures, based on which an inversion scheme to infer fracture compliances is built. A parametric test demonstrates that, for realistic fracture compliance values from${10}^{-14}$to${10}^{-11}$m/Pa, the proposed method is capable of inferring accurate normal and shear compliance estimates for individual fractures with inclinations up to${\sim }89^{\circ }$. The applicability of the plane-wave technique to refracted waves in FWS experiments is validated by performing numerical wave propagation simulations for a fluid-filled borehole in a hard-rock environment embedding an inclined fracture. Finally, we apply the method to FWS data acquired in the Bedretto Underground Laboratory for Geosciences and Geoenergies (BULGG) in Switzerland along a borehole intersected by a fracture inclined at 71° in a granitic formation. The inferred normal and shear compliances are$(4.20\pm 0.93) {}\times {}{10}^{-13}$and$(2.96\pm 1.42) {}\times {}{10}^{-13}$m/Pa, respectively. These values are realistic given the effective fracture scale sensed by refracted sonic waves and consistent with previously reported results for normal compliance of the same fracture. Both the numerical and field experiments show that the accuracy of shear compliance estimates is much more susceptible to fracture inclination than that of normal compliance, thus, underlining the necessity of taking fracture inclination into account for reliable shear compliance estimates.
Zhenya Zhou, Eva Caspari, Nicolás D. Barbosa, Klaus Holliger
IEEE Trans. Geosci. Remote. Sens.1
2023 More Efficient Accuracy-Ensured Waveform Compression for Circuit Simulation Supporting Asynchronous Waveforms
abstract
Efficient and accurate waveform compression is critical for analog circuit simulation. In this work, we propose a waveform compression scheme which supports asynchronous waveforms, while improving the compression ratio (CR) and reducing memory usage based on the techniques of multi-model prediction, residual quantization and random-accessible secondary compression. Experimental results show that the proposed method can achieve up to 7.90X and 35.29X CR for industrial synchronous waveforms and asynchronous waveforms respectively, while keeping absolute error within 10-6 and relative error within 10-3. In comparison with existing work that only supports synchronous waveforms, the CR is improved by 1.23X and the memory usage is reduced by 8.4X on average.
Wenjian Yu, Genhua Guo, Zhenya Zhou
ACM Great Lakes Symposium on VLSI4
2021 SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUs
abstract
Sparse LU factorization is one of the key building blocks of sparse direct solvers and often dominates the computing time of circuit simulation programs. Existing GPU-accelerated sparse LU factorization methods either offload relatively small dense matrix-matrix multiplications to GPU cores, or extract level-set information to parallelize elimination operations in each level. However, because of the insufficient parallelism, neither of the methods can saturate a large amount of compute units on modern GPUs.We in this paper propose a synchronization-free sparse LU factorization algorithm called SFLU. To saturate GPU cores, our method lets each thread block eliminate a column and runs all the thread blocks at the same time. Through communicating dependency information stored on global memory, all the thread blocks either busy wait to run or get updated by their previous columns. Because elimination of all the columns work concurrently, our method avoids any barrier synchronization and saturates GPU resources. By benchmarking over 1000 sparse matrices on an NVIDIA Titan RTX GPU, our SFLU outperforms SuperLU and GLU by a factor of on average 155.71 and 8.21 (up to 3585.62 and 252.66), respectively.
Jianqi Zhao 0001, Zhou Jin 0001, Weifeng Liu 0002, Zhenya Zhou
DAC6
2021 PALBBD: A Parallel ArcLength Method Using Bordered Block Diagonal Form for DC Analysis
abstract
With the increasing complexity of integrated circuits, it is becoming cumulatively challenging to solve the entire large-scale nonlinear algebraic system in DC analysis within reasonable simulation time and without accuracy lost. For this reason, we present an efficient parallel arclength approach called PALBBD to solve DC problems for large capacity and full accuracy in this paper. We process the m+1 dimensions equation of the Newton-Raphson (NR) iteration in an alternative way, which maintains the Jacobian matrix structure. Besides, we exploit the bordered block diagonal (BBD) form to save the matrix for parallel computing. Moreover, we check the convergence of each sub-partition and bypass the calculations of converged ones to reduce the amount of unnecessary computations during the iteration. In order to ensure the accuracy, we use a correction equation to replace the Schur complement updating for the bypassed sub-partitions. The proposed PALBBD is implemented and integrated to the SPICE simulator and verified by 72 real-world circuits. It outperforms the conventional serial arclength method with up to 73.93X speedup and 45% bypass ratio.
Zhou Jin 0001, Tian Feng 0002, Yiru Duan, Minghou Cheng, Zhenya Zhou, Weifeng Liu 0002
ACM Great Lakes Symposium on VLSI6
2020 Empyrean ALPS-GT: GPU-accelerated Analog Circuit Simulation
abstract
SPICE (Simulation Program with Integrated Circuit Emphasis) has become an indispensable tool for the simulation of transistor-level circuits since its introduction in the early 1970s [1]. Over the years, many SPICE simulators have been introduced and their capabilities have been greatly improved. However, as we move into deeper sub-micron designs and circuit sizes keep increasing, the capabilities of current SPICE simulators prove insufficient; continuous performance breakthroughs are needed to keep up with the demands of larger circuits, high accuracy and fast turn-around. SPICE developers have put lots of effort to improve speed. In this paper, we will present how GPU is used to accelerate SPICE simulation with 10 × + performance gain.
Zhenya Zhou, Dake Wu
ICCAD2
2015 MOS Table Models for Fast and Accurate Simulation of Analog and Mixed-Signal Circuits Using Efficient Oscillation-Diminishing Interpolations
abstract
In this paper, we propose an efficient oscillation-diminishing cubic Hermite spline interpolation method for the table-based transistor model approximation. We use the cubic Hermite spline interpolation to ensure the continuity of the derivatives. Oscillation-diminishing techniques are proposed to reduce the oscillations (bumps) of interpolations such that both convergence and accuracy are significantly improved. Further, the oscillation-diminishing schemes do not rely on any real derivatives. Therefore, the proposed method can be used to build table models from measured data of the physical devices, where the real derivatives are not always available. In the proposed method, an adaptive approach is employed to generate the nonuniform interpolation grids such that the interpolation accuracy is guaranteed and the memory requirement is minimized. We also propose a novel combined exponential extrapolation method for off-state (leakage) current, which exactly follows the exponential-decay characteristic of that current. Test simulations on several classic industrial analog and mixed-signal circuits show that the proposed method can achieve high accuracy with lower computational cost compared with existing table-based model approximation methods.
Xiao Li 0002, Fan Yang 0001, Dake Wu, Zhenya Zhou, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4