Dan Niu

dblp:123/8337 · DBLP profile ↗
← Back
39ranked-venue papers
6as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 3 first-author · 25 since 2021Software engineering, systems software and programming languages · 12 · 12 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WaveC2R: Wavelet-Driven Coarse-to-Refined Hierarchical Learning for Radar Retrieval
abstract
Satellite-based radar retrieval methods are widely employed to fill coverage gaps in ground-based radar systems, especially in remote areas affected by terrain blockage and limited detection range. Existing methods predominantly rely on overly simplistic spatial-domain architectures constructed from a single data source, limiting their ability to accurately capture complex precipitation patterns and sharply defined meteorological boundaries. To address these limitations, we propose WaveC2R, a novel wavelet-driven coarse-to-refined framework for radar retrieval. WaveC2R integrates complementary multi-source data and leverages frequency-domain decomposition to separately model low-frequency components for capturing precipitation patterns and high-frequency components for delineating sharply defined meteorological boundaries. Specifically, WaveC2R consists of two stages (i) Intensity-Boundary Decoupled Learning, which leverages wavelet decomposition and frequency-specific loss functions to separately optimize low-frequency intensity and high-frequency boundaries; and (ii) Detail-Enhanced Diffusion Refinement, which employs frequency-aware conditional priors and multi-source data to progressively enhance fine-scale precipitation structures while preserving coarse-scale meteorological consistency. Experimental results on the publicly available SEVIR dataset demonstrate that WaveC2R achieves state-of-the-art performance in satellite-based radar retrieval, particularly excelling at preserving high-intensity precipitation features and sharply defined meteorological boundaries.
Yi-Lin Wei, Yongchao Feng, Yecheng Zhang, Dan Niu
AAAI7
2026 THGB: A Comprehensive Benchmark for Text-attributed Heterogeneous Graphs
Yuan Fang 0001, Dan Niu, Jing Ying
AAAI4
2026 MISP-Net: Significantly Reducing Transient Backward Steppings via Novel Multi-step Irregular Sequence Prediction
abstract
In the post-layout simulation for large-scale integrated circuits, Transient Analysis (TA), determining the time-domain response over a specified time interval, is essential and time-consuming. Especially, a mass of backward steppings and low simulation efficiency occur without proper settings of Newton-Raphson (NR) initial solution and accurate Local Truncation Error (LTE) estimation. In this work, a novel multi-step irregular sequence prediction model (MISP-Net) is proposed to predict multiple NR initial solutions and precise LTE estimations by just one inference step. This model is constructed by an Irregular Multiple Timesteps Prediction Module (IMTP) and a Irregular Multi-step Solution Prediction Module (IMSP). In IMSP, to improve the irregular prediction performance, a Dual-branch Irregular Feature Pyramid (DIFP) equipped with lightweight Multi-Channel Irregular Time Attention (MITA) are designed. We assess the proposed MISP-Net in the real large-scale industrial circuits on a commercial SPICE simulator. Compared with the commercial SPICE and the SOTA ISPT-Net model, significant backward stepping reductions are achieved: up to 78.57% for NR nonconvergence case and 76.62% for LTE overlimit case, respectively. And the prediction time for NR initial solution in our model is remarkably reduced by up to 5.58× compared to the SOTA ISPT-Net model.
Yichao Dong, Dan Niu, Chao Wang 0120, Zhenya Zhou, Zhou Jin 0001, Changyin Sun 0001
DATE2
2026 GE-LLM: Graph-Enhanced Large Language Models for Efficient Transistor-Level Circuit Simulation
abstract
DC analysis holds critical importance in nonlinear circuit simulation, providing the essential precondition for transient and AC analyses. While Pseudo-Transient Analysis (PTA) and its variants excel in DC analysis, selecting the optimal PTA method for specific circuits remains challenging. To address this, we propose GE-LLM, a novel framework for optimal PTA method selection, which integrates Graph Neural Networks (GNNs) with Large Language Models (LLMs). The framework first converts circuit netlists into graph representations and employs a GNN-based graph encoder to capture essential circuit topologies. Subsequently, a novel text-graph alignment strategy bridges circuit topologies and textual descriptions, enabling the LLM to effectively comprehend multimodal information. Finally, we introduce a multi-perspective few-shot prompt that mitigates data scarcity by enabling effective in-context learning from limited circuit examples. Experimental results demonstrate that GE-LLM achieves a high selection accuracy of 0.9714 and improves the efficiency of DC analysis, yielding an average speedup of 2.89× in PTA steps (up to 12.14×) and 3.45× in Newton-Raphson iterations (up to 30.39×) compared to a commercial SPICE-like simulator.
Chao Wang 0120, Dan Niu, Yichao Dong, Dekang Zhang, Changyin Sun 0001, Zhou Jin 0001
DATE2
2026 SCALER: A Stream-Aware Accelerator with Hierarchical Memory for Sparse LU Factorization on HBM FPGAs
abstract
Sparse LU factorization plays a pivotal role in many scientific and engineering applications. However, its inherent high sparsity and random non-zero distribution lead to irregular data dependencies and memory access patterns, leaving efficient acceleration on FPGAs largely unexplored. Recently, high concurrency of High Bandwidth Memory (HBM) has provided new opportunities for accelerating sparse LU factorization. Nonetheless, achieving high bandwidth utilization remains challenging given random dependencies and complex computation patterns.In this paper, we present SCALER, a high-performance sparse LU factorization accelerator on HBM FPGAs. SCALER employs a sparse storage format with vectorized packing for data coalescing, customizing HBM-compatible data streams to boost bandwidth utilization. A two-tier hierarchical memory module enhances access efficiency and data reuse by optimizing memory management and reducing redundant transfers. Furthermore, a multi-stage pipelined data prefetching mechanism hides latency, leveraging the overlap of HBM access stages to improve off-chip memory communication efficiency. Finally, a stream-aware synchronization strategy transforms irregular dependencies into hierarchical streaming access, efficiently maximizing parallelism. Evaluation on 11 matrices demonstrates SCALER’s geometric mean (geomean) throughput, energy efficiency and bandwidth efficiency surpass cuDSS solver on NVIDIA Tesla V100 GPU by 1.79×, 4.20× and 5.12×, respectively. It also outperforms the cuDSS solver on NVIDIA RTX 4090 GPU by 1.44×, 3.05× and 4.12× for the same metrics.
Zishu Li, Dan Niu, Cheng Zhuo, Zhou Jin 0001
DATE4
2026 MinFill: Reinforcement Learning and GNN Guided Reordering for Fill-In Reduction in RF Circuit Matrices
Dan Niu, Cheng Zhuo, Zhou Jin 0001
DATE2
2026 CastDiffuser: Cascaded latent diffusion framework for high-resolution precipitation nowcasting via multi source fusion
abstract
Precipitation nowcasting is critical for meteorological disaster warning, water resource management, and severe rainfall prediction, directly impacting public safety and daily operations. Existing methods based on discriminative modeling tend to produce ambiguous extrapolation maps. While generative models have improved perceptual metrics, they still suffer from limited prediction accuracy (e.g., skills scores such as Critical Success Index (CSI) and Fractions Skill Score (FSS)), high computational costs, and inadequate convective initiation prediction-with the latter stemming from constraints of single-source data. To address these challenges, we propose CastDiffuser, a cascaded latent diffusion framework for high-resolution and high-precision precipitation nowcasting. The framework performs cascaded modeling in the latent space and leverages complementary satellite and radar data, enabling both enhanced predictive accuracy and reduced computational cost. Specifically, CastDiffuser downscales and reconstructs high-resolution radar by variational autocoders. Then we utilize a spatio-temporal translator (ST-Translator) to model the deterministic components of precipitation evolution. Subsequently, a satellite-guided diffusion model is introduced to refine these deterministic features, which are then used to generate high-resolution radar predictions. Experiments on Jiangsu provincial meteorological datasets show CastDiffuser outperforms state-of-the-art methods in prediction accuracy and fine-detail preservation, particularly for heavy rainfall and convective initiation events. • We propose CastDiffuser, a Cascaded Latent Diffusion Framework thatintegrates radar and satellite data in latent space for precipitation forecasting.By modeling in a low-dimensional latent space with a cascaded design, CastDiffuser achieves superior spatio-temporal prediction accuracy while significantly reducing the computational cost of high-resolution forecasting. • We design the Spatio-Temporal Translator composed of hierarchical ST-Inception modules for robust multi-scale feature extraction. This module provides stable and structured representations for subsequent diffusion learning, thereby addressing the issue of unstable input features and improving forecast accuracy. • We introduce FsrFormer, a multi-source fusion denoising network acting as a spatial refinement module. FsrFormer adaptively modulates the influence of satellite-derived features during the diffusion process across different time steps, facilitating efficient and dynamic fusion of multi-source information. • Experimental results on real-world datasets show that the proposed CastDiffuser significantly improves the prediction performance of heavy precipitation and maintains this accuracy over a longer forecast period, especially in the convective incipient prediction task.
Dan Niu, Daben Niu, Yi-Lin Wei, Zengliang Zang, Jun Yang 0011
Neurocomputing1
2026 A Novel Antenna Tracking Method for LEO Satellites Using Bstar Coefficient Dynamic Calibration
abstract
Low Earth Orbit (LEO) satellite communication offers distinct advantages, with precise orbit prediction serving as the foundational requirement for achieving autonomous antenna tracking in these systems. A comparative analysis of the Systems Tool Kit (STK) software and the Simplified General Perturbations 4 (SGP4) model for LEO orbit prediction demonstrates that both methods maintain angular tracking errors below 0.5° in short-term predictions, meeting the accuracy requirements for small wide-beam antenna tracking. Furthermore, the SGP4 model proves more advantageous in terms of cost efficiency and operational flexibility. For medium- to long-term orbit prediction errors, a dynamic optimization framework is proposed. By integrating historical Two Line Elements (TLE) data and space weather parameters, the framework employs a genetic algorithm to optimize the ballistic coefficient globally, minimizing azimuth and elevation prediction errors. Experimental results show that the optimized model improves prediction accuracy by over 12% in both azimuth and elevation angles across a 72-hour forecasting window. Furthermore, experimental validation using a phased-array antenna with a gain of 32 dBi and a half-power beamwidth of 3 ° confirmed the method’s effectiveness. With the beam steered via Ethernet, continuous and stable signal reception was achieved when pointing to the predicted angles. This approach provides a robust and cost-efficient solution for autonomous antenna tracking in LEO satellite communication systems. Our code will be uploaded to https://github.com/miemie-323/SBBP-orbit-code.git after the acceptance of this manuscript.
Laiding Zhao, Dan Niu, Chen Ye 0001
IEEE Internet Things J.2
2025 NeuralMesh: Neural Network For FEM Mesh Generation in 2.5D/3D Chiplet Thermal Simulation
abstract
Advanced integrated circuit (IC) systems increasingly utilize chiplet-based packaging with complex $2.5 \mathrm{D} / 3 \mathrm{D}$ structures and dense Through-Silicon Via (TSV) arrays. While the Finite Element Method (FEM) provides high-fidelity thermal simulation for these systems, its computational efficiency degrades significantly when generating and optimizing meshes for intricate geometries. To address these performance limitations while preserving simulation accuracy, we present NeuralMesh, a novel framework that accelerates thermal analysis of chiplet-based ICs. Our approach integrates deep learning and geometric analysis to optimize mesh generation without the need for iterative refinement steps. NeuralMesh first employs an enhanced segmentation model to predict thermal distributions based on geometric, material, and power parameters. These predictions, combined with key geometric features, guide the optimization of an initial coarse FEM mesh. By eliminating traditional iterative mesh refinement, our framework achieves up to $45.00 \times$ mesh generation speedup while maintaining thermal accuracy within 0.8% of commercial COMSOL simulations. It reduces the number of mesh elements in unimportant areas, which represents a speed improvement of the subsequent thermal simulation. This advancement enables rapid yet precise thermal analysis essential for modern IC package design.
Pengju Chen, Dan Niu, Dekang Zhang, Depeng Xie, Zhou Jin 0001, Wei W. Xing, Lei He 0001
DAC2
2025 PiSPICE: Accelerating Post-Layout SPICE Simulation via Essential Parasitic Identification
abstract
As process nodes scale to more advanced technologies, post-layout simulations for integrated circuits have become increasingly complex, involving billions to trillions of nodes. The growing design complexity and transistor integration require more accurate and efficient post-layout SPICE simulations. However, existing methods for solving large-scale post-layout circuits face significant challenges due to high computational costs. In this paper, we propose a new approach, PiSPICE, which utilizes adjoint sensitivity analysis to identify critical parasitics and eliminate non-critical ones, effectively reducing the simulation scale and improving speed. By modeling parasitics and performing sensitivity analysis on pre-layout circuits, we significantly reduce the computational burden and avoid the overhead of directly analyzing sensitivities in large-scale postlayout circuits. By retaining only the critical parasitics and applying model order reduction to minimize their impact, while eliminating non-critical parasitics, PiSPICE achieves a speedup of up to 17.27 x in simulation with an error margin of less than 0.78% compared to the commercial simulator Spectre.
Jian Xin, Tianjia Zhou, Dan Niu, Zuochang Ye
DAC6
2025 ReChisel: Effective Automatic Chisel Code Generation by LLM with Reflection
abstract
Coding with hardware description languages (HDLs) such as Verilog is a time-intensive and laborious task. With the rapid advancement of large language models (LLMs), there is increasing interest in applying LLMs to assist with HDL coding. Recent efforts have demonstrated the potential of LLMs in translating natural language to traditional HDL Verilog. Chisel, a next-generation HDL based on Scala, introduces higher-level abstractions, facilitating more concise, maintainable, and scalable hardware designs. However, the potential of using LLMs for Chisel code generation remains largely unexplored. This work proposes ReChisel, an LLM-based agentic system designed to enhance the effectiveness of Chisel code generation. ReChisel incorporates a reflection mechanism to iteratively refine the quality of generated code using feedback from compilation and simulation processes, and introduces an escape mechanism to break free from non-progress loops. Experiments demonstrate that ReChisel significantly improves the success rate of Chisel code generation, achieving performance comparable to state-of-the-art LLM-based agentic systems for Verilog code generation.
Juxin Niu, Xiangfeng Liu, Dan Niu, Xi Wang 0009, Zhe Jiang 0004, Nan Guan
DAC3
2025 A Novel Image-Graph Heterogeneous Fusion Framework for Static IR Drop Prediction
abstract
IR drop analysis is crucial for ensuring the reliability and performance of integrated circuits (ICs) but poses computational challenges as the IC designs grow larger, especially for ultra deep-submicron VLSI designs. Deep learnings (DL) as the efficiency-promising solutions, mainly employ various CNN-based networks to achieve image-to-image IR drop predictions. However, they neglect and lose the power delivery network (PDN) global spatial features and cell instance topological information. This paper proposes a novel image-graph heterogeneous fusion framework (IGHF), which integrates the effectiveness and complementarity of dual branches (CNN and GNN) for higher prediction performance. In the CNN-based Power ScaleFusion Unet branch, the proposed long-range and local-detail encoder (LLE) integrates seamlessly with the hierarchical and adjacent compensation group (HACG) module. This design facilitates effective multi-scale global-to-local spatial power feature extraction within the PDN and enables adaptive high-to-low-level feature fusion and compensation in the decoder. Moreover, a cell voltage aware (CVA) module in the GNN branch is designed to adaptively aggregate PDN topological features of heterogeneous neighbors of different orders. Comparative experiments demonstrate that the proposed IGHF achieves significant accuracy improvements, outperforming the state-of-the-art MAUNet and widely-used IREDGe methods by considerable margins of 24.6% and 55.0% reduction in prediction error, while the prediction maps possess higher structural fidelity. Transfer experiments indicate that IGHF with transfer learning can improve the accuracy in real circuits with the few-shot real circuit test cases.
Dan Niu, Dekang Zhang, Yichao Cao, Zhou Jin 0001, Chao Wang 0120, Yichao Dong, Changyin Sun 0001
DAC1
2025 G-SpNN: GPU-Accelerated Passivity Enforcement for S-Parameter Modeling with Neural Networks
abstract
The increasing complexity of high-frequency circuits calls for efficient and accurate passive macromodeling techniques. Existing passivity enforcement methods, including those in commercial tools, often encounter convergence issues or compromise accuracy. The Domain-Alternated Optimization (DAO) framework seeks to restore accuracy through an additional optimization step but is hampered by high memory consumption and slow convergence, particularly for large-scale problems. This paper presents G-SpNN, a novel GPU-accelerated framework that recasts the passivity-enforced macromodeling problem as a neural network training task. This approach significantly enhances both the speed and scalability of passivity enforcement. Experimental results show that G-SpNN achieves an average speedup of $7.63 \times$ in convergence compared to DAO, while reducing memory usage by two orders of magnitude. This enables G-SpNN to handle complex, high-port-count circuits with greater accuracy and efficiency, paving the way for robust high-frequency circuit simulations.
Lijie Zeng, Jiatai Sun, Dan Niu, Yibo Lin, Zuochang Ye, Zhou Jin 0001
DAC4
2025 LaRED: Efficient IR Drop Predictor with Layout-Preserving Rebuilder-Encoder-Decoder Architecture
abstract
In the realm of integrated circuit verification, IR drop analysis plays a crucial role. Recent advancements in machine learning (ML) significantly enhance its efficiency, yet many current approaches fail to fully leverage the input structure of feature maps and the transmission mechanism of Power Delivery Network (PDN) layouts. To bridge these gaps, we introduce Layout-Preserving Rebuilder-Encoder-Decoder Architecture Predictor (LaRED), which employs a novel Rebuilder-Encoder-Decoder (RED) architecture and utilizes an innovative downsampling approach and upsampling framework to optimize its perception of instances and the transmission of features. LaRED captures information from various regions with asymmetric topological structure while preserving and transferring layout characteristics through deformable convolution, hybrid downsampling, cascaded upsampling, and attentional feature fusion. The rebuilder rebuilds raw input, whereas the encoder ensures comprehensive feature transmission across all instances. The decoder then facilitates seamless transfer of feature information across layers. This approach enables LaRED to integrate chip features of varying topologies and scales, enhancing its representational power. Compared to the current State-Of-The-Art (SOTA), MAUnet, LaRED achieves accuracy improvements of 34.6% to 42.6% in benchmark tests, establishing it as the new standard in static IR drop analysis for integrated circuit design with ML techniques. The code is available at https://github.com/Todi85/LaRED.
Chengxuan Yu, Yanshuang Teng, Wenhao Dai, Yongjiang Li, Wei W. Xing, Dan Niu, Zhou Jin 0001
DATE7
2025 A Novel Frequency-Spatial Domain Aware Network for Fast Thermal Prediction in 2.5D ICs
abstract
In the post-Moore era, 2.5D chiplet-based ICs present significant challenges in thermal management due to increased power density and thermal hotspots. Neural network-based thermal prediction models can perform real-time predictions for many unseen new designs. However, existing CNN-based and GCN-based methods cannot effectively capture the global thermal features, especially for high-frequency components, hindering pre-diction accuracy enhancement. In this paper, we propose a novel frequency-spatial dual domain aware prediction network (FSA-Heat) for fast and high-accuracy thermal prediction in 2.5D ICs. It integrates high-to-low frequency and spatial domain encoder (FSTE) module with frequency domain cross-scale interaction module (FCIFormer) to achieve high-to-low frequency and global-to-local thermal dissipation feature extraction. Additionally, a frequency-spatial hybrid loss (FSL) is designed to effectively attenuate high-frequency thermal gradient noise and spatial mis-alignments. The experimental results show that the performance enhancements offered by our proposed method are substantial, outperforming the newly-proposed 2.5D method, GCN+PNA, by considerable margins (over 99% RMSE reduction, 4.23X inference time speedup). Moreover, extensive experiments demonstrate that FSA-Heat also exhibits robust generalization capabilities.
Dekang Zhang, Dan Niu, Zhou Jin 0001, Yichao Dong, Jingweijia Tan, Changyin Sun 0001
DATE2
2025 A Geometry-Material Aware Point Cloud Transformer for Large-scale Unstructured Thermal Analysis in 2.5D ICs
abstract
Thermal management in large-scale unstructured 2.5D ICs faces the challenges due to the integration of complex geometries and heterogeneous materials. Existing deep learning (DL) methods urgently require a memory-efficient and high-fidelity unstructured representation method for multiscale complex ICs to simultaneously model macroscopic components and microscopic structure. Moreover, it further needs to achieve multiscale geometric thermal feature capture and thermal distribution difference adaptation among heterogeneous materials. Combining a multiscale unstructured point-cloud representation, this paper introduces Therm-PCT, a geometry-material aware point-cloud transformer framework to achieve high-accuracy thermal and its gradient prediction. Therm-PCT incorporates three key modules: adaptive multipath-coupled diffusion (AMD), a wavelet-based fine-grained recovery (WFR), and a thermal-aware Mixture-of-Material-Experts (TA-MoME) adapter. AMD adaptively learns heat diffusion path interaction with serialization-gate-based attention. Furthermore, the WFR module recovers fine-grained thermal gradients through high-frequency wavelet domain enhancement, and the TA-MoME adapter adapts to heterogeneous material by dynamically routing material-specific experts. Experiments demonstrate that the Thermal-PCT’s accuracy performance metric improvements are substantial, outperforming the newly proposed method FSA-Heat, by considerable margins of 78.03%, 84.00%, 67.61%, and 78.25% in 80 K-scale point clouds. It also achieves a 147× speed-up compared to the commercial software COMSOL. Additionally, Therm-PCT shows the potential of zero-shot generalization up to 0.4 M-scale points (5.7× than training scale) and robust performance on unseen geometric shapes.
Dekang Zhang, Dan Niu, Yichao Cao, Yichao Dong, Zhenya Zhou, Zhou Jin 0001
ICCAD2
2025 CounterPC: Counterfactual Feature Realignment for Unsupervised Domain Adaptation on Point Clouds
Yichao Cao, Xiu Su, Dan Niu, Xuanpeng Li
ICCV4
2025 M4Caster: Multi-source, multi-spatial, multi-temporal modeling for precipitation nowcasting
Dan Niu, Chunlei Shi 0001, Tianbao Zhang, Zengliang Zang, Mingbo Jiang, Jun Yang 0011
Neurocomputing1
2025 ML-PTA: A Two-Stage ML-Enhanced Framework for Accelerating Nonlinear DC Circuit Simulation With Pseudo-Transient Analysis
abstract
Direct current (DC) analysis lies at the heart of integrated circuit design in seeking DC operating points. Although pseudo-transient analysis (PTA) methods have been widely used in DC analysis in both industry and academia, their initial parameters and stepping strategy require expert knowledge and labor tuning to deliver efficient performance, which hinders their further applications. In this paper, we leverage the latest advancements in machine learning to deploy PTA with more efficient setups for different problems. More specifically, active learning, which automatically draws knowledge from other circuits, is used to provide suitable initial parameters for PTA solver, and then calibrate on-the-fly to further accelerate the simulation process using TD3-based reinforcement learning (RL). To expedite model convergence, we introduce dual agents and a public sampling buffer in our RL method to enhance sample utilization. To further improve the learning efficiency of the RL agent, we incorporate imitation learning to improve reward function and introduce supervised learning to provide a better dual-agent rotation strategy. We make the proposed algorithm a general out-of-the-box SPICE-like solver and assess it on a variety of circuits, demonstrating up to 3.10× reduction in NR iterations for the initial stage and 285.71× for the RL stage.
Zhou Jin 0001, Wenhao Li 0017, Haojie Pei, Xiaru Zha, Yichao Dong, Xiang Jin, Dan Niu, Wei W. Xing
IEEE Trans. Computers8
2024 MSH: A Multi-Stage HiZ-Aware Homotopy Framework for Nonlinear DC Analysis
abstract
Nonlinear DC analysis is one of the most important tasks in transistor-level circuit simulation. Homotopy gains great success to eliminate non-convergence problem occurred in the Newton-Raphson (NR) based methods. However, nonlinear circuits with DC-path available high impedance (HiZ) nodes may fail to converge with homotopy methods due to sufficiently large resistance compared to homotopy insertions, leading to an insufficiently close enough initial-guess. In this paper, we propose a HiZ-aware homotopy framework, MSH, enabling multi-stage continuation for HiZ nodes and others separately to enhance simulation convergence. In addition, a brand-new homotopy function with limited current gain variation for MOS transistors is utilized to ensure smoother solution curve and better efficiency. Moreover, we trace the solution curve with arclength by considering homotopy parameters as unknown variables to better ensure convergence. The effectiveness of our proposed homotopy framework is demon-strated on large-scale industrial-level circuits.
Zhou Jin 0001, Tian Feng 0002, Dan Niu, Zhenya Zhou, Cheng Zhuo
DATE4
2024 Efficient Spectral-Aware Power Supply Noise Analysis for Low-Power Design Verification
abstract
The relentless pursuit of energy-efficient electronic devices necessitates advanced methodologies for low-power design verification, with a particular focus on mitigating power supply noise. The challenges posed by shrinking voltage margins in low-power designs lead to a significant demand for rapid and accurate power supply noise simulation and verification techniques. Too large supply noise inevitably results in the raise of supply level, thereby hurting the lower power design target. Spectral methods have demonstrated as a great alternative to produce a sparse sub-matrix with spectral-similarity property as the preconditioner to efficiently reduce the iteration number and solve the linear system for supply noise verification. However, existing methods either suffer from high computational complexity or rely on approximations to reduce computational time. Therefore, a novel approach is needed to efficiently generate high-quality preconditioners. In this paper, we propose a two-stage spectral-aware algorithm to address these challenges. Our approach has three main highlights. Firstly, by introducing spectral-aware weights, we can better assess the priority of edges and construct high-quality spanning trees with the minimum relative condition number. Secondly, by leveraging eigenvalue transformation strategies, we can quickly and accurately recover off-tree edges that are spectrally critical, avoiding time-consuming iterative computations. Thirdly, we proposed a fast computation method to further decrease the computational complexity of the effective resistance. Compared with two SOTA methods, GRASS and feGRASS, our approach demonstrates higher accuracy and efficiency in preconditioner generation (37.3x and 2.13x speedup, respectively) as well as significant improvements in accelerating the linear solver for power supply noise analysis in power grid simulation and other Laplacian graphs (5.16x and 1.70x speedup, respectively).
Yinuo Bai 0002, Yicheng Lu, Dan Niu, Cheng Zhuo, Zhou Jin 0001, Weifeng Liu 0002
DATE4
2024 TSA-TICER: A Two-Stage TICER Acceleration Framework for Model Order Reduction
abstract
To enhance the post-simulation efficiency of large-scale integrated circuits, various model order reduction (MOR) methods have been proposed. Among these, TICER (Time-Constant Equilibration Reduction) is a widely-used resistor-capacitor (RC) network reduction algorithm. However, the time constant computation for eliminated-node classification in TICER is quite time-consuming. In this work, a two-stage TICER acceleration framework (TSA-TICER) is proposed. First, an improved graph attention network (named BCTu-GAT) equipped with betweenness centrality metric (BCM) based sample selection strategy and bi-level aggregation-based topology updating scheme (BiTu) is proposed to quickly and accurately determine all the eliminated nodes one time in the TICER. Second, an adaptive merging strategy for the new fill-in capacitors are designed to further accelerate the insertion stage. The proposed TSA - TI CER is tested on RC networks with the size from 2k to 2 million nodes. Experimental results show that the proposed TSA-TICER achieves up to 796.21X order reduction speedup and 10.46X fill-in speedup compared to the TICER with 0.574% maximum relative error.
Pengju Chen, Dan Niu, Zhou Jin 0001, Changyin Sun 0001
DATE2
2024 ISPT-Net: A Noval Transient Backward-Stepping Reduction Policy by Irregular Sequential Prediction Transformer
abstract
In the post-layout simulation for large-scale integrated circuits, transient analysis (TA), determining the time-domain response over a specified time interval, is essential and important. However, it tends to be computationally intensive and quite time-consuming without proper settings of NR initial solution and accurate LTE estimation for determining the next transient timestep, which will lead to a mass of backward-steppings. In this paper, an irregular sequential prediction transformer named ISPT-Net is proposed to predict accurately transient solution as NR initial solution and further obtain precise LTE estimation for setting next timestep. The ISPT-Net is strengthened with timestep positional encoding module (TPE), frequency- and timestep-sensitive muti-head self-attention module (FT-MSA) to enhance irregular sequence feature extraction and prediction accuracy. We assess ISPT-Net in the real large-scale industrial circuits on a commercial SPICE simulator, and achieve a remarkable backward stepping reduction: up to 14.43X for NR nonconvergence case and 4.46X for LTE overlimit case while guaranteeing higher solution accuracy.
Yichao Dong, Dan Niu, Zhou Jin 0001, Chuan Zhang 0001, Changyin Sun 0001, Zhenya Zhou
DATE2
2024 Embedded Representation Learning Network for Animating Styled Video Portrait
abstract
The talking head generation recently attracted considerable attention due to its widespread application prospects, especially for digital avatars and 3D animation design. Inspired by this practical demand, several works explored Neural Radiance Fields (NeRF) to synthesize the talking heads. However, these methods based on NeRF face two challenges: (1) Difficulty in generating style-controllable talking heads. (2) Displacement artifacts around the neck in rendered images. To overcome these two challenges, we propose a novel generative paradigm Embedded Representation Learning Network (ERLNet) with two learning stages. First, the audio-driven FLAME (ADF) module is constructed to produce facial expression and head pose sequences synchronized with content audio and style video. Second, given the sequence deduced by the ADF, one novel dual-branch fusion NeRF (DBF-NeRF) explores these contents to render the final images. Extensive empirical studies demonstrate that the collaboration of these two stages effectively facilitates our method to render a more realistic talking head than the existing algorithms.
Tianyong Wang, Xiangyu Liang, Wangguandong Zheng, Dan Niu, Haifeng Xia, Si-Yu Xia
FG4
2024 ISLU: Indexing-Efficient Sparse LU Factorization for Circuit Simulation on GPUs
abstract
Sparse LU factorization is a vital technique in solving circuit linear equations, However, irregular data access patterns contribute to unsatisfactory computational efficiency and excessive memory usage. Conventional LU factorization methods generally involve two approaches: either they utilize space-intensive dense matrices for direct index-to-data mapping, or they inefficiently scour through indices to locate the positions of updated data elements. To resolve these challenges, we propose the Indexing-Efficient Sparse LU factorization (ISLU) in this work. A novel indexing-efficient member union is put forwarded to achieve efficient retrieval of indices within compressed formats, thereby significantly enhancing the LU decomposition efficiency. Furthermore, to expedite the establishment of indexing-efficient member union, we design, for the first time, parallel creating member union strategy for GPU platforms, which remarkably reduces the time overhead associated with constructing the proposed structures. Extensive experimental comparisons on 49 benchmark matrices and real SPICE transient simulations demonstrate that the performance enhancements by our proposed ISLU method are substantial, outperforming various excellent GPU and CPU solvers including commercial solvers.
Dan Niu, Yiyang Tao, Zhou Jin 0001, Yichao Dong, Chao Wang 0120, Changyin Sun 0001
ICCAD1
2024 Pseudo Adjoint Optimization: Harnessing the Solution Curve for SPICE Acceleration
abstract
Pseudo transient analysis (PTA) has been a promising solution for direct current (DC) analysis of transistor-level circuit simulation. Despite its popularity, PTA requires meticulous hyperparameter tuning for optimal performance. In this paper, we propose pseudo adjoint optimization, Soda-PTA, which models the PTA solution curve (which is used to measure convergence) using a neural ordinary differential equation (Neural ODE) and deriving explicit gradients of the Newton-Raphson (NR) iteration w.r.t. the PTA hyperparameters through the classic adjoint method, enabling effective optimization of the PTA hyperparameters. To generalize Soda-PTA for unseen circuits, we further introduce a graph convolution network to transfer optimal PTA hyperparameters from the other circuits to the target one. Soda-PTA is implemented in an out-of-the-box SPICE simulator. Through extensive experiments, Soda-PTA demonstrates superior acceleration performance: an average speedup of 1.53x over the state-of-the-art BoA-PTA while ensuring superior convergence and up to 22.12x speedup compared to the native PTA solver.
Jiatai Sun, Xiaru Zha, Chao Wang 0120, Dan Niu, Wei W. Xing, Zhou Jin 0001
ICCAD5
2024 Leda: Leveraging Tiling Dataflow to Accelerate SpMM on HBM-Equipped FPGAs for GNNs
abstract
Graph neural networks (GNNs) play a pivotal role in extracting insightful representations from graph-structured data, driving advancements across diverse domains. Central to GNNs is the sparse matrix-dense matrix multiplication (SpMM) kernel. However, challenges arise in accelerating SpMM due to the high sparsity and randomly distributed non-zeros in graph matrices. Recently, the high concurrency capability of high bandwidth memory (HBM) has provided a new opportunity for SpMM acceleration. Nonetheless, accelerating SpMM on HBM FPGAs is still non-trivial due to load imbalance and the random memory access patterns.
Enxin Yi, Jiarui Bai, Yijie Nie, Dan Niu, Zhou Jin 0001, Weifeng Liu 0002
ICCAD4
2024 CSP: Comprehensively-Sparsified Preconditioner for Efficient Nonlinear Circuit Simulation
abstract
Solving sparse linear systems dominates the simulation time for nonlinear integrated circuits. Developing an effective preconditioner is crucial for accelerating the iterative solver when dealing with large-scale circuit matrices, yet this remains a challenging task. In this paper, we introduce an efficient sparsification-based preconditioner method that significantly reduces the number of iterations needed in iterative solvers. Our method transforms nonlinear components into symmetric Laplacian matrices, enabling the inclusion of both nonlinear and linear elements in the sparsification process. We then intersect the generated sparsifier with the original Modified Nodal Analysis (MNA) matrix to further reduce the sparsity, thereby decreasing preconditioner factorization time. Furthermore, we enhance the parallelization of the spectral sparsification strategy by integrating block RMQ and point exclusivity algorithms, which substantially speeds up preprocessing. Experiment results demonstrate acceleration of 2.50x, 13.46x, 2.18x on average in serial, 3.72x, 24.23x, 3.86x on average in parallel, and memory reduction of 21.3%, 21.7%, 88.0% on average when solving nonlinear circuit matrices compared to the state-of-the-art solver GPSCP, feGRASS, and direct solver KLU, respectively.
Yinuo Bai 0002, Lijie Zeng, Dan Niu, Weifeng Liu 0002, Zhou Jin 0001
ICCAD5
2024 SCRD: A Spatiotemporal Cues-Guided Residual Diffusion Model for Precipitation Nowcasting
abstract
Precipitation nowcasting is crucial in the field of weather forecasting, and it impacts various public services ranging from rainstorm warnings to flight safety. The existing deterministic model-based methods tend to yield blurry extrapolation maps. In contrast, probabilistic generative models focus on producing realistic predictions with more details, but often have unsatisfactory forecasting accuracy. To address high forecast clarity and high-accuracy balance challenge, we propose a spatiotemporal conditional cues-guided residual diffusion (SCRD) network for precipitation nowcasting, where the spatiotemporal conditional cues-guided (STCG) module and shift window cross-interaction (SWCI) module working with residual prediction strategy extract and adaptively feed the multiscale spatiotemporal (ST) auxiliary cues to the noise generation process, enhancing the prediction accuracy of the diffusion-based model. Experiments on the real-world radar echo dataset demonstrate that the proposed SCRD significantly outperforms typical diffusion-based MCVD and ingenious deterministic models in both heavy rainfall prediction accuracy and image clarity.
Dan Niu, Zengliang Zang, Mingbo Jiang
IEEE Geosci. Remote. Sens. Lett.2
2023 A Multi-Crane Scheduling Scheme with Dynamic Priority in Transit Warehouse
abstract
Aiming at the Multi-Crane Scheduling Problem (MCSP) in the transit warehouse, a scheduling scheme with dynamic priority is proposed in this paper. The problem was modelled with the goal of respectively minimizing the completion time of all tasks and the frequency of crane avoidance. A Simulation-based Genetic Algorithm (SbGA) is designed to solve the model and generate feasible scheduling schemes. Besides, a novel updating strategy for task priority is designed to implement dynamic scheduling. The proposed scheme is applied to an actual industrial case in an iron and steel enterprise. Numerous experiments demonstrate the efficiency of the proposed model and scheme.
Dan Niu, Xisong Chen
CoDIT2
2023 A New Combined Controller for an Industrial Heavy-Duty 3D Overhead Crane System with Load Hoisting or Lowering
abstract
In this paper, a novel combined controller is proposed for an industrial heavy-duty 3D overhead crane system with load hoisting or lowering. It is of great significance to develop the controller of industrial heavy-duty bridge crane to improve production efficiency and reduce safety accidents. Many existing research works have not fully considered many practical factors in the actual industrial overhead crane system, such as actuators, speed limitation, motor current and work efficiency, which has caused difficulties in practical application. The anti-sway module of the proposed controller combines the commonly used the ZVD shaper, the first-order inertial filter and the saturation element to effectively suppress the residual swing of the time-varying rope length, while alleviating the motor current surge and preventing the speed from exceeding the upper limit. The positioning module of the proposed controller optimizes the traditional PID algorithm based on the Multi-dimensional Taylor Network (MTN) to increase efficiency, which saves about 30% of the settling time when the maximum allowable error is 5cm. The overhead crane, vector drive and AC induction motor are modeled on the SIMULINK platform, the proposed algorithm is given in discrete form, and some simulations are performed to verify the effectiveness of the proposed combined controller.
Dan Niu, Xisong Chen
CoDIT2
2023 An Adaptive Event-Triggered Secondary Regulation Strategy for Microgrids with Loss of Effectiveness Actuator Faults
abstract
In this paper, an adaptive event-triggered secondary regulation strategy is investigated for microgrids with loss of effectiveness actuator faults. In order to deal with unknown loss of effectiveness actuator faults, a distributed secondary regulation strategy is proposed, which achieves voltage and frequency regulations, as well as power sharing. Meanwhile, to save system resources and relieve the communication burden, an adaptive event-triggered mechanism is designed. Finally, some simulation results are given to validate the proposed strategy, which indicates that the proposed strategy reduces the controller updates and increases the reliability of system.
Xuechao Qiu, Xiangyu Wang 0003, Dan Niu
IECON3
2023 Ada-CCFNet: Classification of multimodal direct immunofluorescence images for membranous nephropathy via adaptive weighted confidence calibration fusion network
Ruili Wang 0009, Xueyu Liu, Fang Hao, Dan Niu, Yongfei Wu
Eng. Appl. Artif. Intell.7
2023 OSSP-PTA: An Online Stochastic Stepping Policy for PTA on Reinforcement Learning
abstract
The dc analysis is essential and still quite challenging in large-scale nonlinear circuit simulation. Pseudo transient analysis (PTA) is a widely used and has great potential solver in the industry. However, the PTA convergence and simulation efficiency is still seriously affected by its stepping policy. This article proposes an online stochastic stepping policy (OSSP) for PTA based on deep reinforcement learning (DRL). To achieve better policy evaluation and stronger stepping exploration ability, the dual soft Actor–Critic agents work with the proposed valuation splitting and online momental scaling, enabling our OSSP to intelligently encode PTA iteration status and online further adjust forward and backward time-step size for unseen test circuits without human intervention and domain knowledge, trained solely by reinforcement learning from self-search. Our public sample buffer and priority sampling are also introduced to overcome the sparsity and imbalance of sample data. Numerical examples demonstrate that the proposed OSSP achieves a significant efficiency speedup (up to$47.0\times $less Newton–Raphson iterations) and convergence enhancement on unseen test circuits compared with the previous iter-based and switched evolution/relaxation-based stepping methods, in just one stepping iteration.
Dan Niu, Yichao Dong, Zhou Jin 0001, Chuan Zhang 0001, Changyin Sun 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 BoA-PTA: A Bayesian Optimization Accelerated PTA Solver for SPICE Simulation
abstract
One of the greatest challenges in integrated circuit design is the repeated executions of computationally expensive SPICE simulations, particularly when highly complex chip testing/verification is involved. Recently, pseudo-transient analysis (PTA) has shown to be one of the most promising continuation SPICE solvers. However, the PTA efficiency is highly influenced by the inserted pseudo-parameters. In this work, we proposed BoA-PTA, a Bayesian optimization accelerated PTA that can substantially accelerate simulations and improve convergence performance without introducing extra errors. Furthermore, our method does not require any pre-computation data or offline training. The acceleration framework can either speed up ongoing, repeated simulations (e.g., Monte-Carlo simulations) immediately or improve new simulations of completely different circuits. BoA-PTA is equipped with cutting-edge machine learning techniques, such as deep learning, Gaussian process, Bayesian optimization, non-stationary monotonic transformation, and variational inference via reparameterization. We assess BoA-PTA in 43 benchmark circuits and real industrial circuits against other SOTA methods and demonstrate an average of 1.5x (maximum 3.5x) for the benchmark circuits and up to 250x speedup for the industrial circuit designs over the original CEPTA without sacrificing any accuracy.
Wei W. Xing, Xiang Jin, Tian Feng 0002, Dan Niu, Weisheng Zhao 0001, Zhou Jin 0001
ACM Trans. Design Autom. Electr. Syst.4
2022 Accelerating nonlinear DC circuit simulation with reinforcement learning
abstract
DC analysis is the foundation for nonlinear electronic circuit simulation. Pseudo transient analysis (PTA) methods have gained great success among various continuation algorithms. However, PTA tends to be computationally intensive without careful tuning of parameters and proper stepping strategies. In this paper, we harness the latest advancing in machine learning to resolve these challenges simultaneously. Particularly, an active learning is leveraged to provide a fine initial solver environment, in which a TD3-based Reinforcement Learning (RL) is implemented to accelerate the simulation on the fly. The RL agent is strengthen with dual agents, priority sampling, and cooperative learning to enhance its robustness and convergence. The proposed algorithms are implemented in an out-of-the-box SPICElike simulator, which demonstrated a significant speedup: up to 3.1X for the initial stage and 234X for the RL stage.
Zhou Jin 0001, Haojie Pei, Yichao Dong, Xiang Jin, Wei W. Xing, Dan Niu
DAC7
2022 ED-DRAP: Encoder-Decoder Deep Residual Attention Prediction Network for Radar Echoes
abstract
Precipitation nowcasting is quite important and fundamental. It underlies various public services ranging from rainstorm warnings to flight safety. In order to further improve the prediction accuracy for the spatiotemporal sequence forecasting problem, we propose an encoder–decoder deep residual attention prediction network, which adaptively rescales the multiscale sequence- and spatial-wise features and achieves very deep trainable residual prediction by integrating global residual learning and local deep residual sequence and spatial attention blocks (RSSABs). Experiments in a real-world radar echo map dataset of South China show that compared with the ingenious PredRNN++, TrajGRU methods, and newly proposed Unet-based methods, our ED-DRAP network performs better on the precipitation nowcasting metrics, as well as occupies small GPU memory.
Hongshu Che, Dan Niu, Zengliang Zang, Yichao Cao, Xisong Chen
IEEE Geosci. Remote. Sens. Lett.2
2017 MPC for Ozone Dosage in Water Treatment Process based on Disturbance Observer
Dan Niu, Xisong Chen, Jun Yang 0011, Fuchun Jiang, Xing-peng Zhou
ICINCO (2)1
2017 Learning Spatial Constraints using Gaussian Process for Shared Control of Semi-autonomous Mobile Robots
Kun Qian 0005, Dan Niu, Fang Fang 0006, Xudong Ma
ICINCO (2)2