Sifan Wang

dblp:244/2526 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Flux-conserved physics-informed neural networks for electromagnetic scattering computation
Chenyu Peng, Tiaojie Xiao, Sifan Wang, Xinhai Chen 0001, Chunye Gong
Eng. Appl. Artif. Intell.3
2026 Making Gaussian Kolmogorov-Arnold networks reliable and accurate
abstract
Kolmogorov–Arnold Networks (KANs) replace fixed activations with learnable univariate edge functions whose behavior depends strongly on the chosen basis. Gaussian radial basis functions provide a simple and efficient alternative to splines, but their accuracy and stability are highly sensitive to the scale parameter , which has not been studied systematically. We analyze this dependence through the geometry and conditioning of the first-layer feature matrix. Because the first layer is defined directly on the input domain, any loss of feature distinguishability introduced there propagates through the entire network. Based on this analysis, we identify the practical operating interval where is the number of Gaussian centers. This interval is proposed as a stable design rule rather than a universal optimum. Extensive experiments on function approximation and physics-informed problems confirm its reliability across different collocation densities, grid resolutions, architectures, and input dimensions. We also show how the same principle supports fixed-scale selection, variable-scale models, constrained optimization of , and efficient scale search using early-stage training error. These results establish scale selection as a central design principle for reliable and accurate Gaussian KANs. The code and implementation details are available at https://github.com/AmirNoori68/Gaussian-KAN .
Amir Noorizadegan, Sifan Wang, Leevan Ling
Neurocomputing2
2025 M-Rec: Leveraging Multi-View Feature to Reconstruction Incomplete Localization for Image Manipulation Detection
abstract
Image manipulation detection (IMD) serves as a critical technique for identifying forged images, playing a significant role in safeguarding cyberspace security. In this paper, we focus on a challenge that has been overlooked by previous work: incomplete localization. This issue means the network can identify potential tampered regions but lacks the confidence to decisively classify them as forgeries, resulting in only partial detection of tampered regions. When the detected regions and the undetected regions (due to a lack of confidence) differ significantly in semantics, it may mislead observers into believing that the currently localized regions represent the complete tampered regions. To address this issue, we propose M-Rec, a network designed to correct incomplete localization through a reconstruction strategy. Specifically, we introduce a confidence decoder to identify incomplete localization, and a reconstruction decoder to correct the prediction outcomes. Meanwhile, we design a bidirectional filtering strategy: (a) In the training stage, the proposed strategy is used to remove incorrect labels, ensuring that the reconstruction decoder is supervised with reliable ground truth; (b) In the inference stage, the proposed strategy constrains the reconstruction scope to the network prediction. Extensive experiments on four datasets demonstrate that our method effectively corrects incomplete localization and achieves state-of-the-art performance in manipulation detection.
Sifan Wang, Yixun Zhang
ECAI1
2025 CViT: Continuous Vision Transformer for Operator Learning
abstract
Operator learning, which aims to approximate maps between infinite-dimensional function spaces, is an important area in scientific machine learning with applications across various physical domains. Here we introduce the Continuous Vision Transformer (CViT), a novel neural operator architecture that leverages advances in computer vision to address challenges in learning complex physical systems. CViT combines a vision transformer encoder, a novel grid-based coordinate embedding, and a query-wise cross-attention mechanism to effectively capture multi-scale dependencies. This design allows for flexible output representations and consistent evaluation at arbitrary resolutions. We demonstrate CViT's effectiveness across a diverse range of partial differential equation (PDE) systems, including fluid dynamics, climate modeling, and reaction-diffusion processes. Our comprehensive experiments show that CViT achieves state-of-the-art performance on multiple benchmarks, often surpassing larger foundation models, even without extensive pretraining and roll-out fine-tuning. Taken together, CViT exhibits robust handling of discontinuous solutions, multi-scale features, and intricate spatio-temporal dynamics. Our contributions can be viewed as a significant step towards adapting advanced computer vision architectures for building more flexible and accurate machine learning models in the physical sciences.
Sifan Wang, Jacob H. Seidman, Shyam Sankaran, George J. Pappas, Paris Perdikaris
ICLR1
2025 Gradient Alignment in Physics-informed Neural Networks: A Second-Order Optimization Perspective
abstract
Physics-informed neural networks (PINNs) have shown significant promise in computational science and engineering, yet they often face optimization challenges and limited accuracy. In this work, we identify directional gradient conflicts during PINN training as a critical bottleneck. We introduce a novel gradient alignment score to systematically diagnose this issue through both theoretical analysis and empirical experiments. Building on these insights, we show that (quasi) second-order optimization methods inherently mitigate gradient conflicts, thereby consistently outperforming the widely used Adam optimizer. Among them, we highlight the effectiveness of SOAP \cite{vyas2024soap} by establishing its connection to Newton’s method. Empirically, SOAP achieves state-of-the-art results on 10 challenging PDE benchmarks, including the first successful application of PINNs to turbulent flows at Reynolds numbers up to 10,000. It yields 2–10x accuracy improvements over existing methods while maintaining computational scalability, advancing the frontier of neural PDE solvers for real-world, multi-scale physical systems. All code and datasets used in this work are publicly available at: \url{https://github.com/PredictiveIntelligenceLab/jaxpi/tree/pirate}. \end{abstract}
Sifan Wang, Ananyae Kumar Bhartari, Paris Perdikaris
NeurIPS1
2024 PirateNets: Physics-informed Deep Learning with Residual Adaptive Networks
abstract
While physics-informed neural networks (PINNs) have become a popular deep learning framework for tackling forward and inverse problems governed by partial differential equations (PDEs), their performance is known to degrade when larger and deeper neural network architectures are employed. Our study identifies that the root of this counter-intuitive behavior lies in the use of multi-layer perceptron (MLP) architectures with non-suitable initialization schemes, which result in poor trainablity for the network derivatives, and ultimately lead to an unstable minimization of the PDE residual loss. To address this, we introduce Physics-Informed Residual Adaptive Networks (PirateNets), a novel architecture that is designed to facilitate stable and efficient training of deep PINN models. PirateNets leverage a novel adaptive residual connection, which allows the networks to be initialized as shallow networks that progressively deepen during training. We also show that the proposed initialization scheme allows us to encode appropriate inductive biases corresponding to a given PDE system into the network architecture. We provide comprehensive empirical evidence showing that PirateNets are easier to optimize and can gain accuracy from considerably increased depth, ultimately achieving state-of-the-art results across various benchmarks. All code and data accompanying this manuscript will be made publicly available at https://github.com/PredictiveIntelligenceLab/jaxpi/tree/pirate.
Sifan Wang, Paris Perdikaris
J. Mach. Learn. Res.1
2024 Learning Only on Boundaries: A Physics-Informed Neural Operator for Solving Parametric Partial Differential Equations in Complex Geometries
abstract
Recently, deep learning surrogates and neural operators have shown promise in solving partial differential equations (PDEs). However, they often require a large amount of training data and are limited to bounded domains. In this work, we present a novel physics-informed neural operator method to solve parameterized boundary value problems without labeled data. By reformulating the PDEs into boundary integral equations (BIEs), we can train the operator network solely on the boundary of the domain. This approach reduces the number of required sample points from O(Nd) to O(Nd-1), where d is the domain's dimension, leading to a significant acceleration of the training process. Additionally, our method can handle unbounded problems, which are unattainable for existing physics-informed neural networks (PINNs) and neural operators. Our numerical experiments show the effectiveness of parameterized complex geometries and unbounded problems.
Zhiwei Fang, Sifan Wang, Paris Perdikaris
Neural Comput.2
2023 Mitigating Propagation Failures in Physics-informed Neural Networks using Retain-Resample-Release (R3) Sampling
abstract
Despite the success of physics-informed neural networks (PINNs) in approximating partial differential equations (PDEs), PINNs can sometimes fail to converge to the correct solution in problems involving complicated PDEs. This is reflected in several recent studies on characterizing the "failure modes" of PINNs, although a thorough understanding of the connection between PINN failure modes and sampling strategies is missing. In this paper, we provide a novel perspective of failure modes of PINNs by hypothesizing that training PINNs relies on successful "propagation" of solution from initial and/or boundary condition points to interior points. We show that PINNs with poor sampling strategies can get stuck at trivial solutions if there are propagation failures, characterized by highly imbalanced PDE residual fields. To mitigate propagation failures, we propose a novel Retain-Resample-Release sampling (R3) algorithm that can incrementally accumulate collocation points in regions of high PDE residuals with little to no computational overhead. We provide an extension of R3 sampling to respect the principle of causality while solving time-dependent PDEs. We theoretically analyze the behavior of R3 sampling and empirically demonstrate its efficacy and efficiency in comparison with baselines on a variety of PDE problems.
Arka Daw, Jie Bu, Sifan Wang, Paris Perdikaris, Anuj Karpatne
ICML3
2022 A Slide-Save Based Framework for Multi-Source DOA Extraction with Closely Spaced Sources
abstract
In adjacent sources scenarios, the low angular separation between active sources may degrade the performance of direction-of-arrival (DOA) estimation. In this work, we propose a slide-save based framework to address the problem of extracting multi-source DOAs for closely spaced sources. The basic idea is to identify the DOA estimates corresponding to the locally most dominant source within a sliding time-frequency (TF) window. Three different schemes are introduced to determine the critical DOA estimates in each TF window. The final DOAs are extracted using the retained DOA estimates by extending the histogram-based, clustering-based and Gaussian Mixture Model (GMM)-based multi-source DOA extraction methods. In addition, other intensity-based algorithms can also be incorporated into the proposed framework. Simulation results show that the proposed framework is effective to estimate multi-source DOAs in adjacent sources scenarios.
Jianhua Geng, Sifan Wang
ICASSP2
2022 Fine-Grained Feature Enhancement for Object Detection in Remote Sensing Images
abstract
Recently, object detection in aerial images has ushered in a new challenge—a new benchmark for fine-grained object recognition in high-resolution remote sensing imagery called FAIR1M has been proposed. Fine-grained categories usually have smaller inter class differences and intra-class similarities, which is more difficult to classify with existing object detectors. To address this problem, we propose two enhanced strategies on the current two-stage object detection algorithm. The first strategy uses attention-based group feature enhancement called group enhance module (GEM). By extending and grouping feature channels, the model can improve the ability to extract various discriminative features. The second strategy is to emphasize the sub-saliency feature learning, avoiding the network only focusing on the most significant part of the feature and ignoring the other parts. Our method is easy to implement and effective, and experiments show that our method can improve the Oriented regions with convolutional neural networks features (R-CNN) by about 1.45 mAP on the FAIR1M benchmark.
Yong Zhou 0003, Sifan Wang, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006
IEEE Geosci. Remote. Sens. Lett.2
2022 Dual-stream shadow detection network: biologically inspired shadow detection for remote sensing images
Dawei Li 0001, Sifan Wang, Shiyu Xiang, Xue-Song Tang
Neural Comput. Appl.2
2022 Multi-Level Time-Frequency Bins Selection for Direction of Arrival Estimation Using a Single Acoustic Vector Sensor
abstract
In the context of multi-source direction of arrival (DOA) estimation in an enclosed environment, the challenges include reverberation and overlapping of multiple simultaneous active sources. To address these interferences, the identification of time-frequency (TF) bins dominated by the sources signals is essential. In this work, we propose an intensity vector (IV) based TF bins selection technique for DOA estimation using a single acoustic vector sensor (AVS). The proposed technique involves multi-level inliers selection and outliers removal (MLISOR), which is implemented in three steps. In the first step, we derive the distribution of IVs and then select IVs using a norm metric. In the second step, the regions with the highest local IV density in each time frame are identified. In the third step, we cluster the IVs according to their directions and remove the outliers based on the member-to-centroid angle metric. Simulation results show that both the accuracy and the robustness of the proposed technique outperform the existing techniques. The indoor experimental results also verify that the proposed technique is effective and robust in practical situations.
Jianhua Geng, Sifan Wang, Qinglai Liu, Xin Lou 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Reliable Intensity Vector Selection for Multi-Source Direction-of-Arrival Estimation Using a Single Acoustic Vector Sensor
Jianhua Geng, Sifan Wang
Interspeech2
2020 Double-stream atrous network for shadow detection
Dawei Li 0001, Sifan Wang, Xue-Song Tang, Weijian Kong, Guoliang Shi
Neurocomputing2