Fan Zhang 0106

dblp:21/3626-106 · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0003-4944-9442ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-scale hierarchical voxel-aware transformer network for 3D object detection in open-pit mines
Huazhen Zhang, Zhongyu Xie, Fan Zhang 0106, Yuqian Zhao 0001
Expert Syst. Appl.3
2026 Frequency-aware and global-local selective attention network for laser welding spot detection
Ling Gong, Fan Zhang 0106, Yuqian Zhao 0001, Ji'an Duan
Neurocomputing2
2026 Light suppression context fusion network for high-precision laser weld spots recognition
Fan Zhang 0106, Zehua Deng, Shunshun Zhong, Dinghui Luo, Ji'an Duan
Neurocomputing1
2026 Geometry knowledge-embedded self-supervised deep monocular visual odometry for autonomous driving
Donglei Zheng, Yuqian Zhao 0001, Fan Zhang 0106, Gui Gui, Weihua Gui 0001
Knowl. Based Syst.4
2026 Self-Supervised Absolute-Scale Multi-Sensor Fusion Odometry via Multi-Layer Feature Fusion and Pose Refinement for Autonomous Driving
abstract
Accurate odometry is crucial for mapping and localization of autonomous vehicles in unknown environment. Deep neural networks have shown significant promise in self-supervised odometry, enabling pose estimation from consecutive sensor inputs. However, existing self-supervised visual odometry methods face scale ambiguity issue due to the inaccuracy of relative depth estimation, and self-supervised LiDAR odometry methods suffer from sparse data and the difficulty of correspondence searching. To address these challenges, we propose Self-FO, a self-supervised multi-sensor fusion odometry framework that effectively integrates image and point cloud data and overcomes the limitations of single-sensor methods. First, we utilize the inferred depth map as the fusion medium for image and point cloud, and realize multi-layer visual-LiDAR feature extraction and fusion through a novel homogeneous and heterogeneous channel exchange strategies. Then, we introduce a feature alignment and pose refinement module to fine-tune the coarse pose. By leveraging 2D-3D correspondences and a cross-modal attention mechanism, this module guides the image and point cloud features to focus on consistent scene-level observations and aligns features adaptively, significantly improving the accuracy of correspondence searching and enhancing the robustness and consistency of multimodal feature representations. Extensive experiments on KITTI and KITTI-360 dataset demonstrate that Self-FO outperforms existing learning-based odometry methods, delivering superior performance and scalability.
Donglei Zheng, Yuqian Zhao 0001, Fan Zhang 0106, Gui Gui, Weihua Gui 0001
IEEE Trans Autom. Sci. Eng.4
2026 Physics Knowledge-Inspired Scattering Neural Representation for Micro-Adhesive-Spot Segmentation Under Complex Backgrounds
abstract
Segmentation of micro-adhesive spots in high-power laser packaging is challenged by morphological variability, complex backgrounds, and blurred edges, causing traditional models to fail from “feature dilution.” Inspired by physical optics knowledge, we propose a scattering neural representation framework guided by Rayleigh scattering theory. We first pretrain a denoising diffusion model, using light scattering properties, including wavelength, scattering angle, and particle number density as an inductive bias to generate high signal-to-noise ratio target features while suppressing background clutter. Subsequently, three synergistic attention modules, including an adaptive dual-attention module, an edge attention module, and a small object enhancement module, refine target features by dynamically expanding the receptive field, sharpening boundaries, and enhancing microtarget responses. Extensive experiments on proprietary and public datasets demonstrate that our model significantly outperforms state-of-the-art methods in precise segmentation and background interference suppression. This work translates physical insights into architectural advantages, establishing an efficient and interpretable paradigm for addressing the persistent challenge of industrial small object segmentation.
Wenlong Hu, Fan Zhang 0106, Yuqian Zhao 0001, Ji'an Duan
IEEE Trans. Ind. Informatics2
2025 Frequency-domain multi-scale Kolmogorov-Arnold representation attention network for mixed-type wafer defect recognition
Fan Zhang 0106, Yuqian Zhao 0001, Ji'an Duan
Eng. Appl. Artif. Intell.2
2025 Enhancing probabilistic photovoltaic power forecasting with parallel feature interaction and bayesian correction
Fan Zhang 0106, Runmin Zou
Eng. Appl. Artif. Intell.3
2025 CSFIN: A lightweight network for camouflaged object detection via cross-stage feature interaction
Minghong Li, Yuqian Zhao 0001, Fan Zhang 0106, Gui Gui, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang
Expert Syst. Appl.3
2025 R-Net: Recursive decoder with edge refinement network for salient object detection
Hui Wang 0069, Yuqian Zhao 0001, Fan Zhang 0106, Gui Gui, Lingli Yu, Baifan Chen, Miao Liao, Chunhua Yang 0001, Weihua Gui 0001
Expert Syst. Appl.3
2025 Boundary semantic interactive aggregation network for scene segmentation
Fan Zhang 0106, Qijun Lv, Binrong Pan, Yun Wang 0047
Expert Syst. Appl.1
2025 Multi-scale spatio-temporal memory network for semi-supervised video object segmentation
Hui Wang 0069, Yuqian Zhao 0001, Fan Zhang 0106, Lingli Yu, Chunhua Yang 0001
Neurocomputing3
2025 Mine-SSD: Dual-threshold set abstraction and radius-adaptive grouping for 3D object detection in open-pit mines
Zhongyu Xie, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Wenliu Hu, Tenghai Qiu
Neurocomputing3
2025 Weakly supervised free-space segmentation by fusing spatial priors and region features for auto-driving
Dongbo Huang, Hui Wang 0069, Yuqian Zhao 0001, Feifei Guo, Fan Zhang 0106, Chunhua Yang 0001, Weihua Gui 0001
Multim. Syst.5
2024 Asymmetric convolutional multi-level attention network for micro-lens segmentation
Shunshun Zhong, Yixiong Yan, Fan Zhang 0106, Ji'an Duan
Eng. Appl. Artif. Intell.4
2024 Object detection on low-resolution images with two-stage enhancement
Minghong Li, Yuqian Zhao 0001, Gui Gui, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang, Hui Wang 0069
Knowl. Based Syst.4
2024 Multi-scale feature selection network for lightweight image super-resolution
Minghong Li, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang
Neural Networks3
2024 Hierarchical attention-guided multiscale aggregation network for infrared small target detection
Shunshun Zhong, Zhongxu Zheng, Zhu Ma, Fan Zhang 0106, Ji'an Duan
Neural Networks5
2023 TSDTVOS: Target-guided spatiotemporal dual-stream transformers for video object segmentation
Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Lingli Yu, Baifan Chen, Chunhua Yang 0001, Weihua Gui 0001
Neurocomputing3
2023 BASeg: Boundary aware semantic segmentation for autonomous driving
Xiaoyang Xiao, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Ling-Li Yu, Baifan Chen, Chunhua Yang 0001
Neural Networks3
2023 TD-Net: A Hybrid End-to-End Network for Automatic Liver Tumor Segmentation From CT Images
abstract
Liver tumor segmentation plays an essential role in diagnosis and treatment of hepatocellular carcinoma or metastasis. However, accurate and automatic tumor segmentation remains a challenging task, owing to vague boundaries and large variations in shapes, sizes, and locations of liver tumors. In this paper, we propose a novel hybrid end-to-end network, called TD-Net, which incorporates Transformer and direction information into convolution network to segment liver tumor from CT images automatically. The proposed TD-Net is composed of a shared encoder, two decoding branches, four skip connections, and a direction guidance block. The shared encoder is utilized to extract multi-level feature information, and the two decoding branches are respectively designed to produce initial segmentation map and direction information. To preserve spatial information, four skip connections are used to concatenate each encoder layer and its corresponding decoder layer, and in the fourth skip connection a Transformer module is constructed to extract global context. Furthermore, a direction guidance block is well-designed to rectify feature maps to further improve segmentation accuracy. Extensive experiments conducted on public LiTS and 3DIRCADb datasets validate that the proposed TD-Net can effectively segment liver tumor from CT images in an end-to-end manner and its segmentation accuracy surpasses those of many existing methods.
Shuanhu Di, Yuqian Zhao 0001, Miao Liao, Fan Zhang 0106, Xiong Li 0002
IEEE J. Biomed. Health Informatics4
2022 Video object segmentation based on multi-level target models and feature integration
Bocong Gao, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001
Neurocomputing3