Xinnan Fan

dblp:153/8494 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 A novel sonar image instance segmentation algorithm based on multi-scale attention and enhanced feature fusion
Qi He 0010, Xinnan Fan, Yuanxue Xin
Pattern Recognit.4
2025 Boosting Multi-modal Fusion for 3D Vehicle Object Detection
abstract
Significant advancements have been made in neural networks for 3D object detection in autonomous driving. However, these vehicles often encounter small and occluded objects, leading to fewer available features and requirement of high positioning accuracy. Current approaches to 3D vehicle detection frequently overlook this challenge, simply feeding features into existing detection models. This paper introduces an innovative boosting multi-modal fusion for 3D vehicle object detection. Initially, we employ pre-trained 3D and 2D object detection models to generate 3D and 2D bounding boxes. Subsequently, a fusion strategy grounded in the rotation intersection ratio, merges two kinds of bounding boxes. To capture information from small objects, we develop a grouping-splitting residual network enhanced with coordinate attention, facilitating the extraction of more detailed information. Experimental results on KITTI dataset reveal that our method achieves a 90.46% accuracy for hard samples in Bird’s eye view. Compared with the advanced multi-modal 3D object detection performance, such as CLOCs, HMFI and PointPainting, our accuracy in hard samples has improved by 1.1%, 1.84%, and 3.75%.
Xinnan Fan, Yuanxue Xin
ACM Trans. Sens. Networks1
2024 Multi-scale fusion and efficient feature extraction for enhanced sonar image object detection
Qi He 0010, Sisi Zhu, Xinnan Fan, Yuanxue Xin
Expert Syst. Appl.5
2024 An Effective Strategy of Object Instance Segmentation in Sonar Images
abstract
Instance segmentation is a task that involves pixel‐level classification and segmentation of each object instance in images. Various CNN‐based methods have achieved promising results in natural image instance segmentation. However, the noise interference, low resolution, and blurred edges bring more significant challenges for sonar image instance segmentation. To solve these problems, we propose the Effective Strategy for Sonar Images Instance Segmentation (ESSIIS). We introduce ASception, a new network combining Atrous Spatial Pyramid Pooling (ASPP) and Extreme Inception (Xception). By integrating this with ResNet and transforming traditional convolutions into deformable convolutions, we further improve the ability of the network to extract features from sonar images. Additionally, we incorporate a bidirectional feature fusion module to enhance information fusion. Finally, we evaluate the detection accuracy and segmentation accuracy of the proposed method on the public sonar image dataset and the self‐constructed dataset. ESSIIS attains a detection accuracy of 0.981 and a segmentation accuracy of 0.951 on SCTD, further impressively achieving 0.986 in both metrics when appraised on our dataset. The evaluation results demonstrate that the proposed method is more accurate, robust, and considerable for sonar image detection and segmentation.
Huanru Sun, Qi He 0010, Hanren Wang, Xinnan Fan, Yuanxue Xin
IET Signal Process.5
2024 An effective automatic object detection algorithm for continuous sonar image sequences
Huanru Sun, Xinnan Fan
Multim. Tools Appl.3
2024 CrackYOLO: towards efficient dam crack detection for underwater scenes
Shen Shao, Xinnan Fan, Yuanxue Xin, Zhongkai Zhou, Sisi Zhu
Pattern Anal. Appl.3
2024 Recurrent Multiscale Feature Modulation for Geometry Consistent Depth Learning
abstract
The U-Net-like coarse-to-fine network design is currently the dominant choice for dense prediction tasks. Although this design can often achieve competitive performance, it suffers from some inherent limitations, such as training error propagation from low to high resolution and the dependency on the deeper and heavier backbones. To design an effective network that performs better, we instead propose Recurrent Multiscale Feature Modulation (R-MSFM), a new lightweight network design for self-supervised monocular depth estimation. R-MSFM extracts per-pixel features, builds a multiscale feature modulation module, and performs recurrent depth refinement through a parameter-shared decoder at a fixed resolution. This network design enables our R-MSFM to maintain a more lightweight architecture and fundamentally avoid error propagation caused by the coarse-to-fine design. Furthermore, we introduce the mask geometry consistency loss to facilitate our R-MSFM for geometry consistent depth learning. This loss penalizes the inconsistency of the estimated depths between adjacent views within the nonoccluded and nonstationary regions. Experimental results demonstrate the superiority of our proposed R-MSFM both at model size and inference speed, and show state-of-the-art results on two datasets: KITTI and Make3D.
Zhongkai Zhou, Xinnan Fan, Yuanxue Xin, Dongliang Duan, Liuqing Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 R-LKDepth: Recurrent Depth Learning With Larger Kernel
abstract
Monocular depth estimation is a critical task with significant potential for real-time applications. However, current networks for monocular depth estimation exhibit a trade-off between computational overhead and accuracy. Current top-performing coarse-to-fine networks excel in building multi-scale features and thus get a satisfactory result, but suffer from an excessive count of model parameters of its deeper backbone. In contrast, recurrent networks significantly save the use of model parameters, but their performance is often restricted by the smaller receptive field, leading to insufficient local evidence to handle complex objects and homogeneous regions. In this letter, we aim to alleviate the limitation and thus propose an efficient recurrent network R-LKDepth. Specifically, we first introduce the larger kernel scheme with a careful design to effectively extract the image features with larger receptive fields. Furthermore, we propose a group-wise multi-scale self-attention based refine-feed (GMSA-RF) module, which efficiently aggregates multi-scale information for recurrent refinement procedure. The resulting network, R-LKDepth, significantly improves the depth estimation performance over complex objects and homogeneous regions and achieves state-of-the-art performance on standard benchmarks with superior cross-dataset generalization capability.
Zhongkai Zhou, Xinnan Fan, Yuanxue Xin
IEEE Signal Process. Lett.2
2022 An underwater dam crack image segmentation method based on multi-level adversarial transfer learning
Xinnan Fan, Qian Gong
Neurocomputing1
2022 A novel sonar target detection and classification algorithm
Xinnan Fan, Xuewu Zhang 0001
Multim. Tools Appl.1
2022 A novel underwater sonar image enhancement algorithm based on approximation spaces of random sets
Xinnan Fan, Yuanxue Xin, Jianjun Ni
Multim. Tools Appl.3
2022 RAFM: Recurrent Atrous Feature Modulation for Accurate Monocular Depth Estimating
abstract
Current top-performing coarse-to-fine monocular depth estimation systems mainly depend on deeper backbones, such as a full ResNet50. These systems benefit from the powerful multiscale feature representations but suffer from the high computational costs and memory overheads. Conversely, the recurrent depth refinement systems fulfill the lightweight architecture, while the local multiscale context often limits their accuracies. To handle this problem, we propose an atrous spatial (AS) module by utilizing atrous convolutions with different rates, which efficiently captures the multiscale context. We further assemble our AS module to the current recurrent depth refinement system R-MSFM, constructing a powerful monocular depth estimation system called RAFM. To evaluate its effectiveness, we conduct the comprehensive experiments and show that our RAFM outperforms the previous SOTA methods in all metrics on KITTI while keeping the real-time speed. For accuracy, our RAFM achieves a square relative error (Sq Rel) of 0.702 on KITTI, a 12.6% error reduction from the previous SOTA method (0.802). For model size and inference speed, our RAFM gets a frame rate of 33fps on a GPU with 7.5 M parameters, which is faster and more lightweight than the recent SOTA method (5fps and 128 M). The code is available athttps://github.com/jsczzzk/RAFM.
Xinnan Fan, Zhongkai Zhou, Yuanxue Xin
IEEE Signal Process. Lett.1
2021 R-MSFM: Recurrent Multi-Scale Feature Modulation for Monocular Depth Estimating
abstract
In this paper, we propose Recurrent Multi-Scale Feature Modulation (R-MSFM), a new deep network architecture for self-supervised monocular depth estimation. R-MSFM extracts per-pixel features, builds a multi-scale feature modulation module, and iteratively updates an inverse depth through a parameter-shared decoder at the fixed resolution. This architecture enables our R-MSFM to maintain semantically richer while spatially more precise representations and avoid the error propagation caused by the traditional U-Net-like coarse-to-fine architecture widely used in this domain, resulting in strong generalization and efficient parameter count. Experimental results demonstrate the superiority of our proposed R-MSFM both at model size and inference speed, and show the state-of-the-art results on the KITTI benchmark. Code is available at https://github.com/jsczzzk/R-MSFM
Zhongkai Zhou, Xinnan Fan, Yuanxue Xin
ICCV2
2020 Stable positioning for mobile targets using distributed fusion correction strategy of heterogeneous data
Gaifang Xin, Xinnan Fan, Chengming Luo, Hai Yang 0001, Xuewu Zhang 0001
Ad Hoc Networks2
2020 Spectral efficiency analysis of multi-cell multi-user massive MIMO over channel aging
abstract
To achieve the expected attractive potential of massive multiple‐input multiple‐output (MIMO), accurate match between the receiver and the actual channel is necessary. Due to the channel aging, this kind of mismatch appears in the vast majority of practical propagations. Consequently, the impact of channel aging for massive MIMO is investigated in this study. The authors first propose an equivalent estimated channel model by considering the inevitable time‐varying nature of channels in realistic mobile communications. Then, they obtain the achievable uplink rate based on the proposed equivalent channel models. In particular, the closed‐form expression of the asymptotic uplink rate is derived. Different from the existing works, this study considers the joint multi‐cell minimum‐mean‐squared‐error (MMSE) channel estimator and MMSE receiver for massive MIMO systems over aging channels. Through the simulation, it demonstrates that adopting massive MIMO can compensate for the performance loss caused by channel aging. In addition, this pstudy gives some insights into the system design: to achieve the best performance, system with higher‐mobility terminals should be operated with more pilot resources, and it can support more users, moreover, the optimal number of users supported in a cell is less than half of the coherence time interval.
Yuanxue Xin, Xinjiang Xia, Xinnan Fan
IET Commun.4
2018 Contrast Limited Adaptive Histogram Equalization-Based Fusion in YIQ and HSI Color Spaces for Underwater Image Enhancement
abstract
To improve contrast and restore color for underwater images without suffering from insufficient details and color cast, this paper proposes a fusion algorithm for different color spaces based on contrast limited adaptive histogram equalization (CLAHE). The original color image is first converted from RGB space to two different spaces: YIQ and HSI. Then, the algorithm separately applies CLAHE in YIQ and HSI color spaces to obtain two different enhanced images. After that, the YIQ and HSI enhanced images are respectively converted back to RGB space. When the three components of red, green, and blue are not coherent in the YIQ-RGB or HSI-RGB images, the three components will have to be harmonized with the CLAHE algorithm in RGB space. Finally, using a 4-direction Sobel edge detector in the bounded general logarithm ratio operation, a self-adaptive weight selection nonlinear image enhancement is carried out to fuse the YIQ-RGB and HSI-RGB images together to achieve the final image. The experimental results showed that the proposed algorithm provided more detail enhancement and higher values of color restoration than other image enhancement algorithms. The proposed algorithm can effectively reduce noise interference and observably improve the image quality for underwater images.
Jinxiang Ma, Xinnan Fan, Simon X. Yang, Xuewu Zhang 0001, Xifang Zhu
Int. J. Pattern Recognit. Artif. Intell.2
2018 A novel automatic dam crack detection algorithm based on local-global clustering
Xinnan Fan, Xuewu Zhang 0001, Yingjuan Xie
Multim. Tools Appl.1
2017 Positioning technology of mobile vehicle using self-repairing heterogeneous sensor networks
Chengming Luo, Wei Li 0223, Xinnan Fan, Hai Yang 0001, Jianjun Ni, Xuewu Zhang 0001, Gaifang Xin
J. Netw. Comput. Appl.3
2016 Multi-aperture anomaly detector for clutter background
abstract
Without priori information, anomaly detector has more important utility compared with supervised target detection. Many classical anomaly detectors have obtained perfect performance in many situations. However, there still have two problems which are correlated with accuracy of anomaly detector. Firstly, clutter background induced more and more difficult pixel which have moderate statistical difference. Then, ideal uncontaminated subset of clutter background is hard to be obtain which is used to estimate background model. Secondly, difference of spectral content of different background objects will effect salience of anomaly targets. And it is arbitrary that uncertain pixels is nominated as non-anomalies by one threshold. For above two problems, a multi-aperture anomaly detector is proposed in this paper. Without selection of anomaly-free pixels and accurate statistical model, the proposed anomaly detector is expected to decrease false alarm rate with clutter background. A multi-aperture division for hyperspectral cube is conducted by iterative process. Statistical data of ever subaperture will be named as basis, which represent spectral characteristic of a certain range of spectral cube. Then, anomaly salience is proposed to measure the difference between pixels and sub-aperture basis. On the other hand, continuity of membership value based on fuzzy logical theory is more suitable to nominate difficulty pixels which has moderate anomaly salience. At last defuzzification ruler can be used to fuse different detection results from multi-aperture.
Xinnan Fan, Xuewu Zhang 0001, Puhuang Li
IGARSS2
2014 A multi-sensor image fusion algorithm based on multi-scale feature analysis
abstract
The multi-sensor image fusion technology can obtain a more accurate and reliable image to understand the scene or recognize the target more easily. However, most existing algorithms are mainly based on optical images, which are highly susceptible by media interference, and cannot save the textural feature and color information at the same time. In view of these problems, this paper presents a multi-sensor image fusion algorithm based on multi-scale feature analysis. After preprocessing, the SAR image is divided into irregular area and regular area based on an adaptive segmentation method. Then, these two areas are separately decomposed using SIDWT algorithm to extract image texture. A multi-scale feature analysis rule is proposed to be used on wavelet coefficients of SAR and PAN image. The fused image is finally got after inverse wavelet transform with new wavelet coefficients. Experimental results showed the effectiveness of the proposed algorithm by subjective and objective evaluation.
Xinnan Fan, Bingbin Zheng, Xuewu Zhang 0001
IGARSS1