Xianfa Xu

dblp:237/3929 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2022
0000-0002-8538-8280ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 50% Segmentation and scene understanding · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
depth estimation
0.512021
Multi-Scale Spatial Attention-Guided Monocular Depth Estimation With Semantic Enhancement · IEEE Trans. Image Process. 2021
Computer vision › Segmentation and scene understanding › semantic segmentation
joint depth and semantic prediction
0.512021
Multi-Scale Spatial Attention-Guided Monocular Depth Estimation With Semantic Enhancement · IEEE Trans. Image Process. 2021
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.512021
Multi-Scale Spatial Attention-Guided Monocular Depth Estimation With Semantic Enhancement · IEEE Trans. Image Process. 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.512021
Multi-Scale Spatial Attention-Guided Monocular Depth Estimation With Semantic Enhancement · IEEE Trans. Image Process. 2021

Methods — techniques the papers use, named apart from their topics

spatial attention · 0.5mutual information · 0.5atrous spatial pyramid pooling · 0.5
YearPublicationVenuePosition
2022 Deep mutual information multi-view representation for visual recognition
Xianfa Xu, Zhe Chen 0005, Fuliang Yin
Appl. Intell.1
2021 Monocular Depth Estimation With Multi-Scale Feature Fusion
abstract
Depth estimation from a single image is a crucial but challenging task for reconstructing 3D structures and inferring scene geometry. However, most existing methods fail to extract more detailed information and estimate the distant small-scale objects well. In this paper, we propose a monocular depth estimation based on multi-scale feature fusion. Specifically, to obtain input features of different scales, we first feed the input images of different scales to pre-trained residual networks with sharing weights. Then, an attention mechanism is used to learn the salient features at different scales, which can integrate detailed information at large scale feature maps and scene information at small scale feature maps. Furthermore, inspired by the dense atrous spatial pyramid pooling in semantic segmentation, we build a multi-scale feature fusion dense pyramid to further improve the ability of the feature extraction. Last, a scale-invariant error loss is used to predict depth maps in log space. We evaluate our method on several public benchmark datasets (including NYU Depth V2 and KITTI). The experiment results show that the proposed method obtains better performance than the existing methods and achieves state-of-the-art results.
Xianfa Xu, Zhe Chen 0005, Fuliang Yin
IEEE Signal Process. Lett.1
2021 Multi-Scale Spatial Attention-Guided Monocular Depth Estimation With Semantic Enhancement
abstract
Depth estimation from single monocular image is a vital but challenging task in 3D vision and scene understanding. Previous unsupervised methods have yielded impressive results, but the predicted depth maps still have several disadvantages such as missing small objects and object edge blurring. To address these problems, a multi-scale spatial attention guided monocular depth estimation method with semantic enhancement is proposed. Specifically, we first construct a multi-scale spatial attention-guided block based on atrous spatial pyramid pooling and spatial attention. Then, the correlation between the left and right views is fully explored by mutual information to obtain a more robust feature representation. Finally, we design a double-path prediction network to simultaneously generate depth maps and semantic labels. The proposed multi-scale spatial attention-guided block can focus more on the objects, especially on small objects. Moreover, the additional semantic information also enables the objects edge in the predicted depth maps more sharper. We conduct comprehensive evaluations on public benchmark datasets, such as KITTI and Make3D. The experiment results well demonstrate the effectiveness of the proposed method and achieve better performance than other self-supervised methods.
Xianfa Xu, Zhe Chen 0005, Fuliang Yin
IEEE Trans. Image Process.1