Hao Fan 0004

dblp:23/2535-4 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0001-9133-8135ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Part-aware cross-integration transformer with proximity guided regularization for domain generalizable animal re-identification
Zeyuan Sun, Junyu Dong, Xiaowei Zhou 0003, Huiyu Zhou 0001, Hao Fan 0004
Expert Syst. Appl.5
2025 FNIN: A Fourier Neural Operator-based Numerical Integration Network for Surface-from-gradients
abstract
Surface-from-gradients (SfG) aims to recover a three-dimensional (3D) surface from its gradients. Traditional methods encounter significant challenges in achieving high accuracy and handling high-resolution inputs, particularly facing the complex nature of discontinuities and the inefficiencies associated with large-scale linear solvers. Although recent advances in deep learning, such as photometric stereo, have enhanced normal estimation accuracy, they do not fully address the intricacies of gradient-based surface reconstruction. To overcome these limitations, we propose a Fourier neural operator-based Numerical Integration Network (FNIN) within a two-stage optimization framework. In the first stage, our approach employs an iterative architecture for numerical integration, harnessing an advanced Fourier neural operator to approximate the solution operator in Fourier space. Additionally, a self-learning attention mechanism is incorporated to effectively detect and handle discontinuities. In the second stage, we refine the surface reconstruction by formulating a weighted least squares problem, addressing the identified discontinuities rationally. Extensive experiments demonstrate that our method achieves significant improvements in both accuracy and efficiency compared to current state-of-the-art solvers. This is particularly evident in handling high-resolution images with complex data, achieving errors of fewer than 0.1 mm on tested objects.
Jiaqi Leng 0002, Yakun Ju, Yuanxu Duan, Jiangnan Zhang, Qingxuan Lv, Zuxuan Wu, Hao Fan 0004
AAAI7
2025 Out-of-distribution monocular depth estimation with local invariant regression
Yeqi Hu, Yuan Rao 0001, Hui Yu 0001, Gaige Wang, Hao Fan 0004, Wei Pang 0001, Junyu Dong
Knowl. Based Syst.5
2025 Real-World Multi-View Stereo via Learning RGB-D Structural Consistency From Depth Super-Resolution
abstract
Learning-based Multi-View Stereo (MVS) methods, typically reliant on cascaded cost volume formulations, perform well on small-scale scenes. However, as the depth range of captured images becomes broader and more varied, the coarse-to-fine depth sampling process, which depends solely on feature matching, is increasingly prone to local optima. Despite recent advancements in feature representation, depth sampling patterns, and cost aggregation techniques, challenges related to model generalization and computational efficiency persist. In this paper, we propose SR-MVSNet, a novel framework that integrates multi-view feature matching and RGB-D cross-modal structural consistency learning to achieve high-quality 3D reconstruction. Our approach begins with the construction of Low-Resolution (LR) cost volumes for initial LR depth estimation, which are then enhanced to full-resolution via a tailored uncertainty-aware guided depth super-resolution module. To ensure cross-view consistency, the depth maps undergo further refinement through multi-view feature matching. By avoiding high-resolution cost volume processing, our framework improves depth estimation robustness and efficiency. Additionally, we introduce an iterative depth fusion post-processing strategy during inference to improve reconstruction in ambiguous matching regions, a critical challenge for MVS methods. Experiments show that our method achieves top-3 performance on the DTU and Tanks & Temples datasets and ranks first on the ETH3D dataset. Furthermore, it uses significantly fewer GPU resources than most high performing methods, offering a favorable trade-off between reconstruction quality and computational efficiency.
Yimei Liu, Jingchao Cao, Hao Fan 0004, Junyu Dong, Sheng Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Learning Semantic-Aware Point-Line Features for Localization and Reconstruction
abstract
High-precision image matching and localization technology in a 3D environment map is essential for many tasks, such as marine engineering detection, robotics, and autonomous navigation. However, current visual localization and reconstruction methods overly depend on point features, which lack robustness in low-texture environments. To address this limitation, we propose a novel framework for point and line localization and 3D reconstruction with semantic constraints, which integrates multiple innovative components to achieve superior performance. Firstly, we design a point-localization optimization strategy with uniform point sampling and point-based instance segmentation constraints, significantly improving image matching and camera localization accuracy. Secondly, we optimize the selection of 2D-3D lines and line matching using instance segment constraints, leveraging the structural and semantic richness of line features to complement point features. Thirdly, we perform a joint point and line feature 3D reconstruction, enabling the creation of accurate 3D environment maps even in challenging low-texture marine scenes.Our approach has been extensively tested on popular datasets and compared with state-of-the-art methods. This work significantly advances current visual localization and 3D reconstruction techniques by addressing their limitations in low-texture environments, while also providing a robust foundation for future research and applications in marine engineering, robotics, and autonomous navigation.
Jian Yang 0036, Yuan Rao 0001, Hao Fan 0004, Junyu Dong, Hui Yu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Aerial Multiview Stereo via Adaptive Depth Range Inference and Normal Cues
abstract
Three-dimensional digital urban reconstruction from multi-view aerial images is a critical application where deep multi-view stereo (MVS) methods outperform traditional techniques. However, existing methods commonly overlook the key differences between aerial and close-range settings, such as varying depth ranges along epipolar lines and insensitive feature-matching associated with low-detailed aerial images. To address these issues, we propose an Adaptive Depth Range MVS (ADR-MVS), which integrates monocular geometric cues to improve multi-view depth estimation accuracy. The key component of ADR-MVS is the depth range predictor, which generates adaptive range maps from depth and normal estimates using cross-attention discrepancy learning. In the first stage, the range map derived from monocular cues breaks through predefined depth boundaries, improving feature-matching discriminability and mitigating convergence to local optima. In later stages, the inferred range maps are progressively narrowed, ultimately aligning with the cascaded MVS framework for precise depth regression. Moreover, a normal-guided cost aggregation operation is specially devised for aerial stereo images to improve geometric awareness within the cost volume. Finally, we introduce a normal-guided depth refinement module that surpasses existing RGB-guided techniques. Experimental results demonstrate that ADR-MVS achieves state-of-the-art performance on the WHU, LuoJia-MVS, and München datasets, while exhibits superior computational complexity.
Yimei Liu, Yakun Ju, Yuan Rao 0001, Hao Fan 0004, Junyu Dong, Feng Gao 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Consistent Image Inpainting With Pre-Perception and Cross-Perception Collaborative Processes
abstract
It has been proven that introducing multiple guidance sources boosts image inpainting performance. However, existing methods primarily focus on local relationships and neglect the holistic interplay between guidance and texture information. Moreover, they lack an effective feedback mechanism to adaptively update the guidance process as corrupted texture information is progressively restored, potentially resulting in inconsistent inpainting. To tackle this issue, we propose a novel scheme aligned with pre-perception and cross-perception collaborative processes in human drawing. To mimic the pre-perception process, we introduce a pre-perceptual transformer block that captures long-range contextual dependencies and activates meaningful information to individually optimize image structures, semantic layouts, and textures, thereby effectively controlling their respective generation. To mimic the cross-perception collaborative process, we propose a cyclic cross-perceptual interaction to maintain consistency across the entire image regarding structure, layout, and texture while progressively refining their details. This interaction accounts for the global attention relationship between texture and other guidance sources (including image structure and semantic layout) to enhance image texture, alongside integrating a dedicated feedback mechanism to update guidance information. The proposed components are alternately deployed in three-branch decoders of the new scheme from rough to fine-grained levels to achieve these two iterative processes of human drawing. Experimental results prove the superiority of the proposed scheme over state-of-the-art methods across three datasets.
Yongle Zhang 0001, Yimin Liu 0001, Hao Fan 0004, Ruotong Hu, Jian Zhang 0002, Qiang Wu 0001
IEEE Trans. Image Process.3
2024 GaitMA: Pose-guided Multi-modal Feature Fusion for Gait Recognition
abstract
Gait recognition is a biometric technology that recognizes the identity of humans through their walking patterns. Existing appearance-based methods utilize CNN or Transformer to extract spatial and temporal features from silhouettes, while model-based methods employ GCN to focus on the special topological structure of skeleton points. However, the quality of silhouettes is limited by complex occlusions, and skeletons lack dense semantic features of the human body. To tackle these problems, we propose a novel gait recognition framework, dubbed Gait Multi-model Aggregation Network (GaitMA), which effectively combines two modalities to obtain a more robust and comprehensive gait representation for recognition. First, skeletons are represented by joint/limb-based heatmaps, and features from silhouettes and skeletons are respectively extracted using two CNN-based feature extractors. Second, a co-attention alignment module is proposed to align the features by element-wise attention. Finally, we propose a mutual learning module, which achieves feature fusion through cross-attention, Wasser-stein loss is further introduced to ensure the effective fusion of two modalities. Extensive experimental results demonstrate the superiority of our model on Gait3D, OU-MVLP, and CASIA-B.
Fanxu Min, Shaoxiang Guo, Hao Fan 0004, Junyu Dong
ICME3
2024 Learning context-aware local feature descriptors for 3D reconstruction
Jian Yang 0036, Hao Fan 0004, Junyu Dong, Hui Yu 0001
Neurocomputing3
2024 MLNet: An multi-scale line detector and descriptor network for 3D reconstruction
Jian Yang 0036, Yuan Rao 0001, Eric Rigall, Hao Fan 0004, Junyu Dong, Hui Yu 0001
Knowl. Based Syst.5
2024 Geometry-Enhanced Attentive Multi-View Stereo for Challenging Matching Scenarios
abstract
Deep networks have made remarkable progress in Multi-View Stereo (MVS) task in recent years. However, the problem of finding accurate correspondences across different views under ill-posed matching situations remains unresolved and crucial. To address this issue, this paper proposes a Geometry-enhanced Attentive Multi-View Stereo (GA-MVS) network, which can access multi-view consistent feature representation and achieve accurate depth estimation in challenging situations. Specifically, we propose a geometry-enhanced feature extractor to explore illumination-invariant geometric features and incorporate them with common texture features to improve matching accuracy when dealing with view-dependent photometric effects, such as shadow and specularity. Then, we design a novel attentive learning framework to explore per-pixel adaptive supervision, effectively improving the depth estimation performance of textureless regions. The experimental results on the DTU and Tanks & Temples benchmarks demonstrate that our method achieves state-of-the-art results compared to other advanced MVS models.
Yimei Liu, Jian Yang 0036, Hao Fan 0004, Junyu Dong, Sheng Chen 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Deep Color Compensation for Generalized Underwater Image Enhancement
abstract
Underwater images suffer from quality degradation due to the underwater light absorption and scattering. It remains challenging to enhance underwater images using deep learning-based methods since the scarcity of real-world underwater images and their enhanced counterparts. Although existing works manually select well-enhanced images as reference images to train enhancement networks in an end-to-end manner, their performance tends to be inferior in some scenarios. We argue that the manually selected reference images cannot approximate their ground truth perfectly, leading to imbalanced learning and domain shift in enhancement networks. To address this issue, we analyse widely used underwater datasets from the perspective of color spectrum distribution and surprisingly find the sound color spectrum distribution of the enhanced reference images compared to in-air datasets. Based on this perceptive observation, instead of directly learning the enhancement mapping, we propose a novel methodology to learn color compensation for general purposes. Specifically, we present a probabilistic color compensation network that estimates the probabilistic distribution of colors by multi-scale volumetric fusion of texture and color features. We further propose a novel two-stage enhancement framework that first performs color compensation and then enhancement, which is highly flexible to be integrated with an existing enhancement method without tuning. Extensive experiments on underwater image enhancement across various challenging scenarios show that our proposed approach consistently improves the results of the popular conventional and learning-based methods by a significant margin. Moreover, our enhanced images achieve superior performance on underwater salient object detection and visual 3D reconstruction, demonstrating that our method can successfully break through the generalization bottleneck of existing learning-based enhancement models. Our implementation will be made available at https://github.com/Ray2OUC/P2CNet.
Yuan Rao 0001, Kunqian Li, Hao Fan 0004, Sen Wang 0002, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.4
2023 Learning General Descriptors for Image Matching With Regression Feedback
abstract
Recent advances on feature descriptors for image matching put more emphasis on encoding invariances (e.g. illumination invariance) to promote the descriptors’ discriminative power. However, according to the information entropy, more invariance implies greater certainty and less informativeness in a descriptor. Consequently, descriptors encoding too many invariances usually show poor generalization to unknown image changes, lacking enough informativeness to cover the large uncertainty in unseen scenes. This limits the application scenarios of learned descriptors. In this paper, we propose to alleviate this issue from the perspective of informativeness and we thus design hierarchical consistent constraint by introducing regression feedback in a self-supervised manner. Combined with the hardest-within-batch matching constraint, we form a novel dual supervision framework, to encourage the descriptor to learn an informative representation while maintaining a good discriminative power. Moreover, to fully mine the context information hidden in image and boost the informativeness in turn, we present AANet, a descriptor network that efficiently predicts dense description by the powerful Attentional Aggregation of multi-level features. Experiments across challenging feature matching on HPatches, RDNIM datasets, and visual localization tasks on Aachen Day-night dataset show that our method outperforms recent state-of-the-art descriptors while keeping encouraging efficiency. The application of visual 3D reconstruction on various scenarios also demonstrates the high generalization ability of our method.
Yuan Rao 0001, Yakun Ju, Eric Rigall, Jian Yang 0036, Hao Fan 0004, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.6
2022 Spatio-Temporal Representation Learning with Social Tie for Personalized POI Recommendation
abstract
Abstract Recommending a limited number of Point-of-Interests (POIs) a user will visit next has become increasingly important to both users and POI holders for Location-Based Social Networks (LBSNs). However, POI recommendation is a challenging task since complex sequential patterns and rich contexts are contained in extremely sparse user check-in data. Recent studies show that embedding techniques effectively incorporate POI contextual information to alleviate the data sparsity issue, and Recurrent Neural Network (RNN) has been successfully employed for sequential prediction. Nevertheless, existing POI recommendation approaches are still limited in capturing user personalized preference due to separate embedding learning or network modeling. To this end, we propose a novel unified spatio-temporal neural network framework, named PPR, which leverages users’ check-in records and social ties to recommend personalized POIs for querying users by joint embedding and sequential modeling. Specifically, PPR first learns user and POI representations by joint modeling User-POI relation, sequential patterns, geographical influence, and social ties in a heterogeneous graph and then models user personalized sequential patterns using the designed spatio-temporal neural network based on LSTM model for the personalized POI recommendation. Furthermore, we extend PPR to an end-to-end recommendation model by jointly learning node representations and modeling user personalized sequential preference. Extensive experiments on three real-world datasets demonstrate that our model significantly outperforms state-of-the-art baselines for successive POI recommendation in terms of Accuracy, Precision, Recall and NDCG. The source code is available at: https://www.anonymous.4open.science/r/DSE-1BEC .
Shaojie Dai, Yanwei Yu, Hao Fan 0004, Junyu Dong
Data Sci. Eng.3
2022 Near-field photometric stereo using a ring-light imaging device
Hao Fan 0004, Yuan Rao 0001, Eric Rigall, Lin Qi 0004, Zhile Wang, Junyu Dong
Signal Process. Image Commun.1
2021 Personalized POI Recommendation: Spatio-Temporal Representation Learning with Social Tie
Shaojie Dai, Yanwei Yu, Hao Fan 0004, Junyu Dong
DASFAA (1)3
2020 A joint guidance-enhanced perceptual encoder and atrous separable pyramid-convolutions for image inpainting
Yongle Zhang 0001, Yingyu Wang, Junyu Dong, Lin Qi 0004, Hao Fan 0004, Xinghui Dong, Muwei Jian, Hui Yu 0001
Neurocomputing5
2018 Dynamic 3D Surface Reconstruction Using a Hand-Held Camera
abstract
This paper proposes a dynamic 3D reconstruction method for recovering a surface shape from a set of images that are captured by a hand-held camera. A light source is attached to the camera as a photometric constraint. Thus, we can effectively calculate photometric stereo using the relative moving camera. The key contributions of our work are a robust pixel matching method to build effective correspondences between images for normal estimation, and an optimization method to correct the deviation in the recovered surface shape that is caused by the nonideal illumination in a close-range lighting condition. Specially we correct the recovered shape by adding an interpolation surface that is estimated using sparse control points from the structure from motion. The effectiveness of our method is verified on real datasets with a digital camera and a smart phone.
Hao Fan 0004, Lin Qi 0004, Junyu Dong, Gongfa Li, Hui Yu 0001
IECON1
2016 Robust Photometric Stereo in a scattering medium via Low-Rank Matrix Completion and Recovery
abstract
Photometric Stereo is a popular method for 3D reconstruction from images due to its high level of details handling. However, when it is used in a scattering medium such as lakes and oceans, the recovery result will be negatively impacted by the light absorption, light scattering and the impurities in the water. In this paper, we present a new method to solve the problem of better 3D reconstruction via Low-Rank Matrix Completion and Recovery. First, we use the dark points, like shadows and darkness in the water to fit the scattering effect distribution and then remove the scattering from the image. Next, we use the Robust Principal Component Analysis method (RPCA) to recover the image by removing the sparse noise including shadows, impurities and some corrupted points caused by backscatter compensation. Finally, we combine the RPCA results and the least-squares (LS) results to get the surface normal and accomplish the 3D reconstruction. Extensive experimental results demonstrate that our method achieves more accurate estimates of surface normal and 3D reconstruction than previous techniques.
Hao Fan 0004, Yisong Luo, Lin Qi 0004, Junyu Dong, Hui Yu 0001
HSI1