VLDB 2026 Research / reviewers in the wild / expert
Yuan Rao 0001
dblp:73/4103-1
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0003-1111-9210ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Out-of-distribution monocular depth estimation with local invariant regression
Yeqi Hu, Yuan Rao 0001, Hui Yu 0001, Gaige Wang, Hao Fan 0004, Wei Pang 0001, Junyu Dong |
Knowl. Based Syst. | 2 |
| 2025 | SeaDiff: Underwater Image Enhancement With Degradation-Aware Diffusion ModelabstractLight propagation in underwater scenes is significantly hindered by wavelength- and distance-dependent attenuation and scattering, leading to low contrast and severe color distortion in underwater images. Recent advancements in diffusion models have shown impressive performance in image restoration by learning data distribution prior knowledge (diffusion prior) from large amounts of paired data. However, due to the difficulties in collecting paired underwater images, the available data for underwater image enhancement is limited in both quality and quantity. This scarcity leads to a biased diffusion prior and suboptimal performance of diffusion models. To address this issue, we propose a novel method, termed SeaDiff, to learn underwater diffusion prior with wavelength- and distance-dependent degradation awareness. Specifically, we introduce a Prior Knowledge Mining Model (PKMM), which includes two key components: (1) the Physical Prior Embedding Module (PPEM) that simulates the underwater imaging process through a distance-dependent physical model and embeds physical prior by incorporating generalizable distance-aware cues from a large vision foundation model; and (2) the Color Prior Embedding Module (CPEM) that extracts wavelength-dependent color distribution prior from a log-chroma color space. Additionally, we propose a Degradation-Aware Diffusion Model (DADM) that seamlessly integrates degradation prior with diffusion prior and enhances the underwater images with high visual quality. Extensive experiments on popular UIE benchmarks and downstream tasks demonstrate that the proposed SeaDiff achieves state-of-the-art performance in terms of both visual quality and quantitative metrics. The code will be released at https://github.com/Henry-Bi/SeaDiff. Hengyue Bi, Long Chen 0019, Jingchao Cao, Jinghao Sun, Yuan Rao 0001, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Learning Semantic-Aware Point-Line Features for Localization and ReconstructionabstractHigh-precision image matching and localization technology in a 3D environment map is essential for many tasks, such as marine engineering detection, robotics, and autonomous navigation. However, current visual localization and reconstruction methods overly depend on point features, which lack robustness in low-texture environments. To address this limitation, we propose a novel framework for point and line localization and 3D reconstruction with semantic constraints, which integrates multiple innovative components to achieve superior performance. Firstly, we design a point-localization optimization strategy with uniform point sampling and point-based instance segmentation constraints, significantly improving image matching and camera localization accuracy. Secondly, we optimize the selection of 2D-3D lines and line matching using instance segment constraints, leveraging the structural and semantic richness of line features to complement point features. Thirdly, we perform a joint point and line feature 3D reconstruction, enabling the creation of accurate 3D environment maps even in challenging low-texture marine scenes.Our approach has been extensively tested on popular datasets and compared with state-of-the-art methods. This work significantly advances current visual localization and 3D reconstruction techniques by addressing their limitations in low-texture environments, while also providing a robust foundation for future research and applications in marine engineering, robotics, and autonomous navigation. Jian Yang 0036, Yuan Rao 0001, Hao Fan 0004, Junyu Dong, Hui Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Aerial Multiview Stereo via Adaptive Depth Range Inference and Normal CuesabstractThree-dimensional digital urban reconstruction from multi-view aerial images is a critical application where deep multi-view stereo (MVS) methods outperform traditional techniques. However, existing methods commonly overlook the key differences between aerial and close-range settings, such as varying depth ranges along epipolar lines and insensitive feature-matching associated with low-detailed aerial images. To address these issues, we propose an Adaptive Depth Range MVS (ADR-MVS), which integrates monocular geometric cues to improve multi-view depth estimation accuracy. The key component of ADR-MVS is the depth range predictor, which generates adaptive range maps from depth and normal estimates using cross-attention discrepancy learning. In the first stage, the range map derived from monocular cues breaks through predefined depth boundaries, improving feature-matching discriminability and mitigating convergence to local optima. In later stages, the inferred range maps are progressively narrowed, ultimately aligning with the cascaded MVS framework for precise depth regression. Moreover, a normal-guided cost aggregation operation is specially devised for aerial stereo images to improve geometric awareness within the cost volume. Finally, we introduce a normal-guided depth refinement module that surpasses existing RGB-guided techniques. Experimental results demonstrate that ADR-MVS achieves state-of-the-art performance on the WHU, LuoJia-MVS, and München datasets, while exhibits superior computational complexity. Yimei Liu, Yakun Ju, Yuan Rao 0001, Hao Fan 0004, Junyu Dong, Feng Gao 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MLNet: An multi-scale line detector and descriptor network for 3D reconstruction
Jian Yang 0036, Yuan Rao 0001, Eric Rigall, Hao Fan 0004, Junyu Dong, Hui Yu 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Deep Color Compensation for Generalized Underwater Image EnhancementabstractUnderwater images suffer from quality degradation due to the underwater light absorption and scattering. It remains challenging to enhance underwater images using deep learning-based methods since the scarcity of real-world underwater images and their enhanced counterparts. Although existing works manually select well-enhanced images as reference images to train enhancement networks in an end-to-end manner, their performance tends to be inferior in some scenarios. We argue that the manually selected reference images cannot approximate their ground truth perfectly, leading to imbalanced learning and domain shift in enhancement networks. To address this issue, we analyse widely used underwater datasets from the perspective of color spectrum distribution and surprisingly find the sound color spectrum distribution of the enhanced reference images compared to in-air datasets. Based on this perceptive observation, instead of directly learning the enhancement mapping, we propose a novel methodology to learn color compensation for general purposes. Specifically, we present a probabilistic color compensation network that estimates the probabilistic distribution of colors by multi-scale volumetric fusion of texture and color features. We further propose a novel two-stage enhancement framework that first performs color compensation and then enhancement, which is highly flexible to be integrated with an existing enhancement method without tuning. Extensive experiments on underwater image enhancement across various challenging scenarios show that our proposed approach consistently improves the results of the popular conventional and learning-based methods by a significant margin. Moreover, our enhanced images achieve superior performance on underwater salient object detection and visual 3D reconstruction, demonstrating that our method can successfully break through the generalization bottleneck of existing learning-based enhancement models. Our implementation will be made available at https://github.com/Ray2OUC/P2CNet. Yuan Rao 0001, Kunqian Li, Hao Fan 0004, Sen Wang 0002, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Learning Deep Photometric Stereo Network with Reflectance PriorsabstractPhotometric stereo recovers the surface normals of an object from images with varying shading cues. Conventional photometric stereo methods attempt to use handcrafted reflectance models to approximate surface normals, while deep learning-based networks have shown a much more powerful ability to handle non-Lambertian objects. However, none of the existing deep learning methods explores how prior reflectance information can be used to optimize surface-normal prediction. In this paper, we first present the introduction of reflectance prior to deep photometric stereo models. Our explorations include how the reflectance prior can simplify the optimization of deep networks by reparametrizing the weights, and (2) eliminate the impacts of surfaces with spatially varying reflectance for all-pixel input photometric stereo methods. To achieve these goals, we propose a residual fusion module (RFM) in our method, which explicitly extracts features useful for surface-normal recovery and removes those features influenced by reflectance. Additionally, we design a shading extractor with multi-scale and global-local feature fusion operations, which can fuse features with different receptive fields and better utilize the non-maximum features missing in the max-pooling operation. Experiments and ablation studies verify the accuracy and effectiveness of the proposed reflectance prior network on a widely used benchmark. Yakun Ju, Songsong Huang, Yuan Rao 0001, Kin-Man Lam 0001 |
ICME | 4 |
| 2023 | Gated-Cross Aggregation Network for Hyperspectral and LiDAR Data ClassificationabstractExisting hyperspectral image (HSI) and LiDAR data joint classification methods commonly treat LiDAR data equally with HSI in the network. As a result, these methods may fail to effectively leverage the spectral features from HSI and elevation information from LiDAR. In this paper, we show that better cross-modal alignments can be achieved through an HSI encoder for jointly embedding elevation features from LiDAR during spectral feature encoding. To this end, we propose a Gated-Cross Aggregation Network (GCA-Net) to fully investigate the complementary clues hidden in multi-source data progressively. The beneficial spectral-elevation cues are then exploited by cross-attention feature fusion. Then, useful LiDAR features are integrated into the HSI features via an elevation gating module, which occurs at each stage of the network. Experimental results on the Houston 2013 dataset and Trento dataset reveal that the proposed GCA-Net achieves better performance than several closely related methods. Xiaochen Shi, Junyan Lin, Yuan Rao 0001, Feng Gao 0005 |
IGARSS | 3 |
| 2023 | Learning General Descriptors for Image Matching With Regression FeedbackabstractRecent advances on feature descriptors for image matching put more emphasis on encoding invariances (e.g. illumination invariance) to promote the descriptors’ discriminative power. However, according to the information entropy, more invariance implies greater certainty and less informativeness in a descriptor. Consequently, descriptors encoding too many invariances usually show poor generalization to unknown image changes, lacking enough informativeness to cover the large uncertainty in unseen scenes. This limits the application scenarios of learned descriptors. In this paper, we propose to alleviate this issue from the perspective of informativeness and we thus design hierarchical consistent constraint by introducing regression feedback in a self-supervised manner. Combined with the hardest-within-batch matching constraint, we form a novel dual supervision framework, to encourage the descriptor to learn an informative representation while maintaining a good discriminative power. Moreover, to fully mine the context information hidden in image and boost the informativeness in turn, we present AANet, a descriptor network that efficiently predicts dense description by the powerful Attentional Aggregation of multi-level features. Experiments across challenging feature matching on HPatches, RDNIM datasets, and visual localization tasks on Aachen Day-night dataset show that our method outperforms recent state-of-the-art descriptors while keeping encouraging efficiency. The application of visual 3D reconstruction on various scenarios also demonstrates the high generalization ability of our method. Yuan Rao 0001, Yakun Ju, Eric Rigall, Jian Yang 0036, Hao Fan 0004, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | A deep-shallow and global-local multi-feature fusion network for photometric stereo
Yanru Liu, Yakun Ju, Muwei Jian, Feng Gao 0005, Yuan Rao 0001, Yeqi Hu, Junyu Dong |
Image Vis. Comput. | 5 |
| 2022 | Near-field photometric stereo using a ring-light imaging device
Hao Fan 0004, Yuan Rao 0001, Eric Rigall, Lin Qi 0004, Zhile Wang, Junyu Dong |
Signal Process. Image Commun. | 2 |