Yanlong Cao

dblp:122/1650 · DBLP profile ↗
← Back
31ranked-venue papers
2as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Rotational accuracy modeling of gas hydrostatic bearings under multi-source deviations using Copula-PCA joint optimization
Zhaozhe Huang, Yanlong Cao
Adv. Eng. Informatics4
2026 HNI-SLAM: Neural implicit SLAM for high quality reconstruction
Lina Duan, Yejun Shou, Lingfeng Shen, Yanlong Cao
Comput. Graph.7
2026 RGD-SLAM: Robust Gaussian splatting SLAM for dynamic environments
Yejun Shou, Lingfeng Shen, Yanlong Cao
Pattern Recognit.5
2026 Unsupervised Large-Scale Point Cloud Registration via Spherical Projection Consistency
abstract
In real-world applications, large-scale outdoor Li- DAR point cloud registration faces significant challenges, including data sparsity, occlusions, and reliance on expensive ground-truth pose labels. To address these issues, this paper proposes an unsupervised registration framework based on spherical projection consistency. Specifically, both the input point cloud and its spatially transformed counterpart are projected into range images, and their spatial consistency is exploited as a supervision signal. Furthermore, a multi-scale patch-topatch feature fusion module is introduced to effectively integrate features from point clouds and range images, thereby enhancing feature discriminability. In addition, a dynamic masking strategy is applied in the range image domain to mitigate the impact of sparsity variations and further strengthen the spatial consistency of the supervision signal. Extensive experiments on the KITTI and NuScenes datasets demonstrate that the proposed method outperforms state-of-the-art unsupervised approaches, validating its effectiveness and robustness in large-scale outdoor scenarios.
Yejun Shou, Shuai Liu 0009, Zhijie Xu, Yanlong Cao
IEEE Signal Process. Lett.5
2026 SepViT: A Dual-Path Transformer-Convolution Framework for Apex Frame-Based Microexpression Recognition
Hafiz Khizer Bin Talib, Yanlong Cao, Muhammad Zaman, Kaiwei Xu, Adnan Akhunzada
IEEE Trans. Comput. Soc. Syst.2
2026 Cost-Efficient Open Vocabulary 3D Scene Understanding Based on Semantic Probability
abstract
Traditional 3D scene understanding methods heavily depend on 3D annotation and training, which allow for the identification of seen classes but struggle to recognize unseen classes. In this paper, we leverage the open vocabulary inference capabilities of pre-trained models, enabling the encoding of open vocabulary concepts. However, unlike existing open vocabulary 3D scene understanding methods, we propose a framework based on semantic probability. This innovation significantly reduces computational cost and is compatible with state-of-the-art two-stage 2D pre-trained models. Specifically, we align the text features from the CLIP model with the pixel features from the 2D pre-trained models, inferring semantic probability of image pixels based on similarity and projecting it onto 3D points. Subsequently, we introduce a point cloud pairs semantic fusion method to merge the point clouds, reducing the semantic probability of erroneous 3D points. Based on probability scores, we achieve 3D semantic segmentation on open vocabularies without any supervision or training. In addition, the semantic probability of 3D points can serve as pseudo-labels for 3D distillation, and the geometric features of the 3D scene can be exploited to improve the segmentation performance. Experimental results demonstrate that the proposed method exhibits competitive performance on publicly available benchmark datasets, including ScanNet, Matterport3D, and nuScenes.
Lingfeng Shen, Xiaoyao Wei, Gang Pan 0001, Yanlong Cao
IEEE Trans. Image Process.5
2026 Fast Photometric Stereo by Time- and Spectral-Multiplexing With Crosstalk Handling
abstract
Photometric stereo is widely used to recover detailed surface normals. However, previous methods fail to balance the accuracy and efficiency. Conventional photometric stereo achieves high accuracy but suffers from low efficiency due to spectral-multiplexing and inefficient algorithms. In contrast, multispectral photometric stereo captures images efficiently with spectral-multiplexing, but its accuracy is harmed by crosstalk. In this paper, we aim to resolve the crosstalk issue to achieve fast photometric stereo (FPS) at low cost. First, we analyze the formulation and impact of crosstalk, showing that it significantly affects normal estimation, with external factors being primary contributors to crosstalk and internal factors being the secondary. Subsequently, we propose the FPS framework with a fast data capture scheme that combines time- and spectral-multiplexing to introduce constraints on crosstalk regarding both internal and external factors, along with a lightweight network, FPS-Net, to remove crosstalk caused by those factors based on constraints under such scheme. Finally, we build a real-world crosstalk-affected FPS dataset to evaluate the performance in handling crosstalk for normal estimation. Experimental results show the superior accuracy and efficiency of our method. The code and dataset are available at https://github.com/wxy-zju/FPS-Net.
Xiaoyao Wei, Lingfeng Shen, Zhijie Xu, Yanlong Cao
IEEE Trans. Image Process.5
2025 Unsupervised Rgb-D Point Cloud Registration for Scenes With Low Overlap and Photometric Inconsistency
Yejun Shou, Lingfeng Shen, Gang Pan 0001, Yanlong Cao
ICCV6
2025 Enhanced multi-scale feature adaptive fusion sparse convolutional network for large-scale scenes semantic segmentation
Lingfeng Shen, Yanlong Cao, Yejun Shou, Zhijie Xu
Comput. Graph.2
2025 Weakly supervised point cloud semantic segmentation using pseudo-label reliability and consistency regularization
Lingfeng Shen, Yanlong Cao, Xiaoyao Wei
Neurocomputing2
2025 Unleashing powerful generalization for point cloud registration
Yejun Shou, Lingfeng Shen, Yanlong Cao
Knowl. Based Syst.5
2025 Revisiting Supervised Learning-Based Photometric Stereo Networks
abstract
Deep learning has significantly propelled the development of photometric stereo by handling the challenges posed by unknown reflectance and global illumination effects. However, how supervised learning-based photometric stereo networks resolve these challenges remains to be elucidated. In this paper, we aim to reveal how existing methods address these challenges by revisiting their deep features, deep feature encoding strategies, and network architectures. Based on the insights gained from our analysis, we propose ESSENCE-Net, which effectively encodes deep shading features with an easy-first-encoding strategy, enhances shading features with shading supervision, and accurately decodes normal with spatial context-aware attention. The experimental results verify that the proposed method outperforms state-of-the-art methods on three benchmark datasets, whether with dense or sparse inputs.
Xiaoyao Wei, Zongrui Li 0001, Binjie Ding, Boxin Shi, Xudong Jiang 0001, Gang Pan 0001, Yanlong Cao
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 iS-MAP: Neural Implicit Mapping and Positioning for Structural Environments
Yanlong Cao, Yejun Shou, Lingfeng Shen, Xiaoyao Wei, Zhijie Xu
ACCV (9)2
2024 Structerf-SLAM: Neural implicit representation SLAM for structural environments
Yanlong Cao, Xiaoyao Wei, Yejun Shou, Lingfeng Shen, Zhijie Xu
Comput. Graph.2
2023 Light field angular super-resolution based on structure and scene information
Jiangxin Yang, Lingyu Wang 0005, Lifei Ren, Yanpeng Cao, Yanlong Cao
Appl. Intell.5
2023 A sub-region Unet for weak defects segmentation with global information and mask-aware loss
Jiangxin Yang, Yanlong Cao, Guizhong Fu, Yanpeng Cao
Eng. Appl. Artif. Intell.4
2023 Infrared and visible image fusion based on a two-stage class conditioned auto-encoder network
Yanpeng Cao, Xi Tong, Jiangxin Yang, Yanlong Cao
Neurocomputing5
2023 A deep thermal-guided approach for effective low-light visible image enhancement
Yanpeng Cao, Xi Tong, Fan Wang 0022, Jiangxin Yang, Yanlong Cao, Sabin Tiberius Strat, Christel-Loïc Tisse
Neurocomputing5
2023 View position prior-supervised light field angular super-resolution network with asymmetric feature extraction and spatial-angular interaction
Yanlong Cao, Lingyu Wang 0005, Lifei Ren, Jiangxin Yang, Yanpeng Cao
Neurocomputing1
2023 Light field angular super-resolution based on intrinsic and geometric information
Lingyu Wang 0005, Lifei Ren, Xiaoyao Wei, Jiangxin Yang, Yanlong Cao, Yanpeng Cao
Knowl. Based Syst.5
2023 Single image super-resolution based on progressive fusion of orientation-aware features
Zewei He, Yanpeng Cao, Jiangxin Yang, Yanlong Cao, Xin Li 0003, Siliang Tang, Yueting Zhuang, Zheming Lu 0001
Pattern Recognit.5
2023 Multi-Modal Image Fusion via Deep Laplacian Pyramid Hybrid Network
abstract
Fusion of images acquired using different sensors generates a single output with enhanced information for high-level visual perception applications. The transformer architecture has demonstrated its powerful ability to obtain important global contextual dependencies for multi-modal image fusion tasks. However, transformer-based image fusion methods face many critical issues, such as incurring huge computational burdens, limited ability to learn local features, and the difficulty of handling images of arbitrary sizes. To address the above limits, we proposed a novel Laplacian Pyramid Hybrid (LapH) network to combine the advantages of CNN and transformer architectures for multi-modal image fusion tasks. With the divide-and-conquer philosophy, we first build a light-weight CNN-based branch, performing effective extraction and fusion of texture/edge features via central difference convolutions, to process the high-resolution components with abundant details encoded in the lower pyramid levels of the Laplacian pyramid. Then, we design a transformer-based branch to process the low-resolution base components, learning long-range dependencies of global-contextual features without incurring extensive computational loads. Here, we design a multi-scale recurrent modulation mechanism to integrate the edge/texture features from the CNN branch as guidance to progressively refine the feature extraction and fusion on low-frequency components. Finally, we propose a new multi-scale spatial consistency loss term based on the neighbor contrast in source images, generating fused images with more natural and realistic appearances. Extensive experiments on two different multi-modal image fusion tasks verify the superiority of our method. The source codes are made publicly available athttps://github.com/rgttadv/LapH.
Guizhong Fu, Jiangxin Yang, Yanlong Cao, Yanpeng Cao
IEEE Trans. Circuits Syst. Video Technol.4
2023 PPI Edge Infused Spatial-Spectral Adaptive Residual Network for Multispectral Filter Array Image Demosaicing
abstract
Multispectral filter array (MSFA) sensors provide a cost-effective and one-shot acquisition solution to obtain well-aligned multi-band images, which are helpful for various optical and remote sensing applications. However, the sparse spatial sampling rate and strong spectral cross-correlation make MSFA image demosaicing a challenging problem. Therefore, it is essential to develop effective MSFA demosaicing solutions to reconstruct full-resolution and high-fidelity multispectral images from the raw mosaic image. In this paper, we present a Pseudo-panchromatic Image (PPI) Edge infused Spatial-Spectral Adaptive Residual Network (PPIE-SSARN) for multispectral filter array image demosaicing. The proposed two-branch model deploys a residual sub-branch to adaptively compensate for the spatial and spectral differences of reconstructed multispectral images and a PPI edge infusion sub-branch to enrich the edge-related information. Moreover, we design an effective mosaic initial feature extraction module with a spatial- and spectral-adaptive weight-sharing strategy whose kernel weights can change adaptively with spatial locations and spectral bands to avoid artifacts and aliasing problems. Experimental results demonstrate the superiority of our proposed method, outperforming the state-of-the-art MSFA demosaicing approaches and achieving satisfying demosaicing results in terms of spatial accuracy and spectral fidelity. Our models and code will be publicly available.
Jiesi Zheng, Yafei Dong, Jiangxin Yang, Yanlong Cao, Yanpeng Cao
IEEE Trans. Geosci. Remote. Sens.6
2022 Spatio-Temporal 3-D Residual Networks for Simultaneous Detection and Depth Estimation of CFRP Subsurface Defects in Lock-In Thermography
abstract
Nondestructive thermography is a high-speed, low-cost, and safe solution for subsurface defects detection of carbon fiber reinforced polymer (CFRP) materials, providing essential quality control in aerospace, automobile, and sports industries. In this article, we build a reflective lock-in thermography system and construct a dataset that contains real-captured thermal image sequences of CFRP samples with various simulated internal defects under different excitation frequencies. Then, we present a novel 3-D convolutional neural network (CNN) model incorporating a combination of spatial and temporal convolutional filters and batch-size independent group normalization (GN) as a unified framework to process thermal image sequences captured by lock-in thermography for simultaneous subsurface defect detection and depth estimation. Finally, we define a multitask loss function to perform end-to-end training of both defect detection and depth estimation tasks based on the real-captured infrared sequences. Comparative experiments are carried out on CFRP specimens with artificial defects of various sizes/shapes and at different depths. Qualitative and quantitative results illustrate that our 3-D CNN model is capable of predicting accurate locations and depths of subsurface defects and performs favorably against the hand-crafted and CNN-based methods in lock-in thermography for individual defect detection and depth estimation tasks. The captured dataset and the source codes will be made publicly available.
Yafei Dong, Chenjie Xia, Jiangxin Yang, Yanlong Cao, Yanpeng Cao, Xin Li 0003
IEEE Trans. Ind. Informatics4
2021 ESKN: Enhanced selective kernel network for single image super-resolution
Zewei He, Guizhong Fu, Yanpeng Cao, Yanlong Cao, Jiangxin Yang, Xin Li 0003
Signal Process.4
2020 MRFN: Multi-Receptive-Field Network for Fast and Accurate Single Image Super-Resolution
abstract
Recently, convolutional neural network (CNN) based models have shown great potential in the task of single image superresolution (SISR). However, many state-of-the-art SISR solutions are reproducing some tricks proven effective in other vision tasks, such as pursuing a deeper model. In this paper, we propose a new solution (named as Multi-Receptive-Field Network - MRFN), which outperforms existing SISR solutions in three different aspects. First, from receptive field: a novel multi-receptive-field (MRF) module is proposed to extract and fuse features in different receptive fields from local to global. Integrating these hierarchical features can generate better mappings on recovering high-fidelity details at different scales. Second, from network architectures: both dense skip connections and deep supervision are utilized to combine features from the current MRF module and preceding ones for training more representative features. Moreover, a deconvolution layer is embedded at the end of the network to avoid artificial priors induced by numerical data pre-processing (e.g., bicubic stretching), and speed up the restoration process. Finally, from error modeling: different from L1 and L2 loss functions, we proposed a novel two-parameter training loss called Weighted Huber loss function which can adaptively adjust the value of back-propagated derivative according to the residual value, thus fit the reconstruction error more effectively. Extensive qualitative and quantitative evaluation results on benchmark datasets demonstrate that our proposed MRFN can achieve more accurate recovering results than most state-of-the-art methods with significantly less complexity.
Zewei He, Yanpeng Cao, Baobei Xu, Jiangxin Yang, Yanlong Cao, Siliang Tang, Yueting Zhuang
IEEE Trans. Multim.6
2019 Fast and accurate single image super-resolution via an energy-aware improved deep residual network
Yanpeng Cao, Zewei He, Zhangyu Ye, Xin Li 0003, Yanlong Cao, Jiangxin Yang
Signal Process.5
2019 Accurate salient object detection via dense recurrent connections and residual-based hierarchical feature integration
Yanpeng Cao, Guizhong Fu, Jiangxin Yang, Yanlong Cao, Michael Ying Yang
Signal Process. Image Commun.4
2019 Cascaded Deep Networks With Multiple Receptive Fields for Infrared Image Super-Resolution
abstract
Infrared images have a wide range of military and civilian applications, including night vision, surveillance, and robotics. However, high-resolution infrared detectors are difficult to fabricate and their manufacturing cost is expensive. In this paper, we present a cascaded architecture of deep neural networks with multiple receptive fields to increase the spatial resolution of infrared images by a large scale factor (x8). Instead of reconstructing a high-resolution image from its low-resolution version using a single complex deep network, the key idea of our approach is to set up a mid-point (scale x2) between scale x1 and x8 such that lost information can be divided into two components. Lost information within each component contains similar patterns thus can be more accurately recovered even using a simpler deep network. In our proposed cascaded architecture, two consecutive deep networks with different receptive fields are jointly trained through a multi-scale loss function. The first network with a large receptive field is applied to recover largescale structure information, while the second one uses a relatively smaller receptive field to reconstruct small-scale image details. Our proposed method is systematically evaluated using realistic infrared images. Compared with state-of-the-art super-resolution methods, our proposed cascaded approach achieves improved reconstruction accuracy using significantly fewer parameters.
Zewei He, Siliang Tang, Jiangxin Yang, Yanlong Cao, Michael Ying Yang, Yanpeng Cao
IEEE Trans. Circuits Syst. Video Technol.4
2018 A multi-scale non-uniformity correction method based on wavelet decomposition and guided filtering for uncooled long wave infrared camera
Yanlong Cao, Zewei He, Jiangxin Yang, Xiaoping Ye, Yanpeng Cao
Signal Process. Image Commun.1
2017 A Geometric Structure-Based Particle Swarm Optimization Algorithm for Multiobjective Problems
abstract
This paper presents a novel evolutionary strategy for multiobjective optimization in which a population's evolution is guided by exploiting the geometric structure of its Pareto front. Specifically, the Pareto front of a particle population is regarded as a set of scattered points on which interpolation is performed using a geometric curve/surface model to construct a geometric parameter space. On this basis, the normal direction of this space can be obtained and the solutions located exactly in this direction are chosen as the guiding points. Then, the dominated solutions are processed by using a local optimization technique with the help of these guiding points. Particle populations can thus evolve toward optimal solutions with the guidance of such a geometric structure. The strategy is employed to develop a fast and robust algorithm based on correlation analysis for solving the optimization problems with more than three objectives. A number of computational experiments have been conducted to compare the algorithm to another three popular multiobjective algorithms. As demonstrated in the experiments, the proposed algorithm achieves remarkable performance in terms of the solutions obtained, robustness, and speed of convergence.
Wenqiang Yuan, Yusheng Liu 0006, Hongwei Wang 0001, Yanlong Cao
IEEE Trans. Syst. Man Cybern. Syst.4