VLDB 2026 Research / reviewers in the wild / expert
Si-Yuan Cao
dblp:264/6889
· DBLP profile ↗
29ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0001-5143-4501ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT FrameworkabstractWe present MoP-UAV, a new benchmark for UAV-based cross-view object geo-localization guided by multi-modal prompts. MoP-UAV supports fine-grained object-level cross-view localization under diverse prompt modalities, including natural language, bounding boxes, and click points. It offers potential for incorporating large foundation models like large language models (LLMs) and promotes the building of more flexible and intelligent UAV agents. Based on the benchmark, we propose MoPT, a multi-modal-prompt-guided tansformer that embeds prompts as token sequences and extract object location from UAV and satellite features via cross-attention. To enhance semantic consistency and performance, we further adopt a cross-view contrastive loss and propose a RefCOCOg-based pre-training strategy. Extensive experiments show that MoPT achieves robust localization under arbitrary prompt combinations. Notably, multi-modal-prompt training significantly boosts unimodal-prompt inference performance, highlighting the generalization benefits of multi-modal learning. MoPT trained with multi-modal prompts outperforms prior unimodal prompt works under the same setting. Zhangkai Shen, Si-Yuan Cao, Xiaokai Bai, Zheheng Han, Qi Ming |
AAAI | 3 |
| 2026 | Learning to correct unevenly exposed face images using RGB-NIR pairs
Jiacheng Ying, Runmin Zhang, Zhu Yu 0001, Beinan Yu, Si-Yuan Cao, Heng Yu 0001, Bailin Yang |
Pattern Recognit. | 8 |
| 2026 | SADet: A semantic-aware tiny object detection network against missed detection
Dongqi Yuan, Yihan Hu 0006, Beinan Yu, Xiaokai Bai, Si-Yuan Cao, Yuncheng Jin, Bailin Yang |
Pattern Recognit. | 8 |
| 2025 | SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split OptimizationabstractWe propose a novel unsupervised cross-modal homography estimation learning framework, named Split Supervised Homography estimation Network (SSHNet). SSHNet reformulates the unsupervised cross-modal homography estimation into two supervised sub-problems, each addressed by its specialized network: a homography estimation network and a modality transfer network. To realize stable training, we introduce an effective split optimization strategy to train each network separately within its respective sub-problem. We also formulate an extra homography feature space supervision to enhance feature consistency, further boosting the estimation accuracy. Moreover, we employ a simple yet effective distillation training technique to reduce model parameters and improve cross-domain generalization ability while maintaining comparable performance. The training stability of SSHNet enables its cooperation with various homography estimation architectures. Experiments reveal that the SSHNet using IHN as homography estimation network, namely SSHNet-IHN, outperforms previous unsupervised approaches by a significant margin. Even compared to supervised approaches MHN and LocalTrans, SSHNet-IHN achieves 47.4% and 85.8% mean average corner errors (MACEs) reduction on the challenging OPT-SAR dataset. Source code is available at https://github.com/Junchen-Yu/SSHNet. Junchen Yu, Si-Yuan Cao, Runmin Zhang, Chenghao Zhang 0002, Zhu Yu 0001, Shujie Chen 0001, Bailin Yang |
CVPR | 2 |
| 2025 | Language Driven Occupancy PredictionabstractWe introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model-view projections. To alleviate the inaccurate supervision, we propose a semantic transitive labeling pipeline to generate dense and fine-grained 3D language occupancy ground truth. Our pipeline presents a feasible way to dig into the valuable semantic information of images, transferring text labels from images to LiDAR point clouds and ultimately to voxels, to establish precise voxel-to-text correspondences. By replacing the original prediction head of supervised occupancy models with a geometry head for binary occupancy states and a language head for language features, LOcc effectively uses the generated language ground truth to guide the learning of 3D language volume. Through extensive experiments, we demonstrate that our transitive semantic labeling pipeline can produce more accurate pseudo-labeled ground truth, diminishing labor-intensive human annotations. Additionally, we validate LOcc across various architectures, where all models consistently outperform state-of-the-art zero-shot occupancy prediction approaches on the Occ3D-nuScenes dataset. Zhu Yu 0001, Lizhe Liu, Runmin Zhang, Si-Yuan Cao, Maochun Luo, Mingxia Chen |
ICCV | 6 |
| 2025 | Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionabstractThis work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate geometric and contextual information within adaptive regions in each image and dynamically adjust the contributions from different views, enhancing the representation capability of voxel features. Furthermore, we propose a sparse volume construction strategy that adaptively identifies and selects voxels with high occupancy probabilities for feature refinement, minimizing redundant computation in free space. Benefiting from the above designs, our framework achieves effective and efficient volume construction in an adaptive way. Better still, our network can be supervised using only 3D bounding boxes, eliminating the dependence on ground-truth scene geometry. Experimental results demonstrate that SGCDet achieves state-of-the-art performance on the ScanNet, ScanNet200 and ARKitScenes datasets. The source code is available at https://github.com/RM-Zhang/SGCDet. Runmin Zhang, Zhu Yu 0001, Si-Yuan Cao, Lingyu Zhu 0006, Guangyi Zhang 0005, Xiaokai Bai |
ICCV | 3 |
| 2025 | EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image RegistrationabstractPrevious deep image registration methods that employ single homography, multi-grid homography, or thin-plate spline often struggle with real scenes containing depth disparities due to their inherent limitations. To address this, we propose an Exponential-Decay Free-Form Deformation Network (EDFFDNet), which employs free-form deformation with an exponential-decay basis function. This design achieves higher efficiency and performs well in scenes with depth disparities, benefiting from its inherent locality. We also introduce an Adaptive Sparse Motion Aggregator (ASMA), which replaces the MLP motion aggregator used in previous methods. By transforming dense interactions into sparse ones, ASMA reduces parameters and improves accuracy. Additionally, we propose a progressive correlation refinement strategy that leverages global-local correlation patterns for coarse-to-fine motion estimation, further enhancing efficiency and accuracy. Experiments demonstrate that EDFFDNet reduces parameters, memory, and total runtime by 70.5%, 32.6%, and 33.7%, respectively, while achieving a 0.5 dB PSNR gain over the state-of-the-art method. With an additional local refinement stage,EDFFDNet-2 further improves PSNR by 1.06 dB while maintaining lower computational costs. Our method also demonstrates strong generalization ability across datasets, outperforming previous deep learning methods. Haokai Zhu, Bo Qu, Si-Yuan Cao, Runmin Zhang, Shujie Chen 0001, Bailin Yang |
ICCV | 3 |
| 2025 | Structure-Aware Radar-Camera Depth EstimationabstractRadar has gained much attention in autonomous driving due to its accessibility and robustness. However, its standalone application for depth perception is constrained by issues of sparsity and noise. Radar-camera depth estimation offers a more promising complementary solution. Despite significant progress, current approaches fail to produce satisfactory dense depth maps, due to the unsatisfactory processing of the sparse and noisy radar data. They constrain the regions of interest for radar points in rigid rectangular regions, which may introduce unexpected errors and confusions. To address these issues, we develop a structure-aware strategy for radar depth enhancement, which provides more targeted regions of interest by leveraging the structural priors of RGB images. Furthermore, we design a Multi-Scale Structure Guided Network to enhance radar features and preserve detailed structures, achieving accurate and structure-detailed dense metric depth estimation. Building on these, we propose a structure-aware radar-camera depth estimation framework, named SA-RCD. Extensive experiments demonstrate that our SA-RCD achieves state-of-the-art performance on the nuScenes dataset. Our code will be available at https://github.com/FreyZhangYeh/SA-RCD. Fuyi Zhang, Zhu Yu 0001, Chunhao Li, Runmin Zhang, Xiaokai Bai, Si-Yuan Cao |
ICRA | 7 |
| 2025 | LGDD: Local-Global Synergistic Dual-Branch 3D Object Detection Using 4D Radarabstract4D millimeter-wave radar plays a pivotal role in autonomous driving due to its cost-effectiveness and robustness in adverse weather. However, the application of 4D radar point cloud in 3D perception tasks is hindered by its inherent sparsity and noise. To address these challenges, we propose LGDD, a novel local-global synergistic dual-branch 3D object detection framework using 4D radar. Specifically, we first introduce a point-based branch, which utilize a voxel-attended point feature extractor (VPE) to integrate semantic segmentation with cluster voting, thereby mitigating radar noise and extracting local-clustered instances features. Then, for the conventional pillar-based branch, we design a query-based feature pre-fusion (QFP) to address the sparsity and enhance global context representation. Additionally, we devise a proposal mask to filter out noisy points, enabling more focused clustering on regions of interest. Finally, we align the local instances with global context through semantics-geometry aware fusion (SGF) module to achieve comprehensive scene understanding. Extensive experiments demonstrate that LGDD achieves state-of-the-art performance on the public View-of-Delft and TJ4DRadSet datasets. Source code is available at https://github.com/shawnnnkb/LGDD. Xiaokai Bai, Fuyi Zhang, Si-Yuan Cao, Lianqing Zheng, Beinan Yu |
IROS | 6 |
| 2025 | Beyond Registration: Self-supervised Unknown Border Completion
Xiaokai Bai, Leyuan Yu, Runmin Zhang, Beinan Yu, Si-Yuan Cao |
PRCV (8) | 7 |
| 2025 | Boosting Conditional Diffusion Models Using Intermediate Segmentation Map for Infrared Small Target Detection
Wenkai Zhao, Xiaokai Bai, Yicheng Tong, Si-Yuan Cao |
PRCV (4) | 5 |
| 2025 | STARNet: Low-light video enhancement using spatio-temporal consistency aggregation
Zehua Sheng, Si-Yuan Cao, Runmin Zhang, Beinan Yu, Chenghao Zhang 0002, Bailin Yang |
Pattern Recognit. | 4 |
| 2025 | TEFormer: Thermal Infrared Image Enhancement by Preserving Spatial Consistency and DetailsabstractThermal infrared (TIR) images suffer from low contrast due to the atmospheric thermal radiation effect, especially under extreme conditions like low temperature. TIR image enhancement aims to improve image contrast, but previous enhancement approaches usually produce enhanced results with two limitations: spatial inconsistency and detail blurring. To deal with the limitations, we propose a novel TIR image enhancement method, named TEFormer, to preserve spatial consistency and restore fine-grained details during image enhancement. To preserve spatial consistency, we devise the global enhancement module (GEM) to enhance the low-resolution representation. The GEM performs long-range interactions across spatial dimensions and channel dimensions to condition the enhancement curve fitting. To keep details clear, we design the local enhancement module (LEM) as the decoding unit. The LEM injects additional detail structures into the enhanced low-resolution representation for high-resolution reconstruction. Besides, we further apply histogram-based supervision to facilitate learning in intensity distribution of clear images. Extensive experimental results on three challenging benchmarks demonstrate that the proposed method outperforms other state-of-the-art approaches. Yunxin Li, Runmin Zhang, Si-Yuan Cao, Jiacheng Ying, Xiaokai Bai, Shujie Chen 0001, Bailin Yang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | GeoFormer: Boosting Object Distinguishing and Prompt Understanding for Cross-View Object Geo-LocalizationabstractCross-view object geo-localization (CVOGL) determines the geographic location of an object on the satellite view reference image. The object is indicated by a point prompt in a ground- or drone-view query image. Despite wide applications, the current CVOGL method still suffers from limited performance, and the reason for this limitation remains unknown. In this work, we analyze CVOGL and find two primary challenges that hinder its performance,i.e., 1) point prompt understanding and 2) similar-appearance object distinguishing. Therefore, we propose a novel end-to-end framework called object geo-localization transformer (GeoFormer). Specifically, we leverage the knowledge of the segment anything model (SAM) through two settings for accurate point prompt understanding,i.e., using SAM during training and inference (GeoFormer) or using SAM only in training via knowledge distillation (GeoFormer-KD). Additionally, we devise an information aggregation module (IAM) to leverage local and global perception for similar-appearance object distinguishing. Except for the public dataset, we manually annotated a new dataset that contains 1642 image pairs for further comparison. Experiments show that our method significantly outperforms the previous work. Notably, the tiny versions of our method (GeoFormer-t and GeoFormer-t-KD) maintain state-of-the-art performance while substantially reducing parameter costs. Our code and the new dataset will be made available at https://github.com/Temperature-ai/GeoFormer. Si-Yuan Cao, Zhu Yu 0001, Xiaokai Bai, Beinan Yu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | BEVPlace++: Fast, Robust, and Lightweight LiDAR Global Localization for Autonomous Ground VehiclesabstractThis article introduces BEVPlace++, a novel, fast, and robust LiDAR global localization method for unmanned ground vehicles. It uses lightweight convolutional neural networks (CNNs) on Bird's Eye View (BEV) image-like representations of LiDAR data to achieve accurate global localization through place recognition, followed by 3-DoF pose estimation. Our detailed analyses reveal an interesting fact that CNNs are inherently effective at extracting distinctive features from LiDAR BEV images. Remarkably, keypoints of two BEV images with large translations can be effectively matched using CNN-extracted features. Building on this insight, we design a Rotation Equivariant Module (REM) to obtain distinctive features while enhancing robustness to rotational changes. A Rotation Equivariant and Invariant Network (REIN) is then developed by cascading REM and a descriptor generator, NetVLAD, to sequentially generate rotation equivariant local features and rotation invariant global descriptors. The global descriptors are used first to achieve robust place recognition, and then local features are used for accurate pose estimation. Experimental results on seven public datasets and our UGV platform demonstrate that BEVPlace++, even when trained on a small dataset (3000 frames of KITTI) only with place labels, generalizes well to unseen environments, performs consistently across different days and years, and adapts to various types of LiDAR scanners. BEVPlace++ achieves state-of-theart performance in multiple tasks, including place recognition, loop closure detection, and global localization. Additionally, BEVPlace++ is lightweight, runs in real-time, and does not require accurate pose supervision, making it highly convenient for deployment. The source codes are publicly available at https://github.com/zjuluolun/BEVPlace2. Lun Luo, Si-Yuan Cao, Jintao Xu 0001, Rui Ai 0001, Zhu Yu 0001, Xieyuanli Chen |
IEEE Trans. Robotics | 2 |
| 2024 | MCNet: Rethinking the Core Ingredients for Accurate and Efficient Homography EstimationabstractWe propose Multiscale Correlation searching homogra-phy estimation Network, namely MCNet, an iterative deep homography estimation architecture. Different from previous approaches that achieve iterative refinement by correlation searching within a single scale, MCNet combines the multiscale strategy with correlation searching incur-ring nearly ignored computational overhead. Moreover, MCNet adopts a Fine-Grained Optimization loss function, named FGO loss, to further boost the network training at the convergent stage, which can improve the estimation accuracy without additional computational overhead. Ac-cording to our experiments, using the above two simple strategies can produce significant homography estimation accuracy with considerable efficiency. We show that MC-Net achieves state-of-the-art performance on a variety of datasets, including common scene MSCOCO, cross-modal scene GoogleEarth and GoogleMap, and dynamic scene SPID. Compared to the previous SOTA method, 2-scale RHWF, our MCNet reduces inference time, FLOPs, parameter cost, and memory cost by 78.9%, 73.5%, 34.1%, and 33.2% respectively, while achieving 20.5% (MSCOCO), 43.4% (GoogleEarth), and 41.1% (GoogleMap) mean average corner error (MACE) reduction. Source code is available at https://github.com/zjuzhk/MCNet. Haokai Zhu, Si-Yuan Cao, Jianxin Hu, Sitong Zuo, Beinan Yu, Jiacheng Ying, Junwei Li 0009 |
CVPR | 2 |
| 2024 | SCPNet: Unsupervised Cross-Modal Homography Estimation via Intra-modal Self-supervised Learning
Runmin Zhang, Si-Yuan Cao, Lun Luo, Beinan Yu, Shujie Chen 0001, Junwei Li 0009 |
ECCV (23) | 3 |
| 2024 | Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionabstractVision-based Semantic Scene Completion (SSC) has gained much attention due to its widespread applications in various 3D perception tasks. Existing sparse-to-dense approaches typically employ shared context-independent queries across various input images, which fails to capture distinctions among them as the focal regions of different inputs vary and may result in undirected feature aggregation of cross-attention. Additionally, the absence of depth information may lead to points projected onto the image plane sharing the same 2D position or similar sampling points in the feature map, resulting in depth ambiguity. In this paper, we present a novel context and geometry aware voxel transformer. It utilizes a context aware query generator to initialize context-dependent queries tailored to individual input images, effectively capturing their unique characteristics and aggregating information within the region of interest. Furthermore, it extend deformable cross-attention from 2D to 3D pixel space, enabling the differentiation of points with similar image coordinates based on their depth coordinates. Building upon this module, we introduce a neural network named CGFormer to achieve semantic scene completion. Simultaneously, CGFormer leverages multiple 3D representations (i.e., voxel and TPV) to boost the semantic and geometric representation abilities of the transformed 3D volume from both local and global perspectives. Experimental results demonstrate that CGFormer achieves state-of-the-art performance on the SemanticKITTI and SSCBench-KITTI-360 benchmarks, attaining a mIoU of 16.87 and 20.05, as well as an IoU of 45.99 and 48.07, respectively. Remarkably, CGFormer even outperforms approaches employing temporal images as inputs or much larger image backbone networks. Zhu Yu 0001, Runmin Zhang, Jiacheng Ying, Junchen Yu, Xiaohai Hu, Lun Luo, Si-Yuan Cao |
NeurIPS | 7 |
| 2024 | MRF3Net: An Infrared Small Target Detection Network Using Multireceptive Field Perception and Effective Feature FusionabstractInfrared small target detection (IRSTD) has made remarkable achievements in recent years. However, the core focus of current works lies on the philosophy of “increasing network complexity,” which leaves the crucial strategies behind performance improvement unclear. To handle this, we highlight two strategies of IRSTD: 1) multireceptive field perception and 2) effective feature fusion. Focusing on these strategies, we propose a multireceptive field perception and effective feature fusion network, named MRF3Net. Specifically, for multireceptive field perception, we devise a multiple perception encoder (MPE). For effective feature fusion, we devise a feature fusion encoder (FFE) and a feature fusion decoder (FFD). The former improves the encoding efficiency by reducing the interference information and preserving target details, and the latter fuses the target information of low-level and high-level features while eliminating noise. Experiments demonstrate that MRF3Net achieves state-of-the-art performance on popular public datasets while maintaining a fast inference speed of approximately 0.011 s/frame on a single NVIDIA GeForce 3070Ti GPU and 0.049 s/frame on the NVIDIA Jetson Orin. Notably, it significantly reduces parameter costs compared with the previous state-of-the-art approaches, validating its efficiency. Moreover, our MPE, FFE, and FFD have proved to be effective in enhancing other IRSTD approaches. Our code will be made available at:https://github.com/Temperature-ai/MRF3Net. Si-Yuan Cao, Beinan Yu, Chenghao Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerabstractWe propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the attention-focusing mechanism in FocusFormer aggregates the intra-inter correspondence in a global→nonlocal→local manner. Thanks to the above strategies, RHWF ranks top in accuracy on a variety of datasets, including the challenging cross-resolution and cross-modal ones. Meanwhile, benefiting from the recurrent framework, RHWF achieves parameter efficiency despite the transformer architecture. Compared to previous state-of-the-art approaches LocalTrans and IHN, RHWF reduces the mean average corner error (MACE) by about 70% and 38.1% on the MSCOCO dataset, while saving the parameter costs by 86.5% and 24.6%. Similar to the previous works, RHWF can also be arranged in 1-scale for efficiency and 2-scale for accuracy, with the 1-scale RHWF already outperforming most of the previous methods. Source code is available at https://github.com/imdump178/RHWF. Si-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu, Zehua Sheng, Junwei Li 0009 |
CVPR | 1 |
| 2023 | Structure Aggregation for Cross-Spectral Stereo Image Guided DenoisingabstractTo obtain clean images with salient structures from noisy observations, a growing trend in current denoising studies is to seek the help of additional guidance images with high signal-to-noise ratios, which are often acquired in different spectral bands such as near infrared. Although previous guided denoising methods basically require the input images to be well-aligned, a more common way to capture the paired noisy target and guidance images is to exploit a stereo camera system. However, current studies on cross-spectral stereo matching cannot fully guarantee the pixel-level registration accuracy, and rarely consider the case of noise contamination. In this work, for the first time, we propose a guided denoising framework for cross-spectral stereo images. Instead of aligning the input images via conventional stereo matching, we aggregate structures from the guidance image to estimate a clean structure map for the noisy target image, which is then used to regress the final denoising result with a spatially variant linear representation model. Based on this, we design a neural network, called as SANet, to complete the entire guided denoising process. Experimental results show that, our SANet can effectively transfer structures from an unaligned guidance image to the restoration result, and outperforms state-of-the-art denoisers on various stereo image datasets. Besides, our structure aggregation strategy also shows its potential to handle other unaligned guided restoration tasks such as super-resolution and deblurring. The source code is available at https://github.com/lustrouselixir/SANet. Zehua Sheng, Zhu Yu 0001, Xiongwei Liu, Si-Yuan Cao, Yuqi Liu 0005, Huaqi Zhang |
CVPR | 4 |
| 2023 | BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesabstractPlace recognition is a key module for long-term SLAM systems. Current LiDAR-based place recognition methods usually use representations of point clouds such as unordered points or range images. These methods achieve high recall rates of retrieval, but their performance may degrade in the case of view variation or scene changes. In this work, we explore the potential of a different representation in place recognition, i.e. bird’s eye view (BEV) images. We validate that, in scenes of slight viewpoint changes, a simple NetVLAD network trained on BEV images achieves comparable performance to the state-of-the-art place recognition methods. For robustness to view variations, we propose a rotation-invariant network called BEVPlace. We use group convolution to extract rotation-equivariant local features from the images and NetVLAD for global feature aggregation. In addition, we observe that the distance between BEV features is correlated with the geometry distance of point clouds. Based on the observation, we develop a method to estimate the position of the query cloud, extending the usage of place recognition. The experiments conducted on large-scale public datasets show that our method 1) achieves state-of-the-art performance in terms of recall rates, 2) is robust to view changes, 3) shows strong generalization ability, and 4) can estimate the positions of query point clouds. Source codes are publicly available at https://github.com/zjuluolun/BEVPlace. Lun Luo, Shuhang Zheng, Yongzhi Fan, Beinan Yu, Si-Yuan Cao, Junwei Li 0009 |
ICCV | 6 |
| 2023 | Aggregating Feature Point Cloud for Depth CompletionabstractGuided depth completion aims to recover dense depth maps by propagating depth information from the given pixels to the remaining ones under the guidance of RGB images. However, most of the existing methods achieve this using a large number of iterative refinements or stacking repetitive blocks. Due to the limited receptive field of conventional convolution, the generalizability with respect to different sparsity levels of input depth maps is impeded. To tackle these problems, we propose a feature point cloud aggregation framework to directly propagate 3D depth information between the given points and the missing ones. We extract 2D feature map from images and transform the sparse depth map to point cloud to extract sparse 3D features. By regarding the extracted features as two sets of feature point clouds, the depth information for a target location can be reconstructed by aggregating adjacent sparse 3D features from the known points using cross attention. Based on this, we design a neural network, called as PointDC, to complete the entire depth information reconstruction process. Experimental results show that, our PointDC achieves superior or competitive results on the KITTI benchmark and NYUv2 dataset. In addition, the proposed PointDC demonstrates its higher generalizability to different sparsity levels of the input depth maps and cross-dataset evaluation. Zhu Yu 0001, Zehua Sheng, Lun Luo, Si-Yuan Cao, Huaqi Zhang |
ICCV | 5 |
| 2023 | I2P-Rec: Recognizing Images on Large-Scale Point Cloud Maps Through Bird's Eye View ProjectionsabstractPlace recognition is an important technique for autonomous cars to achieve full autonomy since it can provide an initial guess to online localization algorithms. Although current methods based on images or point clouds have achieved satisfactory performance, localizing the images on a large-scale point cloud map remains a fairly unexplored problem. This cross-modal matching task is challenging due to the difficulty in extracting consistent descriptors from images and point clouds. In this paper, we propose the I2P-Rec method to solve the problem by transforming the cross-modal data into the same modality. Specifically, we leverage on the recent success of depth estimation networks to recover point clouds from images. We then project the point clouds into Bird's Eye View (BEV) images. Using the BEV image as an intermediate representation, we extract global features with a Convolutional Neural Network followed by a NetVLAD layer to perform matching. The experimental results evaluated on the KITTI dataset show that, with only a small set of training data, I2P-Rec achieves recall rates at Top-l % over 80% and 90%, when localizing monocular and stereo images on point cloud maps, respectively. We further evaluate I2P-Rec on a 1 km trajectory dataset collected by an autonomous logistics car and show that I2P- Rec can generalize well to previously unseen environments. Shuhang Zheng, Zhu Yu 0001, Beinan Yu, Si-Yuan Cao, Minhang Wang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Lun Luo |
IROS | 5 |
| 2023 | Region-aware RGB and near-infrared image fusion
Jiacheng Ying, Can Tong, Zehua Sheng, Bo-Wen Yao, Si-Yuan Cao, Heng Yu 0001 |
Pattern Recognit. | 5 |
| 2023 | Frequency-Domain Deep Guided Image DenoisingabstractDespite the tremendous advances in denoising techniques, it's still challenging to restore a clean image with salient structures based on one noisy observation, especially at high noise levels. In this work, we propose a frequency-domain guided denoising algorithm to conduct denoising with the help of a well-aligned guidance image. Thanks to their structural correlations, the frequency characteristics of the guidance image can indicate whether the frequency coefficients of the noisy target image are contributed by noise or textures. Therefore, the explicit frequency decomposition enables our denoising model to avoid over-smoothing detailed contents. However, as two input images are usually captured in different fields, their structures are not always consistent. Therefore, we model guided denoising with an optimization problem which considers both the representation model of the guidance image and the fidelity to the noisy target. Further, we design a convolutional neural network, called as FGDNet, to explore the optimal solution. Due to the visual masking phenomenon, human eyes are sensitive to noise in the flat areas, but may not perceive noise around edges or textures. Therefore, we expect to remove as much noise as possible to guarantee the spatial smoothness of flat contents, while also preserving high-frequency structures. Through frequency decomposition, our model can process the low-frequency and high-frequency contents separately. We also adopt a frequency-relevant loss function to train the network. Experimental results show that, compared with state-of-the-art guided and non-guided denoisers, our FGDNet achieves higher denoising accuracy and better visual quality in both flat and texture-rich regions. Zehua Sheng, Xiongwei Liu, Si-Yuan Cao, Huaqi Zhang |
IEEE Trans. Multim. | 3 |
| 2022 | Iterative Deep Homography EstimationabstractWe propose Iterative Homography Network, namely IHN, a new deep homography estimation architecture. Different from previous works that achieve iterative refinement by network cascading or untrainable IC-LK iterator; the iterator of IHN has tied weights and is completely trainable. IHN achieves state-of-the-art accuracy on several datasets including challenging scenes. We propose 2 versions of IHN: (1) IHN for static scenes, (2) IHN-mov for dynamic scenes with moving objects. Both versions can be arranged in 1-scale for efficiency or 2-scale for accuracy. We show that the basic 1-scale IHN already outperforms most of the existing methods. On a variety of datasets, the 2-scale IHN outperforms all competitors by a large gap. We introduce IHN-mov by producing an inlier mask to further improve the estimation accuracy of moving-objects scenes. We experimentally show that the iterative framework of IHN can achieve 95% error reduction while considerably saving network parameters. When processing sequential image pairs, IHN can achieve 32.7 fps, which is about 8× the speed of IC-LK iterator: Source code is available at https://github.com/imdump178/IHN. Si-Yuan Cao, Jianxin Hu, Zehua Sheng |
CVPR | 1 |
| 2022 | Unaligned Hyperspectral Image Fusion via Registration and Interpolation ModelingabstractIn satellite remote sensing, the hyperspectral sensor acquires high-spectral-resolution and low-spatial-resolution hyperspectral images (HSIs). Conversely, the multispectral sensor acquires low-spectral-resolution and high-spatial-resolution multispectral images (MSIs). Thus, HSI and MSI fusion is required to promote both spatial and spectral resolutions. Currently, most algorithms are based on the assumption that the HSI and MSI are perfectly aligned. However, this is hardly achievable in real scenarios when the two sensors acquire images from different viewpoints. In this article, we propose a fusion algorithm that consists of two stages, i.e., image registration and image fusion. For image registration, we introduce the normalized edge difference (NED) for image similarity measure considering the different resolutions of the original images. For image fusion, we incorporate the interpolation process in the spatial degradation model to compensate for the interpolation error. Experimental results show that our algorithm performs better than the state of the arts for unaligned image fusion. Jiacheng Ying, Si-Yuan Cao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Boosting Structure Consistency for Multispectral and Multimodal Image RegistrationabstractMultispectral imaging plays a vital role in the area of computer vision and computational photography. As spectral band images can be misaligned due to imaging device movement or alternation, image registration is necessary to avoid spectral information distortion. The current registration measures specialized for multispectral data are typically robust yet complex, requiring excessive computation. The common measures such as sum of squared differences (SSD) and sum of absolute differences (SAD) are computationally efficient whereas they perform poorly on multispectral data. To cope with this challenge, we propose a structure consistency boosting (SCB) transform that aims at boosting the structural similarity of multispectral images. With SCB, the common measures can be employed for multispectral image registration. The SCB transform exploits the fact that inherent edge structures maintain relative saliency locally despite the nonlinear variation between band images. A statistical prior of the natural image, which is based on the gradient-intensity correlation, is explored to build a parametric form of SCB. Experimental results validate that the SCB transform outperforms current similarity enhancement algorithms, and performs better than the state-of-the-art multispectral registration measures. Thanks to the generality of the statistical prior, the SCB transform is also applicable to various multimodal data such as flash/no-flash images and medical images. Si-Yuan Cao, Shujie Chen 0001, Chunguang Li 0001 |
IEEE Trans. Image Process. | 1 |