VLDB 2026 Research / reviewers in the wild / expert
Beinan Yu
dblp:318/8632
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-1557-7166ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to correct unevenly exposed face images using RGB-NIR pairs
Jiacheng Ying, Runmin Zhang, Zhu Yu 0001, Beinan Yu, Si-Yuan Cao, Heng Yu 0001, Bailin Yang |
Pattern Recognit. | 7 |
| 2026 | SADet: A semantic-aware tiny object detection network against missed detection
Dongqi Yuan, Yihan Hu 0006, Beinan Yu, Xiaokai Bai, Si-Yuan Cao, Yuncheng Jin, Bailin Yang |
Pattern Recognit. | 6 |
| 2025 | CDFormer: Cross-Domain Few-Shot Object Detection Transformer Against Feature ConfusionabstractCross-domain few-shot object detection (CD-FSOD) aims to detect novel objects across different domains with limited class instances. Feature confusion, including object-background confusion and object-object confusion, presents significant challenges in both cross-domain and few-shot settings. In this work, we introduce CDFormer, a cross-domain few-shot object detection transformer against feature confusion, to address these challenges. The method specifically tackles feature confusion through two key modules: object-background distinguishing (OBD) and object-object distinguishing (OOD). The OBD module leverages a learnable background token to differentiate between objects and background, while the OOD module enhances the distinction between objects of different classes. Experimental results demonstrate that CDFormer outperforms previous state-of-the-art approaches, achieving 12.9% mAP, 11.0% mAP, and 10.4% mAP improvements under the 1/5/10 shot settings, respectively, when fine-tuned. Code is available at https://longxuanx.github.io/CDFormer/. Boyuan Meng, Wenkai Zhao, Beinan Yu |
ICME | 7 |
| 2025 | LGDD: Local-Global Synergistic Dual-Branch 3D Object Detection Using 4D Radarabstract4D millimeter-wave radar plays a pivotal role in autonomous driving due to its cost-effectiveness and robustness in adverse weather. However, the application of 4D radar point cloud in 3D perception tasks is hindered by its inherent sparsity and noise. To address these challenges, we propose LGDD, a novel local-global synergistic dual-branch 3D object detection framework using 4D radar. Specifically, we first introduce a point-based branch, which utilize a voxel-attended point feature extractor (VPE) to integrate semantic segmentation with cluster voting, thereby mitigating radar noise and extracting local-clustered instances features. Then, for the conventional pillar-based branch, we design a query-based feature pre-fusion (QFP) to address the sparsity and enhance global context representation. Additionally, we devise a proposal mask to filter out noisy points, enabling more focused clustering on regions of interest. Finally, we align the local instances with global context through semantics-geometry aware fusion (SGF) module to achieve comprehensive scene understanding. Extensive experiments demonstrate that LGDD achieves state-of-the-art performance on the public View-of-Delft and TJ4DRadSet datasets. Source code is available at https://github.com/shawnnnkb/LGDD. Xiaokai Bai, Fuyi Zhang, Si-Yuan Cao, Lianqing Zheng, Beinan Yu |
IROS | 8 |
| 2025 | Beyond Registration: Self-supervised Unknown Border Completion
Xiaokai Bai, Leyuan Yu, Runmin Zhang, Beinan Yu, Si-Yuan Cao |
PRCV (8) | 6 |
| 2025 | STARNet: Low-light video enhancement using spatio-temporal consistency aggregation
Zehua Sheng, Si-Yuan Cao, Runmin Zhang, Beinan Yu, Chenghao Zhang 0002, Bailin Yang |
Pattern Recognit. | 6 |
| 2025 | GeoFormer: Boosting Object Distinguishing and Prompt Understanding for Cross-View Object Geo-LocalizationabstractCross-view object geo-localization (CVOGL) determines the geographic location of an object on the satellite view reference image. The object is indicated by a point prompt in a ground- or drone-view query image. Despite wide applications, the current CVOGL method still suffers from limited performance, and the reason for this limitation remains unknown. In this work, we analyze CVOGL and find two primary challenges that hinder its performance,i.e., 1) point prompt understanding and 2) similar-appearance object distinguishing. Therefore, we propose a novel end-to-end framework called object geo-localization transformer (GeoFormer). Specifically, we leverage the knowledge of the segment anything model (SAM) through two settings for accurate point prompt understanding,i.e., using SAM during training and inference (GeoFormer) or using SAM only in training via knowledge distillation (GeoFormer-KD). Additionally, we devise an information aggregation module (IAM) to leverage local and global perception for similar-appearance object distinguishing. Except for the public dataset, we manually annotated a new dataset that contains 1642 image pairs for further comparison. Experiments show that our method significantly outperforms the previous work. Notably, the tiny versions of our method (GeoFormer-t and GeoFormer-t-KD) maintain state-of-the-art performance while substantially reducing parameter costs. Our code and the new dataset will be made available at https://github.com/Temperature-ai/GeoFormer. Si-Yuan Cao, Zhu Yu 0001, Xiaokai Bai, Beinan Yu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | MCNet: Rethinking the Core Ingredients for Accurate and Efficient Homography EstimationabstractWe propose Multiscale Correlation searching homogra-phy estimation Network, namely MCNet, an iterative deep homography estimation architecture. Different from previous approaches that achieve iterative refinement by correlation searching within a single scale, MCNet combines the multiscale strategy with correlation searching incur-ring nearly ignored computational overhead. Moreover, MCNet adopts a Fine-Grained Optimization loss function, named FGO loss, to further boost the network training at the convergent stage, which can improve the estimation accuracy without additional computational overhead. Ac-cording to our experiments, using the above two simple strategies can produce significant homography estimation accuracy with considerable efficiency. We show that MC-Net achieves state-of-the-art performance on a variety of datasets, including common scene MSCOCO, cross-modal scene GoogleEarth and GoogleMap, and dynamic scene SPID. Compared to the previous SOTA method, 2-scale RHWF, our MCNet reduces inference time, FLOPs, parameter cost, and memory cost by 78.9%, 73.5%, 34.1%, and 33.2% respectively, while achieving 20.5% (MSCOCO), 43.4% (GoogleEarth), and 41.1% (GoogleMap) mean average corner error (MACE) reduction. Source code is available at https://github.com/zjuzhk/MCNet. Haokai Zhu, Si-Yuan Cao, Jianxin Hu, Sitong Zuo, Beinan Yu, Jiacheng Ying, Junwei Li 0009 |
CVPR | 5 |
| 2024 | SCPNet: Unsupervised Cross-Modal Homography Estimation via Intra-modal Self-supervised Learning
Runmin Zhang, Si-Yuan Cao, Lun Luo, Beinan Yu, Shujie Chen 0001, Junwei Li 0009 |
ECCV (23) | 5 |
| 2024 | MRF3Net: An Infrared Small Target Detection Network Using Multireceptive Field Perception and Effective Feature FusionabstractInfrared small target detection (IRSTD) has made remarkable achievements in recent years. However, the core focus of current works lies on the philosophy of “increasing network complexity,” which leaves the crucial strategies behind performance improvement unclear. To handle this, we highlight two strategies of IRSTD: 1) multireceptive field perception and 2) effective feature fusion. Focusing on these strategies, we propose a multireceptive field perception and effective feature fusion network, named MRF3Net. Specifically, for multireceptive field perception, we devise a multiple perception encoder (MPE). For effective feature fusion, we devise a feature fusion encoder (FFE) and a feature fusion decoder (FFD). The former improves the encoding efficiency by reducing the interference information and preserving target details, and the latter fuses the target information of low-level and high-level features while eliminating noise. Experiments demonstrate that MRF3Net achieves state-of-the-art performance on popular public datasets while maintaining a fast inference speed of approximately 0.011 s/frame on a single NVIDIA GeForce 3070Ti GPU and 0.049 s/frame on the NVIDIA Jetson Orin. Notably, it significantly reduces parameter costs compared with the previous state-of-the-art approaches, validating its efficiency. Moreover, our MPE, FFE, and FFD have proved to be effective in enhancing other IRSTD approaches. Our code will be made available at:https://github.com/Temperature-ai/MRF3Net. Si-Yuan Cao, Beinan Yu, Chenghao Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerabstractWe propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the attention-focusing mechanism in FocusFormer aggregates the intra-inter correspondence in a global→nonlocal→local manner. Thanks to the above strategies, RHWF ranks top in accuracy on a variety of datasets, including the challenging cross-resolution and cross-modal ones. Meanwhile, benefiting from the recurrent framework, RHWF achieves parameter efficiency despite the transformer architecture. Compared to previous state-of-the-art approaches LocalTrans and IHN, RHWF reduces the mean average corner error (MACE) by about 70% and 38.1% on the MSCOCO dataset, while saving the parameter costs by 86.5% and 24.6%. Similar to the previous works, RHWF can also be arranged in 1-scale for efficiency and 2-scale for accuracy, with the 1-scale RHWF already outperforming most of the previous methods. Source code is available at https://github.com/imdump178/RHWF. Si-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu, Zehua Sheng, Junwei Li 0009 |
CVPR | 4 |
| 2023 | BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesabstractPlace recognition is a key module for long-term SLAM systems. Current LiDAR-based place recognition methods usually use representations of point clouds such as unordered points or range images. These methods achieve high recall rates of retrieval, but their performance may degrade in the case of view variation or scene changes. In this work, we explore the potential of a different representation in place recognition, i.e. bird’s eye view (BEV) images. We validate that, in scenes of slight viewpoint changes, a simple NetVLAD network trained on BEV images achieves comparable performance to the state-of-the-art place recognition methods. For robustness to view variations, we propose a rotation-invariant network called BEVPlace. We use group convolution to extract rotation-equivariant local features from the images and NetVLAD for global feature aggregation. In addition, we observe that the distance between BEV features is correlated with the geometry distance of point clouds. Based on the observation, we develop a method to estimate the position of the query cloud, extending the usage of place recognition. The experiments conducted on large-scale public datasets show that our method 1) achieves state-of-the-art performance in terms of recall rates, 2) is robust to view changes, 3) shows strong generalization ability, and 4) can estimate the positions of query point clouds. Source codes are publicly available at https://github.com/zjuluolun/BEVPlace. Lun Luo, Shuhang Zheng, Yongzhi Fan, Beinan Yu, Si-Yuan Cao, Junwei Li 0009 |
ICCV | 5 |
| 2023 | I2P-Rec: Recognizing Images on Large-Scale Point Cloud Maps Through Bird's Eye View ProjectionsabstractPlace recognition is an important technique for autonomous cars to achieve full autonomy since it can provide an initial guess to online localization algorithms. Although current methods based on images or point clouds have achieved satisfactory performance, localizing the images on a large-scale point cloud map remains a fairly unexplored problem. This cross-modal matching task is challenging due to the difficulty in extracting consistent descriptors from images and point clouds. In this paper, we propose the I2P-Rec method to solve the problem by transforming the cross-modal data into the same modality. Specifically, we leverage on the recent success of depth estimation networks to recover point clouds from images. We then project the point clouds into Bird's Eye View (BEV) images. Using the BEV image as an intermediate representation, we extract global features with a Convolutional Neural Network followed by a NetVLAD layer to perform matching. The experimental results evaluated on the KITTI dataset show that, with only a small set of training data, I2P-Rec achieves recall rates at Top-l % over 80% and 90%, when localizing monocular and stereo images on point cloud maps, respectively. We further evaluate I2P-Rec on a 1 km trajectory dataset collected by an autonomous logistics car and show that I2P- Rec can generalize well to previously unseen environments. Shuhang Zheng, Zhu Yu 0001, Beinan Yu, Si-Yuan Cao, Minhang Wang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Lun Luo |
IROS | 4 |