EDBT 2026 Demo / reviewers in the wild / expert
Lin Bie
dblp:221/0697
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-4844-1353ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph LearningabstractAlthough the fusion of images and LiDAR point clouds is crucial to many applications in computer vision, the relative poses of cameras and LiDAR scanners are often unknown. However, due to the modality and domain gap between images and LiDAR point clouds, Image-to-Point Cloud Registration is a significant challenge, especially when the image and point cloud come from non-synchronized frames. To tackle these issues, we introduce the virtual point cloud as a bridge to alleviate the cross-modality gap between images and LiDAR point clouds. In this way, the modality gap is converted to the domain gap of point clouds. Moreover, we introduce a virtual-spherical representation achieving orthogonal decoupling between pixel location and predicted depth. As for the domain gap, we propose a distribution-based adaptive sample module to generate a unified distribution of two types of point clouds. Then, we explore the correct correspondence pattern consistency and prune the false correspondences through a graph-based selection process. Experimental results demonstrate that our method outperforms the state-of-the-art methods by more than 10.77% and 12.53% performance on the KITTI Odometry and nuScenes datasets, respectively. The results demonstrate that our method can effectively solve non-synchronized random-frame registration. Lin Bie, Shouan Pan, Siqi Li 0001, Yue Gao 0002 |
CVPR | 1 |
| 2025 | Hyper-Depth: Hypergraph-Based Multi-Scale Representation Fusion for Monocular Depth Estimation
Lin Bie, Siqi Li 0001, Yifan Feng 0001, Yue Gao 0002 |
ICCV | 1 |
| 2025 | Channel pruning on frequency response
Lin Bie, Chenggang Yan 0001, Xibin Zhao, Yue Gao 0002 |
Sci. China Inf. Sci. | 3 |
| 2025 | Filter Pruning by High-Order Spectral ClusteringabstractLarge amount of redundancy is widely present in convolutional neural networks (CNNs). Identifying the redundancy in the network and removing the redundant filters is an effective way to compress the CNN model size with a minimal reduction in performance. However, most of the existing redundancy-based pruning methods only consider the distance information between two filters, which can only model simple correlations between filters. Moreover, we point out that distance-based pruning methods are not applicable for high-dimensional features in CNN models by our experimental observations and analysis. To tackle this issue, we propose a new pruning strategy based on high-order spectral clustering. In this approach, we use hypergraph structure to construct complex correlations among filters, and obtain high-order information among filters by hypergraph structure learning. Finally, based on the high-order information, we can perform better clustering on the filters and remove the redundant filters in each cluster. Experiments on various CNN models and datasets demonstrate that our proposed method outperforms the recent state-of-the-art works. For example, with ResNet50, we achieve a 57.1% FLOPs reduction with no accuracy drop on ImageNet, which is the first to achieve lossless pruning with such a high compression ratio. Yubo Zhang 0006, Lin Bie, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Multi-space Representation Fusion Enhanced Monocular Depth Estimation via Virtual Point CloudabstractMonocular Depth Estimation (MDE) is a fundamental problem in computer vision with broad applications in various downstream tasks. While recent studies focus on designing increasingly complex and powerful deep learning methods to regress depth maps directly, we propose a novel approach by introducing the Virtual Point Cloud (VPC) as an intermediate representation to provide the approximate geometric prior for the MDE task. In this article, we design a multi-scale multi-space representation fusion-enhanced MDE framework to address the challenges of MDE. Specifically, to resolve the issue of scale ambiguity, we design a VPC feature extraction module to learn multi-scale 3D geometric information for the depth prior. Then, we explicitly introduce geometric constraints for global depth prediction by incorporating a multi-space representation fusion from both the texture features in 2D space and the geometric features in 3D space. To mitigate errors at object boundaries, we introduce a confidence map generated based on the quality of the VPC to refine the predicted depth map. Specifically, we construct convolution receptive fields based on 3D spatial distances in spherical coordinates, ensuring that the confidence map provides reliable geometric guidance at object boundaries. Furthermore, we propose an independent confidence geometric consistency loss to supervise the refinement process. Experimental results demonstrate that our method significantly outperforms state-of-the-art approaches across all evaluation metrics on the KITTI and NYU-Depth-v2 datasets, achieving RMSE improvements of 9.2% and 2.8%, respectively. Moreover, zero-shot evaluations on the nuScenes and SUN-RGBD datasets further validate the generalizability of our approach. Lin Bie, Siqi Li 0001, Xiaopin Zhong, Zongze Wu 0001, Yue Gao 0012 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Arbitrary Large-Scale Scene Reconstruction without Annotated Block PartitionsabstractLarge-scale scene reconstruction is a challenging problem. As different parts of the scene could be visible from different collected image frames, previous works manually use distance or geography to decompose the scene into parts and reconstruct each part of the scene separately. However, such manual decomposition is a laborious and time-consuming task when applied to large-scale scene reconstruction in real-world applications. To address this, we propose VisibleNeRF automatically reconstructs large-scale scenes by decomposing scenes into parts based on the part visibility. More specifically, we propose a visibility judgment strategy to decompose the scenes into visible and invisible parts. Then we reconstruct the visible part with the corresponding collected images and continue to decompose the rest of the invisible parts with the proposed visibility judgment strategy. New NeRF modules are re-established for the decomposed invisible parts until the entire scene is reconstructed. To the best of our knowledge, we are the first to propose an online reconstruction of large-scale scenes without manual decomposition. Experimental results on three datasets show that our method successfully reconstructs large-scale scenes in a fully automatic manner. Besides, in the widely used Mission Bay dataset, our model outperforms other state-of-the-art methods by a large margin. Lin Bie, Siqi Li 0001, Dejian Guo, Shaoyi Du, Yue Gao 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | ColorPCR: Color Point Cloud Registration with Multi-Stage Geometric-Color FusionabstractPoint cloud registration is still a challenging and open problem. For example, when the overlap between two point clouds is extremely low, geo-only features may be not suf-ficient. Therefore, it is important to further explore how to utilize color data in this task. Under such circumstances, we propose ColorPCR for color point cloud registration with multi-stage geometric-color fusion. We design a Hier-archical Color Enhanced Feature Extraction module to ex-tract multi-level geometric-color features, and a GeoColor Superpoint Matching Module to encode transformation-invariant geo-color global context for robust patch corre-spondences. In this way, both geometric and color data can be used, thus leading to robust performance even under extremely challenging scenarios, such as low overlap between two point clouds. To evaluate the performance of our method, we colorize 3DMatch/3DLoMatch datasets as Color3DMatch/Color3DLoMatch and evaluations on these datasets demonstrate the effectiveness of our proposed method. Our method achieves state-of-the-art registration recall of 97.5%/88.9% on them. Juncheng Mu, Lin Bie, Shaoyi Du, Yue Gao 0002 |
CVPR | 2 |
| 2024 | Build a Cross-modality Bridge for Image-to-Point Cloud RegistrationabstractAlthough the fusion of images and LiDAR point clouds is crucial to many applications in computer vision, the relative poses of cameras and LiDAR scanners often need to be discovered. The general 2D-3D registration pipeline establishes correspondences and performs pose estimation based on the generated matches. However, 2D-3D correspondences are inherently challenging to establish due to the large gap between images and LiDAR point clouds. To this end, we build a bridge to alleviate the 2D-3D gap and propose a practical framework to align LiDAR point clouds to the virtual points generated by images. In this way, the modality gap is converted to the domain gap of point clouds. Furthermore, we propose a registration method with cross-domain feature extraction, frustum classification, and domain-agnostic correspondence pruning to narrow the domain gap to establish 3D-3D correspondences between real and virtual point clouds. Experimental results demonstrate that our image-to-point cloud registration method achieves state-of-theart performance on the KITTI Odometry and Oxford RobotCar datasets. Lin Bie, Shouan Pan |
ICME | 1 |
| 2024 | Image-to-Point Registration via Cross-Modality Correspondence RetrievalabstractImage-to-Point Cloud registration between 2D images and 3D LiDAR point clouds is a significant task in computer vision. The traditional registration pipeline first establishes correspondences between images and point clouds and then performs pose estimation based on the generated matches. However, 2D-3D correspondences are inherently difficult to be established due to the large modality gap between images and LiDAR point clouds. To this end, we build a bridge to alleviate the 2D-3D modality gap, which aligns LiDAR point clouds to the virtual points generated by images. In this way, the modality gap can be alleviated to the domain gap of different types of point clouds, i.e. original point clouds and virtual point clouds. Concretely, our framework conducts feature fusion from the LiDAR and virtual point cloud by utilizing the Transformer layer. To relieve the domain gap, a frustum points retrieval module and a combined correspondences retrieval module are proposed based on the consistency of the feature and position descriptor to select the correct correspondences among the candidates, which are generated from the simultaneous retrieval of features and position descriptors. In the implementation procedure, we design a frustum retrieval loss and a combined correspondence retrieval loss for cross-modality correspondence retrieval. Experimental results and comparison with state-of-the-art Image-to-Point Cloud methods on KITTI and nuScenes datasets demonstrate our proposed method has achieved superior performance. Lin Bie, Siqi Li 0001 |
ICMR | 1 |
| 2024 | Triadic Elastic Structure Representation for Open-Set Incremental 3D Object RetrievalabstractIn open-set environments, the introduction of unseen classes adds a third stage of data for incremental learning, alongside seen classes of old and new. Existing methods by binary optimization between old and new classes struggle to model the complex triadic relationships among old, new, and unseen classes. In this paper, we introduce the Triadic Elastic Structure Learning (TESR) framework for open-set incremental 3D object retrieval. Specifically, to overcome the Bi-Directional Catastrophic Forgetting (BDCF) problem in open-set incremental learning, we employ the Triadic Regularized Embedding (TRE) module. This module helps to globally preserve the triadic relationships among the three stages of classes, while simultaneously facilitating semantic-specific learning for new classes from a local perspective. To address the challenge of Bi-Directional Knowledge Transfer (BDKT), our method leverages high-order correlations among objects of different classes through the Elastic Structure Distillation (ESD) module, by constructing an elastic hypergraph based on incremental and global correlations. We construct four multi-modal datasets for this open-set incremental 3D object retrieval: OIES, OINT, OIMN, and OIAB. Extensive experiments and ablation studies on these four benchmarks demonstrate the superiority of our method over current state-of-the-art approaches. Yang Xu 0064, Yifan Feng 0001, Lin Bie |
ICMR | 3 |
| 2021 | The hierarchical task network planning method based on Monte Carlo Tree Search
Tianhao Shao, Lin Bie |
Knowl. Based Syst. | 5 |