VLDB 2026 Research / reviewers in the wild / expert
Lei Fan 0005
dblp:40/759-5
· DBLP profile ↗
16ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0001-9426-7029ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint ChangesabstractVisual localization, the task of determining the position and orientation of a camera, typically involves three core components: offline construction of a keyframe database, efficient online keyframes retrieval, and robust local feature matching. However, significant challenges arise when there are large viewpoint disparities between the query view and the database, such as attempting localization in a corridor previously build from an opposing direction. Intuitively, this issue can be addressed by synthesizing a set of virtual keyframes that cover all viewpoints. However, existing methods for synthesizing novel views to assist localization often fail to ensure geometric accuracy under large viewpoint changes. In this paper, we introduce a confidence-aware geometric prior into 2D Gaussian splatting to ensure the geometric accuracy of the scene. Then we can render novel views through the mesh with clear structures and accurate geometry, even under significant viewpoint changes, enabling the synthesis of a comprehensive set of virtual keyframes. Incorporating this geometry-preserving virtual keyframe database into the localization pipeline significantly enhances the robustness of visual localization. Yunxuan Li, Lei Fan 0005, Xiaoying Xing, Jianxiong Zhou, Ying Wu 0001 |
CVPR | 2 |
| 2024 | Evidential Active Recognition: Intelligent and Prudent Open-World Embodied PerceptionabstractActive recognition enables robots to intelligently explore novel observations, thereby acquiring more information while circumventing undesired viewing conditions. Recent approaches favor learning policies from simulated or collected data, wherein appropriate actions are more frequently selected when the recognition is accurate. However, most recognition modules are developed under the closed-world assumption, which makes them ill-equipped to handle unexpected inputs, such as the absence of the target object in the current observation. To address this issue, we propose treating active recognition as a sequential evidence-gathering process, providing by-step uncertainty quantification and reliable prediction under the evidence combination theory. Additionally, the reward function developed in this paper effectively characterizes the merit of actions when operating in open-world environments. To evaluate the performance, we collect a dataset from an indoor simulator, encompassing various recognition challenges such as distance, occlusion levels, and visibility. Through a series of experiments on recognition and robustness analysis, we demonstrate the necessity of introducing uncertainties to active recognition and the superior performance of the proposed method. Lei Fan 0005, Mingfu Liang, Yunxuan Li, Gang Hua 0001, Ying Wu 0001 |
CVPR | 1 |
| 2024 | Active Open-Vocabulary Recognition: Let Intelligent Moving Mitigate CLIP LimitationsabstractActive recognition, which allows intelligent agents to explore observations for better recognition performance, serves as a prerequisite for various embodied AI tasks, such as grasping, navigation and room arrangements. Given the evolving environment and the multitude of object classes, it is impractical to include all possible classes during the training stage. In this paper, we aim at advancing active open-vocabulary recognition, empowering embodied agents to actively perceive and classify arbitrary objects. However, directly adopting recent open-vocabulary classification models, like Contrastive Language Image Pretraining (CLIP), poses its unique challenges. Specifically, we observe that CLIP's performance is heavily affected by the viewpoint and occlusions, compromising its reliability in unconstrained embod-ied perception scenarios. Further, the sequential nature of observations in agent-environment interactions necessitates an effective method for integrating features that maintains discriminative strength for open-vocabulary classification. To address these issues, we introduce a novel agent for active open-vocabulary recognition. The proposed method leverages inter-frame and inter-concept similarities to navigate agent movements and to fuse features, without relying on class-specific knowledge. Compared to baseline CLIP model with 29.6% accuracy on ShapeNet dataset, the proposed agent could achieve 53.3% accuracy for open-vocabulary recognition, without any fine-tuning to the equipped CLIP model. Additional experiments conducted with the Habitat simulator further affirm the efficacy of our method. Lei Fan 0005, Jianxiong Zhou, Xiaoying Xing, Ying Wu 0001 |
CVPR | 1 |
| 2023 | Flexible Visual Recognition by Evidential Modeling of Confusion and IgnoranceabstractIn real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on unknown-class images. To tackle these deficiencies, flexible visual recognition should dynamically predict multiple classes when they are unconfident between choices and reject making predictions when the input is entirely out of the training distribution. Two challenges emerge along with this novel task. First, prediction uncertainty should be separately quantified as confusion depicting inter-class uncertainties and ignorance identifying out-of-distribution samples. Second, both confusion and ignorance should be comparable between samples to enable effective decision-making. In this paper, we propose to model these two sources of uncertainty explicitly with the theory of Subjective Logic. Regarding recognition as an evidence-collecting process, confusion is then defined as conflicting evidence, while ignorance is the absence of evidence. By predicting Dirichlet concentration parameters for singletons, comprehensive subjective opinions, including confusion and ignorance, could be achieved via further evidence combinations. Through a series of experiments on synthetic data analysis, visual recognition, and open-set detection, we demonstrate the effectiveness of our methods in quantifying two sources of uncertainties and dealing with flexible recognition. Lei Fan 0005, Bo Liu 0043, Ying Wu 0001, Gang Hua 0001 |
ICCV | 1 |
| 2023 | Avoiding Lingering in Learning Active Recognition by Adversarial DisturbanceabstractThis paper considers the active recognition scenario, where the agent is empowered to intelligently acquire observations for better recognition. The agents usually compose two modules, i.e., the policy and the recognizer, to select actions and predict the category. While using ground-truth class labels to supervise the recognizer, the policy is typically updated with rewards determined by the current in-training recognizer, like whether achieving correct predictions. However, this joint learning process could lead to unintended solutions, like a collapsed policy that only visits views that the recognizer is already sufficiently trained to obtain rewards, which harms the generalization ability. We call this phenomenon lingering to depict the agent being reluctant to explore challenging views during training. Existing approaches to tackle the exploration-exploitation trade-off could be ineffective as they usually assume reliable feedback during exploration to update the estimate of rarely-visited states. This assumption is invalid here as the reward from the recognizer could be insufficiently trained.To this end, our approach integrates another adversarial policy to constantly disturb the recognition agent during training, forming a competing game to promote active explorations and avoid lingering. The reinforced adversary, rewarded when the recognition fails, contests the recognition agent by turning the camera to challenging observations. Extensive experiments across two datasets validate the effectiveness of the proposed approach regarding its recognition performances, learning efficiencies, and especially robustness in managing environmental noises. Lei Fan 0005, Ying Wu 0001 |
WACV | 1 |
| 2022 | Unsupervised Depth Completion and Denoising for RGB-D SensorsabstractDepth information is considered valuable as it describes geometric structures, which benefits various robotic tasks. However, the depth acquired by RGB-D sensors still suffers from two deficiencies, i.e., incompletion and noises. Previous methods complete depth by exploring hand-tuned models or raising surface assumptions, while nowadays, deep approaches intend to solve this problem with rendered image pairs. For depth denoising, as a consequence of different sensor mechanisms, most methods can only work under specific devices. With existing methods, three challenges emerge: the onerous training set collecting process, the mismatch between existing models and present RGB-D sensors, and the non-real-time computation. In this paper, we first state depth completion and denoising are inherently different and without the need to collect or render complete and noiseless ground truths. We address all mentioned challenges with two separate un-supervised learning procedures. The completion network takes color and incomplete depth as input and predicts values to the unobserved area, which combines prior knowledge and color-depth correlations. The denoising step exploits image sequences to construct noise models in a self-supervised manner with the ability to cater to different sensors. Experimental comparisons and ablation studies demonstrate that even without human-labeled ground truths, the proposed method could produce better completion results and also reduce noises in real-time. Lei Fan 0005, Yunxuan Li, Ying Wu 0001 |
ICRA | 1 |
| 2021 | FLAR: A Unified Prototype Framework for Few-sample Lifelong Active RecognitionabstractIntelligent agents with visual sensors are allowed to actively explore their observations for better recognition performance. This task is referred to as Active Recognition (AR). Currently, most methods toward AR are implemented under a fixed-category setting, which constrains their applicability in realistic scenarios that need to incrementally learn new classes without retraining from scratch. Further, collecting massive data for novel categories is expensive. To address this demand, in this paper, we propose a unified framework towards Few-sample Lifelong Active Recognition (FLAR), which aims at performing active recognition on progressively arising novel categories that only have few training samples. Three difficulties emerge with FLAR: the lifelong recognition policy learning, the knowledge preservation of old categories, and the lack of training samples. To this end, our approach integrates prototypes, a robust representation for limited training samples, into a reinforcement learning solution, which motivates the agent to move towards views resulting in more discriminative features. Catastrophic forgetting during lifelong learning is then alleviated with knowledge distillation. Extensive experiments across two datasets, respectively for object and scene recognition, demonstrate that even without large training samples, the proposed approach could learn to actively recognize novel categories in a class-incremental behavior. Lei Fan 0005, Peixi Xiong, Ying Wu 0001 |
ICCV | 1 |
| 2021 | Semasuperpixel: A Multi-Channel Probability-Driven Superpixel Segmentation MethodabstractSuperpixel, an efficient image segmentation approach, aggregates a group of similar pixels into the same cluster. Existing superpixel algorithms still mainly focus on the color information while ignoring the semantic distribution knowledge. In this paper, we propose a semantic information-driven method that adopts multi-channel semantic probabilities for superpixel segmentation. By conducting statistical analysis on the semantic output and then formulating the distance measure, the prior knowledge of the semantic with a dynamic confidence value could be utilized by our method during the global update effectively. Extensive experimental evaluations show that our method achieves a leading segmentation quality and convergence speed, compared to other five state-of-the-art algorithms, as measured by boundary recall, undersegmentation error, and explained variation. Qingyun Zhao, Lei Fan 0005, Yuzhi Zhao, Qiong Yan, Long Chen 0005 |
ICIP | 3 |
| 2020 | Lightweight Single-Image Super-Resolution Network with Attentive Auxiliary Feature Learning
Qing Wang 0025, Yuzhi Zhao, Junchi Yan, Lei Fan 0005, Long Chen 0005 |
ACCV (2) | 5 |
| 2020 | Toward the Ghosting Phenomenon in a Stereo-Based Map With a Collaborative RGB-D RepairabstractAlthough 3-D reconstruction of dynamic road environment by moving cameras has been broadly applied in recognition and navigation systems, this task is still considered challenging, especially under circumstances with moving objects, where the reconstruction precision is strongly harassed by the ghosting problem. To address this issue, in this paper, we propose a novel approach for reconstructing 3-D maps of complete static scenes, based on a combination of an elaborately designed moving-object filtering mechanism and a map repairing and blank refilling procedure, where both plausible color and depth information from stereo image pairs are utilized. In this approach, first, we employ the planarity knowledge into the initial depth map based on the simple linear iterative cluster (SLIC) superpixel segmentation. The dynamic area in the image is determined under the supervision of odometry calculation. After wiping off moving objects, by collaboratively repairing color and depth information, the final 3-D map containing only static scene is obtained. The experimental results on extensive challenging real-world scenarios demonstrate the effectiveness and robustness of our approach. Jiasong Zhu, Lei Fan 0005, Wei Tian 0001, Long Chen 0005, Dongpu Cao, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | A Full Density Stereo Matching System Based on the Combination of CNNs and Slanted-PlanesabstractStereo matching methods consist of matching cost computation and several post processing steps. Deep learning methods have greatly raised the accuracy of matching cost and achieved the lowest error rate on several public datasets. However, their generality capabilities are not the best due to potential overfitting, which is the common problem of supervised learning approaches. This paper proposes a convolutional neural network (CNN) based cost estimation method for computing the similarity of image patches. In consideration of accuracy and generalization capability, small size convolution kernels are chosen in the convolution layer and dropout in the decision layer is used for preventing overfitting. After obtaining stereo matching cost from the output of the CNN, several post-processing operations are adopted for disparity optimization, which includes semi-global matching in 1-D from different directions, a left-right consistency check, and the slanted plane smoothing method. The method is evaluated on KITTI 2012, KITTI 2015, and Middlebury stereo datasets and the experimental results on the KITTI benchmark demonstrate the competitive accuracy performance of the approach. Additionally, to test the generalization of the method, a series of extended crossover experiments are conducted in which the training samples and testing samples come from different datasets. The results indicate superior generalization capability of our method than other supervised learning methods. Long Chen 0005, Lei Fan 0005, Jianda Chen, Dongpu Cao, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Monocular Outdoor Semantic Mapping with a Multi-task NetworkabstractIn many robotic applications, especially for the autonomous driving, understanding the semantic information and the geometric structure of surroundings are both essential. Semantic 3D maps, as a carrier of the environmental knowledge, are then intensively studied for their abilities and applications. However, it is still challenging to produce a dense outdoor semantic map from a monocular image stream. Motivated by this target, in this paper, we propose a method for large-scale 3D reconstruction from consecutive monocular images. First, with the correlation of underlying information between depth and semantic prediction, a novel multi-task Convolutional Neural Network (CNN) is designed for joint prediction. Given a single image, the network learns low-level information with a shared encoder and separately predicts with decoders containing additional Atrous Spatial Pyramid Pooling (ASPP) layers and the residual connection which merits disparities and semantic mutually. To overcome the inconsistency of monocular depth prediction for reconstruction, post-processing steps with the superpixelization and the effective 3D representation approach are obtained to give the final semantic map. Experiments are compared with other methods on both semantic labeling and depth prediction. We also qualitatively demonstrate the map reconstructed from large-scale, difficult monocular image sequences to prove the effectiveness and superiority. Yucai Bai, Lei Fan 0005, Ziyu Pan, Long Chen 0005 |
IROS | 2 |
| 2018 | Planecell: Representing Structural Space with Plane ElementsabstractReconstruction based on the stereo camera has received considerable attention recently, but two particular challenges still remain. The first concerns the need to present and compress data in an effective way, and the second is to maintain as much of the available information as possible while ensuring sufficient accuracy. To overcome these issues, we propose a new 3D representation method, namely, planecell, that extracts planarity from the depth-assisted image segmentation and then directly projects these depth planes into the 3D world. The proposed method demonstrates its advancement especially dealing with large-scale structural environment, such as autonomous driving scene. The reconstruction result of our method achieves equal accuracy compared to dense point clouds and compresses the output file 200 times. To further obtain global surfaces, an energy function formulated from Conditional Random Field that generalizes the planar relationships is maximized. We evaluate our method with reconstruction baselines on the KITTI outdoor scene dataset, and the results indicate the superiorities compared to other 3D space representation methods in accuracy, memory requirements and the scope of applications. Lei Fan 0005, Long Chen 0005, Kai Huang 0001, Dongpu Cao |
Intelligent Vehicles Symposium | 1 |
| 2017 | RGB-T SLAM: A flexible SLAM framework by combining appearance and thermal informationabstractVisual SLAM in low illumination scenes remains a considerably challenging task since the available amount of appearance information frequently stays insufficient. To tackle with this problem, we propose a novel SLAM framework by using both appearance information and thermal information, which possesses illumination-free recognizable contents, in a flexible manner. The key idea is to continuously update a RGB-T map, which contains both RGB and thermal map points to implement location and mapping. More specifically, in our SLAM system, we detect features in both RGB and thermal images and combine them together to update the RGB-T map and implement simultaneous location and mapping. Both quantitative and qualitative results demonstrate the effectiveness of our framework, especially under low illumination environments. Long Chen 0005, Libo Sun 0002, Lei Fan 0005, Kai Huang 0001, Zhe Xuanyuan |
ICRA | 4 |
| 2017 | Let the robot tell: Describe car image with natural language via LSTM
Long Chen 0005, Lei Fan 0005 |
Pattern Recognit. Lett. | 3 |
| 2017 | Moving-Object Detection From Consecutive Stereo Pairs Using Slanted Plane SmoothingabstractDetecting moving objects is of great importance for autonomous unmanned vehicle systems, and a challenging task especially in complex dynamic environments. This paper proposes a novel approach for the detection of moving objects and the estimation of their motion states using consecutive stereo image pairs on mobile platforms. First, we use a variant of the semi-global matching algorithm to compute initial disparity maps. Second, assisted by the initial disparities, boundaries in the image segmentation produced by simple linear iterative clustering are classified into coplanar, hinge, and occlusion. Moving points are obtained during ego-motion estimation by a modified random sample consensus) algorithm without resorting to time-consuming dense optical flow. Finally, the moving objects are extracted by merging superpixels according to the boundary types and their movements. The proposed method is accelerated on the GPU at 20 frames per second. The data which we use for testing and benchmarking is released, thus completing similar data sets. It includes 812 image pairs and 924 moving objects with ground truth for better algorithms evaluation. Experimental results demonstrate that the proposed method achieves competitive results in terms of moving-object detection and their motion state estimation in challenging urban scenarios. Long Chen 0005, Lei Fan 0005, Guodong Xie, Kai Huang 0001, Andreas Nüchter |
IEEE Trans. Intell. Transp. Syst. | 2 |