EDBT 2026 Demo / reviewers in the wild / expert
Zhipeng Cai 0003
dblp:14/5155-3
· DBLP profile ↗
22ranked-venue papers
6as first author
15since 2021 · last 2026
0009-0008-2469-1316ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniEvent: Unified Event Representation LearningabstractEvent cameras have gained increasing popularity in computer vision due to their ultra-high dynamic range and temporal resolution. However, event networks heavily rely on task-specific designs due to the unstructured data distribution and spatial-temporal (S-T) inhomogeneity, making it hard to reuse existing architectures for new tasks. We propose OmniEvent, an innovative unified event representation learning framework that achieves SOTA performance across diverse tasks, fully removing the need for task-specific designs. Unlike previous methods that treat event data as 3D point clouds with manually tuned S-T scaling weights, OmniEvent proposes a decouple-enhance-fuse paradigm, where the local feature aggregation and enhancement are done independently on the spatial and temporal domains to avoid inhomogeneity issues. Space-filling curves are applied to enable large receptive fields while improving memory and compute efficiency. The features from individual domains are then fused by attention to learn S-T interactions. The output of OmniEvent is a grid-shaped tensor, which enables standard vision models to process event data without architectural changes. With a unified framework and similar hyperparameters, OmniEvent outperforms (task-specific) SOTA by up to 68.2% across 3 representative tasks and 10 datasets (Fig. 1). Weiqi Yan 0005, Chenlu Lin, Youbiao Wang, Zhipeng Cai 0003, Xiuhong Lin, Yangyang Shi, Weiquan Liu |
AAAI | 4 |
| 2025 | ConDo: Continual Domain Expansion for Absolute Pose RegressionabstractVisual localization is a fundamental machine learning problem. Absolute Pose Regression (APR) trains a scene-dependent model to efficiently map an input image to the camera pose in a pre-defined scene. However, many applications have continually changing environments, where inference data at novel poses or scene conditions (weather, geometry) appear after deployment. Training APR on a fixed dataset leads to overfitting, making it fail catastrophically on challenging novel data. This work proposes Continual Domain Expansion (ConDo), which continually collects unlabeled inference data to update the deployed APR. Instead of applying standard unsupervised domain adaptation methods which are ineffective for APR, ConDo effectively learns from unlabeled data by distilling knowledge from scene-agnostic localization methods. By sampling data uniformly from historical and newly collected data, ConDo can effectively expand the generalization domain of APR. Large-scale benchmarks with various scene types are constructed to evaluate models under practical (long-term) data changes. ConDo consistently and significantly outperforms baselines across architectures, scene types, and data changes. On challenging scenes (Fig.1), it reduces the localization error by >7x (14.8m vs 1.7m). Analysis shows the robustness of ConDo against compute budgets, replay buffer sizes and teacher prediction noise. Comparing to model re-training, ConDo achieves similar performance up to 25x faster. Zijun Li 0006, Zhipeng Cai 0003, Bochun Yang, Xuelun Shen, Xiaoliang Fan, Michael Paulitsch, Cheng Wang 0003 |
AAAI | 2 |
| 2025 | MonSter: Marry Monodepth to Stereo Unleashes PowerabstractStereo matching recovers depth from image correspondences. Existing methods struggle to handle ill-posed regions with limited matching cues, such as occlusions and textureless areas. To address this, we propose MonSter, a novel method that leverages the complementary strengths of monocular depth estimation and stereo matching. MonSter integrates monocular depth and stereo matching into a dual-branch architecture to iteratively improve each other. Confidence-based guidance adaptively selects reliable stereo cues for monodepth scale-shift recovery. The refined monodepth is in turn guides stereo effectively at ill-posed regions. Such iterative mutual enhancement enables MonSter to evolve monodepth priors from coarse object-level structures to pixel-level geometry, fully unlocking the potential of stereo matching. As shown in Fig. 2, MonSter ranks 1stacross five most commonly used leaderboards — SceneFlow, KITTI 2012, KITTI 2015, Middlebury, and ETH3D. Achieving up to 49.5% improvements (Bad 1.0 on ETH3D) over the previous best method. Comprehensive analysis verifies the effectiveness of MonSter in ill-posed regions. In terms of zero-shot generalization, MonSter significantly and consistently outperforms state-of-the-art across the board. The code is publicly available at: https://github.com/Junda24/MonSter. Junda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang 0001, Zhaoxing Zhang, Jinliang Zang, Yurui Chen, Zhipeng Cai 0003, Xin Yang 0008 |
CVPR | 9 |
| 2024 | SimCS: Simulation for Domain Incremental Online Continual SegmentationabstractContinual Learning is a step towards lifelong intelligence where models continuously learn from recently collected data without forgetting previous knowledge. Existing continual learning approaches mostly focus on image classification in the class-incremental setup with clear task boundaries and unlimited computational budget. This work explores the problem of Online Domain-Incremental Continual Segmentation (ODICS), where the model is continually trained over batches of densely labeled images from different domains, with limited computation and no information about the task boundaries. ODICS arises in many practical applications. In autonomous driving, this may correspond to the realistic scenario of training a segmentation model over time on a sequence of cities. We analyze several existing continual learning methods and show that they perform poorly in this setting despite working well in class-incremental segmentation. We propose SimCS, a parameter-free method complementary to existing ones that uses simulated data to regularize continual learning. Experiments show that SimCS provides consistent improvements when combined with different CL methods. Motasem Alfarra, Zhipeng Cai 0003, Adel Bibi, Bernard Ghanem, Matthias Müller 0001 |
AAAI | 2 |
| 2024 | L-MAGIC: Language Model Assisted Generation of Images with CoherenceabstractIn the current era of generative AI breakthroughs, generating panoramic scenes from a single input image remains a key challenge. Most existing methods use diffusion-based iterative or simultaneous multi-view inpainting. However, the lack of global scene layout priors leads to subpar outputs with duplicated objects (e.g., multiple beds in a bedroom) or requires time-consuming human text inputs for each view. We propose L-MAGIC, a novel method leveraging large language models for guidance while diffusing multiple coherent views of 360° panoramic scenes. L-MAGIC harnesses pretrained diffusion and language models without fine-tuning, ensuring zero-shot performance. The output quality is further enhanced by super-resolution and multi-view fusion techniques. Extensive experiments demonstrate that the resulting panoramic scenes feature better scene layouts and perspective view rendering quality compared to related works, with >70% preference in human evaluations. Combined with conditional diffusion models, L-MAGIC can accept various input modalities, including but not limited to text, depth maps, sketches, and colored scripts. Applying depth estimation further enables 3D point cloud generation and dynamic scene exploration with fluid camera motion. Code is available at https://github.com/ZhipengCai/L-MAGIC-code-release. Zhipeng Cai 0003, Matthias Müller 0011, Reiner Birkl, Diana Wofk, Shao-Yen Tseng, Junda Cheng, Gabriela Ben Melech Stan, Vasudev Lal, Michael Paulitsch |
CVPR | 1 |
| 2024 | LiSA: LiDAR Localization with Semantic AwarenessabstractLiDAR localization is a fundamental task in robotics and computer vision, which estimates the pose of a Li- DAR point cloud within a global map. Scene Coordinate Regression (SCR) has demonstrated state-of-the-art performance in this task. In SCR, a scene is represented as a neural network, which outputs the world coordinates for each point in the input point cloud. However, SCR treats all points equally during localization, ignoring the fact that not all objects are beneficial for localization. For exam-ple, dynamic objects and repeating structures often negatively impact SCR. To address this problem, we introduce LiSA, the first method that incorporates semantic aware-ness into SCR to boost the localization robustness and accuracy. To avoid extra computation or network parame-ters during inference, we distill the knowledge from a seg-mentation model to the original SCR network. Experi-ments show the superior performance of LiSA on standard LiDAR localization benchmarks compared to state-of-the- art methods. Applying knowledge distillation not only pre-serves high efficiency but also achieves higher localization accuracy than introducing extra semantic segmentation modules. We also analyze the benefit of semantic in-formation for LiDAR localization. Our code is released at https://github.com/Ybchun/LiSA. Bochun Yang, Zijun Li 0006, Wen Li 0005, Zhipeng Cai 0003, Chenglu Wen, Matthias Müller 0011, Cheng Wang 0003 |
CVPR | 4 |
| 2024 | GIM: Learning Generalizable Image Matcher From Internet VideosabstractImage matching is a fundamental computer vision problem. While learning-based methods achieve state-of-the-art performance on existing benchmarks, they generalize poorly to in-the-wild images. Such methods typically need to train separate models for different scene types (e.g., indoor vs. outdoor) and are impractical when the scene type is unknown in advance. One of the underlying problems is the limited scalability of existing data construction pipelines, which limits the diversity of standard image matching datasets. To address this problem, we propose GIM, a self-training framework for learning a single generalizable model based on any image matching architecture using internet videos, an abundant and diverse data source. Given an architecture, GIM first trains it on standard domain-specific datasets and then combines it with complementary matching methods to create dense labels on nearby frames of novel videos. These labels are filtered by robust fitting, and then enhanced by propagating them to distant frames. The final model is trained on propagated data with strong augmentations. Not relying on complex 3D reconstruction makes GIM much more efficient and less likely to fail than standard SfM-and-MVS based frameworks. We also propose ZEB, the first zero-shot evaluation benchmark for image matching. By mixing data from diverse domains, ZEB can thoroughly assess the cross-domain generalization performance of different methods. Experiments demonstrate the effectiveness and generality of GIM. Applying GIM consistently improves the zero-shot performance of 3 state-of-the-art image matching architectures as the number of downloaded videos increases (Fig. 1 (a)); with 50 hours of YouTube videos, the relative zero-shot performance improves by 6.9% − 18.1%. GIM also enables generalization to extreme cross-domain data such as Bird Eye View (BEV) images of projected 3D point clouds (Fig. 1 (c)). More importantly, our single zero-shot model consistently outperforms domain-specific baselines when evaluated on downstream tasks inherent to their respective domains. The code will be released upon acceptance. Xuelun Shen, Zhipeng Cai 0003, Wei Yin 0006, Matthias Müller 0011, Zijun Li 0006, Xiaozhi Chen, Cheng Wang 0003 |
ICLR | 2 |
| 2024 | Evaluation of Test-Time Adaptation Under Computational Time ConstraintsabstractThis paper proposes a novel online evaluation protocol for Test Time Adaptation (TTA) methods, which penalizes slower methods by providing them with fewer samples for adaptation. TTA methods leverage unlabeled data at test time to adapt to distribution shifts. Though many effective methods have been proposed, their impressive performance usually comes at the cost of significantly increased computation budgets. Current evaluation protocols overlook the effect of this extra computation cost, affecting their real-world applicability. To address this issue, we propose a more realistic evaluation protocol for TTA methods, where data is received in an online fashion from a constant-speed data stream, thereby accounting for the method’s adaptation speed. We apply our proposed protocol to benchmark several TTA methods on multiple datasets and scenarios. Extensive experiments shows that, when accounting for inference speed, simple and fast approaches can outperform more sophisticated but slower methods. For example, SHOT from 2020, outperforms the state-of-the-art method SAR from 2023 under our online setting. Our results reveal the importance of developing practical TTA methods that are both accurate and efficient. Motasem Alfarra, Hani Itani, Alejandro Pardo, Shyma Alhuwaider, Merey Ramazanova, Juan C. Pérez, Zhipeng Cai 0003, Matthias Müller 0011, Bernard Ghanem |
ICML | 7 |
| 2024 | Slack-Free Spiking Neural Network Formulation for Hypergraph Minimum Vertex CoverabstractNeuromorphic computers open up the potential of energy-efficient computation using spiking neural networks (SNN), which consist of neurons that exchange spike-based information asynchronously. In particular, SNNs have shown promise in solving combinatorial optimization. Underpinning the SNN methods is the concept of energy minimization of an Ising model, which is closely related to quadratic unconstrained binary optimization (QUBO). Thus, the starting point for many SNN methods is reformulating the target problem as QUBO, then executing an SNN-based QUBO solver. For many combinatorial problems, the reformulation entails introducing penalty terms, potentially with slack variables, that implement feasibility constraints in the QUBO objective. For more complex problems such as hypergraph minimum vertex cover (HMVC), numerous slack variables are introduced which drastically increase the search domain and reduce the effectiveness of the SNN solver. In this paper, we propose a novel SNN formulation for HMVC. Rather than using penalty terms with slack variables, our SNN architecture introduces additional spiking neurons with a constraint checking and correction mechanism that encourages convergence to feasible solutions. In effect, our method obviates the need for reformulating HMVC as QUBO. Experiments on neuromorphic hardware show that our method consistently yielded high quality solutions for HMVC on real and synthetic instances where the SNN-based QUBO solver often failed, while consuming measurably less energy than global solvers on CPU. Anh-Dzung Doan, Zhipeng Cai 0003, Tat-Jun Chin |
NeurIPS | 3 |
| 2024 | Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal EstimationabstractWe introduce Metric3D v2, a geometric foundation model designed for zero-shot metric depth and surface normal estimation from single images, critical for accurate 3D recovery. Depth and normal estimation, though complementary, present distinct challenges. State-of-the-art monocular depth methods achieve zero-shot generalization through affine-invariant depths, but fail to recover real-world metric scale. Conversely, current normal estimation techniques struggle with zero-shot performance due to insufficient labeled data. We propose targeted solutions for both metric depth and normal estimation. For metric depth, we present a canonical camera space transformation module that resolves metric ambiguity across various camera models and large-scale datasets, which can be easily integrated into existing monocular models. For surface normal estimation, we introduce a joint depth-normal optimization module that leverages diverse data from metric depth, allowing normal estimators to improve beyond traditional labels. Our model, trained on over 16 million images from thousands of camera models with varied annotations, excels in zero-shot generalization to new camera settings. As shown in Fig. 1, It ranks the 1st in multiple zero-shot and standard benchmarks for metric depth and surface normal prediction. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. Our model also relieves the scale drift issues of monocular-SLAM (Fig. 3), leading to high-quality metric scale dense mapping. Such applications highlight the versatility of Metric3D v2 models as geometric foundation models. Mu Hu, Wei Yin 0006, Chi Zhang 0007, Zhipeng Cai 0003, Xiaoxiao Long, Hao Chen 0041, Gang Yu 0002, Chunhua Shen, Shaojie Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | GSDDet: Ground Sample Distance-Guided Object Detection for Remote Sensing ImagesabstractObject detection for remote sensing images (ODRSI) is an important task in computer vision. Effective algorithms inspired by oriented object detection have been proposed recently. However, a major challenge still remains. Different categories of objects may be similar under different scales, causing cross-scale confusion. Different from natural images, remote sensing images have a consistent scale within the same image, which is generally referred to as Ground Sample Distance (GSD). In this paper, we show, that GSD can be utilized to address the cross-scale confusion problem, and effectively boost the performance of ODRSI. Specifically, we propose GSDDet, which embeds the deep features that represent GSD constraints to decrease the cross-scale confusion between different object categories. In GSDDet, a deep GSD classification network is first designed to extract the GSD deep features from remote sensing images. Then, the GSD deep feature is coupled with an attention framework to detect multiple categories of objects. Due to the simplicity of our framework, GSDDet can be applied to improve both one-stage and two-stage methods. Experiments demonstrate that GSDDet outperforms state-of-the-art methods on challenging benchmarks, including DOTA-v1.0, DOTA-v1.5, and HRSC2016. The source code will be released upon publication. Yunuo Yang, Cheng Wang 0003, Zhipeng Cai 0003, Pinqing Song, Guanjie Huang, Ming Cheng 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageabstractReconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA) monocular metric depth estimation methods can only handle a single camera model and are unable to perform mixed-data training due to metric ambiguity. Meanwhile, SOTA monocular methods trained on large mixed datasets achieve zero-shot generalization by learning affine-invariant depths, which cannot recover real-world metrics. In this work, we show that the key to a zero-shot single-view metric depth model lies in the combination of large-scale data training and resolving the metric ambiguity from various camera models. We propose a canonical camera space transformation module, which explicitly addresses the ambiguity problems and can be effortlessly plugged into existing monocular models. Equipped with our module, monocular models can be stably trained over 8 millions of images with thousands of camera models, resulting in zero-shot generalization to in-the-wild images with unseen camera settings. Experiments demonstrate SOTA performance of our method on 7 zero-shot benchmarks. Notably, our method won the championship in the 2nd Monocular Depth Estimation Challenge. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. The potential benefits extend to downstream tasks, which can be significantly improved by simply plugging in our model. For example, our model relieves the scale drift issues of monocular-SLAM (Fig. 1), leading to high-quality metric scale dense mapping. The code is available at https://github.com/YvanYin/Metric3D. Wei Yin 0006, Chi Zhang 0007, Hao Chen 0041, Zhipeng Cai 0003, Gang Yu 0002, Xiaozhi Chen, Chunhua Shen |
ICCV | 4 |
| 2023 | CLNeRF: Continual Learning Meets NeRFabstractNovel view synthesis aims to render unseen views given a set of calibrated images. In practical applications, the coverage, appearance or geometry of the scene may change over time, with new images continuously being captured. Efficiently incorporating such continuous change is an open challenge. Standard NeRF benchmarks only involve scene coverage expansion. To study other practical scene changes, we propose a new dataset, World Across Time (WAT), consisting of scenes that change in appearance and geometry over time. We also propose a simple yet effective method, CLNeRF, which introduces continual learning (CL) to Neural Radiance Fields (NeRFs). CLNeRF combines generative replay and the Instant Neural Graphics Primitives (NGP) architecture to effectively prevent catastrophic forgetting and efficiently update the model when new data arrives. We also add trainable appearance and geometry embeddings to NGP, allowing a single compact model to handle complex scene changes. Without the need to store historical images, CLNeRF trained sequentially over multiple scans of a changing scene performs on-par with the upper bound model trained on all scans at once. Compared to other CL baselines CLNeRF performs much better across standard benchmarks and WAT. The source code, a demo, and the WAT dataset are available at https://github.com/IntelLabs/CLNeRF. Zhipeng Cai 0003, Matthias Müller 0001 |
ICCV | 1 |
| 2023 | E2PNet: Event to Point Cloud Registration with Spatio-Temporal Representation LearningabstractEvent cameras have emerged as a promising vision sensor in recent years due to their unparalleled temporal resolution and dynamic range. While registration of 2D RGB images to 3D point clouds is a long-standing problem in computer vision, no prior work studies 2D-3D registration for event cameras. To this end, we propose E2PNet, the first learning-based method for event-to-point cloud registration.
The core of E2PNet is a novel feature representation network called Event-Points-to-Tensor (EP2T), which encodes event data into a 2D grid-shaped feature tensor. This grid-shaped feature enables matured RGB-based frameworks to be easily used for event-to-point cloud registration, without changing hyper-parameters and the training procedure. EP2T treats the event input as spatio-temporal point clouds. Unlike standard 3D learning architectures that treat all dimensions of point clouds equally, the novel sampling and information aggregation modules in EP2T are designed to handle the inhomogeneity of the spatial and temporal dimensions. Experiments on the MVSEC and VECtor datasets demonstrate the superiority of E2PNet over hand-crafted and other learning-based methods. Compared to RGB-based registration, E2PNet is more robust to extreme illumination or fast motion due to the use of event data. Beyond 2D-3D registration, we also show the potential of EP2T for other vision tasks such as flow estimation, event-to-image reconstruction and object recognition. The source code can be found at: https://github.com/Xmu-qcj/E2PNet. Xiuhong Lin, Changjie Qiu, Zhipeng Cai 0003, Weiquan Liu, Xuesheng Bian, Matthias Müller 0011, Cheng Wang 0003 |
NeurIPS | 3 |
| 2021 | Online Continual Learning with Natural Distribution Shifts: An Empirical Study with Visual DataabstractContinual learning is the problem of learning and retaining knowledge through time over multiple tasks and environments. Research has primarily focused on the incremental classification setting, where new tasks/classes are added at discrete time intervals. Such an "offline" setting does not evaluate the ability of agents to learn effectively and efficiently, since an agent can perform multiple learning epochs without any time limitation when a task is added. We argue that "online" continual learning, where data is a single continuous stream without task boundaries, enables evaluating both information retention and online learning efficacy. In online continual learning, each incoming small batch of data is first used for testing and then added to the training set, making the problem truly online. Trained models are later evaluated on historical data to assess information retention. We introduce a new benchmark for online continual visual learning that exhibits large scale and natural distribution shifts. Through a large-scale analysis, we identify critical and previously unobserved phenomena of gradient-based optimization in continual learning, and propose effective strategies for improving gradient-based online continual learning with real data. The source code and dataset are available in: https://github.com/IntelLabs/continuallearning. Zhipeng Cai 0003, Ozan Sener, Vladlen Koltun |
ICCV | 1 |
| 2020 | Robust Fitting in Computer Vision: Easy or Hard?
Tat-Jun Chin, Zhipeng Cai 0003, Frank Neumann 0001 |
Int. J. Comput. Vis. | 2 |
| 2019 | Consensus Maximization Tree Search RevisitedabstractConsensus maximization is widely used for robust fitting in computer vision. However, solving it exactly, i.e., finding the globally optimal solution, is intractable. A* tree search, which has been shown to be fixed-parameter tractable, is one of the most efficient exact methods, though it is still limited to small inputs. We make two key contributions towards improving A* tree search. First, we show that the consensus maximization tree structure used previously actually contains paths that connect nodes at both adjacent and non-adjacent levels. Crucially, paths connecting non-adjacent levels are redundant for tree search, but they were not avoided previously. We propose a new acceleration strategy that avoids such redundant paths. In the second contribution, we show that the existing branch pruning technique also deteriorates quickly with the problem dimension. We then propose a new branch pruning technique that is less dimension-sensitive to address this issue. Experiments show that both new techniques can significantly accelerate A* tree search, making it reasonably efficient on inputs that were previously out of reach. Demo code is available at https://github.com/ZhipengCai/MaxConTreeSearch. Zhipeng Cai 0003, Tat-Jun Chin, Vladlen Koltun |
ICCV | 1 |
| 2018 | Deterministic Consensus Maximization with Biconvex Programming
Zhipeng Cai 0003, Tat-Jun Chin, Huu Le, David Suter |
ECCV (12) | 1 |
| 2018 | Robust Fitting in Computer Vision: Easy or Hard?
Tat-Jun Chin, Zhipeng Cai 0003, Frank Neumann 0001 |
ECCV (12) | 2 |
| 2016 | Patch-Based Semantic Labeling of Road Scene Using Colorized Mobile LiDAR Point CloudsabstractSemantic labeling of road scenes using colorized mobile LiDAR point clouds is of great significance in a variety of applications, particularly intelligent transportation systems. However, many challenges, such as incompleteness of objects caused by occlusion, overlapping between neighboring objects, interclass local similarities, and computational burden brought by a huge number of points, make it an ongoing open research area. In this paper, we propose a novel patch-based framework for labeling road scenes of colorized mobile LiDAR point clouds. In the proposed framework, first, three-dimensional (3-D) patches extracted from point clouds are used to construct a 3-D patch-based match graph structure (3D-PMG), which transfers category labels from labeled to unlabeled point cloud road scenes efficiently. Then, to rectify the transferring errors caused by local patch similarities in different categories, contextual information among 3-D patches is exploited by combining 3D-PMG with Markov random fields. In the experiments, the proposed framework is validated on colorized mobile LiDAR point clouds acquired by the RIEGL VMX-450 mobile LiDAR system. Comparative experiments show the superior performance of the proposed framework for accurate semantic labeling of road scenes. Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Zhipeng Cai 0003, Ziyi Chen 0001, Hanyun Wang, Yongtao Yu, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2016 | Spatial-Related Traffic Sign Inspection for Inventory Purposes Using Mobile Laser Scanning DataabstractThis paper presents a spatial-related traffic sign inspection process for sign type, position, and placement using mobile laser scanning (MLS) data acquired by a RIEGL VMX-450 system and presents its potential for traffic sign inventory applications. First, the paper describes an algorithm for traffic sign detection in complicated road scenes based on the retroreflectivity properties of traffic signs in MLS point clouds. Then, a point cloud-to-image registration process is proposed to project the traffic sign point clouds onto a 2-D image plane. Third, based on the extracted traffic sign points, we propose a traffic sign position and placement inspection process by creating geospatial relations between the traffic signs and road environment. For further inventory applications, we acquire several spatial-related inventory measurements. Finally, a traffic sign recognition process is conducted to assign sign type. With the acquired sign type, position, and placement data, a spatial-associated sign network is built. Experimental results indicate satisfactory performance of the proposed detection, recognition, position, and placement inspection algorithms. The experimental results also prove the potential of MLS data for automatic traffic sign inventory applications. Chenglu Wen, Jonathan Li 0001, Huan Luo 0001, Yongtao Yu, Zhipeng Cai 0003, Hanyun Wang, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2015 | Occluded Boundary Detection for Small-Footprint Groundborne LIDAR Point Cloud Guided by Last EchoabstractOccluded boundary detection in a 3-D point cloud is an indispensable preprocessing step for many applications, such as point cloud completion. Meanwhile, existing methods do not have the ability of distinguishing occluded and complete surface borders, such as the border of a sign board. To solve this problem, this letter presents an occluded boundary detection method for small-footprint LIDAR point clouds. The main novelty of this letter is using the last-echo information for occluded boundary detection. Seed boundary (SB) points are subsequently detected using this last-echo information. Finally, the SB points are grown into neighboring points using an occluded boundary growth algorithm. To the best of our knowledge, this method is the first method that uses the last-echo information to detect occluded boundaries. Experimental results with comparisons indicate that the proposed method can accurately and efficiently detect an occluded boundary without contamination from a complete surface border. These advantages allow the proposed method to benefit further applications such as point cloud completion, as demonstrated in the application section. Zhipeng Cai 0003, Cheng Wang 0003, Chenglu Wen, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |