Pengyu Yin

dblp:257/1484 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Large-Scale UWB Anchor Calibration and One-Shot Localization Using Gaussian Process
abstract
Ultra-wideband (UWB) is gaining popularity with devices like AirTags for precise home item localization but faces significant challenges when scaled to large environments like seaports. The main challenges are calibration and localization under obstructed conditions, which are common in logistics environments. Traditional calibration methods, dependent on line-of-sight (LoS), are slow, costly, and unreliable in seaports and warehouses, making large-scale localization a significant pain point in the industry. To overcome these challenges, we propose a one-shot calibration and localization framework based on UWB-LiDAR fusion. Our method uses Gaussian processes to estimate the anchor position from continuous-time LiDAR Inertial Odometry with sampled UWB ranges. This approach ensures accurate and reliable calibration with only one round of sampling in large-scale areas, i.e.,$600 \times 450 ~\mathrm{m}^{2}$. With LoS issues, UWB-only localization can be problematic, even when anchor positions are known. We demonstrate that by applying a UWB-range filter, the search range for LiDAR loop closure descriptors is significantly reduced, improving both accuracy and speed. This concept can be applied to other loop closure detection methods, enabling cost-effective localization in large-scale warehouses and seaports. It significantly improves precision in challenging environments where the UWB-only and LiDAR-Inertial methods fail, as shown in the video https://https://youtu.be/oY8jQKdM7lU. We will open-source our datasets and calibration codes for community use.
Shenghai Yuan 0001, Boyang Lou, Thien-Minh Nguyen, Pengyu Yin, Muqing Cao, Xinghang Xu, Jianping Li 0004, Jie Xu 0066, Siyu Chen 0036, Lihua Xie 0001
ICRA4
2024 MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception
abstract
Perception plays a crucial role in various robot applications. However, existing well-annotated datasets are biased towards autonomous driving scenarios, while unlabelled SLAM datasets are quickly over-fitted, and often lack environment and domain variations. To expand the frontier of these fields, we introduce a comprehensive dataset named MCD (Multi-Campus Dataset), featuring a wide range of sensing modalities, high-accuracy ground truth, and diverse challenging environments across three Eurasian university campuses. MCD comprises both CCS (Classical Cylindrical Spinning) and NRE (Non-Repetitive Epicyclic) lidars, high-quality IMUs (Inertial Measurement Units), cameras, and UWB (Ultra-WideBand) sensors. Further-more, in a pioneering effort, we introduce semantic annotations of 29 classes over 59k sparse NRE lidar scans across three domains, thus providing a novel challenge to existing semantic segmentation research upon this largely unexplored modality. Finally, we propose, for the first time to the best of our knowledge, continuous-time ground truth based on optimization-based registration of lidar-inertial data on three survey-grade prior maps, each several times larger than the next largest publicly available ones. We conduct a rigorous evaluation of numerous state-of-the-art algorithms on MCD, report their performance, and highlight the challenges awaiting solutions from the research community.
Thien-Minh Nguyen, Shenghai Yuan 0001, Thien Hoang Nguyen, Pengyu Yin, Haozhi Cao, Lihua Xie 0001, Maciej Wozniak 0001, Patric Jensfelt, Marko Thiel 0002, Justin Ziegenbein, Noel Blunder
CVPR4
2024 Reliable Spatial-Temporal Voxels For Multi-modal Test-Time Adaptation
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Xingyu Ji, Shenghai Yuan 0001, Lihua Xie 0001
ECCV (28)4
2024 MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation
abstract
Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-imbalanced performance, restricting their adoption in real applications. This imbalanced performance is mainly caused by: 1) self-training with imbalanced data and 2) the lack of pixel-wise 2D supervision signals. In this work, we propose Multi-modal Prior Aided (MoPA) domain adaptation to improve the performance of rare objects. Specifically, we develop Valid Ground-based Insertion (VGI) to rectify the imbalance supervision signals by inserting prior rare objects collected from the wild while avoiding introducing artificial artifacts that lead to trivial solutions. Meanwhile, our SAM consistency loss leverages the 2D prior semantic masks from SAM as pixel-wise supervision signals to encourage consistent predictions for each object in the semantic mask. The knowledge learned from modal-specific prior is then shared across modalities to achieve better rare object segmentation. Extensive experiments show that our method achieves state-of-the-art performance on the challenging MM-UDA benchmark. Code will be available at https://github.com/AronCao49/MoPA.
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001
ICRA4
2024 Outram: One-shot Global Localization via Triangulated Scene Graph and Global Outlier Pruning
abstract
One-shot LiDAR localization refers to the ability to estimate the robot pose from one single point cloud, which yields significant advantages in initialization and relocalization processes. In the point cloud domain, the topic has been extensively studied as a global descriptor retrieval (i.e., loop closure detection) and pose refinement (i.e., point cloud registration) problem both in isolation or combined. However, few have explicitly considered the relationship between candidate retrieval and correspondence generation in pose estimation, leaving them brittle to substructure ambiguities. To this end, we propose a hierarchical one-shot localization algorithm called Outram that leverages substructures of 3D scene graphs for locally consistent correspondence searching and global substructure-wise outlier pruning. Such a hierarchical process couples the feature retrieval and the correspondence extraction to resolve the substructure ambiguities by conducting a local-to-global consistency refinement. We demonstrate the capability of Outram in a variety of scenarios in multiple large-scale outdoor datasets. Our implementation is open-sourced: https://github.com/Pamphlett/Outram.
Pengyu Yin, Haozhi Cao, Thien-Minh Nguyen, Shenghai Yuan 0001, Kangcheng Liu, Lihua Xie 0001
ICRA1
2024 An Image Acquisition Scheme for Visual Odometry based on Image Bracketing and Online Attribute Control
abstract
Visual odometry (VO) system is challenged by complex illumination environments. Image quality and its consistency in the time domain directly determine feature detection and tracking performance, which further affect the robustness and accuracy of the entire system. In this paper, an image acquisition scheme with image bracketing patterns is proposed. Images with different exposure levels are continuously captured to sufficiently explore the scene under varying illumination. An attribute control method is designed to adjust image exposures within the brackets online. Gaussian process regression fits the relationship between image quality metric and exposure via image synthesis technique. The optimal exposures for the next bracket are obtained directly without attempts to ensure a quick response. Experiments show our acquisition system’s effectiveness and performance improvement for VO tasks in complex illumination scenes.
Jinhao He, Bohuan Xue, Jin Wu 0002, Pengyu Yin, Jianhao Jiao, Ming Liu 0001
ICRA5
2024 Multi-Robot Active Graph Exploration with Reduced Pose-SLAM Uncertainty via Submodular Optimization
abstract
This paper considers the multi-robot active graph exploration problem, where robots need to collaboratively cover a graph environment while maintaining reliable pose estimation in collaborative Simultaneous Localization and Mapping (SLAM). Considering both objectives presents challenges for multi-robot pathfinding, as it involves the expensive covariance propagation for SLAM uncertainty evaluation, especially when considering various combinations of robots’ paths. To reduce the computational complexity, we propose an efficient two-stage strategy where exploration paths are first generated for quick coverage, and then enhanced by adding informative loop-closing actions along the paths for reliable pose estimation. We formulate the latter problem as a non-monotone submodular maximization problem by relating SLAM uncertainty with pose graph topology, which (1) facilitates a more efficient evaluation of SLAM uncertainty than covariance inference, and (2) allows the employment of approximation algorithms in submodular optimization to provide suboptimality guarantees. We further introduce ordering heuristics to improve the objective values while preserving the optimality bound. Simulation experiments over randomly generated graph environments verify the effectiveness of our methods to achieve quick coverage and enhanced pose graph reliability, and benchmark the performance of the approximation algorithms and the greedy-based algorithm in the loop edge selection problem. Our implementations will be open-source at https://github.com/bairuofei/CGE.
Ruofei Bai, Shenghai Yuan 0001, Hongliang Guo 0003, Pengyu Yin, Weiyun Yau, Lihua Xie 0001
IROS4
2024 Key Object Detection: Unifying Salient and Camouflaged Object Detection Into One Task
Pengyu Yin, Keren Fu, Qijun Zhao
PRCV (12)1
2023 Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation
abstract
Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmentation. The key to MMCTTA is to adaptively attend to the reliable modality while avoiding catastrophic forgetting during continual domain shifts, which is out of the capability of previous TTA or CTTA methods. To fulfill this gap, we propose an MM-CTTA method called Continual Cross-Modal Adaptive Clustering (CoMAC) that addresses this task from two perspectives. On one hand, we propose an adaptive dual-stage mechanism to generate reliable cross-modal predictions by attending to the reliable modality based on the class-wise feature-centroid distance in the latent space. On the other hand, to perform test-time adaptation without catastrophic forgetting, we design class-wise momentum queues that capture confident target features for adaptation while stochastically restoring pseudo-source features to revisit source knowledge. We further introduce two new benchmarks to facilitate the exploration of MM-CTTA in the future. Our experimental results show that our method achieves state-of-the-art performance on both benchmarks. Visit our project website at https://sites.google.com/view/mmcotta.
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001
ICCV4
2023 Segregator: Global Point Cloud Registration with Semantic and Geometric Cues
abstract
This paper presents Segregator, a global point cloud registration framework that exploits both semantic information and geometric distribution to efficiently build up outlier-robust correspondences and search for inliers. Current state-of-the-art algorithms rely on point features to set up putative correspondences and refine them by employing pair-wise distance consistency checks. However, such a scheme suffers from degenerate cases, where the descriptive capability of local point features downgrades, and unconstrained cases, where length-preserving (1-TRIMs)-based checks cannot sufficiently constrain whether the current observation is consistent with others, resulting in a complexified NP-complete problem to solve. To tackle these problems, on the one hand, we propose a novel degeneracy-robust and efficient corresponding procedure consisting of both instance-level semantic clusters and geometric-level point features. On the other hand, Gaussian distribution-based translation and rotation invariant measurements (G-TRIMs) are proposed to conduct the consistency check and further constrain the problem size. We validated our proposed algorithm on extensive real-world data-based experiments. The code is available: https://github.com/Pamphlett/Segregator.
Pengyu Yin, Shenghai Yuan 0001, Haozhi Cao, Xingyu Ji, Lihua Xie 0001
ICRA1
2023 PRFNet: Progressive Region Focusing Network for Polyp Segmentation
Jilong Chen, Junlong Cheng, Pengyu Yin, Guoan Wang
PRCV (5)4
2020 CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence
abstract
In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective function through bidirectional correspondence. Globally, error metric of correntropy is introduced as noise model to resist outliers. Importantly, the close resemblance between normal-distributions transform (NDT) and correntropy is revealed. To ease the minimization step, an on-manifold parameterization of the special Euclidean group is proposed. Extensive experiments validate that CoBigICP outperforms several well-known and state-of-the-art methods.
Pengyu Yin, Di Wang 0028, Shaoyi Du, Shihui Ying, Yue Gao 0002, Nanning Zheng 0001
IROS1