EDBT 2026 Demo / reviewers in the wild / expert
Yufeng Yue
dblp:194/9143
· DBLP profile ↗
66ranked-venue papers
9as first author
53since 2021 · last 2026
0000-0001-6628-7946ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 7 first-author · 40 since 2021Systems, architecture and hardware · 34 · 5 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaptRGB-t: Adaptive RGB-t semantic segmentation via efficient parameter-tuning with textual guidance
Yufeng Yue, Yi Yang 0009, Mengyin Fu |
Neurocomputing | 2 |
| 2026 | Conditional diffusion model for infrared and visible image fusion in open environments with few denoising steps
Luojie Yang, Chunming Li, Guangyan Chen, Yufeng Yue |
Signal Process. | 5 |
| 2026 | DiCriTest: Testing Scenario Generation for Decision-Making Agents Considering Diversity and Criticality
Qitong Chu, Yufeng Yue, Danya Yao, Huaxin Pei |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Coupling Structural Descriptors With a Novel Semantic Graph Matching Approach for LiDAR Loop DetectionabstractOutdoor loop closure detection is essential for correcting odometry drift and constructing a globally consistent map. Semantic-graph-based approaches effectively model object-level topology and achieve strong loop closure performance; however, their effectiveness degrades in background-dominated scenes with few distinctive objects, and establishing accurate injective node correspondences remains challenging. In contrast, structural descriptor methods, though offering stronger environmental generality through spatial-distribution modeling, remain susceptible to LiDAR noise and the discriminative power of point-level features. These limitations motivate the need for a more robust method that combines adaptability with enhanced descriptive power. We propose a novel loop-closure detection framework, SAGE, that integrates highly adaptable point-cloud shape-distribution features and generally reliable semantic graph topology, adaptively combining their similarity measures to improve detection performance. Specifically, we design a semantic graph matching module with dual constraints, local graph feature consistency and global spatial consistency, to achieve more accurate injective node correspondences. In addition, we extract point-cloud shape-distribution features and introduce a fusion mechanism that integrates them with the semantic graph module, assessing reliability and adaptively weighting their contributions. Extensive loop closure detection and pose estimation experiments on various datasets demonstrate that SAGE achieves superior performance over strong baselines. We provide the code at https://github.com/SAGE-11/SAGE. Meiling Wang 0002, Sibo Zuo, Chengxi Yang, Jinhao Jiang, Xieyuanli Chen, Yufeng Yue |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2026 | ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching With Collaborative LLM-Agents
Hao Liang 0016, Hao Li 0075, Jiyuan Guo, Yufeng Yue, Mengyin Fu, Yi Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic ManipulationabstractLearning from demonstration is a powerful method for robotic skill acquisition. Nevertheless, a critical limitation lies in the substantial costs associated with gathering demonstration datasets, typically action-labeled robot data, which creates a fundamental constraint in the field. Video data offer a compelling solution as an alternative rich data source, containing diverse behavioral and physical knowledge. This study introduces G3M, an innovative framework that exploits video data viaGraph-to-GraphsGenerativeModeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. The proposed G3M abstracts video frame into graph representations by identifying object and visual action vertices for capturing state information. It then effectively models internal structures and spatial relationships present in these graph constructions, with the objective of predicting forthcoming graphs. The generated graphs function as conditional inputs that guide the control policy in determining robotic behaviors. This concise method effectively encodes critical spatial relationships while facilitating accurate prediction of subsequent graph sequences, thus allowing the development of resilient control policy despite constraints in action-annotated training samples. Furthermore, these transferable graph representations enable the effective extraction of manipulation knowledge through human videos as well as recordings from robots with different embodiments. The experimental results demonstrate that G3M attains superior performance using merely 20% action-labeled data relative to comparable approaches. Moreover, our method outperforms the state-of-the-art method, showing performance gains exceeding 19% in simulated environments and 23% in real-world experiments, while delivering improvements of over 35% in cross-embodiment transfer experiments and exhibiting strong performance on long-horizon tasks. Our project page is available athttps://g3m-project.github.io/. Guangyan Chen, Meiling Wang 0002, Te Cui, Chengcai Yang, Mengxiao Hu, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
IEEE Trans. Robotics | 11 |
| 2025 | GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningabstractLearning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alternative. In this paper, we present GraphMimic, a novel paradigm that leverages video data via graph-to-graphs generative modeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. Specifically, GraphMimic abstracts video frames into object and visual action vertices, and constructs graphs for state representations. The graph generative modeling network then effectively models internal structures and spatial relationships within the constructed graphs, aiming to generate future graphs. The generated graphs serve as conditions for the control policy, mapping to robot actions. Our concise approach captures important spatial relations and enhances future graph generation accuracy, enabling the acquisition of robust policies from limited action-labeled data. Furthermore, the transferable graph representations facilitate the effective learning of manipulation skills from cross-embodiment videos. Our experiments exhibit that GraphMimic achieves superior performance using merely 20% action-labeled data. Moreover, our method outperforms the state-of-the-art method by over 17% and 23% in simulation and real-world experiments, and delivers improvements of over 33% in cross-embodiment transfer experiments. Guangyan Chen, Te Cui, Meiling Wang 0002, Chengcai Yang, Mengxiao Hu, Yao Mu 0001, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
CVPR | 12 |
| 2025 | DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone ImageryabstractDrones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rendering quality, providing a new avenue for 3D reconstruction from drone imagery. However, dynamic distractors in wild environments challenge the static scene assumption in radiance fields, while limited view constraints hinder the accurate capture of underlying scene geometry. To address these challenges, we introduce DroneSplat, a novel framework designed for robust 3D reconstruction from in-the-wild drone imagery. Our method adaptively adjusts masking thresholds by integrating local-global segmentation heuristics with statistical approaches, enabling precise identification and elimination of dynamic distractors in static scenes. We enhance 3D Gaussian Splatting with multi-view stereo predictions and a voxel-guided optimization strategy, supporting high-quality rendering under limited view constraints. For comprehensive evaluation, we provide a drone-captured 3D reconstruction dataset encompassing both dynamic and static scenes. Extensive experiments demonstrate that DroneSplat outperforms both 3DGS and NeRF baselines in handling in-the-wild drone imagery. Project page: https://bityia.github.io/DroneSplat/. Jiadong Tang, Yu Gao 0040, Dianyi Yang, Liqi Yan, Yufeng Yue, Yi Yang 0009 |
CVPR | 5 |
| 2025 | High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic GraspingabstractIn various robotic applications, understanding accurate object poses for robots is essential for high-precision tasks such as factory assembly or daily insertions. Tactile sensing, which compensates for visual information, offers rich texture-based or force-based data for object pose estimation. However, previous methods for pose estimation typically over-look dynamic situations, such as slippage of grasped objects or movement of contacted objects during interactions with the environment, thus increasing the complexity of pose estimation. To address these challenges, we propose an efficient method that utilizes visual and tactile sensing to estimate object poses through particle filtering. We leverage visual information to track the pose of the contacted object in real-time and estimate the pose changes of the grasped object using displacement data obtained from tactile sensors. Our experimental evaluation on 13 objects with diverse geometric shapes demonstrated the ability to estimate high-precision poses, which revealed the robot's powerful ability to cope with dynamic scenes for compelled motion of objects, proving our framework's adaptability in practical scenarios with uncertainty. Zicai Peng, Te Cui, Guangyan Chen, Yi Yang 0009, Yufeng Yue |
ICRA | 6 |
| 2025 | OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene UnderstandingabstractRecent advancements in 3D Gaussian Splatting have significantly improved the efficiency and quality of dense semantic SLAM. However, previous methods are generally constrained by limited-category pre-trained classifiers and implicit semantic representation, which hinder their performance in open-set scenarios and restrict 3D object-level scene understanding. To address these issues, we propose OpenGS-SLAM, an innovative framework that utilizes 3D Gaussian representation to perform dense semantic SLAM in open-set environments. Our system integrates explicit semantic labels derived from 2D foundational models into the 3D Gaussian framework, facilitating robust 3D object-level scene understanding. We introduce Gaussian Voting Splatting to enable fast 2D label map rendering and scene updating. Additionally, we propose a Confidence-based 2D Label Consensus method to ensure consistent labeling across multiple views. Furthermore, we employ a Segmentation Counter Pruning strategy to improve the accuracy of semantic scene representation. Extensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of our method in scene understanding, tracking, and mapping, achieving 10× faster semantic rendering and 2× lower storage costs compared to existing methods. Project page: https://young-bit.github.io/opengs-github.github.io/. Dianyi Yang, Yu Gao 0040, Xihan Wang, Yufeng Yue, Yi Yang 0009, Mengyin Fu |
ICRA | 4 |
| 2025 | CDMFusion: RGB-T Image Fusion Based on Conditional Diffusion Models via Few Denoising Steps in Open Environments
Luojie Yang, Lijin Fang, Yi Yang 0009, Yufeng Yue |
ICRA | 5 |
| 2025 | Open-RGBT: Open-Vocabulary RGB-T Zero-Shot Semantic Segmentation in Open-World EnvironmentsabstractSemantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenarios due to their reliance on pretrained models and predefined categories. Recent advancements in Visual Language Models (VLMs) have facilitated a shift from closedset to open-vocabulary semantic segmentation methods. However, these models face challenges in dealing with intricate scenes, primarily due to the heterogeneity between RGB and thermal modalities. To address this gap, we present Open-RGBT, a novel open-vocabulary RGB-T semantic segmentation model. Specifically, we obtain instance-level detection proposals by incorporating visual prompts to enhance category understanding. Additionally, we employ the CLIP model to assess image-text similarity, which helps correct semantic consistency and mitigates ambiguities in category identification. Empirical evaluations demonstrate that Open-RGBT achieves superior performance in diverse and challenging real-world scenarios, even in the wild, significantly advancing the field of RGB-T semantic segmentation. The project page of Open-RGBT is available at https://OpenRGBT.github.io/. Yufeng Yue, Luojie Yang, Xunjie He, Yi Yang 0009, Mengyin Fu |
ICRA | 2 |
| 2025 | ORA-NET: Enhancing Image Feature Matching through Oriented Overlapping Region AlignmentabstractImage feature matching is a fundamental task in computer vision. Existing local feature matching methods can establish robust correspondences between image pairs. However, these methods heavily rely on dense local image features, making them susceptible to significant perspective differences, characterized by rotation and scale changes. To alleviate this limitation, we introduce a novel oriented Overlapping Region Alignment method, named ORA-NET, which presents a concise and efficient approach to enhance the performance of image feature matching methods. We introduce the Multidirectional Cross-scale Feature Aggregation module to aggregate rotation-equivariant features across multiple scales and model long-range dependencies. Additionally, the Oriented Overlap Alignment module estimates scale and rotation differences within overlapping regions using a coarse-to-fine rotation correction approach. Importantly, our method serves as a plug-and-play module that can be seamlessly integrated into other correspondence matching pipelines. Experimental results demonstrate that ORA-NET significantly enhances the matching performance of existing local feature matching methods, particularly in scenarios involving substantial perspective differences. Te Cui, Meiling Wang 0002, Guangyan Chen, Yufeng Yue |
IROS | 5 |
| 2025 | Human Demonstrations are Generalizable Knowledge for RobotsabstractLearning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object instances. In this paper, we propose a different perspective, considering human demonstration videos not as mere instructions, but as a source of knowledge for robots. Motivated by this perspective and the remarkable comprehension and generalization capabilities exhibited by large language models (LLMs), we propose DigKnow, a method that DIstills Generalizable KNOWledge with a hierarchical structure. Specifically, DigKnow begins by converting human demonstration video frames into observation knowledge. This knowledge is then subjected to analysis to extract human action knowledge and further distilled into pattern knowledge that comprises task and object instances, resulting in the acquisition of generalizable knowledge with a hierarchical structure. In settings with different tasks or object instances, DigKnow retrieves relevant knowledge for the current task and object instances. Subsequently, the LLM-based planner conducts planning based on the retrieved knowledge, and the policy executes actions in line with the plan to achieve the designated task. Utilizing the retrieved knowledge, we validate and rectify planning and execution outcomes, resulting in a substantial enhancement of the success rate. Experimental results across a range of tasks and scenes demonstrate the effectiveness of this approach in facilitating real-world robots to accomplish tasks with the knowledge derived from human demonstrations. Te Cui, Tianxing Zhou, Mengxiao Hu, Zicai Peng, Haizhou Li 0004, Guangyan Chen, Meiling Wang 0002, Yufeng Yue |
IROS | 9 |
| 2025 | OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel RepresentationabstractIn recent years, vision-language models (VLMs) have advanced open-vocabulary mapping, enabling mobile robots to simultaneously achieve environmental reconstruction and high-level semantic understanding. While integrated object cognition helps mitigate semantic ambiguity in point-wise feature maps, efficiently obtaining rich semantic understanding and robust incremental reconstruction at the instance-level remains challenging. To address these challenges, we introduce OpenVox, a real-time incremental open-vocabulary probabilistic instance voxel representation. In the front-end, we design an efficient instance segmentation and comprehension pipeline that enhances language reasoning through encoding captions. In the back-end, we implement probabilistic instance voxels and formulate the cross-frame incremental fusion process into two subtasks: instance association and live map evolution, ensuring robustness to sensor and segmentation noise. Extensive evaluations across multiple datasets demonstrate that OpenVox achieves state-of-the-art performance in zero-shot instance segmentation, semantic segmentation, and open-vocabulary retrieval. The project page of OpenVox is available at https://open-vox.github.io/. Yinan Deng, Bicheng Yao, Yihang Tang, Tianxing Zhou, Yi Yang 0009, Yufeng Yue |
IROS | 6 |
| 2025 | PartGrasp: Generalizable Part-level Grasping via Semantic-Geometric AlignmentabstractThe ability to perform generalizable and precise grasping on functional object parts is a prerequisite for robotic manipulation in open environments. Recent foundation models have demonstrated promising semantic correspondence capabilities in guiding robots to grasp similar parts across objects with resembling shapes and poses. However, existing works struggle to generalize precise grasp poses when the target objects exhibit substantial geometric and positional variations. To tackle this challenge, we present PartGrasp, a method that achieves precise part grasping through hierarchical integration of highly generalizable semantic correspondence and precise geometric registration. Specifically, we first build a grasp knowledge bank by extracting grasp poses and object meshes from demonstrations. Upon retrieving a reference from this bank, we initially perform a coarse alignment using semantic correspondence, followed by a fine registration that adapts to geometric variations. This approach achieves fine-grained generalization of part grasping that is robust to both shape and pose variations. Extensive experiments demonstrate the efficacy of our method in terms of both generalization capability and accuracy. Videos and more details are available on our project site: https://part-grasp.github.io/partgrasp/. Chengcai Yang, Guangyan Chen, Yufeng Yue |
IROS | 4 |
| 2025 | OpenObject-NAV: Open-Vocabulary Object-Oriented Navigation Based on Dynamic Carrier-Relationship Scene GraphabstractIn everyday life, frequently used objects like cups often have unfixed positions and multiple instances within the same category, and their carriers frequently change as well. As a result, it becomes challenging for a robot to efficiently navigate to a specific instance. To tackle this challenge, the robot must capture and update scene changes and plans continuously. However, current object navigation approaches primarily focus on the semantic level and lack the ability to dynamically update scene representation. To address these limitations, this paper captures the relationships between frequently used objects and their static carriers. Specifically, it constructs an open-vocabulary Carrier-Relationship Scene Graph (CRSG) and updates the carrying status during robot navigation to reflect the dynamic changes of the scene. Based on the CRSG, we further propose an instance navigation strategy that models the navigation process as a Markov Decision Process. At each step, decisions are informed by Large Language Model’s commonsense knowledge and visual-language feature similarity. We designed a series of long-sequence navigation tasks for frequently used everyday items in the Habitat simulator. The results demonstrate that by updating the CRSG, the robot can efficiently navigate to moved targets. Additionally, we deployed our algorithm on a real robot and validated its practical effectiveness. The project page can be found here: https://OpenObject-Nav.github.io. Meiling Wang 0002, Yinan Deng, Zibo Zheng, Jiagui Zhong, Chenjie Zhao, Yufeng Yue |
IROS | 8 |
| 2025 | SLOOP: Aligned Coordinate System-aided LiDAR LOOP Closure Detection based on Semantic Node Graph MatchingabstractLoop closure detection and pose estimation play a significant role in correcting odometry trajectories and generating globally consistent point cloud maps. Geometric feature descriptor methods neglect object-level spatial topology features, resulting in inadequate performance in loop closure detection. Semantic graph-based loop closing methods improve upon this, however, they still follow the paradigm of "first generating descriptors, then comparing similarity, and finally achieving alignment (6D pose)". Specifically, they compare two semantic graphs that are not spatially aligned, which makes direct node correspondences impossible and necessitates extensive descriptor extraction and comparison. This decouples similarity comparison from 6D pose estimation, resulting in a cumbersome process that limits practicality and scalability. This paper proposes SLOOP, a novel descriptor-free semantic graph matching method that "aligns two graphs first, followed by efficient similarity comparison". Specifically, we first design a dedicated neighborhood semantic feature module to extract high-quality matched node pairs. Next, we seek the aligned coordinate systems for candidate loops based on the robust ground normal vectors and two suitable node pairs examined by the two-stage global geometric consistency metrics. Finally, the aligned coordinate systems enable efficient extraction and comparison of node spatial distributions. We conducted extensive outdoor loop detection experiments and compared with various loop closure detection approaches, demonstrating the improved performance of SLOOP in loop closure detection and its practicality. The code and related materials are available at https://github.com/bit-tyj/sloop_c. Meiling Wang 0002, Jiagui Zhong, Sibo Zuo, Yinan Deng, Yufeng Yue |
IROS | 7 |
| 2025 | GaussianGraph: 3D Gaussian-Based Scene Graph Generation for Open-World Scene UnderstandingabstractRecent advancements in 3D Gaussian Splatting(3DGS) have significantly improved semantic scene understanding, enabling natural language queries to localize objects within a scene. However, existing methods primarily focus on embedding compressed CLIP features to 3D Gaussians, suffering from low object segmentation accuracy and lack spatial reasoning capabilities. To address these limitations, we propose GaussianGraph, a novel framework that enhances 3DGS-based scene understanding by integrating adaptive semantic clustering and scene graph generation. We introduce a ‘Control-Follow’ clustering strategy, which dynamically adapts to scene scale and feature distribution, avoiding feature compression and significantly improving segmentation accuracy. Additionally, we enrich scene representation by integrating object attributes and spatial relations extracted from 2D foundation models. To address inaccuracies in spatial relationships, we propose 3D correction modules that filter implausible relations through spatial consistency verification, ensuring reliable scene graph construction. Extensive experiments on three datasets demonstrate that GaussianGraph outperforms state-of-the-art methods in both semantic segmentation and object grounding tasks, providing a robust solution for complex scene understanding and interaction. We provide supplementary video and code at https://wangxihan-bit.github.io/GaussianGraph. Xihan Wang, Dianyi Yang, Yu Gao 0040, Yufeng Yue, Yi Yang 0009, Mengyin Fu |
IROS | 4 |
| 2025 | OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level UnderstandingabstractRecent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by rigid offline pipelines and the inability to provide precise 3D object-level understanding given open-ended queries. In this paper, we present OpenGS-Fusion, an innovative open-vocabulary dense mapping framework that improves semantic modeling and refines object-level understanding. OpenGS-Fusion combines 3D Gaussian representation with a Truncated Signed Distance Field to facilitate lossless fusion of semantic features on-the-fly. Furthermore, we introduce a novel multimodal language-guided approach named MLLM-Assisted Adaptive Thresholding, which refines the segmentation of 3D objects by adaptively adjusting similarity thresholds, achieving an improvement 17% in 3D mIoU compared to the fixed threshold strategy. Extensive experiments demonstrate that our method outperforms existing methods in 3D object understanding and scene reconstruction quality, as well as showcasing its effectiveness in language-guided scene interaction. The code is available at https://young-bit.github.io/opengs-fusion.github.io/. Dianyi Yang, Xihan Wang, Yu Gao 0040, Shiyang Liu, Bohan Ren, Yufeng Yue, Yi Yang 0009 |
IROS | 6 |
| 2025 | TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation ModelsabstractReliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when categories share similar visual characteristics. 2) While SAM excels in instance-level segmentation, integrating it with thermal images and text is hindered by modality heterogeneity and computational inefficiency. To address these, we propose TASeg, a text-aware RGB-T segmentation framework by using Low-Rank Adaptation (LoRA) fine-tuning technology to adapt vision foundation models. Specifically, we propose a Dynamic Feature Fusion Module (DFFM) in the image encoder, which effectively merges features from multiple visual modalities while freezing SAM’s original transformer blocks. Additionally, we incorporate CLIP-generated text embeddings in the mask decoder to enable semantic alignment, which further rectifies the classification error and improves the semantic understanding accuracy. Experimental results across diverse datasets demonstrate that our method achieves superior performance in challenging scenarios with fewer trainable parameters. Te Cui, Qitong Chu, Wenjie Song 0001, Yi Yang 0009, Yufeng Yue |
IROS | 6 |
| 2025 | OpenMIGS: Multi-granularity Information-preserving Open-Vocabulary 3D Gaussian SplattingabstractOpen-vocabulary scene understanding is critical for robotics, yet existing 3D Gaussian Splatting (3DGS) methods rely on compressed feature embeddings, compromising semantic fidelity and fine-grained interpretation. Although utilizing uncompressed high-dimensional features offers a potential solution, their direct integration imposes prohibitive memory and computational costs. To address this challenge, we propose OpenMIGS, a novel 3DGS-based framework for multi-granularity, information-preserving open-vocabulary understanding across both object and part levels. Specifically, OpenMIGS first constructs object-level Gaussian fields as structured carriers where a two-stage clustering strategy ensures global consistency in object labeling, and a code-book subsequently associates these object label with their uncompressed high-dimensional features. Building on this, a lightweight implicit field processes the geometric coordinates of object Gaussians to regress part-level high-dimensional features, enabling multi-granularity understanding. Experimental results on multiple datasets show that OpenMIGS outperforms existing methods in open-vocabulary understanding and retrieval tasks. It also supports multi-granularity scene editing for flexible semantic manipulation. The code is available at https://github.com/jingyuzhao1010/OpenMIGS. Yinan Deng, Yufeng Yue |
IROS | 4 |
| 2025 | STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task plannerabstractThe ability to perform reliable long-horizon task planning is crucial for deploying robots in real-world environments. However, directly employing Large Language Models (LLMs) as action sequence generators often results in low success rates due to their limited reasoning ability for long-horizon embodied tasks. In the STEP framework, we construct a subgoal tree through a pair of closed-loop models: a subgoal decomposition model and a leaf node termination model. Within this framework, we develop a hierarchical tree structure that spans from coarse to fine resolutions. The subgoal decomposition model leverages a foundation LLM to break down complex goals into manageable subgoals, thereby spanning the subgoal tree. The leaf node termination model provides real-time feedback based on environmental states, determining when to terminate the tree spanning and ensuring each leaf node can be directly converted into a primitive action. Experiments conducted in both the VirtualHome WAH-NL benchmark and on real robots demonstrate that STEP achieves long-horizon embodied task completion with success rates up to 34% (WAH-NL) and 25% (real robot) outperforming SOTA methods. Tianxing Zhou, Haojia Ao, Guangyan Chen, Boyang Xing, Cheng Jingwen, Yi Yang 0009, Yufeng Yue |
IROS | 8 |
| 2025 | Unifying Latent Action and Latent State Pre-training for Policy Learning from VideosabstractVideo data provides an accessible and rich source beyond expensive action-labeled robot data for advancing robotic learning paradigms. Motivated by this potential, researchers investigate methods to exploit video data in robotic learning. Recent approaches can be primarily divided into two categories: Action-based approaches tokenize latent actions from videos for policy pre-training. State-based approaches pre-train models to predict subsequent states. The former establishes rich motion priors, while the latter empowers the robot to anticipate future events. These complementary capabilities suggest significant potential for integration into a unified framework. In this paper, we propose UniMimic, a novel approach unifying latent action and latent state pre-training from videos. We first train a unified tokenizer to learn latent states from video frames while deriving latent actions between state tokens. Subsequently, the policy is pre-trained on videos to predict these latent actions and subsequent latent states. Finally, the policy is fine-tuned on an action-labeled robot dataset to transfer the learned priors to precise robot execution. Experiments exhibit that our pre-training stage enhances the performance by 19% in the Libero benchmark and improves the average number of tasks completed in a row of 5 from 2.50 and 2.35 to 3.89 and 3.73 in the CALVIN benchmark. In the real-world experiments, our method still delivers improvements exceeding 36%. Guangyan Chen, Meiling Wang 0002, Te Cui, Luojie Yang, Lin Zhao 0016, Yi Yang 0009, Yufeng Yue |
SIGGRAPH Asia | 10 |
| 2025 | Joint Conditional Diffusion Model for image restoration with mixed degradations
Yufeng Yue, Luojie Yang |
Neurocomputing | 1 |
| 2025 | SeGraM: Aligned Coordinate System Aided Semantic Graph Matching Method for Loop Closure DetectionabstractCapturing object semantics and their spatial relationships is crucial to estimating scene similarity for loop closure detection. Existing semantic loop closure detection methods generally treat semantics as landmarks or extract the object topology to compare the similarity of frames. However, they often neglect the absolute spatial distribution of objects, which is essential to capture distinctive features of the scene. A fundamental requirement is to register the spatial coordinates of both frames in a unified reference frame. To address this, we construct aligned coordinate systems between two frames and extract absolute spatial distribution features of objects for loop closure detection. Building on this, we introduce SeGraM, a unified semantic graph matching approach applicable to both indoor and outdoor environments. Specifically, for each pair of semantic graphs, we first establish correspondences between nodes, referred to as node pairs. We then evaluate the geometric and semantic consistency of these pairs, along with the local graph features in the surrounding. To facilitate meaningful comparisons, two node pairs are carefully selected to establish aligned spherical coordinate systems, with ground normals to define the Z axes outdoors. SeGraM is validated in both indoor and outdoor scenarios and is compared with multiple algorithms, demonstrating improvements in loop closure detection accuracy. The code is accessible at https://github.com/BIT-TYJ/SeGraM. Meiling Wang 0002, Yinan Deng, Sibo Zuo, Yufeng Yue |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Point Tree Transformer for Point Cloud RegistrationabstractPoint cloud registration is a fundamental task in the fields of computer vision and robotics. Recent advancements in transformer-based methods have demonstrated enhanced performance in this domain. However, the standard attention mechanisms employed in these approaches tend to incorporate numerous points of low relevance, and therefore struggle to focus their attention weights on sparse yet meaningful points. This inefficiency leads to limited local structure modeling capabilities and quadratic computational complexity. To overcome these limitations, we propose the Point Tree Transformer (PTT), a novel transformer-based approach for point cloud registration that efficiently extracts comprehensive local and global features while maintaining linear computational complexity. The PTT constructs hierarchical feature trees from point clouds in a coarse-to-dense manner, and introduces a novel Point Tree Attention (PTA) mechanism. This mechanism adheres to the tree structure to facilitate the progressive convergence of attended regions toward salient points. Specifically, each tree layer selectively identifies a subset of relevant points with the highest attention scores, and subsequent layers focus attention on areas of significant relevance, derived from the child points of the selected point set. The feature extraction process additionally incorporates coarse point features that capture high-level semantic information, thus facilitating local structure modeling and the progressive integration of multiscale information. Consequently, the PTA enables the model to focus on essential local structures and extract intricate local information while maintaining linear computational complexity. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate that our method outperforms state-of-the-art methods in terms of performance. The code for our method is publicly available at https://github.com/CGuangyan-BIT/PTT. Meiling Wang 0002, Guangyan Chen, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | CoSTFE: Spatio-Temporal Feature Enhancement for Collaborative PerceptionabstractCollaborative perception enables a more comprehensive and precise representation of the environment, owing to the complementary information shared among different agents. However, spatio-temporal disturbances, including localization errors (spatial) and time delays (temporal), are prevalent in practical applications and significantly impair detection performance. To improve both the accuracy and robustness, a novel framework called Spatio-Temporal Feature Enhancement for Collaborative Perception (CoSTFE) is proposed. Specifically, we present a Histogram-based Spatial Correction (HSC) module to optimize the transformation matrix and promote the robustness when localization errors happen. In addition, the Deformable Temporal Augmentation (DTA) module is introduced to predict and enhance the current characteristic with long-term historical dynamics. Compared with existing methods on three publicly available collaborative perception datasets, our approach exhibits superior performance and robustness in the collaborative 3D object detection task. Meiling Wang 0002, Xunjie He, Yufeng Yue |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | OmniMap: A General Mapping Framework Integrating Optics, Geometry, and SemanticsabstractRobotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of these requirements while exhibiting optical blurring, geometric irregularities, and semantic ambiguities. To address these challenges, we propose OmniMap. Overall, OmniMap represents the first online mapping framework that simultaneously captures optical, geometric, and semantic scene attributes while maintaining real-time performance and model compactness. At the architectural level, OmniMap employs a tightly coupled 3DGS-Voxel hybrid representation that combines fine-grained modeling with structural stability. At the implementation level, OmniMap identifies key challenges across different modalities and introduces several innovations: adaptive camera modeling for motion blur and exposure compensation, hybrid incremental representation with normal constraints, and probabilistic fusion for robust instance-level understanding. Extensive experiments show OmniMap's superior performance in rendering fidelity, geometric accuracy, and zero-shot semantic segmentation compared to state-of-the-art methods across diverse scenes. The framework's versatility is further evidenced through a variety of downstream applications, including multi-domain scene Q&A, interactive editing, perception-guided manipulation, and map-assisted navigation. Yinan Deng, Yufeng Yue, Jianyu Dou, Yi Yang 0009, Mengyin Fu |
IEEE Trans. Robotics | 2 |
| 2025 | MC-NeRF: Multi-Camera Neural Radiance Fields for Multi-Camera Image Acquisition SystemsabstractNeural Radiance Fields (NeRF) use multi-view images for 3D scene representation, demonstrating remarkable performance. As one of the primary sources of multi-view images, multi-camera systems encounter challenges such as varying intrinsic parameters and frequent pose changes. Most previous NeRF-based methods assume a unique camera and rarely consider multi-camera scenarios. Besides, some NeRF methods that can optimize intrinsic and extrinsic parameters still remain susceptible to suboptimal solutions when these parameters are poor initialized. In this paper, we propose MC-NeRF, a method for joint optimization of both intrinsic and extrinsic parameters alongside NeRF, allowing individual camera parameters for each image. First, we analyze the coupling issue that arises from the joint optimization between intrinsics and extrinsics, and propose a decoupling constraint utilizing auxiliary images. To further address the degenerate cases in the decoupling process, we introduce an efficient auxiliary image acquisition scheme to mitigate these effects. Furthermore, recognizing that most existing datasets are designed for a unique camera, we provided a new dataset that includes both simulated data and real-world data. Experiments demonstrate the effectiveness of our method in scenarios where each image corresponds to different camera parameters. Specifically, our approach outperforms the baselines favorably in terms of intrinsics estimation, extrinsics estimation, scale estimation, and rendering quality. Yu Gao 0040, Lutong Su, Hao Liang 0016, Yufeng Yue, Yi Yang 0009, Mengyin Fu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | MM4MM: Map Matching Framework for Multi-Session Mapping in Ambiguous and Perceptually-Degraded EnvironmentsabstractMulti-session mapping serves as the pre-requisite for autonomous robots to fulfill various long-term tasks (e.g., map updating, navigation, collaboration). However, it is challenging to implement multi-session mapping in enclosed or partially enclosed ambiguous environments (e.g., long corridors, industrial warehouses). Existing solutions either depend heavily on the matching of elementary geometric features (e.g., points, lines, and planes), which tends to fail in environments with ambiguous geometric features; or depend on the given guess of the initial transformation matrix of multiple single-session maps, which is not always obtainable and accurate enough. The ambient magnetic field has exhibited ubiquity and high distinctiveness at different location, which makes it suitable for estimating the initial transformation matrix. Thus, this paper proposes a novel probabilistic magnetic-aware Map Matching framework for Multi-session Mapping, namely MM4MM, to estimate the relative transformation of multiple single-session maps and to build the globally consistent maps in ambiguous and perceptually-degraded environments. The key novelties of this work are the designing of the hierarchical probabilistic map matching framework and the Particle Swarm Optimization strategy to associate the magnetic data of multiple sessions. Evaluations on both simulated and real world experiments demonstrate the greatly improved utility, accuracy, and robustness of multi-session mapping over the comparative methods. Zhenyu Wu 0001, Yufeng Yue, Jun Zhang 0042, Hongming Shen, Danwei Wang |
ICRA | 4 |
| 2024 | Fast and Robust Point Cloud Registration with Tree-based TransformerabstractPoint cloud registration is essential in computer vision and robotics. Recently, transformer-based methods have achieved advanced point cloud registration performance. However, the standard attention mechanism utilized in these methods considers many low-relevance points, and it has difficulty focusing its attention weights on sparse and meaningful points, leading to limited local structure modeling capabilities and quadratic computational complexity. To address these limitations, we present the Tree-based Transformer (TrT), which is able to extract abundant local and global features with linear computational complexity. Specifically, the TrT builds coarse-to-dense feature trees, and a novel Tree-based Attention (TrA) is proposed to guide the progressive convergence of the attended regions toward meaningful points and to structurize point clouds following tree structures. In each layer, the top ${\mathcal{S}}$ key points with the highest attention scores are selected, such that in the next layer, attention is evaluated only within the specified high-relevance regions, corresponding to the child points of these selected ${\mathcal{S}}$ points. Additionally, coarse features containing high-level semantic information are incorporated into the child points to guide the feature extraction process, facilitating local structure modeling and multiscale information integration. Consequently, TrA enables the model to focus on critical local structures and extract rich local information with linear computational complexity. Experiments demonstrate that our method achieves state-of-the-art performance on 3DMatch and KITTI benchmarks. The code for our method is publicly available at https://github.com/CGuangyan-BIT/TrT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
ICRA | 5 |
| 2024 | Robust Collaborative Perception against Temporal Information DisturbanceabstractCollaborative perception facilitates a more comprehensive representation of the environment by leveraging complementary information shared among various agents and sensors. However, practical applications often encounter information disturbance which includes perception packet loss and time delays, and a comprehensive framework that can simultaneously address such issues is absent. In addition, the feature extraction process prior to fusion is not sufficient, as it lacks exploration of the local semantics and context dependencies of individual features. To enhance both accuracy and robustness, this paper introduces a novel framework named Robust Collaborative Perception against Temporal Information Disturbance, which predicts perception information when disturbance occurs. Specifically, the Historical Frame Prediction (HFP) module is introduced to make compensation for information loss with temporal association excavation of historical features. Based on the predicted features generated by the HFP module, the Pyramid Attention Integration (PAI) module is introduced to augment local semantics and incorporate global long-range dependencies through multi-scale window attention. Compared with existing methods on the publicly available dataset OPV2V, our approach exhibits superior performance and expanded robustness in the 3D object detection task. The code will be publicly available at https://github.com/hexunjie/Ro-temd. Xunjie He, Te Cui, Yufeng Yue |
ICRA | 6 |
| 2024 | HERO-SLAM: Hybrid Enhanced Robust Optimization of Neural SLAMabstractSimultaneous Localization and Mapping (SLAM) is a fundamental task in robotics, driving numerous applications such as autonomous driving and virtual reality. Recent progress on neural implicit SLAM has shown encouraging and impressive results. However, the robustness of neural SLAM, particularly in challenging or data-limited situations, remains an unresolved issue. This paper presents HERO-SLAM, a Hybrid Enhanced Robust Optimization method for neural SLAM, which combines the benefits of neural implicit field and feature-metric optimization. This hybrid method optimizes a multi-resolution implicit field and enhances robustness in challenging environments with sudden viewpoint changes or sparse data collection. Our comprehensive experimental results on benchmarking datasets validate the effectiveness of our hybrid approach, demonstrating its superior performance over existing implicit field-based methods in challenging scenarios. HERO-SLAM provides a new pathway to enhance the stability, performance, and applicability of neural SLAM in real-world scenarios. Project page: https://hero-slam.github.io. Zhe Xin, Yufeng Yue, Liangjun Zhang, Chenming Wu |
ICRA | 2 |
| 2024 | LCP-Fusion: A Neural Implicit SLAM with Enhanced Local Constraints and Computable PriorabstractRecently the dense Simultaneous Localization and Mapping (SLAM) based on neural implicit representation has shown impressive progress in hole filling and high-fidelity mapping. Nevertheless, existing methods either heavily rely on known scene bounds or suffer inconsistent reconstruction due to drift in potential loop-closure regions, or both, which can be attributed to the inflexible representation and lack of local constraints. In this paper, we present LCP-Fusion, a neural implicit SLAM system with enhanced local constraints and computable prior, which takes the sparse voxel octree structure containing feature grids and SDF priors as hybrid scene representation, enabling the scalability and robustness during mapping and tracking. To enhance the local constraints, we propose a novel sliding window selection strategy based on visual overlap to address the loop-closure, and a practical warping loss to constrain relative poses. Moreover, we estimate SDF priors as coarse initialization for implicit features, which brings additional explicit constraints and robustness, especially when a light but efficient adaptive early ending is adopted. Experiments demonstrate that our method achieve better localization accuracy and reconstruction consistency than existing RGB-D implicit SLAM, especially in challenging real scenes (ScanNet) as well as self-captured scenes with unknown scene bounds. The code is available at https://github.com/laliwang/LCP-Fusion. Yinan Deng, Yi Yang 0009, Yufeng Yue |
IROS | 4 |
| 2024 | VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsabstractVisual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by more than 37% in long-horizon tasks. Code and videos are available on our anonymous homepage. Guangyan Chen, Meiling Wang 0002, Te Cui, Yao Mu 0001, Tianxing Zhou, Zicai Peng, Mengxiao Hu, Haizhou Li 0004, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
NeurIPS | 12 |
| 2024 | VIFNet: An end-to-end visible-infrared fusion network for image dehazing
Te Cui, Yufeng Yue |
Neurocomputing | 4 |
| 2024 | Full Transformer Framework for Robust Point Cloud Registration With Deep Information InteractionabstractPoint cloud registration is an essential technology in computer vision and robotics. Recently, transformer-based methods have achieved advanced performance in point cloud registration by utilizing the advantages of the transformer in order-invariance and modeling dependencies to aggregate information. However, they still suffer from indistinct feature extraction, sensitivity to noise, and outliers, owing to three major limitations: 1) the adoption of CNNs fails to model global relations due to their local receptive fields, resulting in extracted features susceptible to noise; 2) the shallow-wide architecture of transformers and the lack of positional information lead to indistinct feature extraction due to inefficient information interaction; and 3) the insufficient consideration of geometrical compatibility leads to the ambiguous identification of incorrect correspondences. To address the above-mentioned limitations, a novel full transformer network for point cloud registration is proposed, named the deep interaction transformer (DIT), which incorporates: 1) a point cloud structure extractor (PSE) to retrieve structural information and model global relations with the local feature integrator (LFI) and transformer encoders; 2) a deep-narrow point feature transformer (PFT) to facilitate deep information interaction across a pair of point clouds with positional information, such that transformers establish comprehensive associations and directly learn the relative position between points; and 3) a geometric matching-based correspondence confidence evaluation (GMCCE) method to measure spatial consistency and estimate correspondence confidence by the designed triangulated descriptor. Extensive experiments on the ModelNet40, ScanObjectNN, and 3DMatch datasets demonstrate that our method is capable of precisely aligning point clouds, consequently, achieving superior performance compared with state-of-the-art methods. The code is publicly available at https://github.com/CGuangyan-BIT/DIT. Guangyan Chen, Meiling Wang 0002, Qingxiang Zhang, Li Yuan 0007, Yufeng Yue |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Rethinking Point Cloud Registration as Masking and ReconstructionabstractPoint cloud registration is essential in computer vision and robotics. In this paper, a critical observation is made that the invisible parts of each point cloud can be directly utilized as inherent masks, and the aligned point cloud pair can be regarded as the reconstruction target. Motivated by this observation, we rethink the point cloud registration problem as a masking and reconstruction task. To this end, a generic and concise auxiliary training network, the Masked Reconstruction Auxiliary Network (MRA), is proposed. The MRA reconstructs the complete point cloud by separately using the encoded features of each point cloud obtained from the backbone, guiding the contextual features in the backbone to capture fine-grained geometric details and the overall structures of point cloud pairs. Unlike recently developed high-performing methods that incorporate specific encoding methods into transformer models, which sacrifice versatility and introduce significant computational complexity during the inference process, our MRA can be easily inserted into other methods to further improve registration accuracy. Additionally, the MRA is detached after training, thereby avoiding extra computational complexity during the inference process. Building upon the MRA, we present a novel transformer-based method, the Masked Reconstruction Transformer (MRT), which achieves both precise and efficient alignment using standard transformers. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate the superior performance of our MRT over state-of-the-art methods. Codes are available at https://github.com/CGuangyan-BIT/MRA. Guangyan Chen, Meiling Wang 0002, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
ICCV | 5 |
| 2023 | Deep Interactive Full Transformer Framework for Point Cloud RegistrationabstractPoint cloud registration is a crucial technology in the fields of robotics and computer vision. Despite the significant advances in point cloud registration enabled by Transformer-based methods, limitations persist due to indistinct feature extraction, noise sensitivity, and outlier handling. These limitations stem from three factors: (1) the inefficiency of convolutional neural networks (CNNs) to capture global relationships due to their local receptive fields, resulting in extracted features susceptible to noise; (2) the shallow-wide architecture of Transformers, coupled with a lack of positional information, leading to inefficient information interaction and indistinct feature extraction; and (3) the omission of geometrical compatibility leads to ambiguous identification of incorrect correspondences. To overcome these limitations, we propose the Deep Interactive Full Transformer (DIFT) network for point cloud registration, which consists of three key components: (1) a Point Cloud Structure Extractor (PSE) for modeling global relationships and retrieving structural information; (2) a Point Feature Transformer (PFT) for establishing comprehensive associations and directly learning the relative positions between points; and (3) a Geometric Matching-based Correspondence Confidence Evaluation (GMCCE) method for measuring spatial consistency and estimating correspondence confidence. Experimental results on ModelNet40 and 3DMatch datasets demonstrate the superior performance of our proposed method compared to existing state-of-the-art methods. The code for our method is publicly available at https://github.com/CGuangyan-BIT/DIFT. Guangyan Chen, Meiling Wang 0002, Qingxiang Zhang, Li Yuan 0007, Tong Liu 0009, Yufeng Yue |
ICRA | 6 |
| 2023 | Multi-View Robust Collaborative Localization in High Outlier Ratio Scenes Based on Semantic FeaturesabstractFiltering out outlier data associations between local maps can improve the robustness and accuracy of multi-robot localization. When the overlap is low and the field of view difference is large, it is likely to produce outlier data associations between local maps, which will reduce the matching accuracy and even lead to the failure of collaborative localization. To solve this problem, this paper proposes a novel outdoor robust collaborative localization algorithm (HORCL) capable for high outlier ratio scenes. The Mixture Probability Model (MPM) and the Hierarchical EM (Expectation Maximization) algorithm in HORCL are applied to screen two levels of outliers (loop closure constraints and point pairs) and improve localization performance. Specifically, the inlier probabilities of data associations are calculated in MPM to identify outliers by considering geometric distances, semantic consistency, and spatial consistency. Then, outlier loop closures and outlier point pairs in inlier constraints are filtered by applying the Hierarchical EM algorithm, thereby relieving the adverse effect of outliers on localization accuracy. The proposed algorithm is validated on public datasets and compared with the latest methods, demonstrating the improvement in localization accuracy and robustness. The code is available at https://github.com/BIT-TYJ/HORCL. Meiling Wang 0002, Yinan Deng, Yi Yang 0009, Ziquan Lan, Yufeng Yue |
IROS | 6 |
| 2023 | SSGM: Spatial Semantic Graph Matching for Loop Closure Detection in Indoor EnvironmentsabstractCapturing the semantics of objects and the topological relationship allows the robot to describe the scene more intelligently like a human and measure the similarity between scenes (loop closure detection) more accurately. However, many current semantic graph matching methods are based on walk descriptors, which only extract adjacency relations between objects. In such way, the comprehensive information in the semantic graph is not fully exploited, which may lead to false closed-loop detection. This paper proposes a novel spatial semantic graph matching method (SSGM) in indoor environments, which considers multifaceted information of the semantic graphs. Firstly, two semantic graphs are aligned in the same coordinate space contributed by the second-order spatial compatibility metric between objects and local graph features of objects in semantic graphs. Secondly, the similarity of the spatial distribution of overall semantic graphs is further evaluated. The proposed algorithm is validated on public datasets and compared with the latest semantic graph matching methods, demonstrating improved accuracy and efficiency in loop closure detection. The code is available at https://github.com/BIT-TYJ/SSGM. Meiling Wang 0002, Yinan Deng, Yi Yang 0009, Yufeng Yue |
IROS | 5 |
| 2023 | L2V2T2Calib: Automatic and Unified Extrinsic Calibration Toolbox for Different 3D LiDAR, Visual Camera and Thermal CameraabstractExtrinsic calibration between LiDAR-Camera and LiDAR-LiDAR has been researched extensively, because it is the foundation for sensor fusion. Meanwhile, many projects are open-sourced and significantly promote related research. However, limited solutions can unify the calibration between repetitive scanning and non-repetitive scanning 3D LiDAR, sparse and dense 3D LiDAR, visual and thermal camera. Currently, to achieve that, we normally need to use different targets and extract different features for different sensor combinations. Sometimes, human intervention is required to locate the target. It is inconvenient and time-consuming. In this paper, L2V2T2Calib is introduced and open-sourced as a trial to unify the calibration. 1). A four-circular-holes board is adopted for all sensors. The four circle centers can be detected by all the sensors, thus are ideal common features. Previous works also use this target, but the algorithms don’t consider non-repetitive scanning LiDARs, thus cannot be directly applied. 2). To unify the process, an important step is to automatically and robustly detect the target from different types of LiDARs. However, this does not receive enough attention. We propose a method based on template matching. It is simple, but effective and general to different depth sensors. 3). We provide two types of output, minimizing 2D re-projection error (Min2D) and minimizing 3D matching error (Min3D), for different users. And their performance is compared. Extensive experiments conducted in both simulation and real environment demonstrate L2V2T2Calib is accurate, robust, more importantly, unified. The code will be open-sourced to promote related research at: https://github.com/Clothooo/lvt2calib Jun Zhang 0042, Yiyao Liu, Mingxing Wen, Yufeng Yue, Danwei Wang |
IV | 4 |
| 2023 | PointGPT: Auto-regressively Generative Pre-training from Point CloudsabstractLarge language models (LLMs) based on the generative pre-training transformer (GPT) have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, addressing the challenges associated with disorder properties, low information density, and task gaps. Specifically, a point cloud auto-regressive generation task is proposed to pre-train transformer models. Our method partitions the input point cloud into multiple point patches and arranges them in an ordered sequence based on their spatial proximity. Then, an extractor-generator based transformer decode, with a dual masking strategy, learns latent representations conditioned on the preceding point patches, aiming to predict the next one in an auto-regressive manner. To explore scalability and enhance performance, a larger pre-training dataset is collected. Additionally, a subsequent post-pre-training stage is introduced, incorporating a labeled hybrid dataset. Our scalable approach allows for learning high-capacity models that generalize well, achieving state-of-the-art performance on various downstream tasks. In particular, our approach achieves classification accuracies of 94.9% on the ModelNet40 dataset and 93.4% on the ScanObjectNN dataset, outperforming all other transformer models. Furthermore, our method also attains new state-of-the-art accuracies on all four few-shot learning benchmarks. Codes are available at https://github.com/CGuangyan-BIT/PointGPT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
NeurIPS | 6 |
| 2023 | PA3DNet: 3-D Vehicle Detection With Pseudo Shape Segmentation and Adaptive Camera-LiDAR Fusionabstract3-D vehicle detection is a key perception technique in autonomous driving. In this article, a novel 3-D vehicle detection framework that fuses camera images and LiDAR point clouds is proposed, named PA3DNet. The key novelties of PA3DNet is the proposing of a Pseudo Shape Segmentation (PSS) model and an Adaptive Camera-LiDAR Fusion (ACLF) module. The PSS model leverages self-assembled vehicle prototypes to learn shape-aware vehicle features. In order to achieve the adaptive fusion between visual semantics and LiDAR point features, learnable weight parameters are developed in the ACLF module to formulate an implicit complementarity between the two modalities. Extensive experiments on the widely used autonomous driving KITTI dataset demonstrate that PA3DNet achieves competitive accuracy when compared to advanced methods. It achieves 5.37% higher AP on Easy difficulty of 30-50m and 9.67% higher AP on Moderate difficulty of 50m. Meiling Wang 0002, Lin Zhao 0016, Yufeng Yue |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Integrated Localization and Planning for Cruise Control of UGV Platoons in Infrastructure-Free EnvironmentsabstractThis paper investigates the cruise control problem of unmanned ground vehicle (UGV) platoons from the implementation perspective. Unlike most existing works related to platoon cruise control which rely on positioning infrastructures such as lane markings, roadside units, and global navigation satellite systems (GNSS), this paper explores a new problem: platoon cruise control in environments without positioning infrastructures. The introduction of this constraint disables most existing cruise control approaches. To address this problem, an integrated localization and planning framework is proposed, which is composed of three modular algorithms. Firstly, to localize multiple vehicles in a common coordinate system, a collaborative localization algorithm is developed through matching local perceptions of different vehicles. Secondly, to maintain the desired platoon configuration, the historical trajectory of the preceding vehicle is reconstructed, based on which the target state is planned for the following vehicle. Finally, a virtual controller based algorithm is designed to generate feasible trajectories for the following vehicle in real time. The proposed framework has two salient features. Firstly, it does not depend on positioning infrastructures and does not introduce additional positioning sensors, such as GNSS/INS modules, ultra-wideband (UWB) devices, magnetic meters and so on, as long as each vehicle is equipped with a perception sensor (Lidar, radar or camera), which however is essential equipment for nowaday autonomous systems. Secondly, the proposed framework does not depend on direct observations between vehicles to achieve relative localization, making it applicable in non-line-of-sight (non-LOS) situations. Real-world experiments have been conducted to validate the effectiveness, robustness and practicality of the proposed framework. Yuanzhe Wang, Mingxing Wen, Yufeng Yue, Danwei Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Aerial-Ground Robots Collaborative 3D Mapping in GNSS-Denied EnvironmentsabstractCollaborative heterogeneous robots are expected to perform comprehensive perception, mapping and coordination in search and rescue scenarios. The challenge of collaboration between heterogeneous robots lies in their huge differences in perception, mobility and processing capabilities. In this paper, a novel collaborative UAV-UGV mapping framework is proposed in GNSS-denied and unknown environments. The key novelty of this work is the proposing of a unified framework to formulate the UAV-UGV collaborative mapping problem with a continuous-discrete model, as well as its realization in real robotic systems. In order to project continuous space into discrete space, a novel information gain trigger scheme is pro-posed. The continuous space allows each robot to perform high frequency local map estimation, while discrete space describes the problem of multi-resolution hybrid map fusion. Considering the nature of data heterogeneity, a flexible probabilistic fusion algorithm is proposed that addresses the multi-resolution hybrid map fusion problem, where the local maps generated by UAV and UGV are fused based on Bayesian rule. The proposed UAV-UGV hybrid system is validated in various challenging scenarios, demonstrating its accuracy and utility in practical tasks. Yufeng Yue, Yuanzhe Wang, Yi Yang 0009, Danwei Wang |
ICRA | 1 |
| 2022 | S-MKI: Incremental Dense Semantic Occupancy Reconstruction Through Multi-Entropy Kernel InferenceabstractAutonomous robots are often required to acquire high-level prior knowledge by continuously reconstructing the semantics and geometry of the surrounding scene, which is the basis of exploration and planning. Most existing continuous semantic mapping algorithms cannot distinguish potential differences in voxels, resulting in an over-inflated map. Furthermore, fixed-size query ranges introduce high computational complexity. Based on the limitation of over-inflation and inefficiency, this paper proposes a novel incremental continuous semantic occupancy mapping algorithm (S-MKI). The key innovation of this work comes from the two models in the preprocessing stage. On the one hand, Redundant Voxel Filter Model utilizes context entropy to filter out redundant voxels to improve the confidence of the final map, where objects have accurate boundaries with sharp edges. On the other hand, Adaptive Kernel Length Model adaptively adjusts the kernel length with class entropy, which reduces the inherent amount of training data. The final multientropy kernel inference function is formulated to integrate these two models to infer sparse noisy sensor data into dense accurate 3D maps. Experimental results conducted in both indoors and outdoors datasets validate that S-MKI outperforms existing methods. Yinan Deng, Meiling Wang 0002, Danwei Wang, Yufeng Yue |
IROS | 4 |
| 2022 | HD-CCSOM: Hierarchical and Dense Collaborative Continuous Semantic Occupancy Mapping through Label DiffusionabstractThe collaborative operation of multiple robots can make up for the shortcomings of a single robot, such as limited field of perception or sensor failure. multirobots collaborative semantic mapping can enhance their comprehensive contextual understanding of the environment. However, existing multirobots collaborative semantic mapping algorithms mainly apply discrete occupancy map inference, and do not compensate for inconsistent labels of local maps caused by differences in robot perspectives, which leads to greatly reduced availability and accuracy of the final global map. To address the challenges of discontinuous maps and inconsistent semantic labels, this paper proposes a novel hierarchical and dense collaborative continuous semantic occupancy mapping algorithm (HD-CCSOM). This work decomposes and formulates robot collaborative continuous semantic occupancy mapping problem at two levels. At the single robot level, the multi-entropy kernel inference method smoothly processes the registered semantic point cloud and infers a local continuous semantic occupancy map for each robot. At the collaborative robots level, the local maps are fused into a global enhanced and consistent semantic map via the label diffusion method based on a graph model. The proposed algorithm has been validated on public datasets and in simulated and real scenes, demonstrating significant improvements in mapping accuracy and efficiency. Yinan Deng, Meiling Wang 0002, Yi Yang 0009, Yufeng Yue |
IROS | 4 |
| 2021 | Semantic Reinforced Attention Learning for Visual Place RecognitionabstractLarge-scale visual place recognition (VPR) is inherently challenging because not all visual cues in the image are beneficial to the task. In order to highlight the task-relevant visual cues in the feature embedding, the existing attention mechanisms are either based on artificial rules or trained in a thorough data-driven manner. To fill the gap between the two types, we propose a novel Semantic Reinforced Attention Learning Network (SRALNet), in which the inferred attention can benefit from both semantic priors and data-driven fine-tuning. The contribution lies in two-folds. (1) To suppress misleading local features, an interpretable local weighting scheme is proposed based on hierarchical feature distribution. (2) By exploiting the interpretability of the local weighting scheme, a semantic constrained initialization is proposed so that the local attention can be reinforced by semantic priors. Experiments demonstrate that our method outperforms state-of-the-art techniques on city-scale VPR benchmark datasets. Guohao Peng, Yufeng Yue, Jun Zhang 0042, Zhenyu Wu 0001, Danwei Wang |
ICRA | 2 |
| 2021 | MSTSL: Multi-Sensor Based Two-Step Localization in Geometrically Symmetric EnvironmentsabstractSymmetric environment is one of the most intractable and challenging scenarios for mobile robots to accomplish global localization tasks, due to the highly similar geometrical structures and insufficient distinctive features. Existing localization solutions in such scenarios either depend on pre-deployed infrastructures which are expensive, inflexible, and hard to maintain; or rely on single sensor-based methods whose initialization module is incapable to provide enough unique information. Thus, this paper proposes a novel Multi-Sensor based Two-Step Localization framework named MSTSL, which addresses the problem of mobile robot global localization in geometrically symmetric environments by utilizing the measured magnetic field, 2-D LiDAR, and wheel odometry information. The proposed system mainly consists of two steps: 1) Magnetic Field-based Initialization, and 2) LiDAR-based Localization. Based on the pre-built magnetic field database, multiple initial hypotheses poses can firstly be determined by the proposed two-stage initialization algorithm. Then, utilizing the obtained multiple initial hypotheses, the robot can be localized more accurately by LiDAR-based localization. Extensive experiments demonstrate the practical utility and accuracy of the proposed system over the alternative approaches in real-world scenarios. Zhenyu Wu 0001, Yufeng Yue, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Danwei Wang |
ICRA | 2 |
| 2021 | Tightly-Coupled Perception and Navigation of Heterogeneous Land-Air Robots in Complex ScenariosabstractIn unstructured and unknown environments, heterogeneous robots must be able to perceive the environment, coordinate with each other and complete tasks collaboratively with onboard sensors. In this paper, a tightly-coupled perception and navigation framework is proposed for heterogeneous land-air robots, which forms a closed loop of perception-navigation for heterogeneous robots. The key novelty of this work is the proposing of a unified framework to formulate the cooperative mapping and navigation problem, as well as the derivation of high-level coordination strategy and low-level goal-oriented navigation within a fully integrated approach. To provide a comprehensive understanding of the environment, a flexible probabilistic map fusion algorithm is applied to merge local maps generated by hybrid robots. The proposed UAV-UGV hybrid system is validated in challenging experiments, proving its robustness and effectiveness in practical tasks. Yufeng Yue, Mingxing Wen, Yosmar Putra, Meiling Wang 0002, Danwei Wang |
ICRA | 1 |
| 2021 | Robust Semantic Map Matching Algorithm Based on Probabilistic Registration ModelabstractThe matching and fusing of local maps generated by multiple robots can greatly enhance the performance of relative localization and collaborative mapping. Currently, existing semantic matching methods are partly based on classical iterative closet point (ICP), which typically fail in cases with large initial error. What’s more, current semantic matching algorithms have high computation complexity in optimizing the transformation matrix. To address the challenge of map matching with large initial error, this paper proposes a novel semantic map matching algorithm with large convergence region. The key novelty of this work is the designing of the initial transformation optimization algorithm and the probabilistic registration model to increase the convergence region. To reduce the initial error before the iteration process, the initial transformation matrix is optimized by estimating the credibility of the data association. At the same time, a factor reflecting the uncertainty of the initial error is calculated and introduced to the formulation of the probabilistic registration model, thereby accelerating the convergence process. The proposed algorithm is performed on public datasets and compared with existing methods, demonstrating the significant improvement in terms of matching accuracy and robustness. Qingxiang Zhang, Meiling Wang 0002, Yufeng Yue |
ICRA | 3 |
| 2020 | HILPS: Human-in-Loop Policy Search for Mobile Robot NavigationabstractReinforcement learning has obtained increasing attention in mobile robot mapless navigation in recent years. However, there are still some obvious challenges including the sample efficiency, safety due to dilemma of exploration and exploitation. These problems are addressed in this paper by proposing the Human-in-Loop Policy Search (HILPS) framework, where learning from demonstration, learning from human intervention and Near Optimal Policy strategies are integrated together. Firstly, the former two make sure that expert experience grant mobile robot a more informative and correct decision for accomplishing the task and also maintaining the safety of the mobile robot due to the priority of human control. Then the Near Optimal Policy (NOP) provides a way to selectively store the similar experience with respect to the preexisting human demonstration, in which case the sample efficiency can be improved by eliminating exclusively exploratory behaviors. To verify the performance of the algorithm, the mobile robot navigation experiments are extensively conducted in simulation and real world. Results show that HILPS can improve sample efficiency and safety in comparison to state-of-art reinforcement learning. Mingxing Wen, Yufeng Yue, Zhenyu Wu 0001, Ehsan Mihankhah, Danwei Wang |
ICARCV | 2 |
| 2020 | Multi-Robot Collaborative Reasoning for Unique Person Recognition in Complex EnvironmentsabstractThe discovery of unique or suspicious people is essential for active surveillance of security or patrol robots, and multi-robot collaboration and dynamic reasoning can further enhance their adaptability in large-scale environments. This paper proposes a hierarchical probabilistic reasoning framework for a multi-robot system to actively identify the unique person with distinct motion patterns in large-scale and dynamic environments. Linear and angular velocities are considered typical motion patterns, which are extracted by using heterogeneous sensors to detect and track people. First, single robot reasoning is performed, each robot judges the uniqueness of people by comparing their motion patterns based on local observations. Meanwhile, multi-robot reasoning is also performed, by fusing the perceptual information from each individual robot to form a global observation and then make another judgment based on it. Finally, each robot can decide which result should be adopted by comparing the beliefs of local and global judgments. Experimental results show that the method is feasible in various environments. Chule Yang, Yufeng Yue, Mingxing Wen, Yuanzhe Wang |
ICARCV | 2 |
| 2020 | Human-Robot Teaming and Coordination in Day and Night EnvironmentsabstractAs robots are sharing work spaces with human, human-robot teamwork is becoming increasingly important. It is foreseeable that the daily work team will be composed of human and robots. The integration of the appropriate decision-making process is an essential part to design and develop the team. If robots can understand the activities and intents of human, it is convenient for a person to cooperate with robots in a natural manner. This paper proposes a system that enables robots to understand human pose and execute given command. The system provides two options for different hardware systems: the first one is suitable for powerful computational units; the second model is compact and efficient on a normal robot platform. In order to enrich application scenarios, we propose a method to extract human pose from thermal images so that our system can be used in all-weather scenario. In addition, we collected extensive training data and trained a MLP neural network to classify several human poses. The experimental results show the accuracy and efficiency of the proposed MLP neural network in day and night environments. Yufeng Yue, Yuanzhe Wang, Jun Zhang 0042, Danwei Wang |
ICARCV | 1 |
| 2020 | Day and Night Collaborative Dynamic Mapping in Unstructured Environment Based on Multimodal SensorsabstractEnabling long-term operation during day and night for collaborative robots requires a comprehensive understanding of the unstructured environment. Besides, in the dynamic environment, robots must be able to recognize dynamic objects and collaboratively build a global map. This paper proposes a novel approach for dynamic collaborative mapping based on multimodal environmental perception. For each mission, robots first apply heterogeneous sensor fusion model to detect humans and separate them to acquire static observations. Then, the collaborative mapping is performed to estimate the relative position between robots and local 3D maps are integrated into a globally consistent 3D map. The experiment is conducted in the day and night rainforest with moving people. The results show the accuracy, robustness, and versatility in 3D map fusion missions. Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang |
ICRA | 1 |
| 2020 | A Hierarchical Framework for Collaborative Probabilistic Semantic MappingabstractPerforming collaborative semantic mapping is a critical challenge for cooperative robots to maintain a comprehensive contextual understanding of the surroundings. Most of the existing work either focus on single robot semantic mapping or collaborative geometry mapping. In this paper, a novel hierarchical collaborative probabilistic semantic mapping framework is proposed, where the problem is formulated in a distributed setting. The key novelty of this work is the mathematical modeling of the overall collaborative semantic mapping problem and the derivation of its probability decomposition. In the single robot level, the semantic point cloud is obtained based on heterogeneous sensor fusion model and is used to generate local semantic maps. Since the voxel correspondence is unknown in collaborative robots level, an Expectation-Maximization approach is proposed to estimate the hidden data association, where Bayesian rule is applied to perform semantic and occupancy probability update. The experimental results show the high quality global semantic map, demonstrating the accuracy and utility of 3D semantic map fusion algorithm in real missions. Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Yuanzhe Wang, Danwei Wang |
ICRA | 1 |
| 2020 | Infrastructure-Free Global Localization in Repetitive Environments: An OverviewabstractRepetitive environment is a challenging scenario for mobile robot global localization due to its highly similar structures and lack of distinctive features. Existing solutions in such environments rely heavily on pre-installed infrastructures, which are neither flexible nor cost-effective. Besides, few of the previous research have been focused on the implementation of infrastructure-free localization approaches in repetitive scenarios. Thus, this paper serves as a survey to investigate the problem of infrastructure-free mobile robot global localization with low-cost and efficient sensors in repetitive environments. Three of the most popular infrastructure-free localization methods, namely LiDAR-based localization (LBL), vision-based localization (VBL), and magnetic field-based localization (MFL), are analyzed and evaluated. Extensive global localization experiments are conducted in real-world repetitive scenarios and the results demonstrate that VBL methods perform slightly better than LBL and MFL methods. The overall evaluations indicate that infrastructure-free global localization in repetitive environment is still a challenging problem which deserves more research efforts to develop new solutions. Zhenyu Wu 0001, Jun Zhang 0042, Yufeng Yue, Mingxing Wen, Zichen Jiang, Danwei Wang |
IECON | 3 |
| 2020 | Collaborative Semantic Perception and Relative Localization Based on Map MatchingabstractIn order to enable a team of robots to operate successfully, retrieving accurate relative transformation between robots is the fundamental requirement. So far, most research on relative localization mainly focus on geometry features such as points, lines and planes. To address this problem, collaborative semantic map matching is proposed to perform semantic perception and relative localization. This paper performs semantic perception, probabilistic data association and nonlinear optimization within an integrated framework. Since the voxel correspondence between partial maps is a hidden variable, a probabilistic semantic data association algorithm is proposed based on Expectation-Maximization. Instead of specifying hard geometry data association, semantic and geometry association are jointly updated and estimated. The experimental verification on Semantic KITTI benchmarks demonstrate the improved robustness and accuracy. Yufeng Yue, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang |
IROS | 1 |
| 2020 | Surrounding-aware correlation filter for UAV tracking with selective spatial regularization
Changhong Fu 0001, Weijiang Xiong, Fuling Lin, Yufeng Yue |
Signal Process. | 4 |
| 2019 | Probabilistic Reasoning for Unique Role Recognition Based on the Fusion of Semantic-Interaction and Spatio-Temporal FeaturesabstractThis paper deals with the problem of recognizing the unique role in dynamic environments. Different from social roles, the unique role refers to those who are unusual in their carrying items or movements in the scene. In this paper, we propose a hierarchical probabilistic reasoning method that relates spatial relationships between interested objects and humans with their temporal changes to recognize the unique individual. Two observation models, Object Existence Model (OEM) and Human Action Model (HAM), are established to support role inference by analyzing the corresponding semantic-interaction features and spatio-temporal features. Then, OEM and HAM results of each person are compared with the overall distribution in the scene, respectively. Finally, we can determine the role through the fusion of two observation models. Experiments are conducted in both indoor and outdoor environments concerning different settings, degrees of clutter, and occlusions. The results show that the proposed method can adapt to a variety of scenarios and outperforms other methods on accuracy and robustness, moreover, exhibiting stable performance even in complex scenes. Chule Yang, Yufeng Yue, Jun Zhang 0042, Mingxing Wen, Danwei Wang |
IEEE Trans. Multim. | 2 |
| 2018 | Point Cloud Based Path Planning for Tower Crane LiftingabstractThis paper discusses automatic path planning for tower crane lifting in highly complex environments to be digitized using point cloud representation. A mathematical optimization technique is developed to identify the lifting path with GPU accelerated massively parallel genetic algorithm. A continuous collision detection method is designed for real time application of collision avoidance during the crane lifting process. Lihui Huang, Jianmin Zheng, Panpan Cai, Souravik Dutta, Yufeng Yue, Nadia Magnenat-Thalmann, Yiyu Cai |
CGI | 6 |
| 2018 | Probabilistic Fusion Framework for Collaborative Robots 3D MappingabstractFusion of local 3D maps generated by individual robots to a globally consistent 3D map is one of the fundamental challenges in multi-robot mapping missions. In this paper, we propose a probabilistic mathematical formulation to address the integrated map fusion problem. More specifically, the problem of estimating fused map posterior can be factorized into a product of relative transformation posterior and the global map posterior, which enables us to solve map matching and map merging problems efficiently. In addition, a distributed communication strategy is employed to share map information among robots. The proposed approach is evaluated in indoor and mixed environments, which shows its utility in 3D map fusion for multi-robot mapping missions. Yufeng Yue, P. G. C. N. Senarathne, Chule Yang, Jun Zhang 0042, Mingxing Wen, Danwei Wang |
FUSION | 1 |
| 2018 | A Two-step Method for Extrinsic Calibration between a Sparse 3D LiDAR and a Thermal CameraabstractTo obtain the 6 DOF extrinsic parameters (rotation and translation matrix) between a 3D ranging sensor and a thermal camera, previous methods require a high-resolution 3D ranging sensor to reliably detect features. Although sparse 3D LiDARs are widely used on autonomous robots, to the best of our knowledge, the extrinsic calibration between a sparse 3D LiDAR (particularly Velodyne VLP-16) and a thermal camera has not been considered in the literature. In this paper, we present a two-step method to address the problem, where a monocular visual camera is used to assist the process. The proposed method decomposes the problem into two steps: extrinsic calibration between a sparse 3D LiDAR and a visual camera; extrinsic calibration between a visual camera and a thermal camera. Experiments are conducted to demonstrate the effectiveness of the proposed two-step method. Jun Zhang 0042, Prarinya Siritanawan, Yufeng Yue, Chule Yang, Mingxing Wen, Danwei Wang |
ICARCV | 3 |
| 2016 | A hybrid probabilistic and point set registration approach for fusion of 3D occupancy grid mapsabstractOne of the major challenges in multi-robot exploration is to fuse the partial maps generated by individual robots into a consistent global map. We address 3D volumetric map fusion by extending the well known iterative closest point(ICP) algorithm to include probabilistic distance and surface information. In addition, the relative transformation is evaluated based on Mahalanobis distance and map dissimilarities are integrated using relative entropy filter. The efficiency of the proposed algorithm is evaluated using maps generated from both simulated and real environments and is shown to generate more consistent global maps. Yufeng Yue, Danwei Wang, P. G. C. N. Senarathne, Diluka Moratuwage |
SMC | 1 |