VLDB 2026 Research / reviewers in the wild / expert
Meiling Wang 0002
dblp:17/1320-2 · also Mei-Ling Wang 0002
· DBLP profile ↗
40ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0002-3618-7423ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 20 since 2021Systems, architecture and hardware · 16 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Coupling Structural Descriptors With a Novel Semantic Graph Matching Approach for LiDAR Loop DetectionabstractOutdoor loop closure detection is essential for correcting odometry drift and constructing a globally consistent map. Semantic-graph-based approaches effectively model object-level topology and achieve strong loop closure performance; however, their effectiveness degrades in background-dominated scenes with few distinctive objects, and establishing accurate injective node correspondences remains challenging. In contrast, structural descriptor methods, though offering stronger environmental generality through spatial-distribution modeling, remain susceptible to LiDAR noise and the discriminative power of point-level features. These limitations motivate the need for a more robust method that combines adaptability with enhanced descriptive power. We propose a novel loop-closure detection framework, SAGE, that integrates highly adaptable point-cloud shape-distribution features and generally reliable semantic graph topology, adaptively combining their similarity measures to improve detection performance. Specifically, we design a semantic graph matching module with dual constraints, local graph feature consistency and global spatial consistency, to achieve more accurate injective node correspondences. In addition, we extract point-cloud shape-distribution features and introduce a fusion mechanism that integrates them with the semantic graph module, assessing reliability and adaptively weighting their contributions. Extensive loop closure detection and pose estimation experiments on various datasets demonstrate that SAGE achieves superior performance over strong baselines. We provide the code at https://github.com/SAGE-11/SAGE. Meiling Wang 0002, Sibo Zuo, Chengxi Yang, Jinhao Jiang, Xieyuanli Chen, Yufeng Yue |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic ManipulationabstractLearning from demonstration is a powerful method for robotic skill acquisition. Nevertheless, a critical limitation lies in the substantial costs associated with gathering demonstration datasets, typically action-labeled robot data, which creates a fundamental constraint in the field. Video data offer a compelling solution as an alternative rich data source, containing diverse behavioral and physical knowledge. This study introduces G3M, an innovative framework that exploits video data viaGraph-to-GraphsGenerativeModeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. The proposed G3M abstracts video frame into graph representations by identifying object and visual action vertices for capturing state information. It then effectively models internal structures and spatial relationships present in these graph constructions, with the objective of predicting forthcoming graphs. The generated graphs function as conditional inputs that guide the control policy in determining robotic behaviors. This concise method effectively encodes critical spatial relationships while facilitating accurate prediction of subsequent graph sequences, thus allowing the development of resilient control policy despite constraints in action-annotated training samples. Furthermore, these transferable graph representations enable the effective extraction of manipulation knowledge through human videos as well as recordings from robots with different embodiments. The experimental results demonstrate that G3M attains superior performance using merely 20% action-labeled data relative to comparable approaches. Moreover, our method outperforms the state-of-the-art method, showing performance gains exceeding 19% in simulated environments and 23% in real-world experiments, while delivering improvements of over 35% in cross-embodiment transfer experiments and exhibiting strong performance on long-horizon tasks. Our project page is available athttps://g3m-project.github.io/. Guangyan Chen, Meiling Wang 0002, Te Cui, Chengcai Yang, Mengxiao Hu, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
IEEE Trans. Robotics | 2 |
| 2025 | GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningabstractLearning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alternative. In this paper, we present GraphMimic, a novel paradigm that leverages video data via graph-to-graphs generative modeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. Specifically, GraphMimic abstracts video frames into object and visual action vertices, and constructs graphs for state representations. The graph generative modeling network then effectively models internal structures and spatial relationships within the constructed graphs, aiming to generate future graphs. The generated graphs serve as conditions for the control policy, mapping to robot actions. Our concise approach captures important spatial relations and enhances future graph generation accuracy, enabling the acquisition of robust policies from limited action-labeled data. Furthermore, the transferable graph representations facilitate the effective learning of manipulation skills from cross-embodiment videos. Our experiments exhibit that GraphMimic achieves superior performance using merely 20% action-labeled data. Moreover, our method outperforms the state-of-the-art method by over 17% and 23% in simulation and real-world experiments, and delivers improvements of over 33% in cross-embodiment transfer experiments. Guangyan Chen, Te Cui, Meiling Wang 0002, Chengcai Yang, Mengxiao Hu, Yao Mu 0001, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
CVPR | 3 |
| 2025 | ORA-NET: Enhancing Image Feature Matching through Oriented Overlapping Region AlignmentabstractImage feature matching is a fundamental task in computer vision. Existing local feature matching methods can establish robust correspondences between image pairs. However, these methods heavily rely on dense local image features, making them susceptible to significant perspective differences, characterized by rotation and scale changes. To alleviate this limitation, we introduce a novel oriented Overlapping Region Alignment method, named ORA-NET, which presents a concise and efficient approach to enhance the performance of image feature matching methods. We introduce the Multidirectional Cross-scale Feature Aggregation module to aggregate rotation-equivariant features across multiple scales and model long-range dependencies. Additionally, the Oriented Overlap Alignment module estimates scale and rotation differences within overlapping regions using a coarse-to-fine rotation correction approach. Importantly, our method serves as a plug-and-play module that can be seamlessly integrated into other correspondence matching pipelines. Experimental results demonstrate that ORA-NET significantly enhances the matching performance of existing local feature matching methods, particularly in scenarios involving substantial perspective differences. Te Cui, Meiling Wang 0002, Guangyan Chen, Yufeng Yue |
IROS | 2 |
| 2025 | Human Demonstrations are Generalizable Knowledge for RobotsabstractLearning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object instances. In this paper, we propose a different perspective, considering human demonstration videos not as mere instructions, but as a source of knowledge for robots. Motivated by this perspective and the remarkable comprehension and generalization capabilities exhibited by large language models (LLMs), we propose DigKnow, a method that DIstills Generalizable KNOWledge with a hierarchical structure. Specifically, DigKnow begins by converting human demonstration video frames into observation knowledge. This knowledge is then subjected to analysis to extract human action knowledge and further distilled into pattern knowledge that comprises task and object instances, resulting in the acquisition of generalizable knowledge with a hierarchical structure. In settings with different tasks or object instances, DigKnow retrieves relevant knowledge for the current task and object instances. Subsequently, the LLM-based planner conducts planning based on the retrieved knowledge, and the policy executes actions in line with the plan to achieve the designated task. Utilizing the retrieved knowledge, we validate and rectify planning and execution outcomes, resulting in a substantial enhancement of the success rate. Experimental results across a range of tasks and scenes demonstrate the effectiveness of this approach in facilitating real-world robots to accomplish tasks with the knowledge derived from human demonstrations. Te Cui, Tianxing Zhou, Mengxiao Hu, Zicai Peng, Haizhou Li 0004, Guangyan Chen, Meiling Wang 0002, Yufeng Yue |
IROS | 8 |
| 2025 | OpenObject-NAV: Open-Vocabulary Object-Oriented Navigation Based on Dynamic Carrier-Relationship Scene GraphabstractIn everyday life, frequently used objects like cups often have unfixed positions and multiple instances within the same category, and their carriers frequently change as well. As a result, it becomes challenging for a robot to efficiently navigate to a specific instance. To tackle this challenge, the robot must capture and update scene changes and plans continuously. However, current object navigation approaches primarily focus on the semantic level and lack the ability to dynamically update scene representation. To address these limitations, this paper captures the relationships between frequently used objects and their static carriers. Specifically, it constructs an open-vocabulary Carrier-Relationship Scene Graph (CRSG) and updates the carrying status during robot navigation to reflect the dynamic changes of the scene. Based on the CRSG, we further propose an instance navigation strategy that models the navigation process as a Markov Decision Process. At each step, decisions are informed by Large Language Model’s commonsense knowledge and visual-language feature similarity. We designed a series of long-sequence navigation tasks for frequently used everyday items in the Habitat simulator. The results demonstrate that by updating the CRSG, the robot can efficiently navigate to moved targets. Additionally, we deployed our algorithm on a real robot and validated its practical effectiveness. The project page can be found here: https://OpenObject-Nav.github.io. Meiling Wang 0002, Yinan Deng, Zibo Zheng, Jiagui Zhong, Chenjie Zhao, Yufeng Yue |
IROS | 2 |
| 2025 | SLOOP: Aligned Coordinate System-aided LiDAR LOOP Closure Detection based on Semantic Node Graph MatchingabstractLoop closure detection and pose estimation play a significant role in correcting odometry trajectories and generating globally consistent point cloud maps. Geometric feature descriptor methods neglect object-level spatial topology features, resulting in inadequate performance in loop closure detection. Semantic graph-based loop closing methods improve upon this, however, they still follow the paradigm of "first generating descriptors, then comparing similarity, and finally achieving alignment (6D pose)". Specifically, they compare two semantic graphs that are not spatially aligned, which makes direct node correspondences impossible and necessitates extensive descriptor extraction and comparison. This decouples similarity comparison from 6D pose estimation, resulting in a cumbersome process that limits practicality and scalability. This paper proposes SLOOP, a novel descriptor-free semantic graph matching method that "aligns two graphs first, followed by efficient similarity comparison". Specifically, we first design a dedicated neighborhood semantic feature module to extract high-quality matched node pairs. Next, we seek the aligned coordinate systems for candidate loops based on the robust ground normal vectors and two suitable node pairs examined by the two-stage global geometric consistency metrics. Finally, the aligned coordinate systems enable efficient extraction and comparison of node spatial distributions. We conducted extensive outdoor loop detection experiments and compared with various loop closure detection approaches, demonstrating the improved performance of SLOOP in loop closure detection and its practicality. The code and related materials are available at https://github.com/bit-tyj/sloop_c. Meiling Wang 0002, Jiagui Zhong, Sibo Zuo, Yinan Deng, Yufeng Yue |
IROS | 2 |
| 2025 | Unifying Latent Action and Latent State Pre-training for Policy Learning from VideosabstractVideo data provides an accessible and rich source beyond expensive action-labeled robot data for advancing robotic learning paradigms. Motivated by this potential, researchers investigate methods to exploit video data in robotic learning. Recent approaches can be primarily divided into two categories: Action-based approaches tokenize latent actions from videos for policy pre-training. State-based approaches pre-train models to predict subsequent states. The former establishes rich motion priors, while the latter empowers the robot to anticipate future events. These complementary capabilities suggest significant potential for integration into a unified framework. In this paper, we propose UniMimic, a novel approach unifying latent action and latent state pre-training from videos. We first train a unified tokenizer to learn latent states from video frames while deriving latent actions between state tokens. Subsequently, the policy is pre-trained on videos to predict these latent actions and subsequent latent states. Finally, the policy is fine-tuned on an action-labeled robot dataset to transfer the learned priors to precise robot execution. Experiments exhibit that our pre-training stage enhances the performance by 19% in the Libero benchmark and improves the average number of tasks completed in a row of 5 from 2.50 and 2.35 to 3.89 and 3.73 in the CALVIN benchmark. In the real-world experiments, our method still delivers improvements exceeding 36%. Guangyan Chen, Meiling Wang 0002, Te Cui, Luojie Yang, Lin Zhao 0016, Yi Yang 0009, Yufeng Yue |
SIGGRAPH Asia | 2 |
| 2025 | SeGraM: Aligned Coordinate System Aided Semantic Graph Matching Method for Loop Closure DetectionabstractCapturing object semantics and their spatial relationships is crucial to estimating scene similarity for loop closure detection. Existing semantic loop closure detection methods generally treat semantics as landmarks or extract the object topology to compare the similarity of frames. However, they often neglect the absolute spatial distribution of objects, which is essential to capture distinctive features of the scene. A fundamental requirement is to register the spatial coordinates of both frames in a unified reference frame. To address this, we construct aligned coordinate systems between two frames and extract absolute spatial distribution features of objects for loop closure detection. Building on this, we introduce SeGraM, a unified semantic graph matching approach applicable to both indoor and outdoor environments. Specifically, for each pair of semantic graphs, we first establish correspondences between nodes, referred to as node pairs. We then evaluate the geometric and semantic consistency of these pairs, along with the local graph features in the surrounding. To facilitate meaningful comparisons, two node pairs are carefully selected to establish aligned spherical coordinate systems, with ground normals to define the Z axes outdoors. SeGraM is validated in both indoor and outdoor scenarios and is compared with multiple algorithms, demonstrating improvements in loop closure detection accuracy. The code is accessible at https://github.com/BIT-TYJ/SeGraM. Meiling Wang 0002, Yinan Deng, Sibo Zuo, Yufeng Yue |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Point Tree Transformer for Point Cloud RegistrationabstractPoint cloud registration is a fundamental task in the fields of computer vision and robotics. Recent advancements in transformer-based methods have demonstrated enhanced performance in this domain. However, the standard attention mechanisms employed in these approaches tend to incorporate numerous points of low relevance, and therefore struggle to focus their attention weights on sparse yet meaningful points. This inefficiency leads to limited local structure modeling capabilities and quadratic computational complexity. To overcome these limitations, we propose the Point Tree Transformer (PTT), a novel transformer-based approach for point cloud registration that efficiently extracts comprehensive local and global features while maintaining linear computational complexity. The PTT constructs hierarchical feature trees from point clouds in a coarse-to-dense manner, and introduces a novel Point Tree Attention (PTA) mechanism. This mechanism adheres to the tree structure to facilitate the progressive convergence of attended regions toward salient points. Specifically, each tree layer selectively identifies a subset of relevant points with the highest attention scores, and subsequent layers focus attention on areas of significant relevance, derived from the child points of the selected point set. The feature extraction process additionally incorporates coarse point features that capture high-level semantic information, thus facilitating local structure modeling and the progressive integration of multiscale information. Consequently, the PTA enables the model to focus on essential local structures and extract intricate local information while maintaining linear computational complexity. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate that our method outperforms state-of-the-art methods in terms of performance. The code for our method is publicly available at https://github.com/CGuangyan-BIT/PTT. Meiling Wang 0002, Guangyan Chen, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | CoSTFE: Spatio-Temporal Feature Enhancement for Collaborative PerceptionabstractCollaborative perception enables a more comprehensive and precise representation of the environment, owing to the complementary information shared among different agents. However, spatio-temporal disturbances, including localization errors (spatial) and time delays (temporal), are prevalent in practical applications and significantly impair detection performance. To improve both the accuracy and robustness, a novel framework called Spatio-Temporal Feature Enhancement for Collaborative Perception (CoSTFE) is proposed. Specifically, we present a Histogram-based Spatial Correction (HSC) module to optimize the transformation matrix and promote the robustness when localization errors happen. In addition, the Deformable Temporal Augmentation (DTA) module is introduced to predict and enhance the current characteristic with long-term historical dynamics. Compared with existing methods on three publicly available collaborative perception datasets, our approach exhibits superior performance and robustness in the collaborative 3D object detection task. Meiling Wang 0002, Xunjie He, Yufeng Yue |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Fast and Robust Point Cloud Registration with Tree-based TransformerabstractPoint cloud registration is essential in computer vision and robotics. Recently, transformer-based methods have achieved advanced point cloud registration performance. However, the standard attention mechanism utilized in these methods considers many low-relevance points, and it has difficulty focusing its attention weights on sparse and meaningful points, leading to limited local structure modeling capabilities and quadratic computational complexity. To address these limitations, we present the Tree-based Transformer (TrT), which is able to extract abundant local and global features with linear computational complexity. Specifically, the TrT builds coarse-to-dense feature trees, and a novel Tree-based Attention (TrA) is proposed to guide the progressive convergence of the attended regions toward meaningful points and to structurize point clouds following tree structures. In each layer, the top ${\mathcal{S}}$ key points with the highest attention scores are selected, such that in the next layer, attention is evaluated only within the specified high-relevance regions, corresponding to the child points of these selected ${\mathcal{S}}$ points. Additionally, coarse features containing high-level semantic information are incorporated into the child points to guide the feature extraction process, facilitating local structure modeling and multiscale information integration. Consequently, TrA enables the model to focus on critical local structures and extract rich local information with linear computational complexity. Experiments demonstrate that our method achieves state-of-the-art performance on 3DMatch and KITTI benchmarks. The code for our method is publicly available at https://github.com/CGuangyan-BIT/TrT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
ICRA | 2 |
| 2024 | Self-supervised Monocular Depth Estimation in Challenging Environments Based on Illumination Compensation PoseNetabstractSelf-supervised depth estimation has attracted much attention due to its ability to improve the 3D perception capabilities of unmanned systems. However, existing unsupervised frameworks rely on the assumption of photometric consistency, which may not hold in challenging environments such as night-time, rainy nights, or snowy winters due to complex lighting and reflections, resulting in inconsistent photometry across different frames for the same pixel. To address this problem, we propose a self-supervised monocular depth estimation unified framework that can handle these complex scenarios, which has the following characteristics: (1) an Illumination Compensation PoseNet (ICP) is designed, which is based on the classic Phong illumination theory and compensates for lighting changes in adjacent frames by estimating per-pixel transformations; (2) a Dual-Axis Transformer (DAT) block is proposed as the backbone network of the depth encoder, which infers the depth of local repeat-texture areas through spatial-channel dual-dimensional global context information of images. Experimental results demonstrate that our approach achieves state-of-the-art depth estimation results in complex environments on the challenging Oxford RobotCar dataset. Shengyu Hou, Wenjie Song 0001, Rongchuan Wang, Meiling Wang 0002, Yi Yang 0009, Mengyin Fu |
IROS | 4 |
| 2024 | VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsabstractVisual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by more than 37% in long-horizon tasks. Code and videos are available on our anonymous homepage. Guangyan Chen, Meiling Wang 0002, Te Cui, Yao Mu 0001, Tianxing Zhou, Zicai Peng, Mengxiao Hu, Haizhou Li 0004, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
NeurIPS | 2 |
| 2024 | Full Transformer Framework for Robust Point Cloud Registration With Deep Information InteractionabstractPoint cloud registration is an essential technology in computer vision and robotics. Recently, transformer-based methods have achieved advanced performance in point cloud registration by utilizing the advantages of the transformer in order-invariance and modeling dependencies to aggregate information. However, they still suffer from indistinct feature extraction, sensitivity to noise, and outliers, owing to three major limitations: 1) the adoption of CNNs fails to model global relations due to their local receptive fields, resulting in extracted features susceptible to noise; 2) the shallow-wide architecture of transformers and the lack of positional information lead to indistinct feature extraction due to inefficient information interaction; and 3) the insufficient consideration of geometrical compatibility leads to the ambiguous identification of incorrect correspondences. To address the above-mentioned limitations, a novel full transformer network for point cloud registration is proposed, named the deep interaction transformer (DIT), which incorporates: 1) a point cloud structure extractor (PSE) to retrieve structural information and model global relations with the local feature integrator (LFI) and transformer encoders; 2) a deep-narrow point feature transformer (PFT) to facilitate deep information interaction across a pair of point clouds with positional information, such that transformers establish comprehensive associations and directly learn the relative position between points; and 3) a geometric matching-based correspondence confidence evaluation (GMCCE) method to measure spatial consistency and estimate correspondence confidence by the designed triangulated descriptor. Extensive experiments on the ModelNet40, ScanObjectNN, and 3DMatch datasets demonstrate that our method is capable of precisely aligning point clouds, consequently, achieving superior performance compared with state-of-the-art methods. The code is publicly available at https://github.com/CGuangyan-BIT/DIT. Guangyan Chen, Meiling Wang 0002, Qingxiang Zhang, Li Yuan 0007, Yufeng Yue |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Rethinking Point Cloud Registration as Masking and ReconstructionabstractPoint cloud registration is essential in computer vision and robotics. In this paper, a critical observation is made that the invisible parts of each point cloud can be directly utilized as inherent masks, and the aligned point cloud pair can be regarded as the reconstruction target. Motivated by this observation, we rethink the point cloud registration problem as a masking and reconstruction task. To this end, a generic and concise auxiliary training network, the Masked Reconstruction Auxiliary Network (MRA), is proposed. The MRA reconstructs the complete point cloud by separately using the encoded features of each point cloud obtained from the backbone, guiding the contextual features in the backbone to capture fine-grained geometric details and the overall structures of point cloud pairs. Unlike recently developed high-performing methods that incorporate specific encoding methods into transformer models, which sacrifice versatility and introduce significant computational complexity during the inference process, our MRA can be easily inserted into other methods to further improve registration accuracy. Additionally, the MRA is detached after training, thereby avoiding extra computational complexity during the inference process. Building upon the MRA, we present a novel transformer-based method, the Masked Reconstruction Transformer (MRT), which achieves both precise and efficient alignment using standard transformers. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate the superior performance of our MRT over state-of-the-art methods. Codes are available at https://github.com/CGuangyan-BIT/MRA. Guangyan Chen, Meiling Wang 0002, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
ICCV | 2 |
| 2023 | Conflict-constrained Multi-agent Reinforcement Learning Method for Parking Trajectory PlanningabstractAutomated Valet Parking (AVP) has been exten-sively researched as an important application of autonomous driving. Considering the high dynamics and density of real parking lots, a system that considers multiple vehicles simultaneously is more robust and efficient than a single vehicle setting as in most studies. In this paper, we propose a dis-tributed Multi-agent Reinforcement Learning(MARL) method for coordinating multiple vehicles in the framework of an AVP system. This method utilizes traditional trajectory planning to accelerate the learning process and introduces collision conflict constraints for policy optimization to mitigate the path conflict problem. In contrast to other centralized multi-agent path finding methods, the proposed approach is scalable, distributed, and adapts to dynamic stochastic scenarios. We train the models in random scenarios and validate in several artificially designed complex parking scenarios where vehicles are always disturbed by dynamic and static obstacles. Experimental results show that our approach mitigates path conflicts and excels in terms of success rate and efficiency. Meiling Wang 0002, Yi Yang 0009, Wenjie Song 0001 |
ICRA | 2 |
| 2023 | Deep Interactive Full Transformer Framework for Point Cloud RegistrationabstractPoint cloud registration is a crucial technology in the fields of robotics and computer vision. Despite the significant advances in point cloud registration enabled by Transformer-based methods, limitations persist due to indistinct feature extraction, noise sensitivity, and outlier handling. These limitations stem from three factors: (1) the inefficiency of convolutional neural networks (CNNs) to capture global relationships due to their local receptive fields, resulting in extracted features susceptible to noise; (2) the shallow-wide architecture of Transformers, coupled with a lack of positional information, leading to inefficient information interaction and indistinct feature extraction; and (3) the omission of geometrical compatibility leads to ambiguous identification of incorrect correspondences. To overcome these limitations, we propose the Deep Interactive Full Transformer (DIFT) network for point cloud registration, which consists of three key components: (1) a Point Cloud Structure Extractor (PSE) for modeling global relationships and retrieving structural information; (2) a Point Feature Transformer (PFT) for establishing comprehensive associations and directly learning the relative positions between points; and (3) a Geometric Matching-based Correspondence Confidence Evaluation (GMCCE) method for measuring spatial consistency and estimating correspondence confidence. Experimental results on ModelNet40 and 3DMatch datasets demonstrate the superior performance of our proposed method compared to existing state-of-the-art methods. The code for our method is publicly available at https://github.com/CGuangyan-BIT/DIFT. Guangyan Chen, Meiling Wang 0002, Qingxiang Zhang, Li Yuan 0007, Tong Liu 0009, Yufeng Yue |
ICRA | 2 |
| 2023 | Multi-View Robust Collaborative Localization in High Outlier Ratio Scenes Based on Semantic FeaturesabstractFiltering out outlier data associations between local maps can improve the robustness and accuracy of multi-robot localization. When the overlap is low and the field of view difference is large, it is likely to produce outlier data associations between local maps, which will reduce the matching accuracy and even lead to the failure of collaborative localization. To solve this problem, this paper proposes a novel outdoor robust collaborative localization algorithm (HORCL) capable for high outlier ratio scenes. The Mixture Probability Model (MPM) and the Hierarchical EM (Expectation Maximization) algorithm in HORCL are applied to screen two levels of outliers (loop closure constraints and point pairs) and improve localization performance. Specifically, the inlier probabilities of data associations are calculated in MPM to identify outliers by considering geometric distances, semantic consistency, and spatial consistency. Then, outlier loop closures and outlier point pairs in inlier constraints are filtered by applying the Hierarchical EM algorithm, thereby relieving the adverse effect of outliers on localization accuracy. The proposed algorithm is validated on public datasets and compared with the latest methods, demonstrating the improvement in localization accuracy and robustness. The code is available at https://github.com/BIT-TYJ/HORCL. Meiling Wang 0002, Yinan Deng, Yi Yang 0009, Ziquan Lan, Yufeng Yue |
IROS | 2 |
| 2023 | SSGM: Spatial Semantic Graph Matching for Loop Closure Detection in Indoor EnvironmentsabstractCapturing the semantics of objects and the topological relationship allows the robot to describe the scene more intelligently like a human and measure the similarity between scenes (loop closure detection) more accurately. However, many current semantic graph matching methods are based on walk descriptors, which only extract adjacency relations between objects. In such way, the comprehensive information in the semantic graph is not fully exploited, which may lead to false closed-loop detection. This paper proposes a novel spatial semantic graph matching method (SSGM) in indoor environments, which considers multifaceted information of the semantic graphs. Firstly, two semantic graphs are aligned in the same coordinate space contributed by the second-order spatial compatibility metric between objects and local graph features of objects in semantic graphs. Secondly, the similarity of the spatial distribution of overall semantic graphs is further evaluated. The proposed algorithm is validated on public datasets and compared with the latest semantic graph matching methods, demonstrating improved accuracy and efficiency in loop closure detection. The code is available at https://github.com/BIT-TYJ/SSGM. Meiling Wang 0002, Yinan Deng, Yi Yang 0009, Yufeng Yue |
IROS | 2 |
| 2023 | PointGPT: Auto-regressively Generative Pre-training from Point CloudsabstractLarge language models (LLMs) based on the generative pre-training transformer (GPT) have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, addressing the challenges associated with disorder properties, low information density, and task gaps. Specifically, a point cloud auto-regressive generation task is proposed to pre-train transformer models. Our method partitions the input point cloud into multiple point patches and arranges them in an ordered sequence based on their spatial proximity. Then, an extractor-generator based transformer decode, with a dual masking strategy, learns latent representations conditioned on the preceding point patches, aiming to predict the next one in an auto-regressive manner. To explore scalability and enhance performance, a larger pre-training dataset is collected. Additionally, a subsequent post-pre-training stage is introduced, incorporating a labeled hybrid dataset. Our scalable approach allows for learning high-capacity models that generalize well, achieving state-of-the-art performance on various downstream tasks. In particular, our approach achieves classification accuracies of 94.9% on the ModelNet40 dataset and 93.4% on the ScanObjectNN dataset, outperforming all other transformer models. Furthermore, our method also attains new state-of-the-art accuracies on all four few-shot learning benchmarks. Codes are available at https://github.com/CGuangyan-BIT/PointGPT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
NeurIPS | 2 |
| 2023 | PA3DNet: 3-D Vehicle Detection With Pseudo Shape Segmentation and Adaptive Camera-LiDAR Fusionabstract3-D vehicle detection is a key perception technique in autonomous driving. In this article, a novel 3-D vehicle detection framework that fuses camera images and LiDAR point clouds is proposed, named PA3DNet. The key novelties of PA3DNet is the proposing of a Pseudo Shape Segmentation (PSS) model and an Adaptive Camera-LiDAR Fusion (ACLF) module. The PSS model leverages self-assembled vehicle prototypes to learn shape-aware vehicle features. In order to achieve the adaptive fusion between visual semantics and LiDAR point features, learnable weight parameters are developed in the ACLF module to formulate an implicit complementarity between the two modalities. Extensive experiments on the widely used autonomous driving KITTI dataset demonstrate that PA3DNet achieves competitive accuracy when compared to advanced methods. It achieves 5.37% higher AP on Easy difficulty of 30-50m and 9.67% higher AP on Moderate difficulty of 50m. Meiling Wang 0002, Lin Zhao 0016, Yufeng Yue |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | S-MKI: Incremental Dense Semantic Occupancy Reconstruction Through Multi-Entropy Kernel InferenceabstractAutonomous robots are often required to acquire high-level prior knowledge by continuously reconstructing the semantics and geometry of the surrounding scene, which is the basis of exploration and planning. Most existing continuous semantic mapping algorithms cannot distinguish potential differences in voxels, resulting in an over-inflated map. Furthermore, fixed-size query ranges introduce high computational complexity. Based on the limitation of over-inflation and inefficiency, this paper proposes a novel incremental continuous semantic occupancy mapping algorithm (S-MKI). The key innovation of this work comes from the two models in the preprocessing stage. On the one hand, Redundant Voxel Filter Model utilizes context entropy to filter out redundant voxels to improve the confidence of the final map, where objects have accurate boundaries with sharp edges. On the other hand, Adaptive Kernel Length Model adaptively adjusts the kernel length with class entropy, which reduces the inherent amount of training data. The final multientropy kernel inference function is formulated to integrate these two models to infer sparse noisy sensor data into dense accurate 3D maps. Experimental results conducted in both indoors and outdoors datasets validate that S-MKI outperforms existing methods. Yinan Deng, Meiling Wang 0002, Danwei Wang, Yufeng Yue |
IROS | 2 |
| 2022 | HD-CCSOM: Hierarchical and Dense Collaborative Continuous Semantic Occupancy Mapping through Label DiffusionabstractThe collaborative operation of multiple robots can make up for the shortcomings of a single robot, such as limited field of perception or sensor failure. multirobots collaborative semantic mapping can enhance their comprehensive contextual understanding of the environment. However, existing multirobots collaborative semantic mapping algorithms mainly apply discrete occupancy map inference, and do not compensate for inconsistent labels of local maps caused by differences in robot perspectives, which leads to greatly reduced availability and accuracy of the final global map. To address the challenges of discontinuous maps and inconsistent semantic labels, this paper proposes a novel hierarchical and dense collaborative continuous semantic occupancy mapping algorithm (HD-CCSOM). This work decomposes and formulates robot collaborative continuous semantic occupancy mapping problem at two levels. At the single robot level, the multi-entropy kernel inference method smoothly processes the registered semantic point cloud and infers a local continuous semantic occupancy map for each robot. At the collaborative robots level, the local maps are fused into a global enhanced and consistent semantic map via the label diffusion method based on a graph model. The proposed algorithm has been validated on public datasets and in simulated and real scenes, demonstrating significant improvements in mapping accuracy and efficiency. Yinan Deng, Meiling Wang 0002, Yi Yang 0009, Yufeng Yue |
IROS | 2 |
| 2022 | Trajectory Prediction-Based Local Spatio-Temporal Navigation Map for Autonomous Driving in Dynamic Highway EnvironmentsabstractAutonomous driving, including intelligent decision-making and path planning, in dynamic environments (like highway) is significantly more difficult than the navigation in static scenarios because of the additional time dimension. Therefore, correlating the time dimension and the space dimension through prediction to create a spatio-temporal navigation map can make decision-making and path planning in such kinds of environment much easier. In this article, NGSIM data is analysed and processed from the perspective of the ego-vehicle (using the data as an ego-vehicle’s perception results). Based on the data, we develop an LSTM (Long-Short Term Memory)-based framework to predict possible trajectories of multiple surrounding vehicles within a certain range of the ego-vehicle. Then, the multiple predicted trajectories in a series of continuous dynamic highway scenes are projected into a spatio-temporal domain to create an octree map. Thus, dynamic targets and static obstacles can be unified into the same domain or map so that the dynamic disturbance problem for autonomous driving in highway environments can be resolved. Experimental results show that the proposed model is capable of predicting all the future trajectories around the ego-vehicle efficiently and the corresponding spatio-temporal map can be generated accurately in different dynamic scenarios. Mengyin Fu, Ting Zhang 0014, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Trajectory Planning Based on Spatio-Temporal Map With Collision Avoidance Guaranteed by Safety StripabstractTrajectory planning for the unmanned vehicle in the complex environment has always been a challenging task. Planned trajectory with the corresponding target velocity or acceleration sequence must be collision-free guaranteed and as comfortable as possible on the premise of obeying the traffic rules and interaction with other dynamic social vehicles. To meet this requirement, this paper proposes a framework for trajectory planning based on spatio-temporal map. Due to the time layer architecture in the map, the trajectory can be generated with velocity and acceleration simultaneously, and the whole trajectory is constrained within a ‘safety strip’, resulting in an efficient and safety guaranteed trajectory. The framework is composed of three sections: rough search, fine optimization and safety strip-based collision avoidance. For rough search, we propose an improved A* algorithm implemented in the discrete time layer to find out the suboptimal states efficiently. In fine optimization, the B-spline curve is exploited to connect the searched states into a continuous trajectory. And the optimal control points of B-spline are further grouped into several segments, forming the safety strip which is actually the distribution space of the planned trajectory. If necessary, an adjustment will be applied to keep the strip away from the collision zone, making the entire trajectory completely collision-free. Experiments on both public dataset and self-driving simulator show that the proposed framework can adapt to different kinds of complex traffic scenes well. Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | A Unified Framework Integrating Decision Making and Trajectory Planning Based on Spatio-Temporal Voxels for Highway Autonomous DrivingabstractIntelligent decision making and efficient trajectory planning are closely related in autonomous driving technology, especially in highway environment full of dynamic interactive traffic participants. This work integrates them into a unified hierarchical framework with long-term behavior planning (LTBP) and short-term dynamic planning (STDP) running in two parallel threads with different horizon, consequently forming a closed-loop maneuver and trajectory planning system that can react to the dynamic environment effectively and efficiently. In LTBP, a novel voxel structure and the ‘voxel expansion’ algorithm are proposed for the generation of driving corridors in 3D configuration, which involves the prediction states of surrounding vehicles. By using Dijkstra search, the maneuver with minimal cost is determined in form of voxel sequences, then a quadratic programming (QP) problem is constructed for solving the optimal trajectory. And in STDP, another small-scaled QP problem is performed to track or adjust the reference trajectory from LTBP in response to the dynamic obstacles. Meanwhile, a Responsibility-Sensitive Safety (RSS) Checker keeps running at high frequency for real-time feedback to ensure security. Experiments on real data collected in different highway scenarios demonstrate the effectiveness and efficiency of our work. Ting Zhang 0014, Wenjie Song 0001, Mengyin Fu, Yi Yang 0009, Xiaohui Tian, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Tightly-Coupled Perception and Navigation of Heterogeneous Land-Air Robots in Complex ScenariosabstractIn unstructured and unknown environments, heterogeneous robots must be able to perceive the environment, coordinate with each other and complete tasks collaboratively with onboard sensors. In this paper, a tightly-coupled perception and navigation framework is proposed for heterogeneous land-air robots, which forms a closed loop of perception-navigation for heterogeneous robots. The key novelty of this work is the proposing of a unified framework to formulate the cooperative mapping and navigation problem, as well as the derivation of high-level coordination strategy and low-level goal-oriented navigation within a fully integrated approach. To provide a comprehensive understanding of the environment, a flexible probabilistic map fusion algorithm is applied to merge local maps generated by hybrid robots. The proposed UAV-UGV hybrid system is validated in challenging experiments, proving its robustness and effectiveness in practical tasks. Yufeng Yue, Mingxing Wen, Yosmar Putra, Meiling Wang 0002, Danwei Wang |
ICRA | 4 |
| 2021 | Robust Semantic Map Matching Algorithm Based on Probabilistic Registration ModelabstractThe matching and fusing of local maps generated by multiple robots can greatly enhance the performance of relative localization and collaborative mapping. Currently, existing semantic matching methods are partly based on classical iterative closet point (ICP), which typically fail in cases with large initial error. What’s more, current semantic matching algorithms have high computation complexity in optimizing the transformation matrix. To address the challenge of map matching with large initial error, this paper proposes a novel semantic map matching algorithm with large convergence region. The key novelty of this work is the designing of the initial transformation optimization algorithm and the probabilistic registration model to increase the convergence region. To reduce the initial error before the iteration process, the initial transformation matrix is optimized by estimating the credibility of the data association. At the same time, a factor reflecting the uncertainty of the initial error is calculated and introduced to the formulation of the probabilistic registration model, thereby accelerating the convergence process. The proposed algorithm is performed on public datasets and compared with existing methods, demonstrating the significant improvement in terms of matching accuracy and robustness. Qingxiang Zhang, Meiling Wang 0002, Yufeng Yue |
ICRA | 2 |
| 2021 | Towards Autonomous Parking using Vision-only SensorsabstractExisting autonomous parking solutions usually require special signs, pre-built maps or accurate ranging sensors to achieve reliable perception of the parking environment, but these methods are difficult to popularize because they either require preconditions or are expensive for production cars. In this paper, we propose a vision-only autonomous parking solution based on only six cameras. Through the appropriate depth estimation algorithms, our method obtains the pixel level depth of the image, and constructs a dense point cloud, so as to realize the fine perception of the parking environment. An improved Radon transform based parking space detection method are applied for better parking space detection method. Our proposed method achieves processing speed of above 5 Hz on a intermediate level computing platform. Furthermore, we demonstrate the practicability of the proposed system in real-world parking lots. Yi Yang 0009, Miaoxin Pan, Sitan Jiang, Jianhang Wang, Meiling Wang 0002 |
IROS | 7 |
| 2020 | Dynamic Object Tracking for Self-Driving Cars Using Monocular Camera and LIDARabstractThe detection and tracking of dynamic traffic participants (e.g., pedestrians, cars, and bicyclists) plays an important role in reliable decision-making and intelligent navigation for autonomous vehicles. However, due to the rapid movement of the target, most current vision-based tracking methods, which perform tracking in the image domain or invoke 3D information in parts of their pipeline, have real-life limitations such as lack of the ability to recover tracking after the target is lost. In this work, we overcome such limitations and propose a complete system for dynamic object tracking in 3D space that combines: (1) a 3D position tracking algorithm based on monocular camera and LIDAR for the dynamic object; (2) a re-tracking mechanism (RTM) that restore tracking when the target reappears in camera's field of view. Compared with the existing methods, each sensor in our method is capable of performing its role to preserve reliability, and further extending its functions through a novel multimodality fusion module. We perform experiments in the real-world self-driving environment and achieve a desired 10Hz update rate for real-time performance. Our quantitative and qualitative analysis shows that this system is reliable for dynamic object tracking purposes of self-driving cars. Lin Zhao 0016, Meiling Wang 0002, Sheng Su, Tong Liu 0009, Yi Yang 0009 |
IROS | 2 |
| 2020 | Trajectory Prediction based on Constraints of Vehicle Kinematics and Social Interaction†abstractTrajectory prediction for vehicles is a popular subject since it is beneficial for efficient and secure trajectory planning. In structured traffic scenarios, the behaviour and motion of vehicles are heavily dependent on the social interaction constraints, such as road geometry and surrounding vehicles, and the kinematics model constraints, such as continuous heading and maximum acceleration. To take these factors into account, we analyse the particular characteristics of driving vehicles and propose a model that predicts the possible and feasible trajectory for host vehicle in 3 seconds. In this model, the trajectory of host vehicle takes the center-line as reference, imitates the leader vehicle and focuses on the social vehicles through attention concentration mechanism (ACM) with spatial and temporal information encoded in a fusion hidden state. Furthermore, in order to make the trajectory feasible for vehicle dynamics and kinematics, we introduce a prediction diagnosis method to check the continuous heading and maximum acceleration condition, pruning and adjusting the prediction candidates. Experiments on released public datasets show that this framework can well evaluate the traffic interactions and forecast the trajectory more accurately than common networks. Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002 |
SMC | 5 |
| 2020 | Unifying Analytical Methods With Numerical Methods for Traffic System Modeling and ControlabstractShockwaves lead to speed variation and capacity drop, which hamper the stationarity and throughput of traffic network greatly in reality. In order to dominate or suppress shockwaves, there exist two philosophies: the analytical and numerical methods to investigate various traffic management schemes. However, both are studied completely separately in the existing literature. In this paper, we primarily focus on the uniformity and combination of the two philosophies, especially in terms of traffic evolution and shockwave trajectory extraction. Aiming at exploring the uniformity, numerical methods are equipped with gradient boundary detection and polar-parameter coordinate projection to iteratively calculate traffic states and extract shockwave trajectories as line segments. By contrast, analytical methods derive traffic evolution from fundamental diagrams and present shockwave trajectories as vector graphics in the traffic time-space diagram. Furthermore, to rationally measure the accuracy of extracted trajectories, a calibration model is established to decrease angle and distance errors yielded by the two methods. In the uncontrolled and variable speed limit (VSL)-controlled bottleneck scenarios, the simulation results have shown that: 1) both analytical and numerical methods have the capability to precisely describe traffic evolution and the effects of VSL strategies on recovering traffic throughput; 2) shockwave trajectories extracted by the two methods are coincident, with the significant reduction of relative distance errors from 15%-80% to 0.1%-3%; and 3) it is quite promising to take full advantage of both the visual/intuitive nature of analytical methods and the iterative optimization of numerical methods to investigate efficient traffic management strategies. Yeqing Zhang, Meiling Wang 0002, Ümit Özgüner |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2018 | Real-Time Obstacles Detection and Status Classification for Collision Warning in a Vehicle Active Safety SystemabstractThis paper presents real-time obstacles detection and their status classification method for collision warning in the vehicle active safety system. Specifically, stereo cameras and millimeter wave (mmw)-radar are fused to help the driving ego-vehicle to find “Danger” or “Potential Danger” in a timely way through combining with the vehicle kinematic model. The proposed method makes full use of the unique advantages of stereo cameras and mmw-radar to sense the environment through several modules. Cameras are mainly used to detect the near or lateral dynamic objects and to obtain the obstacles region of interest (ROI) considering its rich information and high sensitivity to the lateral displacement, while far or longitudinal relative dynamic objects are detected by mmw-radar according to its observational ability to make up for the disadvantage of cameras. In detail, a cameras detector utilizes ”error vectors” rather than the optical flow to obtain dynamic classes through two times clustering. Mmw-radar mainly detects relative dynamic objects, whose absolute speed can be computed according to the ego-vehicle's state. Then, the detected objects of these two detectors are integrated in an obstacles ROI map, which is obtained through an UV-disparity obstacles detection algorithm to get the final dynamic and relative dynamic objects. Finally, they are classified by comparing them with a dangerous area that is acquired according to the vehicle kinematic model in a special vehicle coordinate system, which is fixed to the ground temporarily. This method is tested on our mobile platforms and the results prove that it can work effectively even though the ego-vehicle drives quickly. Wenjie Song 0001, Yi Yang 0009, Mengyin Fu, Fan Qiu, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | Real-time lane detection and forward collision warning system based on stereo visionabstractThis paper presents a real-time and robust lane detection and forward collision warning technique based on stereo cameras. First, obstacles image is obtained through stereo matching and UV-disparity segmentation algorithm. Then, Inverse Perspective Mapping(IPM) and Sobel filtering are conducted to generate a low-noise top view of the road by fusing the obstacles image and the original image. Next, Hough Transformation for the top view map is completed and the extreme points(poles) are calculated as the detected lanes according to the traffic lanes model. Besides, the host lane is selected or supplemented among all the detected lanes and the nearest obstacle in this host lane is detected for the forward collision warning. Experimental results on the public data set indicate that our method can work effectively and real-timely in the normal structured environment. Wenjie Song 0001, Mengyin Fu, Yi Yang 0009, Meiling Wang 0002, Xinyu Wang 0018, Alain L. Kornhauser |
Intelligent Vehicles Symposium | 4 |
| 2017 | An efficient decision and planning method for high speed autonomous driving in dynamic environmentabstractThis paper describes an improved decision and planning algorithm based on our previously proposed methods for unmanned ground vehicle (UGV). The new method can be applied to UGV driving both in structured environment and unstructured environment. In the improved method, the prospect of planning is extended from 40m to 100m for safe driving at high speed and some piecewise linear speed functions are designed for the new prospect. After this improvement our UGV now can drive at a maximum speed of 60km/h rather than 40km/h while avoiding obstacles safely. Besides, a velocity feedforward control is added to make the UGV overtake other cars driving at about 25km/h on the road. At last, the collision detection algorithm is improved to make the lane changing maneuver safer. The proposed decision and planning algorithm is implemented both on our old Polaris all terrain vehicle (ATV) and new FAW-H7 car, which exhibited good performance on Across Dangers & Obstacles 2016, Tahe, China and Future Challenge 2016, Changshu, China, respectively. Kai Zhang 0030, Mengyin Fu, Yi Yang 0009, Songtian Shang, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 5 |
| 2015 | Collision-free and kinematically feasible path planning along a reference path for autonomous vehicleabstractFor the local path planning problem of autonomous vehicle in a complicated environment, a method combining cubic hermite spline curves with the kinematic model of autonomous vehicle is developed. And a novel algorithm for obstacle avoidance, called navigation circle, is proposed to take the road structure into account, which is a practical method for real-time path planning. In the new method, one of the trajectory generated by cubic hermite spline curves or navigation circle is optimized through the kinematic model of autonomous vehicle to get the kinematically feasible trajectory. The optimization is actually a numerical forward propagation and is easy to implement. The simulation experiment is conducted on the Robot Operating System (ROS) platform, which is based on replaying the data of the real world obtained from sensors or other modules on autonomous vehicle. Satisfactory simulation results verify the validity and the efficiency of the proposed method as well as the planner's capability to navigate in a realistic scenario. Mengyin Fu, Kai Zhang 0030, Yi Yang 0009, Hao Zhu 0002, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 5 |
| 2014 | Standing-up control and ramp-climbing control of a spherical wheeled robotabstractThis paper proposes a new type of spherical wheeled robot with an annular support leg. It can keep statically stable when powered off and automatically stand up with the assistance of the support leg when powered on. The stability of the robot at equilibrium is verified firstly using the planar simplified model. And the robot is proved to be controllable. Thus a double-closed loop control system is designed to stabilize the robot. Based on it, the standing-up control system and ramp-climbing control system are realized by changing control structure and using fuzzy control strategies. The design of the annular support leg, the experimental results and conclusions are also described in this paper. Jian Jian, Meiling Wang 0002, Ningyi Lv, Yi Yang 0009, Tong Liu 0009 |
ICARCV | 2 |
| 2014 | Moving object detection under dynamic background in 3D range dataabstractWe proposed an unsupervised algorithm to extract profile features and detect moving object under dynamic background in 3D range Data. Moving object detection under dynamic background has become an increasingly popular research topic in mobile robotics. For the characteristics of dynamic background scene, we proposed an online unsupervised moving object detection algorithm, based on Gaussian Mixture Models and Motion Compensation. Furthermore, we did the work of clustering and identifying of the targets. In order to improve the robustness of the algorithm, we used a tracker to track the results of the detection. At last, experimental results on real laser data depicting urban and rural scenes under static and dynamic background are presented. Yi Yang 0009, Yan Guang, Hao Zhu 0002, Mengyin Fu, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 5 |
| 2013 | Lane recognition self-learning scheme of mobile robot based on integrated perception systemabstractIn this paper, a kind of integrated perception system for mobile robot is presented, which consists of 3D Lidar, 2D camera and their spatial registration. Based on the system and support vector machine (SVM), a self-supervised learning scheme between 3D point cloud data and 2D image data has been established, which can identify the traversable lane in driving environments through data association and parameters training. With this approach, vision-based autonomous navigation can be achieved and its effectiveness has been verified by extensive robot experiments. Yi Yang 0009, Hao Zhu 0002, Mengyin Fu, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 4 |