VLDB 2026 Research / reviewers in the wild / expert
Yi Yang 0009
dblp:33/4854-9
· DBLP profile ↗
68ranked-venue papers
7as first author
55since 2021 · last 2026
0000-0003-3964-2433ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 7 first-author · 41 since 2021Systems, architecture and hardware · 35 · 3 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tackling Narrow-Space Parallel Parking: Reeds-Shepp-integrated Reinforcement Learning with Learnable Cost Heuristic
Shuaicong Yang, Mengying Ruan, Yi Yang 0009, Ting Zhang 0014, Mengyin Fu |
IV | 4 |
| 2026 | AdaptRGB-t: Adaptive RGB-t semantic segmentation via efficient parameter-tuning with textual guidance
Yufeng Yue, Yi Yang 0009, Mengyin Fu |
Neurocomputing | 3 |
| 2026 | Pedestrian-aware end-to-end autonomous parking via coupling-regulated multi-task learning
Mengying Ruan, Yuyi Zhou, Yi Yang 0009, Mengyin Fu, Ting Zhang 0014 |
Knowl. Based Syst. | 3 |
| 2026 | ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching With Collaborative LLM-Agents
Hao Liang 0016, Hao Li 0075, Jiyuan Guo, Yufeng Yue, Mengyin Fu, Yi Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic ManipulationabstractLearning from demonstration is a powerful method for robotic skill acquisition. Nevertheless, a critical limitation lies in the substantial costs associated with gathering demonstration datasets, typically action-labeled robot data, which creates a fundamental constraint in the field. Video data offer a compelling solution as an alternative rich data source, containing diverse behavioral and physical knowledge. This study introduces G3M, an innovative framework that exploits video data viaGraph-to-GraphsGenerativeModeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. The proposed G3M abstracts video frame into graph representations by identifying object and visual action vertices for capturing state information. It then effectively models internal structures and spatial relationships present in these graph constructions, with the objective of predicting forthcoming graphs. The generated graphs function as conditional inputs that guide the control policy in determining robotic behaviors. This concise method effectively encodes critical spatial relationships while facilitating accurate prediction of subsequent graph sequences, thus allowing the development of resilient control policy despite constraints in action-annotated training samples. Furthermore, these transferable graph representations enable the effective extraction of manipulation knowledge through human videos as well as recordings from robots with different embodiments. The experimental results demonstrate that G3M attains superior performance using merely 20% action-labeled data relative to comparable approaches. Moreover, our method outperforms the state-of-the-art method, showing performance gains exceeding 19% in simulated environments and 23% in real-world experiments, while delivering improvements of over 35% in cross-embodiment transfer experiments and exhibiting strong performance on long-horizon tasks. Our project page is available athttps://g3m-project.github.io/. Guangyan Chen, Meiling Wang 0002, Te Cui, Chengcai Yang, Mengxiao Hu, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
IEEE Trans. Robotics | 10 |
| 2025 | GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningabstractLearning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alternative. In this paper, we present GraphMimic, a novel paradigm that leverages video data via graph-to-graphs generative modeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. Specifically, GraphMimic abstracts video frames into object and visual action vertices, and constructs graphs for state representations. The graph generative modeling network then effectively models internal structures and spatial relationships within the constructed graphs, aiming to generate future graphs. The generated graphs serve as conditions for the control policy, mapping to robot actions. Our concise approach captures important spatial relations and enhances future graph generation accuracy, enabling the acquisition of robust policies from limited action-labeled data. Furthermore, the transferable graph representations facilitate the effective learning of manipulation skills from cross-embodiment videos. Our experiments exhibit that GraphMimic achieves superior performance using merely 20% action-labeled data. Moreover, our method outperforms the state-of-the-art method by over 17% and 23% in simulation and real-world experiments, and delivers improvements of over 33% in cross-embodiment transfer experiments. Guangyan Chen, Te Cui, Meiling Wang 0002, Chengcai Yang, Mengxiao Hu, Yao Mu 0001, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
CVPR | 11 |
| 2025 | DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone ImageryabstractDrones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rendering quality, providing a new avenue for 3D reconstruction from drone imagery. However, dynamic distractors in wild environments challenge the static scene assumption in radiance fields, while limited view constraints hinder the accurate capture of underlying scene geometry. To address these challenges, we introduce DroneSplat, a novel framework designed for robust 3D reconstruction from in-the-wild drone imagery. Our method adaptively adjusts masking thresholds by integrating local-global segmentation heuristics with statistical approaches, enabling precise identification and elimination of dynamic distractors in static scenes. We enhance 3D Gaussian Splatting with multi-view stereo predictions and a voxel-guided optimization strategy, supporting high-quality rendering under limited view constraints. For comprehensive evaluation, we provide a drone-captured 3D reconstruction dataset encompassing both dynamic and static scenes. Extensive experiments demonstrate that DroneSplat outperforms both 3DGS and NeRF baselines in handling in-the-wild drone imagery. Project page: https://bityia.github.io/DroneSplat/. Jiadong Tang, Yu Gao 0040, Dianyi Yang, Liqi Yan, Yufeng Yue, Yi Yang 0009 |
CVPR | 6 |
| 2025 | High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic GraspingabstractIn various robotic applications, understanding accurate object poses for robots is essential for high-precision tasks such as factory assembly or daily insertions. Tactile sensing, which compensates for visual information, offers rich texture-based or force-based data for object pose estimation. However, previous methods for pose estimation typically over-look dynamic situations, such as slippage of grasped objects or movement of contacted objects during interactions with the environment, thus increasing the complexity of pose estimation. To address these challenges, we propose an efficient method that utilizes visual and tactile sensing to estimate object poses through particle filtering. We leverage visual information to track the pose of the contacted object in real-time and estimate the pose changes of the grasped object using displacement data obtained from tactile sensors. Our experimental evaluation on 13 objects with diverse geometric shapes demonstrated the ability to estimate high-precision poses, which revealed the robot's powerful ability to cope with dynamic scenes for compelled motion of objects, proving our framework's adaptability in practical scenarios with uncertainty. Zicai Peng, Te Cui, Guangyan Chen, Yi Yang 0009, Yufeng Yue |
ICRA | 5 |
| 2025 | LACNS: Language-Assisted Continuous Navigation in Structured SpacesabstractCurrent autonomous driving technology typically relies on high-precision (HD) maps to ensure safe, reliable, and accurate navigation in urban environments. While these maps provide essential road information, their creation and maintenance are costly, limiting their widespread application. To mitigate this reliance, we propose a novel system, Language-Assisted Continuous Navigation in Structured Spaces (LACNS). LACNS facilitates autonomous driving without the need for HD maps by integrating vehicle-centric local perception with real-time language instructions from map software or human navigators. LACNS begins by generating a BEV map using the vehicle's front-facing camera. Simultaneously, a pretrained Visual Language Model (VLM) detects intersections from the camera images, assigning a score to each. Road elements are then extracted from the BEV map and combined with the intersection scores to identify potential navigation frontiers. Language instructions, processed by a pretrained Large Language Model(LLM), are used to select the most suitable frontier. Finally, the chosen frontier and BEV map are employed to plan a safe route and control the vehicle's movement. We evaluated LACNS using the Carla simulator to validate its navigation capabilities in continuous spaces. Initial experiments involved navigating through four intersections with varying directional instructions, where LACNS demonstrated high and consistent success rates across multiple trials. Further simulations in real-time navigation scenarios revealed that LACNS consistently maintained a high success rate across three progressively challenging routes. These results highlight the effectiveness of our novel autonomous driving navigation method without HD maps. Rutong Peng, Yi Yang 0009, Mengyin Fu |
ICRA | 3 |
| 2025 | UDSV: Unsupervised Deep Stitching for Tractor-Trailer Surround ViewabstractIn recent years, with the rapid development of Advanced Driver Assistance Systems (ADAS), the demand for the precise and efficient surround view stitching system has significantly increased. Traditional stitching methods perform well in small single-unit vehicles with stable camera poses. However, the stitching quality sharply degrades when applied to large tractor-trailers due to the continuous pose changes caused by the non-rigid connection between the tractor and trailer. In detail, first, the extended length of tractor-trailers results in low overlap between cameras, making feature extraction and matching challenging. Additionally, the stitched images often appear irregular, detracting from visual quality. Besides, even if static stitching looks natural, it causes jitter in dynamic scenarios due to random feature extraction. In this paper, we propose an unsupervised deep stitching method for tractor-trailer surround view system. We introduce a feature extraction module for tractor-trailer scenarios (FMT) to enhance feature extraction in low-overlap situations. Besides, we design a spatio-temporally consistent control point constraint strategy (STCC) to achieve spatial shape preservation and temporal smoothing effects, resulting in visually consistent and stable stitched sequences. Experimental results from both public and real dataset show that our method efficiently completes tractor-trailer surround view stitching, producing well-aligned and natural panoramic images compared to previous methods. Leyao Sun, Hao Liang 0016, Yi Yang 0009, Mengyin Fu |
ICRA | 4 |
| 2025 | OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene UnderstandingabstractRecent advancements in 3D Gaussian Splatting have significantly improved the efficiency and quality of dense semantic SLAM. However, previous methods are generally constrained by limited-category pre-trained classifiers and implicit semantic representation, which hinder their performance in open-set scenarios and restrict 3D object-level scene understanding. To address these issues, we propose OpenGS-SLAM, an innovative framework that utilizes 3D Gaussian representation to perform dense semantic SLAM in open-set environments. Our system integrates explicit semantic labels derived from 2D foundational models into the 3D Gaussian framework, facilitating robust 3D object-level scene understanding. We introduce Gaussian Voting Splatting to enable fast 2D label map rendering and scene updating. Additionally, we propose a Confidence-based 2D Label Consensus method to ensure consistent labeling across multiple views. Furthermore, we employ a Segmentation Counter Pruning strategy to improve the accuracy of semantic scene representation. Extensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of our method in scene understanding, tracking, and mapping, achieving 10× faster semantic rendering and 2× lower storage costs compared to existing methods. Project page: https://young-bit.github.io/opengs-github.github.io/. Dianyi Yang, Yu Gao 0040, Xihan Wang, Yufeng Yue, Yi Yang 0009, Mengyin Fu |
ICRA | 5 |
| 2025 | Internal-Stably Energy-Saving Cooperative Control of Articulated Wheeled Robot with Distributed Drive UnitsabstractArticulated wheeled robots play a crucial role in the logistics industry. However, conventional tractor-driven articulated wheeled robots exhibit poor internal stability and are prone to jackknifing, while also consuming a significant amount of energy. By deploying distributed drives and coordinating control among multiple drives, these issues can be effectively addressed. However, the flexible connections between the bodies of articulated vehicles pose significant challenges to the coordinated control of distributed drives. This paper proposes a multi-drive unit coordinated control algorithm based on driving force equivalence and allocation. A neural network is used to predict the driving force, and through non-linear driving force equivalence, a feedforward driving force is obtained. This is combined with a closed-loop feedback compensation controller to form a control architecture that integrates feedforward and feedback, resulting in the equivalent total driving force for the vehicle queue. Subsequently, an equivalent distribution strategy allocates the required driving force to each drive, enabling the vehicle bodies to achieve accurate and stable speed tracking while allowing each drive to operate near its efficient operating point, thereby reducing total energy consumption. Experiments demonstrate that our algorithm significantly lowers the total energy consumption of the vehicle queue under standard operating conditions while ensuring speed-tracking accuracy and improving internal stability. Yi Yang 0009, Huishuai Peng, Zhexi Hu |
ICRA | 1 |
| 2025 | CDMFusion: RGB-T Image Fusion Based on Conditional Diffusion Models via Few Denoising Steps in Open Environments
Luojie Yang, Lijin Fang, Yi Yang 0009, Yufeng Yue |
ICRA | 4 |
| 2025 | Open-RGBT: Open-Vocabulary RGB-T Zero-Shot Semantic Segmentation in Open-World EnvironmentsabstractSemantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenarios due to their reliance on pretrained models and predefined categories. Recent advancements in Visual Language Models (VLMs) have facilitated a shift from closedset to open-vocabulary semantic segmentation methods. However, these models face challenges in dealing with intricate scenes, primarily due to the heterogeneity between RGB and thermal modalities. To address this gap, we present Open-RGBT, a novel open-vocabulary RGB-T semantic segmentation model. Specifically, we obtain instance-level detection proposals by incorporating visual prompts to enhance category understanding. Additionally, we employ the CLIP model to assess image-text similarity, which helps correct semantic consistency and mitigates ambiguities in category identification. Empirical evaluations demonstrate that Open-RGBT achieves superior performance in diverse and challenging real-world scenarios, even in the wild, significantly advancing the field of RGB-T semantic segmentation. The project page of Open-RGBT is available at https://OpenRGBT.github.io/. Yufeng Yue, Luojie Yang, Xunjie He, Yi Yang 0009, Mengyin Fu |
ICRA | 5 |
| 2025 | Parking-SG: Open-Vocabulary Hierarchical 3D Scene Graph Representation for Open Parking EnvironmentsabstractAutomatic Valet Parking (AVP) has garnered significant attention from industry and academia due to its potential to enhance traffic efficiency, parking safety, and user experience. While AVP technologies have been successfully applied in standard parking scenarios with clear markings, real-world parking environments are far more diverse and complex, posing challenges for current systems. To address these limitations, we present Parking-SG, an open-vocabulary hierarchical 3D scene graph representation, facilitating the application of AVP in open and complex environments. Our approach builds an object-based, open-vocabulary map that integrates both ground-level and ground-above objects for comprehensive environmental understanding. Leveraging common sense reasoning and object behavior relationships, various standard or non-standard parking spaces are inferred in open environments. Additionally, we extract and analyze path topology to construct a hierarchical map representation, supporting complex AVP tasks. Parking-SG is validated in both simulated and real-world environments, demonstrating its ability to generate rich environmental representations, accurately and flexibly infer parking spaces, and effectively perform complex AVP tasks. Yi Ruan, Miaoxin Pan, Yi Yang 0009, Mengyin Fu |
ICRA | 4 |
| 2025 | UDSH: An Unsupervised Deep Image Stitching and De-Occlusion Method for Heavy Occlusion SceneabstractImage stitching in heavy occlusion scenarios faces the dual challenges of accurate alignment and occlusion removal. On one hand, occlusion causes the loss of key texture and structural information in the image. On the other hand, it affects the image’s integrity. Existing stitching methods perform well in cases with small occlusion coverage, but they often fail in heavy occlusion. This failure is mainly due to three reasons: 1) they cannot identify occluded regions, 2) they cannot suppress interference from the occluded regions, 3) they cannot remove the occluded regions. To address these issues, we propose an unsupervised deep image stitching and de-occlusion method. First, to solve the issue of occluded region identification, we design an Occlusion-Aware Feature Weighted module (OAFW) that explicitly distinguishes between occluded and non-occluded regions by learning the occlusion masks of the images. Second, to address the issue of interference from occlusion, we use the learned occlusion masks to filter out features from the occluded regions. To further suppress the impact of occlusion-induced errors, we design a Mask-Guided Dual-Granularity Alignment loss function (MGDGA) that only calculates alignment errors for non-occluded regions, effectively reducing occlusion error interference during network training. Finally, to resolve the content gap in the occluded regions, we replace the pixels in the occluded areas with those from the aligned overlapping regions and incorporate a Progressive Content Inpainting module (PCI) to recover the missing content in the non-overlapping regions caused by occlusion, ultimately achieving a complete and natural de-occlusion stitched image. Experimental results show that our method improves the mean squared error metric by 17.45% compared to the state-of-the-art stitching method. Hao Li 0075, Rundong Sun, Yi Yang 0009, Mengyin Fu |
IROS | 4 |
| 2025 | OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel RepresentationabstractIn recent years, vision-language models (VLMs) have advanced open-vocabulary mapping, enabling mobile robots to simultaneously achieve environmental reconstruction and high-level semantic understanding. While integrated object cognition helps mitigate semantic ambiguity in point-wise feature maps, efficiently obtaining rich semantic understanding and robust incremental reconstruction at the instance-level remains challenging. To address these challenges, we introduce OpenVox, a real-time incremental open-vocabulary probabilistic instance voxel representation. In the front-end, we design an efficient instance segmentation and comprehension pipeline that enhances language reasoning through encoding captions. In the back-end, we implement probabilistic instance voxels and formulate the cross-frame incremental fusion process into two subtasks: instance association and live map evolution, ensuring robustness to sensor and segmentation noise. Extensive evaluations across multiple datasets demonstrate that OpenVox achieves state-of-the-art performance in zero-shot instance segmentation, semantic segmentation, and open-vocabulary retrieval. The project page of OpenVox is available at https://open-vox.github.io/. Yinan Deng, Bicheng Yao, Yihang Tang, Tianxing Zhou, Yi Yang 0009, Yufeng Yue |
IROS | 5 |
| 2025 | RoadsideSplat: Robust 3D Gaussian Reconstruction from Monocular Roadside SurveillanceabstractReconstructing dynamic roads from roadside traffic surveillance cameras is crucial for smart cities and digital twin applications. While the latest monocular depth estimation methods demonstrate strong performance, they exhibit instability in roadside scenarios. Existing reconstruction approaches for autonomous driving scenes predominantly adopt vehicle-mounted perspectives, accumulating vehicle point clouds from per-frame depth maps using 3D bounding boxes. These point clouds are used to initialize the center positions and colors of 3D Gaussians to improve reconstruction performance. However, the compressed depth discrepancy between vehicles and road surfaces in roadside views leads to model confusion between vehicle and background depth estimations. To address these challenges, we propose a robust reconstruction framework based on a single fixed RGB traffic camera. Differing from conventional frame-wise depth prediction followed by 3D box-based accumulation, our method processes masked vehicle fore-ground sequences through existing models, directly predicting complete vehicle point clouds via local feature matching and global alignment while iteratively refining 3D boxes to enhance reconstruction quality. Leveraging the explicit nature of 3D Gaussians for scene editing, we introduce simple yet effective road constraints to mitigate penetration artifacts during scene manipulation. Extensive evaluations on the TUMTraf-V2X and RCooper datasets under monocular roadside settings validate the effectiveness of our approach. Zhaoxiang Liang, Wenjun Guo, Bohan Ren, Yi Yang 0009 |
IROS | 4 |
| 2025 | Automated 3D-GS Registration and Fusion via Skeleton Alignment and Gaussian-Adaptive FeaturesabstractIn recent years, 3D Gaussian Splatting (3D-GS)based scene representation demonstrates significant potential in real-time rendering and training efficiency. However, most existing methods primarily focus on single-map reconstruction, while the registration and fusion of multiple 3D-GS submaps remain underexplored. Existing methods typically rely on manual intervention to select a reference sub-map as a template and use point cloud matching for registration. Moreover, hard-threshold filtering of 3D-GS primitives often degrades rendering quality after fusion. In this paper, we present a novel approach for automated 3D-GS sub-map alignment and fusion, eliminating the need for manual intervention while enhancing registration accuracy and fusion quality. First, we extract geometric skeletons across multiple scenes and leverage ellipsoid-aware convolution to capture 3D-GS attributes, facilitating robust scene registration. Second, we introduce a multi-factor Gaussian fusion strategy to mitigate the scene element loss caused by rigid thresholding. Experiments on the ScanNet-GSReg and our Coord datasets demonstrate the effectiveness of the proposed method in registration and fusion. For registration, it achieves a 41.9% reduction in RRE on complex scenes, ensuring more precise pose estimation. For fusion, it improves PSNR by 10.11 dB, highlighting superior structural preservation. These results confirm its ability to enhance scene alignment and reconstruction fidelity, ensuring more consistent and accurate 3D scene representation for robotic perception and autonomous navigation. Shiyang Liu, Dianyi Yang, Yu Gao 0040, Bohan Ren, Yi Yang 0009, Mengyin Fu |
IROS | 5 |
| 2025 | GaussianGraph: 3D Gaussian-Based Scene Graph Generation for Open-World Scene UnderstandingabstractRecent advancements in 3D Gaussian Splatting(3DGS) have significantly improved semantic scene understanding, enabling natural language queries to localize objects within a scene. However, existing methods primarily focus on embedding compressed CLIP features to 3D Gaussians, suffering from low object segmentation accuracy and lack spatial reasoning capabilities. To address these limitations, we propose GaussianGraph, a novel framework that enhances 3DGS-based scene understanding by integrating adaptive semantic clustering and scene graph generation. We introduce a ‘Control-Follow’ clustering strategy, which dynamically adapts to scene scale and feature distribution, avoiding feature compression and significantly improving segmentation accuracy. Additionally, we enrich scene representation by integrating object attributes and spatial relations extracted from 2D foundation models. To address inaccuracies in spatial relationships, we propose 3D correction modules that filter implausible relations through spatial consistency verification, ensuring reliable scene graph construction. Extensive experiments on three datasets demonstrate that GaussianGraph outperforms state-of-the-art methods in both semantic segmentation and object grounding tasks, providing a robust solution for complex scene understanding and interaction. We provide supplementary video and code at https://wangxihan-bit.github.io/GaussianGraph. Xihan Wang, Dianyi Yang, Yu Gao 0040, Yufeng Yue, Yi Yang 0009, Mengyin Fu |
IROS | 5 |
| 2025 | Vehicle Drifting Planning and Control Framework for Flexible U-turns in Space-limited EnvironmentsabstractSpace-limited U-shape bend is a safety-critical scenario that requires the high maneuverability of vehicles. However, due to the non-holonomic nature of the vehicle, it is difficult to perform flexible U-turns without intricate adjustments, which is detrimental to the efficient execution of tasks. To address these issues, this work incorporates the drifting maneuver of the vehicle and proposes a planning and control framework for time-space efficient passing in constrained U-shape bends. First, a dual-track, 3-Dof vehicle model is developed, incorporating load transfer effects and nonlinear tire forces to enhance trajectory precision. Based on this model, a nonlinear optimization-based planner generates time-optimal, space-efficient, and drift-compatible trajectories while ensuring dynamic feasibility. Finally, a multilayer controller is designed for precise trajectory tracking, integrating a trajectory error feedback compensator, a dynamic state feedforward-feedback regulator, and a model inversion-based actuator controller. Simulation experiments in CarSim validate the proposed framework, demonstrating significant improvements in spatial efficiency and completion time. The results highlight its effectiveness in enhancing autonomous vehicle maneuverability for high-performance applications in constrained environments. Shuaicong Yang, Yi Yang 0009, Ting Zhang 0014, Mengyin Fu |
IROS | 3 |
| 2025 | OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level UnderstandingabstractRecent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by rigid offline pipelines and the inability to provide precise 3D object-level understanding given open-ended queries. In this paper, we present OpenGS-Fusion, an innovative open-vocabulary dense mapping framework that improves semantic modeling and refines object-level understanding. OpenGS-Fusion combines 3D Gaussian representation with a Truncated Signed Distance Field to facilitate lossless fusion of semantic features on-the-fly. Furthermore, we introduce a novel multimodal language-guided approach named MLLM-Assisted Adaptive Thresholding, which refines the segmentation of 3D objects by adaptively adjusting similarity thresholds, achieving an improvement 17% in 3D mIoU compared to the fixed threshold strategy. Extensive experiments demonstrate that our method outperforms existing methods in 3D object understanding and scene reconstruction quality, as well as showcasing its effectiveness in language-guided scene interaction. The code is available at https://young-bit.github.io/opengs-fusion.github.io/. Dianyi Yang, Xihan Wang, Yu Gao 0040, Shiyang Liu, Bohan Ren, Yufeng Yue, Yi Yang 0009 |
IROS | 7 |
| 2025 | TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation ModelsabstractReliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when categories share similar visual characteristics. 2) While SAM excels in instance-level segmentation, integrating it with thermal images and text is hindered by modality heterogeneity and computational inefficiency. To address these, we propose TASeg, a text-aware RGB-T segmentation framework by using Low-Rank Adaptation (LoRA) fine-tuning technology to adapt vision foundation models. Specifically, we propose a Dynamic Feature Fusion Module (DFFM) in the image encoder, which effectively merges features from multiple visual modalities while freezing SAM’s original transformer blocks. Additionally, we incorporate CLIP-generated text embeddings in the mask decoder to enable semantic alignment, which further rectifies the classification error and improves the semantic understanding accuracy. Experimental results across diverse datasets demonstrate that our method achieves superior performance in challenging scenarios with fewer trainable parameters. Te Cui, Qitong Chu, Wenjie Song 0001, Yi Yang 0009, Yufeng Yue |
IROS | 5 |
| 2025 | Motion Control of a Hybrid Self-Reconfigurable Wheel-Legged Dual-Arm RobotabstractCurrent wheeled bipedal robots face significant mobility challenges when traversing discontinuous terrain such as gaps and step-like obstacles, and suffer from substantial dynamic inefficiencies. This paper presents a hybrid self-reconfigurable wheel-legged dual-arm robot equipped with an active docking mechanism, enabling transitions between wheeled bipedal and multi-wheel-legged configurations. Based on a self-developed robotic platform, this work addresses key control challenges in articulated multi-wheel-legged mode and proposes a novel distributed operation paradigm for wheeled bipedal robots. Each module utilizes its manipulators for stable grasping of elevated objects and collaborative tasks, while the multi-unit system achieves efficient, high-load, and stable locomotion. To manage the control complexities in multimodal operation, we develop a unified modular control architecture integrating Virtual Model Control (VMC) and Linear Quadratic Regulator (LQR). For the articulated multi-wheel-legged mode, a body-posture controller regulates global body configuration, and a turning controller adjusts the wheelbase and roll angle via distributed actuation to manage the passive degrees of freedom (DoF) at the articulation points. Experimental validation using a physical prototype confirms the effectiveness and practicality of the proposed approach. Hong Du, Peng Qiu, Yi Yang 0009, Wenjie Song 0001 |
IROS | 4 |
| 2025 | STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task plannerabstractThe ability to perform reliable long-horizon task planning is crucial for deploying robots in real-world environments. However, directly employing Large Language Models (LLMs) as action sequence generators often results in low success rates due to their limited reasoning ability for long-horizon embodied tasks. In the STEP framework, we construct a subgoal tree through a pair of closed-loop models: a subgoal decomposition model and a leaf node termination model. Within this framework, we develop a hierarchical tree structure that spans from coarse to fine resolutions. The subgoal decomposition model leverages a foundation LLM to break down complex goals into manageable subgoals, thereby spanning the subgoal tree. The leaf node termination model provides real-time feedback based on environmental states, determining when to terminate the tree spanning and ensuring each leaf node can be directly converted into a primitive action. Experiments conducted in both the VirtualHome WAH-NL benchmark and on real robots demonstrate that STEP achieves long-horizon embodied task completion with success rates up to 34% (WAH-NL) and 25% (real robot) outperforming SOTA methods. Tianxing Zhou, Haojia Ao, Guangyan Chen, Boyang Xing, Cheng Jingwen, Yi Yang 0009, Yufeng Yue |
IROS | 7 |
| 2025 | Unifying Latent Action and Latent State Pre-training for Policy Learning from VideosabstractVideo data provides an accessible and rich source beyond expensive action-labeled robot data for advancing robotic learning paradigms. Motivated by this potential, researchers investigate methods to exploit video data in robotic learning. Recent approaches can be primarily divided into two categories: Action-based approaches tokenize latent actions from videos for policy pre-training. State-based approaches pre-train models to predict subsequent states. The former establishes rich motion priors, while the latter empowers the robot to anticipate future events. These complementary capabilities suggest significant potential for integration into a unified framework. In this paper, we propose UniMimic, a novel approach unifying latent action and latent state pre-training from videos. We first train a unified tokenizer to learn latent states from video frames while deriving latent actions between state tokens. Subsequently, the policy is pre-trained on videos to predict these latent actions and subsequent latent states. Finally, the policy is fine-tuned on an action-labeled robot dataset to transfer the learned priors to precise robot execution. Experiments exhibit that our pre-training stage enhances the performance by 19% in the Libero benchmark and improves the average number of tasks completed in a row of 5 from 2.50 and 2.35 to 3.89 and 3.73 in the CALVIN benchmark. In the real-world experiments, our method still delivers improvements exceeding 36%. Guangyan Chen, Meiling Wang 0002, Te Cui, Luojie Yang, Lin Zhao 0016, Yi Yang 0009, Yufeng Yue |
SIGGRAPH Asia | 9 |
| 2025 | Point Tree Transformer for Point Cloud RegistrationabstractPoint cloud registration is a fundamental task in the fields of computer vision and robotics. Recent advancements in transformer-based methods have demonstrated enhanced performance in this domain. However, the standard attention mechanisms employed in these approaches tend to incorporate numerous points of low relevance, and therefore struggle to focus their attention weights on sparse yet meaningful points. This inefficiency leads to limited local structure modeling capabilities and quadratic computational complexity. To overcome these limitations, we propose the Point Tree Transformer (PTT), a novel transformer-based approach for point cloud registration that efficiently extracts comprehensive local and global features while maintaining linear computational complexity. The PTT constructs hierarchical feature trees from point clouds in a coarse-to-dense manner, and introduces a novel Point Tree Attention (PTA) mechanism. This mechanism adheres to the tree structure to facilitate the progressive convergence of attended regions toward salient points. Specifically, each tree layer selectively identifies a subset of relevant points with the highest attention scores, and subsequent layers focus attention on areas of significant relevance, derived from the child points of the selected point set. The feature extraction process additionally incorporates coarse point features that capture high-level semantic information, thus facilitating local structure modeling and the progressive integration of multiscale information. Consequently, the PTA enables the model to focus on essential local structures and extract intricate local information while maintaining linear computational complexity. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate that our method outperforms state-of-the-art methods in terms of performance. The code for our method is publicly available at https://github.com/CGuangyan-BIT/PTT. Meiling Wang 0002, Guangyan Chen, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Focus-TransUnet3D: High-Precision Model for 3D Segmentation of Medical Point TargetsabstractDeep learning has been extensively applied in medical image segmentation, providing significant support for disease diagnosis. However, traditional encoder-decoder networks struggle with segmenting scale-sensitive point target lesions. To address this challenge, this paper proposes an innovative incremental fusion architecture that can integrate different models and achieve significant performance improvements through complementary fusion. Based on this architecture, we developed Focus-TransUnet3D by combining the Trans-FusionNet3D model and the 3D Unet model. This model adopts a global-to-local segmentation strategy, effectively addressing the challenges of medical point target segmentation, thereby expanding the application of deep learning in the field of medical image processing. Furthermore, we design a deep fusion strategy suitable for the transformer model to adapt to multi-scale feature learning. The integration of the transformer model with convolutional neural networks brings improvements in local and global feature extraction capabilities, enhancing the applicability of our model. We evaluate our model on three clinical datasets with different target scales: the Intracranial Artery dataset, the Intracranial Aneurysm dataset, and the LiTS17 dataset. The results indicate that in the external test for intracranial aneurysm auxiliary diagnosis, the model trained with only 47 annotated samples achieved the state-of-the-art performance, attaining a Dice coefficient of 84.14% and a sensitivity of 100%. This effectively addresses the challenges of annotation scarcity and tiny targets. Our code will be released athttps://github.com/caijilia/FTUnet3D. Dihua Zhai, Hao Li 0075, Ke Tian, Yi Yang 0009, Zhenyao Chang, Shuo Wang 0001, Yuanqing Xia |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | OmniMap: A General Mapping Framework Integrating Optics, Geometry, and SemanticsabstractRobotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of these requirements while exhibiting optical blurring, geometric irregularities, and semantic ambiguities. To address these challenges, we propose OmniMap. Overall, OmniMap represents the first online mapping framework that simultaneously captures optical, geometric, and semantic scene attributes while maintaining real-time performance and model compactness. At the architectural level, OmniMap employs a tightly coupled 3DGS-Voxel hybrid representation that combines fine-grained modeling with structural stability. At the implementation level, OmniMap identifies key challenges across different modalities and introduces several innovations: adaptive camera modeling for motion blur and exposure compensation, hybrid incremental representation with normal constraints, and probabilistic fusion for robust instance-level understanding. Extensive experiments show OmniMap's superior performance in rendering fidelity, geometric accuracy, and zero-shot semantic segmentation compared to state-of-the-art methods across diverse scenes. The framework's versatility is further evidenced through a variety of downstream applications, including multi-domain scene Q&A, interactive editing, perception-guided manipulation, and map-assisted navigation. Yinan Deng, Yufeng Yue, Jianyu Dou, Yi Yang 0009, Mengyin Fu |
IEEE Trans. Robotics | 7 |
| 2025 | MC-NeRF: Multi-Camera Neural Radiance Fields for Multi-Camera Image Acquisition SystemsabstractNeural Radiance Fields (NeRF) use multi-view images for 3D scene representation, demonstrating remarkable performance. As one of the primary sources of multi-view images, multi-camera systems encounter challenges such as varying intrinsic parameters and frequent pose changes. Most previous NeRF-based methods assume a unique camera and rarely consider multi-camera scenarios. Besides, some NeRF methods that can optimize intrinsic and extrinsic parameters still remain susceptible to suboptimal solutions when these parameters are poor initialized. In this paper, we propose MC-NeRF, a method for joint optimization of both intrinsic and extrinsic parameters alongside NeRF, allowing individual camera parameters for each image. First, we analyze the coupling issue that arises from the joint optimization between intrinsics and extrinsics, and propose a decoupling constraint utilizing auxiliary images. To further address the degenerate cases in the decoupling process, we introduce an efficient auxiliary image acquisition scheme to mitigate these effects. Furthermore, recognizing that most existing datasets are designed for a unique camera, we provided a new dataset that includes both simulated data and real-world data. Experiments demonstrate the effectiveness of our method in scenarios where each image corresponds to different camera parameters. Specifically, our approach outperforms the baselines favorably in terms of intrinsics estimation, extrinsics estimation, scale estimation, and rendering quality. Yu Gao 0040, Lutong Su, Hao Liang 0016, Yufeng Yue, Yi Yang 0009, Mengyin Fu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Fast and Robust Point Cloud Registration with Tree-based TransformerabstractPoint cloud registration is essential in computer vision and robotics. Recently, transformer-based methods have achieved advanced point cloud registration performance. However, the standard attention mechanism utilized in these methods considers many low-relevance points, and it has difficulty focusing its attention weights on sparse and meaningful points, leading to limited local structure modeling capabilities and quadratic computational complexity. To address these limitations, we present the Tree-based Transformer (TrT), which is able to extract abundant local and global features with linear computational complexity. Specifically, the TrT builds coarse-to-dense feature trees, and a novel Tree-based Attention (TrA) is proposed to guide the progressive convergence of the attended regions toward meaningful points and to structurize point clouds following tree structures. In each layer, the top ${\mathcal{S}}$ key points with the highest attention scores are selected, such that in the next layer, attention is evaluated only within the specified high-relevance regions, corresponding to the child points of these selected ${\mathcal{S}}$ points. Additionally, coarse features containing high-level semantic information are incorporated into the child points to guide the feature extraction process, facilitating local structure modeling and multiscale information integration. Consequently, TrA enables the model to focus on critical local structures and extract rich local information with linear computational complexity. Experiments demonstrate that our method achieves state-of-the-art performance on 3DMatch and KITTI benchmarks. The code for our method is publicly available at https://github.com/CGuangyan-BIT/TrT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
ICRA | 3 |
| 2024 | Risk-Inspired Aerial Active Exploration for Enhancing Autonomous Driving of UGV in Unknown Off-Road EnvironmentsabstractUnknown area exploration is a crucial but challenging task for autonomous driving of unmanned ground vehicles (UGV) in unknown off-road environments. However, the exploration efficiency of a single UGV is low due to its limited sensing range. To solve this problem, this paper proposes a risk-inspired aerial active exploration system, which utilizes the flexibility and field of view advantages of Unmanned Aerial Vehicles (UAV) to guide the UGV in unknown off-road environments. Firstly, a fast terrain risk mapping method that can be used for both UAV and UGV is developed. This method efficiently combines quadtree and hash table data structure to enable UAV to analyze large scale terrain point cloud in real time. Based on the risk mapping result, a risk-inspired active exploration method is proposed to actively search a safe reference path for the UGV, which introduces terrain risk information into the process of travel point selection. Finally, the reference path is gradually generated and optimized, so that the UGV can safely and smoothly follow the path to the target location. Compared with single UGV exploration system, our approach reduces the overall path risk by 26.8% in simulated experiments, showing that the proposed system can enhance autonomous driving of the UGV and help it effectively avoid high-risk areas in unknown off-road environments. Rongchuan Wang, Mengyin Fu, Yi Yang 0009, Wenjie Song 0001 |
ICRA | 4 |
| 2024 | DSVT: Dynamic 3D Surround View for Tractor-Trailer Vehicles Based on Real-Time Pose Estimation with Drop ModelabstractIn recent years, 3D surround view systems have attracted a lot of attention in the field of advanced driver assistance systems (ADAS). However, the foundational assumption of unchanging camera poses in traditional 3D surround view systems, which is designed for single-unit vehicles, results in a failure to manage the non-rigid connections characteristic of tractor-trailer vehicles. Moreover, tractor-trailer vehicles have the feature of long bodies and large wheelbases, leading to severe distortions and abrupt changes in the rendering results of previous 3D texture mapping models. In this paper, we propose DSVT, a dynamic 3D surround view system for tractor-trailer vehicles, designed to address the aforementioned issues. Specifically, we develop a dynamic surround image stitching algorithm based on relative pose estimation, which estimates the relative poses between cameras and stitches all images together to generate a 2D panoramic image. Subsequently, a novel 3D drop model is proposed, mapping the 2D panoramic image onto the 3D model for panoramic viewing. Our system can run in real time on Nvidia AGX Orin. Experimental results in real tractor-trailer scenes show that our system can achieve more accurate and natural visual effects. Mengyin Fu, Hao Liang 0016, Chunhui Zhu, Yi Yang 0009 |
IROS | 5 |
| 2024 | Self-supervised Monocular Depth Estimation in Challenging Environments Based on Illumination Compensation PoseNetabstractSelf-supervised depth estimation has attracted much attention due to its ability to improve the 3D perception capabilities of unmanned systems. However, existing unsupervised frameworks rely on the assumption of photometric consistency, which may not hold in challenging environments such as night-time, rainy nights, or snowy winters due to complex lighting and reflections, resulting in inconsistent photometry across different frames for the same pixel. To address this problem, we propose a self-supervised monocular depth estimation unified framework that can handle these complex scenarios, which has the following characteristics: (1) an Illumination Compensation PoseNet (ICP) is designed, which is based on the classic Phong illumination theory and compensates for lighting changes in adjacent frames by estimating per-pixel transformations; (2) a Dual-Axis Transformer (DAT) block is proposed as the backbone network of the depth encoder, which infers the depth of local repeat-texture areas through spatial-channel dual-dimensional global context information of images. Experimental results demonstrate that our approach achieves state-of-the-art depth estimation results in complex environments on the challenging Oxford RobotCar dataset. Shengyu Hou, Wenjie Song 0001, Rongchuan Wang, Meiling Wang 0002, Yi Yang 0009, Mengyin Fu |
IROS | 5 |
| 2024 | Robust Multi-Camera BEV Perception: An Image-Perceptive Approach to Counter Imprecise Camera CalibrationabstractRecently, Bird’s Eye View (BEV) detection methodologies that utilize surround-view cameras have seen significant advancements in autonomous driving systems. Traditional methods, however, are constrained by their reliance on specific camera parameters, which poses challenges in generalizing across different vehicle-mounted cameras with varying poses and under adverse conditions. To address these challenges, we propose a robust BEV representation network that integrates Dual-Space Positional Encoding (DSPE) and image perception. This network is designed to enhance resilience to calibration errors and pose fluctuations, resulting in reliable detection performance on the Nuscenes dataset, even with imprecise extrinsic inputs. Our approach demonstrates competitive accuracy when compared to other methods that do not rely on temporal data, highlighting the effectiveness of our DSPE strategy in improving the robustness and accuracy of BEV detection in dynamic and challenging environments. Rundong Sun, Mengyin Fu, Hao Liang 0016, Chunhui Zhu, Yi Yang 0009 |
IROS | 6 |
| 2024 | Fine-tuning the Diffusion Model and Distilling Informative Priors for Sparse-view 3D Reconstructionabstract3D reconstruction methods such as Neural Radiance Fields (NeRFs) are capable of optimizing high-quality 3D representation from images. However, NeRF is limited by the requirement for a large number of multi-view images, making its application to real-world scenarios challenging. In this work, we propose a method that can reconstruct real-world scenes from a few input images and a simple text prompt. Specifically, we fine-tune a pretrained diffusion model to constrain its powerful priors to the visual inputs and generate 3D-aware images, leveraging the coarse renderings obtained from input images as the image condition, along with the text prompt as the text condition. Our fine-tuning method saves a significant amount of training time and GPU memory usage while also generating credible results. Moreover, to enable our method to have self-evaluation capabilities, we design a semantic switch to filter out generated images that do not match real scenes, ensuring that only informative priors from the fine-tuned diffusion model are distilled into the 3D model. The semantic switch we designed can be used as a plug-in and improve performance by 13%. We perform our approach on a real-world dataset and demonstrate competitive results compared to existing sparse-view 3D reconstruction methods. Please see our project page for more visualizations and code: https://bityia.github.io/FDfusion. Jiadong Tang, Yu Gao 0040, Tianji Jiang, Yi Yang 0009, Mengyin Fu |
IROS | 4 |
| 2024 | LCP-Fusion: A Neural Implicit SLAM with Enhanced Local Constraints and Computable PriorabstractRecently the dense Simultaneous Localization and Mapping (SLAM) based on neural implicit representation has shown impressive progress in hole filling and high-fidelity mapping. Nevertheless, existing methods either heavily rely on known scene bounds or suffer inconsistent reconstruction due to drift in potential loop-closure regions, or both, which can be attributed to the inflexible representation and lack of local constraints. In this paper, we present LCP-Fusion, a neural implicit SLAM system with enhanced local constraints and computable prior, which takes the sparse voxel octree structure containing feature grids and SDF priors as hybrid scene representation, enabling the scalability and robustness during mapping and tracking. To enhance the local constraints, we propose a novel sliding window selection strategy based on visual overlap to address the loop-closure, and a practical warping loss to constrain relative poses. Moreover, we estimate SDF priors as coarse initialization for implicit features, which brings additional explicit constraints and robustness, especially when a light but efficient adaptive early ending is adopted. Experiments demonstrate that our method achieve better localization accuracy and reconstruction consistency than existing RGB-D implicit SLAM, especially in challenging real scenes (ScanNet) as well as self-captured scenes with unknown scene bounds. The code is available at https://github.com/laliwang/LCP-Fusion. Yinan Deng, Yi Yang 0009, Yufeng Yue |
IROS | 3 |
| 2024 | VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsabstractVisual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by more than 37% in long-horizon tasks. Code and videos are available on our anonymous homepage. Guangyan Chen, Meiling Wang 0002, Te Cui, Yao Mu 0001, Tianxing Zhou, Zicai Peng, Mengxiao Hu, Haizhou Li 0004, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
NeurIPS | 11 |
| 2024 | Self-Supervised Monocular Depth Estimation for All-Day Images Based on Dual-Axis TransformerabstractAll-day self-supervised monocular depth estimation has strong practical significance for autonomous systems to continuously perceive the 3D information of the world. However, night-time scenes pose challenges of weak texture and violating the brightness consistency assumption due to low illumination and varying lighting, respectively, which easily leads to most existing self-supervised models only being able to handle day-time scenes. To address this problem, we propose a self-supervised monocular depth estimation unified framework that can handle all-day scenarios, which has three features: (1) an Illumination Compensation PoseNet (ICP) is designed, which is based on the classic Phong illumination theory and compensates for lighting changes in adjacent frames by estimating per-pixel transformations; (2) a Dual-Axis Transformer (DAT) block is proposed as the backbone network of the depth encoder, which infers the depth of local low-illumination areas through spatial-channel dual-dimensional global context information of night-time images; (3) a cross-layer Adaptive Fusion Module (AFM) is introduced between multiple DAT blocks, which learns attention weights between different layer features and adaptively fuses cross-layer features using the learned weights, enhancing the complementarity of different layer features. This work was evaluated on multiple datasets, including: RobotCar, Waymo and KITTI datasets, achieving state-of-the-art results in both day-time and night-time scenarios. Shengyu Hou, Mengyin Fu, Rongchuan Wang, Yi Yang 0009, Wenjie Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A Cognition-Inspired Human-Like Decision-Making Method for Automated VehiclesabstractDrivers’ cognitive mechanisms could benefit the development of human-like automated driving (AD) strategies, which are with high intelligence and comfort levels. The common approach of human-like AD is to learn from human demonstration data, for which it is exhausting and difficult to construct well-rounded and reliable datasets. Therefore, we proposed the human-like AD decision-making method based on drivers’ cognition mechanism. The fundamental and difficult thing of this method is to figure out drivers’ cognitive mechanism systematically and comprehensively, which is still either too rough or too fragmented for AD development. By integrating the abundant studies about drivers’ cognition in multiple fields, we propose two novel conceptual models: Potential Hazard Model (PHM) illustrates the mechanisms of drivers’ reaction in the simple meta-scenarios while Candidate Selection Model (CSM) explains how drivers handle complicated scenarios based on PHM. Based on PHM and CSM, we propose the human-like decision-making method for AD. This method integrates cognitive mechanisms, natural driving data, optimization-based planning, and other techniques profoundly. The experiments in extensive road and traffic scenarios verify that the method show good generalizability, interpretability, and human-likeness. Yi Yang 0009, Mengyin Fu, Jingyue Zheng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Dynamic Voxels Based on Ego-Conditioned Prediction: An Integrated Spatio-Temporal Framework for Motion PlanningabstractPrediction is a vital component of motion planning for autonomous vehicles (AVs). By reasoning about the possible behavior of other target agents, the ego vehicle (EV) can navigate safely, efficiently, and politely. However, most of the existing work overlooks the interdependencies of the prediction and planning module, only connecting them in a sequential pipeline or underexploring the prediction results in the planning module. In this work, we propose a framework that integrates the prediction and planning module with three highlights. First, we propose an ego-conditioned model for causal prediction, with the introduced edge-featured graph transformer model, the impact the ego future maneuver poses to the target vehicles is demonstrated. Second, we develop a motion planner based on ‘dynamic voxels’ in the spatio-temporal domain, enabling the time-to-collision criterion evaluation and the optimal trajectory generation in continuous space. Third, the prediction and planning modules are coupled in a closed-loop and efficient form. Specifically, taking each maneuver as a cluster, representative trajectory primitives are generated for conditional prediction, and conversely, prediction results are used to score the primitives as guidance, which alleviates the duplicated callback of the prediction module. The simulations are conducted in overtaking, merging, unprotected left turns, and also scenarios with imperfect social behaviors. The comparison studies demonstrate the better safety assurance and efficiency of the proposed model, and the ablation experiments further reveal the effectiveness of the new ideas. Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Alexandre Alahi |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Rethinking Point Cloud Registration as Masking and ReconstructionabstractPoint cloud registration is essential in computer vision and robotics. In this paper, a critical observation is made that the invisible parts of each point cloud can be directly utilized as inherent masks, and the aligned point cloud pair can be regarded as the reconstruction target. Motivated by this observation, we rethink the point cloud registration problem as a masking and reconstruction task. To this end, a generic and concise auxiliary training network, the Masked Reconstruction Auxiliary Network (MRA), is proposed. The MRA reconstructs the complete point cloud by separately using the encoded features of each point cloud obtained from the backbone, guiding the contextual features in the backbone to capture fine-grained geometric details and the overall structures of point cloud pairs. Unlike recently developed high-performing methods that incorporate specific encoding methods into transformer models, which sacrifice versatility and introduce significant computational complexity during the inference process, our MRA can be easily inserted into other methods to further improve registration accuracy. Additionally, the MRA is detached after training, thereby avoiding extra computational complexity during the inference process. Building upon the MRA, we present a novel transformer-based method, the Masked Reconstruction Transformer (MRT), which achieves both precise and efficient alignment using standard transformers. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate the superior performance of our MRT over state-of-the-art methods. Codes are available at https://github.com/CGuangyan-BIT/MRA. Guangyan Chen, Meiling Wang 0002, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
ICCV | 4 |
| 2023 | Conflict-constrained Multi-agent Reinforcement Learning Method for Parking Trajectory PlanningabstractAutomated Valet Parking (AVP) has been exten-sively researched as an important application of autonomous driving. Considering the high dynamics and density of real parking lots, a system that considers multiple vehicles simultaneously is more robust and efficient than a single vehicle setting as in most studies. In this paper, we propose a dis-tributed Multi-agent Reinforcement Learning(MARL) method for coordinating multiple vehicles in the framework of an AVP system. This method utilizes traditional trajectory planning to accelerate the learning process and introduces collision conflict constraints for policy optimization to mitigate the path conflict problem. In contrast to other centralized multi-agent path finding methods, the proposed approach is scalable, distributed, and adapts to dynamic stochastic scenarios. We train the models in random scenarios and validate in several artificially designed complex parking scenarios where vehicles are always disturbed by dynamic and static obstacles. Experimental results show that our approach mitigates path conflicts and excels in terms of success rate and efficiency. Meiling Wang 0002, Yi Yang 0009, Wenjie Song 0001 |
ICRA | 3 |
| 2023 | Multi-View Robust Collaborative Localization in High Outlier Ratio Scenes Based on Semantic FeaturesabstractFiltering out outlier data associations between local maps can improve the robustness and accuracy of multi-robot localization. When the overlap is low and the field of view difference is large, it is likely to produce outlier data associations between local maps, which will reduce the matching accuracy and even lead to the failure of collaborative localization. To solve this problem, this paper proposes a novel outdoor robust collaborative localization algorithm (HORCL) capable for high outlier ratio scenes. The Mixture Probability Model (MPM) and the Hierarchical EM (Expectation Maximization) algorithm in HORCL are applied to screen two levels of outliers (loop closure constraints and point pairs) and improve localization performance. Specifically, the inlier probabilities of data associations are calculated in MPM to identify outliers by considering geometric distances, semantic consistency, and spatial consistency. Then, outlier loop closures and outlier point pairs in inlier constraints are filtered by applying the Hierarchical EM algorithm, thereby relieving the adverse effect of outliers on localization accuracy. The proposed algorithm is validated on public datasets and compared with the latest methods, demonstrating the improvement in localization accuracy and robustness. The code is available at https://github.com/BIT-TYJ/HORCL. Meiling Wang 0002, Yinan Deng, Yi Yang 0009, Ziquan Lan, Yufeng Yue |
IROS | 4 |
| 2023 | SSGM: Spatial Semantic Graph Matching for Loop Closure Detection in Indoor EnvironmentsabstractCapturing the semantics of objects and the topological relationship allows the robot to describe the scene more intelligently like a human and measure the similarity between scenes (loop closure detection) more accurately. However, many current semantic graph matching methods are based on walk descriptors, which only extract adjacency relations between objects. In such way, the comprehensive information in the semantic graph is not fully exploited, which may lead to false closed-loop detection. This paper proposes a novel spatial semantic graph matching method (SSGM) in indoor environments, which considers multifaceted information of the semantic graphs. Firstly, two semantic graphs are aligned in the same coordinate space contributed by the second-order spatial compatibility metric between objects and local graph features of objects in semantic graphs. Secondly, the similarity of the spatial distribution of overall semantic graphs is further evaluated. The proposed algorithm is validated on public datasets and compared with the latest semantic graph matching methods, demonstrating improved accuracy and efficiency in loop closure detection. The code is available at https://github.com/BIT-TYJ/SSGM. Meiling Wang 0002, Yinan Deng, Yi Yang 0009, Yufeng Yue |
IROS | 4 |
| 2023 | UVSS: Unified Video Stabilization and Stitching for Surround View of Tractor-Trailer VehiclesabstractAutomotive surround-view camera systems have been commonly employed in automated driving to aid in near-field sensing and other perception tasks. Due to the large size of the body and the presence of multiple blind spots, panoramic surround-view systems are particularly crucial for tractor-trailer vehicles. However, the non-rigid body of tractor-trailer vehicles introduces pose changes between cameras, rendering traditional calibration-based methods inadequate. Additionally, cameras mounted separately on the tractor and the trailer will experience independent vibrations, resulting in undesirable shakiness in captured videos. In this paper, we propose a unified video stabilization and stitching method to address these challenges, which can smooth the unsteady frames and align the images from moving cameras. Delving into video stabilization techniques, we extend mesh-based motion model for unified stitching and leverage deep-learning based modules to handle complex real-world scenarios. Moreover, we design a new optimization framework to estimate the optimal displacements of mesh vertices, enabling simultaneous stabilization and stitching of frames. The experimental results, obtained by public datasets and videos captured from a model tractor-trailer vehicle, demonstrate that our approach outperforms previous methods and is highly effective in real-world applications. Chunhui Zhu, Yi Yang 0009, Hao Liang 0016, Mengyin Fu |
IROS | 2 |
| 2023 | PointGPT: Auto-regressively Generative Pre-training from Point CloudsabstractLarge language models (LLMs) based on the generative pre-training transformer (GPT) have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, addressing the challenges associated with disorder properties, low information density, and task gaps. Specifically, a point cloud auto-regressive generation task is proposed to pre-train transformer models. Our method partitions the input point cloud into multiple point patches and arranges them in an ordered sequence based on their spatial proximity. Then, an extractor-generator based transformer decode, with a dual masking strategy, learns latent representations conditioned on the preceding point patches, aiming to predict the next one in an auto-regressive manner. To explore scalability and enhance performance, a larger pre-training dataset is collected. Additionally, a subsequent post-pre-training stage is introduced, incorporating a labeled hybrid dataset. Our scalable approach allows for learning high-capacity models that generalize well, achieving state-of-the-art performance on various downstream tasks. In particular, our approach achieves classification accuracies of 94.9% on the ModelNet40 dataset and 93.4% on the ScanObjectNN dataset, outperforming all other transformer models. Furthermore, our method also attains new state-of-the-art accuracies on all four few-shot learning benchmarks. Codes are available at https://github.com/CGuangyan-BIT/PointGPT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
NeurIPS | 3 |
| 2022 | Aerial-Ground Robots Collaborative 3D Mapping in GNSS-Denied EnvironmentsabstractCollaborative heterogeneous robots are expected to perform comprehensive perception, mapping and coordination in search and rescue scenarios. The challenge of collaboration between heterogeneous robots lies in their huge differences in perception, mobility and processing capabilities. In this paper, a novel collaborative UAV-UGV mapping framework is proposed in GNSS-denied and unknown environments. The key novelty of this work is the proposing of a unified framework to formulate the UAV-UGV collaborative mapping problem with a continuous-discrete model, as well as its realization in real robotic systems. In order to project continuous space into discrete space, a novel information gain trigger scheme is pro-posed. The continuous space allows each robot to perform high frequency local map estimation, while discrete space describes the problem of multi-resolution hybrid map fusion. Considering the nature of data heterogeneity, a flexible probabilistic fusion algorithm is proposed that addresses the multi-resolution hybrid map fusion problem, where the local maps generated by UAV and UGV are fused based on Bayesian rule. The proposed UAV-UGV hybrid system is validated in various challenging scenarios, demonstrating its accuracy and utility in practical tasks. Yufeng Yue, Yuanzhe Wang, Yi Yang 0009, Danwei Wang |
ICRA | 4 |
| 2022 | HD-CCSOM: Hierarchical and Dense Collaborative Continuous Semantic Occupancy Mapping through Label DiffusionabstractThe collaborative operation of multiple robots can make up for the shortcomings of a single robot, such as limited field of perception or sensor failure. multirobots collaborative semantic mapping can enhance their comprehensive contextual understanding of the environment. However, existing multirobots collaborative semantic mapping algorithms mainly apply discrete occupancy map inference, and do not compensate for inconsistent labels of local maps caused by differences in robot perspectives, which leads to greatly reduced availability and accuracy of the final global map. To address the challenges of discontinuous maps and inconsistent semantic labels, this paper proposes a novel hierarchical and dense collaborative continuous semantic occupancy mapping algorithm (HD-CCSOM). This work decomposes and formulates robot collaborative continuous semantic occupancy mapping problem at two levels. At the single robot level, the multi-entropy kernel inference method smoothly processes the registered semantic point cloud and infers a local continuous semantic occupancy map for each robot. At the collaborative robots level, the local maps are fused into a global enhanced and consistent semantic map via the label diffusion method based on a graph model. The proposed algorithm has been validated on public datasets and in simulated and real scenes, demonstrating significant improvements in mapping accuracy and efficiency. Yinan Deng, Meiling Wang 0002, Yi Yang 0009, Yufeng Yue |
IROS | 3 |
| 2022 | Fisheye object detection based on standard image datasets with 24-points regression strategyabstractFisheye object detection is a difficult task in robotics and autonomous driving. One of the reasons is that the fisheye datasets are inferior to standard image datasets in scale and quantity, which inspires the idea of using standard image datasets for fisheye object detection. However, the models trained on standard image datasets do not perform well with fisheye data. In this work, we explore the effect of fisheye images on different stages of the YOLOX with published weights generated by standard image datasets. We also propose a new regression strategy for 24-points object representation method, which is insensitive to image distortion. The experiments show that the feature extraction part is robust to fisheye image features, while the regression part of location and category performs poorly. The strategy can achieve the position of discrete points without calculating the IOU of irregular-shaped boxes. Theoretically, the strategy can be widely adopted to regress the irregular bounding boxes composed of discrete points. Source code is at https://github.com/IN2-ViAUn/Exploration-of-Potential. Yu Gao 0040, Hao Liang 0016, Yi Yang 0009, Mengyin Fu |
IROS | 4 |
| 2022 | Trajectory Prediction-Based Local Spatio-Temporal Navigation Map for Autonomous Driving in Dynamic Highway EnvironmentsabstractAutonomous driving, including intelligent decision-making and path planning, in dynamic environments (like highway) is significantly more difficult than the navigation in static scenarios because of the additional time dimension. Therefore, correlating the time dimension and the space dimension through prediction to create a spatio-temporal navigation map can make decision-making and path planning in such kinds of environment much easier. In this article, NGSIM data is analysed and processed from the perspective of the ego-vehicle (using the data as an ego-vehicle’s perception results). Based on the data, we develop an LSTM (Long-Short Term Memory)-based framework to predict possible trajectories of multiple surrounding vehicles within a certain range of the ego-vehicle. Then, the multiple predicted trajectories in a series of continuous dynamic highway scenes are projected into a spatio-temporal domain to create an octree map. Thus, dynamic targets and static obstacles can be unified into the same domain or map so that the dynamic disturbance problem for autonomous driving in highway environments can be resolved. Experimental results show that the proposed model is capable of predicting all the future trajectories around the ego-vehicle efficiently and the corresponding spatio-temporal map can be generated accurately in different dynamic scenarios. Mengyin Fu, Ting Zhang 0014, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Action-State Joint Learning-Based Vehicle Taillight Recognition in Diverse Actual Traffic ScenesabstractAs the vital factor of vehicle behavior understanding and prediction, vehicle taillight recognition is an important technology for autonomous driving, especially in diverse actual traffic scenes full of dynamic interactive traffic participants. However, in practical application, it always faces many challenges, such as ‘variable lighting conditions’, ‘non-uniform taillight standards’ and ‘random relative observation pose’, which lead to few mature solutions in current common autopilot systems. This work proposes an action-state joint learning-based vehicle taillight recognition method on the basis of vehicles detection and tracking, which takes both taillight state features and time series features into account, consequently getting practicable results even in complex actual scenes. In detail, vehicle tracking sequence is used as input and split into pieces through a sliding window. Then, a CNN-LSTM model is applied to simultaneously identify the action features of brake lights and turn signals, dividing taillight actions into five categories: None, Brake_on, Brake_off, Left_turn, Right_turn. Next, the brightness of high-position brake light is extracted through semantic segmentation and combined with taillight actions to form higher-level features for taillight state sequence analysis. Finally, an undirected graph model is used to establish the long-term dependence between successive pieces by analysing the higher-level features, thus inferring the continuous taillight state into:$off$,$brake$,$left$,$right$. Datasets including daytime, nighttime, congested road, highway, etc. were collected, tested and published in our work to demonstrate its effectiveness and practicability. Wenjie Song 0001, Shixian Liu, Ting Zhang 0014, Yi Yang 0009, Mengyin Fu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Trajectory Planning Based on Spatio-Temporal Map With Collision Avoidance Guaranteed by Safety StripabstractTrajectory planning for the unmanned vehicle in the complex environment has always been a challenging task. Planned trajectory with the corresponding target velocity or acceleration sequence must be collision-free guaranteed and as comfortable as possible on the premise of obeying the traffic rules and interaction with other dynamic social vehicles. To meet this requirement, this paper proposes a framework for trajectory planning based on spatio-temporal map. Due to the time layer architecture in the map, the trajectory can be generated with velocity and acceleration simultaneously, and the whole trajectory is constrained within a ‘safety strip’, resulting in an efficient and safety guaranteed trajectory. The framework is composed of three sections: rough search, fine optimization and safety strip-based collision avoidance. For rough search, we propose an improved A* algorithm implemented in the discrete time layer to find out the suboptimal states efficiently. In fine optimization, the B-spline curve is exploited to connect the searched states into a continuous trajectory. And the optimal control points of B-spline are further grouped into several segments, forming the safety strip which is actually the distribution space of the planned trajectory. If necessary, an adjustment will be applied to keep the strip away from the collision zone, making the entire trajectory completely collision-free. Experiments on both public dataset and self-driving simulator show that the proposed framework can adapt to different kinds of complex traffic scenes well. Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | A Unified Framework Integrating Decision Making and Trajectory Planning Based on Spatio-Temporal Voxels for Highway Autonomous DrivingabstractIntelligent decision making and efficient trajectory planning are closely related in autonomous driving technology, especially in highway environment full of dynamic interactive traffic participants. This work integrates them into a unified hierarchical framework with long-term behavior planning (LTBP) and short-term dynamic planning (STDP) running in two parallel threads with different horizon, consequently forming a closed-loop maneuver and trajectory planning system that can react to the dynamic environment effectively and efficiently. In LTBP, a novel voxel structure and the ‘voxel expansion’ algorithm are proposed for the generation of driving corridors in 3D configuration, which involves the prediction states of surrounding vehicles. By using Dijkstra search, the maneuver with minimal cost is determined in form of voxel sequences, then a quadratic programming (QP) problem is constructed for solving the optimal trajectory. And in STDP, another small-scaled QP problem is performed to track or adjust the reference trajectory from LTBP in response to the dynamic obstacles. Meanwhile, a Responsibility-Sensitive Safety (RSS) Checker keeps running at high frequency for real-time feedback to ensure security. Experiments on real data collected in different highway scenarios demonstrate the effectiveness and efficiency of our work. Ting Zhang 0014, Wenjie Song 0001, Mengyin Fu, Yi Yang 0009, Xiaohui Tian, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Towards Autonomous Parking using Vision-only SensorsabstractExisting autonomous parking solutions usually require special signs, pre-built maps or accurate ranging sensors to achieve reliable perception of the parking environment, but these methods are difficult to popularize because they either require preconditions or are expensive for production cars. In this paper, we propose a vision-only autonomous parking solution based on only six cameras. Through the appropriate depth estimation algorithms, our method obtains the pixel level depth of the image, and constructs a dense point cloud, so as to realize the fine perception of the parking environment. An improved Radon transform based parking space detection method are applied for better parking space detection method. Our proposed method achieves processing speed of above 5 Hz on a intermediate level computing platform. Furthermore, we demonstrate the practicability of the proposed system in real-world parking lots. Yi Yang 0009, Miaoxin Pan, Sitan Jiang, Jianhang Wang, Meiling Wang 0002 |
IROS | 1 |
| 2020 | Dynamic Object Tracking for Self-Driving Cars Using Monocular Camera and LIDARabstractThe detection and tracking of dynamic traffic participants (e.g., pedestrians, cars, and bicyclists) plays an important role in reliable decision-making and intelligent navigation for autonomous vehicles. However, due to the rapid movement of the target, most current vision-based tracking methods, which perform tracking in the image domain or invoke 3D information in parts of their pipeline, have real-life limitations such as lack of the ability to recover tracking after the target is lost. In this work, we overcome such limitations and propose a complete system for dynamic object tracking in 3D space that combines: (1) a 3D position tracking algorithm based on monocular camera and LIDAR for the dynamic object; (2) a re-tracking mechanism (RTM) that restore tracking when the target reappears in camera's field of view. Compared with the existing methods, each sensor in our method is capable of performing its role to preserve reliability, and further extending its functions through a novel multimodality fusion module. We perform experiments in the real-world self-driving environment and achieve a desired 10Hz update rate for real-time performance. Our quantitative and qualitative analysis shows that this system is reliable for dynamic object tracking purposes of self-driving cars. Lin Zhao 0016, Meiling Wang 0002, Sheng Su, Tong Liu 0009, Yi Yang 0009 |
IROS | 5 |
| 2020 | Lane Detection in Low-light Conditions Using an Efficient Data Enhancement: Light Conditions Style TransferabstractNowadays, deep learning techniques are widely used for lane detection, but application in low-light conditions remains a challenge until this day. Although multi-task learning and contextual-information-based methods have been proposed to solve the problem, they either require additional manual annotations or introduce extra inference overhead respectively. In this paper, we propose a style-transfer-based data enhancement method, which uses Generative Adversarial Networks (GANs) to generate images in low-light conditions, that increases the environmental adaptability of the lane detector. Our solution consists of three parts: the proposed SIM-CycleGAN, light conditions style transfer and lane detection network. It does not require additional manual annotations nor extra inference overhead. We validated our methods on the lane detection benchmark CULane using ERFNet. Empirically, lane detection model trained using our method demonstrated adaptability in low-light conditions and robustness in complex scenarios. Our code for this paper will be publicly available. Tong Liu 0009, Zhaowei Chen, Yi Yang 0009 |
IV | 3 |
| 2020 | Trajectory Prediction based on Constraints of Vehicle Kinematics and Social Interaction†abstractTrajectory prediction for vehicles is a popular subject since it is beneficial for efficient and secure trajectory planning. In structured traffic scenarios, the behaviour and motion of vehicles are heavily dependent on the social interaction constraints, such as road geometry and surrounding vehicles, and the kinematics model constraints, such as continuous heading and maximum acceleration. To take these factors into account, we analyse the particular characteristics of driving vehicles and propose a model that predicts the possible and feasible trajectory for host vehicle in 3 seconds. In this model, the trajectory of host vehicle takes the center-line as reference, imitates the leader vehicle and focuses on the social vehicles through attention concentration mechanism (ACM) with spatial and temporal information encoded in a fusion hidden state. Furthermore, in order to make the trajectory feasible for vehicle dynamics and kinematics, we introduce a prediction diagnosis method to check the continuous heading and maximum acceleration condition, pruning and adjusting the prediction candidates. Experiments on released public datasets show that this framework can well evaluate the traffic interactions and forecast the trajectory more accurately than common networks. Ting Zhang 0014, Mengyin Fu, Wenjie Song 0001, Yi Yang 0009, Meiling Wang 0002 |
SMC | 4 |
| 2018 | Underwater Modeling, Experiments and Control Strategies of FroBotabstractFroBot can locomote both on land and underwater based on its dual swing-legs propulsion mechanism. This paper presents the dynamic model, experimental studies, and control strategies of FroBot underwater. In this work, an experimental setup consisting of two-degree-of-freedom(2DOF) robotic swing-legs is built to study the model of FroBot underwater. We first improve the dynamic model of caudal fins based on the Morison equation. Combined with experimental data, we optimize the model parameters and then obtain the optimal control strategy of uniform swing. In addition, we apply the CPGs control strategy and improve it based on the FroBot model. These two control strategies have their advantages and demonstrate the potential for future use in control applications. Yi Yang 0009, Zhenhui Fan, Zhongjing Zhu, Jianqing Zhang |
IROS | 1 |
| 2018 | Real-Time Obstacles Detection and Status Classification for Collision Warning in a Vehicle Active Safety SystemabstractThis paper presents real-time obstacles detection and their status classification method for collision warning in the vehicle active safety system. Specifically, stereo cameras and millimeter wave (mmw)-radar are fused to help the driving ego-vehicle to find “Danger” or “Potential Danger” in a timely way through combining with the vehicle kinematic model. The proposed method makes full use of the unique advantages of stereo cameras and mmw-radar to sense the environment through several modules. Cameras are mainly used to detect the near or lateral dynamic objects and to obtain the obstacles region of interest (ROI) considering its rich information and high sensitivity to the lateral displacement, while far or longitudinal relative dynamic objects are detected by mmw-radar according to its observational ability to make up for the disadvantage of cameras. In detail, a cameras detector utilizes ”error vectors” rather than the optical flow to obtain dynamic classes through two times clustering. Mmw-radar mainly detects relative dynamic objects, whose absolute speed can be computed according to the ego-vehicle's state. Then, the detected objects of these two detectors are integrated in an obstacles ROI map, which is obtained through an UV-disparity obstacles detection algorithm to get the final dynamic and relative dynamic objects. Finally, they are classified by comparing them with a dangerous area that is acquired according to the vehicle kinematic model in a special vehicle coordinate system, which is fixed to the ground temporarily. This method is tested on our mobile platforms and the results prove that it can work effectively even though the ego-vehicle drives quickly. Wenjie Song 0001, Yi Yang 0009, Mengyin Fu, Fan Qiu, Meiling Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Real-time lane detection and forward collision warning system based on stereo visionabstractThis paper presents a real-time and robust lane detection and forward collision warning technique based on stereo cameras. First, obstacles image is obtained through stereo matching and UV-disparity segmentation algorithm. Then, Inverse Perspective Mapping(IPM) and Sobel filtering are conducted to generate a low-noise top view of the road by fusing the obstacles image and the original image. Next, Hough Transformation for the top view map is completed and the extreme points(poles) are calculated as the detected lanes according to the traffic lanes model. Besides, the host lane is selected or supplemented among all the detected lanes and the nearest obstacle in this host lane is detected for the forward collision warning. Experimental results on the public data set indicate that our method can work effectively and real-timely in the normal structured environment. Wenjie Song 0001, Mengyin Fu, Yi Yang 0009, Meiling Wang 0002, Xinyu Wang 0018, Alain L. Kornhauser |
Intelligent Vehicles Symposium | 3 |
| 2017 | Intersection scan model and probability inference for vision based small-scale urban intersection detectionabstractLarge-scale intersections stamped on maps have diverse visual features for detection, while small-scale urban intersections are hard to be identified especially when GPS signals are missing. In this paper, we propose a Hidden Markov Model (HMM) based small-scale intersection detection method utilizing monocular vision. We extract visual cues of road transformations and dynamic vehicles' tracks, and then design an Intersection Scan Model to obtain the potential traversable direction of the current road, which is the primary criterion of the intersection estimation. For better performances, we take the detections of consecutive frames into consideration and finally integrate them into HMM to estimate the probabilities of intersections. Results from KITTI datasets and real-world experiments have shown the functionality of the presented approach. Yi Yang 0009, Hao Li 0075, Hao Zhu 0002, Songtian Shang, Ningyi Lyu, Wenjie Song 0001 |
Intelligent Vehicles Symposium | 1 |
| 2017 | Smooth path planning for autonomous parking systemabstractIn this paper, we present a path planning algorithm for autonomous parking system. We focus on the kinematics of the car-like vehicle and improve the conventional geometric parking algorithm by proposing a new curve element named linearly steering spiral. A path planning algorithm based on smooth path searching and optimizing are presented. This method can generate smooth paths incrementally, and the reference control signals can be deduced directly once the path is determined. A simple closed-loop controller is designed in order to deal with the uncertainty from various aspects. Simulations are implemented in different scenarios including obstacle-free and cluttered environments, moreover, a real-world online experiment is executed. The results indicate that the proposed method achieves good performance on both accuracy and computational cost. Yi Yang 0009, Lu Zhang 0047, Xin Qu, Jinzhou Lei, Yijin Li, Jianhang Wang |
Intelligent Vehicles Symposium | 1 |
| 2017 | An efficient decision and planning method for high speed autonomous driving in dynamic environmentabstractThis paper describes an improved decision and planning algorithm based on our previously proposed methods for unmanned ground vehicle (UGV). The new method can be applied to UGV driving both in structured environment and unstructured environment. In the improved method, the prospect of planning is extended from 40m to 100m for safe driving at high speed and some piecewise linear speed functions are designed for the new prospect. After this improvement our UGV now can drive at a maximum speed of 60km/h rather than 40km/h while avoiding obstacles safely. Besides, a velocity feedforward control is added to make the UGV overtake other cars driving at about 25km/h on the road. At last, the collision detection algorithm is improved to make the lane changing maneuver safer. The proposed decision and planning algorithm is implemented both on our old Polaris all terrain vehicle (ATV) and new FAW-H7 car, which exhibited good performance on Across Dangers & Obstacles 2016, Tahe, China and Future Challenge 2016, Changshu, China, respectively. Kai Zhang 0030, Mengyin Fu, Yi Yang 0009, Songtian Shang, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 3 |
| 2015 | Collision-free and kinematically feasible path planning along a reference path for autonomous vehicleabstractFor the local path planning problem of autonomous vehicle in a complicated environment, a method combining cubic hermite spline curves with the kinematic model of autonomous vehicle is developed. And a novel algorithm for obstacle avoidance, called navigation circle, is proposed to take the road structure into account, which is a practical method for real-time path planning. In the new method, one of the trajectory generated by cubic hermite spline curves or navigation circle is optimized through the kinematic model of autonomous vehicle to get the kinematically feasible trajectory. The optimization is actually a numerical forward propagation and is easy to implement. The simulation experiment is conducted on the Robot Operating System (ROS) platform, which is based on replaying the data of the real world obtained from sensors or other modules on autonomous vehicle. Satisfactory simulation results verify the validity and the efficiency of the proposed method as well as the planner's capability to navigate in a realistic scenario. Mengyin Fu, Kai Zhang 0030, Yi Yang 0009, Hao Zhu 0002, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 3 |
| 2014 | Standing-up control and ramp-climbing control of a spherical wheeled robotabstractThis paper proposes a new type of spherical wheeled robot with an annular support leg. It can keep statically stable when powered off and automatically stand up with the assistance of the support leg when powered on. The stability of the robot at equilibrium is verified firstly using the planar simplified model. And the robot is proved to be controllable. Thus a double-closed loop control system is designed to stabilize the robot. Based on it, the standing-up control system and ramp-climbing control system are realized by changing control structure and using fuzzy control strategies. The design of the annular support leg, the experimental results and conclusions are also described in this paper. Jian Jian, Meiling Wang 0002, Ningyi Lv, Yi Yang 0009, Tong Liu 0009 |
ICARCV | 4 |
| 2014 | Moving object detection under dynamic background in 3D range dataabstractWe proposed an unsupervised algorithm to extract profile features and detect moving object under dynamic background in 3D range Data. Moving object detection under dynamic background has become an increasingly popular research topic in mobile robotics. For the characteristics of dynamic background scene, we proposed an online unsupervised moving object detection algorithm, based on Gaussian Mixture Models and Motion Compensation. Furthermore, we did the work of clustering and identifying of the targets. In order to improve the robustness of the algorithm, we used a tracker to track the results of the detection. At last, experimental results on real laser data depicting urban and rural scenes under static and dynamic background are presented. Yi Yang 0009, Yan Guang, Hao Zhu 0002, Mengyin Fu, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 1 |
| 2013 | Lane recognition self-learning scheme of mobile robot based on integrated perception systemabstractIn this paper, a kind of integrated perception system for mobile robot is presented, which consists of 3D Lidar, 2D camera and their spatial registration. Based on the system and support vector machine (SVM), a self-supervised learning scheme between 3D point cloud data and 2D image data has been established, which can identify the traversable lane in driving environments through data association and parameters training. With this approach, vision-based autonomous navigation can be achieved and its effectiveness has been verified by extensive robot experiments. Yi Yang 0009, Hao Zhu 0002, Mengyin Fu, Meiling Wang 0002 |
Intelligent Vehicles Symposium | 1 |