Xiong You

dblp:59/7385 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic Path Planning for Unmanned Ground Vehicles Based on LSTM and Distributed PPO in Off-Road Environments
abstract
Path planning for unmanned ground vehicles (UGVs) in off-road environments faces challenges such as inadequate trafficability assessment, local optima issues, and limited dynamic obstacle avoidance. This paper presents a dynamic path planning algorithm called LSTM-DPPO, which combines long short-term memory (LSTM) and distributed proximal policy optimization (DPPO). Firstly, a UGV trafficability map was created that incorporates three environmental factors, which provides reliable environmental representation for UGV. Secondly, a comprehensive reward function was developed, which combines dynamic obstacle avoidance with trafficability assessment. This function facilitates the collaborative optimization of both path safety and overall trafficability. Additionally, to address the issues of ineffective dynamic obstacle avoidance and the challenges of model convergence commonly found in reinforcement learning algorithms, we introduced an LSTM network to build a dynamic obstacle prediction module. We also employed a distributed PPO architecture to enhance the training speed of the model. Finally, a comparative simulation experiment was conducted between the LSTM-DPPO algorithm and six path planning algorithms, and the generalization ability of the LSTM-DPPO algorithm was verified in three new scenarios. The results indicate that the average inference time of the complete LSTM-DPPO algorithm is 1.74ms, which meets the requirements of realtime path planning. Although its planning time is slightly longer than DPPO algorithm and CV-DPPO algorithm (Constant Velocity Model, CV), it is much faster than the other four algorithms. The LSTM-DPPO algorithm excels in obstacle avoidance, has the highest success rate, and demonstrates good generalization in new scenarios.
Qingyun Liu 0016, Xiong You, Jian Yang 0034, Jiwei Zuo, Xiangtian Bai
IEEE Internet Things J.2
2025 Dual-BEV Nav: Dual-Layer BEV-Based Heuristic Path Planning for Robotic Navigation in Unstructured Outdoor Environments
abstract
Path planning with strong environmental adaptability plays a crucial role in robotic navigation in unstructured outdoor environments, especially in the case of low-quality location and map information. The path planning ability of a robot depends on the identification of the traversability of global and local ground areas. In real-world scenarios, the complexity of outdoor open environments makes it difficult for robots to identify the traversability of ground areas that lack a clearly defined structure. Moreover, most existing methods have rarely analyzed the integration of local and global traversability identifications in unstructured outdoor scenarios. To address this problem, we propose a novel method, Dual-BEV Nav, first introducing Bird's Eye View (BEV) representations into local planning to generate high-quality traversable paths. Then, these paths are projected into the global traversability probability map generated by the global BEV planning model to obtain the optimal path. By integrating the traversability from both local and global BEV, we establish a dual-layer BEV heuristic planning paradigm, enabling long-distance navigation in unstructured outdoor environments. We test our approach through both public dataset evaluations and real-world robot deployments, yielding promising results. Compared to baselines, the Dual-BEV Nav improved temporal distance prediction accuracy by up to 18.26%. In the real-world deployment, under conditions significantly different from the training set and with notable occlusions in the global BEV, the Dual-BEV Nav successfully achieved a 65-meter-long outdoor navigation. Further analysis demonstrates that the local BEV representation significantly enhances the rationality of the planning, while the global BEV probability map ensures the robustness of the overall planning.
Jian Yang 0034, Shibo Huang, Ke Li 0005, Xian Wei, Xiong You
ICRA9
2025 KiteRunner: Language-Driven Cooperative Local-Global Navigation Policy with UAV Mapping in Outdoor Environments
abstract
Autonomous navigation in open-world outdoor environments faces challenges in integrating dynamic conditions, long-distance spatial reasoning, and semantic understanding. Traditional methods struggle to balance local planning, global planning, and semantic task execution, while existing large language models (LLMs) enhance semantic comprehension but lack spatial reasoning capabilities. Although diffusion models excel in local optimization, they fall short in large-scale long-distance navigation. To address these gaps, this paper proposes KiteRunner, a language-driven cooperative local-global navigation strategy that combines UAV orthophoto-based global planning with diffusion model-driven local path generation for long-distance navigation in open-world scenarios. Our method innovatively leverages real-time UAV orthophotography to construct a global probability map, providing traversability guidance for the local planner, while integrating large models like CLIP and GPT to interpret natural language instructions. Experiments demonstrate that KiteRunner achieves 5.6% and 12.8% improvements in path efficiency over state-of-the-art methods in structured and unstructured environments, respectively, with significant reductions in human interventions and execution time.
Shibo Huang, Chenfan Shi, Jian Yang 0034, Jinpeng Mi, Ke Li 0005, Miao Ding, Peidong Liang, Xiong You, Xian Wei
IROS10
2025 NeuroLoc: Encoding Navigation Cells for 6-DOF Camera Localization
abstract
Recently, camera localization has been widely adopted in autonomous robotic navigation due to its efficiency and convenience. However, autonomous navigation in unknown environments often suffers from scene ambiguity, environmental disturbances, and dynamic object transformation in camera localization. To address this problem, inspired by the brain cognitive navigation mechanism (such as grid cells, place cells, and head direction cells), we propose a novel neurobiological camera location method, namely NeuroLoc. Firstly, we designed a Hebbian learning module driven by place cells to save and replay historical information, aiming to restore the details of historical representations and solve the issue of scene fuzziness. Secondly, we utilized the head direction cell-inspired internal direction learning as multi-head attention embedding to help restore the true orientation in similar scenes. Finally, we added a 3D grid center prediction in the pose regression module to reduce the final wrong prediction. We evaluate the proposed NeuroLoc on commonly used benchmark indoor and outdoor datasets. The experimental results show that our NeuroLoc can enhance the robustness in complex environments and improve the performance of pose regression by using only a single image.
Jian Yang 0034, Fenli Jia, Muyu Wang, Jinpeng Mi, Jilin Hu, Peidong Liang, Ke Li 0005, Xiong You, Xian Wei
IROS11
2025 Deeper and Broader Multimodal Fusion: Cascaded Forest-of-Experts for Land Cover Classification
abstract
Multimodal land cover classification (LCC) of optical and SAR images has become a research hotspot. However, there are still two unsolved problems: the lack of a deep fusion mechanism and the neglect of the diversity of multimodal features. Inspired by ensemble learning, this letter proposes the cascaded multimodal forest-of-experts (CM2FEs) for deeper and broader fusion to further improve the performance of LCC. The proposed method first establishes the expert tree, then combines multiple trees at the same level into a forest, and finally forms a cascaded forest across different levels. Specifically, the novel designs include three points: 1) the multimodal expert tree is built based on linear projection and dynamic routing, with multiple layers of experts; it can acquire more discriminative multimodal features through deeper fusion; 2) the cascaded forest is formed by combining expert trees at the same level and different levels, which can effectively ensemble the knowledge learned by different trees; it can generate more diverse multimodal features through broader fusion; and 3) two expert exchange strategies are proposed to transfer knowledge between different trees and further optimize the feature fusion effect. Experiments show that the proposed method performs better than existing methods, and the mean IoU (mIoU) has been improved by at least 1.60%–3.25%.
Guangxia Wang, Kuiliang Gao, Xiong You
IEEE Geosci. Remote. Sens. Lett.3
2025 Multi-Scale Oriented Object Detection With Focus Error Ellipse Loss
abstract
The loss function and feature extraction framework are essential parts of the algorithm design and significantly affect the accuracy of oriented object detection in remote sensing images. Though considerable progress has been made, there are still challenges left to be explored, e.g., large variations in scales, arbitrary direction, and dense distribution of the objects, which may have some undesirable effects, such as inaccurate object position regression, high false alarm, and miss rate. To address the above problems, we propose a Focus Error Ellipse (FEE) loss function. This function bolsters the detection accuracy by narrowing the distance between the center points of the labeled and predicted bounding boxes based on the Error Ellipse. For the network part, we carefully crafted two unit modules: a Fine-grained and Context-augmented Module (FCM) and a Semantic Information Regrouping Module (SIRM). The FCM aligns fine-grained information with contextual information to establish dependencies between local and global features, which helps grasp the more holistic characteristics of objects. The SIRM reorganizes the acquired deep semantic features in the channel dimension, enhances the weight of task-beneficial semantic information, and further derives the optimal combination method of feature subsets for object detection. Based on the aforementioned work, we developed an oriented object detection framework, which further improves the detection accuracy of large aspect ratio objects and dense scenes. Experimental results show that the proposed method can produce competitive performance in oriented object detection compared to other state-of-the-art models.
Xuanbei Lu, Ke Li 0005, Gong Cheng 0003, Xiong You
IEEE Trans. Geosci. Remote. Sens.5
2025 Fast Path Planning of Multienvironmental Factors Comprehensive Constraints in Off-Road Environment Constructed by Aggregation Vector Point Method
abstract
The accuracy and efficiency of path planning in off-road environments depend on the construction of off-road environment map information. Previous studies have used the grid method to represent off-road environments, but as the number of grids increases, the path planning time significantly increases. Moreover, environmental factors such as elevation, slope, aspect, and land cover type are all factors that affect the trafficability of unmanned vehicles. Therefore, this paper proposes a fast path planning method for multi-environmental factor comprehensive constraints in off-road environment constructed by aggregation vector point method. First, the required off-road environment image data are converted into vector points, the trafficability of the point is quantified based on its environmental attribute information, and the impassable points are aggregated. The aggregated vector points are used to construct a vector environment model (VEM). Second, based on the environmental attribute information of the vector points, the slope influence layer, land cover type influence layer, and goal guidance layer are constructed to quantitatively express the comprehensive trafficability of each point. Finally, to verify the effectiveness of the VEM constructed by the proposed aggregation vector method, the Dijkstra algorithm with comprehensive constraints from multiple environmental factors is used to perform path planning in both the VEM and the grid environment model (GEM), which is constructed using the traditional grid method. The results indicate that the VEM constructed using the aggregation vector point method can accurately reflect environmental information while improving path planning efficiency. The path planning time required in grid environments was longer than that in vector environments. In the grid environment of small, medium and large research areas, the planning path time is about 4.5, 1.9, and 1.26 times of that in the vector environment of small, medium and large research areas, respectively. The path planned by Dijkstra algorithm after comprehensive constraints is safer and more feasible.
Qingyun Liu 0016, Xiong You, Jiwei Zuo
IEEE Trans. Geosci. Remote. Sens.2
2025 Rethinking Semantic Segmentation With Multi-Grained Logical Prototype
abstract
The last decade has witnessed significant advances in semantic segmentation brought about by deep learning. However, existing methods only fit the data-label correspondence in a data-driven manner and do not fully conform to the abstraction and structuralization characteristics of the human visual cognition process, which limits the upper bounds of their performance. To this end, a multi-grained logical prototype (MGLP) method is proposed to rethink semantic segmentation based on these two key characteristics. Its novel design can be summarized as follows. 1) For abstraction, prototypes of the same class at different grain levels are established: a label generation method is proposed to automatically generate a multi-grained label space, which can guide the learning of the multi-grained prototypes for each class. 2) For structuralization, the intrinsic logical structure across different semantic levels is explicitly modeled: the horizontal metric relationships are established via metric relation operations on prototypes at the same grain level, to improve the discriminability between classes while taking the vertical semantic hierarchy into account. Moveover, the vertical logical relationships are established as the sub-to-super positive and super-to-sub negative constraints, to strengthen the semantic dependencies among prototypes at different grain levels. 3)MGLP is plug-and-play and can be directly combined with existing segmentation methods. Extensive experimental results indicate that MGLP can significantly improve the segmentation performance of existing methods, which opens up a new avenue for future research.
Anzhu Yu, Kuiliang Gao, Xiong You, Yanfei Zhong, Bing Liu 0018, Chunping Qiu
IEEE Trans. Image Process.3
2025 FuseFormer: A Manifold Metric Fusing Attention for Pedestrian Trajectory Prediction
abstract
Accurate pedestrian trajectory prediction is critical for ensuring the safety of autonomous vehicles and advancing higher levels of driving automation. However, the complex interpersonal interactions and highly dynamic trajectory patterns in real-world scenarios pose significant challenges to achieving precise predictions. Recently, Transformers have shown remarkable success in pedestrian trajectory prediction, primarily due to their effective modeling of temporal and spatial dependencies via Multi-Head Self-Attention (MHA) mechanisms. Despite these advancements, existing self-attention methods often rely on Euclidean distance-based metrics and dot-product operations, which are inadequate for capturing interaction-induced trajectory curvatures. To address this limitation, we propose a novel hybrid Transformer architecture, FuseFormer, that incorporates Geodesic Self-Attention (GSA) mechanisms. GSA utilizes geodesic distances to characterize interaction features effectively, complementing MHA, which excels in capturing local features and maintaining temporal correlations. FuseFormer employs a gating network to adaptively combine GSA and MHA embeddings, leveraging their complementary strengths. Additionally, FuseFormer integrates a Transformer-based Neural Ordinary Differential Equation (ODE) decoder to model trajectory temporal dynamics. This design enables the generation of future trajectories that align closely with motion trends while adapting the network depth to input sequence lengths. Experimental results demonstrate that FuseFormer achieves state-of-the-art performance across widely used pedestrian trajectory prediction datasets, including ETH/UCY, SDD, and NBA. These results underscore the model’s effectiveness and generalization capability in capturing complex interaction patterns and handling diverse scenarios.
Kohsin Ko, Jian Yang 0034, Ke Li 0005, Xiong You, Jinpeng Mi, Mingsong Chen 0001, Xian Wei
IEEE Trans. Intell. Transp. Syst.6
2024 ProEqBEV: Product Group Equivariant BEV Network for 3D Object Detection in Road Scenes of Autonomous Driving
abstract
With the rapid development of autonomous driving systems, 3D object detection based on Bird’s Eye View (BEV) in road scenes has witnessed great progress over the past few years. As a road scene exhibits a part-whole hierarchy between the within objects and the scene itself, simple parts (e.g., roads, lane lines, vehicles and pedestrians) can be assembled into progressively more complex shapes to form a BEV representation of the whole road scene. Therefore, a BEV often has multiple levels of freedom on motion, i.e., the rotation and the moving shift of the whole BEV, and the random movements of objects (e.g., pedestrians and vehicles) inside the BEV. However, most of the current single-sensor or multi-sensor fusion-based BEV object detection methods have not yet taken into account capturing such multi-level motion in a BEV. To address this problem, we propose a product group equivariant object detection network framework that is equivariant with respect to multiple levels of symmetry groups based on multi-sensor fusion. The proposed framework extracts local equivariant features of objects in point clouds, while global equivariant features are extracted in both point clouds and images. Furthermore, the network learns diverse rotation-equivariant features and mitigates a significant amount of detection errors caused by rotations of BEV and objects inside a BEV, thereby further enhancing the performance of object detection. The experiment results show that the network architecture significantly improves object detection on mAP and NDS, respectively. In addition, in order to demonstrate the effectiveness of the proposed local-multi-global equivariant components, we conduct sufficient ablation experiments. The results show that the individual components are indispensable for the object detection performance improvement of the overall network architecture.
Jian Yang 0034, Ke Li 0005, Jianzhang Zheng, Xihao Wang, Mingsong Chen 0001, Xiong You, Xian Wei
ICRA9
2024 Attention Prompt-Driven Source-Free Adaptation for Remote Sensing Images Semantic Segmentation
abstract
Recently, remote sensing images (RSIs) domain adaptation segmentation has been extensively studied. However, existing methods generally assume that source RSIs must be available, which is obviously an overly demanding condition and will increase unnecessary costs in practice. To this end, this letter takes the lead in exploring RSIs source-free adaptation segmentation, where only the offline model pretrained on the source domain and target RSIs are available. A novel method featuring prompt learning and vision foundation models is proposed, and the novelty design includes two aspects. First, to better adapt the general-purpose knowledge in the foundation model to different target RSIs, an attention-guided prompt tuning strategy is proposed, which can dynamically steer the knowledge at different layers and positions through prompts with different weights. Second, a feature alignment strategy with similarity distance is proposed for source-free domain adaptation by taking full advantage of the representation ability of the foundation model and the flexibility of prompt learning. Extensive experiments indicate that the performance of the proposed method is significantly superior to that of existing methods. Specifically, the mIoU of target RSIs has been improved by at least 3.14%~4.18%.
Kuiliang Gao, Xiong You, Ke Li 0005, Juan Lei, Xibing Zuo
IEEE Geosci. Remote. Sens. Lett.2
2024 Integrating Multiple Sources Knowledge for Class Asymmetry Domain Adaptation Segmentation of Remote Sensing Images
abstract
In the existing unsupervised domain adaptation (UDA) methods for remote sensing images (RSIs) semantic segmentation, class symmetry is a widely followed ideal assumption, where the source and target RSIs have exactly the same class space. In practice, however, it is often very difficult to find a source RSI with exactly the same classes as the target RSI. More commonly, there are multiple source RSIs available. And there is always an intersection or inclusion relationship between the class spaces of each source–target pair, which can be referred to as class asymmetry. Nevertheless, the class asymmetry domain adaptation segmentation of RSIs with multiple sources has not yet been explored. To this end, a novel class asymmetry RSIs domain adaptation method is proposed for the first time in this article, which consists of four key components. First, a multibranch segmentation network is built to learn an expert for each source RSI. Second, a novel collaborative learning method with the cross-domain mixing strategy is proposed, to supplement the class information for each source while achieving the domain adaptation of each source–target pair. Third, a pseudolabel generation strategy is proposed to effectively combine the strengths of different experts, which can be flexibly applied to two cases where the source class union is equal to or includes the target class set. Fourth, a multiview-enhanced knowledge integration module is developed for high-level knowledge routing and transfer from multiple domains to target predictions. The experimental results of six different class settings on airborne and spaceborne RSIs show that the proposed method can effectively perform the multisource domain adaptation in the case of class asymmetry, and the obtained segmentation performance of target RSIs is significantly better than the existing relevant methods.
Kuiliang Gao, Anzhu Yu, Xiong You, Wenyue Guo, Ke Li 0005, Ningbo Huang
IEEE Trans. Geosci. Remote. Sens.3
2023 Prototype and Context-Enhanced Learning for Unsupervised Domain Adaptation Semantic Segmentation of Remote Sensing Images
abstract
In unsupervised domain adaptation (UDA) of remote sensing images (RSIs), the huge inter-domain discrepancies and intra-domain variances lead to complicated class-level relations. Specifically, the instances of the same class differ greatly while instances of different classes are similar, whether across different RSIs domains or within the same RSIs domain. However, existing methods cannot fully consider these problems, limiting the performance of UDA semantic segmentation of RSIs. To this end, this paper proposes a novel cross-domain multi-prototypes learning method, the core idea of which is to abstract the cross-and intra-domain class-level relations into multiple prototypes. Specifically, the multiple prototypes belonging to different classes can detailedly describe complex inter-class relations, and the multiple prototypes within the same class can better model rich intra-class relations. Further, the source and target samples are jointly used for prototypes calculation, to fully fuse the feature information of different RSIs. In a nutshell, utilizing the samples from different RSIs domains to learn multiple prototypes for each class can achieve better domain alignment at the class level. In addition, considering that RSIs simultaneously contain large targets with wide coverage and important small targets, two masked consistency learning strategies are designed to better explore the contextual structure of target RSIs and improve the quality of pseudo labels for prototype updating. The global consistency strategy can strengthen the utilization of global context relations, while the local consistency strategy can further improve the learning of local context details. Therefore, the proposed method is actually a prototype and context enhanced learning method for UDA semantic segmentation of RSIs. Extensive experiments demonstrate that the proposed method can achieve better performance than existing state-of-the-art UDA methods.
Kuiliang Gao, Anzhu Yu, Xiong You, Chunping Qiu, Bing Liu 0018
IEEE Trans. Geosci. Remote. Sens.3
2021 PL-VSCN: Patch-level vision similarity compares network for image matching
abstract
Abstract Image matching plays an important role in various computer vision tasks, such as image retrieval and loop closure detection in Simultaneous Localization and Mapping. The authors propose a discriminative patch‐based image matching method that converts the problem of whole image matching to that of local patch matching. To construct the patch representation, the Patch‐Level Vision Similarity Compare Network (PL‐VSCN) is proposed to produce the patch feature. In the image matching process, local patches that potentially contain objects within images are initially detected, and the discriminative feature of each patch is extracted based on the pre‐trained PL‐VSCN. Then, the similarities between the patch pairs are calculated to construct the similarity matrix, and the corresponding patch pairs are detected based on the mutual matching mechanism on the similarity matrix. Experimental results indicate that the proposed PL‐VSCN can generate the discriminative patch feature, which can accurately match the patch pairs with the corresponding content and distinguish those with non‐corresponding content. In addition, the comparison experiments demonstrate that the proposed image matching method outperforms existing approaches on most datasets and effectively completes the image matching task.
Xiong You, Qin Li 0005, Ke Li 0005, Anzhu Yu, Shuhui Bu
IET Comput. Vis.1
2018 Rotation-Insensitive and Context-Augmented Object Detection in Remote Sensing Images
abstract
Most of the existing deep-learning-based methods are difficult to effectively deal with the challenges faced for geospatial object detection such as rotation variations and appearance ambiguity. To address these problems, this paper proposes a novel deep-learning-based object detection framework including region proposal network (RPN) and local-contextual feature fusion network designed for remote sensing images. Specifically, the RPN includes additional multiangle anchors besides the conventional multiscale and multiaspect-ratio ones, and thus can deal with the multiangle and multiscale characteristics of geospatial objects. To address the appearance ambiguity problem, we propose a double-channel feature fusion network that can learn local and contextual properties along two independent pathways. The two kinds of features are later combined in the final layers of processing in order to form a powerful joint representation. Comprehensive evaluations on a publicly available ten-class object detection data set demonstrate the effectiveness of the proposed method.
Ke Li 0005, Gong Cheng 0003, Shuhui Bu, Xiong You
IEEE Trans. Geosci. Remote. Sens.4
2016 Place recognition based on deep feature and adaptive weighting of similarity matrix
Qin Li 0005, Ke Li 0005, Xiong You, Shuhui Bu, Zhenbao Liu
Neurocomputing3