Jian Yang 0034

dblp:181/2854-34 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0003-2373-8799ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic Path Planning for Unmanned Ground Vehicles Based on LSTM and Distributed PPO in Off-Road Environments
abstract
Path planning for unmanned ground vehicles (UGVs) in off-road environments faces challenges such as inadequate trafficability assessment, local optima issues, and limited dynamic obstacle avoidance. This paper presents a dynamic path planning algorithm called LSTM-DPPO, which combines long short-term memory (LSTM) and distributed proximal policy optimization (DPPO). Firstly, a UGV trafficability map was created that incorporates three environmental factors, which provides reliable environmental representation for UGV. Secondly, a comprehensive reward function was developed, which combines dynamic obstacle avoidance with trafficability assessment. This function facilitates the collaborative optimization of both path safety and overall trafficability. Additionally, to address the issues of ineffective dynamic obstacle avoidance and the challenges of model convergence commonly found in reinforcement learning algorithms, we introduced an LSTM network to build a dynamic obstacle prediction module. We also employed a distributed PPO architecture to enhance the training speed of the model. Finally, a comparative simulation experiment was conducted between the LSTM-DPPO algorithm and six path planning algorithms, and the generalization ability of the LSTM-DPPO algorithm was verified in three new scenarios. The results indicate that the average inference time of the complete LSTM-DPPO algorithm is 1.74ms, which meets the requirements of realtime path planning. Although its planning time is slightly longer than DPPO algorithm and CV-DPPO algorithm (Constant Velocity Model, CV), it is much faster than the other four algorithms. The LSTM-DPPO algorithm excels in obstacle avoidance, has the highest success rate, and demonstrates good generalization in new scenarios.
Qingyun Liu 0016, Xiong You, Jian Yang 0034, Jiwei Zuo, Xiangtian Bai
IEEE Internet Things J.3
2025 Relaxed Rotational Equivariance via G-Biases in Vision
abstract
Group Equivariant Convolution (GConv) can capture rotational equivariance from original data. It assumes uniform and strict rotational equivariance across all features as the transformations under the specific group. However, the presentation or distribution of real-world data rarely conforms to strict rotational equivariance, commonly referred to as Rotational Symmetry-Breaking (RSB) in the system or dataset, making GConv unable to adapt effectively to this phenomenon. Motivated by this, we propose a simple but highly effective method to address this problem, which utilizes a set of learnable biases called G-Biases under the group order to break strict group constraints and then achieve a Relaxed Rotational Equivariant Convolution (RREConv). To validate the efficiency of RREConv, we conduct extensive ablation experiments on the discrete rotational group Cn. Experiments demonstrate that the proposed RREConv-based methods achieve excellent performance compared to existing GConv-based methods in both classification and 2D object detection tasks on the natural image datasets.
Licheng Sun, Jian Yang 0034, Shing-Ho J. Lin, Jinpeng Mi, Xian Wei
AAAI4
2025 R2Det: Exploring Relaxed Rotation Equivariance in 2D Object Detection
abstract
Group Equivariant Convolution (GConv) empowers models to explore underlying symmetry in data, improving performance. However, real-world scenarios often deviate from ideal symmetric systems caused by physical permutation, characterized by non-trivial actions of a symmetry group, resulting in asymmetries that affect the outputs, a phenomenon known as Symmetry Breaking. Traditional GConv-based methods are constrained by rigid operational rules within group space, assuming data remains strictly symmetry after limited group transformations. This limitation makes it difficult to adapt to Symmetry-Breaking and non-rigid transformations. Motivated by this, we mainly focus on a common scenario: Rotational Symmetry-Breaking. By relaxing strict group transformations within Strict Rotation-Equivariant group $\mathbf{C}_n$, we redefine a Relaxed Rotation-Equivariant group $\mathbf{R}_n$ and introduce a novel Relaxed Rotation-Equivariant GConv (R2GConv) with only a minimal increase of $4n$ parameters compared to GConv. Based on R2GConv, we propose a Relaxed Rotation-Equivariant Network (R2Net) as the backbone and develop a Relaxed Rotation-Equivariant Object Detector (R2Det) for 2D object detection. Experimental results demonstrate the effectiveness of the proposed R2GConv in natural image classification, and R2Det achieves excellent performance in 2D object detection with improved generalization capabilities and robustness. The code is available in \texttt{https://github.com/wuer5/r2det}.
Jian Yang 0034, Mingsong Chen 0001, Xian Wei
ICLR5
2025 Supplementary Material for STTODE: Spatio-Temporal Transformer Ordinary Differential Equation Networks for Pedestrian Trajectory Forecasting
abstract
Dataset 1: ETH-UCY [2] [3] dataset: ETH-UCY consists of the following sub-datasets: ETH, HOTEL, UNIV, ZARA1, ZARA2. In accordance with previous research experimental setups, we split each trajectory sample into 8-second segments, using a time interval of 0.4 seconds. We use the first 3.2 seconds (8 steps) of trajectory data to predict the next 4.8 seconds (12 steps). For training, we use a leave-one-out strategy with 4 subsets and the remaining subset for testing.
YingJie Liu, Jian Yang 0034, Mingsong Chen 0001, Xian Wei
ICME3
2025 Dual-BEV Nav: Dual-Layer BEV-Based Heuristic Path Planning for Robotic Navigation in Unstructured Outdoor Environments
abstract
Path planning with strong environmental adaptability plays a crucial role in robotic navigation in unstructured outdoor environments, especially in the case of low-quality location and map information. The path planning ability of a robot depends on the identification of the traversability of global and local ground areas. In real-world scenarios, the complexity of outdoor open environments makes it difficult for robots to identify the traversability of ground areas that lack a clearly defined structure. Moreover, most existing methods have rarely analyzed the integration of local and global traversability identifications in unstructured outdoor scenarios. To address this problem, we propose a novel method, Dual-BEV Nav, first introducing Bird's Eye View (BEV) representations into local planning to generate high-quality traversable paths. Then, these paths are projected into the global traversability probability map generated by the global BEV planning model to obtain the optimal path. By integrating the traversability from both local and global BEV, we establish a dual-layer BEV heuristic planning paradigm, enabling long-distance navigation in unstructured outdoor environments. We test our approach through both public dataset evaluations and real-world robot deployments, yielding promising results. Compared to baselines, the Dual-BEV Nav improved temporal distance prediction accuracy by up to 18.26%. In the real-world deployment, under conditions significantly different from the training set and with notable occlusions in the global BEV, the Dual-BEV Nav successfully achieved a 65-meter-long outdoor navigation. Further analysis demonstrates that the local BEV representation significantly enhances the rationality of the planning, while the global BEV probability map ensures the robustness of the overall planning.
Jian Yang 0034, Shibo Huang, Ke Li 0005, Xian Wei, Xiong You
ICRA3
2025 PDDFormer: Pairwise Distance Distribution Graph Transformer for Crystal Material Property Prediction
abstract
Crystal structures can be simplified as a periodic point set that repeats across three-dimensional space along an underlying lattice. Traditionally, crystal representation methods rely on descriptors such as lattice parameters, symmetry, and space groups to characterize the structure. However, in reality, atoms in materials always vibrate above absolute zero, causing their positions to fluctuate continuously. This dynamic behavior disrupts the fundamental periodicity of the lattice, making crystal graphs based on static lattice parameters and conventional descriptors discontinuous under slight perturbations. Chemists proposed the pairwise distance distribution (PDD) method to address this. However, the completeness of PDD requires defining a large number of neighboring atoms, leading to high computational costs. Additionally, PDD does not account for atomic information, making it challenging to apply it directly to crystal material property prediction tasks. To tackle these challenges, we introduce the atom-weighted Pairwise Distance Distribution (WPDD) and Unit cell Pairwise Distance Distribution (UPDD) for the first time, applying them to the construction of multi-edge crystal graphs. We demonstrate the continuity and general completeness of crystal graphs under slight atomic position perturbations. Moreover, by modeling PDD as global information and integrating it into matrix-based message passing, we significantly reduce computational costs. Comprehensive evaluation results show that WPDDFormer achieves state-of-the-art predictive accuracy across tasks on benchmark datasets such as the Materials Project and JARVIS-DFT.
Xiangxiang Shen, Lingfeng Wen, Licheng Sun, Jian Yang 0034, Shing-Ho J. Lin, Xiao He 0004, Mingsong Chen 0001, Xian Wei
IJCAI5
2025 KiteRunner: Language-Driven Cooperative Local-Global Navigation Policy with UAV Mapping in Outdoor Environments
abstract
Autonomous navigation in open-world outdoor environments faces challenges in integrating dynamic conditions, long-distance spatial reasoning, and semantic understanding. Traditional methods struggle to balance local planning, global planning, and semantic task execution, while existing large language models (LLMs) enhance semantic comprehension but lack spatial reasoning capabilities. Although diffusion models excel in local optimization, they fall short in large-scale long-distance navigation. To address these gaps, this paper proposes KiteRunner, a language-driven cooperative local-global navigation strategy that combines UAV orthophoto-based global planning with diffusion model-driven local path generation for long-distance navigation in open-world scenarios. Our method innovatively leverages real-time UAV orthophotography to construct a global probability map, providing traversability guidance for the local planner, while integrating large models like CLIP and GPT to interpret natural language instructions. Experiments demonstrate that KiteRunner achieves 5.6% and 12.8% improvements in path efficiency over state-of-the-art methods in structured and unstructured environments, respectively, with significant reductions in human interventions and execution time.
Shibo Huang, Chenfan Shi, Jian Yang 0034, Jinpeng Mi, Ke Li 0005, Miao Ding, Peidong Liang, Xiong You, Xian Wei
IROS3
2025 NeuroLoc: Encoding Navigation Cells for 6-DOF Camera Localization
abstract
Recently, camera localization has been widely adopted in autonomous robotic navigation due to its efficiency and convenience. However, autonomous navigation in unknown environments often suffers from scene ambiguity, environmental disturbances, and dynamic object transformation in camera localization. To address this problem, inspired by the brain cognitive navigation mechanism (such as grid cells, place cells, and head direction cells), we propose a novel neurobiological camera location method, namely NeuroLoc. Firstly, we designed a Hebbian learning module driven by place cells to save and replay historical information, aiming to restore the details of historical representations and solve the issue of scene fuzziness. Secondly, we utilized the head direction cell-inspired internal direction learning as multi-head attention embedding to help restore the true orientation in similar scenes. Finally, we added a 3D grid center prediction in the pose regression module to reduce the final wrong prediction. We evaluate the proposed NeuroLoc on commonly used benchmark indoor and outdoor datasets. The experimental results show that our NeuroLoc can enhance the robustness in complex environments and improve the performance of pose regression by using only a single image.
Jian Yang 0034, Fenli Jia, Muyu Wang, Jinpeng Mi, Jilin Hu, Peidong Liang, Ke Li 0005, Xiong You, Xian Wei
IROS2
2025 Hyperbolic prototype rectification for few-shot 3D point cloud classification
Yuanzhi Feng, Shing-Ho J. Lin, Mu-Yu Wang, Jianzhang Zheng, Ziyao He, Zi-Yi Pang, Jian Yang 0034, Mingsong Chen 0001, Xian Wei
Pattern Recognit.8
2025 A Graph Representation Learning Approach for Imbalanced Ship Type Recognition Using AIS Trajectory Data
abstract
Marine transportation constitutes a vital segment of international trade logistics. Recognizing marine carrier ship types is essential for the governance and efficiency of the marine transportation sector. Recognizing ship types by analyzing Automatic Identification System (AIS) trajectory data represents a fundamental aspect of trajectory classification within Intelligent Transportation Systems (ITS). Due to their data representation limitations, traditional CNNs and RNNs struggle to simultaneously process the topological relationships and the attributes of trajectory points. This restriction significantly hinders the learning of micro-activity behaviors of ships. To overcome this, we introduce Graph Isomorphic Network (GIN), a new variant of graphical neural networks with enhanced representation capabilities. Therefore, this study develops a GIN-based framework for graph representation learning. Initially, ship trajectories are represented as vector graph structures, with both original and derived features incorporated into the node embeddings. Subsequently, a GIN-based model for graph representation learning is employed to extract features necessary for ship type recognition. The model’s performance is evaluated through experiments on ground truth data from the Gulf of Mexico and the New Jersey Bight, achieving recognition accuracies of 93.95% and 92.33%, respectively. The high accuracy and robustness of the model substantiate the effectiveness of the proposed method in ship type recognition. This research illustrates that GIN-based graph representation learning can efficiently capture trajectory features, offering a valuable reference for motion trajectory representation extraction and pattern recognition in diverse geographical contexts.
Jiale Pan, Rui Xin 0001, Jian Yang 0034, Fanlin Yang, Bingchao Xu, Fenli Jia
IEEE Trans. Intell. Transp. Syst.3
2025 FuseFormer: A Manifold Metric Fusing Attention for Pedestrian Trajectory Prediction
abstract
Accurate pedestrian trajectory prediction is critical for ensuring the safety of autonomous vehicles and advancing higher levels of driving automation. However, the complex interpersonal interactions and highly dynamic trajectory patterns in real-world scenarios pose significant challenges to achieving precise predictions. Recently, Transformers have shown remarkable success in pedestrian trajectory prediction, primarily due to their effective modeling of temporal and spatial dependencies via Multi-Head Self-Attention (MHA) mechanisms. Despite these advancements, existing self-attention methods often rely on Euclidean distance-based metrics and dot-product operations, which are inadequate for capturing interaction-induced trajectory curvatures. To address this limitation, we propose a novel hybrid Transformer architecture, FuseFormer, that incorporates Geodesic Self-Attention (GSA) mechanisms. GSA utilizes geodesic distances to characterize interaction features effectively, complementing MHA, which excels in capturing local features and maintaining temporal correlations. FuseFormer employs a gating network to adaptively combine GSA and MHA embeddings, leveraging their complementary strengths. Additionally, FuseFormer integrates a Transformer-based Neural Ordinary Differential Equation (ODE) decoder to model trajectory temporal dynamics. This design enables the generation of future trajectories that align closely with motion trends while adapting the network depth to input sequence lengths. Experimental results demonstrate that FuseFormer achieves state-of-the-art performance across widely used pedestrian trajectory prediction datasets, including ETH/UCY, SDD, and NBA. These results underscore the model’s effectiveness and generalization capability in capturing complex interaction patterns and handling diverse scenarios.
Kohsin Ko, Jian Yang 0034, Ke Li 0005, Xiong You, Jinpeng Mi, Mingsong Chen 0001, Xian Wei
IEEE Trans. Intell. Transp. Syst.3
2024 Hyperbolic Graph Diffusion Model
abstract
Diffusion generative models (DMs) have achieved promising results in image and graph generation. However, real-world graphs, such as social networks, molecular graphs, and traffic graphs, generally share non-Euclidean topologies and hidden hierarchies. For example, the degree distributions of graphs are mostly power-law distributions. The current latent diffusion model embeds the hierarchical data in a Euclidean space, which leads to distortions and interferes with modeling the distribution. Instead, hyperbolic space has been found to be more suitable for capturing complex hierarchical structures due to its exponential growth property. In order to simultaneously utilize the data generation capabilities of diffusion models and the ability of hyperbolic embeddings to extract latent hierarchical distributions, we propose a novel graph generation method called, Hyperbolic Graph Diffusion Model (HGDM), which consists of an auto-encoder to encode nodes into successive hyperbolic embeddings, and a DM that operates in the hyperbolic latent space. HGDM captures the crucial graph structure distributions by constructing a hyperbolic potential node space that incorporates edge information. Extensive experiments show that HGDM achieves better performance in generic graph and molecule generation benchmarks, with a 48% improvement in the quality of graph generation with highly hierarchical structures.
Lingfeng Wen, Mingjie Ouyang, Xiangxiang Shen, Jian Yang 0034, Daxin Zhu, Mingsong Chen 0001, Xian Wei
AAAI5
2024 Social Lode: Human Trajectory Prediction with Latent Odes
abstract
Human trajectory prediction is crucial in human-computer interaction and even in the safety of autonomous driving. In this work, A new method, called Social Latent Ordinary Differential Equation (Social LODE), is introduced for predicting human trajectories. The backbone of Social LODE consists of a conditional Variational Autoencoder (VAE) architecture based on Recurrent Neural Network (RNN). The hidden state updated by RNN is often discrete, but the human trajectory is continuous and uncertain. Thus, we use Latent ODEs as the decoder of VAE to overcome the limitation of RNN. Finally, we demonstrate that Social LODE achieves state-of-the-art compared to other methods, such as those involving the ETH/UCY and SDD datasets.
Kexin Ke, Jian Yang 0034, Mingsong Chen 0001, Xian Wei
ICASSP2
2024 ProEqBEV: Product Group Equivariant BEV Network for 3D Object Detection in Road Scenes of Autonomous Driving
abstract
With the rapid development of autonomous driving systems, 3D object detection based on Bird’s Eye View (BEV) in road scenes has witnessed great progress over the past few years. As a road scene exhibits a part-whole hierarchy between the within objects and the scene itself, simple parts (e.g., roads, lane lines, vehicles and pedestrians) can be assembled into progressively more complex shapes to form a BEV representation of the whole road scene. Therefore, a BEV often has multiple levels of freedom on motion, i.e., the rotation and the moving shift of the whole BEV, and the random movements of objects (e.g., pedestrians and vehicles) inside the BEV. However, most of the current single-sensor or multi-sensor fusion-based BEV object detection methods have not yet taken into account capturing such multi-level motion in a BEV. To address this problem, we propose a product group equivariant object detection network framework that is equivariant with respect to multiple levels of symmetry groups based on multi-sensor fusion. The proposed framework extracts local equivariant features of objects in point clouds, while global equivariant features are extracted in both point clouds and images. Furthermore, the network learns diverse rotation-equivariant features and mitigates a significant amount of detection errors caused by rotations of BEV and objects inside a BEV, thereby further enhancing the performance of object detection. The experiment results show that the network architecture significantly improves object detection on mAP and NDS, respectively. In addition, in order to demonstrate the effectiveness of the proposed local-multi-global equivariant components, we conduct sufficient ablation experiments. The results show that the individual components are indispensable for the object detection performance improvement of the overall network architecture.
Jian Yang 0034, Ke Li 0005, Jianzhang Zheng, Xihao Wang, Mingsong Chen 0001, Xiong You, Xian Wei
ICRA2
2024 Continuous Geodesic Self-Attention Models with Gated Fusion for Trajectory Prediction
abstract
Driven by the rapid development of intelligent driving vehicles, predicting the trajectories of pedestrians on the road is crucial for decision-making during driving and even road safety. In this paper, we propose a novel method for trajectory prediction, namely, Continuous Geodesic Self-Attention Models with Gated Fusion (CGSAG). We use geodesic attention to measure the similarity between trajectory points, and utilize a gating mechanism to fuse the geodesic features extracted by multi-layer graph convolution. We then use Neural Ordinary Differential Equations (Neural ODE) to model the continuous-time dynamics of the trajectory. We show that CGSAG improves state-of-the-art performances on several human trajectory prediction datasets, including ETH/UCY, SDD, and Ind. At the same time, we conduct ablation studies to prove the effectiveness and efficiency of our proposed method.
Kexin Ke, Huining Chen, Xian Wei, Jian Yang 0034
IJCNN6
2023 RECO: Rotation Equivariant COnvolutional Neural Network for Human Trajectory Forecasting
Jijun Cheng, Dongheng Shao, Jian Yang 0034, Mingsong Chen 0001, Xian Wei
PRCV (3)4
2023 Continual Learning via Manifold Expansion Replay
abstract
In continual learning, the learner learns multiple tasks in sequence, with data being acquired only once for each task. Catastrophic forgetting is a major challenge to continual learning. To reduce forgetting, some existing rehearsal-based methods use episodic memory to replay samples of previous tasks. However, in the process of knowledge integration when learning a new task, this strategy also suffers from catastrophic forgetting due to an imbalance between old and new knowledge. To address this problem, we propose a novel replay strategy called Manifold Expansion Replay (MaER). We argue that expanding the implicit manifold of the knowledge representation in the episodic memory helps to improve the robustness and expressiveness of the model. To this end, we propose a greedy strategy to keep increasing the diameter of the implicit manifold represented by the knowledge in the buffer during memory management. In addition, we introduce Wasserstein distance instead of cross entropy as distillation loss to preserve previous knowledge. With extensive experimental validation on MNIST, CIFAR10, CIFAR100, and TinyImageNet, we show that the proposed method significantly improves the accuracy in continual learning setup, outperforming the state of the arts.
Zihao Xu 0002, Jian Yang 0034, Mingsong Chen 0001, Xian Wei
SMC5