Xiaobo Chen 0001

dblp:21/4778-1 · DBLP profile ↗
← Back
58ranked-venue papers
28as first author
30since 2021 · last 2026
0000-0001-9940-1637ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 15 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Egocentric-view pedestrian crossing intention prediction with limited observation: An approach based on knowledge distillation and feature decoupling
Wei Xu 0052, Xiaobo Chen 0001, Fuwen Deng
Knowl. Based Syst.2
2026 TIME: Trajectory and Interaction-based Memory Enhancement network for multi-agent trajectory prediction
Xiangzheng Zhou, Xiaobo Chen 0001, Jian Yang 0003
Pattern Recognit.2
2026 Learning From Past and Future: A Unified Instantaneous Pedestrian Intent Prediction Framework Based on Privileged Knowledge Distillation for Autonomous Driving
Xiaobo Chen 0001, Wei Xu 0052, Jianjun Qian
IEEE Trans Autom. Sci. Eng.1
2026 Evidential Multimodal Fusion Network for Trusted Pedestrian Crossing Intent Prediction
abstract
Accurate prediction of pedestrians’ behavior poses formidable challenges for autonomous vehicles in urban environments. Multimodal data, such as pedestrians’ motion data, context images, and ego vehicle speed, offer complementary and comprehensive information that can significantly enhance prediction performance. However, the previous methods, despite yielding promising results, integrate different modalities to form a uniform representation, which falls short of fully exploiting the heterogeneity and complementarity of all modalities. Besides, the uncertainty inherent in predictions is also a major concern for safety-critical systems such as autonomous vehicles. In light of the above concerns, this study puts forward a novel evidential multimodal fusion network called EMFNet, which leverages multimodal data and evidence fusion techniques for trusted pedestrian crossing intention prediction. Specifically, modal-specific embedding is first developed to project raw multimodal data into high-dimensional space by considering the unique properties of each modality. Then, we devise intra-modal and cross-modal feature learning to capture the temporal correlation within each modality and the interaction across different modalities, respectively. By doing so, modal-invariant features and modal-specific features can be effectively extracted. Subsequently, we introduce the Dempster–Shafer’s evidence theory (DST) to amalgamate evidence associated with different modalities, thus allowing the model to estimate the uncertainty and achieve trusted crossing prediction. Finally, we design an adaptive multiloss function that can effectively supervise the learning of modal-invariant and model-specific and facilitate the evidence fusion process. Experimental evaluation on real-world benchmark datasets demonstrates the improved prediction performance and enhanced reliability of the proposed method.
Xiaobo Chen 0001, Wei Xu 0052, Lei Yang 0048, Jian Yang 0003
IEEE Trans. Comput. Soc. Syst.2
2026 Learning Robust Discriminant Projections via Double Capped Lp-Norm Distance Metrics With "Min" Constraints
abstract
Recently, there has been a surge in the development of robust norm distance-based linear discriminant analysis (LDA) techniques, which have garnered significant attention in the field of feature extraction. However, a persistent issue that has yet to be resolved is that the successful suppression of outliers may inadvertently impede the accurate discrimination of normal points. To solve this problem, we, in this article, study a novel robust LDA measured by double capped $L_{p}$ -norm distance (CLD) metrics with min constraints (DCLDA) to learn robust discriminant projections, in which normal points and outliers are separately treated. To be specific, it takes a double capped $L_{p}$ -norm with "Min" constraints in the proposed model to measure the distances for between- and within-class dispersions. The proposed model effectively ensures accurate discrimination of normal points by $L_{p}$ -norm, while also eliminating the exaggerated effect of outliers that may arise from larger $p$ values. The resulted objective is not trivial because of its nonconvexity and nonsmoothness. As one of the major contributions of this article, we introduce a new reformulation that provides an objective problem theoretically equivalent to the original. By this reformulation, we develop an effective iterative algorithm to solve the proposed model. The algorithm is proven to be convergent through rigorous theoretical analysis. Extensive experiments were conducted on several real-world datasets across different image classification tasks to showcase the effectiveness of the proposed method.
Xiaobo Chen 0001, Zhao Zhang 0001, Liyong Fu, Qiaolin Ye
IEEE Trans. Neural Networks Learn. Syst.2
2025 Diff-Refiner: Enhancing Multi-Agent Trajectory Prediction with a Plug-and-Play Diffusion Refiner
abstract
The inherent stochasticity of the agents' behavior presents a challenge to trajectory prediction models, which are required to generate multiple plausible future trajectories. Recently, diffusion models have been applied to implement multimodal trajectory prediction. Existing approaches typically employ a standard diffusion process, denoising from a sample drawn from a Gaussian distribution. However, we identify that most agents exhibit an obvious movement trend, rendering many initial denoising steps redundant-primarily transitioning from pure noise to an initial coarse trajectory. To conquer this challenge, this paper innovatively proposes a diffusion refiner that can be used along with existing multi-agent trajectory prediction models to improve their performance. Specifically, we first leverage a baseline model for predicting the coarse future trajectory. Then, the diffusion model is applied as a refiner to reduce the prediction error. Moreover, our method is naturally plug-and-play, allowing convenient integration with existing models. To achieve this, we improve the traditional diffusion process to not only converge towards noise but also the coarse predictions from the baseline model. In such a case, standard step-skipping sampling techniques is inapplicable and we further propose an ordinary differential equation (ODE)-based fast sampling method. Extensive experiments with selected baseline models demonstrate the effectiveness of our approach.
Xiangzheng Zhou, Xiaobo Chen 0001, Jian Yang 0003
ICRA2
2025 A hypergraph-based dual-path multi-agent trajectory prediction model with topology inferring
Yu Hu 0010, Xiaobo Chen 0001, Yongjie Zhou, Jun Liang 0004
Eng. Appl. Artif. Intell.2
2025 Heterogeneous hypergraph transformer network with cross-modal future interaction for multi-agent trajectory prediction
Xiangzheng Zhou, Xiaobo Chen 0001, Jian Yang 0003
Eng. Appl. Artif. Intell.2
2025 Diversified Distillation Fusion Network for vehicle re-identification
Huaming Zhang, Xiaobo Chen 0001, Haoze Yu, Kok Lay Teo
Expert Syst. Appl.2
2025 Learning a multi-cluster memory prototype for unsupervised video anomaly detection
Yuntao Wu, Zhonghua Peng, Xiaobo Chen 0001
Inf. Sci.5
2025 Adaptive graph transformer with future interaction modeling for multi-agent trajectory prediction
Xiaobo Chen 0001, Fuwen Deng
Knowl. Based Syst.1
2025 Traffic Agents Trajectory Prediction Based on Enhanced Bidirectional Recurrent Network and Adaptive Social Interaction Model
abstract
Accurate prediction of the future trajectory of traffic agents is imperative to the effective motion planning of autonomous vehicles and mobile robots. Despite enormous progress that has been made toward trajectory prediction, dynamic and crowded traffic scenarios pose major challenges to the understanding and forecasting of traffic agents’ motion behavior. In this paper, we propose a novel trajectory prediction method from the perspective of temporal modeling and social interaction. Specifically, we first put forward a recurrent modeling approach to learn temporal features in favor of capturing long-range and short-range temporal dependencies of individual agents. Then, we construct a social feature learning module to capture the sparse and directional interactions among agents while suppressing the spurious connections. Finally, to reduce the accumulated error during prediction, a coordinated bidirectional decoding module is developed where temporal and social features can be properly integrated into the forward and backward prediction processes. Extensive experiments are performed on four real-world trajectory prediction benchmarks, and the results demonstrate the superiority of our method compared with other competing approaches. Detailed ablation studies are also performed to evaluate the effectiveness of each model component. Note to Practitioners—Motion planning is one of the crucial components of autonomous systems, such as intelligent vehicles and mobile robots. For example, the safety and efficiency of motion planning can be drastically improved if the future trajectories of surrounding agents, e.g., pedestrians, bicyclists, cars, etc., can be accurately forecasted. Motivated by the above requirements, this article develops an advanced deep learning model that can learn temporal and social features from trajectory data and perform accuracy prediction. This work aims to enhance the bidirectional recurrent network for dealing with trajectory data with evident temporal characteristics. In addition, this work introduces a novel adaptive social interaction modeling approach that overcomes the inherent defect of the fixed threshold method. This work also addresses the error accumulation problem in the prediction. The proposed model is evaluated on several datasets, and the results demonstrate its effectiveness. Our approach has broad application prospects in autonomous driving and mobile robots.
Xiaobo Chen 0001, Yuwen Liang, Chuan Hu 0003, Hai Wang 0003, Qiaolin Ye
IEEE Trans Autom. Sci. Eng.1
2025 Multibranch Attentive Transformer With Joint Temporal and Social Correlations for Traffic Agents Trajectory Prediction
abstract
Accurately predicting the future trajectories of traffic agents is paramount for autonomous unmanned systems, such as self-driving cars and mobile robotics. Extracting abundant temporal and social features from trajectory data and integrating the resulting features effectively pose great challenges for predictive models. To address these issues, this article proposes a novel multibranch attentive transformer (MBAT) trajectory prediction network for traffic agents. Specifically, to explore and reveal diverse correlations of agents, we propose a decoupled temporal and spatial feature learning module with multibranch to extract temporal, spatial, as well as spatiotemporal features. Such design ensures each branch can be specifically tailored for different types of correlations, thus enhancing the flexibility and representation ability of features. Besides, we put forward an attentive transformer architecture that simultaneously models the complex correlations possibly occurring in historical and future timesteps. Moreover, the temporal, spatial, and spatiotemporal features can be effectively integrated based on different types of attention mechanisms. Empirical results demonstrate that our model achieves outstanding performance on public ETH, UCY, SDD, and INTERACTION datasets. Detailed ablation studies are conducted to verify the effectiveness of the model components.
Xiaobo Chen 0001, Yuwen Liang, Qiaolin Ye, Yingfeng Cai
IEEE Trans. Comput. Soc. Syst.1
2025 Robust Multiple Flat Projections Clustering With Truncated Distance Maximization Constraints
abstract
Recently, interest in flat-type projection clustering methods has grown as they improve learner's performance by exploring multiple projection subspaces. However, solvers used in previous representative works predominantly rely on greedy search strategies, which incur high computational costs and fail to consider interdependencies between projections. Moreover, these methods do not simultaneously guarantee the effective suppression of outliers and noisy data at cluster boundaries, ultimately compromising data discrimination. To address these limitations and discover a more effective subspace for each flat, we propose robust multiple flat projections clustering (RMFPC). This method computes within- and between-cluster distances using the L2,1-norm to enhance robustness against outliers. Furthermore, we propose a truncated distance maximization constraint (TDMC) to eliminate the influence of noisy data on cluster separability. The resulting objective is presented in a ratio form, which is not trivial. We provide a novel formulation to achieve a theoretically equivalent problem. Based on this reformulation, we develop an efficient non-greedy solution algorithm. In addition, a cluster center optimization mechanism is incorporated into the solution process to accurately estimate the distribution of each cluster center. The convergence analysis and proof of the proposed algorithm are provided. Experiments on both toy and real-world datasets demonstrate the effectiveness of the proposed method.
Zhao Zhang 0001, Xiaobo Chen 0001, Zhongqi Xu, Liyong Fu, Qiaolin Ye
IEEE Trans. Cybern.3
2025 Pedestrian Crossing Intention Prediction via Progressive Multimodal Token Fusion for Autonomous Driving
abstract
Pedestrians’ intention to cross the street exercises a substantial influence on the decision-making process of autonomous vehicles in urban traffic environments. However, accurately predicting pedestrian crossing intention is non-trivial due to the interweaving of pedestrian personalities and traffic scene elements. Despite the significant achievement of previous studies, challenges remain in effectively extracting and integrating diverse features from different modalities of observation data. In response, this paper proposes a novel model leveraging pedestrian bounding boxes, poses, and ego-vehicle speed to predict crossing intention. We introduce mixture expert feature embedding (MEFE) to project raw data into high-dimensional space based on different types of inputs. A multi-branch spatial and temporal graph convolutional network (MB-STGCN) is applied to capture multi-scale spatial and temporal features of pedestrian pose skeleton joints. A multi-token temporal aggregation (MTTA) method is devised to preserve abundant temporal information of observation data. Additionally, a progressive multimodal feature fusion (PMFF) method based on symmetric channel split attention (SCSA) is employed to enhance the interaction between different modalities when integrating different features. Extensive comparison experiments and ablation studies on public benchmark datasets substantiate the effectiveness of our approach, showing significant improvements over existing models.
Xiaobo Chen 0001, Wei Xu 0052, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.1
2025 Nonconvex Transform-Based Low-Rank Tensor Completion With Coupled Spatiotemporal Relation Learning for Traffic Data Recovery
abstract
With the rapid development of sensor technology, Intelligent Transportation Systems (ITS) are capable of collecting vast amounts of traffic data. However, unforeseen interruptions during the process of data collection, transmission, and storage often lead to data loss, posing significant challenges to data accuracy and integrity. To address this issue, we have conducted an in-depth analysis of the unique physical characteristics of traffic data and optimized model design based on these characteristics to improve the accuracy of data recovery. This article introduces an innovative low-rank tensor completion model that leverages both global and local features of traffic data to accurately fill in missing values. Specifically, we propose using a weighted composite tensor (WCT) norm as a non-convex alternative to capture the multi-dimensional low-rank properties of traffic data in the transform domain. Additionally, to further enhance the precision of data recovery, we introduce a coupled structure regression method that combines temporal smoothness with sample similarity, aiding in revealing complex spatiotemporal data correlation patterns. To solve the resulting non-convex optimization problem, we have designed an efficient iterative algorithm based on the Alternating Direction Method of Multipliers (ADMM) and conducted a detailed theoretical analysis of its convergence and computational complexity. Experimental results show that compared to other competing algorithms, our model demonstrates significant advantages across three real-world traffic datasets, proving its superior performance in data recovery.
Xiaobo Chen 0001, Qiaolin Ye
IEEE Trans. Intell. Transp. Syst.2
2025 Edge-Enhanced Heterogeneous Graph Transformer With Priority-Based Feature Aggregation for Multi-Agent Trajectory Prediction
abstract
Trajectory prediction, which aims to predict the future positions of all agents in a crowd scene, given their past trajectories, plays a vital role in improving the safety of autonomous driving vehicles. For heterogeneous agents, it is imperative to account for the gap in feature distribution differences between agents in different categories. Besides, exploring the reference relationship between the future motions of agents is crucial yet overlooked in previous trajectory prediction methods. To tackle these challenges, we propose an edge-enhanced heterogeneous graph Transformer with priority-based feature aggregation for multi-modal trajectory prediction. Specifically, a new edge-enhanced heterogeneous interaction module that carries relative position information via edges is proposed to explore the complex interaction among agents. Additionally, we propose the concept of priority during the decoding phase and the corresponding measuring method, based on which a priority-based feature aggregation module is presented to enable referencing between agents, allowing for a more reasonable trajectory generation process. Additionally, we design an effective feature fusion method based on state refinement LSTM so that temporal and social features can be well integrated while accounting for their roles in trajectory prediction. Extensive experimental results on public datasets demonstrate that our approach outperforms the state-of-the-art baseline methods, confirming the effectiveness of our proposed method. The source code of our EPHGT model will be publicly released athttps://github.com/xbchen82/EPHGT.
Xiangzheng Zhou, Xiaobo Chen 0001, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.2
2025 Deformable Cross-Attention Transformer for Weakly Aligned RGB-T Pedestrian Detection
abstract
Pedestrian detection plays a crucial role in autonomous driving systems. To ensure reliable and effective detection in challenging conditions, researchers have proposed RGB–T (RGB–thermal) detectors that integrate thermal images with color images for more complementary feature representations. However, existing methods face challenges in capturing the spatial and geometric correlations between different modalities, as well as in assuming perfect synchronization of the two modalities, which is unrealistic in real-world scenarios. In response to these challenges, we present a new deformable-attention-based approach for weakly aligned RGB–T pedestrian detection. The proposed method uses a dual-branch cross-attention mechanism to capture the inherent spatial and geometric correlations between color and thermal images. Furthermore, it incorporates positional information for each image pixel into the sampling offset generation to enhance robustness in scenarios where modalities are not precisely aligned or registered. To reduce computational complexity, we introduce a local attention mechanism that samples only a small set of keys and values within a limited region in the feature maps for each query. Extensive experiments and ablation studies conducted on multiple public datasets confirm the effectiveness of the proposed framework.
Yu Hu 0010, Xiaobo Chen 0001, Hengyang Shi, Lihong Fan, Jun Liang 0004
IEEE Trans. Multim.2
2024 Unsupervised Cross-Scenario Abnormal Driving Behavior Recognition Using Smartphone Sensor Data
abstract
Accurately recognizing abnormal behavior of drivers (e.g., aggressive driving and fatigued driving) based on multivariate sensor data is vital for human-centric assistive driving systems. Existing data-driven deep learning models for abnormal driving behavior recognition (ADBR) achieve promising performance under specific driving scenes with sufficient labeled data. However, in the real world, dynamic driving scenes and unlabeled data pose a great challenge to the adaptability of models. In light of this, we put forward a novel unsupervised cross-scenario ADBR approach that can transfer domain knowledge in the source scenario with labeled data to the target scenario with only unlabeled data, thus considerably enhancing the adaptability of our model. Specifically, we first propose a feature extraction module that can obtain domain-shared and domain-specific features from raw sensor data derived from different driving scenes. Then, adversarial learning is presented to align the feature distribution of source and target domains to reduce the domain shift. A self-training strategy is further developed to boost the target domain classification performance by iteratively using the pseudo labels. Moreover, prediction uncertainty and ensemble classification are proposed to enhance the quality of pseudo labels. Extensive experiments on cross-scenario ADBR are conducted to evaluate the effectiveness of our model. The results manifest that our model significantly improves the recognition performance for the target domain and outperforms the competing algorithms.
Xiaobo Chen 0001, Yong Wang 0059, Xiaodong Sun 0001, Yingfeng Cai
IEEE Internet Things J.1
2024 Rectifying inaccurate unsupervised learning for robust time series anomaly detection
Zejian Chen, Xiaobo Chen 0001, Haoyi Fan
Inf. Sci.4
2024 Composite Nonconvex Low-Rank Tensor Completion With Joint Structural Regression for Traffic Sensor Networks Data Recovery
abstract
Traffic sensor networks allow convenient collection of travel data that are of great significance for intelligent transportation systems (ITSs). However, the universality of missing data impedes the application of ITS and thus accurate missing data recovery is indispensable in practice. Typically, the global low-rankness and local spatiotemporal smoothness exist in underlying traffic tensor data. In light of this, this article proposes an improved low-rank tensor completion (LRTC) model by exploiting abundant structural information from incomplete tensors. Specifically, a logarithm power composite (LPC)-norm is first proposed as a nonconvex substitute of the rank function, leading to a flexible characterization of tensor multidimensional correlation. Then, a joint structural regression (JSR) model is presented to simultaneously leverage the intrinsic temporal continuity and profile similarity of traffic data. By doing so, we construct a novel nonconvex LRTC model by integrating the global low-rankness and fine-grained spatiotemporal structure that are complementary to each other. To solve the proposed model, following the optimization framework of the alternating direction method of multipliers (ADMMs), we develop an efficient iterative algorithm where each step can be solved in a closed form. Extensive experiments on four real-world traffic data are conducted to evaluate the effectiveness of the proposed approach. The results demonstrate that compared with other tensor completion methods, our model significantly improves the recovery performance.
Xiaobo Chen 0001, Feng Zhao 0006, Fuwen Deng, Qiaolin Ye
IEEE Trans. Comput. Soc. Syst.1
2024 Stochastic Non-Autoregressive Transformer-Based Multi-Modal Pedestrian Trajectory Prediction for Intelligent Vehicles
abstract
Pedestrian trajectory prediction, which aims at predicting the future positions of all pedestrians in a crowd scene given their past trajectories, is the cornerstone of autonomous driving and intelligent transportation systems. Accurate prediction and fast inference are both indispensable for real-world applications. In this paper, we propose a stochastic non-autoregressive Transformer-based multi-modal trajectory prediction model to address the two challenges. Specifically, a novel graph attention module dedicated to joint learning of social and temporal interaction is proposed to explore the complex interaction among pedestrians while integrating sparse attention mechanism, pedestrian identity, and temporal order contained in the trajectory data. By doing so, the interaction across temporal and social dimensions can be simultaneously processed to extract abundant context features for prediction. Besides, to accelerate inference speed, we put forward a stochastic non-autoregressive Transformer model with multi-modal prediction capability where each future trajectory can be inferred in a parallel fashion, therefore, resulting in diverse trajectory predictions and less computational cost. Extensive experiments and ablation studies are performed to evaluate our approach. The empirical results demonstrate that the proposed model not only produces high prediction accuracy but also infers with fast speed. The code of the proposed method will be publicly available at https://github.com/xbchen82/SNARTF.
Xiaobo Chen 0001, Huanjia Zhang, Fuwen Deng, Jun Liang 0004, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.1
2024 Pedestrian Crossing Intention Prediction Based on Cross-Modal Transformer and Uncertainty-Aware Multi-Task Learning for Autonomous Driving
abstract
Accurate prediction of whether pedestrians will cross the street is prevalently recognized as an indispensable function of autonomous driving systems, especially in urban environments. How to utilize the complementary information present in different types of data (or modalities) is one of the major challenges. This paper makes the first attempt to develop a cross-modal transformer-based crossing intention prediction model merely using bounding boxes and ego-vehicle speed as input features. The cross-modal transformer can leverage self-attention and cross-modal attention to mine the modality-specific and complementary correlation. A bottleneck feature fusion is presented to obtain the compressed feature representation. To facilitate the network training, we further put forward a novel uncertainty-aware multi-task learning method that jointly predicts the future bounding box as well as crossing action such that the commonalities and differences across two tasks can be exploited. To evaluate the proposed method, extensive comparative experiments and ablation studies are performed on two benchmark datasets. The results demonstrate that by only using the bounding box and ego-vehicle speed as input features, our model is on a par with other state-of-the-art approaches that rely on more inputs, and even achieves superior performance in most cases. The source code will be released at https://github.com/xbchen82/PedCMT.
Xiaobo Chen 0001, Jun Li 0027, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.1
2023 Driving Style Feature Extraction and Recognition Based on Hyperdimensional Computing and Semi-Supervised Twin Projection Vector Machine
abstract
Driving style recognition is one of the most crucial requirements for human-centric autonomous and assistive driving systems. Existing studies either require a large amount of labeled data or suffer from hand-crafted features, thus hindering the practical application of these methods. To address the above challenges, this paper proposes a driving style recognition approach from the perspective of brain-inspired hyperdimensional computing and semi-supervised learning. Specifically, raw sensor signals are treated as multivariate time series which are encoded by hyperdimensional computing, thus generating a large holistic feature representation. Then, considering the high dimension of the resulting representation, we put forward a semi-supervised twin projection vector machine (SSTPVM) model that can take full advantage of unlabeled data and jointly optimize multiple projection directions tailored for dimension reduction. In addition, a heuristic method based on particle swarm optimization is developed for the parameter selection of SSTPVM. Finally, extensive experiment comparisons with other related methods are performed on naturalistic driving style data. The results show that hyperdimensional computing is rather suitable for semi-supervised learning while our proposed SSTPVM can significantly improve recognition performance even with a small portion of labeled data.
Xiaobo Chen 0001, Yuxiang Gao, Haoze Yu, Hai Wang 0003, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.1
2022 Intention-aware Transformer with Adaptive Social and Temporal Learning for Vehicle Trajectory Prediction
abstract
Accurate prediction of the trajectory of surrounding vehicles is crucial to autonomous driving for path planning and collision avoidance. In this paper, we propose a novel transformer-based model with adaptive social and temporal learning for trajectory prediction. In order to model social interaction between vehicles at each historical timestamp, an enhanced graph attention feature aggregation mechanism combing hidden feature and explicit relative spatial relation is developed. Further, the social and temporal dependency across different timestamps is captured by multi-head self-attention with an extra learnable "intention token". To achieve multi-modal trajectory prediction, we implement intention-aware transformer decoder of driving behavior, the intention recognition and trajectory. Experiments on large-scale benchmark datasets verify that our model achieves better performance comparing with some state-of-the-art trajectory prediction models.
Yu Hu 0010, Xiaobo Chen 0001
ICPR2
2022 A Novel Spatiotemporal Data Low-Rank Imputation Approach for Traffic Sensor Network
abstract
The Internet of Things (IoT) has enormous potential to transform the transport industry by improving passenger experiences, safety, and efficiency. However, the collected spatiotemporal data by traffic sensor network often suffer from missing values (MVs), which affect the overall performance of the system. As a result, accurate recovery of MVs is essential for the successful application of IoT in transportation. In this article, we propose a novel MVs imputation model by integrating low-rank tensor completion (LRTC) and sparse self-representation into a unified framework. In doing so, the global multidimensional correlation, as well as sample self-similarity, can be well leveraged for imputation. In order to solve the proposed model, an elaborate solution algorithm is developed, following the principle of alternating direction method of multipliers (ADMMs). Importantly, each step in ADMM can be implemented efficiently by analyzing the problem structure. Moreover, in order to select proper parameters for the model, an improved harmony search heuristic algorithm based on dual harmonies generation strategy is developed, thus sufficiently considering the information contained in current harmony memory. The experiments on two real-world traffic data are carried out to evaluate the proposed approach. The results verify that in comparison with the classic matrix/tensor completion and other competing algorithms, our method significantly improves the imputation performance.
Xiaobo Chen 0001, Shurong Liang, Feng Zhao 0006
IEEE Internet Things J.1
2022 Multiview Feature Learning With Multiatlas-Based Functional Connectivity Networks for MCI Diagnosis
abstract
Functional connectivity (FC) networks built from resting-state functional magnetic resonance imaging (rs-fMRI) has shown promising results for the diagnosis of Alzheimer's disease and its prodromal stage, that is, mild cognitive impairment (MCI). FC is usually estimated as a temporal correlation of regional mean rs-fMRI signals between any pair of brain regions, and these regions are traditionally parcellated with a particular brain atlas. Most existing studies have adopted a predefined brain atlas for all subjects. However, the constructed FC networks inevitably ignore the potentially important subject-specific information, particularly, the subject-specific brain parcellation. Similar to the drawback of the "single view" (versus the "multiview" learning) in medical image-based classification, FC networks constructed based on a single atlas may not be sufficient to reveal the underlying complicated differences between normal controls and disease-affected patients due to the potential bias from that particular atlas. In this study, we propose a multiview feature learning method with multiatlas-based FC networks to improve MCI diagnosis. Specifically, a three-step transformation is implemented to generate multiple individually specified atlases from the standard automated anatomical labeling template, from which a set of atlas exemplars is selected. Multiple FC networks are constructed based on these preselected atlas exemplars, providing multiple views of the FC network-based feature representations for each subject. We then devise a multitask learning algorithm for joint feature selection from the constructed multiple FC networks. The selected features are jointly fed into a support vector machine classifier for multiatlas-based MCI diagnosis. Extensive experimental comparisons are carried out between the proposed method and other competing approaches, including the traditional single-atlas-based method. The results indicate that our method significantly improves the MCI classification, demonstrating its promise in the brain connectome-based individualized diagnosis of brain diseases.
Yu Zhang 0009, Han Zhang 0002, Ehsan Adeli-Mosabbeb, Xiaobo Chen 0001, Mingxia Liu 0001, Dinggang Shen
IEEE Trans. Cybern.4
2022 Intention-Aware Vehicle Trajectory Prediction Based on Spatial-Temporal Dynamic Attention Network for Internet of Vehicles
abstract
Vehicle trajectory prediction is a keystone for the application of the internet of vehicles (IoV). With the help of deep learning and big data, it is possible to understand the between-vehicle interaction pattern hidden in the complex traffic environment. In this paper, we propose a novel spatial-temporal dynamic attention network for vehicle trajectory prediction, which can comprehensively capture temporal and social patterns in a hierarchical manner. The social relation between vehicles is captured at each timestamp and thus retains the dynamic variation of interaction. The temporal correlation in terms of individual motion state as well as social interaction is captured by different sequential models. Furthermore, a driving intention-specific feature fusion mechanism is proposed such that the extracted temporal and social features can be integrated adaptively for the maneuver-based multi-modal trajectory prediction. Experimental results on two real-world datasets show that compared with the state-of-the-art algorithms, our proposal achieves comparable prediction performance for short-term prediction, however, works much better for long-term prediction. Additionally, various ablation analysis is provided to evaluate the effectiveness of our proposed network components. The code will be available athttps://xbchen82.github.io/resource/.
Xiaobo Chen 0001, Huanjia Zhang, Feng Zhao 0006, Yu Hu 0010, Chenkai Tan, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.1
2022 Learning Robust Discriminant Subspace Based on Joint L₂, ₚ- and L₂, ₛ-Norm Distance Metrics
abstract
-norm as the distance metric. However, both of their robustness and discriminant power are limited. In this article, we present a new robust discriminant subspace (RDS) learning method for feature extraction, with an objective function formulated in a different form. To guarantee the subspace to be robust and discriminative, we measure the within-class distances based on [Formula: see text]-norm and use [Formula: see text]-norm to measure the between-class distances. This also makes our method include rotational invariance. Since the proposed model involves both [Formula: see text]-norm maximization and [Formula: see text]-norm minimization, it is very challenging to solve. To address this problem, we present an efficient nongreedy iterative algorithm. Besides, motivated by trace ratio criterion, a mechanism of automatically balancing the contributions of different terms in our objective is found. RDS is very flexible, as it can be extended to other existing feature extraction techniques. An in-depth theoretical analysis of the algorithm's convergence is presented in this article. Experiments are conducted on several typical databases for image classification, and the promising results indicate the effectiveness of RDS.
Liyong Fu, Zechao Li, Qiaolin Ye, Qingwang Liu, Xiaobo Chen 0001, Xijian Fan, Wankou Yang, Guowei Yang 0002
IEEE Trans. Neural Networks Learn. Syst.6
2021 Geometric projection twin support vector machine for pattern classification
Xiaobo Chen 0001
Multim. Tools Appl.1
2020 Multi-Class ASD Classification Based on Functional Connectivity and Functional Correlation Tensor via Multi-Source Domain Adaptation and Multi-View Sparse Representation
abstract
The resting-state functional magnetic resonance imaging (rs-fMRI) reflects functional activity of brain regions by blood-oxygen-level dependent (BOLD) signals. Up to now, many computer-aided diagnosis methods based on rs-fMRI have been developed for Autism Spectrum Disorder (ASD). These methods are mostly the binary classification approaches to determine whether a subject is an ASD patient or not. However, the disease often consists of several sub-categories, which are complex and thus still confusing to many automatic classification methods. Besides, existing methods usually focus on the functional connectivity (FC) features in grey matter regions, which only account for a small portion of the rs-fMRI data. Recently, the possibility to reveal the connectivity information in the white matter regions of rs-fMRI has drawn high attention. To this end, we propose to use the patch-based functional correlation tensor (PBFCT) features extracted from rs-fMRI in white matter, in addition to the traditional FC features from gray matter, to develop a novel multi-class ASD diagnosis method in this work. Our method has two stages. Specifically, in the first stage of multi-source domain adaptation (MSDA), the source subjects belonging to multiple clinical centers (thus called as source domains) are all transformed into the same target feature space. Thus each subject in the target domain can be linearly reconstructed by the transformed subjects. In the second stage of multi-view sparse representation (MVSR), a multi-view classifier for multi-class ASD diagnosis is developed by jointly using both views of the FC and PBFCT features. The experimental results using the ABIDE dataset verify the effectiveness of our method, which is capable of accurately classifying each subject into a respective ASD sub-category.
Jun Wang 0024, Lichi Zhang, Qian Wang 0001, Lei Chen 0011, Jun Shi 0004, Xiaobo Chen 0001, Dinggang Shen
IEEE Trans. Medical Imaging6
2019 Strength and similarity guided group-level brain functional network construction for MCI diagnosis
Yu Zhang 0009, Han Zhang 0002, Xiaobo Chen 0001, Mingxia Liu 0001, Xiaofeng Zhu 0001, Seong-Whan Lee, Dinggang Shen
Pattern Recognit.3
2018 Multi-Layer Multi-View Classification for Alzheimer's Disease Diagnosis
abstract
In this paper, we propose a novel multi-view learning method for Alzheimer's Disease (AD) diagnosis, using neuroimaging and genetics data. Generally, there are several major challenges associated with traditional classification methods on multi-source imaging and genetics data. First, the correlation between the extracted imaging features and class labels is generally complex, which often makes the traditional linear models ineffective. Second, medical data may be collected from different sources (i.e., multiple modalities of neuroimaging data, clinical scores or genetics measurements), therefore, how to effectively exploit the complementarity among multiple views is of great importance. In this paper, we propose a Multi-Layer Multi-View Classification (ML-MVC) approach, which regards the multi-view input as the first layer, and constructs a latent representation to explore the complex correlation between the features and class labels. This captures the high-order complementarity among different views, as we exploit the underlying information with a low-rank tensor regularization. Intrinsically, our formulation elegantly explores the nonlinear correlation together with complementarity among different views, and thus improves the accuracy of classification. Finally, the minimization problem is solved by the Alternating Direction Method of Multipliers (ADMM). Experimental results on Alzheimer's Disease Neuroimaging Initiative (ADNI) data sets validate the effectiveness of our proposed method.
Changqing Zhang 0002, Ehsan Adeli-Mosabbeb, Tao Zhou 0002, Xiaobo Chen 0001, Dinggang Shen
AAAI4
2018 Graph regularized local self-representation for missing value imputation with applications to on-road traffic sensor data
Xiaobo Chen 0001, Yingfeng Cai, Qiaolin Ye, Lei Chen 0011
Neurocomputing1
2017 Ensemble correlation-based low-rank matrix completion with applications to traffic data imputation
Xiaobo Chen 0001, Zhongjie Wei, Jun Liang 0004, Yingfeng Cai, Bob Zhang 0001
Knowl. Based Syst.1
2017 Improved combined invariant moment for moving targets classification
Xiaojun Chen 0005, Jia Ke, Xiaobo Chen 0001, Qian-Qian Zhang, Xiao-Ming Jiang, Xin-Ping Song, Bao-Ding Chen
Multim. Tools Appl.4
2017 A cloud computing architecture for characterization and classification of moving object
Xiaojun Chen 0005, Jia Ke, Tianming Zhan, Wen-Xin Wang, Xiaobo Chen 0001, Xin-Ping Song
Multim. Tools Appl.6
2016 Ensemble Hierarchical High-Order Functional Connectivity Networks for MCI Classification
Xiaobo Chen 0001, Han Zhang 0002, Dinggang Shen
MICCAI (2)1
2016 Outcome Prediction for Patient with High-Grade Gliomas from Brain Functional and Structural Networks
Luyan Liu, Han Zhang 0002, Islem Rekik, Xiaobo Chen 0001, Qian Wang 0001, Dinggang Shen
MICCAI (2)4
2016 Correlation-Weighted Sparse Group Representation for Brain Network Construction in MCI Classification
Renping Yu, Han Zhang 0002, Xiaobo Chen 0001, Zhihui Wei, Dinggang Shen
MICCAI (1)4
2016 Multilevel framework to handle object occlusions for real-time tracking
abstract
This study proposes an efficient method to handle the object occlusions seen in monocular traffic image sequences. The motivation of this study is different methods perform differently in occlusion segmentation and the authors’ idea is to use a situation‐driven approach to aggregate different methods in order to get a good performance. This study classifies occlusion into four categories according to the foreground situation and a multilevel occlusion handling framework is utilised. First, the image segmentation algorithm based on convex hull analysis is utilised for intra‐frame level occlusion segmentation. The segmentation algorithm is established by the compactness ratio and interior distance ratio of the foreground. Second, an online sample‐based classification algorithm is utilised for tracking level occlusion segmentation. Training samples are extracted from the historical frames before occlusion and testing samples are extracted from the current frame by an adaptive searching strategy. The segmentation of occlusion is transferred into the online classification of testing samples. Such algorithm is established by the similarity and coherence of target's property between continuous frames. Experiments on video sequences illustrate the good performance of the proposed method under different conditions with low computational cost.
Yingfeng Cai, Hai Wang 0003, Xiaobo Chen 0001, Long Chen 0003
IET Image Process.3
2016 Complex video event detection via pairwise fusion of trajectory and multi-label hypergraphs
Xiaojun Chen 0005, Jia Ke, Xiaobo Chen 0001
Multim. Tools Appl.4
2016 Occluded vehicle detection with local connected deep model
Hai Wang 0003, Yingfeng Cai, Xiaobo Chen 0001, Long Chen 0003
Multim. Tools Appl.3
2015 Discriminant feature extraction for image recognition using complete robust maximum margin criterion
Xiaobo Chen 0001, Yingfeng Cai, Long Chen 0003
Mach. Vis. Appl.1
2014 An Improved Linear Discriminant Analysis with L1-Norm for Robust Feature Extraction
abstract
Feature extraction plays an important role in analyzing data with multivariate features. Linear discriminant analysis based on L1-norm (LDA-L1) is a recently developed technique for enhancing the robustness of the classic LDA against outliers. However, LDA-L1 employs a greedy strategy to find all the discriminant vectors, which may lead to suboptimal solution. To address this issue, we develop a novel algorithm termed as ILDA-L1 in this paper, which can optimize all the discriminant vectors simultaneously in a unified framework. Specifically, we introduce an orthonormal constraint on the discriminant vectors and convert the objective function of LDA-L1 into a difference formula. To solve the resulting nonconvex and nonsmooth problem, we first construct a successive concave approximation to the objective function at current solution and then use projected sub gradient method, thus leading to a convergent iterative algorithm. The experimental results on several benchmark datasets confirm the effectiveness of ILDA-L1 in extracting robust features.
Xiaobo Chen 0001, Jian Yang 0003, Zhong Jin
ICPR1
2014 Structural max-margin discriminant analysis for feature extraction
Xiaobo Chen 0001, Yinfeng Cai, Long Chen 0003
Knowl. Based Syst.1
2014 An improved robust and sparse twin support vector regression via linear programming
Xiaobo Chen 0001, Jian Yang 0003, Long Chen 0003
Soft Comput.1
2013 Regularized least squares fisher linear discriminant with applications to image recognition
Xiaobo Chen 0001, Jian Yang 0003, Qirong Mao, Fei Han 0001
Neurocomputing1
2013 Complete large margin linear discriminant analysis using mathematical programming approach
Xiaobo Chen 0001, Jian Yang 0003, David Zhang 0001, Jun Liang 0004
Pattern Recognit.1
2012 Large margin null space discriminant analysis with applications to face recognition
Xiaobo Chen 0001, Jian Yang 0003, Wankou Yang
ICPR1
2012 Recursive robust least squares support vector regression based on maximum correntropy criterion
Xiaobo Chen 0001, Jian Yang 0003, Jun Liang 0004, Qiaolin Ye
Neurocomputing1
2012 A flexible support vector machine for regression
Xiaobo Chen 0001, Jian Yang 0003, Jun Liang 0004
Neural Comput. Appl.1
2012 Smooth twin support vector regression
Xiaobo Chen 0001, Jian Yang 0003, Jun Liang 0004, Qiaolin Ye
Neural Comput. Appl.1
2012 Discriminant Kernel Learning Using Hybrid Regularization
Jun Liang 0004, Long Chen 0003, Xiaobo Chen 0001
Neural Process. Lett.3
2012 Recursive "concave-convex" Fisher Linear Discriminant with applications to face, handwritten digit and terrain recognition
Qiaolin Ye, Chunxia Zhao, Haofeng Zhang 0001, Xiaobo Chen 0001
Pattern Recognit.4
2011 Localized twin SVM via convex minimization
Qiaolin Ye, Chunxia Zhao, Xiaobo Chen 0001
Neurocomputing4
2011 Optimal Locality Regularized Least Squares Support Vector Machine via Alternating Optimization
Xiaobo Chen 0001, Jian Yang 0003, Jun Liang 0004
Neural Process. Lett.1
2011 Recursive projection twin support vector machine via within-class variance minimization
Xiaobo Chen 0001, Jian Yang 0003, Qiaolin Ye, Jun Liang 0004
Pattern Recognit.1