Beihao Xia

dblp:260/6382 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
26since 2021 · last 2026
0000-0001-6156-6946ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
abstract
Spatio-Temporal Video Grounding (STVG) aims to localize a spatio-temporal tube that corresponds to a given language query in an untrimmed video. This is a challenging task since it involves complex vision-language understanding and spatiotemporal reasoning. Recent works have explored weakly-supervised setting in STVG to eliminate reliance on fine-grained annotations like bounding boxes or temporal stamps. However, they typically follow a simple late-fusion manner, which generates tubes independent of the text description, often resulting in failed target identification and inconsistent target tracking. To address this limitation, we propose a Tube-conditioned Reconstruction with Mutual Constraints (TubeRMC) framework that generates text-conditioned candidate tubes with pre-trained visual grounding models and further refine them via tube-conditioned reconstruction with spatio-temporal constraints. Specifically, we design three reconstruction strategies from temporal, spatial, and spatio-temporal perspectives to comprehensively capture rich tube-text correspondences. Each strategy is equipped with a Tube-conditioned Reconstructor, utilizing spatio-temporal tubes as condition to reconstruct the key clues in the query. We further introduce mutual constraints between spatial and temporal proposals to enhance their quality for reconstruction. TubeRMC outperforms existing methods on two public benchmarks VidSTG and HCSTVG. Further visualization shows that TubeRMC effectively mitigates both target identification errors and inconsistent tracking.
Jinxuan Li, Jianfang Hu, Chaolei Tan, Tianming Liang, Beihao Xia
AAAI6
2026 DH-MSVM: A hybrid algorithm for seeking quality support vectors in distributed learning
Jiawen Gong, Beihao Xia, Qinmu Peng, Bin Zou 0002, Xinge You
Neural Networks2
2026 DHS-AE: A Distributed Support Vector Machine With Adaptive Regularization Parameters for Different Data Distributions
abstract
In distributed machine learning scenarios, the difference in data distribution among different nodes is a key issue that cannot be ignored. However, existing methods make it difficult to autonomously adjust model parameters for dynamically changing data distributions, leading to inflexible global decision boundaries with insufficient local adaptation. To address this problem, we propose a distributed hybrid support vector machine (SVM) based on the adaptive ensemble selection of regularization parameters, DHS-AE. The model utilizes the data structure information to cut the data space and thus identify data distribution characteristics. The SVM, integrated with regularization parameters that are adaptively determined within specific ranges, is utilized in the local subspace to enable real-time adjustment of decision boundaries in response to distribution changes, thereby further reducing the computational overhead. The generalization bound of DHS-AE is theoretically established using covering numbers, and the fast convergence speed and consistency are derived. In practical applications, we verify the excellent performance of the DHS-AE using a large number of real datasets.
Jiawen Gong, Beihao Xia, Qinmu Peng, Bin Zou 0002, Xinge You
IEEE Trans. Cybern.2
2026 Why Not Diversify Triggers? APK-Specific Backdoor Attack Against Android Malware Detection
abstract
Machine learning-based Android malware detection (AMD) models require abundant data for training robust app classifiers, creating vulnerability to poisoning attacks. Attackers inject poisoned samples into Android app markets, leading to the insertion of a backdoor into the model upon adoption in the training process. Subsequently, attackers can generate evasive malware by embedding a backdoor trigger in malware samples. Currently, research on backdoor attacks towards AMD has just begun to emerge. Existing attacks produce a fixed trigger and apply it to various malware. Once the trigger is discovered by static analysis methods (e.g., software similarity analysis), however, multiple malware carrying this trigger will be simultaneously exposed. To diversify the trigger, we propose anAPK-SpecificBackdoorAttack algorithm (ASBA), which trains a generative adversarial network to generate a specific trigger for every malware sample. Moreover, ASBA manages to make the generated triggers as different as possible, in order to further reduce the likelihood of malware being collectively captured. Extensive experiments have demonstrated that ASBA achieves a 94.6% average attack success rate (ASR) on three datasets, five feature extraction methods and three classification models. Furthermore, compared to state-of-the-art poisoning attack algorithms, ASBA produces more diverse and more effective triggers.
Heng Li 0008, Bang Wu 0002, Cuiying Gao, Wei Yuan 0001, Beihao Xia, Xiapu Luo
IEEE Trans. Dependable Secur. Comput.6
2026 DHL-FLD: A Distributed Hybrid Learning Based on Fisher Linear Discriminant for Data Classification
abstract
Distributed machine learning provides an efficient solution for large-scale data processing through parallel computing. However, current distributed learning relies on global or local paradigms and cannot adaptively adjust decision boundaries in complex data environments. To address this problem, we propose a Distributed Hybrid Learning algorithm based on Fisher Linear Discriminant (DHL-FLD). Specifically, DHL-FLD consists of a global pre-learning phase and a subspace local learning phase. On the one hand, the global pre-learning phase is designed to divide the data space, which can obtain the data structure information. On the other hand, the local learning phase dynamically adjusts and optimizes the decision boundaries, guided by the structural information and distributional properties of the data. Theoretically, we establish the generalization bound of DHL-FLD using the integral operator technique and verify the scalability and robustness of DHL-FLD. The effectiveness of DHL-FLD is demonstrated through extensive experiments on real datasets.
Jiawen Gong, Beihao Xia, Qinmu Peng, Bin Zou 0002, Xinge You
IEEE Trans. Knowl. Data Eng.2
2025 Resonance: Learning to Predict Social-Aware Pedestrian Trajectories as Co-Vibrations
abstract
Learning to forecast trajectories of intelligent agents has caught much more attention recently. However, it remains a challenge to accurately account for agents' intentions and social behaviors when forecasting, and in particular, to simulate the unique randomness within each of those components in an explainable and decoupled way. Inspired by vibration systems and their resonance properties, we propose the Resonance (short for Re) model to encode and forecast pedestrian trajectories in the form of ``co-vibrations''. It decomposes trajectory modifications and randomnesses into multiple vibration portions to simulate agents' reactions to each single cause, and forecasts trajectories as the superposition of these independent vibrations separately. Also, benefiting from such vibrations and their spectral properties, representations of social interactions can be learned by emulating the resonance phenomena, further enhancing its explainability. Experiments on multiple datasets have verified its usefulness both quantitatively and qualitatively.
Conghao Wong, Ziqian Zou, Beihao Xia
ICCV3
2025 A Multi-modal Hand Imitation Dataset for Dexterous Hand
abstract
Multimodal data is indispensable for advancing imitation learning, particularly in the context of dexterous hands. However, existing datasets predominantly rely on single-modality inputs, such as RGB images, which inherently lack the capacity to capture the spatial and temporal dynamics essential for achieving human-like dexterity. To address this limitation, we introduce Multi-Modal Dex, a dataset that integrates multimodal sensory data to enable the effective learning of dexterous skills from human demonstrations. By combining visual, point cloud, and kinematic modalities, our dataset provides a richer representation of hand interactions, thereby facilitating a more nuanced understanding of dexterous imitation. Our framework leverages neural rendering and kinematic optimization to align human and robotic hand poses in a shared canonical space, enabling geometrically consistent skill transfer. Furthermore, we analyze the dataset’s potential to advance dexterous robots in perception, imitation learning, and real-world dexterous skill transfer. The data is available at https://github.com/WangShaoSUN/MutliDex.
Beihao Xia
IROS6
2025 UAVD-Mamba: Deformable Token Fusion Vision Mamba for Multimodal UAV Detection
abstract
Unmanned Aerial Vehicle (UAV) object detection has been widely used in traffic management, agriculture, emergency rescue, etc. However, it faces significant challenges, including occlusions, small object sizes, and irregular shapes. These challenges highlight the necessity for a robust and efficient multimodal UAV object detection method. Mamba has demonstrated considerable potential in multimodal image fusion. Leveraging this, we propose UAVD-Mamba, a multimodal UAV object detection framework based on Mamba architectures. To improve geometric adaptability, we propose the Deformable Token Mamba Block (DTMB) to generate deformable tokens by incorporating adaptive patches from deformable convolutions alongside normal patches from normal convolutions, which serve as the inputs to the Mamba Block. To optimize the multimodal feature complementarity, we design two separate DTMBs for the RGB and infrared (IR) modalities, with the outputs from both DTMBs integrated into the Mamba Block for feature extraction and into the Fusion Mamba Block for feature fusion. Additionally, to improve multiscale object detection, especially for small objects, we stack four DTMBs at different scales to produce multiscale feature representations, which are then sent to the Detection Neck for Mamba (DNM). The DNM module, inspired by the YOLO series, includes modifications to the SPPF and C3K2 of YOLOv11 to better handle the multiscale features. In particular, we employ cross-enhanced spatial attention before the DTMB and cross-channel attention after the Fusion Mamba Block to extract more discriminative features. Experimental results on the DroneVehicle dataset show that our method outperforms the baseline OAFA method by 3.6% in the mAP metric. Codes will be released at https://github.com/GreatPlum-hnu/UAVD-Mamba.git.
Jiaman Tang, Yang Li 0093, Beihao Xia, Ligang Tan, Hongmao Qin
IV4
2025 LTMSformer: A Local Trend-Aware Attention and Motion State Encoding Transformer for Multi-Agent Trajectory Prediction
abstract
It has been challenging to model the complex temporal-spatial dependencies between agents for trajectory prediction. As each state of an agent is closely related to the states of adjacent time steps, capturing the local temporal dependency is beneficial for prediction, while most studies often overlook it. Besides, learning the high-order motion state attributes is expected to enhance spatial interaction modeling, but it is rarely seen in previous works. To address this, we propose a lightweight framework, i.e., LTMSformer, to extract temporal-spatial interaction features for multi-modal trajectory prediction. Specifically, we introduce a Local Trend-Aware Attention mechanism to capture the local temporal dependency by leveraging a convolutional attention mechanism with hierarchical local time boxes. Next, to model the spatial interaction dependency, we build a Motion State Encoder to incorporate high-order motion state attributes, such as acceleration, jerk, heading, etc. To further refine the trajectory prediction, we propose a Lightweight Proposal Refinement Module that leverages Multi-Layer Perceptrons for trajectory embedding and generates the refined trajectories with fewer model parameters. Experiment results on the Argoverse 1 dataset demonstrate that our method outperforms the baseline HiVT-64, reducing the minADE by approximately 4.35%, the minFDE by 8.74%, and the MR by 20%. We also achieve higher accuracy than HiVT-128 with a 68% reduction in model size.
Yixin Yan, Yang Li 0093, Yuanfan Wang, Beihao Xia, Manjiang Hu, Hongmao Qin
IV5
2025 Another Vertical View: A Hierarchical Network for Heterogeneous Trajectory Prediction via Spectrums
abstract
With the fast development of AI-related techniques, the applications of trajectory prediction are no longer limited to easier scenes and trajectories. More and more trajectories with different forms, such as coordinates, bounding boxes, and even high-dimensional human skeletons, need to be analyzed and forecasted. Among these heterogeneous trajectories, interactions between different elements within a frame of trajectory, which we call "Dimension-wise Interactions", would be more complex and challenging. However, most previous approaches focus mainly on a specific form of trajectories, and potential dimension-wise interactions are less concerned. In this work, we expand the trajectory prediction task by introducing the trajectory dimensionality $M$M, thus extending its application scenarios to heterogeneous trajectories. We first introduce the Haar transform as an alternative to the Fourier transform to better capture the time-frequency properties of each trajectory-dimension. Then, we adopt the bilinear structure to model and fuse two factors simultaneously, including the time-frequency response and the dimension-wise interaction, to forecast heterogeneous trajectories via trajectory spectrums hierarchically in a generic way. Experiments show that the proposed model outperforms most state-of-the-art methods on ETH-UCY, SDD, nuScenes, and Human3.6 M with heterogeneous trajectories, including 2D coordinates, 2D/3D bounding boxes, and 3D human skeletons.
Beihao Xia, Conghao Wong, Duanquan Xu, Qinmu Peng, Xinge You
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Decoupling Objectives for Segmented Path Planning: A Subtask-Oriented Trajectory Planning Approach
abstract
Local trajectory planning (TP) for collision avoidance typically comprises path planning (PP) and velocity planning (VP). Various objectives must be fulfilled in a PP task, and the majority of prior works integrate all objectives into a unified cost function. To prioritize the dominant objectives of each PP stage, we propose to decouple the PP task into two separate subtasks, enabling the logical establishment of subtask-oriented segmented PP methods. First, based on risk evaluation of four vehicle vertices, an improved artificial potential field was established. Second, a novel transit point selection method was applied to decouple the PP task into two segmented subtasks. Then, the optimization problem was converted into a multi-attribute decision-making (MADM) problem and the technique for order preference by similarity to ideal solution (TOPSIS) method was utilized to obtain two optimal segmented paths. Finally, a velocity planner based on cubic polynomial, in conjunction with a support vector machine-based stability classifier, was designed. The proposed trajectory planner was then verified in six typical driving scenarios, including both simulation and real vehicle studies. Verification results demonstrate that the proposed planner effectively decouples the PP task and achieves a safe, comfortable, efficient, and trackable trajectory.
Guangliang Liao, Chunyun Fu, Yinghong Yu, Kexue Lai, Beihao Xia, Jingkang Xia
IEEE Trans. Intell. Transp. Syst.5
2025 Exploring Spatio-Temporal Carbon Emission Across Passenger Car Trajectory Data
abstract
Carbon emissions caused by passenger cars in cities are essentially responsible for severe climate change and serious environmental problems. Exploring carbon emissions from passenger cars helps to control urban pollution and achieve urban sustainability. However, it is a challenging task to foresee the spatio-temporal distribution of carbon emission from passenger cars, as the following technical issues remain. i) Vehicle carbon emissions contain complex spatial interactions and temporal dynamics. How to collaboratively integrate such spatial-temporal correlations for carbon emission prediction is not yet resolved. ii) Given the mobility of passenger cars, the hidden dependencies inherent in traffic density are not properly addressed in predicting carbon emissions from passenger cars. To tackle these issues, we propose a Collaborative Spatial-temporal Network (CSTNet) for implementing carbon emissions prediction by using passenger car trajectory data. Within the proposed method, we devote to extract collaborative properties that stem from a multi-view graph structure together with parallel input of carbon emission and traffic density. Then, we design a spatial-temporal convolutional block for both carbon emission and traffic density, which constitutes of temporal gate convolution, spatial convolution and temporal attention mechanism. Following that, an interaction layer between carbon emission and traffic density is proposed to handle their internal dependencies, and further model spatial relationships between the features. Besides, we identify several global factors and embed them for final prediction with a collaborative fusion. Experimental results on the real-world passenger car trajectory dataset demonstrate that the proposed method outperforms the baselines with a roughly 7%-11% improvement.
Zhu Xiao, Bo Liu 0104, Linshan Wu, Hongbo Jiang 0001, Beihao Xia, Tao Li 0056, Cassandra C. Wang
IEEE Trans. Intell. Transp. Syst.5
2024 Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
abstract
This paper focuses on open-ended video question answering, which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task, since a question may have multiple answers. However, due to annotation costs, the labels in existing benchmarks are always extremely insufficient, typically one answer per question. As a result, existing works tend to directly treat all the unlabeled answers as negative labels, leading to limited ability for generalization. In this work, we introduce a simple yet effective ranking distillation framework (RADI) to mitigate this problem without additional manual annotation. RADI employs a teacher model trained with incomplete labels to generate rankings for potential answers, which contain rich knowledge about label priority as well as label-associated visual cues, thereby enriching the insufficient labeling information. To avoid overconfidence in the imperfect teacher model, we further present two robust and parameter-free ranking distillation approaches: a pairwise approach which introduces adaptive soft margins to dynamically refine the optimization constraints on various pairwise rankings, and a listwise approach which adopts sampling-based partial listwise learning to resist the bias in teacher ranking. Extensive experiments on five popular benchmarks consistently show that both our pairwise and listwise RADIs outperform state-of-the-art methods. Further analysis demonstrates the effectiveness of our methods on the insufficient labeling problem.
Tianming Liang, Chaolei Tan, Beihao Xia, Wei-Shi Zheng 0001, Jianfang Hu
CVPR3
2024 SocialCircle: Learning the Angle-based Social Interaction Representation for Pedestrian Trajectory Prediction
abstract
Analyzing and forecasting trajectories of agents like pedestrians and cars in complex scenes has become more and more significant in many intelligent systems and ap-plications. The diversity and uncertainty in socially inter-active behaviors among a rich variety of agents make this task more challenging than other deterministic computer vision tasks. Researchers have made a lot of efforts to quan-tify the effects of these interactions on future trajectories through different mathematical models and network structures, but this problem has not been well solved. Inspired by marine animals that localize the positions of their com-panions underwater through echoes, we build a new angle-based trainable social interaction representation, named SocialCircle, for continuously reflecting the context of social interactions at different angular orientations relative to the target agent. We validate the effect of the proposed So-ciaiCircle by training it along with several newly released trajectory prediction models, and experiments show that the SocialCircle not only quantitatively improves the prediction performance, but also qualitatively helps better simulate social interactions when forecasting pedestrian trajectories in a way that is consistent with human intuitions.
Conghao Wong, Beihao Xia, Ziqian Zou, Xinge You
CVPR2
2024 Efficient Backdoor Attacks for Deep Neural Networks in Real-world Scenarios
abstract
Recent deep neural networks (DNNs) have came to rely on vast amounts of training data, providing an opportunity for malicious attackers to exploit and contaminate the data to carry out backdoor attacks. However, existing backdoor attack methods make unrealistic assumptions, assuming that all training data comes from a single source and that attackers have full access to the training data. In this paper, we introduce a more realistic attack scenario where victims collect data from multiple sources, and attackers cannot access the complete training data. We refer to this scenario as $\textbf{data-constrained backdoor attacks}$. In such cases, previous attack methods suffer from severe efficiency degradation due to the $\textbf{entanglement}$ between benign and poisoning features during the backdoor injection process. To tackle this problem, we introduce three CLIP-based technologies from two distinct streams: $\textit{Clean Feature Suppression}$ and $\textit{Poisoning Feature Augmentation}$. The results demonstrate remarkable improvements, with some settings achieving over $\textbf{100}$% improvement compared to existing attacks in data-constrained scenarios.
Ziqiang Li 0001, Heng Li 0008, Beihao Xia, Yi Wu 0018, Bin Li 0025
ICLR5
2024 Enhancing robustness of person detection: A universal defense filter against adversarial patch attacks
Zimin Mao, Shuiyan Chen, Zhuang Miao, Heng Li 0008, Beihao Xia, Junzhe Cai, Wei Yuan 0001, Xinge You
Comput. Secur.5
2024 A Proxy Attack-Free Strategy for Practically Improving the Poisoning Efficiency in Backdoor Attacks
abstract
Poisoning efficiency is crucial in poisoning-based backdoor attacks, as attackers aim to minimize the number of poisoning samples while maximizing attack efficacy. Recent studies have sought to enhance poisoning efficiency by selecting effective samples. However, these studies typically rely on a proxy backdoor injection task to identify an efficient set of poisoning samples. This proxy attack-based approach can lead to performance degradation if the proxy attack settings differ from those of the actual victims, due to the shortcut nature of backdoor learning. Furthermore, proxy attack-based methods are extremely time-consuming, as they require numerous complete backdoor injection processes for sample selection. To address these concerns, we present a Proxy attack-Free Strategy (PFS) designed to identify efficient poisoning samples based on the similarity between clean samples and their corresponding poisoning samples, as well as the diversity of the poisoning set. The proposed PFS is motivated by the observation that selecting samples with high similarity between clean and corresponding poisoning samples results in significantly higher attack success rates compared to using samples with low similarity. Additionally, we provide theoretical foundations to explain the proposed PFS. We comprehensively evaluate the proposed strategy across various datasets, triggers, poisoning rates, architectures, and training hyperparameters. Our experimental results demonstrate that PFS enhances backdoor attack efficiency while also offering a remarkable speed advantage over previous proxy attack-based selection methodologies.
Ziqiang Li 0001, Beihao Xia, Xue Rui, Wei Zhang 0251, Qinglang Guo, Zhangjie Fu 0001, Bin Li 0025
IEEE Trans. Inf. Forensics Secur.4
2024 Pedestrian Trajectory Prediction Based on Social Interactions Learning With Random Weights
abstract
Pedestrian trajectory prediction is a critical technology in the evolution of self-driving cars toward complete artificial intelligence. Over recent years, focusing on the trajectories of pedestrians to model their social interactions has surged with great interest in more accurate trajectory predictions. However, existing methods for modeling pedestrian social interactions rely on pre-defined rules, struggling to capture non-explicit social interactions. In this work, we propose a novel framework named DTGAN, which extends the application of Generative Adversarial Networks (GANs) to graph sequence data, with the primary objective of automatically capturing implicit social interactions and achieving precise predictions of pedestrian trajectory. DTGAN innovatively incorporates random weights within each graph to eliminate the need for pre-defined interaction rules. We further enhance the performance of DTGAN by exploring diverse task loss functions during adversarial training, which yields improvements of 16.7% and 39.3% on metrics ADE and FDE, respectively. The effectiveness and accuracy of our framework are verified on two public datasets. The experimental results show that our proposed DTGAN achieves superior performance and is well able to understand pedestrians' intentions.
Jiajia Xie, Sheng Zhang 0006, Beihao Xia, Zhu Xiao, Hongbo Jiang 0001, Siwang Zhou, Zheng Qin 0001, Hongyang Chen 0001
IEEE Trans. Multim.3
2023 TODE-Trans: Transparent Object Depth Estimation with Transformer
abstract
Transparent objects are widely used in industrial automation and daily life. However, robust visual recognition and perception of transparent objects have always been a major challenge. Currently, most commercial-grade depth cameras are still not good at sensing the surfaces of transparent objects due to the refraction and reflection of light. In this work, we present a transformer-based transparent object depth estimation approach from a single RGB-D input. We observe that the global characteristics of the transformer make it easier to extract contextual information to perform depth estimation of transparent areas. In addition, to better enhance the fine-grained features, a feature fusion module (FFM) is designed to assist coherent prediction. Our empirical evidence demonstrates that our model delivers significant improvements in recent popular datasets, e.g., 25% gain on RMSE and 21% gain on REL compared to previous state-of-the-art convolutional-based counterparts in ClearGrasp dataset. Extensive results show that our transformer-based model enables better aggregation of the object's RGB and inaccurate depth information to obtain a better depth representation. Our code and the pre-trained model are available at https://github.com/yuchendoudou/TODE.
Beihao Xia, Zhen Kan, Bin Li 0025
ICRA3
2023 HARP: Let Object Detector Undergo Hyperplasia to Counter Adversarial Patches
abstract
Adversarial patches can mislead object detectors to produce erroneous predictions. To defend against adversarial patches, one can take two types of protections on the model side, including modifying the detector itself (e.g., adversarial training) or attaching a new model in front of the detector. However, the former often deteriorates clean performance of detectors, and the latter may have high deployment costs caused by too many training parameters. Inspired by the phenomenon of "bone hyperplasia" in human bodies, we present a novel model-side adversarial patch defense, called HARP (Hyperplasia based Adversarial Patch defense). Just as bone hyperplasia can enhance bone strength and skeletal stability, the hyperostosia of detectors can also help to resist adversarial patches. Following this idea, HARP chooses to improve adversarial robustness by "growing" lightweight CNN modules (i.e., hyperplasia modules) on the pre-trained object detectors. We conduct extensive experiments on the PASCAL VOC and COCO datasets to compare HARP with the data-side defense JPEG and the model-side defenses adversarial training, SAC and FNC. Experimental results show that HARP provides excellent defense against adversarial patches while maintaining clean performance, outperforming the compared defense methods. Under PGD-based adaptive attacks, HARP surpasses the recently proposed defense method SAC by 12.5% in mean average precision (mAP) on PASCAL VOC, and 13.2% on COCO dataset. In addition, experiments confirm that the increase in model inference time caused by HARP is almost negligible.
Junzhe Cai, Shuiyan Chen, Heng Li 0008, Beihao Xia, Zimin Mao, Wei Yuan 0001
ACM Multimedia4
2023 MSN: Multi-Style Network for Trajectory Prediction
abstract
Trajectory prediction aims to forecast agents’ possible future locations considering their observations along with the video context. It is strongly needed by many autonomous platforms like tracking, detection, robot navigation, and self-driving cars. Whether it is agents’ internal personality factors, interactive behaviors with the neighborhood, or the influence of surroundings, they all impact agents’ future planning. However, many previous methods model and predict agents’ behaviors with the same strategy or feature distribution, making them challenging to make predictions with sufficient style differences. This paper proposes the Multi-Style Network (MSN), which utilizes style proposal and stylized prediction using two sub-networks, to provide multi-style predictions in a novel categorical way adaptively. The proposed network contains a series of style channels, and each channel is bound to a unique and specific behavior style. We use agents’ end-point plannings and their interaction context as the basis for the behavior classification, so as to adaptively learn multiple diverse behavior styles through these channels. Then, we assume that the target agents may plan their future behaviors according to each of these categorized styles, thus utilizing different style channels to make predictions with significant style differences in parallel. Experiments show that the proposed MSN outperforms current state-of-the-art methods up to 10% quantitatively on two widely used datasets, and presents better multi-style characteristics qualitatively.
Conghao Wong, Beihao Xia, Qinmu Peng, Wei Yuan 0001, Xinge You
IEEE Trans. Intell. Transp. Syst.2
2022 View Vertically: A Hierarchical Network for Trajectory Prediction via Fourier Spectrums
Conghao Wong, Beihao Xia, Ziming Hong, Qinmu Peng, Wei Yuan 0001, Qiong Cao, Xinge You
ECCV (22)2
2022 Recent Advances in Concept Drift Adaptation Methods for Deep Learning
abstract
In the ``Big Data'' age, the amount and distribution of data have increased wildly and changed over time in various time-series-based tasks, e.g weather prediction, network intrusion detection. However, deep learning models may become outdated facing variable input data distribution, which is called concept drift. To address this problem, large number of samples are usually required to update deep learning models, which is impractical in many realistic applications. This challenge drives researchers to explore the effective ways to adapt deep learning models to concept drift. In this paper, we first mathematically describe the categories of concept drift including abrupt drift, gradual drift, recurrent drift, incremental drift. We then divide existing studies into two categories (i.e., model parameter updating and model structure updating), and analyze the pros and cons of representative methods in each category. Finally, we evaluate the performance of these methods, and point out the future directions of concept drift adaptation for deep learning.
Liheng Yuan, Heng Li 0008, Beihao Xia, Cuiying Gao, Wei Yuan 0001, Xinge You
IJCAI3
2022 CSCNet: Contextual semantic consistency network for trajectory prediction in crowded spaces
Beihao Xia, Conghao Wong, Qinmu Peng, Wei Yuan 0001, Xinge You
Pattern Recognit.1
2021 FREE: Feature Refinement for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) has achieved significant progress, with many efforts dedicated to over-coming the problems of visual-semantic domain gap and seen-unseen bias. However, most existing methods directly use feature extraction models trained on ImageNet alone, ignoring the cross-dataset bias between ImageNet and GZSL benchmarks. Such a bias inevitably results in poor-quality visual features for GZSL tasks, which potentially limits the recognition performance on both seen and unseen classes. In this paper, we propose a simple yet effective GZSL method, termed feature refinement for generalized zero-shot learning (FREE), to tackle the above problem. FREE employs a feature refinement (FR) module that in-corporates semantic→visual mapping into a unified generative model to refine the visual features of seen and unseen class samples. Furthermore, we propose a self-adaptive margin center loss (SAMC-loss) that cooperates with a semantic cycle-consistency loss to guide FR to learn class- and semantically-relevant representations, and concatenate the features in FR to extract the fully refined features. Extensive experiments on five benchmark datasets demonstrate the significant performance gain of FREE over its baseline and current state-of-the-art methods. The code is available at https://github.com/shiming-chen/FREE.
Shiming Chen 0002, Beihao Xia, Qinmu Peng, Xinge You, Feng Zheng 0001, Ling Shao 0001
ICCV3
2021 CDE-GAN: Cooperative Dual Evolution-Based Generative Adversarial Network
abstract
Generative adversarial networks (GANs) have been a popular deep generative model for real-world applications. Despite many recent efforts on GANs that have been contributed, mode collapse and instability of GANs are still open problems caused by their adversarial optimization difficulties. In this article, motivated by the cooperative co-evolutionary algorithm, we propose a cooperative dual evolution-based GAN (CDE-GAN) to circumvent these drawbacks. In essence, CDE-GAN incorporates dual evolution with respect to the generator(s) and discriminators into a unified evolutionary adversarial framework to conduct effective adversarial multiobjective optimization. Thus, it exploits the complementary properties and injects dual mutation diversity into the training, to steadily diversify the estimated density in capturing multimodes and improve generative performance. Specifically, CDE-GAN decomposes the complex adversarial optimization problem into two subproblems (generation and discrimination), and each subproblem is solved with a separated subpopulation (E-GeneratorsandE-Discriminators), evolved by its own evolutionary algorithm. Additionally, we further propose aSoft Mechanismto balance the tradeoff between E-Generators and E-Discriminators to conduct steady training for CDE-GAN. Extensive experiments on one synthetic dataset and three real-world benchmark image datasets demonstrate that the proposed CDE-GAN achieves a competitive and superior performance in generating good quality and diverse samples over baselines. The code and more generated results are available at our project homepagehttps://shiming-chen.github.io/CDE-GAN-website/CDE-GAN.html.
Shiming Chen 0002, Beihao Xia, Xinge You, Qinmu Peng, Zehong Cao, Weiping Ding 0001
IEEE Trans. Evol. Comput.3