Yuecong Xu

dblp:242/7964 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0002-4292-7379ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 13 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Minute-Long Videos with Dual Parallelisms
abstract
Diffusion Transformer (DiT)-based video diffusion models generate high-quality videos at scale but incur prohibitive processing latency and memory costs for long videos. To address this, we propose a novel distributed inference strategy, termed DualParal. The core idea is that, instead of generating an entire video on a single GPU, we parallelize computation by partitioning both video frames and model layers across multiple GPUs. However, a naive parallel implementation is not feasible. Because all frames need to share the same noise level, they can't be processed independently. Instead, every step must wait for all others to finish, which cancels out the speed benefits of parallel processing. We overcome this obstacle with a block-wise denoising scheme. Namely, we segment the video into sequential blocks, each with a different noise level. As a result, we process them in a pipeline across the GPUs. Each GPU, holding a subset of the model layers, processes a specific block of frames and passes the results to the next GPU, enabling asynchronous computation and communication. To further optimize performance, we incorporate two key enhancements. Firstly, each GPU uses a feature cache technique to reduce the overhead of smooth transitions by reusing only features involved in cross-frame computation from the prior block, minimizing inter-GPU communication and redundant computation. Secondly, we employ a coordinated noise initialization strategy, ensuring globally consistent temporal dynamics by sharing initial noise patterns across GPUs. Together, these enable fast, artifact-free, and infinitely long video generation. Applied to the latest diffusion transformer video generator, our method efficiently produces 1,025-frame videos with up to 6.54x lower latency and 1.48x lower memory cost on 8xRTX 4090 GPUs.
Zeqing Wang, Xingyi Yang, Zhenxiong Tan, Yuecong Xu, Xinchao Wang
AAAI5
2025 Overlap-Aware Feature Learning for Robust Unsupervised Domain Adaptation for 3D Semantic Segmentation
abstract
3D point cloud semantic segmentation (PCSS) is a cornerstone for environmental perception in robotic systems and autonomous driving, enabling precise scene understanding through point-wise classification. While unsupervised domain adaptation (UDA) mitigates label scarcity in PCSS, existing methods critically overlook the inherent vulnerability to real-world perturbations (e.g., snow, fog, rain) and adversarial distortions. This work first identifies two intrinsic limitations that undermine current PCSS-UDA robustness: (a) unsupervised features overlap from unaligned boundaries in shared-class regions and (b) feature structure erosion caused by domain-invariant learning that suppresses target-specific patterns. To address the proposed problems, we propose a tripartite framework consisting of: 1) a robustness evaluation model quantifying resilience against adversarial attack/corruption types through robustness metrics; 2) an invertible attention alignment module (IAAM) enabling bidirectional domain mapping while preserving discriminative structure via attention-guided overlap suppression; and 3) a quality-guided contrastive memory bank that progressively refines pseudo-labels with feature quality for more discriminative representations. Extensive experiments on SynLiDAR-to-SemanticPOSS adaptation demonstrate a maximum mIoU improvement of 14.3% under adversarial attack.
Yuecong Xu, Haosheng Li, Kemi Ding
IROS2
2025 Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models
abstract
With the growing power of text-to-image diffusion models, their potential to generate harmful or biased content has become a pressing concern, motivating the development of concept erasure techniques. Existing approaches, whether relying on retraining or not, frequently compromise the generative capabilities of the target model in achieving concept erasure. Here, we introduce **Semantic Surgery**, a novel training-free framework for zero-shot concept erasure. Semantic Surgery directly operates on text embeddings *before* the diffusion process, aiming to neutralize undesired concepts at their semantic origin with dynamism to enhance both erasure completeness and the locality of generation. Specifically, Semantic Surgery dynamically estimates the presence of target concepts in an input prompt, based on which it performs a calibrated, scaled vector subtraction to neutralize their influence at the source. The overall framework consists of a Co-Occurrence Encoding module for robust multi-concept erasure and a visual feedback loop to address latent concept persistence, thereby reinforcing erasure throughout the subsequent denoising process. Our proposed Semantic Surgery requires no model retraining and adapts dynamically to the specific concepts and their intensity detected in each input prompt, ensuring precise and context-aware interventions. Extensive experiments are conducted on object, explicit content, artistic style, and multi-celebrity erasure tasks, demonstrating that our method significantly outperforms state-of-the-art approaches. That is, our proposed concept erasure framework achieves superior completeness and robustness while preserving locality and general image quality (e.g., achieving a 93.58 H-score in object erasure, reducing explicit content to just 1 instance with a 12.2 FID, and attaining an 8.09 H_a in style erasure with no MS-COCO FID/CLIP degradation). Crucially, this robustness enables our framework to function as a built-in threat detection system by monitoring concept presence scores, offering a highly effective and practical solution for safer text-to-image generation. Our code is publicly available at: https://github.com/Lexiang-Xiong/Semantic-Surgery
Lexiang Xiong, Jingwen Ye, Yuecong Xu
NeurIPS5
2024 Fully-Connected Spatial-Temporal Graph for Multivariate Time-Series Data
abstract
Multivariate Time-Series (MTS) data is crucial in various application fields. With its sequential and multi-source (multiple sensors) properties, MTS data inherently exhibits Spatial-Temporal (ST) dependencies, involving temporal correlations between timestamps and spatial correlations between sensors in each timestamp. To effectively leverage this information, Graph Neural Network-based methods (GNNs) have been widely adopted. However, existing approaches separately capture spatial dependency and temporal dependency and fail to capture the correlations between Different sEnsors at Different Timestamps (DEDT). Overlooking such correlations hinders the comprehensive modelling of ST dependencies within MTS data, thus restricting existing GNNs from learning effective representations. To address this limitation, we propose a novel method called Fully-Connected Spatial-Temporal Graph Neural Network (FC-STGNN), including two key components namely FC graph construction and FC graph convolution. For graph construction, we design a decay graph to connect sensors across all timestamps based on their temporal distances, enabling us to fully model the ST dependencies by considering the correlations between DEDT. Further, we devise FC graph convolution with a moving-pooling GNN layer to effectively capture the ST dependencies for learning effective representations. Extensive experiments show the effectiveness of FC-STGNN on multiple MTS datasets compared to SOTA methods. The code is available at https://github.com/Frank-Wang-oss/FCSTGNN.
Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen
AAAI2
2024 Graph-Aware Contrasting for Multivariate Time-Series Classification
abstract
Contrastive learning, as a self-supervised learning paradigm, becomes popular for Multivariate Time-Series (MTS) classification. It ensures the consistency across different views of unlabeled samples and then learns effective representations for these samples. Existing contrastive learning methods mainly focus on achieving temporal consistency with temporal augmentation and contrasting techniques, aiming to preserve temporal patterns against perturbations for MTS data. However, they overlook spatial consistency that requires the stability of individual sensors and their correlations. As MTS data typically originate from multiple sensors, ensuring spatial consistency becomes essential for the overall performance of contrastive learning on MTS data. Thus, we propose Graph-Aware Contrasting for spatial consistency across MTS data. Specifically, we propose graph augmentations including node and edge augmentations to preserve the stability of sensors and their correlations, followed by graph contrasting with both node- and graph-level contrasting to extract robust sensor- and global-level features. We further introduce multi-window temporal contrasting to ensure temporal consistency in the data for each sensor. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on various MTS classification tasks. The code is available at https://github.com/Frank-Wang-oss/TS-GAC.
Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen
AAAI2
2024 Reliable Spatial-Temporal Voxels For Multi-modal Test-Time Adaptation
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Xingyu Ji, Shenghai Yuan 0001, Lihua Xie 0001
ECCV (28)2
2024 Diffusion Model Is a Good Pose Estimator from 3D RF-Vision
Junqiao Fan, Jianfei Yang 0001, Yuecong Xu, Lihua Xie 0001
ECCV (16)3
2024 Can We Evaluate Domain Adaptation Models Without Target-Domain Labels?
abstract
Unsupervised domain adaptation (UDA) involves adapting a model trained on a label-rich source domain to an unlabeled target domain. However, in real-world scenarios, the absence of target-domain labels makes it challenging to evaluate the performance of UDA models. Furthermore, prevailing UDA methods relying on adversarial training and self-training could lead to model degeneration and negative transfer, further exacerbating the evaluation problem. In this paper, we propose a novel metric called the Transfer Score to address these issues. The proposed metric enables the unsupervised evaluation of UDA models by assessing the spatial uniformity of the classifier via model parameters, as well as the transferability and discriminability of deep representations. Based on the metric, we achieve three novel objectives without target-domain labels: (1) selecting the best UDA method from a range of available options, (2) optimizing hyperparameters of UDA models to prevent model degeneration, and (3) identifying which checkpoint of UDA model performs optimally. Our work bridges the gap between data-level UDA research and practical UDA scenarios, enabling a realistic assessment of UDA model performance. We validate the effectiveness of our metric through extensive empirical studies on UDA datasets of different scales and imbalanced distributions. The results demonstrate that our metric robustly achieves the aforementioned goals.
Jianfei Yang 0001, Hanjie Qian, Yuecong Xu, Kai Wang 0036, Lihua Xie 0001
ICLR3
2024 MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation
abstract
Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-imbalanced performance, restricting their adoption in real applications. This imbalanced performance is mainly caused by: 1) self-training with imbalanced data and 2) the lack of pixel-wise 2D supervision signals. In this work, we propose Multi-modal Prior Aided (MoPA) domain adaptation to improve the performance of rare objects. Specifically, we develop Valid Ground-based Insertion (VGI) to rectify the imbalance supervision signals by inserting prior rare objects collected from the wild while avoiding introducing artificial artifacts that lead to trivial solutions. Meanwhile, our SAM consistency loss leverages the 2D prior semantic masks from SAM as pixel-wise supervision signals to encourage consistent predictions for each object in the semantic mask. The knowledge learned from modal-specific prior is then shared across modalities to achieve better rare object segmentation. Extensive experiments show that our method achieves state-of-the-art performance on the challenging MM-UDA benchmark. Code will be available at https://github.com/AronCao49/MoPA.
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001
ICRA2
2024 Going Deeper into Recognizing Actions in Dark Environments: A Comprehensive Benchmark Study
Yuecong Xu, Haozhi Cao, Jianxiong Yin, Zhenghua Chen, Xiaoli Li 0001, Zhengguo Li, Qianwen Xu 0001, Jianfei Yang 0001
Int. J. Comput. Vis.1
2024 SEA++: Multi-Graph-Based Higher-Order Sensor Alignment for Multivariate Time-Series Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) methods have been successful in reducing label dependency by minimizing the domain discrepancy between labeled source domains and unlabeled target domains. However, these methods face challenges when dealing with Multivariate Time-Series (MTS) data. MTS data typically originates from multiple sensors, each with its unique distribution. This property poses difficulties in adapting existing UDA techniques, which mainly focus on aligning global features while overlooking the distribution discrepancies at the sensor level, thus limiting their effectiveness for MTS data. To address this issue, a practical domain adaptation scenario is formulated as Multivariate Time-Series Unsupervised Domain Adaptation (MTS-UDA). In this paper, we propose SEnsor Alignment (SEA) for MTS-UDA, aiming to address domain discrepancy at both local and global sensor levels. At the local sensor level, we design endo-feature alignment, which aligns sensor features and their correlations across domains. To reduce domain discrepancy at the global sensor level, we design exo-feature alignment that enforces restrictions on global sensor features. We further extend SEA to SEA++ by enhancing the endo-feature alignment. Particularly, we incorporate multi-graph-based higher-order alignment for both sensor features and their correlations. Extensive empirical results have demonstrated the state-of-the-art performance of our SEA and SEA++ on six public MTS datasets for MTS-UDA.
Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Self-Supervised Video Representation Learning by Video Incoherence Detection
abstract
This article introduces a novel self-supervised method that leverages incoherence detection for video representation learning. It stems from the observation that the visual system of human beings can easily identify video incoherence based on their comprehensive understanding of videos. Specifically, we construct the incoherent clip by multiple subclips hierarchically sampled from the same raw video with various lengths of incoherence. The network is trained to learn the high-level representation by predicting the location and length of incoherence given the incoherent clip as input. Additionally, we introduce intravideo contrastive learning to maximize the mutual information between incoherent clips from the same raw video. We evaluate our proposed method through extensive experiments on action recognition and video retrieval using various backbone networks. Experiments show that our proposed method achieves remarkable performance across different backbone networks and different datasets compared to previous coherence-based methods.
Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie 0001, Jianxiong Yin, Simon See, Qianwen Xu 0001, Jianfei Yang 0001
IEEE Trans. Cybern.2
2024 Aligning Correlation Information for Domain Adaptation in Action Recognition
abstract
Domain adaptation (DA) approaches address domain shift and enable networks to be applied to different scenarios. Although various image DA approaches have been proposed in recent years, there is limited research toward video DA. This is partly due to the complexity in adapting the different modalities of features in videos, which includes the correlation features extracted as long-range dependencies of pixels across spatiotemporal dimensions. The correlation features are highly associated with action classes and proven their effectiveness in accurate video feature extraction through the supervised action recognition task. Yet correlation features of the same action would differ across domains due to domain shift. Therefore, we propose a novel adversarial correlation adaptation network (ACAN) to align action videos by aligning pixel correlations. ACAN aims to minimize the distribution of correlation information, termed as pixel correlation discrepancy (PCD). Additionally, video DA research is also limited by the lack of cross-domain video datasets with larger domain shifts. We, therefore, introduce a novel HMDB-ARID dataset with a larger domain shift caused by a larger statistical difference between domains. This dataset is built in an effort to leverage current datasets for dark video classification. Empirical results demonstrate the state-of-the-art performance of our proposed ACAN for both existing and the new video DA datasets.
Yuecong Xu, Haozhi Cao, Kezhi Mao, Zhenghua Chen, Lihua Xie 0001, Jianfei Yang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 SEnsor Alignment for Multivariate Time-Series Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) methods can reduce label dependency by mitigating the feature discrepancy between labeled samples in a source domain and unlabeled samples in a similar yet shifted target domain. Though achieving good performance, these methods are inapplicable for Multivariate Time-Series (MTS) data. MTS data are collected from multiple sensors, each of which follows various distributions. However, most UDA methods solely focus on aligning global features but cannot consider the distinct distributions of each sensor. To cope with such concerns, a practical domain adaptation scenario is formulated as Multivariate Time-Series Unsupervised Domain Adaptation (MTS-UDA). In this paper, we propose SEnsor Alignment (SEA) for MTS-UDA to reduce the domain discrepancy at both the local and global sensor levels. At the local sensor level, we design the endo-feature alignment to align sensor features and their correlations across domains, whose information represents the features of each sensor and the interactions between sensors. Further, to reduce domain discrepancy at the global sensor level, we design the exo-feature alignment to enforce restrictions on the global sensor features. Meanwhile, MTS also incorporates the essential spatial-temporal dependencies information between sensors, which cannot be transferred by existing UDA methods. Therefore, we model the spatial-temporal information of MTS with a multi-branch self-attention mechanism for simple and effective transfer across domains. Empirical results demonstrate the state-of-the-art performance of our proposed SEA on two public MTS datasets for MTS-UDA. The code is available at https://github.com/Frank-Wang-oss/SEA
Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001
AAAI2
2023 Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation
abstract
Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmentation. The key to MMCTTA is to adaptively attend to the reliable modality while avoiding catastrophic forgetting during continual domain shifts, which is out of the capability of previous TTA or CTTA methods. To fulfill this gap, we propose an MM-CTTA method called Continual Cross-Modal Adaptive Clustering (CoMAC) that addresses this task from two perspectives. On one hand, we propose an adaptive dual-stage mechanism to generate reliable cross-modal predictions by attending to the reliable modality based on the class-wise feature-centroid distance in the latent space. On the other hand, to perform test-time adaptation without catastrophic forgetting, we design class-wise momentum queues that capture confident target features for adaptation while stochastically restoring pseudo-source features to revisit source knowledge. We further introduce two new benchmarks to facilitate the exploration of MM-CTTA in the future. Our experimental results show that our method achieves state-of-the-art performance on both benchmarks. Visit our project website at https://sites.google.com/view/mmcotta.
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001
ICCV2
2023 Augmenting and Aligning Snippets for Few-Shot Video Domain Adaptation
abstract
For video models to be transferred and applied seamlessly across video tasks in varied environments, Video Unsupervised Domain Adaptation (VUDA) has been introduced to improve the robustness and transferability of video models. However, current VUDA methods rely on a vast amount of high-quality unlabeled target data, which may not be available in real-world cases. We thus consider a more realistic Few-Shot Video-based Domain Adaptation (FSVDA) scenario where we adapt video models with only a few target video samples. While a few methods have touched upon Few-Shot Domain Adaptation (FSDA) in images and in FSVDA, they rely primarily on spatial augmentation for target domain expansion with alignment performed statistically at the instance level. However, videos contain more knowledge in terms of rich temporal and semantic information, which should be fully considered while augmenting target domains and performing alignment in FSVDA. We propose a novel SSA2lign to address FSVDA at the snippet level, where the target domain is expanded through a simple snippet-level augmentation followed by the attentive alignment of snippets both semantically and statistically, where semantic alignment of snippets is conducted through multiple perspectives. Empirical results demonstrate state-of-the-art performance of SSA2lign across multiple cross-domain action recognition benchmarks. Code will be provided at: https://github.com/xuyu0010/SSA2lign.
Yuecong Xu, Jianfei Yang 0001, Yunjiao Zhou, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001
ICCV1
2023 MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensing
abstract
4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensing has emerged as a promising alternative, leveraging LiDAR, mmWave radar, and WiFi signals for device-free human sensing. In this paper, we propose MM-Fi, the first multi-modal non-intrusive 4D human dataset with 27 daily or rehabilitation action categories, to bridge the gap between wireless sensing and high-level human perception tasks. MM-Fi consists of over 320k synchronized frames of five modalities from 40 human subjects. Various annotations are provided to support potential sensing tasks, e.g., human pose estimation and action recognition. Extensive experiments have been conducted to compare the sensing capacity of each or several modalities in terms of multiple tasks. We envision that MM-Fi can contribute to wireless sensing research with respect to action recognition, human pose estimation, multi-modal learning, cross-modal supervision, and interdisciplinary healthcare research.
Jianfei Yang 0001, Yunjiao Zhou, Xinyan Chen 0002, Yuecong Xu, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001
NeurIPS5
2023 Multi-Source Video Domain Adaptation With Temporal Attentive Moment Alignment Network
abstract
Multi-Source Domain Adaptation (MSDA) is a more practical domain adaptation scenario in real-world scenarios, which relaxes the assumption in conventional Unsupervised Domain Adaptation (UDA) that source data are sampled from a single domain and match a uniform data distribution. The MSDA is more challenging due to the existence of different domain shifts between distinct domain pairs. When considering videos, the negative transfer would be provoked by spatial-temporal features and can be formulated into a more challenging Multi-Source Video Domain Adaptation (MSVDA) problem. In this paper, we address the MSVDA problem by proposing a novel Temporal Attentive Moment Alignment Network (TAMAN) which aims for effective feature transfer by dynamically aligning both spatial and temporal feature moments. The TAMAN further constructs robust global temporal features by attending to dominant domain-invariant local temporal features with high local classification confidence and low disparity between global and local feature discrepancies. To facilitate future research on the MSVDA problem, we introduce comprehensive benchmarks, covering extensive MSVDA scenarios. Empirical results demonstrate a superior performance of the proposed TAMAN across multiple MSVDA benchmarks.
Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Keyu Wu 0002, Min Wu 0008, Zhengguo Li, Zhenghua Chen
IEEE Trans. Circuits Syst. Video Technol.1
2022 Generalizing Reinforcement Learning through Fusing Self-Supervised Learning into Intrinsic Motivation
abstract
Despite the great potential of reinforcement learning (RL) in solving complex decision-making problems, generalization remains one of its key challenges, leading to difficulty in deploying learned RL policies to new environments. In this paper, we propose to improve the generalization of RL algorithms through fusing Self-supervised learning into Intrinsic Motivation (SIM). Specifically, SIM boosts representation learning through driving the cross-correlation matrix between the embeddings of augmented and non-augmented samples close to the identity matrix. This aims to increase the similarity between the embedding vectors of a sample and its augmented version while minimizing the redundancy between the components of these vectors. Meanwhile, the redundancy reduction based self-supervised loss is converted to an intrinsic reward to further improve generalization in RL via an auxiliary objective. As a general paradigm, SIM can be implemented on top of any RL algorithm. Extensive evaluations have been performed on a diversity of tasks. Experimental results demonstrate that SIM consistently outperforms the state-of-the-art methods and exhibits superior generalization capability and sample efficiency.
Keyu Wu 0002, Min Wu 0008, Zhenghua Chen, Yuecong Xu, Xiaoli Li 0001
AAAI4
2022 Source-Free Video Domain Adaptation by Learning Temporal Consistency for Action Recognition
Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Keyu Wu 0002, Min Wu 0008, Zhenghua Chen
ECCV (34)1
2022 Calibrating Class Weights with Multi-Modal Information for Partial Video Domain Adaptation
abstract
Assuming the source label space subsumes the target one, Partial Video Domain Adaptation (PVDA) is a more general and practical scenario for cross-domain video classification problems. The key challenge of PVDA is to mitigate the negative transfer caused by the source-only outlier classes. To tackle this challenge, a crucial step is to aggregate target predictions to assign class weights by up-weighing target classes and down-weighing outlier classes. However, the incorrect predictions of class weights can mislead the network and lead to negative transfer. Previous works improve the class weight accuracy by utilizing temporal features and attention mechanisms, but these methods may fall short when trying to generate accurate class weight when domain shifts are significant, as in most real-world scenarios. To deal with these challenges, we first propose the Multi-modality partial Adversarial Network (MAN), which utilizes multi-scale and multi-modal information to enhance PVDA performance. Based on MAN, we then propose Multi-modality Cluster-calibrated partial Adversarial Network (MCAN). It utilizes a novel class weight calibration method to alleviate the negative transfer caused by incorrect class weights. Specifically, the calibration method tries to identify and weigh correct and incorrect predictions using distributional information implied by unsupervised clustering. Extensive experiments are conducted on prevailing PVDA benchmarks, and the proposed MCAN achieves significant improvements when compared to state-of-the-art PVDA methods.
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao
ACM Multimedia2
2022 A novel end-to-end neural network for simultaneous filtering of task-unrelated named entities and fine-grained typing of task-related named entities
Kezhi Mao, Yuecong Xu, Edmond Yat-Man Lo
Expert Syst. Appl.4
2021 Partial Video Domain Adaptation with Partial Adversarial Temporal Attentive Network
abstract
Partial Domain Adaptation (PDA) is a practical and general domain adaptation scenario, which relaxes the fully shared label space assumption such that the source label space subsumes the target one. The key challenge of PDA is the issue of negative transfer caused by source-only classes. For videos, such negative transfer could be triggered by both spatial and temporal features, which leads to a more challenging Partial Video Domain Adaptation (PVDA) problem. In this paper, we propose a novel Partial Adversarial Temporal Attentive Network (PATAN) to address the PVDA problem by utilizing both spatial and temporal features for filtering source-only classes. Besides, PATAN constructs effective overall temporal features by attending to local temporal features that contribute more toward the class filtration process. We further introduce new benchmarks to facilitate research on PVDA problems, covering a wide range of PVDA scenarios. Empirical results demonstrate the state-of-the-art performance of our proposed PATAN across the multiple PVDA benchmarks. Code will be provided at: https://github.com/xuyu0010/PATAN.
Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Zhenghua Chen, Kezhi Mao
ICCV1
2021 Exploiting inter-frame regional correlation for efficient action recognition
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Expert Syst. Appl.1
2021 PNL: Efficient long-range dependencies extraction with pyramid non-local module for action recognition
Yuecong Xu, Haozhi Cao, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Neurocomputing1
2021 Effective action recognition with embedded key point shifts
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Pattern Recognit.2
2020 Extracting Temporal Features by Key Points Transfer for Effective Action Recognition
abstract
The extraction of temporal features in video is an essential task for effective action recognition. Previous networks utilizes optical flow as effective temporal features, which utilizes positional relationship between pixels to extract temporal features. However, not all the pixels are meaningful in a video frame, with most pixels related to background information rather than the action itself. In this paper, we propose a novel temporal feature extraction model, Key points inter-Frame Transfer Module (KFTM), which extracts the temporal feature by extracting the transfer of multiple key points along the temporal axis. Such information can be equivalent to the temporal feature of the video since it also represents the positional relationship between pixels. Yet such method is more efficient due to the use of key points. Meanwhile, to extract the temporal feature more effectively, we add the attention mechanism which pays more attention to the transfer of key points most relevant to the action. Our proposed module obtains competitive performance on both UCF101 and HMDB51 datasets with 96.49% accuracy on UCF101 and 77.48% accuracy on HMDB51 datasets.
Yuecong Xu
ICARCV2
2020 Domain Adaptation for Degraded Remote Scene Classification
abstract
Remote scene classification serves a vital role in many applications. However, satellite images are often blurred and degraded due to aerosol scattering under fog, haze, and other weather conditions, reducing the image contrast and color fidelity. State-of-the-art remote sensing classification models building upon convolutional neural networks (CNNs) are mostly trained on annotated datasets of clear satellite images. When applied to blurred images, they will suffer a great degradation in performance. To address this problem, we adopt the domain adaptation algorithm TADA and propose Transferable Attention enhanced Adversarial Adaptation Network (TA3N), which utilizes annotated data in clear images by applying knowledge transferring from clear image domain to blurred image domain. Our TA3N first integrates spatial attention to focus on salient areas which are discriminative and transferable. In addition, domain discriminator and adversarial training via gradient reversal layer are used to minimize the discrepancies in extracted features from clear and degraded domains. We synthesize degraded remote scene classification dataset SSI based on FoHIS model. Experiments on degraded SSI showed that TA3N significantly outperforms baseline and other state-of-the-art domain adaptation methods.
Jianfei Yang 0001, Hailin Chen, Yuecong Xu, Ziji Shi, Ruikang Luo, Lihua Xie 0001, Rong Su 0001
ICARCV3
2020 Bag-of-Concepts representation for document classification based on automatic knowledge acquisition from probabilistic knowledge base
Kezhi Mao, Yuecong Xu, Jiaheng Zhang
Knowl. Based Syst.3
2019 Semantic-filtered Soft-Split-Aware video captioning with audio-augmented feature
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao
Neurocomputing1