Xue Lin 0003

dblp:94/7236-3 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-3088-2767ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Collaborative Model and Data Adaptation at Test Time
Chunyun Zhang, Fujun Yang, Chaoran Cui, Shuai Gong, Wenna Wang, Xue Lin 0003, Yonggang Qi, Lei Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.6
2025 Human-object interaction detection via recycling of ground-truth annotations
Xue Lin 0003, Qi Zou 0001, Xixia Xu
Pattern Recognit.1
2024 Adversarial Source Generation for Source-Free Domain Adaptation
abstract
Unsupervised domain adaptation aims to transfer the knowledge learned from a labeled source domain to an unlabeled target domain with different data distributions. However, in practice, source samples are not always available due to privacy protection and storage resource limitations. To address this concern, Source-Free Domain Adaptation (SFDA) has recently attracted growing research attention, as it only needs a pre-trained source model without direct access to source data. In this paper, we propose a novel Adversarial SOurce GEneration (ASOGE) method for SFDA, which introduces an additional generative module to produce synthetic labeled source samples and uses them to facilitate cross-domain adaptation. Unlike early studies that train the generator independently and perform the adaptation only after the generator is finished, ASOGE integrates the generation and adaptation stages within a collaborative framework by making them play an adversarial game. In the generation stage, the labeled source samples are not produced blindly; instead, they are hard-to-align samples that provide knowledge more worth learning for the adaptation stage. To achieve a fine-grained domain alignment, a class-aware discrepancy between source and target domains is measured via contrastive learning. Extensive experiments on benchmark datasets demonstrate the effectiveness of ASOGE compared to the state-of-the-art methods.
Chaoran Cui, Fan'an Meng, Chunyun Zhang, Lei Zhu 0002, Shuai Gong, Xue Lin 0003
IEEE Trans. Circuits Syst. Video Technol.7
2023 Structure-Enriched Topology Learning For Cross-Domain Multi-Person Pose Estimation
abstract
Human pose estimation has been widely studied with much focus on supervised learning. However, in real applications, a pretrained pose estimation model usually needs be adapted to a novel domain without labels or with sparse labels. Existing domain adaptation methods cannot well deal with it since poses have flexible topological structures and need fine-grained local features. Aiming at the characteristics of human pose, we propose a novel domain adaptation method for multi-person pose estimation (MPPE) to alleviate the human-level shift. Firstly, the training samples of human poses are clustered into groups according to the posture similarity. Within the clustered space, we conduct three adaptation modules: Cross-Attentive Feature Alignment (CAFA), Intra-domain Structure Adaptation (ISA) and Adaptive Human-Topology Adaptation (AHTA). The CAFA adopts a bidirectional spatial attention mechanism to explore fine-grained local feature correlation between two humans, and thus to adaptively aggregate consistent features for adaptation. ISA only works in semi-supervised domain adaptation (SSDA) to exploit semantic relationship of corresponding keypoints for reducing the intra-domain bias. Importantly, we creatively propose an AHTA to enrich human topological knowledge for reducing the inter-domain discrepancy. Specifically, the pose structure and the cross-instance topological relations are modeled via graph networks. This flexible topology learning benefits the occluded or extreme pose inference. Extensive experiments are conducted on two popular benchmarks and additional two challenging datasets. Results demonstrate the competency of our method, which works in unsupervised or semi-supervised modes, compared with the existing supervised approaches.
Xixia Xu, Qi Zou 0001, Xue Lin 0003
IEEE Trans. Multim.3
2023 Effects of Motion-Relevant Knowledge From Unlabeled Video to Human-Object Interaction Detection
abstract
The existing works on human-object interaction (HOI) detection usually rely on expensive large-scale labeled image datasets. However, in real scenes, labeled data may be insufficient, and some rare HOI categories have few samples. This poses great challenges for deep-learning-based HOI detection models. Existing works tackle it by introducing compositional learning or word embedding but still need large-scale labeled data or extremely rely on the well-learned knowledge. In contrast, the freely available unlabeled videos contain rich motion-relevant information that can help infer rare HOIs. In this article, we creatively propose a multitask learning (MTL) perspective to assist in HOI detection with the aid of motion-relevant knowledge learning on unlabeled videos. Specifically, we design the appearance reconstruction loss (ARL) and sequential motion mining module in a self-supervised manner to learn more generalizable motion representations for promoting the detection of rare HOIs. Moreover, to better transfer motion-related knowledge from unlabeled videos to HOI images, a domain discriminator is introduced to decrease the domain gap between two domains. Extensive experiments on the HICO-DET dataset with rare categories and the V-COCO dataset with minimum supervision demonstrate the effectiveness of motion-aware knowledge implied in unlabeled videos for HOI detection.
Xue Lin 0003, Qi Zou 0001, Xixia Xu, Ding Ding 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Adaptive Hypergraph Neural Network for Multi-Person Pose Estimation
abstract
This paper proposes a novel two-stage hypergraph-based framework, dubbed ADaptive Hypergraph Neural Network (AD-HNN) to estimate multiple human poses from a single image, with a keypoint localization network and an Adaptive-Pose Hypergraph Neural Network (AP-HNN) added onto the former network. For providing better guided representations of AP-HNN, we employ a Semantic Interaction Convolution (SIC) module within the initial localization network to acquire more explicit predictions. Build upon this, we design a novel adaptive hypergraph to represent a human body for capturing high-order semantic relations among different joints. Notably, it can adaptively adjust the relations between joints and seek the most reasonable structure for the variable poses to benefit the keypoint localization. These two stages are combined to be trained in an end-to-end fashion. Unlike traditional Graph Convolutional Networks (GCNs) that are based on a fixed tree structure, AP-HNN can deal with ambiguity in human pose estimation. Experimental results demonstrate that the AD-HNN achieves state-of-the-art performance both on the MS-COCO, MPII and CrowdPose datasets.
Xixia Xu, Qi Zou 0001, Xue Lin 0003
AAAI3
2022 Location-Free Human Pose Estimation
abstract
Human pose estimation (HPE) usually requires large-scale training data to reach high performance. However, it is rather time-consuming to collect high-quality and fine-grained annotations for human body. To alleviate this issue, we revisit HPE and propose a location-free framework without supervision of keypoint locations. We reformulate the regression-based HPE from the perspective of classification. Inspired by the CAM-based weakly-supervised object localization, we observe that the coarse keypoint locations can be acquired through the part-aware CAMs but unsatisfactory due to the gap between the fine-grained HPE and the object-level localization. To this end, we propose a customized transformer framework to mine the fine-grained representation of human context, equipped with the structural relation to capture subtle differences among keypoints. Concretely, we design a Multi-scale Spatial-guided Context Encoder to fully capture the global human context while focusing on the part-aware regions and a Relation-encoded Pose Prototype Generation module to encode the structural relations. All these works together for strengthening the weak supervision from image-level category labels on locations. Our model achieves competitive performance on three datasets when only supervised at a category-level and importantly, it can achieve comparable results with fully-supervised methods with only 25% location labels on MS-COCO and MPII.
Xixia Xu, Yingguo Gao, Xue Lin 0003, Qi Zou 0001
CVPR4
2022 CFENet: Content-aware feature enhancement network for multi-person pose estimation
Xixia Xu, Qi Zou 0001, Xue Lin 0003
Appl. Intell.3
2021 Motion-Aware Feature Enhancement Network for Video Prediction
abstract
Video prediction is challenging, due to the pixel-level precision requirement and the difficulty in capturing scene dynamics. Most approaches tackle the problems by pixel-level reconstruction objectives and two decomposed branches, which still suffer from blurry generations or dramatic degradations in long-term prediction. In this paper, we propose a Motion-Aware Feature Enhancement (MAFE) network for video prediction to produce realistic future frames and achieve relatively long-term predictions. First, a Channel-wise and Spatial Attention (CSA) module is designed to extract motion-aware features, which enhances the contribution of important motion details during encoding, and subsequently improves the discriminability of attention map for the frame refinement. Second, a Motion Perceptual Loss (MPL) is proposed to guide the learning of temporal cues, which benefits to robust long-term video prediction. Extensive experiments on three human activity video datasets: KTH, Human3.6M, and PennAction demonstrate the effectiveness of the proposed video prediction model compared with the state-of-the-art approaches.
Xue Lin 0003, Qi Zou 0001, Xixia Xu
IEEE Trans. Circuits Syst. Video Technol.1
2020 Action-Guided Attention Mining and Relation Reasoning Network for Human-Object Interaction Detection
abstract
Human-object interaction (HOI) detection is important to understand human-centric scenes and is challenging due to subtle difference between fine-grained actions, and multiple co-occurring interactions. Most approaches tackle the problems by considering the multi-stream information and even introducing extra knowledge, which suffer from a huge combination space and the non-interactive pair domination problem. In this paper, we propose an Action-Guided attention mining and Relation Reasoning (AGRR) network to solve the problems. Relation reasoning on human-object pairs is performed by exploiting contextual compatibility consistency among pairs to filter out the non-interactive combinations. To better discriminate the subtle difference between fine-grained actions, an action-aware attention based on class activation map is proposed to mine the most relevant features for recognizing HOIs. Extensive experiments on V-COCO and HICO-DET datasets demonstrate the effectiveness of the proposed model compared with the state-of-the-art approaches.
Xue Lin 0003, Qi Zou 0001, Xixia Xu
IJCAI1
2020 Alleviating Human-level Shift: A Robust Domain Adaptation Method for Multi-person Pose Estimation
abstract
Human pose estimation has been widely studied with much focus on supervised learning requiring sufficient annotations. However, in real applications, a pretrained pose estimation model usually need be adapted to a novel domain with no labels or sparse labels. Such domain adaptation for 2D pose estimation hasn't been explored. The main reason is that a pose, by nature, has typical topological structure and needs fine-grained features in local keypoints. While existing adaptation methods do not consider topological structure of object-of-interest and they align the whole images coarsely. Therefore, we propose a novel domain adaptation method for multi-person pose estimation to conduct the human-level topological structure alignment and fine-grained feature alignment. Our method consists of three modules: Cross-Attentive Feature Alignment (CAFA), Intra-domain Structure Adaptation (ISA) and Inter-domain Human-Topology Alignment (IHTA) module. The CAFA adopts a bidirectional spatial attention module (BSAM) that focuses on fine-grained local feature correlation between two humans to adaptively aggregate consistent features for adaptation. We adopt ISA only in semi-supervised domain adaptation (SSDA) to exploit the corresponding keypoint semantic relationship for reducing the intra-domain bias. Most importantly, we propose an IHTA to learn more domain-invariant human topological representation for reducing the inter-domain discrepancy. We model the human topological structure via the graph convolution network (GCN), by passing messages on which, high-order relations can be considered. This structure preserving alignment based on GCN is beneficial to the occluded or extreme pose inference. Extensive experiments are conducted on two popular benchmarks and results demonstrate the competency of our method compared with existing supervised approaches.
Xixia Xu, Qi Zou 0001, Xue Lin 0003
ACM Multimedia3
2020 Integral Knowledge Distillation for Multi-Person Pose Estimation
abstract
Both accuracy and efficiency are of equal importance to the human pose estimation. Most of the existing methods simply pursue excellent performance, sacrificing massive computing resources and memory. Out of this consideration, we present a novel compact and lightweight framework to train more efficient estimators using knowledge distillation. Three distillation mechanisms are proposed in our method from different perspectives, including logit distillation, feature distillation and structure distillation. Concretely, the logit distillation regards the output of teacher model as soft target to stimulate the student model. The feature distillation distills the high-level features of the teacher model to assist the student. Unlike the above strategies, the structure distillation considers the problem in a global view, aiming at ensuring the student prediction contains quite abundant structure knowledge like the teacher. We empirically demonstrate the effectiveness and efficiency of our methods on two multi-person pose estimation datasets (COCO and MPII). Specifically, our model can achieve competitive performance with the most state-of-the-art methods and consume only 35% model parameters and GFLOPs of our baseline (SimpleBaseline-ResNet-50) on the COCO dataset.
Xixia Xu, Qi Zou 0001, Xue Lin 0003
IEEE Signal Process. Lett.3