EDBT 2026 Demo / reviewers in the wild / expert
Xixia Xu
dblp:261/2906
· DBLP profile ↗
18ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0001-6305-475XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Contrastive Consistency for Human Pose Estimation in Images and VideosabstractHuman pose estimation (HPE) is an invaluable task in computer vision with various practical applications. This paper proposes a novel Hierarchical Contrastive Consistensy constraint (HICCON) to improve the HPE in both images and videos, which describes the input into multi-granular representations at spatial and temporal domain and performs multi-level feature consistency by exploring the characteristic of human structure and time sequence. The hierarchical contrast is conducted at four levels: keypoint-level, part-level, instance-level and clip-level. In spatial, we consider keypoint-level and part-level consistency across instances within frame to enhance the fine-grained keypoint robustness. The former conducts the single keypoint feature contrast across instances to improve the category-specific keypoint features. The latter explores the specific pair-wise features for preserving the instructive relation. In temporal, we develop the instance-level and clip-level feature consistency across frames to capture more discriminative temporal representations. The former discriminates instance features across frames within the same video, whereas the clip-level constraint aims to discriminate consistent features from different videos in order to capture more distinctive temporal features. Extensive experiments on kinds of architectures across datasets i.e, PoseTrack2017, PoseTrack2018 and PoseTrack2021 show the HICCON achieves about 1.5% improvement than baseline. Besides, the proposed method unleashes the potential of the contrastive learning in HPE field. Xixia Xu, Qi Zou 0001, Jiamao Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | MMCPose: Multimodal Condition-Driven 3D Human Pose Estimation Via Diffusion ModelsabstractNowadays, diffusion-based methods for monocular 3D human pose estimation (3D HPE) have achieved state-of-the-art performance by directly regressing the 3D joint coordinates from the 2D observations. Although some methods incorporated the human body prior to improve the denoising quality, the absense of the structural relation and pose-aware guidance make these models prone to generating unreasonable poses. The challenge is noticeable in complex conditions such as occlusions and crowded scenarios. To alleviate this, we present MMCPose, a novel Multi-modal Condition-driven 3D HPE framework via diffusion models that capitalizes on the benefits of the multi-modal conditioning input. Specifically, we propose Multi-modal Condition Learning (MCL) strategy to incorporate multi-modal conditions such as joint- wise relation, part-aware prompt and pose-aware mask to improve the generation quality. The MCL block consists of (i) Joint- wise Relation Condition Learning (JRCL) models the flexible joint- wise relation via GCN to mitigate disturbances arising from confused joints. (ii) Part-aware Prompt Condition Learning (PPCL) constructs multi-granular prompts via accessible texts and feasible knowledge of body parts with learnable prompts to model implicit textual guidance. (iii) Pose-aware Mask Condition Learning (PMCL) designs a pose-specific mask to increase the model's emphasis to the pose region, augmenting the precision in capturing intricate pose details. Furthermore, we explore a multi-modal condition-pose interaction learning (MCPI) mechanism to establish interaction between the learned multi-modal conditions and poses to maximize the power of condition effect. This method fully unleashes the perceptual capability of the multi-modal conditions in diffusion-based 3D HPE. Extensive evaluations conducted on two popular benchmarks (e.g., Human3.6 M, MPI-INF-3DHP) and achieve new state-of-the-art performance. Xixia Xu, Jiamao Li |
IEEE Trans. Multim. | 1 |
| 2025 | The Motion in the Details: Adapting CLIP for Action Recognition via Dual-prompt GuidanceabstractRecently, large-scale pre-trained language-image models like CLIP have shown extraordinary capabilities for understanding image-level objects, but naively transferring such models to video recognition is still unsatisfactory. Existing methods design plugged temporal modules into the pre-trained model or explore the vision-text relation to improve the performance, which either demonstrated insufficient attention to the frame-wise action subject or is limited by the unreliable prompt guidance, struggling to achieve better adaptation. In this work, we present DP-CLIP, a novel dual-prompt guidance mechanism to disentangle the adaptation in textual and temporal aspects. Specifically, DP-CLIP consists of an explicit instruction-filtered caption prompt guidance (ECPG) module and an implicit action subject prompt mining (ISPM) module to maximize the textual-visual alignment and enhance the temporal reasoning ability. The former designs an instruction-filtered strategy to generate reliable and reasonable semantic captions matching with the video details, narrowing the gap between videos and labels. Further, the ISPM emphasizes action-related discriminative clues in temporal via highlighting the action subject and refining the motion cues across frames, immune to irrelevant interference. Extensive experiments on Kinetics-400, HMDB-51 and UCF-101 demonstrate that our method achieves state-of-the-art performance across fully-supervised and zero-shot settings. Longjuan Sun, Xixia Xu, Dongchen Zhu, Jiamao Li |
ICME | 2 |
| 2025 | PCGE: Boosting 3D Visual Grounding via Progressive Comprehension and Geometric-topology Perception EnhancementabstractThe 3D visual grounding task aims to establish correspondences between the 3D physical world and textual descriptions. Despite significant progress having been made, it still suffers from some challenges that need to be solved. a) Scene-agnostic text reasoning causes misaligned target region concentration. b) The regional pseudo-center interferences result in an inaccurate geometric center. c) Multi-modal features overemphasize semantics, leading to degradation in geometric topological perception for size regression. To address these issues, we creatively propose a Progressive Comprehension and Geometric-topology Perception Enhancement (PCGE) one-stage framework, which decouples the task into keypoint estimation and size regression under textual constraints. Specifically, to enable coarse-to-fine keypoint estimation, we propose the STAR module to focus the target region approximately with a scene-specific reasoning mechanism, while the K2C module performs geometric calibration to alleviate pseudo-center bias. For size regression, we propose GTE to enhance the geometric boundary perception during the decoding process, improving size regression via establishing topological matrices. Compared with previous methods, our approach achieves state-of-the-art performance on ScanRefer and Sr3D, with 3.94% leads of [email protected] on ScanRefer, and 3.7% leads on Sr3D. Zeyue Wang 0002, Xixia Xu, Dongchen Zhu, Jiamao Li |
IROS | 2 |
| 2025 | Human-object interaction detection via recycling of ground-truth annotations
Xue Lin 0003, Qi Zou 0001, Xixia Xu |
Pattern Recognit. | 3 |
| 2025 | Multi-Person Pose Estimation with Feature Enhancement and Decoupling Based on Contrastive LearningabstractMost methods of multi-person pose estimation (MPPE) treat the human detection and keypoint localization separately. They need additional supervision like instance bounding boxes, or complex hand-crafted processes like RoI cropping or grouping. In this article, we propose a novel one-stage MPPE method, named COPE, which unifies human detection and keypoint regression into an end-to-end learnable framework. To handle the challenges plague one-stage MPPE, i.e., instance overlapping and misalignment of local and global context, we design contrastive constraints at two levels of semantic granularity and feature sampling strategies. Based on a whole-process differentiable pipeline, COPE establishes a simple yet effective framework for MPPE without additional instance-level supervision and resource-intensive modules like transformer. Benefit from specially designed contrastive constraints and sampling strategies, COPE can better handle occluded scenes and correct keypoint localization errors. Extensive experiments demonstrate COPE’s superiority. It attains 71.3 AP and 18.0 FPS on COCO val2017, effectively balancing accuracy and speed. Particularly in crowded and occluded scenarios, COPE achieves state-of-the-art performance on CrowdPose and OCHuman, surpassing CID by 0.6 AP and 1.7 AP, respectively. Furthermore, COPE strongly improves generalization performance on the Human-Art benchmark, outperforming ED-Pose by 6.7 AP and ClickPose by 3.7 AP. Qi Zou 0001, Xixia Xu, Yanting Pei |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Rethinking the Sparse End-to-End Multiperson Pose EstimationabstractCurrent methods of multiperson pose estimation (MPPE) typically treat the human detection and association of joints separately. They introduce complex hand-crafted pose-processes like RoI cropping, NMS and grouping or rely on dense representations to preserve the spatial features. In this article, we dive a deeper thought into this task and propose a simpler and effective framework, termed SparsePose, which can directly predict multiperson joint coordinates from the full image without any post-processes and dense representations. In SparsePose, the full-body instances are decoupled by exploring spatial-aware feature learning (SFL) without box and classification supervision. For improving the quality of instance map, the instance contrastive constraint (ICC) and center correction (CC) strategy are proposed to make the instance-wise spatial feature more discriminative. Importantly, we propose a visibility-guided weighting mechanism to enable model be confident to the visible joint predictions and insensitive to the occlusions or partial bodies. In general, SparsePose is conceptually simpler and plays favorably against the existing counterparts on three benchmarks in terms of both accuracy and efficiency. Xixia Xu, Qi Zou 0001, Jiamao Li |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Inter-image Contrastive Consistency for Multi-Person Pose EstimationabstractMulti-person pose estimation (MPPE) has achieved impressive progress in recent years. However, due to the large variance of appearances among images or occlusions, the model can hardly learn consistent patterns enough, which leads to severe location jitter and missing issues. In this study, we propose a novel framework, termed Inter-image Contrastive consistency (ICON), to strengthen the keypoint consistency among images for MPPE. Concretely, we consider two-fold consistency constraints, which include single keypoint contrastive consistency (SKCC) and pair relation contrastive consistency (PRCC). The SKCC learns to strengthen the consistency of individual keypoints across images in the same category to improve the category-specific robustness. Only with SKCC, the model can effectively reduce location errors caused by large appearance variations, but remains challenging with extreme postures (e.g., occlusions) due to lack of relational guidance. Therefore, PRCC is proposed to strengthen the consistency of pair-wise joint relation between images to preserve the instructive relation. Cooperating with SKCC, PRCC further improves structure aware robustness in handling extreme postures. Extensive experiments on kinds of architectures across three datasets (i.e., MS-COCO, MPII, CrowdPose) show the proposed ICON achieves substantial improvements over baselines. Furthermore, ICON under the semi-supervised setup can obtain comparable results with the fully-supervised methods using only 30% labeled data. Xixia Xu, Yingguo Gao, Xingjia Pan, Qi Zou 0001 |
AAAI | 1 |
| 2023 | Structure-Enriched Topology Learning For Cross-Domain Multi-Person Pose EstimationabstractHuman pose estimation has been widely studied with much focus on supervised learning. However, in real applications, a pretrained pose estimation model usually needs be adapted to a novel domain without labels or with sparse labels. Existing domain adaptation methods cannot well deal with it since poses have flexible topological structures and need fine-grained local features. Aiming at the characteristics of human pose, we propose a novel domain adaptation method for multi-person pose estimation (MPPE) to alleviate the human-level shift. Firstly, the training samples of human poses are clustered into groups according to the posture similarity. Within the clustered space, we conduct three adaptation modules: Cross-Attentive Feature Alignment (CAFA), Intra-domain Structure Adaptation (ISA) and Adaptive Human-Topology Adaptation (AHTA). The CAFA adopts a bidirectional spatial attention mechanism to explore fine-grained local feature correlation between two humans, and thus to adaptively aggregate consistent features for adaptation. ISA only works in semi-supervised domain adaptation (SSDA) to exploit semantic relationship of corresponding keypoints for reducing the intra-domain bias. Importantly, we creatively propose an AHTA to enrich human topological knowledge for reducing the inter-domain discrepancy. Specifically, the pose structure and the cross-instance topological relations are modeled via graph networks. This flexible topology learning benefits the occluded or extreme pose inference. Extensive experiments are conducted on two popular benchmarks and additional two challenging datasets. Results demonstrate the competency of our method, which works in unsupervised or semi-supervised modes, compared with the existing supervised approaches. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | Effects of Motion-Relevant Knowledge From Unlabeled Video to Human-Object Interaction DetectionabstractThe existing works on human-object interaction (HOI) detection usually rely on expensive large-scale labeled image datasets. However, in real scenes, labeled data may be insufficient, and some rare HOI categories have few samples. This poses great challenges for deep-learning-based HOI detection models. Existing works tackle it by introducing compositional learning or word embedding but still need large-scale labeled data or extremely rely on the well-learned knowledge. In contrast, the freely available unlabeled videos contain rich motion-relevant information that can help infer rare HOIs. In this article, we creatively propose a multitask learning (MTL) perspective to assist in HOI detection with the aid of motion-relevant knowledge learning on unlabeled videos. Specifically, we design the appearance reconstruction loss (ARL) and sequential motion mining module in a self-supervised manner to learn more generalizable motion representations for promoting the detection of rare HOIs. Moreover, to better transfer motion-related knowledge from unlabeled videos to HOI images, a domain discriminator is introduced to decrease the domain gap between two domains. Extensive experiments on the HICO-DET dataset with rare categories and the V-COCO dataset with minimum supervision demonstrate the effectiveness of motion-aware knowledge implied in unlabeled videos for HOI detection. Xue Lin 0003, Qi Zou 0001, Xixia Xu, Ding Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Adaptive Hypergraph Neural Network for Multi-Person Pose EstimationabstractThis paper proposes a novel two-stage hypergraph-based framework, dubbed ADaptive Hypergraph Neural Network (AD-HNN) to estimate multiple human poses from a single image, with a keypoint localization network and an Adaptive-Pose Hypergraph Neural Network (AP-HNN) added onto the former network. For providing better guided representations of AP-HNN, we employ a Semantic Interaction Convolution (SIC) module within the initial localization network to acquire more explicit predictions. Build upon this, we design a novel adaptive hypergraph to represent a human body for capturing high-order semantic relations among different joints. Notably, it can adaptively adjust the relations between joints and seek the most reasonable structure for the variable poses to benefit the keypoint localization. These two stages are combined to be trained in an end-to-end fashion. Unlike traditional Graph Convolutional Networks (GCNs) that are based on a fixed tree structure, AP-HNN can deal with ambiguity in human pose estimation. Experimental results demonstrate that the AD-HNN achieves state-of-the-art performance both on the MS-COCO, MPII and CrowdPose datasets. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
AAAI | 1 |
| 2022 | Location-Free Human Pose EstimationabstractHuman pose estimation (HPE) usually requires large-scale training data to reach high performance. However, it is rather time-consuming to collect high-quality and fine-grained annotations for human body. To alleviate this issue, we revisit HPE and propose a location-free framework without supervision of keypoint locations. We reformulate the regression-based HPE from the perspective of classification. Inspired by the CAM-based weakly-supervised object localization, we observe that the coarse keypoint locations can be acquired through the part-aware CAMs but unsatisfactory due to the gap between the fine-grained HPE and the object-level localization. To this end, we propose a customized transformer framework to mine the fine-grained representation of human context, equipped with the structural relation to capture subtle differences among keypoints. Concretely, we design a Multi-scale Spatial-guided Context Encoder to fully capture the global human context while focusing on the part-aware regions and a Relation-encoded Pose Prototype Generation module to encode the structural relations. All these works together for strengthening the weak supervision from image-level category labels on locations. Our model achieves competitive performance on three datasets when only supervised at a category-level and importantly, it can achieve comparable results with fully-supervised methods with only 25% location labels on MS-COCO and MPII. Xixia Xu, Yingguo Gao, Xue Lin 0003, Qi Zou 0001 |
CVPR | 1 |
| 2022 | CFENet: Content-aware feature enhancement network for multi-person pose estimation
Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
Appl. Intell. | 1 |
| 2022 | A Stronger Baseline for Seismic Facies Classification With Less DataabstractWith the great success of deep learning in computer vision, the application of convolution neural network (CNN) in seismic facies classification is growing rapidly. However, most of the previous works based on pure state-of-the-art CNN architectures still suffer from coarse segmentation results. In this article, we study the challenges of seismic facies classification and propose a stronger baseline. More specifically, we propose a simple yet effective unsupervised approach named spatial pyramid sampling (SPS) to choose representative samples for training to reduce the labeling costs. Next, we propose a multimodal fusion (M2F) module to extract and fuse the edge and frequency information from selected seismic images to build a stable multimodal representation. Finally, we propose a local-to-global (L2G) module, which improves the recognition power by capturing the local relationship between pixels and enhancing the global context representation. Experimental results demonstrate that the proposed method achieves a superior performance with less labeled training data, especially for small categories. Qi Zou 0001, Xixia Xu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Motion-Aware Feature Enhancement Network for Video PredictionabstractVideo prediction is challenging, due to the pixel-level precision requirement and the difficulty in capturing scene dynamics. Most approaches tackle the problems by pixel-level reconstruction objectives and two decomposed branches, which still suffer from blurry generations or dramatic degradations in long-term prediction. In this paper, we propose a Motion-Aware Feature Enhancement (MAFE) network for video prediction to produce realistic future frames and achieve relatively long-term predictions. First, a Channel-wise and Spatial Attention (CSA) module is designed to extract motion-aware features, which enhances the contribution of important motion details during encoding, and subsequently improves the discriminability of attention map for the frame refinement. Second, a Motion Perceptual Loss (MPL) is proposed to guide the learning of temporal cues, which benefits to robust long-term video prediction. Extensive experiments on three human activity video datasets: KTH, Human3.6M, and PennAction demonstrate the effectiveness of the proposed video prediction model compared with the state-of-the-art approaches. Xue Lin 0003, Qi Zou 0001, Xixia Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Action-Guided Attention Mining and Relation Reasoning Network for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection is important to understand human-centric scenes and is challenging due to subtle difference between fine-grained actions, and multiple co-occurring interactions. Most approaches tackle the problems by considering the multi-stream information and even introducing extra knowledge, which suffer from a huge combination space and the non-interactive pair domination problem. In this paper, we propose an Action-Guided attention mining and Relation Reasoning (AGRR) network to solve the problems. Relation reasoning on human-object pairs is performed by exploiting contextual compatibility consistency among pairs to filter out the non-interactive combinations. To better discriminate the subtle difference between fine-grained actions, an action-aware attention based on class activation map is proposed to mine the most relevant features for recognizing HOIs. Extensive experiments on V-COCO and HICO-DET datasets demonstrate the effectiveness of the proposed model compared with the state-of-the-art approaches. Xue Lin 0003, Qi Zou 0001, Xixia Xu |
IJCAI | 3 |
| 2020 | Alleviating Human-level Shift: A Robust Domain Adaptation Method for Multi-person Pose EstimationabstractHuman pose estimation has been widely studied with much focus on supervised learning requiring sufficient annotations. However, in real applications, a pretrained pose estimation model usually need be adapted to a novel domain with no labels or sparse labels. Such domain adaptation for 2D pose estimation hasn't been explored. The main reason is that a pose, by nature, has typical topological structure and needs fine-grained features in local keypoints. While existing adaptation methods do not consider topological structure of object-of-interest and they align the whole images coarsely. Therefore, we propose a novel domain adaptation method for multi-person pose estimation to conduct the human-level topological structure alignment and fine-grained feature alignment. Our method consists of three modules: Cross-Attentive Feature Alignment (CAFA), Intra-domain Structure Adaptation (ISA) and Inter-domain Human-Topology Alignment (IHTA) module. The CAFA adopts a bidirectional spatial attention module (BSAM) that focuses on fine-grained local feature correlation between two humans to adaptively aggregate consistent features for adaptation. We adopt ISA only in semi-supervised domain adaptation (SSDA) to exploit the corresponding keypoint semantic relationship for reducing the intra-domain bias. Most importantly, we propose an IHTA to learn more domain-invariant human topological representation for reducing the inter-domain discrepancy. We model the human topological structure via the graph convolution network (GCN), by passing messages on which, high-order relations can be considered. This structure preserving alignment based on GCN is beneficial to the occluded or extreme pose inference. Extensive experiments are conducted on two popular benchmarks and results demonstrate the competency of our method compared with existing supervised approaches. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
ACM Multimedia | 1 |
| 2020 | Integral Knowledge Distillation for Multi-Person Pose EstimationabstractBoth accuracy and efficiency are of equal importance to the human pose estimation. Most of the existing methods simply pursue excellent performance, sacrificing massive computing resources and memory. Out of this consideration, we present a novel compact and lightweight framework to train more efficient estimators using knowledge distillation. Three distillation mechanisms are proposed in our method from different perspectives, including logit distillation, feature distillation and structure distillation. Concretely, the logit distillation regards the output of teacher model as soft target to stimulate the student model. The feature distillation distills the high-level features of the teacher model to assist the student. Unlike the above strategies, the structure distillation considers the problem in a global view, aiming at ensuring the student prediction contains quite abundant structure knowledge like the teacher. We empirically demonstrate the effectiveness and efficiency of our methods on two multi-person pose estimation datasets (COCO and MPII). Specifically, our model can achieve competitive performance with the most state-of-the-art methods and consume only 35% model parameters and GFLOPs of our baseline (SimpleBaseline-ResNet-50) on the COCO dataset. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
IEEE Signal Process. Lett. | 1 |