EDBT 2026 Demo / reviewers in the wild / expert
Zhuoyue Xu
dblp:398/2571
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Face, body and person analysis · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
human pose estimation |
1.7 | 2 | 2025 | SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos · AAAI 2025 Optimizing Human Pose Estimation Through Focused Human and Joint Regions · AAAI 2025 |
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation |
1.7 | 2 | 2025 | SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos · AAAI 2025 Optimizing Human Pose Estimation Through Focused Human and Joint Regions · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9spatiotemporal representation learning · 0.9dynamic-aware mask · 0.9deformable cross-attention · 0.9bidirectional separation strategy · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing Human Pose Estimation Through Focused Human and Joint RegionsabstractHuman pose estimation has given rise to a broad spectrum of novel and compelling applications, including action recognition, sports analysis, as well as surveillance. However, accurate video pose estimation remains an open challenge. One aspect that has been overlooked so far is that existing methods learn motion clues from all pixels rather than focusing on the target human body, making them easily misled and disrupted by unimportant information such as background changes or movements of other people. Additionally, while the current Transformer-based pose estimation methods has demonstrated impressive performance with global modeling, they struggle with local context perception and precise positional identification. In this paper, we try to tackle these challenges from three aspects: (1) We propose a bilayer Human-Keypoint Mask module that performs coarse-to-fine visual token refinement, which gradually zooms in on the target human body and keypoints while masking out unimportant figure regions. (2) We further introduce a novel deformable cross attention mechanism and a bidirectional separation strategy to adaptively aggregate spatial and temporal motion clues from constrained surrounding contexts. (3) We mathematically formulate the deformable cross attention, constraining that the model focuses solely on the regions centered at the target person body. Empirically, our method achieves state-of-the-art performance on three large-scale benchmark datasets. A remarkable highlight is that our method achieves an 84.8 mean Average Precision (mAP) on the challenging wrist joint, which significantly outperforms the 81.5 mAP achieved by the current state-of-the-art method on the PoseTrack2017 dataset. Yingying Jiao, Zhenguang Liu, Shaojing Fan, Sifan Wu 0001, Zheqi Wu, Zhuoyue Xu |
AAAI | 7 |
| 2025 | SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled VideosabstractHuman pose estimation in videos remains a challenge, largely due to the reliance on extensive manual annotation of large datasets, which is expensive and labor-intensive. Furthermore, existing approaches often struggle to capture long-range temporal dependencies and overlook the complementary relationship between temporal pose heatmaps and visual features. To address these limitations, we introduce STDPose, a novel framework that enhances human pose estimation by learning spatiotemporal dynamics in sparsely-labeled videos. STDPose incorporates two key innovations: 1) A novel Dynamic-Aware Mask to capture long-range motion context, allowing for a nuanced understanding of pose changes. 2) A system for encoding and aggregating spatiotemporal representations and motion dynamics to effectively model spatiotemporal relationships, improving the accuracy and robustness of pose estimation. STDPose establishes a new performance benchmark for both video pose propagation (i.e., propagating pose annotations from labeled frames to unlabeled frames) and pose estimation tasks, across three large-scale evaluation datasets. Additionally, utilizing pseudo-labels generated by pose propagation, STDPose achieves competitive performance with only 26.7% labeled data. Yingying Jiao, Sifan Wu 0001, Shaojing Fan, Zhenguang Liu, Zhuoyue Xu, Zheqi Wu |
AAAI | 6 |