Xueqing Ma

dblp:317/3205 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Face, body and person analysis · 64% Video understanding and tracking · 16% Generative modeling · 16%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
human pose estimation
1.322023
DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation · ICCV 2023
Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video · CVPR 2023
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation
1.322023
DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation · ICCV 2023
Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video · CVPR 2023
Machine learning › Generative modeling
diffusion model
0.712023
DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation · ICCV 2023
Computer vision › Video understanding and tracking
temporal modeling
0.712023
Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.212023
Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video · CVPR 2023

Methods — techniques the papers use, named apart from their topics

temporal difference learning · 0.7spatiotemporal representation learning · 0.7mutual information · 0.7multi-scale feature interaction · 0.7deformable convolution · 0.7
YearPublicationVenuePosition
2026 Joint Optimization of Task Offloading, Resource Allocation, and Trajectory Design in Cooperative Multi-UAV MEC Networks
abstract
Uncrewed Aerial Vehicle (UAV)-assisted Mobile Edge Computing (MEC) systems provide flexible and resilient computing capabilities for mobile users by leveraging UAVs as edge servers (ESs). However, in practical deployments, user devices typically exhibit spatially non-uniform distributions, which impose significant challenges for achieving optimal UAV placement. Conventional fixed or pre-determined deployment strategies cannot dynamically adapt to heterogeneous user distributions. To overcome these challenges, this article investigates an air–ground cooperative MEC architecture comprising multiple UAVs and terrestrial ESs. A multi-objective optimization problem incorporating UAV trajectory optimization is formulated to jointly minimize task offloading latency and total system energy consumption. As the problem is an NP-hard mixed-integer nonlinear programming model, a Joint Alternating Optimization framework for Task Offloading, Resource Allocation, and UAV Trajectory control (JAOTRU) is proposed. The JAOTRU framework adopts an iterative block coordinate descent structure to decouple the highly coupled optimization variables into two subproblems, which are efficiently solved via a differential evolution algorithm and a successive convex approximation technique, respectively. The simulation results show that the proposed JAOTRU method substantially surpasses several benchmark methods regarding total offloading latency and system energy efficiency, validating its effectiveness and robustness for MEC systems assisted by UAVs.
Qin Wang 0002, Xueqing Ma, Yongxu Zhu, Wenchao Xia, Hangsheng Zhao
IEEE Internet Things J.2
2023 Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video
abstract
Temporal modeling is crucial for multi-frame human pose estimation. Most existing methods directly employ optical flow or deformable convolution to predict full-spectrum motion fields, which might incur numerous irrelevant cues, such as a nearby person or background. Without further efforts to excavate meaningful motion priors, their results are suboptimal, especially in complicated spatio-temporal interactions. On the other hand, the temporal difference has the ability to encode representative motion information which can potentially be valuable for pose estimation but has not been fully exploited. In this paper, we present a novel multi-frame human pose estimation framework, which employs temporal differences across frames to model dynamic contexts and engages mutual information objectively to facilitate useful motion information disentanglement. To be specific, we design a multi-stage Temporal Difference Encoder that performs incremental cascaded learning conditioned on multi-stage feature difference sequences to derive informative motion representation. We further propose a Representation Disentanglement module from the mutual information perspective, which can grasp discriminative task-relevant motion signals by explicitly defining useful and noisy constituents of the raw motion features and minimizing their mutual information. These place us to rank No.1 in the Crowd Pose Estimation in Complex Events Challenge on benchmark dataset HiEve, and achieve state-of-the-art performance on three benchmarks PoseTrack2017, PoseTrack2018, and PoseTrack21.
Runyang Feng, Yixing Gao 0001, Xueqing Ma, Tze Ho Elden Tse, Hyung Jin Chang
CVPR3
2023 DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation
abstract
Denoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining attention in computer vision. However, extending such models to multi-frame human pose estimation is non-trivial due to the presence of the additional temporal dimension in videos. More importantly, learning representations that focus on keypoint regions is crucial for accurate localization of human joints. Nevertheless, the adaptation of the diffusion-based methods remains unclear on how to achieve such objective. In this paper, we present DiffPose, a novel diffusion architecture that formulates video-based human pose estimation as a conditional heatmap generation problem. First, to better leverage temporal information, we propose SpatioTemporal Representation Learner which aggregates visual evidences across frames and uses the resulting features in each denoising step as a condition. In addition, we present a mechanism called Lookup-based Multi-Scale Feature Interaction that determines the correlations between local joints and global contexts across multiple scales. This mechanism generates delicate representations that focus on keypoint regions. Altogether, by extending diffusion models, we show two unique characteristics from DiffPose on pose estimation task: (i) the ability to combine multiple sets of pose estimates to improve prediction accuracy, particularly for challenging joints, and (ii) the ability to adjust the number of iterative steps for feature refinement without retraining the model. DiffPose sets new state-of-the-art results on three benchmarks: PoseTrack2017, PoseTrack2018, and PoseTrack21.
Runyang Feng, Yixing Gao 0001, Tze Ho Elden Tse, Xueqing Ma, Hyung Jin Chang
ICCV4
2023 Knowing Before Seeing: Incorporating Post-retrieval Information into Pre-retrieval Query Intention Classification
Xueqing Ma, Xiaochi Wei, Yixing Gao 0001, Runyang Feng, Dawei Yin 0001, Yi Chang 0001
KSEM (2)1