VLDB 2026 Research / reviewers in the wild / expert
Jie Xu 0021
dblp:37/5126-21
· DBLP profile ↗
9ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-7123-8919ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Adaptive Pedestrian Trajectory Prediction via Target-Directed Angle AugmentationabstractPedestrian trajectory prediction is an important task for many applications such as autonomous driving and surveillance systems. Yet the prediction performance drops dramatically when applying a model trained on the source domain to a new target domain. Therefore, it is of great importance to adapt a predictor to a new domain. Previous works mainly focus on feature-level alignment to solve this problem. In contrast, we solve it from a new perspective of instance-level alignment. Specifically, we first point out one key factor of the domain gaps, i.e., trajectory angles, and then augment the source training data by target-directed orientation augmentation so that its distribution matches with that of the target data. In this way, the trajectory predictor trained on the aligned source data performs better on the target domain. Experiments on standard baselines show that our method improves the state of the art by a large margin. The source code is available at https://github.com/NeoKH/PTP-DA. Jie Xu 0021, Shenjian Gong, Jian Yang 0003, Shanshan Zhang 0001 |
ICASSP | 2 |
| 2023 | Adaptive Decoupled Pose Knowledge DistillationabstractExisting state-of-the-art human pose estimation approaches require heavy computational resources for accurate prediction. One promising technique to obtain an accurate yet lightweight pose estimator is Knowledge Distillation (KD), which distills the pose knowledge from a powerful teacher model to a lightweight student model. However, existing human pose KD methods focus more on designing paired student and teacher network architectures, yet ignore the mechanism of pose knowledge distillation. In this work, we reformulate the human pose KD to a coarse to fine process and decouple the classical KD loss into three terms: Binary Keypoint vs. Non-Keypoint Distillation (BiKD), Keypoint Area Distillation (KAD) and Non-keypoint Area Distillation (NAD). Observing the decoupled formulation, we point out an important limitation of the classical pose KD, i.e. the bias between different loss terms limits the performance gain of the student network. To address the biased knowledge distillation problem, we present a novel KD method named Adaptive Decoupled Pose knowledge Distillation (ADPD), enabling BiKD, KAD and NAD to play their roles more effectively and flexibly. Extensive experiments on two standard human pose datasets, MPII and MS COCO, demonstrate that our proposed method outperforms previous KD methods and is generalizable to different teacher-student pairs. The code will be available at https://github.com/SuperJay1996/ADPD. Jie Xu 0021, Shanshan Zhang 0001, Jian Yang 0003 |
ACM Multimedia | 1 |
| 2022 | Towards High Performance One-Stage Human Pose EstimationabstractMaking top-down human pose estimation method present both good performance and high efficiency is appealing. Mask RCNN can largely improve the efficiency by conducting person detection and pose estimation in a single framework, as the features provided by the backbone are able to be shared by the two tasks. However, the performance is not as good as traditional two-stage methods. In this paper, we aim to largely advance the human pose estimation results of Mask-RCNN and still keep the efficiency. Specifically, we make improvements on the whole process of pose estimation, which contains feature extraction and keypoint detection. The part of feature extraction is ensured to get enough and valuable information of pose. Then, we introduce a Global Context Module into the keypoints detection branch to enlarge the receptive field, as it is crucial to successful human pose estimation. On the COCO val2017 set, our model using the ResNet-50 backbone achieves an AP of 68.1, which is 2.6 higher than Mask RCNN (AP of 65.5). Compared to the classic two-stage top-down method SimpleBaseline, our model largely narrows the performance gap (68.1 APkp vs. 68.9 APkp) with a much faster inference speed (77 ms vs. 168 ms), demonstrating the effectiveness of the proposed method. Code is available at: https://github.com/lingl_space/maskrcnn_keypoint_refined. Lin Zhao 0003, Linhao Xu, Jie Xu 0021 |
MMAsia | 4 |
| 2021 | Tiny Person Pose Estimation via Image and Feature Super Resolution
Jie Xu 0021, Yunan Liu 0001, Lin Zhao 0003, Shanshan Zhang 0001, Jian Yang 0003 |
ICIG (3) | 1 |
| 2021 | Single image super-resolution via hybrid resolution NSST prediction
Yunan Liu 0001, Shanshan Zhang 0001, Chunpeng Wang 0001, Jie Xu 0021 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Learning to Acquire the Quality of Human Pose EstimationabstractMaking human poses serve high-level computer vision tasks such as action recognition, recognizing the quality of estimated poses is of critical importance. Conventionally, the mean confidence of each keypoint is used as pose quality in most human pose estimation frameworks. However, because different types of keypoint are not identical in visibility and size, they should not contribute equally, which produces biased quality scores. In the paper, we propose end-to-end human pose quality learning, which adds a quality prediction block alongside pose regression. The proposed block learns the object keypoint similarity (OKS) between the estimated pose and its corresponding ground truth by sharing the pose features with heatmap regression. The predicted OKS correlates well with pose quality, making the selection of reliable poses straightforward. Moreover, utilizing the learned quality as pose score improves pose estimation performance during COCO AP evaluation, because it ranks more accurate ones high among all pose detections. We conduct extensive experiments based on the three most popular human pose estimation frameworks, including Hourglass, SimpleBaseline and HRNet. Adding the proposed quality learning block is able to consistently bring nearly 1 percent AP improvement on all the frameworks. Lin Zhao 0003, Jie Xu 0021, Chen Gong 0002, Jian Yang 0003, Wangmeng Zuo, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | An Accurate and Lightweight Method for Human Body Image Super-ResolutionabstractIn this paper, we propose a new method to super-resolve low resolution human body images by learning efficient multi-scale features and exploiting useful human body prior. Specifically, we propose a lightweight multi-scale block (LMSB) as basic module of a coherent framework, which contains an image reconstruction branch and a prior estimation branch. In the image reconstruction branch, the LMSB aggregates features of multiple receptive fields so as to gather rich context information for low-to-high resolution mapping. In the prior estimation branch, we adopt the human parsing maps and nonsubsampled shearlet transform (NSST) sub-bands to represent the human body prior, which is expected to enhance the details of reconstructed human body images. When evaluated on the newly collected HumanSR dataset, our method outperforms state-of-the-art image super-resolution methods with ∼ 8× fewer parameters; moreover, our method significantly improves the performance of human image analysis tasks (e.g. human parsing and pose estimation) for low-resolution inputs. Yunan Liu 0001, Shanshan Zhang 0001, Jie Xu 0021, Jian Yang 0003, Yu-Wing Tai |
IEEE Trans. Image Process. | 3 |
| 2020 | Perceiving heavily occluded human poses by assigning unbiased score
Lin Zhao 0003, Jie Xu 0021, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003, Xinbo Gao 0001 |
Inf. Sci. | 2 |
| 2020 | Multi-task learning for object keypoints detection and classification
Jie Xu 0021, Lin Zhao 0003, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003 |
Pattern Recognit. Lett. | 1 |