Ying Huang 0003

dblp:62/2964-3 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-6239-192XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Face, body and person analysis · 50% Efficient and distributed learning · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis › human pose estimation
efficient pose estimation
0.512021
Online Knowledge Distillation for Efficient Pose Estimation · ICCV 2021
Computer vision › Face, body and person analysis
human pose estimation
0.512021
Online Knowledge Distillation for Efficient Pose Estimation · ICCV 2021
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.512021
Online Knowledge Distillation for Efficient Pose Estimation · ICCV 2021
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
online knowledge distillation
0.512021
Online Knowledge Distillation for Efficient Pose Estimation · ICCV 2021

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 0.5feature aggregation · 0.5KL divergence · 0.5
YearPublicationVenuePosition
2025 LKConvPose: A Pose Estimation Model with Large Receptive Field
abstract
Recently, significant progress has been made in 2D human pose estimation. While some research has focused on enhancing the accuracy of keypoint detection, others have aimed at reducing model size. However, most models excel in either one aspect or the other, but rarely both simultaneously. In this paper, we address the challenge of balancing accuracy and inference speed. Inspired by large-kernel convolutions and attention mechanisms, we introduce LKConvPose, a hybrid CNN architecture that achieves high keypoint detection accuracy with low computational cost. Specifically, LKConvPose-S attains 76.5 AP on the COCO validation dataset using only 4 GFLOPs, making it the most efficient model at its scale.
Ying Huang 0003, Xiu-Xiu Zhan, Jianzhang Zhang, Chuang Liu 0001
ICASSP1
2024 MIMIC-Pose: Implicit Membership Discrimination of Body Joints for Human Pose Estimation
abstract
The objective of human pose estimation is to accurately identify the positions of individual body joints for each person seen in an image for human kinematics modelling and analysis. A majority of current deep-learning approaches focus on utilizing feature learning to regress the coordinates of each individual joint. This can be thought of as a constrained point distribution optimization problem in the image plane. However, in the high-dimensional feature space, the feature distribution of joints as a significant constraint and discriminative condition, has not received enough attention yet. Here we propose a novel approach (MIMIC-Pose) that applies an implicit pair-wise keypoint membership constraint using those features to individual joint regression to ensure that the joints of the same person have higher mutual feature similarities compared with those of different people or the image backgrounds. We propose a novel link prediction integrated learning architecture with high-throughput tokens to learn such similarity feature embeddings to improve the prediction accuracy of individual joints in a single forward inference. Experimental results show that our approach can achieve competitive performance with much fewer model parameters and lower computational cost compared to state-of-the-art methods.
Ying Huang 0003, Shanfeng Hu
FG1
2024 TOPOMA: Time-Series Orthogonal Projection Operator with Moving Average for Interpretable and Training-Free Anomaly Detection
Shanfeng Hu, Ying Huang 0003
PAKDD (1)2
2022 Structured Spatial Reasoning for Human Pose Estimation
Ying Huang 0003, Shanfeng Hu, Zi-Ke Zhang
BMVC1
2021 Online Knowledge Distillation for Efficient Pose Estimation
abstract
Existing state-of-the-art human pose estimation methods require heavy computational resources for accurate predictions. One promising technique to obtain an accurate yet lightweight pose estimator is knowledge distillation, which distills the pose knowledge from a powerful teacher model to a less-parameterized student model. However, existing pose distillation works rely on a heavy pre-trained estimator to perform knowledge transfer and require a complex two-stage learning procedure. In this work, we investigate a novel Online Knowledge Distillation framework by distilling Human Pose structure knowledge in a one-stage manner to guarantee the distillation efficiency, termed OKDHP. Specifically, OKDHP trains a single multi-branch network and acquires the predicted heatmaps from each, which are then assembled by a Feature Aggregation Unit (FAU) as the target heatmaps to teach each branch in reverse. Instead of simply averaging the heatmaps, FAU which consists of multiple parallel transformations with different receptive fields, leverages the multi-scale information, thus obtains target heatmaps with higher-quality. Specifically, the pixel-wise Kullback-Leibler (KL) divergence is utilized to mini-mize the discrepancy between the target heatmaps and the predicted ones, which enables the student network to learn the implicit keypoint relationship. Besides, an unbalanced OKDHP scheme is introduced to customize the student networks with different compression rates. The effectiveness of our approach is demonstrated by extensive experiments on two common benchmark datasets, MPII and COCO.
Zheng Li 0028, Jingwen Ye, Mingli Song, Ying Huang 0003
ICCV4
2020 Online Knowledge Distillation via Multi-branch Diversity Enhancement
Zheng Li 0028, Ying Huang 0003, Defang Chen 0001, Tianren Luo
ACCV (4)2
2020 High-speed multi-person pose estimation with deep feature transfer
Ying Huang 0003, Hubert P. H. Shum, Edmond S. L. Ho, Nauman Aslam
Comput. Vis. Image Underst.1
2019 Multi-Level Network for High-Speed Multi-Person Pose Estimation
abstract
In multi-person pose estimation, the left/right joint type discrimination is always a hard problem because of the similar appearance. Traditionally, we solve this problem by stacking multiple refinement modules to increase network's receptive fields and capture more global context, which can also increase a great amount of computation. In this paper, we propose a Multi-level Network (MLN) that learns to aggregate features from lower-level (left/right information), upper-level (localization information), joint-limb level (complementary information) and global-level (context) information for discrimination of joint type. Through feature reuse and its intra-relation, MLN can attain comparable performance to other conventional methods while runtime speed retains at 42 FPS.
Ying Huang 0003, Jiankai Zhuang, Zengchang Qin
ICIP1
2019 FollowMeUp Sports: New Benchmark for 2D Human Keypoint Recognition
Ying Huang 0003, Haipeng Kan, Jiankai Zhuang, Zengchang Qin
PRCV (3)1
2016 Improving an object tracker for infrared flying bird tracking
abstract
We propose an approach to improve the tracking performance of a generic tracker when it is applied to infrared flying bird tracking task. Since the flying bird is a fast-moving small object, its drastic changes of shape and scale, and cluttered background can all cause the performance degradation of generic trackers. Moreover, the gray intensity also weakens the discriminative power of object or background model. In our approach, we apply edge information and segmentation to refine the output of a generic tracker. Edge information as heuristic cues to delimit object region, and a high-order appearance separation term is used to segment finer object region. Their combination enables tracking algorithm to accommodate bird deformation and scale changes. In general, our method takes estimated object location of a generic tracker as input and generates the segmentation output with higher tracking precision and exact object regions. In the framework of online learning, we use segmentation output as the input of online learner to increase the accuracy of learning and enhance the discriminative power of classifier. The experiments are performed on our BIRDSITE-IR dataset. The results demonstrate the effectiveness of our approach and improves the state-of-the-art trackers by 5% at least in average tracking precision.
Ying Huang 0003
ICIP1