EDBT 2026 Demo / reviewers in the wild / expert
Haolun Li 0001
dblp:221/0626
· DBLP profile ↗
25ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0003-3829-2849ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SNS-Grasp: Semantic-guided Noise Scaling for Grasp GenerationabstractWhile diffusion models show promise for intent-based grasp generation, their isotropic noise schedules struggle with joint-specific sensitivity and task-aware variability. This limitation leads to grasps with suboptimal semantic alignment or physical feasibility. To address this challenge, we propose Semantic-guided Noise Scaling for grasp generation (SNS-Grasp), a novel framework that integrates two key innovations. First, the Semantic-guided Noise Scaling Diffusion (SNS-Diff) module generates intent-aware grasps by replacing isotropic noise with anisotropic modulation, dynamically adapting to task semantics and joint-specific sensitivity. Specifically, SNS-Diff leverages a pretrained Intent Recognizer to extract task-aware confidence scores and joint-specific gradient sensitivities from the interaction context. These signals adjust the noise scaling during denoising, downweighting perturbations for semantically critical joints to ensure semantic alignment. Second, the Fine-grained Grasp Refinement (FGR) module establishes dynamic joint-vertex coupling through fine-grained hand-object spatial relationships, enabling iterative optimization of physically executable grasps. Extensive experiments on OakInk and GRAB demonstrate SNS-Grasp's superior performance in semantic accuracy and physical feasibility, with robust generalization to unseen objects. Zhenhua Tang 0001, Yudian Zheng, Yuzhang Zhong, Haolun Li 0001, Yanbin Hao, Chi-Man Pun |
AAAI | 4 |
| 2026 | GPDPose: Self-supervised transformer with geometry, pose, and depth consistency for multi-view 3D human pose estimation
Jucheng Song, Jie Zhang 0090, Xu Yang 0010, Yapeng Wang 0001, Hao Gao 0005, Haolun Li 0001, Sio Kei Im |
Expert Syst. Appl. | 6 |
| 2026 | An adaptive fuzzy mathematical morphology image denoising method using semi-overlap functions
Haolun Li 0001 |
Inf. Sci. | 1 |
| 2026 | Effective Gaussian Management for High-Fidelity Scene ReconstructionabstractThis paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent Gaussian Splatting (GS) pipelines that treat all primitives uniformly during optimization, our framework explicitly manages the attribute activation, representation and pruning of Gaussian. Specifically, our framework first introduces GauSep, a novel densification strategy that selectively activates Gaussian color or normal attributes to alleviate destructive gradient conflicts arising from dual supervision. We further propose GauRep, an adaptive Gaussian representation that dynamically adjusts spherical harmonics (SHs) orders and performs task-decoupled pruning to reduce redundancy at both the individual and global levels. To provide reliable geometric supervision for above mangement process, we additionally introduce CoRe, an regularized surface reconstruction module that distills robust normal fields from an SDF branch to the Gaussian representation through a confidence mechanism. Notably, the proposed Gaussian management is compatible with various reconstruction architectures and can be seamlessly integrated to improve performance while reducing size of the model. Extensive experiments demonstrate that our approach achieves superior or comparable performance in appearance and geometry reconstruction compared with state-of-the-art methods, while using significantly fewer parameters. Jiateng Liu, Hao Gao 0005, Jiucheng Xie, Chi-Man Pun, Jian Xiong 0005, Haolun Li 0001, Junxin Chen 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Refined Mamba-Based Lower Limbs Motor Estimator for Parkinson's Disease DiagnosticabstractBradykinesia, a key Parkinson's disease (PD) symptom, requires accurate lower limbs assessment, yet current clinical assessments are subjective and biased, while computer vision methods lack precision in skeleton extraction and PD-specific movement analysis. Furthermore, clothing-induced foot occlusion further aggravates keypoint localization errors. To address these gaps, we propose a vision-assisted diagnostic framework for PD lower limbs assessment. Our approach incorporates the CSDPose model, which employs CNN for local feature extraction, SSM (State Space Model) for global feature capture, and DCT (Discrete Cosine Transform) for frequency domain analysis, to enhance 2D pose estimation accuracy. These keypoints are used to compute objective PD motor indicators that we proposed, which are subsequently analyzed by a classification model to grade lower limbs dysfunction severity. Experiments demonstrate the algorithm achieves over$90\%$accuracy in classifying lower limbs motions. Validated clinical trials confirm that the automated severity ratings and motor indicators effectively support diagnostic decision-making. Xinyuan Dong, Hao Gao 0005, Yikang He, Yiqin Yao, Chi-Man Pun, Haolun Li 0001, Feng Xu 0005 |
BIBM | 7 |
| 2025 | EEMS: Edge-Prompt Enhanced Medical Image Segmentation Based on Learnable Gating MechanismabstractMedical image segmentation is vital for diagnosis, treatment planning, and disease monitoring but is challenged by complex factors like ambiguous edges and background noise. We introduce EEMS, a new model for segmentation, combining an Edge-Aware Enhancement Unit (EAEU) and a Multi-scale Prompt Generation Unit (MSPGU). EAEU enhances edge perception via multi-frequency feature extraction, accurately defining boundaries. MSPGU integrates high-level semantic and low-level spatial features using a prompt-guided approach, ensuring precise target localization. The Dual-Source Adaptive Gated Fusion Unit (DAGFU) merges edge features from EAEU with semantic features from MSPGU, enhancing segmentation accuracy and robustness. Tests on datasets like ISIC2018 confirm EEMS's superior performance and reliability as a clinical tool. Quanjun Li, Zimeng Li 0001, Hongbin Ye, Yupeng Liu 0003, Haolun Li 0001, Xuhang Chen 0002 |
BIBM | 7 |
| 2025 | A Noise-Resistant 3D Hand Motion Estimator Framework for Parkinson's Tremor AssessmentabstractParkinson's disease (PD) is a progressive neurodegenerative disorder, with tremor being one of its representative motor symptoms. Current clinical evaluations primarily rely on subjective scales such as the MDS-UPDRS, which often introduce significant inter-rater variability. Although vision-based evaluation offers objective motion analysis, existing pose tracking frameworks struggle with accurate tremor quantification due to inter-frame jitter and limited precision, failing to capture fine-grained spatiotemporal dynamics. To address this, we propose Motion-aware Hierarchical Grouping Mamba Network (MHG-Mamba), a non-contact video-based framework for automated evaluation of the 'finger-to-nose' task. Our approach first employs feature extraction to decouple high-degree-of-freedom finger joint movements from stable palm joint motions, enabling accurate finger pose estimation. Second, by incorporating with the hierarchical spatiotemporal scanning mechanism in Mamba's SSM, the model captures global motion features while preserving anatomical constraints, resulting in a temporally smooth and plausible skeletal sequence. Finally, based on the predicted skeletal sequences, we introduce several objective metrics to quantify motion features and apply a classifier for precise objective severity rating. Experimental results demonstrate that MHG-Mamba significantly improves the accuracy of 3D hand pose estimation and reduces noise in the motion sequences. The system achieved a classification accuracy of 93.2 % on the 'finger-to-nose' task. Moreover, clinicians using our system exhibited reduced variability in their assessments, highlighting its high clinical value. Yixing Ye, Hao Gao 0005, Yikang He, Haolun Li 0001, Chi-Man Pun, Feng Xu 0005 |
BIBM | 5 |
| 2025 | Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose EstimationabstractInconsistency of distributions in human actions and camera viewpoints can lead to significant deviations when the pre-trained 3D pose estimators are tested on cross-datasets. In practical applications, the estimators usually follow the standard full fine-tuning paradigm on the target dataset, which requires updating and saving a complete set of training parameters for different tasks, resulting in a large waste of resources and distorting pre-trained features. Taking inspiration from the widely used prompt learning in NLP, we explore the parameter-efficient fine-tuning solution of 3D pose estimators for the first time and propose the Adaptive Skeleton Prompt Tuning (ASP-Tuning) method, which freezes the backbone of the pre-trained model and generates a series of pose generic promptings as well as adaptive promptings specific to the input skeleton features to learn distribution transformation. Extensive experiments on multiple estimator backbones and datasets show that our method is superior to other fine-tuning methods and achieves state-of-the-art performance. Haolun Li 0001, Fuchen Zheng, Ye Liu 0005, Jian Xiong 0005, Haidong Hu, Hao Gao 0005 |
ICASSP | 1 |
| 2025 | RFEM: Remote Feature Enhancement Module for Target DetectionabstractThe research and development of dense crowd detection technology have always been one of the hot and challenging topics in the field of computer vision. DETR-like models have shown good performance in both training efficiency and inference capabilities. Nevertheless, as the optimization proceeds, these models can demonstrate sparse long-range feature correlations. This paper presents a specialized long-range feature enhancement module intended for optimizing DETR-like detection models. By utilizing an optimized PVM clustering algorithm, the robustness of the model is enhanced, and linear attention is incorporated into the aggregated tokens to reinforce long-range feature relationships. Besides, our method maintains the connections between occluded segmentation features during both training and inference phases. It also enhances the detection accuracy of small targets without increasing computational overhead. We conducted experiments on the COCO 2017 and the CrowdHuman datasets, and extensive experimental results demonstrate the effectiveness of our proposed method. Chuangye Wang, Jian Xiong 0005, Haolun Li 0001, Hao Gao 0005 |
ICASSP | 5 |
| 2025 | A Multi-Branch Network for Pose Trajectory Smoothing and RefinementabstractAlthough human pose estimation in video has achieved significant advancements, challenges such as occlusions often result in pronounced jitters and substantial errors. Addressing these challenges is critical for the further development of this field. In this study, we propose a multi-branch network designed for trajectory and pose refinement to mitigate these issues. The network comprises four branches, with three branches focus on capturing structural and motion features of human skeletal sequences, while one branch focuses on analyzing frequency features. More specifically, the first three branches model joint, bone vector, and motion data respectively to capture the structural and dynamic motion characteristics of human skeletal sequences. The frequency branch extracts frequency features to distinguish jitter from normal motion. Furthermore, the network incorporates a spatio-temporal residual block to capture both long-range temporal dependencies of individual joints and the spatial interrelationships among joints. Our method demonstrates competitive performance across three challenging datasets involving 2D, 3D, and SMPL pose representations. Haidong Hu, Chuangye Wang, Haolun Li 0001, Hao Gao 0005 |
ICME | 5 |
| 2025 | MotionRefineNet: Fine-Grained Pose Sequence Smoothing and RefinementabstractCapturing human motion with existing monocular estimators often results in large errors when dealing with rare poses, occlusions, truncations, and frame blurring, leading to jitter and long-term drift. Although previous methods have introduced post-processing networks for pose refinement, they struggle to balance global smoothing and fine-grained correction. In this work, we propose MotionRefineNet, which leverages the synergy and complementarity between long- and short-term features in the temporal domain and high- and low-frequency features in the frequency domain to address these challenges. The temporal branch is designed as a hierarchical motion structure to learn multi-time scale features, where long-term features learn motion smoothness, and short-term features capture local rapid changes. The frequency branch employs different frequency band learning strategies based on the degrees of freedom (DoF) of body parts. For body parts with low DoF, the focus is on low-frequency features that represent overall motion trends and regular actions. For body parts with high DoF, we design a filter to adaptively extract useful information from all frequency bands, including subtle motion changes in the high-frequency bands. Extensive experiments on multiple datasets and estimators demonstrate that MotionRefineNet outperforms existing methods in refining 2D, 3D, and SMPL poses, achieving superior pose smoothing and deviation correction. Our code is available at: https://github.com/Wheels319/MotionRefineNet. Haolun Li 0001, Weihuang Liu, Jiateng Liu, Zhenhua Tang 0001, Chi-Man Pun, Qiguang Miao, Feng Xu 0005, Hao Gao 0005 |
ACM Multimedia | 1 |
| 2025 | Hierarchical Local Temporal Network for 2D-to-3D Human Pose EstimationabstractRecent advancements in transformer-based methods have yielded substantial success in 2D-to-3D human pose estimation. Transformer-based estimators possess inherent advantages like the global receptive field. Nevertheless, existing transformer approaches ignore the differences among local contexts, resulting in insufficient learning of local information. To address this issue, we introduce nonuniform graph convolution to extract spatial local relationships in skeletons, remedying the limitations of traditional transformers in learning human body topology effectively. Additionally, our proposed hierarchical local temporal network (HLTN) models local temporal associations across three hierarchical levels: 1) joints; 2) body-parts; and 3) poses, effectively addressing the constraint of traditional transformers in learning localized human movements. We connect these two modules in parallel with the spatial and temporal transformer to obtain better features of skeleton sequences. Furthermore, we integrate nonuniform graph convolution with spatial Transformer methods to achieve interaction between local and global features at the attention level. Through these improved methods, our network not only effectively identifies global trends but also exhibits stronger sensitivity to local variations. Compared with the latest methods, our method achieves state-of-the-art performance on multiple datasets (Human3.6M and Mpi-Inf-3DHP). Jiucheng Xie, Haolun Li 0001, Hao Gao 0005 |
IEEE Internet Things J. | 4 |
| 2024 | An Automatic Assessment of Parkinson's Disease in Arising from Chair Task via Refined Diffusion-based Pose EstimatorabstractParkinson’s disease (PD) is a progressively common neurodegenerative disorder characterized by a decline in motor function. The diagnosis of PD typically relies on the Movement Disorder Society-Unified Parkinson’s Disease Rating Scale (MDS-UPDRS), which involves subjective scoring through observation of targeted movements. However, this objective method heavily depends on professional experience and has relatively high misdiagnosis rates. In this paper, we introduce a novel vision-based architecture for automated assessment of the ‘arising from chair’ task, which is one of the key MDS-UPDRS components. First, a diffusion-based 2D pose estimator is proposed to enhance keypoint accuracy by iteratively learning the distribution of ground-truth data and then denoising noisy poses. Second, a keypoint trajectory refinement network is introduced to eliminate the jitter error by considering motion information such as position, velocity, acceleration, and jerk. Finally, based on the predicted skeleton keypoint trajectories, we propose several objective indicators to assess the movement characteristics and perform the final rating using the classifier. The experiment substantiates the proposed algorithm, achieving a precision of 98.7% and an accuracy of 95.8% in classifying the ‘arising from chair’ task. Furthermore, the classification results and the proposed objective indicators have been validated as effective aids for neurologists to provide more precise diagnoses. Chi-Man Pun, Haolun Li 0001, Mingliang Zhai, Feng Xu 0005, Hao Gao 0005 |
BIBM | 3 |
| 2024 | SMAFormer: Synergistic Multi-Attention Transformer for Medical Image SegmentationabstractIn medical image segmentation, specialized computer vision techniques, notably transformers grounded in attention mechanisms and residual networks employing skip connections, have been instrumental in advancing performance. Nonetheless, previous models often falter when segmenting small, irregularly shaped tumors. To this end, we introduce SMAFormer, an efficient, Transformer-based architecture that fuses multiple attention mechanisms for enhanced segmentation of small tumors and organs. SMAFormer can capture both local and global features for medical image segmentation. The architecture comprises two pivotal components. First, a Synergistic Multi-Attention (SMA) Transformer block is proposed, which has the benefits of Pixel Attention, Channel Attention, and Spatial Attention for feature enrichment. Second, addressing the challenge of information loss incurred during attention mechanism transitions and feature fusion, we design a Feature Fusion Modulator. This module bolsters the integration between the channel and spatial attention by mitigating reshaping-induced information attrition. To evaluate our method, we conduct extensive experiments on various medical image segmentation tasks, including multi-organ, liver tumor, and bladder tumor segmentation, achieving state-of-the-art results. Code and models are available at: https://github.com/lzeeorno/SMAFormer. Fuchen Zheng, Xuhang Chen 0002, Weihuang Liu, Haolun Li 0001, Yingtie Lei, Chi-Man Pun, Shoujun Zhou |
BIBM | 4 |
| 2024 | Depth-Aware Test-Time Training for Zero-Shot Video Object SegmentationabstractZero-shot Video Object Segmentation (ZSVOS) aims at segmenting the primary moving object without any human annotations. Mainstream solutions mainly focus on learning a single model on large-scale video datasets, which struggle to generalize to unseen videos. In this work, we introduce a test-time training (TTT) strategy to address the problem. Our key insight is to enforce the model to predict consistent depth during the TTT process. In detail, we first train a single network to perform both segmentation and depth prediction tasks. This can be effectively learned with our specifically designed depth modulation layer. Then, for the TTT process, the model is updated by predicting consistent depth maps for the same frame under different data augmentations. In addition, we explore different TTT weight updating strategies. Our empirical results suggest that the momentum-based weight initialization and looping-based training scheme lead to more stable improvements. Experiments show that the proposed method achieves clear improvements on ZSVOS. Our proposed video TTT strategy provides significant superiority over state-of-the-art TTT methods. Our code is available at: https://nifangbaage.github.io/DATTT/. Weihuang Liu, Xi Shen 0001, Haolun Li 0001, Xiuli Bi, Bo Liu 0047, Chi-Man Pun, Xiaodong Cun |
CVPR | 3 |
| 2024 | DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion PredictionabstractPredicting human motion requires addressing dependencies and errors for pose forecasting from sequences. The transformer’s self-attention aids this, but its complexity poses computational challenges. We present an efficient DeformMLP network without self-attention, using fully connected layers. DeformMLP includes DeformFCs, DeformFCt, and DeformFCst layers for spatial temporal modeling and calibration. DeformFCs capture semantics, DeformFCt learns relationships by summarizing time tokens, and DeformFCst assigns significance to dimensions to reduce computation. Our method balances efficiency and accuracy through decomposition and weight allocation. Evaluation on Human3.6M, 3DPW, CMU-MoCap datasets shows state-of-the-art prediction performance by benchmarks. The code is publicly available at https://github.com/HHT-98/DeformMLP. Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Hao Gao 0005 |
ICASSP | 3 |
| 2024 | Local Optimization Networks for Multi-View Multi-Person Human Posture EstimationabstractWith the growing applicability of multi-view multi-person 3D human pose estimation across diverse scenarios, the impact of external environmental factors and occlusion on accuracy has garnered substantial attention. In this research, we introduce a novel approach to multi-view multi-person 3D human pose estimation, leveraging a localized optimization strategy. Specifically, our method enhances the interplay of feature information from different channels and fine-tunes the optimal feature weights to capture intricate dependencies among joints. This refinement leads to improved accuracy in handling external environmental factors. Experimental evaluations were conducted on two prominent benchmark datasets, namely Campus and Shelf. The proposed method achieved a remarkable performance, with a Percentage of Correct Parts (PCP) score of 97.4% and 98.2% for the Campus and Shelf datasets, respectively. Jucheng Song, Chi-Man Pun, Haolun Li 0001, Rushi Lan, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 3 |
| 2024 | Hierarchical Local Temporal Feature Enhancing for Transformer-Based 3D Human Pose EstimationabstractRecent advancements in transformer-based methods have yielded substantial success in 2D-to-3D human pose estimation. Transformer-based estimators have their inherent advantages like global receptive field. Nevertheless, existing transformer approaches ignore the differences among local contexts, resulting in insufficient learning of local information. To address this issue, we introduce non-uniform graph convolution to extract spatial local relationships in skeletons, remedying the limitations of traditional transformers in learning human body topology effectively. Additionally, our proposed Hierarchical Local Temporal Network (HLTN) models local temporal associations across three hierarchical levels: joints, body-parts and poses, effectively addressing the constraint of traditional transformers in learning localized human movements. We connect these two modules in parallel with the spatial and temporal transformer to obtain better features of skeleton sequences. Compared with the latest methods, our method achieves state-of-the-art performance on multiple datasets. Chi-Man Pun, Haolun Li 0001, Hao Gao 0005 |
ICME | 3 |
| 2024 | Adaptive Spatial-Temporal Graph-Mixer for Human Motion PredictionabstractThe Graph Convolutional Network (GCN) has recently achieved promising performance in human motion prediction by modeling the nodes and edges of the human skeleton. However, most previous methods still suffer from two unaddressed drawbacks. First, in the inference stage, their graph topologies are static and fixed, resulting in dependencies between nodes that cannot be dynamically adjusted for different actions. Second, the implicit relationships between pose sequences are ignored, which makes the prior advantages of the graph structure invalid in temporal feature fusion. To address these limitations, we propose an adaptive spatial-temporal graph-mixer (GraphMixer) for human motion prediction, which consists of a series of fully separated spatial-temporal graph convolution structures. In spatial GCN, we construct an additional adaptive skeleton graph to capture the node features of action-specific poses. In temporal GCN, we introduce a variety of graph topologies to enhance feature fusion between pose sequences. Comparing state-of-the-art algorithms on the Human3.6 M and the 3DPW datasets and ablation studies shows that our GraphMixer and the proposed multiple graph topologies are effective and critical. The code is publicly available athttps://github.com/young0304/Adaptive-Spatial-Temporal-Graph-Mixer. Haolun Li 0001, Chi-Man Pun, Chun Du, Hao Gao 0005 |
IEEE Signal Process. Lett. | 2 |
| 2023 | CEE-Net: Complementary End-to-End Network for 3D Human Pose Generation and EstimationabstractThe limited number of actors and actions in existing datasets make 3D pose estimators tend to overfit, which can be seen from the performance degradation of the algorithm on cross-datasets, especially for rare and complex poses. Although previous data augmentation works have increased the diversity of the training set, the changes in camera viewpoint and position play a dominant role in improving the accuracy of the estimator, while the generated 3D poses are limited and still heavily rely on the source dataset. In addition, these works do not consider the adaptability of the pose estimator to generated data, and complex poses will cause training collapse. In this paper, we propose the CEE-Net, a Complementary End-to-End Network for 3D human pose generation and estimation. The generator extremely expands the distribution of each joint-angle in the existing dataset and limits them to a reasonable range. By learning the correlations within and between the torso and limbs, the estimator can combine different body-parts more effectively and weaken the influence of specific joint-angle changes on the global pose, improving the generalization ability. Extensive ablation studies show that our pose generator greatly strengthens the joint-angle distribution, and our pose estimator can utilize these poses positively. Compared with the state-of-the-art methods, our method can achieve much better performance on various cross-datasets, rare and complex poses. Haolun Li 0001, Chi-Man Pun |
AAAI | 1 |
| 2023 | An Optimized-Skeleton-Based Parkinsonian Gait Auxiliary Diagnosis Method with Both Monitoring Indicators and Assisted RatingsabstractAbnormal gait is one of the indispensable diagnostic sources of Parkinson’s disease (PD) diagnosis, typically presenting as small shuffling steps and gait bradykinesia. However, its diagnostic accuracy is lower due to the subjective judgments of doctors. To assist in improving the accuracy and reducing the subjectivity of the doctors, we propose an optimized-skeleton-based Parkinsonian gait auxiliary diagnosis method with both monitoring indicators and assisted ratings. By inputting a patient gait video captured from the side, our PD symptom-applicable pose trajectory model will extract a more precise and stable 2D skeleton sequence of patients. Next, the sequence will be used to calculate our proposed five monitoring indicators: gait frequency, ankle speed, whole speed, ankle angle speed, ankle acceleration, and previous work indicators: arm swing angle, leg angle, two feet x-axis distance to record the patient’s gait details at every moment. The extracted gait frequency will then be input into a random forest model to obtain the gait rating. Lastly, doctors can make more accurate judgments by referring to our objective monitoring indicators and assisted ratings. Experimental results show that our monitored indicators improve the doctors’ diagnosis accuracy by 16%, the skeleton speed and acceleration error of our optimized-skeleton extraction method achieve 4.18 cm/s and 5.71 cm/s2, and our random forest model has reached a classification accuracy of 95.8%. Gaoqi Li, Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Feng Xu 0005, Hao Gao 0005 |
BIBM | 3 |
| 2022 | Monocular Robust 3D Human Localization by Global and Body-Parts Depth AwarenessabstractLearning the human depth localization in camera coordinate space plays a crucial role in understanding the behavior and activities of multi-person in 3D scenes. However, existing monocular-based methods rarely combine the global image features and the human body-parts features effectively, resulting in a large gap from the actual location in some cases, e.g., the special body-sized persons and mutual occlusion between humans in the image. This paper presents a novel Robust 3D Human Localization (R3HL) network consisting of two stages: global depth awareness and body-parts depth awareness, to significantly improve the robustness and accuracy of the 3D location. In the first stage, the front-back and far-near relationship estimation module based on multi-person are proposed to make the network extract depth features from the global perspective. In the second stage, the network focuses on the target human. We propose a Pose-guided Multi-person Repulsion (PMR) module to enhance the target human’s features and reduce the interference features produced by the background and other people. In addition, an Adaptive Body-parts Attention (ABA) module is designed to assign different feature weights to each joint. Finally, the human’s absolute depth is obtained through global pooling and fully connected layers. The experimental results show that the attention from the whole image to a single person helps find the absolute location of different body-sized and poses people from diverse scenes. Our method can achieve better performance than other state-of-the-art methods on both indoor and outdoor 3D multi-person datasets. Haolun Li 0001, Chi-Man Pun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | A Hybrid Feature Selection Algorithm Based on a Discrete Artificial Bee Colony for Parkinson's DiagnosisabstractParkinson's disease is a neurodegenerative disease that affects millions of people around the world and cannot be cured fundamentally. Automatic identification of early Parkinson's disease on feature data sets is one of the most challenging medical tasks today. Many features in these datasets are useless or suffering from problems like noise, which affect the learning process and increase the computational burden. To ensure the optimal classification performance, this article proposes a hybrid feature selection algorithm based on an improved discrete artificial bee colony algorithm to improve the efficiency of feature selection. The algorithm combines the advantages of filters and wrappers to eliminate most of the uncorrelated or noisy features and determine the optimal subset of features. In the filter, three different variable ranking methods are employed to pre-rank the candidate features, then the population of artificial bee colony is initialized based on the significance degree of the re-rank features. In the wrapper part, the artificial bee colony algorithm evaluates individuals (feature subsets) based on the classification accuracy of the classifier to achieve the optimal feature subset. In addition, for the first time, we introduce a strategy that can automatically select the best classifier in the search framework more quickly. By comparing with several publicly available datasets, the proposed method achieves better performance than other state-of-the-art algorithms and can extract fewer effective features. Haolun Li 0001, Chi-Man Pun, Feng Xu 0005, Longsheng Pan, Rui Zong, Hao Gao 0005, Huimin Lu 0001 |
ACM Trans. Internet Techn. | 1 |
| 2020 | Training Feed-Forward Artificial Neural Networks with a modified artificial bee colony algorithm
Feiyi Xu, Chi-Man Pun, Haolun Li 0001, Yushu Zhang 0001, Yurong Song, Hao Gao 0005 |
Neurocomputing | 3 |
| 2020 | High-quality-guided artificial bee colony algorithm for designing loudspeaker
Hao Gao 0005, Haolun Li 0001, Ye Liu 0005, Huimin Lu 0001, Hyoungseop Kim, Chi-Man Pun |
Neural Comput. Appl. | 2 |