EDBT 2026 Demo / reviewers in the wild / expert
Hao Gao 0005
dblp:69/8912-5
· DBLP profile ↗
73ranked-venue papers
6as first author
55since 2021 · last 2026
0000-0003-0148-3713ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 36 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 8 since 2021Computer networks · 6 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point Cloud Quality Assessment via Multi-View Structure-Aware Feature FusionabstractPoint cloud quality assessment (PCQA) is essential for reliable 3D visual applications. While point-based methods face challenges in characterizing distortions due to point cloud disorder, projection-based approaches offer better efficiency but suffer from geometric distortion insensitivity and texture representation blind spots. This study proposes SAF-Net, a multi-view structure-aware feature fusion network for PCQA. We first identify two key limitations in projection-based methods: insufficient geometric distortion perception and representation blind spots (RBS) in texture images. To address these issues, SAF-Net innovatively integrates object mask maps and local binary pattern (LBP) maps. The mask maps enhance geometric distortion perception by extracting edge sharpness and curvature variations, while LBP maps capture essential structural information to overcome RBS and align with human visual system (HVS) sensitivity. SAF-Net employs a hybrid CNN-ViT architecture to balance local feature extraction and global context modeling, along with a progressive fusion strategy to optimize cross-modal feature interaction. Extensive experiments demonstrate the superior performance of SAF-Net on multiple benchmarks, establishing new state-of-the-art results in PCQA. Jian Xiong 0005, Lingxia Jiang, Xianzhong Long, Miaohui Wang, Hao Gao 0005 |
AAAI | 5 |
| 2026 | GPDPose: Self-supervised transformer with geometry, pose, and depth consistency for multi-view 3D human pose estimation
Jucheng Song, Jie Zhang 0090, Xu Yang 0010, Yapeng Wang 0001, Hao Gao 0005, Haolun Li 0001, Sio Kei Im |
Expert Syst. Appl. | 5 |
| 2026 | The Role of Digital Twin in Advancing Industrial Internet of Things: Insights, Applications, and Future DirectionsabstractTo Date, the application of digital twin (DT) in the industrial internet of things (IIoT) has been continuously promoted and deepened, and has become the focus of the industry. IIoT serves as the foundational infrastructure that enables pervasive connectivity, real-time data acquisition, and intelligent control within industrial environments. DTs provide enterprises with an empathetic, virtual environment that enables them to manage and operate their production facilities in a more efficient and intelligent manner. However, there is not a special summary and analysis of the combinability and combination mode of the two. Therefore, this paper firstly sorted out the professional definitions, characteristics and frameworks of IIoT and DT, and deeply analyzed the semantic context of data flow. Secondly, this paper discusses the combinability and combination mode of IIoT and DT, and summarizes the enabling technologies and tools at each layers. Finally, the applications status of DT empowered IIoT in different fields was summarized, and the challenges of the combined application of the two were analyzed. Junxin Chen 0001, Hao Gao 0005, Qiang He 0002, Jun Mou, Wei Wang 0077 |
IEEE Internet Things J. | 3 |
| 2026 | Cellular Aggregation Graph Convolutional Network for Point Cloud Quality AssessmentabstractPoint cloud quality assessment (PCQA) is a challenging task due to the inherently disordered nature of points. Existing point-based methods, such as sparse convolution and PointNet, are limited by local spatial modeling and structural feature extraction. Although 3D graph convolutional networks (GCNs) offer advantages in capturing local structural features through explicit geometric modeling and deformable kernels, their scalability is hindered by the high memory consumption associated with storing neighborhood matrices, particularly for large-scale point clouds. In this paper, to better extract hierarchical structural information and maintain efficiency in computational memory, we propose a novel point-based no-reference PCQA method, namely cellular aggregation network (CANet). The method effectively and efficiently extracts the quality-aware features of large patches in a divide-and-conquer manner. Specifically, a cellular sampling (CS) module is introduced to divide large patches into smaller cells, effectively avoiding the problem of memory explosion. A cellular aggregation (CA) module is proposed to extract intra-cell features and fuse inter-cell features. Moreover, a global aggregation (GA) module is presented to extract global sketch information. Finally, a long-term fusion (LTF) module is introduced to capture long-term dependencies between the features of the CA and GA modules. Experimental results on benchmark datasets demonstrate that the proposed model achieves state-of-the-art performance. Jian Xiong 0005, Lingxia Jiang, Qiang Hu 0003, Jiucheng Xie, Hao Gao 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Effective Gaussian Management for High-Fidelity Scene ReconstructionabstractThis paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent Gaussian Splatting (GS) pipelines that treat all primitives uniformly during optimization, our framework explicitly manages the attribute activation, representation and pruning of Gaussian. Specifically, our framework first introduces GauSep, a novel densification strategy that selectively activates Gaussian color or normal attributes to alleviate destructive gradient conflicts arising from dual supervision. We further propose GauRep, an adaptive Gaussian representation that dynamically adjusts spherical harmonics (SHs) orders and performs task-decoupled pruning to reduce redundancy at both the individual and global levels. To provide reliable geometric supervision for above mangement process, we additionally introduce CoRe, an regularized surface reconstruction module that distills robust normal fields from an SDF branch to the Gaussian representation through a confidence mechanism. Notably, the proposed Gaussian management is compatible with various reconstruction architectures and can be seamlessly integrated to improve performance while reducing size of the model. Extensive experiments demonstrate that our approach achieves superior or comparable performance in appearance and geometry reconstruction compared with state-of-the-art methods, while using significantly fewer parameters. Jiateng Liu, Hao Gao 0005, Jiucheng Xie, Chi-Man Pun, Jian Xiong 0005, Haolun Li 0001, Junxin Chen 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Refined Mamba-Based Lower Limbs Motor Estimator for Parkinson's Disease DiagnosticabstractBradykinesia, a key Parkinson's disease (PD) symptom, requires accurate lower limbs assessment, yet current clinical assessments are subjective and biased, while computer vision methods lack precision in skeleton extraction and PD-specific movement analysis. Furthermore, clothing-induced foot occlusion further aggravates keypoint localization errors. To address these gaps, we propose a vision-assisted diagnostic framework for PD lower limbs assessment. Our approach incorporates the CSDPose model, which employs CNN for local feature extraction, SSM (State Space Model) for global feature capture, and DCT (Discrete Cosine Transform) for frequency domain analysis, to enhance 2D pose estimation accuracy. These keypoints are used to compute objective PD motor indicators that we proposed, which are subsequently analyzed by a classification model to grade lower limbs dysfunction severity. Experiments demonstrate the algorithm achieves over$90\%$accuracy in classifying lower limbs motions. Validated clinical trials confirm that the automated severity ratings and motor indicators effectively support diagnostic decision-making. Xinyuan Dong, Hao Gao 0005, Yikang He, Yiqin Yao, Chi-Man Pun, Haolun Li 0001, Feng Xu 0005 |
BIBM | 2 |
| 2025 | A Noise-Resistant 3D Hand Motion Estimator Framework for Parkinson's Tremor AssessmentabstractParkinson's disease (PD) is a progressive neurodegenerative disorder, with tremor being one of its representative motor symptoms. Current clinical evaluations primarily rely on subjective scales such as the MDS-UPDRS, which often introduce significant inter-rater variability. Although vision-based evaluation offers objective motion analysis, existing pose tracking frameworks struggle with accurate tremor quantification due to inter-frame jitter and limited precision, failing to capture fine-grained spatiotemporal dynamics. To address this, we propose Motion-aware Hierarchical Grouping Mamba Network (MHG-Mamba), a non-contact video-based framework for automated evaluation of the 'finger-to-nose' task. Our approach first employs feature extraction to decouple high-degree-of-freedom finger joint movements from stable palm joint motions, enabling accurate finger pose estimation. Second, by incorporating with the hierarchical spatiotemporal scanning mechanism in Mamba's SSM, the model captures global motion features while preserving anatomical constraints, resulting in a temporally smooth and plausible skeletal sequence. Finally, based on the predicted skeletal sequences, we introduce several objective metrics to quantify motion features and apply a classifier for precise objective severity rating. Experimental results demonstrate that MHG-Mamba significantly improves the accuracy of 3D hand pose estimation and reduces noise in the motion sequences. The system achieved a classification accuracy of 93.2 % on the 'finger-to-nose' task. Moreover, clinicians using our system exhibited reduced variability in their assessments, highlighting its high clinical value. Yixing Ye, Hao Gao 0005, Yikang He, Haolun Li 0001, Chi-Man Pun, Feng Xu 0005 |
BIBM | 2 |
| 2025 | Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose EstimationabstractInconsistency of distributions in human actions and camera viewpoints can lead to significant deviations when the pre-trained 3D pose estimators are tested on cross-datasets. In practical applications, the estimators usually follow the standard full fine-tuning paradigm on the target dataset, which requires updating and saving a complete set of training parameters for different tasks, resulting in a large waste of resources and distorting pre-trained features. Taking inspiration from the widely used prompt learning in NLP, we explore the parameter-efficient fine-tuning solution of 3D pose estimators for the first time and propose the Adaptive Skeleton Prompt Tuning (ASP-Tuning) method, which freezes the backbone of the pre-trained model and generates a series of pose generic promptings as well as adaptive promptings specific to the input skeleton features to learn distribution transformation. Extensive experiments on multiple estimator backbones and datasets show that our method is superior to other fine-tuning methods and achieves state-of-the-art performance. Haolun Li 0001, Fuchen Zheng, Ye Liu 0005, Jian Xiong 0005, Haidong Hu, Hao Gao 0005 |
ICASSP | 7 |
| 2025 | Enhancing Autonomous Vehicle Planning With a Robust Fault-Tolerant Mechanism for Action-Induced Agent DetectionabstractIn autonomous driving, accurately identifying traffic participants that may influence vehicle behavior is crucial for effective system planning. To address this challenge, we propose a fault-tolerant mechanism for detecting action-induced objects, which significantly improves decision-making performance and system explainability. Since these objects are often linked to the vehicle’s driving intentions, we introduce a top-down attention network that adjusts attention weights for traffic participants based on navigational information. Additionally, we define potentially hazardous objects in the driving environment and employ supervised training with a classification head to detect them. To further enhance detection accuracy, we integrate a fault-tolerant process that merges attention maps with classification results, effectively reducing false positives and false negatives in identifying action-induced objects. Extensive testing validates the robustness and effectiveness of our approach, demonstrating its ability to improve both planning and interpretability in autonomous vehicles. Zheng Fu, Hezhe Lin, Kangan Qian, Tuopu Wen, Hao Gao 0005, Diange Yang |
ICASSP | 5 |
| 2025 | RFEM: Remote Feature Enhancement Module for Target DetectionabstractThe research and development of dense crowd detection technology have always been one of the hot and challenging topics in the field of computer vision. DETR-like models have shown good performance in both training efficiency and inference capabilities. Nevertheless, as the optimization proceeds, these models can demonstrate sparse long-range feature correlations. This paper presents a specialized long-range feature enhancement module intended for optimizing DETR-like detection models. By utilizing an optimized PVM clustering algorithm, the robustness of the model is enhanced, and linear attention is incorporated into the aggregated tokens to reinforce long-range feature relationships. Besides, our method maintains the connections between occluded segmentation features during both training and inference phases. It also enhances the detection accuracy of small targets without increasing computational overhead. We conducted experiments on the COCO 2017 and the CrowdHuman datasets, and extensive experimental results demonstrate the effectiveness of our proposed method. Chuangye Wang, Jian Xiong 0005, Haolun Li 0001, Hao Gao 0005 |
ICASSP | 6 |
| 2025 | Diffusion Models are Good Unsupervised Class-agnostic Shape Part SegmentatorsabstractShape part segmentation is a critical task in computer graphics and robotics. However, traditional supervised methods rely heavily on large amounts of labeled data, which poses significant challenges in many real-world scenarios where such data is often scarce or difficult to obtain. To address this issue, we propose an unsupervised, class-agnostic part segmentation method called Point Diffusion Segmentation (PDS). Our research demonstrates that unconditional point cloud diffusion models can capture abstract object concepts within their sub-attention layers. By extracting preliminary point cloud features from these attention maps, PDS generates efficacious segmentation results. This method fully leverages unlabeled data and proves to be highly applicable in various downstream tasks, including zero-shot part segmentation. Without resorting to any labeled data, PDS improves the zero-shot part segmentation performance of PointClipV2 by 3.1% on the ShapeNet Part dataset, setting a new state-of-the-art baseline and demonstrating significant potential of PDS. Zhongbin Jiang, Tianhao Shi, Hao Gao 0005, Jun Liu 0036, Ye Liu 0005 |
ICASSP | 3 |
| 2025 | A Multi-Branch Network for Pose Trajectory Smoothing and RefinementabstractAlthough human pose estimation in video has achieved significant advancements, challenges such as occlusions often result in pronounced jitters and substantial errors. Addressing these challenges is critical for the further development of this field. In this study, we propose a multi-branch network designed for trajectory and pose refinement to mitigate these issues. The network comprises four branches, with three branches focus on capturing structural and motion features of human skeletal sequences, while one branch focuses on analyzing frequency features. More specifically, the first three branches model joint, bone vector, and motion data respectively to capture the structural and dynamic motion characteristics of human skeletal sequences. The frequency branch extracts frequency features to distinguish jitter from normal motion. Furthermore, the network incorporates a spatio-temporal residual block to capture both long-range temporal dependencies of individual joints and the spatial interrelationships among joints. Our method demonstrates competitive performance across three challenging datasets involving 2D, 3D, and SMPL pose representations. Haidong Hu, Chuangye Wang, Haolun Li 0001, Hao Gao 0005 |
ICME | 6 |
| 2025 | MMPX: Multi-modal Mamba Prompter to Large Vision Foundation Model for RGB-X Semantic SegmentationabstractMulti-modal semantic segmentation leverages multiple types of input data to perform pixel-level classification of images, enhancing the accuracy and robustness of segmentation tasks. Mainstream methods use small-scale models which have limited generalization ability. Training large-scale multi-modal models requires massive multi-modal data, which is difficult to obtain. Thus, fine-tuning large-scale vision foundation models (LVFMs) trained with abundant RGB data for multi-modal segmentation is a more practical solution. Existing methods only tap into the potential of LVFMs in the RGB modality which ignores their potential in non-RGB modalities. In this paper, we propose an innovative universal prompting framework, MMPX. Specifically, the effective multi-modal Mamba fuser (MMF) explores the potential of LVFMs in the integrated representation of RGB+X. On the other hand, we introduce multi-modal Mamba prompters (MMPs) to fine-tune large-scale foundation models. This prompter takes the integrated RGB+X representation as input and dynamically adjusts the model parameters using Mamba, eliminating redundant information while retaining key features, thus achieving efficient prompt generation. The proposed method achieves SOTA performance on five multi-modal benchmarks, including RGB+Depth, RGB+Thermal, RGB+Event, which fully validate the effectiveness and generalization ability of the approach. The code and results are available at: https://github.com/CauchyCat/MMPX. Ye Liu 0005, Hao Gao 0005, Jun Liu 0036 |
ICME | 3 |
| 2025 | CANet: Cellular Aggregation Network for Point Cloud Quality AssessmentabstractThe concept of visual masking reveals that human visual perception is influenced by content and distortion information. Existing projection-based methods lose depth information and intrinsic topological structures. Due to the limitations of computational memory, the existing point-based methods tend to deal with small patches with little content information. In this paper, we propose a novel point-based no-reference quality assessment method, namely cellular aggregation network (CANet). The method effectively extracts the quality-aware features of large patches in a divide-and-conquer manner. Specifically, the cellular sampling module is used to divide large patches into smaller cells, which effectively avoids the memory explosion problem. The cellular aggregation module is proposed to obtain more content information from small cells. A global aggregation module is proposed to extract global sketch information. Furthermore, a long-term fusion module is introduced to capture long-term dependencies, which can better receive content-aware semantic features. Experimental results on benchmark databases demonstrate that CANet achieves competitive performances. Lingxia Jiang, Jian Xiong 0005, Jiucheng Xie, Hao Gao 0005 |
ISCAS | 5 |
| 2025 | MotionRefineNet: Fine-Grained Pose Sequence Smoothing and RefinementabstractCapturing human motion with existing monocular estimators often results in large errors when dealing with rare poses, occlusions, truncations, and frame blurring, leading to jitter and long-term drift. Although previous methods have introduced post-processing networks for pose refinement, they struggle to balance global smoothing and fine-grained correction. In this work, we propose MotionRefineNet, which leverages the synergy and complementarity between long- and short-term features in the temporal domain and high- and low-frequency features in the frequency domain to address these challenges. The temporal branch is designed as a hierarchical motion structure to learn multi-time scale features, where long-term features learn motion smoothness, and short-term features capture local rapid changes. The frequency branch employs different frequency band learning strategies based on the degrees of freedom (DoF) of body parts. For body parts with low DoF, the focus is on low-frequency features that represent overall motion trends and regular actions. For body parts with high DoF, we design a filter to adaptively extract useful information from all frequency bands, including subtle motion changes in the high-frequency bands. Extensive experiments on multiple datasets and estimators demonstrate that MotionRefineNet outperforms existing methods in refining 2D, 3D, and SMPL poses, achieving superior pose smoothing and deviation correction. Our code is available at: https://github.com/Wheels319/MotionRefineNet. Haolun Li 0001, Weihuang Liu, Jiateng Liu, Zhenhua Tang 0001, Chi-Man Pun, Qiguang Miao, Feng Xu 0005, Hao Gao 0005 |
ACM Multimedia | 8 |
| 2025 | FGRFlow: Learning Fine-Grained Rigidity Scene Flow from 4D Radar Point CloudabstractScene flow estimation using 4D millimeter-wave radar has emerged as a prominent research focus for 3D dynamic perception. However, compared to LiDAR point clouds, the drastic sparsity of radar point clouds poses challenges in enforcing local rigidity constraints, which are crucial for accurate 3D motion estimation. To address this issue, we propose a novel Gaussian-based pseudo-point generation method that fully leverages two distinct yet complementary data modalities, 3D coordinates and Doppler velocity, to support multi-body rigidity assumptions, effectively capturing fine-grained and structured motion patterns from highly sparse radar point clouds. Furthermore, a velocity calibration mechanism is designed to improve the reliability of fine-grained rigid motion velocity estimation. In addition, a progressive fusion strategy is introduced to systematically integrate fine-grained rigid motion priors at multiple levels, enhancing the robustness of matching costs and motion features while effectively compensating for coarse flows. Experimental results on real-world radar scans from the View-of-Delft (VoD) dataset demonstrate the promising performance of our FGRFlow compared to other leading 4D radar-based approaches, validating the advantages of our design choices. Mingliang Zhai, Haidong Hu, Chi-Man Pun, Hao Gao 0005 |
ACM Multimedia | 5 |
| 2025 | Learning optical flow from spiking camera with direction disassemblyabstractAbstract Conventional optical flow estimation methods typically recover two‐dimensional motion from RGB image sequences. Recently, due to the rise and widespread use of spike cameras, learning optical flow from spiking cameras has become a hot topic in the field of two‐dimensional motion estimation. Although existing methods have been designed to learn optical flow by designing feature processing methods for spike streams, there is still insufficient consideration for flow field post‐processing, resulting in limited accuracy of optical flow estimation. To address this problem, an optical flow estimation method based on directional disassembly is proposed. Specifically, the estimated flow fields along the horizontal and vertical directions are disassembled and the motion vectors along the two directions are denoised separately to reduce the burden of post‐processing for complex two‐dimensional motion information. In addition, contextual information is introduced in the post‐processing so that the scene information can effectively contribute to the results of the flow post‐processing. Experimental results show that this proposed method is capable of achieving comparable performance on spike‐based public datasets. Mingliang Zhai, Xuezhi Xiang, Kang Ni, Hao Gao 0005 |
IET Image Process. | 5 |
| 2025 | Pedestrian Trajectory Prediction for Autonomous Vehicles With Multiple InteractionsabstractPedestrian trajectory prediction is significant for autonomous vehicles, but the difficulty of pedestrian trajectory prediction lies in the accurate modeling of pedestrian multiple interactions. In this paper, we attempt to explore the essential features of pedestrian interaction and propose a pedestrian trajectory prediction method based on multiple interactions. Firstly, considering that the interaction between self-driving cars and pedestrians resembles a dynamic game process involving sequential adaptation, we map them to the same feature space and design a temporal cross-attention mechanism to model the interaction between pedestrians and vehicles. Meanwhile, pedestrian-scene interaction is affected by the global environment as well as the local environment. To capture the global information while preserving the spatial location of pedestrians in the scene, we design a pedestrian-scene heatmap fusion (PSHF) framework to model the pedestrian-scene interaction features. We validate the effectiveness of our algorithm on the publicly available JAAD and PIE datasets, achieving better performance than existing representative methods in both single-trajectory and multi-trajectory prediction tasks. We conducted a thorough ablation study, cross-dataset validation, and qualitative visualization experiments, demonstrating the effectiveness and robustness of our method. Zheng Fu, Mengmeng Yang 0001, Kun Jiang 0002, Jin Huang 0002, Hao Gao 0005, Diange Yang |
IEEE Internet Things J. | 7 |
| 2025 | Hierarchical Local Temporal Network for 2D-to-3D Human Pose EstimationabstractRecent advancements in transformer-based methods have yielded substantial success in 2D-to-3D human pose estimation. Transformer-based estimators possess inherent advantages like the global receptive field. Nevertheless, existing transformer approaches ignore the differences among local contexts, resulting in insufficient learning of local information. To address this issue, we introduce nonuniform graph convolution to extract spatial local relationships in skeletons, remedying the limitations of traditional transformers in learning human body topology effectively. Additionally, our proposed hierarchical local temporal network (HLTN) models local temporal associations across three hierarchical levels: 1) joints; 2) body-parts; and 3) poses, effectively addressing the constraint of traditional transformers in learning localized human movements. We connect these two modules in parallel with the spatial and temporal transformer to obtain better features of skeleton sequences. Furthermore, we integrate nonuniform graph convolution with spatial Transformer methods to achieve interaction between local and global features at the attention level. Through these improved methods, our network not only effectively identifies global trends but also exhibits stronger sensitivity to local variations. Compared with the latest methods, our method achieves state-of-the-art performance on multiple datasets (Human3.6M and Mpi-Inf-3DHP). Jiucheng Xie, Haolun Li 0001, Hao Gao 0005 |
IEEE Internet Things J. | 5 |
| 2025 | Lifespan age synthesis on human faces with decorrelation constraints and geometry guidance
Jiucheng Xie, Lingqing Zhang, Hao Gao 0005, Chi-Man Pun |
Pattern Recognit. Lett. | 3 |
| 2025 | Multi-Task Learning Model for V-PCC Geometry Compression Artifact RemovalabstractIn video-based point cloud compression (V-PCC), point clouds are projected as videos using a patch projection method and then compressed using video coding techniques. However, the lossy video compression and the down-sampling of occupancy maps (OMs) can lead to geometry compression artifacts, i.e., depth errors and OM errors, respectively. These errors can significantly affect the reconstruction quality of the point clouds. Existing methods can only eliminate one type of error and therefore have limited quality improvement. In this paper, to improve the quality maximally, a multi-task learning-based geometry compression artifact removal method is proposed to reduce both types of errors simultaneously. Considering the differences between the two tasks, the proposed method deals with the challenges of shared feature extraction and heterogeneous objective optimization. First, we propose a context-aware multi-task learning (CAML) model. The proposed CAML model can extract shared features that are context-aware and satisfy both tasks. Second, an improved optimization scheme is presented to train the proposed model. The improved optimization can fix the gradient imbalance of model updating. Cross-validation experiments show that the proposed method saves an average of over 45% Bjϕntegaard Delta bitrate in terms of the D2 metric. Jian Xiong 0005, Jiucheng Xie, Hui Yuan 0001, Hao Gao 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Coupled Noise Suppression and Feature Enhancement Network for Skeleton-Based Action RecognitionabstractIn recent years, remarkable progress has been made in skeleton-based action recognition. However, there is a significant amount of noise in skeleton data, which is simply overlooked by most existing methods. Some methods have designed specialized mechanisms to handle noise, but these mechanisms are either based on prior knowledge or require additional supervision information. To overcome these problems, we propose in this article a fully implicit solution, which embeds a soft-thresholding-based denoising module into existing networks, which can automatically learn to remove noise without any prior knowledge or additional supervision information. In addition, by relaxing the nonnegative constraint, the module gains the ability to adaptively enhance key features. Based on this, we further propose a two-staged method for coupled noise suppression and feature enhancement. The proposed method achieves state-of-the-art performance on public datasets. Moreover, on noise polluted datasets, the proposed method demonstrates significant performance advantages over existing methods. Ye Liu 0005, Tianyong Wu, Tianhao Shi, Miaohui Wang, Hao Gao 0005, Jun Liu 0036 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Top-Down Attention-Based Mechanisms for Interpretable Autonomous DrivingabstractDespite the remarkable advancements in autonomous driving, the challenge persists in achieving interpretable action decision-making, primarily owing to the intricate and ambiguous relationship between detected agents and driving intention. In this study, we introduce an interpretable action prediction model, denoted as the Prediction-Driven Attention Network (PDANet), designed to undertake action decisions and provide corresponding interpretations cohesively. The PDANet is inspired by the perceptual mechanisms inherent in human drivers, who allocate attention according to their driving intentions. Specifically, we elaborate a prediction module to generate vehicle prospective trajectories to characterize driving intentions. Subsequently, the features of this predicted trajectory are utilized to modulate the attention distribution among agents through the top-down attention module, yielding an attention map. Finally, two distinct task tokens are applied to aggregate agent features and generate the final output according to the derived attention map. Extensive experiments conducted on the publicly available BDD-OIA and nu-AR datasets demonstrate that our proposed method outperforms all prior works in terms of both action prediction and behavior interpretation tasks. Remarkably, our method attains a noteworthy enhancement in the behavior interpretation task, surpassing the previous state-of-the-art by a substantial margin of +10.8% in terms of F1-score on the nu-AR dataset. We also validate our algorithm on Carla Town05 long in a closed-loop decision-making scenario, highlighting the generality and robustness of our approach. Furthermore, qualitative results show that the agents selected by our model are more closely aligned with human cognitive processes. Zheng Fu, Kun Jiang 0002, Yunlong Wang 0009, Tuopu Wen, Hao Gao 0005, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | GaussianHead: High-Fidelity Head Avatars With Learnable Gaussian DerivationabstractCreating lifelike 3D head avatars and generating compelling animations for diverse subjects remain challenging in computer vision. This paper presents GaussianHead, which models the active head based on anisotropic 3D Gaussians. Our method integrates a motion deformation field and a single-resolution tri-plane to capture the head's intricate dynamics and detailed texture. Notably, we introduce a customized derivation scheme for each 3D Gaussian, facilitating the generation of multiple "doppelgangers" through learnable parameters for precise position transformation. This approach enables efficient representation of diverse Gaussian attributes and ensures their precision. Additionally, we propose an inherited derivation strategy for newly added Gaussians to expedite training. Extensive experiments demonstrate GaussianHead's efficacy, achieving high-fidelity visual results with a remarkably compact model size ($\approx 12$≈12 MB). Our method outperforms state-of-the-art alternatives in tasks such as reconstruction, cross-identity reenactment, and novel view synthesis. Jie Wang 0137, Jiucheng Xie, Xianyan Li, Feng Xu 0005, Chi-Man Pun, Hao Gao 0005 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | An Automatic Assessment of Parkinson's Disease in Arising from Chair Task via Refined Diffusion-based Pose EstimatorabstractParkinson’s disease (PD) is a progressively common neurodegenerative disorder characterized by a decline in motor function. The diagnosis of PD typically relies on the Movement Disorder Society-Unified Parkinson’s Disease Rating Scale (MDS-UPDRS), which involves subjective scoring through observation of targeted movements. However, this objective method heavily depends on professional experience and has relatively high misdiagnosis rates. In this paper, we introduce a novel vision-based architecture for automated assessment of the ‘arising from chair’ task, which is one of the key MDS-UPDRS components. First, a diffusion-based 2D pose estimator is proposed to enhance keypoint accuracy by iteratively learning the distribution of ground-truth data and then denoising noisy poses. Second, a keypoint trajectory refinement network is introduced to eliminate the jitter error by considering motion information such as position, velocity, acceleration, and jerk. Finally, based on the predicted skeleton keypoint trajectories, we propose several objective indicators to assess the movement characteristics and perform the final rating using the classifier. The experiment substantiates the proposed algorithm, achieving a precision of 98.7% and an accuracy of 95.8% in classifying the ‘arising from chair’ task. Furthermore, the classification results and the proposed objective indicators have been validated as effective aids for neurologists to provide more precise diagnoses. Chi-Man Pun, Haolun Li 0001, Mingliang Zhai, Feng Xu 0005, Hao Gao 0005 |
BIBM | 6 |
| 2024 | Geometry Compression Artifact Removal for V-PCC over a Wide Bitrate RangeabstractIn video-based point cloud compression (V-PCC), point clouds are generated as videos via patch projection to be compressed using video coding techniques. However, a large number of filled empty pixels in the videos creates a fake context, which reduces the noise prediction accuracy in compression artifact removal. Moreover, mean square error (MSE)-based trained models perform better on low-bitrates than on high-bitrates due to the unbalanced parameter updates. This paper proposes an learning-based geometry compression artifact removal for V-PCC over a wide range of bitrates. Firstly, an occupancy map-based contextual feature extraction is proposed to eliminate the interference of empty pixels on the neighboring non-empty pixels. Secondly, an incremental Peak Signal to Noise Ratio (PSNR)-based training scheme is presented to balance the error differences. Experimental results show the effectiveness of the proposed method. Jian Xiong 0005, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 5 |
| 2024 | DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion PredictionabstractPredicting human motion requires addressing dependencies and errors for pose forecasting from sequences. The transformer’s self-attention aids this, but its complexity poses computational challenges. We present an efficient DeformMLP network without self-attention, using fully connected layers. DeformMLP includes DeformFCs, DeformFCt, and DeformFCst layers for spatial temporal modeling and calibration. DeformFCs capture semantics, DeformFCt learns relationships by summarizing time tokens, and DeformFCst assigns significance to dimensions to reduce computation. Our method balances efficiency and accuracy through decomposition and weight allocation. Evaluation on Human3.6M, 3DPW, CMU-MoCap datasets shows state-of-the-art prediction performance by benchmarks. The code is publicly available at https://github.com/HHT-98/DeformMLP. Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Hao Gao 0005 |
ICASSP | 6 |
| 2024 | Local Optimization Networks for Multi-View Multi-Person Human Posture EstimationabstractWith the growing applicability of multi-view multi-person 3D human pose estimation across diverse scenarios, the impact of external environmental factors and occlusion on accuracy has garnered substantial attention. In this research, we introduce a novel approach to multi-view multi-person 3D human pose estimation, leveraging a localized optimization strategy. Specifically, our method enhances the interplay of feature information from different channels and fine-tunes the optimal feature weights to capture intricate dependencies among joints. This refinement leads to improved accuracy in handling external environmental factors. Experimental evaluations were conducted on two prominent benchmark datasets, namely Campus and Shelf. The proposed method achieved a remarkable performance, with a Percentage of Correct Parts (PCP) score of 97.4% and 98.2% for the Campus and Shelf datasets, respectively. Jucheng Song, Chi-Man Pun, Haolun Li 0001, Rushi Lan, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 6 |
| 2024 | Hierarchical Local Temporal Feature Enhancing for Transformer-Based 3D Human Pose EstimationabstractRecent advancements in transformer-based methods have yielded substantial success in 2D-to-3D human pose estimation. Transformer-based estimators have their inherent advantages like global receptive field. Nevertheless, existing transformer approaches ignore the differences among local contexts, resulting in insufficient learning of local information. To address this issue, we introduce non-uniform graph convolution to extract spatial local relationships in skeletons, remedying the limitations of traditional transformers in learning human body topology effectively. Additionally, our proposed Hierarchical Local Temporal Network (HLTN) models local temporal associations across three hierarchical levels: joints, body-parts and poses, effectively addressing the constraint of traditional transformers in learning localized human movements. We connect these two modules in parallel with the spatial and temporal transformer to obtain better features of skeleton sequences. Compared with the latest methods, our method achieves state-of-the-art performance on multiple datasets. Chi-Man Pun, Haolun Li 0001, Hao Gao 0005 |
ICME | 5 |
| 2024 | Towards Distortion-Debiased Blind Image Quality AssessmentabstractExisting blind image quality assessment (BIQA) models are susceptible to biases related to distortion intensity and domain. Intensity bias refers to the relatively accurate perception of severe distortions but larger estimation errors for mild distortions, while domain bias stems from the discrepancies between synthetic and authentic distortion properties. This work introduces a unified learning framework towards addressing these distortion biases. We integrate distortion perception and restoration methods to mitigate intensity bias, where images with minor distortions, which are easily restorable, serve as references for mildly distorted images, while severe distortions benefit directly from distortion perception. The restoration modules employ a combined image-level and feature-level denoising approach, and then an intensity-aware cross-attention mechanism is designed for adaptive handling of intensity bias. To tackle domain bias, we introduce a distortion domain recognition task based on the intrinsic differences between distortion domains and use intra-domain similarity for weighting the quality scores from these domains. Experimental results show that the proposed method achieves state-of-the-art performance on multiple synthetic and authentic distortion datasets. Code and models will be available at https://github.com/xxVENTAZEDxx/Distortion-Debiased-BIQA Lize Zhou, Jian Xiong 0005, Xianzhong Long, Hao Gao 0005 |
ACM Multimedia | 5 |
| 2024 | Scene flow estimation from 3D point clouds based on dual-branch implicit neural representationsabstractAbstract Recently, online optimisation‐based scene flow estimation has attracted significant attention due to its strong domain adaptivity. Although online optimisation‐based methods have made significant advances, the performance is far from satisfactory as only flow priors are considered, neglecting scene priors that are crucial for the representations of dynamic scenes. To address this problem, the authors introduce a dual‐branch MLP‐based architecture to encode implicit scene representations from a source 3D point cloud, which can additionally synthesise a target 3D point cloud. Thus, the mapping function between the source and synthesised target 3D point clouds is established as an extra implicit regulariser to capture scene priors. Moreover, their model infers both flow and scene priors in a stronger bidirectional manner. It can effectively establish spatiotemporal constraints among the synthesised, source, and target 3D point clouds. Experiments on four challenging datasets, including KITTI scene flow, FlyingThings3D, Argoverse, and nuScenes, show that our method can achieve potential and comparable results, proving its effectiveness and generality. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
IET Comput. Vis. | 4 |
| 2024 | Geometry-guided generalizable NeRF for human rendering
Jiucheng Xie, Yiqin Yao, Xun Lv, Shuliang Zhu, Yijing Guo, Hao Gao 0005 |
Multim. Tools Appl. | 6 |
| 2024 | Learning graph-based representations for scene flow estimation
Mingliang Zhai, Hao Gao 0005, Ye Liu 0005, Jianhui Nie, Kang Ni |
Multim. Tools Appl. | 2 |
| 2024 | Adaptive Spatial-Temporal Graph-Mixer for Human Motion PredictionabstractThe Graph Convolutional Network (GCN) has recently achieved promising performance in human motion prediction by modeling the nodes and edges of the human skeleton. However, most previous methods still suffer from two unaddressed drawbacks. First, in the inference stage, their graph topologies are static and fixed, resulting in dependencies between nodes that cannot be dynamically adjusted for different actions. Second, the implicit relationships between pose sequences are ignored, which makes the prior advantages of the graph structure invalid in temporal feature fusion. To address these limitations, we propose an adaptive spatial-temporal graph-mixer (GraphMixer) for human motion prediction, which consists of a series of fully separated spatial-temporal graph convolution structures. In spatial GCN, we construct an additional adaptive skeleton graph to capture the node features of action-specific poses. In temporal GCN, we introduce a variety of graph topologies to enhance feature fusion between pose sequences. Comparing state-of-the-art algorithms on the Human3.6 M and the 3DPW datasets and ablation studies shows that our GraphMixer and the proposed multiple graph topologies are effective and critical. The code is publicly available athttps://github.com/young0304/Adaptive-Spatial-Temporal-Graph-Mixer. Haolun Li 0001, Chi-Man Pun, Chun Du, Hao Gao 0005 |
IEEE Signal Process. Lett. | 5 |
| 2024 | A Parkinson's Auxiliary Diagnosis Algorithm Based on a Hyperparameter Optimization Method of Deep LearningabstractParkinson's disease is a common mental disease in the world, especially in the middle-aged and elderly groups. Today, clinical diagnosis is the main diagnostic method of Parkinson's disease, but the diagnosis results are not ideal, especially in the early stage of the disease. In this paper, a Parkinson's auxiliary diagnosis algorithm based on a hyperparameter optimization method of deep learning is proposed for the Parkinson's diagnosis. The diagnosis system uses ResNet50 to achieve feature extraction and Parkinson's classification, mainly including speech signal processing part, algorithm improvement part based on Artificial Bee Colony algorithm (ABC) and optimizing the hyperparameters of ResNet50 part. The improved algorithm is called Gbest Dimension Artificial Bee Colony algorithm (GDABC), proposing "Range pruning strategy" which aims at narrowing the scope of search and "Dimension adjustment strategy" which is to adjust gbest dimension by dimension. The accuracy of the diagnosis system in the verification set of Mobile Device Voice Recordings at King's College London (MDVR-CKL) dataset can reach more than 96%. Compared with current Parkinson's sound diagnosis methods and other optimization algorithms, our auxiliary diagnosis system shows better classification performance on the dataset within limited time and resources. Shujuan Li, Chi-Man Pun, Yijing Guo, Feng Xu 0005, Hao Gao 0005, Huimin Lu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | An Optimized-Skeleton-Based Parkinsonian Gait Auxiliary Diagnosis Method with Both Monitoring Indicators and Assisted RatingsabstractAbnormal gait is one of the indispensable diagnostic sources of Parkinson’s disease (PD) diagnosis, typically presenting as small shuffling steps and gait bradykinesia. However, its diagnostic accuracy is lower due to the subjective judgments of doctors. To assist in improving the accuracy and reducing the subjectivity of the doctors, we propose an optimized-skeleton-based Parkinsonian gait auxiliary diagnosis method with both monitoring indicators and assisted ratings. By inputting a patient gait video captured from the side, our PD symptom-applicable pose trajectory model will extract a more precise and stable 2D skeleton sequence of patients. Next, the sequence will be used to calculate our proposed five monitoring indicators: gait frequency, ankle speed, whole speed, ankle angle speed, ankle acceleration, and previous work indicators: arm swing angle, leg angle, two feet x-axis distance to record the patient’s gait details at every moment. The extracted gait frequency will then be input into a random forest model to obtain the gait rating. Lastly, doctors can make more accurate judgments by referring to our objective monitoring indicators and assisted ratings. Experimental results show that our monitored indicators improve the doctors’ diagnosis accuracy by 16%, the skeleton speed and acceleration error of our optimized-skeleton extraction method achieve 4.18 cm/s and 5.71 cm/s2, and our random forest model has reached a classification accuracy of 95.8%. Gaoqi Li, Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Feng Xu 0005, Hao Gao 0005 |
BIBM | 6 |
| 2023 | Learning Hybrid Representations of Semantics and Distortion for Blind Image Quality AssessmentabstractRecently, some studies have shown that semantic and distortion representations both benefit the evaluation of image quality. However, the images of existing synthetic distortion databases are annotated with subjective quality scores and distortion types, lacking labels with semantic objects. Therefore, it is virtually infeasible to learn the representations of image semantics and distortion by co-guiding with semantic and distortion labels. To address this issue, we propose a dual-perception network (DPNet) via an end-to-end multi-task learning method, where knowledge distillation is lever-aged as a semantic label-free strategy. Specifically, semantic representation derived from pre-trained ResNet152 is applied to supervise the output of DPNet, while the output is utilized to construct a distortion recognition task. In this way, image semantics and distortion can be hybridly represented in an identical feature map. Finally, image quality is regressed based on the hybrid representations. Experimental results conducted on five benchmark databases validate that the proposed method can achieve state-of-the-art performance. Jian Xiong 0005, Jin-Li Suo, Hao Gao 0005 |
ICASSP | 5 |
| 2023 | ψ-Net: Point Structural Information Network for No-Reference Point Cloud Quality AssessmentabstractThe human vision system is highly adapted to extract structural information from the viewed scenes. The irregularity of point clouds makes the extraction of structural information containing both color and geometry an important challenge for point cloud quality assessment (PCQA). This paper proposes a point structural information (PSI) network (ψ-Net) for no-reference PCQA. Firstly, a PSI module is proposed to map the position vectors of neighboring points to weights for the calculation of color and geometric structure information. Secondly, a dual-stream network is presented to introduce distortion-related features for PCQA. Experimental results show the effectiveness of the proposed method. Jian Xiong 0005, Jin-Li Suo, Hao Gao 0005 |
ICASSP | 5 |
| 2023 | Spike-Based Optical Flow Estimation Via Contrastive LearningabstractSpiking cameras have shown promising advantages for optical flow estimation in high-speed scenarios. The recent work SCFlow [1] attempts to train an optical flow model using spike frames based on a multi-scale flow reconstruction loss. However, only using the flow reconstruction loss is unable to effectively deal with the details of motion, which may lead to noise and blur in the estimated flow fields. To address this issue, we introduce a contrastive loss into spike-based optical flow estimation, which exploits both the information of positive samples and negative samples. Moreover, we propose a refinement step with flexible reception fields to effectively refine the initial flow fields. Experiments on the spiking optical flow dataset PHM demonstrate that the proposed network is effective for spike-based optical flow estimation. In addition, our method achieves competitive performance compared to recent spike-based, frame-based, and event-based methods. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 4 |
| 2023 | Cross-Modal Optical Flow Estimation via Modality Compensation and AlignmentabstractCross-modal optical flow estimation aims to predict motion fields between two frames collected from different modalities, recently attracting intensive attention. However, a substantial yet challenging problem is how to match images across a large modal discrepancy. In this paper, we propose a modality compensation module (MCM) to extract complementary features from different modalities adaptively. Moreover, a cross-modal feature alignment loss is introduced into our network, pulling the compensative features of two cross-modal frames closer and effectively reducing the modal discrepancy. The experimental results demonstrate that our method can achieve competitive performance on the cross-modal optical flow dataset CrossKITTI. Moreover, we experimentally verify that the proposed MCM and cross-modal feature alignment loss are effective for cross-modal optical flow estimation. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 4 |
| 2023 | Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion CuesabstractScene flow estimation is critical for real-world vision problems such as autonomous driving and augmented reality. Due to the popularity of 3D LiDAR sensors, scene flow estimation from 3D point clouds arouses increasing attention. Existing methods usually use a flow embedding-based layer to find correspondences between point pairs. However, only using a flow embedding-based layer is not enough to model the global mutual relationship between two features due to local matching. In this paper, we introduce a cross-transformer to capture more reliable dependencies for point pairs. Moreover, a global motion-aware module is adopted to learn large displacements with a non-local approach. The experimental results demonstrate that the proposed method achieves comparable performance on public datasets and confirm the effectiveness of exploiting the cross-transformer and global motion cues for scene flow estimation. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 4 |
| 2023 | Scene Flow Estimation from Point Clouds with Contrastive Loss and Dual Pseudo LabelsabstractScene flow estimation aims to extract the 3D motion vector between each surface point in two consecutive point clouds. Pseudo-label-based approaches usually exploit point-to-point relations and 3D geometry information to generate the pseudo label for self-supervised learning. However, unreasonable results are still obtained due to the unexploited information of negative samples. Moreover, previous approaches are limited by the fact that pseudo labels are only generated along the forward direction, ignoring the backward direction that has strong spatiotemporal correlations with the forward direction. In this paper, we address these issues in a simple yet effective manner. Specifically, we introduce a contrastive loss to exploit both the information of positive samples and negative samples. Furthermore, we design a dual pseudo labels generation strategy to provide a bidirectional self-supervision for scene flow estimation. Experiments on FlyingThings3D and KITTI datasets show that our method can achieve competitive performance compared to recent self-supervised methods. Mingliang Zhai, Kang Ni, Jiucheng Xie, Xuezhi Xiang, Hao Gao 0005 |
ICIP | 5 |
| 2023 | Non-Local Geometry and Color Gradient Aggregation Graph Model for No-Reference Point Cloud Quality AssessmentabstractNo-Reference point cloud quality assessment (NR-PCQA) is a challenging task in computer vision due to the irregularity of point cloud structures and the unavailability of reference information. Existing point-based and projection-based NR-PCQA models are limited by the representation of point cloud distortion and the modeling of spatial topological structure. To address these limitations, we first propose two visual quality-related gradients: local-maximum geometry gradient and distance-weighted color gradient, which can effectively represent local variations in terms of spatial structure and color intensities between adjacent points. We further propose a non-local geometry and color gradient aggregation graph model for evaluating the perceptual quality of point clouds. Specifically, local graph convolutions are designed to model the topological relationship across neighboring points by aggregating the geometry and color gradients. Furthermore, a position-adaptive self-attention mechanism is introduced to expand the receptive field for modeling the global dependencies of point clouds. Experimental results on two benchmark databases demonstrate that the proposed model outperforms existing state-of-the-art methods. Hao Gao 0005, Jian Xiong 0005 |
ACM Multimedia | 3 |
| 2023 | Efficient Geometry Surface Coding in V-PCCabstractIn recent video-based point cloud compression (V-PCC), 3D point clouds are projected onto 2D images and compressed by High-Efficiency Video Coding (HEVC). However, HEVC was originally designed for natural visual signals, which is a suboptimal framework for point clouds. Therefore, there are still problems in geometry information compression in V-PCC: (1) The distortion based on the sum of squared error (SSE) in the existing rate-distortion optimization (RDO) is inconsistent with the geometric quality measurement; (2) The existing prediction cannot explore the fixed relationship between the corresponding far layer and near layer depth, which means that the far layer depth can be always not less than the corresponding near layer depth. In this paper, we present an efficient geometry surface coding (EGSC) method for V-PCC to address the problems. Firstly, an error projection (EP) model is designed to establish the relationship between the SSE-based distortion and the geometry quality metric. Secondly, an EP-based RDO is employed to improve the geometry information compression by estimating the point normals with gradients. Finally, an occupancy-map driven scheme is proposed to improve the prediction accuracy of merge modes. Experimental results show that the proposed method achieves an average of over 10% bit-rate saving compared with the V-PCC reference software. Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, King Ngi Ngan, Weisi Lin |
IEEE Trans. Multim. | 2 |
| 2022 | SingleMatch: a point cloud coarse registration method with single match point and deep-learning describer
Jianhui Nie, Hao Gao 0005, Ye Liu 0005, Haotian Lu 0006 |
Multim. Tools Appl. | 3 |
| 2022 | Occupancy Map Guided Fast Video-Based Dynamic Point Cloud CodingabstractIn video-based dynamic point cloud compression (V-PCC), 3D point clouds are projected into patches, and then the patches are padded into 2D images suitable for the video compression framework. However, the patch projection-based method produces a large number of empty pixels; the far and near components are projected to generate different 2D images (video frames), respectively. As a result, the generated video is with high resolutions and double frame rates, so the V-PCC has huge computational complexity. This paper proposes an occupancy map guided fast V-PCC method. Firstly, the relationship between the prediction coding and block complexity is studied based on a local linear image gradient model. Secondly, according to the V-PCC strategies of patch projection and block generation, we investigate the differences of rate-distortion characteristics between different types of blocks, and the temporal correlations between the far and near layers. Finally, by taking advantage of the fact that occupancy maps can explicitly indicate the block types, we propose an occupancy map guided fast coding method, in which coding is performed on the different types of blocks. Experiments have tested typical dynamic point clouds, and shown that the proposed method achieves an average 43.66% time-saving at the cost of only 0.27% and 0.16% Bjontegaard Delta (BD) rate increment under the geometry Point-to-Point (D1) error and attribute Luma Peak-Signal-Noise-Ratio (PSNR), respectively. Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | An Efficient Artificial Bee Colony Algorithm With an Improved Linkage Identification MethodabstractThe artificial colony (ABC) algorithm shows a relatively powerful exploration search capability but is constrained by the curse of dimensionality, especially on nonseparable functions, where its convergence speed slows dramatically. In this article, based on an analysis of the difference between updating mechanisms that include both all-variable and one-variable updating mechanisms, we find that when equipped with the former strategy, the algorithm rapidly converges to an optimal region, while with the latter strategy, it searches the solution space thoroughly. To utilize multivariable and one-variable updating mechanisms on nonseparable and separable functions, respectively, we embed an improved linkage identification strategy into the ABC by detecting the linkage between variables more effectively. Then, we propose three common strategies for ABC to improve its performance. First, a new approach that considers the historic experiences of the population is proposed to balance exploration and exploitation. Second, a new strategy for initializing scout bees is used to reduce the number of function evaluations. Finally, the individual with the worst performance is updated with a defined probability on multiple dimensions instead of one dimension, causing it to follow the population steps on nonseparable functions. This article is the first to propose all these concepts, which could be adopted for other ABC variants. The effectiveness of our algorithm is validated through basic, CEC2010, CEC2013, and CEC2014 functions and real-world problems. Hao Gao 0005, Zheng Fu, Chi-Man Pun, Jun Zhang 0003, Sam Kwong |
IEEE Trans. Cybern. | 1 |
| 2022 | Action Recognition Framework in Traffic Scene for Autonomous Driving SystemabstractFor the autonomous driving system, accurately recognizing the actions of different roles in the traffic scene is the prerequisite for realizing this kind of human-vehicle information interaction. In this paper, we propose a complete framework based on 3D human pose estimation to recognize the actions of different roles on the road. The main objects recognized include traffic police, cyclists, and some passersby in need. We perform action recognition based on a dynamic adaptive graph convolutional network, which can realize the action recognition of objects based on 3D human pose. In addition to the action recognition module, we have optimized both the object detection module and the human pose estimation module in the framework so that the framework can handle multiple objects at the same time, which can be closer to the real traffic scene. To realize complex and changeable human action recognition, we built a multi-view camera system to collect responsible 3D human pose datasets containing traffic police gestures, cyclist gestures, and pedestrians’ body movements. In the experiments, compared to other state-of-the-art researches, the proposed framework can achieve comparable results with the same dataset. Satisfactory performance has also been obtained on the real data we collected, which can handle a variety of different action recognition tasks at the same time. Feiyi Xu, Feng Xu 0005, Jiucheng Xie, Chi-Man Pun, Huimin Lu 0001, Hao Gao 0005 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Virtual Reality Aided High-Quality 3D Reconstruction by Remote DronesabstractArtificial intelligence including deep learning and 3D reconstruction methods is changing the daily life of people. Now, an unmanned aerial vehicle that can move freely in the air and avoid harsh ground conditions has been commonly adopted as a suitable tool for 3D reconstruction. The traditional 3D reconstruction mission based on drones usually consists of two steps: image collection and offline post-processing. But there are two problems: one is the uncertainty of whether all parts of the target object are covered, and another is the tedious post-processing time. Inspired by modern deep learning methods, we build a telexistence drone system with an onboard deep learning computation module and a wireless data transmission module that perform incremental real-time dense reconstruction of urban cities by itself. Two technical contributions are proposed to solve the preceding issues. First, based on the popular depth fusion surface reconstruction framework, we combine it with a visual-inertial odometry estimator that integrates the inertial measurement unit and allows for robust camera tracking as well as high-accuracy online 3D scan. Second, the capability of real-time 3D reconstruction enables a new rendering technique that can visualize the reconstructed geometry of the target as navigation guidance in the HMD. Therefore, it turns the traditional path-planning-based modeling process into an interactive one, leading to a higher level of scan completeness. The experiments in the simulation system and our real prototype demonstrate an improved quality of the 3D model using our artificial intelligence leveraged drone system. Feng Xu 0005, Chi-Man Pun, Yang Yang 0002, Rushi Lan, Yujie Li 0001, Hao Gao 0005 |
ACM Trans. Internet Techn. | 8 |
| 2022 | Localizing and tracking dense crowd of microbes by joint association and detection refinement
Ye Liu 0005, Shuohong Wang, Jianhui Nie, Hao Gao 0005 |
Vis. Comput. | 4 |
| 2021 | A Rate-based Drone Control with Adaptive Origin Update in TelexistenceabstractA new form of telexistence is achieved by recording videos with a camera on an Uncrewed aerial vehicle (UAV) and playing the videos to a user via a head-mounted display (HMD). One key problem here is how to let the user freely and naturally control the UAV and thus the viewpoint. In this paper, we develop an HMD-based telexistence technique that achieves full 6- DOF control of the viewpoint. The core of our technique is an improved rate-based control technique with our adaptive origin update (AOU), in which the origin of the coordinate system of the user changes adaptively. This makes the user naturally perceive the origin and thus easily perform the control motion to get his/her desired viewpoint changing. As a consequence, without the aid of any auxiliary equipment, the AOU scheme handles the well known self-centering problem in the rate-based control methods. A real prototype is also built to evaluate this feature of our technique. To explore the advantage of our telexistence technique, we further use it as an interactive tool to perform the task of 3D scene reconstruction. User studies demonstrate that comparing with other telexistence solutions and the widely used joystick-based solutions, our solution largely reduces the workload and saves time and moving distance for the user. Chi-Man Pun, Yang Yang 0002, Hao Gao 0005, Feng Xu 0005 |
VR | 4 |
| 2021 | Enhancement of ridge-valley features in point cloud based on position and normal guidance
Jianhui Nie, Zhaochen Zhang, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Wenkai Shi |
Comput. Graph. | 4 |
| 2021 | An improved artificial bee colony algorithm based on elite search strategy with segmentation application on robot vision systemabstractSummary Aiming at accelerating the convergence speed and enhancing relative poor local search ability of the traditional artificial bee colony algorithm (ABC), this article introduces an ABC with a new elite search strategy. First, we propose a strategy of recording individuals with high performance. Then bees have more chances to learn from a real elite. In the onlooked bee phase, its updating equation is changed for having more opportunities to search in a valuable area. Furthermore, for saving the value of function evaluations, a new learning equation for the best onlooked bee is proposed. The image segmentation of a robot binocular stereo vision system is a key problem in mechanical robot vision system, but the computation time limits its application. The experimental results show that the proposed algorithm achieves better performance on 10 benchmark functions and the image segmentation problem of mechanical robot in comparison with several other state of the art algorithms. Chuyi Gao, Maolong Xi, Jian Xiong 0005, Chi-Man Pun, Hao Gao 0005 |
Concurr. Comput. Pract. Exp. | 8 |
| 2021 | Bas-relief generation from point clouds based on normal space compression with real-time adjustment on CPU
Jianhui Nie, Wenkai Shi, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Zhaochen Zhang |
Graph. Model. | 4 |
| 2021 | A Hybrid Feature Selection Algorithm Based on a Discrete Artificial Bee Colony for Parkinson's DiagnosisabstractParkinson's disease is a neurodegenerative disease that affects millions of people around the world and cannot be cured fundamentally. Automatic identification of early Parkinson's disease on feature data sets is one of the most challenging medical tasks today. Many features in these datasets are useless or suffering from problems like noise, which affect the learning process and increase the computational burden. To ensure the optimal classification performance, this article proposes a hybrid feature selection algorithm based on an improved discrete artificial bee colony algorithm to improve the efficiency of feature selection. The algorithm combines the advantages of filters and wrappers to eliminate most of the uncorrelated or noisy features and determine the optimal subset of features. In the filter, three different variable ranking methods are employed to pre-rank the candidate features, then the population of artificial bee colony is initialized based on the significance degree of the re-rank features. In the wrapper part, the artificial bee colony algorithm evaluates individuals (feature subsets) based on the classification accuracy of the classifier to achieve the optimal feature subset. In addition, for the first time, we introduce a strategy that can automatically select the best classifier in the search framework more quickly. By comparing with several publicly available datasets, the proposed method achieves better performance than other state-of-the-art algorithms and can extract fewer effective features. Haolun Li 0001, Chi-Man Pun, Feng Xu 0005, Longsheng Pan, Rui Zong, Hao Gao 0005, Huimin Lu 0001 |
ACM Trans. Internet Techn. | 6 |
| 2020 | New multi-view human motion capture frameworkabstractEstimating human pose and shape without markers is a challenging problem. This study proposes a multiple‐view markerless human motion capture framework. Firstly, a multi‐view camera system is built for capturing real‐time images of moving humans on multiple views. Secondly, by employing the OpenPose method, the authors calculate robust 3D key points from 2D key points of the human body, which are estimated from the multi‐view images. And dense 3D point cloud is reconstructed from images. Thirdly, they propose a novel SMPL‐based method to represent human motion by fitting the SMPL model to 3D key points and 3D point clouds. In order to achieve a more accurate human pose, a penalty term is utilised to solve the problem of error accumulation in the process of human motion capture. In addition, they present a dense mesh template‐based SMPL that can be deformed to point cloud to recover a real human body shape. Finally, they map multi‐view colour images onto the human mesh model to acquire rendered mesh. The experimental results show that the proposed method improves the accuracy of human pose and realises the 3D human body model more realistic. Feiyi Xu, Chi-Man Pun, Wenqi Xiao, Jianhui Nie, Jian Xiong 0005, Hao Gao 0005, Feng Xu 0005 |
IET Image Process. | 7 |
| 2020 | Training Feed-Forward Artificial Neural Networks with a modified artificial bee colony algorithm
Feiyi Xu, Chi-Man Pun, Haolun Li 0001, Yushu Zhang 0001, Yurong Song, Hao Gao 0005 |
Neurocomputing | 6 |
| 2020 | Endmember Extraction of Hyperspectral Remote Sensing Images Based on an Improved Discrete Artificial Bee Colony Algorithm and Genetic Algorithm
Zheng Fu, Chi-Man Pun, Hao Gao 0005, Huimin Lu 0001 |
Mob. Networks Appl. | 3 |
| 2020 | An artificial bee algorithm with a leading group and its application into image registration
Haidong Hu, Chi-Man Pun, Ye Liu 0005, Xiangjing Lai, Hao Gao 0005 |
Multim. Tools Appl. | 6 |
| 2020 | Vision-based position and pose determination of non-cooperative target for on-orbit servicing
Haidong Hu, Dayi Wang, Hao Gao 0005, Chunling Wei, Yingzi He |
Multim. Tools Appl. | 3 |
| 2020 | Vehicle power train optimization using multi-objective bird swarm algorithm
Dongmei Wu, Chi-Man Pun, Bin Xu 0014, Hao Gao 0005, Zhenghua Wu |
Multim. Tools Appl. | 4 |
| 2020 | High-quality-guided artificial bee colony algorithm for designing loudspeaker
Hao Gao 0005, Haolun Li 0001, Ye Liu 0005, Huimin Lu 0001, Hyoungseop Kim, Chi-Man Pun |
Neural Comput. Appl. | 1 |
| 2020 | Face image super-resolution with pose via nuclear norm regularized structural orthogonal Procrustes regression
Guangwei Gao, Meng Yang 0001, Huimin Lu 0001, Wankou Yang, Hao Gao 0005 |
Neural Comput. Appl. | 6 |
| 2019 | A 6-DOF Telexistence Drone Controlled by a Head Mounted DisplayabstractRecently, a new form of telexistence is achieved by recording images with cameras on an unmanned aerial vehicle (UAV) and displaying them to the user via a head mounted display (HMD). A key problem here is how to provide a free and natural mechanism for the user to control the viewpoint and watch a scene. To this end, we propose an improved rate-control method with an adaptive origin update (AOU) scheme. Without the aid of any auxiliary equipment, our scheme handles the self-centering problem. In addition, we present a full 6-DOF viewpoint control method to manipulate the motion of a stereo camera, and we build a real prototype to realize this by utilizing a pan-tilt-zoom (PTZ) which not only provides 2-DOF to the camera but also compensates the jittering motion of the UAV to record more stable image streams. Xingyu Xia, Chi-Man Pun, Yang Yang 0002, Huimin Lu 0001, Hao Gao 0005, Feng Xu 0005 |
VR | 6 |
| 2019 | Automatic Medical Image Registration Based on an Integrated Method Combining Feature and Area Information
Jiucheng Xie, Chi-Man Pun, Zhaoqing Pan, Hao Gao 0005, Baoyun Wang |
Neural Process. Lett. | 4 |
| 2019 | An Improved Artificial Bee Colony Algorithm With its ApplicationabstractThe artificial bee colony is a popular evolutionary algorithm that exhibits strong exploration ability but slow convergence. This paper proposes two new updating equations to boost the performances of employed and onlooker bees, respectively. In the new updating equations, two intelligent learning strategies give bees a chance to learn from individuals with better performances. New control operators are also utilized to balance global and local searches. Second, we define a new search direction mechanism to overcome the oscillation phenomenon in employed bees. Finally, an intelligent learning mechanism is proposed to accelerate the convergence rate of the worst employed bee. To test the effectiveness of our algorithm, a series of benchmark functions and two industrial problems are utilized. Experimental results demonstrate that our proposed algorithm performs more favorably on both theoretical and practical problems. Hao Gao 0005, Yujiao Shi 0002, Chi-Man Pun, Sam Kwong |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Context-Aware Three-Dimensional Mean-Shift With Occlusion Handling for Robust Object Tracking in RGB-D VideosabstractDepth cameras have recently become popular and many vision problems can be better solved with depth information. But, how to integrate depth information into a visual tracker to overcome the challenges such as occlusion and background distraction is still underinvestigated in current literature on visual tracking. In this paper, we investigate a 3-D extension of a classical mean-shift tracker whose greedy gradient ascend strategy is generally considered as unreliable in conventional 2-D tracking. However, through careful study of the physical property of 3-D point clouds, we reveal that objects which may appear to be adjacent on a 2-D image will form distinctive modes in the 3-D probability distribution approximated by kernel density estimation, and finding the nearest mode using 3-D mean-shift can always work in tracking. Based on the understanding of 3-D mean-shift, we propose two important mechanisms to further boost the tracker's robustness: one is to enable the tracker to be aware of potential distractions and make corresponding adjustments to the appearance model; and the other is to enable the tracker to detect and recover from tracking failures caused by total occlusion. The proposed method is both effective and computationally efficient. On a conventional personal computer, it runs at more than 60 FPS without graphical processing unit acceleration. Ye Liu 0005, Xiaoyuan Jing, Jianhui Nie, Hao Gao 0005, Jun Liu 0036, Guoping Jiang |
IEEE Trans. Multim. | 4 |
| 2018 | Applying stochastic second-order entropy images to multi-modal image registration
Xiaodong Cun, Chi-Man Pun, Hao Gao 0005 |
Signal Process. Image Commun. | 3 |
| 2016 | An efficient image segmentation method based on a hybrid particle swarm algorithm with learning strategy
Hao Gao 0005, Chi-Man Pun, Sam Kwong |
Inf. Sci. | 1 |
| 2016 | An improved artificial bee colony and its application
Yujiao Shi 0002, Chi-Man Pun, Haidong Hu, Hao Gao 0005 |
Knowl. Based Syst. | 4 |
| 2014 | Online discriminative dictionary learning via label information for multi task object trackingabstractIn this paper, a supervised approach to online learn a structured sparse and discriminative representation for object tracking is presented. Label information from training data is incorporated into the dictionary learning process to construct a compact and discriminative dictionary. This is accomplished by adding an ideal-code regularization term and classification error term to the total objective function. By minimizing the total objective function, we learn the high quality dictionary and optimal linear multi-classifier simultaneously. Combined with multi task sparse learning, the learned classifier is employed directly to separate the object from background. As the tracking continues, the proposed algorithm alternates between multi task sparse coding and dictionary updating. Experimental evaluations on the challenging sequences show that the proposed algorithm performs favorably against state-of-the-art methods in terms of effectiveness, accuracy and robustness. Baojie Fan, Yingkui Du, Hao Gao 0005, Baoyun Wang |
ICME | 3 |
| 2014 | A Hybrid Particle-Swarm Tabu Search Algorithm for Solving Job Shop Scheduling ProblemsabstractThis paper proposes a method for the job shop scheduling problem (JSSP) based on the hybrid metaheuristic method. This method makes use of the merits of an improved particle swarm optimization (PSO) and a tabu search (TS) algorithm. In this work, based on scanning a valuable region thoroughly, a balance strategy is introduced into the PSO for enhancing its exploration ability. Then, the improved PSO could provide diverse and elite initial solutions to the TS for making a better search in the global space. We also present a new local search strategy for obtaining better results in JSSP. A real-integer encode and decode scheme for associating a solution in continuous space to a discrete schedule solution is designed for the improved PSO and the tabu algorithm to directly apply their solutions for intensifying the search of better solutions. Experimental comparisons with several traditional metaheuristic methods demonstrate the effectiveness of the proposed PSO-TS algorithm. Hao Gao 0005, Sam Kwong, Baojie Fan, Ran Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2013 | Particle swarm optimization based on intermediate disturbance strategy algorithm and its application in multi-threshold image segmentation
Hao Gao 0005, Sam Kwong, Jingjing Cao |
Inf. Sci. | 1 |