Yong Wang 0053

dblp:84/2694-53 · DBLP profile ↗
← Back
29ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0002-7847-3807ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MuRE: Multi-Relationship Encoder for 3D human pose estimation
Yong Wang 0053, Doudou Wu, Hongbo Kang, Wenming Yang
Comput. Vis. Image Underst.1
2026 Hyperspectral image denoising via enhanced Laplacian total variation regularizer
Yusen Tan, Yong Wang 0053, Tao Jia 0001, Zhi Wang 0015
Expert Syst. Appl.3
2026 Fidelity-Preserving Concept Stylization with ST-LoRA and multimodal conditions
Suoyu Zhang, Eric C. C. Tsang, Yong Wang 0053
Expert Syst. Appl.3
2026 A structured document understanding model based on gate mechanism and cross attention
Yong Wang 0053
Int. J. Document Anal. Recognit.3
2026 SkyPose: Skeleton consistency guided diffusion model for 3D human pose estimation
abstract
Although diffusion models have demonstrated promising potential in 3D human pose estimation, existing approaches commonly adopt fixed noise variance scheduling strategies. This design overlooks the skeletal consistency constraints that should be maintained between 2D observations and 3D predictions during iterative denoising. Consequently, the generated 3D poses often exhibit geometric inconsistencies with the input 2D poses, thereby degrading the overall reconstruction quality. To address this limitation, we propose SkyPose: Skeleton Consistency Guided Diffusion Model for 3D Human Pose Estimation. The proposed method enforces skeletal consistency by reprojecting the previously generated 3D pose onto the 2D plane, decomposing it into skeletal direction and length, and comparing it with the input 2D skeletal structure to construct a dynamic skeletal consistency metric. This metric adaptively adjusts the noise variance σt during DDIM iterations, explicitly preserving skeletal structural coherence throughout the denoising process and enhancing the geometric fidelity of the generated 3D poses. Furthermore, we redesign the denoising network architecture to integrate skeletal structure features with joint coordinate features at the local joint, body-part, and full-body levels, thereby enabling joint–skeleton collaborative denoising under multi-granularity geometric constraints. Extensive experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that our approach achieves superior accuracy compared with existing methods. The code is open-sourced at here.
Xuguang Liu, Yong Wang 0053, Wenxiu Dan, Wenming Yang
Knowl. Based Syst.2
2026 SARL: Structure-aware representation learning for 3D human pose estimation from point clouds
Yong Wang 0053, Mengyuan Liu 0001
Pattern Recognit.2
2026 DBMambaPose: Decoupled spatial-temporal bidirectional state space model for efficient 3D human pose estimation
Yong Wang 0053, Xuguang Liu, Hongbo Kang, Wenming Yang
Pattern Recognit.2
2026 Improving Fine-Grained Understanding for Retrieval in Human Motion and Text
abstract
This work focuses on human motion-text retrieval (MTR), a task recently proposed for motion understanding. Unlike traditional visual-text retrieval, human motion can be understood as the superposition of numerous atomic actions, and its description is also limited to human-centered themes. Considering this characteristic, directly mapping similar samples into a joint embedding space and conducting naive contrastive training is suboptimal, as it lacks cognition of fine-grained human language descriptions and fails to alleviate semantic conflicts between similar samples. To address this, we propose a meticulous Cross-perceptual Salience Mapping, highlighting fine-grained poses or words to provide more accurate similarity measurement. Additionally, a novel Drop-then-Contrast scheme is designed for MTR, discarding false negative samples from the negative set and mining the remaining sample for contrastive training to reduce violations they caused. Our framework, termed as improving fine-grained understanding for Retrieval in Human Motion and Text or Rehamot for short, outperforms previous works by a recall of 58.6% and 56.5% on HumanML3D and KIT-ML respectively (motion retrieval, R@10).
Yong Wang 0053, Hongchang Jin, Mengyuan Liu 0001
IEEE Signal Process. Lett.2
2026 DRPose: A Diffusion-Based Pose Refinement Framework for 3D Human Pose Estimation
abstract
Recently, two-stage 3D human pose estimation using monocular cameras has gained significant attention. However, the inherent uncertainty in the upscaling process from 2D to 3D often compromises the accuracy of deterministic methods. To address this, we propose a novel diffusion-based refinement framework (DRPose) which models the uncertainty during the upscaling process by introducing stochastic noise to the initially predicted 3D poses. This approach facilitates the generation of more realistic predictions through iterative refinement with multiple noise samples, ultimately producing multi-hypothesis predictions that better align with ground truth. Our framework incorporates two key components: a Graph Convolution Transformer module (SGCT), which integrates scaling and displacement adjustments based on conditional information with a joint temporal-spatial feature separation mechanism, and a Pose Refinement Module (PRM), which balances the initial and refined poses. This design allows DRPose to effectively refine pose estimation for both individual frames and sequential data. Furthermore, our framework establishes new benchmarks for performance in bothframe2frameandseq2framescenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the Human3.6M and MPI-INF-3DHP datasets. Notably, when applied to the current state-of-the-art single-frame 3D pose extractor, our multi-hypothesis optimization achieves an 18.8% reduction in Mean Per Joint Position Error (MPJPE) and a 16.9% reduction in Procrustes MPJPE (P-MPJPE). Code is available at https://github.com/KHB1698/DRPose.
Yong Wang 0053, Xuguang Liu, Doudou Wu, Wenming Yang, Hongbo Kang
IEEE Trans. Circuits Syst. Video Technol.1
2026 Multi-Scale Local-Global Fusion for Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) is a formidable computer vision challenge due to the striking resemblance between camouflaged objects and their surroundings. Despite progress in existing methods, they still face significant limitations, particularly in addressing the issues of fuzzy boundaries and the inadequate fusion of local and global features. To address these challenges, we present a multi-scale COD network named Multi-Scale Local-Global Fusion (MSLGF). MSLGF incorporates a Multi-Scale Fusion Module (MSFM), which skillfully integrates feature maps at multiple scales to produce high-fidelity edge features. Additionally, to refine the detection process, a Local-Global Feature Fusion Module (LGFFM) combines the local edge details with global semantic information of camouflaged targets, significantly enhancing the accuracy of COD. Experimental results show that MSLGF achieves remarkable performance across 3 benchmark datasets, i.e., Camouflaged Object Dataset (CAMO), Camouflaged Object Dataset with 10,000 Images (COD10K), and NC4K. Specifically, MSLGF attains a structure-measure from 0.879 to 0.894 and a weighted F-measure between 0.817 and 0.856. The source code is publicly available at https://github.com/tc-fro/MSLGF.
Boran Yang, Yong Wang 0053, Duoqian Miao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Double-Chain Graph Convolution Transformer for 3D Human Pose Estimation
abstract
Reconstructing 3D poses from 2D poses lacking depth information is particularly challenging due to the complexity and diversity of human motion. The key is to effectively model the spatial constraints between joints to leverage their inherent dependencies. Thus, we propose a novel model, called Double-chain Graph Convolution Transformer (DC-GCT), to constrain the pose through a double-chain design consisting of local-to-global and global-to-local chains to obtain a complex representation more suitable for the current human pose. Specifically, we combine the advantages of GCN and Transformer and design a Local Constraint Module (LCM) based on GCN and a Global Constraint Module (GCM) based on self-attention mechanism as well as a Feature Interaction Module (FIM). The proposed method fully captures the multi-level dependencies between human body joints to optimize the modeling capability of the model. Moreover, we propose a method to use temporal information into the single-frame model by guiding the video sequence embedding through the joint embedding of the target frame, with negligible increase in computational cost. Experimental results demonstrate that DC-GCT achieves state-of-the-art performance on two challenging datasets (Human3.6 M and MPI-INF-3DHP). Notably, our model achieves state-of-the-art performance on all action categories in the Human3.6 M dataset using detected 2D poses from CPN, and our code is available at:https://github.com/KHB1698/DC-GCT.
Hongbo Kang, Yong Wang 0053, Mengyuan Liu 0001, Doudou Wu, Wenming Yang
IEEE Trans. Multim.2
2025 Offset attention with seed generation for point cloud completion
abstract
Abstract Point cloud data acquired through 3D scanning is frequently subject to fragmentation due to the constraints of the scanner’s field of view and occlusions within the scanned object. The ensuing incompleteness in the data can significantly degrade the accuracy of subsequent computational tasks. Traditional methods for predicting complete point clouds from these fragments often fail to capture the fine-grained local details, leading to inaccurate reconstructions. In this work, we introduce a novel neural network architecture designed for point cloud completion that addresses these limitations.Our network accepts an incomplete point cloud and employs a multi-scale feature extraction module, which integrates an offset attention mechanism alongside a feature aggregation module operating across various scales. This dual approach significantly bolsters the network’s capacity to discern both local and global features inherent in the point cloud data. Furthermore, we incorporate a seed generation module within our missing point cloud generator, harnessing a hierarchical feature pyramid network to forecast the entirety of the point cloud. This innovative strategy allows our network to accurately predict the structure of missing regions.Empirical evaluations conducted on the Shapenet-Part and ModelNet10 datasets substantiate the efficacy of our proposed methodology. Our approach outperforms the state-of-the-art PF-Net algorithm, achieving a remarkable reduction in chamfer distance by 16.15$\%$ and 41.87$\%$ on the respective datasets. Visual inspection of the results underscores the robust generalization capabilities of our algorithm, which is particularly evident in scenarios with limited dataset sizes. It adeptly predicts the contours of the missing regions and synthesizes a more comprehensive point cloud shape.
Yong Wang 0053
Comput. J.2
2025 Multi-scale frequency attention fusion network for infrared and visible image fusion
Yong Wang 0053, Xueyuan Zhao, Jianfei Pu, Duoqian Miao 0001
Eng. Appl. Artif. Intell.1
2025 Camouflaged Object Detection with boundary localization in complex backgrounds
Guangjian Zhang, Zhengming Yang, Yong Wang 0053, Yuliang Chen, Duoqian Miao 0001
Eng. Appl. Artif. Intell.3
2025 ICFNet: Interactive-complementary fusion network for monocular 3D human pose estimation
Yong Wang 0053, Hongbo Kang, Doudou Wu, Duoqian Miao 0001
Neurocomputing1
2025 GMS-YOLO: A Lightweight Real-Time Object Detection Algorithm for Pedestrians and Vehicles Under Foggy Conditions
abstract
In conditions of foggy weather, challenges, such as low light, blurred imagery, and dense fog that obscures target objects are prevalent. Moreover, computing resources are limited on edge devices. To tackle these challenges, a novel real-time detection algorithm GMS-YOLO for pedestrians and vehicles is proposed based on YOLOv10, which overcomes the semantic bottleneck of the model and enhances its detection performance. A novel ghost multiscale convolution (GMSConv) module is constructed, serving as the ghost multiScale feature extraction backbone network (GMS-Net). The Shape Consistent Intersection over Union (SCIoU) is introduced as the localization loss function, which takes into account the influence of the attributes of the regression box in the loss computation. Additionally, a compensatory consistency matching metric (CCMM) formula is designed to reduce the sensitivity of the original metric to IoU and regression scores. The GMS-YOLO algorithm has a lightweight structure, achieving FPS of 94 and 92 during the detection phase at the “‘n”’ and “‘s”’ sizes, respectively. Furthermore, we have deployed the model on Jetson Nano hardware, and the inference speed is also quite encouraging. We validated the effectiveness of the algorithm on the Foggy Cityscapes, RTTS, VOC2007-fog, and VOC2012-fog datasets. Experimental results indicate that GMS-YOLO outperforms the baseline model, with a mean average precision (mAP) improvement of 6.3% and 5.5% for the “n” and “s” scales, respectively. Consequently, the proposed GMS-YOLO algorithm not only demonstrates superior detection performance but also maintains a relatively low model complexity, significantly enhancing the efficiency and accuracy of object detection tasks in foggy environments. The source code for our algorithm is available at:https://github.com/Fwdchina/GMS-YOLO.
Yafei Chen, Yong Wang 0053, Wenxiu Dan
IEEE Internet Things J.2
2025 CSBNet: Leveraging Edge Intelligence for Multigranularity Low-Light Image Enhancement
abstract
Low-light (LOL) conditions constantly restrict the performance of Internet of Things (IoT) image sensors, thereby impacting image quality and the precision of visual data analysis. The emerging edge intelligence is crucial for LOL image enhancement in improving image quality and data support reliability for IoT systems, which in turn fosters the intelligence and automation progress of the IoT. The enhancement of LOL images necessitates the restoration of both contextual information and spatial details, maintaining the semantic content of the original image and the point-to-point correspondence between inputs and outputs. However, existing methods predominantly concentrate on one aspect, either contextual information or spatial details, making it difficult to simultaneously balance both. To overcome this challenge, we introduce a novel two-branch network, the context-space balance network (CSBNet), and tailored for LOL image enhancement. It comprises a contextual information recovery network (CIRNet), which adeptly extracts contextual information from multiscale LOL images, and a spatial information recovery network (SIRNet), which is designed to preserve spatial details at the original resolution. We also implement a context-space feature fusion (CSFF) module to seamlessly integrate contextual information with spatial details. Qualitative and quantitative experimental results demonstrate that our CSBNet can better handle various kinds of degradations in lowlight images compared with state-of-the-art solutions on the benchmark LOL dataset. The source code of CSBNet is available athttps://github.com/Loong161/CSBNet.
Yong Wang 0053, Lijun Jiang, Zilong Du, Bo Li 0115, Wenming Yang
IEEE Internet Things J.1
2025 mLANet: An efficient recurrent neural network for long-term time series forecasting
Jihua Jiang, Yong Wang 0053, Duoqian Miao 0001
Knowl. Based Syst.3
2025 FM-RTDETR: Small Object Detection Algorithm Based on Enhanced Feature Fusion With Mamba
abstract
Traditional real-time object detection networks deployed in autonomous aerial vehicles (AAVs) struggle to extract features from small objects in complex backgrounds with occlusions and overlapping objects. To address this challenge, we propose FM-RTDETR, a real-time object detection algorithm optimized for small object detection. We redesign the encoder of RT-DETRv2 by integrating the Feature Aggregation and Diffusion Network (FADN), improving the algorithm's ability to capture contextual information. Subsequently, we introduce the Parallel Atrous Mamba Feature Fusion Module (PAMFFM), which combines shallow and deep semantic information to better capture small object features. Furthermore, we propose the Cross-stage Enhanced Feature Fusion Module (CEFFM), merging features for small objects to provide richer and more detailed information. Finally, we propose STIoU Loss, which incorporates a penalty term to adjust the scaling of the loss function, improving detection granularity for small objects. FM-RTDETR achieves AP$_{50}$scores of 54.0% and 56.3% on the VisDrone2019-DET and AI-TOD datasets. Compared with other state-of-the-art methods, our method shows great potential in small object detection from drones.
Jiahui Dai, Yong Wang 0053, Yafei Chen
IEEE Signal Process. Lett.3
2024 Diffusion-Based Pose Refinement and Multi-Hypothesis Generation for 3D Human Pose Estimation
abstract
Previous probabilistic models for 3D Human Pose Estimation (3DHPE) aimed to enhance pose accuracy by generating multiple hypotheses. However, most of the hypotheses generated deviate substantially from the true pose. Compared to deterministic models, the excessive uncertainty in probabilistic models leads to weaker performance in single-hypothesis prediction. To address these two challenges, we propose a diffusion-based refinement framework called DRPose, which refines the output of deterministic models by reverse diffusion and achieves more suitable multi-hypothesis prediction for the current pose benchmark by multi-step refinement with multiple noises. To this end, we propose a Scalable Graph Convolution Transformer (SGCT) and a Pose Refinement Module (PRM) for denoising and refining. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets demonstrate that our method achieves state-of-the-art performance on both single and multi-hypothesis 3DHPE. Code is available at https://github.com/KHB1698/DRPose.
Hongbo Kang, Yong Wang 0053, Mengyuan Liu 0001, Doudou Wu, Xinlin Yuan, Wenming Yang
ICASSP2
2024 SCGRFuse: An infrared and visible image fusion network based on spatial/channel attention mechanism and gradient aggregation residual dense blocks
Yong Wang 0053, Jianfei Pu, Duoqian Miao 0001, Longbin Zhang
Eng. Appl. Artif. Intell.1
2024 TFAN: Twin-Flow Axis Normalization for Human Motion Prediction
abstract
Human motion prediction involves forecasting upcoming body poses from historically observed sequences. Presently, various methods focus on modeling the positions of skeletal joints to generate future movements. Beyond solely using the joint position information, this letter further explores the informative bone vector information between joints for human motion prediction. Therefore, we propose TFAN, which integrates joint position and bone vector information to achieve more precise predictions. Additionally, joint positions and bone vectors, both represented as 3D vectors, were frequently imprecisely modeled due to neglect of variations in data distribution across axes in prior methods. Instead, we introduce Axis Normalization, which standardizes each coordinate axis individually, enhancing the model's sensitivity to data distribution disparities. Through experimental evaluations on the Human3.6M, AMASS, and 3DPW datasets, we consistently demonstrate that TFAN outperforms other existing methods. Our code will be available athttps://github.com/Deante-dx/TFAN.
Yong Wang 0053, Zongying Li, Mengyuan Liu 0001
IEEE Signal Process. Lett.2
2024 MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human Motions
abstract
In this paper, we address the unexplored question of temporal sentence localization in human motions (TSLM), aiming to locate a target moment from a 3D human motion that semantically corresponds to a text query. Considering that 3D human motions are captured using specialized motion capture devices, motions with only a few joints lack complex scene information like objects and lighting. Due to this character, motion data has low contextual richness and semantic ambiguity between frames, which limits the accuracy of predictions made by current video localization frameworks extended to TSLM to only a rough level. To refine this, we devise two novel label-prior-assisted training schemes: one embed prior knowledge of foreground and background to highlight the localization chances of target moments, and the other forces the originally rough predictions to overlap with the more accurate predictions obtained from the flipped start/end prior label sequences during recovery training. We show that injecting label-prior knowledge into the model is crucial for improving performance at high IoU. In our constructed TSLM benchmark, our model termedMLPachieves a recall of 44.13 at [email protected] on the BABEL dataset and 71.17 on HumanML3D (Restore), outperforming prior works. Finally, we showcase the potential of our approach in corpus-level moment retrieval. Our source code is openly accessible athttps://github.com/eanson023/mlp.
Mengyuan Liu 0001, Yong Wang 0053, Yang Liu 0264, Hong Liu 0008
IEEE Trans. Circuits Syst. Video Technol.3
2024 Global and Local Spatio-Temporal Encoder for 3D Human Pose Estimation
abstract
Transformers have been used for 3D human pose estimation with excellent performance; however, most transformers focus on encoding the global spatio-temporal correlation of all joints in the human body and there are few studies on the local Spatio-temporal correlation of each joint in the human body. In this article, we propose a Global and Local Spatio-Temporal Encoder (GLSTE) to model the Spatio-temporal correlation. Specifically, a Global Spatial Encoder (GSE) and a Global Temporal Encoder (GTE) are constructed to capture the global spatial information of all joints in a single frame and the global temporal information of all frames, respectively. A Local Spatio-Temporal Encoder (LSTE) is constructed to capture the spatial and temporal information of each joint in the local N frames. Furthermore, we propose a parallel attention module with weight sharing to better incorporate spatial and temporal information into each node simultaneously. Extensive experiments show that GLSTE outperforms state-of-the-art methods with fewer parameters and less computational overhead on two challenging datasets: Human3.6 M and MPI-INF-3DHP. Especially in the evaluation of Human3.6 M dataset, the results of our method with 27 frames as input are better than the vast majority of recent SOTA methods with 81 and 243 frames as input, which indicates that the model can learn more useful information with smaller inputs.
Yong Wang 0053, Hongbo Kang, Doudou Wu, Wenming Yang, Longbin Zhang
IEEE Trans. Multim.1
2023 Edge Intelligence Computing Power Collaboration Framework for Connected Health
abstract
Connected health is a rapidly advancing field that encompasses wireless, digital, mobile, and telehealth technologies. It aims to improve healthcare management by leveraging abundant health data shared by individuals for proactive and efficient care. Edge intelligence (EI), which integrates artificial intelligence (AI) with edge computing, has emerged as a transformative approach to connected health. This paper proposes an EI computing power collaboration framework for connected health. The framework leverages the computing power of edge devices, including consumer electronics, to enhance connected health services. In addition, the framework incorporates edge caching and blockchain technology for efficient healthcare data storage and secure EI computing power collaborations. Deep reinforcement learning, specifically the MuZero algorithm, is used to generate optimal strategies for adjusting the supply-demand relationship of EI computing power. The proposed framework enables responsive and economic healthcare services and empowers various compute-intensive connected health applications.
Boran Yang, Yong Wang 0053
HealthCom2
2023 Dual-channel and multi-granularity gated graph attention network for aspect-based sentiment analysis
Yong Wang 0053, Ningchuang Yang, Duoqian Miao 0001, Qiuyi Chen
Appl. Intell.1
2023 BrightFormer: A transformer to brighten the image
Yong Wang 0053, Bo Li 0115, Xinlin Yuan
Comput. Graph.1
2022 Combining attention mechanism and Retinex model to enhance low-light images
Yong Wang 0053, Yujuan Han, Duoqian Miao 0001
Comput. Graph.1
2022 R2Net: Relight the restored low-light image based on complementarity of illumination and reflection
Yong Wang 0053, Bo Li 0115, Lijun Jiang, Wenming Yang
Signal Process. Image Commun.1