Dong Wang 0043

dblp:40/3934-43 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-9457-263XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MLENet: Multi-level efficient network based on single-scale feature extraction for human keypoint estimation
Dong Wang 0043, Youcheng Cai, Yiming Tang 0001, Wenjun Xie, Xiaoping Liu 0003
Expert Syst. Appl.1
2026 MatPose: A 2D Human Pose Estimation Model with Hybrid Mamba-Transformer
abstract
Recently, Mamba has gained widespread attention due to its ability to model long-range dependencies with linear computational complexity. To explore the application of Mamba in 2D human pose estimation, we propose MatPose, a Mamba-Transformer hybrid model specifically designed for efficient 2D human pose estimation. The model aims to combine Mamba’s efficient modeling of long-range dependencies with the powerful global context modeling capabilities of the Transformer to effectively extract human pose keypoints. Initially, to address the lack of local features when Mamba is applied to computer vision tasks, we design a Cross-Stage Multi-Scale Convolution (CSMSC) module by integrating multi-scale convolution, cross-stage feature fusion, and spatial attention mechanisms to effectively extract local features. Then, to mitigate the long-range forgetting issue inherent in Mamba, we shorten the sequence length using the Conv-Reduce operation. In addition, we design a Channel Selection Attention (CSA) mechanism to compensate for the feature loss caused by the Conv-Reduce operation. Finally, to explore a suitable integration method for the Mamba-Transformer hybrid model in 2D human pose estimation, we conduct a comprehensive ablation study on the feasibility of integrating Mamba and Transformer models. Experimental results show that the proposed method, compared to the baseline model, improves performance while reducing computational overhead. On the COCO val2017 dataset, MatPose achieves an AP of 74.6 with only 5.18 GFLOPs, outperforming most existing human pose estimation models.
Wenjun Xie, Kejun Chen, Dong Wang 0043, Xiaoping Liu 0003
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Audio2AB: Audio-driven collaborative generation of virtual character animation
abstract
Considerable research has been conducted in the areas of audio-driven virtual character gestures and facial animation with some degree of success. However, few methods exist for generating full-body animations, and the portability of virtual character gestures and facial animations has not received sufficient attention. Therefore, we propose a deep-learning-based audio-to-animation-and-blendshape (Audio2AB) network that generates gesture animations andARK it’s 52 facial expression parameter blendshape weights based on audio, audio-corresponding text, emotion labels, and semantic relevance labels to generate parametric data for full- body animations. This parameterization method can be used to drive full-body animations of virtual characters and improve their portability. In the experiment, we first downsampled the gesture and facial data to achieve the same temporal resolution for the input, output, and facial data. The Audio2AB network then encoded the audio, audio- corresponding text, emotion labels, and semantic relevance labels, and then fused the text, emotion labels, and semantic relevance labels into the audio to obtain better audio features. Finally, we established links between the body, gestures, and facial decoders and generated the corresponding animation sequences through our proposed GAN-GF loss function. By using audio, audio-corresponding text, and emotional and semantic relevance labels as input, the trained Audio2AB network could generate gesture animation data containing blendshape weights. Therefore, different 3D virtual character animations could be created through parameterization. The experimental results showed that the proposed method could generate significant gestures and facial animations.
Lichao Niu, Wenjun Xie, Dong Wang 0043, Zhongrui Cao, Xiaoping Liu 0003
Virtual Real. Intell. Hardw.3
2023 MFNet: Multi-level fusion aware feature pyramid based multi-view stereo network for 3D reconstruction
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xiaoping Liu 0003
Appl. Intell.3
2023 Transformer-based rapid human pose estimation network
Dong Wang 0043, Wenjun Xie, Youcheng Cai, Xinjie Li 0006, Xiaoping Liu 0003
Comput. Graph.1
2023 HTMatch: An efficient hybrid transformer based graph neural network for local feature matching
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xinjie Li 0006, Xiaoping Liu 0003
Signal Process.3
2023 GlcMatch: global and local constraints for reliable feature matching
Youcheng Cai, Lin Li 0053, Dong Wang 0043, Xintao Huang, Xiaoping Liu 0003
Vis. Comput.3
2022 A Fast and Effective Transformer for Human Pose Estimation
abstract
Most of the existing human pose estimation methods improve accuracy by constantly increasing computational resources. However, balancing the efficiency and efficacy of the model is the key to enhancing the real application value. In this work, we present a Fast and Effective Transformer model to ensure the efficiency and efficacy of the model, called FET. Specifically, the FET consists of three parts: Feature Extraction Module (FEM), Feature Interaction Module (FIM) and Feature Decode Module (FDM). The FEM is used to efficiently extract low-level features from input images. Unlike CNN-based strategies, the FIM enables our model to capture global dependencies by self-attention, thus improving the accuracy for human pose estimation. The FDM is a multistage way that gradually recovers the size of the features to obtain a higher-quality target heatmap. In addition, Feature Squeeze Attention is introduced in the FET to further improve the overall performance of our model. Extensive experiments show that our method is 1.7× and 7× faster than SimpleBaseline and HRNet-32, respectively, while achieving comparable or even better results with the most state-of-the-art methods on the COCO dataset and the MPII dataset.
Dong Wang 0043, Wenjun Xie, Youcheng Cai, Xiaoping Liu 0003
IEEE Signal Process. Lett.1