Mengxiao Yin

dblp:154/2524 · also Meng-Xiao Yin · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0001-8327-4813ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Wave-DETR: Real-Time UAV Small Object Detection with Wavelet Feature Fusion and Progressive Query Pruning
Mengxiao Yin, Jiachao Li, Junyuan Huang
ICIC (20)2
2026 DMNet: Dual-Decomposition and Multi-frequency Differentiated Learning Architecture for Long-term Time Series Forecasting
Peizhao Zheng, Zhijie Liang, Fancui Xie, Mengxiao Yin
ICIC (14)4
2026 A dual-branch multi-scale encoding and fusion model for multivariate time series forecasting
Jiachao Li, Mengxiao Yin, Junyuan Huang
Eng. Appl. Artif. Intell.2
2025 AG-Mask: Augmented 3D Generative Masked Motion Model for Text-to-Motion
abstract
Generating natural 3D human motions that are consistent with textual descriptions is a key task in text-to-motion generation. The Transformer-based text-conditional mask motion generation model relies on the multi-head attention mechanism and mask training strategy, making significant progress in generating high-quality and high-fidelity 3D human motions. However, these models only use the multi-head attention mechanism to capture the long-distance dependence of features, lacking the modeling of spatio-temporal relationship of motion sequence, which affects the coherence of the generated motion. In addition, the multi-head attention mechanism has limited ability to capture the detailed information of motions, which affects the authenticity and accuracy of generated motions. Therefore, we propose an augmented 3D generative masked motion model (AG-Mask), which significantly enhances the model’s ability to capture spatiotemporal feature and detailed feature of motion sequence, and effectively generates high-quality motions consistent with text descriptions. Specifically, we design two augmented bidirectional transformer in AG-Mask: STM-Transformer and MDR-Transformer, which are used to process the basic and detailed information of motions respectively. STM-Transformer can boost the information extraction ability of the model in the motion channel and spatial dimension. MDR-Transformer can model the spatio-temporal relationship of motions and extract rich multidimensional features. Their combined processing promotes the generation of motion sequence from coarse to fine, optimizing the quality of generated motions. Experiments on the HumanML3D and KIT-ML datasets show that AG-Mask achieves state-of-the-art performance in generating high-quality motions. In addition, AG-Mask refines the motion editing function, making it more flexible for practical applications.
Zixin Su, Mengxiao Yin, Fancui Xie, Peihong Wu, Bei Hua, Feng Zhan
IJCNN2
2025 MD-Mono: Lightweight Self-Supervised Monocular Depth Estimation Based on Multi-Scale Adaptive Detail Enhancement
abstract
Self-supervised monocular depth estimation has garnered widespread attention because it does not require hard-to-obtain depth labels during training. Many existing studies have focused on the design of depth encoders, often neglecting the potential of decoders, which results in decoders that struggle to utilize the multi-scale features extracted by the encoder fully, lack the ability to capture the features comprehensively, and also fall short in recovering local details. To address these issues, this paper proposes a lightweight self-supervised monocular depth estimation architecture called MD-Mono. MD-Mono employs a hybrid depth encoder combining Convolutional Neural Networks (CNNs) and Transformers, aiming to capture both local features and global semantic information. In the depth decoder, we propose an Adaptive Depth Focus (ADF) module and an Implicit Detail Enhancement (IDE) module. The ADF module adaptively adjusts each stage of the decoding process according to the input features, effectively integrating and utilizing multi-scale features. The IDE module implicitly maps the input to a high-dimensional, nonlinear feature space, capturing more detailed feature information for recovering local details. The synergy of these two modules enables our architecture to achieve a semantically richer and spatially more accurate representation with fewer parameters. Experimental results show that MD-Mono significantly outperforms Monodepth2 in terms of accuracy and exhibits good generalization ability on the Make3D and DrivingStereo datasets.
Peihong Wu, Mengxiao Yin, Pengfei Lai, Zixin Su, Feng Zhan, Bei Hua
IJCNN2
2025 SAGA-Feat: A semantic- and geometry-aware network for sparse local feature learning
Yanhan Mo, Mengxiao Yin, Guiqing Li, Zhijie Liang
Neurocomputing2
2024 GPNF:A Point Cloud Registration Framework Using Sharp Global Linear Attention Prior and Neighborhood Filtering Strategy
Congyang Zhu, Mengxiao Yin, Zhijie Liang, Kan Chang
ACCV (10)2
2024 ALFC-Point: Adaptive Laplacian Feature Convolution Network for 3D Point Cloud Understanding
Mengxiao Yin, Congyang Zhu, Feng Zhan
CGI (3)2
2024 Object and spatial discrimination makes weakly supervised local feature better
Mengxiao Yin, Yunhui Xiong, Pengfei Lai, Kan Chang, Feng Yang 0014
Neural Networks2
2024 PCMG:3D point cloud human motion generation based on self-attention and transformer
Weizhao Ma, Mengxiao Yin, Guiqing Li, Feng Yang 0014, Kan Chang
Vis. Comput.2
2023 SwinFusion: Channel Query-Response Based Feature Fusion for Monocular Depth Estimation
Pengfei Lai, Mengxiao Yin
PRCV (2)2
2023 3D mesh pose transfer based on skeletal deformation
abstract
Abstract For 3D mesh pose transfer, the target model is obtained by transferring the pose of the reference mesh to the source mesh, where the shape and pose of the source are usually different from that of the reference. In this paper, pose transfer is considered as a deformation process of the source mesh, and we propose a 3D mesh pose transfer method based on skeletal deformation. First, we design a neural network based on the edge convolution operator to extract the skeleton of the 3D mesh and bind the rigid weights; then, we calculate the bone transformations between the two skeletons with different poses and use the diffusion equation to smooth the rigid weights; finally, the source mesh is deformed according to the bone transformations and the smooth weights to get the target mesh. Experiment results on different datasets show that the pose of the reference mesh can be effectively transferred to the source one while maintaining the shape and high‐quality geometric details of the source mesh by using our method.
Shigeng Yang, Mengxiao Yin, Guiqing Li, Kan Chang, Feng Yang 0014
Comput. Animat. Virtual Worlds2
2022 A Self-improving Skin Lesions Diagnosis Framework Via Pseudo-labeling and Self-distillation
Shaochang Deng, Mengxiao Yin, Feng Yang 0014
ACML2
2021 Three-view generation based on a single front view image for car
Zixuan Qin, Mengxiao Yin, Zhenfeng Lin, Feng Yang 0014
Vis. Comput.2
2020 SP-Flow: Self-supervised optical flow correspondence point prediction for real-time SLAM
Zixuan Qin, Mengxiao Yin, Guiqing Li, Feng Yang 0014
Comput. Aided Geom. Des.2
2015 Spectral pose transfer
Mengxiao Yin, Guiqing Li, Huina Lu, Yaobin Ouyang, Zhibang Zhang, Chuhua Xian
Comput. Aided Geom. Des.1
2015 EC-CageR: Error controllable cage reverse for animated meshes
Huina Lu, Guiqing Li, Chuhua Xian, Zhibang Zhang, Mengxiao Yin
Comput. Graph.5
2015 Fast as-isometric-as-possible shape interpolation
Zhibang Zhang, Guiqing Li, Huina Lu, Yaobin Ouyang, Mengxiao Yin, Chuhua Xian
Comput. Graph.5