VLDB 2026 Research / reviewers in the wild / expert
Mengxiao Yin
dblp:154/2524 · also Meng-Xiao Yin
· DBLP profile ↗
18ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0001-8327-4813ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Wave-DETR: Real-Time UAV Small Object Detection with Wavelet Feature Fusion and Progressive Query Pruning
Mengxiao Yin, Jiachao Li, Junyuan Huang |
ICIC (20) | 2 |
| 2026 | DMNet: Dual-Decomposition and Multi-frequency Differentiated Learning Architecture for Long-term Time Series Forecasting
Peizhao Zheng, Zhijie Liang, Fancui Xie, Mengxiao Yin |
ICIC (14) | 4 |
| 2026 | A dual-branch multi-scale encoding and fusion model for multivariate time series forecasting
Jiachao Li, Mengxiao Yin, Junyuan Huang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | AG-Mask: Augmented 3D Generative Masked Motion Model for Text-to-MotionabstractGenerating natural 3D human motions that are consistent with textual descriptions is a key task in text-to-motion generation. The Transformer-based text-conditional mask motion generation model relies on the multi-head attention mechanism and mask training strategy, making significant progress in generating high-quality and high-fidelity 3D human motions. However, these models only use the multi-head attention mechanism to capture the long-distance dependence of features, lacking the modeling of spatio-temporal relationship of motion sequence, which affects the coherence of the generated motion. In addition, the multi-head attention mechanism has limited ability to capture the detailed information of motions, which affects the authenticity and accuracy of generated motions. Therefore, we propose an augmented 3D generative masked motion model (AG-Mask), which significantly enhances the model’s ability to capture spatiotemporal feature and detailed feature of motion sequence, and effectively generates high-quality motions consistent with text descriptions. Specifically, we design two augmented bidirectional transformer in AG-Mask: STM-Transformer and MDR-Transformer, which are used to process the basic and detailed information of motions respectively. STM-Transformer can boost the information extraction ability of the model in the motion channel and spatial dimension. MDR-Transformer can model the spatio-temporal relationship of motions and extract rich multidimensional features. Their combined processing promotes the generation of motion sequence from coarse to fine, optimizing the quality of generated motions. Experiments on the HumanML3D and KIT-ML datasets show that AG-Mask achieves state-of-the-art performance in generating high-quality motions. In addition, AG-Mask refines the motion editing function, making it more flexible for practical applications. Zixin Su, Mengxiao Yin, Fancui Xie, Peihong Wu, Bei Hua, Feng Zhan |
IJCNN | 2 |
| 2025 | MD-Mono: Lightweight Self-Supervised Monocular Depth Estimation Based on Multi-Scale Adaptive Detail EnhancementabstractSelf-supervised monocular depth estimation has garnered widespread attention because it does not require hard-to-obtain depth labels during training. Many existing studies have focused on the design of depth encoders, often neglecting the potential of decoders, which results in decoders that struggle to utilize the multi-scale features extracted by the encoder fully, lack the ability to capture the features comprehensively, and also fall short in recovering local details. To address these issues, this paper proposes a lightweight self-supervised monocular depth estimation architecture called MD-Mono. MD-Mono employs a hybrid depth encoder combining Convolutional Neural Networks (CNNs) and Transformers, aiming to capture both local features and global semantic information. In the depth decoder, we propose an Adaptive Depth Focus (ADF) module and an Implicit Detail Enhancement (IDE) module. The ADF module adaptively adjusts each stage of the decoding process according to the input features, effectively integrating and utilizing multi-scale features. The IDE module implicitly maps the input to a high-dimensional, nonlinear feature space, capturing more detailed feature information for recovering local details. The synergy of these two modules enables our architecture to achieve a semantically richer and spatially more accurate representation with fewer parameters. Experimental results show that MD-Mono significantly outperforms Monodepth2 in terms of accuracy and exhibits good generalization ability on the Make3D and DrivingStereo datasets. Peihong Wu, Mengxiao Yin, Pengfei Lai, Zixin Su, Feng Zhan, Bei Hua |
IJCNN | 2 |
| 2025 | SAGA-Feat: A semantic- and geometry-aware network for sparse local feature learning
Yanhan Mo, Mengxiao Yin, Guiqing Li, Zhijie Liang |
Neurocomputing | 2 |
| 2024 | GPNF:A Point Cloud Registration Framework Using Sharp Global Linear Attention Prior and Neighborhood Filtering Strategy
Congyang Zhu, Mengxiao Yin, Zhijie Liang, Kan Chang |
ACCV (10) | 2 |
| 2024 | ALFC-Point: Adaptive Laplacian Feature Convolution Network for 3D Point Cloud Understanding
Mengxiao Yin, Congyang Zhu, Feng Zhan |
CGI (3) | 2 |
| 2024 | Object and spatial discrimination makes weakly supervised local feature better
Mengxiao Yin, Yunhui Xiong, Pengfei Lai, Kan Chang, Feng Yang 0014 |
Neural Networks | 2 |
| 2024 | PCMG:3D point cloud human motion generation based on self-attention and transformer
Weizhao Ma, Mengxiao Yin, Guiqing Li, Feng Yang 0014, Kan Chang |
Vis. Comput. | 2 |
| 2023 | SwinFusion: Channel Query-Response Based Feature Fusion for Monocular Depth Estimation
Pengfei Lai, Mengxiao Yin |
PRCV (2) | 2 |
| 2023 | 3D mesh pose transfer based on skeletal deformationabstractAbstract For 3D mesh pose transfer, the target model is obtained by transferring the pose of the reference mesh to the source mesh, where the shape and pose of the source are usually different from that of the reference. In this paper, pose transfer is considered as a deformation process of the source mesh, and we propose a 3D mesh pose transfer method based on skeletal deformation. First, we design a neural network based on the edge convolution operator to extract the skeleton of the 3D mesh and bind the rigid weights; then, we calculate the bone transformations between the two skeletons with different poses and use the diffusion equation to smooth the rigid weights; finally, the source mesh is deformed according to the bone transformations and the smooth weights to get the target mesh. Experiment results on different datasets show that the pose of the reference mesh can be effectively transferred to the source one while maintaining the shape and high‐quality geometric details of the source mesh by using our method. Shigeng Yang, Mengxiao Yin, Guiqing Li, Kan Chang, Feng Yang 0014 |
Comput. Animat. Virtual Worlds | 2 |
| 2022 | A Self-improving Skin Lesions Diagnosis Framework Via Pseudo-labeling and Self-distillation
Shaochang Deng, Mengxiao Yin, Feng Yang 0014 |
ACML | 2 |
| 2021 | Three-view generation based on a single front view image for car
Zixuan Qin, Mengxiao Yin, Zhenfeng Lin, Feng Yang 0014 |
Vis. Comput. | 2 |
| 2020 | SP-Flow: Self-supervised optical flow correspondence point prediction for real-time SLAM
Zixuan Qin, Mengxiao Yin, Guiqing Li, Feng Yang 0014 |
Comput. Aided Geom. Des. | 2 |
| 2015 | Spectral pose transfer
Mengxiao Yin, Guiqing Li, Huina Lu, Yaobin Ouyang, Zhibang Zhang, Chuhua Xian |
Comput. Aided Geom. Des. | 1 |
| 2015 | EC-CageR: Error controllable cage reverse for animated meshes
Huina Lu, Guiqing Li, Chuhua Xian, Zhibang Zhang, Mengxiao Yin |
Comput. Graph. | 5 |
| 2015 | Fast as-isometric-as-possible shape interpolation
Zhibang Zhang, Guiqing Li, Huina Lu, Yaobin Ouyang, Mengxiao Yin, Chuhua Xian |
Comput. Graph. | 5 |