VLDB 2026 Research / reviewers in the wild / expert
Mengyin Wang
dblp:172/4647
· DBLP profile ↗
22ranked-venue papers
3as first author
19since 2021 · last 2026
0009-0001-1985-7026ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DGD-CCAFNet: Dual-level graph distillation and continuous cross-modal attention fusion network for multimodal sentiment analysis
Houjie Li, Fenglin Li 0002, Mengyin Wang |
Expert Syst. Appl. | 5 |
| 2026 | Wild Animal Tracking with High-Quality Segment Anything Model and Domain Adaptation
Ganggang Huang, Fasheng Wang, Hanwei Li, Mingshu Zhang, Mengyin Wang, Fuming Sun |
Int. J. Comput. Vis. | 6 |
| 2026 | HGLFFNet: Hierarchical global-local feature fusion network for facial expression recognition
Houjie Li, Fuming Sun, Mengyin Wang |
Neurocomputing | 6 |
| 2026 | Modality Interaction Decoupling: Spatiotemporal Memory-Driven Unified Multimodal Tracking FrameworkabstractCurrent designs of multimodal tracking networks primarily focus on spatial feature interaction within the backbone and lack the exploitation of temporal information. Although some approaches incorporate temporal cues by introducing sequential information from adjacent frames or employ updated temporal features during the feature extraction stage, they struggle to capture dynamic object variations and motion information in complex scenarios. To address these limitations, this paper proposes a Spatiotemporal Memory-Driven Unified Multimodal Tracking Framework (SMMTrack). Unlike existing works relying on feature interaction paradigms within the backbone network, this study innovatively introduces a decoupled backbone feature extraction framework. It deploys parallel, independent Vision Transformer (ViT) networks dedicated to extracting information from RGB and X modalities (RGB-T, RGB-D, RGB-E), abandoning the conventional intra-backbone feature interaction. Furthermore, we introduce a memory mechanism during the feature extraction stage to enable long-term object modeling. In addition, a Long-term Memory storage and retrieval module is designed to dynamically update the Memory-list, thereby allowing the model to capture object appearance variations and motion trends comprehensively. SMMTrack is a unified framework across three tasks (RGB-T, RGB-D, and RGB-E tracking). Experimental results demonstrate that SMMTrack outperforms the state-of-the-art (SOTA) models, achieving outstanding performance in diverse multimodal tracking scenarios. Codes and results are released on https://github.com/qfxb/SMMTrack. Fasheng Wang, Mengyin Wang, Fuming Sun |
IEEE Internet Things J. | 4 |
| 2026 | SPMNet: Self-prompt mask-guided network for Camouflaged Object Detection
Mengyin Wang, Fenglin Li 0002, Houjie Li |
Knowl. Based Syst. | 3 |
| 2026 | Rethinking RGB-D salient object detection
Chang Kou, Jinyu Han, Mengyin Wang |
Multim. Syst. | 3 |
| 2026 | Progressive edge-aware multi-scale fusion network for camouflaged object detection
Jinqi Wang, Mengyin Wang |
Multim. Syst. | 2 |
| 2026 | CERNet: Real-time stereo matching via collaborative enhancement and refinement
Zhisheng Zhu, Mengyin Wang, Fuming Sun |
Pattern Recognit. | 2 |
| 2025 | Knowledge-guided and Collaborative Learning Network for Camouflaged Object Detection
Mengyin Wang, Jing Sun 0012 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Perceptual localization and focus refinement network for RGB-D salient object detection
Jinyu Han, Mengyin Wang, Weiyi Wu |
Expert Syst. Appl. | 2 |
| 2025 | Asymmetric cross-modality interaction network for RGB-D salient object detection
Yiming Su, Mengyin Wang, Fasheng Wang |
Expert Syst. Appl. | 3 |
| 2025 | Highly Efficient RGB-D Salient Object Detection With Adaptive Fusion and Attention RegulationabstractExisting RGB-D salient object detection (SOD) models have large numbers of parameters, high computational complexity, and slow inference speeds, limiting their deployment on edge devices. To address this issue, we propose a highly efficient network (HENet), focusing on developing lightweight RGB-D SOD models. Specifically, to fairly handle multimodal inputs and capture long-range dependencies of features, we employ a dual-stream structure and use MobileViT as the network encoder. We introduce the Adaptive Edge-Aware Fusion Module (AEFM) that adaptively adjusts the contribution of features during the fusion process based on the amount of feature information, and perceives the edges of the fused features at the pixel level. To compensate for the insufficient feature extraction capability of the lightweight backbone network, we propose the Dual-Branch Feature Enhancement Module (DFEM) to enhance the representation capability of the fused features. Finally, we design the Feature Attention Regulation Module (FARM) to adjust the model’s focus in real time. HENet has fewer parameters (11.9M) and lower computational complexity (10.7 GFLOPs), achieving an inference speed of 121 FPS for images with size$384\times 384$. Extensive experiments are conducted on seven challenging RGB-D SOD datasets. The experimental results demonstrate that HENet outperforms 16 state-of-the-art methods and shows great potential in downstream computer vision tasks. Codes and results are available onhttps://github.com/BojueGao/HENet. Fasheng Wang, Mengyin Wang, Fuming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Rethinking How to Capture Long-Range Dependency in 3D Object DetectionabstractLiDAR-based 3D object detection is essential for autonomous driving. Existing high-performance 3D object detectors usually design complex structures in the 3D backbone to capture long-range dependencies among features. However, introducing these complex structures into the 3D backbone significantly increases computational cost and inference latency, limiting the efficiency and feasibility of detectors in practical applications. In this work, we rethink the long-range dependency capturing problem from a new perspective, that is transferring this task from 3D backbone to 2D feature space. To accomplish this goal, we propose a Long-Range Dense Feature Capture Network (LDFCNet). LDFCNet retains the basic structure of the 3D backbone to extract preliminary 3D features but shifts the complex long-range dependency capturing task to be processed on a 2D dense feature map, thereby enhancing the detection performance while reducing the computational cost. Importantly, a robust 2D dense feature capture (2D-DFC) backbone is devised to effectively and efficiently capture the long-range dependencies. In addition, we introduce a re-parameterization technique to decouple the training and inference of the 2D backbone, further reducing inference latency. We conduct extensive experiments on the Waymo Open and nuScenes datasets and the experimental results show that LDFCNet demonstrates competitive performance. Notably, LDFCNet is$1.5\times $faster than the state-of-the-art hybrid detector HEDNet and$2.1\times $faster than the transformer-based detector DSVT. Codes and results are released onhttps://github.com/asd291614761/LDFCNet. Fasheng Wang, Mengyin Wang, Fuming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Lightweight Edge-Aware Mamba-Fusion Network for Weakly Supervised Salient Object Detection in Optical Remote Sensing ImagesabstractDespite the significant progress made in fully supervised salient object detection in optical remote sensing images (ORSI-SOD), these methods rely heavily on pixel-level annotations, which are time-consuming and labor-intensive. This situation has driven the development of weakly supervised ORSI-SOD methods. However, existing weakly supervised ORSI-SOD methods still face excessive model parameters and high computational complexity, hindering their flexibility and deployment in edge devices. To address these challenges, we propose the LightEMNet, a scribble-based, lightweight, and high-performance edge-aware network for ORSI-SOD. The network employs MobileNetV2 as its lightweight encoder backbone. To mitigate the suboptimal feature extraction performance caused by the lightweight architecture, we design a feature refinement layer (FRL) to refine the features extracted from the backbone, thereby generating guidance information while achieving better structural awareness and object localization. To realize better detail optimization, we introduce edge information extracted by a multiscale edge perception module (MEP) to regulate high-level features. Finally, considering the shortcomings of traditional convolution in global-awareness, we propose a Mamba-based cross-scale edge-semantic interaction (CESI) module to achieve efficient alignment of semantics and edges, which consequently enhances the representation consistency of the fused features and improves the model’s adaptability to complex scenes. We verify the effectiveness of the LightEMNet through extensive experiments. The results demonstrate that the proposed LightEMNet exhibits competitive detection performance with only 4.81 M parameters. Codes and results are available athttps://github.com/xingggao/LightEMNet Gaojie Xing, Mengyin Wang, Fasheng Wang, Fuming Sun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | A UNet-Like Transformer Network for Camouflaged Object DetectionabstractThe role of Camouflaged Object Detection (COD) is to identify the objects that integrate seamlessly with the surrounding environment. Due to the high intrinsic similarity between the objects and their background, this task presents greater challenges than traditional object detection. Most existing COD methods often have a large number of parameters and high computational complexity in the pursuit of detection accuracy, which hinders the application of COD in practical scenarios. To address this issue, we propose a UNet-like Transformer Network for COD, termed UTNet, which achieves competitive detection accuracy with a smaller parameter set. Specifically, we propose a Camouflaged Region Awareness Module (CRAM) consisting of a Hierarchical Attention Mechanism (HAM) that groups features to reveal intrinsic consistency between sub-features. This CRAM can be embedded into the backbone network, giving it powerful modeling capabilities. And, we present a Contextual Knowledge Collector (CKC) that exploits a cross-aggregation approach for neighboring feature layers, promoting the flow of semantic information from high-level to low-level features, and ensuring the integrity of camouflaged objects at each level of features. Furthermore, we introduce a progressive decoder that utilizes a cascade of attention units to filter noise and explores knowledge aggregation to emphasize features from different levels, ensuring that camouflaged objects have complete spatial details at the local level. Extensive experimental results show that UTNet achieves competitive results compared to 20 state-of-the-art methods. Codes and results are released onhttps://github.com/hjy0518/UTNet. Fuming Sun, Jinyu Han, Weiyi Wu, Jing Sun 0012, Mengyin Wang |
IEEE Trans. Multim. | 5 |
| 2025 | Spatial-Frequency Collaborative Learning for Camouflaged Object DetectionabstractCamouflaged object detection (COD) is a challenging task that struggles to accurately detect the objects concealed in the surrounding environment. This is largely attributed to the intrinsic similarity of the camouflaged objects with the surrounding environment. To address this challenge, we propose a Spatial-Frequency Collaborative Learning network for COD (SFCNet). Specifically, we propose a Domain Transformation Fusion (DTF) module to handle the similarity between the camouflaged objects and the background, because when processed in the frequency domain, the features of the camouflaged object and the background become easy to discriminate. Then, we design a Cross-domain Integration Unit (CIU) to integrate the high-level features progressively through a Spatial-Frequency Coordinated Fusion (SFCF) module and a Multi-scale Feature Enhancement (MFE) module. Finally, the low-level features are combined with the high-level features from different decoding stages to correct the camouflaged objects in detail. In addition, an Edge Amplification (EA) module is designed to enable the model to pay attention to the global contour of the camouflaged object. It can facilitate the generation of prediction maps with accurate object boundaries. Extensive experiments on four benchmark COD datasets show that SFCNet outperforms state-of-the-art (SOTA) COD models. Meanwhile, it also has the characteristics of low parameters (21.01 M) and low computational complexity (24.14 G). Codes and results are released onhttps://github.com/Zhaorui328/SFCNet. Mengyin Wang, Fasheng Wang, Fuming Sun |
IEEE Trans. Multim. | 2 |
| 2024 | OmniStyleGAN for Style-Guided Image-to-Image Translation
Qianyi Zhao, Mengyin Wang, Fasheng Wang, Fuming Sun |
PRCV (11) | 2 |
| 2024 | Unsupervised image-to-image translation with multiscale attention generative adversarial network
Fasheng Wang, Qianyi Zhao, Mengyin Wang, Fuming Sun |
Appl. Intell. | 4 |
| 2024 | OR2Net: Online Re-weighting Relation Network for kinship verificationabstractKinship verification aims to infer whether there is a kin relation between different individuals from facial images . However, popular kinship datasets are often small and suffer from data imbalance. Most existing methods build complex networks to extract features but ignore some implicit information, like family information. They use balanced datasets with fixed negative samples for training, which overlooks valuable information from multiple negative samples, leading to poor performance and robustness. To address these issues, we propose a novel end-to-end framework for kinship verification called Online Re-weighting Relation Network (OR 2 Net) based on an online re-weighting strategy of meta-learning and relation network. Our novel relation network aims to extract fine-grained features and reduce differences between generations by using multi-scale features and mining family information from kinship datasets. Additionally, we design a lightweight meta re-weighting network that uses a small, clean meta-set to guide the adaptive weighing of training examples. This is done using one-step stochastic gradient descent (SGD) based on an online re-weighting strategy from meta-learning. This helps find effective hard negative samples and reduces the imbalance problem. Extensive experiments on three public kinship verification datasets show that our proposed method is more effective compared to state-of-the-art methods. The code is publicly available on https://github.com/XinZhao-dlnu/OR2N . Houjie Li, Mengyin Wang, Haiyu Song 0002, Fuming Sun |
Expert Syst. Appl. | 3 |
| 2020 | Deep multi-person kinship matching and recognition for family photos
Mengyin Wang, Xiangbo Shu, Jiashi Feng, Xun Wang 0007, Jinhui Tang 0001 |
Pattern Recognit. | 1 |
| 2020 | Deep supervised feature selection for social relationship recognition
Mengyin Wang, Xiaoyu Du 0002, Xiangbo Shu, Xun Wang 0007, Jinhui Tang 0001 |
Pattern Recognit. Lett. | 1 |
| 2015 | Deep kinship verificationabstractTo improve the performance of kinship verification, we propose a novel deep kinship verification (DKV) model by integrating excellent deep learning architecture into metric learning. Unlike most existing shallow models based on metric learning for kinship verification, we employ a deep learning model followed by a metric learning formulation to select nonlinear features, which can find the appropriate project space to ensure the margin of negative sample pairs (i.e. parent and child without kinship relation) as large as possible and the margin of positive sample pairs (i.e. parent and child with kinship relation) as small as possible. Experimental results show that our method achieves satisfactory performance on two widely-used benchmarks, i.e. KFW-I and KFW-II. Mengyin Wang, Zechao Li, Xiangbo Shu, Jingdong Wang 0001, Jinhui Tang 0001 |
MMSP | 1 |