Wenbin Wang 0001

dblp:27/3361-1 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-4394-0145ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Salience Feature Guided Decoupling Network for UAV Forests Flame Detection
Zerui Wang, Li Liu 0059, Wenbin Wang 0001
Expert Syst. Appl.5
2025 Bidirectional-Modulation Frequency-Heterogeneous Network for Remote Sensing Image Dehazing
abstract
Recently, deep neural networks have been extensively explored in remote sensing image haze removal and achieved remarkable performance. However, existing methods fail to effectively fuse the features extracted from Convolutional Neural Networks (CNNs) and Transformer networks, leading to performance degradation. Moreover, most dehazing methods lack further exploration of the distinct properties of high- and low-frequency features, which are crucial for texture restoration and haze removal. To address these issues, we propose a Bidirectional-Modulation Frequency-Heterogeneous Network (BMFH-Net). Specifically, we propose a Differential-Expert Guided Bidirectional Modulation (DGBM) module that incorporates Differential experts and physical inversion models to exploit the complementarity of CNN-Transformer features and extract their latent haze-related physical characteristics, thereby enabling more effective bidirectional alignment. Furthermore, a Wavelet Frequency Heterogeneous Enhancement (WFHE) Module is designed to capture the most representative high-frequency features to refine image texture details, while enhancing the global perception of haze and reconstructing structural information during low-frequency processing. Experiments on challenging remote sensing image datasets demonstrate that our BMFH-Net outperforms several state-of-the-art haze removal methods. The code is released publicly at https://github.com/zqf2024/BMFH-Net.
Qingfei Zhong, Bo Du 0001, Zhigang Tu 0001, Jun Wan 0005, Wenbin Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 HOIEdit: Human-object interaction editing with text-to-image diffusion model
Tang Xu, Wenbin Wang 0001, Alin Zhong
Vis. Comput.2
2024 Towards Robust Multimodal Sentiment Analysis with Incomplete Data
abstract
The field of Multimodal Sentiment Analysis (MSA) has recently witnessed an emerging direction seeking to tackle the issue of data incompleteness. Recognizing that the language modality typically contains dense sentiment information, we consider it as the dominant modality and present an innovative Language-dominated Noise-resistant Learning Network (LNLN) to achieve robust MSA. The proposed LNLN features a dominant modality correction (DMC) module and dominant modality based multimodal learning (DMML) module, which enhances the model's robustness across various noise scenarios by ensuring the quality of dominant modality representations. Aside from the methodical design, we perform comprehensive experiments under random data missing scenarios, utilizing diverse and meaningful settings on several popular datasets (e.g., MOSI, MOSEI, and SIMS), providing additional uniformity, transparency, and fairness compared to existing evaluations in the literature. Empirically, LNLN consistently outperforms existing baselines, demonstrating superior performance across these challenging and extensive evaluation metrics.
Haoyu Zhang 0001, Wenbin Wang 0001, Tianshu Yu 0001
NeurIPS2
2023 Pose-disentangled Contrastive Learning for Self-supervised Facial Representation
abstract
Self-supervised facial representation has recently attracted increasing attention due to its ability to perform face understanding without relying on large-scale annotated datasets heavily. However, analytically, current contrastive-based self-supervised learning (SSL) still performs unsatisfactorily for learning facial representation. More specifically, existing contrastive learning (CL) tends to learn pose-invariant features that cannot depict the pose details of faces, compromising the learning performance. To conquer the above limitation of CL, we propose a novel Pose-disentangled Contrastive Learning (PCL) method for general self-supervised facial representation. Our PCL first devises a pose-disentangled decoder (PDD) with a delicately designed orthogonalizing regulation, which disentangles the pose-related features from the face-aware features; therefore, pose-related and other pose-unrelated facial information could be performed in individual subnetworks and do not affect each other's training. Furthermore, we introduce a pose-related contrastive learning scheme that learns pose-related information based on data augmentation of the same image, which would deliver more effective face-aware representation for various downstream tasks. We conducted linear evaluation on four challenging downstream facial understanding tasks, i.e., facial expression recognition, face recognition, AU detection and head pose estimation. Experimental results demonstrate that PCL significantly outperforms cuttingedge SSL methods. Our Code is available at https://github.com/DreamMr/PCL.
Yuanyuan Liu 0004, Wenbin Wang 0001, Yibing Zhan, Shaoze Feng, Kejun Liu, Zhe Chen 0013
CVPR2
2023 Importance First: Generating Scene Graph of Human Interest
Wenbin Wang 0001, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001
Int. J. Comput. Vis.1
2023 APSL: Action-positive separation learning for unsupervised temporal action localization
Yuanyuan Liu 0004, Fayong Zhang, Wenbin Wang 0001, Yu Wang 0246, Kejun Liu, Ziyuan Liu 0005
Inf. Sci.4
2023 Expression snippet transformer for robust video-based facial expression recognition
Yuanyuan Liu 0004, Wenbin Wang 0001, Chuanxu Feng, Haoyu Zhang 0001, Zhe Chen 0013, Yibing Zhan
Pattern Recognit.2
2022 MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild
abstract
Dynamic facial expression recognition (FER) databases provide important data support for affective computing and applications. However, most FER databases are annotated with several basic mutually exclusive emotional categories and contain only one modality, e.g., videos. The monotonous labels and modality cannot accurately imitate human emotions and fulfill applications in the real world. In this paper, we propose MAFW, a large-scale multi-modal compound affective database with 10,045 video-audio clips in the wild. Each clip is annotated with a compound emotional category and a couple of sentences that describe the subjects' affective behaviors in the clip. For the compound emotion annotation, each clip is categorized into one or more of the 11 widely-used emotions, i.e., anger, disgust, fear, happiness, neutral, sadness, surprise, contempt, anxiety, helplessness, and disappointment. To ensure high quality of the labels, we filter out the unreliable annotations by an Expectation Maximization (EM) algorithm, and then obtain 11 single-label emotion categories and 32 multi-label emotion categories. To the best of our knowledge, MAFW is the first in-the-wild multi-modal database annotated with compound emotion annotations and emotion-related captions. Additionally, we also propose a novel Transformer-based expression snippet feature learning method to recognize the compound emotions leveraging the expression-change relations among different emotions and modalities. Extensive experiments on MAFW database show the advantages of the proposed method over other state-of-the-art methods for both uni- and multi-modal FER. Our MAFW database is publicly available from https://mafw-database.github.io/MAFW.
Yuanyuan Liu 0004, Chuanxu Feng, Wenbin Wang 0001, Guanghao Yin, Jiabei Zeng, Shiguang Shan
ACM Multimedia4
2022 Clip-aware expressive feature learning for video-based facial expression recognition
Yuanyuan Liu 0004, Chuanxu Feng, Xiaohui Yuan 0001, Lin Zhou 0017, Wenbin Wang 0001, Zhongwen Luo
Inf. Sci.5
2021 Topic Scene Graph Generation by Attention Distillation from Caption
abstract
If an image tells a story, the image caption is the briefest narrator. Generally, a scene graph prefers to be an omniscient "generalist", while the image caption is more willing to be a "specialist", which outlines the gist. Lots of previous studies have found that a scene graph is not as practical as expected unless it can reduce the trivial contents and noises. In this respect, the image caption is a good tutor. To this end, we let the scene graph borrow the ability from the image caption so that it can be a specialist on the basis of remaining all-around, resulting in the socalled Topic Scene Graph. What an image caption pays attention to is distilled and passed to the scene graph for estimating the importance of partial objects, relationships, and events. Specifically, during the caption generation, the attention about individual objects in each time step is collected, pooled, and assembled to obtain the attention about relationships, which serves as weak supervision for regularizing the estimated importance scores of relationships. In addition, as this attention distillation process provides an opportunity for combining the generation of image caption and scene graph together, we further transform the scene graph into linguistic form with rich and free-form expressions by sharing a single generation model with image caption. Experiments show that attention distillation brings significant improvements in mining important relationships without strong supervision, and the topic scene graph shows great potential in subsequent applications.
Wenbin Wang 0001, Ruiping Wang 0001, Xilin Chen 0001
ICCV1
2020 Sketching Image Gist: Human-Mimetic Hierarchical Scene Graph Generation
Wenbin Wang 0001, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001
ECCV (13)1
2019 Exploring Context and Visual Pattern of Relationship for Scene Graph Generation
abstract
Relationship is the core of scene graph, but its prediction is far from satisfying because of its complex visual diversity. To alleviate this problem, we treat relationship as an abstract object, exploring not only significative visual pattern but contextual information for it, which are two key aspects when considering object recognition. Our observation on current datasets reveals that there exists intimate association among relationships. Therefore, inspired by the successful application of context to object-oriented tasks, we especially construct context for relationships where all of them are gathered so that the recognition could benefit from their association. Moreover, accurate recognition needs discriminative visual pattern for object, and so does relationship. In order to discover effective pattern for relationship, traditional relationship feature extraction methods such as using union region or combination of subject-object feature pairs are replaced with our proposed intersection region which focuses on more essential parts. Therefore, we present our so-called Relationship Context - InterSeCtion Region (CISC) method. Experiments for scene graph generation on Visual Genome dataset and visual relationship prediction on VRD dataset indicate that both the relationship context and intersection region improve performances and realize anticipated functions.
Wenbin Wang 0001, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001
CVPR1