Mengzhu Wang

dblp:279/5316 · DBLP profile ↗
← Back
78ranked-venue papers
18as first author
78since 2021 · last 2026
0000-0002-1059-0441ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 8 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 9 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning to Cluster Rare Cell Types: Implicit Semantic Data Augmentation for Spatial Multi-modal Omics Analysis
abstract
Spatial multi-modal omics technologies have transformed biological research by enabling the simultaneous profiling of gene expression, protein abundance, and chromatin accessibility within their native spatial contexts. Despite these advances, accurately clustering rare cell types remains a major challenge due to data sparsity, high dimensionality, and limited annotated samples. While Graph Neural Networks (GNNs) have shown potential in modeling spatial omics data, their effectiveness is often constrained by the use of fixed K-nearest neighbor (KNN) graph structures, which fail to capture latent semantic relationships masked by sequencing noise. To overcome these limitations, we propose CRCT (Clustering Rare Cell Types): a novel framework that combines Implicit Semantic Data Augmentation (ISDA) with adaptive graph learning for spatial multi-modal omics analysis. Unlike traditional augmentation strategies that generate explicit synthetic samples, CRCT operates in the deep feature space by dynamically estimating intra-class covariance matrices and implicitly perturbing features along semantically meaningful directions. This enables effective augmentation for rare cell populations while preserving biological fidelity. Extensive experiments across four real-world datasets (HLN, MB, Stereo‑CITE‑seq, and SPOTS) and one synthetic benchmark demonstrate the state-of-the-art performance of CRCT, achieving improvements of up to +1.7 NMI and +7.8 ARI over strong baseline methods.
Daixian Liu, Hau-Sing So, Shanshan Wang 0008, Mengzhu Wang, Jingcai Guo
AAAI6
2026 Nested Graph Pseudo-Label Refinement for Noisy Label Domain Adaptation Learning
abstract
Graph Domain Adaptation (GDA) facilitates knowledge transfer from labeled source graphs to unlabeled target graphs by learning domain-invariant representations, which is essential in applications such as molecular property prediction and social network analysis. However, most existing GDA methods rely on the assumption of clean source labels, which rarely holds in real-world scenarios where annotation noise is pervasive. This label noise severely impairs feature alignment and degrades adaptation performance under domain shifts. To address this challenge, we propose Nested Graph Pseudo-Label Refinement (NeGPR), a novel framework tailored for graph-level domain adaptation with noisy labels. NeGPR first pretrains dual branches, i.e., semantic and topology branches, by enforcing neighborhood consistency in the feature space, thereby reducing the influence of noisy supervision. To bridge domain gaps, NeGPR employs a nested refinement mechanism in which one branch selects high-confidence target samples to guide the adaptation of the other, enabling progressive cross-domain learning. Furthermore, since pseudo-labels may still contain noise and the pre-trained branches are already overfitted to the noisy labels in the source domain, NeGPR incorporates a noise-aware regularization strategy. This regularization is theoretically proven to mitigate the adverse effects of pseudo-label noise, even under the presence of source overfitting, thus enhancing the robustness of the adaptation process. Extensive experiments on benchmark datasets demonstrate that NeGPR consistently outperforms state-of-the-art methods under severe label noise.
Mengzhu Wang, Suyu Liu
AAAI2
2026 DeFT-LoRA: Decoupled and Fused Tuning with LoRA Experts for Universal Cross-Domain Retrieval
abstract
Universal Cross-Domain Retrieval (UCDR) aims to retrieve images across unseen domains and categories, a critical capability for real-world applications. While large-scale Vision-Language Models (VLMs) like CLIP offer strong zero-shot category generalization, they struggle with domain shifts. Existing methods often improve domain robustness at the cost of high computational overhead or by compromising the VLM's inherent knowledge. To address this, we propose Decoupled and Fused Tuning with LoRA (DeFT-LoRA), a novel and parameter-efficient framework that integrates Low-Rank Adaptation (LoRA) with a Mixture-of-Experts (MoE) mechanism. This approach resolves the intrinsic conflict between domain-invariant and domain-specific knowledge in a single adapter, enabling our model to construct a domain adapters for each input image. We propose a three-stage training strategy, which first learns a shared Base LoRA for domain-invariant features, then derives Domain-Specific Experts to capture specific styles, and finally fuses them dynamically with a lightweight gating network. Extensive experiments on three UCDR benchmarks demonstrate that DeFT-LoRA achieves comparable or superior performance to state-of-the-art methods while requiring only 1.46 percent of CLIP's image-encoder parameters and reducing computational overhead, thereby establishing an exceptional balance between accuracy and efficiency.
Ke Xu 0011, Xiaozheng Shen, Shanshan Wang 0008, Mengzhu Wang, Xun Yang 0001
AAAI4
2026 Semi-Supervised Medical Image Segmentation via ${\varvec{f}}$ -Divergence Regularized Sharpness-Aware Learning
Mengzhu Wang, Hongcheng Su
ICIC (29)2
2026 Diffusion-based Kriging Model with Graph-enhanced Attention
abstract
In web-based systems, elements are commonly organized within a graph structure, with each node collecting essential spatio-temporal data. Examples include websites on the World Wide Web, traffic monitors in transportation networks, or sensors in the Internet of Things (IoT). However, sensors are typically deployed sparsely and unevenly, leaving the remaining nodes unobserved. The spatio-temporal kriging task, which infers values at unobserved nodes from observed ones, has thus attracted significant research interest. Due to limitations such as reliance on static graph structures and iterative Graph Convolution Network (GCN) frameworks, accurate kriging remains challenging. To address these issues, we propose a Diffusion-based Kriging Model with Graph-enhanced Attention (DKM-GA). Our approach first introduces a graph-enhanced attention mechanism that dynamically learns more accurate graph structures by combining predefined graph knowledge with global node value similarities. It is then integrated into a diffusion-based framework, which is tailored for the reliance of attention on known values. Therefore, the framework progressively refines the target values using correlated nodes, and the graph-enhanced attention selects more relevant neighbors based on the refined values. Furthermore, a node-based rescaling strategy is introduced to align the inference phase graphs to the training ones. Experiments on eight real-world datasets demonstrate that DKM-GA achieves superior performance, reducing estimation errors by up to 12.66%. Moreover, our analysis identifies three practical scenarios where the model delivers greater performance gains, even achieving 19.51% improvements on datasets that show minor gains under standard settings. These results highlight the effectiveness and potential of our model, while the scenarios provide settings for more comprehensive evaluations in terms of performance and robustness.
Guoli Yang, Zhanxing Zhu, Guangyin Jin, Mengzhu Wang, Xiaoying Bai
WWW5
2026 Let Synthetic Data Shine: Domain Reassembly and Soft-Fusion for Single Domain Generalization
Hao Li 0025, Yubin Xiao, Ke Liang 0006, Mengzhu Wang, Long Lan, Kenli Li 0001, Xinwang Liu 0002
Int. J. Comput. Vis.4
2026 D3HRL: A distributed hierarchical reinforcement learning approach based on causal discovery and spurious correlation detection
Chenran Zhao, Dian-xi Shi, Mengzhu Wang, Jianqiang Xia, Huanhuan Yang, Songchang Jin, Shaowu Yang, Chunping Qiu
Neural Networks3
2026 Object style diffusion for generalized object detection in urban scene
Hao Li 0025, Xiangyuan Yang, Mengzhu Wang, Long Lan, Ke Liang 0006, Xinwang Liu 0002, Kenli Li 0001
Pattern Recognit.3
2026 scKAN: Integration of multi-modal single-cell data via adaptive Kolmogorov-Arnold networks
Yu Zhang 0268, Siqi Tian, Anxin Gu, Liang Yang 0002, Mengzhu Wang
Pattern Recognit.9
2026 Probability-Guided Contrastive Learning for Long-Tailed Domain Generalization
abstract
After training on a specific source domain, models can leverage domain generalization (DG) techniques to achieve superior and broader performance on new, unseen target domains. Existing DG often utilizes contrastive learning to learn domain-invariant features. The goal of contrastive learning is to learn effective representations of data, causing samples from the same category to cluster together in feature space, while samples from different categories are dispersed. Traditional contrastive learning is limited to a finite set of contrastive pairs for DG. To handle this problem, we consider sampling from an infinite number of contrastive pairs using a mixture of von Mises-Fisher (vMF) distributions on the unit hypersphere. We propose a novel method called Probability-guided Contrastive Learning (PgCL), which selects contrastive pairs based on estimated data distributions of samples from each category in feature space. Additionally, we derive the exact analytical formula for the expected contrastive loss. We conduct an empirical investigation of the error bounds of PgCL and demonstrate its performance by comparing it with several leading methods across a range of DG datasets.
Mengzhu Wang, Houcheng Su, Shanshan Wang 0008, Long Lan, Liang Yang 0002, Li Shen 0008
IEEE Trans. Big Data1
2026 Phrase Grounding-Based Style Transfer for Single-Domain Generalized Object Detection
abstract
Single-domain generalized object detection aims to enhance a model’s generalization to multiple unseen target domains using only data from a single source domain during training. This is a practical yet challenging scenario, as it requires the model to address domain shift without incorporating target domain data into the training process. In this paper, we propose a novel phrase-grounding-based style transfer (PGST) approach for the task. Specifically, we first define textual prompts to describe objects for potential unseen target domains. Then, we leverage the grounded language-image pre-training (GLIP) model to capture the styles of these target domains and perform style transfer from the source to the target domains. The style-transferred visual features from the source domain are semantically rich and closely approximate those of their hypothetical counterparts in the target domain. Finally, we employ these style-transferred visual features to fine-tune GLIP. By introducing these imaginary counterparts, the detector can be effectively generalized to unseen target domains using only a single source domain during training. Our method significantly improves mean average precision (mAP), with an average increase of 8.8% across five diverse weather-driving benchmarks. Notably, our approach outperforms or matches the performance of domain-adaptive object detection methods, which require target domain data for training, in several challenging scenarios.
Wei Wang 0335, Cong Wang 0018, Mengzhu Wang, Xiang Zhang 0008, Long Lan, Xinwang Liu 0002, Kenli Li 0001, Xiaochun Cao
IEEE Trans. Circuits Syst. Video Technol.4
2026 Explainability-Guided Untargeted Attacks on Knowledge Graph Embedding
Yawei Lin, Hao Yu 0017, Ke Liang 0006, Mengzhu Wang, Liang Yang 0002, Xinwang Liu 0002
IEEE Trans. Inf. Forensics Secur.4
2026 TypiCD: Cognitive Diagnosis via Problem-Type-Guided Bias Correction
Shanshan Wang 0008, Yali Ye, Xun Yang 0001, Pichao Wang, Mengzhu Wang, Xingyi Zhang 0001
IEEE Trans. Knowl. Data Eng.5
2025 Pano3R: Training Free Panoramic 3D Reconstruction
abstract
Panoramic 3D reconstruction is essential for immersive scene understanding in robotics, AR, and autonomous driving. However, most existing methods are designed for pinhole images and generalize poorly to 360° inputs due to the scarcity of panoramic training data and the high cost of retraining. We present Pano3R, the first training-free framework for panoramic 3D reconstruction that adapts existing pinhole-based models without any retraining. Pano3R consists of two stages. Specifically, the pre-processing stage applies a position-aware pairing strategy to decompose each panorama into a minimal set of perspective views. These views are selected to ensure sufficient co-visible regions while minimizing the number of projections. The test-time optimization stage incorporates a pose-prior-guided global alignment strategy to improve global consistency and mitigate accumulated errors. Our method enables accurate 360° reconstruction under both single- and multi-view input conditions. Extensive experiments demonstrate that Pano3R consistently improves reconstruction accuracy and pose estimation quality, establishing a strong and practical benchmark for training-free panoramic 3D reconstruction.
Shiming Song 0003, Yongjun Zhang 0006, Yuanze Wang, Mengzhu Wang, Yuetian Wang, Zhuojing Tian, Jinming Song, Dian-xi Shi
ECAI4
2025 GraphCL: Graph-based Clustering for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning (SSL) has made notable advancements in medical image segmentation (MIS), particularly in scenarios with limited labeled data and significantly enhancing data utilization efficiency. Previous methods primarily focus on complex training strategies to utilize unlabeled data but neglect the importance of graph structural information. Different from existing methods, we propose a graph-based clustering for semi-supervised medical image segmentation (GraphCL) by jointly modeling graph data structure in a unified deep model. The proposed GraphCL model enjoys several advantages. Firstly, to the best of our knowledge, this is the first work to model the data structure information for semi-supervised medical image segmentation (SSMIS). Secondly, to get the clustered features across different graphs, we integrate both pairwise affinities between local image features and raw features as inputs. Extensive experimental results on three standard benchmarks show that the proposed GraphCL algorithm outperforms state-of-the-art semi-supervised medical image segmentation methods.
Mengzhu Wang, Houcheng Su, Li Shen 0008, Jingcai Guo
ICML1
2025 ESBN: Estimation Shift of Batch Normalization for Source-free Universal Domain Adaptation
abstract
Domain adaptation (DA) is crucial for transferring models trained in one domain to perform well in a different, often unseen domain. Traditional methods, including unsupervised domain adaptation (UDA) and source-free domain adaptation (SFDA), have made significant progress. However, most existing DA methods rely heavily on Batch Normalization (BN) layers, which are not optimal in source-free settings, where the source domain is unavailable for comparison. In this study, we propose a novel method, ESBN, which addresses the challenge of domain shift by adjusting the placement of normalization layers and replacing BN with Batch-free Normalization (BFN). Unlike BN, BFN is less dependent on batch statistics and provides more robust feature representations through instance-specific statistics. We systematically investigate the effects of different BN layer placements across various network configurations and demonstrate that selective replacement with BFN improves generalization performance. Extensive experiments on multiple domain adaptation benchmarks show that our approach outperforms state-of-the-art methods, particularly in challenging scenarios such as Open-Partial Domain Adaptation (OPDA).
Houcheng Su, Bingli Wang, Yuandong Min, Mengzhu Wang, Shanshan Wang 0008, Jingcai Guo
IJCAI5
2025 Exploring Transferable Homogenous Groups for Compositional Zero-Shot Learning
abstract
Conditional dependency present one of the trickiest problems in Compositional Zero-Shot Learning, leading to significant property variations of the same state (object) across different objects (states). To address this problem, existing approaches often adopt either all-to-one or one-to-one representation paradigms. However, these extremes create an imbalance in the seesaw between transferability and discriminability, favoring one at the expense of the other. Comparatively, humans are adept at analogizing and reasoning in a hierarchical clustering manner, intuitively grouping categories with similar properties to form cohesive concepts. Motivated by this, we propose Homogeneous Group Representation Learning (HGRL), a new perspective formulates state (object) representation learning as multiple homogeneous sub-group representation learning. HGRL seeks to achieve a balance between semantic transferability and discriminability by adaptively discovering and aggregating categories with shared properties, learning distributed group centers that retain group-specific discriminative features. Our method integrates three core components designed to simultaneously enhance both the visual and prompt representation capabilities of the model. Extensive experiments on three benchmark datasets validate the effectiveness of our method. Code is available at https://github.com/zjrao/HGRL.
Zhijie Rao, Jingcai Guo, Miaoge Li, Yang Chen 0039, Mengzhu Wang
IJCAI5
2025 In-Context Meta LoRA Generation
abstract
Low-rank Adaptation (LoRA) has demonstrated remarkable capabilities for task specific fine-tuning. However, in scenarios that involve multiple tasks, training a separate LoRA model for each one results in considerable inefficiency in terms of storage and inference. Moreover, existing parameter generation methods fail to capture the correlations among these tasks, making multi-task LoRA parameter generation challenging. To address these limitations, we propose In-Context Meta LoRA (ICM-LoRA), a novel approach that efficiently achieves task-specific customization of large language models (LLMs). Specifically, we use training data from all tasks to train a tailored generator, Conditional Variational Autoencoder (CVAE). CVAE takes task descriptions as inputs and produces task-aware LoRA weights as outputs. These LoRA weights are then merged with LLMs to create task-specialized models without the need for additional fine-tuning. Furthermore, we utilize in-context meta-learning for knowledge enhancement and task mapping, to capture the relationship between tasks and parameter distributions. As a result, our method achieves more accurate LoRA parameter generation for diverse tasks using CVAE. ICM-LoRA enables more accurate LoRA parameter reconstruction than current parameter reconstruction methods and is useful for implementing task-specific enhancements of LoRA parameters. At the same time, our method occupies 283MB, only 1% storage compared with the original LoRA. The code is available at https://github.com/YihuaJerry/ICM-LoRA.
Yihua Shao, Minxi Yan, Yang Liu 0360, Siyu Chen 0021, Xinwei Long, Ziyang Yan, Lei Li 0050, Nicu Sebe, Hao Tang 0005, Yan Wang 0068, Hao Zhao 0002, Mengzhu Wang, Jingcai Guo
IJCAI14
2025 Wave-wise Discriminative Tracking by Phase-Amplitude Separation, Augmentation and Mixture
abstract
Distinguishing key features in complex visual tasks is challenging. A novel approach treats image patches (tokens) as waves. By using both phase and amplitude, it captures richer semantics and specific invariances compared to pixel-based methods, and allows for feature fusion across regions for a holistic image representation. Based on this, we propose the Wave-wise Discriminative Transformer Tracker (WDT). During tracking, WDT represents features via phase-amplitude separation, enhancement, and mixture. First, we designed a Mutual Exclusive Phase-Amplitude Extractor (MEPAE) to separate phase and amplitude features with distinct semantics, representing spatial target info and background brightness respectively. Then, Wave-wise Feature Augmentation is carried out with two submodules: Phase-Amplitude Feature Augmentation and Mixture. The augmentation module disrupts the separated features in the same batch, and the mixture module recombines them to generate positive and negative waves. The original features are aggregated into the original wave. Positive waves have the same phase but different amplitudes, and negative waves have different phase components. Finally, self-supervised and tracking-supervised losses guide the global and local representation learning for original, positive, and negative waves, enhancing wave-level discrimination. Experiments on five benchmarks prove the effectiveness of our method.
Huibin Tan, Mingyu Cao, Xihuai He, Hao Li 0025, Long Lan, Mengzhu Wang
IJCAI8
2025 Gaussian Mixture Model for Graph Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) has been widely studied with the goal of transferring knowledge from a label-rich source domain to a related but unlabeled target domain. Most UDA techniques achieve this by reducing the feature discrepancies between the two domains to learn domain-invariant feature representations. While domain-invariant feature representations can reduce the differences between the source and target domains, excessively simplifying these differences may cause the model to overlook important domain-specific features, resulting in a decline in transfer learning effectiveness. To address this issue, this paper proposes a novel Gaussian Mixture Model for graph domain adaptation (GMM). This model effectively reduces the distributional bias between the source and target domains by modeling the distribution differences on a graph structure. GMM leverages the local structural information of the graph and the clustering capability of the Gaussian mixture model to automatically learn the latent mapping relationships between the source and target domains. To the best of our knowledge, this is the first work to introduce a Gaussian mixture model into UDA. Extensive experimental results on three standard benchmarks demonstrate that the proposed GMM algorithm outperforms state-of-the-art unsupervised domain adaptation methods in terms of performance.
Mengzhu Wang, Wenhao Ren, Yu Zhang 0268, Yanlong Fan, Dian-xi Shi, Luoxi Jing
IJCAI1
2025 Coupling Category Alignment for Graph Domain Adaptation
abstract
Graph domain adaptation (GDA), which transfers knowledge from a labeled source domain to an unlabeled target graph domain, attracts considerable attention in numerous fields. However, existing methods commonly employ message-passing neural networks (MPNNs) to learn domain-invariant representations by aligning the entire domain distribution, inadvertently neglecting category-level distribution alignment and potentially causing category confusion. To address the problem, we propose an effective framework named Coupling Category Alignment (CoCA) for GDA, which effectively addresses the category alignment issue with theoretical guarantees. CoCA incorporates a graph convolutional network branch and a graph kernel network branch, which explore graph topology in implicit and explicit manners. To mitigate category-level domain shifts, we leverage knowledge from both branches, iteratively filtering highly reliable samples from the target domain using one branch and fine-tuning the other accordingly. Furthermore, with these reliable target domain samples, we incorporate the coupled branches into a holistic contrastive learning framework. This framework includes multi-view contrastive learning to ensure consistent representations across the dual branches, as well as cross-domain contrastive learning to achieve category-level domain consistency. Theoretically, we establish a sharper generalization bound, which ensures the effectiveness of category alignment. Extensive experiments on benchmark datasets validate the superiority of the proposed CoCA compared with baselines.
Xiao Teng, Zhiguang Cao, Mengzhu Wang
IJCAI4
2025 DREAM: A Dual Variational Framework for Unsupervised Graph Domain Adaptation
abstract
Graph classification has been a prominent problem in graph machine learning fields. This problem has been investigated by leveraging message passing neural networks (MPNNs) to learn powerful graph representations. However, MPNNs extract topological semantics implicitly under label supervision, which could suffer from domain shift and label scarcity in unsupervised domain adaptation settings. In this paper, we propose an effective solution named Dual Variational Semantics Graph Mining (DREAM) for unsupervised graph domain adaptation by combining graph structural semantics from complementary perspectives. Besides a message passing branch to learn implicit semantics, our DREAM trains a path aggregation branch, which can provide explicit high-order structural semantics as a supplement. To train these two branches conjointly, we employ an expectation-maximization (EM) style variational framework for the maximization of likelihood. In the E-step, we fix the message passing branch and construct a graph-of-graph to indicate the geometric correlation between source and target domains, which would be adopted for the optimization of the other branch. In the M-step, we train the message passing branch and update the graph neural networks on the graph-of-graph with the other branch fixed. The alternative optimization improves the collaboration of knowledge from two branches. Extensive experiments on several benchmark datasets validate the superiority of the proposed DREAM compared with various baselines.
Li Shen 0008, Mengzhu Wang, Xinwang Liu 0002, Chong Chen 0002, Xian-Sheng Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Graph Convolutional Mixture-of-Experts Learner Network for Long-Tailed Domain Generalization
abstract
The goal of single domain generalization is to use data from a single domain (source domain) to train a model, which is then deployed over several unknown domains for testing (target domains). This study introduces a practical approach diverging from traditional DG, which typically relies on multiple source domains. We focus on Single Long-Tailed Domain Generalization, which refers to a scenario in the context of long-tail distribution, where although minority classes may have fewer samples in a single domain, these minority classes could become more prevalent and dominant in other domains. We introduce the Graph Convolutional Mixture-of-Experts Learners Network for Long-Tailed Domain Generalization (GCML) as a solution to this problem. Our approach presents two novel tactics. Initially, we utilize an expert learning technique that is skill-diverse. In order to properly manage the unknown target domain, this entails training multiple specialists inside a single long-tailed source domain and combining their knowledge. Then, we use a graph convolutional network to facilitate domain generalization, leveraging joint data structure modeling to learn more domain-invariant feature. Experiments conducted on four established benchmarks reveal that our GCML algorithm outperforms contemporary domain generalization techniques, demonstrating its efficacy in this complex task.
Mengzhu Wang, Houcheng Su, Shanshan Wang 0008, Li Shen 0008, Long Lan, Liang Yang 0002, Xiaochun Cao
IEEE Trans. Circuits Syst. Video Technol.1
2025 Optimal Graph Learning-Based Label Propagation for Cross-Domain Image Classification
abstract
Label propagation (LP) is a popular semi-supervised learning technique that propagates labels from a training dataset to a test one using a similarity graph, assuming that nearby samples should have similar labels. However, the recent cross-domain problem assumes that training (source domain) and test data sets (target domain) follow different distributions, which may unexpectedly degrade the performance of LP due to small similarity weights connecting the two domains. To address this problem, we propose optimal graph learning-based label propagation (OGL2P), which optimizes one cross-domain graph and two intra-domain graphs to connect the two domains and preserve domain-specific structures, respectively. During label propagation, the cross-domain graph draws two labels close if they are nearby in feature space and from different domains, while the intra-domain graph pulls two labels close if they are nearby in feature space and from the same domain. This makes label propagation more insensitive to cross-domain problems. During graph embedding, we optimize the three graphs using features and labels in the embedded subspace to extract locally discriminative and domain-invariant features and make the graph construction process robust to noise in the original feature space. Notably, as a more relaxed constraint, locally discriminative and domain-invariant can somewhat alleviate the contradiction between discriminability and domain-invariance. Finally, we conduct extensive experiments on five cross-domain image classification datasets to verify that OGL2P outperforms some state-of-the-art cross-domain approaches.
Wei Wang 0335, Mengzhu Wang, Chao Huang 0008, Cong Wang 0018, Jie Mu, Feiping Nie 0001, Xiaochun Cao
IEEE Trans. Image Process.2
2025 STFormer: Spatial-Temporal-Aware Transformer for Video Instance Segmentation
abstract
Video instance segmentation (VIS) is a challenging task, requiring handling object classification, segmentation, and tracking in videos. Existing Transformer-based VIS approaches have shown remarkable success, combining encoded features and instance queries as decoder inputs. However, their decoder inputs are low-resolution due to computational cost, resulting in a loss of fine-grained information, sensitivity to background interference, and poor handling of small objects. Moreover, the queries are randomly initialized without location information, hindering convergence efficiency and accurate object instance localization. To address these issues, we propose a novel VIS approach, STFormer, with a spatial-temporal feature aggregation (STFA) module and spatial-temporal-aware Transformer (STT). Specifically, STFA obtains robust high-resolution masked features efficiently for the decoder, while STT's location-guided instance query (LGIQ) improves initial instance queries. STFormer preserves more fine-grained information, improves convergence efficiency, and localizes object instance features accurately. Extensive experiments on YouTube-VIS 2019, YouTube-VIS 2021, and OVIS datasets show that STFormer outperforms mainstream VIS methods.
Wei Wang 0335, Mengzhu Wang, Huibin Tan, Long Lan, Zhigang Luo, Xinwang Liu 0002, Kenli Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Smooth-Guided Implicit Data Augmentation for Domain Generalization
abstract
The training process of a domain generalization (DG) model involves utilizing one or more interrelated source domains to attain optimal performance on an unseen target domain. Existing DG methods often use auxiliary networks or require high computational costs to improve the model's generalization ability by incorporating a diverse set of source domains. In contrast, this work proposes a method called Smooth-Guided Implicit Data Augmentation (SGIDA) that operates in the feature space to capture the diversity of source domains. To amplify the model's generalization capacity, a distance metric learning (DML) loss function is incorporated. Additionally, rather than depending on deep features, the suggested approach employs logits produced from cross entropy (CE) losses with infinite augmentations. A theoretical analysis shows that logits are effective in estimating distances defined on original features, and the proposed approach is thoroughly analyzed to provide a better understanding of why logits are beneficial for DG. Moreover, to increase the diversity of the source domain, a sampling-based method called smooth is introduced to obtain semantic directions from interclass relations. The effectiveness of the proposed approach is demonstrated through extensive experiments on widely used DG, object detection, and remote sensing datasets, where it achieves significant improvements over existing state-of-the-art methods across various backbone networks.
Mengzhu Wang, Junze Liu, Ge Luo 0003, Shanshan Wang 0008, Wei Wang 0335, Long Lan, Ye Wang 0023, Feiping Nie 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 HCL: A Hierarchical Contrastive Learning Framework for Zero-Shot Relation Extraction
abstract
Zero-shot relation extraction (ZSRE) is shown to become more significant in the current information extraction system, which aims at predicting relation classes that lack annotations or have just never appeared during training. Previous works focus on projecting sentences with their corresponding relation descriptions to an intermediate semantic space and searching the nearest semantic for predicting unseen classes. Though these methods can achieve sound performance, they only obtain inferior semantic information via a trivial distance metric and neglect the interaction in the instance representations. We are thus motivated to tackle these issues and propose a hierarchical contrastive learning (HCL) framework for ZSRE including projection-level and instance-level modules. Specifically, the projection-level component replaces the distance score function by contrastive loss to connect the input sentence with the relation semantic space. And the instance-level component integrates the external knowledge from sentence entities to establish new contrastive pairs for efficiently learning representations from mutual information. The experimental results on three well-known datasets demonstrate that our model surpasses the existing SOTA by at most 18.97% improvement on the F1 score when unseen classes are 15. Moreover, our model can achieve more competitive performance alone with the increasing number of unseen classes.
Tianwei Yan 0001, Shan Zhao 0002, Minghao Hu 0001, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Correlation Matching Transformation Transformers for UHD Image Restoration
abstract
This paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution space and (b) learning in low-resolution space. The former learns multi-level high-resolution features and fuses low-high features and reconstructs the residual images, while the latter explores more representative features learning from the high-resolution ones to facilitate better restoration. To better improve feature representation in low-resolution space, we propose to build feature transformation from the high-resolution space to the low-resolution one. To that end, we propose two new modules: Dual-path Correlation Matching Transformation module (DualCMT) and Adaptive Channel Modulator (ACM). The DualCMT selects top C/r (r is greater or equal to 1 which controls the squeezing level) correlation channels from the max-pooling/mean-pooling high-resolution features to replace low-resolution ones in Transformers, which can effectively squeeze useless content to improve the feature representation in low-resolution space to facilitate better recovery. The ACM is exploited to adaptively modulate multi-level high-resolution features, enabling to provide more useful features to low-resolution space for better learning. Experimental results show that our UHDformer reduces about ninety-seven percent model sizes compared with most state-of-the-art methods while significantly improving performance under different training sets on 3 UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring. The source codes will be made available at https://github.com/supersupercong/UHDformer.
Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Mengzhu Wang, Xiao-Ming Wu 0003, Jun Liu 0036
AAAI6
2024 Sharpness-Aware Model-Agnostic Long-Tailed Domain Generalization
abstract
Domain Generalization (DG) aims to improve the generalization ability of models trained on a specific group of source domains, enabling them to perform well on new, unseen target domains. Recent studies have shown that methods that converge to smooth optima can enhance the generalization performance of supervised learning tasks such as classification. In this study, we examine the impact of smoothness-enhancing formulations on domain adversarial training, which combines task loss and adversarial loss objectives. Our approach leverages the fact that converging to a smooth minimum with respect to task loss can stabilize the task loss and lead to better performance on unseen domains. Furthermore, we recognize that the distribution of objects in the real world often follows a long-tailed class distribution, resulting in a mismatch between machine learning models and our expectations of their performance on all classes of datasets with long-tailed class distributions. To address this issue, we consider the domain generalization problem from the perspective of the long-tail distribution and propose using the maximum square loss to balance different classes which can improve model generalizability. Our method's effectiveness is demonstrated through comparisons with state-of-the-art methods on various domain generalization datasets. Code: https://github.com/bamboosir920/SAMALTDG.
Houcheng Su, Weihao Luo, Daixian Liu, Mengzhu Wang, Junyang Chen 0001, Cong Wang 0018, Zhenghan Chen
AAAI4
2024 Dynamic Spiking Graph Neural Networks
abstract
The integration of Spiking Neural Networks (SNNs) and Graph Neural Networks (GNNs) is gradually attracting attention due to the low power consumption and high efficiency in processing the non-Euclidean data represented by graphs. However, as a common problem, dynamic graph representation learning faces challenges such as high complexity and large memory overheads. Current work often uses SNNs instead of Recurrent Neural Networks (RNNs) by using binary features instead of continuous ones for efficient training, which overlooks graph structure information and leads to the loss of details during propagation. Additionally, optimizing dynamic spiking models typically requires the propagation of information across time steps, which increases memory requirements. To address these challenges, we present a framework named Dynamic Spiking Graph Neural Networks (Dy-SIGN). To mitigate the information loss problem, Dy-SIGN propagates early-layer information directly to the last layer for information compensation. To accommodate the memory requirements, we apply the implicit differentiation on the equilibrium state, which does not rely on the exact reverse of the forward computation. While traditional implicit differentiation methods are usually used for static situations, Dy-SIGN extends it to the dynamic graph setting. Extensive experiments on three large-scale real-world dynamic graph datasets validate the effectiveness of Dy-SIGN on dynamic node classification tasks with lower computational costs.
Mengzhu Wang, Zhenghan Chen, Giulia De Masi, Huan Xiong, Bin Gu 0001
AAAI2
2024 Improving Graph Contrastive Learning via Adaptive Positive Sampling
abstract
Graph Contrastive Learning (GCL), a Self-Supervised Learning (SSL) architecture tailored for graphs, has shown notable potential for mitigating label scarcity. Its core idea is to amplify feature similarities between the positive sample pairs and reduce them between the negative sample pairs. Unfortunately, most existing GCLs consistently present sub-optimal performances on both homophilic and heterophilic graphs. This is primarily attributed to two limitations of positive sampling, that is, incomplete local sampling and blind sampling. To address these limitations, this paper introduces a novel GCL framework with an adaptive positive sampling module, named grapH contrastivE Adaptive Positive Samples (HEATS). Motivated by the observation that the affinity matrix corresponding to optimal positive sample sets has a block-diagonal structure with equal weights within each block, a self-expressive learning objective incorporating the block and idempotent constraint is presented. This learning objective and the contrastive learning objective are iteratively optimized to improve the adaptability and robustness of HEATS. Extensive experiments on graphs and images validate the effectiveness and generality of HEATS.
Jiaming Zhuo, Feiyang Qin, Can Cui 0005, Bingxin Niu, Mengzhu Wang, Yuanfang Guo, Chuan Wang 0002, Zhen Wang 0004, Xiaochun Cao, Liang Yang 0002
CVPR6
2024 SBM: Smoothness-Based Minimization for Domain Generalization
abstract
In topical domain generalization (DG), trained models are asked to perform well on an unknown target domain with different data statistics. In order to improve domain generalization, adversarial learning has proven to be one of the most effective methods. Existing approaches, however, rely primarily on adversarial learning, which can only generalize within a limited range of domains. We argue that smoothness- based minimization (SBM) is a more promising direction for adversarial domain generalization. Our findings indicate that achieving a smoothness-based minimization of task loss stabilizes adversarial training, resulting in better domain generalization performance. This method has been shown to achieve remarkable domain generalization performance on three publicly available benchmarks including PACS, Office- Home and DomainNet.
Chunqing Ruan, Mengzhu Wang, Shanshan Wang 0008, Tianyi Liang 0001, Wei Yu 0029
ICASSP2
2024 DREAM: Dual Structured Exploration with Mixup for Open-set Graph Domain Adaption
abstract
Recently, numerous graph neural network methods have been developed to tackle domain shifts in graph data. However, these methods presuppose that unlabeled target graphs belong to categories previously seen in the source domain. This assumption could not hold true for in-the-wild target graphs. In this paper, we delve deeper to explore a more realistic problem open-set graph domain adaptation. Our objective is to not only identify target graphs from new categories but also accurately classify remaining target graphs into their respective categories under domain shift and label scarcity. To solve this challenging problem, we introduce a new method named Dual Structured Exploration with Mixup (DREAM). DREAM incorporates a graph-level representation learning branch as well as a subgraph-enhanced branch, which jointly explores graph topological structures from both global and local viewpoints. To maximize the use of unlabeled target graphs, we train these two branches simultaneously using posterior regularization to enhance their inter-module consistency. To accommodate the open-set setting, we amalgamate dissimilar samples to generate virtual unknown samples belonging to novel classes. Moreover, to alleviate domain shift, we establish a k nearest neighbor-based graph-of-graphs and blend multiple neighbors of each sample to produce cross-domain virtual samples for inter-domain consistency learning. Extensive experiments validate the effectiveness of the proposed DREAM in comparison to various state-of-the-art approaches in different settings.
Mengzhu Wang, Zhenghan Chen, Li Shen 0008, Huan Xiong, Bin Gu 0001, Xiao Luo 0001
ICLR2
2024 Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?
abstract
Transformer-based trajectory optimization methods have demonstrated exceptional performance in offline Reinforcement Learning (offline RL). Yet, it poses challenges due to substantial parameter size and limited scalability, which is particularly critical in sequential decision-making scenarios where resources are constrained such as in robots and drones with limited computational power. Mamba, a promising new linear-time sequence model, offers performance on par with transformers while delivering substantially fewer parameters on long sequences. As it remains unclear whether Mamba is compatible with trajectory optimization, this work aims to conduct comprehensive experiments to explore the potential of Decision Mamba (dubbed DeMa) in offline RL from the aspect of data structures and essential components with the following insights: (1) Long sequences impose a significant computational burden without contributing to performance improvements since DeMa's focus on sequences diminishes approximately exponentially. Consequently, we introduce a Transformer-like DeMa as opposed to an RNN-like DeMa. (2) For the components of DeMa, we identify the hidden attention mechanism as a critical factor in its success, which can also work well with other residual structures and does not require position embedding. Extensive evaluations demonstrate that our specially designed DeMa is compatible with trajectory optimization and surpasses previous methods, outperforming Decision Transformer (DT) with higher performance while using 30\% fewer parameters in Atari, and exceeding DT with only a quarter of the parameters in MuJoCo.
Oubo Ma, Xingxing Liang, Shengchao Hu, Mengzhu Wang, Shouling Ji, Jincai Huang 0001, Li Shen 0008
NeurIPS6
2024 Consistency-constrained unsupervised video anomaly detection framework based on Co-teaching
Wenhao Shao, Praboda Rajapaksha, Noël Crespi, Xuechen Zhao, Mengzhu Wang, Xinwang Liu 0002, Zhigang Luo
Neurocomputing5
2024 Compressing the Multiobject Tracking Model via Knowledge Distillation
abstract
Recent multiobject tracking (MOT) methods usually use very deep neural networks to achieve competitive accuracy, which inevitably results in degraded inference speed. To strike a better balance between tracking accuracy and speed, in this work, we propose to compress the MOT model via knowledge distillation (KD), enabling the more lightweight student model to obtain similar performance as the teacher model. Nonetheless, despite KD has been well studied for simpler tasks such as image classification, the complexity of MOT poses new challenges because the MOT model is more sensitive to foreground information than the classification model. To deal with that, we first propose attention-guided feature distillation, which focuses the student model on the crucial region (foreground and the region with strong discrepancy against itself) of the teacher’s feature map. Moreover, we propose foreground mask, which leverages the knowledge from the teacher model to filter out the low-quality soft labels from the background, thereby reducing their negative effects for distillation. Evaluations on several benchmarks demonstrate that the proposed KD method can make the student network achieve leading performance, meanwhile running faster than the teacher network 20.0%–27.4% and reducing the parameters 28.5%–87.1%. To the best of our knowledge, this is the first work to compress the MOT model via KD.
Tianyi Liang 0001, Mengzhu Wang, Junyang Chen 0001, Dingyao Chen, Zhigang Luo, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.2
2024 TFC: Transformer Fused Convolution for Adversarial Domain Adaptation
abstract
In unsupervised domain adaptation (UDA), a classifier is applied to the target domain without or with limited labels, when the target domain has no or few labels. Recently, inspired by their capabilities of long-distance feature dependencies, vision transformer (ViT)-based methods have been used in UDA, however, they ignore the fact that ViT lacks strength in extracting local feature details. To handle the above problems, the purpose of this article is to demonstrate how to take advantage of both convolutional operations and transformer mechanisms for adversarial UDA by using a hybrid network structure called transformer fused convolution (TFC). TFC integrates local features with global features to boost the representation capacity for UDA which can enhance the discrimination between foreground and background. Moreover, to improve the robustness of the TFC, we leverage an uncertainty penalty loss to make incorrect classes have consistently lower scores. Extensive experiments validate the significant performance gains compared to conditional adversarial domain adaptation (CDAN) on all five datasets including DomainNet ($\uparrow ~8.5$%), VisDA-2017 ($\uparrow ~14.9$%), Office-Home ($\uparrow ~18.9$%), Office-31 ($\uparrow ~11.5$%), and ImageCLEF-DA ($\uparrow ~5.5$%).
Mengzhu Wang, Junyang Chen 0001, Ye Wang 0023, Zhiguo Gong, Kaishun Wu, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.1
2024 Multidocument Aspect Classification for Aspect-Based Abstractive Summarization
abstract
Multidocument aspect-based summarization (AspSumm) aims to generate focused summaries based on the target aspects from a cluster of relevant documents. Generating such summaries can better satisfy readers’ specific points of interest, as readers may have different concerns about the same articles. However, previous methods usually generate aspect-based summaries based on the given aspects without using the relationship among aspects to assist in the summarization. In this work, we propose a two-stage general framework for multidocument AspSumm. The model first discovers the latent relationship among aspects and then uses relevant sentences selected by aspect discovery to generate abstractive summaries. We exploit latent dependencies among aspects using a tag mask training (TMT) strategy, which increases the interpretability of the model. In addition to improvements in summarization over aspect-based strong baselines, experimental results show that our proposed model can accurately discover multidomain aspects on the WikiAsp dataset.
Ye Wang 0023, Mengzhu Wang, Zhenghan Chen, Zhiping Cai, Junyang Chen 0001, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.3
2024 Equity in Unsupervised Domain Adaptation by Nuclear Norm Maximization
abstract
Nuclear norm maximization has shown the power to enhance the transferability of unsupervised domain adaptation model (UDA) in an empirical scheme. In this paper, we identify a new property termedequity, which indicates the balance degree of predicted classes, to demystify the efficacy of nuclear norm maximization for UDA theoretically. With this in mind, we offer a new discriminability-and-equity maximization paradigm built on squares loss, such that predictions are equalized explicitly. To verify its feasibility and flexibility, two new losses termed Class Weighted Squares Maximization (CWSM) and Normalized Squares Maximization (NSM), are proposed to maximize both predictive discriminability and equity, from the class level and the sample level, respectively. Importantly, we theoretically relate these two novel losses (i.e., CWSM and NSM) to the equity maximization under mild conditions, and empirically suggest the importance of the predictive equity in UDA. Moreover, it is very efficient to realize the equity constraints in both losses. Experiments of cross-domain image classification on three popular benchmark datasets show that both CWSM and NSM contribute to outperforming the corresponding counterparts.
Mengzhu Wang, Shanshan Wang 0008, Xun Yang 0001, Jianlong Yuan, Wenju Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2024 Inter-Class and Inter-Domain Semantic Augmentation for Domain Generalization
abstract
The domain generalization approach seeks to develop a universal model that performs well on unknown target domains with the aid of diverse source domains. Data augmentation has proven to be an effective method to enhance domain generalization in computer vision. Recently, semantic-level based data augmentation has yielded remarkable results. However, these methods focus on sampling semantic directions on feature space from intra-class and intra-domain, limiting the diversity of the source domain. To address this issue, we propose a novel approach called Inter-Class and Inter-Domain Semantic Augmentation (CDSA) for domain generalization. We first introduce a sampling-based method called CrossSmooth to obtain semantic directions from inter-class. Then, CrossVariance obtains the styles of different domains by sampling semantic directions. Our experiments on four well-known domain generalization benchmark datasets (Digits-DG, PACS, Office-Home, and DomainNet) demonstrate the effectiveness of our approach. We also validate our approach on commonly-used semantic segmentation datasets, namely GTAV, SYNTHIA, Cityscapes, Mapillary, and BDDS which also show significant improvements.
Mengzhu Wang, Yuehua Liu, Jianlong Yuan, Shanshan Wang 0008, Zhibin Wang 0004, Wei Wang 0335
IEEE Trans. Image Process.1
2023 A Topic-Aware Graph-Based Neural Network for User Interest Summarization and Item Recommendation in Social Media
Junyang Chen 0001, Ge Fan, Zhiguo Gong, Xueliang Li 0002, Victor C. M. Leung, Mengzhu Wang
DASFAA (2)6
2023 Enhanced Dcf Tracker Regularized by Reliable Sample Construction
abstract
Discriminative correlation filter (DCF) is a highly efficient tracking technique using the circulant shifted samples of search images to update the template, so the reliability of input samples determines template quality. In this paper, we rethink the reliability problem of input samples in advance during template updating and propose an enhanced DCF tracking method regularized by a novel sparse representation based reliable sample construction term, called enhanced sparse correlation filter (ESCF). Specifically, the reconstructed reliable samples are the sparse representation of circulant shifted samples of unfiltered input samples, in which the target will approach the center to preserve target visual cues into the template when using the cosine window. Besides, we jointly perform template learning and reliable sample construction into a unified learning paradigm to benefit from each other, which further can be carried out in the frequency domain without incurring excessive time cost by skillful decomposition. Experiments on several popular visual tracking datasets verify the efficacy of ESCF and show that ESCF performs favorably against several well-established representative counterparts.
Mingyu Cao, Mengzhu Wang, Long Lan, Wenjing Yang 0002, Huibin Tan
ICASSP3
2023 Progressive Perception Learning for Distribution Modulation in Siamese Tracking
abstract
We explore an innovative view on distribution modulation to boost Siamese trackers. Specially, we observed two cases of possible distribution inconsistency in Siamese tracking: 1) Two branches with different sizes may be in different distribution ranges after a shared backbone (including BN layers). 2) The background data may affect the total feature distribution of the search branch. To address these issues, we proposed a plug-and-play component named Progressive Perception Learning Module (P2LM) to modulate the distribution using three feature normalization blocks successively, i.e., Self-Aware Block (SAB), Target-Aware Block (TAB), and Region-Aware Block (RAB). SAB regulates the distribution of each branch independently for the first issue. TAB uses the target information to guide the distribution adjustments of the two branches. RAB divides the search image into foreground and background with a region mask and normalizes them separately to filter the background distractors for robust tracking. TAB and RAB synergistically alleviate the distribution shifts caused by environmental variance. Experiments on OTB100, UAV123, LaSOT, and GOT-10k verify the compelling effects of our module.
Xianchen Zhou, Mingyu Cao, Mengzhu Wang, Guangjie Gao, Wenjing Yang 0002, Huibin Tan
ICASSP4
2023 Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking
abstract
Tensor decomposition and reconstruction attention is a promising global context learning approach because it can remain efficient while avoiding feature compression. To exploit its potential even further in visual tracking, we redesign a 3D tensor modeling paradigm, namely tensor Decomposition, Interaction, Reconstruction attention (DIR), respectively corresponding to three function components, Tensor Decomposition Module (TDM), Tensor Interaction Module (TIM) and Context Reconstruction Module (CRM). Specifically, TDM decomposes a 3D tensor feature into rank-1 context fragments in different dimension views. The ingenuity here lies in the introduction of Circular Convolution for processing features at arbitrary scales and channel-sharing segments to enhance the interaction of the two branches in the Siamese network architecture. TIM obtains the tensor planes of each dimension by the Cross-Similarity operation of rank-1 tensors and fused cubic features, which brings more interactions between all feature dimensions. CRM reconstructs 3D context representations with the outputs of the above modules. In experiments, DIR is embedded into the tracker to verify its effectiveness.
Huibin Tan, Mingyu Cao, Mengzhu Wang, Wenjing Yang 0002
ICASSP4
2023 Atomic-action-based Contrastive Network for Weakly Supervised Temporal Language Grounding
abstract
As one knows, an event often consists of several actions while each action is atomic. Inspired by this insight, we propose a novel framework named Atomic-action-based Contrastive Network model (ACN) for weakly supervised temporal language grounding task to localize the query-related event moment in an untrimmed video, without access to any temporal annotations. Specifically, ACN first determines the accurate moment boundary of each action in a query-agnostic way. This can adequately exploit homogeneous visual cues while impeding the heterogeneity of the query from hurting the atomicity of visual action, i.e., action boundary. To effectively localize the query-related event, we seek the discriminative words in the given query, and explore a composite-grained contrastive module to retrieve those corresponding atomic actions in the common latent space across modalities. This boosts feature discrimination of visual event segment to remove irrelevant action video segments. Experiments on two popular datasets show the efficacy of our model.
Hongzhou Wu, Xuechen Zhao, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo
ICME5
2023 CoCo: A Coupled Contrastive Framework for Unsupervised Domain Adaptive Graph Classification
abstract
Although graph neural networks (GNNs) have achieved impressive achievements in graph classification, they often need abundant task-specific labels, which could be extensively costly to acquire. A credible solution is to explore additional labeled graphs to enhance unsupervised learning on the target domain. However, how to apply GNNs to domain adaptation remains unsolved owing to the insufficient exploration of graph topology and the significant domain discrepancy. In this paper, we propose Coupled Contrastive Graph Representation Learning (CoCo), which extracts the topological information from coupled learning branches and reduces the domain discrepancy with coupled contrastive learning. CoCo contains a graph convolutional network branch and a hierarchical graph kernel network branch, which explore graph topology in implicit and explicit manners. Besides, we incorporate coupled branches into a holistic multi-view contrastive learning framework, which not only incorporates graph representations learned from complementary views for enhanced understanding, but also encourages the similarity between cross-domain example pairs with the same semantics for domain alignment. Extensive experiments on popular datasets show that our CoCo outperforms these competing baselines in different settings generally.
Li Shen 0008, Mengzhu Wang, Long Lan, Zeyu Ma 0001, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001
ICML3
2023 Disentangled Representation Learning with Causality for Unsupervised Domain Adaptation
abstract
Most efforts in unsupervised domain adaptation (UDA) focus on learning the domain-invariant representations between the two domains. However, such representations may still confuse two patterns due to the domain gap. Considering that semantic information is useful for the final task and domain information always indicates the discrepancy between two domains, to address this issue, we propose to decouple the representations of semantic features from domain features to reduce domain bias. Different from traditional methods, we adopt a simple but effective module with only one domain discriminator to decouple the representations, offering two benefits. Firstly, it eliminates the need for labeled sample pairs, making it more suitable for UDA. Secondly, without adversarial learning, our model can achieve a more stable training phase. Moreover, to further enhance the task-specific features, we employ a causal mechanism to separate semantic features related to causal factors from the overall feature representations. Specially, we utilize a dual-classifier strategy, where each classifier is fed with the entire features and the semantic features, respectively. By minimizing the discrepancy between the outputs of the two classifiers, the causal influence of the semantic features is accentuated. Experiments on several public datasets demonstrate the proposed model can outperform the state-of-the-art methods. Our code is available at: https://github.com/qzxRtY37/DRLC https://github.com/qzxRtY37/DRLC.
Shanshan Wang 0008, Zhenwei He, Xun Yang 0001, Mengzhu Wang, Quanzeng You, Xingyi Zhang 0001
ACM Multimedia5
2023 Zero-shot Micro-video Classification with Neural Variational Inference in Graph Prototype Network
abstract
Micro-video classification plays a central role in online content recommendation platforms, such as Kwai and Tik-Tok. Existing works on video classification largely exploit the interactions between users and items as well as the item labels to provide quality recommendation services. However, scarce or even no labeled data of emerging videos is a great challenge for existing classification methods. In this paper, we propose a zero-shot micro-video classification model (NVIGPN) by exploiting the hidden topics behind items to guide the representation learning in user-item interactions. Specifically, we study this zero-shot classification in two stages: (1) exploiting a generalized semantic hidden topic descriptions for transferable knowledge learning, and (2) designing a graph-based learning model for guiding the minor seen class information to the unseen ones. Through mining the transferable knowledge between the hidden topics and the small number of the seen classes, NVIGPN can achieves state-of-the-art performances in predicting the unseen classes of micro-videos. We conduct extensive experiments to demonstrate the effectiveness of our method.
Junyang Chen 0001, Zhijiang Dai, Huisi Wu, Mengzhu Wang, Qin Zhang 0011, Huan Wang 0005
ACM Multimedia5
2023 A Closer Look at Classifier in Adversarial Domain Generalization
abstract
The task of domain generalization is to learn a classification model from multiple source domains and generalize it to unknown target domains. The key to domain generalization is learning discriminative domain-invariant features. Invariant representations are achieved using adversarial domain generalization as one of the primary techniques. For example, generative adversarial networks have been widely used, but suffer from the problem of low intra-class diversity, which can lead to poor generalization ability. To address this issue, we propose a new method called auxiliary classifier in adversarial domain generalization (CloCls). CloCls improve the diversity of the source domain by introducing auxiliary classifier. Combining typical task-related losses, e.g., cross-entropy loss for classification and adversarial loss for domain discrimination, our overall goal is to guarantee the learning of condition-invariant features for all source domains while increasing the diversity of source domains. Further, inspired by smoothing optima have improved generalization for supervised learning tasks like classification. We leverage that converging to a smooth minima with respect task loss stabilizes the adversarial training leading to better performance on unseen target domain which can effectively enhances the performance of domain adversarial methods. We have conducted extensive image classification experiments on benchmark datasets in domain generalization, and our model exhibits sufficient generalization ability and outperforms state-of-the-art DG methods.
Ye Wang 0023, Junyang Chen 0001, Mengzhu Wang, Hao Li 0058, Wei Wang 0335, Houcheng Su, Zhihui Lai 0001, Wei Wang 0077, Zhenghan Chen
ACM Multimedia3
2023 Interpolation Normalization for Contrast Domain Generalization
abstract
Domain generalization refers to the challenge of training a model from various source domains that can generalize well to unseen target domains. Contrastive learning is a promising solution that aims to learn domain-invariant representations by utilizing rich semantic relations among sample pairs from different domains. One simple approach is to bring positive sample pairs from different domains closer, while pushing negative pairs further apart. However, in this paper, we find that directly applying contrastive-based methods is not effective in domain generalization. To overcome this limitation, we propose to leverage a novel contrastive learning approach that promotes class-discriminative and class-balanced features from source domains. Essentially, clusters of sample representations from the same category are encouraged to cluster, while those from different categories are spread out, thus enhancing the model's generalization capability. Furthermore, most existing contrastive learning methods use batch normalization, which may prevent the model from learning domain-invariant features. Inspired by recent research on universal representations for neural networks, we propose a simple emulation of this mechanism by utilizing batch normalization layers to distinguish visual classes and formulating a way to combine them for domain generalization tasks. Our experiments demonstrate a significant improvement in classification accuracy over state-of-the-art techniques on popular domain generalization benchmarks, including Digits-DG, PACS, Office-Home and DomainNet.
Mengzhu Wang, Junyang Chen 0001, Huan Wang 0005, Huisi Wu, Zhidan Liu 0001, Qin Zhang 0011
ACM Multimedia1
2023 Mixture-of-Experts Learner for Single Long-Tailed Domain Generalization
abstract
Domain generalization (DG) refers to the task of training a model on multiple source domains and test it on a different target domain with different distribution. In this paper, we address a more challenging and realistic scenario known as Single Long-Tailed Domain Generalization, where only one source domain is available and the minority class in this domain has an abundance of instances in other domains. To tackle this task, we propose a novel approach called Mixture-of-Experts Learner for Single Long-Tailed Domain Generalization (MoEL), which comprises two key strategies. The first strategy is a simple yet effective data augmentation technique that leverages saliency maps to identify important regions on the original images and preserves these regions during augmentation. The second strategy is a new skill-diverse expert learning approach that trains multiple experts from a single long-tailed source domain and leverages mutual learning to aggregate their learned knowledge for the unknown target domain. We evaluate our method on various benchmark datasets, including Digits-DG, CIFAR-10-C, PACS, and DomainNet, and demonstrate its superior performance compared to previous single domain generalization methods. Additionally, the ablation study is also conducted to illustrate the inner workings of our approach.
Mengzhu Wang, Jianlong Yuan, Zhibin Wang 0004
ACM Multimedia1
2023 PromptRestorer: A Prompting Image Restoration Method with Degradation Perception
abstract
We show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely hindered. To address this, we propose a Prompting image Restorer, termed as PromptRestorer. Specifically, PromptRestorer contains two branches: a restoration branch and a prompting branch. The former is used to restore images, while the latter perceives degradation priors to prompt the restoration branch with reliable perceived content to guide the restoration process for better recovery. To better perceive the degradation which is extracted by a pre-trained model from given degradation observations, we propose a prompting degradation perception modulator, which adequately considers the characters of the self-attention mechanism and pixel-wise modulation, to better perceive the degradation priors from global and local perspectives. To control the propagation of the perceived content for the restoration branch, we propose gated degradation perception propagation, enabling the restoration branch to adaptively learn more useful features for better recovery. Extensive experimental results show that our PromptRestorer achieves state-of-the-art results on 4 image restoration tasks, including image deraining, deblurring, dehazing, and desnowing.
Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Jiangxin Dong, Mengzhu Wang, Yakun Ju, Junyang Chen 0001
NeurIPS5
2023 Self-aware circular response-guided attention for robust siamese tracking
Huibin Tan, Mengzhu Wang, Tianyi Liang 0001, Yuhua Tang, Long Lan, Wenjing Yang 0002
Appl. Intell.2
2023 A Neural Inference of User Social Interest for Item Recommendation
abstract
Abstract User-generated content is daily produced in social media, as such user interest summarization is critical to distill salient information from massive information for recommendation tasks. While the interested messages (e.g., tags or posts) from a single user are usually sparse becoming a bottleneck for existing methods, we propose a neural inference method (NIGraphNet) by mining user social interest for item recommendation. It can unearth user latent topics combined with user relation learning. Specifically, we exploit a neural variational inference approach to learn the distributions between user interests and hidden topics. (We denote it as interest-topic distributions in the following.) Then, we adopt a unified graph-based training loss that jointly learns the hidden topics and user relations for item recommendation. Experiments on two datasets collected from well-known social media platforms demonstrate the superior performance of our model in the tasks of user interest summarization and item recommendation. Further discussions also show that exploiting the latent topic representations and user relations is conducive to the user’s automatic language understanding.
Junyang Chen 0001, Mengzhu Wang, Ge Fan, Guo Zhong, Ou Liu, Wenfeng Du, Zhenghua Xu 0001, Zhiguo Gong
Data Sci. Eng.3
2023 Domain-specific feature recalibration and alignment for multi-source unsupervised domain adaptation
abstract
Abstract Traditional unsupervised domain adaptation (UDA) usually assumes that the source domain has labels and the target domain has no labels. In a real environment, labelled source domain data usually comes from multiple different distributions. To handle this problem, multi‐source unsupervised domain adaptation (MUDA) is proposed. Multi‐source unsupervised domain adaptation aims to adapt the model trained on multi‐labelled source domains to the unlabelled target domain. In this paper, a novel MUDA method by domain‐specific feature recalibration and alignment (FRA) is proposed. Specifically, to achieve feature recalibration, the authors leverage channel attention to pick out significant channels and spatial attention to focus on important features in different channels. Such integration of channel and spatial attention can lead to effective domain‐specific feature recalibration that may be of great importance to MUDA. In addition, to achieve better MUDA, the authors propose domain‐specific feature alignment which consists of Maximum Mean Discrepancy and JS‐divergence loss. Maximum Mean Discrepancy can reduce the difference between the source domain and target domain. Meanwhile, JS‐divergence loss may ensure the prediction consistency of different classifiers in the source domains. Four experiments have proved that FRA can achieve significantly better results in popular benchmarks for MUDA.
Mengzhu Wang, Dingyao Chen, Fangzhou Tan, Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Zhigang Luo
IET Comput. Vis.1
2023 Importance filtered soft label-based deep adaptation network
Wei Wang 0335, Mengzhu Wang, Zhihui Wang 0001
Knowl. Based Syst.3
2023 Boosting unsupervised domain adaptation: A Fourier approach
Mengzhu Wang, Shanshan Wang 0008, Ye Wang 0023, Wei Wang 0335, Tianyi Liang 0001, Junyang Chen 0001, Zhigang Luo
Knowl. Based Syst.1
2023 Video anomaly detection with NTCN-ML: A novel TCN for multi-instance learning
Wenhao Shao, Ruliang Xiao, Praboda Rajapaksha, Mengzhu Wang, Noël Crespi, Zhigang Luo, Roberto Minerva
Pattern Recognit.4
2023 Class-specific and self-learning local manifold structure for domain adaptation
Wei Wang 0335, Mengzhu Wang, Long Lan, Quannan Zu, Xiang Zhang 0008, Cong Wang 0018
Pattern Recognit.2
2023 Reducing bi-level feature redundancy for unsupervised domain adaptation
Mengzhu Wang, Shanshan Wang 0008, Wei Wang 0335, Li Shen 0008, Xiang Zhang 0008, Long Lan, Zhigang Luo
Pattern Recognit.1
2023 BP-triplet net for unsupervised domain adaptation: A Bayesian perspective
Shanshan Wang 0008, Lei Zhang 0038, Pichao Wang, Mengzhu Wang, Xingyi Zhang 0001
Pattern Recognit.4
2023 Bravely Say I Don't Know: Relational Question-Schema Graph for Text-to-SQL Answerability Classification
abstract
Recently, the Text-to-SQL task has received much attention. Many sophisticated neural models have been invented that achieve significant results. Most current work assumes that all the inputs are legal and the model should generate an SQL query for any input. However, in the real scenario, users are allowed to enter the arbitrary text that may not be answered by an SQL query. In this article, we focus on the issue–answerability classification for the Text-to-SQL system, which aims to distinguish the answerability of the question according to the given database schema. Existing methods concatenate the question and the database schema into a sentence, then fine-tune the pre-trained language model on the answerability classification task. In this way, the database schema is regarded as sequence text that may ignore the intrinsic structure relationship of the schema data, and the attention that represents the correlation between the question token and the database schema items is not well designed. To this end, we propose a relational Question-Schema graph framework that can effectively model the attention and relation between question and schema. In addition, a conditional layer normalization mechanism is employed to modulate the pre-trained language model to generate better question representation. Experiments demonstrate that the proposed framework outperforms all existing models by large margins, achieving new state of the art on the benchmark TRIAGESQL. Specifically, the model attains 88.41%, 78.24%, and 75.98% in Precision, Recall, and F1, respectively. Additionally, it outperforms the baseline by approximately 4.05% in Precision, 6.96% in Recall, and 6.01% in F1.
Wei Yu 0029, Mengzhu Wang, Xiaodong Wang 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 IRLM: Inductive Representation Learning Model for Personalized POI Recommendation
abstract
With the rapid development of the Internet of Things technology, the concept of smart cities that aims to help residents improve their quality of life has raised much attention in several application areas. In the context of smart cities, the provision of point of interest (POI) recommendations become an important requirement because a wide range of POIs are available for urban dwellers. Location-based social networks (LBSNs) such as Foursquare and Gowalla provide a massive volume of user check-in records that can assist users in choosing new POIs. However, user trajectories are mostly sparse in the real world. For example, users only check in a few POIs, and this makes it difficult to provide recommendations based on limited history trajectories. Though some attempts have adopted auxiliary geographical information to enhance POI recommendation, they still encounter the following problems: 1) the geographical trajectories of users are usually sparse in real-world datasets; 2) users may be more interested in the remote POIs; and 3) the previous models inherently perform transductive learning that cannot handle well the recommendation of unseen users and POIs. To address these problems, we propose an inductive representation learning model (IRLM) for location recommendation. IRLM contains two parts, namely geographic feature extraction and inductive representation learning. IRLM first captures global geographical influences among POIs through a standard Gaussian mixture model (GMM). Then IRLM adopts an attention neural network for the recommendation. Experimental results indicate that our proposed model can achieve superior performance over state-of-the-art models.
Junyang Chen 0001, Mengzhu Wang, Zhenghua Xu 0001, Xueliang Li 0002, Zhiguo Gong, Kaishun Wu, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.2
2023 A Closer Look at the Joint Training of Object Detection and Re-Identification in Multi-Object Tracking
abstract
Unifying object detection and re-identification (ReID) into a single network enables faster multi-object tracking (MOT), while this multi-task setting poses challenges for training. In this work, we dissect the joint training of detection and ReID from two dimensions: label assignment and loss function. We find previous works generally overlook them and directly borrow the practices from object detection, inevitably causing inferior performance. Specifically, we identify a qualified label assignment for MOT should: 1) have the assignment cost aware of ReID cost, not just detection cost; 2) provide sufficient positive samples for robust feature learning while avoiding ambiguous positives (i.e., the positives shared by different ground-truth objects). To achieve the above goals, we first propose Identity-aware Label Assignment, which jointly considers the assignment cost of detection and ReID to select positive samples for each instance without ambiguities. Moreover, we advance a novel Discriminative Focal Loss that integrates ReID predictions with Focal Loss to focus the training on the discriminative samples. Finally, we upgrade the strong baseline FairMOT with our techniques and achieve up to 7.0 MOTA / 54.1% IDs improvements on MOT16/17/20 benchmarks under favorable inference speed, which verifies our tailored label assignment and loss function for MOT are superior to those inherited from object detection.
Tianyi Liang 0001, Baopu Li, Mengzhu Wang, Huibin Tan, Zhigang Luo
IEEE Trans. Image Process.3
2023 OMG: Towards Effective Graph Classification Against Label Noise
abstract
Graph classification is a fundamental problem with diverse applications in bioinformatics and chemistry. Due to the intricate procedures of manual annotations in graphical domains, there may be abundant noisy labels of graphs in practice, resulting in poor performance for existing supervised methods. Thus, it is necessary and urgent to study the problem of graph classification with label noise. However, this problem is challenging due to the overfitting of noisy data as well as complicated relational structures of graphs. To handle this problem, we present a simple but effective approach called cOupledMix forGraph Contrast (OMG), which combines coupled Mixup with graph contrastive learning in the feature space. On the one hand, to improve the model generalization, we take convex combination of sample pairs in the feature space for positive pair construction. On the other hand, to accomplish effective optimization, we offer challenging negatives by multiple sample Mixup with different emphasis. To further reduce the impact of noisy data, we develop a neighbour-aware noise removal strategy, which promotes the smoothness in the neighbourhood of samples following the principle of curriculum learning. Extensive experiments on a range of benchmark datasets demonstrate the superiority of our proposed OMG.
Li Shen 0008, Mengzhu Wang, Xiao Luo 0001, Zhigang Luo, Dacheng Tao
IEEE Trans. Knowl. Data Eng.3
2022 Attention-based Adversarial Partial Domain Adaptation
abstract
With the rapid development of vision-based deep learning (DL), it is an effective method to generate large-scale synthetic data to supplement real data to train the DL models for domain adaptation. However, previous vanilla domain adaptation methods generally assume the same label space, and such an assumption is no longer valid for a more realistic scenario where it requires adaptation from a larger and more diverse source domain to a smaller target domain with less number of classes. To handle this problem, we propose an attention-based adversarial partial domain adaptation (AADA). Specifically, we leverage adversarial domain adaptation to augment the target domain by using source domain, then we can readily turn this task into a vanilla domain adaptation. Meanwhile, to accurately focus on the transferable features, we apply attention-based method to train the adversarial networks to obtain better transferable semantic features. Experiments on four benchmarks demonstrate that the proposed method outperforms existing methods by a large margin, especially on the tough domain adaptation tasks, e.g. VisDA-2017.
Mengzhu Wang, Shan An, Xiao Luo 0001, Xiong Peng, Wei Yu 0029, Junyang Chen 0001, Zhigang Luo
ICASSP1
2022 Logit Distillation via Student Diversity
Dingyao Chen, Long Lan, Mengzhu Wang, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICONIP (5)3
2022 AAT: Non-local Networks for Sim-to-Real Adversarial Augmentation Transfer
Mengzhu Wang, Shanshan Wang 0008, Tianwei Yan 0001, Zhigang Luo
ICONIP (4)1
2022 Implicit Feature Alignment For Knowledge Distillation
abstract
Knowledge distillation is a technique of transferring knowledge from a large teacher network to a light student one. Existing studies purely use immediate layers' features for distillation and may fail to gain insufficient semantic knowledge from the teacher. Inspired by recent advances in contrastive learning, we propose to introduce extra light embedding layers of the teacher to enforce its generalization ability and further align the mixup-type features for knowledge distillation in an implicit fashion (IFKD). IFKD allows the student to learn richer structural knowledge, thanks to the learned embedding layers of the teacher. Crucially, benefitting from a plethora of mixed samples, we can further adequately mine much semantic knowledge of the teacher. For efficiency, we propose a simple reversed mixup scheme to organize images and implicitly ensure complete positive information comparisons. Extensive experiments on image classification on two popular datasets including CIFAR-100 and ImageNet verify the effectiveness of our approach as compared to the previous methods.
Dingyao Chen, Mengzhu Wang, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICTAI2
2022 Semantic Data Augmentation based Distance Metric Learning for Domain Generalization
abstract
Domain generalization (DG) aims to learn a model on one or more different but related source domains that could be generalized into an unseen target domain. Existing DG methods try to prompt the diversity of source domains for the model's generalization ability, while they may have to introduce auxiliary networks or striking computational costs. On the contrary, this work applies the implicit semantic augmentation in feature space to capture the diversity of source domains. Concretely, an additional loss function of distance metric learning (DML) is included to optimize the local geometry of data distribution. Besides, the logits from cross entropy loss with infinite augmentations is adopted as input features for the DML loss in lieu of the deep features. We also provide a theoretical analysis to show that the logits can approximate the distances defined on original features well. Further, we provide an in-depth analysis of the mechanism and rational behind our approach, which gives us a better understanding of why leverage logits in lieu of features can help domain generalization. The proposed DML loss with the implicit augmentation is incorporated into a recent DG method, that is, Fourier Augmented Co-Teacher framework (FACT). Meanwhile, our method also can be easily plugged into various DG methods. Extensive experiments on three benchmarks (Digits-DG, PACS and Office-Home) have demonstrated that the proposed method is able to achieve the state-of-the-art performance.
Mengzhu Wang, Jianlong Yuan, Qi Qian 0001, Zhibin Wang 0004, Hao Li 0030
ACM Multimedia1
2022 DEAL: An Unsupervised Domain Adaptive Framework for Graph-level Classification
abstract
Graph neural networks (GNNs) have achieved state-of-the-art results on graph classification tasks. They have been primarily studied in cases of supervised end-to-end training, which requires abundant task-specific labels. Unfortunately, annotating labels of graph data could be prohibitively expensive or even impossible in many applications. An effective solution is to incorporate labeled graphs from a different, but related source domain, to develop a graph classification model for the target domain. However, the problem of unsupervised domain adaptation for graph classification is challenging due to potential domain discrepancy in graph space as well as the label scarcity in the target domain. In this paper, we present a novel GNN framework named DEAL by incorporating both source graphs and target graphs, which is featured by two modules, i.e., adversarial perturbation and pseudo-label distilling. Specifically, to overcome domain discrepancy, we equip source graphs with target semantics by applying to them adaptive perturbations which are adversarially trained against a domain discriminator. Additionally, DEAL explores distinct feature spaces at different layers of the GNN encoder, which emphasize global and local semantics respectively. Then, we distill the consistent predictions from two spaces to generate reliable pseudo-labels for sufficiently utilizing unlabeled data, which further improves the performance of graph classification. Extensive experiments on a wide range of graph classification datasets reveal the effectiveness of our proposed DEAL.
Li Shen 0008, Baopu Li, Mengzhu Wang, Xiao Luo 0001, Chong Chen 0002, Zhigang Luo, Xian-Sheng Hua 0001
ACM Multimedia4
2022 Informative pairs mining based adaptive metric learning for adversarial domain adaptation
Mengzhu Wang, Paul Li, Li Shen 0008, Ye Wang 0023, Shanshan Wang 0008, Wei Wang 0335, Xiang Zhang 0008, Junyang Chen 0001, Zhigang Luo
Neural Networks1
2022 Confidence Regularized Label Propagation Based Domain Adaptation
abstract
In domain adaptation (DA), label-induced losses generally occupy a dominant position and most previous models regard hard or soft labels as their inputs. However, these two types of labels may mislead the modeling process of label-induced losses since hard label is sensitive to a wrongly-predicted sample while soft label may introduce label noise, thus they may cause negative transfer. To relieve this problem, we propose a novel label learning approach namely confidence regularized label propagation (CRLP) that regularizes the confidence of predicted soft labels with constraints of F-norm or L21-norm. It is validated that maximizing either one of these two constraints equals to minimizing entropy loss. Specially, we illustrate that L21-norm is more suitable for DA than F-norm when the dataset contain a large number of categories. Then, we leverage the regularized soft labels produced by CRLP to reformulate some popular label-induced losses that consider feature transferability and discriminability such as class-wise maximum mean discrepancy, intra-class compactness and inter-class dispersion in a probability manner to present a novel DA method (i.e., CRLP-DA). Comprehensive analysis and experiments on four cross-domain object recognition datasets verify that the proposed CRLP-DA outperforms some state-of-the-art methods, especially 59.5% for Office10+Caltech10 dataset with SURF features. For others to better reproduce, our preliminary Matlab code will be available athttps://github.com/WWLoveTransfer/CRLP-DA/.
Wei Wang 0335, Baopu Li, Mengzhu Wang, Feiping Nie 0001, Zhihui Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Similar Questions Correspond to Similar SQL Queries: A Case-Based Reasoning Approach for Text-to-SQL Translation
Wei Yu 0029, Tao Chang, Mengzhu Wang, Xiaodong Wang 0002
ICCBR5
2021 A Focally Discriminative Loss for Unsupervised Domain Adaptation
Dongting Sun, Mengzhu Wang, Xurui Ma, Tianming Zhang, Wei Yu 0029, Zhigang Luo
ICONIP (1)2
2021 Spatio-Temporal Action Detector with Self-Attention
abstract
In the field of spatio-temporal action detection, some current studies attempt to solve the problem of action detection by using the one-stage object detectors based on anchor-free. Albeit efficiency, more performance boosts are expected. Towards this goal, a Self-Attention MovingCenter Detector (SAMOC) is proposed, which is blessed with two attractive aspects: 1) to effectively capture motion cues, a spatio-temporal self-attention block is explored to reinforce feature representation by aggregating motion-dependent global contexts, and 2) a link branch serves to model the frame-level object dependency, which promotes the confidence scores of correct actions. Experiments on two benchmark datasets show that SAMOC with the proposed two aspects achieves the state-of-the-art and works in real-time as well.
Xurui Ma, Zhigang Luo, Xiang Zhang 0008, Qing Liao 0001, Mengzhu Wang
IJCNN6
2021 InterBN: Channel Fusion for Adversarial Unsupervised Domain Adaptation
abstract
A classifier trained on one dataset rarely works on other datasets obtained under different conditions because of domain shifting. Such a problem is usually solved by domain adaptation methods. In this paper, we propose a novel unsupervised domain adaptation (UDA) method based on Interchangeable Batch Normalization (InterBN) to fuse different channels in deep neural networks for adversarial domain adaptation.Specifically, we first observe that the channels with small batch normalization scaling factor have less influence on the whole domain adaption, followed by a theoretical proof that the scaling factors for some channels will definitely come close to zero when imposing a sparsity regularization. Then, we replace the channels that have smaller scaling factors in the source domain with the mean of the channels which have larger scaling factors in the target domain or vice versa. Such a simple but effective channel fusion scheme can drastically increase the domain adaption ability.Extensive experimental results show that our InterBN significantly outperforms the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, InterBN achieves a remarkable improvement of 7.7% over the conditional adversarial adaptation networks (CDAN) on VisDA-2017 benchmark.
Mengzhu Wang, Wei Wang 0335, Baopu Li, Xiang Zhang 0008, Long Lan, Huibin Tan, Tianyi Liang 0001, Wei Yu 0029, Zhigang Luo
ACM Multimedia1
2021 An interaction-modeling mechanism for context-dependent Text-to-SQL translation based on heterogeneous graph aggregation
Wei Yu 0029, Tao Chang, Mengzhu Wang, Xiaodong Wang 0002
Neural Networks4