VLDB 2026 Research / reviewers in the wild / expert
Shanshan Wang 0008
dblp:62/3650-8
· DBLP profile ↗
46ranked-venue papers
22as first author
40since 2021 · last 2026
0000-0002-3824-687XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 13 first-author · 22 since 2021Artificial intelligence and machine learning · 21 · 9 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to Cluster Rare Cell Types: Implicit Semantic Data Augmentation for Spatial Multi-modal Omics AnalysisabstractSpatial multi-modal omics technologies have transformed biological research by enabling the simultaneous profiling of gene expression, protein abundance, and chromatin accessibility within their native spatial contexts. Despite these advances, accurately clustering rare cell types remains a major challenge due to data sparsity, high dimensionality, and limited annotated samples. While Graph Neural Networks (GNNs) have shown potential in modeling spatial omics data, their effectiveness is often constrained by the use of fixed K-nearest neighbor (KNN) graph structures, which fail to capture latent semantic relationships masked by sequencing noise. To overcome these limitations, we propose CRCT (Clustering Rare Cell Types): a novel framework that combines Implicit Semantic Data Augmentation (ISDA) with adaptive graph learning for spatial multi-modal omics analysis. Unlike traditional augmentation strategies that generate explicit synthetic samples, CRCT operates in the deep feature space by dynamically estimating intra-class covariance matrices and implicitly perturbing features along semantically meaningful directions. This enables effective augmentation for rare cell populations while preserving biological fidelity. Extensive experiments across four real-world datasets (HLN, MB, Stereo‑CITE‑seq, and SPOTS) and one synthetic benchmark demonstrate the state-of-the-art performance of CRCT, achieving improvements of up to +1.7 NMI and +7.8 ARI over strong baseline methods. Daixian Liu, Hau-Sing So, Shanshan Wang 0008, Mengzhu Wang, Jingcai Guo |
AAAI | 5 |
| 2026 | DeFT-LoRA: Decoupled and Fused Tuning with LoRA Experts for Universal Cross-Domain RetrievalabstractUniversal Cross-Domain Retrieval (UCDR) aims to retrieve images across unseen domains and categories, a critical capability for real-world applications. While large-scale Vision-Language Models (VLMs) like CLIP offer strong zero-shot category generalization, they struggle with domain shifts. Existing methods often improve domain robustness at the cost of high computational overhead or by compromising the VLM's inherent knowledge. To address this, we propose Decoupled and Fused Tuning with LoRA (DeFT-LoRA), a novel and parameter-efficient framework that integrates Low-Rank Adaptation (LoRA) with a Mixture-of-Experts (MoE) mechanism. This approach resolves the intrinsic conflict between domain-invariant and domain-specific knowledge in a single adapter, enabling our model to construct a domain adapters for each input image. We propose a three-stage training strategy, which first learns a shared Base LoRA for domain-invariant features, then derives Domain-Specific Experts to capture specific styles, and finally fuses them dynamically with a lightweight gating network. Extensive experiments on three UCDR benchmarks demonstrate that DeFT-LoRA achieves comparable or superior performance to state-of-the-art methods while requiring only 1.46 percent of CLIP's image-encoder parameters and reducing computational overhead, thereby establishing an exceptional balance between accuracy and efficiency. Ke Xu 0011, Xiaozheng Shen, Shanshan Wang 0008, Mengzhu Wang, Xun Yang 0001 |
AAAI | 3 |
| 2026 | Towards personalized long-term learning modeling in knowledge tracing
Shanshan Wang 0008, Jianqi Qiu, Jiaxin Pang, Xun Yang 0001, Ke Xu 0011, Zhangling Duan, Yuanhong Zhong, Xingyi Zhang 0001 |
Expert Syst. Appl. | 1 |
| 2026 | Gradually Vanishing Gap in Prototypical Network for unsupervised domain adaptation
Shanshan Wang 0008, Alusi, Hao Zhou 0001, Keyang Wang, Xun Yang 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2026 | Probability-Guided Contrastive Learning for Long-Tailed Domain GeneralizationabstractAfter training on a specific source domain, models can leverage domain generalization (DG) techniques to achieve superior and broader performance on new, unseen target domains. Existing DG often utilizes contrastive learning to learn domain-invariant features. The goal of contrastive learning is to learn effective representations of data, causing samples from the same category to cluster together in feature space, while samples from different categories are dispersed. Traditional contrastive learning is limited to a finite set of contrastive pairs for DG. To handle this problem, we consider sampling from an infinite number of contrastive pairs using a mixture of von Mises-Fisher (vMF) distributions on the unit hypersphere. We propose a novel method called Probability-guided Contrastive Learning (PgCL), which selects contrastive pairs based on estimated data distributions of samples from each category in feature space. Additionally, we derive the exact analytical formula for the expected contrastive loss. We conduct an empirical investigation of the error bounds of PgCL and demonstrate its performance by comparing it with several leading methods across a range of DG datasets. Mengzhu Wang, Houcheng Su, Shanshan Wang 0008, Long Lan, Liang Yang 0002, Li Shen 0008 |
IEEE Trans. Big Data | 3 |
| 2026 | Feature Responsive LoRA: Toward Parameter-Efficient Transfer Learning for Self-Supervised Visual ModelsabstractLow-Rank Adaptation (LoRA) is a widely utilized technique in topic of Parameter-Efficient Transfer Learning (PETL) which could use a limited number of trainable parameters to adapt the model to various downstream tasks. However, the setting of the locations and low-rank sizes in traditional LoRA relies heavily on the fixed and empirical values, which may hinder adaptability and lead to sharply decreasing performance, especially on some self-supervised pre-trained models. To alleviate this dilemma, we introduce a feature responsive LoRA (ResLoRA) method, a resource-efficient algorithm that automatically determines the LoRA modules’ required size based on the downstream task’s response. Firstly, we propose a Feature Decomposition loss (FD-loss) which leverages the feature singular values to mine the corresponding features of different downstream tasks, making the model parameters able to adequately represent downstream tasks. Subsequently, we leverage the Taylor expansion to measure the salience of the model parameters, then some high-efficient parameters with high significance could be leveraged to design a dynamically responsive LoRA. Specifically, the location and low-rank sizes of LoRA are determined based on the response parameters of the features for downstream tasks. Extensive experiments show that our ResLoRA achieves state-of-the-art performance, especially in the transfer capability of self-supervised models based on MoCo v3 and MAE. Our code is available at: https://github.com/wildboarman/ResLoRA. Shanshan Wang 0008, Xiaozheng Shen, Xun Yang 0001, Ke Xu 0011, Xingyi Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Local-Global Feature Fusion for Enhancing 3D Human Pose EstimationabstractBased on its excellent capability to extract temporal features, transformer has been widely used in monocular 3D human pose estimation. However, due to its global perspective, it performs inadequately in extracting spatial features, which hinders breakthroughs in performance. In this paper, we propose a local-global feature fusion method based on GCN and transformer for 3D human pose estimation. Our method integrates GCN with multiscale transformer to extract local spatiotemporal features of poses. These are then integrated with the global spatiotemporal features extracted by vanilla transformer to reconstruct 3D human poses accurately. In addition, we introduce a hierarchical feature fusion method to better capturing the underlying 3D pose structure. It blends deep abstract features with shallow raw features. We evaluate our model on the Human3.6M and MPI-INF-3DHP datasets, and experimental results demonstrate that our approach outperforms existing state-of-the-art methods. We achieve advanced performance on both datasets with errors of 37.7mm and 16.4mm under MPJPE, respectively. The code and model are available at https://github.com/ygx7/LG3DPose. Yuanhong Zhong, Guangxia Yang, Daidi Zhong, Xun Yang 0001, Shanshan Wang 0008, Zhangling Duan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Style-Aware Contrastive Test-Time Adaptation: A Dual-Cache Model for Robust Vision-Language AlignmentabstractTest-time adaptation (TTA) has emerged as a key strategy to enhance vision-language models (VLMs) under real-world distribution changes. However, existing methods always face two problems: 1) The fundamental trade-off dilemma: parameter-free TTA retains inference efficiency but fails to correct modality misalignment, while prompt tuning adapts to shifts, incurs high computational costs, and lacks knowledge retention. 2) Discriminative collapse also exists in TTA when faced with fine-grained downstream tasks. To alleviate these two bottlenecks, we introduce Style-aware Contrastive Test-Time Adaptation (SCTTA), a novel framework that jointly addresses modality misalignment and discriminative collapse. Firstly, we introduce Style-aware Embedding Adaptation (SEA), which dynamically refines text embeddings by incorporating domain-specific style attributes, improving alignment between visual and textual modalities. Secondly, we propose Fine-grained Contrastive Adaptation (FCA), which enhances feature separation by enforcing contrastive learning with adaptive prototypes, reducing inter-class feature overlap in fine-grained tasks. In addition, we introduce Dual-Cache Model (DCM), which extends prior unimodal cache model to a multimodal cache for the first time. Eventually, it accumulates adaptation knowledge through a visual-cache (capturing evolving domain styles) and a textual-cache (retaining discriminative semantics), enabling long-term adaptation without additional overhead. Extensive experiments on 15 datasets demonstrate that our approach achieves state-of-the-art performance for both fine-grained and out-of-distribution dataset benchmarks. Furthermore, SCTTA continuously improves as more test samples accumulate, validating its sustainable adaptation capacity. Our code is available at https://github.com/alusi123/SCTTA. Shanshan Wang 0008, ALuSi, Xun Yang 0001, Pichao Wang, Ke Xu 0011, Xingyi Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | TypiCD: Cognitive Diagnosis via Problem-Type-Guided Bias Correction
Shanshan Wang 0008, Yali Ye, Xun Yang 0001, Pichao Wang, Mengzhu Wang, Xingyi Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | ASCD: An Adaptive Framework for Long-Tailed Cognitive DiagnosisabstractCognitive Diagnosis Modeling is a fundamental task in intelligent education, intending to assess students’ mastery levels on knowledge concepts through interactions. Previous methodologies prioritized enhancing average diagnostic accuracy. However, they often neglected the Long-Tailed issue in interactions. To address this issue, we propose an Adaptive Self-Supervised Graph Learning for Cognitive Diagnosis (ASCD) framework that leverages relatively balanced sparse views to forcing the graph network to focus on long-tailed nodes, aiming to tackle the long-tailed problem in graph-based cognitive diagnosis. Additionally, we employ self-supervised manners to mitigate the impact of dropped information on other nodes. Our approach leverages the adaptive graph confusion method to create sparse views of the original student-exercise interaction graph. In these sparse views, both long-tailed and head students carry similar weights, pushing the graph network to allocate attention impartially across all students. We integrate two different graph confusion techniques into adaptive graph confusion to accommodate varying degrees of data sparsity: edge dropout and feature masking. ASCD can serve as a plug-and-play module integrated into any graph-based cognitive diagnosis model, enhancing its performance regarding long-tailed scenarios. Extensive experiments on real-world datasets show the effectiveness of our approach, especially on the students with much sparser interaction records. Shanshan Wang 0008, Xun Yang 0001, Yuanhong Zhong, Xingyi Zhang 0001, Meng Wang 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2025 | ESBN: Estimation Shift of Batch Normalization for Source-free Universal Domain AdaptationabstractDomain adaptation (DA) is crucial for transferring models trained in one domain to perform well in a different, often unseen domain. Traditional methods, including unsupervised domain adaptation (UDA) and source-free domain adaptation (SFDA), have made significant progress. However, most existing DA methods rely heavily on Batch Normalization (BN) layers, which are not optimal in source-free settings, where the source domain is unavailable for comparison. In this study, we propose a novel method, ESBN, which addresses the challenge of domain shift by adjusting the placement of normalization layers and replacing BN with Batch-free Normalization (BFN). Unlike BN, BFN is less dependent on batch statistics and provides more robust feature representations through instance-specific statistics. We systematically investigate the effects of different BN layer placements across various network configurations and demonstrate that selective replacement with BFN improves generalization performance. Extensive experiments on multiple domain adaptation benchmarks show that our approach outperforms state-of-the-art methods, particularly in challenging scenarios such as Open-Partial Domain Adaptation (OPDA). Houcheng Su, Bingli Wang, Yuandong Min, Mengzhu Wang, Shanshan Wang 0008, Jingcai Guo |
IJCAI | 7 |
| 2025 | Multimodal Adapter-Driven Source-Free Domain Adaptation
Shanshan Wang 0008, Houmeng He, Keyang Wang, Xun Yang 0001 |
PRCV (8) | 1 |
| 2025 | Learning states enhanced Knowledge Tracing: Simulating the diversity in real-world learning process
Shanshan Wang 0008, Xun Yang 0001, Xingyi Zhang 0001, Keyang Wang |
Expert Syst. Appl. | 1 |
| 2025 | Graph Convolutional Mixture-of-Experts Learner Network for Long-Tailed Domain GeneralizationabstractThe goal of single domain generalization is to use data from a single domain (source domain) to train a model, which is then deployed over several unknown domains for testing (target domains). This study introduces a practical approach diverging from traditional DG, which typically relies on multiple source domains. We focus on Single Long-Tailed Domain Generalization, which refers to a scenario in the context of long-tail distribution, where although minority classes may have fewer samples in a single domain, these minority classes could become more prevalent and dominant in other domains. We introduce the Graph Convolutional Mixture-of-Experts Learners Network for Long-Tailed Domain Generalization (GCML) as a solution to this problem. Our approach presents two novel tactics. Initially, we utilize an expert learning technique that is skill-diverse. In order to properly manage the unknown target domain, this entails training multiple specialists inside a single long-tailed source domain and combining their knowledge. Then, we use a graph convolutional network to facilitate domain generalization, leveraging joint data structure modeling to learn more domain-invariant feature. Experiments conducted on four established benchmarks reveal that our GCML algorithm outperforms contemporary domain generalization techniques, demonstrating its efficacy in this complex task. Mengzhu Wang, Houcheng Su, Shanshan Wang 0008, Li Shen 0008, Long Lan, Liang Yang 0002, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Exploring Invariance Matters for Domain GeneralizationabstractDomain generalization (DG) aims to solve the problem of significant performance degradation when target domain data collected from the Out-Of-Distribution (O.O.D). Previous efforts try to exploit invariant features in the source domain through CNN networks. However, inspired by causal mechanisms, we find that the complex spurious-invariant information is still hidden in this view invariant features, and the impact of domain and class discrepancies on extracting invariance has not been effectively mitigated. To alleviate these issues, we propose a self-weighted multi-view mining invariance domain generalization framework (SMIDG). On the one hand, to make up for the insufficiency of traditional single-view convolutional feature extraction networks, we propose to mine features from another frequency view and use the self-adaptive adversarial masks to eliminate some spurious correlations, ensuring causal invariance in the coarse-grained generalization. However, due to inconsistencies in discriminative information between inter-domain and intra-domain samples, as well as inter-class and intra-class samples, the coarse-grained elimination of spurious associations does not fully resolve this issue. On the other hand, we also consider the fine-grained generalization from two aspects. Firstly, to better tackle the domain discrepancies, we propose a novel progressive contrastive learning strategy that learns the underlying specific features of samples while gradually mitigating domain discrepancies, thereby ensuring domain invariance in fine-grained generalization. Secondly, due to the issue of feature inconsistency, we adopt a self-adaptive hard sample mining method with information gain to ensure that the model pays more attention on hard disentangled samples, thus maintaining feature invariance. Extensive experiments on five benchmark datasets demonstrate that our method outperforms state-of-the-art approaches. Our code is available at https://github.com/bihhm/SMIDG. Shanshan Wang 0008, Houmeng He, Xun Yang 0001, Zhipu Liu, Yuanhong Zhong, Xingyi Zhang 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Dual-State Personalized Knowledge Tracing With Emotional IncorporationabstractKnowledge tracing has been widely used in online learning systems to guide the students’ future learning. However, most existing KT models primarily focus on extracting abundant information from the question sets and explore the relationships between them, but ignore the personalized student behavioral information in the learning process. This will limit the model’s ability to accurately capture the personalized knowledge states of students and reasonably predict their performances. To alleviate this limitation, we explicitly models the personalized learning process by incorporating the emotions, a representative personalized behavior in the learning process, into KT framework. Specifically, we present a novel Dual-State Personalized Knowledge Tracing with Emotional Incorporation model to achieve this goal: First, we incorporate emotional information into the modeling process of knowledge state, resulting in the Knowledge State Boosting Module. Second, we design an Emotional State Tracing Module to monitor students’ personalized emotional states, and propose an emotion prediction method based on personalized emotional states. Finally, we apply the predicted emotions to enhance students’ response prediction. Furthermore, to extend the generalization capability of our model across different datasets, we design a transferred version of DEKT, named Transfer Learning-based Self-loop model (T-DEKT). Extensive experiments show our method achieves the state-of-the-art performance. Shanshan Wang 0008, Fangzheng Yuan, Keyang Wang, Xun Yang 0001, Xingyi Zhang 0001, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | UniAda: Domain Unifying and Adapting Network for Generalizable Medical Image SegmentationabstractLearning a generalizable medical image segmentation model is an important but challenging task since the unseen (testing) domains may have significant discrepancies from seen (training) domains due to different vendors and scanning protocols. Existing segmentation methods, typically built upon domain generalization (DG), aim to learn multi-source domain-invariant features through data or feature augmentation techniques, but the resulting models either fail to characterize global domains during training or cannot sense unseen domain information during testing. To tackle these challenges, we propose a domain Unifying and Adapting network (UniAda) for generalizable medical image segmentation, a novel "unifying while training, adapting while testing" paradigm that can learn a domain-aware base model during training and dynamically adapt it to unseen target domains during testing. First, we propose to unify the multi-source domains into a global inter-source domain via a novel feature statistics update mechanism, which can sample new features for the unseen domains, facilitating the training of a domain base model. Second, we leverage the uncertainty map to guide the adaptation of the trained model for each testing sample, considering the specific target domain may be outside the global inter-source domain. Extensive experimental results on two public cross-domain medical datasets and one in-house cross-domain dataset demonstrate the strong generalization capacity of the proposed UniAda over state-of-the-art DG methods. The source code of our UniAda is available at https://github.com/ZhouZhang233/UniAda. Zhongzhou Zhang, Zhiwen Wang 0002, Shanshan Wang 0008, Fenglei Fan, Hongming Shan, Yi Zhang 0018 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Smooth-Guided Implicit Data Augmentation for Domain GeneralizationabstractThe training process of a domain generalization (DG) model involves utilizing one or more interrelated source domains to attain optimal performance on an unseen target domain. Existing DG methods often use auxiliary networks or require high computational costs to improve the model's generalization ability by incorporating a diverse set of source domains. In contrast, this work proposes a method called Smooth-Guided Implicit Data Augmentation (SGIDA) that operates in the feature space to capture the diversity of source domains. To amplify the model's generalization capacity, a distance metric learning (DML) loss function is incorporated. Additionally, rather than depending on deep features, the suggested approach employs logits produced from cross entropy (CE) losses with infinite augmentations. A theoretical analysis shows that logits are effective in estimating distances defined on original features, and the proposed approach is thoroughly analyzed to provide a better understanding of why logits are beneficial for DG. Moreover, to increase the diversity of the source domain, a sampling-based method called smooth is introduced to obtain semantic directions from interclass relations. The effectiveness of the proposed approach is demonstrated through extensive experiments on widely used DG, object detection, and remote sensing datasets, where it achieves significant improvements over existing state-of-the-art methods across various backbone networks. Mengzhu Wang, Junze Liu, Ge Luo 0003, Shanshan Wang 0008, Wei Wang 0335, Long Lan, Ye Wang 0023, Feiping Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Boosting Neural Cognitive Diagnosis with Student's Affective State ModelingabstractCognitive Diagnosis Modeling aims to infer students' proficiency level on knowledge concepts from their response logs. Existing methods typically model students’ response processes as the interaction between students and exercises or concepts based on hand-crafted or deeply-learned interaction functions. Despite their promising achievements, they fail to consider the relationship between students' cognitive states and affective states in learning, e.g., the feelings of frustration, boredom, or confusion with the learning content, which is insufficient for comprehensive cognitive diagnosis in intelligent education. To fill the research gap, we propose a novel Affect-aware Cognitive Diagnosis (ACD) model which can effectively diagnose the knowledge proficiency levels of students by taking into consideration the affective factors. Specifically, we first design a student affect perception module under the assumption that the affective state is jointly influenced by the student's affect trait and the difficulty of the exercise. Then, our inferred affective distribution is further used to estimate the student's subjective factors, i.e., guessing and slipping, respectively. Finally, we integrate the estimated guessing and slipping parameters with the basic neural cognitive diagnosis framework based on the DINA model, which facilitates the modeling of complex exercising interactions in a more accurate and interpretable fashion. Besides, we also extend our affect perception module in an unsupervised learning setting based on contrastive learning, thus significantly improving the compatibility of our ACD. To the best of our knowledge, we are the first to unify the cognition modeling and affect modeling into the same framework for student cognitive diagnosis. Extensive experiments on real-world datasets clearly demonstrate the effectiveness of our ACD. Our code is available at https://github.com/zeng-zhen/ACD. Shanshan Wang 0008, Xun Yang 0001, Ke Xu 0011, Xingyi Zhang 0001 |
AAAI | 1 |
| 2024 | PTMQ: Post-training Multi-Bit Quantization of Neural NetworksabstractThe ability of model quantization with arbitrary bit-width to dynamically meet diverse bit-width requirements during runtime has attracted significant attention. Recent research has focused on optimizing large-scale training methods to achieve robust bit-width adaptation, which is a time-consuming process requiring hundreds of GPU hours. Furthermore, converting bit-widths requires recalculating statistical parameters of the norm layers, thereby impeding real-time switching of the bit-width. To overcome these challenges, we propose an efficient Post-Training Multi-bit Quantization (PTMQ) scheme that requires only a small amount of calibration data to perform block-wise reconstruction of multi-bit quantization errors. It eliminates the influence of statistical parameters by fusing norm layers, and supports real-time switching bit-widths in uniform quantization and mixed-precision quantization. To improve quantization accuracy and robustness, we propose a Multi-bit Feature Mixer technique (MFM) for fusing features of different bit-widths to enhance robustness across varying bit-widths. Moreover, we introduced the Group-wise Distillation Loss (GD-Loss) to enhance the correlation between different bit-width groups and further improve the overall performance of PTMQ. Extensive experiments demonstrate that PTMQ achieves comparable performance to existing state-of-the-art post-training quantization methods, while optimizing it speeds up by 100$\times$ compared to recent multi-bit quantization works. Code can be available at https://github.com/xuke225/PTMQ. Ke Xu 0011, Zhongcheng Li, Shanshan Wang 0008, Xingyi Zhang 0001 |
AAAI | 3 |
| 2024 | SBM: Smoothness-Based Minimization for Domain GeneralizationabstractIn topical domain generalization (DG), trained models are asked to perform well on an unknown target domain with different data statistics. In order to improve domain generalization, adversarial learning has proven to be one of the most effective methods. Existing approaches, however, rely primarily on adversarial learning, which can only generalize within a limited range of domains. We argue that smoothness- based minimization (SBM) is a more promising direction for adversarial domain generalization. Our findings indicate that achieving a smoothness-based minimization of task loss stabilizes adversarial training, resulting in better domain generalization performance. This method has been shown to achieve remarkable domain generalization performance on three publicly available benchmarks including PACS, Office- Home and DomainNet. Chunqing Ruan, Mengzhu Wang, Shanshan Wang 0008, Tianyi Liang 0001, Wei Yu 0029 |
ICASSP | 3 |
| 2024 | Self-Training with Contrastive Learning for Adversarial Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to transfer the knowledge learned from the labeled source domain to the unlabeled target domain. Traditional methods often focus on minimizing the distribution gap between feature spaces of the two domains to achieve domain-invariant representations. However, this approach may not fully leverage the inherent class-specific information and could adversely affect decision boundaries in the target domain. In this paper, we introduce an adversarial contrastive self-training framework. The model integrates both intra-class and inter-class domain differences for better alignment between the source and target domains. Initially, we augment target domain samples using Fourier transformations, obtaining reliable samples and pseudo-labels through self-training. We then employ intra-domain contrastive learning on the reliable samples and pseudo-labels, alongside labeled source domain samples, to achieve intra-domain compactness. Finally, we utilize adversarial learning to achieve global alignment. This comprehensive approach ensures improved adaptation performance while also addressing the limitations of previous methods. Shanshan Wang 0008, Minbin Hu, Xingyi Zhang 0001 |
IJCNN | 1 |
| 2024 | Dual-stream Feature Augmentation for Domain Generalization
Shanshan Wang 0008, ALuSi, Xun Yang 0001, Ke Xu 0011, Huibin Tan, Xingyi Zhang 0001 |
ACM Multimedia | 1 |
| 2024 | Robust video question answering via contrastive cross-modality representation learning
Xun Yang 0001, Jianming Zeng, Dan Guo 0001, Shanshan Wang 0008, Jianfeng Dong, Meng Wang 0001 |
Sci. China Inf. Sci. | 4 |
| 2024 | Learning Hierarchical Visual Transformation for Domain Generalizable Visual Matching and Recognition
Xun Yang 0001, Tianzhu Zhang 0001, Shanshan Wang 0008, Richang Hong, Meng Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | FaSRnet: a feature and semantics refinement network for human pose estimationabstractDue to factors such as motion blur, video out-of-focus, and occlusion, multi-frame human pose estimation is a challenging task. Exploiting temporal consistency between consecutive frames is an efficient approach for addressing this issue. Currently, most methods explore temporal consistency through refinements of the final heatmaps. The heatmaps contain the semantics information of key points, and can improve the detection quality to a certain extent. However, they are generated by features, and feature-level refinements are rarely considered. In this paper, we propose a human pose estimation framework with refinements at the feature and semantics levels. We align auxiliary features with the features of the current frame to reduce the loss caused by different feature distributions. An attention mechanism is then used to fuse auxiliary features with current features. In terms of semantics, we use the difference information between adjacent heatmaps as auxiliary features to refine the current heatmaps. The method is validated on the large-scale benchmark datasets PoseTrack2017 and PoseTrack2018, and the results demonstrate the effectiveness of our method. Yuanhong Zhong, Qianfeng Xu, Daidi Zhong, Xun Yang 0001, Shanshan Wang 0008 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2024 | Mutual-weighted feature disentanglement for unsupervised domain adaptation
Shanshan Wang 0008, Keyang Wang, Xun Yang 0001, Xingyi Zhang 0001 |
Multim. Syst. | 1 |
| 2024 | Equity in Unsupervised Domain Adaptation by Nuclear Norm MaximizationabstractNuclear norm maximization has shown the power to enhance the transferability of unsupervised domain adaptation model (UDA) in an empirical scheme. In this paper, we identify a new property termedequity, which indicates the balance degree of predicted classes, to demystify the efficacy of nuclear norm maximization for UDA theoretically. With this in mind, we offer a new discriminability-and-equity maximization paradigm built on squares loss, such that predictions are equalized explicitly. To verify its feasibility and flexibility, two new losses termed Class Weighted Squares Maximization (CWSM) and Normalized Squares Maximization (NSM), are proposed to maximize both predictive discriminability and equity, from the class level and the sample level, respectively. Importantly, we theoretically relate these two novel losses (i.e., CWSM and NSM) to the equity maximization under mild conditions, and empirically suggest the importance of the predictive equity in UDA. Moreover, it is very efficient to realize the equity constraints in both losses. Experiments of cross-domain image classification on three popular benchmark datasets show that both CWSM and NSM contribute to outperforming the corresponding counterparts. Mengzhu Wang, Shanshan Wang 0008, Xun Yang 0001, Jianlong Yuan, Wenju Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Inter-Class and Inter-Domain Semantic Augmentation for Domain GeneralizationabstractThe domain generalization approach seeks to develop a universal model that performs well on unknown target domains with the aid of diverse source domains. Data augmentation has proven to be an effective method to enhance domain generalization in computer vision. Recently, semantic-level based data augmentation has yielded remarkable results. However, these methods focus on sampling semantic directions on feature space from intra-class and intra-domain, limiting the diversity of the source domain. To address this issue, we propose a novel approach called Inter-Class and Inter-Domain Semantic Augmentation (CDSA) for domain generalization. We first introduce a sampling-based method called CrossSmooth to obtain semantic directions from inter-class. Then, CrossVariance obtains the styles of different domains by sampling semantic directions. Our experiments on four well-known domain generalization benchmark datasets (Digits-DG, PACS, Office-Home, and DomainNet) demonstrate the effectiveness of our approach. We also validate our approach on commonly-used semantic segmentation datasets, namely GTAV, SYNTHIA, Cityscapes, Mapillary, and BDDS which also show significant improvements. Mengzhu Wang, Yuehua Liu, Jianlong Yuan, Shanshan Wang 0008, Zhibin Wang 0004, Wei Wang 0335 |
IEEE Trans. Image Process. | 4 |
| 2024 | Frame-Padded Multiscale Transformer for Monocular 3D Human Pose EstimationabstractMonocular 3D human pose estimation is an ill-posed problem in computer vision due to its depth ambiguity. Most existing works supplement the depth information by extracting temporal pose features from video frames, and they have made notable progress. However, these approaches divide a long sequence of video frames into multiple short sequences for separate processing, which leads to the loss of complementary information between sequences. Furthermore, the short-term temporal correlation among frames in a sequence is often not fully exploited. To model temporal dependencies efficiently, we propose the frame-padded multiscale transformer approach, which includes a frame-padded video sequence preprocessing step and a multiscale temporal transformer backbone. Our approach addresses the omission of the temporal features of edge frames in existing approaches by padding video frames in the shallow layer. In addition, we extract the temporal information of 3D human poses using a multiscale transformer to enhance the short-term correlation of human pose skeleton keypoints. Extensive experiments validate the effectiveness of our approach on two popular datasets: Human3.6M and MPI-INF-3DHP. The results show that our approach achieves state-of-the-art performance. Yuanhong Zhong, Guangxia Yang, Daidi Zhong, Xun Yang 0001, Shanshan Wang 0008 |
IEEE Trans. Multim. | 5 |
| 2024 | Video Compressed Sensing Reconstruction via an Untrained Network with Low-Rank RegularizationabstractDeep image prior (DIP) is an emerging technology that indicates that the structure of an untrained network can serve as an excellent prior for image restoration. It bridges the gap between training-based and training-free methods and exhibits considerable potential in image compressed sensing (CS) reconstruction. In this article, we extend DIP and propose a novel Low-Rank Regularization Video Compressed Sensing Network for CS video reconstruction (dubbed LRR-VCSNet). We explore the application of a low-rank latent tensor with an untrained network for global low-rank regularization on video reconstruction, and the interframe low-rank approximation for framewise nonlocal low-rank regularization in the data space is also exploited. In addition, we design the structure of the untrained network based on the encoder-decoder architecture to improve the performance. Extensive experiments on six standard CIF video sequences show that LLR-VCSNet significantly outperforms traditional video CS methods and achieves competitive results when compared with the state-of-the-art training-based video CS method. Yuanhong Zhong, Chenxu Zhang 0001, Xun Yang 0001, Shanshan Wang 0008 |
IEEE Trans. Multim. | 4 |
| 2023 | Self-Supervised Graph Learning for Long-Tailed Cognitive DiagnosisabstractCognitive diagnosis is a fundamental yet critical research task in the field of intelligent education, which aims to discover the proficiency level of different students on specific knowledge concepts. Despite the effectiveness of existing efforts, previous methods always considered the mastery level on the whole students, so they still suffer from the Long Tail Effect. A large number of students who have sparse interaction records are usually wrongly diagnosed during inference. To relieve the situation, we proposed a Self-supervised Cognitive Diagnosis (SCD) framework which leverages the self-supervised manner to assist the graph-based cognitive diagnosis, then the performance on those students with sparse data can be improved. Specifically, we came up with a graph confusion method that drops edges under some special rules to generate different sparse views of the graph. By maximizing the cross-view consistency of node representations, our model could pay more attention on long-tailed students. Additionally, we proposed an importance-based view generation rule to improve the influence of long-tailed students. Extensive experiments on real-world datasets show the effectiveness of our approach, especially on the students with much sparser interaction records. Our code is available at https://github.com/zeng-zhen/SCD. Shanshan Wang 0008, Xun Yang 0001, Xingyi Zhang 0001 |
AAAI | 1 |
| 2023 | Disentangled Representation Learning with Causality for Unsupervised Domain AdaptationabstractMost efforts in unsupervised domain adaptation (UDA) focus on learning the domain-invariant representations between the two domains. However, such representations may still confuse two patterns due to the domain gap. Considering that semantic information is useful for the final task and domain information always indicates the discrepancy between two domains, to address this issue, we propose to decouple the representations of semantic features from domain features to reduce domain bias. Different from traditional methods, we adopt a simple but effective module with only one domain discriminator to decouple the representations, offering two benefits. Firstly, it eliminates the need for labeled sample pairs, making it more suitable for UDA. Secondly, without adversarial learning, our model can achieve a more stable training phase. Moreover, to further enhance the task-specific features, we employ a causal mechanism to separate semantic features related to causal factors from the overall feature representations. Specially, we utilize a dual-classifier strategy, where each classifier is fed with the entire features and the semantic features, respectively. By minimizing the discrepancy between the outputs of the two classifiers, the causal influence of the semantic features is accentuated. Experiments on several public datasets demonstrate the proposed model can outperform the state-of-the-art methods. Our code is available at: https://github.com/qzxRtY37/DRLC https://github.com/qzxRtY37/DRLC. Shanshan Wang 0008, Zhenwei He, Xun Yang 0001, Mengzhu Wang, Quanzeng You, Xingyi Zhang 0001 |
ACM Multimedia | 1 |
| 2023 | Boosting unsupervised domain adaptation: A Fourier approach
Mengzhu Wang, Shanshan Wang 0008, Ye Wang 0023, Wei Wang 0335, Tianyi Liang 0001, Junyang Chen 0001, Zhigang Luo |
Knowl. Based Syst. | 2 |
| 2023 | Reducing bi-level feature redundancy for unsupervised domain adaptation
Mengzhu Wang, Shanshan Wang 0008, Wei Wang 0335, Li Shen 0008, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Pattern Recognit. | 2 |
| 2023 | BP-triplet net for unsupervised domain adaptation: A Bayesian perspective
Shanshan Wang 0008, Lei Zhang 0038, Pichao Wang, Mengzhu Wang, Xingyi Zhang 0001 |
Pattern Recognit. | 1 |
| 2022 | AAT: Non-local Networks for Sim-to-Real Adversarial Augmentation Transfer
Mengzhu Wang, Shanshan Wang 0008, Tianwei Yan 0001, Zhigang Luo |
ICONIP (4) | 2 |
| 2022 | Refining pseudo labels for unsupervised Domain Adaptive Re-Identification
Shanshan Wang 0008, Lei Zhang 0038, Fan Wang 0019, Hao Li 0030 |
Knowl. Based Syst. | 1 |
| 2022 | Informative pairs mining based adaptive metric learning for adversarial domain adaptation
Mengzhu Wang, Paul Li, Li Shen 0008, Ye Wang 0023, Shanshan Wang 0008, Wei Wang 0335, Xiang Zhang 0008, Junyang Chen 0001, Zhigang Luo |
Neural Networks | 5 |
| 2022 | Video Moment Retrieval With Cross-Modal Neural Architecture SearchabstractThe task of video moment retrieval (VMR) is to retrieve the specific video moment from an untrimmed video, according to a textual query. It is a challenging task that requires effective modeling of complex cross-modal matching relationship. Recent efforts primarily model the cross-modal interactions by hand-crafted network architectures. Despite their effectiveness, they rely heavily on expert experience to select architectures and have numerous hyperparameters that need to be carefully tuned, which significantly limit their applications in real-world scenarios. How to design flexible architectures for modeling cross-modal interactions with less manual effort is crucial for the task of VMR but has received limited attention so far. To address this issue, we present a novel VMR approach that automatically searches for an optimal architecture to learn cross-modal matching relationship. Specifically, we develop a cross-modal architecture searching method. It first searches for repeatable cell network architectures based on a directed acyclic graph, which performs operation sampling over a customized task-specific operation set. Then, we adaptively modulate the edge importance in the graph by a query-aware attention network, which performs edge sampling softly in the searched cell. Different from existing neural architecture search methods, our approach can effectively exploit the query information to reach query-conditioned architectures for modeling cross modal matching. Extensive experiments on three benchmark datasets show that our approach can not only significantly outperform the state-of-the-art methods but also run more efficiently and robustly than manually crafted network architectures. Xun Yang 0001, Shanshan Wang 0008, Jian Dong 0011, Jianfeng Dong, Meng Wang 0001, Tat-Seng Chua |
IEEE Trans. Image Process. | 2 |
| 2020 | Self-adaptive Re-weighted Adversarial Domain AdaptationabstractExisting adversarial domain adaptation methods mainly consider the marginal distribution and these methods may lead to either under transfer or negative transfer. To address this problem, we present a self-adaptive re-weighted adversarial domain adaptation approach, which tries to enhance domain alignment from the perspective of conditional distribution. In order to promote positive transfer and combat negative transfer, we reduce the weight of the adversarial loss for aligned features while increasing the adversarial force for those poorly aligned measured by the conditional entropy. Additionally, triplet loss leveraging source samples and pseudo-labeled target samples is employed on the confusing domain. Such metric loss ensures the distance of the intra-class sample pairs closer than the inter-class pairs to achieve the class-level alignment. In this way, the high accurate pseudolabeled target samples and semantic alignment can be captured simultaneously in the co-training process. Our method achieved low joint error of the ideal source and target hypothesis. The expected target error can then be upper bounded following Ben-David’s theorem. Empirical evidence demonstrates that the proposed model outperforms state of the arts on standard domain adaptation datasets. Shanshan Wang 0008, Lei Zhang 0038 |
IJCAI | 1 |
| 2020 | Adversarial transfer learning for cross-domain visual recognition
Shanshan Wang 0008, Lei Zhang 0038, Jingru Fu |
Knowl. Based Syst. | 1 |
| 2020 | Class-Specific Reconstruction Transfer Learning for Visual Recognition Across DomainsabstractSubspace learning and reconstruction have been widely explored in recent transfer learning work. Generally, a specially designed projection and reconstruction transfer functions bridging multiple domains for heterogeneous knowledge sharing are wanted. However, we argue that the existing subspace reconstruction based domain adaptation algorithms neglect the class prior, such that the learned transfer function is biased, especially when data scarcity of some class is encountered. Different from those previous methods, in this paper, we propose a novel class-wise reconstruction-based adaptation method called Class-specific Reconstruction Transfer Learning (CRTL), which optimizes a well modeled transfer loss function by fully exploiting intra-class dependency and inter-class independency. The merits of the CRTL are three-fold. 1) Using a class-specific reconstruction matrix to align the source domain with the target domain fully exploits the class prior in modeling the domain distribution consistency, which benefits the cross-domain classification. 2) Furthermore, to keep the intrinsic relationship between data and labels after feature augmentation, a projected Hilbert-Schmidt Independence Criterion (pHSIC), that measures the dependency between data and label, is first proposed in transfer learning community by mapping the data from raw space to RKHS. 3) In addition, by imposing low-rank and sparse constraints on the class-specific reconstruction coefficient matrix, the global and local data structure that contributes to domain correlation can be effectively preserved. Extensive experiments on challenging benchmark datasets demonstrate the superiority of the proposed method over state-of-the-art representation-based domain adaptation methods. The demo code is available in https://github.com/wangshanshanCQU/CRTL. Shanshan Wang 0008, Lei Zhang 0038, Wangmeng Zuo, Bob Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Guide Subspace Learning for Unsupervised Domain AdaptationabstractA prevailing problem in many machine learning tasks is that the training (i.e., source domain) and test data (i.e., target domain) have different distribution [i.e., non-independent identical distribution (i.i.d.)]. Unsupervised domain adaptation (UDA) was proposed to learn the unlabeled target data by leveraging the labeled source data. In this article, we propose a guide subspace learning (GSL) method for UDA, in which an invariant, discriminative, and domain-agnostic subspace is learned by three guidance terms through a two-stage progressive training strategy. First, the subspace-guided term reduces the discrepancy between the domains by moving the source closer to the target subspace. Second, the data-guided term uses the coupled projections to map both domains to a unified subspace, where each target sample can be represented by the source samples with a low-rank coefficient matrix that can preserve the global structure of data. In this way, the data from both domains can be well interlaced and the domain-invariant features can be obtained. Third, for improving the discrimination of the subspaces, the label-guided term is constructed for prediction based on source labels and pseudo-target labels. To further improve the model tolerance to label noise, a label relaxation matrix is introduced. For the solver, a two-stage learning strategy with teacher teaches and student feedbacks mode is proposed to obtain the discriminative domain-agnostic subspace. In addition, for handling nonlinear domain shift, a nonlinear GSL (NGSL) framework is formulated with kernel embedding, such that the unified subspace is imposed with nonlinearity. Experiments on various cross-domain visual benchmark databases show that our methods outperform many state-of-the-art UDA methods. The source code is available at https://github.com/Fjr9516/GSL. Lei Zhang 0038, Jingru Fu, Shanshan Wang 0008, David Zhang 0001, Zhao Yang Dong, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Manifold Criterion Guided Transfer Learning via Intermediate Domain GenerationabstractIn many practical transfer learning scenarios, the feature distribution is different across the source and target domains (i.e., nonindependent identical distribution). Maximum mean discrepancy (MMD), as a domain discrepancy metric, has achieved promising performance in unsupervised domain adaptation (DA). We argue that the MMD-based DA methods ignore the data locality structure, which, up to some extent, would cause the negative transfer effect. The locality plays an important role in minimizing the nonlinear local domain discrepancy underlying the marginal distributions. For better exploiting the domain locality, a novel local generative discrepancy metric-based intermediate domain generation learning called Manifold Criterion guided Transfer Learning (MCTL) is proposed in this paper. The merits of the proposed MCTL are fourfold: 1) the concept of manifold criterion (MC) is first proposed as a measure validating the distribution matching across domains, and DA is achieved if the MC is satisfied; 2) the proposed MC can well guide the generation of the intermediate domain sharing similar distribution with the target domain, by minimizing the local domain discrepancy; 3) a global generative discrepancy metric is presented, such that both the global and local discrepancies can be effectively and positively reduced; and 4) a simplified version of MCTL called MCTL-S is presented under a perfect domain generation assumption for more generic learning scenario. Experiments on a number of benchmark visual transfer tasks demonstrate the superiority of the proposed MC guided generative transfer method, by comparing with the other state-of-the-art methods. The source code is available in https://github.com/wangshanshanCQU/MCTL. Lei Zhang 0038, Shanshan Wang 0008, Guang-Bin Huang, Wangmeng Zuo, Jian Yang 0003, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | LSTN: Latent Subspace Transfer Network for Unsupervised Domain Adaptation
Shanshan Wang 0008, Lei Zhang 0038 |
PRCV (2) | 1 |