VLDB 2026 Research / reviewers in the wild / expert
Shuang Li 0008
dblp:43/6294-8
· DBLP profile ↗
67ranked-venue papers
21as first author
55since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 13 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 10 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPSC: Sparse and Scalable Multi-Modal 3D Occupancy Prediction for Autonomous Drivingabstract3D semantic occupancy prediction offers a nuanced representation of the surrounding environment, which is crucial for ensuring the safety of autonomous driving. However, fine-grained scene representations inevitably result in cubic growth in data scale, which imposes substantial demands on model architecture and computational complexity, especially in high-resolution scenarios. Existing approaches for handling high-resolution scenes typically obtain fine-grained features by grid sampling on low-resolution feature map, resulting in limited sparsity and insufficient feature interaction. This paper presents a framework leveraging SParse representation and SCalable feature interaction to address the aforementioned challenges, called SPSC. Specifically, we maintain sparsity by progressively pruning unoccupied queries during the coarse-to-fine process, thereby reducing the scale of data that the model needs to handle. Subsequently, we introduce query serialization, which transforms queries into an ordered sequence while preserving their spatial structure, This enables fine-grained feature interaction while maintaining linear computational complexity and a larger receptive field. Without complex architectural designs, SPSC significantly outperforms SOTA approaches, relatively enhances the mIoU by 12.0%, 11.0% and 4.8% on nuScenes-Occupancy dataset under the muli-modal, LiDAR and camera settings, respectively. Qingju Guo, Shuang Li 0008, Binhui Xie, Jing Geng 0002, Wei Li 0111 |
AAAI | 2 |
| 2026 | Cluster-based Pseudo-labeling for Semi-Supervised LiDAR Semantic SegmentationabstractThe costly annotation process has driven the development of semi-supervised learning (SSL) approaches. Existing semi-supervised LiDAR segmentation methods typically process entire point clouds directly, aiming to assign labels to all points at the scene scale. However, the large number of points, combined with their sparse and irregular nature, makes it challenging to learn scene-level optimization objectives, especially in SSL settings where labeled data are insufficient. This paper presents a Cluster-based pseudo-LAbeling Semi-Supervised technique, called CLASS. CLASS is designed to divide point clouds into several small, pure clusters, thereby decomposing challenging scene-scale segmentation task into more manageable cluster-scale classification and segmentation tasks, enabling the generation of high-quality pseudo labels for unlabeled data. CLASS possesses three key properties. i) Task simplicity: our pseudo-labeling process is based on simpler cluster-scale classification and segmentation tasks, resulting in ease of learning. ii) Labeling effectiveness: CLASS can generate pseudo-labels comparable to ground truth using only approximately 10% labeled data. iii) Universal versatility: CLASS exhibits flexibility regarding LiDAR representations (e.g., BEV, voxel, and range view). Comprehensive experiments on popular LiDAR segmentation benchmarks demonstrate its superiority. Qingju Guo, Shuang Li 0008, Jing Geng 0002, Binhui Xie, Jiawei Shan, Wei Li 0111 |
WACV | 2 |
| 2026 | FairNS: Fair Negative Sampling in Collaborative Filtering via Diffusion ModelsabstractCollaborative Filtering (CF) methods commonly use negative sampling to improve preference learning by contrasting observed interactions with unobserved items. While effective, conventional practice implicitly treats all unclicked items as equally informative, regardless of their semantic group (e.g., genre or category). This overlooks a critical limitation: models may exploit coarse group distinctions rather than genuine fine-grained preferences, thereby inducing exposure imbalances across item groups. Such imbalances constitute a violation of item-side fairness, which seeks equitable exposure and evaluation for items from different semantic groups. When negative samples are drawn predominantly from groups semantically distant from a user’s positives, the learning signal becomes biased and comparisons unfair. We therefore revisit negative sampling through the lens of item-side fairness and argue that genuine fairness requires context-aware sampling that ensures like-for-like comparisons within each semantic group. To this end, we introduce FairNS, a diffusion-based sampling framework that generates negative samples within the same semantic group as the user’s positives, encouraging fair intra-group contrasts that respect group integrity. By centering training on these intra-group comparisons, FairNS mitigates cross-group bias and enables the recommender to learn more precise user preferences. FairNS is optimized via a bi-level objective that jointly refines the sampling mechanism and the recommendation model. Experiments on three benchmark datasets show that FairNS achieves a favorable fairness–accuracy tradeoff. Shuang Li 0008, Zhao Zhang 0011, Yakun Wang 0001, Deqing Wang 0001, Fuzhen Zhuang |
ACM Trans. Inf. Syst. | 3 |
| 2026 | Beyond Texts: Incorporating Co-occurrences into the Review-based Conversation Recommendation SystemsabstractConversational Recommender Systems (CRSs) interact with users through natural language to provide recommendations and generate responses. Due to limited information in conversation, existing works utilize KGs or reviews to improve CRS. Despite achievements, they overlook co-occurrence relations which have shown effectiveness in collaborative filtering systems. In this work, we first propose a novel framework named CoCRS , aiming to incorporate Co-occurrences into the Review-based Conversation Recommendation Systems . In CoCRS, we mine co-occurrences from two aspects: (1) item and entity , (2) user and item . For the first one, we extract entities from redundant review texts by KG and construct a relation-aware item-entity heterogeneous graph. In the second aspect, we analyze review sentiments and construct a sentiment-aware user-item bipartite graph. We encode two graphs to obtain user and entity embeddings. Since users in CRS are anonymous, we generate a virtual similar user representation to match reviews with users. Besides, we capture time-aware preference representation from two-time dimensions. Finally, we generate word-level user representation with word-oriented KG and model user preference by integrating the above representations. Extensive experiments demonstrate that CoCRS outperforms baselines and the cold-start experiment highlights its robustness. The Large Language Model (LLM) experiment illustrates the significant role of co-occurrence relationships in LLM-based CRS. Our code are available at https://github.com/Qin-lab-code/CoCRS . Haoyao Zhang, Zhida Qin, Xufeng Liang, Shuang Li 0008, John C. S. Lui |
ACM Trans. Inf. Syst. | 5 |
| 2025 | From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based SelectionabstractPretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping, in alignment with fine-grained class descriptions generated by large language models (LLMs), significantly enhancing zero-shot performance by incorporating multi-view information. However, the inherent randomness of these augmentations can inevitably introduce background artifacts and cause models to overly focus on local details, compromising global semantic understanding. To address these issues, we propose an Attention-Based Selection (ABS) method from local details to global context, which applies attention-guided cropping in both raw images and feature space, supplement global semantic information through strategic feature selection. Additionally, we introduce a soft matching technique to effectively filter LLM descriptions for better alignment. ABS achieves state-of-the-art performance on out-of-distribution generalization and zero-shot classification tasks. Notably, ABS is training-free and even rivals few-shot and test-time adaptation methods. Lincan Cai, Jingxuan Kang, Shuang Li 0008, Wenxuan Ma 0001, Binhui Xie, Zhida Qin, Jian Liang 0002 |
ICML | 3 |
| 2025 | Learning Time-Aware Causal Representation for Model Generalization in Evolving DomainsabstractEndowing deep models with the ability to generalize in dynamic scenarios is of vital significance for real-world deployment, given the continuous and complex changes in data distribution. Recently, evolving domain generalization (EDG) has emerged to address distribution shifts over time, aiming to capture evolving patterns for improved model generalization. However, existing EDG methods may suffer from spurious correlations by modeling only the dependence between data and targets across domains, creating a shortcut between task-irrelevant factors and the target, which hinders generalization. To this end, we design a time-aware structural causal model (SCM) that incorporates dynamic causal factors and the causal mechanism drifts, and propose Static-DYNamic Causal Representation Learning (SYNC), an approach that effectively learns time-aware causal representations. Specifically, it integrates specially designed information-theoretic objectives into a sequential VAE framework which captures evolving patterns, and produces the desired representations by preserving intra-class compactness of causal factors both across and within domains. Moreover, we theoretically show that our method can yield the optimal causal predictor for each time domain. Results on both synthetic and real-world datasets exhibit that SYNC can achieve superior temporal generalization performance. Zhuo He, Shuang Li 0008, Wenze Song, Longhui Yuan, Jian Liang 0002, Han Li 0005, Kun Gai |
ICML | 2 |
| 2025 | Dual-GT: Dual-scale Spatial Dependency for Grid-based Traffic Flow Prediction
Xufeng Liang, Zhida Qin, Pengzhan Zhou, Shuang Li 0008 |
INFOCOM | 4 |
| 2025 | Time Matters: Enhancing Sequential Recommendations with Time-Guided Graph Neural ODEsabstractSequential recommendation (SR) is widely deployed in e-commerce platforms, streaming services, etc., revealing significant potential to enhance user experience. The core of SR lies in exploring the sequential relationships in historical user-item interactions. However, existing methods often overlook two critical factors: irregular user interests between interactions and highly uneven item distributions over time. The former factor implies that actual user preferences are not always continuous, and long-term historical interactions may not be relevant to current purchasing behavior. Therefore, relying only on these historical interactions for recommendations may result in a lack of user interest at the target time. The latter factor, characterized by peaks and valleys in interaction frequency, may result from seasonal trends, special events, or promotions. These externally driven distributions may not align with individual user interests, leading to inaccurate recommendations. To address these deficiencies, we propose TGODE to both enhance and capture the long-term historical interactions. Specifically, we first construct the user time graph and item evolution graph, which utilize user personalized preferences and global item distribution information, respectively. To tackle the temporal sparsity caused by irregular user interactions, we design a time-guided diffusion generator to automatically obtain an augmented time-aware user graph. Additionally, we devise a user interest truncation factor to efficiently identify sparse time intervals and achieve balanced preference inference. After that, the augmented user graph and item graph are fed into a generalized graph neural ordinary differential equation (ODE) to align with the evolution of user preferences and item distributions. This allows two patterns of information evolution to be matched over time. Experimental results demonstrate that TGODE outperforms baseline methods across five datasets, with improvements ranging from 10% to 46%. The code is available at https://github.com/Qin-lab-code/TGODE. Haoyan Fu, Zhida Qin, Shixiao Yang, Haoyao Zhang, Bin Lu 0005, Shuang Li 0008, John C. S. Lui |
KDD (2) | 6 |
| 2025 | AdvMixUp: Adversarial MixUp Regularization for Deep LearningabstractDeep neural networks (DNNs) have shown significant progress in many application fields. However, overfitting remains a significant challenge in their development. While existing data-augmentation techniques such as MixUp have been successful in preventing overfitting, they often fail to generate hard mixed samples near the decision boundary, impeding model optimization. In this article, we present adversarial MixUp (AdvMixUp), a novel sample-dependent method for regularizing DNNs. AdvMixUp addresses this issue by incorporating adversarial training (AT) to create sample-dependent and feature-level interpolation masks, generating more challenging mixed samples. These virtual samples enable DNNs to learn more robust features, ultimately reducing overfitting. Empirical evaluations on CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet demonstrate that AdvMixUp outperforms existing MixUp variants. Jun Fu 0001, Xianrui Ji, Dexiong Chen, Guosheng Hu, Shuang Li 0008, Xiating Feng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Domain Adaptation via Prompt LearningabstractUnsupervised domain adaptation (UDA) aims to adapt models learned from a well-annotated source domain to a target domain, where only unlabeled samples are given. Current UDA approaches learn domain-invariant features by aligning source and target feature spaces through statistical discrepancy minimization or adversarial training. However, these constraints could lead to the distortion of semantic feature structures and loss of class discriminability. In this article, we introduce a novel prompt learning paradigm for UDA, named domain adaptation via prompt learning (DAPrompt). In contrast to prior works, our approach learns the underlying label distribution for target domain rather than aligning domains. The main idea is to embed domain information into prompts, a form of representation generated from natural language, which is then used to perform classification. This domain information is shared only by images from the same domain, thereby dynamically adapting the classifier according to each domain. By adopting this paradigm, we show that our model not only outperforms previous methods on several cross-domain benchmarks but also is very efficient to train and easy to implement. Chunjiang Ge, Rui Huang 0012, Mixue Xie, Zihang Lai, Shiji Song, Shuang Li 0008, Gao Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Source-Free Active Domain Adaptation via Augmentation-Based Sample Query and Progressive Model AdaptationabstractActive domain adaptation (ADA), which enormously improves the performance of unsupervised domain adaptation (UDA) at the expense of annotating limited target data, has attracted a surge of interest. However, in real-world applications, the source data in conventional ADA are not always accessible due to data privacy and security issues. To alleviate this dilemma, we introduce a more practical and challenging setting, dubbed as source-free ADA (SFADA), where one can select a small quota of target samples for label query to assist the model learning, but labeled source data are unavailable. Therefore, how to query the most informative target samples and mitigate the domain gap without the aid of source data are two key challenges in SFADA. To address SFADA, we propose a unified method SQAdapt via augmentation-based ample uery and progressive model Adapt ation. In specific, an active selection module (ASM) is built for target label query, which exploits data augmentation to select the most informative target samples with high predictive sensitivity and uncertainty. Then, we further introduce a classifier adaptation module (CAM) to leverage both the labeled and unlabeled target data for progressively calibrating the classifier weights. Meanwhile, the source-like target samples with low selection scores are taken as source surrogates to realize the distribution alignment in the source-free scenario by the proposed distribution alignment module (DAM). Moreover, as a general active label query method, SQAdapt can be easily integrated into other source-free UDA (SFUDA) methods, and improve their performance. Comprehensive experiments on multiple benchmarks have shown that SQAdapt can achieve superior performance and even surpass most of the ADA methods. Shuang Li 0008, Rui Zhang 0113, Kaixiong Gong, Mixue Xie, Wenxuan Ma 0001, Guangyu Gao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Semantic Correlation Transfer for Heterogeneous Domain AdaptationabstractHeterogeneous domain adaptation (HDA) is expected to achieve effective knowledge transfer from a label-rich source domain to a heterogeneous target domain with scarce labeled data. Most prior HDA methods strive to align the cross-domain feature distributions by learning domain invariant representations without considering the intrinsic semantic correlations among categories, which inevitably results in the suboptimal adaptation performance across domains. Therefore, to address this issue, we propose a novel semantic correlation transfer (SCT) method for HDA, which not only matches the marginal and conditional distributions between domains to mitigate the large domain discrepancy, but also transfers the category correlation knowledge underlying the source domain to target by maximizing the pairwise class similarity across source and target. Technically, the domainwise and classwise centroids (prototypes) are first computed and aligned according to the feature embeddings. Then, based on the derived classwise prototypes, we leverage the cosine similarity of each two classes in both domains to transfer the supervised source semantic correlation knowledge among different categories to target effectively. As a result, the feature transferability and category discriminability can be simultaneously improved during the adaptation process. Comprehensive experiments and ablation studies on standard HDA tasks, such as text-to-image, image-to-image, and text-to-text, have demonstrated the superiority of our proposed SCT against several state-of-the-art HDA methods. Shuang Li 0008, Rui Zhang 0113, Chi Harold Liu, Weipeng Cao, Xizhao Wang, Song Tian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Learning Modality Knowledge Alignment for Cross-Modality TransferabstractCross-modality transfer aims to leverage large pretrained models to complete tasks that may not belong to the modality of pretraining data. Existing works achieve certain success in extending classical finetuning to cross-modal scenarios, yet we still lack understanding about the influence of modality gap on the transfer. In this work, a series of experiments focusing on the source representation quality during transfer are conducted, revealing the connection between larger modality gap and lesser knowledge reuse which means ineffective transfer. We then formalize the gap as the knowledge misalignment between modalities using conditional distribution $P(Y|X)$. Towards this problem, we present Modality kNowledge Alignment (MoNA), a meta-learning approach that learns target data transformation to reduce the modality knowledge discrepancy ahead of the transfer. Experiments show that the approach significantly improves upon cross-modal finetuning methods, and most importantly leads to better reuse of source modality knowledge. Wenxuan Ma 0001, Shuang Li 0008, Lincan Cai, Jingxuan Kang |
ICML | 2 |
| 2024 | Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality GenerationabstractLarge-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scarcity of labeled data. In this paper, we propose an end-to-end method, PaRe, to enhance cross-modal fine-tuning, aiming to transfer a large-scale pretrained model to various target modalities. PaRe employs a gating mechanism to select key patches from both source and target data. Through a modality-agnostic Patch Replacement scheme, these patches are preserved and combined to construct data-rich intermediate modalities ranging from easy to hard. By gradually intermediate modality generation, we can not only effectively bridge the modality gap to enhance stability and transferability of cross-modal fine-tuning, but also address the challenge of limited data in the target modality by leveraging enriched intermediate modality data. Compared with hand-designed, general-purpose, task-specific, and state-of-the-art cross-modal fine-tuning approaches, PaRe demonstrates superior performance across three challenging benchmarks, encompassing more than ten modalities. Lincan Cai, Shuang Li 0008, Wenxuan Ma 0001, Jingxuan Kang, Binhui Xie, Zixun Sun, Chengwei Zhu |
ICML | 2 |
| 2024 | Weighted Multiple Source-Free Domain Adaptation Ensemble Network in Intelligent Machinery Fault Diagnosis
Renhu Bu, Shuang Li 0008, Chi Harold Liu |
KSEM (2) | 2 |
| 2024 | Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time AdaptationabstractCapitalizing on the complementary advantages of generative and discriminative models has always been a compelling vision in machine learning, backed by a growing body of research. This work discloses the hidden semantic structure within score-based generative models, unveiling their potential as effective discriminative priors. Inspired by our theoretical findings, we propose DUSA to exploit the structured semantic priors underlying diffusion score to facilitate the test-time adaptation of image classifiers or dense predictors. Notably, DUSA extracts knowledge from a single timestep of denoising diffusion, lifting the curse of Monte Carlo-based likelihood estimation over timesteps. We demonstrate the efficacy of our DUSA in adapting a wide variety of competitive pre-trained discriminative models on diverse test-time scenarios. Additionally, a thorough ablation study is conducted to dissect the pivotal elements in DUSA. Code is publicly available at https://github.com/BIT-DA/DUSA. Mingjia Li 0003, Shuang Li 0008, Tongrui Su, Longhui Yuan, Jian Liang 0002, Wei Li 0111 |
NeurIPS | 2 |
| 2024 | Weight Diffusion for Future: Learn to Generalize in Non-Stationary EnvironmentsabstractEnabling deep models to generalize in non-stationary environments is vital for real-world machine learning, as data distributions are often found to continually change. Recently, evolving domain generalization (EDG) has emerged to tackle the domain generalization in a time-varying system, where the domain gradually evolves over time in an underlying continuous structure. Nevertheless, it typically assumes multiple source domains simultaneously ready. It still remains an open problem to address EDG in the domain-incremental setting, where source domains are non-static and arrive sequentially to mimic the evolution of training domains. To this end, we propose Weight Diffusion (W-Diff), a novel framework that utilizes the conditional diffusion model in the parameter space to learn the evolving pattern of classifiers during the domain-incremental training process. Specifically, the diffusion model is conditioned on the classifier weights of different historical domain (regarded as a reference point) and the prototypes of current domain, to learn the evolution from the reference point to the classifier weights of current domain (regarded as the anchor point). In addition, a domain-shared feature encoder is learned by enforcing prediction consistency among multiple classifiers, so as to mitigate the overfitting problem and restrict the evolving pattern to be reflected in the classifier as much as possible. During inference, we adopt the ensemble manner based on a great number of target domain-customized classifiers, which are cheaply obtained via the conditional diffusion model, for robust prediction. Comprehensive experiments on both synthetic and real-world datasets show the superior generalization performance of W-Diff on unseen domains in the future. Mixue Xie, Shuang Li 0008, Binhui Xie, Chi Harold Liu, Jian Liang 0002, Zixun Sun, Chengwei Zhu |
NeurIPS | 2 |
| 2024 | Adapting Across Domains via Target-Oriented Transferable Semantic Augmentation Under Prototype Constraint
Mixue Xie, Shuang Li 0008, Kaixiong Gong, Yulin Wang 0002, Gao Huang 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | VBLC: Visibility Boosting and Logit-Constraint Learning for Domain Adaptive Semantic Segmentation under Adverse ConditionsabstractGeneralizing models trained on normal visual conditions to target domains under adverse conditions is demanding in the practical systems. One prevalent solution is to bridge the domain gap between clear- and adverse-condition images to make satisfactory prediction on the target. However, previous methods often reckon on additional reference images of the same scenes taken from normal conditions, which are quite tough to collect in reality. Furthermore, most of them mainly focus on individual adverse condition such as nighttime or foggy, weakening the model versatility when encountering other adverse weathers. To overcome the above limitations, we propose a novel framework, Visibility Boosting and Logit-Constraint learning (VBLC), tailored for superior normal-toadverse adaptation. VBLC explores the potential of getting rid of reference images and resolving the mixture of adverse conditions simultaneously. In detail, we first propose the visibility boost module to dynamically improve target images via certain priors in the image level. Then, we figure out the overconfident drawback in the conventional cross-entropy loss for self-training method and devise the logit-constraint learning, which enforces a constraint on logit outputs during training to mitigate this pain point. To the best of our knowledge, this is a new perspective for tackling such a challenging task. Extensive experiments on two normal-to-adverse domain adaptation benchmarks, i.e., Cityscapes to ACDC and Cityscapes to FoggyCityscapes + RainCityscapes, verify the effectiveness of VBLC, where it establishes the new state of the art. Code is available at https://github.com/BIT-DA/VBLC. Mingjia Li 0003, Binhui Xie, Shuang Li 0008, Chi Harold Liu, Xinjing Cheng |
AAAI | 3 |
| 2023 | Improving Generalization with Domain Convex GameabstractDomain generalization (DG) tends to alleviate the poor generalization capability of deep neural networks by learning model with multiple source domains. A classical solution to DG is domain augmentation, the common belief of which is that diversifying source domains will be conducive to the out-of-distribution generalization. However, these claims are understood intuitively, rather than mathematically. Our explorations empirically reveal that the correlation between model generalization and the diversity of domains may be not strictly positive, which limits the effectiveness of domain augmentation. This work therefore aim to guarantee and further enhance the validity of this strand. To this end, we propose a new perspective on DG that recasts it as a convex game between domains. We first encourage each diversified domain to enhance model generalization by elaborately designing a regularization term based on supermodularity. Meanwhile, a sample filter is constructed to eliminate low-quality samples, thereby avoiding the impact of potentially harmful information. Our framework presents a new avenue for the formal analysis of DG, heuristic analysis and extensive experiments demonstrate the rationality and effectiveness.11Code is available at “https://github.com/BIT-DA/DCG”. Fangrui Lv, Jian Liang 0002, Shuang Li 0008, Di Liu 0029 |
CVPR | 3 |
| 2023 | On the Difficulty of Unpaired Infrared-to-Visible Video Translation: Fine-Grained Content-Rich Patches TransferabstractExplicit visible videos can provide sufficient visual information and facilitate vision applications. Unfortunately, the image sensors of visible cameras are sensitive to light conditions like darkness or overexposure. To make up for this, recently, infrared sensors capable of stable imaging have received increasing attention in autonomous driving and monitoring. However, most prosperous vision models are still trained on massive clear visible data, facing huge visual gaps when deploying to infrared imaging scenarios. In such cases, transferring the infrared video to a distinct visible one with fine-grained semantic patterns is a worthwhile endeavor. Previous works improve the outputs by equally optimizing each patch on the translated visible results, which is unfair for enhancing the details on content-rich patches due to the long-tail effect of pixel distribution. Here we propose a novel CPTrans framework to tackle the challenge via balancing gradients of different patches, achieving the fine-grained Content-rich Patches Transferring. Specifically, the content-aware optimization module encourages model optimization along gradients of target patches, ensuring the improvement of visual details. Additionally, the content-aware temporal normalization module enforces the generator to be robust to the motions of target patches. Moreover, we extend the existing dataset InfraredCity to more challenging adverse weather conditions (rain and snow), dubbed as InfraredCity-Adverse11The code and dataset are available at https://github.com/BIT-DA/12V-Processing, Extensive experiments show that the proposed CPTrans achieves state-of-the-art performance under diverse scenes while requiring less training time than competitive methods. Zhenjie Yu, Shuang Li 0008, Yirui Shen, Chi Harold Liu, Shuigen Wang |
CVPR | 2 |
| 2023 | Robust Test-Time Adaptation in Dynamic ScenariosabstractTest-time adaptation (TTA) intends to adapt the pretrained model to test distributions with only unlabeled test data streams. Most of the previous TTA methods have achieved great success on simple test data streams such as independently sampled data from single or multiple distributions. However, these attempts may fail in dynamic scenarios of real-world applications like autonomous driving, where the environments gradually change and the test data is sampled correlatively over time. In this work, we explore such practical test data streams to deploy the model on the fly, namely practical test-time adaptation (PTTA). To do so, we elaborate a Robust Test-Time Adaptation (RoTTA) method against the complex data stream in PTTA. More specifically, we present a robust batch normalization scheme to estimate the normalization statistics. Meanwhile, a memory bank is utilized to sample category-balanced data with consideration of timeliness and uncertainty. Further, to stabilize the training procedure, we develop a time-aware reweighting strategy with a teacher-student model. Extensive experiments prove that RoTTA enables continual test-time adaptation on the correlatively sampled data streams. Our method is easy to implement, making it a good choice for rapid deployment. The code is publicly available at https://github.com/BIT-DA/RoTTA Longhui Yuan, Binhui Xie, Shuang Li 0008 |
CVPR | 3 |
| 2023 | Borrowing Knowledge From Pre-trained Language Model: A New Data-efficient Visual Learning ParadigmabstractThe development of vision models for real-world applications is hindered by the challenge of annotated data scarcity, which has necessitated the adoption of dataefficient visual learning techniques such as semi-supervised learning. Unfortunately, the prevalent cross-entropy supervision is limited by its focus on category discrimination while disregarding the semantic connection between concepts, which ultimately results in the suboptimal exploitation of scarce labeled data. To address this issue, this paper presents a novel approach that seeks to leverage linguistic knowledge for data-efficient visual learning. The proposed approach, BorLan, Borrows knowledge from off-theshelf pretrained Language models that are already endowed with rich semantics extracted from large corpora, to compensate the semantic deficiency due to limited annotation in visual training. Specifically, we design a distribution alignment objective, which guides the vision model to learn both semantic-aware and domain-agnostic representations for the task through linguistic knowledge. One significant advantage of this paradigm is its flexibility in combining various visual and linguistic models. Extensive experiments on semi-supervised learning, single domain generalization and few-shot learning validate its effectiveness. Code is available at https://github.com/BIT-DA/BorLan. Wenxuan Ma 0001, Shuang Li 0008, Chi Harold Liu, Jingxuan Kang, Yulin Wang 0002, Gao Huang 0001 |
ICCV | 2 |
| 2023 | Dirichlet-based Uncertainty Calibration for Active Domain Adaptation
Mixue Xie, Shuang Li 0008, Rui Zhang 0113, Chi Harold Liu |
ICLR | 2 |
| 2023 | Style Transfer Meets Super-Resolution: Advancing Unpaired Infrared-to-Visible Image Translation with Detail EnhancementabstractThe problem of unpaired infrared-to-visible image translation has gained significant attention due to its ability to generate visible images with color information from low-detail grayscale infrared inputs. However, current methodologies often depend on conventional style transfer techniques, which constrain the spatial resolution of the visible output to be equivalent to that of the input infrared image. The fixed generation pattern results in blurry generated results when translating low-resolution infrared inputs, and utilizing high-resolution infrared inputs as a solution necessitates greater computational resources. This spurs us to investigate the challenging unpaired image translation from low-resolution infrared inputs to high-resolution visible outputs, with the ultimate goal of enhancing image details while reducing computational costs. Therefore, we propose a unified framework that integrates the super-resolution process into our unpaired infrared-to-visible image transfer, yielding realistic and high-resolution results. Specifically, we propose the Detail Consistency Loss to establish a connection between the two aforementioned modules, thereby enhancing the quality of visual detail in style transfer results through the super-resolution module. Furthermore, our Texture Perceptual Loss is designed to ensure that the generator generates high-quality visual details accurately and reliably. Experimental results indicate that our method outperforms other comparative approaches when utilizing low-resolution infrared inputs. Remarkably, our approach even surpasses techniques that use high-resolution infrared inputs to generate visible images. Last but equally important, we propose a new and challenging dataset, dubbed as InfraredCity-HD, which comprises 512X512 resolution images, to advance research on high-resolution infrared-related fields. Yirui Shen, Jingxuan Kang, Shuang Li 0008, Zhenjie Yu, Shuigen Wang |
ACM Multimedia | 3 |
| 2023 | Language Semantic Graph Guided Data-Efficient LearningabstractDeveloping generalizable models that can effectively learn from limited data and with minimal reliance on human supervision is a significant objective within the machine learning community, particularly in the era of deep neural networks. Therefore, to achieve data-efficient learning, researchers typically explore approaches that can leverage more related or unlabeled data without necessitating additional manual labeling efforts, such as Semi-Supervised Learning (SSL), Transfer Learning (TL), and Data Augmentation (DA).
SSL leverages unlabeled data in the training process, while TL enables the transfer of expertise from related data distributions. DA broadens the dataset by synthesizing new data from existing examples. However, the significance of additional knowledge contained within labels has been largely overlooked in research. In this paper, we propose a novel perspective on data efficiency that involves exploiting the semantic information contained in the labels of the available data. Specifically, we introduce a Language Semantic Graph (LSG) which is constructed from labels manifest as natural language descriptions. Upon this graph, an auxiliary graph neural network is trained to extract high-level semantic relations and then used to guide the training of the primary model, enabling more adequate utilization of label knowledge. Across image, video, and audio modalities, we utilize the LSG method in both TL and SSL scenarios and illustrate its versatility in significantly enhancing performance compared to other data-efficient learning approaches. Additionally, our in-depth analysis shows that the LSG method also expedites the training process. Wenxuan Ma 0001, Shuang Li 0008, Lincan Cai, Jingxuan Kang |
NeurIPS | 2 |
| 2023 | Annotator: A Generic Active Learning Baseline for LiDAR Semantic SegmentationabstractActive learning, a label-efficient paradigm, empowers models to interactively query an oracle for labeling new data. In the realm of LiDAR semantic segmentation, the challenges stem from the sheer volume of point clouds, rendering annotation labor-intensive and cost-prohibitive. This paper presents Annotator, a general and efficient active learning baseline, in which a voxel-centric online selection strategy is tailored to efficiently probe and annotate the salient and exemplar voxel girds within each LiDAR scan, even under distribution shift. Concretely, we first execute an in-depth analysis of several common selection strategies such as Random, Entropy, Margin, and then develop voxel confusion degree (VCD) to exploit the local topology relations and structures of point clouds. Annotator excels in diverse settings, with a particular focus on active learning (AL), active source-free domain adaptation (ASFDA), and active domain adaptation (ADA). It consistently delivers exceptional performance across LiDAR semantic segmentation benchmarks, spanning both simulation-to-real and real-to-real scenarios. Surprisingly, Annotator exhibits remarkable efficiency, requiring significantly fewer annotations, e.g., just labeling five voxels per scan in the SynLiDAR → SemanticKITTI task. This results in impressive performance, achieving 87.8% fully-supervised performance under AL, 88.5% under ASFDA, and 94.4% under ADA. We envision that Annotator will offer a simple, general, and efficient solution for label-efficient 3D applications. Binhui Xie, Shuang Li 0008, Qingju Guo, Chi Harold Liu, Xinjing Cheng |
NeurIPS | 2 |
| 2023 | Evolving Standardization for Continual Domain Generalization over Temporal DriftabstractThe capability of generalizing to out-of-distribution data is crucial for the deployment of machine learning models in the real world. Existing domain generalization (DG) mainly embarks on offline and discrete scenarios, where multiple source domains are simultaneously accessible and the distribution shift among domains is abrupt and violent. Nevertheless, such setting may not be universally applicable to all real-world applications, as there are cases where the data distribution gradually changes over time due to various factors, e.g., the process of aging. Additionally, as the domain constantly evolves, new domains will continually emerge. Re-training and updating models with both new and previous domains using existing DG methods can be resource-intensive and inefficient. Therefore, in this paper, we present a problem formulation for Continual Domain Generalization over Temporal Drift (CDGTD). CDGTD addresses the challenge of gradually shifting data distributions over time, where domains arrive sequentially and models can only access the data of the current domain. The goal is to generalize to unseen domains that are not too far into the future. To this end, we propose an Evolving Standardization (EvoS) method, which characterizes the evolving pattern of feature distribution and mitigates the distribution shift by standardizing features with generated statistics of corresponding domain. Specifically, inspired by the powerful ability of transformers to model sequence relations, we design a multi-scale attention module (MSAM) to learn the evolving pattern under sliding time windows of different lengths. MSAM can generate statistics of current domain based on the statistics of previous domains and the learned evolving pattern. Experiments on multiple real-world datasets including images and texts validate the efficacy of our EvoS. Mixue Xie, Shuang Li 0008, Longhui Yuan, Chi Harold Liu, Zehui Dai |
NeurIPS | 2 |
| 2023 | Joint Semantic Transfer Network for IoT Intrusion DetectionabstractIn this article, we propose a joint semantic transfer network (JSTN) toward effective intrusion detection (ID) for large-scale scarcely labeled Internet of Things (IoT) domain. As a multisource heterogeneous domain adaptation (MS-HDA) method, the JSTN integrates a knowledge-rich network intrusion (NI) domain and another small-scale IoT intrusion (II) domain as source domains and preserves intrinsic semantic properties to assist target II domain ID. The JSTN jointly transfers the following three semantics to learn a domain-invariant and discriminative feature representation. The scenario semantic endows source NI and II domains with characteristics from each other to ease the knowledge transfer process via a confused domain discriminator and categorical distribution knowledge preservation. It also reduces the source–target discrepancy to make the shared feature space domain invariant. Meanwhile, the weighted implicit semantic transfer boosts discriminability via a fine-grained knowledge preservation, which transfers the source categorical distribution to the target domain. The source–target divergence guides the importance weighting during knowledge preservation to reflect the degree of knowledge learning. Additionally, the hierarchical explicit semantic alignment performs centroid-level and representative-level alignment with the help of a geometric similarity-aware pseudo-label refiner, which exploits the value of the unlabeled target II domain and explicitly aligns feature representations from a global and local perspective in a concentrated manner. Comprehensive experiments on various tasks verify the superiority of the JSTN against state-of-the-art comparing methods, on average a 10.3% of accuracy boost is achieved. The statistical soundness of each constituting component and the computational efficiency is also verified. Jiashu Wu, Yang Wang 0006, Binhui Xie, Shuang Li 0008, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Internet Things J. | 4 |
| 2023 | SePiCo: Semantic-Guided Pixel Contrast for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation attempts to make satisfactory dense predictions on an unlabeled target domain by utilizing the supervised model trained on a labeled source domain. One popular solution is self-training, which retrains the model with pseudo labels on target instances. Plenty of approaches tend to alleviate noisy pseudo labels, however, they ignore the intrinsic connection of the training data, i.e., intra-class compactness and inter-class dispersion between pixel representations across and within domains. In consequence, they struggle to handle cross-domain semantic variations and fail to build a well-structured embedding space, leading to less discrimination and poor generalization. In this work, we propose emantic-Guided Pixel Contrast (SePiCo), a novel one-stage adaptation framework that highlights the semantic concepts of individual pixels to promote learning of class-discriminative and class-balanced pixel representations across domains, eventually boosting the performance of self-training methods. Specifically, to explore proper semantic concepts, we first investigate a centroid-aware pixel contrast that employs the category centroids of the entire source domain or a single source image to guide the learning of discriminative features. Considering the possible lack of category diversity in semantic concepts, we then blaze a trail of distributional perspective to involve a sufficient quantity of instances, namely distribution-aware pixel contrast, in which we approximate the true distribution of each semantic category from the statistics of labeled source data. Moreover, such an optimization objective can derive a closed-form upper bound by implicitly involving an infinite number of (dis)similar pairs, making it computationally efficient. Extensive experiments show that SePiCo not only helps stabilize training but also yields discriminative representations, making significant progress on both synthetic-to-real and daytime-to-nighttime adaptation scenarios. The code and models are available at https://github.com/BIT-DA/SePiCo. Binhui Xie, Shuang Li 0008, Mingjia Li 0003, Chi Harold Liu, Gao Huang 0001, Guoren Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Critical Classes and Samples Discovering for Partial Domain AdaptationabstractPartial domain adaptation (PDA) attempts to learn transferable models from a large-scale labeled source domain to a small unlabeled target domain with fewer classes, which has attracted a recent surge of interest in transfer learning. Most conventional PDA approaches endeavor to design delicate source weighting schemes by leveraging target predictions to align cross-domain distributions in the shared class space. Accordingly, two crucial issues are overlooked in these methods. First, target prediction is a double-edged sword, and inaccurate predictions will result in negative transfer inevitably. Second, not all target samples have equal transferability during the adaptation; thus, "ambiguous" target data predicted with high uncertainty should be paid more attentions. In this article, we propose a critical classes and samples discovering network (CSDN) to identify the most relevant source classes and critical target samples, such that more precise cross-domain alignment in the shared label space could be enforced by co-training two diverse classifiers. Specifically, during the training process, CSDN introduces an adaptive source class weighting scheme to select the most relevant classes dynamically. Meanwhile, based on the designed target ambiguous score, CSDN emphasizes more on ambiguous target samples with larger inconsistent predictions to enable fine-grained alignment. Taking a step further, the weighting schemes in CSDN can be easily coupled with other PDA and DA methods to further boost their performance, thereby demonstrating its flexibility. Extensive experiments verify that CSDN attains excellent results compared to state of the arts on four highly competitive benchmark datasets. Shuang Li 0008, Kaixiong Gong, Binhui Xie, Chi Harold Liu, Weipeng Cao, Song Tian |
IEEE Trans. Cybern. | 1 |
| 2023 | Domain Adaptive Remote Sensing Scene Recognition via Semantic Relationship Knowledge TransferabstractScene recognition has attracted rising attentions of many researchers in the remote sensing fields, owing to the rapidly advancing of remote sensing devices in recent years. However, images obtained from various sensors dominate diverse sensor-specific characteristics, which will dramatically weaken the model transferability trained on a source data domain to a different target domain on account of the domain shift issues. To mitigate the domain discrepancy, most existing methods attend to align the cross-domain distributions. While the valuable knowledge of semantic relationships between different scenes is generally overlooked, and the underlying correlation across scenes cannot be fully discovered. For the sake of tackling this challenge, we propose an adaptive remote sensing scene recognition network, which can successfully transfer both the discriminative knowledge and cross-scene relationship from source to target. Specifically, in this paper, we acquire sensor-invariant representations in an adversarial manner and realize fine-grained conditional distribution alignment contrastively. In such a way, the tremendous domain gap can be mitigated to a large extent, and the discriminative and well-matched representations will be derived favorably. In addition, we explicitly construct class-wise relationship distributions belonging to two domains respectively and minimize their divergence to conduct semantic relationship knowledge transfer (SRKT), for the purpose of sufficiently unearthing the intrinsic semantic relative structures that can prompt generality of the model in the target domain. Finally, we conduct multiple experiments on representative multi-domain remote sensing benchmarks, and the extensive experimental results demonstrate the superiority of our proposed approach. Shuang Li 0008, Chi Harold Liu, Yuqi Han, Hao Shi 0006, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | End-to-End Transferable Anomaly Detection via Multi-Spectral Cross-Domain Representation AlignmentabstractAnomaly detection (AD) aims to distinguish abnormal instances from what is defined as normal, which strongly correlates with the safe and robust applications of machine learning. A well-performed anomaly detector often relies on the training on massive labeled data, while it is of high cost to annotate data in practice. Fortunately, this dilemma can be solved by transferring the knowledge of a label-rich dataset (source domain) to assist the learning on the label-scarce dataset (target domain), which is known as domain adaptation in transfer learning. In this paper, we propose a Multi-spectral Cross-domain Representation Alignment (MsRA) method for the anomaly detection in the domain adaptation setting, where we can only access normal source data andlimitednormal target data. Specifically, MsRA first constructs multi-spectral feature representations by fusing different frequency components of the original features, which mitigates the information scarcity due to limited target training data by capturing richer input pattern information. Then we employ the adversarial training strategy to learn domain-invariant features and force the features of normal data to be more compact by the center clustering. Finally, the distance of each sample to the prototype of normal class can be used as its anomaly score, where the prototype is the center of both source and target data. In this way, we achieve anomaly detection in an end-to-end manner, without two-stage training for feature extraction and anomaly detection. Comprehensive experiments on cross-domain anomaly detection benchmarks validate the effectiveness of MsRA. Shuang Li 0008, Shugang Li 0002, Mixue Xie, Kaixiong Gong, Jianxin Zhao 0001, Chi Harold Liu, Guoren Wang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Meta-Reweighted Regularization for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation enables knowledge transfer from a labeled source domain to an unlabeled target domain by reducing the cross-domain distribution discrepancy, and the adversarial learning based paradigm has achieved remarkable success. On top of this, recent works seeks to further regularize the classification decision boundary via self-training to learn target adaptive classifier with pseudo-labeled target samples. However, since pseudo labels are inevitably noisy, most of prior methods focus on manually designing elaborate target selection algorithms or optimization objectives. Different from them, we propose a meta-learning based target-reweighting regularization algorithm called MetaReg. Specifically, MetaReg is motivated by the intuition that an ideal target classifier trained on correct target pseudo labels should make small classification errors on target-like source samples. Therefore, we explicitly define a meta reweighting problem that aims to find optimal weights for different samples by minimizing the classification loss on a class-balanced set consisting of source samples that are most similar to target ones. The optimization problem is solved efficiently with a simplified approximation technique. As a result, the automatically learned optimal weights are utilized to reweight pseudo-labeled target samples and regularize the model learning. Comprehensive experiments verify that MetaReg outperforms the non-regularized UDA counterparts with state-of-the-art performance. Shuang Li 0008, Wenxuan Ma 0001, Chi Harold Liu, Jian Liang 0002, Guoren Wang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | A Collaborative Alignment Framework of Transferable Knowledge Extraction for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to utilize knowledge from a label-rich source domain to understand a similar yet distinct unlabeled target domain. Notably, global distribution statistics across domains and local semantic characteristics across samples, are two essential factors of data analysis that should be fully explored. Most existing UDA approaches either harness only one of them or fail to closely associate them for efficient adaptation. In this work, we propose a unified framework, called Collaborative Alignment Framework (CAF), which simultaneously reduces the global domain discrepancy and preserves the local semantic consistency for cross-domain knowledge transfer in a collaborative manner. Specifically, for domain-oriented alignment, we utilize adversarial training or minimize the Wasserstein distance between the two distributions to learn domain-level invariant representations. For semantic-oriented matching, we capture the semantic discrepancy between the predictions of two diverse task-specific classifiers and enhance the features of target data to be near the support of the source data class-wisely, which promotes semantic consistency across domains effectively. These two adaptation processes can be deeply intertwined in CAF via collaborative training, thus CAF can learn domain-invariant and semantic-consistent feature representations. Extensive experiments on four popular benchmarks, including DomainNet, VisDA-2017, Office-31, and ImageCLEF, demonstrate the proposed methods significantly outperform the existing methods, especially on the large-scale dataset. The code is available athttps://github.com/BIT-DA/CAF. Binhui Xie, Shuang Li 0008, Fangrui Lv, Chi Harold Liu, Guoren Wang, Dapeng Oliver Wu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Active Learning for Domain Adaptation: An Energy-Based ApproachabstractUnsupervised domain adaptation has recently emerged as an effective paradigm for generalizing deep neural networks to new target domains. However, there is still enormous potential to be tapped to reach the fully supervised performance. In this paper, we present a novel active learning strategy to assist knowledge transfer in the target domain, dubbed active domain adaptation. We start from an observation that energy-based models exhibit free energy biases when training (source) and test (target) data come from different distributions. Inspired by this inherent mechanism, we empirically reveal that a simple yet efficient energy-based sampling strategy sheds light on selecting the most valuable target samples than existing approaches requiring particular architectures or computation of the distances. Our algorithm, Energy-based Active Domain Adaptation (EADA), queries groups of target data that incorporate both domain characteristic and instance uncertainty into every selection round. Meanwhile, by aligning the free energy of target data compact around the source domain via a regularization term, domain gap can be implicitly diminished. Through extensive experiments, we show that EADA surpasses state-of-the-art methods on well-known challenging benchmarks with substantial improvements, making it a useful option in the open world. Code is available at https://github.com/BIT-DA/EADA. Binhui Xie, Longhui Yuan, Shuang Li 0008, Chi Harold Liu, Xinjing Cheng, Guoren Wang |
AAAI | 3 |
| 2022 | Causality Inspired Representation Learning for Domain GeneralizationabstractDomain generalization (DG) is essentially an out-of-distribution problem, aiming to generalize the knowledge learned from multiple source domains to an unseen target domain. The mainstream is to leverage statistical models to model the dependence between data and labels, intending to learn representations independent of domain. Nevertheless, the statistical models are superficial descriptions of reality since they are only required to model dependence instead of the intrinsic causal mechanism. When the dependence changes with the target distribution, the statistic models may fail to generalize. In this regard, we introduce a general structural causal model to formalize the DG problem. Specifically, we assume that each input is constructed from a mix of causal factors (whose relationship with the label is invariant across domains) and non-causal factors (category-independent), and only the former cause the classification judgments. Our goal is to extract the causal factors from inputs and then reconstruct the invariant causal mechanisms. However, the theoretical idea is far from practical of DG since the required causal/non-causal factors are unobserved. We highlight that ideal causal factors should meet three basic properties: separated from the non-causal ones, jointly independent, and causally sufficient for the classification. Based on that, we propose a Causality Inspired Representation Learning (CIRL) algorithm that enforces the representations to satisfy the above properties and then uses them to simulate the causal factors, which yields improved generalization ability. Extensive experimental results on several widely used datasets verify the effectiveness of our approach.11Code is available at “https://github.com/BIT-DA/CIRL”. Fangrui Lv, Jian Liang 0002, Shuang Li 0008, Bin Zang, Chi Harold Liu, Di Liu 0029 |
CVPR | 3 |
| 2022 | Towards Fewer Annotations: Active Learning via Region Impurity and Prediction Uncertainty for Domain Adaptive Semantic SegmentationabstractSelf-training has greatly facilitated domain adaptive semantic segmentation, which iteratively generates pseudo labels on unlabeled target data and retrains the network. However, realistic segmentation datasets are highly imbalanced, pseudo labels are typically biased to the majority classes and basically noisy, leading to an error-prone and suboptimal model. In this paper, we propose a simple region-based active learning approach for semantic segmentation under a domain shift, aiming to automatically query a small partition of image regions to be labeled while maximizing segmentation performance. Our algorithm, Region Impurity and Prediction Uncertainty (RIPU), introduces a new acquisition strategy characterizing the spatial adjacency of image regions along with the prediction confidence. We show that the proposed region-based selection strategy makes more efficient use of a limited budget than image-based or point-based counterparts. Further, we enforce local prediction consistency between a pixel and its nearest neighbors on a source image. Alongside, we develop a negative learning loss to make the features more discriminative. Extensive experiments demonstrate that our method only requires very few annotations to almost reach the supervised performance and substantially outperforms state-of-the-art methods. The code is available at https://github.com/BIT-DA/RIPU. Binhui Xie, Longhui Yuan, Shuang Li 0008, Chi Harold Liu, Xinjing Cheng |
CVPR | 3 |
| 2022 | Improving Transferability for Domain Adaptive Detection TransformersabstractDETR-style detectors stand out amongst in-domain scenarios, but their properties in domain shift settings are under-explored. This paper aims to build a simple but effective baseline with a DETR-style detector on domain shift settings based on two findings. For one, mitigating the domain shift on the backbone and the decoder output features excels in getting favorable results. For another, advanced domain alignment methods in both parts further enhance the performance. Thus, we propose the Object-Aware Alignment (OAA) module and the Optimal Transport based Alignment (OTA) module to achieve comprehensive domain alignment on the outputs of the backbone and the detector. The OAA module aligns the foreground regions identified by pseudo-labels in the backbone outputs, leading to domain-invariant base features. The OTA module utilizes sliced Wasserstein distance to maximize the retention of location information while minimizing the domain gap in the decoder outputs. We implement the findings and the alignment modules into our adaptation method, and it benchmarks the DETR-style detector on the domain shift settings. Experiments on various domain adaptive scenarios validate the effectiveness of our method. Kaixiong Gong, Shuang Li 0008, Shugang Li 0002, Rui Zhang 0113, Chi Harold Liu |
ACM Multimedia | 2 |
| 2022 | Making The Best of Both Worlds: A Domain-Oriented Transformer for Unsupervised Domain AdaptationabstractExtensive studies on Unsupervised Domain Adaptation (UDA) have propelled the deployment of deep learning from limited experimental datasets into real-world unconstrained domains. Most UDA approaches align features within a common embedding space and apply a shared classifier for target prediction. However, since a perfectly aligned feature space may not exist when the domain discrepancy is large, these methods suffer from two limitations. First, the coercive domain alignment deteriorates target domain discriminability due to lacking target label supervision. Second, the source-supervised classifier is inevitably biased to source data, thus it may underperform in target domain. To alleviate these issues, we propose to simultaneously conduct feature alignment in two individual spaces focusing on different domains, and create for each space a domain-oriented classifier tailored specifically for that domain. Specifically, we design a Domain-Oriented Transformer (DOT) that has two individual classification tokens to learn different domain-oriented representations, and two classifiers to preserve domain-wise discriminability. Theoretical guaranteed contrastive-based alignment and the source-guided pseudo-label refinement strategy are utilized to explore both domain-invariant and specific information. Comprehensive experiments validate that our method achieves state-of-the-art on several benchmarks. Code is released at https://github.com/BIT-DA/Domain-Oriented-Transformer. Wenxuan Ma 0001, Shuang Li 0008, Chi Harold Liu, Yulin Wang 0002, Wei Li 0111 |
ACM Multimedia | 3 |
| 2022 | ROMA: Cross-Domain Region Similarity Matching for Unpaired Nighttime Infrared to Daytime Visible Video TranslationabstractInfrared cameras are often utilized to enhance the night vision since the visible light cameras exhibit inferior efficacy without sufficient illumination. However, infrared data possesses inadequate color contrast and representation ability attributed to its intrinsic heat-related imaging principle, which hinders its application. Although, the domain gaps between unpaired nighttime infrared and daytime visible videos are even huger than paired ones that captured at the same time, establishing an effective translation mapping will greatly contribute to various fields. In this case, the structural knowledge within nighttime infrared videos and semantic information contained in the translated daytime visible pairs could be utilized simultaneously. To this end, we propose a tailored framework ROMA that couples with our introduced cRoss-domain regiOn siMilarity mAtching technique for bridging the huge gaps. To be specific, ROMA could efficiently translate the unpaired nighttime infrared videos into fine-grained daytime visible ones, meanwhile maintain the spatiotemporal consistency via matching the cross-domain region similarity. Furthermore, we design a multiscale region-wise discriminator to distinguish the details from synthesized visible results and real references. Moreover, we provide a new and challenging dataset encouraging further research for unpaired nighttime infrared and daytime visible video translation, named InfraredCity, which is $20$ times larger than the recently released infrared-related dataset IRVI. Codes and datasets are available https://github.com/BIT-DA/ROMA here. Zhenjie Yu, Kai Chen 0030, Shuang Li 0008, Bingfeng Han, Chi Harold Liu, Shuigen Wang |
ACM Multimedia | 3 |
| 2022 | Generalized Domain Conditioned Adaptation NetworkabstractDomain adaptation (DA) attempts to transfer knowledge learned in the labeled source domain to the unlabeled but related target domain without requiring large amounts of target supervision. Recent advances in DA mainly proceed by aligning the source and target distributions. Despite the significant success, the adaptation performance still degrades accordingly when the source and target domains encounter a large distribution discrepancy. We consider this limitation may attribute to the insufficient exploration of domain-specialized features because most studies merely concentrate on domain-general feature learning in task-specific layers and integrate totally-shared convolutional networks (convnets) to generate common features for both domains. In this paper, we relax the completely-shared convnets assumption adopted by previous DA methods and propose Domain Conditioned Adaptation Network (DCAN), which introduces domain conditioned channel attention module with a multi-path structure to separately excite channel activation for each domain. Such a partially-shared convnets module allows domain-specialized features in low-level to be explored appropriately. Further, given the knowledge transferability varying along with convolutional layers, we develop Generalized Domain Conditioned Adaptation Network (GDCAN) to automatically determine whether domain channel activations should be separately modeled in each attention module. Afterward, the critical domain-specialized knowledge could be adaptively extracted according to the domain statistic gaps. As far as we know, this is the first work to explore the domain-wise convolutional channel activations separately for deep DA networks. Additionally, to effectively match high-level feature distributions across domains, we consider deploying feature adaptation blocks after task-specific layers, which can explicitly mitigate the domain discrepancy. Extensive experiments on four cross-domain benchmarks, including DomainNet, Office-Home, Office-31, and ImageCLEF, demonstrate the proposed approaches outperform the existing methods by a large margin, especially on the large-scale challenging dataset. The code and models are available at https://github.com/BIT-DA/GDCAN. Shuang Li 0008, Binhui Xie, Qiuxia Lin, Chi Harold Liu, Gao Huang 0001, Guoren Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Bi-Classifier Determinacy Maximization for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation challenges the problem of transferring knowledge from a well-labelled source domain to an unlabelled target domain. Recently, adversarial learning with bi-classifier has been proven effective in pushing cross-domain distributions close. Prior approaches typically leverage the disagreement between bi-classifier to learn transferable representations, however, they often neglect the classifier determinacy in the target domain, which could result in a lack of feature discriminability. In this paper, we present a simple yet effective method, namely Bi-Classifier Determinacy Maximization (BCDM), to tackle this problem. Motivated by the observation that target samples cannot always be separated distinctly by the decision boundary, here in the proposed BCDM, we design a novel classifier determinacy disparity (CDD) metric, which formulates classifier discrepancy as the class relevance of distinct target predictions and implicitly introduces constraint on the target feature discriminability. To this end, the BCDM can generate discriminative representations by encouraging target predictive outputs to be consistent and determined, meanwhile, preserve the diversity of predictions in an adversarial manner. Furthermore, the properties of CDD as well as the theoretical guarantees of BCDM's generalization bound are both elaborated. Extensive experiments show that BCDM compares favorably against the existing state-of-the-art domain adaptation methods. Shuang Li 0008, Fangrui Lv, Binhui Xie, Chi Harold Liu, Jian Liang 0002, Chen Qin |
AAAI | 1 |
| 2021 | One-shot Face Reenactment Using Appearance Adaptive NormalizationabstractThe paper proposes a novel generative adversarial network for one-shot face reenactment, which can animate a single face image to a different pose-and-expression (provided by a driving image) while keeping its original appearance. The core of our network is a novel mechanism called appearance adaptive normalization, which can effectively integrate the appearance information from the input image into our face generator by modulating the feature maps of the generator using the learned adaptive parameters. Furthermore, we specially design a local net to reenact the local facial components (i.e., eyes, nose and mouth) first, which is a much easier task for the network to learn and can in turn provide explicit anchors to guide our face generator to learn the global appearance and pose-and-expression. Extensive quantitative and qualitative experiments demonstrate the significant efficacy of our model compared with prior one-shot methods. Guangming Yao, Yi Yuan 0002, Tianjia Shao, Shuang Li 0008, Shanqi Liu, Yong Liu 0007, Mengmeng Wang 0005, Kun Zhou 0001 |
AAAI | 4 |
| 2021 | MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual RecognitionabstractReal-world training data usually exhibits long-tailed distribution, where several majority classes have a significantly larger number of samples than the remaining minority classes. This imbalance degrades the performance of typical supervised learning algorithms designed for balanced training sets. In this paper, we address this issue by augmenting minority classes with a recently proposed implicit semantic data augmentation (ISDA) algorithm [37], which produces diversified augmented samples by translating deep features along many semantically meaningful directions. Importantly, given that ISDA estimates the class-conditional statistics to obtain semantic directions, we find it ineffective to do this on minority classes due to the insufficient training data. To this end, we propose a novel approach to learn transformed semantic directions with meta-learning automatically. In specific, the augmentation strategy during training is dynamically optimized, aiming to minimize the loss on a small balanced validation set, which is approximated via a meta update step. Extensive empirical results on CIFAR-LT-10/100, ImageNet-LT, and iNaturalist 2017/2018 validate the effectiveness of our method. Shuang Li 0008, Kaixiong Gong, Chi Harold Liu, Yulin Wang 0002, Feng Qiao 0001, Xinjing Cheng |
CVPR | 1 |
| 2021 | Transferable Semantic Augmentation for Domain AdaptationabstractDomain adaptation has been widely explored by transferring the knowledge from a label-rich source domain to a related but unlabeled target domain. Most existing domain adaptation algorithms attend to adapting feature representations across two domains with the guidance of a shared source-supervised classifier. However, such classifier limits the generalization ability towards unlabeled target recognition. To remedy this, we propose a Transferable Semantic Augmentation (TSA) approach to enhance the classifier adaptation ability through implicitly generating source features towards target semantics. Specifically, TSA is inspired by the fact that deep feature transformation towards a certain direction can be represented as meaningful semantic altering in the original input space. Thus, source features can be augmented to effectively equip with target semantics to train a more transferable classifier. To achieve this, for each class, we first use the inter-domain feature mean difference and target intra-class feature covariance to construct a multivariate normal distribution. Then we augment source features with random directions sampled from the distribution class-wisely. Interestingly, such source augmentation is implicitly implemented through an expected transferable cross-entropy loss over the augmented source distribution, where an upper bound of the expected loss is derived and minimized, introducing negligible computational overhead. As a light-weight and general technique, TSA can be easily plugged into various domain adaptation methods, bringing remarkable improvements. Comprehensive experiments on cross-domain benchmarks validate the efficacy of TSA. Shuang Li 0008, Mixue Xie, Kaixiong Gong, Chi Harold Liu, Yulin Wang 0002, Wei Li 0111 |
CVPR | 1 |
| 2021 | Dynamic Domain Adaptation for Efficient InferenceabstractDomain adaptation (DA) enables knowledge transfer from a labeled source domain to an unlabeled target domain by reducing the cross-domain distribution discrepancy. Most prior DA approaches leverage complicated and powerful deep neural networks to improve the adaptation capacity and have shown remarkable success. However, they may have a lack of applicability to real-world situations such as real-time interaction, where low target inference latency is an essential requirement under limited computational budget. In this paper, we tackle the problem by proposing a dynamic domain adaptation (DDA) framework, which can simultaneously achieve efficient target inference in low-resource scenarios and inherit the favorable cross-domain generalization brought by DA. In contrast to static models, as a simple yet generic method, DDA can integrate various domain confusion constraints into any typical adaptive network, where multiple intermediate classifiers can be equipped to infer “easier” and “harder” target data dynamically. Moreover, we present two novel strategies to further boost the adaptation performance of multiple prediction exits: 1) a confidence score learning strategy to derive accurate target pseudo labels by fully exploring the prediction consistency of different classifiers; 2) a class-balanced self-training strategy to explicitly adapt multi-stage classifiers from source to target without losing prediction diversity. Extensive experiments on multiple benchmarks are conducted to verify that DDA can consistently improve the adaptation performance and accelerate target inference under domain shift and limited resources scenarios. Shuang Li 0008, Wenxuan Ma 0001, Chi Harold Liu, Wei Li 0111 |
CVPR | 1 |
| 2021 | SOE-Net: A Self-Attention and Orientation Encoding Network for Point Cloud Based Place RecognitionabstractWe tackle the problem of place recognition from point cloud data and introduce a self-attention and orientation encoding network (SOE-Net) that fully explores the relationship between points and incorporates long-range context into point-wise local descriptors. Local information of each point from eight orientations is captured in a PointOE module, whereas long-range feature dependencies among local descriptors are captured with a self-attention unit. Moreover, we propose a novel loss function called Hard Positive Hard Negative quadruplet loss (HPHN quadruplet), that achieves better performance than the commonly used metric learning loss. Experiments on various benchmark datasets demonstrate superior performance of the proposed network over the current state-of-the-art approaches. Our code is released publicly at https://github.com/Yan-Xia/SOE-Net. Yan Xia 0003, Yusheng Xu, Shuang Li 0008, Rui Wang 0037, Juan Du 0012, Daniel Cremers, Uwe Stilla |
CVPR | 3 |
| 2021 | Semantic Concentration for Domain AdaptationabstractDomain adaptation (DA) paves the way for label annotation and dataset bias issues by the knowledge transfer from a label-rich source domain to a related but unlabeled target domain. A mainstream of DA methods is to align the feature distributions of the two domains. However, the majority of them focus on the entire image features where irrelevant semantic information, e.g., the messy background, is inevitably embedded. Enforcing feature alignments in such case will negatively influence the correct matching of objects and consequently lead to the semantically negative transfer due to the confusion of irrelevant semantics. To tackle this issue, we propose Semantic Concentration for Domain Adaptation (SCDA), which encourages the model to concentrate on the most principal features via the pair-wise adversarial alignment of prediction distributions. Specifically, we train the classifier to class-wisely maximize the prediction distribution divergence of each sample pair, which enables the model to find the region with large differences among the same class of samples. Meanwhile, the feature extractor attempts to minimize that discrepancy, which suppresses the features of dissimilar regions among the same class of samples and accentuates the features of principal parts. As a general method, SCDA can be easily integrated into various DA methods as a regularizer to further boost their performance. Extensive experiments on the cross-domain benchmarks show the efficacy of SCDA. Shuang Li 0008, Mixue Xie, Fangrui Lv, Chi Harold Liu, Jian Liang 0002, Chen Qin, Wei Li 0111 |
ICCV | 1 |
| 2021 | I2V-GAN: Unpaired Infrared-to-Visible Video TranslationabstractHuman vision is often adversely affected by complex environmental factors, especially in night vision scenarios. Thus, infrared cameras are often leveraged to help enhance the visual effects via detecting infrared radiation in the surrounding environment, but the infrared videos are undesirable due to the lack of detailed semantic information. In such a case, an effective video-to-video translation method from the infrared domain to the visible light counterpart is strongly needed by overcoming the intrinsic huge gap between infrared and visible fields. To address this challenging problem, we propose an infrared-to-visible (I2V) video translation method I2V-GAN to generate fine-grained and spatial-temporal consistent visible light videos by given unpaired infrared videos. Technically, our model capitalizes on three types of constraints: 1) adversarial constraint to generate synthetic frames that are similar to the real ones, 2) cyclic consistency with the introduced perceptual loss for effective content conversion as well as style preservation, and 3) similarity constraints across and within domains to enhance the content and motion consistency in both spatial and temporal spaces at a fine-grained level. Furthermore, the current public available infrared and visible light datasets are mainly used for object detection or tracking, and some are composed of discontinuous images which are not suitable for video tasks. Thus, we provide a new dataset for infrared-to-visible video translation, which is named IRVI. Specifically, it has 12 consecutive video clips of vehicle and monitoring scenes, and both infrared and visible light videos could be apart into 24352 frames. Comprehensive experiments on IRVI validate that I2V-GAN is superior to the compared state-of-the-art methods in the translation of infrared-to-visible videos with higher fluency and finer semantic details. Moreover, additional experimental results on the flower-to-flower dataset indicate I2V-GAN is also applicable to other video translation tasks. The code and IRVI dataset are available at https://github.com/BIT-DA/I2V-GAN. Shuang Li 0008, Bingfeng Han, Zhenjie Yu, Chi Harold Liu, Kai Chen 0030, Shuigen Wang |
ACM Multimedia | 1 |
| 2021 | ZiGAN: Fine-grained Chinese Calligraphy Font Generation via a Few-shot Style Transfer ApproachabstractChinese character style transfer is a very challenging problem because of the complexity of the glyph shapes or underlying structures and large numbers of existed characters, when comparing with English letters. Moreover, the handwriting of calligraphy masters has a more irregular stroke and is difficult to obtain in real-world scenarios. Recently, several GAN-based methods have been proposed for font synthesis, but some of them require numerous reference data and the other part of them have cumbersome preprocessing steps to divide the character into different parts to be learned and transferred separately. In this paper, we propose a simple but powerful end-to-end Chinese calligraphy font generation framework ZiGAN, which does not require any manual operation or redundant preprocessing to generate fine-grained target style characters with few-shot references. To be specific, a few paired samples from different character styles are leveraged to attain fine-grained correlation between structures underlying different glyphs. To capture valuable style knowledge in target and strengthen the coarse-grained understanding of character content, we utilize multiple unpaired samples to align the feature distributions belonging to different character styles. By doing so, only a few target Chinese calligraphy characters are needed to generated expected style transferred characters. Experiments demonstrate that our method has a state-of-the-art generalization ability in few-shot Chinese character style transfer. Shuang Li 0008, Bingfeng Han, Yi Yuan 0002 |
ACM Multimedia | 2 |
| 2021 | Pareto Domain AdaptationabstractDomain adaptation (DA) attempts to transfer the knowledge from a labeled source domain to an unlabeled target domain that follows different distribution from the source. To achieve this, DA methods include a source classification objective to extract the source knowledge and a domain alignment objective to diminish the domain shift, ensuring knowledge transfer. Typically, former DA methods adopt some weight hyper-parameters to linearly combine the training objectives to form an overall objective. However, the gradient directions of these objectives may conflict with each other due to domain shift. Under such circumstances, the linear optimization scheme might decrease the overall objective value at the expense of damaging one of the training objectives, leading to restricted solutions. In this paper, we rethink the optimization scheme for DA from a gradient-based perspective. We propose a Pareto Domain Adaptation (ParetoDA) approach to control the overall optimization direction, aiming to cooperatively optimize all training objectives. Specifically, to reach a desirable solution on the target domain, we design a surrogate loss mimicking target classification. To improve target-prediction accuracy to support the mimicking, we propose a target-prediction refining mechanism which exploits domain labels via Bayes’ theorem. On the other hand, since prior knowledge of weighting schemes for objectives is often unavailable to guide optimization to approach the optimal solution on the target domain, we propose a dynamic preference mechanism to dynamically guide our cooperative optimization by the gradient of the surrogate loss on a held-out unlabeled target dataset. Our theoretical analyses show that the held-out data can guide but will not be over-fitted by the optimization. Extensive experiments on image classification and semantic segmentation benchmarks demonstrate the effectiveness of ParetoDA Fangrui Lv, Jian Liang 0002, Kaixiong Gong, Shuang Li 0008, Chi Harold Liu, Han Li 0005, Di Liu 0029, Guoren Wang |
NeurIPS | 4 |
| 2021 | Deep Residual Correction Network for Partial Domain AdaptationabstractDeep domain adaptation methods have achieved appealing performance by learning transferable representations from a well-labeled source domain to a different but related unlabeled target domain. Most existing works assume source and target data share the identical label space, which is often difficult to be satisfied in many real-world applications. With the emergence of big data, there is a more practical scenario called partial domain adaptation, where we are always accessible to a more large-scale source domain while working on a relative small-scale target domain. In this case, the conventional domain adaptation assumption should be relaxed, and the target label space tends to be a subset of the source label space. Intuitively, reinforcing the positive effects of the most relevant source subclasses and reducing the negative impacts of irrelevant source subclasses are of vital importance to address partial domain adaptation challenge. This paper proposes an efficiently-implemented Deep Residual Correction Network (DRCN) by plugging one residual block into the source network along with the task-specific feature layer, which effectively enhances the adaptation from source to target and explicitly weakens the influence from the irrelevant source classes. Specifically, the plugged residual block, which consists of several fully-connected layers, could deepen basic network and boost its feature representation capability correspondingly. Moreover, we design a weighted class-wise domain alignment loss to couple two domains by matching the feature distributions of shared classes between source and target. Comprehensive experiments on partial, traditional and fine-grained cross-domain visual recognition demonstrate that DRCN is superior to the competitive deep domain adaptation approaches. Shuang Li 0008, Chi Harold Liu, Qiuxia Lin, Limin Su, Gao Huang 0001, Zhengming Ding |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Discriminative Dimension Reduction via Maximin Separation Probability AnalysisabstractIn this paper, we propose a novel discriminative dimension reduction (DR) method, maximin separation probability analysis (MSPA), which maximizes the minimum separation probability of all classes in the reduced low-dimensional subspace. Separation probability is a novel class separability measure, which gives a lower bound of the generalization accuracy for a learned linear classifier in a binary classification problem. The proposed MSPA duly considers the separation of all class pairs in multiclass linear discriminant analysis (LDA) and thus improves the subsequent classification performance. DR via MSPA leads to a nonconvex optimization problem. We develop an algorithm to solve the problem and the global optimal solution can be found by converting the original problem into a series of second-order cone programming problems. A low-computational cost extension and a non-LDA with kernel mapping of MSPA are also provided in this paper. The experimental results on 14 real-world datasets show our methods are superior to other state-of-the-art algorithms in discriminative DR tasks. Le Yang 0007, Shiji Song, Shuang Li 0008, Yiming Chen 0005, C. L. Philip Chen |
IEEE Trans. Cybern. | 3 |
| 2021 | Graph Embedding-Based Dimension Reduction With Extreme Learning MachineabstractDimension reduction (DR)-based on extreme learning machine auto-encoder (ELM-AE) has achieved many successes in recent years. By minimizing the self-reconstruction error, the ELM-AE-based DR algorithms learn the compressed representations which facilitate the subsequent classification. However, the existing ELM-AEs only consider the DR problem in an unsupervised manner and ignore the valuable supervised information when these information is available. To find discriminative features of the original data, in this paper, we propose a graph embedding-based DR framework with ELM (GDR-ELM) for DR problems. Instead of self-reconstruction, the proposed GDR-ELM reconstructs all samples according to the weights in a graph matrix containing the supervised information. Furthermore, GDR-ELM can be stacked as building blocks to construct a multilayer framework like other ELM-AEs for more complicated representation learning tasks. Experiments on various datasets demonstrate the effectiveness of the proposed GDR-ELM and its multilayer framework. Le Yang 0007, Shiji Song, Shuang Li 0008, Yiming Chen 0005, Gao Huang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Domain Conditioned Adaptation NetworkabstractTremendous research efforts have been made to thrive deep domain adaptation (DA) by seeking domain-invariant features. Most existing deep DA models only focus on aligning feature representations of task-specific layers across domains while integrating a totally shared convolutional architecture for source and target. However, we argue that such strongly-shared convolutional layers might be harmful for domain-specific feature learning when source and target data distribution differs to a large extent. In this paper, we relax a shared-convnets assumption made by previous DA methods and propose a Domain Conditioned Adaptation Network (DCAN), which aims to excite distinct convolutional channels with a domain conditioned channel attention mechanism. As a result, the critical low-level domain-dependent knowledge could be explored appropriately. As far as we know, this is the first work to explore the domain-wise convolutional channel activation for deep DA networks. Moreover, to effectively align high-level feature distributions across two domains, we further deploy domain conditioned feature correction blocks after task-specific layers, which will explicitly correct the domain discrepancy. Extensive experiments on three cross-domain benchmarks demonstrate the proposed approach outperforms existing methods by a large margin, especially on very tough cross-domain learning tasks. Shuang Li 0008, Chi Harold Liu, Qiuxia Lin, Binhui Xie, Zhengming Ding, Gao Huang 0001, Jian Tang 0008 |
AAAI | 1 |
| 2020 | Simultaneous Semantic Alignment Network for Heterogeneous Domain AdaptationabstractHeterogeneous domain adaptation (HDA) transfers knowledge across source and target domains that present heterogeneities e.g., distinct domain distributions and difference in feature type or dimension. Most previous HDA methods tackle this problem through learning a domain-invariant feature subspace to reduce the discrepancy between domains. However, the intrinsic semantic properties contained in data are under-explored in such alignment strategy, which is also indispensable to achieve promising adaptability. In this paper, we propose a Simultaneous Semantic Alignment Network (SSAN) to simultaneously exploit correlations among categories and align the centroids for each category across domains. In particular, we propose an implicit semantic correlation loss to transfer the correlation knowledge of source categorical prediction distributions to target domain. Meanwhile, by leveraging target pseudo-labels, a robust triplet-centroid alignment mechanism is explicitly applied to align feature representations for each category. Notably, a pseudo-label refinement procedure with geometric similarity involved is introduced to enhance the target pseudo-label assignment accuracy. Comprehensive experiments on various HDA tasks across text-to-image, image-to-image and text-to-text successfully validate the superiority of our SSAN against state-of-the-art HDA methods. The code is publicly available at https://github.com/BIT-DA/SSAN. Shuang Li 0008, Binhui Xie, Jiashu Wu, Chi Harold Liu, Zhengming Ding |
ACM Multimedia | 1 |
| 2020 | A Graph Embedding Framework for Maximum Mean Discrepancy-Based Domain Adaptation AlgorithmsabstractDomain adaptation aims to deal with learning problems in which the labeled training data and unlabeled testing data are differently distributed. Maximum mean discrepancy (MMD), as a distribution distance measure, is minimized in various domain adaptation algorithms for eliminating domain divergence. We analyze empirical MMD from the point of view of graph embedding. It is discovered from the MMD intrinsic graph that, when the empirical MMD is minimized, the compactness within each domain and each class is simultaneously reduced. Therefore, points from different classes may mutually overlap, leading to unsatisfactory classification results. To deal with this issue, we present a graph embedding framework with intrinsic and penalty graphs for MMD-based domain adaptation algorithms. In the framework, we revise the intrinsic graph of MMD-based algorithms such that the within-class scatter is minimized, and thus, the new features are discriminative. Two strategies are proposed. Based on the strategies, we instantiate the framework by exploiting four models. Each model has a penalty graph characterizing certain similarity property that should be avoided. Comprehensive experiments on visual cross-domain benchmark datasets demonstrate that the proposed models can greatly enhance the classification performance compared with the state-of-the-art methods. Yiming Chen 0005, Shiji Song, Shuang Li 0008, Cheng Wu 0002 |
IEEE Trans. Image Process. | 3 |
| 2020 | Discriminative Transfer Feature and Label Consistency for Cross-Domain Image ClassificationabstractVisual domain adaptation aims to seek an effective transferable model for unlabeled target images by benefiting from the well-labeled source images following different distributions. Many recent efforts focus on extracting domain-invariant image representations via exploring target pseudo labels, predicted by the source classifier, to further mitigate the conditional distribution shift across domains. However, two essential factors are overlooked by most existing methods: 1) the learned transferable features should be not only domain invariant but also category discriminative; and 2) the target pseudo label is a two-edged sword to cross-domain alignment. In other words, the wrongly predicted target labels may hinder the class-wise domain matching. In this article, to address these two issues simultaneously, we propose a discriminative transfer feature and label consistency (DTLC) approach for visual domain adaptation problems, which can naturally unify cross-domain alignment with discriminative information preserved and label consistency of source and target data into one framework. To be specific, DTLC first incorporates class discriminative information by penalizing the maximum distance of data pair in the same class and the minimum distance of data pair sharing the different labels for each data into the distribution alignment of both domains. The target pseudo labels are then refined based on the label consistency within the domains. Thus, the transfer feature learning and coarse-to-fine target labels would be coupled to benefit each other in an iterative way. Comprehensive experiments on several visual cross-domain benchmarks verify that DTLC can gain remarkable margins over state-of-the-art (SOTA) nondeep visual domain adaptation methods and even be comparable to competitive deep domain adaptation ones. Shuang Li 0008, Chi Harold Liu, Limin Su, Binhui Xie, Zhengming Ding, C. L. Philip Chen, Dapeng Oliver Wu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Joint Adversarial Domain AdaptationabstractDomain adaptation aims to transfer the enriched label knowledge from large amounts of source data to unlabeled target data. It has raised significant interest in multimedia analysis. Existing researches mainly focus on learning domain-wise transferable representations via statistical moment matching or adversarial adaptation techniques, while ignoring the class-wise mismatch across domains, resulting in inaccurate distribution alignment. To address this issue, we propose a Joint Adversarial Domain Adaptation (JADA) approach to simultaneously align domain-wise and class-wise distributions across source and target in a unified adversarial learning process. Specifically, JADA attempts to solve two complementary minimax problems jointly. The feature generator aims to not only fool the well-trained domain discriminator to learn domain-invariant features, but also minimize the disagreement between two distinct task-specific classifiers' predictions to synthesize target features near the support of source class-wisely. As a result, the learned transferable features will be equipped with more discriminative structures, and effectively avoid mode collapse. Additionally, JADA enables an efficient end-to-end training manner via a simple back-propagation scheme. Extensive experiments on several real-world cross-domain benchmarks, including VisDA-2017, ImageCLEF, Office-31 and digits, verify that JADA can gain remarkable improvements over other state-of-the-art deep domain adaptation approaches. Shuang Li 0008, Chi Harold Liu, Binhui Xie, Limin Su, Zhengming Ding, Gao Huang 0001 |
ACM Multimedia | 1 |
| 2019 | Domain Space Transfer Extreme Learning Machine for Domain AdaptationabstractExtreme learning machine (ELM) has been applied in a wide range of classification and regression problems due to its high accuracy and efficiency. However, ELM can only deal with cases where training and testing data are from identical distribution, while in real world situations, this assumption is often violated. As a result, ELM performs poorly in domain adaptation problems, in which the training data (source domain) and testing data (target domain) are differently distributed but somehow related. In this paper, an ELM-based space learning algorithm, domain space transfer ELM (DST-ELM), is developed to deal with unsupervised domain adaptation problems. To be specific, through DST-ELM, the source and target data are reconstructed in a domain invariant space with target data labels unavailable. Two goals are achieved simultaneously. One is that, the target data are input into an ELM-based feature space learning network, and the output is supposed to approximate the input such that the target domain structural knowledge and the intrinsic discriminative information can be preserved as much as possible. The other one is that, the source data are projected into the same space as the target data and the distribution distance between the two domains is minimized in the space. This unsupervised feature transformation network is followed by an adaptive ELM classifier which is trained from the transferred labeled source samples, and is used for target data label prediction. Moreover, the ELMs in the proposed method, including both the space learning ELM and the classifier, require just a small number of hidden nodes, thus maintaining low computation complexity. Extensive experiments on real-world image and text datasets are conducted and verify that our approach outperforms several existing domain adaptation methods in terms of accuracy while maintaining high efficiency. Yiming Chen 0005, Shiji Song, Shuang Li 0008, Le Yang 0007, Cheng Wu 0002 |
IEEE Trans. Cybern. | 3 |
| 2019 | Cross-Domain Extreme Learning Machines for Domain AdaptationabstractExtreme learning machines (ELMs), as “generalized” single hidden layer feedforward networks, have been proved to be effective and efficient for classification and regression problems. Traditional ELMs assume that the training and testing data are drawn from the same distribution, which however is often violated in real-world applications. In this paper, we propose a unified cross-domain ELM (CDELM) framework to address domain adaptation problems, in which the distributions of training data (source domain) and testing data (target domain) are different but related. CDELM not only fully leverages labeled source data and unlabeled target data simultaneously to construct an adaptive target classifier but also maintains the computational efficiency of ELMs. Specifically, CDELM adapts the source classifier to target domain by matching the projected means of both domains, and explores the structure property of target domain by using manifold regularization to make the final classifier more adaptable to target data. Based on the framework, two algorithms CDELM-M and CDELM-C are proposed, which aim at minimizing the marginal and conditional distribution distance between source and target domains, respectively. Moreover, CDELM-C can further enhance the classification accuracy by multiple iterations. Comprehensive experimental studies on artificial datasets and public text and image datasets demonstrate that both CDELM-M and CDELM-C are competitive with several state-of-the-art domain adaptation learning methods in terms of the classification accuracy and efficiency. Shuang Li 0008, Shiji Song, Gao Huang 0001, Cheng Wu 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2018 | Laplacian twin extreme learning machine for semi-supervised classification
Shuang Li 0008, Shiji Song, Yihe Wan |
Neurocomputing | 1 |
| 2018 | Layer-wise domain correction for unsupervised domain adaptationabstractDeep neural networks have been successfully applied to numerous machine learning tasks because of their impressive feature abstraction capabilities. However, conventional deep networks assume that the training and test data are sampled from the same distribution, and this assumption is often violated in real-world scenarios. To address the domain shift or data bias problems, we introduce layer-wise domain correction (LDC), a new unsupervised domain adaptation algorithm which adapts an existing deep network through additive correction layers spaced throughout the network. Through the additive layers, the representations of source and target domains can be perfectly aligned. The corrections that are trained via maximum mean discrepancy, adapt to the target domain while increasing the representational capacity of the network. LDC requires no target labels, achieves state-of-the-art performance across several adaptation benchmarks, and requires significantly less training time than existing adaptation methods. Shuang Li 0008, Shiji Song, Cheng Wu 0002 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2018 | Domain Invariant and Class Discriminative Feature Learning for Visual Domain AdaptationabstractDomain adaptation manages to build an effective target classifier or regression model for unlabeled target data by utilizing the well-labeled source data but lying different distributions. Intuitively, to address domain shift problem, it is crucial to learn domain invariant features across domains, and most existing approaches have concentrated on it. However, they often do not directly constrain the learned features to be class discriminative for both source and target data, which is of vital importance for the final classification. Therefore, in this paper, we put forward a novel feature learning method for domain adaptation to construct both domain invariant and class discriminative representations, referred to as DICD. Specifically, DICD is to learn a latent feature space with important data properties preserved, which reduces the domain difference by jointly matching the marginal and class-conditional distributions of both domains, and simultaneously maximizes the inter-class dispersion and minimizes the intra-class scatter as much as possible. Experiments in this paper have demonstrated that the class discriminative properties will dramatically alleviate the cross-domain distribution inconsistency, which further boosts the classification performance. Moreover, we show that exploring both domain invariance and class discriminativeness of the learned representations can be integrated into one optimization framework, and the optimal solution can be derived effectively by solving a generalized eigen-decomposition problem. Comprehensive experiments on several visual cross-domain classification tasks verify that DICD can outperform the competitors significantly. Shuang Li 0008, Shiji Song, Gao Huang 0001, Zhengming Ding, Cheng Wu 0002 |
IEEE Trans. Image Process. | 1 |
| 2017 | Twin extreme learning machines for pattern classification
Yihe Wan, Shiji Song, Gao Huang 0001, Shuang Li 0008 |
Neurocomputing | 4 |
| 2017 | Prediction Reweighting for Domain AdaptationabstractThere are plenty of classification methods that perform well when training and testing data are drawn from the same distribution. However, in real applications, this condition may be violated, which causes degradation of classification accuracy. Domain adaptation is an effective approach to address this problem. In this paper, we propose a general domain adaptation framework from the perspective of prediction reweighting, from which a novel approach is derived. Different from the major domain adaptation methods, our idea is to reweight predictions of the training classifier on testing data according to their signed distance to the domain separator, which is a classifier that distinguishes training data (from source domain) and testing data (from target domain). We then propagate the labels of target instances with larger weights to ones with smaller weights by introducing a manifold regularization method. It can be proved that our reweighting scheme effectively brings the source and target domains closer to each other in an appropriate sense, such that classification in target domain becomes easier. The proposed method can be implemented efficiently by a simple two-stage algorithm, and the target classifier has a closed-form solution. The effectiveness of our approach is verified by the experiments on artificial datasets and two standard benchmarks, a visual object recognition task and a cross-domain sentiment analysis of text. Experimental results demonstrate that our method is competitive with the state-of-the-art domain adaptation algorithms. Shuang Li 0008, Shiji Song, Gao Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |