Zhekai Du

dblp:259/2953 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0002-9406-3920ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Generalizing Vision-Language Models with Dedicated Prompt Guidance
abstract
Fine-tuning large pretrained vision-language models (VLMs) has emerged as a prevalent paradigm for downstream adaptation, yet it faces a critical trade-off between domain specificity and domain generalization (DG) ability. Current methods typically fine-tune a universal model on the entire dataset, which potentially compromises the ability to generalize to unseen domains. To fill this gap, we provide a theoretical understanding of the generalization ability for VLM fine-tuning, which reveals that training multiple parameter-efficient expert models on partitioned source domains leads to better generalization than fine-tuning a universal model. Inspired by this finding, we propose a two-step domain-expert-Guided DG (GuiDG) framework. GuiDG first employs prompt tuning to obtain source domain experts, then introduces a Cross-Modal Attention module to guide the fine-tuning of the vision encoder via adaptive expert integration. To better evaluate few-shot DG, we construct ImageNet-DG from ImageNet and its variants. Extensive experiments on standard DG benchmarks and ImageNet-DG demonstrate that GuiDG improves upon state-of-the-art fine-tuning methods while maintaining efficiency.
Yinjie Min, Zhekai Du, Fengling Li 0001, Jingjing Li 0001
AAAI4
2026 Sharpness-Consistent Cross-Domain Recommendation for Cold-Start Items
abstract
Cold-start remains a fundamental challenge in recommendation systems due to the scarcity of interaction data. Recent methods address this issue by leveraging semantic ID embeddings and cross-domain transfer techniques, achieving notable progress. However, the common practice of learning semantic ID embeddings and training the recommendation model in separate stages hinders the generalization capability of semantic IDs throughout the training process. In this work, we propose Sharpness-Consistent Cross-Domain Recommendation (SC2 Rec), a novel framework designed to enhance the generalization of semantic ID-based models in cold-start scenarios. SC2 Rec alternately optimizes the sharpness of the loss landscape and enforces landscape consistency between warm and cold domains, leading to unified and flatter minima and improved generalization. Extensive experiments on industrial datasets demonstrate the effectiveness of SC2 Rec. Furthermore, we release a high-quality dataset to facilitate further research in this area.
Ke Fei, Jingjing Li 0001, Zhekai Du
WWW3
2026 FAST: Foreground-aware active self-training for domain adaptive object detection
Hongmin Deng, Hailin Wang 0002, Zhekai Du, Guisong Liu, Jingjing Li 0001, Mao Ye 0001
Neural Networks4
2025 LoCA: Location-Aware Cosine Adaptation for Parameter-Efficient Fine-Tuning
abstract
Low-rank adaptation (LoRA) has become a prevalent method for adapting pre-trained large language models to downstream tasks. However, the simple low-rank decomposition form may constrain the optimization flexibility. To address this limitation, we introduce Location-aware Cosine Adaptation (LoCA), a novel frequency-domain parameter-efficient fine-tuning method based on inverse Discrete Cosine Transform (iDCT) with selective locations of learnable components. We begin with a comprehensive theoretical comparison between frequency-domain and low-rank decompositions for fine-tuning pre-trained large models. Our analysis reveals that frequency-domain decomposition with carefully selected frequency components can surpass the expressivity of traditional low-rank-based methods. Furthermore, we demonstrate that iDCT offers a more efficient implementation compared to inverse Discrete Fourier Transform (iDFT), allowing for better selection and tuning of frequency components while maintaining equivalent expressivity to the optimal iDFT-based adaptation. By employing finite-difference approximation to estimate gradients for discrete locations of learnable coefficients on the DCT spectrum, LoCA dynamically selects the most informative frequency components during training. Experiments on diverse language and vision fine-tuning tasks demonstrate that LoCA offers enhanced parameter efficiency while maintains computational feasibility comparable to low-rank-based methods.
Zhekai Du, Yinjie Min, Jingjing Li 0001, Ke Lu 0001, Changliang Zou, Liuhua Peng, Tingjin Chu, Mingming Gong
ICLR1
2025 PatAug: Augmentation of Augmentation for Test-Time Adaptation
abstract
The rich pretrained knowledge in vision-language models (VLMs) endows them with the ability to discriminate common objects given only category names, but may be challenged by out-of-distribution unlabeled samples. To address this limitation, test-time adaptation (TTA) dynamically adjusts VLMs to target distributions during inference. Current TTA frameworks rely heavily on unsupervised data augmentations to enhance sample informativeness, but remain vulnerable to naive augmented views. This work introduces Patch Augmentation (PatAug), a pixel-level perturbation framework that optimizes the benefits of informative augmentations and mitigates negative transformation impacts. Implemented as trainable pixels, PatAug are prepared given only category names before inference, introducing few additional overheads. The patches encode class-related semantic information. They assist VLMs in emphasizing on the compatible visual information in the images, restoring perturbed image details, while retaining unrecognized information. Such merits inspire the design of an augmentation of augmentation framework, where PatAug is applied to standard augmentation views for reliable TTA inference results. To better fit the target distributions, we adjust patches with a cross-modal similarity alignment loss and learnable patching weights. Experiments on natural and specialized domain shifts confirm the effectiveness of PatAug.
Zhekai Du, Lei Zhu 0002, Zhi Chen 0010, Jingjing Li 0001
ACM Multimedia3
2025 Unified Modality Separation: A Vision-Language Framework for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) enables models trained on a labeled source domain to handle new unlabeled domains. Recently, pre-trained vision-language models (VLMs) have demonstrated promising zero-shot performance by leveraging semantic information to facilitate target tasks. By aligning vision and text embeddings, VLMs have shown notable success in bridging domain gaps. However, inherent differences naturally exist between modalities, which is known as modality gap. Our findings reveal that direct UDA with the presence of modality gap only transfers modality-invariant knowledge, leading to suboptimal target performance. To address this limitation, we propose a unified modality separation framework that accommodates both modality-specific and modality-invariant components. During training, different modality components are disentangled from VLM features then handled separately in a unified manner. At test time, modality-adaptive ensemble weights are automatically determined to maximize the synergy of different components. To evaluate instance-level modality characteristics, we design a modality discrepancy metric to categorize samples into modality-invariant, modality-specific, and uncertain ones. The modality-invariant samples are exploited to facilitate cross-modal alignment, while uncertain ones are annotated to enhance model capabilities. Building upon prompt tuning techniques, our methods achieve up to 9% performance gain with 9 times of computational efficiencies. Extensive experiments and analysis across various backbones, baselines, datasets and adaptation settings demonstrate the efficacy of our design.
Jingjing Li 0001, Zhekai Du, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Fourier-Based Frequency Space Disentanglement and Augmentation for Generalizable Face Anti-Spoofing
abstract
Generalizing face anti-spoofing (FAS) models to unseen distributions is challenging due to domain shifts. Previous domain generalization (DG) based FAS methods focus on learning invariant features across domains in the spatial space, which may be ineffective in detecting subtle spoof patterns. In this paper, we propose a novel approach called Frequency Space Disentanglement and Augmentation (FSDA) for generalizable FAS. Specifically, we leverage Fourier transformation to analyze face images in the frequency space, where the amplitude spectrum captures low-level texture information that forms distinct visual appearances, and the phase spectrum corresponds to the content information. We hypothesize that the liveness of a face is more related to these low-level patterns rather than high-level content information. To locate spoof traces, we disentangle the amplitude spectrum into domain-related and spoof-related components using either empirical or learnable strategies. We then propose a frequency space augmentation technique that mixes the disentangled components of two images to synthesize new variations. By imposing a distillation loss and a consistency loss on the augmented samples, our model learns to capture spoof patterns that are robust to both domain and spoof type variations. Extensive experiments on four FAS datasets demonstrate the superiority of our method in improving the generalization ability of FAS models in various unseen scenarios.
Zhekai Du, Chengwei Xiao
IEEE J. Biomed. Health Informatics2
2024 Domain-Agnostic Mutual Prompting for Unsupervised Domain Adaptation
abstract
Conventional Unsupervised Domain Adaptation (UDA) strives to minimize distribution discrepancy between do-mains, which neglects to harness rich semantics from data and struggles to handle complex domain shifts. A promising technique is to leverage the knowledge of large-scale pretrained vision-language models for more guided adaptation. Despite some endeavors, current methods often learn textual prompts to embed domain semantics for source and target domains separately and perform classification within each domain, limiting cross-domain knowledge transfer. Moreover, prompting only the language branch lacks flex-ibility to adapt both modalities dynamically. To bridge this gap, we propose Domain-Agnostic Mutual Prompting (DAMP) to exploit domain-invariant semantics by mutually aligning visual and textual embeddings. Specifically, the image contextual information is utilized to prompt the language branch in a domain-agnostic and instance-conditioned way. Meanwhile, visual prompts are im-posed based on the domain-agnostic textual prompt to elicit domain-invariant visual embeddings. These two branches of prompts are learned mutually with a cross-attention module and regularized with a semantic-consistency loss and an instance-discrimination contrastive loss. Experiments on three UDA benchmarks demonstrate the superiority of DAMP over state-of-the-art approaches1.
Zhekai Du, Fengling Li 0001, Ke Lu 0001, Lei Zhu 0002, Jingjing Li 0001
CVPR1
2024 Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
abstract
Large vision-language models (VLMs) like CLIP have demonstrated good zero-shot learning performance in the unsupervised domain adaptation task. Yet, most transfer approaches for VLMs focus on either the language or visual branches, overlooking the nuanced interplay between both modalities. In this work, we introduce a Unified Modality Separation (UniMoS) framework for unsupervised domain adaptation. Leveraging insights from modality gap studies, we craft a nimble modality separation network that distinctly disentangles CLIP's features into language-associated and vision-associated components. Our proposed Modality-Ensemble Training (MET) method fosters the exchange of modality-agnostic information while maintaining modality-specific nuances. We align features across domains using a modality discriminator. Comprehensive evaluations on three benchmarks reveal our approach sets a new state-of-the-art with minimal computational costs. Code: https://github.com/TL-UESTC/UniMoS.
Zhekai Du, Fengling Li 0001, Ke Lu 0001, Jingjing Li 0001
CVPR3
2024 A Comprehensive Survey on Source-Free Domain Adaptation
abstract
Over the past decade, domain adaptation has become a widely studied branch of transfer learning which aims to improve performance on target domains by leveraging knowledge from the source domain. Conventional domain adaptation methods often assume access to both source and target domain data simultaneously, which may not be feasible in real-world scenarios due to privacy and confidentiality concerns. As a result, the research of Source-Free Domain Adaptation (SFDA) has drawn growing attention in recent years, which only utilizes the source-trained model and unlabeled target data to adapt to the target domain. Despite the rapid explosion of SFDA work, there has been no timely and comprehensive survey in the field. To fill this gap, we provide a comprehensive survey of recent advances in SFDA and organize them into a unified categorization scheme based on the framework of transfer learning. Instead of presenting each approach independently, we modularize several components of each method to more clearly illustrate their relationships and mechanisms in light of the composite properties of each method. Furthermore, we compare the results of more than 30 representative SFDA methods on three popular classification benchmarks, namely Office-31, Office-home, and VisDA, to explore the effectiveness of various technical routes and the combination effects among them. Additionally, we briefly introduce the applications of SFDA and related fields. Drawing on our analysis of the challenges confronting SFDA, we offer some insights into future research directions and potential settings.
Jingjing Li 0001, Zhiqi Yu, Zhekai Du, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Domain-Adaptive Energy-Based Models for Generalizable Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) plays a crucial role in securing face recognition systems against presentation attacks. However, existing FAS methods often struggle to generalize to unseen attacks and domains. Existing generalizable FAS studies generally leverage domain generalization (DG) techniques for exploiting intermediate features that support generalization while neglecting the task-specific nature of FAS. In this paper, we argue that the FAS task is an imbalanced classification problem, which renders it unsuitable to be handled by a standard discriminative classifier. In contrast, we propose a novel approach for FAS by modeling the problem from a generative perspective using an energy-based model (EBM). The EBM captures the distribution of genuine faces and detects spoofing attempts as deviations from this distribution. We train the EBM using a discriminative objective and an energy regularization term to shape the learned distribution and improve generalization. To enhance the robustness to unseen domains, we introduce an energy-based domain augmentation technique that explores the latent space around the source distribution guided by the EBM. We further leverage a meta-learning framework and a gradient-based variant to leverage the augmented data for domain generalization. For practicability, we consider a practical setting where samples are holistically collected under different environments without distinct domain labels, and show that our method can naturally harness this challenging setting by training with cluster labels. Extensive experiments on four FAS datasets demonstrate the superiority of our method in both intra- and cross-dataset settings, outperforming state-of-the-art approaches.
Zhekai Du, Jingjing Li 0001, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Multim.2
2024 Cross-domain Recommendation via Dual Adversarial Adaptation
abstract
Data scarcity is a perpetual challenge of recommendation systems, and researchers have proposed a variety of cross-domain recommendation methods to alleviate the problem of data scarcity in target domains. However, in many real-world cross-domain recommendation systems, the source domain and the target domain are sampled from different data distributions, which obstructs the cross-domain knowledge transfer. In this article, we propose to specifically align the data distributions between the source domain and the target domain to alleviate imbalanced sample distribution and thus challenge the data scarcity issue in the target domain. Technically, our proposed approach builds a dual adversarial adaptation (DAA) framework to adversarially train the target model together with a pre-trained source model. Two domain discriminators play the two-player minmax game with the target model and guide the target model to learn reliable domain-invariant features that can be transferred across domains. At the same time, the target model is calibrated to learn domain-specific information of the target domain. In addition, we formulate our approach as a plug-and-play module to boost existing recommendation systems. We apply the proposed method to address the issues of insufficient data and imbalanced sample distribution in real-world Click-through Rate/Conversion Rate predictions on two large-scale industrial datasets. We evaluate the proposed method in scenarios with and without overlapping users/items, and extensive experiments verify that the proposed method is able to significantly improve the prediction performance on the target domain. For instance, our method can boost PLE with a performance improvement of 15.4% in terms of Area Under Curve compared with single-domain PLE on our private game dataset. In addition, our method is able to surpass single-domain MMoE by 6.85% on the public datasets. Code: https://github.com/TL-UESTC/DAA .
Hongzu Su, Jingjing Li 0001, Zhekai Du, Lei Zhu 0002, Ke Lu 0001, Heng Tao Shen
ACM Trans. Inf. Syst.3
2023 Cross-Domain Adaptative Learning for Online Advertisement Customer Lifetime Value Prediction
abstract
Accurate estimation of customer lifetime value (LTV), which reflects the potential consumption of a user over a period of time, is crucial for the revenue management of online advertising platforms. However, predicting LTV in real-world applications is not an easy task since the user consumption data is usually insufficient within a specific domain. To tackle this problem, we propose a novel cross-domain adaptative framework (CDAF) to leverage consumption data from different domains. The proposed method is able to simultaneously mitigate the data scarce problem and the distribution gap problem caused by data from different domains. To be specific, our method firstly learns a LTV prediction model from a different but related platform with sufficient data provision. Subsequently, we exploit domain-invariant information to mitigate data scarce problem by minimizing the Wasserstein discrepancy between the encoded user representations of two domains. In addition, we design a dual-predictor schema which not only enhances domain-invariant information in the semantic space but also preserves domain-specific information for accurate target prediction. The proposed framework is evaluated on five datasets collected from real historical data on the advertising platform of Tencent Games. Experimental results verify that the proposed framework is able to significantly improve the LTV prediction performance on this platform. For instance, our method can boost DCNv2 with the improvement of 13.7% in terms of AUC on dataset G2. Code: https://github.com/TL-UESTC/CDAF.
Hongzu Su, Zhekai Du, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001
AAAI2
2023 Exploring Low-Dimensional Manifolds of Deep Neural Network Parameters for Improved Model Optimization
abstract
Manifold learning techniques have significantly enhanced the comprehension of massive data by exploring the geometric properties of the data manifold in low-dimensional subspaces. However, existing research on manifold learning primarily focuses on understanding the intricate data, overlooking the explosive growth of the scale and complexity of deep neural networks (DNNs), which presents a significant challenge for model optimization. In this work, we propose to explore the intrinsic low-dimensional manifold of network parameters for efficient model optimization. Specifically, we analyze parameter distributions in a deep model and perform sampling to map them onto a low-dimensional parameter manifold using the local tangent space alignment (LTSA). Since our focus is on studying parameter manifolds to guide model optimization, we therefore select dynamic optimal training trajectories for sampling and approximate tangent spaces to obtain low-dimensional representations of DNNs. By applying manifold learning techniques and employing a two-step alternate optimization method, we achieve a fixed subspace that reduces training time and resource costs for commonly used deep networks. The trained low-dimensional network can be mapped back to the original parameter space for further use. We demonstrate the benefits of learning low-dimensional parameterization of DNNs on both noisy label learning and federated learning tasks. Extensive experimental results on various benchmarks show the effectiveness of our method concerning both superior accuracy and reduced resource consumption.
Ke Lu 0001, Xiaotong He, Ze Qin, Zhekai Du
CIKM5
2023 Noise-Robust Continual Test-Time Domain Adaptation
abstract
Continual test-time domain adaptation (TTA) is a challenging topic in the field of source-free domain adaptation, which focuses on addressing cross-domain multimedia information during inference with a continuously changing data distribution. Previous methods have been found to lack noise robustness, leading to a significant increase in errors under strong noise. In this paper, we address the noise-robustness problem in continual TTA by offering three effective recipes to mitigate it. At the category level, we employ the Taylor cross-entropy loss to alleviate the low confidence category bias commonly associated with cross-entropy. At the sample level, we reweight the target samples based on uncertainty to prevent the model from overfitting on noisy samples. Finally, to reduce pseudo-label noise, we propose a soft ensemble negative learning mechanism to guide the model optimization using ensemble complementary pseudo labels. Our method achieves state-of-the-art performance on three widely used continual TTA datasets, particularly in the strong noise setting that we introduced.
Zhiqi Yu, Jingjing Li 0001, Zhekai Du, Fengling Li 0001, Lei Zhu 0002, Yang Yang 0002
ACM Multimedia3
2023 Diffusion-Based Probabilistic Uncertainty Estimation for Active Domain Adaptation
abstract
Active Domain Adaptation (ADA) has emerged as an attractive technique for assisting domain adaptation by actively annotating a small subset of target samples. Most ADA methods focus on measuring the target representativeness beyond traditional active learning criteria to handle the domain shift problem, while leaving the uncertainty estimation to be performed by an uncalibrated deterministic model. In this work, we introduce a probabilistic framework that captures both data-level and prediction-level uncertainties beyond a point estimate. Specifically, we use variational inference to approximate the joint posterior distribution of latent representation and model prediction. The variational objective of labeled data can be formulated by a variational autoencoder and a latent diffusion classifier, and the objective of unlabeled data can be implemented in a knowledge distillation framework. We utilize adversarial learning to ensure an invariant latent space. The resulting diffusion classifier enables efficient sampling of all possible predictions for each individual to recover the predictive distribution. We then leverage a t-test-based criterion upon the sampling and select informative unlabeled target samples based on the p-value, which encodes both prediction variability and cross-category ambiguity. Experiments on both ADA and Source-Free ADA settings show that our method provides more calibrated predictions than previous ADA methods and achieves favorable performance on three domain adaptation datasets.
Zhekai Du, Jingjing Li 0001
NeurIPS1
2023 Alleviating the generalization issue in adversarial domain adaptation networks
Zhekai Du, Chunwei Lou, Jingjing Li 0001
Image Vis. Comput.2
2023 Online continual learning with declarative memory
Zhekai Du, Ruijin Wang, Ruimeng Gan, Jingjing Li 0001
Neural Networks2
2022 Interpretable Open-Set Domain Adaptation via Angular Margin Separation
Xinhao Li 0002, Jingjing Li 0001, Zhekai Du, Lei Zhu 0002, Wen Li 0001
ECCV (34)3
2022 Towards Distributed Communication and Control in Real-World Multi-Agent Reinforcement Learning
abstract
Multi-agent system investigates the problem of designing a complex system composed of multiple autonomous agents with limited ability and partial observability. As a milestone, AlphaStar has achieved remarkable success in StarCraft II, which is a significant breakthrough in the competitive environments with complex strategic spaces and real-time decisions. However, it poses new challenges for deploying these centralized control models in real-world environments because many of them in such competitive environments were not designed to accommodate the requirements of real-world communication networks, e.g., the problems of high latency and large traffic are inevitable when they are actually deployed. To alleviate this issue, we propose a distributed control paradigm that explicitly splits the control power between the centralized meta-agent and agent units through a combination of centralized and decentralized paradigms. The units can autonomously decide to follow the decisions of the meta-agent or adapt to environment variations immediately by themselves in a decentralized manner. We simulate real-world network environments based on the Mininet platform, experiments based on the StarCraft II Learning Environment (SC2LE) show that our approach achieves a better adaptation in real-world network environments.
Jieyan Liu, Zhekai Du, Ke Lu 0001
ICC3
2022 Energy-Based Domain Generalization for Face Anti-Spoofing
abstract
With various unforeseeable face presentation attacks (PA) springing up, face anti-spoofing (FAS) urgently needs to generalize to unseen scenarios. Research on generalizable FAS has lately attracted growing attention. Existing methods cast FAS as a vanilla binary classification problem and address it by a standard discriminative classifier p(y|x) under a domain generalization framework. However, discriminative models are unreliable for samples far away from the training distribution. In this paper, we resort to an energy-based model (EBM) to tackle FAS in a generative perspective. Our motivation is to model the joint density p(x,y), which allows to compute not only p(y|x) but also p(x). Due to the intractability of direct modeling, we use EBMs as an alternative to probabilistic estimation. With energy-based training, real faces are encouraged to get low free energy associated with the marginal probability p(x) of real faces, and all samples with high free energy are regarded as fake faces, thus rejecting any kind of PA out of the distribution of real faces. To learn to generalize to unseen domains, we generate diverse and novel populations in feature space under the guidance of energy model. Our model is updated in a meta-learning schema, where the original source samples are utilized for meta-training and the generated ones for meta-testing. We validate our method on four widely used FAS datasets. Comprehensive experimental results demonstrate the effectiveness of our method compared with state-of-the-arts.
Zhekai Du, Jingjing Li 0001, Lin Zuo, Lei Zhu 0002, Ke Lu 0001
ACM Multimedia1
2022 Source-Free Active Domain Adaptation via Energy-Based Locality Preserving Transfer
abstract
Unsupervised domain adaptation (UDA) aims at transferring knowledge from one labeled source domain to a related but unlabeled target domain. Recently, active domain adaptation (ADA) has been proposed as a new paradigm which significantly boosts performance of UDA with minor additional labeling. However, existing ADA methods require source data to explicitly measure the domain gap between the source domain and the target domain, which is restricted in many real-world scenarios. In this work, we handle ADA with only a source-pretrained model and unlabeled target data, proposing a new setting named source-free active domain adaptation. Specifically, we propose a Locality Preserving Transfer (LPT) framework which preserves and utilizes locality structures on target data to achieve adaptation without source data. Meanwhile, a label propagation strategy is adopted to improve the discriminability for better adaptation. After LPT, unique samples with insignificant locality structure are identified by an energy-based approach for active annotation. An energy-based pseudo labeling strategy is further applied to generate labels for reliable samples. Finally, with supervision from the annotated samples and pseudo labels, a well adapted model is obtained. Extensive experiments on three widely used UDA benchmarks show that our method is comparable or superior to current state-of-the-art active domain adaptation methods even without access to source data.
Zhekai Du, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001
ACM Multimedia2
2022 Divergence-Agnostic Unsupervised Domain Adaptation by Adversarial Attacks
abstract
Conventional machine learning algorithms suffer the problem that the model trained on existing data fails to generalize well to the data sampled from other distributions. To tackle this issue, unsupervised domain adaptation (UDA) transfers the knowledge learned from a well-labeled source domain to a different but related target domain where labeled data is unavailable. The majority of existing UDA methods assume that data from the source domain and the target domain are available and complete during training. Thus, the divergence between the two domains can be formulated and minimized. In this paper, we consider a more practical yet challenging UDA setting where either the source domain data or the target domain data are unknown. Conventional UDA methods would fail this setting since the domain divergence is agnostic due to the absence of the source data or the target data. Technically, we investigate UDA from a novel view-adversarial attack-and tackle the divergence-agnostic adaptive learning problem in a unified framework. Specifically, we first report the motivation of our approach by investigating the inherent relationship between UDA and adversarial attacks. Then we elaborately design adversarial examples to attack the training model and harness these adversarial examples. We argue that the generalization ability of the model would be significantly improved if it can defend against our attack, so as to improve the performance on the target domain. Theoretically, we analyze the generalization bound for our method based on domain adaptation theories. Extensive experimental results on multiple UDA benchmarks under conventional, source-absent and target-absent UDA settings verify that our method is able to achieve a favorable performance compared with previous ones. Notably, this work extends the scope of both domain adaptation and adversarial attack, and expected to inspire more ideas in the community.
Jingjing Li 0001, Zhekai Du, Lei Zhu 0002, Zhengming Ding, Ke Lu 0001, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Resource-Efficient Distributed Deep Neural Networks Empowered by Intelligent Software-Defined Networking
abstract
Contemporary machine learning methods have evolved from conventional algorithms to deep neural networks (DNNs) that are computation- and data- intensive. Thus, they are suitable to be deployed in the cloud that can offer high computational capacity and scalable resources. However, the cloud computing paradigm is not optimal for delay- and energy-sensitive applications. To mitigate these problems, a battery of distributed DNNs have been proposed to allow a fast inference with device-edge-cloud synergy. Furthermore, although distributed deployment of DNNs on real communication networks is an important research topic, the legacy network architecture cannot meet the requirements of these distributed deep neural networks due to the complicated management and manual configuration, etc. To cope with these requirements, we develop a novel and explicit Intelligent Software Defined Networking (ISDN) that aims to manage the bandwidth and computing resources across the network via the SDN paradigm. We first identify the difficulties of deploying distributed intelligent computing in the current network architecture. Then, we explain how to address these problems by introducing the ISDN architecture. Specifically, we develop a dynamic routing method to enable Quality-of-Service (QoS) communication based on the SDN paradigm and propose a Markov Decision Process (MDP) based dynamic task offloading model to achieve the optimal offloading policy of DNN tasks. We develop a simulation platform based on Mininet to measure its performance advantages over traditional architectures. Extensive experimental results show that compared with the traditional network architecture, our architecture based on the SDN paradigm can perform better in terms of both network throughput and resource utilization.
Ke Lu 0001, Zhekai Du, Jingjing Li 0001, Geyong Min
IEEE Trans. Netw. Serv. Manag.2
2021 Cross-Domain Gradient Discrepancy Minimization for Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to generalize the knowledge learned from a well-labeled source domain to an unlabled target domain. Recently, adversarial domain adaptation with two distinct classifiers (biclassifier) has been introduced into UDA which is effective to align distributions between different domains. Previous bi-classifier adversarial learning methods only focus on the similarity between the outputs of two distinct classifiers. However, the similarity of the outputs cannot guarantee the accuracy of target samples, i.e., traget samples may match to wrong categories even if the discrepancy between two classifiers is small. To challenge this issue, in this paper, we propose a cross-domain gradient discrepancy minimization (CGDM) method which explicitly minimizes the discrepancy of gradients generated by source samples and target samples. Specifically, the gradient gives a cue for the semantic information of target samples so it can be used as a good supervision to improve the accuracy of target samples. In order to compute the gradient signal of target smaples, we further obtain target pseudo labels through a clustering-based self-supervised learning. Extensive experiments on three widely used UDA datasets show that our method surpasses many previous state-of-the-arts.
Zhekai Du, Jingjing Li 0001, Hongzu Su, Lei Zhu 0002, Ke Lu 0001
CVPR1
2021 Learning Transferrable and Interpretable Representations for Domain Generalization
abstract
Conventional machine learning models are often vulnerable to samples with different distributions from the ones of training samples, which is known as domain shift. Domain Generalization (DG) challenges this issue by training a model based on multiple source domains and generalizing it to arbitrary unseen target domains. In spite of remarkable results made in DG, a majority of existing works lack a deep understanding of the feature representations learned in DG models, resulting in limited generalization ability when facing domainsout-of-distribution. In this paper, we aim to learn a domain transformation space via a domain transformer network (DTN) which explicitly mines the relationship among multiple domains and constructs transferable feature representations for down-stream tasks by interpreting each feature as a semantically weighted combination of multiple domain-specific features. Our DTN is encouraged to meta-learn the properties and characteristics of domains during the training process based on multiple seen domains, making transformed feature representations more semantical, thus generalizing better to unseen domains. Once the model is constructed, the feature representations of unseen target domains can also be inferred adaptively by selectively combining the feature representations from the diverse set of seen domains. We conduct extensive experiments on five DG benchmarks and the results strongly demonstrate the effectiveness of our approach.
Zhekai Du, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Zi Huang
ACM Multimedia1
2021 CMRD-Net: An Improved Method for Underwater Image Enhancement
abstract
Underwater image enhancement is a challenging task due to the degradation of image quality in underwater complicated lighting conditions and scenes. In recent years, most methods improve the visual quality of underwater images by using deep Convolutional Neural Networks and Generative Adversarial Networks. However, the majority of existing methods do not consider that the attenuation degrees of R, G, B channels of the underwater image are different, leading to a sub-optimal performance. Based on this observation, we propose a Channel-wise Multi-scale Residual Dense Network called CMRD-Net, which learns the weights of different color channels instead of treating all the channels equally. More specifically, the Channel-wise Multi-scale Fusion Residual Attention Block (CMFRAB) is involved in the CMRD-Net to obtain a better ability of feature extraction and representation. Notably, we evaluate the effectiveness of our model by comparing it with recent state-of-the-art methods. Extensive experimental results show that our method can achieve a satisfactory performance on a popular public dataset.
Fengjie Xu, Changhua Zhang, Zhongshu Chen, Zhekai Du, Lin Zuo
MMAsia4
2021 Local-Global Attentive Adaptation for Object Detection
Jingjing Li 0001, Xingpeng Li, Zhekai Du, Mao Ye 0001
Eng. Appl. Artif. Intell.4
2020 Joint metric and feature representation learning for unsupervised domain adaptation
Zhekai Du, Jingjing Li 0001, Mengmeng Jing, Erpeng Chen, Ke Lu 0001
Knowl. Based Syst.2