EDBT 2026 Demo / reviewers in the wild / expert
Ke Lu 0001
dblp:33/1254-1
· DBLP profile ↗
74ranked-venue papers
10as first author
38since 2021 · last 2026
0000-0002-3456-4993ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 18 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 8 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Open-Set Domain Adaptation With Vision-Language ModelsabstractWith the prevalence of pre-trained vision-language models like CLIP, leveraging the generic knowledge embedded in CLIP for domain adaptation has proved to be a promising direction. However, most existing CLIP-based methods are limited to closed-set settings. This is primarily because CLIP needs the semantic labels of unknown classes for inference, thus making it not applicable to Open-Set Domain Adaptation (OSDA). To utilize the complementary roles of CLIP and the source model, our paper proposes a novel Semantic-guided Target Adaptation (SemTA) framework for OSDA in a training-free manner. Specifically, we introduce an unknown semantic discovery module. It uses the cluster centroids of the target data to obtain the semantic labels of unknown classes from the worldwide corpus. Then, the semantic-based inference can be performed with CLIP. Additionally, the dual sample attention mechanism is implemented to output sample-based inference. Representative features from both the source model and CLIP serve as the key to improve task specificity. Compared to previous OSDA methods which reject unknown data by confidence threshold, the proposed approach is more practical and offers better interpretability. Comprehensive evaluations on four benchmarks reveal our method sets a new state-of-the-art even without training. Our code will be publicly available soon. Zhiqi Yu, Ke Lu 0001, Kangkai Wu, Fengling Li 0001, Jingjing Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | UGLP: Unifying Global and Local Preferences for Multi-behavior RecommendationabstractMulti-behavior recommender systems have demonstrated their effectiveness in mitigating issues such as data sparsity by incorporating auxiliary behaviors into the target behavior. However, existing multi-behavior recommendation approaches typically take one of two directions: (1) fusing behavior-specific preference features from various behavior interaction graphs explicitly or implicitly for recommendation; or (2) utilizing behavior-unified preference features from the unified interaction graph for recommendation or to initialize features for subsequent modeling. These methods fail to exploit the integration of behavior-unified global and behavior-specific local preference features, resulting in incomplete preference modeling. To address this issue, in this work, we propose a novel method calledUnifyingGlobal andLocalPreferences (UGLP) for multi-behavior recommendation. In UGLP, we design a behavior feature fusion network that consists of global and local fusion modules for comprehensive and fine-grained user preferences. The global fusion module performs graph convolution on behavior-unified global and behavior-specific local interaction graphs to obtain behavior-unified and behavior-specific features. The behavior-unified and behavior-specific features are then fused into globally fused features via a gating network. The local fusion module then performs cross-behavior fusion on these globally fused features via another gating network. We introduce a contrastive learning module to promote preference alignment and knowledge transfer from auxiliary behaviors to the target behavior. Additionally, we incorporate a GCN refinement module to adjust the fused features to ensure that both global and local user preferences are learned. Experimental results on three real-world datasets verify that our method is able to surpass various state-of-the-art models. For instance, our method outperforms the best baseline by an average of 16.43% and 14.36% in terms of HR@10 and NDCG@10, respectively. Zhichao Liao, Ke Lu 0001, Jingxi Xie, Jingjing Li 0001, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | LoCA: Location-Aware Cosine Adaptation for Parameter-Efficient Fine-TuningabstractLow-rank adaptation (LoRA) has become a prevalent method for adapting pre-trained large language models to downstream tasks. However, the simple low-rank decomposition form may constrain the optimization flexibility. To address this limitation, we introduce Location-aware Cosine Adaptation (LoCA), a novel frequency-domain parameter-efficient fine-tuning method based on inverse Discrete Cosine Transform (iDCT) with selective locations of learnable components. We begin with a comprehensive theoretical comparison between frequency-domain and low-rank decompositions for fine-tuning pre-trained large models. Our analysis reveals that frequency-domain decomposition with carefully selected frequency components can surpass the expressivity of traditional low-rank-based methods. Furthermore, we demonstrate that iDCT offers a more efficient implementation compared to inverse Discrete Fourier Transform (iDFT), allowing for better selection and tuning of frequency components while maintaining equivalent expressivity to the optimal iDFT-based adaptation. By employing finite-difference approximation to estimate gradients for discrete locations of learnable coefficients on the DCT spectrum, LoCA dynamically selects the most informative frequency components during training. Experiments on diverse language and vision fine-tuning tasks demonstrate that LoCA offers enhanced parameter efficiency while maintains computational feasibility comparable to low-rank-based methods. Zhekai Du, Yinjie Min, Jingjing Li 0001, Ke Lu 0001, Changliang Zou, Liuhua Peng, Tingjin Chu, Mingming Gong |
ICLR | 4 |
| 2025 | Online Adaptive Fault Diagnosis With Test-Time Domain AdaptationabstractCross-domain bearing fault diagnosis algorithms have garnered considerable attention in recent years due to their robust ability to address domain bias. However, prevailing methods often grapple with two key challenges: the absence of privacy preservation (necessitating access to source domain data) and the inability to facilitate real-time predictions (requiring iterative training on complete target domain data). In response to these issues, this article introduces an algorithm designed to adapt a pretrained model to the target domain in an online fashion. Notably, data augmentation is employed for pretraining the source domain model, enhancing the generalization capabilities. Subsequently, self-supervised learning is integrated through weight average updating. Furthermore, a memory bank-based approach is introduced to augment the compactness of features within the same class. Evaluation on several public datasets demonstrates that our model not only effectively enhances the diagnostic accuracy of the source model, but also achieves state-of-the-art results compared to other test-time adaptation methods. Kangkai Wu, Jingjing Li 0001, Lichao Meng, Fengling Li 0001, Ke Lu 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Agile Multi-Source-Free Domain AdaptationabstractEfficiently utilizing rich knowledge in pretrained models has become a critical topic in the era of large models. This work focuses on adaptively utilize knowledge from multiple source-pretrained models to an unlabeled target domain without accessing the source data. Despite being a practically useful setting, existing methods require extensive parameter tuning over each source model, which is computationally expensive when facing abundant source domains or larger source models. To address this challenge, we propose a novel approach which is free of the parameter tuning over source backbones. Our technical contribution lies in the Bi-level ATtention ENsemble (Bi-ATEN) module, which learns both intra-domain weights and inter-domain ensemble weights to achieve a fine balance between instance specificity and domain consistency. By slightly tuning source bottlenecks, we achieve comparable or even superior performance on a challenging benchmark DomainNet with less than 3% trained parameters and 8 times of throughput compared with SOTA method. Furthermore, with minor modifications, the proposed module can be easily equipped to existing methods and gain more than 4% performance boost. Code is available at https://github.com/TL-UESTC/Bi-ATEN. Jingjing Li 0001, Fengling Li 0001, Lei Zhu 0002, Ke Lu 0001 |
AAAI | 5 |
| 2024 | Domain-Agnostic Mutual Prompting for Unsupervised Domain AdaptationabstractConventional Unsupervised Domain Adaptation (UDA) strives to minimize distribution discrepancy between do-mains, which neglects to harness rich semantics from data and struggles to handle complex domain shifts. A promising technique is to leverage the knowledge of large-scale pretrained vision-language models for more guided adaptation. Despite some endeavors, current methods often learn textual prompts to embed domain semantics for source and target domains separately and perform classification within each domain, limiting cross-domain knowledge transfer. Moreover, prompting only the language branch lacks flex-ibility to adapt both modalities dynamically. To bridge this gap, we propose Domain-Agnostic Mutual Prompting (DAMP) to exploit domain-invariant semantics by mutually aligning visual and textual embeddings. Specifically, the image contextual information is utilized to prompt the language branch in a domain-agnostic and instance-conditioned way. Meanwhile, visual prompts are im-posed based on the domain-agnostic textual prompt to elicit domain-invariant visual embeddings. These two branches of prompts are learned mutually with a cross-attention module and regularized with a semantic-consistency loss and an instance-discrimination contrastive loss. Experiments on three UDA benchmarks demonstrate the superiority of DAMP over state-of-the-art approaches1. Zhekai Du, Fengling Li 0001, Ke Lu 0001, Lei Zhu 0002, Jingjing Li 0001 |
CVPR | 4 |
| 2024 | Split to Merge: Unifying Separated Modalities for Unsupervised Domain AdaptationabstractLarge vision-language models (VLMs) like CLIP have demonstrated good zero-shot learning performance in the unsupervised domain adaptation task. Yet, most transfer approaches for VLMs focus on either the language or visual branches, overlooking the nuanced interplay between both modalities. In this work, we introduce a Unified Modality Separation (UniMoS) framework for unsupervised domain adaptation. Leveraging insights from modality gap studies, we craft a nimble modality separation network that distinctly disentangles CLIP's features into language-associated and vision-associated components. Our proposed Modality-Ensemble Training (MET) method fosters the exchange of modality-agnostic information while maintaining modality-specific nuances. We align features across domains using a modality discriminator. Comprehensive evaluations on three benchmarks reveal our approach sets a new state-of-the-art with minimal computational costs. Code: https://github.com/TL-UESTC/UniMoS. Zhekai Du, Fengling Li 0001, Ke Lu 0001, Jingjing Li 0001 |
CVPR | 5 |
| 2024 | SOIL: Contrastive Second-Order Interest Learning for Multimodal RecommendationabstractMainstream multimodal recommender systems are designed to learn user interest by analyzing user-item interaction graphs. However, what they learn about user interest needs to be completed because historical interactions only record items that best match user interest (i.e., the first-order interest), while suboptimal items are absent. To fully exploit user interest, we propose a Second-Order Interest Learning (SOIL) framework to retrieve second-order interest from unrecorded suboptimal items. In this framework, we build a user-item interaction graph augmented by second-order interest, an interest-aware item-item graph for the visual modality, and a similar graph for the textual modality. In our work, all three graphs are constructed from user-item interaction records and multimodal feature similarity. Similarly to other graph-based approaches, we apply graph convolutional networks to each of the three graphs to learn representations of users and items. To improve the exploitation of both first-order and second-order interest, we optimize the model by implementing contrastive learning modules for user and item representations at both the user-item and item-item levels. The proposed framework is evaluated on three real-world public datasets in online shopping scenarios. Experimental results verify that our method is able to significantly improve prediction performance. For instance, our method outperforms the previous state-of-the-art method MGCN by an average of 8.1% in terms of Recall@10. Code: https://github.com/TL-UESTC/SOIL. Hongzu Su, Jingjing Li 0001, Fengling Li 0001, Ke Lu 0001, Lei Zhu 0002 |
ACM Multimedia | 4 |
| 2024 | DDPO: Direct Dual Propensity Optimization for Post-Click Conversion Rate EstimationabstractIn online advertising, the sample selection bias problem is a major cause of inaccurate conversion rate estimates. Current mainstream solutions only perform causality-based optimization in the click space since the conversion labels in the non-click space are absent. However, optimization for unclicked samples is equally essential because the non-click space contains more samples and user characteristics than the click space. To exploit the unclicked samples, we propose a Direct Dual Propensity Optimization (DDPO) framework to optimize the model directly in impression space with both clicked and unclicked samples. In this framework, we specifically design a click propensity network and a conversion propensity network. The click propensity network is dedicated to ensuring that optimization in the click space is unbiased. The conversion propensity network is designed to generate pseudo-conversion labels for unclicked samples, thus overcoming the challenge of absent labels in non-click space. With these two propensity networks, we are able to perform causality-based optimization in both click space and non-click space. In addition, to strengthen the causal relationship, we design two causal transfer modules for the conversion rate prediction model with the attention mechanism. The proposed framework is evaluated on five real-world public datasets and one private Tencent advertising dataset. Experimental results verify that our method is able to improve the prediction performance significantly. For instance, our method outperforms the previous state-of-the-art method by 7.0% in terms of the Area Under the Curve on the Ali-CCP dataset. Hongzu Su, Lichao Meng, Lei Zhu 0002, Ke Lu 0001, Jingjing Li 0001 |
SIGIR | 4 |
| 2024 | Visually Source-Free Domain Adaptation via Adversarial Style MatchingabstractThe majority of existing works explore Unsupervised Domain Adaptation (UDA) with an ideal assumption that samples in both domains are available and complete. In real-world applications, however, this assumption does not always hold. For instance, data-privacy is becoming a growing concern, the source domain samples may be not publicly available for training, leading to a typical Source-Free Domain Adaptation (SFDA) problem. Traditional UDA methods would fail to handle SFDA since there are two challenges in the way: the data incompleteness issue and the domain gaps issue. In this paper, we propose a visually SFDA method named Adversarial Style Matching (ASM) to address both issues. Specifically, we first train a style generator to generate source-style samples given the target images to solve the data incompleteness issue. We use the auxiliary information stored in the pre-trained source model to ensure that the generated samples are statistically aligned with the source samples, and use the pseudo labels to keep semantic consistency. Then, we feed the target domain samples and the corresponding source-style samples into a feature generator network to reduce the domain gaps with a self-supervised loss. An adversarial scheme is employed to further expand the distributional coverage of the generated source-style samples. The experimental results verify that our method can achieve comparative performance even compared with the traditional UDA methods with source samples for training. Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 3 |
| 2024 | Cross-domain Recommendation via Dual Adversarial AdaptationabstractData scarcity is a perpetual challenge of recommendation systems, and researchers have proposed a variety of cross-domain recommendation methods to alleviate the problem of data scarcity in target domains. However, in many real-world cross-domain recommendation systems, the source domain and the target domain are sampled from different data distributions, which obstructs the cross-domain knowledge transfer. In this article, we propose to specifically align the data distributions between the source domain and the target domain to alleviate imbalanced sample distribution and thus challenge the data scarcity issue in the target domain. Technically, our proposed approach builds a dual adversarial adaptation (DAA) framework to adversarially train the target model together with a pre-trained source model. Two domain discriminators play the two-player minmax game with the target model and guide the target model to learn reliable domain-invariant features that can be transferred across domains. At the same time, the target model is calibrated to learn domain-specific information of the target domain. In addition, we formulate our approach as a plug-and-play module to boost existing recommendation systems. We apply the proposed method to address the issues of insufficient data and imbalanced sample distribution in real-world Click-through Rate/Conversion Rate predictions on two large-scale industrial datasets. We evaluate the proposed method in scenarios with and without overlapping users/items, and extensive experiments verify that the proposed method is able to significantly improve the prediction performance on the target domain. For instance, our method can boost PLE with a performance improvement of 15.4% in terms of Area Under Curve compared with single-domain PLE on our private game dataset. In addition, our method is able to surpass single-domain MMoE by 6.85% on the public datasets. Code: https://github.com/TL-UESTC/DAA . Hongzu Su, Jingjing Li 0001, Zhekai Du, Lei Zhu 0002, Ke Lu 0001, Heng Tao Shen |
ACM Trans. Inf. Syst. | 5 |
| 2023 | Cross-Domain Adaptative Learning for Online Advertisement Customer Lifetime Value PredictionabstractAccurate estimation of customer lifetime value (LTV), which reflects the potential consumption of a user over a period of time, is crucial for the revenue management of online advertising platforms. However, predicting LTV in real-world applications is not an easy task since the user consumption data is usually insufficient within a specific domain. To tackle this problem, we propose a novel cross-domain adaptative framework (CDAF) to leverage consumption data from different domains. The proposed method is able to simultaneously mitigate the data scarce problem and the distribution gap problem caused by data from different domains. To be specific, our method firstly learns a LTV prediction model from a different but related platform with sufficient data provision. Subsequently, we exploit domain-invariant information to mitigate data scarce problem by minimizing the Wasserstein discrepancy between the encoded user representations of two domains. In addition, we design a dual-predictor schema which not only enhances domain-invariant information in the semantic space but also preserves domain-specific information for accurate target prediction. The proposed framework is evaluated on five datasets collected from real historical data on the advertising platform of Tencent Games. Experimental results verify that the proposed framework is able to significantly improve the LTV prediction performance on this platform. For instance, our method can boost DCNv2 with the improvement of 13.7% in terms of AUC on dataset G2. Code: https://github.com/TL-UESTC/CDAF. Hongzu Su, Zhekai Du, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001 |
AAAI | 5 |
| 2023 | Exploring Low-Dimensional Manifolds of Deep Neural Network Parameters for Improved Model OptimizationabstractManifold learning techniques have significantly enhanced the comprehension of massive data by exploring the geometric properties of the data manifold in low-dimensional subspaces. However, existing research on manifold learning primarily focuses on understanding the intricate data, overlooking the explosive growth of the scale and complexity of deep neural networks (DNNs), which presents a significant challenge for model optimization. In this work, we propose to explore the intrinsic low-dimensional manifold of network parameters for efficient model optimization. Specifically, we analyze parameter distributions in a deep model and perform sampling to map them onto a low-dimensional parameter manifold using the local tangent space alignment (LTSA). Since our focus is on studying parameter manifolds to guide model optimization, we therefore select dynamic optimal training trajectories for sampling and approximate tangent spaces to obtain low-dimensional representations of DNNs. By applying manifold learning techniques and employing a two-step alternate optimization method, we achieve a fixed subspace that reduces training time and resource costs for commonly used deep networks. The trained low-dimensional network can be mapped back to the original parameter space for further use. We demonstrate the benefits of learning low-dimensional parameterization of DNNs on both noisy label learning and federated learning tasks. Extensive experimental results on various benchmarks show the effectiveness of our method concerning both superior accuracy and reduced resource consumption. Ke Lu 0001, Xiaotong He, Ze Qin, Zhekai Du |
CIKM | 1 |
| 2023 | Task-Adversarial Adaptation for Multi-modal RecommendationabstractAn ideal multi-modal recommendation system is supposed to be timely updated with the latest modality information and interaction data because the distribution discrepancy between new data and historical data will lead to severe recommendation performance deterioration. However, upgrading a recommendation system with numerous new data consumes much time and computing resources. To mitigate this problem, we propose a Task-Adversarial Adaptation (TAA) framework, which is able to align data distributions and reduce resource consumption at the same time. This framework is specifically designed to align distributions of embedded features for different recommendation tasks between the source domain (i.e., historical data) and the target domain (i.e., new data). Technically, we design a domain feature discriminator for each task to distinguish which domain a feature comes from. By the two-player min-max game between the feature discriminator and the feature embedding network, the feature embedding network is able to align the source and target data distributions. With the ability to align source and target distributions, we are able to reduce the number of training samples by random sampling. In addition, we formulate the proposed approach as a plug-and-play module to accelerate the model training and improve the performance of mainstream multi-modal multi-task recommendation systems. We evaluate our method by predicting the Click-Through Rate (CTR) in e-commerce scenarios. Extensive experiments verify that our method is able to significantly improve prediction performance and accelerate model training on the target domain. For instance, our method is able to surpass the previous state-of-the-art method by 2.45% in terms of Area Under Curve (AUC) on AliExpress_US dataset while only utilizing one percent of the target data in training. Code: https://github.com/TL-UESTC/TAA. Hongzu Su, Jingjing Li 0001, Fengling Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002 |
ACM Multimedia | 5 |
| 2023 | Continuous-time graph directed information maximization for temporal network representation
Chenming Yang, Jingjing Li 0001, Ke Lu 0001, Bryan Hooi, Liang Zhou 0003 |
Inf. Sci. | 3 |
| 2023 | Dual-Aligned Feature Confusion Alleviation for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) aims to recognize both seen and unseen samples by leveraging the connections between semantic and visual representations. Recently, a majority of GZSL methods focus on generating visual features for unseen categories conditioned on category-level semantic attributes. However, there is a considerable gap between generated features and real unseen features since the generator is trained with only seen samples. The final classifier may get confused by the unfaithful generated features and make misclassification. To alleviate this issue, we propose a dual-aligned feature confusion alleviation (DFCA) framework that simultaneously generates faithful and discriminative features for unseen categories. Specifically, our DFCA attains the faithfulness via a conditional invertible neural network (cINN) and aligns the generated visual features and reconstructed semantic conditions with their real counterparts, respectively. To further encourage distinguishable synthetic features, we learn discriminative category-level semantic conditions for cINN with an attributes mapping layer. To verify the proposed method, we conduct extensive experiments on five widely used benchmarks. Experimental results show that our method outperforms previous state-of-the-arts and successfully alleviates features confusion problem in GZSL. For instance, our method achieves the best performance in terms of seen accuracy, unseen accuracy and harmonic mean accuracy on FLO. Hongzu Su, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Uneven Bi-Classifier Learning for Domain AdaptationabstractThe bi-classifier paradigm is widely adopted as an adversarial method to address domain shift challenge in unsupervised domain adaptation (UDA) by evenly training two classifiers. In this paper, we report that although the generalization ability of the feature extractor can be strengthened by the two even classifiers, the decision boundaries of the two classifiers would be shrank to the source domain in the adversarial process, which weakens the discriminative ability of the learned model. To tame this dilemma, we disentangle the function of the two classifiers and introduce uneven bi-classifier learning for domain adaptation. Specifically, we leverage the F-norm (Frobenius Norm) of classifier predictions instead of the classifier disagreement to achieve adversarial learning. By this way, our feature extractor can be adversarially trained with a single classifier and the other classifier is used for preserving the target-specific decision boundaries. The proposed uneven bi-classifier learning protocol can simultaneously enhance the generalization ability of the feature extractor and expand the decision boundary of the target classifier. Extensive experiments on large-scale datasets prove that our method can significantly surpass previous domain adaptation methods, even with only a single classifier being involved. Zhiqi Yu, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Classification Certainty Maximization for Unsupervised Domain AdaptationabstractThe bi-classifier paradigm is a common practice in unsupervised domain adaptation (UDA), where two classifiers are leveraged to guide the model to learn domain invariant features. Previous approaches only focused on the consistency of the outputs between classifiers, but ignored the classification certainty of each classifier. Therefore, existing methods in some cases may mislead the classifiers into the wrong direction of ambiguous outputs and subsequently undermine the discriminability. To challenge this problem, in this paper we propose Classification Certainty Maximization (CCM) which considers both the joint certainty between classifiers and the individual certainty of each classifier via a novel formulation, and derived their optimal weight ratios by theoretical derivation from the perspective of gradient. In addition, we also propose a dynamic centroid update strategy to mitigate the domain gaps at the feature level. Extensive experiments on four widely used UDA datasets show that CCM performs better than the existing state-of-the-art domain adaptation methods. Notably, our dynamic centroid update strategy can be used as a plug-and-play module for existing bi-classifier domain adaptation methods to boost classification accuracy. Zhiqi Yu, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Manifold Regularized Joint Transfer for Open Set Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to leverage knowledge of a well-labeled source domain to learn an effective classifier for an unlabeled target domain. However, a common scenario in real-world applications is that the target domain contains unknown categories that are not observed in the source domain. This setting is termed as open set domain adaptation (OSDA). Most existing approaches of OSDA can only classify known classes well but fail to recognize unknown samples effectively. In this paper, we propose an effective method, named manifold regularized joint transfer (MRJT), for OSDA. MRJT learns new feature representations by simultaneously reducing distribution discrepancy between domains, increasing compactness of within-class, discriminating different known classes, and distinguishing the unknown from the known. The learned new features are projected onto reproducing kernel Hilbert space. In this space, a weighted structural risk minimization method is integrated with manifold regularization to utilize geometric information sufficiently to learn an effective classifier. Extensive experimental results on four real-world datasets verify the superiority of our method. It can not only classify known samples into the right known classes but also recognize unknown samples effectively. Jieyan Liu, Hongcai He, Mingzhu Liu, Jingjing Li 0001, Ke Lu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Open Set Domain Adaptation via Joint Alignment and Category SeparationabstractPrevalent domain adaptation approaches are suitable for a close-set scenario where the source domain and the target domain are assumed to share the same data categories. However, this assumption is often violated in real-world conditions where the target domain usually contains samples of categories that are not presented in the source domain. This setting is termed as open set domain adaptation (OSDA). Most existing domain adaptation approaches do not work well in this situation. In this article, we propose an effective method, named joint alignment and category separation (JACS), for OSDA. Specifically, JACS learns a latent shared space, where the marginal and conditional divergence of feature distributions for commonly known classes across domains is alleviated (Joint Alignment), the distribution discrepancy between the known classes and the unknown class is enlarged, and the distance between different known classes is also maximized (Category Separation). These two aspects are unified into an objective to reinforce the optimization of each part simultaneously. The classifier is achieved based on the learned new feature representations by minimizing the structural risk in the reproducing kernel Hilbert space. Extensive experiment results verify that our method outperforms other state-of-the-art approaches on several benchmark datasets. Jieyan Liu, Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Distinguishing Unseen from Seen for Generalized Zero-shot LearningabstractGeneralized zero-shot learning (GZSL) aims to recognize samples whose categories may not have been seen at training. Recognizing unseen classes as seen ones or vice versa often leads to poor performance in GZSL. Therefore, distinguishing seen and unseen domains is naturally an effective yet challenging solution for GZSL. In this paper, we present a novel method which leverages both visual and semantic modalities to distinguish seen and unseen categories. Specifically, our method deploys two variational autoencoders to generate latent representations for visual and semantic modalities in a shared latent space, in which we align latent representations of both modalities by Wasserstein distance and reconstruct two modalities with the representations of each other. In order to learn a clearer boundary between seen and unseen classes, we propose a two-stage training strategy which takes advantage of seen and unseen semantic descriptions and searches a threshold to separate seen and unseen visual samples. At last, a seen expert and an unseen expert are used for final classification. Extensive experiments on five widely used benchmarks verify that the proposed method can significantly improve the results of GZSL. For instance, our method correctly recognizes more than 99% samples when separating domains and improves the final classification accuracy from 72.6% to 82.9% on AWA1. Hongzu Su, Jingjing Li 0001, Zhi Chen 0010, Lei Zhu 0002, Ke Lu 0001 |
CVPR | 5 |
| 2022 | Towards Distributed Communication and Control in Real-World Multi-Agent Reinforcement LearningabstractMulti-agent system investigates the problem of designing a complex system composed of multiple autonomous agents with limited ability and partial observability. As a milestone, AlphaStar has achieved remarkable success in StarCraft II, which is a significant breakthrough in the competitive environments with complex strategic spaces and real-time decisions. However, it poses new challenges for deploying these centralized control models in real-world environments because many of them in such competitive environments were not designed to accommodate the requirements of real-world communication networks, e.g., the problems of high latency and large traffic are inevitable when they are actually deployed. To alleviate this issue, we propose a distributed control paradigm that explicitly splits the control power between the centralized meta-agent and agent units through a combination of centralized and decentralized paradigms. The units can autonomously decide to follow the decisions of the meta-agent or adapt to environment variations immediately by themselves in a decentralized manner. We simulate real-world network environments based on the Mininet platform, experiments based on the StarCraft II Learning Environment (SC2LE) show that our approach achieves a better adaptation in real-world network environments. Jieyan Liu, Zhekai Du, Ke Lu 0001 |
ICC | 4 |
| 2022 | Cross-Domain Communications Between Agents Via Adversarial-Based Domain Adaptation in Reinforcement LearningabstractReinforcement learning is suitable for solving sequential decision-making problems, and deep reinforcement learning methods have shown excellent performance in many fields. However, agents often face the challenge of a large number of interactions with the environment, which means that it is unrealistic to train agents from scratch in each new domain. In order to overcome this problem, our paper introduces the cross-domain communications between RL agents, so that agents in new domains can receive and use the information sent by agents trained in related domains to assist decision-making. Specifically, this paper uses adversarial-based domain adaptation methods and multi-granular loss constraints to realize implicit communications between cross-domain agents and encourage agents in different domains to extract domain-invariant information for communication and sharing, thereby the optimal behavior policy of the agent training based on the shared information in the source domain can be transferred to the related domain and achieve the expected performance. Finally, we evaluate our method on various variants of Car-Racing games, and the results show that this method can achieve efficient information communications between cross-domain agents and better performance than previous methods. Lichao Meng, Jingjing Li 0001, Ke Lu 0001 |
ICC | 3 |
| 2022 | Energy-Based Domain Generalization for Face Anti-SpoofingabstractWith various unforeseeable face presentation attacks (PA) springing up, face anti-spoofing (FAS) urgently needs to generalize to unseen scenarios. Research on generalizable FAS has lately attracted growing attention. Existing methods cast FAS as a vanilla binary classification problem and address it by a standard discriminative classifier p(y|x) under a domain generalization framework. However, discriminative models are unreliable for samples far away from the training distribution. In this paper, we resort to an energy-based model (EBM) to tackle FAS in a generative perspective. Our motivation is to model the joint density p(x,y), which allows to compute not only p(y|x) but also p(x). Due to the intractability of direct modeling, we use EBMs as an alternative to probabilistic estimation. With energy-based training, real faces are encouraged to get low free energy associated with the marginal probability p(x) of real faces, and all samples with high free energy are regarded as fake faces, thus rejecting any kind of PA out of the distribution of real faces. To learn to generalize to unseen domains, we generate diverse and novel populations in feature space under the guidance of energy model. Our model is updated in a meta-learning schema, where the original source samples are utilized for meta-training and the generated ones for meta-testing. We validate our method on four widely used FAS datasets. Comprehensive experimental results demonstrate the effectiveness of our method compared with state-of-the-arts. Zhekai Du, Jingjing Li 0001, Lin Zuo, Lei Zhu 0002, Ke Lu 0001 |
ACM Multimedia | 5 |
| 2022 | Source-Free Active Domain Adaptation via Energy-Based Locality Preserving TransferabstractUnsupervised domain adaptation (UDA) aims at transferring knowledge from one labeled source domain to a related but unlabeled target domain. Recently, active domain adaptation (ADA) has been proposed as a new paradigm which significantly boosts performance of UDA with minor additional labeling. However, existing ADA methods require source data to explicitly measure the domain gap between the source domain and the target domain, which is restricted in many real-world scenarios. In this work, we handle ADA with only a source-pretrained model and unlabeled target data, proposing a new setting named source-free active domain adaptation. Specifically, we propose a Locality Preserving Transfer (LPT) framework which preserves and utilizes locality structures on target data to achieve adaptation without source data. Meanwhile, a label propagation strategy is adopted to improve the discriminability for better adaptation. After LPT, unique samples with insignificant locality structure are identified by an energy-based approach for active annotation. An energy-based pseudo labeling strategy is further applied to generate labels for reliable samples. Finally, with supervision from the annotated samples and pseudo labels, a well adapted model is obtained. Extensive experiments on three widely used UDA benchmarks show that our method is comparable or superior to current state-of-the-art active domain adaptation methods even without access to source data. Zhekai Du, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001 |
ACM Multimedia | 5 |
| 2022 | Domain adaptive state representation alignment for reinforcement learning
Lichao Meng, Jingjing Li 0001, Ke Lu 0001, Yang Yang 0002 |
Inf. Sci. | 4 |
| 2022 | Divergence-Agnostic Unsupervised Domain Adaptation by Adversarial AttacksabstractConventional machine learning algorithms suffer the problem that the model trained on existing data fails to generalize well to the data sampled from other distributions. To tackle this issue, unsupervised domain adaptation (UDA) transfers the knowledge learned from a well-labeled source domain to a different but related target domain where labeled data is unavailable. The majority of existing UDA methods assume that data from the source domain and the target domain are available and complete during training. Thus, the divergence between the two domains can be formulated and minimized. In this paper, we consider a more practical yet challenging UDA setting where either the source domain data or the target domain data are unknown. Conventional UDA methods would fail this setting since the domain divergence is agnostic due to the absence of the source data or the target data. Technically, we investigate UDA from a novel view-adversarial attack-and tackle the divergence-agnostic adaptive learning problem in a unified framework. Specifically, we first report the motivation of our approach by investigating the inherent relationship between UDA and adversarial attacks. Then we elaborately design adversarial examples to attack the training model and harness these adversarial examples. We argue that the generalization ability of the model would be significantly improved if it can defend against our attack, so as to improve the performance on the target domain. Theoretically, we analyze the generalization bound for our method based on domain adaptation theories. Extensive experimental results on multiple UDA benchmarks under conventional, source-absent and target-absent UDA settings verify that our method is able to achieve a favorable performance compared with previous ones. Notably, this work extends the scope of both domain adaptation and adversarial attack, and expected to inspire more ideas in the community. Jingjing Li 0001, Zhekai Du, Lei Zhu 0002, Zhengming Ding, Ke Lu 0001, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Investigating the Bilateral Connections in Generative Zero-Shot LearningabstractZero-shot learning (ZSL) is a pretty intriguing topic in the computer vision community since it handles novel instances and unseen categories. In a typical ZSL setting, there is a main visual space and an auxiliary semantic space. Most existing ZSL methods handle the problem by learning either a visual-to-semantic mapping or a semantic-to-visual mapping. In other words, they investigate a unilateral connection from one end to the other. However, the connection between the visual space and the semantic space are bilateral in reality, that is, the visual space depicts the semantic space; the semantic space, on the other hand, describes the visual space. In this article, therefore, we investigate the bilateral connections in ZSL and present a novel model, called Boomerang-GAN, by taking advantage of conditional generative adversarial networks (GANs). Specifically, we generate unseen visual samples from their category semantic embeddings by a conditional GAN. Different from the existing generative ZSL methods that only consider generating visual features from class descriptions, our method also considers that the generated visual features can be translated back to their corresponding semantic embeddings by introducing a multimodal cycle-consistent loss. Extensive experiments of both ZSL and generalized ZSL on five widely used datasets verify that our method is able to outperform previous state-of-the-art approaches in both recognition and segmentation tasks. Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Cybern. | 3 |
| 2022 | Faster Domain Adaptation NetworksabstractIt is widely acknowledged that the success of deep learning is built upon large-scale training data and tremendous computing power. However, the data and computing power are not always available for many real-world applications. In this paper, we address the machine learning problem where it lacks training data and limits computing power. Specifically, we investigate domain adaptation which is able to transfer knowledge from one labeled source domain to an unlabeled target domain, so that we do not need much training data from the target domain. At the same time, we consider the situation that the running environment is confined, e.g., in edge computing the end device has very limited running resources. Technically, we present the Faster Domain Adaptation (FDA) protocol and further report two paradigms of FDA: early stopping and amid skipping. The former accelerates domain adaptation by multiple early exit points. The latter speeds up the adaptation by wisely skip several amid neural network blocks. Extensive experiments on standard benchmarks verify that our method is able to achieve the comparable and even better accuracy but employ much less computing resources. To the best of our knowledge, there are very few works which investigated accelerating knowledge adaptation in the community. This work is expected to inspire the topic for more discussion. Jingjing Li 0001, Mengmeng Jing, Hongzu Su, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Adversarial Entropy Optimization for Unsupervised Domain AdaptationabstractDomain adaptation is proposed to deal with the challenging problem where the probability distribution of the training source is different from the testing target. Recently, adversarial learning has become the dominating technique for domain adaptation. Usually, adversarial domain adaptation methods simultaneously train a feature learner and a domain discriminator to learn domain-invariant features. Accordingly, how to effectively train the domain-adversarial model to learn domain-invariant features becomes a challenge in the community. To this end, we propose in this article a novel domain adaptation scheme named adversarial entropy optimization (AEO) to address the challenge. Specifically, we minimize the entropy when samples are from the independent distributions of source domain or target domain to improve the discriminability of the model. At the same time, we maximize the entropy when features are from the combined distribution of source domain and target domain so that the domain discriminator can be confused and the transferability of representations can be promoted. This minimax regime is well matched with the core idea of adversarial learning, empowering our model with transferability as well as discriminability for domain adaptation tasks. Also, AEO is flexible and compatible with different deep networks and domain adaptation frameworks. Experiments on five data sets show that our method can achieve state-of-the-art performance across diverse domain adaptation tasks. Ao Ma 0001, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Resource-Efficient Distributed Deep Neural Networks Empowered by Intelligent Software-Defined NetworkingabstractContemporary machine learning methods have evolved from conventional algorithms to deep neural networks (DNNs) that are computation- and data- intensive. Thus, they are suitable to be deployed in the cloud that can offer high computational capacity and scalable resources. However, the cloud computing paradigm is not optimal for delay- and energy-sensitive applications. To mitigate these problems, a battery of distributed DNNs have been proposed to allow a fast inference with device-edge-cloud synergy. Furthermore, although distributed deployment of DNNs on real communication networks is an important research topic, the legacy network architecture cannot meet the requirements of these distributed deep neural networks due to the complicated management and manual configuration, etc. To cope with these requirements, we develop a novel and explicit Intelligent Software Defined Networking (ISDN) that aims to manage the bandwidth and computing resources across the network via the SDN paradigm. We first identify the difficulties of deploying distributed intelligent computing in the current network architecture. Then, we explain how to address these problems by introducing the ISDN architecture. Specifically, we develop a dynamic routing method to enable Quality-of-Service (QoS) communication based on the SDN paradigm and propose a Markov Decision Process (MDP) based dynamic task offloading model to achieve the optimal offloading policy of DNN tasks. We develop a simulation platform based on Mininet to measure its performance advantages over traditional architectures. Extensive experimental results show that compared with the traditional network architecture, our architecture based on the SDN paradigm can perform better in terms of both network throughput and resource utilization. Ke Lu 0001, Zhekai Du, Jingjing Li 0001, Geyong Min |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Balanced Open Set Domain Adaptation via Centroid AlignmentabstractOpen Set Domain Adaptation (OSDA) is a challenging domain adaptation setting which allows the existence of unknown classes on the target domain. Although existing OSDA methods are good at classifying samples of known classes, they ignore the classification ability for the unknown samples, making them unbalanced OSDA methods. To alleviate this problem, we propose a balanced OSDA methods which could recognize the unknown samples while maintain high classification performance for the known samples. Specifically, to reduce the domain gaps, we first project the features to a hyperspherical latent space. In this space, we propose to bound the centroid deviation angles to not only increase the intra-class compactness but also enlarge the inter-class margins. With the bounded centroid deviation angles, we employ the statistical Extreme Value Theory to recognize the unknown samples that are misclassified into known classes. In addition, to learn better centroids, we propose an improved centroid update strategy based on sample reweighting and adaptive update rate to cooperate with centroid alignment. Experimental results on three OSDA benchmarks verify that our method can significantly outperform the compared methods and reduce the proportion of the unknown samples being misclassified into known classes. Mengmeng Jing, Jingjing Li 0001, Lei Zhu 0002, Zhengming Ding, Ke Lu 0001, Yang Yang 0002 |
AAAI | 5 |
| 2021 | Cross-Domain Gradient Discrepancy Minimization for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to generalize the knowledge learned from a well-labeled source domain to an unlabled target domain. Recently, adversarial domain adaptation with two distinct classifiers (biclassifier) has been introduced into UDA which is effective to align distributions between different domains. Previous bi-classifier adversarial learning methods only focus on the similarity between the outputs of two distinct classifiers. However, the similarity of the outputs cannot guarantee the accuracy of target samples, i.e., traget samples may match to wrong categories even if the discrepancy between two classifiers is small. To challenge this issue, in this paper, we propose a cross-domain gradient discrepancy minimization (CGDM) method which explicitly minimizes the discrepancy of gradients generated by source samples and target samples. Specifically, the gradient gives a cue for the semantic information of target samples so it can be used as a good supervision to improve the accuracy of target samples. In order to compute the gradient signal of target smaples, we further obtain target pseudo labels through a clustering-based self-supervised learning. Extensive experiments on three widely used UDA datasets show that our method surpasses many previous state-of-the-arts. Zhekai Du, Jingjing Li 0001, Hongzu Su, Lei Zhu 0002, Ke Lu 0001 |
CVPR | 5 |
| 2021 | Learning Transferrable and Interpretable Representations for Domain GeneralizationabstractConventional machine learning models are often vulnerable to samples with different distributions from the ones of training samples, which is known as domain shift. Domain Generalization (DG) challenges this issue by training a model based on multiple source domains and generalizing it to arbitrary unseen target domains. In spite of remarkable results made in DG, a majority of existing works lack a deep understanding of the feature representations learned in DG models, resulting in limited generalization ability when facing domainsout-of-distribution. In this paper, we aim to learn a domain transformation space via a domain transformer network (DTN) which explicitly mines the relationship among multiple domains and constructs transferable feature representations for down-stream tasks by interpreting each feature as a semantically weighted combination of multiple domain-specific features. Our DTN is encouraged to meta-learn the properties and characteristics of domains during the training process based on multiple seen domains, making transformed feature representations more semantical, thus generalizing better to unseen domains. Once the model is constructed, the feature representations of unseen target domains can also be inferred adaptively by selectively combining the feature representations from the diverse set of seen domains. We conduct extensive experiments on five DG benchmarks and the results strongly demonstrate the effectiveness of our approach. Zhekai Du, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Zi Huang |
ACM Multimedia | 3 |
| 2021 | Learning a Weighted Classifier for Conditional Domain Adaptation
Fuming You, Hongzu Su, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002 |
Knowl. Based Syst. | 5 |
| 2021 | Maximum Density Divergence for Domain AdaptationabstractUnsupervised domain adaptation addresses the problem of transferring knowledge from a well-labeled source domain to an unlabeled target domain where the two domains have distinctive data distributions. Thus, the essence of domain adaptation is to mitigate the distribution divergence between the two domains. The state-of-the-art methods practice this very idea by either conducting adversarial training or minimizing a metric which defines the distribution gaps. In this paper, we propose a new domain adaptation method named adversarial tight match (ATM) which enjoys the benefits of both adversarial training and metric learning. Specifically, at first, we propose a novel distance loss, named maximum density divergence (MDD), to quantify the distribution divergence. MDD minimizes the inter-domain divergence ("match" in ATM) and maximizes the intra-class density ("tight" in ATM). Then, to address the equilibrium challenge issue in adversarial domain adaptation, we consider leveraging the proposed MDD into adversarial domain adaptation framework. At last, we tailor the proposed MDD as a practical learning loss and report our ATM. Both empirical evaluation and theoretical analysis are reported to verify the effectiveness of the proposed method. The experimental results on four benchmarks, both classical and large-scale, show that our method is able to achieve new state-of-the-art performance on most evaluations. Jingjing Li 0001, Erpeng Chen, Zhengming Ding, Lei Zhu 0002, Ke Lu 0001, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Challenging tough samples in unsupervised domain adaptation
Lin Zuo, Mengmeng Jing, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002 |
Pattern Recognit. | 5 |
| 2021 | On Both Cold-Start and Long-Tail Recommendation with Social DataabstractThe number of “hits” has been widely regarded as the lifeblood of many web systems, e.g., e-commerce systems, advertising systems and multimedia consumption systems. However, users would not hit an item if they cannot see it, or they are not interested in the item. Recommender system plays a critical role of discovering interesting items from near-infinite inventory and exhibiting them to potential users. Yet, two issues are crippling the recommender systems. One is “how to handle new users”, and the other is “how to surprise users”. The former is well-known as cold-start recommendation. In this paper, we show that the latter can be investigated as long-tail recommendation. We also exploit the benefits of jointly challenging both cold-start and long-tail recommendation, and propose a novel approach which can simultaneously handle both of them in a unified objective. For the cold-start problem, we learn from side information, e.g., user attributes, user social relationships, etc. Then, we transfer the learned knowledge to new users. For the long-tail recommendation, we decompose the overall interesting items into two parts: a low-rank part for short-head items and a sparse part for long-tail items. The two parts are independently revealed in the training stage, and transfered into the final recommendation for new users. Furthermore, we effectively formulate the two problems into a unified objective and present an iterative optimization algorithm. A fast extension of the method is proposed to reduce the complexity, and extensive theoretical analysis are provided to proof the bounds of our approach. At last, experiments of social recommendation on various real-world datasets, e.g., images, blogs, videos and musics, verify the superiority of our approach compared with the state-of-the-art work. Jingjing Li 0001, Ke Lu 0001, Zi Huang, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersabstractLearning the relationship between the multi-modal data, e.g., texts, images and videos, is a classic task in the multimedia community. Cross-modal retrieval (CMR) is a typical example where the query and the corresponding results are in different modalities. Yet, a majority of existing works investigate CMR with an ideal assumption that the training samples in every modality are sufficient and complete. In real-world applications, however, this assumption does not always hold. Mismatch is common in multi-modal datasets. There is a high chance that samples in some modalities are either missing or corrupted. As a result, incomplete CMR has become a challenging issue. In this paper, we propose a Dual-Aligned Variational Autoencoders (DAVAE) to address the incomplete CMR problem. Specifically, we propose to learn modality-invariant representations for different modalities and use the learned representations for retrieval. We train multiple autoencoders, one for each modality, to learn the latent factors among different modalities. These latent representations are further dual-aligned at the distribution level and the semantic level to alleviate the modality gaps and enhance the discriminability of representations. For missing instances, we leverage generative models to synthesize latent representations for them. Notably, we test our method with different ratios of random incompleteness.Extensive experiments on three datasets verify that our method can consistently outperform the state-of-the-arts. Mengmeng Jing, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002, Zi Huang |
ACM Multimedia | 4 |
| 2020 | Learning Modality-Invariant Latent Representations for Generalized Zero-shot LearningabstractRecently, feature generating methods have been successfully applied to zero-shot learning (ZSL). However, most previous approaches only generate visual representations for zero-shot recognition. In fact, typical ZSL is a classic multi-modal learning protocol which consists of a visual space and a semantic space. In this paper, therefore, we present a new method which can simultaneously generate both visual representations and semantic representations so that the essential multi-modal information associated with unseen classes can be captured. Specifically, we address the most challenging issue in such a paradigm, i.e., how to handle the domain shift and thus guarantee that the learned representations are modality-invariant. To this end, we propose two strategies: 1) leveraging the mutual information between the latent visual representations and the semantic representations; 2) maximizing the entropy of the joint distribution of the two latent representations. By leveraging the two strategies, we argue that the two modalities can be well aligned. At last, extensive experiments on five widely used datasets verify that the proposed method is able to significantly outperform previous the state-of-the-arts. Jingjing Li 0001, Mengmeng Jing, Lei Zhu 0002, Zhengming Ding, Ke Lu 0001, Yang Yang 0002 |
ACM Multimedia | 5 |
| 2020 | Robust optimal graph clustering
Fei Wang 0055, Lei Zhu 0002, Cheng Liang 0001, Jingjing Li 0001, Xiaojun Chang, Ke Lu 0001 |
Neurocomputing | 6 |
| 2020 | Multi-source domain adaptation with graph embedding and adaptive label prediction
Ao Ma 0001, Fuming You, Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001 |
Inf. Process. Manag. | 5 |
| 2020 | Joint metric and feature representation learning for unsupervised domain adaptation
Zhekai Du, Jingjing Li 0001, Mengmeng Jing, Erpeng Chen, Ke Lu 0001 |
Knowl. Based Syst. | 6 |
| 2020 | Learning explicitly transferable representations for domain adaptation
Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Yang Yang 0002 |
Neural Networks | 3 |
| 2019 | From Zero-Shot Learning to Cold-Start RecommendationabstractZero-shot learning (ZSL) and cold-start recommendation (CSR) are two challenging problems in computer vision and recommender system, respectively. In general, they are independently investigated in different communities. This paper, however, reveals that ZSL and CSR are two extensions of the same intension. Both of them, for instance, attempt to predict unseen classes and involve two spaces, one for direct feature representation and the other for supplementary description. Yet there is no existing approach which addresses CSR from the ZSL perspective. This work, for the first time, formulates CSR as a ZSL problem, and a tailor-made ZSL method is proposed to handle CSR. Specifically, we propose a Lowrank Linear Auto-Encoder (LLAE), which challenges three cruxes, i.e., domain shift, spurious correlations and computing efficiency, in this paper. LLAE consists of two parts, a low-rank encoder maps user behavior into user attributes and a symmetric decoder reconstructs user behavior from user attributes. Extensive experiments on both ZSL and CSR tasks verify that the proposed method is a win-win formulation, i.e., not only can CSR be handled by ZSL models with a significant performance improvement compared with several conventional state-of-the-art methods, but the consideration of CSR can benefit ZSL as well. Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Yang Yang 0002, Zi Huang |
AAAI | 3 |
| 2019 | Leveraging the Invariant Side of Generative Zero-Shot LearningabstractConventional zero-shot learning (ZSL) methods generally learn an embedding, e.g., visual-semantic mapping, to handle the unseen visual samples via an indirect manner. In this paper, we take the advantage of generative adversarial networks (GANs) and propose a novel method, named leveraging invariant side GAN (LisGAN), which can directly generate the unseen features from random noises which are conditioned by the semantic descriptions. Specifically, we train a conditional Wasserstein GANs in which the generator synthesizes fake unseen features from noises and the discriminator distinguishes the fake from real via a minimax game. Considering that one semantic description can correspond to various synthesized visual samples, and the semantic description, figuratively, is the soul of the generated features, we introduce soul samples as the invariant side of generative zero-shot learning in this paper. A soul sample is the meta-representation of one class. It visualizes the most semantically-meaningful aspects of each sample in the same category. We regularize that each generated sample (the varying side of generative ZSL) should be close to at least one soul sample (the invariant side) which has the same class label with it. At the zero-shot recognition stage, we propose to use two classifiers, which are deployed in a cascade way, to achieve a coarse-to-fine result. Experiments on five popular benchmarks verify that our proposed approach can outperform state-of-the-art methods with significant improvements. Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Zhengming Ding, Lei Zhu 0002, Zi Huang |
CVPR | 3 |
| 2019 | Adaptive Component Embedding for Unsupervised Domain AdaptationabstractDomain adaptation has obtained considerable interest from the literatures of multimedia, especially in cross-domain knowledge transfer problems. In this paper, we propose an effective yet time-saving approach, named Adaptive Component Embedding (ACE), for unsupervised domain adaptation. Specifically, ACE learns adaptive components across domains to embed all data in a shared subspace where the distribution divergence is mitigated and the underlying geometric structures in the local manifold are preserved. Then, an adaptive classifier is learned by using Representer Theorem in the Reproducing Kernel Hilbert Space (RKHS). The objective of our method can be efficiently solved in a closed form. Comprehensive experiments on both standard and large-scale datasets verify that ACE significantly outperforms previous state-of-the-art methods in terms of the classification accuracy and training time. Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Jieyan Liu, Zi Huang |
ICME | 3 |
| 2019 | Agile Domain AdaptationabstractDomain adaptation investigates the problem of leveraging knowledge from a well-labeled source domain to an unlabeled target domain, where the two domains are drawn from different data distributions. Because of the distribution shifts, different target samples have distinct degrees of difficulty in adaptation. However, existing domain adaptation approaches overwhelmingly neglect the degrees of difficulty and deploy exactly the same framework for all of the target samples. Generally, a simple or shadow framework is fast but rough. A sophisticated or deep framework, on the contrary, is accurate but slow. In this paper, we aim to challenge the fundamental contradiction between the accuracy and speed in domain adaptation tasks. We propose a novel approach, named agile domain adaptation, which agilely applies optimal frameworks to different target samples and classifies the target samples according to their adaptation difficulties. Specifically, we propose a paradigm which performs several early detections before the final classification. If a sample can be classified at one of the early stage with enough confidence, the sample would exit without the subsequent processes. Notably, the proposed method can significantly reduce the running cost of domain adaptation approaches, which can extend the application scenarios of domain adaptation to even mobile devices and real-time systems. Extensive experiments on two open benchmarks verify the effectiveness and efficiency of the proposed method. Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Zi Huang |
IJCNN | 4 |
| 2019 | Cycle-consistent Conditional Adversarial Transfer NetworksabstractDomain adaptation investigates the problem of cross-domain knowledge transfer where the labeled source domain and unlabeled target domain have distinctive data distributions. Recently, adversarial training have been successfully applied to domain adaptation and achieved state-of-the-art performance. However, there is still a fatal weakness existing in current adversarial models which is raised from the equilibrium challenge of adversarial training. Specifically, although most of existing methods are able to confuse the domain discriminator, they cannot guarantee that the source domain and target domain are sufficiently similar. In this paper, we propose a novel approach named cycle-consistent conditional adversarial transfer networks (3CATN) to handle this issue. Our approach takes care of the domain alignment by leveraging adversarial training. Specifically, we condition the adversarial networks with the cross-covariance of learned features and classifier predictions to capture the multimodal structures of data distributions. However, since the classifier predictions are not certainty information, a strong condition with the predictions is risky when the predictions are not accurate. We, therefore, further propose that the truly domain-invariant features should be able to be translated from one domain to the other. To this end, we introduce two feature translation losses and one cycle-consistent loss into the conditional adversarial domain adaptation networks. Extensive experiments on both classical and large-scale datasets verify that our model is able to outperform previous state-of-the-arts with significant improvements. Jingjing Li 0001, Erpeng Chen, Zhengming Ding, Lei Zhu 0002, Ke Lu 0001, Zi Huang |
ACM Multimedia | 5 |
| 2019 | Alleviating Feature Confusion for Generative Zero-shot LearningabstractLately, generative adversarial networks (GANs) have been successfully applied to zero-shot learning (ZSL) and achieved state-of-the-art performance. By synthesizing virtual unseen visual features, GAN-based methods convert the challenging ZSL task into a supervised learning problem. However, since real unseen visual features are not available at the training stage, GAN-based ZSL methods have to train the GAN generator on the seen categories and further apply it to unseen instances. An inevitable issue of such a paradigm is that the synthesized unseen features are prone to seen references and incapable to reflect the novelty and diversity of real unseen instances. In a nutshell, the synthesized features are confusing. One cannot tell unseen categories from seen ones using the synthesized features. As a result, the synthesized features are too subtle to be classified in generalized zero-shot learning (GZSL) which involves both seen and unseen categories at the test stage. In this paper, we first introduce the feature confusion issue. Then, we propose a new feature generating network, named alleviating feature confusion GAN (AFC-GAN), to challenge the issue. Specifically, we present a boundary loss which maximizes the decision boundary of seen categories and unseen ones. Furthermore, a novel metric named feature confusion score (FCS) is proposed to quantify the feature confusion. Extensive experiments on five widely used datasets verify that our method is able to outperform previous state-of-the-arts under both ZSL and GZSL protocols. Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Yang Yang 0002, Zi Huang |
ACM Multimedia | 3 |
| 2019 | Transfer Independently Together: A Generalized Framework for Domain AdaptationabstractCurrently, unsupervised heterogeneous domain adaptation in a generalized setting, which is the most common scenario in real-world applications, is under insufficient exploration. Existing approaches either are limited to special cases or require labeled target samples for training. This paper aims to overcome these limitations by proposing a generalized framework, named as transfer independently together (TIT). Specifically, we learn multiple transformations, one for each domain (independently), to map data onto a shared latent space, where the domains are well aligned. The multiple transformations are jointly optimized in a unified framework (together) by an effective formulation. In addition, to learn robust transformations, we further propose a novel landmark selection algorithm to reweight samples, i.e., increase the weight of pivot samples and decrease the weight of outliers. Our landmark selection is based on graph optimization. It focuses on sample geometric relationship rather than sample features. As a result, by abstracting feature vectors to graph vertices, only a simple and fast integer arithmetic is involved in our algorithm instead of matrix operations with float point arithmetic in existing approaches. At last, we effectively optimize our objective via a dimensionality reduction procedure. TIT is applicable to arbitrary sample dimensionality and does not need labeled target samples for training. Extensive evaluations on several standard benchmarks and large-scale datasets of image classification, text categorization and text-to-image recognition verify the superiority of our approach. Jingjing Li 0001, Ke Lu 0001, Zi Huang, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Cybern. | 2 |
| 2019 | Locality Preserving Joint Transfer for Domain AdaptationabstractDomain adaptation aims to leverage knowledge from a well-labeled source domain to a poorly labeled target domain. A majority of existing works transfer the knowledge at either feature level or sample level. Recent studies reveal that both of the paradigms are essentially important, and optimizing one of them can reinforce the other. Inspired by this, we propose a novel approach to jointly exploit feature adaptation with distribution matching and sample adaptation with landmark selection. During the knowledge transfer, we also take the local consistency between the samples into consideration so that the manifold structures of samples can be preserved. At last, we deploy label propagation to predict the categories of new instances. Notably, our approach is suitable for both homogeneous- and heterogeneous-domain adaptations by learning domain-specific projections. Extensive experiments on five open benchmarks, which consist of both standard and large-scale datasets, verify that our approach can significantly outperform not only conventional approaches but also end-to-end deep models. The experiments also demonstrate that we can leverage handcrafted features to promote the accuracy on deep features by heterogeneous adaptation. Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 3 |
| 2019 | Heterogeneous Domain Adaptation Through Progressive AlignmentabstractIn real-world transfer learning tasks, especially in cross-modal applications, the source domain and the target domain often have different features and distributions, which are well known as the heterogeneous domain adaptation (HDA) problem. Yet, existing HDA methods focus on either alleviating the feature discrepancy or mitigating the distribution divergence due to the challenges of HDA. In fact, optimizing one of them can reinforce the other. In this paper, we propose a novel HDA method that can optimize both feature discrepancy and distribution divergence in a unified objective function. Specifically, we present progressive alignment, which first learns a new transferable feature space by dictionary-sharing coding, and then aligns the distribution gaps on the new space. Different from previous HDA methods that are limited to specific scenarios, our approach can handle diverse features with arbitrary dimensions. Extensive experiments on various transfer learning tasks, such as image classification, text categorization, and text-to-image recognition, verify the superiority of our method against several state-of-the-art approaches. Jingjing Li 0001, Ke Lu 0001, Zi Huang, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Learning Distribution-Matched Landmarks for Unsupervised Domain Adaptation
Mengmeng Jing, Jingjing Li 0001, Jidong Zhao, Ke Lu 0001 |
DASFAA (2) | 4 |
| 2018 | I read, I saw, I tell: Texts Assisted Fine-Grained Visual ClassificationabstractIn visual classification tasks, it is hard to tell the subtle differences from one species to another similar breeds. Such a challenging problem is generally known as Fine-Grained Visual Classification (FGVC). In this paper, we propose a novel FGVC approach called Texts Assisted Fine-Grained Visual Classification (TA-FGVC). TA-FGVC reads from texts to gain attention, sees the images with the gained attention and then tells the subtle differences. Technically, we propose a deep neural network which learns a visual-semantic embedding model. The proposed deep architecture mainly consists of two parts: one for visual localization, and the other for visual to semantic projection. The model is fed with both visual features which are extracted from raw images and semantic information which are learned from two sources: gleaned from unannotated texts and gathered from image attributes. At the very last layer of the model, each image is embedded into the semantic space which is related to class labels. Finally, the categorization results from both visual stream and visual-semantic stream are combined to achieve the ultimate decision. Extensive experiments on open standard benchmarks verify the superiority of our model against several state of the art work. Jingjing Li 0001, Lei Zhu 0002, Zi Huang, Ke Lu 0001, Jidong Zhao |
ACM Multimedia | 4 |
| 2018 | Coupled local-global adaptation for multi-source transfer learning
Jieyan Liu, Jingjing Li 0001, Ke Lu 0001 |
Neurocomputing | 3 |
| 2017 | Two Birds One Stone: On both Cold-Start and Long-Tail RecommendationabstractThe number of "hits" has been widely regarded as the lifeblood of many web systems, e.g., e-commerce systems, advertising systems and multimedia consumption systems. However, users would not hit an item if they cannot see it, or they are not interested in the item. Recommender system plays a critical role of discovering interested items from near-infinite inventory and exhibiting them to potential users. Yet, two issues are crippling the recommender systems. One is "how to handle new users", and the other is "how to surprise users". The former is well-known as cold-start recommendation, and the latter can be investigated as long-tail recommendation. This paper, for the first time, proposes a novel approach which can simultaneously handle both cold-start and long-tail recommendation in a unified objective. Jingjing Li 0001, Ke Lu 0001, Zi Huang, Heng Tao Shen |
ACM Multimedia | 2 |
| 2017 | Structured Domain AdaptationabstractIn many real-world applications, labeled data are either expensive or too scarce to be used to train an accurate classifier. Therefore, it is worth exploring and often essential to make full use of existing resources. Domain adaptation is one of the most promising techniques of leveraging an existing well-labeled source domain and a limited labeled target domain. With the aim of better understanding new or unknown domains through well-labeled ones, this paper proposes an efficient and powerful algorithm, named structured domain adaptation (SDA), which transfers knowledge across two domains. Specifically, SDA aims to seek a discriminate subspace shared by two domains where the well-learned knowledge of the source domain can be transferred to the target domain. In SDA, samples from both domains are combined together to reveal more shared information across two domains. Furthermore, an iteratively structured matrix$H$is learned to bridge the domain shift so that the marginal and conditional distributions are mitigated. Since the target is labeled to a limited extent or even totally unlabeled, we adopt pseudolabels of the target data to optimize$H$iteratively so that a reconstruction coefficient matrix is learned in order to guarantee local awareness. Our approach is robust to outliers as it applies an$\ell _{2,1}$-norm on the error term. SDA can work in either an unsupervised or a semisupervised manner. Extensive experiments on five data sets including faces, objects, digits, and visual events, demonstrate that SDA outperforms several state-of-the-art approaches with significant advantages. Notably, SDA can even achieve 100% accuracy on several popular benchmarks. Jingjing Li 0001, Ke Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Low-Rank Discriminant Embedding for Multiview LearningabstractThis paper focuses on the specific problem of multiview learning where samples have the same feature set but different probability distributions, e.g., different viewpoints or different modalities. Since samples lying in different distributions cannot be compared directly, this paper aims to learn a latent subspace shared by multiple views assuming that the input views are generated from this latent subspace. Previous approaches usually learn the common subspace by either maximizing the empirical likelihood, or preserving the geometric structure. However, considering the complementarity between the two objectives, this paper proposes a novel approach, named low-rank discriminant embedding (LRDE), for multiview learning by taking full advantage of both sides. By further considering the duality between data points and features of multiview scene, i.e., data points can be grouped based on their distribution on features, while features can be grouped based on their distribution on the data points, LRDE not only deploys low-rank constraints on both sample level and feature level to dig out the shared factors across different views, but also preserves geometric information in both the ambient sample space and the embedding feature space by designing a novel graph structure under the framework of graph embedding. Finally, LRDE jointly optimizes low-rank representation and graph embedding in a unified framework. Comprehensive experiments in both multiview manner and pairwise manner demonstrate that LRDE performs much better than previous approaches proposed in recent literatures. Jingjing Li 0001, Jidong Zhao, Ke Lu 0001 |
IEEE Trans. Cybern. | 4 |
| 2016 | Joint Feature Selection and Structure Preservation for Domain Adaptation
Jingjing Li 0001, Jidong Zhao, Ke Lu 0001 |
IJCAI | 3 |
| 2016 | Multi-manifold Sparse Graph Embedding for Multi-modal Image Classification
Jingjing Li 0001, Jidong Zhao, Ke Lu 0001 |
Neurocomputing | 4 |
| 2013 | Locally connected graph for visual tracking
Ke Lu 0001, Zhengming Ding, Shuzhi Sam Ge |
Neurocomputing | 1 |
| 2013 | G-Optimal Feature Selection with Laplacian regularization
Guanhong Yao, Ke Lu 0001, Xiaofei He 0001 |
Neurocomputing | 2 |
| 2012 | Sparse-Representation-Based Graph Embedding for Traffic Sign RecognitionabstractResearchers have proposed various machine learning algorithms for traffic sign recognition, which is a supervised multicategory classification problem with unbalanced class frequencies and various appearances. We present a novel graph embedding algorithm that strikes a balance between local manifold structures and global discriminative information. A novel graph structure is designed to depict explicitly the local manifold structures of traffic signs with various appearances and to intuitively model between-class discriminative information. Through this graph structure, our algorithm effectively learns a compact and discriminative subspace. Moreover, by using$L_{2, 1}$-norm, the proposed algorithm can preserve the sparse representation property in the original space after graph embedding, thereby generating a more accurate projection matrix. Experiments demonstrate that the proposed algorithm exhibits better performance than the recent state-of-the-art methods. Ke Lu 0001, Zhengming Ding, Shuzhi Sam Ge |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2011 | Neighborhood preserving regression for image retrieval
Ke Lu 0001, Jidong Zhao |
Neurocomputing | 1 |
| 2011 | Hessian optimal design for image retrieval
Ke Lu 0001, Jidong Zhao |
Pattern Recognit. | 1 |
| 2010 | Approximately harmonic projection: Theoretical analysis and an algorithm
Binbin Lin 0001, Xiaofei He 0001, Ke Lu 0001 |
Pattern Recognit. | 5 |
| 2009 | An efficient semi-blind source extraction algorithm and its applications to biomedical signal extraction
Yalan Ye, Phillip C.-Y. Sheu, Jiazhi Zeng, Ke Lu 0001 |
Sci. China Ser. F Inf. Sci. | 5 |
| 2008 | Locality sensitive semi-supervised feature selection
Jidong Zhao, Ke Lu 0001, Xiaofei He 0001 |
Neurocomputing | 2 |
| 2006 | Semi-supervised Support Vector Learning for Face Recognition
Ke Lu 0001, Xiaofei He 0001, Jidong Zhao |
ISNN (2) | 1 |
| 2006 | An algorithm for semi-supervised learning in image retrieval
Ke Lu 0001, Jidong Zhao, Deng Cai 0001 |
Pattern Recognit. | 1 |
| 2005 | Semi-supervised Learning for Image Retrieval Using Support Vector Machines
Ke Lu 0001, Jidong Zhao, Mengqin Xia, Jiazhi Zeng |
ISNN (1) | 1 |
| 2005 | Image retrieval based on incremental subspace learning
Ke Lu 0001, Xiaofei He 0001 |
Pattern Recognit. | 1 |
| 2004 | Locality pursuit embedding
Wanli Min, Ke Lu 0001, Xiaofei He 0001 |
Pattern Recognit. | 2 |