Luojun Lin

dblp:213/7881 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
31since 2021 · last 2026
0000-0002-1141-2487ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 15 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prior Refinement Is Better: Diffusion-Driven Graph Harmonization for Federated Graph Learning
abstract
Federated Graph Learning (FGL) has emerged as a compelling paradigm for collaboratively training a global model while preserving the privacy of multi-source graphs. Nonetheless, FGL faces a critical challenge of data heterogeneity, where semantic and structural discrepancies across clients significantly degrade its performance. Although existing methods attempt to calibrate client-specific graph distributions during federated training, they inevitably fall short in aligning the optimization behaviors across clients due to dynamic parameter updates, thereby inducing a bottleneck in generalization improvement. To tackle this challenge, we propose a solution from a new perspective of prior refinement, which seeks to proactively harmonize client graph distributions before the federated training. In particular, we propose a Federated Graph Harmonization (FedGH) framework that exploits the generative strengths of graph diffusion models to perform prior refinement of local graphs. In a nutshell, FedGH designs a conditional diffusion mechanism on each client that synthesizes pseudo-graphs encapsulating both feature and structural priors, thereby facilitating explicit correction of inter-client distributional bias. On the server side, we employ the graph contrastive learning between various client-specific pseudo-graphs to incorporate the global information, subsequently guiding local data reconstruction. Importantly, model-agnostic FedGH can be seamlessly deployed as a plug-and-play module to be easily integrated with existing FGL architectures. Extensive experiments demonstrate that FedGH consistently outperforms state-of-the-art FGL baselines.
Shuman Zhuang, Zhihao Wu 0003, Wei Huang 0013, Luojun Lin, Jiali Yin, Lele Fu, Hongning Dai
AAAI4
2026 Uncertainty-aware multi-instance partial-label learning via evidential deep model
Gaowen Jie, Fumiao Wang, Gaojie Song, Luojun Lin, Yuanlong Yu 0001, Qinghai Zheng
Neurocomputing4
2025 A Tiny Change, a Giant Leap: Long-Tailed Class-Incremental Learning via Geometric Prototype Alignment
Xinyi Lai, Luojun Lin, Weijie Chen 0006, Yuanlong Yu 0001
ICCV2
2025 Assessing the Generalizability of Deep Models without Out-of-Distribution Data
abstract
Existing domain generalization (DG) methods typically rely on source domain data for model training and unknown target domain data to evaluate generalization ability. However, collecting target domain data is challenging because it must be out-of-distribution from the source domain data. Thus, developing Target-Free Evaluation (TFEval) methods is essential to assess the generalization performance of deep models without requiring unknown target domain data, relying only on trained data. In this paper, we estimate generalization performance for TFEval by training a regression predictor. The predictor takes model representations as input and predicts generalization performance. Inspired by the observation that Batch Normalization (BN) parameters strongly reflect domain/distribution information, we extract model representations by calculating BN parameter differences between the trained model and its proxy models. The proxy models are trained on a significantly smaller subset of the source domain data. We construct multiple TFEval datasets by collecting numerous deep models, and experimental results on these datasets demonstrate that our evaluation method effectively predicts the generalization performance of deep models.
Xiaojie Gan, Lingye Zhao, Luojun Lin
ICME4
2025 Synthetic-to-Real Camouflaged Object Detection
abstract
Due to the high cost of collection and labeling, there are relatively few datasets for camouflaged object detection (COD). In particular, for certain specialized categories, the available image dataset is insufficiently populated. Synthetic datasets can be utilized to alleviate the problem of limited data to some extent. However, directly training with synthetic datasets compared to real datasets can lead to a degradation in model performance. To tackle this problem, in this work, we investigate a new task, namely Syn-to-Real Camouflaged Object Detection (S2R-COD). In order to improve the model performance in real world scenarios, a set of annotated synthetic camouflaged images and a limited number of unannotated real images must be utilized. We propose the Cycling Syn-to-Real Domain Adaptation Framework (CSRDA), a method based on the student-teacher model. Specially, CSRDA propagates class information from the labeled source domain to the unlabeled target domain through pseudo labeling combined with consistency regularization. Considering that narrowing the intra-domain gap can improve the quality of pseudo labeling, CSRDA utilizes a recurrent learning framework to build an evolving real domain for bridging the source and target domain. Extensive experiments demonstrate the effectiveness of our framework, mitigating the problem of limited data and handcraft annotations in COD. Our code is publicly available at https://github.com/Muscape/S2R-COD.
Luojun Lin, Zheng Lin 0005
ACM Multimedia2
2025 Probabilistic Visual Prompt Tuning
Minghong Sun, Lingye Zhao, Luojun Lin
PRCV (6)3
2025 Supervised Contrastive Learning With Mixed Samples for Long-Tailed Recognition
abstract
In the domain of signal processing and deep learning, long-tailed data distributions present significant challenges due to the class imbalance in which a few classes contain a large number of samples, while most classes have far fewer. This imbalance hinders the ability of traditional models to effectively learn from minority classes. In this work, we focus on long-tailed supervised contrastive learning and introduce a novel approach termed Mixture-based Supervised Contrastive Learning (MixSCL), which integrates image mixing techniques into the supervised contrastive learning framework. By focusing on intra-class diversity and inter-class separability, our method aims to enhance the global uniformity of feature representations and improve model robustness. Specifically, MixSCL employs dual-stream projection heads designed to optimize separately for original and mixed samples, ensuring that the introduction of mixed samples does not distort the representations of original samples. We conduct extensive evaluations on bench mark datasets including CIFAR-100-LT and ImageNet-LT, which demonstrate that MixSCL achieves superior and more balanced performance in long-tailed scenarios.
Peihuan Song, Luojun Lin, Yuanlong Yu 0001, Wenjie Yang 0005, Qinghai Zheng
IEEE Signal Process. Lett.2
2025 Adapt Anything: Tailor Any Image Classifier Across Domains and Categories Using Text-to-Image Diffusion Models
abstract
We study a novel problem in this paper, that is, if a modern text-to-image diffusion model can tailor any image classifier across domains and categories. Existing domain adaption works exploit both source and target data for domain alignment so as to transfer the knowledge from the labeled source data to the unlabeled target data. However, as the development of text-to-image diffusion models, we wonder if the high-fidelity synthetic data can serve as a surrogate of the source data in real world. In this way, we do not need to collect and annotate the source data for each image classification task in a one-for-one manner. Instead, we utilize only one off-the-shelf text-to-image model to synthesize images with labels derived from text prompts, and then leverage them as a bridge to dig out the knowledge from the task-agnostic text-to-image generator to the task-oriented image classifier via domain adaptation. Such a one-for-all adaptation paradigm allows us to adapt anything in the world using only one text-to-image generator as well as any unlabeled target data. Extensive experiments validate the feasibility of this idea, which even surprisingly surpasses the state-of-the-art domain adaptation works using the source data collected and annotated in real world.
Weijie Chen 0006, Haoyu Wang 0016, Shicai Yang, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Luojun Lin, Di Xie, Yueting Zhuang
IEEE Trans. Big Data7
2024 MEAT: Median-Ensemble Adversarial Training for Improving Robustness and Generalization
abstract
Self-ensemble adversarial training methods improve model robustness by ensembling models at different training epochs, such as model weight averaging (WA). However, previous research has shown that self-ensemble defense methods in adversarial training (AT) still suffer from robust overfitting, which severely affects the generalization performance. Empirically, in the late phases of training, the AT becomes more overfitting to the extent that the individuals for weight averaging also suffer from overfitting and produce anomalous weight values, which causes the self-ensemble model to continue to undergo robust overfitting due to the failure in removing the weight anomalies. To solve this problem, we aim to tackle the influence of outliers in the weight space in this work and propose an easy-to-operate and effective Median-Ensemble Adversarial Training (MEAT) method to solve the robust overfitting phenomenon existing in self-ensemble defense from the source by searching for the median of the historical model weights. Experimental results show that MEAT achieves the best robustness against the powerful AutoAttack and can effectively allievate the robust overfitting. We further demonstrate that most defense methods can improve robust generalization and robustness by combining with MEAT.
Zhaozhe Hu, Jia-Li Yin, Bin Chen 0020, Luojun Lin, Ximeng Liu
ICASSP4
2024 No-Reference Segmentation Annotation Quality Assessment
abstract
Image segmentation tasks aim to separate the image into masks that represent different objects or regions, where deep-learning-based methods have become mainstream. In the common practice, researchers utilize large-scale datasets including images along with their annotations to train their models, and evaluate the predictions with evaluation metrics. However, to our knowledge, no metrics have been proposed to assess the quality of the segmentation annotations, which will bring benefits to both the labeling and experimental process. In this paper, we fill this research gap and propose the first no-reference segmentation annotation quality assessment named SAQ. Based on our observation, we utilize the normal gradients of pixels on the annotation contours to represent the degree of fitting the real contours, which reflect the annotation accuracy. To alleviate the image differences, we adopt the gradient ranking score rather than directly using the gradient value. The multi-scale strategy is introduced to accommodate annotations of objects with different structures. Extensive experiments on datasets for various segmentation tasks have demonstrated the rationality of our proposed SAQ, and the assessment results of their annotation quality can serve as significant references for researchers.
Zheng Lin 0005, Zheng-Peng Duan, Xuying Zhang, Luojun Lin
ICME4
2024 Slow-Fast Adaptation for Source-Free Object Detection
abstract
Unsupervised Domain Adaptive Object Detection (DAOD) task can relax the domain shift problem between source and target domains, which requires to train models on labeled source and unlabeled target domains jointly. However, due to limitations of data privacy protection, the source domain data is usually inaccessible, which poses significant challenges for the DAOD task. Hence, Source-Free Object Detection (SFOD) task has been developed that aims to fine-tune a pre-trained source model with only unlabeled target domain data. Most of the existing SFOD methods are based on pseudo labeling using the student-teacher framework, where the teacher model is the Exponential Moving Average (EMA) of the student models in different time steps. However, these methods always exist a knowledge bias problem due to class imbalance, and therefore, a fixed EMA update rate is no longer suitable for different classes. For high-quality classes, a fast EMA rate can accelerate knowledge updating and promote model convergence, while for low-quality classes, a fast EMA rate can accelerate the accumulation of knowledge bias and lead to the collapse of such categories. To solve this problem, we propose a novel SFOD method called Slow-Fast Adaptation which develops two different teacher models, a slow teacher, and a fast teacher model, to jointly guide the student training. The slow and fast teacher models can provide richer supervision information and complement each other. The experiments on four benchmark datasets show that our method achieves state-of-the-art results and even outperforms DAOD methods in some cases, which demonstrate the effectiveness of our method on the SFOD task.
Luojun Lin, Qipeng Liu 0004, Xiangwei Zheng 0003, Zheng Lin 0005
ICME1
2024 Learning feature alignment across attribute domains for improving facial beauty prediction
abstract
Facial beauty prediction (FBP) aims to develop a system to assess facial attractiveness automatically. Through prior research and our own observations, it has become evident that attribute information, such as gender and race, is a key factor leading to the distribution discrepancy in the FBP data. Such distribution discrepancy hinders current conventional FBP models from generalizing effectively to unseen attribute domain data, thereby discounting further performance improvement . To address this problem, in this paper, we exploit the attribute information to guide the training of convolutional neural networks (CNNs), with the final purpose of implicit feature alignment across various attribute domain data. To this end, we introduce the attribute information into convolution layer and batch normalization (BN) layer, respectively, as they are the most crucial parts for representation learning in CNNs. Specifically, our method includes: 1) Attribute-guided convolution (AgConv) that dynamically updates convolutional filters based on attributes by parameter tuning or parameter rebirth; 2) Attribute-guided batch normalization (AgBN) is developed to compute the attribute-specific statistics through an attribute guided batch sampling strategy; 3) To benefit from both approaches, we construct an integrated framework by combining AgConv and AgBN to achieve a more thorough feature alignment across different attribute domains. Extensive qualitative and quantitative experiments have been conducted on the SCUT-FBP, SCUT-FBP5500 and HotOrNot benchmark datasets. The results show that AgConv significantly improves the attribute-guided representation learning capacity and AgBN provides more stable optimization. Owing to the combination of AgConv and AgBN, the proposed framework (Ag-Net) achieves further performance improvement and is superior to other state-of-the-art approaches for FBP.
Zhishu Sun, Luojun Lin, Yuanlong Yu 0001
Expert Syst. Appl.2
2024 You only label once: A self-adaptive clustering-based method for source-free active domain adaptation
abstract
Abstract With the growing significance of data privacy protection, Source‐Free Domain Adaptation (SFDA) has gained attention as a research topic that aims to transfer knowledge from a labeled source domain to an unlabeled target domain without accessing source data. However, the absence of source data often leads to model collapse or restricts the performance improvements of SFDA methods, as there is insufficient true‐labeled knowledge for each category. To tackle this, Source‐Free Active Domain Adaptation (SFADA) has emerged as a new task that aims to improve SFDA by selecting a small set of informative target samples labeled by experts. Nevertheless, existing SFADA methods impose a significant burden on human labelers, requiring them to continuously label a substantial number of samples throughout the training period. In this paper, a novel approach is proposed to alleviate the labeling burden in SFADA by only necessitating the labeling of an extremely small number of samples on a one‐time basis . Moreover, considering the inherent sparsity of these selected samples in the target domain, a Self‐adaptive Clustering‐based Active Learning (SCAL) method is proposed that propagates the labels of selected samples to other datapoints within the same cluster. To further enhance the accuracy of SCAL, a self‐adaptive scale search method is devised that automatically determines the optimal clustering scale, using the entropy of the entire target dataset as a guiding criterion. The experimental evaluation presents compelling evidence of our method's supremacy. Specifically, it outstrips previous SFDA methods, delivering state‐of‐the‐art (SOTA) results on standard benchmarks. Remarkably, it accomplishes this with less than 0.5% annotation cost, in stark contrast to the approximate 5% required by earlier techniques. The approach thus not only sets new performance benchmarks but also offers a markedly more practical and cost‐effective solution for SFADA, making it an attractive choice for real‐world applications where labeling resources are limited.
Zhishu Sun, Luojun Lin, Yuanlong Yu 0001
IET Image Process.2
2024 Better Together: Data-Free Multi-Student Coevolved Distillation
Weijie Chen 0006, Yunyi Xuan, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang
Knowl. Based Syst.5
2024 Semi-supervised domain generalization with evolving intermediate domain
Luojun Lin, Zhishu Sun, Weijie Chen 0006, Wenxi Liu, Yuanlong Yu 0001, Lei Zhang 0038
Pattern Recognit.1
2023 FBPFormer: Dynamic Convolutional Transformer for Global-Local-Contexual Facial Beauty Prediction
Qipeng Liu 0004, Luojun Lin, Zhifeng Shen, Yuanlong Yu 0001
ICANN (10)2
2023 Customized Automatic Face Beautification
abstract
In the age of social media, posting attractive mugshots is commonplace, leading to an urgent need for automatic facial beautification techniques. To better meet the esthetic preferences of users, we devise a customized automatic face beautification task that can retouch the face adaptively to match the user-entered target score whilst preserving the ID information as much as possible. To accomplish this task, we propose a Human Esthetics Guided StyleGAN Inversion method to retouch each face in the embedding space using StyleGAN inversion. This process is guided by a pre-trained facial beauty prediction model that measures the difference between the target score and the predicted score of the retouched face. We conduct extensive experiments on various faces with different attributes, where the experimental results show that our method achieves the competitive performance, both in terms of visual effect and the proposed criterion.
Wang Chen 0005, Peizhen Chen, Weijie Chen 0006, Luojun Lin
ICASSP4
2023 Periodically Exchange Teacher-Student for Source-Free Object Detection
abstract
Source-free object detection (SFOD) aims to adapt the source detector to unlabeled target domain data in the absence of source domain data. Most SFOD methods follow the same self-training paradigm using mean-teacher (MT) framework where the student model is guided by only one single teacher model. However, such paradigm can easily fall into a training instability problem that when the teacher model collapses uncontrollably due to the domain shift, the student model also suffers drastic performance degradation. To address this issue, we propose the Periodically Exchange Teacher-Student (PETS) method, a simple yet novel approach that introduces a multiple-teacher framework consisting of a static teacher, a dynamic teacher, and a student model. During the training phase, we periodically exchange the weights between the static teacher and the student model. Then, we update the dynamic teacher using the moving average of the student model that has already been exchanged by the static teacher. In this way, the dynamic teacher can integrate knowledge from past periods, effectively reducing error accumulation and enabling a more stable training process within the MT-based framework. Further, we develop a consensus mechanism to merge the predictions of two teacher models to provide higher-quality pseudo labels for student model. Extensive experiments on multiple SFOD benchmarks show that the proposed method achieves state-of-the-art performance compared with other related methods, demonstrating the effectiveness and superiority of our method on SFOD task.
Qipeng Liu 0004, Luojun Lin, Zhifeng Shen, Zhifeng Yang
ICCV2
2023 Unsupervised Prompt Tuning for Text-Driven Object Detection
abstract
Grounded language-image pre-trained models have shown strong zero-shot generalization to various downstream object detection tasks. Despite their promising performance, the models rely heavily on the laborious prompt engineering. Existing works typically address this problem by tuning text prompts using downstream training data in a few-shot or fully supervised manner. However, a rarely studied problem is to optimize text prompts without using any annotations. In this paper, we delve into this problem and propose an Unsupervised Prompt Tuning framework for text-driven object detection, which is composed of two novel mean teaching mechanisms. In conventional mean teaching, the quality of pseudo boxes is expected to optimize better as the training goes on, but there is still a risk of overfitting noisy pseudo boxes. To mitigate this problem, 1) we propose Nested Mean Teaching, which adopts nested-annotation to supervise teacher-student mutual learning in a bi-level optimization manner; 2) we propose Dual Complementary Teaching, which employs an offline pre-trained teacher and an online mean teacher via data-augmentation-based complementary labeling so as to ensure learning without accumulating confirmation bias. By integrating these two mechanisms, the proposed unsupervised prompt tuning framework achieves significant performance improvement on extensive object detection datasets.
Weizhen He, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Donglian Qi, Yueting Zhuang
ICCV6
2023 Run and Chase: Towards Accurate Source-Free Domain Adaptive Object Detection
abstract
Recently, there has been increasing interest in the Source-Free Domain Adaptive Object Detection task, which involves training an object detector on the unlabeled target data using a pre-trained source model without accessing the source data. Most related methods are developed from the mean-teacher framework, which aims to train the student model closer to the teacher model via a pseudo labeling manner, where the teacher model is the exponential-moving-average of the student models at different time-steps. Following this line of works, we propose a Run-and-Chase Mutual-Learning method to strengthen the interactions between the student model and the teacher model in both feature and prediction levels. In our method, the student model is optimized to run away from the teacher model at the feature level, while chasing the teacher model at the prediction level. In this way, the student model is forced to be distinguishable at different time-steps, so that the teacher model can acquire more diverse task-related information and produce higher-accuracy pseudo labels. As the training goes, the student and teacher models are updated iteratively and promoted mutually, which can prevent the model collapse problem. Extensive experiments are conducted to validate the effectiveness of our method.
Luojun Lin, Zhifeng Yang, Qipeng Liu 0004, Yuanlong Yu 0001, Qifeng Lin
ICME1
2023 Adapt then Generalize: A Simple Two-Stage Framework for Semi-Supervised Domain Generalization
abstract
Semi-supervised domain generalization (SSDG) focuses on training a model on a single labeled source domain and several unlabeled source domains simultaneously for the purpose of generalizing to out-of-distribution domain. To prevent model overfitting on the unlabeled source domains and enhance model generalization, we propose an effective two-stage framework that disentangle the SSDG task into a burn-in stage and a mutual-training stage. In the burn-in stage, we train a domain adaptation model to reduce domain gap and produce high-accuracy pseudo labels for the unlabeled data. In the second stage, we devise two peer networks equipped with style randomization modules to take advantage of both the style and class information from the pseudo-labeled data. The peer networks are mutually teaching each other, which allows them to avoid overfitting to noisy data and ultimately improves generalization ability. Extensive experiments show that our method achieves state-of-the-art performance on various benchmarks, demonstrating the effectiveness of our method on SSDG.
Zhifeng Shen, Shicai Yang, Weijie Chen 0006, Luojun Lin
ICME5
2023 A Multiple Prediction Mechanisms Ensemble for Complex Remote Sensing Scenes
abstract
Facing complex remote sensing scenes, detection models with single detection mechanisms cannot always provide satisfactory detection capabilities. In order to obtain better detection performance in various remote sensing scenes, this paper constructs a novel ensemble model, namely: the multiple prediction mechanisms ensemble (MPME). In order to improve the feature representation ability and region recognition ability of the ensemble model, we build the ensemble of feature pyramids (EFP) and the ensemble of detection heads (EDH) respectively. In order to further improve the detection accuracy of the ensemble model, we propose a training strategy (k-Nearest Loss Learning), so that each sub-detector does not need to learn a trade-off among all training samples, and also reduces the possibility of model over-fitting. The experimental results show that our MPME is a more efficient and effective ensemble model. Compared with other ensemble models, our MPME has a faster detection speed and better detection accuracy. Compared with other state-of-the-art detectors, our detector also achieves superior detection performance.
Qifeng Lin, Luojun Lin, Yuanlong Yu 0001, Gang Fu 0003
ACM Multimedia2
2023 Parameter Exchange for Robust Dynamic Domain Generalization
abstract
Agnostic domain shift is the main reason of model degradation on the unknown target domains, which brings an urgent need to develop Domain generalization (DG). Recent advances at DG use dynamic networks to achieve training-free adaptation on the unknown target domains, termed Dynamic Domain Generalization (DDG), which compensates for the lack of self-adaptability in static models with fixed weights. The parameters of dynamic networks can be decoupled into a static and a dynamic component, which are designed to learn domain-invariant and domain-specific features, respectively. Based on the existing arts, in this work, we try to push the limits of DDG by disentangling the static and dynamic components more thoroughly from an optimization perspective. Our main consideration is that we can enable the static component to learn domain-invariant features more comprehensively by augmenting the domain-specific information. As a result, the more comprehensive domain-invariant features learned by the static component can then enforce the dynamic component to focus more on learning adaptive domain-specific features. To this end, we propose a simple yet effective Parameter Exchange (PE) method to perturb the combination between the static and dynamic components. We optimize the model using the gradients from both the perturbed and non-perturbed feed-forward jointly to implicitly achieve the aforementioned disentanglement. In this way, the two components can be optimized in a mutually-beneficial manner, which can resist the agnostic domain shifts and improve the self-adaptability on the unknown target domain. Extensive experiments show that PE can be easily plugged into existing dynamic networks to improve their generalization ability without bells and whistles.
Luojun Lin, Zhifeng Shen, Zhishu Sun, Yuanlong Yu 0001, Lei Zhang 0038, Weijie Chen 0006
ACM Multimedia1
2023 MetaFBP: Learning to Learn High-Order Predictor for Personalized Facial Beauty Prediction
abstract
Predicting individual aesthetic preferences holds significant practical applications and academic implications for human society. However, existing studies mainly focus on learning and predicting the commonality of facial attractiveness, with little attention given to Personalized Facial Beauty Prediction (PFBP). PFBP aims to develop a machine that can adapt to individual aesthetic preferences with only a few images rated by each user. In this paper, we formulate this task from a meta-learning perspective that each user corresponds to a meta-task. To address such PFBP task, we draw inspiration from the human aesthetic mechanism that visual aesthetics in society follows a Gaussian distribution, which motivates us to disentangle user preferences into a commonality and an individuality part. To this end, we propose a novel MetaFBP framework, in which we devise a universal feature extractor to capture the aesthetic commonality and then optimize to adapt the aesthetic individuality by shifting the decision boundary of the predictor via a meta-learning mechanism. Unlike conventional meta-learning methods that may struggle with slow adaptation or overfitting to tiny support sets, we propose a novel approach that optimizes a high-order predictor for fast adaptation. In order to validate the performance of the proposed method, we build several PFBP benchmarks by using existing facial beauty prediction datasets rated by numerous users. Extensive experiments on these benchmarks demonstrate the effectiveness of the proposed MetaFBP method.
Luojun Lin, Zhifeng Shen, Jia-Li Yin, Qipeng Liu 0004, Yuanlong Yu 0001, Weijie Chen 0006
ACM Multimedia1
2023 Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
abstract
Data-Free Knowledge Distillation (DFKD) has shown great potential in creating a compact student model while alleviating the dependency on real training data by synthesizing surrogate data. However, prior arts are seldom discussed under distribution shifts, which may be vulnerable in real-world applications. Recent Vision-Language Foundation Models, e.g., CLIP, have demonstrated remarkable performance in zero-shot out-of-distribution generalization, yet consuming heavy computation resources. In this paper, we discuss the extension of DFKD to Vision-Language Foundation Models without access to the billion-level image-text datasets. The objective is to customize a student model for distribution-agnostic downstream tasks with given category concepts, inheriting the out-of-distribution generalization capability from the pre-trained foundation models. In order to avoid generalization degradation, the primary challenge of this task lies in synthesizing diverse surrogate images driven by text prompts. Since not only category concepts but also style information are encoded in text prompts, we propose three novel Prompt Diversification methods to encourage image synthesis with diverse styles, namely Mix-Prompt, Random-Prompt, and Contrastive-Prompt. Experiments on out-of-distribution generalization datasets demonstrate the effectiveness of the proposed methods, with Contrastive-Prompt performing the best.
Yunyi Xuan, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang
ACM Multimedia5
2022 Slimmable Domain Adaptation
abstract
Vanilla unsupervised domain adaptation methods tend to optimize the model with fixed neural architecture, which is not very practical in real-world scenarios since the target data is usually processed by different resource-limited devices. It is therefore of great necessity to facilitate architecture adaptation across various devices. In this paper, we introduce a simple framework, Slimmable Domain Adaptation, to improve cross-domain generalization with a weight-sharing model bank, from which models of different capacities can be sampled to accommodate different accuracy-efficiency trade-offs. The main challenge in this frame-work lies in simultaneously boosting the adaptation performance of numerous models in the model bank. To tackle this problem, we develop a Stochastic EnsEmble Distillation method to fully exploit the complementary knowledge in the model bank for inter-model interaction. Nevertheless, considering the optimization conflict between inter-model interaction and intra-model adaptation, we augment the existing bi-classifier domain confusion architecture into an Optimization-Separated Tri-Classifier counterpart. After optimizing the model bank, architecture adaptation is leveraged via our proposed Unsupervised Performance Evaluation Metric. Under various resource constraints, our framework surpasses other competing approaches by a very large margin on multiple benchmarks. It is also worth emphasizing that our framework can preserve the performance improvement against the source-only model even when the computing complexity is reduced to 1/64. Code will be available at https://github.com/HIK-LAB/SlimDA.
Rang Meng, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Luojun Lin, Di Xie, Shiliang Pu, Xinchao Wang, Mingli Song, Yueting Zhuang
CVPR5
2022 Dynamic Domain Generalization
abstract
Domain generalization (DG) is a fundamental yet very challenging research topic in machine learning. The existing arts mainly focus on learning domain-invariant features with limited source domains in a static model. Unfortunately, there is a lack of training-free mechanism to adjust the model when generalized to the agnostic target domains. To tackle this problem, we develop a brand-new DG variant, namely Dynamic Domain Generalization (DDG), in which the model learns to twist the network parameters to adapt to the data from different domains. Specifically, we leverage a meta-adjuster to twist the network parameters based on the static model with respect to different data from different domains. In this way, the static model is optimized to learn domain-shared features, while the meta-adjuster is designed to learn domain-specific features. To enable this process, DomainMix is exploited to simulate data from diverse domains during teaching the meta-adjuster to adapt to the agnostic target domains. This learning mechanism urges the model to generalize to different agnostic target domains via adjusting the model without training. Extensive experiments demonstrate the effectiveness of our proposed method. Code is available: https://github.com/MetaVisionLab/DDG
Zhishu Sun, Zhifeng Shen, Luojun Lin, Yuanlong Yu 0001, Zhifeng Yang, Shicai Yang, Weijie Chen 0006
IJCAI3
2022 Self-Supervised Noisy Label Learning for Source-Free Unsupervised Domain Adaptation
abstract
Domain adaptation is an important property in robot vision, which enables the neural networks pre-trained on source domains to adapt target domains automatically without any annotation efforts. During this process, source data is not always accessible due to the constraints of expensive storage overhead and data privacy protection. Therefore, the source domain pre-trained model is expected to optimize with only unlabeled target data, termed as source-free unsupervised domain adaptation. In this paper, we view this problem as a special case of noisy label learning, since the given pre-trained model can generate noisy labels for unlabeled target data via network inference. The potential semantic cues for unsupervised domain adaptation exactly lie on these noisy labels. Inspired by this problem modeling, we propose a simple yet effective Self-Supervised Noisy Label Learning method, which injects self-supervised learning to impose the intrinsic data structure and facilitate label-denoising. Extensive experiments have been conducted on diverse benchmarks to validate the effectiveness. Our method achieves state-of-the-art performance.
Weijie Chen 0006, Luojun Lin, Shicai Yang, Di Xie, Shiliang Pu, Yueting Zhuang
IROS2
2022 SynSig2Vec: Forgery-Free Learning of Dynamic Signature Representations by Sigma Lognormal-Based Synthesis and 1D CNN
abstract
Handwritten signature verification is a challenging task because signatures of a writer may be skillfully imitated by a forger. As skilled forgeries are generally difficult to acquire for training, in this paper, we propose a deep learning-based dynamic signature verification framework, SynSig2Vec, to address the skilled forgery attack without training with any skilled forgeries. Specifically, SynSig2Vec consists of a novel learning-by-synthesis method for training and a 1D convolutional neural network model, called Sig2Vec, for signature representation extraction. The learning-by-synthesis method first applies the Sigma Lognormal model to synthesize signatures with different distortion levels for genuine template signatures, and then learns to rank these synthesized samples in a learnable representation space based on average precision optimization. The representation space is achieved by the proposed Sig2Vec model, which is designed to extract fixed-length representations from dynamic signatures of arbitrary lengths. Through this training method, the Sig2Vec model can extract extremely effective signature representations for verification. Our SynSig2Vec framework requires only genuine signatures for training, yet achieves state-of-the-art performance on the largest dynamic signature database to date, DeepSignDB, in both skilled forgery and random forgery scenarios. Source codes of SynSig2Vec will be available at https://github.com/LaiSongxuan/SynSig2Vec.
Songxuan Lai, Yecheng Zhu, Zhe Li 0046, Luojun Lin
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Regression Guided by Relative Ranking Using Convolutional Neural Network (R$^3$3CNN) for Facial Beauty Prediction
abstract
Facial beauty prediction (FBP) aims to automatically assess facial attractiveness consistently with judgements based on human perception. Most of previous methods formulate FBP as a classification, regression or ranking problem of machine learning. However, humans not only represent facial attractiveness as a score, but also perceive the relative aesthetics of faces. Inspired by this observation, we formulate FBP as a specific regression problem guided by ranking information. Specifically, we propose a general CNN architecture, called R$^3$CNN, to integrate the relative ranking of faces in terms of aesthetics to improve performance of FBP. As R$^3$CNN consists of both regression and ranking components, it is challenging to train and fine-tune it by existing techniques. To tackle this problem, we propose the following learning schemes for R$^3$CNN: 1) a hard pair sampling strategy that generates challenging-to-predicted image pairs and pseudo ranking labels from true rating scores; 2) an assemble loss function that combines regression loss and pairwise ranking loss (PR-Loss), where PR-Loss can be a hinge-form loss or a log-sum-exp pairwise loss; 3) a cascaded fine-tuning method that further improves prediction. Moreover, we build a benchmark dataset, called SCUT-FBP5500, containing 5,500 facial images with diverse properties (male/female, Asian/Caucasian, ages) and labels (face landmarks, rating scores within [1, 5], rating score distribution). Experiments were performed on both the SCUT-FBP and the SCUT-FBP5500 benchmark datasets, where our method achieves state-of-the-art performance on different evaluation settings. Comparisons with related CNN models highlight the effectiveness of the R$^3$CNN architecture for FBP.
Luojun Lin, Lingyu Liang
IEEE Trans. Affect. Comput.1
2021 Skeleton-based action recognition using sparse spatio-temporal GCN with edge effective resistance
Tasweer Ahmad, Luojun Lin, Guozhi Tang
Neurocomputing3
2020 SynSig2Vec: Learning Representations from Synthetic Dynamic Signatures for Real-World Verification
abstract
An open research problem in automatic signature verification is the skilled forgery attacks. However, the skilled forgeries are very difficult to acquire for representation learning. To tackle this issue, this paper proposes to learn dynamic signature representations through ranking synthesized signatures. First, a neuromotor inspired signature synthesis method is proposed to synthesize signatures with different distortion levels for any template signature. Then, given the templates, we construct a lightweight one-dimensional convolutional network to learn to rank the synthesized samples, and directly optimize the average precision of the ranking to exploit relative and fine-grained signature similarities. Finally, after training, fixed-length representations can be extracted from dynamic signatures of variable lengths for verification. One highlight of our method is that it requires neither skilled nor random forgeries for training, yet it surpasses the state-of-the-art by a large margin on two public benchmarks.
Songxuan Lai, Luojun Lin, Yecheng Zhu, Huiyun Mao
AAAI3
2019 Attribute-Aware Convolutional Neural Networks for Facial Beauty Prediction
abstract
Facial beauty prediction (FBP) aims to develop a machine that automatically makes facial attractiveness assessment. To a large extent, the perception of facial beauty for a human is involved with the attributes of facial appearance, which provides some significant visual cues for FBP. Deep convolution neural networks (CNNs) have shown its power for FBP, but convolution filters with fixed parameters cannot take full advantage of the facial attributes for FBP. To address this problem, we propose an Attribute-aware Convolutional Neural Network (AaNet) that modulates the filters of the main network, adaptively, using parameter generators that take beauty-related attributes as extra inputs. The parameter generators update the filters in the main network in two different manners: filter tuning or filter rebirth. However, AaNet takes attributes information as prior knowledge, that is ill-suited to those datasets merely with task-oriented labels. Therefore, imitating the design of AaNet, we further propose a Pseudo Attribute-aware Convolutional Neural Network (P-AaNet) that modulates filters conditioned on global context embeddings (pseudo attributes) of input faces learnt by a lightweight pseudo attribute distiller. Extensive ablation studies show that the AaNet and P-AaNet improve the performance of FBP when compared to conventional convolution and attention scheme, which validates the effectiveness of our method.
Luojun Lin, Lingyu Liang, Weijie Chen 0006
IJCAI1
2018 SCUT-FBP5500: A Diverse Benchmark Dataset for Multi-Paradigm Facial Beauty Prediction
abstract
Facial beauty prediction (FBP) is a significant visual recognition problem to make assessment of facial attractiveness that is consistent to human perception. To tackle this problem, various data-driven models, especially state-of-the-art deep learning techniques, were introduced, and benchmark dataset become one of the essential elements to achieve FBP. Previous works have formulated the recognition of facial beauty as a specific supervised learning problem of classification, regression or ranking, which indicates that FBP is intrinsically a computation problem with multiple paradigms. However, most of FBP benchmark datasets were built under specific computation constrains, which limits the performance and flexibility of the computational model trained on the dataset. In this paper, we argue that FBP is a multi-paradigm computation problem, and propose a new diverse benchmark dataset, called SCUT-FBP5500, to achieve multi-paradigm facial beauty prediction. The SCUT-FBP5500 dataset has totally 5500 frontal faces with diverse properties (male/female, Asian/Caucasian, ages) and diverse labels (face landmarks, beauty scores within [1], [5], beauty score distribution), which allows different computational models with different FBP paradigms, such as appearance-based/shape-based facial beauty classification/regression model for male/female of Asian/Caucasian. We evaluated the SCUT-FBP5500 dataset for FBP using different combinations of feature and predictor, and various deep learning methods. The results indicates the improvement of FBP and the potential applications based on the SCUT-FBP5500.
Lingyu Liang, Luojun Lin, Duorui Xie, Mengru Li
ICPR2
2018 R2-ResNeXt: A ResNeXt-Based Regression Model with Relative Ranking for Facial Beauty Prediction
abstract
The purpose of facial beauty prediction (FBP) is to develop a machine that automatically evaluates facial attractiveness in a human perceptual manner. One of the essential problem of facial beauty prediction is the discriminative facial representation of the prediction model. Previous methods formulate FBP as a specific supervised learning of classification, regression, or ranking. We find that the relative ranking information is useful to improve the regression model of FBP. Based on this observation, this paper proposes a regression model guided by the relative ranking with the state-of-the-art Res NeXt structure to achieve FBP, and we call the model as R2-ResNeXt. The R2-ResNeXt facilitates to learn the representation and predictor guided by relative ranking for facial attractiveness assessment in an end-to-end manner. To train the R2-ResNeXt, we develop an aggregated loss that combines regression loss and pairwise ranking loss linearly. We also design a method to construct a dataset containing relatively -labelled image pairs whose individual images are sampled from the SCUT-FBP benchmark database. The experimental results on the SCUT-FBP benchmark show that our R2-ResNeXt achieves the state-of-the-art performance compared with related literatures, and further indicates the effectiveness of the deep residual architecture and relative beauty ranking into regression task for facial beauty prediction.
Luojun Lin, Lingyu Liang
ICPR1
2017 Region-aware scattering convolution networks for facial beauty prediction
abstract
This paper proposes a scattering convolutional network with region-aware facial attributes to obtain a mid-level representation for facial beauty prediction (FBP). Different from the previous works that only focus on the discriminative representation for prediction, this paper also considers the invariant properties of the facial representation that reduces the variances caused by the image transformations, such as rotations and translation. The proposed region-aware scattering convolution network (RegionScatNet) is based on a deep convolution network of scattering transforms (ScatNet) integrated with facial texture and shape features. It consists of three components, including: 1) Region Extraction to obtain the significant facial perception region with implicit shape features using region-aware mask; 2) Attributes Decomposition to separate the extracted region into detail and structure facial layers by a guided filter; 3) Scattering Convolution that computes the roto-translation invariant representation of facial detail and structure for FBP by cascading three-layer wavelets filters and non-linear modulus pooling. The comparisons with related deep learning-based methods illustrate the effectiveness of RegionScatNet for FBP. The evaluations with various prediction model (like SVR and Gaussian process) and with different facial variances (like rotaion) indicate the robustness of the RegionScatNet-based features.
Lingyu Liang, Duorui Xie, Jie Xu 0041, Mengru Li, Luojun Lin
ICIP6