VLDB 2026 Research / reviewers in the wild / expert
Xin-Chun Li
dblp:246/2947
· DBLP profile ↗
22ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0001-9417-7971ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring dark knowledge under various teacher capacities and addressing capacity mismatch
Wen-Shu Fan, Xin-Chun Li, De-Chuan Zhan |
Frontiers Comput. Sci. | 2 |
| 2025 | Maximizing the Effectiveness of Larger BERT Models for CompressionabstractKnowledge distillation (KD) is a widely used approach for BERT compression, where a larger BERT model serves as a teacher to transfer knowledge to a smaller student model.Prior works have found that distilling a larger BERT with superior performance may degrade student's performance than a smaller BERT.In this paper, we investigate the limitations of existing KD methods for larger BERT models.Through Canonical Correlation Analysis, we identify that these methods fail to fully exploit the potential advantages of larger teachers.To address this, we propose an improved distillation approach that effectively enhances knowledge transfer.Comprehensive experiments demonstrate the effectiveness of our method in enabling larger BERT models to distill knowledge more efficiently. Wen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li, De-Chuan Zhan |
ACL (1) | 4 |
| 2024 | Weight Scope Alignment: A Frustratingly Easy Method for Model MergingabstractMerging models becomes a fundamental procedure in some applications that consider model efficiency and robustness. The training randomness or Non-I.I.D. data poses a huge challenge for averaging-based model fusion. Previous research efforts focus on element-wise regularization or neural permutations to enhance model averaging while overlooking weight scope variations among models, which can significantly affect merging effectiveness. In this paper, we reveal variations in weight scope under different training conditions, shedding light on its influence on model merging. Fortunately, the parameters in each layer basically follow the Gaussian distribution, which inspires a novel and simple regularization approach named Weight Scope Alignment (WSA). It contains two key components: 1) leveraging a target weight scope to guide the model training process for ensuring weight scope matching in the subsequent model merging. 2) fusing the weight scope of two or more models into a unified one for multi-stage model fusion. We extend the WSA regularization to two different scenarios, including Mode Connectivity and Federated Learning. Abundant experimental studies validate the effectiveness of our approach. Yichu Xu, Xin-Chun Li, Le Gan, De-Chuan Zhan |
ECAI | 2 |
| 2024 | CLAF: Contrastive Learning with Augmented Features for Imbalanced Semi-Supervised LearningabstractDue to the advantages of leveraging unlabeled data and learning meaningful representations, semi-supervised learning and contrastive learning have been progressively combined to achieve better performances in popular applications with few labeled data and abundant unlabeled data. One common manner is assigning pseudo-labels to unlabeled samples and selecting positive and negative samples from pseudo-labeled samples to apply contrastive learning. However, the real-world data may be imbalanced, causing pseudo-labels to be biased toward the majority classes and further undermining the effectiveness of contrastive learning. To address the challenge, we propose Contrastive Learning with Augmented Features (CLAF). We design a class-dependent feature augmentation module to alleviate the scarcity of minority class samples in contrastive learning. For each pseudo-labeled sample, we select positive and negative samples from labeled data instead of unlabeled data to compute contrastive loss. Comprehensive experiments on imbalanced image classification datasets demonstrate the effectiveness of CLAF in the context of imbalanced semi-supervised learning. Bowen Tao, Lan Li 0001, Xin-Chun Li, De-Chuan Zhan |
ICASSP | 3 |
| 2024 | Revisit the Essence of Distilling Knowledge through CalibrationabstractKnowledge Distillation (KD) has evolved into a practical technology for transferring knowledge from a well-performing model (teacher) to a weak model (student). A counter-intuitive phenomenon known as capacity mismatch has been identified, wherein KD performance may not be good when a better teacher instructs the student. Various preliminary methods have been proposed to alleviate capacity mismatch, but a unifying explanation for its cause remains lacking. In this paper, we propose a unifying analytical framework to pinpoint the core of capacity mismatch based on calibration. Through extensive analytical experiments, we observe a positive correlation between the calibration of the teacher model and the KD performance with original KD methods. As this correlation arises due to the sensitivity of metrics (e.g., KL divergence) to calibration, we recommend employing measurements insensitive to calibration such as ranking-based loss. Our experiments demonstrate that ranking-based loss can effectively replace KL divergence, aiding large models with poor calibration to teach better. Wen-Shu Fan, Su Lu, Xin-Chun Li, De-Chuan Zhan, Le Gan |
ICML | 3 |
| 2024 | Enhancing Class-Imbalanced Learning with Pre-Trained Guidance through Class-Conditional Knowledge DistillationabstractIn class-imbalanced learning, the scarcity of information about minority classes presents challenges in obtaining generalizable features for these classes. Leveraging large-scale pre-trained models with powerful generalization capabilities as teacher models can help fill this information gap. Traditional knowledge distillation transfers the label distribution $p(\boldsymbol{y}|\boldsymbol{x})$ predicted by the teacher model to the student model. However, this method falls short on imbalanced data as it fails to capture the class-conditional probability distribution $p(\boldsymbol{x}|\boldsymbol{y})$ from the teacher model, which is crucial for enhancing generalization. To overcome this, we propose Class-Conditional Knowledge Distillation (CCKD), a novel approach that enables learning of the teacher model’s class-conditional probability distribution during the distillation process. Additionally, we introduce Augmented CCKD (ACCKD), which involves distillation on a constructed class-balanced dataset (formed through data mixing) and feature imitation on the entire dataset to further facilitate the learning of features. Experimental results on various imbalanced datasets demonstrate an average accuracy improvement of 7.4% using our method. Lan Li 0001, Xin-Chun Li, Han-Jia Ye, De-Chuan Zhan |
ICML | 2 |
| 2024 | MLI Formula: A Nearly Scale-Invariant Solution with Noise PerturbationabstractMonotonic Linear Interpolation (MLI) refers to the peculiar phenomenon that the error between the initial and converged model monotonically decreases along the linear interpolation, i.e., $(1-\alpha)\boldsymbol{\theta}_0 + \alpha \boldsymbol{\theta}_F$. Previous works focus on paired initial and converged points, relating MLI to the smoothness of the optimization trajectory. In this paper, we find a shocking fact that the error curves still exhibit a monotonic decrease when $\boldsymbol{\theta}_0$ is replaced with noise or even zero values, implying that the decreasing curve may be primarily related to the property of the converged model rather than the optimization trajectory. We further explore the relationship between $\alpha\boldsymbol{\theta}_F$ and $\boldsymbol{\theta}_F$ and propose scale invariance properties in various cases, including Generalized Scale Invariance (GSI), Rectified Scale Invariance (RSI), and Normalized Scale Invariance (NSI). From an inverse perspective, the MLI formula is essentially an equation that adds varying levels of noise (i.e., $(1-\alpha)\boldsymbol{\epsilon}$) to a nearly scale-invariant network (i.e., $\alpha \boldsymbol{\theta}_F$), resulting in a monotonically increasing error as the noise level rises. MLI is a special case where $\boldsymbol{\epsilon}$ is equal to $\boldsymbol{\theta}_0$. Bowen Tao, Xin-Chun Li, De-Chuan Zhan |
ICML | 2 |
| 2024 | Exploring and Exploiting the Asymmetric Valley of Deep Neural NetworksabstractExploring the loss landscape offers insights into the inherent principles of deep neural networks (DNNs). Recent work suggests an additional asymmetry of the valley beyond the flat and sharp ones, yet without thoroughly examining its causes or implications. Our study methodically explores the factors affecting the symmetry of DNN valleys, encompassing (1) the dataset, network architecture, initialization, and hyperparameters that influence the convergence point; and (2) the magnitude and direction of the noise for 1D visualization. Our major observation shows that the {\it degree of sign consistency} between the noise and the convergence point is a critical indicator of valley symmetry. Theoretical insights from the aspects of ReLU activation and softmax function could explain the interesting phenomenon. Our discovery propels novel understanding and applications in the scenario of Model Fusion: (1) the efficacy of interpolating separate models significantly correlates with their sign consistency ratio, and (2) imposing sign alignment during federated learning emerges as an innovative approach for model parameter alignment. Xin-Chun Li, Jin-Lin Tang, De-Chuan Zhan |
NeurIPS | 1 |
| 2024 | Aligning model outputs for class imbalanced non-IID federated learning
Lan Li 0001, De-Chuan Zhan, Xin-Chun Li |
Mach. Learn. | 3 |
| 2024 | MAP: Model Aggregation and Personalization in Federated Learning With Incomplete ClassesabstractIn some real-world applications, data samples are usually distributed on local devices, where federated learning (FL) techniques are proposed to coordinate decentralized clients without directly sharing users’ private data. FL commonly follows the parameter server architecture and contains multiple personalization and aggregation procedures. The natural data heterogeneity across clients, i.e., Non-I.I.D. data, challenges both the aggregation and personalization goals in FL. In this paper, we focus on a special kind of Non-I.I.D. scene where clients own incomplete classes, i.e., each client can only access a partial set of the whole class set. The server aims to aggregate a complete classification model that could generalize to all classes, while the clients are inclined to improve the performance of distinguishing their observed classes. For better model aggregation, we point out that the standard softmax will encounter several problems caused by missing classes and propose “restricted softmax” as an alternative. For better model personalization, we point out that the hard-won personalized models are not well exploited and propose “inherited private model” to store the personalization experience. Our proposed algorithm named MAP could simultaneously achieve the aggregation and personalization goals in FL. Abundant experimental studies verify the superiorities of our algorithm. Xin-Chun Li, Shaoming Song, Yinchuan Li, Bingshuai Li, Yunfeng Shao 0001, Yang Yang 0074, De-Chuan Zhan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | MrTF: model refinery for transductive federated learning
Xin-Chun Li, Yang Yang 0074, De-Chuan Zhan |
Data Min. Knowl. Discov. | 1 |
| 2022 | Federated Learning with Position-Aware NeuronsabstractFederated Learning (FL) fuses collaborative models from local nodes without centralizing users' data. The permutation invariance property of neural networks and the non-i.i.d. data across clients make the locally updated parameters imprecisely aligned, disabling the coordinate-based parameter averaging. Traditional neurons do not explicitly consider position information. Hence, we propose Position-Aware Neurons (PANs) as an alternative, fusing position-related values (i.e., position encodings) into neuron outputs. PANs couple themselves to their positions and minimize the possibility of dislocation, even updating on heterogeneous data. We turn on/off PANs to disable/enable the permutation invariance property of neural networks. PANs are tightly coupled with positions when applied to FL, making parameters across clients pre-aligned and facilitating coordinate-based parameter averaging. PANs are algorithm-agnostic and could universally improve existing FL algorithms. Furthermore, “FL with PANs” is simple to implement and computationally friendly. Xin-Chun Li, Yichu Xu, Shaoming Song, Bingshuai Li, Yinchuan Li, Yunfeng Shao 0001, De-Chuan Zhan |
CVPR | 1 |
| 2022 | Exploring Transferability Measures and Domain Selection in Cross-Domain Slot FillingabstractAs an essential task for natural language understanding, slot filling aims to identify the contiguous spans of specific slots in an utterance. In real-world applications, the labeling costs of utterances may be expensive, and transfer learning techniques have been developed to ease this problem. However, cross-domain slot filling could significantly suffer from negative transfer due to non-targeted or zero-shot slots. Originally, this paper explores several ways to measure transferability across slot filling domains and finds that the shared slot number could serve as an efficient and effective estimator. First, this frustratingly easy measure requires no training data and is efficient to calculate. Second, it guides us heuristically select source domains that contain more shared slots with the target domain, which obtains SOTA results on Snips benchmark. Third, a dynamic transfer procedure based on this estimator clearly shows the negative transfer in cross-domain slot filling. We finally explore a source-free scene that we could only obtain black-box source models and propose to weight source domains based on prediction entropy. Xin-Chun Li, Yan-Jia Wang, Le Gan, De-Chuan Zhan |
ICASSP | 1 |
| 2022 | Avoid Overfitting User Specific Information in Federated Keyword SpottingabstractKeyword spotting (KWS) aims to discriminate a specific wakeup word from other signals precisely and efficiently for different users.Recent works utilize various deep networks to train KWS models with all users' speech data centralized without considering data privacy.Federated KWS (FedKWS) could serve as a solution without directly sharing users' data.However, the small amount of data, different user habits, and various accents could lead to fatal problems, e.g., overfitting or weight divergence.Hence, we propose several strategies to encourage the model not to overfit user-specific information in FedKWS.Specifically, we first propose an adversarial learning strategy, which updates the downloaded global model against an overfitted local model and explicitly encourages the global model to capture user-invariant information.Furthermore, we propose an adaptive local training strategy, letting clients with more training data and more uniform class distributions undertake more local update steps.Equivalently, this strategy could weaken the negative impacts of those users whose data is less qualified.Our proposed FedKWS-UI could explicitly and implicitly learn user-invariant information in FedKWS.Abundant experimental results on federated Google Speech Commands verify the effectiveness of FedKWS-UI. Xin-Chun Li, Jin-Lin Tang, Shaoming Song, Bingshuai Li, Yinchuan Li, Yunfeng Shao 0001, Le Gan, De-Chuan Zhan |
INTERSPEECH | 1 |
| 2022 | Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainabstractKnowledge Distillation (KD) aims at transferring the knowledge of a well-performed neural network (the {\it teacher}) to a weaker one (the {\it student}). A peculiar phenomenon is that a more accurate model doesn't necessarily teach better, and temperature adjustment can neither alleviate the mismatched capacity. To explain this, we decompose the efficacy of KD into three parts: {\it correct guidance}, {\it smooth regularization}, and {\it class discriminability}. The last term describes the distinctness of {\it wrong class probabilities} that the teacher provides in KD. Complex teachers tend to be over-confident and traditional temperature scaling limits the efficacy of {\it class discriminability}, resulting in less discriminative wrong class probabilities. Therefore, we propose {\it Asymmetric Temperature Scaling (ATS)}, which separately applies a higher/lower temperature to the correct/wrong class. ATS enlarges the variance of wrong class probabilities in the teacher's label and makes the students grasp the absolute affinities of wrong classes to the target class as discriminative as possible. Both theoretical analysis and extensive experimental results demonstrate the effectiveness of ATS. The demo developed in Mindspore is available at \url{https://gitee.com/lxcnju/ats-mindspore} and will be available at \url{https://gitee.com/mindspore/models/tree/master/research/cv/ats}. Xin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li, Bingshuai Li, Yunfeng Shao 0001, De-Chuan Zhan |
NeurIPS | 1 |
| 2021 | Task Cooperation for Semi-Supervised Few-Shot LearningabstractTraining a model with limited data is an essential task for machine learning and visual recognition. Few-shot learning approaches meta-learn a task-level inductive bias from SEEN class few-shot tasks, and the meta-model is expected to facilitate the few-shot learning with UNSEEN classes. Inspired by the idea that unlabeled data can be utilized to smooth the model space in traditional semi-supervised learning, we propose TAsk COoperation (TACO) which takes advantage of unsupervised tasks to smooth the meta-model space. Specifically, we couple the labeled support set in a few-shot task with easily-collected unlabeled instances, prediction agreement on which encodes the relationship between tasks. The learned smooth meta-model promotes the generalization ability on supervised UNSEEN few-shot tasks. The state-of-the-art few-shot classification results on MiniImageNet and TieredImageNet verify the superiority of TACO to leverage unlabeled data and task relationship in meta-learning. Han-Jia Ye, Xin-Chun Li, De-Chuan Zhan |
AAAI | 2 |
| 2021 | FedRS: Federated Learning with Restricted Softmax for Label Distribution Non-IID DataabstractFederated Learning (FL) aims to generate a global shared model via collaborating decentralized clients with privacy considerations. Unlike standard distributed optimization, FL takes multiple optimization steps on local clients and then aggregates the model updates via a parameter server. Although this significantly reduces communication costs, the non-iid property across heterogeneous devices could make the local update diverge a lot, posing a fundamental challenge to aggregation. In this paper, we focus on a special kind of non-iid scene, i.e., label distribution skew, where each client can only access a partial set of the whole class set. Considering top layers of neural networks are more task-specific, we advocate that the last classification layer is more vulnerable to the shift of label distribution. Hence, we in-depth study the classifier layer and point out that the standard softmax will encounter several problems caused by missing classes. As an alternative, we propose "Restricted Softmax" to limit the update of missing classes' weights during the local procedure. Our proposed FedRS is very easy to implement with only a few lines of code. We investigate our methods on both public datasets and a real-world service awareness application. Abundant experimental results verify the superiorities of our methods. Xin-Chun Li, De-Chuan Zhan |
KDD | 1 |
| 2021 | FedPHP: Federated Personalization with Inherited Private Models
Xin-Chun Li, De-Chuan Zhan, Yunfeng Shao 0001, Bingshuai Li, Shaoming Song |
ECML/PKDD (1) | 1 |
| 2021 | Deep multiple instance selection
Xin-Chun Li, De-Chuan Zhan, Jia-Qi Yang 0001 |
Sci. China Inf. Sci. | 1 |
| 2020 | Towards Understanding Transfer Learning Algorithms Using Meta Transfer Features
Xin-Chun Li, De-Chuan Zhan, Jia-Qi Yang 0001, Cheng Hang, Yi Lu 0007 |
PAKDD (2) | 1 |
| 2020 | Bottom-Up and Top-Down Graph Pooling
Jia-Qi Yang 0001, De-Chuan Zhan, Xin-Chun Li |
PAKDD (2) | 3 |
| 2019 | Automatic Successive Reinforcement Learning with Multiple Auxiliary RewardsabstractReinforcement learning has played an important role in decision making related applications, e.g., robotics motion, self-driving, recommendation, etc. The reward function, as a crucial component, affects the efficiency and effectiveness of reinforcement learning to a large extent. In this paper, we focus on the investigation of reinforcement learning with more than one auxiliary reward. It is found that different auxiliary rewards can boost up the learning rate and effectiveness in different stages, and consequently we propose the Automatic Successive Reinforcement Learning (ASR) for auxiliary rewards grading selection for efficient reinforcement learning by stages. Experiments and simulations have shown the superiority of our proposed ASR on a range of environments, including OpenAI classical control domains and video games; Freeway and Catcher. Zhao-Yang Fu, De-Chuan Zhan, Xin-Chun Li, Yi-Xing Lu |
IJCAI | 3 |