EDBT 2026 Demo / reviewers in the wild / expert
Le Gan
dblp:199/0588
· DBLP profile ↗
20ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0002-8260-6932ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningabstractClass-Incremental Learning (CIL) requires models to continually acquire knowledge of new classes without forgetting old ones. Despite Pre-trained Models (PTMs) have shown excellent performance in CIL, catastrophic forgetting still occurs as the model learns new concepts. Existing work seeks to utilize lightweight components to adjust the PTM, while the forgetting phenomenon still comes from parameter and retrieval levels. Specifically, iterative updates of the model result in parameter drift, while mistakenly retrieving irrelevant modules leads to the mismatch during inference. To this end, we propose MOdel Surgery (MOS) to rescue the model from forgetting previous knowledge. By training task-specific adapters, we continually adjust the PTM to downstream tasks. To mitigate parameter-level forgetting, we present an adapter merging approach to learn task-specific adapters, which aims to bridge the gap between different components while reserve task-specific information. Besides, to address retrieval-level forgetting, we introduce a training-free self-refined adapter retrieval mechanism during inference, which leverages the model's inherent ability for better adapter retrieval. By jointly rectifying the model with those steps, MOS can robustly resist catastrophic forgetting in the learning process. Extensive experiments on seven benchmark datasets validate MOS's state-of-the-art performance. Hai-Long Sun, Da-Wei Zhou 0001, Hanbin Zhao, Le Gan, De-Chuan Zhan, Han-Jia Ye |
AAAI | 4 |
| 2025 | FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision MakingabstractFoundation Models (FMs) and World Models (WMs) offer complementary strengths in task generalization at different levels. In this work, we propose FOUNDER, a framework that integrates the generalizable knowledge embedded in FMs with the dynamic modeling capabilities of WMs to enable open-ended task solving in embodied environments in a reward-free manner. We learn a mapping function that grounds FM representations in the WM state space, effectively inferring the agent’s physical states in the world simulator from external observations. This mapping enables the learning of a goal-conditioned policy through imagination during behavior learning, with the mapped task serving as the goal state. Our method leverages the predicted temporal distance to the goal state as an informative reward signal. FOUNDER demonstrates superior performance on various multi-task offline visual control benchmarks, excelling in capturing the deep-level semantics of tasks specified by text or videos, particularly in scenarios involving complex observations or domain gaps where prior methods struggle. The consistency of our learned reward function with the ground-truth reward is also empirically validated. Our project website is https://sites.google.com/view/founder-rl. Yucen Wang, Shenghua Wan, Le Gan, De-Chuan Zhan |
ICML | 4 |
| 2025 | Reward Models in Deep Reinforcement Learning: A SurveyabstractIn reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are introduced as proxies of the desired objectives, such that when the agent maximizes the accumulated reward, it also fulfills the task designer's intentions. Recently, significant attention from both academic and industrial researchers has focused on developing reward models that not only align closely with the true objectives but also facilitate policy optimization. In this survey, we provide a comprehensive review of reward modeling techniques within the RL literature. We begin by outlining the background and preliminaries in reward modeling. Next, we present an overview of recent reward modeling approaches, categorizing them based on the source, the mechanism, and the reward learning paradigm. Building on this understanding, we discuss various applications of these reward modeling techniques and review methods for evaluating reward models. Finally, we conclude by highlighting promising research directions in reward modeling. Altogether, this survey includes both established and emerging methods, filling the vacancy of a systematic review of reward models in current literature. Shenghua Wan, Yucen Wang, Chenxiao Gao, Le Gan, Zongzhang Zhang, De-Chuan Zhan |
IJCAI | 5 |
| 2025 | Leveraging Conditional Dependence for Efficient World Model DenoisingabstractEffective denoising is critical for managing complex visual inputs contaminated with noisy distractors in model-based reinforcement learning (RL). Current methods often oversimplify the decomposition of observations by neglecting the conditional dependence between task-relevant and task-irrelevant components given an observation. To address this limitation, we introduce CsDreamer, a model-based RL approach built upon the world model of Collider-structure Recurrent State-Space Model (CsRSSM). CsRSSM incorporates colliders to comprehensively model the denoising inference process and explicitly capture the conditional dependence. Furthermore, it employs a decoupling regularization to balance the influence of this conditional dependence. By accurately inferring a task-relevant state space, CsDreamer improves learning efficiency during rollouts. Experimental results demonstrate the effectiveness of CsRSSM in extracting task-relevant information, leading to CsDreamer outperforming existing approaches in environments characterized by complex noise interference. Shaowei Zhang 0001, Jiahan Cao, Dian Cheng, Xunlan Zhou, Shenghua Wan, Le Gan, De-Chuan Zhan |
NeurIPS | 6 |
| 2025 | Generalized Conditional Similarity Learning via Semantic MatchingabstractThe inherent complexity of image semantics engenders a fascinating variability in relationships between images. For instance, under a certain condition, two images may demonstrate similarity, while under different circumstances, the same pair could exhibit absolute dissimilarity. A singular feature space is therefore insufficient for capturing the nuanced semantic relationships that exist between samples. Conditional Similarity Learning (CSL) aims to address this gap by learning multiple, distinct feature spaces. Existing approaches in CSL often fail to capture the intricate similarity relationships between samples across different semantic conditions, particularly in weakly-supervised settings where condition labels are absent during training. To address this limitation, we introduce Distance Induced Semantic COndition VERification NETwork (DiscoverNet), a unified framework designed to cater to a range of CSL scenarios- supervised CSL (sCSL), weakly-supervised CSL (wsCSL), and semi-supervised CSL (ssCSL). In addition to traditional linear projections, we also introduce a prompt learning technique utilizing transformer encoding layer to create diverse embedding spaces. Our framework incorporates a Condition Match Module (CMM) that dynamically matches different training triplets with corresponding embedding spaces, adapting to varying levels of supervision. We also shed light on existing evaluation biases in wsCSL and introduce two novel criteria for a more robust evaluation. Through extensive experiments and visualizations on benchmark datasets such as UT-Zappos-50 k and Celeb-A, we substantiate the efficacy and interpretability of DiscoverNet. Rui-Xiang Li, Le Gan, De-Chuan Zhan, Han-Jia Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Weight Scope Alignment: A Frustratingly Easy Method for Model MergingabstractMerging models becomes a fundamental procedure in some applications that consider model efficiency and robustness. The training randomness or Non-I.I.D. data poses a huge challenge for averaging-based model fusion. Previous research efforts focus on element-wise regularization or neural permutations to enhance model averaging while overlooking weight scope variations among models, which can significantly affect merging effectiveness. In this paper, we reveal variations in weight scope under different training conditions, shedding light on its influence on model merging. Fortunately, the parameters in each layer basically follow the Gaussian distribution, which inspires a novel and simple regularization approach named Weight Scope Alignment (WSA). It contains two key components: 1) leveraging a target weight scope to guide the model training process for ensuring weight scope matching in the subsequent model merging. 2) fusing the weight scope of two or more models into a unified one for multi-stage model fusion. We extend the WSA regularization to two different scenarios, including Mode Connectivity and Federated Learning. Abundant experimental studies validate the effectiveness of our approach. Yichu Xu, Xin-Chun Li, Le Gan, De-Chuan Zhan |
ECAI | 3 |
| 2024 | Revisit the Essence of Distilling Knowledge through CalibrationabstractKnowledge Distillation (KD) has evolved into a practical technology for transferring knowledge from a well-performing model (teacher) to a weak model (student). A counter-intuitive phenomenon known as capacity mismatch has been identified, wherein KD performance may not be good when a better teacher instructs the student. Various preliminary methods have been proposed to alleviate capacity mismatch, but a unifying explanation for its cause remains lacking. In this paper, we propose a unifying analytical framework to pinpoint the core of capacity mismatch based on calibration. Through extensive analytical experiments, we observe a positive correlation between the calibration of the teacher model and the KD performance with original KD methods. As this correlation arises due to the sensitivity of metrics (e.g., KL divergence) to calibration, we recommend employing measurements insensitive to calibration such as ranking-based loss. Our experiments demonstrate that ranking-based loss can effectively replace KL divergence, aiding large models with poor calibration to teach better. Wen-Shu Fan, Su Lu, Xin-Chun Li, De-Chuan Zhan, Le Gan |
ICML | 5 |
| 2024 | SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual DatasetsabstractModel-based offline reinforcement Learning (RL) is a promising approach that leverages existing data effectively in many real-world applications, especially those involving high-dimensional inputs like images and videos. To alleviate the distribution shift issue in offline RL, existing model-based methods heavily rely on the uncertainty of learned dynamics. However, the model uncertainty estimation becomes significantly biased when observations contain complex distractors with non-trivial dynamics. To address this challenge, we propose a new approach - *Separated Model-based Offline Policy Optimization* (SeMOPO) - decomposing latent states into endogenous and exogenous parts via conservative sampling and estimating model uncertainty on the endogenous states only. We provide a theoretical guarantee of model uncertainty and performance bound of SeMOPO. To assess the efficacy, we construct the Low-Quality Vision Deep Data-Driven Datasets for RL (LQV-D4RL), where the data are collected by non-expert policy and the observations include moving distractors. Experimental results show that our method substantially outperforms all baseline methods, and further analytical experiments validate the critical designs in our method. The project website is https://sites.google.com/view/semopo. Shenghua Wan, Le Gan, De-Chuan Zhan |
ICML | 3 |
| 2024 | AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual DistractorsabstractModel-based methods have significantly contributed to distinguishing task-irrelevant distractors for visual control. However, prior research has primarily focused on heterogeneous distractors like noisy background videos, leaving homogeneous distractors that closely resemble controllable agents largely unexplored, which poses significant challenges to existing methods. To tackle this problem, we propose Implicit Action Generator (IAG) to learn the implicit actions of visual distractors, and present a new algorithm named implicit Action-informed Diverse visual Distractors Distinguisher (AD3), that leverages the action inferred by IAG to train separated world models. Implicit actions effectively capture the behavior of background distractors, aiding in distinguishing the task-irrelevant components, and the agent can optimize the policy within the task-relevant state space. Our method achieves superior performance on various visual control tasks featuring both heterogeneous and homogeneous distractors. The indispensable role of implicit actions learned by IAG is also empirically validated. Yucen Wang, Shenghua Wan, Le Gan, De-Chuan Zhan |
ICML | 3 |
| 2024 | MOSER: Learning Sensory Policy for Task-specific Viewpoint via View-conditional World Model
Shenghua Wan, Hai-Hang Sun, Le Gan, De-Chuan Zhan |
IJCAI | 3 |
| 2024 | Leveraging Separated World Model for Exploration in Visually Distracted EnvironmentsabstractModel-based unsupervised reinforcement learning (URL) has gained prominence for reducing environment interactions and learning general skills using intrinsic rewards. However, distractors in observations can severely affect intrinsic reward estimation, leading to a biased exploration process, especially in environments with visual inputs like images or videos. To address this challenge, we propose a bi-level optimization framework named Separation-assisted eXplorer (SeeX). In the inner optimization, SeeX trains a separated world model to extract exogenous and endogenous information, minimizing uncertainty to ensure task relevance. In the outer optimization, it learns a policy on imaginary trajectories generated within the endogenous state space to maximize task-relevant uncertainty. Evaluations on multiple locomotion and manipulation tasks demonstrate SeeX's effectiveness. Kaichen Huang, Shenghua Wan, Minghao Shao, Hai-Hang Sun, Le Gan, De-Chuan Zhan |
NeurIPS | 5 |
| 2023 | Beyond probability partitions: Calibrating neural networks with semantic aware groupingabstractResearch has shown that deep networks tend to be overly optimistic about their predictions, leading to an underestimation of prediction errors. Due to the limited nature of data, existing studies have proposed various methods based on model prediction probabilities to bin the data and evaluate calibration error. We propose a more generalized definition of calibration error called Partitioned Calibration Error (PCE), revealing that the key difference among these calibration error metrics lies in how the data space is partitioned. We put forth an intuitive proposition that an accurate model should be calibrated across any partition, suggesting that the input space partitioning can extend beyond just the partitioning of prediction probabilities, and include partitions directly related to the input. Through semantic-related partitioning functions, we demonstrate that the relationship between model accuracy and calibration lies in the granularity of the partitioning function. This highlights the importance of partitioning criteria for training a calibrated and accurate model. To validate the aforementioned analysis, we propose a method that involves jointly learning a semantic aware grouping function based on deep model features and logits to partition the data space into subsets. Subsequently, a separate calibration function is learned for each subset. Experimental results demonstrate that our approach achieves significant performance improvements across multiple datasets and network architectures, thus highlighting the importance of the partitioning function for calibration. De-Chuan Zhan, Le Gan |
NeurIPS | 3 |
| 2022 | Exploring Transferability Measures and Domain Selection in Cross-Domain Slot FillingabstractAs an essential task for natural language understanding, slot filling aims to identify the contiguous spans of specific slots in an utterance. In real-world applications, the labeling costs of utterances may be expensive, and transfer learning techniques have been developed to ease this problem. However, cross-domain slot filling could significantly suffer from negative transfer due to non-targeted or zero-shot slots. Originally, this paper explores several ways to measure transferability across slot filling domains and finds that the shared slot number could serve as an efficient and effective estimator. First, this frustratingly easy measure requires no training data and is efficient to calculate. Second, it guides us heuristically select source domains that contain more shared slots with the target domain, which obtains SOTA results on Snips benchmark. Third, a dynamic transfer procedure based on this estimator clearly shows the negative transfer in cross-domain slot filling. We finally explore a source-free scene that we could only obtain black-box source models and propose to weight source domains based on prediction entropy. Xin-Chun Li, Yan-Jia Wang, Le Gan, De-Chuan Zhan |
ICASSP | 3 |
| 2022 | Avoid Overfitting User Specific Information in Federated Keyword SpottingabstractKeyword spotting (KWS) aims to discriminate a specific wakeup word from other signals precisely and efficiently for different users.Recent works utilize various deep networks to train KWS models with all users' speech data centralized without considering data privacy.Federated KWS (FedKWS) could serve as a solution without directly sharing users' data.However, the small amount of data, different user habits, and various accents could lead to fatal problems, e.g., overfitting or weight divergence.Hence, we propose several strategies to encourage the model not to overfit user-specific information in FedKWS.Specifically, we first propose an adversarial learning strategy, which updates the downloaded global model against an overfitted local model and explicitly encourages the global model to capture user-invariant information.Furthermore, we propose an adaptive local training strategy, letting clients with more training data and more uniform class distributions undertake more local update steps.Equivalently, this strategy could weaken the negative impacts of those users whose data is less qualified.Our proposed FedKWS-UI could explicitly and implicitly learn user-invariant information in FedKWS.Abundant experimental results on federated Google Speech Commands verify the effectiveness of FedKWS-UI. Xin-Chun Li, Jin-Lin Tang, Shaoming Song, Bingshuai Li, Yinchuan Li, Yunfeng Shao 0001, Le Gan, De-Chuan Zhan |
INTERSPEECH | 7 |
| 2021 | Towards Enabling Meta-Learning from Target ModelsabstractMeta-learning can extract an inductive bias from previous learning experience and assist the training of new tasks. It is often realized through optimizing a meta-model with the evaluation loss of task-specific solvers. Most existing algorithms sample non-overlapping $\mathit{support}$ sets and $\mathit{query}$ sets to train and evaluate the solvers respectively due to simplicity ($\mathcal{S}$/$\mathcal{Q}$ protocol). Different from $\mathcal{S}$/$\mathcal{Q}$ protocol, we can also evaluate a task-specific solver by comparing it to a target model $\mathcal{T}$, which is the optimal model for this task or a model that behaves well enough on this task ($\mathcal{S}$/$\mathcal{T}$ protocol). Although being short of research, $\mathcal{S}$/$\mathcal{T}$ protocol has unique advantages such as offering more informative supervision, but it is computationally expensive. This paper looks into this special evaluation method and takes a step towards putting it into practice. We find that with a small ratio of tasks armed with target models, classic meta-learning algorithms can be improved a lot without consuming many resources. We empirically verify the effectiveness of $\mathcal{S}$/$\mathcal{T}$ protocol in a typical application of meta-learning, $\mathit{i.e.}$, few-shot learning. In detail, after constructing target models by fine-tuning the pre-trained network on those hard tasks, we match the task-specific solvers and target models via knowledge distillation. Su Lu, Han-Jia Ye, Le Gan, De-Chuan Zhan |
NeurIPS | 3 |
| 2018 | Multisource Earth Observation Data for Land-Cover Classification Using Random ForestabstractIn this letter, multisource earth observation (EO) data sets, including multitemporal Landsat-8, digital surface model, and spatial information, were integrated for land-cover classification by random forest (RF) and support vector machines (SVMs). We demonstrated in this letter that both RF and SVM are useful tools for classification of land cover in the local climate zones featured with highly heterogeneous landscape. Classification of land cover by RF was with an overall accuracy (OA) of 86.2%, while the OA was 85.5% for SVM. However, we found that RF was more stable than SVM for multisource EO data in classifying land cover without normalizing different feature data sets. Experiments showed that the thermal features were more important than temporal and spatial ones in discriminating impervious objects, while the temporal and spatial features were generally better than thermal ones in separating the distinct vegetation categories. Another finding was that our experiments indicated that spectral features were the most important in classification of land cover, followed by temporal, thermal, and spatial features, respectively. As to the spectral features, red channels were the most important, followed by short-wave infrared, near-infrared, and green channels. Thus, it could be concluded that the combination of spectral, thermal, spatial, and temporal information would be an optimal approach to increase the OA of land-cover classification in the zones featured with highly heterogeneous landscape. Jike Chen, Junshi Xia, Peijun Du, Hongrui Zheng, Le Gan |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2018 | Multikernel Adaptive Collaborative Representation for Hyperspectral Image ClassificationabstractTo adequately represent the nonlinearities in the high-dimensional feature space for hyperspectral images (HSIs), we propose a multiple kernel collaborative representation-based classifier (CRC) in this paper. Extended morphological profiles are first extracted from the original HSIs, because they can efficiently capture the spatial and spectral information. In the proposed method, a novel multiple kernel learning (MKL) model is embedded into CRC. Multiple kernel patterns, e.g., Naive, Multimetric, and Multiscale are adopted for the optimal set of basic kernels, which are helpful to capture the useful information from different pixel distributions, kernel metric spaces, and kernel scales. To learn an optimal linear combination of the predefined basic kernels, we add an extra training stage to the typical CRC where kernel weights are jointly learned with the representation coefficients from the training samples by minimizing the representation error. Moreover, by considering different contributions of dictionary atoms, the adaptive representation strategy is applied to the MKL framework via a dissimilarity-weighted regularizer to obtain a more robust representation of test pixels in the fused kernel space. Experimental results on three real HSIs confirm that the proposed classifiers outperform the other state-of-the-art representation-based classifiers. Peijun Du, Le Gan, Junshi Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Multiple Feature Kernel Sparse Representation Classifier for Hyperspectral ImageryabstractMultiple types of features, e.g., spectral, filtering, texture, and shape features, are helpful for hyperspectral image (HSI) classification tasks. Combining multiple features can describe the characteristics of pixels from different perspectives, and always results in better classification performance. Recently, multifeature combination learning has been widely employed to the multitask-learning-based representation-based model to obtain a multifeature representation vector. However, the linear sparse representation-based classifier (SRC) cannot handle the HSI with highly nonlinear distribution, and kernel sparse representation-based classifier (KSRC) can remedy the drawback of linear SRC. By adopting nonlinear mapping, the samples in kernel space are often of high or even infinite dimensionality. In this paper, we integrate kernel principal component analysis into multifeature-based KSRC and propose a novel multiple feature kernel sparse representation-based classifier (namely, MFKSRC) for hyperspectral imagery. More specifically, spatial features, Gabor textures, local binary patterns, and difference morphological profiles are adopted and then each kind of feature is transformed nonlinearly into a new low-dimensional kernel space. The proposed framework can handle data with nonlinear distribution and add a dimensionality reduction stage in kernel space before optimizing the corresponding cost function. Experimental results on different HSIs demonstrate that the proposed MFKSRC algorithm outperforms the state-of-the-art classifiers. Le Gan, Junshi Xia, Peijun Du, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Kernel Fused Representation-Based Classifier for Hyperspectral ImageryabstractIn this letter, we propose a kernel fused representation-based classifier (KFRC) for hyperspectral images (HSIs), which combines sparse representation (SR) and collaborative representation (CR) into a unified kernel representation-based classification framework. First, we present two individual kernel methods, i.e., kernel SR (KSR) and kernel CR (KCR), which kernelize the representation methods by projecting the samples into a high-dimensional kernel space to improve the samples separability between different classes. Once obtaining the two kernel representation coefficients, KFRC attempts to achieve a balance between KSR and KCR via an adjusting parameter $\theta $ in the kernel residual domain. Subsequently, the class label of each test sample is determined by the minimum residual for each class. Experimental results on two HSIs demonstrate the proposed kernel fused method performs better than the other state-of-the-art representation-based classifiers. Le Gan, Peijun Du, Junshi Xia, Yaping Meng |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Dissimilarity-Weighted Sparse Representation for Hyperspectral Image ClassificationabstractTo improve the capability of a traditional sparse representation-based classifier (SRC), we propose a novel dissimilarity-weighted SRC (DWSRC) for hyperspectral image (HSI) classification. In particular, DWSRC computes the weights for each atom according to the distance or dissimilarity information between the test pixel and the atoms. First, a locality constraint dictionary set is constructed by the Gaussian kernel distance with a suitable distance metric (e.g., Euclidean distance). Second, the test pixel is sparsely coded over the new weighted dictionary set based on the 11-norm minimization problem. Finally, the test pixel is classified by using the obtained sparse coefficients with the minimal residual rule. Experimental results on two widely used public HSIs demonstrate that the proposed DWSRC is more efficient and accurate than other state-of-the-art SRCs. Le Gan, Junshi Xia, Peijun Du |
IEEE Geosci. Remote. Sens. Lett. | 1 |