Chunlin Chen 0001

dblp:68/6992-1 · DBLP profile ↗
← Back
92ranked-venue papers
5as first author
73since 2021 · last 2026
0000-0003-3929-4707ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 4 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 20 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 14 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021
YearPublicationVenuePosition
2026 Online Cross-Modal Hashing with Expanding Label Space
abstract
Due to the continuous increase of multimedia data on the internet, online hashing has garnered considerable attention for handling multi-modal data streams. However, most existing online hashing approaches focus solely on data growth of samples, overlooking the dynamics of classes. In this paper, we simultaneously address the challenges of both sample-level and class-level growth, and propose a novel Online Hashing method with Expanding Label Space (OH-ELS) for cross-modal retrieval. In OH-ELS, multi-modal data arrives continuously, and incoming data may introduce new classes. To avoid catastrophic forgetting, we transfer the historical knowledge at both the sample and class levels. At the sample-level, a small subset of anchor codes from old data are replayed to preserve the similarities between new data and old data. At the class-level, a consistency regularizer is applied to new classifiers to leverage the priors of historical classes. To ensure both efficiency and accuracy, a discrete optimization algorithm is proposed to solve the binary-constrained optimization problem without relaxation. Experimental results illustrate the effectiveness and superiority of OH-ELS in class-incremental cross-modal retrieval compared with the state-of-the-art methods.
Wentao Fan 0003, Chao Zhang 0078, Chunlin Chen 0001, Huaxiong Li
AAAI3
2026 Semantic-Aware Feature Enhancement for Partial Label Learning
abstract
Partial label learning (PLL) aims to learn from the data where each instance is associated with a candidate label set, with only one being valid. Most existing approaches are designed to eliminate noisy labels and use the remaining reliable ones for model training, following a label-centric learning paradigm. In this paper, we propose a new PLL method called Semantic-Aware Feature Enhancement (SAFE), which tackles the problem through a novel feature-centric learning paradigm. SAFE presumes that the candidate labels are correct while the observed features are partial, and thus seeks to recover the underlying missing features. In this manner, a desired predictive model is constructed by integrating the observed and recovered features, which are responsible for predicting the true label and the remaining candidate labels, respectively. To ensure the quality of recovered features, SAFE jointly explores the intrinsic topological structures via dynamic graphs in both feature and label spaces as guidance for semantic-aware feature enhancement. Extensive experimental results on some popular datasets demonstrate the effectiveness and superiority of the proposed method over state-of-the-art PLL approaches.
Haowei Mei, Chao Zhang 0078, Wentao Fan 0003, Xiuyi Jia, Chunlin Chen 0001, Huaxiong Li
AAAI5
2026 Semantic-Augmented Image Clustering via Adaptive Multi-Modal Collaboration
abstract
Image clustering is a fundamental task in unsupervised visual learning. While recent self-supervised methods have explored various pretext tasks to generate supervision signals for clustering, they typically depend exclusively on raw images, resulting in insufficient supervision signals that are inherently constrained by limited visual semantics. In this paper, we propose a novel Semantic-Augmented image Clustering (SAC) method, which transcends the inherent limitations of purely visual representations through the integration of external knowledge. Specifically, SAC utilizes Vision-Language pre-trained Models (VLMs) to flexibly generate textual descriptions for each image, providing external semantic cues to supplement the visual information. By integrating both visual and textual information, SAC achieves image clustering through a multi-modal learning framework. To mitigate the negative impact of inaccurate textual information, SAC designs an uncertainty-driven adaptive weighting mechanism that explores both intra-modal and inter-modal neighborhood structures, and incorporates the adaptive weights into intra-modal and inter-modal contrastive learning, which improves the robustness against noisy image-text correspondences. Experiments on several popular datasets demonstrate the superiority of SAC compared to state-of-the-art methods.
Chao Zhang 0078, Deng Xu, Hong Yu 0007, Chunlin Chen 0001, Huaxiong Li
AAAI5
2026 Conditional Diffusion Model for Multi-Agent Dynamic Task Decomposition
abstract
Task decomposition has shown promise in complex cooperative multi-agent reinforcement learning (MARL) tasks, which enables efficient hierarchical learning for long-horizon tasks in dynamic and uncertain environments. However, learning dynamic task decomposition from scratch generally requires a large number of training samples, especially exploring the large joint action space under partial observability. In this paper, we present the Conditional Diffusion Model for Dynamic Task Decomposition (CD3T), a novel two-level hierarchical MARL framework designed to automatically infer subtask and coordination patterns. The high-level policy learns subtask representation to generate a subtask selection strategy based on subtask effects. To capture the effects of subtasks on the environment, CD3T predicts the next observation and reward using a conditional diffusion model. At the low level, agents collaboratively learn and share specialized skills within their assigned subtasks. Moreover, the learned subtask representation is also used as additional semantic information in a multi-head attention mixing network to enhance value decomposition and provide an efficient reasoning bridge between individual and joint value functions. Experimental results on various benchmarks demonstrate that CD3T achieves better performance than existing baselines.
Yanda Zhu, Yuanyang Zhu, Daoyi Dong, Caihua Chen, Chunlin Chen 0001
AAAI5
2026 Diverse embeddings and consensus pseudo-supervision learning for unsupervised feature selection
Ziqi Meng, Wentao Fan 0003, Bo Wang 0027, Chunlin Chen 0001, Huaxiong Li
Inf. Sci.4
2026 GCM: Interpretable Multiagent Reinforcement Learning via Graph Cooperation Modeling
abstract
Multiagent reinforcement learning (MARL) has been widely investigated, ranging from theoretical analysis to real-life applications. However, the utilization of existing non-transparent neural network architectures has resulted in opaque decision-making processes, making it difficult for humans to understand and trust the models being used. Fundamentally, all data is a topological structure, which provides reliable transparency for MARL tasks due to its powerful relational expression capability, scalability, and explicit structural relationships. In this article, we propose a novel approach of graph cooperation modeling (GCM), explicitly capturing and comprehending the complex dynamics of collaborative relationships among agents with the graph structure. GCM learns a metric function to discern beneficial interactions among agents, integrating it into the agent aggregation strategy of a graph neural network (GNN) capable of modeling arbitrary-order interactions. Furthermore, GCM utilizes identity semantics together with global state and individual value functions to estimate the credit of each agent, enhancing each agent's distinct focus on task-related regions. Extensive experiments on a range of challenging MARL benchmarks demonstrate that GCM not only delivers up to 28.75% relative performance gains on super-hard maps but also offers clear interpretability that provides insights into the underlying cooperative patterns.
Xuefei Wu, Yuanyang Zhu, Caihua Chen, Chunlin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Fast Incomplete Multi-view Clustering with Adaptive Similarity Completion and Reconstruction
abstract
Recently, anchor-based incomplete multi-view clustering (IMVC) has been widely adopted for fast clustering, but most existing approaches still encounter some issues: (1) They generally rely on the observed samples to construct anchor graphs, ignoring the potentially useful information of missing instances. (2) Most methods attempt to learn a consensus anchor graph, failing to fully excavate the complementary information and high-order correlations across views. (3) They generally apply post-processing on learned anchor graph to seek latent embeddings, making them not globally-optimal. To address these issues, this paper proposes a novel fast IMVC approach with Adaptive Similarity Completion and Reconstruction (ASCR), which unifies anchor learning, anchor-sample similarity construction and completion, and latent multi-view embedding learning in a joint framework. Specifically, ASCR learns an anchor-sample similarity graph for each view, and the missing values are fulfilled to mitigate the adverse effects. To explore the consistent and complementary information across views, ASCR simultaneously seeks the view-specific anchor embeddings and sample embeddings in a latent subspace by similarity reconstruction, which not only preserves the semantic information into latent embeddings but also enhances the low-rank property of similarity graphs, achieving a reliable graph completion process. Furthermore, the high-order cross-view correlations are explored with tensor-based regularization. Extensive experimental results demonstrate the superiority and efficiency of ASCR compared with SOTA approaches.
Deng Xu, Chao Zhang 0078, Cong Guo 0008, Chunlin Chen 0001, Huaxiong Li
AAAI4
2025 PN-GAIL: Leveraging Non-optimal Information from Imperfect Demonstrations
abstract
Imitation learning aims at constructing an optimal policy by emulating expert demonstrations. However, the prevailing approaches in this domain typically presume that the demonstrations are optimal, an assumption that seldom holds true in the complexities of real-world applications. The data collected in practical scenarios often contains imperfections, encompassing both optimal and non-optimal examples. In this study, we propose Positive-Negative Generative Adversarial Imitation Learning (PN-GAIL), a novel approach that falls within the framework of Generative Adversarial Imitation Learning (GAIL). PN-GAIL innovatively leverages non-optimal information from imperfect demonstrations, allowing the discriminator to comprehensively assess the positive and negative risks associated with these demonstrations. Furthermore, it requires only a small subset of labeled confidence scores. Theoretical analysis indicates that PN-GAIL deviates from the non-optimal data while mimicking imperfect demonstrations. Experimental results demonstrate that PN-GAIL surpasses conventional baseline methods in dealing with imperfect demonstrations, thereby significantly augmenting the practical utility of imitation learning in real-world contexts. Our codes are available at https://github.com/QiangLiuT/PN-GAIL.
Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001, Daoyi Dong
ICLR4
2025 Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning
abstract
While showing sophisticated reasoning abilities, large language models (LLMs) still struggle with long-horizon decision-making tasks due to deficient exploration and long-term credit assignment, especially in sparse-reward scenarios. Inspired by the divide-and-conquer principle, we propose an innovative framework GLIDER (Grounding Language Models as EffIcient Decision-Making Agents via Offline HiErarchical Reinforcement Learning) that introduces a parameter-efficient and generally applicable hierarchy to LLM policies. We develop a scheme where the low-level controller is supervised with abstract, step-by-step plans that are learned and instructed by the high-level policy. This design decomposes complicated problems into a series of coherent chain-of-thought reasoning sub-tasks, providing flexible temporal abstraction to significantly enhance exploration and learning for long-horizon tasks. Furthermore, GLIDER facilitates fast online adaptation to non-stationary environments owing to the strong transferability of its task-agnostic low-level skills. Experiments on ScienceWorld and ALFWorld benchmarks show that GLIDER achieves consistent performance gains, along with enhanced generalization capabilities.
Zican Hu, Wei Liu 0131, Xiaoye Qu, Xiangyu Yue 0001, Chunlin Chen 0001, Zhi Wang 0001, Yu Cheng 0001
ICML5
2025 Semi-supervised Multi-view Clustering with Active Constraints
abstract
Multi-view clustering has attracted increasing attention in recent years. However, most existing multi-view clustering approaches are performed in a purely unsupervised manner, while ignoring the valuable weak supervision information that can be obtained (e.g., active query) in many real applications. This paper considers the weak pairwise constraints among samples to enhance the clustering performance, and proposes a Semi-supervised Multi-view Clustering method with Active Constraints, SMCAC for short. SMCAC consists of two stages, clustering (C-stage) and active query (A-stage). In the C-stage, we design a tensor based multi-view graph learning model equipped with sample pairwise constraints regularization to facilitate the discriminative graph learning and fusion. An effective optimization algorithm based on alternating direction minimization is devised to solve the clustering model. In the A-stage, the most uncertain or difficult sample pairs are actively selected to query the constraints, based on the divergence of multi-view similarities learned in the C-stage. The two processes alternate iteratively until the maximum number of queries is reached. Extensive experiments on several popular datasets well validate the effectiveness of the proposed method.
Chao Zhang 0078, Deng Xu, Chunlin Chen 0001, Huaxiong Li
KDD (1)3
2025 Online Cross-Modal Hashing with Multi-Level Memory
abstract
Online cross-modal hashing has recently gained significant attention due to its remarkable capability to handle cross-modal streaming data retrieval. Despite promising progress, existing methods still face challenges in fully exploiting the intricate relations across heterogeneous modalities and streaming data chunks, limiting the retrieval performance. In this paper, a novel Online Cross-modal Hashing method with Multi-level Memory (OCH-MM) is proposed. OCH-MM captures the cross-modal consistency and sample semantic correlations for discrete hash learning with latent feature disentanglement, and designs a multi-level memory framework for effective knowledge transfer. Specifically, for discriminative hash learning, OCH-MM maps the multi-modal data into a latent feature space that is further disentangled into a common Hamming space and a modality-specific feature space. The semantic correlations among samples are also preserved into discrete hash codes without relaxation in a nonlinear manner. For effectively learning from streaming data, OCH-MM designs an intra-space feature association memory, an inter-space feature association memory, and a hash codes memory, which encode the historical feature correlations within original multi-modal spaces, the feature correlations between original and latent space, and a subset of hash codes, respectively. By dynamically updating and utilizing the multi-level memory, the data correlations between different chunks are well explored and the historical knowledge is effectively reused to guide future learning. The proposed model is solved by an efficient discrete optimization algorithm. Experimental results on three benchmark datasets demonstrate that our proposed method achieves better retrieval accuracy over the state-of-the-art baselines.
Wentao Fan 0003, Chao Zhang 0078, Chunlin Chen 0001, Huaxiong Li
ACM Multimedia3
2025 Mixture-of-Experts Meets In-Context Reinforcement Learning
abstract
In-context reinforcement learning (ICRL) has emerged as a promising paradigm for adapting RL agents to downstream tasks through prompt conditioning. However, two notable challenges remain in fully harnessing in-context learning within RL domains: the intrinsic multi-modality of the state-action-reward data and the diverse, heterogeneous nature of decision tasks. To tackle these challenges, we propose **T2MIR** (**T**oken- and **T**ask-wise **M**oE for **I**n-context **R**L), an innovative framework that introduces architectural advances of mixture-of-experts (MoE) into transformer-based decision models. T2MIR substitutes the feedforward layer with two parallel layers: a token-wise MoE that captures distinct semantics of input tokens across multiple modalities, and a task-wise MoE that routes diverse tasks to specialized experts for managing a broad task distribution with alleviated gradient conflicts. To enhance task-wise routing, we introduce a contrastive learning method that maximizes the mutual information between the task and its router representation, enabling more precise capture of task-relevant information. The outputs of two MoE components are concatenated and fed into the next layer. Comprehensive experiments show that T2MIR significantly facilitates in-context learning capacity and outperforms various types of baselines. We bring the potential and promise of MoE to ICRL, offering a simple and scalable architectural enhancement to advance ICRL one step closer toward achievements in language and vision communities. Our code is available at [https://github.com/NJU-RL/T2MIR](https://github.com/NJU-RL/T2MIR).
Fuhong Liu, Haoru Li, Zican Hu, Daoyi Dong, Chunlin Chen 0001, Zhi Wang 0001
NeurIPS6
2025 DEAL: Diffusion Evolution Adversarial Learning for Sim-to-Real Transfer
abstract
Training Reinforcement Learning (RL) controllers in simulation offers cost-efficiency and safety advantages. However, the resultant policies often suffer significant performance degradation during real-world deployment due to the reality gap. Previous works like System Identification (Sys-Id) have attempted to bridge this discrepancy by improving simulator fidelity, but encounter challenges including the collapse of high-dimensional parameter identification, low identification accuracy, and unstable convergence dynamics. To address these challenges, we propose a novel Sys-Id framework that combines Diffusion Evolution with Adversarial Learning (DEAL) to iteratively infer physical parameters with limited real-world data, which makes the state transitions between simulation and reality as similar as possible. Specifically, our method iteratively refines physical parameters through a dual mechanism: a discriminator network evaluates the similarity of state transitions between parameterized simulations and target environment as fitness guidance, while diffusion evolution adaptively modulates noise prediction and denoising processes to optimize parameter distributions. We validate DEAL in both simulated and real-world environments. Compared to baseline methods, DEAL demonstrates state-of-the-art stability and identification accuracy in high-dimensional parameter identification tasks, and significantly enhances sim-to-real transfer performance while requiring minimal real-world data.
Huiqiao Fu, Zhehao Zhou, Chunlin Chen 0001
NeurIPS5
2025 High-order Interactions Modeling for Interpretable Multi-Agent Q-Learning
abstract
The ability to model interactions among agents is crucial for effective coordination and understanding their cooperation mechanisms in multi-agent reinforcement learning (MARL). However, previous efforts to model high-order interactions have been primarily hindered by the combinatorial explosion or the opaque nature of their black-box network structures. In this paper, we propose a novel value decomposition framework, called Continued Fraction Q-Learning (QCoFr), which can flexibly capture arbitrary-order agent interactions with only linear complexity $\mathcal{O}\left({n}\right)$ in the number of agents, thus avoiding the combinatorial explosion when modeling rich cooperation. Furthermore, we introduce the variational information bottleneck to extract latent information for estimating credits. This latent information helps agents filter out noisy interactions, thereby significantly enhancing both cooperation and interpretability. Extensive experiments demonstrate that QCoFr not only consistently achieves better performance but also provides interpretability that aligns with our theoretical analysis.
Qinyu Xu, Yuanyang Zhu, Xuefei Wu, Chunlin Chen 0001
NeurIPS4
2025 Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language Supervision
abstract
Offline meta-RL usually tackles generalization by inferring task beliefs from high-quality samples or warmup explorations. The restricted form limits their generality and usability since these supervision signals are expensive and even infeasible to acquire in advance for unseen tasks. Learning directly from the raw text about decision tasks is a promising alternative to leverage a much broader source of supervision. In the paper, we propose **T**ext-to-**D**ecision **A**gent (**T2DA**), a simple and scalable framework that supervises offline meta-RL with natural language. We first introduce a generalized world model to encode multi-task decision data into a dynamics-aware embedding space. Then, inspired by CLIP, we predict which textual description goes with which decision embedding, effectively bridging their semantic gap via contrastive language-decision pre-training and aligning the text embeddings to comprehend the environment dynamics. After training the text-conditioned generalist policy, the agent can directly realize zero-shot text-to-decision generation in response to language instructions. Comprehensive experiments on MuJoCo and Meta-World benchmarks show that T2DA facilitates high-capacity zero-shot generalization and outperforms various types of baselines. Our code is available at [https://github.com/NJU-RL/T2DA](https://github.com/NJU-RL/T2DA).
Zican Hu, Jianxiang Tang, Chunlin Chen 0001, Daoyi Dong, Yu Cheng 0001, Zhenhong Sun, Zhi Wang 0001
NeurIPS6
2025 Constrained Policy Optimization with Approximately Monotonically Increasing Rewards
abstract
In reinforcement learning (RL), agents maximize accumulated rewards through trial-and-error in the environment to obtain high-performing policies. However, in some situations, loopholes in the purely synthetic reward signals are often exploited by agents, leading to unsafe behaviors, which necessitates the incorporation of safety constraints. In this paper, we propose a safe RL algorithm called Constrained Policy Optimization with Approximately Monotonically Increasing Rewards (CPO-AMIR) to address the safe policy learning in different scenarios and provide practical solutions. We present a novel update formula to achieve a better balance between increasing the reward and decreasing the cost. Furthermore, a theoretical analysis is provided to demonstrate that our approach guarantees the approximately monotonic improvement in rewards when learning constraint-satisfying policies. Our empirical results illustrate the effectiveness and superiority of CPO-AMIR on a set of constrained control tasks.
Yuanyang Lu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001
SMC4
2025 Imitation Learning with Process Adversarial Diffusion
abstract
Generative Adversarial Imitation Learning (GAIL) replicates expert behaviors by employing adversarial training involving the discriminator and the generator. In theory, GAIL can balance discriminator and generator through considerable online interactions to learn well. But real-world adversarial training is tricky and unstable. Early mistakes by the discriminator can lead to bad learning outcomes, or the generator might just copy average expert actions (mode collapse) instead of diverse strategies. To address these, we propose Process Adversarial Diffusion Imitation Learning (PADIL). Employing a conditional diffusion model as the generator facilitates the generation of a multi-modal strategy, thereby reducing the likelihood of mode collapse and enhancing imitation learning performance with fewer online interactions. At the same time, it lessens the generator’s vulnerability to the fluctuating signals or errors conveyed by the discriminator throughout the training process. Furthermore, we have revised the sample extraction approach of the diffusion policy to tackle the concern of iterative instability under rapidly fluctuating reward signals. Experimental results demonstrate that our method achieves better performance compared to baseline methods.
Yiming Qi, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001
SMC4
2025 Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise
abstract
Efficient exploration in multi-agent reinforcement learning (MARL) is a challenging problem when receiving only a team reward, especially in environments with sparse rewards. A powerful method to mitigate this issue involves crafting dense individual rewards to guide the agents toward efficient exploration. However, individual rewards generally rely on manually engineered shaping-reward functions that lack high-order intelligence, thus it behaves ineffectively than humans regarding learning and generalization in complex problems. To tackle these issues, we combine the above two paradigms and propose a novel framework, LIGHT (Learning Individual Intrinsic reward via Incorporating Generalized Human experTise), which can integrate human knowledge into MARL algorithms in an end-to-end manner. LIGHT guides each agent to avoid unnecessary exploration by considering both individual action distribution and human expertise preference distribution. Then, LIGHT designs individual intrinsic rewards for each agent based on actionable representational transformation relevant to Q-learning so that the agents align their action preferences with the human expertise while maximizing the joint action value. Experimental results demonstrate the superiority of our method over representative baselines regarding performance and better knowledge reusability across different sparse-reward tasks on challenging scenarios.
Xuefei Wu, Yuanyang Zhu, Chunlin Chen 0001
SMC4
2025 Charging Scheduling Optimization of Electric Buses Considering Battery Degradation
abstract
Due to the depletion of fossil fuels and the rise of electric vehicles, electric buses have emerged as a new means of transportation, helping to reduce air pollution and energy consumption. However, the limited battery capacity of electric buses makes on-the-go charging a major concern, as installing chargers at every bus stop is impractical due to costs. Therefore, employing appropriate charging strategies for electric buses is crucial. A critical issue that urgently needs resolution is determining the optimal timing and quantity for charging electric buses, considering the current size of the bus fleet, routes, and charging station facilities. Battery aging, which affects battery capacity and thereby influences charging decisions and the operating costs of bus routes, must be considered. This paper proposes a hybrid integer programming model to describe bus line operations and uses a semi-empirical method to estimate battery aging. The Gurobi solver is used to select an appropriate solution strategy to solve the dualobjective mathematical model in this paper, so as to reduce the operating cost of bus lines and the aging of batteries. The results show that the model proposed in this paper can effectively reduce the operating cost of the bus route and the aging of the battery.
Jingwen Wei, Chunlin Chen 0001, Guangzhong Dong
SMC3
2025 MIXRTs: Toward Interpretable Multi-Agent Reinforcement Learning via Mixing Recurrent Soft Decision Trees
abstract
While achieving tremendous success in various fields, existing multi-agent reinforcement learning (MARL) with a black-box neural network makes decisions in an opaque manner that hinders humans from understanding the learned knowledge and how input observations influence decisions. In contrast, existing interpretable approaches usually suffer from weak expressivity and low performance. To bridge this gap, we propose MIXing Recurrent soft decision Trees (MIXRTs), a novel interpretable architecture that can represent explicit decision processes via the root-to-leaf path and reflect each agent's contribution to the team. Specifically, we construct a novel soft decision tree using a recurrent structure and demonstrate which features influence the decision-making process. Then, based on the value decomposition framework, we linearly assign credit to each agent by explicitly mixing individual action values to estimate the joint action value using only local observations, providing new insights into interpreting the cooperation mechanism. Theoretical analysis confirms that MIXRTs guarantee additivity and monotonicity in the factorization of joint action values. Evaluations on complex tasks like Spread and StarCraft II demonstrate that MIXRTs compete with existing methods while providing clear explanations, paving the way for interpretable and high-performing MARL systems.
Zichuan Liu, Yuanyang Zhu, Zhi Wang 0001, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Tomography of Quantum States From Structured Measurements via Quantum-Aware Transformer
abstract
Quantum state tomography (QST) is the process of reconstructing the state of a quantum system (mathematically described as a density matrix) through a series of different measurements, which can be solved by learning a parameterized function to translate experimentally measured statistics into physical density matrices. However, the specific structure of quantum measurements for characterizing a quantum state has been neglected in previous work. In this article, we explore the similarity between highly structured sentences in natural language and intrinsically structured measurements in QST. To fully leverage the intrinsic quantum characteristics involved in QST, we design a quantum-aware transformer (QAT) model to capture the complex relationship between measured frequencies and density matrices. In particular, we query quantum operators in the architecture to facilitate informative representations of quantum data and integrate the Bures distance into the loss function to evaluate quantum state fidelity, thereby enabling the reconstruction of quantum states from measured data with high fidelity. Extensive simulations and experiments (on IBM quantum computers) demonstrate the superiority of the QAT in reconstructing quantum states with favorable robustness against experimental noise.
Hailan Ma, Zhenhong Sun, Daoyi Dong, Chunlin Chen 0001, Herschel Rabitz
IEEE Trans. Cybern.4
2025 Deep Learning-Enabled Fault Diagnosis of Lithium-Ion Batteries Using Real-World Vehicle Data With Gramian Angular Difference Fields
abstract
Battery failure represents one of the most common threats to electric vehicles (EVs). Existing diagnosis methods for onboard Lithium-ion batteries are heavily limited by complex real-world scenarios and struggle to handle early faults. With the aim of detecting battery faults in an early stage, this article proposes a deep learning-enabled fault diagnosis framework that blends the advantages of Gramian angular difference fields (GADF) and Transformer-based networks. First, a median-difference-process is developed to capture the dynamic electrical behaviors of cells. Then, the voltages of cells are converted into GADF matrices and presented as grayscale images. Afterward, a Vision Transformer is introduced to extract and learn the features of different battery fault patterns. Experimental verification in a real-world EV battery pack indicates that the proposed method achieves fault diagnosis for early battery failures with an accuracy of 97.30$\%$. Moreover, it outperforms existing methods and maintains a high recall rate of 94.77$\%$. Consequently, the proposed strategy proves to be effective for real-world applications.
Ling Xie, Jingwen Wei, Xiaoke Li, Chunlin Chen 0001, Guangzhong Dong
IEEE Trans. Ind. Informatics4
2025 Multi-View Clustering With Incremental Instances and Views
abstract
Multi-view clustering (MVC) has attracted increasing attention with the emergence of various data collected from multiple sources. In real-world dynamic environment, instances are continually gathered, and the number of views expands as new data sources become available. Learning for such simultaneous increment of instances and views, particularly in unsupervised scenarios, is crucial yet underexplored. In this paper, we address this problem by proposing a novel MVC method with Incremental Instances and Views, MVC-IIV for short. MVC-IIV contains two stages, an initial stage and an incremental stage. In the initial stage, a basic latent multi-view subspace clustering model is constructed to handle existing data, which can be viewed as traditional static MVC. In the incremental stage, the previously trained model is reused to guide learning for newly arriving instances with new views, transferring historical knowledge while avoiding redundant computations. In specific, we design and reuse two modules, i.e., multi-view embedding module for low-dimensional representation learning, and consensus centroids module for cluster probability learning. By adding consistency regularization on the two modules, the knowledge acquired from previous data is used, which not only enhances the exploration within current data batch, but also extracts the between-batch data correlations. The proposed model can be efficiently solved with linear space and time complexity. Extensive experiments demonstrate the effectiveness and efficiency of our method compared with the state-of-the-art approaches.
Chao Zhang 0078, Zhi Wang 0001, Xiuyi Jia, Zechao Li, Chunlin Chen 0001, Huaxiong Li
IEEE Trans. Image Process.5
2025 Label Distribution Guided Hashing for Cross-Modal Retrieval
abstract
Hashing methods have recently attracted extensive attention in cross-modal retrieval. Most supervised hashing methods attempt to preserve the semantic information into hash codes by leveraging the original logical label matrix. However, they generally treat all labels equally, and ignore the relative significance of different labels due to the variety of data features. In this article, we argue that exploring the relative importance of labels benefits the enhancement of semantic information, and we propose a novel LAbel Distribution Guided Hashing (LADH) method for cross-modal retrieval. In particular, LADH first learns a feature-induced label distribution for each sample to weigh different labels, which leverages the multi-modal feature information to enrich the semantic label information. By jointly using the learned label distributions and multi-modal features, the latent representation and hash codes are obtained with multi-modal feature selection and enhanced semantic similarities embedded. An efficient algorithm is designed to solve the proposed method whose time complexity is linear to the number of the training instances. Experimental results on several public benchmark datasets verify the effectiveness and efficiency of our method compared with the state-of-the-art methods.
Fatang Lei, Chao Zhang 0078, Huaxiong Li, Yang Gao 0001, Chunlin Chen 0001
ACM Trans. Knowl. Discov. Data5
2025 Fast Disentangled Slim Tensor Learning for Multi-View Clustering
abstract
Tensor-based multi-view clustering has recently received significant attention due to its exceptional ability to explore cross-view high-order correlations. However, most existing methods still encounter some limitations. (1) Most of them explore the correlations among different affinity matrices, making them unscalable to large-scale data. (2) Although some methods address it by introducing bipartite graphs, they may result in sub-optimal solutions caused by an unstable anchor selection process. (3) They generally ignore the negative impact of latent semantic-unrelated information in each view. To tackle these issues, we propose a new approach termed fast Disentangled Slim Tensor Learning (DSTL) for multi-view clustering. Instead of focusing on the multi-view graph structures, DSTL directly explores the high-order correlations among multi-view latent semantic representations based on matrix factorization. To alleviate the negative influence of feature redundancy, inspired by robust PCA, DSTL disentangles the latent low-dimensional representation into a semantic-unrelated part and a semantic-related part for each view. Subsequently, two slim tensors are constructed with tensor-based regularization. To further enhance the quality of feature disentanglement, the semantic-related representations are aligned across views through a consensus alignment indicator. Our proposed model is computationally efficient and can be solved effectively. Extensive experiments demonstrate the superiority and efficiency of DSTL over state-of-the-art approaches.
Deng Xu, Chao Zhang 0078, Zechao Li, Chunlin Chen 0001, Huaxiong Li
IEEE Trans. Multim.4
2025 Discretizing Continuous Action Space With Unimodal Probability Distributions for On-Policy Reinforcement Learning
abstract
For on-policy reinforcement learning (RL), discretizing action space for continuous control can easily express multiple modes and is straightforward to optimize. However, without considering the inherent ordering between the discrete atomic actions, the explosion in the number of discrete actions can possess undesired properties and induce a higher variance for the policy gradient (PG) estimator. In this article, we introduce a straightforward architecture that addresses this issue by constraining the discrete policy to be unimodal using Poisson probability distributions. This unimodal architecture can better leverage the continuity in the underlying continuous action space using explicit unimodal probability distributions. We conduct extensive experiments to show that the discrete policy with the unimodal probability distribution provides significantly faster convergence and higher performance for on-policy RL algorithms in challenging control tasks, especially in highly complex tasks such as Humanoid. We provide theoretical analysis on the variance of the PG estimator, which suggests that our attentively designed unimodal discrete policy can retain a lower variance and yield a stable learning process.
Yuanyang Zhu, Zhi Wang 0001, Yuanheng Zhu, Chunlin Chen 0001, Dongbin Zhao
IEEE Trans. Neural Networks Learn. Syst.4
2024 Learning Cluster-Wise Anchors for Multi-View Clustering
abstract
Due to its effectiveness and efficiency, anchor based multi-view clustering (MVC) has recently attracted much attention. Most existing approaches try to adaptively learn anchors to construct an anchor graph for clustering. However, they generally focus on improving the diversity among anchors by using orthogonal constraint and ignore the underlying semantic relations, which may make the anchors not representative and discriminative enough. To address this problem, we propose an adaptive Cluster-wise Anchor learning based MVC method, CAMVC for short. We first make an anchor cluster assumption that supposes the prior cluster structure of target anchors by pre-defining a consensus cluster indicator matrix. Based on the prior knowledge, an explicit cluster structure of latent anchors is enforced by learning diverse cluster centroids, which can explore both inter-cluster diversity and intra-cluster consistency of anchors, and improve the subspace representation discrimination. Extensive results demonstrate the effectiveness and superiority of our proposed method compared with some state-of-the-art MVC approaches.
Chao Zhang 0078, Xiuyi Jia, Zechao Li, Chunlin Chen 0001, Huaxiong Li
AAAI4
2024 Attention-Guided Contrastive Role Representations for Multi-agent Reinforcement Learning
abstract
Real-world multi-agent tasks usually involve dynamic team composition with the emergence of roles, which should also be a key to efficient cooperation in multi-agent reinforcement learning (MARL). Drawing inspiration from the correlation between roles and agent's behavior patterns, we propose a novel framework of **A**ttention-guided **CO**ntrastive **R**ole representation learning for **M**ARL (**ACORM**) to promote behavior heterogeneity, knowledge transfer, and skillful coordination across agents. First, we introduce mutual information maximization to formalize role representation learning, derive a contrastive learning objective, and concisely approximate the distribution of negative pairs. Second, we leverage an attention mechanism to prompt the global state to attend to learned role representations in value decomposition, implicitly guiding agent coordination in a skillful role space to yield more expressive credit assignment. Experiments on challenging StarCraft II micromanagement and Google research football tasks demonstrate the state-of-the-art performance of our method and its advantages over existing approaches. Our code is available at [https://github.com/NJU-RL/ACORM](https://github.com/NJU-RL/ACORM).
Zican Hu, Zongzhang Zhang, Huaxiong Li, Chunlin Chen 0001, Hongyu Ding, Zhi Wang 0001
ICLR4
2024 Continual Multi-View Clustering with Consistent Anchor Guidance
Chao Zhang 0078, Deng Xu, Xiuyi Jia, Chunlin Chen 0001, Huaxiong Li
IJCAI4
2024 Transparent Projection Networks for Interpretable Image Recognition
abstract
The absence of interpretability in deep convolutional neural networks (CNNs) leaves us with the dilemma of how to explain their decision mechanism in terms of human-understandable semantics due to the black-box nature. In this paper, we present a novel network architecture called Transparent Projection Networks (ProNets) for interpretable image recognition, by constructing a meaningful latent space through transparent projection learning. Specifically, we exploit the expressiveness of NNs to learn a group of input-dependent pixel-level weights, which project input images into the latent space, thus allowing for a transparent weighted connection between raw pixel-level features and latent high-level representations. We organize our layers in light of ResNets with valid modifications, where we propose to develop new block designs to learn projection weights and use shortcuts to enable linear projection flexibly every few layers. Our model inherits the properties of linear models, decomposing the output into a linear combination of contributions from each input feature, which can deliver transparent interpretations of the reasoning process along with high visual quality. Further, as an interpretable model for image classification, through experimental studies on several benchmark datasets, ProNets achieve competitive accuracy results with classic CNNs like VGG and ResNets, while demonstrating a remarkably high level of interpretability.
Chao Zhang 0078, Chunlin Chen 0001, Huaxiong Li
IJCNN3
2024 EASI: Evolutionary Adversarial Simulator Identification for Sim-to-Real Transfer
abstract
Reinforcement Learning (RL) controllers have demonstrated remarkable performance in complex robot control tasks. However, the presence of reality gap often leads to poor performance when deploying policies trained in simulation directly onto real robots. Previous sim-to-real algorithms like Domain Randomization (DR) requires domain-specific expertise and suffers from issues such as reduced control performance and high training costs. In this work, we introduce Evolutionary Adversarial Simulator Identification (EASI), a novel approach that combines Generative Adversarial Network (GAN) and Evolutionary Strategy (ES) to address sim-to-real challenges. Specifically, we consider the problem of sim-to-real as a search problem, where ES acts as a generator in adversarial competition with a neural network discriminator, aiming to find physical parameter distributions that make the state transitions between simulation and reality as similar as possible. The discriminator serves as the fitness function, guiding the evolution of the physical parameter distributions. EASI features simplicity, low cost, and high fidelity, enabling the construction of a more realistic simulator with minimal requirements for real-world data, thus aiding in transferring simulated-trained policies to the real world. We demonstrate the performance of EASI in both sim-to-sim and sim-to-real tasks, showing superior performance compared to existing sim-to-real algorithms.
Huiqiao Fu, Zhehao Zhou, Chunlin Chen 0001
NeurIPS5
2024 Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model Disentanglement
abstract
A longstanding goal of artificial general intelligence is highly capable generalists that can learn from diverse experiences and generalize to unseen tasks. The language and vision communities have seen remarkable progress toward this trend by scaling up transformer-based models trained on massive datasets, while reinforcement learning (RL) agents still suffer from poor generalization capacity under such paradigms. To tackle this challenge, we propose Meta Decision Transformer (Meta-DT), which leverages the sequential modeling ability of the transformer architecture and robust task representation learning via world model disentanglement to achieve efficient generalization in offline meta-RL. We pretrain a context-aware world model to learn a compact task representation, and inject it as a contextual condition to the causal transformer to guide task-oriented sequence generation. Then, we subtly utilize history trajectories generated by the meta-policy as a self-guided prompt to exploit the architectural inductive bias. We select the trajectory segment that yields the largest prediction error on the pretrained world model to construct the prompt, aiming to encode task-specific information complementary to the world model maximally. Notably, the proposed framework eliminates the requirement of any expert demonstration or domain knowledge at test time. Experimental results on MuJoCo and Meta-World benchmarks across various dataset types show that Meta-DT exhibits superior few and zero-shot generalization capacity compared to strong baselines while being more practical with fewer prerequisites. Our code is available at https://github.com/NJU-RL/Meta-DT.
Zhi Wang 0001, Yuanheng Zhu, Dongbin Zhao, Chunlin Chen 0001
NeurIPS6
2024 Quantum Robust Control for Time-Varying Noises Based on Adversarial Learning
abstract
Time-varying noises are one of the reasons that make it difficult for quantum systems to complete control tasks. How to quantify the influence of time-varying noises on control results and how to design a control law that can resist time-varying noises are two important problems. In this paper, the adversarial learning is introduced into quantum control and the loss function under the worst-case noise is used as a way to quantify the impact of time-varying noises on control performance. We utilize the Gradient Ascent Pulse Engineering (GRAPE) technique to search the worst-case noise and meanwhile offer a strategy to improve the robustness of the control law. Simulation experiments on a two-qubit system and a four-qubit system show that the found noises indeed can act as worst-case noises. Furthermore, the optimized control laws demonstrate good robustness to time-varying noises in state preparation tasks.
Haotian Ji, Sen Kuang, Daoyi Dong, Chunlin Chen 0001
SMC4
2024 Dynamic Capacitated Vehicle Routing Problem with Stochastic Requests Using Deep Reinforcement Learning
abstract
With the rapid growth of industries like e-commerce, food delivery, and ride-hailing, research on the Vehicle Routing Problem (VRP) is becoming increasingly relevant. However, most of the research in the field of VRP is based on static delivery tasks. In such tasks, the information about customer and order requests is provided before the delivery vehicle departs from the depot, and it remains constant during delivery. However, in the real world, the delivery tasks are often dynamic, where only some orders are known before the vehicle departs from the depot, and the rest of the orders are disclosed over time [1].
Kaiqiang Tang, Huiqiao Fu, Jiasheng Liu, Guizhou Deng, Yuanyang Lu, Chunlin Chen 0001
SMC6
2024 Multi-level graph regularized robust multi-modal feature selection for Alzheimer's disease classification
Chao Zhang 0078, Wentao Fan 0003, Huaxiong Li, Chunlin Chen 0001
Knowl. Based Syst.4
2024 Learning latent disentangled embeddings and graphs for multi-view clustering
Chao Zhang 0078, Haoxing Chen, Huaxiong Li, Chunlin Chen 0001
Pattern Recognit.4
2024 Efficient Bayesian Policy Reuse With a Scalable Observation Model in Deep Reinforcement Learning
abstract
Bayesian policy reuse (BPR) is a general policy transfer framework for selecting a source policy from an offline library by inferring the task belief based on some observation signals and a trained observation model. In this article, we propose an improved BPR method to achieve more efficient policy transfer in deep reinforcement learning (DRL). First, most BPR algorithms use the episodic return as the observation signal that contains limited information and cannot be obtained until the end of an episode. Instead, we employ the state transition sample, which is informative and instantaneous, as the observation signal for faster and more accurate task inference. Second, BPR algorithms usually require numerous samples to estimate the probability distribution of the tabular-based observation model, which may be expensive and even infeasible to learn and maintain, especially when using the state transition sample as the signal. Hence, we propose a scalable observation model based on fitting state transition functions of source tasks from only a small number of samples, which can generalize to any signals observed in the target task. Moreover, we extend the offline-mode BPR to the continual learning setting by expanding the scalable observation model in a plug-and-play fashion, which can avoid negative transfer when faced with new unknown tasks. Experimental results show that our method can consistently facilitate faster and more efficient policy transfer.
Jinmei Liu, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong
IEEE Trans. Neural Networks Learn. Syst.3
2024 Joint Projection Learning and Tensor Decomposition-Based Incomplete Multiview Clustering
abstract
Incomplete multiview clustering (IMVC) has received increasing attention since it is often that some views of samples are incomplete in reality. Most existing methods learn similarity subgraphs from original incomplete multiview data and seek complete graphs by exploring the incomplete subgraphs of each view for spectral clustering. However, the graphs constructed on the original high-dimensional data may be suboptimal due to feature redundancy and noise. Besides, previous methods generally ignored the graph noise caused by the interclass and intraclass structure variation during the transformation of incomplete graphs and complete graphs. To address these problems, we propose a novel joint projection learning and tensor decomposition (JPLTD)-based method for IMVC. Specifically, to alleviate the influence of redundant features and noise in high-dimensional data, JPLTD introduces an orthogonal projection matrix to project the high-dimensional features into a lower-dimensional space for compact feature learning. Meanwhile, based on the lower-dimensional space, the similarity graphs corresponding to instances of different views are learned, and JPLTD stacks these graphs into a third-order low-rank tensor to explore the high-order correlations across different views. We further consider the graph noise of projected data caused by missing samples and use a tensor-decomposition-based graph filter for robust clustering. JPLTD decomposes the original tensor into an intrinsic tensor and a sparse tensor. The intrinsic tensor models the true data similarities. An effective optimization algorithm is adopted to solve the JPLTD model. Comprehensive experiments on several benchmark datasets demonstrate that JPLTD outperforms the state-of-the-art methods. The code of JPLTD is available at https://github.com/weilvNJU/JPLTD.
Chao Zhang 0078, Huaxiong Li, Xiuyi Jia, Chunlin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Depthwise Convolution for Multi-Agent Communication With Enhanced Mean-Field Approximation
abstract
Multi-Agent settings remain a fundamental challenge in the reinforcement learning (RL) domain due to the partial observability and the lack of accurate real-time interactions across agents. In this article, we propose a new method based on local communication learning to tackle the multi-agent RL (MARL) challenge within a large number of agents coexisting. First, we design a new communication protocol that exploits the ability of depthwise convolution to efficiently extract local relations and learn local communication between neighboring agents. To facilitate multi-agent coordination, we explicitly learn the effect of joint actions by taking the policies of neighboring agents as inputs. Second, we introduce the mean-field approximation into our method to reduce the scale of agent interactions. To more effectively coordinate behaviors of neighboring agents, we enhance the mean-field approximation by a supervised policy rectification network (PRN) for rectifying real-time agent interactions and by a learnable compensation term for correcting the approximation bias. The proposed method enables efficient coordination as well as outperforms several baseline approaches on the adaptive traffic signal control (ATSC) task and the StarCraft II multi-agent challenge (SMAC).
Donghan Xie, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong
IEEE Trans. Neural Networks Learn. Syst.3
2024 Hamiltonian Identification via Quantum Ensemble Classification
abstract
Identifying the Hamiltonian of an unknown quantum system is a critical task in the area of quantum information. In this article, we propose a systematic Hamiltonian identification approach via quantum ensemble multiclass classification (HI-QEMC). This approach is implemented by a three-step iterative refining process, i.e., parameter interval guess, verification, and judgment. In the parameter interval guess step, the parameter interval is divided into several sub-intervals and the true Hamiltonian parameter is guessed in one of them. In the parameter interval verification step, cross verification is applied to verify the accuracy of the guess. In the parameter interval judgment step, an adaptive interval judgment (AIJ) algorithm is designed to determine the sub-interval containing the true Hamiltonian parameter. Numerical results on two typical quantum systems, i.e., two-level quantum systems and three-level quantum systems, demonstrate the effectiveness and superior performance of the proposed approach for quantum Hamiltonian identification.
Haixu Yu, Xudong Zhao 0001, Daoyi Dong, Chunlin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Low-Rank Tensor Regularized Views Recovery for Incomplete Multiview Clustering
abstract
In real applications, it is often that the collected multiview data contain missing views. Most existing incomplete multiview clustering (IMVC) methods cannot fully utilize the underlying information of missing data or sufficiently explore the consistent and complementary characteristics. In this article, we propose a novel Low-rAnk Tensor regularized viEws Recovery (LATER) method for IMVC, which jointly reconstructs and utilizes the missing views and learns multilevel graphs for comprehensive similarity discovery in a unified model. The missing views are recovered from a common latent representation, and the recovered views conversely improve the learning of shared patterns. Based on the shared subspace representations and recovered complete multiview data, the multilevel graphs are learned by self-representation to fully exploit the consistent and complementary information among views. Besides, a tensor nuclear norm regularizer is introduced to pursue the global low-rank property and explore the interview correlations. An alternating direction minimization algorithm is presented to optimize the proposed model. Moreover, a new initialization method is proposed to promote the effectiveness of our method for latent representation learning and missing data recovery. Extensive experiments demonstrate that our method outperforms the state-of-the-art approaches.
Chao Zhang 0078, Huaxiong Li, Caihua Chen, Xiuyi Jia, Chunlin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Enhanced Tensor Low-Rank and Sparse Representation Recovery for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) has attracted remarkable attention due to the emergence of multi-view data with missing views in real applications. Recent methods attempt to recover the missing information to address the IMVC problem. However, they generally cannot fully explore the underlying properties and correlations of data similarities across views. This paper proposes a novel Enhanced Tensor Low-rank and Sparse Representation Recovery (ETLSRR) method, which reformulates the IMVC problem as a joint incomplete similarity graphs learning and complete tensor representation recovery problem. Specifically, ETLSRR learns the intra-view similarity graphs and constructs a 3-way tensor by stacking the graphs to explore the inter-view correlations. To alleviate the negative influence of missing views and data noise, ETLSRR decomposes the tensor into two parts: a sparse tensor and an intrinsic tensor, which models the noise and underlying true data similarities, respectively. Both global low-rank and local structured sparse characteristics of the intrinsic tensor are considered, which enhances the discrimination of similarity matrix. Moreover, instead of using the convex tensor nuclear norm, ETLSRR introduces a generalized non-convex tensor low-rank regularization to alleviate the biased approximation. Experiments on several datasets demonstrate the effectiveness of our method compared with the state-of-the-art methods.
Chao Zhang 0078, Huaxiong Li, Zizheng Huang, Yang Gao 0001, Chunlin Chen 0001
AAAI6
2023 BiERL: A Meta Evolutionary Reinforcement Learning Framework via Bilevel Optimization
abstract
Evolutionary reinforcement learning (ERL) algorithms recently raise attention in tackling complex reinforcement learning (RL) problems due to high parallelism, while they are prone to insufficient exploration or model collapse without carefully tuning hyperparameters (aka meta-parameters). In the paper, we propose a general meta ERL framework via bilevel optimization (BiERL) to jointly update hyperparameters in parallel to training the ERL model within a single agent, which relieves the need for prior domain knowledge or costly optimization procedure before model deployment. We design an elegant meta-level architecture that embeds the inner-level’s evolving experience into an informative population representation and introduce a simple and feasible evaluation of the meta-level fitness function to facilitate learning efficiency. We perform extensive experiments in MuJoCo and Box2D tasks to verify that as a general framework, BiERL outperforms various baselines and consistently improves the learning performance for a diversity of ERL algorithms.
Yuanyang Zhu, Zhi Wang 0001, Yan Zheng 0002, Jianye Hao, Chunlin Chen 0001
ECAI6
2023 Model-Aware Contrastive Learning: Towards Escaping the Dilemmas
abstract
Contrastive learning (CL) continuously achieves significant breakthroughs across multiple domains. However, the most common InfoNCE-based methods suffer from some dilemmas, such as uniformity-tolerance dilemma (UTD) and gradient reduction, both of which are related to a $\mathcal{P}_{ij}$ term. It has been identified that UTD can lead to unexpected performance degradation. We argue that the fixity of temperature is to blame for UTD. To tackle this challenge, we enrich the CL loss family by presenting a Model-Aware Contrastive Learning (MACL) strategy, whose temperature is adaptive to the magnitude of alignment that reflects the basic confidence of the instance discrimination task, then enables CL loss to adjust the penalty strength for hard negatives adaptively. Regarding another dilemma, the gradient reduction issue, we derive the limits of an involved gradient scaling factor, which allows us to explain from a unified perspective why some recent approaches are effective with fewer negative samples, and summarily present a gradient reweighting to escape this dilemma. Extensive remarkable empirical results in vision, sentence, and graph modality validate our approach’s general improvement for representation learning and downstream tasks.
Zizheng Huang, Haoxing Chen, Ziqi Wen, Chao Zhang 0078, Huaxiong Li, Bo Wang 0027, Chunlin Chen 0001
ICML7
2023 NA2Q: Neural Attention Additive Model for Interpretable Multi-Agent Q-Learning
Zichuan Liu, Yuanyang Zhu, Chunlin Chen 0001
ICML3
2023 Robust Spectral Embedding Completion Based Incomplete Multi-view Clustering
abstract
Graph based methods have been widely used in incomplete multi-view clustering (IMVC). Most recent methods try to fill the original missing samples or incomplete affinity matrices to obtain a complete similarity graph for the subsequent spectral clustering. However, recovering the original high-dimensional data or complete n X n similarity matrix is usually time-consuming and noise-sensitive. Besides, they generally separate the cluster indicator learning into an individual step, which may result in sub-optimal graphs or spectral embeddings for clustering. To address these problems, this paper proposes a robust Spectral Embedding Completion based IMVC (SEC-IMVC) method, which incorporates spectral embedding completion and discrete cluster indicator learning into a unified framework. SEC-IMVC performs completion on spectral embeddings, and the embedding noise is eliminated to reduce the negative influence of original data noise. The discrete cluster indicator matrix is seamlessly learned by using spectral rotation, and it can explore the first-order feature consistency among different views. To further improve the completion robustness, the second-order correlation consistency is also captured by pairwise relations alignment. We compare our method with some state-of-the-art approaches on several datasets, and the experimental results show the effectiveness and advantages of our method.
Chao Zhang 0078, Jingwen Wei, Bo Wang 0027, Zechao Li, Chunlin Chen 0001, Huaxiong Li
ACM Multimedia5
2023 Ess-InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations
abstract
Imitation learning aims to reproduce expert behaviors without relying on an explicit reward signal. However, real-world demonstrations often present challenges, such as multi-modal, data imbalance, and expensive labeling processes. In this work, we propose a novel semi-supervised imitation learning architecture that learns disentangled behavior representations from imbalanced demonstrations using limited labeled data. Specifically, our method consists of three key components. First, we adapt the concept of semi-supervised generative adversarial networks to the imitation learning context. Second, we employ a learnable latent distribution to align the generated and expert data distributions. Finally, we utilize a regularized information maximization approach in conjunction with an approximate label prior to further improve the semi-supervised learning performance. Experimental results demonstrate the efficiency of our method in learning multi-modal behaviors from imbalanced demonstrations compared to baseline methods.
Huiqiao Fu, Kaiqiang Tang, Yuanyang Lu, Yiming Qi, Guizhou Deng, Flood Sung, Chunlin Chen 0001
NeurIPS7
2023 Rolling horizon wind-thermal unit commitment optimization based on deep reinforcement learning
Jinhao Shi, Bo Wang 0027, Ran Yuan, Zhi Wang 0001, Chunlin Chen 0001, Junzo Watada
Appl. Intell.5
2023 Sparse spatial transformers for few-shot learning
Haoxing Chen, Huaxiong Li, Chunlin Chen 0001
Sci. China Inf. Sci.4
2023 A robust mixed error coding method based on nonconvex sparse representation
Chao Zhang 0078, Huaxiong Li, Bo Wang 0027, Chunlin Chen 0001
Inf. Sci.5
2023 Extracting Decision Tree From Trained Deep Reinforcement Learning in Traffic Signal Control
abstract
Deep reinforcement learning (DRL) has achieved impressive success in traffic signal control systems (TSCS). However, since a key component of many DRL models is the complex deep neural networks (DNNs), it hinders humans or experts from understanding and explaining the learned policy. Recently, many works have focused on developing interpretable techniques to compress or distill complex DNNs into smaller, faster, or more understandable models. The decision trees (DTs) are viewed as the de facto technique for interpretable and transparent machine learning, and can provide an easy-understanding decision path from the root to the leaf node. In this work, we utilize modified DTs to extract models with simpler hierarchical structures from premium policy achieved by DRL methods. First, we use a DRL algorithm to learn a premium policy for traffic signal control. Then, we collect a dataset with the learned premium policy by interacting with the environment. Finally, we extract the DTs with the collected dataset of state-actions pairs. We evaluate our method on a Simulation of Urban Mobility simulator in a simulation way. Simulation results show that the extracted DTs can generate human-understandable decision processes and provide explicit knowledge from the DNNs reference of the DRL.
Yuanyang Zhu, Chunlin Chen 0001
IEEE Trans. Comput. Soc. Syst.3
2023 Multi-Level Cascade Sparse Representation Learning for Small Data Classification
abstract
Deep learning (DL) methods have recently captured much attention for image classification. However, such methods may lead to a suboptimal solution for small-scale data since the lack of training samples. Sparse representation stands out with its efficiency and interpretability, but its precision is not so competitive. We develop a Multi-Level Cascade Sparse Representation (ML-CSR) learning method to combine both advantages when processing small-scale data. ML-CSR is proposed using a pyramid structure to expand the training data size. It adopts two core modules, the Error-To-Feature (ETF) module, and the Generate-Adaptive-Weight (GAW) module, to further improve the precision. ML-CSR calculates the inter-layer differences by the ETF module to increase the diversity of samples and obtains adaptive weights based on the layer accuracy in the GAW module. This helps ML-CSR learn more discriminative features. State-of-the-art results on the benchmark face databases validate the effectiveness of the proposed ML-CSR. Ablation experiments demonstrate that the proposed pyramid structure, ETF, and GAW module can improve the performance of ML-CSR. The code is available athttps://github.com/Zhongwenyuan98/ML-CSR.
Wenyuan Zhong, Huaxiong Li, Qinghua Hu, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 A Dirichlet Process Mixture of Robust Task Models for Scalable Lifelong Reinforcement Learning
abstract
While reinforcement learning (RL) algorithms are achieving state-of-the-art performance in various challenging tasks, they can easily encounter catastrophic forgetting or interference when faced with lifelong streaming information. In this article, we propose a scalable lifelong RL method that dynamically expands the network capacity to accommodate new knowledge while preventing past memories from being perturbed. We use a Dirichlet process mixture to model the nonstationary task distribution, which captures task relatedness by estimating the likelihood of task-to-cluster assignments and clusters the task models in a latent space. We formulate the prior distribution of the mixture as a Chinese restaurant process (CRP) that instantiates new mixture components as needed. The update and expansion of the mixture are governed by the Bayesian nonparametric framework with an expectation maximization (EM) procedure, which dynamically adapts the model complexity without explicit task boundaries or heuristics. Moreover, we use the domain randomization technique to train robust prior parameters for the initialization of each task model in the mixture; thus, the resulting model can better generalize and adapt to unseen tasks. With extensive experiments conducted on robot navigation and locomotion domains, we show that our method successfully facilitates scalable lifelong RL and outperforms relevant existing methods.
Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong
IEEE Trans. Cybern.2
2023 Hierarchical Free Gait Motion Planning for Hexapod Robots Using Deep Reinforcement Learning
abstract
This paper addresses the problem of legged locomotion in unstructured environments, and a novel Hierarchical multi-contact motion planning method for hexapod robots is proposed by combining Free Gait motion planning and Deep Reinforcement Learning (HFG-DRL). We structurally decompose the complex free gait multi-contact motion planning task into path planning in discrete state space and gait planning in continuous state space. Firstly, the Soft Deep Q-Network (SDQN) is used to obtain the global prior path information in the Path Planner (PP). Secondly, a Free Gait Planner (FGP) is proposed to obtain the gait sequence. Finally, based on the PP and the FGP, the Center-of-Mass (CoM) sequence is generated by the trained optimal policy using the designed Deep Reinforcement Learning (DRL) algorithm. Experimental results in different environments demonstrate the feasibility, effectiveness, and advancement of the proposed method. Videos are shown athttp://www.hexapod.cn/hfg-drl.html.
Xinpeng Wang 0006, Huiqiao Fu, Guizhou Deng, Canghai Liu, Kaiqiang Tang, Chunlin Chen 0001
IEEE Trans. Ind. Informatics6
2023 Adaptive Label Correlation Based Asymmetric Discrete Hashing for Cross-Modal Retrieval
abstract
Hashing methods have captured much attention for cross-modal retrieval in recent years. Most existing approaches mainly focus on preserving the semantic similarity across heterogeneous modalities in a shared Hamming subspace, while the label information and potential correlations of multi-label semantics are not fully excavated. In this article, a novel Adaptive Label correlation based asymmEtric Cross-modal Hashing method, i.e., ALECH, is proposed for cross-modal retrieval. ALECH decomposes hash learning into two steps, hash codes learning and hash functions learning. For hash codes learning, the high-order semantic label correlations are adaptively exploited to guide the latent feature learning, while simultaneously generating the binary codes in a discrete manner. The asymmetric strategy is utilized to connect the latent feature space and Hamming space, and preserve the pairwise semantic similarity. Different from other two-step methods that directly adopt simple least-squares regression to learn hash functions based on binary codes, ALECH leverages both hash codes and semantic labels for hash functions learning which further preserves the similarity. Experiments on several benchmark datasets demonstrate that the proposed ALECH method outperforms the state-of-the-art cross-hashing methods.
Huaxiong Li, Chao Zhang 0078, Xiuyi Jia, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Knowl. Data Eng.5
2023 Weakly-Supervised Enhanced Semantic-Aware Hashing for Cross-Modal Retrieval
abstract
Owing to its query and storage efficiency, hash learning has sparked much interest for Cross-Modal Retrieval (CMR) task. Previous literatures have proved the superiority of supervised Cross-Modal Hashing (CMH) methods over unsupervised ones. Nevertheless, most existing supervised CMH methods still suffer from some limitations: 1) it is assumed that the observed labels of training data are complete and accurate, which may be impractical due to the missing and wrong class assignments in real applications, and 2) the semantic information is not fully excavated, especially for the semantic correlations among labels. To address these issues, this paper proposes a Weakly-supervised enhAnced Semantic-aware Hashing (WASH) method which simultaneously estimates the label noises and performs enhanced semantic-aware hash learning. WASH employs the low-rank and sparse decomposition to alleviate the label noises, and a high-level semantic factor as well as a semantic correlation matrix is obtained by low-rank factorization on the noise-reduced labels. The low-rank semantic factors and multi-modal features are jointly factorized into a common subspace to reduce the heterogeneity gaps, so as to enhance the semantic awareness of shared representation. In this way, the hash codes can be obtained by binarizing the shared representation with pairwise semantic similarity preserved. Experiments on several benchmark datasets verify the effectiveness of the proposed method in comparison with the state-of-the-art CMH approaches.
Chao Zhang 0078, Huaxiong Li, Yang Gao 0001, Chunlin Chen 0001
IEEE Trans. Knowl. Data Eng.4
2023 Adaptive Marginalized Semantic Hashing for Unpaired Cross-Modal Retrieval
abstract
In recent years, Cross-Modal Hashing (CMH) has attracted much attention due to its fast query speed and efficient storage. Previous studies have achieved promising results for Cross-Modal Retrieval (CMR) by discovering discriminative hash codes and modality-specific hash functions. Nonetheless, most existing CMR works are subjected to some restrictions: 1) It is assumed that data of different modalities are fully paired, which is impractical in real applications due to sample missing and false data alignment, and 2) binary regression targets including the label matrix and binary codes are too rigid to effectively learn semantic-preserving hash codes and hash functions. To address these problems, this paper proposes an Adaptive Marginalized Semantic Hashing (AMSH) method which not only enhances the discrimination of latent representations and hash codes by adaptive margins, but can also be used for both paired and unpaired CMR. As a two-step method, in the first step, AMSH generates semantic-aware modality-specific latent representations with adaptively marginalized labels, thereby enlarging the distances between different classes, and exploiting the labels to preserve the inter-modal and intra-modal semantic similarities into latent representations and hash codes. In the second step, adaptive margin matrices are embedded into the hash codes, and enlarge the gaps between positive and negative bits, which improves the discrimination and robustness of hash functions. On this basis, AMSH generates similarity-preserving hash codes and robust hash functions without the strict one-to-one data correspondence requirement. Experiments are conducted on several benchmark datasets to demonstrate the superiority and flexibility of AMSH over some state-of-the-art CMR methods. The source code is available athttps://github.com/LKYLKYZ/AMSH.
Kaiyi Luo, Chao Zhang 0078, Huaxiong Li, Xiuyi Jia, Chunlin Chen 0001
IEEE Trans. Multim.5
2023 Curriculum-Based Deep Reinforcement Learning for Quantum Control
abstract
Deep reinforcement learning (DRL) has been recognized as an efficient technique to design optimal strategies for different complex systems without prior knowledge of the control landscape. To achieve a fast and precise control for quantum systems, we propose a novel DRL approach by constructing a curriculum consisting of a set of intermediate tasks defined by fidelity thresholds, where the tasks among a curriculum can be statically determined before the learning process or dynamically generated during the learning process. By transferring knowledge between two successive tasks and sequencing tasks according to their difficulties, the proposed curriculum-based DRL (CDRL) method enables the agent to focus on easy tasks in the early stage, then move onto difficult tasks, and eventually approaches the final task. Numerical comparison with the traditional methods [gradient method (GD), genetic algorithm (GA), and several other DRL methods] demonstrates that CDRL exhibits improved control performance for quantum systems and also provides an efficient way to identify optimal strategies with few control pulses.
Hailan Ma, Daoyi Dong, Steven X. Ding, Chunlin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Instance Weighted Incremental Evolution Strategies for Reinforcement Learning in Dynamic Environments
abstract
Evolution strategies (ESs), as a family of black-box optimization algorithms, recently emerge as a scalable alternative to reinforcement learning (RL) approaches such as Q-learning or policy gradient and are much faster when many central processing units (CPUs) are available due to better parallelization. In this article, we propose a systematic incremental learning method for ES in dynamic environments. The goal is to adjust previously learned policy to a new one incrementally whenever the environment changes. We incorporate an instance weighting mechanism with ES to facilitate its learning adaptation while retaining scalability of ES. During parameter updating, higher weights are assigned to instances that contain more new knowledge, thus encouraging the search distribution to move toward new promising areas of parameter space. We propose two easy-to-implement metrics to calculate the weights: instance novelty and instance quality. Instance novelty measures an instance's difference from the previous optimum in the original environment, while instance quality corresponds to how well an instance performs in the new environment. The resulting algorithm, instance weighted incremental evolution strategies (IW-IESs), is verified to achieve significantly improved performance on challenging RL tasks ranging from robot navigation to locomotion. This article thus introduces a family of scalable ES algorithms for RL domains that enables rapid learning adaptation to dynamic environments.
Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong
IEEE Trans. Neural Networks Learn. Syst.2
2022 Multi-level Metric Learning for Few-Shot Image Recognition
Haoxing Chen, Huaxiong Li, Chunlin Chen 0001
ICANN (1)4
2022 Multi-Scale Adaptive Task Attention Network for Few-Shot Learning
abstract
Few-shot learning has aroused considerable interest in recent years, which aims to recognize unseen categories by using a few labeled samples. In various few-shot methods, pixel-level metric-learning based methods have achieved promising performance. However, most of these methods deal with each category in the support set independently, which may be insufficient to measure the relations among features, especially in a specific task. Besides, the coexistence of dominant objects at different scales may degrade the performance of these methods. To address these issues, a novel Multi-Scale Adaptive Task Attention Network, MATANet for short, is proposed for few-shot learning. In MATANet, a multi-scale feature generator is first constructed to extract the image features at different scales. Then, an adaptive task attention module is built to select the most important local representations among the entire task. Finally, a similarity-to-class module is adapted to measure the similarities between query and support set. Extensive experiments on popular benchmarks show the effectiveness of the proposed MATANet compared with state-of-the-art methods. Our source code is available at: https://github.com/chenhaoxing/MATANet.
Haoxing Chen, Huaxiong Li, Chunlin Chen 0001
ICPR4
2022 Fast Probabilistic Policy Reuse via Reward Function Fitting
abstract
Transfer learning has shown great potential to accelerate reinforcement learning (RL) by utilizing prior knowledge of relevant task that has been learned in the past. Policy Reuse Q-learning (PRQL) is a general policy transfer framework, which speeds up the learning process of the target task by probabilistically reusing source policies from the policy library. In this paper, we propose an improved PRQL method to achieve more fast probabilistic policy reuse in deep reinforcement learning (DRL). First, we extend the basic PRQL algorithm to DRL, proposing a probability policy reuse algorithm that builds on DRL to solve more complex problems. Second, PRQL algorithms usually use a metric based on the average gain to measure the similarity between tasks. However, it contains very limited information and must be delayed until the end of an episode to update, which is inefficient. Instead, we propose a new metric based on fitting the reward function, which can make the agent converge to the most suitable reuse policy more quickly and accurately. We demonstrate the detection accuracy, received cumulative reward, and speed of convergence of our method in three complex Markov tasks. Experimental results show that our method can consistently achieve efficient policy transfer in these tasks.
Jinmei Liu, Zhi Wang 0001, Chunlin Chen 0001
IJCNN3
2022 HRL2E: Hierarchical Reinforcement Learning with Low-level Ensemble
abstract
Goal-conditioned hierarchical reinforcement learning (HRL) is a promising approach to solve challenging tasks with sparse rewards and long horizons. However, it suffers from the non-stationary problem due to the updating and unstable low level. To stabilize the low level more quickly and accelerate the non-stationary stage, we propose a novel HRL method: Hierarchical Reinforcement Learning with Low-level Ensemble (HRL2E). In HRL2E, the high level generates goals as high-level actions based on current states. Then the low level made up of several homogeneous policies attempts to complete these goals within a specific timestep budget. The improvement of our approach to the general goal-conditioned HRL algorithms can be summarized in two aspects. First, we estimate the target value function with the ensemble, stabilizing the training process. Second, we propose the Gates module composed of several scoring machines to score each low-level policy and judge which one has the most success potential to execute a specific goal. We adopt Twin Delayed Deep Deterministic Policy Gradient (TD3) in each level. Experimental comparison between our method and state-of-the-art goal-conditioned HRL methods on challenging continuous control tasks in MuJoCo domains shows our method can significantly accelerate training.
You Qin, Zhi Wang 0001, Chunlin Chen 0001
IJCNN3
2022 Enhanced Group Sparse Regularized Nonconvex Regression for Face Recognition
abstract
Regression analysis based methods have shown strong robustness and achieved great success in face recognition. In these methods, convex$l_1$-norm and nuclear norm are usually utilized to approximate the$l_0$-norm and rank function. However, such convex relaxations may introduce a bias and lead to a suboptimal solution. In this paper, we propose a novel Enhanced Group Sparse regularized Nonconvex Regression (EGSNR) method for robust face recognition. An upper bounded nonconvex function is introduced to replace$l_1$-norm for sparsity, which alleviates the bias problem and adverse effects caused by outliers. To capture the characteristics of complex errors, we propose a mixed model by combining$\gamma$-norm and matrix$\gamma$-norm induced from the nonconvex function. Furthermore, an$l_{2,\gamma }$-norm based regularizer is designed to directly seek the interclass sparsity or group sparsity instead of traditional$l_{2,1}$-norm. The locality of data, i.e., the distance between the query sample and multi-subspaces, is also taken into consideration. This enhanced group sparse regularizer enables EGSNR to learn more discriminative representation coefficients. Comprehensive experiments on several popular face datasets demonstrate that the proposed EGSNR outperforms the state-of-the-art regression based methods for robust face recognition.
Chao Zhang 0078, Huaxiong Li, Chunlin Chen 0001, Xianzhong Zhou
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Shaping Visual Representations With Attributes for Few-Shot Recognition
abstract
Few-shot recognition aims to recognize novel categories under low-data regimes. Some recent few-shot recognition methods introduce auxiliary semantic modality, i.e., category attribute information, into representation learning, which enhances the feature discrimination and improves the recognition performance. Most of these existing methods only consider the attribute information of support set while ignoring the query set, resulting in a potential loss of performance. In this letter, we propose a novel attribute-shaped learning (ASL) framework, which can jointly perform query attributes generation and discriminative visual representation learning for few-shot recognition. Specifically, a visual-attribute predictor (VAP) is constructed to predict the attributes of queries. By leveraging the attributes information, an attribute-visual attention module (AVAM) is designed, which can adaptively utilize attributes and visual representations to learn more discriminative features. Under the guidance of attribute modality, our method can learn enhanced semantic-aware representation for classification. Experiments demonstrate that our method can achieve competitive results on CUB and SUN benchmarks. Our source code is available at:https://github.com/chenhaoxing/ASL.
Haoxing Chen, Huaxiong Li, Chunlin Chen 0001
IEEE Signal Process. Lett.4
2022 Deep Reinforcement Learning With Quantum-Inspired Experience Replay
abstract
In this article, a novel training paradigm inspired by quantum computation is proposed for deep reinforcement learning (DRL) with experience replay. In contrast to the traditional experience replay mechanism in DRL, the proposed DRL with quantum-inspired experience replay (DRL-QER) adaptively chooses experiences from the replay buffer according to the complexity and the replayed times of each experience (also called transition), to achieve a balance between exploration and exploitation. In DRL-QER, transitions are first formulated in quantum representations and then the preparation operation and depreciation operation are performed on the transitions. In this process, the preparation operation reflects the relationship between the temporal-difference errors (TD-errors) and the importance of the experiences, while the depreciation operation is taken into account to ensure the diversity of the transitions. The experimental results on Atari 2600 games show that DRL-QER outperforms state-of-the-art algorithms, such as DRL-PER and DCRL on most of these games with improved training efficiency and is also applicable to such memory-based DRL approaches as double network and dueling network.
Hailan Ma, Chunlin Chen 0001, Daoyi Dong
IEEE Trans. Cybern.3
2022 Lifelong Incremental Reinforcement Learning With Online Bayesian Inference
abstract
A central capability of a long-lived reinforcement learning (RL) agent is to incrementally adapt its behavior as its environment changes and to incrementally build upon previous experiences to facilitate future learning in real-world scenarios. In this article, we propose lifelong incremental reinforcement learning (LLIRL), a new incremental algorithm for efficient lifelong adaptation to dynamic environments. We develop and maintain a library that contains an infinite mixture of parameterized environment models, which is equivalent to clustering environment parameters in a latent space. The prior distribution over the mixture is formulated as a Chinese restaurant process (CRP), which incrementally instantiates new environment models without any external information to signal environmental changes in advance. During lifelong learning, we employ the expectation-maximization (EM) algorithm with online Bayesian inference to update the mixture in a fully incremental manner. In EM, the E-step involves estimating the posterior expectation of environment-to-cluster assignments, whereas the M-step updates the environment parameters for future learning. This method allows for all environment models to be adapted as necessary, with new models instantiated for environmental changes and old models retrieved when previously seen environments are encountered again. Simulation experiments demonstrate that LLIRL outperforms relevant existing methods and enables effective incremental adaptation to various dynamic environments for lifelong learning.
Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong
IEEE Trans. Neural Networks Learn. Syst.2
2022 Locality-Constrained Discriminative Matrix Regression for Robust Face Identification
abstract
Regression-based methods have been widely applied in face identification, which attempts to approximately represent a query sample as a linear combination of all training samples. Recently, a matrix regression model based on nuclear norm has been proposed and shown strong robustness to structural noises. However, it may ignore two important issues: the label information and local relationship of data. In this article, a novel robust representation method called locality-constrained discriminative matrix regression (LDMR) is proposed, which takes label information and locality structure into account. Instead of focusing on the representation coefficients, LDMR directly imposes constraints on representation components by fully considering the label information, which has a closer connection to identification process. The locality structure characterized by subspace distances is used to learn class weights, and the correct class is forced to make more contribution to representation. Furthermore, the class weights are also incorporated into a competitive constraint on the representation components, which reduces the pairwise correlations between different classes and enhances the competitive relationships among all classes. An iterative optimization algorithm is presented to solve LDMR. Experiments on several benchmark data sets demonstrate that LDMR outperforms some state-of-the-art regression-based methods.
Chao Zhang 0078, Huaxiong Li, Chunlin Chen 0001, Xianzhong Zhou
IEEE Trans. Neural Networks Learn. Syst.4
2022 Perspective-Corrected Spatial Referring Expression Generation for Human-Robot Interaction
abstract
Intelligent robots designed to interact with humans in real scenarios need to be able to refer to entities actively by natural language. In spatial referring expression generation (REG), the ambiguity is unavoidable due to the diversity of reference frames, which will lead to an understanding gap between humans and robots. To narrow this gap, in this article, we propose a novel perspective-corrected spatial REG (PcSREG) approach for human–robot interaction (HRI) by considering the selection of reference frames. The task of REG is simplified into the process of generating diverse spatial relation units. First, we pick out all landmarks in these spatial relation units according to the entropy of preference and allow its updating through a stack model. Then, all possible referring expressions are generated according to different reference frame strategies. Finally, we evaluate every expression using a probabilistic referring expression resolution model and find the best expression that satisfies both of the appropriateness and effectiveness. We implement the proposed approach on a robot system and empirical experimental results show that our approach can generate more effective spatial referring expressions for practical applications.
Mingjiang Liu, Chengli Xiao, Chunlin Chen 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod Robots
abstract
Legged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of hexapod robots is formulated as a Markov Decision Process (MDP) with a specified reward function. Second, a transition feasibility model is proposed for hexapod robots, which describes the feasibility of the state transition under the condition of satisfying kinematics and dynamics, and in turn determines the rewards. Third, the footholds and Center-of-Mass (CoM) sequences are sampled from a diagonal Gaussian distribution and the sequences are optimized through learning the optimal policies using the designed DRL algorithm. Both of the simulation and experimental results on physical systems demonstrate the feasibility and efficiency of the proposed method. Videos are shown at https://videoviewpage.wixsite.com/mcrl.
Huiqiao Fu, Kaiqiang Tang, Peng Li 0031, Wenqi Zhang 0001, Xinpeng Wang 0006, Guizhou Deng, Tao Wang 0004, Chunlin Chen 0001
IJCAI8
2021 Local Mutual Metric Network for Few-Shot Image Classification
Huaxiong Li, Haoxing Chen, Chunlin Chen 0001
PRCV (1)4
2021 A Guided Differential Evolution Algorithm for Control Design of Quantum Gates
abstract
Constructing high fidelity quantum gates is a key task in quantum information technology. Gradient-based methods (e.g., GRAPE) have been widely used for numerous implementations of quantum gate control tasks but are not practical for in-situ laboratory experiments. In this paper, for the design of optimal control fields for constructing quantum gates, we propose a Guided Differential Evolution (GDE) algorithm. Numerical results verify the effectiveness of the proposed GDE algorithm for the control design of quantum gates with short evolution time and few control pulses.
Shouliang Hu, Hailan Ma, Chunlin Chen 0001
SMC3
2021 Pairwise Relations Oriented Discriminative Regression
abstract
Linear Regression (LR) is a popular and effective technique in pattern recognition area, which aims to find a transform matrix between source data and target data (usually label matrix). However, a binary zero-one label matrix may be too strict and inappropriate for regression. Besides, directly projecting source data to target data by one transform matrix may lose some intrinsic data information. To address these issues, this paper proposes a novel Pairwise Relations oriented Discriminative Regression (PRDR) method. In PRDR, the source data is regressed into a latent space instead of label space. To supervise the discriminative projection learning, the pairwise relations in source data space and label space are exploited in the latent space simultaneously. The pairwise label relations are transferred into the latent subspace by solving a distance-distance difference minimization problem, and the intraclass instance relations are also preserved in latent space. These two constraints ensure the pairwise similarity of data points after transformation which is beneficial for classification. By further enlarging the margins between true and false classes, PRDR is extended to a robust version, i.e., R-PRDR. An efficient algorithm is presented to solve the PRDR model. Extensive experiments on several popular image datasets demonstrate the effectiveness and efficiency of the proposed method compared with some state-of-the-art regression approaches.
Chao Zhang 0078, Huaxiong Li, Chunlin Chen 0001, Yang Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 IEDQN: Information Exchange DQN with a Centralized Coordinator for Traffic Signal Control
abstract
Finding the optimal control strategy for traffic signals, especially for multi-intersection traffic signals, is still a difficult task. The use of reinforcement learning (RL) algorithms to this problem is greatly limited because of the partially observable and nonstationary environment. In this paper, we study how to eliminate the above influence from the environment through communication among agents. The proposed method, called Information Exchange Deep Q-Network (IEDQN), has a learning communication protocol, which makes each local agent pay unbalanced and asymmetric attention to other agents' information. Besides the protocol, each agent has the ability to abstract local information from its own history data for interacting, which means that the communication can avoid the dependent instant information and it is robust to the potential time delay of communication. Specifically, by alleviating the effects of partial observation, experience replay can recover to good performance. We evaluate IEDQN via simulation experiments in the simulation of urban mobility (SUMO) in a traffic grid, and it outperforms the comparative multi-agent RL (MARL) methods in both efficiency and effectiveness.
Donghan Xie, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong
IJCNN3
2020 Several developments in learning control of quantum systems
abstract
This paper summarizes several recent achievements in the area of learning control of quantum systems and draw several new directions for future research. Three learning algorithms including gradient method, differential evolution and reinforcement learning are introduced for quantum control. Quantum state control in closed and open quantum systems is analyzed, where gradient method and differential evolution are employed, respectively. The approach of deep reinforcement learning for quantum gate control is introduced, and a sampling-based learning control is illustrated for robust control of quantum gates.
Hailan Ma, Chunlin Chen 0001
SMC2
2020 Learning-Based Quantum Robust Control: Algorithm, Applications, and Experiments
abstract
Robust control design for quantum systems has been recognized as a key task in quantum information technology, molecular chemistry, and atomic physics. In this paper, an improved differential evolution algorithm, referred to as multiple-samples and mixed-strategy DE (msMS_DE), is proposed to search robust fields for various quantum control problems. In msMS_DE, multiple samples are used for fitness evaluation and a mixed strategy is employed for the mutation operation. In particular, the msMS_DE algorithm is applied to the control problems of: 1) open inhomogeneous quantum ensembles and 2) the consensus goal of a quantum network with uncertainties. Numerical results are presented to demonstrate the excellent performance of the improved machine learning algorithm for these two classes of quantum robust control problems. Furthermore, msMS_DE is experimentally implemented on femtosecond (fs) laser control applications to optimize two-photon absorption and control fragmentation of the molecule CH2BrI. The experimental results demonstrate the excellent performance of msMS_DE in searching for effective fs laser pulses for various tasks.
Daoyi Dong, Xi Xing, Hailan Ma, Chunlin Chen 0001, Zhixin Liu 0003, Herschel Rabitz
IEEE Trans. Cybern.4
2020 Reinforcement Learning-Based Optimal Sensor Placement for Spatiotemporal Modeling
abstract
A reinforcement learning-based method is proposed for optimal sensor placement in the spatial domain for modeling distributed parameter systems (DPSs). First, a low-dimensional subspace, derived by Karhunen-Loève decomposition, is identified to capture the dominant dynamic features of the DPS. Second, a spatial objective function is proposed for the sensor placement. This function is defined in the obtained low-dimensional subspace by exploiting the time-space separation property of distributed processes, and in turn aims at minimizing the modeling error over the entire time and space domain. Third, the sensor placement configuration is mathematically formulated as a Markov decision process (MDP) with specified elements. Finally, the sensor locations are optimized through learning the optimal policies of the MDP according to the spatial objective function. The experimental results of a simulated catalytic rod and a real snap curing oven system are provided to demonstrate the feasibility and efficiency of the proposed method in solving the combinatorial optimization problems, such as optimal sensor placement.
Zhi Wang 0001, Han-Xiong Li, Chunlin Chen 0001
IEEE Trans. Cybern.3
2020 Incremental Reinforcement Learning in Continuous Spaces via Policy Relaxation and Importance Weighting
abstract
In this paper, a systematic incremental learning method is presented for reinforcement learning in continuous spaces where the learning environment is dynamic. The goal is to adjust the previously learned policy in the original environment to a new one incrementally whenever the environment changes. To improve the adaptability to the ever-changing environment, we propose a two-step solution incorporated with the incremental learning procedure: policy relaxation and importance weighting. First, the behavior policy is relaxed to a random one in the initial learning episodes to encourage a proper exploration in the new environment. It alleviates the conflict between the new information and the existing knowledge for a better adaptation in the long term. Second, it is observed that episodes receiving higher returns are more in line with the new environment, and hence contain more new information. During parameter updating, we assign higher importance weights to the learning episodes that contain more new information, thus encouraging the previous optimal policy to be faster adapted to a new one that fits in the new environment. Empirical studies on continuous controlling tasks with varying configurations verify that the proposed method achieves a significantly faster adaptation to various dynamic environments than the baselines.
Zhi Wang 0001, Han-Xiong Li, Chunlin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 Self-Paced Prioritized Curriculum Learning With Coverage Penalty in Deep Reinforcement Learning
abstract
In this paper, a new training paradigm is proposed for deep reinforcement learning using self-paced prioritized curriculum learning with coverage penalty. The proposed deep curriculum reinforcement learning (DCRL) takes the most advantage of experience replay by adaptively selecting appropriate transitions from replay memory based on the complexity of each transition. The criteria of complexity in DCRL consist of self-paced priority as well as coverage penalty. The self-paced priority reflects the relationship between the temporal-difference error and the difficulty of the current curriculum for sample efficiency. The coverage penalty is taken into account for sample diversity. With comparison to deep Q network (DQN) and prioritized experience replay (PER) methods, the DCRL algorithm is evaluated on Atari 2600 games, and the experimental results show that DCRL outperforms DQN and PER on most of these games. More results further show that the proposed curriculum training paradigm of DCRL is also applicable and effective for other memory-based deep reinforcement learning approaches, such as double DQN and dueling network. All the experimental results demonstrate that DCRL can achieve improved training efficiency and robustness for deep reinforcement learning.
Zhipeng Ren, Daoyi Dong, Huaxiong Li, Chunlin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.4
2017 Robust Learning Control Design for Quantum Unitary Transformations
abstract
Robust control design for quantum unitary transformations has been recognized as a fundamental and challenging task in the development of quantum information processing due to unavoidable decoherence or operational errors in the experimental implementation of quantum operations. In this paper, we extend the systematic methodology of sampling-based learning control (SLC) approach with a gradient flow algorithm for the design of robust quantum unitary transformations. The SLC approach first uses a "training" process to find an optimal control strategy robust against certain ranges of uncertainties. Then a number of randomly selected samples are tested and the performance is evaluated according to their average fidelity. The approach is applied to three typical examples of robust quantum transformation problems including robust quantum transformations in a three-level quantum system, in a superconducting quantum circuit, and in a spin chain system. Numerical results demonstrate the effectiveness of the SLC approach and show its potential applications in various implementation of quantum unitary transformations.
Chengzhi Wu, Chunlin Chen 0001, Daoyi Dong
IEEE Trans. Cybern.3
2017 Multiagent Reinforcement Learning With Sparse Interactions by Negotiation and Knowledge Transfer
abstract
Reinforcement learning has significant applications for multiagent systems, especially in unknown dynamic environments. However, most multiagent reinforcement learning (MARL) algorithms suffer from such problems as exponential computation complexity in the joint state-action space, which makes it difficult to scale up to realistic multiagent problems. In this paper, a novel algorithm named negotiation-based MARL with sparse interactions (NegoSIs) is presented. In contrast to traditional sparse-interaction-based MARL algorithms, NegoSI adopts the equilibrium concept and makes it possible for agents to select the nonstrict equilibrium-dominating strategy profile (nonstrict EDSP) or meta equilibrium for their joint actions. The presented NegoSI algorithm consists of four parts: 1) the equilibrium-based framework for sparse interactions; 2) the negotiation for the equilibrium set; 3) the minimum variance method for selecting one joint action; and 4) the knowledge transfer of local Q -values. In this integrated algorithm, three techniques, i.e., unshared value functions, equilibrium solutions, and sparse interactions are adopted to achieve privacy protection, better coordination and lower computational complexity, respectively. To evaluate the performance of the presented NegoSI algorithm, two groups of experiments are carried out regarding three criteria: 1) steps of each episode; 2) rewards of each episode; and 3) average runtime. The first group of experiments is conducted using six grid world games and shows fast convergence and high scalability of the presented algorithm. Then in the second group of experiments NegoSI is applied to an intelligent warehouse problem and simulated results demonstrate the effectiveness of the presented NegoSI algorithm compared with other state-of-the-art MARL algorithms.
Luowei Zhou, Chunlin Chen 0001, Yang Gao 0001
IEEE Trans. Cybern.3
2017 Quantum Ensemble Classification: A Sampling-Based Learning Control Approach
abstract
Quantum ensemble classification (QEC) has significant applications in discrimination of atoms (or molecules), separation of isotopes, and quantum information extraction. However, quantum mechanics forbids deterministic discrimination among nonorthogonal states. The classification of inhomogeneous quantum ensembles is very challenging, since there exist variations in the parameters characterizing the members within different classes. In this paper, we recast QEC as a supervised quantum learning problem. A systematic classification methodology is presented by using a sampling-based learning control (SLC) approach for quantum discrimination. The classification task is accomplished via simultaneously steering members belonging to different classes to their corresponding target states (e.g., mutually orthogonal states). First, a new discrimination method is proposed for two similar quantum systems. Then, an SLC method is presented for QEC. Numerical results demonstrate the effectiveness of the proposed approach for the binary classification of two-level quantum ensembles and the multiclass classification of multilevel quantum ensembles.
Chunlin Chen 0001, Daoyi Dong, Ian R. Petersen, Herschel Rabitz
IEEE Trans. Neural Networks Learn. Syst.1
2015 Differential Evolution with Equally-Mixed Strategies for Robust Control of Open Quantum Systems
abstract
Robust control of open quantum systems from one state to another is much more difficult than closed quantum systems as a result of system-environment interactions. In this paper, we adopt the sampling-based learning control approach with the motivation of utilizing some artificial samples instead of unknown uncertainties to design an optimal control field against parameter fluctuations. To enhance the learning performance, we introduce an improved differential evolution (DE) algorithm with equally-mixed strategies in the training step of the control design for open quantum systems. Numerical results verify the effectiveness of the proposed equally-mixed strategies DE (EMSDE) algorithm regarding the control design for open quantum systems with uncertainties.
Hailan Ma, Chunlin Chen 0001, Daoyi Dong
SMC2
2015 A Selective Harmonic Optimization Method for STATCOM in Steady State Based on the Sliding DFT
abstract
The low-order harmonic components in the output current of a three-phase three-wire Static Synchronous Compensator (STATCOM) with an LCL filter are inevitable due to the dead time of the switching devices and the tracking speed of the PI controllers when using the inverter-side inductance current as the control object. In order to suppress the selective harmonic components in the steady-state output current, a control method for STATCOM based on the sliding DFT harmonic detection, closed-loop control by the inverter-side inductance current and feed-forward compensation by the selective harmonic current of the grid-side inductance is proposed. The simulation results in Mat lab demonstrate the effectiveness of the proposed selective harmonic optimization method.
Dachuan Tian, Zhangqing Zhu, Chunlin Chen 0001
SMC3
2015 Robust Quantum Operation for Two-Level Systems Using Sampling-Based Learning Control
abstract
Robust control design for operation of quantum systems has been considered as a demanding and challenging task in the development of quantum technologies. In this paper, we apply the sampling-based learning control (SLC) approach to design a control law for manipulating two-level quantum systems with uncertainties. The gradient-based learning and optimization algorithm is adopted to find the optimal piece-wise control fields for an augmented system by sampling the domain of uncertainties. Numerical results demonstrate the effectiveness of the proposed method for unitary operation of two-level quantum systems even when there are large uncertainties.
Chengzhi Wu, Chunlin Chen 0001, Daoyi Dong
SMC2
2014 Sampling-based learning control for quantum discrimination and ensemble classification
abstract
Quantum ensemble classification has significant applications in discrimination of atoms (or molecules), separation of isotopic molecules and quantum information extraction. In this paper, we recast quantum ensemble classification as a supervised quantum learning problem. A systematic classification methodology is presented by using a sampling-based learning control (SLC) approach for quantum discrimination. The classification task is accomplished via simultaneously steering members belonging to different classes to their corresponding target states (e.g., mutually orthogonal states). Numerical results demonstrate the effectiveness of the proposed approach for the discrimination of two quantum systems and the binary classification of two-level quantum ensembles.
Chunlin Chen 0001, Daoyi Dong, Ian R. Petersen, Herschel Rabitz
IJCNN1
2014 Coordinated standoff tracking of moving targets using differential geometry
abstract
This research is concerned with coordinated standoff tracking, and a guidance law against a moving target is proposed by using differential geometry. We first present the geometry between the unmanned aircraft (UA) and the target to obtain the convergent solution of standoff tracking when the speed ratio of the UA to the target is larger than one. Then, the convergent solution is used to guide the UA onto the standoff tracking geometry. We propose an improved guidance law by adding a derivative term to the relevant algorithm. To keep the phase angle difference of multiple UAs, we add a second derivative term to the relevant control law. Simulations are done to demonstrate the feasibility and performance of the proposed approach. The proposed algorithm can achieve coordinated control of multiple UAs with its simplicity and stability in terms of the standoff distance and phase angle difference.
Zhi-qiang Song, Huaxiong Li, Chunlin Chen 0001, Xianzhong Zhou
J. Zhejiang Univ. Sci. C3
2014 Fidelity-Based Probabilistic Q-Learning for Control of Quantum Systems
abstract
The balance between exploration and exploitation is a key problem for reinforcement learning methods, especially for Q-learning. In this paper, a fidelity-based probabilistic Q-learning (FPQL) approach is presented to naturally solve this problem and applied for learning control of quantum systems. In this approach, fidelity is adopted to help direct the learning process and the probability of each action to be selected at a certain state is updated iteratively along with the learning process, which leads to a natural exploration strategy instead of a pointed one with configured parameters. A probabilistic Q-learning (PQL) algorithm is first presented to demonstrate the basic idea of probabilistic action selection. Then the FPQL algorithm is presented for learning control of quantum systems. Two examples (a spin-1/2 system and a Λ-type atomic system) are demonstrated to test the performance of the FPQL algorithm. The results show that FPQL algorithms attain a better balance between exploration and exploitation, and can also avoid local optimal policies and accelerate the learning process.
Chunlin Chen 0001, Daoyi Dong, Han-Xiong Li, Jian Chu, Tzyh Jong Tarn
IEEE Trans. Neural Networks Learn. Syst.1
2012 Control Design of Uncertain Quantum Systems With Fuzzy Estimators
abstract
An approach of control design using fuzzy estimators (FEs) is proposed for quantum systems with uncertainties. Two types of quantum control problems are considered: 1) control of a pure-state quantum system in the presence of uncertainties and 2) control design of quantum systems with initial mixed states and uncertainties. For the first type of tasks, a partial feedback control scheme with an FE is presented to design controllers. In this scheme, an FE is trained to estimate the quantum state for feedback control of a quantum system, and controlled projective measurement is used to assist in controlling the system. For the second type of quantum control tasks, a probabilistic fuzzy estimator (PFE) is trained to estimate the quantum state for control design of a quantum system with an initial mixed state, and a corresponding control algorithm is proposed to design a control law that drives the system from the mixed state to a target pure state. Two examples of two-spin-1/2systems are also presented and analyzed to demonstrate the process of control design and potential applications of the proposed approach.
Chunlin Chen 0001, Daoyi Dong, James Lam, Jian Chu, Tzyh Jong Tarn
IEEE Trans. Fuzzy Syst.1
2011 Hybrid MDP based integrated hierarchical Q-learning
Chunlin Chen 0001, Daoyi Dong, Han-Xiong Li, Tzyh Jong Tarn
Sci. China Inf. Sci.1
2008 Quantum Reinforcement Learning
abstract
The key approaches for machine learning, particularly learning in unknown probabilistic environments, are new representations and computation mechanisms. In this paper, a novel quantum reinforcement learning (QRL) method is proposed by combining quantum theory and reinforcement learning (RL). Inspired by the state superposition principle and quantum parallelism, a framework of a value-updating algorithm is introduced. The state (action) in traditional RL is identified as the eigen state (eigen action) in QRL. The state (action) set can be represented with a quantum superposition state, and the eigen state (eigen action) can be obtained by randomly observing the simulated quantum state according to the collapse postulate of quantum measurement. The probability of the eigen action is determined by the probability amplitude, which is updated in parallel according to rewards. Some related characteristics of QRL such as convergence, optimality, and balancing between exploration and exploitation are also analyzed, which shows that this approach makes a good tradeoff between exploration and exploitation using the probability amplitude and can speedup learning through the quantum parallelism. To evaluate the performance and practicability of QRL, several simulated experiments are given, and the results demonstrate the effectiveness and superiority of the QRL algorithm for some complex problems. This paper is also an effective exploration on the application of quantum computation to artificial intelligence.
Daoyi Dong, Chunlin Chen 0001, Han-Xiong Li, Tzyh Jong Tarn
IEEE Trans. Syst. Man Cybern. Part B2
2008 Incoherent Control of Quantum Systems With Wavefunction-Controllable Subspaces via Quantum Reinforcement Learning
abstract
In this paper, an incoherent control scheme for accomplishing the state control of a class of quantum systems which have wavefunction-controllable subspaces is proposed. This scheme includes the following two steps: projective measurement on the initial state and learning control in the wavefunction-controllable subspace. The first step probabilistically projects the initial state into the wavefunction-controllable subspace. The probability of success is sensitive to the initial state; however, it can be greatly improved through multiple experiments on several identical initial states even in the case with a small probability of success for an individual measurement. The second step finds a local optimal control sequence via quantum reinforcement learning and drives the controlled system to the objective state through a set of suitable controls. In this strategy, the initial states can be unknown identical states, the quantum measurement is used as an effective control, and the controlled system is not necessarily unitarily controllable. This incoherent control scheme provides an alternative quantum engineering strategy for locally controllable quantum systems.
Daoyi Dong, Chunlin Chen 0001, Tzyh Jong Tarn, Alexander N. Pechen, Herschel Rabitz
IEEE Trans. Syst. Man Cybern. Part B2