VLDB 2026 Research / reviewers in the wild / expert
Daoyi Dong
dblp:27/3317
· DBLP profile ↗
63ranked-venue papers
3as first author
39since 2021 · last 2026
0000-0002-7425-3559ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 11 since 2021Human-computer interaction and ubiquitous computing · 18 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Theory of computation · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Diffusion Model for Multi-Agent Dynamic Task DecompositionabstractTask decomposition has shown promise in complex cooperative multi-agent reinforcement learning (MARL) tasks, which enables efficient hierarchical learning for long-horizon tasks in dynamic and uncertain environments. However, learning dynamic task decomposition from scratch generally requires a large number of training samples, especially exploring the large joint action space under partial observability. In this paper, we present the Conditional Diffusion Model for Dynamic Task Decomposition (CD3T), a novel two-level hierarchical MARL framework designed to automatically infer subtask and coordination patterns. The high-level policy learns subtask representation to generate a subtask selection strategy based on subtask effects. To capture the effects of subtasks on the environment, CD3T predicts the next observation and reward using a conditional diffusion model. At the low level, agents collaboratively learn and share specialized skills within their assigned subtasks. Moreover, the learned subtask representation is also used as additional semantic information in a multi-head attention mixing network to enhance value decomposition and provide an efficient reasoning bridge between individual and joint value functions. Experimental results on various benchmarks demonstrate that CD3T achieves better performance than existing baselines. Yanda Zhu, Yuanyang Zhu, Daoyi Dong, Caihua Chen, Chunlin Chen 0001 |
AAAI | 3 |
| 2026 | Select Before Use: On the Importance of Reference Model Selection in Preference AlignmentabstractThe post-training stage of Large Language Models (LLMs) typically involves Supervised Fine-Tuning (SFT) followed by preference alignment to ensure LLMs generate safe, helpful, and instruction-aligned content.The SFT model critically serves as both the initialization and reference model for subsequent preference alignment.However, an essential yet often neglected question is the optimal selection of the SFT checkpoint for this role.We show that checkpoint selection substantially affects final performance, and that the common practice of choosing the minimum validation-loss checkpoint often fails, due to a fundamental conflict between SFT's focus on imitation and alignment's goal of response discriminability.To this end, we propose RewardRank, a simple, effective, training-free metric for estimating initial implicit alignment between reference model and preference objective.Empirical evidence suggests that, using our selected model as reference can gain up to 67.6% relative increase on length-controlled win rate on the popular Zephyr recipe comparing to baselines. Runze Wu 0001, Xiangyu Zhao 0001, Bo Han 0003, Daoyi Dong, Tongliang Liu |
ACL (1) | 5 |
| 2026 | Learning hierarchical time-frequency representation for long-term time series forecastingabstractTime series forecasting is essential for planning and management across various domains. Existing models struggle to maintain long-term trends in extended predictions and overlook the interplay between time and frequency-domain dependencies. To address these challenges, we propose TFformer, a hierarchical time–frequency representation architecture with Transformer, involving two key innovations: (i) spectrum decomposition isolates long-term patterns from short-term fluctuations and (ii) sequence aggregation integrates two categories of features distinguished by different energy intensities in a hierarchical manner. Experiments on six real-world datasets show that TFformer outperforms the frequency-domain baseline (FreTS) with an average 16.54% improvement in Mean Squared Error (MSE) and surpasses the time-domain baseline (iTransformer) with an average 5.91% MSE improvement, highlighting its effectiveness in capturing both time and frequency-domain patterns. Zhongju Wang 0001, Zhenhong Sun, Yatao Bian, Huadong Mo, Daoyi Dong |
Inf. Process. Manag. | 5 |
| 2026 | TemSoGraph: Learning temporal social graphs for cyberbullying predictionabstractCyberbullying is a pervasive issue on online platforms, yet early intervention via predictive modeling remains an open challenge. This challenge is compounded by the temporal dynamics of user interactions and the sparsity of such interactions in real-world social networks, making reliable modeling difficult. Current methods predominantly focus on detecting cyberbullying after it occurs through user content and profiles, while overlooking the temporal patterns and struggling when social interaction data is limited. We propose TemSoGraph, a unified temporal social graph learning model for cyberbullying detection and prediction. The model leverages a temporal self-attention mechanism to capture time-evolving user interactions and employs joint global and local node updates to represent users with limited interactions. It further incorporates a domain adaptor that learns domain-invariant features, enhancing generalization across datasets even when labeled target data is scarce. Experiments on two real-world datasets, Instagram and Vine, show that TemSoGraph outperforms eight cyberbullying detection models in detection task and six dynamic graph neural networks in prediction task. On the prediction task, TemSoGraph achieves a recall of 97.18% on Instagram with 2.53% improvement and 93.38% on Vine with 6.25% improvement. The model supports both detection and future prediction and provides a strong benchmark for cyberbullying modeling. • We propose TemSoGraph model for cyberbullying detection and prediction. • TemSoGraph works effectively under real-world data sparsity problem. • TemSoGraph integrates domain adaptor for cross-dataset generalization. Wensi Jiang, Min Wang 0009, Huadong Mo, Daoyi Dong, Yu Zhang 0217, Wenjie Zhang 0001 |
Inf. Sci. | 4 |
| 2026 | Evolutionary Optimization-Based Design of LQG Controllers in Quantum Coherent FeedbackabstractIn this article, we propose a differential evolution (DE) algorithm specifically tailored for the design of linear-quadratic-Gaussian (LQG) controllers in quantum systems. Building upon the foundational DE framework, the algorithm incorporates specialized modules, including relaxed feasibility rules, a scheduled penalty function, adaptive search range adjustment, and the "bet-and-run" initialization strategy. These enhancements improve the algorithm's exploration and exploitation capabilities while addressing the unique physical realizability requirements of quantum systems. The proposed method is applied to a quantum optical system, where three distinct controllers with varying configurations relative to the plant are designed. The resulting controllers demonstrate superior performance, achieving lower LQG performance indices compared to existing approaches. In addition, the algorithm ensures that the designs comply with physical realizability constraints, guaranteeing compatibility with practical quantum platforms. The proposed approach holds significant potential for application to other linear quantum systems in performance optimization tasks subject to physically feasible constraints. Chunxiang Song, Guofeng Zhang 0003, Huadong Mo, Daoyi Dong |
IEEE Trans. Cybern. | 5 |
| 2025 | PN-GAIL: Leveraging Non-optimal Information from Imperfect DemonstrationsabstractImitation learning aims at constructing an optimal policy by emulating expert demonstrations. However, the prevailing approaches in this domain typically presume that the demonstrations are optimal, an assumption that seldom holds true in the complexities of real-world applications. The data collected in practical scenarios often contains imperfections, encompassing both optimal and non-optimal examples. In this study, we propose Positive-Negative Generative Adversarial Imitation Learning (PN-GAIL), a novel approach that falls within the framework of Generative Adversarial Imitation Learning (GAIL). PN-GAIL innovatively leverages non-optimal information from imperfect demonstrations, allowing the discriminator to comprehensively assess the positive and negative risks associated with these demonstrations. Furthermore, it requires only a small subset of labeled confidence scores. Theoretical analysis indicates that PN-GAIL deviates from the non-optimal data while mimicking imperfect demonstrations. Experimental results demonstrate that PN-GAIL surpasses conventional baseline methods in dealing with imperfect demonstrations, thereby significantly augmenting the practical utility of imitation learning in real-world contexts. Our codes are available at https://github.com/QiangLiuT/PN-GAIL. Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001, Daoyi Dong |
ICLR | 5 |
| 2025 | Mixture-of-Experts Meets In-Context Reinforcement LearningabstractIn-context reinforcement learning (ICRL) has emerged as a promising paradigm for adapting RL agents to downstream tasks through prompt conditioning. However, two notable challenges remain in fully harnessing in-context learning within RL domains: the intrinsic multi-modality of the state-action-reward data and the diverse, heterogeneous nature of decision tasks. To tackle these challenges, we propose **T2MIR** (**T**oken- and **T**ask-wise **M**oE for **I**n-context **R**L), an innovative framework that introduces architectural advances of mixture-of-experts (MoE) into transformer-based decision models. T2MIR substitutes the feedforward layer with two parallel layers: a token-wise MoE that captures distinct semantics of input tokens across multiple modalities, and a task-wise MoE that routes diverse tasks to specialized experts for managing a broad task distribution with alleviated gradient conflicts. To enhance task-wise routing, we introduce a contrastive learning method that maximizes the mutual information between the task and its router representation, enabling more precise capture of task-relevant information. The outputs of two MoE components are concatenated and fed into the next layer. Comprehensive experiments show that T2MIR significantly facilitates in-context learning capacity and outperforms various types of baselines. We bring the potential and promise of MoE to ICRL, offering a simple and scalable architectural enhancement to advance ICRL one step closer toward achievements in language and vision communities. Our code is available at [https://github.com/NJU-RL/T2MIR](https://github.com/NJU-RL/T2MIR). Fuhong Liu, Haoru Li, Zican Hu, Daoyi Dong, Chunlin Chen 0001, Zhi Wang 0001 |
NeurIPS | 5 |
| 2025 | Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language SupervisionabstractOffline meta-RL usually tackles generalization by inferring task beliefs from high-quality samples or warmup explorations. The restricted form limits their generality and usability since these supervision signals are expensive and even infeasible to acquire in advance for unseen tasks. Learning directly from the raw text about decision tasks is a promising alternative to leverage a much broader source of supervision. In the paper, we propose **T**ext-to-**D**ecision **A**gent (**T2DA**), a simple and scalable framework that supervises offline meta-RL with natural language. We first introduce a generalized world model to encode multi-task decision data into a dynamics-aware embedding space. Then, inspired by CLIP, we predict which textual description goes with which decision embedding, effectively bridging their semantic gap via contrastive language-decision pre-training and aligning the text embeddings to comprehend the environment dynamics. After training the text-conditioned generalist policy, the agent can directly realize zero-shot text-to-decision generation in response to language instructions. Comprehensive experiments on MuJoCo and Meta-World benchmarks show that T2DA facilitates high-capacity zero-shot generalization and outperforms various types of baselines. Our code is available at [https://github.com/NJU-RL/T2DA](https://github.com/NJU-RL/T2DA). Zican Hu, Jianxiang Tang, Chunlin Chen 0001, Daoyi Dong, Yu Cheng 0001, Zhenhong Sun, Zhi Wang 0001 |
NeurIPS | 7 |
| 2025 | A Prescription-Centric Estimation Framework for Bi-Level Power System Operations with Analytical RepresentationabstractThe increasing integration of renewable energy sources and battery energy storage systems has amplified uncertainties in power system operations, necessitating advanced estimation methods that transcend traditional quality-oriented approaches. In this paper, we propose a prescription-centric estimation framework that embeds decision-making insights directly into the anticipation process. By leveraging parametric programming and implicit gradient representations, our approach establishes an analytical mapping between uncertain parameters and optimal operational decisions, thereby addressing the inherent asymmetry between uncertainty estimation and system re-balancing costs. Notably, the proposed method breaks through the limitations imposed by linearization constraints in conventional estimation networks and optimization models, paving the way for more accurate and robust decision-making. An iterative bi-level nonlinear optimization strategy is also introduced to overcome the shortcomings of purely data-driven methods. The effectiveness of the framework is demonstrated through case studies, underscoring its potential to enhance both decision accuracy and efficiency in power system operations. Yuhao Jing, Fusen Guo, Huadong Mo, Daoyi Dong |
SMC | 6 |
| 2025 | A Comparative Study of Battery SOH Prediction Models: Exploration of Transformer Method with Reversible Instance NormalisationabstractAccurate estimation of the state of health (SOH) of batteries is of great importance for the safe and efficient operation of energy storage systems. However, data-driven methods are often affected by limited generalisation due to sample distribution shift between different batteries, which make them difficult to extent models to unknown domains. To this end, a Transformer-based architecture integrated with RevIN is proposed in this study, which has presented a significant enhancement on the adaptability to distribution shifts of the input data. The normalisation and denormalisation process of Reversible Instance Normalisation reduces statistical bias whilst maintaining trend information. A comparative evaluation is conducted on the NASA battery datasets across six representative models, including Random Forest, eXtreme Gradient Boosting, Multilayer Perceptron, Long Short-Term Memory, Transformer, and the proposed RevIN-Transformer, under both intra-battery and cross-battery prediction settings. Moreover, the analysis of variance is employed to assess the consistency of SOH prediction errors across different batteries. The results indicate that whilst deep learning methods generally outperform traditional models, the RevIN-Transformer method achieves superior accuracy and stability in cross-battery SOH prediction tasks against distribution shifts. In addition, they also validate the effectiveness of integrating lightweight normalisation modules in a Transformer-based architecture for battery SOH time-series forecasting. Runpu Wang, Zongjun Li, Fusen Guo, Daoyi Dong, Huadong Mo |
SMC | 4 |
| 2025 | Entanglement Measure-Based Sliding Mode Control for Quantum State PreparationabstractEntangled states are fundamental to quantum information processing. However, many existing quantum control methods rely on predefined target states, limiting their flexibility in accommodating diverse entanglement structures. This article introduces a sliding mode control framework that utilizes an entanglement measure as the sliding surface, enabling the generation of entangled states without specifying a fixed target. By adjusting the desired entanglement level, the proposed method can generate a wide range of states, including both bipartite and multipartite configurations, as well as pure and mixed states. Since the entanglement measure is scalar-valued, the resulting control law is inherently independent of the number of subsystems-an important advantage of the proposed approach. Among various entangled states, maximally entangled states (MESs) are of particular interest. Lyapunov stability of the control scheme is established, and numerical simulations confirm its effectiveness in robustly generating MESs in both bipartite and multipartite systems. Yunyan Lee, Ciann-Dong Yang, Daoyi Dong |
IEEE Trans. Cybern. | 3 |
| 2025 | Tomography of Quantum States From Structured Measurements via Quantum-Aware TransformerabstractQuantum state tomography (QST) is the process of reconstructing the state of a quantum system (mathematically described as a density matrix) through a series of different measurements, which can be solved by learning a parameterized function to translate experimentally measured statistics into physical density matrices. However, the specific structure of quantum measurements for characterizing a quantum state has been neglected in previous work. In this article, we explore the similarity between highly structured sentences in natural language and intrinsically structured measurements in QST. To fully leverage the intrinsic quantum characteristics involved in QST, we design a quantum-aware transformer (QAT) model to capture the complex relationship between measured frequencies and density matrices. In particular, we query quantum operators in the architecture to facilitate informative representations of quantum data and integrate the Bures distance into the loss function to evaluate quantum state fidelity, thereby enabling the reconstruction of quantum states from measured data with high fidelity. Extensive simulations and experiments (on IBM quantum computers) demonstrate the superiority of the QAT in reconstructing quantum states with favorable robustness against experimental noise. Hailan Ma, Zhenhong Sun, Daoyi Dong, Chunlin Chen 0001, Herschel Rabitz |
IEEE Trans. Cybern. | 3 |
| 2025 | Auxiliary Task-Based Deep Reinforcement Learning for Quantum ControlabstractDue to its property of not requiring prior knowledge of the environment, reinforcement learning (RL) has significant potential for solving quantum control problems. In this work, we investigate the effectiveness of continuous control policies based on deep deterministic policy gradient. To achieve good control of quantum systems with high fidelity, we propose an auxiliary task-based deep RL (AT-DRL) for quantum control. In particular, we design an auxiliary task to predict the fidelity value, sharing partial parameters with the main network (from the main RL task). The auxiliary task learns synchronously with the main task, allowing one to extract intrinsic features of the environment, thus aiding the agent to achieve the desired state with high fidelity. To further enhance the control performance, we also design a guided reward function based on the fidelity of quantum states that enables gradual fidelity improvement. Numerical simulations demonstrate that the proposed AT-DRL can provide a good solution to the exploration of quantum dynamics. It not only achieves high task fidelities but also demonstrates fast learning rates. Moreover, AT-DRL has great potential in designing control pulses that achieve effective quantum state preparation. Shumin Zhou, Hailan Ma, Sen Kuang, Daoyi Dong |
IEEE Trans. Cybern. | 4 |
| 2025 | A Two-Stage Solution to Quantum Process Tomography: Error Analysis and Optimal DesignabstractQuantum process tomography is a critical task for characterizing the dynamics of quantum systems and achieving precise quantum control. In this paper, we propose a two-stage solution for both trace-preserving and non-trace-preserving quantum process tomography. Utilizing a tensor structure, our algorithm exhibits a computational complexity of$O(MLd^{2})$where d is the dimension of the quantum system and$M, L~(M\geq d^{2}, L\geq d^{2})$represent the numbers of different input states and measurement operators, respectively. We establish an analytical error upper bound and then design the optimal input states and the optimal measurement operators, which are both based on minimizing the error upper bound and maximizing the robustness characterized by the condition number. Numerical examples and testing on IBM quantum devices are presented to demonstrate the performance and efficiency of our algorithm. Shuixin Xiao, Yuanlong Wang 0001, Jun Zhang 0090, Daoyi Dong, Gary J. Mooney, Ian R. Petersen, Hidehiro Yonezawa |
IEEE Trans. Inf. Theory | 4 |
| 2025 | Power Characterization of Noisy Quantum KernelsabstractQuantum kernel methods have been widely recognized as one of the promising quantum machine learning (QML) algorithms that have the potential to achieve quantum advantages. However, their capabilities may be severely degraded by inevitable noises in the current noisy intermediate-scale quantum (NISQ) era. In this article, we theoretically characterize the power of noisy quantum kernels and demonstrate that under depolarizing noise, quantum kernel methods may only have very poor prediction capability, even when the generalization error is small. Specifically, we quantitatively describe the decreasing of the prediction capability of noisy quantum kernels in terms of the rate of quantum noise, the size of training samples, the number of qubits, and the number of layers affected by quantum noises. Our results clearly demonstrate that for a given number of training samples, once the number of layers affected by noise exceeds some threshold, the prediction capability of noisy kernels is very poor. Thus, we provide a crucial warning to employ noisy quantum kernel methods for quantum computation and the theoretical results can also serve as guidelines when developing practical quantum kernel algorithms for achieving quantum advantages. Xin Wang 0094, Tongliang Liu, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | EGGen: Image Generation with Multi-entity Prior Learning through Entity GuidanceabstractDiffusion models have shown remarkable prowess in text-to-image synthesis and editing, yet they often stumble when tasked with interpreting complex prompts that describe multiple entities with specific attributes and interrelations. The generated images often contain inconsistent multi-entity representation (IMR), reflected as inaccurate presentations of the multiple entities and their attributes. Although providing spatial layout guidance improves the multi-entity generation quality in existing works, it is still challenging to handle the leakage attributes and avoid unnatural characteristics. To address the IMR challenge, we first conduct in-depth analyses of the diffusion process and attention operation, revealing that the IMR challenges largely stem from the process of cross-attention mechanisms. According to the analyses, we introduce the entity guidance generation mechanism, which maintains the integrity of the original diffusion model parameters by integrating plug-in networks. Our work advances the stable diffusion model by segmenting comprehensive prompts into distinct entity-specific prompts with bounding boxes, enabling a transition from multi-entity to single-entity generation in cross-attention layers. More importantly, we introduce entity-centric cross-attention layers that focus on individual entities to preserve their uniqueness and accuracy, alongside global entity alignment layers that refine cross-attention maps using multi-entity priors for precise positioning and attribute accuracy. Additionally, a linear attenuation module is integrated to progressively reduce the influence of these layers during inference, preventing oversaturation and preserving generation fidelity. Our comprehensive experiments demonstrate that this entity guidance generation enhances existing text-to-image models in generating detailed, multi-entity images. Zhenhong Sun, Junyan Wang 0001, Zhiyu Tan, Daoyi Dong, Hailan Ma, Hao Li 0030, Dong Gong |
ACM Multimedia | 4 |
| 2024 | Quantum Robust Control for Time-Varying Noises Based on Adversarial LearningabstractTime-varying noises are one of the reasons that make it difficult for quantum systems to complete control tasks. How to quantify the influence of time-varying noises on control results and how to design a control law that can resist time-varying noises are two important problems. In this paper, the adversarial learning is introduced into quantum control and the loss function under the worst-case noise is used as a way to quantify the impact of time-varying noises on control performance. We utilize the Gradient Ascent Pulse Engineering (GRAPE) technique to search the worst-case noise and meanwhile offer a strategy to improve the robustness of the control law. Simulation experiments on a two-qubit system and a four-qubit system show that the found noises indeed can act as worst-case noises. Furthermore, the optimized control laws demonstrate good robustness to time-varying noises in state preparation tasks. Haotian Ji, Sen Kuang, Daoyi Dong, Chunlin Chen 0001 |
SMC | 3 |
| 2024 | Distributed Charging Scheduling and Pricing Strategy for Plug-in Electric Vehicles Based on Stackelberg-Nash and Multi-Cluster Aggregative GamesabstractIn this paper, we propose a distributed and interactive Plug-in Electric Vehicle (PEV) charging scheduling approach, which is also combined with an optimal pricing strategy. This method tackles challenges such as fluctuations in charging currents, potential supply congestion, and uneven demand distribution that arise as PEV penetration increases. The objective is to improve the robust stability of the charging system while also reducing the costs for PEV users. This study designs a multi-cluster aggregative game mechanism to handle the competitive dynamics among operational clusters and the collective behavior of individual PEVs. Additionally, a strategic pricing method, based on Stackelberg game theory, is designed to refine the determination of basic electricity prices. We further introduce a distributed update method that efficiently seeks the Nash Equilibrium (NE) of the hierarchical game described. The effectiveness of the proposed architecture and solution methodology is validated through experimental studies. Yuhao Jing, Huadong Mo, Daoyi Dong |
SMC | 5 |
| 2024 | Sustainable Energy Planning for Community Microgrids Considering Economic, Environmental, and Resilience FactorsabstractThis study presents a framework for sustainable energy planning of community microgrids (MGs), integrating optimal design and decision-support tools. A rural community in New South Wales, Australia, is considered as a case study for this investigation. The proposed microgrid framework is evaluated based on economic viability, environmental sustainability, and community resilience. The economic analysis reveals an attractive net present cost of $3.26 million over the MG's 25-year lifetime, with a competitive levelized cost of energy of $0.196 per kWh. The environmental impact assessment quantifies a significant reduction of 394.429 tonnes of$CO_{2}{-}$equivalent greenhouse gas emissions annually through the integration of 200 kW of solar photovoltaic and 258 kW of wind turbines. The resilience assessment demonstrates a high energy reliability with zero unmet loads facilitated by backup systems and decision-making tools. The findings contribute to the field of sustainable energy planning by providing a comprehensive and integrated approach that addresses the complex interplay of economic, environmental, and resilience factors in the context of community MGs. Moslem Uddin, Huadong Mo, Daoyi Dong |
SMC | 3 |
| 2024 | Quantum Bandit With Amplitude Amplification Exploration in an Adversarial EnvironmentabstractThe rapid proliferation of learning systems in an arbitrarily changing environment mandates the need to manage tensions between exploration and exploitation. This work proposes a quantum-inspired bandit learning approach for the learning-and-adapting-based offloading problem where a client observes and learns the costs of each task offloaded to the candidate resource providers, e.g., fog nodes. In this approach, a new action update strategy and novel probabilistic action selection are adopted, provoked by the amplitude amplification and collapse postulate in quantum computation theory. We devise a locally linear mapping between a quantum-mechanical phase in a quantum domain, e.g., Grover-type search algorithm, and a distilled probability-magnitude in a value-based decision-making domain, e.g., adversarial multi-armed bandit algorithm. The proposed algorithm is generalized, via the devised mapping, for better learning weight adjustments on favorable/unfavorable actions, and its effectiveness is verified via simulation. Byungjin Cho, Yu Xiao 0001, Pan Hui 0001, Daoyi Dong |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Efficient Bayesian Policy Reuse With a Scalable Observation Model in Deep Reinforcement LearningabstractBayesian policy reuse (BPR) is a general policy transfer framework for selecting a source policy from an offline library by inferring the task belief based on some observation signals and a trained observation model. In this article, we propose an improved BPR method to achieve more efficient policy transfer in deep reinforcement learning (DRL). First, most BPR algorithms use the episodic return as the observation signal that contains limited information and cannot be obtained until the end of an episode. Instead, we employ the state transition sample, which is informative and instantaneous, as the observation signal for faster and more accurate task inference. Second, BPR algorithms usually require numerous samples to estimate the probability distribution of the tabular-based observation model, which may be expensive and even infeasible to learn and maintain, especially when using the state transition sample as the signal. Hence, we propose a scalable observation model based on fitting state transition functions of source tasks from only a small number of samples, which can generalize to any signals observed in the target task. Moreover, we extend the offline-mode BPR to the continual learning setting by expanding the scalable observation model in a plug-and-play fashion, which can avoid negative transfer when faced with new unknown tasks. Experimental results show that our method can consistently facilitate faster and more efficient policy transfer. Jinmei Liu, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Depthwise Convolution for Multi-Agent Communication With Enhanced Mean-Field ApproximationabstractMulti-Agent settings remain a fundamental challenge in the reinforcement learning (RL) domain due to the partial observability and the lack of accurate real-time interactions across agents. In this article, we propose a new method based on local communication learning to tackle the multi-agent RL (MARL) challenge within a large number of agents coexisting. First, we design a new communication protocol that exploits the ability of depthwise convolution to efficiently extract local relations and learn local communication between neighboring agents. To facilitate multi-agent coordination, we explicitly learn the effect of joint actions by taking the policies of neighboring agents as inputs. Second, we introduce the mean-field approximation into our method to reduce the scale of agent interactions. To more effectively coordinate behaviors of neighboring agents, we enhance the mean-field approximation by a supervised policy rectification network (PRN) for rectifying real-time agent interactions and by a learnable compensation term for correcting the approximation bias. The proposed method enables efficient coordination as well as outperforms several baseline approaches on the adaptive traffic signal control (ATSC) task and the StarCraft II multi-agent challenge (SMAC). Donghan Xie, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Hamiltonian Identification via Quantum Ensemble ClassificationabstractIdentifying the Hamiltonian of an unknown quantum system is a critical task in the area of quantum information. In this article, we propose a systematic Hamiltonian identification approach via quantum ensemble multiclass classification (HI-QEMC). This approach is implemented by a three-step iterative refining process, i.e., parameter interval guess, verification, and judgment. In the parameter interval guess step, the parameter interval is divided into several sub-intervals and the true Hamiltonian parameter is guessed in one of them. In the parameter interval verification step, cross verification is applied to verify the accuracy of the guess. In the parameter interval judgment step, an adaptive interval judgment (AIJ) algorithm is designed to determine the sub-interval containing the true Hamiltonian parameter. Numerical results on two typical quantum systems, i.e., two-level quantum systems and three-level quantum systems, demonstrate the effectiveness and superior performance of the proposed approach for quantum Hamiltonian identification. Haixu Yu, Xudong Zhao 0001, Daoyi Dong, Chunlin Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Guided Reward Design in Continuous Reinforcement Learning for Quantum ControlabstractReinforcement learning has been intensively applied to tackle complex quantum control problems owing to its adaptability in dynamic environments. However, commonly used reinforcement learning algorithms are restricted to selecting actions from a discrete action space, which may result in inaccurate control. In order to surpass this constraint, we propose an improved continuous reinforcement learning algorithm that can generate control policies in a continuous action space, enabling precise control of quantum systems. Moreover, a guided reward function design method is proposed to guide the learning process toward higher fidelity. Numerical results demonstrate that the proposed continuous reinforcement learning algorithm with a guided reward function can capably prepare states on one-qubit and two-qubit systems. Shumin Zhou, Hailan Ma, Sen Kuang, Daoyi Dong |
SMC | 4 |
| 2023 | Demand response management of smart grid based on Stackelberg-evolutionary joint game
Daoyi Dong |
Sci. China Inf. Sci. | 3 |
| 2023 | Quantum Language Model With Entanglement Embedding for Question AnsweringabstractQuantum language models (QLMs) in which words are modeled as a quantum superposition of sememes have demonstrated a high level of model transparency and good post-hoc interpretability. Nevertheless, in the current literature, word sequences are basically modeled as a classical mixture of word states, which cannot fully exploit the potential of a quantum probabilistic description. A quantum-inspired neural network (NN) module is yet to be developed to explicitly capture the nonclassical correlations within the word sequences. We propose a NN model with a novel entanglement embedding (EE) module, whose function is to transform the word sequence into an entangled pure state representation. Strong quantum entanglement, which is the central concept of quantum information and an indication of parallelized correlations among the words, is observed within the word sequences. The proposed QLM with EE (QLM-EE) is proposed to implement on classical computing devices with a quantum-inspired NN structure, and numerical experiments show that QLM-EE achieves superior performance compared with the classical deep NN models and other QLMs on question answering (QA) datasets. In addition, the post-hoc interpretability of the model can be improved by quantifying the degree of entanglement among the word states. Yiwei Chen 0002, Yu Pan 0001, Daoyi Dong |
IEEE Trans. Cybern. | 3 |
| 2023 | A Dirichlet Process Mixture of Robust Task Models for Scalable Lifelong Reinforcement LearningabstractWhile reinforcement learning (RL) algorithms are achieving state-of-the-art performance in various challenging tasks, they can easily encounter catastrophic forgetting or interference when faced with lifelong streaming information. In this article, we propose a scalable lifelong RL method that dynamically expands the network capacity to accommodate new knowledge while preventing past memories from being perturbed. We use a Dirichlet process mixture to model the nonstationary task distribution, which captures task relatedness by estimating the likelihood of task-to-cluster assignments and clusters the task models in a latent space. We formulate the prior distribution of the mixture as a Chinese restaurant process (CRP) that instantiates new mixture components as needed. The update and expansion of the mixture are governed by the Bayesian nonparametric framework with an expectation maximization (EM) procedure, which dynamically adapts the model complexity without explicit task boundaries or heuristics. Moreover, we use the domain randomization technique to train robust prior parameters for the initialization of each task model in the mixture; thus, the resulting model can better generalize and adapt to unseen tasks. With extensive experiments conducted on robot navigation and locomotion domains, we show that our method successfully facilitates scalable lifelong RL and outperforms relevant existing methods. Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Cybern. | 3 |
| 2023 | Curriculum-Based Deep Reinforcement Learning for Quantum ControlabstractDeep reinforcement learning (DRL) has been recognized as an efficient technique to design optimal strategies for different complex systems without prior knowledge of the control landscape. To achieve a fast and precise control for quantum systems, we propose a novel DRL approach by constructing a curriculum consisting of a set of intermediate tasks defined by fidelity thresholds, where the tasks among a curriculum can be statically determined before the learning process or dynamically generated during the learning process. By transferring knowledge between two successive tasks and sequencing tasks according to their difficulties, the proposed curriculum-based DRL (CDRL) method enables the agent to focus on easy tasks in the early stage, then move onto difficult tasks, and eventually approaches the final task. Numerical comparison with the traditional methods [gradient method (GD), genetic algorithm (GA), and several other DRL methods] demonstrates that CDRL exhibits improved control performance for quantum systems and also provides an efficient way to identify optimal strategies with few control pulses. Hailan Ma, Daoyi Dong, Steven X. Ding, Chunlin Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Instance Weighted Incremental Evolution Strategies for Reinforcement Learning in Dynamic EnvironmentsabstractEvolution strategies (ESs), as a family of black-box optimization algorithms, recently emerge as a scalable alternative to reinforcement learning (RL) approaches such as Q-learning or policy gradient and are much faster when many central processing units (CPUs) are available due to better parallelization. In this article, we propose a systematic incremental learning method for ES in dynamic environments. The goal is to adjust previously learned policy to a new one incrementally whenever the environment changes. We incorporate an instance weighting mechanism with ES to facilitate its learning adaptation while retaining scalability of ES. During parameter updating, higher weights are assigned to instances that contain more new knowledge, thus encouraging the search distribution to move toward new promising areas of parameter space. We propose two easy-to-implement metrics to calculate the weights: instance novelty and instance quality. Instance novelty measures an instance's difference from the previous optimum in the original environment, while instance quality corresponds to how well an instance performs in the new environment. The resulting algorithm, instance weighted incremental evolution strategies (IW-IESs), is verified to achieve significantly improved performance on challenging RL tasks ranging from robot navigation to locomotion. This article thus introduces a family of scalable ES algorithms for RL domains that enables rapid learning adaptation to dynamic environments. Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | A Modified Deep Q-Learning Algorithm for Optimal and Robust Quantum Gate Design of a Single Qubit System*abstractPrecise and resilient quantum gate design is important for the building of quantum devices. In this paper, we consider the optimal and robust quantum gate design problem for three classes of two-level quantum systems. The aim is to construct quantum gates in a given fixed time with limited control resources. A modified dueling deep Q-learning (MDuDQL) is employed for the optimal and robust gate design problem. To improve the performance of the classical DuDQL method, we propose a unique semi-Markov DuDQL algorithm based on a modified action selection procedure, modified replay memory, and soft update procedure. The proposed algorithm outperforms ordinary DuDQL in terms of discovering global optimal or near-global optimal control protocols and faster convergence to a better policy. Moreover, the modified DuDQL agent shows improved performance in finding robust control protocols which achieve high-fidelity quantum gate design for varying uncertainties in a certain range. The effectiveness of the proposed algorithm for the optimal and robust gate design problems has been illustrated by numerical results. Omar Shindi, Qi Yu 0006, Parth Girdhar, Daoyi Dong |
SMC | 4 |
| 2022 | Learning Control with Evolution Strategy for Inhomogeneous Open Quantum EnsemblesabstractThis paper investigates the application of an evolutionary algorithm, evolution strategy (ES)($\mu+\lambda$) to the control design in several inhomogeneous open quantum ensembles. We apply the ES ($\mu+\lambda$) to assist the sampling-based learning control (SLC) technique, by which a set of control signals is designed to drive the inhomogeneous open quantum ensemble to a given target state. We illustrate our algorithm in two-level and four-level inhomogeneous open quantum ensembles. Numerical results show the effectiveness of the proposed control algorithm. The comparison with other evolutionary algorithms such as differential evolution (DE) and genetic algorithm (GA) shows the superiority of our ES ($\mu+\lambda$) both in average fidelity and stability. In a four-level open quantum ensemble, for example, the fitness error after optimization using the ES ($\mu+\lambda$) is decreased by around 59% compared to DE, and the standard deviation is lowered by about 47%. Chunxiang Song, David McManus, Daoyi Dong |
SMC | 4 |
| 2022 | Multi-channel quantum parameter estimation
Liying Bao, Daoyi Dong, Rebing Wu |
Sci. China Inf. Sci. | 4 |
| 2022 | Deep Reinforcement Learning With Quantum-Inspired Experience ReplayabstractIn this article, a novel training paradigm inspired by quantum computation is proposed for deep reinforcement learning (DRL) with experience replay. In contrast to the traditional experience replay mechanism in DRL, the proposed DRL with quantum-inspired experience replay (DRL-QER) adaptively chooses experiences from the replay buffer according to the complexity and the replayed times of each experience (also called transition), to achieve a balance between exploration and exploitation. In DRL-QER, transitions are first formulated in quantum representations and then the preparation operation and depreciation operation are performed on the transitions. In this process, the preparation operation reflects the relationship between the temporal-difference errors (TD-errors) and the importance of the experiences, while the depreciation operation is taken into account to ensure the diversity of the transitions. The experimental results on Atari 2600 games show that DRL-QER outperforms state-of-the-art algorithms, such as DRL-PER and DCRL on most of these games with improved training efficiency and is also applicable to such memory-based DRL approaches as double network and dueling network. Hailan Ma, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Cybern. | 4 |
| 2022 | Hybrid Filtering for a Class of Nonlinear Quantum Systems Subject to Classical Stochastic DisturbancesabstractA hybrid quantum-classical filtering problem, where a qubit system is disturbed by a classical stochastic process, is investigated. The strategy is to model the classical disturbance by using an optical cavity. The relations between classical disturbances and the cavity analog system are analyzed. The dynamics of the enlarged quantum network system, which includes a qubit system and a cavity system, are derived. A stochastic master equation for the qubit-cavity hybrid system is given, based on which estimates for the state of the cavity system and the classical signal are obtained. The quantum-extended Kalman filter is employed to achieve efficient computation. The numerical results are presented to illustrate the effectiveness of our methods. Qi Yu 0006, Daoyi Dong, Ian R. Petersen |
IEEE Trans. Cybern. | 2 |
| 2022 | Lifelong Incremental Reinforcement Learning With Online Bayesian InferenceabstractA central capability of a long-lived reinforcement learning (RL) agent is to incrementally adapt its behavior as its environment changes and to incrementally build upon previous experiences to facilitate future learning in real-world scenarios. In this article, we propose lifelong incremental reinforcement learning (LLIRL), a new incremental algorithm for efficient lifelong adaptation to dynamic environments. We develop and maintain a library that contains an infinite mixture of parameterized environment models, which is equivalent to clustering environment parameters in a latent space. The prior distribution over the mixture is formulated as a Chinese restaurant process (CRP), which incrementally instantiates new environment models without any external information to signal environmental changes in advance. During lifelong learning, we employ the expectation-maximization (EM) algorithm with online Bayesian inference to update the mixture in a fully incremental manner. In EM, the E-step involves estimating the posterior expectation of environment-to-cluster assignments, whereas the M-step updates the environment parameters for future learning. This method allows for all environment models to be adapted as necessary, with new models instantiated for environmental changes and old models retrieved when previously seen environments are encountered again. Simulation experiments demonstrate that LLIRL outperforms relevant existing methods and enables effective incremental adaptation to various dynamic environments for lifelong learning. Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Path Planning for Cellular-Connected UAV: A DRL Solution With Quantum-Inspired Experience ReplayabstractIn cellular-connected unmanned aerial vehicle (UAV) network, a minimization problem on the weighted sum of time cost and expected outage duration is considered. Taking advantage of UAV’s adjustable mobility, a UAV navigation approach is formulated to achieve the aforementioned optimization goal. Conventional offline optimization techniques suffer from inefficiency in accomplishing the formulated UAV navigation task due to the practical consideration of local building distribution and directional antenna radiation pattern. Alternatively, after mapping the navigation task into a Markov decision process (MDP), a deep reinforcement learning (DRL)-aided solution is proposed to help the UAV find the optimal flying direction within each time slot, and thus the designed trajectory towards the destination can be generated. To help the DRL agent commit a better trade-off between sampling priority and diversity, a novel quantum-inspired experience replay (QiER) framework is proposed, via relating experienced transition’s importance to its associated quantum bit (qubit) and applying Grover iteration based amplitude amplification technique. Compared to several representative DRL-related and non-learning baselines, the effectiveness and supremacy of the proposed DRL-QiER solution are demonstrated and validated in numerical results. Yuanjian Li, Hamid Aghvami, Daoyi Dong |
IEEE Trans. Wirel. Commun. | 3 |
| 2021 | A Modified Deep Q-Learning Algorithm for Control of Two-qubit SystemsabstractQuantum control refers to the manipulation of dynamical quantum systems to force them to complete given tasks such as preparing a desired state and tracking a designed trajectory. We consider the state preparation problem of a two-qubit closed quantum system from an initial to a desired state. The aim is to achieve a high fidelity in a given fixed time with limited control resources. The Deep Q-learning (DQL) for solving the quantum state preparation problem is explored in this paper. We propose a novel semi-Markov DQL algorithm based on a modified action selection procedure and improved replay memory to enhance performance of standard DQL algorithm. The proposed algorithm shows high performance for discovering high-fidelity control protocols and for converging to a good policy, compared with standard DQL. The proposed modifications enhance the exploration-exploitation ability for DQL agent and the robustness for solving quantum control problem with high-fidelity at different numbers of control steps. Numerical results on a two-qubit closed system show effectiveness of the proposed algorithm. Omar Shindi, Qi Yu 0006, Daoyi Dong |
SMC | 3 |
| 2021 | Design of a Discrete-Time Fault-Tolerant Quantum Filter and Fault DetectorabstractThis paper solves the problem of discrete-time fault-tolerant quantum filtering for a class of laser-atom open quantum systems subject to the stochastic faults. We show that by using the discrete-time quantum measurements, optimal estimates of both the atomic observables and the classical fault process can be simultaneously determined in terms of recursive quantum stochastic difference equations. A dispersive interaction quantum system example is used to demonstrate the proposed filtering approach. Qing Gao 0001, Daoyi Dong, Ian R. Petersen, Steven X. Ding |
IEEE Trans. Cybern. | 2 |
| 2021 | Two-Stage Estimation for Quantum Detector Tomography: Error Analysis, Numerical and Experimental ResultsabstractQuantum detector tomography is a fundamental technique for calibrating quantum devices and performing quantum engineering tasks. In this paper, a novel quantum detector tomography method is proposed. First, a series of different probe states are used to generate measurement data. Then, using constrained linear regression estimation, a stage-1 estimation of the detector is obtained. Finally, the positive semidefinite requirement is added to guarantee a physical stage-2 estimation. This Two-stage Estimation (TSE) method has computational complexity O(nd2M), where n is the number of d-dimensional detector matrices and M is the number of different probe states. An error upper bound is established, and optimization on the coherent probe states is investigated. We perform simulation and a quantum optical experiment to testify the effectiveness of the TSE method. Yuanlong Wang 0001, Shota Yokoyama, Daoyi Dong, Ian R. Petersen, Elanor Huntington, Hidehiro Yonezawa |
IEEE Trans. Inf. Theory | 3 |
| 2020 | IEDQN: Information Exchange DQN with a Centralized Coordinator for Traffic Signal ControlabstractFinding the optimal control strategy for traffic signals, especially for multi-intersection traffic signals, is still a difficult task. The use of reinforcement learning (RL) algorithms to this problem is greatly limited because of the partially observable and nonstationary environment. In this paper, we study how to eliminate the above influence from the environment through communication among agents. The proposed method, called Information Exchange Deep Q-Network (IEDQN), has a learning communication protocol, which makes each local agent pay unbalanced and asymmetric attention to other agents' information. Besides the protocol, each agent has the ability to abstract local information from its own history data for interacting, which means that the communication can avoid the dependent instant information and it is robust to the potential time delay of communication. Specifically, by alleviating the effects of partial observation, experience replay can recover to good performance. We evaluate IEDQN via simulation experiments in the simulation of urban mobility (SUMO) in a traffic grid, and it outperforms the comparative multi-agent RL (MARL) methods in both efficiency and effectiveness. Donghan Xie, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IJCNN | 4 |
| 2020 | Adaptive Quantum Process Tomography via Linear Regression EstimationabstractThis paper proposes a recursively adaptive tomography protocol to improve the precision of quantum process estimation for finite dimensional systems. The problem of quantum process tomography is firstly formulated as a parameter estimation problem which can then be solved by the linear regression estimation method. An adaptive algorithm is proposed for the selection of subsequent input states given the previous estimation results. Numerical results show that the proposed adaptive process tomography protocol can achieve an improved level of estimation performance. Qi Yu 0006, Daoyi Dong, Yuanlong Wang 0001, Ian R. Petersen |
SMC | 2 |
| 2020 | Learning-Based Quantum Robust Control: Algorithm, Applications, and ExperimentsabstractRobust control design for quantum systems has been recognized as a key task in quantum information technology, molecular chemistry, and atomic physics. In this paper, an improved differential evolution algorithm, referred to as multiple-samples and mixed-strategy DE (msMS_DE), is proposed to search robust fields for various quantum control problems. In msMS_DE, multiple samples are used for fitness evaluation and a mixed strategy is employed for the mutation operation. In particular, the msMS_DE algorithm is applied to the control problems of: 1) open inhomogeneous quantum ensembles and 2) the consensus goal of a quantum network with uncertainties. Numerical results are presented to demonstrate the excellent performance of the improved machine learning algorithm for these two classes of quantum robust control problems. Furthermore, msMS_DE is experimentally implemented on femtosecond (fs) laser control applications to optimize two-photon absorption and control fragmentation of the molecule CH2BrI. The experimental results demonstrate the excellent performance of msMS_DE in searching for effective fs laser pulses for various tasks. Daoyi Dong, Xi Xing, Hailan Ma, Chunlin Chen 0001, Zhixin Liu 0003, Herschel Rabitz |
IEEE Trans. Cybern. | 1 |
| 2019 | Generation of accessible sets for a class of quantum spin networksabstractIn this paper, we consider the modeling of dynamical spin network systems. The system Hamiltonian governs the evolution of the network system and shows the structure of how the element systems are coupled. Probes are employed to measure a set of spins of the network. For a variety of applications, the state space model is a useful way to describe the system dynamics. One important task in establishing a state space model is to obtain an accessible set containing all the operators coupled to the measurement operators. We provide analytic results on simplifying the process of generating accessible sets. An example is provided to demonstrate the generation of an accessible set under a certain measurement scheme. Qi Yu 0006, Yuanlong Wang 0001, Daoyi Dong, Guo-Yong Xiang |
SMC | 3 |
| 2018 | Stability of a Class of Linear Quantum Feedback Systems with Time DelaysabstractThe bounded stability problem of linear mixed quantum-classical feedback control systems with time delays is investigated in this paper. The whole feedback system is composed of a quantum plant, a classical controller, and some interconnection devices. Two classes of feedback delays in the feedback loop are considered: constant delays and time-varying delays. The stability of the closed-loop systems under these two classes of time delays is analyzed by constructing different Lyapunov-Krasovskii functions and the corresponding stability criteria are derived, respectively. In particular, a weighting matrix is introduced into stability analysis and the obtained stability results are less conservative than relevant results for this class of quantum systems. Sen Kuang, Xiujuan Lu, Daoyi Dong |
SMC | 3 |
| 2018 | Quantum Filtering for a Qubit System Subject to Classical DisturbancesabstractIn this paper, we consider the filtering problem for a hybrid system where a quantum qubit system is disturbed by a classical signal. The quantum filtering theory, which is based on quantum probability theory, can not be directly applied to a hybrid system where a classical stochastic process is also needed in describing the system dynamics. An optical cavity system is employed to model the classical disturbance. By designing the parameters of the auxiliary cavity system, the expectation of the quadrature operator of the cavity shares the same dynamics with the classical signal. With this correspondence guaranteed, one can obtain the real time expectation of the classical signal. The quantum concatenation product is adopted to describe the quantum system which contains both the qubit subsystem and the cavity subsystem. A stochastic master equation, which provides estimates for the quantum state and the classical signal, is given. To reduce the computational complexity, the quantum extended Kalman filter is also applied to this system. Qi Yu 0006, Daoyi Dong, Ian R. Petersen |
SMC | 2 |
| 2018 | Self-Paced Prioritized Curriculum Learning With Coverage Penalty in Deep Reinforcement LearningabstractIn this paper, a new training paradigm is proposed for deep reinforcement learning using self-paced prioritized curriculum learning with coverage penalty. The proposed deep curriculum reinforcement learning (DCRL) takes the most advantage of experience replay by adaptively selecting appropriate transitions from replay memory based on the complexity of each transition. The criteria of complexity in DCRL consist of self-paced priority as well as coverage penalty. The self-paced priority reflects the relationship between the temporal-difference error and the difficulty of the current curriculum for sample efficiency. The coverage penalty is taken into account for sample diversity. With comparison to deep Q network (DQN) and prioritized experience replay (PER) methods, the DCRL algorithm is evaluated on Atari 2600 games, and the experimental results show that DCRL outperforms DQN and PER on most of these games. More results further show that the proposed curriculum training paradigm of DCRL is also applicable and effective for other memory-based deep reinforcement learning approaches, such as double DQN and dueling network. All the experimental results demonstrate that DCRL can achieve improved training efficiency and robustness for deep reinforcement learning. Zhipeng Ren, Daoyi Dong, Huaxiong Li, Chunlin Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Rapid control of two-qubit systems based on measurement feedbackabstractFor two-qubit systems with a degenerate measurement operator, this paper presents a control approach to achieve rapid stabilization to a target Bell state. This approach uses two control channels where the control laws associated with these two channels are designed as a constant and a switching control law, respectively. In the control scheme, the system state space is divided into two parts: the set containing the target state, and its complementary set. We design corresponding control laws in these two sets via different Lyapunov functions. In this scheme, the system trajectory switches only one or two times between the two sets and the convergence rate to the target state can be improved. We provide the conditions on the control Hamiltonians to stabilize the target state by analyzing the stability of the closed-loop system. Xiaqing Sun, Sen Kuang, Daoyi Dong |
SMC | 3 |
| 2017 | Robust Learning Control Design for Quantum Unitary TransformationsabstractRobust control design for quantum unitary transformations has been recognized as a fundamental and challenging task in the development of quantum information processing due to unavoidable decoherence or operational errors in the experimental implementation of quantum operations. In this paper, we extend the systematic methodology of sampling-based learning control (SLC) approach with a gradient flow algorithm for the design of robust quantum unitary transformations. The SLC approach first uses a "training" process to find an optimal control strategy robust against certain ranges of uncertainties. Then a number of randomly selected samples are tested and the performance is evaluated according to their average fidelity. The approach is applied to three typical examples of robust quantum transformation problems including robust quantum transformations in a three-level quantum system, in a superconducting quantum circuit, and in a spin chain system. Numerical results demonstrate the effectiveness of the SLC approach and show its potential applications in various implementation of quantum unitary transformations. Chengzhi Wu, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Cybern. | 4 |
| 2017 | Quantum Ensemble Classification: A Sampling-Based Learning Control ApproachabstractQuantum ensemble classification (QEC) has significant applications in discrimination of atoms (or molecules), separation of isotopes, and quantum information extraction. However, quantum mechanics forbids deterministic discrimination among nonorthogonal states. The classification of inhomogeneous quantum ensembles is very challenging, since there exist variations in the parameters characterizing the members within different classes. In this paper, we recast QEC as a supervised quantum learning problem. A systematic classification methodology is presented by using a sampling-based learning control (SLC) approach for quantum discrimination. The classification task is accomplished via simultaneously steering members belonging to different classes to their corresponding target states (e.g., mutually orthogonal states). First, a new discrimination method is proposed for two similar quantum systems. Then, an SLC method is presented for QEC. Numerical results demonstrate the effectiveness of the proposed approach for the binary classification of two-level quantum ensembles and the multiclass classification of multilevel quantum ensembles. Chunlin Chen 0001, Daoyi Dong, Ian R. Petersen, Herschel Rabitz |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Lower Bounds on the Proportion of Leaders Needed for Expected Consensus of 3-D FlocksabstractThis paper considers the consensus behavior of a spatially distributed 3-D dynamical network composed of heterogeneous agents: leaders and followers, in which the leaders have the preferred information about the destination, while the followers do not have. All followers move in a 3-D Euclidean space with a given speed and with their headings updated according to the average velocity of the corresponding neighbors. Compared with the 2-D model, a key point lies in how to analyze the dynamical behavior of a "linear" nonhomogeneous equation where the nonhomogeneous term strongly nonlinearly depends on the states of all agents. Using the network structure and the estimation of some characteristics for the initial states, we present a proper decaying rate for the nonhomogeneous term and then establish lower bounds on the ratio of the number of leaders to the number of followers that is needed for the expected consensus by considering two cases: 1) fixed speed and neighborhood radius and 2) variable speed and neighborhood radius with respect to the population size. Some simulation examples are given to justify the theoretical results.This paper considers the consensus behavior of a spatially distributed 3-D dynamical network composed of heterogeneous agents: leaders and followers, in which the leaders have the preferred information about the destination, while the followers do not have. All followers move in a 3-D Euclidean space with a given speed and with their headings updated according to the average velocity of the corresponding neighbors. Compared with the 2-D model, a key point lies in how to analyze the dynamical behavior of a "linear" nonhomogeneous equation where the nonhomogeneous term strongly nonlinearly depends on the states of all agents. Using the network structure and the estimation of some characteristics for the initial states, we present a proper decaying rate for the nonhomogeneous term and then establish lower bounds on the ratio of the number of leaders to the number of followers that is needed for the expected consensus by considering two cases: 1) fixed speed and neighborhood radius and 2) variable speed and neighborhood radius with respect to the population size. Some simulation examples are given to justify the theoretical results. Xuejing Li, Lin Wang 0022, Zhixin Liu 0003, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | Learning control of population transfer between subspaces of quantum systems using an adaptive target schemeabstractAn adaptive target scheme is implemented for learning control of population transfer between subspaces of quantum systems. In this control scheme, the target state is updated according to the renormalized yield in the desired subspace throughout the learning iterations, to obtain the desired laser control field. In the numerical experiments, we perform learning control simulations based on a V-type three-subspace quantum system. The field obtained by learning control can transfer the population to the target subspace with high probability. In comparison with a fixed target state, this adaptive target scheme proves to be more efficient for the quantum control problem under consideration. Daoyi Dong, Ian R. Petersen |
IJCNN | 2 |
| 2016 | Synchronization of a Group of Mobile Agents With Variable Speeds Over Proximity NetsabstractThis paper focuses on the synchronization analysis of a class of multiagent systems, where both speed and heading of each agent depend on the states of its local neighbors. The neighbors are defined through the distance between agents and all agents are interconnected via proximity nets. In the variable speed model, the speed of each agent depends on the polarization order of its neighbors in a power-law manner, and the heading is updated according to the average heading of its neighbors. Therefore, the speeds, headings, and positions of all agents are strongly coupled together. For the uniformly and independently distributed initial states, we provide sufficient conditions, imposed only on model parameters, to guarantee synchronization of the variable speed model in the following two cases: 1) the maximum speed and the neighborhood radius are fixed constants and 2) the maximum speed and the neighborhood radius are changing with the population size. Our results reveal that the permitted maximum speed in the variable speed model can be larger than that in the relevant constant speed model. Zhixin Liu 0003, Lin Wang 0022, Daoyi Dong |
IEEE Trans. Cybern. | 4 |
| 2015 | Differential Evolution with Equally-Mixed Strategies for Robust Control of Open Quantum SystemsabstractRobust control of open quantum systems from one state to another is much more difficult than closed quantum systems as a result of system-environment interactions. In this paper, we adopt the sampling-based learning control approach with the motivation of utilizing some artificial samples instead of unknown uncertainties to design an optimal control field against parameter fluctuations. To enhance the learning performance, we introduce an improved differential evolution (DE) algorithm with equally-mixed strategies in the training step of the control design for open quantum systems. Numerical results verify the effectiveness of the proposed equally-mixed strategies DE (EMSDE) algorithm regarding the control design for open quantum systems with uncertainties. Hailan Ma, Chunlin Chen 0001, Daoyi Dong |
SMC | 3 |
| 2015 | Robust Quantum Operation for Two-Level Systems Using Sampling-Based Learning ControlabstractRobust control design for operation of quantum systems has been considered as a demanding and challenging task in the development of quantum technologies. In this paper, we apply the sampling-based learning control (SLC) approach to design a control law for manipulating two-level quantum systems with uncertainties. The gradient-based learning and optimization algorithm is adopted to find the optimal piece-wise control fields for an augmented system by sampling the domain of uncertainties. Numerical results demonstrate the effectiveness of the proposed method for unitary operation of two-level quantum systems even when there are large uncertainties. Chengzhi Wu, Chunlin Chen 0001, Daoyi Dong |
SMC | 4 |
| 2015 | Universal Fuzzy Models and Universal Fuzzy Controllers for Discrete-Time Nonlinear SystemsabstractThis paper investigates the problems of universal fuzzy model and universal fuzzy controller for discrete-time nonaffine nonlinear systems (NNSs). It is shown that a kind of generalized T-S fuzzy model is the universal fuzzy model for discrete-time NNSs satisfying a sufficient condition. The results on universal fuzzy controllers are presented for two classes of discrete-time stabilizable NNSs. Constructive procedures are provided to construct the model reference fuzzy controllers. The simulation example of an inverted pendulum is presented to illustrate the effectiveness and advantages of the proposed method. These results significantly extend the approach for potential applications in solving complex engineering problems. Qing Gao 0001, Gang Feng 0001, Daoyi Dong, Lu Liu 0002 |
IEEE Trans. Cybern. | 3 |
| 2014 | Sampling-based learning control for quantum discrimination and ensemble classificationabstractQuantum ensemble classification has significant applications in discrimination of atoms (or molecules), separation of isotopic molecules and quantum information extraction. In this paper, we recast quantum ensemble classification as a supervised quantum learning problem. A systematic classification methodology is presented by using a sampling-based learning control (SLC) approach for quantum discrimination. The classification task is accomplished via simultaneously steering members belonging to different classes to their corresponding target states (e.g., mutually orthogonal states). Numerical results demonstrate the effectiveness of the proposed approach for the discrimination of two quantum systems and the binary classification of two-level quantum ensembles. Chunlin Chen 0001, Daoyi Dong, Ian R. Petersen, Herschel Rabitz |
IJCNN | 2 |
| 2014 | Fidelity-Based Probabilistic Q-Learning for Control of Quantum SystemsabstractThe balance between exploration and exploitation is a key problem for reinforcement learning methods, especially for Q-learning. In this paper, a fidelity-based probabilistic Q-learning (FPQL) approach is presented to naturally solve this problem and applied for learning control of quantum systems. In this approach, fidelity is adopted to help direct the learning process and the probability of each action to be selected at a certain state is updated iteratively along with the learning process, which leads to a natural exploration strategy instead of a pointed one with configured parameters. A probabilistic Q-learning (PQL) algorithm is first presented to demonstrate the basic idea of probabilistic action selection. Then the FPQL algorithm is presented for learning control of quantum systems. Two examples (a spin-1/2 system and a Λ-type atomic system) are demonstrated to test the performance of the FPQL algorithm. The results show that FPQL algorithms attain a better balance between exploration and exploitation, and can also avoid local optimal policies and accelerate the learning process. Chunlin Chen 0001, Daoyi Dong, Han-Xiong Li, Jian Chu, Tzyh Jong Tarn |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Agent-Based Self-Adaptable Context-Aware Network Vulnerability AssessmentabstractImmunology inspired computer security has attracted enormous attention as its potential impacts on the next generation service-oriented network operation system. In this paper, we propose a new agent-based threat awareness assessment strategy inspired by the human immune system to dynamically adapt against attacks. Specifically, this approach is based on the dynamic reconfiguration of the file access right for system calls or logs (e.g., file rewritability) with balanced adaptability and vulnerability. Based on an information-theoretic analysis on the coherently associations of adaptability, autonomy as well as vulnerability, a generic solution is suggested to break down their coherent links. The principle is to maximize context-situation awared systems' adaptability and reduce systems' vulnerability simultaneously. Experimental results show the efficiency of the proposed biological behaviour-inspired vulnerability awareness system. Frank Jiang 0001, Daoyi Dong, Longbing Cao, Michael R. Frater |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2012 | Control Design of Uncertain Quantum Systems With Fuzzy EstimatorsabstractAn approach of control design using fuzzy estimators (FEs) is proposed for quantum systems with uncertainties. Two types of quantum control problems are considered: 1) control of a pure-state quantum system in the presence of uncertainties and 2) control design of quantum systems with initial mixed states and uncertainties. For the first type of tasks, a partial feedback control scheme with an FE is presented to design controllers. In this scheme, an FE is trained to estimate the quantum state for feedback control of a quantum system, and controlled projective measurement is used to assist in controlling the system. For the second type of quantum control tasks, a probabilistic fuzzy estimator (PFE) is trained to estimate the quantum state for control design of a quantum system with an initial mixed state, and a corresponding control algorithm is proposed to design a control law that drives the system from the mixed state to a target pure state. Two examples of two-spin-1/2systems are also presented and analyzed to demonstrate the process of control design and potential applications of the proposed approach. Chunlin Chen 0001, Daoyi Dong, James Lam, Jian Chu, Tzyh Jong Tarn |
IEEE Trans. Fuzzy Syst. | 2 |
| 2011 | Hybrid MDP based integrated hierarchical Q-learning
Chunlin Chen 0001, Daoyi Dong, Han-Xiong Li, Tzyh Jong Tarn |
Sci. China Inf. Sci. | 2 |
| 2008 | Quantum Reinforcement LearningabstractThe key approaches for machine learning, particularly learning in unknown probabilistic environments, are new representations and computation mechanisms. In this paper, a novel quantum reinforcement learning (QRL) method is proposed by combining quantum theory and reinforcement learning (RL). Inspired by the state superposition principle and quantum parallelism, a framework of a value-updating algorithm is introduced. The state (action) in traditional RL is identified as the eigen state (eigen action) in QRL. The state (action) set can be represented with a quantum superposition state, and the eigen state (eigen action) can be obtained by randomly observing the simulated quantum state according to the collapse postulate of quantum measurement. The probability of the eigen action is determined by the probability amplitude, which is updated in parallel according to rewards. Some related characteristics of QRL such as convergence, optimality, and balancing between exploration and exploitation are also analyzed, which shows that this approach makes a good tradeoff between exploration and exploitation using the probability amplitude and can speedup learning through the quantum parallelism. To evaluate the performance and practicability of QRL, several simulated experiments are given, and the results demonstrate the effectiveness and superiority of the QRL algorithm for some complex problems. This paper is also an effective exploration on the application of quantum computation to artificial intelligence. Daoyi Dong, Chunlin Chen 0001, Han-Xiong Li, Tzyh Jong Tarn |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2008 | Incoherent Control of Quantum Systems With Wavefunction-Controllable Subspaces via Quantum Reinforcement LearningabstractIn this paper, an incoherent control scheme for accomplishing the state control of a class of quantum systems which have wavefunction-controllable subspaces is proposed. This scheme includes the following two steps: projective measurement on the initial state and learning control in the wavefunction-controllable subspace. The first step probabilistically projects the initial state into the wavefunction-controllable subspace. The probability of success is sensitive to the initial state; however, it can be greatly improved through multiple experiments on several identical initial states even in the case with a small probability of success for an individual measurement. The second step finds a local optimal control sequence via quantum reinforcement learning and drives the controlled system to the objective state through a set of suitable controls. In this strategy, the initial states can be unknown identical states, the quantum measurement is used as an effective control, and the controlled system is not necessarily unitarily controllable. This incoherent control scheme provides an alternative quantum engineering strategy for locally controllable quantum systems. Daoyi Dong, Chunlin Chen 0001, Tzyh Jong Tarn, Alexander N. Pechen, Herschel Rabitz |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2006 | Grey Reinforcement Learning for Incomplete Information Processing
Chunlin Chen 0002, Daoyi Dong, Zonghai Chen |
TAMC | 2 |