Qing-Shan Jia

dblp:09/3139 · also Qing-Shan (Samuel) Jia · DBLP profile ↗
← Back
57ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-4683-7215ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 7 since 2021Computer networks · 4Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Consensus-Based Distributed Reinforcement Learning With Primal-Dual Update for Networked Microgrids On-Line Coordination
abstract
This paper develops a distributed reinforcement learning (RL) method to coordinate cooperative microgrids (MGs). The high uncertainty of power loads and renewable energy sources motivate the operator to perform real-time dispatch. On the one hand, the existing online methods usually utilize approximate models that result in intractable constraint violation. A common method is to relax it as a chance constraint, while it is still hard to ensure its satisfaction in practice. On the other hand, some MGs may hope to preserve the private information on their local costs and states. To address these problems, we make the following contributions. First, the coordination problem is reformulated as a constrained multi-agent Markov decision process. Second, the distributed RL algorithm with a theoretical convergence guarantee is developed. Third, to further preserve the local private information and improve the performance, this algorithm is modified by adding a local feature extraction module for each agent. This module could also be regarded as an encryption module for the local state information. Fourth, numerical experiments are carried out to validate the effectiveness of the modified algorithm.
Gaochen Cui, Qing-Shan Jia, Xiaohong Guan, Qiaozhu Zhai, Xianping Guo
IEEE Trans Autom. Sci. Eng.2
2026 A Bayesian Optimization Method for Design Space Exploration of Processors With Efficient Evaluation Budget Allocation
abstract
Design space exploration (DSE), which aims to optimize the parameter configurations for the processor with a limited evaluation budget, is an important but challenging problem. Processor performance for a given parameter configuration can only be obtained through expensive evaluations. A key characteristic of processor DSE is that the optimization objective is the benchmark score, which is typically computed as a weighted combination of multiple sub-benchmark scores. Each sub-benchmark score can only be obtained through a costly evaluation. Bayesian optimization (BO) is widely used for DSE through an iterative process, where a surrogate model is constructed and updated to guide the selection of parameter configurations in each iteration. However, during the optimization process, it is not always necessary to evaluate all subbenchmarks for every recommended parameter configuration. In some cases, evaluating only a subset of sub-benchmarks can still provide significant improvements to the surrogate model. This paper proposes a BO-based algorithm that effectively utilizes the evaluation budget for efficient optimization. In each iteration, the evaluation budget is divided into two parts: exploitation and exploration. For the exploitation part, we propose a diversity-aware horse racing selection rule and use ordinal optimization to identify configurations with a high likelihood of being optimal. They are evaluated on all sub-benchmarks to obtain the total scores. For the exploration part, we propose a conditional entropy-based greedy algorithm to allocate the evaluation budget across parameter configurations and their corresponding sub-benchmarks, which can effectively improve the surrogate model. We conduct a comprehensive numerical experiment to validate the effectiveness of both components of our algorithm. Furthermore, experiments on three real-world DSE problems demonstrate that our algorithm outperforms state-of-the-art Bayesian optimization methods.
Yuhang Zhu 0001, Mingyuan Hou, Xiaoliang Lv, Qing-Shan Jia, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.4
2025 CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
abstract
Preference-based reinforcement learning (PbRL) bypasses explicit reward engineering by inferring reward functions from human preference comparisons, enabling better alignment with human intentions. However, humans often struggle to label a clear preference between similar segments, reducing label efficiency and limiting PbRL’s real-world applicability. To address this, we propose an offline PbRL method: Contrastive LeArning for ResolvIng Ambiguous Feedback (CLARIFY), which learns a trajectory embedding space that incorporates preference information, ensuring clearly distinguished segments are spaced apart, thus facilitating the selection of more unambiguous queries. Extensive experiments demonstrate that CLARIFY outperforms baselines in both non-ideal teachers and real human feedback settings. Our approach not only selects more distinguished queries but also learns meaningful trajectory embeddings.
Ni Mu, Hao Hu 0006, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia
ICML6
2025 S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning
abstract
Preference-based reinforcement learning (PbRL) stands out by utilizing human preferences as a direct reward signal, eliminating the need for intricate reward engineering. However, despite its potential, traditional PbRL methods are often constrained by the indistinguishability of segments, which impedes the learning process. In this paper, we introduce Skill-Enhanced Preference Optimization Algorithm (S-EPOA), which addresses the segment indistinguishability issue by integrating skill mechanisms into the preference learning framework. Specifically, we first conduct the unsupervised pretraining to learn useful skills. Then, we propose a novel query selection mechanism to balance the information gain and distinguishability over the learned skill space. Experimental results on a range of tasks, including robotic manipulation and locomotion, demonstrate that S-EPOA significantly outperforms conventional PbRL methods in terms of both robustness and learning efficiency. The results highlight the effectiveness of skill-driven learning in overcoming the challenges posed by segment indistinguishability.
Ni Mu, Yao Luan 0001, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia
IJCAI5
2025 STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
abstract
Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where agents sequentially perform sub-tasks (e.g., navigation, grasping), is limited by **stage misalignment**: Comparing segments from mismatched stages, such as movement versus manipulation, results in uninformative feedback, thus hindering policy learning. In this paper, we validate the stage misalignment issue through theoretical analysis and empirical experiments. To address this issue, we propose **ST**age-**A**l**I**gned **R**eward learning (STAIR), which first learns a stage approximation based on temporal distance, then prioritizes comparisons within the same stage. Temporal distance is learned via contrastive learning, which groups temporally close states into coherent stages, without predefined task knowledge, and adapts dynamically to policy changes. Extensive experiments demonstrate STAIR's superiority in multi-stage tasks and competitive performance in single-stage tasks. Furthermore, human studies show that stages approximated by STAIR are consistent with human cognition, confirming its effectiveness in mitigating stage misalignment.
Yao Luan 0001, Ni Mu, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia
NeurIPS5
2025 Multi-Agent Reinforcement Learning With Decentralized Distribution Correction
abstract
This work considers decentralized multi-agent reinforcement learning (MARL), where the global states and rewards are assumed to be fully observable, while the local behavior policy is preserved locally for resisting adversarial attack. In order to cooperatively accumulate more rewards, the agents exchange messages among a time-varying communication network to reach consensus. For these cooperative tasks, we propose a decentralized actor-critic algorithm, where the agents make individual decisions, but the joint behavior policy is optimized towards more cumulative rewards. We provide the theoretical analysis towards the convergence under the tabular setting and then expand it to nonlinear function approximations. Furthermore, by incorporating decentralized distribution correction, the agents are trained in an off-policy manner for higher sample efficiency. Finally, we conduct experiments to evaluate the algorithms, where the proposed algorithm performs competitively in both stability and asymptotic performance.Note to Practitioners—Fully decentralized MARL algorithms are widely applied in multi-agent systems for generating cooperative behaviors, e.g., multiple unmanned aerial vehicles (UAV) cooperatively performing search and rescue tasks, multiple vehicles efficiently passing a crowded intersection, and multiple robots cooperatively handling cargo or obstacles. Focusing on these potential applications, this work is motivated to improve the sample efficiency of recent decentralized MARL algorithms by incorporating off-policy training approaches. In this work, we reweight historical trajectories via a decentralized average consensus step and develop corresponding policy-optimization procedures, with which previous trajectories could be used to stabilize later iterations. Since the training materials are augmented by historical samples, the sample efficiency is significantly improved, and the training process is stabilized. With the fully decentralized training approach, the proposed algorithms are expected to be applied in large-scale systems, e.g., vehicle teams and UAV groups, for effective real-time control.
Qing-Shan Jia
IEEE Trans Autom. Sci. Eng.2
2025 An Efficient Real-Time Railway Container Yard Management Method Based on Partial Decoupling
abstract
Sea-rail intermodal transportation is an essential infrastructure in global supply chains nowadays. Since the yard is the interface between sea and land, optimizing the transportation process in yards is of significant interest for increasing transportation efficiency. However, the yard management problem usually suffers from the large state space and cannot be solved effectively. We consider this important problem from a decomposition perspective, focusing on real-time scheduling in railway container yards, and make the following contributions. We show the partial decomposition property of the yard management problem. This property establishes the equivalence between joint optimization of yard components and independent optimization of each component under mild assumptions, achieving optimality and effectiveness simultaneously. Utilizing this property, we further simplify each subproblem while preserving the optima, and develop algorithms for each simplified subproblem. Numerical experiments demonstrate the performance of proposed algorithms and the contribution of optimizing each subproblem from various perspectives.
Yao Luan 0001, Qing-Shan Jia, Yi Xing
IEEE Trans Autom. Sci. Eng.2
2025 Proactive Robust Hardening of Resilient Power Distribution Network: Decision-Dependent Uncertainty Modeling and Fast Solution Strategy
abstract
To address the power system hardening problem, traditional approaches often adopt robust optimization (RO) that considers a fixed set of concerned contingencies, regardless of the fact that hardening some components actually renders relevant contingencies impractical. In this paper, we directly adopt a dynamic uncertainty set that explicitly incorporates the impact of hardening decisions on the worst-case contingencies, which leads to a decision-dependent uncertainty (DDU) set. Then, a DDU-based robust-stochastic optimization (DDU-RSO) model is proposed to support the hardening decisions on distribution lines and distributed generators (DGs). Also, the randomness of load variations and available storage levels is considered through stochastic programming (SP) in the innermost level problem. Various corrective measures (e.g., the joint scheduling of DGs and energy storage) are included, coupling with a finite support of stochastic scenarios, for resilience enhancement. To relieve the computation burden of this new hardening formulation, an enhanced customization of parametric column-and-constraint generation (P-C&CG) algorithm is developed. By leveraging the network structural information, the enhancement strategies based onresilience importance indicesare designed to improve the convergence performance. Numerical results on 33-bus and 118-bus test distribution networks have demonstrated the effectiveness of DDU-RSO aided hardening scheme. Furthermore, in comparison to existing solution methods, the enhanced P-C&CG has achieved a superior performance by reducing the solution time by a few orders of magnitude.
Donglai Ma, Bo Zeng 0001, Qing-Shan Jia, Chen Chen 0007, Qiaozhu Zhai, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.4
2025 Preference-Based Multi-Objective Reinforcement Learning
abstract
Multi-objective reinforcement learning (MORL) is a structured approach for optimizing tasks with multiple objectives. However, it often relies on pre-defined reward functions, which can be hard to design for balancing conflicting goals and may lead to oversimplification. Preferences can serve as more flexible and intuitive decision-making guidance, eliminating the need for complicated reward design. This paper introduces preference-based MORL (Pb-MORL), which formalizes the integration of preferences into the MORL framework. We theoretically prove that preferences can derive policies across the entire Pareto frontier. To guide policy optimization using preferences, our method constructs a multi-objective reward model that aligns with the given preferences. We further provide theoretical proof to show that optimizing this reward model is equivalent to training the Pareto optimal policy. Extensive experiments in benchmark multi-objective tasks, a multi-energy management task, and an autonomous driving task on a multi-line highway show that our method performs competitively, surpassing the oracle method, which uses the ground truth reward function. This highlights its potential for practical applications in complex real-world systems.
Ni Mu, Yao Luan 0001, Qing-Shan Jia
IEEE Trans Autom. Sci. Eng.3
2024 Query-Policy Misalignment in Preference-Based Reinforcement Learning
abstract
Preference-based reinforcement learning (PbRL) provides a natural way to align RL agents’ behavior with human desired outcomes, but is often restrained by costly human feedback. To improve feedback efficiency, most existing PbRL methods focus on selecting queries to maximally improve the overall quality of the reward model, but counter-intuitively, we find that this may not necessarily lead to improved performance. To unravel this mystery, we identify a long-neglected issue in the query selection schemes of existing PbRL studies: Query-Policy Misalignment. We show that the seemingly informative queries selected to improve the overall quality of reward model actually may not align with RL agents’ interests, thus offering little help on policy learning and eventually resulting in poor feedback efficiency. We show that this issue can be effectively addressed via policy-aligned query and a specially designed hybrid experience replay, which together enforce the bidirectional query-policy alignment. Simple yet elegant, our method can be easily incorporated into existing approaches by changing only a few lines of code. We showcase in comprehensive experiments that our method achieves substantial gains in both human feedback and RL sample efficiency, demonstrating the importance of addressing query-policy misalignment in PbRL tasks.
Xianyuan Zhan, Qing-Shan Jia, Ya-Qin Zhang
ICLR4
2024 Constrained reinforcement learning with statewise projection: a control barrier function approach
Xinze Jin, Qing-Shan Jia
Sci. China Inf. Sci.3
2024 Two-Phase on-Line Joint Scheduling for Welfare Maximization of Charging Station
abstract
The widespread adoption of EVs brings practical interest to the operation optimization of the charging station. This paper considers the joint scheduling of pricing and charging control to enhance the operational capability of the station. The following contributions are made. First, a joint scheduling model of pricing and charging control is developed to maximize the expected social welfare of the charging station considering the quality of service and the price fluctuation sensitivity of EV drivers. It is formulated as a Markov decision process with a variance criterion to capture uncertainties during operation. Second, a two-phase on-line policy learning algorithm is proposed to solve this joint scheduling problem. In the first phase, it implements event-based policy iteration to find the optimal pricing scheme. In the second phase, it implements scenario-based model predictive control for smart charging under the updated pricing scheme. Third, by leveraging the performance difference theory, the optimality of the proposed algorithm is theoretically analyzed. The feasibility of the proposed method and the improved social welfare of the charging station are numerically demonstrated based on a typical charging station with distributed generation and storage.Note to Practitioners—The popularization of EVs requires the high-efficiency operation of the charging station. The joint scheduling of pricing and charging control can provide a promising way to achieve the balance between EV drivers’ satisfaction and the profit maximization of the charging station. However, it suffers from EV drivers’ uncertain responses to the pricing scheme and the coupled relationship between pricing and charging control. This multi-stage stochastic programming is non-trivial to solve. In this paper, we propose a two-phase on-line policy learning method for this joint scheduling problem. This algorithm can be implemented in the controller of the charging station. In the first phase, the event-based policy iteration can be implemented to iteratively improve the current best pricing scheme until convergence. In the second phase, the scenario-based model predictive control, which is reformulated as a mixed integer linear programming, can be quickly solved for smart charging. Case studies demonstrate the operation enhancement of the station.
Qilong Huang, Qing-Shan Jia, Xiang Wu 0008, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.2
2024 An OCBA-Based Method for Efficient Sample Collection in Reinforcement Learning
abstract
This work focuses on the sample collection in reinforcement learning (RL), where the interaction with the environment is typically time-consuming and extravagantly expensive. In order to collect samples in a more valuable way, we propose a confidence-based sampling strategy based on the optimal computing budget allocation algorithm (OCBA), which actively allocates the computing efforts to actions with different predictive uncertainties. We estimate the uncertainty with ensembles and generalize them from tabular representations to function approximations. The OCBA-based sampling strategy could be easily integrated into various off-policy RL algorithms, where we take Q-learning, DQN, and SAC as examples to show the incorporation. Besides, we provide the theoretical analysis towards convergence and evaluate the algorithms experimentally. According to the experiments, the incorporated algorithms obtain remarkable gains compared with modern ensemble-based RL algorithms.Note to Practitioners—Reinforcement learning is a powerful tool for handling sequential decision-making problems, e.g., autonomous driving and robotics control, where the behaviors typically have a long-term effect on future events. However, although RL achieves human-level control in some tasks, it severely suffers from low sample efficiency. Therefore, implementing RL in some practical areas, e.g., healthcare and rescue, is extremely hard due to the requirement of massive samples. This work aims to enhance the exploration of RL by incorporating OCBA, which provides an asymptotically optimal data-collection strategy for simulation-based optimization. Based on ensemble-based uncertainty estimation and OCBA-based action selection, the incorporated RL algorithms show competitive performance on many benchmarks and significantly reduce the sampling efforts during iterations.
Xinze Jin, Qing-Shan Jia, Dongchun Ren, Huaxia Xia
IEEE Trans Autom. Sci. Eng.3
2024 Complexity-Based Structural Optimization of Deep Belief Network and Application in Wastewater Treatment Process
abstract
Deep belief network (DBN) is an effective deep learning model, which can learn the complex data by extracting features hierarchically. However, the successful application of DBN depends on the suitable size of the structure (the number of hidden neurons), which is still an open problem. Currently, the network structure size is basically determined by experience with a time-consuming process. In this article, a complexity-based structural optimization (CBSO) algorithm, based on multiobjective ordinal optimization (MOO), is developed for designing the DBN structure. First, the problem formulation of structural optimization of DBN is given, where the multiple objectives are to minimize the fitting error and complexity. Second, the lower bound for alignment probability in optimizing DBN structure is developed according to MOO. Finally, an effective method to maximize the probability of correct select is given to pursue the good tradeoff between the complexity and the performance. The performance of proposed CBSO algorithm is demonstrated via predicting and controlling water quality of wastewater treatment process (WWTP) using the CBSO-DBN-based model predictive control (MPC) strategy. The simulation results show that the resulting CBSO-DBN can find the better structure design by using CBSO algorithm with smaller fitting error and limited computational complexity, and thereby achieve the better performance in WWTP than its peers. Especially, the CBSO-DBN-MPC improves the control accuracy by 76.16% and computational complexity by 50.45%, respectively.
Gongming Wang, Guanghui Yuan, Yuanying Chi, Qing-Shan Jia, Junfei Qiao 0001
IEEE Trans. Ind. Informatics5
2023 Mind the Gap: Offline Policy Optimization for Imperfect Rewards
Haoran Xu 0003, Xianyuan Zhan, Qing-Shan Jia, Ya-Qin Zhang
ICLR6
2023 A Simulation-Based Primal-Dual Approach for Constrained V2G Scheduling in a Microgrid of Building
abstract
The electric vehicle (EV) utilizing vehicle-to-grid (V2G) technology can serve as an important mobile storage in a microgrid of building. Since more and more buildings are equipped with charging piles and distributed renewable energy, it can potentially improve the operation efficiency of building microgrid and reduce the impact of EVs to the grid by V2G scheduling considering the uncertain supply and EV charging demand in the building. We consider this important problem in this paper and make the following contributions. First, we formulate this V2G scheduling problem as a constrained Markov decision process (CMDP). The objective function is to improve the overall building energy operation cost while constraining the expected cycling time to ensure the participation enthusiasm of the EV users. Second, a simulation-based primal-dual approach is developed to decompose the original problem into a continuous optimization subproblem on the supply side and a discrete optimization subproblem on the demand side. The demand side optimization can be further decoupled as a distributed single EV scheduling problem where simulation-based regularized rollout method can be applied to improve from existing base policies. Third, the structural property of the problem and the cost improvement of the proposed method are analyzed to speed up the optimization and ensure the solution quality. Numerical experiments based on real building load data and EV data are conducted to demonstrate the performance of this method.Note to Practitioners—With the rapid adoption of EVs in the microgrid of building, there goes the challenge of how to reduce their charging impact to the microgrid. The V2G scheduling between EVs and building microgrid can provide a promising way to reduce the impact and improve the building operation efficiency. However, it suffers from the uncertain EV charging demand and the conflict between synthesized coordination and ensuring the participation enthusiasm of drivers. In this paper, we propose a constrained V2G scheduling model to address this issue. After formulating the problem as a constrained Markov decision process, a simulation-based primal-dual approach is developed to decompose the problem where the derived control policy in the supply side can be deployed in the building energy controller and the derived control policy in the demand side can be deployed in the controller of each charging pile. Case studies show the performance of the proposed method and its win-win property for the building and EV users.
Qilong Huang, Qing-Shan Jia, Yaowen Qi, Cangqi Zhou, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.3
2023 Event-Driven Model Predictive Control With Deep Learning for Wastewater Treatment Process
abstract
Wastewater treatment processes (WWTPs) have been considered as complex control problems, because effluent water standard, stability and multioperational conditions need to be taken into account. In this article, an event-driven model predictive control with deep learning (EMPC-DL) is proposed for the control problems to improve the running performance of WWTPs. First, several events are defined based on different operational conditions reflected by operational data. Then, an event-driven deep belief network (EDBN) is developed based on deep learning to approximate the nonlinear characteristics of the WWTPs. Second, a quadratic optimization is designed to solve the control law of MPC based on the predictive output of the EDBN. The major advantage of quadratic optimization is its efficiency, which is achieved by an efficient strategy that only needs one-step prediction of EDBN during one-time rolling optimization. Third, this article gives convergence and stability analysis of EMPC-DL. Finally, the feasibility and applicability of EMPC-DL are demonstrated on the benchmark simulation model No. 1 (BSM1). The experimental results show that EMPC-DL achieves the more satisfactory performance in modeling, controlling, and tracking water quality parameters than its peers.
Gongming Wang, Jing Bi 0001, Qing-Shan Jia, Junfei Qiao 0001, Lei Wang 0173
IEEE Trans. Ind. Informatics3
2022 On large action space in EV charging scheduling optimization
Zhaoyu Jiang, Qing-Shan Jia, Xiaohong Guan
Sci. China Inf. Sci.2
2022 Multiagent Dynamic Task Assignment Based on Forest Fire Point Model
abstract
Multiagent dynamic task assignment of forest fires is a complicated optimization problem because it requires the consideration of multiple factors, such as the spread speed of fires, firefighting speed of agents, the movement speed of agents, and the number of deployed agents. In this article, we investigate multiagent dynamic task assignment based on a forest fire point model, the objective of which is to minimize task completion time. First, we establish a model for the spread of fire and dynamic task assignments. Second, we prove that the optimal static task assignment always makes all task completion times the same under certain assumptions. Furthermore, we calculate the optimal solution to the static task assignment problem assuming no travel time for the agents, which provides the theoretical basis for the initial deployment and dynamic deployment. Third, we propose a dynamic task assignment scheme based on the global information, which ensures that every reassignment reduces the task completion time and makes all task completion times close to each other. Finally, the simulation is carried out on the MATLAB platform to verify the performance of the proposed dynamic task assignment scheme by comparing with a multistage global auction algorithm. We hope that this work provides insight for decision-makers designing reasonable assignment strategies based on the model and solving assignment optimization problem in different situations.Note to practitioners—The forest firefighting problem considered in this article is a typical multitask and multistage optimization problem. Many searching algorithms for multistage optimization problem are available in the existing literature. However, one of the main challenges is that the time of searching increases exponentially with the number of stages. This work first proves that the tasks are completed in the minimum amount of time, under the constraint of one-shot assignment. This finding helps us to evaluate the gap between the searching algorithm and the optimal solution. In addition, in practice, if the underlying dynamic process can be modeled or partially modeled, then we can predict the behavior of future stages and reduce the searching domain. If a model is available, then we can also adjust the assignment scheme dynamically based on the principle that each adjustment would reduce the total time of tasks completion. In this article, we establish a dynamical fire-spreading model and propose a model-based solution to the multistage optimization problems. The findings in this work can serve as a supplement to the existing optimization algorithms.
Jie Chen 0079, Yuqian Guo, Zhifeng Qiu, Bin Xin 0002, Qing-Shan Jia, Weihua Gui 0001
IEEE Trans Autom. Sci. Eng.5
2022 Special Issue on the 2020 International Conference on Automation Science and Engineering
abstract
We are pleased to present this Special Issue of TASE, including 12 extended articles selected from the technical program of the 2020 International Conference on Automation Science and Engineering (CASE2020). CASE2020 was held virtually due to the COVID19 pandemics, August 20–21, 2020, and was originally scheduled in Hong Kong, China. CASE is an offspring of TASE and is the flagship automation conference of the IEEE Robotics and Automation Society, constituting the primary forum for cross-industry and multidisciplinary research in automation. The 2020 CASE theme was Automation Analytics, a global challenge emphasized at the conference by several invited and regular sessions, as well as specific workshops.
Mariagrazia Dotoli, Weiming Shen 0001, Qing-Shan Jia, Ray Y. Zhong
IEEE Trans Autom. Sci. Eng.3
2021 Soft-sensing of Wastewater Treatment Process via Deep Belief Network with Event-triggered Learning
Gongming Wang, Qing-Shan Jia, MengChu Zhou, Jing Bi 0001, Junfei Qiao 0001
Neurocomputing2
2021 A Computing Budget Allocation Method for Minimizing EV Charging Cost Using Uncertain Wind Power
abstract
The idea of using wind power to charge electric vehicles (EVs) has attracted more and more attention nowadays due to the potential in significantly reducing air pollution. However, this problem is challenging on account of the uncertainty in the wind power generation and the charging demand from the EVs. Simulation-based policy improvement (SBPI) has been an important method for decision-making in stochastic dynamic programming and, in particular, for charging decisions of EVs in microgrids. However, the problem of allocating the limited computing budget for the best decision-making in online applications is less discussed. We consider this important problem in this work and make the following three major contributions. First, we show that the significant uncertainty in wind power generation forecasting could make the policy that is the outcome of an SBPI worse than the base policy. Second, we apply two existing methods to address this issue, namely, the optimal computing budget allocation (OCBA) for maximizing the probability of correct selection (OCBA_PCS) and the OCBA for minimizing the expected opportunity cost (OCBA_EOC). The asymptotic optimality is briefly reviewed. Third, we numerically compare the performance of OCBA_PCS and OCBA_EOC with the equal allocation (EA), a principle-based method, and a stochastic scenario-based method on small-scale and large-scale experiments. This work sheds light on the EV charging decision in general. Note to Practitioners-Together with the growing adoption of EVs in modern societies, there goes the challenge of how to satisfy the charging demand. Given the high uncertainty both in the wind power generation and in the charging demand, it is important to make decisions online using up-to-date estimation on the renewable power generation and the charging demand. Simulation-based policy improvement (SBPI) is shown both theoretically and practically to be useful to improve a given base policy in various applications, including this EV charging problem. However, the high uncertainty in forecasting could sometimes make the output of SBPI worse than that of the base policy. In this work, we first use numerical experiments to demonstrate the risk for such scenarios. Then, we propose to use two computing budget allocation procedures to address this issue. The asymptotic optimality of both algorithms is briefly reviewed. We demonstrate their performance on numerical experiments when there are only several EVs and when there are 100 EVs.
Zhaoyu Jiang, Qing-Shan Jia, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.2
2021 Deep Learning-Based Model Predictive Control for Continuous Stirred-Tank Reactor System
abstract
A continuous stirred-tank reactor (CSTR) system is widely applied in wastewater treatment processes. Its control is a challenging industrial-process-control problem due to great difficulty to achieve accurate system identification. This work proposes a deep learning-based model predictive control (DeepMPC) to model and control the CSTR system. The proposed DeepMPC consists of a growing deep belief network (GDBN) and an optimal controller. First, GDBN can automatically determine its size with transfer learning to achieve high performance in system identification, and it serves just as a predictive model of a controlled system. The model can accurately approximate the dynamics of the controlled system with a uniformly ultimately bounded error. Second, quadratic optimization is conducted to obtain an optimal controller. This work analyzes the convergence and stability of DeepMPC. Finally, the DeepMPC is used to model and control a second-order CSTR system. In the experiments, DeepMPC shows a better performance in modeling, tracking, and antidisturbance than the other state-of-the-art methods.
Gongming Wang, Qing-Shan Jia, Junfei Qiao 0001, Jing Bi 0001, MengChu Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2020 Optimization of Large-Scale Commercial Electric Vehicles Fleet Charging Location Schedule Under the Distributed wind Power Supply
abstract
The electrification of commercial vehicles is underway which brings practical interest for transportation companies to optimize the charging scheduling of electric vehicles (EVs) to reduce the operating cost. This paper considers the joint optimization problem for the charging location schedule of a commercial EV fleet considering the distributed wind power supply and makes three main contributions. First, a bi-level bipartite graph (BBG) model is first proposed for this problem where the upper level tries to coordinate the wind power among the local wind power generators, and the lower level tries to find out the optimal charging location schedule for all EVs. Second, a bi-level iteration optimization method combining linear programming and Kuhn-Munkres (KM) algorithm is developed for the real-time market operation, and its optimality is proved mathematically. Third, the effectiveness of our method on reducing the operating cost and maintaining a high service rate is demonstrated by a case study compared with other policies.
Qing-Shan Jia
COMPSAC2
2020 Event-based optimization with random packet dropping
Qing-Shan Jia, Jing-Xian Tang, Zhenning Lang
Sci. China Inf. Sci.1
2020 A sparse deep belief network with efficient fuzzy learning framework
Gongming Wang, Qing-Shan Jia, Junfei Qiao 0001, Jing Bi 0001
Neural Networks2
2020 Guest Editorial Special Section on the 2017 International Conference on Automation Science and Engineering
abstract
We are very pleased to present this Special Section of these Transactions, including six extended articles selected from the technical program of the 2017 International Conference on Automation Science and Engineering (CASE 2017), held in Xi’an, China, August 20–23, 2017. CASE is an offspring of the Transactions on Automation Science and Engineering and is the flagship automation conference of the IEEE Robotics and Automation Society, constituting the primary forum for cross-industry and multidisciplinary research in automation.
Mariagrazia Dotoli, Qing-Shan Jia
IEEE Trans Autom. Sci. Eng.2
2020 An Adaptive Deep Belief Network With Sparse Restricted Boltzmann Machines
abstract
Deep belief network (DBN) is an efficient learning model for unknown data representation, especially nonlinear systems. However, it is extremely hard to design a satisfactory DBN with a robust structure because of traditional dense representation. In addition, backpropagation algorithm-based fine-tuning tends to yield poor performance since its ease of being trapped into local optima. In this article, we propose a novel DBN model based on adaptive sparse restricted Boltzmann machines (AS-RBM) and partial least square (PLS) regression fine-tuning, abbreviated as ARP-DBN, to obtain a more robust and accurate model than the existing ones. First, the adaptive learning step size is designed to accelerate an RBM training process, and two regularization terms are introduced into such a process to realize sparse representation. Second, initial weight derived from AS-RBM is further optimized via layer-by-layer PLS modeling starting from the output layer to input one. Third, we present the convergence and stability analysis of the proposed method. Finally, our approach is tested on Mackey-Glass time-series prediction, 2-D function approximation, and unknown system identification. Simulation results demonstrate that it has higher learning accuracy and faster learning speed. It can be used to build a more robust model than the existing ones.
Gongming Wang, Junfei Qiao 0001, Jing Bi 0001, Qing-Shan Jia, MengChu Zhou
IEEE Trans. Neural Networks Learn. Syst.4
2019 Decentralized EV-Based Charging Optimization With Building Integrated Wind Energy
abstract
Electric vehicles (EVs) have experienced a rapid growth due to the economic and environmental benefits. However, the substantial charging load brings challenging issues to the power grid. Modern technological advances and the huge number of high-rise buildings have promoted the development of distributed energy resources, such as building integrated/mounted wind turbines. The issue to coordinate EV charging with locally generated wind power of buildings can potentially reduce the impacts of EV charging demand on the power grid. As a result, this paper investigates this important problem and three contributions are made. First, the real-time scheduling of EV charging is addressed in a centralized framework based on the ideas of model predictive control, which incorporates the volatile wind power supply of buildings and the random daily driving cycles of EVs among different buildings. Second, an EV-based decentralized charging algorithm (EBDC) is developed to overcome the difficulties due to: 1) the possible lack of global information regarding the charging requirements of all EVs and 2) the computational burden with the increasing number of EVs. Third, we prove that the EBDC method can converge to the optimal solution of the centralized problem over each planning horizon. Moreover, the performance of the EBDC method is assessed through numeric comparisons with an optimal and two heuristic charging strategies (i.e., myopic and greedy). The results demonstrate that the EBDC method can achieve a satisfactory performance in improving the scalability and the balance between the EV charging demand and wind power supply of buildings.Note to Practitioners—This paper is motivated by the challenging problem due to the substantial charging load of electric vehicles (EVs) on the power grid. Nowadays, modern technological advances and the rapid increase of high-rise buildings have promoted the development of building integrated/mounted wind turbines. As the EVs are usually parked in buildings for a large proportion of time every day, the issue to best utilize locally generated wind power of buildings to suffice EV travelling requirements shows vital significance in reducing their dependence on the power grid. However, there exist two main challenges including: 1) the multiple uncertainties regarding the uncertain wind power generation and the random driving behaviors of EVs and 2) the scalability of the solution method. To tackle the first challenge, the idea of model predictive control is introduced to make charging decisions at each stage based on a short-term prediction of the on-site wind power and the current collection of EVs parked there. To consider the scalability and overcome the lack of global charging information of all EVs in practical deployment, an iterative EV-based decentralized charging algorithm (EBDC) is derived, in which each EV can dynamically update its own charging decisions according to a dynamic charging “price” announced by the buildings. Alternatively, the buildings dynamically adjust the charging “price” to motivate the EVs to get charged during the time periods with sufficient wind power supply. Numeric results demonstrate that the EBDC method is scalable and performs well in improving the balance between the EV charging demand and the wind power supply of buildings.
Yu Yang 0008, Qing-Shan Jia, Xiaohong Guan, Xuan Zhang 0004, Zhifeng Qiu, Geert Deconinck
IEEE Trans Autom. Sci. Eng.2
2018 Cyber-physical model for efficient and secured operation of CPES or energy Internet
Xiaohong Guan, Zhanbo Xu, Qing-Shan Jia, Kun Liu 0017
Sci. China Inf. Sci.3
2018 On distributed event-based optimization for shared economy in cyber-physical energy systems
Qing-Shan Jia, Junjie Wu 0004
Sci. China Inf. Sci.1
2018 Guest Editorial: Special Issue on (Industrial) Internet-of-Things for Smart and Sensing Systems: Issues, Trends, and Applications
abstract
Smart and sensing systems represent a significant research theme that has motivated a number of research initiatives around the world. Smart systems typically consist of different components, including sensors for signal acquisition, communication units for data transmission between components, control, and management units for decision-making, and actuators to perform the appropriate actions. They have impacted applications in many fields, such as manufacturing, healthcare, energy, environment, logistic, defence, monitoring, and mobility. As the complexity of such systems continue to grow, the challenge of developing integrated smart and sensing systems has surpassed the design complexity of their individual components. A smart and sensing system may include a large number of heterogeneous components, subsystems, smart and intelligent sensors, and may combine several correlated functionalities. Thus, the main issue of developing smart and sensing systems lies in the complexity to integrate and to manage these different components, technologies, and goals across a wide spectrum. In recent years, the emergence of the Internet of Things (IoT) has amplified the capacity of sensing the world through a network of connected devices using the existing network infrastructure. Grouping together smart and sensing systems in an IoT setting to form large-scale distributed cyber-physical systems has tremendous potential in bringing smart systems to many application domains.
Hervé Panetto, Paulo Cézar Stadzisz, Wenchao Li 0001, Qing-Shan Jia
IEEE Internet Things J.4
2018 Event-Based HVAC Control - A Complexity-Based Approach
abstract
The optimal control of the heating, ventilation, and air-conditioning (HVAC) system in buildings has a significant energy saving potential and therefore is of a great practical interest. An event-based HVAC control adjusts control actions when certain events occur, which may be faster and more scalable than state-based or time-driven control methods. However, events may capture either local or global changes in the rooms. The choice of events is a tradeoff between the computational efficiency and the control performance. This challenging problem remains open. We consider this as an important problem in this paper and make three major contributions. First, we define local and global events for the HVAC control problem. The complexity of these event-based control policies is defined. Second, based on hypothesis testing, we develop a method to select events that capture a sufficient state information and with a relatively small event space. Third, we demonstrate the performance of this method on two groups of examples, including one group of small-scale problems for the proof of concept and the other group of large-scale problems in the HVAC control. It is shown that our method outperforms the Levin search, which is a traditional complexity-based search method and finds event-based HVAC control policies with a good performance.Note to Practitioners—When there are multiple rooms in a building, the HVAC control may achieve a significant energy saving and an indoor comfort satisfaction in the same time through exploring the coupling among the rooms. By appropriately defining the events, the size of the event space is usually much smaller than the state space. Therefore, an event-based control is more scalable and preferred in practice. Local events capture the state changes of rooms in a small neighborhood, which leads to a small event space but limited information. Global events capture the state changes of rooms in a large neighborhood, which leads to more information but a large event space. It remains open how to select events in large-scale HVAC control problems, especially when the computing budget is limited. In this paper, we define the complexity of an event-based control policy by the number of neighboring rooms considered. We develop a method based on hypothesis testing to select events with a proper complexity in order to achieve a good system performance. The performance of this method is demonstrated on an HVAC control problem.
Qing-Shan Jia, Junjie Wu 0004, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.1
2017 A Multi-Timescale and Bilevel Coordination Approach for Matching Uncertain Wind Supply With EV Charging Demand
abstract
The matching between random wind supply and electric vehicle (EV) charging demand can reduce the requirement of traditional power sources and the emission of CO2. This problem is of great practical interest but involves system dynamics in multiple timescales. We consider this an important problem in this paper. In order to capture the randomness in the wind supply and EV charging demand, we formulate the problem as a bilevel Markov decision process. At the upper level, the charging demand of EVs in different locations is aggregated into multiple aggregators. The system operator dispatches power among the aggregators in a coarse timescale to maximize the wind power utilization. At the lower level, the aggregator schedules the charging process of individual EVs at a finer timescale to minimize the charging cost. In order to solve this large-scale problem, a bilevel simulation-based policy improvement (SBPI) method is developed. It is mathematically proved that the SBPI can improve from base policies in both levels. The performance of this multi-timescale and bilevel coordination approach is demonstrated through case studies in the city of Beijing.
Qilong Huang, Qing-Shan Jia, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.2
2015 A Decentralized Stay-Time Based Occupant Distribution Estimation Method for Buildings
abstract
Zonal occupant level is of great practical interest for building energy saving under normal operations and for fast evacuation under emergency. Though there are many existing sensing systems to estimate this information, the problem is still challenging due to the privacy concerns, the random human movement, and the accumulative error. In this paper, we consider this important problem and focus on infrared beam systems that monitor the zonal arrival and departure events. We make the following contributions. First, a rule (i.e., Rule 1) based on the stay time is developed to reduce the accumulated estimation error in each zone. Second, a rule (i.e., Rule 2) is designed to coordinate the estimation among neighboring zones. A decentralized estimation method is then developed using these two rules. Third, the advantage of this method is demonstrated through simulation results and field tests. We hope this work brings insight to zonal occupant level estimation in buildings in more general situations.
Qing-Shan Jia, Hengtao Wang, Yulin Lei, Qianchuan Zhao, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.1
2015 Event-Based Optimization Within the Lagrangian Relaxation Framework for Energy Savings in HVAC Systems
abstract
Optimizing HVAC operation becomes increasingly important because of the rising energy cost and comfort requirements. In this paper, an innovative event-based approach is developed within the Lagrangian relaxation framework to minimize an HVAC's day-ahead energy cost. To solve the HVAC optimization problem based on events is challenging since with time-dependent uncertainties in weather, cooling load, etc., the optimal policy is not stationary. The nonstationary policy space is extremely large, and it is time consuming to find the optimal policy. To overcome the challenge, we develop an event-based approach to make the nonstationary optimal policy stationary in the planning horizon. The key idea is to augment state variables to include the time-dependent variables that make the optimal policy nonstationary and then define events based on the extended state variables. In addition, we develop within the Lagrangian relaxation framework a Q-learning method where Q-factors are used to evaluate event-action pairs and to obtain the optimal policy. Numerical results demonstrate that, as compared with time-based approaches, the event-based approach maintains similar levels of energy costs and human comfort, but reduces computational efforts significantly and has a much faster response to events.
Peter B. Luh, Qing-Shan Jia, Bing Yan 0003
IEEE Trans Autom. Sci. Eng.3
2015 Supply Demand Coordination for Building Energy Saving: Explore the Soft Comfort
abstract
Due to the large amount of energy consumed in buildings, building energy savings has attracted more and more attention recently. The total energy consumed during building operations is determined by the building energy efficiency and the total demand. On the one hand, though most existing studies focus on improving building energy efficiency, there are limits. On the other hand, the demand grows fast and without limit. Therefore it is important to coordinate the supply and demand in buildings. We consider this important problem in this paper and make the following major contributions. First, the concept of average price of electricity (APE) is defined to measure the average generation cost of electricity using multiple devices. Second, a comfort model of occupant is developed to capture the tradeoff between thermal comfort and cost. Human building interaction allows the user to adjust their temperature set ranges according to the APE in real time. Third, an iterative solution method is developed to solve the supply demand coordination optimization problem. Numerical examples show that significant energy saving is possible through exploring the soft comfort requirement of the occupants, and the iterative method achieves a solution which is close to that of the centralized method, but in a much faster way. We hope this work brings insight to building energy saving in general.
Zhanbo Xu, Qing-Shan Jia, Xiaohong Guan
IEEE Trans Autom. Sci. Eng.2
2014 Indoor Occupant Positioning System Using Active RFID Deployment and Particle Filters
abstract
This article describes a method for indoor positioning of human-carried active Radio Frequency Identification (RFID) tags based on the Sampling Importance Resampling (SIR) particle filtering algorithm. To use particle filtering methods, it is necessary to furnish statistical state transition and observation distributions. The state transition distribution is obstacle-aware and sampled from a precomputed accessibility map. The observation distribution is empirically determined by ground truth RSS measurements while moving the RFID tags along a known trajectory. From this data, we generate estimates of the sensor measurement distributions, grouped by distance, between the tag and sensor. A grid of 24 sensors is deployed in an office environment, measuring Received Signal Strength (RSS) from the tags, and a multithreaded program is written to implement the method. We discuss the accuracy of the method using a verification data set collected during a field-operational test.
Kevin Weekly, Han Zou, Lihua Xie 0001, Qing-Shan Jia, Alexandre M. Bayen
DCOSS4
2014 Sample path sharing in simulation-based policy improvement
abstract
Simulation-based policy improvement (SBPI) has been widely used to improve given base policies through simulation. The basic idea of SBPI is to estimate all the Q-factors for a given state using simulation, and then select the action that achieves the minimal cost. It is therefore of great importance to efficiently use the given budget in order to select the best action with high probability. Different from existing budget allocation algorithms that estimate Q-factors by independent simulation, we share the sample paths to improve the probability of correctly selecting the best action. Our method can be combined with equal allocation, Successive Rejects, and optimal computing budget allocation to enhance their probabilities of correct selection as well as to achieve better policies in SBPI. Such improvement depends on the overlap in reachable states under different actions. Numerical results show that with such overlap, combining our method with equal allocation, Successive Rejects and optimal computing budget allocation produces higher probability of selection as well as better policies in SBPI.
Qing-Shan Jia, Chun-Hung Chen
ICRA2
2014 Building occupant level estimation based on heterogeneous information fusion
Hengtao Wang, Qing-Shan Jia, Ruixi Yuan, Xiaohong Guan
Inf. Sci.2
2014 Building Energy Doctors: An SPC and Kalman Filter-Based Method for System-Level Fault Detection in HVAC Systems
abstract
Buildings worldwide account for nearly 40% of global energy consumption. The biggest energy consumer in buildings is the Heating, Ventilation and Air Conditioning (HVAC) systems. HVAC also ranks top in terms of number of complaints by tenants. Maintaining HVAC systems in good conditions through early fault detection is thus a critical problem. The problem, however, is difficult since HVAC systems are large in scale, consisting of many coupling subsystems, building and equipment dependent, and working under time-varying conditions. In this paper, a model-based and data-driven method is presented for robust system-level fault detection with potential for large-scale implementation. It is a synergistic integration of: ) Statistical Process Control (SPC) for measuring and analyzing variations; 2) Kalman filtering based on gray-box models to provide predictions and to determine SPC control limits; and (3) system analysis for analyzing propagation of faults' effects across subsystems. In the method, two new SPC rules are developed for detecting sudden and gradual faults. The method has been tested against a simulation model of the HVAC system for a 420-meter-high building. It detects both sudden faults and gradual degradation, and both device and sensor faults. Furthermore, the method is simple and generic, and has potential replicability and scalability.
Peter B. Luh, Qing-Shan Jia, Zheng O'Neill, Fangting Song
IEEE Trans Autom. Sci. Eng.3
2013 An integrative Weighted Path Loss and Extreme Learning Machine approach to Rfid based Indoor Positioning
abstract
In recent years, applying RFID technology to develop an Indoor Positioning System (IPS) has become a hot research topic. The most prominent advantage of active RFID IPS comes from its unique identification of different objects in indoor environment. However, certain drawbacks of existing RFID IPSs, such as high cost of RFID readers and active tags, as well as heavy dependence on the density of reference tags to provide the location based service, largely limit the applications of active RFID IPS. In order to overcome these drawbacks, we develop a cost-efficient RFID IPS by using cheaper active RFID tags, sensors and reader. In addition, one localization algorithm: integrated Weighted Path Loss (WPL) - Extreme Learning Machine (ELM) which combines the fast estimation of WPL and the high localization accuracy of ELM is proposed. According to the algorithm, an indoor environment is divided into small zones firstly and an ELM model is developed for each zone during the offline phase. During the online phase, the WPL approach is used to determine the zone of the target primarily, then the ELM model of that zone is deployed to provide the final estimated location of the target. Based on our experimental result, this integrated algorithm provides a higher localization efficiency and accuracy than existing approaches.
Han Zou, Lihua Xie 0001, Qing-Shan Jia, Hengtao Wang
IPIN3
2013 Building Energy Management: Integrated Control of Active and Passive Heating, Cooling, Lighting, Shading, and Ventilation Systems
abstract
Buildings account for nearly 40% of global energy consumption. About 40% and 15% of that are consumed, respectively, by HVAC and lighting. These energy uses can be reduced by integrated control of active and passive sources of heating, cooling, lighting, shading and ventilation. However, rigorous studies of such control strategies are lacking since computationally tractable models are not available. In this paper, a novel formulation capturing key interactions of the above building functions is established to minimize the total daily energy cost. To obtain effective integrated strategies in a timely manner, a methodology that combines stochastic dynamic programming (DP) and the rollout technique is developed within the price-based coordination framework. For easy implementation, DP-derived heuristic rules are developed to coordinate shading blinds and natural ventilation, with simplified optimization strategies for HVAC and lighting systems. Numerical simulation results show that these strategies are scalable, and can effectively reduce energy costs and improve human comfort.
Peter B. Luh, Qing-Shan Jia, Ziyan Jiang
IEEE Trans Autom. Sci. Eng.3
2013 Smart Management of Multiple Energy Systems in Automotive Painting Shop
abstract
Automotive painting shops consume electricity and natural gas to provide the required temperature and humidity for painting processes. The painting shop is not only responsible for a significant portion of energy consumption with automobile manufacturers, but also affects the quality of the product. Various storage devices play a crucial role in the management of multiple energy systems. It is thus of great practical interest to manage the storage devices together with other energy systems to provide the required environment with minimal cost. In this paper, we formulate the scheduling problem of these multiple energy systems as a Markov decision process (MDP) and then provide two approximate solution methods. Method 1 is dynamic programming with value function approximation. Method 2 is mixed integer programming with mean value approximation. The performance of the two methods is demonstrated on numerical examples. The results show that method 2 provides good solutions fast and with little performance degradation comparing with method 1. Then, we apply method 2 to optimize the capacity and to select the combination of the storage devices, and demonstrate the performance by numerical examples.
Zhanbo Xu, Qing-Shan Jia, Xiaohong Guan, Jian-Xiang Shen
IEEE Trans Autom. Sci. Eng.2
2013 Reconfiguring Networked Infrastructures by Adding Wireless Communication Capabilities to Selected Nodes
abstract
Robustness is an important requirement on many critical networked infrastructures. One of the most important ways to achieve robustness is to keep an appropriate level of redundancy by adding redundant resources for the critical networked infrastructures. Considering wireless with good reconfigurabilities, we add wireless communication capacities to selected nodes of a wired networked system in this paper, where we firstly focus on adding minimum wireless communication capacities to achieve biconnectivity requirements, secondly optimize the wireless capacities installation that maximizes the efficiency of reconfiguration, and thirdly propose an approach of topology reconfiguration with wireless communication to keep system connectivity if there are node failures. To resolve the computational complexity issue in searching for the optimal node combination to add wireless capacities, we propose the concept of the equivalent network that reduces the system topology as a simplified acyclic network. The topology optimization is under a metric of reconfiguration distance which guarantees the maximum of the shortest distances between any one node and that in the set of reconfigurable nodes to be minimal such that the efficiency of topology reconfiguration with node failure is the highest. The performances of the methods presented in the paper are demonstrated through numerical experiments and testing.
Hengtao Wang, Qianchuan Zhao, Xiaohong Guan, Qing-Shan Jia
IEEE Trans. Wirel. Commun.4
2012 Event-based sensor activation for indoor occupant distribution estimation
abstract
The information of the distribution of occupant in an indoor environment is important for building energy saving under normal conditions and for evacuation under emergent conditions, and thus is of great practical interest. Due to low set-up cost, wireless sensor networks powered by batteries are usually used for such estimation. The question is how to activate the sensors to minimize the estimation error within a given period of time. In this paper we develop an event-based activation policy which can be easily implemented in a decentralized way. We use numerical experiments to compare this policy with four other policies and to demonstrate the impact of various factors on the performance of these policies, such as the topology, sensor accuracy, battery capacity, occupant movement model, and occupant population. Our method outperforms the other policies in all the tested scenarios.
Qing-Shan Jia, Zhe Wen
ICARCV1
2012 A systematic method for network topology reconfiguration with limited link additions
Qing-Shan Jia, Hengtao Wang, Ruixi Yuan, Xiaohong Guan
J. Netw. Comput. Appl.2
2012 Efficient Computing Budget Allocation for Simulation-Based Policy Improvement
abstract
Policy improvement in discrete event dynamic systems is usually based on simulation, which is time-consuming and provides only noisy performance evaluation. For a given system state, it is of great practical interest to understand how to allocate the computing budget among action candidates so that the best action is correctly selected with high probability. Despite the abundant studies on simulation-based policy optimization, few consider this important allocation problem, which is considered in this paper. We develop the method of optimal computing budget allocation for policy improvement (OCBAPI) which is shown to asymptotically maximize a lower bound of the probability of correctly selecting the best action. OCBAPI can also be used when there are multiple base policies available. This allocation procedure is compared with equal allocation and proportional-to-variance on an academic toy example and an engine maintenance policy optimization problem. The numerical results show that even when there are only finite computing budget to allocate, OCBAPI performs well. We hope this work brings insight to computing budget allocation for simulation-based policy improvement in more general situations.
Qing-Shan Jia
IEEE Trans Autom. Sci. Eng.1
2011 Ordinal Optimization-Based Multi-energy System Scheduling for Building Energy Saving
Zhong-Hua Su, Qing-Shan Jia
ICIC (2)2
2011 Tracking a moving object via a sensor network with a partial information broadcasting scheme
Jianghai Li, Qing-Shan Jia, Xiaohong Guan, Xi Chen 0036
Inf. Sci.2
2011 An Adaptive Sampling Algorithm for Simulation-Based Optimization With Descriptive Complexity Preference
abstract
Many systems nowadays follow not only physical laws but also manmade rules. These systems are known as discrete-event dynamic systems (DEDSs), where simulation is the only faithful way for performance evaluation. Due to various advantages in practice, designs (or solution candidates) with low descriptive complexity (called simple designs) are usually preferred over complex ones when their performances are close. However, the descriptive complexity (DC) is usually nonlinear and takes discrete value, which makes traditional methods such as linear programming and gradient-based local search not applicable. Existing methods for simulation-based optimization (SBO) do not explore the preference on descriptive complexity and thus cannot solve the problem efficiently. The major contributions of this paper are to point out the importance of considering SBO problems with DC preference, and to develop an adaptive sampling algorithm (ASA) to find the simplest good design. It is shown that ASA terminates within finite iterations and with controllable probability of making mistake. The computational complexity of ASA and its dependence on various parameters are discussed. ASA is then applied to three parameter optimization problems and a node activation policy optimization problem in a wireless sensor network. Numerical results show that ASA is more efficient than blind picking and Levin search in most cases. We hope this work can shed some insight to how to find simple and good designs in general.
Qing-Shan Jia
IEEE Trans Autom. Sci. Eng.1
2010 Prediction-Based Activation in WSN Tracking
abstract
In Wireless Sensor Network (WSN) based tracking, the random factors in the target's movement and the sensors' measurement together with the limited sensing ability will cause target loss. Therefore, target loss probability is important for the WSN design and performance evaluation. In this paper, a theoretical analysis of the target loss probability with other key parameters is carried out. Based on the analysis, an adaptive-radius activation policy is developed to reduce target loss. Numerical examples demonstrate adaptive-radius activation policy is effective.
Jianghai Li, Xi Chen 0036, Qing-Shan Jia, Xiaohong Guan
MSN3
2010 A Structural Property of Optimal Policies for Multi-Component Maintenance Problems
abstract
In this paper, we focus on the opportunistic maintenance of an asset which is composed of multiple nonidentical life-limited components with both economic and structural dependence. Besides the random asset failure, each component also has independent failure with constant rate. Both finite and infinite horizons are considered. We first prove the optimality of the Generalized Strict Shortest-Remaining-Lifetime-First (GSSRLF) rule to efficiently reduce the size of the action space from O(2n) to O(IIi=1m= ni), where n is the number of components, m is the number of modules, and n_i is the number of components in module i . Then, we show that the GSSRLF rule is more general than the existing SRLF rule, and has close relationship with several other rules and properties known in literature. Finally, we discuss the limitations of the GSSRLF rule and use numerical results to show that even when the rule is not optimal, it helps to identify good policies.
Qing-Shan Jia
IEEE Trans Autom. Sci. Eng.1
2010 Long-Term Scheduling for Cascaded Hydro Energy Systems With Annual Water Consumption and Release Constraints
abstract
Long-term scheduling for cascaded hydro energy systems is very important for low carbon energy production. It aims at determining the water release over a planning horizon to meet water resource requirements and the system demands for electric power. The problem is challenging in view of the complicated and stochastic system dynamics, nonlinear marginal cost, coupled hydraulic constraints, and the large problem size. In this paper we formulate the long-term scheduling problem of cascaded hydro energy systems with annual consumption and release constraints as a finite horizon constrained Markov decision process (CMDP), and develop a new rollout algorithm to optimize the policies. Numerical results demonstrate the effectiveness and the efficiency of the formulation and the new algorithm.
Yanjia Zhao, Xi Chen 0036, Qing-Shan Jia, Xiaohong Guan, Shuanghu Zhang, Yunzhong Jiang
IEEE Trans Autom. Sci. Eng.3
2009 Strategy optimization for controlled Markov process with descriptive complexity constraint
Qing-Shan Jia, Qianchuan Zhao
Sci. China Ser. F Inf. Sci.1
2008 A Structure Property of Optimal Policies for Maintenance Problems WithSafety-Critical Components
abstract
The maintenance problem with safety-critical components is significant for the economical benefit of companies. Motivated by a practical asset maintenance project, a new joint replacement maintenance problem is introduced in this paper. The dynamics of the problem are modelled as a Markov decision process, whose action space increases exponentially with the number of safety-critical components in the asset. To deal with the curse of dimensionality, we identify a key property of the optimal solution: the optimal performance can always be achieved in a class of policies which satisfy the so-called shortest-remaining-lifetime-first (SRLF) rule. It reduces the action space from 0(2n) to O(n), where n is the number of safety-critical components. To further speed up the optimization procedure, some interesting properties of the optimal policy are derived. Combining the SRLF rule and the neuro-dynamic programming (NDP) methodology, we develop an efficient on-line algorithm to optimize this maintenance problem. This algorithm can handle the difficulties of large state space and large action space. Besides the theoretical proof, the optimality and efficiency of the SRLF rule and the properties of the optimal policy are also illustrated by numerical examples. This work can shed some insights to the maintenance problems in a more general situation.
Qianchuan Zhao, Qing-Shan Jia
IEEE Trans Autom. Sci. Eng.3
2006 A SVM-based Method for Engine Maintenance Strategy Optimization
abstract
Due to the abundant application background, the optimization of maintenance problem has been extensively studied in the past decades. Besides the well-known difficulty of large state space and large action space, the pervasive application of digital computers forces us to consider the new constraint of limited memory space. The given memory space restricts what strategies can be explored during the optimization procedure. By explicitly quantifying the minimal memory space to store a strategy using support vector machine, we propose to describe simple strategies exactly and only approximate complex strategies. This selective approximation can best utilize the given memory space for any description mechanism. We use numerical results on illustrative examples to show how the selective approximation improves the solution quality. We hope this work sheds some insights to best utilize the memory space for practical engine maintenance strategy optimization problems
Qing-Shan Jia, Qianchuan Zhao
ICRA1