VLDB 2026 Research / reviewers in the wild / expert
Xianzhong Ding
dblp:198/7795
· DBLP profile ↗
12ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0001-6114-2801ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data CentersabstractManaging a vast number of virtual machines (VMs) efficiently is a critical challenge in modern large-scale data centers. The continuous creation and termination of VMs lead to resource fragmentation across physical machines (PMs), necessitating periodic VM rescheduling to optimize resource utilization. Despite its significance, VM rescheduling has received limited attention in the literature. A key challenge is that, unlike conventional combinatorial optimization problems, the efficiency of rescheduling algorithms is heavily impacted by inference time, as VM states evolve dynamically during execution. This scalability bottleneck hampers existing methods. To address this, we propose VMR$^{2}$L, a reinforcement learning framework tailored for VM rescheduling. VMR$^{2}$L integrates a two-stage decision-making process to accommodate complex operational constraints, a feature extraction mechanism that captures critical relational information for rescheduling, and a risk-aware evaluation strategy that enables users to balance execution speed and rescheduling accuracy. Extensive experiments using real-world data from a production-scale data center demonstrate that VMR$^{2}$L achieves near-optimal performance while reducing inference time to a matter of seconds. To facilitate reproducibility, we provide access to our implementation and datasets. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement LearningabstractModern industry-scale data centers need to manage a large number of virtual machines (VMs). Due to the continual creation and release of VMs, many small resource fragments are scattered across physical machines (PMs). To handle these fragments, data centers periodically reschedule some VMs to alternative PMs, a practice commonly referred to as VM rescheduling. Despite the increasing importance of VM rescheduling as data centers grow in size, the problem remains understudied. We first show that, unlike most combinatorial optimization tasks, the inference time of VM rescheduling algorithms significantly influences their performance, due to dynamic VM state changes during this period. This causes existing methods to scale poorly. Therefore, we develop a reinforcement learning system for VM rescheduling, VMR2L, which incorporates a set of customized techniques, such as a two-stage framework that accommodates diverse constraints and workload conditions, a feature extraction module that captures relational information specific to rescheduling, as well as a risk-seeking evaluation enabling users to optimize the trade-off between latency and accuracy. We conduct extensive experiments with data from an industry-scale data center. Our results show that VMR2L can achieve a performance comparable to the optimal solution but with a running time of seconds. Code12 and datasets3 are open-sourced. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
EuroSys | 1 |
| 2025 | A Safe and Data-Efficient Model-Based Reinforcement Learning System for HVAC ControlabstractModel-based reinforcement learning (MBRL) is widely studied for heating, ventilation, and air conditioning (HVAC) control in buildings. One of the critical challenges is the large amount of data required to effectively train neural networks for modeling building dynamics. This article presents CLUE, an MBRL system for HVAC control in buildings. CLUE optimizes HVAC operations by integrating a Gaussian process (GP) model to model building dynamics with uncertainty awareness. CLUE utilizes GP to predict state transitions as Gaussian distributions, effectively capturing prediction uncertainty and enhancing decision-making under sparse data conditions. Our approach employs a meta-kernel learning technique to efficiently set GP kernel hyperparameters using domain knowledge from diverse buildings. This drastically reduces the data requirements typically associated with GP models in HVAC applications. Additionally, CLUE incorporates these uncertainty estimates into a model predictive path integral (MPPI) algorithm, enabling the selection of safe, energy-efficient control actions. This uncertainty-aware control strategy evaluates and selects action trajectories based on their predicted impact on energy consumption and human comfort, optimizing operations even under uncertain conditions. Extensive simulations in a five-zone office building demonstrate that CLUE reduces the required training data from hundreds of days to just seven while maintaining robust control performance. It reduces comfort violations by an average of 12.07% compared to existing MBRL methods, without compromising on energy efficiency. Our code and dataset are available athttps://github.com/ryeii/CLUE. Xianzhong Ding, Zhiyu An, Arya Rathee, Wan Du |
IEEE Internet Things J. | 1 |
| 2025 | Multi-Zone HVAC Control With Model-Based Deep Reinforcement LearningabstractThe application of reinforcement learning in controlling Heating, Ventilation, and Air Conditioning (HVAC) systems has been extensively researched. Existing studies primarily focus on Model-Free Reinforcement Learning (MFRL), which involves trial-and-error interactions with real buildings to train the agent. However, MFRL encounters a significant challenge: it requires a large amount of training data to achieve satisfactory performance. While simulation models have been used to generate training data and expedite the training process, they necessitate high-fidelity building models that are difficult to calibrate. As a result, Model-Based Reinforcement Learning (MBRL) has been employed for HVAC control. Although MBRL demonstrates remarkable sample efficiency, it often falls short in terms of asymptotic control performance, particularly in achieving substantial energy savings while ensuring occupants’ thermal comfort. In this study, we conduct experiments to analyze the limitations of current MBRL-based HVAC control methods, focusing on model uncertainty and controller effectiveness. Leveraging the insights gained from these experiments, we develop MB2C, an innovative MBRL-based HVAC control system that combines high control performance with exceptional sample efficiency. MB2C learns the dynamics of the building by employing an ensemble of environment-conditioned neural networks and utilizes a novel control method called Model Predictive Path Integral (MPPI) for HVAC control. MPPI generates candidate action sequences using an importance sampling weighted algorithm, which is well-suited for multi-zone buildings with high state and action dimensions. We evaluate MB2C using EnergyPlus simulations in a five-zone office building, and the results demonstrate that MB2C achieves 8.23% higher energy savings compared to the state-of-the-art MBRL solution while maintaining comparable thermal comfort. Moreover, MB2C significantly reduces the required training data set by an order of magnitude ($10.52\times $) while delivering performance on par with MFRL approaches. Note to Practitioners—Our research addresses a critical challenge in HVAC control, offering an innovative solution to enhance the data efficiency of HVAC systems while optimizing energy usage. Traditional approaches, such as Model-Free Reinforcement Learning, often require a large volume of real-world data. Our primary focus is improving the effectiveness of HVAC control, a vital aspect of building management that directly affects energy consumption and occupant well-being. We introduce MB2C, a Model-Based Reinforcement Learning system designed to significantly improve energy savings while maintaining thermal comfort. MB2C achieves remarkable results, offering exceptional sample efficiency and substantially reducing the required training data. Our research leverages an ensemble of environment-conditioned neural networks and employs Model Predictive Path Integral in HVAC control. While MB2C presents notable benefits, it also has limitations. Further research and development are required to optimize its performance across different building environments and specific use cases. Future directions should focus on addressing the safety challenges associated with real-world deployment. Beyond HVAC control, the principles and methods explored in this research have potential applications in various automation domains, such as robotics, industrial automation, and manufacturing processes. Xianzhong Ding, Alberto Cerpa, Wan Du |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Go Beyond Black-box Policies: Rethinking the Design of Learning Agent for Interpretable and Verifiable HVAC ControlabstractRecent research has shown the potential of Model-based Reinforcement Learning (MBRL) to enhance energy efficiency of Heating, Ventilation, and Air Conditioning (HVAC) systems. However, existing methods rely on black-box thermal dynamics models and stochastic optimizers, lacking reliability guarantees and posing risks to occupant health. In this work, we overcome the reliability bottleneck by redesigning HVAC controllers using decision trees extracted from existing thermal dynamics models and historical data. Our decision tree-based policies are deterministic, verifiable, interpretable, and more energy-efficient than current MBRL methods. First, we introduce a novel verification criterion for RL agents in HVAC control based on domain knowledge. Second, we develop a policy extraction procedure that produces a verifiable decision tree policy. We found that the high dimensionality of the thermal dynamics model input hinders the efficiency of policy extraction. To tackle the dimensionality challenge, we leverage importance sampling conditioned on historical data distributions, significantly improving policy extraction efficiency. Lastly, we present an offline verification algorithm that guarantees the reliability of a control policy. Extensive experiments show that our method saves 68.4% more energy and increases human comfort gain by 14.8% compared to the state-of-the-art method, in addition to an 1127× reduction in computation overhead. Our code and data are available at https://github.com/ryeii/Veri_HVAC. Zhiyu An, Xianzhong Ding, Wan Du |
DAC | 2 |
| 2024 | Exploring Deep Reinforcement Learning for Holistic Smart Building ControlabstractIn recent years, the focus has been on enhancing user comfort in commercial buildings while cutting energy costs. Efforts have mainly centered on improving HVAC systems, the central control system. However, it’s evident that HVAC alone can’t ensure occupant comfort. Lighting, blinds, and windows, often overlooked, also impact energy use and comfort. This paper introduces a holistic approach to managing the delicate balance between energy efficiency and occupant comfort in commercial buildings. We present OCTOPUS , a system employing a deep reinforcement learning (DRL) framework using data-driven techniques to optimize control sequences for all building subsystems, including HVAC, lighting, blinds, and windows. OCTOPUS ’s DRL architecture features a unique reward function facilitating the exploration of tradeoffs between energy usage and user comfort, effectively addressing the high-dimensional control problem resulting from interactions among these four building subsystems. To meet data training requirements, we emphasize the importance of calibrated simulations that closely replicate target-building operational conditions. We train OCTOPUS using 10-year weather data and a calibrated building model in the EnergyPlus simulator. Extensive simulations demonstrate that OCTOPUS achieves substantial energy savings, outperforming state-of-the-art rule-based and DRL-based methods by 14.26% and 8.1%, respectively, in a LEED Gold Certified building while maintaining desired human comfort levels. Xianzhong Ding, Alberto Cerpa, Wan Du |
ACM Trans. Sens. Networks | 1 |
| 2024 | Optimizing Irrigation Efficiency using Deep Reinforcement Learning in the FieldabstractAgricultural irrigation is a significant contributor to freshwater consumption. However, the current irrigation systems used in the field are not efficient. They rely mainly on soil moisture sensors and the experience of growers but do not account for future soil moisture loss. Predicting soil moisture loss is challenging because it is influenced by numerous factors, including soil texture, weather conditions, and plant characteristics. This article proposes a solution to improve irrigation efficiency, which is called DRLIC (deep reinforcement learning for irrigation control). DRLIC is a sophisticated irrigation system that uses deep reinforcement learning (DRL) to optimize its performance. The system employs a neural network, known as the DRL control agent, which learns an optimal control policy that considers both the current soil moisture measurement and the future soil moisture loss. We introduce an irrigation reward function that enables our control agent to learn from previous experiences. However, there may be instances in which the output of our DRL control agent is unsafe, such as irrigating too much or too little. To avoid damaging the health of the plants, we implement a safety mechanism that employs a soil moisture predictor to estimate the performance of each action. If the predicted outcome is deemed unsafe, we perform a relatively conservative action instead. To demonstrate the real-world application of our approach, we develop an irrigation system that comprises sprinklers, sensing and control nodes, and a wireless network. We evaluate the performance of DRLIC by deploying it in a testbed consisting of six almond trees. During a 15-day in-field experiment, we compare the water consumption of DRLIC with a widely used irrigation scheme. Our results indicate that DRLIC outperforms the traditional irrigation method by achieving water savings of up to 9.52%. Xianzhong Ding, Wan Du |
ACM Trans. Sens. Networks | 1 |
| 2023 | Poster Abstract: Data Efficient HVAC Control using Gaussian Process-based Reinforcement LearningabstractModel-based Reinforcement Learning (MBRL) has been widely studied for energy-efficient control of the Heating, Ventilation, and Air Conditioning (HVAC) systems. One of the fundamental issues of the current approaches is the large amount of data required to train an accurate building system dynamics model. In this work, we developed a data-efficient system capable of excellent HVAC control performance with only days of training data. We use a Gaussian Process (GP) as the dynamics model which provides uncertainty for each prediction. To improve the data efficiency, we designed a meta kernel learning technique for GP kernel selection. To incorporate uncertainty in the control decisions, we designed a model predictive control method that considers the uncertainty of every prediction. Simulation experiments show that our method achieves excellent data efficiency, yielding similar energy savings and 12.07% less human comfort violation compared with the state-of-the-art MBRL method, while only trained on a seven-day training dataset. Zhiyu An, Xianzhong Ding, Wan Du |
SenSys | 2 |
| 2022 | DRLIC: Deep Reinforcement Learning for Irrigation ControlabstractAgricultural irrigation is a major consumer of freshwater. Current irrigation systems used in the field are not efficient, since they are mainly based on soil moisture sensors' measurement and growers' experience, but not future soil moisture loss. It is hard to predict soil moisture loss, as it depends on a variety of factors, such as soil texture, weather and plants' characteristics. To improve irrigation efficiency, this paper presents DRLIC, a deep reinforcement learning (DRL)-based irrigation system. DRLIC uses a neural network (DRL control agent) to learn an optimal control policy that takes both current soil moisture measurement and future soil moisture loss into account. We define an irrigation reward function that facilitates the control agent to learn from past experience. Sometimes, our DRL control agent may output an unsafe action (e.g., irrigating too much water or too little). To prevent any possible damage to plants' health, we adopt a safe mechanism that leverages a soil moisture predictor to estimate each action's performance. If it is unsafe, we will perform a relatively-conservative action instead. Finally, we develop a real-world irrigation system that is composed of sprinklers, sensing and control nodes, and a wireless network. We deploy DRLIC in our testbed composed of six almond trees. Through a IS-day in-field experiment, we find that DRLIC can save up to 9.52% of water over a widely-used irrigation scheme. Xianzhong Ding, Wan Du |
IPSN | 1 |
| 2022 | Poster Abstract: Smart Irrigation Control Using Deep Reinforcement LearningabstractAgriculture is a major user of ground and surface water in the United States. Improving the efficiency of agricultural irrigation systems is critical for sustainable agriculture. We propose an IoT-based irrigation system, which includes two major components, i.e., an IoT wireless network of sensing and actuation nodes, and a DRL-based control algorithm. Given the collected soil moisture data and weather data, the DRL-based algorithm finds an optimal irrigation schedule, which uses the minimum amount of water to guarantee the soil-water content above the required level before the next irrigation cycle. We deploy the system in our testbed composed of six almond trees. Through a 12-day in-field experiment, we find that our proposed system can save up to 7.8% of water over a widely-used irrigation scheme. Xianzhong Ding, Wan Du |
IPSN | 1 |
| 2020 | Continuous, Real-Time Object Detection on Mobile Devices without OffloadingabstractThis paper presents AdaVP, a continuous and real-time video processing system for mobile devices without offloading. AdaVP uses Deep Neural Network (DNN) based tools like YOLOv3 for object detection. Since DNN computation is time-consuming, multiple frames may be captured by the camera during the processing of one frame. To support real-time video processing, we develop a mobile parallel detection and tracking (MPDT) pipeline that executes object detection and tracking in parallel. When the object detector is processing a new frame, a light-weight object tracker is used to track the objects in the accumulated frames. As the tracking accuracy decreases gradually, due to the accumulation of tracking error and the appearance of new objects, new object detection results are used to calibrate the tracking accuracy periodically. In addition, a large DNN model produces high accuracy, but requires long processing latency, resulting in a great degradation for tracking accuracy. Based on our experiments, we find that the tracking accuracy degradation is also related to the variation of video content, e.g., for a dynamically changing video, the tracking accuracy degrades fast. A model adaptation algorithm is thus developed to adapt the DNN models according to the change rate of video content. We implement AdaVP on Jetson TX2 and conduct a variety of experiments on a large video dataset. The experiment results reveal that AdaVP improves the accuracy of the state-of-the-art solution by up to 43.9%. Xianzhong Ding, Wan Du |
ICDCS | 2 |
| 2017 | Unified nvTCAM and sTCAM architecture for improving packet matching performanceabstractSoftware-Defined Networking (SDN) allows controlling applications to install fine-grained forwarding policies in the underlying switches. Ternary Content Addressable Memory (TCAM) enables fast lookups in hardware switches with flexible wildcard rule patterns. However, the performance of packet processing is severely constrained by the capacity of TCAM, which aggravates the processing burden and latency issues. In this paper, we propose a hybrid TCAM architecture which consists of NVM-based TCAM (nvTCAM) and SRAM-based TCAM (sTCAM), utilizing nvTCAM to cache the most popular rules to improve cache-hit-ratio while relying on a very small-size sTCAM to handle cache-miss traffic to effectively decrease update latency. Considering the special rule dependency, we present an efficient Rule Migration Replacement (RMR) policy to make full utilization of both nvTCAM and sTCAM to obtain better performance. Experimental results show that the proposed architecture outperforms current TCAM architectures. Xianzhong Ding, Zhiyong Zhang 0006, Zhiping Jia, Lei Ju 0001, Mengying Zhao, Huawei Huang |
LCTES | 1 |