Xin Zhou 0003

dblp:05/3403-3 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0001-5405-9890ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 since 2021Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Polarization information restoration for visual reflection removal via cross dual-stream network
Lijun Deng, Hedong Liu, Zixian Liu, Zhen Yang 0012, Xin Zhou 0003, Haofeng Hu
Knowl. Based Syst.6
2026 OACI: Object-aware contextual integration for image captioning
Shuhan Xu, Mengya Han, Wei Yu 0004, Zheng He 0001, Xin Zhou 0003, Yong Luo 0002
Knowl. Based Syst.5
2026 ChatDC: Geometric-aware Data Center Digital Twin Generation via Large Language Models
abstract
A modern data center is supported by an Internet of Assets (IoA), a specialized Internet of Things (IoT), where the sim-ready assets encompass physical properties and connected data streams for passive data collection and proactive simulation. A digital twin (DT) is a virtual replica of the IoA system with a proper level of abstraction, integrating asset-level dynamics models to simulate system-wide behavior, prototype designs, and conduct what-if analysis. However, creating high-fidelity DTs is hindered by the manual effort needed to encode complex geometric layouts and domain-specific design constraints. While large language models (LLMs) offer automation potential, existing methods struggle to generate geometrically plausible and functionally valid DT scenes due to limited domain integration. To bridge the gaps, we propose ChatDC, a conversational system that leverages LLMs to automate data center digital twin generation through a S egment- G enerate- O ptimize (SGO) workflow. ChatDC integrates domain knowledge via a dedicated code library named DCBuild and employs SGO to decompose user prompts, generate initial structures, and optimize layouts in compliance with data center design constraints. Evaluation shows that ChatDC outperforms other baselines with 98% success rate on scratch generation tasks and reduces the average makespan by 10×. Ablation study reveals that the SGO design increases the generation success rate by 65% at most. Furthermore, computational fluid dynamics simulations validate the physical plausibility, confirming the readiness of generated DTs for real-world analysis.
Minghao Li 0005, Ruihang Wang, Xin Zhou 0003, Zhaomeng Zhu, Yonggang Wen 0001, Rui Tan 0001, Huiwen Zheng, Stuart Kennedy
ACM Trans. Internet Things3
2026 D3T: Dual-Timescale Optimization of Task Scheduling and Thermal Management for Energy Efficient Geo-Distributed Data Centers
abstract
The surge of artificial intelligence (AI) has intensified compute-intensive tasks, sharply increasing the need for energy-efficient management in geo-distributed data centers. Existing approaches struggle to coordinate task scheduling and cooling control due to mismatched time constants, stochastic Information Technology (IT) workloads, variable renewable energy, and fluctuating electricity prices. To address these challenges, we propose D3T, a dual-timescale deep reinforcement learning (DRL) framework that jointly optimizes task scheduling and thermal management for energy-efficient geo-distributed data centers. At the fast timescale, D3T employs Deep Q-Network (DQN) to schedule tasks, reducing operational expenditure (OPEX) and task sojourn time. At the slow timescale, a QMIX-based multi-agent DRL method regulates cooling across distributed data centers by dynamically adjusting airflow rates, thereby preventing hotspots and reducing energy waste. Extensive experiments were conducted using TRNSYS with real-world traces, and the results demonstrate that, compared to baseline algorithms, D3T reduces OPEX by 13% in IT subsystems and 29% in cooling subsystems, improves power usage effectiveness (PUE) by 7%, and maintains more stable thermal safety across geo-distributed data centers.
Yongyi Ran, Tongyao Sun, Xin Zhou 0003, Jiangtao Luo, Shuangwu Chen
IEEE Trans. Parallel Distributed Syst.4
2025 EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates
abstract
Active Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Additionally, Ego4D poses a novel task: detecting the speaking activities of the camera wearer, who never appears in the field of view. Existing methods treat these two tasks separately, and treat candidates out of sight as noise. In contrast, we propose EgoNet, a framework that uniformly models all candidates, including those not visible. By capturing interactions among all candidates and modeling broader temporal context, EgoNet reduces uncertainty and improves performance in egocentric active speaker detection.
Yongqian Li, Xin Zhou 0003, Zheng He 0001, Wei Yu 0004, Yong Luo 0002
ICASSP2
2025 CLEST-IQA: Contrastive Learning-Enhanced Swin Transformer for Image Quality Assessment
Yongqian Li, Dixiao Tao, Yong Luo 0002, Xin Zhou 0003, Dehua Cao
ICIG (3)5
2025 Robust Active Speaker Detection in Challenging Environments Using GNN-Fused Multi-modal Cues and Body Language
Yongqian Li, Yong Luo 0002, Xin Zhou 0003
MMM (3)3
2025 Adaptive Capacity Provisioning for Carbon-Aware Data Centers: A Digital Twin-Based Approach
abstract
This paper considers the carbon-aware data center (DC) capacity provisioning problem under uncertain green energy availability and computing demand. To address it, accurate carbon emissions estimation and robust capacity provisioning are necessary. Existing studies mainly consider the carbon footprint of the computing system and merely consider that of the physical facilities, which also contribute significant carbon emissions. Furthermore, their capacity provisioning is neither uncertainty-aware nor adaptive to the dynamic computing demand. To bridge these gaps, we propose an adaptive capacity provisioning framework based on the physics-informed digital twin. We design the digital twin to holistically capture a DC's operational carbon footprint, including both the computing system and the physical facilities. The digital twin is differentiable and established with a collection of physics-informed learnable models that are learned with online operational data. We further address the challenge of capacity provisioning under uncertainties by designing a shrinking horizon model predictive control. The designed capacity planner updates its estimation of future computing demand based on the observable computing system states. At each capacity provisioning round, we solve the capacity provisioning problem using a gradient-based optimization technique with the gradient provided by the digital twin. We extensively evaluate our approach usingrealoperational data from a large-scale production data center. First, our digital twin accurately predicts holistic DC energy usage with a relative absolute error of less than 5%, which is accurate according to the industrial rule of thumb. Second, we show that our solution is comparable to the oracle solution with perfect knowledge about all uncertainties, outperforming the state-of-the-art Predict-then-Plan approach significantly in terms of SLO violation reduction. Furthermore, our approach reduces carbon footprint by 27% compared with the over-provisioning scheme currently adopted by the industry.
Ruihang Wang, Xin Zhou 0003, Rui Tan 0001, Yonggang Wen 0001, Yuejun Yan
IEEE Trans. Sustain. Comput.3
2024 EDPS-SST: Enhanced Dynamic Path Stitching with Structural Similarity Thresholding for Large-Scale Medical Image Stitching Under Sparse Pixel Overlap
Zhuan Han, Dixiao Tao, Bohan Yang 0015, Yong Luo 0002, Dehua Cao, Xin Zhou 0003
ICANN (8)8
2024 SCST: Spatial Consistent Swin Transformer for Multi-focus Biomedical Microscopic Image Fusion
Dengpan Liu, Bohan Yang 0015, Yong Luo 0002, Dehua Cao, Xin Zhou 0003
ICANN (8)8
2024 Depression Diagnosis and Analysis via Multimodal Multi-order Factor Fusion
Chengbo Yuan, Xuxu Liu, Yongqian Li, Yong Luo 0002, Xin Zhou 0003
ICANN (8)6
2024 Green Data Center Cooling Control via Physics-guided Safe Reinforcement Learning
abstract
Deep reinforcement learning (DRL) has shown good performance in tackling Markov decision process (MDP) problems. As DRL optimizes a long-term reward, it is a promising approach to improving the energy efficiency of data-center cooling. However, enforcement of thermal safety constraints during DRL’s state exploration is a main challenge. The widely adopted reward-shaping approach adds negative reward when the exploratory action results in unsafety. Thus, it needs to experience sufficient unsafe states before it learns how to prevent unsafety. In this article, we propose a safety-aware DRL framework for data-center cooling control. It applies offline imitation learning and online post-hoc rectification to holistically prevent thermal unsafety during online DRL. In particular, the post-hoc rectification searches for the minimum modification to the DRL-recommended action such that the rectified action will not result in unsafety. The rectification is designed based on a thermal state transition model that is fitted using historical safe operation traces and able to extrapolate the transitions to unsafe states explored by DRL. Extensive evaluation for chilled water and direct expansion-cooled data centers in two climate conditions show that our approach saves 18% to 26.6% of total data-center power compared with conventional control and reduces safety violations by 94.5% to 99% compared with reward shaping. We also extend the proposed framework to address data centers with non-uniform temperature distributions for detailed safety considerations. The evaluation shows that our approach saves 14% power usage compared with the PID control while addressing safety compliance during the training.
Ruihang Wang, Xin Zhou 0003, Yonggang Wen 0001, Rui Tan 0001
ACM Trans. Cyber Phys. Syst.3
2024 EdgeVision: Towards Collaborative Video Analytics on Distributed Edges for Performance Maximization
abstract
Deep Neural Network (DNN)-based video analytics significantly improves recognition accuracy in computer vision applications. Deploying DNN models at edge nodes, closer to end users, reduces inference delay and minimizes bandwidth costs. However, these resource-constrained edge nodes may experience substantial delays under heavy workloads, leading to imbalanced workload distribution. While previous efforts focused on optimizing hierarchical device-edge-cloud architectures or centralized clusters for video analytics, we propose addressing these challenges through collaborative distributed and autonomous edge nodes. Despite the intricate control involved, we introduce EdgeVision, a Multiagent Reinforcement Learning (MARL)-based framework for collaborative video analytics on distributed edges. EdgeVision enables edge nodes to autonomously learn policies for video preprocessing, model selection, and request dispatching. Our approach utilizes an actor-critic-based MARL algorithm enhanced with an attention mechanism to learn optimal policies. To validate EdgeVision, we construct a multi-edge testbed and conduct experiments with real-world datasets. Results demonstrate a performance enhancement of 33.6% to 86.4% compared to baseline methods.
Guanyu Gao, Yuqi Dong, Ran Wang 0004, Xin Zhou 0003
IEEE Trans. Multim.4
2024 Data Center Sustainability: Revisits and Outlooks
abstract
As energy-intensive entities, data centers are associated with significant environmental impacts, making their sustainability a subject of growing interest in recent years. In this article, we revisit data center sustainability and propose a forward-looking vision for improving data center sustainability. We argue that data center sustainability encompasses more than just energy efficiency and must be evaluated and optimized through a multi-faceted approach. To this end, we first present an overview of the sustainability metrics from five aspects. After that, we demonstrate the sustainability status of the latest data centers utilizing publicly available data center sustainability ratings. Furthermore, we examine the evolution of data center sustainability standards in Singapore to highlight several trending features. Based on the analysis, we identify several key elements of sustainable data centers. We then propose the Cognitive Digital Twin (CDT) architecture, which incorporates a digital twin engine for system-wide simulation and a decision engine for optimal control to improve data center sustainability. A case study is performed to optimize the chiller plant efficiency of a production data center in Singapore. The results demonstrate that the CDT can improve chiller plant energy efficiency by 5%, indicating around 140 metric tons of annual carbon emission savings.
Xin Zhou 0003, Zhaomeng Zhu, Tracy Liu, Jeffery Neng, Yonggang Wen 0001
IEEE Trans. Sustain. Comput.2
2023 Spatially Invariant and Frequency-Aware CycleGAN for Unsupervised MR-to-CT Synthesis
Wenbin Hu 0001, Yong Luo 0002, Xin Zhou 0003
ICANN (9)5
2023 MBMS-GAN: Multi-Band Multi-Scale Adversarial Learning for Enhancement of Coded Speech at Very Low Rate
Weiping Tu, Yong Luo 0002, Xin Zhou 0003, Li Xiao 0007, Youqiang Zheng
ICANN (7)4
2023 MFT: Multi-scale Fusion Transformer for Infrared and Visible Image Fusion
Chen-Ming Zhang, Chengbo Yuan, Yong Luo 0002, Xin Zhou 0003
ICANN (6)4
2023 Optimizing Energy Efficiency for Data Center via Parameterized Deep Reinforcement Learning
abstract
The rapid advancements in cloud computing, Big Data and their related applications have led to a skyrocketing increase in data center energy consumption year by year. The prior approaches for improving data center energy efficiency mostly suffer from high system dynamics or the complexity of data centers. In this paper, we propose an optimization framework based on deep reinforcement learning, named DeepEE, to jointly optimize energy consumption from the perspectives of task scheduling and cooling control. In DeepEE, a PArameterized action space based Deep Q-Network (PADQN) algorithm is proposed to tackle the hybrid action space problem. Then, a dynamic time factor mechanism for adjusting cooling control interval is introduced into PADQN (PADQN-D) to achieve more accurate and efficient coordination of IT and cooling subsystems. Finally, in order to train and evaluate the proposed algorithms safely and quickly, a simulation platform is built to model the dynamics of IT and cooling subsystems. Extensive real-trace based experiments illustrate that: 1) the proposed PADQN algorithm can save up to 15% and 10% energy consumption compared with the baseline siloed and joint optimization approaches respectively; 2) the proposed PADQN-D algorithm with dynamic cooling control interval can better adapt to the change of IT workload; 3) our proposed algorithms achieve more stable performance gain in terms of power consumption by adopting the parameterized action space.
Yongyi Ran, Han Hu 0003, Yonggang Wen 0001, Xin Zhou 0003
IEEE Trans. Serv. Comput.4
2023 Optimizing Data Center Energy Efficiency via Event-Driven Deep Reinforcement Learning
abstract
To reduce the skyrocketing energy consumption of data centers, the prevailing approaches adopt the time-driven manner to control IT and cooling subsystems. These methods suffer from highly dynamic system states, complex action spaces and the risk of instability caused by frequent and unnecessary control operations. To tackle these problems, we propose a novel event-driven control paradigm and an optimization algorithm, under the deep reinforcement learning (DRL) framework. The principle is to make decisions based on certain critical events (e.g., overheating), rather than fixed periodic control. Specifically, we design an event-driven optimization framework to trigger control operations. Then, we present several models to describe IT and cooling subsystems, and mathematically define events to capture four types of prior factors that impact system performance. Furthermore, we develop an event-driven DRL (E-DRL) optimization algorithm to dispatch jobs and regulate cooling facilities for energy efficiency. Using two different types of real workload traces, we conduct extensive experiments to demonstrate that: 1) E-DRL reduces the number of regulating decisions by 70%$\sim$95% while achieving a comparable or even better energy efficiency in comparison with the state-of-the-art algorithm; and 2) E-DRL can adapt the control frequency to the changing operational conditions and diverse workloads.
Yongyi Ran, Xin Zhou 0003, Han Hu 0003, Yonggang Wen 0001
IEEE Trans. Serv. Comput.2
2022 Robust Metric Boosts Transfer
abstract
Transfer metric learning (TML) aims to improve the metric learning in target domains by transferring knowledge from related tasks, where the distance metrics are strong and reliable. Existing TML approaches only focus on how to transfer the source metric knowledge, which is often prone to be over-fitting to the source domain. In this paper, we study how to train a source metric that is appropriate for transfer and then design a general deep TML method for effective metric transfer. In particular, we propose to learn the source metric parameterized by a deep neural network in an adversarial way and then transfer the metric to the target domain by embedding imitation, which allows the inputs of source and target domains to be heterogeneous. Besides, we restrict the size of the target metric network to be small so that the inference is efficient in the target domain. Results in the popular face verification application demonstrate the effectiveness of our method.
Qiancheng Yang, Yong Luo 0002, Han Hu 0003, Xin Zhou 0003, Bo Du 0001, Dacheng Tao
MMSP4
2022 Optimizing Data Centre Energy Efficiency via Event Driven Deep Reinforcement Learning
abstract
[J1C2 Presentation Abstract at IEEE SERVICES 2022 for IEEE Transactions on Services Computing DOI 10.1109/TSC.2022.3157145]
Yongyi Ran, Xin Zhou 0003, Han Hu 0003, Yonggang Wen 0001
SERVICES2
2021 Intelligent Trainer for Dyna-Style Model-Based Deep Reinforcement Learning
abstract
Model-based reinforcement learning (MBRL) has been proposed as a promising alternative solution to tackle the high sampling cost challenge in the canonical RL, by leveraging a system dynamics model to generate synthetic data for policy training purpose. The MBRL framework, nevertheless, is inherently limited by the convoluted process of jointly optimizing control policy, learning system dynamics, and sampling data from two sources controlled by complicated hyperparameters. As such, the training process involves overwhelmingly manual tuning and is prohibitively costly. In this research, we propose a "reinforcement on reinforcement" (RoR) architecture to decompose the convoluted tasks into two decoupled layers of RL. The inner layer is the canonical MBRL training process which is formulated as a Markov decision process, called training process environment (TPE). The outer layer serves as an RL agent, called intelligent trainer, to learn an optimal hyperparameter configuration for the inner TPE. This decomposition approach provides much-needed flexibility to implement different trainer designs, referred to "train the trainer." In our research, we propose and optimize two alternative trainer designs: 1) an unihead trainer and 2) a multihead trainer. Our proposed RoR framework is evaluated for five tasks in the OpenAI gym. Compared with three other baseline methods, our proposed intelligent trainer methods have a competitive performance in autotuning capability, with up to 56% expected sampling cost saving without knowing the best parameter configurations in advance. The proposed trainer framework can be easily extended to tasks that require costly hyperparameter tuning.
Linsen Dong, Xin Zhou 0003, Yonggang Wen 0001, Kyle Guan
IEEE Trans. Neural Networks Learn. Syst.3
2020 Efficient Compute-Intensive Job Allocation in Data Centers via Deep Reinforcement Learning
abstract
Reducing the energy consumption of the servers in a data center via proper job allocation is desirable. Existing advanced job allocation algorithms, based on constrained optimization formulations capturing servers' complex power consumption and thermal dynamics, often scale poorly with the data center size and optimization horizon. This article applies deep reinforcement learning to build an allocation algorithm for long-lasting and compute-intensive jobs that are increasingly seen among today's computation demands. Specifically, a deep Q-network is trained to allocate jobs, aiming to maximize a cumulative reward over long horizons. The training is performed offline using a computational model based on long short-term memory networks that capture the servers' power and thermal dynamics. This offline training approach avoids slow online convergence, low energy efficiency, and potential server overheating during the agent's extensive state-action space exploration if it directly interacts with the physical data center in the usually adopted online learning scheme. At run time, the trained Q-network is forward-propagated with little computation to allocate jobs. Evaluation based on eight months' physical state and job arrival records from a national supercomputing data center hosting 1,152 processors shows that our solution reduces computing power consumption by more than 10 percent and processor temperature by more than 4°C without sacrificing job processing throughput.
Deliang Yi, Xin Zhou 0003, Yonggang Wen 0001, Rui Tan 0001
IEEE Trans. Parallel Distributed Syst.2
2019 DeepEE: Joint Optimization of Job Scheduling and Cooling Control for Data Center Energy Efficiency Using Deep Reinforcement Learning
abstract
The past decade witnessed the tremendous growth of power consumption in data centers due to the rapid development of cloud computing, big data analytics, and machine learning, etc. The prior approaches that optimize the power consumption of the information technology (IT) system and/or the cooling system always fail to capture the system dynamics or suffer from the complexity of system states and action spaces. In this paper, we propose a Deep Reinforcement Learning (DRL) based optimization framework, named DeepEE, to improve the energy efficiency for data centers by considering the IT and cooling systems concurrently. In DeepEE, we first propose a PArameterized action space based Deep Q-Network (PADQN) algorithm to solve the hybrid action space problem and jointly optimize the job scheduling for the IT system and the airflow rate adjustment for the cooling system. Then, a two-time-scale control mechanism is applied in PADQN to coordinate the IT and cooling systems more accurately and efficiently. In addition, to train and evaluate the proposed PADQN in a safe and quick way, we build a simulation platform to model the dynamics of IT workload and cooling systems simultaneously. Through extensive real-trace based simulations, we demonstrate that: 1) our algorithm can save up to 15% and 10% energy consumption in comparison with the baseline siloed and joint optimization approaches respectively; 2) our algorithm achieves more stable performance gain in terms of power consumption by adopting the parameterized action space; and 3) our algorithm leads to a better tradeoff between energy saving and service quality.
Yongyi Ran, Han Hu 0003, Xin Zhou 0003, Yonggang Wen 0001
ICDCS3
2019 Toward Efficient Compute-Intensive Job Allocation for Green Data Centers: A Deep Reinforcement Learning Approach
abstract
Reducing the energy consumption of the servers in a data center via proper job allocation is desirable. Existing advanced job allocation algorithms, based on constrained optimization formulations capturing servers' complex power consumption and thermal dynamics, often scale poorly with the data center size and optimization horizon. This paper applies deep reinforcement learning (DRL) to build an allocation algorithm for long-lasting and compute-intensive jobs that are increasingly seen among today's computation demands. Specifically, a deep Q-network is trained to allocate jobs, aiming to maximize a cumulative reward over long horizons. The training is performed offline using a computational model based on long short-term memory networks that capture the servers' power and thermal dynamics. This offline training approach avoids slow online convergence, low energy efficiency, and potential server overheating during the DRL's extensive state-action space exploration if it directly interacts with the physical data center in the usually adopted online learning scheme. At run time, the trained Q-network is forward-propagated with little computation to allocate jobs. Evaluation based on 8 months' physical state and job arrival records from a national supercomputing data center hosting 1,152 processors shows that our solution reduces computing power consumption by nearly 10% and processor temperature by more than 3°C without sacrificing job processing throughput.
Deliang Yi, Xin Zhou 0003, Yonggang Wen 0001, Rui Tan 0001
ICDCS2
2016 An Efficient Implementation of LZW Compression in the FPGA
Xin Zhou 0003, Yasuaki Ito, Koji Nakano
ICA3PP1