Zhewen Xu

dblp:260/7744 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-8309-6084ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model Training
Hongliang Li 0003, Jie Wu 0001, Zhewen Xu, Hairui Zhao 0002, Haixiao Xu
INFOCOM4
2026 Sift: Channel-Wise Historical Embedding for High Efficiency Distributed Graph Neural Network Training with Accuracy Guarantee
abstract
Distributed Graph Neural Network (DGNN) is a powerful tool in large-scale graph representation learning. However, high data-transfer overhead among workers in a DGNN training job confines its scalability and thus the overall performance. Vertex-wise historical embedding methods have demonstrated high potential to alleviate the problems, but still suffer from severe accuracy loss and limited performance scalability, which has been attributed to the information loss of critical channels in historical vertices and redundant information in local channels. This article explores the optimization of channel level and construct a quantitative accuracy model for channel-wise historical embedding. We propose Sift, a novel DGNN training framework, supporting channel-wise partial historical embedding with accuracy guarantee. Sift has three components: a historical embedding evaluator with channel-wise quantitative accuracy model, a sawtooth-like matrix rearrangement for accelerating message passing, and a hybrid parallel framework for overlapping communication overhead. Comprehensive experimental results show that Sift achieves near-linear parallel convergence speedup, outperforming the state-of-the-art baselines by up to 72% in total training performance and up to 21% in convergence speed.
Zhewen Xu, Hongliang Li 0003, Junze Han, Hengshan Yue, Hairui Zhao 0002, Dongyuan Tian, Zijian Li 0007, Xiaohui Wei 0002
ACM Trans. Archit. Code Optim.1
2025 ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model Exploration
Hairui Zhao 0002, Hongliang Li 0003, Jie Wu 0001, Zhewen Xu, Xiang Li 0197, Haixiao Xu
INFOCOM6
2025 Alleviating straggler impacts for data parallel deep learning with hybrid parameter update
Hongliang Li 0003, Hairui Zhao 0002, Zhewen Xu
Future Gener. Comput. Syst.5
2025 Harnessing dynamic graph differential operators for efficient data-driven wind prediction
Xiaohui Wei 0002, Zhewen Xu, Hongliang Li 0003, Jieyun Hao, Hengshan Yue, Changzheng Liu
GeoInformatica2
2025 ACCNet: Adaptive cross-frequency coupling graph attention for EEG emotion recognition
Dongyuan Tian, Yucheng Wang 0001, Peiliang Gong, Zhewen Xu, Zhenghua Chen, Min Wu 0008
Neural Networks4
2025 Accurate Sea Surface Parameter Retrieval via a Dynamic Graph Physical Residual Network
abstract
Satellite data often exhibit spatial discontinuities, and the unstructured nature of these data makes it difficult for researchers to use structured spatial analysis methods such as convolutional neural networks (CNNs) in studies of sea surface parameters retrieval. To address these issues, we propose a dynamic graph physical residual network (DGPRN) for retrieving sea surface vector winds and sea surface temperatures (SSTs) via data from the microwave radiation imager (MWRI) onboard the Fengyun-3 (FY-3) satellite. This model consists of a position encoder, a dynamic adjacency matrix generator, a graph convolution module, and a physical residual module. The results indicate that the DGPRN model performs better than traditional retrieval algorithms under complex wind conditions. The DGPRN incorporates the radiative transfer equation as an inductive bias to convert relative wind direction information into brightness temperature residuals, thereby increasing the retrieval accuracy. Moreover, we capture the spatial distribution characteristics of sea surface wind speed (SSWS) and sea surface wind direction (SSWD) and retrieve continuous wind fields in areas with complex topography and cyclonic circulation. Extensive experiments indicate that the DGPRN model outperforms other competing methods in both retrieval accuracy and generalization performance. Specifically, DGPRN model achieves root mean square error (RMSE) values of 0.87 m/s for the wind speed, 14.80° for the wind direction, and 0.81 K for the SST, with correlation coefficients exceeding 0.93.
Renge Zhou, Zhewen Xu, Changzheng Liu
IEEE Trans. Geosci. Remote. Sens.2
2024 HiRM: Hierarchical resource management for earth system models on many-core clusters
Zhewen Xu, Xiaohui Wei 0002, Jieyun Hao, Hongliang Li 0003, Zhaohui Ding
CCF Trans. High Perform. Comput.1
2024 DGFormer: a physics-guided station level weather forecasting model with dynamic spatial-temporal graph neural network
Zhewen Xu, Xiaohui Wei 0002, Jieyun Hao, Junze Han, Hongliang Li 0003, Changzheng Liu, Zijian Li 0007, Dongyuan Tian, Nong Zhang
GeoInformatica1
2024 Exploiting Complex Network-Based Clustering for Personalization-Enhanced Hierarchical Federated Edge Learning
abstract
Federated Learning (FL) has been extensively applied in urban environmental prediction tasks of mobile edge computing by training a global machine learning model without data sharing. However, the training of FL faces the challenges such as the poor generalization capability of a single global model over heterogeneous data and hefty communication overhead caused by the frequent model exchange between massive edge servers and remote cloud servers. To address such issues, we propose HPFL-CN, a novel communication-efficient Hierarchical Personalized Federated edge Learning framework with Complex Network clustering. HPFL-CN introduces Privacy-preserving Feature Clustering (PFC) to extract privacy-preserving low-dimensional feature representations of each edge server via mapping the environmental data to different complex network domains for clustering similar edge servers accurately. Based on the clustering results of PFC, anedge-mediator-cloudhierarchical architecture is proposed to realize personalization at the cluster level by Effective Hierarchical Scheduling (EHS). Furthermore, to adapt to dynamic scenarios of new edge servers joining and streaming data generation, we further extend HPFL-CN to Adaptive personalized federated learning with dynamic grouping (Ada-HPFL-CN), which can flexibly re-group edge servers and adjust mixed model weights and the model aggregation frequency adaptively. Our extensive experiments on real-world datasets demonstrate the efficacy of our framework, which outperforms state-of-the-art FL methods regarding personalization and communication efficiency performance.
Zijian Li 0007, Zihan Chen 0001, Xiaohui Wei 0002, Shang Gao 0005, Hengshan Yue, Zhewen Xu, Tony Q. S. Quek
IEEE Trans. Mob. Comput.6
2023 ExplSched: Maximizing Deep Learning Cluster Efficiency for Exploratory Jobs
abstract
Resource management for Deep Learning (DL) clusters is essential for system efficiency and model training quality. Existing schedulers provided by DL frameworks are mostly adaptations from traditional HPC clusters and usually work on jobs’ makespan, assuming that DL training jobs finish completely. Unfortunately, it is reported that a fair amount of training jobs are exploratory jobs and often finish unsuccessfully (over 30%) in production clusters. This is due to the distinct characteristic of Deep Neural Network (DNN) training that it is an exploratory process of frequent user interventions, such as adjusting model structures, tuning hyperparameters, and exploring feature validity. Existing DL cluster schedulers using offline algorithms are not suitable for exploratory jobs when unexpected early terminations can cause noticeable resource waste. Moreover, DL training jobs are iterative and usually yield diminishing returns as they progress. Equally allocating resource among training iterations is not efficient, especially when dealing with exploratory jobs where it can worsen the degradation of system efficiency. The fundamental goal of a DL training job is to gain model quality improvement, usually indicated by the loss reduction (job profit) of a DNN model. This paper introduces a novel scheduling problem for exploratory jobs that seeks to maximize the overall training profit of a DL cluster. We propose ExplSched, an online scheduling solution based on the primal-dual framework, resulting in a competitive ratio of 2α that belongs to O(ln n). It uses a resource price function that emphasizes the importance of job profit to resource consumption ratio to make quick resource allocation decisions. Experimental results show that ExplSched achieved an average system utility improvement of 87.28% compared with other related work.
Hongliang Li 0003, Hairui Zhao 0002, Zhewen Xu, Xiang Li 0197, Haixiao Xu
CLUSTER3
2021 Coordinated process scheduling algorithms for coupled earth system models
abstract
Abstract It is becoming increasingly significant for humans to predict and understand future climate changes using coupled climate system models. Although the performance and scalability of individual physical components have improved over the past few years, coupled climate systems still suffer from low efficiency. This paper focuses on the process scheduling problem for the widely applied coupled earth system model (CESM). The proposed resource allocation strategies allow components to execute on a compromised suboptimal setup and still maintain approximately the best parallel speedup. With this flexible resource allocation strategy, we further propose a coordinated process scheduling algorithm (CPSA). More notably, we propose an upgraded version called CPSA‐B, which makes efficient resource sharing configurations, including resource allocation and process layout of components. We integrate CPSA and CPSA‐B as pre‐arrangement tools into the CESM program and deploy them on the Huawei Kunpeng platform. The speedup curves of the CESM components are prepared in advance, based on sampling tests. Experimental data show that CPSA‐B reduces up to 58% of the execution time compared with the CESM default strategy. The algorithm has low complexity and can efficiently find solutions for large input sizes.
Xiaohui Wei 0002, Zhewen Xu, Hongliang Li 0003, Zhaohui Ding
Concurr. Comput. Pract. Exp.2
2020 CPSA: A Coordinated Process Scheduling Algorithm for Coupled Earth System Model
abstract
Coupled climate system models are important tools for climatologists to predict and understand future climate. These models are usually resource-consuming due to the large number of processors required and long execution time. Although the performance and scalability of individual physical system model have been improved over the past years, coupled climate systems still suffer from low efficiency when sharing resource across models. This paper focuses on the process scheduling strategy of Coupled Earth System Model (CESM), a widely applied coupled system model. Instead of pursuing best speedup efficiency for individual component, the proposed resource allocation strategy allows components to execute on compromised sub-optimal setup and still maintains relatively high parallel speedup. With this flexible resource allocation strategy, we further propose a Coordinated Process Scheduling Algorithm (CPSA) to make efficient resource sharing configurations, including resource allocation and process layout of components. We integrate CPSA as a tool into CESM program, and deploy it on Huawei Kunpeng Platform. Speedup curves of CESM components are prepared in advance based on sampling tests. Experimental data show that our algorithm reduces up to 52.6% of execution time compared with CESM default strategy. We also present simulation data to show that our algorithm is efficient for the platforms with up to a million cores.
Hongliang Li 0003, Zhewen Xu, Fangyu Tang, Xiaohui Wei 0002, Zhaohui Ding
ICCCN2