Wenxiang Yang

dblp:191/3081 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-7912-0181ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Promoting Resource Utilization in HPC via Scheduling
abstract
ABSTRACT The increasing complexity of supercomputing workloads poses challenges to efficient resource management, especially in balancing computational and I/O demands. Shared burst buffers, as high‐speed intermediate storage, offer a promising avenue to mitigate I/O bottlenecks. However, naive job scheduling strategies often neglect the potential of burst buffers, relying on heuristic methods with limited adaptability. To address this, we propose an innovative burst‐buffer‐aware scheduling framework that integrates burst buffer capacity into job scheduling as a key resource. Through multi‐objective optimization, the framework intelligently balances trade‐offs among job waiting time, slowdown, and completion time, surpassing the rigidity of conventional approaches. Leveraging real‐world workload traces, the framework dynamically adapts to varying windows and job characteristics, combining computation and burst buffer demands to optimize scheduling decisions. Experimental results reveal that the proposed framework enhances scheduling efficiency and system adaptability, establishing a smarter and more effective approach to supercomputing job scheduling. This work underscores the importance of burst‐buffer‐aware strategies in advancing high‐performance computing, offering novel insights into intelligent resource management.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang, Bao Li 0002
Concurr. Comput. Pract. Exp.4
2024 Examining SSD-correlated Failures within Racks in Production Data Centers
abstract
With flash-based solid-state drives (SSDs) becoming the primary storage medium in data centers, SSD failures are emerging as a significant factor affecting the reliability of data center storage systems. However, our understanding of the temporal and spatial distribution of field SSD failures remains limited, constraining the effectiveness of redundancy protection schemes in data center storage. To achieve high-quality rack-level fault tolerance, we conduct an in-depth analysis of correlated failures among 100K SSDs in Alibaba’s data center. The SSD-hosted applications exert a constant influence on the operation of the data center. We categorize intra-rack failures into heterogeneous application (hete-app) and homogeneous application (homo-app) failures based on the applications. We conduct a thorough investigation into the spatiotemporal trends of these two types of in-rack failures, exploring their causes and assessing the feasibility of rearranging racks to mitigate the occurrence of intra-rack failure chains. Furthermore, we employ a trace-driven simulator to verify the impact of different redundancy schemes on the reliability in clusters of hete-app racks and homo-app racks under high-failure-percentage environments.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang
HPCC4
2024 Trade-off topology design for hierarchical network based on job characteristics
Wenxiang Yang, Jie Yu 0006
CCF Trans. High Perform. Comput.1
2024 FP-JSC: Job failure prediction on supercomputers through job application sequence correlation
abstract
Summary Supercomputers are advanced computing systems interconnected through high‐speed communication networks, consisting of independent computational nodes. During the unfolding of the big data era, the potent computational capabilities of these supercomputers play a pivotal role in scientific computing. Despite executing numerous advanced computational science and engineering tasks on supercomputers, many submitted jobs fail due to various factors, resulting in user inefficiencies. These failures not only consume system resources but also reduce the overall efficiency of the system. Previous research often couples job performance features with a single machine learning method for predicting job failure. However, a primary hurdle emerges from the high cost of gathering these features, complicating their real‐world applicability. To address this challenge, our study establishes correlations among job applications through extensive job log analysis. Leveraging correlations, we propose a predictive framework based on job application sequence correlation (called FP‐JSC). This innovative framework employs multiple machine learning models to offer holistic predictions, selecting the most suitable model based on its learning effectiveness. Moreover, the framework optimizes feature collection expenses without adversely affecting job execution. We determine job applications using both job paths and job names, with the former emerging as a novel feature derived from supplementary monitoring data. Empirical results underscore FP‐JSC's effectiveness, accurately identifying over 89% of jobs with 95% specificity and 89% sensitivity—outperforming single prediction methods employed in related works.
Gang Xian, Wenxiang Yang, Jie Yu 0006
Concurr. Comput. Pract. Exp.2
2024 Mobilizing underutilized storage nodes via job path: A job-aware file striping approach
Gang Xian, Wenxiang Yang, Yusong Tan, Jinghua Feng, Jie Yu 0006
Parallel Comput.2
2023 Exploring job running path to predict runtime on multiple production supercomputers
Wenxiang Yang, Xiangke Liao, Dezun Dong, Jie Yu 0006
J. Parallel Distributed Comput.1
2023 An Intelligent Method for Predicting the Pressure Coefficient Curve of Airfoil-Based Conditional Generative Adversarial Networks
abstract
Recently, extensive studies have focused on analyzing aerodynamic performance due to its important impact on aircraft design. Most of these works compute the aerodynamic coefficient of the airfoil through computational fluid dynamics (CFD) simulation, which is too time-consuming. To reduce the computational time required, some intelligence-based methods have been presented. However, these methods also suffer from certain issues. First, most of them directly implement existing machine learning methods used to predict the aerodynamic coefficient without adding any improvements. Second, some methods convert the airfoil shape and aerodynamic curves into images, which may lead to curve distortion and the introduction of noise. Third, some methods learn the relationship between the airfoil shape and aerodynamic coefficients but ignore the influence of initial inflow conditions. Accordingly, to address these issues, we propose an intelligent method for predicting the pressure coefficients (Cp) of airfoil based on a conditional generative adversarial network (cGAN). More specifically, we first present a two-step data augmentation strategy designed to expand the original airfoil dataset. Subsequently, we design a novel cGAN-based neural network to predict the Cp curve. To the best of our knowledge, this is the first work to apply generative adversarial network (GAN) to aerodynamic coefficient prediction. Moreover, we design a new loss function to train our network. Extensive experimental results demonstrate that the Cp curve predicted by our method is very close to that generated via CFD simulation. More importantly, our method achieves a speedup close to 1000x compared with CFD simulation.
Yueqing Wang, Liang Deng, Yunbo Wan, Zhigong Yang, Wenxiang Yang, Cheng Chen 0005, Dan Zhao 0002, Fang Wang 0004
IEEE Trans. Neural Networks Learn. Syst.5
2022 A Quantitative Study of the Spatiotemporal I/O Burstiness of HPC Application
abstract
Understanding the I/O characteristics of applications on supercomputers is crucial to paving the path for application optimization and system resource allocation. We collect and analyze I/O traces of applications on a production supercomputer and reconfirm that I/O bursts exist in most applications. What's more, we find that the I/O bursts not only occur in short periods of time but also originate from a minority of adjacent compute nodes allocated to the applications, which we call spatiotemporal I/O burstiness. The concentration of I/O traffic in both time and space dimension will make applications experience poor I/O performance and incur I/O inefficiency of the storage system. Although there are some solutions, such as burst buffer, can help alleviate such inefficiency, there is still no work that measures, analyzes and further predicts the application I/O characteristic in terms of spatiotemporal burstiness, which we think is vital for application-aware optimizations, including but not limited to burst buffer allocation and job scheduling. In this paper, we first propose a mathematical model to measure the spatiotemporal I/O burstiness. Then a thorough analysis on the spatiotemporal I/O characteristic of all applications on the system is elaborated. We further make use of the job's submitting path to explore the I/O characteristic similarity among jobs, based on which a machine learning classification algorithm is proposed to accurately predict the job spatiotemporal I/O burstiness in advance. With accurate job I/O characteristic at hand, some useful suggestions are put forward to hedge the impacts of the spatiotemporal I/O burstiness.
Wenxiang Yang, Xiangke Liao, Dezun Dong, Jie Yu 0006
IPDPS1
2022 FastCredit: Expediting credit-based congestion control in datacenters
Shan Huang 0002, Dezun Dong, Zejia Zhou, Hanyi Shi, Wenxiang Yang, Xiangke Liao
Comput. Networks5
2022 PreF: Predicting job failure on supercomputers with job path and user behavior
abstract
Abstract Large numbers of jobs are executed on supercomputers almost every day. Unfortunately, many jobs would fail for various reasons, resulting in the waste of resources and the prolonged waiting time for queuing jobs. Job failure prediction can guide adjustment measures in advance, which is vital to the system's overall execution efficiency and reliability. Aiming at the problem that the existing job failure prediction methods are single, the collection of job features is complex and challenging to apply. This article strives to study whether these failed jobs can be predicted with known and synthetic features. We perform a comprehensive analysis of large amounts of historical data and various features and find that two novel features (running path and retry count) can predict job failure well. The running path indicates the application type a job belongs to, and the retry count reflects the user's behavior when the job fails. We propose a job failure prediction framework called PreF on supercomputers using machine learning based on the novel features. The experimental results show that PreF can correctly identify over 89% of jobs, outperforming the latest related methods on the comprehensive evaluation indicator (S_score) by around 4%.
Gang Xian, Jie Yu 0006, Wenxiang Yang, Longfang Zhou
Concurr. Comput. Pract. Exp.5
2021 PREP: Predicting Job Runtime with Job Running Path on Supercomputers
abstract
Supercomputers serve a lot of parallel jobs by scheduling jobs and allocating computing resources. One popular scheduling strategy is First Come First Serve (FCFS). However, there are always some idle resources not being effectively utilized, since they are not enough and are reserved for the head job in the waiting queue. To improve resource utilization, a common solution is to use backfilling, which allocates the reserved computing resources to a small, short job selected from the queue, on the premise of not delaying the original head job. Unfortunately, the estimated job runtime provided by users is often overestimated. Previous studies extract features from historical job logs and predict runtime based on machine learning. However, traditional features (e.g. CPU, user, submitting time, etc.) are insufficient to describe the characteristics of jobs. In this paper, we propose a novel runtime prediction framework called PREP. It explores a new feature named job running path, which encodes important implications about the job’s characteristics, such as the project it belongs to, data sets and parameters it uses, etc. As there is a strong correlation between job runtime and its running path. PREP groups jobs into separate clusters according to their running paths and trains a runtime prediction model for each job cluster. Final results demonstrate that adding the new feature can achieve high prediction accuracy of 88% and has a better prediction effect than other methods, such as Last-2 and IRPA.
Longfang Zhou, Wenxiang Yang, Yongguo Han, Jie Yu 0006
ICPP3
2020 FastCredit: Expediting Credit-based Proactive Transports in Datacenters
abstract
Recent proposals have leveraged emerging credit-based proactive transports to achieve high throughput low latency datacenter network transports. Particularly, those transports that employ hop-by-hop credits have the merits of fast convergence, low buffer occupancy, and strong congestion avoidability. However, they fairly transmit long flows and latency-sensitive short flows, which will cause the transmission latency of short flows and the average flow completion time increased. Although flow scheduling mechanisms have studied extensively to accelerate short flow transmission, they are hard to be directly applied in credit-based transports. The root cause is that most traditional flow scheduling mechanisms mainly work in the long queue containing flows in various sizes, while credit-based proactive transports maintain the extremely short bounded queue, near zero. Based on this observation, this paper makes the first attempt to accelerate short-flow scheduling in credit-based proactive transport, and proposed FastCredit. FastCredit can be used as a general building block to expedite short flows in credit-based proactive transports. In FastCredit, we schedule credit transmission at both receivers and switches to indirectly perform flow scheduling, and develop a mechanism to mitigate credit waste and improve network goodput. Compared to the state-of-the-art credit-based transport protocol, FastCredit reduces average flow completion time to 0.78x and greatly improves the short flow transmission latency to 0.51x in realistic workloads. Especially, FastCredit reduces average flow completion time to 0.76x under incast circumstances and 0.62x in many-to-one traffic mode. Furthermore, FastCredit still maintains the advantages of short queue and high throughput.
Dezun Dong, Shan Huang 0002, Zejia Zhou, Wenxiang Yang, Hanyi Shi
ICPADS4
2020 Spatially Bursty I/O on Supercomputers: Causes, Impacts and Solutions
abstract
Understanding the I/O characteristics of supercomputers is crucial for grasping accurate I/O workloads and uncovering potential I/O inefficiency. We collect and analyze I/O traces from two production supercomputers, and find that the I/O traffic peaks in the system not only occur in short periods of time but also originate from a minority of adjacent compute nodes, which we call spatially bursty I/O. Since modern supercomputers widely adopt I/O forwarding architecture, in which an I/O node performs I/O on behalf of a subset of compute nodes in the vicinity, spatially bursty I/O will cause significant load imbalance and underutilization on the I/O nodes. To address such problems, we quantitatively analyze the two causes of spatially bursty I/O, including uneven I/O distribution on job's processes and uneven job nodes distribution on the system. Two different solutions are proposed to mobilize more I/O nodes to participate in job's I/O activity. (1) We change the I/O node mapping, making adjacent compute nodes use different I/O nodes instead of a same one. (2) According to the job's I/O characteristics extracted from history I/O traces, we distribute the compute nodes of data-intensive jobs more sparsely to utilize more I/O nodes. Extensive evaluations of both solutions show that they can further exploit the potential of I/O forwarding layer. We have deployed the proposed I/O node mapping on a production supercomputer for 11 months. Our experience finds that it can effectively promote I/O performance, balance loads, and alleviate I/O interference.
Jie Yu 0006, Wenxiang Yang, Dezun Dong, Jinghua Feng
IEEE Trans. Parallel Distributed Syst.2
2016 MBL: A Multi-stage Bufferless High-radix Router
abstract
There is a pressing need for high-radix routers in modern HPC (High Performance Computing) interconnects and to build the exascale computers with massive clusters. In this paper, we propose MBL, a high-radix router with a multi-stage bufferless switch Clos network inside. Booksim interconnection network simulator is used to implement our arbitrating designs for the architecture and it runs well under different traffic patterns in a flattened butterfly network, with 136 ports for each router.
Wenxiang Yang, Dezun Dong, Jingyue Zhao, Cunlu Li
CLUSTER1