Gang Xian

dblp:330/3516 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-0388-3406ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Promoting Resource Utilization in HPC via Scheduling
abstract
ABSTRACT The increasing complexity of supercomputing workloads poses challenges to efficient resource management, especially in balancing computational and I/O demands. Shared burst buffers, as high‐speed intermediate storage, offer a promising avenue to mitigate I/O bottlenecks. However, naive job scheduling strategies often neglect the potential of burst buffers, relying on heuristic methods with limited adaptability. To address this, we propose an innovative burst‐buffer‐aware scheduling framework that integrates burst buffer capacity into job scheduling as a key resource. Through multi‐objective optimization, the framework intelligently balances trade‐offs among job waiting time, slowdown, and completion time, surpassing the rigidity of conventional approaches. Leveraging real‐world workload traces, the framework dynamically adapts to varying windows and job characteristics, combining computation and burst buffer demands to optimize scheduling decisions. Experimental results reveal that the proposed framework enhances scheduling efficiency and system adaptability, establishing a smarter and more effective approach to supercomputing job scheduling. This work underscores the importance of burst‐buffer‐aware strategies in advancing high‐performance computing, offering novel insights into intelligent resource management.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang, Bao Li 0002
Concurr. Comput. Pract. Exp.1
2024 Examining SSD-correlated Failures within Racks in Production Data Centers
abstract
With flash-based solid-state drives (SSDs) becoming the primary storage medium in data centers, SSD failures are emerging as a significant factor affecting the reliability of data center storage systems. However, our understanding of the temporal and spatial distribution of field SSD failures remains limited, constraining the effectiveness of redundancy protection schemes in data center storage. To achieve high-quality rack-level fault tolerance, we conduct an in-depth analysis of correlated failures among 100K SSDs in Alibaba’s data center. The SSD-hosted applications exert a constant influence on the operation of the data center. We categorize intra-rack failures into heterogeneous application (hete-app) and homogeneous application (homo-app) failures based on the applications. We conduct a thorough investigation into the spatiotemporal trends of these two types of in-rack failures, exploring their causes and assessing the feasibility of rearranging racks to mitigate the occurrence of intra-rack failure chains. Furthermore, we employ a trace-driven simulator to verify the impact of different redundancy schemes on the reliability in clusters of hete-app racks and homo-app racks under high-failure-percentage environments.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang
HPCC1
2024 FP-JSC: Job failure prediction on supercomputers through job application sequence correlation
abstract
Summary Supercomputers are advanced computing systems interconnected through high‐speed communication networks, consisting of independent computational nodes. During the unfolding of the big data era, the potent computational capabilities of these supercomputers play a pivotal role in scientific computing. Despite executing numerous advanced computational science and engineering tasks on supercomputers, many submitted jobs fail due to various factors, resulting in user inefficiencies. These failures not only consume system resources but also reduce the overall efficiency of the system. Previous research often couples job performance features with a single machine learning method for predicting job failure. However, a primary hurdle emerges from the high cost of gathering these features, complicating their real‐world applicability. To address this challenge, our study establishes correlations among job applications through extensive job log analysis. Leveraging correlations, we propose a predictive framework based on job application sequence correlation (called FP‐JSC). This innovative framework employs multiple machine learning models to offer holistic predictions, selecting the most suitable model based on its learning effectiveness. Moreover, the framework optimizes feature collection expenses without adversely affecting job execution. We determine job applications using both job paths and job names, with the former emerging as a novel feature derived from supplementary monitoring data. Empirical results underscore FP‐JSC's effectiveness, accurately identifying over 89% of jobs with 95% specificity and 89% sensitivity—outperforming single prediction methods employed in related works.
Gang Xian, Wenxiang Yang, Jie Yu 0006
Concurr. Comput. Pract. Exp.1
2024 Mobilizing underutilized storage nodes via job path: A job-aware file striping approach
Gang Xian, Wenxiang Yang, Yusong Tan, Jinghua Feng, Jie Yu 0006
Parallel Comput.1
2024 Hybrid Decision-Making for Intelligent High-Speed Train Operation: A Boundary Constraint and Pre-Evaluation Reinforcement Learning Approach
abstract
Deep Reinforcement Learning (DRL) is the most promising technology for improving high-speed train’s energy efficiency and operation quality. Existing solutions, however, suffer from three significant limitations: 1) They cannot effectively constrain the huge exploration space generated by high-speed trains under high temporal deformability and long-distance trips; 2) The reward function has no adaptability to the different energy-efficiency difficulties of different travel schedules, and the agent will receive incorrect reward signals, requiring manual adjustment; 3) They do not avoid the invalid action sequences of the agent well. To address this challenge, we propose a revolutionary Boundary Constrained and Pre-evaluated Reinforcement Learning (BCPRL) approach to alleviate these issues. This approach combines the Shrink Trajectory Exploration Space (STES) module, the Pre-evaluated Energy-efficiency Scenario Complexity (PESC) module, and the Twin Delayed Deep Deterministic Policy Gradient (TD3) module and uses a hybrid of STES and TD3 to make train operation decision-making to improve the operation quality and learning efficiency of the agent. Numerical experiments validate the effectiveness of the BCPRL approach, which, by drastically reducing the exploration space and getting the agent the correct reward signal, not only maintains excellence in efficiency and punctuality but also far surpasses the other baseline approaches in learning efficiency and robustness.
Haotong Zhang 0004, Deqing Huang, Deqiang He, Shixun Wu, Gang Xian
IEEE Trans. Intell. Transp. Syst.6
2023 ASTPSI: Allocating Spare Time and Planning Speed Interval for Intelligent Train Control of Sparse Reward
Haotong Zhang 0004, Gang Xian
ICONIP (1)2
2022 PreF: Predicting job failure on supercomputers with job path and user behavior
abstract
Abstract Large numbers of jobs are executed on supercomputers almost every day. Unfortunately, many jobs would fail for various reasons, resulting in the waste of resources and the prolonged waiting time for queuing jobs. Job failure prediction can guide adjustment measures in advance, which is vital to the system's overall execution efficiency and reliability. Aiming at the problem that the existing job failure prediction methods are single, the collection of job features is complex and challenging to apply. This article strives to study whether these failed jobs can be predicted with known and synthetic features. We perform a comprehensive analysis of large amounts of historical data and various features and find that two novel features (running path and retry count) can predict job failure well. The running path indicates the application type a job belongs to, and the retry count reflects the user's behavior when the job fails. We propose a job failure prediction framework called PreF on supercomputers using machine learning based on the novel features. The experimental results show that PreF can correctly identify over 89% of jobs, outperforming the latest related methods on the comprehensive evaluation indicator (S_score) by around 4%.
Gang Xian, Jie Yu 0006, Wenxiang Yang, Longfang Zhou
Concurr. Comput. Pract. Exp.1