Jie Yu 0006

dblp:74/3437-6 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-9912-2627ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
YearPublicationVenuePosition
2025 Promoting Resource Utilization in HPC via Scheduling
abstract
ABSTRACT The increasing complexity of supercomputing workloads poses challenges to efficient resource management, especially in balancing computational and I/O demands. Shared burst buffers, as high‐speed intermediate storage, offer a promising avenue to mitigate I/O bottlenecks. However, naive job scheduling strategies often neglect the potential of burst buffers, relying on heuristic methods with limited adaptability. To address this, we propose an innovative burst‐buffer‐aware scheduling framework that integrates burst buffer capacity into job scheduling as a key resource. Through multi‐objective optimization, the framework intelligently balances trade‐offs among job waiting time, slowdown, and completion time, surpassing the rigidity of conventional approaches. Leveraging real‐world workload traces, the framework dynamically adapts to varying windows and job characteristics, combining computation and burst buffer demands to optimize scheduling decisions. Experimental results reveal that the proposed framework enhances scheduling efficiency and system adaptability, establishing a smarter and more effective approach to supercomputing job scheduling. This work underscores the importance of burst‐buffer‐aware strategies in advancing high‐performance computing, offering novel insights into intelligent resource management.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang, Bao Li 0002
Concurr. Comput. Pract. Exp.3
2024 Examining SSD-correlated Failures within Racks in Production Data Centers
abstract
With flash-based solid-state drives (SSDs) becoming the primary storage medium in data centers, SSD failures are emerging as a significant factor affecting the reliability of data center storage systems. However, our understanding of the temporal and spatial distribution of field SSD failures remains limited, constraining the effectiveness of redundancy protection schemes in data center storage. To achieve high-quality rack-level fault tolerance, we conduct an in-depth analysis of correlated failures among 100K SSDs in Alibaba’s data center. The SSD-hosted applications exert a constant influence on the operation of the data center. We categorize intra-rack failures into heterogeneous application (hete-app) and homogeneous application (homo-app) failures based on the applications. We conduct a thorough investigation into the spatiotemporal trends of these two types of in-rack failures, exploring their causes and assessing the feasibility of rearranging racks to mitigate the occurrence of intra-rack failure chains. Furthermore, we employ a trace-driven simulator to verify the impact of different redundancy schemes on the reliability in clusters of hete-app racks and homo-app racks under high-failure-percentage environments.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang
HPCC3
2024 Trade-off topology design for hierarchical network based on job characteristics
Wenxiang Yang, Jie Yu 0006
CCF Trans. High Perform. Comput.2
2024 FP-JSC: Job failure prediction on supercomputers through job application sequence correlation
abstract
Summary Supercomputers are advanced computing systems interconnected through high‐speed communication networks, consisting of independent computational nodes. During the unfolding of the big data era, the potent computational capabilities of these supercomputers play a pivotal role in scientific computing. Despite executing numerous advanced computational science and engineering tasks on supercomputers, many submitted jobs fail due to various factors, resulting in user inefficiencies. These failures not only consume system resources but also reduce the overall efficiency of the system. Previous research often couples job performance features with a single machine learning method for predicting job failure. However, a primary hurdle emerges from the high cost of gathering these features, complicating their real‐world applicability. To address this challenge, our study establishes correlations among job applications through extensive job log analysis. Leveraging correlations, we propose a predictive framework based on job application sequence correlation (called FP‐JSC). This innovative framework employs multiple machine learning models to offer holistic predictions, selecting the most suitable model based on its learning effectiveness. Moreover, the framework optimizes feature collection expenses without adversely affecting job execution. We determine job applications using both job paths and job names, with the former emerging as a novel feature derived from supplementary monitoring data. Empirical results underscore FP‐JSC's effectiveness, accurately identifying over 89% of jobs with 95% specificity and 89% sensitivity—outperforming single prediction methods employed in related works.
Gang Xian, Wenxiang Yang, Jie Yu 0006
Concurr. Comput. Pract. Exp.3
2024 Mobilizing underutilized storage nodes via job path: A job-aware file striping approach
Gang Xian, Wenxiang Yang, Yusong Tan, Jinghua Feng, Jie Yu 0006
Parallel Comput.7
2023 Exploring job running path to predict runtime on multiple production supercomputers
Wenxiang Yang, Xiangke Liao, Dezun Dong, Jie Yu 0006
J. Parallel Distributed Comput.4
2022 A Quantitative Study of the Spatiotemporal I/O Burstiness of HPC Application
abstract
Understanding the I/O characteristics of applications on supercomputers is crucial to paving the path for application optimization and system resource allocation. We collect and analyze I/O traces of applications on a production supercomputer and reconfirm that I/O bursts exist in most applications. What's more, we find that the I/O bursts not only occur in short periods of time but also originate from a minority of adjacent compute nodes allocated to the applications, which we call spatiotemporal I/O burstiness. The concentration of I/O traffic in both time and space dimension will make applications experience poor I/O performance and incur I/O inefficiency of the storage system. Although there are some solutions, such as burst buffer, can help alleviate such inefficiency, there is still no work that measures, analyzes and further predicts the application I/O characteristic in terms of spatiotemporal burstiness, which we think is vital for application-aware optimizations, including but not limited to burst buffer allocation and job scheduling. In this paper, we first propose a mathematical model to measure the spatiotemporal I/O burstiness. Then a thorough analysis on the spatiotemporal I/O characteristic of all applications on the system is elaborated. We further make use of the job's submitting path to explore the I/O characteristic similarity among jobs, based on which a machine learning classification algorithm is proposed to accurately predict the job spatiotemporal I/O burstiness in advance. With accurate job I/O characteristic at hand, some useful suggestions are put forward to hedge the impacts of the spatiotemporal I/O burstiness.
Wenxiang Yang, Xiangke Liao, Dezun Dong, Jie Yu 0006
IPDPS4
2022 PreF: Predicting job failure on supercomputers with job path and user behavior
abstract
Abstract Large numbers of jobs are executed on supercomputers almost every day. Unfortunately, many jobs would fail for various reasons, resulting in the waste of resources and the prolonged waiting time for queuing jobs. Job failure prediction can guide adjustment measures in advance, which is vital to the system's overall execution efficiency and reliability. Aiming at the problem that the existing job failure prediction methods are single, the collection of job features is complex and challenging to apply. This article strives to study whether these failed jobs can be predicted with known and synthetic features. We perform a comprehensive analysis of large amounts of historical data and various features and find that two novel features (running path and retry count) can predict job failure well. The running path indicates the application type a job belongs to, and the retry count reflects the user's behavior when the job fails. We propose a job failure prediction framework called PreF on supercomputers using machine learning based on the novel features. The experimental results show that PreF can correctly identify over 89% of jobs, outperforming the latest related methods on the comprehensive evaluation indicator (S_score) by around 4%.
Gang Xian, Jie Yu 0006, Wenxiang Yang, Longfang Zhou
Concurr. Comput. Pract. Exp.3
2021 PREP: Predicting Job Runtime with Job Running Path on Supercomputers
abstract
Supercomputers serve a lot of parallel jobs by scheduling jobs and allocating computing resources. One popular scheduling strategy is First Come First Serve (FCFS). However, there are always some idle resources not being effectively utilized, since they are not enough and are reserved for the head job in the waiting queue. To improve resource utilization, a common solution is to use backfilling, which allocates the reserved computing resources to a small, short job selected from the queue, on the premise of not delaying the original head job. Unfortunately, the estimated job runtime provided by users is often overestimated. Previous studies extract features from historical job logs and predict runtime based on machine learning. However, traditional features (e.g. CPU, user, submitting time, etc.) are insufficient to describe the characteristics of jobs. In this paper, we propose a novel runtime prediction framework called PREP. It explores a new feature named job running path, which encodes important implications about the job’s characteristics, such as the project it belongs to, data sets and parameters it uses, etc. As there is a strong correlation between job runtime and its running path. PREP groups jobs into separate clusters according to their running paths and trains a runtime prediction model for each job cluster. Final results demonstrate that adding the new feature can achieve high prediction accuracy of 88% and has a better prediction effect than other methods, such as Last-2 and IRPA.
Longfang Zhou, Wenxiang Yang, Yongguo Han, Jie Yu 0006
ICPP7
2020 Spatially Bursty I/O on Supercomputers: Causes, Impacts and Solutions
abstract
Understanding the I/O characteristics of supercomputers is crucial for grasping accurate I/O workloads and uncovering potential I/O inefficiency. We collect and analyze I/O traces from two production supercomputers, and find that the I/O traffic peaks in the system not only occur in short periods of time but also originate from a minority of adjacent compute nodes, which we call spatially bursty I/O. Since modern supercomputers widely adopt I/O forwarding architecture, in which an I/O node performs I/O on behalf of a subset of compute nodes in the vicinity, spatially bursty I/O will cause significant load imbalance and underutilization on the I/O nodes. To address such problems, we quantitatively analyze the two causes of spatially bursty I/O, including uneven I/O distribution on job's processes and uneven job nodes distribution on the system. Two different solutions are proposed to mobilize more I/O nodes to participate in job's I/O activity. (1) We change the I/O node mapping, making adjacent compute nodes use different I/O nodes instead of a same one. (2) According to the job's I/O characteristics extracted from history I/O traces, we distribute the compute nodes of data-intensive jobs more sparsely to utilize more I/O nodes. Extensive evaluations of both solutions show that they can further exploit the potential of I/O forwarding layer. We have deployed the proposed I/O node mapping on a production supercomputer for 11 months. Our experience finds that it can effectively promote I/O performance, balance loads, and alleviate I/O interference.
Jie Yu 0006, Wenxiang Yang, Dezun Dong, Jinghua Feng
IEEE Trans. Parallel Distributed Syst.1
2019 WatCache: a workload-aware temporary cache on the compute side of HPC systems
Jie Yu 0006, Wenrui Dong, Xiaoyong Li 0002
J. Supercomput.1
2018 Rethinking Node Allocation Strategy for Data-intensive Applications in Consideration of Spatially Bursty I/O
abstract
Job scheduling in HPC systems by default allocate adjacent compute nodes for jobs for lower communication overhead. However, it is no longer applicable to data-intensive jobs running on systems with I/O forwarding layer, where each I/O node performs I/O on behalf of a subset of compute nodes in the vicinity. Under the default node allocation strategy a job's nodes are located close to each other and thus it only uses a limited number of I/O nodes. Since the I/O activities of jobs are bursty, at any moment only a minority of jobs in the system are busy processing I/O. Consequently, the bursty I/O traffic in the system is also concentrated in space, making the load on I/O nodes highly unbalanced. In this paper, we use the job logs and I/O traces collected from Tianhe-1A to quantitatively analyze the two causes of spatially bursty I/O, including uneven I/O traffic of job's processes and uneven distribution of job's nodes. Based on the analysis we propose a node allocation strategy that takes account of processes' different amounts of I/O traffic, so that the I/O traffic can be processed by more I/O nodes more evenly. Our evaluations on Tianhe-1A with synthetic benchmarks and realistic applications show that the proposed strategy can further exploit the potential of I/O forwarding layer and promote the I/O performance.
Jie Yu 0006, Xin Liu 0018, Wenrui Dong, Xiaoyong Li 0002
ICS1
2018 Cross-layer coordination in the I/O software stack of extreme-scale systems
abstract
Summary I/O forwarding layer has now become a standard storage layer in today's HPC systems in order to scale current storage systems to new levels of concurrency. With the deepening of storage hierarchy, I/O requests must traverse through several types of nodes to access required data, including compute nodes, I/O nodes, and storage nodes. It becomes difficult to control the data path and apply cross‐layer I/O optimization. In this paper, we propose a well coordinated I/O stack, which coordinates the data path between compute nodes and I/O nodes for better load balancing and data locality with a job‐level I/O node mapping, and coordinates data path between I/O nodes and storage nodes for lighter I/O interference. We implement and evaluate our ideas on Tianhe‐1A by leveraging an open‐source I/O forwarding layer named IOFSL. The experimental results show that our proposals can significantly accelerate I/O performance of multiple I/O kernels and real applications.
Jie Yu 0006, Xiaoyong Li 0002, Wenrui Dong
Concurr. Comput. Pract. Exp.1
2018 Erratum to: ONFS: a hierarchical hybrid file system based on memory, SSD, and HDD for high performance computers
abstract
In the original version of this article, the abbreviation ‘OWDM’ was incorrectly defined. The phrase ‘orthogonal wavelength division multiplexing’ should all be changed to ‘one-way wave depth migration’.
Xin Liu 0018, Yutong Lu, Jie Yu 0006, Jieting Wu, Ying Lu 0002
Frontiers Inf. Technol. Electron. Eng.3
2017 Efficient skyline computation over distributed interval data
abstract
Summary The increasing volume of uncertain data has resulted in a dire need for supporting efficient uncertain data management. The skyline query as an important aspect of data management has received considerable attention in recent years, because of its importance in making intelligent decisions over complex data. Moreover, data collection and storage have become increasingly distributed, which makes the central assembly of data for storage and query infeasible and inefficient. Although many research efforts have been conducted to address the skyline query problem in various distributed scenarios, we still lack algorithms to address the queries over interval data, which is a special kind of attribute‐level uncertain data that widely exists in many applications. In this paper, we extensively study the skyline query over distributed interval data. We model the skyline query problem and define the distributed skyline query over interval data. Particularly, 2 efficient algorithms are proposed to retrieve the skylines progressively from distributed local sites with a highly optimized feedback framework. Moreover, we exploit 2 strategies for further improving the queries. Extensive experiments on synthetic and real datasets with real deployment are conducted to validate the effectiveness and efficiency of our proposals.
Xiaoyong Li 0002, Kaijun Ren, Xiaoling Li 0002, Jie Yu 0006
Concurr. Comput. Pract. Exp.4
2017 ONFS: a hierarchical hybrid file system based on memory, SSD, and HDD for high performance computers
abstract
With supercomputers developing towards exascale, the number of compute cores increases dramatically, making more complex and larger-scale applications possible. The input/output (I/O) requirements of large-scale applications, workflow applications, and their checkpointing include substantial bandwidth and an extremely low latency, posing a serious challenge to high performance computing (HPC) storage systems. Current hard disk drive (HDD) based underlying storage systems are becoming more and more incompetent to meet the requirements of next-generation exascale supercomputers. To rise to the challenge, we propose a hierarchical hybrid storage system, on-line and near-line file system (ONFS). It leverages dynamic random access memory (DRAM) and solid state drive (SSD) in compute nodes, and HDD in storage servers to build a three-level storage system in a unified namespace. It supports portable operating system interface (POSIX) semantics, and provides high bandwidth, low latency, and huge storage capacity. In this paper, we present the technical details on distributed metadata management, the strategy of memory borrow and return, data consistency, parallel access control, and mechanisms guiding downward and upward migration in ONFS. We implement an ONFS prototype on the TH-1A supercomputer, and conduct experiments to test its I/O performance and scalability. The results show that the bandwidths of single-thread and multi-thread ‘read’/‘write’ are 6-fold and 5-fold better than HDD-based Lustre, respectively. The I/O bandwidth of data-intensive applications in ONFS can be 6.35 times that in Lustre.
Xin Liu 0018, Yutong Lu, Jie Yu 0006, Jieting Wu, Ying Lu 0002
Frontiers Inf. Technol. Electron. Eng.3
2015 Characterizing I/O workloads of HPC applications through online analysis
abstract
The performance of storage subsystem of super-computers can not meet the demands of complex applications running on them. One of its major causes is that the bandwidth of storage hardware has not been utilized efficiently due to the complex and changing application I/O behavior. Therefore, I/O characterization tools are vital to application development and orchestration of storage system. This paper proposes an I/O characterization tool called FTracer. It captures I/O traces and performs traces analysis at runtime. In order to provide more flexible analysis, this FTracer allows users to vary the analysis instances at runtime. This mechanism ensures users get what exactly they want about the I/O characteristics of their applications when applications are running. In this work, we characterize MADbench2 benchmark to demonstrate the ability of FTracer.
Wenrui Dong, Jie Yu 0006, You Zuo
IPCCC3