EDBT 2026 Demo / reviewers in the wild / expert
Yusheng Hua
dblp:282/7813
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0001-7017-4941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RuYi: Optimizing Burst Buffer Through Automated, Fine-Grained Process-to-BB MappingabstractCurrent supercomputers use an SSD-based storage layer called Burst Buffer (BB) to provide I/O-intensive applications with accelerated storage access. However, efficiently utilizing this limited and expensive storage remains a critical issue, creating an urgent need for implementing Quality of Service (QoS) in BB. To address this, we propose RuYi, a QoS-aware method to provide applications with bandwidth guarantees in the BB file system. RuYi tackles two main issues. First, it quantitatively profiles available bandwidth resources in BB to ensure reliable QoS, a crucial aspect seldom studied in the literature. Second, RuYi offers fine-grained process-level QoS via an innovative process-to-BB mapping, maximizing resource utilization—something not achievable with conventional coarse-grained compute-to-BB mapping. We evaluated RuYi on a subsystem of the leading exascale supercomputer Sunway, consisting of 4,000 compute nodes and 200 BB nodes. The experimental results demonstrate that RuYi achieves an impressive end-to-end bandwidth control accuracy of 97%, while improving BB utilization by up to 116% compared to conventional coarse-grained compute-to-BB mapping. Yusheng Hua, Xuanhua Shi, Ligang He, Teng Zhang 0001, Hai Jin 0001, Yong Chen 0001 |
IEEE Trans. Computers | 1 |
| 2024 | MMDataLoader: Reusing Preprocessed Data Among Concurrent Model Training TasksabstractData preprocessing plays an important role in deep learning, which directly affects the training efficiency. Data preprocessing is performed on the CPU. The preprocessed data are then fed to the models that are trained on the GPU. We observe that data preprocessing on the CPU can potentially create a bottleneck in the entire process of a model training task. In order to tackle this issue, we have developed MMDataLoader, which enables reusing preprocessed data among multiple model training tasks. MMDataLoader automatically constructs a data preprocessing pipeline based on each task's specific preprocessing workflow, allowing for maximum data reuse and reduced computing workload on the CPU. Unlike conventional data loaders that operate at the task level and provide data provision services to specific training tasks, MMDataLoader operates at the server level and provides data for all concurrently running tasks. We have conducted extensive experiments. The results show that MMDataLoader can significantly increase preprocessing throughput without affecting model convergence when compared to conventional methods where model training tasks are executed concurrently. For instance, with three tasks running, the preprocessing throughput can increase by 1.6x to 3.15x, depending on the tasks being executed and the proportion of preprocessing operations that are shared among them. Hai Jin 0001, Zhanyang Zhu, Ligang He, Yusheng Hua, Xuanhua Shi |
IEEE Trans. Computers | 5 |
| 2022 | LoomIO: Object-Level Coordination in Distributed File SystemsabstractDevice-level interference is recognized as a major cause of the performance degradation in distributed file systems. Although the approaches of mitigating interference through coordination at application-level, middleware-level, and server-level have shown beneficial results in previous studies, we find their effectiveness is largely reduced since I/O requests are re-arranged by underlying object file systems. In this research study, we prove that object-level coordination is critical and often the key to address the interference issue, as the scheduling of object requests determines the device-level accesses and thus determines the actual I/O bandwidth and latency. This article proposes an object-level coordination system, LoomIO, which uses an OBOP (One-Broadcast-One-Propagate) method and a time-limited coordination process to deliver highly efficient coordination service. Specifically, LoomIO enables object requests to achieve an optimized scheduling decision within a few milliseconds and largely mitigates the device-level interference. We have implemented a LoomIO prototye and integrated it into Ceph file system. The evaluation results show that LoomIO achieved the considerable improvements in resource utilization (by up to 35%), in I/O throughput (by up to 31%), and in 99th percentile latency (by up to 54%) compared to the K-optimal method which uses the same scheduling algorithm as LoomIO but does not have the coordination support. Yusheng Hua, Xuanhua Shi, Hai Jin 0001, Wei Xie 0017, Ligang He, Yong Chen 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | DDL-QoS: A dynamic I/O scheduling strategy of QoS for HPC applicationsabstractSummary With the increasing cloud‐trend of high‐performance computing (HPC), more users submit their applications simultaneously to the platform and wish they could finish before the deadline. Moreover, due to the severe holistic performance degradation caused by I/O contention, a deadline‐sensitive I/O scheduler is needed to allocate storage resources according to the requirements of applications and resultantly guarantee the quality of service (QoS) of concurrently running applications. In this paper, we first explore the bandwidth allocation phenomenon caused by interference in applications through the modeling of historical data, and then we quote a metric called random percentage that can represent the random degree of the applications and be used to guide I/O scheduling in the later stage. We design a dynamic I/O scheduler named DDL‐QoS that uses solid state drives(SSDs) as QoS guarantee to minimize interference and ensure applications meet their deadline. The potential of our design is that the greater the I/O interference, the greater the performance improvement, but this performance improvement will be limited by the physical properties of the storage hardware. Xuanhua Shi, Wei Liu 0004, Hai Jin 0001, Yusheng Hua |
Concurr. Comput. Pract. Exp. | 5 |
| 2019 | Software-defined QoS for I/O in exascale computing
Yusheng Hua, Xuanhua Shi, Hai Jin 0001, Wei Liu 0004, Yong Chen 0001, Ligang He |
CCF Trans. High Perform. Comput. | 1 |