EDBT 2026 Demo / reviewers in the wild / expert
Bin Yang 0043
dblp:77/377-43
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-3783-2228ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture
Qixin Chang, Xiaohui Duan, Huihai An, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Xiting Ju, Haopeng Huang, Wei Xue 0003, Lin Gan 0008, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu, Ren Hu |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2025 | Efficiency Optimization Under Spatiotemporal Sharing Fairness for Deep Learning Workloads in Heterogeneous GPU ClustersabstractModern GPU clusters increasingly comprise diverse heterogeneous GPUs, driven by the continuous release of new GPU models. Achieving a balance between fairness and efficiency when scheduling multi-tenant Deep Learning (DL) training jobs on such clusters is inherently challenging. Existing DL training schedulers largely emphasize fairness through GPU temporal sharing, while the spatial dimension of resource allocation is often underexplored. This oversight can lead to GPU fragmentation and suboptimal system performance. In this paper, we propose STS-Fairness, a spatiotemporal sharing fairness scheduler. STS-Fairness partitions each GPU into multiple isolated slots under a novel spatiotemporal fairness constraint and allocates jobs using a round-based allocation mechanism. We guarantee that STS-Fairness achieves overall performance optimality while satisfying spatiotemporal fairness constraints. The scheduling problem is formulated as an integer nonlinear program (INLP) that is solved to optimality in polynomial time via dynamic programming. We deployed the STS-Fairness framework on both physical and simulated heterogeneous clusters and conducted large-scale experiments. These results demonstrate that STS-Fairness reduces average JCT by$1.2 \times$, shortens makespan by$1.24 \times$, and increases throughput by$\mathbf{1. 2 5} \times$compared to state-of-the-art (SoTA) schedulers. Chunhong Du, Mengyu Shi, Shanjiang Tang, Jianhang Tang, Ce Yu, Jian Xiao 0001, Chao Sun 0008, Bin Yang 0043 |
ICPADS | 8 |
| 2025 | MixLoRA: An Efficient Multi-Tenant Framework for Concurrently Serving Diverse LoRA Models in Large Language ModelsabstractThe rapid advancement of large language models (LLMs) has driven the widespread adoption of Low-Rank Adaptation (LoRA) for efficient fine-tuning. However, existing multi-LoRA inference systems, such as Punica, face significant challenges in GPU utilization when handling concurrent requests with diverse ranks, leading to suboptimal throughput and increased latency. To address these limitations, we propose MixLoRA, a novel CUDA stream-based framework that enhances GPU efficiency by enabling parallel execution of mixed-rank LoRA requests. MixLoRA introduces a multi-worker scheduling architecture that selects the optimal number of workers for concurrent execution based on the model size, inference request characteristics, and device information. Each worker manages requests of a fixed rank, eliminating the need for rank-aligned batching and maximizing resource utilization through independent CUDA streams. Experimental results demonstrate that MixLoRA achieves up to 1.2× higher throughput compared to state-of-the-art solutions, offering a scalable and efficient approach for multi-tenant LLM inference services. Ronghuai Chen, Ce Yu, Hao Fu 0021, Xiaoteng Hu, Bin Yang 0043 |
ICPP | 5 |
| 2025 | An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million CoresabstractGlobal Storm Resolving Models (GSRMs) is crucial for understanding extreme weather events under the climate change background. In this study, we optimize Global-Regional Integrated Forecast System (GRIST), which is a unified weather-climate modeling system designed for research and operation, for the next-generation Sunway supercomputer, incorporating AI-enhanced physics suite, OpenMP-based parallelization, and mixed-precision optimizations to enhance both efficiency and performance portability, as well as the unified modeling capability. Our experiments successfully capture significant events during the "23.7" extreme rainfall over northern China influenced by super Typhoon Doksuri, at 1km resolution. Notably, our work scales to 34 million cores, enabling simulation speeds at 491 SDPD (3km) and 181 SDPD (1km). Xiaohui Duan, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Huihai An, Xiting Ju, Haopeng Huang, Wei Xue 0003, Jianye Hou, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu |
PPoPP | 5 |
| 2025 | Task Scheduling in Geo-Distributed Computing: A SurveyabstractGeo-distributed computing, a paradigm that assigns computational tasks to globally distributed nodes, has emerged as a promising approach in cloud computing, edge computing, cloud-edge computing, and supercomputer computing (SC). It enables low-latency services, ensures data locality, and handles large-scale applications. As global computing capacity and task demands increase rapidly, scheduling tasks for efficient execution in geo-distributed computing systems has become an increasingly critical research challenge. It arises from the inherent characteristics of geographic distribution, including heterogeneous network conditions, region-specific resource pricing, and varying computational capabilities across locations. Researchers have developed diverse task scheduling methods tailored to geo-distributed scenarios, aiming to achieve objectives such as performance enhancement, fairness assurance, and fault-tolerance improvement. This survey provides a comprehensive and systematic review of task scheduling techniques across four major distributed computing environments, with an in-depth analysis of these approaches based on their core scheduling objectives. Through our analysis, we identify key research challenges and outline promising directions for advancing task scheduling in geo-distributed computing. Yujian Wu, Shanjiang Tang, Ce Yu, Bin Yang 0043, Chao Sun 0008, Jian Xiao 0001, Hutong Wu, Jinghua Feng |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Full Lifecycle Data Analysis on a Large-scale and Leadership Supercomputer: What Can We Learn from It?
Bin Yang 0043, Wei Xue 0003 |
USENIX ATC | 1 |
| 2023 | HadaFS: A File System Bridging the Local and Shared Burst Buffer for Exascale Supercomputers
Xiaobin He, Bin Yang 0043, Shupeng Shi, Dexun Chen, Wei Xue 0003, Zuoning Chen |
FAST | 2 |
| 2023 | Rapid simulations of atmospheric data assimilation of hourly-scale phenomena with modern neural networksabstractAtmospheric data assimilation is essential for numerical weather prediction. Ensemble data assimilation connects multiple instances of an atmospheric model through a Kalman filter-based algorithm, which is regarded as a challenging computing task today. In this work, we build a fast, low-cost, and scalable atmospheric data assimilation prototype, DIDA, for the new-generation Sunway supercomputer, including: (1) a framework that enables flexible deployment of components, and manages and optimizes data communication among modules, achieving maximum resource efficiency; (2) an accurate, robust, UNet-based surrogate model for atmospheric dynamic simulation to generate the background ensemble; (3) a batch-LETKF algorithm with high-performance eigenvalue decomposition, which is up to 7.37 times faster than existing numerical libraries while exhibiting almost linear scalability. Experimental evaluations show that our AI-integrated ensemble data assimilation prototype can complete hour-cycle assimilation in minutes, maintain linear scalability, and save an order of magnitude of computing resources, compared with the traditional method. Yiyuan Li, Xiting Ju, Qilong Jia, Yongxiao Zhou, Simeng Qian, Rongfen Lin, Bin Yang 0043, Shupeng Shi, Xin Liu 0081, Jian Tan 0005, Zhengding Hu, Limin Yan, Wei Xue 0003 |
SC | 8 |
| 2023 | End-to-end I/O Monitoring on Leading SupercomputersabstractThis paper offers a solution to overcome the complexities of production system I/O performance monitoring. We present Beacon, an end-to-end I/O resource monitoring and diagnosis system for the 40960-node Sunway TaihuLight supercomputer, currently the fourth-ranked supercomputer in the world. Beacon simultaneously collects and correlates I/O tracing/profiling data from all the compute nodes, forwarding nodes, storage nodes, and metadata servers. With mechanisms such as aggressive online and offline trace compression and distributed caching/storage, it delivers scalable, low-overhead, and sustainable I/O diagnosis under production use. With Beacon’s deployment on TaihuLight for more than three years, we demonstrate Beacon’s effectiveness with real-world use cases for I/O performance issue identification and diagnosis. It has already successfully helped center administrators identify obscure design or configuration flaws, system anomaly occurrences, I/O performance interference, and resource under- or over-provisioning problems. Several of the exposed problems have already been fixed, with others being currently addressed. Encouraged by Beacon’s success in I/O monitoring, we extend it to monitor interconnection networks, which is another contention point on supercomputers. In addition, we demonstrate Beacon’s generality by extending it to other supercomputers. Both Beacon codes and part of collected monitoring data are released. 1 Bin Yang 0043, Wei Xue 0003, Shichao Liu 0004, Xiaosong Ma, Xiyang Wang 0003 |
ACM Trans. Storage | 1 |
| 2022 | DFBuffer: High-performance data forwarding software optimized for single-process I/O scenariosabstractMost supercomputers adopt a data forwarding architecture to achieve storage scalability. However, it results in a significant reduction in single-process bandwidth compared to direct file system access. Moreover, considering that a majority of applications uses only a single process for writing and reading data, the low single-process performance also leads to a time overhead for these applications. This paper proposes an userspace forwarding mechanism DFBUFFER with two performance optimization methods: user-space multi-thread request processing and data write buffer in a unit of file. The client of DFBUFFER is embedded in the application as a library reducing the software overhead, and the server implements multi-thread I/O request processing to improve bandwidth efficiency. The data write buffer can asynchronously handle write requests, which accelerates the write bandwidth of compute nodes. We evaluate DFBUFFER on the Sunway exascale prototype system. The results indicate that in the regular mode of DFBUFFER, both the write and read latency are reduced, and the write bandwidth and large-block read bandwidth of single-process are increased by 1.8 times and 2.8 times respectively. The DFBUFFER buffer mode increases the write bandwidth of a single process by 0.8 times over the regular mode. Although the performance advantage of the regular mode of DFBUFFER gradually weakens with the increase of concurrent processes, the DFBUFFER buffer mode has the effect of improving the write bandwidth, the 64-IO-processes application is increased by 0.2 times. Xiaobin He, Bin Yang 0043, Zuoning Chen |
ICPADS | 5 |
| 2022 | An End-to-end and Adaptive I/O Optimization Tool for Modern HPC Storage SystemsabstractReal-world large-scale applications expose more and more pressures to storage services of modern supercomputers. Supercomputers have been introducing new storage devices and technologies to meet the performance requirements of various applications, leading to more complicated architectures. High I/O demand of applications and the complicated and shared storage architectures make the issues, such as unbalanced load, I/O interference, system parameter configuration error, and node performance degradation, more frequently observed. And it is challenging to both achieve high I/O performance on application level and efficiently utilize scarce storage resources. We propose AIOT, an end-to-end and adaptive I/O optimization tool for HPC storage systems, which introduces effective I/O performance modeling and several active tuning strategies to improve both the I/O performance of applications and the utilization of storage resources. AIOT provides a global view of the whole storage system and searches for the optimal end-to-end I/O path through flow network modeling. Moreover, AIOT tunes system parameters across multiple layers of the storage system by using the automated identified application I/O behaviors and the instant status of the workload of storage system. We verified the effectiveness of AIOT for balancing I/O load, resolving I/O interference, improving I/O performance by configuring appropriate system parameters, and avoiding I/O performance degradation caused by abnormal nodes through quite a few real-world cases. AIOT has helped to save over ten millions of core-hours during the deployment on Sunway TaihuLight since July 2021. It's worth mentioning that our proposed AIOT is capable of managing other I/O optimization methods across various storage platforms. Bin Yang 0043, Yanliang Zou, Wei Xue 0003 |
IPDPS | 1 |
| 2020 | Lessons Learned from Optimizing the Sunway Storage System for Higher Application I/O Performance
Zuoning Chen, Wei Xue 0003, Bin Yang 0043 |
J. Comput. Sci. Technol. | 6 |
| 2019 | Automatic, Application-Aware I/O Forwarding Resource Allocation
Bin Yang 0043, Xiaosong Ma, Xiupeng Zhu, Xiyang Wang 0003, Nosayba El-Sayed, Jidong Zhai, Wei Xue 0003 |
FAST | 2 |
| 2019 | End-to-end I/O Monitoring on a Leading Supercomputer
Bin Yang 0043, Xiaosong Ma, Xiyang Wang 0003, Xiupeng Zhu, Nosayba El-Sayed, Haidong Lan, Jidong Zhai, Wei Xue 0003 |
NSDI | 1 |