VLDB 2026 Research / reviewers in the wild / expert
Mingtian Shao
dblp:255/4921
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-2368-4946ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fully Decentralized Data Distribution for Large-Scale HPC SystemsabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In large-scale, especially exascale HPC systems, this mode of decoupling the demander and provider presents significant scalability limitations and incurs substantial costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Highest ranking and Longest consecutive piece segment First (HLF) policy based on the features of the HPC networking environment to improve the performance of FD3. In addition, we design a torrent-tree to accelerate the distribution of seed file data and the aggregation of distribution state, and release the tracker load with neighborhood local-generation algorithm. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 8-15 times. FD3 highlights the considerable potential of the all-to-all model in HPC data distribution scenarios. Furthermore, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Ruibo Wang, Mingtian Shao, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | From Islands to Archipelago: Towards Collaborative and Adaptive Burst Buffer for HPC SystemsabstractModern supercomputers increasingly use node-local storage as burst buffers (BB) to address I/O bottlenecks.However, current BBs do not naturally support workflows, a common workload in HPC consisting of many interconnected subtasks.Workflow I/O can be divided into three types: intratask I/O, inter-task I/O, and stage-in/out I/O.While BBs can accelerate intra-task I/O, they often overlook the other two.Inter-task I/O relies on migrating data through the Parallel File System (PFS), which can slow down overall performance.Although allocating more resources to create larger BBs could help, it increases costs.Additionally, temporary BBs lack permanent storage, requiring data migration between the PFS and BB for stage-in and stage-out I/O.This process often involves multiple data copies and reduces I/O efficiency.Even for intra-task I/O, unbalanced data distribution can cause bottlenecks on heavily loaded nodes.To improve workflow acceleration in BB systems, it is important to address the needs of all the above-mentioned three Mingtian Shao, Ruibo Wang, Kai Lu 0001, Yiqin Dai, Huijun Wu 0001 |
ICS | 1 |
| 2025 | MIST: Towards MPI Instant Startup and Termination on Tianhe HPC SystemsabstractAs the size of MPI programs grows with expanding HPC resources and parallelism demands, the overhead of MPI startup and termination escalates due to the inclusion of less scalable global operations. Global operations involving extensive cross-machine communication and synchronization are crucial for ensuring semantic correctness. The current focus is on optimizing and accelerating these global operations rather than removing them, as the latter involves systematic changes to the system software stack and may impact program semantics. Given this background, we propose a systematic solution named MIST to safely eliminate global operations in MPI startup and termination. Through optimizing the generation of communication addresses, designing reliable communication protocols, and exploiting the resource release mechanism, MIST eliminates all global operations to achieve MPI instant startup and termination while ensuring correct program execution. Experiments on Tianhe-2A supercomputer demonstrate that MIST can reduce theMPI_Init()time by 32.5-77.6% and theMPI_Finalize()time by 28.9-85.0%. Yiqin Dai, Ruibo Wang, Yong Dong, Juan Chen 0001, Huijun Wu 0001, Mingtian Shao, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2024 | Fully Decentralized Data Distribution for Exascale-HPC: End of the Provider-Demander Matching PuzzleabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In the era of Exascale Computing, this mode of decoupling the demander and provider has limited scalability and huge costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Nearest and Longest consecutive piece Segment First (NLSF) policy based on the features of the HPC networking environment to improve the performance of FD3. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 7–11 times. FD3 shows the great potential of the all-to-all model in HPC data distribution scenarios. At the same time, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Mingtian Shao, Ruibo Wang, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
CLUSTER | 1 |
| 2024 | Faster and Scalable MPI Applications LaunchingabstractDistributed parallel MPI applications are the dominant workload in many high-performance computing systems. While optimizing MPI application execution is a well-studied field, little work has considered optimizing the initial MPI application launching phase, which incurs extensive cross-machine communications and synchronization. The overhead of MPI application launching can be expensive, accounting for more than million core hours per 10K nodes annually on the production Tianhe-2A supercomputer, which will increase as the number of parallel machines used grows. Therefore, it is critical to optimize the MPI application launching process. This paper presents a novel approach to optimizing the MPI application launch. Our approach adopts a location-aware address generation rule to eliminate the need for address exchange and a topology-aware global communication scheme to optimize cross-machine synchronization. We then design a new application launch procedure to support the proposed optimizations to further reduce the pressure of the shared I/O system. Our techniques have been deployed to production in the Tianhe-2A supercomputer and the Next Generation Tianhe Supercomputer. Experimental results show that our approach scales well and outperforms alternative schemes, reducing the MPI application launching time by over 29% with 320K MPI processes. Yong Dong, Yiqin Dai, Kai Lu 0001, Ruibo Wang, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | The Fast and Scalable MPI Application Launch of the Tianhe HPC systemabstractFast and scalable MPI application launch helps achieve exascale performance and is becoming a common goal in high-performance computing. However, the traditional launch technique suffers from scalability deficiencies in the global information exchange and the global barrier operation. This drawback makes it challenging to launch MPI applications quickly in large-scale systems. In this paper, we propose a fast and scalable application launch technique and details its associated hardware and software support. The optimized launch technique includes a locality-aware static address generation rule for eliminating the need for address exchange and a topology-aware global communication scheme for improving global communication efficiency. We also propose an optimized application launch sequence for supporting the above launch technique. We implement and evaluate the proposed launch technique on the Tianhe-2A supercomputer and the Tianhe Exascale Prototype Upgrade System. Experimental results show that our technique can reduce the launch time by 26.1% when launching an application with 256K processes. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Mingtian Shao, Juan Chen 0001 |
IPDPS | 6 |
| 2022 | Towards Scalable Resource Management for SupercomputersabstractToday's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Wei Zhang 0027, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
SC | 7 |
| 2022 | TEES: topology-aware execution environment service for fast and agile application deployment in HPCabstractHigh-performance computing (HPC) systems are about to reach a new height: exascale. Application deployment is becoming an increasingly prominent problem. Container technology solves the problems of encapsulation and migration of applications and their execution environment. However, the container image is too large, and deploying the image to a large number of compute nodes is time-consuming. Although the peer-to-peer (P2P) approach brings higher transmission efficiency, it introduces larger network load. All of these issues lead to high startup latency of the application. To solve these problems, we propose the topology-aware execution environment service (TEES) for fast and agile application deployment on HPC systems. TEES creates a more lightweight execution environment for users, and uses a more efficient topology-aware P2P approach to reduce deployment time. Combined with a split-step transport and launch-in-advance mechanism, TEES reduces application startup latency. In the Tianhe HPC system, TEES realizes the deployment and startup of a typical application on 17 560 compute nodes within 3 s. Compared to container-based application deployment, the speed is increased by 12-fold, and the network load is reduced by 85%. Mingtian Shao, Kai Lu 0001, Wanqing Chi, Ruibo Wang, Yiqin Dai |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2022 | Self-deployed execution environment for high performance computingabstractTraditional high performance computing (HPC) systems provide a standard preset environment to support scientific computation. However, HPC development needs to provide support for more and more diverse applications, such as artificial intelligence and big data. The standard preset environment can no longer meet these diverse requirements. If users still run these emerging applications on HPC systems, they need to manually maintain the specific dependencies (libraries, environment variables, and so on) of their applications. This increases the development and deployment burden for users. Moreover, the multi-user mode brings about privacy problems among users. Containers like Docker and Singularity can encapsulate the job’s execution environment, but in a highly customized HPC system, cross-environment application deployment of Docker and Singularity is limited. The introduction of container images also imposes a maintenance burden on system administrators. Facing the above-mentioned problems, in this paper we propose a self-deployed execution environment (SDEE) for HPC. SDEE combines the advantages of traditional virtualization and modern containers. SDEE provides an isolated and customizable environment (similar to a virtual machine) to the user. The user is the root user in this environment. The user develops and debugs the application and deploys its special dependencies in this environment. Then the user can load the job to compute nodes directly through the traditional HPC job management system. The job and its dependencies are analyzed, packaged, deployed, and executed automatically. This process enables transparent and rapid job deployment, which not only reduces the burden on users, but also protects user privacy. Experiments show that the overhead introduced by SDEE is negligible and lower than those of both Docker and Singularity. Mingtian Shao, Kai Lu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |