VLDB 2026 Research / reviewers in the wild / expert
Yiqin Dai
dblp:225/0795
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-5629-1578ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A survey of anomaly detection in HPC systems using machine learningabstractAbstract High-performance computing (HPC) systems must remain stable and reliable to consistently deliver robust computational power and ensure the proper execution of user jobs. Anomaly detection is a key means to ensure the stability and reliability of these systems. With the expansion of HPC systems and changes in their architecture, accurately identifying anomalies in dynamic environments has become increasingly challenging. Traditional detection methods rely on experience and rules, which could be inefficient and inaccurate. To address these issues, researchers have proposed machine learning-based methods to automatically process large amounts of complex data, improving the efficiency of anomaly identification and diagnosis. In this survey, we conduct a comprehensive and in-depth investigation of machine learning-based anomaly detection methods in HPC systems. Firstly, we summarize and introduce the background and challenges of anomaly detection in HPC systems. Secondly, we compare a series of machine learning-based anomaly detection works in detail and summarize their frameworks. We conclude their advantages and disadvantages and application scenarios. Finally, we discuss several promising development trends of machine learning-based HPC system anomaly detection. Wei Zhang 0027, Yiqin Dai, Huijun Wu 0001, Zhenwei Wu, Hongyun Tian, Juan Chen 0001, Chubo Liu, Yong Dong |
CCF Trans. High Perform. Comput. | 4 |
| 2026 | A Survey on Machine Learning-Based HPC I/O Analysis and OptimizationabstractThe soaring computing power of HPC systems supports numerous large-scale applications, which generate massive data volumes and diverse I/O patterns, leading to severe I/O bottlenecks. Analyzing and optimizing HPC I/O is therefore critical. However, traditional approaches are typically customized and lack the adaptability required to cope with dynamic changes in HPC environments. To address the challenge, Machine Learning (ML) has been increasingly adopted to automate and enhance I/O analysis and optimization. Given sufficient I/O traces from HPC systems, ML can learn underlying I/O behaviors, extract actionable insights, and dynamically adapt to evolving workloads to improve performance. In this survey, we propose a novel taxonomy that aligns HPC I/O problems with learning tasks to systematically review existing studies. Through this taxonomy, we synthesize key findings on research distribution, data preparation, and model selection. Finally, we discuss several directions to advance the effective integration of ML in HPC I/O systems. Jingxian Peng, Huijun Wu 0001, Zhenwei Wu, Wei Zhang 0027, Yiqin Dai, Yong Dong |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2026 | Fully Decentralized Data Distribution for Large-Scale HPC SystemsabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In large-scale, especially exascale HPC systems, this mode of decoupling the demander and provider presents significant scalability limitations and incurs substantial costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Highest ranking and Longest consecutive piece segment First (HLF) policy based on the features of the HPC networking environment to improve the performance of FD3. In addition, we design a torrent-tree to accelerate the distribution of seed file data and the aggregation of distribution state, and release the tracker load with neighborhood local-generation algorithm. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 8-15 times. FD3 highlights the considerable potential of the all-to-all model in HPC data distribution scenarios. Furthermore, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Ruibo Wang, Mingtian Shao, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2025 | From Islands to Archipelago: Towards Collaborative and Adaptive Burst Buffer for HPC SystemsabstractModern supercomputers increasingly use node-local storage as burst buffers (BB) to address I/O bottlenecks.However, current BBs do not naturally support workflows, a common workload in HPC consisting of many interconnected subtasks.Workflow I/O can be divided into three types: intratask I/O, inter-task I/O, and stage-in/out I/O.While BBs can accelerate intra-task I/O, they often overlook the other two.Inter-task I/O relies on migrating data through the Parallel File System (PFS), which can slow down overall performance.Although allocating more resources to create larger BBs could help, it increases costs.Additionally, temporary BBs lack permanent storage, requiring data migration between the PFS and BB for stage-in and stage-out I/O.This process often involves multiple data copies and reduces I/O efficiency.Even for intra-task I/O, unbalanced data distribution can cause bottlenecks on heavily loaded nodes.To improve workflow acceleration in BB systems, it is important to address the needs of all the above-mentioned three Mingtian Shao, Ruibo Wang, Kai Lu 0001, Yiqin Dai, Huijun Wu 0001 |
ICS | 5 |
| 2025 | LingXi: An Architecture for COM/MON-Based High-Integrity TSN/TTE Switch
Pengye Xia, Weiliang Li, Yiqin Dai, Jiabo Zhang, Xuyan Jiang |
NPC (1) | 4 |
| 2025 | MIST: Towards MPI Instant Startup and Termination on Tianhe HPC SystemsabstractAs the size of MPI programs grows with expanding HPC resources and parallelism demands, the overhead of MPI startup and termination escalates due to the inclusion of less scalable global operations. Global operations involving extensive cross-machine communication and synchronization are crucial for ensuring semantic correctness. The current focus is on optimizing and accelerating these global operations rather than removing them, as the latter involves systematic changes to the system software stack and may impact program semantics. Given this background, we propose a systematic solution named MIST to safely eliminate global operations in MPI startup and termination. Through optimizing the generation of communication addresses, designing reliable communication protocols, and exploiting the resource release mechanism, MIST eliminates all global operations to achieve MPI instant startup and termination while ensuring correct program execution. Experiments on Tianhe-2A supercomputer demonstrate that MIST can reduce theMPI_Init()time by 32.5-77.6% and theMPI_Finalize()time by 28.9-85.0%. Yiqin Dai, Ruibo Wang, Yong Dong, Juan Chen 0001, Huijun Wu 0001, Mingtian Shao, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Fully Decentralized Data Distribution for Exascale-HPC: End of the Provider-Demander Matching PuzzleabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In the era of Exascale Computing, this mode of decoupling the demander and provider has limited scalability and huge costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Nearest and Longest consecutive piece Segment First (NLSF) policy based on the features of the HPC networking environment to improve the performance of FD3. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 7–11 times. FD3 shows the great potential of the all-to-all model in HPC data distribution scenarios. At the same time, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Mingtian Shao, Ruibo Wang, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
CLUSTER | 5 |
| 2024 | Towards Highly Compatible I/O-Aware Workflow Scheduling on HPC SystemsabstractScientific workflows on High-Performance Computing (HPC) consist of multiple data processing and computing tasks with dependencies. Efficiently scheduling computing resources and multi-tier storage across workflow tasks is crucial for optimizing performance. Existing solutions often fall short in achieving the co-scheduling of computing and 1/O resources and lack compatibility with HPC system software. In this paper, we introduce a performance model for scheduling workflows on HPC systems to enhance the understanding of workflow scheduling and facilitate the testing of the scheduling algorithm. Additionally, we propose THman, an open-source scientific workflow scheduler featuring our heuristic scheduling algorithm, Highest Contribution First (HCF). THman achieves online co-scheduling of computing and I/O resources for workflows and is designed to work with traditional HPC batch schedulers for high compatibility. We evaluate THman using simulated workloads and real-world workflow applications. Experimental results show that THman reduces workflow makespan by up to 30.9% compared to alternative methods. Yiqin Dai, Ruibo Wang, Yong Dong, Kai Lu 0001 |
SC | 1 |
| 2024 | Faster and Scalable MPI Applications LaunchingabstractDistributed parallel MPI applications are the dominant workload in many high-performance computing systems. While optimizing MPI application execution is a well-studied field, little work has considered optimizing the initial MPI application launching phase, which incurs extensive cross-machine communications and synchronization. The overhead of MPI application launching can be expensive, accounting for more than million core hours per 10K nodes annually on the production Tianhe-2A supercomputer, which will increase as the number of parallel machines used grows. Therefore, it is critical to optimize the MPI application launching process. This paper presents a novel approach to optimizing the MPI application launch. Our approach adopts a location-aware address generation rule to eliminate the need for address exchange and a topology-aware global communication scheme to optimize cross-machine synchronization. We then design a new application launch procedure to support the proposed optimizations to further reduce the pressure of the shared I/O system. Our techniques have been deployed to production in the Tianhe-2A supercomputer and the Next Generation Tianhe Supercomputer. Experimental results show that our approach scales well and outperforms alternative schemes, reducing the MPI application launching time by over 29% with 320K MPI processes. Yong Dong, Yiqin Dai, Kai Lu 0001, Ruibo Wang, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | The Fast and Scalable MPI Application Launch of the Tianhe HPC systemabstractFast and scalable MPI application launch helps achieve exascale performance and is becoming a common goal in high-performance computing. However, the traditional launch technique suffers from scalability deficiencies in the global information exchange and the global barrier operation. This drawback makes it challenging to launch MPI applications quickly in large-scale systems. In this paper, we propose a fast and scalable application launch technique and details its associated hardware and software support. The optimized launch technique includes a locality-aware static address generation rule for eliminating the need for address exchange and a topology-aware global communication scheme for improving global communication efficiency. We also propose an optimized application launch sequence for supporting the above launch technique. We implement and evaluate the proposed launch technique on the Tianhe-2A supercomputer and the Tianhe Exascale Prototype Upgrade System. Experimental results show that our technique can reduce the launch time by 26.1% when launching an application with 256K processes. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Mingtian Shao, Juan Chen 0001 |
IPDPS | 1 |
| 2022 | Few-Shot Non-Parametric Learning with Deep Latent Variable ModelabstractMost real-world problems that machine learning algorithms are expected to solve face the situation with (1) unknown data distribution; (2) little domain-specific knowledge; and (3) datasets with limited annotation. We propose Non-Parametric learning by Compression with Latent Variables (NPC-LV), a learning framework for any dataset with abundant unlabeled data but very few labeled ones. By only training a generative model in an unsupervised way, the framework utilizes the data distribution to build a compressor. Using a compressor-based distance metric derived from Kolmogorov complexity, together with few labeled data, NPC-LV classifies without further training. We show that NPC-LV outperforms supervised methods on all three datasets on image classification in the low data regime and even outperforms semi-supervised learning methods on CIFAR-10. We demonstrate how and when negative evidence lowerbound (nELBO) can be used as an approximate compressed length for classification. By revealing the correlation between compression rate and classification accuracy, we illustrate that under NPC-LV how the improvement of generative models can enhance downstream classification accuracy. Zhiying Jiang, Yiqin Dai, Ji Xin, Ming Li 0001, Jimmy Lin |
NeurIPS | 2 |
| 2022 | Towards Scalable Resource Management for SupercomputersabstractToday's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Wei Zhang 0027, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
SC | 1 |
| 2022 | TEES: topology-aware execution environment service for fast and agile application deployment in HPCabstractHigh-performance computing (HPC) systems are about to reach a new height: exascale. Application deployment is becoming an increasingly prominent problem. Container technology solves the problems of encapsulation and migration of applications and their execution environment. However, the container image is too large, and deploying the image to a large number of compute nodes is time-consuming. Although the peer-to-peer (P2P) approach brings higher transmission efficiency, it introduces larger network load. All of these issues lead to high startup latency of the application. To solve these problems, we propose the topology-aware execution environment service (TEES) for fast and agile application deployment on HPC systems. TEES creates a more lightweight execution environment for users, and uses a more efficient topology-aware P2P approach to reduce deployment time. Combined with a split-step transport and launch-in-advance mechanism, TEES reduces application startup latency. In the Tianhe HPC system, TEES realizes the deployment and startup of a typical application on 17 560 compute nodes within 3 s. Compared to container-based application deployment, the speed is increased by 12-fold, and the network load is reduced by 85%. Mingtian Shao, Kai Lu 0001, Wanqing Chi, Ruibo Wang, Yiqin Dai |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2019 | Goldilocks: Learning Pattern-Based Task Assignment in Mobile Crowdsensing
Jinghan Jiang, Yiqin Dai, Kui Wu 0001, Rong Zheng 0001 |
QSHINE | 2 |