Wanwang Yin

dblp:202/6644 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
6since 2021 · last 2023
0000-0003-3680-2681ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 since 2021
YearPublicationVenuePosition
2023 5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway Supercomputer
abstract
HPL-MxP is an emerging high performance benchmark used to measure the mixed-precision computing capability of leading supercomputers. In this work, we present our efforts on the new Sunway that linearly scales the benchmark to over 40 million cores, sustains an overall mixed-precision performance exceeding 5 ExaFlop/s, and achieves over 85% of peak performance, which is the highest efficiency reached among all heterogeneous systems on the HPL-MxP list. The optimizations of our HPL-MxP implementation include the following: (1) a Two-Direction Look-Ahead and Overlap algorithm that enables overlaps of all communications with computation; (2) a multi-level process-mapping and communication scheduling method that uses the entire network as best as possible while maintaining conflict-free algorithm-flow; and (3) a CG-Fusion computing framework that eliminates up to 60% of inter-chip communications and removes the memory access bottleneck while serving both computation and communication simultaneously. This work could also provide useful insights for tuning cutting-edge applications on Sunway supercomputers as well as other heterogeneous supercomputers.
Rongfen Lin, Xinhui Yuan, Wei Xue 0003, Wanwang Yin, Jienan Yao, Junda Shi, Chaobo Song, Fei Wang 0096
SC4
2023 xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002
CCF Trans. High Perform. Comput.7
2023 Publisher Correction: xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002
CCF Trans. High Perform. Comput.7
2022 Scaling graph traversal to 281 trillion edges with 40 million cores
abstract
Graph processing, especially high-performance graph traversal, plays a more and more important role in data analytics. The successor of Sunway TaihuLight, New Sunway, is equipped with nearly 10 PB memory and over 40 million cores, which brings the opportunity to process hundreds of trillions of edges graphs. However, the graph with an unprecedented scale also brings severe performance challenges, including load imbalance, poor locality, and irregular access of graph traversal workload.
Huanqi Cao, Yuanwei Wang, Haojie Wang 0004, Heng Lin, Zixuan Ma, Wanwang Yin
PPoPP6
2022 Scaling Graph 500 SSSP to 140 Trillion Edges with over 40 Million Cores
abstract
The SSSP kernel was first introduced into the Graph 500 benchmark in 2017. However, there has been no result from a full-scale world-top supercomputer. The primary reason is the poor work-inefficiency of existing algorithms at large scales. In this paper, we propose an SSSP implementation for The Newest Generation Sunway Supercomputer,including an SSSP algorithm to achieve work-efficiency, along with an adaptive dense/sparse-mode selection approach to achieve communication-efficiency. Our implementation reaches 7638 GTEPS, with 103158 processors (over 40 million cores), and achieves 3.7× in performance and 512× in graph size compared with the current top one on the Graph 500 SSSP list. Based on our experience of running extreme-scale SSSP, we uncover the root cause of its poor scalability: the weight distribution allows edges with weights close to zero, making the SSSP tree deeper on larger graphs. We further explore a scalability-friendly weight distribution by setting a non-zero lower bound to the edge weights.
Yuanwei Wang, Huanqi Cao, Zixuan Ma, Wanwang Yin
SC4
2021 Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous cores
abstract
We study and evaluate performance optimization techniques for the HPCG benchmark on the newest generation Sunway supercomputer. Specifically, a two-level blocking scheme is proposed to expose adequate parallelism in the symmetric Gauss-Seidel kernel while keeping a fast convergence rate, a fine-grained kernel fusion technique is developed to alleviate the bandwidth load on local storage with small capacity, and a low overhead thread collaboration method is presented to efficiently move data between threads and hide its cost with data transfer operations. Test results show that the optimized HPCG code is able to exploit 73.0% of the theoretical memory bandwidth, and scale to over 42 million heterogeneous cores with 95.5% weak-scaling efficiency and 5.91 Pflop/s performance. We also study how the performance can be improved if the specific rules of HPCG are not fully obeyed, and design dependency preserving parallelization and vectorization methods, further boosting performance to 27.6 Pflop/s.
Qianchao Zhu, Hao Luo 0015, Chao Yang 0002, Mingshuo Ding, Wanwang Yin, Xinhui Yuan
SC5
2018 Performance Optimization of the HPCG Benchmark on the Sunway TaihuLight Supercomputer
abstract
In this article, we present some key techniques for optimizing HPCG on Sunway TaihuLight and demonstrate how to achieve high performance in memory-bound applications by exploiting specific characteristics of the hardware architecture. In particular, we utilize a block multicoloring approach for parallelization and propose methods such as requirement-based data mapping and customized gather collective to enhance the effective memory bandwidth. Experiments indicate that the optimized HPCG code can sustain 77% of the theoretical memory bandwidth and scale to the full system of more than 10 million cores, with an aggregated performance of 480.8 Tflop/s and a weak scaling efficiency of 87.3%.
Yulong Ao, Chao Yang 0002, Fangfang Liu 0004, Wanwang Yin, Lijuan Jiang, Qiao Sun 0005
ACM Trans. Archit. Code Optim.4
2017 Towards Highly Efficient DGEMM on the Emerging SW26010 Many-Core Processor
abstract
The matrix-matrix multiplication is an essential building block that can be found in various scientific and engineering applications. High-performance implementations of the matrix-matrix multiplication on state-of-the-art processors may be of great importance for both the vendors and the users. In this paper, we present a detailed methodology of implementing and optimizing the double-precision general format matrix-matrix multiplication (DGEMM) kernel on the emerging SW26010 processor, which is used to build the Sunway TaihuLight supercomputer. We propose a three level blocking algorithm to orchestrate data on the memory hierarchy and expose parallelism on different hardware levels, and design a collective data sharing scheme by using the register communication mechanism to exchange data efficiently among different cores. On top of those, further optimizations are done based on a data-thread mapping method for efficient data distribution, a double buffering scheme for asynchronous DMA data transfer, and an instruction scheduling method for maximizing the pipeline usage. Experiment results show that the proposed DGEMM implementation can fully exploit the unique hardware features provided by SW26010 and can sustain up to 95% of the peak performance.
Lijuan Jiang, Chao Yang 0002, Yulong Ao, Wanwang Yin, Wenjing Ma, Qiao Sun 0005, Fangfang Liu 0004, Rongfen Lin
ICPP4
2017 Scalable Graph Traversal on Sunway TaihuLight with Ten Million Cores
abstract
Interest has recently grown in efficiently analyzing unstructured data such as social network graphs and protein structures. A fundamental graph algorithm for doing such task is the Breadth-First Search (BFS) algorithm, the foundation for many other important graph algorithms such as calculating the shortest path or finding the maximum flow in graphs. In this paper, we share our experience of designing and implementing the BFS algorithm on Sunway TaihuLight, a newly released machine with 40,960 nodes and 10.6 million accelerator cores. It tops the Top500 list of June 2016 with a 93.01 petaflops Linpack performance [1]. Designed for extremely large-scale computation and power efficiency, processors on Sunway TaihuLight employ a unique heterogeneous many-core architecture and memory hierarchy. With its extremely large size, the machine provides both opportunities and challenges for implementing high-performance irregular algorithms, such as BFS. We propose several techniques, including pipelined module mapping, contention-free data shuffling, and group-based message batching, to address the challenges of efficiently utilizing the features of this large scale heterogeneous machine. We ultimately achieved 23755.7 giga-traversed edges per second (GTEPS), which is the best among heterogeneous machines and the second overall in the Graph500s June 2016 list [2].
Heng Lin, Xiongchao Tang, Bowen Yu 0003, Youwei Zhuo, Jidong Zhai, Wanwang Yin
IPDPS7
2017 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenarios
abstract
This paper reports our large-scale nonlinear earthquake simulation software on Sunway TaihuLight. Our innovations include: (1) a customized parallelization scheme that employs the 10 million cores efficiently at both the process and the thread levels; (2) an elaborate memory scheme that integrates on-chip halo exchange through register communcation, optimized blocking configuration guided by an analytic model, and coalesced DMA access with array fusion; (3) on-the-fly compression that doubles the maximum problem size and further improves the performance by 24%. With these innovations to remove the memory constraints of Sunway TaihuLight, our software achieves over 15% of the system's peak, better than the 11.8% efficiency achieved by a similar software running on Titan, whose byte to flop ratio is 5 times better than TaihuLight. The extreme cases demonstrate a sustained performance of over 18.9 Pflops, enabling the simulation of Tangshan earthquake as an 18-Hz scenario with an 8-meter resolution.
Haohuan Fu, Conghui He, Bingwei Chen, Zekun Yin, Tingjian Zhang, Wei Xue 0003, Wanwang Yin, Guangwen Yang 0002
SC10