Yujin Yan

dblp:259/5030 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 47% GPUs and heterogeneous computing · 43% Hardware accelerators and domain-specific architectures · 5%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
performance optimization at scale
0.612022
Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atoms · PPoPP 2022
High-performance computing
scientific computing systems
0.612022
Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atoms · PPoPP 2022
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous CPU-GPU cluster
0.512021
Optimizing the LINPACK Algorithm for Large-Scale PCIe-Based CPU-GPU Heterogeneous Systems · IEEE Trans. Parallel Distributed Syst. 2021
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems
0.412020
Revisiting linpack algorithm on large-scale CPU-GPU heterogeneous systems · PPoPP 2020
High-performance computing › numerical linear algebra
linpack
0.412020
Revisiting linpack algorithm on large-scale CPU-GPU heterogeneous systems · PPoPP 2020
Performance modeling and evaluation
benchmarking
0.112021
Optimizing the LINPACK Algorithm for Large-Scale PCIe-Based CPU-GPU Heterogeneous Systems · IEEE Trans. Parallel Distributed Syst. 2021

Methods — techniques the papers use, named apart from their topics

fine-grained pipelining · 0.9redundancy removal · 0.6model tabulation · 0.6kernel fusion · 0.6customized kernel acceleration · 0.6algorithmic modeling · 0.5
YearPublicationVenuePosition
2024 10-Million Atoms Simulation of First-Principle Package LS3DF
Yujin Yan, Haibo Li 0007, Lin-Wang Wang, Guangming Tan, Weile Jia, Ninghui Sun
J. Comput. Sci. Technol.1
2022 Routine mining on location sequences
abstract
In this article, we propose a novel routine pattern extraction architecture to analyze the daily behaviors of mobile users. The key component of the proposed architecture is a dynamic programming-based sequence dissimilarity calculation method, which aims to measure the dissimilarity between trajectories and then extract patterns using the clustering method. The method exploits three different information: (1) spatial-temporal information, (2) information in the continuous same location in the sequence and (3) the probabilities of location occurrences in the data. We conduct experiments on a synthetic dataset and two real-world datasets. The obtained results demonstrate that our proposed method is efficient in extracting hidden routine patterns from users’ trajectory data.
Yujin Yan, Alexandre Pauchet, Arnaud Knippel
KES1
2022 Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atoms
abstract
High-performance computing, together with a neural network model trained from data generated with first-principles methods, has greatly boosted applications of ab initio molecular dynamics in terms of spatial and temporal scales on modern supercomputers. Previous state-of-the-art can achieve 1 -- 2 nanoseconds molecular dynamics simulation per day for 100-million atoms on the entire Summit supercomputer. In this paper, we have significantly reduced the memory footprint and computational time by a comprehensive approach with both algorithmic and system innovations. The neural network model is compressed by model tabulation, kernel fusion, and redundancy removal. Then optimizations such as acceleration of customized kernel, tabulation of activation function, MPI+OpenMP parallelization are implemented on GPU and ARM architectures. Testing results of the copper system show that the optimized code can scale up to the entire machine of both Fugaku and Summit, and the corresponding system size can be extended by a factor of 134 to an unprecedented 17 billion atoms. The strong scaling of a 13.5-million atom copper system shows that the time-to-solution can be 7 times faster, reaching 11.2 nanoseconds per day. This work opens the door for unprecedentedly large-scale molecular dynamics simulations based on ab initio accuracy and can be potentially utilized in studying more realistic applications such as mechanical properties of metals, semiconductor devices, batteries, etc. The optimization techniques detailed in this paper also provide insight for relevant high-performance computing applications.
Zhuoqiang Guo, Denghui Lu, Yujin Yan, Siyu Hu, Guangming Tan, Ninghui Sun, Wanrun Jiang, Linfeng Zhang 0002, Mohan Chen 0002, Han Wang 0006, Weile Jia
PPoPP3
2021 Optimizing the LINPACK Algorithm for Large-Scale PCIe-Based CPU-GPU Heterogeneous Systems
abstract
There is a widening gap between GPU and other components (CPU, PCIe bus and communication network) in heterogeneous parallel system. The gap forces us to orchestrate cooperative execution among these components much more carefully than ever before. By taking the LINPACK benchmark as a case study, this article proposes a fine-grained pipelining algorithm on large-scale CPU-GPU heterogeneous cluster systems. First, we build an algorithmic model that reveals a new approach to GPU-centric and fine-grained pipelining algorithm design. Then, we present four model-driven pipelining algorithms that incrementally squeeze bubbles in the pipeline so that it is occupied by more useful floating-point calculations. The algorithms are implemented on both the AMD and NVIDIA GPU platforms. The finally optimized LINPACK program achieves 107 PFlops on 25, 600 GPUs (70 percent floating-point efficiency). Several insights have been drawn to suggest tradeoff of algorithm design, programming support, and architecture design.
Guangming Tan, Chaoyang Shui, Yinshan Wang, Xianzhi Yu, Yujin Yan
IEEE Trans. Parallel Distributed Syst.5
2020 Revisiting linpack algorithm on large-scale CPU-GPU heterogeneous systems
abstract
As the widening gap between GPU computing capability and other components (CPU, PCIe bus and communication network), it's increasingly challenging to design high performance parallel algorithms for large CPU-GPU heterogeneous systems. There are mainly two reasons. Firstly, simply offloading the kernel library to GPU incurs large volume data transfer through low-speed PCIe bus. Secondly, communication overheads through network severely affects scalability. To solve the above issues, we advocate a paradigm shift to CPU-centric and fine-grained pipelining algorithm design. By taking Linpack benchmark as a case study, the new algorithm design paradigm shows its effectiveness. Our optimized Linpack program achieves 63.79PFlops on 16384 GPUs. Its floating-point efficiency outperforms the NVIDIA proprietary counterparts by 5% on average.
Chaoyang Shui, Xianzhi Yu, Yujin Yan, Yinshan Wang, Guangming Tan
PPoPP3