Ji Hoon Kang 0002

dblp:71/10537-2 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0001-8797-8746ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2022 Efficient Task-Mapping of Parallel Applications Using a Space-Filling Curve
abstract
Improving the communication performance of parallel programs is an important but difficult problem in a large-scale distributed memory-based cluster. Efforts to improve parallel scalability often face severe huddles in managing communication overheads. This paper proposes a framework of a space-filing curve(SFC)-based task-remapping for communication intensive parallel applications. An SFC-based mapping, when applied for task-mapping of parallel applications preserves locality in terms of communications and produce a less fragmented task-mapping, reducing communication overheads. The framework also provides tools for performance analysis to see if the proposed task-mapping is appropriate for a given application running on a target system. It further develops a binary classifier as a predictor to decide whether or not to apply the proposed mapping before run-time. We evaluate the framework with three communication intensive applications in Cartesian coordinates: P3DFFT solver and Channel code using 2D domain decomposition model, and Poisson solver using 3D domain decomposition. The evaluation is conducted on a large-scale cluster system of fat-tree topology with up to 1,024 compute nodes. The proposed task-mapping achieves the overall performance improvement ranging from ~30% to ~66% over the baseline approach depending on the workloads. Also, when used in combination with the binary classifier-based predictor, it achieves the expected performance gains from 4% to 8%.
Oh-Kyoung Kwon, Ji Hoon Kang 0002, Seungchul Lee, Wonjung Kim 0002, Junehwa Song
PACT2
2022 Empirical Study on the GPU-accelerated HPL Performance: Effects of PCIe Communication
abstract
Computing resources equipped with GPU devices are popularly adopted in various application areas, and the HPL benchmark is used to evaluate their standard performance. The performance of GPU-accelerated HPL however cannot be free from PCIe communication even in the most optimal case. Here, we investigate the PCIe communication overhead and its effect on the NVIDIA HPL performance using a simple model. This preliminary study intends to derive a way that can estimate the HPL performance in upcoming computing resources that support full cache coherence between CPU and GPU memory.
Jieun Choi, Yosang Jeong, Ji Hoon Kang 0002, Gibeom Gu, Hoon Ryu
CLUSTER3
2022 Scalable implementation of multigrid methods using partial semi-aggregation of coarse grids
Ji Hoon Kang 0002
J. Supercomput.1
2021 High-performance simulations of turbulent boundary layer flow using Intel Xeon Phi many-core processors
Ji Hoon Kang 0002, Jinyul Hwang, Hyung Jin Sung, Hoon Ryu
J. Supercomput.1
2020 An HPC-based Prediction on the Practicality of Long-distance Quantum Key Distributions
abstract
Practicality of long-distance quantum communications based on a BB84 quantum key distribution (QKD) protocol is examined with Monte Carlo simulations coupled to a parallel computing. Given a quantum channel that is not free from noises, optimal sizes of the shared key information and corresponding chances for detecting the eavesdropper are calculated to present clues that can be utilized to figure out the utility of the protocol. Delivering simple but sound principles that have not been focused quite well, this work serves as a useful case study that shows the need of high performance computing for QKD modeling, and can trigger potential efforts for further modeling studies that involve more realistic and complicated factors.
Hoon Ryu, Ji Hoon Kang 0002
CLUSTER2
2017 Acceleration of Turbulent Flow Simulations with Intel Xeon Phi(TM) Manycore Processors
abstract
Enhancing the performance of turbulent flow simulations is important as the size of simulations grows with higher Reynolds number. We discuss the performance of our in-house turbulent flow simulation solver, named as DNS-TBL (Direct Numerical Simulation: Turbulent Boundary Layer), on the Intel Xeon Phi™ manycore processors. With bootable Knights Landing processors, the DNS-TBL solver shows excellent parallel scalability and, in particular, shows a 1.6 times better performance in solver time than the original CPU-based version. Current intermediate results serve as a practical case-study which directly shows how much turbulent flow simulations can be accelerated on manycore processors, providing a good reference for the parallelization and optimization schemes in the filed of computational fluid dynamics.
Ji Hoon Kang 0002, Hoon Ryu
CLUSTER1