Kien Trung Pham

dblp:243/8803 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0001-8386-3917ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 A Bandwidth-Optimal All-to-All Communication in Two-Dimensional Fully Connected Network
abstract
High-radix direct interconnection networks enable many simultaneous, non-blocking point-to-point communications at every compute node of high-performance computing systems. This feature provides a novel opportunity to improve the performance of collective communications. In this study, we focus on two-dimensional fully connected network topologies (2dfc) which are basic blocks of cutting-edge high-radix network topologies such as Dragonfly and HyperX. We aim to design a bandwidth-optimal algorithm for the All-to-All collective communication in the 2dfc network. By using parallel point-to-point communications, the proposed algorithm enables parallel non-blocking intra- and inter-group communications with striping data placement. Simulation results using SimGrid illustrated that the proposed algorithm outperforms the conventional algorithms significantly. Compared with topology-aware algorithms, the proposed approach reduces 1.66× and 1.58× communication time of an existing and a topology-aware collective algorithm in a 16 × 16 2dfc network, respectively.
Kien Trung Pham, Truong Thao Nguyen, Michihiro Koibuchi
CCGrid1
2023 Effective switchless inter-FPGA memory networks
Truong Thao Nguyen, Kien Trung Pham, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi
J. Parallel Distributed Comput.2
2022 Scalable Low-Latency Inter-FPGA Networks
abstract
A cutting-edge FPGA card can be equipped with many high-bandwidth I/Os by means of high-density optical integration, e.g., onboard Si-photonics transceivers, to provide high network bandwidth for memory-to-memory inter-FPGA communication. This study presents its scalable switchless net-work architecture by exploiting an indirect path, consisting of two one-hop paths, for enabling a diameter-2 network topology. It then takes a Kautz network topology with a diameter of two for connecting d(d + 1) FPGAs with a degree of$d$, which is close to the theoretical upper bound. The Kautz network topologies have bi-directional links and uni-directional links which form triangles. Uni-directional links introduce difficulty in avoiding channel buffer overflow because the existing link-level flow control assumes a bi-directional link. This study presents an indirect flow control along a uni-directional triangle embedded in the Kautz network topology. It then develops a combination of unicasts that forms multi-port collective communications to mitigate the influence of the startup latency on the execution time. Since a high-degree FPGA card introduces difficulty in storing many I/O ports at the panel of a 1- U compute server, we propose using WDM (Wavelength Division Multiplexing) as an alternative and present its efficient mapping onto arrayed waveguide grating (AWG). The required number of wavelengths becomes d on d+ 1 AWG equipments. Based on our experimental results with OPTWEB of custom Stratix10 FPGA cards, SimGrid simulation results show that our collective communication is 7 × faster than that of Dragonfly with 272 FPGAs.
Kien Trung Pham, Truong Thao Nguyen, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi
IPDPS1
2019 A time-stamping system to detect memory consistency errors in MPI one-sided applications
Thanh-Dang Diep, Kien Trung Pham, Karl Fürlinger, Nam Thoai
Parallel Comput.2