EDBT 2026 Demo / reviewers in the wild / expert
Hyojun Son
dblp:249/4775
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-5199-9139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMabstractProcessing-in-memory (PIM), where compute is moved closer to memory or data, has been explored to accelerate emerging workloads. Different PIM-based systems have been announced, each offering a unique microarchitectural organization of their compute units, ranging from fixed functional units to programmable general-purpose compute cores near memory. However, one fundamental limitation of PIM is that each compute unit can only access its local memory; access to “remote” memory must occur through the host CPU - potentially limiting application performance scalability. In this work, we first characterize the scalability of real PIM architectures using the UPMEM PIM system. We analyze how the overhead of communicating through the host (instead of providing direct communication between the PIM compute units) can become a bottleneck for collective communications that are commonly used in many workloads. To overcome this inter-PIM bank communication, we propose PIMnet - a PIM interconnection network for PIM banks that provides direct connectivity between compute units and removes the overhead of communicating through the host. PIMnet exploits bandwidth parallelism where communication across the different PIM bank/chips can occur in parallel to maximize communication performance. PIMnet also matches the DRAM packaging hierarchy with a multi-tier network architecture. Unlike traditional interconnection networks, PIMnet is a PIMcontrolled network where communication is managed by the PIM logic, optimizing collective communications and minimizing the hardware overhead of PIMnet. Our evaluation of PIMnet shows that it provides up to $85 \times$ speedup on collective communications and achieves a $11.8 \times$ improvement on real applications compared to the baseline PIM. Hyojun Son, Gilbert Jonatan, Haeyoon Cho 0002, Kaustubh Shivdikar, José L. Abellán, Ajay Joshi, David R. Kaeli, John Kim 0001 |
HPCA | 1 |
| 2023 | Logical/Physical Topology-Aware Collective Communication in Deep Learning TrainingabstractTraining is an important aspect of deep learning to enable network models to be deployed. To scale training, multiple GPUs are commonly used with data parallelism to exploit the additional GPU compute and memory capacity. However, one challenge in scalability is the collective communication between GPUs. In this work, we propose to accelerate the AllReduce collective. AllReduce communication is often based on a logical topology (e.g., ring or tree algorithms) that is mapped to a physical topology or the physical connectivity between the nodes. In this work, we propose a logical/physical topology-aware collective communication that we refer to as C-Cube architecture – Chaining Collective Communication with Computation. C-Cube exploits the opportunity to overlap or chain different phases of collective communication as well as forward computation in a tree algorithm AllReduce. We exploit the communication pattern in a logical tree topology to overlap the different phases of communication. Since ordering is maintained in the tree collective algorithm, we propose gradient queuing to enable chaining of communication with forward computation to accelerate overall performance while having no impact on training accuracy. We also exploit the physical topology characteristics to further improve the performance, including proposing detour connections for collective communication while leveraging the additional connectivity to enable a double-tree C-Cube implementation. We implement a C-Cube proof-of-concept on a real system (8-GPU NVIDIA DGX-1) and show C-Cube results in performance improvement in communication performance compared to non-overlapped tree algorithms as well as overall performance. Jo Sanghoon, Hyojun Son, John Kim 0001 |
HPCA | 2 |
| 2021 | BoomGate: Deadlock Avoidance in Non-Minimal Routing for High-Radix NetworksabstractAvoiding routing deadlock is an important component of an interconnection network. For large-scale systems with high-radix topologies that leverage non-minimal adaptive routing, virtual channels (VCs) are commonly used to prevent routing deadlock. However, VCs in large-scale networks can be costly because of deep buffers and restrict VC usage. In this work, we propose BoomGATE for deadlock avoidance in large-scale networks. In particular, BoomGATE consists of two components - Restricted Intermediate-node Non-minimal Routing (RINR) algorithm and opportunistic flow control (OFC) which both exploit the low-diameter characteristics of high-radix networks while maximizing path diversity within the topology. We identify how routing deadlock in fully-connected topologies are caused by non-minimal routes and propose to restrict the non-minimal routing to ensure deadlock freedom without additional virtual channels. We also propose an algorithm that ensures path diversity is load-balanced across all nodes in the system. However, since path diversity is restricted with the RINR algorithm, complement RINR algorithm with opportunistic flow control (OFC) where “illegal routes” are allowed if and only if sufficient buffer can be guaranteed to ensure cyclical dependency does not occur. We propose both a static and dynamic OFC implementation. We evaluate the performance of BoomGATE and demonstrate there is minimal performance loss compared to global adaptive routing, while reducing the amount of buffers required by 50%. Gyuyoung Kwauk, Seungkwan Kang, Hans Kasan, Hyojun Son, John Kim 0001 |
HPCA | 4 |
| 2019 | Ghost routers: energy-efficient asymmetric multicore processors with symmetric NoCsabstractAsymmetric multicore architectures have been proposed to exploit the benefits of heterogeneous cores. However, asymmetric cores present challenge to network-on-chip (NoC) designers since the floorplan is not necessarily regular with "nodes" being different size. In contrast, most of the previously proposed NoC topologies commonly assume a regular or symmetric floorplan with equal size nodes. In this work, we first describe how asymmetric floorplan leads to asymmetric topology and can limit overall performance. To overcome the asymmetric floorplan, we present Ghost Routers - extra "dummy" routers that are added to the NoC to create a symmetric NoC architecture for asymmetric multicore architectures. Ghost router provides higher network path diversity and provides higher network performance that leads to higher system performance. Ghost routers also enable simpler routing algorithms because of the symmetric NoC architecture. While ghost routers is a simplistic modification to the NoC architecture, it does increase NoC cost. However, ghost routers exploit the observations that in realistic systems, the cost of NoC is not a significant fraction of overall system cost. Our evaluations show that ghost routers can improve performance by up to 21% while improving overall energy-efficiency of the system by up to 26%. Hyojun Son, Hanjoon Kim, Hao Wang 0011, Nam Sung Kim, John Kim 0001 |
NOCS | 1 |