Jiyun Jeong

dblp:157/2798 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Interconnection networks and networks-on-chip · 42% Memory systems · 35% GPUs and heterogeneous computing · 15%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip › routing algorithms
adaptive routing
0.212016
Contention-based congestion management in large-scale networks · MICRO 2016
Interconnection networks and networks-on-chip
congestion control
0.212016
Contention-based congestion management in large-scale networks · MICRO 2016
Memory systems
memory management
0.212014
Multi-GPU System Design with Memory Networks · MICRO 2014
Memory systems › memory disaggregation
memory network
0.212014
Multi-GPU System Design with Memory Networks · MICRO 2014
GPUs and heterogeneous computing
multi-GPU computing
0.212014
Multi-GPU System Design with Memory Networks · MICRO 2014
Interconnection networks and networks-on-chip
network topology
0.212014
Multi-GPU System Design with Memory Networks · MICRO 2014
Memory systems › virtual memory management
unified memory
0.212014
Multi-GPU System Design with Memory Networks · MICRO 2014
Distributed systems
large-scale network
0.112016
Contention-based congestion management in large-scale networks · MICRO 2016
GPUs and heterogeneous computing
GPU memory access
0.112014
Multi-GPU System Design with Memory Networks · MICRO 2014
Distributed systems › distributed communication
remote memory access
0.112014
Multi-GPU System Design with Memory Networks · MICRO 2014

Methods — techniques the papers use, named apart from their topics

virtual channels · 0.2throttling · 0.2overlay network · 0.2hybrid memory cubes · 0.2
YearPublicationVenuePosition
2016 Automatically Exploiting Implicit Pipeline Parallelism from Multiple Dependent Kernels for GPUs
abstract
Execution of GPGPU workloads consists of different stages including data I/O on the CPU, memory copy between the CPU and GPU, and kernel execution. While GPU can remain idle during I/O and memory copy, prior work has shown that overlapping data movement (I/O and memory copies) with kernel execution can improve performance. However, when there are multiple dependent kernels, the execution of the kernels is serialized and the benefit of overlapping data movement can be limited. In order to improve the performance of workloads that have multiple dependent kernels, we propose to automatically overlap the execution of kernels by exploiting implicit pipeline parallelism. We first propose Coarse-grained Reference Counting-based Scoreboarding (CRCS) to guarantee correctness during overlapped execution of multiple kernels. However, CRCS alone does not necessarily improve overall performance if the thread blocks (or CTAs) are scheduled sequentially. Thus, we propose an alternative CTA scheduler -- Pipeline Parallelism-aware CTA Scheduler (PPCS) that takes available pipeline parallelism into account in CTA scheduling to maximize pipeline parallelism and improve overall performance. Our evaluation results show that the proposed mechanisms can improve performance by up to 67% (33% on average). To the best of our knowledge, this is one of the first work that enables overlapped execution of multiple dependent kernels without any kernel modification or explicitly expressing dependency by the programmer.
Gwangsun Kim, Jiyun Jeong, John Kim 0001, Mark Stephenson
PACT2
2016 Contention-based congestion management in large-scale networks
abstract
Global adaptive routing exploits non-minimal paths to improve performance on adversarial traffic patterns and load-balance network channels in large-scale networks. However, most prior work on global adaptive routing have assumed admissible traffic pattern where no endpoint node is oversubscribed. In the presence of a greedy flow or hotspot traffic, we show how exploiting path diversity with global adaptive routing can spread network congestion and degrade performance. When global adaptive routing is combined with congestion management, the two types of congestion - network congestion that occurs within the interconnection network channels and endpoint congestion that occurs from oversubscribed endpoint nodes - are not properly differentiated. As a result, previously proposed congestion management mechanisms that are effective in addressing endpoint congestion are not necessarily effective when global adaptive routing is also used in the network. Thus, we propose a novel, low-cost contention-based congestion management (CBCM) to identify endpoint congestion based on the contention within the intermediate routers and at the endpoint nodes. While contention also occurs for network congestion, the endpoint nodes or the destination determines whether the congestion is endpoint congestion or network congestion. If it is only network congestion, CBCM ignores the network congestion and adaptive routing is allowed to minimize network congestion. However, if endpoint congestion occurs, CBCM throttles the hotspot senders and minimally route the traffic through a separate VC. Our evaluation across different traffic patterns and network sizes demonstrates that our approach is more robust in identifying endpoint congestion in the network while complementing global adaptive routing to avoid network congestion.
Gwangsun Kim, Jiyun Jeong, Mike Parker, John Kim 0001
MICRO3
2014 Multi-GPU System Design with Memory Networks
abstract
GPUs are being widely used to accelerate different workloads and multi-GPU systems can provide higher performance with multiple discrete GPUs interconnected together. However, there are two main communication bottlenecks in multi-GPU systems -- accessing remote GPU memory and the communication between GPU and the host CPU. Recent advances in multi-GPU programming, including unified virtual addressing and unified memory from NVIDIA, has made programming simpler but the costly remote memory access still makes multi-GPU programming difficult. In order to overcome the communication limitations, we propose to leverage the memory network based on hybrid memory cubes (HMCs) to simplify multi-GPU memory management and improve programmability. In particular, we propose scalable kernel execution (SKE) where multiple GPUs are viewed as a single virtual GPU as a single kernel can be executed across multiple GPUs without modifying the source code. To fully enable the benefits of SKE, we explore alternative memory network designs in a multi-GPU system. We propose a GPU memory network (GMN) to simplify data sharing between the discrete GPUs while a CPU memory network (CMN) is used to simplify data communication between the host CPU and the discrete GPUs. These two types of networks can be combined to create a unified memory network (UMN) where the communication bottleneck in multi-GPU can be significantly minimized as both the CPU and GPU share the memory network. We evaluate alternative network designs and propose a sliced flattened butterfly topology for the memory network that scales better than previously proposed alternative topologies by removing local HMC channels. In addition, we propose an overlay network organization for unified memory network to minimize the latency for CPU access while providing high bandwidth for the GPUs. We evaluate trade-offs between the different memory network organization and show how UMN significantly reduces the communication bottleneck in multi-GPU systems.
Gwangsun Kim, Minseok Lee, Jiyun Jeong, John Kim 0001
MICRO3