Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jason Jong Kyu Park

dblp:124/7047 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
0since 2021 · last 2017
0000-0002-5457-219XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
GPUs and heterogeneous computing · 68% Cloud and datacenter computing · 12% Memory systems · 9%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU sharing
0.522017
Dynamic Resource Management for Efficient Utilization of Multitasking GPUs · ASPLOS 2017
Chimera: Collaborative Preemption for Multitasking on a Shared GPU · ASPLOS 2015
Cloud and datacenter computing › resource management
dynamic resource management
0.312017
Dynamic Resource Management for Efficient Utilization of Multitasking GPUs · ASPLOS 2017
GPUs and heterogeneous computing › GPU sharing
spatial multitasking
0.312017
Dynamic Resource Management for Efficient Utilization of Multitasking GPUs · ASPLOS 2017
GPUs and heterogeneous computing › GPU sharing
GPU preemption
0.212015
Chimera: Collaborative Preemption for Multitasking on a Shared GPU · ASPLOS 2015
Memory systems › memory access optimization
memory-level parallelism
0.212015
ELF: maximizing memory-level parallelism for GPUs with coordinated warp and fetch scheduling · SC 2015
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.212015
ELF: maximizing memory-level parallelism for GPUs with coordinated warp and fetch scheduling · SC 2015
Hardware accelerators and domain-specific architectures › data-parallel accelerator
SIMD accelerator
0.112012
Libra: Tailoring SIMD Execution Using Heterogeneous Hardware and Dynamic Configurability · MICRO 2012
Processor architecture and microarchitecture
latency hiding
0.112015
ELF: maximizing memory-level parallelism for GPUs with coordinated warp and fetch scheduling · SC 2015
Embedded and real-time systems
mobile computing
0.012012
Libra: Tailoring SIMD Execution Using Heterogeneous Hardware and Dynamic Configurability · MICRO 2012

Methods — techniques the papers use, named apart from their topics

simulation · 0.4resource partitioning · 0.3direct measurement · 0.3collaborative preemption · 0.2
YearPublicationVenuePosition
2017 Dynamic Resource Management for Efficient Utilization of Multitasking GPUs
abstract
As graphics processing units (GPUs) are broadly adopted, running multiple applications on a GPU at the same time is beginning to attract wide attention. Recent proposals on multitasking GPUs have focused on either spatial multitasking, which partitions GPU resource at a streaming multiprocessor (SM) granularity, or simultaneous multikernel (SMK), which runs multiple kernels on the same SM. However, multitasking performance varies heavily depending on the resource partitions within each scheme, and the application mixes. In this paper, we propose GPU Maestro that performs dynamic resource management for efficient utilization of multitasking GPUs. GPU Maestro can discover the best performing GPU resource partition exploiting both spatial multitasking and SMK. Furthermore, dynamism within a kernel and interference between the kernels are automatically considered because GPU Maestro finds the best performing partition through direct measurements. Evaluations show that GPU Maestro can improve average system throughput by 20.2% and 13.9% over the baseline spatial multitasking and SMK, respectively.
Jason Jong Kyu Park, Yongjun Park 0001, Scott A. Mahlke
ASPLOS1
2015 Fine Grain Cache Partitioning Using Per-Instruction Working Blocks
abstract
A traditional least-recently used (LRU) cache replacement policy fails to achieve the performance of the optimal replacement policy when cache blocks with diverse reuse characteristics interfere with each other. When multiple applications share a cache, it is often partitioned among the applications because cache blocks show similar reuse characteristics within each application. In this paper, we extend the idea to a single application by viewing a cache as a shared resource between individual memory instructions. To that end, we propose Instruction-based LRU (ILRU), a fine grain cache partitioning that way-partitions individual cache sets based on per-instruction working blocks, which are cache blocks required by an instruction to satisfy all the reuses within a set. In ILRU, a memory instruction steals a block from another only when it requires more blocks than it currently has. Otherwise, a memory instruction victimizes among the cache blocks inserted by itself. Experiments show that ILRU can improve the cache performance in all levels of cache, reducing the number of misses by an average of 7.0% for L1, 9.1% for L2, and 8.7% for L3, which results in a geometric mean performance improvement of 5.3%. ILRU for a three-level cache hierarchy imposes a modest 1.3% storage overhead over the total cache size.
Jason Jong Kyu Park, Yongjun Park 0001, Scott A. Mahlke
PACT1
2015 Chimera: Collaborative Preemption for Multitasking on a Shared GPU
abstract
The demand for multitasking on graphics processing units (GPUs) is constantly increasing as they have become one of the default components on modern computer systems along with traditional processors (CPUs). Preemptive multitasking on CPUs has been primarily supported through context switching. However, the same preemption strategy incurs substantial overhead due to the large context in GPUs. The overhead comes in two dimensions: a preempting kernel suffers from a long preemption latency, and the system throughput is wasted during the switch. Without precise control over the large preemption overhead, multitasking on GPUs has little use for applications with strict latency requirements.
Jason Jong Kyu Park, Yongjun Park 0001, Scott A. Mahlke
ASPLOS1
2015 ELF: maximizing memory-level parallelism for GPUs with coordinated warp and fetch scheduling
abstract
Graphics processing units (GPUs) are increasingly utilized as throughput engines in the modern computer systems. GPUs rely on fast context switching between thousands of threads to hide long latency operations, however, they still stall due to the memory operations. To minimize the stalls, memory operations should be overlapped with other operations as much as possible to maximize memory-level parallelism (MLP). In this paper, we propose Earliest Load First (ELF) warp scheduling, which maximizes the MLP by giving higher priority to the warps that have the fewest instructions to the next memory load. ELF utilizes the same warp priority for the fetch scheduling so that both are coordinated. We also show that ELF reveals its full benefits when there are fewer memory conflicts and fetch stalls. Evaluations show that ELF can improve the performance by 4.1% and achieve total improvement of 11.9% when used with other techniques over commonly-used greedy-then-oldest scheduling.
Jason Jong Kyu Park, Yongjun Park 0001, Scott A. Mahlke
SC1
2013 Efficient execution of augmented reality applications on mobile programmable accelerators
abstract
Mobile devices are ubiquitous in daily lives. From smartphones to tablets, customers are constantly demanding richer user experiences through more visual and interactive interface with prolonged battery life. To meet the demands, accelerators are commonly adopted in system-on-chip (SoC) for various applications. Coarse-grained reconfigurable architecture (CGRA) is a promising solution, which accelerates hot loops with software pipelining. Although CGRAs have shown that they can support multimedia applications efficiently, more interactive applications such as augmented reality put much more pressure on performance and energy requirements. In this paper, we extend heterogeneous CGRA to provide SIMD capabilities, which improves performance and energy efficiency significantly for augmented reality applications. We show that if we can exploit data level parallelism (DLP), it is more beneficial to run on SIMD natively than to transform it into instruction level parallelism (ILP) and run on CGRA. To utilize this property, multiple processing elements in CGRA are grouped to form homogeneous SIMD cores. To reduce the hardware overhead of fetching and replicating configuration in SIMD mode, we propose a ring network and a recycle buffer to pass the configuration around as well as to temporarily store it, which has minimized impact on throughput. Also, we modify memory access units and memory banks to support split memory transactions with forwarding for handling SIMD data access. To adapt to the proposed extension, we introduce a compile technique for SIMD mode code generation to maximize the resource utilization of each SIMD core. Experimental results show that it is possible to achieve an average of 17.6% performance improvement while saving 16.9% energy over heterogeneous CGRA.
Jason Jong Kyu Park, Yongjun Park 0001, Scott A. Mahlke
FPT1
2012 Efficient performance scaling of future CGRAs for mobile applications
abstract
Mobile computing as exemplified by the smart phone has become an integral part of our daily lives. The next generation of these devices will be driven by providing richer user experiences and compelling capabilities: higher definition multimedia, 3D graphics, augmented reality, and voice interfaces. To meet these goals, the core computing capabilities of mobile terminals must be scaled within highly constrained energy budgets. Coarse-grained reconfigurable architectures (CGRAs) are an appealing hardware platform for mobile systems by providing programmability with the potential for high computational throughput, low cost, and energy efficiency. CGRAs are most commonly used for innermost loops that contain an abundance of instruction-level parallelism. Unfortunately, current CGRAs fail to meet future performance requirements due to their inability to scale. Simply increasing the size of the array is too expensive in terms of power and area. In this paper, we first perform a deep analysis of several mobile applications from the domains of multimedia and gaming. We then explore potential solutions in the context of these applications for scaling the array performance in an energy efficient manner: homogeneous versus heterogeneous functionality, interconnect topologies, simple versus complex processing elements, and scalar versus vector memory support.
Yongjun Park 0001, Jason Jong Kyu Park, Scott A. Mahlke
FPT2
2012 Libra: Tailoring SIMD Execution Using Heterogeneous Hardware and Dynamic Configurability
abstract
Mobile computing as exemplified by the smart phone has become an integral part of our daily lives. The next generation of these devices will be driven by providing an even richer user experience and compelling capabilities: higher definition multimedia, 3D graphics, augmented reality, games, and voice interfaces. To address these goals, the core computing capabilities of the smart phone must be scaled. However, the energy budgets are increasing at a much lower rate, requiring fundamental improvements in computing efficiency. SIMD accelerators offer the combination of high performance and low energy consumption through low control and interconnect overhead. However, SIMD accelerators are not a panacea. Many applications lack sufficient vector parallelism to effectively utilize a large number of SIMD lanes. Further, the use of symmetric hardware lanes leads to low utilization and high static power dissipation as SIMD width is scaled. To address these inefficiencies, this paper focuses on breaking two traditional rules of SIMD processing: homogeneity and static configuration. The Libra accelerator increases SIMD utility by blurring the divide between vector and instruction parallelism to support efficient execution of a wider range of loops, and it increases hardware utilization through the use of heterogeneous hardware across the SIMD lanes. Experimental results show that the 32-lane Libra outperforms traditional SIMD accelerators by an average of 1.58x performance improvement due to higher loop coverage with 29% less energy consumption through heterogeneous hardware.
Yongjun Park 0001, Jason Jong Kyu Park, Hyunchul Park 0001, Scott A. Mahlke
MICRO2