Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hyunchul Park 0001

dblp:60/776-1 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
0since 2021 · last 2012
0000-0002-4025-6153ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 27% Processor architecture and microarchitecture · 27% Hardware accelerators and domain-specific architectures · 22%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
data parallelism
0.112012
SIMD defragmenter: efficient ILP realization on data-parallel architectures · ASPLOS 2012
Processor architecture and microarchitecture
instruction-level parallelism
0.112012
SIMD defragmenter: efficient ILP realization on data-parallel architectures · ASPLOS 2012
Hardware accelerators and domain-specific architectures › data-parallel accelerator
SIMD accelerator
0.112012
Libra: Tailoring SIMD Execution Using Heterogeneous Hardware and Dynamic Configurability · MICRO 2012
Parallel and multicore computing › data parallelism
SIMD vectorization
0.112012
SIMD defragmenter: efficient ILP realization on data-parallel architectures · ASPLOS 2012
Hardware accelerators and domain-specific architectures › video coding accelerator
multimedia accelerators
0.112009
Polymorphic pipeline array: a flexible multicore accelerator with virtualized execution for mobile multimedia applications · MICRO 2009
Electronic design automation
high-level synthesis
0.112005
Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005
Reconfigurable computing and FPGAs
modulo scheduling
0.112005
Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005
Electronic design automation › high-level synthesis
scheduling
0.112005
Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005
Processor architecture and microarchitecture › instruction set architecture › instruction set extension
application-specific instruction set extension
0.012004
Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization · MICRO 2004
Processor architecture and microarchitecture
instruction set architecture
0.012004
Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization · MICRO 2004
Processor architecture and microarchitecture › instruction set architecture
instruction set customization
0.012004
Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization · MICRO 2004
Compilers and program optimization › code generation
SIMD code generation
0.012012
SIMD defragmenter: efficient ILP realization on data-parallel architectures · ASPLOS 2012
Compilers and program optimization
vectorization
0.012012
SIMD defragmenter: efficient ILP realization on data-parallel architectures · ASPLOS 2012
Embedded and real-time systems
mobile computing
0.012012
Libra: Tailoring SIMD Execution Using Heterogeneous Hardware and Dynamic Configurability · MICRO 2012
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.012009
Polymorphic pipeline array: a flexible multicore accelerator with virtualized execution for mobile multimedia applications · MICRO 2009
Embedded and real-time systems › mobile computing
mobile multimedia
0.012009
Polymorphic pipeline array: a flexible multicore accelerator with virtualized execution for mobile multimedia applications · MICRO 2009
Compilers and program optimization
instruction scheduling
0.012005
Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005

Methods — techniques the papers use, named apart from their topics

subgraph-level SIMDization · 0.3data packing/unpacking · 0.3simulation · 0.2integer linear programming · 0.1decomposition · 0.1branch-and-bound · 0.1trace cache · 0.1subgraph identification · 0.1
YearPublicationVenuePosition
2012 SIMD defragmenter: efficient ILP realization on data-parallel architectures
abstract
Single-instruction multiple-data (SIMD) accelerators provide an energy-efficient platform to scale the performance of mobile systems while still retaining post-programmability. The central challenge is translating the parallel resources of the SIMD hardware into real application performance. In scientific applications, automatic vectorization techniques have proven quite effective at extracting large levels of data-level parallelism (DLP). However, vectorization is often much less effective for media applications due to low trip count loops, complex control flow, and non-uniform execution behavior. As a result, SIMD lanes remain idle due to insufficient DLP. To attack this problem, this paper proposes a new vectorization pass called SIMD Defragmenter to uncover hidden DLP that lurks below the surface in the form of instruction-level parallelism (ILP). The difficulty is managing the data packing/unpacking overhead that can easily exceed the benefits gained through SIMD execution. The SIMD degragmenter overcomes this problem by identifying groups of compatible instructions (subgraphs) that can be executed in parallel across the SIMD lanes. By SIMDizing in bulk at the subgraph level, packing/unpacking overhead is minimized. On a 16-lane SIMD processor, experimental results show that SIMD defragmentation achieves a mean 1.6x speedup over traditional loop vectorization and a 31% gain over prior research approaches for converting ILP to DLP.
Yongjun Park 0001, Sangwon Seo, Hyunchul Park 0001, Hyoun Kyu Cho, Scott A. Mahlke
ASPLOS3
2012 Libra: Tailoring SIMD Execution Using Heterogeneous Hardware and Dynamic Configurability
abstract
Mobile computing as exemplified by the smart phone has become an integral part of our daily lives. The next generation of these devices will be driven by providing an even richer user experience and compelling capabilities: higher definition multimedia, 3D graphics, augmented reality, games, and voice interfaces. To address these goals, the core computing capabilities of the smart phone must be scaled. However, the energy budgets are increasing at a much lower rate, requiring fundamental improvements in computing efficiency. SIMD accelerators offer the combination of high performance and low energy consumption through low control and interconnect overhead. However, SIMD accelerators are not a panacea. Many applications lack sufficient vector parallelism to effectively utilize a large number of SIMD lanes. Further, the use of symmetric hardware lanes leads to low utilization and high static power dissipation as SIMD width is scaled. To address these inefficiencies, this paper focuses on breaking two traditional rules of SIMD processing: homogeneity and static configuration. The Libra accelerator increases SIMD utility by blurring the divide between vector and instruction parallelism to support efficient execution of a wider range of loops, and it increases hardware utilization through the use of heterogeneous hardware across the SIMD lanes. Experimental results show that the 32-lane Libra outperforms traditional SIMD accelerators by an average of 1.58x performance improvement due to higher loop coverage with 29% less energy consumption through heterogeneous hardware.
Yongjun Park 0001, Jason Jong Kyu Park, Hyunchul Park 0001, Scott A. Mahlke
MICRO3
2010 Resource recycling: putting idle resources to work on a composable accelerator
abstract
Mobile computing platforms in the form of smart phones, netbooks, and personal digital assistants have become an integral part of our everyday lives. Moving ahead to the future, mobile multimedia support will become a key differentiating factor for customers. Features such as high-definition audio and video, video conferencing, 3D graphics, and image projection will lead to the adoption of one phone over another. However, in contrast to wireless signal processing which is dominated by vectorizable computation, mobile multimedia applications often contain complex control flow and variable computational requirements. Moreover, data access is more complex where media applications typically operate on multi-dimensional vectors of data rather than single-dimensional vectors with simple strides. To handle these complexities, composable accelerators such as the Polymorphic Pipeline Array, or PPA, present an appealing hardware platform by adding a degree of hardware configurability over existing accelerators. Hardware resources can be both statically as well as dynamically partitioned among executing tasks to maximize execution efficiency. However, an effective compilation framework is essential to partition and assign resources to make intelligent use of the available hardware. In this paper, a compilation framework is introduced that maximizes application throughput with hybrid resource partitioning of a PPA system. Static partitioning handles part of the resource assignment, but this is followed up by dynamic partitioning to identify idle resources and put them to use -- resource recycling. Experimental results show that real-time media applications can take advantage of the static and dynamic configurability of the PPA for increase.
Yongjun Park 0001, Hyunchul Park 0001, Scott A. Mahlke, Sukjin Kim
CASES2
2009 CGRA express: accelerating execution using dynamic operation fusion
abstract
Coarse-grained reconfigurable architectures (CGRAs) present an appealing hardware platform by providing programmability with the potential for high computation throughput, scalability, low cost, and energy efficiency. CGRAs have been effectively used for innermost loops that contain an abundant of instruction-level parallelism. Conversely, non-loop and outer-loop code are latency constrained and do not offer significant amounts of instruction-level parallelism. In these situations, CGRAs are ineffective as the majority of the resources remain idle. In this paper, dynamic operation fusion is introduced to enable CGRAs to effectively accelerate latency-constrained code regions. Dynamic operation fusion is enabled through the combination of a small bypass network added between function units in a conventional CGRA and a sub-cycle modulo scheduler to automatically identify opportunities for fusion. Results show that dynamic operation fusion reduced total application run-time by up to 17% on a 4x4 CGRA.
Yongjun Park 0001, Hyunchul Park 0001, Scott A. Mahlke
CASES2
2009 Recurrence cycle aware modulo scheduling for coarse-grained reconfigurable architectures
abstract
In high-end embedded systems, coarse-grained reconfigurable architectures (CGRA) continue to replace traditional ASIC designs. CGRAs offer high performance at a low power consumption, yet provide flexibility through programmability. In this paper we introduce a recurrence cycle-aware scheduling technique for CGRAs. Our modulo scheduler groups operations belonging to a recurrence cycle into a clustered node and then computes a scheduling order for those clustered nodes. Deadlocks that arise when two or more recurrence cycles depend on each other are resolved by using heuristics that favor recurrence cycles with long recurrence delays. While with previous work one had to sacrifice either a fast compilation speed in order to get good quality results, or vice versa, this is not necessary anymore with the proposed recurrence cycle-aware scheduling technique. We have implemented the proposed method into our in-house CGRA chip and compiler solution and show that the technique achieves better quality schedules than schedulers based on simulated annealing at a 170-fold speed increase.
Taewook Oh, Bernhard Egger 0002, Hyunchul Park 0001, Scott A. Mahlke
LCTES3
2009 Polymorphic pipeline array: a flexible multicore accelerator with virtualized execution for mobile multimedia applications
abstract
Mobile computing in the form of smart phones, netbooks, and personal digital assistants has become an integral part of our everyday lives. Moving ahead to the next generation of mobile devices, we believe that multimedia will become a more critical and product-differentiating feature. High definition audio and video as well as 3D graphics provide richer interfaces and compelling capabilities. However, these algorithms also bring different computational challenges than wireless signal processing. Multimedia algorithms are more complex featuring more control flow and variable computational requirements where execution time is not dominated by innermost vector loops. Further, data access is more complex where media applications typically operate on multi-dimensional vectors of data rather than single-dimensional vectors with simple strides. Thus, the design of current mobile platforms requires re-examination to account for these new application domains. In this work, we focus on the design of a programmable, low-power accelerator for multimedia algorithms referred to as a Polymorphic Pipeline Array, or PPA. The PPA is designed with flexibility and programmability as first-order requirements to enable the hardware to be dynamically customizable to the application. PPAs exploit pipeline parallelism found in streaming applications to create a coarse-grain hardware pipeline to execute streaming media applications. PPA resources are allocated to each stage depending on its size and ability to exploit fine-grain parallelism. Experimental results show that real-time media applications can take advantage of the static and dynamic configurability for increased power efficiency.
Hyunchul Park 0001, Yongjun Park 0001, Scott A. Mahlke
MICRO1
2008 Edge-centric modulo scheduling for coarse-grained reconfigurable architectures
abstract
Coarse-grained reconfigurable architectures (CGRAs) present an appealing hardware platform by providing the potential for high computation throughput, scalability, low cost, and energy efficiency. CGRAs consist of an array of function units and register files often organized as a two dimensional grid. The most difficult challenge in deploying CGRAs is compiler scheduling technology that can efficiently map software implementations of compute intensive loops onto the array. Traditional schedulers focus on the placement of operations in time and space. With CGRAs, the challenge of placement is compounded by the need to explicitly route operands from producers to consumers. To systematically attack this problem, we take an edge-centric approach to modulo scheduling that focuses on the routing problem as its primary objective. With edge-centric modulo scheduling (EMS), placement is a by-product of the routing process, and the schedule is developed by routing each edge in the dataflow graph. Routing cost metrics provide the scheduler with a global perspective to guide selection. Experiments on a wide variety of compute-intensive loops from the multimedia domain show that EMS improves throughput by 25% over traditional iterative modulo scheduling, and achieves 98% of the throughput of simulated annealing techniques at a fraction of the compilation time.
Hyunchul Park 0001, Kevin Fan, Scott A. Mahlke, Taewook Oh
PACT1
2008 Modulo scheduling for highly customized datapaths to increase hardware reusability
abstract
In the embedded domain, custom hardware in the form of ASICs is often used to implement critical parts of applications when performance and energy efficiency goals cannot be met with software implementations on a general purpose processor or DSP. The downsides of using ASICs include high non-recurring engineering costs, inability to accommodate changes in the application after production, and inability to reuse hardware for new applications. However, by allowing a degree of post-programmability, the hardware can retain high performance and energy efficiency while increasing flexibility and reusability. The difficulty with programmable custom hardware lies in mapping new applications onto an existing datapath that is both sparse and irregular. This paper proposes a constraint-driven modulo scheduler that maps software-pipelineable loops onto programmable loop accelerator hardware. The scheduler is able to target accelerators with widely varying levels of datapath functional capability and connectivity, and thus, varying degrees of programmability. The paper investigates the ability of the scheduler to map new loops onto existing hardware, which depends on both the degree of programmability of the hardware as well as the similarity of the new loop to the original loop for which the hardware was designed.
Kevin Fan, Hyunchul Park 0001, Manjunath Kudlur, Scott A. Mahlke
CGO2
2006 Modulo graph embedding: mapping applications onto coarse-grained reconfigurable architectures
abstract
Coarse-grained reconfigurable architectures (CGRAs) present an appealing hardware platform by providing the potential for high computation throughput, scalability, low cost and energy efficiency. CGRAs consist of an array of function units and register files generally organized as a two dimensional grid. The most difficult challenge with deploying CGRAs is compiler scheduling technology that can map software implementations of compute intensive loops onto the array. Traditional schedulers are not suitable because they do not take into account the explicit routing of operand values that is necessary. In essence, the problem of binding operations to time slots and resources is extended to also include explicit routing of operands from producers to consumers. To tackle this problem, this paper introduces a software pipelining technique for mapping loop bodies onto CGRAs, referred to as modulo graph embedding. We leverage graph embedding from graph theory, which is used to draw graphs onto a target space. The loop body is essentially drawn onto the CGRA mesh, subject to modulo resource usage constraints. Modulo graph embedding is effective because it can take into account the communication structure of the loop body during mapping. On average, a compute utilization of 56-68% is achieved for a set of loop kernels across three 4x4 CGRA designs.
Hyunchul Park 0001, Kevin Fan, Manjunath Kudlur, Scott A. Mahlke
CASES1
2005 Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System
abstract
Scheduling algorithms used in compilers traditionally focus on goals such as reducing schedule length and register pressure or producing compact code. In the context of a hardware synthesis system where the schedule is used to determine various components of the hardware, including datapath, storage, and interconnect, the goals of a scheduler change drastically. In addition to achieving the traditional goals, the scheduler must proactively make decisions to ensure efficient hardware is produced. This paper proposes two exact solutions for cost sensitive modulo scheduling, one based on an integer linear programming formulation and another based on branch-and-bound search. To achieve reasonable compilation times, decomposition techniques to break down the complex scheduling problem into phase ordered sub-problems are proposed. The decomposition techniques work either by partitioning the dataflow graph into smaller subgraphs and optimally scheduling the subgraphs, or by splitting the scheduling problem into two phases, time slot and resource assignment. The effectiveness of cost sensitive modulo scheduling in minimizing the costs of function units, register structures, and interconnection wires are evaluated within a fully automatic synthesis system for loop accelerators. The cost sensitive modulo scheduler increases the efficiency of the resulting hardware significantly compared to both traditional cost unaware and greedy cost aware modulo schedulers.
Kevin Fan, Manjunath Kudlur, Hyunchul Park 0001, Scott A. Mahlke
MICRO3
2004 Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization
abstract
Application-specific instruction set extensions are an effective way of improving the performance of processors. Critical computation subgraphs can be accelerated by collapsing them into new instructions that are executed on specialized function units. Collapsing the subgraphs simultaneously reduces the length of computation as well as the number of intermediate results stored in the register file. The main problem with this approach is that a new processor must be generated for each application domain. While new instructions can be designed automatically, there is a substantial amount of engineering cost incurred to verify and to implement the final custom processor. In this work, we propose a strategy to transparent customization of the core computation capabilities of the processor without changing its instruction set. A congurable array of function units is added to the baseline processor that enables the acceleration of a wide range of data flow subgraphs. To exploit the array, the microarchitecture performs subgraph identification at run-time, replacing them with new microcode instructions to configure and utilize the array. We compare the effectiveness of replacing subgraphs in the fill unit of a trace cache versus using a translation table during decode, and evaluate the tradeoffs between static and dynamic identification of subgraphs for instruction set customization.
Nathan Clark, Manjunath Kudlur, Hyunchul Park 0001, Scott A. Mahlke, Krisztián Flautner
MICRO3