Yeonghun Jeong

dblp:129/7633 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Reconfigurable computing and FPGAs · 38% Electronic design automation · 38% GPUs and heterogeneous computing · 12%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop transformation
0.212013
Evaluator-executor transformation for efficient pipelining of loops with conditionals · ACM Trans. Archit. Code Optim. 2013
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.212013
Evaluator-executor transformation for efficient pipelining of loops with conditionals · ACM Trans. Archit. Code Optim. 2013
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining
0.212013
Evaluator-executor transformation for efficient pipelining of loops with conditionals · ACM Trans. Archit. Code Optim. 2013
GPUs and heterogeneous computing
control flow divergence
0.012013
Evaluator-executor transformation for efficient pipelining of loops with conditionals · ACM Trans. Archit. Code Optim. 2013
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.012013
Evaluator-executor transformation for efficient pipelining of loops with conditionals · ACM Trans. Archit. Code Optim. 2013

Methods — techniques the papers use, named apart from their topics

software transformation · 0.3hardware extension · 0.3
YearPublicationVenuePosition
2013 Fast shared on-chip memory architecture for efficient hybrid computing with CGRAs
abstract
While Coarse-Grained Reconfigurable Architectures (CGRAs) are very efficient at handling regular, compute-intensive loops, their weakness at control-intensive processing and the need for frequent reconfiguration require another processor, for which usually a main processor is used. To minimize the overhead arising in such collaborative execution, we integrate a dedicated sequential processor (SP) with a reconfigurable array (RA), where the crucial problem is how to share the memory between SP and RA while keeping the SP's memory access latency very short. We present a detailed architecture, control, and program example of our approach, focusing on our optimized on-chip shared memory organization between SP and RA. Our preliminary results demonstrate that our optimized memory architecture is very effective in reducing kernel execution times (23.5% compared to a more straightforward alternative), and our approach can reduce the RA control overhead and other sequential code execution time in kernels significantly, resulting in up to 23.1% reduction in kernel execution time, compared to the conventional system using the main processor for sequential code execution.
Jongeun Lee, Yeonghun Jeong, Sungsok Seo
DATE2
2013 Evaluator-executor transformation for efficient pipelining of loops with conditionals
abstract
Control divergence poses many problems in parallelizing loops. While predicated execution is commonly used to convert control dependence into data dependence, it often incurs high overhead because it allocates resources equally for both branches of a conditional statement regardless of their execution frequencies. For those loops with unbalanced conditionals, we propose a software transformation that divides a loop into two or three smaller loops so that the condition is evaluated only in the first loop, while the less frequent branch is executed in the second loop in a way that is much more efficient than in the original loop. To reduce the overhead of extra data transfer caused by the loop fission, we also present a hardware extension for a class of Coarse-Grained Reconfigurable Architectures (CGRAs). Our experiments using MiBench and computer vision benchmarks on a CGRA demonstrate that our techniques can improve the performance of loops over predicated execution by up to 65% (37.5%, on average), when the hardware extension is enabled. Without any hardware modification, our software-only version can improve performance by up to 64% (33%, on average), while simultaneously reducing the energy consumption of the entire CGRA including configuration and data memory by 22%, on average.
Yeonghun Jeong, Seongseok Seo, Jongeun Lee
ACM Trans. Archit. Code Optim.1