VLDB 2026 Research / reviewers in the wild / expert
Gang-Ryung Uh
dblp:26/2665
· DBLP profile ↗
18ranked-venue papers
5as first author
2since 2021 · last 2026
0009-0008-2811-3724ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 4 first-author · 2 since 2021Systems, architecture and hardware · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can Fine-Grain Multi-threading Subsume VLIW?abstractWe explore the question: "Can a fine-grain multi-threaded architecture form the basis for an efficient, VLIW style, statically scheduled architecture?" We illustrate that operations comprising a VLIW instruction can indeed be viewed as belonging to separate threads, such that the number of such operations is equivalent to the number of threads representing the program's semantics. On the other hand, a more efficient synchronization mechanism than data synchronization is needed to realize the lock-step execution model of VLIW processors. This synchronization is accomplished through the instruction space, by using a small number of bits in each instruction under the compiler control. We call the resulting architecture a "Synchronized Lane Architecture (SLA)". Scott Pomerville, Soner Önder, Gang-Ryung Uh, David B. Whalley |
LCTES | 3 |
| 2023 | Facilitating the Bootstrapping of a New ISAabstractImplementation of a new instruction set architecture (ISA) is a non-trivial task that involves significant modifications to the system software, such as the compiler, the assembler, and the linker. This task also includes modifying and verifying functional and cycle accurate simulators to facilitate performance evaluation of programs under the new ISA. Isolating errors in these software components becomes extremely challenging and demands automated and semi-automated mechanisms since neither the compilation infrastructure nor the simulation infrastructure can be trusted as both parties have been heavily modified. Bootstrapping a new ISA is very common in embedded systems since there is a greater variety of embedded ISAs due to often not having a need to support backward compatibility of executables. In this paper, we present the tools and the verification mechanisms we have implemented to support the development of a number of related, but distinct ISAs. Our work in developing the system software and simulators for these ISAs demonstrate that a step-by-step semi-automated approach which relies on simple invariants can facilitate effective bootstrapping of the complete system software and the simulator infrastructure. Abigail Mortensen, Scott Pomerville, David B. Whalley, Soner Önder, Gang-Ryung Uh |
LCTES | 5 |
| 2015 | Scheduling instruction effects for a statically pipelined processorabstractStatically pipelined processors have a fully exposed datapath where all portions of the pipeline are directly controlled by effects within an instruction, which simplifies hardware and enables a new level of compiler optimizations. This paper describes an effect scheduling strategy to aggressively compact instructions, which has a critical impact on code size and performance. Unique scheduling challenges include more frequent name dependences and fewer renaming opportunities due to static pipeline (SP) registers being dedicated for specific operations. We also realized the SP in a hardware implementation language (VHDL) to evaluate the real energy benefits. Despite the compiler challenges, we achieve performance, code size, and energy improvements compared to a conventional MIPS processor. B. Davis, Ryan Baird, Peter Gavin, Magnus Själander, Ian Finlayson, F. Rasapour, G. Cook, Gang-Ryung Uh, David B. Whalley, Gary S. Tyson |
CASES | 8 |
| 2015 | Optimizing Transfers of Control in the Static Pipeline ArchitectureabstractStatically pipelined processors offer a new way to improve the performance beyond that of a traditional in-order pipeline while simultaneously reducing energy usage by enabling the compiler to control more fine-grained details of the program execution. This paper describes how a compiler can exploit the features of the static pipeline architecture to apply optimizations on transfers of control that are not possible on a conventional architecture. The optimizations presented in this paper include hoisting the target address calculations for branches, jumps, and calls out of loops, performing branch chaining between calls and jumps, hoisting the setting of return addresses out of loops, and exploiting conditional calls and returns. The benefits of performing these transfer of control optimizations include a 6.8% reduction in execution time and a 3.6% decrease in estimated energy usage. Ryan Baird, Peter Gavin, Magnus Själander, David B. Whalley, Gang-Ryung Uh |
LCTES | 5 |
| 2013 | Improving processor efficiency by statically pipelining instructionsabstractA new generation of applications requires reduced power consumption without sacrificing performance. Instruction pipelining is commonly used to meet application performance requirements, but some implementation aspects of pipelining are inefficient with respect to energy usage. We propose static pipelining as a new instruction set architecture to enable more efficient instruction flow through the pipeline, which is accomplished by exposing the pipeline structure to the compiler. While this approach simplifies hardware pipeline requirements, significant modifications to the compiler are required. This paper describes the code generation and compiler optimizations we implemented to exploit the features of this architecture. We show that we can achieve performance and code size improvements despite a very low-level instruction representation. We also demonstrate that static pipelining of instructions reduces energy usage by simplifying hardware, avoiding many unnecessary operations, and allowing the compiler to perform optimizations that are not possible on traditional architectures. Ian Finlayson, Brandon Davis, Peter Gavin, Gang-Ryung Uh, David B. Whalley, Magnus Själander, Gary S. Tyson |
LCTES | 4 |
| 2007 | Preprocessing Strategy for Effective Modulo Scheduling on Multi-issue Digital Signal Processors
Doosan Cho, Ravi Ayyagari, Gang-Ryung Uh, Yunheung Paek |
CC | 3 |
| 2005 | Branch elimination by condition mergingabstractConditional branches are expensive. Branches require a significant percentage of execution cycles since they occur frequently and cause pipeline flushes when mispredicted. In addition, branches result in forks in the control flow, which can prevent other code-improving transformations from being applied. In this paper we describe profile-based techniques for replacing the execution of a set of two or more branches with a single branch on a conventional scalar processor. These sets of branches can include tests of multiple variables. For instance, the test if (p1 != 0 && p2 != 0), which is testing for NULL pointers, can be replaced with if (p1 & p2 != 0). Program profiling is performed to target condition merging along frequently executed paths. The results show that eliminating branches by merging conditions can significantly reduce the number of conditional branches executed in non-numerical applications. Copyright © 2004 John Wiley & Sons, Ltd. William C. Kreahling, David B. Whalley, Mark W. Bailey, Xin Yuan 0001, Gang-Ryung Uh, Robert A. van Engelen |
Softw. Pract. Exp. | 5 |
| 2005 | Compiler transformations for effectively exploiting a zero overhead loop bufferabstractAbstract A Zero Overhead Loop Buffer (ZOLB) is an architectural feature that is commonly found in DSP processors. This buffer can be viewed as a compiler managed cache that contains a sequence of instructions that will be executed a specified number of times without incurring any loop overhead. Unlike loop unrolling, a loop buffer can be used to minimize loop overhead without the penalty of increasing code size. In addition, a ZOLB requires relatively little space and power, which are both important considerations for most DSP applications. This paper describes strategies for generating code to effectively use a ZOLB. We have found that many common code improving transformations used by optimizing compilers on conventional architectures can be easily used to (1) allow more loops to be placed in a ZOLB, (2) further reduce loop overhead of the loops placed in a ZOLB, and (3) avoid redundant loading of ZOLB loops. The results given in this paper demonstrate that this architectural feature can often be exploited with substantial improvements in execution time and slight reductions in code size for various signal processing applications. Copyright © 2004 John Wiley & Sons, Ltd. Gang-Ryung Uh, David B. Whalley, Sanjay Jinturkar, Yunheung Paek, Vincent Cao |
Softw. Pract. Exp. | 1 |
| 2004 | Tuning the WCET of Embedded ApplicationsabstractIt is advantageous to not only calculate the WCET of an application, but to also perform transformations to reduce the WCET since an application with a lower WCET is less likely to violate its timing constraints. In this paper we describe an environment consisting of an interactive compilation system and a timing analyzer, where a user can interactively tune the WCET of an application. After each optimization phase is applied, the timing analyzer is automatically invoked to calculate the WCET of the function being tuned. Thus, a user can easily gauge the progress of reducing the WCET. In addition, the user can apply a genetic algorithm to search for an effective optimization sequence that best reduces the WCET. Using the genetic algorithm, we show that the WCET for a number of applications can be reduced by 7% on average as compared to the default batch optimization sequence. Wankang Zhao, Prasad A. Kulkarni, David B. Whalley, Christopher A. Healy, Frank Mueller 0001, Gang-Ryung Uh |
IEEE Real-Time and Embedded Technology and Applications Symposium | 6 |
| 2004 | Code optimizations for a VLIW-style network processing unitabstractAbstract The explosive growth in network bandwidth and Internet services such as QoS (quality of service) and SLA (service level agreement) monitoring have created the need for new networking hardware called aNetwork Processing Unit (NPU). In order to rapidly reconfigure the NPU for frequently varying Internet services and technologies, a high‐performance C compiler is urgently needed. Several code generation techniques, which are intended to meet the high code quality demands of other types ofapplication specific instruction‐set processors(ASIPs) likedigital signal processors(DSPs), have already been developed. However, these techniques are insufficient for NPUs due to striking architectural differences such as asymmetric data paths. The main purpose of this paper is to discuss our recent experience with the development of a commercial compiler for a new NPU called thePaion PPII, which is basically apacket enginefor NPU to meet the growing need for new high‐bandwidth communication equipment targeted for Internet routers and ethernet adapters. For this purpose, we will first show the architectural challenges posed by the target NPU. Then, we will describe several compiler techniques that we found to be effective for the target NPU with various unorthogonal architectural features. The current implementations of the PPII use a VLIW (Very Long Instruction Word) architecture. So, we handled this VLIW‐style architecture by employing a simplecode compactionscheme which packs multiple parallel instructions into one long instruction word. The experimental results show that our techniques are effective for significantly reducing the dynamic instruction count. Copyright © 2004 John Wiley & Sons, Ltd. Jinhwan Kim, Yunheung Paek, Gang-Ryung Uh |
Softw. Pract. Exp. | 3 |
| 2003 | Branch Elimination via Multi-variable Condition Merging
William C. Kreahling, David B. Whalley, Mark W. Bailey, Xin Yuan 0001, Gang-Ryung Uh, Robert A. van Engelen |
Euro-Par | 5 |
| 2003 | Tailoring Software Pipelining for Effective Exploitation of Zero Overhead Loop Buffer
Gang-Ryung Uh |
SCOPES | 1 |
| 2002 | Experience with a retargetable compiler for a commercial network processorabstractThe Paion PPII network processor is designed to meet the growing need for new high bandwidth network equipment. In order to rapidly reconfigure the processor for frequently varying internet services and technologies, a high performance compiler is urgently needed. Albeit various code generation techniques have been proposed for DSPs or ASIPs, we experienced these techniques are not easily tailored towards the target Paion PPII processor due to striking architectural differences. First, we will show the architectural challenges posed by the target processor. Second, novel compiler techniques will be described that effectively exploit unorthogonal architectural features. The techniques include virtual data path, compiler intrinsics, and interprocedural register allocation. Third, intermediate benchmark results will be presented to demonstrate the effectiveness of our techniques. Jinhwan Kim, Sungjoon Jung, Yunheung Paek, Gang-Ryung Uh |
CASES | 4 |
| 2002 | Efficient and effective branch reordering using profile dataabstractThe conditional branch has long been considered an expensive operation. The relative cost of conditional branches has increased as recently designed machines are now relying on deeper pipelines and higher multiple issue. Reducing the number of conditional branches executed often results in a substantial performance benefit. This paper describes a code-improving transformation to reorder sequences of conditional branches that compare a common variable to constants. The goal is to obtain an ordering where the fewest average number of branches in the sequence will be executed. First, sequences of branches that can be reordered are detected in the control flow. Second, profiling information is collected to predict the probability that each branch will transfer control out of the sequence. Third, the cost of performing each conditional branch is estimated. Fourth, the most beneficial ordering of the branches based on the estimated probability and cost is selected. The most beneficial ordering often includes the insertion of additional conditional branches that did not previously exist in the sequence. Finally, the control flow is restructured to reflect the new ordering. The results of applying the transformation are on average reductions of about 8% fewer instructions executed and 13% branches performed, as well as about a 4% decrease in execution time. Gang-Ryung Uh, David B. Whalley |
ACM Trans. Program. Lang. Syst. | 2 |
| 2000 | Techniques for Effectively Exploiting a Zero Overhead Loop Buffer
Gang-Ryung Uh, David B. Whalley, Sanjay Jinturkar, Vincent Cao |
CC | 1 |
| 1999 | Effectively Exploiting Indirect JumpsabstractThis paper describes a general code-improving transformation that can coalesce conditional branches into an indirect jump from a table. Applying this transformation allows an optimizer to exploit indirect jumps for many other coalescing opportunities besides the translation of multiway branch statements. First, dataflow analysis is performed to detect a set of coalescent conditional branches, which are often separated by blocks of intervening instructions. Secondly, several techniques are applied to reduce the cost of performing an indirect jump operation, often requiring the execution of only two instructions on a SPARC. Finally, the control flow is restructured using code duplication to replace the set of branches with an indirect jump. Thus, the transformation essentially provides early resolution of conditional branches that may originally have been some distance from the point where the indirect jump is inserted. The transformation can be frequently applied with often significant reductions in the number of instructions executed, total cache work, and execution time. In addition, we show that with branch target buffer support, indirect jumps improve branch prediction since they cause fewer mispredictions than the set of branches they replaced. Copyright © 1999 John Wiley & Sons, Ltd. Gang-Ryung Uh, David B. Whalley |
Softw. Pract. Exp. | 1 |
| 1998 | Improving Performance by Branch ReorderingabstractThe conditional branch has long been considered an expensive operation. The relative cost of conditional branches has increased as recently designed machines are now relying on deeper pipelines and higher multiple issue. Reducing the number of conditional branches executed can often result in a substantial performance benefit. This paper describes a code-improving transformation to reorder sequences of conditional branches. First, sequences of branches that can be reordered are detected in the control flow. Second, profiling information is collected to predict the probability that each branch will transfer control out of the sequence. Third, the cost of performing each conditional branch is estimated. Fourth, the most beneficial ordering of the branches based on the estimated probability and cost is selected. The most beneficial ordering often included the insertion of additional conditional branches that did not previously exist in the sequence. Finally, the control flow is restructured to refflect the new ordering. The results of applying the transformation were significant reductions in the dynamic number of instructions and branches, as well as decreases in execution time. Gang-Ryung Uh, David B. Whalley |
PLDI | 2 |
| 1997 | Coalescing Conditional Branches into Efficient Indirect Jumps
Gang-Ryung Uh, David B. Whalley |
SAS | 1 |