EDBT 2026 Demo / reviewers in the wild / expert
John C. Gyllenhaal
dblp:50/3479
· DBLP profile ↗
14ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 1 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Parallel and multicore computing · 51% Processor architecture and microarchitecture · 19% High-performance computing · 17% | |
| Software engineering, system software, and programming languages
10 papers |
Compilers and program optimization · 87% Runtime systems and virtual machines · 13% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
performance optimization at scale |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
synchronization |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
transactional memory |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 3 | 2001 | An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001 Architectural Support for Compiler-Synthesized Dynamic Branch Prediction Strategies: Rationale and Initial Results · HPCA 1997 Characterizing the impact of predicated execution on branch prediction · MICRO 1994 |
Parallel and multicore computing
runtime optimization |
0.1 | 2 | 2001 | An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001 A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.0 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 4 | 1995 | Characterizing the impact of predicated execution on branch prediction · MICRO 1994 Dynamic Memory Disambiguation Using the Memory Conflict Buffer · ASPLOS 1994 Compiler technology for future microprocessors · Proc. IEEE 1995 |
Performance modeling and evaluation › profiling
hardware profiling |
0.0 | 1 | 1999 | A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999 |
Electronic design automation › physical design › lithography
lithography hotspot detection |
0.0 | 1 | 1999 | A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999 |
Compilers and program optimization › instruction scheduling
instruction-level parallelism |
0.0 | 2 | 1995 | Compiler technology for future microprocessors · Proc. IEEE 1995 Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.0 | 2 | 1997 | Characterizing the impact of predicated execution on branch prediction · MICRO 1994 Architectural Support for Compiler-Synthesized Dynamic Branch Prediction Strategies: Rationale and Initial Results · HPCA 1997 |
Compilers and program optimization
dynamic optimization |
0.0 | 2 | 2001 | An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001 A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999 |
Runtime systems and virtual machines › binary translation
bytecode translation |
0.0 | 1 | 1996 | Java Bytecode to Native Code Translation: The Caffeine Prototype and Preliminary Results · MICRO 1996 |
Memory systems
cache |
0.0 | 2 | 1994 | Data relocation and prefetching for programs with large data sets · MICRO 1994 Speculative execution exception recovery using write-back suppression · MICRO 1993 |
Compilers and program optimization
dependence analysis |
0.0 | 1 | 1995 | Compiler technology for future microprocessors · Proc. IEEE 1995 |
Compilers and program optimization
predicated compilation |
0.0 | 1 | 1995 | Compiler technology for future microprocessors · Proc. IEEE 1995 |
Memory systems › cache
conflict miss reduction |
0.0 | 1 | 1994 | Data relocation and prefetching for programs with large data sets · MICRO 1994 |
Memory systems › cache › prefetching
data prefetching |
0.0 | 1 | 1994 | Data relocation and prefetching for programs with large data sets · MICRO 1994 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 2 | 1992 | Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 Code scheduling for VLIW/superscalar processors with limited register files · MICRO 1992 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 2 | 1992 | Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 Code scheduling for VLIW/superscalar processors with limited register files · MICRO 1992 |
Compilers and program optimization › instruction scheduling
software pipelining |
0.0 | 1 | 1993 | Superblock formation using static program analysis · MICRO 1993 |
Processor architecture and microarchitecture › instruction-level parallelism
multiple instruction issue |
0.0 | 1 | 1993 | Speculative execution exception recovery using write-back suppression · MICRO 1993 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 1 | 1993 | Speculative execution exception recovery using write-back suppression · MICRO 1993 |
Compilers and program optimization › program transformation
compiler transformations |
0.0 | 1 | 1992 | Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 |
Compilers and program optimization
instruction scheduling |
0.0 | 1 | 1992 | Code scheduling for VLIW/superscalar processors with limited register files · MICRO 1992 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 2 | 1996 | Optimization of Machine Descriptions for Efficient Use · MICRO 1996 Superblock formation using static program analysis · MICRO 1993 |
Runtime systems and virtual machines › managed runtime
java runtime |
0.0 | 1 | 1996 | Java Bytecode to Native Code Translation: The Caffeine Prototype and Preliminary Results · MICRO 1996 |
Memory systems › data movement
data relocation |
0.0 | 1 | 1994 | Data relocation and prefetching for programs with large data sets · MICRO 1994 |
Memory systems › cache › CPU cache
data cache |
0.0 | 1 | 1993 | Speculative execution exception recovery using write-back suppression · MICRO 1993 |
Methods — techniques the papers use, named apart from their topics
performance characterization · 0.1best practices · 0.1out-of-order execution · 0.1instruction profiling · 0.1hardware speculation · 0.1hardware performance counters · 0.0predicate generation · 0.0compiler synthesis · 0.0hardware memory conflict buffer · 0.0compiler repair code · 0.0stack-to-register mapping · 0.0exception handling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | What scientific applications can benefit from hardware transactional memory?abstractAchieving efficient and correct synchronization of multiple threads is a difficult and error-prone task at small scale and, as we march towards extreme scale computing, will be even more challenging when the resulting application is supposed to utilize millions of cores efficiently. Transactional Memory (TM) is a promising technique to ease the burden on the programmer, but only recently has become available on commercial hardware in the new Blue Gene/Q system and hence the real benefit for realistic applications has not been studied yet. This paper presents the first performance results of TM embedded into OpenMP on a prototype system of BG/Q and characterizes code properties that will likely lead to benefits when augmented with TM primitives. We first study the influence of thread count, environment variables and memory layout on TM performance and identify code properties that will yield performance gains with TM. Second, we evaluate the combination of OpenMP with multiple synchronization primitives on top of MPI to determine suitable task to thread ratios per node. Finally, we condense our findings into a set of best practices. These are applied to a Monte Carlo Benchmark and a Smoothed Particle Hydrodynamics method. In both cases an optimized TM version, executed with 64 threads on one node, outperforms a simple TM implementation. MCB with optimized TM yields a speedup of 27.45 over baseline. Martin Schindewolf, Barna L. Bihari, John C. Gyllenhaal, Martin Schulz 0001, Amy Wang, Wolfgang Karl |
SC | 3 |
| 2001 | An Architectural Framework for Runtime OptimizationabstractWide-issue processors continue to achieve higher performance by exploiting greater instruction-level parallelism. Dynamic techniques such as out-of-order execution and hardware speculation have proven effective at increasing instruction throughput. Runtime optimization promises to provide an even higher level of performance by adaptively applying aggressive code transformations on a larger scope. This paper presents a new hardware mechanism for generating and deploying runtime optimized code. The mechanism can be viewed as a filtering system that resides in the retirement stage of the processor pipeline, accepts an instruction execution stream as input, and produces instruction profiles and sets of linked, optimized traces as output. The code deployment mechanism uses an extension to the branch prediction mechanism to migrate execution into the new code without modifying the original code. These new components do not add delay to the execution of the program except during short bursts of reoptimization. This technique provides a strong platform for runtime optimization because the hot execution regions are extracted, optimized, and written to main memory for execution and because these regions persist across context switches. The current design of the framework supports a suite of optimizations, including partial function inlining (even into shared libraries), code straightening optimizations, loop unrolling, and peephole optimizations. Matthew C. Merten, Andrew R. Trick, Ronald D. Barnes, Erik M. Nystrom, Christopher N. George, John C. Gyllenhaal, Wen-Mei W. Hwu |
IEEE Trans. Computers | 6 |
| 1999 | A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime OptimizationabstractThis paper presents a novel hardware-based approach for identifying, profiling, and monitoring hot spots in order to support runtime optimization of general-purpose programs. The proposed approach consists of a set of tightly coupled hardware tables and control logic modules that are placed in the retirement stage of a processor pipeline removed from the critical path. The features of the proposed design include rapid detection of program hot spots after changes in execution behavior, runtime-tunable selection criteria for hot spot detection, and negligible overhead during application execution. Experiments using several SPEC95 benchmarks, as well as several large WindowsNT applications, demonstrate the promise of the proposed design. Matthew C. Merten, Andrew R. Trick, Christopher N. George, John C. Gyllenhaal, Wen-Mei W. Hwu |
ISCA | 4 |
| 1997 | Architectural Support for Compiler-Synthesized Dynamic Branch Prediction Strategies: Rationale and Initial ResultsabstractThis paper introduces a new architectural approach that supports compiler-synthesized dynamic branch predication. In compiler-synthesized dynamic branch prediction, the compiler generates code sequences that, when executed, digest relevant state information and execution statistics into a condition bit, or predicate. The hardware then utilizes this information to make predictions. Two categories of such architectures are proposed and evaluated. In Predicate Only Prediction (POP), the hardware simply uses the condition generated by the code sequence as a prediction. In Predicate Enhanced Prediction (PEP), the hardware uses the generated condition to enhance the accuracy of conventional branch prediction hardware. The IMPACT compiler currently provides a minimal level of compiler support for the proposed approach. Experiments based on current predicated code show that the proposed predictors achieve better performance than conventional branch predictors. Furthermore, they enable future compiler techniques which have the potential to achieve extremely high branch prediction accuracies. Several such compiler techniques are proposed in this paper. David I. August, Daniel A. Connors, John C. Gyllenhaal, Wen-Mei W. Hwu |
HPCA | 3 |
| 1996 | Optimization of Machine Descriptions for Efficient Use
John C. Gyllenhaal, Wen-Mei W. Hwu, Bob Rau |
MICRO | 1 |
| 1996 | Java Bytecode to Native Code Translation: The Caffeine Prototype and Preliminary ResultsabstractThe Java bytecode language is emerging as a software distribution standard. With major vendors committed to porting the Java run-time environment to their platforms, programs in Java bytecode are expected to run without modification on multiple platforms. These first generation runtime environments rely on an interpreter to bridge the gap between the bytecode instructions and the native hardware. This interpreter approach is sufficient for specialized applications such as Internet browsers where application performance is often limited by network delays rather than processor speed. It is, however not sufficient for executing general applications distributed in Java bytecode. This paper presents our initial prototyping experience with Caffeine, an optimizing translator from Java bytecode to native machine code. We discuss the major technical issues involved in stack to register mapping, run-time memory structure mapping and exception handlers. Encouraging initial results based on our X86 port are presented. Cheng-Hsueh A. Hsieh, John C. Gyllenhaal, Wen-Mei W. Hwu |
MICRO | 2 |
| 1995 | Compiler technology for future microprocessorsabstractAdvances in hardware technology have made it possible for microprocessors to execute a large number of instructions concurrently (i.e., in parallel). These microprocessors take advantage of the opportunity to execute instructions in parallel to increase the execution speed of a program. As in other forms of parallel processing, the performance of these microprocessors can vary greatly depending on the qualify of the software. In particular the quality of compilers can make an order of magnitude difference in performance. This paper presents a new generation of compiler technology that has emerged to deliver the large amount of instruction-level-parallelism that is already required by some current state-of-the-art microprocessors and will be required by more future microprocessors. We introduce critical components of the technology which deal with difficult problems that are encountered when compiling programs for a high degree of instruction-level-parallelism. We present examples to illustrate the functional requirements of these components. To provide more insight into the challenges involved, we present in-depth case studies on predicated compilation and maintenance of dependence information, two of the components that are largely missing from most current commercial compilers. Wen-Mei W. Hwu, Richard E. Hank, David M. Gallagher, Scott A. Mahlke, Daniel M. Lavery, Grant E. Haab, John C. Gyllenhaal, David I. August |
Proc. IEEE | 7 |
| 1994 | Dynamic Memory Disambiguation Using the Memory Conflict BufferabstractTo exploit instruction level parallelism, compilers for VLIW and superscalar processors often employ static code scheduling. However, the available code reordering may be severely restricted due to ambiguous dependences between memory instructions. This paper introduces a simple hardware mechanism, referred to as the memory conflict buffer, which facilitates static code scheduling in the presence of memory store/load dependences. Correct program execution is ensured by the memory conflict buffer and repair code provided by the compiler. With this addition, significant speedup over an aggressive code scheduling model can be achieved for both non-numerical and numerical programs. David M. Gallagher, William Y. Chen, Scott A. Mahlke, John C. Gyllenhaal, Wen-Mei W. Hwu |
ASPLOS | 4 |
| 1994 | Characterizing the impact of predicated execution on branch predictionabstractBranch instructions are recognized as a major impediment to exploiting instruction level parallelism. Even with sophisticated branch prediction techniques, many frequently executed branches remain difficult to predict. An architecture supporting predicated execution may allow the compiler to remove many of these hard-to-predict branches, reducing the number of branch mispredictions and thereby improving performance. We present an in-depth analysis of the characteristics of those branches which are frequently mispredicted and examine the effectiveness of an advanced compiler to eliminate these branches. Over the benchmarks studied, an average of 27% of the dynamic branches and 56% of the dynamic branch mispredictions are eliminated with predicated execution support. Scott A. Mahlke, Richard E. Hank, Roger A. Bringmann, John C. Gyllenhaal, David M. Gallagher, Wen-Mei W. Hwu |
MICRO | 4 |
| 1994 | Data relocation and prefetching for programs with large data setsabstractNumerical applications frequently contain nested loop structures that process large arrays of data. The execution of these loop structures often produces memory reference patterns that poorly utilize data caches. Limited associativity and cache capacity result in cache conflict misses. Also, non-unit stride access patterns can cause low utilization of cache lines. Data copying has been proposed and investigated in order to reduce cache conflict misses [1][2], but this technique has a high execution overhead since it performs the copy operations entirely in software. We propose a combined hardware and software technique called data relocation and prefetching which eliminates much of the overhead of data copying through the else of special hardware. Furthermore, by relocating the data while performing software prefetching, the overhead of copying the data can be reduced further. Experimental results for data relocation and prefetching are encouraging and show a large improvement in cache performance. Yoji Yamada, John C. Gyllenhaal, Grant E. Haab, Wen-Mei W. Hwu |
MICRO | 2 |
| 1993 | Speculative execution exception recovery using write-back suppressionabstractOne of the key design concerns of multiple instruction issue (MII) processors is deciding how many memory ports need to be provided, considering performance and efficiency of the target processor. For an MII processor that exploits instruction-level parallelism (ILP) in non-numerical code, this decision is difficult to make due to its irregularity. The authors perform an empirical study aimed at characterizing a suitable MII organization that best exploits irregular ILP. The study is based on the selective scheduling compiler that performs precise memory disambiguation for concurrent execution of multiple memory operations, along with renaming, speculation, and software pipelining. The result indicates that a small number of memory ports (i.e. less than half of the issue rate) is enough for exploiting most of irregular ILP. The authors also examine related issues such as the utilization of memory ports and additional data cache misses caused by speculative loads.> Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, John C. Gyllenhaal, Wen-Mei W. Hwu |
MICRO | 4 |
| 1993 | Superblock formation using static program analysisabstractTo achieve higher instruction-level parallelism, the constraint imposed by a single control flow must be relaxed. Control operations should execute in parallel just like data operations. We present a new software pipelining method called GPMB (Global Pipelining with Multiple Branches) which is based on architectures supporting multi-way branching and multiple control flows. Preliminary experimental results show that, for IFless loops, GPMB performs as well as modulo scheduling, and for branch-intensive loops, GPMB performs much better than software pipelining assuming the constraint of one two-way branch per cycle.> Richard E. Hank, Scott A. Mahlke, Roger A. Bringmann, John C. Gyllenhaal, Wen-Mei W. Hwu |
MICRO | 4 |
| 1992 | Code scheduling for VLIW/superscalar processors with limited register files
Tokuzo Kiyohara, John C. Gyllenhaal |
MICRO | 2 |
| 1992 | Compiler Code Transformations for Superscalar-Based High Performance SystemsabstractA set of compiler transformations designed to increase instruction-level parallelism is described. The effectiveness of these transformations is evaluated using 40 loop nests extracted from a range of supercomputer applications. This evaluation shows that increasing execution resources in superscalar/VLIW node processors yields little performance improvement unless loop unrolling and register renaming are applied. It also reveals that these two transformations are sufficient for DOALL loops. However, more advanced transformations are required in order for serial and DOACROSS loops to fully benefit from the increased execution resources. The results show that the six additional transformations studied satisfy most of this need.> Scott A. Mahlke, William Y. Chen, John C. Gyllenhaal, Wen-Mei W. Hwu |
SC | 3 |