Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

John C. Gyllenhaal

dblp:50/3479 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
13 papers
Parallel and multicore computing · 51% Processor architecture and microarchitecture · 19% High-performance computing · 17%
Software engineering, system software, and programming languages
10 papers
Compilers and program optimization · 87% Runtime systems and virtual machines · 13%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
performance optimization at scale
0.112012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Parallel and multicore computing
synchronization
0.112012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Parallel and multicore computing
transactional memory
0.112012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Processor architecture and microarchitecture
branch prediction
0.132001
An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001
Architectural Support for Compiler-Synthesized Dynamic Branch Prediction Strategies: Rationale and Initial Results · HPCA 1997
Characterizing the impact of predicated execution on branch prediction · MICRO 1994
Parallel and multicore computing
runtime optimization
0.122001
An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.012012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Parallel and multicore computing
parallel programming models
0.012012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Processor architecture and microarchitecture
instruction-level parallelism
0.041995
Characterizing the impact of predicated execution on branch prediction · MICRO 1994
Dynamic Memory Disambiguation Using the Memory Conflict Buffer · ASPLOS 1994
Compiler technology for future microprocessors · Proc. IEEE 1995
Performance modeling and evaluation › profiling
hardware profiling
0.011999
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999
Electronic design automation › physical design › lithography
lithography hotspot detection
0.011999
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.021995
Compiler technology for future microprocessors · Proc. IEEE 1995
Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.021997
Characterizing the impact of predicated execution on branch prediction · MICRO 1994
Architectural Support for Compiler-Synthesized Dynamic Branch Prediction Strategies: Rationale and Initial Results · HPCA 1997
Compilers and program optimization
dynamic optimization
0.022001
An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999
Runtime systems and virtual machines › binary translation
bytecode translation
0.011996
Java Bytecode to Native Code Translation: The Caffeine Prototype and Preliminary Results · MICRO 1996
Memory systems
cache
0.021994
Data relocation and prefetching for programs with large data sets · MICRO 1994
Speculative execution exception recovery using write-back suppression · MICRO 1993
Compilers and program optimization
dependence analysis
0.011995
Compiler technology for future microprocessors · Proc. IEEE 1995
Compilers and program optimization
predicated compilation
0.011995
Compiler technology for future microprocessors · Proc. IEEE 1995
Memory systems › cache
conflict miss reduction
0.011994
Data relocation and prefetching for programs with large data sets · MICRO 1994
Memory systems › cache › prefetching
data prefetching
0.011994
Data relocation and prefetching for programs with large data sets · MICRO 1994
Processor architecture and microarchitecture
superscalar processor
0.021992
Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992
Code scheduling for VLIW/superscalar processors with limited register files · MICRO 1992
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.021992
Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992
Code scheduling for VLIW/superscalar processors with limited register files · MICRO 1992
Compilers and program optimization › instruction scheduling
software pipelining
0.011993
Superblock formation using static program analysis · MICRO 1993
Processor architecture and microarchitecture › instruction-level parallelism
multiple instruction issue
0.011993
Speculative execution exception recovery using write-back suppression · MICRO 1993
Processor architecture and microarchitecture
speculative execution
0.011993
Speculative execution exception recovery using write-back suppression · MICRO 1993
Compilers and program optimization › program transformation
compiler transformations
0.011992
Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992
Compilers and program optimization
instruction scheduling
0.011992
Code scheduling for VLIW/superscalar processors with limited register files · MICRO 1992
Processor architecture and microarchitecture
instruction set architecture
0.021996
Optimization of Machine Descriptions for Efficient Use · MICRO 1996
Superblock formation using static program analysis · MICRO 1993
Runtime systems and virtual machines › managed runtime
java runtime
0.011996
Java Bytecode to Native Code Translation: The Caffeine Prototype and Preliminary Results · MICRO 1996
Memory systems › data movement
data relocation
0.011994
Data relocation and prefetching for programs with large data sets · MICRO 1994
Memory systems › cache › CPU cache
data cache
0.011993
Speculative execution exception recovery using write-back suppression · MICRO 1993

Methods — techniques the papers use, named apart from their topics

performance characterization · 0.1best practices · 0.1out-of-order execution · 0.1instruction profiling · 0.1hardware speculation · 0.1hardware performance counters · 0.0predicate generation · 0.0compiler synthesis · 0.0hardware memory conflict buffer · 0.0compiler repair code · 0.0stack-to-register mapping · 0.0exception handling · 0.0
YearPublicationVenuePosition
2012 What scientific applications can benefit from hardware transactional memory?
abstract
Achieving efficient and correct synchronization of multiple threads is a difficult and error-prone task at small scale and, as we march towards extreme scale computing, will be even more challenging when the resulting application is supposed to utilize millions of cores efficiently. Transactional Memory (TM) is a promising technique to ease the burden on the programmer, but only recently has become available on commercial hardware in the new Blue Gene/Q system and hence the real benefit for realistic applications has not been studied yet. This paper presents the first performance results of TM embedded into OpenMP on a prototype system of BG/Q and characterizes code properties that will likely lead to benefits when augmented with TM primitives. We first study the influence of thread count, environment variables and memory layout on TM performance and identify code properties that will yield performance gains with TM. Second, we evaluate the combination of OpenMP with multiple synchronization primitives on top of MPI to determine suitable task to thread ratios per node. Finally, we condense our findings into a set of best practices. These are applied to a Monte Carlo Benchmark and a Smoothed Particle Hydrodynamics method. In both cases an optimized TM version, executed with 64 threads on one node, outperforms a simple TM implementation. MCB with optimized TM yields a speedup of 27.45 over baseline.
Martin Schindewolf, Barna L. Bihari, John C. Gyllenhaal, Martin Schulz 0001, Amy Wang, Wolfgang Karl
SC3
2001 An Architectural Framework for Runtime Optimization
abstract
Wide-issue processors continue to achieve higher performance by exploiting greater instruction-level parallelism. Dynamic techniques such as out-of-order execution and hardware speculation have proven effective at increasing instruction throughput. Runtime optimization promises to provide an even higher level of performance by adaptively applying aggressive code transformations on a larger scope. This paper presents a new hardware mechanism for generating and deploying runtime optimized code. The mechanism can be viewed as a filtering system that resides in the retirement stage of the processor pipeline, accepts an instruction execution stream as input, and produces instruction profiles and sets of linked, optimized traces as output. The code deployment mechanism uses an extension to the branch prediction mechanism to migrate execution into the new code without modifying the original code. These new components do not add delay to the execution of the program except during short bursts of reoptimization. This technique provides a strong platform for runtime optimization because the hot execution regions are extracted, optimized, and written to main memory for execution and because these regions persist across context switches. The current design of the framework supports a suite of optimizations, including partial function inlining (even into shared libraries), code straightening optimizations, loop unrolling, and peephole optimizations.
Matthew C. Merten, Andrew R. Trick, Ronald D. Barnes, Erik M. Nystrom, Christopher N. George, John C. Gyllenhaal, Wen-Mei W. Hwu
IEEE Trans. Computers6
1999 A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization
abstract
This paper presents a novel hardware-based approach for identifying, profiling, and monitoring hot spots in order to support runtime optimization of general-purpose programs. The proposed approach consists of a set of tightly coupled hardware tables and control logic modules that are placed in the retirement stage of a processor pipeline removed from the critical path. The features of the proposed design include rapid detection of program hot spots after changes in execution behavior, runtime-tunable selection criteria for hot spot detection, and negligible overhead during application execution. Experiments using several SPEC95 benchmarks, as well as several large WindowsNT applications, demonstrate the promise of the proposed design.
Matthew C. Merten, Andrew R. Trick, Christopher N. George, John C. Gyllenhaal, Wen-Mei W. Hwu
ISCA4
1997 Architectural Support for Compiler-Synthesized Dynamic Branch Prediction Strategies: Rationale and Initial Results
abstract
This paper introduces a new architectural approach that supports compiler-synthesized dynamic branch predication. In compiler-synthesized dynamic branch prediction, the compiler generates code sequences that, when executed, digest relevant state information and execution statistics into a condition bit, or predicate. The hardware then utilizes this information to make predictions. Two categories of such architectures are proposed and evaluated. In Predicate Only Prediction (POP), the hardware simply uses the condition generated by the code sequence as a prediction. In Predicate Enhanced Prediction (PEP), the hardware uses the generated condition to enhance the accuracy of conventional branch prediction hardware. The IMPACT compiler currently provides a minimal level of compiler support for the proposed approach. Experiments based on current predicated code show that the proposed predictors achieve better performance than conventional branch predictors. Furthermore, they enable future compiler techniques which have the potential to achieve extremely high branch prediction accuracies. Several such compiler techniques are proposed in this paper.
David I. August, Daniel A. Connors, John C. Gyllenhaal, Wen-Mei W. Hwu
HPCA3
1996 Optimization of Machine Descriptions for Efficient Use
John C. Gyllenhaal, Wen-Mei W. Hwu, Bob Rau
MICRO1
1996 Java Bytecode to Native Code Translation: The Caffeine Prototype and Preliminary Results
abstract
The Java bytecode language is emerging as a software distribution standard. With major vendors committed to porting the Java run-time environment to their platforms, programs in Java bytecode are expected to run without modification on multiple platforms. These first generation runtime environments rely on an interpreter to bridge the gap between the bytecode instructions and the native hardware. This interpreter approach is sufficient for specialized applications such as Internet browsers where application performance is often limited by network delays rather than processor speed. It is, however not sufficient for executing general applications distributed in Java bytecode. This paper presents our initial prototyping experience with Caffeine, an optimizing translator from Java bytecode to native machine code. We discuss the major technical issues involved in stack to register mapping, run-time memory structure mapping and exception handlers. Encouraging initial results based on our X86 port are presented.
Cheng-Hsueh A. Hsieh, John C. Gyllenhaal, Wen-Mei W. Hwu
MICRO2
1995 Compiler technology for future microprocessors
abstract
Advances in hardware technology have made it possible for microprocessors to execute a large number of instructions concurrently (i.e., in parallel). These microprocessors take advantage of the opportunity to execute instructions in parallel to increase the execution speed of a program. As in other forms of parallel processing, the performance of these microprocessors can vary greatly depending on the qualify of the software. In particular the quality of compilers can make an order of magnitude difference in performance. This paper presents a new generation of compiler technology that has emerged to deliver the large amount of instruction-level-parallelism that is already required by some current state-of-the-art microprocessors and will be required by more future microprocessors. We introduce critical components of the technology which deal with difficult problems that are encountered when compiling programs for a high degree of instruction-level-parallelism. We present examples to illustrate the functional requirements of these components. To provide more insight into the challenges involved, we present in-depth case studies on predicated compilation and maintenance of dependence information, two of the components that are largely missing from most current commercial compilers.
Wen-Mei W. Hwu, Richard E. Hank, David M. Gallagher, Scott A. Mahlke, Daniel M. Lavery, Grant E. Haab, John C. Gyllenhaal, David I. August
Proc. IEEE7
1994 Dynamic Memory Disambiguation Using the Memory Conflict Buffer
abstract
To exploit instruction level parallelism, compilers for VLIW and superscalar processors often employ static code scheduling. However, the available code reordering may be severely restricted due to ambiguous dependences between memory instructions. This paper introduces a simple hardware mechanism, referred to as the memory conflict buffer, which facilitates static code scheduling in the presence of memory store/load dependences. Correct program execution is ensured by the memory conflict buffer and repair code provided by the compiler. With this addition, significant speedup over an aggressive code scheduling model can be achieved for both non-numerical and numerical programs.
David M. Gallagher, William Y. Chen, Scott A. Mahlke, John C. Gyllenhaal, Wen-Mei W. Hwu
ASPLOS4
1994 Characterizing the impact of predicated execution on branch prediction
abstract
Branch instructions are recognized as a major impediment to exploiting instruction level parallelism. Even with sophisticated branch prediction techniques, many frequently executed branches remain difficult to predict. An architecture supporting predicated execution may allow the compiler to remove many of these hard-to-predict branches, reducing the number of branch mispredictions and thereby improving performance. We present an in-depth analysis of the characteristics of those branches which are frequently mispredicted and examine the effectiveness of an advanced compiler to eliminate these branches. Over the benchmarks studied, an average of 27% of the dynamic branches and 56% of the dynamic branch mispredictions are eliminated with predicated execution support.
Scott A. Mahlke, Richard E. Hank, Roger A. Bringmann, John C. Gyllenhaal, David M. Gallagher, Wen-Mei W. Hwu
MICRO4
1994 Data relocation and prefetching for programs with large data sets
abstract
Numerical applications frequently contain nested loop structures that process large arrays of data. The execution of these loop structures often produces memory reference patterns that poorly utilize data caches. Limited associativity and cache capacity result in cache conflict misses. Also, non-unit stride access patterns can cause low utilization of cache lines. Data copying has been proposed and investigated in order to reduce cache conflict misses [1][2], but this technique has a high execution overhead since it performs the copy operations entirely in software. We propose a combined hardware and software technique called data relocation and prefetching which eliminates much of the overhead of data copying through the else of special hardware. Furthermore, by relocating the data while performing software prefetching, the overhead of copying the data can be reduced further. Experimental results for data relocation and prefetching are encouraging and show a large improvement in cache performance.
Yoji Yamada, John C. Gyllenhaal, Grant E. Haab, Wen-Mei W. Hwu
MICRO2
1993 Speculative execution exception recovery using write-back suppression
abstract
One of the key design concerns of multiple instruction issue (MII) processors is deciding how many memory ports need to be provided, considering performance and efficiency of the target processor. For an MII processor that exploits instruction-level parallelism (ILP) in non-numerical code, this decision is difficult to make due to its irregularity. The authors perform an empirical study aimed at characterizing a suitable MII organization that best exploits irregular ILP. The study is based on the selective scheduling compiler that performs precise memory disambiguation for concurrent execution of multiple memory operations, along with renaming, speculation, and software pipelining. The result indicates that a small number of memory ports (i.e. less than half of the issue rate) is enough for exploiting most of irregular ILP. The authors also examine related issues such as the utilization of memory ports and additional data cache misses caused by speculative loads.>
Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, John C. Gyllenhaal, Wen-Mei W. Hwu
MICRO4
1993 Superblock formation using static program analysis
abstract
To achieve higher instruction-level parallelism, the constraint imposed by a single control flow must be relaxed. Control operations should execute in parallel just like data operations. We present a new software pipelining method called GPMB (Global Pipelining with Multiple Branches) which is based on architectures supporting multi-way branching and multiple control flows. Preliminary experimental results show that, for IFless loops, GPMB performs as well as modulo scheduling, and for branch-intensive loops, GPMB performs much better than software pipelining assuming the constraint of one two-way branch per cycle.>
Richard E. Hank, Scott A. Mahlke, Roger A. Bringmann, John C. Gyllenhaal, Wen-Mei W. Hwu
MICRO4
1992 Code scheduling for VLIW/superscalar processors with limited register files
Tokuzo Kiyohara, John C. Gyllenhaal
MICRO2
1992 Compiler Code Transformations for Superscalar-Based High Performance Systems
abstract
A set of compiler transformations designed to increase instruction-level parallelism is described. The effectiveness of these transformations is evaluated using 40 loop nests extracted from a range of supercomputer applications. This evaluation shows that increasing execution resources in superscalar/VLIW node processors yields little performance improvement unless loop unrolling and register renaming are applied. It also reveals that these two transformations are sufficient for DOALL loops. However, more advanced transformations are required in order for serial and DOACROSS loops to fully benefit from the increased execution resources. The results show that the six additional transformations studied satisfy most of this need.>
Scott A. Mahlke, William Y. Chen, John C. Gyllenhaal, Wen-Mei W. Hwu
SC3