Andrew R. Trick

dblp:49/345 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2001
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 40% Parallel and multicore computing · 38% Electronic design automation · 11%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
runtime optimization
0.132001
An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001
A hardware mechanism for dynamic extraction and relayout of program hot spots · ISCA 2000
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999
Processor architecture and microarchitecture
branch prediction
0.012001
An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001
Processor architecture and microarchitecture
instruction fetch
0.012000
A hardware mechanism for dynamic extraction and relayout of program hot spots · ISCA 2000
Processor architecture and microarchitecture › instruction fetch
trace cache
0.012000
A hardware mechanism for dynamic extraction and relayout of program hot spots · ISCA 2000
Performance modeling and evaluation › profiling
hardware profiling
0.011999
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999
Electronic design automation › physical design › lithography
lithography hotspot detection
0.011999
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999
Compilers and program optimization
dynamic optimization
0.022001
An Architectural Framework for Runtime Optimization · IEEE Trans. Computers 2001
A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization · ISCA 1999

Methods — techniques the papers use, named apart from their topics

out-of-order execution · 0.1instruction profiling · 0.1hardware speculation · 0.1branch target buffer · 0.1hardware performance counters · 0.0
YearPublicationVenuePosition
2001 An Architectural Framework for Runtime Optimization
abstract
Wide-issue processors continue to achieve higher performance by exploiting greater instruction-level parallelism. Dynamic techniques such as out-of-order execution and hardware speculation have proven effective at increasing instruction throughput. Runtime optimization promises to provide an even higher level of performance by adaptively applying aggressive code transformations on a larger scope. This paper presents a new hardware mechanism for generating and deploying runtime optimized code. The mechanism can be viewed as a filtering system that resides in the retirement stage of the processor pipeline, accepts an instruction execution stream as input, and produces instruction profiles and sets of linked, optimized traces as output. The code deployment mechanism uses an extension to the branch prediction mechanism to migrate execution into the new code without modifying the original code. These new components do not add delay to the execution of the program except during short bursts of reoptimization. This technique provides a strong platform for runtime optimization because the hot execution regions are extracted, optimized, and written to main memory for execution and because these regions persist across context switches. The current design of the framework supports a suite of optimizations, including partial function inlining (even into shared libraries), code straightening optimizations, loop unrolling, and peephole optimizations.
Matthew C. Merten, Andrew R. Trick, Ronald D. Barnes, Erik M. Nystrom, Christopher N. George, John C. Gyllenhaal, Wen-Mei W. Hwu
IEEE Trans. Computers2
2000 A hardware mechanism for dynamic extraction and relayout of program hot spots
abstract
This paper presents a new mechanism for collecting and deploying runtime optimized code. The code-collecting component resides in the instruction retirement stage and lays out hot execution paths to improve instruction fetch rate as well as enable further code optimization. The code deployment component uses an extension to the Branch Target Buffer to migrate execution into the new code without modifying the original code. No significant delay is added to the total execution of the program due to these components. The code collection scheme enables safe runtime optimization along paths that span function boundaries. This technique provides a better platform for runtime optimization than trace caches, because the traces are longer and persist in main memory across context switches. Additionally, these traces are not as susceptible to transient behavior because they are restricted to frequently executed code. Empirical results show that on average this mechanism can achieve better instruction fetch rates using only 12KB of hardware than a trace cache requiring 15KB of hardware, while producing long, persistent traces more suited to optimization.
Matthew C. Merten, Andrew R. Trick, Erik M. Nystrom, Ronald D. Barnes, Wen-Mei W. Hwu
ISCA2
1999 A Hardware-Driven Profiling Scheme for Identifying Program Hot Spots to Support Runtime Optimization
abstract
This paper presents a novel hardware-based approach for identifying, profiling, and monitoring hot spots in order to support runtime optimization of general-purpose programs. The proposed approach consists of a set of tightly coupled hardware tables and control logic modules that are placed in the retirement stage of a processor pipeline removed from the critical path. The features of the proposed design include rapid detection of program hot spots after changes in execution behavior, runtime-tunable selection criteria for hot spot detection, and negligible overhead during application execution. Experiments using several SPEC95 benchmarks, as well as several large WindowsNT applications, demonstrate the promise of the proposed design.
Matthew C. Merten, Andrew R. Trick, Christopher N. George, John C. Gyllenhaal, Wen-Mei W. Hwu
ISCA2