Brian A. Fields

dblp:25/1767 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
0since 2021 · last 2004
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Performance modeling and evaluation · 51% Processor architecture and microarchitecture · 37% Electronic design automation · 9%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction scheduling
0.122002
Slack: Maximizing Performance Under Technological Constraints · ISCA 2002
Focusing processor policies via critical-path prediction · ISCA 2001
Performance modeling and evaluation
bottleneck analysis
0.012004
Interaction cost and shotgun profiling · ACM Trans. Archit. Code Optim. 2004
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.012004
Interaction cost and shotgun profiling · ACM Trans. Archit. Code Optim. 2004
Performance modeling and evaluation › profiling
microarchitectural profiling
0.012004
Interaction cost and shotgun profiling · ACM Trans. Archit. Code Optim. 2004
Electronic design automation › timing prediction
critical path prediction
0.012001
Focusing processor policies via critical-path prediction · ISCA 2001
Processor architecture and microarchitecture › speculation
speculation control
0.012001
Focusing processor policies via critical-path prediction · ISCA 2001
Energy-efficient computing
power management
0.012002
Slack: Maximizing Performance Under Technological Constraints · ISCA 2002
Processor architecture and microarchitecture
clustered architecture
0.012001
Focusing processor policies via critical-path prediction · ISCA 2001
Processor architecture and microarchitecture › clustered architecture
instruction steering
0.012001
Focusing processor policies via critical-path prediction · ISCA 2001

Methods — techniques the papers use, named apart from their topics

dependence graph construction · 0.0performance monitoring · 0.0slack prediction · 0.0control policy · 0.0value prediction · 0.0token-passing algorithm · 0.0dependence-graph model · 0.0
YearPublicationVenuePosition
2004 Interaction cost and shotgun profiling
abstract
We observe that the challenges software optimizers and microarchitects face every day boil down to a single problem: bottleneck analysis. A bottleneck is any event or resource that contributes to execution time, such as a critical cache miss or window stall. Tasks such as tuning processors for energy efficiency and finding the right loads to prefetch all require measuring the performance costs of bottlenecks.In the past, simple event counts were enough to find the important bottlenecks. Today, the parallelism of modern processors makes such analysis much more difficult, rendering traditional performance counters less useful. If two microarchitectural events (such as a fetch stall and a cache miss) occur in the same cycle, which event should we blame for the cycle? What cost should we assign to each event? In this paper, we introduce a new model for understanding event costs to facilitate processor design and optimization.First, we observe that all instructions, hardware structures, and events in a machine can interact in only one of two ways (in parallel or serially). We quantify these interactions by defining interaction cost , which can be zero (independent, no interaction), positive (parallel), or negative (serial).Second, we illustrate the value of using interaction costs in processor design and optimization. In a processor with a long pipeline, we show how to mitigate the negative performance effect of long latency "critical" loops, such as the level-one cache access and issue-wakeup, by optimizing seemingly unrelated resources that interact with them.Finally, we propose shotgun profiling , a class of hardware profiling infrastructures that are parallelism-aware, in contrast to traditional event counters. Our recommended design requires only modest extensions to current hardware counters, while enabling the construction of full-featured dependence graphs of the microexecution. With these dependence graphs, many types of analyses can be performed, including identifying critical instructions, finding slack, as well as computing costs and interaction costs.
Brian A. Fields, Rastislav Bodík, Mark D. Hill, Chris J. Newburn
ACM Trans. Archit. Code Optim.1
2003 Using Interaction Costs for Microarchitectural Bottleneck Analysis
abstract
Attacking bottlenecks in modern processors is difficult because many microarchitectural events overlap with each other. This parallelism makes it difficult to both: (a) assign a cost to an event (e.g., to one of two overlapping cache misses); and (b) assign blame for each cycle (e.g., for a cycle where many, overlapping resources are active). This paper introduces a new model for understanding event costs to facilitate processor design and optimization. First, we observe that everything in a machine (instructions, hardware structures, events) can interact in only one of two ways (in parallel or serially). We quantify these interactions by defining interaction cost, which can be zero (independent, no interaction), positive (parallel), or negative (serial). Second, we illustrate the value of using interaction costs in processor design and optimization. Finally, we propose performance-monitoring hardware for measuring interaction costs that is suitable for modern processors.
Brian A. Fields, Rastislav Bodík, Mark D. Hill, Chris J. Newburn
MICRO1
2002 Slack: Maximizing Performance Under Technological Constraints
abstract
Many emerging processor microarchitectures seek to manage technological constraints (e.g., wire, delay, power, and circuit complexity) by resorting to non-uniform designs that provide resources at multiple quality levels (e.g., fast/slow bypass paths, multi-speed functional units, and grid architectures). In such designs, the constraint problem becomes a control problem, and the challenge becomes designing a control policy that mitigates the performance penalty of the non-uniformity. Given the increasing importance of non-uniform control policies, we believe it is appropriate to examine them, in their own right. To this end, we develop slack for use in creating control policies that match program execution behavior to machine design. Intuitively, the slack of a dynamic instruction i is the number of cycles i can be delayed with no effect on execution time. This property makes slack a natural candidate for hiding non-uniform latencies. We make three contributions in our exploration of slack. First, we formally define slack, distinguish three variants (local, global and apportioned), and perform a limit study to show that slack is prevalent in our SPEC2000 workload. Second, we show how to predict slack in hardware. Third, we illustrate how to create a control policy based on slack for steering instructions among fast (high power) and slow (lower power) pipelines.
Brian A. Fields, Rastislav Bodík, Mark D. Hill
ISCA1
2001 Focusing processor policies via critical-path prediction
abstract
Although some instructions hurt performance more than others, current processors typically apply scheduling and speculation as if each instruction was equally costly. Instruction cost can be naturally expressed through the critical path: if we could predict it at run-time, egalitarian policies could be replaced with cost-sensitive strategies that will grow increasingly effective as processors become more parallel. This paper introduces a hardware predictor of instruction criticality and uses it to improve performance. The predictor is both effective and simple in its hardware implementation. The effectiveness at improving performance stems from us-ing a dependence-graph model of the microarchitectural critical path that identifies execution bottlenecks by incorporating both data and machine-specific dependences. The simplicity stems from a token-passing algorithm that computes the critical path without actually building the dependence graph. By focusing processor policies on critical instructions, our predictor enables a large class of optimizations. It can (i) give priority to critical instructions for scarce resources (functional units, ports, predictor entries); and (ii) suppress speculation on non-critical instructions, thus reducing "useless" misspecula-tions. We present two case studies that illustrate the potential of the two types of optimization, we show that (i) critical-path-based dynamic instruction scheduling and steering in a clus-tered architecture improves performance by as much as 21% (10 % on average); and (ii) focusing value prediction only on critical instructions improves performance by as much as 5%, due to removing nearly half of the misspeculations.
Brian A. Fields, Shai Rubin, Rastislav Bodík
ISCA1