Brian D. Koblenz

dblp:50/142 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 1998
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 61% Performance modeling and evaluation · 30% Interconnection networks and networks-on-chip · 9%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
memory latency tolerance
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Processor architecture and microarchitecture
multithreading
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Performance modeling and evaluation
parallel performance evaluation
0.011998
Multi-processor Performance on the Tera MTA · SC 1998
Compilers and program optimization › register allocation
graph coloring register allocation
0.011991
Register Allocation via Hierarchical Graph Coloring · PLDI 1991
Compilers and program optimization
register allocation
0.011991
Register Allocation via Hierarchical Graph Coloring · PLDI 1991

Methods — techniques the papers use, named apart from their topics

amdahl's law analysis · 0.0NAS benchmarks · 0.0hierarchical graph coloring · 0.0
YearPublicationVenuePosition
1998 Multi-processor Performance on the Tera MTA
abstract
The Tera MTA is a revolutionary commercial computer based on a multithreaded processor architecture. In contrast to many other parallel architectures, the Tera MTA can effectively use high amounts of parallelism on a single processor. By running multiple threads on a single processor, it can tolerate memory latency and to keep the processor saturated. If the computation is sufficiently large, it can benefit from running on multiple processors. A primary architectural goal of the MTA is that it provide scalable performance over multiple processors. This paper is a preliminary investigation of the first multi-processor Tera MTA. In a previous paper [1] we reported that on the kernel NAS 2 benchmarks [2], a single-processor MTA system running at the architected clock speed would be similar in performance to a single processor of the Cray T90. We found that the compilers of both machines were able to find the necessary threads or vector operations, after making standard changes to the random number generator. In this paper we update the single-processor results in two ways: we use only actual clock speeds, and we report improvements given by further tuning of the MTA codes. We then investigate the performance of the best single-processor codes when run on a two-processor MTA, making no further tuning effort. The parallel efficiency of the codes range from 77% to 99%. An analysis shows that the "serial bottlenecks" -- unparallelized code sections and the cost of allocating and freeing the parallel hardware resources -- account for less than a percent of the runtimes. Thus, Amdahl's Law needn't take effect on the NAS benchmarks until there are hundreds of processors running thousands of threads. Instead, the major source of inefficiency appears to be an imperfect network connecting the processors to the memory. Ideally, the network can support one memory reference per instruction. The current hardware has defects that reduce the throughput to about 85% of this rate. Except for the EP benchmark, the tuned codes issue memory references at nearly the peak rate of one per instruction. Consequently, the network can support the memory references issued by one, but not two, processors. As a result, the parallel efficiency of EP is near- perfect, but the others are reduced accordingly. Another reason for imperfect speedup pertains to the compiler. While the definition of a thread in a single processor or multi-processor mode is essentially the same, there is a different implementation and an associated overhead with running on multiple processors. We characterize the overhead of running "frays" (a collection of threads running on a single processor) and "crews" (a collection of frays, one per processor.)
Allan Snavely, Larry Carter, Jay Boisseau, Amitava Majumdar 0001, Kang Su Gatlin, Nick Mitchell, John Feo, Brian D. Koblenz
SC8
1992 Exploiting heterogeneous parallelism on a multithreaded multiprocessor
abstract
This paper describes an integrated architecture, compiler, runtime, and operating system solution to exploiting heterogeneous parallelism. The architecture is a pipelined multi-threaded multiprocessor, enabling the execution of very fine (multiple operations within an instruction) to very coarse (multiple jobs) parallel activities. The compiler and runtime focus on managing parallelism within a job, while the operating system focuses on managing parallelism across jobs. By considering the entire system in the design, we were able to smoothly interface its four components. While each component is primarily responsible for managing its own level of parallel activity, feedback mechanisms between components enable resource allocation and usage to be dynamically updated. This dynamic adaptation to changing requirements and available resources fosters both high utilization of the machine and the efficient expression and execution of parallelism.
Gail A. Alverson, Robert Alverson, David Callahan, Brian D. Koblenz, Allan Porterfield, Burton J. Smith
ICS4
1991 Register Allocation via Hierarchical Graph Coloring
abstract
We present a graph coloring register allocator de-signed to minimize the number of dynamic memory references. We cover the program with sets of blocks called tiles and group these tiles into a tree reflecting the program’s hierarchical control structure. Registers are allocated for each tile using standard graph coloring techniques and the local allocation and conflict information is passed around the tree in a two phase algorithm. This results in an allocation of reg-isters that is sensitive to local usage patterns while retaining a global perspective. Spill code is placed in less frequently executed portions of the program and the choice of variables to spill is based on usage pat-terns between the spills and the reloads rather than usage patterns over the entire program. 1
David Callahan, Brian D. Koblenz
PLDI2
1990 The Tera computer system
abstract
Article Free Access Share on The Tera computer system Authors: Robert Alverson Tera Computer Company, Seattle, Washington Tera Computer Company, Seattle, WashingtonView Profile , David Callahan Tera Computer Company, Seattle, Washington Tera Computer Company, Seattle, WashingtonView Profile , Daniel Cummings Tera Computer Company, Seattle, Washington Tera Computer Company, Seattle, WashingtonView Profile , Brian Koblenz Tera Computer Company, Seattle, Washington Tera Computer Company, Seattle, WashingtonView Profile , Allan Porterfield Tera Computer Company, Seattle, Washington Tera Computer Company, Seattle, WashingtonView Profile , Burton Smith Tera Computer Company, Seattle, Washington Tera Computer Company, Seattle, WashingtonView Profile Authors Info & Claims ICS '90: Proceedings of the 4th international conference on SupercomputingJune 1990 Pages 1–6https://doi.org/10.1145/77726.255132Online:01 June 1990Publication History 553citation2,281DownloadsMetricsTotal Citations553Total Downloads2,281Last 12 Months382Last 6 weeks81 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Robert Alverson, David Callahan, Daniel Cummings, Brian D. Koblenz, Allan Porterfield, Burton J. Smith
ICS4
1989 Constraint based vectorization
abstract
The constraint tree provides a uniform framework for representing many loop transformations. It allows us to estimate the performance of several alternative execution methods before committing to any of the transformations.
Brian D. Koblenz, William B. Noyce
ICS1