Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Eric Sprangle

dblp:94/1807 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 54% Parallel and multicore computing · 27% GPUs and heterogeneous computing · 15%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
many-core architecture
0.112008
Larrabee: a many-core x86 architecture for visual computing · ACM Trans. Graph. 2008
Runtime systems and virtual machines
language runtime
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Processor architecture and microarchitecture
chip multiprocessor
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Parallel and multicore computing
parallel programming runtimes
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Processor architecture and microarchitecture
branch prediction
0.122002
Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002
The Agree Predictor: A Mechanism for Reducing Negative Branch History Interference · ISCA 1997
Processor architecture and microarchitecture › pipelining
pipeline depth
0.012002
Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002
Processor architecture and microarchitecture › branch prediction
two-level adaptive branch prediction
0.011997
The Agree Predictor: A Mechanism for Reducing Negative Branch History Interference · ISCA 1997
Processor architecture and microarchitecture
superscalar processor
0.021997
Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994
The Agree Predictor: A Mechanism for Reducing Negative Branch History Interference · ISCA 1997
Processor architecture and microarchitecture
register file
0.011994
Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994
Processor architecture and microarchitecture › out-of-order execution
register renaming
0.011994
Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994
Memory systems › cache design
cache sizing
0.012002
Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002
Memory systems › cache
on-chip cache
0.012002
Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.011994
Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994

Methods — techniques the papers use, named apart from their topics

experimental evaluation · 0.1vector processing · 0.1software task scheduling · 0.1binning · 0.1performance modeling · 0.0agree prediction · 0.0compiler-specified renaming tags · 0.0
YearPublicationVenuePosition
2008 Larrabee: a many-core x86 architecture for visual computing
abstract
This paper presents a many-core visual computing architecture code named Larrabee, a new software rendering pipeline, a manycore programming model, and performance analysis for several applications. Larrabee uses multiple in-order x86 CPU cores that are augmented by a wide vector processor unit, as well as some fixed function logic blocks. This provides dramatically higher performance per watt and per unit of area than out-of-order CPUs on highly parallel workloads. It also greatly increases the flexibility and programmability of the architecture as compared to standard GPUs. A coherent on-die 2 nd level cache allows efficient inter-processor communication and high-bandwidth local data access by CPU cores. Task scheduling is performed entirely with software in Larrabee, rather than in fixed function logic. The customizable software graphics rendering pipeline for this architecture uses binning in order to reduce required memory bandwidth, minimize lock contention, and increase opportunities for parallelism relative to standard GPUs. The Larrabee native programming model supports a variety of highly parallel applications that use irregular data structures. Performance analysis on those applications demonstrates Larrabee's potential for a broad range of parallel computation.
Larry Seiler, Doug Carmean, Eric Sprangle, Tom Forsyth, Michael Abrash, Pradeep Dubey, Stephen Junkins, Adam T. Lake, Jeremy Sugerman, Robert Cavin, Roger Espasa, Ed Grochowski, Toni Juan, Pat Hanrahan
ACM Trans. Graph.3
2007 Enabling scalability and performance in a large scale CMP environment
abstract
Hardware trends suggest that large-scale CMP architectures, with tens to hundreds of processing cores on a single piece of silicon, are iminent within the next decade. While existing CMP machines have traditionally been handled in the same way as SMPs, this magnitude of parallelism introduces several fundamental challenges at the architectural level and this, in turn, translates to novel challenges in the design of the software stack for these platforms. This paper presents the "Many Core Run Time" (McRT), a software prototype of an integrated language runtime that was designed to explore configurations of the software stack for enabling performance and scalability on large scale CMP platforms. This paper presents the architecture of McRT and discusses our experiences with the system, including experimental evaluation that lead to several interesting, non-intuitive findings, providing key insights about the structure of the system stack at this scale. A key contribution of this paper is to demonstrate how McRT enables near linear improvements in performance and scalability for desktop workloads such as the popular XviD encoder and a set of RMS (recognition, mining, and synthesis) applications. Another key contribution of this work is its use of McRT to explore non-traditional system configurations such as a light-weight executive in which McRT runs on "bare metal" and replaces the traditional OS. Such configurations are becoming an increasingly attractive alternative to leverage heterogeneous computing uints as seen in today's CPU-GPU configurations.
Bratin Saha, Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Mohan Rajagopalan, Richard L. Hudson, Leaf Petersen, Vijay Menon 0002, Brian R. Murphy, Tatiana Shpeisman, Eric Sprangle, Anwar Rohillah, Doug Carmean, Jesse Fang
EuroSys10
2002 Increasing Processor Performance by Implementing Deeper Pipelines
abstract
One architectural method for increasing processor performance involves increasing the frequency by implementing deeper pipelines. This paper explores the relationship between performance and pipeline depth using a Pentium/sup (R)/ 4 processor like architecture as a baseline, and shows that deeper pipelines can continue to increase the performance. This paper shows that the branch misprediction latency is the single largest contributor to performance degradation as pipelines are stretched, and therefore branch prediction and fast branch recovery will continue to increase in importance. We show that higher performance cores, implemented with longer pipelines, for example, will put more pressure on the memory system, and therefore require larger on-chip caches. Finally, we show that in the same process technology, designing deeper pipelines can increase the processor frequency by 100%, which, when combined with larger on-chip caches can yield performance improvements of 35% to 90% over a Pentium 4 like processor.
Eric Sprangle, Doug Carmean
ISCA1
1997 The Agree Predictor: A Mechanism for Reducing Negative Branch History Interference
abstract
Deeply pipelined, superscalar processors require accurate branch prediction to achieve high performance. Two-level branch predictors have been shown to achieve high prediction accuracy. It has also been shown that branch interference is a major contributor to the number of branches mispredicted by two-level predictors.This paper presents a new method to reduce the interference problem called agree prediction, which reduces the chance that two branches aliasing the same PHT entry will interfere negatively. We evaluate the performance of this scheme using full traces (both user and supervisor) of the SPECint95 benchmarks. The result is a reduction in the misprediction rate of gcc ranging from 8.62% with a 64K-entry PHT up to 33.3% with a 1K-entry PHT.
Eric Sprangle, Robert Chappell, Mitch Alsup, Yale N. Patt
ISCA1
1994 Facilitating superscalar processing via a combined static/dynamic register renaming scheme
abstract
A superscalar implementation of a conventional instruction set architecture (ISA) requires N(N-1) comparators to determine dependencies between the N instructions issuing concurrently and 2N register file read ports to handle the 2 operands that each instruction can potentially source. On the other hand, if the compiler is allowed to specify part of the renaming tag, we show that we can eliminate the comparators needed to detect data dependencies between instructions issuing concurrently, and we can reduce the number of read ports from 16 to about 7 without losing performance. Finally, we show that this approach more efficiently implements predicated execution than can be done with a conventional ISA on a machine that renames registers.
Eric Sprangle, Yale N. Patt
MICRO1