EDBT 2026 Demo / reviewers in the wild / expert
Eric Sprangle
dblp:94/1807
· DBLP profile ↗
5ranked-venue papers
3as first author
0since 2021 · last 2008
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 54% Parallel and multicore computing · 27% GPUs and heterogeneous computing · 15% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2008 | Larrabee: a many-core x86 architecture for visual computing · ACM Trans. Graph. 2008 |
Runtime systems and virtual machines
language runtime |
0.1 | 1 | 2007 | Enabling scalability and performance in a large scale CMP environment · EuroSys 2007 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2007 | Enabling scalability and performance in a large scale CMP environment · EuroSys 2007 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2007 | Enabling scalability and performance in a large scale CMP environment · EuroSys 2007 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 2 | 2002 | Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002 The Agree Predictor: A Mechanism for Reducing Negative Branch History Interference · ISCA 1997 |
Processor architecture and microarchitecture › pipelining
pipeline depth |
0.0 | 1 | 2002 | Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002 |
Processor architecture and microarchitecture › branch prediction
two-level adaptive branch prediction |
0.0 | 1 | 1997 | The Agree Predictor: A Mechanism for Reducing Negative Branch History Interference · ISCA 1997 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 2 | 1997 | Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994 The Agree Predictor: A Mechanism for Reducing Negative Branch History Interference · ISCA 1997 |
Processor architecture and microarchitecture
register file |
0.0 | 1 | 1994 | Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994 |
Processor architecture and microarchitecture › out-of-order execution
register renaming |
0.0 | 1 | 1994 | Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994 |
Memory systems › cache design
cache sizing |
0.0 | 1 | 2002 | Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002 |
Memory systems › cache
on-chip cache |
0.0 | 1 | 2002 | Increasing Processor Performance by Implementing Deeper Pipelines · ISCA 2002 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.0 | 1 | 1994 | Facilitating superscalar processing via a combined static/dynamic register renaming scheme · MICRO 1994 |
Methods — techniques the papers use, named apart from their topics
experimental evaluation · 0.1vector processing · 0.1software task scheduling · 0.1binning · 0.1performance modeling · 0.0agree prediction · 0.0compiler-specified renaming tags · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | Larrabee: a many-core x86 architecture for visual computingabstractThis paper presents a many-core visual computing architecture code named Larrabee, a new software rendering pipeline, a manycore programming model, and performance analysis for several applications. Larrabee uses multiple in-order x86 CPU cores that are augmented by a wide vector processor unit, as well as some fixed function logic blocks. This provides dramatically higher performance per watt and per unit of area than out-of-order CPUs on highly parallel workloads. It also greatly increases the flexibility and programmability of the architecture as compared to standard GPUs. A coherent on-die 2 nd level cache allows efficient inter-processor communication and high-bandwidth local data access by CPU cores. Task scheduling is performed entirely with software in Larrabee, rather than in fixed function logic. The customizable software graphics rendering pipeline for this architecture uses binning in order to reduce required memory bandwidth, minimize lock contention, and increase opportunities for parallelism relative to standard GPUs. The Larrabee native programming model supports a variety of highly parallel applications that use irregular data structures. Performance analysis on those applications demonstrates Larrabee's potential for a broad range of parallel computation. Larry Seiler, Doug Carmean, Eric Sprangle, Tom Forsyth, Michael Abrash, Pradeep Dubey, Stephen Junkins, Adam T. Lake, Jeremy Sugerman, Robert Cavin, Roger Espasa, Ed Grochowski, Toni Juan, Pat Hanrahan |
ACM Trans. Graph. | 3 |
| 2007 | Enabling scalability and performance in a large scale CMP environmentabstractHardware trends suggest that large-scale CMP architectures, with tens to hundreds of processing cores on a single piece of silicon, are iminent within the next decade. While existing CMP machines have traditionally been handled in the same way as SMPs, this magnitude of parallelism introduces several fundamental challenges at the architectural level and this, in turn, translates to novel challenges in the design of the software stack for these platforms. This paper presents the "Many Core Run Time" (McRT), a software prototype of an integrated language runtime that was designed to explore configurations of the software stack for enabling performance and scalability on large scale CMP platforms. This paper presents the architecture of McRT and discusses our experiences with the system, including experimental evaluation that lead to several interesting, non-intuitive findings, providing key insights about the structure of the system stack at this scale. A key contribution of this paper is to demonstrate how McRT enables near linear improvements in performance and scalability for desktop workloads such as the popular XviD encoder and a set of RMS (recognition, mining, and synthesis) applications. Another key contribution of this work is its use of McRT to explore non-traditional system configurations such as a light-weight executive in which McRT runs on "bare metal" and replaces the traditional OS. Such configurations are becoming an increasingly attractive alternative to leverage heterogeneous computing uints as seen in today's CPU-GPU configurations. Bratin Saha, Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Mohan Rajagopalan, Richard L. Hudson, Leaf Petersen, Vijay Menon 0002, Brian R. Murphy, Tatiana Shpeisman, Eric Sprangle, Anwar Rohillah, Doug Carmean, Jesse Fang |
EuroSys | 10 |
| 2002 | Increasing Processor Performance by Implementing Deeper PipelinesabstractOne architectural method for increasing processor performance involves increasing the frequency by implementing deeper pipelines. This paper explores the relationship between performance and pipeline depth using a Pentium/sup (R)/ 4 processor like architecture as a baseline, and shows that deeper pipelines can continue to increase the performance. This paper shows that the branch misprediction latency is the single largest contributor to performance degradation as pipelines are stretched, and therefore branch prediction and fast branch recovery will continue to increase in importance. We show that higher performance cores, implemented with longer pipelines, for example, will put more pressure on the memory system, and therefore require larger on-chip caches. Finally, we show that in the same process technology, designing deeper pipelines can increase the processor frequency by 100%, which, when combined with larger on-chip caches can yield performance improvements of 35% to 90% over a Pentium 4 like processor. Eric Sprangle, Doug Carmean |
ISCA | 1 |
| 1997 | The Agree Predictor: A Mechanism for Reducing Negative Branch History InterferenceabstractDeeply pipelined, superscalar processors require accurate branch prediction to achieve high performance. Two-level branch predictors have been shown to achieve high prediction accuracy. It has also been shown that branch interference is a major contributor to the number of branches mispredicted by two-level predictors.This paper presents a new method to reduce the interference problem called agree prediction, which reduces the chance that two branches aliasing the same PHT entry will interfere negatively. We evaluate the performance of this scheme using full traces (both user and supervisor) of the SPECint95 benchmarks. The result is a reduction in the misprediction rate of gcc ranging from 8.62% with a 64K-entry PHT up to 33.3% with a 1K-entry PHT. Eric Sprangle, Robert Chappell, Mitch Alsup, Yale N. Patt |
ISCA | 1 |
| 1994 | Facilitating superscalar processing via a combined static/dynamic register renaming schemeabstractA superscalar implementation of a conventional instruction set architecture (ISA) requires N(N-1) comparators to determine dependencies between the N instructions issuing concurrently and 2N register file read ports to handle the 2 operands that each instruction can potentially source. On the other hand, if the compiler is allowed to specify part of the renaming tag, we show that we can eliminate the comparators needed to detect data dependencies between instructions issuing concurrently, and we can reduce the number of read ports from 16 to about 7 without losing performance. Finally, we show that this approach more efficiently implements predicated execution than can be done with a conventional ISA on a machine that renames registers. Eric Sprangle, Yale N. Patt |
MICRO | 1 |