VLDB 2026 Research / reviewers in the wild / expert
Jeremy Sugerman
dblp:67/2242
· DBLP profile ↗
7ranked-venue papers
2as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorSystems, architecture and hardware · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 32% GPUs and heterogeneous computing · 28% Cloud and datacenter computing · 21% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% | |
| Software engineering, system software, and programming languages
2 papers |
Operating systems · 69% Compilers and program optimization · 31% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
virtualization |
0.2 | 2 | 2012 | Bringing Virtualization to the x86 Architecture with the Original VMware Workstation · ACM Trans. Comput. Syst. 2012 Virtualizing I/O Devices on VMware Workstation's Hosted Virtual Machine Monitor · USENIX ATC, General Track 2001 |
Processor architecture and microarchitecture › binary translation
dynamic binary translation |
0.1 | 1 | 2012 | Bringing Virtualization to the x86 Architecture with the Original VMware Workstation · ACM Trans. Comput. Syst. 2012 |
Rendering
real-time rendering |
0.1 | 1 | 2009 | GRAMPS: A programming model for graphics pipelines · ACM Trans. Graph. 2009 |
GPUs and heterogeneous computing › GPU rendering
programmable graphics pipeline |
0.1 | 1 | 2009 | GRAMPS: A programming model for graphics pipelines · ACM Trans. Graph. 2009 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2008 | Larrabee: a many-core x86 architecture for visual computing · ACM Trans. Graph. 2008 |
Parallel and multicore computing › parallel algorithms › parallel primitives
data-parallel primitives |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
GPUs and heterogeneous computing
GPU computing |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
GPUs and heterogeneous computing › GPU programming
GPU programming models |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
Distributed systems
stream processing |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 2012 | Bringing Virtualization to the x86 Architecture with the Original VMware Workstation · ACM Trans. Comput. Syst. 2012 |
Processor architecture and microarchitecture › instruction set architecture › CISC
x86 |
0.0 | 1 | 2012 | Bringing Virtualization to the x86 Architecture with the Original VMware Workstation · ACM Trans. Comput. Syst. 2012 |
Operating systems › virtualization
hypervisor |
0.0 | 1 | 2001 | Virtualizing I/O Devices on VMware Workstation's Hosted Virtual Machine Monitor · USENIX ATC, General Track 2001 |
Cloud and datacenter computing › virtualization
i/o virtualization |
0.0 | 1 | 2001 | Virtualizing I/O Devices on VMware Workstation's Hosted Virtual Machine Monitor · USENIX ATC, General Track 2001 |
Compilers and program optimization › accelerator compilation
GPU compiler |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
Methods — techniques the papers use, named apart from their topics
queue-based data exchange · 0.2pipeline scheduling · 0.2trap-and-emulate · 0.1partial evaluation · 0.1dynamic binary translation · 0.1adaptive retranslation · 0.1compiler and runtime abstraction · 0.1vector processing · 0.1software task scheduling · 0.1binning · 0.1stream programming · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Bringing Virtualization to the x86 Architecture with the Original VMware WorkstationabstractThis article describes the historical context, technical challenges, and main implementation techniques used by VMware Workstation to bring virtualization to the x86 architecture in 1999. Although virtual machine monitors (VMMs) had been around for decades, they were traditionally designed as part of monolithic, single-vendor architectures with explicit support for virtualization. In contrast, the x86 architecture lacked virtualization support, and the industry around it had disaggregated into an ecosystem, with different vendors controlling the computers, CPUs, peripherals, operating systems, and applications, none of them asking for virtualization. We chose to build our solution independently of these vendors. As a result, VMware Workstation had to deal with new challenges associated with (i) the lack of virtualization support in the x86 architecture, (ii) the daunting complexity of the architecture itself, (iii) the need to support a broad combination of peripherals, and (iv) the need to offer a simple user experience within existing environments. These new challenges led us to a novel combination of well-known virtualization techniques, techniques from other domains, and new techniques. VMware Workstation combined a hosted architecture with a VMM. The hosted architecture enabled a simple user experience and offered broad hardware compatibility. Rather than exposing I/O diversity to the virtual machines, VMware Workstation also relied on software emulation of I/O devices. The VMM combined a trap-and-emulate direct execution engine with a system-level dynamic binary translator to efficiently virtualize the x86 architecture and support most commodity operating systems. By relying on x86 hardware segmentation as a protection mechanism, the binary translator could execute translated code at near hardware speeds. The binary translator also relied on partial evaluation and adaptive retranslation to reduce the overall overheads of virtualization. Written with the benefit of hindsight, this article shares the key lessons we learned from building the original system and from its later evolution. Edouard Bugnion, Scott Devine, Mendel Rosenblum, Jeremy Sugerman, Edward Y. Wang |
ACM Trans. Comput. Syst. | 4 |
| 2011 | Dynamic Fine-Grain Scheduling of Pipeline ParallelismabstractScheduling pipeline-parallel programs, defined as a graph of stages that communicate explicitly through queues, is challenging. When the application is regular and the underlying architecture can guarantee predictable execution times, several techniques exist to compute highly optimized static schedules. However, these schedules do not admit run-time load balancing, so variability introduced by the application or the underlying hardware causes load imbalance, hindering performance. On the other hand, existing schemes for dynamic fine-grain load balancing (such as task-stealing) do not work well on pipeline-parallel programs: they cannot guarantee memory footprint bounds, and do not adequately schedule complex graphs or graphs with ordered queues. We present a scheduler implementation for pipeline-parallel programs that performs fine-grain dynamic load balancing efficiently. Specifically, we implement the first real runtime for GRAMPS, a recently proposed programming model that focuses on supporting irregular pipeline and data-parallel applications (in contrast to classical stream programming models and schedulers, which require programs to be regular). Task-stealing with per-stage queues and queuing policies, coupled with a backpressure mechanism, allow us to maintain strict footprint bounds, and a buffer management scheme based on packet-stealing allows low-overhead and locality-aware dynamic allocation of queue data. We evaluate our runtime on a multi-core SMP and find that it provides low-overhead scheduling of irregular workloads while maintaining locality. We also show that the GRAMPS scheduler outperforms several other commonly used scheduling approaches. Specifically, while a typical task-stealing scheduler performs on par with GRAMPS on simple graphs, it does significantly worse on complex ones, a canonical GPGPU scheduler cannot exploit pipeline parallelism and suffers from large memory footprints, and a typical static, streaming scheduler achieves somewhat better locality, but suffers significant load imbalance on a general-purpose multi-core due to fine-grain architecture variability (e.g., cache misses and SMT). Daniel Sánchez 0003, David Lo 0003, Richard M. Yoo, Jeremy Sugerman, Christoforos E. Kozyrakis |
PACT | 4 |
| 2009 | GRAMPS: A programming model for graphics pipelinesabstractWe introduce GRAMPS, a programming model that generalizes concepts from modern real-time graphics pipelines by exposing a model of execution containing both fixed-function and application-programmable processing stages that exchange data via queues. GRAMPS allows the number, type, and connectivity of these processing stages to be defined by software, permitting arbitrary processing pipelines or even processing graphs. Applications achieve high performance using GRAMPS by expressing advanced rendering algorithms as custom pipelines, then using the pipeline as a rendering engine. We describe the design of GRAMPS, then evaluate it by implementing three pipelines, that is, Direct3D, a ray tracer, and a hybridization of the two, and running them on emulations of two different GRAMPS implementations: a traditional GPU-like architecture and a CPU-like multicore architecture. In our tests, our GRAMPS schedulers run our pipelines with 500 to 1500KB of queue usage at their peaks. Jeremy Sugerman, Kayvon Fatahalian, Solomon Boulos, Kurt Akeley, Pat Hanrahan |
ACM Trans. Graph. | 1 |
| 2008 | Larrabee: a many-core x86 architecture for visual computingabstractThis paper presents a many-core visual computing architecture code named Larrabee, a new software rendering pipeline, a manycore programming model, and performance analysis for several applications. Larrabee uses multiple in-order x86 CPU cores that are augmented by a wide vector processor unit, as well as some fixed function logic blocks. This provides dramatically higher performance per watt and per unit of area than out-of-order CPUs on highly parallel workloads. It also greatly increases the flexibility and programmability of the architecture as compared to standard GPUs. A coherent on-die 2 nd level cache allows efficient inter-processor communication and high-bandwidth local data access by CPU cores. Task scheduling is performed entirely with software in Larrabee, rather than in fixed function logic. The customizable software graphics rendering pipeline for this architecture uses binning in order to reduce required memory bandwidth, minimize lock contention, and increase opportunities for parallelism relative to standard GPUs. The Larrabee native programming model supports a variety of highly parallel applications that use irregular data structures. Performance analysis on those applications demonstrates Larrabee's potential for a broad range of parallel computation. Larry Seiler, Doug Carmean, Eric Sprangle, Tom Forsyth, Michael Abrash, Pradeep Dubey, Stephen Junkins, Adam T. Lake, Jeremy Sugerman, Robert Cavin, Roger Espasa, Ed Grochowski, Toni Juan, Pat Hanrahan |
ACM Trans. Graph. | 9 |
| 2007 | Interactive k-d tree GPU raytracingabstractOver the past few years, the powerful computation rates and high memory bandwidth of GPUs have attracted efforts to run raytracing on GPUs. Our work extends Foley et al.'s GPU k-d tree research. We port their kd-restart algorithm from multi-pass, using CPU load balancing, to single pass, using current GPUs' branching and looping abilities. We introduce three optimizations: a packetized formulation, a technique for restarting partially down the tree instead of at the root, and a small, fixed-size stack that is checked before resorting to restart. Our optimized implementation achieves 15 - 18 million primary rays per second and 16 - 27 million shadow rays per second on our test scenes. Daniel Reiter Horn, Jeremy Sugerman, Mike Houston, Pat Hanrahan |
SI3D | 2 |
| 2004 | Brook for GPUs: stream computing on graphics hardwareabstractIn this paper, we present Brook for GPUs, a system for general-purpose computation on programmable graphics hardware. Brook extends C to include simple data-parallel constructs, enabling the use of the GPU as a streaming co-processor. We present a compiler and runtime system that abstracts and virtualizes many aspects of graphics hardware. In addition, we present an analysis of the effectiveness of the GPU as a compute engine compared to the CPU, to determine when the GPU can outperform the CPU for a particular algorithm. We evaluate our system with five applications, the SAXPY and SGEMV BLAS operators, image segmentation, FFT, and ray tracing. For these applications, we demonstrate that our Brook implementations perform comparably to hand-written GPU code and up to seven times faster than their CPU counterparts. Ian Buck, Theresa Foley, Daniel Reiter Horn, Jeremy Sugerman, Kayvon Fatahalian, Mike Houston, Pat Hanrahan |
ACM Trans. Graph. | 4 |
| 2001 | Virtualizing I/O Devices on VMware Workstation's Hosted Virtual Machine Monitor
Jeremy Sugerman, Ganesh Venkitachalam, Beng-Hong Lim |
USENIX ATC, General Track | 1 |