Daniel Reiter Horn

dblp:h/DRHorn · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 59% Distributed systems · 17% GPUs and heterogeneous computing · 13%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Virtual and augmented reality · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
file systems
0.312017
The Design, Implementation, and Deployment of a System to Transparently Compress Hundreds of Petabytes of Image Files for a File-Storage Service · NSDI 2017
Storage systems › data compression
transparent compression
0.312017
The Design, Implementation, and Deployment of a System to Transparently Compress Hundreds of Petabytes of Image Files for a File-Storage Service · NSDI 2017
Compilers and program optimization › memory optimization
memory hierarchy optimization
0.112006
Sequoia: programming the memory hierarchy · SC 2006
Parallel and multicore computing
parallel programming models
0.112006
Sequoia: programming the memory hierarchy · SC 2006
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.112005
ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005
GPUs and heterogeneous computing
GPU-accelerated bioinformatics
0.112005
ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005
Parallel and multicore computing › parallel algorithms › parallel primitives
data-parallel primitives
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
GPUs and heterogeneous computing
GPU computing
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
GPUs and heterogeneous computing › GPU programming
GPU programming models
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
Distributed systems
stream processing
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
Virtual and augmented reality
immersive interaction
0.012012
A Scalable Server for 3D Metaverses · USENIX ATC 2012
High-performance computing › data-intensive computing
streaming computation
0.012005
ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005
Compilers and program optimization › accelerator compilation
GPU compiler
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004

Methods — techniques the papers use, named apart from their topics

viterbi algorithm · 0.1streaming algorithm · 0.1stream programming · 0.1compiler and runtime abstraction · 0.1
YearPublicationVenuePosition
2017 The Design, Implementation, and Deployment of a System to Transparently Compress Hundreds of Petabytes of Image Files for a File-Storage Service
Daniel Reiter Horn, Ken Elkabany, Chris Lesniewski-Laas, Keith Winstein
NSDI1
2012 A Scalable Server for 3D Metaverses
Ewen Cheslack-Postava, Tahir Azim, Behram F. T. Mistree, Daniel Reiter Horn, Jeff Terrace, Philip Alexander Levis, Michael J. Freedman
USENIX ATC4
2007 LightShop: interactive light field manipulation and rendering
abstract
Light fields can be used to represent an object's appearance with a high degree of realism. However, unlike their geometric counterparts, these image-based representations lack user control for manipulating them. We present a system that allows a user to interactively manipulate, composite and render multiple light fields. LightShop is a modular system consisting of three parts: 1) a set of functions that allow a user to model a scene containing multiple light fields, 2) a ray-shading language that describes how an image should be constructed from a set of light fields, and 3) a real-time light field rendering system in OpenGL that can plug into existing 3D engines as a GLSL shader.
Daniel Reiter Horn, Billy Chen
SI3D1
2007 Interactive k-d tree GPU raytracing
abstract
Over the past few years, the powerful computation rates and high memory bandwidth of GPUs have attracted efforts to run raytracing on GPUs. Our work extends Foley et al.'s GPU k-d tree research. We port their kd-restart algorithm from multi-pass, using CPU load balancing, to single pass, using current GPUs' branching and looping abilities. We introduce three optimizations: a packetized formulation, a technique for restarting partially down the tree instead of at the root, and a small, fixed-size stack that is checked before resorting to restart. Our optimized implementation achieves 15 - 18 million primary rays per second and 16 - 27 million shadow rays per second on our test scenes.
Daniel Reiter Horn, Jeremy Sugerman, Mike Houston, Pat Hanrahan
SI3D1
2006 Sequoia: programming the memory hierarchy
abstract
We present Sequoia, a programming language designed to facilitate the development of memory hierarchy aware parallel programs that remain portable across modern machines featuring different memory hierarchy configurations. Sequoia abstractly exposes hierarchical memory in the programming model and provides language mechanisms to describe communication vertically through the machine and to localize computation to particular memory locations within it. We have implemented a complete programming system, including a compiler and runtime systems for Cell processor-based blade systems and distributed memory clusters, and demonstrate efficient performance running Sequoia programs on both of these platforms.
Kayvon Fatahalian, Daniel Reiter Horn, Timothy J. Knight, Larkhoon Leem, Mike Houston, Ji Young Park, Mattan Erez, Manman Ren, Alex Aiken, William J. Dally, Pat Hanrahan
SC2
2005 ClawHMMER: A Streaming HMMer-Search Implementation
abstract
The proliferation of biological sequence data has motivated the need for an extremely fast probabilistic sequence search. One method for performing this search involves evaluating the Viterbi probability of a hidden Markov model (HMM) of a desired sequence family for each sequence in a protein database. However, one of the difficulties with current implementations is the time required to search large databases. Many current and upcoming architectures offering large amounts of compute power are designed with data-parallel execution and streaming in mind. We present a streaming algorithm for evaluating an HMM’s Viterbi probability and refine it for the specific HMM used in biological sequence search. We implement our streaming algorithm in the Brook language, allowing us to execute the algorithm on graphics processors. We demonstrate that this streaming algorithm on graphics processors can outperform available CPU implementations. We also demonstrate this implementation running on a 16 node graphics cluster.
Daniel Reiter Horn, Mike Houston, Pat Hanrahan
SC1
2004 Brook for GPUs: stream computing on graphics hardware
abstract
In this paper, we present Brook for GPUs, a system for general-purpose computation on programmable graphics hardware. Brook extends C to include simple data-parallel constructs, enabling the use of the GPU as a streaming co-processor. We present a compiler and runtime system that abstracts and virtualizes many aspects of graphics hardware. In addition, we present an analysis of the effectiveness of the GPU as a compute engine compared to the CPU, to determine when the GPU can outperform the CPU for a particular algorithm. We evaluate our system with five applications, the SAXPY and SGEMV BLAS operators, image segmentation, FFT, and ray tracing. For these applications, we demonstrate that our Brook implementations perform comparably to hand-written GPU code and up to seven times faster than their CPU counterparts.
Ian Buck, Theresa Foley, Daniel Reiter Horn, Jeremy Sugerman, Kayvon Fatahalian, Mike Houston, Pat Hanrahan
ACM Trans. Graph.3