Mike Houston

dblp:46/3385 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
GPUs and heterogeneous computing · 48% High-performance computing · 20% Parallel and multicore computing · 19%
Computer graphics and multimedia
3 papers
Rendering · 70% Visualization and visual analytics · 23% Virtual and augmented reality · 7%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 23 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.342008
GPU Computing · Proc. IEEE 2008
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006
Poster reception - N-Body simulation on GPUs · SC 2006
Parallel and multicore computing
parallel programming runtimes
0.112008
A portable runtime interface for multi-level memory hierarchies · PPoPP 2008
Compilers and program optimization › memory optimization
memory hierarchy optimization
0.112006
Sequoia: programming the memory hierarchy · SC 2006
GPUs and heterogeneous computing
GPU performance analysis
0.112006
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006
GPUs and heterogeneous computing
GPU programming
0.112006
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006
High-performance computing › scientific computing systems
molecular dynamics simulation
0.112006
Poster reception - N-Body simulation on GPUs · SC 2006
High-performance computing
n-body simulation
0.112006
Poster reception - N-Body simulation on GPUs · SC 2006
Parallel and multicore computing
parallel programming models
0.112006
Sequoia: programming the memory hierarchy · SC 2006
High-performance computing
scientific computing
0.112006
Poster reception - N-Body simulation on GPUs · SC 2006
Distributed systems
stream processing
0.122004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
Chromium: a stream-processing framework for interactive rendering on clusters · ACM Trans. Graph. 2002
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.112005
ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005
GPUs and heterogeneous computing
GPU-accelerated bioinformatics
0.112005
ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005
Parallel and multicore computing › parallel algorithms › parallel primitives
data-parallel primitives
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
GPUs and heterogeneous computing › GPU programming
GPU programming models
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
Visualization and visual analytics
interactive visualization
0.012003
Non-invasive interactive visualization of dynamic architectural environments · ACM Trans. Graph. 2003
Rendering › parallel rendering › distributed rendering
cluster rendering
0.012002
Chromium: a stream-processing framework for interactive rendering on clusters · ACM Trans. Graph. 2002
Rendering
parallel rendering
0.012002
Chromium: a stream-processing framework for interactive rendering on clusters · ACM Trans. Graph. 2002
Memory systems
data movement
0.012008
A portable runtime interface for multi-level memory hierarchies · PPoPP 2008
Memory systems
memory hierarchy
0.012008
A portable runtime interface for multi-level memory hierarchies · PPoPP 2008
Rendering
graphics hardware
0.012006
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006
High-performance computing › data-intensive computing
streaming computation
0.012005
ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005
Compilers and program optimization › accelerator compilation
GPU compiler
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
Virtual and augmented reality
immersive visualization
0.012003
Non-invasive interactive visualization of dynamic architectural environments · ACM Trans. Graph. 2003

Methods — techniques the papers use, named apart from their topics

compiler scheduling · 0.1bulk operation manipulation · 0.1performance analysis · 0.1GPU programming · 0.1viterbi algorithm · 0.1streaming algorithm · 0.1compiler and runtime abstraction · 0.1runtime composition · 0.1parallel programming · 0.1compiler target interface · 0.1stream programming · 0.0graphics API stream filtering · 0.0
YearPublicationVenuePosition
2009 AMD and OpenCL
Mike Houston
Hot Chips Symposium1
2008 A tuning framework for software-managed memory hierarchies
abstract
Achieving good performance on a modern machine with a multi-level memory hierarchy, and in particular on a machine with software-managed memories, requires precise tuning of programs to the machine's particular characteristics. A large program on a multi-level machine can easily expose tens or hundreds of inter-dependent parameters which require tuning, and manually searching the resultant large, non-linear space of program parameters is a tedious process of trial-and-error. In this paper we present a general framework for automatically tuning general applications to machines with software-managed memory hierarchies. We evaluate our framework by measuring the performance of benchmarks that are tuned for a range of machines with different memory hierarchy configurations: a cluster of Intel P4 Xeon processors, a single Cell processor, and a cluster of Sony Playstation3's.
Manman Ren, Ji Young Park, Mike Houston, Alex Aiken, William J. Dally
PACT3
2008 A portable runtime interface for multi-level memory hierarchies
abstract
We present a platform independent runtime interface for moving data and computation through parallel machines with multi-level memory hierarchies. We show that this interface can be used as a compiler target and can be implemented easily and efficiently on a variety of platforms. The interface design allows us to compose multiple runtimes, achieving portability across machines with multiple memory levels. We demonstrate portability of programs across machines with two memory levels with runtime implementations for multi-core/SMP machines, the STI Cell Broadband Engine, a distributed memory cluster, and disk systems. We also demonstrate portability across machines with multiple memory levels by composing runtimes and running on a cluster of SMP nodes, out-of-core algorithms on a Sony Playstation 3 pulling data from disk, and a cluster of Sony Playstation 3's. With this uniform interface, we achieve good performance for our applications and maximize bandwidth and computational resources on these system configurations.
Mike Houston, Ji Young Park, Manman Ren, Timothy J. Knight, Kayvon Fatahalian, Alex Aiken, William J. Dally, Pat Hanrahan
PPoPP1
2008 GPU Computing
abstract
The graphics processing unit (GPU) has become an integral part of today's mainstream computing systems. Over the past six years, there has been a marked increase in the performance and capabilities of GPUs. The modern GPU is not only a powerful graphics engine but also a highly parallel programmable processor featuring peak arithmetic and memory bandwidth that substantially outpaces its CPU counterpart. The GPU's rapid increase in both programmability and capability has spawned a research community that has successfully mapped a broad range of computationally demanding, complex problems to the GPU. This effort in general-purpose computing on the GPU, also known as GPU computing, has positioned the GPU as a compelling alternative to traditional microprocessors in high-performance computer systems of the future. We describe the background, hardware, and programming model for GPU computing, summarize the state of the art in tools and techniques, and present four GPU computing successes in game physics and computational biophysics that deliver order-of-magnitude performance gains over optimized CPU applications.
John D. Owens, Mike Houston, David P. Luebke, Simon Green, John E. Stone, James C. Phillips
Proc. IEEE2
2007 Compilation for explicitly managed memory hierarchies
abstract
We present a compiler for machines with an explicitly managed memory hierarchy and suggest that a primary role of any compiler for such architectures is to manipulate and schedule a hierarchy of bulk operations at varying scales of the application and of the machine. We evaluate the performance of our compiler using several benchmarks running on a Cell processor.
Timothy J. Knight, Ji Young Park, Manman Ren, Mike Houston, Mattan Erez, Kayvon Fatahalian, Alex Aiken, William J. Dally, Pat Hanrahan
PPoPP4
2007 Interactive k-d tree GPU raytracing
abstract
Over the past few years, the powerful computation rates and high memory bandwidth of GPUs have attracted efforts to run raytracing on GPUs. Our work extends Foley et al.'s GPU k-d tree research. We port their kd-restart algorithm from multi-pass, using CPU load balancing, to single pass, using current GPUs' branching and looping abilities. We introduce three optimizations: a packetized formulation, a technique for restarting partially down the tree instead of at the root, and a small, fixed-size stack that is checked before resorting to restart. Our optimized implementation achieves 15 - 18 million primary rays per second and 16 - 27 million shadow rays per second on our test scenes.
Daniel Reiter Horn, Jeremy Sugerman, Mike Houston, Pat Hanrahan
SI3D3
2006 Poster reception - N-Body simulation on GPUs
abstract
Commercial graphics processors (GPUs) have high compute capacity at very low cost, which makes them attractive for general purpose scientic computing. In this poster we show how graphics processors can be used for N-body simulations to obtain large improvements in performance over current generation CPUs. We have developed a highly optimized algorithm for performing the O(N^2) force calculations that constitute the major part of stellar and molecular dynamics simulations. In the calculations, we achieve sustained performance of nearly 100 GFlops on an ATI X1900XTX. The performance on GPUs 25x an Intel Pentium4, and 2x specialized hardware such as GRAPE-6A, but at a fraction of the cost. Furthermore, the wide availability of GPUs has signicant implications for cluster computing and distributed computing efforts like [email protected]
Erich Elsen, Mike Houston, Vaidyanathan Vishal, Eric Darve, Pat Hanrahan, Vijay S. Pande
SC2
2006 Sequoia: programming the memory hierarchy
abstract
We present Sequoia, a programming language designed to facilitate the development of memory hierarchy aware parallel programs that remain portable across modern machines featuring different memory hierarchy configurations. Sequoia abstractly exposes hierarchical memory in the programming model and provides language mechanisms to describe communication vertically through the machine and to localize computation to particular memory locations within it. We have implemented a complete programming system, including a compiler and runtime systems for Cell processor-based blade systems and distributed memory clusters, and demonstrate efficient performance running Sequoia programs on both of these platforms.
Kayvon Fatahalian, Daniel Reiter Horn, Timothy J. Knight, Larkhoon Leem, Mike Houston, Ji Young Park, Mattan Erez, Manman Ren, Alex Aiken, William J. Dally, Pat Hanrahan
SC5
2006 S07 - GPGPU: general-purpose computation on graphics hardware
abstract
The graphics processor (GPU) on today's commodity video cards has evolved into an extremely powerful and flexible processor. Modern graphics architectures provide tremendous memory bandwidth and computational horsepower, with dozens of fully programmable shading units that support vector operations and IEEE floating point precision. High-level languages have emerged for graphics hardware, making this computational power accessible. GPGPU stands for "General-Purpose Computation on GPUs". GPGPU researchers have achieved over an order of magnitude speedup over modern CPUs on some non-graphics problems.This course provides detailed coverage of general-purpose computation on graphics hardware. We emphasize core computational building blocks, ranging from linear algebra to database queries, and review the tools, perils, and strategies in GPU programming. We present analysis of GPU performance characteristics, and use this analysis to provide insight into how to build efficient GPGPU algorithms. Finally we present a set of case studies on general-purpose applications of graphics hardware.
David P. Luebke, Mark J. Harris, Naga K. Govindaraju, Aaron E. Lefohn, Mike Houston, John D. Owens, Mark Segal, Matthew Papakipos, Ian Buck
SC5
2005 ClawHMMER: A Streaming HMMer-Search Implementation
abstract
The proliferation of biological sequence data has motivated the need for an extremely fast probabilistic sequence search. One method for performing this search involves evaluating the Viterbi probability of a hidden Markov model (HMM) of a desired sequence family for each sequence in a protein database. However, one of the difficulties with current implementations is the time required to search large databases. Many current and upcoming architectures offering large amounts of compute power are designed with data-parallel execution and streaming in mind. We present a streaming algorithm for evaluating an HMM’s Viterbi probability and refine it for the specific HMM used in biological sequence search. We implement our streaming algorithm in the Brook language, allowing us to execute the algorithm on graphics processors. We demonstrate that this streaming algorithm on graphics processors can outperform available CPU implementations. We also demonstrate this implementation running on a 16 node graphics cluster.
Daniel Reiter Horn, Mike Houston, Pat Hanrahan
SC2
2004 Brook for GPUs: stream computing on graphics hardware
abstract
In this paper, we present Brook for GPUs, a system for general-purpose computation on programmable graphics hardware. Brook extends C to include simple data-parallel constructs, enabling the use of the GPU as a streaming co-processor. We present a compiler and runtime system that abstracts and virtualizes many aspects of graphics hardware. In addition, we present an analysis of the effectiveness of the GPU as a compute engine compared to the CPU, to determine when the GPU can outperform the CPU for a particular algorithm. We evaluate our system with five applications, the SAXPY and SGEMV BLAS operators, image segmentation, FFT, and ray tracing. For these applications, we demonstrate that our Brook implementations perform comparably to hand-written GPU code and up to seven times faster than their CPU counterparts.
Ian Buck, Theresa Foley, Daniel Reiter Horn, Jeremy Sugerman, Kayvon Fatahalian, Mike Houston, Pat Hanrahan
ACM Trans. Graph.6
2003 Non-invasive interactive visualization of dynamic architectural environments
abstract
We present a system for interactively producing exploded views of 3D architectural environments such as multi-story buildings. These exploded views allow viewers to simultaneously see the internal and external structures of such environments. To create an exploded view we analyze the geometry of the environment to locate individual stories. We then use clipping planes and multipass rendering to separately render each story of the environment in exploded form. Our system operates at the graphics driver level and therefore can be applied to existing OpenGL applications, such as first-person multi-player video games, without modification. The resulting visualization allows users to understand the global structure of architectural environments and to observe the actions of dynamic characters and objects interacting within such environments.
Christopher Niederauer, Mike Houston, Maneesh Agrawala, Greg Humphreys
SI3D2
2003 Non-invasive interactive visualization of dynamic architectural environments
abstract
No abstract available.
Christopher Niederauer, Mike Houston, Maneesh Agrawala, Greg Humphreys
ACM Trans. Graph.2
2002 Chromium: a stream-processing framework for interactive rendering on clusters
abstract
We describe Chromium, a system for manipulating streams of graphics API commands on clusters of workstations. Chromium's stream filters can be arranged to create sort-first and sort-last parallel graphics architectures that, in many cases, support the same applications while using only commodity graphics accelerators. In addition, these stream filters can be extended programmatically, allowing the user to customize the stream transformations performed by nodes in a cluster. Because our stream processing mechanism is completely general, any cluster-parallel rendering algorithm can be either implemented on top of or embedded in Chromium. In this paper, we give examples of real-world applications that use Chromium to achieve good scalability on clusters of workstations, and describe other potential uses of this stream processing technology. By completely abstracting the underlying graphics architecture, network topology, and API command processing semantics, we allow a variety of applications to run in different environments.
Greg Humphreys, Mike Houston, Ren Ng, Randall Frank, Sean Ahern, Peter D. Kirchner, James T. Klosowski
ACM Trans. Graph.2