VLDB 2026 Research / reviewers in the wild / expert
Mike Houston
dblp:46/3385
· DBLP profile ↗
14ranked-venue papers
2as first author
0since 2021 · last 2009
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
GPUs and heterogeneous computing · 48% High-performance computing · 20% Parallel and multicore computing · 19% | |
| Computer graphics and multimedia
3 papers |
Rendering · 70% Visualization and visual analytics · 23% Virtual and augmented reality · 7% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU computing |
0.3 | 4 | 2008 | GPU Computing · Proc. IEEE 2008 S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006 Poster reception - N-Body simulation on GPUs · SC 2006 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2008 | A portable runtime interface for multi-level memory hierarchies · PPoPP 2008 |
Compilers and program optimization › memory optimization
memory hierarchy optimization |
0.1 | 1 | 2006 | Sequoia: programming the memory hierarchy · SC 2006 |
GPUs and heterogeneous computing
GPU performance analysis |
0.1 | 1 | 2006 | S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006 |
GPUs and heterogeneous computing
GPU programming |
0.1 | 1 | 2006 | S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006 |
High-performance computing › scientific computing systems
molecular dynamics simulation |
0.1 | 1 | 2006 | Poster reception - N-Body simulation on GPUs · SC 2006 |
High-performance computing
n-body simulation |
0.1 | 1 | 2006 | Poster reception - N-Body simulation on GPUs · SC 2006 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2006 | Sequoia: programming the memory hierarchy · SC 2006 |
High-performance computing
scientific computing |
0.1 | 1 | 2006 | Poster reception - N-Body simulation on GPUs · SC 2006 |
Distributed systems
stream processing |
0.1 | 2 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 Chromium: a stream-processing framework for interactive rendering on clusters · ACM Trans. Graph. 2002 |
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search |
0.1 | 1 | 2005 | ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005 |
GPUs and heterogeneous computing
GPU-accelerated bioinformatics |
0.1 | 1 | 2005 | ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005 |
Parallel and multicore computing › parallel algorithms › parallel primitives
data-parallel primitives |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
GPUs and heterogeneous computing › GPU programming
GPU programming models |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
Visualization and visual analytics
interactive visualization |
0.0 | 1 | 2003 | Non-invasive interactive visualization of dynamic architectural environments · ACM Trans. Graph. 2003 |
Rendering › parallel rendering › distributed rendering
cluster rendering |
0.0 | 1 | 2002 | Chromium: a stream-processing framework for interactive rendering on clusters · ACM Trans. Graph. 2002 |
Rendering
parallel rendering |
0.0 | 1 | 2002 | Chromium: a stream-processing framework for interactive rendering on clusters · ACM Trans. Graph. 2002 |
Memory systems
data movement |
0.0 | 1 | 2008 | A portable runtime interface for multi-level memory hierarchies · PPoPP 2008 |
Memory systems
memory hierarchy |
0.0 | 1 | 2008 | A portable runtime interface for multi-level memory hierarchies · PPoPP 2008 |
Rendering
graphics hardware |
0.0 | 1 | 2006 | S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006 |
High-performance computing › data-intensive computing
streaming computation |
0.0 | 1 | 2005 | ClawHMMER: A Streaming HMMer-Search Implementation · SC 2005 |
Compilers and program optimization › accelerator compilation
GPU compiler |
0.0 | 1 | 2004 | Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004 |
Virtual and augmented reality
immersive visualization |
0.0 | 1 | 2003 | Non-invasive interactive visualization of dynamic architectural environments · ACM Trans. Graph. 2003 |
Methods — techniques the papers use, named apart from their topics
compiler scheduling · 0.1bulk operation manipulation · 0.1performance analysis · 0.1GPU programming · 0.1viterbi algorithm · 0.1streaming algorithm · 0.1compiler and runtime abstraction · 0.1runtime composition · 0.1parallel programming · 0.1compiler target interface · 0.1stream programming · 0.0graphics API stream filtering · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | AMD and OpenCL
Mike Houston |
Hot Chips Symposium | 1 |
| 2008 | A tuning framework for software-managed memory hierarchiesabstractAchieving good performance on a modern machine with a multi-level memory hierarchy, and in particular on a machine with software-managed memories, requires precise tuning of programs to the machine's particular characteristics. A large program on a multi-level machine can easily expose tens or hundreds of inter-dependent parameters which require tuning, and manually searching the resultant large, non-linear space of program parameters is a tedious process of trial-and-error. In this paper we present a general framework for automatically tuning general applications to machines with software-managed memory hierarchies. We evaluate our framework by measuring the performance of benchmarks that are tuned for a range of machines with different memory hierarchy configurations: a cluster of Intel P4 Xeon processors, a single Cell processor, and a cluster of Sony Playstation3's. Manman Ren, Ji Young Park, Mike Houston, Alex Aiken, William J. Dally |
PACT | 3 |
| 2008 | A portable runtime interface for multi-level memory hierarchiesabstractWe present a platform independent runtime interface for moving data and computation through parallel machines with multi-level memory hierarchies. We show that this interface can be used as a compiler target and can be implemented easily and efficiently on a variety of platforms. The interface design allows us to compose multiple runtimes, achieving portability across machines with multiple memory levels. We demonstrate portability of programs across machines with two memory levels with runtime implementations for multi-core/SMP machines, the STI Cell Broadband Engine, a distributed memory cluster, and disk systems. We also demonstrate portability across machines with multiple memory levels by composing runtimes and running on a cluster of SMP nodes, out-of-core algorithms on a Sony Playstation 3 pulling data from disk, and a cluster of Sony Playstation 3's. With this uniform interface, we achieve good performance for our applications and maximize bandwidth and computational resources on these system configurations. Mike Houston, Ji Young Park, Manman Ren, Timothy J. Knight, Kayvon Fatahalian, Alex Aiken, William J. Dally, Pat Hanrahan |
PPoPP | 1 |
| 2008 | GPU ComputingabstractThe graphics processing unit (GPU) has become an integral part of today's mainstream computing systems. Over the past six years, there has been a marked increase in the performance and capabilities of GPUs. The modern GPU is not only a powerful graphics engine but also a highly parallel programmable processor featuring peak arithmetic and memory bandwidth that substantially outpaces its CPU counterpart. The GPU's rapid increase in both programmability and capability has spawned a research community that has successfully mapped a broad range of computationally demanding, complex problems to the GPU. This effort in general-purpose computing on the GPU, also known as GPU computing, has positioned the GPU as a compelling alternative to traditional microprocessors in high-performance computer systems of the future. We describe the background, hardware, and programming model for GPU computing, summarize the state of the art in tools and techniques, and present four GPU computing successes in game physics and computational biophysics that deliver order-of-magnitude performance gains over optimized CPU applications. John D. Owens, Mike Houston, David P. Luebke, Simon Green, John E. Stone, James C. Phillips |
Proc. IEEE | 2 |
| 2007 | Compilation for explicitly managed memory hierarchiesabstractWe present a compiler for machines with an explicitly managed memory hierarchy and suggest that a primary role of any compiler for such architectures is to manipulate and schedule a hierarchy of bulk operations at varying scales of the application and of the machine. We evaluate the performance of our compiler using several benchmarks running on a Cell processor. Timothy J. Knight, Ji Young Park, Manman Ren, Mike Houston, Mattan Erez, Kayvon Fatahalian, Alex Aiken, William J. Dally, Pat Hanrahan |
PPoPP | 4 |
| 2007 | Interactive k-d tree GPU raytracingabstractOver the past few years, the powerful computation rates and high memory bandwidth of GPUs have attracted efforts to run raytracing on GPUs. Our work extends Foley et al.'s GPU k-d tree research. We port their kd-restart algorithm from multi-pass, using CPU load balancing, to single pass, using current GPUs' branching and looping abilities. We introduce three optimizations: a packetized formulation, a technique for restarting partially down the tree instead of at the root, and a small, fixed-size stack that is checked before resorting to restart. Our optimized implementation achieves 15 - 18 million primary rays per second and 16 - 27 million shadow rays per second on our test scenes. Daniel Reiter Horn, Jeremy Sugerman, Mike Houston, Pat Hanrahan |
SI3D | 3 |
| 2006 | Poster reception - N-Body simulation on GPUsabstractCommercial graphics processors (GPUs) have high compute capacity at very low cost, which makes them attractive for general purpose scientic computing. In this poster we show how graphics processors can be used for N-body simulations to obtain large improvements in performance over current generation CPUs. We have developed a highly optimized algorithm for performing the O(N^2) force calculations that constitute the major part of stellar and molecular dynamics simulations. In the calculations, we achieve sustained performance of nearly 100 GFlops on an ATI X1900XTX. The performance on GPUs 25x an Intel Pentium4, and 2x specialized hardware such as GRAPE-6A, but at a fraction of the cost. Furthermore, the wide availability of GPUs has signicant implications for cluster computing and distributed computing efforts like [email protected] Erich Elsen, Mike Houston, Vaidyanathan Vishal, Eric Darve, Pat Hanrahan, Vijay S. Pande |
SC | 2 |
| 2006 | Sequoia: programming the memory hierarchyabstractWe present Sequoia, a programming language designed to facilitate the development of memory hierarchy aware parallel programs that remain portable across modern machines featuring different memory hierarchy configurations. Sequoia abstractly exposes hierarchical memory in the programming model and provides language mechanisms to describe communication vertically through the machine and to localize computation to particular memory locations within it. We have implemented a complete programming system, including a compiler and runtime systems for Cell processor-based blade systems and distributed memory clusters, and demonstrate efficient performance running Sequoia programs on both of these platforms. Kayvon Fatahalian, Daniel Reiter Horn, Timothy J. Knight, Larkhoon Leem, Mike Houston, Ji Young Park, Mattan Erez, Manman Ren, Alex Aiken, William J. Dally, Pat Hanrahan |
SC | 5 |
| 2006 | S07 - GPGPU: general-purpose computation on graphics hardwareabstractThe graphics processor (GPU) on today's commodity video cards has evolved into an extremely powerful and flexible processor. Modern graphics architectures provide tremendous memory bandwidth and computational horsepower, with dozens of fully programmable shading units that support vector operations and IEEE floating point precision. High-level languages have emerged for graphics hardware, making this computational power accessible. GPGPU stands for "General-Purpose Computation on GPUs". GPGPU researchers have achieved over an order of magnitude speedup over modern CPUs on some non-graphics problems.This course provides detailed coverage of general-purpose computation on graphics hardware. We emphasize core computational building blocks, ranging from linear algebra to database queries, and review the tools, perils, and strategies in GPU programming. We present analysis of GPU performance characteristics, and use this analysis to provide insight into how to build efficient GPGPU algorithms. Finally we present a set of case studies on general-purpose applications of graphics hardware. David P. Luebke, Mark J. Harris, Naga K. Govindaraju, Aaron E. Lefohn, Mike Houston, John D. Owens, Mark Segal, Matthew Papakipos, Ian Buck |
SC | 5 |
| 2005 | ClawHMMER: A Streaming HMMer-Search ImplementationabstractThe proliferation of biological sequence data has motivated the need for an extremely fast probabilistic sequence search. One method for performing this search involves evaluating the Viterbi probability of a hidden Markov model (HMM) of a desired sequence family for each sequence in a protein database. However, one of the difficulties with current implementations is the time required to search large databases. Many current and upcoming architectures offering large amounts of compute power are designed with data-parallel execution and streaming in mind. We present a streaming algorithm for evaluating an HMM’s Viterbi probability and refine it for the specific HMM used in biological sequence search. We implement our streaming algorithm in the Brook language, allowing us to execute the algorithm on graphics processors. We demonstrate that this streaming algorithm on graphics processors can outperform available CPU implementations. We also demonstrate this implementation running on a 16 node graphics cluster. Daniel Reiter Horn, Mike Houston, Pat Hanrahan |
SC | 2 |
| 2004 | Brook for GPUs: stream computing on graphics hardwareabstractIn this paper, we present Brook for GPUs, a system for general-purpose computation on programmable graphics hardware. Brook extends C to include simple data-parallel constructs, enabling the use of the GPU as a streaming co-processor. We present a compiler and runtime system that abstracts and virtualizes many aspects of graphics hardware. In addition, we present an analysis of the effectiveness of the GPU as a compute engine compared to the CPU, to determine when the GPU can outperform the CPU for a particular algorithm. We evaluate our system with five applications, the SAXPY and SGEMV BLAS operators, image segmentation, FFT, and ray tracing. For these applications, we demonstrate that our Brook implementations perform comparably to hand-written GPU code and up to seven times faster than their CPU counterparts. Ian Buck, Theresa Foley, Daniel Reiter Horn, Jeremy Sugerman, Kayvon Fatahalian, Mike Houston, Pat Hanrahan |
ACM Trans. Graph. | 6 |
| 2003 | Non-invasive interactive visualization of dynamic architectural environmentsabstractWe present a system for interactively producing exploded views of 3D architectural environments such as multi-story buildings. These exploded views allow viewers to simultaneously see the internal and external structures of such environments. To create an exploded view we analyze the geometry of the environment to locate individual stories. We then use clipping planes and multipass rendering to separately render each story of the environment in exploded form. Our system operates at the graphics driver level and therefore can be applied to existing OpenGL applications, such as first-person multi-player video games, without modification. The resulting visualization allows users to understand the global structure of architectural environments and to observe the actions of dynamic characters and objects interacting within such environments. Christopher Niederauer, Mike Houston, Maneesh Agrawala, Greg Humphreys |
SI3D | 2 |
| 2003 | Non-invasive interactive visualization of dynamic architectural environmentsabstractNo abstract available. Christopher Niederauer, Mike Houston, Maneesh Agrawala, Greg Humphreys |
ACM Trans. Graph. | 2 |
| 2002 | Chromium: a stream-processing framework for interactive rendering on clustersabstractWe describe Chromium, a system for manipulating streams of graphics API commands on clusters of workstations. Chromium's stream filters can be arranged to create sort-first and sort-last parallel graphics architectures that, in many cases, support the same applications while using only commodity graphics accelerators. In addition, these stream filters can be extended programmatically, allowing the user to customize the stream transformations performed by nodes in a cluster. Because our stream processing mechanism is completely general, any cluster-parallel rendering algorithm can be either implemented on top of or embedded in Chromium. In this paper, we give examples of real-world applications that use Chromium to achieve good scalability on clusters of workstations, and describe other potential uses of this stream processing technology. By completely abstracting the underlying graphics architecture, network topology, and API command processing semantics, we allow a variety of applications to run in different environments. Greg Humphreys, Mike Houston, Ren Ng, Randall Frank, Sean Ahern, Peter D. Kirchner, James T. Klosowski |
ACM Trans. Graph. | 2 |