Mark J. Harris

dblp:45/6132 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.112006
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006
GPUs and heterogeneous computing
GPU performance analysis
0.112006
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006
GPUs and heterogeneous computing
GPU programming
0.112006
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006
Rendering
graphics hardware
0.012006
S07 - GPGPU: general-purpose computation on graphics hardware · SC 2006

Methods — techniques the papers use, named apart from their topics

performance analysis · 0.1GPU programming · 0.1
YearPublicationVenuePosition
2009 Designing efficient sorting algorithms for manycore GPUs
abstract
We describe the design of high-performance parallel radix sort and merge sort routines for manycore GPUs, taking advantage of the full programmability offered by CUDA. Our radix sort is the fastest GPU sort and our merge sort is the fastest comparison-based sort reported in the literature. Our radix sort is up to 4 times faster than the graphics-based GPUSort and greater than 2 times faster than other CUDA-based radix sorts. It is also 23% faster, on average, than even a very carefully optimized multicore CPU sorting routine. To achieve this performance, we carefully design our algorithms to expose substantial fine-grained parallelism and decompose the computation into independent tasks that perform minimal global communication. We exploit the high-speed onchip shared memory provided by NVIDIA's GPU architecture and efficient data-parallel primitives, particularly parallel scan. While targeted at GPUs, these algorithms should also be well-suited for other manycore processors.
Nadathur Satish, Mark J. Harris, Michael Garland
IPDPS2
2008 Many-core GPU computing with NVIDIA CUDA
abstract
In the past, graphics processors were special-purpose hardwired application accelerators, suitable only for conventional graphics applications. Modern GPUs are fully programmable, massively parallel floating point processors. In this talk I will describe NVIDIA's scalable, highly parallel many-core GPU architecture and how CUDA software for GPU computing delivers high throughput for data-intensive processing. I will discuss how CUDA is reinvigorating research on data-parallel algorithms, reducing time to scientific discovery, and enabling a variety of compute-intensive industrial applications of GPUs beyond computer graphics.
Mark J. Harris
ICS1
2006 S07 - GPGPU: general-purpose computation on graphics hardware
abstract
The graphics processor (GPU) on today's commodity video cards has evolved into an extremely powerful and flexible processor. Modern graphics architectures provide tremendous memory bandwidth and computational horsepower, with dozens of fully programmable shading units that support vector operations and IEEE floating point precision. High-level languages have emerged for graphics hardware, making this computational power accessible. GPGPU stands for "General-Purpose Computation on GPUs". GPGPU researchers have achieved over an order of magnitude speedup over modern CPUs on some non-graphics problems.This course provides detailed coverage of general-purpose computation on graphics hardware. We emphasize core computational building blocks, ranging from linear algebra to database queries, and review the tools, perils, and strategies in GPU programming. We present analysis of GPU performance characteristics, and use this analysis to provide insight into how to build efficient GPGPU algorithms. Finally we present a set of case studies on general-purpose applications of graphics hardware.
David P. Luebke, Mark J. Harris, Naga K. Govindaraju, Aaron E. Lefohn, Mike Houston, John D. Owens, Mark Segal, Matthew Papakipos, Ian Buck
SC2
2004 Radiosity on Graphics Hardware
Greg Coombe, Mark J. Harris, Anselmo Lastra
Graphics Interface2
2001 Real-Time Cloud Rendering
abstract
This paper presents a method for realistic real-time rendering of clouds suitable for flight simulation and games. It provides a cloud shading algorithm that approximates multiple forward scattering in a preprocess, and first order anisotropic scattering at runtime. Impostors are used to accelerate cloud rendering by exploiting frame-to-frame coherence in an interactive flight simulation. Impostors are shown to be particularly well suited to clouds, even in circumstances under which they cannot be applied to the rendering of polygonal geometry. The method allows hundreds of clouds and hundreds of thousands of particles to be rendered at high frame rates, and improves interaction with clouds by reducing artifacts introduced by direct particle rendering techniques.
Mark J. Harris, Anselmo Lastra
Comput. Graph. Forum1