Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

David Blair Kirk

dblp:46/5249 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
0since 2021 · last 2015
0000-0002-4887-5098ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-authorHuman-computer interaction and ubiquitous computing · 7 · 3 first-authorSystems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
12 papers
GPUs and heterogeneous computing · 68% Memory systems · 14% Integrated circuit design · 10%
Computer graphics and multimedia
6 papers
Rendering · 66% Geometric modeling and processing · 20% Visual content generation and editing · 14%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
heterogeneous parallel programming
0.212015
Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015
GPUs and heterogeneous computing › GPU kernel optimization
CUDA optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
GPU performance optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
GPU programming
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
Memory systems
memory access optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
Integrated circuit design
interconnect
0.112015
Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015
Integrated circuit design › analog and mixed-signal circuits
analog VLSI
0.021993
Implementing rotation matrix constraints in Analog VLSI · SIGGRAPH 1993
An Analog VLSI Chip for Radial Basis Functions · NIPS 1992
Memory systems
cache
0.031990
SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990
SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989
Process Dependent Static Cache Partitioning for Real-Time Systems · RTSS 1988
Integrated circuit design
analog and mixed-signal circuits
0.021993
Implementing rotation matrix constraints in Analog VLSI · SIGGRAPH 1993
An Analog VLSI Chip for Radial Basis Functions · NIPS 1992
Memory systems › cache management
cache allocation
0.021990
SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990
SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989
Memory systems
cache design
0.021990
SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990
SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989
Electronic design automation › timing analysis
predictable cache performance
0.021990
SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990
SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989
Embedded and real-time systems
real-time scheduling
0.031990
Priority-Driven, Preemptive I/O Controllers for Real-Time Systems · ISCA 1988
SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990
SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989
Rendering › sampling
adaptive sampling
0.011991
Unbiased sampling techniques for image synthesis · SIGGRAPH 1991
Visual content generation and editing
image generation
0.011991
Unbiased sampling techniques for image synthesis · SIGGRAPH 1991
Geometric modeling and processing
unbiased sampling
0.011991
Unbiased sampling techniques for image synthesis · SIGGRAPH 1991
Integrated circuit design › analog and mixed-signal circuits
analog circuit design
0.011991
Constrained Optimization Applied to the Parameter Setting Problem for Analog Circuits · NIPS 1991
Mathematical optimization
constrained optimization
0.011991
Constrained Optimization Applied to the Parameter Setting Problem for Analog Circuits · NIPS 1991
Rendering › rasterization
hardware rasterization
0.011990
The rendering architecture of the DN10000VS · SIGGRAPH 1990
Rendering
monte carlo rendering
0.011990
Particle transport and image synthesis · SIGGRAPH 1990
Rendering
rasterization
0.011990
The rendering architecture of the DN10000VS · SIGGRAPH 1990
Rendering
ray tracing
0.021990
Fast ray tracing by ray classification · SIGGRAPH 1987
Particle transport and image synthesis · SIGGRAPH 1990
Hardware accelerators and domain-specific architectures › domain-specific accelerator
application-specific processor
0.011988
The White Dwarf: A High-Performance Application-Specific Processor · ISCA 1988
Memory systems › cache management
cache partitioning
0.011988
Process Dependent Static Cache Partitioning for Real-Time Systems · RTSS 1988
Memory systems › cache › CPU cache
instruction cache
0.011988
Process Dependent Static Cache Partitioning for Real-Time Systems · RTSS 1988
Embedded and real-time systems › real-time scheduling › priority scheduling
preemptive priority scheduling
0.021990
SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990
SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989
Geometric modeling and processing
rotation representation
0.011993
Implementing rotation matrix constraints in Analog VLSI · SIGGRAPH 1993
Rendering › ray tracing
monte carlo ray tracing
0.011990
Particle transport and image synthesis · SIGGRAPH 1990
Embedded and real-time systems
real-time system design
0.011988
Priority-Driven, Preemptive I/O Controllers for Real-Time Systems · ISCA 1988
High-performance computing
scientific computing systems
0.011988
The White Dwarf: A High-Performance Application-Specific Processor · ISCA 1988

Methods — techniques the papers use, named apart from their topics

pinned buffers · 0.2peer DMA · 0.2double buffering · 0.2register allocation · 0.1memory coalescing · 0.1latency hiding · 0.1least squares · 0.0adaptive analog VLSI · 0.0radial basis functions · 0.0analog VLSI · 0.0statistical bias analysis · 0.0constrained optimization · 0.0anti-aliasing · 0.0texture mapping · 0.0splitting · 0.0russian roulette · 0.0quadratic interpolation · 0.0importance sampling · 0.0
YearPublicationVenuePosition
2015 Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications
abstract
Heterogeneous parallel computing applications often process large data sets that require multiple GPUs to jointly meet their needs for physical memory capacity and compute throughput. However, the lack of high-level abstractions in previous heterogeneous parallel programming models force programmers to resort to multiple code versions, complex data copy steps and synchronization schemes when exchanging data between multiple GPU devices, which results in high software development cost, poor maintainability, and even poor performance. This paper describes the HPE runtime system, and the associated architecture support, which enables a simple, efficient programming interface for exchanging data between multiple GPUs through either interconnects or cross-node network interfaces. The runtime and architecture support presented in this paper can also be used to support other types of accelerators. We show that the simplified programming interface reduces programming complexity. The research presented in this paper started in 2009. It has been implemented and tested extensively in several generations of HPE runtime systems as well as adopted into the NVIDIA GPU hardware and drivers for CUDA 4.0 and beyond since 2011. The availability of real hardware that support key HPE features gives rise to a rare opportunity for studying the effectiveness of the hardware support by running important benchmarks on real runtime and hardware. Experimental results show that in a exemplar heterogeneous system, peer DMA and double-buffering, pinned buffers, and software techniques can improve the inter-accelerator data communication bandwidth by 2×. They can also improve the execution speed by 1.6× for a 3D finite difference, 2.5× for 1D FFT, and 1.6× for merge sort, all measured on real hardware. The proposed architecture support enables the HPE runtime to transparently deploy these optimizations under simple portable user code, allowing system designers to freely employ devices of different capabilities. We further argue that simple interfaces such as HPE are needed for most applications to benefit from advanced hardware features in practice.
Javier Cabezas, Isaac Gelado, John E. Stone, Nacho Navarro, David Blair Kirk, Wen-Mei W. Hwu
IEEE Trans. Parallel Distributed Syst.5
2008 Visualization and Analysis of GPU Summer School Applicants and Participants
abstract
With the development of petascale computing systems, a long-term effort is needed to educate and train the next generation of researchers. As part of its graduate education component, the Virtual School of Computational Science and Engineering held a summer school in August 2008 entitled "Accelerators for Science and Engineering Applications," providing participants with knowledge and hands-on experience with graphics processing units (GPUs). In this paper, we present visualizations exploring the broad spectrum of summer school applicants and participants. We examine demographic information of the overall applicant pool, accepted and attending applicants, and remote participants, as well as apply hierarchical clustering and rule association techniques to all applicant data. These statistical and data mining analyses demonstrate the wide range of fields of study where research applications can be readily accelerated through the use of massively parallel computing resources.
Elaine Wah, Loretta Auvil, Umesh Thakkar, Wen-Mei W. Hwu, David Blair Kirk, Thom H. Dunning, Sharon C. Glotzer
eScience6
2008 Optimization principles and application performance evaluation of a multithreaded GPU using CUDA
abstract
GPUs have recently attracted the attention of many application developers as commodity data-parallel coprocessors. The newest generations of GPU architecture provide easier programmability and increased generality while maintaining the tremendous memory bandwidth and computational power of traditional GPUs. This opportunity should redirect efforts in GPGPU research from ad hoc porting of applications to establishing principles and strategies that allow efficient mapping of computation to graphics hardware. In this work we discuss the GeForce 8800 GTX processor's organization, features, and generalized optimization strategies. Key to performance on this platform is using massive multithreading to utilize the large number of cores and hide global memory latency. To achieve this, developers face the challenge of striking the right balance between each thread's resource usage and the number of simultaneously active threads. The resources to manage include the number of registers and the amount of on-chip memory used per thread, number of threads per multiprocessor, and global memory bandwidth. We also obtain increased performance by reordering accesses to off-chip memory to combine requests to the same or contiguous memory locations and apply classical optimizations to reduce the number of executed operations. We apply these strategies across a variety of applications and domains and achieve between a 10.5X to 457X speedup in kernel codes and between 1.16X to 431X total application speedup.
Shane Ryoo, Christopher I. Rodrigues, Sara S. Baghsorkhi, Sam S. Stone, David Blair Kirk, Wen-Mei W. Hwu
PPoPP5
2007 NVIDIA cuda software and gpu parallel computing architecture
abstract
In the past, graphics processors were special purpose hardwired application accelerators, suitable only for conventional rasterization-style graphics applications. Modern GPUs are now fully programmable, massively parallel floating point processors. This talk will describe NVIDIA's massively multithreaded computing architecture and CUDA software for GPU computing. The architecture is a scalable, highly parallel architecture that delivers high throughput for data-intensive processing. Although not truly general-purpose processors, GPUs can now be used for a wide variety of compute-intensive applications beyond graphic.
David Blair Kirk
ISMM1
2006 Processor architecture: too much parallelism?
abstract
CPUs and GPUs have evolved considerably in the past few years, and the pace of change and evolution in processor architecture is likely to increase. Constraints of excess heat dissipation and power consumption have forced a radical rethinking of microprocessor architecture, from the headlong pursuit of GHz clock rates to multicore and multithreaded approaches. The demands of graphics vertex and pixel processing as well as more general non-graphics applications have driven GPUs to be powerful data-parallel floating point processing engines. Research and development in programming languages and environments has not kept pace with the changes in processors. Consequently, computer science and computer engineering research and education is not addressing important problems, or preparing students well for today's computer industr.This talk will provide some historical and architectural perspective on data-parallel GPU architectures, and will attempt to make some trend predictions for the future of GPUs. We will then provide some examples of successes and failures in mapping parallel algorithms to these architectures. Finally, we will conclude with some calls to action in research and education, to improve the utilization of these ubiquitous and powerful parallel machines.
David Blair Kirk
PACT1
2004 Panel 3: The Future Visualization Platform
abstract
Advances in graphics hardware and rendering methods are shaping the future of visualization. For example, programmable graphics processors are redefining the traditional visualization cycle. In some cases it is now possible to run the computational simulation and associated visualization side-by-side on the same chip. Moreover, global illumination and non-photorealistic effects promise to deliver imagery which enables greater insight into high resolution, multivariate, and higher-dimensional data. The panelists will offer distinct viewpoints on the direction of future graphics hardware and its potential impact on visualization, and on the nature of advanced visualizationrelated tools and techniques. Presentation of these viewpoints will be followed by audience participation in the form of a question and answer period moderated by the panel organizer.
Greg Johnson, David S. Ebert, Charles D. Hansen, David Blair Kirk, Bill Mark, Hanspeter Pfister
IEEE Visualization4
1993 Implementing rotation matrix constraints in Analog VLSI
abstract
We describe an algorithm for continuously producing a 3x3 rotation matrix from 9 changing input values that form an approximate rotation matrix, and we describe the implementation of that constraint in analog VLSI circuits. This constraint is useful when some source (e.g., sensors, a modeling system, other analog VLSI circuits), produces a potentially "imperfect" matrix, to be used as a rotation. The9 values are continuously adjustedover time to find the "nearest" true rotation matrix, based on a leastsquares metric. The constraint solution is implemented in analog VLSI circuitry; with appropriate design methodology [Kirk 93], adaptive analog VLSI is a fast, accurate, and low-power computational medium. The implementation is potentially interesting to the graphics community because there is an opportunity to apply adaptive analog VLSI to many other graphics problems. CR Categories and Subject Descriptors: C.1.2---[Processor Architectures]: Multiprocessors - parallel processors; C.1.3...
David Blair Kirk, Alan H. Barr
SIGGRAPH1
1992 An Analog VLSI Chip for Radial Basis Functions
Janeen Anderson, John C. Platt, David Blair Kirk
NIPS3
1991 Constrained Optimization Applied to the Parameter Setting Problem for Analog Circuits
David Blair Kirk, Kurt W. Fleischer, Lloyd Watts, Alan H. Barr
NIPS1
1991 Unbiased sampling techniques for image synthesis
abstract
We examine a class of adaptive sampling techniques employed in image synthesis and show that those commonly used for efficient anti-aliasing are statistically biased. This bias is dependent upon the image function being sampled as well as the strategy for determining the number of samples to use. It is most prominent in areas of high contrast and is attributable to early stages of sampling systematically favoring one extreme or the other. If the expected outcome of the entire adaptive sampling algorithm is considered, we find that the bias of the early decisions is still present in the final estimator. We propose an alternative strategy for performing adaptive sampling that is unbiased but potentially more costly. We conclude that it may not always be practical to mitigate this source of bias, but as a source of error it should be considered when high accuracy and image fidelity are a central concern.
David Blair Kirk, James Arvo
SIGGRAPH1
1990 SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000
abstract
SMART, a technique for providing predictable cache performance for real-time systems with priority-based preemptive scheduling, is presented. The technique is implemented in a R3000 cache design. The value density acceleration (VDA) cache allocation algorithm is also introduced, and shown to be suitable for run-time cache allocation.>
David Blair Kirk, Jay K. Strosnider
RTSS1
1990 Particle transport and image synthesis
abstract
The rendering equation is similar to the linear Boltzmann equation which has been widely studied in physics and nuclear engineering. Consequently, many of the powerful techniques which have been developed in these fields can be applied to problems in image synthesis. In this paper we adapt several statistical techniques commonly used in neutron transport to stochastic ray tracing and, more generally, to Monte Carlo solution of the rendering equation. First, we describe a technique known as Russian roulette which can be used to terminate the recursive tracing of rays without introducing statistical bias. We also examine the practice of creating ray trees in classical ray tracing in the light of a well-known technique in particle transport known as splitting. We show that neither ray trees nor paths as described in [10] constitute an optimal sampling plan in themselves and that a hybrid may be more efficient.
James Arvo, David Blair Kirk
SIGGRAPH2
1990 The rendering architecture of the DN10000VS
abstract
The Appollo DN10000VS treats graphics as an integral part of the system architecture. Graphics requirements influence the entire system design. All floating-point computations for graphics are performed by the CPU(s), while rasterizing is handled by simplified hardware having no microcode. We decided to support alpha buffering, quadratic interpolation, and texture mapping directly in hardware. This partitioning reduces the cost of a high-end workstation, without sacrificing high rendering quality and performance. This paper describes some of the design trade-offs which led to the final system design.
David Blair Kirk, Douglas Voorhies
SIGGRAPH1
1989 SMART (Strategic Memory Allocation for Real-Time) Cache Design
abstract
A discussion is presented as to why the present approach to cache architecture design results in unpredictable performance improvements in real-time systems with priority-based preemptive scheduling algorithms. The SMART cache design is shown to be compatible with the goals of scheduling in a real-time system. The results of this research provide a scheme not only for utilizing the performance enhancement provided by hierarchical memory designs, but also for fine tuning these enhancements to provide increased benefit to the desired scheduling goal.>
David Blair Kirk
RTSS1
1988 The White Dwarf: A High-Performance Application-Specific Processor
abstract
The design and implementation of a high-performance special-purpose processor, called the White Dwarf, or accelerating finite-element analysis algorithms is presented. The White Dwarf CPU contains two Am2935 32-bit floating-point processors and one Am29332 32-bit arithmetic logic unit (ALU), and uses a wide-instruction-word architecture in which the application algorithm is directly implemented in microcode. The entire system is VME-bus compatible and interfaces with a Sun 3/160 host. The system's potential peak performance is 20 MFLOPS (million floating-point operations per second) a sustained computation rate in excess of 15 MFLOPS is expected. A potential speedup of between one and two orders of magnitude is possible. With a fully populated memory subsystem, the White Dwarf can accommodate finite-element problems involving up to half a million nodes. The system is designed using an approach called application-specific processor design (ASPD). A retargetable compiler has been developed which is capable of generating highly parallel and efficient code for the White Dwarf and other processors with similar architecture. System debug/integration is in progress; a highly useful system is expected.>
Andrew Wolfe, Maurício Breternitz, Chriss Stephens, A. L. Ting, David Blair Kirk, Ronald P. Bianchini Jr., John Paul Shen
ISCA5
1988 Priority-Driven, Preemptive I/O Controllers for Real-Time Systems
abstract
The effect of three I/O controller architectures on schedulable utilization, which is the highest attainable resource utilization at or below which all deadlines can be guaranteed, is examined. FIFO (first-in-first-out) request queuing, priority queuing, and priority queuing with preemptable service are simulated for a range of CPU computation to I/O traffic ratios. The results show that, for I/O-bound task sets and zero preemption costs, priority queuing with preemptable service can provide a level of schedulable utilization 35% higher than that attainable with FIFO queuing, and 20% higher than priority queuing and nonpreemptable service. Although the potential gain for priority queuing with preemptable service is large, further simulations that incorporate a time penalty for each preemption show that the gain is very sensitive to preemption cost. With preemption cost represented as a ratio of preemption time to the minimum-task period, the level of schedulable utilization for priority queuing with preemptable service degrades to that of priority queuing with nonpreemptible service, for a preemption cost ratio of 0.04. A high-level design of a preemptable I/O controller is described and the issues determining preemption cost are detailed, along with techniques for its minimization.>
Brinkley Sprunt, David Blair Kirk, Lui Sha
ISCA2
1988 Process Dependent Static Cache Partitioning for Real-Time Systems
abstract
The author investigates the use of a priori knowledge of program behavior to partition an instruction cache of size C into a static partition of size S and an LRU partition of size C-S. The value of S is task-dependent and is nonzero for most programs running on the system. Example programs are presented, and their behavior in various size caches is discussed. Cache partitions are generated and evaluated to determine the increase in cache performance and predictability. A high-level hardware design is presented that provides the desired partitioning scheme.>
David Blair Kirk
RTSS1
1988 Virtual graphics
abstract
Graphics can be implemented as a virtual system resource. This abstraction appears to each application on a multiprocessing workstation as a dedicated rendering and display pipeline. A variety of simple mechanisms support the simultaneous display of different types of images and eliminate the need for low-level device driver software. They permit applications to embed graphics instructions directly in their code. The abstraction allows for cleaner software design, higher performance, and effective concurrent use of the display by several applications.
Douglas Voorhies, David Blair Kirk, Olin Lathrop
SIGGRAPH2
1987 Fast ray tracing by ray classification
abstract
We describe a new approach to ray tracing which drastically reduces the number of ray-object and ray-bounds intersection calculations by means of 5-dimensional space subdivision. Collections of rays originating from a common 3D rectangular volume and directed through a 2D solid angle are represented as hypercubes in 5-space. A 5D volume bounding the space of rays is dynamically subdivided into hypercubes, each linked to a set of objects which are candidates for intersection. Rays are classified into unique hypercubes and checked for intersection with the associated candidate object set. We compare several techniques for object extent testing, including boxes, spheres, plane-sets, and convex polyhedra. In addition, we examine optimizations made possible by the directional nature of the algorithm, such as sorting, caching and backface culling. Results indicate that this algorithm significantly outperforms previous ray tracing techniques, especially for comples environments.
James Arvo, David Blair Kirk
SIGGRAPH2
1987 The simulation of natural features using cone tracing
David Blair Kirk
Vis. Comput.1