EDBT 2026 Demo / reviewers in the wild / expert
David Blair Kirk
dblp:46/5249
· DBLP profile ↗
20ranked-venue papers
10as first author
0since 2021 · last 2015
0000-0002-4887-5098ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-authorHuman-computer interaction and ubiquitous computing · 7 · 3 first-authorSystems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
12 papers |
GPUs and heterogeneous computing · 68% Memory systems · 14% Integrated circuit design · 10% | |
| Computer graphics and multimedia
6 papers |
Rendering · 66% Geometric modeling and processing · 20% Visual content generation and editing · 14% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
heterogeneous parallel programming |
0.2 | 1 | 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015 |
GPUs and heterogeneous computing › GPU kernel optimization
CUDA optimization |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
GPUs and heterogeneous computing
GPU performance optimization |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
GPUs and heterogeneous computing
GPU programming |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
Memory systems
memory access optimization |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
Integrated circuit design
interconnect |
0.1 | 1 | 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015 |
Integrated circuit design › analog and mixed-signal circuits
analog VLSI |
0.0 | 2 | 1993 | Implementing rotation matrix constraints in Analog VLSI · SIGGRAPH 1993 An Analog VLSI Chip for Radial Basis Functions · NIPS 1992 |
Memory systems
cache |
0.0 | 3 | 1990 | SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990 SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989 Process Dependent Static Cache Partitioning for Real-Time Systems · RTSS 1988 |
Integrated circuit design
analog and mixed-signal circuits |
0.0 | 2 | 1993 | Implementing rotation matrix constraints in Analog VLSI · SIGGRAPH 1993 An Analog VLSI Chip for Radial Basis Functions · NIPS 1992 |
Memory systems › cache management
cache allocation |
0.0 | 2 | 1990 | SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990 SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989 |
Memory systems
cache design |
0.0 | 2 | 1990 | SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990 SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989 |
Electronic design automation › timing analysis
predictable cache performance |
0.0 | 2 | 1990 | SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990 SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989 |
Embedded and real-time systems
real-time scheduling |
0.0 | 3 | 1990 | Priority-Driven, Preemptive I/O Controllers for Real-Time Systems · ISCA 1988 SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990 SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989 |
Rendering › sampling
adaptive sampling |
0.0 | 1 | 1991 | Unbiased sampling techniques for image synthesis · SIGGRAPH 1991 |
Visual content generation and editing
image generation |
0.0 | 1 | 1991 | Unbiased sampling techniques for image synthesis · SIGGRAPH 1991 |
Geometric modeling and processing
unbiased sampling |
0.0 | 1 | 1991 | Unbiased sampling techniques for image synthesis · SIGGRAPH 1991 |
Integrated circuit design › analog and mixed-signal circuits
analog circuit design |
0.0 | 1 | 1991 | Constrained Optimization Applied to the Parameter Setting Problem for Analog Circuits · NIPS 1991 |
Mathematical optimization
constrained optimization |
0.0 | 1 | 1991 | Constrained Optimization Applied to the Parameter Setting Problem for Analog Circuits · NIPS 1991 |
Rendering › rasterization
hardware rasterization |
0.0 | 1 | 1990 | The rendering architecture of the DN10000VS · SIGGRAPH 1990 |
Rendering
monte carlo rendering |
0.0 | 1 | 1990 | Particle transport and image synthesis · SIGGRAPH 1990 |
Rendering
rasterization |
0.0 | 1 | 1990 | The rendering architecture of the DN10000VS · SIGGRAPH 1990 |
Rendering
ray tracing |
0.0 | 2 | 1990 | Fast ray tracing by ray classification · SIGGRAPH 1987 Particle transport and image synthesis · SIGGRAPH 1990 |
Hardware accelerators and domain-specific architectures › domain-specific accelerator
application-specific processor |
0.0 | 1 | 1988 | The White Dwarf: A High-Performance Application-Specific Processor · ISCA 1988 |
Memory systems › cache management
cache partitioning |
0.0 | 1 | 1988 | Process Dependent Static Cache Partitioning for Real-Time Systems · RTSS 1988 |
Memory systems › cache › CPU cache
instruction cache |
0.0 | 1 | 1988 | Process Dependent Static Cache Partitioning for Real-Time Systems · RTSS 1988 |
Embedded and real-time systems › real-time scheduling › priority scheduling
preemptive priority scheduling |
0.0 | 2 | 1990 | SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000 · RTSS 1990 SMART (Strategic Memory Allocation for Real-Time) Cache Design · RTSS 1989 |
Geometric modeling and processing
rotation representation |
0.0 | 1 | 1993 | Implementing rotation matrix constraints in Analog VLSI · SIGGRAPH 1993 |
Rendering › ray tracing
monte carlo ray tracing |
0.0 | 1 | 1990 | Particle transport and image synthesis · SIGGRAPH 1990 |
Embedded and real-time systems
real-time system design |
0.0 | 1 | 1988 | Priority-Driven, Preemptive I/O Controllers for Real-Time Systems · ISCA 1988 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1988 | The White Dwarf: A High-Performance Application-Specific Processor · ISCA 1988 |
Methods — techniques the papers use, named apart from their topics
pinned buffers · 0.2peer DMA · 0.2double buffering · 0.2register allocation · 0.1memory coalescing · 0.1latency hiding · 0.1least squares · 0.0adaptive analog VLSI · 0.0radial basis functions · 0.0analog VLSI · 0.0statistical bias analysis · 0.0constrained optimization · 0.0anti-aliasing · 0.0texture mapping · 0.0splitting · 0.0russian roulette · 0.0quadratic interpolation · 0.0importance sampling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator ApplicationsabstractHeterogeneous parallel computing applications often process large data sets that require multiple GPUs to jointly meet their needs for physical memory capacity and compute throughput. However, the lack of high-level abstractions in previous heterogeneous parallel programming models force programmers to resort to multiple code versions, complex data copy steps and synchronization schemes when exchanging data between multiple GPU devices, which results in high software development cost, poor maintainability, and even poor performance. This paper describes the HPE runtime system, and the associated architecture support, which enables a simple, efficient programming interface for exchanging data between multiple GPUs through either interconnects or cross-node network interfaces. The runtime and architecture support presented in this paper can also be used to support other types of accelerators. We show that the simplified programming interface reduces programming complexity. The research presented in this paper started in 2009. It has been implemented and tested extensively in several generations of HPE runtime systems as well as adopted into the NVIDIA GPU hardware and drivers for CUDA 4.0 and beyond since 2011. The availability of real hardware that support key HPE features gives rise to a rare opportunity for studying the effectiveness of the hardware support by running important benchmarks on real runtime and hardware. Experimental results show that in a exemplar heterogeneous system, peer DMA and double-buffering, pinned buffers, and software techniques can improve the inter-accelerator data communication bandwidth by 2×. They can also improve the execution speed by 1.6× for a 3D finite difference, 2.5× for 1D FFT, and 1.6× for merge sort, all measured on real hardware. The proposed architecture support enables the HPE runtime to transparently deploy these optimizations under simple portable user code, allowing system designers to freely employ devices of different capabilities. We further argue that simple interfaces such as HPE are needed for most applications to benefit from advanced hardware features in practice. Javier Cabezas, Isaac Gelado, John E. Stone, Nacho Navarro, David Blair Kirk, Wen-Mei W. Hwu |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2008 | Visualization and Analysis of GPU Summer School Applicants and ParticipantsabstractWith the development of petascale computing systems, a long-term effort is needed to educate and train the next generation of researchers. As part of its graduate education component, the Virtual School of Computational Science and Engineering held a summer school in August 2008 entitled "Accelerators for Science and Engineering Applications," providing participants with knowledge and hands-on experience with graphics processing units (GPUs). In this paper, we present visualizations exploring the broad spectrum of summer school applicants and participants. We examine demographic information of the overall applicant pool, accepted and attending applicants, and remote participants, as well as apply hierarchical clustering and rule association techniques to all applicant data. These statistical and data mining analyses demonstrate the wide range of fields of study where research applications can be readily accelerated through the use of massively parallel computing resources. Elaine Wah, Loretta Auvil, Umesh Thakkar, Wen-Mei W. Hwu, David Blair Kirk, Thom H. Dunning, Sharon C. Glotzer |
eScience | 6 |
| 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDAabstractGPUs have recently attracted the attention of many application developers as commodity data-parallel coprocessors. The newest generations of GPU architecture provide easier programmability and increased generality while maintaining the tremendous memory bandwidth and computational power of traditional GPUs. This opportunity should redirect efforts in GPGPU research from ad hoc porting of applications to establishing principles and strategies that allow efficient mapping of computation to graphics hardware. In this work we discuss the GeForce 8800 GTX processor's organization, features, and generalized optimization strategies. Key to performance on this platform is using massive multithreading to utilize the large number of cores and hide global memory latency. To achieve this, developers face the challenge of striking the right balance between each thread's resource usage and the number of simultaneously active threads. The resources to manage include the number of registers and the amount of on-chip memory used per thread, number of threads per multiprocessor, and global memory bandwidth. We also obtain increased performance by reordering accesses to off-chip memory to combine requests to the same or contiguous memory locations and apply classical optimizations to reduce the number of executed operations. We apply these strategies across a variety of applications and domains and achieve between a 10.5X to 457X speedup in kernel codes and between 1.16X to 431X total application speedup. Shane Ryoo, Christopher I. Rodrigues, Sara S. Baghsorkhi, Sam S. Stone, David Blair Kirk, Wen-Mei W. Hwu |
PPoPP | 5 |
| 2007 | NVIDIA cuda software and gpu parallel computing architectureabstractIn the past, graphics processors were special purpose hardwired application accelerators, suitable only for conventional rasterization-style graphics applications. Modern GPUs are now fully programmable, massively parallel floating point processors. This talk will describe NVIDIA's massively multithreaded computing architecture and CUDA software for GPU computing. The architecture is a scalable, highly parallel architecture that delivers high throughput for data-intensive processing. Although not truly general-purpose processors, GPUs can now be used for a wide variety of compute-intensive applications beyond graphic. David Blair Kirk |
ISMM | 1 |
| 2006 | Processor architecture: too much parallelism?abstractCPUs and GPUs have evolved considerably in the past few years, and the pace of change and evolution in processor architecture is likely to increase. Constraints of excess heat dissipation and power consumption have forced a radical rethinking of microprocessor architecture, from the headlong pursuit of GHz clock rates to multicore and multithreaded approaches. The demands of graphics vertex and pixel processing as well as more general non-graphics applications have driven GPUs to be powerful data-parallel floating point processing engines. Research and development in programming languages and environments has not kept pace with the changes in processors. Consequently, computer science and computer engineering research and education is not addressing important problems, or preparing students well for today's computer industr.This talk will provide some historical and architectural perspective on data-parallel GPU architectures, and will attempt to make some trend predictions for the future of GPUs. We will then provide some examples of successes and failures in mapping parallel algorithms to these architectures. Finally, we will conclude with some calls to action in research and education, to improve the utilization of these ubiquitous and powerful parallel machines. David Blair Kirk |
PACT | 1 |
| 2004 | Panel 3: The Future Visualization PlatformabstractAdvances in graphics hardware and rendering methods are shaping the future of visualization. For example, programmable graphics processors are redefining the traditional visualization cycle. In some cases it is now possible to run the computational simulation and associated visualization side-by-side on the same chip. Moreover, global illumination and non-photorealistic effects promise to deliver imagery which enables greater insight into high resolution, multivariate, and higher-dimensional data. The panelists will offer distinct viewpoints on the direction of future graphics hardware and its potential impact on visualization, and on the nature of advanced visualizationrelated tools and techniques. Presentation of these viewpoints will be followed by audience participation in the form of a question and answer period moderated by the panel organizer. Greg Johnson, David S. Ebert, Charles D. Hansen, David Blair Kirk, Bill Mark, Hanspeter Pfister |
IEEE Visualization | 4 |
| 1993 | Implementing rotation matrix constraints in Analog VLSIabstractWe describe an algorithm for continuously producing a 3x3 rotation matrix from 9 changing input values that form an approximate rotation matrix, and we describe the implementation of that constraint in analog VLSI circuits. This constraint is useful when some source (e.g., sensors, a modeling system, other analog VLSI circuits), produces a potentially "imperfect" matrix, to be used as a rotation. The9 values are continuously adjustedover time to find the "nearest" true rotation matrix, based on a leastsquares metric. The constraint solution is implemented in analog VLSI circuitry; with appropriate design methodology [Kirk 93], adaptive analog VLSI is a fast, accurate, and low-power computational medium. The implementation is potentially interesting to the graphics community because there is an opportunity to apply adaptive analog VLSI to many other graphics problems. CR Categories and Subject Descriptors: C.1.2---[Processor Architectures]: Multiprocessors - parallel processors; C.1.3... David Blair Kirk, Alan H. Barr |
SIGGRAPH | 1 |
| 1992 | An Analog VLSI Chip for Radial Basis Functions
Janeen Anderson, John C. Platt, David Blair Kirk |
NIPS | 3 |
| 1991 | Constrained Optimization Applied to the Parameter Setting Problem for Analog Circuits
David Blair Kirk, Kurt W. Fleischer, Lloyd Watts, Alan H. Barr |
NIPS | 1 |
| 1991 | Unbiased sampling techniques for image synthesisabstractWe examine a class of adaptive sampling techniques employed in image synthesis and show that those commonly used for efficient anti-aliasing are statistically biased. This bias is dependent upon the image function being sampled as well as the strategy for determining the number of samples to use. It is most prominent in areas of high contrast and is attributable to early stages of sampling systematically favoring one extreme or the other. If the expected outcome of the entire adaptive sampling algorithm is considered, we find that the bias of the early decisions is still present in the final estimator. We propose an alternative strategy for performing adaptive sampling that is unbiased but potentially more costly. We conclude that it may not always be practical to mitigate this source of bias, but as a source of error it should be considered when high accuracy and image fidelity are a central concern. David Blair Kirk, James Arvo |
SIGGRAPH | 1 |
| 1990 | SMART (Strategic Memory Allocation for Real-Time) Cache Design Using the MIPS R3000abstractSMART, a technique for providing predictable cache performance for real-time systems with priority-based preemptive scheduling, is presented. The technique is implemented in a R3000 cache design. The value density acceleration (VDA) cache allocation algorithm is also introduced, and shown to be suitable for run-time cache allocation.> David Blair Kirk, Jay K. Strosnider |
RTSS | 1 |
| 1990 | Particle transport and image synthesisabstractThe rendering equation is similar to the linear Boltzmann equation which has been widely studied in physics and nuclear engineering. Consequently, many of the powerful techniques which have been developed in these fields can be applied to problems in image synthesis. In this paper we adapt several statistical techniques commonly used in neutron transport to stochastic ray tracing and, more generally, to Monte Carlo solution of the rendering equation. First, we describe a technique known as Russian roulette which can be used to terminate the recursive tracing of rays without introducing statistical bias. We also examine the practice of creating ray trees in classical ray tracing in the light of a well-known technique in particle transport known as splitting. We show that neither ray trees nor paths as described in [10] constitute an optimal sampling plan in themselves and that a hybrid may be more efficient. James Arvo, David Blair Kirk |
SIGGRAPH | 2 |
| 1990 | The rendering architecture of the DN10000VSabstractThe Appollo DN10000VS treats graphics as an integral part of the system architecture. Graphics requirements influence the entire system design. All floating-point computations for graphics are performed by the CPU(s), while rasterizing is handled by simplified hardware having no microcode. We decided to support alpha buffering, quadratic interpolation, and texture mapping directly in hardware. This partitioning reduces the cost of a high-end workstation, without sacrificing high rendering quality and performance. This paper describes some of the design trade-offs which led to the final system design. David Blair Kirk, Douglas Voorhies |
SIGGRAPH | 1 |
| 1989 | SMART (Strategic Memory Allocation for Real-Time) Cache DesignabstractA discussion is presented as to why the present approach to cache architecture design results in unpredictable performance improvements in real-time systems with priority-based preemptive scheduling algorithms. The SMART cache design is shown to be compatible with the goals of scheduling in a real-time system. The results of this research provide a scheme not only for utilizing the performance enhancement provided by hierarchical memory designs, but also for fine tuning these enhancements to provide increased benefit to the desired scheduling goal.> David Blair Kirk |
RTSS | 1 |
| 1988 | The White Dwarf: A High-Performance Application-Specific ProcessorabstractThe design and implementation of a high-performance special-purpose processor, called the White Dwarf, or accelerating finite-element analysis algorithms is presented. The White Dwarf CPU contains two Am2935 32-bit floating-point processors and one Am29332 32-bit arithmetic logic unit (ALU), and uses a wide-instruction-word architecture in which the application algorithm is directly implemented in microcode. The entire system is VME-bus compatible and interfaces with a Sun 3/160 host. The system's potential peak performance is 20 MFLOPS (million floating-point operations per second) a sustained computation rate in excess of 15 MFLOPS is expected. A potential speedup of between one and two orders of magnitude is possible. With a fully populated memory subsystem, the White Dwarf can accommodate finite-element problems involving up to half a million nodes. The system is designed using an approach called application-specific processor design (ASPD). A retargetable compiler has been developed which is capable of generating highly parallel and efficient code for the White Dwarf and other processors with similar architecture. System debug/integration is in progress; a highly useful system is expected.> Andrew Wolfe, Maurício Breternitz, Chriss Stephens, A. L. Ting, David Blair Kirk, Ronald P. Bianchini Jr., John Paul Shen |
ISCA | 5 |
| 1988 | Priority-Driven, Preemptive I/O Controllers for Real-Time SystemsabstractThe effect of three I/O controller architectures on schedulable utilization, which is the highest attainable resource utilization at or below which all deadlines can be guaranteed, is examined. FIFO (first-in-first-out) request queuing, priority queuing, and priority queuing with preemptable service are simulated for a range of CPU computation to I/O traffic ratios. The results show that, for I/O-bound task sets and zero preemption costs, priority queuing with preemptable service can provide a level of schedulable utilization 35% higher than that attainable with FIFO queuing, and 20% higher than priority queuing and nonpreemptable service. Although the potential gain for priority queuing with preemptable service is large, further simulations that incorporate a time penalty for each preemption show that the gain is very sensitive to preemption cost. With preemption cost represented as a ratio of preemption time to the minimum-task period, the level of schedulable utilization for priority queuing with preemptable service degrades to that of priority queuing with nonpreemptible service, for a preemption cost ratio of 0.04. A high-level design of a preemptable I/O controller is described and the issues determining preemption cost are detailed, along with techniques for its minimization.> Brinkley Sprunt, David Blair Kirk, Lui Sha |
ISCA | 2 |
| 1988 | Process Dependent Static Cache Partitioning for Real-Time SystemsabstractThe author investigates the use of a priori knowledge of program behavior to partition an instruction cache of size C into a static partition of size S and an LRU partition of size C-S. The value of S is task-dependent and is nonzero for most programs running on the system. Example programs are presented, and their behavior in various size caches is discussed. Cache partitions are generated and evaluated to determine the increase in cache performance and predictability. A high-level hardware design is presented that provides the desired partitioning scheme.> David Blair Kirk |
RTSS | 1 |
| 1988 | Virtual graphicsabstractGraphics can be implemented as a virtual system resource. This abstraction appears to each application on a multiprocessing workstation as a dedicated rendering and display pipeline. A variety of simple mechanisms support the simultaneous display of different types of images and eliminate the need for low-level device driver software. They permit applications to embed graphics instructions directly in their code. The abstraction allows for cleaner software design, higher performance, and effective concurrent use of the display by several applications. Douglas Voorhies, David Blair Kirk, Olin Lathrop |
SIGGRAPH | 2 |
| 1987 | Fast ray tracing by ray classificationabstractWe describe a new approach to ray tracing which drastically reduces the number of ray-object and ray-bounds intersection calculations by means of 5-dimensional space subdivision. Collections of rays originating from a common 3D rectangular volume and directed through a 2D solid angle are represented as hypercubes in 5-space. A 5D volume bounding the space of rays is dynamically subdivided into hypercubes, each linked to a set of objects which are candidates for intersection. Rays are classified into unique hypercubes and checked for intersection with the associated candidate object set. We compare several techniques for object extent testing, including boxes, spheres, plane-sets, and convex polyhedra. In addition, we examine optimizations made possible by the directional nature of the algorithm, such as sorting, caching and backface culling. Results indicate that this algorithm significantly outperforms previous ray tracing techniques, especially for comples environments. James Arvo, David Blair Kirk |
SIGGRAPH | 2 |
| 1987 | The simulation of natural features using cone tracing
David Blair Kirk |
Vis. Comput. | 1 |