Theresa Foley

dblp:65/2380 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
8 papers
Rendering · 100%
Software engineering, system software, and programming languages
4 papers
Programming languages and type systems · 60% Compilers and program optimization · 33% Requirements engineering and software design · 7%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 44% Parallel and multicore computing · 34% Distributed systems · 22%

Topics — the 14 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
real-time rendering
1.162018
Shader components: modular and high performance shader development · ACM Trans. Graph. 2017
A system for rapid exploration of shader optimization choices · ACM Trans. Graph. 2016
Efficient GPU rendering of subdivision surfaces using adaptive quadtrees · ACM Trans. Graph. 2016
Rendering › shader programming
shading languages
0.932018
Slang: language mechanisms for extensible real-time shading systems · ACM Trans. Graph. 2018
Shader components: modular and high performance shader development · ACM Trans. Graph. 2017
A system for rapid exploration of shader optimization choices · ACM Trans. Graph. 2016
Rendering › rendering optimization
shader optimization
0.522016
A system for rapid exploration of shader optimization choices · ACM Trans. Graph. 2016
A system for rapid, automatic shader level-of-detail · ACM Trans. Graph. 2015
Rendering › level of detail
shader level-of-detail
0.322016
A system for rapid, automatic shader level-of-detail · ACM Trans. Graph. 2015
A system for rapid exploration of shader optimization choices · ACM Trans. Graph. 2016
Rendering
GPU rendering
0.212016
Efficient GPU rendering of subdivision surfaces using adaptive quadtrees · ACM Trans. Graph. 2016
Rendering › surface rendering
subdivision surface rendering
0.212016
Efficient GPU rendering of subdivision surfaces using adaptive quadtrees · ACM Trans. Graph. 2016
Requirements engineering and software design › modularity
modularity mechanisms
0.112017
Shader components: modular and high performance shader development · ACM Trans. Graph. 2017
Rendering
parallel rendering
0.112008
Parallel computing for graphics · SIGGRAPH ASIA Courses 2008
Rendering
graphics pipeline
0.112015
A system for rapid, automatic shader level-of-detail · ACM Trans. Graph. 2015
Parallel and multicore computing › parallel algorithms › parallel primitives
data-parallel primitives
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
GPUs and heterogeneous computing
GPU computing
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
GPUs and heterogeneous computing › GPU programming
GPU programming models
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
Distributed systems
stream processing
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004
Compilers and program optimization › accelerator compilation
GPU compiler
0.012004
Brook for GPUs: stream computing on graphics hardware · ACM Trans. Graph. 2004

Methods — techniques the papers use, named apart from their topics

design space exploration · 0.8interface extensions · 0.7generics with interface constraints · 0.7associated types · 0.7static specialization · 0.6shader components · 0.6metaprogramming · 0.4meta-programming · 0.4quadtree traversal · 0.2compiler framework · 0.2GPU tessellation hardware · 0.2stream programming · 0.1compiler and runtime abstraction · 0.1parallel programming · 0.1
YearPublicationVenuePosition
2019 Staged metaprogramming for shader system development
abstract
The shader system for a modern game engine comprises much more than just compilation of source code to executable kernels. Shaders must also be exposed to art tools, interfaced with engine code, and specialized for performance. Engines typically address each of these tasks in an ad hoc fashion, without a unifying abstraction. The alternative of developing a more powerful compiler framework is prohibitive for most engines. In this paper, we identify staged metaprogramming as a unifying abstraction and implementation strategy to develop a powerful shader system with modest effort. By using a multi-stage language to perform metaprogramming at compile time, engine-specific code can consume, analyze, transform, and generate shader code that will execute at runtime. Staged metaprogramming reduces the effort required to implement a shader system that provides earlier error detection, avoids repeat declarations of shader parameters, and explores opportunities to improve performance. To demonstrate the value of this approach, we design and implement a shader system, called Selos, built using staged metaprogramming. In our system, shader and application code are written in the same language and can share types and functions. We implement a design space exploration framework for Selos that investigates static versus dynamic composition of shader features, exploring the impact of shader specialization in a deferred renderer. Staged metaprogramming allows Selos to provide compelling features with a simple implementation.
Kerry A. Seitz Jr., Theresa Foley, Serban D. Porumbescu, John D. Owens
ACM Trans. Graph.2
2018 Slang: language mechanisms for extensible real-time shading systems
abstract
Designers of real-time rendering engines must balance the conflicting goals of maintaining clear, extensible shading systems and achieving high rendering performance. In response, engine architects have established effective design patterns for authoring shading systems, and developed engine-specific code synthesis tools, ranging from preprocessor hacking to domain-specific shading languages, to productively implement these patterns. The problem is that proprietary tools add significant complexity to modern engines, lack advanced language features, and create additional challenges for learning and adoption. We argue that the advantages of engine-specific code generation tools can be achieved using the underlying GPU shading language directly, provided the shading language is extended with a small number of best-practice principles from modern, well-established programming languages. We identify that adding generics with interface constraints, associated types, and interface/structure extensions to existing C-like GPU shading languages enables real-time Tenderer developers to build shading systems that are extensible, maintainable, and execute efficiently on modern GPUs without the need for additional domain-specific tools. We embody these ideas in an extension of HLSL called Slang, and provide a reference design for a large, extensible shader library implemented using Slang's features. We rearchitect an open source Tenderer to use this library and Slang's compiler services, and demonstrate the resulting shading system is substantially simpler, easier to extend with new features, and achieves higher rendering performance than the original HLSL-based implementation.
Yong He 0013, Kayvon Fatahalian, Theresa Foley
ACM Trans. Graph.3
2017 Shader components: modular and high performance shader development
abstract
Modern game engines seek to balance the conflicting goals of high rendering performance and productive software development. To improve CPU performance, the most recent generation of real-time graphics APIs provide new primitives for performing efficient batch updates to shader parameters. However, modern game engines featuring large shader codebases have struggled to take advantage of these benefits. The problem is that even though shader parameters can be organized into efficient modules bound to the pipeline at various frequencies, modern shading languages lack corresponding primitives to organize shader logic (requiring these parameters) into modules as well. The result is that complex shaders are typically compiled to use a monolithic block of parameters, defeating the design, and performance benefits, of the new parameter binding API. In this paper we propose to resolve this mismatch by introducing shader components , a first-class unit of modularity in a shader program that encapsulates a unit of shader logic and the parameters that must be bound when that logic is in use. We show that by building sophisticated shaders out of components, we can retain essential aspects of performance (static specialization of the shader logic in use and efficient update of parameters at component granularity) while maintaining the modular shader code structure that is desirable in today's high-end game engines.
Yong He 0013, Theresa Foley, Teguh Hofstee, Haomin Long, Kayvon Fatahalian
ACM Trans. Graph.2
2016 Efficient GPU rendering of subdivision surfaces using adaptive quadtrees
abstract
We present a novel method for real-time rendering of subdivision surfaces whose goal is to make subdivision faces as easy to render as triangles, points, or lines. Our approach uses standard GPU tessellation hardware and processes each face of a base mesh independently, thus allowing an entire model to be rendered in a single pass. The key idea of our method is to subdivide the u, v domain of each face ahead of time, generating a quadtree structure, and then submit one tessellated primitive per input face. By traversing the quadtree for each post-tessellation vertex, we are able to accurately and efficiently evaluate the limit surface. Our method yields a more uniform tessellation of the surface, and faster rendering, as fewer primitives are submitted. We evaluate our method on a variety of assets, and realize performance that can be three times faster than state-of-the-art approaches. In addition, our streaming formulation makes it easier to integrate subdivision surfaces into applications and shader code written for polygonal models. We illustrate integration of our technique into a full-featured video game engine.
Wade Brainerd, Theresa Foley, Manuel Kraemer, Henry P. Moreton, Matthias Nießner
ACM Trans. Graph.2
2016 A system for rapid exploration of shader optimization choices
abstract
We present Spire, a shading language and compiler framework that facilitates rapid exploration of shader optimization choices (such as frequency reduction and algorithmic approximation) afforded by modern real-time graphics engines. Our design combines ideas from rate-based shader programming with new language features that expand the scope of shader execution beyond traditional GPU hardware pipelines, and enable a diverse set of shader optimizations to be described by a single mechanism: overloading shader terms at various spatio-temporal computation rates provided by the pipeline. In contrast to prior work, neither the shading language's design, nor our compiler framework's implementation, is specific to the capabilities of any one rendering pipeline, thus Spire establishes architectural separation between the shading system and the implementation of modern rendering engines (allowing different rendering pipelines to utilize its services). We demonstrate use of Spire to author complex shaders that are portable across different rendering pipelines and to rapidly explore shader optimization decisions that span multiple compute and graphics passes and even offline asset preprocessing. We further demonstrate the utility of Spire by developing a shader level-of-detail library and shader auto-tuning system on top of its abstractions, and demonstrate rapid, automatic re-optimization of shaders for different target hardware platforms.
Yong He 0013, Theresa Foley, Kayvon Fatahalian
ACM Trans. Graph.2
2015 A system for rapid, automatic shader level-of-detail
abstract
Level-of-detail (LOD) rendering is a key optimization used by modern video game engines to achieve high-quality rendering with fast performance. These LOD systems require simplified shaders, but generating simplified shaders remains largely a manual optimization task for game developers. Prior efforts to automate this process have taken hours to generate simplified shader candidates, making them impractical for use in modern shader authoring workflows for complex scenes. We present an end-to-end system for automatically generating a LOD policy for an input shader. The system operates on shaders used in both forward and deferred rendering pipelines, requires no additional semantic information beyond input shader source code, and in only seconds to minutes generates LOD policies (consisting of simplified shader, the desired LOD distance set, and transition generation) with performance and quality characteristics comparable to custom hand-authored solutions. Our design contributes new shader simplification transforms such as approximate common subexpression elimination and movement of GPU logic to parameter bind-time processing on the CPU, and it uses a greedy search algorithm that employs extensive caching and upfront collection of input shader statistics to rapidly identify simplified shaders with desirable performance-quality trade-offs.
Yong He 0013, Theresa Foley, Natalya Tatarchuk, Kayvon Fatahalian
ACM Trans. Graph.2
2011 Spark: modular, composable shaders for graphics hardware
abstract
In creating complex real-time shaders, programmers should be able to decompose code into independent, localized modules of their choosing. Current real-time shading languages, however, enforce a fixed decomposition into per-pipeline-stage procedures. Program concerns at other scales -- including those that cross-cut multiple pipeline stages -- cannot be expressed as reusable modules. We present a shading language, Spark, and its implementation for modern graphics hardware that improves support for separation of concerns into modules. A Spark shader class can encapsulate code that maps to more than one pipeline stage, and can be extended and composed using object-oriented inheritance. In our tests, shaders written in Spark achieve performance within 2% of HLSL.
Theresa Foley, Pat Hanrahan
ACM Trans. Graph.1
2008 Parallel computing for graphics
abstract
This course provides an introduction to parallel-programming architectures and environments for interactive graphics and demonstrates how to combine traditional rendering API with advanced parallel computation.
Theresa Foley, Justin Hensley, Jason C. Yang
SIGGRAPH ASIA Courses1
2004 Brook for GPUs: stream computing on graphics hardware
abstract
In this paper, we present Brook for GPUs, a system for general-purpose computation on programmable graphics hardware. Brook extends C to include simple data-parallel constructs, enabling the use of the GPU as a streaming co-processor. We present a compiler and runtime system that abstracts and virtualizes many aspects of graphics hardware. In addition, we present an analysis of the effectiveness of the GPU as a compute engine compared to the CPU, to determine when the GPU can outperform the CPU for a particular algorithm. We evaluate our system with five applications, the SAXPY and SGEMV BLAS operators, image segmentation, FFT, and ray tracing. For these applications, we demonstrate that our Brook implementations perform comparably to hand-written GPU code and up to seven times faster than their CPU counterparts.
Ian Buck, Theresa Foley, Daniel Reiter Horn, Jeremy Sugerman, Kayvon Fatahalian, Mike Houston, Pat Hanrahan
ACM Trans. Graph.2