EDBT 2026 Demo / reviewers in the wild / expert
Justin Hensley
dblp:74/6675
· DBLP profile ↗
14ranked-venue papers
4as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-authorHuman-computer interaction and ubiquitous computing · 8 · 2 first-authorSystems, architecture and hardware · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
GPUs and heterogeneous computing · 52% Memory systems · 35% Parallel and multicore computing · 8% | |
| Computer graphics and multimedia
3 papers |
Rendering · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Rendering
texture mapping |
0.1 | 1 | 2011 | A space-efficient and hardware-friendly implementation of Ptex · SIGGRAPH Asia Sketches 2011 |
GPUs and heterogeneous computing › GPU programming
GPU programming models |
0.1 | 1 | 2010 | Physical and graphical effects in OpenCL by example · SIGGRAPH ASIA (Courses) 2010 |
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL |
0.1 | 1 | 2010 | Physical and graphical effects in OpenCL by example · SIGGRAPH ASIA (Courses) 2010 |
Rendering
parallel rendering |
0.1 | 1 | 2008 | Parallel computing for graphics · SIGGRAPH ASIA Courses 2008 |
Memory systems › processing-in-memory
intelligent memory |
0.1 | 2 | 2003 | Cache Coherence in Intelligent Memory Systems · IEEE Trans. Computers 2003 Exploiting ILP in Page-based Intelligent Memory · MICRO 1999 |
Memory systems
processing-in-memory |
0.1 | 2 | 2003 | Cache Coherence in Intelligent Memory Systems · IEEE Trans. Computers 2003 Exploiting ILP in Page-based Intelligent Memory · MICRO 1999 |
Memory systems
cache coherence |
0.0 | 1 | 2003 | Cache Coherence in Intelligent Memory Systems · IEEE Trans. Computers 2003 |
GPUs and heterogeneous computing
graphics hardware |
0.0 | 1 | 2011 | A space-efficient and hardware-friendly implementation of Ptex · SIGGRAPH Asia Sketches 2011 |
Rendering › graphics pipeline
programmable graphics pipeline |
0.0 | 1 | 2010 | Physical and graphical effects in OpenCL by example · SIGGRAPH ASIA (Courses) 2010 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 1999 | Exploiting ILP in Page-based Intelligent Memory · MICRO 1999 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 1 | 2003 | Cache Coherence in Intelligent Memory Systems · IEEE Trans. Computers 2003 |
Methods — techniques the papers use, named apart from their topics
parallel programming · 0.4protocol design · 0.0simulation · 0.0VLIW processor design · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | A framework for rendering complex scattering effects on hairabstractThe appearance of hair plays a critical role in synthesizing realistic looking human characters. However, due to the high complexity in hair geometry and the scattering nature of hair fibers, rendering hair with photorealistic quality and at interactive speeds remains as an open problem in computer graphics. Previous approaches attempt to simplify the scattering model to only tackle a specific aspect of the scattering effects. In this paper, we present a new approach to simultaneously render complex scattering effects including volumetric shadows, transparency, and antialiasing under a unified framework. Our solution uses a shadow-ray path to produce volumetric self-shadows and an additional view-ray path to produce transparency. To compute and accumulate the contribution of individual hair fibers along each (shadow or view) path, we develop a new GPU-based k-buffer technique that can efficiently locate the K nearest scattering locations and combine them in the correct order. Compared with existing multi-layer based approaches[Kim and Neumann 2001; Yuksel and Keyser 2008; Sintorn and Assarsson 2009], we show that our k-buffer solution can more accurately reproduce the shadowing and transparency effects. Further, we present an anti-aliasing scheme that directly builds upon the k-buffer. We implement all three effects (volumetric shadows, transparency, and anti-aliasing) under a unified rendering pipeline. Experiments on complex hair models demonstrate that our new solution produces near photorealistic hair rendering at very interactive speed. Jason C. Yang, Justin Hensley, Takahiro Harada, Jingyi Yu 0001 |
I3D | 3 |
| 2011 | A space-efficient and hardware-friendly implementation of PtexabstractWe introduce a method to pack Ptex per-face texture data that is both space-efficient and hardware-friendly. Recently presented real-time implementations of Ptex have been wasteful with space and required a storage cost many times higher than the size of the original texture data. Our method packs multiple levels of Ptex data together, and requires only around 8% increase in storage for our test textures. Additionally, because of efficient data packing, our method wastes less space than a typical texture atlas, which requires buffer regions to be added between the separate charts within the texture. Sujeong Kim, Karl E. Hillesland, Justin Hensley |
SIGGRAPH Asia Sketches | 3 |
| 2011 | HK-2207abstractNo abstract available. Abe Wiley, Jay McKee, Jason C. Yang, Dan Roeger, Takahiro Harada, Justin Hensley, Saif Ali |
SIGGRAPH Asia Computer Animation Festival | 6 |
| 2010 | What is OpenCL™?
Justin Hensley |
SIGGRAPH ASIA (Courses) | 1 |
| 2010 | Physical and graphical effects in OpenCL by exampleabstractThere are strong indications that the future of interactive graphics involves a more flexible programming model than today's OpenGL/Direct3D pipelines. That means that graphics developers will need a basic understanding of how to combine emerging parallel-programming techniques with the traditional interactive rendering pipeline. Justin Hensley, Derek K. Gerstmann, Jason C. Yang |
SIGGRAPH ASIA (Courses) | 1 |
| 2010 | Real-Time Concurrent Linked List Construction on the GPUabstractAbstract We introduce a method to dynamically construct highly concurrent linked lists on modern graphics processors. Once constructed, these data structures can be used to implement a host of algorithms useful in creating complex rendering effects in real time. We present a straightforward way to create these linked lists using generic atomic operations available in APIs such as OpenGL 4.0 and DirectX 11. We also describe several possible applications of our algorithm. One example uses per‐pixel linked lists for order‐independent transparency; as a consequence, we are able to directly implement fully programmable blending, which frees developers from the restrictions imposed by current graphics APIs. The second uses linked lists to implement real‐time indirect shadows. Jason C. Yang, Justin Hensley, Holger Grün, Nicolas Thibieroz |
Comput. Graph. Forum | 2 |
| 2008 | Parallel computing for graphicsabstractThis course provides an introduction to parallel-programming architectures and environments for interactive graphics and demonstrates how to combine traditional rendering API with advanced parallel computation. Theresa Foley, Justin Hensley, Jason C. Yang |
SIGGRAPH ASIA Courses | 2 |
| 2007 | Efficient histogram generation using scattering on GPUsabstractWe present an efficient algorithm to compute image histograms entirely on the GPU. Unlike previous implementations that use a gather approach, we take advantage of scattering data through the vertex shader and of high-precision blending available on modern GPUs. This results in fewer operations executed per pixel and speeds up the computation. Thorsten Scheuermann, Justin Hensley |
SI3D | 2 |
| 2005 | Fast Summed-Area Table Generation and its ApplicationsabstractWe introduce a technique to rapidly generate summed-area tables using graphics hardware. Summed area tables, originally introduced by Crow, provide a way to filter arbitrarily large rectangular regions of an image in a constant amount of time. Our algorithm for generating summed-area tables, similar to a technique used in scientific computing called recursive doubling, allows the generation of a summed-area table in O(log n) time. We also describe a technique to mitigate the precision requirements of summed-area tables. The ability to calculate and use summed-area tables at interactive rates enables numerous interesting rendering effects. We present several possible applications. First, the use of summed-area tables allows real-time rendering of interactive, glossy environmental reflections. Second, we present glossy planar reflections with varying blurriness dependent on a reflected object’s distance to the reflector. Third, we show a technique that uses a summed-area table to render glossy transparent objects. The final application demonstrates an interactive depth-of-field effect using summedarea tables. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three-Dimensional Justin Hensley, Thorsten Scheuermann, Greg Coombe, Montek Singh, Anselmo Lastra |
Comput. Graph. Forum | 1 |
| 2004 | An Area- and Energy-Efficient Asynchronous Booth Multiplier for Mobile DevicesabstractThe recent explosion in the number of handheld multimedia devices has created a need for energy-efficient computation due to limited battery lifetimes. We focus on multiplication, which is needed in several application domains, e.g., 3D graphics, signal processing, and cryptography. We introduce an asynchronous implementation of a plain Booth multiplier (i.e., radix-2), which is both area- and energy-efficient, and therefore suitable for mobile applications. This paper makes the following contributions. First, a novel counterflow organization is introduced, in which the data bits flow in one direction, and the Booth commands piggyback on the acknowledgments flowing in the opposite direction. Second, the arithmetic and shifter units are merged together to obtain significant improvement in area, energy as well as speed. Third, our design performs overlapped execution of multiple iterations of the Booth algorithm. Finally, the design is quite modular, which allows scaling to arbitrary operand widths, without gate resizing or cycle time overheads. Spice simulations in a 0.18 /spl mu/m TSMC process at 1.8 V, indicate promising performance: the multiplier takes 1.08 ns per Booth iteration, regardless of the operand widths, thereby demonstrating the scalability of our approach. In addition, the multiplier is fully functional at reduced supply voltages (e.g., 1.0 V), and thus capable of dynamically trading off performance for energy efficiency. Justin Hensley, Anselmo Lastra, Montek Singh |
ICCD | 1 |
| 2003 | Cache Coherence in Intelligent Memory SystemsabstractThe Active Pages model of intelligent memory can speed up data-intensive applications by up to two to three orders of magnitude over conventional systems. A fundamental problem with intelligent memory, however, arises when data cached by the processor is modified by logic in the memory. The Active Page model inherently limits sharing, keeping coherence tractable, but exacerbates saturation problems. We first present a hybrid snoopy/directory protocol for use in Active Pages. Limited sharing allows for a low-latency, low-bandwidth hybrid protocol. A transparent remapping mechanism is added for efficient caching. On smaller data sizes, explicit flushing and hardware coherence exhibit similar performance, but hardware coherence is easier to program and uses less bandwidth. Finally, we examine SMP multiprocessor systems to mitigate saturation effects. As the number of threads increases, the bandwidth needs increase, making hardware coherence even more attractive. Diana Franklin, Mark Oskin, Justin Hensley, Fred Chong |
IEEE Trans. Computers | 3 |
| 2001 | PixelFlex: A Reconfigurable Multi-Projector Display SystemabstractThis paper presents PixelFlex - a spatially reconfigurable multi-projector display system. The PixelFlex system is composed of ceiling-mounted projectors, each with computer-controlled pan, tilt, zoom and focus; and a camera for closed-loop calibration. Working collectively, these controllable projectors function as a single logical display capable of being easily modified into a variety of spatial formats of differing pixel density, size and shape. New layouts are automatically calibrated within minutes to generate the accurate warping and blending functions needed to produce seamless imagery across planar display surfaces, thus giving the user the flexibility to quickly create, save and restore multiple screen configurations. Overall, PixelFlex provides a new level of automatic reconfigurability and usage, departing from the static, one-size-fits-all design of traditional large-format displays. As a front-projection system, PixelFlex can be installed in most environments with space constraints and requires little or no post-installation mechanical maintenance because of the closed-loop calibration. Ruigang Yang, David Gotz, Justin Hensley, Herman Towles, Michael S. Brown |
IEEE Visualization | 3 |
| 2000 | Reducing Cost and Tolerating Defects in Page-based Intelligent MemoryabstractActive Pages is a page-based model of intelligent memory specifically designed to support virtualized hardware resources. Previous work has shown substantial performance benefits from off loading data-intensive tasks to a memory system that implements Active Pages. With a simple VLIW processor embedded near each page on DRAM, Active Page memory systems achieve up to 1000X speedups over conventional memory systems. In this study, we examine Active Page memories that share, or multiplex, embedded VLIW processors across multiple physical Active Pages. We explore the trade-off between individual page-processor performance and page-level multiplexing. We find that hardware costs of computational logic can be reduced from 31% of DRAM chip area to 12%, through multiplexing, without significant loss in performance. Furthermore, manufacturing defects that disable up to 50% of the page processors can be tolerated through efficient resource allocation and associative multiplexing. Mark Oskin, Diana Franklin, Justin Hensley, Lucian Vlad Lita, Fred Chong |
ICCD | 3 |
| 1999 | Exploiting ILP in Page-based Intelligent MemoryabstractThis study compares the speed, area, and power of different implementations of Active Pages, an intelligent memory system which helps bridge the growing gap between processor and memory performance by associating simple functions with each page of data. Previous investigations have shown up to 1000X speedups using a block of reconfigurable logic to implement these functions next to each subarray on a DRAM chip. In this study, we show that instruction-level parallelism, not hardware specialization, is the key to the previous success with reconfigurable logic. In order to demonstrate this fact, an Active Page implementation based upon a simplified VLIW processor was developed. Unlike conventional VLIW processors, power and area constraints lead to a design which has a small number of pipeline stages. Our results demonstrate that a four-wide VLIW processor attains comparable performance to that of pure FPGA logic but requires significantly less area and power. Mark Oskin, Justin Hensley, Diana Franklin, Fred Chong, Matthew K. Farrens, Aneet Chopra |
MICRO | 2 |