Qiming Hou

dblp:17/6598 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
11 papers
Rendering · 72% Geometric modeling and processing · 13% Visual content generation and editing · 12%
Artificial intelligence
3 papers
Generative modeling · 33% Face, body and person analysis · 22% 3D vision · 19%
Computer architecture, parallel and distributed computing, and storage systems
6 papers
GPUs and heterogeneous computing · 100%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 64% Debugging and program repair · 36%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation · ICCV 2025
Rendering
ray tracing
0.652014
Cone Tracing for Furry Object Rendering · IEEE Trans. Vis. Comput. Graph. 2014
Memory-Scalable GPU Spatial Hierarchy Construction · IEEE Trans. Vis. Comput. Graph. 2011
A shading reuse method for efficient micropolygon ray tracing · ACM Trans. Graph. 2011
Rendering
image-based rendering
0.512021
Neural compositing for real-time augmented reality rendering in low-frequency lighting environments · Sci. China Inf. Sci. 2021
Visual content generation and editing › image editing
image compositing
0.512021
Neural compositing for real-time augmented reality rendering in low-frequency lighting environments · Sci. China Inf. Sci. 2021
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
3d convolutional neural network
0.412020
H-CNN: Spatial Hashing Based CNN for 3D Shape Analysis · IEEE Trans. Vis. Comput. Graph. 2020
Computer vision › 3D vision
3d shape analysis
0.412020
H-CNN: Spatial Hashing Based CNN for 3D Shape Analysis · IEEE Trans. Vis. Comput. Graph. 2020
Computer vision › Vision and language
vision-language model
0.312025
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation · ICCV 2025
Rendering
antialiasing
0.212016
Efficient GPU path rendering using scanline rasterization · ACM Trans. Graph. 2016
Rendering › geometric rendering › vector graphics rendering
path rendering
0.212016
Efficient GPU path rendering using scanline rasterization · ACM Trans. Graph. 2016
Geometric modeling and processing
winding number
0.212016
Efficient GPU path rendering using scanline rasterization · ACM Trans. Graph. 2016
Rendering
light transport
0.212015
Unbiased photon gathering for light transport simulation · ACM Trans. Graph. 2015
Rendering › global illumination
photon mapping
0.212015
Unbiased photon gathering for light transport simulation · ACM Trans. Graph. 2015
Computer vision › Face, body and person analysis
face tracking
0.212014
Displaced dynamic expression regression for real-time facial tracking and animation · ACM Trans. Graph. 2014
Computer vision › Face, body and person analysis
facial animation
0.212014
Displaced dynamic expression regression for real-time facial tracking and animation · ACM Trans. Graph. 2014
Computer vision › Face, body and person analysis › face tracking
real-time face tracking
0.212014
Displaced dynamic expression regression for real-time facial tracking and animation · ACM Trans. Graph. 2014
GPUs and heterogeneous computing
GPU programming
0.222009
Debugging GPU stream programs through automatic dataflow recording and visualization · ACM Trans. Graph. 2009
BSGP: bulk-synchronous GPU programming · ACM Trans. Graph. 2008
Virtual and augmented reality › augmented reality
augmented reality rendering
0.112021
Neural compositing for real-time augmented reality rendering in low-frequency lighting environments · Sci. China Inf. Sci. 2021
GPUs and heterogeneous computing › GPU computing
GPU algorithms
0.122011
Memory-Scalable GPU Spatial Hierarchy Construction · IEEE Trans. Vis. Comput. Graph. 2011
Real-time KD-tree construction on graphics hardware · ACM Trans. Graph. 2008
Geometric modeling and processing
shape representation
0.112020
H-CNN: Spatial Hashing Based CNN for 3D Shape Analysis · IEEE Trans. Vis. Comput. Graph. 2020
Rendering › light transport
precomputed light transport
0.112011
Radiance Transfer Biclustering for Real-Time All-Frequency Biscale Rendering · IEEE Trans. Vis. Comput. Graph. 2011
Rendering › global illumination
precomputed radiance transfer
0.112011
Radiance Transfer Biclustering for Real-Time All-Frequency Biscale Rendering · IEEE Trans. Vis. Comput. Graph. 2011
Rendering
real-time rendering
0.112011
Radiance Transfer Biclustering for Real-Time All-Frequency Biscale Rendering · IEEE Trans. Vis. Comput. Graph. 2011
Rendering
shading
0.112011
A shading reuse method for efficient micropolygon ray tracing · ACM Trans. Graph. 2011
Rendering › shading
shading reuse
0.112011
A shading reuse method for efficient micropolygon ray tracing · ACM Trans. Graph. 2011
Rendering
GPU rendering
0.112009
RenderAnts: interactive Reyes rendering on GPUs · ACM Trans. Graph. 2009
Rendering › parallel rendering
multi-GPU rendering
0.112009
RenderAnts: interactive Reyes rendering on GPUs · ACM Trans. Graph. 2009
Debugging and program repair
fault localization
0.112009
Debugging GPU stream programs through automatic dataflow recording and visualization · ACM Trans. Graph. 2009
Geometric modeling and processing › spatial data structures
bounding volume hierarchy
0.122014
Cone Tracing for Furry Object Rendering · IEEE Trans. Vis. Comput. Graph. 2014
Memory-Scalable GPU Spatial Hierarchy Construction · IEEE Trans. Vis. Comput. Graph. 2011
Geometric modeling and processing
spatial data structures
0.112008
Real-time KD-tree construction on graphics hardware · ACM Trans. Graph. 2008
Compilers and program optimization › code generation › parallel code generation
GPU kernel generation
0.112008
BSGP: bulk-synchronous GPU programming · ACM Trans. Graph. 2008

Methods — techniques the papers use, named apart from their topics

spatial hashing · 1.3perfect spatial hashing · 1.3hash2col · 1.3col2hash · 1.3vision-language model · 0.9open-vocabulary detection · 0.9large language model · 0.9neural rendering · 0.5bounding volume hierarchy · 0.3ray shooting · 0.2parallelization over boundary fragments · 0.2multiple importance sampling · 0.2monte carlo estimation · 0.2compiler instrumentation · 0.2GPU interrupt · 0.2landmark regression · 0.2displaced dynamic expression regression · 0.2partial breadth-first search · 0.1
YearPublicationVenuePosition
2025 ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
abstract
We present ROVI, a high-quality synthetic dataset for instance-grounded text-to-image generation, created by labeling 1M curated web images. Our key innovation is a strategy called re-captioning, focusing on the pre-detection stage, where a VLM (Vision-Language Model) generates comprehensive visual descriptions that are then processed by an LLM (Large Language Model) to extract a flat list of potential categories for OVDs (Open-Vocabulary Detectors) to detect. This approach yields a global prompt inherently linked to instance annotations while capturing secondary visual elements humans typically overlook. Evaluations show that ROVI exceeds existing detection datasets in image quality and resolution while containing two orders of magnitude more categories with an open-vocabulary nature. For demonstrative purposes, a text-to-image model GLIGEN trained on ROVI significantly outperforms state-of-the-art alternatives in instance grounding accuracy, prompt fidelity, and aesthetic quality. Our dataset and reproducible pipeline are available at https://github.com/CihangPeng/ROVI.
Cihang Peng, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001
ICCV2
2025 Efficient Self-Adaptive Pseudo-Resistor with Rapid Settling and High Linearity for Neurorecording Front-End Circuits
abstract
In this paper, we present a novel self-adaptive pseudo-resistor (A-PR) designed to enhance the performance of neurorecording front-end circuits in terms of settling time, linearity, and tunability. We validate the effectiveness of the proposed A-PR through the implementation of a capacitively- coupled instrumentation amplifier (CCIA) recording front-end using TSMC 40-nm process technology. The results demonstrate that the A-PR enables continuous recording with minimal interruptions, enhancing the system’s robustness and enabling more reliable acquisition of neural signals. Notably, the A-PR achieves a significant reduction in settling time, reaching the millisecond level—1000 times faster than conventional pseudo-resistors—while also exhibiting wide linear characteristics and easy tunability.
Hui Wu 0010, Xing Liu 0014, Jinbo Chen 0002, Wenjun Zou, Qiming Hou, Yutao Mao, Xiaofei Kuang, Jie Yang 0033, Mohamad Sawan
ISCAS6
2024 A Low-Power Level-Crossing Analog-to-Spike Converter Intended for Neuromorphic Biomedical Applications
abstract
The increasing interests in building bio-signal recording and processing systems for personal healthcare applications have been hindered by the critical sampling energy consumption issues of conventional biomedical systems. To address these limits, we propose a comprehensive strategy centered around a low-power level-crossing analog-to-spike converter (LC-ASC). This strategy enables event-driven compressive sampling by leveraging signal sparsity, achieving lower average sampling rates than Nyquist sampling. Our strategy includes universal VerilogA LC-ASC models, evaluation tools, and a reconfigurable data interface for versatile digital processing. Specifically, we introduce an online open-source VerilogA LC-ASC model and compression performance calculation tools for evaluating its performance with different bio-signals. The implemented LC-ASC chip demonstrates very-low power consumption of 31.5125.3 nW validated through chip measurements. Additionally, the proposed reconfigurable data interface ensures seamless integration with synchronous and asynchronous digital processing modules without sacrificing system-level performance. These advancements pave the way for energy-efficient neuromorphic biomedical circuits and systems.
Jinbo Chen 0002, Hui Wu 0010, Fengshi Tian, Qiming Hou, Jie Yang 0033, Mohamad Sawan
ISCAS4
2023 InstantTrace: fast parallel neuron tracing on GPUs
Yuxuan Hou, Zhong Ren 0001, Qiming Hou, Yubo Tao, Yankai Jiang 0001, Wei Chen 0001
Vis. Comput.3
2021 Neural compositing for real-time augmented reality rendering in low-frequency lighting environments
Shengjie Ma, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001
Sci. China Inf. Sci.3
2020 H-CNN: Spatial Hashing Based CNN for 3D Shape Analysis
abstract
We present a novel spatial hashing based data structure to facilitate 3D shape analysis using convolutional neural networks (CNNs). Our method builds hierarchical hash tables for an input model under different resolutions that leverage the sparse occupancy of 3D shape boundary. Based on this data structure, we design two efficient GPU algorithms namely hash2col and col2hash so that the CNN operations like convolution and pooling can be efficiently parallelized. The perfect spatial hashing is employed as our spatial hashing scheme, which is not only free of hash collision but also nearly minimal so that our data structure is almost of the same size as the raw input. Compared with existing 3D CNN methods, our data structure significantly reduces the memory footprint during the CNN training. As the input geometry features are more compactly packed, CNN operations also run faster with our data structure. The experiment shows that, under the same network structure, our method yields comparable or better benchmark results compared with the state-of-the-art while it has only one-third memory consumption when under high resolutions (i.e., 2563).
Tianjia Shao, Yin Yang 0002, Yanlin Weng, Qiming Hou, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.4
2016 Variance Analysis and Adaptive Sampling for Indirect Light Path Reuse
Xin Sun 0014, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001
J. Comput. Sci. Technol.4
2016 Efficient GPU path rendering using scanline rasterization
abstract
We introduce a novel GPU path rendering method based on scan-line rasterization, which is highly work-efficient but traditionally considered as GPU hostile. Our method is parallelized over boundary fragments , i.e., pixels directly intersecting the path boundary. Non-boundary pixels are processed in bulk as horizontal spans like in CPU scanline rasterizers, which saves a significant amount of winding number computation workload. The distinction also allows the majority of our algorithmic steps to focus on boundary fragments only, which leads to highly balanced workload among the GPU threads. In addition, we develop a ray shooting pattern that minimizes the global data dependency when computing winding numbers at anti-aliasing samples. This allows us to shift the majority of winding-number-related workload to the same kernel that consumes its result, which saves a significant amount of GPU memory bandwidth. Experiments show that our method gives a consistent 2.5X speedup over state-of-the-art alternatives for high-quality rendering at Ultra HD resolution, which can increase to more than 30X in extreme cases. We can also get a consistent 10X speedup on animated input.
Qiming Hou, Kun Zhou 0001
ACM Trans. Graph.2
2015 Unbiased photon gathering for light transport simulation
abstract
Photon mapping (PM) has been widely regarded as an efficient solution for light transport simulation, including challenging caustics paths and many-bounce indirect lighting. The efficiency of PM comes from reusing traced photons. However, the handling of photon gathering in existing PM algorithms is universally biased -- the expected value of their results does not necessarily agree with the true solution of the rendering equation. We present a novel photon gathering method to efficiently achieve unbiased rendering with photon mapping. Instead of aggregating the gathered photons into an estimated density as in classical photon mapping, we process each photon individually and connect the corresponding light sub-path with the eye sub-path that generates the gather point, creating an unbiased path sample. The Monte Carlo estimate for such a path sample is calculated by evaluating all relevant terms in a strict and unbiased way, leading to a self-contained unbiased sampling technique. We further develop a set of multiple importance sampling (MIS) weights that allow our method to be optimally combined with bidirectional path tracing (BDPT), resulting in an unbiased rendering algorithm that can efficiently handle a wide variety of light paths and that compares favorably with previous algorithms. Experiments demonstrate the efficacy and robustness of our method.
Xin Sun 0014, Qiming Hou, Baining Guo, Kun Zhou 0001
ACM Trans. Graph.3
2014 Real-time facial animation on mobile devices
Yanlin Weng, Qiming Hou, Kun Zhou 0001
Graph. Model.3
2014 Displaced dynamic expression regression for real-time facial tracking and animation
abstract
We present a fully automatic approach to real-time facial tracking and animation with a single video camera. Our approach does not need any calibration for each individual user. It learns a generic regressor from public image datasets, which can be applied to any user and arbitrary video cameras to infer accurate 2D facial landmarks as well as the 3D facial shape from 2D video frames. The inferred 2D landmarks are then used to adapt the camera matrix and the user identity to better match the facial expressions of the current user. The regression and adaptation are performed in an alternating manner. With more and more facial expressions observed in the video, the whole process converges quickly with accurate facial tracking and animation. In experiments, our approach demonstrates a level of robustness and accuracy on par with state-of-the-art techniques that require a time-consuming calibration step for each individual user, while running at 28 fps on average. We consider our approach to be an attractive solution for wide deployment in consumer-level applications.
Qiming Hou, Kun Zhou 0001
ACM Trans. Graph.2
2014 Cone Tracing for Furry Object Rendering
abstract
We present a cone-based ray tracing algorithm for high-quality rendering of furry objects with reflection, refraction and defocus effects. By aggregating many sampling rays in a pixel as a single cone, we significantly reduce the high supersampling rate required by the thin geometry of fur fibers. To reduce the cost of intersecting fur fibers with cones, we construct a bounding volume hierarchy for the fiber geometry to find the fibers potentially intersecting with cones, and use a set of connected ribbons to approximate the projections of these fibers on the image plane. The computational cost of compositing and filtering transparent samples within each cone is effectively reduced by approximating away in-cone variations of shading, opacity and occlusion. The result is a highly efficient ray tracing algorithm for furry objects which is able to render images of quality comparable to those generated by alternative methods, while significantly reducing the rendering time. We demonstrate the rendering quality and performance of our algorithm using several examples and a user study.
Menglei Chai, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.3
2012 Scalable Programmable Motion Effects on GPUs
abstract
Abstract We present an efficient and scalable system that enables programmable motion effects on GPUs. Our system is based on the framework proposed by Schmid et al. [ SSBG10 ] that extends the concept of a surface shader to that of a programmable motion effect. While capable of expressing a variety of motion depiction styles, the execution of motion effect programs requires global knowledge about all portions of an object's surface that passes in front of a pixel during an arbitrarily long period of time, resulting in extremely high memory usage and significantly restricting the degree of parallelism of typical GPU rendering algorithms that parallelize computations over pixels in each frame of animations. To address this problem, we design our system to process multiple frames of a pixel in parallel. This new parallelization approach enables better utilization of GPU memory and also makes it possible to design an efficient out‐of‐core algorithm required in rendering real‐world animations. We also develop an analytical visibility algorithm to resolve depth conflicts of objects, reducing the required temporal resampling rate and further exposing parallelism. Experiments show that we are able to handle very large scenes and improve runtime performance up to an order of magnitude.
Xuezhen Huang, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001
Comput. Graph. Forum2
2011 A shading reuse method for efficient micropolygon ray tracing
abstract
We present a shading reuse method for micropolygon ray tracing. Unlike previous shading reuse methods that require an explicit object-to-image space mapping for shading density estimation or shading accuracy, our method performs shading density control and actual shading reuse in different spaces with uncorrelated criterions. Specifically, we generate the shading points by shooting a user-controlled number of shading rays from the image space, while the evaluated shading values are assigned to antialiasing samples through object-space nearest neighbor searches. Shading samples are generated in separate layers corresponding to first bounce ray paths to reduce spurious reuse from very different ray paths. This method eliminates the necessity of an explicit object-to-image space mapping, enabling the elegant handling of ray tracing effects such as reflection and refraction. The overhead of our shading reuse operations is minimized by a highly parallel implementation on the GPU. Compared to the state-of-the-art micropolygon ray tracing algorithm, our method is able to reduce the required shading evaluations by an order of magnitude and achieve significant performance gains.
Qiming Hou, Kun Zhou 0001
ACM Trans. Graph.1
2011 Memory-Scalable GPU Spatial Hierarchy Construction
abstract
Recent GPU algorithms for constructing spatial hierarchies have achieved promising performance for moderately complex models by using the breadth-first search (BFS) construction order. While being able to exploit the massive parallelism on the GPU, the BFS order also consumes excessive GPU memory, which becomes a serious issue for interactive applications involving very complex models with more than a few million triangles. In this paper, we propose to use the partial breadth-first search (PBFS) construction order to control memory consumption while maximizing performance. We apply the PBFS order to two hierarchy construction algorithms. The first algorithm is for kd-trees that automatically balances between the level of parallelism and intermediate memory usage. With PBFS, peak memory consumption during construction can be efficiently controlled without costly CPU-GPU data transfer. We also develop memory allocation strategies to effectively limit memory fragmentation. The resulting algorithm scales well with GPU memory and constructs kd-trees of models with millions of triangles at interactive rates on GPUs with 1 GB memory. Compared with existing algorithms, our algorithm is an order of magnitude more scalable for a given GPU memory bound. The second algorithm is for out-of-core bounding volume hierarchy (BVH) construction for very large scenes based on the PBFS construction order. At each iteration, all constructed nodes are dumped to the CPU memory, and the GPU memory is freed for the next iteration's use. In this way, the algorithm is able to build trees that are too large to be stored in the GPU memory. Experiments show that our algorithm can construct BVHs for scenes with up to 20 M triangles, several times larger than previous GPU algorithms.
Qiming Hou, Xin Sun 0014, Kun Zhou 0001, Christian Lauterbach, Dinesh Manocha
IEEE Trans. Vis. Comput. Graph.1
2011 Radiance Transfer Biclustering for Real-Time All-Frequency Biscale Rendering
abstract
We present a real-time algorithm to render all-frequency radiance transfer at both macroscale and mesoscale. At a mesoscale, the shading is computed on a per-pixel basis by integrating the product of the local incident radiance and a bidirectional texture function. While at a macroscale, the precomputed transfer matrix, which transfers the global incident radiance to the local incident radiance at each vertex, is losslessly compressed by a novel biclustering technique. The biclustering is directly applied on the radiance transfer represented in a pixel basis, on which the BTF is naturally defined. It exploits the coherence in the transfer matrix and a property of matrix element values to reduce both storage and runtime computation cost. Our new algorithm renders at real-time frame rates realistic materials and shadows under all-frequency direct environment lighting. Comparisons show that our algorithm is able to generate images that compare favorably with reference ray tracing results, and has obvious advantages over alternative methods in storage and preprocessing time.
Xin Sun 0014, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001, Baining Guo
IEEE Trans. Vis. Comput. Graph.2
2010 Micropolygon ray tracing with defocus and motion blur
abstract
We present a micropolygon ray tracing algorithm that is capable of efficiently rendering high quality defocus and motion blur effects. A key component of our algorithm is a BVH (bounding volume hierarchy) based on 4D hyper-trapezoids that project into 3D OBBs (oriented bounding boxes) in spatial dimensions. This acceleration structure is able to provide tight bounding volumes for scene geometries, and is thus efficient in pruning intersection tests during ray traversal. More importantly, it can exploit the natural coherence on the time dimension in motion blurred scenes. The structure can be quickly constructed by utilizing the micropolygon grids generated during micropolygon tessellation. Ray tracing of defocused and motion blurred scenes is efficiently performed by traversing the structure. Both the BVH construction and ray traversal are easily implemented on GPUs and integrated into a GPU-based micropolygon renderer. In our experiments, our ray tracer performs up to an order of magnitude faster than the state-of-art rasterizers while consistently delivering an image quality equivalent to a maximum-quality rasterizer. We also demonstrate that the ray tracing algorithm can be extended to handle a variety of effects, such as secondary ray effects and transparency.
Qiming Hou, Baining Guo, Kun Zhou 0001
ACM Trans. Graph.1
2009 Debugging GPU stream programs through automatic dataflow recording and visualization
abstract
We present a novel framework for debugging GPU stream programs through automatic dataflow recording and visualization. Our debugging system can help programmers locate errors that are common in general purpose stream programs but very difficult to debug with existing tools. A stream program is first compiled into an instrumented program using a compiler. This instrumenting compiler automatically adds to the original program dataflow recording code that saves the information of all GPU memory operations into log files. The resulting stream program is then executed on the GPU. With dataflow recording, our debugger automatically detects common memory errors such as out-of-bound access, uninitialized data access, and race conditions -- these errors are extremely difficult to debug with existing tools. When the instrumented program terminates, either normally or due to an error, a dataflow visualizer is launched and it allows the user to examine the memory operation history of all threads and values in all streams. Thus the user can analyze error sources by tracing through relevant threads and streams using the recorded dataflow. A key ingredient of our debugging framework is the GPU interrupt , a novel mechanism that we introduce to support CPU function calls from inside GPU code. We enable interrupts on the GPU by designing a specialized compilation algorithm that translates these interrupts into GPU kernels and CPU management code. Dataflow recording involving disk I/O operations can thus be implemented as interrupt handlers. The GPU interrupt mechanism also allows the programmer to discover errors in more active ways by developing customized debugging functions that can be directly used in GPU code. As examples we show two such functions: assert for data verification and watch for visualizing intermediate results.
Qiming Hou, Kun Zhou 0001, Baining Guo
ACM Trans. Graph.1
2009 RenderAnts: interactive Reyes rendering on GPUs
abstract
We present RenderAnts, the first system that enables interactive Reyes rendering on GPUs. Taking RenderMan scenes and shaders as input, our system first compiles RenderMan shaders to GPU shaders. Then all stages of the basic Reyes pipeline, including bounding/splitting, dicing, shading, sampling, compositing and filtering, are executed on GPUs using carefully designed data-parallel algorithms. Advanced effects such as shadows, motion blur and depth-of-field can also be rendered. In order to avoid exhausting GPU memory, we introduce a novel dynamic scheduling algorithm to bound the memory consumption during rendering. The algorithm automatically adjusts the amount of data being processed in parallel at each stage so that all data can be maintained in the available GPU memory. This allows our system to maximize the parallelism in all individual stages of the pipeline and achieve superior performance. We also propose a multi-GPU scheduling technique based on work stealing so that the system can support scalable rendering on multiple GPUs. The scheduler is designed to minimize inter-GPU communication and balance workloads among GPUs. We demonstrate the potential of RenderAnts using several complex RenderMan scenes and an open source movie entitled Elephants Dream. Compared to Pixar's PRMan, our system can generate images of comparably high quality, but is over one order of magnitude faster. For moderately complex scenes, the system allows the user to change the viewpoint, lights and materials while producing photorealistic results at interactive speed.
Kun Zhou 0001, Qiming Hou, Zhong Ren 0001, Minmin Gong, Xin Sun 0014, Baining Guo
ACM Trans. Graph.2
2008 BSGP: bulk-synchronous GPU programming
abstract
We present BSGP, a new programming language for general purpose computation on the GPU. A BSGP program looks much the same as a sequential C program. Programmers only need to supply a bare minimum of extra information to describe parallel processing on GPUs. As a result, BSGP programs are easy to read, write, and maintain. Moreover, the ease of programming does not come at the cost of performance. A well-designed BSGP compiler converts BSGP programs to kernels and combines them using optimally allocated temporary streams. In our benchmark, BSGP programs achieve similar or better performance than well-optimized CUDA programs, while the source code complexity and programming time are significantly reduced. To test BSGP's code efficiency and ease of programming, we implemented a variety of GPU applications, including a highly sophisticated X3D parser that would be extremely difficult to develop with existing GPU programming languages.
Qiming Hou, Kun Zhou 0001, Baining Guo
ACM Trans. Graph.1
2008 Real-time KD-tree construction on graphics hardware
abstract
We present an algorithm for constructing kd-trees on GPUs. This algorithm achieves real-time performance by exploiting the GPU's streaming architecture at all stages of kd-tree construction. Unlike previous parallel kd-tree algorithms, our method builds tree nodes completely in BFS (breadth-first search) order. We also develop a special strategy for large nodes at upper tree levels so as to further exploit the fine-grained parallelism of GPUs. For these nodes, we parallelize the computation over all geometric primitives instead of nodes at each level. Finally, in order to maintain kd-tree quality, we introduce novel schemes for fast evaluation of node split costs. As far as we know, ours is the first real-time kd-tree algorithm on the GPU. The kd-trees built by our algorithm are of comparable quality as those constructed by off-line CPU algorithms. In terms of speed, our algorithm is significantly faster than well-optimized single-core CPU algorithms and competitive with multi-core CPU algorithms. Our algorithm provides a general way for handling dynamic scenes on the GPU. We demonstrate the potential of our algorithm in applications involving dynamic scenes, including GPU ray tracing, interactive photon mapping, and point cloud modeling.
Kun Zhou 0001, Qiming Hou, Rui Wang 0004, Baining Guo
ACM Trans. Graph.2
2007 Fogshop: Real-Time Design and Rendering of Inhomogeneous, Single-Scattering Media
abstract
We describe a new, analytic approximation to the airlight integral from scattering media whose density is modeled as a sum of Gaussians. The approximation supports real-time rendering of inhomogeneous media including their shadowing and scattering effects. For each Gaussian, this approximation samples the scattering integrand at the projection of its center along the view ray but models attenuation and shadowing with respect to the other Gaussians by integrating density along the fixed path from light source to 3D center to view point. Our method handles isotropic, single-scattering media illuminated by point light sources or low-frequency lighting environments. We also generalize models for reflectance of surfaces from constant-density to inhomogeneous media, using simple optical depth averaging in the direction of the light source or all around the receiver point. Our real-time renderer is incorporated into a system for real-time design and preview of realistic animated fog, steam, or smoke.
Kun Zhou 0001, Qiming Hou, Minmin Gong, John Snyder, Baining Guo, Harry Shum
PG2