Yazhen Yuan

dblp:138/3442 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-8857-8373ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 StereoFG: Generating Stereo Frames from Centered Feature Stream
abstract
In recent years, the community has seen the emergence of neural-based super-resolution and frame generation techniques. These methods have effectively sped up high-resolution rendering by exploiting the spatial and temporal coherence between sequential frames, but none of them are designed specifically for improving the rendering performance in VR applications, where stereo rendering doubles the rendering cost.
Chenyu Zuo, Yazhen Yuan, Zhizhen Wu, Jingzhen Lan, Ming Fu, Yuchi Huo, Rui Wang 0004
SIGGRAPH Asia2
2025 Streaming-Aware Neural Monte Carlo Rendering Framework with Unified Denoising-Compression and Client Collaboration
abstract
Recent advances in cloud rendering have brought us a promising alternative for interactive photorealistic rendering on lightweight devices, which used to be only available on high-end platforms equipped with powerful graphic cards. This technique enables users to perform rendering-related creative tasks, such as 3D product visualization and lighting design, from the comfort of any location using handheld devices, rather than being confined to the front of a noisy heat-generating workstation. However, existing large-scale cloud rendering systems that stream path-traced frames from the server to the client present extremely high rendering costs and transmission bandwidth requirements, even with advanced path-tracing acceleration and video compression techniques. To alleviate these problems, we propose a novel streaming-aware rendering framework that is able to learn a joint optimal model integrating two path-tracing acceleration techniques (adaptive sampling and denoising) and video compression technique. Our joint model can fully exploit the inherent connections between these techniques and thus achieve substantially reduced rendering costs and enhanced compression quality. We also introduce the collaboration of client rendering ability to assist the frame decoding by rendering G-buffers as the shared side information. We demonstrate that appropriately incorporating the geometry and material priors from G-buffers into a neural compression pipeline can significantly reduce the streaming bandwidth in a cloud rendering system, and lighten the compression module design for computation efficiency. Our experiments show that our method delivers the best quality at various bitrates compared to existing Monte Carlo rendering streaming schemes, while remaining lightweight and efficient for cross-platform thin clients, including mobiles and tablets.
Hangming Fan, Yuchi Huo, Chuankun Zheng, Chonghao Hu, Yazhen Yuan, Rui Wang 0004
ACM Trans. Graph.5
2025 Consecutive Frame Extrapolation with Predictive Sparse Shading
abstract
The demand for high-frame-rate rendering keeps increasing in modern displays. Existing frame generation and super-resolution techniques accelerate rendering by reducing rendering samples across space or time. However, they rely on a uniform sampling reduction strategy, which undersamples areas with complex details or dynamic shading. To address this, we propose to sparsely shade critical areas while reusing generated pixels in low-variation areas for neural extrapolation. Specifically, we introduce the Predictive Error-Flow-eXtrapolation Network (EFXNet)-an architecture that predicts extrapolation errors, estimates flows, and extrapolates frames at once. Firstly, EFXNet leverages temporal coherence to predict extrapolation error and guide the sparse shading of dynamic areas. In addition, EFXNet employs a target-grid correlation module to estimate robust optical flows from pixel correlations rather than pixel values. Finally, EFXNet uses dedicated motion representations for the historical geometric and lighting components, respectively, to extrapolate temporally stable frames. Extensive experimental results show that, compared with state-of-the-art methods, our frame extrapolation method exhibits superior visual quality and temporal stability under a low rendering budget.
Zhizhen Wu, Yazhen Yuan, Zhilong Yuan, Rui Wang 0004, Yuchi Huo
ACM Trans. Graph.3
2025 MoFlow: Motion-Guided Flows for Recurrent Rendered Frame Prediction
abstract
Rendering realistic images in real-time on high-frame-rate display devices poses considerable challenges, even with advanced graphics cards. This stimulates a demand for frame prediction technologies to boost frame rates. The key to these algorithms is to exploit spatiotemporal coherence by warping rendered pixels with motion representations. However, existing motion estimation methods can suffer from low precision, high overhead, and incomplete support for visual effects. In this article, we present a rendered frame prediction framework with a novel motion representation, dubbed motion-guided flow (MoFlow) , aiming at overcoming the intrinsic limitations of optical flow and motion vectors and precisely capture the dynamics of intricate geometries, lighting, and translucent objects. Notably, we construct MoFlows using a recurrent feature streaming network, which specializes in learning latent motion features from multiple frames. The results of extensive experiments demonstrate that, compared to state-of-the-art methods, our method achieves superior visual quality and temporal stability with lower latency. The recurrent mechanism allows our method to predict single or multiple consecutive frames, increasing the frame rate by over 2×. The proposed approach represents a flexible pipeline to meet the demands of various graphics applications, devices, and scenarios.
Zhizhen Wu, Zhilong Yuan, Chenyu Zuo, Yazhen Yuan, Yifan Peng 0001, Guiyang Pu, Rui Wang 0004, Yuchi Huo
ACM Trans. Graph.4
2024 RobIR: Robust Inverse Rendering for High-Illumination Scenes
abstract
Implicit representation has opened up new possibilities for inverse rendering. However, existing implicit neural inverse rendering methods struggle to handle strongly illuminated scenes with significant shadows and slight reflections. The existence of shadows and reflections can lead to an inaccurate understanding of the scene, making precise factorization difficult. To this end, we present RobIR, an implicit inverse rendering approach that uses ACES tone mapping and regularized visibility estimation to reconstruct accurate BRDF of the object. By accurately modeling the indirect radiance field, normal, visibility, and direct light simultaneously, we are able to accurately decouple environment lighting and the object's PBR materials without imposing strict constraints on the scene. Even in high-illumination scenes with shadows and specular reflections, our method can recover high-quality albedo and roughness with no shadow interference. RobIR outperforms existing methods in both quantitative and qualitative evaluations.
Ziyi Yang 0008, Chenyanzhen, Yazhen Yuan, Xiaogang Jin 0001
NeurIPS4
2023 Adaptive Recurrent Frame Prediction with Learnable Motion Vectors
abstract
The utilization of dedicated ray tracing graphics cards has revolutionized the production of stunning visual effects in real-time rendering. However, the demand for high frame rates and high resolutions remains a challenge. The pixel warping approach is a crucial technique for increasing frame rate and resolution by exploiting the spatio-temporal coherence. To this end, existing super-resolution and frame prediction methods rely heavily on motion vectors from rendering engine pipelines to track object movements. This work builds upon state-of-the-art heuristic approaches by exploring a novel adaptive recurrent frame prediction framework that integrates learnable motion vectors. Our framework supports the prediction of transparency, particles, and texture animations, with improved motion vectors that capture shading, reflections, and occlusions, in addition to geometry movements. In addition, we introduce a feature streaming neural network, dubbed FSNet, that allows for the adaptive prediction of one or multiple sequential frames. Extensive experiments against state-of-the-art methods demonstrate that FSNet can operate at lower latency with significant visual enhancements and can upscale frame rates by at least two times. This approach offers a flexible pipeline to improve the rendering frame rates of various graphics applications and devices.
Zhizhen Wu, Chenyu Zuo, Yuchi Huo, Yazhen Yuan, Yifan Peng 0001, Guiyang Pu, Rui Wang 0004, Hujun Bao
SIGGRAPH Asia4
2022 Multirate Shading with Piecewise Interpolatory Approximation
abstract
Abstract Evaluating shading functions on geometry surfaces dominates the rendering computation. A high‐quality but time‐consuming estimate is usually achieved with a dense sampling rate for pixels or sub‐pixels. In this paper, we leverage sparsely sampled points on vertices of dynamically‐generated subdivision surfaces to approximate the ground‐truth shading signal by piecewise linear reconstruction. To control the introduced interpolation error at runtime, we analytically derive an L∞ error bound and compute the optimal subdivision surfaces based on a user‐specified error threshold. We apply our analysis on multiple shading functions including Lambertian, Blinn‐Phong, Microfacet BRDF and also extend it to handle textures, yielding easy‐to‐compute formulas. To validate our derivation, we design a forward multirate shading algorithm powered by hardware tessellator that moves shading computation at pixels to the vertices of subdivision triangles on the fly. We show our approach significantly reduces the sampling rates on various test cases, reaching a speedup ratio of 134% ~ 283% compared to dense per‐pixel shading in current graphics hardware.
Yazhen Yuan, Rui Wang 0004, Hujun Bao
Comput. Graph. Forum2
2020 Tile Pair-Based Adaptive Multi-Rate Stereo Shading
abstract
This work proposes a new stereo shading architecture that enables adaptive shading rates and automatic shading reuse among triangles and between two views. The proposed pipeline presents several novel features. First, the present sort-middle/bin shading is extended to tile pair-based shading to rasterize and shade pixels at two views simultaneously. A new rasterization algorithm utilizing epipolar geometry is then proposed to schedule tile pairs and perform rasterization at stereo views efficiently. Second, this work presents an adaptive multi-rate shading framework to compute shading on pixels at different rates. A novel tile-based screen space cache and a new cache reuse shader are proposed to perform such multi-rate shading across triangles and views. The results show that the newly proposed method outperforms the standard sort-middle shading and the state-of-the-art multi-rate shading by achieving considerably lower shading cost and memory bandwidth.
Yazhen Yuan, Rui Wang 0004, Hujun Bao
IEEE Trans. Vis. Comput. Graph.1
2018 Runtime Shader Simplification via Instant Search in Reduced Optimization Space
abstract
Abstract Traditional automatic shader simplification simplifies shaders in an offline process, which is typically carried out in a context‐oblivious manner or with the use of some example contexts, e.g., certain hardware platforms, scenes, and uniform parameters, etc. As a result, these pre‐simplified shaders may fail at adapting to runtime changes of the rendering context that were not considered in the simplification process. In this paper, we propose a new automatic shader simplification technique, which explores two key aspects of a runtime simplification framework: the optimization space and the instant search for optimal simplified shaders with runtime context. The proposed technique still requires a preprocess stage to process the original shader. However, instead of directly computing optimal simplified shaders, the proposed preprocess generates a reduced shader optimization space. In particular, two heuristic estimates of the quality and performance of simplified shaders are presented to group similar variants into representative ones, which serve as basic graph nodes of the simplification dependency graph (SDG), a new representation of the optimization space. At the runtime simplification stage, a parallel discrete optimization algorithm is employed to instantly search in the SDG for optimal simplified shaders. New data‐driven cost models are proposed to predict the runtime quality and performance of simplified shaders on the basis of data collected during runtime. Results show that the selected simplifications of complex shaders achieve 1.6 to 2.5 times speedup and still retain high rendering quality.
Yazhen Yuan, Rui Wang 0004, Tianlei Hu, Hujun Bao
Comput. Graph. Forum1
2016 Simplified and tessellated mesh for realtime high quality rendering
Yazhen Yuan, Rui Wang 0004, Jin Huang 0001, Yanming Jia, Hujun Bao
Comput. Graph.1
2014 Automatic shader simplification using surface signal approximation
abstract
In this paper, we present a new automatic shader simplification method using surface signal approximation. We regard the entire multi-stage rendering pipeline as a process that generates signals on surfaces, and we formulate the simplification of the fragment shader as a global simplification problem across multi-shader stages. Three new shader simplification rules are proposed to solve the problem. First, the code transformation rule transforms fragment shader code to other shader stages in order to redistribute computations on pixels up to the level of geometry primitives. Second, the surface-wise approximation rule uses high-order polynomial basis functions on surfaces to approximate pixel-wise computations in the fragment shader. These approximations are pre-cached and simplify computations at runtime. Third, the surface subdivision rule tessellates surfaces into smaller patches. It combines with the previous two rules to approximate pixel-wise signals at different levels of tessellations with different computation times and visual errors. To evaluate simplified shaders using these simplification rules, we introduce a new cost model that includes the visual quality, rendering time and memory consumption. With these simplification rules and the cost model, we present an integrated shader simplification algorithm that is capable of automatically generating variants of simplified shaders and selecting a sequence of preferable shaders. Results show that the sequence of selected simplified shaders balance performance, accuracy and memory consumption well.
Rui Wang 0004, Xianjin Yang, Yazhen Yuan, Wei Chen 0001, Kavita Bala, Hujun Bao
ACM Trans. Graph.3
2013 GPU-based out-of-core many-lights rendering
abstract
In this paper, we present a GPU-based out-of-core rendering approach under the many-lights rendering framework. Many-lights rendering is an efficient and scalable rendering framework for a large number of lights. But when the data sizes of lights and geometry are both beyond the in-core memory storage size, the data management of these two out-of-core data becomes critical and challenging. In our approach, we formulate such a data management as a graph traversal optimization problem that first builds out-of-core lights and geometry data into a graph, and then guides shading computations by finding a shortest path to visit all vertices in the graph. Based on the proposed data management, we develop a GPU-based out-of-GPU-core rendering algorithm that manages data between the CPU host memory and the GPU device memory. Two main steps are taken in the algorithm: the out-of-core data preparation to pack data into optimal data layouts for the many-lights rendering, and the out-of-core shading using graph-based data management. We demonstrate our algorithm on scenes with out-of-core detailed geometry and out-of-core lights. Results show that our approach generates complex global illumination effects with increased data access coherence and has one order of magnitude performance gain over the CPU-based approach.
Rui Wang 0004, Yuchi Huo, Yazhen Yuan, Kun Zhou 0001, Wei Hua 0002, Hujun Bao
ACM Trans. Graph.3