Duan Gao

dblp:245/4838 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-0647-6160ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Rendering · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
global illumination
0.712023
Neural Global Illumination: Interactive Indirect Illumination Prediction Under Dynamic Area Lights · IEEE Trans. Vis. Comput. Graph. 2023
Rendering › relighting
image-based relighting
0.412020
Deferred neural lighting: free-viewpoint relighting from unstructured photographs · ACM Trans. Graph. 2020
Rendering
neural rendering
0.412020
Deferred neural lighting: free-viewpoint relighting from unstructured photographs · ACM Trans. Graph. 2020
Rendering
appearance modeling
0.412019
Deep inverse rendering for high-resolution SVBRDF estimation from an arbitrary number of images · ACM Trans. Graph. 2019
Rendering
inverse rendering
0.412019
Deep inverse rendering for high-resolution SVBRDF estimation from an arbitrary number of images · ACM Trans. Graph. 2019
Rendering › appearance acquisition › material acquisition
SVBRDF estimation
0.412019
Deep inverse rendering for high-resolution SVBRDF estimation from an arbitrary number of images · ACM Trans. Graph. 2019
Rendering
interactive rendering
0.212023
Neural Global Illumination: Interactive Indirect Illumination Prediction Under Dynamic Area Lights · IEEE Trans. Vis. Comput. Graph. 2023

Methods — techniques the papers use, named apart from their topics

screen-space neural buffer · 0.7positional encoding · 0.7deep rendering network · 0.7radiance cues · 0.4neural rendering network · 0.4augmentation refinement · 0.4latent embedding · 0.4fully convolutional auto-encoder · 0.4deep inverse rendering · 0.4
YearPublicationVenuePosition
2025 UniFaceGAN: High-Quality 3D Face Editing With a Unified Latent Space
abstract
Recent advancements in 3D face generation have explored various representation and generative models. However, these methods often offer limited 3D face editing capabilities. In this paper, we introduce UniFaceGAN, a novel framework for 3D facial editing, leveraging a unified latent space to facilitate diverse and user-friendly 3D facial manipulation. The key to efficient 3D facial editing lies in establishing a representation space that offers essential facial priors. To achieve this, we propose encoding high-dimensional 3D faces into a compact, disentangled latent space which is learned through conditional 3D GANs guided by text descriptions. With the help of the GAN inversion techniques, UniFaceGAN allows us to edit existing 3D faces, accompanied by a residual editing strategy to mitigate inversion errors efficiently. We demonstrate UniFaceGAN can generate high-quality 3D faces and supports various 3D face editing applications, including CLIP-based stylizations, multiple-point-based drag manipulation, and local blending among multiple faces.
Jinfu Wei, Ran Liao, Duan Gao
ICASSP4
2025 Few-Shot 3D Face Generation via a Controllable Diffusion Model Guided by Text and Images
abstract
Recent advancements in text-to-3D generation have relied on large 3D datasets or expensive optimization processes during inference. In this paper, we introduce ControlFace, a novel framework designed for the creation of computer graphics-friendly 3D faces under the guidance of text and images. We utilize a controllable diffusion model to generate physically-based facial assets in texture space. The key to achieving few-shot generation lies in 3D-aware controls: a texture-space facial representation of geometry proxy. The main distinguishing feature of our framework is the effective integration of 3D facial priors with the diversity inherited from text-to-image diffusion models through few-shot learning, requiring only 36 3D faces for training. Once trained, ControlFace can generate diverse 3D faces in a feed-forward manner within 5 seconds and perform editing and stylization without 3D labeled data. We have demonstrated the effectiveness of our method in generating and editing various digital characters, guided by multi-modal controls.
Jinfu Wei, Qinchuan Zhang, Ran Liao, Duan Gao
ICME5
2025 DreamPBR: Text-driven High-Resolution SVBRDF Generation with Multimodal Guidance
abstract
Existing material creation methods are limited in diversity due to the scarcity of real-world data. To enhance controllability and diversity, we propose DreamPBR, a diffusion-based generative framework that creates spatially varying appearance properties guided by text and multimodal controls. By integrating large-scale vision-language models trained on billions of text-image pairs with material priors from hundreds of Physically Based Rendering (PBR) samples, we achieve high-quality PBR material generation. We employ a material Latent Diffusion Model (m-LDM) to map albedo maps to latent space, which is then decoded into full Spatially Varying Bidirectional Reflectance Distribution Function (SVBRDF) parameter maps via a rendering-aware PBR decoder. To achieve diverse control, we introduce a multimodal guidance module that includes image and 3D shape guidance. We demonstrate DreamPBR’s effectiveness in material creation, showcasing its versatility and user-friendliness across various controllable generation and editing applications.
Linxuan Xin, Zhiyi Pan 0001, Jinfu Wei, Duan Gao, Wei Gao 0003
ICME5
2023 Neural Global Illumination: Interactive Indirect Illumination Prediction Under Dynamic Area Lights
abstract
We propose neural global illumination, a novel method for fast rendering full global illumination in static scenes with dynamic viewpoint and area lighting. The key idea of our method is to utilize a deep rendering network to model the complex mapping from each shading point to global illumination. To efficiently learn the mapping, we propose a neural-network-friendly input representation including attributes of each shading point, viewpoint information, and a combinational lighting representation that enables high-quality fitting with a compact neural network. To synthesize high-frequency global illumination effects, we transform the low-dimension input to higher-dimension space by positional encoding and model the rendering network as a deep fully-connected network. Besides, we feed a screen-space neural buffer to our rendering network to share global information between objects in the screen-space to each shading point. We have demonstrated our neural global illumination method in rendering a wide variety of scenes exhibiting complex and all-frequency global illumination effects such as multiple-bounce glossy interreflection, color bleeding, and caustics.
Duan Gao, Haoyuan Mu, Kun Xu 0003
IEEE Trans. Vis. Comput. Graph.1
2020 Deferred neural lighting: free-viewpoint relighting from unstructured photographs
abstract
We present deferred neural lighting, a novel method for free-viewpoint relighting from unstructured photographs of a scene captured with handheld devices. Our method leverages a scene-dependent neural rendering network for relighting a rough geometric proxy with learnable neural textures. Key to making the rendering network lighting aware are radiance cues: global illumination renderings of a rough proxy geometry of the scene for a small set of basis materials and lit by the target lighting. As such, the light transport through the scene is never explicitely modeled, but resolved at rendering time by a neural rendering network. We demonstrate that the neural textures and neural renderer can be trained end-to-end from unstructured photographs captured with a double hand-held camera setup that concurrently captures the scene while being lit by only one of the cameras' flash lights. In addition, we propose a novel augmentation refinement strategy that exploits the linearity of light transport to extend the relighting capabilities of the neural rendering network to support other lighting types (e.g., environment lighting) beyond the lighting used during acquisition (i.e., flash lighting). We demonstrate our deferred neural lighting solution on a variety of real-world and synthetic scenes exhibiting a wide range of material properties, light transport effects, and geometrical complexity.
Duan Gao, Yue Dong 0001, Pieter Peers, Kun Xu 0003, Xin Tong 0001
ACM Trans. Graph.1
2019 Deep inverse rendering for high-resolution SVBRDF estimation from an arbitrary number of images
abstract
In this paper we present a unified deep inverse rendering framework for estimating the spatially-varying appearance properties of a planar exemplar from an arbitrary number of input photographs, ranging from just a single photograph to many photographs. The precision of the estimated appearance scales from plausible when the input photographs fails to capture all the reflectance information, to accurate for large input sets. A key distinguishing feature of our framework is that it directly optimizes for the appearance parameters in a latent embedded space of spatially-varying appearance, such that no handcrafted heuristics are needed to regularize the optimization. This latent embedding is learned through a fully convolutional auto-encoder that has been designed to regularize the optimization. Our framework not only supports an arbitrary number of input photographs, but also at high resolution. We demonstrate and evaluate our deep inverse rendering solution on a wide variety of publicly available datasets.
Duan Gao, Xiao Li 0030, Yue Dong 0001, Pieter Peers, Kun Xu 0003, Xin Tong 0001
ACM Trans. Graph.1