Xueming Yu

dblp:41/7188 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0009-8189-6024ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
11 papers
Rendering · 50% Computational photography and imaging · 22% Image and video processing · 10%
Artificial intelligence
2 papers
Generative modeling · 92% 3D vision · 8%

Topics — the 24 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
relighting
2.042026
BodyReLux: Temporally Consistent Full-Body Video Relighting · ACM Trans. Graph. 2026
DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024
Deep reflectance fields: high-quality facial reflectance field inference from color gradient illumination · ACM Trans. Graph. 2019
Rendering › relighting
portrait relighting
1.222025
Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset · CVPR 2025
Single image portrait relighting · ACM Trans. Graph. 2019
Visual content generation and editing › visual effects
video relighting
1.012026
BodyReLux: Temporally Consistent Full-Body Video Relighting · ACM Trans. Graph. 2026
Machine learning › Generative modeling
diffusion model
0.912025
Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset · CVPR 2025
Image and video processing › image enhancement
detail enhancement
0.912025
Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture · ACM Trans. Graph. 2025
Rendering
gaussian splatting
0.912025
Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture · ACM Trans. Graph. 2025
Virtual and augmented reality
volumetric capture
0.912025
Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture · ACM Trans. Graph. 2025
Computational photography and imaging
high dynamic range imaging
0.812024
DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024
Computational photography and imaging
image relighting
0.412019
Single image portrait relighting · ACM Trans. Graph. 2019
Computational photography and imaging › 3d scanning
volumetric performance capture
0.412019
The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019
Image and video processing › video processing
temporal consistency
0.312025
Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset · CVPR 2025
Computational photography and imaging › illumination estimation
illumination capture
0.212016
Practical multispectral lighting reproduction · ACM Trans. Graph. 2016
Rendering › global illumination
image-based lighting
0.212016
Practical multispectral lighting reproduction · ACM Trans. Graph. 2016
Rendering › novel view synthesis
free-viewpoint rendering
0.212024
DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024
Rendering
image-based rendering
0.212024
DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024
Rendering
novel view synthesis
0.212024
DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024
Rendering › appearance acquisition
shape and reflectance capture
0.212013
Acquiring reflectance and shape from continuous spherical harmonic illumination · ACM Trans. Graph. 2013
Rendering › lighting
spherical harmonic lighting
0.212013
Acquiring reflectance and shape from continuous spherical harmonic illumination · ACM Trans. Graph. 2013
Computer animation and physical simulation › performance capture
facial performance capture
0.112011
Multiview face capture using polarized spherical gradient illumination · ACM Trans. Graph. 2011
Rendering › appearance acquisition
reflectance map estimation
0.112019
The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019
Virtual and augmented reality › telepresence
3d teleconferencing
0.112009
Achieving eye contact in a one-to-many 3D video teleconferencing system · ACM Trans. Graph. 2009
Rendering
reflectance modeling
0.012013
Acquiring reflectance and shape from continuous spherical harmonic illumination · ACM Trans. Graph. 2013
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.012011
Multiview face capture using polarized spherical gradient illumination · ACM Trans. Graph. 2011
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction
0.012011
Multiview face capture using polarized spherical gradient illumination · ACM Trans. Graph. 2011

Methods — techniques the papers use, named apart from their topics

lighting injection · 1.7conditional video diffusion · 1.7video diffusion model · 1.0masked attention · 1.0data augmentation · 1.0diffusion-based detail enhancement · 0.94d gaussian splatting · 0.9diffusion model · 0.8color gradient illumination · 0.83d gaussian splatting · 0.8photometric stereo · 0.1message passing stereo · 0.1
YearPublicationVenuePosition
2026 BodyReLux: Temporally Consistent Full-Body Video Relighting
abstract
Being able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject-specific video diffusion-based framework for relighting full-body human performances in a temporally consistent way. Our model is trained on a hybrid dataset of pixel-aligned video relighting pairs, covering a diverse combination of lighting conditions, performances and viewpoints. To acquire such dataset, we combine traditional static One-Light-at-a-Time (OLAT) capture and a novel dynamic performance capture in which two smoothly varying lighting sequences are rapidly interleaved. Because the lighting operates above the human flicker-fusion threshold, the interleaving does not appears to strobe. We train our video relighting model from a pretrained text-to-video model to fully leverage the generative priors for producing high quality videos. To achieve accurate lighting control, we introduce a new lighting conditioning method that represents each light source as a token. We further condition on sequences of lighting using masked attention to support dynamic lighting control. Together with a carefully designed data augmentation pipeline, we achieve photorealistic, robust, and temporally consistent video relighting of subject-specific human performances.
Mingming He, Xueming Yu, David M. George, Ahmet Levent Tasel, Paul E. Debevec, Julien Philip
ACM Trans. Graph.3
2025 Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset
abstract
Video portrait relighting remains challenging because the results need to be both photorealistic and temporally stable. This typically requires a strong model design that can capture complex facial reflections as well as intensive training on a high-quality paired video dataset, such as dynamic one-light-at-a-time (OLAT). In this work, we introduce Lux Post Facto, a novel portrait video relighting method that produces both photorealistic and temporally consistent lighting effects. From the model side, we design a new conditional video diffusion model built upon state-of-the-art pre-trained video diffusion model, alongside a new lighting injection mechanism to enable precise control. This way we leverage strong spatial and temporal generative capability to generate plausible solutions to the ill-posed relighting problem. Our technique uses a hybrid dataset consisting of static expression OLAT data and in-the-wild portrait performance videos to jointly learn relighting and temporal modeling. This avoids the need to acquire paired video data in different lighting conditions. Our extensive experiments show that our model produces state-of-the-art results both in terms of photorealism and temporal consistency. Video results can be found on our project page.
Yiqun Mei, Mingming He, Julien Philip, Wenqi Xian, David M. George, Xueming Yu, Gabriel Dedic, Ahmet Levent Tasel, Ning Yu 0006, Vishal M. Patel, Paul E. Debevec
CVPR7
2025 Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture
abstract
We present a unique system for large-scale, multi-performer, high resolution 4D volumetric capture providing realistic free-viewpoint video up to and including 4K resolution facial closeups. To achieve this, we employ a novel volumetric capture, reconstruction and rendering pipeline based on Dynamic Gaussian Splatting and Diffusion-based Detail Enhancement. We design our pipeline specifically to meet the demands of high-end media production. We employ two capture rigs: the Scene Rig , which captures multi-actor performances at a resolution which falls short of 4K production quality, and the Face Rig , which records high-fidelity single-actor facial detail to serve as a reference for detail enhancement. We first reconstruct dynamic performances from the Scene Rig using 4D Gaussian Splatting, incorporating new model designs and training strategies to improve reconstruction, dynamic range, and rendering quality. Then to render high-quality images for facial closeups, we introduce a diffusion-based detail enhancement model. This model is fine-tuned with high-fidelity data from the same actors recorded in the Face Rig. We train on paired data generated from low- and high-quality Gaussian Splatting (GS) models, using the low-quality input to match the quality of the Scene Rig , with the high-quality GS as ground truth. Our results demonstrate the effectiveness of this pipeline in bridging the gap between the scalable performance capture of a large-scale rig and the high-resolution standards required for film and media production.
Julien Philip, Pascal Clausen, Wenqi Xian, Ahmet Levent Tasel, Mingming He, Xueming Yu, David M. George, Ning Yu 0006, Oliver Pilarski, Paul E. Debevec
ACM Trans. Graph.7
2024 DifFRelight: Diffusion-Based Facial Performance Relighting
abstract
We present a novel framework for free-viewpoint facial performance relighting using diffusion-based image-to-image translation. Leveraging a subject-specific dataset containing diverse facial expressions captured under various lighting conditions, including flat-lit and one-light-at-a-time (OLAT) scenarios, we train a diffusion model for precise lighting control, enabling high-fidelity relit facial images from flat-lit inputs. Our framework includes spatially-aligned conditioning of flat-lit captures and random noise, along with integrated lighting information for global control, utilizing prior knowledge from the pre-trained Stable Diffusion model. This model is then applied to dynamic facial performances captured in a consistent flat-lit environment and reconstructed for novel-view synthesis using a scalable dynamic 3D Gaussian Splatting method to maintain quality and consistency in the relit results. In addition, we introduce unified lighting control by integrating a novel area lighting representation with directional lighting, allowing for joint adjustments in light size and direction. We also enable high dynamic range imaging (HDRI) composition using multiple directional lights to produce dynamic sequences under complex lighting conditions. Our evaluations demonstrate the models efficiency in achieving precise lighting control and generalizing across various facial expressions while preserving detailed features such as skintexture andhair. The model accurately reproduces complex lighting effects like eye reflections, subsurface scattering, self-shadowing, and translucency, advancing photorealism within our framework.
Mingming He, Pascal Clausen, Ahmet Levent Tasel, Oliver Pilarski, Wenqi Xian, Laszlo Rikker, Xueming Yu, Ryan D. Burgert, Ning Yu 0006, Paul E. Debevec
SIGGRAPH Asia8
2019 The relightables: volumetric performance capture of humans with realistic relighting
abstract
We present "The Relightables", a volumetric capture system for photorealistic and high quality relightable full-body performance capture. While significant progress has been made on volumetric capture systems, focusing on 3D geometric reconstruction with high resolution textures, much less work has been done to recover photometric properties needed for relighting. Results from such systems lack high-frequency details and the subject's shading is prebaked into the texture. In contrast, a large body of work has addressed relightable acquisition for image-based approaches, which photograph the subject under a set of basis lighting conditions and recombine the images to show the subject as they would appear in a target lighting environment. However, to date, these approaches have not been adapted for use in the context of a high-resolution volumetric capture system. Our method combines this ability to realistically relight humans for arbitrary environments, with the benefits of free-viewpoint volumetric capture and new levels of geometric accuracy for dynamic performances. Our subjects are recorded inside a custom geodesic sphere outfitted with 331 custom color LED lights, an array of high-resolution cameras, and a set of custom high-resolution depth sensors. Our system innovates in multiple areas: First, we designed a novel active depth sensor to capture 12.4 MP depth maps, which we describe in detail. Second, we show how to design a hybrid geometric and machine learning reconstruction pipeline to process the high resolution input and output a volumetric video. Third, we generate temporally consistent reflectance maps for dynamic performers by leveraging the information contained in two alternating color gradient illumination images acquired at 60Hz. Multiple experiments, comparisons, and applications show that The Relightables significantly improves upon the level of realism in placing volumetrically captured human performances into arbitrary CG scenes.
Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Ryan Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor 0001, Paul E. Debevec, Shahram Izadi
ACM Trans. Graph.5
2019 Deep reflectance fields: high-quality facial reflectance field inference from color gradient illumination
abstract
We present a novel technique to relight images of human faces by learning a model of facial reflectance from a database of 4D reflectance field data of several subjects in a variety of expressions and viewpoints. Using our learned model, a face can be relit in arbitrary illumination environments using only two original images recorded under spherical color gradient illumination. The output of our deep network indicates that the color gradient images contain the information needed to estimate the full 4D reflectance field, including specular reflections and high frequency details. While capturing spherical color gradient illumination still requires a special lighting setup, reduction to just two illumination conditions allows the technique to be applied to dynamic facial performance capture. We show side-by-side comparisons which demonstrate that the proposed system outperforms the state-of-the-art techniques in both realism and speed.
Abhimitra Meka, Christian Häne, Rohit Pandey, Michael Zollhöfer, Sean Ryan Fanello, Graham Fyffe, Adarsh Kowdle, Xueming Yu, Jay Busch, Jason Dourgarian, Peter Denny, Sofien Bouaziz, Peter Lincoln, Matt Whalen, Geoff Harvey, Jonathan Taylor 0001, Shahram Izadi, Andrea Tagliasacchi, Paul E. Debevec, Christian Theobalt, Julien P. C. Valentin, Christoph Rhemann
ACM Trans. Graph.8
2019 Single image portrait relighting
abstract
Lighting plays a central role in conveying the essence and depth of the subject in a portrait photograph. Professional photographers will carefully control the lighting in their studio to manipulate the appearance of their subject, while consumer photographers are usually constrained to the illumination of their environment. Though prior works have explored techniques for relighting an image, their utility is usually limited due to requirements of specialized hardware, multiple images of the subject under controlled or known illuminations, or accurate models of geometry and reflectance. To this end, we present a system for portrait relighting : a neural network that takes as input a single RGB image of a portrait taken with a standard cellphone camera in an unconstrained environment, and from that image produces a relit image of that subject as though it were illuminated according to any provided environment map. Our method is trained on a small database of 18 individuals captured under different directional light sources in a controlled light stage setup consisting of a densely sampled sphere of lights. Our proposed technique produces quantitatively superior results on our dataset's validation set compared to prior works, and produces convincing qualitative relighting results on a dataset of hundreds of real-world cellphone portraits. Because our technique can produce a 640 × 640 image in only 160 milliseconds, it may enable interactive user-facing photographic applications in the future.
Tiancheng Sun, Jonathan T. Barron, Yun-Ta Tsai, Zexiang Xu, Xueming Yu, Graham Fyffe, Christoph Rhemann, Jay Busch, Paul E. Debevec, Ravi Ramamoorthi
ACM Trans. Graph.5
2016 Practical multispectral lighting reproduction
abstract
We present a practical framework for reproducing omnidirectional incident illumination conditions with complex spectra using a light stage with multispectral LED lights. For lighting acquisition, we augment standard RGB panoramic photography with one or more observations of a color chart with numerous reflectance spectra. We then solve for how to drive the multispectral light sources so that they best reproduce the appearance of the color charts in the original lighting. Even when solving for non-negative intensities, we show that accurate lighting reproduction is achievable using just four or six distinct LED spectra for a wide range of incident illumination spectra. A significant benefit of our approach is that it does not require the use of specialized equipment (other than the light stage) such as monochromators, spectroradiometers, or explicit knowledge of the LED power spectra, camera spectral response functions, or color chart reflectance spectra. We describe two simple devices for multispectral lighting capture, one for slow measurements of detailed angular spectral detail, and one for fast measurements with coarse angular detail. We validate the approach by realistically compositing real subjects into acquired lighting environments, showing accurate matches to how the subject would actually look within the environments, even for those including complex multispectral illumination. We also demonstrate dynamic lighting capture and playback using the technique.
Chloe LeGendre, Xueming Yu, Dai Liu, Jay Busch, Val Jones 0002, Sumanta N. Pattanaik, Paul E. Debevec
ACM Trans. Graph.2
2013 Measurement-Based Synthesis of Facial Microgeometry
abstract
Abstract We present a technique for generating microstructure‐level facial geometry by augmenting a mesostructure‐level facial scan with detail synthesized from a set of exemplar skin patches scanned at much higher resolution. Additionally, we make point‐source reflectance measurements of the skin patches to characterize the specular reflectance lobes at this smaller scale and analyze facial reflectance variation at both the mesostructure and microstructure scales. We digitize the exemplar patches with a polarization‐based computational illumination technique which considers specular reflection and single scattering. The recorded microstructure patches can be used to synthesize full‐facial microstructure detail for either the same subject or to a different subject. We show that the technique allows for greater realism in facial renderings including more accurate reproduction of skin's specular reflection effects.
Paul Graham, Borom Tunwattanapong, Jay Busch, Xueming Yu, Val Jones 0002, Paul E. Debevec, Abhijeet Ghosh
Comput. Graph. Forum4
2013 Acquiring reflectance and shape from continuous spherical harmonic illumination
abstract
We present a novel technique for acquiring the geometry and spatially-varying reflectance properties of 3D objects by observing them under continuous spherical harmonic illumination conditions. The technique is general enough to characterize either entirely specular or entirely diffuse materials, or any varying combination across the surface of the object. We employ a novel computational illumination setup consisting of a rotating arc of controllable LEDs which sweep out programmable spheres of incident illumination during 1-second exposures. We illuminate the object with a succession of spherical harmonic illumination conditions, as well as photographed environmental lighting for validation. From the response of the object to the harmonics, we can separate diffuse and specular reflections, estimate world-space diffuse and specular normals, and compute anisotropic roughness parameters for each view of the object. We then use the maps of both diffuse and specular reflectance to form correspondences in a multiview stereo algorithm, which allows even highly specular surfaces to be corresponded across views. The algorithm yields a complete 3D model and a set of merged reflectance maps. We use this technique to digitize the shape and reflectance of a variety of objects difficult to acquire with other techniques and present validation renderings which match well to photographs in similar lighting.
Borom Tunwattanapong, Graham Fyffe, Paul Graham, Jay Busch, Xueming Yu, Abhijeet Ghosh, Paul E. Debevec
ACM Trans. Graph.5
2011 Single-shot photometric stereo by spectral multiplexing
abstract
We propose a novel method for single-shot photometric stereo by spectral multiplexing. The output of our method is a simultaneous per-pixel estimate of the surface normal and full-color reflectance. Our method is well suited to materials with varying color and texture, requires no time-varying illumination, and no high-speed cameras. Being a single-shot method, it may be applied to dynamic scenes without any need for optical flow. Our key contributions are a generalization of three-color photometric stereo to more than three color channels, and the design of a practical six-color-channel system using off-the-shelf parts.
Graham Fyffe, Xueming Yu, Paul E. Debevec
ICCP2
2011 Multiview face capture using polarized spherical gradient illumination
abstract
We present a novel process for acquiring detailed facial geometry with high resolution diffuse and specular photometric information from multiple viewpoints using polarized spherical gradient illumination. Key to our method is a new pair of linearly polarized lighting patterns which enables multiview diffuse-specular separation under a given spherical illumination condition from just two photographs. The patterns -- one following lines of latitude and one following lines of longitude -- allow the use of fixed linear polarizers in front of the cameras, enabling more efficient acquisition of diffuse and specular albedo and normal maps from multiple viewpoints. In a second step, we employ these albedo and normal maps as input to a novel multi-resolution adaptive domain message passing stereo reconstruction algorithm to create high resolution facial geometry. To do this, we formulate the stereo reconstruction from multiple cameras in a commonly parameterized domain for multiview reconstruction. We show competitive results consisting of high-resolution facial geometry with relightable reflectance maps using five DSLR cameras. Our technique scales well for multiview acquisition without requiring specialized camera systems for sensing multiple polarization states.
Abhijeet Ghosh, Graham Fyffe, Borom Tunwattanapong, Jay Busch, Xueming Yu, Paul E. Debevec
ACM Trans. Graph.5
2009 Achieving eye contact in a one-to-many 3D video teleconferencing system
abstract
We present a set of algorithms and an associated display system capable of producing correctly rendered eye contact between a three-dimensionally transmitted remote participant and a group of observers in a 3D teleconferencing system. The participant's face is scanned in 3D at 30Hz and transmitted in real time to an autostereoscopic horizontal-parallax 3D display, displaying him or her over more than a 180° field of view observable to multiple observers. To render the geometry with correct perspective, we create a fast vertex shader based on a 6D lookup table for projecting 3D scene vertices to a range of subject angles, heights, and distances. We generalize the projection mathematics to arbitrarily shaped display surfaces, which allows us to employ a curved concave display surface to focus the high speed imagery to individual observers. To achieve two-way eye contact, we capture 2D video from a cross-polarized camera reflected to the position of the virtual participant's eyes, and display this 2D video feed on a large screen in front of the real participant, replicating the viewpoint of their virtual self. To achieve correct vertical perspective, we further leverage this image to track the position of each audience member's eyes, allowing the 3D display to render correct vertical perspective for each of the viewers around the device. The result is a one-to-many 3D teleconferencing system able to reproduce the effects of gaze, attention, and eye contact generally missing in traditional teleconferencing systems.
Val Jones 0002, Magnus Lang, Graham Fyffe, Xueming Yu, Jay Busch, Ian McDowall, Mark T. Bolas, Paul E. Debevec
ACM Trans. Graph.4