VLDB 2026 Research / reviewers in the wild / expert
Xueming Yu
dblp:41/7188
· DBLP profile ↗
13ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0009-8189-6024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
11 papers |
Rendering · 50% Computational photography and imaging · 22% Image and video processing · 10% | |
| Artificial intelligence
2 papers |
Generative modeling · 92% 3D vision · 8% |
Topics — the 24 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Rendering
relighting |
2.0 | 4 | 2026 | BodyReLux: Temporally Consistent Full-Body Video Relighting · ACM Trans. Graph. 2026 DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024 Deep reflectance fields: high-quality facial reflectance field inference from color gradient illumination · ACM Trans. Graph. 2019 |
Rendering › relighting
portrait relighting |
1.2 | 2 | 2025 | Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset · CVPR 2025 Single image portrait relighting · ACM Trans. Graph. 2019 |
Visual content generation and editing › visual effects
video relighting |
1.0 | 1 | 2026 | BodyReLux: Temporally Consistent Full-Body Video Relighting · ACM Trans. Graph. 2026 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset · CVPR 2025 |
Image and video processing › image enhancement
detail enhancement |
0.9 | 1 | 2025 | Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture · ACM Trans. Graph. 2025 |
Rendering
gaussian splatting |
0.9 | 1 | 2025 | Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture · ACM Trans. Graph. 2025 |
Virtual and augmented reality
volumetric capture |
0.9 | 1 | 2025 | Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture · ACM Trans. Graph. 2025 |
Computational photography and imaging
high dynamic range imaging |
0.8 | 1 | 2024 | DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024 |
Computational photography and imaging
image relighting |
0.4 | 1 | 2019 | Single image portrait relighting · ACM Trans. Graph. 2019 |
Computational photography and imaging › 3d scanning
volumetric performance capture |
0.4 | 1 | 2019 | The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019 |
Image and video processing › video processing
temporal consistency |
0.3 | 1 | 2025 | Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset · CVPR 2025 |
Computational photography and imaging › illumination estimation
illumination capture |
0.2 | 1 | 2016 | Practical multispectral lighting reproduction · ACM Trans. Graph. 2016 |
Rendering › global illumination
image-based lighting |
0.2 | 1 | 2016 | Practical multispectral lighting reproduction · ACM Trans. Graph. 2016 |
Rendering › novel view synthesis
free-viewpoint rendering |
0.2 | 1 | 2024 | DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024 |
Rendering
image-based rendering |
0.2 | 1 | 2024 | DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024 |
Rendering
novel view synthesis |
0.2 | 1 | 2024 | DifFRelight: Diffusion-Based Facial Performance Relighting · SIGGRAPH Asia 2024 |
Rendering › appearance acquisition
shape and reflectance capture |
0.2 | 1 | 2013 | Acquiring reflectance and shape from continuous spherical harmonic illumination · ACM Trans. Graph. 2013 |
Rendering › lighting
spherical harmonic lighting |
0.2 | 1 | 2013 | Acquiring reflectance and shape from continuous spherical harmonic illumination · ACM Trans. Graph. 2013 |
Computer animation and physical simulation › performance capture
facial performance capture |
0.1 | 1 | 2011 | Multiview face capture using polarized spherical gradient illumination · ACM Trans. Graph. 2011 |
Rendering › appearance acquisition
reflectance map estimation |
0.1 | 1 | 2019 | The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019 |
Virtual and augmented reality › telepresence
3d teleconferencing |
0.1 | 1 | 2009 | Achieving eye contact in a one-to-many 3D video teleconferencing system · ACM Trans. Graph. 2009 |
Rendering
reflectance modeling |
0.0 | 1 | 2013 | Acquiring reflectance and shape from continuous spherical harmonic illumination · ACM Trans. Graph. 2013 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.0 | 1 | 2011 | Multiview face capture using polarized spherical gradient illumination · ACM Trans. Graph. 2011 |
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction |
0.0 | 1 | 2011 | Multiview face capture using polarized spherical gradient illumination · ACM Trans. Graph. 2011 |
Methods — techniques the papers use, named apart from their topics
lighting injection · 1.7conditional video diffusion · 1.7video diffusion model · 1.0masked attention · 1.0data augmentation · 1.0diffusion-based detail enhancement · 0.94d gaussian splatting · 0.9diffusion model · 0.8color gradient illumination · 0.83d gaussian splatting · 0.8photometric stereo · 0.1message passing stereo · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BodyReLux: Temporally Consistent Full-Body Video RelightingabstractBeing able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject-specific video diffusion-based framework for relighting full-body human performances in a temporally consistent way. Our model is trained on a hybrid dataset of pixel-aligned video relighting pairs, covering a diverse combination of lighting conditions, performances and viewpoints. To acquire such dataset, we combine traditional static One-Light-at-a-Time (OLAT) capture and a novel dynamic performance capture in which two smoothly varying lighting sequences are rapidly interleaved. Because the lighting operates above the human flicker-fusion threshold, the interleaving does not appears to strobe. We train our video relighting model from a pretrained text-to-video model to fully leverage the generative priors for producing high quality videos. To achieve accurate lighting control, we introduce a new lighting conditioning method that represents each light source as a token. We further condition on sequences of lighting using masked attention to support dynamic lighting control. Together with a carefully designed data augmentation pipeline, we achieve photorealistic, robust, and temporally consistent video relighting of subject-specific human performances. Mingming He, Xueming Yu, David M. George, Ahmet Levent Tasel, Paul E. Debevec, Julien Philip |
ACM Trans. Graph. | 3 |
| 2025 | Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid DatasetabstractVideo portrait relighting remains challenging because the results need to be both photorealistic and temporally stable. This typically requires a strong model design that can capture complex facial reflections as well as intensive training on a high-quality paired video dataset, such as dynamic one-light-at-a-time (OLAT). In this work, we introduce Lux Post Facto, a novel portrait video relighting method that produces both photorealistic and temporally consistent lighting effects. From the model side, we design a new conditional video diffusion model built upon state-of-the-art pre-trained video diffusion model, alongside a new lighting injection mechanism to enable precise control. This way we leverage strong spatial and temporal generative capability to generate plausible solutions to the ill-posed relighting problem. Our technique uses a hybrid dataset consisting of static expression OLAT data and in-the-wild portrait performance videos to jointly learn relighting and temporal modeling. This avoids the need to acquire paired video data in different lighting conditions. Our extensive experiments show that our model produces state-of-the-art results both in terms of photorealism and temporal consistency. Video results can be found on our project page. Yiqun Mei, Mingming He, Julien Philip, Wenqi Xian, David M. George, Xueming Yu, Gabriel Dedic, Ahmet Levent Tasel, Ning Yu 0006, Vishal M. Patel, Paul E. Debevec |
CVPR | 7 |
| 2025 | Detail Enhanced Gaussian Splatting for Large-Scale Volumetric CaptureabstractWe present a unique system for large-scale, multi-performer, high resolution 4D volumetric capture providing realistic free-viewpoint video up to and including 4K resolution facial closeups. To achieve this, we employ a novel volumetric capture, reconstruction and rendering pipeline based on Dynamic Gaussian Splatting and Diffusion-based Detail Enhancement. We design our pipeline specifically to meet the demands of high-end media production. We employ two capture rigs: the Scene Rig , which captures multi-actor performances at a resolution which falls short of 4K production quality, and the Face Rig , which records high-fidelity single-actor facial detail to serve as a reference for detail enhancement. We first reconstruct dynamic performances from the Scene Rig using 4D Gaussian Splatting, incorporating new model designs and training strategies to improve reconstruction, dynamic range, and rendering quality. Then to render high-quality images for facial closeups, we introduce a diffusion-based detail enhancement model. This model is fine-tuned with high-fidelity data from the same actors recorded in the Face Rig. We train on paired data generated from low- and high-quality Gaussian Splatting (GS) models, using the low-quality input to match the quality of the Scene Rig , with the high-quality GS as ground truth. Our results demonstrate the effectiveness of this pipeline in bridging the gap between the scalable performance capture of a large-scale rig and the high-resolution standards required for film and media production. Julien Philip, Pascal Clausen, Wenqi Xian, Ahmet Levent Tasel, Mingming He, Xueming Yu, David M. George, Ning Yu 0006, Oliver Pilarski, Paul E. Debevec |
ACM Trans. Graph. | 7 |
| 2024 | DifFRelight: Diffusion-Based Facial Performance RelightingabstractWe present a novel framework for free-viewpoint facial performance relighting using diffusion-based image-to-image translation. Leveraging a subject-specific dataset containing diverse facial expressions captured under various lighting conditions, including flat-lit and one-light-at-a-time (OLAT) scenarios, we train a diffusion model for precise lighting control, enabling high-fidelity relit facial images from flat-lit inputs. Our framework includes spatially-aligned conditioning of flat-lit captures and random noise, along with integrated lighting information for global control, utilizing prior knowledge from the pre-trained Stable Diffusion model. This model is then applied to dynamic facial performances captured in a consistent flat-lit environment and reconstructed for novel-view synthesis using a scalable dynamic 3D Gaussian Splatting method to maintain quality and consistency in the relit results. In addition, we introduce unified lighting control by integrating a novel area lighting representation with directional lighting, allowing for joint adjustments in light size and direction. We also enable high dynamic range imaging (HDRI) composition using multiple directional lights to produce dynamic sequences under complex lighting conditions. Our evaluations demonstrate the models efficiency in achieving precise lighting control and generalizing across various facial expressions while preserving detailed features such as skintexture andhair. The model accurately reproduces complex lighting effects like eye reflections, subsurface scattering, self-shadowing, and translucency, advancing photorealism within our framework. Mingming He, Pascal Clausen, Ahmet Levent Tasel, Oliver Pilarski, Wenqi Xian, Laszlo Rikker, Xueming Yu, Ryan D. Burgert, Ning Yu 0006, Paul E. Debevec |
SIGGRAPH Asia | 8 |
| 2019 | The relightables: volumetric performance capture of humans with realistic relightingabstractWe present "The Relightables", a volumetric capture system for photorealistic and high quality relightable full-body performance capture. While significant progress has been made on volumetric capture systems, focusing on 3D geometric reconstruction with high resolution textures, much less work has been done to recover photometric properties needed for relighting. Results from such systems lack high-frequency details and the subject's shading is prebaked into the texture. In contrast, a large body of work has addressed relightable acquisition for image-based approaches, which photograph the subject under a set of basis lighting conditions and recombine the images to show the subject as they would appear in a target lighting environment. However, to date, these approaches have not been adapted for use in the context of a high-resolution volumetric capture system. Our method combines this ability to realistically relight humans for arbitrary environments, with the benefits of free-viewpoint volumetric capture and new levels of geometric accuracy for dynamic performances. Our subjects are recorded inside a custom geodesic sphere outfitted with 331 custom color LED lights, an array of high-resolution cameras, and a set of custom high-resolution depth sensors. Our system innovates in multiple areas: First, we designed a novel active depth sensor to capture 12.4 MP depth maps, which we describe in detail. Second, we show how to design a hybrid geometric and machine learning reconstruction pipeline to process the high resolution input and output a volumetric video. Third, we generate temporally consistent reflectance maps for dynamic performers by leveraging the information contained in two alternating color gradient illumination images acquired at 60Hz. Multiple experiments, comparisons, and applications show that The Relightables significantly improves upon the level of realism in placing volumetrically captured human performances into arbitrary CG scenes. Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Ryan Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor 0001, Paul E. Debevec, Shahram Izadi |
ACM Trans. Graph. | 5 |
| 2019 | Deep reflectance fields: high-quality facial reflectance field inference from color gradient illuminationabstractWe present a novel technique to relight images of human faces by learning a model of facial reflectance from a database of 4D reflectance field data of several subjects in a variety of expressions and viewpoints. Using our learned model, a face can be relit in arbitrary illumination environments using only two original images recorded under spherical color gradient illumination. The output of our deep network indicates that the color gradient images contain the information needed to estimate the full 4D reflectance field, including specular reflections and high frequency details. While capturing spherical color gradient illumination still requires a special lighting setup, reduction to just two illumination conditions allows the technique to be applied to dynamic facial performance capture. We show side-by-side comparisons which demonstrate that the proposed system outperforms the state-of-the-art techniques in both realism and speed. Abhimitra Meka, Christian Häne, Rohit Pandey, Michael Zollhöfer, Sean Ryan Fanello, Graham Fyffe, Adarsh Kowdle, Xueming Yu, Jay Busch, Jason Dourgarian, Peter Denny, Sofien Bouaziz, Peter Lincoln, Matt Whalen, Geoff Harvey, Jonathan Taylor 0001, Shahram Izadi, Andrea Tagliasacchi, Paul E. Debevec, Christian Theobalt, Julien P. C. Valentin, Christoph Rhemann |
ACM Trans. Graph. | 8 |
| 2019 | Single image portrait relightingabstractLighting plays a central role in conveying the essence and depth of the subject in a portrait photograph. Professional photographers will carefully control the lighting in their studio to manipulate the appearance of their subject, while consumer photographers are usually constrained to the illumination of their environment. Though prior works have explored techniques for relighting an image, their utility is usually limited due to requirements of specialized hardware, multiple images of the subject under controlled or known illuminations, or accurate models of geometry and reflectance. To this end, we present a system for portrait relighting : a neural network that takes as input a single RGB image of a portrait taken with a standard cellphone camera in an unconstrained environment, and from that image produces a relit image of that subject as though it were illuminated according to any provided environment map. Our method is trained on a small database of 18 individuals captured under different directional light sources in a controlled light stage setup consisting of a densely sampled sphere of lights. Our proposed technique produces quantitatively superior results on our dataset's validation set compared to prior works, and produces convincing qualitative relighting results on a dataset of hundreds of real-world cellphone portraits. Because our technique can produce a 640 × 640 image in only 160 milliseconds, it may enable interactive user-facing photographic applications in the future. Tiancheng Sun, Jonathan T. Barron, Yun-Ta Tsai, Zexiang Xu, Xueming Yu, Graham Fyffe, Christoph Rhemann, Jay Busch, Paul E. Debevec, Ravi Ramamoorthi |
ACM Trans. Graph. | 5 |
| 2016 | Practical multispectral lighting reproductionabstractWe present a practical framework for reproducing omnidirectional incident illumination conditions with complex spectra using a light stage with multispectral LED lights. For lighting acquisition, we augment standard RGB panoramic photography with one or more observations of a color chart with numerous reflectance spectra. We then solve for how to drive the multispectral light sources so that they best reproduce the appearance of the color charts in the original lighting. Even when solving for non-negative intensities, we show that accurate lighting reproduction is achievable using just four or six distinct LED spectra for a wide range of incident illumination spectra. A significant benefit of our approach is that it does not require the use of specialized equipment (other than the light stage) such as monochromators, spectroradiometers, or explicit knowledge of the LED power spectra, camera spectral response functions, or color chart reflectance spectra. We describe two simple devices for multispectral lighting capture, one for slow measurements of detailed angular spectral detail, and one for fast measurements with coarse angular detail. We validate the approach by realistically compositing real subjects into acquired lighting environments, showing accurate matches to how the subject would actually look within the environments, even for those including complex multispectral illumination. We also demonstrate dynamic lighting capture and playback using the technique. Chloe LeGendre, Xueming Yu, Dai Liu, Jay Busch, Val Jones 0002, Sumanta N. Pattanaik, Paul E. Debevec |
ACM Trans. Graph. | 2 |
| 2013 | Measurement-Based Synthesis of Facial MicrogeometryabstractAbstract We present a technique for generating microstructure‐level facial geometry by augmenting a mesostructure‐level facial scan with detail synthesized from a set of exemplar skin patches scanned at much higher resolution. Additionally, we make point‐source reflectance measurements of the skin patches to characterize the specular reflectance lobes at this smaller scale and analyze facial reflectance variation at both the mesostructure and microstructure scales. We digitize the exemplar patches with a polarization‐based computational illumination technique which considers specular reflection and single scattering. The recorded microstructure patches can be used to synthesize full‐facial microstructure detail for either the same subject or to a different subject. We show that the technique allows for greater realism in facial renderings including more accurate reproduction of skin's specular reflection effects. Paul Graham, Borom Tunwattanapong, Jay Busch, Xueming Yu, Val Jones 0002, Paul E. Debevec, Abhijeet Ghosh |
Comput. Graph. Forum | 4 |
| 2013 | Acquiring reflectance and shape from continuous spherical harmonic illuminationabstractWe present a novel technique for acquiring the geometry and spatially-varying reflectance properties of 3D objects by observing them under continuous spherical harmonic illumination conditions. The technique is general enough to characterize either entirely specular or entirely diffuse materials, or any varying combination across the surface of the object. We employ a novel computational illumination setup consisting of a rotating arc of controllable LEDs which sweep out programmable spheres of incident illumination during 1-second exposures. We illuminate the object with a succession of spherical harmonic illumination conditions, as well as photographed environmental lighting for validation. From the response of the object to the harmonics, we can separate diffuse and specular reflections, estimate world-space diffuse and specular normals, and compute anisotropic roughness parameters for each view of the object. We then use the maps of both diffuse and specular reflectance to form correspondences in a multiview stereo algorithm, which allows even highly specular surfaces to be corresponded across views. The algorithm yields a complete 3D model and a set of merged reflectance maps. We use this technique to digitize the shape and reflectance of a variety of objects difficult to acquire with other techniques and present validation renderings which match well to photographs in similar lighting. Borom Tunwattanapong, Graham Fyffe, Paul Graham, Jay Busch, Xueming Yu, Abhijeet Ghosh, Paul E. Debevec |
ACM Trans. Graph. | 5 |
| 2011 | Single-shot photometric stereo by spectral multiplexingabstractWe propose a novel method for single-shot photometric stereo by spectral multiplexing. The output of our method is a simultaneous per-pixel estimate of the surface normal and full-color reflectance. Our method is well suited to materials with varying color and texture, requires no time-varying illumination, and no high-speed cameras. Being a single-shot method, it may be applied to dynamic scenes without any need for optical flow. Our key contributions are a generalization of three-color photometric stereo to more than three color channels, and the design of a practical six-color-channel system using off-the-shelf parts. Graham Fyffe, Xueming Yu, Paul E. Debevec |
ICCP | 2 |
| 2011 | Multiview face capture using polarized spherical gradient illuminationabstractWe present a novel process for acquiring detailed facial geometry with high resolution diffuse and specular photometric information from multiple viewpoints using polarized spherical gradient illumination. Key to our method is a new pair of linearly polarized lighting patterns which enables multiview diffuse-specular separation under a given spherical illumination condition from just two photographs. The patterns -- one following lines of latitude and one following lines of longitude -- allow the use of fixed linear polarizers in front of the cameras, enabling more efficient acquisition of diffuse and specular albedo and normal maps from multiple viewpoints. In a second step, we employ these albedo and normal maps as input to a novel multi-resolution adaptive domain message passing stereo reconstruction algorithm to create high resolution facial geometry. To do this, we formulate the stereo reconstruction from multiple cameras in a commonly parameterized domain for multiview reconstruction. We show competitive results consisting of high-resolution facial geometry with relightable reflectance maps using five DSLR cameras. Our technique scales well for multiview acquisition without requiring specialized camera systems for sensing multiple polarization states. Abhijeet Ghosh, Graham Fyffe, Borom Tunwattanapong, Jay Busch, Xueming Yu, Paul E. Debevec |
ACM Trans. Graph. | 5 |
| 2009 | Achieving eye contact in a one-to-many 3D video teleconferencing systemabstractWe present a set of algorithms and an associated display system capable of producing correctly rendered eye contact between a three-dimensionally transmitted remote participant and a group of observers in a 3D teleconferencing system. The participant's face is scanned in 3D at 30Hz and transmitted in real time to an autostereoscopic horizontal-parallax 3D display, displaying him or her over more than a 180° field of view observable to multiple observers. To render the geometry with correct perspective, we create a fast vertex shader based on a 6D lookup table for projecting 3D scene vertices to a range of subject angles, heights, and distances. We generalize the projection mathematics to arbitrarily shaped display surfaces, which allows us to employ a curved concave display surface to focus the high speed imagery to individual observers. To achieve two-way eye contact, we capture 2D video from a cross-polarized camera reflected to the position of the virtual participant's eyes, and display this 2D video feed on a large screen in front of the real participant, replicating the viewpoint of their virtual self. To achieve correct vertical perspective, we further leverage this image to track the position of each audience member's eyes, allowing the 3D display to render correct vertical perspective for each of the viewers around the device. The result is a one-to-many 3D teleconferencing system able to reproduce the effects of gaze, attention, and eye contact generally missing in traditional teleconferencing systems. Val Jones 0002, Magnus Lang, Graham Fyffe, Xueming Yu, Jay Busch, Ian McDowall, Mark T. Bolas, Paul E. Debevec |
ACM Trans. Graph. | 4 |