Sunil Hadap

dblp:70/4856 · DBLP profile ↗
← Back
34ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 17 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
18 papers
Computational photography and imaging · 57% Rendering · 20% Visual content generation and editing · 10%
Artificial intelligence
11 papers
3D vision · 54% Generative modeling · 25% Video understanding and tracking · 8%

Topics — the 30 heaviest of 50, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational photography and imaging
illumination estimation
0.932019
All-Weather Deep Outdoor Lighting Estimation · CVPR 2019
Fast Spatially-Varying Indoor Lighting Estimation · CVPR 2019
Automatic Scene Inference for 3D Object Compositing · ACM Trans. Graph. 2014
Computational photography and imaging
image relighting
0.932018
Portrait Lighting Transfer Using a Mass Transport Approach · ACM Trans. Graph. 2018
EyeOpener: Editing Eyes in the Wild · ACM Trans. Graph. 2017
Portrait lighting transfer using a mass transport approach · ACM Trans. Graph. 2017
Computer vision › 3D vision
object pose estimation
0.812024
MRC-Net: 6-DoF Pose Estimation with MultiScale Residual Correlation · CVPR 2024
Rendering › relighting
portrait relighting
0.622019
Deep Single-Image Portrait Relighting · ICCV 2019
High-quality hair modeling from a single portrait photo · ACM Trans. Graph. 2015
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.512021
Self-Supervised Multi-View Person Association and its Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Computer vision › Video understanding and tracking › multi-camera tracking
multi-view multi-human tracking
0.512021
Self-Supervised Multi-View Person Association and its Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Computational photography and imaging
depth estimation
0.522017
Shape Estimation from Shading, Defocus, and Correspondence Using Light-Field Angular Coherence · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Automatic Scene Inference for 3D Object Compositing · ACM Trans. Graph. 2014
Computational photography and imaging
light field imaging
0.522017
Shape Estimation from Shading, Defocus, and Correspondence Using Light-Field Angular Coherence · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Depth from Combining Defocus and Correspondence Using Light-Field Cameras · ICCV 2013
Machine learning › Generative modeling › image generation
GAN-based image generation
0.412019
Deep Single-Image Portrait Relighting · ICCV 2019
Computer vision › 3D vision
image-based rendering
0.412019
Deep view synthesis from sparse photometric images · ACM Trans. Graph. 2019
Computer vision › 3D vision
novel view synthesis
0.412019
Deep view synthesis from sparse photometric images · ACM Trans. Graph. 2019
Computer vision › Segmentation and scene understanding
scene understanding
0.412019
Fast Spatially-Varying Indoor Lighting Estimation · CVPR 2019
Computational photography and imaging › illumination estimation
HDR lighting estimation
0.412019
All-Weather Deep Outdoor Lighting Estimation · CVPR 2019
Computational photography and imaging › illumination estimation
indoor lighting estimation
0.412019
Fast Spatially-Varying Indoor Lighting Estimation · CVPR 2019
Computational photography and imaging › illumination estimation
outdoor illumination estimation
0.412019
All-Weather Deep Outdoor Lighting Estimation · CVPR 2019
Computational photography and imaging › image relighting
single-image relighting
0.412019
Deep Single-Image Portrait Relighting · ICCV 2019
Computer vision › 3D vision
camera calibration
0.312018
A Perceptual Measure for Deep Single Image Camera Calibration · CVPR 2018
Machine learning › Generative modeling › diffusion model
human motion generation
0.312018
MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics · ECCV (5) 2018
Machine learning › Generative modeling
motion generation
0.312018
MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics · ECCV (5) 2018
Computer vision › 3D vision › camera calibration
single image calibration
0.312018
A Perceptual Measure for Deep Single Image Camera Calibration · CVPR 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics · ECCV (5) 2018
Rendering › relighting
image-based relighting
0.312018
Deep image-based relighting from optimal sparse samples · ACM Trans. Graph. 2018
Machine learning › Generative modeling
generative adversarial network
0.312017
Neural Face Editing with Intrinsic Image Disentangling · CVPR 2017
Computer vision › 3D vision › inverse rendering
illumination estimation
0.312017
Deep Outdoor Illumination Estimation · CVPR 2017
Computer vision › 3D vision › inverse rendering
outdoor lighting estimation
0.312017
Deep Outdoor Illumination Estimation · CVPR 2017
Rendering › appearance acquisition › material acquisition
BRDF estimation
0.312017
Reflectance Capture Using Univariate Sampling of BRDFs · ICCV 2017
Visual content generation and editing
face editing
0.312017
Neural Face Editing with Intrinsic Image Disentangling · CVPR 2017
Visual content generation and editing
image editing
0.312017
EyeOpener: Editing Eyes in the Wild · ACM Trans. Graph. 2017
Computational photography and imaging
reflectance acquisition
0.312017
Reflectance Capture Using Univariate Sampling of BRDFs · ICCV 2017
Rendering
reflectance modeling
0.312017
Reflectance Capture Using Univariate Sampling of BRDFs · ICCV 2017

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 3.5spherical harmonics · 0.8soft probabilistic labels · 0.8siamese network · 0.8physically-based rendering · 0.8multi-scale residual correlation · 0.8GAN loss · 0.8mass transport · 0.6color histogram matching · 0.63d morphable face model · 0.6optimization · 0.5tracking-by-clustering · 0.5self-supervised learning · 0.5lalonde-mathews illumination model · 0.4attention · 0.4human perception study · 0.3
YearPublicationVenuePosition
2025 Direct and Explicit 3D Generation from a Single Image
abstract
Current image-to-3D approaches suffer from high computational costs and lack scalability for high-resolution outputs. In contrast, we introduce a novel framework to directly generate explicit surface geometry and texture using multi-view 2D depth and RGB images along with 3D Gaussian features using a repurposed Stable Diffusion model. We introduce a depth branch into U-Net for efficient and high quality multi-view, cross-domain generation and incorporate epipolar attention into the latent-to-pixel decoder for pixel-level multi-view consistency. By back-projecting the generated depth pixels into 3D space, we create a structured 3D representation that can be either rendered via Gaussian splatting or extracted to high-quality meshes, thereby leveraging additional novel view synthesis loss to further improve our performance. Extensive experiments demonstrate that our method surpasses existing baselines in geometry and texture quality while achieving significantly faster generation time.
Meher Gitika Karumuri, Chuhang Zou, Seungbae Bang, Dimitris Samaras, Sunil Hadap
3DV7
2024 MRC-Net: 6-DoF Pose Estimation with MultiScale Residual Correlation
abstract
We propose a single-shot approach to determining 6-DoF pose of an object with available 3D computer-aided design (CAD) model from a single RGB image. Our method, dubbed MRC-Net, comprises two stages. The first performs pose classification and renders the 3D object in the classified pose. The second stage performs regression to predict fine-grained residual pose within class. Connecting the two stages is a novel multi-scale residual correlation (MRC) layer that captures high-and-low level correspondences between the input image and rendering from first stage. MRC-Net employs a Siamese network with shared weights between both stages to learn embeddings for input and rendered images. To mitigate ambiguity when predicting discrete pose class labels on symmetric objects, we use soft probabilistic labels to define pose class in the first stage. We demonstrate state-of-the-art accuracy, outperforming all competing RGB-based methods on four challenging BOP benchmark datasets: T-LESS, LM-O, YCB-V, and ITODD. Our method is non-iterative and requires no complex post-processing. Our code and pretrained models are available at https://github.com/amzn/mrc-net-6d-pose.
Yafei Mao, Raja Bala, Sunil Hadap
CVPR4
2021 Self-Supervised Multi-View Person Association and its Applications
abstract
Reliable markerless motion tracking of people participating in a complex group activity from multiple moving cameras is challenging due to frequent occlusions, strong viewpoint and appearance variations, and asynchronous video streams. To solve this problem, reliable association of the same person across distant viewpoints and temporal instances is essential. We present a self-supervised framework to adapt a generic person appearance descriptor to the unlabeled videos by exploiting motion tracking, mutual exclusion constraints, and multi-view geometry. The adapted discriminative descriptor is used in a tracking-by-clustering formulation. We validate the effectiveness of our descriptor learning on WILDTRACK T. Chavdarova et al., "WILDTRACK: A multi-camera HD dataset for dense unscripted pedestrian detection," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2018, pp. 5030-5039. and three new complex social scenes captured by multiple cameras with up to 60 people "in the wild". We report significant improvement in association accuracy (up to 18 percent) and stable and coherent 3D human skeleton tracking (5 to 10 times) over the baseline. Using the reconstructed 3D skeletons, we cut the input videos into a multi-angle video where the image of a specified person is shown from the best visible front-facing camera. Our algorithm detects inter-human occlusion to determine the camera switching moment while still maintaining the flow of the action well. Website: http://www.cs.cmu.edu/~ILIM/projects/IM/Association4Tracking.
Minh Vo, Ersin Yumer, Kalyan Sunkavalli, Sunil Hadap, Yaser Sheikh, Srinivasa G. Narasimhan
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Learning Monocular Face Reconstruction using Multi-View Supervision
abstract
We present a method to reconstruct faces from a single portrait image. While traditional face reconstruction methods fit low-dimensional 3D morphable models to images, we train a deep network to regress depth from a single image directly. We do so by combining supervised losses on synthetic data with indirect supervision on real data using a novel multi-view photo-consistency loss. Furthermore, we regularize the depth estimation using a 3D morphable model (3DMM). We demonstrate that this leads to results that preserve facial features, capture facial geometry that goes beyond 3DMMs, and is also robust to viewpoint conditions. We evaluate our method on various datasets and via ablation studies, and demonstrate that it outperforms previous work significantly.
Zhixin Shu, Duygu Ceylan, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Dimitris Samaras
FG5
2019 Fast Spatially-Varying Indoor Lighting Estimation
abstract
We propose a real-time method to estimate spatially-varying indoor lighting from a single RGB image. Given an image and a 2D location in that image, our CNN estimates a 5th order spherical harmonic representation of the lighting at the given location in less than 20ms on a laptop mobile graphics card. While existing approaches estimate a single, global lighting representation or require depth as input, our method reasons about local lighting without requiring any geometry information. We demonstrate, through quantitative experiments including a user study, that our results achieve lower lighting estimation errors and are preferred by users over the state-of-the-art. Our approach can be used directly for augmented reality applications, where a virtual object is relit realistically at any position in the scene in real-time.
Mathieu Garon, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr 0001, Jean-François Lalonde
CVPR3
2019 All-Weather Deep Outdoor Lighting Estimation
abstract
We present a neural network that predicts HDR outdoor illumination from a single LDR image. At the heart of our work is a method to accurately learn HDR lighting from LDR panoramas under any weather condition. We achieve this by training another CNN (on a combination of synthetic and real images) to take as input an LDR panorama, and regress the parameters of the Lalonde-Mathews outdoor illumination model. This model is trained such that it a) reconstructs the appearance of the sky, and b) renders the appearance of objects lit by this illumination. We use this network to label a large-scale dataset of LDR panoramas with lighting parameters and use them to train our single image outdoor lighting estimation network. We demonstrate, via extensive experiments, that both our panorama and singe image networks outperform the state of the art, and unlike prior work, are able to handle weather conditions ranging from fully sunny to overcast skies.
Kalyan Sunkavalli, Yannick Hold-Geoffroy, Sunil Hadap, Jonathan Eisenmann, Jean-François Lalonde
CVPR4
2019 Deep Single-Image Portrait Relighting
abstract
Conventional physically-based methods for relighting portrait images need to solve an inverse rendering problem, estimating face geometry, reflectance and lighting. However, the inaccurate estimation of face components can cause strong artifacts in relighting, leading to unsatisfactory results. In this work, we apply a physically-based portrait relighting method to generate a large scale, high quality, “in the wild” portrait relighting dataset (DPR). A deep Convolutional Neural Network (CNN) is then trained using this dataset to generate a relit portrait image by using a source image and a target lighting as input. The training procedure regularizes the generated results, removing the artifacts caused by physically-based relighting methods. A GAN loss is further applied to improve the quality of the relit portrait image. Our trained network can relight portrait images with resolutions as high as 1024 × 1024. We evaluate the proposed method on the proposed DPR datset, Flickr portrait dataset and Multi-PIE dataset both qualitatively and quantitatively. Our experiments demonstrate that the proposed method achieves state-of-the-art results. Please refer to https://zhhoper.github.io/dpr.html for dataset and code.
Hao Zhou 0011, Sunil Hadap, Kalyan Sunkavalli, David Jacobs 0001
ICCV2
2019 Deep view synthesis from sparse photometric images
abstract
The goal of light transport acquisition is to take images from a sparse set of lighting and viewing directions, and combine them to enable arbitrary relighting with changing view. While relighting from sparse images has received significant attention, there has been relatively less progress on view synthesis from a sparse set of "photometric" images---images captured under controlled conditions, lit by a single directional source; we use a spherical gantry to position the camera on a sphere surrounding the object. In this paper, we synthesize novel viewpoints across a wide range of viewing directions (covering a 60° cone) from a sparse set of just six viewing directions. While our approach relates to previous view synthesis and image-based rendering techniques, those methods are usually restricted to much smaller baselines, and are captured under environment illumination. At our baselines, input images have few correspondences and large occlusions; however we benefit from structured photometric images. Our method is based on a deep convolutional network trained to directly synthesize new views from the six input views. This network combines 3D convolutions on a plane sweep volume with a novel per-view per-depth plane attention map prediction network to effectively aggregate multi-view appearance. We train our network with a large-scale synthetic dataset of 1000 scenes with complex geometry and material properties. In practice, it is able to synthesize novel viewpoints for captured real data and reproduces complex appearance effects like occlusions, view-dependent specularities and hard shadows. Moreover, the method can also be combined with previous relighting techniques to enable changing both lighting and view, and applied to computer vision problems like multiview stereo from sparse image sets.
Zexiang Xu, Sai Bi, Kalyan Sunkavalli, Sunil Hadap, Hao Su 0001, Ravi Ramamoorthi
ACM Trans. Graph.4
2018 A Perceptual Measure for Deep Single Image Camera Calibration
abstract
Most current single image camera calibration methods rely on specific image features or user input, and cannot be applied to natural images captured in uncontrolled settings. We propose directly inferring camera calibration parameters from a single image using a deep convolutional neural network. This network is trained using automatically generated samples from a large-scale panorama dataset, and considerably outperforms other methods, including recent deep learning-based approaches, in terms of standard L2 error. However, we argue that in many cases it is more important to consider how humans perceive errors in camera estimation. To this end, we conduct a large-scale human perception study where we ask users to judge the realism of 3D objects composited with and without ground truth camera calibration. Based on this study, we develop a new perceptual measure for camera calibration, and demonstrate that our deep calibration network outperforms other methods on this measure. Finally, we demonstrate the use of our calibration network for a number of applications including virtual object insertion, image retrieval and compositing.
Yannick Hold-Geoffroy, Kalyan Sunkavalli, Jonathan Eisenmann, Matthew Fisher, Emiliano Gambaretto, Sunil Hadap, Jean-François Lalonde
CVPR6
2018 Illuminant Spectra-Based Source Separation Using Flash Photography
abstract
Real-world lighting often consists of multiple illuminants with different spectra. Separating and manipulating these illuminants in post-process is a challenging problem that requires either significant manual input or calibrated scene geometry and lighting. In this work, we leverage a flash/no-flash image pair to analyze and edit scene illuminants based on their spectral differences. We derive a novel physics-based relationship between color variations in the observed flash/no-flash intensities and the spectra and surface shading corresponding to individual scene illuminants. Our technique uses this constraint to automatically separate an image into constituent images lit by each illuminant. This separation can be used to support applications like white balancing, lighting editing, and RGB photometric stereo, where we demonstrate results that outperform state-of-the-art techniques on a wide range of images.
Zhuo Hui, Kalyan Sunkavalli, Sunil Hadap, Aswin C. Sankaranarayanan
CVPR3
2018 MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
Xinchen Yan, Akash Rastogi, Ruben Villegas, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Ersin Yumer, Honglak Lee
ECCV (5)6
2018 Portrait Lighting Transfer Using a Mass Transport Approach
abstract
Lighting is a critical element of portrait photography. However, good lighting design typically requires complex equipment and significant time and expertise. Our work simplifies this task using a relighting technique that transfers the desired illumination of one portrait onto another. The novelty in our approach to this challenging problem is our formulation of relighting as a mass transport problem. We start from standard color histogram matching that only captures the overall tone of the illumination, and we show how to use the mass-transport formulation to make it dependent on facial geometry. We fit a three-dimensional (3D) morphable face model to the portrait, and for each pixel, we combine the color value with the corresponding 3D position and normal. We then solve a mass-transport problem in this augmented space to generate a color remapping that achieves localized, geometry-aware relighting. Our technique is robust to variations in facial appearance and small errors in face reconstruction. As we demonstrate, this allows our technique to handle a variety of portraits and illumination conditions, including scenarios that are challenging for previous methods.
Zhixin Shu, Sunil Hadap, Eli Shechtman, Kalyan Sunkavalli, Sylvain Paris, Dimitris Samaras
ACM Trans. Graph.2
2018 Deep image-based relighting from optimal sparse samples
abstract
We present an image-based relighting method that can synthesize scene appearance under novel, distant illumination from the visible hemisphere, from only five images captured under pre-defined directional lights. Our method uses a deep convolutional neural network to regress the relit image from these five images; this relighting network is trained on a large synthetic dataset comprised of procedurally generated shapes with real-world reflectances. We show that by combining a custom-designed sampling network with the relighting network, we can jointly learn both the optimal input light directions and the relighting function. We present an extensive evaluation of our network, including an empirical analysis of reconstruction quality, optimal lighting configurations for different scenarios, and alternative network architectures. We demonstrate, on both synthetic and real scenes, that our method is able to reproduce complex, high-frequency lighting effects like specularities and cast shadows, and outperforms other image-based relighting methods that require an order of magnitude more images.
Zexiang Xu, Kalyan Sunkavalli, Sunil Hadap, Ravi Ramamoorthi
ACM Trans. Graph.3
2017 Deep Outdoor Illumination Estimation
abstract
We present a convolutional neural network-based (CNN-based) technique to estimate high-dynamic range outdoor illumination from a single low dynamic range image. To train the CNN, we leverage a large dataset of outdoor panoramas. We fit a low-dimensional physically-based outdoor illumination model to the skies in these panoramas giving us a compact set of parameters (including sun position, atmospheric conditions, and camera parameters). We extract limited field-of-view images from the panoramas, and train a CNN with this large set of input image–output lighting parameter pairs. Given a test image, this network can be used to infer illumination parameters that can, in turn, be used to reconstruct an outdoor illumination environment map. We demonstrate that our approach allows the recovery of plausible illumination conditions and enables photorealistic virtual object insertion from a single image. An extensive evaluation on both the panorama dataset and captured HDR environment maps shows that our technique significantly outperforms previous solutions to this problem.
Yannick Hold-Geoffroy, Kalyan Sunkavalli, Sunil Hadap, Emiliano Gambaretto, Jean-François Lalonde
CVPR3
2017 Neural Face Editing with Intrinsic Image Disentangling
abstract
Traditional face editing methods often require a number of sophisticated and task specific algorithms to be applied one after the other - a process that is tedious, fragile, and computationally intensive. In this paper, we propose an end-to-end generative adversarial network that infers a face-specific disentangled representation of intrinsic face properties, including shape (i.e. normals), albedo, and lighting, and an alpha matte. We show that this network can be trained on “in-the-wild” images by incorporating an in-network physically-based image formation module and appropriate loss functions. Our disentangling latent representation allows for semantically relevant edits, where one aspect offacial appearance can be manipulated while keeping orthogonal properties fixed, and we demonstrate its use for a number offacial editing applications.
Zhixin Shu, Ersin Yumer, Sunil Hadap, Kalyan Sunkavalli, Eli Shechtman, Dimitris Samaras
CVPR3
2017 Reflectance Capture Using Univariate Sampling of BRDFs
abstract
We propose the use of a light-weight setup consisting of a collocated camera and light source – commonly found on mobile devices – to reconstruct surface normals and spatially-varying BRDFs of near-planar material samples. A collocated setup provides only a 1-D “univariate” sampling of a 3-D isotropic BRDF. We show that a univariate sampling is sufficient to estimate parameters of commonly used analytical BRDF models. Subsequently, we use a dictionary-based reflectance prior to derive a robust technique for per-pixel normal and BRDF estimation. We demonstrate real-world shape and capture, and its application to material editing and classification, using real data acquired using a mobile phone.
Zhuo Hui, Kalyan Sunkavalli, Joon-Young Lee, Sunil Hadap, Jian Wang 0100, Aswin C. Sankaranarayanan
ICCV4
2017 Real-Time Oil Painting on Mobile Hardware
abstract
Abstract This paper presents a realistic digital oil painting system, specifically targeted at the real‐time performance on highly resource‐constrained portable hardware such as tablets and iPads. To effectively use the limited computing power, we develop an efficient adaptation of the shallow water equations that models all the characteristic properties of oil paint. The pigments are stored in a multi‐layered structure to model the peculiar nature of pigment mixing in oil paint. The user experience ranges from thick shape‐retaining strokes to runny diluted paint that reacts naturally to the gravity set by tablet orientation. Finally, the paint is rendered in real time using a combination of carefully chosen efficient rendering techniques. The virtual lighting adapts to the tablet orientation, or alternatively, the front‐facing camera captures the lighting environment, which leads to a truly immersive user experience. Our proposed features are evaluated via a user study. In our experience, our system enables artists to quickly try out ideas and compositions anywhere when inspiration strikes, in a truly ubiquitous way. They do not need to carry expensive and messy oil paint supplies.
Tuur Stuyck, Fang Da, Sunil Hadap, Philip Dutré
Comput. Graph. Forum3
2017 Shape Estimation from Shading, Defocus, and Correspondence Using Light-Field Angular Coherence
abstract
Light-field cameras are quickly becoming commodity items, with consumer and industrial applications. They capture many nearby views simultaneously using a single image with a micro-lens array, thereby providing a wealth of cues for depth recovery: defocus, correspondence, and shading. In particular, apart from conventional image shading, one can refocus images after acquisition, and shift one's viewpoint within the sub-apertures of the main lens, effectively obtaining multiple views. We present a principled algorithm for dense depth estimation that combines defocus and correspondence metrics. We then extend our analysis to the additional cue of shading, using it to refine fine details in the shape. By exploiting an all-in-focus image, in which pixels are expected to exhibit angular coherence, we define an optimization framework that integrates photo consistency, depth consistency, and shading consistency. We show that combining all three sources of information: defocus, correspondence, and shading, outperforms state-of-the-art light-field depth estimation algorithms in multiple scenarios.
Michael W. Tao, Pratul P. Srinivasan, Sunil Hadap, Szymon Rusinkiewicz, Jitendra Malik, Ravi Ramamoorthi
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Portrait lighting transfer using a mass transport approach
abstract
Lighting is a critical element of portrait photography. However, good lighting design typically requires complex equipment and significant time and expertise. Our work simplifies this task using a relighting technique that transfers the desired illumination of one portrait onto another. The novelty in our approach to this challenging problem is our formulation of relighting as a mass transport problem. We start from standard color histogram matching that only captures the overall tone of the illumination, and we show how to use the mass-transport formulation to make it dependent on facial geometry. We fit a three-dimensional (3D) morphable face model to the portrait, and for each pixel, we combine the color value with the corresponding 3D position and normal. We then solve a mass-transport problem in this augmented space to generate a color remapping that achieves localized, geometry-aware relighting. Our technique is robust to variations in facial appearance and small errors in face reconstruction. As we demonstrate, this allows our technique to handle a variety of portraits and illumination conditions, including scenarios that are challenging for previous methods.
Zhixin Shu, Sunil Hadap, Eli Shechtman, Kalyan Sunkavalli, Sylvain Paris, Dimitris Samaras
ACM Trans. Graph.2
2017 EyeOpener: Editing Eyes in the Wild
abstract
Closed eyes and look-aways can ruin precious moments captured in photographs. In this article, we present a new framework for automatically editing eyes in photographs. We leverage a user’s personal photo collection to find a “good” set of reference eyes and transfer them onto a target image. Our example-based editing approach is robust and effective for realistic image editing. A fully automatic pipeline for realistic eye editing is challenging due to the unconstrained conditions under which the face appears in a typical photo collection. We use crowd-sourced human evaluations to understand the aspects of the target-reference image pair that will produce the most realistic results. We subsequently train a model that automatically selects the top-ranked reference candidate(s) by narrowing the gap in terms of pose, local contrast, lighting conditions, and even expressions. Finally, we develop a comprehensive pipeline of three-dimensional face estimation, image warping, relighting, image harmonization, automatic segmentation, and image compositing in order to achieve highly believable results. We evaluate the performance of our method via quantitative and crowd-sourced experiments.
Zhixin Shu, Eli Shechtman, Dimitris Samaras, Sunil Hadap
ACM Trans. Graph.4
2016 Shading-Aware Multi-view Stereo
Fabian Langguth, Kalyan Sunkavalli, Sunil Hadap, Michael Goesele
ECCV (3)3
2016 White balance under mixed illumination using flash photography
abstract
Real-world illumination is often a complex spatially-varying combination of multiple illuminants. In this work, we present a technique to white-balance images captured in such illumination by leveraging flash photography. Even though this problem is severely ill-posed, we show that using two images — captured with and without flash lighting — leads to a closed form solution for spatially-varying mixed illumination. Our solution is completely automatic and makes no assumptions about the number or nature of the illuminants. We also propose an extension of our scheme to handle practical challenges such as shadows, specularities, as well as the camera and scene motion. We evaluate our technique on datasets captured in both the laboratory and the real-world, and show that it significantly outperforms a number of previous white balance algorithms.
Zhuo Hui, Aswin C. Sankaranarayanan, Kalyan Sunkavalli, Sunil Hadap
ICCP4
2015 High-quality hair modeling from a single portrait photo
abstract
We propose a novel system to reconstruct a high-quality hair depth map from a single portrait photo with minimal user input. We achieve this by combining depth cues such as occlusions, silhouettes, and shading, with a novel 3D helical structural prior for hair reconstruction. We fit a parametric morphable face model to the input photo and construct a base shape in the face, hair and body regions using occlusion and silhouette constraints. We then estimate the normals in the hair region via a Shape-from-Shading-based optimization that uses the lighting inferred from the face model and enforces an adaptive albedo prior that models the typical color and occlusion variations of hair. We introduce a 3D helical hair prior that captures the geometric structure of hair, and show that it can be robustly recovered from the input photo in an automatic manner. Our system combines the base shape, the normals estimated by Shape from Shading, and the 3D helical hair prior to reconstruct high-quality 3D hair models. Our single-image reconstruction closely matches the results of a state-of-the-art multi-view stereo applied on a multi-view stereo dataset. Our technique can reconstruct a wide variety of hairstyles ranging from short to long and from straight to messy, and we demonstrate the use of our 3D hair models for high-quality portrait relighting, novel view synthesis and 3D-printed portrait reliefs.
Menglei Chai, Linjie Luo, Kalyan Sunkavalli, Nathan Carr 0001, Sunil Hadap, Kun Zhou 0001
ACM Trans. Graph.5
2014 Automatic Scene Inference for 3D Object Compositing
abstract
We present a user-friendly image editing system that supports a drag-and-drop object insertion (where the user merely drags objects into the image, and the system automatically places them in 3D and relights them appropriately), postprocess illumination editing, and depth-of-field manipulation. Underlying our system is a fully automatic technique for recovering a comprehensive 3D scene model (geometry, illumination, diffuse albedo, and camera parameters) from a single, low dynamic range photograph. This is made possible by two novel contributions: an illumination inference algorithm that recovers a full lighting model of the scene (including light sources that are not directly visible in the photograph), and a depth estimation algorithm that combines data-driven depth transfer with geometric reasoning about the scene layout. A user study shows that our system produces perceptually convincing results, and achieves the same level of realism as techniques that require significant user interaction.
Kevin Karsch, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr 0001, Hailin Jin, Rafael Fonte, Michael Sittig, David A. Forsyth
ACM Trans. Graph.3
2013 Specular Reflection Separation Using Dark Channel Prior
abstract
We present a novel method to separate specular reflection from a single image. Separating an image into diffuse and specular components is an ill-posed problem due to lack of observations. Existing methods rely on a specular-free image to detect and estimate specularity, which however may confuse diffuse pixels with the same hue but a different saturation value as specular pixels. Our method is based on a novel observation that for most natural images the dark channel can provide an approximate specular-free image. We also propose a maximum a posteriori formulation which robustly recovers the specular reflection and chromaticity despite of the hue-saturation ambiguity. We demonstrate the effectiveness of the proposed algorithm on real and synthetic examples. Experimental results show that our method significantly outperforms the state-of-the-art methods in separating specular reflection.
Hyeongwoo Kim, Hailin Jin, Sunil Hadap, In-So Kweon
CVPR3
2013 Depth from Combining Defocus and Correspondence Using Light-Field Cameras
abstract
Light-field cameras have recently become available to the consumer market. An array of micro-lenses captures enough information that one can refocus images after acquisition, as well as shift one's viewpoint within the sub-apertures of the main lens, effectively obtaining multiple views. Thus, depth cues from both defocus and correspondence are available simultaneously in a single capture. Previously, defocus could be achieved only through multiple image exposures focused at different depths, while correspondence cues needed multiple exposures at different viewpoints or multiple cameras, moreover, both cues could not easily be obtained together. In this paper, we present a novel simple and principled algorithm that computes dense depth estimation by combining both defocus and correspondence depth cues. We analyze the x-u 2D epipolar image (EPI), where by convention we assume the spatial x coordinate is horizontal and the angular u coordinate is vertical (our final algorithm uses the full 4D EPI). We show that defocus depth cues are obtained by computing the horizontal (spatial) variance after vertical (angular) integration, and correspondence depth cues by computing the vertical (angular) variance. We then show how to combine the two cues into a high quality depth map, suitable for computer vision applications such as matting, full control of depth-of-field, and surface reconstruction.
Michael W. Tao, Sunil Hadap, Jitendra Malik, Ravi Ramamoorthi
ICCV2
2013 Multiple Light Source Estimation in a Single Image
abstract
Abstract Many high‐level image processing tasks require an estimate of the positions, directions and relative intensities of the light sources that illuminated the depicted scene. In image‐based rendering, augmented reality and computer vision, such tasks include matching image contents based on illumination, inserting rendered synthetic objects into a natural image, intrinsic images, shape from shading and image relighting. Yet, accurate and robust illumination estimation, particularly from a single image, is a highly ill‐posed problem. In this paper, we present a new method to estimate the illumination in a single image as a combination of achromatic lights with their 3D directions and relative intensities. In contrast to previous methods, we base our azimuth angle estimation on curve fitting and recursive refinement of the number of light sources. Similarly, we present a novel surface normal approximation using an osculating arc for the estimation of zenith angles. By means of a new data set of ground‐truth data and images, we demonstrate that our approach produces more robust and accurate results, and show its versatility through novel applications such as image compositing and analysis.
Jorge Lopez-Moreno, Elena Garces 0001, Sunil Hadap, Erik Reinhard, Diego Gutierrez
Comput. Graph. Forum3
2012 Reconstructing Shape from Dictionaries of Shading Primitives
Alexandros Panagopoulos, Sunil Hadap, Dimitris Samaras
ACCV (4)2
2011 Non-photorealistic, depth-based image editing
Jorge Lopez-Moreno, Jorge Jimenez, Sunil Hadap, Ken Anjyo, Erik Reinhard, Diego Gutierrez
Comput. Graph.3
2010 Industrial-strength painting with a virtual bristle brush
abstract
Research in natural media painting has produced impressive images, but those results have not been adopted by commercial applications to date because of the heavy demands of industrial painting workflows. In this paper, we present a new 3D brush model with associated algorithms for stroke generation and bidirectional paint transfer that is suitable for professional use. Our model can reproduce arbitrary brush tip shapes and can be used to generate raster or vector output, none of which was possible in previous simulations. This is achieved by an efficient formulation of bristle behaviors as strand dynamics in a non-inertial reference frame. To demonstrate the robustness and flexibility of our approach, we have integrated our model into major commercial painting and vector editing applications and given it to professional artists to evaluate.
Stephen DiVerdi, Aravind Krishnaswamy, Sunil Hadap
VRST3
2010 Compositing images through light source detection
Jorge Lopez-Moreno, Sunil Hadap, Erik Reinhard, Diego Gutierrez
Comput. Graph.2
2008 Regularized depth from defocus
abstract
In the area of depth estimation from images an interesting approach has been structure recovery from defocus cue. Towards this end, there have been a number of approaches [4,6]. Here we propose a technique to estimate the regularized depth from defocus using diffusion. The coefficient of the diffusion equation is modeled using a pair-wise Markov random field (MRF) ensuring spatial regularization to enhance the robustness of the depth estimated. This framework is solved efficiently using a graph-cuts based techniques. The MRF representation is enhanced by incorporating a smoothness prior that is obtained from a graph based segmentation of the input images. The method is demonstrated on a number of data sets and its performance is compared with state of the art techniques.
Vinay P. Namboodiri, Subhasis Chaudhuri, Sunil Hadap
ICIP3
2001 Modeling Dynamic Hair as a Continuum
abstract
In this paper we address the difficult problem of hair dynamics, particularly hair-hair and hair-air interactions. To model these interactions, we propose to consider hair volume as a continuum. Subsequently, we treat the interaction dynamics to be fluid dynamics. This proves to be a strong as well as viable approach for an otherwise very complex phenomenon. However, we retain the individual character of hair, which is vital to visually realistic rendering of hair animation. For that, we develop an elaborate model for stiffness and inertial dynamics of individual hair strand. Being a reduced coordinate formulation, the stiffness dynamics is numerically stable and fast. We then unify the continuum interaction dynamics and the individual hair's stiffness dynamics.
Sunil Hadap, Nadia Magnenat-Thalmann
Comput. Graph. Forum1
1999 Animating Wrinkles on Clothes
abstract
This paper describes a method to simulate realistic wrinkles on clothes without fine mesh and large computational overheads. Cloth has very little in-plane deformations, as most of the deformations come from buckling. This can be looked at as area conservation property of cloth. The area conservation formulation of the method modulates the user defined wrinkle pattern, based on deformation of individual triangle. The methodology facilitates use of small in-plane deformation stiffnesses and a coarse mesh for the numerical simulation, this makes cloth simulation fast and robust. Moreover, the ability to design wrinkles (even on generalized deformable models) makes this method versatile for synthetic image generation. The method inspired from cloth wrinkling problem, being geometric in nature, can be extended to other wrinkling phenomena.
Sunil Hadap, Endre Bangerter, Pascal Volino, Nadia Magnenat-Thalmann
IEEE Visualization1