VLDB 2026 Research / reviewers in the wild / expert
Nima Khademi Kalantari
dblp:03/4517 · also Nima K. Kalantari, Nima Kalantari
· DBLP profile ↗
46ranked-venue papers
16as first author
17since 2021 · last 2025
0000-0002-2588-9219ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 15 first-author · 17 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RI3D: Few-Shot Gaussian Splatting with Repair and Inpainting Diffusion PriorsabstractIn this paper, we propose RI3D, a novel 3DGS-based approach that harnesses the power of diffusion models to reconstruct high-quality novel views given a sparse set of input images. Our key contribution is separating the view synthesis process into two tasks of reconstructing visible regions and hallucinating missing regions, and introducing two personalized diffusion models, each tailored to one of these tasks. Specifically, one model ('repair') takes a rendered image as input and predicts the corresponding high-quality image, which in turn is used as a pseudo ground truth image to constrain the optimization. The other model ('inpainting') primarily focuses on hallucinating details in unobserved areas. To integrate these models effectively, we introduce a two-stage optimization strategy: the first stage reconstructs visible areas using the repair model, and the second stage reconstructs missing regions with the inpainting model while ensuring coherence through further optimization. Moreover, we augment the optimization with a novel Gaussian initialization method that obtains per-image depth by combining 3D-consistent and smooth depth with highly detailed relative depth. We demonstrate that by separating the process into two tasks and addressing them with the repair and inpainting models, we produce results with detailed textures in both visible and missing regions that outperform state-of-the-art approaches on a diverse set of scenes with extremely sparse inputs. Avinash Paliwal, Xilong Zhou 0001, Wei Ye 0008, Jinhui Xiong, Nima Khademi Kalantari |
ICCV | 6 |
| 2025 | PanoDreamer: Optimization-Based Single Image to 360 3D Scene With DiffusionabstractIn this paper, we present PanoDreamer, a novel method for producing a coherent 360° 3D scene from a single input image. Unlike existing methods that generate the scene sequentially, we frame the problem as single-image panorama and depth estimation. Once the coherent panoramic image and its corresponding depth are obtained, the scene can be reconstructed by inpainting the small occluded regions and projecting them into 3D space. Our key contribution is formulating single-image panorama and depth estimation as two optimization tasks and introducing alternating minimization strategies to effectively solve their objectives. We demonstrate that our approach outperforms existing techniques in single-image 360° 3D scene reconstruction in terms of consistency and overall quality. Avinash Paliwal, Xilong Zhou 0001, Andrii Tsarov, Nima Khademi Kalantari |
SIGGRAPH Asia | 4 |
| 2025 | Analyzing and Improving the Skin Tone Consistency and Bias in Implicit 3D Relightable Face GeneratorsabstractWith the advances in generative adversarial networks (GANs) and neural rendering, 3D relightable face generation has received significant attention. Among the existing methods, a particularly successful technique uses an implicit lighting representation and generates relit images through the product of synthesized albedo and light-dependent shading images. While this approach produces high-quality results with intricate shading details, it often has difficulty producing relit images with consistent skin tones, particularly when the lighting condition is extracted from images of individuals with dark skin. Additionally, this technique is biased towards producing albedo images with lighter skin tones. Our main observation is that this problem is rooted in the biased spherical harmonics (SH) coefficients, used during training. Following this observation, we conduct an analysis and demonstrate that the bias appears not only in band 0 (DC term), but also in the other bands of the estimated SH coefficients. We then propose a simple, but effective, strategy to mitigate the problem. Specifically, we normalize the SH coefficients by their DC term to eliminate the inherent magnitude bias, while statistically align the coefficients in the other bands to alleviate the directional bias. We also propose a scaling strategy to match the distribution of illumination magnitude in the generated images with the training data. Through extensive experiments, we demonstrate the effectiveness of our solution in increasing the skin tone consistency and mitigating bias. Libing Zeng, Nima Khademi Kalantari |
WACV | 2 |
| 2024 | CoherentGS: Sparse Novel View Synthesis with Coherent 3D Gaussians
Avinash Paliwal, Wei Ye 0008, Jinhui Xiong, Dmytro Kotovenko, Vikas Chandra, Nima Khademi Kalantari |
ECCV (30) | 7 |
| 2023 | Implicit View-Time Interpolation of Stereo Videos Using Multi-Plane Disparities and Non-Uniform CoordinatesabstractIn this paper, we propose an approach for view-time interpolation of stereo videos. Specifically, we build upon X-Fields that approximates an interpolatable mapping between the input coordinates and 2D RGB images using a convolutional decoder. Our main contribution is to analyze and identify the sources of the problems with using X-Fields in our application and propose novel techniques to overcome these challenges. Specifically, we observe that X-Fields struggles to implicitly interpolate the disparities for large baseline cameras. Therefore, we propose multi-plane disparities to reduce the spatial distance of the objects in the stereo views. Moreover, we propose non-uniform time coordinates to handle the non-linear and sudden motion spikes in videos. We additionally introduce several simple, but important, improvements over X-Fields. We demonstrate that our approach is able to produce better results than the state of the art, while running in near real-time rates and having low memory and storage costs. Avinash Paliwal, Andrii Tsarov, Nima Khademi Kalantari |
CVPR | 3 |
| 2023 | 3D-aware Facial Landmark Detection via Multi-view Consistent Training on Synthetic DataabstractAccurate facial landmark detection on wild images plays an essential role in human-computer interaction, entertainment, and medical applications. Existing approaches have limitations in enforcing 3D consistency while detecting 3D/2D facial landmarks due to the lack of multi-view in-the-wild training data. Fortunately, with the recent advances in generative visual models and neural rendering, we have witnessed rapid progress towards high quality 3D image synthesis. In this work, we leverage such approaches to construct a synthetic dataset and propose a novel multi-view consistent learning strategy to improve 3D facial landmark detection accuracy on in-the-wild images. The proposed 3D-aware module can be plugged into any learning-based landmark detection algorithm to enhance its accuracy. We demonstrate the superiority of the proposed plug-in module with extensive comparison against state-of-the-art methods on several real and synthetic datasets. Libing Zeng, Wentao Bao, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Nima Khademi Kalantari |
CVPR | 7 |
| 2023 | MyStyle++: A Controllable Personalized Generative PriorabstractIn this paper, we propose an approach to obtain a personalized generative prior with explicit control over a set of attributes. We build upon MyStyle, a recently introduced method, that tunes the weights of a pre-trained StyleGAN face generator on a few images of an individual. This system allows synthesizing, editing, and enhancing images of the target individual with high fidelity to their facial features. However, MyStyle does not demonstrate precise control over the attributes of the generated images. We propose to address this problem through a novel optimization system that organizes the latent space in addition to tuning the generator. Our key contribution is to formulate a loss that arranges the latent codes, corresponding to the input images, along a set of specific directions according to their attributes. We demonstrate that our approach, dubbed MyStyle++, is able to synthesize, edit, and enhance images of an individual with great control over the attributes, while preserving the unique facial characteristics of that individual. Libing Zeng, Yi Xu 0002, Nima Khademi Kalantari |
SIGGRAPH Asia | 4 |
| 2023 | Frame Interpolation for Dynamic Scenes with Implicit Flow EncodingabstractIn this paper, we propose an algorithm to interpolate between a pair of images of a dynamic scene. While in the past years significant progress in frame interpolation has been made, current approaches are not able to handle images with brightness and illumination changes, which are common even when the images are captured shortly apart. We propose to address this problem by taking advantage of the existing optical flow methods that are highly robust to the variations in the illumination. Specifically, using the bidirectional flows estimated using an existing pre-trained flow network, we predict the flows from an intermediate frame to the two input images. To do this, we propose to encode the bidirectional flows into a coordinate-based network, powered by a hypernetwork, to obtain a continuous representation of the flow across time. Once we obtain the estimated flows, we use them within an existing blending network to obtain the final intermediate frame. Through extensive experiments, we demonstrate that our approach is able to produce significantly better results than state-of-the-art frame interpolation algorithms. Pedro Figueirêdo, Avinash Paliwal, Nima Khademi Kalantari |
WACV | 3 |
| 2023 | Variational Pose Prediction with Dynamic Sample Selection from Sparse Tracking SignalsabstractAbstract We propose a learning‐based approach for full‐body pose reconstruction from extremely sparse upper body tracking data, obtained from a virtual reality (VR) device. We leverage a conditional variational autoencoder with gated recurrent units to synthesize plausible and temporally coherent motions from 4‐point tracking (head, hands, and waist positions and orientations). To avoid synthesizing implausible poses, we propose a novel sample selection and interpolation strategy along with an anomaly detection algorithm. Specifically, we monitor the quality of our generated poses using the anomaly detection algorithm and smoothly transition to better samples when the quality falls below a statistically defined threshold. Moreover, we demonstrate that our sample selection and interpolation method can be used for other applications, such as target hitting and collision avoidance, where the generated motions should adhere to the constraints of the virtual environment. Our system is lightweight, operates in real‐time, and is able to produce temporally coherent and realistic motions. Nicholas Milef, Shinjiro Sueda, Nima Khademi Kalantari |
Comput. Graph. Forum | 3 |
| 2023 | Test-Time Optimization for Video Depth Estimation Using Pseudo Reference DepthabstractAbstract In this paper, we propose a learning‐based test‐time optimization approach for reconstructing geometrically consistent depth maps from a monocular video. Specifically, we optimize an existing single image depth estimation network on the test example at hand. We do so by introducing pseudo reference depth maps which are computed based on the observation that the optical flow displacement for an image pair should be consistent with the displacement obtained by depth‐reprojection. Additionally, we discard inaccurate pseudo reference depth maps using a simple median strategy and propose a way to compute a confidence map for the reference depth. We use our pseudo reference depth and the confidence map to formulate a loss function for performing the test‐time optimization in an efficient and effective manner. We compare our approach against the state‐of‐the‐art methods on various scenes both visually and numerically. Our approach is on average 2.5× faster than the state of the art and produces depth maps with higher quality. Libing Zeng, Nima Khademi Kalantari |
Comput. Graph. Forum | 2 |
| 2023 | A Semi-Procedural Convolutional Material PriorabstractAbstract Lightweight material capture methods require a material prior, defining the subspace of plausible textures within the large space of unconstrained texel grids. Previous work has either used deep neural networks (trained on large synthetic material datasets) or procedural node graphs (constructed by expert artists) as such priors. In this paper, we propose a semi‐procedural differentiable material prior that represents materials as a set of (typically procedural) grayscale noises and patterns that are processed by a sequence of lightweight learnable convolutional filter operations. We demonstrate that the restricted structure of this architecture acts as an inductive bias on the space of material appearances, allowing us to optimize the weights of the convolutions per‐material, with no need for pre‐training on a large dataset. Combined with a differentiable rendering step and a perceptual loss, we enable single‐image tileable material capture comparable with state of the art. Our approach does not target the pixel‐perfect recovery of the material, but rather uses noises and patterns as input to match the target appearance. To achieve this, it does not require complex procedural graphs, and has a much lower complexity, computational cost and storage cost. We also enable control over the results, through changing the provided patterns and using guide maps to push the material properties towards a user‐driven objective. Xilong Zhou 0001, Milos Hasan, Valentin Deschaintre, Paul Guerrero 0001, Kalyan Sunkavalli, Nima Khademi Kalantari |
Comput. Graph. Forum | 6 |
| 2023 | ReShader: View-Dependent Highlights for Single Image View-SynthesisabstractIn recent years, novel view synthesis from a single image has seen significant progress thanks to the rapid advancements in 3D scene representation and image inpainting techniques. While the current approaches are able to synthesize geometrically consistent novel views, they often do not handle the view-dependent effects properly. Specifically, the highlights in their synthesized images usually appear to be glued to the surfaces, making the novel views unrealistic. To address this major problem, we make a key observation that the process of synthesizing novel views requires changing the shading of the pixels based on the novel camera, and moving them to appropriate locations. Therefore, we propose to split the view synthesis process into two independent tasks of pixel reshading and relocation. During the reshading process, we take the single image as the input and adjust its shading based on the novel camera. This reshaded image is then used as the input to an existing view synthesis method to relocate the pixels and produce the final novel view image. We propose to use a neural network to perform reshading and generate a large set of synthetic input-reshaded pairs to train our network. We demonstrate that our approach produces plausible novel view images with realistic moving highlights on a variety of real world scenes. Avinash Paliwal, Brandon G. Nguyen, Andrii Tsarov, Nima Khademi Kalantari |
ACM Trans. Graph. | 4 |
| 2022 | TileGen: Tileable, Controllable Material Generation and CaptureabstractRecent methods (e.g. MaterialGAN) have used unconditional GANs to generate per-pixel material maps, or as a prior to reconstruct materials from input photographs. These models can generate varied random material appearance, but do not have any mechanism to constrain the generated material to a specific category or to control the coarse structure of the generated material, such as the exact brick layout on a brick wall. Furthermore, materials reconstructed from a single input photo commonly have artifacts and are generally not tileable, which limits their use in practical content creation pipelines. We propose TileGen, a generative model for SVBRDFs that is specific to a material category, always tileable, and optionally conditional on a provided input structure pattern. TileGen is a variant of StyleGAN whose architecture is modified to always produce tileable (periodic) material maps. In addition to the standard “style” latent code, TileGen can optionally take a condition image, giving a user direct control over the dominant spatial (and optionally color) features of the material. For example, in brick materials, the user can specify a brick layout and the brick color, or in leather materials, the locations of wrinkles and folds. Our inverse rendering approach can find a material perceptually matching a single target photograph by optimization. This reconstruction can also be conditional on a user-provided pattern. The resulting materials are tileable, can be larger than the target image, and are editable by varying the condition. Xilong Zhou 0001, Milos Hasan, Valentin Deschaintre, Paul Guerrero 0001, Kalyan Sunkavalli, Nima Khademi Kalantari |
SIGGRAPH Asia | 6 |
| 2022 | Differentiable Simulation of Inertial MusculotendonsabstractWe propose a simple and practical approach for incorporating the effects of muscle inertia, which has been ignored by previous musculoskeletal simulators in both graphics and biomechanics. We approximate the inertia of the muscle by assuming that muscle mass is distributed along the centerline of the muscle. We express the motion of the musculotendons in terms of the motion of the skeletal joints using a chain of Jacobians, so that at the top level, only the reduced degrees of freedom of the skeleton are used to completely drive both bones and musculotendons. Our approach can handle all commonly used musculotendon path types, including those with multiple path points and wrapping surfaces. For muscle paths involving wrapping surfaces, we use neural networks to model the Jacobians, trained using existing wrapping surface libraries, which allows us to effectively handle the Jacobian discontinuities that occur when musculotendon paths collide with wrapping surfaces. We demonstrate support for higher-order time integrators, complex joints, inverse dynamics, Hill-type muscle models, and differentiability. In the limit, as the muscle mass is reduced to zero, our approach gracefully degrades to traditional simulators without support for muscle inertia. Finally, it is possible to mix and match inertial and non-inertial musculotendons, depending on the application. Jasper Verheul, Sang Hoon Yeo, Nima Khademi Kalantari, Shinjiro Sueda |
ACM Trans. Graph. | 4 |
| 2022 | Look-Ahead Training with Learned Reflectance Loss for Single-Image SVBRDF EstimationabstractIn this paper, we propose a novel optimization-based method to estimate the reflectance properties of a near planar surface from a single input image. Specifically, we perform test-time optimization by directly updating the parameters of a neural network to minimize the test error. Since single image SVBRDF estimation is a highly ill-posed problem, such an optimization is prone to overfitting. Our main contribution is to address this problem by introducing a training mechanism that takes the test-time optimization into account. Specifically, we train our network by minimizing the training loss after one or more gradient updates with the test loss. By training the network in this manner, we ensure that the network does not overfit to the input image during the test-time optimization process. Additionally, we propose a learned reflectance loss to augment the typically used rendering loss during the test-time optimization. We do so by using an auxiliary network that estimates pseudo ground truth reflectance parameters and train it in combination with the main network. Our approach is able to converge with a small number of iterations of the test-time optimization and produces better results compared to the state-of-the-art methods. Xilong Zhou 0001, Nima Khademi Kalantari |
ACM Trans. Graph. | 2 |
| 2021 | Multi-Stage Raw Video Denoising with Adversarial Loss and Gradient MaskabstractIn this paper, we propose a learning-based approach for denoising raw videos captured under low lighting conditions. We propose to do this by first explicitly aligning the neighboring frames to the current frame using a convolutional neural network (CNN). We then fuse the registered frames using another CNN to obtain the final denoised frame. To avoid directly aligning the temporally distant frames, we perform the two processes of alignment and fusion in multiple stages. Specifically, at each stage, we perform the denoising process on three consecutive input frames to generate the intermediate denoised frames which are then passed as the input to the next stage. By performing the process in multiple stages, we can effectively utilize the information of neighboring frames without directly aligning the temporally distant frames. We train our multi-stage system using an adversarial loss with a conditional discriminator. Specifically, we condition the discriminator on a soft gradient mask to prevent introducing high-frequency artifacts in smooth regions. We show that our system is able to produce temporally coherent videos with realistic details. Furthermore, we demonstrate through extensive experiments that our approach outperforms state-of-the-art image and video denoising methods both numerically and visually. Avinash Paliwal, Libing Zeng, Nima Khademi Kalantari |
ICCP | 3 |
| 2021 | Adversarial Single-Image SVBRDF Estimation with Hybrid TrainingabstractAbstract In this paper, we propose a deep learning approach for estimating the spatially‐varying BRDFs (SVBRDF) from a single image. Most existing deep learning techniques use pixel‐wise loss functions which limits the flexibility of the networks in handling this highly unconstrained problem. Moreover, since obtaining ground truth SVBRDF parameters is difficult, most methods typically train their networks on synthetic images and, therefore, do not effectively generalize to real examples. To avoid these limitations, we propose an adversarial framework to handle this application. Specifically, we estimate the material properties using an encoder‐decoder convolutional neural network (CNN) and train it through a series of discriminators that distinguish the output of the network from ground truth. To address the gap in data distribution of synthetic and real images, we train our network on both synthetic and real examples. Specifically, we propose a strategy to train our network on pairs of real images of the same object with different lighting. We demonstrate that our approach is able to handle a variety of cases better than the state‐of‐the‐art methods. Xilong Zhou 0001, Nima Khademi Kalantari |
Comput. Graph. Forum | 2 |
| 2020 | Deep Slow Motion Video Reconstruction With Hybrid Imaging SystemabstractSlow motion videos are becoming increasingly popular, but capturing high-resolution videos at extremely high frame rates requires professional high-speed cameras. To mitigate this problem, current techniques increase the frame rate of standard videos through frame interpolation by assuming linear object motion which is not valid in challenging cases. In this paper, we address this problem using two video streams as input; an auxiliary video with high frame rate and low spatial resolution, providing temporal information, in addition to the standard main video with low frame rate and high spatial resolution. We propose a two-stage deep learning system consisting of alignment and appearance estimation that reconstructs high resolution slow motion video from the hybrid video input. For alignment, we propose to compute flows between the missing frame and two existing frames of the main video by utilizing the content of the auxiliary video frames. For appearance estimation, we propose to combine the warped and auxiliary frames using a context and occlusion aware network. We train our model on synthetically generated hybrid videos and show high-quality results on a variety of test scenes. To demonstrate practicality, we show the performance of our system on two real dual camera setups with small baseline. Avinash Paliwal, Nima Khademi Kalantari |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Synthesizing light field from a single image with variable MPI and two network fusionabstractWe propose a learning-based approach to synthesize a light field with a small baseline from a single image. We synthesize the novel view images by first using a convolutional neural network (CNN) to promote the input image into a layered representation of the scene. We extend the multiplane image (MPI) representation by allowing the disparity of the layers to be inferred from the input image. We show that, compared to the original MPI representation, our representation models the scenes more accurately. Moreover, we propose to handle the visible and occluded regions separately through two parallel networks. The synthesized images using these two networks are then combined through a soft visibility mask to generate the final results. To effectively train the networks, we introduce a large-scale light field dataset of over 2,000 unique scenes containing a wide range of objects. We demonstrate that our approach synthesizes high-quality light fields on a variety of scenes, better than the state-of-the-art methods. Qinbo Li, Nima Khademi Kalantari |
ACM Trans. Graph. | 2 |
| 2020 | Single image HDR reconstruction using a CNN with masked features and perceptual lossabstractDigital cameras can only capture a limited range of real-world scenes' luminance, producing images with saturated pixels. Existing single image high dynamic range (HDR) reconstruction methods attempt to expand the range of luminance, but are not able to hallucinate plausible textures, producing results with artifacts in the saturated areas. In this paper, we present a novel learning-based approach to reconstruct an HDR image by recovering the saturated pixels of an input LDR image in a visually pleasing way. Previous deep learning-based methods apply the same convolutional filters on wellexposed and saturated pixels, creating ambiguity during training and leading to checkerboard and halo artifacts. To overcome this problem, we propose a feature masking mechanism that reduces the contribution of the features from the saturated areas. Moreover, we adapt the VGG-based perceptual loss function to our application to be able to synthesize visually pleasing textures. Since the number of HDR images for training is limited, we propose to train our system in two stages. Specifically, we first train our system on a large number of images for image inpainting task and then fine-tune it on HDR reconstruction. Since most of the HDR examples contain smooth regions that are simple to reconstruct, we propose a sampling strategy to select challenging training patches during the HDR fine-tuning stage. We demonstrate through experimental results that our approach can reconstruct visually pleasing HDR results, better than the current state of the art on a wide range of scenes. Marcel Santana Santos, Ing Ren Tsang, Nima Khademi Kalantari |
ACM Trans. Graph. | 3 |
| 2019 | Deep HDR Video from Sequences with Alternating ExposuresabstractAbstract A practical way to generate a high dynamic range (HDR) video using off‐the‐shelf cameras is to capture a sequence with alternating exposures and reconstruct the missing content at each frame. Unfortunately, existing approaches are typically slow and are not able to handle challenging cases. In this paper, we propose a learning‐based approach to address this difficult problem. To do this, we use two sequential convolutional neural networks (CNN) to model the entire HDR video reconstruction process. In the first step, we align the neighboring frames to the current frame by estimating the flows between them using a network, which is specifically designed for this application. We then combine the aligned and current images using another CNN to produce the final HDR frame. We perform an end‐to‐end training by minimizing the error between the reconstructed and ground truth HDR images on a set of training scenes. We produce our training data synthetically from existing HDR video datasets and simulate the imperfections of standard digital cameras using a simple approach. Experimental results demonstrate that our approach produces high‐quality HDR videos and is an order of magnitude faster than the state‐of‐the‐art techniques for sequences with two and three alternating exposures. Nima Khademi Kalantari, Ravi Ramamoorthi |
Comput. Graph. Forum | 1 |
| 2019 | Learning generative models for rendering specular microgeometryabstractRendering specular material appearance is a core problem of computer graphics. While smooth analytical material models are widely used, the high-frequency structure of real specular highlights requires considering discrete, finite microgeometry. Instead of explicit modeling and simulation of the surface microstructure (which was explored in previous work), we propose a novel direction: learning the high-frequency directional patterns from synthetic or measured examples, by training a generative adversarial network (GAN). A key challenge in applying GAN synthesis to spatially varying BRDFs is evaluating the reflectance for a single location and direction without the cost of evaluating the whole hemisphere. We resolve this using a novel method for partial evaluation of the generator network. We are also able to control large-scale spatial texture using a conditional GAN approach. The benefits of our approach include the ability to synthesize spatially large results without repetition, support for learning from measured data, and evaluation performance independent of the complexity of the dataset synthesis or measurement. Alexandr Kuznetsov, Milos Hasan, Zexiang Xu, Lingqi Yan 0001, Bruce Walter, Nima Khademi Kalantari, Steve Marschner, Ravi Ramamoorthi |
ACM Trans. Graph. | 6 |
| 2019 | Local light field fusion: practical view synthesis with prescriptive sampling guidelinesabstractWe present a practical and robust deep learning solution for capturing and rendering novel views of complex real world scenes for virtual exploration. Previous approaches either require intractably dense view sampling or provide little to no guidance for how users should sample views of a scene to reliably render high-quality novel views. Instead, we propose an algorithm for view synthesis from an irregular grid of sampled views that first expands each sampled view into a local light field via a multiplane image (MPI) scene representation, then renders novel views by blending adjacent local light fields. We extend traditional plenoptic sampling theory to derive a bound that specifies precisely how densely users should sample views of a given scene when using our algorithm. In practice, we apply this bound to capture and render views of real world scenes that achieve the perceptual quality of Nyquist rate view sampling while using up to 4000X fewer views. We demonstrate our approach's practicality with an augmented reality smart-phone app that guides users to capture input images of a scene and viewers that enable realtime virtual exploration on desktop and mobile platforms. Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, Abhishek Kar |
ACM Trans. Graph. | 4 |
| 2018 | Deep Adaptive Sampling for Low Sample Count RenderingabstractAbstract Recently, deep learning approaches have proven successful at removing noise from Monte Carlo (MC) rendered images at extremely low sampling rates, e.g., 1–4 samples per pixel (spp). While these methods provide dramatic speedups, they operate on uniformly sampled MC rendered images. However, the full promise of low sample counts requires both adaptive sampling and reconstruction/denoising. Unfortunately, the traditional adaptive sampling techniques fail to handle the cases with low sampling rates, since there is insufficient information to reliably calculate their required features, such as variance and contrast. In this paper, we address this issue by proposing a deep learning approach for joint adaptive sampling and reconstruction of MC rendered images with extremely low sample counts. Our system consists of two convolutional neural networks (CNN), responsible for estimating the sampling map and denoising, separated by a renderer. Specifically, we first render a scene with one spp and then use the first CNN to estimate a sampling map, which is used to distribute three additional samples per pixel on average adaptively. We then filter the resulting render with the second CNN to produce the final denoised image. We train both networks by minimizing the error between the denoised and ground truth images on a set of training scenes. To use backpropagation for training both networks, we propose an approach to effectively compute the gradient of the renderer. We demonstrate that our approach produces better results compared to other sampling techniques. On average, our 4 spp renders are comparable to 6 spp from uniform sampling with deep learning‐based denoising. Therefore, 50% more uniformly distributed samples are required to achieve equal quality without adaptive sampling. Alexandr Kuznetsov, Nima Khademi Kalantari, Ravi Ramamoorthi |
Comput. Graph. Forum | 2 |
| 2017 | Patch-based optimization for image-based texture mappingabstractImage-based texture mapping is a common way of producing texture maps for geometric models of real-world objects. Although a high-quality texture map can be easily computed for accurate geometry and calibrated cameras, the quality of texture map degrades significantly in the presence of inaccuracies. In this paper, we address this problem by proposing a novel global patch-based optimization system to synthesize the aligned images. Specifically, we use patch-based synthesis to reconstruct a set of photometrically-consistent aligned images by drawing information from the source images. Our optimization system is simple, flexible, and more suitable for correcting large misalignments than other techniques such as local warping. To solve the optimization, we propose a two-step approach which involves patch search and vote, and reconstruction. Experimental results show that our approach can produce high-quality texture maps better than existing techniques for objects scanned by consumer depth cameras such as Intel RealSense. Moreover, we demonstrate that our system can be used for texture editing tasks such as hole-filling and reshuffling as well as multiview camouflage. Sai Bi, Nima Khademi Kalantari, Ravi Ramamoorthi |
ACM Trans. Graph. | 2 |
| 2017 | Deep high dynamic range imaging of dynamic scenesabstractProducing a high dynamic range (HDR) image from a set of images with different exposures is a challenging process for dynamic scenes. A category of existing techniques first register the input images to a reference image and then merge the aligned images into an HDR image. However, the artifacts of the registration usually appear as ghosting and tearing in the final HDR images. In this paper, we propose a learning-based approach to address this problem for dynamic scenes. We use a convolutional neural network (CNN) as our learning model and present and compare three different system architectures to model the HDR merge process. Furthermore, we create a large dataset of input LDR images and their corresponding ground truth HDR images to train our system. We demonstrate the performance of our system by producing high-quality HDR images from a set of three LDR images. Experimental results show that our method consistently produces better results than several state-of-the-art approaches on challenging scenes. Nima Khademi Kalantari, Ravi Ramamoorthi |
ACM Trans. Graph. | 1 |
| 2017 | Light field video capture using a learning-based hybrid imaging systemabstractLight field cameras have many advantages over traditional cameras, as they allow the user to change various camera settings after capture. However, capturing light fields requires a huge bandwidth to record the data: a modern light field camera can only take three images per second. This prevents current consumer light field cameras from capturing light field videos. Temporal interpolation at such extreme scale (10x, from 3 fps to 30 fps) is infeasible as too much information will be entirely missing between adjacent frames. Instead, we develop a hybrid imaging system, adding another standard video camera to capture the temporal information. Given a 3 fps light field sequence and a standard 30 fps 2D video, our system can then generate a full light field video at 30 fps. We adopt a learning-based approach, which can be decomposed into two steps: spatio-temporal flow estimation and appearance estimation. The flow estimation propagates the angular information from the light field sequence to the 2D video, so we can warp input images to the target view. The appearance estimation then combines these warped images to output the final pixels. The whole process is trained end-to-end using convolutional neural networks. Experimental results demonstrate that our algorithm outperforms current video interpolation methods, enabling consumer light field videography, and making applications such as refocusing and parallax view generation achievable on videos for the first time. Ting-Chun Wang, Jun-Yan Zhu, Nima Khademi Kalantari, Alexei A. Efros, Ravi Ramamoorthi |
ACM Trans. Graph. | 3 |
| 2016 | Learning-based view synthesis for light field camerasabstractWith the introduction of consumer light field cameras, light field imaging has recently become widespread. However, there is an inherent trade-off between the angular and spatial resolution, and thus, these cameras often sparsely sample in either spatial or angular domain. In this paper, we use machine learning to mitigate this trade-off. Specifically, we propose a novel learning-based approach to synthesize new views from a sparse set of input views. We build upon existing view synthesis techniques and break down the process into disparity and color estimation components. We use two sequential convolutional neural networks to model these two components and train both networks simultaneously by minimizing the error between the synthesized and ground truth images. We show the performance of our approach using only four corner sub-aperture views from the light fields captured by the Lytro Illum camera. Experimental results show that our approach synthesizes high-quality images that are superior to the state-of-the-art techniques on a variety of challenging real-world scenes. We believe our method could potentially decrease the required angular resolution of consumer light field cameras, which allows their spatial resolution to increase. Nima Khademi Kalantari, Ting-Chun Wang, Ravi Ramamoorthi |
ACM Trans. Graph. | 1 |
| 2015 | Robust Radiometric Calibration for Dynamic Scenes in the WildabstractThe camera response function (CRF) that maps linear irradiance to pixel intensities must be known for computational imaging applications that match features in images with different exposures. This function is scene dependent and is difficult to estimate in scenes with significant motion. In this paper, we present a novel algorithm for radiometric calibration from multiple exposure images of a dynamic scene. Our approach is based on two key ideas from the literature: (1) intensity mapping functions which map pixel values in one image to the other without the need for pixel correspondences, and (2) a rank minimization algorithm for radiometric calibration. Although each method has its problems, we show how to combine them in a formulation that leverages their benefits. Our algorithm recovers the CRFs for dynamic scenes better than previous methods, and we show how it can be applied to existing algorithms such as those for high-dynamic range imaging to improve their results. Abhishek Badki, Nima Khademi Kalantari, Pradeep Sen |
ICCP | 2 |
| 2015 | A machine learning approach for filtering Monte Carlo noiseabstractThe most successful approaches for filtering Monte Carlo noise use feature-based filters (e.g., cross-bilateral and cross non-local means filters) that exploit additional scene features such as world positions and shading normals. However, their main challenge is finding the optimal weights for each feature in the filter to reduce noise but preserve scene detail. In this paper, we observe there is a complex relationship between the noisy scene data and the ideal filter parameters, and propose to learn this relationship using a nonlinear regression model. To do this, we use a multilayer perceptron neural network and combine it with a matching filter during both training and testing. To use our framework, we first train it in an offline process on a set of noisy images of scenes with a variety of distributed effects. Then at run-time, the trained network can be used to drive the filter parameters for new scenes to produce filtered images that approximate the ground truth. We demonstrate that our trained network can generate filtered images in only a few seconds that are superior to previous approaches on a wide range of distributed effects such as depth of field, motion blur, area lighting, glossy reflections, and global illumination. Nima Khademi Kalantari, Steve Bako, Pradeep Sen |
ACM Trans. Graph. | 1 |
| 2014 | Improving patch-based synthesis by learning patch masksabstractPatch-based synthesis is a powerful framework for numerous image and video editing applications such as hole-filling, retargeting, and reshuffling. In all these applications, a patch-based objective function is optimized through a patch search-and-vote process. However, existing techniques typically use fixed-size square patches when comparing the distance between two patches in the search process. This presents a fundamental limitation for these methods, since many patches cover multiple regions that can move, occlude, or otherwise behave independently in source and target images. We address this problem by using masks to down-weight some pixels in the patch-comparison operation. The main challenge is to choose the right mask according to the content during the search-and-vote process. We show how simple user assistance can lead to excellent results in challenging hole-filling examples. In addition, we propose a fully automated solution by learning a model to predict an appropriate mask using a set of features extracted around each patch. The model is trained using a manually annotated dataset, augmented with simulated divergence from ground truth. We demonstrate that our proposed method improves over existing approaches for single-and multi-image hole-filling applications. Nima Khademi Kalantari, Eli Shechtman, Soheil Darabi, Dan B. Goldman, Pradeep Sen |
ICCP | 1 |
| 2013 | Removing the Noise in Monte Carlo Rendering with General Image Denoising AlgorithmsabstractAbstract Monte Carlo rendering systems can produce important visual effects such as depth of field, motion blur, and area lighting, but the rendered images suffer from objectionable noise at low sampling rates. Although years of research in image processing has produced powerful denoising algorithms, most of them assume that the noise is spatially‐invariant over the entire image and cannot be directly applied to denoise Monte Carlo rendering. In this paper, we propose a new approach that enables the use of any spatially‐invariant image denoising technique to remove the noise in Monte Carlo renderings. Our key insight is to use a noise estimation metric to locally identify the amount of noise in different parts of the image, coupled with a multilevel algorithm that denoises the image in a spatially‐varying manner using a standard denoising technique. We also propose a new way to perform adaptive sampling that uses the noise estimation metric to identify the noisy regions in which to place more samples. We show that our framework runs in a few seconds with modern denoising algorithms and produces results that outperform state‐of‐the‐art techniques in Monte Carlo rendering. Nima Khademi Kalantari, Pradeep Sen |
Comput. Graph. Forum | 1 |
| 2013 | Patch-based high dynamic range videoabstractDespite significant progress in high dynamic range (HDR) imaging over the years, it is still difficult to capture high-quality HDR video with a conventional, off-the-shelf camera. The most practical way to do this is to capture alternating exposures for every LDR frame and then use an alignment method based on optical flow to register the exposures together. However, this results in objectionable artifacts whenever there is complex motion and optical flow fails. To address this problem, we propose a new approach for HDR reconstruction from alternating exposure video sequences that combines the advantages of optical flow and recently introduced patch-based synthesis for HDR images. We use patch-based synthesis to enforce similarity between adjacent frames, increasing temporal continuity. To synthesize visually plausible solutions, we enforce constraints from motion estimation coupled with a search window map that guides the patch-based synthesis. This results in a novel reconstruction algorithm that can produce high-quality HDR videos with a standard camera. Furthermore, our method is able to synthesize plausible texture and motion in fast-moving regions, where either patch-based synthesis or optical flow alone would exhibit artifacts. We present results of our reconstructed HDR video sequences that are superior to those produced by current approaches. Nima Khademi Kalantari, Eli Shechtman, Connelly Barnes, Soheil Darabi, Dan B. Goldman, Pradeep Sen |
ACM Trans. Graph. | 1 |
| 2012 | Fast Generation of Approximate Blue Noise Point SetsabstractAbstract Poisson‐disk sampling is a popular sampling method because of its blue noise power spectrum, but generation of these samples is computationally very expensive. In this paper, we propose an efficient method for fast generation of a large number of blue noise samples using a small initial patch of Poisson‐disk samples that can be generated with any existing approach. Our main idea is to convolve this set of samples with another to generate our final set of samples. We use the convolution theorem from signal processing to show that the spectrum of the resulting sample set preserves the blue noise properties. Since our method is approximate, we have error with respect to the true Poisson‐disk samples, but we show both mathematically and practically that this error is only a function of the number of samples in the small initial patch and is therefore bounded. Our method is parallelizable and we demonstrate an implementation of it on a GPU, running more than 10 times faster than any previous method and generating more than 49 million 2D samples per second. We can also use the proposed approach to generate multidimensional blue noise samples. Nima Khademi Kalantari, Pradeep Sen |
Comput. Graph. Forum | 1 |
| 2012 | Robust patch-based hdr reconstruction of dynamic scenesabstractHigh dynamic range (HDR) imaging from a set of sequential exposures is an easy way to capture high-quality images of static scenes, but suffers from artifacts for scenes with significant motion. In this paper, we propose a new approach to HDR reconstruction that draws information from all the exposures but is more robust to camera/scene motion than previous techniques. Our algorithm is based on a novel patch-based energy-minimization formulation that integrates alignment and reconstruction in a joint optimization through an equation we call the HDR image synthesis equation. This allows us to produce an HDR result that is aligned to one of the exposures yet contains information from all of them. We present results that show considerable improvement over previous approaches. Pradeep Sen, Nima Khademi Kalantari, Maziar Yaesoubi, Soheil Darabi, Dan B. Goldman, Eli Shechtman |
ACM Trans. Graph. | 2 |
| 2011 | Efficient Computation of Blue Noise Point Sets through Importance SamplingabstractAbstract Dart‐throwing can generate ideal Poisson‐disk distributions with excellent blue noise properties, but is very computationally expensive if a maximal point set is desired. In this paper, we observe that the Poisson‐disk sampling problem can be posed in terms of importance sampling by representing the available space to be sampled as a probability density function (pdf). This allows us to develop an efficient algorithm for the generation of maximal Poisson‐disk distributions with quality similar to naïve dart‐throwing but without rejection of samples. In our algorithm, we first position samples in one dimension based on its marginal cumulative distribution function (cdf). We then throw samples in the other dimension only in the regions which are available for sampling. After each 2D sample is placed, we update the cdf and data structures to keep track of the available regions. In addition to uniform sampling, our method is able to perform variable‐density sampling with small modifications. Finally, we also propose a new min‐conflict metric for variable‐density sampling which results in better adaptation of samples to the underlying importance field. Nima Khademi Kalantari, Pradeep Sen |
Comput. Graph. Forum | 1 |
| 2010 | Rational Dither Modulation using logarithmic quantization with optimum parameterabstractIn this paper, a new logarithmic quantization for Rational Dither Modulation (RDM) is presented. It can be shown that the μ-law function produces quantization levels which are in best accordance with the noise characteristics in RDM. However, since the μ-law function cannot be used in its original form, as in [5], we use it with slight modifications. In order to obtain the optimum quantizer arrangement, which results in minimum error probability, the analytical error probability is obtained and then minimized with respect to μ, which defines the compression level of the logarithmic function. The embedding distortion for the proposed method is also derived. Simulation results show that, using logarithmic quantization, better performance can be achieved in comparison with conventional RDM. Furthermore, as will be shown, the proposed method is more suitable in practice due to its perceptual advantages. Nima Khademi Kalantari, Khademi Ahadi |
ICASSP | 1 |
| 2010 | A new approach for robust realtime Voice Activity Detection using spectral patternabstractIn this paper a Voice Activity Detection approach is proposed which applies a voting algorithm to decide on the existence of speech in audio signal. For this purpose, the proposed approach uses three different short time features along with the pattern of spectral peaks of every frame. Spectral peaks pattern is appropriate for determining vowel sounds in speech signal even in the presence of noise. Therefore this measure can be applicable in voice activity detection in which the vowels characterize the speech signal. Experiments show that incorporating this measure along with our recently proposed approach for VAD, will improve the results of the algorithm considerably while imposing little computational overhead. The proposed approach is evaluated on different datasets with various noises and SNR levels and satisfying results are achieved. Mohammad Hossein Moattar, Mohammad Mehdi Homayounpour, Nima Khademi Kalantari |
ICASSP | 3 |
| 2010 | Robust audio and speech watermarking using Gaussian and Laplacian modeling
Mohammad Ali Akhaee, Nima Khademi Kalantari, Farrokh Marvasti |
Signal Process. | 2 |
| 2010 | A Robust Image Watermarking in the Ridgelet Domain Using Universally Optimum DecoderabstractA robust image watermarking scheme in the ridgelet transform domain is proposed in this paper. Due to the use of the ridgelet domain, sparse representation of an image which deals with line singularities is obtained. In order to achieve more robustness and transparency, the watermark data is embedded in selected blocks of the host image by modifying the amplitude of the ridgelet coefficients which represent the most energetic direction. Since the probability distribution function of the ridgelet coefficients is not known, we propose a universally optimum decoder to perform the watermark extraction in a distribution-independent fashion. Decoder extracts the watermark data using the variance of the ridgelet coefficients of the most energetic direction in each block. Furthermore, since the decoder needs the noise variance to perform decoding, a robust noise estimation scheme is proposed. Moreover, the implementation of error correction codes on the proposed method is investigated. Analytical derivation of bit error probability is also carried out and experimental results prove its accuracy. Simulation also shows outstanding robustness of the proposed scheme against common attacks, especially additive white noise and JPEG compression. Nima Khademi Kalantari, Seyed Mohammad Ahadi, Mansur Vafadust |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | A Logarithmic Quantization Index Modulation for Perceptually Better Data HidingabstractIn this paper, a novel arrangement for quantizer levels in the Quantization Index Modulation (QIM) method is proposed. Due to perceptual advantages of logarithmic quantization, and in order to solve the problems of a previous logarithmic quantization-based method, we used the compression function of mu-Law standard for quantization. In this regard, the host signal is first transformed into the logarithmic domain using the mu-Law compression function. Then, the transformed data is quantized uniformly and the result is transformed back to the original domain using the inverse function. The scalar method is then extended to vector quantization. For this, the magnitude of each host vector is quantized on the surface of hyperspheres which follow logarithmic radii. Optimum parameter mu for both scalar and vector cases is calculated according to the host signal distribution. Moreover, inclusion of a secret key in the proposed method, similar to the dither modulation in QIM, is introduced. Performance of the proposed method in both cases is analyzed and the analytical derivations are verified through extensive simulations on artificial signals. The method is also simulated on real images and its performance is compared with previous scalar and vector quantization-based methods. Results show that this method features stronger a watermark in comparison with conventional QIM and, as a result, has better performance while it does not suffer from the drawbacks of a previously proposed logarithmic quantization algorithm. Nima Khademi Kalantari, Seyed Mohammad Ahadi |
IEEE Trans. Image Process. | 1 |
| 2009 | Logarithmic Quantization Index Modulation: A perceptually better way to embed data within a cover signalabstractIn this paper, a new method for logarithmic quantization index modulation (QIM) is proposed. In this regard a logarithmic function is first applied to the host signal. Then the transformed signal is quantized using uniform quantization as conventional QIM to embed watermark data within. Finally using inverse transform the watermarked signal is obtained. The watermark extraction is performed using minimum distance decoder. The optimum parameter for data embedding with minimum quantization distortion is derived. Also the probability of error is analytically calculated and verified by simulation. Furthermore data hiding using secret key is proposed and the probability of error is obtained. Simulation results show that the proposed method outperforms the conventional QIM in terms of robustness when the perceptual quality of watermarked image for both methods are similar. Moreover, simulation shows that the proposed scheme has outstanding robustness in comparison with a recent quantization based data hiding method. Nima Khademi Kalantari, Seyed Mohammad Ahadi |
ICASSP | 1 |
| 2009 | Robust Multiplicative Audio and Speech Watermarking Using Statistical ModelingabstractIn this paper, a semi-blind multiplicative watermarking approach for audio and speech signals has been presented. At the receiver end, the optimal maximum likelihood (ML) detector aided by the channel side information for Gaussian and Laplacian signals in noisy environment is designed and implemented. The performance of the proposed scheme is analytically calculated and verified by simulation. Then, we adapt the proposed scheme to speech and audio signals. To improve robustness, the algorithm is applied to low frequency components of the host signal. Besides, the power of the watermark is controlled elegantly to have inaudibility using perceptual evaluation of audio quality (PEAQ) and perceptual evaluation of speech quality (PESQ) algorithms. Experimental results over several audio and speech signals show the higher robustness of the proposed technique in comparison with a recent watermarking scheme. Mohammad Ali Akhaee, Nima Khademi Kalantari, Farrokh Marvasti |
ICC | 2 |
| 2009 | Robust Multiplicative Patchwork Method for Audio WatermarkingabstractThis paper presents a Multiplicative Patchwork Method (MPM) for audio watermarking. The watermark signal is embedded by selecting two subsets of the host signal features and modifying one subset multiplicatively regarding the watermark data, whereas another subset is left unchanged. The method is implemented in wavelet domain and approximation coefficients are used to embed data. In order to have an error-free detection, the watermark data is inserted only in the frames where the ratio of the energy of subsets is between two predefined values. Also in order to control the inaudibility of watermark insertion, we use an iterative algorithm to reach a desired quality for the watermarked audio signal. The quality of watermarked signal is evaluated in each iteration using Perceptual Evaluation of Audio Quality (PEAQ) method. The probability of error is also derived for the watermarking scheme and simulation results prove the validity of the analytical derivations. Simulation results show that MPM is robust against various common attacks such as noise addition, filtering, echo, MP3 compression, etc. In comparison to the original patchwork method and its modified versions, and some recent methods, MPM provides more robustness and inaudibility of the watermark insertion. Nima Khademi Kalantari, Mohammad Ali Akhaee, Seyed Mohammad Ahadi, Hamidreza Amindavar |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Vector Quantization Index Modulation watermarking using concentric hyperspherical codebooksabstractIn this paper, a digital watermarking system based on vector quantization is presented. Each vector containing N samples is mapped on the surface of the hyperspheres each of which are associated with a message to embed the digital watermark. We called this method vector quantization index modulation (VQIM) since it is conventional QIM in the N-dimensional space. The performance of the method and its comparison to orthogonal code-based watermarking is investigated. Furthermore, we implemented the VQIM method on a real audio watermarking system and adopted it with the human auditory system. The experimental results show the robustness of this scheme against common attacks in audio watermarking such as MP3 compression, lowpass filtering, resampling etc. Nima Khademi Kalantari, Seyed Mohammad Ahadi |
ICASSP | 1 |
| 2008 | A universally optimum decoder for multiplicative audio watermarkingabstractIn this paper, we propose a novel detector for multiplicative watermarking. The decoder extracts the watermark data by comparing the variances of the watermarked signal and the original signal. Due to the use of the variance test, the decoder works independent of the distribution of the host signal. This is the major advantage of this decoder per other decoders. The decoder is optimized under the Additive White Gaussian Noise channel by calculating the best threshold value. In order to show the optimal performance of this decoder on audio signals, a Maximum likelihood (ML) decoder is also introduced by considering a Gaussian distribution for the host signal. Furthermore, PEAQ algorithm is used for controlling the inaudibility of the watermark data insertion. Using this algorithm, the watermark strength factor updates automatically every 200 ms, according to the quality which is desired for the watermarked signal. Simulation results showed that the proposed decoder is extremely robust to the common audio watermarking attacks and slightly better than the ML decoder under Additive White Gaussian Noise attack. Nima Khademi Kalantari, Seyed Mohammad Ahadi, Hamidreza Amindavar |
ICME | 1 |