Pradeep Sen

dblp:09/1364 · DBLP profile ↗
← Back
61ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-8042-924XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 59 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021
YearPublicationVenuePosition
2026 Prism: Semi-Supervised Multi-View Stereo with Monocular Structure Priors
abstract
The promise of unsupervised multi-view stereo (MVS) is to leverage large unlabeled datasets, yet current methods underperform when training on difficult data, such as handheld smartphone videos of indoor scenes. Meanwhile, high-quality synthetic datasets are available but MVS networks trained on these datasets fail to generalize to realworld examples. To bridge this gap, we propose a semisupervised learning framework that allows us to train on real and rendered images jointly, capturing structural priors from synthetic data while ensuring parity with the realworld domain. Central to our framework is a novel set of losses that leverages powerful existing monocular relativedepth estimators trained on the synthetic dataset, transferring the rich structure of this relative depth to the MVS predictions on unlabeled data. Inspired by perceptual image metrics, we compare the MVS and monocular predictions via a deep feature loss and a multi-scale statistical loss. Our full framework, which we call Prism, achieves large quantitative and qualitative improvements over current unsupervised and synthetic-supervised MVS networks. This is quite a useful result, opening the door to using both unlabeled smartphone videos and photorealistic synthetic datasets for training MVS networks.
Alexander Rich 0001, Noah Stier, Pradeep Sen, Tobias Höllerer
3DV3
2025 SiCo: An Interactive Size-Controllable Virtual Try-On Approach for Informed Decision-Making
abstract
Figure 1: SiCo Overview.Our system begins by prompting users to upload an image of themselves and provide their actual sizes for tops and bottoms.Users then proceed to the garment selection step, where they can choose different sizes for an item (Steps 1 and 3) and visualize how it would look on them (Steps 2 and 4, respectively).In addition to trying on items individually, users can also style multiple items together (Steps 5 and 6).The features in our system enable users to interact with garment sizes and fits more intuitively, providing valuable insights to help them make informed decisions about which garments and sizes to purchase.Garment images are sourced from the DressCode dataset [41].Model image © Pexels.
Sherry X. Chen, Alex Christopher Lim, Pradeep Sen, Misha Sra
Conference on Designing Interactive Systems4
2025 Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning
abstract
Although natural language instructions offer an intuitive way to guide automated image editing, deep- learning models often struggle to achieve high- quality results, largely due to the difficulty of creating large, high- quality training datasets. To do this, previous approaches have typically relied on text- to- image (T2I) generative models to produce pairs of original and edited images that simulate the input/output of an instruction- guided image- editing model. However, these image pairs often fail to align with the specified edit instructions due to the limitations of T2I models, which negatively impacts models trained on such datasets. To address this, we present Instruct- CLIP (I- CLIP), a self-supervised method that learns the semantic changes between original and edited images to refine and better align the instructions in existing datasets. Furthermore, we adapt Instruct- CLIP to handle noisy latent images and diffusion time-steps so that it can be used to train latent diffusion models (LDMs) and efficiently enforce alignment between the edit instruction and the image changes in latent space at any step of the diffusion pipeline. We use Instruct- CLIP to correct the InstructPix2Pix dataset and get over 120K refined samples we then use to fine- tune their model, guided by our novel I- CLIP- based loss function. The resulting model can produce edits that are more aligned with the given instructions. Our code and dataset are available at https://github.com/SherryXTChen/Instruct-CLIP.git.
Sherry X. Chen, Misha Sra, Pradeep Sen
CVPR3
2025 AniGrad: Anisotropic Gradient-Adaptive Sampling for 3D Reconstruction From Monocular Video
abstract
Recent image-based 3D reconstruction methods have achieved excellent quality for indoor scenes using 3D convolutional neural networks. However, they rely on a high-resolution grid in order to achieve detailed output surfaces, which is quite costly in terms of compute time, and it results in large mesh sizes that are more expensive to store, transmit, and render. In this paper we propose a new solution to this problem, using adaptive sampling. By re-formulating the final layers of the network, we are able to analytically bound the local surface complexity, and set the local sample rate accordingly. Our method, AniGrad1, achieves an order of magnitude reduction in both surface extraction latency and mesh size, while preserving mesh accuracy and detail.
Noah Stier, Alexander Rich 0001, Pradeep Sen, Tobias Höllerer
CVPR3
2025 Appearance-Preserving Scene Aggregation for Level-of-Detail Rendering
abstract
Creating an appearance-preserving level-of-detail (LoD) representation for arbitrary 3D scenes is a challenging problem. The appearance of a scene is an intricate combination of both geometry and material models and is further complicated by correlation due to the spatial configuration of scene elements. We present a novel volumetric representation for the aggregated appearance of complex scenes and a pipeline for LoD generation and rendering. The core of our representation is the Aggregated Bidirectional Scattering Distribution Function (ABSDF) that summarizes the far-field appearance of all surfaces inside a voxel. We propose a closed-form factorization of the ABSDF that accounts for spatially varying and orientation-varying material parameters. We tackle the challenge of capturing the correlation existing locally within a voxel and globally across different parts of the scene. Our method faithfully reproduces appearance and achieves higher quality than existing scene filtering methods. The memory footprint and rendering cost of our representation are decoupled from the original scene complexity.
Tao Huang 0026, Ravi Ramamoorthi, Pradeep Sen, Lingqi Yan 0001
ACM Trans. Graph.4
2024 TiNO-Edit: Timestep and Noise Optimization for Robust Diffusion-Based Image Editing
abstract
Despite many attempts to leverage pre-trained text-to-image models (T2I) like Stable Diffusion (SD) [25] for controllable image editing, producing good predictable results remains a challenge. Previous approaches have focused on either fine-tuning pre-trained T2I models on specific datasets to generate certain kinds of images (e.g., with a specific object or person), or on optimizing the weights, text prompts, and/or learning features for each input image in an attempt to coax the image generator to produce the desired result. However, these approaches all have shortcomings and fail to produce good results in a predictable and controllable manner. To address this problem, we present TiNO-Edit, an SD-based method that focuses on optimizing the noise patterns and diffusion timesteps during editing, something previously unexplored in the liter-ature. With this simple change, we are able to generate results that both better align with the original images and reflect the desired result. Furthermore, we propose a set of new loss functions that operate in the latent domain of SD, greatly speeding up the optimization when compared to prior losses, which operate in the pixel domain. Our method can be easily applied to variations of SD including Textual Inversion [13] and DreamBooth [27] that encode new concepts and incorporate them into the edited results. We present a host of image-editing capabilities enabled by our approach. Our code is publicly available at https://github.com//SherryXTChen/TiNO-Edit.
Sherry X. Chen, Yaron Vaxman, Elad Ben Baruch, David Asulin, Aviad Moreshet, Kuo-Chin Lien, Misha Sra, Pradeep Sen
CVPR8
2024 AID-AppEAL: Automatic Image Dataset and Algorithm for Content Appeal Enhancement and Assessment Labeling
Sherry X. Chen, Yaron Vaxman, Elad Ben Baruch, David Asulin, Aviad Moreshet, Misha Sra, Pradeep Sen
ECCV (19)7
2024 Smoothness, Synthesis, and Sampling: Re-thinking Unsupervised Multi-view Stereo with DIV Loss
Alexander Rich 0001, Noah Stier, Pradeep Sen, Tobias Höllerer
ECCV (63)3
2023 Deep Appearance Prefiltering
abstract
Physically based rendering of complex scenes can be prohibitively costly with a potentially unbounded and uneven distribution of complexity across the rendered image. The goal of an ideal level of detail (LoD) method is to make rendering costs independent of the 3D scene complexity, while preserving the appearance of the scene. However, current prefiltering LoD methods are limited in the appearances they can support due to their reliance of approximate models and other heuristics. We propose the first comprehensive multi-scale LoD framework for prefiltering 3D environments with complex geometry and materials (e.g., the Disney BRDF), while maintaining the appearance with respect to the ray-traced reference. Using a multi-scale hierarchy of the scene, we perform a data-driven prefiltering step to obtain an appearance phase function and directional coverage mask at each scale. At the heart of our approach is a novel neural representation that encodes this information into a compact latent form that is easy to decode inside a physically based renderer. Once a scene is baked out, our method requires no original geometry, materials, or textures at render time. We demonstrate that our approach compares favorably to state-of-the-art prefiltering methods and achieves considerable savings in memory for complex scenes.
Steve Bako, Pradeep Sen, Anton Kaplanyan
ACM Trans. Graph.2
2022 Interactive Segmentation and Visualization for Tiny Objects in Multi-megapixel Images
abstract
We introduce an interactive image segmentation and visualization framework for identifying, inspecting, and editing tiny objects (just a few pixels wide) in large multi-megapixel high-dynamic-range (HDR) images. Detecting cosmic rays (CRs) in astronomical observations is a cum-bersome workflow that requires multiple tools, so we developed an interactive toolkit that unifies model inference, HDR image visualization, segmentation mask inspection and editing into a single graphical user interface. The feature set, initially designed for astronomical data, makes this work a useful research-supporting tool for human-in-the-loop tiny-object segmentation in scientific areas like biomedicine, materials science, remote sensing, etc., as well as computer vision. Our interface features mouse-controlled, synchronized, dual-window visualization of the image and the segmentation mask, a critical feature for locating tiny objects in multi-megapixel images. The browser-based tool can be readily hosted on the web to provide multi-user access and GPU acceleration for any device. The toolkit can also be used as a high-precision annotation tool, or adapted as the frontend for an interactive machine learning framework. Our open-source dataset, CR detection model, and visualization toolkit are available at https://github.com/cy-xu/cosmic-com.
Boning Dong, Noah Stier, Curtis McCully, D. Andrew Howell, Pradeep Sen, Tobias Höllerer
CVPR6
2022 Impact of Annotator Demographics on Sentiment Dataset Labeling
abstract
As machine learning methods become more powerful and capture more nuances of human behavior, biases in the dataset can shape what the model learns and is evaluated on. This paper explores and attempts to quantify the uncertainties and biases due to annotator demographics when creating sentiment analysis datasets. We ask >1000 crowdworkers to provide their demographic information and annotations for multimodal sentiment data and its component modalities. We show that demographic differences among annotators impute a significant effect on their ratings, and that these effects also occur in each component modality. We compare predictions of different state-of-the-art multimodal machine learning algorithms against annotations provided by different demographic groups, and find that changing annotator demographics can cause >4.5 in accuracy difference when determining positive versus negative sentiment. Our findings underscore the importance of accounting for crowdworker attributes, such as demographics, when building datasets, evaluating algorithms, and interpreting results for sentiment analysis.
Yi Ding 0010, Jacob You, Tonja Machulla, Jennifer Jacobs 0001, Pradeep Sen, Tobias Höllerer
Proc. ACM Hum. Comput. Interact.5
2022 Towards practical physical-optics rendering
abstract
Physical light transport (PLT) algorithms can represent the wave nature of light globally in a scene, and are consistent with Maxwell's theory of electromagnetism. As such, they are able to reproduce the wave-interference and diffraction effects of real physical optics. However, the recent works that have proposed PLT are too expensive to apply to real-world scenes with complex geometry and materials. To address this problem, we propose a novel framework for physical light transport based on several key ideas that actually makes PLT practical for complex scenes. First, we restrict the spatial coherence shape of light to an anisotropic Gaussian and justify this restriction with general arguments based on entropy. This restriction serves to simplify the rest of the derivations, without practical loss of generality. To describe partially-coherent light, we present new rendering primitives that generalize the radiometric radiance and irradiance, and are based on the well-known Stokes parameters. We are able to represent light of arbitrary spectral content and states of polarization, and with any coherence volume and anisotropy. We also present the wave BSDF to accurately render diffractions and wave-interference effects. Furthermore, we present an approach to importance sample this wave BSDF to facilitate bi-directional path tracing, which has been previously impossible. We show good agreement with state-of-the-art methods, but unlike them we are able to render complex scenes where all the materials are new, coherence-aware physical optics materials, and with performance approaching that of "classical" rendering methods.
Shlomi Steinberg, Pradeep Sen, Lingqi Yan 0001
ACM Trans. Graph.2
2021 3DVNet: Multi-View Depth Prediction and Volumetric Refinement
abstract
We present 3DVNet, a novel multi-view stereo (MVS) depth-prediction method that combines the advantages of previous depth-based and volumetric MVS approaches. Our key idea is the use of a 3D scene-modeling network that iteratively updates a set of coarse depth predictions, resulting in highly accurate predictions which agree on the underlying scene geometry. Unlike existing depth-prediction techniques, our method uses a volumetric 3D convolutional neural network (CNN) that operates in world space on all depth maps jointly. The network can therefore learn meaningful scene-level priors. Furthermore, unlike existing volumetric MVS techniques, our 3D CNN operates on a feature-augmented point cloud, allowing for effective aggregation of multi-view information and flexible iterative refinement of depth maps. Experimental results show our method exceeds state-of-the-art accuracy in both depth prediction and 3D reconstruction metrics on the ScanNet dataset, as well as a selection of scenes from the TUM-RGBD and ICL-NUIM datasets. This shows that our method is both effective and generalizes to new settings.
Alexander Rich 0001, Noah Stier, Pradeep Sen, Tobias Höllerer
3DV3
2021 VoRTX: Volumetric 3D Reconstruction With Transformers for Voxelwise View Selection and Fusion
abstract
Recent volumetric 3D reconstruction methods can produce very accurate results, with plausible geometry even for unobserved surfaces. However, they face an undesirable trade-off when it comes to multi-view fusion. They can fuse all available view information by global averaging, thus losing fine detail, or they can heuristically cluster views for local fusion, thus restricting their ability to consider all views jointly. Our key insight is that greater detail can be retained without restricting view diversity by learning a view-fusion function conditioned on camera pose and image content. We propose to learn this multi-view fusion using a transformer. To this end, we introduce VoRTX,1an end-to-end volumetric 3D reconstruction network using transformers for wide-baseline, multi-view feature fusion. Our model is occlusion-aware, leveraging the transformer architecture to predict an initial, projective scene geometry estimate. This estimate is used to avoid back-projecting image features through surfaces into occluded regions. We train our model on ScanNet and show that it produces better reconstructions than state-of-the-art methods. We also demonstrate generalization without any fine-tuning, outperforming the same state-of-the-art methods on two other datasets, TUM-RGBD and ICL-NUIM.
Noah Stier, Alexander Rich 0001, Pradeep Sen, Tobias Höllerer
3DV3
2021 Noise-Aware Video Saliency Prediction
Ekta Prashnani, Orazio Gallo, Joohwan Kim, Josef B. Spjut, Pradeep Sen, Iuri Frosio
BMVC5
2021 Binary TTC: A Temporal Geofence for Autonomous Navigation
abstract
Time-to-contact (TTC), the time for an object to collide with the observer’s plane, is a powerful tool for path planning: it is potentially more informative than the depth, velocity, and acceleration of objects in the scene—even for humans. TTC presents several advantages, including requiring only a monocular, uncalibrated camera. However, regressing TTC for each pixel is not straightforward, and most existing methods make over-simplifying assumptions about the scene. We address this challenge by estimating TTC via a series of simpler, binary classifications. We predict with low latency whether the observer will collide with an obstacle within a certain time, which is often more critical than knowing exact, per-pixel TTC. For such scenarios, our method offers a temporal geofence in 6.4 ms—over 25× faster than existing methods. Our approach can also estimate per-pixel TTC with arbitrarily fine quantization (including continuous values), when the computational budget allows for it. To the best of our knowledge, our method is the first to offer TTC information (binary or coarsely quantized) at sufficiently high frame-rates for practical use.
Abhishek Badki, Orazio Gallo, Jan Kautz, Pradeep Sen
CVPR4
2021 Spiral-spectral fluid simulation
abstract
We introduce a fast, expressive method for simulating fluids over radial domains, including discs, spheres, cylinders, ellipses, spheroids, and tori. We do this by generalizing the spectral approach of Laplacian Eigenfunctions, resulting in what we call spiral-spectral fluid simulations. Starting with a set of divergence-free analytical bases for polar and spherical coordinates, we show that their singularities can be removed by introducing a set of carefully selected enrichment functions. Orthogonality is established at minimal cost, viscosity is supported analytically, and we specifically design basis functions that support scalable FFT-based reconstructions. Additionally, we present an efficient way of computing all the necessary advection tensors. Our approach applies to both three-dimensional flows as well as their surface-based, codimensional variants. We establish the completeness of our basis representation, and compare against a variety of existing solvers.
Qiaodong Cui, Timothy R. Langlois, Pradeep Sen, Theodore Kim
ACM Trans. Graph.3
2021 Neural complex luminaires: representation and rendering
abstract
Complex luminaires, such as grand chandeliers, can be extremely costly to render because the light-emitting sources are typically encased in complex refractive geometry, creating difficult light paths that require many samples to evaluate with Monte Carlo approaches. Previous work has attempted to speed up this process, but the methods are either inaccurate, require the storage of very large lightfields, and/or do not fit well into modern path-tracing frameworks. Inspired by the success of deep networks, which can model complex relationships robustly and be evaluated efficiently, we propose to use a machine learning framework to compress a complex luminaire's lightfield into an implicit neural representation. Our approach can easily plug into conventional renderers, as it works with the standard techniques of path tracing and multiple importance sampling (MIS). Our solution is to train three networks to perform the essential operations for evaluating the complex luminaire at a specific point and view direction, importance sampling a point on the luminaire given a shading location, and blending to determine the transparency of luminaire queries to properly composite them with other scene elements. We perform favorably relative to state-of-the-art approaches and render final images that are close to the high-sample-count reference with only a fraction of the computation and storage costs, with no need to store the original luminaire geometry and materials.
Junqiu Zhu, Yaoyi Bai, Zilin Xu, Steve Bako, Edgar Velázquez-Armendáriz, Lu Wang 0007, Pradeep Sen, Milos Hasan, Lingqi Yan 0001
ACM Trans. Graph.7
2020 Meshlet Priors for 3D Mesh Reconstruction
abstract
Estimating a mesh from an unordered set of sparse, noisy 3D points is a challenging problem that requires to carefully select priors. Existing hand-crafted priors, such as smoothness regularizers, impose an undesirable trade-off between attenuating noise and preserving local detail. Recent deep-learning approaches produce impressive results by learning priors directly from the data. However, the priors are learned at the object level, which makes these algorithms class-specific, and even sensitive to the pose of the object. We introduce meshlets, small patches of mesh that we use to learn local shape priors. Meshlets act as a dictionary of local features and thus allow to use learned priors to reconstruct object meshes in any pose and from unseen classes, even when the noise is large and the samples sparse.
Abhishek Badki, Orazio Gallo, Jan Kautz, Pradeep Sen
CVPR4
2020 Bi3D: Stereo Depth Estimation via Binary Classifications
abstract
Stereo-based depth estimation is a cornerstone of computer vision, with state-of-the-art methods delivering accurate results in real time. For several applications such as autonomous navigation, however, it may be useful to trade accuracy for lower latency. We present Bi3D, a method that estimates depth via a series of binary classifications. Rather than testing if objects are at a particular depth D, as existing stereo methods do, it classifies them as being closer or farther than D. This property offers a powerful mechanism to balance accuracy and latency. Given a strict time budget, Bi3D can detect objects closer than a given distance in as little as a few milliseconds, or estimate depth with arbitrarily coarse quantization, with complexity linear with the number of quantization levels. Bi3D can also use the allotted quantization levels to get continuous depth, but in a specific depth range. For standard stereo (i.e., continuous depth on the whole range), our method is close to or on par with state-of-the-art, finely tuned stereo methods.
Abhishek Badki, Alejandro J. Troccoli, Jan Kautz, Pradeep Sen, Orazio Gallo
CVPR5
2020 Fast and Robust Stochastic Structural Optimization
abstract
Abstract Stochastic structural analysis can assess whether a fabricated object will break under real‐world conditions. While this approach is powerful, it is also quite slow, which has previously limited its use to coarse resolutions (e.g., 26 × 34 × 28). We show that this approach can be made asymptotically faster, which in practice reduces computation time by two orders of magnitude, and allows the use of previously‐infeasible resolutions. We achieve this by showing that the probability gradient can be computed in linear time instead of quadratic, and by using a robust new scheme that stabilizes the inertia gradients used by the optimization. Additionally, we propose a constrained restart method that deals with local minima, and a sheathing approach that further reduces the weight of the shape. Together, these components enable the discovery of previously‐inaccessible designs.
Qiaodong Cui, Timothy R. Langlois, Pradeep Sen, Theodore Kim
Comput. Graph. Forum3
2020 Patch-Based Image Hallucination for Super Resolution With Detail Reconstruction From Similar Sample Images
abstract
Image hallucination and super-resolution have been studied for decades, and many approaches have been proposed to upsample low-resolution images using information from the images themselves, multiple example images, or large image databases. However, most of this work has focused exclusively on small magnification levels because the algorithms simply sharpen the blurry edges in the upsampled images - no actual new detail is typically reconstructed in the final result. In this paper, we present a patch-based algorithm for image hallucination which, for the first time, properly synthesizes novel high frequency detail. To do this, we pose the synthesis problem as a patch-based optimization which inserts coherent, high-frequency detail from contextually-similar images of the same physical scene/subject provided from either a personal image collection or a large online database. The resulting image is visually plausible and contains coherent high frequency information. We demonstrate the robustness of our algorithm by testing it on a large number of images and show that its performance is considerably superior to all state-of-the-art approaches, a result that is verified to be statistically significant through a randomized user study.
Chieh-Chi Kao, Yu-Xiang Wang 0003, Jonathan Waltman, Pradeep Sen
IEEE Trans. Multim.4
2019 Offline Deep Importance Sampling for Monte Carlo Path Tracing
abstract
Abstract Although modern path tracers are successfully being applied to many rendering applications, there is considerable interest to push them towards ever‐decreasing sampling rates. As the sampling rate is substantially reduced, however, even Monte Carlo (MC) denoisers–which have been very successful at removing large amounts of noise–typically do not produce acceptable final results. As an orthogonal approach to this, we believe that good importance sampling of paths is critical for producing better‐converged, path‐traced images at low sample counts that can then, for example, be more effectively denoised. However, most recent importance‐sampling techniques for guiding path tracing (an area known as “path guiding”) involve expensive online (per‐scene) training and offer benefits only at high sample counts. In this paper, we propose an offline, scene‐independent deep‐learning approach that can importance sample first‐bounce light paths for general scenes without the need of the costly online training, and can start guiding path sampling with as little as 1 sample per pixel. Instead of learning to “overfit” to the sampling distribution of a specific scene like most previous work, our data‐driven approach is trained a priori on a set of training scenes on how to use a local neighborhood of samples with additional feature information to reconstruct the full incident radiance at a point in the scene, which enables first‐bounce importance sampling for new test scenes. Our solution is easy to integrate into existing rendering pipelines without the need for retraining, as we demonstrate by incorporating it into both the Blender/Cycles and Mitsuba path tracers. Finally, we show how our offline, deep importance sampler (ODIS) increases convergence at low sample counts and improves the results of an off‐the‐shelf denoiser relative to other state‐of‐the‐art sampling techniques.
Steve Bako, Mark Meyer, Tony DeRose, Pradeep Sen
Comput. Graph. Forum4
2018 Localization-Aware Active Learning for Object Detection
Chieh-Chi Kao, Teng-Yok Lee, Pradeep Sen, Ming-Yu Liu 0001
ACCV (6)3
2018 PieAPP: Perceptual Image-Error Assessment Through Pairwise Preference
abstract
The ability to estimate the perceptual error between images is an important problem in computer vision with many applications. Although it has been studied extensively, however, no method currently exists that can robustly predict visual differences like humans. Some previous approaches used hand-coded models, but they fail to model the complexity of the human visual system. Others used machine learning to train models on human-labeled datasets, but creating large, high-quality datasets is difficult because people are unable to assign consistent error labels to distorted images. In this paper, we present a new learning-based method that is the first to predict perceptual image error like human observers. Since it is much easier for people to compare two given images and identify the one more similar to a reference than to assign quality scores to each, we propose a new, large-scale dataset labeled with the probability that humans will prefer one image over another. We then train a deep-learning model using a novel, pairwise-learning framework to predict the preference of one distorted image over the other. Our key observation is that our trained network can then be used separately with only one distorted image and a reference to predict its perceptual error, without ever being trained on explicit human perceptual-error labels. The perceptual error estimated by our new metric, PieAPP, is well-correlated with human opinion. Furthermore, it significantly outperforms existing algorithms, beating the state-of-the-art by almost 3Ã - on our test set in terms of binary error rate, while also generalizing to new kinds of distortions, unlike previous learning-based methods.
Ekta Prashnani, Yasamin Mostofi, Pradeep Sen
CVPR4
2018 Scalable laplacian eigenfluids
abstract
The Laplacian Eigenfunction method for fluid simulation, which we refer to as Eigenfluids , introduced an elegant new way to capture intricate fluid flows with near-zero viscosity. However, the approach does not scale well, as the memory cost grows prohibitively with the number of eigenfunctions. The method also lacks generality, because the dynamics are constrained to a closed box with Dirichlet boundaries, while open, Neumann boundaries are also needed in most practical scenarios. To address these limitations, we present a set of analytic eigenfunctions that supports uniform Neumann and Dirichlet conditions along each domain boundary, and show that by carefully applying the discrete sine and cosine transforms, the storage costs of the eigenfunctions can be made completely negligible. The resulting algorithm is both faster and more memory-efficient than previous approaches, and able to achieve lower viscosities than similar pseudo-spectral methods. We are able to surpass the scalability of the original Laplacian Eigenfunction approach by over two orders of magnitude when simulating rectangular domains. Finally, we show that the formulation allows forward scattering to be directed in a way that is not possible with any other method.
Qiaodong Cui, Pradeep Sen, Theodore Kim
ACM Trans. Graph.2
2018 A framework for developing and benchmarking sampling and denoising algorithms for Monte Carlo rendering
Jonas Deyson Brito dos Santos, Pradeep Sen, Manuel Menezes de Oliveira Neto
Vis. Comput.2
2017 GraphMatch: Efficient Large-Scale Graph Construction for Structure from Motion
abstract
We present GraphMatch, an approximate yet efficient method for building the matching graph for large-scale structure-from-motion~(SfM) pipelines. GraphMatch leverages two priors that can predict which image pairs are likely to match, thereby making the matching process for SfM much more efficient. The first is a score computed from the distance between the Fisher vectors of any two images. The second prior is based on the graph distance between vertices in the underlying matching graph. GraphMatch combines these two priors into an iterative ``sample-and-propagate'' scheme similar to the PatchMatch algorithm. Its sampling stage uses Fisher similarity priors to guide the search for matching image pairs, while its propagation stage explores neighbors of matched pairs to find new ones with a high image similarity score. Our experiments show that GraphMatch finds the most image pairs as compared to competing, approximate methods while at the same time being the most efficient.
Qiaodong Cui, Victor Fragoso, Chris Sweeney, Pradeep Sen
3DV4
2017 ANSAC: Adaptive Non-Minimal Sample and Consensus
Victor Fragoso, Chris Sweeney, Pradeep Sen, Matthew Turk 0001
BMVC3
2017 PanoTrace: interactive 3D modeling of surround-view panoramic images in virtual reality
abstract
Full-surround panoramic imagery can provide a viewer with a high-resolution visual impression of a pictured real or realistically rendered environment, but it does not provide as high a level of immersion as modeled 3D geometry can, when viewed with virtual reality (VR) headsets or projection-based setups. In this paper, we demonstrate that augmenting panorama images with geometrical models can be done simply in VR itself and can significantly increase the feeling of immersion a viewer experiences. We propose a novel interactive modeling tool that allows users to model geometry depicted in a surround-panoramic scene directly in VR, utilizing projection mapping of the panorama on top of the evolving geometry. The user interface is intuitive and allows novice users to produce geometry that approximates ground truth models sufficiently to enhance a user's VR viewing experience. We designed a user study that compares users' self-reported levels of immersion, scene realism, and discomfort on a set of created models and comparison cases. Our results indicate that our modeled scenes produce a significantly higher sense of immersion than a basic dome geometry for the panorama when viewed in VR with head orientation and position tracking.
Ehsan Sayyad, Pradeep Sen, Tobias Höllerer
VRST2
2017 A Phase-Based Approach for Animating Images Using Video Examples
abstract
Abstract We present a novel approach for animating static images that contain objects that move in a subtle, stochastic fashion (e.g. rippling water, swaying trees, or flickering candles). To do this, our algorithm leverages example videos of similar objects, supplied by the user. Unlike previous approaches which estimate motion fields in the example video to transfer motion into the image, a process which is brittle and produces artefacts, we propose an Eulerianphase‐basedapproach which uses the phase information from the sample video to animate the static image. As is well known, phase variations in a signal relate naturally to the displacement of the signal via the Fourier Shift Theorem. To enable local and spatially varying motion analysis, we analyse phase changes in a complex steerable pyramid of the example video. These phase changes are then transferred to the corresponding spatial sub‐bands of the input image to animate it. We demonstrate that this simple, phase‐based approach for transferring small motion is more effective at animating still images than methods which rely on optical flow.
Ekta Prashnani, Maneli Noorkami, Daniel Vaquero, Pradeep Sen
Comput. Graph. Forum4
2017 Computational zoom: a framework for post-capture image composition
abstract
Capturing a picture that "tells a story" requires the ability to create the right composition. The two most important parameters controlling composition are the camera position and the focal length of the lens. The traditional paradigm is for a photographer to mentally visualize the desired picture, select the capture parameters to produce it, and finally take the photograph, thus committing to a particular composition. We propose to change this paradigm. To do this, we introduce computational zoom , a framework that allows a photographer to manipulate several aspects of composition in post-processing from a stack of pictures captured at different distances from the scene. We further define a multi-perspective camera model that can generate compositions that are not physically attainable, thus extending the photographer's control over factors such as the relative size of objects at different depths and the sense of depth of the picture. We show several applications and results of the proposed computational zoom framework.
Abhishek Badki, Orazio Gallo, Jan Kautz, Pradeep Sen
ACM Trans. Graph.4
2017 Kernel-predicting convolutional networks for denoising Monte Carlo renderings
abstract
Regression-based algorithms have shown to be good at denoising Monte Carlo (MC) renderings by leveraging its inexpensive by-products (e.g., feature buffers). However, when using higher-order models to handle complex cases, these techniques often overfit to noise in the input. For this reason, supervised learning methods have been proposed that train on a large collection of reference examples, but they use explicit filters that limit their denoising ability. To address these problems, we propose a novel, supervised learning approach that allows the filtering kernel to be more complex and general by leveraging a deep convolutional neural network (CNN) architecture. In one embodiment of our framework, the CNN directly predicts the final denoised pixel value as a highly non-linear combination of the input features. In a second approach, we introduce a novel, kernel-prediction network which uses the CNN to estimate the local weighting kernels used to compute each denoised pixel from its neighbors. We train and evaluate our networks on production data and observe improvements over state-of-the-art MC denoisers, showing that our methods generalize well to a variety of scenes. We conclude by analyzing various components of our architecture and identify areas of further research in deep learning for MC denoising.
Steve Bako, Thijs Vogels, Brian McWilliams, Mark Meyer, Jan Novák, Alex Harvill, Pradeep Sen, Tony DeRose, Fabrice Rousselle
ACM Trans. Graph.7
2017 User-Perspective AR Magic Lens from Gradient-Based IBR and Semi-Dense Stereo
abstract
We present a new approach to rendering a geometrically-correct user-perspective view for a magic lens interface, based on leveraging the gradients in the real world scene. Our approach couples a recent gradient-domain image-based rendering method with a novel semi-dense stereo matching algorithm. Our stereo algorithm borrows ideas from PatchMatch, and adapts them to semi-dense stereo. This approach is implemented in a prototype device build from off-the-shelf hardware, with no active depth sensing. Despite the limited depth data, we achieve high-quality rendering for the user-perspective magic lens.
Domagoj Baricevic, Tobias Höllerer, Pradeep Sen, Matthew Turk 0001
IEEE Trans. Vis. Comput. Graph.3
2016 Removing Shadows from Images of Documents
Steve Bako, Soheil Darabi, Eli Shechtman, Jue Wang 0001, Kalyan Sunkavalli, Pradeep Sen
ACCV (3)6
2015 Robust Radiometric Calibration for Dynamic Scenes in the Wild
abstract
The camera response function (CRF) that maps linear irradiance to pixel intensities must be known for computational imaging applications that match features in images with different exposures. This function is scene dependent and is difficult to estimate in scenes with significant motion. In this paper, we present a novel algorithm for radiometric calibration from multiple exposure images of a dynamic scene. Our approach is based on two key ideas from the literature: (1) intensity mapping functions which map pixel values in one image to the other without the need for pixel correspondences, and (2) a rank minimization algorithm for radiometric calibration. Although each method has its problems, we show how to combine them in a formulation that leverages their benefits. Our algorithm recovers the CRFs for dynamic scenes better than previous methods, and we show how it can be applied to existing algorithms such as those for high-dynamic range imaging to improve their results.
Abhishek Badki, Nima Khademi Kalantari, Pradeep Sen
ICCP3
2015 Recent Advances in Adaptive Sampling and Reconstruction for Monte Carlo Rendering
abstract
Abstract Monte Carlo integration is firmly established as the basis for most practical realistic image synthesis algorithms because of its flexibility and generality. However, the visual quality of rendered images often suffers from estimator variance, which appears as visually distracting noise. Adaptive sampling and reconstruction algorithms reduce variance by controlling the sampling density and aggregating samples in a reconstruction step, possibly over large image regions. In this paper we survey recent advances in this area. We distinguish between “a priori” methods that analyze the light transport equations and derive sampling rates and reconstruction filters from this analysis, and “a posteriori” methods that apply statistical techniques to sets of samples to drive the adaptive sampling and reconstruction process. They typically estimate the errors of several reconstruction filters, and select the best filter locally to minimize error. We discuss advantages and disadvantages of recent state‐of‐the‐art techniques, and provide visual and quantitative comparisons. Some of these techniques are proving useful in real‐world applications, and we aim to provide an overview for practitioners and researchers to assess these approaches. In addition, we discuss directions for potential further improvements.
Matthias Zwicker, Wojciech Jarosz, Jaakko Lehtinen, Bochang Moon, Ravi Ramamoorthi, Fabrice Rousselle, Pradeep Sen, Cyril Soler, Sung-Eui Yoon
Comput. Graph. Forum7
2015 A Surface Approximation Method for Image and Video Correspondences
abstract
Although finding correspondences between similar images is an important problem in image processing, the existing algorithms cannot find accurate and dense correspondences in images with significant changes in lighting/transformation or with the non-rigid objects. This paper proposes a novel method for finding accurate and dense correspondences between images even in these difficult situations. Starting with the non-rigid dense correspondence algorithm [1] to generate an initial correspondence map, we propose a new geometric filter that uses cubic B-Spline surfaces to approximate the correspondence mapping functions for shared objects in both images, thereby eliminating outliers and noise. We then propose an iterative algorithm which enlarges the region containing valid correspondences. Compared with the existing methods, our method is more robust to significant changes in lighting, color, or viewpoint. Furthermore, we demonstrate how to extend our surface approximation method to video editing by first generating a reliable correspondence map between a given source frame and each frame of a video. The user can then edit the source frame, and the changes are automatically propagated through the entire video using the correspondence map. To evaluate our approach, we examine applications of unsupervised image recognition and video texture editing, and show that our algorithm produces better results than those from state-of-the-art approaches.
Jingwei Huang 0001, Bin Wang 0021, Wenping Wang 0001, Pradeep Sen
IEEE Trans. Image Process.4
2015 A machine learning approach for filtering Monte Carlo noise
abstract
The most successful approaches for filtering Monte Carlo noise use feature-based filters (e.g., cross-bilateral and cross non-local means filters) that exploit additional scene features such as world positions and shading normals. However, their main challenge is finding the optimal weights for each feature in the filter to reduce noise but preserve scene detail. In this paper, we observe there is a complex relationship between the noisy scene data and the ideal filter parameters, and propose to learn this relationship using a nonlinear regression model. To do this, we use a multilayer perceptron neural network and combine it with a matching filter during both training and testing. To use our framework, we first train it in an offline process on a set of noisy images of scenes with a variety of distributed effects. Then at run-time, the trained network can be used to drive the filter parameters for new scenes to produce filtered images that approximate the ground truth. We demonstrate that our trained network can generate filtered images in only a few seconds that are superior to previous approaches on a wide range of distributed effects such as depth of field, motion blur, area lighting, glossy reflections, and global illumination.
Nima Khademi Kalantari, Steve Bako, Pradeep Sen
ACM Trans. Graph.3
2014 Improving patch-based synthesis by learning patch masks
abstract
Patch-based synthesis is a powerful framework for numerous image and video editing applications such as hole-filling, retargeting, and reshuffling. In all these applications, a patch-based objective function is optimized through a patch search-and-vote process. However, existing techniques typically use fixed-size square patches when comparing the distance between two patches in the search process. This presents a fundamental limitation for these methods, since many patches cover multiple regions that can move, occlude, or otherwise behave independently in source and target images. We address this problem by using masks to down-weight some pixels in the patch-comparison operation. The main challenge is to choose the right mask according to the content during the search-and-vote process. We show how simple user assistance can lead to excellent results in challenging hole-filling examples. In addition, we propose a fully automated solution by learning a model to predict an appropriate mask using a set of features extracted around each patch. The model is trained using a manually annotated dataset, augmented with simulated divergence from ground truth. We demonstrate that our proposed method improves over existing approaches for single-and multi-image hole-filling applications.
Nima Khademi Kalantari, Eli Shechtman, Soheil Darabi, Dan B. Goldman, Pradeep Sen
ICCP5
2014 Efficient and robust radiance transfer for probeless photorealistic augmented reality
abstract
Photorealistic Augmented Reality (AR) requires knowledge of the scene geometry and environment lighting to compute photometric registration. Recent work has introduced probeless photometric registration, where environment lighting is estimated directly from observations of reflections in the scene rather than through an invasive probe such as a reflective ball. However, computing the dense radiance transfer of a dynamically changing scene is computationally challenging. In this work, we present an improved radiance transfer sampling approach, which combines adaptive sampling in image and visibility space with robust caching of radiance transfer to yield real time framerates for photorealistic AR scenes with dynamically changing scene geometry and environment lighting.
Lukas Gruber, Tobias Langlotz, Pradeep Sen, Tobias Hoherer, Dieter Schmalstieg
VR3
2014 User-perspective augmented reality magic lens from gradients
abstract
In this paper we present a new approach to creating a geometrically-correct user-perspective magic lens and a prototype device implementing the approach. Our prototype uses just standard color cameras, with no active depth sensing. We achieve this by pairing a recent gradient domain image-based rendering method with a novel semi-dense stereo matching algorithm inspired by PatchMatch. Our stereo algorithm is simple but fast and accurate within its search area. The resulting system is a real-time magic lens that displays the correct user perspective with a high-quality rendering, despite the lack of a dense disparity map.
Domagoj Baricevic, Tobias Höllerer, Pradeep Sen, Matthew Turk 0001
VRST3
2013 EVSAC: Accelerating Hypotheses Generation by Modeling Matching Scores with Extreme Value Theory
abstract
Algorithms based on RANSAC that estimate models using feature correspondences between images can slow down tremendously when the percentage of correct correspondences (inliers) is small. In this paper, we present a probabilistic parametric model that allows us to assign confidence values for each matching correspondence and therefore accelerates the generation of hypothesis models for RANSAC under these conditions. Our framework leverages Extreme Value Theory to accurately model the statistics of matching scores produced by a nearest-neighbor feature matcher. Using a new algorithm based on this model, we are able to estimate accurate hypotheses with RANSAC at low inlier ratios significantly faster than previous state-of-the-art approaches, while still performing comparably when the number of inliers is large. We present results of homography and fundamental matrix estimation experiments for both SIFT and SURF matches that demonstrate that our method leads to accurate and fast model estimations.
Victor Fragoso, Pradeep Sen, Sergio Rodríguez, Matthew Turk 0001
ICCV2
2013 Acceleration methods for radiance transfer in photorealistic augmented reality
abstract
Radiance transfer computation from unknown real-world environments is an intrinsic task in probe-less photometric registration for photorealistic augmented reality, which affects both the accuracy of the real-world light estimation and the quality of the rendering. We discuss acceleration methods that can reduce the overall ray-tracing costs for computing the radiance transfer for photometric registration in order to free up resources for more advanced augmented reality lighting. We also present evaluation metrics for a systematic evaluation.
Lukas Gruber, Pradeep Sen, Tobias Höllerer, Dieter Schmalstieg
ISMAR2
2013 Removing the Noise in Monte Carlo Rendering with General Image Denoising Algorithms
abstract
Abstract Monte Carlo rendering systems can produce important visual effects such as depth of field, motion blur, and area lighting, but the rendered images suffer from objectionable noise at low sampling rates. Although years of research in image processing has produced powerful denoising algorithms, most of them assume that the noise is spatially‐invariant over the entire image and cannot be directly applied to denoise Monte Carlo rendering. In this paper, we propose a new approach that enables the use of any spatially‐invariant image denoising technique to remove the noise in Monte Carlo renderings. Our key insight is to use a noise estimation metric to locally identify the amount of noise in different parts of the image, coupled with a multilevel algorithm that denoises the image in a spatially‐varying manner using a standard denoising technique. We also propose a new way to perform adaptive sampling that uses the noise estimation metric to identify the noisy regions in which to place more samples. We show that our framework runs in a few seconds with modern denoising algorithms and produces results that outperform state‐of‐the‐art techniques in Monte Carlo rendering.
Nima Khademi Kalantari, Pradeep Sen
Comput. Graph. Forum2
2013 Patch-based high dynamic range video
abstract
Despite significant progress in high dynamic range (HDR) imaging over the years, it is still difficult to capture high-quality HDR video with a conventional, off-the-shelf camera. The most practical way to do this is to capture alternating exposures for every LDR frame and then use an alignment method based on optical flow to register the exposures together. However, this results in objectionable artifacts whenever there is complex motion and optical flow fails. To address this problem, we propose a new approach for HDR reconstruction from alternating exposure video sequences that combines the advantages of optical flow and recently introduced patch-based synthesis for HDR images. We use patch-based synthesis to enforce similarity between adjacent frames, increasing temporal continuity. To synthesize visually plausible solutions, we enforce constraints from motion estimation coupled with a search window map that guides the patch-based synthesis. This results in a novel reconstruction algorithm that can produce high-quality HDR videos with a standard camera. Furthermore, our method is able to synthesize plausible texture and motion in fast-moving regions, where either patch-based synthesis or optical flow alone would exhibit artifacts. We present results of our reconstructed HDR video sequences that are superior to those produced by current approaches.
Nima Khademi Kalantari, Eli Shechtman, Connelly Barnes, Soheil Darabi, Dan B. Goldman, Pradeep Sen
ACM Trans. Graph.6
2012 Fast Generation of Approximate Blue Noise Point Sets
abstract
Abstract Poisson‐disk sampling is a popular sampling method because of its blue noise power spectrum, but generation of these samples is computationally very expensive. In this paper, we propose an efficient method for fast generation of a large number of blue noise samples using a small initial patch of Poisson‐disk samples that can be generated with any existing approach. Our main idea is to convolve this set of samples with another to generate our final set of samples. We use the convolution theorem from signal processing to show that the spectrum of the resulting sample set preserves the blue noise properties. Since our method is approximate, we have error with respect to the true Poisson‐disk samples, but we show both mathematically and practically that this error is only a function of the number of samples in the small initial patch and is therefore bounded. Our method is parallelizable and we demonstrate an implementation of it on a GPU, running more than 10 times faster than any previous method and generating more than 49 million 2D samples per second. We can also use the proposed approach to generate multidimensional blue noise samples.
Nima Khademi Kalantari, Pradeep Sen
Comput. Graph. Forum2
2012 Image melding: combining inconsistent images using patch-based synthesis
abstract
Current methods for combining two different images produce visible artifacts when the sources have very different textures and structures. We present a new method for synthesizing a transition region between two source images, such that inconsistent color, texture, and structural properties all change gradually from one source to the other. We call this process image melding . Our method builds upon a patch-based optimization foundation with three key generalizations: First, we enrich the patch search space with additional geometric and photometric transformations. Second, we integrate image gradients into the patch representation and replace the usual color averaging with a screened Poisson equation solver. And third, we propose a new energy based on mixed L 2 /L 0 norms for colors and gradients that produces a gradual transition between sources without sacrificing texture sharpness. Together, all three generalizations enable patch-based solutions to a broad class of image melding problems involving inconsistent sources: object cloning, stitching challenging panoramas, hole filling from multiple photos, and image harmonization. In several cases, our unified method outperforms previous state-of-the-art methods specifically designed for those applications.
Soheil Darabi, Eli Shechtman, Connelly Barnes, Dan B. Goldman, Pradeep Sen
ACM Trans. Graph.5
2012 On filtering the noise from the random parameters in Monte Carlo rendering
abstract
Monte Carlo (MC) rendering systems can produce spectacular images but are plagued with noise at low sampling rates. In this work, we observe that this noise occurs in regions of the image where the sample values are a direct function of the random parameters used in the Monte Carlo system. Therefore, we propose a way to identify MC noise by estimating this functional relationship from a small number of input samples. To do this, we treat the rendering system as a black box and calculate the statistical dependency between the outputs and inputs of the system. We then use this information to reduce the importance of the sample values affected by MC noise when applying an image-space, cross-bilateral filter, which removes only the noise caused by the random parameters but preserves important scene detail. The process of using the functional relationships between sample values and the random parameter inputs to filter MC noise is called Random Parameter Filtering (RPF), and we demonstrate that it can produce images in a few minutes that are comparable to those rendered with a thousand times more samples. Furthermore, our algorithm is general because we do not assign any physical meaning to the random parameters, so it works for a wide range of Monte Carlo effects, including depth of field, area light sources, motion blur, and path-tracing. We present results for still images and animated sequences at low sampling rates that have higher quality than those produced with previous approaches.
Pradeep Sen, Soheil Darabi
ACM Trans. Graph.1
2012 Robust patch-based hdr reconstruction of dynamic scenes
abstract
High dynamic range (HDR) imaging from a set of sequential exposures is an easy way to capture high-quality images of static scenes, but suffers from artifacts for scenes with significant motion. In this paper, we propose a new approach to HDR reconstruction that draws information from all the exposures but is more robust to camera/scene motion than previous techniques. Our algorithm is based on a novel patch-based energy-minimization formulation that integrates alignment and reconstruction in a joint optimization through an equation we call the HDR image synthesis equation. This allows us to produce an HDR result that is aligned to one of the exposures yet contains information from all of them. We present results that show considerable improvement over previous approaches.
Pradeep Sen, Nima Khademi Kalantari, Maziar Yaesoubi, Soheil Darabi, Dan B. Goldman, Eli Shechtman
ACM Trans. Graph.1
2011 Efficient Computation of Blue Noise Point Sets through Importance Sampling
abstract
Abstract Dart‐throwing can generate ideal Poisson‐disk distributions with excellent blue noise properties, but is very computationally expensive if a maximal point set is desired. In this paper, we observe that the Poisson‐disk sampling problem can be posed in terms of importance sampling by representing the available space to be sampled as a probability density function (pdf). This allows us to develop an efficient algorithm for the generation of maximal Poisson‐disk distributions with quality similar to naïve dart‐throwing but without rejection of samples. In our algorithm, we first position samples in one dimension based on its marginal cumulative distribution function (cdf). We then throw samples in the other dimension only in the regions which are available for sampling. After each 2D sample is placed, we update the cdf and data structures to keep track of the available regions. In addition to uniform sampling, our method is able to perform variable‐density sampling with small modifications. Finally, we also propose a new min‐conflict metric for variable‐density sampling which results in better adaptation of samples to the underlying importance field.
Nima Khademi Kalantari, Pradeep Sen
Comput. Graph. Forum2
2011 A versatile HDR video production system
abstract
Although High Dynamic Range (HDR) imaging has been the subject of significant research over the past fifteen years, the goal of acquiring cinema-quality HDR images of fast-moving scenes using available components has not yet been achieved. In this work, we present an optical architecture for HDR imaging that allows simultaneous capture of high, medium, and low-exposure images on three sensors at high fidelity with efficient use of the available light. We also present an HDR merging algorithm to complement this architecture, which avoids undesired artifacts when there is a large exposure difference between the images. We implemented a prototype high-definition HDR-video system and we present still frames from the acquired HDR video, tonemapped with various techniques.
Michael D. Tocci, Chris Kiser, Nora Tocci, Pradeep Sen
ACM Trans. Graph.4
2011 Compressive Rendering: A Rendering Application of Compressed Sensing
abstract
Recently, there has been growing interest in compressed sensing (CS), the new theory that shows how a small set of linear measurements can be used to reconstruct a signal if it is sparse in a transform domain. Although CS has been applied to many problems in other fields, in computer graphics, it has only been used so far to accelerate the acquisition of light transport. In this paper, we propose a novel application of compressed sensing by using it to accelerate ray-traced rendering in a manner that exploits the sparsity of the final image in the wavelet basis. To do this, we raytrace only a subset of the pixel samples in the spatial domain and use a simple, greedy CS-based algorithm to estimate the wavelet transform of the image during rendering. Since the energy of the image is concentrated more compactly in the wavelet domain, less samples are required for a result of given quality than with conventional spatial-domain rendering. By taking the inverse wavelet transform of the result, we compute an accurate reconstruction of the desired final image. Our results show that our framework can achieve high-quality images with approximately 75 percent of the pixel samples using a nonadaptive sampling scheme. In addition, we also perform better than other algorithms that might be used to fill in the missing pixel data, such as interpolation or inpainting. Furthermore, since the algorithm works in image space, it is completely independent of scene complexity.
Pradeep Sen, Soheil Darabi
IEEE Trans. Vis. Comput. Graph.1
2010 Compressed sensing for aperture synthesis imaging
abstract
The theory of compressed sensing has a natural application in interferometric aperture synthesis. As in many real-world applications, however, the assumption of random sampling, which is elementary to many propositions of this theory, is not met. Instead, the induced sampling patterns exhibit a large degree of regularity. In this paper, we statistically quantify the effects of this kind of regularity for the problem of radio interferometry where astronomical images are sparsely sampled in the frequency domain. Based on the favorable results of our statistical evaluation, we present a practical method for interferometric image reconstruction that is evaluated on observational data from the Very Large Array (VLA) telescope.
Stephan Wenger, Soheil Darabi, Pradeep Sen, Karl-Heinz Glassmeier, Marcus A. Magnor
ICIP3
2010 Compressive estimation for signal integration in rendering
abstract
Abstract In rendering applications, we are often faced with the problem of computing the integral of an unknown function. Typical approaches used to estimate these integrals are often based on Monte Carlo methods that slowly converge to the correct answer after many point samples have been taken. In this work, we study this problem under the framework of compressed sensing and reach the conclusion that if the signal is sparse in a transform domain, we can evaluate the integral accurately using a small set of point samples without requiring the lengthy iterations of Monte Carlo approaches. We demonstrate the usefulness of our framework by proposing novel algorithms to address two problems in computer graphics: image antialiasing and motion blur. We show that we can use our framework to generate good results with fewer samples than is possible with traditional approaches.
Pradeep Sen, Soheil Darabi
Comput. Graph. Forum1
2009 A novel framework for imaging using compressed sensing
abstract
Recently, there has been growing interest in using compressed sensing to perform imaging. Most of these algorithms capture the image of a scene by taking projections of the imaged scene with a large set of different random patterns. Unfortunately, these methods require thousands of serial measurements in order to reconstruct a high quality image, which makes them impractical for most real-world imaging applications. In this work, we explore the idea of performing sparse image capture from a single image taken in one moment of time. Our framework measures a subset of the pixels in the photograph and uses compressed sensing algorithms to reconstruct the entire image from this data. The benefit of our approach is that we can get a high-quality image while reducing the bandwidth of the imaging device because we only read a fraction of the pixels, not the entire array. Our approach can also be used to accurately fill in the missing pixel information for sensor arrays with defective pixels. We demonstrate better reconstructions of test images using our approach than with traditional reconstruction methods.
Pradeep Sen, Soheil Darabi
ICIP1
2009 Compressive Dual Photography
abstract
Abstract The accurate measurement of the light transport characteristics of a complex scene is an important goal in computer graphics and has applications in relighting and dual photography. However, since the light transport data sets are typically very large, much of the previous research has focused on adaptive algorithms that capture them efficiently. In this work, we propose a novel, non‐adaptive algorithm that takes advantage of the compressibility of the light transport signal in a transform domain to capture it with less acquisitions than with standard approaches. To do this, we leverage recent work in the area of compressed sensing, where a signal is reconstructed from a few samples assuming that it is sparse in a transform domain. We demonstrate our approach by performing dual photography and relighting by using a much smaller number of acquisitions than would normally be needed. Because our algorithm is not adaptive, it is also simpler to implement than many of the current approaches.
Pradeep Sen, Soheil Darabi
Comput. Graph. Forum1
2008 Accelerating active contour algorithms with the Gradient Diffusion Field
abstract
Active contours were proposed by Kass et al. as a way to represent the contours of an image. Although the method is simple, one of its shortcomings is its inability to converge into concave structures. The gradient vector flow (GVF) algorithm was put forth by Xu and Prince to succesfully address the concave structure problem. Although there has been much research into GVF, little has been done to reduce its computation time, which makes it unsuitable for applications requiring real-time processing of images. In this paper, we propose a method for computing an approximation of the GVF, called the gradient diffusion field (GDF), which exhibits the same useful properties of the GVF but converges faster and requires less resources for implementation. Our proposed method is also more amenable for real-time hardware and we outline a method for implementing an active contour algorithm in FPGA hardware using the GDF.
Willie Kiser, Pradeep Sen, Chris Musial
ICPR2
2007 IStar: A Raster Representation for Scalable Image and Volume Data
abstract
Topology has been an important tool for analyzing scalar data and flow fields in visualization. In this work, we analyze the topology of multivariate image and volume data sets with discontinuities in order to create an efficient, raster-based representation we call IStar. Specifically, the topology information is used to create a dual structure that contains nodes and connectivity information for every segmentable region in the original data set. This graph structure, along with a sampled representation of the segmented data set, is embedded into a standard raster image which can then be substantially downsampled and compressed. During rendering, the raster image is upsampled and the dual graph is used to reconstruct the original function. Unlike traditional raster approaches, our representation can preserve sharp discontinuities at any level of magnification, much like scalable vector graphics. However, because our representation is raster-based, it is well suited to the real-time rendering pipeline. We demonstrate this by reconstructing our data sets on graphics hardware at real-time rates.
Joe Michael Kniss, Warren A. Hunt, Kristi Potter, Pradeep Sen
IEEE Trans. Vis. Comput. Graph.4
2005 Dual photography
abstract
We present a novel photographic technique called dual photography, which exploits Helmholtz reciprocity to interchange the lights and cameras in a scene. With a video projector providing structured illumination, reciprocity permits us to generate pictures from the viewpoint of the projector, even though no camera was present at that location. The technique is completely image-based, requiring no knowledge of scene geometry or surface properties, and by its nature automatically includes all transport paths, including shadows, inter-reflections and caustics. In its simplest form, the technique can be used to take photographs without a camera; we demonstrate this by capturing a photograph using a projector and a photo-resistor. If the photo-resistor is replaced by a camera, we can produce a 4D dataset that allows for relighting with 2D incident illumination. Using an array of cameras we can produce a 6D slice of the 8D reflectance field that allows for relighting with arbitrary light fields. Since an array of cameras can operate in parallel without interference, whereas an array of light sources cannot, dual photography is fundamentally a more efficient way to capture such a 6D dataset than a system based on multiple projectors and one camera. As an example, we show how dual photography can be used to capture and relight scenes.
Pradeep Sen, Billy Chen, Steve Marschner, Mark Horowitz, Marc Levoy, Hendrik P. A. Lensch
ACM Trans. Graph.1
2003 Shadow silhouette maps
abstract
The most popular techniques for interactive rendering of hard shadows areshadow mapsandshadow volumes. Shadow maps work well in regions that are completely in light or in shadow but result in objectionable artifacts near shadow boundaries. In contrast, shadow volumes generate precise shadow boundaries but require high fill rates. In this paper, we propose the method ofsilhouette maps, in which a shadow depth map is augmented by storing the location of points on the geometric silhouette. This allows the shader to construct a piecewise linear approximation to the true shadow silhouette, improving the visual quality over the piecewise constant approximation of conventional shadow maps. We demonstrate an implementation of our approach running on programmable graphics hardware in real-time.
Pradeep Sen, Mike Cammarano, Pat Hanrahan
ACM Trans. Graph.1