EDBT 2026 Demo / reviewers in the wild / expert
Kavita Bala
dblp:b/KavitaBala
· DBLP profile ↗
89ranked-venue papers
3as first author
16since 2021 · last 2025
0000-0001-9761-6503ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 26 · 13 since 2021Systems, architecture and hardware · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 4Software engineering, systems software and programming languages · 3 · 1 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiSciPLE: Learning Interpretable Programs for Scientific Visual DiscoveryabstractVisual data is used in numerous different scientific work-flows ranging from remote sensing to ecology. As the amount of observation data increases, the challenge is not just to make accurate predictions but also to understand the underlying mechanisms for those predictions. Good interpretation is important in scientific workflows, as it allows for better decision-making by providing insights into the data. This paper introduces an automatic way of obtaining such interpretable-by-design models, by learning programs that interleave neural networks. We propose DiSciPLE (Discovering Scientific Programs using LLMs and Evolution) an evolutionary algorithm that leverages common sense and prior knowledge of large language models (LLMs) to create Python programs explaining visual data. Additionally, we propose two improvements: a program critic and a program simplifier to improve our method further to synthesize good programs. On three different real-world problems, DiSciPLE learns state-of-the-art programs on novel tasks with no prior literature. For example, we can learn programs with 35% lower error than the closest non-interpretable baseline for population density estimation. Utkarsh Mall, Cheng Perng Phoo, Mia Chiquier, Bharath Hariharan, Kavita Bala, Carl Vondrick |
CVPR | 5 |
| 2025 | Scale-aware Recognition in Satellite Images under Resource ConstraintsabstractRecognition of features in satellite imagery (forests, swimming pools, etc.) depends strongly on the spatial scale of the concept and therefore the resolution of the images. This poses two challenges:
Which resolution is best suited for recognizing a given concept, and where and when should the costlier higher-resolution (HR) imagery be acquired?
We present a novel scheme to address these challenges by introducing three components: (1) A technique to distill knowledge from models trained on HR imagery to recognition models that operate on imagery of lower resolution (LR), (2) a sampling strategy for HR imagery based on model disagreement, and (3) an LLM-based approach for inferring concept "scale". With these components we present a system to efficiently perform scale-aware recognition in satellite imagery, improving accuracy over single-scale inference while following budget constraints. **Our novel approach offers up to a 26.3\% improvement over entirely HR baselines, using 76.3 \% fewer HR images.** Resources are available at https://www.cs.cornell.edu/~revankar/scale_aware. Shreelekha Revankar, Cheng Perng Phoo, Utkarsh Mall, Bharath Hariharan, Kavita Bala |
ICLR | 5 |
| 2025 | MONITRS: Multimodal Observations of Natural Incidents Through Remote SensingabstractNatural disasters cause devastating damage to communities and infrastructure every year. Effective disaster response is hampered by the difficulty of accessing affected areas during and after events. Remote sensing has allowed us to monitor natural disasters in a remote way. More recently there have been advances in computer vision and deep learning that help automate satellite imagery analysis, However, they remain limited by their narrow focus on specific disaster types, reliance on manual expert interpretation, and lack of datasets with sufficient temporal granularity or natural language annotations for tracking disaster progression. We present MONITRS, a novel multimodal dataset of $\sim$10,000 FEMA disaster events with temporal satellite imagery with natural language annotations from news articles, accompanied by geotagged locations, and question-answer pairs. We demonstrate that fine-tuning existing MLLMs on our dataset yields significant performance improvements for disaster monitoring tasks, establishing a new benchmark for machine learning-assisted disaster response systems. Shreelekha Revankar, Utkarsh Mall, Cheng Perng Phoo, Kavita Bala, Bharath Hariharan |
NeurIPS | 4 |
| 2024 | Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentabstractWe introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for connecting remote-sensing images and language. Specifically, we train an image encoder for remote sensing images to align with the image encoder of CLIP using a large amount of paired internet and satellite images. Our unsupervised approach enables the training of a first-of-its-kind large scale VLM for remote sensing images at two different resolutions. We show that these VLMs enable zero-shot, open-vocabulary image classification, retrieval, segmentation and visual question answering for satellite images. On each of these tasks, our VLM trained without textual annotations outperforms existing VLMs trained with supervision, with gains of up to 20\% for classification and 80\% for segmentation. Utkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick, Bharath Hariharan, Kavita Bala |
ICLR | 6 |
| 2024 | AllClear: A Comprehensive Dataset and Benchmark for Cloud Removal in Satellite ImageryabstractClouds in satellite imagery pose a significant challenge for downstream applications.A major challenge in current cloud removal research is the absence of a comprehensive benchmark and a sufficiently large and diverse training dataset.To address this problem, we introduce the largest public dataset -- *AllClear* for cloud removal, featuring 23,742 globally distributed regions of interest (ROIs) with diverse land-use patterns, comprising 4 million images in total. Each ROI includes complete temporal captures from the year 2022, with (1) multi-spectral optical imagery from Sentinel-2 and Landsat 8/9, (2) synthetic aperture radar (SAR) imagery from Sentinel-1, and (3) auxiliary remote sensing products such as cloud masks and land cover maps.We validate the effectiveness of our dataset by benchmarking performance, demonstrating the scaling law - the PSNR rises from $28.47$ to $33.87$ with $30\times$ more data, and conducting ablation studies on the temporal length and the importance of individual modalities. This dataset aims to provide comprehensive coverage of the Earth's surface and promote better cloud removal results. Hangyu Zhou, Chia-Hsiang Kao, Cheng Perng Phoo, Utkarsh Mall, Bharath Hariharan, Kavita Bala |
NeurIPS | 6 |
| 2023 | Change-Aware Sampling and Contrastive Learning for Satellite ImagesabstractAutomatic remote sensing tools can help inform many large-scale challenges such as disaster management, climate change, etc. While a vast amount of spatio-temporal satellite image data is readily available, most of it remains unlabelled. Without labels, this data is not very useful for supervised learning algorithms. Self-supervised learning instead provides a way to learn effective representations for various downstream tasks without labels. In this work, we leverage characteristics unique to satellite images to learn better self-supervised features. Specifically, we use the temporal signal to contrast images with long-term and short-term differences, and we leverage the fact that satellite images do not change frequently. Using these characteristics, we formulate a new loss contrastive loss called Change-Aware Contrastive (CACo) Loss. Further, we also present a novel method of sampling different geographical regions. We show that leveraging these properties leads to better performance on diverse downstream tasks. For example, we see a 6.5% relative improvement for semantic segmentation and an 8.5% relative improvement for change detection over the best-performing baseline with our method. Utkarsh Mall, Bharath Hariharan, Kavita Bala |
CVPR | 3 |
| 2022 | Change Event Dataset for Discovery from Spatio-temporal Remote Sensing ImageryabstractSatellite imagery is increasingly available, high resolution, and temporally detailed. Changes in spatio-temporal datasets such as satellite images are particularly interesting as they reveal the many events and forces that shape our world. However, finding such interesting and meaningful change events from the vast data is challenging. In this paper, we present new datasets for such change events that include semantically meaningful events like road construction. Instead of manually annotating the very large corpus of satellite images, we introduce a novel unsupervised approach that takes a large spatio-temporal dataset from satellite images and finds interesting change events. To evaluate the meaningfulness on these datasets we create 2 benchmarks namely CaiRoad and CalFire which capture the events of road construction and forest fires. These new benchmarks can be used to evaluate semantic retrieval/classification performance. We explore these benchmarks qualitatively and quantitatively by using several methods and show that these new datasets are indeed challenging for many existing methods. Utkarsh Mall, Bharath Hariharan, Kavita Bala |
NeurIPS | 3 |
| 2022 | Discovering Underground Maps from FashionabstractThe fashion sense—meaning the clothing styles people wear—in a geographical region can reveal information about that region. For example, it can reflect the kind of activities people do there, or the type of crowds that frequently visit the region (e.g., tourist hot spot, student neighborhood, business center). We propose a method to create underground neighborhood maps of cities by analyzing how people dress. Using publicly available images from across a city, our method automatically segments the map into neighborhoods with a similar fashion sense. Our approach further allows discovering insights about a city, such as detecting distinct neighborhoods (what is the most unique region of NYC?) and answering analogy questions between cities (what is the "Downtown LA" of Bogota?). We also present two new underground map benchmarks derived from non-image data for 37 cities worldwide. Our method shows promising results on both these benchmarks as well as experiments with human judges."The map is not the thing mapped."—Eric Temple Bell Utkarsh Mall, Kavita Bala, Tamara Berg, Kristen Grauman |
WACV | 2 |
| 2021 | PiCIE: Unsupervised Semantic Segmentation Using Invariance and Equivariance in ClusteringabstractWe present a new framework for semantic segmentation without annotations via clustering. Off-the-shelf clustering methods are limited to curated, single-label, and objectcentric images yet real-world data are dominantly uncurated, multi-label, and scene-centric. We extend clustering from images to pixels and assign separate cluster membership to different instances within each image. However, solely relying on pixel-wise feature similarity fails to learn high-level semantic concepts and overfits to lowlevel visual cues. We propose a method to incorporate geometric consistency as an inductive bias to learn invariance and equivariance for photometric and geometric variations. With our novel learning objective, our framework can learn high-level semantic concepts. Our method, PiCIE (Pixel-level feature Clustering using Invariance and Equivariance), is the first method capable of segmenting both things and stuff categories without any hyperparameter tuning or task-specific pre-processing. Our method largely outperforms existing baselines on COCO [31] and Cityscapes [8] with +17.5 Acc. and +4.5 mIoU. We show that PiCIE gives a better initialization for standard supervised training. The code is available at https://github.com/janghyuncho/PiCIE. Jang Hyun Cho, Utkarsh Mall, Kavita Bala, Bharath Hariharan |
CVPR | 3 |
| 2021 | What Can Style Transfer and Paintings Do for Model Robustness?abstractA common strategy for improving model robustness is through data augmentations. Data augmentations encourage models to learn desired invariances, such as invariance to horizontal flipping or small changes in color. Recent work has shown that arbitrary style transfer can be used as a form of data augmentation to encourage invariance to textures by creating painting-like images from photographs. However, a stylized photograph is not quite the same as an artist-created painting. Artists depict perceptually meaningful cues in paintings so that humans can recognize salient components in scenes, an emphasis which is not enforced in style transfer. Therefore, we study how style transfer and paintings differ in their impact on model robustness. First, we investigate the role of paintings as style images for stylization-based data augmentation. We find that style transfer functions well even without paintings as style images. Second, we show that learning from paintings as a form of perceptual data augmentation can improve model robustness. Finally, we investigate the invariances learned from stylization and from paintings, and show that models learn different invariances from these differing forms of data. Our results provide insights into how stylization improves model robustness, and provide evidence that artist-created paintings can be a valuable source of data for model robustness. Code and data are available at: https://github.com/hubertsgithub/style_painting_robustness Hubert Lin, Mitchell van Zuijlen, Sylvia C. Pont, Maarten Wijntjes, Kavita Bala |
CVPR | 5 |
| 2021 | PhySG: Inverse Rendering With Spherical Gaussians for Physics-Based Material Editing and RelightingabstractWe present PhySG, an end-to-end inverse rendering pipeline that includes a fully differentiable renderer, and can reconstruct geometry, materials, and illumination from scratch from a set of images. Our framework represents specular BRDFs and environmental illumination using mixtures of spherical Gaussians, and represents geometry as a signed distance function parameterized as a Multi-Layer Perceptron. The use of spherical Gaussians allows us to efficiently solve for approximate light transport, and our method works on scenes with challenging non-Lambertian reflectance captured under natural, static illumination. We demonstrate, with both synthetic and real data, that our reconstructions not only enable rendering of novel viewpoints, but also physics-based appearance editing of materials and illumination. Kai Zhang 0045, Fujun Luan, Qianqian Wang 0002, Kavita Bala, Noah Snavely |
CVPR | 4 |
| 2021 | Field-Guide-Inspired Zero-Shot LearningabstractModern recognition systems require large amounts of supervision to achieve accuracy. Adapting to new domains requires significant data from experts, which is onerous and can become too expensive. Zero-shot learning requires an annotated set of attributes for a novel category. Annotating the full set of attributes for a novel category proves to be a tedious and expensive task in deployment. This is especially the case when the recognition domain is an expert domain. We introduce a new field-guide-inspired approach to zero-shot annotation where the learner model interactively asks for the most useful attributes that define a class. We evaluate our method on classification benchmarks with attribute annotations like CUB, SUN, and AWA2 and show that our model achieves the performance of a model with full annotations at the cost of significantly fewer number of annotations. Since the time of experts is precious, decreasing annotation cost can be very valuable for real-world deployment. Utkarsh Mall, Bharath Hariharan, Kavita Bala |
ICCV | 3 |
| 2021 | AutoPhoto: Aesthetic Photo Capture using Reinforcement LearningabstractThe process of capturing a well-composed photo is difficult and it takes years of experience to master. We propose a novel pipeline for an autonomous agent to automatically capture an aesthetic photograph by navigating within a local region in a scene. Instead of classical optimization over heuristics such as the rule-of-thirds, we adopt a data-driven aesthetics estimator to assess photo quality. A reinforcement learning framework is used to optimize the model with respect to the learned aesthetics metric. We train our model in simulation with indoor scenes, and we demonstrate that our system can capture aesthetic photos in both simulation and real world environments on a ground robot. To our knowledge, this is the first system that can automatically explore an environment to capture an aesthetic photo with respect to a learned aesthetic estimator. Source code is at https://github.com/HadiZayer/AutoPhoto Hadi AlZayer, Hubert Lin, Kavita Bala |
IROS | 3 |
| 2021 | Unified Shape and SVBRDF Recovery using Differentiable Monte Carlo RenderingabstractAbstract Reconstructing the shape and appearance of real‐world objects using measured 2D images has been a long‐standing inverse rendering problem. In this paper, we introduce a new analysis‐by‐synthesis technique capable of producing high‐quality reconstructions through robust coarse‐to‐fine optimization and physics‐based differentiable rendering. Unlike most previous methods that handle geometry and reflectance largely separately, our method unifies the optimization of both by leveraging image gradients with respect to both object reflectance and geometry. To obtain physically accurate gradient estimates, we develop a new GPU‐based Monte Carlo differentiable renderer leveraging recent advances in differentiable rendering theory to offer unbiased gradients while enjoying better performance than existing tools like PyTorch3D [RRN*20] and redner [LADL18]. To further improve robustness, we utilize several shape and material priors as well as a coarse‐to‐fine optimization strategy to reconstruct geometry. Using both synthetic and real input images, we demonstrate that our technique can produce reconstructions with higher quality than previous methods. Fujun Luan, Kavita Bala, Zhao Dong 0001 |
Comput. Graph. Forum | 3 |
| 2021 | Guest Editorial: Introduction to the Special Section on Computational PhotographyabstractThe papers in this special section focus on computational photography. The past year has been significant in many ways. For the scientific community of computational photography, we have a hybrid in-person conference, one of the first in the vision/optics/graphics community post-pandemic. Further, we expanded our community to better include the physical optics community. This move was made consciously to strengthen and expand the span of computational photography and make these different, yet closely related communities, have a common venue to share ideas. This expansion we believe has significantly enriched the papers submitted to this special issue, through the IEEE International Conference on Computational Photography (ICCP’2021). Yoav Y. Schechner, Kavita Bala, Ori Katz, Kalyan Sunkavalli, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Scene Summarization via Motion NormalizationabstractWhen observing the visual world, temporal phenomena are ubiquitous: people walk, cars drive, rivers flow, clouds drift, and shadows elongate. Some of these, like water splashing and cloud motion, occur over time intervals that are either too short or too long for humans to easily observe. High-speed and timelapse videos provide a popular and compelling way to visualize these phenomena, but many real-world scenes exhibit motions occurring at a variety of rates. Once a framerate is chosen, phenomena at other rates are at best invisible, and at worst create distracting artifacts. In this article, we propose to automatically normalize the pixel-space speed of different motions in an input video to produce a seamless output with spatiotemporally varying framerate. To achieve this, we propose to analyze scenes at different timescales to isolate and analyze motions that occur at vastly different rates. Our method optionally allows a user to specify additional constraints according to artistic preferences. The motion normalized output provides a novel way to compactly visualize the changes occurring in a scene over a broad range of timescales. Scott Wehrwein, Kavita Bala, Noah Snavely |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Towards Learning-based Inverse Subsurface ScatteringabstractGiven images of translucent objects, of unknown shape and lighting, we aim to use learning to infer the optical parameters controlling subsurface scattering of light inside the objects. We introduce a new architecture, the inverse transport network (ITN), that aims to improve generalization of an encoder network to unseen scenes, by connecting it with a physically-accurate, differentiable Monte Carlo renderer capable of estimating image derivatives with respect to scattering material parameters. During training, this combination forces the encoder network to predict parameters that not only match groundtruth values, but also reproduce input images. During testing, the encoder network is used alone, without the renderer, to predict material parameters from a single input image. Drawing insights from the physics of radiative transfer, we additionally use material parameterizations that help reduce estimation errors due to ambiguities in the scattering parameter space. Finally, we augment the training loss with pixelwise weight maps that emphasize the parts of the image most informative about the underlying scattering parameters. We demonstrate that this combination allows neural networks to generalize to scenes with completely unseen geometries and illuminations better than traditional networks, with 38.06% reduced parameter error on average. Chengqian Che, Fujun Luan, Kavita Bala, Ioannis Gkioulekas |
ICCP | 4 |
| 2020 | DeepSemanticHPPC: Hypothesis-based Planning over Uncertain Semantic Point CloudsabstractPlanning in unstructured environments is challenging - it relies on sensing, perception, scene reconstruction, and reasoning about various uncertainties. We propose DeepSemanticHPPC, a novel uncertainty-aware hypothesis-based planner for unstructured environments. Our algorithmic pipeline consists of: a deep Bayesian neural network which segments surfaces with uncertainty estimates; a flexible point cloud scene representation; a next-best-view planner which minimizes the uncertainty of scene semantics using sparse visual measurements; and a hypothesis-based path planner that proposes multiple kinematically feasible paths with evolving safety confidences given next-best-view measurements. Our pipeline iteratively decreases semantic uncertainty along planned paths, filtering out unsafe paths with high confidence. We show that our framework plans safe paths in real-world environments where existing path planners typically fail. Yutao Han, Hubert Lin, Jacopo Banfi, Kavita Bala, Mark E. Campbell |
ICRA | 4 |
| 2020 | Langevin monte carlo rendering with gradient-based adaptationabstractWe introduce a suite of Langevin Monte Carlo algorithms for efficient photorealistic rendering of scenes with complex light transport effects, such as caustics, interreflections, and occlusions. Our algorithms operate in primary sample space, and use the Metropolis-adjusted Langevin algorithm (MALA) to generate new samples. Drawing inspiration from state-of-the-art stochastic gradient descent procedures, we combine MALA with adaptive preconditioning and momentum schemes that re-use previously-computed first-order gradients, either in an online or in a cache-driven fashion. This combination allows MALA to adapt to the local geometry of the primary sample space, without the computational overhead associated with previous Hessian-based adaptation algorithms. We use the theory of controlled Markov chain Monte Carlo to ensure that these combinations remain ergodic, and are therefore suitable for unbiased Monte Carlo rendering. Through extensive experiments, we show that our algorithms, MALA with online and cache-driven adaptation, can successfully handle complex light transport in a large variety of scenes, leading to improved performance (on average more than 3× variance reduction at equal time, and 7× for motion blur) compared to state-of-the-art Markov chain Monte Carlo rendering algorithms. Fujun Luan, Kavita Bala, Ioannis Gkioulekas |
ACM Trans. Graph. | 3 |
| 2019 | Block Annotation: Better Image Annotation With Sub-Image DecompositionabstractImage datasets with high-quality pixel-level annotations are valuable for semantic segmentation: labelling every pixel in an image ensures that rare classes and small objects are annotated. However, full-image annotations are expensive, with experts spending up to 90 minutes per image. We propose block sub-image annotation as a replacement for full-image annotation. Despite the attention cost of frequent task switching, we find that block annotations can be crowdsourced at higher quality compared to full-image annotation with equal monetary cost using existing annotation tools developed for full-image annotation. Surprisingly, we find that 50% pixels annotated with blocks allows semantic segmentation to achieve equivalent performance to 100% pixels annotated. Furthermore, as little as 12% of pixels annotated allows performance as high as 98% of the performance with dense annotation. In weakly-supervised settings, block annotation outperforms existing methods by 3-4% (absolute) given equivalent annotation time. To recover the necessary global structure for applications such as characterizing spatial context and affordance relationships, we propose an effective method to inpaint block-annotated images with high-quality labels without additional human effort. As such, fewer annotations can also be used for these applications compared to full-image annotation. Hubert Lin, Paul Upchurch, Kavita Bala |
ICCV | 3 |
| 2019 | GeoStyle: Discovering Fashion Trends and EventsabstractUnderstanding fashion styles and trends is of great potential interest to retailers and consumers alike. The photos people upload to social media are a historical and public data source of how people dress across the world and at different times. While we now have tools to automatically recognize the clothing and style attributes of what people are wearing in these photographs, we lack the ability to analyze spatial and temporal trends in these attributes or make predictions about the future. In this paper we address this need by providing an automatic framework that analyzes large corpora of street imagery to (a) discover and forecast long-term trends of various fashion attributes as well as automatically discovered styles, and (b) identify spatio-temporally localized events that affect what people wear. We show that our framework makes long term trend forecasts that are > 20% more accurate than prior art, and identifies hundreds of socially meaningful events that impact fashion across the globe. Utkarsh Mall, Kevin Matzen, Bharath Hariharan, Noah Snavely, Kavita Bala |
ICCV | 5 |
| 2019 | Context-Aware Asset Search for Graphic DesignabstractGraphic design tools provide powerful controls for expert-level design creation, but the options can often be overwhelming for novices. This paper proposes Context-Aware Asset Search tools that take the current state of the user's design into account, thereby providing search and selections that are compatible with the current design and better fit the user's needs. In particular, we focus on image search and color selection, two tasks that are central to design. We learn a model for compatibility of images and colors within a design, using crowdsourced data. We then use the learned model to rank image search results or color suggestions during design. We found counterintuitive behavior using conventional training with pairwise comparisons for image search, where models with and without compatibility performed similarly. We describe a data collection procedure that alleviates this problem. We show that our method outperforms baseline approaches in quantitative evaluation, and we also evaluate a prototype interactive design tool. Balazs Kovacs, Peter O'Donovan, Kavita Bala, Aaron Hertzmann |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Learning Material-Aware Local Descriptors for 3D ShapesabstractMaterial understanding is critical for design, geometric modeling, and analysis of functional objects. We enable material-aware 3D shape analysis by employing a projective convolutional neural network architecture to learn material-aware descriptors from view-based representations of 3D points for point-wise material classification or material-aware retrieval. Unfortunately, only a small fraction of shapes in 3D repositories are labeled with physical materials, posing a challenge for learning methods. To address this challenge, we crowdsource a dataset of 3080 3D shapes with part-wise material labels. We focus on furniture models which exhibit interesting structure and material variability. In addition, we also contribute a high-quality expert-labeled benchmark of 115 shapes from Herman-Miller and IKEA for evaluation. We further apply a mesh-aware conditional random field, which incorporates rotational and reflective symmetries, to smooth our local material predictions across neighboring surface patches. We demonstrate the effectiveness of our learned descriptors for automatic texturing, material-aware retrieval, and physical simulation. Hubert Lin, Melinos Averkiou, Evangelos Kalogerakis, Balazs Kovacs, Siddhant Ranade, Vladimir G. Kim, Siddhartha Chaudhuri, Kavita Bala |
3DV | 8 |
| 2018 | Deep Painterly HarmonizationabstractAbstract Copying an element from a photo and pasting it into a painting is a challenging task. Applying photo compositing techniques in this context yields subpar results that look like a collage — and existing painterly stylization algorithms, which are global, perform poorly when applied locally. We address these issues with a dedicated algorithm that carefully determines the local statistics to be transferred. We ensure both spatial and inter‐scale statistical consistency and demonstrate that both aspects are key to generating quality results. To cope with the diversity of abstraction levels and types of paintings, we introduce a technique to adjust the parameters of the transfer depending on the painting. We show that our algorithm produces significantly better results than photo compositing or global stylization techniques and that it enables creative painterly edits that would be otherwise difficult to achieve. Fujun Luan, Sylvain Paris, Eli Shechtman, Kavita Bala |
Comput. Graph. Forum | 4 |
| 2017 | Shading Annotations in the WildabstractUnderstanding shading effects in images is critical for a variety of vision and graphics problems, including intrinsic image decomposition, shadow removal, image relighting, and inverse rendering. As is the case with other vision tasks, machine learning is a promising approach to understanding shading - but there is little ground truth shading data available for real-world images. We introduce Shading Annotations in the Wild (SAW), a new large-scale, public dataset of shading annotations in indoor scenes, comprised of multiple forms of shading judgments obtained via crowdsourcing, along with shading annotations automatically generated from RGB-D imagery. We use this data to train a convolutional neural network to predict per-pixel shading information in an image. We demonstrate the value of our data and network in an application to intrinsic images, where we can reduce decomposition artifacts produced by existing algorithms. Our database is available at http://opensurfaces.cs.cornell.edu/saw. Balazs Kovacs, Sean Bell, Noah Snavely, Kavita Bala |
CVPR | 4 |
| 2017 | Deep Photo Style TransferabstractThis paper introduces a deep-learning approach to photographic style transfer that handles a large variety of image content while faithfully transferring the reference style. Our approach builds upon the recent work on painterly transfer that separates style from the content of an image by considering different layers of a neural network. However, as is, this approach is not suitable for photorealistic style transfer. Even when both the input and reference images are photographs, the output still exhibits distortions reminiscent of a painting. Our contribution is to constrain the transformation from the input to the output to be locally affine in colorspace, and to express this constraint as a custom fully differentiable energy term. We show that this approach successfully suppresses distortion and yields satisfying photorealistic style transfers in a broad variety of scenarios, including transfer of the time of day, weather, season, and artistic edits. Fujun Luan, Sylvain Paris, Eli Shechtman, Kavita Bala |
CVPR | 4 |
| 2017 | Deep Feature Interpolation for Image Content ChangesabstractWe propose Deep Feature Interpolation (DFI), a new datadriven baseline for automatic high-resolution image transformation. As the name suggests, DFI relies only on simple linear interpolation of deep convolutional features from pre-trained convnets. We show that despite its simplicity, DFI can perform high-level semantic transformations like “make older/younger”, “make bespectacled”, “add smile”, among others, surprisingly well-sometimes even matching or outperforming the state-of-the-art. This is particularly unexpected as DFI requires no specialized network architecture or even any deep network to be trained for these tasks. DFI therefore can be used as a new baseline to evaluate more complex algorithms and provides a practical answer to the question of which image transformation tasks are still challenging after the advent of deep learning. Paul Upchurch, Jacob R. Gardner, Geoff Pleiss, Robert Pless, Noah Snavely, Kavita Bala, Kilian Q. Weinberger |
CVPR | 6 |
| 2017 | Intrinsic Decompositions for Image EditingabstractIntrinsic images are a mid-level representation of an image that decompose the image into reflectance and illumination layers. The reflectance layer captures the color/texture of surfaces in the scene, while the illumination layer captures shading effects caused by interactions between scene illumination and surface geometry. Intrinsic images have a long history in computer vision and recently in computer graphics, and have been shown to be a useful representation for tasks ranging from scene understanding and reconstruction to image editing. In this report, we review and evaluate past work on this problem. Specifically, we discuss each work in terms of the priors they impose on the intrinsic image problem. We introduce a new synthetic ground-truth dataset that we use to evaluate the validity of these priors and the performance of the methods. Finally, we evaluate the performance of the different methods in the context of image-editing applications. Nicolas Bonneel, Balazs Kovacs, Sylvain Paris, Kavita Bala |
Comput. Graph. Forum | 4 |
| 2017 | Fiber-Level On-the-Fly Procedural TextilesabstractAbstract Procedural textile models are compact, easy to edit, and can achieve state‐of‐the‐art realism with fiber‐level details. However, these complex models generally need to be fully instantiated (aka. realized) into 3D volumes or fiber meshes and stored in memory, We introduce a novel realization‐minimizing technique that enables physically based rendering of procedural textiles, without the need of full model realizations. The key ingredients of our technique are new data structures and search algorithms that look up regular and flyaway fibers on the fly, efficiently and consistently. Our technique works with compact fiber‐level procedural yarn models in their exact form with no approximation imposed. In practice, our method can render very large models that are practically unrenderable using existing methods, while using considerably less memory (60–200× less) and achieving good performance. Fujun Luan, Kavita Bala |
Comput. Graph. Forum | 3 |
| 2017 | Fast rendering of fabric micro-appearance models under directional and spherical gaussian lightsabstractRendering fabrics using micro-appearance models---fiber-level microgeometry coupled with a fiber scattering model---can take hours per frame. We present a fast, precomputation-based algorithm for rendering both single and multiple scattering in fabrics with repeating structure illuminated by directional and spherical Gaussian lights. Precomputed light transport (PRT) is well established but challenging to apply directly to cloth. This paper shows how to decompose the problem and pick the right approximations to achieve very high accuracy, with significant performance gains over path tracing. We treat single and multiple scattering separately and approximate local multiple scattering using precomputed transfer functions represented in spherical harmonics. We handle shadowing between fibers with precomputed per-fiber-segment visibility functions, using two different representations to separately deal with low and high frequency spherical Gaussian lights. Our algorithm is designed for GPU performance and high visual quality. Compared to existing PRT methods, it is more accurate. In tens of seconds on a commodity GPU, it renders high-quality supersampled images that take path tracing tens of minutes on a compute cluster. Pramook Khungurn, Rundong Wu, James Noeckel, Steve Marschner, Kavita Bala |
ACM Trans. Graph. | 5 |
| 2016 | Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural NetworksabstractIt is well known that contextual and multi-scale representations are important for accurate visual recognition. In this paper we present the Inside-Outside Net (ION), an object detector that exploits information both inside and outside the region of interest. Contextual information outside the region of interest is integrated using spatial recurrent neural networks. Inside, we use skip pooling to extract information at multiple scales and levels of abstraction. Through extensive experiments we evaluate the design space and provide readers with an overview of what tricks of the trade are important. ION improves state-of-the-art on PASCAL VOC 2012 object detection from 73.9% to 77.9% mAP. On the new and more challenging MS COCO dataset, we improve state-of-the-art from 19.7% to 33.1% mAP. In the 2015 MS COCO Detection Challenge, our ION model won "Best Student Entry" and finished 3rd place overall. As intuition suggests, our detection results provide strong evidence that context and multi-scale representations improve small object detection. Sean Bell, C. Lawrence Zitnick, Kavita Bala, Ross B. Girshick |
CVPR | 3 |
| 2016 | Interactive Consensus Agreement Games for Labeling ImagesabstractScene understanding algorithms in computer vision are improving dramatically by training deep convolutional neural networks on millions of accurately annotated images. Collecting large-scale datasets for this kind of training is challenging, and the learning algorithms are only as good as the data they train on. Training annotations are often obtained by taking the majority label from independent crowdsourced workers using platforms such as Amazon Mechanical Turk. However, the accuracy of the resulting annotations can vary, with the hardest-to-annotate samples having prohibitively low accuracy. Our insight is that in cases where independent worker annotations are poor more accurate results can be obtained by having workers collaborate. This paper introduces consensus agreement games, a novel method for assigning annotations to images by the agreement of multiple consensuses of small cliques of workers. We demonstrate that this approach reduces error by 37.8% on two different datasets at a cost of $0.10 or $0.17 per annotation. The higher cost is justified because our method does not need to be run on the entire dataset. Ultimately, our method enables us to more accurately annotate images and build more challenging training datasets for learning algorithms. Paul Upchurch, Daniel Sedra, Andrew Mullen, Haym Hirsh, Kavita Bala |
HCOMP | 5 |
| 2016 | Do-it-yourself lighting design for product videographyabstractThe growth of online marketplaces for selling goods has increased the need for product photography by novice users and consumers. Additionally, the increased use of online media and large-screen billboards promotes the adoption of videos for advertising, going beyond just using still imagery. Lighting is a key distinction between professional and casual product videography. Professionals use specialized hardware setups, and bring expert skills to create good lighting that shows off the product's shape and material, while also producing aesthetically pleasing results. In this paper, we introduce a new do-it-yourself (DIY) approach to lighting design that lets novice users create studio quality product videography. We identify design principles to light products through emphasizing highlights, rim lighting, and contours. We devise a set of computational metrics to achieve these design goals. Our workflow is: the user acquires a video of the product by mounting a video camera on a tripod and using a tablet to light objects by waving the tablet around the object. We automatically analyze and split this acquired video into snippets that match our design principles. Finally, we present an interface that lets users easily select snippets with specific characteristics and then assembles them to produce a final pleasing video of the product. Alternatively, they can rely on our template mechanism to automatically assemble a video. Ivaylo Boyadzhiev, Jiawen Chen 0001, Sylvain Paris, Kavita Bala |
ICCP | 4 |
| 2016 | Photometric Ambient Occlusion for Intrinsic Image DecompositionabstractWe present a method for computing ambient occlusion (AO) for a stack of images of a Lambertian scene from a fixed viewpoint. Ambient occlusion, a concept common in computer graphics, characterizes the local visibility at a point: it approximates how much light can reach that point from different directions without getting blocked by other geometry. While AO has received surprisingly little attention in vision, we show that it can be approximated using simple, per-pixel statistics over image stacks, based on a simplified image formation model. We use our derived AO measure to compute reflectance and illumination for objects without relying on additional smoothness priors, and demonstrate state-of-the art performance on the MIT Intrinsic Images benchmark. We also demonstrate our method on several synthetic and real scenes, including 3D printed objects with known ground truth geometry. Daniel Cabrini Hauagge, Scott Wehrwein, Kavita Bala, Noah Snavely |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Fitting procedural yarn models for realistic cloth renderingabstractFabrics play a significant role in many applications in design, prototyping, and entertainment. Recent fiber-based models capture the rich visual appearance of fabrics, but are too onerous to design and edit. Yarn-based procedural models are powerful and convenient, but too regular and not realistic enough in appearance. In this paper, we introduce an automatic fitting approach to create high-quality procedural yarn models of fabrics with fiber-level details. We fit CT data to procedural models to automatically recover a full range of parameters, and augment the models with a measurement-based model of flyaway fibers. We validate our fabric models against CT measurements and photographs, and demonstrate the utility of this approach for fabric modeling and editing. Fujun Luan, Kavita Bala |
ACM Trans. Graph. | 3 |
| 2015 | Shadow Detection and Sun Direction in Photo CollectionsabstractModeling the appearance of outdoor scenes from photo collections is challenging because of appearance variation, especially due to illumination. In this paper we present a simple and robust algorithm for estimating illumination properties-shadows and sun direction-from photo collections. These properties are key to a variety of scene modeling applications, including outdoor intrinsic images, realistic 3D scene rendering, and temporally varying (4D) reconstruction. Our shadow detection method uses illumination ratios to analyze lighting independent of camera effects, and determines shadow labels for each 3D point in a reconstruction. These shadow labels can then be used to detect shadow boundaries and estimate sun direction, as well as to compute dense shadow labels in pixel space. We demonstrate our method on large Internet photo collections of scenes, and show that it outperforms prior multi-image shadow detection and sun direction estimation methods. Scott Wehrwein, Kavita Bala, Noah Snavely |
3DV | 2 |
| 2015 | Material recognition in the wild with the Materials in Context DatabaseabstractRecognizing materials in real-world images is a challenging task. Real-world materials have rich surface texture, geometry, lighting conditions, and clutter, which combine to make the problem particularly difficult. In this paper, we introduce a new, large-scale, open dataset of materials in the wild, the Materials in Context Database (MINC), and combine this dataset with deep learning to achieve material recognition and segmentation of images in the wild. MINC is an order of magnitude larger than previous material databases, while being more diverse and well-sampled across its 23 categories. Using MINC, we train convolutional neural networks (CNNs) for two tasks: classifying materials from patches, and simultaneous material recognition and segmentation in full images. For patch-based classification on MINC we found that the best performing CNN architectures can achieve 85.2% mean class accuracy. We convert these trained CNN classifiers into an efficient fully convolutional framework combined with a fully connected conditional random field (CRF) to predict the material at every pixel in an image, achieving 73.1% mean class accuracy. Our experiments demonstrate that having a large, well-sampled dataset such as MINC is crucial for real-world material recognition and segmentation. Sean Bell, Paul Upchurch, Noah Snavely, Kavita Bala |
CVPR | 4 |
| 2015 | On the appearance of translucent edgesabstractEdges in images of translucent objects are very different from edges in images of opaque objects. The physical causes for these differences are hard to characterize analytically and are not well understood. This paper considers one class of translucency edges-those caused by a discontinuity in surface orientation-and describes the physical causes of their appearance. We simulate thousands of translucency edge profiles using many different scattering material parameters, and we explain the resulting variety of edge patterns by qualitatively analyzing light transport. We also discuss the existence of shape and material metamers, or combinations of distinct shape or material parameters that generate the same edge profile. This knowledge is relevant to visual inference tasks that involve translucent objects, such as shape or material estimation. Ioannis Gkioulekas, Bruce Walter, Edward H. Adelson, Kavita Bala, Todd E. Zickler |
CVPR | 4 |
| 2015 | Learning Visual Clothing Style with Heterogeneous Dyadic Co-OccurrencesabstractWith the rapid proliferation of smart mobile devices, users now take millions of photos every day. These include large numbers of clothing and accessory images. We would like to answer questions like 'What outfit goes well with this pair of shoes?' To answer these types of questions, one has to go beyond learning visual similarity and learn a visual notion of compatibility across categories. In this paper, we propose a novel learning framework to help answer these types of questions. The main idea of this framework is to learn a feature transformation from images of items into a latent space that expresses compatibility. For the feature transformation, we use a Siamese Convolutional Neural Network (CNN) architecture, where training examples are pairs of items that are either compatible or incompatible. We model compatibility based on co-occurrence in large-scale user behavior data, in particular co-purchase data from Amazon.com. To learn cross-category fit, we introduce a strategic method to sample training data, where pairs of items are heterogeneous dyads, i.e., the two elements of a pair belong to different high-level categories. While this approach is applicable to a wide variety of settings, we focus on the representative problem of learning compatible clothing style. Our results indicate that the proposed framework is capable of learning semantic information about visual style and is able to generate outfits of clothes, with items from different categories, that go well together. Andreas Veit, Balazs Kovacs, Sean Bell, Julian J. McAuley, Kavita Bala, Serge J. Belongie |
ICCV | 5 |
| 2015 | Talk abstract: Computational lighting design and band-sifting operatorsabstractIn this talk, I present two projects that are inspired by how photographers work. These projects are in collaboration with Ivo Boyadzhiev and Kavita Bala at Cornell University and Ted Adelson at MIT. Sylvain Paris, Ivaylo Boyadzhiev, Kavita Bala, Edward H. Adelson |
ICIP | 3 |
| 2015 | Computational rim illumination of dynamic subjects using aerial robotsabstractLighting plays a major role in photography. Professional photographers use elaborate installations to light their subjects and achieve sophisticated styles. However, lighting moving subjects performing dynamic tasks presents significant challenges and requires expensive manual intervention. A skilled additional assistant might be needed to reposition lights as the subject changes pose or moves, and the extra logistics significantly raises costs and time. The associated latencies as the assistant lights the subject, and the communication required from the photographer to achieve optimum lighting could mean missing a critical shot. We present a new approach to lighting dynamic subjects where an aerial robot equipped with a portable light source lights the subject to automatically achieve a desired lighting effect. We focus on rim lighting, a particularly challenging effect to achieve with dynamic subjects, and allow the photographer to specify a required rim width. Our algorithm processes the images from the photographer׳s camera and provides necessary motion commands to the aerial robot to achieve the desired rim width in the resulting photographs. With an indoor setup, we demonstrate a control approach that localizes the aerial robot with reference to the subject and tracks the subject to achieve the necessary motion. In addition to indoor experiments, we perform open-loop outdoor experiments in a realistic photo-shooting scenario to understand lighting ergonomics . Our proof-of-concept results demonstrate the utility of robots in computational lighting. Manohar Srikanth, Kavita Bala, Frédo Durand |
Comput. Graph. | 2 |
| 2015 | Learning visual similarity for product design with convolutional neural networksabstractPopular sites like Houzz, Pinterest, and LikeThatDecor, have communities of users helping each other answer questions about products in images. In this paper we learn an embedding for visual search in interior design. Our embedding contains two different domains of product images: products cropped from internet scenes, and products in their iconic form. With such a multi-domain embedding, we demonstrate several applications of visual search including identifying products in scenes and finding stylistically similar products. To obtain the embedding, we train a convolutional neural network on pairs of images. We explore several training architectures including re-purposing object classifiers, using siamese networks, and using multitask learning. We evaluate our search quantitatively and qualitatively and demonstrate high quality results for search across multiple visual domains, enabling new applications in interior design. Sean Bell, Kavita Bala |
ACM Trans. Graph. | 2 |
| 2015 | Band-Sifting Decomposition for Image-Based Material EditingabstractPhotographers often “prep” their subjects to achieve various effects; for example, toning down overly shiny skin, covering blotches, etc. Making such adjustments digitally after a shoot is possible, but difficult without good tools and good skills. Making such adjustments to video footage is harder still. We describe and study a set of 2D image operations, based on multiscale image analysis, that are easy and straightforward and that can consistently modify perceived material properties. These operators first build a subband decomposition of the image and then selectively modify the coefficients within the subbands. We call this selection process band sifting . We show that different siftings of the coefficients can be used to modify the appearance of properties such as gloss, smoothness, pigmentation, or weathering. The band-sifting operators have particularly striking effects when applied to faces; they can provide “knobs” to make a face look wetter or drier, younger or older, and with heavy or light variation in pigmentation. Through user studies, we identify a set of operators that yield consistent subjective effects for a variety of materials and scenes. We demonstrate that these operators are also useful for processing video sequences. Ivaylo Boyadzhiev, Kavita Bala, Sylvain Paris, Edward H. Adelson |
ACM Trans. Graph. | 2 |
| 2015 | Matching Real Fabrics with Micro-Appearance ModelsabstractMicro-appearance models explicitly model the interaction of light with microgeometry at the fiber scale to produce realistic appearance. To effectively match them to real fabrics, we introduce a new appearance matching framework to determine their parameters. Given a micro-appearance model and photographs of the fabric under many different lighting conditions, we optimize for parameters that best match the photographs using a method based on calculating derivatives during rendering. This highly applicable framework, we believe, is a useful research tool because it simplifies development and testing of new models. Using the framework, we systematically compare several types of micro-appearance models. We acquired computed microtomography (micro CT) scans of several fabrics, photographed the fabrics under many viewing/illumination conditions, and matched several appearance models to this data. We compare a new fiber-based light scattering model to the previously used microflake model. We also compare representing cloth microgeometry using volumes derived directly from the micro CT data to using explicit fibers reconstructed from the volumes. From our comparisons, we make the following conclusions: (1) given a fiber-based scattering model, volume- and fiber-based microgeometry representations are capable of very similar quality, and (2) using a fiber-specific scattering model is crucial to good results as it achieves considerably higher accuracy than prior work. Pramook Khungurn, Daniel Schroeder, Kavita Bala, Steve Marschner |
ACM Trans. Graph. | 4 |
| 2014 | Reasoning about Photo Collections using Models of Outdoor Illumination
Daniel Cabrini Hauagge, Scott Wehrwein, Paul Upchurch, Kavita Bala, Noah Snavely |
BMVC | 4 |
| 2014 | High-order similarity relations in radiative transferabstractRadiative transfer equations (RTEs) with different scattering parameters can lead to identical solution radiance fields. Similarity theory studies this effect by introducing a hierarchy of equivalence relations called "similarity relations". Unfortunately, given a set of scattering parameters, it remains unclear how to find altered ones satisfying these relations, significantly limiting the theory's practical value. This paper presents a complete exposition of similarity theory, which provides fundamental insights into the structure of the RTE's parameter space. To utilize the theory in its general high-order form, we introduce a new approach to solve for the altered parameters including the absorption and scattering coefficients as well as a fully tabulated phase function. We demonstrate the practical utility of our work using two applications: forward and inverse rendering of translucent media. Forward rendering is our main application, and we develop an algorithm exploiting similarity relations to offer "free" speedups for Monte Carlo rendering of optically dense and forward-scattering materials. For inverse rendering, we propose a proof-of-concept approach which warps the parameter space and greatly improves the efficiency of gradient descent algorithms. We believe similarity theory is important for simulating and acquiring volume-based appearance, and our approach has the potential to benefit a wide range of future applications in this area. Ravi Ramamoorthi, Kavita Bala |
ACM Trans. Graph. | 3 |
| 2014 | Automatic shader simplification using surface signal approximationabstractIn this paper, we present a new automatic shader simplification method using surface signal approximation. We regard the entire multi-stage rendering pipeline as a process that generates signals on surfaces, and we formulate the simplification of the fragment shader as a global simplification problem across multi-shader stages. Three new shader simplification rules are proposed to solve the problem. First, the code transformation rule transforms fragment shader code to other shader stages in order to redistribute computations on pixels up to the level of geometry primitives. Second, the surface-wise approximation rule uses high-order polynomial basis functions on surfaces to approximate pixel-wise computations in the fragment shader. These approximations are pre-cached and simplify computations at runtime. Third, the surface subdivision rule tessellates surfaces into smaller patches. It combines with the previous two rules to approximate pixel-wise signals at different levels of tessellations with different computation times and visual errors. To evaluate simplified shaders using these simplification rules, we introduce a new cost model that includes the visual quality, rendering time and memory consumption. With these simplification rules and the cost model, we present an integrated shader simplification algorithm that is capable of automatically generating variants of simplified shaders and selecting a sequence of preferable shaders. Results show that the sequence of selected simplified shaders balance performance, accuracy and memory consumption well. Rui Wang 0004, Xianjin Yang, Yazhen Yuan, Wei Chen 0001, Kavita Bala, Hujun Bao |
ACM Trans. Graph. | 5 |
| 2014 | Intrinsic images in the wildabstractIntrinsic image decomposition separates an image into a reflectance layer and a shading layer. Automatic intrinsic image decomposition remains a significant challenge, particularly for real-world scenes. Advances on this longstanding problem have been spurred by public datasets of ground truth data, such as the MIT Intrinsic Images dataset. However, the difficulty of acquiring ground truth data has meant that such datasets cover a small range of materials and objects. In contrast, real-world scenes contain a rich range of shapes and materials, lit by complex illumination. In this paper we introduce Intrinsic Images in the Wild , a large-scale, public dataset for evaluating intrinsic image decompositions of indoor scenes. We create this benchmark through millions of crowdsourced annotations of relative comparisons of material properties at pairs of points in each scene. Crowdsourcing enables a scalable approach to acquiring a large database, and uses the ability of humans to judge material comparisons, despite variations in illumination. Given our database, we develop a dense CRF-based intrinsic image algorithm for images in the wild that outperforms a range of state-of-the-art intrinsic image algorithms. Intrinsic image decomposition remains a challenging problem; we release our code and database publicly to support future research on this problem, available online at http://intrinsic.cs.cornell.edu/. Sean Bell, Kavita Bala, Noah Snavely |
ACM Trans. Graph. | 2 |
| 2014 | A Local Frequency Analysis of Light Scattering and AbsorptionabstractRendering participating media requires significant computation, but the effect of volumetric scattering is often eventually smooth. This article proposes an innovative analysis of absorption and scattering of local light fields in the Fourier domain and derives the corresponding set of operators on the covariance matrix of the power spectrum of the light field. This analysis brings an efficient prediction tool for the behavior of light along a light path in participating media. We leverage this analysis to derive proper frequency prediction metrics in 3D by combining per-light path information in the volume. We demonstrate the use of these metrics to significantly improve the convergence of a variety of existing methods for the simulation of multiple scattering in participating media. First, we propose an efficient computation of second derivatives of the fluence, to be used in methods like irradiance caching. Second, we derive proper filters and adaptive sample densities for image-space adaptive sampling and reconstruction. Third, we propose an adaptive sampling for the integration of scattered illumination to the camera. Finally, we improve the convergence of progressive photon beams by predicting where the radius of light gathering can stop decreasing. Light paths in participating media can be very complex. Our key contribution is to show that analyzing local light fields in the Fourier domain reveals the consistency of illumination in such media and provides a set of simple and useful rules to be used to accelerate existing global illumination methods. Laurent Belcour, Kavita Bala, Cyril Soler |
ACM Trans. Graph. | 2 |
| 2013 | Photometric Ambient OcclusionabstractWe present a method for computing ambient occlusion (AO) for a stack of images of a scene from a fixed viewpoint. Ambient occlusion, a concept common in computer graphics, characterizes the local visibility at a point: it approximates how much light can reach that point from different directions without getting blocked by other geometry. While AO has received surprisingly little attention in vision, we show that it can be approximated using simple, per-pixel statistics over image stacks, based on a simplified image formation model. We use our derived AO measure to compute reflectance and illumination for objects without relying on additional smoothness priors, and demonstrate state-of-the art performance on the MIT Intrinsic Images benchmark. We also demonstrate our method on several synthetic and real scenes, including 3D printed objects with known ground truth geometry. Daniel Cabrini Hauagge, Scott Wehrwein, Kavita Bala, Noah Snavely |
CVPR | 3 |
| 2013 | OpenSurfaces: a richly annotated catalog of surface appearanceabstractThe appearance of surfaces in real-world scenes is determined by the materials, textures, and context in which the surfaces appear. However, the datasets we have for visualizing and modeling rich surface appearance in context, in applications such as home remodeling, are quite limited. To help address this need, we present OpenSurfaces, a rich, labeled database consisting of thousands of examples of surfaces segmented from consumer photographs of interiors, and annotated with material parameters (reflectance, material names), texture information (surface normals, rectified textures), and contextual information (scene category, and object names). Retrieving usable surface information from uncalibrated Internet photo collections is challenging. We use human annotations and present a new methodology for segmenting and annotating materials in Internet photo collections suitable for crowdsourcing (e.g., through Amazon's Mechanical Turk). Because of the noise and variability inherent in Internet photos and novice annotators, designing this annotation engine was a key challenge; we present a multi-stage set of annotation tasks with quality checks and validation. We demonstrate the use of this database in proof-of-concept applications including surface retexturing and material and image browsing, and discuss future uses. OpenSurfaces is a public resource available at http://opensurfaces.cs.cornell.edu/. Sean Bell, Paul Upchurch, Noah Snavely, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2013 | User-assisted image compositing for photographic lightingabstractGood lighting is crucial in photography and can make the difference between a great picture and a discarded image. Traditionally, professional photographers work in a studio with many light sources carefully set up, with the goal of getting a near-final image at exposure time, with post-processing mostly focusing on aspects orthogonal to lighting. Recently, a new workflow has emerged for architectural and commercial photography, where photographers capture several photos from a fixed viewpoint with a moving light source. The objective is not to produce the final result immediately, but rather to capture useful data that are later processed, often significantly, in photo editing software to create the final well-lit image. This new workflow is flexible, requires less manual setup, and works well for time-constrained shots. But dealing with several tens of unorganized layers is painstaking, requiring hours to days of manual effort, as well as advanced photo editing skills. Our objective in this paper is to make the compositing step easier. We describe a set of optimizations to assemble the input images to create a fewbasis lightsthat correspond to common goals pursued by photographers, e.g., accentuating edges and curved regions. We also introducemodifiersthat capture standard photographic tasks, e.g., to alter the lights to soften highlights and shadows, akin to umbrellas and soft boxes. Our experiments with novice and professional users show that our approach allows them to quickly create satisfying results, whereas working with unorganized images requires considerably more time. Casual users particularly benefit from our approach since coping with a large number of layers is daunting for them and requires significant experience. Ivaylo Boyadzhiev, Sylvain Paris, Kavita Bala |
ACM Trans. Graph. | 3 |
| 2013 | Understanding the role of phase function in translucent appearanceabstractMultiple scattering contributes critically to the characteristic translucent appearance of food, liquids, skin, and crystals; but little is known about how it is perceived by human observers. This article explores the perception of translucency by studying the image effects of variations in one factor of multiple scattering: the phase function. We consider an expanded space of phase functions created by linear combinations of Henyey-Greenstein and von Mises-Fisher lobes, and we study this physical parameter space using computational data analysis and psychophysics. Our study identifies a two-dimensional embedding of the physical scattering parameters in a perceptually meaningful appearance space. Through our analysis of this space, we find uniform parameterizations of its two axes by analytical expressions of moments of the phase function, and provide an intuitive characterization of the visual effects that can be achieved at different parts of it. We show that our expansion of the space of phase functions enlarges the range of achievable translucent appearance compared to traditional single-parameter phase function models. Our findings highlight the important role phase function can have in controlling translucent appearance, and provide tools for manipulating its effect in material design applications. Ioannis Gkioulekas, Bei Xiao, Edward H. Adelson, Todd E. Zickler, Kavita Bala |
ACM Trans. Graph. | 6 |
| 2013 | Inverse volume rendering with material dictionariesabstractTranslucent materials are ubiquitous, and simulating their appearance requires accurate physical parameters. However, physically-accurate parameters for scattering materials are difficult to acquire. We introduce an optimization framework for measuring bulk scattering properties of homogeneous materials (phase function, scattering coefficient, and absorption coefficient) that is more accurate, and more applicable to a broad range of materials. The optimization combines stochastic gradient descent with Monte Carlo rendering and a material dictionary to invert the radiative transfer equation. It offers several advantages: (1) it does not require isolating single-scattering events; (2) it allows measuring solids and liquids that are hard to dilute; (3) it returns parameters in physically-meaningful units; and (4) it does not restrict the shape of the phase function using Henyey-Greenstein or any other low-parameter model. We evaluate our approach by creating an acquisition setup that collects images of a material slab under narrow-beam RGB illumination. We validate results by measuring prescribed nano-dispersions and showing that recovered parameters match those predicted by Lorenz-Mie theory. We also provide a table of RGB scattering parameters for some common liquids and solids, which are validated by simulating color images in novel geometric configurations that match the corresponding photographs with less than 5% error. Ioannis Gkioulekas, Kavita Bala, Todd E. Zickler, Anat Levin |
ACM Trans. Graph. | 3 |
| 2013 | Modular flux transfer: efficient rendering of high-resolution volumes with repeated structuresabstractThe highest fidelity images to date of complex materials like cloth use extremely high-resolution volumetric models. However, rendering such complex volumetric media is expensive, with brute-force path tracing often the only viable solution. Fortunately, common volumetric materials (fabrics, finished wood, synthesized solid textures) are structured, with repeated patterns approximated by tiling a small number of exemplar blocks. In this paper, we introduce a precomputation-based rendering approach for such volumetric media with repeated structures based on a modular transfer formulation. We model each exemplar block as a voxel grid and precompute voxel-to-voxel, patch-to-patch, and patch-to-voxel flux transfer matrices. At render time, when blocks are tiled to produce a high-resolution volume, we accurately compute low-order scattering, with modular flux transfer used to approximate higher-order scattering. We achieve speedups of up to 12× over path tracing on extremely complex volumes, with minimal loss of quality. In addition, we demonstrate that our approach outperforms photon mapping on these materials. Milos Hasan, Ravi Ramamoorthi, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2012 | Crowd Light: Evaluating the Perceived Fidelity of Illuminated Dynamic ScenesabstractAbstract Rendering realistic illumination effects for complex animated scenes with many dynamic objects or characters is computationally expensive. Yet, it is not obvious how important such accurate lighting is for the overall perceived realism in these scenes. In this paper, we present a methodology to evaluate the perceived fidelity of illumination in scenes with dynamic aggregates, such as crowds, and explore several factors which may affect this perception. We focus in particular on evaluating how a popular spherical harmonics lighting method can be used to approximate realistic lighting of crowds. We conduct a series of psychophysical experiments to explore how a simple approach to approximating global illumination, using interpolation in the temporal domain, affects the perceived fidelity of dynamic scenes with high geometric, motion, and illumination complexity. We show that the complexity of the geometry and temporal properties of the crowd entities, the motion of the aggregate as a whole, the type of interpolation (i.e., of the direct and/or indirect illumination coefficients), and the presence or absence of colour all affect perceived fidelity. We show that high (i.e., above 75%) levels of perceived scene fidelity can be maintained while interpolating indirect illumination for intervals of up to 30 frames, resulting in a greater than three‐fold rendering speed‐up. Adrián Jarabo, Tom Van Eyck, Veronica Sundstedt, Kavita Bala, Diego Gutierrez, Carol O'Sullivan |
Comput. Graph. Forum | 4 |
| 2012 | User-guided white balance for mixed lighting conditionsabstractProper white balance is essential in photographs to eliminate color casts due to illumination. The single-light case is hard to solve automatically but relatively easy for humans. Unfortunately, many scenes contain multiple light sources such as an indoor scene with a window, or when a flash is used in a tungsten-lit room. The light color can then vary on a per-pixel basis and the problem becomes challenging at best, even with advanced image editing tools. We propose a solution to the ill-posed mixed light white balance problem, based on user guidance. Users scribble on a few regions that should have the same color, indicate one or more regions of neutral color, and select regions where the current color looks correct. We first expand the provided scribble groups to more regions using pixel similarity and a robust voting scheme. We formulate the spatially varying white balance problem as a sparse data interpolation problem in which the user scribbles and their extensions form constraints. We demonstrate that our approach can produce satisfying results on a variety of scenes with intuitive scribbles and without any knowledge about the lights. Ivaylo Boyadzhiev, Kavita Bala, Sylvain Paris, Frédo Durand |
ACM Trans. Graph. | 2 |
| 2012 | Bidirectional lightcutsabstractScenes modeling the real-world combine a wide variety of phenomena including glossy materials, detailed heterogeneous anisotropic media, subsurface scattering, and complex illumination. Predictive rendering of such scenes is difficult; unbiased algorithms are typically too slow or too noisy. Virtual point light (VPL) based algorithms produce low noise results across a wide range of performance/accuracy tradeoffs, from interactive rendering to high quality offline rendering, but their bias means that locally important illumination features may be missing. We introduce a bidirectional formulation and a set of weighting strategies to significantly reduce the bias in VPL-based rendering algorithms. Our approach, bidirectional lightcuts , maintains the scalability and low noise global illumination advantages of prior VPL-based work, while significantly extending their generality to support a wider range of important materials and visual cues. We demonstrate scalable, efficient, and low noise rendering of scenes with highly complex materials including gloss, BSSRDFs, and anisotropic volumetric models. Bruce Walter, Pramook Khungurn, Kavita Bala |
ACM Trans. Graph. | 3 |
| 2012 | Structure-aware synthesis for predictive woven fabric appearanceabstractWoven fabrics have a wide range of appearance determined by their small-scale 3D structure. Accurately modeling this structural detail can produce highly realistic renderings of fabrics and is critical for predictive rendering of fabric appearance. But building these yarn-level volumetric models is challenging. Procedural techniques are manually intensive, and fail to capture the naturally arising irregularities which contribute significantly to the overall appearance of cloth. Techniques that acquire the detailed 3D structure of real fabric samples are constrained only to model the scanned samples and cannot represent different fabric designs. This paper presents a new approach to creating volumetric models of woven cloth, which starts with user-specified fabric designs and produces models that correctly capture the yarn-level structural details of cloth. We create a small database of volumetric exemplars by scanning fabric samples with simple weave structures. To build an output model, our method synthesizes a new volume by copying data from the exemplars at each yarn crossing to match a weave pattern that specifies the desired output structure. Our results demonstrate that our approach generalizes well to complex designs and can produce highly realistic results at both large and small scales. Wenzel Jakob, Steve Marschner, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2011 | Single view reflectance capture using multiplexed scattering and time-of-flight imagingabstractThis paper introduces the concept of time-of-flight reflectance estimation, and demonstrates a new technique that allows a camera to rapidly acquire reflectance properties of objects from a single view-point, over relatively long distances and without encircling equipment. We measure material properties by indirectly illuminating an object by a laser source, and observing its reflected light indirectly using a time-of-flight camera. The configuration collectively acquires dense angular, but low spatial sampling, within a limited solid angle range - all from a single viewpoint. Our ultra-fast imaging approach captures space-time "streak images" that can separate out different bounces of light based on path length. Entanglements arise in the streak images mixing signals from multiple paths if they have the same total path length. We show how reflectances can be recovered by solving for a linear system of equations and assuming parametric material models; fitting to lower dimensional reflectance models enables us to disentangle measurements. We demonstrate proof-of-concept results of parametric reflectance models for homogeneous and discretized heterogeneous patches, both using simulation and experimental hardware. As compared to lengthy or highly calibrated BRDF acquisition techniques, we demonstrate a device that can rapidly, on the order of seconds, capture meaningful reflectance information. We expect hardware advances to improve the portability and speed of this device. Nikhil Naik 0003, Andreas Velten, Ramesh Raskar, Kavita Bala |
ACM Trans. Graph. | 5 |
| 2011 | Building volumetric appearance models of fabric using micro CT imagingabstractThe appearance of complex, thick materials like textiles is determined by their 3D structure, and they are incompletely described by surface reflection models alone. While volume scattering can produce highly realistic images of such materials, creating the required volume density models is difficult. Procedural approaches require significant programmer effort and intuition to design specialpurpose algorithms for each material. Further, the resulting models lack the visual complexity of real materials with their naturally-arising irregularities. This paper proposes a new approach to acquiring volume models, based on density data from X-ray computed tomography (CT) scans and appearance data from photographs under uncontrolled illumination. To model a material, a CT scan is made, resulting in a scalar density volume. This 3D data is processed to extract orientation information and remove noise. The resulting density and orientation fields are used in an appearance matching procedure to define scattering properties in the volume that, when rendered, produce images with texture statistics that match the photographs. As our results show, this approach can easily produce volume appearance models with extreme detail, and at larger scales the distinctive textures and highlights of a range of very different fabrics like satin and velvet emerge automatically---all based simply on having accurate mesoscale geometry. Wenzel Jakob, Steve Marschner, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2011 | Heterogeneous Subsurface Scattering Using the Finite Element MethodabstractMaterials with visually important heterogeneous subsurface scattering, including marble, skin, leaves, and minerals are common in the real world. However, general, accurate, and efficient rendering of these materials is an open problem. In this paper, we describe a finite element (FE) solution of the heterogeneous diffusion equation (DE) that solves this problem. Our algorithm is the first to use the FE method to solve the difficult problem of heterogeneous subsurface rendering. To create our algorithm, we make two contributions. First, we correct previous work and derive an accurate and complete heterogeneous diffusion formulation with two key elements: the diffusive source boundary condition (DSBC)-an accurate model of the reduced intensity (RI) source-and its associated render query function. Second, we solve this formulation accurately and efficiently using the FE method. With these contributions, we can render subsurface scattering with a simple four step algorithm. To demonstrate that our algorithm is simultaneously general, accurate, and efficient, we test its performance on a series of difficult scenes. For a wide range of materials and geometry, it produces, in minutes, images that match path traced references, that required hours. Adam Arbree, Bruce Walter, Kavita Bala |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | Combining global and local virtual lights for detailed glossy illuminationabstractAccurately rendering glossy materials in design applications, where previewing and interactivity are important, remains a major challenge. While many fast global illumination solutions have been proposed, all of them work under limiting assumptions on the materials and lighting in the scene. In the presence of many glossy (directionally scattering) materials, fast solutions either fail or degenerate to inefficient, brute-force simulations of the underlying light transport. In particular, many-light algorithms are able to provide fast approximations by clamping elements of the light transport matrix, but they eliminate the part of the transport that contributes to accurate glossy appearance. In this paper we introduce a solution that separately solves for the global (low-rank, dense) and local (highrank, sparse) illumination components. For the low-rank component we introduce visibility clustering and approximation, while for the high-rank component we introduce a local light technique to correct for the missing illumination. Compared to competing techniques we achieve superior gloss rendering in minutes, making our technique suitable for applications such as industrial design and architecture, where material appearance is critical. Tomás Davidovic, Jaroslav Krivánek, Milos Hasan, Philipp Slusallek, Kavita Bala |
ACM Trans. Graph. | 5 |
| 2010 | A radiative transfer framework for rendering materials with anisotropic structureabstractThe radiative transfer framework that underlies all current rendering of volumes is limited to scattering media whose properties are invariant to rotation. Many systems allow for "anisotropic scattering," in the sense that scattered intensity depends on the scattering angle, but the standard equation assumes that the structure of the medium is isotropic. This limitation impedes physics-based rendering of volume models of cloth, hair, skin, and other important volumetric or translucent materials that do have anisotropic structure. This paper presents an end-to-end formulation of physics-based volume rendering of anisotropic scattering structures, allowing these materials to become full participants in global illumination simulations. We begin with a generalized radiative transfer equation, derived from scattering by oriented non-spherical particles. Within this framework, we propose a new volume scattering model analogous to the well-known family of microfacet surface reflection models; we derive an anisotropic diffusion approximation, including the weak form required for finite element solution and a way to compute the diffusion matrix from the parameters of the scattering model; and we also derive a new anisotropic dipole BSSRDF for anisotropic translucent materials. We demonstrate results from Monte Carlo, finite element, and dipole simulations. All these contributions are readily implemented in existing rendering systems for volumes and translucent materials, and they all reduce to the standard practice in the isotropic case. Wenzel Jakob, Adam Arbree, Jonathan T. Moon, Kavita Bala, Steve Marschner |
ACM Trans. Graph. | 4 |
| 2010 | Effects of global illumination approximations on material appearanceabstractRendering applications in design, manufacturing, ecommerce and other fields are used to simulate the appearance of objects and scenes. Fidelity with respect to appearance is often critical, and calculating global illumination (GI) is an important contributor to image fidelity; but it is expensive to compute. GI approximation methods, such as virtual point light (VPL) algorithms, are efficient, but they can induce image artifacts and distortions of object appearance. In this paper we systematically study the perceptual effects on image quality and material appearance of global illumination approximations made by VPL algorithms. In a series of psychophysical experiments we investigate the relationships between rendering parameters, object properties and image fidelity in a VPL renderer. Using the results of these experiments we analyze how VPL counts and energy clamping levels affect the visibility of image artifacts and distortions of material appearance, and show how object geometry and material properties modulate these effects. We find the ranges of these parameters that produce VPL renderings that are visually equivalent to reference renderings. Further we identify classes of shapes and materials that cannot be accurately rendered using VPL methods with limited resources. Using these findings we propose simple heuristics to guide visually equivalent and efficient rendering, and present a method for correcting energy losses in VPL renderings. This work provides a strong perceptual foundation for a popular and efficient class of GI algorithms. Jaroslav Krivánek, James A. Ferwerda, Kavita Bala |
ACM Trans. Graph. | 3 |
| 2009 | Virtual spherical lights for many-light rendering of glossy scenesabstractIn this paper, we aim to lift the accuracy limitations of many-light algorithms by introducing a new light type, the virtual spherical light (VSL). The illumination contribution of a VSL is computed over a non-zero solid angle, thus eliminating the illumination spikes that virtual point lights used in traditional many-light methods are notorious for. The VSL enables application of many-light approaches in scenes with glossy materials and complex illumination that could previously be rendered only by much slower algorithms. By combining VSLs with the matrix row-column sampling algorithm, we achieve high-quality images in one to four minutes, even in scenes where path tracing or photon mapping take hours to converge. Milos Hasan, Jaroslav Krivánek, Bruce Walter, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2009 | Automatic bounding of programmable shaders for efficient global illuminationabstractThis paper describes a technique to automatically adapt programmable shaders for use in physically-based rendering algorithms. Programmable shading provides great flexibility and power for creating rich local material detail, but only allows the material to be queried in one limited way: point sampling. Physically-based rendering algorithms simulate the complex global flow of light through an environment but rely on higher level information about the material properties, such as importance sampling and bounding, to intelligently solve high dimensional rendering integrals. We propose using a compiler to automatically generate interval versions of programmable shaders that can be used to provide the higher level query functions needed by physically-based rendering without the need for user intervention or expertise. We demonstrate the use of programmable shaders in two such algorithms, multidimensional lightcuts and photon mapping, for a wide range of scenes including complex geometry, materials and lighting. Edgar Velázquez-Armendáriz, Milos Hasan, Bruce Walter, Kavita Bala |
ACM Trans. Graph. | 5 |
| 2009 | Single scattering in refractive media with triangle mesh boundariesabstractLight scattering in refractive media is an important optical phenomenon for computer graphics. While recent research has focused on multiple scattering, there has been less work on accurate solutions for single or low-order scattering. Refraction through a complex boundary allows a single external source to be visible in multiple directions internally with different strengths; these are hard to find with existing techniques. This paper presents techniques to quickly find paths that connect points inside and outside a medium while obeying the laws of refraction. We introduce: a half-vector based formulation to support the most common geometric representation, triangles with interpolated normals; hierarchical pruning to scale to triangular meshes; and, both a solver with strong accuracy guarantees, and a faster method that is empirically accurate. A GPU version achieves interactive frame rates in several examples. Bruce Walter, Nicolas Holzschuch, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2008 | Optimistic parallelism benefits from data partitioningabstractRecent studies of irregular applications such as finite-element mesh generators and data-clustering codes have shown that these applications have a generalized data parallelism arising from the use of iterative algorithms that perform computations on elements of worklists. In some irregular applications, the computations on different elements are independent. In other applications, there may be complex patterns of dependences between these computations. Milind Kulkarni 0001, Keshav Pingali, Ganesh Ramanarayanan, Bruce Walter, Kavita Bala, L. Paul Chew |
ASPLOS | 5 |
| 2008 | Scheduling strategies for optimistic parallel execution of irregular programsabstractRecent application studies have shown that many irregular applications have a generalized data parallelism that manifests itself as iterative computations over worklists of different kinds. In general, there are complex dependencies between iterations. These dependencies cannot be elucidated statically because they depend on the inputs to the program; thus, optimistic parallel execution is the only tractable approach to parallelizing these applications. Milind Kulkarni 0001, Patrick Carribault, Keshav Pingali, Ganesh Ramanarayanan, Bruce Walter, Kavita Bala, L. Paul Chew |
SPAA | 6 |
| 2008 | Single-pass Scalable Subsurface Rendering with LightcutsabstractAbstract This paper presents a new, scalable, single pass algorithm for computing subsurface scattering using the diffusion approximation. Instead of pre‐computing a globally conservative estimate of the surface irradiance like previous two pass methods, the algorithm simultaneously refines hierarchical and adaptive estimates of both the surface irradiance and the subsurface transport. By using an adaptive, top‐down refinement method, the algorithm directs computational effort only to simulating those eye‐surface‐light paths that make significant contributions to the final image. Because the algorithm is driven by image importance, it scales more efficiently than previous methods that have a linear dependence on translucent surface area. We demonstrate that in scenes with many translucent objects and in complex lighting environments, our new algorithm has a significant performance advantage. Adam Arbree, Bruce Walter, Kavita Bala |
Comput. Graph. Forum | 3 |
| 2008 | Tensor Clustering for Rendering Many-Light AnimationsabstractAbstract Rendering animations of scenes with deformable objects, camera motion, and complex illumination, including indirect lighting and arbitrary shading, is a long‐standing challenge. Prior work has shown that complex lighting can be accurately approximated by a large collection of point lights. In this formulation, rendering of animation sequences becomes the problem of efficiently shading many surface samples from many lights across several frames. This paper presents a tensor formulation of the animated many‐light problem, where each element of the tensor expresses the contribution of one light to one pixel in one frame. We sparsely sample rows and columns of the tensor, and introduce a clustering algorithm to select a small number of representative lights to efficiently approximate the animation. Our algorithm achieves efficiency by reusing representatives across frames, while minimizing temporal flicker. We demonstrate our algorithm in a variety of scenes that include deformable objects, complex illumination and arbitrary shading and show that a surprisingly small number of representative lights is sufficient for high quality rendering. We believe out algorithm will find practical use in applications that require fast previews of complex animation. Milos Hasan, Edgar Velázquez-Armendáriz, Fabio Pellacini, Kavita Bala |
Comput. Graph. Forum | 4 |
| 2008 | Perception of complex aggregatesabstractAggregates of individual objects, such as forests, crowds, and piles of fruit, are a common source of complexity in computer graphics scenes. When viewing an aggregate, observers attend less to individual objects and focus more on overall properties such as numerosity, variety, and arrangement. Paradoxically, rendering and modeling costs increase with aggregate complexity, exactly when observers are attending less to individual objects. In this paper we take some first steps to characterize the limits of visual coding of aggregates to efficiently represent their appearance in scenes. We describe psychophysical experiments that explore the roles played by the geometric and material properties of individual objects in observers' abilities to discriminate different aggregate collections. Based on these experiments we derive metrics to predict when two aggregates have the same appearance, even when composed of different objects. In a follow-up experiment we confirm that these metrics can be used to predict the appearance of a range of realistic aggregates. Finally, as a proof-of-concept we show how these new aggregate perception metrics can be applied to simplify scenes by allowing substitution of geometrically simpler aggregates for more complex ones without changing appearance. Ganesh Ramanarayanan, Kavita Bala, James A. Ferwerda |
ACM Trans. Graph. | 2 |
| 2007 | Optimistic parallelism requires abstractionsabstractIrregular applications, which manipulate large, pointer-based data structures like graphs, are difficult to parallelize manually. Automatic tools and techniques such as restructuring compilers and run-time speculative execution have failed to uncover much parallelism in these applications, in spite of a lot of effort by the research community. These difficulties have even led some researchers to wonder if there is any coarse-grain parallelism worth exploiting in irregular applications. Milind Kulkarni 0001, Keshav Pingali, Bruce Walter, Ganesh Ramanarayanan, Kavita Bala, L. Paul Chew |
PLDI | 5 |
| 2007 | Matrix row-column sampling for the many-light problemabstractRendering complex scenes with indirect illumination, high dynamic range environment lighting, and many direct light sources remains a challenging problem. Prior work has shown that all these effects can be approximated by many point lights. This paper presents a scalable solution to the many-light problem suitable for a GPU implementation. We view the problem as a large matrix of sample-light interactions; the ideal final image is the sum of the matrix columns. We propose an algorithm for approximating this sum by sampling entire rows and columns of the matrix on the GPU using shadow mapping. The key observation is that the inherent structure of the transfer matrix can be revealed by sampling just a small number of rows and columns. Our prototype implementation can compute the light transfer within a few seconds for scenes with indirect and environment illumination, area lights, complex geometry and arbitrary shaders. We believe this approach can be very useful for rapid previewing in applications like cinematic and architectural lighting design. Milos Hasan, Fabio Pellacini, Kavita Bala |
ACM Trans. Graph. | 3 |
| 2007 | Visual equivalence: towards a new standard for image fidelityabstractEfficient, realistic rendering of complex scenes is one of the grand challenges in computer graphics. Perceptually based rendering addresses this challenge by taking advantage of the limits of human vision. However, existing methods, based on predicting visible image differences, are too conservative because some kinds of image differences do not matter to human observers. In this paper, we introduce the concept of visual equivalence , a new standard for image fidelity in graphics. Images are visually equivalent if they convey the same impressions of scene appearance, even if they are visibly different. To understand this phenomenon, we conduct a series of experiments that explore how object geometry, material, and illumination interact to provide information about appearance, and we characterize how two kinds of transformations on illumination maps (blurring and warping) affect these appearance attributes. We then derive visual equivalence predictors (VEPs): metrics for predicting when images rendered with transformed illumination maps will be visually equivalent to images rendered with reference maps. We also run a confirmatory study to validate the effectiveness of these VEPs for general scenes. Finally, we show how VEPs can be used to improve the efficiency of two rendering algorithms: Light-cuts and precomputed radiance transfer. This work represents some promising first steps towards developing perceptual metrics based on higher order aspects of visual coding. Ganesh Ramanarayanan, James A. Ferwerda, Bruce Walter, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2007 | Constrained Texture Synthesis via Energy MinimizationabstractThis paper describes CMS (constrained minimization synthesis), a fast, robust texture synthesis algorithm that creates output textures while satisfying constraints. We show that constrained texture synthesis can be posed in a principled way as an energy minimization problem that requires balancing two measures of quality: constraint satisfaction and texture seamlessness. We then present an efficient algorithm for finding good solutions to this problem using an adaptation of graphcut energy minimization. CMS is particularly well suited to detail synthesis, the process of adding high-resolution detail to low-resolution images. It also supports the full Image Analogies framework, while providing superior image quality and performance. CMS is easily extended to handle multiple constraints on a single output, thus enabling novel applications that combine both user-specified and image-based control. Ganesh Ramanarayanan, Kavita Bala |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2006 | Implementing the render cache and the edge-and-point image on graphics hardware
Edgar Velázquez-Armendáriz, Kavita Bala, Bruce Walter |
Graphics Interface | 3 |
| 2006 | Direct-to-indirect transfer for cinematic relightingabstractThis paper presents an interactive GPU-based system for cinematic relighting with multiple-bounce indirect illumination from a fixed view-point. We use a deep frame-buffer containing a set of view samples, whose indirect illumination is recomputed from the direct illumination on a large set of gather samples, distributed around the scene. This direct-to-indirect transfer is a linear transform which is particularly large, given the size of the view and gather sets. This makes it hard to precompute, store and multiply with. We address this problem by representing the transform as a set of sparse matrices encoded in wavelet space. A hierarchical construction is used to impose a wavelet basis on the unstructured gather cloud, and an image-based approach is used to map the sparse matrix computations to the GPU. We precompute the transfer matrices using a hierarchical algorithm and a variation of photon mapping in less than three hours on one processor. We achieve high-quality indirect illumination at 10-20 frames per second for complex scenes with over 2 million polygons, with diffuse and glossy materials, and arbitrary direct lighting models (expressed using shaders). We compute per-pixel indirect illumination without the need of irradiance caching or other subsampling techniques. Milos Hasan, Fabio Pellacini, Kavita Bala |
ACM Trans. Graph. | 3 |
| 2006 | Multidimensional lightcutsabstractMultidimensional lightcuts is a new scalable method for efficiently rendering rich visual effects such as motion blur, participating media, depth of field, and spatial anti-aliasing in complex scenes. It introduces a flexible, general rendering framework that unifies the handling of such effects by discretizing the integrals into large sets of gather and light points and adaptively approximating the sum of all possible gather-light pair interactions.We create an implicit hierarchy, the product graph, over the gather-light pairs to rapidly and accurately approximate the contribution from hundreds of millions of pairs per pixel while only evaluating a tiny fraction (e.g., 200--1,000). We build upon the techniques of the prior Lightcuts method for complex illumination at a point, however, by considering the complete pixel integrals, we achieve much greater efficiency and scalability.Our example results demonstrate efficient handling of volume scattering, camera focus, and motion of lights, cameras, and geometry. For example, enabling high quality motion blur with 256x temporal sampling requires only a 6.7x increase in shading cost in a scene with complex moving geometry, materials, and illumination. Bruce Walter, Adam Arbree, Kavita Bala, Donald P. Greenberg |
ACM Trans. Graph. | 3 |
| 2006 | Accurate Direct Illumination Using Iterative Adaptive SamplingabstractThis paper introduces a new multipass algorithm for efficiently computing direct illumination in scenes with many lights and complex occlusion. Images are first divided into 8 x 8 pixel blocks and for each point to be shaded within a block, a probability density function (PDF) is constructed over the lights and sampled to estimate illumination using a small number of shadow rays. Information from these samples is then aggregated at both the pixel and block level and used to optimize the PDFs for the next pass. Over multiple passes the PDFs and pixel estimates are updated until convergence. Using aggregation and feedback progressively improves the sampling and automatically exploits both visibility and spatial coherence. We also use novel extensions for efficient antialiasing. Our adaptive multipass approach computes accurate direct illumination eight times faster than prior approaches in tests on several complex scenes. Michael Donikian, Bruce Walter, Kavita Bala, Sebastian Fernandez, Donald P. Greenberg |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2005 | Lightcuts: a scalable approach to illuminationabstractLightcuts is a scalable framework for computing realistic illumination. It handles arbitrary geometry, non-diffuse materials, and illumination from a wide variety of sources including point lights, area lights, HDR environment maps, sun/sky models, and indirect illumination. At its core is a new algorithm for accurately approximating illumination from many point lights with a strongly sublinear cost. We show how a group of lights can be cheaply approximated while bounding the maximum approximation error. A binary light tree and perceptual metric are then used to adaptively partition the lights into groups to control the error vs. cost tradeoff.We also introduce reconstruction cuts that exploit spatial coherence to accelerate the generation of anti-aliased images with complex illumination. Results are demonstrated for five complex scenes and show that lightcuts can accurately approximate hundreds of thousands of point lights using only a few hundred shadow rays. Reconstruction cuts can reduce the number of shadow rays to tens. Bruce Walter, Sebastian Fernandez, Adam Arbree, Kavita Bala, Michael Donikian, Donald P. Greenberg |
ACM Trans. Graph. | 4 |
| 2003 | Detail synthesis for image-based texturingabstractImage-based modeling techniques permit the creation of visually interesting geometric models from photographs. But traditional image-based texturing (IBT) techniques often result in extracted textures of poor, uneven quality. This paper introduces a novel technique to improve the quality of image-based textures. We compute a simple and efficient texture quality metric based on the Jacobian of the imaging transform. We identify the correlation between the values of the Jacobian metric and the levels of an image pyramid, allowing us to formulate a novel texture synthesis approach which operates over textures from 3D surfaces in a scene. Our technique allows the creation of uniform, high-resolution textures, relieving the user of the burden of collecting large numbers of images while increasing the visual quality of image-based models. This improved quality is important to create compelling visual experiences in interactive environments. Ryan M. Ismert, Kavita Bala, Donald P. Greenberg |
SI3D | 2 |
| 2003 | Combining edges and points for interactive high-quality renderingabstractThis paper presents a new interactive rendering and display technique for complex scenes with expensive shading, such as global illumination. Our approach combines sparsely sampled shading (points) and analytically computed discontinuities (edges) to interactively generate high-quality images. The edge-and-point image is a new compact representation that combines edges and points such that fast, table-driven interpolation of pixel shading from nearby point samples is possible, while respecting discontinuities.The edge-and-point renderer is extensible, permitting the use of arbitrary shaders to collect shading samples. Shading discontinuities, such as silhouettes and shadow edges, are found at interactive rates. Our software implementation supports interactive navigation and object manipulation in scenes that include expensive lighting effects (such as global illumination) and geometrically complex objects. For interactive rendering we show that high-quality images of these scenes can be rendered at 8--14 frames per second on a desktop PC: a speedup of 20--60 over a ray tracer computing a single sample per pixel. Kavita Bala, Bruce Walter, Donald P. Greenberg |
ACM Trans. Graph. | 1 |
| 2001 | Adaptive shadow mapsabstractShadow maps provide a fast and convenient method of identifying shadows in scenes but can introduce aliasing. This paper introduces the Adaptive Shadow Map (ASM) as a solution to this problem. An ASM removes aliasing by resolving pixel size mismatches between the eye view and the light source view. It achieves this goal by storing the light source view (i.e., the shadow map for the light source) as a hierarchical grid structure as opposed to the conventional flat structure. As pixels are transformed from the eye view to the light source view, the ASM is refined to create higher-resolution pieces of the shadow map when needed. This is done by evaluating the contributions of shadow map pixels to the overall image quality. The improvement process is view-driven, progressive, and confined to a user-specifiable memory footprint. We show that ASMs enable dramatic improvements in shadow quality while maintaining interactive rates. Randima Fernando, Sebastian Fernandez, Kavita Bala, Donald P. Greenberg |
SIGGRAPH | 3 |
| 1999 | Radiance interpolants for accelerated bounded-error ray tracingabstractRay tracers, which sample radiance, are usually regarded as offline rendering algorithms that are too slow for interactive use. In this article we present a system that exploits object-space, ray-space, image-space, and temporal coherence to accelerate ray tracing. Our system uses per-surface interpolants to approximate radiance both interactive and batch ray tracers. Our approach explicity decouples the two primary operations of a ray tracer—shading and visibility determination—and accelerates each of them independently. Shading is accelerated by quadrilinearily interpolating lazily acquired radiance samples. Interpolation error does not exceed a user-specified bound, allowing the user to control performance/quality tradeoffs. Error is bounded by adaptive sampling at discontinuities and radiance nonlinearities. Visibility determination at pixels is accelerated by reprojecting interpolants as the user's viewpoint changes. A fast scan-line alogoithm then achieves high performance without sacrificing image quality. For a smoothly varying viewpoint, the combination of lazy interpolants and projection substantially accelerates the ray tracer. Additionally, an efficient cache management algorithm keeps the memory footprint of the system small with negilgible overhead. Kavita Bala, Julie Dorsey, Seth J. Teller |
ACM Trans. Graph. | 1 |
| 1995 | Routing in a linear lightwave networkabstractDynamic routing of point-to-point connections in a waveband selective linear lightwave network is addressed. Linear lightwave networks are all optical networks in which only linear operations are performed on signals in a waveband selective manner. Special constraints arise because of the linearity in the linear lightwave network. The overall problem of finding a path satisfying all the routing constraints for point-to-point connections is shown to be very complex. Owing to the complexity, the overall routing problem is decomposed into several subproblems. In particular, given a request for a point-to-point connection a waveband is first chosen for the call. Two heuristics, MAXBAND which allocates the most used band to a call and another MINBAND (least used band) are studied. Then, the problem of routing in a given waveband is further divided into smaller subproblems of finding a path in the waveband, checking for feasibility of the path in the chosen waveband and channel allocation (within the waveband). For finding paths in a waveband, K-SP, BLOW-UP and MIN-INT algorithms are proposed. A recursive algorithm checks for feasibility of the path on the waveband. Two channel allocation schemes (within a single waveband) MIN and MAX are presented. Simulations show that using MAXBAND (waveband), MIN-INT (path on waveband) and MIN (channel within waveband) policies resulted in the best performance (least blocking).> Krishna Bala, Thomas E. Stern, David Simchi-Levi, Kavita Bala |
IEEE/ACM Trans. Netw. | 4 |
| 1994 | Software Prefetching and Caching for Translation Lookaside Buffers
Kavita Bala, M. Frans Kaashoek, William E. Weihl |
OSDI | 1 |
| 1991 | Algorithms for Routing in a Linear Lightwave NetworkabstractRouting algorithms are proposed for setting up calls on a circuit-switched basis in linear lightwave networks (LLN), i.e., networks composed only of linear components, including controllable power combiners and dividers, and possibly linear (non-regenerative) optical amplifiers. The overall problem is decomposed into three subproblems: (1) physical path allocation, (2) checking for violations of the special optical constraints on the allocated physical path, and (3) channel assignment. Only point to point connections are considered. The physical path allocation technique uses the K-shortest path algorithm and tries to minimize the number of sources potentially interfering with each other, as a result of the incoming call. A channel assignment heuristic that tends to spread out calls evenly among the available channels works better than one that tries to maximize channel reuse.> Krishna Bala, Thomas E. Stern, Kavita Bala |
INFOCOM | 3 |