Peter N. Belhumeur

dblp:12/5558 · DBLP profile ↗
← Back
82ranked-venue papers
17as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 65 · 16 first-authorGraphics, computer vision, multimedia, augmented reality and games · 59 · 9 first-authorComputer networks · 3Security and privacy · 2Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
46 papers
Image recognition and object detection · 33% Face, body and person analysis · 28% 3D vision · 20%
Computer graphics and multimedia
33 papers
Computational photography and imaging · 36% Rendering · 35% Image and video processing · 22%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%

Topics — the 30 heaviest of 120, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
1.172014
Birdsnap: Large-Scale Fine-Grained Visual Categorization of Birds · CVPR 2014
Bird Part Localization Using Exemplar-Based Models with Enforced Pose and Subcategory Consistency · ICCV 2013
How Do You Tell a Blackbird from a Crow? · ICCV 2013
Computer vision › Image recognition and object detection › object localization
part localization
0.322014
Part-Pair Representation for Part Localization · ECCV (2) 2014
Dog Breed Classification Using Part Localization · ECCV (1) 2012
Computer vision › Face, body and person analysis
face recognition
0.392011
FaceTracer: A Search Engine for Large Collections of Images with Faces · ECCV (4) 2008
Using Eye Reflections for Face Recognition Under Varying Illumination · ICCV 2005
Describable Visual Attributes for Face Verification and Image Search · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › Face, body and person analysis
face alignment
0.322013
Localizing Parts of Faces Using a Consensus of Exemplars · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Localizing parts of faces using a consensus of exemplars · CVPR 2011
Computer vision › Face, body and person analysis › face recognition
face verification
0.332013
Describable Visual Attributes for Face Verification and Image Search · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Attribute and simile classifiers for face verification · ICCV 2009
POOF: Part-Based One-vs.-One Features for Fine-Grained Categorization, Face Verification, and Attribute Estimation · CVPR 2013
Computer vision › Face, body and person analysis › human pose estimation
articulated pose estimation
0.212016
Articulated Pose Estimation Using Hierarchical Exemplar-Based Models · AAAI 2016
Computer vision › Face, body and person analysis
human pose estimation
0.212016
Articulated Pose Estimation Using Hierarchical Exemplar-Based Models · AAAI 2016
Rendering
appearance modeling
0.242007
Time-Varying BRDFs · IEEE Trans. Vis. Comput. Graph. 2007
Time-varying surface appearance: acquisition, modeling and rendering · ACM Trans. Graph. 2006
Reflectance Sharing: Predicting Appearance from a Sparse Set of Images of a Known Shape · IEEE Trans. Pattern Anal. Mach. Intell. 2006
Information retrieval
image retrieval
0.222011
Describable Visual Attributes for Face Verification and Image Search · IEEE Trans. Pattern Anal. Mach. Intell. 2011
FaceTracer: A Search Engine for Large Collections of Images with Faces · ECCV (4) 2008
Computer vision › Image recognition and object detection
attribute recognition
0.222013
Multi-attribute spaces: Calibration for attribute fusion and similarity search · CVPR 2012
POOF: Part-Based One-vs.-One Features for Fine-Grained Categorization, Face Verification, and Attribute Estimation · CVPR 2013
Computer vision › Image recognition and object detection › object detection › fine-grained object detection
bird species detection
0.212014
Birdsnap: Large-Scale Fine-Grained Visual Categorization of Birds · CVPR 2014
Machine learning › Learning theory › classification
classifier design
0.212014
Birdsnap: Large-Scale Fine-Grained Visual Categorization of Birds · CVPR 2014
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.222013
Localizing Parts of Faces Using a Consensus of Exemplars · IEEE Trans. Pattern Anal. Mach. Intell. 2013
A Bayesian approach to binocular steropsis · Int. J. Comput. Vis. 1996
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
non-parametric methods
0.212013
Localizing Parts of Faces Using a Consensus of Exemplars · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Machine learning › Trustworthy machine learning › interpretability
visual explanation
0.212013
How Do You Tell a Blackbird from a Crow? · ICCV 2013
Image and video processing › color image processing
illumination invariance
0.122008
Color Subspaces as Photometric Invariants · Int. J. Comput. Vis. 2008
Color Subspaces as Photometric Invariants · CVPR (2) 2006
Bioinformatics and computational biology › plant biology
plant species identification
0.112012
Leafsnap: A Computer Vision System for Automatic Plant Species Identification · ECCV (2) 2012
Information retrieval
similarity search
0.112012
Multi-attribute spaces: Calibration for attribute fusion and similarity search · CVPR 2012
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.152003
Binocular Helmholtz Stereopsis · ICCV 2003
Helmholtz Stereopsis: Exploiting Reciprocity for Surface Reconstruction · Int. J. Comput. Vis. 2002
Helmholtz Stereopsis: Exploiting Reciprocity for Surface Reconstruction · ECCV (3) 2002
Information retrieval › image retrieval › semantic image retrieval
attribute-based image retrieval
0.112011
Describable Visual Attributes for Face Verification and Image Search · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computational photography and imaging › active illumination
illumination multiplexing
0.122007
Multiplexing for Optimal Lighting · IEEE Trans. Pattern Anal. Mach. Intell. 2007
A Theory of Multiplexed Illumination · ICCV 2003
Image and video processing
image restoration
0.132009
Specularity Removal in Images and Videos: A PDE Approach · ECCV (1) 2006
Removing image artifacts due to dirty camera lenses and thin occluders · ACM Trans. Graph. 2009
Active refocusing of images and videos · ACM Trans. Graph. 2007
Wireless sensing and localization
network localization
0.122006
A Theory of Network Localization · IEEE Trans. Mob. Comput. 2006
Rigidity, Computation, and Randomization in Network Localization · INFOCOM 2004
Computer vision › 3D vision
photometric analysis
0.122005
A Fourier Theory for Cast Shadows · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Using Eye Reflections for Face Recognition Under Varying Illumination · ICCV 2005
Computer vision › Face, body and person analysis › face recognition › robust face recognition
illumination-invariant face recognition
0.142005
Using Eye Reflections for Face Recognition Under Varying Illumination · ICCV 2005
In Search of Illumination Invariants · CVPR 2000
Illumination Cones for Recognition under Variable Lighting: Faces · CVPR 1998
Computer vision › Video understanding and tracking › object tracking › appearance modeling
illumination modeling
0.142001
From Few to Many: Illumination Cone Models for Face Recognition under Variable Lighting and Pose · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Determining Generative Models of Objects Under Varying Illumination: Shape and Albedo from Multiple Images Using SVD and Integrability · Int. J. Comput. Vis. 1999
Comparing Images under Variable Illumination · CVPR 1998
Computer vision › Face, body and person analysis › facial attribute analysis
facial attribute recognition
0.112009
Attribute and simile classifiers for face verification · ICCV 2009
Image and video processing › image restoration › artifact removal
image artifact removal
0.112009
Removing image artifacts due to dirty camera lenses and thin occluders · ACM Trans. Graph. 2009
Image and video processing › video frame interpolation › interpolation
image interpolation
0.112009
Moving gradients: a path-based method for plausible image interpolation · ACM Trans. Graph. 2009
Image and video processing
occlusion handling
0.112009
Moving gradients: a path-based method for plausible image interpolation · ACM Trans. Graph. 2009

Methods — techniques the papers use, named apart from their topics

computer vision · 0.5structured light · 0.3hierarchical exemplar-based model · 0.2deep convolutional neural network · 0.2compressive sensing · 0.2spatio-temporal prior estimation · 0.2part-pair representation · 0.2one-vs-all classifier · 0.2graph rigidity theory · 0.2computational complexity analysis · 0.2visual field guide generation · 0.2similarity tree · 0.2part-based one-vs-one features · 0.2iterative attenuation correction · 0.2discriminative intermediate features · 0.2face detection · 0.2score calibration · 0.1extreme value theory · 0.1
YearPublicationVenuePosition
2020 Stochastic Dynamics for Video Infilling
abstract
In this paper, we introduce a stochastic dynamics video infilling (SDVI) framework to generate frames between long intervals in a video. Our task differs from video interpolation which aims to produce transitional frames for a short interval between every two frames and increase the temporal resolution. Our task, namely video infilling, however, aims to infill long intervals with plausible frame sequences. Our framework models the infilling as a constrained stochastic generation process and sequentially samples dynamics from the inferred distribution. SDVI consists of two parts: (1) a bi-directional constraint propagation module to guarantee the spatial-temporal coherence among frames, (2) a stochastic sampling process to generate dynamics from the inferred distributions. Experimental results show that SDVI can generate clear frame sequences with varying contents. Moreover, motions in the generated sequence are realistic and able to transfer smoothly from the given start frame to the terminal frame.
Qiangeng Xu, Hanwang Zhang, Weiyue Wang 0002, Peter N. Belhumeur, Ulrich Neumann
WACV4
2018 The Minimalist Camera
Parita Pooj, Michael D. Grossberg, Peter N. Belhumeur, Shree K. Nayar
BMVC3
2016 Articulated Pose Estimation Using Hierarchical Exemplar-Based Models
abstract
Exemplar-based models have achieved great success on localizing the parts of semi-rigid objects. However, their efficacy on highly articulated objects such as humans is yet to be explored. Inspired by hierarchical object representation and recent application of Deep Convolutional Neural Networks (DCNNs) on human pose estimation, we propose a novel formulation that incorporates both hierarchical exemplar-based models and DCNNs in the spatial terms. Specifically, we obtain more expressive spatial models by assuming independence between exemplars at different levels in the hierarchy; we also obtain stronger spatial constraints by inferring the spatial relations between parts at the same level. As our method strikes a good balance between expressiveness and strength of spatial models, it is both effective and generalizable, achieving state-of-the-art results on different benchmarks: Leeds Sports Dataset and CUB-200-2011.
Jiongxin Liu, Yinxiao Li, Peter K. Allen, Peter N. Belhumeur
AAAI4
2014 Birdsnap: Large-Scale Fine-Grained Visual Categorization of Birds
abstract
We address the problem of large-scale fine-grained visual categorization, describing new methods we have used to produce an online field guide to 500 North American bird species. We focus on the challenges raised when such a system is asked to distinguish between highly similar species of birds. First, we introduce "one-vs-most classifiers." By eliminating highly similar species during training, these classifiers achieve more accurate and intuitive results than common one-vs-all classifiers. Second, we show how to estimate spatio-temporal class priors from observations that are sampled at irregular and biased locations. We show how these priors can be used to significantly improve performance. We then show state-of-the-art recognition performance on a new, large dataset that we make publicly available. These recognition methods are integrated into the online field guide, which is also publicly available.
Thomas Berg, Jiongxin Liu, Michelle L. Alexander, David Jacobs 0001, Peter N. Belhumeur
CVPR6
2014 Part-Pair Representation for Part Localization
Jiongxin Liu, Yinxiao Li, Peter N. Belhumeur
ECCV (2)3
2013 From Bikers to Surfers: Visual Recognition of Urban Tribes
abstract
Iljung S. Kwak1 [email protected] Ana C. Murillo2 [email protected] Peter N. Belhumeur3 [email protected] David Kriegman1 [email protected] Serge Belongie1 [email protected] 1 Dept. of Computer Science and Engineering University of California, San Diego, USA. 2 Dpt. Informatica e Ing. Sistemas Inst. Investigacion en Ingenieria de Aragon. University of Zaragoza, Spain. 3 Department of Computer Science Columbia University, USA.
Iljung S. Kwak, Ana Cristina Murillo, Peter N. Belhumeur, David J. Kriegman, Serge J. Belongie
BMVC3
2013 POOF: Part-Based One-vs.-One Features for Fine-Grained Categorization, Face Verification, and Attribute Estimation
abstract
From a set of images in a particular domain, labeled with part locations and class, we present a method to automatically learn a large and diverse set of highly discriminative intermediate features that we call Part-based One-vs.-One Features (POOFs). Each of these features specializes in discrimination between two particular classes based on the appearance at a particular part. We demonstrate the particular usefulness of these features for fine-grained visual categorization with new state-of-the-art results on bird species identification using the Caltech UCSD Birds (CUB) dataset and parity with the best existing results in face verification on the Labeled Faces in the Wild (LFW) dataset. Finally, we demonstrate the particular advantage of POOFs when training data is scarce.
Thomas Berg, Peter N. Belhumeur
CVPR2
2013 How Do You Tell a Blackbird from a Crow?
abstract
How do you tell a blackbird from a crow? There has been great progress toward automatic methods for visual recognition, including fine-grained visual categorization in which the classes to be distinguished are very similar. In a task such as bird species recognition, automatic recognition systems can now exceed the performance of non-experts - most people are challenged to name a couple dozen bird species, let alone identify them. This leads us to the question, "Can a recognition system show humans what to look for when identifying classes (in this case birds)?" In the context of fine-grained visual categorization, we show that we can automatically determine which classes are most visually similar, discover what visual features distinguish very similar classes, and illustrate the key features in a way meaningful to humans. Running these methods on a dataset of bird images, we can generate a visual field guide to birds which includes a tree of similarity that displays the similarity relations between all species, pages for each species showing the most similar other species, and pages for each pair of similar species illustrating their differences.
Thomas Berg, Peter N. Belhumeur
ICCV2
2013 Bird Part Localization Using Exemplar-Based Models with Enforced Pose and Subcategory Consistency
abstract
In this paper, we propose a novel approach for bird part localization, targeting fine-grained categories with wide variations in appearance due to different poses (including aspect and orientation) and subcategories. As it is challenging to represent such variations across a large set of diverse samples with tractable parametric models, we turn to individual exemplars. Specifically, we extend the exemplar-based models in [4] by enforcing pose and subcategory consistency at the parts. During training, we build pose-specific detectors scoring part poses across subcategories, and subcategory-specific detectors scoring part appearance across poses. At the testing stage, likely exemplars are matched to the image, suggesting part locations whose pose and subcategory consistency are well-supported by the image cues. From these hypotheses, part configuration can be predicted with very high accuracy. Experimental results demonstrate significant performance gains from our method on an extensive dataset: CUB-200-2011 [30], for both localization and classification tasks.
Jiongxin Liu, Peter N. Belhumeur
ICCV2
2013 Localizing Parts of Faces Using a Consensus of Exemplars
abstract
We present a novel approach to localizing parts in images of human faces. The approach combines the output of local detectors with a nonparametric set of global models for the part locations based on over 1,000 hand-labeled exemplar images. By assuming that the global models generate the part locations as hidden variables, we derive a Bayesian objective function. This function is optimized using a consensus of models for these hidden variables. The resulting localizer handles a much wider range of expression, pose, lighting, and occlusion than prior ones. We show excellent performance on real-world face datasets such as Labeled Faces in the Wild (LFW) and a new Labeled Face Parts in the Wild (LFPW) and show that our localizer achieves state-of-the-art performance on the less challenging BioID dataset.
Peter N. Belhumeur, David Jacobs 0001, David J. Kriegman, Neeraj Kumar 0006
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Compressive Structured Light for Recovering Inhomogeneous Participating Media
abstract
We propose a new method named compressive structured light for recovering inhomogeneous participating media. Whereas conventional structured light methods emit coded light patterns onto the surface of an opaque object to establish correspondence for triangulation, compressive structured light projects patterns into a volume of participating medium to produce images which are integral measurements of the volume density along the line of sight. For a typical participating medium encountered in the real world, the integral nature of the acquired images enables the use of compressive sensing techniques that can recover the entire volume density from only a few measurements. This makes the acquisition process more efficient and enables reconstruction of dynamic volumetric phenomena. Moreover, our method requires the projection of multiplexed coded illumination, which has the added advantage of increasing the signal-to-noise ratio of the acquisition. Finally, we propose an iterative algorithm to correct for the attenuation of the participating medium during the reconstruction process. We show the effectiveness of our method with simulations as well as experiments on the volumetric recovery of multiple translucent layers, 3D point clouds etched in glass, and the dynamic process of milk drops dissolving in water.
Jinwei Gu, Shree K. Nayar, Eitan Grinspun, Peter N. Belhumeur, Ravi Ramamoorthi
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Tom-vs-Pete Classifiers and Identity-Preserving Alignment for Face Verification
abstract
We propose a method of face verification that takes advantage of a reference set of faces, disjoint by identity from the test faces, labeled with identity and face part locations. The reference set is used in two ways. First, we use it to perform an “identity-preserving” alignment, warping the faces in a way that reduces differences due to pose and expression but preserves differences that indicate identity. Second, using the aligned faces, we learn a large set of identity classifiers, each trained on images of just two people. We call these “Tom-vs-Pete” classifiers to stress their binary nature. We assemble a collection of these classifiers able to discriminate among a wide variety of subjects and use their outputs as features in a same-or-different classifier on face pairs. We evaluate our method on the Labeled Faces in the Wild benchmark, achieving an accuracy of 93.10%, significantly improving on the published state of the art.
Thomas Berg, Peter N. Belhumeur
BMVC2
2012 Multi-attribute spaces: Calibration for attribute fusion and similarity search
abstract
Recent work has shown that visual attributes are a powerful approach for applications such as recognition, image description and retrieval. However, fusing multiple attribute scores - as required during multi-attribute queries or similarity searches - presents a significant challenge. Scores from different attribute classifiers cannot be combined in a simple way; the same score for different attributes can mean different things. In this work, we show how to construct normalized “multi-attribute spaces” from raw classifier outputs, using techniques based on the statistical Extreme Value Theory. Our method calibrates each raw score to a probability that the given attribute is present in the image. We describe how these probabilities can be fused in a simple way to perform more accurate multiattribute searches, as well as enable attribute-based similarity searches. A significant advantage of our approach is that the normalization is done after-the-fact, requiring neither modification to the attribute classification system nor ground truth attribute annotations. We demonstrate results on a large data set of nearly 2 million face images and show significant improvements over prior work. We also show that perceptual similarity of search results increases by using contextual attributes.
Walter J. Scheirer, Neeraj Kumar 0006, Peter N. Belhumeur, Terrance E. Boult
CVPR3
2012 Leafsnap: A Computer Vision System for Automatic Plant Species Identification
Neeraj Kumar 0006, Peter N. Belhumeur, Arijit Biswas, David Jacobs 0001, W. John Kress, Ida C. Lopez, João V. B. Soares
ECCV (2)2
2012 Dog Breed Classification Using Part Localization
Jiongxin Liu, Angjoo Kanazawa, David Jacobs 0001, Peter N. Belhumeur
ECCV (1)4
2011 Localizing parts of faces using a consensus of exemplars
abstract
We present a novel approach to localizing parts in images of human faces. The approach combines the output of local detectors with a non-parametric set of global models for the part locations based on over one thousand hand-labeled exemplar images. By assuming that the global models generate the part locations as hidden variables, we derive a Bayesian objective function. This function is optimized using a consensus of models for these hidden variables. The resulting localizer handles a much wider range of expression, pose, lighting and occlusion than prior ones. We show excellent performance on a new dataset gathered from the internet and show that our localizer achieves state-of-the-art performance on the less challenging BioID dataset.
Peter N. Belhumeur, David Jacobs 0001, David J. Kriegman, Neeraj Kumar 0006
CVPR1
2011 Two faces are better than one: Face recognition in group photographs
abstract
Face recognition systems classically recognize people individually. When presented with a group photograph containing multiple people, such systems implicitly assume statistical independence between each detected face. We question this basic assumption and consider instead that there is a dependence between face regions from the same image; after all, the image was acquired with a single camera, under consistent lighting (distribution, direction, spectrum), camera motion, and scene/camera geometry. Such naturally occurring commonalities between face images can be exploited when recognition decisions are made jointly across the faces, rather than independently. Furthermore, when recognizing people in isolation, some features such as color are usually uninformative in unconstrained settings. But by considering pairs of people, the relative color difference provides valuable information. This paper reconsiders the independence assumption, introduces new features and methods for recognizing pairs of individuals in group photographs, and demonstrates a marked improvement when these features are used in joint decision making vs. independent decision making. While these features alone are only moderately discriminative, we combine these new features with state-of art attribute features and demonstrate effective recognition performance. Initial experiments on two datasets show promising improvements in accuracy.
Ohil K. Manyam, Neeraj Kumar 0006, Peter N. Belhumeur, David J. Kriegman
IJCB3
2011 Fusing with context: A Bayesian approach to combining descriptive attributes
abstract
For identity related problems, descriptive attributes can take the form of any information that helps represent an individual, including age data, describable visual attributes, and contextual data. With a rich set of descriptive at- tributes, it is possible to enhance the base matching accuracy of a traditional face identification system through intelligent score weighting. If we can factor any attribute differences between people into our match score calculation, we can deemphasize incorrect results, and ideally lift the correct matching record to a higher rank position. Naturally, the presence of all descriptive attributes during a match instance cannot be expected, especially when considering non-biometric context. Thus, in this paper, we examine the application of Bayesian Attribute Networks to combine descriptive attributes and produce accurate weighting factors to apply to match scores from face recognition systems based on incomplete observations made at match time. We also examine the pragmatic concerns of attribute network creation, and introduce a Noisy-OR formulation for stream- lined truth value assignment and more accurate weighting. Experimental results show that incorporating descriptive attributes into the matching process significantly enhances face identification over the baseline by up to 32.8%.
Walter J. Scheirer, Neeraj Kumar 0006, Karl Ricanek, Peter N. Belhumeur, Terrance E. Boult
IJCB4
2011 Describable Visual Attributes for Face Verification and Image Search
abstract
We introduce the use of describable visual attributes for face verification and image search. Describable visual attributes are labels that can be given to an image to describe its appearance. This paper focuses on images of faces and the attributes used to describe them, although the concepts also apply to other domains. Examples of face attributes include gender, age, jaw shape, nose size, etc. The advantages of an attribute-based representation for vision tasks are manifold: They can be composed to create descriptions at various levels of specificity; they are generalizable, as they can be learned once and then applied to recognize new objects or categories without any further training; and they are efficient, possibly requiring exponentially fewer attributes (and training data) than explicitly naming each category. We show how one can create and label large data sets of real-world images to train classifiers which measure the presence, absence, or degree to which an attribute is expressed in images. These classifiers can then automatically label new images. We demonstrate the current effectiveness--and explore the future potential--of using attributes for face verification and image search via human and computational experiments. Finally, we introduce two new face data sets, named FaceTracer and PubFig, with labeled attributes and identities, respectively.
Neeraj Kumar 0006, Alexander C. Berg, Peter N. Belhumeur, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Towards Full 3D Helmholtz Stereovision Algorithms
Amaël Delaunoy, Emmanuel Prados, Peter N. Belhumeur
ACCV (1)3
2010 Editorial for the Special Issue on Photometric Analysis for Computer Vision
Peter N. Belhumeur, Katsushi Ikeuchi, Emmanuel Prados, Stefano Soatto, Peter F. Sturm
Int. J. Comput. Vis.1
2009 Attribute and simile classifiers for face verification
abstract
We present two novel methods for face verification. Our first method - “attribute” classifiers - uses binary classifiers trained to recognize the presence or absence of describable aspects of visual appearance (e.g., gender, race, and age). Our second method - “simile” classifiers - removes the manual labeling required for attribute classification and instead learns the similarity of faces, or regions of faces, to specific reference people. Neither method requires costly, often brittle, alignment between image pairs; yet, both methods produce compact visual descriptions, and work on real-world images. Furthermore, both the attribute and simile classifiers improve on the current state-of-the-art for the LFW data set, reducing the error rates compared to the current best by 23.92% and 26.34%, respectively, and 31.68% when combined. For further testing across pose, illumination, and expression, we introduce a new data set - termed PubFig - of real-world images of public figures (celebrities and politicians) acquired from the internet. This data set is both larger (60,000 images) and deeper (300 images per individual) than existing data sets of its kind. Finally, we present an evaluation of human performance.
Neeraj Kumar 0006, Alexander C. Berg, Peter N. Belhumeur, Shree K. Nayar
ICCV3
2009 Removing image artifacts due to dirty camera lenses and thin occluders
abstract
Dirt on camera lenses, and occlusions from thin objects such as fences, are two important types of artifacts in digital imaging systems. These artifacts are not only an annoyance for photographers, but also a hindrance to computer vision and digital forensics. In this paper, we show that both effects can be described by a single image formation model, wherein an intermediate layer (of dust, dirt or thin occluders) both attenuates the incoming light and scatters stray light towards the camera. Because of camera defocus, these artifacts are low-frequency and either additive or multiplicative, which gives us the power to recover the original scene radiance pointwise. We develop a number of physics-based methods to remove these effects from digital photographs and videos. For dirty camera lenses, we propose two methods to estimate the attenuation and the scattering of the lens dirt and remove the artifacts -- either by taking several pictures of a structured calibration pattern beforehand, or by leveraging natural image statistics for post-processing existing images. For artifacts from thin occluders, we propose a simple yet effective iterative method that recovers the original scene from multiple apertures. The method requires two images if the depths of the scene and the occluder layer are known, or three images if the depths are unknown. The effectiveness of our proposed methods are demonstrated by both simulated and real experimental results.
Jinwei Gu, Ravi Ramamoorthi, Peter N. Belhumeur, Shree K. Nayar
ACM Trans. Graph.3
2009 Moving gradients: a path-based method for plausible image interpolation
abstract
We describe a method for plausible interpolation of images, with a wide range of applications like temporal up-sampling for smooth playback of lower frame rate video, smooth view interpolation, and animation of still images. The method is based on the intuitive idea, that a given pixel in the interpolated frames traces out a path in the source images. Therefore, we simply move and copy pixel gradients from the input images along this path. A key innovation is to allow arbitrary (asymmetric) transition points , where the path moves from one image to the other. This flexible transition preserves the frequency content of the originals without ghosting or blurring, and maintains temporal coherence. Perhaps most importantly, our framework makes occlusion handling particularly simple. The transition points allow for matches away from the occluded regions, at any suitable point along the path. Indeed, occlusions do not need to be handled explicitly at all in our initial graph-cut optimization. Moreover, a simple comparison of computed path lengths after the optimization, allows us to robustly identify occluded regions, and compute the most plausible interpolation in those areas. Finally, we show that significant improvements are obtained by moving gradients and using Poisson reconstruction.
Dhruv Mahajan 0001, Fu-Chung Huang, Wojciech Matusik, Ravi Ramamoorthi, Peter N. Belhumeur
ACM Trans. Graph.5
2009 Graphical properties of easily localizable sensor networks
Brian D. O. Anderson, Peter N. Belhumeur, Tolga Eren, David Kiyoshi Goldenberg, A. Stephen Morse, Walter Whiteley, Yang Richard Yang
Wirel. Networks2
2008 Searching the World's Herbaria: A System for Visual Identification of Plant Species
Peter N. Belhumeur, Daozheng Chen, Steven K. Feiner, David Jacobs 0001, W. John Kress, Haibin Ling, Ida C. Lopez, Ravi Ramamoorthi, Sameer Sheorey, Sean White
ECCV (4)1
2008 Compressive Structured Light for Recovering Inhomogeneous Participating Media
Jinwei Gu, Shree K. Nayar, Eitan Grinspun, Peter N. Belhumeur, Ravi Ramamoorthi
ECCV (4)4
2008 FaceTracer: A Search Engine for Large Collections of Images with Faces
Neeraj Kumar 0006, Peter N. Belhumeur, Shree K. Nayar
ECCV (4)2
2008 Color Subspaces as Photometric Invariants
Todd E. Zickler, Satya P. Mallick, David J. Kriegman, Peter N. Belhumeur
Int. J. Comput. Vis.4
2008 Face swapping: automatically replacing faces in photographs
abstract
In this paper, we present a complete system for automatic face replacement in images. Our system uses a large library of face images created automatically by downloading images from the internet, extracting faces using face detection software, and aligning each extracted face to a common coordinate system. This library is constructed off-line, once, and can be efficiently accessed during face replacement. Our replacement algorithm has three main stages. First, given an input image, we detect all faces that are present, align them to the coordinate system used by our face library, and select candidate face images from our face library that are similar to the input face in appearance and pose. Second, we adjust the pose, lighting, and color of the candidate face images to match the appearance of those in the input image, and seamlessly blend in the results. Third, we rank the blended candidate replacements by computing a match distance over the overlap region. Our approach requires no 3D model, is fully automatic, and generates highly plausible results across a wide range of skin tones, lighting conditions, and viewpoints. We show how our approach can be used for a variety of applications including face de-identification and the creation of appealing group photographs from a set of images. We conclude with a user study that validates the high quality of our replacement results, and a discussion on the current limitations of our system.
Dmitri Bitouk, Neeraj Kumar 0006, Samreen Dhillon, Peter N. Belhumeur, Shree K. Nayar
ACM Trans. Graph.4
2007 Dirty Glass: Rendering Contamination on Transparent Surfaces
Jinwei Gu, Ravi Ramamoorthi, Peter N. Belhumeur, Shree K. Nayar
Rendering Techniques3
2007 Multiplexing for Optimal Lighting
abstract
Imaging of objects under variable lighting directions is an important and frequent practice in computer vision, machine vision, and image-based rendering. Methods for such imaging have traditionally used only a single light source per acquired image. They may result in images that are too dark and noisy, e.g., due to the need to avoid saturation of highlights. We introduce an approach that can significantly improve the quality of such images, in which multiple light sources illuminate the object simultaneously from different directions. These illumination-multiplexed frames are then computationally demultiplexed. The approach is useful for imaging dim objects, as well as objects having a specular reflection component. We give the optimal scheme by which lighting should be multiplexed to obtain the highest quality output, for signal-independent noise. The scheme is based on Hadamard codes. The consequences of imperfections such as stray light, saturation, and noisy illumination sources are then studied. In addition, the paper analyzes the implications of shot noise, which is signal-dependent, to Hadamard multiplexing. The approach facilitates practical lighting setups having high directional resolution. This is shown by a setup we devise, which is flexible, scalable, and programmable. We used it to demonstrate the benefit of multiplexing in experiments.
Yoav Y. Schechner, Shree K. Nayar, Peter N. Belhumeur
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 A theory of locally low dimensional light transport
abstract
Blockwise or Clustered Principal Component Analysis (CPCA) is commonly used to achieve real-time rendering of shadows and glossy reflections with precomputed radiance transfer (PRT). The vertices or pixels are partitioned into smaller coherent regions, and light transport in each region is approximated by alocally low-dimensional subspaceusing PCA. Many earlier techniques such as surface light field and reflectance field compression use a similar paradigm. However, there has been no clear theoretical understanding of how light transport dimensionality increases with local patch size, nor of the optimal block size or number of clusters. In this paper, we develop a theory of locally low dimensional light transport, by using Szego's eigenvalue theorem to analytically derive the eigenvalues of the covariance matrix for canonical cases. We show mathematically that for symmetric patches of areaA, the number of basis functions for glossy reflections increases linearly withA, while for simple cast shadows, it often increases as √A. These results are confirmed numerically on a number of test scenes. Next, we carry out an analysis of the cost of rendering, trading off local dimensionality and the number of patches, deriving an optimal block size. Based on this analysis, we provide useful practical insights for setting parameters in CPCA and also derive a new adaptive subdivision algorithm. Moreover, we show that rendering time scales sub-linearly with the resolution of the image, allowing for interactive all-frequency relighting of 1024 x 1024 images.
Dhruv Mahajan 0001, Ira Kemelmacher-Shlizerman, Ravi Ramamoorthi, Peter N. Belhumeur
ACM Trans. Graph.4
2007 Active refocusing of images and videos
abstract
We present a system for refocusing images and videos of dynamic scenes using a novel, single-view depth estimation method. Our method for obtaining depth is based on the defocus of a sparse set of dots projected onto the scene. In contrast to other active illumination techniques, the projected pattern of dots can be removed from each captured image and its brightness easily controlled in order to avoid under- or over-exposure. The depths corresponding to the projected dots and a color segmentation of the image are used to compute an approximate depth map of the scene with clean region boundaries. The depth map is used to refocus the acquired image after the dots are removed, simulating realistic depth of field effects. Experiments on a wide variety of scenes, including close-ups and live action, demonstrate the effectiveness of our method.
Francesc Moreno-Noguer, Peter N. Belhumeur, Shree K. Nayar
ACM Trans. Graph.2
2007 A first-order analysis of lighting, shading, and shadows
abstract
The shading in a scene depends on a combination of many factors---how the lighting varies spatially across a surface, how it varies along different directions, the geometric curvature and reflectance properties of objects, and the locations of soft shadows. In this article, we conduct a complete first-order or gradient analysis of lighting, shading, and shadows, showing how each factor separately contributes to scene appearance, and when it is important. Gradients are well-suited to analyzing the intricate combination of appearance effects, since each gradient term corresponds directly to variation in a specific factor. First, we show how the spatialanddirectional gradients of the light field change as light interacts with curved objects. This extends the recent frequency analysis of Durand et al. [2005] to gradients, and has many advantages for operations, like bump mapping, that are difficult to analyze in the Fourier domain. Second, we consider the individual terms responsible for shading gradients, such as lighting variation, convolution with the surface BRDF, and the object's curvature. This analysis indicates the relative importance of various terms, and shows precisely how they combine in shading. Third, we understand the effects of soft shadows, computing accurate visibility gradients, and generalizing previous work to arbitrary curved occluders. As one practical application, our visibility gradients can be directly used with conventional ray-tracing methods in practical gradient interpolation methods for efficient rendering. Moreover, our theoretical framework can be used to adaptively sample images in high-gradient regions for efficient rendering.
Ravi Ramamoorthi, Dhruv Mahajan 0001, Peter N. Belhumeur
ACM Trans. Graph.3
2007 Time-Varying BRDFs
abstract
The properties of virtually all real-world materials change with time, causing their bidirectional reflectance distribution functions (BRDFs) to be time varying. However, none of the existing BRDF models and databases take time variation into consideration; they represent the appearance of a material at a single time instance. In this paper, we address the acquisition, analysis, modeling, and rendering of a wide range of time-varying BRDFs (TVBRDFs). We have developed an acquisition system that is capable of sampling a material's BRDF at multiple time instances, with each time sample acquired within 36 sec. We have used this acquisition system to measure the BRDFs of a wide range of time-varying phenomena, which include the drying of various types of paints (watercolor, spray, and oil), the drying of wet rough surfaces (cement, plaster, and fabrics), the accumulation of dusts (household and joint compound) on surfaces, and the melting of materials (chocolate). Analytic BRDF functions are fit to these measurements and the model parameters' variations with time are analyzed. Each category exhibits interesting and sometimes nonintuitive parameter trends. These parameter trends are then used to develop analytic TVBRDF models. The analytic TVBRDF models enable us to apply effects such as paint drying and dust accumulation to arbitrary surfaces and novel materials.
Kalyan Sunkavalli, Ravi Ramamoorthi, Peter N. Belhumeur, Shree K. Nayar
IEEE Trans. Vis. Comput. Graph.4
2006 Color Subspaces as Photometric Invariants
abstract
Complex reflectance phenomena such as specular reflections confound many vision problems since they produce image ‘features’ that do not correspond directly to intrinsic surface properties such as shape and spectral reflectance. A common approach to mitigate these effects is to explore functions of an image that are invariant to these photometric events. In this paper we describe two such invariants" one invariant to specular reflections, and the other invariant to both specular reflections and diffuse shading" that result from exploiting color information in images of dichromatic surfaces. These invariants are derived from subspaces of RGB color space, and they enable the application of Lambertian-based vision techniques to a broad class of specular, non-Lambertian scenes. Using implementations of recent algorithms taken from the literature, we demonstrate the practical utility of these invariants for a wide variety of applications, including stereo, shape from shading, material-based segmentation, and motion estimation.
Todd E. Zickler, Satya P. Mallick, David J. Kriegman, Peter N. Belhumeur
CVPR (2)4
2006 Specularity Removal in Images and Videos: A PDE Approach
Satya P. Mallick, Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
ECCV (1)3
2006 Reflectance Sharing: Predicting Appearance from a Sparse Set of Images of a Known Shape
abstract
Three-dimensional appearance models consisting of spatially varying reflectance functions defined on a known shape can be used in analysis-by-synthesis approaches to a number of visual tasks. The construction of these models requires the measurement of reflectance, and the problem of recovering spatially varying reflectance from images of known shape has drawn considerable interest. To date, existing methods rely on either: (1) low-dimensional (e.g., parametric) reflectance models, or (2) large data sets involving thousands of images (or more) per object. Appearance models based on the former have limited accuracy and generality since they require the selection of a specific reflectance model a priori, and while approaches based on the latter may be suitable for certain applications, they are generally too costly and cumbersome to be used for image analysis. We present an alternative approach that seeks to combine the benefits of existing methods by enabling the estimation of a nonparametric spatially varying reflectance function from a small number of images. We frame the problem as scattered-data interpolation in a mixed spatial and angular domain, and we present a theory demonstrating that the angular accuracy of a recovered reflectance function can be increased in exchange for a decrease in its spatial resolution. We also present a practical solution to this interpolation problem using a new representation of reflectance based on radial basis functions. This representation is evaluated experimentally by testing its ability to predict appearance under novel view and lighting conditions. Our results suggest that since reflectance typically varies slowly from point to point over much of an object's surface, we can often obtain a nonparametric reflectance function from a sparse set of images. In fact, in some cases, we can obtain reasonable results in the limiting case of only a single input image.
Todd E. Zickler, Ravi Ramamoorthi, Sebastian Enrique, Peter N. Belhumeur
IEEE Trans. Pattern Anal. Mach. Intell.4
2006 A Theory of Network Localization
abstract
In this paper, we provide a theoretical foundation for the problem of network localization in which some nodes know their locations and other nodes determine their locations by measuring the distances to their neighbors. We construct grounded graphs to model network localization and apply graph rigidity theory to test the conditions for unique localizability and to construct uniquely localizable networks. We further study the computational complexity of network localization and investigate a subclass of grounded graphs where localization can be computed efficiently. We conclude with a discussion of localization in sensor networks where the sensors are placed randomly.
James Aspnes, Tolga Eren, David Kiyoshi Goldenberg, A. Stephen Morse, Walter Whiteley, Yang Richard Yang, Brian D. O. Anderson, Peter N. Belhumeur
IEEE Trans. Mob. Comput.8
2006 Time-varying surface appearance: acquisition, modeling and rendering
abstract
For computer graphics rendering, we generally assume that the appearance of surfaces remains static over time. Yet, there are a number of natural processes that cause surface appearance to vary dramatically, such as burning of wood, wetting and drying of rock and fabric, decay of fruit skins, and corrosion and rusting of steel and copper. In this paper, we take a significant step towards measuring, modeling, and rendering time-varying surface appearance. We describe the acquisition of the first time-varying database of 26 samples, encompassing a variety of natural processes including burning, drying, decay, and corrosion. Our main technical contribution is a Space-Time Appearance Factorization (STAF). This model factors space and time-varying effects. We derive an overall temporal appearance variation characteristic curve of the specific process, as well as space-dependent textures, rates, and offsets. This overall temporal curve controls different spatial locations evolve at the different rates, causing spatial patterns on the surface over time. We show that the model accurately represents a variety of phenomena. Moreover, it enables a number of novel rendering applications, such as transfer of the time-varying effect to a new static surface, control to accelerate time evolution in certain areas, extrapolation beyond the acquired sequence, and texture synthesis of time-varying appearance.
Jinwei Gu, Chien-I Tu, Ravi Ramamoorthi, Peter N. Belhumeur, Wojciech Matusik, Shree K. Nayar
ACM Trans. Graph.4
2005 Beyond Lambert: Reconstructing Specular Surfaces Using Color
abstract
We present a photometric stereo method for non-diffuse materials that does not require an explicit reflectance model or reference object. By computing a data-dependent rotation of RGB color space, we show that the specular reflection effects can be separated from the much simpler, diffuse (approximately Lambertian) reflection effects for surfaces that can be modeled with dichromatic reflectance. Images in this transformed color space are used to obtain photometric reconstructions that are independent of the specular reflectance. In contrast to other methods for highlight removal based on dichromatic color separation (e.g., color histogram analysis and/or polarization), we do not explicitly recover the specular and diffuse components of an image. Instead, we simply find a transformation of color space that yields more direct access to shape information. The method is purely local and is able to handle surfaces with arbitrary texture.
Satya P. Mallick, Todd E. Zickler, David J. Kriegman, Peter N. Belhumeur
CVPR (2)4
2005 Using Eye Reflections for Face Recognition Under Varying Illumination
abstract
Face recognition under varying illumination remains a challenging problem. Much progress has been made toward a solution through methods that require multiple gallery images of each subject under varying illumination. Yet for many applications, this requirement is too severe. In this paper, we propose a novel method that requires only a single gallery image per subject taken under unknown lighting. The method builds upon two contributions. We first estimate the lighting from its reflection in the eyes. This allows us to explicitly recover the illumination in the single gallery images as well as the probe image. Next, we exploit the local linearity of face appearance variation across different people. We represent the gallery images as locally linear montages of images of many different faces taken under the same lighting (bootstrap images). Then, we transfer the estimated combination of bootstrap images to synthesize each subject's face under tile probe lighting to accomplish recognition. Finally, we show through tests on the CMU PIE database that we can achieve better recognition results using our lighting estimation method and locally linear montages than the current state-of-the-art.
Ko Nishino, Peter N. Belhumeur, Shree K. Nayar
ICCV2
2005 Reflectance Sharing: Image-based Rendering from a Sparse Set of Images
Todd E. Zickler, Sebastian Enrique, Ravi Ramamoorthi, Peter N. Belhumeur
Rendering Techniques4
2005 A Fourier Theory for Cast Shadows
abstract
Cast shadows can be significant in many computer vision applications, such as lighting-insensitive recognition and surface reconstruction. Nevertheless, most algorithms neglect them, primarily because they involve nonlocal interactions in nonconvex regions, making formal analysis difficult. However, many real instances map closely to canonical configurations like a wall, a V-groove type structure, or a pitted surface. In particular, we experiment with 3D textures like moss, gravel, and a kitchen sponge, whose surfaces include canonical configurations like V-grooves. This paper takes a first step toward a formal analysis of cast shadows, showing theoretically that many configurations can be mathematically analyzed using convolutions and Fourier basis functions. Our analysis exposes the mathematical convolution structure of cast shadows and shows strong connections to recent signal-processing frameworks for reflection and illumination.
Ravi Ramamoorthi, Melissa L. Koudelka, Peter N. Belhumeur
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Making One Object Look Like Another: Controlling Appearance Using a Projector-Camera System
Michael D. Grossberg, Harish Peri, Shree K. Nayar, Peter N. Belhumeur
CVPR (1)4
2004 A Fourier Theory for Cast Shadows
Ravi Ramamoorthi, Melissa L. Koudelka, Peter N. Belhumeur
ECCV (1)3
2004 Rigidity, Computation, and Randomization in Network Localization
abstract
We provide a theoretical foundation for the problem of network localization in which some nodes know their locations and other nodes determine their locations by measuring the distances to their neighbors. We construct grounded graphs to model network localization and apply graph rigidity theory to test the conditions for unique localizability and to construct uniquely localizable networks. We further study the computational complexity of network localization and investigate a subclass of grounded graphs where localization can be computed efficiently. We conclude with a discussion of localization in sensor networks where the sensors are placed randomly.
Tolga Eren, David Kiyoshi Goldenberg, Walter Whiteley, Yang Richard Yang, A. Stephen Morse, Brian D. O. Anderson, Peter N. Belhumeur
INFOCOM7
2004 Lighting sensitive display
abstract
Although display devices have been used for decades, they have functioned without taking into account the illumination of their environment. We present the concept of a lighting sensitive display (LSD)---a display that measures the incident illumination and modifies its content accordingly. An ideal LSD would be able to measure the 4D illumination field incident upon it and generate a 4D light field in response to the illumination. However, current sensing and display technologies do not allow for such an ideal implementation. Our initial LSD prototype uses a 2D measurement of the illumination field and produces a 2D image in response to it. In particular, it renders a 3D scene such that it always appears to be lit by the real environment that the display resides in. The current system is designed to perform best when the light sources in the environment are distant from the display, and a single user in a known location views the display. The displayed scene is represented by compressing a very large set of images (acquired or rendered) of the scene that correspond to different lighting conditions. The compression algorithm is a lossy one that exploits not only image correlations over the illumination dimensions but also coherences over the spatial dimensions of the image. This results in a highly compressed representation of the original image set. This representation enables us to achieve high quality relighting of the scene in real time. Our prototype LSD can render 640 × 480 images of scenes under complex and varying illuminations at 15 frames per second using a 2 GHz processor. We conclude with a discussion on the limitations of the current implementation and potential areas for future research.
Shree K. Nayar, Peter N. Belhumeur, Terrance E. Boult
ACM Trans. Graph.2
2003 Toward a Stratification of Helmholtz Stereopsis
abstract
Helmholtz stereopsis has been previously introduced as a surface reconstruction technique that does not assume a model of surface reflectance. This technique relies on the use of multiple cameras and light sources, and it has been shown to be effective when the camera and source positions are known. Here, we take a stratified look at uncalibrated Helmholtz stereopsis. We derive a photometric matching constraint that can be used to establish correspondence without any knowledge of the cameras and sources (except that they are co-located), and we determine conditions under which we can obtain affine and metric reconstructions. An implementation and experimental results are presented.
Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
CVPR (1)2
2003 A Theory of Multiplexed Illumination
abstract
Imaging of objects under variable lighting directions is an important and frequent practice in computer vision and image-based rendering. We introduce an approach that significantly improves the quality of such images. Traditional methods for acquiring images under variable illumination directions use only a single light source per acquired image. In contrast, our approach is based on a multiplexing principle, in which multiple light sources illuminate the object simultaneously from different directions. Thus, the object irradiance is much higher. The acquired images are then computationally demultiplexed. The number of image acquisitions is the same as in the single-source method. The approach is useful for imaging dim object areas. We give the optimal code by which the illumination should be multiplexed to obtain the highest quality output. For n images corresponding to n light sources, the noise is reduced by /spl radic/(n)/2 relative to the signal. This noise reduction translates to a faster acquisition time or an increase in density of illumination direction samples. It also enables one to use lighting with high directional resolution using practical setups, as we demonstrate in our experiments.
Yoav Y. Schechner, Shree K. Nayar, Peter N. Belhumeur
ICCV3
2003 Binocular Helmholtz Stereopsis
abstract
Helmholtz stereopsis has been introduced recently as a surface reconstruction technique that does not assume a model of surface reflectance. In the reported formulation, correspondence was established using a rank constraint, necessitating at least three viewpoints and three pairs of images. Here, it is revealed that the fundamental Helmholtz stereopsis constraint defines a nonlinear partial differential equation, which can be solved using only two images. It is shown that, unlike conventional stereo, binocular Helmholtz stereopsis is able to establish correspondence (and thereby recover surface depth) for objects having an arbitrary and unknown BRDF and in textureless regions (i.e., regions of constant or slowly varying BRDF). An implementation and experimental results validate the method for specular surfaces with and without texture.
Todd E. Zickler, Jeffrey Ho, David J. Kriegman, Jean Ponce, Peter N. Belhumeur
ICCV5
2002 Helmholtz Stereopsis: Exploiting Reciprocity for Surface Reconstruction
Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
ECCV (3)2
2002 Helmholtz Stereopsis: Exploiting Reciprocity for Surface Reconstruction
Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
Int. J. Comput. Vis.2
2001 Finding Folds: On the Appearance and Identification of Occlusion
abstract
A natural sequel to edge detection is the interpretation of edges. This interpretation can provide useful information to various computer vision processes, including recognition, reconstruction, and tracking. In this paper we consider the problem of identifying occlusion edges in a single image. We examine the appearance of occlusion edges under variable illumination, both analytically and empirically, and find that the pattern of shading in the neighborhood of occlusion edges is a stable feature. Finally, we derive a filter for detecting occlusion and present the results of its application.
Patrick S. Huggins, Hansen F. Chen, Peter N. Belhumeur, Steven W. Zucker
CVPR (2)3
2001 Image-based Modeling and Rendering of Surfaces with Arbitrary BRDFs
abstract
A goal of image-based rendering is to synthesize as realistically as possible man made and natural objects. The paper presents a method for image-based modeling and rendering of objects with arbitrary (possibly anisotropic and spatially varying) BRDFs. An object is modeled by sampling the surface's incident light field to reconstruct a non-parametric apparent BRDF at each visible point on the surface, This can be used to render the object from the same viewpoint but under arbitrarily specified illumination. We demonstrate how these object models can be embedded in synthetic scenes and rendered under global illumination which captures the interreflections between real and synthetic objects. We also show how these image-based models can be automatically composited onto video footage with dynamic illumination so that the effects (shadows and shading) of the lighting on the composited object match those of the scene.
Melissa L. Koudelka, Peter N. Belhumeur, Sebastian Magda, David J. Kriegman
CVPR (1)2
2001 Beyond Lambert: Reconstructing Surfaces with Arbitrary BRDFs
abstract
We address an open and hitherto neglected problem in computer vision, how to reconstruct the geometry of objects with arbitrary and possibly anisotropic bidirectional reflectance distribution functions (BRDFs). Present reconstruction techniques, whether stereo vision, structure from motion, laser range finding, etc. make explicit or implicit assumptions about the BRDF. Here, we introduce two methods that were developed by re-examining the underlying image formation process; the methods make no assumptions about the object's shape, the presence or absence of shadowing, or the nature of the BRDF which may vary over the surface. The first method takes advantage of Helmholtz reciprocity, while the second method exploits the fact that the radiance along a ray of light is constant. In particular, the first method uses stereo pairs of images in which point light sources are co-located at the centers of projection of the stereo cameras. The second method is based on double covering a scene's incident light field; the depths of surface points are estimated using a large collection of images in which the viewpoint remains fixed and a point light source illuminates the object. Results from our implementations lend empirical support to both techniques.
Sebastian Magda, David J. Kriegman, Todd E. Zickler, Peter N. Belhumeur
ICCV4
2001 From Few to Many: Illumination Cone Models for Face Recognition under Variable Lighting and Pose
abstract
We present a generative appearance-based method for recognizing human faces under variation in lighting and viewpoint. Our method exploits the fact that the set of images of an object in fixed pose, but under all possible illumination conditions, is a convex cone in the space of images. Using a small number of training images of each face taken with different lighting directions, the shape and albedo of the face can be reconstructed. In turn, this reconstruction serves as a generative model that can be used to render (or synthesize) images of the face under novel poses and illumination conditions. The pose space is then sampled and, for each pose, the corresponding illumination cone is approximated by a low-dimensional linear subspace whose basis vectors are estimated using the generative model. Our recognition algorithm assigns to a test image the identity of the closest approximated illumination cone. Test results show that the method performs almost without error, except on the most extreme lighting directions.
Athinodoros S. Georghiades, Peter N. Belhumeur, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 In Search of Illumination Invariants
abstract
We consider the problem of determining functions of an image of an object that are insensitive to illumination changes. We first show that for an object with Lambertian reflectance there are no discriminative functions that are invariant to illumination. This result leads as to adopt a probabilistic approach in which we analytically determine a probability distribution for the image gradient as a function of the surface's geometry and reflectance. Our distribution reveals that the direction of the image gradient is insensitive to changes in illumination direction. We verify this empirically by constructing a distribution for the image gradient from more than 20 million samples of gradients in a database of 1,280 images of 20 inanimate objects taken under varying lighting condition. Using this distribution we develop an illumination insensitive measure of image comparison and test it on the problem of face recognition.
Hansen F. Chen, Peter N. Belhumeur, David Jacobs 0001
CVPR2
2000 From Few to Many: Generative Models for Recognition Under Variable Pose and Illumination
abstract
Image variability due to changes in pose and illumination can seriously impair object recognition. This paper presents appearance-based methods which, unlike previous appearance-based approaches, require only a small set of training images to generate a rich representation that models this variability. Specifically, from as few as three images of an object in fixed pose seen under slightly varying but unknown lighting, a surface and an albedo map are reconstructed. These are then used to generate synthetic images with large variations in pose and illumination and thus build a representation useful for object recognition. Our methods have been tested within the domain of face recognition on a subset of the Yale Face Database B containing 4050 images of 10 faces seen under variable pose and illumination. This database was specifically gathered for testing these generative methods. Their performance is shown to exceed that of popular existing methods.
Athinodoros S. Georghiades, Peter N. Belhumeur, David J. Kriegman
FG2
1999 Computational Vision at Yale
Peter N. Belhumeur, James S. Duncan, Gregory D. Hager, Drew McDermott, A. Stephen Morse, Steven W. Zucker
Int. J. Comput. Vis.1
1999 The Bas-Relief Ambiguity
Peter N. Belhumeur, David J. Kriegman, Alan L. Yuille
Int. J. Comput. Vis.1
1999 Determining Generative Models of Objects Under Varying Illumination: Shape and Albedo from Multiple Images Using SVD and Integrability
Alan L. Yuille, Daniel Snow, Russell Epstein, Peter N. Belhumeur
Int. J. Comput. Vis.4
1999 Tracking in 3D: Image Variability Decomposition for Recovering Object Pose and Illumination
Peter N. Belhumeur, Gregory D. Hager
Pattern Anal. Appl.1
1998 Illumination Cones for Recognition under Variable Lighting: Faces
abstract
Due to illumination variability, the same object can appear dramatically different even when viewed in fixed pose. To handle this variability, an object recognition system must employ a representation that is either invariant to, or models this variability. This paper presents an appearance-based method for modeling the variability due to illumination in the images of objects. The method differs from past appearance-based methods, however, in that a small set of training images is used to generate a representation-the illumination cone-which models the complete set of images of an object with Lambertian reflectance map under an arbitrary combination of point light sources at infinity. This method is both an implementation and extension (an extension in that it models cast shadows) of the illumination cone representation proposed in Belhumeur and Kriegman (1996). The method is tested on a database of 660 images of 10 faces, and the results exceed those of popular existing methods.
Athinodoros S. Georghiades, David J. Kriegman, Peter N. Belhumeur
CVPR3
1998 Comparing Images under Variable Illumination
abstract
We consider the problem of determining whether two images come from different objects or the same object in the same pose, but under different illumination conditions. We show that this problem cannot be solved using hard constraints: even using a Lambertian reflectance model, there is always an object and a pair of lighting conditions consistent with any two images. Nevertheless, we show that for point sources and objects with Lambertian reflectance, the ratio of two images from the same object is simpler than the ratio of images from different objects. We also show that the ratio of the two images provides two of the three distinct values in the Hessian matrix of the object's surface. Using these observations, we develop a simple measure for matching images under variable illumination, comparing its performance to other existing methods on a database of 450 images of 10 individuals.
David Jacobs 0001, Peter N. Belhumeur, Ronen Basri
CVPR2
1998 What Shadows Reveal about Object Structure
David J. Kriegman, Peter N. Belhumeur
ECCV (2)2
1998 What Is the Set of Images of an Object Under All Possible Illumination Conditions?
Peter N. Belhumeur, David J. Kriegman
Int. J. Comput. Vis.1
1998 Efficient Region Tracking With Parametric Models of Geometry and Illumination
abstract
As an object moves through the field of view of a camera, the images of the object may change dramatically. This is not simply due to the translation of the object across the image plane; complications arise due to the fact that the object undergoes changes in pose relative to the viewing camera, in illumination relative to light sources, and may even become partially or fully occluded. We develop an efficient general framework for object tracking, which addresses each of these complications. We first develop a computationally efficient method for handling the geometric distortions produced by changes in pose. We then combine geometry and illumination into an algorithm that tracks large image regions using no more computation than would be required to track with no accommodation for illumination changes. Finally, we augment these methods with techniques from robust statistics and treat occluded regions on the object as statistical outliers. Experimental results are given to demonstrate the effectiveness of our methods.
Gregory D. Hager, Peter N. Belhumeur
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 The Bas-Relief Ambiguity
abstract
Since antiquity, artisans have created flattened forms, often called "bas-reliefs,"-which give an exaggerated perception of depth when viewed from a particular vantage point. This paper presents an explanation of this phenomena, showing that the ambiguity in determining the relief of an object is not confined to bas-relief sculpture but is implicit in the determination of the structure of any object. Formally, if the object's true surface is denoted by z/sub true/=f(x, y), then we define the "generalized bas-relief transformation" as z=/spl lambda/f(x, y)+/spl mu/x+/spl nu/y, with a corresponding transformation of the albedo. For each image of a Lambertian surface f(x, y) produced by a point light source at infinity, there exists an identical image of a bas-relief produced by a transformed light source. This equality holds for both shaded and shadowed regions. Thus, the set of possible images (illumination cone) is invariant over generalized bas-relief transformations. When /spl mu/=/spl nu/=0 (e.g. a classical bas-relief sculpture), we show that the set of possible motion fields are also identical. Thus, neither small unknown motions nor changes of illumination can resolve the bas-relief ambiguity. Implications of this ambiguity on structure recovery and shape representation are discussed.
Peter N. Belhumeur, David J. Kriegman, Alan L. Yuille
CVPR1
1997 Eigenfaces vs. Fisherfaces: Recognition Using Class Specific Linear Projection
abstract
We develop a face recognition algorithm which is insensitive to large variation in lighting direction and facial expression. Taking a pattern classification approach, we consider each pixel in an image as a coordinate in a high-dimensional space. We take advantage of the observation that the images of a particular face, under varying illumination but fixed pose, lie in a 3D linear subspace of the high dimensional image space-if the face is a Lambertian surface without shadowing. However, since faces are not truly Lambertian surfaces and do indeed produce self-shadowing, images will deviate from this linear subspace. Rather than explicitly modeling this deviation, we linearly project the image into a subspace in a manner which discounts those regions of the face with large deviation. Our projection method is based on Fisher's linear discriminant and produces well separated classes in a low-dimensional subspace, even under severe variation in lighting and facial expressions. The eigenface technique, another method based on linearly projecting the image space to a low dimensional subspace, has similar computational requirements. Yet, extensive experimental results demonstrate that the proposed "Fisherface" method has error rates that are lower than those of the eigenface technique for tests on the Harvard and Yale face databases.
Peter N. Belhumeur, João Pedro Hespanha, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 What is the set of images of an object under all possible lighting conditions?
abstract
The appearance of a particular object depends on both the viewpoint from which it is observed and the light sources by which it is illuminated. If the appearance of two objects is never identical for any pose or lighting conditions, then-in theory - the objects can always be distinguished or recognized. The question arises: What is the set of images of an object under all lighting conditions and pose? In this paper, we consider only the set of images of an object under variable illumination (including multiple, extended light sources and attached shadows). We prove that the set of n-pixel images of a convex object with a Lambertian reflectance function, illuminated by an arbitrary number of point light sources at infinity, forms a convex polyhedral cone in IR/sup n/ and that the dimension of this illumination cone equals the number of distinct surface normals. Furthermore, we show that the cone for a particular object can be constructed from three properly chosen images. Finally, we prove that the set of n-pixel images of an object of any shape and with an arbitrary reflectance function, seen under all possible illumination conditions, still forms a convex cone in IR/sup n/. These results immediately suggest certain approaches to object recognition. Throughout this paper, we offer results demonstrating the empirical validity of the illumination cone representation.
Peter N. Belhumeur, David J. Kriegman
CVPR1
1996 Real-time tracking of image regions with changes in geometry and illumination
abstract
Historically, SSD or correlation-based visual tracking algorithms have been sensitive to changes in illumination and shading across the target region. This paper describes methods for implementing SSD tracking that is both insensitive to illumination variations and computationally efficient. We first describe a vector-space formulation of the tracking problem, showing how to recover geometric deformations. We then show that the same vector space formulation can be used to account for changes in illumination. We combine geometry and illumination into an algorithm that tracks large image regions on live video sequences using no more computation than would be required to trade with no accommodation for illumination changes. We present experimental results which compare the performance of SSD tracking with and without illumination compensation.
Gregory D. Hager, Peter N. Belhumeur
CVPR2
1996 Eigenfaces vs. Fisherfaces: Recognition Using Class Specific Linear Projection
Peter N. Belhumeur, João Pedro Hespanha, David J. Kriegman
ECCV (1)1
1996 Estimation of motion boundary location and optical flow using dynamic programming
abstract
We present a new method for the estimation of optical flow which uses a dynamic programming based algorithm to simultaneously detect the presence of motion boundaries and to estimate optical flow. This allows for a more accurate estimation of the motion field near discontinuities. The results compare favorably with those produced by other methods.
Xenophon Papademetris, Peter N. Belhumeur
ICIP (1)2
1996 A Bayesian approach to binocular steropsis
Peter N. Belhumeur
Int. J. Comput. Vis.1
1995 Recovering Object Surfaces from Viewed Changes in Surface Texture Patterns
abstract
Explores the reconstruction of object surfaces from viewed changes in surface texture patterns. Our approach differs from those in the past in that instead of simply producing local estimates of the surface orientation, our algorithm recovers complete surfaces. Past approaches only found the surface orientation locally and, therefore, did not take advantage of the surface integrability constraint. Our algorithm does not assume that the surface texture pattern is isotropic, and it does not assume that the viewed surface is at some point fronto-parallel. Furthermore, our algorithm has mechanisms for handling texture boundaries and, consequently, does not produce erratic results in the regions abutting these boundaries. Results on real images are presented demonstrating the potential of our approach.>
Peter N. Belhumeur, Alan L. Yuille
ICCV1
1994 Global Priors for Binocular Stereopsis
abstract
Develops a Bayesian feedback method for incorporating global structure into prior models for binocular stereopsis. Since most stereo scenes contain either background continuation (large background surfaces continuing behind smaller fore-ground surfaces) or transparency continuation (small opaque patches on a transparent surface), highly nonlocal interactions are often present in the scene geometry. The commonly used local prior models which impose piecewise smoothness constraints on the reconstructions do not capture the probabilistic subtleties of global 3D structures. Therefore, the authors develop a hybridized prior which balances the local properties of the scene geometry with the global properties. Experimental results demonstrating the potential of this technique are provided.>
Peter N. Belhumeur
ICIP (2)1
1993 A binocular stereo algorithm for reconstructing sloping, creased, and broken surfaces in the presence of half-occlusion
abstract
The author presents a method for reconstructing the three-dimensional scene geometry, i.e., depth, surface orientation, occluding contours, and surface creases, from a pair of stereo images. This reconstruction is not done as a postprocessing step, but rather all of the quantities are estimated simultaneously as part of the matching algorithm. An energy functional is considered in which each of the quantities in the scene geometry is explicitly represented. For this energy functional, a smoothness prior that, in addition to its ability to detect surface discontinuities and the accompanying half-occluded regions, is able to reconstruct steeply sloping surfaces with sharp creases is used. Experimental results demonstrating the effectiveness of the algorithm are presented.>
Peter N. Belhumeur
ICCV1
1992 A Bayesian treatment of the stereo correspondence problem using half-occluded regions
abstract
A half-occluded region in a stereo pair is a set of pixels in one image representing points in space visible to that camera or eye only, and not to the other. These occur typically as parts of the background immediately to the left and right sides of nearby occluding objects, and are present in most natural scenes. Previous approaches to stereo either ignored these unmatchable points or attempted to weed them out in a second pass. An algorithm that incorporates them from the start as a strong clue to depth discontinuities is presented. The authors first derive a measure for goodness of fit and a prior based on a simplified model of objects in space, which leads to an energy functional depending both on the depth as measured from a central cyclopean eye and on the regions of points occluded from the left and right eye perspectives. They minimize this using dynamic programming along epipolar lines followed by annealing in both dimensions. Experiments indicate that this method is very effective even in difficult scenes.>
Peter N. Belhumeur, David Mumford
CVPR1
1989 Toward a Model-Based Bayesian Theory for Estimating and Recognizing Parameterized 3-D Objects Using Two or More Images Taken from Different Positions
abstract
A parametric modeling and statistical estimation approach is proposed and simulation data are shown for estimating 3-D object surfaces from images taken by calibrated cameras in two positions. The parameter estimation suggested is gradient descent, though other search strategies are also possible. Processing image data in blocks (windows) is central to the approach. After objects are modeled as patches of spheres, cylinders, planes and general quadrics-primitive objects, the estimation proceeds by searching in parameter space to simultaneously determine and use the appropriate pair of image regions, one from each image, and to use these for estimating a 3-D surface patch. The expression for the joint likelihood of the two images is derived and it is shown that the algorithm is a maximum-likelihood parameter estimator. A concept arising in the maximum likelihood estimation of 3-D surfaces is modeled and estimated. Cramer-Rao lower bounds are derived for the covariance matrices for the errors in estimating the a priori unknown object surface shape parameters.>
Bruno Cernuschi-Frías, David B. Cooper, Yi-Ping Hung, Peter N. Belhumeur
IEEE Trans. Pattern Anal. Mach. Intell.4
1986 3-D object position estimation and recognition based on parameterized surfaces and multiple views
abstract
A new approach is introduced to 3-D parameterized object estimation and recognition. Though the theory is applicable for any parameterization, we use a model for which objects are approximated by patches of spheres, cylinders, and planes-primitive objects. These primitive surfaces are special cases of 3-D quadric surfaces. Primitive surface estimation is treated as parameter estimation using data patches in two or more noisy images taken by calibrated cameras in different locations and from different directions. Included is the case of a single moving camera. Though various techniques can be used to implement this nonlinear estimation, we discuss the use of gradient descent. Experiments are run and discussed for the case of a sphere of unknown location. It is shown that the estimation procedure can be viewed geometrically as a cross correlation of nonlinearly transformed image patches in two or more images. Approaches to object surface segmentation into primitive object surfaces, and primitive object-type recognition are briefly presented and discussed. The attractiveness of the approach is that maximum likelihood estimation and all the usual tools of statistical signal analysis can be brought to bear, the information extraction appears to be robust and computationally reasonable, the concepts are geometric and simple, and close to optimal accuracy should result.
Bruno Cernuschi-Frías, Peter N. Belhumeur, David B. Cooper
ICRA2