David J. Kriegman

dblp:k/DavidJKriegman · DBLP profile ↗
← Back
157ranked-venue papers
28as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 138 · 27 first-authorGraphics, computer vision, multimedia, augmented reality and games · 81 · 3 first-authorSystems, architecture and hardware · 11 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4Security and privacy · 2Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
71 papers
3D vision · 44% Face, body and person analysis · 13% Segmentation and scene understanding · 7%
Computer graphics and multimedia
49 papers
Computational photography and imaging · 46% Image and video processing · 25% Rendering · 15%
Theoretical computer science
10 papers
Mathematical optimization · 70% Graph algorithms and graph theory · 24% Computational geometry · 4%

Topics — the 30 heaviest of 184, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational photography and imaging
photometric stereo
0.792015
Photometric Stereo in a Scattering Medium · ICCV 2015
Photometric stereo with non-parametric and spatially-varying reflectance · CVPR 2008
Shape from Varying Illumination and Viewpoint · ICCV 2007
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.532020
Deep 3D Capture: Geometry and Reflectance From Sparse Multi-View Images · CVPR 2020
Shape from Varying Illumination and Viewpoint · ICCV 2007
Toward a Stratification of Helmholtz Stereopsis · CVPR (1) 2003
Computer vision › 3D vision
photometric stereo
0.542017
Photometric Stereo in a Scattering Medium · IEEE Trans. Pattern Anal. Mach. Intell. 2017
ShadowCuts: Photometric Stereo with Shadows · CVPR 2007
Beyond Lambert: Reconstructing Specular Surfaces Using Color · CVPR (2) 2005
Image and video processing
image restoration
0.532017
Depth and Image Restoration from Light Field in a Scattering Medium · ICCV 2017
Personal photo enhancement using example images · ACM Trans. Graph. 2010
Specularity Removal in Images and Videos: A PDE Approach · ECCV (1) 2006
Computer vision › 3D vision
3d reconstruction
0.422020
Deep 3D Capture: Geometry and Reflectance From Sparse Multi-View Images · CVPR 2020
Reconstruction of HOT curves from image sequences · CVPR 1993
Computational photography and imaging › reflectance acquisition
reflectance estimation
0.412020
Deep 3D Capture: Geometry and Reflectance From Sparse Multi-View Images · CVPR 2020
Rendering › appearance acquisition › material acquisition
SVBRDF estimation
0.412020
Deep 3D Capture: Geometry and Reflectance From Sparse Multi-View Images · CVPR 2020
Image and video processing › image restoration
image deblurring
0.442017
Photometric Stereo in a Scattering Medium · ICCV 2015
Image deblurring and denoising using color priors · CVPR 2009
Photometric Stereo in a Scattering Medium · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Computer vision › 3D vision
3d shape reconstruction
0.452017
Photometric Stereo in a Scattering Medium · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Isotropy, Reciprocity and the Generalized Bas-Relief Ambiguity · CVPR 2007
Beyond Lambert: Reconstructing Surfaces with Arbitrary BRDFs · ICCV 2001
Computer vision › Face, body and person analysis
face recognition
0.4102011
Pose, illumination and expression invariant pairwise face-similarity measure via Doppelgänger list comparison · ICCV 2011
Face Recognition Using 3-D Models: Pose and Illumination · Proc. IEEE 2006
Acquiring Linear Subspaces for Face Recognition under Variable Lighting · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Computer vision › Segmentation and scene understanding › semantic segmentation › transfer learning for semantic segmentation
domain adaptive semantic segmentation
0.312018
Image to Image Translation for Domain Adaptation · CVPR 2018
Computer vision › Segmentation and scene understanding
semantic segmentation
0.312018
Image to Image Translation for Domain Adaptation · CVPR 2018
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.312018
Image to Image Translation for Domain Adaptation · CVPR 2018
Computer vision › Face, body and person analysis
face alignment
0.322013
Localizing Parts of Faces Using a Consensus of Exemplars · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Localizing parts of faces using a consensus of exemplars · CVPR 2011
Computational photography and imaging › computational optics
imaging through scattering media
0.312017
Photometric Stereo in a Scattering Medium · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Computational photography and imaging › light field imaging
light field depth estimation
0.312017
Depth and Image Restoration from Light Field in a Scattering Medium · ICCV 2017
Computational photography and imaging
light field imaging
0.312017
Depth and Image Restoration from Light Field in a Scattering Medium · ICCV 2017
Geometric modeling and processing
3d reconstruction
0.342015
Shape from Varying Illumination and Viewpoint · ICCV 2007
Globally Optimal Affine and Metric Upgrades in Stratified Autocalibration · ICCV 2007
Photometric Stereo in a Scattering Medium · ICCV 2015
Computer vision › 3D vision
structure from motion
0.272008
Robust Structure and Motion from Outlines of Smooth Curved Surfaces · IEEE Trans. Pattern Anal. Mach. Intell. 2006
Structure and View Estimation for Tomographic Reconstruction: A Bayesian Approach · CVPR (2) 2006
Structure and Motion from Images of Smooth Textureless Objects · ECCV (2) 2004
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.222013
Localizing Parts of Faces Using a Consensus of Exemplars · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Structure and View Estimation for Tomographic Reconstruction: A Bayesian Approach · CVPR (2) 2006
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › semantic embedding
concept embedding
0.212015
Learning Concept Embeddings with Combined Human-Machine Expertise · ICCV 2015
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.212014
Guess-Averse Loss Functions For Cost-Sensitive Multiclass Boosting · ICML 2014
Machine learning › Learning paradigms › cost-sensitive learning
cost-sensitive classification
0.212014
Guess-Averse Loss Functions For Cost-Sensitive Multiclass Boosting · ICML 2014
Machine learning › Learning theory
loss function
0.212014
Guess-Averse Loss Functions For Cost-Sensitive Multiclass Boosting · ICML 2014
Machine learning › Kernel, tree and ensemble methods › ensemble learning › boosting
multiclass boosting
0.212014
Guess-Averse Loss Functions For Cost-Sensitive Multiclass Boosting · ICML 2014
Computer vision › 3D vision
camera calibration
0.222010
Globally Optimal Algorithms for Stratified Autocalibration · Int. J. Comput. Vis. 2010
Autocalibration via Rank-Constrained Estimation of the Absolute Quadric · CVPR 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
non-parametric methods
0.212013
Localizing Parts of Faces Using a Consensus of Exemplars · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Computational photography and imaging › shape and reflectance estimation
shape from shading
0.222012
Shape from Fluorescence · ECCV (7) 2012
Color Subspaces as Photometric Invariants · CVPR (2) 2006
Computer vision › 3D vision
multi-view geometry
0.222008
Practical Global Optimization for Multiview Geometry · Int. J. Comput. Vis. 2008
Synthetic Aperture Tracking: Tracking through Occlusions · ICCV 2007
Image and video processing › color image processing
illumination invariance
0.122008
Color Subspaces as Photometric Invariants · Int. J. Comput. Vis. 2008
Color Subspaces as Photometric Invariants · CVPR (2) 2006

Methods — techniques the papers use, named apart from their topics

photometric stereo · 1.0reflectance volumes · 0.9photometric optimization · 0.9multi-view reflectance network · 0.9deep multi-view stereo · 0.9deconvolution · 0.8point spread function estimation · 0.6image-to-image translation · 0.3feature regularization · 0.3adversarial learning · 0.3transmission-based depth cue · 0.3light field imaging · 0.3branch-and-bound · 0.3fluorescence imaging · 0.2convex relaxation · 0.2global optimization · 0.1multi-scale texture descriptors · 0.1color descriptors · 0.1
YearPublicationVenuePosition
2023 One-Vote Veto: Semi-Supervised Learning for Low-Shot Glaucoma Diagnosis
abstract
Convolutional neural networks (CNNs) are a promising technique for automated glaucoma diagnosis from images of the fundus, and these images are routinely acquired as part of an ophthalmic exam. Nevertheless, CNNs typically require a large amount of well-labeled data for training, which may not be available in many biomedical image classification applications, especially when diseases are rare and where labeling by experts is costly. This article makes two contributions to address this issue: 1) It extends the conventional Siamese network and introduces a training method for low-shot learning when labeled data are limited and imbalanced, and 2) it introduces a novel semi-supervised learning strategy that uses additional unlabeled training data to achieve greater accuracy. Our proposed multi-task Siamese network (MTSN) can employ any backbone CNN, and we demonstrate with four backbone CNNs that its accuracy with limited training data approaches the accuracy of backbone CNNs trained with a dataset that is 50 times larger. We also introduce One-Vote Veto (OVV) self-training, a semi-supervised learning strategy that is designed specifically for MTSNs. By taking both self-predictions and contrastive predictions of the unlabeled training data into account, OVV self-training provides additional pseudo labels for fine-tuning a pre-trained MTSN. Using a large (imbalanced) dataset with 66,715 fundus photographs acquired over 15 years, extensive experimental results demonstrate the effectiveness of low-shot learning with MTSN and semi-supervised learning with OVV self-training. Three additional, smaller clinical datasets of fundus images acquired under different conditions (cameras, instruments, locations, populations) are used to demonstrate the generalizability of the proposed methods.
Rui Fan 0001, Christopher Bowd, Nicole Brye, Mark Christopher, Robert N. Weinreb, David J. Kriegman, Linda M. Zangwill
IEEE Trans. Medical Imaging6
2020 Deep 3D Capture: Geometry and Reflectance From Sparse Multi-View Images
abstract
We introduce a novel learning-based method to reconstruct the high-quality geometry and complex, spatially-varying BRDF of an arbitrary object from a sparse set of only six images captured by wide-baseline cameras under collocated point lighting. We first estimate per-view depth maps using a deep multi-view stereo network; these depth maps are used to coarsely align the different views. We propose a novel multi-view reflectance estimation network architecture that is trained to pool features from these coarsely aligned images and predict per-view spatially-varying diffuse albedo, surface normals, specular roughness and specular albedo. We do this by jointly optimizing the latent space of our multi-view reflectance network to minimize the photometric error between images rendered with our predictions and the input images. While previous state-of-the-art methods fail on such sparse acquisition setups, we demonstrate, via extensive experiments on synthetic and real data, that our method produces high-quality reconstructions that can be used to render photorealistic images.
Sai Bi, Zexiang Xu, Kalyan Sunkavalli, David J. Kriegman, Ravi Ramamoorthi
CVPR4
2020 Deep Reflectance Volumes: Relightable Reconstructions from Multi-view Photometric Images
Sai Bi, Zexiang Xu, Kalyan Sunkavalli, Milos Hasan, Yannick Hold-Geoffroy, David J. Kriegman, Ravi Ramamoorthi
ECCV (3)6
2020 Detecting the Starting Frame of Actions in Video
abstract
In this work, we address the problem of precisely localizing key frames of an action, for example, the precise time that a pitcher releases a baseball, or the precise time that a crowd begins to applaud. Key frame localization is a largely overlooked and important action-recognition problem, for example in the field of neuroscience, in which we would like to understand the neural activity that produces the start of a bout of an action. To address this problem, we introduce a novel structured loss function that properly weights the types of errors that matter in such applications: it more heavily penalizes extra and missed action start detections over small misalignments. Our structured loss is based on the best matching between predicted and labeled action starts. We train recurrent neural networks (RNNs) to minimize differentiable approximations of this loss. To evaluate these methods, we introduce the Mouse Reach Dataset, a large, annotated video dataset of mice performing a sequence of actions. The dataset was collected and labeled by experts for the purpose of neuroscience research. On this dataset, we demonstrate that our method outperforms related approaches and baseline methods using an unstructured loss.
Iljung S. Kwak, Jian-Zhong Guo, Adam W. Hantman, Kristin Branson, David J. Kriegman
WACV5
2018 Image to Image Translation for Domain Adaptation
abstract
We propose a general framework for unsupervised domain adaptation, which allows deep neural networks trained on a source domain to be tested on a different target domain without requiring any training annotations in the target domain. This is achieved by adding extra networks and losses that help regularize the features extracted by the backbone encoder network. To this end we propose the novel use of the recently proposed unpaired image-to-image translation framework to constrain the features extracted by the encoder network. Specifically, we require that the features extracted are able to reconstruct the images in both domains. In addition we require that the distribution of features extracted from images in the two domains are indistinguishable. Many recent works can be seen as specific cases of our general framework. We apply our method for domain adaptation between MNIST, USPS, and SVHN datasets, and Amazon, Webcam and DSLR Office datasets in classification tasks, and also between GTA5 and Cityscapes datasets for a segmentation task. We demonstrate state of the art performance on each of these datasets.
Zak Murez, Soheil Kolouri, David J. Kriegman, Ravi Ramamoorthi, Kyungnam Kim
CVPR3
2018 Learning to See Through Turbulent Water
abstract
Imaging through dynamic refractive media, such as looking into turbulent water, or through hot air, is challenging since light rays are bent by unknown amounts leading to complex geometric distortions. Inverting these distortions and recovering high quality images is an inherently ill-posed problem, leading previous works to require extra information such as high frame-rate video or a template image, which limits their applicability in practice. This paper proposes training a deep convolution neural network to undistort dynamic refractive effects using only a single image. The neural network is able to solve this ill-posed problem by learning image priors as well as distortion priors. Our network consists of two parts, a warping net to remove geometric distortion and a color predictor net to further refine the restoration. Adversarial loss is used to achieve better visual quality and help the network hallucinate missing and blurred information. To train our network, we collect a large training set of images distorted by a turbulent water surface. Unlike prior works on water undistortion, our method is trained end-to-end, only requires a single image and does not use a ground truth template at test time. Experiments show that by exploiting the structure of the problem, our network outperforms state-of-the-art deep image to image translation.
Zhengqin Li, Zak Murez, David J. Kriegman, Ravi Ramamoorthi, Manmohan Krishna Chandraker
WACV3
2017 Depth and Image Restoration from Light Field in a Scattering Medium
abstract
Traditional imaging methods and computer vision algorithms are often ineffective when images are acquired in scattering media, such as underwater, fog, and biological tissue. Here, we explore the use of light field imaging and algorithms for image restoration and depth estimation that address the image degradation from the medium. Towards this end, we make the following three contributions. First, we present a new single image restoration algorithm which removes backscatter and attenuation from images better than existing methods do, and apply it to each view in the light field. Second, we combine a novel transmission based depth cue with existing correspondence and defocus cues to improve light field depth estimation. In densely scattering media, our transmission depth cue is critical for depth estimation since the images have low signal to noise ratios which significantly degrades the performance of the correspondence and defocus cues. Finally, we propose shearing and refocusing multiple views of the light field to recover a single image of higher quality than what is possible from a single view. We demonstrate the benefits of our method through extensive experimental results in a water tank.
Jiandong Tian, Zak Murez, Tong Cui, David J. Kriegman, Ravi Ramamoorthi
ICCV5
2017 Photometric Stereo in a Scattering Medium
abstract
Photometric stereo is widely used for 3D reconstruction. However, its use in scattering media such as water, biological tissue and fog has been limited until now, because of forward scattered light from both the source and object, as well as light scattered back from the medium (backscatter). Here we make three contributions to address the key modes of light propagation, under the common single scattering assumption for dilute media. First, we show through extensive simulations that single-scattered light from a source can be approximated by a point light source with a single direction. This alleviates the need to handle light source blur explicitly. Next, we model the blur due to scattering of light from the object. We measure the object point-spread function and introduce a simple deconvolution method. Finally, we show how imaging fluorescence emission where available, eliminates the backscatter component and increases the signal-to-noise ratio. Experimental results in a water tank, with different concentrations of scattering media added, show that deconvolution produces higher-quality 3D reconstructions than previous techniques, and that when combined with fluorescence, can produce results similar to that in clear water even for highly turbid media.
Zak Murez, Tali Treibitz, Ravi Ramamoorthi, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Dense Volume-to-Volume Vascular Boundary Detection
Jameson Tyler Merkow, Alison L. Marsden, David J. Kriegman, Zhuowen Tu
MICCAI (3)3
2015 Photometric Stereo in a Scattering Medium
abstract
Photometric stereo is widely used for 3D reconstruction. However, its use in scattering media such as water, biological tissue and fog has been limited until now, because of forward scattered light from both the source and object, as well as light scattered back from the medium (backscatter). Here we make three contributions to address the key modes of light propagation, under the common single scattering assumption for dilute media. First, we show through extensive simulations that single-scattered light from a source can be approximated by a point light source with a single direction. This alleviates the need to handle light source blur explicitly. Next, we model the blur due to scattering of light from the object. We measure the object point-spread function and introduce a simple deconvolution method. Finally, we show how imaging fluorescence emission where available, eliminates the backscatter component and increases the signal-to-noise ratio. Experimental results in a water tank, with different concentrations of scattering media added, show that deconvolution produces higher-quality 3D reconstructions than previous techniques, and that when combined with fluorescence, can produce results similar to that in clear water even for highly turbid media.
Zak Murez, Tali Treibitz, Ravi Ramamoorthi, David J. Kriegman
ICCV4
2015 Learning Concept Embeddings with Combined Human-Machine Expertise
abstract
This paper presents our work on "SNaCK," a low-dimensional concept embedding algorithm that combines human expertise with automatic machine similarity kernels. Both parts are complimentary: human insight can capture relationships that are not apparent from the object's visual similarity and the machine can help relieve the human from having to exhaustively specify many constraints. We show that our SNaCK embeddings are useful in several tasks: distinguishing prime and nonprime numbers on MNIST, discovering labeling mistakes in the Caltech UCSD Birds (CUB) dataset with the help of deep-learned features, creating training datasets for bird classifiers, capturing subjective human taste on a new dataset of 10,000 foods, and qualitatively exploring an unstructured set of pictographic characters. Comparisons with the state-of-the-art in these tasks show that SNaCK produces better concept embeddings that require less human supervision than the leading methods.
Kimberly Wilber, Iljung S. Kwak, David J. Kriegman, Serge J. Belongie
ICCV3
2015 Structural Edge Detection for Cardiovascular Modeling
Jameson Tyler Merkow, Zhuowen Tu, David J. Kriegman, Alison L. Marsden
MICCAI (3)3
2014 Adaptive ranking of facial attractiveness
abstract
As humans, we love to rank things. Top ten lists exist for everything from movie stars to scary animals. Ambiguities (i.e., ties) naturally occur in the process of ranking when people feel they cannot distinguish two items. Human reported rankings derived from star ratings abound on recommendation websites such as Yelp and Netflix. However, those websites differ in star precision which points to the need for ranking systems that adapt to an individual user's preference sensitivity. In this work we propose an adaptive system that allows for ties when collecting ranking data. Using this system, we propose a framework for obtaining computer-generated rankings. We test our system and a computer-generated ranking method on the problem of evaluating human attractiveness. Extensive experimental evaluations and analysis demonstrate the effectiveness and efficiency of our work.
Iljung S. Kwak, Serge J. Belongie, David J. Kriegman, Haizhou Ai
ICME4
2014 Guess-Averse Loss Functions For Cost-Sensitive Multiclass Boosting
abstract
Cost-sensitive multiclass classification has recently acquired significance in several applications, through the introduction of multiclass datasets with well-defined misclassification costs. The design of classification algorithms for this setting is considered. It is argued that the unreliable performance of current algorithms is due to the inability of the underlying loss functions to enforce a certain fundamental underlying property. This property, denoted guess-aversion, is that the loss should encourage correct classifications over the arbitrary guessing that ensues when all classes are equally scored by the classifier. While guess-aversion holds trivially for binary classification, this is not true in the multiclass setting. A new family of cost-sensitive guess-averse loss functions is derived, and used to design new cost-sensitive multiclass boosting algorithms, denoted GEL- and GLL-MCBoost. Extensive experiments demonstrate (1) the general importance of guess-aversion and (2) that the GLL loss function outperforms other loss functions for multiclass boosting.
Oscar Beijbom, Mohammad J. Saberian, David J. Kriegman, Nuno Vasconcelos
ICML3
2013 Match-time covariance for descriptors
Eric M. Christiansen, Vincent C. Rabaud, Andrew Ziegler, David J. Kriegman, Serge J. Belongie
BMVC4
2013 From Bikers to Surfers: Visual Recognition of Urban Tribes
abstract
Iljung S. Kwak1 [email protected] Ana C. Murillo2 [email protected] Peter N. Belhumeur3 [email protected] David Kriegman1 [email protected] Serge Belongie1 [email protected] 1 Dept. of Computer Science and Engineering University of California, San Diego, USA. 2 Dpt. Informatica e Ing. Sistemas Inst. Investigacion en Ingenieria de Aragon. University of Zaragoza, Spain. 3 Department of Computer Science Columbia University, USA.
Iljung S. Kwak, Ana Cristina Murillo, Peter N. Belhumeur, David J. Kriegman, Serge J. Belongie
BMVC4
2013 Localizing Parts of Faces Using a Consensus of Exemplars
abstract
We present a novel approach to localizing parts in images of human faces. The approach combines the output of local detectors with a nonparametric set of global models for the part locations based on over 1,000 hand-labeled exemplar images. By assuming that the global models generate the part locations as hidden variables, we derive a Bayesian objective function. This function is optimized using a consensus of models for these hidden variables. The resulting localizer handles a much wider range of expression, pose, lighting, and occlusion than prior ones. We show excellent performance on real-world face datasets such as Labeled Faces in the Wild (LFW) and a new Labeled Face Parts in the Wild (LFPW) and show that our localizer achieves state-of-the-art performance on the less challenging BioID dataset.
Peter N. Belhumeur, David Jacobs 0001, David J. Kriegman, Neeraj Kumar 0006
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Automated annotation of coral reef survey images
abstract
With the proliferation of digital cameras and automatic acquisition systems, scientists can acquire vast numbers of images for quantitative analysis. However, much image analysis is conducted manually, which is both time consuming and prone to error. As a result, valuable scientific data from many domains sit dormant in image libraries awaiting annotation. This work addresses one such domain: coral reef coverage estimation. In this setting, the goal, as defined by coral reef ecologists, is to determine the percentage of the reef surface covered by rock, sand, algae, and corals; it is often desirable to resolve these taxa at the genus level or below. This is challenging since the data exhibit significant within class variation, the borders between classes are complex, and the viewpoints and image quality vary. We introduce Moorea Labeled Corals, a large multi-year dataset with 400,000 expert annotations, to the computer vision community, and argue that this type of ecological data provides an excellent opportunity for performance benchmarking. We also propose a novel algorithm using texture and color descriptors over multiple scales that outperforms commonly used techniques from the texture classification literature. We show that the proposed algorithm accurately estimates coral coverage across locations and years, thereby taking a significant step towards reliable automated coral reef image annotation.
Oscar Beijbom, Peter J. Edmunds, David I. Kline, B. Greg Mitchell, David J. Kriegman
CVPR5
2012 Shape from Fluorescence
Tali Treibitz, Zak Murez, B. Greg Mitchell, David J. Kriegman
ECCV (7)4
2012 Locally Uniform Comparison Image Descriptor
abstract
Keypoint matching between pairs of images using popular descriptors like SIFT or a faster variant called SURF is at the heart of many computer vision algorithms including recognition, mosaicing, and structure from motion. For real-time mobile applications, very fast but less accurate descriptors like BRIEF and related methods use a random sampling of pairwise comparisons of pixel intensities in an image patch. Here, we introduce Locally Uniform Comparison Image Descriptor (LUCID), a simple description method based on permutation distances between the ordering of intensities of RGB values between two patches. LUCID is computable in linear time with respect to patch size and does not require floating point computation. An analysis reveals an underlying issue that limits the potential of BRIEF and related approaches compared to LUCID. Experiments demonstrate that LUCID is faster than BRIEF, and its accuracy is directly comparable to SURF while being more than an order of magnitude faster.
Andrew Ziegler, Eric M. Christiansen, David J. Kriegman, Serge J. Belongie
NIPS3
2011 Localizing parts of faces using a consensus of exemplars
abstract
We present a novel approach to localizing parts in images of human faces. The approach combines the output of local detectors with a non-parametric set of global models for the part locations based on over one thousand hand-labeled exemplar images. By assuming that the global models generate the part locations as hidden variables, we derive a Bayesian objective function. This function is optimized using a consensus of models for these hidden variables. The resulting localizer handles a much wider range of expression, pose, lighting and occlusion than prior ones. We show excellent performance on a new dataset gathered from the internet and show that our localizer achieves state-of-the-art performance on the less challenging BioID dataset.
Peter N. Belhumeur, David Jacobs 0001, David J. Kriegman, Neeraj Kumar 0006
CVPR3
2011 Wet fingerprint recognition: Challenges and opportunities
abstract
Many fingers wrinkle or shrivel when immersed in water. When used for biometric identification, the recognition rate for wrinkled fingers degrades. The impact of wrinkling has so far not been well-understood. In this study, we present an investigation of how the finger-skin expansion due to wrinkling impacts the quality of scanned finger prints and characterize the qualitative changes that affect recognition. We also introduce the Wet and Wrinkled Finger (WWF) database that we will make available to other researchers. In this database of 300 fingers, 185 are visibly wrinkled after immersion; multiple images of dry and immersed fingerprints were acquired. In this paper, we present baseline recognition rates on WWF using two algorithms a commercial fingerprint recognition algorithm and the publicly available Bozorth3 matcher. Specifically, we show a degradation in accuracy with both algorithms when comparing Dry-finger to Dry finger verification with Dry-finger to Wet-finger verification. We analyze performance on a per-finger basis and note a difference in accuracy amongst fingers, and as consequence make recommendations about which fingers to use in environments where fingers are apt to be wet. Additionally, we propose an implementation of a classifier that can decide if the incoming query is wrinkled.
Prasanna Krishnasamy, Serge J. Belongie, David J. Kriegman
IJCB3
2011 Two faces are better than one: Face recognition in group photographs
abstract
Face recognition systems classically recognize people individually. When presented with a group photograph containing multiple people, such systems implicitly assume statistical independence between each detected face. We question this basic assumption and consider instead that there is a dependence between face regions from the same image; after all, the image was acquired with a single camera, under consistent lighting (distribution, direction, spectrum), camera motion, and scene/camera geometry. Such naturally occurring commonalities between face images can be exploited when recognition decisions are made jointly across the faces, rather than independently. Furthermore, when recognizing people in isolation, some features such as color are usually uninformative in unconstrained settings. But by considering pairs of people, the relative color difference provides valuable information. This paper reconsiders the independence assumption, introduces new features and methods for recognizing pairs of individuals in group photographs, and demonstrates a marked improvement when these features are used in joint decision making vs. independent decision making. While these features alone are only moderately discriminative, we combine these new features with state-of art attribute features and demonstrate effective recognition performance. Initial experiments on two datasets show promising improvements in accuracy.
Ohil K. Manyam, Neeraj Kumar 0006, Peter N. Belhumeur, David J. Kriegman
IJCB4
2011 Pose, illumination and expression invariant pairwise face-similarity measure via Doppelgänger list comparison
abstract
Face recognition approaches have traditionally focused on direct comparisons between aligned images, e.g. using pixel values or local image features. Such comparisons become prohibitively difficult when comparing faces across extreme differences in pose, illumination and expression. The goal of this work is to develop a face-similarity measure that is largely invariant to these differences. We propose a novel data driven method based on the insight that comparing images of faces is most meaningful when they are in comparable imaging conditions. To this end we describe an image of a face by an ordered list of identities from a Library. The order of the list is determined by the similarity of the Library images to the probe image. The lists act as a signature for each face image: similarity between face images is determined via the similarity of the signatures. Here the CMU Multi-PIE database, which includes images of 337 individuals in more than 2000 pose, lighting and illumination combinations, serves as the Library. We show improved performance over state of the art face-similarity measures based on local features, such as FPLBP, especially across large pose variations on FacePix and multi-PIE. On LFW we show improved performance in comparison with measures like SIFT (on fiducials), LBP, FPLBP and Gabor (C1).
Florian Schroff, Tali Treibitz, David J. Kriegman, Serge J. Belongie
ICCV3
2011 Introduction to the Special Section on Real-World Face Recognition
abstract
The motivations for organizing this special section were to better address the challenges of face recognition in real-world scenarios, to promote systematic research and evaluation of promising methods and systems, to provide a snapshot of where we are in this domain, and to stimulate discussion about future directions. We solicited original contributions of research on all aspects of real-world face recognition, including: the design of robust face similarity features and metrics; robust face clustering and sorting algorithms; novel user interaction models and face recognition algorithms for face tagging; novel applications of web face recognition; novel computational paradigms for face recognition; challenges in large scale face recognition tasks, e.g., on the Internet; face recognition with contextual information; face recognition benchmarks and evaluation methodology for moderately controlled or uncontrolled environments; and video face recognition. We received 42 original submissions, four of which were rejected without review; the other 38 papers entered the normal review process. Each paper was reviewed by three reviewers who are experts in their respective topics. More than 100 expert reviewers have been involved in the review process. The papers were equally distributed among the guest editors. A final decision for each paper was made by at least two guest editors assigned to it. To avoid conflict of interest, no guest editor submitted any papers to this special section.
Gang Hua 0001, Ming-Hsuan Yang 0001, Erik G. Learned-Miller, Yi Ma 0001, Matthew Turk 0001, David J. Kriegman, Thomas S. Huang
IEEE Trans. Pattern Anal. Mach. Intell.6
2010 Globally Optimal Algorithms for Stratified Autocalibration
abstract
We present practical algorithms for stratified autocalibration with theoretical guarantees of global optimality. Given a projective reconstruction, we first upgrade it to affine by estimating the position of the plane at infinity. The plane at infinity is computed by globally minimizing a least squares formulation of the modulus constraints. In the second stage, this affine reconstruction is upgraded to a metric one by globally minimizing the infinite homography relation to compute the dual image of the absolute conic (DIAC). The positive semidefiniteness of the DIAC is explicitly enforced as part of the optimization process, rather than as a post-processing step. For each stage, we construct and minimize tight convex relaxations of the highly non-convex objective functions in a branch and bound optimization framework. We exploit the inherent problem structure to restrict the search space for the DIAC and the plane at infinity to a small, fixed number of branching dimensions, independent of the number of views. Chirality constraints are incorporated into our convex relaxations to automatically select an initial region which is guaranteed to contain the global minimum. Experimental evidence of the accuracy, speed and scalability of our algorithm is presented on synthetic and real data.
Manmohan Krishna Chandraker, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
Int. J. Comput. Vis.3
2010 Personal photo enhancement using example images
abstract
We describe a framework for improving the quality of personal photos by using a person's favorite photographs as examples. We observe that the majority of a person's photographs include the faces of a photographer's family and friends and often the errors in these photographs are the most disconcerting. We focus on correcting these types of images and use common faces across images to automatically perform both global and face-specific corrections. Our system achieves this by using face detection to align faces between “good” and “bad” photos such that properties of the good examples can be used to correct a bad photo. These “personal” photos provide strong guidance for a number of operations and, as a result, enable a number of high-quality image processing operations. We illustrate the power and generality of our approach by presenting a novel deblurring algorithm, and we show corrections that perform sharpening, superresolution, in-painting of over- and underexposured regions, and white-balancing.
Neel Joshi, Wojciech Matusik, Edward H. Adelson, David J. Kriegman
ACM Trans. Graph.4
2009 Image deblurring and denoising using color priors
abstract
Image blur and noise are difficult to avoid in many situations and can often ruin a photograph. We present a novel image deconvolution algorithm that deblurs and denoises an image given a known shift-invariant blur kernel. Our algorithm uses local color statistics derived from the image as a constraint in a unified framework that can be used for deblurring, denoising, and upsampling. A pixel's color is required to be a linear combination of the two most prevalent colors within a neighborhood of the pixel. This two-color prior has two major benefits: it is tuned to the content of the particular image and it serves to decouple edge sharpness from edge strength. Our unified algorithm for deblurring and denoising out-performs previous methods that are specialized for these individual applications. We demonstrate this with both qualitative results and extensive quantitative comparisons that show that we can out-perform previous methods by approximately 1 to 3 DB.
Neel Joshi, C. Lawrence Zitnick, Richard Szeliski, David J. Kriegman
CVPR4
2009 Moving in stereo: Efficient structure and motion using lines
abstract
We present a fast and robust system for estimating structure and motion using a stereo pair, with straight lines as features. Our first set of contributions are efficient algorithms to perform this estimation using a few (two or three) lines, which are well-suited for use in a hypothesize-and-test framework. Our second contribution is the design of an efficient structure from motion system that performs robustly in complex indoor environments. By using infinite lines rather than line segments, our approach avoids the issues arising due to uncertain determination of end-points. Our cost function stems from a rank condition on planes backprojected from corresponding image lines. We propose a framework that imposes orthonormality constraints on the rigid body motion and can perform estimation using only two or three lines, through efficient solution of an overdetermined system of polynomials. This is in contrast to simple approaches which first reconstruct 3D lines and then align them, but perform poorly in real-world scenes with narrow baseline stereo. Experiments using synthetic as well as real data demonstrate the speed, accuracy and reliability of our system.
Manmohan Krishna Chandraker, Jongwoo Lim, David J. Kriegman
ICCV3
2009 Introduction of New Editor-in-Chief
abstract
I am pleased to announce that, upon the recommendation of a search committee, the President of the IEEE Computer Society has appointed Professor Ramin Zabih as the new Editor-in-Chief (EIC) of the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Dr. Zabih is well known for his research on discrete optimization techniques and their application to computer vision and medical imaging. He brings to TPAMI leadership experience as a program cochair of the IEEE Conference on Computer Vision and Pattern Recognition and as a guest editor and then as an associate editor (AE) of TPAMI. Over the years, I have worked closely with Ramin and seen that he has a broad view of the fi eld, that he has very high standards, that he adjudicates matters fairly, and that he puts the needs of the community in the forefront. Taken together, I have very high expectations for the direction and success of TPAMI under his leadership. Please join me in congratulating Dr. Zabih on his appointment. His brief bio appears below. Dr. Zabih joins TPAMI when it is in good health with 850 submissions in 2007 and 915 projected for 2008. The acceptance rate is about 24 percent. Papers in general are getting through the review process quickly with an average time of three months from submission to fi rst decision and six months to a fi nal decision. This compares favorably with the review cycle for major conferences. However, there is greater variance at a journal since there isn’t a hard deadline and most accepted papers undergo revisions and a second round of review. While metrics defi ning the importance of an academic journal are contentious, the most widely used one is the impact factor, which is the average number of times papers published in the two previous years are referenced by publications in the given year. Based on the publications tracked by Thomson ISI, TPAMI papers were cited 16,492 times in 2007 and the impact factor in 2007 was 3.579. TPAMI has the second highest impact factor of all IEEE publications; it ranks second in all of electrical engineering, second in artifi cial intelligence, and seventh in all of computer science. December 31, 2008 marks the end of my term as the EIC of TPAMI and of nearly 12 years of continuous service on the TPAMI Editorial Board. I’d like to thank my predecessors who helped create an impressive journal that I have had the privilege of cultivating. I was fi rst appointed as a TPAMI AE under Professor Rangachar Kasturi, continued under Professor Kevin Bowyer, and then served as Associate Editor-in-Chief under Professor Rama Chellappa. Under Professor Kasturi, I learned how to fairly and effectively manage the review process of a paper. Under Professor Bowyer, I learned how critical it is to have an effective process and staff. Under Professor Chellappa, I learned to how to scale the whole thing as the number of submissions soared. There are three people that I have worked most closely with as EIC and who deserve special mention and thanks. First, when I became EIC, I selected David Fleet as the Associate Editor-in-Chief and he has been integral part of TPAMI from making strategic decisions to putting out fi res. Seeing both the growth in TPAMI submissions to more than 900 manuscripts and the increasing importance of Machine Learning to TPAMI, I was pleased when Zoubin Gharamani accepted the appointment as a second AEIC. You have both been amazing. The smooth operations and reduction in review time can largely be attributed to Elaine Stephenson. From its fi rst upload to fi nal decision, 30-40 e-mails are sent for every paper. Elaine’s no-nonsense approach and commitment to process have resulted in a reduction of the time of a submission to fi rst decision from seven months in 2003 to three months today. There have been two notable transitions over the past year at the Computer Society: Jennifer Carruth replaced Suzanne Wagner as Peer Review Supervisor and Kathy Santa Maria replaced Julie Hicks as the Senior Production Editor in charge of the title; their titles indicate their important roles to TPAMI and their support, which is gratefully appreciated. A huge amount of volunteer effort goes into the TPAMI review process and a back of the envelope estimate is that over 34,000 hours (equivalent to 17 person years) are spent each year on the review process (based on three reviews, 8 hours per review, major revisions, 10 hours of Associate Editor effort, etc.). I really appreciate the hard work of the reviewers and the dedication of the Associate Editors who have to make very diffi cult decisions and work with the authors to improve their papers. Finally, I’d like to thank my wife Teresa and children Bryce and Dylan for their patience and forbearance which allowed me to focus on TPAMI, often late into the night.
David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Toward a perceptual space for gloss
abstract
We design and implement a comprehensive study of the perception of gloss. This is the largest study of its kind to date, and the first to use real material measurements. In addition, we develop a novel multi-dimensional scaling (MDS) algorithm for analyzing pairwise comparisons. The data from the psychophysics study and the MDS algorithm is used to construct a low dimensional perceptual embedding of these bidirectional reflectance distribution functions (BRDFs). The embedding is validated by correlating it with nine gloss dimensions, fitted parameters of seven analytical BRDF models, and a perceptual parameterization of Ward's model. We also introduce a novel perceptual interpolation scheme that uses the embedding to provide the user with an intuitive interface for navigating the space of gloss and constructing new materials.
Josh Wills, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
ACM Trans. Graph.3
2008 Photometric stereo with non-parametric and spatially-varying reflectance
abstract
We present a method for simultaneously recovering shape and spatially varying reflectance of a surface from photometric stereo images. The distinguishing feature of our approach is its generality; it does not rely on a specific parametric reflectance model and is therefore purely ldquodata-drivenrdquo. This is achieved by employing novel bi-variate approximations of isotropic reflectance functions. By combining this new approximation with recent developments in photometric stereo, we are able to simultaneously estimate an independent surface normal at each point, a global set of non-parametric ldquobasis materialrdquo BRDFs, and per-point material weights. Our experimental results validate the approach and demonstrate the utility of bi-variate reflectance functions for general non-parametric appearance capture.
Neil Gordon Alldrin, Todd E. Zickler, David J. Kriegman
CVPR3
2008 Globally optimal bilinear programming for computer vision applications
abstract
We present a practical algorithm that provably achieves the global optimum for a class of bilinear programs commonly arising in computer vision applications. Our approach relies on constructing tight convex relaxations of the objective function and minimizing it in a branch and bound framework. A key contribution of the paper is a novel, provably convergent branching strategy that allows us to solve large-scale problems by restricting the branching dimensions to just one set of variables constituting the bilinearity. Experiments with synthetic and real data validate our claims of optimality, speed and convergence. We contrast the optimality of our solutions with those obtained by a traditional singular value decomposition approach. Among several potential applications, we discuss two: exemplar-based face reconstruction and non-rigid structure from motion. In both cases, we compute the best bilinear fit that represents a shape, observed in a single image from an arbitrary viewpoint, as a combination of the elements of a basis.
Manmohan Krishna Chandraker, David J. Kriegman
CVPR2
2008 PSF estimation using sharp edge prediction
abstract
Image blur is caused by a number of factors such as motion, defocus, capturing light over the non-zero area of the aperture and pixel, the presence of anti-aliasing filters on a camera sensor, and limited sensor resolution. We present an algorithm that estimates non-parametric, spatially-varying blur functions (i.e., point-spread functions or PSFs) at subpixel resolution from a single image. Our method handles blur due to defocus, slight camera motion, and inherent aspects of the imaging system. Our algorithm can be used to measure blur due to limited sensor resolution by estimating a sub-pixel, super-resolved PSF even for in-focus images. It operates by predicting a ldquosharprdquo version of a blurry input image and uses the two images to solve for a PSF. We handle the cases where the scene content is unknown and also where a known printed calibration target is placed in the scene. Our method is completely automatic, fast, and produces accurate results.
Neel Joshi, Richard Szeliski, David J. Kriegman
CVPR3
2008 Practical Global Optimization for Multiview Geometry
Fredrik Kahl, Sameer Agarwal 0001, Manmohan Krishna Chandraker, David J. Kriegman, Serge J. Belongie
Int. J. Comput. Vis.4
2008 Color Subspaces as Photometric Invariants
Todd E. Zickler, Satya P. Mallick, David J. Kriegman, Peter N. Belhumeur
Int. J. Comput. Vis.3
2008 Introduction of New Editors
abstract
O support the continuing increase in submission arising from the popularity of TPAMI with authors, we are pleased to announce that Professor Zoubin Ghahramani will be joining David Fleet as an Associate Editor-in-Chief (AEIC) of TPAMI. He will help to maintain the high quality that readers expect and the timeliness and thoroughness of reviewing that authors’ demand. The AEIC works hand-in-hand with the Editor-in-Chief in all aspects of TPAMI’s editorial review process, including establishing policies, selecting special issues, selecting editors, helping the fl ow of papers through the review process, handling appeals, etc. Dr. Ghahramani is a leader in the fi eld of machine learning, an area of increasing importance to TPAMI, for both the quality of his research and his service to the community. We are also happy to announce that the TPAMI editorial board is expanding with the addition of four new Associate Editors, Dr. Sing Bing Kang, Professor Kevin Murphy. Dr. Salil Prabhakar, and Professor Dale Schuurmans. Dr. Kang will handle papers on image-based rendering, vision for graphics, image and video processing/enhancement, and stereopsis and structure from motion. Professor Murphy will oversee papers about graphical models, structure learning, causal inference, unsupervised learning, and Bayesian methods. Dr. Prabhakar will oversee the review process of papers in all aspects of biometric systems and theoretical and empirical evaluation of computer vision algorithms. Professor Schuurmans’ expertise includes machine learning techniques (support vector machines, kernel methods, clustering, dimensionality reduction), graphical models, optimization, and Monte Carlo methods, and he has begun to handle papers in these areas. Their brief biographies are below. Welcome to TPAMI’s editorial board!
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Editorial-State of the Transactions
David J. Kriegman, David J. Fleet, Zoubin Ghahramani
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet, Zoubin Ghahramani
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet, Zoubin Ghahramani
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet, Zoubin Ghahramani
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Resolving the Generalized Bas-Relief Ambiguity by Entropy Minimization
abstract
It is well known in the photometric stereo literature that uncalibrated photometric stereo, where light source strength and direction are unknown, can recover the surface geometry of a Lambertian object up to a 3-parameter linear transform known as the generalized bas relief (GBR) ambiguity. Many techniques have been proposed for resolving the GBR ambiguity, typically by exploiting prior knowledge of the light sources, the object geometry, or non-Lambertian effects such as specularities. A less celebrated consequence of the GBR transformation is that the albedo at each surface point is transformed along with the geometry. Thus, it should be possible to resolve the GBR ambiguity by exploiting priors on the albedo distribution. To the best of our knowledge, the only time the albedo distribution has been used to resolve the GBR is in the case of uniform albedo. We propose a new prior on the albedo distribution : that the entropy of the distribution should be low. This prior is justified by the fact that many objects in the real-world are composed of a small finite set of albedo values.
Neil Gordon Alldrin, Satya P. Mallick, David J. Kriegman
CVPR3
2007 ShadowCuts: Photometric Stereo with Shadows
abstract
We present an algorithm for performing Lambertian photometric stereo in the presence of shadows. The algorithm has three novel features. First, a fast graph cuts based method is used to estimate per pixel light source visibility. Second, it allows images to be acquired with multiple illuminants, and there can be fewer images than light sources. This leads to better surface coverage and improves the reconstruction accuracy by enhancing the signal to noise ratio and the condition number of the light source matrix. The ability to use fewer images than light sources means that the imaging effort grows sublinearly with the number of light sources. Finally, the recovered shadow maps are combined with shading information to perform constrained surface normal integration. This reduces the low frequency bias inherent to the normal integration process and ensures that the recovered surface is consistent with the shadowing configuration The algorithm works with as few as four light sources and four images. We report results for light source visibility detection and high quality surface reconstructions for synthetic and real datasets.
Manmohan Krishna Chandraker, Sameer Agarwal 0001, David J. Kriegman
CVPR3
2007 Autocalibration via Rank-Constrained Estimation of the Absolute Quadric
abstract
We present an autocalibration algorithm for upgrading a projective reconstruction to a metric reconstruction by estimating the absolute dual quadric. The algorithm enforces the rank degeneracy and the positive semidefiniteness of the dual quadric as part of the estimation procedure, rather than as a post-processing step. Furthermore, the method allows the user, if he or she so desires, to enforce conditions on the plane at infinity so that the reconstruction satisfies the chirality constraints. The algorithm works by constructing low degree polynomial optimization problems, which are solved to their global optimum using a series of convex linear matrix inequality relaxations. The algorithm is fast, stable, robust and has time complexity independent of the number of views. We show extensive results on synthetic as well as real datasets to validate our algorithm.
Manmohan Krishna Chandraker, Sameer Agarwal 0001, Fredrik Kahl, David Nistér, David J. Kriegman
CVPR5
2007 Leveraging temporal, contextual and ordering constraints for recognizing complex activities in video
abstract
We present a scalable approach to recognizing and describing complex activities in video sequences. We are interested in long-term, sequential activities that may have several parallel streams of action. Our approach integrates temporal, contextual and ordering constraints with output from low-level visual detectors to recognize complex, long-term activities. We argue that a hierarchical, object-oriented design lends our solution to be scalable in that higher-level reasoning components are independent from the particular low-level detector implementation and that recognition of additional activities and actions can easily be added. Three major components to realize this design are: a dynamic Bayesian network structure for representing activities comprised of partially ordered sub-actions, an object-oriented action hierarchy for building arbitrarily complex action detectors and an approximate Viterbi-like algorithm for inferring the most likely observed sequence of actions. Additionally, this study proposes the Erlang distribution as a comprehensive model of idle time between actions and frequency of observing new actions. We show results for our approach on real video sequences containing complex activities.
Benjamin Laxton, Jongwoo Lim, David J. Kriegman
CVPR3
2007 Isotropy, Reciprocity and the Generalized Bas-Relief Ambiguity
abstract
A set of images of a Lambertian surface under varying lighting directions defines its shape up to a three-parameter generalized bas-relief (GBR) ambiguity. In this paper, we examine this ambiguity in the context of surfaces having an additive non-Lambertian reflectance component, and we show that the GBR ambiguity is resolved by any non-Lambertian reflectance function that is isotropic and spatially invariant. The key observation is that each point on a curved surface under directional illumination is a member of a family of points that are in isotropic or reciprocal configurations. We show that the GBR can be resolved in closed form by identifying members of these families in two or more images. Based on this idea, we present an algorithm for recovering full Euclidean geometry from a set of uncalibrated photometric stereo images, and we evaluate it empirically on a number of examples.
Satya P. Mallick, Long Quan, David J. Kriegman, Todd E. Zickler
CVPR4
2007 Toward Reconstructing Surfaces With Arbitrary Isotropic Reflectance : A Stratified Photometric Stereo Approach
abstract
We consider the problem of reconstructing the shape of a surface with an arbitrary, spatially varying isotropic bidirectional reflectance distribution function (BRDF), and introduce a novel, stratified photometric stereo method. By using a particular configuration of lights, it is possible to use symmetry in the image measurements resulting from BRDF isotropy to estimate at each point a plane containing the surface normal. For differentiable surfaces, this allows us to recover the isocontours of the depth map, but not the actual depth associated with each contour. The isocontour structure provides topological information about the surface (critical points, Reeb graph, etc.). By using additional cues in the image data or imposing additional constraints on the surface (e.g., shadows, specular highlights, Helmholtz reciprocity, uniform BRDF), the unknown height of each isocontour can be estimated and the metric structure is resolved. We validate this technique on real and synthetic data by successfully recovering the isocontours of the depth map from images.
Neil Gordon Alldrin, David J. Kriegman
ICCV2
2007 Globally Optimal Affine and Metric Upgrades in Stratified Autocalibration
abstract
We present a practical, stratified autocalibration algorithm with theoretical guarantees of global optimality. Given a projective reconstruction, the first stage of the algorithm upgrades it to affine by estimating the position of the plane at infinity. The plane at infinity is computed by globally minimizing a least squares formulation of the modulus constraints. In the second stage, the algorithm upgrades this affine reconstruction to a metric one by globally minimizing the infinite homography relation to compute the dual image of the absolute conic (DIAC). The positive semidefiniteness of the DIAC is explicitly enforced as part of the optimization process, rather than as a post-processing step. For each stage, we construct and minimize tight convex relaxations of the highly non-convex objective functions in a branch and bound optimization framework. We exploit the problem structure to restrict the search space for the DIAC and the plane at infinity to a small, fixed number of branching dimensions, independent of the number of views. Experimental evidence of the accuracy, speed and scalability of our algorithm is presented on synthetic and real data. MATLAB code for the implementation is made available to the community.
Manmohan Krishna Chandraker, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
ICCV3
2007 Synthetic Aperture Tracking: Tracking through Occlusions
abstract
Occlusion is a significant challenge for many tracking algorithms. Most current methods can track through transient occlusion, but cannot handle significant extended occlusion when the object's trajectory may change significantly. We present a method to track a 3D object through significant occlusion using multiple nearby cameras (e.g., a camera array). When an occluder and object are at different depths, different parts of the object are visible or occluded in each view due to parallax. By aggregating across these views, the method can track even when any individual camera observes very little of the target object. Implementation- wise, the methods are straightforward and build upon established single-camera algorithms. They do not require explicit modeling or reconstruction of the scene and enable tracking in complex, dynamic scenes with moving cameras. Analysis of accuracy and robustness shows that these methods are successful when upwards of '70% of the object is occluded in every camera view. To the best of our knowledge, this system is the first capable of tracking in the presence of such significant occlusion.
Neel Joshi, Shai Avidan, Wojciech Matusik, David J. Kriegman
ICCV4
2007 Shape from Varying Illumination and Viewpoint
abstract
We address the problem of reconstructing the 3-D shape of a Lambertian surface from multiple images acquired as an object rotates under distant and possibly varying illumination. Using camera projection matrices estimated from point correspondences across views, the algorithm computes a dense correspondence map by minimizing a multi-ocular photometric constraint. Once correspondence across views is established, photometric stereo is applied to estimate a surface normal field and 3-D surface. Conceptually, the algorithm merges multi-view stereo and photometric stereo and uses aspects of both methods to recover shape. The method is straightforward to implement and relies on established principles from the two stereo methods. We empirically validate the method on images of a number of objects and show that it outperforms previous methods.
Neel Joshi, David J. Kriegman
ICCV2
2007 Editorial-State of the Transactions
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Vision in the Small: Reconstructing the Structure of Protein Macromolecules from Cryo-Electron Micrographs
abstract
Single particle reconstruction using Cryo-Electron Microscopy (cryo-EM) is an emerging technique in structural biology for estimating the 3-D structure (density) of protein macromolecules. Unlike tomography where a large number of images of a specimen can be acquired, the number of images of an individual particle is limited because of radiation damage. Instead, the specimen consists of identical copies of the same protein macro-molecule embedded in vitreous ice at random and unknown 3-D orientations. Because the images are extremely noisy, thousands to hundreds-of-thousands of projections are needed to achieve the desired resolution of 5A. Along with differences of the imaging modality compared to photographs, single particle reconstruction provides a unique set of challenges to existing computer vision algorithms. Here, we introduce the challenge and opportunity of reconstruction from transmission electron micrographs, and briefly describe our contributions in areas of particle detection, contrast transfer function (CTF) estimation, and initial 3-D model construction. Reconstructing the Structure of Protein Macromolecules One of the most exciting challenges for biology today is understanding the molecular machinery of the cell as a working, dynamic system. Critical to this understanding is determining the 3-D structure of protein macromolecules, a task that is often accomplished using x-ray crystallography. The technique of cryo electron microscopy (cryo-EM) has a unique role to play in addressing this challenge as it can provide structural information of large macromolecular complexes in a variety of conformational and compositional states while preserved under close to physiological conditions. Traditionally the methods for cryo-EM have been time consuming and labor intensive, involving data acquisition, analysis and averaging of thousands to hundreds of thousands of images (views) of the individual macro-molecular complexes. Thus, over the last few years there has been considerable interest and substantial effort devoted to developing automated methods to improve the accuracy, robustness, ease of use, and throughput of cryo-EM [1, 2, 3, 12, 15, 17], and this abstract considers three aspects originally presented in [9, 10, 11]. 1
Satya P. Mallick, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
BMVC3
2006 A Planar Light Probe
abstract
We develop a novel technique for measuring lighting that exploits the interaction of light with a set of custom BRDFs. This enables the construction of a planar light probe with certain advantages over existing methods for measuring lighting. To facilitate the construction of our light probe, we derive a new class of bi-directional reflectance functions based on the interaction of light through two planar surfaces separated by a transparent medium. Under certain assumptions and proper selection of the two surfaces, we show how to recover Fourier series coefficients of the incident lighting parameterized over the plane. The results are experimentally validated by imaging a sheet of glass with spatially varying patterns printed on either side.
Neil Gordon Alldrin, David J. Kriegman
CVPR (2)2
2006 Structure and View Estimation for Tomographic Reconstruction: A Bayesian Approach
abstract
This paper addresses the problem of reconstructing the density of a scene from multiple projection images produced by modalities such as x-ray, electron microscopy, etc. where an image value is related to the integral of the scene density along a 3D line segment between a radiation source and a point on the image plane. While computed tomography (CT) addresses this problem when the absolute orientation of the image plane and radiation source directions are known, this paper addresses the problem when the orientations are unknown - it is akin to the structure-from-motion (SFM) problem when the extrinsic camera parameters are unknown. We study the problem within the context of reconstructing the density of protein macro-molecules in Cryogenic Electron Microscopy (cryo-EM), where images are very noisy and existing techniques use several thousands of images. In a non-degenerate configuration, the viewing planes corresponding to two projections, intersect in a line in 3D. Using the geometry of the imaging setup, it is possible to determine the projections of this 3D line on the two image planes. In turn, the problem can be formulated as a type of orthographic structure from motion from line correspondences where the line correspondences between two views are unreliable due to image noise. We formulate the task as the problem of denoising a correspondence matrix and present a Bayesian solution to it. Subsequently, the absolute orientation of each projection is determined followed by density reconstruction. We show results on cryo-EM images of proteins and compare our results to that of Electron Micrograph Analysis (EMAN) - a widely used reconstruction tool in cryo-EM.
Satya P. Mallick, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie, Bridget Carragher, Clinton S. Potter
CVPR (2)3
2006 Color Subspaces as Photometric Invariants
abstract
Complex reflectance phenomena such as specular reflections confound many vision problems since they produce image ‘features’ that do not correspond directly to intrinsic surface properties such as shape and spectral reflectance. A common approach to mitigate these effects is to explore functions of an image that are invariant to these photometric events. In this paper we describe two such invariants" one invariant to specular reflections, and the other invariant to both specular reflections and diffuse shading" that result from exploiting color information in images of dichromatic surfaces. These invariants are derived from subspaces of RGB color space, and they enable the application of Lambertian-based vision techniques to a broad class of specular, non-Lambertian scenes. Using implementations of recent algorithms taken from the literature, we demonstrate the practical utility of these invariants for a wide variety of applications, including stereo, shape from shading, material-based segmentation, and motion estimation.
Todd E. Zickler, Satya P. Mallick, David J. Kriegman, Peter N. Belhumeur
CVPR (2)3
2006 Practical Global Optimization for Multiview Geometry
Sameer Agarwal 0001, Manmohan Krishna Chandraker, Fredrik Kahl, David J. Kriegman, Serge J. Belongie
ECCV (1)4
2006 Integrating Surface Normal Vectors Using Fast Marching Method
Jeffrey Ho, Jongwoo Lim, Ming-Hsuan Yang 0001, David J. Kriegman
ECCV (3)4
2006 Specularity Removal in Images and Videos: A PDE Approach
Satya P. Mallick, Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
ECCV (1)4
2006 Reconstruction of Volumetric Surface Textures for Real-Time Rendering
Sebastian Magda, David J. Kriegman
Rendering Techniques2
2006 Robust Structure and Motion from Outlines of Smooth Curved Surfaces
abstract
This paper addresses the problem of estimating the motion of a camera as it observes the outline (or apparent contour) of a solid bounded by a smooth surface in successive image frames. In this context, the surface points that project onto the outline of an object depend on the viewpoint and the only true correspondences between two outlines of the same object are the projections of frontier points where the viewing rays intersect in the tangent plane of the surface. In turn, the epipolar geometry is easily estimated once these correspondences have been identified. Given the apparent contours detected in an image sequence, a robust procedure based on RANSAC and a voting strategy is proposed to simultaneously estimate the camera configurations and a consistent set of frontier point projections by enforcing the redundancy of multiview epipolar geometry. The proposed approach is, in principle, applicable to orthographic, weak-perspective, and affine projection models. Experiments with nine real image sequences are presented for the orthographic projection case, including a quantitative comparison with the ground-truth data for the six data sets for which the latter information is available. Sample visual hulls have been computed from all image sequences for qualitative evaluation.
Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.4
2006 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Editorial-State of the Transactions
abstract
IT was another good year for the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)—a year with more and better papers and a more competitive publication environment. It was also a year of some challenges. There are two major sets of metrics that are commonly used to evaluate the performance of a journal: citation statistics, which assess the significance of the papers published, and editorial statistics, which characterize the flow of papers through the editorial process. Certainly, the former is connected to the latter. A vibrant and healthy journal starts with authors submitting their most important papers. Between conferences and journals, authors have a great deal of choice about where to publish. They want a prestigious journal where their papers will receive a fair review in a timely manner. Starting with citation statistics, the most widely used measure is the impact factor which is the average number of times papers published in the two previous years referenced by journals in the given year. Based on the publications tracked by Thomson ISI, TPAMI papers were cited 14,708 times in 2006 and the impact factor in 2006 was 4.3, up from 3.8 in 2005. TPAMI is the IEEE’s most cited publication, the second most cited journal in all of electrical engineering, and the fifth most cited journal in all of computer science. In 2006, TPAMI received 912 submissions, up from 749 the year before. The acceptance rate is about 25 percent. The peer review cycle has been improving and the average time from submission to first decision is three months, and the average time from submission to acceptance is nine months. It is worth noting how favorably the submission to first decision time compares to that of conferences. The reason for the length of time between submission and acceptance is that many papers undergo a major and a minor revision and, consequently, require additional reviewing. There is an asymmetry between accepted and rejected papers since those that are rejected (the majority) leave the review process more quickly. TPAMI posts accepted papers digitally in the CS Digital Library and IEEE Xplore in advance of the printed version, providing early access to readers. Once papers are posted online, they are considered published and with this new posting online upon acceptance, submission to publication has been dramatically reduced. In 2007, TPAMI published a very successful special issue on Progress and Directions in Biometrics, edited by Josef Kittler, Davide Maltoni, Lawrence O’Gorman, Salil Prabhakar, and Tieniu Tan. The issue received 85 submissions, of which 19 were accepted. Handling this many papers was a daunting task for the guest editors. Looking toward 2008, there are two special issues in the works: First, we are pleased to announce that TPAMI will be having a special section devoted to the award-winning papers from the 2007 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). There have always been very strong ties between the major conferences in our field and TPAMI, and it was high time for the best papers at the conferences to be recognized by TPAMI. Please note that these papers are still fully reviewed by TPAMI. Second, papers are under review for a special issue on Real-World Image Annotation and Retrieval. This issue points to the large scale applications that are becoming possible as the methods of computer vision and machine learning become more effective and robust. The year (2007) started with a conversion of the Webbased software for tracking manuscripts through peer review to a new system (Manuscript Central v3.0). The new version has a backend which is much more secure and a user interface which has been updated to provide more information to authors, reviewers, and editors and a more contemporary look and feel. While security, unless violated, is invisible to the user, the new interface is inescapable—the response has been mixed. While many bugs were fixed and most of the more serious issues have been addressed, there is still much to be done. We apologize to all who were inconvenienced. We would like to thank everyone who helps to make TPAMI a great journal, starting with the authors who submit their best works to TPAMI. The largest group is composed of the reviewers who collectively provide about 2,500 each year. Reviewing can be very rewarding, but it is also a heavy commitment and we appreciate the efforts and time these reviewers put into TPAMI and the community for which they are volunteering. For seasoned researchers, reviewing is a chance to see research results before they hit the press—a sneak preview—and to provide a valuable service to the community. For a novice reviewer, the first few dozen papers will highlight to the reviewer what makes a good versus a poor submission. Papers that are published are among 25 percent of submissions passing through the sieve and can be viewed as “positive training examples.” The other 75 percent of submissions are “negative training examples,” and these are only accessible to reviewers. The implication from a pattern classification perspective should be apparent to readers of this editorial. We would also like to thank the Associate Editors (AEs) who make a long term (four year) and ongoing commitment to the journal. The AEs are among the top researchers in the field, and their time is very valuable. They must make personal decisions about how much time they devote to IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 30, NO. 2, FEBRUARY 2008 193
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Introduction of New Associate Editors
abstract
The EiC and Associate EiC express our gratitude to David Forsyth, Brendan Frey, Venu Govindaraju, and Cordelia Schmid who are retiring as associate editors of TPAMI. While we will miss their dedication to the transactions, we hope that they will be enjoying a bit more free time. We are also pleased to announce that Professor Daniel Lopresti, Professor B.S. Manjunath, Professor Marc Pollefeys, and Professor Ramin Zabih have joined the editorial board. Professor Lopresti will oversee papers in document and handwriting analysis, biometrics, approximate string matching algorithms, and performance evaluation. Professor Manjunath will be considering papers in feature extraction, segmentation, image/video retrieval, and image registration. Professor Pollefeys will be responsible for submissions in structure from motion, stereo, multiple view geometry and camera calibration, 3D and appearance modeling, shape-from-X techniques, and novel sensors. Professor Zabih will handle papers on stereo and medical imaging as well as energy minimization and graph algorithms. We look forward to working with them. Their brief biographies appear herein.
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Face Recognition Using 3-D Models: Pose and Illumination
abstract
Unconstrained illumination and pose variation lead to significant variation in the photographs of faces and constitute a major hurdle preventing the widespread use of face recognition systems. The challenge is to generalize from a limited number of images of an individual to a broad range of conditions. Recently, advances in modeling the effects of illumination and pose have been accomplished using three-dimensional (3-D) shape information coupled with reflectance models. Notable developments in understanding the effects of illumination include the nonexistence of illumination invariants, a characterization of the set of images of objects in fixed pose under variable illumination (the illumination cone), and the introduction of spherical harmonics and low-dimensional linear subspaces for modeling illumination. To generalize to novel conditions, either multiple images must be available to reconstruct 3-D shape or, if only a single image is accessible, prior information about the 3-D shape and appearance of faces in general must be used. The 3-D Morphable Model was introduced as a generative model to predict the appearances of an individual while using a statistical prior on shape and texture allowing its parameters to be estimated from single image. Based on these new understandings, face recognition algorithms have been developed to address the joint challenges of pose and lighting. In this paper, we review these developments and provide a brief survey of the resulting face recognition algorithms and their performance
Sami Romdhani, Jeffrey Ho, Thomas Vetter, David J. Kriegman
Proc. IEEE4
2005 Beyond Pairwise Clustering
abstract
We consider the problem of clustering in domains where the affinity relations are not dyadic (pairwise), but rather triadic, tetradic or higher. The problem is an instance of the hypergraph partitioning problem. We propose a two-step algorithm for solving this problem. In the first step we use a novel scheme to approximate the hypergraph using a weighted graph. In the second step a spectral partitioning algorithm is used to partition the vertices of this graph. The algorithm is capable of handling hyperedges of all orders including order two, thus incorporating information of all orders simultaneously. We present a theoretical analysis that relates our algorithm to an existing hypergraph partitioning algorithm and explain the reasons for its superior performance. We report the performance of our algorithm on a variety of computer vision problems and compare it to several existing hypergraph partitioning algorithms.
Sameer Agarwal 0001, Jongwoo Lim, Lihi Zelnik-Manor, Pietro Perona, David J. Kriegman, Serge J. Belongie
CVPR (2)5
2005 Reflections on the Generalized Bas-Relief Ambiguity
abstract
Prior work has argued that when a Lambertian surface in fixed pose is observed in multiple images under varying distant illumination, there is an equivalence class of surfaces given by the generalized bas-relief (GBR) ambiguity that could have produced these images. In contrast, this paper shows that for general nonconvex surfaces, interreflections completely resolve the GBR ambiguity. In turn, the full Euclidean geometry can be recovered from uncalibrated photometric stereo for which the light source directions and strengths are unknown. Further, we show that surfaces with a translational symmetry do not lend enough constraints to be disambiguated by inter reflections.
Manmohan Krishna Chandraker, Fredrik Kahl, David J. Kriegman
CVPR (1)3
2005 Online Learning of Probabilistic Appearance Manifolds for Video-Based Recognition and Tracking
abstract
This paper presents an online learning algorithm to construct from video sequences an image-based representation that is useful for recognition and tracking. For a class of objects (e.g., human faces), a generic representation of the appearances of the class is learned off-line. From video of an instance of this class (e.g., a particular person), an appearance model is incrementally learned on-line using the prior generic model and successive frames from the video. More specifically, both the generic and individual appearances are represented as an appearance manifold that is approximated by a collection of sub-manifolds (named pose manifolds) and the connectivity between them. In turn, each sub-manifold is approximated by a low-dimensional linear sub-space while the connectivity is modeled by transition probabilities between pairs of sub-manifolds. We demonstrate that our online learning algorithm constructs an effective representation for face tracking, and its use in video-based face recognition compares favorably to the representation constructed with a batch technique.
Kuang-chih Lee, David J. Kriegman
CVPR (1)2
2005 Beyond Lambert: Reconstructing Specular Surfaces Using Color
abstract
We present a photometric stereo method for non-diffuse materials that does not require an explicit reflectance model or reference object. By computing a data-dependent rotation of RGB color space, we show that the specular reflection effects can be separated from the much simpler, diffuse (approximately Lambertian) reflection effects for surfaces that can be modeled with dichromatic reflectance. Images in this transformed color space are used to obtain photometric reconstructions that are independent of the specular reflectance. In contrast to other methods for highlight removal based on dichromatic color separation (e.g., color histogram analysis and/or polarization), we do not explicitly recover the specular and diffuse components of an image. Instead, we simply find a transformation of color space that yields more direct access to shape information. The method is purely local and is able to handle surfaces with arbitrary texture.
Satya P. Mallick, Todd E. Zickler, David J. Kriegman, Peter N. Belhumeur
CVPR (2)3
2005 Passive Photometric Stereo from Motion
abstract
We introduce an iterative algorithm for shape reconstruction from multiple images of a moving (Lambertian) object illuminated by distant (and possibly time varying) lighting. Starting with an initial piecewise linear surface, the algorithm iteratively estimates a new surface based on the previous surface estimate and the photometric information available from the input image sequence. During each iteration, standard photometric stereo techniques are applied to estimate the surface normals up to an unknown generalized bas-relief transform, and a new surface is computed by integrating the estimated normals. The algorithm essentially consists of a sequence of matrix factorizations (of intensity values) followed by minimization using gradient descent (integration of the normals). Conceptually, the algorithm admits a clear geometric interpretation, which is used to provide a qualitative analysis of the algorithm's convergence. Implementation-wise, it is straightforward being based on several established photometric stereo and structure from motion algorithms. We demonstrate experimentally the effectiveness of our algorithm using several videos of hand-held objects moving in front of a fixed light and camera.
Jongwoo Lim, Jeffrey Ho, Ming-Hsuan Yang 0001, David J. Kriegman
ICCV4
2005 Visual tracking and recognition using probabilistic appearance manifolds
Kuang-chih Lee, Jeffrey Ho, Ming-Hsuan Yang 0001, David J. Kriegman
Comput. Vis. Image Underst.4
2005 Introduction of New Associate Editors
abstract
The EiC and Associate EiC express our gratitude to David Forsyth, Brendan Frey, Venu Govindaraju, and Cordelia Schmid who are retiring as associate editors of TPAMI. While we will miss their dedication to the transactions, we hope that they will be enjoying a bit more free time. We are also pleased to announce that Professor Daniel Lopresti, Professor B.S. Manjunath, Professor Marc Pollefeys, and Professor Ramin Zabih have joined the editorial board. Professor Lopresti will oversee papers in document and handwriting analysis, biometrics, approximate string matching algorithms, and performance evaluation. Professor Manjunath will be considering papers in feature extraction, segmentation, image/video retrieval, and image registration. Professor Pollefeys will be responsible for submissions in structure from motion, stereo, multiple view geometry and camera calibration, 3D and appearance modeling, shape-from-X techniques, and novel sensors. Professor Zabih will handle papers on stereo and medical imaging as well as energy minimization and graph algorithms. We look forward to working with them. Their brief biographies appear herein.
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Introduction of New Associate Editors
David J. Kriegman, David J. Fleet
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Acquiring Linear Subspaces for Face Recognition under Variable Lighting
abstract
Previous work has demonstrated that the image variation of many objects (human faces in particular) under variable lighting can be effectively modeled by low-dimensional linear spaces, even when there are multiple light sources and shadowing. Basis images spanning this space are usually obtained in one of three ways: A large set of images of the object under different lighting conditions is acquired, and principal component analysis (PCA) is used to estimate a subspace. Alternatively, synthetic images are rendered from a 3D model (perhaps reconstructed from images) under point sources and, again, PCA is used to estimate a subspace. Finally, images rendered from a 3D model under diffuse lighting based on spherical harmonics are directly used as basis images. In this paper, we show how to arrange physical lighting so that the acquired images of each object can be directly used as the basis vectors of a low-dimensional linear space and that this subspace is close to those acquired by the other methods. More specifically, there exist configurations of k point light source directions, with k typically ranging from 5 to 9, such that, by taking k images of an object under these single sources, the resulting subspace is an effective representation for recognition under a wide range of lighting conditions. Since the subspace is generated directly from real images, potentially complex and/or brittle intermediate steps such as 3D reconstruction can be completely avoided; nor is it necessary to acquire large numbers of training images or to physically construct complex diffuse (harmonic) light fields. We validate the use of subspaces constructed in this fashion within the context of face recognition.
Kuang-chih Lee, Jeffrey Ho, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Visual Tracking Using Learned Linear Subspaces
Jeffrey Ho, Kuang-chih Lee, Ming-Hsuan Yang 0001, David J. Kriegman
CVPR (1)4
2004 On Refractive Optical Flow
Sameer Agarwal 0001, Satya P. Mallick, David J. Kriegman, Serge J. Belongie
ECCV (2)3
2004 Structure and Motion from Images of Smooth Textureless Objects
Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman
ECCV (2)4
2004 Image Clustering with Metric, Local Linear Structure, and Affine Symmetry
Jongwoo Lim, Jeffrey Ho, Ming-Hsuan Yang 0001, Kuang-chih Lee, David J. Kriegman
ECCV (1)5
2004 Curve and Surface Duals and the Recognition of Curved 3D Objects from their Silhouettes
Amit Sethi, David Renaudie, David J. Kriegman, Jean Ponce
Int. J. Comput. Vis.3
2004 Introduction of New Associate Editor
abstract
The EiC and Associate EiC express our gratitude to David Forsyth, Brendan Frey, Venu Govindaraju, and Cordelia Schmid who are retiring as associate editors of TPAMI. While we will miss their dedication to the transactions, we hope that they will be enjoying a bit more free time. We are also pleased to announce that Professor Daniel Lopresti, Professor B.S. Manjunath, Professor Marc Pollefeys, and Professor Ramin Zabih have joined the editorial board. Professor Lopresti will oversee papers in document and handwriting analysis, biometrics, approximate string matching algorithms, and performance evaluation. Professor Manjunath will be considering papers in feature extraction, segmentation, image/video retrieval, and image registration. Professor Pollefeys will be responsible for submissions in structure from motion, stereo, multiple view geometry and camera calibration, 3D and appearance modeling, shape-from-X techniques, and novel sensors. Professor Zabih will handle papers on stereo and medical imaging as well as energy minimization and graph algorithms. We look forward to working with them. Their brief biographies appear herein.
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Introduction of New Associate Editors
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 In Memoriam, Azriel Rosenfeld(1931-2004)
abstract
A brief biography of Azriel Rosenfeld (1931-2004) is given highlighting his professional achievements. Dr. Rosenfeld was widely regarded as the leading researcher in the world in the field of computer image analysis.
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Clustering Appearances of Objects Under Varying Illumination Conditions
abstract
We introduce two appearance-based methods for clustering a set of images of 3D (three-dimensional) objects, acquired under varying illumination conditions, into disjoint subsets corresponding to individual objects. The first algorithm is based on the concept of illumination cones. According to the theory, the clustering problem is equivalent to finding convex polyhedral cones in the high-dimensional image space. To efficiently determine the conic structures hidden in the image data, we introduce the concept of conic affinity, which measures the likelihood of a pair of images belonging to the same underlying polyhedral cone. For the second method, we introduce another affinity measure based on image gradient comparisons. The algorithm operates directly on the image gradients by comparing the magnitudes and orientations of the image gradient at each pixel. Both methods have clear geometric motivations, and they operate directly on the images without the need for feature extraction or computation of pixel statistics. We demonstrate experimentally that both algorithms are surprisingly effective in clustering images acquired under varying illumination conditions with two large, well-known image data sets.
Jeffrey Ho, Ming-Hsuan Yang 0001, Jongwoo Lim, Kuang-chih Lee, David J. Kriegman
CVPR (1)5
2003 Video-Based Face Recognition Using Probabilistic Appearance Manifolds
abstract
This paper presents a method to model and recognize human faces in video sequences. Each registered person is represented by a low-dimensional appearance manifold in the ambient image space, the complex nonlinear appearance manifold expressed as a collection of subsets (named pose manifolds), and the connectivity among them. Each pose manifold is approximated by an affine plane. To construct this representation, exemplars are sampled from videos, and these exemplars are clustered with a K-means algorithm; each cluster is represented as a plane computed through principal component analysis (PCA). The connectivity between the pose manifolds encodes the transition probability between images in each of the pose manifold and is learned from a training video sequences. A maximum a posteriori formulation is presented for face recognition in test video sequences by integrating the likelihood that the input image comes from a particular pose manifold and the transition probability to this pose manifold from the previous frame. To recognize faces with partial occlusion, we introduce a weight mask into the process. Extensive experiments demonstrate that the proposed algorithm outperforms existing frame-based face recognition methods with temporal voting schemes.
Kuang-chih Lee, Jeffrey Ho, Ming-Hsuan Yang 0001, David J. Kriegman
CVPR (1)4
2003 Toward a Stratification of Helmholtz Stereopsis
abstract
Helmholtz stereopsis has been previously introduced as a surface reconstruction technique that does not assume a model of surface reflectance. This technique relies on the use of multiple cameras and light sources, and it has been shown to be effective when the camera and source positions are known. Here, we take a stratified look at uncalibrated Helmholtz stereopsis. We derive a photometric matching constraint that can be used to establish correspondence without any knowledge of the cameras and sources (except that they are co-located), and we determine conditions under which we can obtain affine and metric reconstructions. An implementation and experimental results are presented.
Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
CVPR (1)3
2003 Binocular Helmholtz Stereopsis
abstract
Helmholtz stereopsis has been introduced recently as a surface reconstruction technique that does not assume a model of surface reflectance. In the reported formulation, correspondence was established using a rank constraint, necessitating at least three viewpoints and three pairs of images. Here, it is revealed that the fundamental Helmholtz stereopsis constraint defines a nonlinear partial differential equation, which can be solved using only two images. It is shown that, unlike conventional stereo, binocular Helmholtz stereopsis is able to establish correspondence (and thereby recover surface depth) for objects having an arbitrary and unknown BRDF and in textureless regions (i.e., regions of constant or slowly varying BRDF). An implementation and experimental results validate the method for specular surfaces with and without texture.
Todd E. Zickler, Jeffrey Ho, David J. Kriegman, Jean Ponce, Peter N. Belhumeur
ICCV3
2003 Fast texture synthesis on arbitrary meshes
abstract
No abstract available.
Sebastian Magda, David J. Kriegman
SIGGRAPH2
2003 Special issue on face recognition
Aleix Martinez, Ming-Hsuan Yang 0001, David J. Kriegman
Comput. Vis. Image Underst.3
2003 Introduction of New Associate Editors
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Editorial - State of the Transactions
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Introduction of New Associate Editor
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Introduction of New Associate Editor
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Introduction of New Associate Editors
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 On Pencils of Tangent Planes and the Recognition of Smooth 3D Shapes from Silhouettes
Svetlana Lazebnik, Amit Sethi, Cordelia Schmid, David J. Kriegman, Jean Ponce, Martial Hebert
ECCV (3)4
2002 Helmholtz Stereopsis: Exploiting Reciprocity for Surface Reconstruction
Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
ECCV (3)3
2002 Appearance-based Eye Gaze Estimation
abstract
We present a method for estimating eye gaze direction, which represents a departure from conventional eye gaze estimation methods, the majority of which are based on tracking specific optical phenomena like corneal reflection and the Purkinje images. We employ an appearance manifold model, but instead of using a densely sampled spline to perform the nearest manifold point query, we retain the original set of sparse appearance samples and use linear interpolation among a small subset of samples to approximate the nearest manifold point. The advantage of this approach is that since we are only storing a sparse set of samples, each sample can be a high dimensional vector that retains more representational accuracy than short vectors produced with dimensionality reduction methods. The algorithm was tested with a set of eye images labelled with ground truth point-of-regard coordinates. We have found that the algorithm is capable of estimating eye gaze with a mean angular error of 0.38 degrees, which is comparable to that obtained by commercially available eye trackers.
Kar-Han Tan, David J. Kriegman, Narendra Ahuja
WACV2
2002 A Real-Time Approach to the Spotting, Representation, and Recognition of Hand Gestures for Human-Computer Interaction
Yuanxin Zhu, Guangyou Xu, David J. Kriegman
Comput. Vis. Image Underst.3
2002 Helmholtz Stereopsis: Exploiting Reciprocity for Surface Reconstruction
Todd E. Zickler, Peter N. Belhumeur, David J. Kriegman
Int. J. Comput. Vis.3
2002 State of the Transactions
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Introduction of New Associate Editors
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Introduction of New Associate Editor
Rama Chellappa, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Detecting Faces in Images: A Survey
abstract
Images containing faces are essential to intelligent vision-based human-computer interaction, and research efforts in face processing include face recognition, face tracking, pose estimation and expression recognition. However, many reported methods assume that the faces in an image or an image sequence have been identified and localized. To build fully automated systems that analyze the information contained in face images, robust and efficient face detection algorithms are required. Given a single image, the goal of face detection is to identify all image regions which contain a face, regardless of its 3D position, orientation and lighting conditions. Such a problem is challenging because faces are non-rigid and have a high degree of variability in size, shape, color and texture. Numerous techniques have been developed to detect faces in a single image, and the purpose of this paper is to categorize and evaluate these algorithms. We also discuss relevant issues such as data collection, evaluation metrics and benchmarking. After analyzing these algorithms and identifying their limitations, we conclude with several promising directions for future research.
Ming-Hsuan Yang 0001, David J. Kriegman, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Image-based Modeling and Rendering of Surfaces with Arbitrary BRDFs
abstract
A goal of image-based rendering is to synthesize as realistically as possible man made and natural objects. The paper presents a method for image-based modeling and rendering of objects with arbitrary (possibly anisotropic and spatially varying) BRDFs. An object is modeled by sampling the surface's incident light field to reconstruct a non-parametric apparent BRDF at each visible point on the surface, This can be used to render the object from the same viewpoint but under arbitrarily specified illumination. We demonstrate how these object models can be embedded in synthetic scenes and rendered under global illumination which captures the interreflections between real and synthetic objects. We also show how these image-based models can be automatically composited onto video footage with dynamic illumination so that the effects (shadows and shading) of the lighting on the composited object match those of the scene.
Melissa L. Koudelka, Peter N. Belhumeur, Sebastian Magda, David J. Kriegman
CVPR (1)4
2001 Nine Points of Light: Acquiring Subspaces for Face Recognition under Variable Lighting
abstract
Previous work has demonstrated that the image variations of many objects (human faces in particular) under variable lighting can be effectively modeled by low dimensional linear spaces. Basis images spanning this space are usually obtained in one of two ways: A large number of images of the object under different conditions is acquired, and principal component analysis (PCA) is used to estimate a subspace. Alternatively, a 3D model (perhaps reconstructed from images) is used to render virtual images under either point sources from which a subspace is derived using PCA or more recently under diffuse synthetic lighting based on spherical harmonics. In this paper we show that there exists a configuration of nine point light source directions such that by taking nine images of each individual under these single sources, the resulting subspace is effective at recognition under a wide range of lighting conditions. Since the subspace is generated directly from real images, potentially complex intermediate steps such as PCA and 3D reconstruction can be completely avoided; nor is it necessary to acquire large numbers of training images or physically construct complex diffuse (harmonic) light fields. We provide both theoretical and empirical results to explain why these linear spaces should be good for recognition.
Kuang-chih Lee, Jeffrey Ho, David J. Kriegman
CVPR (1)3
2001 Beyond Lambert: Reconstructing Surfaces with Arbitrary BRDFs
abstract
We address an open and hitherto neglected problem in computer vision, how to reconstruct the geometry of objects with arbitrary and possibly anisotropic bidirectional reflectance distribution functions (BRDFs). Present reconstruction techniques, whether stereo vision, structure from motion, laser range finding, etc. make explicit or implicit assumptions about the BRDF. Here, we introduce two methods that were developed by re-examining the underlying image formation process; the methods make no assumptions about the object's shape, the presence or absence of shadowing, or the nature of the BRDF which may vary over the surface. The first method takes advantage of Helmholtz reciprocity, while the second method exploits the fact that the radiance along a ray of light is constant. In particular, the first method uses stereo pairs of images in which point light sources are co-located at the centers of projection of the stereo cameras. The second method is based on double covering a scene's incident light field; the depths of surface points are estimated using a large collection of images in which the viewpoint remains fixed and a point light source illuminates the object. Results from our implementations lend empirical support to both techniques.
Sebastian Magda, David J. Kriegman, Todd E. Zickler, Peter N. Belhumeur
ICCV2
2001 Compressing Large Polygonal Models
abstract
Presents an algorithm that uses partitioning and gluing to compress large triangular meshes which are too complex to fit in main memory. The algorithm is based largely on the existing mesh compression algorithms, most of which require an 'in-core' representation of the input mesh. Our solution is to partition the mesh into smaller submeshes and compress these submeshes separately using existing mesh compression techniques. Since a direct partition of the input mesh is out of question, instead we partition a simplified mesh and use the partition on the simplified model to obtain a partition on the original model. In order to recover the full connectivity, we present a simple scheme for encoding/decoding the resulting boundary structure from the mesh partition. When compressing large models with few singular vertices, a negligible portion of the compressed output is devoted to gluing information. On desktop computers, we have run experiments on models with millions of vertices, which could not be compressed using standard compression software packages, and have observed compression ratios as high as 17 to 1 using our technique.
Jeffrey Ho, Kuang-chih Lee, David J. Kriegman
IEEE Visualization3
2001 Face Detection Using Multimodal Density Models
Ming-Hsuan Yang 0001, David J. Kriegman, Narendra Ahuja
Comput. Vis. Image Underst.2
2001 From Few to Many: Illumination Cone Models for Face Recognition under Variable Lighting and Pose
abstract
We present a generative appearance-based method for recognizing human faces under variation in lighting and viewpoint. Our method exploits the fact that the set of images of an object in fixed pose, but under all possible illumination conditions, is a convex cone in the space of images. Using a small number of training images of each face taken with different lighting directions, the shape and albedo of the face can be reconstructed. In turn, this reconstruction serves as a generative model that can be used to render (or synthesize) images of the face under novel poses and illumination conditions. The pose space is then sampled and, for each pose, the corresponding illumination cone is approximated by a low-dimensional linear subspace whose basis vectors are estimated using the generative model. Our recognition algorithm assigns to a test image the identity of the closest approximated illumination cone. Test results show that the method performs almost without error, except on the most extreme lighting directions.
Athinodoros S. Georghiades, Peter N. Belhumeur, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.3
2000 Duals, Invariants, and the Recognition of Smooth Objects from their Occlucing Contours
David Renaudie, David J. Kriegman, Jean Ponce
ECCV (1)2
2000 From Few to Many: Generative Models for Recognition Under Variable Pose and Illumination
abstract
Image variability due to changes in pose and illumination can seriously impair object recognition. This paper presents appearance-based methods which, unlike previous appearance-based approaches, require only a small set of training images to generate a rich representation that models this variability. Specifically, from as few as three images of an object in fixed pose seen under slightly varying but unknown lighting, a surface and an albedo map are reconstructed. These are then used to generate synthetic images with large variations in pose and illumination and thus build a representation useful for object recognition. Our methods have been tested within the domain of face recognition on a subset of the Yale Face Database B containing 4050 images of 10 faces seen under variable pose and illumination. This database was specifically gathered for testing these generative methods. Their performance is shown to exceed that of popular existing methods.
Athinodoros S. Georghiades, Peter N. Belhumeur, David J. Kriegman
FG3
2000 Face Detection Using Mixtures of Linear Subspaces
abstract
We present two methods using mixtures of linear sub-spaces for face detection in gray level images. One method uses a mixture of factor analyzers to concurrently perform clustering and, within each cluster, perform local dimensionality reduction. The parameters of the mixture model are estimated using an EM algorithm. A face is detected if the probability of an input sample is above a predefined threshold. The other mixture of subspaces method uses Kohonen's self-organizing map for clustering and Fisher linear discriminant to find the optimal projection for pattern classification, and a Gaussian distribution to model the class-conditioned density function of the projected samples for each class. The parameters of the class-conditioned density functions are maximum likelihood estimates and the decision rule is also based on maximum likelihood. A wide range of face images including ones in different poses, with different expressions and under different lighting conditions are used as the training set to capture the variations of human faces. Our methods have been tested on three sets of 225 images which contain 871 faces. Experimental results on the first two datasets show that our methods perform as well as the best methods in the literature, yet have fewer false detects.
Ming-Hsuan Yang 0001, Narendra Ahuja, David J. Kriegman
FG3
2000 Face Recognition Using Kernel Eigenfaces
abstract
Eigenface or principal component analysis (PCA) methods have demonstrated their success in face recognition, detection, and tracking. The representation in PCA is based on the second order statistics of the image set, and does not address higher order statistical dependencies such as the relationships among three or more pixels. Higher order statistics (HOS) have been used as a more informative low dimensional representation than PCA for face and vehicle detection. We investigate a generalization of PCA, kernel principal component analysis (kernel PCA), for learning low dimensional representations in the context of face recognition. In contrast to HOS, kernel PCA computes the higher order statistics without the combinatorial explosion of time and memory complexity. While PCA aims to find a second order correlation of patterns, kernel PCA provides a replacement which takes into account higher order correlations. We compare the recognition results using kernel methods with eigenface methods on two benchmarks. Empirical results show that kernel PCA outperforms the eigenface method in face recognition.
Ming-Hsuan Yang 0001, Narendra Ahuja, David J. Kriegman
ICIP3
2000 Selecting Promising Landmarks
abstract
Many approaches to visual servoing and mobile robot navigation are based on tracking feature points or landmarks in images, but not all features points are equally effective as landmarks. Here we develop methods for selecting within an image those landmarks which are both perceptually salient and visually distinctive, and consequently are readily recognized in a second image acquired from a different viewpoint. Empirically, we characterize the performance of the recognition method and then demonstrate that the selection process does in fact choose the landmarks which are more likely to be recognized.
Marcus Knapek, Ricardo Swain Oropeza, David J. Kriegman
ICRA3
1999 Face Detection Using a Mixture of Factor Analyzers
abstract
We present a probabilistic method to detect human faces using a mixture of factor analyzers. One characteristic of this mixture model is that it concurrently performs clustering and, within each cluster, local dimensionality reduction. A wide range of face images including ones in different poses, with different expressions and under different lighting conditions are used as the training set to capture the variations of human faces. In order to fit the mixture model to the sample face images, the parameters are estimated using an EM algorithm. Experimental results show that faces in different poses, with different facial expressions, and under different lighting conditions are accurately detected by our method.
Ming-Hsuan Yang 0001, Narendra Ahuja, David J. Kriegman
ICIP (3)3
1999 On Manipulating Polygonal Objects with Three 2-DOF Robots in the Plane
abstract
Addresses the problem of grasping and manipulating a polygonal object with three disc-shaped robots in the plane. These robots may be the fingertips of a gripper or mobile platforms. The proposed approach is based on the characterization of the range of possible object motions when two of the effectors are fixed and the third one is allowed to move in the plane with two degrees of freedom. This technique does not assume that contact is maintained during the execution of the grasping/manipulation task, nor does it rely on detailed (and a priori unverifiable) models of friction or contact dynamics, but it allows the construction of manipulation plans guaranteed to succeed under the weaker assumption that jamming does not occur during the task execution. The proposed approach is validated by simulation examples and preliminary experiments with Nomadic Scout robots.
Attawith Sudsang, Jean Ponce, Mark Hyman, David J. Kriegman
ICRA4
1999 The Bas-Relief Ambiguity
Peter N. Belhumeur, David J. Kriegman, Alan L. Yuille
Int. J. Comput. Vis.2
1998 Illumination Cones for Recognition under Variable Lighting: Faces
abstract
Due to illumination variability, the same object can appear dramatically different even when viewed in fixed pose. To handle this variability, an object recognition system must employ a representation that is either invariant to, or models this variability. This paper presents an appearance-based method for modeling the variability due to illumination in the images of objects. The method differs from past appearance-based methods, however, in that a small set of training images is used to generate a representation-the illumination cone-which models the complete set of images of an object with Lambertian reflectance map under an arbitrary combination of point light sources at infinity. This method is both an implementation and extension (an extension in that it models cast shadows) of the illumination cone representation proposed in Belhumeur and Kriegman (1996). The method is tested on a database of 660 images of 10 faces, and the results exceed those of popular existing methods.
Athinodoros S. Georghiades, David J. Kriegman, Peter N. Belhumeur
CVPR2
1998 What Shadows Reveal about Object Structure
David J. Kriegman, Peter N. Belhumeur
ECCV (2)1
1998 Invariant-Based Recognition of Complex Curved 3D Objects from Image Contours
abstract
This paper addresses the problem of recognizing three-dimensional objects bounded by smooth curved surfaces from image contours found in a single photograph. The proposed approach is based on a viewpoint-invariant relationship between object geometry and certain image features under weak perspective projection. The image features themselves are viewpoint-dependent. Concretely, the set of all possible silhouette bitangents, along with the contour points sharing the same tangent direction, is the projection of a one-dimensional set of surface points where each point lies on the occluding contour for a five-parameter family of viewpoints. These image features form a one-parameter family of equivalence classes, and it is shown that each class can be characterized by a set of numerical attributes that remain constant across the corresponding five-dimensional set of viewpoints. This is the basis for describing objects by “invariant” curves embedded in high-dimensional spaces. Modeling is achieved by moving an object in front of a camera and does not require knowing the object-to-camera transformation; nor does it involve implicit or explicit three-dimensional shape reconstruction. At recognition time, attributes computed from a single image are used to index the model database, and both qualitative and quantitative verification procedures eliminate potential false matches. The approach has been implemented and examples are presented.
B. Vijayakumar, David J. Kriegman, Jean Ponce
Comput. Vis. Image Underst.2
1998 What Is the Set of Images of an Object Under All Possible Illumination Conditions?
Peter N. Belhumeur, David J. Kriegman
Int. J. Comput. Vis.2
1998 Vision-based motion planning and exploration algorithms for mobile robots
abstract
This paper considers the problem of systematically exploring an unfamiliar environment in search of one or more recognizable targets. The proposed exploration algorithm is based on a novel representation of environments containing visual landmarks, called the boundary place graph. This representation records the set of recognizable objects (landmarks) that are visible from the boundary of each configuration space obstacle. The exploration algorithm constructs the boundary place graph incrementally from sensor data. Once the robot has completely explored an environment, it can use the constructed representation to carry out further navigation tasks. We provide a necessary and sufficient condition under which the algorithm is guaranteed to discover all landmarks. This algorithm has been implemented on our mobile robot platform RJ, and results from these experiments are presented.
Camillo J. Taylor, David J. Kriegman
IEEE Trans. Robotics Autom.2
1997 The Bas-Relief Ambiguity
abstract
Since antiquity, artisans have created flattened forms, often called "bas-reliefs,"-which give an exaggerated perception of depth when viewed from a particular vantage point. This paper presents an explanation of this phenomena, showing that the ambiguity in determining the relief of an object is not confined to bas-relief sculpture but is implicit in the determination of the structure of any object. Formally, if the object's true surface is denoted by z/sub true/=f(x, y), then we define the "generalized bas-relief transformation" as z=/spl lambda/f(x, y)+/spl mu/x+/spl nu/y, with a corresponding transformation of the albedo. For each image of a Lambertian surface f(x, y) produced by a point light source at infinity, there exists an identical image of a bas-relief produced by a transformed light source. This equality holds for both shaded and shadowed regions. Thus, the set of possible images (illumination cone) is invariant over generalized bas-relief transformations. When /spl mu/=/spl nu/=0 (e.g. a classical bas-relief sculpture), we show that the set of possible motion fields are also identical. Thus, neither small unknown motions nor changes of illumination can resolve the bas-relief ambiguity. Implications of this ambiguity on structure recovery and shape representation are discussed.
Peter N. Belhumeur, David J. Kriegman, Alan L. Yuille
CVPR2
1997 Image-based prediction of landmark features for mobile robot navigation
abstract
We have been developing an architecture for vision-based navigation which relies on continuous feedback from visual "landmarks" to control robot motion, In this approach, landmarks are consistently located and acquired as they come into view. To make this process efficient and robust, it is important that the image locations of these features can be predicted from available image information. In this article, we discuss methods for direct image-based prediction of point and line features for a mobile system operating on a planar surface. Preliminary experimental results suggest that image-based prediction con be performed efficiently and with sufficient accuracy to ensure robust acquisition of navigational landmarks.
Gregory D. Hager, David J. Kriegman, Erliang Yeh, Christopher Rasmussen
ICRA2
1997 Hot curves for modelling and recognition of smooth curved 3D objects
abstract
: We represent arbitrary smooth curved 3D shapes by a discrete set of HOT curves where a surface admits High Order Tangents. These curves determine the structure of the image contours and its catastrophic changes, and there is a natural correspondence between some of them and monocular contour features such as inflections and bitangents. We present a method for automatically constructing the HOT curves from continuous sequences of video images and describe an approach to object recognition using viewpoint-dependent monocular image features as indices into a database of models and as a basis for pose estimation. We have implemented both the methods and present results obtained from real images. 1 Introduction While implemented recognition systems based on parametric shape representations such as algebraic surfaces or superquadrics have demonstrated their usefulness, the ultimate utility of a representation is limited by its scope. This suggests looking for a more general representation ...
Tanuja Joshi, B. Vijayakumar, David J. Kriegman, Jean Ponce
Image Vis. Comput.3
1997 Eigenfaces vs. Fisherfaces: Recognition Using Class Specific Linear Projection
abstract
We develop a face recognition algorithm which is insensitive to large variation in lighting direction and facial expression. Taking a pattern classification approach, we consider each pixel in an image as a coordinate in a high-dimensional space. We take advantage of the observation that the images of a particular face, under varying illumination but fixed pose, lie in a 3D linear subspace of the high dimensional image space-if the face is a Lambertian surface without shadowing. However, since faces are not truly Lambertian surfaces and do indeed produce self-shadowing, images will deviate from this linear subspace. Rather than explicitly modeling this deviation, we linearly project the image into a subspace in a manner which discounts those regions of the face with large deviation. Our projection method is based on Fisher's linear discriminant and produces well separated classes in a low-dimensional subspace, even under severe variation in lighting and facial expressions. The eigenface technique, another method based on linearly projecting the image space to a low dimensional subspace, has similar computational requirements. Yet, extensive experimental results demonstrate that the proposed "Fisherface" method has error rates that are lower than those of the eigenface technique for tests on the Harvard and Yale face databases.
Peter N. Belhumeur, João Pedro Hespanha, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.3
1996 What is the set of images of an object under all possible lighting conditions?
abstract
The appearance of a particular object depends on both the viewpoint from which it is observed and the light sources by which it is illuminated. If the appearance of two objects is never identical for any pose or lighting conditions, then-in theory - the objects can always be distinguished or recognized. The question arises: What is the set of images of an object under all lighting conditions and pose? In this paper, we consider only the set of images of an object under variable illumination (including multiple, extended light sources and attached shadows). We prove that the set of n-pixel images of a convex object with a Lambertian reflectance function, illuminated by an arbitrary number of point light sources at infinity, forms a convex polyhedral cone in IR/sup n/ and that the dimension of this illumination cone equals the number of distinct surface normals. Furthermore, we show that the cone for a particular object can be constructed from three properly chosen images. Finally, we prove that the set of n-pixel images of an object of any shape and with an arbitrary reflectance function, seen under all possible illumination conditions, still forms a convex cone in IR/sup n/. These results immediately suggest certain approaches to object recognition. Throughout this paper, we offer results demonstrating the empirical validity of the illumination cone representation.
Peter N. Belhumeur, David J. Kriegman
CVPR2
1996 Structure and motion of curved 3D objects from monocular silhouettes
abstract
The silhouette of a smooth 3D object observed by a moving camera changes over time. Past work has shown how surface geometry can be recovered using the deformation of the silhouette when the camera motion is known. This paper addresses the problem of estimating both the full Euclidean surface structure and the camera motion from a dense set of silhouettes captured under orthographic or scaled orthographic projection. The approach relies on a viewpoint-invariant representation of curves swept by viewpoint-dependent features such as bitangents, inflections and contour points with parallel tangents. Feature points, which form stereo frontier points between non-consecutive images, are matched using this representation. The camera's angular velocity is computed from constraints derived from this correspondence along with the image velocity of these features. From the angular velocity, the epipolar geometry is ascertained, and infinitesimal motion frontier points can be detected. In turn, the motion of these frontier points constrains the translation component of camera motion. Finally, the surface is reconstructed using established techniques once the camera motion has been estimated.
B. Vijayakumar, David J. Kriegman, Jean Ponce
CVPR2
1996 Eigenfaces vs. Fisherfaces: Recognition Using Class Specific Linear Projection
Peter N. Belhumeur, João Pedro Hespanha, David J. Kriegman
ECCV (1)3
1996 Complete algorithms for feeding polyhedral parts using pivot grasps
abstract
To rapidly feed industrial parts on an assembly line, Carlisle et al. (1994) proposed a flexible part feeding system that drops parts on a flat conveyor belt, determines position and orientation of each part with a vision system, and then moves them into a desired orientation. When a part is grasped with two hard finger contacts and lifted, it pivots under gravity into a stable configuration. The authors refer to the sequence of picking up the part, allowing it to pivot, and replacing it on the table as a pivot grasp. The authors show that under idealized conditions, a robot arm with four degrees of freedom (DOF) can move (feed) parts arbitrarily in 6 DOF using pivot grasps. This paper considers the following planning problem: Given a polyhedral part shape, coefficient of friction, and a pair of stable configurations as input, find pairs of grasp points that will cause the part to pivot from one stable configuration to the other. For a part with n faces and m stable configurations, the authors give an O(m/sup 2/n log n) algorithm to generate the m/spl times/m matrix of pivot grasps. When the part is star-shaped, this reduces to O(m/sup 2/n). Since pivot grasps may not exist for some transitions, multiple steps may be needed. Alternatively, the authors consider the set of grasps where the part pivots to a configuration within a "capture region" around the stable configuration; when the part is released, it will tumble to the desired configuration. Both algorithms are complete in that they are guaranteed to find pivot grasps when they exist.
Anil S. Rao, David J. Kriegman, Kenneth Y. Goldberg
IEEE Trans. Robotics Autom.2
1995 Complete Algorithms for Reorienting Polyhedral Parts Using a Pivoting Gripper
abstract
No abstract available.
Anil S. Rao, David J. Kriegman, Kenneth Y. Goldberg
SCG2
1995 Invariant-Based Recognition of Complex Curved 3D Objects from Image Contours
abstract
To recognize three-dimensional objects bounded by smooth curved surfaces from monocular image contours, viewpoint-dependent image features must be related to object geometry. Contour bitangents and inflections along with associated parallel tangents points are the projection of surface points that lie on the occluding contour for a five-parameter family of scaled orthographic projection viewpoints. An invariant representation can be computed from these image features and seen for modeling and recognizing objects. Modeling is achieved by moving an object in front of a camera to obtain a curve of possible invariants. The relative camera-object motion is not required, and 3D models are not utilized. At recognition time, invariants computed from a single image are used to index the model database. Using the matched features, independent qualitative and quantitative verification procedures eliminate potential false matches. Examples from an implementation are presented.>
B. Vijayakumar, David J. Kriegman, Jean Ponce
ICCV2
1995 Complete Algorithms for Reorienting Polyhedral Parts Using a Pivoting Gripper
abstract
To rapidly feed industrial parts on an assembly line, Carlisle et. al. (1994) proposed a flexible part feeding system that drops parts on a flat conveyor belt, determines the pose of parts with a vision system and manipulates them into a desired pose. A robot arm with 4-DOF is capable of moving parts through 6-DOF when equipped with a passive pivoting axis between the parallel jaws of its gripper. We refer to these actions as pivot grasps. This paper considers the planning problem. Given a polyhedral part shape, coefficient of friction and a pair of stable configurations as input, find pairs of grasp points that will cause the part to pivot from one stable configuration to the other. For some transitions, pivot grasps may not exist. For a part with n faces and m stable configurations, we give an O(m/sup 2/n log n) algorithm to generate the m/spl times/m matrix of pivot grasps. When the part is star shaped, this reduces to O(m/sup 2/n). We also study a generalization that considers "capture regions" around stable configurations. Both algorithms are complete in that they are guaranteed to find pivot grasps when they exist.
Anil S. Rao, David J. Kriegman, Kenneth Y. Goldberg
ICRA2
1995 Toward selecting and recognizing natural landmarks
abstract
Landmarks are often used as a basis for mobile robot navigation. In this paper, we consider the problem of automatically selecting from a set of 3D features the subset which is most likely to be recognized from noisy monocular image data and is least likely to be confused with any other group of features. Assuming perspective projection, real valued recognition functions are constructed for a set of features. The value returned from such functions are invariant to changes of viewpoint and can be evaluated directly from image measurements without prior knowledge of the position and orientation of the camera. With image noise, the recognition function no longer evaluates to a constant value. Because of the possibility of false matches, a Bayes detector is used to determine the optimal range of values of the recognition function that will be accepted as image features of the model. The model with the lowest Bayes cost is selected as the most distinguishable landmark. We show implementation results for real 3D objects.
Erliang Yeh, David J. Kriegman
IROS (1)2
1995 Structure and Motion from Line Segments in Multiple Images
abstract
This paper presents a new method for recovering the three dimensional structure of a scene composed of straight line segments using the image data obtained from a moving camera. The recovery algorithm is formulated in terms of an objective function which measures the total squared distance in the image plane between the observed edge segments and the projections (perspective) of the reconstructed lines. This objective function is minimized with respect to the line parameters and the camera positions to obtain an estimate for the structure of the scene. The effectiveness of this approach is demonstrated quantitatively through extensive simulations and qualitatively with actual image sequences. The implementation is being made publicly available.>
Camillo J. Taylor, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 HOT curves for modelling and recognition of smooth curved 3D objects
abstract
Arbitrary smooth curved 3D shapes are represented by a discrete set of high-order tangent (HOT) curves, where a surface admits HOTs. These curves determine the structure of the image contours and its catastrophic changes, and there is a natural correspondence between some of them and monocular contour features such as inflections and bitangents. We present a method for automatically constructing the HOT curves from continuous sequences of video images and describe an approach to object recognition using viewpoint-dependent monocular image features as indices into a database of models and as a basis for pose estimation. We have implemented both of the methods, and present results obtained from real images.>
Tanuja Joshi, Jean Ponce, B. Vijayakumar, David J. Kriegman
CVPR4
1994 Let Them Fall Where They May: Capture Regions of Curved 3D Objects
abstract
When an object is placed on a supporting plane, gravitational forces move it to one of a finite set of stable poses. For each stable pose, there is a region in configuration space (a capture region) from which the object is guaranteed to converge to that pose. The problem of computing maximal capture regions for an object with a smooth convex hull is analyzed assuming only that its dynamics are dissipative; the precise equations governing the system are unnecessary. As examples from an implemented algorithm demonstrate, calculating these regions from a geometric model is computationally practical.>
David J. Kriegman
ICRA1
1994 Parameterized Families of Polynomials for Bounded Algebraic Curve and Surface Fitting
abstract
Interest in algebraic curves and surfaces of high degree as geometric models or shape descriptors for different model-based computer vision tasks has increased in recent years, and although their properties make them a natural choice for object recognition and positioning applications, algebraic curve and surface fitting algorithms often suffer from instability problems. One of the main reasons for these problems is that, while the data sets are always bounded, the resulting algebraic curves or surfaces are, in most cases, unbounded. In this paper, the authors propose to constrain the polynomials to a family with bounded zero sets, and use only members of this family in the fitting process. For every even number d the authors introduce a new parameterized family of polynomials of degree d whose level sets are always bounded, in particular, its zero sets. This family has the same number of degrees of freedom as a general polynomial of the same degree. Three methods for fitting members of this polynomial family to measured data points are introduced. Experimental results of fitting curves to sets of points in R/sup 2/ and surfaces to sets of points in R/sup 3/ are presented.>
Gabriel Taubin, Fernando Cukierman, Steve Sullivan, Jean Ponce, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.5
1993 Reconstruction of HOT curves from image sequences
abstract
An approach is presented for reconstructing two types of 3-D higher order tangency (HOT) curves from a sequence of images. These curves are useful for object recognition. The reconstruction results for bitangents are encouraging in comparison to those of inflections. They are probably more accurately reconstructed because they are readily located in images, and their common tangent is very accurately estimated from the point locations. This makes bitangents a good feature choice for recognition.>
David J. Kriegman, B. Vijayakumar, Jean Ponce
CVPR1
1992 Parametrizing and fitting bounded algebraic curves and surfaces
abstract
An approach to fitting of implicit algebraic curves and surfaces to point data is introduced. Two families of polynomials with bounded zero sets are presented. Members of these families have the same number of degrees of freedom as general polynomials of the same degree. Methods for fitting members of these families of polynomials to measured data points are described. Experimental results for sets of points in R/sup 2/ and R/sup 3/ for curves and surfaces, respectively, are presented.>
Gabriel Taubin, Fernando Cukierman, Steve Sullivan, Jean Ponce, David J. Kriegman
CVPR5
1992 Constraints for Recognizing and Locating Curved 3D Objects from Monocular Image Features
David J. Kriegman, B. Vijayakumar, Jean Ponce
ECCV1
1992 Computing Exact Aspect Graphs of Curved Objects: Algebraic Surfaces
Jean Ponce, Sylvain Petitjean, David J. Kriegman
ECCV3
1992 Structure and motion from line segments in multiple images
abstract
An approach for recovering the structure of a rigid scene composed of straight line segments and the pose of moving camera from multiple images under perspective projection is presented. Recovery is formulated as minimizing the total squared image distance between measured segments and the projection of the reconstructed infinite lines with respect to structural and motion parameters. An efficient algorithm for minimizing this nonlinear objective function is presented. The approach can be directly used for accurately locating correspondences in multieyed stereo, as well as for pose estimation from landmarks. Its effectiveness is demonstrated quantitatively through simulation and qualitatively with a real image sequence.>
Camillo J. Taylor, David J. Kriegman
ICRA2
1992 Computing stable poses of piecewise smooth objects
abstract
When a three dimensional object is known to be lying on a planar surface, its pose is restricted from six to three degrees of freedom. Computer vision algorithms can exploit the few stable poses of modeled objects to simplify scene interpretation and more accurately determine object location. This paper presents necessary and sufficient conditions for the pose of a piecewise smooth curved three-dimensional object to be stable. For objects whose surfaces are represented by implicit algebraic equations, these conditions can be expressed as systems of polynomial equations that are readily solved by homotopy continuation. Examples from the implemented algorithm are presented.
David J. Kriegman
CVGIP Image Underst.1
1992 On using CAD models to compute the pose of curved 3D objects
Jean Ponce, Anthony Hoogs, David J. Kriegman
CVGIP Image Underst.3
1992 Computing exact aspect graphs of curved objects: Algebraic surfaces
Sylvain Petitjean, Jean Ponce, David J. Kriegman
Int. J. Comput. Vis.3
1990 Computing Exact Aspect Graphs of Curved Objects: Parametric Surfaces
Jean Ponce, David J. Kriegman
AAAI2
1990 Computing exact aspect graphs of curved objects: Solids of revolution
David J. Kriegman, Jean Ponce
Int. J. Comput. Vis.1
1990 On Recognizing and Positioning Curved 3-D Objects from Image Contours
abstract
An approach for explicitly relating the shape of image contours to models of curved three-dimensional objects is presented. This relationship is used for object recognition and positioning. Object models consist of collections of parametric surface patches and their intersection curves; this includes nearly all representations used in computer-aided geometric design and computer vision. The image contours considered are the projections of surface discontinuities and occluding contours. Elimination theory provides a method for constructing the implicit equation of these contours for an object observed under orthographic or perspective projection. This equation is parameterized by the object's position and orientation with respect to the observer. Determining these parameters is reduced to a fitting problem between the theoretical contour and the observed data points. The proposed approach readily extends to parameterized models. It has been implemented for a simple world composed of various surfaces of revolution and tested on several real images.>
David J. Kriegman, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.1
1989 Stereo vision and navigation in buildings for mobile robots
abstract
A mobile robot that autonomously functions in a complex and previously unknown indoor environment has been developed. The omnidirectional mobile robot uses stereo vision, odometry, and contact bumpers to instantiate a symbolic world model. Finding stereo correspondences across a single epipolar line is adequate for instantiating the model. Uncertainty in sensor data is represented by a multivariate normal distribution, and uncertainty models for motion and stereo are presented. Uncertainty is reduced by extended Kalman filtering. To execute a high-level command such as 'Enter the second door on the left', a model is instantiated from sensing and motions are planned and executed. Experimental results from the fast, running system are presented.>
David J. Kriegman, Ernst E. Triendl, Thomas O. Binford
IEEE Trans. Robotics Autom.1
1988 Generic models for robot navigation
abstract
The authors introduce the concept of an explicit generic building model. This abstract model is valid for a broad class of buildings. It uses mechanisms that have been used to build generic models for other object classes, including machine screws and generalized cylinders. These methods extend the generality of object classes which can be implemented, compared to previous implementations of object classes. Using a combination of epipolar stereo vision from the mobile robot along with monocular vision, the generic model is instantiated. This instantiation can be used for robot navigation.>
David J. Kriegman, Thomas O. Binford
ICRA1
1987 A mobile robot: Sensing, planning and locomotion
abstract
A mobile robot architecture must include sensing, planning, and locomotion which are tied together by a model or map of the world based on sensor information, apriori knowledge and generic models. The architecture of a Stanford's autonomous mobile robot is described including its distributed computing system, locomotion, and sensing. Additionally, some of the issues in the representation of a world model are explored. Sensor models are used to update the world model in a uniform manner, and uncertainty reduction is discussed.
David J. Kriegman, Ernst E. Triendl, Thomas O. Binford
ICRA1
1987 Stereo vision and navigation within buildings
abstract
Soft modeling, stereo vision, motion planning, uncertainty reduction, image processing, and locomotion enable the Mobile Autonomous Robot Stanford to explore a benign indoor environment without human intervention. The modeling system describes rooms in terms of floor, walls, hinged doors and allows for unspecified obstacles. Image processing basically extracts vertical edges along the horizon using an edge appearance model. Stereo vision matches those edges using edge and grey level similarity, constraint propagation and a preference for epipolar ordering. The motion planner tries to move in a way that is likely to increase knowledge about obstacle free space. Results presented are from an autonomous run that included difficult passages such as navigation around a pillar without apriori knowledge.
Ernst E. Triendl, David J. Kriegman
ICRA2
1985 Computational architecture for the Utah/MIT hand
abstract
This paper presents the computational architecture for the Utah-MIT hand, and discusses design issues encountered in its hardware and software development. The large number of linkages, actuators, and sensors offers a potentially severe computational burden for control; a multiprocessor hardware configuration has been developed to distribute this computation. In the interests of efficiency, minimal operating systems were devised for each processor which nevertheless were made sufficiently general for task scheduling, intertask communication, and debugging. This computational architecture is a general system which is potentially useful for other robotics applications.
David J. Kriegman, David M. Siegel, Sundar Narasimhan, John M. Hollerbach, George E. Gerpheide
ICRA1