David Jacobs 0001

dblp:j/DavidWJacobs · also David W. Jacobs · DBLP profile ↗
← Back
135ranked-venue papers
18as first author
20since 2021 · last 2025
0009-0009-2027-1039ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 122 · 18 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 9 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Security and privacy · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Multimodal Agentic Model Predictive Control
Saptarashmi Bandyopadhyay, John (Jack) Cole, Tom Goldstein, David Jacobs 0001
AAMAS4
2025 MyTimeMachine: Personalized Facial Age Transformation
abstract
Facial aging is a complex process, highly dependent on multiple factors like gender, ethnicity, lifestyle, etc., making it extremely challenging to learn a global aging prior to predict aging for any individual accurately. Existing techniques often produce realistic and plausible aging results, but the re-aged images often do not resemble the person's appearance at the target age and thus need personalization. In many practical applications of virtual aging, e.g. VFX in movies and TV shows, access to a personal photo collection of the user depicting aging in a small time interval (20~40 years) is often available. However, naive attempts to personalize global aging techniques on personal photo collections often fail. Thus, we propose MyTimeMachine (MyTM), a method that combines a global aging prior with a personalized photo collection (ranging from as few as 10 images, ideally 50) to learn individualized age transformations. We introduce a novel Adapter Network that combines personalized aging features with global aging features and generates a re-aged image with StyleGAN2. We also introduce three loss functions to personalize the Adapter Network with personalized aging loss, extrapolation regularization, and adaptive w-norm regularization. Our method demonstrates strong performance on fair-use imagery of widely recognizable individuals, producing photorealistic and identity-consistent age transformations that generalize well across diverse appearances. It also extends naturally to video, delivering high-quality, temporally consistent results that closely resemble actual appearances at target ages—outperforming state-of-the-art approaches.
Luchao Qi, Jiaye Wu 0001, Bang Gong, Annie N. Wang, David Jacobs 0001, Roni Sengupta
ACM Trans. Graph.5
2024 Rethinking Score Distillation as a Bridge Between Image Distributions
abstract
Score distillation sampling (SDS) has proven to be an important tool, enabling the use of large-scale diffusion priors for tasks operating in data-poor domains. Unfortunately, SDS has a number of characteristic artifacts that limit its utility in general-purpose applications. In this paper, we make progress toward understanding the behavior of SDS and its variants by viewing them as solving an optimal-cost transport path from some current source distribution to a target distribution. Under this new interpretation, we argue that these methods' characteristic artifacts are caused by (1) linear approximation of the optimal path and (2) poor estimates of the source distribution. We show that by calibrating the text conditioning of the source distribution, we can produce high-quality generation and translation results with little extra overhead. Our method can be easily applied across many domains, matching or beating the performance of specialized methods. We demonstrate its utility in text-to-2D, text-to-3D, translating paintings to real images, optical illusion generation, and 3D sketch-to-real. We compare our method to existing approaches for score distillation sampling and show that it can produce high-frequency details with realistic colors.
David McAllister, Songwei Ge, Jia-Bin Huang 0001, David Jacobs 0001, Alexei A. Efros, Aleksander Holynski, Angjoo Kanazawa
NeurIPS4
2024 CALVIN: Improved Contextual Video Captioning via Instruction Tuning
abstract
The recent emergence of powerful Vision-Language models (VLMs) has significantly improved image captioning. Some of these models are extended to caption videos as well. However, their capabilities to understand complex scenes are limited, and the descriptions they provide for scenes tend to be overly verbose and focused on the superficial appearance of objects. Scene descriptions, especially in movies, require a deeper contextual understanding, unlike general-purpose video captioning. To address this challenge, we propose a model, CALVIN, a specialized video LLM that leverages previous movie context to generate fully "contextual" scene descriptions. To achieve this, we train our model on a suite of tasks that integrate both image-based question-answering and video captioning within a unified framework, before applying instruction tuning to refine the model's ability to provide scene captions. Lastly, we observe that our model responds well to prompt engineering and few-shot in-context learning techniques, enabling the user to adapt it to any new movie with very little additional annotation.
Gowthami Somepalli, Arkabandhu Chowdhury, Jonas Geiping, Ronen Basri, Tom Goldstein, David Jacobs 0001
NeurIPS6
2023 Hyperbolic Contrastive Learning for Visual Representations beyond Objects
abstract
Although self-/un-supervised methods have led to rapid progress in visual representation learning, these methods generally treat objects and scenes using the same lens. In this paper, we focus on learning representations for objects and scenes that preserve the structure among them. Motivated by the observation that visually similar objects are close in the representation space, we argue that the scenes and objects should instead follow a hierarchical structure based on their compositionality. To exploit such a structure, we propose a contrastive learning framework where a Euclidean loss is used to learn object representations and a hyperbolic loss is used to encourage representations of scenes to lie close to representations of their constituent objects in a hyperbolic space. This novel hyperbolic objective encourages the scene-object hypernymy among the representations by optimizing the magnitude of their norms. We show that when pretraining on the COCO and OpenImages datasets, the hyperbolic loss improves downstream performance of several baselines across multiple datasets and tasks, including image classification, object detection, and semantic segmentation. We also show that the properties of the learned representations allow us to solve various vision tasks that involve the interaction between scenes and objects in a zero-shot fashion.
Songwei Ge, Shlok Kumar Mishra, Simon Kornblith, Chun-Liang Li, David Jacobs 0001
CVPR5
2023 HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of Actions
abstract
Supervised learning of skeleton sequence encoders for action recognition has received significant attention in recent times. However, learning such encoders without labels continues to be a challenging problem. While prior works have shown promising results by applying contrastive learning to pose sequences, the quality of the learned representations is often observed to be closely tied to data augmentations that are used to craft the positives. However, augmenting pose sequences is a difficult task as the geometric constraints among the skeleton joints need to be enforced to make the augmentations realistic for that action. In this work, we propose a new contrastive learning approach to train models for skeleton-based action recognition without labels. Our key contribution is a simple module, HaLP - to Hallucinate Latent Positives for contrastive learning. Specifically, HaLP explores the latent space of poses in suitable directions to generate new positives. To this end, we present a novel optimization formulation to solve for the synthetic positives with an explicit control on their hardness. We propose approximations to the objective, making them solvable in closed form with minimal overhead. We show via experiments that using these generated positives within a standard contrastive learning framework leads to consistent improvements across benchmarks such as NTU-60, NTU-120, and PKU-II on tasks like linear evaluation, transfer learning, and kNN evaluation. Our code can be found at https://github.com/anshulbshah/HaLP.
Anshul Shah 0001, Aniket Roy, Ketul Shah, Shlok Kumar Mishra, David Jacobs 0001, Anoop Cherian, Rama Chellappa
CVPR5
2023 Measured Albedo in the Wild: Filling the Gap in Intrinsics Evaluation
abstract
Intrinsic image decomposition and inverse rendering are long-standing problems in computer vision. To evaluate albedo recovery, most algorithms report their quantitative performance with a mean Weighted Human Disagreement Rate (WHDR) metric on the IIW dataset. However, WHDR focuses only on relative albedo values and often fails to capture overall quality of the albedo. In order to comprehensively evaluate albedo, we collect a new dataset, Measured Albedo in the Wild (MAW), and propose three new metrics that complement WHDR: intensity, chromaticity and texture metrics. We show that existing algorithms often improve WHDR metric but perform poorly on other metrics. We then finetune different algorithms on our MAW dataset to significantly improve the quality of the reconstructed albedo both quantitatively and qualitatively. Since the proposed intensity, chromaticity, and texture metrics and the WHDR are all complementary we further introduce a relative performance measure that captures average performance. By analysing existing algorithms we show that there is significant room for improvement. Our dataset and evaluation metrics will enable researchers to develop algorithms that improve albedo reconstruction.
Jiaye Wu 0001, Sanjoy Chowdhury, Hariharmano Shanmugaraja, David Jacobs 0001, Roni Sengupta
ICCP4
2023 Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
abstract
Despite tremendous progress in generating high-quality images using diffusion models, synthesizing a sequence of animated frames that are both photorealistic and temporally coherent is still in its infancy. While off-the-shelf billion-scale datasets for image generation are available, collecting similar video data of the same scale is still challenging. Also, training a video diffusion model is computationally much more expensive than its image counterpart. In this work, we explore finetuning a pretrained image diffusion model with video data as a practical solution for the video synthesis task. We find that naively extending the image noise prior to video noise prior in video diffusion leads to sub-optimal performance. Our carefully designed video noise prior leads to substantially better performance. Extensive experimental validation shows that our model, Preserve Your Own COrrelation (PYoCo), attains SOTA zero-shot text-to-video results on the UCF-101 and MSR-VTT benchmarks. It also achieves SOTA video generation quality on the small-scale UCF-101 benchmark with a 10× smaller model using significantly less computation than the prior art. The project page is available at https://research.nvidia.com/labs/dir/pyoco/.
Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs 0001, Jia-Bin Huang 0001, Ming-Yu Liu 0001, Yogesh Balaji
ICCV7
2023 LD-ZNet: A Latent Diffusion Approach for Text-Based Image Segmentation
abstract
Large-scale pre-training tasks like image classification, captioning, or self-supervised techniques do not incentivize learning the semantic boundaries of objects. However, recent generative foundation models built using text-based latent diffusion techniques may learn semantic boundaries. This is because they have to synthesize intricate details about all objects in an image based on a text description. Therefore, we present a technique for segmenting real and AI-generated images using latent diffusion models (LDMs) trained on internet-scale datasets. First, we show that the latent space of LDMs (z-space) is a better input representation compared to other feature representations like RGB images or CLIP encodings for text-based image segmentation. By training the segmentation models on the latent z-space, which creates a compressed representation across several domains like different forms of art, cartoons, illustrations, and photographs, we are also able to bridge the domain gap between real and AI-generated images. We show that the internal features of LDMs contain rich semantic information and present a technique in the form of LD-ZNet to further boost the performance of text-based segmentation. Overall, we show up to 6% improvement over standard baselines for text-to-image segmentation on natural images. For AI-generated imagery, we show close to 20% improvement compared to state-of-the-art techniques. The project is available at https://koutilya-pnvr.github.io/LD-ZNet/.
Koutilya PNVR, Pallabi Ghosh, Behjat Siddiquie, David Jacobs 0001
ICCV5
2022 Learning visual representations for transfer learning by suppressing texture
Shlok Kumar Mishra, Anshul Shah 0001, Ankan Bansal, Janit Anjaria, Abhinav Shrivastava, Abhishek Sharma 0001, David Jacobs 0001
BMVC8
2022 Fast Light-Weight Near-Field Photometric Stereo
abstract
We introduce the first end-to-end learning-based solution to near-field Photometric Stereo (PS), where the light sources are close to the object of interest. This setup is especially useful for reconstructing large immobile objects. Our method is fast, producing a mesh from 52 512x384 resolution images in about 1 second on a commodity GPU, thus potentially unlocking several AR/VR applications. Existing approaches rely on optimization coupled with a far-field PS network operating on pixels or small patches. Using optimization makes these approaches slow and memory intensive (requiring 17GB GPU and 27GB of CPU memory) while using only pixels or patches makes them highly sus-ceptible to noise and calibration errors. To address these issues, we develop a recursive multi-resolution scheme to estimate surface normal and depth maps of the whole image at each step. The predicted depth map at each scale is then used to estimate 'per-pixel lighting, for the next scale. This design makes our approach almost 45x faster and 2° more accurate (11.3° vs. 13.3° Mean Angular Error) than the state-of-the-art near-field PS reconstruction technique, which uses iterative optimization.
Daniel Lichy, Roni Sengupta, David Jacobs 0001
CVPR3
2022 Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin 0001, Guan Pang, David Jacobs 0001, Jia-Bin Huang 0001, Devi Parikh
ECCV (17)6
2022 Improved Presentation Attack Detection Using Image Decomposition
abstract
Presentation attack detection (PAD) is a critical component in secure face authentication. We present a PAD algorithm to distinguish face spoofs generated by a photograph of a subject from live images. Our method uses an image decomposition network to extract albedo and normal. The domain gap between the real and spoof face images leads to easily identifiable differences, especially between the re-covered albedo maps. We enhance this domain gap by retraining existing methods using supervised contrastive loss. We present empirical and theoretical analysis that demonstrates that contrast and lighting effects can play a significant role in PAD; these show up particularly in the recovered albedo. Finally, we demonstrate that by combining all of these methods we achieve state-of-the-art results on both intra-dataset testing for CelebA-Spoof, OULU, CASIA-SURF datasets and inter-dataset setting on SiW, CASIA-MFSD, Replay-Attack and MSU-MFSD datasets.
Shlok Kumar Mishra, Kuntal Sengupta, Wen-Sheng Chu, Max Horowitz-Gelb, Sofien Bouaziz, David Jacobs 0001
IJCB6
2022 On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process Kernels
abstract
We study the properties of various over-parameterized convolutional neural architectures through their respective Gaussian Process and Neural Tangent kernels. We prove that, with normalized multi-channel input and ReLU activation, the eigenfunctions of these kernels with the uniform measure are formed by products of spherical harmonics, defined over the channels of the different pixels. We next use hierarchical factorizable kernels to bound their respective eigenvalues. We show that the eigenvalues decay polynomially, quantify the rate of decay, and derive measures that reflect the composition of hierarchical features in these networks. Our theory provides a concrete quantitative characterization of the role of locality and hierarchy in the inductive bias of over-parameterized convolutional network architectures.
Amnon Geifman, Meirav Galun, David Jacobs 0001, Ronen Basri
NeurIPS3
2022 Autoregressive Perturbations for Data Poisoning
abstract
The prevalence of data scraping from social media as a means to obtain datasets has led to growing concerns regarding unauthorized use of data. Data poisoning attacks have been proposed as a bulwark against scraping, as they make data ``unlearnable'' by adding small, imperceptible perturbations. Unfortunately, existing methods require knowledge of both the target architecture and the complete dataset so that a surrogate network can be trained, the parameters of which are used to generate the attack. In this work, we introduce autoregressive (AR) poisoning, a method that can generate poisoned data without access to the broader dataset. The proposed AR perturbations are generic, can be applied across different datasets, and can poison different architectures. Compared to existing unlearnable methods, our AR poisons are more resistant against common defenses such as adversarial training and strong data augmentations. Our analysis further provides insight into what makes an effective data poison.
Pedro Sandoval Segura, Vasu Singla, Jonas Geiping, Micah Goldblum, Tom Goldstein, David Jacobs 0001
NeurIPS6
2022 SfSNet: Learning Shape, Reflectance and Illuminance of Faces in the Wild
abstract
We present SfSNet, an end-to-end learning framework for producing an accurate decomposition of an unconstrained human face image into shape, reflectance and illuminance. SfSNet is designed to reflect a physical lambertian rendering model. SfSNet learns from a mixture of labeled synthetic and unlabeled real-world images. This allows the network to capture low-frequency variations from synthetic and high-frequency details from real images through the photometric reconstruction loss. SfSNet consists of a new decomposition architecture with residual blocks that learns a complete separation of albedo and normal. This is used along with the original image to predict lighting. SfSNet produces significantly better quantitative and qualitative results than state-of-the-art methods for inverse rendering and independent normal and illumination estimation. We also introduce a companion network, SfSMesh, that utilizes normals estimated by SfSNet to reconstruct a 3D face mesh. We demonstrate that SfSMesh produces face meshes with greater accuracy than state-of-the-art methods on real-world images.
Roni Sengupta, Daniel Lichy, Angjoo Kanazawa, Carlos Domingo Castillo, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Shape and Material Capture at Home
abstract
In this paper, we present a technique for estimating the geometry and reflectance of objects using only a camera, flashlight, and optionally a tripod. We propose a simple data capture technique in which the user goes around the object, illuminating it with a flashlight and capturing only a few images. Our main technical contribution is the introduction of a recursive neural architecture, which can predict geometry and reflectance at 2k×2kresolution given an input image at 2k×2kand estimated geometry and reflectance from the previous step at 2k−1×2k−1. This recursive architecture, termed RecNet, is trained with 256×256 resolution but can easily operate on 1024×1024 images during inference. We show that our method produces more accurate surface normal and albedo, especially in regions of specular highlights and cast shadows, compared to previous approaches, given three or fewer input images.
Daniel Lichy, Jiaye Wu 0001, Roni Sengupta, David Jacobs 0001
CVPR4
2021 Low Curvature Activations Reduce Overfitting in Adversarial Training
abstract
Adversarial training is one of the most effective defenses against adversarial attacks. Previous works suggest that overfitting is a dominant phenomenon in adversarial training leading to a large generalization gap between test and train accuracy in neural networks. In this work, we show that the observed generalization gap is closely related to the choice of the activation function. In particular, we show that using activation functions with low (exact or approximate) curvature values has a regularization effect that significantly reduces both the standard and robust generalization gaps in adversarial training. We observe this effect for both differentiable/smooth activations such as SiLU as well as non-differentiable/non-smooth activations such as LeakyReLU. In the latter case, the "approximate" curvature of the activation is low. Finally, we show that for activation functions with low curvature, the double descent phenomenon for adversarially trained models does not occur.
Vasu Singla, Sahil Singla 0002, Soheil Feizi, David Jacobs 0001
ICCV4
2021 Robust Contrastive Learning Using Negative Samples with Diminished Semantics
abstract
Unsupervised learning has recently made exceptional progress because of the development of more effective contrastive learning methods. However, CNNs are prone to depend on low-level features that humans deem non-semantic. This dependency has been conjectured to induce a lack of robustness to image perturbations or domain shift. In this paper, we show that by generating carefully designed negative samples, contrastive learning can learn more robust representations with less dependence on such features. Contrastive learning utilizes positive pairs which preserve semantic information while perturbing superficial features in the training images. Similarly, we propose to generate negative samples in a reversed way, where only the superfluous instead of the semantic features are preserved. We develop two methods, texture-based and patch-based augmentations, to generate negative samples. These samples achieve better generalization, especially under out-of-domain settings. We also analyze our method and the generated texture-based samples, showing that texture features are indispensable in classifying particular ImageNet classes and especially finer classes. We also show that the model bias between texture and shape features favors them differently under different test settings.
Songwei Ge, Shlok Kumar Mishra, Chun-Liang Li, Haohan Wang, David Jacobs 0001
NeurIPS5
2021 Shift Invariance Can Reduce Adversarial Robustness
abstract
Shift invariance is a critical property of CNNs that improves performance on classification. However, we show that invariance to circular shifts can also lead to greater sensitivity to adversarial attacks. We first characterize the margin between classes when a shift-invariant {\em linear} classifier is used. We show that the margin can only depend on the DC component of the signals. Then, using results about infinitely wide networks, we show that in some simple cases, fully connected and shift-invariant neural networks produce linear decision boundaries. Using this, we prove that shift invariance in neural networks produces adversarial examples for the simple case of two classes, each consisting of a single image with a black or white dot on a gray background. This is more than a curiosity; we show empirically that with real datasets and realistic architectures, shift invariance reduces adversarial robustness. Finally, we describe initial experiments using synthetic data to probe the source of this connection.
Vasu Singla, Songwei Ge, Ronen Basri, David Jacobs 0001
NeurIPS4
2020 Making L-BFGS Work with Industrial-Strength Nets
Abhay Kumar Yadav, Tom Goldstein, David Jacobs 0001
BMVC3
2020 SharinGAN: Combining Synthetic and Real Data for Unsupervised Geometry Estimation
abstract
We propose a novel method for combining synthetic and real images when training networks to determine geometric information from a single image. We suggest a method for mapping both image types into a single, shared domain. This is connected to a primary network for end-to-end training. Ideally, this results in images from two domains that present shared information to the primary network. Our experiments demonstrate significant improvements over the state-of-the-art in two important domains, surface normal estimation of human faces and monocular depth estimation for outdoor scenes, both in an unsupervised setting.
Koutilya PNVR, Hao Zhou 0011, David Jacobs 0001
CVPR3
2020 Adversarially robust transfer learning
Ali Shafahi, Parsa Saadatpanah, Chen Zhu 0001, Amin Ghiasi, Christoph Studer, David Jacobs 0001, Tom Goldstein
ICLR6
2020 Frequency Bias in Neural Networks for Input of Non-Uniform Density
abstract
Recent works have partly attributed the generalization ability of over-parameterized neural networks to frequency bias – networks trained with gradient descent on data drawn from a uniform distribution find a low frequency fit before high frequency ones. As realistic training sets are not drawn from a uniform distribution, we here use the Neural Tangent Kernel (NTK) model to explore the effect of variable density on training dynamics. Our results, which combine analytic and empirical observations, show that when learning a pure harmonic function of frequency $\kappa$, convergence at a point $x \in \S^{d-1}$ occurs in time $O(\kappa^d/p(x))$ where $p(x)$ denotes the local density at $x$. Specifically, for data in $\S^1$ we analytically derive the eigenfunctions of the kernel associated with the NTK for two-layer networks. We further prove convergence results for deep, fully connected networks with respect to the spectral decomposition of the NTK. Our empirical study highlights similarities and differences between deep and shallow networks in this model.
Ronen Basri, Meirav Galun, Amnon Geifman, David Jacobs 0001, Yoni Kasten, Shira Kritchman
ICML4
2020 On the Similarity between the Laplace and Neural Tangent Kernels
abstract
Recent theoretical work has shown that massively overparameterized neural networks are equivalent to kernel regressors that use Neural Tangent Kernels (NTKs). Experiments show that these kernel methods perform similarly to real neural networks. Here we show that NTK for fully connected networks with ReLU activation is closely related to the standard Laplace kernel. We show theoretically that for normalized data on the hypersphere both kernels have the same eigenfunctions and their eigenvalues decay polynomially at the same rate, implying that their Reproducing Kernel Hilbert Spaces (RKHS) include the same sets of functions. This means that both kernels give rise to classes of functions with the same smoothness properties. The two kernels differ for data off the hypersphere, but experiments indicate that when data is properly normalized these differences are not significant. Finally, we provide experiments on real data comparing NTK and the Laplace kernel, along with a larger class of $\gamma$-exponential kernels. We show that these perform almost identically. Our results suggest that much insight about neural networks can be obtained from analysis of the well-known Laplace kernel, which has a simple closed form.
Amnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun, David Jacobs 0001, Ronen Basri
NeurIPS5
2019 Neural Inverse Rendering of an Indoor Scene From a Single Image
abstract
Inverse rendering aims to estimate physical attributes of a scene, e.g., reflectance, geometry, and lighting, from image(s). Inverse rendering has been studied primarily for single objects or with methods that solve for only one of the scene attributes. We propose the first learning based approach that jointly estimates albedo, normals, and lighting of an indoor scene from a single image. Our key contribution is the Residual Appearance Renderer (RAR), which can be trained to synthesize complex appearance effects (e.g., inter-reflection, cast shadows, near-field illumination, and realistic shading), which would be neglected otherwise. This enables us to perform self-supervised learning on real data using a reconstruction loss, based on re-synthesizing the input image from the estimated components. We finetune with real data after pretraining with synthetic data. To this end, we use physically-based rendering to create a large-scale synthetic dataset, named SUNCG-PBR, which is a significant improvement over prior datasets. Experimental results show that our approach outperforms state-of-the-art methods that estimate one or more scene attributes.
Roni Sengupta, Jinwei Gu, Guilin Liu, David Jacobs 0001, Jan Kautz
ICCV5
2019 Deep Single-Image Portrait Relighting
abstract
Conventional physically-based methods for relighting portrait images need to solve an inverse rendering problem, estimating face geometry, reflectance and lighting. However, the inaccurate estimation of face components can cause strong artifacts in relighting, leading to unsatisfactory results. In this work, we apply a physically-based portrait relighting method to generate a large scale, high quality, “in the wild” portrait relighting dataset (DPR). A deep Convolutional Neural Network (CNN) is then trained using this dataset to generate a relit portrait image by using a source image and a target lighting as input. The training procedure regularizes the generated results, removing the artifacts caused by physically-based relighting methods. A GAN loss is further applied to improve the quality of the relit portrait image. Our trained network can relight portrait images with resolutions as high as 1024 × 1024. We evaluate the proposed method on the proposed DPR datset, Flickr portrait dataset and Multi-PIE dataset both qualitatively and quantitatively. Our experiments demonstrate that the proposed method achieves state-of-the-art results. Please refer to https://zhhoper.github.io/dpr.html for dataset and code.
Hao Zhou 0011, Sunil Hadap, Kalyan Sunkavalli, David Jacobs 0001
ICCV4
2019 GLoSH: Global-Local Spherical Harmonics for Intrinsic Image Decomposition
abstract
Traditional intrinsic image decomposition focuses on decomposing images into reflectance and shading, leaving surfaces normals and lighting entangled in shading. In this work, we propose a Global-Local Spherical Harmonics (GLoSH) lighting model to improve the lighting component, and jointly predict reflectance and surface normals. The global SH models the holistic lighting while local SH account for the spatial variation of lighting. Also, a novel non-negative lighting constraint is proposed to encourage the estimated SH to be physically meaningful. To seamlessly reflect the GLoSH model, we design a coarse-to-fine network structure. The coarse network predicts global SH, reflectance and normals, and the fine network predicts their local residuals. Lacking labels for reflectance and lighting, we apply synthetic data for model pre-training and fine-tune the model with real data in a self-supervised way. Compared to the state-of-the-art methods only targeting normals or reflectance and shading, our method recovers all components and achieves consistently better results on three real datasets, IIW, SAW and NYUv2.
Hao Zhou 0011, David Jacobs 0001
ICCV3
2019 The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
abstract
We study the relationship between the frequency of a function and the speed at which a neural network learns it. We build on recent results that show that the dynamics of overparameterized neural networks trained with gradient descent can be well approximated by a linear system. When normalized training data is uniformly distributed on a hypersphere, the eigenfunctions of this linear system are spherical harmonic functions. We derive the corresponding eigenvalues for each frequency after introducing a bias term in the model. This bias term had been omitted from the linear network model without significantly affecting previous theoretical results. However, we show theoretically and experimentally that a shallow neural network without bias cannot represent or learn simple, low frequency functions with odd frequencies. Our results lead to specific predictions of the time it will take a network to learn functions of varying frequency. These predictions match the empirical behavior of both shallow and deep networks.
Ronen Basri, David Jacobs 0001, Yoni Kasten, Shira Kritchman
NeurIPS2
2018 End-to-End Recovery of Human Shape and Pose
abstract
We describe Human Mesh Recovery (HMR), an end-to-end framework for reconstructing a full 3D mesh of a human body from a single RGB image. In contrast to most current methods that compute 2D or 3D joint locations, we produce a richer and more useful mesh representation that is parameterized by shape and 3D joint angles. The main objective is to minimize the reprojection loss of keypoints, which allows our model to be trained using in-the-wild images that only have ground truth 2D annotations. However, the reprojection loss alone is highly underconstrained. In this work we address this problem by introducing an adversary trained to tell whether human body shape and pose parameters are real or not using a large database of 3D human meshes. We show that HMR can be trained with and without using any paired 2D-to-3D supervision. We do not rely on intermediate 2D keypoint detections and infer 3D pose and shape parameters directly from image pixels. Our model runs in real-time given a bounding box containing the person. We demonstrate our approach on various images in-the-wild and out-perform previous optimization-based methods that output 3D meshes and show competitive results on tasks such as 3D joint location estimation and part segmentation.
Angjoo Kanazawa, Michael J. Black, David Jacobs 0001, Jitendra Malik
CVPR3
2018 SfSNet: Learning Shape, Reflectance and Illuminance of Faces 'in the Wild'
abstract
We present SfSNet, an end-to-end learning framework for producing an accurate decomposition of an unconstrained human face image into shape, reflectance and illuminance. SfSNet is designed to reflect a physical lambertian rendering model. SfSNet learns from a mixture of labeled synthetic and unlabeled real world images. This allows the network to capture low frequency variations from synthetic and high frequency details from real images through the photometric reconstruction loss. SfSNet consists of a new decomposition architecture with residual blocks that learns a complete separation of albedo and normal. This is used along with the original image to predict lighting. SfSNet produces significantly better quantitative and qualitative results than state-of-the-art methods for inverse rendering and independent normal and illumination estimation.
Roni Sengupta, Angjoo Kanazawa, Carlos Domingo Castillo, David Jacobs 0001
CVPR4
2018 Label Denoising Adversarial Network (LDAN) for Inverse Lighting of Faces
abstract
Lighting estimation from faces is an important task and has applications in many areas such as image editing, intrinsic image decomposition, and image forgery detection. We propose to train a deep Convolutional Neural Network (CNN) to regress lighting parameters from a single face image. Lacking massive ground truth lighting labels for face images in the wild, we use an existing method to estimate lighting parameters, which are treated as ground truth with noise. To alleviate the effect of such noise, we utilize the idea of Generative Adversarial Networks (GAN) and propose a Label Denoising Adversarial Network (LDAN). LDAN makes use of synthetic data with accurate ground truth to help train a deep CNN for lighting regression on real face images. Experiments show that our network outperforms existing methods in producing consistent lighting parameters of different faces under similar lighting conditions. To further evaluate the proposed method, we also apply it to regress object 2D key points where ground truth labels are available. Our experiments demonstrate its effectiveness on this application.
Hao Zhou 0011, Jin Sun 0011, Yaser Yacoob, David Jacobs 0001
CVPR4
2018 Stabilizing Adversarial Nets with Prediction Methods
Abhay Kumar Yadav, Sohil Shah, Zheng Xu 0002, David Jacobs 0001, Tom Goldstein
ICLR (Poster)4
2017 Automated Inference with Adaptive Batches
abstract
Classical stochastic gradient methods for optimization rely on noisy gradient approximations that become progressively less accurate as iterates approach a solution. The large noise and small signal in the resulting gradients makes it difficult to use them for adaptive stepsize selection and automatic stopping. We propose alternative “big batch” SGD schemes that adaptively grow the batch size over time to maintain a nearly constant signal-to-noise ratio in the gradient approximation. The resulting methods have similar convergence rates to classical SGD, and do not require convexity of the objective. The high fidelity gradients enable automated learning rate selection and do not require stepsize decay. Big batch methods are thus easily automated and can run with little or no oversight.
Soham De, Abhay Kumar Yadav, David Jacobs 0001, Tom Goldstein
AISTATS3
2017 A New Rank Constraint on Multi-view Fundamental Matrices, and Its Application to Camera Location Recovery
abstract
Accurate estimation of camera matrices is an important step in structure from motion algorithms. In this paper we introduce a novel rank constraint on collections of fundamental matrices in multi-view settings. We show that in general, with the selection of proper scale factors, a matrix formed by stacking fundamental matrices between pairs of images has rank 6. Moreover, this matrix forms the symmetric part of a rank 3 matrix whose factors relate directly to the corresponding camera matrices. We use this new characterization to produce better estimations of fundamental matrices by optimizing an L1-cost function using Iterative Re-weighted Least Squares and Alternate Direction Method of Multiplier. We further show that this procedure can improve the recovery of camera locations, particularly in multi-view settings in which fewer images are available.
Roni Sengupta, Tal Amir, Meirav Galun, Tom Goldstein, David Jacobs 0001, Amit Singer, Ronen Basri
CVPR5
2017 Seeing What is Not There: Learning Context to Determine Where Objects are Missing
abstract
Most of computer vision focuses on what is in an image. We propose to train a standalone object-centric context representation to perform the opposite task: seeing what is not there. Given an image, our context model can predict where objects should exist, even when no object instances are present. Combined with object detection results, we can perform a novel vision task: finding where objects are missing in an image. Our model is based on a convolutional neural network structure. With a specially designed training strategy, the model learns to ignore objects and focus on context only. It is fully convolutional thus highly efficient. Experiments show the effectiveness of the proposed approach in one important accessibility task: finding city street regions where curb ramps are missing, which could help millions of people with mobility disabilities.
Jin Sun 0011, David Jacobs 0001
CVPR2
2017 3D Menagerie: Modeling the 3D Shape and Pose of Animals
abstract
There has been significant work on learning realistic, articulated, 3D models of the human body. In contrast, there are few such models of animals, despite many applications. The main challenge is that animals are much less cooperative than humans. The best human body models are learned from thousands of 3D scans of people in specific poses, which is infeasible with live animals. Consequently, we learn our model from a small set of 3D scans of toy figurines in arbitrary poses. We employ a novel part-based shape model to compute an initial registration to the scans. We then normalize their pose, learn a statistical shape model, and refine the registrations and the model together. In this way, we accurately align animal scans from different quadruped families with very different shapes and poses. With the registration to a common template we learn a shape space representing animals including lions, cats, dogs, horses, cows and hippos. Animal shapes can be sampled from the model, posed, animated, and fit to data. We demonstrate generalization by fitting it to images of real animals including species not seen in training.
Silvia Zuffi, Angjoo Kanazawa, David Jacobs 0001, Michael J. Black
CVPR3
2017 Efficient Representation of Low-Dimensional Manifolds using Deep Networks
Ronen Basri, David Jacobs 0001
ICLR (Poster)2
2016 WarpNet: Weakly Supervised Matching for Single-View Reconstruction
abstract
We present an approach to matching images of objects in fine-grained datasets without using part annotations, with an application to the challenging problem of weakly supervised single-view reconstruction. This is in contrast to prior works that require part annotations, since matching objects across class and pose variations is challenging with appearance features alone. We overcome this challenge through a novel deep learning architecture, WarpNet, that aligns an object in one image with a different object in another. We exploit the structure of the fine-grained dataset to create artificial data for training this network in an unsupervised-discriminative learning approach. The output of the network acts as a spatial prior that allows generalization at test time to match real images across variations in appearance, viewpoint and articulation. On the CUB-200-2011 dataset of bird categories, we improve the AP over an appearance-only network by 13.6%. We further demonstrate that our WarpNet matches, together with the structure of fine-grained datasets, allow single-view reconstructions with quality comparable to using annotated point correspondences.
Angjoo Kanazawa, David Jacobs 0001, Manmohan Krishna Chandraker
CVPR2
2016 Biconvex Relaxation for Semidefinite Programming in Computer Vision
Sohil Shah, Abhay Kumar Yadav, Carlos Domingo Castillo, David Jacobs 0001, Christoph Studer, Tom Goldstein
ECCV (6)4
2016 Frontal to profile face verification in the wild
abstract
We have collected a new face data set that will facilitate research in the problem of frontal to profile face verification `in the wild'. The aim of this data set is to isolate the factor of pose variation in terms of extreme poses like profile, where many features are occluded, along with other `in the wild' variations. We call this data set the Celebrities in Frontal-Profile (CFP) data set. We find that human performance on Frontal-Profile verification in this data set is only slightly worse (94.57% accuracy) than that on Frontal-Frontal verification (96.24% accuracy). However we evaluated many state-of-the-art algorithms, including Fisher Vector, Sub-SML and a Deep learning algorithm. We observe that all of them degrade more than 10% from Frontal-Frontal to Frontal-Profile verification. The Deep learning implementation, which performs comparable to humans on Frontal-Frontal, performs significantly worse (84.91% accuracy) on Frontal-Profile. This suggests that there is a gap between human performance and automatic face recognition methods for large pose variation in unconstrained images.
Roni Sengupta, Jun-Cheng Chen, Carlos Domingo Castillo, Vishal M. Patel, Rama Chellappa, David Jacobs 0001
WACV6
2016 Learning 3D Deformation of Animals from 2D Images
abstract
Abstract Understanding how an animal can deform and articulate is essential for a realistic modification of its 3D model. In this paper, we show that such information can be learned from user‐clicked 2D images and a template 3D model of the target animal. We present a volumetric deformation framework that produces a set of new 3D models by deforming a template 3D model according to a set of user‐clicked images. Our framework is based on a novel locally‐bounded deformation energy, where every local region has its own stiffness value that bounds how much distortion is allowed at that location. We jointly learn the local stiffness bounds as we deform the template 3D mesh to match each user‐clicked image. We show that this seemingly complex task can be solved as a sequence of convex optimization problems. We demonstrate the effectiveness of our approach on cats and horses, which are highly deformable and articulated animals. Our framework produces new 3D models of animals that are significantly more plausible than methods without learned stiffness.
Angjoo Kanazawa, Shahar Z. Kovalsky, Ronen Basri, David Jacobs 0001
Comput. Graph. Forum4
2015 An Efficient Algorithm for Learning Distances that Obey the Triangle Inequality
Arijit Biswas, David Jacobs 0001
BMVC2
2015 Deep hierarchical parsing for semantic segmentation
abstract
This paper proposes a learning-based approach to scene parsing inspired by the deep Recursive Context Propagation Network (RCPN). RCPN is a deep feed-forward neural network that utilizes the contextual information from the entire image, through bottom-up followed by top-down context propagation via random binary parse trees. This improves the feature representation of every super-pixel in the image for better classification into semantic categories. We analyze RCPN and propose two novel contributions to further improve the model. We first analyze the learning of RCPN parameters and discover the presence of bypass error paths in the computation graph of RCPN that can hinder contextual propagation. We propose to tackle this problem by including the classification loss of the internal nodes of the random parse trees in the original RCPN loss function. Secondly, we use an MRF on the parse tree nodes to model the hierarchical dependency present in the output. Both modifications provide performance boosts over the original RCPN and the new system achieves state-of-the-art performance on Stanford Background, SIFT-Flow and Daimler urban datasets.
Abhishek Sharma 0001, Oncel Tuzel, David Jacobs 0001
CVPR3
2015 From Shading to Local Shape
abstract
We develop a framework for extracting a concise representation of the shape information available from diffuse shading in a small image patch. This produces a mid-level scene descriptor, comprised of local shape distributions that are inferred separately at every image patch across multiple scales. The framework is based on a quadratic representation of local shape that, in the absence of noise, has guarantees on recovering accurate local shape and lighting. And when noise is present, the inferred local shape distributions provide useful shape information without over-committing to any particular image explanation. These local shape distributions naturally encode the fact that some smooth diffuse regions are more informative than others, and they enable efficient and robust reconstruction of object-scale shape. Experimental results show that this approach to surface reconstruction compares well against the state-of-art on both synthetic images and captured photographs.
Ayan Chakrabarti, Ronen Basri, Steven J. Gortler, David Jacobs 0001, Todd E. Zickler
IEEE Trans. Pattern Anal. Mach. Intell.5
2014 Birdsnap: Large-Scale Fine-Grained Visual Categorization of Birds
abstract
We address the problem of large-scale fine-grained visual categorization, describing new methods we have used to produce an online field guide to 500 North American bird species. We focus on the challenges raised when such a system is asked to distinguish between highly similar species of birds. First, we introduce "one-vs-most classifiers." By eliminating highly similar species during training, these classifiers achieve more accurate and intuitive results than common one-vs-all classifiers. Second, we show how to estimate spatio-temporal class priors from observations that are sampled at irregular and biased locations. We show how these priors can be used to significantly improve performance. We then show state-of-the-art recognition performance on a new, large dataset that we make publicly available. These recognition methods are integrated into the online field guide, which is also publicly available.
Thomas Berg, Jiongxin Liu, Michelle L. Alexander, David Jacobs 0001, Peter N. Belhumeur
CVPR5
2014 Tohme: detecting curb ramps in google street view using crowdsourcing, computer vision, and machine learning
abstract
Building on recent prior work that combines Google Street View (GSV) and crowdsourcing to remotely collect information on physical world accessibility, we present the first 'smart' system, Tohme, that combines machine learning, computer vision (CV), and custom crowd interfaces to find curb ramps remotely in GSV scenes. Tohme consists of two workflows, a human labeling pipeline and a CV pipeline with human verification, which are scheduled dynamically based on predicted performance. Using 1,086 GSV scenes (street intersections) from four North American cities and data from 403 crowd workers, we show that Tohme performs similarly in detecting curb ramps compared to a manual labeling approach alone (F- measure: 84% vs. 86% baseline) but at a 13% reduction in time cost. Our work contributes the first CV-based curb ramp detection system, a custom machine-learning based workflow controller, a validation of GSV as a viable curb ramp data source, and a detailed examination of why curb ramp detection is a hard problem along with steps forward.
Kotaro Hara, Jin Sun 0011, David Jacobs 0001, Jon Froehlich
UIST4
2014 Active subclustering
Arijit Biswas, David Jacobs 0001
Comput. Vis. Image Underst.2
2014 Active Image Clustering with Pairwise Constraints from Humans
Arijit Biswas, David Jacobs 0001
Int. J. Comput. Vis.2
2014 Feature Matching with Bounded Distortion
abstract
We consider the problem of finding a geometrically consistent set of point matches between two images. We assume that local descriptors have provided a set of candidate matches, which may include many outliers. We then seek the largest subset of these correspondences that can be aligned perfectly using a nonrigid deformation that exerts a bounded distortion. We formulate this as a constrained optimization problem and solve it using a constrained, iterative reweighted least-squares algorithm. In each iteration of this algorithm we solve a convex quadratic program obtaining a globally optimal match over a subset of the bounded distortion transformations. We further prove that a sequence of such iterations converges monotonically to a critical point of our objective function. We show experimentally that this algorithm produces excellent results on a number of test sets, in comparison to several state-of-the-art approaches.
Yaron Lipman, Stav Yagev, Roi Poranne, David Jacobs 0001, Ronen Basri
ACM Trans. Graph.4
2013 Efficient segmentation of leaves in semi-controlled conditions
João V. B. Soares, David Jacobs 0001
Mach. Vis. Appl.2
2013 Localizing Parts of Faces Using a Consensus of Exemplars
abstract
We present a novel approach to localizing parts in images of human faces. The approach combines the output of local detectors with a nonparametric set of global models for the part locations based on over 1,000 hand-labeled exemplar images. By assuming that the global models generate the part locations as hidden variables, we derive a Bayesian objective function. This function is optimized using a consensus of models for these hidden variables. The resulting localizer handles a much wider range of expression, pose, lighting, and occlusion than prior ones. We show excellent performance on real-world face datasets such as Labeled Faces in the Wild (LFW) and a new Labeled Face Parts in the Wild (LFPW) and show that our localizer achieves state-of-the-art performance on the less challenging BioID dataset.
Peter N. Belhumeur, David Jacobs 0001, David J. Kriegman, Neeraj Kumar 0006
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Dynamic changes in motivation in collaborative citizen-science projects
abstract
Online citizen science projects engage volunteers in collecting, analyzing, and curating scientific data. Existing projects have demonstrated the value of using volunteers to collect data, but few projects have reached the full collaborative potential of scientists and volunteers. Understanding the shared and unique motivations of these two groups can help designers establish the technical and social infrastructures needed to promote effective partnerships. We present findings from a study of the motivational factors affecting participation in ecological citizen science projects. We show that volunteers are motivated by a complex framework of factors that dynamically change throughout their cycle of work on scientific projects; this motivational framework is strongly affected by personal interests as well as external factors such as attribution and acknowledgment. Identifying the pivotal points of motivational shift and addressing them in the design of citizen-science systems will facilitate improved collaboration between scientists and volunteers.
Dana Rotman, Jennifer Preece, Jennifer Hammock, Kezee Procita, Derek L. Hansen, Cynthia Sims Parr, Darcy Lewis, David Jacobs 0001
CSCW8
2012 Active image clustering: Seeking constraints from humans to complement algorithms
abstract
We propose a method of clustering images that combines algorithmic and human input. An algorithm provides us with pairwise image similarities. We then actively obtain selected, more accurate pairwise similarities from humans. A novel method is developed to choose the most useful pairs to show a person, obtaining constraints that improve clustering. In a clustering assignment elements in each data pair are either in the same cluster or in different clusters. We simulate inverting these pairwise relations and see how that affects the overall clustering. We choose a pair that maximizes the expected change in the clustering. The proposed algorithm has high time complexity, so we also propose a version of this algorithm that is much faster and exactly replicates our original algorithm. We further improve run-time by adding heuristics, and show that these do not significantly impact the effectiveness of our method. We have run experiments in two different domains, namely leaf images and face images, and show that clustering performance can be improved significantly.
Arijit Biswas, David Jacobs 0001
CVPR2
2012 Generalized Multiview Analysis: A discriminative latent space
abstract
This paper presents a general multi-view feature extraction approach that we call Generalized Multiview Analysis or GMA. GMA has all the desirable properties required for cross-view classification and retrieval: it is supervised, it allows generalization to unseen classes, it is multi-view and kernelizable, it affords an efficient eigenvalue based solution and is applicable to any domain. GMA exploits the fact that most popular supervised and unsupervised feature extraction techniques are the solution of a special form of a quadratic constrained quadratic program (QCQP), which can be solved efficiently as a generalized eigenvalue problem. GMA solves a joint, relaxed QCQP over different feature spaces to obtain a single (non)linear subspace. Intuitively, GMA is a supervised extension of Canonical Correlational Analysis (CCA), which is useful for cross-view classification and retrieval. The proposed approach is general and has the potential to replace CCA whenever classification or retrieval is the purpose and label information is available. We outperform previous approaches for textimage retrieval on Pascal and Wiki text-image data. We report state-of-the-art results for pose and lighting invariant face recognition on the MultiPIE face dataset, significantly outperforming other approaches.
Abhishek Sharma 0001, Abhishek Kumar 0001, Hal Daumé III, David Jacobs 0001
CVPR4
2012 A Fast Illumination and Deformation Insensitive Image Comparison Algorithm Using Wavelet-Based Geodesics
Anne Jorstad, David Jacobs 0001, Alain Trouvé
ECCV (4)2
2012 Leafsnap: A Computer Vision System for Automatic Plant Species Identification
Neeraj Kumar 0006, Peter N. Belhumeur, Arijit Biswas, David Jacobs 0001, W. John Kress, Ida C. Lopez, João V. B. Soares
ECCV (2)4
2012 Dog Breed Classification Using Part Localization
Jiongxin Liu, Angjoo Kanazawa, David Jacobs 0001, Peter N. Belhumeur
ECCV (1)3
2012 Robust pose invariant face recognition using coupled latent space discriminant analysis
Abhishek Sharma 0001, Murad Al Haj, Larry Davis 0001, David Jacobs 0001
Comput. Vis. Image Underst.5
2011 Localizing parts of faces using a consensus of exemplars
abstract
We present a novel approach to localizing parts in images of human faces. The approach combines the output of local detectors with a non-parametric set of global models for the part locations based on over one thousand hand-labeled exemplar images. By assuming that the global models generate the part locations as hidden variables, we derive a Bayesian objective function. This function is optimized using a consensus of models for these hidden variables. The resulting localizer handles a much wider range of expression, pose, lighting and occlusion than prior ones. We show excellent performance on a new dataset gathered from the internet and show that our localizer achieves state-of-the-art performance on the less challenging BioID dataset.
Peter N. Belhumeur, David Jacobs 0001, David J. Kriegman, Neeraj Kumar 0006
CVPR2
2011 Wide-baseline stereo for face recognition with large pose variation
abstract
2-D face recognition in the presence of large pose variations presents a significant challenge. When comparing a frontal image of a face to a near profile image, one must cope with large occlusions, non-linear correspondences, and significant changes in appearance due to viewpoint. Stereo matching has been used to handle these problems, but performance of this approach degrades with large pose changes. We show that some of this difficulty is due to the effect that foreshortening of slanted surfaces has on window-based matching methods, which are needed to provide robustness to lighting change. We address this problem by designing a new, dynamic programming stereo algorithm that accounts for surface slant. We show that on the CMU PIE dataset this method results in significant improvements in recognition performance.
Carlos Domingo Castillo, David Jacobs 0001
CVPR2
2011 A deformation and lighting insensitive metric for face recognition based on dense correspondences
abstract
Face recognition is a challenging problem, complicated by variations in pose, expression, lighting, and the passage of time. Significant work has been done to solve each of these problems separately. We consider the problems of lighting and expression variation together, proposing a method that accounts for both variabilities within a single model. We present a novel deformation and lighting insensitive metric to compare images, and we present a novel framework to optimize over this metric to calculate dense correspondences between images. Typical correspondence cost patterns are learned between face image pairs and a Naïve Bayes classifier is applied to improve recognition accuracy. Very promising results are presented on the AR Face Database, and we note that our method can be extended to a broad set of applications.
Anne Jorstad, David Jacobs 0001, Alain Trouvé
CVPR2
2011 Bypassing synthesis: PLS for face recognition with pose, low-resolution and sketch
abstract
This paper presents a novel way to perform multi-modal face recognition. We use Partial Least Squares (PLS) to linearly map images in different modalities to a common linear subspace in which they are highly correlated. PLS has been previously used effectively for feature selection in face recognition. We show both theoretically and experimentally that PLS can be used effectively across modalities. We also formulate a generic intermediate subspace comparison framework for multi-modal recognition. Surprisingly, we achieve high performance using only pixel intensities as features. We experimentally demonstrate the highest published recognition rates on the pose variations in the PIE data set, and also show that PLS can be used to compare sketches to photos, and to compare images taken at different resolutions.
Abhishek Sharma 0001, David Jacobs 0001
CVPR2
2011 Dynamic Processing Allocation in Video
abstract
Large stores of digital video pose severe computational challenges to existing video analysis algorithms. In applying these algorithms, users must often trade off processing speed for accuracy, as many sophisticated and effective algorithms require large computational resources that make it impractical to apply them throughout long videos. One can save considerable effort by applying these expensive algorithms sparingly, directing their application using the results of more limited processing. We show how to do this for retrospective video analysis by modeling a video using a chain graphical model and performing inference both to analyze the video and to direct processing. We apply our method to problems in background subtraction and face detection, and show in experiments that this leads to significant improvements over baseline algorithms.
Daozheng Chen, Mustafa Bilgic 0001, Lise Getoor, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 Illumination Recovery From Image With Cast Shadows Via Sparse Representation
abstract
In this paper, we propose using sparse representation for recovering the illumination of a scene from a single image with cast shadows, given the geometry of the scene. The images with cast shadows can be quite complex and, therefore, cannot be well approximated by low-dimensional linear subspaces. However, it can be shown that the set of images produced by a Lambertian scene with cast shadows can be efficiently represented by a sparse set of images generated by directional light sources. We first model an image with cast shadows composed of a diffusive part (without cast shadows) and a residual part that captures cast shadows. Then, we express the problem in an l(1)-regularized least-squares formulation, with nonnegativity constraints (as light has to be non-negative at any point in space). This sparse representation enjoys an effective and fast solution thanks to recent advances in compressive sensing. In experiments on synthetic and real data, our approach performs favorably in comparison with several previously proposed methods.
Xue Mei, Haibin Ling, David Jacobs 0001
IEEE Trans. Image Process.3
2010 Comparing and combining lighting insensitive approaches for face recognition
Raghuraman Gopalan, David Jacobs 0001
Comput. Vis. Image Underst.2
2010 Mesh saliency and human eye fixations
abstract
Mesh saliency has been proposed as a computational model of perceptual importance for meshes, and it has been used in graphics for abstraction, simplification, segmentation, illumination, rendering, and illustration. Even though this technique is inspired by models of low-level human vision, it has not yet been validated with respect to human performance. Here, we present a user study that compares the previous mesh saliency approaches with human eye movements. To quantify the correlation between mesh saliency and fixation locations for 3D rendered images, we introduce the normalized chance-adjusted saliency by improving the previous chance-adjusted saliency measure. Our results show that the current computational model of mesh saliency can model human eye movements significantly better than a purely random model or a curvature-based model.
Amitabh Varshney, David Jacobs 0001, François Guimbretière
ACM Trans. Appl. Percept.3
2010 Face verification across age progression using discriminative methods
abstract
Face verification in the presence of age progression is an important problem that has not been widely addressed. In this paper, we study the problem by designing and evaluating discriminative approaches. These directly tackle verification tasks without explicit age modeling, which is a hard problem by itself. First, we find that the gradient orientation, after discarding magnitude information, provides a simple but effective representation for this problem. This representation is further improved when hierarchical information is used, which results in the use of the gradient orientation pyramid (GOP). When combined with a support vector machine GOP demonstrates excellent performance in all our experiments, in comparison with seven different approaches including two commercial systems. Our experiments are conducted on the FGnet dataset and two large passport datasets, one of them being the largest ever reported for recognition tasks. Second, taking advantage of these datasets, we empirically study how age gaps and related issues (including image quality, spectacles, and facial hair) affect recognition algorithms. We found surprisingly that the added difficulty of verification produced by age gaps becomes saturated after the gap is larger than four years, for gaps of up to ten years. In addition, we find that image quality and eyewear present more of a challenge than facial hair.
Haibin Ling, Stefano Soatto, Narayanan Ramanathan, David Jacobs 0001
IEEE Trans. Inf. Forensics Secur.4
2009 Visibility constraints on features of 3D objects
abstract
To recognize three-dimensional objects it is important to model how their appearances can change due to changes in viewpoint. A key aspect of this involves understanding which object features can be simultaneously visible under different viewpoints. We address this problem in an image-based framework, in which we use a limited number of images of an object taken from unknown viewpoints to determine which subsets of features might be simultaneously visible in other views. This leads to the problem of determining whether a set of images, each containing a set of features, is consistent with a single 3D object. We assume that each feature is visible from a disk of viewpoints on the viewing sphere. In this case we show the problem is NP-hard in general, but can be solved efficiently when all views come from a circle on the viewing sphere. We also give iterative algorithms that can handle noisy data and converge to locally optimal solutions in the general case. Our techniques can also be used to recover viewpoint information from the set of features that are visible in different images. We show that these algorithms perform well both on synthetic data and images from the COIL dataset.
Ronen Basri, Pedro F. Felzenszwalb, Ross B. Girshick, David Jacobs 0001, Caroline J. Klivans
CVPR4
2009 Sparse representation of cast shadows via l1-regularized least squares
abstract
Scenes with cast shadows can produce complex sets of images. These images cannot be well approximated by low-dimensional linear subspaces. However, in this paper we show that the set of images produced by a Lambertian scene with cast shadows can be efficiently represented by a sparse set of images generated by directional light sources. We first model an image with cast shadows as composed of a diffusive part (without cast shadows) and a residual part that captures cast shadows. Then, we express the problem in an ℓ1-regularized least squares formulation, with nonnegativity constraints. This sparse representation enjoys an effective and fast solution, thanks to recent advances in compressive sensing. In experiments on both synthetic and real data, our approach performs favorably in comparison to several previously proposed methods.
Xue Mei, Haibin Ling, David Jacobs 0001
ICCV3
2009 Assigning cameras to subjects in video surveillance systems
abstract
We consider the problem of tracking multiple agents moving amongst obstacles, using multiple cameras. Given an environment with obstacles, and many people moving through it, we construct a separate narrow field of view video for as many people as possible, by stitching together video segments from multiple cameras over time. We employ a novel approach to assign cameras to people as a function of time, with camera switches when needed. The problem is modeled as a bipartite graph and the solution corresponds to a maximum matching. As people move, the solution is efficiently updated by computing an augmenting path rather than by solving for a new matching. This reduces computation time by an order of magnitude. In addition, solving for the shortest augmenting path minimizes the number of camera switches at each update. When not all people can be covered by the available cameras, we cluster as many people as possible into small groups, then assign cameras to groups using a minimum cost matching algorithm. We test our method using numerous runs from different simulators.
Hazem El-Alfy, David Jacobs 0001, Larry Davis 0001
ICRA2
2009 Using Stereo Matching with General Epipolar Geometry for 2D Face Recognition across Pose
abstract
Face recognition across pose is a problem of fundamental importance in computer vision. We propose to address this problem by using stereo matching to judge the similarity of two, 2D images of faces seen from different poses. Stereo matching allows for arbitrary, physically valid, continuous correspondences. We show that the stereo matching cost provides a very robust measure of similarity of faces that is insensitive to pose variations. To enable this, we show that, for conditions common in face recognition, the epipolar geometry of face images can be computed using either four or three feature points. We also provide a straightforward adaptation of a stereo matching algorithm to compute the similarity between faces. The proposed approach has been tested on the CMU PIE data set and demonstrates superior performance compared to existing methods in the presence of pose variation. It also shows robustness to lighting variation.
Carlos Domingo Castillo, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Approximate earth mover's distance in linear time
abstract
The earth moverpsilas distance (EMD) is an important perceptually meaningful metric for comparing histograms, but it suffers from high (O(N3logN)) computational complexity. We present a novel linear time algorithm for approximating the EMD for low dimensional histograms using the sum of absolute values of the weighted wavelet coefficients of the difference histogram. EMD computation is a special case of the Kantorovich-Rubinstein transshipment problem, and we exploit the Holder continuity constraint in its dual form to convert it into a simple optimization problem with an explicit solution in the wavelet domain. We prove that the resulting wavelet EMD metric is equivalent to EMD, i.e. the ratio of the two is bounded. We also provide estimates for the bounds. The weighted wavelet transform can be computed in time linear in the number of histogram bins, while the comparison is about as fast as for normal Euclidean distance or chi2statistic. We experimentally show that wavelet EMD is a good approximation to EMD, has similar performance, but requires much less computation.
Sameer Shirdhonkar, David Jacobs 0001
CVPR2
2008 Searching the World's Herbaria: A System for Visual Identification of Plant Species
Peter N. Belhumeur, Daozheng Chen, Steven K. Feiner, David Jacobs 0001, W. John Kress, Haibin Ling, Ida C. Lopez, Ravi Ramamoorthi, Sameer Sheorey, Sean White
ECCV (4)4
2008 Tracking Down Under: Following the Satin Bowerbird
abstract
Socio biologists collect huge volumes of video to study animal behavior (our collaborators work with 30,000 hours of video). The scale of these datasets demands the development of automated video analysis tools. Detecting and tracking animals is a critical first step in this process. However, off-the-shelf methods prove incapable of handling videos characterized by poor quality, drastic illumination changes, non-stationary scenery and foreground objects that become motionless for long stretches of time. We improve on existing approaches by taking advantage of specific aspects of this problem: by using information from the entire video we are able to find animals that become motionless for long intervals of time; we make robust decisions based on regional features; for different parts of the image, we tailor the selection of model features, choosing the features most helpful in differentiating the target animal from the background in that part of the image. We evaluate our method, achieving almost 83% tracking accuracy on a more than 200,000 frame dataset of Satin Bowerbird courtship videos.
Aniruddha Kembhavi, Ryan Farrell, Yuancheng Luo, David Jacobs 0001, Ramani Duraiswami, Larry Davis 0001
WACV4
2008 Using specularities in comparing 3D models and 2D images
Margarita Osadchy, David Jacobs 0001, Ravi Ramamoorthi, David Tucker
Comput. Vis. Image Underst.2
2007 Using Stereo Matching for 2-D Face Recognition Across Pose
abstract
We propose using stereo matching for 2-D face recognition across pose. We match one 2-D query image to one 2-D gallery image without performing 3-D reconstruction. Then the cost of this matching is used to evaluate the similarity of the two images. We show that this cost is robust to pose variations. To illustrate this idea we built a face recognition system on top of a dynamic programming stereo matching algorithm. The method works well even when the epipolar lines we use do not exactly fit the viewpoints. We have tested our approach on the PIE dataset. In all the experiments, our method demonstrates effective performance compared with other algorithms.
Carlos Domingo Castillo, David Jacobs 0001
CVPR2
2007 Efficiently Determining Silhouette Consistency
abstract
Volume intersection is a frequently used technique to solve the Shape-From-Silhouette problem, which constructs a 3D object estimate from a set of silhouettes taken with calibrated cameras. It is natural to develop an efficient algorithm to determine the consistency of a set of silhouettes before performing time-consuming reconstruction, so that inaccurate silhouettes can be omitted. In this paper we first present a fast algorithm to determine the consistency of three silhouettes from known (but arbitrary) viewing directions, assuming the projection is scaled orthographic. The temporal complexity of the algorithm is linear in the number of points of the silhouette boundaries. We further prove that a set of more than three convex silhouettes are consistent if and only if any three of them are consistent. Another possible application of our approach is to determine the miscalibrated cameras in a large camera system. A consistent subset of cameras can be determined on the fly and miscalibrated cameras can also be recalibrated at a coarse scale. Real and synthesized data are used to demonstrate our results.
David Jacobs 0001
CVPR2
2007 A Study of Face Recognition as People Age
abstract
In this paper we study face recognition across ages within a real passport photo verification task. First, we propose using the gradient orientation pyramid for this task. Discarding the gradient magnitude and utilizing hierarchical techniques, we found that the new descriptor yields a robust and discriminative representation. With the proposed descriptor, we model face verification as a two-class problem and use a support vector machine as a classifier. The approach is applied to two passport data sets containing more than 1,800 image pairs from each person with large age differences. Although simple, our approach outperforms previously tested Bayesian technique and other descriptors, including the intensity difference and gradient with magnitude. In addition, it works as well as two commercial systems. Second, for the first time, we empirically study how age differences affect recognition performance. Our experiments show that, although the aging process adds difficulty to the recognition task, it does not surpass illumination or expression as a confounding factor.
Haibin Ling, Stefano Soatto, Narayanan Ramanathan, David Jacobs 0001
ICCV4
2007 Multi-scale video cropping
abstract
We consider the problem of cropping surveillance videos. This process chooses a trajectory that a small sub-window can take through the video, selecting the most important parts of the video for display on a smaller monitor. We model the information content of the video simply, by whether the image changes at each pixel. Then we show that we can find the globally optimal trajectory for a cropping window by using a shortest path algorithm. In practice, we can speed up this process without affecting the results, by stitching together trajectories computed over short intervals. This also reduces system latency. We then show that we can use a second shortest path formulation to find good cuts from one trajectory to another, improving coverage of interesting events in the video. We describe additional techniques to improve the quality and efficiency of the algorithm, and show results on surveillance videos.
Hazem El-Alfy, David Jacobs 0001, Larry Davis 0001
ACM Multimedia2
2007 Photometric Stereo with General, Unknown Lighting
Ronen Basri, David Jacobs 0001, Ira Kemelmacher-Shlizerman
Int. J. Comput. Vis.2
2007 Shape Classification Using the Inner-Distance
abstract
Part structure and articulation are of fundamental importance in computer and human vision. We propose using the inner-distance to build shape descriptors that are robust to articulation and capture part structure. The inner-distance is defined as the length of the shortest path between landmark points within the shape silhouette. We show that it is articulation insensitive and more effective at capturing part structures than the Euclidean distance. This suggests that the inner-distance can be used as a replacement for the Euclidean distance to build more accurate descriptors for complex shapes, especially for those with articulated parts. In addition, texture information along the shortest path can be used to further improve shape classification. With this idea, we propose three approaches to using the inner-distance. The first method combines the inner-distance and multidimensional scaling (MDS) to build articulation invariant signatures for articulated shapes. The second method uses the inner-distance to build a new shape descriptor based on shape contexts. The third one extends the second one by considering the texture information along shortest paths. The proposed approaches have been tested on a variety of shape databases, including an articulated shape data set, MPEG7 CE-Shape-1, Kimia silhouettes, the ETH-80 data set, two leaf data sets, and a human motion silhouette data set. In all the experiments, our methods demonstrate effective performance compared with other algorithms.
Haibin Ling, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Surface Dependent Representations for Illumination Insensitive Image Comparison
abstract
We consider the problem of matching images to tell whether they come from the same scene viewed under different lighting conditions. We show that the surface characteristics determine the type of image comparison method that should be used. Previous work has shown the effectiveness of comparing the image gradient direction for surfaces with material properties that change rapidly in one direction. We show analytically that two other widely used methods, normalized correlation of small windows and comparison of multiscale oriented filters, essentially compute the same thing. Then, we show that for surfaces whose properties change more slowly, comparison of the output of whitening filters is most effective. This suggests that a combination of these strategies should be employed to compare general objects. We discuss indications that Gabor jets use such a mixed strategy effectively, and we propose a new mixed strategy. We validate our results on synthetic and real images.
Margarita Osadchy, David Jacobs 0001, Michael Lindenbaum
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Appearance Characterization of Linear Lambertian Objects, Generalized Photometric Stereo, and Illumination-Invariant Face Recognition
abstract
Traditional photometric stereo algorithms employ a Lambertian reflectance model with a varying albedo field and involve the appearance of only one object. In this paper, we generalize photometric stereo algorithms to handle all appearances of all objects in a class, in particular the human face class, by making use of the linear Lambertian property. A linear Lambertian object is one which is linearly spanned by a set of basis objects and has a Lambertian surface. The linear property leads to a rank constraint and, consequently, a factorization of an observation matrix that consists of exemplar images of different objects (e.g., faces of different subjects) under different, unknown illuminations. Integrability and symmetry constraints are used to fully recover the subspace bases using a novel linearized algorithm that takes the varying albedo field into account. The effectiveness of the linear Lambertian property is further investigated by using it for the problem of illumination-invariant face recognition using just one image. Attached shadows are incorporated in the model by a careful treatment of the inherent nonlinearity in Lambert's law. This enables us to extend our algorithm to perform face recognition in the presence of multiple illumination sources. Experimental results using standard data sets are presented.
Shaohua Kevin Zhou, Gaurav Aggarwal, Rama Chellappa, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2005 Using the Inner-Distance for Classification of Articulated Shapes
abstract
We propose using the inner-distance between landmark points to build shape descriptors. The inner-distance is defined as the length of the shortest path between landmark points within the shape silhouette. We show that the inner-distance is articulation insensitive and more effective at capturing complex shapes with part structures than Euclidean distance. To demonstrate this idea, it is used to build a new shape descriptor based on shape contexts. After that, we design a dynamic programming based method for shape matching and comparison. We have tested our approach on a variety of shape databases including an articulated shape dataset, MPEG7 CE-Shape-1, Kimia silhouettes, a Swedish leaf database and a human motion silhouette dataset. In all the experiments, our method demonstrates effective performance compared with other algorithms.
Haibin Ling, David Jacobs 0001
CVPR (2)2
2005 Deformation Invariant Image Matching
abstract
We propose a novel framework to build descriptors of local intensity that are invariant to general deformations. In this framework, an image is embedded as a 2D surface in 3D space, with intensity weighted relative to distance in x-y. We show that as this weight increases, geodesic distances on the embedded surface are less affected by image deformations. In the limit, distances are deformation invariant. We use geodesic sampling to get neighborhood samples for interest points, and then use a geodesic-intensity histogram (GIH) as a deformation invariant local descriptor. In addition to its invariance, the new descriptor automatically finds its support region. This means it can safely gather information from a large neighborhood to improve discriminability. Furthermore, we propose a matching method for this descriptor that is invariant to affine lighting changes. We have tested this new descriptor on interest point matching for two data sets, one with synthetic deformation and lighting change, and another with real non-affine deformations. Our method shows promising matching results compared to several other approaches
Haibin Ling, David Jacobs 0001
ICCV2
2005 On the Equivalence of Common Approaches to Lighting Insensitive Recognition
abstract
Lighting variation is commonly handled by methods invariant to additive and multiplicative changes in image intensity. It has been demonstrated that comparing images using the direction of the gradient can produce broader insensitivity to changes in lighting conditions, even for 3D scenes. We analyze two common approaches to image comparison that are invariant, normalized correlation using small correlation windows, and comparison based on a large set of oriented difference of Gaussian filters. We show analytically that these methods calculate a monotonic (cosine) function of the gradient direction difference and hence are equivalent to the direction of gradient method. Our analysis is supported with experiments on both synthetic and real scenes
Margarita Osadchy, David Jacobs 0001, Michael Lindenbaum
ICCV2
2005 Non-Negative Lighting and Specular Object Recognition
abstract
Recognition of specular objects is particularly difficult because their appearance is much more sensitive to lighting changes than that of Lambertian objects. We consider an approach in which we use a 3D model to deduce the lighting that best matches the model to the image. In this case, an important constraint is that incident lighting should be non-negative everywhere. In this paper, we propose a new method to enforce this constraint and explore its usefulness in specular object recognition, using the spherical harmonic representation of lighting. The method follows from a novel extension of Szego's eigenvalue distribution theorem to spherical harmonics, and uses semidefinite programming to perform a constrained optimization. The new method is faster as well as more accurate than previous methods. Experiments on both synthetic and real data indicate that the constraint can improve recognition of specular objects by better separating the correct and incorrect models
Sameer Shirdhonkar, David Jacobs 0001
ICCV2
2005 Mesh saliency
abstract
Research over the last decade has built a solid mathematical foundation for representation and analysis of 3D meshes in graphics and geometric modeling. Much of this work however does not explicitly incorporate models of low-level human visual attention. In this paper we introduce the idea of mesh saliency as a measure of regional importance for graphics meshes. Our notion of saliency is inspired by low-level human visual system cues. We define mesh saliency in a scale-dependent manner using a center-surround operator on Gaussian-weighted mean curvatures. We observe that such a definition of mesh saliency is able to capture what most would classify as visually interesting regions on a mesh. The human-perception-inspired importance measure computed by our mesh saliency operator results in more visually pleasing results in processing and viewing of 3D meshes. compared to using a purely geometric measure of shape. such as curvature. We discuss how mesh saliency can be incorporated in graphics applications such as mesh simplification and viewpoint selection and present examples that show visually appealing results from using mesh saliency.
Chang Ha Lee, Amitabh Varshney, David Jacobs 0001
ACM Trans. Graph.3
2004 Whitening for Photometric Comparison of Smooth Surfaces under Varying Illumination
Margarita Osadchy, Michael Lindenbaum, David Jacobs 0001
ECCV (4)3
2004 Characterization of Human Faces under Illumination Variations Using Rank, Integrability, and Symmetry Constraints
Shaohua Kevin Zhou, Rama Chellappa, David Jacobs 0001
ECCV (1)3
2003 Using Specularities for Recognition
abstract
Recognition systems have generally treated specular highlights as noise. We show how to use these highlights as a positive source of information that improves recognition of shiny objects. This also enables us to recognize very challenging shiny transparent objects, such as wine glasses. Specifically, we show how to find highlights that are consistent with a hypothesized pose of an object of known 3D shape. We do this using only a qualitative description of highlight formation that is consistent with most models of specular reflection, so no specific knowledge of an object's reflectance properties is needed. We first present a method that finds highlights produced by a dominant compact light source, whose position is roughly known. We then show how to estimate the lighting automatically for objects whose reflection is part specular and part Lambertian. We demonstrate this method for two classes of objects. First, we show that specular information alone can suffice to identify objects with no Lambertian reflectance, such as transparent wine glasses. Second, we use our complete system to recognize shiny objects, such as pottery.
Margarita Osadchy, David Jacobs 0001, Ravi Ramamoorthi
ICCV2
2003 Automatic thumbnail cropping and its effectiveness
abstract
Thumbnail images provide users of image retrieval and browsing systems with a method for quickly scanning large numbers of images. Recognizing the objects in an image is important in many retrieval tasks, but thumbnails generated by shrinking the original image often render objects illegible. We study the ability of computer vision systems to detect key components of images so that automated cropping, prior to shrinking, can render objects more recognizable. We evaluate automatic cropping techniques 1) based on a general method that detects salient portions of images, and 2) based on automatic face detection. Our user study shows that these methods result in small thumbnails that are substantially more recognizable and easier to find in the context of visual search.
Bongwon Suh, Haibin Ling, Benjamin B. Bederson, David Jacobs 0001
UIST4
2003 Lambertian Reflectance and Linear Subspaces
abstract
We prove that the set of all Lambertian reflectance functions (the mapping from surface normals to intensities) obtained with arbitrary distant light sources lies close to a 9D linear subspace. This implies that, in general, the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace, explaining prior empirical results. We also provide a simple analytic characterization of this linear space. We obtain these results by representing lighting using spherical harmonics and describing the effects of Lambertian materials as the analog of a convolution. These results allow us to construct algorithms for object recognition based on linear methods as well as algorithms that use convex optimization to enforce nonnegative lighting functions. We also show a simple way to enforce nonnegative lighting when the images of an object lie near a 4D linear space. We apply these algorithms to perform face recognition by finding the 3D model that best matches a 2D query image.
Ronen Basri, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Guest Editors' Introduction to the Special Section on Perceptual Organization in Computer Vision
David Jacobs 0001, Michael Lindenbaum
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Guest Editors' Introduction to the Special Section on Perceptual Organization in Computer Vision
abstract
THIS issue contains the second installment of the special issue on perceptual organization in computer vision. In the April issue of TPAMI, we published 11 papers addressing basic principles, algorithms, and applications. This month’s papers continue along these lines and raising new issues as well. The first paper in the special section contributes to one of the most significant trends in perceptual organization, the use of graph algorithms to efficiently find regions that globally optimize, or approximately optimize, a grouping property. P. Soundararajan and S. Sarkar present an analysis and experimental evaluation of a set of graph algorithms, showing that some of the common algorithms are practically equivalent with respect to performance. In the next paper, J.H. Elder, A. Krupnik, and L.A. Johnston describe a probabilistic formulation for contour grouping that combines Gestalt grouping cues with prior knowledge about the shape of the objects whose boundary is being detected. They apply this approach to the problem of detecting exact lake boundaries in satellite imagery. S. Wang and J.M. Siskind also study graph algorithms as a grouping mechanism. They present a novel algorithm based on the ratio of two properties of graph edges that are cut in a segmentation. This leads to a polynomial time algorithm that finds globally optimal groupings, which they then enhance using a set of heuristics. Finally, it is fitting that our issue concludes with S.-C. Zhu’s paper, which reviews a large set of past work on modeling visual patterns, an area fundamental to perceptual organization. Professor Zhu presents a broad taxonomy of methods, stressing the importance of generative models in image modeling. The 15 papers we have presented provide a wide range of viewpoints and attack a variety of problems using diverse tools. We find this appropriate since the challenges of perceptual organization are great and many angles on the problem should continue to be explored. Hopefully, the special section will provide a useful snapshot of the state of many of these approaches.
David Jacobs 0001, Michael Lindenbaum
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Photometric Stereo with General, Unknown Lighting
abstract
Work on photometric stereo has shown how to recover the shape and reflectance properties of an object using multiple images taken with a fixed viewpoint and variable lighting conditions. This work has primarily relied on the presence of a single point source of light in each image. The authors show how to perform photometric stereo, assuming that all lights in a scene are isotropic and distant from the object but otherwise unconstrained. Lighting in each image may be an unknown and arbitrary combination of diffuse, point and extended sources. Our work is based on recent results showing that for Lambertian objects, general lighting conditions can be represented using low order spherical harmonics. Using this representation, we can recover shape by performing a simple optimization in a low-dimensional space. We also analyze the shape ambiguities that arise in such a representation.
Ronen Basri, David Jacobs 0001
CVPR (2)2
2001 Lambertian Reflectance and Linear Subspaces
abstract
We prove that the set of all reflectance functions (the mapping from surface normals to intensities) produced by Lambertian objects under distant, isotropic lighting lies close to a 9D linear subspace. This implies that the images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately with a low-dimensional linear subspace, explaining prior empirical results. We also provide a simple analytic characterization of this linear space. We obtain these results by representing lighting using spherical harmonics and describing the effects of Lambertian materials as the analog of a convolution. These results allow us to construct algorithms for object recognition based on linear methods as well as algorithms that use convex optimization to enforce non-negative lighting functions.
Ronen Basri, David Jacobs 0001
ICCV2
2001 Fragment Completion in Humans and Machines
abstract
Partial information can trigger a complete memory. At the same time, human memory is not perfect. A cue can contain enough information to specify an item in memory, but fail to trigger that item. In the context of word memory, we present experiments that demonstrate some basic patterns in human memory errors. We use cues that consist of word frag- ments. We show that short and long cues are completed more accurately than medium length ones and study some of the factors that lead to this behavior. We then present a novel computational model that shows some of the flexibility and patterns of errors that occur in human memory. This model iterates between bottom-up and top-down computations. These are tied together using a Markov model of words that allows memory to be accessed with a simple feature set, and enables a bottom-up process to compute a probability distribution of possible completions of word frag- ments, in a manner similar to models of visual perceptual completion.
David Jacobs 0001, Bas Rokers, Archisman Rudra
NIPS1
2001 Linear Fitting with Missing Data for Structure-from-Motion
David Jacobs 0001
Comput. Vis. Image Underst.1
2001 Projective Alignment with Regions
abstract
We have previously proposed (Basri and Jacobs, 1999, and Jacobs and Basri, 1999) an approach to recognition that uses regions to determine the pose of objects while allowing for partial occlusion of the regions. Regions introduce an attractive alternative to existing global and local approaches, since, unlike global features, they can handle occlusion and segmentation errors, and unlike local features they are not as sensitive to sensor errors, and they are easier to match. The region-based approach also uses image information directly, without the construction of intermediate representations, such as algebraic descriptions, which may be difficult to reliably compute. We further analyze properties of the method for planar objects undergoing projective transformations. In particular, we prove that three visible regions are sufficient to determine the transformation uniquely and that for a large class of objects, two regions are insufficient for this purpose. However, we show that when several regions are available, the pose of the object can generally be recovered even when some or all regions are significantly occluded. Our analysis is based on investigating the flow patterns of points under projective transformations in the presence of fixed points.
Ronen Basri, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 In Search of Illumination Invariants
abstract
We consider the problem of determining functions of an image of an object that are insensitive to illumination changes. We first show that for an object with Lambertian reflectance there are no discriminative functions that are invariant to illumination. This result leads as to adopt a probabilistic approach in which we analytically determine a probability distribution for the image gradient as a function of the surface's geometry and reflectance. Our distribution reveals that the direction of the image gradient is insensitive to changes in illumination direction. We verify this empirically by constructing a distribution for the image gradient from more than 20 million samples of gradients in a database of 1,280 images of 20 inanimate objects taken under varying lighting condition. Using this distribution we develop an illumination insensitive measure of image comparison and test it on the problem of face recognition.
Hansen F. Chen, Peter N. Belhumeur, David Jacobs 0001
CVPR3
2000 Classification with Nonmetric Distances: Image Retrieval and Class Representation
abstract
A key problem in appearance-based vision is understanding how to use a set of labeled images to classify new images. Systems that model human performance, or that use robust image matching methods, often use nonmetric similarity judgments; but when the triangle inequality is not obeyed, most pattern recognition techniques are not applicable. Exemplar-based (nearest-neighbor) methods can be applied to a wide class of nonmetric similarity functions. The key issue, however, is to find methods for choosing good representatives of a class that accurately characterize it. We show that existing condensing techniques are ill-suited to deal with nonmetric dataspaces. We develop techniques for solving this problem, emphasizing two points: First, we show that the distance between images is not a good measure of how well one image can represent another in nonmetric spaces. Instead, we use the vector correlation between the distances from each image to other previously seen images. Second, we show that in nonmetric spaces, boundary points are less significant for capturing the structure of a class than in Euclidean spaces. We suggest that atypical points may be more important in describing classes. We demonstrate the importance of these ideas to learning that generalizes from experience by improving performance. We also suggest ways of applying parametric techniques to supervised learning problems that involve a specific nonmetric distance functions, showing how to generalize the idea of linear discriminant functions in a way that may be more useful in nonmetric spaces.
David Jacobs 0001, Daphna Weinshall, Yoram Gdalyahu
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Projective Alignment with Regions
abstract
We consider a recent approach to recognition that uses regions to determine the pose of objects while allowing for partial occlusion of the regions. We further analyze properties of the method for planar objects undergoing projective transformations. We prove that three visible regions are sufficient to determine the transformation uniquely, and that for a large class of objects two regions are insufficient. However, we show that when several regions are available, the pose of the object can generally be recovered even when all but two regions are significantly occluded. Our analysis is based on investigating the flow patterns of points under projective transformations in the presence of fixed points.
Ronen Basri, David Jacobs 0001
ICCV2
1999 3-D to 2-D Pose Determination with Regions
David Jacobs 0001, Ronen Basri
Int. J. Comput. Vis.1
1998 Clustering Appearances of 3D Objects
abstract
We introduce a method for unsupervised clustering of images of 3D objects. Our method examines the space of all images and partitions the images into sets that form smooth and parallel surfaces in this space. It further uses sequences of images to obtain more reliable clustering. Finally, since our method relies on a non-Euclidean similarity measure we introduce algebraic techniques for estimating local properties of these surfaces without first embedding the images in a Euclidean space. We demonstrate our method by applying it to a large database of images.
Ronen Basri, Dan Roth 0001, David Jacobs 0001
CVPR3
1998 Comparing Images under Variable Illumination
abstract
We consider the problem of determining whether two images come from different objects or the same object in the same pose, but under different illumination conditions. We show that this problem cannot be solved using hard constraints: even using a Lambertian reflectance model, there is always an object and a pair of lighting conditions consistent with any two images. Nevertheless, we show that for point sources and objects with Lambertian reflectance, the ratio of two images from the same object is simpler than the ratio of images from different objects. We also show that the ratio of the two images provides two of the three distinct values in the Hessian matrix of the object's surface. Using these observations, we develop a simple measure for matching images under variable illumination, comparing its performance to other existing methods on a database of 450 images of 10 individuals.
David Jacobs 0001, Peter N. Belhumeur, Ronen Basri
CVPR1
1998 Condensing Image Databases when Retrieval is Based on Non-Metric Distances
abstract
One of the key problems in appearance-based vision is understanding how to use a set of labeled images to classify new images. Classification systems that can model human performance, or that use robust image matching methods, often make use of similarity judgments that are non-metric but when the triangle inequality is not obeyed, most existing pattern recognition techniques are not applicable. We note that exemplar-based (or nearest-neighbor) methods can be applied naturally when using a wide class of non-metric similarity functions. The key issue, however, is to find methods for choosing good representatives of a class that accurately characterize it. We note that existing condensing techniques for finding class representatives are ill-suited to deal with non-metric dataspaces. We then focus on developing techniques for solving this problem, emphasizing two points: First, we show that the distance between two images is not a good measure of how well one image can represent another in non-metric spaces. Instead, we use the vector correlation between the distances from each image to other previously seen images. Second, we show that in non-metric spaces, boundary points are less significant for capturing the structure of a class than they are in Euclidean spaces. We suggest that atypical points may be more important in describing classes. We demonstrate the importance of these ideas to learning that generalizes from experience by improving performance using both synthetic and real images.
David Jacobs 0001, Daphna Weinshall, Yoram Gdalyahu
ICCV1
1998 Classification in Non-Metric Spaces
Daphna Weinshall, David Jacobs 0001, Yoram Gdalyahu
NIPS2
1998 Uncertainty Propagation in Model-Based Recognition
Tao Daniel Alter, David Jacobs 0001
Int. J. Comput. Vis.2
1998 Efficient determination of shape from multiple images containing partial information
Ronen Basri, Adam J. Grove, David Jacobs 0001
Pattern Recognit.3
1997 Linear Fitting with Missing Data: Applications to Structure-from-Motion and to Characterizing Intensity Images
abstract
Several vision problems can be reduced to the problem of fitting a linear surface of low dimension to data, including the problems of structure-from-affine-motion, and of characterizing the intensity images of a Lambertian scene by constructing the intensity manifold. For these problems, one must deal with a data matrix with some missing elements. In structure-from-motion, missing elements will occur if some point features are not visible in some frames. To construct the intensity manifold missing matrix elements will arise when the surface normals of some scene points do not face the light source in some images. We propose a novel method for fitting a low rank matrix to a matrix with missing elements. We show experimentally that our method produces good results in the presence of noise. These results can be either used directly, or can serve as an excellent starting point for an iterative method.
David Jacobs 0001
CVPR1
1997 3-D to 2-D Recognition with Regions
abstract
This paper presents a novel approach to parts-based object recognition in the presence of occlusion. We focus on the problem of determining the pose of a 3-D object from a single 2-D image when convex parts of the object have been matched to corresponding regions in the image. We consider three types of occlusions: self-occlusion, occlusions whose locus is identified in the image, and completely arbitrary occlusions. We derive efficient algorithms for the first two cases, and characterize their performance. For the last case, we prove that the problem of finding valid poses is computationally hard, but provide an efficient, approximate algorithm. This work generalizes our previous work on region-based object recognition, which focused on the case of planar models.
David Jacobs 0001, Ronen Basri
CVPR1
1997 Constancy and Similarity
Ronen Basri, David Jacobs 0001
Comput. Vis. Image Underst.2
1997 Recognition Using Region Correspondences
Ronen Basri, David Jacobs 0001
Int. J. Comput. Vis.2
1997 Matching 3-D Models to 2-D Images
David Jacobs 0001
Int. J. Comput. Vis.1
1997 Stochastic Completion Fields: A Neural Model of Illusory Contour Shape and Salience
abstract
We describe an algorithm- and representation-level theory of illusory contour shape and salience. Unlike previous theories, our model is derived from a single assumption: that the prior probability distribution of boundary completion shape can be modeled by a random walk in a lattice whose points are positions and orientations in the image plane (i.e., the space that one can reasonably assume is represented by neurons of the mammalian visual cortex). Our model does not employ numerical relaxation or other explicit minimization, but instead relies on the fact that the probability that a particle following a random walk will pass through a given position and orientation on a path joining two boundary fragments can be computed directly as the product of two vector-field convolutions. We show that for the random walk we define, the maximum likelihood paths are curves of least energy, that is, on average, random walks follow paths commonly assumed to model the shape of illusory contours. A computer model is demonstrated on numerous illusory contour stimuli from the literature.
Lance R. Williams, David Jacobs 0001
Neural Comput.2
1997 Local Parallel Computation of Stochastic Completion Fields
abstract
We describe a local parallel method for computing the stochastic completion field introduced in the previous article (Williams and Jacobs, 1997). The stochastic completion field represents the likelihood that a completion joining two contour fragments passes through any given position and orientation in the image plane. It is based on the assumption that the prior probability distribution of completion shape can be modeled as a random walk in a lattice of discrete positions and orientations. The local parallel method can be interpreted as a stable finite difference scheme for solving the underlying Fokker-Planck equation identified by Mumford (1994). The resulting algorithm is significantly faster than the previously employed method, which relied on convolution with large-kernel filters computed by Monte Carlo simulation. The complexity of the new method is O (n3m), while that of the previous algorithm was O(n4m2 (for an n × n image with m discrete orientations). Perhaps most significant, the use of a local method allows us to model the probability distribution of completion shape using stochastic processes that are neither homogeneous nor isotropic. For example, it is possible to modulate particle decay rate by a directional function of local image brightnesses (i.e., anisotropic decay). The effect is that illusory contours can be made to respect the local image brightness structure. Finally, we note that the new method is more plausible as a neural model since (1) unlike the previous method, it can be computed in a sparse, locally connected network, and (2) the network dynamics are consistent with psychophysical measurements of the time course of illusory contour formation.
Lance R. Williams, David Jacobs 0001
Neural Comput.2
1996 Local Parallel Computation of Stochastic Completion Fields
abstract
We describe a local parallel method for computing the stochastic completion field introduced in an earlier paper Williams and Jacobs (1995). The stochastic completion field represents the likelihood that a completion joining two contour fragments passes through any given position and orientation in the image plane. It is based upon the assumption that the prior probability distribution of completion shape can be modeled as a random walk in a lattice of discrete positions and orientations. The local parallel method can be interpreted as a stable finite difference scheme for solving the underlying Fokker-Planck equation identified by Mumford (1994). The resulting algorithm is significantly faster than the previously employed method which relied on convolution with large-kernel filters computed by Monte Carlo simulation. The complexity of the new method is Of(n/sup 3/m) while that of the previous algorithm was 0(n/sup 4/m/sup 2/) (for an n x n image with m discrete orientations). Perhaps most significantly, the use of a local method allows us to model the probability distribution of completion shape using stochastic processes which are neither homogenous nor isotropic.
Lance R. Williams, David Jacobs 0001
CVPR2
1996 Efficient determination of shape from multiple images containing partial information
abstract
We consider the problem of reconstructing the shape of an object from multiple images related by translations, when only small portions of the object can be observed in each image. Lindenbaum and Bruckstein (1988) have considered this problem in the specific case where the translating object is seen by small sensors, for application to the understanding of insect vision. Their solution is limited by the fact that its run time is exponential in the number of images and sensors. We show that the problem can be solved in time that is polynomial in the number of sensors, but is in fact NP complete when the number of images is unbounded. We therefore consider the special case of convex objects, which we can solve efficiently even when many images are used.
Ronen Basri, Adam J. Grove, David Jacobs 0001
ICPR3
1996 Space/time trade-offs for associative memory
abstract
In any storage scheme, there is some trade-off between the space used (size of memory) and access time. However, the nature of this trade-off seems to depend on more than just what is being stored-it also depends the types of queries we consider. We justify this claim by considering a particular memory model and contrast recognition (membership queries) with associative recall. We show that the latter task can require exponentially larger memories even when identical information is stored.
Adam J. Grove, David Jacobs 0001
ICPR2
1996 Robust and Efficient Detection of Salient Convex Groups
abstract
This paper describes an algorithm that robustly locates salient convex collections of line segments in an image. The algorithm is guaranteed to find all convex sets of line segments in which the length of the gaps between segments is smaller than some fixed proportion of the total length of the lines. This enables the algorithm to find convex groups whose contours are partially occluded or missing due to noise. We give an expected case analysis of the algorithm performance. This demonstrates that salient convexity is unlikely to occur at random, and hence is a strong clue that grouped line segments reflect underlying structure in the scene. We also show that our algorithm run time is O(n/sup 2/log(n)+nm), when we wish to find the m most salient groups in an image with n line segments. We support this analysis with experiments on real data, and demonstrate the grouping system as part of a complete recognition system.
David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 The Space Requirements of Indexing Under Perspective Projections
abstract
Object recognition systems can be made more efficient through the use of table lookup to match features. The cost of this indexing process depends on the space required to represent groups of model features in such a lookup table. We determine the space required to perform indexing of arbitrary sets of 3D model points for lookup from a single 2D image formed under perspective projection. We show that in this case, one must use a 3D surface to represent model groups, and we provide an analytic description of such a surface. This is in contrast to the cases of scaled-orthographic or affine projection, in which only a 2D surface is required to represent a group of model features. This demonstrates a fundamental way in which the recognition of objects under perspective projection is more complex than is recognition under other projection models.
David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
1995 Recognition Using Region Correspondences
abstract
A central problem in object recognition is to determine the transformation that relates the model to the image, given some partial correspondence between the two. This is useful in determining whether an object is present in an image, and if so, in determining where the object is. We present a novel method of solving this problem that uses region information. In our approach, the model is divided into volumes and the image is divided into regions. Given a match between subsets of volumes and regions (without any explicit correspondence between different pieces of the regions), the alignment transformation is computed. The method applies to planar objects under similarity, affine and projective transformations and to projections of 3D objects undergoing affine and projective transformations.>
Ronen Basri, David Jacobs 0001
ICCV2
1995 Stochastic Completion Fields: A Neural Model of Illusory Contour Shape and Salience
abstract
We describe an algorithm and representation level theory of illusory contour shape and salience. Unlike previous theories, our model is derived from a single assumption-namely, that the prior probability distribution of boundary completion shape can be modeled by a random walk in a lattice whose points are positions and orientations in the image plane (i.e. the space which one can reasonably assume is represented by neurons of the mammalian visual cortex). Our model does not employ numerical relaxation or other explicit minimization, but instead relies on the fact that the probability that a particle following a random walk will pass through a given position and orientation on a path joining two boundary fragments can be computed directly as the product of two vector-field convolutions. We show that for the random walk we define, the maximum likelihood paths are curves of least energy, that is, on average, random walks follow paths commonly assumed to model the shape of illusory contours. A computer model is demonstrated on numerous illusory contour stimuli from the literature.>
Lance R. Williams, David Jacobs 0001
ICCV2
1994 Error propagation in full 3D-from-2D object recognition
abstract
Robust recognition systems require a careful understanding of the effects of error in sensed features. Error in these image features results in uncertainty in the possible image location of each additional model feature. We present an accurate, analytic approximation for this uncertainty when model poses are based on matching three image and model points. This result applies to objects that are fully three-dimensional, where past results considered only two-dimensional objects. Further, we introduce a linear programming algorithm to compute this uncertainty when poses are based on any number of initial matches.>
Tao Daniel Alter, David Jacobs 0001
CVPR2
1994 Finding structurally consistent motion correspondences
abstract
Much work on deriving scene structure and motion from features assumes as input a set of tracked image features that share a common 3D motion. Producing this input requires segmenting independent motions, and detecting image features that do not correspond to 3D features, originating instead, for example in occlusion boundaries or specularities. We derive a linear program that tells when a set of tracked points might have come from 3D points that share a single motion, assuming affine motion and bounded error. We can also use linear programming to place conservative bounds on the structure of the scene that corresponds to tracked points. We implement and test this algorithm on real images.
David Jacobs 0001, S. Chakra Chennubhotla
ICPR (1)1
1994 A study of affine matching with bounded sensor error
W. Eric L. Grimson, Daniel P. Huttenlocher, David Jacobs 0001
Int. J. Comput. Vis.3
1993 2D images of 3-D oriented points
abstract
A number of vision problems have been shown to become simpler when one models projection from 3-D to 2-D as a nonrigid linear transformation. These results have been largely restricted to models and scenes that consist only of 3-D points. It is shown that, with this projection model, several vision tasks become fundamentally more complex in the somewhat more complicated domain of oriented points. More space is required for indexing models in a database, more images are required to derive structure from motion, and new views of an object cannot be synthesized linearly from old views.>
David Jacobs 0001
CVPR1
1993 Robust and efficient detection of convex groups
abstract
An algorithm is presented that finds all convex sets of line segments in an image, such that the length of the line segments account for at least some fixed proportion of the length of the convex hull. This enables the algorithm to find convex groups whose contours are partially occluded or missing due to noise. An expected time analysis of the algorithm's performance is performed, together with experiments on real images that show that the algorithm is efficient and that tell when the groups found are unlikely to occur at random, and are likely to capture the underlying structure of a scene.>
David Jacobs 0001
CVPR1
1992 Space efficient 3-D model indexing
abstract
It is shown that the set of 2-D images produced by a group of 3-D point features of a rigid model can be optimally represented with two lines in two high-dimensional spaces. This result is used to match images and model groups by table lookup. The table is efficiently built and accessed through analytic methods that account for the effect of sensing error. In real images, it reduces the set of potential matches by a factor of several thousand. This representation of a model's images is used to analyze two other approaches to recognition. It is determined when invariants exist in several domains, and it is shown that there is an infinite set of qualitatively similar nonaccidental properties.>
David Jacobs 0001
CVPR1
1992 A Study of Affine Matching With Bounded Sensor Error
W. Eric L. Grimson, Daniel P. Huttenlocher, David Jacobs 0001
ECCV3
1991 Model group indexing for recognition
abstract
It is shown that an index space can be a powerful tool for reducing the image-model match search by a factor of k/sup G-3/, but only when accompanied by some mechanism, such as grouping, that prevents the system from having to consider all matches between image groups of size G and model groups of size G. It is also shown that if image groups are to index a single point at recognition time, then the index space must contain pointers to each model group over a 2-D sheet, and should therefore be 2G-4 dimensional. A simple indexing system has been implemented to demonstrate these concepts, and a series of experiments have been conducted to investigate the tradeoffs between space and time. They indicate that the speedups are achievable, but require a large amount of space.>
David T. Clemens, David Jacobs 0001
CVPR2
1991 Optimal matching of planar models in 3D scenes
abstract
The problem of matching a model consisting of the point features of a flat object to point features found in an image that contains the object in an arbitrary three-dimensional pose is addressed. Once three points are matched, it is possible to determine the pose of the object. Assuming bounded sensing error, the author presents a solution to the problem of determining the range of possible locations in the image at which any additional model points may appear. This solution leads to an algorithm for determining the largest possible matching between image and model features that includes this initial hypothesis. The author implements a close approximation to this algorithm, which is O(nm in /sup 6/), where n is the number of image points, m is the number of model points, and in is the maximum sensing error. This algorithm is compared to existing methods, and it is shown that it produces more accurate results.>
David Jacobs 0001
CVPR1
1991 Space and Time Bounds on Indexing 3D Models from 2D Images
abstract
Model-based visual recognition systems often match groups of image features to groups of model features to form initial hypotheses, which are then verified. In order to accelerate recognition considerably, the model groups can be arranged in an index space (hashed) offline such that feasible matches are found by indexing into this space. For the case of 2D images and 3D models consisting of point features, bounds on the space required for indexing and on the speedup that such indexing can achieve are demonstrated. It is proved that, even in the absence of image error, each model must be represented by a 2D surface in the index space. This places an unexpected lower bound on the space required to implement indexing and proves that no quantity is invariant for all projections of a model into the image. Theoretical bounds on the speedup achieved by indexing in the presence of image error are also determined, and an implementation of indexing for measuring this speedup empirically is presented. It is found that indexing can produce only a minimal speedup on its own. However, when accompanied by a grouping operation, indexing can provide significant speedups that grow exponentially with the number of features in the groups.>
David T. Clemens, David Jacobs 0001
IEEE Trans. Pattern Anal. Mach. Intell.2