Jonathan Brandt

dblp:03/8 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 28 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorTheory of computation · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
20 papers
Face, body and person analysis · 28% Generative modeling · 20% Image recognition and object detection · 8%
Databases, data mining, and information retrieval
8 papers
Information retrieval · 100%
Computer graphics and multimedia
5 papers
Multimedia analysis and retrieval · 46% Visual content generation and editing · 37% Image and video processing · 11%

Topics — the 30 heaviest of 52, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.812024
TurboEdit: Instant Text-Based Image Editing · ECCV (80) 2024
Visual content generation and editing › image editing
text-guided image editing
0.812024
TurboEdit: Instant Text-Based Image Editing · ECCV (80) 2024
Information retrieval
image retrieval
0.742017
Spatial-Semantic Image Search by Visual Feature Synthesis · CVPR 2017
Detecting and Aligning Faces by Image Retrieval · CVPR 2013
Mobile Product Image Search by Automatic Query Object Extraction · ECCV (4) 2012
Computer vision › Face, body and person analysis
face detection
0.742015
A convolutional neural network cascade for face detection · CVPR 2015
Efficient Boosted Exemplar-Based Face Detection · CVPR 2014
Probabilistic Elastic Part Model for Unsupervised Face Detector Adaptation · ICCV 2013
Machine learning › Deep learning architectures and training
attention mechanism
0.622018
Top-Down Neural Attention by Excitation Backprop · Int. J. Comput. Vis. 2018
Top-Down Neural Attention by Excitation Backprop · ECCV (4) 2016
Machine learning › Trustworthy machine learning
interpretability
0.622018
Top-Down Neural Attention by Excitation Backprop · Int. J. Comput. Vis. 2018
Top-Down Neural Attention by Excitation Backprop · ECCV (4) 2016
Computer vision › Face, body and person analysis
face alignment
0.532014
Nonparametric Context Modeling of Local Appearance for Pose- and Expression-Robust Facial Landmark Localization · CVPR 2014
Exemplar-Based Graph Matching for Robust Facial Landmark Localization · ICCV 2013
Detecting and Aligning Faces by Image Retrieval · CVPR 2013
Machine learning › Generative modeling › cross-modal generation
story visualization
0.512021
AESOP: Abstract Encoding of Stories, Objects, and Pictures · ICCV 2021
Natural language and speech › Language models and text generation
text summarization
0.512021
StreamHover: Livestream Transcript Summarization and Annotation · EMNLP (1) 2021
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.512021
AESOP: Abstract Encoding of Stories, Objects, and Pictures · ICCV 2021
Multimedia analysis and retrieval › multimedia analysis › multimedia content description
multimedia annotation
0.512021
StreamHover: Livestream Transcript Summarization and Annotation · EMNLP (1) 2021
Information retrieval › image retrieval
object retrieval
0.432014
Spatially-Constrained Similarity Measurefor Large-Scale Object Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Object retrieval and localization with spatially-constrained similarity measure and k-NN re-ranking · CVPR 2012
A Local Bag-of-Features Model for Large-Scale Object Retrieval · ECCV (6) 2010
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.422016
Shortlist Selection with Residual-Aware Distance Estimator for K-Nearest Neighbor Search · CVPR 2016
Transform coding for fast approximate nearest neighbor search in high dimensions · CVPR 2010
Information retrieval › similarity search
nearest neighbor search
0.422016
Shortlist Selection with Residual-Aware Distance Estimator for K-Nearest Neighbor Search · CVPR 2016
Transform coding for fast approximate nearest neighbor search in high dimensions · CVPR 2010
Information retrieval
reranking
0.322014
Spatially-Constrained Similarity Measurefor Large-Scale Object Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Object retrieval and localization with spatially-constrained similarity measure and k-NN re-ranking · CVPR 2012
Computer vision › Face, body and person analysis
person identification
0.212016
A Multi-level Contextual Model for Person Recognition in Photo Albums · CVPR 2016
Information retrieval › similarity search
vector quantization
0.212016
Shortlist Selection with Residual-Aware Distance Estimator for K-Nearest Neighbor Search · CVPR 2016
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.212015
DeepFont: Identify Your Font from An Image · ACM Multimedia 2015
Machine learning › Efficient and distributed learning › adaptive computation
model cascading
0.212015
A convolutional neural network cascade for face detection · CVPR 2015
Machine learning › Efficient and distributed learning
model compression
0.212015
DeepFont: Identify Your Font from An Image · ACM Multimedia 2015
Image and video processing › document image analysis
font recognition
0.212015
DeepFont: A System for Font Recognition and Similarity · ACM Multimedia 2015
Computer vision › Face, body and person analysis › face detection
boosting-based face detection
0.212014
Efficient Boosted Exemplar-Based Face Detection · CVPR 2014
Computer vision › Face, body and person analysis › face detection
facial feature localization
0.212014
Consensus of Regression for Occlusion-Robust Facial Feature Localization · ECCV (4) 2014
Natural language and speech › Information extraction and text analysis › document understanding › document image analysis
font recognition
0.212014
Large-Scale Visual Font Recognition · CVPR 2014
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
detector adaptation
0.212013
Probabilistic Elastic Part Model for Unsupervised Face Detector Adaptation · ICCV 2013
Computer vision › Segmentation and scene understanding › part parsing
face parsing
0.212013
Exemplar-Based Face Parsing · CVPR 2013
Computer vision › Face, body and person analysis
face recognition
0.212013
Probabilistic Elastic Matching for Pose Variant Face Verification · CVPR 2013
Computer vision › Segmentation and scene understanding › object segmentation
face segmentation
0.212013
Exemplar-Based Face Parsing · CVPR 2013
Computer vision › Face, body and person analysis › face recognition
face verification
0.212013
Probabilistic Elastic Matching for Pose Variant Face Verification · CVPR 2013
Machine learning › Graph learning
graph matching
0.212013
Exemplar-Based Graph Matching for Robust Facial Landmark Localization · ICCV 2013

Methods — techniques the papers use, named apart from their topics

text conditioning · 1.5diffusion model · 1.5convolutional neural network · 1.2data augmentation · 0.8excitation backprop · 0.6discriminative classifier · 0.5inverted file · 0.3visual feature synthesis · 0.3vector quantization · 0.2unsupervised photo grouping · 0.2residual-aware distance estimation · 0.2identity prior · 0.2clothing and body appearance · 0.2boosting · 0.2stacked convolutional auto-encoder · 0.2model compression · 0.2voting-based localization · 0.2RANSAC · 0.2
YearPublicationVenuePosition
2024 TurboEdit: Instant Text-Based Image Editing
Zongze Wu 0002, Nicholas I. Kolkin, Jonathan Brandt, Richard Zhang 0001, Eli Shechtman
ECCV (80)3
2021 StreamHover: Livestream Transcript Summarization and Annotation
abstract
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu 0004
EMNLP (1)8
2021 AESOP: Abstract Encoding of Stories, Objects, and Pictures
Hareesh Ravi, Kushal Kafle, Scott Cohen, Jonathan Brandt, Mubbasir Kapadia
ICCV4
2019 Hotels-50K: A Global Hotel Recognition Dataset
abstract
Recognizing a hotel from an image of a hotel room is important for human trafficking investigations. Images directly link victims to places and can help verify where victims have been trafficked, and where their traffickers might move them or others in the future. Recognizing the hotel from images is challenging because of low image quality, uncommon camera perspectives, large occlusions (often the victim), and the similarity of objects (e.g., furniture, art, bedding) across different hotel rooms. To support efforts towards this hotel recognition task, we have curated a dataset of over 1 million annotated hotel room images from 50,000 hotels. These images include professionally captured photographs from travel websites and crowd-sourced images from a mobile application, which are more similar to the types of images analyzed in real-world investigations. We present a baseline approach based on a standard network architecture and a collection of data-augmentation approaches tuned to this problem domain.
Abby Stylianou, Hong Xuan, Maya Shende, Jonathan Brandt, Richard Souvenir, Robert Pless
AAAI4
2018 Learning to Doodle with Stroke Demonstrations and Deep Q-Networks
Jimei Yang, Jonathan Brandt, Demetri Terzopoulos
BMVC7
2018 Top-Down Neural Attention by Excitation Backprop
Jianming Zhang 0001, Sarah Adel Bargal, Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Stan Sclaroff
Int. J. Comput. Vis.4
2017 Spatial-Semantic Image Search by Visual Feature Synthesis
abstract
The performance of image retrieval has been improved tremendously in recent years through the use of deep feature representations. Most existing methods, however, aim to retrieve images that are visually similar or semantically relevant to the query, irrespective of spatial configuration. In this paper, we develop a spatial-semantic image search technology that enables users to search for images with both semantic and spatial constraints by manipulating concept text-boxes on a 2D query canvas. We train a convolutional neural network to synthesize appropriate visual features that captures the spatial-semantic constraints from the user canvas query. We directly optimize the retrieval performance of the visual features when training our deep neural network. These visual features then are used to retrieve images that are both spatially and semantically relevant to the user query. The experiments on large-scale datasets such as MS-COCO and Visual Genome show that our method outperforms other baseline and state-of-the-art methods in spatial-semantic image search.
Long Mai, Hailin Jin, Zhe Lin 0001, Jonathan Brandt, Feng Liu 0015
CVPR5
2016 Shortlist Selection with Residual-Aware Distance Estimator for K-Nearest Neighbor Search
abstract
In this paper, we introduce a novel shortlist computation algorithm for approximate, high-dimensional nearest neighbor search. Our method relies on a novel distance estimator: the residual-aware distance estimator, that accounts for the residual distances of data points to their respective quantized centroids, and uses it for accurate short-list computation. Furthermore, we perform the residual-aware distance estimation with little additional memory and computational cost through simple pre-computation methods for inverted index and multi-index schemes. Because it modifies the initial shortlist collection phase, our new algorithm is applicable to most inverted indexing methods that use vector quantization. We have tested the proposed method with the inverted index and multi-index on a diverse set of benchmarks including up to one billion data points with varying dimensions, and found that our method robustly improves the accuracy of shortlists (up to 127% relatively higher) over the state-of-the-art techniques with a comparable or even faster computational cost.
Jae-Pil Heo, Zhe Lin 0001, Xiaohui Shen, Jonathan Brandt, Sung-Eui Yoon
CVPR4
2016 A Multi-level Contextual Model for Person Recognition in Photo Albums
abstract
In this work, we present a new framework for person recognition in photo albums that exploits contextual cues at multiple levels, spanning individual persons, individual photos, and photo groups. Through experiments, we show that the information available at each of these distinct contextual levels provides complementary cues as to person identities. At the person level, we leverage clothing and body appearance in addition to facial appearance, and to compensate for instances where the faces are not visible. At the photo level we leverage a learned prior on the joint distribution of identities on the same photo to guide the identity assignments. Going beyond a single photo, we are able to infer natural groupings of photos with shared context in an unsupervised manner. By exploiting this shared contextual information, we are able to reduce the identity search space and exploit higher intra-personal appearance consistency within photo groups. Our new framework enables efficient use of these complementary multi-level contextual cues to improve overall recognition rates on the photo album person recognition task, as demonstrated through state-of-theart results on a challenging public dataset. Our results outperform competing methods by a significant margin, while being computationally efficient and practical in a real world application.
Jonathan Brandt, Zhe Lin 0001, Xiaohui Shen, Gang Hua 0001
CVPR2
2016 Top-Down Neural Attention by Excitation Backprop
Jianming Zhang 0001, Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Stan Sclaroff
ECCV (4)3
2016 Customized expression recognition for performance-driven cutout character animation
abstract
Performance-driven character animation enables users to create expressive results by performing the desired motion of the character with their face and/or body. However, for cutout animations where continuous motion is combined with discrete artwork replacements, supporting a performance-driven workflow has some unique requirements. To trigger the appropriate artwork replacements, the system must reliably detect a wide range of customized facial expressions that are challenging for existing recognition methods, which focus on a few canonical expressions (e.g., angry, disgusted, scared, happy, sad and surprised). Also, real usage scenarios require the system to work in realtime with minimal training. In this paper, we propose a novel customized expression recognition technique that meets all of these requirements. We first use a set of handcrafted features combining geometric features derived from facial landmarks and patch-based appearance features through group sparsity-based facial component learning. To improve discrimination and generalization, these handcrafted features are integrated into a custom-designed Deep Convolutional Neural Network (CNN) structure trained from publicly available facial expression datasets. The combined features are fed to an online ensemble of SVMs designed for the few training sample problem and performs in realtime. To improve temporal coherence, we also apply a Hidden Markov Model (HMM) to smooth the recognition results. Our system achieves state-of-the-art performance on canonical expression datasets and promising results on our collected dataset of customized expressions.
Xiang Yu 0002, Jianchao Yang, Linjie Luo, Wilmot Li, Jonathan Brandt, Dimitris N. Metaxas
WACV5
2016 Discovering Primary Objects in Videos by Saliency Fusion and Iterative Appearance Estimation
abstract
In this paper, we propose a new method for detecting primary objects in unconstrained videos in a completely automatic setting. Here, we define the primary object in a video as the object that presents saliently in most of the frames. Unlike previous works considering only local saliency detection or common pattern discovery, the proposed method integrates the local visual/motion saliency extracted from each frame, global appearance consistency throughout the video, and spatiotemporal smoothness constraint on object trajectories. We first identify a temporal coherent salient region throughout the whole video, and then explicitly learn a global appearance model to distinguish the primary object against the background. In order to obtain high-quality saliency estimations from both appearance and motion cues, we propose a novel self-adaptive saliency map fusion method by learning the reliability of saliency maps from labeled data. As a whole, our method can robustly localize and track primary objects in diverse video content, and handle the challenges such as fast object and camera motion, large scale and appearance variation, background clutter, and pose deformation. Moreover, compared with some existing approaches that assume the object is present in all the frames, our approach can naturally handle the case where the object is present only in part of the frames, e.g., the object enters the scene in the middle of the video or leaves the scene before the video ends. We also propose a new video data set containing 51 videos for primary object detection with per-frame ground-truth labeling. Quantitative experiments on several challenging video data sets demonstrate the superiority of our method compared with the recent state of the arts.
Gangqiang Zhao, Junsong Yuan 0001, Xiaohui Shen, Zhe Lin 0001, Brian L. Price, Jonathan Brandt
IEEE Trans. Circuits Syst. Video Technol.7
2015 A convolutional neural network cascade for face detection
abstract
In real-world face detection, large visual variations, such as those due to pose, expression, and lighting, demand an advanced discriminative model to accurately differentiate faces from the backgrounds. Consequently, effective models for the problem tend to be computationally prohibitive. To address these two conflicting challenges, we propose a cascade architecture built on convolutional neural networks (CNNs) with very powerful discriminative capability, while maintaining high performance. The proposed CNN cascade operates at multiple resolutions, quickly rejects the background regions in the fast low resolution stages, and carefully evaluates a small number of challenging candidates in the last high resolution stage. To improve localization effectiveness, and reduce the number of candidates at later stages, we introduce a CNN-based calibration stage after each of the detection stages in the cascade. The output of each calibration stage is used to adjust the detection window position for input to the subsequent stage. The proposed method runs at 14 FPS on a single CPU core for VGA-resolution images and 100 FPS using a GPU, and achieves state-of-the-art detection performance on two public face detection benchmarks.
Zhe Lin 0001, Xiaohui Shen, Jonathan Brandt, Gang Hua 0001
CVPR4
2015 DeepFont: A System for Font Recognition and Similarity
abstract
We develop the DeepFont system, a large-scale learning-based solution for automatic font identification, organization and selection. In this proposed technical demonstration, we will give our audience a tour to the DeepFont system, with the focus on its impacts on real consumer products, including but not limited to: 1) a cloud-based iOS App for font recognition; 2) a web-based tool for font similarity evaluation and discovery.
Zhangyang Wang, Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Yuyan Song, Joseph Hsieh, Sarah Kong, Thomas S. Huang
ACM Multimedia4
2015 DeepFont: Identify Your Font from An Image
abstract
As font is one of the core design concepts, automatic font identification and similar font suggestion from an image or photo has been on the wish list of many designers. We study the Visual Font Recognition (VFR) problem [4] LFE, and advance the state-of-the-art remarkably by developing the DeepFont system. First of all, we build up the first available large-scale VFR dataset, named AdobeVFR, consisting of both labeled synthetic data and partially labeled real-world data. Next, to combat the domain mismatch between available training and testing data, we introduce a Convolutional Neural Network (CNN) decomposition approach, using a domain adaptation technique based on a Stacked Convolutional Auto-Encoder (SCAE) that exploits a large corpus of unlabeled real-world text images combined with synthetic data preprocessed in a specific way. Moreover, we study a novel learning-based model compression approach, in order to reduce the DeepFont model size without sacrificing its performance. The DeepFont system achieves an accuracy of higher than 80% (top-5) on our collected dataset, and also produces a good font similarity measure for font selection and suggestion. We also achieve around 6 times compression of the model without any visible loss of recognition accuracy.
Zhangyang Wang, Jianchao Yang, Hailin Jin, Eli Shechtman, Aseem Agarwala, Jonathan Brandt, Thomas S. Huang
ACM Multimedia6
2015 Selective Pooling Vector for Fine-Grained Recognition
abstract
We propose a new framework for image recognition by selectively pooling local visual descriptors, and show its superior discriminative power on fine-grained image classification tasks. The representation is based on selecting the most confident local descriptors for nonlinear function learning using a linear approximation in an embedded higher dimensional space. The advantage of our Selective Pooling Vector over the previous state-of-the-art Super Vector and Fisher Vector representations, is that it ensures a more accurate learning function, which proves to be important for classifying details in fine-grained image recognition. Our experimental results corroborate this claim: with a simple linear SVM as the classifier, the selective pooling vector achieves significant performance gains on standard benchmark datasets for various fine-grained tasks such as the CMU Multi-PIE dataset for face recognition, the Caltech-UCSD Bird dataset and the Stanford Dogs dataset for fine-grained object categorization. On all datasets we outperform the state of the arts and boost the recognition rates to 96.4%, 48.9%, 52.0% respectively.
Jianchao Yang, Hailin Jin, Eli Shechtman, Jonathan Brandt, Tony X. Han
WACV5
2015 Scalable Similarity Learning Using Large Margin Neighborhood Embedding
abstract
Classifying large-scale image data into object categories is an important problem that has received increasing research attention. Given the huge amount of data, non-parametric approaches such as nearest neighbor classifiers have shown promising results, especially when they are underpinned by a learned distance or similarity measurement. Although metric learning has been well studied in the past decades, most existing algorithms are impractical to handle large-scale data sets. In this paper, we present an image similarity learning method that can scale well in both the number of images and the dimensionality of image descriptors. To this end, similarity comparison is restricted to each sample's local neighbors and a discriminative similarity measure is induced from large margin neighborhood embedding. We also exploit the ensemble of projections so that high-dimensional features can be processed in a set of lower-dimensional subspaces in parallel. The efficiency and scalability of our proposed model are validated on several data sets with scales varying from tens of thousands to one million images.
Jianchao Yang, Zhe Lin 0001, Jonathan Brandt, Shiyu Chang, Thomas S. Huang
WACV4
2014 Eigen-PEP for Video Face Recognition
Gang Hua 0001, Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt
ACCV (3)5
2014 Large-Scale Visual Font Recognition
abstract
This paper addresses the large-scale visual font recognition (VFR) problem, which aims at automatic identification of the typeface, weight, and slope of the text in an image or photo without any knowledge of content. Although visual font recognition has many practical applications, it has largely been neglected by the vision community. To address the VFR problem, we construct a large-scale dataset containing 2,420 font classes, which easily exceeds the scale of most image categorization datasets in computer vision. As font recognition is inherently dynamic and open-ended, i.e., new classes and data for existing categories are constantly added to the database over time, we propose a scalable solution based on the nearest class mean classifier (NCM). The core algorithm is built on local feature embedding, local feature metric learning and max-margin template selection, which is naturally amenable to NCM and thus to such open-ended classification problems. The new algorithm can generalize to new classes and new data at little added cost. Extensive experiments demonstrate that our approach is very effective on our synthetic test images, and achieves promising results on real world test images.
Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Tony X. Han
CVPR4
2014 Efficient Boosted Exemplar-Based Face Detection
abstract
Despite the fact that face detection has been studied intensively over the past several decades, the problem is still not completely solved. Challenging conditions, such as extreme pose, lighting, and occlusion, have historically hampered traditional, model-based methods. In contrast, exemplar-based face detection has been shown to be effective, even under these challenging conditions, primarily because a large exemplar database is leveraged to cover all possible visual variations. However, relying heavily on a large exemplar database to deal with the face appearance variations makes the detector impractical due to the high space and time complexity. We construct an efficient boosted exemplar-based face detector which overcomes the defect of the previous work by being faster, more memory efficient, and more accurate. In our method, exemplars as weak detectors are discriminatively trained and selectively assembled in the boosting framework which largely reduces the number of required exemplars. Notably, we propose to include non-face images as negative exemplars to actively suppress false detections to further improve the detection accuracy. We verify our approach over two public face detection benchmarks and one personal photo album, and achieve significant improvement over the state-of-the-art algorithms in terms of both accuracy and efficiency.
Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Gang Hua 0001
CVPR3
2014 Nonparametric Context Modeling of Local Appearance for Pose- and Expression-Robust Facial Landmark Localization
abstract
We propose a data-driven approach to facial landmark localization that models the correlations between each landmark and its surrounding appearance features. At runtime, each feature casts a weighted vote to predict landmark locations, where the weight is precomputed to take into account the feature's discriminative power. The feature voting-based landmark detection is more robust than previous local appearance-based detectors, we combine it with nonparametric shape regularization to build a novel facial landmark localization pipeline that is robust to scale, in-plane rotation, occlusion, expression, and most importantly, extreme head pose. We achieve state-of-the-art performance on two especially challenging in-the-wild datasets populated by faces with extreme head pose and expression.
Brandon M. Smith 0001, Jonathan Brandt, Zhe Lin 0001, Li Zhang 0003
CVPR2
2014 Consensus of Regression for Occlusion-Robust Facial Feature Localization
Xiang Yu 0002, Zhe Lin 0001, Jonathan Brandt, Dimitris N. Metaxas
ECCV (4)3
2014 Spatially-Constrained Similarity Measurefor Large-Scale Object Retrieval
abstract
One fundamental problem in object retrieval with the bag-of-words model is its lack of spatial information. Although various approaches are proposed to incorporate spatial constraints into the model, most of them are either too strict or too loose so that they are only effective in limited cases. In this paper, a new spatially-constrained similarity measure (SCSM) is proposed to handle object rotation, scaling, view point change and appearance deformation. The similarity measure can be efficiently calculated by a voting-based method using inverted files. During the retrieval process, object localization in the database images can also be simultaneously achieved using SCSM without post-processing. Furthermore, based on the retrieval and localization results of SCSM, we introduce a novel and robust re-ranking method with the k-nearest neighbors of the query for automatically refining the initial search results. Extensive performance evaluations on six public data sets show that SCSM significantly outperforms other spatial models including RANSAC-based spatial verification, while k-NN re-ranking outperforms most state-of-the-art approaches using query expansion. We also adapted SCSM for mobile product image search with an iterative algorithm to simultaneously extract the product instance from the mobile query image, identify the instance, and retrieve visually similar product images. Experiments on two product image search data sets show that our approach can robustly localize and extract the product in the query image, and hence drastically improve the retrieval accuracy over baseline methods.
Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Ying Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Probabilistic Elastic Matching for Pose Variant Face Verification
abstract
Pose variation remains to be a major challenge for real-world face recognition. We approach this problem through a probabilistic elastic matching method. We take a part based representation by extracting local features (e.g., LBP or SIFT) from densely sampled multi-scale image patches. By augmenting each feature with its location, a Gaussian mixture model (GMM) is trained to capture the spatial-appearance distribution of all face images in the training corpus. Each mixture component of the GMM is confined to be a spherical Gaussian to balance the influence of the appearance and the location terms. Each Gaussian component builds correspondence of a pair of features to be matched between two faces/face tracks. For face verification, we train an SVM on the vector concatenating the difference vectors of all the feature pairs to decide if a pair of faces/face tracks is matched or not. We further propose a joint Bayesian adaptation algorithm to adapt the universally trained GMM to better model the pose variations between the target pair of faces/face tracks, which consistently improves face verification accuracy. Our experiments show that our method outperforms the state-of-the-art in the most restricted protocol on Labeled Face in the Wild (LFW) and the YouTube video face database by a significant margin.
Gang Hua 0001, Zhe Lin 0001, Jonathan Brandt, Jianchao Yang
CVPR4
2013 Detecting and Aligning Faces by Image Retrieval
abstract
Detecting faces in uncontrolled environments continues to be a challenge to traditional face detection methods due to the large variation in facial appearances, as well as occlusion and clutter. In order to overcome these challenges, we present a novel and robust exemplar-based face detector that integrates image retrieval and discriminative learning. A large database of faces with bounding rectangles and facial landmark locations is collected, and simple discriminative classifiers are learned from each of them. A voting-based method is then proposed to let these classifiers cast votes on the test image through an efficient image retrieval technique. As a result, faces can be very efficiently detected by selecting the modes from the voting maps, without resorting to exhaustive sliding window-style scanning. Moreover, due to the exemplar-based framework, our approach can detect faces under challenging conditions without explicitly modeling their variations. Evaluation on two public benchmark datasets shows that our new face detection approach is accurate and efficient, and achieves the state-of-the-art performance. We further propose to use image retrieval for face validation (in order to remove false positives) and for face alignment/landmark localization. The same methodology can also be easily generalized to other face-related tasks, such as attribute recognition, as well as general object detection.
Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Ying Wu 0001
CVPR3
2013 Exemplar-Based Face Parsing
abstract
In this work, we propose an exemplar-based face image segmentation algorithm. We take inspiration from previous works on image parsing for general scenes. Our approach assumes a database of exemplar face images, each of which is associated with a hand-labeled segmentation map. Given a test image, our algorithm first selects a subset of exemplar images from the database, Our algorithm then computes a nonrigid warp for each exemplar image to align it with the test image. Finally, we propagate labels from the exemplar images to the test image in a pixel-wise manner, using trained weights to modulate and combine label maps from different exemplars. We evaluate our method on two challenging datasets and compare with two face parsing algorithms and a general scene parsing algorithm. We also compare our segmentation results with contour-based face alignment results, that is, we first run the alignment algorithms to extract contour points and then derive segments from the contours. Our algorithm compares favorably with all previous works on all datasets evaluated.
Brandon M. Smith 0001, Li Zhang 0003, Jonathan Brandt, Zhe Lin 0001, Jianchao Yang
CVPR3
2013 Probabilistic Elastic Part Model for Unsupervised Face Detector Adaptation
abstract
We propose an unsupervised detector adaptation algorithm to adapt any offline trained face detector to a specific collection of images, and hence achieve better accuracy. The core of our detector adaptation algorithm is a probabilistic elastic part (PEP) model, which is offline trained with a set of face examples. It produces a statistically aligned part based face representation, namely the PEP representation. To adapt a general face detector to a collection of images, we compute the PEP representations of the candidate detections from the general face detector, and then train a discriminative classifier with the top positives and negatives. Then we re-rank all the candidate detections with this classifier. This way, a face detector tailored to the statistics of the specific image collection is adapted from the original detector. We present extensive results on three datasets with two state-of-the-art face detectors. The significant improvement of detection accuracy over these state of-the-art face detectors strongly demonstrates the efficacy of the proposed face detector adaptation algorithm.
Gang Hua 0001, Zhe Lin 0001, Jonathan Brandt, Jianchao Yang
ICCV4
2013 Exemplar-Based Graph Matching for Robust Facial Landmark Localization
abstract
Localizing facial landmarks is a fundamental step in facial image analysis. However, the problem is still challenging due to the large variability in pose and appearance, and the existence of occlusions in real-world face images. In this paper, we present exemplar-based graph matching (EGM), a robust framework for facial landmark localization. Compared to conventional algorithms, EGM has three advantages: (1) an affine-invariant shape constraint is learned online from similar exemplars to better adapt to the test face, (2) the optimal landmark configuration can be directly obtained by solving a graph matching problem with the learned shape constraint, (3) the graph matching problem can be optimized efficiently by linear programming. To our best knowledge, this is the first attempt to apply a graph matching technique for facial landmark localization. Experiments on several challenging datasets demonstrate the advantages of EGM over state-of-the-art methods.
Jonathan Brandt, Zhe Lin 0001
ICCV2
2012 Object retrieval and localization with spatially-constrained similarity measure and k-NN re-ranking
abstract
One fundamental problem in object retrieval with the bag-of-visual words (BoW) model is its lack of spatial information. Although various approaches are proposed to incorporate spatial constraints into the BoW model, most of them are either too strict or too loose so that they are only effective in limited cases. We propose a new spatially-constrained similarity measure (SCSM) to handle object rotation, scaling, view point change and appearance deformation. The similarity measure can be efficiently calculated by a voting-based method using inverted files. Object retrieval and localization are then simultaneously achieved without post-processing. Furthermore, we introduce a novel and robust re-ranking method with the k-nearest neighbors of the query for automatically refining the initial search results. Extensive performance evaluations on six public datasets show that SCSM significantly outperforms other spatial models, while k-NN re-ranking outperforms most state-of-the-art approaches using query expansion.
Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Shai Avidan, Ying Wu 0001
CVPR3
2012 Interactive Facial Feature Localization
Vuong Le, Jonathan Brandt, Zhe Lin 0001, Lubomir D. Bourdev, Thomas S. Huang
ECCV (3)2
2012 Mobile Product Image Search by Automatic Query Object Extraction
Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Ying Wu 0001
ECCV (4)3
2010 Transform coding for fast approximate nearest neighbor search in high dimensions
abstract
We examine the problem of large scale nearest neighbor search in high dimensional spaces and propose a new approach based on the close relationship between nearest neighbor search and that of signal representation and quantization. Our contribution is a very simple and efficient quantization technique using transform coding and product quantization. We demonstrate its effectiveness in several settings, including large-scale retrieval, nearest neighbor classification, feature matching, and similarity search based on the bag-of-words representation. Through experiments on standard data sets we show it is competitive with state-of-the-art methods, with greater speed, simplicity, and generality. The resulting compact representation can be the basis for more elaborate hierarchical search structures for sub-linear approximate search. However, we demonstrate that optimized linear search using the quantized representation is extremely fast and trivially parallelizable on modern computer architectures, with further acceleration possible by way of GPU implementation.
Jonathan Brandt
CVPR1
2010 A Local Bag-of-Features Model for Large-Scale Object Retrieval
Zhe Lin 0001, Jonathan Brandt
ECCV (6)2
2005 Robust Object Detection via Soft Cascade
abstract
We describe a method for training object detectors using a generalization of the cascade architecture, which results in a detection rate and speed comparable to that of the best published detectors while allowing for easier training and a detector with fewer features. In addition, the method allows for quickly calibrating the detector for a target detection rate, false positive rate or speed. One important advantage of our method is that it enables systematic exploration of the ROC surface, which characterizes the trade-off between accuracy and speed for a given classifier.
Lubomir D. Bourdev, Jonathan Brandt
CVPR (2)2
1992 Corrigendum: An Algorithm for the Computation of the Hutchinson Distance
Jonathan Brandt, Carlos Cabrelli, Ursula Molter
Inf. Process. Lett.1
1991 An Algorithm for the Computation of the Hutchinson Distance
Jonathan Brandt, Carlos Cabrelli, Ursula Molter
Inf. Process. Lett.1