EDBT 2026 Demo / reviewers in the wild / expert
Alexander C. Berg
dblp:73/2348
· DBLP profile ↗
68ranked-venue papers
4as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 4 first-author · 7 since 2021Systems, architecture and hardware · 5Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improved Visual Grounding through Self-Consistent ExplanationsabstractVision-and-language models trained to match images with text can be combined with visual explanation methods to point to the locations of specific objects in an image. Our work shows that the localization - “grounding”- abilities of these models can be further improved by fine-tuning for self-consistent visual explanations. We propose a strategy for augmenting existing text-image datasets with paraphrases using a large language model, and SelfEQ, a weakly-supervised strategy on visual explanation maps for paraphrases that encourages self-consistency. Specifically, for an input textual phrase, we attempt to generate a paraphrase and finetune the model so that the phrase and paraphrase map to the same region in the image. We posit that this both expands the vocabulary that the model is able to handle, and improves the quality of the object locations highlighted by gradient-based visual explanation methods (e.g. GradCAM). We demonstrate that SelfEQ improves performance on Flickr30k, ReferIt, and RefCOCO+ over a strong baseline method and several prior works. Particularly, comparing to other methods that do not use any type of box annotations, we obtain 84.07% on Flickr30k (an absolute improvement of 4.69%), 67.40% on ReferIt (an absolute improvement of 7.68%), and 75.10%, 55.49% on RefCOCO+ test sets A and B respectively (an absolute improvement of 3.74% on average). Ruozhen He, Paola Cascante-Bonilla, Alexander C. Berg, Vicente Ordonez |
CVPR | 4 |
| 2024 | Joint Depth Prediction and Semantic Segmentation with Multi-View SAMabstractMulti-task approaches to joint depth and segmentation prediction are well-studied for monocular images. Yet, predictions from a single-view are inherently limited, while multiple views are available in many robotics applications. On the other end of the spectrum, video-based and full 3D methods require numerous frames to perform reconstruction and segmentation. With this work we propose a Multi-View Stereo (MVS) technique for depth prediction that benefits from rich semantic features of the Segment Anything Model (SAM). This enhanced depth prediction, in turn, serves as a prompt to our Transformer-based semantic segmentation decoder. We report the mutual benefit that both tasks enjoy in our quantitative and qualitative studies on the ScanNet dataset. Our approach consistently outperforms single-task MVS and segmentation models, along with multi-task monocular methods. Mykhailo Shvets, Dongxu Zhao 0001, Marc Niethammer, Roni Sengupta, Alexander C. Berg |
WACV | 5 |
| 2023 | Segment AnythingabstractWe introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest segmentation dataset to date (by far), with over 1 billion masks on 11M licensed and privacy respecting images. The model is designed and trained to be promptable, so it can transfer zero-shot to new image distributions and tasks. We evaluate its capabilities on numerous tasks and find that its zero-shot performance is impressive – often competitive with or even superior to prior fully supervised results. We are releasing the Segment Anything Model (SAM) and corresponding dataset (SA-1B) of 1B masks and 11M images at segment-anything.com to foster research into foundation models for computer vision. We recommend reading the full paper at: arxiv.org/abs/2304.02643. Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloé Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, Ross B. Girshick |
ICCV | 9 |
| 2022 | Similarity Search for Efficient Active Learning and Search of Rare ConceptsabstractMany active learning and search approaches are intractable for large-scale industrial settings with billions of unlabeled examples. Existing approaches search globally for the optimal examples to label, scaling linearly or even quadratically with the unlabeled data. In this paper, we improve the computational efficiency of active learning and search methods by restricting the candidate pool for labeling to the nearest neighbors of the currently labeled set instead of scanning over all of the unlabeled data. We evaluate several selection strategies in this setting on three large-scale computer vision datasets: ImageNet, OpenImages, and a de-identified and aggregated dataset of 10 billion publicly shared images provided by a large internet company. Our approach achieved similar mAP and recall as the traditional global approach while reducing the computational cost of selection by up to three orders of magnitude, enabling web-scale active learning. Cody Coleman, Edward Chou, Julian Katz-Samuels, Sean Culatana, Peter Bailis, Alexander C. Berg, Robert D. Nowak, Roshan Sumbaly, Matei Zaharia, Ismet Zeki Yalniz |
AAAI | 6 |
| 2022 | Point-Level Region Contrast for Object Detection Pre-TrainingabstractIn this work we present point-level region contrast, a self-supervised pre-training approach for the task of object detection. This approach is motivated by the two key factors in detection: localization and recognition. While accurate localization favors models that operate at the pixel- or point-level, correct recognition typically relies on a more holistic, region-level view of objects. Incorporating this perspective in pre-training, our approach performs contrastive learning by directly sampling individual point pairs from different regions. Compared to an aggregated representation per region, our approach is more robust to the change in input region quality, and further enables us to implicitly improve initial region assignments via online knowledge distillation during training. Both advantages are important when dealing with imperfect regions encountered in the unsupervised setting. Experiments show point-level region contrast improves on state-of-the-art pre-training methods for object detection and segmentation across multiple tasks and datasets, and we provide extensive ablation studies and visualizations to aid understanding. Code will be made available. Yutong Bai, Xinlei Chen, Alexander Kirillov, Alan L. Yuille, Alexander C. Berg |
CVPR | 5 |
| 2021 | Boundary IoU: Improving Object-Centric Image Segmentation EvaluationabstractWe present Boundary IoU (Intersection-over-Union), a new segmentation evaluation measure focused on boundary quality. We perform an extensive analysis across different error types and object sizes and show that Boundary IoU is significantly more sensitive than the standard Mask IoU measure to boundary errors for large objects and does not over-penalize errors on smaller objects. The new quality measure displays several desirable characteristics like symmetry w.r.t. prediction/ground truth pairs and balanced responsiveness across scales, which makes it more suitable for segmentation evaluation than other boundary-focused measures like Trimap IoU and F-measure. Based on Boundary IoU, we update the standard evaluation protocols for instance and panoptic segmentation tasks by proposing the Boundary AP (Average Precision) and Boundary PQ (Panoptic Quality) metrics, respectively. Our experiments show that the new evaluation metrics track boundary quality improvements that are generally overlooked by current Mask IoU-based evaluation metrics. We hope that the adoption of the new boundary-sensitive evaluation metrics will lead to rapid progress in segmentation methods that improve boundary quality.1 Bowen Cheng, Ross B. Girshick, Piotr Dollár, Alexander C. Berg, Alexander Kirillov |
CVPR | 4 |
| 2021 | Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single ImageabstractWe present Worldsheet, a method for novel view synthesis using just a single RGB image as input. The main insight is that simply shrink-wrapping a planar mesh sheet onto the input image, consistent with the learned intermediate depth, captures underlying geometry sufficient to generate photorealistic unseen views with large viewpoint changes. To operationalize this, we propose a novel differentiable texture sampler that allows our wrapped mesh sheet to be textured and rendered differentiably into an image from a target viewpoint. Our approach is category-agnostic, end-to-end trainable without using any 3D supervision, and requires a single image at test time. We also explore a simple extension by stacking multiple layers of Worldsheets to better handle occlusions. Worldsheet consistently outperforms prior state-of-the-art methods on single-image view synthesis across several datasets. Furthermore, this simple idea captures novel views surprisingly well on a wide range of high-resolution in-the-wild images, converting them into navigable 3D pop-ups. Video results and code are available at https://worldsheet.github.io. Ronghang Hu, Nikhila Ravi, Alexander C. Berg, Deepak Pathak |
ICCV | 3 |
| 2021 | Neural Pseudo-Label Optimism for the Bank Loan ProblemabstractWe study a class of classification problems best exemplified by the \emph{bank loan} problem, where a lender decides whether or not to issue a loan. The lender only observes whether a customer will repay a loan if the loan is issued to begin with, and thus modeled decisions affect what data is available to the lender for future decisions. As a result, it is possible for the lender's algorithm to ``get stuck'' with a self-fulfilling model. This model never corrects its false negatives, since it never sees the true label for rejected data, thus accumulating infinite regret. In the case of linear models, this issue can be addressed by adding optimism directly into the model predictions. However, there are few methods that extend to the function approximation case using Deep Neural Networks. We present Pseudo-Label Optimism (PLOT), a conceptually and computationally simple method for this setting applicable to DNNs. \PLOT{} adds an optimistic label to the subset of decision points the current model is deciding on, trains the model on all data so far (including these points along with their optimistic labels), and finally uses the resulting \emph{optimistic} model for decision making. \PLOT{} achieves competitive performance on a set of three challenging benchmark problems, requiring minimal hyperparameter tuning. We also show that \PLOT{} satisfies a logarithmic regret guarantee, under a Lipschitz and logistic mean label model, and under a separability condition on the data. Aldo Pacchiano, Shaun Singh, Edward Chou, Alexander C. Berg, Jakob N. Foerster |
NeurIPS | 4 |
| 2020 | SymGAN: Orientation Estimation without Annotation for Symmetric ObjectsabstractTraining a computer vision system to predict an object's pose is crucial to improving robotic manipulation, where robots can easily locate and then grasp objects. Some of the key challenges in pose estimation lie in obtaining labeled data and handling objects with symmetries. We explore both these problems of viewpoint estimation (object 3D orientation) by proposing a novel unsupervised training paradigm that only requires a 3D model of the object of interest. We show that we can successfully train an orientation detector, which simply consumes an RGB image, in an adversarial training framework, where the discriminator learns to provide a learning signal to retrieve the object orientation using a black-box non differentiable renderer. In order to overcome this non differentiability, we introduce a randomized sampling method to obtain training gradients. To our knowledge this is the first time an adversarial framework is employed to successfully train a viewpoint detector that can handle symmetric objects. Using this training framework we show state of the art results on 3D orientation prediction on T-LESS [12], a challenging dataset for texture-less and symmetric objects. Phil Ammirato, Jonathan Tremblay, Ming-Yu Liu 0001, Alexander C. Berg, Dieter Fox |
WACV | 4 |
| 2019 | IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of ThingsabstractIn this work, we present a new operator, called Instance Mask Projection (IMP), which projects a predicted instance segmentation as a new feature for semantic segmentation. It also supports back propagation and is trainable end-to end. By adding this operator, we introduce a new way to combine top-down and bottom-up information in semantic segmentation. Our experiments show the effectiveness of IMP on both clothing parsing (with complex layering, large deformations, and non-convex objects), and on street scene segmentation (with many overlapping instances and small objects). On the Varied Clothing Parsing dataset (VCP), we show instance mask projection can improve mIOU by 3 points over a state-of-the-art Panoptic FPN segmentation approach. On the ModaNet clothing parsing dataset, we show a dramatic improvement of 20.4% compared to existing baseline semantic segmentation results. In addition, the Instance Mask Projection operator works well on other (non-clothing) datasets, providing an improvement in mIOU of 3 points on “thing” classes of Cityscapes, a self-driving dataset, over a state-of-the-art approach. Cheng-Yang Fu, Tamara L. Berg, Alexander C. Berg |
ICCV | 3 |
| 2019 | Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionabstractSingle-frame object detectors perform well on videos sometimes, even without temporal context. However, challenges such as occlusion, motion blur, and rare poses of objects are hard to resolve without temporal awareness. Thus, there is a strong need to improve video object detection by considering long-range temporal dependencies. In this paper, we present a light-weight modification to a single-frame detector that accounts for arbitrary long dependencies in a video. It improves the accuracy of a single-frame detector significantly with negligible compute overhead. The key component of our approach is a novel temporal relation module, operating on object proposals, that learns the similarities between proposals from different frames and selects proposals from past and/or future to support current proposals. Our final “causal" model, without any offline post-processing steps, runs at a similar speed as a single-frame detector and achieves state-of-the-art video object detection on ImageNet VID dataset. Mykhailo Shvets, Wei Liu 0015, Alexander C. Berg |
ICCV | 3 |
| 2019 | Combining Multiple Cues for Visual Madlibs Question Answering
Tatiana Tommasi, Arun Mallya, Bryan A. Plummer, Svetlana Lazebnik, Alexander C. Berg, Tamara L. Berg |
Int. J. Comput. Vis. | 5 |
| 2018 | Three years of low-power image recognition challenge: Introduction to special sessionabstractReducing power consumption has been one of the most important goals since the creation of electronic systems. Energy efficiency is increasingly important as battery-powered systems (such as smartphones, drones, and body cameras) are widely used. It is desirable using the on-board computers to recognize objects in the images captured by these cameras. The Low-Power Image Recognition Challenge (LPIRC) is an annual competition started in 2015. The special session includes presentations given by the winners of the first three years of LPIRC. This paper explains the rules of the competition and the rationale, summarizes the teams' scores, and describes the lessons learned in the first three years. The paper suggests possible improvements of future challenges. Kent Gauen, Ryan Dailey, Yung-Hsiang Lu, Eunbyung Park, Wei Liu 0015, Alexander C. Berg, Yiran Chen 0001 |
DATE | 6 |
| 2018 | Meta-tracker: Fast and Robust Online Adaptation for Visual Object Trackers
Eunbyung Park, Alexander C. Berg |
ECCV (3) | 2 |
| 2017 | Low-power image recognition challengeabstractSignificant progress has been made in recent years using computer programs recognizing objects in images. Meanwhile, many cameras are embedded in battery-powered systems (such as mobile phones, wearable devices, and drones) and energy efficiency is essential. Even though many research papers have been published on the topics related to low power and image recognition, there does not exist a common metric for comparing different solutions in terms of (1) energy efficiency and (2) accuracy in recognition. Low-Power Image Recognition Challenge (LPIRC) is, to our knowledge, the only on-site competition that considers both energy consumption and recognition accuracy. LPIRC was held as one-day workshops in the Design Automation Conference in 2015 and 2016. Each participating team brought their own system to the workshops. The referee system of LPIRC includes (1) an intranet, (2) a power meter, and (3) an HTTP server that provided the images and accepted the answers from the contestants' systems. The scores were the ratio of recognition accuracy and the energy consumption. The winner of 2016 was able to analyze 7,347 images and achieve 9.44% normalized mAP (mean average precision) with average power consumption of 4.7 W. Another team analyzed 1,020 images and achieved 25.7% normalized mAP. Kent Gauen, Rohit Rangan, Anup Mohan, Yung-Hsiang Lu, Wei Liu 0015, Alexander C. Berg |
ASP-DAC | 6 |
| 2017 | Transformation-Grounded Image Generation Network for Novel 3D View SynthesisabstractWe present a transformation-grounded image generation network for novel 3D view synthesis from a single image. Our approach first explicitly infers the parts of the geometry visible both in the input and novel views and then casts the remaining synthesis problem as image completion. Specifically, we both predict a flow to move the pixels from the input to the novel view along with a novel visibility map that helps deal with occulsion/disocculsion. Next, conditioned on those intermediate results, we hallucinate (infer) parts of the object invisible in the input image. In addition to the new network structure, training with a combination of adversarial and perceptual loss results in a reduction in common artifacts of novel view synthesis such as distortions and holes, while successfully generating high frequency details and preserving visual aspects of the input image. We evaluate our approach on a wide range of synthetic and real examples. Both qualitative and quantitative results show our method achieves significantly better results compared to existing methods. Eunbyung Park, Jimei Yang, Ersin Yumer, Duygu Ceylan, Alexander C. Berg |
CVPR | 5 |
| 2017 | Video Highlight Prediction Using Audience Chat ReactionsabstractSports channel video portals offer an exciting domain for research on multimodal, multilingual analysis.We present methods addressing the problem of automatic video highlight prediction based on joint visual features and textual analysis of the real-world audience discourse with complex slang, in both English and traditional Chinese.We present a novel dataset based on League of Legends championships recorded from North American and Taiwanese Twitch.tvchannels (will be released for further research), and demonstrate strong results on these using multimodal, character-level CNN-RNN model architectures. Cheng-Yang Fu, Joon Lee, Mohit Bansal, Alexander C. Berg |
EMNLP | 4 |
| 2017 | A dataset for developing and benchmarking active visionabstractWe present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured in 9 unique scenes. We train a fast object category detector for instance detection on our data. Using the dataset we show that, although increasingly accurate and fast, the state of the art for object detection is still severely impacted by object scale, occlusion, and viewing direction all of which matter for robotics applications. We next validate the dataset for simulating active vision, and use the dataset to develop and evaluate a deep-network-based system for next best move prediction for object classification using reinforcement learning. Our dataset is available for download at cs.unc.edu/~ammirato/active_vision_dataset_website/. Phil Ammirato, Patrick Poirson, Eunbyung Park, Jana Kosecka, Alexander C. Berg |
ICRA | 5 |
| 2017 | An Evaluation of the NVIDIA TX1 for Supporting Real-Time Computer-Vision WorkloadsabstractAutonomous vehicles are an exemplar for forward-looking safety-critical real-time systems where significant computing capacity must be provided within strict size, weight, and power (SWaP) limits. A promising way forward in meeting these needs is to leverage multicore platforms augmented with graphics processing units (GPUs) as accelerators. Such an approach is being strongly advocated by NVIDIA, whose Jetson TX1 board is currently a leading multicore+GPU solution marketed for autonomous systems. Unfortunately, no study has ever been published that expressly evaluates the effectiveness of the TX1, or any other comparable platform, in hosting safety-critical real-time workloads. In this paper, such a study is presented. Specifically, the TX1 is evaluated via benchmarking efforts, blackbox evaluations of GPU behavior, and case-study evaluations involving computer-vision workloads inspired by autonomousdriving use cases. Autonomous vehicles are an exemplar for forward-looking safety-critical real-time systems where significant computing capacity must be provided within strict size, weight, and power (SWaP) limits. A promising way forward in meeting these needs is to leverage multicore platforms augmented with graphics processing units (GPUs) as accelerators. Such an approach is being strongly advocated by NVIDIA, whose Jetson TX1 board is currently a leading multicore+GPU solution marketed for autonomous systems. Unfortunately, no study has ever been published that expressly evaluates the effectiveness of the TX1, or any other comparable platform, in hosting safety-critical real-time workloads. In this paper, such a study is presented. Specifically, the TX1 is evaluated via benchmarking efforts, blackbox evaluations of GPU behavior, and case-study evaluations involving computer-vision workloads inspired by autonomous-driving use cases. Nathan Otterness, Ming Yang 0036, Sarah Rust, Eunbyung Park, James H. Anderson, F. Donelson Smith, Alexander C. Berg, Shige Wang |
RTAS | 7 |
| 2017 | When Was That Made?abstractIn this paper, we explore deep learning methods for estimating when the objects were made. Temporal estimation of objects is a challenging task which requires expertise in the object domain. With temporal information of objects, historian, genealogists, sociologist, archaeologist or conservationists can study the past through the objects. Toward this goal, we utilize features from existing deep networks and fine-tune new networks for temporal estimation task. The results demonstrate that the deep learning approach outperforms both a color-based baseline and visual data mining approach which is the previous state of the art method for the temporal estimation. To gain the insights into the deep network performance, we provide the analyses of neuron activations and their entropy including neuron temporal sensitivity, neuron activity and the correlation between discriminative parts from the deep network and the data mining approach. Finally, we demonstrate the potential of the temporal estimation pipeline for an interesting application such as fashion trend analysis. Sirion Vittayakorn, Alexander C. Berg, Tamara L. Berg |
WACV | 2 |
| 2016 | Fast Single Shot Detection and Pose EstimationabstractFor applications in navigation and robotics, estimating the 3D pose of objects is as important as detection. Many approaches to pose estimation rely on detecting or tracking parts or keypoints [11, 21]. In this paper we build on a recent state-of-the-art convolutional network for sliding-window detection [10] to provide detection and rough pose estimation in a single shot, without intermediate stages of detecting parts or initial bounding boxes. While not the first system to treat pose estimation as a categorization problem, this is the first attempt to combine detection and pose estimation at the same level using a deep learning approach. The key to the architecture is a deep convolutional network where scores for the presence of an object category, the offset for its location, and the approximate pose are all estimated on a regular grid of locations in the image. The resulting system is as accurate as recent work on pose estimation (42.4% 8 View mAVP on Pascal 3D+ [21] ) and significantly faster (46 frames per second (FPS) on a TITAN X GPU). This approach to detection and rough pose estimation is fast and accurate enough to be widely applied as a pre-processing step for tasks including high-accuracy pose estimation, object tracking and localization, and vSLAM. Patrick Poirson, Phil Ammirato, Cheng-Yang Fu, Wei Liu 0015, Jana Kosecka, Alexander C. Berg |
3DV | 6 |
| 2016 | Solving VIsual Madlibs with Multiple Cues
Tatiana Tommasi, Arun Mallya, Bryan A. Plummer, Svetlana Lazebnik, Alexander C. Berg, Tamara L. Berg |
BMVC | 5 |
| 2016 | SSD: Single Shot MultiBox Detector
Wei Liu 0015, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, Alexander C. Berg |
ECCV (1) | 7 |
| 2016 | Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Alexander C. Berg, Tamara L. Berg |
ECCV (2) | 4 |
| 2016 | Combining multiple sources of knowledge in deep CNNs for action recognitionabstractAlthough deep convolutional neural networks (CNNs) have shown remarkable results for feature learning and prediction tasks, many recent studies have demonstrated improved performance by incorporating additional handcrafted features or by fusing predictions from multiple CNNs. Usually, these combinations are implemented via feature concatenation or by averaging output prediction scores from several CNNs. In this paper, we present new approaches for combining different sources of knowledge in deep learning. First, we propose feature amplification, where we use an auxiliary, hand-crafted, feature (e.g. optical flow) to perform spatially varying soft-gating on intermediate CNN feature maps. Second, we present a spatially varying multiplicative fusion method for combining multiple CNNs trained on different sources that results in robust prediction by amplifying or suppressing the feature activations based on their agreement. We test these methods in the context of action recognition where information from spatial and temporal cues is useful, obtaining results that are comparable with state-of-the-art methods and outperform methods using only CNNs and optical flow features. Eunbyung Park, Xufeng Han, Tamara L. Berg, Alexander C. Berg |
WACV | 4 |
| 2016 | Large Scale Retrieval and Generation of Image Descriptions
Vicente Ordonez, Xufeng Han, Polina Kuznetsova, Girish Kulkarni, Margaret Mitchell, Kota Yamaguchi, Karl Stratos, Amit Goyal 0001, Jesse Dodge, Alyssa C. Mensch, Hal Daumé III, Alexander C. Berg, Yejin Choi 0001, Tamara L. Berg |
Int. J. Comput. Vis. | 12 |
| 2015 | Refer-to-as Relations as Semantic KnowledgeabstractWe study Refer-to-as relations as a new type of semanticknowledge. Compared to the much studied Is-a relation,which concerns factual taxonomy knowledge, Refer-to-as relationsaim to address pragmatic semantic knowledge. Forexample, a “penguin” is a “bird” from a taxonomy point ofview, but people rarely refer to a “penguin” as a “bird” invernacular use. This observation closely relates to the entrylevelcategorization studied in Prototype Theory in Psychology.We posit that Refer-to-as relations can be learned fromdata, and that both textual and visual information would behelpful in inferring the relations. By integrating existing lexicalstructure knowledge with language statistics and visualsimilarities, we formulate a collective inference approach tomap all object names in an encyclopedia to commonly usednames for each object. Our contributions include a new labeleddata set, the inference and optimization approach, andthe computed mappings and similarities. Song Feng 0002, Sujith Ravi, Ravi Kumar 0001, Polina Kuznetsova, Wei Liu 0015, Alexander C. Berg, Tamara L. Berg, Yejin Choi 0001 |
AAAI | 6 |
| 2015 | MatchNet: Unifying feature and metric learning for patch-based matchingabstractMotivated by recent successes on learning feature representations and on learning feature comparison functions, we propose a unified approach to combining both for training a patch matching system. Our system, dubbed Match-Net, consists of a deep convolutional network that extracts features from patches and a network of three fully connected layers that computes a similarity between the extracted features. To ensure experimental repeatability, we train MatchNet on standard datasets and employ an input sampler to augment the training set with synthetic exemplar pairs that reduce overfitting. Once trained, we achieve better computational efficiency during matching by disassembling MatchNet and separately applying the feature computation and similarity networks in two sequential stages. We perform a comprehensive set of experiments on standard datasets to carefully study the contributions of each aspect of MatchNet, with direct comparisons to established methods. Our results confirm that our unified approach improves accuracy over previous state-of-the-art results on patch matching datasets, while reducing the storage requirement for descriptors. We make pre-trained MatchNet publicly available. Xufeng Han, Thomas K. Leung, Yangqing Jia, Rahul Sukthankar, Alexander C. Berg |
CVPR | 5 |
| 2015 | PAIGE: PAirwise Image Geometry Encoding for improved efficiency in Structure-from-MotionabstractLarge-scale Structure-from-Motion systems typically spend major computational effort on pairwise image matching and geometric verification in order to discover connected components in large-scale, unordered image collections. In recent years, the research community has spent significant effort on improving the efficiency of this stage. In this paper, we present a comprehensive overview of various state-of-the-art methods, evaluating and analyzing their performance. Based on the insights of this evaluation, we propose a learning-based approach, the PAirwise Image Geometry Encoding (PAIGE), to efficiently identify image pairs with scene overlap without the need to perform exhaustive putative matching and geometric verification. PAIGE achieves state-of-the-art performance and integrates well into existing Structure-from-Motion pipelines. Johannes L. Schönberger, Alexander C. Berg, Jan-Michael Frahm |
CVPR | 2 |
| 2015 | Rebooting Computing and Low-Power Image Recognition Challengeabstract“Rebooting Computing” (RC) is an effort in the IEEE to rethink future computers. RC started in 2012 by the co-chairs, Elie Track (IEEE Council on Superconductivity) and Tom Conte (Computer Society). RC takes a holistic approach, considering revolutionary as well as evolutionary solutions needed to advance computer technologies. Three summits have been held in 2013 and 2014, discussing different technologies, from emerging devices to user interface, from security to energy efficiency, from neuromorphic to reversible computing. The first part of this paper introduces RC to the design automation community and solicits revolutionary ideas from the community for the directions of future computer research. Energy efficiency is identified as one of the most important challenges in future computer technologies. The importance of energy efficiency spans from miniature embedded sensors to wearable computers, from individual desktops to data centers. To gauge the state of the art, the RC Committee organized the first Low Power Image Recognition Challenge (LPIRC). Each image contains one or multiple objects, among 200 categories. A contestant has to provide a working system that can recognize the objects and report the bounding boxes of the objects. The second part of this paper explains LPIRC and the solutions from the top two winners. Yung-Hsiang Lu, Alan M. Kadin, Alexander C. Berg, Thomas M. Conte, Erik DeBenedictis, Ganesh Gingade, Bichlien Hoang, Yongzhen Huang, Boxun Li, Jingyu Liu 0004, Wei Liu 0015, Huizi Mao, Junran Peng, Tianqi Tang 0001, Elie K. Track, Jingqiu Wang, Tao Wang 0004, Yu Wang 0002 |
ICCAD | 3 |
| 2015 | Where to Buy It: Matching Street Clothing Photos in Online ShopsabstractIn this paper, we define a new task, Exact Street to Shop, where our goal is to match a real-world example of a garment item to the same item in an online shop. This is an extremely challenging task due to visual differences between street photos (pictures of people wearing clothing in everyday uncontrolled settings) and online shop photos (pictures of clothing items on people, mannequins, or in isolation, captured by professionals in more controlled settings). We collect a new dataset for this application containing 404,683 shop photos collected from 25 different online retailers and 20,357 street photos, providing a total of 39,479 clothing item matches between street and shop photos. We develop three different methods for Exact Street to Shop retrieval, including two deep learning baseline methods, and a method to learn a similarity measure between the street and shop domains. Experiments demonstrate that our learned similarity significantly outperforms our baselines that use existing deep learning based representations. M. Hadi Kiapour, Xufeng Han, Svetlana Lazebnik, Alexander C. Berg, Tamara L. Berg |
ICCV | 4 |
| 2015 | Visual Madlibs: Fill in the Blank Description Generation and Question AnsweringabstractIn this paper, we introduce a new dataset consisting of 360,001 focused natural language descriptions for 10,738 images. This dataset, the Visual Madlibs dataset, is collected using automatically produced fill-in-the-blank templates designed to gather targeted descriptions about: people and objects, their appearances, activities, and interactions, as well as inferences about the general scene or its broader context. We provide several analyses of the Visual Madlibs dataset and demonstrate its applicability to two new description generation tasks: focused description generation, and multiple-choice question-answering for images. Experiments using joint-embedding and deep learning methods show promising results on these tasks. Licheng Yu, Eunbyung Park, Alexander C. Berg, Tamara L. Berg |
ICCV | 3 |
| 2015 | Runway to Realway: Visual Analysis of FashionabstractClothing and fashion are an integral part of our everyday lives. In this paper we present an approach to studying fashion both on the runway and in more real-world settings, computationally, and at large scale, using computer vision. Our contributions include collecting a new runway dataset, designing features suitable for capturing outfit appearance, collecting human judgments of outfit similarity, and learning similarity functions on the features to mimic those judgments. We provide both intrinsic and extrinsic evaluations of our learned models to assess performance on outfit similarity prediction as well as season, year, and brand estimation. An example application tracks visual trends as runway fashions filter down to "real way" street fashions. Sirion Vittayakorn, Kota Yamaguchi, Alexander C. Berg, Tamara L. Berg |
WACV | 3 |
| 2015 | Predicting Entry-Level Categories
Vicente Ordonez, Wei Liu 0015, Jia Deng 0001, Yejin Choi 0001, Alexander C. Berg, Tamara L. Berg |
Int. J. Comput. Vis. | 5 |
| 2015 | ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng 0001, Hao Su 0001, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, Li Fei-Fei 0001 |
Int. J. Comput. Vis. | 11 |
| 2014 | Scalable multi-label annotationabstractWe study strategies for scalable multi-label annotation, or for efficiently acquiring multiple labels from humans for a collection of items. We propose an algorithm that exploits correlation, hierarchy, and sparsity of the label distribution. A case study of labeling 200 objects using 20,000 images demonstrates the effectiveness of our approach. The algorithm results in up to 6x reduction in human computation time compared to the naive method of querying a human annotator for the presence of every object in every image. Jia Deng 0001, Olga Russakovsky, Jonathan Krause, Michael S. Bernstein, Alexander C. Berg, Li Fei-Fei 0001 |
CHI | 5 |
| 2014 | Hipster Wars: Discovering Elements of Fashion Styles
M. Hadi Kiapour, Kota Yamaguchi, Alexander C. Berg, Tamara L. Berg |
ECCV (1) | 3 |
| 2014 | Materials discovery: Fine-grained classification of X-ray scattering imagesabstractWe explore the use of computer vision methods for organizing, searching, and classifying x-ray scattering images. X-ray scattering is a technique that shines an intense beam of x-rays through a sample of interest. By recording the intensity of x-ray deflection as a function of angle, scientists can measure the structure of materials at the molecular and nano-scale. Current and planned synchrotron instruments are producing x-ray scattering data at an unprecedented rate, making the design of automatic analysis techniques crucial for future research. In this paper, we devise an attribute-based approach to recognition in x-ray scattering images and demonstrate applications to image annotation and retrieval. M. Hadi Kiapour, Kevin G. Yager, Alexander C. Berg, Tamara L. Berg |
WACV | 3 |
| 2013 | From Large Scale Image Categorization to Entry-Level CategoriesabstractEntry level categories - the labels people will use to name an object - were originally defined and studied by psychologists in the 1980s. In this paper we study entry-level categories at a large scale and learn the first models for predicting entry-level categories for images. Our models combine visual recognition predictions with proxies for word "naturalness" mined from the enormous amounts of text on the web. We demonstrate the usefulness of our models for predicting nouns (entry-level words) associated with images by people. We also learn mappings between concepts predicted by existing visual recognition systems and entry-level concepts that could be useful for improving human-focused applications such as natural language image description or retrieval. Vicente Ordonez, Jia Deng 0001, Yejin Choi 0001, Alexander C. Berg, Tamara L. Berg |
ICCV | 4 |
| 2013 | Detecting Avocados to Zucchinis: What Have We Done, and Where Are We Going?abstractThe growth of detection datasets and the multiple directions of object detection research provide both an unprecedented need and a great opportunity for a thorough evaluation of the current state of the field of categorical object detection. In this paper we strive to answer two key questions. First, where are we currently as a field: what have we done right, what still needs to be improved? Second, where should we be going in designing the next generation of object detectors? Inspired by the recent work of Hoiem et al. on the standard PASCAL VOC detection dataset, we perform a large-scale study on the Image Net Large Scale Visual Recognition Challenge (ILSVRC) data. First, we quantitatively demonstrate that this dataset provides many of the same detection challenges as the PASCAL VOC. Due to its scale of 1000 object categories, ILSVRC also provides an excellent test bed for understanding the performance of detectors as a function of several key properties of the object classes. We conduct a series of analyses looking at how different detection methods perform on a number of image-level and object-class-level properties such as texture, color, deformation, and clutter. We learn important lessons of the current object detection methods and propose a number of insights for designing the next generation object detectors. Olga Russakovsky, Jia Deng 0001, Zhiheng Huang, Alexander C. Berg, Li Fei-Fei 0001 |
ICCV | 4 |
| 2013 | BabyTalk: Understanding and Generating Simple Image DescriptionsabstractWe present a system to automatically generate natural language descriptions from images. This system consists of two parts. The first part, content planning, smooths the output of computer vision-based detection and recognition algorithms with statistics mined from large pools of visually descriptive text to determine the best content words to use to describe an image. The second step, surface realization, chooses words to construct natural language sentences based on the predicted content and general statistics from natural language. We present multiple approaches for the surface realization step and evaluate each using automatic measures of similarity to human generated reference descriptions. We also collect forced choice human evaluations between descriptions from the proposed generation system and descriptions from competing approaches. The proposed system is very effective at producing relevant sentences for images. It also generates descriptions that are notably more true to the specific image content than previous work. Girish Kulkarni, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi 0001, Alexander C. Berg, Tamara L. Berg |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2013 | Efficient Classification for Additive Kernel SVMsabstractWe show that a class of nonlinear kernel SVMs admits approximate classifiers with runtime and memory complexity that is independent of the number of support vectors. This class of kernels, which we refer to as additive kernels, includes widely used kernels for histogram-based image comparison like intersection and chi-squared kernels. Additive kernel SVMs can offer significant improvements in accuracy over linear SVMs on a wide variety of tasks while having the same runtime, making them practical for large-scale recognition or real-time detection tasks. We present experiments on a variety of datasets, including the INRIA person, Daimler-Chrysler pedestrians, UIUC Cars, Caltech-101, MNIST, and USPS digits, to demonstrate the effectiveness of our method for efficient evaluation of SVMs with additive kernels. Since its introduction, our method has become integral to various state-of-the-art systems for PASCAL VOC object detection/image classification, ImageNet Challenge, TRECVID, etc. The techniques we propose can also be applied to settings where evaluation of weighted additive kernels is required, which include kernelized versions of PCA, LDA, regression, k-means, as well as speeding up the inner loop of SVM classifier training algorithms. Subhransu Maji, Alexander C. Berg, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Collective Generation of Natural Image Descriptions
Polina Kuznetsova, Vicente Ordonez, Alexander C. Berg, Tamara L. Berg, Yejin Choi 0001 |
ACL (1) | 3 |
| 2012 | Understanding and predicting importance in imagesabstractWhat do people care about in an image? To drive computational visual recognition toward more human-centric outputs, we need a better understanding of how people perceive and judge the importance of content in images. In this paper, we explore how a number of factors relate to human perception of importance. Proposed factors fall into 3 broad types: 1) factors related to composition, e.g. size, location, 2) factors related to semantics, e.g. category of object or scene, and 3) contextual factors related to the likelihood of attribute-object, or object-scene pairs. We explore these factors using what people describe as a proxy for importance. Finally, we build models to predict what will be described about an image given either known image content, or image content estimated automatically by recognition systems. Alexander C. Berg, Tamara L. Berg, Hal Daumé III, Jesse Dodge, Amit Goyal 0001, Xufeng Han, Alyssa C. Mensch, Margaret Mitchell, Aneesh Sood, Karl Stratos, Kota Yamaguchi |
CVPR | 1 |
| 2012 | Hedging your bets: Optimizing accuracy-specificity trade-offs in large scale visual recognitionabstractAs visual recognition scales up to ever larger numbers of categories, maintaining high accuracy is increasingly difficult. In this work, we study the problem of optimizing accuracy-specificity trade-offs in large scale recognition, motivated by the observation that object categories form a semantic hierarchy consisting of many levels of abstraction. A classifier can select the appropriate level, trading off specificity for accuracy in case of uncertainty. By optimizing this trade-off, we obtain classifiers that try to be as specific as possible while guaranteeing an arbitrarily high accuracy. We formulate the problem as maximizing information gain while ensuring a fixed, arbitrarily small error rate with a semantic hierarchy. We propose the Dual Accuracy Reward Trade-off Search (DARTS) algorithm and prove that, under practical conditions, it converges to an optimal solution. Experiments demonstrate the effectiveness of our algorithm on datasets ranging from 65 to over 10,000 categories. Jia Deng 0001, Jonathan Krause, Alexander C. Berg, Li Fei-Fei 0001 |
CVPR | 3 |
| 2012 | DCMSVM: Distributed parallel training for single-machine multiclass classifiersabstractWe present an algorithm and implementation for distributed parallel training of single-machine multiclass SVMs. While there is ongoing and healthy debate about the best strategy for multiclass classification, there are some features of the single-machine approach that are not available when training alternatives such as one-vs-all, and that are quite complex for tree based methods. One obstacle to exploring single-machine approaches on large datasets is that they are usually limited to running on a single machine! We build on a framework borrowed from parallel convex optimization - the alternating direction method of multipliers (ADMM) - to develop a new consensus based algorithm for distributed training of single-machine approaches. This is demonstrated with an implementation of our novel sequential dual algorithm (DCMSVM) which allows distributed parallel training with small communication requirements. Benchmark results show significant reduction in wall clock time compared to current state of the art multiclass SVM implementation (Liblinear) on a single node. Experiments are performed on large scale image classification including results with modern high-dimensional features. Xufeng Han, Alexander C. Berg |
CVPR | 2 |
| 2012 | Midge: Generating Image Descriptions From Computer Vision Detections
Margaret Mitchell, Jesse Dodge, Amit Goyal 0001, Kota Yamaguchi, Karl Stratos, Xufeng Han, Alyssa C. Mensch, Alexander C. Berg, Tamara L. Berg, Hal Daumé III |
EACL | 8 |
| 2012 | Detecting Visual Text
Jesse Dodge, Amit Goyal 0001, Xufeng Han, Alyssa C. Mensch, Margaret Mitchell, Karl Stratos, Kota Yamaguchi, Yejin Choi 0001, Hal Daumé III, Alexander C. Berg, Tamara L. Berg |
HLT-NAACL | 10 |
| 2011 | Composing Simple Image Descriptions using Web-scale N-grams
Siming Li, Girish Kulkarni, Tamara L. Berg, Alexander C. Berg, Yejin Choi 0001 |
CoNLL | 4 |
| 2011 | Hierarchical semantic indexing for large scale image retrievalabstractThis paper addresses the problem of similar image retrieval, especially in the setting of large-scale datasets with millions to billions of images. The core novel contribution is an approach that can exploit prior knowledge of a semantic hierarchy. When semantic labels and a hierarchy relating them are available during training, significant improvements over the state of the art in similar image retrieval are attained. While some of this advantage comes from the ability to use additional information, experiments exploring a special case where no additional data is provided, show the new approach can still outperform OASIS, the current state of the art for similarity learning. Exploiting hierarchical relationships is most important for larger scale problems, where scalability becomes crucial. The proposed learning approach is fundamentally parallelizable and as a result scales more easily than previous work. An additional contribution is a novel hashing scheme (for bilinear similarity on vectors of probabilities, optionally taking into account hierarchy) that is able to reduce the computational cost of retrieval. Experiments are performed on Caltech256 and the larger ImageNet dataset. Jia Deng 0001, Alexander C. Berg, Li Fei-Fei 0001 |
CVPR | 2 |
| 2011 | Baby talk: Understanding and generating simple image descriptionsabstractWe posit that visually descriptive language offers computer vision researchers both information about the world, and information about how people describe the world. The potential benefit from this source is made more significant due to the enormous amount of language data easily available today. We present a system to automatically generate natural language descriptions from images that exploits both statistics gleaned from parsing large quantities of text data and recognition algorithms from computer vision. The system is very effective at producing relevant sentences for images. It also generates descriptions that are notably more true to the specific image content than previous work. Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi 0001, Alexander C. Berg, Tamara L. Berg |
CVPR | 6 |
| 2011 | Who are you with and where are you going?abstractWe propose an agent-based behavioral model of pedestrians to improve tracking performance in realistic scenarios. In this model, we view pedestrians as decision-making agents who consider a plethora of personal, social, and environmental factors to decide where to go next. We formulate prediction of pedestrian behavior as an energy minimization on this model. Two of our main contributions are simple, yet effective estimates of pedestrian destination and social relationships (groups). Our final contribution is to incorporate these hidden properties into an energy formulation that results in accurate behavioral prediction. We evaluate both our estimates of destination and grouping, as well as our accuracy at prediction and tracking against state of the art behavioral model and show improvements, especially in the challenging observational situation of infrequent appearance observations-something that might occur in thousands of webcams available on the Internet. Kota Yamaguchi, Alexander C. Berg, Luis E. Ortiz, Tamara L. Berg |
CVPR | 2 |
| 2011 | Fast and Balanced: Efficient Label Tree Learning for Large Scale Object RecognitionabstractWe present a novel approach to efficiently learn a label tree for large scale classification with many classes. The key contribution of the approach is a technique to simultaneously determine the structure of the tree and learn the classifiers for each node in the tree. This approach also allows fine grained control over the efficiency vs accuracy trade-off in designing a label tree, leading to more balanced trees. Experiments are performed on large scale image classification with 10184 classes and 9 million images. We demonstrate significant improvements in test accuracy and efficiency with less training time and more balanced trees compared to the previous state of the art by Bengio et al. Jia Deng 0001, Sanjeev Satheesh, Alexander C. Berg, Li Fei-Fei 0001 |
NIPS | 3 |
| 2011 | Describable Visual Attributes for Face Verification and Image SearchabstractWe introduce the use of describable visual attributes for face verification and image search. Describable visual attributes are labels that can be given to an image to describe its appearance. This paper focuses on images of faces and the attributes used to describe them, although the concepts also apply to other domains. Examples of face attributes include gender, age, jaw shape, nose size, etc. The advantages of an attribute-based representation for vision tasks are manifold: They can be composed to create descriptions at various levels of specificity; they are generalizable, as they can be learned once and then applied to recognize new objects or categories without any further training; and they are efficient, possibly requiring exponentially fewer attributes (and training data) than explicitly naming each category. We show how one can create and label large data sets of real-world images to train classifiers which measure the presence, absence, or degree to which an attribute is expressed in images. These classifiers can then automatically label new images. We demonstrate the current effectiveness--and explore the future potential--of using attributes for face verification and image search via human and computational experiments. Finally, we introduce two new face data sets, named FaceTracer and PubFig, with labeled attributes and identities, respectively. Neeraj Kumar 0006, Alexander C. Berg, Peter N. Belhumeur, Shree K. Nayar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Automatic Attribute Discovery and Characterization from Noisy Web Data
Tamara L. Berg, Alexander C. Berg, Jonathan Shih |
ECCV (1) | 2 |
| 2010 | What Does Classifying More Than 10, 000 Image Categories Tell Us?
Jia Deng 0001, Alexander C. Berg, Kai Li 0001, Li Fei-Fei 0001 |
ECCV (5) | 2 |
| 2009 | Attribute and simile classifiers for face verificationabstractWe present two novel methods for face verification. Our first method - “attribute” classifiers - uses binary classifiers trained to recognize the presence or absence of describable aspects of visual appearance (e.g., gender, race, and age). Our second method - “simile” classifiers - removes the manual labeling required for attribute classification and instead learns the similarity of faces, or regions of faces, to specific reference people. Neither method requires costly, often brittle, alignment between image pairs; yet, both methods produce compact visual descriptions, and work on real-world images. Furthermore, both the attribute and simile classifiers improve on the current state-of-the-art for the LFW data set, reducing the error rates compared to the current best by 23.92% and 26.34%, respectively, and 31.68% when combined. For further testing across pose, illumination, and expression, we introduce a new data set - termed PubFig - of real-world images of public figures (celebrities and politicians) acquired from the internet. This data set is both larger (60,000 images) and deeper (300 images per individual) than existing data sets of its kind. Finally, we present an evaluation of human performance. Neeraj Kumar 0006, Alexander C. Berg, Peter N. Belhumeur, Shree K. Nayar |
ICCV | 2 |
| 2009 | Max-margin additive classifiers for detectionabstractWe present methods for training high quality object detectors very quickly. The core contribution is a pair of fast training algorithms for piece-wise linear classifiers, which can approximate arbitrary additive models. The classifiers are trained in a max-margin framework and significantly outperform linear classifiers on a variety of vision datasets. We report experimental results quantifying training time and accuracy on image classification tasks and pedestrian detection, including detection results better than the best previous on the INRIA dataset with faster training. Subhransu Maji, Alexander C. Berg |
ICCV | 2 |
| 2008 | Classification using intersection kernel support vector machines is efficientabstractStraightforward classification using kernelized SVMs requires evaluating the kernel for a test vector and each of the support vectors. For a class of kernels we show that one can do this much more efficiently. In particular we show that one can build histogram intersection kernel SVMs (IKSVMs) with runtime complexity of the classifier logarithmic in the number of support vectors as opposed to linear for the standard approach. We further show that by precomputing auxiliary tables we can construct an approximate classifier with constant runtime and space requirements, independent of the number of support vectors, with negligible loss in classification accuracy on various tasks. This approximation also applies to 1 - chi2and other kernels of similar form. We also introduce novel features based on a multi-level histograms of oriented edge energy and present experiments on various detection datasets. On the INRIA pedestrian dataset an approximate IKSVM classifier based on these features has the current best performance, with a miss rate 13% lower at 10-6False Positive Per Window than the linear SVM detector of Dalal & Triggs. On the Daimler Chrysler pedestrian dataset IKSVM gives comparable accuracy to the best results (based on quadratic SVM), while being 15times faster. In these experiments our approximate IKSVM is up to 2000times faster than a standard implementation and requires 200times less memory. Finally we show that a 50times speedup is possible using approximate IKSVM based on spatial pyramid features on the Caltech 101 dataset with negligible loss of accuracy. Subhransu Maji, Alexander C. Berg, Jitendra Malik |
CVPR | 2 |
| 2007 | Parsing Images of Architectural ScenesabstractWe address image parsing in the setting of architectural scenes. Our goal is to parse an image into regions of various types such as sky, foliage, buildings, and street. Furthermore we parse the building regions at a finer level of detail, identifying the positions of windows, doors, and rooflines, the colors of walls, and the spatial extent of particular buildings. Recognizing these individual elements is often impossible without the context provided by the initial parsing of the image, for instance a roofline is only defined in relation to the building below and the sky above. Our approach is driven by recognition of generic classes of visual appearance, e.g. for foliage. The generic recognition results boot-strap an image specific model that provides refined estimates to use for matting, segmentation, and more detailed parsing. Alexander C. Berg, Floraine Grabler, Jitendra Malik |
ICCV | 1 |
| 2006 | SVM-KNN: Discriminative Nearest Neighbor Classification for Visual Category RecognitionabstractWe consider visual category recognition in the framework of measuring similarities, or equivalently perceptual distances, to prototype examples of categories. This approach is quite flexible, and permits recognition based on color, texture, and particularly shape, in a homogeneous framework. While nearest neighbor classifiers are natural in this setting, they suffer from the problem of high variance (in bias-variance decomposition) in the case of limited sampling. Alternatively, one could use support vector machines but they involve time-consuming optimization and computation of pairwise distances. We propose a hybrid of these two methods which deals naturally with the multiclass setting, has reasonable computational complexity both in training and at run time, and yields excellent results in practice. The basic idea is to find close neighbors to a query sample and train a local support vector machine that preserves the distance function on the collection of neighbors. Our method can be applied to large, multiclass data sets for which it outperforms nearest neighbor and support vector machines, and remains efficient when the problem becomes intractable for support vector machines. A wide variety of distance functions can be used and our experiments show state-of-the-art performance on a number of benchmark data sets for shape and texture classification (MNIST, USPS, CUReT) and object recognition (Caltech- 101). On Caltech-101 we achieved a correct classification rate of 59.05%(±0.56%) at 15 training images per class, and 66.23%(±0.48%) at 30 training images. Hao Zhang 0003, Alexander C. Berg, Michael Maire, Jitendra Malik |
CVPR (2) | 2 |
| 2005 | Shape Matching and Object Recognition Using Low Distortion CorrespondencesabstractWe approach recognition in the framework of deformable shape matching, relying on a new algorithm for finding correspondences between feature points. This algorithm sets up correspondence as an integer quadratic programming problem, where the cost function has terms based on similarity of corresponding geometric blur point descriptors as well as the geometric distortion between pairs of corresponding feature points. The algorithm handles outliers, and thus enables matching of exemplars to query images in the presence of occlusion and clutter. Given the correspondences, we estimate an aligning transform, typically a regularized thin plate spline, resulting in a dense correspondence between the two shapes. Object recognition is then handled in a nearest neighbor framework where the distance between exemplar and query is the matching cost between corresponding points. We show results on two datasets. One is the Caltech 101 dataset (Fei-Fei, Fergus and Perona), an extremely challenging dataset with large intraclass variation. Our approach yields a 48% correct classification rate, compared to Fei-Fei et al 's 16%. We also show results for localizing frontal and profile faces that are comparable to special purpose approaches tuned to faces. Alexander C. Berg, Tamara L. Berg, Jitendra Malik |
CVPR (1) | 1 |
| 2005 | Recovering Human Body Configurations Using Pairwise Constraints between PartsabstractThe goal of this work is to recover human body configurations from static images. Without assuming a priori knowledge of scale, pose or appearance, this problem is extremely challenging and demands the use of all possible sources of information. We develop a framework which can incorporate arbitrary pairwise constraints between body parts, such as scale compatibility, relative position, symmetry of clothing and smooth contour connections between parts. We detect candidate body parts from bottom-up using parallelism, and use various pairwise configuration constraints to assemble them together into body configurations. To find the most probable configuration, we solve an integer quadratic programming problem with a standard technique using linear approximations. Approximate IQP allows us to incorporate much more information than the traditional dynamic programming and remains computationally efficient. 15 hand-labeled images are used to train the low-level part detector and learn the pairwise constraints. We show test results on a variety of images. Xiaofeng Ren, Alexander C. Berg, Jitendra Malik |
ICCV | 2 |
| 2004 | Names and Faces in the News
Tamara L. Berg, Alexander C. Berg, Jaety Edwards, Michael Maire, Ryan White, Yee Whye Teh, Erik G. Learned-Miller, David A. Forsyth |
CVPR (2) | 2 |
| 2004 | Whos In the PictureabstractThe context in which a name appears in a caption provides powerful cues as to who is depicted in the associated image. We obtain 44,773 face im- ages, using a face detector, from approximately half a million captioned news images and automatically link names, obtained using a named en- tity recognizer, with these faces. A simple clustering method can pro- duce fair results. We improve these results significantly by combining the clustering process with a model of the probability that an individual is depicted given its context. Once the labeling procedure is over, we have an accurately labeled set of faces, an appearance model for each individual depicted, and a natural language model that can produce ac- curate results on captions in isolation. Tamara L. Berg, Alexander C. Berg, Jaety Edwards, David A. Forsyth |
NIPS | 2 |
| 2003 | Recognizing Action at a DistanceabstractOur goal is to recognize human action at a distance, at resolutions where a whole person may be, say, 30 pixels tall. We introduce a novel motion descriptor based on optical flow measurements in a spatiotemporal volume for each stabilized human figure, and an associated similarity measure to be used in a nearest-neighbor framework. Making use of noisy optical flow measurements is the key challenge, which is addressed by treating optical flow not as precise pixel displacements, but rather as a spatial pattern of noisy measurements which are carefully smoothed and aggregated to form our spatiotemporal motion descriptor. To classify the action being performed by a human figure in a query sequence, we retrieve nearest neighbor(s) from a database of stored, annotated video sequences. We can also use these retrieved exemplars to transfer 2D/3D skeletons onto the figures in the query sequence, as well as two forms of data-based action synthesis "do as I do" and "do as I say". Results are demonstrated on ballet, tennis as well as football datasets. Alexei A. Efros, Alexander C. Berg, Greg Mori, Jitendra Malik |
ICCV | 2 |
| 2001 | Geometric Blur for Template MatchingabstractWe address the problem of finding point correspondences in images by way of an approach to template matching that is robust under affine distortions. This is achieved by applying "geometric blur" to both the template and the image, resulting in a fall-off in similarity that is close to linear in the norm of the distortion between the template and the image. Results in wide baseline stereo correspondence, face detection, and feature correspondence are included. Alexander C. Berg, Jitendra Malik |
CVPR (1) | 1 |
| 2000 | Approximating Aggregate Queries about Web Pages via Random Walks
Ziv Bar-Yossef, Alexander C. Berg, Steve Chien, Jittat Fakcharoenphol, Dror Weitz |
VLDB | 2 |