Serge J. Belongie

dblp:b/SJBelongie · DBLP profile ↗
← Back
197ranked-venue papers
7as first author
41since 2021 · last 2025
0000-0002-0388-5217ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 163 · 7 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 140 · 4 first-author · 25 since 2021Databases, data management, data science and information retrieval · 8Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 1
YearPublicationVenuePosition
2025 Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
abstract
Generalized few-shot 3D point cloud segmentation (GFS-PCS) adapts models to new classes with few support samples while retaining base class segmentation. Existing GFS-PCS methods enhance prototypes via interacting with support or query features but remain limited by sparse knowledge from few-shot samples. Meanwhile, 3D vision-language models (3D VLMs), generalizing across open-world novel classes, contain rich but noisy novel class knowledge. In this work, we introduce a GFS-PCS framework that synergizes dense but noisy pseudo-labels from 3D VLMs with precise yet sparse few-shot samples to maximize the strengths of both, named GFS-VL. Specifically, we present a prototype-guided pseudo-label selection to filter low-quality regions, followed by an adaptive infilling strategy that combines knowledge from pseudo-label contexts and few-shot samples to adaptively label the filtered, un-labeled areas. Additionally, we design a novel-base mix strategy to embed few-shot samples into training scenes, preserving essential context for improved novel class learning. Moreover, recognizing the limited diversity in current GFS-PCS benchmarks, we introduce two challenging benchmarks with diverse novel classes for comprehensive generalization evaluation. Experiments validate the effectiveness of our framework across models and datasets. Our approach and benchmarks provide a solid foundation for advancing GFS-PCS in the real world. The code is at here.
Zhaochong An, Guolei Sun, Yun Liu 0011, Runjia Li, Junlin Han, Ender Konukoglu, Serge J. Belongie
CVPR7
2025 Taxonomy-Aware Evaluation of Vision-Language Models
abstract
When a vision–language model (VLM) is prompted to identify an entity depicted in an image, it may answer "I see a conifer," rather than the specific label NORWAY SPRUCE. This raises two issues for evaluation: Firstly, the unconstrained generated text needs to be mapped to the evaluation label space (i.e., CONIFER). Secondly, a useful classification measure should give partial credit to lessspecific, but not incorrect, answers (NORWAY SPRUCE being a type of CONIFER). To meet these requirements, we propose a framework for evaluating unconstrained text predictions such as those generated from a vision–language model against a taxonomy. Specifically, we propose the use of hierarchical precision and recall measures to assess the level of correctness and specificity of predictions with regard to a taxonomy. Experimentally, we first show that existing text similarity measures do not capture taxonomic similarity well. We then develop and compare different methods to map textual VLM predictions onto a taxonomy. This allows us to compute hierarchical similarity measures between the generated text and the ground truth labels. Finally, we analyze modern VLMs on fine-grained visual classification tasks based on our proposed taxonomic evaluation scheme. Data and code are made available at https://github.com/vesteinn/vlm-eval.
Vésteinn Snæbjarnarson, Kevin Du, Niklas Stoehr, Serge J. Belongie, Ryan Cotterell, Nico Lang, Stella Frank
CVPR4
2025 Multi-Modal Framing Analysis of News
abstract
Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its reception.So far, such studies have been narrow, in that they use a fixed set of pre-defined frames and focus only on the text, ignoring the visual contexts in which those texts appear.Especially for framing in the news, this leaves out valuable information about editorial choices, which include not just the written article but also accompanying photographs.To overcome such limitations, we present a method for conducting multi-modal, multi-label framing analysis at scale using large (vision-) language models.Grounding our work in framing theory, we extract latent meaning embedded in images used to convey a certain point and contrast that to the text by comparing the respective frames used.We also identify highly partisan framing of topics with issue-specific frame analysis found in prior qualitative work.We demonstrate a method for doing scalable integrative framing analysis of both text and image in news, providing a more complete picture for understanding media bias.
Arnav Arora, Srishti Yadav, Maria Antoniak, Serge J. Belongie, Isabelle Augenstein
EMNLP4
2025 Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
abstract
Meta-learning is a powerful paradigm for tackling few-shot tasks. However, recent studies indicate that models trained with the whole-class training strategy can achieve comparable performance to those trained with meta-learning in few-shot classification tasks. To demonstrate the value of meta-learning, we establish an entropy-limited supervised setting for fair comparisons. Through both theoretical analysis and experimental validation, we establish that meta-learning has a tighter generalization bound compared to whole-class training. We unravel that meta-learning is more efficient with limited entropy and is more robust to label noise and heterogeneous tasks, making it well-suited for unsupervised tasks. Based on these insights, We propose MINO, a meta-learning framework designed to enhance unsupervised performance. MINO utilizes the adaptive clustering algorithm DBSCAN with a dynamic head for unsupervised task construction and a stability-based meta-scaler for robustness against label noise. Extensive experiments confirm its effectiveness in multiple unsupervised few-shot and zero-shot tasks.
Yunchuan Guan, Yu Liu 0040, Ke Zhou 0001, Zhiqi Shen 0001, Jenq-Neng Hwang, Serge J. Belongie, Lei Li 0050
ICCV6
2025 Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
abstract
Few-shot 3D point cloud segmentation (FS-PCS) aims at generalizing models to segment novel categories with minimal annotated support samples. While existing FS-PCS methods have shown promise, they primarily focus on unimodal point cloud inputs, overlooking the potential benefits of leveraging multimodal information. In this paper, we address this gap by introducing a multimodal FS-PCS setup, utilizing textual labels and the potentially available 2D image modality. Under this easy-to-achieve setup, we present the MultiModal Few-Shot SegNet (MM-FSS), a model effectively harnessing complementary information from multiple modalities. MM-FSS employs a shared backbone with two heads to extract intermodal and unimodal visual features, and a pretrained text encoder to generate text embeddings. To fully exploit the multimodal information, we propose a Multimodal Correlation Fusion (MCF) module to generate multimodal correlations, and a Multimodal Semantic Fusion (MSF) module to refine the correlations using text-aware semantic guidance. Additionally, we propose a simple yet effective Test-time Adaptive Cross-modal Calibration (TACC) technique to mitigate training bias, further improving generalization. Experimental results on S3DIS and ScanNet datasets demonstrate significant performance improvements achieved by our method. The efficacy of our approach indicates the benefits of leveraging commonly-ignored free modalities for FS-PCS, providing valuable insights for future research. The code is available at github.com/ZhaochongAn/Multimodality-3D-Few-Shot.
Zhaochong An, Guolei Sun, Yun Liu 0011, Runjia Li, Min Wu 0008, Ming-Ming Cheng, Ender Konukoglu, Serge J. Belongie
ICLR8
2025 Unlearning-based Neural Interpretations
abstract
Gradient-based interpretations often require an anchor point of comparison to avoid saturation in computing feature importance. We show that current baselines defined using static functions—constant mapping, averaging or blurring—inject harmful colour, texture or frequency assumptions that deviate from model behaviour. This leads to accumulation of irregular gradients, resulting in attribution maps that are biased, fragile and manipulable. Departing from the static approach, we propose $\texttt{UNI}$ to compute an (un)learnable, debiased and adaptive baseline by perturbing the input towards an $\textit{unlearning direction}$ of steepest ascent. Our method discovers reliable baselines and succeeds in erasing salient features, which in turn locally smooths the high-curvature decision boundaries. Our analyses point to unlearning as a promising avenue for generating faithful, efficient and robust interpretations.
Ching Lam Choi, Alexandre Duplessis, Serge J. Belongie
ICLR3
2025 Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
abstract
Sparse Autoencoders (SAEs) have recently gained attention as a means to improve the interpretability and steerability of Large Language Models (LLMs), both of which are essential for AI safety. In this work, we extend the application of SAEs to Vision-Language Models (VLMs), such as CLIP, and introduce a comprehensive framework for evaluating monosemanticity at the neuron-level in visual representations. To ensure that our evaluation aligns with human perception, we propose a benchmark derived from a large-scale user study. Our experimental results reveal that SAEs trained on VLMs significantly enhance the monosemanticity of individual neurons, with sparsity and wide latents being the most influential factors. Further, we demonstrate that applying SAE interventions on CLIP's vision encoder directly steers multimodal LLM outputs (e.g., LLaVA), without any modifications to the underlying language model. These findings emphasize the practicality and efficacy of SAEs as an unsupervised tool for enhancing both interpretability and control of VLMs. Code and benchmark data are available at https://github.com/ExplainableML/sae-for-vlm.
Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge J. Belongie, Zeynep Akata
NeurIPS4
2025 RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
abstract
The rapid advancement of diffusion models has enabled high-fidelity and semantically rich text-to-image generation; however, ensuring fairness and safety remains an open challenge. Existing methods typically improve fairness and safety at the expense of semantic fidelity and image quality. In this work, we propose RespoDiff, a novel framework for responsible text-to-image generation that incorporates a dual-module transformation on the intermediate bottleneck representations of diffusion models. Our approach introduces two distinct learnable modules: one focused on capturing and enforcing responsible concepts, such as fairness and safety, and the other dedicated to maintaining semantic alignment with neutral prompts. To facilitate the dual learning process, we introduce a novel score-matching objective that enables effective coordination between the modules. Our method outperforms state-of-the-art methods in responsible generation by ensuring semantic alignment while optimizing both objectives without compromising image fidelity. Our approach improves responsible and semantically coherent generation by \textasciitilde20\% across diverse, unseen prompts. Moreover, it integrates seamlessly into large-scale models like SDXL, enhancing fairness and safety. The project page is available at https://vssilpa.github.io/respodiff_project_page.
Silpa Vadakkeeveetil Sreelatha, Sauradip Nag, Serge J. Belongie, Anjan Dutta 0001
NeurIPS4
2025 Noise-Coded Illumination for Forensic and Photometric Video Analysis
abstract
The proliferation of advanced tools for manipulating video has led to an arms race, pitting those who wish to sow disinformation against those who want to detect and expose it. Unfortunately, time favors the ill-intentioned in this race, with fake videos growing increasingly difficult to distinguish from real ones. At the root of this trend is a fundamental advantage held by those manipulating media: equal access to a distribution of what we consider authentic (i.e., “natural”) video. In this paper, we show how coding very subtle, noise-like modulations into the illumination of a scene can help combat this advantage by creating an information asymmetry that favors verification. Our approach effectively adds a temporal watermark to any video recorded under coded illumination. However, rather than encoding a specific message, this watermark encodes an image of the unmanipulated scene as it would appear lit only by the coded illumination. We show that even when an adversary knows that our technique is being used, creating a plausible coded fake video amounts to solving a second, more difficult version of the original adversarial content creation problem at an information disadvantage. This is a promising avenue for protecting high-stakes settings like public events and interviews, where the content on display is a likely target for manipulation, and while the illumination can be controlled, the cameras capturing video cannot.
Peter F. Michael, Zekun Hao, Serge J. Belongie, Abe Davis
ACM Trans. Graph.3
2024 Rethinking Few-shot 3D Point Cloud Semantic Segmentation
abstract
This paper revisits few-shot 3D point cloud semantic segmentation (FS-PCS), with a focus on two significant is-sues in the state-of-the-art: foreground leakage and sparse point distribution. The former arises from non-uniform point sampling, allowing models to distinguish the density disparities between foreground and background for easier segmentation. The latter results from sampling only 2,048 points, limiting semantic information and deviating from the real-world practice. To address these issues, we in-troduce a standardized FS-PCS setting, upon which a new benchmark is built. Moreover, we propose a novel FS-PCS model. While previous methods are based on feature op-timization by mainly refining support features to enhance prototypes, our method is based on correlation optimization, referred to as Correlation Optimization Segmentation (COSeg). Specifically, we compute Class-specific Multi-prototypical Correlation (CMC) for each query point, rep-resenting its correlations to category prototypes. Then, we propose the Hyper Correlation Augmentation (HCA) mod-ule to enhance CMC. Furthermore, tackling the inherent property of few-shot training to incur base susceptibility for models, we propose to learn non-parametric prototypes for the base classes during training. The learned base proto-types are used to calibrate correlations for the background class through a Base Prototypes Calibration (BPC) module. Experiments on popular datasets demonstrate the superior-ity of COSeg over existing methods. The code is available at github.com/ZhaochongAnICOSeg.
Zhaochong An, Guolei Sun, Yun Liu 0011, Fayao Liu, Zongwei Wu, Dan Wang 0011, Luc Van Gool, Serge J. Belongie
CVPR8
2024 From Coarse to Fine-Grained Open-Set Recognition
abstract
Open-set recognition (OSR) methods aim to identify whether or not a test example belongs to a category observed during training. Depending on how visually similar a test example is to the training categories, the OSR task can be easy or extremely challenging. However, the vast majority of previous work has studied OSR in the presence of large, coarse-grained semantic shifts. In contrast, many real-world problems are inherently fine-grained, which means that test examples may be highly visually similar to the training categories. Motivated by this observation, we investigate three aspects of OSR: label granularity, similarity between the open- and closed-sets, and the role of hierarchical supervision during training. To study these dimensions, we curate new open-set splits of a large fine-grained visual categorization dataset. Our analysis results in several interesting findings, including: (i) the best OSR method to use is heavily dependent on the degree of semantic shift present, and (ii) hierarchical representation learning can improve coarse-grained OSR, but has little effect on fine-grained OSR performance. To further enhance fine-grained OSR performance, we propose a hierarchy-adversarial learning method to discourage hierarchical structure in the representation space, which results in a perhaps counter-intuitive behaviour, and a relative improvement in fine-grained OSR of up to 2% in AUROC and 7% in AUPR over standard training. Code and data are available: langnico.github. io/fine-grained-osr.
Nico Lang, Vésteinn Snæbjarnarson, Elijah Cole, Oisin Mac Aodha, Christian Igel, Serge J. Belongie
CVPR6
2024 MMEarth: Exploring Multi-modal Pretext Tasks for Geospatial Representation Learning
Vishal Nedungadi, Ankit Kariryaa, Stefan Oehmcke, Serge J. Belongie, Christian Igel, Nico Lang
ECCV (64)4
2024 Labeled Data Selection for Category Discovery
Bingchen Zhao, Nico Lang, Serge J. Belongie, Oisin Mac Aodha
ECCV (55)3
2024 Coarse-To-Fine Tensor Trains for Compact Visual Representations
abstract
The ability to learn compact, high-quality, and easy-to-optimize representations for visual data is paramount to many applications such as novel view synthesis and 3D reconstruction. Recent work has shown substantial success in using tensor networks to design such compact and high-quality representations. However, the ability to optimize tensor-based representations, and in particular, the highly compact tensor train representation, is still lacking. This has prevented practitioners from deploying the full potential of tensor networks for visual data. To this end, we propose ’Prolongation Upsampling Tensor Train (PuTT)’, a novel method for learning tensor train representations in a coarse-to-fine manner. Our method involves the prolonging or ‘upsampling’ of a learned tensor train representation, creating a sequence of ’coarse-to-fine’ tensor trains that are incrementally refined. We evaluate our representation along three axes: (1). compression, (2). denoising capability, and (3). image completion capability. To assess these axes, we consider the tasks of image fitting, 3D fitting, and novel view synthesis, where our method shows an improved performance compared to state-of-the-art tensor-based methods.
Sebastian Loeschcke, Dan Wang 0011, Christian Leth-Espensen, Serge J. Belongie, Michael J. Kastoryano, Sagie Benaim
ICML4
2024 LoQT: Low-Rank Adapters for Quantized Pretraining
abstract
Despite advances using low-rank adapters and quantization, pretraining of large models on consumer hardware has not been possible without model sharding, offloading during training, or per-layer gradient updates. To address these limitations, we propose Low-Rank Adapters for Quantized Training (LoQT), a method for efficiently training quantized models. LoQT uses gradient-based tensor factorization to initialize low-rank trainable weight matrices that are periodically merged into quantized full-rank weight matrices. Our approach is suitable for both pretraining and fine-tuning models. We demonstrate this for language modeling and downstream task adaptation, finding that LoQT enables efficient training of models up to 7B parameters on a 24GB GPU. We also demonstrate the feasibility of training a 13B model using per-layer gradient updates on the same hardware.
Sebastian Loeschcke, Mads Bech Toftrup, Michael J. Kastoryano, Serge J. Belongie, Vésteinn Snæbjarnarson
NeurIPS4
2024 Volumetric Disentanglement for 3D Scene Manipulation
abstract
Recently, advances in differential volumetric rendering enabled significant breakthroughs in the photo-realistic and fine-detailed reconstruction of complex 3D scenes, which is key for many virtual reality applications. However, in the context of augmented reality, one may also wish to effect semantic manipulations or augmentations of objects within a scene. To this end, we propose a volumetric framework for (i) disentangling or separating, the volumetric representation of a given foreground object from the background, and (ii) semantically manipulating the foreground object, as well as the background. Our framework takes as input a set of 2D masks specifying the desired foreground object for training views, together with the associated 2D views and poses, and produces a foreground-background disentanglement that respects the surrounding illumination, reflections, and partial occlusions, which can be applied to both training and novel views. Our method enables the separate control of pixel color and depth as well as 3D similarity transformations of both the foreground and background objects. We subsequently demonstrate our framework’s applicability on several downstream manipulation tasks, going beyond the placement and movement of foreground objects. These tasks include object camouflage, non-negative 3D object in-painting, 3D object translation, 3D object inpainting, and 3D text-based object manipulation. The project webpage is provided in https://sagiebenaim.github.io/volumetric-disentanglement/.
Sagie Benaim, Frederik Warburg, Peter Ebert Christensen, Serge J. Belongie
WACV4
2024 Assessing Neural Network Robustness via Adversarial Pivotal Tuning
abstract
The robustness of image classifiers is essential to their deployment in the real world. The ability to assess this resilience to manipulations or deviations from the training data is thus crucial. These modifications have traditionally consisted of minimal changes that still manage to fool classifiers, and modern approaches are increasingly robust to them. Semantic manipulations that modify elements of an image in meaningful ways have thus gained traction for this purpose. However, they have primarily been limited to style, color, or attribute changes. While expressive, these manipulations do not make use of the full capabilities of a pretrained generative model. In this work, we aim to bridge this gap. We show how a pretrained image generator can be used to semantically manipulate images in a detailed, diverse, and photorealistic way while still preserving the class of the original image. Inspired by recent GAN-based image inversion methods, we propose a method called Adversarial Pivotal Tuning (APT). Given an image, APT first finds a pivot latent space input that reconstructs the image using a pretrained generator. It then adjusts the generator’s weights to create small yet semantic manipulations in order to fool a pretrained classifier. APT preserves the full expressive editing capabilities of the generative model. We demonstrate that APT is capable of a wide range of class-preserving semantic image manipulations that fool a variety of pretrained classifiers. Finally, we show that classifiers that are robust to other benchmarks are not robust to APT manipulations and suggest a method to improve them.
Peter Ebert Christensen, Vésteinn Snæbjarnarson, Andrea Dittadi, Serge J. Belongie, Sagie Benaim
WACV4
2023 Discriminative Class Tokens for Text-to-Image Diffusion Models
abstract
Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in the input text. One way of alleviating these issues is to train diffusion models on class-labeled datasets. This approach has two disadvantages: (i) supervised datasets are generally small compared to large-scale scraped text-image datasets on which text-to-image models are trained, affecting the quality and diversity of the generated images, or (ii) the input is a hard-coded label, as opposed to free-form text, limiting the control over the generated images.In this work, we propose a non-invasive fine-tuning technique that capitalizes on the expressive potential of freeform text while achieving high accuracy through discriminative signals from a pretrained classifier. This is done by iteratively modifying the embedding of an added input token of a text-to-image diffusion model, by steering generated images toward a given target class according to a classifier. Our method is fast compared to prior fine-tuning methods and does not require a collection of in-class images or retraining of a noise-tolerant classifier. We evaluate our method extensively, showing that the generated images are: (i) more accurate and of higher quality than standard diffusion models, (ii) can be used to augment training data in a low-resource setting, and (iii) reveal information about the data used to train the guiding classifier. The code is available at https://github.com/idansc/discriminative_class_tokens.
Idan Schwartz, Vésteinn Snæbjarnarson, Hila Chefer, Serge J. Belongie, Lior Wolf, Sagie Benaim
ICCV4
2023 Descriptive Attributes for Language-Based Object Keypoint Detection
Jerod J. Weinman, Serge J. Belongie, Stella Frank
ICVS2
2023 Learning to Taste: A Multimodal Wine Dataset
abstract
We present WineSensed, a large multimodal wine dataset for studying the relations between visual perception, language, and flavor. The dataset encompasses 897k images of wine labels and 824k reviews of wines curated from the Vivino platform. It has over 350k unique vintages, annotated with year, region, rating, alcohol percentage, price, and grape composition. We obtained fine-grained flavor annotations on a subset by conducting a wine-tasting experiment with 256 participants who were asked to rank wines based on their similarity in flavor, resulting in more than 5k pairwise flavor distances. We propose a low-dimensional concept embedding algorithm that combines human experience with automatic machine similarity kernels. We demonstrate that this shared concept embedding space improves upon separate embedding spaces for coarse flavor classification (alcohol percentage, country, grape, price, rating) and representing human perception of flavor.
Thoranna Bender, Simon Møe Sørensen, Alireza Kashani, Kristjan Eldjarn Hjorleifsson, Grethe Hyldig, Søren Hauberg, Serge J. Belongie, Frederik Warburg
NeurIPS7
2022 On Temporal Granularity in Self-Supervised Video Representation Learning
Rui Qian 0003, Yeqing Li, Liangzhe Yuan, Boqing Gong, Ting Liu 0005, Matthew Brown 0001, Serge J. Belongie, Ming-Hsuan Yang 0001, Hartwig Adam, Yin Cui
BMVC7
2022 When Does Contrastive Visual Representation Learning Work?
abstract
Recent self-supervised representation learning techniques have largely closed the gap between supervised and unsupervised learning on ImageNet classification. While the particulars of pretraining on ImageNet are now relatively well understood, the field still lacks widely accepted best practices for replicating this success on other datasets. As a first step in this direction, we study contrastive self-supervised learning on four diverse large-scale datasets. By looking through the lenses of data quantity, data domain, data quality, and task granularity, we provide new insights into the necessary conditions for successful self-supervised learning. Our key findings include observations such as: (i) the benefit of additional pretraining data beyond 500k images is modest, (ii) adding pretraining images from another domain does not lead to more general representations, (iii) corrupted pretraining images have a disparate impact on supervised and self-supervised pretraining, and (iv) contrastive learning lags far behind supervised learning on finegrained visual classification tasks.
Elijah Cole, Kimberly Wilber, Oisin Mac Aodha, Serge J. Belongie
CVPR5
2022 On Label Granularity and Object Localization
Elijah Cole, Kimberly Wilber, Grant Van Horn, Marco Fornoni, Pietro Perona, Serge J. Belongie, Andrew G. Howard, Oisin Mac Aodha
ECCV (10)7
2022 Exploring Fine-Grained Audiovisual Categorization with the SSW60 Dataset
Grant Van Horn, Rui Qian 0003, Kimberly Wilber, Hartwig Adam, Oisin Mac Aodha, Serge J. Belongie
ECCV (8)6
2022 Visual Prompt Tuning
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge J. Belongie, Bharath Hariharan, Ser-Nam Lim
ECCV (33)5
2022 Language-driven Semantic Segmentation
Boyi Li 0001, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun, René Ranftl
ICLR3
2022 Implicit Neural Representations with Levels-of-Experts
abstract
Coordinate-based networks, usually in the forms of MLPs, have been successfully applied to the task of predicting high-frequency but low-dimensional signals using coordinate inputs. To scale them to model large-scale signals, previous works resort to hybrid representations, combining a coordinate-based network with a grid-based representation, such as sparse voxels. However, such approaches lack a compact global latent representation in its grid, making it difficult to model a distribution of signals, which is important for generalization tasks. To address the limitation, we propose the Levels-of-Experts (LoE) framework, which is a novel coordinate-based representation consisting of an MLP with periodic, position-dependent weights arranged hierarchically. For each linear layer of the MLP, multiple candidate values of its weight matrix are tiled and replicated across the input space, with different layers replicating at different frequencies. Based on the input, only one of the weight matrices is chosen for each layer. This greatly increases the model capacity without incurring extra computation or compromising generalization capability. We show that the new representation is an efficient and competitive drop-in replacement for a wide range of tasks, including signal fitting, novel view synthesis, and generative modeling.
Zekun Hao, Arun Mallya, Serge J. Belongie, Ming-Yu Liu 0001
NeurIPS3
2022 Polynomial Neural Fields for Subband Decomposition and Manipulation
abstract
Neural fields have emerged as a new paradigm for representing signals, thanks to their ability to do it compactly while being easy to optimize. In most applications, however, neural fields are treated like a black box, which precludes many signal manipulation tasks. In this paper, we propose a new class of neural fields called basis-encoded polynomial neural fields (PNFs). The key advantage of a PNF is that it can represent a signal as a composition of a number of manipulable and interpretable components without losing the merits of neural fields representation. We develop a general theoretical framework to analyze and design PNFs. We use this framework to design Fourier PNFs, which match state-of-the-art performance in signal representation tasks that use neural fields. In addition, we empirically demonstrate that Fourier PNFs enable signal manipulation applications such as texture transfer and scale-space interpolation. Code is available at https://github.com/stevenygd/PNF.
Guandao Yang, Sagie Benaim, Varun Jampani, Kyle Genova, Jonathan T. Barron, Thomas A. Funkhouser, Bharath Hariharan, Serge J. Belongie
NeurIPS8
2022 Occluded Video Instance Segmentation: A Benchmark
abstract
Abstract Can our video understanding systems perceive objects when a heavy occlusion exists in a scene? To answer this question, we collect a large-scale dataset called OVIS for occluded video instance segmentation, that is, to simultaneously detect, segment, and track instances in occluded scenes. OVIS consists of 296k high-quality instance masks from 25 semantic categories, where object occlusions usually occur. While our human vision systems can understand those occluded instances by contextual reasoning and association, our experiments suggest that current video understanding systems cannot. On the OVIS dataset, the highest AP achieved by state-of-the-art algorithms is only 16.3, which reveals that we are still at a nascent stage for understanding objects, instances, and videos in a real-world scenario. We also present a simple plug-and-play module that performs temporal feature calibration to complement missing object cues caused by occlusion. Built upon MaskTrack R-CNN and SipMask, we obtain a remarkable AP improvement on the OVIS dataset. The OVIS dataset and project code are available at http://songbai.site/ovis .
Jiyang Qi, Yan Gao 0017, Yao Hu 0002, Xinggang Wang, Xiang Bai, Serge J. Belongie, Alan L. Yuille, Philip Torr 0001, Song Bai 0001
Int. J. Comput. Vis.7
2022 Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges
abstract
In he past decade, object detection has achieved significant progress in natural images but not in aerial images, due to the massive variations in the scale and orientation of objects caused by the bird's-eye view of aerial images. More importantly, the lack of large-scale benchmarks has become a major obstacle to the development of object detection in aerial images (ODAI). In this paper, we present a large-scale Dataset of Object deTection in Aerial images (DOTA) and comprehensive baselines for ODAI. The proposed DOTA dataset contains 1,793,658 object instances of 18 categories of oriented-bounding-box annotations collected from 11,268 aerial images. Based on this large-scale and well-annotated dataset, we build baselines covering 10 state-of-the-art algorithms with over 70 configurations, where the speed and accuracy performances of each model have been evaluated. Furthermore, we provide a code library for ODAI and build a website for evaluating different algorithms. Previous challenges run on DOTA have attracted more than 1300 teams worldwide. We believe that the expanded large-scale DOTA dataset, the extensive baselines, the code library and the challenges can facilitate the designs of robust algorithms and reproducible research on the problem of object detection in aerial images.
Jian Ding 0001, Nan Xue 0001, Gui-Song Xia, Xiang Bai, Wen Yang 0001, Michael Ying Yang, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 Guest Editorial: Introduction to the Special Section on Fine-Grained Visual Categorization
abstract
This special section on fine-grained visual categorization has attracted many research works on fine-grain related topics. We thank all authors for submitting their papers to the special section and all reviewers who have provided professional, insightful, and timely reviews, leading to the high quality of accepted papers. We also thank TPAMI EIC Sven Dickinson and the Associate EICs for recognizing the widespread interest in this field, which warrants this special section. The accepted papers are divided into four groups based on their different focuses: Fine-grained image recognition, Fine-grained human analysis, Fine-grained video action recognition, and Fine-grained vision-language reasoning. We briefly review the accepted papers in each of the groups.
Jingdong Wang 0001, Zhuowen Tu, Jianlong Fu, Nicu Sebe, Serge J. Belongie
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Fine-Grained Image Analysis With Deep Learning: A Survey
abstract
Fine-grained image analysis (FGIA) is a longstanding and fundamental problem in computer vision and pattern recognition, and underpins a diverse set of real-world applications. The task of FGIA targets analyzing visual objects from subordinate categories, e.g., species of birds or models of cars. The small inter-class and large intra-class variation inherent to fine-grained image analysis makes it a challenging problem. Capitalizing on advances in deep learning, in recent years we have witnessed remarkable progress in deep learning powered FGIA. In this paper we present a systematic survey of these advances, where we attempt to re-define and broaden the field of FGIA by consolidating two fundamental fine-grained research areas - fine-grained image recognition and fine-grained image retrieval. In addition, we also review other key issues of FGIA, such as publicly available benchmark datasets and related domain-specific applications. We conclude by highlighting several research directions and open problems which need further exploration from the community.
Xiu-Shen Wei, Yi-Zhe Song, Oisin Mac Aodha, Jianxin Wu 0001, Yuxin Peng 0001, Jinhui Tang 0001, Jian Yang 0003, Serge J. Belongie
IEEE Trans. Pattern Anal. Mach. Intell.8
2021 Benchmarking Representation Learning for Natural World Image Collections
abstract
Recent progress in self-supervised learning has resulted in models that are capable of extracting rich representations from image collections without requiring any explicit label supervision. However, to date the vast majority of these approaches have restricted themselves to training on standard benchmark datasets such as ImageNet. We argue that fine-grained visual categorization problems, such as plant and animal species classification, provide an informative testbed for self-supervised learning. In order to facilitate progress in this area we present two new natural world visual classification datasets, iNat2021 and NeWT. The former consists of 2.7M images from 10k different species up-loaded by users of the citizen science application iNaturalist. We designed the latter, NeWT, in collaboration with domain experts with the aim of benchmarking the performance of representation learning algorithms on a suite of challenging natural world binary classification tasks that go beyond standard species classification. These two new datasets allow us to explore questions related to large-scale representation and transfer learning in the context of fine-grained categories. We provide a comprehensive analysis of feature extractors trained with and without supervision on ImageNet and iNat2021, shedding light on the strengths and weaknesses of different learned features across a diverse set of tasks. We find that features produced by standard supervised methods still outperform those produced by self-supervised approaches such as SimCLR. However, improved self-supervised learning methods are constantly being released and the iNat2021 and NeWT datasets are a valuable resource for tracking their progress.
Grant Van Horn, Elijah Cole, Sara Beery, Kimberly Wilber, Serge J. Belongie, Oisin Mac Aodha
CVPR5
2021 Intentonomy: A Dataset and Study Towards Human Intent Understanding
abstract
An image is worth a thousand words, conveying information that goes beyond the mere visual content therein. In this paper, we study the intent behind social media images with an aim to analyze how visual information can facilitate recognition of human intent. Towards this goal, we introduce an intent dataset, Intentonomy, comprising 14K images covering a wide range of everyday scenes. These images are manually annotated with 28 intent categories derived from a social psychology taxonomy. We then systematically study whether, and to what extent, commonly used visual information, i.e., object and context, contribute to human motive understanding. Based on our findings, we conduct further study to quantify the effect of attending to object and context classes as well as textual information in the form of hashtags when training an intent classifier. Our results quantitatively and qualitatively shed light on how visual and textual information can produce observable effects when predicting intent.1
Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge J. Belongie, Ser-Nam Lim
CVPR5
2021 On Feature Normalization and Data Augmentation
abstract
The moments (a.k.a., mean and standard deviation) of latent features are often removed as noise when training image recognition models, to increase stability and reduce training time. However, in the field of image generation, the moments play a much more central role. Studies have shown that the moments extracted from instance normalization and positional normalization can roughly capture style and shape information of an image. Instead of being discarded, these moments are instrumental to the generation process. In this paper we propose Moment Exchange, an implicit data augmentation method that encourages the model to utilize the moment information also for recognition models. Specifically, we replace the moments of the learned features of one training image by those of another, and also interpolate the target labels—forcing the model to extract training signal from the moments in addition to the normalized features. As our approach is fast, operates entirely in feature space, and mixes different signals than prior methods, one can effectively combine it with existing augmentation approaches. We demonstrate its efficacy across several recognition benchmark data sets where it improves the generalization capability of highly competitive baseline networks with remarkable consistency.
Boyi Li 0001, Felix Wu, Ser-Nam Lim, Serge J. Belongie, Kilian Q. Weinberger
CVPR4
2021 Stay Positive: Non-Negative Image Synthesis for Augmented Reality
abstract
In applications such as optical see-through and projector augmented reality, producing images amounts to solving non-negative image generation, where one can only add light to an existing image. Most image generation methods, however, are ill-suited to this problem setting, as they make the assumption that one can assign arbitrary color to each pixel. In fact, naive application of existing methods fails even in simple domains such as MNIST digits, since one cannot create darker pixels by adding light. We know, however, that the human visual system can be fooled by optical illusions involving certain spatial configurations of brightness and contrast. Our key insight is that one can leverage this behavior to produce high quality images with negligible artifacts. For example, we can create the illusion of darker patches by brightening surrounding pixels. We propose a novel optimization procedure to produce images that satisfy both semantic and non-negativity constraints. Our approach can incorporate existing state-of-the-art methods, and exhibits strong performance in a variety of tasks including image-to-image translation and style transfer.
Katie Luo, Guandao Yang, Wenqi Xian, Harald Haraldsson, Bharath Hariharan, Serge J. Belongie
CVPR6
2021 Spatiotemporal Contrastive Video Representation Learning
abstract
We present a self-supervised Contrastive Video Representation Learning (CVRL) method to learn spatiotemporal visual representations from unlabeled videos. Our representations are learned using a contrastive loss, where two augmented clips from the same short video are pulled together in the embedding space, while clips from different videos are pushed away. We study what makes for good data augmentations for video self-supervised learning and find that both spatial and temporal information are crucial. We carefully design data augmentations involving spatial and temporal cues. Concretely, we propose a temporally consistent spatial augmentation method to impose strong spatial augmentations on each frame of the video while maintaining the temporal consistency across frames. We also propose a sampling-based temporal augmentation method to avoid overly enforcing invariance on clips that are distant in time. On Kinetics-600, a linear classifier trained on the representations learned by CVRL achieves 70.4% top-1 accuracy with a 3D-ResNet-50 (R3D-50) backbone, outperforming ImageNet supervised pre-training by 15.7% and SimCLR unsupervised pre-training by 18.8% using the same inflated R3D-50. The performance of CVRL can be further improved to 72.9% with a larger R3D-152 (2× filters) backbone, significantly closing the gap between unsupervised and supervised video representation learning. Our code and models will be available at https://github.com/tensorflow/models/tree/master/official/.
Rui Qian 0003, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang 0001, Huisheng Wang, Serge J. Belongie, Yin Cui
CVPR6
2021 GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds
abstract
We present GANcraft, an unsupervised neural rendering framework for generating photorealistic images of large 3D block worlds such as those created in Minecraft. Our method takes a semantic block world as input, where each block is assigned a semantic label such as dirt, grass, or water. We represent the world as a continuous volumetric function and train our model to render view-consistent photorealistic images for a user-controlled camera. In the absence of paired ground truth real images for the block world, we devise a training technique based on pseudo-ground truth and adversarial training. This stands in contrast to prior work on neural rendering for view synthesis, which requires ground truth images to estimate scene geometry and view-dependent appearance. In addition to camera trajectory, GANcraft allows user control over both scene semantics and output style. Experimental results with comparison to strong baselines show the effectiveness of GANcraft on this novel task of photorealistic 3D block world synthesis. The project website is available at https://nvlabs.github.io/GANcraft/.
Zekun Hao, Arun Mallya, Serge J. Belongie, Ming-Yu Liu 0001
ICCV3
2021 Exploring Visual Engagement Signals for Representation Learning
abstract
Visual engagement in social media platforms comprises interactions with photo posts including comments, shares, and likes. In this paper, we leverage such Visual Engagement clues as supervisory signals for representation learning. However, learning from engagement signals is non-trivial as it is not clear how to bridge the gap between low-level visual information and high-level social interactions. We present VisE, a weakly supervised learning approach, which maps social images to pseudo labels derived by clustered engagement signals. We then study how models trained in this way benefit subjective downstream computer vision tasks such as emotion recognition or political bias detection. Through extensive studies, we empirically demonstrate the effectiveness of VisE across a diverse set of classification tasks beyond the scope of conventional recognition1.
Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge J. Belongie, Ser-Nam Lim
ICCV5
2021 Robustness and Generalization via Generative Adversarial Training
abstract
While deep neural networks have achieved remarkable success in various computer vision tasks, they often fail to generalize to new domains and subtle variations of input images. Several defenses have been proposed to improve the robustness against these variations. However, current defenses can only withstand the specific attack used in training, and the models often remain vulnerable to other input variations. Moreover, these methods often degrade performance of the model on clean images and do not generalize to out-of-domain samples. In this paper we present Generative Adversarial Training, an approach to simultaneously improve the model’s generalization to the test set and out-of-domain samples as well as its robustness to unseen adversarial attacks. Instead of altering a low-level pre-defined aspect of images, we generate a spectrum of low-level, mid-level and high-level changes using generative models with a disentangled latent space. Adversarial training with these examples enable the model to withstand a wide range of attacks by observing a variety of input alterations during training. We show that our approach not only improves performance of the model on clean images and out-of-domain samples but also makes it robust against unforeseen attacks and outperforms prior work. We validate effectiveness of our method by demonstrating results on various tasks such as classification, segmentation and object detection.
Omid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie, Ser-Nam Lim
ICCV4
2021 Geometry Processing with Neural Fields
abstract
Most existing geometry processing algorithms use meshes as the default shape representation. Manipulating meshes, however, requires one to maintain high quality in the surface discretization. For example, changing the topology of a mesh usually requires additional procedures such as remeshing. This paper instead proposes the use of neural fields for geometry processing. Neural fields can compactly store complicated shapes without spatial discretization. Moreover, neural fields are infinitely differentiable, which allows them to be optimized for objectives that involve higher-order derivatives. This raises the question: can geometry processing be done entirely using neural fields? We introduce loss functions and architectures to show that some of the most challenging geometry processing tasks, such as deformation and filtering, can be done with neural fields. Experimental results show that our methods are on par with the well-established mesh-based methods without committing to a particular surface discretization. Code is available at https://github.com/stevenygd/NFGP.
Guandao Yang, Serge J. Belongie, Bharath Hariharan, Vladlen Koltun
NeurIPS2
2020 DualSDF: Semantic Shape Manipulation Using a Two-Level Representation
abstract
We are seeing a Cambrian explosion of 3D shape representations for use in machine learning. Some representations seek high expressive power in capturing high-resolution detail. Other approaches seek to represent shapes as compositions of simple parts, which are intuitive for people to understand and easy to edit and manipulate. However, it is difficult to achieve both fidelity and interpretability in the same representation. We propose DualSDF, a representation expressing shapes at two levels of granularity, one capturing fine details and the other representing an abstracted proxy shape using simple and semantically consistent shape primitives. To achieve a tight coupling between the two representations, we use a variational objective over a shared latent space. Our two-level model gives rise to a new shape manipulation technique in which a user can interactively manipulate the coarse proxy shape and see the changes instantly mirrored in the high-resolution shape. Moreover, our model actively augments and guides the manipulation towards producing semantically meaningful shapes, making complex manipulations possible with minimal user input.
Zekun Hao, Hadar Averbuch-Elor, Noah Snavely, Serge J. Belongie
CVPR4
2020 End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection
abstract
Reliable and accurate 3D object detection is a necessity for safe autonomous driving. Although LiDAR sensors can provide accurate 3D point cloud estimates of the environment, they are also prohibitively expensive for many settings. Recently, the introduction of pseudo-LiDAR (PL) has led to a drastic reduction in the accuracy gap between methods based on LiDAR sensors and those based on cheap stereo cameras. PL combines state-of-the-art deep neural networks for 3D depth estimation with those for 3D object detection by converting 2D depth map outputs to 3D point cloud inputs. However, so far these two networks have to be trained separately. In this paper, we introduce a new framework based on differentiable Change of Representation (CoR) modules that allow the entire PL pipeline to be trained end-to-end. The resulting framework is compatible with most state-of-the-art networks for both tasks and in combination with PointRCNN improves over PL consistently across all benchmarks --- yielding the highest entry on the KITTI image-based 3D object detection leaderboard at the time of submission. Our code will be made available at https://github.com/mileyan/pseudo-LiDAR_e2e.
Rui Qian 0003, Divyansh Garg, Yan Wang 0051, Yurong You, Serge J. Belongie, Bharath Hariharan, Mark E. Campbell, Kilian Q. Weinberger, Wei-Lun Chao
CVPR5
2020 Learning Gradient Fields for Shape Generation
Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge J. Belongie, Noah Snavely, Bharath Hariharan
ECCV (3)5
2020 Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
Menglin Jia, Mengyun Shi, Mikhail Sirotenko, Yin Cui, Claire Cardie, Bharath Hariharan, Hartwig Adam, Serge J. Belongie
ECCV (1)8
2020 A Metric Learning Reality Check
Kevin Musgrave, Serge J. Belongie, Ser-Nam Lim
ECCV (25)2
2020 Differentiating through the Fréchet Mean
abstract
Recent advances in deep representation learning on Riemannian manifolds extend classical deep learning operations to better capture the geometry of the manifold. One possible extension is the Fr{é}chet mean, the generalization of the Euclidean mean; however, it has been difficult to apply because it lacks a closed form with an easily computable derivative. In this paper, we show how to differentiate through the Fr{é}chet mean for arbitrary Riemannian manifolds. Then, focusing on hyperbolic space, we derive explicit gradient expressions and a fast, accurate, and hyperparameter-free Fr{é}chet mean solver. This fully integrates the Fr{é}chet mean into the hyperbolic neural network pipeline. To demonstrate this integration, we present two case studies. First, we apply our Fr{é}chet mean to the existing Hyperbolic Graph Convolutional Network, replacing its projected aggregation to obtain state-of-the-art results on datasets with high hyperbolicity. Second, to demonstrate the Fr{é}chet mean’s capacity to generalize Euclidean neural network operations, we develop a hyperbolic batch normalization method that gives an improvement parallel to the one observed in the Euclidean setting.
Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge J. Belongie, Ser-Nam Lim, Christopher De Sa
ICML4
2020 Neural Puppet: Generative Layered Cartoon Characters
abstract
We propose a learning based method for generating new animations of a cartoon character given a few example images. Our method is designed to learn from a traditionally animated sequence, where each frame is drawn by an artist, and thus the input images lack any common structure, correspondences, or labels. We express pose changes as a deformation of a layered 2.5D template mesh, and devise a novel architecture that learns to predict mesh deformations matching the template to a target image. This enables us to extract a common low-dimensional structure from a diverse set of character poses. We combine recent advances in differentiable rendering as well as mesh-aware models to successfully align common template even if only a few character images are available during training. In addition to coarse poses, character appearance also varies due to shading, out-of-plane motions, and artistic effects. We capture these subtle changes by applying an image translation network to refine the mesh rendering, providing an end-to-end model to generate new animations of a character with high visual quality. We demonstrate that our generative model can be used to synthesize in-between frames and to create data-driven deformation. Our template fitting procedure outperforms state-of-the-art generic techniques for detecting image correspondences.
Omid Poursaeed, Vladimir G. Kim, Eli Shechtman, Jun Saito, Serge J. Belongie
WACV5
2020 Convolutional Networks with Adaptive Inference Graphs
Andreas Veit, Serge J. Belongie
Int. J. Comput. Vis.2
2020 Transfer learning in computer vision tasks: Remember where you come from
Xuhong Li 0002, Yves Grandvalet, Franck Davoine, Jingchun Cheng, Yin Cui, Han Zhang 0005, Serge J. Belongie, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
Image Vis. Comput.7
2019 Class-Balanced Loss Based on Effective Number of Samples
abstract
With the rapid increase of large-scale, real-world datasets, it becomes critical to address the problem of long-tailed data distribution (i.e., a few classes account for most of the data, while most classes are under-represented). Existing solutions typically adopt class re-balancing strategies such as re-sampling and re-weighting based on the number of observations for each class. In this work, we argue that as the number of samples increases, the additional benefit of a newly added data point will diminish. We introduce a novel theoretical framework to measure data overlap by associating with each sample a small neighboring region rather than a single point. The effective number of samples is defined as the volume of samples and can be calculated by a simple-formula (1-βn)/(1-β), where n is the number of samples and β ∈ [0, 1) is a hyperparameter. We design a re-weighting scheme that uses the effective number of samples for each class to re-balance the loss, thereby yielding a class-balanced loss. Comprehensive experiments are conducted on artificially induced long-tailed CIFAR datasets and large-scale datasets including ImageNet and iNaturalist. Our results show that when trained with the proposed class-balanced loss, the network is able to achieve significant performance gains on long-tailed datasets.
Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song 0009, Serge J. Belongie
CVPR5
2019 Neural Naturalist: Generating Fine-Grained Image Comparisons
abstract
Maxwell Forbes, Christine Kaeser-Chen, Piyush Sharma, Serge Belongie. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Maxwell Forbes, Christine Kaeser-Chen, Piyush Sharma, Serge J. Belongie
EMNLP/IJCNLP (1)4
2019 Enhancing Adversarial Example Transferability With an Intermediate Level Attack
abstract
Neural networks are vulnerable to adversarial examples, malicious inputs crafted to fool trained models. Adversarial examples often exhibit black-box transfer, meaning that adversarial examples for one model can fool another model. However, adversarial examples are typically overfit to exploit the particular architecture and feature representation of a source model, resulting in sub-optimal black-box transfer attacks to other target models. We introduce the Intermediate Level Attack (ILA), which attempts to fine-tune an existing adversarial example for greater black-box transferability by increasing its perturbation on a pre-specified layer of the source model, improving upon state-of-the-art methods. We show that we can select a layer of the source model to perturb without any knowledge of the target models while achieving high transferability. Additionally, we provide some explanatory insights regarding our method and the effect of optimizing for adversarial examples using intermediate feature maps.
Isay Katsman, Zeqi Gu, Horace He, Serge J. Belongie, Ser-Nam Lim
ICCV5
2019 PointFlow: 3D Point Cloud Generation With Continuous Normalizing Flows
abstract
As 3D point clouds become the representation of choice for multiple vision and graphics applications, the ability to synthesize or reconstruct high-resolution, high-fidelity point clouds becomes crucial. Despite the recent success of deep learning models in discriminative tasks of point clouds, generating point clouds remains challenging. This paper proposes a principled probabilistic framework to generate 3D point clouds by modeling them as a distribution of distributions. Specifically, we learn a two-level hierarchy of distributions where the first level is the distribution of shapes and the second level is the distribution of points given a shape. This formulation allows us to both sample shapes and sample an arbitrary number of points from a shape. Our generative model, named PointFlow, learns each level of the distribution with a continuous normalizing flow. The invertibility of normalizing flows enables the computation of the likelihood during training and allows us to train our model in the variational inference framework. Empirically, we demonstrate that PointFlow achieves state-of-the-art performance in point cloud generation. We additionally show that our model can faithfully reconstruct point clouds and learn useful representations in an unsupervised manner. The code is available at https://github.com/stevenygd/PointFlow.
Guandao Yang, Xun Huang 0002, Zekun Hao, Ming-Yu Liu 0001, Serge J. Belongie, Bharath Hariharan
ICCV5
2019 Positional Normalization
abstract
A widely deployed method for reducing the training time of deep neural networks is to normalize activations at each layer. Although various normalization schemes have been proposed, they all follow a common theme: normalize across spatial dimensions and discard the extracted statistics. In this paper, we propose a novel normalization method that deviates from this theme. Our approach, which we refer to as Positional Normalization (PONO), normalizes exclusively across channels, which allows us to capture structural information of the input image in the first and second moments. Instead of disregarding this information, we inject it into later layers to preserve or transfer structural information in generative networks. We show that PONO significantly improves the performance of deep networks across a wide range of model architectures and image generation tasks.
Boyi Li 0001, Felix Wu, Kilian Q. Weinberger, Serge J. Belongie
NeurIPS4
2019 Understanding Image Quality and Trust in Peer-to-Peer Marketplaces
abstract
As any savvy online shopper knows, second-hand peer-to-peer marketplaces are filled with images of mixed quality. How does image quality impact marketplace outcomes, and can quality be automatically predicted? In this work, we conducted a large-scale study on the quality of user-generated images in peer-to-peer marketplaces. By gathering a dataset of common second-hand products (≈75,000 images) and annotating a subset with human-labeled quality judgments, we were able to model and predict image quality with decent accuracy (≈87%). We then conducted two studies focused on understanding the relationship between these image quality scores and two marketplace outcomes: sales and perceived trustworthiness. We show that image quality is associated with higher likelihood that an item will be sold, though other factors such as view count were better predictors of sales. Nonetheless, we show that high quality user-generated images selected by our models outperform stock imagery in eliciting perceptions of trust from users. Our findings can inform the design of future marketplaces and guide potential sellers to take better product images.
Xiao Ma 0010, Lina Mezghani, Kimberly Wilber, Hui Hong, Robinson Piramuthu, Mor Naaman, Serge J. Belongie
WACV7
2018 Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning
abstract
Transferring the knowledge learned from large scale datasets (e.g., ImageNet) via fine-tuning offers an effective solution for domain-specific fine-grained visual categorization (FGVC) tasks (e.g., recognizing bird species or car make & model). In such scenarios, data annotation often calls for specialized domain knowledge and thus is difficult to scale. In this work, we first tackle a problem in large scale FGVC. Our method won first place in iNaturalist 2017 large scale species classification challenge. Central to the success of our approach is a training scheme that uses higher image resolution and deals with the long-tailed distribution of training data. Next, we study transfer learning via fine-tuning from large scale datasets to small scale, domain-specific FGVC datasets. We propose a measure to estimate domain similarity via Earth Mover's Distance and demonstrate that transfer learning benefits from pre-training on a source domain that is similar to the target domain by this measure. Our proposed transfer learning outperforms ImageNet pre-training and obtains state-of-the-art results on multiple commonly used FGVC datasets.
Yin Cui, Yang Song 0009, Chen Sun 0002, Andrew G. Howard, Serge J. Belongie
CVPR5
2018 Learning to Evaluate Image Captioning
abstract
Evaluation metrics for image captioning face two challenges. Firstly, commonly used metrics such as CIDEr, METEOR, ROUGE and BLEU often do not correlate well with human judgments. Secondly, each metric has well known blind spots to pathological caption constructions, and rule-based metrics lack provisions to repair such blind spots once identified. For example, the newly proposed SPICE correlates well with human judgments, but fails to capture the syntactic structure of a sentence. To address these two challenges, we propose a novel learning based discriminative evaluation metric that is directly trained to distinguish between human and machine-generated captions. In addition, we further propose a data augmentation scheme to explicitly incorporate pathological transformations as negative examples during training. The proposed metric is evaluated with three kinds of robustness tests and its correlation with human judgments. Extensive experiments show that the proposed data augmentation scheme not only makes our metric more robust toward several pathological transformations, but also improves its correlation with human judgments. Our metric outperforms other metrics on both caption level human correlation in Flickr 8k and system level human correlation in COCO. The proposed approach could be served as a learning based evaluation metric that is complementary to existing rule-based metrics.
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang 0002, Serge J. Belongie
CVPR5
2018 Controllable Video Generation With Sparse Trajectories
abstract
Video generation and manipulation is an important yet challenging task in computer vision. Existing methods usually lack ways to explicitly control the synthesized motion. In this work, we present a conditional video generation model that allows detailed control over the motion of the generated video. Given the first frame and sparse motion trajectories specified by users, our model can synthesize a video with corresponding appearance and motion. We propose to combine the advantage of copying pixels from the given frame and hallucinating the lightness difference from scratch which help generate sharp video while keeping the model robust to occlusion and lightness change. We also propose a training paradigm that calculate trajectories from video clips, which eliminated the need of annotated training data. Experiments on several standard benchmarks demonstrate that our approach can generate realistic videos comparable to state-of-the-art video generation and video prediction methods while the motion of the generated videos can correspond well with user input.
Zekun Hao, Xun Huang 0002, Serge J. Belongie
CVPR3
2018 The INaturalist Species Classification and Detection Dataset
abstract
Existing image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imbalanced, as some species are more abundant and easier to photograph than others. To encourage further progress in challenging real world conditions we present the iNaturalist species classification and detection dataset, consisting of 859,000 images from over 5,000 different species of plants and animals. It features visually similar species, captured in a wide variety of situations, from all over the world. Images were collected with different camera types, have varying image quality, feature a large class imbalance, and have been verified by multiple citizen scientists. We discuss the collection of the dataset and present extensive baseline experiments using state-of-the-art computer vision classification and detection models. Results show that current non-ensemble based methods achieve only 67% top one classification accuracy, illustrating the difficulty of the dataset. Specifically, we observe poor results for classes with small numbers of training examples suggesting more attention is needed in low-shot learning.
Grant Van Horn, Oisin Mac Aodha, Yang Song 0009, Yin Cui, Chen Sun 0002, Alexander Shepard, Hartwig Adam, Pietro Perona, Serge J. Belongie
CVPR9
2018 Lean Multiclass Crowdsourcing
abstract
We introduce a method for efficiently crowdsourcing multiclass annotations in challenging, real world image datasets. Our method is designed to minimize the number of human annotations that are necessary to achieve a desired level of confidence on class labels. It is based on combining models of worker behavior with computer vision. Our method is general: it can handle a large number of classes, worker labels that come from a taxonomy rather than a flat list, and can model the dependence of labels when workers can see a history of previous annotations. Our method may be used as a drop-in replacement for the majority vote algorithms used in online crowdsourcing services that aggregate multiple human annotations into a final consolidated label. In experiments conducted on two real-life applications we find that our method can reduce the number of required annotations by as much as a factor of 5.4 and can reduce the residual annotation error by up to 90% when compared with majority voting. Furthermore, the online risk estimates of the models may be used to sort the annotated collection and minimize subsequent expert review effort.
Grant Van Horn, Steve Branson, Scott Loarie, Serge J. Belongie, Pietro Perona
CVPR4
2018 Generative Adversarial Perturbations
abstract
In this paper, we propose novel generative models for creating adversarial examples, slightly perturbed images resembling natural images but maliciously crafted to fool pre-trained models. We present trainable deep neural networks for transforming images to adversarial perturbations. Our proposed models can produce image-agnostic and image-dependent perturbations for both targeted and non-targeted attacks. We also demonstrate that similar architectures can achieve impressive results in fooling classification and semantic segmentation models, obviating the need for hand-crafting attack methods for each task. Using extensive experiments on challenging high-resolution datasets such as ImageNet and Cityscapes, we show that our perturbations achieve high fooling rates with small perturbation norms. Moreover, our attacks are considerably faster than current iterative methods at inference time.
Omid Poursaeed, Isay Katsman, Bicheng Gao, Serge J. Belongie
CVPR4
2018 Separating Self-Expression and Visual Content in Hashtag Supervision
abstract
The variety, abundance, and structured nature of hashtags make them an interesting data source for training vision models. For instance, hashtags have the potential to significantly reduce the problem of manual supervision and annotation when learning vision models for a large number of concepts. However, a key challenge when learning from hashtags is that they are inherently subjective because they are provided by users as a form of self-expression. As a consequence, hashtags may have synonyms (different hashtags referring to the same visual content) and may be polysemous (the same hashtag referring to different visual content). These challenges limit the effectiveness of approaches that simply treat hashtags as image-label pairs. This paper presents an approach that extends upon modeling simple image-label pairs with a joint model of images, hashtags, and users. We demonstrate the efficacy of such approaches in image tagging and retrieval experiments, and show how the joint model can be used to perform user-conditional retrieval and tagging.
Andreas Veit, Maximilian Nickel, Serge J. Belongie, Laurens van der Maaten
CVPR3
2018 DOTA: A Large-Scale Dataset for Object Detection in Aerial Images
abstract
Object detection is an important and challenging problem in computer vision. Although the past decade has witnessed major advances in object detection in natural scenes, such successes have been slow to aerial imagery, not only because of the huge variation in the scale, orientation and shape of the object instances on the earth's surface, but also due to the scarcity of well-annotated datasets of objects in aerial scenes. To advance object detection research in Earth Vision, also known as Earth Observation and Remote Sensing, we introduce a large-scale Dataset for Object deTection in Aerial images (DOTA). To this end, we collect 2806 aerial images from different sensors and platforms. Each image is of the size about 4000 × 4000 pixels and contains objects exhibiting a wide variety of scales, orientations, and shapes. These DOTA images are then annotated by experts in aerial image interpretation using 15 common object categories. The fully annotated DOTA images contains 188, 282 instances, each of which is labeled by an arbitrary (8 d.o.f.) quadrilateral. To build a baseline for object detection in Earth Vision, we evaluate state-of-the-art object detection algorithms on DOTA. Experiments demonstrate that DOTA well represents real Earth Vision applications and are quite challenging.
Gui-Song Xia, Xiang Bai, Jian Ding 0001, Zhen Zhu 0006, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001
CVPR5
2018 Multimodal Unsupervised Image-to-Image Translation
Xun Huang 0002, Ming-Yu Liu 0001, Serge J. Belongie, Jan Kautz
ECCV (3)3
2018 Convolutional Networks with Adaptive Inference Graphs
Andreas Veit, Serge J. Belongie
ECCV (1)2
2018 Learning Single-View 3D Reconstruction with Limited Pose Supervision
Guandao Yang, Yin Cui, Serge J. Belongie, Bharath Hariharan
ECCV (15)3
2018 ICPR2018 Contest on Object Detection in Aerial Images (ODAI-18)
abstract
Object detection in aerial images plays a significant role in intelligent interpretation of aerial images. Hence many effective methods, especially the new-generation data-driven methods, have been developed for this task. Here, we hold the ODAI, a new contest that focused on object detection in aerial images, based on a new large-scale aerial image dataset called DOTA [1]. This contest contains over 3000 large-size images ( 4k×4k pixels), which cover 211,581 instances divided into 15 categories. Each instance is labeled by an arbitrary (8 d.o.f.) quadrilateral. Besides, we propose two tasks for this contest, named object detection with the horizontal bounding box (OD-HBB) and object detection with the oriented bounding box (OD-OBB). The contest was opened on February 7, 2018, and ended on April 30, 2018. A website is open to the public, which provides links to download data and evaluation server. We have totally received 60 registrations. There are 8 teams that have successfully submitted results on the OD-HBB task with the top mAP as 0.719, and 9 teams that have successfully submitted results on the OD-OBB task with the top mAP as 0.705. Through the contest, we hope to draw extensive attention from a wide range of communities and call for more future research and efforts for the task of object detection in aerial images.
Jian Ding 0001, Zhen Zhu 0006, Gui-Song Xia, Xiang Bai, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001
ICPR5
2018 Unbiased offline recommender evaluation for missing-not-at-random implicit feedback
abstract
Implicit-feedback Recommenders (ImplicitRec) leverage positive only user-item interactions, such as clicks, to learn personalized user preferences. Recommenders are often evaluated and compared offline using datasets collected from online platforms. These platforms are subject to popularity bias (i.e., popular items are more likely to be presented and interacted with), and therefore logged ground truth data are Missing-Not-At-Random (MNAR). As a result, the widely used Average-Over-All (AOA) evaluator is biased toward accurately recommending trendy items. In this paper, we (a) investigate evaluation bias of AOA and (b) develop an unbiased and practical offline evaluator for implicit MNAR datasets using the Inverse-Propensity-Scoring (IPS) technique. Through extensive experiments using four real-world datasets and four widely used algorithms, we show that (a) popularity bias is widely manifested in item presentation and interaction; (b) evaluation bias due to MNAR data pervasively exists in most cases where AOA is used to evaluate ImplicitRec; and (c) the unbiased estimator significantly reduces the AOA evaluation bias by more than 30% in the Yahoo! music dataset in terms of the Mean Absolute Error (MAE).
Longqi Yang 0001, Yin Cui, Yuan Xuan, Serge J. Belongie, Deborah Estrin
RecSys5
2018 Vision-based real estate price estimation
Omid Poursaeed, Tomas Matera, Serge J. Belongie
Mach. Vis. Appl.3
2017 Kernel Pooling for Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) with Bilinear Pooling, initially in their full form and later using compact representations, have yielded impressive performance gains on a wide range of visual tasks, including fine-grained visual categorization, visual question answering, face recognition, and description of texture and style. The key to their success lies in the spatially invariant modeling of pairwise (2ndorder) feature interactions. In this work, we propose a general pooling framework that captures higher order interactions of features in the form of kernels. We demonstrate how to approximate kernels such as Gaussian RBF up to a given order using compact explicit feature maps in a parameter-free manner. Combined with CNNs, the composition of the kernel can be learned from data in an end-to-end fashion via error back-propagation. The proposed kernel pooling scheme is evaluated in terms of both kernel approximation error and visual recognition accuracy. Experimental evaluations demonstrate state-of-the-art performance on commonly used fine-grained recognition datasets.
Yin Cui, Feng Zhou 0002, Xiao Liu 0022, Yuanqing Lin, Serge J. Belongie
CVPR6
2017 Stacked Generative Adversarial Networks
abstract
In this paper, we propose a novel generative model named Stacked Generative Adversarial Networks (SGAN), which is trained to invert the hierarchical representations of a bottom-up discriminative network. Our model consists of a top-down stack of GANs, each learned to generate lower-level representations conditioned on higher-level representations. A representation discriminator is introduced at each feature hierarchy to encourage the representation manifold of the generator to align with that of the bottom-up discriminative network, leveraging the powerful discriminative representations to guide the generative model. In addition, we introduce a conditional loss that encourages the use of conditional information from the layer above, and a novel entropy loss that maximizes a variational lower bound on the conditional entropy of generator outputs. We first train each stack independently, and then train the whole model end-to-end. Unlike the original GAN that uses a single noise vector to represent all the variations, our SGAN decomposes variations into multiple levels and gradually resolves uncertainties in the top-down generative process. Based on visual inspection, Inception scores and visual Turing test, we demonstrate that SGAN is able to generate images of much higher quality than GANs without stacking.
Xun Huang 0002, Yixuan Li 0001, Omid Poursaeed, John E. Hopcroft, Serge J. Belongie
CVPR5
2017 Feature Pyramid Networks for Object Detection
abstract
Feature pyramids are a basic component in recognition systems for detecting objects at different scales. But pyramid representations have been avoided in recent object detectors that are based on deep convolutional networks, partially because they are slow to compute and memory intensive. In this paper, we exploit the inherent multi-scale, pyramidal hierarchy of deep convolutional networks to construct feature pyramids with marginal extra cost. A top-down architecture with lateral connections is developed for building high-level semantic feature maps at all scales. This architecture, called a Feature Pyramid Network (FPN), shows significant improvement as a generic feature extractor in several applications. Using a basic Faster R-CNN system, our method achieves state-of-the-art single-model results on the COCO detection benchmark without bells and whistles, surpassing all existing single-model entries including those from the COCO 2016 challenge winners. In addition, our method can run at 5 FPS on a GPU and thus is a practical and accurate solution to multi-scale object detection. Code will be made publicly available.
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, Serge J. Belongie
CVPR6
2017 Detecting Oriented Text in Natural Images by Linking Segments
abstract
Most state-of-the-art text detection methods are specific to horizontal Latin text and are not fast enough for real-time applications. We introduce Segment Linking (SegLink), an oriented text detection method. The main idea is to decompose text into two locally detectable elements, namely segments and links. A segment is an oriented box covering a part of a word or text line, A link connects two adjacent segments, indicating that they belong to the same word or text line. Both elements are detected densely at multiple scales by an end-to-end trained, fully-convolutional neural network. Final detections are produced by combining segments connected by links. Compared with previous methods, SegLink improves along the dimensions of accuracy, speed, and ease of training. It achieves an f-measure of 75.0% on the standard ICDAR 2015 Incidental (Challenge 4) benchmark, outperforming the previous best by a large margin. It runs at over 20 FPS on 512x512 images. Moreover, without modification, SegLink is able to detect long lines of non-Latin text, such as Chinese.
Baoguang Shi, Xiang Bai, Serge J. Belongie
CVPR3
2017 Learning from Noisy Large-Scale Datasets with Minimal Supervision
abstract
We present an approach to effectively use millions of images with noisy annotations in conjunction with a small subset of cleanly-annotated images to learn powerful image representations. One common approach to combine clean and noisy data is to first pre-train a network using the large noisy dataset and then fine-tune with the clean dataset. We show this approach does not fully leverage the information contained in the clean set. Thus, we demonstrate how to use the clean annotations to reduce the noise in the large dataset before fine-tuning the network using both the clean set and the full set with reduced noise. The approach comprises a multi-task network that jointly learns to clean noisy annotations and to accurately classify images. We evaluate our approach on the recently released Open Images dataset, containing ~9 million images, multiple annotations per image and over 6000 unique classes. For the small clean set of annotations we use a quarter of the validation set with ~40k images. Our results demonstrate that the proposed approach clearly outperforms direct fine-tuning across all major categories of classes in the Open Image dataset. Further, our approach is particularly effective for a large number of classes with wide range of noise in annotations (20-80% false positive annotations).
Andreas Veit, Neil Alldrin, Gal Chechik, Ivan Krasin, Abhinav Gupta 0001, Serge J. Belongie
CVPR6
2017 Conditional Similarity Networks
abstract
What makes images similar? To measure the similarity between images, they are typically embedded in a feature-vector space, in which their distance preserve the relative dissimilarity. However, when learning such similarity embeddings the simplifying assumption is commonly made that images are only compared to one unique measure of similarity. A main reason for this is that contradicting notions of similarities cannot be captured in a single space. To address this shortcoming, we propose Conditional Similarity Networks (CSNs) that learn embeddings differentiated into semantically distinct subspaces that capture the different notions of similarities. CSNs jointly learn a disentangled embedding where features for different similarities are encoded in separate dimensions as well as masks that select and reweight relevant dimensions to induce a subspace that encodes a specific similarity notion. We show that our approach learns interpretable image representations with visually relevant semantic subspaces. Further, when evaluating on triplet questions from multiple similarity notions our model even outperforms the accuracy obtained by training individual specialized networks for each notion separately.
Andreas Veit, Serge J. Belongie, Theofanis Karaletsos
CVPR2
2017 Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization
abstract
Gatys et al. recently introduced a neural algorithm that renders a content image in the style of another image, achieving so-called style transfer. However, their framework requires a slow iterative optimization process, which limits its practical application. Fast approximations with feed-forward neural networks have been proposed to speed up neural style transfer. Unfortunately, the speed improvement comes at a cost: the network is usually tied to a fixed set of styles and cannot adapt to arbitrary new styles. In this paper, we present a simple yet effective approach that for the first time enables arbitrary style transfer in real-time. At the heart of our method is a novel adaptive instance normalization (AdaIN) layer that aligns the mean and variance of the content features with those of the style features. Our method achieves speed comparable to the fastest existing approach, without the restriction to a pre-defined set of styles. In addition, our approach allows flexible user controls such as content-style trade-off, style interpolation, color & spatial controls, all using a single feed-forward neural network.
Xun Huang 0002, Serge J. Belongie
ICCV2
2017 BAM! The Behance Artistic Media Dataset for Recognition Beyond Photography
abstract
Computer vision systems are designed to work well within the context of everyday photography. However, artists often render the world around them in ways that do not resemble photographs. Artwork produced by people is not constrained to mimic the physical world, making it more challenging for machines to recognize.,,This work is a step toward teaching machines how to categorize images in ways that are valuable to humans. First, we collect a large-scale dataset of contemporary artwork from Behance, a website containing millions of portfolios from professional and commercial artists. We annotate Behance imagery with rich attribute labels for content, emotions, and artistic media. Furthermore, we carry out baseline experiments to show the value of this dataset for artistic style prediction, for improving the generality of existing object classifiers, and for the study of visual domain adaptation. We believe our Behance Artistic Media dataset will be a good starting point for researchers wishing to study artistic imagery and relevant problems. This dataset can be found at https://bam-dataset.org/.
Kimberly Wilber, Hailin Jin, Aaron Hertzmann, John P. Collomosse, Serge J. Belongie
ICCV6
2017 ICDAR2017 Robust Reading Challenge on COCO-Text
abstract
This report presents the final results of the ICDAR 2017 Robust Reading Challenge on COCO-Text. A challenge on scene text detection and recognition based on the largest real scene text dataset currently available: the COCO-Text dataset. The competition is structured around three tasks: Text Localization, Cropped Word Recognition and End-To-End Recognition. The competition received a total of 27 submissions over the different opened tasks. This report describes the datasets and the ground truth, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods.
Raul Gomez, Baoguang Shi, Lluís Gómez i Bigorda, Lukás Neumann, Andreas Veit, Jiri Matas, Serge J. Belongie, Dimosthenis Karatzas
ICDAR7
2017 ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)
abstract
Chinese is the most widely used language in the world. Algorithms that read Chinese text in natural images facilitate applications of various kinds. Despite the large potential value, datasets and competitions in the past primarily focus on English, which bares very different characteristics than Chinese. This report introduces RCTW, a new competition that focuses on Chinese text reading. The competition features a large-scale dataset with over 12,000 annotated images. Two tasks, namely text localization and end-to-end recognition, are set up. The competition took place from January 20 to May 31, 2017. 23 valid submissions were received from 19 teams. This report includes dataset description, task definitions, evaluation protocols, and results summaries and analysis. Through this competition, we call for more future research on the Chinese text reading problem.
Baoguang Shi, Cong Yao, Minghui Liao, Pei Xu 0006, Linyan Cui, Serge J. Belongie, Shijian Lu, Xiang Bai
ICDAR7
2017 Crowd Research: Open and Scalable University Laboratories
abstract
Research experiences today are limited to a privileged few at select universities. Providing open access to research experiences would enable global upward mobility and increased diversity in the scientific workforce. How can we coordinate a crowd of diverse volunteers on open-ended research? How could a PI have enough visibility into each person's contributions to recommend them for further study? We present Crowd Research, a crowdsourcing technique that coordinates open-ended research through an iterative cycle of open contribution, synchronous collaboration, and peer assessment. To aid upward mobility and recognize contributions in publications, we introduce a decentralized credit system: participants allocate credits to each other, which a graph centrality algorithm translates into a collectively-created author order. Over 1,500 people from 62 countries have participated, 74% from institutions with low access to research. Over two years and three projects, this crowd has produced articles at top-tier Computer Science venues, and participants have gone on to leading graduate programs.
Rajan Vaish, Snehalkumar (Neil) S. Gaikwad, Geza Kovacs, Andreas Veit, Ranjay Krishna, Imanol Arrieta Ibarra, Camelia Simoiu, Kimberly Wilber, Serge J. Belongie, Sharad Goel, James Davis 0001, Michael S. Bernstein
UIST9
2017 Collaborative Metric Learning
abstract
Metric learning algorithms produce distance metrics that capture the important relationships among data. In this work, we study the connection between metric learning and collaborative filtering. We propose Collaborative Metric Learning (CML) which learns a joint metric space to encode not only users' preferences but also the user-user and item-item similarity. The proposed algorithm outperforms state-of-the-art collaborative filtering algorithms on a wide range of recommendation tasks and uncovers the underlying spectrum of users' fine-grained preferences. CML also achieves significant speedup for Top-K recommendation tasks using off-the-shelf, approximate nearest-neighbor search, with negligible accuracy reduction.
Cheng-Kang Hsieh, Longqi Yang 0001, Yin Cui, Tsung-Yi Lin, Serge J. Belongie, Deborah Estrin
WWW5
2017 Yum-Me: A Personalized Nutrient-Based Meal Recommender System
abstract
Nutrient-based meal recommendations have the potential to help individuals prevent or manage conditions such as diabetes and obesity. However, learning people’s food preferences and making recommendations that simultaneously appeal to their palate and satisfy nutritional expectations are challenging. Existing approaches either only learn high-level preferences or require a prolonged learning period. We propose Yum-me , a personalized nutrient-based meal recommender system designed to meet individuals’ nutritional expectations, dietary restrictions, and fine-grained food preferences. Yum-me enables a simple and accurate food preference profiling procedure via a visual quiz-based user interface and projects the learned profile into the domain of nutritionally appropriate food options to find ones that will appeal to the user. We present the design and implementation of Yum-me and further describe and evaluate two innovative contributions. The first contriution is an open source state-of-the-art food image analysis model, named FoodDist . We demonstrate FoodDist’s superior performance through careful benchmarking and discuss its applicability across a wide array of dietary applications. The second contribution is a novel online learning framework that learns food preference from itemwise and pairwise image comparisons. We evaluate the framework in a field study of 227 anonymous users and demonstrate that it outperforms other baselines by a significant margin. We further conducted an end-to-end validation of the feasibility and effectiveness of Yum-me through a 60-person user study, in which Yum-me improves the recommendation acceptance rate by 42.63%.
Longqi Yang 0001, Cheng-Kang Hsieh, Hongjian Yang, John P. Pollak, Nicola Dell, Serge J. Belongie, Curtis L. Cole, Deborah Estrin
ACM Trans. Inf. Syst.6
2016 Learning to Detect and Match Keypoints with Deep Architectures
Hani Altwaijry, Andreas Veit, Serge J. Belongie
BMVC3
2016 Boosted Convolutional Neural Networks
Mohammad Moghimi, Serge J. Belongie, Mohammad J. Saberian, Jian Yang 0003, Nuno Vasconcelos, Li-Jia Li 0001
BMVC2
2016 Context Matters: Refining Object Detection in Video with Recurrent Neural Networks
Subarna Tripathi, Zachary C. Lipton, Serge J. Belongie, Truong Q. Nguyen
BMVC3
2016 Learning to Match Aerial Images with Deep Attentive Architectures
abstract
Image matching is a fundamental problem in Computer Vision. In the context of feature-based matching, SIFT and its variants have long excelled in a wide array of applications. However, for ultra-wide baselines, as in the case of aerial images captured under large camera rotations, the appearance variation goes beyond the reach of SIFT and RANSAC. In this paper we propose a data-driven, deep learning-based approach that sidesteps local correspondence by framing the problem as a classification task. Furthermore, we demonstrate that local correspondences can still be useful. To do so we incorporate an attention mechanism to produce a set of probable matches, which allows us to further increase performance. We train our models on a dataset of urban aerial imagery consisting of 'same' and 'different' pairs, collected for this purpose, and characterize the problem via a human study with annotations from Amazon Mechanical Turk. We demonstrate that our models outperform the state-of-the-art on ultra-wide baseline matching and approach human accuracy.
Hani Altwaijry, Eduard Trulls, James Hays, Pascal Fua, Serge J. Belongie
CVPR5
2016 Fine-Grained Categorization and Dataset Bootstrapping Using Deep Metric Learning with Humans in the Loop
abstract
Existing fine-grained visual categorization methods often suffer from three challenges: lack of training data, large number of fine-grained categories, and high intraclass vs. low inter-class variance. In this work we propose a generic iterative framework for fine-grained categorization and dataset bootstrapping that handles these three challenges. Using deep metric learning with humans in the loop, we learn a low dimensional feature embedding with anchor points on manifolds for each category. These anchor points capture intra-class variances and remain discriminative between classes. In each round, images with high confidence scores from our model are sent to humans for labeling. By comparing with exemplar images, labelers mark each candidate image as either a "true positive" or a "false positive." True positives are added into our current dataset and false positives are regarded as "hard negatives" for our metric learning model. Then the model is retrained with an expanded dataset and hard negatives for the next round. To demonstrate the effectiveness of the proposed framework, we bootstrap a fine-grained flower dataset with 620 categories from Instagram images. The proposed deep metric learning scheme is evaluated on both our dataset and the CUB-200-2001 Birds dataset. Experimental evaluations show significant performance gain using dataset bootstrapping and demonstrate state-of-the-art results achieved by the proposed deep metric learning methods.
Yin Cui, Feng Zhou 0002, Yuanqing Lin, Serge J. Belongie
CVPR4
2016 Residual Networks Behave Like Ensembles of Relatively Shallow Networks
abstract
In this work we propose a novel interpretation of residual networks showing that they can be seen as a collection of many paths of differing length. Moreover, residual networks seem to enable very deep networks by leveraging only the short paths during training. To support this observation, we rewrite residual networks as an explicit collection of paths. Unlike traditional models, paths through residual networks vary in length. Further, a lesion study reveals that these paths show ensemble-like behavior in the sense that they do not strongly depend on each other. Finally, and most surprising, most paths are shorter than one might expect, and only the short paths are needed during training, as longer paths do not contribute any gradient. For example, most of the gradient in a residual network with 110 layers comes from paths that are only 10-34 layers deep. Our results reveal one of the key characteristics that seem to enable the training of very deep networks: Residual networks avoid the vanishing gradient problem by introducing short paths which can carry gradient throughout the extent of very deep networks.
Andreas Veit, Kimberly Wilber, Serge J. Belongie
NIPS3
2016 Detecting temporally consistent objects in videos through object class label propagation
abstract
Object proposals for detecting moving or static video objects need to address issues such as speed, memory complexity and temporal consistency. We propose an efficient Video Object Proposal (VOP) generation method and show its efficacy in learning a better video object detector A deep-learning based video object detector learned using the proposed VOP achieves state-of-the-art detection performance on the Youtube-Objects dataset. We further propose a clustering of VOPs which can efficiently be used for detecting objects in video in a streaming fashion. As opposed to applying per-frame convolutional neural network (CNN) based object detection, our proposed method called Objects in Video Enabler thRough LAbel Propagation (OVERLAP) needs to classify only a small fraction of all candidate proposals in every video frame through streaming clustering of object proposals and class-label propagation. Source code for VOP clustering is available at https://github. com/subtri/streaming_VOP_clustering.
Subarna Tripathi, Serge J. Belongie, Youngbae Hwang, Truong Q. Nguyen
WACV2
2016 Can we still avoid automatic face detection?
abstract
After decades of study, automatic face detection and recognition systems are now accurate and widespread. Naturally, this means users who wish to avoid automatic recognition are becoming less able to do so. Where do we stand in this cat-and-mouse race? We currently live in a society where everyone carries a camera in their pocket. Many people willfully upload most or all of the pictures they take to social networks which invest heavily in automatic face recognition systems. In this setting, is it still possible for privacy-conscientious users to avoid automatic face detection and recognition? If so, how? Must evasion techniques be obvious to be effective, or are there still simple measures that users can use to protect themselves? In this work, we find ways to evade face detection on Facebook, a representative example of a popular social network that uses automatic face detection to enhance their service. We challenge widely-held beliefs about evading face detection: do our old techniques such as blurring the face region or wearing "privacy glasses" still work? We show that in general, state-of-the-art detectors can often find faces even if the subject wears occluding clothing or even if the uploader damages the photo to prevent faces from being detected.
Kimberly Wilber, Vitaly Shmatikov, Serge J. Belongie
WACV3
2016 Visipedia circa 2015
Serge J. Belongie, Pietro Perona
Pattern Recognit. Lett.1
2015 PlateClick: Bootstrapping Food Preferences Through an Adaptive Visual Interface
abstract
Food preference learning is an important component of wellness applications and restaurant recommender systems as it provides personalized information for effective food targeting and suggestions. However, existing systems require some form of food journaling to create a historical record of an individual's meal selections. In addition, current interfaces for food or restaurant preference elicitation rely extensively on text-based descriptions and rating methods, which can impose high cognitive load, thereby hampering wide adoption.
Longqi Yang 0001, Yin Cui, Fan Zhang 0022, John P. Pollak, Serge J. Belongie, Deborah Estrin
CIKM5
2015 Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection
abstract
We introduce tools and methodologies to collect high quality, large scale fine-grained computer vision datasets using citizen scientists - crowd annotators who are passionate and knowledgeable about specific domains such as birds or airplanes. We worked with citizen scientists and domain experts to collect NABirds, a new high quality dataset containing 48,562 images of North American birds with 555 categories, part annotations and bounding boxes. We find that citizen scientists are significantly more accurate than Mechanical Turkers at zero cost. We worked with bird experts to measure the quality of popular datasets like CUB-200-2011 and ImageNet and found class label error rates of at least 4%. Nevertheless, we found that learning algorithms are surprisingly robust to annotation errors and this level of training data corruption can lead to an acceptably small increase in test error if the training set has sufficient size. At the same time, we found that an expert-curated high quality test set like NABirds is necessary to accurately measure the performance of fine-grained computer vision systems. We used NABirds to train a publicly available bird recognition service deployed on the web site of the Cornell Lab of Ornithology.
Grant Van Horn, Steve Branson, Ryan Farrell, Scott Haber, Jessie Barry, Panagiotis G. Ipeirotis, Pietro Perona, Serge J. Belongie
CVPR8
2015 Learning deep representations for ground-to-aerial geolocalization
abstract
The recent availability of geo-tagged images and rich geospatial data has inspired a number of algorithms for image based geolocalization. Most approaches predict the location of a query image by matching to ground-level images with known locations (e.g., street-view data). However, most of the Earth does not have ground-level reference photos available. Fortunately, more complete coverage is provided by oblique aerial or “bird's eye” imagery. In this work, we localize a ground-level query image by matching it to a reference database of aerial imagery. We use publicly available data to build a dataset of 78K aligned crossview image pairs. The primary challenge for this task is that traditional computer vision approaches cannot handle the wide baseline and appearance variation of these cross-view pairs. We use our dataset to learn a feature representation in which matching views are near one another and mismatched views are far apart. Our proposed approach, Where-CNN, is inspired by deep learning success in face verification and achieves significant improvements over traditional hand-crafted features and existing deep features learned from other large-scale databases. We show the effectiveness of Where-CNN in finding matches between street view and aerial view imagery and demonstrate the ability of our learned features to generalize to novel locations.
Tsung-Yi Lin, Yin Cui, Serge J. Belongie, James Hays
CVPR3
2015 Tropel: Crowdsourcing Detectors with Minimal Training
abstract
This paper introduces the Tropel system which enables non-technical users to create arbitrary visual detectors without first annotating a training set. Our primary contribution is a crowd active learning pipeline that is seeded with only a single positive example and an unlabeled set of training images. We examine the crowd's ability to train visual detectors given severely limited training themselves. This paper presents a series of experiments that reveal the relationship between worker training, worker consensus and the average precision of detectors trained by crowd-in-the-loop active learning. In order to verify the efficacy of our system, we train detectors for bird species that work nearly as well as those trained on the exhaustively labeled CUB 200 dataset at significantly lower cost and with little effort from the end user. To further illustrate the usefulness of our pipeline, we demonstrate qualitative results on unlabeled datasets containing fashion images and street-level photographs of Paris.
Genevieve Patterson, Grant Van Horn, Serge J. Belongie, Pietro Perona, James Hays
HCOMP3
2015 Learning Visual Clothing Style with Heterogeneous Dyadic Co-Occurrences
abstract
With the rapid proliferation of smart mobile devices, users now take millions of photos every day. These include large numbers of clothing and accessory images. We would like to answer questions like 'What outfit goes well with this pair of shoes?' To answer these types of questions, one has to go beyond learning visual similarity and learn a visual notion of compatibility across categories. In this paper, we propose a novel learning framework to help answer these types of questions. The main idea of this framework is to learn a feature transformation from images of items into a latent space that expresses compatibility. For the feature transformation, we use a Siamese Convolutional Neural Network (CNN) architecture, where training examples are pairs of items that are either compatible or incompatible. We model compatibility based on co-occurrence in large-scale user behavior data, in particular co-purchase data from Amazon.com. To learn cross-category fit, we introduce a strategic method to sample training data, where pairs of items are heterogeneous dyads, i.e., the two elements of a pair belong to different high-level categories. While this approach is applicable to a wide variety of settings, we focus on the representative problem of learning compatible clothing style. Our results indicate that the proposed framework is capable of learning semantic information about visual style and is able to generate outfits of clothes, with items from different categories, that go well together.
Andreas Veit, Balazs Kovacs, Sean Bell, Julian J. McAuley, Kavita Bala, Serge J. Belongie
ICCV6
2015 Learning Concept Embeddings with Combined Human-Machine Expertise
abstract
This paper presents our work on "SNaCK," a low-dimensional concept embedding algorithm that combines human expertise with automatic machine similarity kernels. Both parts are complimentary: human insight can capture relationships that are not apparent from the object's visual similarity and the machine can help relieve the human from having to exhaustively specify many constraints. We show that our SNaCK embeddings are useful in several tasks: distinguishing prime and nonprime numbers on MNIST, discovering labeling mistakes in the Caltech UCSD Birds (CUB) dataset with the help of deep-learned features, creating training datasets for bird classifiers, capturing subjective human taste on a new dataset of 10,000 foods, and qualitatively exploring an unstructured set of pictographic characters. Comparisons with the state-of-the-art in these tasks show that SNaCK produces better concept embeddings that require less human supervision than the leading methods.
Kimberly Wilber, Iljung S. Kwak, David J. Kriegman, Serge J. Belongie
ICCV4
2015 Discriminative Regions: A Substrate for Analyzing Life-Logging Image Sequences
Mohammad Moghimi, Jacqueline Kerr, Eileen Johnson, Suneeta Godbole, Serge J. Belongie
MMM (2)5
2015 Learning Localized Perceptual Similarity Metrics for Interactive Categorization
abstract
Current similarity-based approaches to interactive fine grained categorization rely on learning metrics from holistic perceptual measurements of similarity between objects or images. However, making a single judgment of similarity at the object level can be a difficult or overwhelming task for the human user to perform. Secondly, a single general metric of similarity may not be able to adequately capture the minute differences that discriminate fine-grained categories. In this work, we propose a novel approach to interactive categorization that leverages multiple perceptual similarity metrics learned from localized and roughly aligned regions across images, reporting state-of-the-art results and outperforming methods that use a single nonlocalized similarity metric.
Catherine Wah, Subhransu Maji, Serge J. Belongie
WACV3
2014 Improved Bird Species Recognition Using Pose Normalized Deep Convolutional Nets
Steve Branson, Grant Van Horn, Pietro Perona, Serge J. Belongie
BMVC4
2014 Similarity Comparisons for Interactive Fine-Grained Categorization
abstract
Current human-in-the-loop fine-grained visual categorization systems depend on a predefined vocabulary of attributes and parts, usually determined by experts. In this work, we move away from that expert-driven and attribute-centric paradigm and present a novel interactive classification system that incorporates computer vision and perceptual similarity metrics in a unified framework. At test time, users are asked to judge relative similarity between a query image and various sets of images, these general queries do not require expert-defined terminology and are applicable to other domains and basic-level categories, enabling a flexible, efficient, and scalable system for fine-grained categorization with humans in the loop. Our system outperforms existing state-of-the-art systems for relevance feedback-based image retrieval as well as interactive classification, resulting in a reduction of up to 43% in the average number of questions needed to correctly classify an image.
Catherine Wah, Grant Van Horn, Steve Branson, Subhransu Maji, Pietro Perona, Serge J. Belongie
CVPR6
2014 Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, C. Lawrence Zitnick
ECCV (5)3
2014 Cost-Effective HITs for Relative Similarity Comparisons
abstract
Similarity comparisons of the form "Is object a more similar to b than to c?" form a useful foundation in several computer vision and machine learning applications. Unfortunately, an embedding of n points is only uniquely specified by n3 triplets, making collecting every triplet an expensive task. In noticing this difficulty, other researchers investigated more intelligent triplet sampling techniques, but they do not study their effectiveness or their potential drawbacks. Although it is important to reduce the number of collected triplets to generate a good embedding, it is also important to understand how best to display a triplet collection task to the user to better respect the worker's human constraints. In this work, we explore an alternative method for collecting triplets and analyze its financial cost, collection speed, and worker happiness as a function of the final embedding quality. We propose best practices for creating cost effective human intelligence tasks for collecting triplets. We show that rather than changing the sampling algorithm, simple changes to the crowdsourcing UI can drastically decrease the cost of collecting similarity comparisons. Finally, we provide a food similarity dataset as well as the labels collected from crowd workers.
Kimberly Wilber, Iljung S. Kwak, Serge J. Belongie
HCOMP3
2014 Analyzing sedentary behavior in life-logging images
abstract
We describe a study that aims to understand physical activity and sedentary behavior in free-living settings. We employed a wearable camera to record 3 to 5 days of imaging data with 40 participants, resulting in over 360,000 images. These images were then fully annotated by experienced staff with a rigorous coding protocol. We designed a deep learning based classifier in which we adapted a model that was originally trained for ImageNet [1]. We then added a spatio-temporal pyramid to our deep learning based classifier. Our results show our proposed method performs better than the state-of-the-art visual classification methods on our dataset. For most of the labels our system achieves more than 90% average accuracy across different individuals for frequent labels and more than 80% average accuracy for rare labels.
Mohammad Moghimi, Wanmin Wu, Jacqueline Chen, Suneeta Godbole, Simon J. Marshall, Jacqueline Kerr, Serge J. Belongie
ICIP7
2014 Adaptive ranking of facial attractiveness
abstract
As humans, we love to rank things. Top ten lists exist for everything from movie stars to scary animals. Ambiguities (i.e., ties) naturally occur in the process of ranking when people feel they cannot distinguish two items. Human reported rankings derived from star ratings abound on recommendation websites such as Yelp and Netflix. However, those websites differ in star precision which points to the need for ranking systems that adapt to an individual user's preference sensitivity. In this work we propose an adaptive system that allows for ties when collecting ranking data. Using this system, we propose a framework for obtaining computer-generated rankings. We test our system and a computer-generated ranking method on the problem of evaluating human attractiveness. Extensive experimental evaluations and analysis demonstrate the effectiveness and efficiency of our work.
Iljung S. Kwak, Serge J. Belongie, David J. Kriegman, Haizhou Ai
ICME3
2014 Recognizing locations with Google Glass: A case study
abstract
Wearable computers are rapidly gaining popularity as more people incorporate them into their everyday lives. The introduction of these devices allows for wider deployment of Computer Vision based applications. In this paper, we describe a system developed to deliver users of wearable computers a tour guide experience. In building our system, we compare and contrast different techniques towards achieving our goals. Those techniques include using various descriptor types, such as HOG, SIFT and SURF, under different encoding models, such as holistic approaches, Bag-of-Words, and Fisher Vectors. We evaluate those approaches using classification methods including Nearest Neighbor and Support Vector Machines. We also show how to incorporate information external to images, specifically GPS, to improve the user experience.
Hani Altwaijry, Mohammad Moghimi, Serge J. Belongie
WACV3
2014 Video text detection and recognition: Dataset and benchmark
abstract
This paper focuses on the problem of text detection and recognition in videos. Even though text detection and recognition in images has seen much progress in recent years, relatively little work has been done to extend these solutions to the video domain. In this work, we extend an existing end-to-end solution for text recognition in natural images to video. We explore a variety of methods for training local character models and explore methods to capitalize on the temporal redundancy of text in video. We present detection performance using the Video Analysis and Content Extraction (VACE) benchmarking framework on the ICDAR 2013 Robust Reading Challenge 3 video dataset and on a new video text dataset. We also propose a new performance metric based on precision-recall curves to measure the performance of text recognition in videos. Using this metric, we provide early video text recognition results on the above mentioned datasets.
Phuc Xuan Nguyen, Kai Wang 0059, Serge J. Belongie
WACV3
2014 Improving streaming video segmentation with early and mid-level visual processing
abstract
Despite recent advances in video segmentation, many opportunities remain to improve it using a variety of low and mid-level visual cues. We propose improvements to the leading streaming graph-based hierarchical video segmentation (streamGBH) method based on early and mid level visual processing. The extensive experimental analysis of our approach validates the improvement of hierarchical supervoxel representation by incorporating motion and color with effective filtering. We also pose and illuminate some open questions towards intermediate level video analysis as further extension to streamGBH. We exploit the supervoxels as an initialization towards estimation of dominant affine motion regions, followed by merging of such motion regions in order to hierarchically segment a video in a novel motion-segmentation framework which aims at subsequent applications such as foreground recognition.
Subarna Tripathi, Youngbae Hwang, Serge J. Belongie, Truong Q. Nguyen
WACV3
2014 The Ignorant Led by the Blind: A Hybrid Human-Machine Vision System for Fine-Grained Categorization
Steve Branson, Grant Van Horn, Catherine Wah, Pietro Perona, Serge J. Belongie
Int. J. Comput. Vis.5
2014 Editorial: Special Issue on Active and Interactive Methods in Computer Vision
Kristen Grauman, Serge J. Belongie
Int. J. Comput. Vis.2
2014 Fast Feature Pyramids for Object Detection
abstract
Multi-resolution image features may be approximated via extrapolation from nearby scales, rather than being computed explicitly. This fundamental insight allows us to design object detection algorithms that are as accurate, and considerably faster, than the state-of-the-art. The computational bottleneck of many modern detectors is the computation of features at every scale of a finely-sampled image pyramid. Our key insight is that one may compute finely sampled feature pyramids at a fraction of the cost, without sacrificing performance: for a broad family of features we find that features computed at octave-spaced scale intervals are sufficient to approximate features on a finely-sampled pyramid. Extrapolation is inexpensive as compared to direct feature computation. As a result, our approximation yields considerable speedups with negligible loss in detection accuracy. We modify three diverse visual recognition systems to use fast feature pyramids and show results on both pedestrian detection (measured on the Caltech, INRIA, TUD-Brussels and ETH data sets) and general object detection (measured on the PASCAL VOC). The approach is general and is widely applicable to vision algorithms requiring fine-grained multi-scale analysis. Our approximation is valid for images with broad spectra (most natural images) and fails for images with narrow band-pass spectra (e.g., periodic textures).
Piotr Dollár, Ron Appel, Serge J. Belongie, Pietro Perona
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Ultra-wide Baseline Aerial Imagery Matching in Urban Environments
Hani Altwaijry, Serge J. Belongie
BMVC2
2013 Match-time covariance for descriptors
Eric M. Christiansen, Vincent C. Rabaud, Andrew Ziegler, David J. Kriegman, Serge J. Belongie
BMVC5
2013 From Bikers to Surfers: Visual Recognition of Urban Tribes
abstract
Iljung S. Kwak1 [email protected] Ana C. Murillo2 [email protected] Peter N. Belhumeur3 [email protected] David Kriegman1 [email protected] Serge Belongie1 [email protected] 1 Dept. of Computer Science and Engineering University of California, San Diego, USA. 2 Dpt. Informatica e Ing. Sistemas Inst. Investigacion en Ingenieria de Aragon. University of Zaragoza, Spain. 3 Department of Computer Science Columbia University, USA.
Iljung S. Kwak, Ana Cristina Murillo, Peter N. Belhumeur, David J. Kriegman, Serge J. Belongie
BMVC5
2013 Efficient Large-Scale Structured Learning
abstract
We introduce an algorithm, SVM-IS, for structured SVM learning that is computationally scalable to very large datasets and complex structural representations. We show that structured learning is at least as fast-and often much faster-than methods based on binary classification for problems such as deformable part models, object detection, and multiclass classification, while achieving accuracies that are at least as good. Our method allows problem-specific structural knowledge to be exploited for faster optimization by integrating with a user-defined importance sampling function. We demonstrate fast train times on two challenging large scale datasets for two very different problems: Image Net for multiclass classification and CUB-200-2011 for deformable part model training. Our method is shown to be 10-50 times faster than SVMstructfor cost-sensitive multiclass classification while being about as fast as the fastest 1-vs-all methods for multiclass classification. For deformable part model training, it is shown to be 50-1000 times faster than methods based on SVMstruct, mining hard negatives, and Pegasos-style stochastic gradient descent. Source code of our method is publicly available.
Steve Branson, Oscar Beijbom, Serge J. Belongie
CVPR3
2013 Cross-View Image Geolocalization
abstract
The recent availability of large amounts of geotagged imagery has inspired a number of data driven solutions to the image geolocalization problem. Existing approaches predict the location of a query image by matching it to a database of georeferenced photographs. While there are many geotagged images available on photo sharing and street view sites, most are clustered around landmarks and urban areas. The vast majority of the Earth's land area has no ground level reference photos available, which limits the applicability of all existing image geolocalization methods. On the other hand, there is no shortage of visual and geographic data that densely covers the Earth - we examine overhead imagery and land cover survey data - but the relationship between this data and ground level query photographs is complex. In this paper, we introduce a cross-view feature translation approach to greatly extend the reach of image geolocalization methods. We can often localize a query even if it has no corresponding ground level images in the database. A key idea is to learn the relationship between ground level appearance and overhead appearance and land cover attributes from sparsely available geotagged ground-level images. We perform experiments over a 1600 km2 region containing a variety of scenes and land cover types. For each query, our algorithm produces a probability density over the region of interest.
Tsung-Yi Lin, Serge J. Belongie, James Hays
CVPR2
2013 Attribute-Based Detection of Unfamiliar Classes with Humans in the Loop
abstract
Recent work in computer vision has addressed zero-shot learning or unseen class detection, which involves categorizing objects without observing any training examples. However, these problems assume that attributes or defining characteristics of these unobserved classes are known, leveraging this information at test time to detect an unseen class. We address the more realistic problem of detecting categories that do not appear in the dataset in any form. We denote such a category as an unfamiliar class, it is neither observed at train time, nor do we possess any knowledge regarding its relationships to attributes. This problem is one that has received limited attention within the computer vision community. In this work, we propose a novel approach to the unfamiliar class detection task that builds on attribute-based classification methods, and we empirically demonstrate how classification accuracy is impacted by attribute noise and dataset "difficulty," as quantified by the separation of classes in the attribute space. We also present a method for incorporating human users to overcome deficiencies in attribute detection. We demonstrate results superior to existing methods on the challenging CUB-200-2011 dataset.
Catherine Wah, Serge J. Belongie
CVPR2
2013 Example based depth from fog
abstract
The presence of fog in an image reduces contrast which can be considered a nuisance in imaging applications, however, we consider this useful information for image enhancement and scene understanding. In this paper, we present a new method for estimating depth from fog in a single image and single image fog removal. We use an example based approach that is trained from data with known fog and depth. A data driven method and physics based model are used to develop the example based learning framework for single image fog removal. In addition, we account for various colors of fog by using a linear transformation of the RGB colorspace. This approach has the flexibility to learn from various scenes and relaxes the common constraint of fixed camera position. We present depth estimations and fog removal from a single image with good results.
Kristofor B. Gibson, Serge J. Belongie, Truong Q. Nguyen
ICIP2
2013 Relative ranking of facial attractiveness
abstract
Automatic evaluation of human facial attractiveness is a challenging problem that has received relatively little attention from the computer vision community. Previous work in this area have posed attractiveness as a classification problem. However, for applications that require fine-grained relationships between objects, learning to rank has been shown to be superior over the direct interpretation of classifier scores as ranks [27]. In this paper, we propose and implement a personalized relative beauty ranking system. Given training data of faces sorted based on a subject's personal taste, we learn how to rank novel faces according to that person's taste. Using a blend of Facial Geometric Relations, HOG, GIST, L*a*b* Color Histograms, and Dense-SIFT + PCA feature types, our system achieves an average accuracy of 63% on pairwise comparisons of novel test faces. We examine the effectiveness of our method through lesion testing and find that the most effective feature types for predicting beauty preferences are HOG, GIST, and Dense-SIFT + PCA features.
Hani Altwaijry, Serge J. Belongie
WACV2
2012 Locally Uniform Comparison Image Descriptor
abstract
Keypoint matching between pairs of images using popular descriptors like SIFT or a faster variant called SURF is at the heart of many computer vision algorithms including recognition, mosaicing, and structure from motion. For real-time mobile applications, very fast but less accurate descriptors like BRIEF and related methods use a random sampling of pairwise comparisons of pixel intensities in an image patch. Here, we introduce Locally Uniform Comparison Image Descriptor (LUCID), a simple description method based on permutation distances between the ordering of intensities of RGB values between two patches. LUCID is computable in linear time with respect to patch size and does not require floating point computation. An analysis reveals an underlying issue that limits the potential of BRIEF and related approaches compared to LUCID. Experiments demonstrate that LUCID is faster than BRIEF, and its accuracy is directly comparable to SURF while being more than an order of magnitude faster.
Andrew Ziegler, Eric M. Christiansen, David J. Kriegman, Serge J. Belongie
NIPS4
2012 Non-rigid surface detection for gestural interaction with applicable surfaces
abstract
In this work we present a novel application of non-rigid surface detection to enable gestural interaction with applicable surfaces. This method can add interactivity to traditionally passive media such as books, newspapers, restaurant menus, or anything else printed on paper. We allow a user to interact with these surfaces in a natural manner and present basic gestures based on pointing and touching. This technique was developed as part of an ongoing effort to create an assisted reading device for the visually impaired. However, it is suited to general applications and can be used as a practical mechanism for interaction with screen-less wearable devices. Our key contributions are a unique application of non-rigid surface detection, a basic gesturing paradigm, and a proof of concept system.
Andrew Ziegler, Serge J. Belongie
WACV2
2011 JBoost Optimization of Color Detectors for Autonomous Underwater Vehicle Navigation
Christopher Barngrover, Serge J. Belongie, Ryan Kastner
CAIP (2)2
2011 From region similarity to category discovery
abstract
The goal of object category discovery is to automatically identify groups of image regions which belong to some new, previously unseen category. This task is typically performed in a purely unsupervised setting, and as a result, performance depends critically upon accurate assessments of similarity between unlabeled image regions. To improve the accuracy of category discovery, we develop a novel multiple kernel learning algorithm based on structural SVM, which optimizes a similarity space for nearest-neighbor prediction. The optimized space is then used to cluster unlabeled data and identify new categories. Experimental results on the MSRC and PASCAL VOC2007 data sets indicate that using an optimized similarity metric can improve clustering for category discovery. Furthermore, we demonstrate that including both labeled and unlabeled training data when optimizing the similarity metric can improve the overall quality of the system.
Carolina Galleguillos, Brian McFee, Serge J. Belongie, Gert R. G. Lanckriet
CVPR3
2011 Wet fingerprint recognition: Challenges and opportunities
abstract
Many fingers wrinkle or shrivel when immersed in water. When used for biometric identification, the recognition rate for wrinkled fingers degrades. The impact of wrinkling has so far not been well-understood. In this study, we present an investigation of how the finger-skin expansion due to wrinkling impacts the quality of scanned finger prints and characterize the qualitative changes that affect recognition. We also introduce the Wet and Wrinkled Finger (WWF) database that we will make available to other researchers. In this database of 300 fingers, 185 are visibly wrinkled after immersion; multiple images of dry and immersed fingerprints were acquired. In this paper, we present baseline recognition rates on WWF using two algorithms a commercial fingerprint recognition algorithm and the publicly available Bozorth3 matcher. Specifically, we show a degradation in accuracy with both algorithms when comparing Dry-finger to Dry finger verification with Dry-finger to Wet-finger verification. We analyze performance on a per-finger basis and note a difference in accuracy amongst fingers, and as consequence make recommendations about which fingers to use in environments where fingers are apt to be wet. Additionally, we propose an implementation of a classifier that can decide if the incoming query is wrinkled.
Prasanna Krishnasamy, Serge J. Belongie, David J. Kriegman
IJCB2
2011 Strong supervision from weak annotation: Interactive training of deformable part models
abstract
We propose a framework for large scale learning and annotation of structured models. The system interleaves interactive labeling (where the current model is used to semi-automate the labeling of a new example) and online learning (where a newly labeled example is used to update the current model parameters). This framework is scalable to large datasets and complex image models and is shown to have excellent theoretical and practical properties in terms of train time, optimality guarantees, and bounds on the amount of annotation effort per image. We apply this framework to part-based detection, and introduce a novel algorithm for interactive labeling of deformable part models. The labeling tool updates and displays in real-time the maximum likelihood location of all parts as the user clicks and drags the location of one or more parts. We demonstrate that the system can be used to efficiently and robustly train part and pose detectors on the CUB Birds-200-a challenging dataset of birds in unconstrained pose and environment.
Steve Branson, Pietro Perona, Serge J. Belongie
ICCV3
2011 Pose, illumination and expression invariant pairwise face-similarity measure via Doppelgänger list comparison
abstract
Face recognition approaches have traditionally focused on direct comparisons between aligned images, e.g. using pixel values or local image features. Such comparisons become prohibitively difficult when comparing faces across extreme differences in pose, illumination and expression. The goal of this work is to develop a face-similarity measure that is largely invariant to these differences. We propose a novel data driven method based on the insight that comparing images of faces is most meaningful when they are in comparable imaging conditions. To this end we describe an image of a face by an ordered list of identities from a Library. The order of the list is determined by the similarity of the Library images to the probe image. The lists act as a signature for each face image: similarity between face images is determined via the similarity of the signatures. Here the CMU Multi-PIE database, which includes images of 337 individuals in more than 2000 pose, lighting and illumination combinations, serves as the Library. We show improved performance over state of the art face-similarity measures based on local features, such as FPLBP, especially across large pose variations on FacePix and multi-PIE. On LFW we show improved performance in comparison with measures like SIFT (on fiducials), LBP, FPLBP and Gabor (C1).
Florian Schroff, Tali Treibitz, David J. Kriegman, Serge J. Belongie
ICCV4
2011 Multiclass recognition and part localization with humans in the loop
abstract
We propose a visual recognition system that is designed for fine-grained visual categorization. The system is composed of a machine and a human user. The user, who is unable to carry out the recognition task by himself, is interactively asked to provide two heterogeneous forms of information: clicking on object parts and answering binary questions. The machine intelligently selects the most informative question to pose to the user in order to identify the object's class as quickly as possible. By leveraging computer vision and analyzing the user responses, the overall amount of human effort required, measured in seconds, is minimized. We demonstrate promising results on a challenging dataset of uncropped images, achieving a significant average reduction in human effort over previous methods.
Catherine Wah, Steve Branson, Pietro Perona, Serge J. Belongie
ICCV4
2011 End-to-end scene text recognition
abstract
This paper focuses on the problem of word detection and recognition in natural images. The problem is significantly more challenging than reading text in scanned documents, and has only recently gained attention from the computer vision community. Sub-components of the problem, such as text detection and cropped image word recognition, have been studied in isolation [7, 4, 20]. However, what is unclear is how these recent approaches contribute to solving the end-to-end problem of word recognition. We fill this gap by constructing and evaluating two systems. The first, representing the de facto state-of-the-art, is a two stage pipeline consisting of text detection followed by a leading OCR engine. The second is a system rooted in generic object recognition, an extension of our previous work in [20]. We show that the latter approach achieves superior performance. While scene text recognition has generally been treated with highly domain-specific methods, our results demonstrate the suitability of applying generic computer vision methods. Adopting this approach opens the door for real world scene text recognition to benefit from the rapid advances that have been taking place in object recognition.
Kai Wang 0059, Boris Babenko, Serge J. Belongie
ICCV3
2011 Multiple Instance Learning with Manifold Bags
Boris Babenko, Nakul Verma, Piotr Dollár, Serge J. Belongie
ICML4
2011 Adaptively Learning the Crowd Kernel
Omer Tamuz, Ce Liu 0001, Serge J. Belongie, Ohad Shamir, Adam Tauman Kalai
ICML3
2011 Robust Object Tracking with Online Multiple Instance Learning
abstract
In this paper, we address the problem of tracking an object in a video given its location in the first frame and no other information. Recently, a class of tracking techniques called "tracking by detection" has been shown to give promising results at real-time speeds. These methods train a discriminative classifier in an online manner to separate the object from the background. This classifier bootstraps itself by using the current tracker state to extract positive and negative examples from the current frame. Slight inaccuracies in the tracker can therefore lead to incorrectly labeled training examples, which degrade the classifier and can cause drift. In this paper, we show that using Multiple Instance Learning (MIL) instead of traditional supervised learning avoids these problems and can therefore lead to a more robust tracker with fewer parameter tweaks. We propose a novel online MIL algorithm for object tracking that achieves superior results with real-time performance. We present thorough experimental results (both qualitative and quantitative) on a number of challenging video clips.
Boris Babenko, Ming-Hsuan Yang 0001, Serge J. Belongie
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 The Fastest Pedestrian Detector in the West
abstract
We demonstrate a multiscale pedestrian detector operating in near real time (~6 fps on 640x480 images) with state-of-the-art detection performance. The computational bottleneck of many modern detectors is the construction of an image pyramid, typically sampled at 8-16 scales per octave, and associated feature computations at each scale. We propose a technique to avoid constructing such a finely sampled image pyramid without sacrificing performance: our key insight is that for a broad family of features, including gradient histograms, the feature responses computed at a single scale can be used to approximate feature responses at nearby scales. The approximation is accurate within an entire scale octave. This allows us to decouple the sampling of the image pyramid from the sampling of detection scales. Overall, our approximation yields a speedup of 10-100 times over competing methods with only a minor loss in detection accuracy of about 1-2% on the Caltech Pedestrian dataset across a wide range of evaluation settings. The results are confirmed on three additional datasets (INRIA, ETH, and TUD-Brussels) where our method always scores within a few percent of the state-of-the-art while being 1-2 orders of magnitude faster. The approach is general and should be widely applicable.
Piotr Dollár, Serge J. Belongie, Pietro Perona
BMVC2
2010 Multi-class object localization by combining local contextual interactions
abstract
Recent work in object localization has shown that the use of contextual cues can greatly improve accuracy over models that use appearance features alone. Although many of these models have successfully explored different types of contextual sources, they only consider one type of contextual interaction (e.g., pixel, region or object level interactions), leaving open questions about the true potential contribution of context. Furthermore, contributions across object classes and over appearance features still remain unknown. In this work, we introduce a novel model for multi-class object localization that incorporates different levels of contextual interactions. We study contextual interactions at pixel, region and object level by using three different sources of context: semantic, boundary support and contextual neighborhoods. Our framework learns a single similarity metric from multiple kernels, combining pixel and region interactions with appearance features, and then uses a conditional random field to incorporate object level interactions. We perform experiments on two challenging image databases: MSRC and PASCAL VOC 2007. Experimental results show that our model outperforms current state-of-the-art contextual frameworks and reveals individual contributions for each contextual interaction level, as well as the importance of each type of feature in object localization.
Carolina Galleguillos, Brian McFee, Serge J. Belongie, Gert R. G. Lanckriet
CVPR3
2010 Visual Recognition with Humans in the Loop
Steve Branson, Catherine Wah, Florian Schroff, Boris Babenko, Peter Welinder, Pietro Perona, Serge J. Belongie
ECCV (4)7
2010 Word Spotting in the Wild
Kai Wang 0059, Serge J. Belongie
ECCV (1)2
2010 The Multidimensional Wisdom of Crowds
abstract
Distributing labeling tasks among hundreds or thousands of annotators is an increasingly important method for annotating large datasets. We present a method for estimating the underlying value (e.g. the class) of each image from (noisy) annotations provided by multiple annotators. Our method is based on a model of the image formation and annotation process. Each image has different characteristics that are represented in an abstract Euclidean space. Each annotator is modeled as a multidimensional entity with variables representing competence, expertise and bias. This allows the model to discover and represent groups of annotators that have different sets of skills and knowledge, as well as groups of images that differ qualitatively. We find that our model predicts ground truth labels on both synthetic and real data more accurately than state of the art methods. Experiments also show that our model, starting from a set of binary labels, may discover rich information, such as different "schools of thought" amongst the annotators, and can group together images belonging to separate categories.
Peter Welinder, Steve Branson, Serge J. Belongie, Pietro Perona
NIPS3
2010 Context based object categorization: A critical survey
Carolina Galleguillos, Serge J. Belongie
Comput. Vis. Image Underst.2
2010 Globally Optimal Algorithms for Stratified Autocalibration
abstract
We present practical algorithms for stratified autocalibration with theoretical guarantees of global optimality. Given a projective reconstruction, we first upgrade it to affine by estimating the position of the plane at infinity. The plane at infinity is computed by globally minimizing a least squares formulation of the modulus constraints. In the second stage, this affine reconstruction is upgraded to a metric one by globally minimizing the infinite homography relation to compute the dual image of the absolute conic (DIAC). The positive semidefiniteness of the DIAC is explicitly enforced as part of the optimization process, rather than as a post-processing step. For each stage, we construct and minimize tight convex relaxations of the highly non-convex objective functions in a branch and bound optimization framework. We exploit the inherent problem structure to restrict the search space for the DIAC and the plane at infinity to a small, fixed number of branching dimensions, independent of the number of views. Chirality constraints are incorporated into our convex relaxations to automatically select an initial region which is guaranteed to contain the global minimum. Experimental evidence of the accuracy, speed and scalability of our algorithm is presented on synthetic and real data.
Manmohan Krishna Chandraker, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
Int. J. Comput. Vis.4
2009 Integral Channel Features
abstract
We study the performance of ‘integral channel features’ for image classification tasks, \nfocusing in particular on pedestrian detection. The general idea behind integral channel features is that multiple registered image channels are computed using linear and \nnon-linear transformations of the input image, and then features such as local sums, histograms, and Haar features and their various generalizations are efficiently computed \nusing integral images. Such features have been used in recent literature for a variety of \ntasks – indeed, variations appear to have been invented independently multiple times. \nAlthough integral channel features have proven effective, little effort has been devoted to \nanalyzing or optimizing the features themselves. In this work we present a unified view \nof the relevant work in this area and perform a detailed experimental evaluation. We \ndemonstrate that when designed properly, integral channel features not only outperform \nother features including histogram of oriented gradient (HOG), they also (1) naturally \nintegrate heterogeneous sources of information, (2) have few parameters and are insensitive to exact parameter settings, (3) allow for more accurate spatial localization during \ndetection, and (4) result in fast detectors when coupled with cascade classifiers.
Piotr Dollár, Zhuowen Tu, Pietro Perona, Serge J. Belongie
BMVC4
2009 Visual tracking with online Multiple Instance Learning
abstract
In this paper, we address the problem of learning an adaptive appearance model for object tracking. In particular, a class of tracking techniques called “tracking by detection” have been shown to give promising results at real-time speeds. These methods train a discriminative classifier in an online manner to separate the object from the background. This classifier bootstraps itself by using the current tracker state to extract positive and negative examples from the current frame. Slight inaccuracies in the tracker can therefore lead to incorrectly labeled training examples, which degrades the classifier and can cause further drift. In this paper we show that using Multiple Instance Learning (MIL) instead of traditional supervised learning avoids these problems, and can therefore lead to a more robust tracker with fewer parameter tweaks. We present a novel online MIL algorithm for object tracking that achieves superior results with real-time performance.
Boris Babenko, Ming-Hsuan Yang 0001, Serge J. Belongie
CVPR3
2009 Linear embeddings in non-rigid structure from motion
abstract
This paper proposes a method to recover the embedding of the possible shapes assumed by a deforming nonrigid object by comparing triplets of frames from an orthographic video sequence. We assume that we are given features tracked with no occlusions and no outliers but possible noise, an orthographic camera and that any 3D shape of a deforming object is a linear combination of several canonical shapes. By exploiting any repetition in the object motion and defining an ordering between triplets of frames in a generalized non-metric multi-dimensional scaling framework, our approach recovers the shape coefficients of the linear combination, independently from other structure and motion parameters. From this point, a good estimate of the remaining unknowns is obtained for a final optimization to perform full non-rigid structure from motion. Results are presented on synthetic and real image sequences and our method is found to perform better than current state of the art.
Vincent C. Rabaud, Serge J. Belongie
CVPR2
2009 Similarity metrics for categorization: From monolithic to category specific
abstract
Similarity metrics that are learned from labeled training data can be advantageous in terms of performance and/or efficiency. These learned metrics can then be used in conjunction with a nearest neighbor classifier, or can be plugged in as kernels to an SVM. For the task of categorization two scenarios have thus far been explored. The first is to train a single “monolithic” similarity metric that is then used for all examples. The other is to train a metric for each category in a 1-vs-all manner. While the former approach seems to be at a disadvantage in terms of performance, the latter is not practical for large numbers of categories. In this paper we explore the space in between these two extremes. We present an algorithm that learns a few similarity metrics, while simultaneously grouping categories together and assigning one of these metrics to each group. We present promising results and show how the learned metrics generalize to novel categories.
Boris Babenko, Steve Branson, Serge J. Belongie
ICCV3
2009 Fingerprint recognition system performance in the maritime environment
abstract
This study assesses the effects of prolonged exposure of fingers to water on the performance of existing fingerprint recognition systems. The dataset used in this research is collected using a high-end, multispectral fingerprint scanner. To perform a data acquisition, we recruited volunteers to contribute their fingerprint samples to the dataset in multiple sessions. Once the dataset is filled with both fingerprints under normal and wrinkled conditions, we use a minutiae-based fingerprint verification system to retrieve the match scores between all combinations of prints. Finally, we use receiver operating characteristic (ROC) curve to measure the behavior of such systems under maritime environment. Using the equal error rate (EER), we successfully quantify the degradation in performance due to water-induced skin pruning, which is approximately 1% reduction in EER.
Hourieh Fakourfar, Serge J. Belongie
WACV2
2009 Toward a perceptual space for gloss
abstract
We design and implement a comprehensive study of the perception of gloss. This is the largest study of its kind to date, and the first to use real material measurements. In addition, we develop a novel multi-dimensional scaling (MDS) algorithm for analyzing pairwise comparisons. The data from the psychophysics study and the MDS algorithm is used to construct a low dimensional perceptual embedding of these bidirectional reflectance distribution functions (BRDFs). The embedding is validated by correlating it with nine gloss dimensions, fitted parameters of seven analytical BRDF models, and a perceptual parameterization of Ward's model. We also introduce a novel perceptual interpolation scheme that uses the embedding to provide the user with an intuitive interface for navigating the space of gloss and constructing new materials.
Josh Wills, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
ACM Trans. Graph.4
2008 Object categorization using co-occurrence, location and appearance
abstract
In this work we introduce a novel approach to object categorization that incorporates two types of context-co-occurrence and relative location - with local appearance-based features. Our approach, named CoLA (for co-occurrence, location and appearance), uses a conditional random field (CRF) to maximize object label agreement according to both semantic and spatial relevance. We model relative location between objects using simple pairwise features. By vector quantizing this feature space, we learn a small set of prototypical spatial relationships directly from the data. We evaluate our results on two challenging datasets: PASCAL 2007 and MSRC. The results show that combining co-occurrence and spatial context improves accuracy in as many as half of the categories compared to using co-occurrence alone.
Carolina Galleguillos, Andrew Rabinovich, Serge J. Belongie
CVPR3
2008 Re-thinking non-rigid structure from motion
abstract
We present a novel approach to non-rigid structure from motion (NRSFM) from an orthographic video sequence, based on a new interpretation of the problem. Existing approaches assume the object shape space is well-modeled by a linear subspace. Our approach only assumes that small neighborhoods of shapes are well-modeled with a linear subspace. This constrains the shapes to belong to a manifold of dimensionality equal to the number of degrees of freedom of the object. After showing that the problem is still overconstrained, we present a solution composed of a novel initialization algorithm, followed by a robust extension of the Locally Smooth Manifold Learning algorithm tailored to the NRSFM problem. We finally present some test cases where the linear basis method fails (and is actually not meant to work) while the proposed approach is successful.
Vincent C. Rabaud, Serge J. Belongie
CVPR2
2008 Multiple Component Learning for Object Detection
Piotr Dollár, Boris Babenko, Serge J. Belongie, Pietro Perona, Zhuowen Tu
ECCV (2)3
2008 Weakly Supervised Object Localization with Stable Segmentations
Carolina Galleguillos, Boris Babenko, Andrew Rabinovich, Serge J. Belongie
ECCV (1)4
2008 Practical Global Optimization for Multiview Geometry
Fredrik Kahl, Sameer Agarwal 0001, Manmohan Krishna Chandraker, David J. Kriegman, Serge J. Belongie
Int. J. Comput. Vis.5
2007 Feature Mining for Image Classification
abstract
The efficiency and robustness of a vision system is often largely determined by the quality of the image features available to it. In data mining, one typically works with immense volumes of raw data, which demands effective algorithms to explore the data space. In analogy to data mining, the space of meaningful features for image analysis is also quite vast. Recently, the challenges associated with these problem areas have become more tractable through progress made in machine learning and concerted research effort in manual feature design by domain experts. In this paper, we propose a feature mining paradigm for image classification and examine several feature mining strategies. We also derive a principled approach for dealing with features with varying computational demands. Our goal is to alleviate the burden of manual feature design, which is a key problem in computer vision and machine learning. We include an in-depth empirical study on three typical data sets and offer theoretical explanations for the performance of various feature mining strategies. As a final confirmation of our ideas, we show results of a system, that utilizing feature mining strategies matches or outperforms the best reported results on pedestrian classification (where considerable effort has been devoted to expert feature design).
Piotr Dollár, Zhuowen Tu, Serge J. Belongie
CVPR4
2007 Recognizing Groceries in situ Using in vitro Training Data
abstract
The problem of using pictures of objects captured under ideal imaging conditions (here referred to as in vitro) to recognize objects in natural environments (in situ) is an emerging area of interest in computer vision and pattern recognition. Examples of tasks in this vein include assistive vision systems for the blind and object recognition for mobile robots; the proliferation of image databases on the web is bound to lead to more examples in the near future. Despite its importance, there is still a need for a freely available database to facilitate study of this kind of training/testing dichotomy. In this work one of our contributions is a new multimedia database of 120 grocery products, GroZi-120. For every product, two different recordings are available: in vitro images extracted from the web, and in situ images extracted from camcorder video collected inside a grocery store. As an additional contribution, we present the results of applying three commonly used object recognition/detection algorithms (color histogram matching, SIFT matching, and boosted Haar-like features) to the dataset. Finally, we analyze the successes and failures of these algorithms against product type and imaging conditions, both in terms of recognition rate and localization accuracy, in order to suggest ways forward for further research in this domain.
Michele Merler, Carolina Galleguillos, Serge J. Belongie
CVPR3
2007 Task Specific Local Region Matching
abstract
Many problems in computer vision require the knowledge of potential point correspondences between two images. The usual approach for automatically determining correspondences begins by comparing small neighborhoods of high saliency in both images. Since speed is of the essence, most current approaches for local region matching involve the computation of a feature vector that is invariant to various geometric and photometric transformations, followed by fast distance computations using standard vector norms. These algorithms include many parameters, and choosing an algorithm and setting its parameters for a given problem is more an art than a science. Furthermore, although invariance of the resulting feature space is in general desirable, there is necessarily a tradeoff between invariance and descriptiveness for any given task. In this paper we pose local region matching as a classification problem, and use powerful machine learning techniques to train a classifier that selects features from a much larger pool. Our algorithm can be trained on specific domains or tasks, and performs better than the state of the art in such cases. Since our method is an application of boosting, we refer to it as boosted region matching (BOOM).
Boris Babenko, Piotr Dollár, Serge J. Belongie
ICCV3
2007 Globally Optimal Affine and Metric Upgrades in Stratified Autocalibration
abstract
We present a practical, stratified autocalibration algorithm with theoretical guarantees of global optimality. Given a projective reconstruction, the first stage of the algorithm upgrades it to affine by estimating the position of the plane at infinity. The plane at infinity is computed by globally minimizing a least squares formulation of the modulus constraints. In the second stage, the algorithm upgrades this affine reconstruction to a metric one by globally minimizing the infinite homography relation to compute the dual image of the absolute conic (DIAC). The positive semidefiniteness of the DIAC is explicitly enforced as part of the optimization process, rather than as a post-processing step. For each stage, we construct and minimize tight convex relaxations of the highly non-convex objective functions in a branch and bound optimization framework. We exploit the problem structure to restrict the search space for the DIAC and the plane at infinity to a small, fixed number of branching dimensions, independent of the number of views. Experimental evidence of the accuracy, speed and scalability of our algorithm is presented on synthetic and real data. MATLAB code for the implementation is made available to the community.
Manmohan Krishna Chandraker, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
ICCV4
2007 Objects in Context
abstract
In the task of visual object categorization, semantic context can play the very important role of reducing ambiguity in objects' visual appearance. In this work we propose to incorporate semantic object context as a post-processing step into any off-the-shelf object categorization model. Using a conditional random field (CRF) framework, our approach maximizes object label agreement according to contextual relevance. We compare two sources of context: one learned from training data and another queried from Google Sets. The overall performance of the proposed framework is evaluated on the PASCAL and MSRC datasets. Our findings conclude that incorporating context into object categorization greatly improves categorization accuracy.
Andrew Rabinovich, Andrea Vedaldi, Carolina Galleguillos, Eric Wiewiora, Serge J. Belongie
ICCV5
2007 Soylent Grid: it's Made of People
abstract
The ground truth labeling of an image dataset is a task that often requires a large amount of human time and labor. We present an infrastructure for distributed human labeling that can exploit the modularity of common vision problems involving segmentation and recognition. We present the different elements of this infrastructure in detail, in particular the different vision human computational tasks (HCTs) and machine computable tasks (MCTs). We also discuss the impact of such a system on Internet security vs. the current state of the art. Finally, we present our prototype implementation of such a system, named soylent grid, on typical problems.
Stephan Steinbach, Vincent C. Rabaud, Serge J. Belongie
ICCV3
2007 Non-isometric manifold learning: analysis and an algorithm
abstract
In this work we take a novel view of nonlinear manifold learning. Usually, manifold learning is formulated in terms of finding an embedding or 'unrolling' of a manifold into a lower dimensional space. Instead, we treat it as the problem of learning a representation of a nonlinear, possibly non-isometric manifold that allows for the manipulation of novel points. Central to this view of manifold learning is the concept of generalization beyond the training data. Drawing on concepts from supervised learning, we establish a framework for studying the problems of model assessment, model complexity, and model selection for manifold learning. We present an extension of a recent algorithm, Locally Smooth Manifold Learning (LSML), and show it has good generalization properties. LSML learns a representation of a manifold or family of related manifolds and can be used for computing geodesic distances, finding the projection of a point onto a manifold, recovering a manifold from points corrupted by noise, generating novel points on a manifold, and more.
Piotr Dollár, Vincent C. Rabaud, Serge J. Belongie
ICML3
2006 Vision in the Small: Reconstructing the Structure of Protein Macromolecules from Cryo-Electron Micrographs
abstract
Single particle reconstruction using Cryo-Electron Microscopy (cryo-EM) is an emerging technique in structural biology for estimating the 3-D structure (density) of protein macromolecules. Unlike tomography where a large number of images of a specimen can be acquired, the number of images of an individual particle is limited because of radiation damage. Instead, the specimen consists of identical copies of the same protein macro-molecule embedded in vitreous ice at random and unknown 3-D orientations. Because the images are extremely noisy, thousands to hundreds-of-thousands of projections are needed to achieve the desired resolution of 5A. Along with differences of the imaging modality compared to photographs, single particle reconstruction provides a unique set of challenges to existing computer vision algorithms. Here, we introduce the challenge and opportunity of reconstruction from transmission electron micrographs, and briefly describe our contributions in areas of particle detection, contrast transfer function (CTF) estimation, and initial 3-D model construction. Reconstructing the Structure of Protein Macromolecules One of the most exciting challenges for biology today is understanding the molecular machinery of the cell as a working, dynamic system. Critical to this understanding is determining the 3-D structure of protein macromolecules, a task that is often accomplished using x-ray crystallography. The technique of cryo electron microscopy (cryo-EM) has a unique role to play in addressing this challenge as it can provide structural information of large macromolecular complexes in a variety of conformational and compositional states while preserved under close to physiological conditions. Traditionally the methods for cryo-EM have been time consuming and labor intensive, involving data acquisition, analysis and averaging of thousands to hundreds of thousands of images (views) of the individual macro-molecular complexes. Thus, over the last few years there has been considerable interest and substantial effort devoted to developing automated methods to improve the accuracy, robustness, ease of use, and throughput of cryo-EM [1, 2, 3, 12, 15, 17], and this abstract considers three aspects originally presented in [9, 10, 11]. 1
Satya P. Mallick, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie
BMVC4
2006 Supervised Learning of Edges and Object Boundaries
abstract
Edge detection is one of the most studied problems in computer vision, yet it remains a very challenging task. It is difficult since often the decision for an edge cannot be made purely based on low level cues such as gradient, instead we need to engage all levels of information, low, middle, and high, in order to decide where to put edges. In this paper we propose a novel supervised learning algorithm for edge and object boundary detection which we refer to as Boosted Edge Learning or BEL for short. A decision of an edge point is made independently at each location in the image; a very large aperture is used providing significant context for each decision. In the learning stage, the algorithm selects and combines a large number of features across different scales in order to learn a discriminative model using an extended version of the Probabilistic Boosting Tree classification algorithm. The learning based framework is highly adaptive and there are no parameters to tune. We show applications for edge detection in a number of specific image domains as well as on natural images. We test on various datasets including the Berkeley dataset and the results obtained are very good.
Piotr Dollár, Zhuowen Tu, Serge J. Belongie
CVPR (2)3
2006 Structure and View Estimation for Tomographic Reconstruction: A Bayesian Approach
abstract
This paper addresses the problem of reconstructing the density of a scene from multiple projection images produced by modalities such as x-ray, electron microscopy, etc. where an image value is related to the integral of the scene density along a 3D line segment between a radiation source and a point on the image plane. While computed tomography (CT) addresses this problem when the absolute orientation of the image plane and radiation source directions are known, this paper addresses the problem when the orientations are unknown - it is akin to the structure-from-motion (SFM) problem when the extrinsic camera parameters are unknown. We study the problem within the context of reconstructing the density of protein macro-molecules in Cryogenic Electron Microscopy (cryo-EM), where images are very noisy and existing techniques use several thousands of images. In a non-degenerate configuration, the viewing planes corresponding to two projections, intersect in a line in 3D. Using the geometry of the imaging setup, it is possible to determine the projections of this 3D line on the two image planes. In turn, the problem can be formulated as a type of orthographic structure from motion from line correspondences where the line correspondences between two views are unreliable due to image noise. We formulate the task as the problem of denoising a correspondence matrix and present a Bayesian solution to it. Subsequently, the absolute orientation of each projection is determined followed by density reconstruction. We show results on cryo-EM images of proteins and compare our results to that of Electron Micrograph Analysis (EMAN) - a widely used reconstruction tool in cryo-EM.
Satya P. Mallick, Sameer Agarwal 0001, David J. Kriegman, Serge J. Belongie, Bridget Carragher, Clinton S. Potter
CVPR (2)4
2006 Counting Crowded Moving Objects
abstract
In its full generality, motion analysis of crowded objects necessitates recognition and segmentation of each moving entity. The difficulty of these tasks increases considerably with occlusions and therefore with crowding. When the objects are constrained to be of the same kind, however, partitioning of densely crowded semi-rigid objects can be accomplished by means of clustering tracked feature points. We base our approach on a highly parallelized version of the KLT tracker in order to process the video into a set of feature trajectories. While such a set of trajectories provides a substrate for motion analysis, their unequal lengths and fragmented nature present difficulties for subsequent processing. To address this, we propose a simple means of spatially and temporally conditioning the trajectories. Given this representation, we integrate it with a learned object descriptor to achieve a segmentation of the constituent motions. We present experimental results for the problem of estimating the number of moving objects in a dense crowd as a function of time.
Vincent C. Rabaud, Serge J. Belongie
CVPR (1)2
2006 Model Order Selection and Cue Combination for Image Segmentation
abstract
Model order selection and cue combination are both difficult open problems in the area of clustering. In this work we build upon stability-based approaches to develop a new method for automatic model order selection and cue combination with applications to visual grouping. Novel features of our approach include the ability to detect multiple stable clusterings (instead of only one), a simpler means of calculating stability that does not require training a classifier, and a new characterization of the space of stabilities for a continuum of segmentations that provides for an efficient sampling scheme. Our contribution is a framework for visual grouping that frees the user from the hassles of parameter tuning and model order selection: the input is an image, the output is a shortlist of segmentations.
Andrew Rabinovich, Serge J. Belongie, Tilman Lange, Joachim M. Buhmann
CVPR (1)2
2006 Practical Global Optimization for Multiview Geometry
Sameer Agarwal 0001, Manmohan Krishna Chandraker, Fredrik Kahl, David J. Kriegman, Serge J. Belongie
ECCV (1)5
2006 Higher order learning with graphs
abstract
Recently there has been considerable interest in learning with higher order relations (i.e., three-way or higher) in the unsupervised and semi-supervised settings. Hypergraphs and tensors have been proposed as the natural way of representing these relations and their corresponding algebra as the natural tools for operating on them. In this paper we argue that hypergraphs are not a natural representation for higher order relations, indeed pairwise as well as higher order relations can be handled using graphs. We show that various formulations of the semi-supervised and the unsupervised learning problem on hypergraphs result in the same graph theoretic problem and can be analyzed using existing tools.
Sameer Agarwal 0001, Kristin Branson, Serge J. Belongie
ICML3
2006 Learning to Traverse Image Manifolds
abstract
We present a new algorithm, Locally Smooth Manifold Learning (LSML), that learns a warping function from a point on an manifold to its neighbors. Important characteristics of LSML include the ability to recover the structure of the manifold in sparsely populated regions and beyond the support of the provided data. Appli- cations of our proposed technique include embedding with a natural out-of-sample extension and tasks such as tangent distance estimation, frame rate up-conversion, video compression and motion transfer.
Piotr Dollár, Serge J. Belongie, Vincent C. Rabaud
NIPS2
2006 A Feature-based Approach for Dense Segmentation and Estimation of Large Disparity Motion
Josh Wills, Sameer Agarwal 0001, Serge J. Belongie
Int. J. Comput. Vis.3
2006 Framework for Parsing, Visualizing and Scoring Tissue Microarray Images
abstract
Increasingly automated techniques for arraying, immunostaining, and imaging tissue sections led us to design software for convenient management, display, and scoring. Demand for molecular marker data derived in situ from tissue has driven histology informatics automation to the point where one can envision the computer, rather than the microscope, as the primary viewing platform for histopathological scoring and diagnoses. Tissue microarrays (TMAs), with hundreds or even thousands of patients' tissue sections on each slide, were the first step in this wave of automation. Via TMAs, increasingly rapid identification of the molecular patterns of cancer that define distinct clinical outcome groups among patients has become possible. TMAs have moved the bottleneck of acquiring molecular pattern information away from sampling and processing the tissues to the tasks of scoring and results analyses. The need to read large numbers of new slides, primarily for research purposes, is driving continuing advances in commercially available automated microscopy instruments that already do or soon will automatically image hundreds of slides per day. We reviewed strategies for acquiring, collating, and storing histological images with the goal of streamlining subsequent data analyses. As a result of this work, we report an implementation of software for automated preprocessing, organization, storage, and display of high resolution composite TMA images.
Andrew Rabinovich, S. Krajewski, M. Krajewska, Ahmed Shabaik, Stephen M. Hewitt, Serge J. Belongie, J. C. Reed, Jeffrey H. Price
IEEE Trans. Inf. Technol. Biomed.6
2005 Beyond Pairwise Clustering
abstract
We consider the problem of clustering in domains where the affinity relations are not dyadic (pairwise), but rather triadic, tetradic or higher. The problem is an instance of the hypergraph partitioning problem. We propose a two-step algorithm for solving this problem. In the first step we use a novel scheme to approximate the hypergraph using a weighted graph. In the second step a spectral partitioning algorithm is used to partition the vertices of this graph. The algorithm is capable of handling hyperedges of all orders including order two, thus incorporating information of all orders simultaneously. We present a theoretical analysis that relates our algorithm to an existing hypergraph partitioning algorithm and explain the reasons for its superior performance. We report the performance of our algorithm on a variety of computer vision problems and compare it to several existing hypergraph partitioning algorithms.
Sameer Agarwal 0001, Jongwoo Lim, Lihi Zelnik-Manor, Pietro Perona, David J. Kriegman, Serge J. Belongie
CVPR (2)6
2005 Tracking Multiple Mouse Contours (without Too Many Samples)
abstract
We present a particle filtering algorithm for robustly tracking the contours of multiple deformable objects through severe occlusions. Our algorithm combines a multiple blob tracker with a contour tracker in a manner that keeps the required number of samples small. This is a natural combination because both algorithms have complementary strengths. The multiple blob tracker uses a natural multi-target model and searches a smaller and simpler space. On the other hand, contour tracking gives more fine-tuned results and relies on cues that are available during severe occlusions. Our choice of combination of these two algorithms accentuates the advantages of each. We demonstrate good performance on challenging video of three identical mice that contains multiple instances of severe occlusion.
Kristin Branson, Serge J. Belongie
CVPR (1)2
2005 Periodic Motion Detection and Segmentation via Approximate Sequence Alignment
abstract
A method for detecting and segmenting periodic motion is presented. We exploit periodicity as a cue and detect periodic motion in complex scenes where common methods for motion segmentation are likely to fail. We note that periodic motion detection can be seen as an approximate case of sequence alignment where an image sequence is matched to itself over one or more periods of time. To use this observation, we first consider alignment of two video sequences obtained by independently moving cameras. Under assumption of constant translation, the fundamental matrices and the homographies are shown to be time-linear matrix functions. These dynamic quantities can be estimated by matching corresponding space-time points with similar local motion and shape. For periodic motion, we match corresponding points across periods and develop a RANSAC procedure to simultaneously estimate the period and the dynamic geometric transformations between periodic views. Using this method, we demonstrate detection and segmentation of human periodic motion in complex scenes with nonrigid backgrounds, moving camera and motion parallax.
Ivan Laptev, Serge J. Belongie, Patrick Pérez, Josh Wills
ICCV2
2005 Spatio-temporal texture synthesis and image inpainting for video applications
abstract
In this paper we investigate the application of texture synthesis and image inpainting techniques for video applications. Working in the non-parametric framework, we use 3D patches for matching and copying. This ensures temporal continuity to some extent which is not possible to obtain by working with individual frames. Since, in present application, patches might contain arbitrary shaped and multiple disconnected holes, fast Fourier transform (FFT) and summed area table based sum of squared difference (SSD) calculation (S.L. Kilthau, et al, 2002) cannot be used. We propose a modification of above scheme which allows its use in present application. This results in significant gain of efficiency since search space is typically huge for video applications.
Sanjeev Kumar 0003, Mainak Biswas, Serge J. Belongie, Truong Q. Nguyen
ICIP (2)3
2005 Efficient Shape Matching Using Shape Contexts
abstract
We demonstrate that shape contexts can be used to quickly prune a search for similar shapes. We present two algorithms for rapid shape retrieval: representative shape contexts, performing comparisons based on a small number of shape contexts, and shapemes, using vector quantization in the space of shape contexts to obtain prototypical shape pieces.
Greg Mori, Serge J. Belongie, Jitendra Malik
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 On Refractive Optical Flow
Sameer Agarwal 0001, Satya P. Mallick, David J. Kriegman, Serge J. Belongie
ECCV (2)4
2004 A Feature-Based Approach for Determining Dense Long Range Correspondences
Josh Wills, Serge J. Belongie
ECCV (3)2
2004 Spectral Grouping Using the Nyström Method
abstract
Spectral graph theoretic methods have recently shown great promise for the problem of image segmentation. However, due to the computational demands of these approaches, applications to large problems such as spatiotemporal data and high resolution imagery have been slow to appear. The contribution of this paper is a method that substantially reduces the computational requirements of grouping algorithms based on spectral partitioning making it feasible to apply them to very large grouping problems. Our approach is based on a technique for the numerical solution of eigenfunction problems known as the Nyström method. This method allows one to extrapolate the complete grouping solution using only a small number of samples. In doing so, we leverage the fact that there are far fewer coherent groups in a scene than pixels.
Charless C. Fowlkes, Serge J. Belongie, Fan Chung Graham, Jitendra Malik
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Normalized cuts in 3-D for spinal MRI segmentation
abstract
Segmentation of medical images has become an indispensable process to perform quantitative analysis of images of human organs and their functions. Normalized Cuts (NCut) is a spectral graph theoretic method that readily admits combinations of different features for image segmentation. The computational demand imposed by NCut has been successfully alleviated with the Nyström approximation method for applications different than medical imaging. In this paper we discuss the application of NCut with the Nyström approximation method to segment vertebral bodies from sagittal T1-weighted magnetic resonance images of the spine. The magnetic resonance images were preprocessed by the anisotropic diffusion algorithm, and three-dimensional local histograms of brightness was chosen as the segmentation feature. Results of the segmentation as well as limitations and challenges in this area are presented.
Julio Carballido-Gamio, Serge J. Belongie, Sharmila Majumdar
IEEE Trans. Medical Imaging2
2003 What Went Where
abstract
We present a framework for motion segmentation that combines the concepts of layer-based methods and feature-based motion estimation. We estimate the initial correspondences by comparing vectors of filter outputs at interest points, from which we compute candidate scene relations via random sampling of minimal subsets of correspondences. We achieve a dense, piecewise smooth assignment of pixels to motion layers using a fast approximate graph-cut algorithm based on a Markov random field formulation. We demonstrate our approach on image pairs containing large inter-frame motion and partial occlusion. The approach is efficient and it successfully segments scenes with inter-frame disparities previously beyond the scope of layer-based motion segmentation methods.
Josh Wills, Sameer Agarwal 0001, Serge J. Belongie
CVPR (1)3
2003 Unsupervised Color Decomposition Of Histologically Stained Tissue Samples
abstract
Accurate spectral decomposition is essential for the analysis and diagnosis of histologically stained tissue sections. In this paper we present the first automated system for performing this decompo- sition. We compare the performance of our system with ground truth data and report favorable results.
Andrew Rabinovich, Sameer Agarwal 0001, Casey Laris, Jeffrey H. Price, Serge J. Belongie
NIPS5
2003 Structured importance sampling of environment maps
abstract
We introduce structured importance sampling , a new technique for efficiently rendering scenes illuminated by distant natural illumination given in an environment map. Our method handles occlusion, high-frequency lighting, and is significantly faster than alternative methods based on Monte Carlo sampling. We achieve this speedup as a result of several ideas. First, we present a new metric for stratifying and sampling an environment map taking into account both the illumination intensity as well as the expected variance due to occlusion within the scene. We then present a novel hierarchical stratification algorithm that uses our metric to automatically stratify the environment map into regular strata. This approach enables a number of rendering optimizations, such as pre-integrating the illumination within each stratum to eliminate noise at the cost of adding bias, and sorting the strata to reduce the number of sample rays. We have rendered several scenes illuminated by natural lighting, and our results indicate that structured importance sampling is better than the best previous Monte Carlo techniques, requiring one to two orders of magnitude fewer samples for the same image quality.
Sameer Agarwal 0001, Ravi Ramamoorthi, Serge J. Belongie, Henrik Wann Jensen
ACM Trans. Graph.3
2002 Spectral Partitioning with Indefinite Kernels Using the Nyström Extension
Serge J. Belongie, Charless C. Fowlkes, Fan Chung Graham, Jitendra Malik
ECCV (3)1
2002 Approximate Thin Plate Spline Mappings
Gianluca Donato, Serge J. Belongie
ECCV (3)2
2002 On the non-optimality of four color coding of image partitions
abstract
The increased interest in region based image coding has given rise to graph coloring based partition encoding methods. These methods are based on the four color theorem for planar graphs, and assume that a coloring for a graph with the minimum possible number of colors will result in the most compressible representation. We show that this assumption is wrong. We show that there exist graphs with chromatic number k that can be colored with k + 1 colors resulting in bitmaps representing image partitions which are more compressible than, the corresponding bitmaps generated using any k coloring of the same graph. We conclude with some conjectures on optimal coloring of weighted graphs.
Sameer Agarwal 0001, Serge J. Belongie
ICIP (2)2
2002 Shape Matching and Object Recognition Using Shape Contexts
abstract
We present a novel approach to measuring similarity between shapes and exploit it for object recognition. In our framework, the measurement of similarity is preceded by: (1) solving for correspondences between points on the two shapes; (2) using the correspondences to estimate an aligning transform. In order to solve the correspondence problem, we attach a descriptor, the shape context, to each point. The shape context at a reference point captures the distribution of the remaining points relative to it, thus offering a globally discriminative characterization. Corresponding points on two similar shapes will have similar shape contexts, enabling us to solve for correspondences as an optimal assignment problem. Given the point correspondences, we estimate the transformation that best aligns the two shapes; regularized thin-plate splines provide a flexible class of transformation maps for this purpose. The dissimilarity between the two shapes is computed as a sum of matching errors between corresponding points, together with a term measuring the magnitude of the aligning transform. We treat recognition in a nearest-neighbor classification framework as the problem of finding the stored prototype shape that is maximally similar to that in the image. Results are presented for silhouettes, trademarks, handwritten digits, and the COIL data set.
Serge J. Belongie, Jitendra Malik, Jan Puzicha
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Blobworld: Image Segmentation Using Expectation-Maximization and Its Application to Image Querying
abstract
Retrieving images from large and varied collections using image content as a key is a challenging and important problem. We present a new image representation that provides a transformation from the raw pixel data to a small set of image regions that are coherent in color and texture. This "Blobworld" representation is created by clustering pixels in a joint color-texture-position feature space. The segmentation algorithm is fully automatic and has been run on a collection of 10,000 natural images. We describe a system that uses the Blobworld representation to retrieve images from this collection. An important aspect of the system is that the user is allowed to view the internal representation of the submitted image and the query results. Similar systems do not offer the user this view into the workings of the system; consequently, query results from these systems can be inexplicable, despite the availability of knobs for adjusting the similarity metrics. By finding image regions that roughly correspond to objects, we allow querying at the level of objects rather than global image properties. We present results indicating that querying for images using Blobworld produces higher precision than does querying using color and texture histograms of the entire image in cases where the image contains distinctive objects.
Chad Carson, Serge J. Belongie, Hayit Greenspan, Jitendra Malik
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Efficient Spatiotemporal Grouping Using the Nystro"m Method
abstract
Spectral graph theoretic methods have recently shown great promise for the problem of image segmentation, but due to the computational demands, applications of such methods to spatiotemporal data have been slow to appear For even a short video sequence, the set of all pairwise voxel similarities is a huge quantity of data: one second of a 256/spl times/384 sequence captured at 30 Hz entails on the order of 10/sup 13/ pairwise similarities. The contribution of this paper is a method that substantially reduces the computational requirements of grouping algorithms based on spectral partitioning, making it feasible to apply them to very large spatiotemporal grouping problems. Our approach is based on a technique for the numerical solution of eigenfunction problems known as the Nystrom method This method allows extrapolation of the complete grouping solution using only a small number of "typical" samples. In doing so, we successfully exploit the fact that there are far fewer coherent groups in an image sequence than pixels.
Charless C. Fowlkes, Serge J. Belongie, Jitendra Malik
CVPR (1)2
2001 Shape contexts enable efficient retrieval of similar shapes
abstract
In this paper we demonstrate that a recently introduced shape descriptor, the "shape context", can be used to quickly prune a search for similar shapes. Our representation for a shape is a discrete set of n points sampled from its internal and external contours. For each of these points, the shape context is a histogram of the relative positions of the n - 1 remaining points. We present two methods for rapid shape retrieval: one that does comparisons based on a small number of shape contexts and another that uses vector quantization in the space of shape contexts. We verify the discriminative power of these methods with tests on the Columbia (COIL-100) 3D object database and the Snodgrass and Vanderwart line drawings. The shape context-based methods are shown to quickly produce an accurate shortlist of candidates suitable for a more exact matching engine in spite of pose variation and occlusion.
Greg Mori, Serge J. Belongie, Jitendra Malik
CVPR (1)2
2001 Matching Shapes
Serge J. Belongie, Jitendra Malik, Jan Puzicha
ICCV1
2001 Contour and Texture Analysis for Image Segmentation
Jitendra Malik, Serge J. Belongie, Thomas K. Leung, Jianbo Shi
Int. J. Comput. Vis.2
2000 Model-Based Halftoning for Color Image Segmentation
abstract
Grouping algorithms based on histograms over measured image features have very successfully been applied to textured image segmentation. However, the competing goals of statistical estimation significance demanding few quantization levels versus the necessary richness in representation often prevent a successful application for the color cue, since quantization may result in contouring. We combine a halftoning technique called spatial quantization with distribution-based grouping algorithms to synthesize a powerful color image segmentation technique. The spatial quantization simultaneously determines color palette and halftoning by optimization of a joint cost function. It therefore allows for a highly adapted image representation with a smooth transition of color distributions for non-constant image surfaces.
Jan Puzicha, Serge J. Belongie
ICPR2
2000 Shape Context: A New Descriptor for Shape Matching and Object Recognition
abstract
We develop an approach to object recognition based on match(cid:173) ing shapes and using a resulting measure of similarity in a nearest neighbor classifier. The key algorithmic problem here is that of finding pointwise correspondences between an image shape and a stored prototype shape. We introduce a new shape descriptor, the shape context, which makes this possible, using a simple and robust algorithm. The shape context at a point captures the distri(cid:173) bution over relative positions of other shape points and thus sum(cid:173) marizes global shape in a rich, local descriptor. We demonstrate that shape contexts greatly simplify recovery of correspondences between points of two given shapes. Once shapes are aligned, shape contexts are used to define a robust score for measuring shape sim(cid:173) ilarity. We have used this score in a nearest-neighbor classifier for recognition of hand written digits as well as 3D objects, using exactly the same distance function. On the benchmark MNIST dataset of handwritten digits, this yields an error rate of 0.63%, outperforming other published techniques.
Serge J. Belongie, Jitendra Malik, Jan Puzicha
NIPS1
1999 Textons, Contours and Regions: Cue Integration in Image Segmentation
abstract
The paper makes two contributions: it provides (1) an operational definition of textons, the putative elementary units of texture perception, and (2) an algorithm for partitioning the image into disjoint regions of coherent brightness and texture, where boundaries of regions are defined by peaks in contour orientation energy and differences in texton densities across the contour. B. Julesz (1981) introduced the term texton, analogous to a phoneme in speech recognition, but did not provide an operational definition for gray-level images. We re-invent textons as frequently co-occurring combinations of oriented linear filter outputs. These can be learned using a K-means approach. By mapping each pixel to its nearest texton, the image can be analyzed into texton channels, each of which is a point set where discrete techniques such as Voronoi diagrams become applicable. Local histograms of texton frequencies can be used with a /spl chi//sup 2/ test for significant differences to find texture boundaries. Natural images contain both textured and untextured regions, so we combine this cue with that of the presence of peaks of contour energy derived from outputs of odd- and even-symmetric oriented Gaussian derivative filters. Each of these cues has a domain of applicability, so to facilitate cue combination we introduce a gating operator based on a statistical test for isotropy of Delaunay neighbors. Having obtained a local measure of how likely two nearby pixels are to belong to the same region, we use the spectral graph theoretic framework of normalized cuts to find partitions of the image into regions of coherent texture and brightness. Experimental results on a wide range of images are shown.
Jitendra Malik, Serge J. Belongie, Jianbo Shi, Thomas K. Leung
ICCV2
1998 Finding Boundaries in Natural Images: A New Method Using Point Descriptors and Area Completion
Serge J. Belongie, Jitendra Malik
ECCV (1)1
1998 Color- and Texture-based Image Segmentation Using the Expectation-Maximization Algorithm and its Application to Content-Based Image Retrieval
abstract
Retrieving images from large and varied collections using image content as a key is a challenging and important problem. In this paper we present a new image representation which provides a transformation from the raw pixel data to a small set of image regions which are coherent in color and texture space. This so-called "blobworld" representation is based on segmentation using the expectation-maximization algorithm on combined color and texture features. The texture features we use for the segmentation arise from a new approach to texture description and scale selection. We describe a system that uses the blobworld representation to retrieve images. An important and unique aspect of the system is that, in the context of similarity-based querying, the user is allowed to view the internal representation of the submitted image and the query results. Similar systems do not offer the user this view into the workings of the system; consequently, the outcome of many queries on these systems can be quite inexplicable, despite the availability of knobs for adjusting the similarity metric.
Serge J. Belongie, Chad Carson, Hayit Greenspan, Jitendra Malik
ICCV1
1998 Image and Video Segmentation: The Normalized Cut Framework
abstract
We propose a segmentation system based on the normalized cut framework proposed by Shi and Malik (see Proc. IEEE Conf. Computer Vision and Pattern Recognition, San Juan, Puerto Rico, p.731-7, 1997). The goal is to partition the image from a big picture point of view. Perceptually significant groups are detected first while small variations and details are treated later. Different image features-intensity, color, texture, contour continuity, motion and stereo disparity are treated in one uniform framework.
Jianbo Shi, Serge J. Belongie, Thomas K. Leung, Jitendra Malik
ICIP (1)2
1996 Finding objects in image databases by grouping
abstract
Retrieving images from very large collections, using image content as a key, is becoming an important problem. Finding objects in image databases is a big challenge in the field. The paper describes our approach to object recognition, which is distinguished by: a rich involvement of early visual primitives, including color and texture; hierarchical grouping and learning strategies in the classification process; the ability to deal with rather general objects in uncontrolled configurations and contexts. We illustrate these properties with three case studies: one demonstrating the use of color and texture descriptors; one learning scenery concepts using grouped features; and one demonstrating a possible application domain in detecting naked people in a scene.
Jitendra Malik, David A. Forsyth, Margaret M. Fleck, Hayit Greenspan, Thomas K. Leung, Chad Carson, Serge J. Belongie, Christoph Bregler
ICIP (2)7
1994 Overcomplete steerable pyramid filters and rotation invariance
abstract
A given (overcomplete) discrete oriented pyramid may be converted into a steerable pyramid by interpolation. We present a technique for deriving the optimal interpolation functions (otherwise called 'steering coefficients'). The proposed scheme is demonstrated on a computationally efficient oriented pyramid, which is a variation on the Burt and Adelson (1983) pyramid. We apply the generated steerable pyramid to orientation-invariant texture analysis in order to demonstrate its excellent rotational isotropy. High classification rates and precise rotation identification are demonstrated.>
Hayit Greenspan, Serge J. Belongie, Rodney M. Goodman, Pietro Perona, Subrata Rakshit, Charles H. Anderson
CVPR2
1994 Rotation invariant texture recognition using a steerable pyramid
abstract
A rotation-invariant texture recognition system is presented. A steerable oriented pyramid is used to extract representative features for the input textures. The steerability of the filter set allows a shift to an invariant representation via a DFT-encoding step. Supervised classification follows. State-of-the-art recognition results are presented on a 30 texture database with a comparison across the performance of the k-NN, backpropagation and rule-based classifiers. In addition, high accuracy estimation of the input rotation angle is demonstrated.
Hayit Greenspan, Serge J. Belongie, Rodney M. Goodman, Pietro Perona
ICPR (2)2