Steve Branson

dblp:56/8610 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Image recognition and object detection · 57% Segmentation and scene understanding · 9% Probabilistic and Bayesian machine learning · 9%
Human-computer interaction and pervasive computing
7 papers
Human-AI interaction · 61% Collaborative and social computing · 39%
Computer graphics and multimedia
1 paper
Image and video coding · 100%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 78% Data mining · 22%

Topics — the 29 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.742015
Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection · CVPR 2015
The Ignorant Led by the Blind: A Hybrid Human-Machine Vision System for Fine-Grained Categorization · Int. J. Comput. Vis. 2014
Similarity Comparisons for Interactive Fine-Grained Categorization · CVPR 2014
Computer vision › Image recognition and object detection
object detection
0.532016
Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016
Efficient Large-Scale Structured Learning · CVPR 2013
Strong supervision from weak annotation: Interactive training of deformable part models · ICCV 2011
Computer vision › Image recognition and object detection
image annotation
0.522017
Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017
Active Annotation Translation · CVPR 2014
Collaborative and social computing
crowdsourcing
0.422018
Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017
Lean Multiclass Crowdsourcing · CVPR 2018
Image and video coding › video compression
learned video compression
0.412019
Learned Video Compression · ICCV 2019
Image and video coding
video compression
0.412019
Learned Video Compression · ICCV 2019
Computer vision › Image recognition and object detection › object detection › part-based object detection
deformable part model
0.322013
Efficient Large-Scale Structured Learning · CVPR 2013
Strong supervision from weak annotation: Interactive training of deformable part models · ICCV 2011
Knowledge, reasoning and agents › Multi-agent systems
crowdsourcing
0.312017
Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017
Human-AI interaction
human-AI collaboration
0.312017
Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017
Computer vision › Image recognition and object detection › object detection
multi-view object detection
0.212016
Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016
Computer vision › Segmentation and scene understanding
scene understanding
0.212016
Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016
Machine learning › Efficient and distributed learning
active learning
0.212014
Active Annotation Translation · CVPR 2014
Computer vision › Video understanding and tracking › video analytics › behavior analysis
animal behavior analysis
0.212014
Detecting Social Actions of Fruit Flies · ECCV (2) 2014
Computer vision › Segmentation and scene understanding › part parsing
part labeling
0.212014
Active Annotation Translation · CVPR 2014
Information retrieval › image retrieval › user-centric image retrieval
interactive image retrieval
0.212014
Similarity Comparisons for Interactive Fine-Grained Categorization · CVPR 2014
Information retrieval
relevance feedback
0.212014
Similarity Comparisons for Interactive Fine-Grained Categorization · CVPR 2014
Machine learning › Efficient and distributed learning › large-scale learning
scalable training
0.212013
Efficient Large-Scale Structured Learning · CVPR 2013
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.212013
Efficient Large-Scale Structured Learning · CVPR 2013
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured SVM
0.212013
Efficient Large-Scale Structured Learning · CVPR 2013
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112010
The Multidimensional Wisdom of Crowds · NIPS 2010
Computer vision › Image recognition and object detection
visual recognition
0.112010
Visual Recognition with Humans in the Loop · ECCV (4) 2010
Data mining
crowdsourcing
0.112010
The Multidimensional Wisdom of Crowds · NIPS 2010
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.112009
Similarity metrics for categorization: From monolithic to category specific · ICCV 2009
Computer vision › 3D vision
multi-view geometry
0.112016
Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016
Natural language and speech › Information extraction and text analysis
dataset construction
0.112015
Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection · CVPR 2015
Computer vision › Image recognition and object detection
image classification
0.122010
The Multidimensional Wisdom of Crowds · NIPS 2010
Similarity metrics for categorization: From monolithic to category specific · ICCV 2009
Computer vision › Video understanding and tracking
multi-object tracking
0.112014
Detecting Social Actions of Fruit Flies · ECCV (2) 2014
Human-AI interaction › human-in-the-loop
human-in-the-loop annotation
0.112014
Active Annotation Translation · CVPR 2014
Human-AI interaction
human-in-the-loop
0.012010
Visual Recognition with Humans in the Loop · ECCV (4) 2010

Methods — techniques the papers use, named apart from their topics

probabilistic modeling · 0.8worker behavior modeling · 0.7sequential risk estimation · 0.6rate control · 0.4motion estimation · 0.4learned compensation · 0.4perceptual similarity metrics · 0.4incremental learning · 0.4majority-vote aggregation · 0.3majority vote aggregation · 0.3object detector · 0.2convolutional neural network · 0.2classifier · 0.2crowdsourcing · 0.2citizen science annotation · 0.2structured prediction · 0.2hybrid human-machine vision · 0.2part localization · 0.1
YearPublicationVenuePosition
2019 Learned Video Compression
abstract
We present a new algorithm for video coding, learned end-to-end for the low-latency mode. In this setting, our approach outperforms all existing video codecs across nearly the entire bitrate range. To our knowledge, this is the first ML-based method to do so. We evaluate our approach on standard video compression test sets of varying resolutions, and benchmark against all mainstream commercial codecs in the low-latency mode. On standard-definition videos, HEVC/H.265, AVC/H.264 and VP9 typically produce codes up to 60% larger than our algorithm. On high-definition 1080p videos, H.265 and VP9 typically produce codes up to 20% larger, and H.264 up to 35% larger. Furthermore, our approach does not suffer from blocking artifacts and pixelation, and thus produces videos that are more visually pleasing. We propose two main contributions. The first is a novel architecture for video compression, which (1) generalizes motion estimation to perform any learned compensation beyond simple translations, (2) rather than strictly relying on previously transmitted reference frames, maintains a state of arbitrary information learned by the model, and (3) enables jointly compressing all transmitted signals (such as optical flow and residual). Secondly, we present a framework for ML-based spatial rate control - a mechanism for assigning variable bitrates across space for each frame. This is a critical component for video coding, which to our knowledge had not been developed within a machine learning setting.
Oren Rippel, Sanjay Nair, Carissa Lew, Steve Branson, Alexander G. Anderson, Lubomir D. Bourdev
ICCV4
2018 Lean Multiclass Crowdsourcing
abstract
We introduce a method for efficiently crowdsourcing multiclass annotations in challenging, real world image datasets. Our method is designed to minimize the number of human annotations that are necessary to achieve a desired level of confidence on class labels. It is based on combining models of worker behavior with computer vision. Our method is general: it can handle a large number of classes, worker labels that come from a taxonomy rather than a flat list, and can model the dependence of labels when workers can see a history of previous annotations. Our method may be used as a drop-in replacement for the majority vote algorithms used in online crowdsourcing services that aggregate multiple human annotations into a final consolidated label. In experiments conducted on two real-life applications we find that our method can reduce the number of required annotations by as much as a factor of 5.4 and can reduce the residual annotation error by up to 90% when compared with majority voting. Furthermore, the online risk estimates of the models may be used to sort the annotated collection and minimize subsequent expert review effort.
Grant Van Horn, Steve Branson, Scott Loarie, Serge J. Belongie, Pietro Perona
CVPR2
2017 Lean Crowdsourcing: Combining Humans and Machines in an Online System
abstract
We introduce a method to greatly reduce the amount of redundant annotations required when crowdsourcing annotations such as bounding boxes, parts, and class labels. For example, if two Mechanical Turkers happen to click on the same pixel location when annotating a part in a given image-an event that is very unlikely to occur by random chance-, it is a strong indication that the location is correct. A similar type of confidence can be obtained if a single Turker happened to agree with a computer vision estimate. We thus incrementally collect a variable number of worker annotations per image based on online estimates of confidence. This is done using a sequential estimation of risk over a probabilistic model that combines worker skill, image difficulty, and an incrementally trained computer vision model. We develop specialized models and algorithms for binary annotation, part keypoint annotation, and sets of bounding box annotations. We show that our method can reduce annotation time by a factor of 4-11 for binary filtering of websearch results, 2-4 for annotation of boxes of pedestrians in images, while in many cases also reducing annotation error. We will make an end-to-end version of our system publicly available.
Steve Branson, Grant Van Horn, Pietro Perona
CVPR1
2016 Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees
abstract
Each corner of the inhabited world is imaged from multiple viewpoints with increasing frequency. Online map services like Google Maps or Here Maps provide direct access to huge amounts of densely sampled, georeferenced images from street view and aerial perspective. There is an opportunity to design computer vision systems that will help us search, catalog and monitor public infrastructure, buildings and artifacts. We explore the architecture and feasibility of such a system. The main technical challenge is combining test time information from multiple views of each geographic location (e.g., aerial and street views). We implement two modules: det2geo, which detects the set of locations of objects belonging to a given category, and geo2cat, which computes the fine-grained category of the object at a given location. We introduce a solution that adapts state-of the-art CNN-based object detectors and classifiers. We test our method on "Pasadena Urban Trees", a new dataset of 80,000 trees with geographic and species annotations, and show that combining multiple views significantly improves both tree detection and tree species classification, rivaling human performance.
Jan Dirk Wegner, Steve Branson, David Hall 0002, Konrad Schindler, Pietro Perona
CVPR2
2015 Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection
abstract
We introduce tools and methodologies to collect high quality, large scale fine-grained computer vision datasets using citizen scientists - crowd annotators who are passionate and knowledgeable about specific domains such as birds or airplanes. We worked with citizen scientists and domain experts to collect NABirds, a new high quality dataset containing 48,562 images of North American birds with 555 categories, part annotations and bounding boxes. We find that citizen scientists are significantly more accurate than Mechanical Turkers at zero cost. We worked with bird experts to measure the quality of popular datasets like CUB-200-2011 and ImageNet and found class label error rates of at least 4%. Nevertheless, we found that learning algorithms are surprisingly robust to annotation errors and this level of training data corruption can lead to an acceptably small increase in test error if the training set has sufficient size. At the same time, we found that an expert-curated high quality test set like NABirds is necessary to accurately measure the performance of fine-grained computer vision systems. We used NABirds to train a publicly available bird recognition service deployed on the web site of the Cornell Lab of Ornithology.
Grant Van Horn, Steve Branson, Ryan Farrell, Scott Haber, Jessie Barry, Panagiotis G. Ipeirotis, Pietro Perona, Serge J. Belongie
CVPR2
2014 Improved Bird Species Recognition Using Pose Normalized Deep Convolutional Nets
Steve Branson, Grant Van Horn, Pietro Perona, Serge J. Belongie
BMVC1
2014 Active Annotation Translation
abstract
We introduce a general framework for quickly annotating an image dataset when previous annotations exist. The new annotations (e.g. part locations) may be quite different from the old annotations (e.g. segmentations). Human annotators may be thought of as helping translate the old annotations into the new ones. As annotators label images, our algorithm incrementally learns a translator from source to target labels as well as a computer-vision-based structured predictor. These two components are combined to form an improved prediction system which accelerates the annotators' work through a smart GUI. We show how the method can be applied to translate between a wide variety of annotation types, including bounding boxes, segmentations, 2D and 3D part-based systems, and class and attribute labels. The proposed system will be a useful tool toward exploring new types of representations beyond simple bounding boxes, object segmentations, and class labels, and toward finding new ways to exploit existing large datasets with traditional types of annotations like SUN [36], Image Net [11], and Pascal VOC [12]. Experiments on the CUB-200-2011 and H3D datasets demonstrate 1) our method accelerates collection of part annotations by a factor of 3-20 compared to manual labeling, 2) our system can be used effectively in a scheme where definitions of part, attribute, or action vocabularies are evolved interactively without relabeling the entire dataset, and 3) toward collecting pose annotations, segmentations are more useful than bounding boxes, and part-level annotations are more effective than segmentations.
Steve Branson, Kristjan Eldjarn Hjorleifsson, Pietro Perona
CVPR1
2014 Similarity Comparisons for Interactive Fine-Grained Categorization
abstract
Current human-in-the-loop fine-grained visual categorization systems depend on a predefined vocabulary of attributes and parts, usually determined by experts. In this work, we move away from that expert-driven and attribute-centric paradigm and present a novel interactive classification system that incorporates computer vision and perceptual similarity metrics in a unified framework. At test time, users are asked to judge relative similarity between a query image and various sets of images, these general queries do not require expert-defined terminology and are applicable to other domains and basic-level categories, enabling a flexible, efficient, and scalable system for fine-grained categorization with humans in the loop. Our system outperforms existing state-of-the-art systems for relevance feedback-based image retrieval as well as interactive classification, resulting in a reduction of up to 43% in the average number of questions needed to correctly classify an image.
Catherine Wah, Grant Van Horn, Steve Branson, Subhransu Maji, Pietro Perona, Serge J. Belongie
CVPR3
2014 Detecting Social Actions of Fruit Flies
Eyrun Eyjolfsdottir, Steve Branson, Xavier Paolo Burgos-Artizzu, Eric D. Hoopfer, Jonathan Schor, David J. Anderson, Pietro Perona
ECCV (2)2
2014 The Ignorant Led by the Blind: A Hybrid Human-Machine Vision System for Fine-Grained Categorization
Steve Branson, Grant Van Horn, Catherine Wah, Pietro Perona, Serge J. Belongie
Int. J. Comput. Vis.1
2013 Efficient Large-Scale Structured Learning
abstract
We introduce an algorithm, SVM-IS, for structured SVM learning that is computationally scalable to very large datasets and complex structural representations. We show that structured learning is at least as fast-and often much faster-than methods based on binary classification for problems such as deformable part models, object detection, and multiclass classification, while achieving accuracies that are at least as good. Our method allows problem-specific structural knowledge to be exploited for faster optimization by integrating with a user-defined importance sampling function. We demonstrate fast train times on two challenging large scale datasets for two very different problems: Image Net for multiclass classification and CUB-200-2011 for deformable part model training. Our method is shown to be 10-50 times faster than SVMstructfor cost-sensitive multiclass classification while being about as fast as the fastest 1-vs-all methods for multiclass classification. For deformable part model training, it is shown to be 50-1000 times faster than methods based on SVMstruct, mining hard negatives, and Pegasos-style stochastic gradient descent. Source code of our method is publicly available.
Steve Branson, Oscar Beijbom, Serge J. Belongie
CVPR1
2011 Strong supervision from weak annotation: Interactive training of deformable part models
abstract
We propose a framework for large scale learning and annotation of structured models. The system interleaves interactive labeling (where the current model is used to semi-automate the labeling of a new example) and online learning (where a newly labeled example is used to update the current model parameters). This framework is scalable to large datasets and complex image models and is shown to have excellent theoretical and practical properties in terms of train time, optimality guarantees, and bounds on the amount of annotation effort per image. We apply this framework to part-based detection, and introduce a novel algorithm for interactive labeling of deformable part models. The labeling tool updates and displays in real-time the maximum likelihood location of all parts as the user clicks and drags the location of one or more parts. We demonstrate that the system can be used to efficiently and robustly train part and pose detectors on the CUB Birds-200-a challenging dataset of birds in unconstrained pose and environment.
Steve Branson, Pietro Perona, Serge J. Belongie
ICCV1
2011 Multiclass recognition and part localization with humans in the loop
abstract
We propose a visual recognition system that is designed for fine-grained visual categorization. The system is composed of a machine and a human user. The user, who is unable to carry out the recognition task by himself, is interactively asked to provide two heterogeneous forms of information: clicking on object parts and answering binary questions. The machine intelligently selects the most informative question to pose to the user in order to identify the object's class as quickly as possible. By leveraging computer vision and analyzing the user responses, the overall amount of human effort required, measured in seconds, is minimized. We demonstrate promising results on a challenging dataset of uncropped images, achieving a significant average reduction in human effort over previous methods.
Catherine Wah, Steve Branson, Pietro Perona, Serge J. Belongie
ICCV2
2010 Visual Recognition with Humans in the Loop
Steve Branson, Catherine Wah, Florian Schroff, Boris Babenko, Peter Welinder, Pietro Perona, Serge J. Belongie
ECCV (4)1
2010 The Multidimensional Wisdom of Crowds
abstract
Distributing labeling tasks among hundreds or thousands of annotators is an increasingly important method for annotating large datasets. We present a method for estimating the underlying value (e.g. the class) of each image from (noisy) annotations provided by multiple annotators. Our method is based on a model of the image formation and annotation process. Each image has different characteristics that are represented in an abstract Euclidean space. Each annotator is modeled as a multidimensional entity with variables representing competence, expertise and bias. This allows the model to discover and represent groups of annotators that have different sets of skills and knowledge, as well as groups of images that differ qualitatively. We find that our model predicts ground truth labels on both synthetic and real data more accurately than state of the art methods. Experiments also show that our model, starting from a set of binary labels, may discover rich information, such as different "schools of thought" amongst the annotators, and can group together images belonging to separate categories.
Peter Welinder, Steve Branson, Serge J. Belongie, Pietro Perona
NIPS2
2009 Similarity metrics for categorization: From monolithic to category specific
abstract
Similarity metrics that are learned from labeled training data can be advantageous in terms of performance and/or efficiency. These learned metrics can then be used in conjunction with a nearest neighbor classifier, or can be plugged in as kernels to an SVM. For the task of categorization two scenarios have thus far been explored. The first is to train a single “monolithic” similarity metric that is then used for all examples. The other is to train a metric for each category in a 1-vs-all manner. While the former approach seems to be at a disadvantage in terms of performance, the latter is not practical for large numbers of categories. In this paper we explore the space in between these two extremes. We present an algorithm that learns a few similarity metrics, while simultaneously grouping categories together and assigning one of these metrics to each group. We present promising results and show how the learned metrics generalize to novel categories.
Boris Babenko, Steve Branson, Serge J. Belongie
ICCV2