VLDB 2026 Research / reviewers in the wild / expert
Steve Branson
dblp:56/8610
· DBLP profile ↗
16ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Image recognition and object detection · 57% Segmentation and scene understanding · 9% Probabilistic and Bayesian machine learning · 9% | |
| Human-computer interaction and pervasive computing
7 papers |
Human-AI interaction · 61% Collaborative and social computing · 39% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 78% Data mining · 22% |
Topics — the 29 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
0.7 | 4 | 2015 | Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection · CVPR 2015 The Ignorant Led by the Blind: A Hybrid Human-Machine Vision System for Fine-Grained Categorization · Int. J. Comput. Vis. 2014 Similarity Comparisons for Interactive Fine-Grained Categorization · CVPR 2014 |
Computer vision › Image recognition and object detection
object detection |
0.5 | 3 | 2016 | Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016 Efficient Large-Scale Structured Learning · CVPR 2013 Strong supervision from weak annotation: Interactive training of deformable part models · ICCV 2011 |
Computer vision › Image recognition and object detection
image annotation |
0.5 | 2 | 2017 | Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017 Active Annotation Translation · CVPR 2014 |
Collaborative and social computing
crowdsourcing |
0.4 | 2 | 2018 | Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017 Lean Multiclass Crowdsourcing · CVPR 2018 |
Image and video coding › video compression
learned video compression |
0.4 | 1 | 2019 | Learned Video Compression · ICCV 2019 |
Image and video coding
video compression |
0.4 | 1 | 2019 | Learned Video Compression · ICCV 2019 |
Computer vision › Image recognition and object detection › object detection › part-based object detection
deformable part model |
0.3 | 2 | 2013 | Efficient Large-Scale Structured Learning · CVPR 2013 Strong supervision from weak annotation: Interactive training of deformable part models · ICCV 2011 |
Knowledge, reasoning and agents › Multi-agent systems
crowdsourcing |
0.3 | 1 | 2017 | Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017 |
Human-AI interaction
human-AI collaboration |
0.3 | 1 | 2017 | Lean Crowdsourcing: Combining Humans and Machines in an Online System · CVPR 2017 |
Computer vision › Image recognition and object detection › object detection
multi-view object detection |
0.2 | 1 | 2016 | Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.2 | 1 | 2016 | Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016 |
Machine learning › Efficient and distributed learning
active learning |
0.2 | 1 | 2014 | Active Annotation Translation · CVPR 2014 |
Computer vision › Video understanding and tracking › video analytics › behavior analysis
animal behavior analysis |
0.2 | 1 | 2014 | Detecting Social Actions of Fruit Flies · ECCV (2) 2014 |
Computer vision › Segmentation and scene understanding › part parsing
part labeling |
0.2 | 1 | 2014 | Active Annotation Translation · CVPR 2014 |
Information retrieval › image retrieval › user-centric image retrieval
interactive image retrieval |
0.2 | 1 | 2014 | Similarity Comparisons for Interactive Fine-Grained Categorization · CVPR 2014 |
Information retrieval
relevance feedback |
0.2 | 1 | 2014 | Similarity Comparisons for Interactive Fine-Grained Categorization · CVPR 2014 |
Machine learning › Efficient and distributed learning › large-scale learning
scalable training |
0.2 | 1 | 2013 | Efficient Large-Scale Structured Learning · CVPR 2013 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.2 | 1 | 2013 | Efficient Large-Scale Structured Learning · CVPR 2013 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured SVM |
0.2 | 1 | 2013 | Efficient Large-Scale Structured Learning · CVPR 2013 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.1 | 1 | 2010 | The Multidimensional Wisdom of Crowds · NIPS 2010 |
Computer vision › Image recognition and object detection
visual recognition |
0.1 | 1 | 2010 | Visual Recognition with Humans in the Loop · ECCV (4) 2010 |
Data mining
crowdsourcing |
0.1 | 1 | 2010 | The Multidimensional Wisdom of Crowds · NIPS 2010 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.1 | 1 | 2009 | Similarity metrics for categorization: From monolithic to category specific · ICCV 2009 |
Computer vision › 3D vision
multi-view geometry |
0.1 | 1 | 2016 | Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees · CVPR 2016 |
Natural language and speech › Information extraction and text analysis
dataset construction |
0.1 | 1 | 2015 | Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection · CVPR 2015 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 2 | 2010 | The Multidimensional Wisdom of Crowds · NIPS 2010 Similarity metrics for categorization: From monolithic to category specific · ICCV 2009 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.1 | 1 | 2014 | Detecting Social Actions of Fruit Flies · ECCV (2) 2014 |
Human-AI interaction › human-in-the-loop
human-in-the-loop annotation |
0.1 | 1 | 2014 | Active Annotation Translation · CVPR 2014 |
Human-AI interaction
human-in-the-loop |
0.0 | 1 | 2010 | Visual Recognition with Humans in the Loop · ECCV (4) 2010 |
Methods — techniques the papers use, named apart from their topics
probabilistic modeling · 0.8worker behavior modeling · 0.7sequential risk estimation · 0.6rate control · 0.4motion estimation · 0.4learned compensation · 0.4perceptual similarity metrics · 0.4incremental learning · 0.4majority-vote aggregation · 0.3majority vote aggregation · 0.3object detector · 0.2convolutional neural network · 0.2classifier · 0.2crowdsourcing · 0.2citizen science annotation · 0.2structured prediction · 0.2hybrid human-machine vision · 0.2part localization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Learned Video CompressionabstractWe present a new algorithm for video coding, learned end-to-end for the low-latency mode. In this setting, our approach outperforms all existing video codecs across nearly the entire bitrate range. To our knowledge, this is the first ML-based method to do so. We evaluate our approach on standard video compression test sets of varying resolutions, and benchmark against all mainstream commercial codecs in the low-latency mode. On standard-definition videos, HEVC/H.265, AVC/H.264 and VP9 typically produce codes up to 60% larger than our algorithm. On high-definition 1080p videos, H.265 and VP9 typically produce codes up to 20% larger, and H.264 up to 35% larger. Furthermore, our approach does not suffer from blocking artifacts and pixelation, and thus produces videos that are more visually pleasing. We propose two main contributions. The first is a novel architecture for video compression, which (1) generalizes motion estimation to perform any learned compensation beyond simple translations, (2) rather than strictly relying on previously transmitted reference frames, maintains a state of arbitrary information learned by the model, and (3) enables jointly compressing all transmitted signals (such as optical flow and residual). Secondly, we present a framework for ML-based spatial rate control - a mechanism for assigning variable bitrates across space for each frame. This is a critical component for video coding, which to our knowledge had not been developed within a machine learning setting. Oren Rippel, Sanjay Nair, Carissa Lew, Steve Branson, Alexander G. Anderson, Lubomir D. Bourdev |
ICCV | 4 |
| 2018 | Lean Multiclass CrowdsourcingabstractWe introduce a method for efficiently crowdsourcing multiclass annotations in challenging, real world image datasets. Our method is designed to minimize the number of human annotations that are necessary to achieve a desired level of confidence on class labels. It is based on combining models of worker behavior with computer vision. Our method is general: it can handle a large number of classes, worker labels that come from a taxonomy rather than a flat list, and can model the dependence of labels when workers can see a history of previous annotations. Our method may be used as a drop-in replacement for the majority vote algorithms used in online crowdsourcing services that aggregate multiple human annotations into a final consolidated label. In experiments conducted on two real-life applications we find that our method can reduce the number of required annotations by as much as a factor of 5.4 and can reduce the residual annotation error by up to 90% when compared with majority voting. Furthermore, the online risk estimates of the models may be used to sort the annotated collection and minimize subsequent expert review effort. Grant Van Horn, Steve Branson, Scott Loarie, Serge J. Belongie, Pietro Perona |
CVPR | 2 |
| 2017 | Lean Crowdsourcing: Combining Humans and Machines in an Online SystemabstractWe introduce a method to greatly reduce the amount of redundant annotations required when crowdsourcing annotations such as bounding boxes, parts, and class labels. For example, if two Mechanical Turkers happen to click on the same pixel location when annotating a part in a given image-an event that is very unlikely to occur by random chance-, it is a strong indication that the location is correct. A similar type of confidence can be obtained if a single Turker happened to agree with a computer vision estimate. We thus incrementally collect a variable number of worker annotations per image based on online estimates of confidence. This is done using a sequential estimation of risk over a probabilistic model that combines worker skill, image difficulty, and an incrementally trained computer vision model. We develop specialized models and algorithms for binary annotation, part keypoint annotation, and sets of bounding box annotations. We show that our method can reduce annotation time by a factor of 4-11 for binary filtering of websearch results, 2-4 for annotation of boxes of pedestrians in images, while in many cases also reducing annotation error. We will make an end-to-end version of our system publicly available. Steve Branson, Grant Van Horn, Pietro Perona |
CVPR | 1 |
| 2016 | Cataloging Public Objects Using Aerial and Street-Level Images - Urban TreesabstractEach corner of the inhabited world is imaged from multiple viewpoints with increasing frequency. Online map services like Google Maps or Here Maps provide direct access to huge amounts of densely sampled, georeferenced images from street view and aerial perspective. There is an opportunity to design computer vision systems that will help us search, catalog and monitor public infrastructure, buildings and artifacts. We explore the architecture and feasibility of such a system. The main technical challenge is combining test time information from multiple views of each geographic location (e.g., aerial and street views). We implement two modules: det2geo, which detects the set of locations of objects belonging to a given category, and geo2cat, which computes the fine-grained category of the object at a given location. We introduce a solution that adapts state-of the-art CNN-based object detectors and classifiers. We test our method on "Pasadena Urban Trees", a new dataset of 80,000 trees with geographic and species annotations, and show that combining multiple views significantly improves both tree detection and tree species classification, rivaling human performance. Jan Dirk Wegner, Steve Branson, David Hall 0002, Konrad Schindler, Pietro Perona |
CVPR | 2 |
| 2015 | Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collectionabstractWe introduce tools and methodologies to collect high quality, large scale fine-grained computer vision datasets using citizen scientists - crowd annotators who are passionate and knowledgeable about specific domains such as birds or airplanes. We worked with citizen scientists and domain experts to collect NABirds, a new high quality dataset containing 48,562 images of North American birds with 555 categories, part annotations and bounding boxes. We find that citizen scientists are significantly more accurate than Mechanical Turkers at zero cost. We worked with bird experts to measure the quality of popular datasets like CUB-200-2011 and ImageNet and found class label error rates of at least 4%. Nevertheless, we found that learning algorithms are surprisingly robust to annotation errors and this level of training data corruption can lead to an acceptably small increase in test error if the training set has sufficient size. At the same time, we found that an expert-curated high quality test set like NABirds is necessary to accurately measure the performance of fine-grained computer vision systems. We used NABirds to train a publicly available bird recognition service deployed on the web site of the Cornell Lab of Ornithology. Grant Van Horn, Steve Branson, Ryan Farrell, Scott Haber, Jessie Barry, Panagiotis G. Ipeirotis, Pietro Perona, Serge J. Belongie |
CVPR | 2 |
| 2014 | Improved Bird Species Recognition Using Pose Normalized Deep Convolutional Nets
Steve Branson, Grant Van Horn, Pietro Perona, Serge J. Belongie |
BMVC | 1 |
| 2014 | Active Annotation TranslationabstractWe introduce a general framework for quickly annotating an image dataset when previous annotations exist. The new annotations (e.g. part locations) may be quite different from the old annotations (e.g. segmentations). Human annotators may be thought of as helping translate the old annotations into the new ones. As annotators label images, our algorithm incrementally learns a translator from source to target labels as well as a computer-vision-based structured predictor. These two components are combined to form an improved prediction system which accelerates the annotators' work through a smart GUI. We show how the method can be applied to translate between a wide variety of annotation types, including bounding boxes, segmentations, 2D and 3D part-based systems, and class and attribute labels. The proposed system will be a useful tool toward exploring new types of representations beyond simple bounding boxes, object segmentations, and class labels, and toward finding new ways to exploit existing large datasets with traditional types of annotations like SUN [36], Image Net [11], and Pascal VOC [12]. Experiments on the CUB-200-2011 and H3D datasets demonstrate 1) our method accelerates collection of part annotations by a factor of 3-20 compared to manual labeling, 2) our system can be used effectively in a scheme where definitions of part, attribute, or action vocabularies are evolved interactively without relabeling the entire dataset, and 3) toward collecting pose annotations, segmentations are more useful than bounding boxes, and part-level annotations are more effective than segmentations. Steve Branson, Kristjan Eldjarn Hjorleifsson, Pietro Perona |
CVPR | 1 |
| 2014 | Similarity Comparisons for Interactive Fine-Grained CategorizationabstractCurrent human-in-the-loop fine-grained visual categorization systems depend on a predefined vocabulary of attributes and parts, usually determined by experts. In this work, we move away from that expert-driven and attribute-centric paradigm and present a novel interactive classification system that incorporates computer vision and perceptual similarity metrics in a unified framework. At test time, users are asked to judge relative similarity between a query image and various sets of images, these general queries do not require expert-defined terminology and are applicable to other domains and basic-level categories, enabling a flexible, efficient, and scalable system for fine-grained categorization with humans in the loop. Our system outperforms existing state-of-the-art systems for relevance feedback-based image retrieval as well as interactive classification, resulting in a reduction of up to 43% in the average number of questions needed to correctly classify an image. Catherine Wah, Grant Van Horn, Steve Branson, Subhransu Maji, Pietro Perona, Serge J. Belongie |
CVPR | 3 |
| 2014 | Detecting Social Actions of Fruit Flies
Eyrun Eyjolfsdottir, Steve Branson, Xavier Paolo Burgos-Artizzu, Eric D. Hoopfer, Jonathan Schor, David J. Anderson, Pietro Perona |
ECCV (2) | 2 |
| 2014 | The Ignorant Led by the Blind: A Hybrid Human-Machine Vision System for Fine-Grained Categorization
Steve Branson, Grant Van Horn, Catherine Wah, Pietro Perona, Serge J. Belongie |
Int. J. Comput. Vis. | 1 |
| 2013 | Efficient Large-Scale Structured LearningabstractWe introduce an algorithm, SVM-IS, for structured SVM learning that is computationally scalable to very large datasets and complex structural representations. We show that structured learning is at least as fast-and often much faster-than methods based on binary classification for problems such as deformable part models, object detection, and multiclass classification, while achieving accuracies that are at least as good. Our method allows problem-specific structural knowledge to be exploited for faster optimization by integrating with a user-defined importance sampling function. We demonstrate fast train times on two challenging large scale datasets for two very different problems: Image Net for multiclass classification and CUB-200-2011 for deformable part model training. Our method is shown to be 10-50 times faster than SVMstructfor cost-sensitive multiclass classification while being about as fast as the fastest 1-vs-all methods for multiclass classification. For deformable part model training, it is shown to be 50-1000 times faster than methods based on SVMstruct, mining hard negatives, and Pegasos-style stochastic gradient descent. Source code of our method is publicly available. Steve Branson, Oscar Beijbom, Serge J. Belongie |
CVPR | 1 |
| 2011 | Strong supervision from weak annotation: Interactive training of deformable part modelsabstractWe propose a framework for large scale learning and annotation of structured models. The system interleaves interactive labeling (where the current model is used to semi-automate the labeling of a new example) and online learning (where a newly labeled example is used to update the current model parameters). This framework is scalable to large datasets and complex image models and is shown to have excellent theoretical and practical properties in terms of train time, optimality guarantees, and bounds on the amount of annotation effort per image. We apply this framework to part-based detection, and introduce a novel algorithm for interactive labeling of deformable part models. The labeling tool updates and displays in real-time the maximum likelihood location of all parts as the user clicks and drags the location of one or more parts. We demonstrate that the system can be used to efficiently and robustly train part and pose detectors on the CUB Birds-200-a challenging dataset of birds in unconstrained pose and environment. Steve Branson, Pietro Perona, Serge J. Belongie |
ICCV | 1 |
| 2011 | Multiclass recognition and part localization with humans in the loopabstractWe propose a visual recognition system that is designed for fine-grained visual categorization. The system is composed of a machine and a human user. The user, who is unable to carry out the recognition task by himself, is interactively asked to provide two heterogeneous forms of information: clicking on object parts and answering binary questions. The machine intelligently selects the most informative question to pose to the user in order to identify the object's class as quickly as possible. By leveraging computer vision and analyzing the user responses, the overall amount of human effort required, measured in seconds, is minimized. We demonstrate promising results on a challenging dataset of uncropped images, achieving a significant average reduction in human effort over previous methods. Catherine Wah, Steve Branson, Pietro Perona, Serge J. Belongie |
ICCV | 2 |
| 2010 | Visual Recognition with Humans in the Loop
Steve Branson, Catherine Wah, Florian Schroff, Boris Babenko, Peter Welinder, Pietro Perona, Serge J. Belongie |
ECCV (4) | 1 |
| 2010 | The Multidimensional Wisdom of CrowdsabstractDistributing labeling tasks among hundreds or thousands of annotators is an increasingly important method for annotating large datasets. We present a method for estimating the underlying value (e.g. the class) of each image from (noisy) annotations provided by multiple annotators. Our method is based on a model of the image formation and annotation process. Each image has different characteristics that are represented in an abstract Euclidean space. Each annotator is modeled as a multidimensional entity with variables representing competence, expertise and bias. This allows the model to discover and represent groups of annotators that have different sets of skills and knowledge, as well as groups of images that differ qualitatively. We find that our model predicts ground truth labels on both synthetic and real data more accurately than state of the art methods. Experiments also show that our model, starting from a set of binary labels, may discover rich information, such as different "schools of thought" amongst the annotators, and can group together images belonging to separate categories. Peter Welinder, Steve Branson, Serge J. Belongie, Pietro Perona |
NIPS | 2 |
| 2009 | Similarity metrics for categorization: From monolithic to category specificabstractSimilarity metrics that are learned from labeled training data can be advantageous in terms of performance and/or efficiency. These learned metrics can then be used in conjunction with a nearest neighbor classifier, or can be plugged in as kernels to an SVM. For the task of categorization two scenarios have thus far been explored. The first is to train a single “monolithic” similarity metric that is then used for all examples. The other is to train a metric for each category in a 1-vs-all manner. While the former approach seems to be at a disadvantage in terms of performance, the latter is not practical for large numbers of categories. In this paper we explore the space in between these two extremes. We present an algorithm that learns a few similarity metrics, while simultaneously grouping categories together and assigning one of these metrics to each group. We present promising results and show how the learned metrics generalize to novel categories. Boris Babenko, Steve Branson, Serge J. Belongie |
ICCV | 2 |