Sean Bell

dblp:132/3980 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
2since 2021 · last 2023
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Image recognition and object detection · 23% Representation and self-supervised learning · 22% Learning paradigms · 18%
Computer graphics and multimedia
5 papers
Visual content generation and editing · 42% Computational photography and imaging · 22% Rendering · 19%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
masked generative modeling
0.712023
Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation · CVPR 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data · ICLR 2023
Machine learning › Representation and self-supervised learning › representation learning
robust representation learning
0.712023
RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data · ICLR 2023
Machine learning › Learning paradigms
semi-supervised learning
0.712023
RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data · ICLR 2023
Visual content generation and editing
video generation
0.712023
Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
image embedding
0.412020
GrokNet: Unified Computer Vision Model Trunk and Embeddings For Commerce · KDD 2020
Machine learning › Learning paradigms
multi-task learning
0.412020
GrokNet: Unified Computer Vision Model Trunk and Embeddings For Commerce · KDD 2020
Computer vision › 3D vision › inverse rendering
intrinsic image decomposition
0.312017
Shading Annotations in the Wild · CVPR 2017
Computer vision › Image recognition and object detection › object detection
contextual reasoning
0.212016
Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks · CVPR 2016
Computer vision › Image recognition and object detection
object detection
0.212016
Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks · CVPR 2016
Computer vision › Image recognition and object detection › texture classification
material recognition
0.212015
Material recognition in the wild with the Materials in Context Database · CVPR 2015
Computer vision › Segmentation and scene understanding › semantic segmentation
material segmentation
0.212015
Material recognition in the wild with the Materials in Context Database · CVPR 2015
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212015
Material recognition in the wild with the Materials in Context Database · CVPR 2015
Multimedia analysis and retrieval
visual search
0.212015
Learning visual similarity for product design with convolutional neural networks · ACM Trans. Graph. 2015
Computational photography and imaging
intrinsic image decomposition
0.212014
Intrinsic images in the wild · ACM Trans. Graph. 2014
Rendering
appearance modeling
0.212013
OpenSurfaces: a richly annotated catalog of surface appearance · ACM Trans. Graph. 2013
Machine learning › Deep learning architectures and training
convolutional neural network
0.122015
Learning visual similarity for product design with convolutional neural networks · ACM Trans. Graph. 2015
Material recognition in the wild with the Materials in Context Database · CVPR 2015
Rendering
inverse rendering
0.112017
Shading Annotations in the Wild · CVPR 2017
Machine learning › Deep learning architectures and training
multi-scale representation
0.112016
Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks · CVPR 2016
Recommender systems
fashion recommendation
0.112015
Learning Visual Clothing Style with Heterogeneous Dyadic Co-Occurrences · ICCV 2015
Visualization and visual analytics
crowdsourced annotation
0.112014
Intrinsic images in the wild · ACM Trans. Graph. 2014
Rendering
physically based rendering
0.012013
OpenSurfaces: a richly annotated catalog of surface appearance · ACM Trans. Graph. 2013

Methods — techniques the papers use, named apart from their topics

masked modeling · 1.3discrete visual tokenization · 1.3crowdsourcing · 0.9multi-task learning · 0.9convolutional neural network · 0.8semi-supervised learning · 0.7self-supervised learning · 0.7embedding loss · 0.4categorical loss · 0.4recurrent neural network · 0.2siamese network · 0.2siamese convolutional neural network · 0.2metric learning · 0.2conditional random field · 0.2segmentation · 0.2human annotation · 0.2
YearPublicationVenuePosition
2023 Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation
abstract
Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head and tail is also crucial, but they have rarely been explored for video completion. Since there could be different outcomes from the hints of just a few frames, a system that can follow natural language to perform video completion may significantly improve controllability. Inspired by this, we introduce a novel task, text-guided video completion (TVC), which requests the model to generate a video from partial frames guided by an instruction. We then propose Multimodal Masked Video Generation (MMVG) to address this TVC task. During training, MMVG discretizes the video frames into visual tokens and masks most of them to perform video completion from any time point. At inference time, a single MMVG model can address all 3 cases of TVC, including video prediction, rewind, and infilling, by applying corresponding masking conditions. We evaluate MMVG in various video scenarios, including egocentric, animation, and gaming. Extensive experimental results indicate that MMVG is effective in generating high-quality visual appearances with text guidance for TVC.
Tsu-Jui Fu, Licheng Yu, Ning Zhang 0014, Cheng-Yang Fu, Jong-Chyi Su, William Yang Wang, Sean Bell
CVPR7
2023 RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data
Sangwoo Mo, Jong-Chyi Su, Chih-Yao Ma, Mido Assran, Ishan Misra, Licheng Yu, Sean Bell
ICLR7
2020 GrokNet: Unified Computer Vision Model Trunk and Embeddings For Commerce
abstract
In this paper, we present GrokNet, a deployed image recognition system for commerce applications. GrokNet leverages a multi-task learning approach to train a single computer vision trunk. We achieve a 2.1x improvement in exact product match accuracy when compared to the previous state-of-the-art Facebook product recognition system. We achieve this by training on 7 datasets across several commerce verticals, using 80 categorical loss functions and 3 embedding losses. We share our experience of combining diverse sources with wide-ranging label semantics and image statistics, including learning from human annotations, user-generated tags, and noisy search engine interaction data. GrokNet has demonstrated gains in production applications and operates at Facebook scale.
Sean Bell, Yiqun Liu 0006, Sami Alsheikh, Yina Tang, Edward Pizzi, M. Henning, Karun Singh, Omkar Parkhi, Fedor Borisyuk
KDD1
2017 Shading Annotations in the Wild
abstract
Understanding shading effects in images is critical for a variety of vision and graphics problems, including intrinsic image decomposition, shadow removal, image relighting, and inverse rendering. As is the case with other vision tasks, machine learning is a promising approach to understanding shading - but there is little ground truth shading data available for real-world images. We introduce Shading Annotations in the Wild (SAW), a new large-scale, public dataset of shading annotations in indoor scenes, comprised of multiple forms of shading judgments obtained via crowdsourcing, along with shading annotations automatically generated from RGB-D imagery. We use this data to train a convolutional neural network to predict per-pixel shading information in an image. We demonstrate the value of our data and network in an application to intrinsic images, where we can reduce decomposition artifacts produced by existing algorithms. Our database is available at http://opensurfaces.cs.cornell.edu/saw.
Balazs Kovacs, Sean Bell, Noah Snavely, Kavita Bala
CVPR2
2016 Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
abstract
It is well known that contextual and multi-scale representations are important for accurate visual recognition. In this paper we present the Inside-Outside Net (ION), an object detector that exploits information both inside and outside the region of interest. Contextual information outside the region of interest is integrated using spatial recurrent neural networks. Inside, we use skip pooling to extract information at multiple scales and levels of abstraction. Through extensive experiments we evaluate the design space and provide readers with an overview of what tricks of the trade are important. ION improves state-of-the-art on PASCAL VOC 2012 object detection from 73.9% to 77.9% mAP. On the new and more challenging MS COCO dataset, we improve state-of-the-art from 19.7% to 33.1% mAP. In the 2015 MS COCO Detection Challenge, our ION model won "Best Student Entry" and finished 3rd place overall. As intuition suggests, our detection results provide strong evidence that context and multi-scale representations improve small object detection.
Sean Bell, C. Lawrence Zitnick, Kavita Bala, Ross B. Girshick
CVPR1
2015 Material recognition in the wild with the Materials in Context Database
abstract
Recognizing materials in real-world images is a challenging task. Real-world materials have rich surface texture, geometry, lighting conditions, and clutter, which combine to make the problem particularly difficult. In this paper, we introduce a new, large-scale, open dataset of materials in the wild, the Materials in Context Database (MINC), and combine this dataset with deep learning to achieve material recognition and segmentation of images in the wild. MINC is an order of magnitude larger than previous material databases, while being more diverse and well-sampled across its 23 categories. Using MINC, we train convolutional neural networks (CNNs) for two tasks: classifying materials from patches, and simultaneous material recognition and segmentation in full images. For patch-based classification on MINC we found that the best performing CNN architectures can achieve 85.2% mean class accuracy. We convert these trained CNN classifiers into an efficient fully convolutional framework combined with a fully connected conditional random field (CRF) to predict the material at every pixel in an image, achieving 73.1% mean class accuracy. Our experiments demonstrate that having a large, well-sampled dataset such as MINC is crucial for real-world material recognition and segmentation.
Sean Bell, Paul Upchurch, Noah Snavely, Kavita Bala
CVPR1
2015 Learning Visual Clothing Style with Heterogeneous Dyadic Co-Occurrences
abstract
With the rapid proliferation of smart mobile devices, users now take millions of photos every day. These include large numbers of clothing and accessory images. We would like to answer questions like 'What outfit goes well with this pair of shoes?' To answer these types of questions, one has to go beyond learning visual similarity and learn a visual notion of compatibility across categories. In this paper, we propose a novel learning framework to help answer these types of questions. The main idea of this framework is to learn a feature transformation from images of items into a latent space that expresses compatibility. For the feature transformation, we use a Siamese Convolutional Neural Network (CNN) architecture, where training examples are pairs of items that are either compatible or incompatible. We model compatibility based on co-occurrence in large-scale user behavior data, in particular co-purchase data from Amazon.com. To learn cross-category fit, we introduce a strategic method to sample training data, where pairs of items are heterogeneous dyads, i.e., the two elements of a pair belong to different high-level categories. While this approach is applicable to a wide variety of settings, we focus on the representative problem of learning compatible clothing style. Our results indicate that the proposed framework is capable of learning semantic information about visual style and is able to generate outfits of clothes, with items from different categories, that go well together.
Andreas Veit, Balazs Kovacs, Sean Bell, Julian J. McAuley, Kavita Bala, Serge J. Belongie
ICCV3
2015 Learning visual similarity for product design with convolutional neural networks
abstract
Popular sites like Houzz, Pinterest, and LikeThatDecor, have communities of users helping each other answer questions about products in images. In this paper we learn an embedding for visual search in interior design. Our embedding contains two different domains of product images: products cropped from internet scenes, and products in their iconic form. With such a multi-domain embedding, we demonstrate several applications of visual search including identifying products in scenes and finding stylistically similar products. To obtain the embedding, we train a convolutional neural network on pairs of images. We explore several training architectures including re-purposing object classifiers, using siamese networks, and using multitask learning. We evaluate our search quantitatively and qualitatively and demonstrate high quality results for search across multiple visual domains, enabling new applications in interior design.
Sean Bell, Kavita Bala
ACM Trans. Graph.1
2014 Intrinsic images in the wild
abstract
Intrinsic image decomposition separates an image into a reflectance layer and a shading layer. Automatic intrinsic image decomposition remains a significant challenge, particularly for real-world scenes. Advances on this longstanding problem have been spurred by public datasets of ground truth data, such as the MIT Intrinsic Images dataset. However, the difficulty of acquiring ground truth data has meant that such datasets cover a small range of materials and objects. In contrast, real-world scenes contain a rich range of shapes and materials, lit by complex illumination. In this paper we introduce Intrinsic Images in the Wild , a large-scale, public dataset for evaluating intrinsic image decompositions of indoor scenes. We create this benchmark through millions of crowdsourced annotations of relative comparisons of material properties at pairs of points in each scene. Crowdsourcing enables a scalable approach to acquiring a large database, and uses the ability of humans to judge material comparisons, despite variations in illumination. Given our database, we develop a dense CRF-based intrinsic image algorithm for images in the wild that outperforms a range of state-of-the-art intrinsic image algorithms. Intrinsic image decomposition remains a challenging problem; we release our code and database publicly to support future research on this problem, available online at http://intrinsic.cs.cornell.edu/.
Sean Bell, Kavita Bala, Noah Snavely
ACM Trans. Graph.1
2013 OpenSurfaces: a richly annotated catalog of surface appearance
abstract
The appearance of surfaces in real-world scenes is determined by the materials, textures, and context in which the surfaces appear. However, the datasets we have for visualizing and modeling rich surface appearance in context, in applications such as home remodeling, are quite limited. To help address this need, we present OpenSurfaces, a rich, labeled database consisting of thousands of examples of surfaces segmented from consumer photographs of interiors, and annotated with material parameters (reflectance, material names), texture information (surface normals, rectified textures), and contextual information (scene category, and object names). Retrieving usable surface information from uncalibrated Internet photo collections is challenging. We use human annotations and present a new methodology for segmenting and annotating materials in Internet photo collections suitable for crowdsourcing (e.g., through Amazon's Mechanical Turk). Because of the noise and variability inherent in Internet photos and novice annotators, designing this annotation engine was a key challenge; we present a multi-stage set of annotation tasks with quality checks and validation. We demonstrate the use of this database in proof-of-concept applications including surface retexturing and material and image browsing, and discuss future uses. OpenSurfaces is a public resource available at http://opensurfaces.cs.cornell.edu/.
Sean Bell, Paul Upchurch, Noah Snavely, Kavita Bala
ACM Trans. Graph.1