EDBT 2026 Demo / reviewers in the wild / expert
Sean Bell
dblp:132/3980
· DBLP profile ↗
10ranked-venue papers
6as first author
2since 2021 · last 2023
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Image recognition and object detection · 23% Representation and self-supervised learning · 22% Learning paradigms · 18% | |
| Computer graphics and multimedia
5 papers |
Visual content generation and editing · 42% Computational photography and imaging · 22% Rendering · 19% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
masked generative modeling |
0.7 | 1 | 2023 | Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation · CVPR 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data · ICLR 2023 |
Machine learning › Representation and self-supervised learning › representation learning
robust representation learning |
0.7 | 1 | 2023 | RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data · ICLR 2023 |
Machine learning › Learning paradigms
semi-supervised learning |
0.7 | 1 | 2023 | RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data · ICLR 2023 |
Visual content generation and editing
video generation |
0.7 | 1 | 2023 | Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
image embedding |
0.4 | 1 | 2020 | GrokNet: Unified Computer Vision Model Trunk and Embeddings For Commerce · KDD 2020 |
Machine learning › Learning paradigms
multi-task learning |
0.4 | 1 | 2020 | GrokNet: Unified Computer Vision Model Trunk and Embeddings For Commerce · KDD 2020 |
Computer vision › 3D vision › inverse rendering
intrinsic image decomposition |
0.3 | 1 | 2017 | Shading Annotations in the Wild · CVPR 2017 |
Computer vision › Image recognition and object detection › object detection
contextual reasoning |
0.2 | 1 | 2016 | Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks · CVPR 2016 |
Computer vision › Image recognition and object detection
object detection |
0.2 | 1 | 2016 | Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks · CVPR 2016 |
Computer vision › Image recognition and object detection › texture classification
material recognition |
0.2 | 1 | 2015 | Material recognition in the wild with the Materials in Context Database · CVPR 2015 |
Computer vision › Segmentation and scene understanding › semantic segmentation
material segmentation |
0.2 | 1 | 2015 | Material recognition in the wild with the Materials in Context Database · CVPR 2015 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.2 | 1 | 2015 | Material recognition in the wild with the Materials in Context Database · CVPR 2015 |
Multimedia analysis and retrieval
visual search |
0.2 | 1 | 2015 | Learning visual similarity for product design with convolutional neural networks · ACM Trans. Graph. 2015 |
Computational photography and imaging
intrinsic image decomposition |
0.2 | 1 | 2014 | Intrinsic images in the wild · ACM Trans. Graph. 2014 |
Rendering
appearance modeling |
0.2 | 1 | 2013 | OpenSurfaces: a richly annotated catalog of surface appearance · ACM Trans. Graph. 2013 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 2 | 2015 | Learning visual similarity for product design with convolutional neural networks · ACM Trans. Graph. 2015 Material recognition in the wild with the Materials in Context Database · CVPR 2015 |
Rendering
inverse rendering |
0.1 | 1 | 2017 | Shading Annotations in the Wild · CVPR 2017 |
Machine learning › Deep learning architectures and training
multi-scale representation |
0.1 | 1 | 2016 | Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks · CVPR 2016 |
Recommender systems
fashion recommendation |
0.1 | 1 | 2015 | Learning Visual Clothing Style with Heterogeneous Dyadic Co-Occurrences · ICCV 2015 |
Visualization and visual analytics
crowdsourced annotation |
0.1 | 1 | 2014 | Intrinsic images in the wild · ACM Trans. Graph. 2014 |
Rendering
physically based rendering |
0.0 | 1 | 2013 | OpenSurfaces: a richly annotated catalog of surface appearance · ACM Trans. Graph. 2013 |
Methods — techniques the papers use, named apart from their topics
masked modeling · 1.3discrete visual tokenization · 1.3crowdsourcing · 0.9multi-task learning · 0.9convolutional neural network · 0.8semi-supervised learning · 0.7self-supervised learning · 0.7embedding loss · 0.4categorical loss · 0.4recurrent neural network · 0.2siamese network · 0.2siamese convolutional neural network · 0.2metric learning · 0.2conditional random field · 0.2segmentation · 0.2human annotation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video GenerationabstractGenerating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head and tail is also crucial, but they have rarely been explored for video completion. Since there could be different outcomes from the hints of just a few frames, a system that can follow natural language to perform video completion may significantly improve controllability. Inspired by this, we introduce a novel task, text-guided video completion (TVC), which requests the model to generate a video from partial frames guided by an instruction. We then propose Multimodal Masked Video Generation (MMVG) to address this TVC task. During training, MMVG discretizes the video frames into visual tokens and masks most of them to perform video completion from any time point. At inference time, a single MMVG model can address all 3 cases of TVC, including video prediction, rewind, and infilling, by applying corresponding masking conditions. We evaluate MMVG in various video scenarios, including egocentric, animation, and gaming. Extensive experimental results indicate that MMVG is effective in generating high-quality visual appearances with text guidance for TVC. Tsu-Jui Fu, Licheng Yu, Ning Zhang 0014, Cheng-Yang Fu, Jong-Chyi Su, William Yang Wang, Sean Bell |
CVPR | 7 |
| 2023 | RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data
Sangwoo Mo, Jong-Chyi Su, Chih-Yao Ma, Mido Assran, Ishan Misra, Licheng Yu, Sean Bell |
ICLR | 7 |
| 2020 | GrokNet: Unified Computer Vision Model Trunk and Embeddings For CommerceabstractIn this paper, we present GrokNet, a deployed image recognition system for commerce applications. GrokNet leverages a multi-task learning approach to train a single computer vision trunk. We achieve a 2.1x improvement in exact product match accuracy when compared to the previous state-of-the-art Facebook product recognition system. We achieve this by training on 7 datasets across several commerce verticals, using 80 categorical loss functions and 3 embedding losses. We share our experience of combining diverse sources with wide-ranging label semantics and image statistics, including learning from human annotations, user-generated tags, and noisy search engine interaction data. GrokNet has demonstrated gains in production applications and operates at Facebook scale. Sean Bell, Yiqun Liu 0006, Sami Alsheikh, Yina Tang, Edward Pizzi, M. Henning, Karun Singh, Omkar Parkhi, Fedor Borisyuk |
KDD | 1 |
| 2017 | Shading Annotations in the WildabstractUnderstanding shading effects in images is critical for a variety of vision and graphics problems, including intrinsic image decomposition, shadow removal, image relighting, and inverse rendering. As is the case with other vision tasks, machine learning is a promising approach to understanding shading - but there is little ground truth shading data available for real-world images. We introduce Shading Annotations in the Wild (SAW), a new large-scale, public dataset of shading annotations in indoor scenes, comprised of multiple forms of shading judgments obtained via crowdsourcing, along with shading annotations automatically generated from RGB-D imagery. We use this data to train a convolutional neural network to predict per-pixel shading information in an image. We demonstrate the value of our data and network in an application to intrinsic images, where we can reduce decomposition artifacts produced by existing algorithms. Our database is available at http://opensurfaces.cs.cornell.edu/saw. Balazs Kovacs, Sean Bell, Noah Snavely, Kavita Bala |
CVPR | 2 |
| 2016 | Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural NetworksabstractIt is well known that contextual and multi-scale representations are important for accurate visual recognition. In this paper we present the Inside-Outside Net (ION), an object detector that exploits information both inside and outside the region of interest. Contextual information outside the region of interest is integrated using spatial recurrent neural networks. Inside, we use skip pooling to extract information at multiple scales and levels of abstraction. Through extensive experiments we evaluate the design space and provide readers with an overview of what tricks of the trade are important. ION improves state-of-the-art on PASCAL VOC 2012 object detection from 73.9% to 77.9% mAP. On the new and more challenging MS COCO dataset, we improve state-of-the-art from 19.7% to 33.1% mAP. In the 2015 MS COCO Detection Challenge, our ION model won "Best Student Entry" and finished 3rd place overall. As intuition suggests, our detection results provide strong evidence that context and multi-scale representations improve small object detection. Sean Bell, C. Lawrence Zitnick, Kavita Bala, Ross B. Girshick |
CVPR | 1 |
| 2015 | Material recognition in the wild with the Materials in Context DatabaseabstractRecognizing materials in real-world images is a challenging task. Real-world materials have rich surface texture, geometry, lighting conditions, and clutter, which combine to make the problem particularly difficult. In this paper, we introduce a new, large-scale, open dataset of materials in the wild, the Materials in Context Database (MINC), and combine this dataset with deep learning to achieve material recognition and segmentation of images in the wild. MINC is an order of magnitude larger than previous material databases, while being more diverse and well-sampled across its 23 categories. Using MINC, we train convolutional neural networks (CNNs) for two tasks: classifying materials from patches, and simultaneous material recognition and segmentation in full images. For patch-based classification on MINC we found that the best performing CNN architectures can achieve 85.2% mean class accuracy. We convert these trained CNN classifiers into an efficient fully convolutional framework combined with a fully connected conditional random field (CRF) to predict the material at every pixel in an image, achieving 73.1% mean class accuracy. Our experiments demonstrate that having a large, well-sampled dataset such as MINC is crucial for real-world material recognition and segmentation. Sean Bell, Paul Upchurch, Noah Snavely, Kavita Bala |
CVPR | 1 |
| 2015 | Learning Visual Clothing Style with Heterogeneous Dyadic Co-OccurrencesabstractWith the rapid proliferation of smart mobile devices, users now take millions of photos every day. These include large numbers of clothing and accessory images. We would like to answer questions like 'What outfit goes well with this pair of shoes?' To answer these types of questions, one has to go beyond learning visual similarity and learn a visual notion of compatibility across categories. In this paper, we propose a novel learning framework to help answer these types of questions. The main idea of this framework is to learn a feature transformation from images of items into a latent space that expresses compatibility. For the feature transformation, we use a Siamese Convolutional Neural Network (CNN) architecture, where training examples are pairs of items that are either compatible or incompatible. We model compatibility based on co-occurrence in large-scale user behavior data, in particular co-purchase data from Amazon.com. To learn cross-category fit, we introduce a strategic method to sample training data, where pairs of items are heterogeneous dyads, i.e., the two elements of a pair belong to different high-level categories. While this approach is applicable to a wide variety of settings, we focus on the representative problem of learning compatible clothing style. Our results indicate that the proposed framework is capable of learning semantic information about visual style and is able to generate outfits of clothes, with items from different categories, that go well together. Andreas Veit, Balazs Kovacs, Sean Bell, Julian J. McAuley, Kavita Bala, Serge J. Belongie |
ICCV | 3 |
| 2015 | Learning visual similarity for product design with convolutional neural networksabstractPopular sites like Houzz, Pinterest, and LikeThatDecor, have communities of users helping each other answer questions about products in images. In this paper we learn an embedding for visual search in interior design. Our embedding contains two different domains of product images: products cropped from internet scenes, and products in their iconic form. With such a multi-domain embedding, we demonstrate several applications of visual search including identifying products in scenes and finding stylistically similar products. To obtain the embedding, we train a convolutional neural network on pairs of images. We explore several training architectures including re-purposing object classifiers, using siamese networks, and using multitask learning. We evaluate our search quantitatively and qualitatively and demonstrate high quality results for search across multiple visual domains, enabling new applications in interior design. Sean Bell, Kavita Bala |
ACM Trans. Graph. | 1 |
| 2014 | Intrinsic images in the wildabstractIntrinsic image decomposition separates an image into a reflectance layer and a shading layer. Automatic intrinsic image decomposition remains a significant challenge, particularly for real-world scenes. Advances on this longstanding problem have been spurred by public datasets of ground truth data, such as the MIT Intrinsic Images dataset. However, the difficulty of acquiring ground truth data has meant that such datasets cover a small range of materials and objects. In contrast, real-world scenes contain a rich range of shapes and materials, lit by complex illumination. In this paper we introduce Intrinsic Images in the Wild , a large-scale, public dataset for evaluating intrinsic image decompositions of indoor scenes. We create this benchmark through millions of crowdsourced annotations of relative comparisons of material properties at pairs of points in each scene. Crowdsourcing enables a scalable approach to acquiring a large database, and uses the ability of humans to judge material comparisons, despite variations in illumination. Given our database, we develop a dense CRF-based intrinsic image algorithm for images in the wild that outperforms a range of state-of-the-art intrinsic image algorithms. Intrinsic image decomposition remains a challenging problem; we release our code and database publicly to support future research on this problem, available online at http://intrinsic.cs.cornell.edu/. Sean Bell, Kavita Bala, Noah Snavely |
ACM Trans. Graph. | 1 |
| 2013 | OpenSurfaces: a richly annotated catalog of surface appearanceabstractThe appearance of surfaces in real-world scenes is determined by the materials, textures, and context in which the surfaces appear. However, the datasets we have for visualizing and modeling rich surface appearance in context, in applications such as home remodeling, are quite limited. To help address this need, we present OpenSurfaces, a rich, labeled database consisting of thousands of examples of surfaces segmented from consumer photographs of interiors, and annotated with material parameters (reflectance, material names), texture information (surface normals, rectified textures), and contextual information (scene category, and object names). Retrieving usable surface information from uncalibrated Internet photo collections is challenging. We use human annotations and present a new methodology for segmenting and annotating materials in Internet photo collections suitable for crowdsourcing (e.g., through Amazon's Mechanical Turk). Because of the noise and variability inherent in Internet photos and novice annotators, designing this annotation engine was a key challenge; we present a multi-stage set of annotation tasks with quality checks and validation. We demonstrate the use of this database in proof-of-concept applications including surface retexturing and material and image browsing, and discuss future uses. OpenSurfaces is a public resource available at http://opensurfaces.cs.cornell.edu/. Sean Bell, Paul Upchurch, Noah Snavely, Kavita Bala |
ACM Trans. Graph. | 1 |